Tidyverse for development data
Clean, join, reshape, weight and chart household survey data with R's tidyverse: readr, dplyr, tidyr and ggplot2, running live in your browser. For people who have finished the basics in R & Python for Development.
households.csv and districts.csv. Both are illustrative data, invented for teaching. The district names are real places, but no number here describes them.Reading data into a tibble
The tidyverse is a family of R packages that share one way of working: every dataset is a table with one row per observation and one column per variable, and every function takes that table as its first argument. This course uses four of them. readr reads files, dplyr transforms tables, tidyr reshapes them and ggplot2 draws them. We load each one by name instead of the tidyverse meta-package, which would pull in about thirty packages you do not need here.
The data is a household survey of 240 households across 10 districts in four states. Illustrative data, invented for teaching read_csv() from readr reads it into a tibble, the tidyverse version of a data frame.
A tibble prints only the first 10 rows and as many columns as fit, then tells you what it left out: here 230 more rows and 5 more variables. Under each column name is its type: <dbl> for numbers and <chr> for text. Printing a 50,000-row NFHS extract this way is safe; printing it as a base R data frame floods the console.
glimpse() turns the table on its side: one line per column, with its type and first few values. It is the fastest way to see every variable in a wide file. Here it confirms 240 rows and 13 columns, and that the four yes/no indicators were read as text.
The class line shows that a tibble is still a data.frame, so base R functions you already know keep working. distinct() lists the 10 districts.
"households.csv" to "districts.csv" in the first cell and run it. How many rows and columns does the district table have?Filter, select, arrange and mutate
dplyr gives you one verb per job. filter() keeps rows, select() keeps columns, arrange() sorts and mutate() adds columns. The pipe |> passes the table on the left into the function on the right, so a chain reads top to bottom as a list of steps.
Rural households headed by a woman, poorest first:
Two conditions separated by a comma inside filter() must both be true. 26 households pass, and the poorest, in Gaya, spends ₹880 per person per month.
mutate() computes new columns from existing ones, row by row. below_2000 is a logical column (TRUE or FALSE), which is how you build a flag for a threshold such as a poverty line. Household 7 has one member, so its household total equals its per-person figure.
%in% matches any value in a list, and between() is shorthand for x >= 0 & x <= 5. The three Bihar districts have 41 households whose head had five years of schooling or fewer. desc() sorts from largest down.
big_hh that is TRUE when hh_size >= 6.Group, summarise and count
count() is the quickest table of frequencies. Give it two columns for a cross-tab in long form.
OBC households are the largest group (108 of 240), and 20 are ST, of whom only 5 are urban.
group_by() splits the table and summarise() reduces each piece to one row. Every function inside summarise() must return one number per group. A share is the mean of a TRUE/FALSE condition, so 100 * mean(has_toilet == "Yes") is a percentage.
Look at Udaipur: a mean of ₹3,611 and a median of ₹2,795. A few high-spending households pull the mean up, which is normal for consumption and income data. Report the median for a typical household. Kozhikode has the highest median, ₹6,080, and Rewa the lowest toilet coverage, 50 per cent.
.groups = "drop" removes the grouping after summarising, so the next step does not quietly run within groups. Check the n column before reading any median: the SC female-headed cell holds 2 households, so its median of ₹4,240 says almost nothing. In a real report you would suppress or flag cells that small.
district, area instead, and add pct_bank for bank accounts. Which district has the widest rural-urban gap in median spending?Joining households to districts
Survey files rarely carry every variable you need. Here the household file has the district name, and a separate district file has the state, region, programme phase and field team. A join matches rows on a shared key.
left_join() keeps every household and adds the district columns where the names match. Each district appears once in districts.csv, so every household gets exactly one match and the row count stays at 240.
Joins fail silently when keys are spelt differently. Purnia is often written Purnea. Watch what happens when the district file uses the other spelling:
anti_join() returns the rows that found no partner: all 24 Purnia households. After a left join those households are still there, but with NA for state, so every state total would be quietly short. Run an anti_join() after every join; it should return zero rows.
With the join done you can summarise at any level the district table offers. Kerala has rows for phases 1 and 3 only, because none of its districts is in phase 2. A missing row is a fact about the design, worth a footnote in a report.
field_team and compute the share of SHG members for each team.Reshaping indicator tables with tidyr
The four yes/no indicators sit in four columns. That is wide form. To summarise all four with one line of code, turn them into long form: one row per household per indicator. pivot_longer() does this.
240 households times 4 indicators gives 960 rows. Each row now holds a household, an indicator name and its answer, so a single group_by(district, indicator) computes every share at once: 40 rows, one per district per indicator.
Reports want the opposite shape: districts down the side, indicators across the top. pivot_wider() spreads the long table back out.
This is the standard indicator table. Rewa has 92 per cent of households with a bank account and 50 per cent with a toilet, and the lowest SHG membership, 4 per cent. Long form is for computing and charting, wide form is for reading.
head_gender next to district in group_by() and run the second cell again. The wide table gets one row per district per gender.Factors, bands and labels
Text columns sort alphabetically, which puts General before SC and ST. A factor stores a fixed set of categories in the order you choose, and that order carries through to tables and charts.
The second count follows the order in levels. Choose an order that means something for the reader, such as a fixed order used across all your tables, or ranked by the indicator.
cut() turns a number into bands. The breaks are the edges: (-Inf, 0] is no schooling, (0, 5] is one to five years, and so on. 36 heads had no schooling and 25 had more than 10 years. Always check that the bands cover every value; a value outside every break becomes NA.
across() applies the same function to several columns at once: first to turn each Yes/No into TRUE/FALSE, then to compute four shares in one summarise(). case_when() rewrites values by rule and is how you replace codes with readable labels. Here 63 per cent of rural households and 92 per cent of urban households have a toilet.
land_acres into bands (landless, under 1 acre, 1 to 2.5 acres, over 2.5 acres) with cut() and count each band.Weighted summaries and survey weights
Most large surveys do not give every household the same chance of selection. A household drawn with probability 1 in 500 stands for 500 households in the population, so its design weight is 500. Unweighted means describe the sample; weighted means estimate the population.
Our illustrative file has no weight column, so we invent a design for teaching: rural households drawn 1 in 500 and urban households 1 in 1,000. Urban households are under-sampled, so each one carries a weight of 1,000.
The weighted mean spending, ₹3,857, is higher than the unweighted ₹3,401, and toilet coverage rises from 72.5 to 77.2 per cent, because urban households (which spend more and more often have toilets) now count for twice as much. Same data, different question answered.
NFHS weights
The DHS Program lists NFHS-5 (2019-21) as India's Standard DHS survey. The DHS Program's Guide to DHS Statistics (checked October 2026) explains that the women's weight v005 is stored without its decimal point and must be divided by 1,000,000 before use. Three made-up rows show the step:
Dividing by a constant does not change a weighted mean, so weighted.mean(anaemic, v005) gives the same 0.789. The division matters when you sum weights to estimate a count of women.
weighted.mean() gives the right point estimate but knows nothing about clusters and strata, so any confidence interval built from it will be too narrow. For standard errors use the survey package by Thomas Lumley (version 4.5 on CRAN as of October 2026): declare the design once with svydesign(ids = ~psu, strata = ~strata, weights = ~wt, data = df), then use svymean() and svyby().Charts with ggplot2
ggplot2 builds a chart in layers: the data, an aes() mapping that says which column goes on which axis, and a geom that says how to draw it. Add layers with +. The first run of this module downloads ggplot2 and its dependencies, so give it a minute.
geom_col() draws bars from values you computed. Bars always start at zero in ggplot2, which is right: a bar's length is its value.
Each point is a household. Spending is skewed, so a log scale spreads out the crowded lower end; say so in the axis label, as here, because a reader will otherwise read the gaps as equal.
facet_wrap(~ district) draws one small panel per district on a shared scale, so the panels can be compared at a glance. Avoid scales = "free_y" when readers will compare panels: it gives each panel its own axis, and a short bar in one panel can stand for a larger value than a tall bar in another. Kozhikode shows no rural bar because the value is zero: none of its 3 rural households has a toilet. Three households are far too few to report a percentage for.
Honest axes
The median is ₹2,380 for rural and ₹4,515 for urban households, so urban spending is about 1.9 times rural. In the first chart the axis starts at 2,000 and the urban bar looks more than six times as tall. The second chart, from zero, shows the real ratio. For bar charts, start at zero. See Data Visualization 101 for more.
x = area for x = caste and add caste to group_by(). Which districts have caste groups with very few households?A reproducible pipeline
Put the whole analysis in one script that runs from the raw files to the final table with no hand edits. Anyone with the files can then reproduce every number, including you in six months.
The script reads both files, stops with an error if any household fails to match a district, builds the weights and indicators, and writes a four-row state table: Bihar 80.9 per cent toilet coverage, Kerala 82.2, Madhya Pradesh 72.9 and Rajasthan 71.7 (weighted, on illustrative data). It then reads the file back and confirms 4 rows.
stopifnot()turns a silent problem into a loud one. Add a check after every join and every filter you depend on.- Keep raw files read-only. Write outputs to new files, as
write_csv()does here. - Keep names, phone numbers and other identifiers out of files you share. Data Protection & the DPDP Act 101 covers what India's Digital Personal Data Protection Act, 2023 asks of survey teams.
- On your own laptop, save the script as a
.Rfile inside an RStudio project, or write it up in Quarto so the text and numbers come from the same run.
region to the group_by() in step 4, then change the output file name. Run it and check the row count printed at the end.Where next
pandas for Development Data
The same analysis on the same files, in Python.
R & Python for Development
Back to the basics, side by side in both languages.
Exploratory Data Analysis 101
What to look for in a household survey before you model it.
Survey Design 101
Where sampling weights come from.
Data Visualization 101
Design principles for the charts you can now draw.