Skip to content
ImpactMojo ImpactMojo
Browse Membership
Back to Code Studio
Code Studio · Runs in your browser

Tidyverse for development data

Clean, join, reshape, weight and chart household survey data with R's tidyverse: readr, dplyr, tidyr and ggplot2, running live in your browser. For people who have finished the basics in R & Python for Development.

How the live code works. Your first Run downloads WebR (R for the browser, about 7 MB) and then each package a cell asks for: readr and dplyr first, tidyr and ggplot2 when you reach them. Give the first click of each module up to a minute on a slow connection; after that it is quick. Two files are already in the working folder: households.csv and districts.csv. Both are illustrative data, invented for teaching. The district names are real places, but no number here describes them.
Module 1 of 9

Reading data into a tibble

The tidyverse is a family of R packages that share one way of working: every dataset is a table with one row per observation and one column per variable, and every function takes that table as its first argument. This course uses four of them. readr reads files, dplyr transforms tables, tidyr reshapes them and ggplot2 draws them. We load each one by name instead of the tidyverse meta-package, which would pull in about thirty packages you do not need here.

The data is a household survey of 240 households across 10 districts in four states. Illustrative data, invented for teaching read_csv() from readr reads it into a tibble, the tidyverse version of a data frame.

A tibble prints only the first 10 rows and as many columns as fit, then tells you what it left out: here 230 more rows and 5 more variables. Under each column name is its type: <dbl> for numbers and <chr> for text. Printing a 50,000-row NFHS extract this way is safe; printing it as a base R data frame floods the console.

glimpse() turns the table on its side: one line per column, with its type and first few values. It is the fastest way to see every variable in a wide file. Here it confirms 240 rows and 13 columns, and that the four yes/no indicators were read as text.

The class line shows that a tibble is still a data.frame, so base R functions you already know keep working. distinct() lists the 10 districts.

Try it. Change "households.csv" to "districts.csv" in the first cell and run it. How many rows and columns does the district table have?
Module 2 of 9

Filter, select, arrange and mutate

dplyr gives you one verb per job. filter() keeps rows, select() keeps columns, arrange() sorts and mutate() adds columns. The pipe |> passes the table on the left into the function on the right, so a chain reads top to bottom as a list of steps.

Rural households headed by a woman, poorest first:

Two conditions separated by a comma inside filter() must both be true. 26 households pass, and the poorest, in Gaya, spends ₹880 per person per month.

mutate() computes new columns from existing ones, row by row. below_2000 is a logical column (TRUE or FALSE), which is how you build a flag for a threshold such as a poverty line. Household 7 has one member, so its household total equals its per-person figure.

%in% matches any value in a list, and between() is shorthand for x >= 0 & x <= 5. The three Bihar districts have 41 households whose head had five years of schooling or fewer. desc() sorts from largest down.

Try it. Filter for urban households in Kozhikode and Wayanad with more than 10 years of schooling. Then add a column big_hh that is TRUE when hh_size >= 6.
Module 3 of 9

Group, summarise and count

count() is the quickest table of frequencies. Give it two columns for a cross-tab in long form.

OBC households are the largest group (108 of 240), and 20 are ST, of whom only 5 are urban.

group_by() splits the table and summarise() reduces each piece to one row. Every function inside summarise() must return one number per group. A share is the mean of a TRUE/FALSE condition, so 100 * mean(has_toilet == "Yes") is a percentage.

Look at Udaipur: a mean of ₹3,611 and a median of ₹2,795. A few high-spending households pull the mean up, which is normal for consumption and income data. Report the median for a typical household. Kozhikode has the highest median, ₹6,080, and Rewa the lowest toilet coverage, 50 per cent.

.groups = "drop" removes the grouping after summarising, so the next step does not quietly run within groups. Check the n column before reading any median: the SC female-headed cell holds 2 households, so its median of ₹4,240 says almost nothing. In a real report you would suppress or flag cells that small.

Try it. Group by district, area instead, and add pct_bank for bank accounts. Which district has the widest rural-urban gap in median spending?
Module 4 of 9

Joining households to districts

Survey files rarely carry every variable you need. Here the household file has the district name, and a separate district file has the state, region, programme phase and field team. A join matches rows on a shared key.

left_join() keeps every household and adds the district columns where the names match. Each district appears once in districts.csv, so every household gets exactly one match and the row count stays at 240.

Joins fail silently when keys are spelt differently. Purnia is often written Purnea. Watch what happens when the district file uses the other spelling:

anti_join() returns the rows that found no partner: all 24 Purnia households. After a left join those households are still there, but with NA for state, so every state total would be quietly short. Run an anti_join() after every join; it should return zero rows.

With the join done you can summarise at any level the district table offers. Kerala has rows for phases 1 and 3 only, because none of its districts is in phase 2. A missing row is a fact about the design, worth a footnote in a report.

Try it. Change the grouping to field_team and compute the share of SHG members for each team.
Module 5 of 9

Reshaping indicator tables with tidyr

The four yes/no indicators sit in four columns. That is wide form. To summarise all four with one line of code, turn them into long form: one row per household per indicator. pivot_longer() does this.

240 households times 4 indicators gives 960 rows. Each row now holds a household, an indicator name and its answer, so a single group_by(district, indicator) computes every share at once: 40 rows, one per district per indicator.

Reports want the opposite shape: districts down the side, indicators across the top. pivot_wider() spreads the long table back out.

This is the standard indicator table. Rewa has 92 per cent of households with a bank account and 50 per cent with a toilet, and the lowest SHG membership, 4 per cent. Long form is for computing and charting, wide form is for reading.

Try it. Add head_gender next to district in group_by() and run the second cell again. The wide table gets one row per district per gender.
Module 6 of 9

Factors, bands and labels

Text columns sort alphabetically, which puts General before SC and ST. A factor stores a fixed set of categories in the order you choose, and that order carries through to tables and charts.

The second count follows the order in levels. Choose an order that means something for the reader, such as a fixed order used across all your tables, or ranked by the indicator.

cut() turns a number into bands. The breaks are the edges: (-Inf, 0] is no schooling, (0, 5] is one to five years, and so on. 36 heads had no schooling and 25 had more than 10 years. Always check that the bands cover every value; a value outside every break becomes NA.

across() applies the same function to several columns at once: first to turn each Yes/No into TRUE/FALSE, then to compute four shares in one summarise(). case_when() rewrites values by rule and is how you replace codes with readable labels. Here 63 per cent of rural households and 92 per cent of urban households have a toilet.

Try it. Recode land_acres into bands (landless, under 1 acre, 1 to 2.5 acres, over 2.5 acres) with cut() and count each band.
Module 7 of 9

Weighted summaries and survey weights

Most large surveys do not give every household the same chance of selection. A household drawn with probability 1 in 500 stands for 500 households in the population, so its design weight is 500. Unweighted means describe the sample; weighted means estimate the population.

Our illustrative file has no weight column, so we invent a design for teaching: rural households drawn 1 in 500 and urban households 1 in 1,000. Urban households are under-sampled, so each one carries a weight of 1,000.

The weighted mean spending, ₹3,857, is higher than the unweighted ₹3,401, and toilet coverage rises from 72.5 to 77.2 per cent, because urban households (which spend more and more often have toilets) now count for twice as much. Same data, different question answered.

NFHS weights

The DHS Program lists NFHS-5 (2019-21) as India's Standard DHS survey. The DHS Program's Guide to DHS Statistics (checked October 2026) explains that the women's weight v005 is stored without its decimal point and must be divided by 1,000,000 before use. Three made-up rows show the step:

Dividing by a constant does not change a weighted mean, so weighted.mean(anaemic, v005) gives the same 0.789. The division matters when you sum weights to estimate a count of women.

Weights fix the estimate; they do not fix the standard error. weighted.mean() gives the right point estimate but knows nothing about clusters and strata, so any confidence interval built from it will be too narrow. For standard errors use the survey package by Thomas Lumley (version 4.5 on CRAN as of October 2026): declare the design once with svydesign(ids = ~psu, strata = ~strata, weights = ~wt, data = df), then use svymean() and svyby().
Try it. Change the urban weight to 500 so both areas match, and run the weights cell again. The weighted and unweighted results should now agree.
Module 8 of 9

Charts with ggplot2

ggplot2 builds a chart in layers: the data, an aes() mapping that says which column goes on which axis, and a geom that says how to draw it. Add layers with +. The first run of this module downloads ggplot2 and its dependencies, so give it a minute.

geom_col() draws bars from values you computed. Bars always start at zero in ggplot2, which is right: a bar's length is its value.

Each point is a household. Spending is skewed, so a log scale spreads out the crowded lower end; say so in the axis label, as here, because a reader will otherwise read the gaps as equal.

facet_wrap(~ district) draws one small panel per district on a shared scale, so the panels can be compared at a glance. Avoid scales = "free_y" when readers will compare panels: it gives each panel its own axis, and a short bar in one panel can stand for a larger value than a tall bar in another. Kozhikode shows no rural bar because the value is zero: none of its 3 rural households has a toilet. Three households are far too few to report a percentage for.

Honest axes

The median is ₹2,380 for rural and ₹4,515 for urban households, so urban spending is about 1.9 times rural. In the first chart the axis starts at 2,000 and the urban bar looks more than six times as tall. The second chart, from zero, shows the real ratio. For bar charts, start at zero. See Data Visualization 101 for more.

Try it. In the facet chart, swap x = area for x = caste and add caste to group_by(). Which districts have caste groups with very few households?
Module 9 of 9

A reproducible pipeline

Put the whole analysis in one script that runs from the raw files to the final table with no hand edits. Anyone with the files can then reproduce every number, including you in six months.

The script reads both files, stops with an error if any household fails to match a district, builds the weights and indicators, and writes a four-row state table: Bihar 80.9 per cent toilet coverage, Kerala 82.2, Madhya Pradesh 72.9 and Rajasthan 71.7 (weighted, on illustrative data). It then reads the file back and confirms 4 rows.

Try it. Add region to the group_by() in step 4, then change the output file name. Run it and check the row count printed at the end.

Where next

→

pandas for Development Data

The same analysis on the same files, in Python.

→

R & Python for Development

Back to the basics, side by side in both languages.

→

Exploratory Data Analysis 101

What to look for in a household survey before you model it.

→

Survey Design 101

Where sampling weights come from.

→

Data Visualization 101

Design principles for the charts you can now draw.