Stata syntax for development data
Write Stata do-files from your first import to a weighted regression: describe, generate, recode, labels, egen, collapse, merge, reshape, svyset for NFHS-style surveys, regress and margins, graphs and logs. Follow along in your own copy of Stata with two small CSV files, and check each step in R on this page.
Stata, the do-file and your project folder
Stata is a paid statistics package from StataCorp, used widely in development economics, evaluation teams and survey firms. Most published replication files for NFHS, PLFS and IHDS analyses are Stata do-files, so reading Stata is useful even if you work in R.
The five main windows
Stata's own Getting Started manual names five main windows: History (commands you have run), Results (output), Command (where you type one command and press Enter), Variables (the variables in memory) and Properties (details of the selected variable). The menus can build most commands for you, and the command they build appears in the Results window, which is a good way to learn syntax.
Work in a do-file from the first day
A do-file is a plain text file of commands that Stata runs from top to bottom. It is your record of what you did, and anyone with the same data can re-run it and get the same numbers. Open the Do-file Editor by clicking its button on the toolbar or by typing doedit in the Command window.
- Make a project folder, for example
C:\impactmojo\stataon Windows or~/impactmojo/stataon a Mac. - Download the two course files into it: households.csv (240 households) and districts.csv (10 districts). Illustrative data, invented for teaching The district names are real places; every number is made up.
- In Stata, type
doeditin the Command window and press Enter. - Type the lines below into the new do-file and save it in your project folder as
01_setup.do. - Click the Do button (Execute (do)) on the Do-file Editor toolbar. If you highlight some lines first, the same button runs only those lines.
* 01_setup.do : first do-file for the Code Studio Stata course version 19 // run as Stata 19, even in a later Stata clear all // start with nothing in memory cd "C:\impactmojo\stata" // change to YOUR project folder pwd // print the working folder, to check dir *.csv // you should see households.csv and districts.csv
Lines starting with * are comments, and // starts a comment at the end of a line. Look in the Results window: pwd should print your folder, and dir should list both CSV files. If cd fails, the path is wrong; copy it from your file manager's address bar.
cd line to your own folder, run the do-file, then type help import delimited in the Command window. Every Stata command has a help file like this, with the syntax at the top and worked examples at the bottom.Import a CSV and look at it
import delimited reads comma-separated and tab-separated text files. With the variable names in the first row, as in our files, it needs only the file name. Then look before you change anything.
version 19 clear all cd "C:\impactmojo\stata" import delimited using "households.csv", clear describe // variables, types, number of observations codebook district caste, compact list in 1/5 // first five households summarize monthly_pc_exp hh_size head_edu_years summarize monthly_pc_exp, detail // adds percentiles and the median tabulate caste tabulate caste area // two-way table tabulate caste area, row // row percentages
What to look for
describeshould report 240 observations and 13 variables. Variables holding text (district,caste,has_toiletand the other Yes/No columns) show astrstorage type. Stata cannot run numeric commands on them until you convert them, which is the next module.- In
tabulate castethe Freq. column should sum to 240, with four rows: General, OBC, SC and ST. - In
summarize, detailcompare the mean with the 50% percentile (the median). Expenditure is usually skewed to the right, so the mean sits above the median.
The same look in R, on the same file, runs here. Use it to check the counts in your Stata tables.
tabulate district area and find the district with the most urban households. In the R cell, change the last line to table(hh$district, hh$area), run it, and check that the two tables agree.generate, replace, recode, encode and labels
generate makes a new variable; replace changes values in one that exists. Stata refuses to generate a name already in use, which protects you from overwriting a variable by accident.
version 19
clear all
cd "C:\impactmojo\stata"
import delimited using "households.csv", clear
* 0/1 indicators from Yes/No text
generate toilet = (has_toilet == "Yes")
generate bank = (has_bank_account == "Yes")
generate annual_pc_exp = monthly_pc_exp * 12
* replace changes values in an existing variable
generate low_exp = 0
replace low_exp = 1 if monthly_pc_exp < 2000
* recode a number into bands, with value labels, into a new variable
recode head_edu_years (0 = 0 "None") (1/5 = 1 "Primary (1-5)") ///
(6/10 = 2 "Secondary (6-10)") (11/max = 3 "Higher (11+)"), generate(edu_cat)
tabulate edu_cat
* encode turns text into numbers with labels (codes follow alphabetical order)
encode caste, generate(caste_n)
encode area, generate(area_n)
tabulate caste_n
tabulate caste_n, nolabel // see the numbers behind the labels
Labels: say what a variable means
A variable label describes the variable. A value label names each code. You define a value label once with label define, then attach it with label values.
label variable toilet "Household has a toilet (1 = yes)" label variable monthly_pc_exp "Monthly per capita expenditure (Rs)" label define yesno 0 "No" 1 "Yes" label values toilet bank low_exp yesno describe toilet bank low_exp edu_cat caste_n tabulate toilet save households_clean, replace // Stata's own .dta format, labels included
- After
encode,tabulate caste_n, nolabelshould show codes 1 to 4, with 1 for General and 4 for ST, becauseencodeassigns codes in alphabetical order. tabulate toiletshould now show the words No and Yes, and the two frequencies should sum to 240.///at the end of a line tells a do-file that the command continues on the next line.
The R version of the indicator and the education bands:
large_hh equal to 1 when hh_size is 6 or more, give it the yesno label, and tabulate it against area_n. In the R cell, add table(hh$hh_size >= 6, hh$area) and compare.Group summaries: egen and collapse
Two commands summarise by group, and they differ in what they leave in memory. egen ..., by() adds a column and keeps every household. collapse replaces the data with one row per group. Use egen when you want to compare each household with its district; use collapse when you want a district table.
version 19 clear all cd "C:\impactmojo\stata" use households_clean, clear * egen: one new column, all 240 rows kept egen district_mean = mean(monthly_pc_exp), by(district) egen district_n = count(hh_id), by(district) generate above_district = monthly_pc_exp > district_mean list hh_id district monthly_pc_exp district_mean in 1/5 * collapse: one row per district (preserve/restore brings the households back) preserve collapse (mean) monthly_pc_exp hh_size toilet (count) n = hh_id, by(district) list export delimited using "district_summary.csv", replace restore describe, short // back to 240 observations
- After
collapse,listshould show 10 rows, one per district, and thencolumn should sum to 240. preservetakes a copy of the data andrestoreputs it back, so one do-file can make a summary table and carry on with the household data.
In R, ave() plays the part of egen, by() and aggregate() the part of collapse. Check the district means against your Stata list.
by() in the Stata collapse to by(caste) and run again; you should get four rows. In the R cell, change ~ district to ~ caste and hh$district to hh$caste in both places.merge (1:1 and m:1) and reshape
merge adds variables from a second Stata file (the using file) to the data in memory (the master), matching on key variables. Each household belongs to one district, and each district has many households, so attaching the district file is a many-to-one, m:1, merge. The using file must be a .dta, so save the districts first.
version 19 clear all cd "C:\impactmojo\stata" * the using file must be a Stata .dta import delimited using "districts.csv", clear isid district // stops with an error if district is not unique save districts, replace use households_clean, clear merge m:1 district using districts tabulate _merge assert _merge == 3 // stop the do-file if any row failed to match drop _merge
Read _merge every time
merge creates a variable _merge: 1 means the row was only in the master (a household whose district is missing from the district file), 2 means only in the using file, and 3 means matched. Here tabulate _merge should show 240 matched rows and nothing else. In real files a misspelt district (Purnea for Purnia) shows up as a 1 and a 2, which is why the assert line is worth keeping.
A 1:1 merge joins two files with one row per key on each side, such as two survey modules for the same households. To practise one, split the household file into two and join it back:
use households_clean, clear preserve keep hh_id shg_member received_transfer // the 'programme module' save hh_programme, replace restore drop shg_member received_transfer // the 'roster module' merge 1:1 hh_id using hh_programme tabulate _merge // expect 240 matched assert _merge == 3 drop _merge
The same district merge in R, with the R version of the _merge check:
reshape: wide and long
Long data has one row per unit per category (district by area); wide data has one row per unit and one column per category. reshape moves between them. You name the stub (the variable that repeats), i() (the unit) and j() (the category); add string when j is text.
use households_clean, clear collapse (mean) monthly_pc_exp, by(district area) list in 1/4 // long: 20 rows reshape wide monthly_pc_exp, i(district) j(area) string list // wide: 10 rows generate urban_rural_ratio = monthly_pc_expUrban / monthly_pc_expRural reshape long monthly_pc_exp, i(district) j(area) string // and back
After reshape wide you should have 10 rows and two new columns, monthly_pc_expRural and monthly_pc_expUrban. Every district in this file has both rural and urban households, so neither column has missing values. The R version runs here:
collapse (mean) monthly_pc_exp (count) n = hh_id, by(district head_gender), then reshape wide monthly_pc_exp n, i(district) j(head_gender) string. Two stubs reshape together. Look at nFemale: every district has some female-headed households, but in seven districts the mean rests on three or four of them. Say in a comment why you would not publish those seven means.Weights, and svyset for NFHS-style surveys
Stata takes weights in square brackets after the variable list. The two you will use most are [pw=...] (probability or sampling weights, the kind a survey supplies) and [aw=...] (analytic weights, used for averages of group means). Our course file has no sampling weight, but it shows why weights matter: each row is a household, and a household of eight people counts for more people than a household of one. Weighting by household size turns a per-household average into a per-person average.
use households_clean, clear mean monthly_pc_exp // average household mean monthly_pc_exp [pw = hh_size] // average person mean monthly_pc_exp [pw = hh_size], over(caste_n)
The two means differ when large households are poorer or richer than small ones. The R cell computes both, overall and by caste, so you can compare with your Stata output:
Declaring an NFHS design with svyset
NFHS is India's Demographic and Health Survey, and its files use the DHS recode variable names. The DHS Recode VII manual defines the three you need: v005 is the sample weight, "an 8 digit variable with 6 implied decimal places", to be divided by 1,000,000 before use; v021 is the primary sampling unit; and v022 is the "sample strata for sampling errors", the grouping of PSUs used for Taylor series variance estimates, which is the method svy uses by default. svyset declares this once, and every svy: command afterwards uses it.
* Use the NFHS women's (IR) Stata file you downloaded from the DHS Program use "your_nfhs_women_file.dta", clear generate wt = v005 / 1000000 svyset v021 [pw = wt], strata(v022) svyset // prints the design you declared * any 0/1 indicator you have built, e.g. currently using a modern method generate modern_use = (v313 == 3) svy: mean modern_use svy: mean modern_use, over(v024) // v024 is the DHS region variable; codebook v024 shows its labels * subpopulations: keep the full design, estimate for a subgroup codebook v025 // read the urban/rural codes first generate rural = (v025 == 2) // if codebook shows 2 = rural svy, subpop(rural): mean modern_use
- The DHS Program's Guide to DHS Statistics uses the same
v313 == 3definition of modern method use in its own Stata example, and notes that strata are not defined the same way in every survey: its example usesv023, and it suggests region by urban/rural (v024byv025) when that is how the sample was drawn. Check Appendix A of the survey report before you choose. - Use
subpop()for a subgroup. Dropping the other rows withkeep ifthrows away design information, and the standard errors come out wrong. - The output of
svy: meanreports the number of strata, the number of PSUs and the design degrees of freedom above the estimates. Read them: a PSU count far below what the survey report says means the wrong variable went intosvyset.
mean x [pw=wt] gives the right point estimate for NFHS, but only svy: mean x after svyset gives confidence intervals that allow for clustering and stratification.mean toilet and mean toilet [pw = hh_size]. Write one sentence in a comment saying which is the share of households with a toilet and which is the share of people living in one.regress, margins and graphs
regress fits a linear regression. The prefix i. tells Stata a variable is categorical, so it makes the dummy variables for you and leaves out the lowest code as the base (General, code 1, for caste_n). margins then turns the coefficients into predicted values you can explain to a programme manager.
use households_clean, clear regress monthly_pc_exp i.caste_n head_edu_years i.area_n margins caste_n // average predicted expenditure for each caste group marginsplot // graph of the margins with confidence intervals * with a 0/1 outcome the same command gives a linear probability model regress toilet i.caste_n head_edu_years i.area_n
By default margins caste_n sets every household to General, predicts and averages, then does the same for OBC, SC and ST, keeping each household's own education and area. The R cell below does exactly that by hand, so your Stata coefficients and margins should match it to the rupee.
Graphs
graph bar (mean) monthly_pc_exp, over(caste_n) ///
ytitle("Mean monthly per capita expenditure (Rs)")
graph export "exp_by_caste.png", replace width(1200)
twoway (scatter monthly_pc_exp head_edu_years) ///
(lfit monthly_pc_exp head_edu_years), ///
ytitle("Monthly per capita expenditure (Rs)") xtitle("Years of schooling, head")
graph export "exp_by_edu.png", replace width(1200)
The bar chart in R, for comparison with yours:
land_acres to the Stata regression and run margins area_n. In the R cell above, add + land_acres to the lm() formula and check that the coefficients still match Stata's.Logs, a master do-file and reproducible work
A log file records the commands you ran and everything in the Results window. Keep one for every analysis you report, so you can show where a number came from months later. log using name, text writes plain text that any editor can open.
* 00_master.do : runs the whole analysis from raw CSV to tables and graphs version 19 clear all set more off cd "C:\impactmojo\stata" capture log close log using "analysis_log", text replace do 03_newvars.do // import, clean, label, save households_clean.dta do 04_groups.do // district summary do 05_merge.do // attach district information do 07_regress.do // regression, margins, graphs log close
version 19at the top makes later releases of Stata interpret the do-file as Stata 19 did.capture log closecloses a log left open by an earlier run without stopping on an error if none is open.- Never edit the raw CSV. Every change goes in a do-file, so the path from raw data to result can be re-run.
- If you draw a random sample or bootstrap, put
set seed 20261006(any fixed number) before it, so the same draw comes out every time. - Use
assertfor things that must be true (assert hh_size >= 1,isid hh_id). A do-file that stops loudly is better than one that finishes with wrong numbers.
00_master.do from the do-files you wrote in this course, run it, then open analysis_log.log in a text editor and find the tabulate _merge table in it.Stata's own documentation
- Stata documentation: every manual as a free PDF, including Getting Started and the User's Guide.
- [SVY] svyset, [D] merge, [D] reshape and [R] margins.
- New in Stata 19.
Where next
R & Python for Development
Run the same steps in free software, live in your browser.
SPSS Syntax for Development Data
The same arc in SPSS syntax, including Complex Samples for NFHS.
Survey Design 101
Why surveys are stratified and clustered, and what that does to your standard errors.
Econometrics 101
What a regression coefficient means, and when it does not mean what you hope.