Skip to content
ImpactMojo ImpactMojo
Browse Membership
Back to Code Studio
Code Studio · Guided tool course

Stata syntax for development data

Write Stata do-files from your first import to a weighted regression: describe, generate, recode, labels, egen, collapse, merge, reshape, svyset for NFHS-style surveys, regress and margins, graphs and logs. Follow along in your own copy of Stata with two small CSV files, and check each step in R on this page.

Stata does not run in a browser. The grey boxes are do-file lines for you to type into your own Stata. This page does not show Stata output, because it cannot produce any; each module tells you what to look for on your screen instead. Where the same step can be done in R, an R cell runs it here on the same data, so you can check your Stata numbers against it. The first R run downloads the R engine once (about 7 MB).
Module 1 of 8

Stata, the do-file and your project folder

Stata is a paid statistics package from StataCorp, used widely in development economics, evaluation teams and survey firms. Most published replication files for NFHS, PLFS and IHDS analyses are Stata do-files, so reading Stata is useful even if you work in R.

Version and cost, checked 6 October 2026. The current release is Stata 19 (StataCorp's citation for it reads "StataCorp. 2025. Stata 19"). Annual licences also carry StataNow, which adds new features during the release. Stata comes in three editions, Stata/BE (basic, mid-sized datasets), Stata/SE (larger datasets) and Stata/MP (multicore, the largest datasets), and all three run on Windows, macOS and Linux. StataCorp's order page asks for your country before it shows a price, and buyers in many countries go through an authorised international reseller. Check with your university or organisation before you buy: many hold a site licence.

The five main windows

Stata's own Getting Started manual names five main windows: History (commands you have run), Results (output), Command (where you type one command and press Enter), Variables (the variables in memory) and Properties (details of the selected variable). The menus can build most commands for you, and the command they build appears in the Results window, which is a good way to learn syntax.

Work in a do-file from the first day

A do-file is a plain text file of commands that Stata runs from top to bottom. It is your record of what you did, and anyone with the same data can re-run it and get the same numbers. Open the Do-file Editor by clicking its button on the toolbar or by typing doedit in the Command window.

  1. Make a project folder, for example C:\impactmojo\stata on Windows or ~/impactmojo/stata on a Mac.
  2. Download the two course files into it: households.csv (240 households) and districts.csv (10 districts). Illustrative data, invented for teaching The district names are real places; every number is made up.
  3. In Stata, type doedit in the Command window and press Enter.
  4. Type the lines below into the new do-file and save it in your project folder as 01_setup.do.
  5. Click the Do button (Execute (do)) on the Do-file Editor toolbar. If you highlight some lines first, the same button runs only those lines.
Stata do-file: 01_setup.do
* 01_setup.do : first do-file for the Code Studio Stata course
version 19            // run as Stata 19, even in a later Stata
clear all             // start with nothing in memory
cd "C:\impactmojo\stata"   // change to YOUR project folder
pwd                   // print the working folder, to check
dir *.csv             // you should see households.csv and districts.csv

Lines starting with * are comments, and // starts a comment at the end of a line. Look in the Results window: pwd should print your folder, and dir should list both CSV files. If cd fails, the path is wrong; copy it from your file manager's address bar.

Exercise. Change the cd line to your own folder, run the do-file, then type help import delimited in the Command window. Every Stata command has a help file like this, with the syntax at the top and worked examples at the bottom.
Module 2 of 8

Import a CSV and look at it

import delimited reads comma-separated and tab-separated text files. With the variable names in the first row, as in our files, it needs only the file name. Then look before you change anything.

Stata do-file: 02_look.do
version 19
clear all
cd "C:\impactmojo\stata"
import delimited using "households.csv", clear

describe                         // variables, types, number of observations
codebook district caste, compact
list in 1/5                      // first five households
summarize monthly_pc_exp hh_size head_edu_years
summarize monthly_pc_exp, detail // adds percentiles and the median
tabulate caste
tabulate caste area              // two-way table
tabulate caste area, row         // row percentages

What to look for

The same look in R, on the same file, runs here. Use it to check the counts in your Stata tables.

Exercise. In Stata, run tabulate district area and find the district with the most urban households. In the R cell, change the last line to table(hh$district, hh$area), run it, and check that the two tables agree.
Module 3 of 8

generate, replace, recode, encode and labels

generate makes a new variable; replace changes values in one that exists. Stata refuses to generate a name already in use, which protects you from overwriting a variable by accident.

Stata do-file: 03_newvars.do
version 19
clear all
cd "C:\impactmojo\stata"
import delimited using "households.csv", clear

* 0/1 indicators from Yes/No text
generate toilet = (has_toilet == "Yes")
generate bank   = (has_bank_account == "Yes")
generate annual_pc_exp = monthly_pc_exp * 12

* replace changes values in an existing variable
generate low_exp = 0
replace  low_exp = 1 if monthly_pc_exp < 2000

* recode a number into bands, with value labels, into a new variable
recode head_edu_years (0 = 0 "None") (1/5 = 1 "Primary (1-5)") ///
    (6/10 = 2 "Secondary (6-10)") (11/max = 3 "Higher (11+)"), generate(edu_cat)
tabulate edu_cat

* encode turns text into numbers with labels (codes follow alphabetical order)
encode caste, generate(caste_n)
encode area,  generate(area_n)
tabulate caste_n
tabulate caste_n, nolabel        // see the numbers behind the labels

Labels: say what a variable means

A variable label describes the variable. A value label names each code. You define a value label once with label define, then attach it with label values.

Stata do-file: add to 03_newvars.do
label variable toilet "Household has a toilet (1 = yes)"
label variable monthly_pc_exp "Monthly per capita expenditure (Rs)"

label define yesno 0 "No" 1 "Yes"
label values toilet bank low_exp yesno

describe toilet bank low_exp edu_cat caste_n
tabulate toilet
save households_clean, replace   // Stata's own .dta format, labels included

The R version of the indicator and the education bands:

Exercise. Make a variable large_hh equal to 1 when hh_size is 6 or more, give it the yesno label, and tabulate it against area_n. In the R cell, add table(hh$hh_size >= 6, hh$area) and compare.
Module 4 of 8

Group summaries: egen and collapse

Two commands summarise by group, and they differ in what they leave in memory. egen ..., by() adds a column and keeps every household. collapse replaces the data with one row per group. Use egen when you want to compare each household with its district; use collapse when you want a district table.

Stata do-file: 04_groups.do
version 19
clear all
cd "C:\impactmojo\stata"
use households_clean, clear

* egen: one new column, all 240 rows kept
egen district_mean = mean(monthly_pc_exp), by(district)
egen district_n    = count(hh_id), by(district)
generate above_district = monthly_pc_exp > district_mean
list hh_id district monthly_pc_exp district_mean in 1/5

* collapse: one row per district (preserve/restore brings the households back)
preserve
collapse (mean) monthly_pc_exp hh_size toilet (count) n = hh_id, by(district)
list
export delimited using "district_summary.csv", replace
restore
describe, short    // back to 240 observations

In R, ave() plays the part of egen, by() and aggregate() the part of collapse. Check the district means against your Stata list.

Exercise. Change the by() in the Stata collapse to by(caste) and run again; you should get four rows. In the R cell, change ~ district to ~ caste and hh$district to hh$caste in both places.
Module 5 of 8

merge (1:1 and m:1) and reshape

merge adds variables from a second Stata file (the using file) to the data in memory (the master), matching on key variables. Each household belongs to one district, and each district has many households, so attaching the district file is a many-to-one, m:1, merge. The using file must be a .dta, so save the districts first.

Stata do-file: 05_merge.do
version 19
clear all
cd "C:\impactmojo\stata"

* the using file must be a Stata .dta
import delimited using "districts.csv", clear
isid district                 // stops with an error if district is not unique
save districts, replace

use households_clean, clear
merge m:1 district using districts
tabulate _merge
assert _merge == 3            // stop the do-file if any row failed to match
drop _merge

Read _merge every time

merge creates a variable _merge: 1 means the row was only in the master (a household whose district is missing from the district file), 2 means only in the using file, and 3 means matched. Here tabulate _merge should show 240 matched rows and nothing else. In real files a misspelt district (Purnea for Purnia) shows up as a 1 and a 2, which is why the assert line is worth keeping.

A 1:1 merge joins two files with one row per key on each side, such as two survey modules for the same households. To practise one, split the household file into two and join it back:

Stata do-file: a 1:1 merge
use households_clean, clear
preserve
keep hh_id shg_member received_transfer      // the 'programme module'
save hh_programme, replace
restore
drop shg_member received_transfer            // the 'roster module'
merge 1:1 hh_id using hh_programme
tabulate _merge                              // expect 240 matched
assert _merge == 3
drop _merge

The same district merge in R, with the R version of the _merge check:

reshape: wide and long

Long data has one row per unit per category (district by area); wide data has one row per unit and one column per category. reshape moves between them. You name the stub (the variable that repeats), i() (the unit) and j() (the category); add string when j is text.

Stata do-file: reshape
use households_clean, clear
collapse (mean) monthly_pc_exp, by(district area)
list in 1/4                                        // long: 20 rows
reshape wide monthly_pc_exp, i(district) j(area) string
list                                               // wide: 10 rows
generate urban_rural_ratio = monthly_pc_expUrban / monthly_pc_expRural
reshape long monthly_pc_exp, i(district) j(area) string   // and back

After reshape wide you should have 10 rows and two new columns, monthly_pc_expRural and monthly_pc_expUrban. Every district in this file has both rural and urban households, so neither column has missing values. The R version runs here:

Exercise. Repeat the reshape by head of household's gender: collapse (mean) monthly_pc_exp (count) n = hh_id, by(district head_gender), then reshape wide monthly_pc_exp n, i(district) j(head_gender) string. Two stubs reshape together. Look at nFemale: every district has some female-headed households, but in seven districts the mean rests on three or four of them. Say in a comment why you would not publish those seven means.
Module 6 of 8

Weights, and svyset for NFHS-style surveys

Stata takes weights in square brackets after the variable list. The two you will use most are [pw=...] (probability or sampling weights, the kind a survey supplies) and [aw=...] (analytic weights, used for averages of group means). Our course file has no sampling weight, but it shows why weights matter: each row is a household, and a household of eight people counts for more people than a household of one. Weighting by household size turns a per-household average into a per-person average.

Stata do-file: 06_weights.do
use households_clean, clear
mean monthly_pc_exp                    // average household
mean monthly_pc_exp [pw = hh_size]     // average person
mean monthly_pc_exp [pw = hh_size], over(caste_n)

The two means differ when large households are poorer or richer than small ones. The R cell computes both, overall and by caste, so you can compare with your Stata output:

Declaring an NFHS design with svyset

NFHS is India's Demographic and Health Survey, and its files use the DHS recode variable names. The DHS Recode VII manual defines the three you need: v005 is the sample weight, "an 8 digit variable with 6 implied decimal places", to be divided by 1,000,000 before use; v021 is the primary sampling unit; and v022 is the "sample strata for sampling errors", the grouping of PSUs used for Taylor series variance estimates, which is the method svy uses by default. svyset declares this once, and every svy: command afterwards uses it.

Stata do-file: NFHS women's file (not part of the course data)
* Use the NFHS women's (IR) Stata file you downloaded from the DHS Program
use "your_nfhs_women_file.dta", clear

generate wt = v005 / 1000000
svyset v021 [pw = wt], strata(v022)
svyset                                  // prints the design you declared

* any 0/1 indicator you have built, e.g. currently using a modern method
generate modern_use = (v313 == 3)
svy: mean modern_use
svy: mean modern_use, over(v024)        // v024 is the DHS region variable; codebook v024 shows its labels

* subpopulations: keep the full design, estimate for a subgroup
codebook v025                           // read the urban/rural codes first
generate rural = (v025 == 2)            // if codebook shows 2 = rural
svy, subpop(rural): mean modern_use
Weights change estimates; the design changes standard errors. mean x [pw=wt] gives the right point estimate for NFHS, but only svy: mean x after svyset gives confidence intervals that allow for clustering and stratification.
Exercise. In the course file, run mean toilet and mean toilet [pw = hh_size]. Write one sentence in a comment saying which is the share of households with a toilet and which is the share of people living in one.
Module 7 of 8

regress, margins and graphs

regress fits a linear regression. The prefix i. tells Stata a variable is categorical, so it makes the dummy variables for you and leaves out the lowest code as the base (General, code 1, for caste_n). margins then turns the coefficients into predicted values you can explain to a programme manager.

Stata do-file: 07_regress.do
use households_clean, clear
regress monthly_pc_exp i.caste_n head_edu_years i.area_n
margins caste_n          // average predicted expenditure for each caste group
marginsplot              // graph of the margins with confidence intervals

* with a 0/1 outcome the same command gives a linear probability model
regress toilet i.caste_n head_edu_years i.area_n

By default margins caste_n sets every household to General, predicts and averages, then does the same for OBC, SC and ST, keeping each household's own education and area. The R cell below does exactly that by hand, so your Stata coefficients and margins should match it to the rupee.

Graphs

Stata do-file: graphs
graph bar (mean) monthly_pc_exp, over(caste_n) ///
    ytitle("Mean monthly per capita expenditure (Rs)")
graph export "exp_by_caste.png", replace width(1200)

twoway (scatter monthly_pc_exp head_edu_years) ///
       (lfit monthly_pc_exp head_edu_years),   ///
       ytitle("Monthly per capita expenditure (Rs)") xtitle("Years of schooling, head")
graph export "exp_by_edu.png", replace width(1200)

The bar chart in R, for comparison with yours:

Exercise. Add land_acres to the Stata regression and run margins area_n. In the R cell above, add + land_acres to the lm() formula and check that the coefficients still match Stata's.
Module 8 of 8

Logs, a master do-file and reproducible work

A log file records the commands you ran and everything in the Results window. Keep one for every analysis you report, so you can show where a number came from months later. log using name, text writes plain text that any editor can open.

Stata do-file: 00_master.do
* 00_master.do : runs the whole analysis from raw CSV to tables and graphs
version 19
clear all
set more off
cd "C:\impactmojo\stata"

capture log close
log using "analysis_log", text replace

do 03_newvars.do      // import, clean, label, save households_clean.dta
do 04_groups.do       // district summary
do 05_merge.do        // attach district information
do 07_regress.do      // regression, margins, graphs

log close
Exercise. Build 00_master.do from the do-files you wrote in this course, run it, then open analysis_log.log in a text editor and find the tabulate _merge table in it.

Stata's own documentation

Where next

→

R & Python for Development

Run the same steps in free software, live in your browser.

→

SPSS Syntax for Development Data

The same arc in SPSS syntax, including Complex Samples for NFHS.

→

Survey Design 101

Why surveys are stratified and clustered, and what that does to your standard errors.

→

Econometrics 101

What a regression coefficient means, and when it does not mean what you hope.