SPSS syntax for development data
Use SPSS through syntax files you can save and re-run: GET DATA, variable and value labels, RECODE, COMPUTE, FREQUENCIES, DESCRIPTIVES, CROSSTABS, AGGREGATE, MATCH FILES, WEIGHT BY and the Complex Samples procedures for NFHS, and REGRESSION. Learn to paste syntax from the menus, follow along in your own SPSS with two small CSV files, and check each step in R on this page.
SPSS, the Syntax Editor and why syntax
IBM SPSS Statistics is a paid statistics package used across government departments, NGOs, universities and market research firms in South Asia. Most people learn it through the menus. This course teaches the syntax underneath the menus, because a syntax file is a record of your analysis that you, a colleague or a reviewer can run again and get the same tables.
Open a syntax window
- Make a project folder, for example
C:\impactmojo\spss. - Download the two course files into it: households.csv (240 households) and districts.csv (10 districts). Illustrative data, invented for teaching The district names are real places; every number is made up.
- In SPSS choose File > New > Syntax. The Syntax Editor opens.
- Type or paste a block of syntax, select it, and click the Run button (the right-pointing triangle) on the Syntax Editor toolbar. The Run menu also has All, Selection, To End and Step Through.
- Save the file with File > Save; SPSS syntax files end in
.sps.
Rules of SPSS syntax
- Every command ends with a full stop. A missing full stop is the most common reason a block does nothing or errors.
- Subcommands start with a slash:
/TABLES=,/STATISTICS=. - A comment starts with an asterisk at the beginning of a line and ends with a full stop.
- Commands such as
COMPUTEandRECODEwait for the next procedure before they change the data.EXECUTE.makes them run straight away, so you can see the new column in the Data Editor.
* My first SPSS syntax file. on the first line, and save it as 01_import.sps in your project folder. You will fill it in the next module.Read a CSV with GET DATA
GET DATA /TYPE=TXT reads a delimited text file. You tell it the delimiter, which row the data start on (row 2, after the header), and the name and format of each column. F4.0 is a number four characters wide with no decimals; A12 is text up to 12 characters.
* 01_import.sps : read the course household file.
* Change the FILE path to your own folder.
GET DATA
/TYPE=TXT
/FILE='C:\impactmojo\spss\households.csv'
/DELIMITERS=","
/QUALIFIER='"'
/ARRANGEMENT=DELIMITED
/FIRSTCASE=2
/VARIABLES=
hh_id F4.0
district A12
area A5
caste A8
head_gender A6
head_edu_years F2.0
hh_size F2.0
monthly_pc_exp F6.0
land_acres F5.2
has_toilet A3
has_bank_account A3
shg_member A3
received_transfer A3.
DATASET NAME households WINDOW=FRONT.
DISPLAY DICTIONARY.
LIST VARIABLES=hh_id district caste monthly_pc_exp /CASES=FROM 1 TO 5.
- The Data Editor should show 240 rows.
DISPLAY DICTIONARYlists the 13 variables with their formats. - If a column is blank or wrong, check its format: text read as
Fcomes out as system-missing, shown as a dot. - Dialogs that read and change data, and most analysis dialogs, have a Paste button beside OK. The rest of this module explains it.
Paste syntax from any dialog
IBM's documentation describes this as the easiest way to build a syntax file: make your selections in a dialog, then click Paste. The command goes into the syntax window (a new one opens if none is open), where you can run it, edit it and save it. Pasting at each step of an analysis builds a file that repeats the whole analysis later. Use the menus to find a command, and keep the syntax as your record.
caste into the Variable(s) box and click Paste. Compare the pasted command with the FREQUENCIES line in the next module.FREQUENCIES, DESCRIPTIVES and CROSSTABS
FREQUENCIES VARIABLES=caste area head_gender. DESCRIPTIVES VARIABLES=monthly_pc_exp hh_size head_edu_years /STATISTICS=MEAN STDDEV MIN MAX. FREQUENCIES VARIABLES=monthly_pc_exp /FORMAT=NOTABLE /STATISTICS=MEAN MEDIAN /PERCENTILES=25 75. CROSSTABS /TABLES=caste BY area /CELLS=COUNT ROW /STATISTICS=CHISQ.
What to look for
- The caste frequency table should have four rows (General, OBC, SC, ST) and a Total of 240.
/FORMAT=NOTABLEstops SPSS printing one row per distinct expenditure value and keeps only the statistics.- In the crosstab,
ROWgives the percentage of each caste group living in rural and urban areas. The Chi-Square Tests table follows it; read the Pearson Chi-Square row. - The menu route for the crosstab is Analyze > Descriptive Statistics > Crosstabs.
The same three summaries in R, on the same file. Your SPSS counts, means and Pearson chi-square should match these.
/TABLES=head_gender BY caste with /CELLS=COUNT COLUMN. In the R cell, change tab to table(hh$head_gender, hh$caste) and prop.table(tab, 1) to prop.table(tab, 2), then run both.COMPUTE, RECODE and labels
COMPUTE makes or changes a variable from an expression. RECODE ... INTO maps old values to new ones in a new variable, which keeps the original intact. Then give both variables and values readable labels.
* 0/1 indicators from Yes/No text.
COMPUTE toilet = (has_toilet = 'Yes').
COMPUTE bank = (has_bank_account = 'Yes').
COMPUTE annual_pc_exp = monthly_pc_exp * 12.
* Education bands in a new variable.
RECODE head_edu_years (0=0) (1 THRU 5=1) (6 THRU 10=2) (11 THRU HIGHEST=3) INTO edu_cat.
* Text categories to numeric codes.
RECODE caste ('General'=1) ('OBC'=2) ('SC'=3) ('ST'=4) INTO caste_n.
EXECUTE.
VARIABLE LABELS
toilet 'Household has a toilet (1 = yes)'
edu_cat 'Schooling of household head, banded'
caste_n 'Caste group'
monthly_pc_exp 'Monthly per capita expenditure (Rs)'.
VALUE LABELS
toilet bank 0 'No' 1 'Yes'
/edu_cat 0 'None' 1 'Primary (1-5)' 2 'Secondary (6-10)' 3 'Higher (11+)'
/caste_n 1 'General' 2 'OBC' 3 'SC' 4 'ST'.
FREQUENCIES VARIABLES=toilet edu_cat caste_n.
SAVE OUTFILE='C:\impactmojo\spss\households_clean.sav'.
(has_toilet = 'Yes')is a logical expression: it is 1 when true and 0 when false. String comparisons in SPSS are case-sensitive, so'yes'would match nothing in this file.- The frequency table for
edu_catshould show the four labels, summing to 240, and no missing values. - A
.savfile keeps your labels, which a CSV cannot.
The R version of the indicator and the bands, to check your frequencies:
COMPUTE large_hh = (hh_size >= 6)., give it the No/Yes value labels, and run CROSSTABS /TABLES=large_hh BY area. In the R cell, add table(hh$hh_size >= 6, hh$area) and compare.AGGREGATE, MATCH FILES and restructuring
AGGREGATE summarises by group. With MODE=ADDVARIABLES it adds the group value to every case and keeps all 240 households; with an output file it writes one row per group.
GET FILE='C:\impactmojo\spss\households_clean.sav'. * District mean on every household row. AGGREGATE /OUTFILE=* MODE=ADDVARIABLES /BREAK=district /district_mean=MEAN(monthly_pc_exp). * One row per district, in a separate file. AGGREGATE /OUTFILE='C:\impactmojo\spss\district_summary.sav' /BREAK=district /mean_exp=MEAN(monthly_pc_exp) /mean_size=MEAN(hh_size) /n=N.
Open district_summary.sav afterwards: it should have 10 rows and the n column should sum to 240. The R cell gives the same district means:
MATCH FILES: attach district information
Each household belongs to one district and each district has many households, so the district file is a table lookup: /TABLE in MATCH FILES. Both files must be sorted by the key, and you should give a text key the same width in both files, which is why both GET DATA commands give district the format A12.
GET DATA
/TYPE=TXT
/FILE='C:\impactmojo\spss\districts.csv'
/DELIMITERS=","
/QUALIFIER='"'
/ARRANGEMENT=DELIMITED
/FIRSTCASE=2
/VARIABLES=
district A12
state A16
region A8
programme_phase F1.0
field_team A8.
SORT CASES BY district.
SAVE OUTFILE='C:\impactmojo\spss\districts.sav'.
GET FILE='C:\impactmojo\spss\households_clean.sav'.
SORT CASES BY district.
MATCH FILES
/FILE=*
/TABLE='C:\impactmojo\spss\districts.sav'
/IN=in_district
/BY district.
EXECUTE.
FREQUENCIES VARIABLES=in_district state.
/IN=in_districtmakes a 0/1 flag: 1 when the household found its district in the table. Its frequency table should show 240 cases coded 1. In real files, a misspelt district (Purnea for Purnia) appears as a 0.- For two files with one row per household each (a
1:1join, such as two modules of the same survey), use two/FILEsubcommands:MATCH FILES /FILE='roster.sav' /FILE='programme.sav' /BY hh_id.
Restructure: cases to variables and back
GET FILE='C:\impactmojo\spss\households_clean.sav'. AGGREGATE /OUTFILE=* /BREAK=district area /monthly_pc_exp=MEAN(monthly_pc_exp). SORT CASES BY district area. CASESTOVARS /ID=district /INDEX=area.
After CASESTOVARS the active file should have 10 rows and one expenditure column per area. VARSTOCASES goes the other way; read the new column names in Variable View and list them after /MAKE monthly_pc_exp FROM.
/BREAK of the district summary to caste and add /toilet_share=MEAN(toilet). You should get four rows. Change the R cell to group by caste and check the means.WEIGHT BY, and Complex Samples for NFHS
WEIGHT BY makes every following procedure count each case as many times as its weight. Our course file has no sampling weight, but weighting each household by its size shows what a weight does: the per-household average becomes a per-person average.
GET FILE='C:\impactmojo\spss\households_clean.sav'. DESCRIPTIVES VARIABLES=monthly_pc_exp /STATISTICS=MEAN. WEIGHT BY hh_size. DESCRIPTIVES VARIABLES=monthly_pc_exp /STATISTICS=MEAN. WEIGHT OFF.
Compare the two means and the two N values. With WEIGHT BY hh_size the N becomes the number of people, because SPSS treats each household as hh_size copies of itself. The R cell computes both:
An NFHS analysis plan
NFHS uses the DHS recode variable names. The DHS Recode VII manual defines v005 as the sample weight with six implied decimal places (divide by 1,000,000), v021 as the primary sampling unit and v022 as the "sample strata for sampling errors". In SPSS you save these in an analysis plan file with CSPLAN, then name the plan in every CS procedure. The pattern below follows the SPSS example in the DHS Program's Guide to DHS Statistics.
* Open the NFHS women's (IR) SPSS file you downloaded from the DHS Program. GET FILE='C:\nfhs\your_nfhs_women_file.sav'. COMPUTE wt = v005 / 1000000. COMPUTE modern_use = (v313 = 3). EXECUTE. CSPLAN ANALYSIS /PLAN FILE='C:\nfhs\nfhs_ir.csplan' /PLANVARS ANALYSISWEIGHT=wt /DESIGN STRATA=v022 CLUSTER=v021 /ESTIMATOR TYPE=WR. CSDESCRIPTIVES /PLAN FILE='C:\nfhs\nfhs_ir.csplan' /SUMMARY VARIABLES=modern_use /SUBPOP TABLE=v024 DISPLAY=LAYERED /MEAN /STATISTICS SE CIN /MISSING SCOPE=ANALYSIS CLASSMISSING=EXCLUDE.
- The menu route to build the plan is Analyze > Complex Samples > Prepare for Analysis; its wizard can paste the
CSPLANcommand. The estimates are under Analyze > Complex Samples > Descriptives, Frequencies, Crosstabs and General Linear Model. - The DHS guide's own example uses
v023as the stratum and notes that strata are not defined the same way in every survey; it suggests region by urban/rural (v024byv025) when that is how the sample was drawn. Check Appendix A of the survey report before you choose. /SUBPOP TABLE=v024estimates for each value ofv024while keeping the whole design, which is the right way to get subgroup estimates. Filtering cases out withSELECT IFfirst would drop design information.- IBM's Complex Samples manual (version 31) says these procedures are included in SPSS Statistics Premium Edition or the Complex Samples option. Check your licence before you plan an NFHS analysis around them.
DESCRIPTIVES VARIABLES=toilet with and without WEIGHT BY hh_size. Write a comment in your syntax file saying which mean is the share of households with a toilet and which is the share of people.REGRESSION and a chart
SPSS REGRESSION takes numeric predictors only, so a categorical variable needs dummy variables, one for each group except the reference group (General here). The menu route is Analyze > Regression > Linear.
GET FILE='C:\impactmojo\spss\households_clean.sav'. COMPUTE obc = (caste = 'OBC'). COMPUTE sc = (caste = 'SC'). COMPUTE st = (caste = 'ST'). COMPUTE urban = (area = 'Urban'). EXECUTE. REGRESSION /STATISTICS COEFF OUTS R ANOVA CI(95) /DEPENDENT monthly_pc_exp /METHOD=ENTER obc sc st head_edu_years urban. GRAPH /BAR(SIMPLE)=MEAN(monthly_pc_exp) BY caste.
In the Coefficients table, read the column labelled B: each caste coefficient is the difference from General households with the same schooling and area. The R cell fits the same model with the same dummies, so the B column in your output should match its Estimate column, and R Square should match too.
land_acres to /METHOD=ENTER and run again. In the R cell, add + land_acres to the formula and check that both give the same coefficients.A master syntax file and saved output
Keep each step in its own .sps file and run them in order from one master file. INSERT runs another syntax file.
* 00_master.sps : the whole analysis, from CSV to tables. INSERT FILE='C:\impactmojo\spss\01_import.sps'. INSERT FILE='C:\impactmojo\spss\03_recode.sps'. INSERT FILE='C:\impactmojo\spss\04_aggregate.sps'. INSERT FILE='C:\impactmojo\spss\05_match.sps'. INSERT FILE='C:\impactmojo\spss\07_regression.sps'. OUTPUT SAVE OUTFILE='C:\impactmojo\spss\analysis_output.spv'.
- Never edit the raw CSV, and never fix a value by typing in the Data Editor. Put every change in syntax so it can be re-run and checked.
- Comment each block with what it does and why:
* Drop test interviews done before 1 March. - Use full paths, or set one folder at the top and keep every file in it, so the syntax runs on a colleague's computer after one edit.
- Save output (
.spv) alongside the syntax that made it, with the same date in both file names.
00_master.sps from the files you wrote, close SPSS, open it again, and run only the master file. If every table comes back, your analysis is reproducible.IBM's own documentation
- IBM SPSS Statistics, with a free trial and links to the documentation.
- CSPLAN examples in the Command Syntax Reference.
- Guide to DHS Statistics: analyzing DHS data, with weights and complex designs in SPSS, Stata and R.
Where next
jamovi and JASP
Free point-and-click statistics with an SPSS-like layout, built on R.
Stata Syntax for Development Data
The same arc in Stata do-files, including svyset for NFHS.
R & Python for Development
Run the same steps in free software, live in your browser.
Statistics Without Code 101
The ideas behind the tables, without software.
Survey Design 101
Why NFHS is stratified and clustered, and what that does to your standard errors.