ImpactMojoQuantitative Methods 101www.impactmojo.in
ImpactMojo 101 Series · Free Forever
Quantitative
Methods
101
Reading and commissioning numbers in development work: measurement, sampling, survey weights, intervals, tests, regression and power, worked through with NFHS, PLFS and HCES data
Quantitative MethodsSouth Asia Focus100 SlidesFree Forever
ImpactMojoQuantitative Methods 101www.impactmojo.in
What we cover
01
Why quantitative methods
Slides 3–9
02
Variables and measurement
Slides 10–18
03
Describing data
Slides 19–27
04
Distributions
Slides 28–35
05
Sampling and sampling error
Slides 36–44
06
Survey weights and design effects
Slides 45–54
07
Intervals, tests and p-values
Slides 55–64
08
Correlation, causation, regression
Slides 65–74
09
Effect sizes and power
Slides 75–83
10
Reading and commissioning analysis
Slides 84–92
11
Errors in published statistics
Slides 93–99
ImpactMojoQuantitative Methods 101www.impactmojo.in
01
Section One
Why quantitative methods
ImpactMojoQuantitative Methods 101www.impactmojo.in
What this course is for
Most development practitioners will never run a regression themselves. Almost all of them will read one, approve one, fund one or defend a budget on the strength of one. A district officer quotes NFHS stunting figures, a programme manager reads an evaluation report, a foundation officer signs off a baseline survey of 2,400 households. Each of those moments asks for the same skill: knowing what a number can and cannot carry.
This deck teaches that skill at the level of intuition and arithmetic. Every worked example shows its sums, so you can redo them on paper and check the next report you read.
By the end you should be able to
  • Say what kind of variable an indicator is and which summaries suit it
  • Apply NFHS and PLFS weights correctly and explain why they exist
  • Compute a standard error and a 95% confidence interval for a proportion
  • Read a p-value and a regression table without overclaiming
  • Ask for a power calculation and judge whether it is credible
  • Spot the common errors in published Indian statistics
ImpactMojoQuantitative Methods 101www.impactmojo.in
Three numbers that already shape policy
29.3%
children under 5 stunted, India
NFHS-6 (2023–24) India Fact Sheet, IIPS/MoHFW
40.3%
female labour force participation, age 15+, usual status
MoSPI, PLFS annual estimate, calendar year 2024
₹4,122
average monthly per capita consumption, rural India
MoSPI, HCES 2023–24, without imputation
Each of these figures drives budgets: Poshan Abhiyaan targets, labour policy debates, poverty line arguments. Each is also an estimate. It rests on a definition (stunting is height-for-age below −2 standard deviations of the WHO growth standard), a sample (679,238 households in NFHS-6), a reference period (the 365 days before the PLFS interview for usual status) and a margin of error. A reader who knows those four things can use the number well. A reader who does not will eventually compare two figures that were never measuring the same thing.
Quantitative literacy here means asking of every figure: defined how, counted among whom, when, and how precisely.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Four questions to ask of any number
QuestionWhat you are checkingExample of the trap
What exactly was counted?The definition, the indicator formula, the reference periodPLFS reports unemployment on usual status and on current weekly status; the two differ by almost 2 points
Who was asked, and who was left out?The population, the sampling frame, non-responseNFHS covers households; people in hostels, barracks and prisons are outside the frame
How precise is it?Sample size, standard error, design effect, confidence intervalA district estimate from 300 children carries an interval of roughly ±5 to ±8 points
Compared with what?Baseline, other places, other rounds, a counterfactualA fall between two survey rounds can come from a changed questionnaire
These four questions organise the whole deck. Sections 2 to 4 deal with what was counted and how to summarise it. Sections 5 to 7 deal with who was asked and how precise the answer is. Sections 8 and 9 deal with comparisons, causes and the size of effects. Sections 10 and 11 turn the questions into a checklist you can use on a report this week.
ImpactMojoQuantitative Methods 101www.impactmojo.in
From a question to a defensible number
01
QUESTION: what do we need to know?
→
02
MEASURE: define and operationalise
→
03
SAMPLE: select who is asked
→
04
ESTIMATE: weight and summarise
→
05
QUANTIFY ERROR: SE, CI, design effect
→
06
INTERPRET: compare, explain, decide
Errors enter at every stage, and later stages cannot repair earlier ones. A brilliant regression on a badly worded question is still an answer to the wrong question. A large sample drawn from an incomplete frame is a precise estimate of the wrong population. Most arguments about a published figure turn out to be about the first three boxes, which is why this deck spends as long on measurement and sampling as on testing.
Where reviewers look first
When you review a quantitative report, start at the left of this chain. Read the questionnaire item before the result table. Read the sampling section before the conclusions. Ask how the weights were built before you ask whether the coefficient is significant. It is quicker, and it catches more.
ImpactMojoQuantitative Methods 101www.impactmojo.in
One survey, two unemployment rates
Unemployment rate, age 15+, India (%), by reference period
MoSPI, PLFS annual estimates (calendar years), via eSankhyiki API
The same PLFS interviews yield two unemployment rates. Usual status (principal plus subsidiary, PS+SS) classifies a person by their activity over the 365 days before the survey; anyone who worked for 30 days or more in a subsidiary capacity counts as employed. Current weekly status looks only at the seven days before the interview.
For calendar year 2024 the first gives 3.2% and the second 4.9%. Neither is wrong. They answer different questions: chronic joblessness versus joblessness last week, which catches seasonal gaps in farm work.
Quote the status with the number, every time.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Reader, commissioner, analyst
Reader
Reads published statistics and evaluation reports. Needs to know what an indicator means, how precise it is, and whether a claimed difference is real. This is most practitioners, most of the time.
Commissioner
Writes terms of reference, approves sample sizes, accepts deliverables. Needs to ask for a power calculation, a weighting note and the code, and to judge the answers.
Analyst
Runs the numbers. Needs software skills on top of everything here; the linked decks on Econometrics, Exploratory Data Analysis and Multivariate Analysis go further.
Anyone holding unit-level data about people is a data fiduciary under the Digital Personal Data Protection Act 2023, whose duties apply, with the 2025 Rules, from 13 May 2027. From that date section 17(2)(b) exempts processing for research, archiving or statistical purposes only where its conditions are met, including that the data is not used for decisions about a specific person.
ImpactMojoQuantitative Methods 101www.impactmojo.in
02
Section Two
Variables and measurement
ImpactMojoQuantitative Methods 101www.impactmojo.in
Cases, variables and the unit of analysis
Case and variable
A case is one row of a dataset: a household, a person, a child, a village. A variable is one column: a characteristic recorded for every case, such as age, caste category or monthly consumption.
The unit of analysis is the kind of case your question is about. It decides which file you open and which weight you use. A question about how many households have piped water is answered on the household file; a question about how many people live in such households is answered on the person file, or by weighting households by their size.
How Indian surveys are organised
NFHS data, like all DHS data, comes as separate recode files: household (HR), household member (PR), women (IR), men (MR) and children (KR, BR). PLFS ships a household-level file and a person-level file linked by a common key of quarter, visit, FSU serial number, hamlet group, second-stage stratum and household number. Merging them wrongly, or analysing children from the women's file without understanding that each row is a birth, is one of the most frequent errors in student and consultant work alike.
Write the unit of analysis into the first line of every analysis plan.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Nominal, ordinal, interval, ratio
LevelWhat the values tell youSouth Asian exampleSensible summaries
NominalCategories with no orderReligion; state; social group (SC, ST, OBC, others); type of cooking fuelCounts, percentages, mode
OrdinalOrdered categories, unequal or unknown gapsWealth quintile; education level completed; a five-point agreement scaleMedian, percentiles, percentages by category
IntervalEqual gaps, arbitrary zeroTemperature in °C; a test score on a scaled metric; a height-for-age z-scoreMean, SD, differences
RatioEqual gaps and a true zeroMonthly consumption in ₹; land owned in acres; number of children; days workedEverything above, plus ratios and percentage change
The level limits the arithmetic. Saying one household spends twice as much as another makes sense for consumption (ratio) and makes no sense for wealth quintile (ordinal): the richest quintile is not "five times" the poorest. The classification was set out by the psychologist S.S. Stevens in 1946; in practice, the useful question is simply which operations the numbers will bear.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Codes are labels, and averaging them is meaningless
Survey files store categories as numbers. In a typical layout social group might be coded 1 for ST, 2 for SC, 3 for OBC and 9 for others. Software will happily compute the mean of that column and report 2.7. The number describes nothing: change the coding order and it changes, while the households stay the same.
For nominal variables report the share of cases in each category. For ordinal variables report the distribution across categories or the median category. A cross-tabulation, covered in Section 3, is usually the right first table.
Ordinal scales in programme surveys
Satisfaction and attitude items ("strongly agree" to "strongly disagree") are ordinal. Averaging them as if the gaps were equal is common and sometimes defensible when many items are summed into a scale, but a single item should be reported as a distribution. A mean of 3.4 hides whether half the respondents strongly disagreed.
Watch the special codes
Codes like 98 ("don't know") and 99 ("missing") sit inside numeric columns. Leave them in and a mean age or a mean landholding can jump. DHS files flag implausible anthropometric values with special high codes in the z-score variables; the recode manual lists them, and they must be dropped first.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Interval and ratio variables in practice
Interval scale
Equal distances mean equal differences, but zero is a convention. A z-score of 0 means "at the reference median"; it does not mean "no height". You can say a child moved from −2.4 to −1.9, a gain of 0.5 SD; you cannot say one child is twice as stunted as another.
Ratio scale
Zero means none of the quantity. Consumption, income, hours worked, acres owned and number of antenatal visits are ratio variables, so percentage changes and ratios between groups are meaningful.
Why this matters for reporting
HCES 2023–24 reports average monthly per capita consumption of ₹6,996 in urban India and ₹4,122 in rural India (MoSPI). Because consumption is a ratio variable you can say urban spending per head is about 1.7 times rural (6,996 ÷ 4,122 = 1.70). Learning-assessment scale scores are closer to interval: a gap of 20 points is meaningful, a statement that one district scores "10% higher" usually is not, because the zero of the scale was set by the test designer.
Before computing a percentage change, check the variable has a true zero.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Discrete, continuous, zero and missing
Discrete variables take whole values: children ever born, rooms in a dwelling, days of MGNREGA work in a year (the Act was repealed from 1 July 2026 and replaced by the Viksit Bharat G RAM G Act 2025, which provides for 125 days, so series that span both need care). Continuous variables can take any value in a range: height, weight, consumption. Counts with many zeros, such as days of hired labour, need their own summaries; a mean of 4 days may be most households at zero and a few at 40.
Report the share with a zero next to the mean whenever zeros are common.
Zero and missing are different facts
"Earned nothing last month" is information. "Did not answer" is an absence of information. A dataset that stores both as 0 will understate average earnings and overstate the share with no income. A dataset that stores both as blank will do the reverse once the software drops blanks. Check the codebook for how each was recorded, then decide and document.
A missing value is never silently a zero.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Most indicators are built, and the build is the definition
IndicatorBuilt fromFormulaSource
StuntingChild height, age, sexShare of children under 5 with height-for-age z-score below −2 (WHO Child Growth Standards)NFHS / DHS
Labour force participation rateActivity status of each person(employed + unemployed) ÷ population of that age, × 100PLFS
Unemployment rateActivity statusunemployed ÷ labour force, × 100PLFS
MPCEHousehold consumption, household sizehousehold monthly consumption ÷ household sizeHCES
Sex ratioCounts by sexfemales per 1,000 malesCensus, NFHS
Two notes on this table. The unemployment rate divides by the labour force, while the worker population ratio divides by the whole population, so a fall in unemployment can coexist with a fall in the share of people working if many leave the labour force. And MPCE assigns each member the household average, which hides inequality inside the household, between men and women in particular.
ImpactMojoQuantitative Methods 101www.impactmojo.in
From a concept to a column
01
CONCEPT: women's work
→
02
DEFINITION: economic activity in a reference period
→
03
QUESTION: activity codes asked of each member
→
04
VARIABLE: usual principal and subsidiary status
→
05
INDICATOR: female LFPR
Each arrow is a decision that someone made, and each can be argued with. Unpaid work producing goods for the household's own use, such as fetching water or tending poultry, sits on a contested boundary between economic and domestic activity. Moving that boundary changes measured female participation by many points without any woman doing anything different.
The practitioner's check
When a report states a figure for a concept (women's agency, food security, learning), find the operational definition in its annex. If there is no annex, ask for one. The Time Use Survey, run by NSO in 2019 and again in 2024 (report released 25 February 2025), exists precisely because a labour force survey's activity codes cannot see most unpaid work.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Validity and reliability
Validity
Does the variable measure the concept it claims to? A household asset index is a reasonable proxy for long-run wealth; it is a poor measure of last month's income shock.
Reliability
Would the same person, asked again under the same conditions, give the same answer? Recall of small, frequent food purchases fades over a month, which is one reason consumption surveys use shorter recall periods for frequently bought items.
A measure can be reliable and invalid: a scale that always reads 2 kg light is consistent and wrong.
Signs of trouble in the data itself
Age heaping, where reported ages pile up on numbers ending in 0 and 5, shows that people are estimating. Heights recorded mostly to the whole centimetre show measurer shortcuts. A sudden change in an indicator between two survey phases that used different field agencies is worth checking before it is believed. These patterns are visible in a frequency table, and they are why analysts look at raw distributions before summarising.
Validity is argued; reliability can be tested with repeat measurement and back-checks.
ImpactMojoQuantitative Methods 101www.impactmojo.in
03
Section Three
Describing data
ImpactMojoQuantitative Methods 101www.impactmojo.in
Mean, median and mode
MeasureHow it is computedStrengthWeakness
MeanSum of values ÷ number of valuesUses every observation; adds up (mean × n = total)Pulled hard by extreme values
MedianMiddle value once sorted (average of the two middle values if n is even)Unmoved by a few very large or small valuesIgnores how far the tails stretch; cannot be summed across groups
ModeMost frequent value or categoryThe only centre for nominal dataCan be unstable; there may be several
Choose by the question. A state finance department estimating total spending needs the mean, because only the mean multiplies back to a total. A report on what a typical household experiences should lead with the median, because in skewed data most households sit below the mean.
When mean and median are far apart, report both. The gap itself tells the reader the distribution is lopsided, and which way it leans.
ImpactMojoQuantitative Methods 101www.impactmojo.in
One rich household moves the mean
Illustrative: monthly per capita consumption (₹) for ten households in a hamlet, sorted:
2,000 · 2,200 · 2,400 · 2,600 · 2,800 · 3,000 · 3,200 · 3,400 · 3,800 · 30,000
Mean = 55,400 ÷ 10 = ₹5,540.
Median = (2,800 + 3,000) ÷ 2 = ₹2,900.
Drop the landowner's household and the mean of the other nine is 25,400 ÷ 9 = ₹2,822, while the median of nine becomes the fifth value, ₹2,800.
What the reader would conclude
A report that says "average consumption in the hamlet is ₹5,540" is arithmetically right and describes nobody: nine of ten households consume less than ₹3,900. The median of ₹2,900 is a far better description of the typical household. One observation changed the mean by ₹2,718 and the median by ₹100.
Illustrative figures. The same arithmetic applies to land, income and loan sizes in real surveys, which are almost always skewed to the right.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Consumption in India is skewed to the right
Average MPCE (₹) by fractile class, rural India, 2023–24
MoSPI, HCES 2023–24, without imputation, via eSankhyiki API
The rural mean is ₹4,122. The class containing the median person, the 40th to 60th percentiles, averages ₹3,498 and ₹3,866, so the median lies somewhere between them and below the mean. The mean sits up at roughly the 60th percentile because the top classes pull it upward: the richest 5% average ₹10,137.
Urban India shows the same pattern more sharply: a mean of ₹6,996 against ₹5,622 and ₹6,334 in the 40th–60th percentile classes, and ₹20,310 in the top 5%.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Range, interquartile range and standard deviation
A centre without a spread is half a description. Two districts with the same mean test score can differ completely: one with every child near the mean, another with a cluster of high scorers and a cluster who cannot read.
  • Range: maximum minus minimum. Simple, and decided by the two most extreme cases.
  • Interquartile range: 75th percentile minus 25th. The spread of the middle half, resistant to outliers; pairs with the median.
  • Variance: the average squared deviation from the mean.
  • Standard deviation (SD): the square root of the variance, in the original units; pairs with the mean.
Worked: SD of five scores
Scores 4, 6, 8, 10, 12. Mean = 40 ÷ 5 = 8.
Deviations: −4, −2, 0, 2, 4.
Squared: 16, 4, 0, 4, 16; sum = 40.
Sample variance = 40 ÷ (5 − 1) = 10.
SD = √10 = 3.16.

We divide by n − 1 for a sample because deviations measured from the sample's own mean are slightly too small on average; the correction removes that bias.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Spread between groups: ratios and the Gini
6.0×
rural: top 5% average MPCE ÷ bottom 5% (10,137 ÷ 1,677)
MoSPI, HCES 2023–24
8.5×
urban: top 5% ÷ bottom 5% (20,310 ÷ 2,376)
MoSPI, HCES 2023–24
Ratios between percentile groups are easy to explain to a non-specialist and say exactly which comparison is being made.
The Gini coefficient
The Gini runs from 0 (everyone consumes the same) to 1 (one person consumes everything). MoSPI reports a consumption Gini of 0.237 for rural and 0.284 for urban India in 2023–24, against 0.266 and 0.314 in 2022–23. A single summary number for a whole distribution is convenient and lossy: very different distributions can share a Gini.
Consumption Ginis run lower than income or wealth Ginis because households smooth consumption, and survey consumption under-records the richest. Never compare a consumption Gini for India with an income Gini elsewhere.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Percentages and percentage points
Stunting among children under five in India was 38.4% in NFHS-4 (2015–16) and 35.5% in NFHS-5 (2019–21), according to the NFHS-5 India Fact Sheet. That change can be described two ways, and both are correct:
  • Absolute: 38.4 − 35.5 = a fall of 2.9 percentage points.
  • Relative: 2.9 ÷ 38.4 = 0.0755, a fall of 7.6 per cent.
Writing "stunting fell by 2.9%" is wrong on both readings. Press releases and reports make this slip often, and advocates on each side of a debate pick whichever version suits them.
Rules for writing about change
Use "percentage points" for the difference between two percentages. Use "per cent" only for a relative change, and say what it is relative to. Report the two levels as well, so readers can do either calculation. Relative changes in small numbers look dramatic: a rise from 0.5% to 1.0% is "doubling" and half a percentage point.
Levels first, then the absolute change, then the relative change if it helps.
ImpactMojoQuantitative Methods 101www.impactmojo.in
A ratio is only as good as its denominator
943
females per 1,000 males, total population
Census of India 2011, Primary Census Abstract
1,020
females per 1,000 males, de jure household population
NFHS-5 (2019–21) India Fact Sheet
The NFHS-5 figure circulated widely as evidence that India had "more women than men". The two numbers do not measure the same population.
Why they differ
The Census counts everyone, including people in institutions and the houseless. NFHS counts usual members of sampled households only. Men who live away for work in hostels, dormitories, worksites or other states are more likely to fall outside a household roster, which raises the female share among those who remain. The same fact sheet gives a sex ratio at birth of 929 for children born in the five years before the survey, which points the other way.
Before comparing two ratios, write down the numerator and denominator of each.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Cross-tabulations and which way to percentage
IllustrativeToilet usedNot usedTotal
Received IEC visit24060300
No visit350350700
Total5904101,000
Row percentages: 240 ÷ 300 = 80% of visited households use the toilet, against 350 ÷ 700 = 50% of the rest.
Column percentages: 240 ÷ 590 = 41% of users had a visit.
Percentage across the cause
The rule of thumb: compute percentages within categories of the variable you treat as the explanation (here, the visit), and compare across them. The row percentages answer "do visited households behave differently?". The column percentages answer a different question, about the make-up of users. Both are legitimate; mixing them up produces sentences that sound like findings and mean nothing.
A 30-point gap is an association. Whether the visit caused it is a separate question, taken up in Section 8.
ImpactMojoQuantitative Methods 101www.impactmojo.in
04
Section Four
Distributions
ImpactMojoQuantitative Methods 101www.impactmojo.in
What a distribution is, and the four things to read in it
Distribution
The pattern of how often each value, or range of values, occurs. A histogram draws it for a continuous variable; a bar chart of category shares draws it for a categorical one.
Plot a variable before summarising it. Summaries assume a shape, and the plot shows whether the assumption holds. The Exploratory Data Analysis 101 deck spends a whole course on this habit.
ReadQuestionExample
CentreWhere is the bulk?Median MPCE
SpreadHow wide?IQR of test scores
ShapeSymmetric or skewed? One peak or two?Landholding: long right tail
OutliersValues far from the rest, real or errors?A height of 210 cm for a two-year-old
Two peaks usually mean two populations mixed together, such as rural and urban, that should be analysed apart.
ImpactMojoQuantitative Methods 101www.impactmojo.in
The normal distribution and the 68–95–99.7 rule
Standard normal density (mean 0, SD 1)
Computed from the normal density formula
Many measured quantities, such as adult heights within one sex, and above all the averages of repeated samples, follow a bell-shaped curve. In a normal distribution about 68% of values lie within 1 SD of the mean, about 95% within 2 SD (1.96 exactly) and about 99.7% within 3 SD.
The 1.96 is the number behind every 95% confidence interval in Section 7.
Income, consumption, land and loan sizes are not normal. Applying normal-based rules to the raw values misleads.
ImpactMojoQuantitative Methods 101www.impactmojo.in
z-scores, and how stunting is measured
z = value − reference median / reference SD
A z-score says how many standard deviations a value sits from a reference. It puts measurements on a common scale, so a two-year-old and a four-year-old can be compared even though their heights differ.
Worked (illustrative reference values)
A child measures 82.0 cm. Suppose the reference median for her age and sex is 87.1 cm with an SD of 3.2 cm.
z = (82.0 − 87.1) ÷ 3.2 = −5.1 ÷ 3.2 = −1.59. She is short for her age and above the −2 stunting cut-off.
How NFHS uses it
NFHS and every DHS compute height-for-age z-scores against the WHO Child Growth Standards. A child below −2 is counted as stunted, below −3 as severely stunted. The 29.3% stunting figure for India in NFHS-6 (2023–24), like 35.5% in NFHS-5 (2019–21), is therefore a share of children whose z-score fell below a fixed line.
Because stunting is a cut-off on a continuous measure, a whole population can grow taller without the share below −2 moving much. Report the mean z-score alongside the prevalence when you can.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Skewed variables and the log scale
Consumption, earnings, loan sizes, farm output and firm size are skewed to the right: many small values and a long tail of large ones. For such variables the median and percentile ratios describe better than the mean and SD.
Analysts often take the natural logarithm. Equal ratios become equal distances: ₹2,000 to ₹4,000 and ₹10,000 to ₹20,000 are the same step on a log scale (ln 2 = 0.693 each time). The log of a right-skewed variable is often close to symmetric, which suits regression.
What this means for reading results
When a report models log consumption, its coefficients describe proportional changes. A coefficient of 0.08 on a programme dummy means consumption about 8% higher (exactly e0.08 − 1 = 8.3%). Section 8 returns to this.
Zeros break logs
The log of zero is undefined. Earnings or harvest values with many zeros cannot be logged directly; adding 1 before logging is common and changes the result depending on the units. Ask how zeros were handled.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Binary outcomes and counts
Many development indicators are binary: stunted or not, enrolled or not, institutional delivery or not. Coded 1 and 0, the mean of a binary variable is the proportion with the characteristic. Its variance is p(1 − p).
pp(1 − p)SD
0.050.04750.22
0.200.160.40
0.3550.2290.48
0.500.250.50
Variance peaks at p = 0.5. A survey designed to measure a 50% indicator needs the largest sample; this is why sample size calculations often assume p = 0.5 when nothing better is known.
Counts
Counts such as antenatal visits, children ever born or days ill last month are whole numbers starting at zero and usually right-skewed. Their mean is meaningful; their SD often exceeds what a symmetric distribution would give. Specialised models (Poisson, negative binomial) exist for them, covered in the Econometrics 101 deck.
For a binary outcome, the mean is the prevalence. Read a "mean of 0.36" as 36%.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Heaping, digit preference and impossible values
Distributions carry fingerprints of how the data were collected. Where people do not know their exact age, reported ages pile up on numbers ending in 0 and 5. Where enumerators round, heights cluster on whole centimetres. Where a questionnaire skip is mis-programmed, a whole category goes missing for one field team.
None of these show in a table of means. All of them show in a frequency table or histogram, which is why reviewers ask to see one.
Why age heaping matters downstream
Stunting depends on height for age. If a child's age is wrong by six months, the z-score is wrong, and heaping on whole years shifts children across age bands. Child age in DHS surveys comes from birth dates recorded in the birth history, which is more accurate than asking age directly; small local surveys that ask age in years produce noisier z-scores.
Impossible values
A 2-year-old at 130 cm, a household of 31, 400 days worked last year: check them against the questionnaire and the field notes, then correct, flag or exclude, and record which.
ImpactMojoQuantitative Methods 101www.impactmojo.in
The sampling distribution: why averages behave
Sampling distribution
The distribution an estimate would have if you drew the sample again and again from the same population. Its SD has a special name: the standard error.
The central limit theorem says that for reasonably large samples, the sampling distribution of a mean or a proportion is close to normal, even when the variable itself is skewed. Consumption is lopsided; the average consumption of many repeated samples of 1,000 households is bell-shaped.
Why this is the hinge of the whole deck
We only ever see one sample. The central limit theorem tells us how far a single sample estimate is likely to sit from the true value, using the normal curve and the standard error. Confidence intervals, p-values and power calculations in Sections 7 and 9 all rest on this one result.
Two different SDs: the SD of the variable describes people; the standard error describes the estimate. Reports confuse them often.
ImpactMojoQuantitative Methods 101www.impactmojo.in
05
Section Five
Sampling and sampling error
ImpactMojoQuantitative Methods 101www.impactmojo.in
How a sample can describe a country
679,238
households
NFHS-6 India Fact Sheet
716,397
women
NFHS-6 India Fact Sheet
100,977
men
NFHS-6 India Fact Sheet
NFHS-6 interviewed these households in two phases, 28 May 2023 to 26 February 2024 and 7 February 2024 to 31 December 2024, through 27 field agencies. That is well under 1% of India's households, and it supports estimates for every state. NFHS-5 (2019–21), with 636,699 households, also published a fact sheet for every district.
Why it works
Precision depends mainly on the size of the sample and very little on the share of the population sampled. A well-drawn random sample of 1,000 households measures a proportion about as precisely in a district of 2 lakh households as in a state of 2 crore. What makes this work is probability selection: every unit in the population has a known, non-zero chance of selection, so the sample can be weighted back to the population and its error quantified.
A convenience sample (whoever came to the meeting, whoever answered the phone) has no known selection probabilities, so its error cannot be computed at all.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Probability sampling designs used in South Asian surveys
DesignHow units are chosenWhy usedCost to precision
Simple randomEvery unit equally likely, drawn from a full listBenchmark for all formulasNone, but a full list of households rarely exists
SystematicEvery k-th unit from a list after a random startEasy in the field from a household listingUsually close to simple random
StratifiedPopulation split into groups (state, rural/urban); sample drawn in eachGuarantees coverage; separate estimates per stratumUsually improves precision
ClusterGroups (villages, urban blocks) selected, then units within themCuts travel and listing costUsually worsens precision
MultistageClusters first, then households within selected clustersThe standard for NFHS, PLFS and HCESCombines both effects
NFHS and other DHS surveys select villages and urban census enumeration blocks as primary sampling units, list the households in each, and then select a fixed number of households. PLFS calls its first-stage units FSUs and stratifies within them; its README refers to second-stage strata (SSS), and the weights are calculated at that level. A survey's sampling chapter tells you which units are clusters and which are strata, and you need both facts to compute correct standard errors.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Sampling error and non-sampling error
Sampling error
The chance difference between a sample estimate and the population value, arising because only some units were observed. It is random in direction, shrinks as the sample grows, and can be estimated from the data: the standard error is its measure. Every confidence interval in a report describes this error and only this error.
A report that gives a margin of error has addressed sampling error. It has said nothing yet about the other kind.
Non-sampling error
Everything else: an incomplete frame, refusals, households not found at home, questions misunderstood, interviewer effects, recall error, data entry slips, processing mistakes. It can be systematic, so it does not average out. It does not shrink with a larger sample; a larger survey may even have more of it if field supervision thins out.
In large national surveys non-sampling error often exceeds sampling error. A national interval of ±0.3 points can sit on top of a measurement problem several times that size.
ImpactMojoQuantitative Methods 101www.impactmojo.in
The standard error of a proportion
SE(p) = √[ p(1 − p) / n ]
This is the simple random sampling formula. Section 6 shows how clustering inflates it.
Nepal DHS 2022, stunting
p = 24.8% = 0.248; children measured (unweighted) n = 2,687.
p(1 − p) = 0.248 × 0.752 = 0.1865.
0.1865 ÷ 2,687 = 0.0000694.
√0.0000694 = 0.00833, so SE ≈ 0.83 percentage points.
95% interval: 24.8 ± 1.96 × 0.83 = 24.8 ± 1.63, i.e. 23.2% to 26.4% (before any design effect).
Source for p and n: The DHS Program API, indicator CN_NUTS_C_HA2, survey NP2022DHS. The published DHS report gives its own standard errors computed with the full design; they are larger than this simple version, for reasons Section 6 explains.
Three habits from this slide: square-root of p(1−p)/n; multiply by 1.96 for 95%; and treat the result as a floor until the design effect is applied.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Stunting across South Asia, with the sample behind each figure
Children under 5 stunted (%), most recent survey in the DHS Program API
The DHS Program API, indicator CN_NUTS_C_HA2; Pakistan DHS 2017–18, NFHS-5 2019–21, Bangladesh DHS 2022, Nepal DHS 2022
Surveyn (unweighted)SRS SE
Pakistan 2017–183,4920.82
India 2019–21206,4070.11
Nepal 20222,6870.83
Bangladesh 20224,2600.65
India's sample of children is about 48 times Bangladesh's, because NFHS-5 was designed for district estimates. The surveys are also four to five years apart, so the bars are not a single moment in time. India's NFHS-6 (2023–24) puts stunting at 29.3%; it is not yet in the DHS API, so its sample size is not available here and this deck's worked standard errors use NFHS-5.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Precision collapses below the state
National: India, NFHS-5
p = 0.355, n = 206,407.
SE = √(0.355 × 0.645 ÷ 206,407) = √0.00000111 = 0.00105, about 0.11 points.
95% interval ±0.21 points: 35.3% to 35.7%.
District: 300 measured children (illustrative)
Same p = 0.355, n = 300.
SE = √(0.229 ÷ 300) = √0.000763 = 0.0276, about 2.8 points.
95% interval ±5.4 points. With a design effect of 1.95, multiply by √1.95 = 1.40: ±7.6 points, roughly 28% to 43%.
District fact sheets are among the most quoted products of NFHS-5, and they rest on samples of a few hundred children per district. Two districts that differ by 6 points, or one district that "improved" by 6 points between rounds, may show nothing more than sampling noise.
Rankings of districts are especially fragile: a district can move ten places on noise alone. Ask for the interval before acting on a district rank, and prefer groupings (top third, bottom third) to exact positions.
ImpactMojoQuantitative Methods 101www.impactmojo.in
To halve the error, quadruple the sample
nSE of p = 0.35 (points)95% margin (points)
1004.77±9.3
4002.38±4.7
1,6001.19±2.3
6,4000.60±1.2
Computed as √(0.35 × 0.65 ÷ n) × 100, then × 1.96, simple random sampling. Each fourfold increase in n halves the error.
What this means for budgets
Precision is expensive at the margin. Going from ±4.7 to ±2.3 points costs 1,200 extra interviews; going from ±2.3 to ±1.2 costs 4,800 more. A commissioner should decide in advance what precision the decision needs. If the programme would act the same way whether coverage is 60% or 66%, there is no point paying for a survey that can tell them apart.
The question to ask a survey firm: what margin of error will each reported subgroup have, after the design effect?
ImpactMojoQuantitative Methods 101www.impactmojo.in
Who is outside the frame, and who did not answer
Frame gaps
Household surveys sample from lists of households. People living in institutions (hostels, barracks, prisons, care homes), the houseless, many seasonal migrants at worksites and residents of unlisted settlements are under-covered or excluded by design. For some topics, such as migrant labour, disability or homelessness, the excluded groups are exactly the ones the question is about.
Read the survey's definition of its population before generalising from it.
Non-response bias, illustrated
Illustrative: a phone survey reaches 70% of sampled households. If the 30% not reached are poorer, the estimate of poverty is biased downward, and interviewing more of the reachable households only gives a more precise estimate of the wrong number. Weighting adjustments help only to the extent that the variables used to adjust (say, district and household size) predict both response and the outcome.
Ask for response rates by stratum and by key characteristics, and for a comparison of respondents with a known benchmark such as the Census.
ImpactMojoQuantitative Methods 101www.impactmojo.in
06
Section Six
Survey weights and design effects
ImpactMojoQuantitative Methods 101www.impactmojo.in
Why survey records carry weights
Design weight
The inverse of a unit's probability of selection. A household selected with probability 1 in 500 stands for 500 households; its weight is 500.
National surveys deliberately sample some groups at higher rates than others: small states and union territories, urban areas, or districts, so that each gets a usable sample. That makes the raw sample unrepresentative of the country by construction. Weights undo it. Most survey weights then add adjustments for non-response and for agreement with known population totals.
What happens without them
If Goa, Sikkim and Delhi each receive a sample big enough for a state estimate, an unweighted all-India figure gives each of them far more influence than their population warrants. Means, proportions and totals all come out wrong, and the error has a direction: toward whatever the over-sampled groups look like.
Weighted estimates are the default for any population figure from NFHS, PLFS or HCES. An unweighted figure needs a reason.
ImpactMojoQuantitative Methods 101www.impactmojo.in
A weight changes the answer
IllustrativeUrban stratumRural stratum
Households sampled400400
Selection probability1 in 5001 in 2,000
Weight5002,000
Mean MPCE in sample₹7,000₹4,000
Unweighted mean = (400 × 7,000 + 400 × 4,000) ÷ 800 = ₹5,500.
Weighted mean = (400 × 500 × 7,000 + 400 × 2,000 × 4,000) ÷ (400 × 500 + 400 × 2,000) = 4,600,000,000 ÷ 1,000,000 = ₹4,600.
Reading the arithmetic
The sample has equal numbers of urban and rural households, but the population has four rural households for every urban one (2,000 ÷ 500 = 4). The weights restore that balance. The unweighted mean overstates consumption by ₹900, about 20%, purely because cheaper-to-reach urban households were over-represented in the sample.
The sum of the weights, 1,000,000, is the estimated number of households in the population. Totals need weights that sum to the population; NFHS weights do not, as the next slide shows.
ImpactMojoQuantitative Methods 101www.impactmojo.in
NFHS weights: divide v005 by 1,000,000
The DHS Program's guide Using Datasets for Analysis is explicit: decimal points are not stored in the weight variables, and "analysts need to divide the sampling weight they are using by 1,000,000".
Unit of analysisWeight variable
Households, household membershv005
Women, childrenv005
Menmv005
Domestic violence moduled005
generate wgt = v005/1000000
tab stunted [iweight=wgt]
These are relative weights
DHS weights are scaled so that the weighted sample size stays close to the unweighted one. For NFHS-5 stunting the DHS API reports 206,407 children measured and a weighted count of 201,276. Such weights give correct proportions and means; they cannot give a population total. To estimate how many million children are stunted, multiply the weighted proportion by a population figure from another source.
The quiet error
Forgetting to divide by 1,000,000 leaves weighted means and proportions unchanged, because the constant cancels. It multiplies every weighted count by a million, and in some software it shrinks standard errors toward zero. The figure looks fine; the significance stars do not.
ImpactMojoQuantitative Methods 101www.impactmojo.in
PLFS weights: the MLTS rules from the README
MoSPI's PLFS unit-level README defines MLTS as the "weight or multiplier (in two places of decimal) calculated at the level of Second Stage Stratum", and gives three rules:
Estimate wantedFinal weight
One sub-sample onlyMLTS ÷ 100
Both sub-samples combined, where NSS = NSCMLTS ÷ 100
Both sub-samples combined, where NSS ≠ NSCMLTS ÷ 200
Annual, from all quartersthe above, divided by the number of quarters
NSS is the number of FSUs surveyed within a quarter × visit × sector × state × stratum × sub-stratum for the sub-sample; NSC is the same count for the combined sub-samples.
Worked (illustrative MLTS)
A record carries MLTS = 2,450,000.
Sub-sample estimate: 2,450,000 ÷ 100 = 24,500.
Combined, NSS ≠ NSC: 2,450,000 ÷ 200 = 12,250.
Annual combined from four quarters: 12,250 ÷ 4 = 3,062.5.

Apply the rule record by record: within one dataset some strata have NSS = NSC and others do not.
Using MLTS/100 everywhere roughly doubles weighted totals in the strata where NSS ≠ NSC. Rates barely move, so the error survives casual checking.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Common weighting mistakes and how they show
MistakeWhat goes wrongHow to catch it
No weights at allOver-sampled states, sectors or groups dominateCompare unweighted and weighted shares of states with the published report
NFHS weight not divided by 1,000,000Counts inflated a million-fold; SEs can be wrongWeighted N should be close to unweighted N
PLFS: MLTS/100 for a combined estimate everywhereTotals inflated where NSS ≠ NSCWeighted population should match the PLFS report's estimate
PLFS: quarters pooled without dividingAnnual totals four times too largeSame check against the report
Women's weight used for menWrong population representedMatch weight to file: v005 women, mv005 men
Dropping rows to get a subgroupSEs too small (see the slide on software)Use subpopulation options
The fastest single check on any weighted analysis is to reproduce one headline number from the official report, to the decimal, before computing anything new. NFHS fact sheets and PLFS annual reports publish dozens of such numbers. If your code cannot reproduce them, it is wrong somewhere, and the fault is usually in the weights, the filters or the merge.
ImpactMojoQuantitative Methods 101www.impactmojo.in
The design effect: why clustered samples are less precise
Households in the same village share water sources, health workers, prices and weather, so they resemble each other. Twenty households from one village carry less independent information than twenty households from twenty villages.
deff ≈ 1 + (m − 1) ρ
m is the number of interviews per cluster and ρ (rho) the intra-cluster correlation, the share of variation that lies between clusters. This is Kish's approximation for equal-sized clusters.
Worked (illustrative)
m = 20 children per village, ρ = 0.05.
deff = 1 + 19 × 0.05 = 1.95.
Effective sample size = 2,000 ÷ 1.95 = 1,026.
Standard errors grow by √1.95 = 1.40.
Even a small ρ matters when clusters are large. With m = 40 and the same ρ, deff = 2.95: almost two-thirds of the interviews add no independent information about that indicator. Outcomes tied to place, such as water source or electrification, have high ρ and larger design effects.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Stratification usually helps, clustering usually hurts
Stratification
Sampling separately within states, or within rural and urban sectors, removes the chance that the sample happens to contain too many of one kind. Where the outcome differs between strata, as consumption differs between rural and urban India, stratification reduces the standard error. Its design effect is usually below 1.
Unequal weights, by contrast, add variance: a sample where a few households carry very large weights is less precise than one with similar weights. Kish's rule of thumb is a factor of 1 + CV², where CV is the coefficient of variation of the weights.
Clustering
Selecting villages and then households within them saves listing and travel costs, which is why every national household survey in South Asia does it. The price is the design effect on the previous slide.
A survey report that gives one overall design effect is giving an average. Design effects differ by indicator: they are small for individual traits like age and large for community traits like water source. Look for the sampling-errors appendix, which NFHS reports carry for selected indicators.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Tell the software about the design
Weights alone fix the point estimate. Correct standard errors also need the strata and the primary sampling units. Stata's svyset and R's survey package take all three and compute standard errors by Taylor linearisation.
* Stata, NFHS women's file
gen wgt = v005/1000000
svyset v021 [pweight=wgt], strata(v022)
svy: mean anaemic
In DHS files v021 is the PSU and v022 the sampling stratum; check the recode manual for the survey you use.
Subgroups: keep the whole sample
To estimate for Dalit women or for one state, do not delete the other rows first. Use svy, subpop() in Stata or subset() on the design object in R. Dropping rows discards the PSUs that contain no members of the subgroup, and the software then understates the standard error.
Ask any analyst: which variable did you use as PSU, which as strata, and how did you handle subgroups? Three short answers tell you whether the standard errors can be trusted.
ImpactMojoQuantitative Methods 101www.impactmojo.in
A weighting and design checklist
  • Right file for the unit of analysis (household, person, woman, child)
  • Right weight for that file, scaled as the documentation says
  • PSU and strata declared before any standard error is computed
  • Subgroups estimated with subpopulation options, never by deleting rows
  • One published headline figure reproduced exactly before new analysis
  • Report weighted estimates and unweighted sample sizes together
  • State the design effect for key indicators, or the design-based SEs
  • Flag small denominators: DHS tables put figures from 25–49 unweighted cases in parentheses and suppress those under 25
  • Do not add weights from two different surveys or rounds
  • Keep the weighting code in the deliverable
Most weighting mistakes leave percentages almost right and standard errors or totals badly wrong, so the check must cover all three.
ImpactMojoQuantitative Methods 101www.impactmojo.in
07
Section Seven
Intervals, tests and p-values
ImpactMojoQuantitative Methods 101www.impactmojo.in
What a 95% confidence interval says
estimate ± 1.96 × SE
The interval is a statement about the procedure. If the survey were repeated many times and an interval built the same way each time, about 95% of those intervals would contain the true population value. Any single interval either contains it or does not; we cannot know which.
For 90% intervals use 1.645; for 99% use 2.576. Wider intervals buy more confidence with less precision.
How to say it in a report
"Stunting is estimated at 37.6% (95% CI 35.3% to 39.9%)." A reader can see at once that 36% and 39% are both consistent with the data, and that 30% is not. That is far more useful than a lone point estimate given to one decimal place, which suggests a precision the survey does not have.
What the interval leaves out
It covers sampling error under the assumed design. It does not cover frame gaps, non-response or measurement error. A narrow interval around a biased estimate is precisely wrong.
ImpactMojoQuantitative Methods 101www.impactmojo.in
A confidence interval with the design effect
Pakistan DHS 2017–18, stunting
p = 0.376, n = 3,492 (DHS API, CN_NUTS_C_HA2).
Simple random SE = √(0.376 × 0.624 ÷ 3,492) = √(0.2346 ÷ 3,492) = √0.0000672 = 0.0082, i.e. 0.82 points.
SRS interval: 37.6 ± 1.96 × 0.82 = 37.6 ± 1.6 → 36.0% to 39.2%.
Now allow for clustering (illustrative deff = 2)
SE × √2 = 0.82 × 1.414 = 1.16 points.
Interval: 37.6 ± 1.96 × 1.16 = 37.6 ± 2.27 → 35.3% to 39.9%.
The interval widened by about 40%, from ±1.6 to ±2.3 points, once clustering was allowed for. A design effect of 2 is an assumption made for teaching; the Pakistan DHS final report publishes design-based standard errors for its key indicators, and those should be used in real work.
When a report gives a confidence interval, check whether it says "design-based", "accounting for clustering" or "svy". If it says nothing, assume it is too narrow.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Is Nepal different from Bangladesh?
Bangladesh DHS 2022: 23.6% stunted, SE 0.65 (n = 4,260). Nepal DHS 2022: 24.8%, SE 0.83 (n = 2,687). Simple random SEs computed from the DHS API denominators.
Test the difference directly
Difference = 24.8 − 23.6 = 1.2 points.
SE of the difference for two independent samples = √(0.65² + 0.83²) = √(0.42 + 0.69) = √1.11 = 1.05.
z = 1.2 ÷ 1.05 = 1.14, below 1.96. Even before any design effect, the data cannot distinguish the two countries.
The overlap shortcut and its trap
If two 95% intervals do not overlap, the difference is significant. The reverse does not hold: two intervals can overlap a little while the difference is still significant, because the SE of a difference is smaller than the sum of the two SEs. Compute the difference and its own SE whenever the comparison matters.
For two rounds of the same panel, or two groups in the same clusters, the samples are not independent and this formula needs a covariance term. Ask the analyst.
ImpactMojoQuantitative Methods 101www.impactmojo.in
How a hypothesis test reasons
01
NULL: assume no difference
→
02
STATISTIC: how far is the data from the null, in SEs?
→
03
REFERENCE: how often would that happen by chance?
→
04
P-VALUE: that probability
→
05
JUDGEMENT: with effect size and context
A test asks a narrow question: if there were in fact no difference in the population, how surprising would a sample difference this large be? The test statistic (z or t) measures the observed difference in units of its standard error. A value beyond about ±2 would occur by chance less than 5% of the time if the null were true.
Two-sided by default
Most programme questions allow an effect in either direction: a scheme might raise or lower enrolment. A two-sided test counts surprises at both ends. One-sided tests halve the p-value and should be chosen before seeing the data, with a stated reason, or not at all.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Six principles from the American Statistical Association
  • P-values can indicate how incompatible the data are with a specified statistical model.
  • P-values do not measure the probability that the studied hypothesis is true, or the probability that the data were produced by random chance alone.
  • Scientific conclusions and business or policy decisions should not be based only on whether a p-value passes a specific threshold.
  • Proper inference requires full reporting and transparency.
  • A p-value, or statistical significance, does not measure the size of an effect or the importance of a result.
  • By itself, a p-value does not provide a good measure of evidence regarding a model or hypothesis.
American Statistical Association, Statement on Statistical Significance and P-Values, released 7 March 2016 (Wasserstein and Lazar, The American Statistician 70(2)).
The p-value was never intended to be a substitute for scientific reasoning.
Ron Wasserstein, ASA executive director, press release of 7 March 2016
Principle 5 is the one most violated in development reporting: a tiny, unimportant effect in a huge sample can have p < 0.001.
ImpactMojoQuantitative Methods 101www.impactmojo.in
A two-group comparison, step by step
Illustrative: reading scores
100 children in tutored schools average 52 marks; 100 in comparison schools average 47. SD in both groups = 20 marks.
Difference = 5 marks.
SE of difference = 20 × √(1/100 + 1/100) = 20 × 0.1414 = 2.83.
t = 5 ÷ 2.83 = 1.77; two-sided p ≈ 0.08.
95% CI for the difference: 5 ± 1.96 × 2.83 = 5 ± 5.5 → −0.5 to 10.5 marks.
The conventional verdict is "not significant at 5%". The better reading is that the data are consistent with anything from no effect to a gain of about 10 marks, which is half a standard deviation. The study was too small to settle the question; it did not show the tutoring failed.
Report the difference, its interval and the p-value together. The interval carries most of the information; p = 0.08 versus p = 0.04 is a small difference in evidence, and the 0.05 line is a convention.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Did stunting in India fall between NFHS rounds?
Children under 5 stunted, India (%)
The DHS Program API (CN_NUTS_C_HA2): NFHS-3 2005–06, NFHS-4 2015–16, NFHS-5 2019–21; NFHS-6 2023–24 India Fact Sheet (provisional)
Testing NFHS-4 against NFHS-5
SE (NFHS-4) = √(0.384 × 0.616 ÷ 232,440) = 0.10 points.
SE (NFHS-5) = √(0.355 × 0.645 ÷ 206,407) = 0.11 points.
SE of the difference = √(0.10² + 0.11²) = 0.15.
z = 2.9 ÷ 0.15 ≈ 20; with a design effect of 2, about 14. The fall is far beyond sampling noise.
Statistical significance settles only chance. Whether 2.9 points over about four years is a large change for policy, and what caused it, are separate questions. The bigger fall, 9.6 points, came between NFHS-3 and NFHS-4, a gap of ten years. NFHS-6 (2023–24) reports a further fall to 29.3%, 6.2 points below NFHS-5; its sample size is not yet in the DHS API, so the test is shown on the earlier pair.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Type I and type II errors
In truth: no effectIn truth: a real effect
Test says: significantType I error (false positive). Probability α, usually 5%.Correct detection. Probability = power, usually designed at 80%.
Test says: not significantCorrect. Probability 1 − α.Type II error (false negative). Probability β = 1 − power.
The two errors trade off. Demanding stronger evidence (α = 1%) lowers false positives and raises false negatives unless the sample grows. A programme evaluation with 50% power will miss a real effect half the time, and its "no significant impact" headline will be read as failure.
Which error costs more?
Scaling up an ineffective programme wastes money (type I). Abandoning an effective one wastes the benefit (type II). The balance is a policy judgement, and the conventional 5% and 80% are defaults, open to change with a stated reason.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Multiple comparisons and the garden of forking paths
Worked: twenty outcomes, no true effects
Each test has a 5% chance of a false positive.
Chance that none of 20 independent tests is significant = 0.9520 = 0.358.
Chance of at least one false positive = 1 − 0.358 = 64%.
Expected number of false positives = 20 × 0.05 = 1.
An evaluation that tests twenty outcomes across four subgroups and reports the three significant ones has told you very little.
Defences a commissioner can require
A pre-analysis plan registered before data arrive (the AEA RCT Registry and 3ie's RIDIE are used for development studies) naming primary outcomes. Indices that combine related outcomes into one. Adjusted p-values (Bonferroni divides α by the number of tests; false discovery rate methods are less conservative). Full reporting of every outcome tested, significant or not.
Trying specifications until one crosses 0.05 is p-hacking, the practice the ASA statement names.
ImpactMojoQuantitative Methods 101www.impactmojo.in
08
Section Eight
Correlation, causation, regression
ImpactMojoQuantitative Methods 101www.impactmojo.in
The correlation coefficient
Pearson's r
A number from −1 to +1 measuring how closely two numeric variables follow a straight line. +1 is a perfect rising line, −1 a perfect falling line, 0 no linear pattern.
r² is the share of the variation in one variable that a straight line in the other accounts for. An r of 0.5 sounds large and gives r² = 0.25: three-quarters of the variation lies elsewhere. An r of 0.3 gives 0.09.
What r cannot see
r measures linear association only. A U-shaped relation, such as the one often found between age and many health outcomes, can have r near 0 while the two variables are tightly linked. r is also sensitive to outliers: one extreme district can create or destroy a correlation among thirty. And r has no units, so it says nothing about how much y changes when x changes. For that you need the regression slope.
Always look at the scatter plot before reporting r. The Bivariate Analysis 101 deck works through rank correlations and other alternatives.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Reading a scatter plot
Illustrative: female literacy and institutional births, 14 districts
Illustrative data for teaching; not survey estimates
Thirteen districts lie close to a rising line; one (62% literacy, 95% institutional births) sits well above it. Read a scatter in this order: direction (rising), form (roughly straight), strength (tight), exceptions (one outlier worth a field visit).
The outlier may be a district with a strong conditional cash transfer, a large private hospital sector, or a data error. The plot cannot say which; it says where to look.
The pattern does not show that literacy causes institutional delivery. Richer districts tend to have both.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Three reasons an association can mislead
Confounding
A third factor drives both. Households with more schooling also tend to have more land, better roads and more connections; any of these may explain their higher earnings. Comparing educated and uneducated households compares all of these at once.
Reverse causation
The outcome drives the "cause". Villages with more self-help groups may have higher incomes because better-off villages form more groups, as well as, or instead of, the reverse.
Selection
Who receives the programme is not random. If NGOs choose villages that are easier to work in, programme villages will look better whether or not the programme did anything. If the neediest are targeted, programme villages may look worse even when it works.
Each of these produces a real, reproducible, statistically significant association. Significance cannot tell you which story is true; the design of the comparison can. The Causal Inference and Impact Evaluation decks take this further.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Simpson's paradox: the aggregate can reverse the parts
IllustrativeScheme usersNon-users
Urban: institutional births90 of 100 = 90%160 of 200 = 80%
Rural: institutional births120 of 400 = 30%25 of 100 = 25%
All areas210 of 500 = 42%185 of 300 = 62%
Users do better in urban areas (90% against 80%) and in rural areas (30% against 25%), yet worse overall (42% against 62%). The scheme enrolled mostly rural women, where institutional births are low for everyone.
The classic case
Bickel, Hammel and O'Connell (1975, Science 187: 398–404) examined graduate admissions at Berkeley. Aggregated, women appeared to be admitted at a lower rate; department by department the bias largely disappeared, because women applied more often to departments that admitted few applicants of either sex.
Whenever a comparison mixes groups with very different baselines (rural and urban, states, castes), look at the comparison within groups before trusting the total.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Regression as the best-fitting line
y = a + b x + error
Ordinary least squares picks the intercept a and slope b that make the squared vertical distances from the points to the line as small as possible. The slope b answers the question r cannot: by how much does y change, on average, when x is one unit higher?
Worked (illustrative)
Monthly earnings (₹) = 4,200 + 812 × years of schooling.
Predicted for 10 years: 4,200 + 8,120 = ₹12,320.
Predicted for 12 years: 4,200 + 9,744 = ₹13,944. The difference, ₹1,624, is 2 × 812.
How to read the slope
"Among these workers, each additional year of schooling is associated with ₹812 higher monthly earnings on average." Associated: the slope describes the data, and becomes a causal effect only under the conditions on the next slides.
The intercept is the prediction at x = 0, here a worker with no schooling. It is often outside the data or meaningless (a household of size zero) and rarely worth interpreting.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Multiple regression and omitted variable bias
Adding variables lets each coefficient be read as the association holding the others constant. In an earnings regression with schooling, age, sex and urban residence, the schooling coefficient compares workers of the same age, sex and location who differ in schooling.
This is statistical control, comparing within groups the way the Simpson's paradox table did, done for many variables at once.
Omitted variable bias
If a variable that affects earnings and is correlated with schooling is left out (family wealth, ability, social networks), the schooling coefficient absorbs part of its effect. The bias has a predictable sign: a left-out factor that raises earnings and goes with more schooling pushes the schooling coefficient up.
Controls can also hurt
Controlling for a variable the treatment itself changes (occupation, in a schooling regression) removes part of the effect you want to measure. Choose controls that were fixed before the cause, and say why each is there.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Reading coefficients: units, dummies and logs
ModelCoefficientPlain-language reading
y in ₹, x in yearsb = 812Each extra year: ₹812 more per month on average
y in ₹, x a 0/1 dummy (female)b = −2,950Women earn ₹2,950 less than otherwise similar men
ln(y), x in yearsb = 0.08Each extra year: about 8% more (exactly e0.08 − 1 = 8.3%)
ln(y), ln(x)b = 0.5A 1% rise in x goes with a 0.5% rise in y (an elasticity)
y binary (0/1), linear probability modelb = 0.066 percentage points higher probability
Interaction female × urbanb = 1,100The urban gap differs for women by ₹1,100
All coefficients are illustrative. Two habits: always find the units of y and x before reading b, and for a dummy variable always find the omitted category, because the coefficient is a comparison with it. A "Scheduled Caste" dummy compared with "all others" and one compared with "Others (general)" answer different questions. For log approximations, the simple reading works well below about 0.1; above it, compute eb − 1.
ImpactMojoQuantitative Methods 101www.impactmojo.in
When can a coefficient be read as an effect?
DesignSource of comparisonKey assumptionIndian example of use
Randomised trialLottery decides who is treatedRandomisation done and keptRemedial education in Mumbai and Vadodara (Banerjee et al., NBER WP 11904)
Difference-in-differencesChange over time, treated against untreatedParallel trends without the programmePhased roll-out of a state scheme
Regression discontinuityUnits just either side of a cut-offNo sorting around the cut-offEligibility by a score or population threshold
Instrumental variablesA factor shifting treatment onlyInstrument affects outcome only through treatmentDistance to a facility (often disputed)
Regression with controlsSimilar units, measured traitsNo unmeasured confoundersThe weakest; state it plainly
The design, decided before the data, is what makes the causal claim; regression is the arithmetic that estimates it. Duflo, Glennerster and Kremer's toolkit (NBER Technical Working Paper 0333, December 2006) remains the standard practical guide to randomised designs in development. Econometrics 101 and Impact Evaluation 101 cover these designs in depth.
ImpactMojoQuantitative Methods 101www.impactmojo.in
The ecological fallacy: district patterns are not individual facts
Many quick analyses correlate district or state averages because those are what fact sheets publish. A correlation across districts between female literacy and institutional births describes districts. It does not show that literate women are more likely to deliver in hospitals, though that may also be true.
Associations at the group level are often much stronger than at the individual level, because averaging removes individual noise. They can even have the opposite sign.
Practical rules
State the unit in every sentence about a correlation ("districts with higher female literacy have more institutional births"). Use unit-level data from NFHS or PLFS when the question is about people. Weight districts by population if the question is about the population, since an unweighted district analysis gives Lakshadweep and Thane equal votes.
Multivariate Analysis 101 covers multilevel models, which handle people within districts properly.
ImpactMojoQuantitative Methods 101www.impactmojo.in
09
Section Nine
Effect sizes and power
ImpactMojoQuantitative Methods 101www.impactmojo.in
Effect size: the answer to how much
A p-value says whether an effect is distinguishable from zero. An effect size says how large it is. Policy decisions turn on the second: a nutrition programme that lowers stunting by 0.5 points with p < 0.001 and one that lowers it by 5 points with p = 0.08 call for different conversations.
  • Raw effects keep the units: percentage points, rupees, days, marks. They are the clearest for policy.
  • Standardised effects divide by a standard deviation, so outcomes on different scales can be compared or pooled.
  • Relative effects (risk ratios, odds ratios, percentage changes) compare against a baseline.
d = meanT − meanC / SD
From the worked test in Section 7
A 5-mark difference with an SD of 20 marks gives d = 5 ÷ 20 = 0.25. That study's interval ran from −0.5 to 10.5 marks, or about −0.03 to 0.53 SD.
Every result should come with its effect size in natural units first, standardised second.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Cohen's conventions, and why to use them cautiously
0.2
small
Cohen (1988)
0.5
medium
Cohen (1988)
0.8
large
Cohen (1988)
Jacob Cohen's Statistical Power Analysis for the Behavioral Sciences (2nd edition, 1988) proposed these thresholds for d, with 0.1, 0.3 and 0.5 for a correlation r. He offered them as rough guides for psychology where nothing better was known.
Context beats convention
In school-based programmes in South Asia an effect of 0.2 SD on learning is substantial, and many well-run programmes achieve less. Calling it "small" because of Cohen's labels misleads a funder. A cheap programme with d = 0.1 that reaches millions of children can matter more than an expensive one with d = 0.5 that reaches a few hundred.
Judge an effect against effects of other programmes on the same outcome, against its cost (Cost Effectiveness 101), and against what would change a decision.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Effect sizes from two Indian experiments
0.14 → 0.28 SD
remedial teaching by young women from the community, all children in treatment schools, year 1 and year 2
Banerjee, Cole, Duflo & Linden, NBER WP 11904 (2005)
0.35 → 0.47 SD
computer-assisted learning, mathematics, year 1 and year 2
Banerjee, Cole, Duflo & Linden, NBER WP 11904 (2005)
Both experiments ran in Mumbai and Vadodara. The authors report that the gains persisted for at least one year after children left the programmes.
Turning SD into marks (illustrative)
If the test's SD were 20 marks, 0.28 SD would be 0.28 × 20 = 5.6 marks, and 0.47 SD would be 9.4 marks. Translating back to the test's own units, or to "share of children who can read a paragraph", makes a result legible to a district education officer in a way that "0.28 SD" is not.
Standardised effects depend on the SD used. An effect divided by the SD of a narrow, homogeneous sample looks larger than the same effect divided by a national SD. Check which SD a paper used before comparing.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Statistical power and its four ingredients
Power
The probability that a study detects an effect of a given size, if that effect really exists. Conventionally designed at 80%, at a 5% significance level.
Power is calculated before data collection, to choose a sample size, or to state the smallest effect a fixed budget can detect. A power calculation done after a null result to explain it away is of little use.
IngredientRaise it and power
Sample size (and number of clusters)rises
True effect sizerises
Variance of the outcomefalls
Intra-cluster correlationfalls
Strictness of α (5% to 1%)falls
In clustered designs the number of clusters matters more than the number of households per cluster. Adding a 21st household to each of 40 villages buys little; adding villages buys a lot.
ImpactMojoQuantitative Methods 101www.impactmojo.in
The minimum detectable effect
MDE = (z1−α/2 + zpower) × √[1 / (P(1−P))] × √(σ²/N)
The form used in Duflo, Glennerster and Kremer (2006) for individual randomisation. P is the share treated, N the total sample, σ the outcome SD. In SD units, set σ = 1.
N = 1,000, half treated, 5% two-sided, 80% power
z values: 1.96 + 0.84 = 2.80.
√[1 ÷ (0.5 × 0.5)] = √4 = 2.
√(1 ÷ 1,000) = 0.0316.
MDE = 2.80 × 2 × 0.0316 = 0.177 SD.
Now cluster it (illustrative)
Randomise 50 villages of 20 households instead, with ρ = 0.05: deff = 1.95.
MDE × √1.95 = 0.177 × 1.40 = 0.247 SD.
The same 1,000 households can now detect only effects about 40% larger.
Read backwards: if the programme can plausibly move the outcome by 0.15 SD, neither design is big enough, and the options are a larger sample, a less noisy outcome, baseline covariates that cut variance, or not running the evaluation.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Sample size to detect a fall in stunting
n per arm = (zα/2 + zβ)² × [p1(1−p1) + p2(1−p2)] / (p1−p2)²
From 35% to 30%, 5% two-sided, 80% power
(1.960 + 0.842)² = 2.802² = 7.85.
0.35 × 0.65 + 0.30 × 0.70 = 0.2275 + 0.2100 = 0.4375.
(0.35 − 0.30)² = 0.0025.
n = 7.85 × 0.4375 ÷ 0.0025 = 1,374 children per arm.
With clustering (deff 1.95, illustrative)
1,374 × 1.95 = 2,679 children per arm, about 5,360 in all. A baseline of 35% is close to India's NFHS-5 (2019–21) figure of 35.5%; NFHS-6 (2023–24) reports 29.3% nationally.
A 5-point fall over a programme cycle is ambitious; the national fall was 2.9 points between NFHS-4 and NFHS-5 and 6.2 points between NFHS-5 and NFHS-6, each over about four years. Halving the target effect to 2.5 points roughly quadruples the sample, to about 10,900 children per arm with the same design effect. Most programme evaluations cannot afford that, which is why many choose a more sensitive outcome such as mean height-for-age z-score.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Underpowered studies mislead twice
First: false reassurance
A study with 30% power usually finds nothing even when the programme works. Its "no significant effect" becomes, in a ministry note, "the programme did not work". Absence of evidence from a small study is weak evidence of absence. Report the confidence interval, which will show that large benefits were not ruled out.
Check the reported minimum detectable effect against what the programme could plausibly achieve.
Second: exaggerated discoveries
When a low-powered study does cross p < 0.05, the estimate that got it there is usually much larger than the true effect, because only the lucky draws clear the bar. Gelman and Carlin (2014, Perspectives on Psychological Science 9(6)) call this a type M (magnitude) error, and also describe type S errors, where the sign is wrong. Scale-ups planned on such estimates disappoint.
Distrust a striking effect from a small pilot until a larger study repeats it.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Power in a commissioning conversation
Ask the evaluatorA good answer includesA warning sign
What is the primary outcome?One named indicator, measured the same way in both arms"Several outcomes" with no ranking
What effect can the design detect?An MDE in natural units and SD units, with the formula"The sample is large"
Where do the variance and ρ come from?A named prior survey, such as NFHS district data or a baselineDefaults with no source
How many clusters?A number, with the design effect it impliesOnly a household count
What if take-up is 60%?MDE adjusted for partial compliance (it grows by 1 ÷ 0.6)No mention of take-up
Is there a pre-analysis plan?Registered before endline dataAnalysis decided after seeing results
Partial take-up matters more than it looks: if only 60% of the treatment group participates, the intention-to-treat effect is diluted to 60% of the effect on participants, and the sample needed rises by a factor of 1 ÷ 0.6² = 2.8.
ImpactMojoQuantitative Methods 101www.impactmojo.in
10
Section Ten
Reading and commissioning analysis
ImpactMojoQuantitative Methods 101www.impactmojo.in
A results table, as it appears in a report
Dependent variable: monthly earnings (₹)Coef.SEtp
Years of schooling8121405.80<0.001
Female (ref: male)−2,950610−4.84<0.001
Age (years)95303.170.002
Urban (ref: rural)1,8405203.54<0.001
Constant4,2009004.67<0.001
N = 4,812; R² = 0.21
Illustrative table. Notes: OLS; standard errors clustered at village level; survey weights applied.
Read a table in this order. First the title and the dependent variable, with its units. Then N, and whether it matches the sample described in the methods. Then the notes: weights, clustering, the estimator. Only then the coefficients, each with its unit and its reference category.
The t-statistic is the coefficient divided by its standard error: 812 ÷ 140 = 5.80. Check one or two rows yourself; errors in transcription are common.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Reading the table: intervals and meaning
95% intervals from coef ± 1.96 × SE
Schooling: 812 ± 274.4 → ₹538 to ₹1,086 per year.
Female: −2,950 ± 1,195.6 → −₹4,146 to −₹1,754.
Age: 95 ± 58.8 → ₹36 to ₹154 per year.
Urban: 1,840 ± 1,019.2 → ₹821 to ₹2,859.
Every interval excludes zero, consistent with every p-value being below 0.05. The intervals are wide: the schooling return could be two-thirds or one-and-a-third of the point estimate.
What the table supports
"Among these workers, women earn about ₹2,950 a month less than men of the same schooling, age and location (95% CI ₹1,754 to ₹4,146)." This is a conditional gap. It does not identify discrimination as the cause, since occupation, hours and sector are not in the model, and it is silent on the women who are not in paid work at all, a large group given a female LFPR of 40.3% in PLFS 2024.
R² = 0.21 means the model accounts for a fifth of the variation in earnings. A low R² does not make the coefficients wrong; it means much else matters too.
ImpactMojoQuantitative Methods 101www.impactmojo.in
The notes under the table decide whether to trust it
  • Clustered standard errors: required when observations share villages, schools or clinics. Unclustered SEs from clustered data are too small.
  • Heteroskedasticity-consistent SEs (the "sandwich" option in most software): allow the error variance to differ across observations.
  • Weights: say which weight and why.
  • Fixed effects: "district fixed effects" means comparisons are made within districts only.
  • Stars: *, **, *** usually mark p below 0.10, 0.05 and 0.01; check the legend, as conventions vary.
Signs to stop and ask
N changes from column to column with no explanation (rows silently dropped for missing values). Coefficients reported without standard errors. Only starred results discussed. A control list that includes variables the treatment could have changed. A sample described in the methods that does not match the N in the table.
Ask for the code and a log file with any commissioned analysis. Reproducibility is a deliverable, and the DPDP Act's research exemption still requires that personal data be handled with safeguards.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Commissioning a district baseline: a worked example
Illustrative brief: a state nutrition mission wants a baseline for anaemia among women aged 15–49 in one district, precise enough to detect later change, and separate figures for SC and ST women.
Sizing it
Assume prevalence near 57% (NFHS-5, 2019–21, all women 15–49: 57.0%; NFHS-6 anaemia figures were not yet released as of October 2026).
Target margin ±4 points at 95%: n = 1.96² × 0.57 × 0.43 ÷ 0.04² = 3.842 × 0.2451 ÷ 0.0016 = 589 women.
Design effect 1.95: 589 × 1.95 = 1,148.6, round up to 1,149.
Allow 10% non-response: 1,149 ÷ 0.9 = 1,276.7, so 1,277 women.
The subgroup problem
If SC women are 20% of the district's women, the 1,149 completed interviews include about 230 of them, and their estimate carries a margin of about ±8.9 points after the design effect. To report SC and ST figures at ±4 points, each group needs its own 1,149 completed interviews, which means oversampling and weights.
Decide the subgroups that must be reported before the sample is drawn. Adding them afterwards cannot be fixed by analysis.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Which tool for which question
Your questionSummary or methodReport with
What share of households have X?Weighted proportion95% CI, unweighted n
What does a typical household spend?Weighted median (and mean)Percentiles or IQR
How unequal is it?Percentile ratios, GiniDefinition of the welfare measure
Do two groups differ?Difference with its own SE; t or z testDifference, CI, p-value
Did it change between rounds?Difference across rounds, checked for comparable methodsBoth levels, change in points, CI
Are X and Y related?Scatter, correlation, simple regressionSlope with units, plot
Is X related to Y, other things equal?Multiple regressionCoefficients, SEs, controls listed
Did the programme cause the change?A causal design (RCT, DiD, RD, IV)Effect size, CI, design assumptions
How big a sample do we need?Power or precision calculationMDE or margin, deff, assumptions
Most disputes about analysis are disputes about which row of this table a question belongs in. Settle that in the terms of reference.
ImpactMojoQuantitative Methods 101www.impactmojo.in
A checklist before you quote a number
About the number
  • Source named: survey, round, table, year of fieldwork
  • Definition stated: usual status or CWS, with or without imputation, age range
  • Population stated: all India, rural, women 15–49, children under 5
  • Precision known: CI or SE, or at least the unweighted n
  • Weighted, and the right weight
About the comparison
  • Same definition and method in both figures
  • Change given in percentage points and levels
  • Difference tested, with its own standard error
  • Causal language only with a causal design
  • Effect size in natural units, beside any p-value
If an item cannot be ticked, the number may still be quoted, with the gap stated.
ImpactMojoQuantitative Methods 101www.impactmojo.in
The quantitative clauses worth writing into a ToR
  • A sampling plan with frame, stages, strata, clusters and selection probabilities
  • A power or precision calculation with sourced assumptions
  • A weighting note: design weights, non-response adjustment, final scaling
  • Questionnaires and codebook, with skip logic, in every field language
  • A pre-analysis plan for any impact claim
  • Design-based standard errors for all reported estimates
  • De-identified data, code and log files as deliverables
  • A data protection plan consistent with the DPDP Act 2023 and the 2025 Rules, and with the consent given
Each clause costs little at the start and is nearly impossible to obtain after the final payment.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Working with NFHS, PLFS and HCES unit-level data
Unit-level PLFS and HCES data are published by MoSPI through its microdata portal (microdata.gov.in), with README files and layouts; NFHS data are distributed through the DHS Program after registration. Both come with documentation that answers most weighting and design questions, and both ask users to accept terms of use.
MoSPI's eSankhyiki portal also serves published aggregates through an API; the PLFS and HCES figures in this deck were retrieved that way in October 2026.
Order of work
Read the README and the questionnaire. Reproduce one published table exactly. Only then compute new estimates, using the design variables. Keep a log of every filter and recode. When an estimate surprises you, suspect your code first: the published figure has been checked by many people.
Public microdata are de-identified, but linking them to other sources can re-identify people in small districts. Do not attempt it, and do not publish cells built from a handful of respondents.
ImpactMojoQuantitative Methods 101www.impactmojo.in
11
Section Eleven
Errors in published statistics
ImpactMojoQuantitative Methods 101www.impactmojo.in
Comparing across survey rounds
Changes in questionnaire, recall period, sample design or processing can move an indicator as much as real change does. Consumption surveys are the clearest Indian case. On 15 November 2019 MoSPI announced that the results of the 2017–18 Consumer Expenditure Survey would not be released, citing "data quality issues", after a leaked draft had reported a fall in real spending (Business Today, 15 November 2019).
The next surveys, HCES 2022–23 and 2023–24, changed the design, so comparisons with 2011–12 need care and the survey reports' own notes on comparability.
Questions before comparing rounds
Was the questionnaire the same? The recall periods? The age range or the reference standard (growth references changed in 2006)? The sample frame and the season of fieldwork? Was any part of fieldwork interrupted, as NFHS-5's was split across 2019–21? If any answer is no, report the change with that caveat beside it.
A trend line drawn through non-comparable points is the most common error in published development charts.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Averages hide the groups inside them
63.7%
ST women
MoSPI, PLFS 2024, LFPR 15+, usual status
40.3%
all women
MoSPI, PLFS 2024
31.3%
women of "others" social group
MoSPI, PLFS 2024
Female labour force participation varies by more than 30 points across social groups. A national change can come from changes within groups or from changes in which groups make up the population, or the sample.
Reading group differences carefully
Higher participation among Adivasi women reflects, among other things, more agricultural self-employment and wage work out of economic necessity; lower participation among women in better-off groups reflects income effects and social norms. A single national figure averages these together. Policy reading needs the breakdown, and the breakdown needs the standard error, because ST samples are small in many states.
Before attributing a national trend to a policy, check whether it holds within groups.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Ten errors a reviewer should catch
ErrorTypical formThe fix
Points read as per cent"Stunting fell 2.9%"2.9 percentage points (7.6% relative)
Status unstatedUnemployment "3.2%" set against "4.9%"Name usual status or CWS
Unweighted national figureRaw sample shares reportedApply the survey weight
NFHS weight unscaledCounts in the hundreds of billionsDivide v005 by 1,000,000
PLFS combined weightMLTS/100 used where NSS ≠ NSCMLTS/200 there; divide by quarters for annual
SRS errors on cluster dataIntervals too narrowDeclare PSU and strata
District ranks as fact"District X ranks 3rd"Show intervals; use bands
Mismatched denominatorsCensus and NFHS sex ratios comparedName both populations
Non-comparable roundsTrend across a method changeCaveat or drop the point
Causal verbs on correlations"Literacy drives institutional births""Is associated with", or a causal design
Every example in this table appears earlier in the deck with its sources and arithmetic.
ImpactMojoQuantitative Methods 101www.impactmojo.in
New data to watch, as of October 2026
Census 2027
The reference date for the next Census is 1 March 2027, and it includes caste enumeration. It will refresh the sampling frames and population totals that every household survey uses for weights, sixteen years after the 2011 Census. Expect survey estimates of totals, and some ratios, to be revised once the new frame is in use.
Monthly labour data
MoSPI now publishes PLFS labour force indicators monthly from 2025, alongside quarterly bulletins and annual reports. Monthly estimates carry wider margins than annual ones; a 0.2-point monthly move is usually noise.
Rules that change measurement
The four Labour Codes came into force on 21 November 2025, and MGNREGA was replaced from 1 July 2026 by the Viksit Bharat G RAM G Act 2025, with 125 days of work. Where administrative series track entitlements under these laws, check the date of any break before reading a trend. The Income-tax Act 2025, in force from 1 April 2026, may affect tax-based income series in the same way.
Every time-sensitive statement in this deck is as of October 2026.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Ten ideas to keep
  • Every number is a definition, a population, a period and a margin of error
  • The level of measurement limits the arithmetic
  • Lead with the median in skewed data; report both when they differ
  • Use percentage points for differences between percentages
  • Precision depends on sample size, and it collapses in small areas
  • Weight every population estimate: v005/1,000,000 in NFHS, the MLTS rules in PLFS
  • Clustering inflates standard errors; declare PSU and strata
  • A p-value measures surprise under the null, never the size or importance of an effect
  • Association becomes causation only through design
  • Size a study by its minimum detectable effect before collecting data
All ten fit in one sentence: know what was measured, among whom, how precisely, and against what.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Where next in the 101 series
This deck covered the reading and commissioning toolkit. Each of these decks goes deeper into one part of it:
Suggested order for a practitioner: Survey Design, then Exploratory Data Analysis, then Bivariate Analysis and Impact Evaluation.
ImpactMojoQuantitative Methods 101www.impactmojo.in
Quantitative Methods 101 · Complete
Know what was measured,
among whom, how precisely
CC BY-NC-ND 4.0·Free Forever·ImpactMojo 101 Series