fullscreen
ImpactMojoData Literacy 101www.impactmojo.in
ImpactMojo 101 Series · Free Forever
Data
Literacy
101
Reading, Questioning & Using Data Responsibly — a Foundational Course for Development Practitioners in South Asia
Research-BackedSouth Asia Focus100 SlidesFree Access
ImpactMojoData Literacy 101www.impactmojo.in
What We Cover
01
What Is Data Literacy?
Slides 3–9
02
Types & Sources of Data
Slides 10–19
03
From Concept to Indicator
Slides 20–28
04
Describing Data
Slides 29–38
05
Visualising Data
Slides 39–48
06
Relationships & Correlation
Slides 49–58
07
Sampling & Surveys
Slides 59–68
08
Data Quality & Cleaning
Slides 69–77
09
Reading Data Critically
Slides 78–86
10
Data Ethics, Privacy & Equity
Slides 87–94
11
Tools & Further Reading
Slides 95–99
ImpactMojoData Literacy 101www.impactmojo.in
01
Section One
What Is Data Literacy?
ImpactMojoData Literacy 101www.impactmojo.in
Data literacy is a survival skill
Development work runs on numbers — targets, indicators, budgets, surveys, dashboards. Data literacy is the ability to read, question, interpret and communicate data, and to use it to make better decisions. It is not statistics for its own sake; it is judgement.
Data literacy
The capacity to find, read, interpret, critically evaluate and communicate with data — and to recognise when data is being misused. It sits between raw numbers and good decisions.
You do not need to be a statistician. You need to ask the right questions of any number that lands on your desk.
Not data literacyData literacy
Knowing the formula for a standard deviationKnowing when an average is the wrong summary
Making a chartNoticing that the axis starts at 90
Quoting a survey figureAsking what level the survey was designed to report at
Running a significance testKnowing what “significant” does and does not claim
The left column is technique and the right column is judgement. Technique can be looked up; judgement is what stops a competent analysis from answering the wrong question, and it is the part this course is about.
Most costly data errors in this sector are not computational. They are a state-level estimate used for a block, a mean reported for a skewed distribution, or a denominator nobody asked for — each of which is a question, not a calculation.
ImpactMojoData Literacy 101www.impactmojo.in
Decisions you already make with data
Programme decisions
  • Which blocks or wards to prioritise
  • Whether an intervention is working
  • How to set a realistic target
  • Where the budget actually goes
Daily judgement calls
  • Is this survey finding trustworthy?
  • Does this chart mislead?
  • Is the sample like our population?
  • Who is missing from this count?
Every one of these is a data-literacy question before it is a technical one.
DecisionThe data question underneath it
Which blocks to prioritiseIs the estimate reliable at block level, or borrowed from state?
Whether the intervention workedCompared with what — before, or a comparable group?
Setting a targetWhat was the trend before we arrived?
Reporting reachOut of how many — who is the denominator?
Every row is a decision someone in your organisation made this quarter, usually without treating it as a statistical question at all. That is the point: data literacy is mostly noticing that a judgement has been made.
The second row produces the most inflated claims. Before-and-after with no comparison group attributes to the programme everything that happened in that period, including the good harvest and the new road.

A sixth row worth adding yourself: when you quote a survey finding, who was sampled and who was missed?
ImpactMojoData Literacy 101www.impactmojo.in
Data → Information → Knowledge → Decision
01
DATA: raw records — 1,240 children weighed
02
INFORMATION: 18% are underweight
03
KNOWLEDGE: underweight is concentrated in 3 hamlets
04
DECISION: target supplementary feeding there
Data literacy is what moves you up the ladder without slipping — each step adds interpretation, and each step can introduce error.
StepWhat can go wrong here
Data — 1,240 children weighedScale uncalibrated; some hamlets never visited
Information — 18% underweightWhich reference standard? Denominator = weighed, or eligible?
Knowledge — concentrated in 3 hamletsAre the numbers there large enough to be more than noise?
Decision — feed those hamletsWas concentration caused by the problem, or by who got measured?
Each step up the ladder adds interpretation, and interpretation is where error enters. By the fourth step the original caveats have usually been dropped, because a decision needs a single number and a number carries no footnotes.
The last row is the trap worth naming. If three hamlets were visited more thoroughly, they will show more of everything — more malnutrition, more disease, more need. Measurement effort and measured burden are easy to confuse.
ImpactMojoData Literacy 101www.impactmojo.in
Five habits of a data-literate practitioner
  • Ask where it came from. Who collected it, when, how, and why?
  • Ask what it measures. Is the indicator really capturing the concept?
  • Ask who is missing. Whom does this number leave out?
  • Ask how sure we are. What is the uncertainty, the sample, the error?
  • Ask what decision it serves. Data without a question is noise.
HabitThe question in the room
Where did it come from?“Who collected this, when, and for what purpose?”
What does it measure?“What exactly was counted, and for whom?”
Who is missing?“Who would not have appeared in this dataset at all?”
How sure are we?“What is the sample size behind this cell?”
What decision does it serve?“What would we do differently if it were half as large?”
These five questions need no statistical training and catch most real errors. They are also socially easy to ask, which matters: the reason bad numbers survive meetings is usually that nobody wants to look ignorant.
The last one is the most underrated. A number that would not change any decision at any plausible value is not evidence; it is decoration, and collecting it cost someone their afternoon.
ImpactMojoData Literacy 101www.impactmojo.in
Numbers feel objective. They are not neutral.
Not everything that counts can be counted, and not everything that can be counted counts.
— commonly attributed to William Bruce Cameron
Every dataset embeds choices: what to measure, what category to use, whom to ask, what to ignore. Those choices carry power. A data-literate practitioner reads the choices, not just the digits.
Choice made when data is createdWhat it decides
What to measureUnmeasured problems are politically invisible
Which categories to useWho has a box to tick, and who is “other”
Whom to askHousehold head interviews report the head’s view
What counts as a caseA definition change can halve a rate overnight
What unit to reportState averages hide district collapse
Read the choices, not just the digits. Each row above was decided by a person, usually years earlier, and is invisible in the spreadsheet you receive — which is exactly why numbers feel objective.
The third row is a persistent problem in household surveys. Asking one person about everyone’s income, decisions or food intake produces data about what that person believes and is willing to say — and the gap is systematically gendered.
ImpactMojoData Literacy 101www.impactmojo.in
How this course is built
Foundations
  • Types and sources of data
  • Turning concepts into indicators
  • Describing and visualising data
Judgement
  • Correlation, sampling and surveys
  • Data quality and critical reading
  • Ethics, privacy and equity
Throughout, examples come from India and the wider region — the data you will actually meet at work.
SectionWhat you will be able to do
2–3 · Sources and indicatorsName the right national dataset; turn a concept into a defensible indicator
4–5 · Describing and visualisingChoose mean or median correctly; spot a truncated axis
6 · RelationshipsExplain confounding, ecological fallacy and Simpson’s paradox to a colleague
7–8 · Sampling and qualitySay what a sample can and cannot support; run a cleaning pipeline
9–11 · Reading and ethicsInterrogate a statistic; handle data about people
If you have one hour: Section 6 on correlation, Section 9 on reading critically, and slide 38 on denominators. Those three prevent most of the errors that reach a published report.
The examples are Indian throughout — NFHS, PLFS, Census, the National MPI. Knowing which dataset answers which question is half of applied data literacy, and it is local knowledge.
ImpactMojoData Literacy 101www.impactmojo.in
02
Section Two
Types & Sources of Data
ImpactMojoData Literacy 101www.impactmojo.in
Quantitative and qualitative
Quantitative
Numbers and counts — how many, how much, how often. Strong for measuring scale, comparing groups, tracking change.
Qualitative
Words, meanings, experiences — why, how, in what context. Strong for understanding process, mechanism and lived reality.
They are partners, not rivals. Numbers tell you that something changed; stories tell you why. The best evidence usually uses both.
QuestionWhich family answers it
How many girls dropped out this year?Quantitative
Why did they drop out?Qualitative
Did dropout fall after the intervention?Quantitative
What did families think the intervention was for?Qualitative
Is the fall large enough to be real?Quantitative
Why did it work here and not there?Qualitative
The alternating pattern is the argument. Almost every serious programme question needs both families in sequence, and an evaluation with only one of them will be able to say either that something changed or why things happen, but not both.
Beware qualitative work used as decoration. Three quotes chosen to illustrate a finding already decided is not qualitative research; it is illustration, and it carries none of the corrective power the method is meant to supply.
ImpactMojoData Literacy 101www.impactmojo.in
Primary vs secondary data
PrimarySecondary
SourceYou collect itSomeone else collected it
ExampleYour baseline survey, FGDsCensus, NFHS, district HMIS
ControlHigh — you design itLow — you take it as given
Cost / timeHighUsually low
RiskFieldwork error, biasMay not fit your question
Rule of thumb: exhaust good secondary data before collecting primary data. Much of what you need already exists — and re-collecting it wastes respondents' time.
Exhaust secondary data first — it is an ethical rule as well as a budgetary one. Collecting what already exists spends respondents’ time, which they gave once already, and survey fatigue in frequently studied districts is real and measurable.
The usual reason teams skip this step is not that the data is absent but that finding it takes a day and designing a survey feels like progress. Budget the day.
ImpactMojoData Literacy 101www.impactmojo.in
Structured, semi-structured, unstructured
Structured
Neat rows & columns — survey tables, registers, spreadsheets
Semi
Some structure — forms with open text, tagged records, JSON
Unstructured
Free text, audio, images, video, field notes
Most development M&E lives in structured data, but a growing share — call-centre logs, photos, social media, satellite imagery — is unstructured and needs different tools.
TypeCommon in our workWhat it needs
StructuredSurvey tables, MIS exports, registersSpreadsheets, standard statistics
Semi-structuredForms with open-text fields, tagged recordsCoding of the free text before analysis
UnstructuredField notes, call logs, photographs, voice notesQualitative method, or tooling most teams lack
The third row is where most organisations are sitting on unused evidence. Years of field reports, helpline recordings and monitoring photographs exist and are never analysed, because nobody owns them and no method is assigned.
A cheap first step: code one quarter of helpline or grievance text against a short list of categories. It is the closest thing most programmes have to a continuous, unprompted account of what is going wrong.
ImpactMojoData Literacy 101www.impactmojo.in
Cross-section, time series, panel
  • Cross-section: many units at one time (one NFHS round)
  • Time series: one unit over time (national TFR, 1990–2024)
  • Panel / longitudinal: same units tracked over time (a cohort re-surveyed every year)
Panel data is powerful — it can follow the same household as it changes, separating real change from differences between households.
FormCan answerCannot answer
Cross-sectionHow things stand now, across placesWhether this household improved
Time seriesWhether the aggregate movedWho moved, and who was replaced
PanelWhether the same units changedMuch, if attrition is high and non-random
The second row hides a common error. A district poverty rate can fall because poor households escaped poverty, or because poor households migrated out. The time series looks identical; the policy conclusion is opposite.
Panel attrition is the price of the panel’s power. The households hardest to find at round two — migrants, the very poor, dissolved households — are rarely lost at random, so a panel can look like improvement simply by losing the people who did worst.
ImpactMojoData Literacy 101www.impactmojo.in
India's official data ecosystem
SourceWhat it coversFrequency
Census of IndiaEvery person — population, literacy, housing, migrationDecennial (2011 latest)
NFHSHealth, nutrition, fertility, anaemia, women's status~5 years (NFHS-5: 2019–21)
NSS / PLFSConsumption, employment, unemploymentPLFS annual since 2017–18
SRSBirth & death rates, infant mortality, life expectancyAnnual
HMISFacility-level health service deliveryMonthly
SECC 2011Socio-economic & caste deprivation indicatorsOne-off (2011)
Know these by name. Most of your secondary-data needs are met by one of them — free and downloadable.
QuestionGo to
Anaemia, immunisation, child nutritionNFHS
Employment, unemployment, wagesPLFS
Population by village or wardCensus
Infant mortality, life expectancySRS
Facility-level service delivery, monthlyHMIS
Knowing which dataset answers which question is most of applied data literacy, and it is entirely learnable. All of these are free, downloadable, and documented — and each has a lowest level at which its estimates are designed to hold, which is the next slide.
ImpactMojoData Literacy 101www.impactmojo.in
NFHS, NSS and the Census do different jobs
Census = everyone
Counts every person. Best for small-area detail (a village, a ward). Expensive, so it is rare.
Surveys = a sample
NFHS & NSS interview a carefully chosen sample and infer the whole. Cheaper, frequent — but only reliable down to the level they were designed for (usually state or district).
Common error: using a state-level survey estimate to make claims about a single block. The sample was never designed to say anything that local.
ClaimSupportable?
“Anaemia among women in Bihar is X%” (NFHS)Yes — designed for state estimates
“Anaemia in this district is X%” (NFHS)Usually yes, with wider uncertainty
“Anaemia in this block is X%” (NFHS)No — the sample was never designed for it
“Population of this village” (Census)Yes — a census counts everyone
“Village anaemia” (Census)No — the Census does not measure it
This is the single most common misuse of national data in programme documents. A state estimate is quoted for a block because it is the only number available, and the caveat drops off between the proposal and the presentation.
The honest move when no local estimate exists is to say so and use the state figure as context, explicitly labelled — not to present it as a measurement of the place you work in.
ImpactMojoData Literacy 101www.impactmojo.in
How big are these datasets?
1.21 bn
people enumerated in Census 2011
Census of India 2011
~636,000
households interviewed in NFHS-5
NFHS-5, 2019–21
707
districts covered by NFHS-5
IIPS / MoHFW
These are among the largest demographic and health surveys in the world. Their size is what lets them speak reliably about districts — but not about your single panchayat.
Why size buys reliabilityWhat it does not buy
District-level estimates with usable precisionBlock or village estimates
Disaggregation by sex, residence, wealth quintileEvery cross-tabulation you might want
Comparison across roundsComparison where the question changed between rounds
The second row is where analysts overreach. A survey powered for state estimates by sex is not powered for district estimates by sex and caste and age together; the cell sizes collapse quickly, and the software will still print a number.
Always look at the unweighted count behind a percentage. “62% of women in this category” means something different when the category holds 900 respondents and when it holds nine.
ImpactMojoData Literacy 101www.impactmojo.in
The data your programme already generates
Every scheme produces administrative data as a by-product of delivery: MGNREGA muster rolls, school enrolment (UDISE+), health records (HMIS), ration transactions, immunisation registers.
Strengths
  • Continuous, cheap, already collected
  • Universal coverage of beneficiaries
  • Real-time-ish monitoring
Watch-outs
  • Records who is served, not who is missed
  • Incentives to over- or under-report
  • Gaps, duplicates, stale entries
Administrative data recordsIt cannot tell you
Who received the serviceWho was eligible and did not come
What was reported by the providerWhat actually happened, where reporting is incentivised
Transactions completedAttempts that failed at the counter
EnrolmentAttendance, and whether anything was learned
The first row is the defining limitation. Administrative data is generated by service delivery, so it is a record of the served population — and the excluded are absent from it by construction, not by oversight.
Watch the incentive attached to each field. Where a number is used to judge the person entering it, it will drift in the flattering direction — which is Goodhart’s law operating quietly inside every MIS.
ImpactMojoData Literacy 101www.impactmojo.in
Big data and digital traces
Mobile-phone records, satellite night-lights, transaction logs and remote sensing increasingly supplement official statistics — useful where surveys are slow or coverage is thin.
But digital traces over-represent the connected and under-represent the poor, women, the elderly and remote areas. Big data can deepen exclusion if read uncritically.
Digital traceWho it under-represents
Mobile-phone recordsWomen, the elderly, the poorest — who own phones less
Transaction logsCash economies, which is most of the informal sector
Satellite night-lightsActivity that is not electrified
Social mediaAlmost everyone this sector works with
Big data is fast, cheap and systematically biased toward the connected. Used to supplement a survey it can fill gaps; used to replace one it silently redefines the population as “people with digital footprints”.
The useful test: would the person you are most worried about appear in this dataset at all? If not, the dataset can tell you about many things but not about them.
ImpactMojoData Literacy 101www.impactmojo.in
03
Section Three
From Concept to Indicator
ImpactMojoData Literacy 101www.impactmojo.in
You cannot measure 'wellbeing' directly
Most things we care about — poverty, empowerment, health, learning — are concepts, not numbers. Measurement is the bridge from an abstract concept to an observable indicator.
01
CONCEPT: women's empowerment
02
DIMENSIONS: mobility, decision-making, assets
03
INDICATORS: % who can visit a health centre alone
04
DATA: survey responses
ConceptA defensible indicatorWhat it still misses
Women’s empowermentCan visit a health centre aloneWhether she wants to; whether there is one
LearningCan read a Class 2 textComprehension, reasoning, everything untested
Food securityMonths of adequate food provisioningQuality, diversity, who eats last
Access to waterImproved source within 30 minutesReliability, seasonality, queueing, who fetches it
The third column is the part that vanishes from reports. Every indicator is narrower than its concept, and once the number exists people stop saying “proxy for” and start saying the concept’s name.
The last row is a good example of an indicator changing behaviour. “Within 30 minutes” makes distance the target, so investment goes to placing sources rather than keeping them working — and a broken tap 200 metres away still counts.
ImpactMojoData Literacy 101www.impactmojo.in
Indicators, defined
Indicator
An observable, measurable marker that stands in for something we cannot observe directly. A good indicator is a faithful proxy for the concept — no more, no less.
Operationalisation
The precise rule that turns a concept into a measurement: exactly what to count, for whom, over what period, in what units.
Operationalisation must fixExample of it going wrong
Exactly what to count“Trained” — attended, completed, or passed?
For whom“Children” — under 5, under 6, or school-age?
Over what period“Last year” — calendar, financial, or recall?
In what unitsHouseholds or individuals — a 4× difference
Counted by whomSelf-report, observation, or register
Most disagreements about numbers turn out to be disagreements about definitions. Two teams reporting different coverage for the same district are usually both right, and have not compared their operationalisations.
Write the definition down before collection, in one sentence per indicator. It takes an hour, it is the document everyone asks for a year later, and it is what makes comparison across rounds possible at all.
ImpactMojoData Literacy 101www.impactmojo.in
Four kinds of variable
LevelMeaningExampleValid maths
NominalLabels, no orderDistrict, caste, religionCounts, mode
OrdinalOrdered, unequal gapsWealth quintile, Likert scaleMedian, rank
IntervalEqual gaps, no true zeroTemperature (°C), calendar yearMean, difference
RatioEqual gaps, true zeroIncome, age, children ever bornAll, ratios
Why it matters: you cannot take a meaningful average of caste categories, and a wealth quintile is a rank, not a rupee amount. The level decides which statistics are legal.
Illegal moveWhy
Average of district codesNominal — the numbers are names
“Mean wealth quintile = 3.2”Ordinal — gaps between quintiles are not equal
“Satisfaction rose 0.4 points” from a Likert scaleOrdinal treated as interval; common, and contested
“Twice as hot” in °CInterval — no true zero, so ratios are meaningless
Software will compute all four without complaint, which is why the level of measurement has to be carried in your head. The spreadsheet does not know that district codes are labels.
The third row is genuinely disputed. Treating Likert scales as interval is standard practice in much applied research and is defensible with enough categories; it is not defensible on a three-point scale, and reporting a median instead costs nothing.
ImpactMojoData Literacy 101www.impactmojo.in
What makes an indicator trustworthy?
Valid
Measures what it claims to measure
Reliable
Gives the same answer on repeat measurement
Sensitive
Moves when the real thing moves
Feasible
Can actually be collected, affordably
ImpactMojoData Literacy 101www.impactmojo.in
Accurate is not the same as consistent
Reliable, not valid
A miscalibrated weighing scale: it reads 2 kg high every time. Perfectly consistent — consistently wrong.
Valid and reliable
A calibrated scale: same answer each time, and the right answer. This is the target.
You can have reliability without validity, but never validity without reliability. Check both.
CaseReliable?Valid?
Scale reads 2 kg high, every timeYesNo
Scale drifts randomly by 3 kgNoNo
Calibrated scaleYesYes
Asking men about women’s decision-makingOften yesNo — consistently the wrong respondent
The fourth row is the version that matters in our work. A consistent measurement procedure can be consistently measuring the wrong thing, and consistency is often mistaken for quality — a stable series looks trustworthy precisely because it does not move.
The bias in row one is at least correctable. A known constant offset can be subtracted; unknown bias in who was asked cannot be, which is why validity problems are worse than reliability problems even though they look tidier.
ImpactMojoData Literacy 101www.impactmojo.in
When you measure A to learn about B
A proxy stands in for something hard to measure. Household assets proxy for wealth; night-light intensity proxies for economic activity; mid-upper-arm circumference proxies for acute malnutrition.
Every proxy leaks. Asset indices miss debt; night-lights miss the informal economy. Name the gap between your proxy and the concept — and report it.
ProxyForWhat it leaks
Asset indexWealthDebt, income flow, and urban/rural comparability
Night-lightsEconomic activityInformal and unelectrified activity
MUACAcute malnutritionChronic undernutrition; oedema
EnrolmentEducationAttendance, and whether anything was learned
Bank account openedFinancial inclusionWhether it is ever used
Every proxy leaks, and the leak is usually systematic rather than random. Asset indices understate the position of an asset-rich, cash-poor farming household in exactly the seasons when that matters most.
Name the gap in the report, once, in a sentence. “We measure enrolment, which is not attendance” costs nothing and prevents a reader from drawing the conclusion the proxy cannot support.
ImpactMojoData Literacy 101www.impactmojo.in
Bundling many indicators into one number
Indices like the Human Development Index or the Multidimensional Poverty Index (MPI) combine several indicators into a single score for easy comparison.
Upside
One memorable number; ranks and headlines; captures several dimensions at once.
Downside
Weights are value judgements; aggregation hides trade-offs; a good score can mask a terrible component.
Question to ask of any indexWhy it matters
What are the weights, and who chose them?Weights are values, presented as arithmetic
Can a bad component be masked?Aggregation hides exactly what you need to act on
Is the cut-off a natural break or a choice?Move the threshold, move the headline
Are the components correlated?If so, one thing is being counted several times
An index buys attention and spends detail. A single ranked number gets into a headline and a cabinet note; the component that would tell you what to fix does not travel with it.
Always publish the components alongside the score. The index is for attention; the components are for decisions, and a district that improved its rank by moving on the cheapest indicator has told you something the rank hides.
ImpactMojoData Literacy 101www.impactmojo.in
India's National MPI
NITI Aayog's National MPI bundles 12 indicators across health, education and standard of living — nutrition, child mortality, schooling, cooking fuel, sanitation, housing, assets and more.
12
indicators in 3 dimensions
NITI Aayog National MPI
Headcount × Intensity
MPI = share who are poor × how deeply poor they are
Notice the design choice: a household is 'MPI poor' if deprived in a weighted third or more of indicators. Change that threshold and the poverty rate changes.
Design choice in the MPIWhat changes if you change it
Deprived in a weighted third or more = poorThe headcount rises or falls, with no change on the ground
Which 12 indicators are includedWhich deprivations count as poverty at all
Equal weight to three dimensionsHealth, education and living standards treated as equally important
Headcount × intensityRewards moving people just over the line as much as deep gains
None of these is wrong; all of them are choices. The MPI is a careful, well-documented index, and reading its methodology note is the fastest way to see that a poverty rate is a definition applied to data rather than a fact discovered in it.
The fourth row has a policy consequence. Because the measure combines how many are poor with how deeply, a programme can improve the score fastest by helping those closest to the threshold — which is not the same as helping those worst off.
ImpactMojoData Literacy 101www.impactmojo.in
04
Section Four
Describing Data
ImpactMojoData Literacy 101www.impactmojo.in
Where is the centre? How spread out?
Before any fancy analysis, describe the data. Two questions answer most of it: what is typical (central tendency) and how much do values vary (dispersion).
01
CENTRE: mean, median, mode
02
SPREAD: range, IQR, standard deviation
03
SHAPE: skew, peaks, outliers
Before any modelling, reportBecause
N, and N per subgroup you will discussHalf of all overreach is a small cell
Missingness per variableA clean-looking mean may rest on 60% of cases
Centre and spreadTwo districts with one mean can be nothing alike
Shape, by eyeSkew decides whether the mean is honest
Range and extremesImpossible values are found here, not later
Descriptive statistics are not the boring preliminary; they are where nearly every data-entry error, definition mismatch and impossible value is caught. Analysts who skip them find the same problems later, after the conclusions are written.
Plot every variable once before analysing it. Five minutes of histograms catches things no summary statistic will show — a spike at zero, a wall at 99, an age distribution with peaks on round numbers.
ImpactMojoData Literacy 101www.impactmojo.in
Mean, median and mode
MeasureWhat it isBest when
MeanArithmetic averageRoughly symmetric data, no wild outliers
MedianMiddle value when sortedSkewed data — income, land, wealth
ModeMost frequent valueCategories — commonest crop, caste, response
For money — income, consumption, landholding — prefer the median. A few crorepatis drag the mean far above what a typical household actually has.
VariableUseWhy
Household income or consumptionMedianRight-skewed; a few large values dominate the mean
LandholdingMedianSame, usually more extreme
Child height or weightMeanRoughly symmetric
Days of work in a monthBothBounded; report the distribution
Commonest crop or casteModeCategorical — no other option is meaningful
For money, prefer the median unless you have a reason not to. Mean income describes a household that may not exist anywhere in the sample, and it is the statistic most often quoted precisely because it is larger.
When you must report a mean on skewed data, report the median beside it. The gap between the two is itself informative — it is a rough measure of how concentrated the top of the distribution is.
ImpactMojoData Literacy 101www.impactmojo.in
Mean vs median: the same village, two stories
Monthly income, 11 households (₹000s)
Illustrative example
Median = ₹12k (typical household). Mean = ₹31k, pulled up by one rich household. Report the mean here and you describe a village that does not exist.
ImpactMojoData Literacy 101www.impactmojo.in
Spread: range, IQR and standard deviation
  • Range: max − min. Simple, but one outlier wrecks it.
  • IQR (interquartile range): the middle 50% — robust to outliers.
  • Standard deviation: typical distance from the mean. The everyday measure of variability.
Two districts can share the same average income yet feel completely different — one equal, one polarised. The mean hides that; the spread reveals it.
MeasureRobust to outliers?Use when
RangeNo — defined by themQuick sanity check for impossible values
IQRYesSkewed data; reporting alongside a median
Standard deviationNoRoughly symmetric data
p90/p10 ratioFairlyComparing inequality across places
Two districts with identical mean income can be entirely different places. One with everyone near the average, one split between a landed few and the rest — and only the spread distinguishes them. Reporting a centre without a spread conceals that by default.
A practical habit: never report an average alone. Median with IQR, or mean with standard deviation, is the same number of words and a different amount of information.
ImpactMojoData Literacy 101www.impactmojo.in
Percentiles, quartiles and quintiles
A percentile is the value below which a given share of cases fall. The 25th percentile (Q1) has a quarter of households below it. Wealth quintiles — five 20% bands — are how NFHS and NSS routinely report inequality.
Q1–Q5
Quintiles: five equal-size groups
p50
The 50th percentile is the median
p90/p10
A common inequality ratio
TermMeansWatch out for
PercentileValue below which x% of cases fallNot the same as “x% higher”
QuartileFour equal-sized groupsGroups are equal in count, not in width
Wealth quintileFive equal-sized groups, by asset rankRelative to the survey population, not to a rupee value
p90/p10Ratio of the 90th to the 10th percentileIgnores everything above and below
The third row causes constant confusion in NFHS analysis. The poorest quintile is the poorest fifth of that survey’s population, so “the poorest quintile in Kerala” and “the poorest quintile in Bihar” are not the same living standard, and comparing them as if they were is a common error.
Quintiles also cannot show change over time in absolute terms. The bottom fifth is always exactly a fifth, so a country can grow richer without the poorest quintile share moving at all.
ImpactMojoData Literacy 101www.impactmojo.in
The shape of the data matters
Symmetric vs right-skewed distributions
Illustrative
Income, landholding and firm size are almost always right-skewed — a long tail of large values. That is exactly when the mean misleads.
ImpactMojoData Literacy 101www.impactmojo.in
The bell curve and the 68–95–99.7 rule
Many natural measurements (height, birth weight, measurement error) follow a roughly normal distribution — symmetric, bell-shaped, defined by its mean and standard deviation (SD).
68%
of values fall within 1 SD of the mean
95%
fall within 2 SD
99.7%
fall within 3 SD
But do not assume normality. Development data — income, expenditure, programme size — is usually skewed, so the rule does not apply. Always look at the shape first.
Roughly normalUsually not
Adult height; birth weightIncome, consumption, landholding
Measurement error around a true valueFirm size, farm size, city size
Sample means of large samplesCounts of rare events
Many small independent influencesAnything with a floor at zero and a long tail
Most development data lives in the right column, so the 68–95–99.7 rule — and every technique that assumes normality — needs checking rather than assuming. The check is a histogram, and it takes seconds.
The third row is why the normal curve still matters. Sample means tend toward normality even when the underlying data does not, which is what makes confidence intervals work on skewed variables — a subtlety worth knowing before someone tells you the rule never applies.
ImpactMojoData Literacy 101www.impactmojo.in
Outliers: error, or the most important case?
Could be an error
A 9-foot-tall respondent, a household of 80, an income of ₹0 — check for data-entry slips before analysing.
Could be the story
The one block with triple the dropout rate may be precisely where the programme is needed. Do not delete it — investigate it.
Never silently drop outliers. Flag them, explain them, and decide transparently.
OutlierFirst question
Household of 80 membersIs it a joint family, a hostel, or a typing error?
Income of zeroNo income, refused to answer, or not asked?
One block with triple the dropoutIs it real — and if so, that is the finding
Age 999A missing-value code someone forgot to declare
The last row is worth its own warning. Codes like 99, 999 and −1 are used for “missing” or “refused” in many official datasets, and if they are not declared to the software they enter every average silently and enormously.
Never silently drop an outlier. Flag it, investigate it, and state in the report what you did and why. An analysis whose conclusion depends on which extreme cases were removed should say so, because that is a finding about the analysis.
ImpactMojoData Literacy 101www.impactmojo.in
Always ask: out of how many?
A raw count without its denominator is almost meaningless. '500 dropouts' could be a crisis or a triumph depending on whether the base is 600 or 60,000.
Count
How many? (numerator alone)
Rate
How many out of how many? (numerator ÷ denominator)
The denominator is where most honest comparison lives — per 1,000 people, per eligible child, per year. Demand it.
Count reportedThe denominator that changes its meaning
“500 dropouts”Out of 600, or out of 60,000
“40 maternal deaths”Per 100,000 live births, not per district
“12,000 beneficiaries reached”Out of how many eligible
“Cases doubled”From 3 to 6, or from 3,000 to 6,000
“Most crime happens in this district”Most people also live in this district
The last row is the map version of this error and it is everywhere: a map coloured by raw counts is, in most cases, a population map wearing a different label. Rates fix it; counts almost never belong on a choropleth.
“Out of how many?” is the single highest-yield question in this deck. It is short, nobody can be offended by it, and it catches more bad reasoning than any technique taught later.
ImpactMojoData Literacy 101www.impactmojo.in
05
Section Five
Visualising Data
ImpactMojoData Literacy 101www.impactmojo.in
A good chart is an argument you can see
Visualisation is not decoration. A well-made chart reveals patterns the table hides — trends, gaps, outliers, relationships — and lets a busy decision-maker grasp them in seconds.
The greatest value of a picture is when it forces us to notice what we never expected to see.
— John Tukey, pioneer of exploratory data analysis
A chart is doing its job whenIt has failed when
The reader sees the comparison you intendedThey have to read the caption to know what to look at
The visual size matches the numeric differenceA 4-point rise looks like a tripling
Every axis and unit is labelled“Index” with no baseline or source
It survives being printed in black and whiteThe only distinction is red versus green
A chart is an argument, not an illustration. Choosing what to put on each axis, what to hold constant and what to leave out is analysis, and it is done before any software is opened.
Say the sentence first. If you cannot state in one line what the chart is meant to show, the chart will not show it — and that sentence usually belongs in the title, where most charts put the variable name instead.
ImpactMojoData Literacy 101www.impactmojo.in
Pick the chart for the question
You want to show…UseAvoid
Change over timeLine chartPie chart
Comparison across categoriesBar chart3-D anything
Composition / shares of a wholeStacked bar (or 1 pie, few slices)Many pies
Relationship between two variablesScatter plotDual-axis tricks
Distribution of one variableHistogram / box plotSingle average
Geographic patternChoropleth mapMap coloured by raw counts
ImpactMojoData Literacy 101www.impactmojo.in
What every honest chart needs
  • A clear title that states the takeaway, not just the topic
  • Labelled axes with units
  • An honest baseline — usually zero for bar charts
  • A source and a date
  • A note on the denominator and any exclusions
ElementWhy it is not optional
Title stating the findingMost readers read only this
Axis labels with units“Rate” per what, over what period?
Source and dateAn unsourced number is not evidence
N, or the sample basePercentages of nine cases look like percentages
A note on what is excludedMissing categories change the reading
The fourth row is the most-skipped and the most consequential. A bar chart with no base makes 2 out of 3 and 2,000 out of 3,000 look identical, and only one of them is a finding.
Charts travel further than the documents they appear in. Screenshotted into a slide, a message or a report, a chart carries only what is inside its own frame — so the source belongs on the chart, not in the paragraph beneath it.
ImpactMojoData Literacy 101www.impactmojo.in
The truncated axis
Y-axis starts at 90 (misleading)
Illustrative
Y-axis starts at 0 (honest)
Illustrative
Same data. The left chart makes a 4-point rise look like a tripling. Truncating the axis is the most common way charts lie.
ImpactMojoData Literacy 101www.impactmojo.in
More ink is not more information
Edward Tufte's principle: maximise the data-ink ratio. Every gradient, shadow, 3-D effect and clip-art icon competes with the data for attention — and usually wins.
  • No 3-D bars — they distort the very lengths you are comparing
  • No pie charts with 8 slices — the eye cannot rank angles
  • No rainbow palettes — colour should carry meaning, not noise
RemoveKeep
3-D effects, shadows, gradientsDirect labels on the lines or bars
Heavy gridlines and bordersA light gridline where reading a value matters
Decorative backgrounds and iconsAnnotation that names the point
A legend the reader must cross-referenceThe source, the units and the N
3-D is not merely ugly; it distorts. Perspective makes the front slice of a pie larger than the back one at equal value, which means a 3-D chart is misreporting even when the data is right.
Replacing a legend with direct labels is the single fastest improvement to most charts in this sector. It removes a lookup step, and the lookup step is where readers give up.
ImpactMojoData Literacy 101www.impactmojo.in
Use colour with intent — and for everyone
Good colour
  • Sequential for ordered data (light→dark)
  • One accent colour to highlight the point
  • Consistent meaning across charts
Accessibility
  • ~8% of men have colour-vision deficiency
  • Never rely on red-vs-green alone
  • Add labels, patterns or direct text
Data typePaletteCommon error
CategoriesDistinct huesA rainbow that implies an order
Sequential magnitudeOne hue, light to darkUsing a diverging scale with no midpoint
Diverging around a midpointTwo hues from a neutral centrePlacing the neutral point arbitrarily
AnyRed/green as the only distinction
Roughly one in twelve men has some colour-vision deficiency, most commonly red–green, so a chart whose meaning rests on that pair is unreadable for a predictable share of any audience. Vary lightness as well as hue and the problem disappears.
Print it in greyscale as a test. If the categories are still distinguishable, the chart will survive photocopying, projection and a low-quality screen — which is where most of your charts will actually be read.
ImpactMojoData Literacy 101www.impactmojo.in
Repeat a small chart to compare many
Instead of cramming ten states into one tangled line chart, draw ten tiny identical charts side by side — small multiples. The eye compares shapes effortlessly when scale and layout are shared.
Rule: when a single chart gets crowded, split it into a grid of small, identical ones rather than adding more colours.
Instead ofUse small multiples
Twelve lines on one chartTwelve small charts, one per district
A dual-axis chartTwo stacked panels sharing an x-axis
An animated chartA row of frames the reader can compare at once
Many piesA single sorted bar chart
Small multiples work because comparison is easiest side by side, not overlaid. The rule that makes them work is a shared scale on every panel — without it, the layout implies a comparison the axes do not support.
Dual-axis charts deserve their reputation. Two series on different scales can be made to cross wherever the designer chooses, so the apparent relationship is a formatting decision. If you must, label both axes prominently and expect scepticism.
ImpactMojoData Literacy 101www.impactmojo.in
Sometimes a table beats a chart
  • Use a table for exact values people will look up or quote
  • Use a chart for pattern, trend and comparison
  • Right-align numbers, fix decimal places, and add row/column totals
  • A heat-shaded table can do both — precise and patterned
Use a table whenUse a chart when
Exact values matterThe pattern matters more than the values
There are few rows and several columnsThere are many observations
Readers will look up their own districtReaders need one comparison
The numbers will be quotedThe shape will be remembered
A well-made table is often the better choice and is treated as a failure of effort. Six numbers in a table are read accurately; six numbers in a bar chart are read approximately, and the chart takes longer to make.
Sort the table by the column that matters, not alphabetically by district. Alphabetical order is the default and hides every pattern the table contains.
ImpactMojoData Literacy 101www.impactmojo.in
Before you publish a chart, ask…
  • Does the title state the finding?
  • Is the baseline honest (zero where it should be)?
  • Are axes, units and the denominator labelled?
  • Could a colour-blind reader read it?
  • Is the source and date on the chart?
  • Have I removed everything that is not data?
CheckFail state
Does the axis start at zero, or is the truncation justified and labelled?The most common way charts lie
Is the title the finding, not the variable name?“Figure 3: Enrolment”
Are source, date and N present?Unsourced, undated, unbased
Is it readable in greyscale?Red/green only
Would someone who disagrees find it fair?The test that catches the rest
The last check is the one worth internalising. Most misleading charts are made by people arguing for something true, who chose the framing that made their case look strongest and never asked how it would read to a sceptic.
Truncating an axis is sometimes right — for a series that varies within a narrow band, zero-anchoring hides everything. The rule is not “never truncate”; it is “never truncate silently”.
ImpactMojoData Literacy 101www.impactmojo.in
06
Section Six
Relationships & Correlation
ImpactMojoData Literacy 101www.impactmojo.in
Do two things move together?
Correlation
A measure of how strongly two variables move together. Positive: both rise together. Negative: one rises as the other falls. Zero: no linear relationship.
The correlation coefficient r runs from −1 (perfect negative) through 0 (none) to +1 (perfect positive).
rMeansDoes not mean
+0.9Strong positive linear associationThat one causes the other
0No linear relationshipNo relationship — a U-shape gives r near 0
−0.5Moderate negative associationThat the effect is moderate in size
Any valueSomething about the sampleAnything about an individual case
The second row is the one that catches people. A correlation near zero is often reported as “no relationship”, when a strong non-linear pattern — benefit rising then falling — produces exactly that number. Plot it before concluding.
Strength is not effect size. A tight correlation can describe a tiny effect, and a noisy one can describe a large effect. r tells you how closely the points follow a line, not how steep the line is.
ImpactMojoData Literacy 101www.impactmojo.in
Female literacy and fertility across Indian states
Female literacy (%) vs total fertility rate, major states
Illustrative, patterned on Census 2011 & NFHS-5
A clear negative correlation: states with higher female literacy tend to have lower fertility. But does literacy cause lower fertility? Hold that thought.
ImpactMojoData Literacy 101www.impactmojo.in
Correlation is not causation
Two variables can move together for several reasons, only one of which is 'A causes B'.
  • Reverse causation: B might cause A
  • Confounding: a third factor C drives both
  • Selection: the sample was chosen in a way that creates the link
  • Chance: with enough variables, some correlate by luck
ExplanationExample
A causes BThe one usually assumed
B causes AProgramme presence correlates with need — because need attracted the programme
C causes bothIce-cream sales and drowning; the confounder is summer
SelectionVillages that volunteered differ from those that did not
ChanceTest enough pairs and some will correlate
The second row is endemic in programme evaluation. Interventions are placed where the problem is worst, so a naive comparison shows programme areas doing worse than non-programme areas — and the programme looks harmful.
The practical discipline: before accepting a causal reading, say the other four aloud and explain why each is unlikely here. If you cannot, the claim is an association, and should be written as one.
ImpactMojoData Literacy 101www.impactmojo.in
The lurking third variable
Ice-cream sales correlate with drowning deaths. Ice cream does not cause drowning — summer heat drives both. The confounder is the real story.
01
Hot weather (confounder C)
02
drives ice-cream sales (A)
03
AND drives swimming & drowning (B)
04
so A and B correlate — with no causal link
Apparent relationshipPlausible confounder
Toilet ownership and child heightHousehold wealth, which drives both
Attending a training and higher yieldsFarmers who attend differ in land, literacy and motivation
Mobile ownership and women’s mobilityUrban residence
Private schooling and test scoresParental education and income
Confounding is the default state of observational data, not an occasional hazard. Anything associated with wealth is associated with everything else associated with wealth, which in this sector is nearly everything.
Controlling for a confounder helps only if you measured it. “We controlled for household characteristics” addresses the ones in the dataset; motivation, information and connections are usually not among them.
ImpactMojoData Literacy 101www.impactmojo.in
Patterns appear in pure noise
Test enough pairs of unrelated variables and some will correlate strongly by sheer chance. A tight correlation is a clue, never a proof.
Before believing a correlation, ask: is there a plausible mechanism? Could a confounder explain it? Does it survive in other data?
How spurious findings ariseGuard
Many variables compared pairwiseState the number of comparisons made
Both series trend over timeCompare changes, not levels
Small samples produce large r by chanceReport N with every correlation
Only the interesting result is written upPre-specify what you will test
Two variables that both rise over time will correlate regardless of any connection, which is why so many striking time-series correlations are meaningless. Differencing — looking at year-on-year change rather than level — removes most of them.
The fourth row is Section 9’s forking paths, arriving early. The correlation that made it into the report is rarely the only one computed, and the ones that did not are what a reader needs in order to judge it.
ImpactMojoData Literacy 101www.impactmojo.in
Never trust the number without the picture
Anscombe's quartet is four datasets with identical means, variances and correlation (r = 0.82) — yet utterly different shapes: one linear, one curved, one a single outlier driving everything.
The lesson, proven in 1973 and true today: always plot your data. Summary statistics alone can hide the truth.
Anscombe’s four datasets shareAnd look like
The same mean in x and yA clean linear relationship
The same varianceA curve
The same correlationA line plus one extreme outlier
The same regression lineA vertical cluster plus one distant point
Four datasets, identical summary statistics, four completely different pictures. It is the most economical demonstration in statistics that summaries can be identical while the data is not, and it takes one glance to absorb.
The instruction is simply: plot it. Before quoting any correlation or fitting any line, look at the scatter. It costs seconds and it is the only reliable protection against the four cases above.
ImpactMojoData Literacy 101www.impactmojo.in
Group patterns ≠ individual truths
Ecological fallacy
Wrongly inferring something about individuals from a pattern seen only at the group level.
A district with higher average income may have higher literacy on average — that does not mean the richer individuals within it are the literate ones. What holds for districts need not hold for people.
True at group levelNot therefore true of individuals
States with higher female literacy have lower fertilityThat any literate woman has fewer children
Districts with more migrants report more remittancesThat migrant households receive more
Richer states have higher obesityThat richer individuals are heavier
The ecological fallacy is inferring about individuals from group averages, and it is common because group data is what is published. Aggregate patterns can even reverse at individual level, which is the next slide.
The reverse error exists too. Assuming a group behaves like its typical member — the individualistic fallacy — underlies a good deal of policy built from case studies.
ImpactMojoData Literacy 101www.impactmojo.in
A trend can reverse when you split the data
Simpson's paradox: a relationship visible in the whole dataset can flip when you break it into subgroups. A scheme can look worse overall yet be better in every region — if the regions differ in size and baseline.
Always disaggregate before concluding. The aggregate can point the opposite way to the truth.
Where the paradox shows upThe hidden variable
Hospital A has worse survival than B overallA takes the severe cases
A scheme looks less effective overall than by districtUptake differs by district size
Wages fall overall while rising in every sectorEmployment shifted toward lower-paying sectors
A trend can genuinely reverse when the data is split, and neither the aggregate nor the split is wrong. Which one answers your question depends on what you are deciding — and that has to be argued, not assumed.
The practical habit: always look at an aggregate result broken down by the obvious grouping variable at least once. If the direction changes, you have found something worth understanding rather than reporting.
ImpactMojoData Literacy 101www.impactmojo.in
Extreme results drift back to average
When you pick the worst-performing districts and re-measure them later, they usually look better — even with no intervention. Extreme values contain extra luck that does not repeat. This is regression to the mean.
It is a notorious trap in evaluation: target the bottom 10% of schools, see them improve, and credit your programme — when much of the gain would have happened anyway. A comparison group is the cure.
SituationWhy improvement appears without any cause
Targeting the worst-performing blocksExtreme values are partly luck; luck does not repeat
Selecting schools with the lowest scoresSome were having a bad year, and would recover anyway
Intervening after a spike in casesSpikes subside; the intervention gets the credit
This is one of the most consequential statistical facts for programme evaluation, because targeting the worst cases is exactly what good programmes do — and it guarantees that a before-and-after comparison will overstate the effect.
The fix is a comparison group selected the same way: equally extreme units that did not receive the intervention. They will improve too, and the difference between the two improvements is the part the programme can claim.
ImpactMojoData Literacy 101www.impactmojo.in
07
Section Seven
Sampling & Surveys
ImpactMojoData Literacy 101www.impactmojo.in
You rarely need to ask everyone
A well-chosen sample of a few thousand can describe a population of millions — the principle behind NFHS, NSS and every opinion poll. The magic is not size; it is representativeness.
Representative sample
A sample whose composition mirrors the population on the characteristics that matter, so findings can be generalised back to the whole.
Sampling buysAt the cost of
Speed — weeks instead of yearsUncertainty, which must be reported
Depth — longer interviews, better trainingSmall-area estimates
Repeatability — annual roundsComparability if the design changes
Lower cost per respondentNothing, if the sample is well drawn
A well-drawn sample often beats a census in quality, not merely in cost. Fewer interviews means better-trained enumerators, more supervision and shorter fieldwork — and a census carried out badly can be less accurate than a sample carried out well.
The second row is the real trade-off in our work. Programme staff want block-level numbers; sample surveys are designed for state or district. Wanting the number does not create the precision.
ImpactMojoData Literacy 101www.impactmojo.in
Define the population first
01
TARGET POPULATION: whom you want to learn about
02
SAMPLING FRAME: the list you can actually draw from
03
SAMPLE: who you end up measuring
04
RESPONDENTS: who actually answers
Each gap — frame missing people, non-response, refusals — is a place bias creeps in. The frame is often the weakest link: a list of phone owners is not a list of citizens.
TermMeaningWhere it goes wrong
Target populationEveryone you want to describeStated vaguely: “the community”
Sampling frameThe list you actually sample fromExcludes the homeless, migrants, new arrivals
SampleThose selectedConfused with those who responded
RespondentsThose who actually answeredNon-response treated as random
The gap between rows one and two is where most exclusion enters. A frame built from a voter list, a ration list or a village register omits precisely the people least attached to institutions — and no amount of random sampling from that frame recovers them.
Report the frame, always. “A random sample of households” is not a description; “a random sample from the 2021 gram panchayat household list” is, and it tells the reader who could never have been selected.
ImpactMojoData Literacy 101www.impactmojo.in
Let chance choose — it removes bias
MethodHowUse when
Simple randomEvery unit equal chanceYou have a full list
SystematicEvery k-th unit from a listOrdered list, no hidden cycle
StratifiedSplit into groups, sample eachYou must represent subgroups
ClusterSample whole groups (villages)People are geographically spread
MultistageClusters, then units withinLarge national surveys (NFHS)
Only probability sampling lets you calculate a margin of error and generalise honestly.
DesignWhen it fitsWatch
Simple randomA good complete frame existsRarely practical over a large area
SystematicAn ordered list, easy in the fieldHidden periodicity in the list order
StratifiedYou need estimates for subgroupsStrata must be defined before sampling
ClusterTravel cost dominatesLarger uncertainty for the same N
Multi-stageNational surveys — NFHS, NSSWeights become essential
Cluster sampling is cheap and statistically expensive. People within a village resemble each other, so 30 households in one village carry less information than 30 spread across 30 villages — the design effect, and it is why cluster surveys need larger samples.
Stratify by what you will report on. If the report must speak about Scheduled Tribe households separately, that has to be a stratum at design time; it cannot be recovered by filtering afterwards.
ImpactMojoData Literacy 101www.impactmojo.in
Convenient, but you cannot generalise
  • Convenience: whoever is easy to reach — the people at the camp
  • Purposive: hand-picked for a reason — key informants
  • Snowball: respondents refer others — hidden populations
  • Quota: fill fixed counts per group, but non-randomly
These are legitimate for qualitative depth and hard-to-reach groups — but you cannot attach a margin of error or claim population-level numbers from them.
MethodLegitimate useIllegitimate use
ConveniencePiloting a questionnaireAny prevalence estimate
PurposiveSelecting information-rich cases for qualitative work“Representative” claims
SnowballReaching hidden or stigmatised populationsEstimating population size
QuotaFast market-style pollingConfidence intervals
Non-probability sampling is not bad sampling; it is different sampling. It answers “what exists” and “how does this work”, and cannot answer “how many” — and the failure is almost always in the claim, not the method.
Never attach a margin of error to a non-probability sample. The arithmetic will produce one and it means nothing, because the formula assumes a known probability of selection that does not exist here.
ImpactMojoData Literacy 101www.impactmojo.in
How many do I actually need?
Margin of error vs sample size (95% confidence, p=0.5)
Standard sampling theory
Note the curve flattens: ~1,067 gives ±3%, but halving the error to ±1.5% needs ~4,000. Precision gets expensive fast.
ImpactMojoData Literacy 101www.impactmojo.in
It's the sample size, not the fraction
A counter-intuitive truth: for a large population, accuracy depends on the absolute sample size, not the share of the population sampled. 1,500 people describe a state and a country about equally well.
This is why a national survey of ~600,000 households can speak about all of India — and why your block of 2,000 households still needs a few hundred interviews, not twenty.
PopulationSample for ±3%Fraction sampled
10,000~1,00010%
1,000,000~1,0700.1%
1,400,000,000~1,0700.00008%
Precision depends on the sample size, not on the share of the population. This is the most counter-intuitive fact in survey work, and it is why a national poll of 1,200 people can be as precise as a district poll of 1,200 — and why “they only asked 1,000 people” is not, by itself, an objection.
The catch is subgroups. A sample of 1,000 that supports a national estimate supports almost nothing about a district within it, because the district cell holds a handful of cases. Sample size is needed per estimate, not per survey.
ImpactMojoData Literacy 101www.impactmojo.in
The errors that size cannot fix
  • Selection bias: the frame or method systematically misses people
  • Non-response bias: those who refuse differ from those who answer
  • Survivorship bias: you only see who remained (drop-outs vanish)
  • Social-desirability bias: people answer how they think they should
A bigger biased sample is just a more confident wrong answer. Size fixes noise, never bias.
BiasHow it enters
CoverageThe frame omits a group entirely
Non-responseThose absent or refusing differ systematically
Selection by the enumeratorThe nearest, easiest, most welcoming households
Social desirabilityAnswers shaped by what is acceptable to say
RecallDistant events remembered selectively
Bias does not shrink as the sample grows. A larger biased sample gives a more precise estimate of the wrong number, which is why sample size is the wrong thing to argue about when the design is flawed.
The third row is under-policed in our sector. Enumerators sent to a village with a target and no fixed selection rule will complete the easy interviews, and the households they skip are the ones furthest out, most transient, and least at home during the day.
ImpactMojoData Literacy 101www.impactmojo.in
Why survey results come 'weighted'
When some groups are deliberately over-sampled (to study them reliably) or respond less, surveys apply weights so each respondent represents the right number of real people.
Practical warning: using NFHS or PLFS unit data without the survey weights gives wrong totals. Always weight when the documentation says to.
Weights correct forConsequence of ignoring them
Unequal probability of selectionOver-sampled groups dominate every estimate
Deliberate over-sampling of small groupsNational figures skewed toward the boosted stratum
Non-response, adjusted post hocThe respondents stand in for everyone
Unweighted analysis of a weighted survey is a common and silent error. The software runs, the numbers look plausible, and they describe the sample rather than the population — which is not what anyone intended to report.
Two practical rules: use the weight variable the survey documentation specifies, and report unweighted counts alongside weighted percentages so a reader can see how many real cases sit behind each figure.
ImpactMojoData Literacy 101www.impactmojo.in
A survey is only as good as its questions
  • Avoid leading questions ('Don't you agree that…?')
  • Avoid double-barrelled ones ('clean and safe?' — which one?)
  • Use language and units respondents actually use locally
  • Pilot every instrument before the real round — always
Bad questionWhy
“Do you agree that the scheme has improved your life?”Leading; invites agreement
“How satisfied are you with health and education services?”Double-barrelled — two questions, one answer
“How much did you spend on food last year?”Recall period far too long
“Do you own assets?”Undefined; every respondent answers a different question
“You do send your daughter to school, don’t you?”Social desirability, at maximum
Question wording is measurement. A badly worded item produces clean, analysable, precisely wrong data — and no statistical technique downstream can recover what the wording destroyed.
Pilot with twenty people and listen to the confusion. Where respondents ask a clarifying question, the item is ambiguous; where they answer instantly with the socially approved reply, the item is leading. Both are visible in an afternoon.
ImpactMojoData Literacy 101www.impactmojo.in
08
Section Eight
Data Quality & Cleaning
ImpactMojoData Literacy 101www.impactmojo.in
Most data work is cleaning
Analysts often spend the majority of a project just preparing data — finding errors, reconciling formats, handling gaps. Glamorous analysis sits on a large, unglamorous foundation of cleaning.
Garbage in, garbage out. No model rescues bad data.
— computing proverb, truer than ever
Where the time actually goesRoughly
Finding, understanding and merging dataThe largest single share
Cleaning, recoding, handling missingnessThe next largest
AnalysisMuch less than anyone expects
Writing up and making chartsMore than budgeted
Cleaning is the work, not the preliminary, and budgeting a fortnight of analysis with two days of cleaning is how deadlines are missed. The proportions above are folklore in the field rather than a measured statistic, and every experienced analyst recognises them.
It is also where the findings are. The anomalies you resolve while cleaning — duplicate households, impossible dates, a block with no records for March — are frequently more informative than the analysis they were preparing for.
ImpactMojoData Literacy 101www.impactmojo.in
What 'dirty' data looks like
ProblemExampleRisk
Missing valuesBlank income fieldBiased averages if not random
DuplicatesSame beneficiary twiceInflated counts
Inconsistent codes'F' / 'Female' / '2'Broken grouping
Outliers / impossibleAge = 200, −5 childrenDistorted statistics
Format driftDD/MM vs MM/DD datesSilent miscalculation
Typos in keysMisspelt village nameFailed merges
SymptomTypical cause
“Bihar”, “bihar”, “BIHAR”, “Bihar ”Free-text entry with no controlled list
Ages of 0 and 999 in the same columnUndeclared missing codes
Dates in three formatsExcel, and different data-entry operators
Duplicate household IDsRe-visits recorded as new records
Numbers stored as textLeading zeros, commas, or a stray space
Almost all of this is preventable at collection. Dropdowns instead of free text, range checks on numeric fields, and a required ID format eliminate four of the five rows before any data reaches an analyst.
The fourth row is the dangerous one, because duplicates inflate counts silently and are invisible in any summary statistic. Check unique-ID counts against record counts as the first step of every merge.
ImpactMojoData Literacy 101www.impactmojo.in
Why values are missing matters more than how many
  • Missing at random: gaps unrelated to the value — least harmful
  • Missing not at random: the richest refuse to state income — this biases results
  • Dropping rows with gaps can quietly delete the very people you care about
Before deleting or filling missing values, ask why they are missing. The pattern of absence is itself data.
Why it is missingWhat you can do
Missing at random — a skipped pageAnalysis is largely unaffected
Related to something you measuredCan be adjusted for, with care
Related to the answer itselfThe dangerous case: no fix from the data alone
The third row is common and usually unremarked. Households that refuse to state income are not a random subset of households; people absent at interview are absent for reasons connected to work and migration. Dropping them silently assumes they resemble everyone else.
Report missingness per variable, every time. A table of “% missing” beside each indicator is two lines of code, and it lets a reader judge which of your findings rest on most of the sample and which on half of it.
ImpactMojoData Literacy 101www.impactmojo.in
A repeatable cleaning workflow
01
INSPECT: look at every column's range & uniques
02
VALIDATE: rules (age 0–120, % in 0–100)
03
FIX: standardise codes, dates, units
04
DOCUMENT: log every change
05
FREEZE: keep raw data untouched
Golden rule: never edit the raw file. Clean in a script or a copy so every change is reversible and visible.
StepRule
Keep the raw file untouchedRead-only; never edited, ever
Every change in a scriptNot by hand in a spreadsheet
One script, run start to finishReproduces the clean file from the raw one
Log what was dropped and whyCounts before and after each filter
Output a documented clean datasetWith a codebook that matches it
The second row is the one that separates work you can defend from work you cannot. Manual edits in a spreadsheet are invisible, unrepeatable and impossible to review — and they are how most cleaning in this sector is still done.
Row four saves careers. “We started with 4,200 records and analysed 3,180” invites the obvious question, and having the answer — with the counts at each step — is the difference between rigour and an awkward silence.
ImpactMojoData Literacy 101www.impactmojo.in
If you can't redo it, you can't trust it
Fragile
Manual edits in a spreadsheet, no record of what changed. Next month, nobody can reproduce the number — including you.
Robust
A documented script from raw to result. Re-run it any time, audit every step, hand it to a colleague.
TestIf you fail it
Could a colleague rerun this from the raw data?The result cannot be checked
Could you, in six months?The result cannot be updated
Does the number in the report match the script’s output?Something was changed by hand
Is the exact data version recorded?Re-running gives a different answer
If you cannot redo it, you cannot trust it — and neither can anyone else. Reproducibility is not an academic virtue here; it is what allows a finding to survive the analyst leaving, which in this sector happens frequently.
The minimum viable version needs no new tools: a folder with the raw data, one script, one output, and a short README naming the source file and the date. That covers most of the benefit.
ImpactMojoData Literacy 101www.impactmojo.in
Build checks in, don't hope
  • Range checks: can this value exist at all?
  • Logic checks: a 6-year-old cannot be married with children
  • Cross-checks: do parts sum to the reported total?
  • Sense checks: does the headline number pass the smell test?
CheckCatches
Range checks on every numeric fieldAges of 200; incomes of −5
Controlled lists for categoriesFour spellings of one district
Skip logic enforced in the formPregnancy answers from men
Unique-ID constraintDuplicates, at entry rather than at analysis
Daily review of the first week’s dataAn enumerator misunderstanding a question
The last row is worth more than all the others combined. Reviewing the first days of fieldwork catches systematic misunderstandings while they can still be corrected; the same error found at analysis is unfixable and costs the whole variable.
Build checks in rather than hoping. Every constraint above is a setting in ordinary survey software, costs minutes to configure, and prevents an error class permanently rather than case by case.
ImpactMojoData Literacy 101www.impactmojo.in
Document so future-you can understand
Metadata is data about your data: what each variable means, its units, allowed values, how and when it was collected, and what you changed.
A dataset without a data dictionary is a puzzle with no key. The six months it takes to forget your own coding is shorter than you think.
DocumentBecause in six months
What each variable means, and its units“inc2” will mean nothing to anyone
Category codes and missing codes99 will quietly enter an average
How the data was collected, and whenComparability depends on it
Known limitations and gapsThe caveats live only in your head
Who to askThat person will have left
The audience for a codebook is future-you, and future-you will not remember. In a sector with high staff turnover, an undocumented dataset is effectively lost when the person who built it moves on.
Write it while collecting, not after. Documentation written at the end is reconstructed from memory and is wrong in exactly the places that matter — the odd recodes and the exceptions.
ImpactMojoData Literacy 101www.impactmojo.in
Keep the trail
  • Keep the raw extract read-only and dated
  • Name files with versions and dates, not 'final_FINAL_v3'
  • Record the source, download date and any filters applied
  • Save the cleaning script alongside the data
PracticePrevents
Dated, immutable raw files“Which version produced the report?”
Clear naming — no “final_v3_FINAL”Analysing the wrong file
A change log, even a text fileUnexplained differences between runs
Version control for scriptsLosing a working analysis to an edit
The second row is a joke that costs real money. Every organisation has a folder of near-identical files, and the recurring cost is not storage but the hour spent establishing which one the published figure came from — sometimes unsuccessfully.
A dated folder per extract is enough to start. Full version control is better and takes learning; naming discipline and a change log capture most of the benefit on day one.
ImpactMojoData Literacy 101www.impactmojo.in
09
Section Nine
Reading Data Critically
ImpactMojoData Literacy 101www.impactmojo.in
Every estimate has a range
A survey figure of 42% is shorthand for 'about 42%, give or take'. The confidence interval — say 39–45% — is the honest version. A point estimate without its range overstates certainty.
If two groups' confidence intervals overlap heavily, a difference between them may be noise, not signal. Look for the range, not just the dot.
Reported asShould be read as
“Anaemia is 52.1%”“Between about 50 and 54, on this survey’s design”
“Up from 51.8% last round”Possibly unchanged — the intervals overlap
“District A is worse than B”Check whether the intervals separate
A single decimal placeUsually false precision
Every survey estimate is a range presented as a point. The decimal place makes it look exact, and the confidence interval — which is published in the report almost nobody opens — is what actually bounds the claim.
The practical rule for rankings: before saying one district is worse than another, check whether their intervals overlap. Most district league tables in this sector are ranking noise for the middle two-thirds of the list.
ImpactMojoData Literacy 101www.impactmojo.in
'Significant' has a narrow technical meaning
Statistical significance
A result unlikely to have arisen by chance alone if there were truly no effect. It says nothing about whether the effect is large or important.
Statistically significant ≠ practically important. With a huge sample, a trivially small difference can be 'significant'. Always ask: how big is the effect, and does it matter?
“Statistically significant” meansIt does not mean
Unlikely to have arisen by chance aloneLarge, or important
Under a specific null hypothesisThat the effect is real in the world
At a conventional threshold someone choseProven
A statement about the dataA statement about the decision to act
With a large enough sample, trivial differences become significant; with a small one, important differences do not. Significance answers “could this be noise?” and is silent on “is this worth anything?”, which is the question a programme actually has.
Ask for the effect size, always. “Significantly higher” with no magnitude attached is a claim about arithmetic; “three percentage points higher, significantly” is a claim you can weigh against the cost of acting on it.
ImpactMojoData Literacy 101www.impactmojo.in
The base-rate trap
Even a 99%-accurate test for a rare condition produces mostly false positives — because the healthy vastly outnumber the sick. Ignoring the underlying rate is one of the commonest reasoning errors.
Whenever you read 'X% accurate', ask how common the thing is to begin with. The base rate changes everything.
A test that is 99% accurateFor a condition affecting 1 in 1,000
Test 100,000 people100 have it
True positives~99
False positives (1% of 99,900)~999
So a positive result meansRoughly a 1 in 11 chance of having it
This is the single most useful piece of arithmetic in the deck. When a condition is rare, most positives from an accurate test are false, and the intuition that “99% accurate” means “99% likely” is wrong by an order of magnitude.
It generalises far beyond testing. Fraud flags, risk scores and targeting models applied to rare outcomes produce mostly false positives — which is why a screening system must be judged on how many wrongly flagged people it burdens, not on its accuracy rate.
ImpactMojoData Literacy 101www.impactmojo.in
Percent vs percentage points
If coverage rises from 40% to 50%, that is a 10 percentage-point increase — but a 25 percent increase. Mixing the two is a classic way to exaggerate or hide change.
+10 pp
percentage-point change (50 − 40)
+25%
relative change (10 ÷ 40)
From 20% to 25% isNot
A rise of 5 percentage pointsA rise of 5 per cent
A rise of 25 per cent, relativelyInterchangeable with the above
Written “pp” where space is tightSomething to leave ambiguous
The ambiguity is exploited routinely, because “a 25 per cent increase” sounds five times more impressive than “5 percentage points” and both describe the same change. Which one appears in a press release tells you what the release is for.
Write both, once: “from 20% to 25%, a rise of 5 percentage points (25% relative)”. It removes the ambiguity permanently and costs one clause.
ImpactMojoData Literacy 101www.impactmojo.in
'Doubled' can hide tiny numbers
'Cases doubled!' sounds alarming — but a rise from 2 to 4 is a doubling of almost nothing. Relative change without the absolute base is designed to impress, not inform.
Always pair the relative figure with the raw counts. '100% increase, from 2 to 4 cases' tells the honest story.
HeadlineWhat the absolute numbers were
“Risk doubled”From 1 in 100,000 to 2 in 100,000
“50% reduction in cases”From 4 to 2
“Three times more likely”0.3% versus 0.1%
“A 30% improvement”On an index nobody defines
Relative change without absolute numbers is not informative, and the choice to report it that way is usually deliberate. Both framings are true; only one lets a reader decide whether to care.
The same trick runs in reverse for harms. A small relative rise on a large base is reported in absolute terms to look minor, and a large relative rise on a tiny base is reported relatively to look alarming. Ask for both, in every direction.
ImpactMojoData Literacy 101www.impactmojo.in
Beware the chosen baseline
  • Start the time axis at a low year to exaggerate growth
  • Quote the one indicator that improved, ignore the rest
  • Compare to an unusual reference period (a drought, a peak)
  • Report only the subgroup that helps the argument
Ask: why this start date, this indicator, this comparison? What is left out?
MoveWhat to ask
The baseline is an unusual yearWhat does the series look like over ten years?
The window ends at a convenient pointWhat happened in the months after?
One district is highlightedHow did the others do?
One indicator is reportedWhat else was measured, and not shown?
Cherry-picking rarely involves a false number. Every figure presented is accurate; the selection is the argument, and selection leaves no trace in the data shown — only in the data withheld.
The defence is to ask for the full series, and to notice when a chart starts at an odd year. A baseline chosen because it was a drought year makes any subsequent year look like recovery.
ImpactMojoData Literacy 101www.impactmojo.in
Test enough things and something 'works'
If you slice the data twenty ways, roughly one slice will show a 'significant' result by chance. Reporting only that slice — p-hacking — manufactures false findings.
Trust analyses that were specified before seeing the data, and findings that replicate. Be wary of a single surprising subgroup result.
Choice made during analysisEach one multiplies the paths
Which outcome to focus onFive outcomes, five chances
Which subgroup to examineSex, caste, region, age — dozens of cells
Which controls to includeMany defensible specifications
Where to cut the sampleAbove/below median, quintiles, thresholds
Nobody needs to cheat for this to produce a false finding. An honest analyst making defensible choices, one at a time, guided by which result looks interesting, will find something — and will not have run a hundred tests, so will not think of it as multiple testing.
The guard is writing the analysis plan first. Naming the primary outcome and the subgroups before looking at the data converts a garden of paths into one path, and makes any later exploration honestly labelled as exploration.
ImpactMojoData Literacy 101www.impactmojo.in
Eight questions for any statistic
  • Who produced it, and what is their interest?
  • How was it measured — and what is the denominator?
  • Is it a sample? How big, how chosen?
  • What is the uncertainty / margin of error?
  • Percent or percentage points? Relative or absolute?
  • Who is missing from the count?
  • Correlation, or genuine causation?
  • Does it pass the common-sense smell test?
AskCatches
Out of how many?Counts masquerading as rates
Who collected it, and why?Interested measurement
Who is missing from it?Frame and coverage bias
Compared with what?Before-and-after with no counterfactual
How precise is it?Rankings built from noise
Percentage or percentage points?Inflated headlines
Why this baseline, this window?Cherry-picking
What else was tested?Forking paths
These eight questions require no statistics and catch most of what goes wrong. They are the practical residue of the whole course, and they can be asked in any meeting by anyone, which is the point of a literacy rather than a technique.
ImpactMojoData Literacy 101www.impactmojo.in
10
Section Ten
Data Ethics, Privacy & Equity
ImpactMojoData Literacy 101www.impactmojo.in
Every data point is a person
In development data, rows are people — often poor, often without power over how their information is used. Data ethics is not paperwork; it is respect made operational.
Data are not just numbers; they are people reduced to numbers. The reduction is never neutral.
— a principle of feminist data practice
The row recordsThe person experienced
hh_income = 0A month with nothing coming in
child_status = deceasedA death, recounted to a stranger with a tablet
violence_last12m = 1A disclosure that may be dangerous to have made
refused = 1A decision the dataset treats as an inconvenience
Data about people is collected from people, usually at a cost in their time and sometimes at a risk to their safety. That obligation does not weaken as the dataset grows; it is simply harder to feel at row 40,000.
The third row deserves specific care. A record of disclosed violence, stored with a name and a village, is a document that can put someone in danger. Such fields need stricter handling than the rest of the dataset, and usually get the same.
ImpactMojoData Literacy 101www.impactmojo.in
Informed consent is the floor
  • People should know what is collected and why
  • How it will be used, stored and shared — and for how long
  • That they can refuse or stop, with no penalty
  • Consent must be in a language and form they genuinely understand
A thumbprint on a form nobody explained is not consent. For children and other vulnerable groups, extra safeguards apply.
Consent is real whenNot when
The person knows what is collected and whyA form is read out at speed in a second language
Declining costs them nothingThe interviewer also decides programme eligibility
It names a purpose and a period“For research and related purposes”
Withdrawal is possible and explainedNobody knows how, including staff
The second row is the one our sector structurally fails. When the person seeking consent is connected to the organisation providing the service, refusal carries an implied cost whatever anyone says — and the remedy is to ask for less rather than to word the request better.
A quick self-test: how many people declined last round? If the answer is nobody, the process is collecting signatures rather than consent.
ImpactMojoData Literacy 101www.impactmojo.in
Anonymisation is harder than deleting names
Removing names is not enough. A combination of village + age + caste + occupation can re-identify one person — especially in small areas where few people share those traits.
Direct IDs
Name, Aadhaar, phone — remove
Quasi-IDs
Age + place + caste can re-identify — aggregate or coarsen
StepProtects against
Remove direct identifiersCasual lookup only
Coarsen quasi-identifiers — age bands, block not villageCombination attacks, the real risk
Suppress small cellsIdentifying the one person in a category
Restrict access rather than publishEverything, and it is the most-skipped option
Village plus age plus caste plus occupation identifies people in the small populations development data usually covers. Removing names is the beginning of anonymisation, not the whole of it, and “it is anonymised” describes an intention rather than a property.
Publishing at a coarser level costs analytical detail and is frequently the honest choice. A district-level release that cannot identify anyone is more useful than a village-level release that has to be withdrawn.
ImpactMojoData Literacy 101www.impactmojo.in
India's Digital Personal Data Protection Act, 2023
The DPDP Act, 2023 is India's first comprehensive data-protection law. It sets duties for anyone handling personal digital data — including NGOs and researchers.
  • Collect only what you need, for a stated purpose (purpose limitation)
  • Obtain free, informed, specific consent
  • Protect data with reasonable security safeguards
  • Stronger protections for children's data
Know your obligations before you collect. 'We're a small NGO' is not an exemption.
ObligationWhat it means for a research team
Lawful purpose and noticeRecord why each field is collected, in plain language
Access, correction, erasureSomeone must be able to act on a request
Breach notificationKnow who you would tell, and how fast
Children’s dataStricter treatment; verify before collecting
Being small or non-profit is not an exemption. The obligations attach to processing digital personal data, which is what a beneficiary database is, and an organisation holding one is covered whatever its size.
Take advice on specifics. Rules and exemptions under the Act have been evolving since it passed; this is orientation rather than legal guidance, and the detail is exactly the kind that changes between a deck being written and read.
ImpactMojoData Literacy 101www.impactmojo.in
What you don't disaggregate, you can't see
An average hides the people behind it. A programme can report good overall numbers while failing Dalit, Adivasi, disabled, or women beneficiaries. Disaggregation is how inequity becomes visible.
An overall average can mask large gaps between groups
Illustrative
ImpactMojoData Literacy 101www.impactmojo.in
Count the people the data forgets
  • Homeless and pavement-dwellers absent from household frames
  • Migrants counted nowhere — neither origin nor destination
  • Informal workers invisible to formal employment statistics
  • Trans and non-binary people erased by binary-only forms
'Missing data' is rarely random. The uncounted are usually the most marginalised — and policy built on the count leaves them out twice.
Systematically under-countedWhy
Homeless and pavement-dwelling peopleNo address, so absent from most frames
Seasonal migrantsCounted at neither origin nor destination
People with disabilitiesUnder-reported by proxy respondents; stigma
Transgender and gender-diverse peopleNo category, or one nobody selects
Unpaid care workNot counted as work by most instruments
What is not counted is not funded. Absence from data becomes absence from budgets and plans, so under-counting is not a technical shortfall but a political outcome with a technical cause.
Name the missing in your own reports. A sentence stating which groups the frame could not reach is honest, costs nothing, and is the only record that will exist of the gap.
ImpactMojoData Literacy 101www.impactmojo.in
Data colonialism and who benefits
Communities are often extracted from — surveyed repeatedly, with the knowledge and value flowing to outside institutions while respondents see nothing back.
  • Give findings back to the community in usable form
  • Involve people in defining what gets measured
  • Ask: who owns this data, and who profits from it?
Who typically holdsWho typically supplies
The dataset and the analysisThe time, the answers, the risk
The publication and the creditRarely a copy of the findings
The decision about reuseNo say in it
The asymmetry is ordinary rather than exceptional, and most of it is unintended: nobody decides to withhold findings from a village, it simply never gets scheduled. That is what makes it worth putting in a plan.
Two cheap corrections: budget a return visit to share results in the local language, and give the community a copy of what was collected about it. Both are rare, both are inexpensive, and both change the relationship for the next survey.
ImpactMojoData Literacy 101www.impactmojo.in
11
Section Eleven
Tools & Further Reading
ImpactMojoData Literacy 101www.impactmojo.in
Tools to grow into
ToolGood forNote
Spreadsheets (Excel, Google Sheets)Most everyday analysisStart here; learn pivot tables
RStatistics, reproducible analysis, graphicsFree, powerful, steeper curve
Python (pandas)Cleaning, large data, automationFree, general-purpose
KoboToolbox / ODKMobile survey data collectionFree, offline-capable
QGISMaps and spatial dataFree, open-source GIS
Power BI / Looker StudioDashboardsQuick visual reporting
Tools matter less than habits. A clear spreadsheet beats a confused script. Master the thinking first.
ToolWorth it when
SpreadsheetsSmall, one-off, and shared with non-analysts
R or PythonThe analysis will be repeated or must be reproducible
Stata / SPSSYour sector or collaborators already use it
QGISAnything genuinely spatial
The tool matters far less than the habits. A documented, reproducible spreadsheet analysis beats an undocumented script, and the reason to move beyond spreadsheets is repeatability rather than sophistication.
ImpactMojoData Literacy 101www.impactmojo.in
Open data you can use today
  • data.gov.in — India's open government data portal
  • censusindia.gov.in — Census tables & maps
  • NFHS / DHS Program — health & demographic data
  • MoSPI — NSS, PLFS, national accounts
  • World Bank Open Data & Our World in Data — global comparisons
SourceBest for
data.gov.inMinistry datasets across sectors
censusindia.gov.inVillage and ward-level population data
rchiips.org (NFHS)Health, nutrition, women’s status by district
mospi.gov.inPLFS, NSS, national accounts
UDISE+ and HMIS portalsSchool and facility administrative data
Read the methodology note before the data. Every one of these publishes one, and it contains the sample design, the definitions and the lowest level at which estimates hold — which is to say, everything Sections 3 and 7 said you must know.
Save the file with its download date. These portals update and revise; an analysis that cannot say which vintage it used cannot be reproduced, and revisions are common enough to matter.
ImpactMojoData Literacy 101www.impactmojo.in
A short, honest reading list
  • How to Lie with Statistics — Darrell Huff (still the classic primer)
  • The Visual Display of Quantitative Information — Edward Tufte
  • Data Feminism — D'Ignazio & Klein (power and data)
  • Factfulness — Hans Rosling (reading global data well)
  • Poor Economics — Banerjee & Duflo (evidence in development)
Pair this deck with ImpactMojo's Exploratory Data Analysis, Qualitative Methods and Research Ethics 101 courses.
ReadFor
Tufte, The Visual Display of Quantitative InformationSection 5, from the source
Rosling, FactfulnessReading global statistics without panic or complacency
D’Ignazio & Klein, Data FeminismWho is counted, and who decides the categories
Wheelan, Naked StatisticsThe concepts in Sections 4 and 6, informally
Survey methodology notes (NFHS, PLFS)The most useful reading on this list, and free
The last row is not a joke. An afternoon with the NFHS methodology note teaches more applied data literacy than most textbooks, because it is the document behind numbers you will actually quote.
ImpactMojoData Literacy 101www.impactmojo.in
If you remember five things
  • Always ask where the data came from — and who is missing
  • Plot it before you trust any summary number
  • Median over mean for skewed things like money
  • Correlation is not causation — look for the confounder
  • Behind every row is a person — handle with care
If you remember one thing per sectionSection
Ask “out of how many?”4
Use the median for money4
Plot it before you quote a correlation6
Check what level the survey was designed for2 and 7
Say who is missing from the data7 and 10
None of these requires software or statistical training. They are questions, and they are the ones that catch the errors that actually reach published reports in this sector — which is why a data literacy course ends with questions rather than techniques.
And the habit underneath all five: treat every number as the end of a process with choices in it, and ask what they were.
ImpactMojoData Literacy 101www.impactmojo.in
Data Literacy 101 · Complete
Now go question
the next number.
CC BY-NC-ND 4.0·Free Forever·ImpactMojo 101 Series