| Not data literacy | Data literacy |
|---|---|
| Knowing the formula for a standard deviation | Knowing when an average is the wrong summary |
| Making a chart | Noticing that the axis starts at 90 |
| Quoting a survey figure | Asking what level the survey was designed to report at |
| Running a significance test | Knowing what “significant” does and does not claim |
| Decision | The data question underneath it |
|---|---|
| Which blocks to prioritise | Is the estimate reliable at block level, or borrowed from state? |
| Whether the intervention worked | Compared with what — before, or a comparable group? |
| Setting a target | What was the trend before we arrived? |
| Reporting reach | Out of how many — who is the denominator? |
| Step | What can go wrong here |
|---|---|
| Data — 1,240 children weighed | Scale uncalibrated; some hamlets never visited |
| Information — 18% underweight | Which reference standard? Denominator = weighed, or eligible? |
| Knowledge — concentrated in 3 hamlets | Are the numbers there large enough to be more than noise? |
| Decision — feed those hamlets | Was concentration caused by the problem, or by who got measured? |
| Habit | The question in the room |
|---|---|
| Where did it come from? | “Who collected this, when, and for what purpose?” |
| What does it measure? | “What exactly was counted, and for whom?” |
| Who is missing? | “Who would not have appeared in this dataset at all?” |
| How sure are we? | “What is the sample size behind this cell?” |
| What decision does it serve? | “What would we do differently if it were half as large?” |
| Choice made when data is created | What it decides |
|---|---|
| What to measure | Unmeasured problems are politically invisible |
| Which categories to use | Who has a box to tick, and who is “other” |
| Whom to ask | Household head interviews report the head’s view |
| What counts as a case | A definition change can halve a rate overnight |
| What unit to report | State averages hide district collapse |
| Section | What you will be able to do |
|---|---|
| 2–3 · Sources and indicators | Name the right national dataset; turn a concept into a defensible indicator |
| 4–5 · Describing and visualising | Choose mean or median correctly; spot a truncated axis |
| 6 · Relationships | Explain confounding, ecological fallacy and Simpson’s paradox to a colleague |
| 7–8 · Sampling and quality | Say what a sample can and cannot support; run a cleaning pipeline |
| 9–11 · Reading and ethics | Interrogate a statistic; handle data about people |
| Question | Which family answers it |
|---|---|
| How many girls dropped out this year? | Quantitative |
| Why did they drop out? | Qualitative |
| Did dropout fall after the intervention? | Quantitative |
| What did families think the intervention was for? | Qualitative |
| Is the fall large enough to be real? | Quantitative |
| Why did it work here and not there? | Qualitative |
| Primary | Secondary | |
|---|---|---|
| Source | You collect it | Someone else collected it |
| Example | Your baseline survey, FGDs | Census, NFHS, district HMIS |
| Control | High — you design it | Low — you take it as given |
| Cost / time | High | Usually low |
| Risk | Fieldwork error, bias | May not fit your question |
| Type | Common in our work | What it needs |
|---|---|---|
| Structured | Survey tables, MIS exports, registers | Spreadsheets, standard statistics |
| Semi-structured | Forms with open-text fields, tagged records | Coding of the free text before analysis |
| Unstructured | Field notes, call logs, photographs, voice notes | Qualitative method, or tooling most teams lack |
| Form | Can answer | Cannot answer |
|---|---|---|
| Cross-section | How things stand now, across places | Whether this household improved |
| Time series | Whether the aggregate moved | Who moved, and who was replaced |
| Panel | Whether the same units changed | Much, if attrition is high and non-random |
| Source | What it covers | Frequency |
|---|---|---|
| Census of India | Every person — population, literacy, housing, migration | Decennial (2011 latest) |
| NFHS | Health, nutrition, fertility, anaemia, women's status | ~5 years (NFHS-5: 2019–21) |
| NSS / PLFS | Consumption, employment, unemployment | PLFS annual since 2017–18 |
| SRS | Birth & death rates, infant mortality, life expectancy | Annual |
| HMIS | Facility-level health service delivery | Monthly |
| SECC 2011 | Socio-economic & caste deprivation indicators | One-off (2011) |
| Question | Go to |
|---|---|
| Anaemia, immunisation, child nutrition | NFHS |
| Employment, unemployment, wages | PLFS |
| Population by village or ward | Census |
| Infant mortality, life expectancy | SRS |
| Facility-level service delivery, monthly | HMIS |
| Claim | Supportable? |
|---|---|
| “Anaemia among women in Bihar is X%” (NFHS) | Yes — designed for state estimates |
| “Anaemia in this district is X%” (NFHS) | Usually yes, with wider uncertainty |
| “Anaemia in this block is X%” (NFHS) | No — the sample was never designed for it |
| “Population of this village” (Census) | Yes — a census counts everyone |
| “Village anaemia” (Census) | No — the Census does not measure it |
| Why size buys reliability | What it does not buy |
|---|---|
| District-level estimates with usable precision | Block or village estimates |
| Disaggregation by sex, residence, wealth quintile | Every cross-tabulation you might want |
| Comparison across rounds | Comparison where the question changed between rounds |
| Administrative data records | It cannot tell you |
|---|---|
| Who received the service | Who was eligible and did not come |
| What was reported by the provider | What actually happened, where reporting is incentivised |
| Transactions completed | Attempts that failed at the counter |
| Enrolment | Attendance, and whether anything was learned |
| Digital trace | Who it under-represents |
|---|---|
| Mobile-phone records | Women, the elderly, the poorest — who own phones less |
| Transaction logs | Cash economies, which is most of the informal sector |
| Satellite night-lights | Activity that is not electrified |
| Social media | Almost everyone this sector works with |
| Concept | A defensible indicator | What it still misses |
|---|---|---|
| Women’s empowerment | Can visit a health centre alone | Whether she wants to; whether there is one |
| Learning | Can read a Class 2 text | Comprehension, reasoning, everything untested |
| Food security | Months of adequate food provisioning | Quality, diversity, who eats last |
| Access to water | Improved source within 30 minutes | Reliability, seasonality, queueing, who fetches it |
| Operationalisation must fix | Example of it going wrong |
|---|---|
| Exactly what to count | “Trained” — attended, completed, or passed? |
| For whom | “Children” — under 5, under 6, or school-age? |
| Over what period | “Last year” — calendar, financial, or recall? |
| In what units | Households or individuals — a 4× difference |
| Counted by whom | Self-report, observation, or register |
| Level | Meaning | Example | Valid maths |
|---|---|---|---|
| Nominal | Labels, no order | District, caste, religion | Counts, mode |
| Ordinal | Ordered, unequal gaps | Wealth quintile, Likert scale | Median, rank |
| Interval | Equal gaps, no true zero | Temperature (°C), calendar year | Mean, difference |
| Ratio | Equal gaps, true zero | Income, age, children ever born | All, ratios |
| Illegal move | Why |
|---|---|
| Average of district codes | Nominal — the numbers are names |
| “Mean wealth quintile = 3.2” | Ordinal — gaps between quintiles are not equal |
| “Satisfaction rose 0.4 points” from a Likert scale | Ordinal treated as interval; common, and contested |
| “Twice as hot” in °C | Interval — no true zero, so ratios are meaningless |
| Case | Reliable? | Valid? |
|---|---|---|
| Scale reads 2 kg high, every time | Yes | No |
| Scale drifts randomly by 3 kg | No | No |
| Calibrated scale | Yes | Yes |
| Asking men about women’s decision-making | Often yes | No — consistently the wrong respondent |
| Proxy | For | What it leaks |
|---|---|---|
| Asset index | Wealth | Debt, income flow, and urban/rural comparability |
| Night-lights | Economic activity | Informal and unelectrified activity |
| MUAC | Acute malnutrition | Chronic undernutrition; oedema |
| Enrolment | Education | Attendance, and whether anything was learned |
| Bank account opened | Financial inclusion | Whether it is ever used |
| Question to ask of any index | Why it matters |
|---|---|
| What are the weights, and who chose them? | Weights are values, presented as arithmetic |
| Can a bad component be masked? | Aggregation hides exactly what you need to act on |
| Is the cut-off a natural break or a choice? | Move the threshold, move the headline |
| Are the components correlated? | If so, one thing is being counted several times |
| Design choice in the MPI | What changes if you change it |
|---|---|
| Deprived in a weighted third or more = poor | The headcount rises or falls, with no change on the ground |
| Which 12 indicators are included | Which deprivations count as poverty at all |
| Equal weight to three dimensions | Health, education and living standards treated as equally important |
| Headcount × intensity | Rewards moving people just over the line as much as deep gains |
| Before any modelling, report | Because |
|---|---|
| N, and N per subgroup you will discuss | Half of all overreach is a small cell |
| Missingness per variable | A clean-looking mean may rest on 60% of cases |
| Centre and spread | Two districts with one mean can be nothing alike |
| Shape, by eye | Skew decides whether the mean is honest |
| Range and extremes | Impossible values are found here, not later |
| Measure | What it is | Best when |
|---|---|---|
| Mean | Arithmetic average | Roughly symmetric data, no wild outliers |
| Median | Middle value when sorted | Skewed data — income, land, wealth |
| Mode | Most frequent value | Categories — commonest crop, caste, response |
| Variable | Use | Why |
|---|---|---|
| Household income or consumption | Median | Right-skewed; a few large values dominate the mean |
| Landholding | Median | Same, usually more extreme |
| Child height or weight | Mean | Roughly symmetric |
| Days of work in a month | Both | Bounded; report the distribution |
| Commonest crop or caste | Mode | Categorical — no other option is meaningful |
| Measure | Robust to outliers? | Use when |
|---|---|---|
| Range | No — defined by them | Quick sanity check for impossible values |
| IQR | Yes | Skewed data; reporting alongside a median |
| Standard deviation | No | Roughly symmetric data |
| p90/p10 ratio | Fairly | Comparing inequality across places |
| Term | Means | Watch out for |
|---|---|---|
| Percentile | Value below which x% of cases fall | Not the same as “x% higher” |
| Quartile | Four equal-sized groups | Groups are equal in count, not in width |
| Wealth quintile | Five equal-sized groups, by asset rank | Relative to the survey population, not to a rupee value |
| p90/p10 | Ratio of the 90th to the 10th percentile | Ignores everything above and below |
| Roughly normal | Usually not |
|---|---|
| Adult height; birth weight | Income, consumption, landholding |
| Measurement error around a true value | Firm size, farm size, city size |
| Sample means of large samples | Counts of rare events |
| Many small independent influences | Anything with a floor at zero and a long tail |
| Outlier | First question |
|---|---|
| Household of 80 members | Is it a joint family, a hostel, or a typing error? |
| Income of zero | No income, refused to answer, or not asked? |
| One block with triple the dropout | Is it real — and if so, that is the finding |
| Age 999 | A missing-value code someone forgot to declare |
| Count reported | The denominator that changes its meaning |
|---|---|
| “500 dropouts” | Out of 600, or out of 60,000 |
| “40 maternal deaths” | Per 100,000 live births, not per district |
| “12,000 beneficiaries reached” | Out of how many eligible |
| “Cases doubled” | From 3 to 6, or from 3,000 to 6,000 |
| “Most crime happens in this district” | Most people also live in this district |
| A chart is doing its job when | It has failed when |
|---|---|
| The reader sees the comparison you intended | They have to read the caption to know what to look at |
| The visual size matches the numeric difference | A 4-point rise looks like a tripling |
| Every axis and unit is labelled | “Index” with no baseline or source |
| It survives being printed in black and white | The only distinction is red versus green |
| You want to show… | Use | Avoid |
|---|---|---|
| Change over time | Line chart | Pie chart |
| Comparison across categories | Bar chart | 3-D anything |
| Composition / shares of a whole | Stacked bar (or 1 pie, few slices) | Many pies |
| Relationship between two variables | Scatter plot | Dual-axis tricks |
| Distribution of one variable | Histogram / box plot | Single average |
| Geographic pattern | Choropleth map | Map coloured by raw counts |
| Element | Why it is not optional |
|---|---|
| Title stating the finding | Most readers read only this |
| Axis labels with units | “Rate” per what, over what period? |
| Source and date | An unsourced number is not evidence |
| N, or the sample base | Percentages of nine cases look like percentages |
| A note on what is excluded | Missing categories change the reading |
| Remove | Keep |
|---|---|
| 3-D effects, shadows, gradients | Direct labels on the lines or bars |
| Heavy gridlines and borders | A light gridline where reading a value matters |
| Decorative backgrounds and icons | Annotation that names the point |
| A legend the reader must cross-reference | The source, the units and the N |
| Data type | Palette | Common error |
|---|---|---|
| Categories | Distinct hues | A rainbow that implies an order |
| Sequential magnitude | One hue, light to dark | Using a diverging scale with no midpoint |
| Diverging around a midpoint | Two hues from a neutral centre | Placing the neutral point arbitrarily |
| Any | — | Red/green as the only distinction |
| Instead of | Use small multiples |
|---|---|
| Twelve lines on one chart | Twelve small charts, one per district |
| A dual-axis chart | Two stacked panels sharing an x-axis |
| An animated chart | A row of frames the reader can compare at once |
| Many pies | A single sorted bar chart |
| Use a table when | Use a chart when |
|---|---|
| Exact values matter | The pattern matters more than the values |
| There are few rows and several columns | There are many observations |
| Readers will look up their own district | Readers need one comparison |
| The numbers will be quoted | The shape will be remembered |
| Check | Fail state |
|---|---|
| Does the axis start at zero, or is the truncation justified and labelled? | The most common way charts lie |
| Is the title the finding, not the variable name? | “Figure 3: Enrolment” |
| Are source, date and N present? | Unsourced, undated, unbased |
| Is it readable in greyscale? | Red/green only |
| Would someone who disagrees find it fair? | The test that catches the rest |
| r | Means | Does not mean |
|---|---|---|
| +0.9 | Strong positive linear association | That one causes the other |
| 0 | No linear relationship | No relationship — a U-shape gives r near 0 |
| −0.5 | Moderate negative association | That the effect is moderate in size |
| Any value | Something about the sample | Anything about an individual case |
| Explanation | Example |
|---|---|
| A causes B | The one usually assumed |
| B causes A | Programme presence correlates with need — because need attracted the programme |
| C causes both | Ice-cream sales and drowning; the confounder is summer |
| Selection | Villages that volunteered differ from those that did not |
| Chance | Test enough pairs and some will correlate |
| Apparent relationship | Plausible confounder |
|---|---|
| Toilet ownership and child height | Household wealth, which drives both |
| Attending a training and higher yields | Farmers who attend differ in land, literacy and motivation |
| Mobile ownership and women’s mobility | Urban residence |
| Private schooling and test scores | Parental education and income |
| How spurious findings arise | Guard |
|---|---|
| Many variables compared pairwise | State the number of comparisons made |
| Both series trend over time | Compare changes, not levels |
| Small samples produce large r by chance | Report N with every correlation |
| Only the interesting result is written up | Pre-specify what you will test |
| Anscombe’s four datasets share | And look like |
|---|---|
| The same mean in x and y | A clean linear relationship |
| The same variance | A curve |
| The same correlation | A line plus one extreme outlier |
| The same regression line | A vertical cluster plus one distant point |
| True at group level | Not therefore true of individuals |
|---|---|
| States with higher female literacy have lower fertility | That any literate woman has fewer children |
| Districts with more migrants report more remittances | That migrant households receive more |
| Richer states have higher obesity | That richer individuals are heavier |
| Where the paradox shows up | The hidden variable |
|---|---|
| Hospital A has worse survival than B overall | A takes the severe cases |
| A scheme looks less effective overall than by district | Uptake differs by district size |
| Wages fall overall while rising in every sector | Employment shifted toward lower-paying sectors |
| Situation | Why improvement appears without any cause |
|---|---|
| Targeting the worst-performing blocks | Extreme values are partly luck; luck does not repeat |
| Selecting schools with the lowest scores | Some were having a bad year, and would recover anyway |
| Intervening after a spike in cases | Spikes subside; the intervention gets the credit |
| Sampling buys | At the cost of |
|---|---|
| Speed — weeks instead of years | Uncertainty, which must be reported |
| Depth — longer interviews, better training | Small-area estimates |
| Repeatability — annual rounds | Comparability if the design changes |
| Lower cost per respondent | Nothing, if the sample is well drawn |
| Term | Meaning | Where it goes wrong |
|---|---|---|
| Target population | Everyone you want to describe | Stated vaguely: “the community” |
| Sampling frame | The list you actually sample from | Excludes the homeless, migrants, new arrivals |
| Sample | Those selected | Confused with those who responded |
| Respondents | Those who actually answered | Non-response treated as random |
| Method | How | Use when |
|---|---|---|
| Simple random | Every unit equal chance | You have a full list |
| Systematic | Every k-th unit from a list | Ordered list, no hidden cycle |
| Stratified | Split into groups, sample each | You must represent subgroups |
| Cluster | Sample whole groups (villages) | People are geographically spread |
| Multistage | Clusters, then units within | Large national surveys (NFHS) |
| Design | When it fits | Watch |
|---|---|---|
| Simple random | A good complete frame exists | Rarely practical over a large area |
| Systematic | An ordered list, easy in the field | Hidden periodicity in the list order |
| Stratified | You need estimates for subgroups | Strata must be defined before sampling |
| Cluster | Travel cost dominates | Larger uncertainty for the same N |
| Multi-stage | National surveys — NFHS, NSS | Weights become essential |
| Method | Legitimate use | Illegitimate use |
|---|---|---|
| Convenience | Piloting a questionnaire | Any prevalence estimate |
| Purposive | Selecting information-rich cases for qualitative work | “Representative” claims |
| Snowball | Reaching hidden or stigmatised populations | Estimating population size |
| Quota | Fast market-style polling | Confidence intervals |
| Population | Sample for ±3% | Fraction sampled |
|---|---|---|
| 10,000 | ~1,000 | 10% |
| 1,000,000 | ~1,070 | 0.1% |
| 1,400,000,000 | ~1,070 | 0.00008% |
| Bias | How it enters |
|---|---|
| Coverage | The frame omits a group entirely |
| Non-response | Those absent or refusing differ systematically |
| Selection by the enumerator | The nearest, easiest, most welcoming households |
| Social desirability | Answers shaped by what is acceptable to say |
| Recall | Distant events remembered selectively |
| Weights correct for | Consequence of ignoring them |
|---|---|
| Unequal probability of selection | Over-sampled groups dominate every estimate |
| Deliberate over-sampling of small groups | National figures skewed toward the boosted stratum |
| Non-response, adjusted post hoc | The respondents stand in for everyone |
| Bad question | Why |
|---|---|
| “Do you agree that the scheme has improved your life?” | Leading; invites agreement |
| “How satisfied are you with health and education services?” | Double-barrelled — two questions, one answer |
| “How much did you spend on food last year?” | Recall period far too long |
| “Do you own assets?” | Undefined; every respondent answers a different question |
| “You do send your daughter to school, don’t you?” | Social desirability, at maximum |
| Where the time actually goes | Roughly |
|---|---|
| Finding, understanding and merging data | The largest single share |
| Cleaning, recoding, handling missingness | The next largest |
| Analysis | Much less than anyone expects |
| Writing up and making charts | More than budgeted |
| Problem | Example | Risk |
|---|---|---|
| Missing values | Blank income field | Biased averages if not random |
| Duplicates | Same beneficiary twice | Inflated counts |
| Inconsistent codes | 'F' / 'Female' / '2' | Broken grouping |
| Outliers / impossible | Age = 200, −5 children | Distorted statistics |
| Format drift | DD/MM vs MM/DD dates | Silent miscalculation |
| Typos in keys | Misspelt village name | Failed merges |
| Symptom | Typical cause |
|---|---|
| “Bihar”, “bihar”, “BIHAR”, “Bihar ” | Free-text entry with no controlled list |
| Ages of 0 and 999 in the same column | Undeclared missing codes |
| Dates in three formats | Excel, and different data-entry operators |
| Duplicate household IDs | Re-visits recorded as new records |
| Numbers stored as text | Leading zeros, commas, or a stray space |
| Why it is missing | What you can do |
|---|---|
| Missing at random — a skipped page | Analysis is largely unaffected |
| Related to something you measured | Can be adjusted for, with care |
| Related to the answer itself | The dangerous case: no fix from the data alone |
| Step | Rule |
|---|---|
| Keep the raw file untouched | Read-only; never edited, ever |
| Every change in a script | Not by hand in a spreadsheet |
| One script, run start to finish | Reproduces the clean file from the raw one |
| Log what was dropped and why | Counts before and after each filter |
| Output a documented clean dataset | With a codebook that matches it |
| Test | If you fail it |
|---|---|
| Could a colleague rerun this from the raw data? | The result cannot be checked |
| Could you, in six months? | The result cannot be updated |
| Does the number in the report match the script’s output? | Something was changed by hand |
| Is the exact data version recorded? | Re-running gives a different answer |
| Check | Catches |
|---|---|
| Range checks on every numeric field | Ages of 200; incomes of −5 |
| Controlled lists for categories | Four spellings of one district |
| Skip logic enforced in the form | Pregnancy answers from men |
| Unique-ID constraint | Duplicates, at entry rather than at analysis |
| Daily review of the first week’s data | An enumerator misunderstanding a question |
| Document | Because in six months |
|---|---|
| What each variable means, and its units | “inc2” will mean nothing to anyone |
| Category codes and missing codes | 99 will quietly enter an average |
| How the data was collected, and when | Comparability depends on it |
| Known limitations and gaps | The caveats live only in your head |
| Who to ask | That person will have left |
| Practice | Prevents |
|---|---|
| Dated, immutable raw files | “Which version produced the report?” |
| Clear naming — no “final_v3_FINAL” | Analysing the wrong file |
| A change log, even a text file | Unexplained differences between runs |
| Version control for scripts | Losing a working analysis to an edit |
| Reported as | Should be read as |
|---|---|
| “Anaemia is 52.1%” | “Between about 50 and 54, on this survey’s design” |
| “Up from 51.8% last round” | Possibly unchanged — the intervals overlap |
| “District A is worse than B” | Check whether the intervals separate |
| A single decimal place | Usually false precision |
| “Statistically significant” means | It does not mean |
|---|---|
| Unlikely to have arisen by chance alone | Large, or important |
| Under a specific null hypothesis | That the effect is real in the world |
| At a conventional threshold someone chose | Proven |
| A statement about the data | A statement about the decision to act |
| A test that is 99% accurate | For a condition affecting 1 in 1,000 |
|---|---|
| Test 100,000 people | 100 have it |
| True positives | ~99 |
| False positives (1% of 99,900) | ~999 |
| So a positive result means | Roughly a 1 in 11 chance of having it |
| From 20% to 25% is | Not |
|---|---|
| A rise of 5 percentage points | A rise of 5 per cent |
| A rise of 25 per cent, relatively | Interchangeable with the above |
| Written “pp” where space is tight | Something to leave ambiguous |
| Headline | What the absolute numbers were |
|---|---|
| “Risk doubled” | From 1 in 100,000 to 2 in 100,000 |
| “50% reduction in cases” | From 4 to 2 |
| “Three times more likely” | 0.3% versus 0.1% |
| “A 30% improvement” | On an index nobody defines |
| Move | What to ask |
|---|---|
| The baseline is an unusual year | What does the series look like over ten years? |
| The window ends at a convenient point | What happened in the months after? |
| One district is highlighted | How did the others do? |
| One indicator is reported | What else was measured, and not shown? |
| Choice made during analysis | Each one multiplies the paths |
|---|---|
| Which outcome to focus on | Five outcomes, five chances |
| Which subgroup to examine | Sex, caste, region, age — dozens of cells |
| Which controls to include | Many defensible specifications |
| Where to cut the sample | Above/below median, quintiles, thresholds |
| Ask | Catches |
|---|---|
| Out of how many? | Counts masquerading as rates |
| Who collected it, and why? | Interested measurement |
| Who is missing from it? | Frame and coverage bias |
| Compared with what? | Before-and-after with no counterfactual |
| How precise is it? | Rankings built from noise |
| Percentage or percentage points? | Inflated headlines |
| Why this baseline, this window? | Cherry-picking |
| What else was tested? | Forking paths |
| The row records | The person experienced |
|---|---|
| hh_income = 0 | A month with nothing coming in |
| child_status = deceased | A death, recounted to a stranger with a tablet |
| violence_last12m = 1 | A disclosure that may be dangerous to have made |
| refused = 1 | A decision the dataset treats as an inconvenience |
| Consent is real when | Not when |
|---|---|
| The person knows what is collected and why | A form is read out at speed in a second language |
| Declining costs them nothing | The interviewer also decides programme eligibility |
| It names a purpose and a period | “For research and related purposes” |
| Withdrawal is possible and explained | Nobody knows how, including staff |
| Step | Protects against |
|---|---|
| Remove direct identifiers | Casual lookup only |
| Coarsen quasi-identifiers — age bands, block not village | Combination attacks, the real risk |
| Suppress small cells | Identifying the one person in a category |
| Restrict access rather than publish | Everything, and it is the most-skipped option |
| Obligation | What it means for a research team |
|---|---|
| Lawful purpose and notice | Record why each field is collected, in plain language |
| Access, correction, erasure | Someone must be able to act on a request |
| Breach notification | Know who you would tell, and how fast |
| Children’s data | Stricter treatment; verify before collecting |
| Systematically under-counted | Why |
|---|---|
| Homeless and pavement-dwelling people | No address, so absent from most frames |
| Seasonal migrants | Counted at neither origin nor destination |
| People with disabilities | Under-reported by proxy respondents; stigma |
| Transgender and gender-diverse people | No category, or one nobody selects |
| Unpaid care work | Not counted as work by most instruments |
| Who typically holds | Who typically supplies |
|---|---|
| The dataset and the analysis | The time, the answers, the risk |
| The publication and the credit | Rarely a copy of the findings |
| The decision about reuse | No say in it |
| Tool | Good for | Note |
|---|---|---|
| Spreadsheets (Excel, Google Sheets) | Most everyday analysis | Start here; learn pivot tables |
| R | Statistics, reproducible analysis, graphics | Free, powerful, steeper curve |
| Python (pandas) | Cleaning, large data, automation | Free, general-purpose |
| KoboToolbox / ODK | Mobile survey data collection | Free, offline-capable |
| QGIS | Maps and spatial data | Free, open-source GIS |
| Power BI / Looker Studio | Dashboards | Quick visual reporting |
| Tool | Worth it when |
|---|---|
| Spreadsheets | Small, one-off, and shared with non-analysts |
| R or Python | The analysis will be repeated or must be reproducible |
| Stata / SPSS | Your sector or collaborators already use it |
| QGIS | Anything genuinely spatial |
| Source | Best for |
|---|---|
| data.gov.in | Ministry datasets across sectors |
| censusindia.gov.in | Village and ward-level population data |
| rchiips.org (NFHS) | Health, nutrition, women’s status by district |
| mospi.gov.in | PLFS, NSS, national accounts |
| UDISE+ and HMIS portals | School and facility administrative data |
| Read | For |
|---|---|
| Tufte, The Visual Display of Quantitative Information | Section 5, from the source |
| Rosling, Factfulness | Reading global statistics without panic or complacency |
| D’Ignazio & Klein, Data Feminism | Who is counted, and who decides the categories |
| Wheelan, Naked Statistics | The concepts in Sections 4 and 6, informally |
| Survey methodology notes (NFHS, PLFS) | The most useful reading on this list, and free |
| If you remember one thing per section | Section |
|---|---|
| Ask “out of how many?” | 4 |
| Use the median for money | 4 |
| Plot it before you quote a correlation | 6 |
| Check what level the survey was designed for | 2 and 7 |
| Say who is missing from the data | 7 and 10 |