| If you are | Start at |
|---|---|
| Commissioning a survey | Sections 1–3 |
| Writing the instrument | Sections 4–6 |
| Running fieldwork | Sections 7, 8, 10 |
| This section answers | Slide |
|---|---|
| What a survey actually is | 4 |
| When one is the right tool — and when it is not | 5–6 |
| Why to exhaust free data first | 7 |
| What a survey can and cannot give you | 8–9 |
| What you owe the respondent | 10 |
| Property | What it buys you | What it costs |
|---|---|---|
| Standardisation | Answers are comparable across people | Anything not anticipated is unrecordable |
| Sampling | A few thousand describe millions | An estimate with a margin, not a fact |
| Closed options | Fast, countable data | You have decided the answer space in advance |
| Repetition | Change over time is measurable | Improving the question breaks the trend |
| The question you have | Right tool |
|---|---|
| What share of households have a toilet? | Survey |
| Why do people with toilets not use them? | Qualitative |
| How many children enrolled last year? | Administrative (UDISE+) |
| Did our programme cause the change? | Survey plus a design (control, baseline) |
| Survey | Qualitative | Administrative | |
|---|---|---|---|
| Answers | How many, how much | Why, how, meaning | What the system records |
| Strength | Generalisable, comparable | Depth, context, mechanism | Cheap, continuous |
| Sample | Representative sample | Small, purposive | Everyone served |
| Cost | High — fieldwork | Moderate | Already collected |
| Blind spot | Misses the 'why' | Cannot generalise | Misses who is not served |
| Failure mode | Survey | Qualitative | Administrative |
|---|---|---|---|
| Wrong people | Frame gaps, non-response | Convenience selection | Only those the system touched |
| Wrong answers | Bad wording, desirability | Interviewer steer | Reporting incentives |
| Cost of fixing | High — re-field | Low — ask again | None — you get what exists |
| Source | Covers | Watch for |
|---|---|---|
| Census | Every household, decadal | 2011 is the last full round; badly dated |
| NFHS | Health, nutrition, women’s status; district level | Sample sizes thin below district |
| PLFS / NSS | Employment, consumption; quarterly urban | Definitions of ‘work’ are technical |
| HMIS · UDISE+ · MGNREGA MIS | Continuous, administrative | Records the system, not the population |
| Claim you can make | Only because |
|---|---|
| ‘38% of households, ±3 points’ | The sample was drawn probabilistically |
| ‘Higher in Block A than Block B’ | The same instrument was used in both |
| ‘Up from 31% in 2023’ | The question and reference period did not change |
| Limit | What it looks like in a report |
|---|---|
| Meaning is flattened | ‘Satisfied’ hides four different reasons |
| Only the asked is known | No line for the thing you did not anticipate |
| Self-report | ‘Reported open defecation’ ≠ observed |
| The missing | Migrants, homeless, institutionalised — absent, uncounted |
| Question you are about to add | Ask instead |
|---|---|
| ‘Might be useful for analysis’ | Which planned table needs it? |
| ‘The donor may want it’ | Will anyone read it? What decision changes? |
| ‘We have always asked it’ | Has it ever appeared in a report? |
| Stage | The mistake that is expensive here |
|---|---|
| Objectives | Starting from questions, not the decision |
| Indicators | Leaving the definition to the analyst |
| Instrument | Skipping cognitive pretesting |
| Sample | A frame nobody inspected |
| Fieldwork | No back-checks while data is still collectable |
| Cleaning & analysis | Discovering a missing variable |
| Work backwards | Example |
|---|---|
| Decision | Whether to extend the cash transfer to Block C |
| Finding needed | Whether Block C’s food-insecurity rate resembles A and B |
| Indicator | FIES moderate-or-severe, past 12 months, household |
| Questions | The eight standard FIES items, translated and tested |
| Component | Vague version | Answerable version |
|---|---|---|
| Population | ‘women in the district’ | women aged 18–45 in Block X |
| Indicator | ‘livelihoods’ | earned any cash income |
| Reference period | — | in the last 30 days |
| Comparison | — | vs the same measure in 2024 |
| Choice | If you get it wrong |
|---|---|
| Unit of analysis | Household and individual rows mixed — denominators wrong |
| Respondent vs subject | ‘Proxy reporting’ unlabelled — mother’s report read as the child’s |
| Reference period | ‘Usually’ means something different to every respondent |
| Inclusion rule | Who counts as a household member differs by enumerator |
| Dummy table | Variables it demands |
|---|---|
| Food insecurity by block, 2024 vs 2026 | FIES score, block, round |
| Income by sex of respondent | Cash income, sex, reference period |
| Toilet use by caste group | Use item, social group, weights |
| Constraint | Pushes toward | Watch-out |
|---|---|---|
| Tight budget | Shorter instrument, phone mode | Coverage & quality loss |
| Short timeline | Smaller sample, fewer items | Underpowered estimates |
| Few trained staff | Simpler skips, CAPI logic | Enumerator error |
| Remote terrain | Cluster sampling, offline tools | Travel cost & fatigue |
| Constraint | Honest response | Dishonest response |
|---|---|---|
| Tight budget | Cut items, keep the sample and training | Keep everything, cut training |
| Short timeline | Smaller sample, state the wider margin | Same sample, rushed fieldwork |
| Few trained staff | Simpler instrument, CAPI logic | Complex paper skips |
| Document | Answers | Who needs it |
|---|---|---|
| Protocol | Why, whom, how many, how | Reviewers, ethics, successors |
| Questionnaire with variable names | What was asked, in what order | Analyst |
| Field manual | What to do at every edge case | Enumerators, supervisors |
| Data dictionary | What each variable, code and unit means | Anyone reusing the data |
| What sends you back | To fix |
|---|---|
| Cognitive pretest | Wording, recall, option sets |
| Pilot | Routing, length, logistics, sample |
| Back-checks | Enumerator practice, mid-field |
| Analysis | What the next round should ask |
| Concept | A concrete stand-in | What the stand-in misses |
|---|---|---|
| Food security | FIES: 8 experience items, 12 months | Quality and dietary diversity |
| Empowerment | Who decides on large purchases | Whether she wanted to decide |
| Trust in the panchayat | Would you approach it with a problem | Whether approaching it works |
| Rule must fix | Example |
|---|---|
| Exactly what to ask | ‘In the last 30 days, did you or anyone…’ |
| Of whom | The most knowledgeable adult member |
| Over what period | 30 days, ending yesterday |
| In what units | Rupees; local units converted at entry |
| How it maps to the indicator | Any ‘yes’ → coded 1 |
| Change this | And you have changed |
|---|---|
| A word in the stem | The concept respondents answer about |
| The option set | What answers are possible at all |
| The order | What the previous question primed |
| The translation | The instrument, in that language |
| Branch | Error | Concrete example |
|---|---|---|
| Representation | Coverage | Frame is the ration-card list; the unlisted are invisible |
| Representation | Sampling | Chance variation in who was drawn |
| Representation | Non-response | Working men never at home in the day |
| Measurement | Validity | The item does not capture the concept |
| Measurement | Response | Desirability, recall, interviewer effect |
| Measurement | Processing | Coding and entry errors |
| Representation (wrong people) | Measurement (wrong answers) | |
|---|---|---|
| Coverage | Frame misses part of the population | — |
| Sampling | Chance variation in who is drawn | — |
| Non-response | Those who answer differ from those who don't | — |
| Validity | — | Question doesn't capture the concept |
| Response | — | Recall, social desirability, bad wording |
| Processing | — | Data entry & coding mistakes |
| Error | Fixable after fieldwork? |
|---|---|
| Sampling | Yes — reported as a margin of error |
| Coverage | No — the missing were never eligible to be drawn |
| Non-response | Partly — weights, if you know who is missing |
| Response bias | No — it is baked into the answers |
| Processing | Yes — if the raw export was kept |
| Extra rupee spent on | Reduces | At the cost of |
|---|---|---|
| A bigger sample | Sampling error | Training, follow-up, quality |
| Enumerator training | Response & processing error | Sample size |
| Revisits & refusal conversion | Non-response bias | Coverage of new areas |
| Pretesting | Validity & response error | Fieldwork days |
| Fault | Slide |
|---|---|
| Double-barrelled | 30 |
| Leading | 31 |
| Loaded or assuming | 32 |
| Jargon | 33 |
| Vague reference period | 34–35 |
| Test | Failing version | Passing version |
|---|---|---|
| Clear | ‘Do you avail ANC services?’ | ‘Did you go for check-ups when pregnant?’ |
| Specific | ‘Do you earn much?’ | ‘Last month, how much did you earn in cash?’ |
| One idea | ‘Clean and well-staffed?’ | Two separate questions |
| Hidden second idea | Split into |
|---|---|
| ‘Clean and well-staffed?’ | Clean? / Adequately staffed? |
| ‘Safe and affordable?’ | Safe? / Affordable? |
| ‘Do you save and invest?’ | Save? / Invest? |
| Leading device | Neutral repair |
|---|---|
| ‘Don’t you agree…’ | ‘Do you agree or disagree…’ |
| One-sided stem | Name both directions in the stem |
| Unbalanced options | Equal number either side of the midpoint |
| Prestige cue (‘experts say’) | Delete it |
| Question | Assumption smuggled in | Open it |
|---|---|---|
| ‘How much did the official demand?’ | That a demand was made | ‘Did you pay anything beyond the official fee?’ |
| ‘How many children do you still want?’ | That she wants more | ‘Do you want any more children?’ → filter |
| ‘When did you stop attending?’ | That she attended, and stopped | Two filtered questions |
| Insider term | What respondents say |
|---|---|
| institutional delivery | gave birth at the hospital / at home |
| ANC | check-ups when you were pregnant |
| livelihood diversification | other work besides farming |
| open defecation | going outside / to the field |
| Vague | Anchored | Why it's better |
|---|---|---|
| “Do you usually…?” | “In the last 7 days…?” | Defines 'usually' for everyone |
| “recently” | “in the last 30 days” | Same window for all respondents |
| “your income” | “income last month” | Fixes the unit of time |
| “often sick” | “ill in the past 2 weeks” | Countable, comparable |
| Period | Suits | Risk |
|---|---|---|
| 7 days | Frequent, small events (meals, wage work) | Seasonal events missed entirely |
| 30 days | Income, expenditure, health visits | Telescoping in from outside |
| 12 months | Rare events (birth, migration, shock) | Heavy forgetting |
| Since a local landmark | Recall-friendly anchoring | Date differs slightly by respondent |
| Recall failure | Direction of the error |
|---|---|
| Forgetting small routine events | Under-report |
| Telescoping (pulling events in) | Over-report |
| Rounding and heaping (0, 5, 10) | Distribution distorted, mean roughly intact |
| Reconstructing from a rule | Smoothed, less variance than reality |
| Fault | Where it is in the ‘before’ |
|---|---|
| Leading | ‘Don’t you think you should…’ |
| Double-barrelled | education and health |
| Vague | ‘usually’, ‘more’ |
| Loaded | presumes current spending is too low |
| Open-ended | Closed-ended | |
|---|---|---|
| Answer | In their own words | Pick from fixed options |
| Best for | Exploring, unknown answers | Counting, comparing |
| Analysis | Slow — needs coding | Fast — ready to tabulate |
| Risk | Vague, hard to compare | Misses the unlisted answer |
| Use when | Few cases, discovery | Most quantitative items |
| Use open when | Use closed when |
|---|---|
| You do not yet know the answer space | The options are known and stable |
| Piloting, to build the option list | Fielding at scale |
| The point is the respondent’s words | The point is a comparable count |
| Fault | Example | Fix |
|---|---|---|
| Overlapping | 0–18, 18–30, 30–50 | 0–17, 18–29, 30–49 |
| Not exhaustive | Married / Unmarried | + Widowed, Separated, Divorced |
| Unbalanced | Excellent–Good–Fair | Equal points either side |
| Missing escape | Fixed list only | + Other (specify) |
| Design choice | Convention |
|---|---|
| Points | 5 or 7, odd if a neutral is genuine |
| Labels | Label every point, not just the ends |
| Balance | Same number and intensity either side |
| Direction | Keep it constant within a block |
| Choice | Pro | Con |
|---|---|---|
| 5 points | Simple, fast, fits small screens | Less fine-grained |
| 7 points | More discrimination | Harder to read aloud |
| With midpoint | Allows a genuine neutral | Some hide there to avoid choosing |
| Forced (no midpoint) | Pushes a stance | Fabricates an opinion that isn't there |
| Question | Choose | Because |
|---|---|---|
| Read aloud by an enumerator? | 5 points | 7 is hard to hold in memory |
| Self-completed on screen? | 5 or 7 | Respondent can see the scale |
| Is neutral a real position? | Keep the midpoint | Forcing hides genuine indifference |
| Is neutral an escape here? | Consider forced choice + explicit DK | Separates indifference from avoidance |
| Rating | Ranking | |
|---|---|---|
| Cognitive load | Low | High — grows fast with list length |
| Reveals trade-offs | No — all items can be ‘important’ | Yes |
| List length | Can be long | Cap at 4–5 items |
| Analysis | Straightforward | Ordinal, awkward to average |
| Response | Means | Handle as |
|---|---|---|
| Don’t know | Genuinely lacks the information | Valid answer — report the share |
| Refused | Declines to say | Valid; often informative on sensitive items |
| Not applicable | Filtered out by design | Structural missing, not missing data |
| Blank | Enumerator skipped it | A data-quality problem |
| Filter | Routes to | Trap |
|---|---|---|
| Any land cultivated? | Farming module | ‘Cultivated’ undefined — leased in? homestead? |
| Ever been pregnant? | Maternal module | Asked of the wrong respondent |
| Any member migrated? | Migration module | ‘Member’ rule not trained |
| CAPI gives you | Paper cannot |
|---|---|
| Enforced routing | — skips are the enumerator’s job |
| Real-time range and consistency checks | — caught weeks later, if ever |
| Timestamps and GPS per interview | — |
| Same-day sync for back-checks | — |
| Stage | Purpose | Typical share of the interview |
|---|---|---|
| Introduction & consent | Explain, obtain consent | 2–3 minutes |
| Warm-up | Easy, relevant, builds rapport | Short |
| Core modules | The substance | The bulk |
| Sensitive items | Placed late, after trust | Short, careful |
| Close | Thanks, contact, next steps | 1 minute |
| Ordering rule | Why |
|---|---|
| Group by topic, signpost each module | Reduces context-switching cost |
| General before specific | Specific items narrow later general answers |
| Same scale within a block | Builds rhythm, cuts errors |
| Filter before the module it gates | Avoids asking then discarding |
| Order | Effect |
|---|---|
| Crime questions, then life satisfaction | Satisfaction reported lower |
| Specific items, then a general one | General answer narrows to what was just asked |
| Programme questions, then trust in government | Trust reflects the programme |
| Open with | Never open with |
|---|---|
| Household roster, simple facts | Income |
| An easy, obviously relevant item | Caste |
| Something the respondent can answer confidently | A long grid |
| Technique | What it does |
|---|---|
| Place late | Trades on rapport already built |
| Normalising preamble | ‘Many people find…’ lowers the threat |
| Self-completion / sealed response | Removes the interviewer from the answer |
| Explicit right to refuse | Makes refusal safe, so answers given are truer |
| Fatigue shows as | Detect it by |
|---|---|
| Straight-lining a grid | Zero variance across a block |
| Element | Rule |
|---|---|
| Variable names | Fixed, meaningful, in the instrument itself |
| Enumerator instructions | Visually distinct from what is read aloud |
| Skip instructions | At the branching question, in bold |
| CAPI screens | One idea per screen; no scrolling grids |
| Belief | Reality |
|---|---|
| ‘We need 10% of the population’ | Precision depends on n, not the share sampled |
| ‘Bigger is always better’ | A biased sample gets more confidently wrong |
| ‘We surveyed everyone available’ | That is a convenience sample, not a sample |
| Method | How | Use when |
|---|---|---|
| Simple random | Every unit equal chance | You have a full list |
| Systematic | Every k-th unit from a list | Ordered list, no hidden cycle |
| Stratified | Split into groups, sample each | Must represent subgroups |
| Cluster | Sample whole groups (villages) | People geographically spread |
| Multistage | Clusters, then units within | Large national surveys (NFHS) |
| Method | Needs | Fails when |
|---|---|---|
| Simple random | A complete list | No usable frame exists |
| Systematic | An ordered list | The list has a hidden cycle |
| Stratified | Strata known in advance | Strata variable is missing or wrong |
| Cluster / multi-stage | Village or PSU list only | Design effect ignored in analysis |
| PPS | Population sizes per unit | Sizes are stale (2011 Census) |
| Frame | Systematically omits |
|---|---|
| Ration-card list | The unlisted, recent migrants, the newly poor |
| SHG membership register | Non-members — often the poorest |
| Voter roll | Under-18s, the recently moved |
| Programme beneficiary list | Everyone the programme did not reach |
| Situation | Do |
|---|---|
| Small subgroup, need its own estimate | Over-sample the stratum; weight back |
| Comparing two blocks | Allocate roughly equally, not proportionally |
| Rare population | Screen, or use a specialised design |
| n | Margin (95%, p=0.5) |
|---|---|
| 384 · 1,067 | ±5% · ±3% |
| If you want… | You need… | Note |
|---|---|---|
| Tighter margin of error | A larger sample | ±3% needs ~3× the n of ±5% |
| Estimates for subgroups | More per subgroup | Each cell needs its own n |
| To detect a small change | More power | Effect size drives this |
| Cluster (not random) design | An inflation factor | The 'design effect' |
| You want | Cost |
|---|---|
| ±3% instead of ±5% | About 3× the sample |
| A separate estimate per district | A full sample per district |
| To detect a 3-point change, not 10 | Roughly 10× the sample |
| To sample clusters, not households | Inflate n by the design effect |
| Bias | A bigger n does |
|---|---|
| Coverage | Nothing — the missing were never in the frame |
| Non-response | Nothing — unless the refusers change |
| Selection | Nothing — it repeats the same skew |
| Sampling error | Reduces it — this one only |
| Weight component | Corrects for |
|---|---|
| Design weight | Unequal probability of selection |
| Non-response adjustment | Groups that answered less |
| Post-stratification | Drift from known population totals |
| Mode shapes | How |
|---|---|
| Who you reach | Phone ownership, internet access, being at home |
| How long they stay | Phone drops off sharply past 20 minutes |
| How honest they are | Less desirability bias without a human present |
| What you can ask | Show-cards, observation and long grids need presence |
| CAPI in practice | Note |
|---|---|
| KoboToolbox · ODK · SurveyCTO | Offline-capable; sync when a signal appears |
| Enforced routing | Removes the largest source of paper error |
| GPS & timestamps | Enable back-checks and duration monitoring |
| Device logistics | Charging, theft, breakage, data plans |
| Phone survey | Constraint |
|---|---|
| Length | 15–20 minutes is the practical ceiling |
| Who answers | Phone owner ≠ household; often the man |
| Who is missed | No phone, no network, no charge, no literacy for IVR |
| Response rates | Low; and refusers differ systematically |
| Web works for | Web fails for |
|---|---|
| Staff, students, professionals | General rural populations |
| Lists with verified emails/numbers | Open links shared onward |
| Sensitive self-completion | Anyone without a device or data |
| Mode | Cost | Coverage | Length | Best for |
|---|---|---|---|---|
| Face-to-face / CAPI | High | Broadest | Long | Rigorous, rural, complex |
| Phone / IVR | Low | Phone owners | Short | Speed, crises, monitoring |
| Web | Lowest | Connected only | Medium | Staff & literate audiences |
| Self-completion | Medium | Literate | Medium | Sensitive topics |
| Mode | Roughly, per completed interview | Typical response rate |
|---|---|---|
| Face-to-face / CAPI | Highest | Highest |
| Phone / IVR | Low | Low to moderate |
| Web | Lowest | Lowest |
| As cost falls | Coverage narrows to |
|---|---|
| Face-to-face → phone → web | Everyone → phone owners → the connected |
| Item type | Reported more honestly to |
|---|---|
| Stigmatised behaviour | A screen, not a person |
| Socially approved behaviour | Over-reported to a person |
| Complex recall | A person who can probe |
| Household detail | A person who can see the dwelling |
| Mixed-mode gains | Mixed-mode costs |
|---|---|
| Higher response | Blended, hard-to-separate biases |
| Lower average cost | Mode effects confounded with real differences |
| Reaches refusers | Complex weighting |
| Decision | Consequence |
|---|---|
| Which languages to field in | Who can answer at all |
| Who translates | Whether jargon returns |
| Whether translation is tested | Whether it measures the same thing |
| On-the-spot oral translation | Every enumerator becomes a translator |
| Equivalence type | Question to ask |
|---|---|
| Semantic | Do the words mean the same? |
| Conceptual | Does it evoke the same idea? |
| Difficulty | Is it equally easy to answer? |
| Connotation | Does it carry the same social charge? |
| Step | Who |
|---|---|
| Forward translation | Translator A, into the target language |
| Back-translation | Translator B, blind to the original |
| Comparison | The design team, against the source |
| Reconciliation | All three, resolving each discrepancy |
| Method | Catches |
|---|---|
| Back-translation | Literal errors and omissions |
| Committee review | Register, naturalness, cultural misfit |
| Field-staff review | What is actually sayable aloud |
| Cognitive pretest in the target language | How respondents really interpret it |
| Imported category | Local reality |
|---|---|
| ‘Employed / unemployed’ | Daily wage, seasonal, unpaid family work, several at once |
| ‘Nuclear / extended family’ | Joint households that split and re-form seasonally |
| Monthly income bands | Paid daily, weekly, in kind, or at harvest |
| ‘Head of household’ | Contested; often nominal rather than actual |
| Ask in | Convert to |
|---|---|
| Bigha, katha, guntha | Acres or hectares — with a district-specific factor |
| Seer, tin, bundle | Kilograms, at listed local weights |
| ‘Since kharif harvest’ | A dated reference period |
| Daily or piece-rate wage | A monthly equivalent, showing the assumption |
| Anchor | Fixes recall to |
|---|---|
| ‘Since Diwali’ | A shared, vivid date |
| ‘Since the kharif harvest’ | The agricultural cycle |
| ‘Since the school reopened’ | A local, verifiable event |
| Hold identical | Adapt freely |
|---|---|
| The concept | Wording and idiom |
| The indicator definition | Examples used |
| The reference period | Units, then converted |
| The response scale | Labels in the local language |
| Pretest | Pilot | |
|---|---|---|
| Asks | Do the questions work? | Does the operation work? |
| Scale | 8–20 respondents | A realistic mini-round |
| Method | Think-aloud, probing | Full protocol, end-to-end |
| Output | Rewritten questions | Fixed routing, timings, logistics |
| Probe | Reveals |
|---|---|
| ‘What did that question mean to you?’ | Comprehension |
| ‘How did you work out the answer?’ | Recall strategy and estimation |
| ‘Was any option hard to choose between?’ | Whether the scale discriminates |
| ‘Would that be awkward to answer honestly?’ | Sensitivity and desirability |
| Pilot check | Pass condition |
|---|---|
| Interview length | Within the planned ceiling for most respondents |
| Routing | Every branch reached by at least one case |
| Sync & GPS | Data arrives complete and located |
| Consent & refusal | Handled per protocol, recorded |
| Pilot data run through the analysis | Every dummy table populates |
| Training must cover | Or you get |
|---|---|
| Each question’s intent | Enumerators inventing interpretations |
| Neutral probing | Leading, and answers that match expectations |
| Skips and edge cases | Silent routing errors |
| Consent and refusal | Ethics failures and coerced participation |
| Practice interviews, observed | Errors discovered in the field |
| Interviewer attribute | Where it matters most |
|---|---|
| Gender | Violence, reproductive health, autonomy |
| Caste / community | Discrimination, access to services |
| Age | Deference and youth topics |
| Manner and dress | Whether the respondent reads you as an official |
| Direction | Typical items |
|---|---|
| Over-reported | Hand-washing, toilet use, school attendance, voting |
| Under-reported | Alcohol, tobacco, violence, income, caste practice |
| Check | Sample | Catches |
|---|---|---|
| Back-check call/visit | 5–10% of interviews | Fabrication; key answers not matching |
| Spot-check (live observation) | A few per enumerator | Protocol drift, leading, skipped items |
| High-frequency data checks | All records, daily | Duration outliers, heaping, DK spikes |
| Record | Why it matters |
|---|---|
| Refusal | May differ systematically from responders |
| Not at home after 3 visits | Working households under-represented |
| Ineligible | Affects the frame, not the response rate |
| Partial completion | Item non-response, treated separately |
| This section covers | Slide |
|---|---|
| Why ethics is respect made operational | 92 |
| Consent as a floor, not a formality | 93 |
| What the DPDP Act, 2023 asks of you | 94 |
| De-identification and small-cell risk | 95 |
| Clean data, documentation and tools | 96–97 |
| Respondent gives | You owe them |
|---|---|
| 30–60 minutes of time | An instrument with nothing spare in it |
| Personal information | Storage, limits, deletion |
| Trust in a stranger | Honesty about what happens next |
| Nothing in return, usually | At minimum, findings shared back |
| Consent must state | Commonly skipped |
|---|---|
| What is collected and why | Rushed into one sentence |
| How it is stored, shared, for how long | Omitted entirely |
| That refusal carries no penalty | Implied but not said |
| Who to contact afterwards | No contact given |
| DPDP duty | In a CAPI survey |
|---|---|
| Purpose limitation | Collect only what the protocol names |
| Notice & consent | The consent script, in the respondent’s language |
| Security safeguards | Device encryption, access control, secure sync |
| Retention limits | A deletion date, and someone responsible for it |
| Rights of the data principal | A route to correction and erasure |
| Identifier | Risk |
|---|---|
| Name, phone, exact GPS | Direct — remove from the analysis file |
| Village + age + caste + occupation | Indirect — can single out one person |
| Small-cell tables | A cell of 1 or 2 is an identification |
| Free-text answers | Often contain names and places |
| Practice | Payoff |
|---|---|
| Validation in the CAPI form | Errors caught at the doorstep |
| Raw export kept untouched | Every cleaning step is reproducible |
| Cleaning done in a script, not by hand | You can say what you changed and why |
| Data dictionary shipped with the data | The file survives its author |
| Tool | Good for | Note |
|---|---|---|
| KoboToolbox | Humanitarian & NGO CAPI | Free, offline, widely used |
| ODK (Open Data Kit) | Open-source mobile collection | Free, flexible, technical |
| SurveyCTO | Rigorous research surveys | Paid; strong quality controls |
| Survey Solutions | Large official surveys | Free (World Bank) |
| Tool | Best fit | Watch |
|---|---|---|
| KoboToolbox | NGO and humanitarian CAPI | Free, offline, large user base |
| ODK | Open-source, self-hosted | Flexible; needs technical capacity |
| SurveyCTO | Rigorous research at scale | Paid; strong quality-control features |
| CSPro | Census-style large operations | Heavier, statistical-office lineage |
| Read | For |
|---|---|
| Groves et al., Survey Methodology | Total Survey Error, in full |
| Fowler, Improving Survey Questions | Question wording, practically |
| Bradburn, Sudman & Wansink, Asking Questions | Worked examples across topic types |
| NFHS · PLFS instruments and reports | How this is done at national scale |
| Takeaway | What it rules out |
|---|---|
| Objectives before questions | Drafting an instrument first |
| One question, one idea | ‘Clean and well-staffed?’ |
| The questionnaire is the measurement | Treating wording as presentation |
| Manage total error | Reporting only the sampling margin |
| Pretest and pilot | Trusting a first draft |
| The respondent is a person | An instrument with spare questions in it |