fullscreen
ImpactMojoBivariate Analysis 101www.impactmojo.in
ImpactMojo 101 Series · Free Forever
Bivariate
Analysis
101
How Two Variables Move Together — Cross-Tabs, Correlation, Group Comparisons & Simple Regression for Development Practitioners in South Asia
Research-BackedSouth Asia Focus~90 SlidesFree Access
ImpactMojoBivariate Analysis 101www.impactmojo.in
What We Cover
01
What Bivariate Analysis Is
Slides 3–10
02
Choosing a Method by Variable Type
Slides 11–17
03
Cross-Tabulation & Contingency Tables
Slides 18–25
04
Visualising Relationships
Slides 26–34
05
Correlation
Slides 35–44
06
Comparing Two Groups
Slides 45–53
07
Comparing Several Groups (ANOVA)
Slides 54–61
08
Association Between Categories (Chi-Square)
Slides 62–71
09
Simple Linear Regression
Slides 72–81
10
Pitfalls
Slides 82–91
11
Reporting, Practice & Tools
Slides 92–99
ImpactMojoBivariate Analysis 101www.impactmojo.in
01
Section One
What Bivariate Analysis Is
ImpactMojoBivariate Analysis 101www.impactmojo.in
From one variable to two
Univariate analysis describes one variable at a time — the average household size, the spread of incomes. Bivariate analysis is the very next step: it asks how two variables relate to each other.
Bivariate analysis
The study of the relationship between two variables — whether they move together, how strongly, and in what direction. 'Bi' = two; 'variate' = variable.
Most interesting development questions are bivariate at heart: does this go with that? Does schooling go with earnings? Does the programme go with better outcomes?
Univariate asksBivariate asks
What share of children are underweight?Does it differ by wealth quintile?
What is median household income?Does it differ between migrant and non-migrant households?
How many women completed secondary school?Is completion linked to age at marriage?
Bivariate analysis is where programme questions actually live. Nobody funds a programme because 18 per cent of children are underweight; they fund it because the 18 per cent is concentrated somewhere, and finding where is a two-variable question.
It is also where overclaiming begins. A univariate statistic is hard to misread; a bivariate one invites a causal sentence, and Section 9 exists because that sentence is usually written before the analysis supports it.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Where it sits in the analysis ladder
01
UNIVARIATE: describe one variable (centre, spread, shape)
02
BIVARIATE: relate two variables (direction, strength)
03
MULTIVARIATE: many variables at once (control, adjust)
This course lives squarely in the middle rung. Master it before reaching for regression with twenty controls — the two-variable picture is where most reasoning errors are caught.
LevelAnswersTypical tool
UnivariateWhat is typical, how spread outMean, median, histogram
BivariateDo these two move togetherCross-tab, scatter, r, t-test
MultivariateDoes it hold once other things are held constantMultiple regression
The third row is where a bivariate finding usually goes to die. Almost every two-variable relationship in development data shrinks once wealth or education is held constant, which is not a failure of the bivariate method — it is the method doing its job as a first look.
Do not skip the first row. Analysts who jump straight to relationships miss the impossible values, the spike at zero and the missingness that would have changed everything downstream.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Direction and strength
Direction
Do they move the same way (positive) or opposite ways (negative)? More literacy, fewer births — that is a negative direction.
Strength
How tightly do they track each other? A loose cloud is weak; a near-straight line is strong.
Almost every bivariate tool you will meet is just a precise way to answer these two questions — plus a third: could this be chance?
QuestionWhat answers itWhat does not
Which way does it go?The sign of r; the direction of the differenceA p-value
How strong is the pattern?The size of r; how far apart the group means areSignificance alone
How big is the effect?The slope, in real unitsr, which is unitless
Could it be chance?The p-value or the confidence intervalThe size of r
These four questions are distinct and are constantly conflated. Strength, size and significance answer different things, and a report that supplies only one of them has left the reader unable to judge the other two.
Report direction, magnitude and uncertainty together. “Three percentage points lower (95% CI 1 to 5)” answers all four rows in one clause; “significantly lower” answers one.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Explanatory and response variables
Explanatory variable (X)
The variable you think does the explaining or predicting — also called the independent or predictor variable. Conventionally on the horizontal axis.
Response variable (Y)
The outcome you want to understand or predict — also called the dependent variable. Conventionally on the vertical axis.
Naming X and Y does not prove X causes Y. It only states which one you are treating as the outcome. The causal claim must be earned separately.
Also calledRole
Explanatory, independent, predictor, XThe thing you think does the explaining
Response, dependent, outcome, YThe thing you are trying to account for
The labels are a decision you make, not a property of the data. Nothing in a dataset says which variable is the cause; calling one “independent” is an assumption you are importing, and the arithmetic will run identically if you swap them.
Which is why the naming deserves a moment’s thought. Regressing income on toilet ownership and toilet ownership on income both produce a line; only one of them matches a story anyone believes, and the software has no opinion.
ImpactMojoBivariate Analysis 101www.impactmojo.in
The questions practitioners actually ask
  • Do villages with self-help groups have higher women's savings?
  • Is anaemia more common among Adivasi women than others?
  • Does distance to a health centre predict institutional delivery?
  • Did test scores differ between the treated and control schools?
  • Is caste associated with whether a household has a toilet?
Each is a relationship between two variables — and each maps to a specific bivariate method, which Section Two helps you choose.
Programme questionVariable pairMethod
Do poorer households use the service less?Wealth quintile × used/notCross-tab, chi-square
Do trained farmers get higher yields?Trained/not × yieldDifference in means, t-test
Does distance predict attendance?Km × days attendedScatter, correlation, regression
Do outcomes differ across five blocks?Block × scoreANOVA, then pairwise
Every one of these is a real question a programme asks in its first year, and each has exactly one appropriate first method — determined by the variable types, not by preference. That mapping is Section 2, and it removes most of the anxiety about “which test”.
Note what none of them establishes. Trained farmers differ from untrained farmers in more than training, so the second row measures a difference between two self-selected groups — useful, and not an impact estimate.
ImpactMojoBivariate Analysis 101www.impactmojo.in
A relationship is a clue, not a verdict
Bivariate analysis can reveal a relationship, quantify its strength, and tell you whether it is likely real or just noise. What it cannot do, on its own, is prove that one variable causes the other.
The plural of anecdote is not data; and the presence of correlation is not the presence of cause.
— a working principle of careful analysis
Bivariate analysis can showIt cannot show
That two things move togetherThat one causes the other
Where to look nextWhat to conclude
That a difference is bigger than noiseThat the difference is due to your programme
A pattern worth explainingThe explanation
A bivariate relationship is a clue. In observational development data it almost always has at least one plausible confounder, and often several — wealth, education and location are correlated with nearly everything else we measure.
That does not make it worthless. Clues direct effort: they tell you which blocks to visit, which subgroup to investigate and which hypothesis is worth the cost of testing properly. The error is treating the clue as the verdict.
ImpactMojoBivariate Analysis 101www.impactmojo.in
How this course is built
Tools
  • Cross-tabs for category × category
  • Correlation for number × number
  • t-tests and ANOVA for comparing groups
  • Chi-square and simple regression
Judgement
  • Choosing the right method
  • Reading effect size, not just p-values
  • Spotting confounders and fallacies
  • Reporting a relationship honestly
Examples come from India and the wider region — the data you actually meet at work.
SectionMethod it gives you
2–3 · Choosing and cross-tabsPick a method from variable types; read a contingency table
4 · VisualisingScatter, box plot, grouped bars — and when each lies
5 · CorrelationPearson, Spearman, and what r is not
6–7 · Comparing groupst-test, ANOVA, effect size
8–9 · Chi-square, regressionAssociation between categories; a line and its slope
10–11 · Pitfalls and reportingWhat to check, and how to write it honestly
Sections 10 and 11 are the ones that change practice. The methods are standard and available in any software; what separates a defensible analysis from a misleading one is the checking and the wording, which is where most training stops.
ImpactMojoBivariate Analysis 101www.impactmojo.in
02
Section Two
Choosing a Method by Variable Type
ImpactMojoBivariate Analysis 101www.impactmojo.in
First, classify both variables
The single most useful habit in bivariate analysis: before choosing any method, ask what kind of variable each of your two is — categorical or numeric. The pair of answers points straight to the right tool.
Categorical
Labels or groups — sex, caste, religion, district, yes/no. Includes ordered categories (wealth quintile, Likert scale).
Numeric
Counts and measurements — age, income, test score, fertility rate, distance in km. You can meaningfully average them.
VariableTypeNote
District, caste, sexCategorical (nominal)No order
Wealth quintile, Likert scaleCategorical (ordinal)Ordered, unequal gaps
Age, income, yieldNumeric (continuous)Ratios meaningful
Number of childrenNumeric (count)Discrete; often skewed
Used the service (yes/no)Categorical (binary)The commonest outcome in our work
Classify both variables first and the method chooses itself. Nearly all confusion about “which test” is confusion about variable types, and it is resolved by looking at the codebook rather than by reading more about tests.
Ordinal is the awkward middle. Wealth quintile is ordered but its gaps are not equal, so treating it as numeric assumes something the variable does not carry. Treating it as categorical is always defensible; treating it as numeric sometimes is.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Method by the two variable types
X typeY typeDescribe withTest with
CategoricalCategoricalCross-tab, % & barsChi-square test
Categorical (2 groups)NumericGroup means, box plotst-test
Categorical (3+ groups)NumericGroup means, box plotsOne-way ANOVA
NumericNumericScatter plotCorrelation / regression
Pin this table to your wall. Nine times out of ten, classifying your two variables tells you exactly which method to reach for.
XYStart with
CategoricalCategoricalCross-tab with row %; chi-square
Categorical (2 groups)NumericBox plots; difference in means; t-test
Categorical (3+ groups)NumericBox plots; ANOVA, then pairwise
NumericNumericScatter; correlation; simple regression
NumericBinaryBeyond this deck — logistic regression
Every row begins with a picture, not a test. The plot is not a courtesy to the reader; it is how you discover the curve, the outlier or the two clusters that would make the test meaningless.
The last row is worth knowing exists. A yes/no outcome with a numeric predictor is extremely common in our work — did she deliver in a facility, given distance — and linear regression handles it badly. That is a signal to seek help, not to proceed.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Two categories: is membership linked?
When both variables are categories — say caste group and has a household toilet (yes/no) — you ask whether knowing one tells you anything about the other.
Describe it with a cross-tabulation and percentages; test it with a chi-square test of independence. Covered in Sections Three and Eight.
QuestionCross-tab of
Do SC/ST households use the scheme less?Social group × used/not
Is dropout linked to sex?Sex × enrolled/dropped
Does facility delivery vary by wealth?Wealth quintile × place of delivery
Is complaint type linked to block?Block × complaint category
Cross-tabulation is the most used and least respected tool in this list. Most equity findings a programme will ever produce are a two-way table with row percentages, and no more sophisticated method is required to see them.
Always print the counts alongside the percentages. A row showing 67 per cent uptake is a different claim when the row holds 300 households and when it holds three.
ImpactMojoBivariate Analysis 101www.impactmojo.in
A group label and a number: compare means
When one variable is a group label and the other is numeric — treatment vs control and test score — you compare the numeric outcome across the groups.
Two groups
Difference in means — the t-test idea. Section Six.
Three or more groups
One-way ANOVA compares all the group means at once. Section Seven.
ReportWhy
The mean or median for each groupThe quantity anyone actually wants
N per groupSmall groups produce unstable means
Spread within each groupTwo groups can differ in variability, not level
The difference, in units“3.2 kg lower” beats “significantly lower”
Use the median when the numeric variable is skewed, which for income, expenditure and landholding it always is. The comparison of medians is less standard and more honest, and a box plot shows both.
The third row is under-reported and often the finding. A programme that raises average yield while widening the spread has helped some farmers and left others where they were — visible in the spread, invisible in the mean.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Two numbers: do they track each other?
When both variables are numeric — female literacy and fertility rate across districts — plot one against the other on a scatter and measure how tightly they move together.
Describe the strength with correlation (Section Five); model the line with simple linear regression (Section Nine).
StepWhat you are looking for
Plot the scatterShape: line, curve, clusters, nothing
Look for outliersPoints far from the cloud, which will dominate any fit
Compute rOnly if the shape is roughly linear
Fit a line, if the question needs a rateSlope in real units, with its uncertainty
The order matters and is routinely reversed. Computing r first and plotting afterwards means the number frames how you read the picture; plotting first means the picture tells you whether the number is meaningful at all.
ImpactMojoBivariate Analysis 101www.impactmojo.in
The method serves the question, not the reverse
Do not pick a fancy test and then hunt for variables to feed it. Start from the real question, classify the two variables it involves, and let the decision table hand you the method.
Far better an approximate answer to the right question than an exact answer to the wrong one.
— John Tukey
Wrong orderRight order
“Let us run a regression”“What are we trying to find out?”
Method chosen for familiarityMethod chosen by variable types
Recode a variable so the test runsChoose the test that fits the variable
Report whatever the software printedReport what answers the question
The third row is the most damaging habit. Collapsing a five-category variable into two so a familiar test applies throws away information and usually flatters the result, because the collapse point is chosen after seeing the data.
Write the question in a sentence before opening the software. It takes a minute, it determines the method, and it is the sentence the report will need anyway.
ImpactMojoBivariate Analysis 101www.impactmojo.in
03
Section Three
Cross-Tabulation & Contingency Tables
ImpactMojoBivariate Analysis 101www.impactmojo.in
Counting two categories together
Cross-tabulation (contingency table)
A table that counts how many cases fall into each combination of two categorical variables — rows for one variable, columns for the other.
It is the most-used tool in applied development analysis: a single table that shows, at a glance, how two group memberships line up.
A contingency table showsReport alongside
Counts in each combination of categoriesRow or column percentages
Row and column totals (marginals)The total N
Where cells are unexpectedly full or emptyAny cell with fewer than five cases
The marginals are the part people skip and the part that carries the context. Without the totals, a reader cannot tell whether an imbalanced table reflects a relationship or simply that one category is much larger than the other.
Small cells are the standard trap. Once any cell drops below about five expected cases, chi-square becomes unreliable — and the table will still print a percentage that looks like all the others.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Toilet ownership by location (raw counts)
Has toiletNo toiletTotal
Rural320280600
Urban34060400
Total6603401,000
Illustrative figures. The four inner cells are the joint counts; the right and bottom margins are the marginal totals for each variable on its own.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Raw counts mislead when groups differ in size
Rural and urban have different totals (600 vs 400), so comparing raw counts is unfair. To compare fairly, convert to percentages — but in which direction?
This is the single most common cross-tab error: percentaging the wrong way and reading a relationship backwards. Get the direction right and the table tells the truth.
Raw count comparisonWhy it misleads
“More non-users are rural”Most of the population is rural
“Most dropouts are boys”There may be more boys enrolled
“This block has the most cases”It also has the most people
Percentages are how a table becomes comparable when the groups differ in size — which they always do. A count table answers “how many”; only a percentage table answers “is membership linked”.
Keep both in the published table. Percentages for the comparison, counts so the reader can see the base — the standard format is “62.4% (n=187)” and it answers both questions in one cell.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Percent within each row
Has toiletNo toiletRow total
Rural53%47%100%
Urban85%15%100%
Each row sums to 100%. This answers: of rural households, what share have a toilet? 53% rural vs 85% urban — a clear location gap. (Illustrative.)
ImpactMojoBivariate Analysis 101www.impactmojo.in
Percent within each column
Has toiletNo toilet
Rural48%82%
Urban52%18%
Column total100%100%
Each column sums to 100%. This answers a different question: of households without a toilet, what share are rural? 82%. Same table, different story.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Percentage in the direction of the cause
Convention: percentage within categories of the explanatory variable, then compare across them. If location (X) might shape toilet ownership (Y), percentage within rural and within urban — that is row percentaging here.
01
Put X in the rows
02
Percentage so each row = 100%
03
Compare the same Y-column across rows
If your question isPercentage within
“Do rural households use it less than urban?”Location — so each location sums to 100
“Among users, how many are rural?”Use status — so users sum to 100
“Does uptake differ by wealth?”Wealth quintile
Percentage in the direction of the presumed cause, so that each category of the explanatory variable sums to 100 and the comparison across those categories is the finding. Getting this backwards produces a table that answers a question nobody asked.
The two directions answer genuinely different questions, and both can be true and interesting. State which one you have used in the table heading; readers assume whichever suits their expectation otherwise.
ImpactMojoBivariate Analysis 101www.impactmojo.in
What a difference in percentages means
85% − 53%
= 32 percentage-point gap in toilet ownership, urban vs rural
Direction
Urban households far more likely to have a toilet — a positive urban–toilet link
A gap in the table suggests an association. Whether it is bigger than chance is what the chi-square test in Section Eight decides.
Way to express a 2×2 differenceExample
Percentage-point difference62% vs 48% — a gap of 14 points
Ratio of proportions62/48 = 1.3 times as likely
Relative change29% higher
All three describe the same table and land very differently on a reader. Reporting only the largest-sounding one is the commonest way an honest table becomes a misleading sentence.
Give the percentage-point gap and the two underlying percentages. From those a reader can compute any of the others, and cannot be led by the choice of framing.
ImpactMojoBivariate Analysis 101www.impactmojo.in
04
Section Four
Visualising Relationships
ImpactMojoBivariate Analysis 101www.impactmojo.in
Always look before you compute
Before any coefficient or test, draw the relationship. A picture reveals direction, strength, curvature, clusters and outliers that a single number can hide entirely.
The greatest value of a picture is when it forces us to notice what we never expected to see.
— John Tukey
The plot revealsThe statistic hides it
A curve rather than a liner near zero suggests “no relationship”
One extreme point driving everythingr looks strong
Two distinct clustersA line is fitted through empty space
A ceiling or floor in the dataThe fit is distorted at the ends
Every row above produces a plausible-looking number. The software does not warn you, and none of these conditions is rare in field data — ceilings and clusters are the normal shape of programme datasets.
Plot first, every time, even when you are confident. It costs one line of code and it is the only step in this deck that catches all four failures at once.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Which chart for which pair
Variable pairBest plotShows
Numeric × numericScatter plotDirection, strength, shape
Categorical × numericBox plot by groupSpread & median per group
Categorical × numericGrouped / clustered barsMean per group
Categorical × categoricalStacked / grouped barsShares within groups
PairPlotAvoid
Numeric × numericScatterTwo lines on one time axis
Categorical × numericBox plot; strip plot for small NA bar of means with no spread
Categorical × categoricalGrouped or stacked barsPies, especially several
Ordered category × numericBox plots in orderAlphabetical ordering
The second row’s “avoid” is the most common chart in this sector. A bar chart of group means shows the centres and conceals everything else — the spread, the N, the outliers — and it is what most reporting templates default to.
With fewer than about twenty points per group, show the points. A box plot summarising six observations implies more than it knows; the raw dots are both more honest and easier to read.
ImpactMojoBivariate Analysis 101www.impactmojo.in
The workhorse for two numbers
A scatter plot puts the explanatory variable on the X-axis, the response on the Y-axis, and one dot per case. The cloud's tilt shows direction; its tightness shows strength.
Read it like this: upward cloud = positive; downward = negative; round blob = no linear relationship; tight line = strong; fat cloud = weak.
When reading a scatter, askBecause
Is the pattern linear?Everything that follows assumes it
How tight is the cloud?That is strength, and it is visible
Are there points far from the rest?They will dominate any fitted line
Does the spread change along X?Fanning out breaks a standard assumption
What is the unit — person, village, district?Decides what the pattern can be about
The last row is the one that becomes the ecological fallacy in Section 10. A scatter of state averages is a statement about states; nothing in the picture prevents a reader from taking it as a statement about people, and most do.
With large N, overplotting hides density. Thousands of points become a solid mass; transparency, jitter or a hexbin plot restores the information that the raw scatter has already lost.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Three scatters: positive, negative, none
Same axes, three relationships
Illustrative
Green rises, red falls, indigo wanders. Your eye reads direction and strength instantly — before any number.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Female literacy vs child mortality, by state
Female literacy (%) vs under-5 mortality (per 1,000), major states
Illustrative, patterned on Census 2011 & NFHS-5
A strong negative pattern: states with higher female literacy tend to have lower child mortality. Illustrative, but it mirrors the real Census–NFHS picture closely.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Comparing a number across groups
For a categorical X and numeric Y, a box plot per group beats a single average. Each box shows the median (the line), the middle 50% (the box), the range (the whiskers) and outliers (the dots).
Box plots let you see not just whether the centres differ between groups, but whether the spreads do — two distributions with the same mean can look completely different.
The box plot showsWhat that tells you
The median lineThe typical value — robust to skew
The box (Q1 to Q3)Where the middle half sits
The whiskersThe bulk of the range, by a stated rule
Points beyond themCandidates for investigation, not deletion
Box width, if drawn to scaleRelative group size
A box plot shows five numbers where a bar chart shows one, in the same space, and it is the single best default for comparing a numeric outcome across groups. Overlapping boxes are also a visual signal that a difference in means may not survive testing.
State the whisker rule in the caption. Different software uses different conventions — 1.5×IQR is common but not universal — and an unlabelled box plot is ambiguous about which points count as outliers.
ImpactMojoBivariate Analysis 101www.impactmojo.in
A mean compared across groups
Mean monthly earnings by education level (illustrative, ₹000)
Illustrative, patterned on PLFS-style data
A clean grouped bar shows the mean of a numeric outcome rising across ordered categories — the visual form of a category × number relationship.
ImpactMojoBivariate Analysis 101www.impactmojo.in
The same rules still apply
  • Start bar axes at zero — a truncated axis fakes a gap
  • Label both axes with units
  • Note the denominator and sample size
  • Put the source and date on the chart
  • Do not let one outlier dominate a scatter unremarked
A bivariate chart is an argument about a relationship. Truncated axes and missing denominators turn an honest pattern into a misleading one.
RuleBivariate-specific version
Axis from zero, or say why notTruncating a Y-axis exaggerates a group gap
Label units“Score” on what scale, out of what?
Show NPer group, not just overall
Order categories meaningfullyBy value or by natural order, never alphabetically
Do not imply causation in the title“Yields by training status”, not “Training raises yields”
The last row is where most causal overclaiming actually enters a report. Nobody writes “training causes yields” in the analysis section; the chart title says it instead, and the chart is what gets screenshotted into the presentation.
Titles are the cheapest place to be disciplined. A neutral title costs nothing and removes the claim you cannot support, while leaving the finding entirely visible.
ImpactMojoBivariate Analysis 101www.impactmojo.in
05
Section Five
Correlation
ImpactMojoBivariate Analysis 101www.impactmojo.in
Putting a value on 'move together'
Correlation coefficient (r)
A single number summarising how strongly and in which direction two numeric variables move together in a straight-line sense. It ranges from −1 to +1.
Where a scatter shows the relationship, r measures it — one number for direction and strength combined.
r tells your does not tell you
Direction of a linear associationWhich variable causes which
How tightly points follow a lineHow steep that line is
A property of this sampleAnything about an individual case
Nothing about unitsThe size of the effect in real terms
The second row is the distinction most worth holding. Correlation measures scatter around a line, not the line’s gradient — so a very strong correlation can accompany a practically negligible effect, and a modest one can accompany a large effect with noise around it.
When the question is “how much”, you want a slope, not r. “Each additional kilometre is associated with 0.4 fewer days attended” is usable by a programme; r = −0.62 is not.
ImpactMojoBivariate Analysis 101www.impactmojo.in
r runs from −1 to +1
−1
Perfect negative — points fall exactly on a downward line
0
No linear relationship — a shapeless cloud
+1
Perfect positive — points fall exactly on an upward line
The sign gives direction; the distance from zero gives strength. r = −0.8 and r = +0.8 are equally strong, opposite in direction.
ImpactMojoBivariate Analysis 101www.impactmojo.in
The default: Pearson correlation
The everyday correlation is Pearson's r. Crucially, it measures only the linear (straight-line) component of a relationship between two numeric variables.
If the true relationship is curved — rising then falling — Pearson's r can be near zero even when the two variables are tightly related. Pearson sees lines, not curves.
Pearson’s r assumesBroken when
Both variables are numericOne is an ordinal scale
The relationship is linearThe pattern curves or plateaus
No extreme outliersOne district is far from the rest
Roughly symmetric distributionsIncome, landholding, firm size
Three of the four are broken routinely in development data, which is why Spearman’s rank correlation is often the better default here. It is not a weaker method; it is a method that assumes less.
Compute both and compare. If Pearson and Spearman disagree sharply, something in the four rows above is happening — usually an outlier or a curve — and the disagreement is the diagnostic.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Interpreting the size of r
|r|Rough strengthWhat the scatter looks like
0.0 – 0.1NegligibleShapeless cloud
0.1 – 0.3WeakFaint tilt
0.3 – 0.5ModerateClear tilt, wide scatter
0.5 – 0.7StrongTight tilt
0.7 – 1.0Very strongNear a straight line
These bands are conventions, not laws. In messy social data, an r of 0.3 can be a genuinely important signal.
|r|LooselyCaveat
0.0–0.2Negligible to weakThese bands are conventions, not facts. What counts as strong depends entirely on the field: an r of 0.3 is large in social data and trivial in a physical measurement
0.2–0.4Weak to moderate
0.4–0.7Moderate to strong
0.7–1.0Strong to very strong
Be suspicious of a very high r in survey data. Above about 0.9 between two supposedly distinct social variables, the usual explanation is that they measure the same thing — or that one was derived from the other during cleaning.
Report r with N and the scatter. The band label is the least informative part; a reader who sees the plot and the sample size can judge the strength without any adjective at all.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Literacy and fertility, quantified
Female literacy (%) vs total fertility rate, major states
Illustrative, patterned on Census 2011 & NFHS-5
This cloud has a clear downward tilt — an r near −0.9 (illustrative). Strong and negative. But strength is not the same as cause: literacy may proxy for income, health and much else.
ImpactMojoBivariate Analysis 101www.impactmojo.in
When to use rank correlation instead
Spearman's rank correlation (ρ)
Pearson's correlation computed on the ranks of the data rather than the raw values. It measures whether two variables move together in the same order, not strictly in a straight line.
  • Use it for ordinal data — wealth quintiles, Likert scales
  • Use it for skewed data — income, landholding
  • Use it when the relationship is monotonic but curved
  • It is robust to outliers that would swing Pearson's r
Use Spearman whenBecause
A variable is ordinalRanks are all the variable actually carries
The data is heavily skewedRanking removes the influence of the long tail
There are extreme outliersAn outlier becomes just the top rank
The relationship is monotonic but curvedSpearman detects consistent direction without linearity
Spearman is Pearson computed on the ranks, which is why it inherits the interpretation and drops the linearity requirement. It answers “do they move in the same direction consistently” rather than “do they follow a straight line”.
The cost is information. Ranking discards magnitude, so a Spearman correlation cannot tell you that the gap between the first and second districts is enormous while the rest are close together.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Two correlations, side by side
Pearson rSpearman ρ
MeasuresLinear associationMonotonic (rank) association
Best dataNumeric, roughly symmetricOrdinal or skewed numeric
Outlier-sensitive?Yes — one point can swing itNo — uses ranks
Range−1 to +1−1 to +1
When Pearson and Spearman disagree sharply, suspect non-linearity or an outlier — and go back to the scatter.
If they agreeIf they differ substantially
The relationship is roughly linear and outlier-freeLook for an outlier, a curve, or heavy skew
Report either; say whichReport Spearman, and explain why
Nothing further neededThe disagreement is itself worth a sentence
Running both is a two-line diagnostic that catches most of the ways a correlation misleads, and it costs nothing. Where they diverge, the plot will show you why within seconds.
Never choose between them after seeing which is larger. That is the forking-paths problem in miniature: decide by the variable types and the shape of the data, and state the rule you used.
ImpactMojoBivariate Analysis 101www.impactmojo.in
A big r and a small p are different things
Strength (effect size)
How big is the relationship? That is what r itself tells you — the practical magnitude.
Significance (p-value)
How sure are we it is not zero by chance? That depends heavily on sample size.
With thousands of cases, a tiny, uninteresting r = 0.05 can be 'statistically significant'. Always report the size, not just the star.
SituationrpReading
N = 20,0000.03TinyReal, and far too small to matter
N = 250.55Not significantPossibly important; sample too small to tell
N = 4000.42SmallThe useful case
Strength and significance answer different questions, and sample size drives the second far more than the first. With a large enough N almost any correlation becomes significant, which makes “significant correlation” a nearly empty phrase in big datasets.
The second row is the one that gets discarded wrongly. A non-significant result in a small sample is not evidence of no relationship; it is an absence of evidence, and reporting it as “no association” is a real error.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Square it for 'variance explained'
Square the correlation and you get — the share of the variation in one variable that is statistically accounted for by the other. r = 0.7 means R² = 0.49: about half the variation is shared.
We return to R² properly in the regression section — it is the natural bridge from correlation to a fitted line.
rShared variation
0.30.099%
0.50.2525%
0.70.4949%
0.90.8181%
Squaring is deflating, which is the point of doing it. A correlation of 0.5 sounds like half the story and accounts for a quarter of the variation — a useful corrective when a moderate r is being described as a strong finding.
“Variance explained” is a statistical phrase, not a causal one. It means the line accounts for that share of the scatter; it does not mean the X variable produces that share of the outcome, and the wording invites exactly that reading.
ImpactMojoBivariate Analysis 101www.impactmojo.in
06
Section Six
Comparing Two Groups
ImpactMojoBivariate Analysis 101www.impactmojo.in
One group label, one number
A categorical X with two values (treatment/control, girls/boys, SHG/non-SHG) and a numeric Y (score, income, weight). The question: do the two group means differ — and is the difference real?
01
Split the data into two groups by X
02
Compute each group's mean of Y
03
Ask: is the gap bigger than chance?
Before comparing two groups, checkWhy
How the groups were formedSelf-selected groups differ in more than the label
N in eachAn unbalanced comparison is dominated by the small group’s noise
The distribution in eachSkew or bimodality makes a mean the wrong summary
Whether observations are independentHouseholds within a village are not
The first row decides what the comparison can mean. Randomly assigned groups support a causal reading; groups that chose themselves — trained farmers, scheme participants, migrants — differ in motivation, resources and information as well as in treatment.
The last row is quietly violated in almost all field data. Cluster sampling makes households within a village correlated, which makes a standard t-test overstate precision — the p-value is smaller than it should be.
ImpactMojoBivariate Analysis 101www.impactmojo.in
The quantity of interest
Mean test score: control vs treatment schools (illustrative)
Illustrative
The raw difference is 7 points (61 − 54). But two questions remain: could a 7-point gap arise by chance, and is 7 points big enough to matter?
ImpactMojoBivariate Analysis 101www.impactmojo.in
Signal divided by noise
t-test
A test of whether the difference between two group means is larger than we would expect from sampling variation alone — essentially, the size of the gap relative to how noisy the data are.
Intuitively, t ≈ difference in means ÷ uncertainty in that difference. A big, clean gap gives a big t; a small gap drowned in scatter gives a small t.
ComponentEffect on the t statistic
Larger difference between meansLarger t — more signal
Larger spread within groupsSmaller t — more noise
Larger sampleLarger t — the noise estimate tightens
That is the whole idea: signal divided by noise, scaled by how much data you have. Everything else about the t-test is machinery for turning that ratio into a probability, and understanding the ratio is enough to read a result critically.
It also explains the two disappointments. A real difference in a noisy, small sample fails to reach significance; a trivial difference in a huge sample reaches it easily. Both are the formula working correctly.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Three things make a gap convincing
  • A larger difference between the two means
  • Less spread (variation) within each group
  • A larger sample in each group
The same 7-point gap is unconvincing with 20 noisy pupils per arm, but compelling with 2,000 tightly clustered ones. Sample size and spread decide the verdict.
Makes a gap convincingIn practice
A big gapThe part you cannot control
Low variability within groupsBetter measurement; more homogeneous comparison
A large sampleThe part you can plan for, at design stage
Only the third row is under your control after the fact, and it is decided before any data is collected. This is why power calculations belong at design and are useless at analysis, when the sample is whatever it is.
The second row is undervalued. Reducing measurement error — better questions, trained enumerators, a cleaner instrument — raises your ability to detect a difference just as surely as adding respondents, and usually costs less.
ImpactMojoBivariate Analysis 101www.impactmojo.in
What a p-value does and doesn't say
p-value
The probability of seeing a difference at least this large if there were truly no difference between the groups. Small p (say < 0.05) means the gap is unlikely to be pure chance.
A p-value is not the probability the effect is real, nor a measure of its size. It only addresses 'could this be chance?' — nothing about importance.
A p-value isA p-value is not
The probability of data this extreme if there were no real differenceThe probability that there is no difference
A statement about the data, given a hypothesisA statement about the hypothesis, given the data
Sensitive to sample sizeA measure of importance
Compared to a threshold someone choseA natural boundary at 0.05
The first row is the definition and is almost universally misstated, including in published papers. The conditional runs one way: it assumes no difference and asks how surprising the data is, not the reverse.
0.05 is a convention, not a discovery. Treating 0.049 and 0.051 as categorically different results is arithmetic superstition, and reporting the actual value with a confidence interval avoids the whole problem.
ImpactMojoBivariate Analysis 101www.impactmojo.in
How big, in plain terms
Significance asks whether there is a difference; effect size asks how big. Report it in real units (7 marks, ₹400/month) or as a standardised measure (Cohen's d) so others can judge whether it matters.
p-value
Is the gap likely real, not chance?
Effect size
Is the gap big enough to act on?
Way to state the sizeExample
Raw difference in units“0.4 kg heavier on average”
With a confidence interval“0.4 kg (95% CI 0.1 to 0.7)”
Standardised (Cohen’s d)“d = 0.3” — for comparing across studies
In programme terms“About one additional child per 20 households”
The last row is the one decision-makers use and the one analysts most often omit. Translating an effect into the unit a programme operates in — households, children, rupees per beneficiary — is part of the analysis, not a communications afterthought.
Prefer raw units with an interval to a standardised measure. Cohen’s d is useful when comparing across studies with different scales; within one programme it converts something concrete into something abstract.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Significant but trivial — and vice versa
Significant but trivial
Huge sample → a 0.3-mark difference is 'significant' but means nothing for any child.
Real but 'not significant'
Tiny pilot → a promising 9-point gap fails the test purely for lack of sample. Absence of evidence is not evidence of absence.
Always read the p-value and the effect size together. Neither alone tells the whole story.
CaseWhat to write
Significant, tiny“A difference of 0.2 points, statistically detectable and too small to act on”
Large, not significant“A gap of 8 points that this sample cannot confirm; worth a larger study”
Large and significantThe straightforward case — report both numbers
Small and not significant“No evidence of a difference” — not “no difference”
The fourth row’s distinction is not pedantry. “No evidence of a difference” leaves open that the study was too small to see one; “no difference” asserts something the data cannot support, and it is how negative findings get overstated.
Reporting the confidence interval resolves all four rows at once. An interval from −0.1 to 0.3 says “small, whatever else”; one from −5 to 12 says “we do not know”, and neither can be misread as the other.
ImpactMojoBivariate Analysis 101www.impactmojo.in
When the simple t-test is appropriate
  • The two groups are independent (different people)
  • Y is roughly symmetric within each group (or n is large)
  • For paired data — same people before/after — use a paired t-test instead
  • For severely skewed data, consider a rank-based alternative
AssumptionCheckIf broken
Independent observationsHow was the sample drawn?Account for clustering
Roughly normal, or large NHistogram per groupMann–Whitney, or transform
Similar spread in both groupsBox plots side by sideWelch’s t-test — often the better default
No extreme outliersLook at the plotInvestigate before deciding
Welch’s version is a sensible default for group comparisons in field data, because it does not assume equal variances and loses very little when they are equal. Many packages now use it automatically; check which one yours ran.
The first assumption is the one that cannot be patched at analysis. If the design was clustered, the analysis must account for it — and no choice of test statistic repairs a p-value computed as though villages were households.
ImpactMojoBivariate Analysis 101www.impactmojo.in
07
Section Seven
Comparing Several Groups (ANOVA)
ImpactMojoBivariate Analysis 101www.impactmojo.in
When there are three or more groups
A categorical X with three or more values — four caste groups, five wealth quintiles, six districts — and a numeric Y. We want to know whether the group means differ overall.
Tempting shortcut: run a t-test on every pair. Don't. With many pairs, the chance of a false 'significant' result piles up fast — the multiple-comparisons problem.
Why not just run many t-tests?
Five groups gives ten pairwise comparisons
Each carries its own chance of a false positive
At 5% each, the chance of at least one false positive is around 40%
So “one pair differed” means very little without adjustment
This is the multiple-comparisons problem, and it is why ANOVA exists. One overall test asks whether the groups differ at all, at a controlled error rate, before you go looking for which pair.
The same arithmetic applies to subgroup analysis generally. Testing an outcome across twelve districts and reporting the one that reached significance is the forking-paths problem with a respectable name.
ImpactMojoBivariate Analysis 101www.impactmojo.in
One-way ANOVA, conceptually
One-way ANOVA
Analysis of Variance — a single test of whether the means of three or more groups differ by more than chance, by comparing variation BETWEEN groups to variation WITHIN groups.
Despite the name, ANOVA is about means. It uses variances as the yardstick for deciding whether those means are really apart.
ANOVA answersIt does not answer
Is there a difference somewhere among the groups?Which groups differ
At a controlled overall error rateBy how much
Under stated assumptionsWhether the difference matters
A significant ANOVA is a permission slip, not a finding. It licenses you to look for which pairs differ, using a method that adjusts for having looked — and on its own it tells a programme nothing it can act on.
Report the group means and their intervals whatever the F test says. Those are the numbers a reader needs; the test statistic is a gatekeeping device, not the result.
ImpactMojoBivariate Analysis 101www.impactmojo.in
The heart of ANOVA
Between-group variation
How far apart the group means are from each other — the signal.
Within-group variation
How much individuals scatter inside each group — the noise.
The F-ratio = between ÷ within. If groups sit far apart relative to their internal scatter, F is large — evidence the means genuinely differ.
Source of variationWhat it captures
Between groupsHow far the group means sit from the overall mean
Within groupsHow much individuals vary around their own group’s mean
F = between ÷ withinLarge when the groups are far apart relative to internal spread
It is the t-test’s signal-to-noise idea generalised to several groups. Nothing conceptually new is happening; the ratio simply compares variation attributable to group membership with variation that has nothing to do with it.
Which is why within-group spread matters so much. Five quintiles with widely different means but enormous internal scatter will not produce a large F — and looking at the box plots shows you that before the test does.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Mean nutrition score by wealth quintile
Mean child height-for-age z-score by wealth quintile (illustrative)
Illustrative, patterned on NFHS-5
Five group means rising steadily across quintiles. ANOVA asks: taken together, is this spread of means bigger than the scatter within each quintile? (Illustrative; the real gradient is well documented.)
ImpactMojoBivariate Analysis 101www.impactmojo.in
It says 'somewhere', not 'where'
A significant ANOVA tells you the groups are not all equal — but not which ones differ. To locate the differences, follow up with post-hoc pairwise comparisons that correct for multiple testing.
Think of ANOVA as the smoke alarm: it tells you there is a fire somewhere, then post-hoc tests find the room.
After a significant ANOVAWhy
Run a post-hoc test (Tukey, Bonferroni)Adjusts for the many comparisons you are about to make
Report every group mean with its intervalThe pattern matters, not just the significant pair
Say how many comparisons were madeA reader cannot judge otherwise
Look at the orderingA monotonic gradient across quintiles is a stronger finding than one odd pair
The last row is the one that carries real weight in our work. If nutrition scores rise steadily from the poorest quintile to the richest, that gradient is far more informative than any single pairwise test, and it is visible without one.
Do not run unadjusted t-tests after ANOVA. It reintroduces exactly the error rate the ANOVA was protecting you from, and it is the commonest way this sequence is misused.
ImpactMojoBivariate Analysis 101www.impactmojo.in
When one-way ANOVA fits
  • Groups are independent
  • Y is roughly symmetric within each group (or n is large)
  • Group spreads are not wildly different
  • For badly skewed data, a rank-based alternative (Kruskal–Wallis) may suit
ANOVA is the natural extension of the two-group t-test to many groups — same logic, scaled up.
AssumptionIf broken
Independent observationsAccount for clustering; ANOVA cannot fix it
Roughly normal within groupsKruskal–Wallis, the rank-based alternative
Similar spread across groupsWelch’s ANOVA
Groups defined before lookingOtherwise the test is meaningless
The fourth row is not usually listed as an assumption and matters more than the other three. Groups formed after inspecting the outcome — “the blocks that did badly” — guarantee a difference, because the difference is how the groups were chosen.
Kruskal–Wallis is a safe fallback for the skewed outcomes common here, and costs little. As with Spearman, choose it from the shape of the data rather than from which result you prefer.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Comparing groups, at a glance
GroupsY typeMethod
2 (independent)NumericTwo-sample t-test
2 (same units, paired)NumericPaired t-test
3 or moreNumericOne-way ANOVA
3+, skewed dataNumericKruskal–Wallis
GroupsParametricRank-based
Twot-test (Welch by default)Mann–Whitney
Three or moreOne-way ANOVAKruskal–Wallis
ThenPost-hoc pairwise, adjustedAdjusted pairwise ranks
Pick the row from the number of groups and the column from the shape of the data. That is the entire decision, and it removes almost all of the anxiety about choosing a test for group comparisons.
Whichever cell you land in, report the group medians or means with intervals. The test decides whether to believe the pattern; the summary statistics are the pattern.
ImpactMojoBivariate Analysis 101www.impactmojo.in
08
Section Eight
Association Between Categories (Chi-Square)
ImpactMojoBivariate Analysis 101www.impactmojo.in
Two categories: are they independent?
Back to two categorical variables and a cross-tab. The chi-square test of independence asks: is the pattern in this table bigger than we would expect if the two variables were completely unrelated?
Independence
Two variables are independent if knowing one tells you nothing about the other — the distribution of Y is the same within every category of X.
Chi-square testsTypical use here
Whether two categorical variables are independentSocial group × scheme uptake
Using counts, not percentagesFeed it the raw table
Against what independence would predictThe expected counts on the next slide
Chi-square answers a yes/no question and is silent on strength. It tells you the pattern is unlikely under independence; the percentage-point gap in the cross-tab tells you whether it matters, and only one of those belongs in a recommendation.
Feed it counts, never percentages. Running chi-square on a table of percentages produces a number that depends on whether you scaled to 100 or to 1,000 — a mistake that is easy to make and invisible in the output.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Observed vs expected counts
Chi-square compares what you observed in each cell with what you would expect if the two variables were independent. Big gaps between observed and expected = evidence of association.
01
Compute expected counts (if independent)
02
Compare with observed counts, cell by cell
03
Sum the scaled squared gaps → χ²
04
Large χ² → reject independence
StepWhat happens
Compute what each cell would hold if the variables were independentRow total × column total ÷ grand total
Compare with what was observedCell by cell
Square the gaps, scale by expected, add upLarge when observed departs from expected
Refer to the distributionGives the p-value
The scaling in step three is what makes the statistic sensible. A gap of ten cases matters enormously in a cell expected to hold twelve and hardly at all in one expected to hold two thousand.
Look at the individual cell contributions, not just the total. Most software will report them, and they tell you which combination is actually driving the result — which is the finding a programme can use.
ImpactMojoBivariate Analysis 101www.impactmojo.in
What 'no relationship' would predict
Each cell's expected count is (row total × column total) ÷ grand total. It is the count you'd see if Y were distributed identically across every category of X.
Example: with 660 of 1,000 households owning a toilet (66%), independence predicts 66% of the 600 rural households — 396 — would own one. We observed only 320. That gap is the signal.
Expected count for a cellMeaning
Row total × column total ÷ NWhat that cell would hold if the two variables were unrelated
Uses only the marginsIt preserves how many are in each category overall
Rarely a whole numberThat is fine; it is an expectation, not a count
Expected counts are the whole idea of the test made explicit. Independence is not an abstraction here: it is a specific table you can compute, and the statistic simply measures how far the real table sits from it.
Check them before trusting the p-value. The validity rule on the next slides is about expected counts, not observed ones — a cell can hold seven real cases and still be too sparse if only two were expected.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Side by side (illustrative)
CellObservedExpectedGap
Rural, toilet320396−76
Rural, no toilet280204+76
Urban, toilet340264+76
Urban, no toilet60136−76
Rural areas have far fewer toilets than independence predicts, urban far more. Consistent gaps like these push χ² up. (Illustrative figures.)
ImpactMojoBivariate Analysis 101www.impactmojo.in
What chi-square tells you
Small p
Reject independence — the two variables ARE associated
Large p
No evidence of association — consistent with independence
Chi-square tells you whether there is an association, not how strong it is. For strength, report a measure like Cramér's V alongside the test.
Chi-square gives youYou still need
Evidence against independenceThe cross-tab, to see the direction
A p-valueA measure of strength — Cramér’s V, or the percentage gap
Nothing about which cellCell contributions, or the percentages
Nothing causalEverything in Section 10
Reporting a chi-square result alone is close to reporting nothing. “Uptake was associated with social group (p < 0.01)” tells a reader that a relationship exists and not which group is disadvantaged or by how much — the only two facts a programme needs.
Lead with the table. “Uptake was 61% among general-category households and 44% among SC/ST households (p < 0.01)” carries the test and the finding in one sentence.
ImpactMojoBivariate Analysis 101www.impactmojo.in
When chi-square is valid
  • Each case counted once (independent observations)
  • Cells hold counts, not percentages or means
  • Expected count ≥ 5 in (almost) every cell
  • Categories are mutually exclusive and exhaustive
The expected-count rule is the one most often broken: with thin cells the chi-square approximation fails. Use Fisher's exact test for small tables instead.
RequirementIf it fails
Counts, not percentages or meansThe statistic is meaningless
Independent observationsClustering inflates the statistic
Expected count of about 5 or more in most cellsCombine categories, or use Fisher’s exact test
Each case in exactly one cellMultiple-response questions break this
The last row catches people with survey data. A “tick all that apply” question puts one respondent in several cells, which violates the test’s basic structure — and the software will happily produce a p-value anyway.
Combining sparse categories is legitimate if the combination is defensible on substantive grounds and decided before seeing the result. Merging categories until the p-value crosses a threshold is not.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Three ways chi-square goes wrong
  • Feeding it percentages instead of raw counts
  • Ignoring tiny expected counts in sparse cells
  • Reading a significant result as proof of causation
  • Forgetting it says nothing about the strength of association
MistakeConsequence
Running it on percentagesThe result scales with your choice of base
Ignoring sparse cellsAn unreliable p-value that looks normal
Reporting significance without directionA finding nobody can act on
Reading it as causalThe error the whole of Section 10 addresses
All four are avoidable by looking at the cross-tab first. The table shows the counts, the sparse cells and the direction; the test adds one number to a picture you should already have understood.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Chi-square in a development workflow
It is the natural partner to the cross-tab: you describe the two categories with row percentages, then test whether the visible gap is bigger than chance with chi-square.
Worked flow
Caste × toilet ownership
Build the cross-tab → percentage within caste → eyeball the gap → check expected counts ≥ 5 → run chi-square → report the test and a strength measure → resist the leap to 'caste causes…'.
StepOutput
Build the cross-tab with countsThe raw table
Add row percentages in the causal directionThe comparison
Check expected countsWhether the test is valid
Run chi-squareEvidence against chance
Report percentages, gap, N and pSomething a programme can use
List plausible confoundersHonest framing
Steps one and two do most of the work, and are frequently skipped in favour of going straight to the test. The table is the finding; the test is a check on whether the table is worth discussing.
Step six costs a sentence and buys credibility. Naming the obvious confounder yourself — usually wealth or location — is what distinguishes an analyst who understands the limits from one who has not thought about them.
ImpactMojoBivariate Analysis 101www.impactmojo.in
The category × category toolkit
StepToolAnswers
DescribeCross-tab + % What does the pattern look like?
TestChi-squareIs it bigger than chance?
Quantify strengthCramér's VHow strong is it?
Small table?Fisher's exactSame question, thin cells
NeedUse
See the patternCross-tab with row percentages
Test independenceChi-square, on counts
Small or sparse tableFisher’s exact test
Strength of associationCramér’s V, or the percentage-point gap
Two binary variables, risk framingRisk ratio or odds ratio
The last row is worth knowing for health work. Risk ratios and odds ratios are the standard language of epidemiology, and they are not interchangeable — odds ratios exaggerate when the outcome is common, which is a frequent source of misreporting.
For most programme reporting, the percentage-point gap is enough. It is understood without training, it is in the units of the decision, and it does not need a footnote explaining what it means.
ImpactMojoBivariate Analysis 101www.impactmojo.in
09
Section Nine
Simple Linear Regression
ImpactMojoBivariate Analysis 101www.impactmojo.in
Drawing the best line through a scatter
Correlation gives a number for a numeric–numeric relationship. Simple linear regression goes one step further: it fits the single straight line of best fit through the scatter, summarising the relationship as an equation.
Simple linear regression
Modelling a numeric response Y as a straight-line function of one numeric predictor X: Y = intercept + slope × X, plus error.
Regression adds to correlationWhy it matters
A slope in real units“0.4 fewer days per extra kilometre”
A prediction for any XFitted values, with their uncertainty
An explicit directionYou must nominate X and Y
Residuals to inspectWhere the model fails, and on which cases
The line is chosen to minimise squared vertical distances, which is why extreme points pull it so hard: an error twice as large counts four times as much. That single fact explains most of the outlier behaviour in Section 10.
The fourth row is the underused output. Plotting residuals against fitted values reveals curvature, changing spread and influential cases in one picture — and takes one line of code that almost nobody runs.
ImpactMojoBivariate Analysis 101www.impactmojo.in
A regression line over a scatter
Years of schooling vs monthly earnings, with line of best fit
Illustrative
The red line is the one that sits closest to all the points at once — it minimises the total squared vertical distance from points to line (least squares). Illustrative data.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Y = a + bX
a (intercept)
Predicted Y when X = 0 — where the line crosses the vertical axis
b (slope)
Change in Y for each 1-unit rise in X — the rate of the relationship
The slope carries the action. Its sign matches the correlation's sign; its size is in real units of Y per unit of X.
TermReads as
a (intercept)Predicted Y when X is zero
b (slope)Change in Y for a one-unit rise in X
ResidualWhat the line got wrong for this case
Fitted valueWhat the line predicts for this case
The slope is the number that belongs in the report, because it is in the units the programme thinks in. Everything else in the output is machinery for deciding how much to trust it.
Write the units into the sentence. “b = −0.42” means nothing on its own; “0.42 fewer days attended per additional kilometre” is the same number and a usable finding.
ImpactMojoBivariate Analysis 101www.impactmojo.in
When the intercept is meaningless
The intercept is the predicted Y at X = 0 — but X = 0 may be nonsensical or far outside your data. The earnings at zero years of schooling may be a mathematical artefact, not a real group.
Interpret the intercept only when X = 0 is both meaningful and within the range of your data. Otherwise treat it as a mathematical anchor for the line.
XIs X = 0 meaningful?
Distance to facilityYes — living at the facility
Years of schoolingYes — no schooling
Mother’s ageNo — there are no newborn mothers
Year (e.g. 2015–2024)No — the intercept refers to year zero
Where X = 0 is impossible, the intercept is arithmetic rather than a finding. It is required to position the line and should not be interpreted, and reporting it as though it means something is a small but common embarrassment.
Centring fixes it. Subtract the mean of X before fitting, and the intercept becomes the predicted Y at the average X — a quantity that exists and is worth quoting.
ImpactMojoBivariate Analysis 101www.impactmojo.in
What b means in plain words
In our example the line rises about ₹1,250 per extra year of schooling (illustrative). So b ≈ 1.25 in ₹000: ‘each additional year of schooling is associated with about ₹1,250 more monthly earnings.’
Note the phrase ‘associated with’, not ‘causes’. Regression fits a line; it does not, by itself, license a causal claim.
SayDo not say
“Associated with 0.4 fewer days”“Causes 0.4 fewer days”
“On average, across this sample”“For any given household”
“Within the observed range of X”“At any distance”
“Before accounting for wealth and location”Nothing about what was not controlled
A slope from observational data is an association per unit, not an effect of intervening. If you halved everyone’s distance tomorrow, the change in attendance would not be the slope — because distance is entangled with everything else that differs between near and far households.
The second row matters for targeting. A slope describes an average across the sample; individual households vary around it enormously, and a prediction for one household carries far wider uncertainty than the line suggests.
ImpactMojoBivariate Analysis 101www.impactmojo.in
How much variation the line explains
R² (coefficient of determination)
The proportion of the variation in Y that is explained by X through the fitted line. It runs from 0 (line explains nothing) to 1 (line explains everything). In simple regression, R² = r².
R² = 0.8
Line explains 80% of the variation in Y; 20% is left unexplained
R² = 0.1
Line explains just 10% — X is a weak guide to Y
Reads as
0.05The line accounts for 5% of the variation — typical for social data
0.30A substantial share, by the standards of this field
0.95Suspicious in survey data — check for a definitional link
Low R² is normal and not a failure. Human behaviour varies for reasons no two-variable model captures, and a slope can be precisely estimated, important and worth acting on while the model explains a small fraction of the scatter.
The third row is a warning, not a triumph. An R² that high between two social variables usually means one was computed from the other — income and income-per-capita, or a score and one of its own components.
ImpactMojoBivariate Analysis 101www.impactmojo.in
High R² is not always good
  • A high R² does not prove the model is correct or causal
  • In social data, an R² of 0.2 can still be a useful finding
  • A low R² with a clear slope can still matter for policy
  • R² says nothing about whether a straight line fits
Report R² for context, but never let it crowd out the slope, its uncertainty, and a look at the scatter.
High R² can come fromWhich is not a good sign
Both variables trending over timeAlmost any two series will fit well
Aggregated data — districts, statesAveraging removes individual noise mechanically
One variable derived from the otherA definitional relationship, not a finding
A few extreme pointsLeverage, not fit
The second row explains why state-level scatters look so convincing. Aggregation averages away individual variation, so district or state data routinely produces R² values that would be impossible at household level — and the tighter fit says nothing about individuals.
Judge a model by the residual plot, not by R². A high R² with a clear curve in the residuals is a worse model than a low R² with residuals scattered evenly.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Two different goals for one line
Explanation
Understand the relationship: how does Y change with X, and how strong is it? The slope is the prize.
Prediction
Estimate Y for a new X you haven't measured. Accuracy, not interpretation, is the prize.
Be clear which you are doing — the same line, read for different purposes, demands different cautions.
PredictionExplanation
GoalGuess Y accuratelyUnderstand how X relates to Y
Cares aboutFit and out-of-sample accuracyThe slope and its uncertainty
ConfoundersDo not matter if fit holdsMatter completely
Useful forTargeting, forecastingDeciding what to change
The third row is the distinction that matters most. A model that predicts well can be built on pure proxies and still be useful for finding households to visit; the same model says nothing about what would happen if you changed one of those variables.
Trouble arrives when a predictive model is read as an explanatory one. A targeting model that flags households by roof type is not a finding that roofs cause deprivation, and acting on it as though it were produces roof-replacement programmes.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Don't predict beyond your data
Extrapolation — using the line far outside the range of X you actually observed — is one of regression's classic traps. A line fitted on 0–16 years of schooling says nothing reliable about 30 years.
The relationship may bend, plateau or break entirely outside your data. Predict within the range you measured, and flag any step beyond it.
Why extrapolation failsExample
The relationship may curve outside the rangeDose response that plateaus
No data means no evidence of the shapeThe line continues because it is a line
Predictions become physically impossibleNegative attendance; percentages above 100
Uncertainty widens sharply at the edgesAnd keeps widening beyond them
State the range of X your model was fitted on, and refuse predictions outside it. This is a one-sentence discipline and it prevents the most confident category of error — a number computed by the software with nothing behind it.
Time-series projection is the commonest case. A trend fitted to five years and projected to twenty is arithmetic, not forecasting, and the intervals around it — if anyone computed them — would usually be wide enough to include almost anything.
ImpactMojoBivariate Analysis 101www.impactmojo.in
10
Section Ten
Pitfalls
ImpactMojoBivariate Analysis 101www.impactmojo.in
Correlation is not causation
A relationship between two variables can arise for several reasons, only one of which is ‘X causes Y’. This is the single most important caution in all of bivariate analysis.
  • Reverse causation: Y might cause X
  • Confounding: a third factor C drives both
  • Selection: how the sample was chosen creates the link
  • Chance: with enough variables, some correlate by luck
Before writing a causal sentence, rule out
Reverse causation — could Y have produced X?
Confounding — what third factor drives both?
Selection — how did units end up in each group?
Chance — how many relationships were examined?
Measurement — do both variables share a common error?
The exercise is to say each aloud and explain why it is unlikely here. If you cannot, the finding is an association and should be written as one — which is not a weaker claim, only an accurate one.
The fifth row is rarely considered. If the same enumerator recorded both variables, or both come from one respondent’s recall, a shared bias can create a correlation with no relationship behind it at all.
ImpactMojoBivariate Analysis 101www.impactmojo.in
The lurking third variable
A confounder is a variable linked to both X and Y that creates — or distorts — the relationship between them. It is the commonest reason a bivariate link is misleading.
01
Household wealth (confounder C)
02
drives girls' schooling (X)
03
AND drives child survival (Y)
04
so X and Y correlate — partly via C
Apparent bivariate linkThe usual confounder
Toilet ownership and child heightHousehold wealth
Training attendance and yieldsWho chooses to attend
Private schooling and test scoresParental education and income
Distance and facility deliveryRemoteness, which brings poverty and fewer roads
Wealth is the confounder in most development data, because it is correlated with nearly everything we measure. A bivariate relationship that has not been checked against wealth has not really been checked.
The cheapest partial fix is stratification. Repeat the analysis within each wealth quintile: if the relationship survives inside every stratum, it is not purely a wealth artefact — and if it disappears, you have learned that in an afternoon.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Bivariate links almost always hide confounders
Recall our literacy–fertility scatter. Female literacy really does track lower fertility — but richer states also have better health systems, later marriage and more contraception. Wealth and development confound the simple two-variable picture.
This is exactly why bivariate analysis is a starting point. To isolate one variable's effect you need to control for confounders — the job of multivariate regression.
What bivariate analysis is good forWhat it should not be used for
Describing who is worse offClaiming why they are worse off
Finding where to look nextDeciding what to change
Screening many candidate relationshipsConcluding from the one that survived
Communicating a pattern simplyAttributing it to a programme
The left column is genuinely valuable and often undersold. Equity analysis — who is being reached, who is not — is descriptive by nature, needs no causal claim, and is the most useful thing most programmes could do with their own data.
Moving to the right column requires design, not statistics. A comparison group, randomisation, or a credible natural experiment — none of which can be added at analysis stage.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Group patterns ≠ individual truths
Ecological fallacy
Wrongly inferring something about individuals from a relationship observed only at the group (district, state) level.
States with higher average literacy have lower average fertility — that does not mean the literate women within a state are the ones with fewer children. A relationship across districts need not hold across people.
Unit of the scatterWhat a pattern can support
StatesA statement about states
DistrictsA statement about districts
HouseholdsA statement about households
IndividualsA statement about individuals
The rule is that simple and is broken constantly. A famous state-level scatter of female literacy against fertility is a fact about states; every reader takes it as a fact about women, and the two can even point in opposite directions.
Put the unit in the chart title. “By state” or “by household” costs two words and is the only thing standing between an honest chart and an ecological inference.
ImpactMojoBivariate Analysis 101www.impactmojo.in
One point can swing everything
A single extreme point can dominate a correlation or drag a regression line toward itself — high leverage. The whole ‘relationship’ may rest on one unusual case.
Always check: does the pattern survive if I remove the most extreme point? If r or the slope collapses, the finding was that one point, not a real trend.
TermMeaning
OutlierFar from the cloud in Y
LeverageFar from the cloud in X — can swing the slope hard
InfluenceBoth: removing it changes the conclusion
Because the line minimises squared distances, one far-out point can dominate the fit. In a scatter of Indian states this is routine — a very large or very small state sits alone at one end and quietly determines the slope.
The honest test is to refit without it and report both. If the conclusion changes, that is the finding: your result depends on one observation, and the reader is entitled to know which.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Straight-line tools miss curves
Pearson's r and linear regression assume a straight line. A U-shaped or saturating relationship — common in real data — can show a near-zero correlation while being strongly, but non-linearly, related.
The fix is the cheapest in statistics: plot the scatter. A curve is obvious to the eye and invisible to r.
Common non-linear shapes hereExample
Diminishing returnsEach extra year of schooling adds less
ThresholdNothing happens until a minimum dose is reached
CeilingCoverage cannot exceed 100%, so gains flatten
U-shapeBenefit rises then falls — r near zero
Straight-line tools will fit a straight line to any of these and report a slope, an r and a p-value. Only the scatter reveals that the summary is describing a shape the data does not have.
Diminishing returns is the most common and the most consequential. A linear model averages the steep early gains with the flat later ones, understating what a programme achieves at the bottom and overstating it at the top.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Four datasets, identical statistics
Anscombe-style: same r, four very different shapes
After F. J. Anscombe (1973)
Anscombe's quartet (1973): four datasets with near-identical means, variances and correlation, yet utterly different shapes. The lesson, still true: always plot your data.
ImpactMojoBivariate Analysis 101www.impactmojo.in
A trend can reverse when you split
Simpson's paradox: a relationship in the pooled data can flip direction within every subgroup. A scheme can look worse overall yet be better in every district — if districts differ in size and baseline.
Always disaggregate before concluding. The aggregate two-variable relationship can point the opposite way to the truth.
Aggregate saysEvery subgroup saysBecause
Hospital A does worseA does better in each severity bandA takes the severe cases
Scheme B looks weaker overallB is stronger in every districtB operates where uptake is hardest
Neither view is wrong; they answer different questions. Which one you want depends on the decision — a patient choosing a hospital and a regulator assessing one need different tables, and the choice has to be argued rather than assumed.
Always split an aggregate result by the obvious grouping variable once. A reversal is a finding worth understanding; discovering it after publication is not.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Before you believe a relationship
  • Did I plot it — is the shape really linear?
  • Could a confounder explain it?
  • Does it survive removing the outlier?
  • Is this a group pattern I'm reading onto individuals?
  • Does it reverse when I disaggregate?
  • Is there a plausible mechanism, or just a coincidence?
CheckCatches
Did I plot it?Curves, clusters, outliers
What is the unit of analysis?Ecological fallacy
What is the obvious confounder?Wealth, almost always
Does it survive stratification?Spurious associations
Does one point drive it?Leverage
How many relationships did I test?Forking paths
How big is it, in units?Significant-but-trivial
Seven checks, none requiring new statistics. Running them on every bivariate finding before it leaves your desk is what separates analysis that survives scrutiny from analysis that is quietly withdrawn later.
ImpactMojoBivariate Analysis 101www.impactmojo.in
11
Section Eleven
Reporting, Practice & Tools
ImpactMojoBivariate Analysis 101www.impactmojo.in
How to report a relationship
  • State the direction and strength in plain words
  • Give the effect size in real units, not just a p-value
  • Report the uncertainty — confidence interval or range
  • Name the sample size and how it was drawn
  • Flag the obvious confounders you could not rule out
ReportExample
The pattern in plain words“Uptake was lower among SC/ST households”
The numbers, with units“44% vs 61% — a gap of 17 points”
N, per group“(n = 312 and 908)”
Uncertainty“95% CI for the gap: 11 to 23 points”
The obvious confounder“Not adjusted for wealth or block”
All five fit in two sentences, and a report that contains them cannot easily be misread. The version that says only “uptake was significantly lower (p < 0.01)” has withheld everything a reader needs to judge it.
Naming the confounder yourself is a strength, not a weakness. It is what a competent reviewer will ask first, and answering it in advance is the difference between careful and defensive.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Mind your verbs
Safe phrasing
  • ‘is associated with’
  • ‘tends to go with’
  • ‘is correlated with’
Reserve for evidence
  • ‘causes’
  • ‘leads to’
  • ‘increases / reduces’
Causal verbs need a causal design — an experiment or a careful quasi-experiment. A bivariate correlation does not earn them.
Safe with observational dataReserve for causal designs
Associated with, linked toCauses, produces, leads to
Higher among, lower inRaises, reduces, drives
Accompanied by, varies withImpacts, results in
Predicts (statistically)Improves, prevents
Verbs are where causal claims enter a report unnoticed. The analysis section says “associated”, the executive summary says “improves”, and the press release says “proves” — each step made by someone who did not run the analysis.
“Predicts” is the treacherous one. It has a precise statistical meaning and an everyday causal one, and a reader outside the field will hear the second. Where it matters, spell out “statistically predicts, without implying cause”.
ImpactMojoBivariate Analysis 101www.impactmojo.in
What to avoid in write-ups
  • Reporting a p-value with no effect size
  • Claiming causation from a correlation
  • Hiding the scatter behind a single coefficient
  • Testing many pairs and reporting only the one that worked
  • Using a state-level relationship to make individual claims
SinCorrection
Reporting p without an effect sizeGive the gap in units, with an interval
“No difference” from a non-significant result“No evidence of a difference”
Reporting only the relationship that workedSay how many were examined
Causal verbs on observational dataUse the left column of the previous slide
Percentages with no NAlways “62% (n = 187)”
A chart title that states a causeName the variables, not the mechanism
None of these is dishonesty; all of them are habits. They survive because they make findings sound stronger and nobody in the review chain is assigned to catch them — which is why a personal checklist is the practical remedy.
ImpactMojoBivariate Analysis 101www.impactmojo.in
A reliable bivariate workflow
01
CLASSIFY: what type is each variable?
02
PLOT: scatter, box plot or bars
03
DESCRIBE: direction, strength, shape
04
TEST: the matching method
05
INTERPRET: effect size + uncertainty
06
CAUTION: confounders, fallacies
Six steps, every time. The discipline is what turns a number into a defensible finding.
StepOutput
Write the question in one sentenceDetermines everything after
Classify both variablesChooses the method
Plot itShape, outliers, clusters
Describe it — means, medians, percentages, NThe finding
Test, if the question needs a testEvidence against chance
Check the pitfall listWhat could still be wrong
Write it with hedged verbs and the caveatSomething defensible
Steps three and four are the analysis; step five is a check on it. Reversing that order — running the test first and looking afterwards — is how most of the errors in Section 10 survive to publication.
Step one prevents the most waste. An analysis without a written question expands to fill the available time and produces a folder of tables nobody can act on.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Software for bivariate analysis
ToolGood forNote
Excel / Google SheetsCross-tabs, scatter, CORREL, t-testStart here; pivot tables for cross-tabs
REvery test, publication graphicsFree; cor.test, t.test, aov, chisq.test, lm
Python (pandas, scipy, statsmodels)Cleaning + tests + regressionFree, general-purpose
Jamovi / JASPPoint-and-click statsFree, friendly for learners
Stata / SPSSSurvey data, weightsCommon in research shops
Tools matter less than habits: classify, plot, then test. A clear scatter in a spreadsheet beats a misread coefficient in any package.
ToolGood for
SpreadsheetsCross-tabs, simple charts, sharing with non-analysts
REverything here, reproducibly; free
PythonSame, if the team already uses it
Stata / SPSSWhere your collaborators already work
Every method in this deck exists in all four. The reason to move beyond a spreadsheet is reproducibility — being able to rerun the whole analysis from raw data after a correction — not sophistication.
ImpactMojoBivariate Analysis 101www.impactmojo.in
A short, honest reading list
  • Statistics — Freedman, Pisani & Purves (the gold standard primer)
  • The Art of Statistics — David Spiegelhalter
  • How to Lie with Statistics — Darrell Huff
  • Naked Statistics — Charles Wheelan
  • Mostly Harmless Econometrics — Angrist & Pischke (going causal)
Pair this deck with ImpactMojo's Data Literacy, Exploratory Data Analysis and Research Methods 101 courses.
ReadFor
Wheelan, Naked StatisticsThese methods, informally and well
Reinhart, Statistics Done WrongSections 6, 7 and 10, in detail
Huff, How to Lie with StatisticsOld, short, and still the best on charts
Pearl & Mackenzie, The Book of WhyWhy correlation is not causation, properly
Your survey’s methodology noteWhat your data can actually support
Reinhart is the most useful for a working analyst. It is short, it is about the errors people actually make — multiple comparisons, misread p-values, underpowered studies — and it is available free online.
ImpactMojoBivariate Analysis 101www.impactmojo.in
If you remember five things
  • Classify both variables first — it picks the method
  • Plot before you compute — shape, outliers, curves
  • Report effect size, not just the p-value
  • Correlation is not causation — hunt the confounder
  • Group patterns are not individual truths — mind the fallacies
If you remember five thingsSection
Classify both variables; the method follows2
Plot it before you compute anything4
Report the effect size, not just the p-value6
Name the confounder yourself10
Match the verb to the design11
None of these is a technique; all five are habits. The methods in this deck are standard and available everywhere, and what distinguishes an analysis that holds up is the checking around them rather than the choice among them.
And the one underneath all five: a bivariate relationship is where the investigation starts. Treating it as where the investigation ends is the error this whole deck is arranged against.
ImpactMojoBivariate Analysis 101www.impactmojo.in
Bivariate Analysis 101 · Complete
Two variables.
One honest story.
CC BY-NC-ND 4.0·Free Forever·ImpactMojo 101 Series