fullscreen
ImpactMojoItem Response Theory 101www.impactmojo.in
ImpactMojo 101 Series · Free Forever
Item
Response
Theory 101
Measuring Latent Traits Well — a Foundational Course on IRT for Assessment & M&E Practitioners in South Asia
Research-BackedSouth Asia Focus100 SlidesFree Access
ImpactMojoItem Response Theory 101www.impactmojo.in
What We Cover
01
The Measurement Problem
Slides 3–10
02
Classical Test Theory & Its Limits
Slides 11–19
03
The Big Idea of IRT
Slides 20–28
04
The Item Characteristic Curve
Slides 29–37
05
Item Difficulty — the b Parameter
Slides 38–45
06
Item Discrimination — the a Parameter
Slides 46–54
07
Guessing & the Model Family
Slides 55–63
08
Information & Precision
Slides 64–72
09
Validating Scales & DIF
Slides 73–81
10
Applications in Development
Slides 82–90
11
Assumptions, Limits & Tools
Slides 91–99
ImpactMojoItem Response Theory 101www.impactmojo.in
01
Section One
The Measurement Problem
ImpactMojoItem Response Theory 101www.impactmojo.in
You cannot see what you most want to measure
Development work is full of things we care about but cannot observe directly: a child's reading ability, a woman's empowerment, a household's food insecurity. We only ever see responses — answers to items, ticks on a scale.
Latent trait
An unobservable characteristic of a person — ability, attitude, deprivation — that we infer from their answers to a set of items. By convention it is written as theta (θ).
Latent traitObservedWhat can go wrong
Reading abilityRight / wrong on tasksTasks measure vocabulary instead
EmpowermentAgree / disagree itemsThe construct is contested
Food insecurityYes / no experiencesSeasonality changes the answers
Health knowledgeCorrect answersRecall of a campaign slogan
The word "latent" is doing precise work: it means the trait is inferred from responses under a model, not observed. Every number you report is a model output, not a measurement.
That is not a weakness peculiar to IRT. A total score is also a model output; it just hides its assumptions instead of stating them.
ImpactMojoItem Response Theory 101www.impactmojo.in
We observe answers, not the trait
01
LATENT TRAIT: reading ability (θ) — unseen
02
ITEMS: a child reads words, sentences, a paragraph
03
RESPONSES: correct / incorrect on each item
04
INFERENCE: a measurement model estimates θ from the pattern
A measurement model is the bridge: it links the unseen trait to the seen responses, with explicit assumptions you can check.
StageThe assumption it rests on
Trait existsThe construct is coherent and one-dimensional
Items tap itContent validity — the items are about the trait
Responses recordedAdministration is consistent and honest
Model estimatesThe chosen model fits the data
Errors at the first two stages cannot be repaired at the fourth. A well-fitting model on items that measure the wrong thing gives you a precise number about nothing in particular.
This is why psychometricians spend more time on item writing and dimensionality than on estimation. The estimation is the easy part.
ImpactMojoItem Response Theory 101www.impactmojo.in
The same problem, three development settings
Ability
ASER-style early-grade reading & arithmetic — a child either can or cannot do each task
Empowerment
A woman's decision-making, mobility & asset control — agree / disagree items
Food insecurity
FIES — eight yes/no experiences from worry to going a whole day without eating
In each case the trait is on a hidden continuum and the items are graded markers along it.
SettingResponse typeModel usually used
ASER-style readingOrdered tasks, pass/failRasch or 2PL
Empowerment itemsAgree / disagree2PL or graded response
Food insecurity (FIES)Yes / no experienceRasch / 1PL
Multiple-choice testCorrect / incorrect3PL where guessing matters
Response format largely determines model choice. Yes/no experience items with no guessing route are natural Rasch candidates; four-option multiple choice is where the guessing parameter earns its place.
Ordered response categories (never / sometimes / often) call for a graded response or partial credit model, which are extensions of the same logic rather than different theories.
ImpactMojoItem Response Theory 101www.impactmojo.in
Why not just count right answers?
The obvious approach — add up correct answers, or count 'yes' responses — treats every item as equal and every score gap as the same size. But a hard item is not worth the same as an easy one, and the jump from 4 to 5 correct is not the same as 8 to 9.
A measurement model lets items differ — in difficulty, in how well they sort people — and places people and items on one common scale.
Sum score assumesWhen it fails
Every item is worth the sameA hard item and an easy one both count 1
Every gap is the same sizeThe jump from 4 to 5 is not the jump from 9 to 10
The pattern does not matterTwo people with 6/12 are treated as identical
The scale is intervalIt is ordinal at best
The last row is the formal statement of the problem. A total score orders people reliably and does not tell you how far apart they are, which means differences and averages of it are on shaky ground.
This matters most for change over time. "Scores rose by 2 points" means something different at the bottom of the scale than at the top, and a sum score cannot distinguish them.
ImpactMojoItem Response Theory 101www.impactmojo.in
Two families of measurement theory
Classical Test Theory (CTT)
Works on the total score. Simple, familiar, everywhere — but its statistics depend on the particular sample and the particular test.
Item Response Theory (IRT)
Models each item and each person separately, on a shared scale. More demanding, but the properties travel across samples and forms.
This course starts from CTT — what it is, and exactly where it strains — then builds IRT as the answer.
CTTIRT
Unit of analysisThe test as a wholeEach item
Person scoreSum or percentageEstimated theta
Item statisticsDepend on the sampleSample-free, if the model fits
PrecisionOne number for everyoneVaries along the scale
Data neededModestSubstantial
The last row is the honest trade-off and the reason CTT is not obsolete. With 80 respondents you cannot estimate item parameters reliably, and a well-constructed sum score is the better tool.
The two are not rival theories of measurement so much as different resolutions. Where you have the data, IRT tells you more; where you do not, it tells you noise.
ImpactMojoItem Response Theory 101www.impactmojo.in
Why measurement quality is not a luxury
  • Comparability: compare a child this year with last year, or one state's empowerment score with another's, on the same ruler
  • Fairness: detect items that behave differently for girls, or for one language group
  • Precision where it matters: know exactly where the test measures well and where it is guessing
  • Honest scores: a number that means the same thing for everyone
BenefitWhat it enables concretely
ComparabilityTrack learning across years with different items
FairnessDetect items biased against a group
EfficiencyDrop items that carry no information
Targeted precisionMeasure accurately at the decision point
Honest errorReport uncertainty that varies by person
Only two of these are available at all under CTT, and neither cleanly. That is the case for the extra data and complexity, and it should be made in those terms rather than as sophistication.
If none of the five is something your programme needs, you probably do not need IRT. That is a legitimate conclusion and saves a great deal of work.
ImpactMojoItem Response Theory 101www.impactmojo.in
How this course is built
Foundations
  • The measurement problem & CTT's limits
  • The big idea of IRT and the ICC
  • Difficulty (b), discrimination (a), guessing (c)
Using IRT well
  • Information, precision and test targeting
  • Validating scales: fit, dimensionality, DIF
  • Applications, assumptions and tools
Examples are drawn from learning assessments, empowerment, wealth and food-security scales used across the region.
ImpactMojoItem Response Theory 101www.impactmojo.in
02
Section Two
Classical Test Theory & Its Limits
ImpactMojoItem Response Theory 101www.impactmojo.in
The familiar sum or percentage score
Classical Test Theory is the world of the total score: number correct, percentage right, the count of 'yes' responses on a scale. It is what nearly every report and dashboard already uses.
01
Observed score = True score + Error
02
X = T + E
03
Goal: estimate T, shrink E
CTT termMeans
Observed scoreWhat the person actually got
True scoreTheir expected score over infinite retests
ErrorThe difference, assumed random
ReliabilityShare of variance that is true score
The true score in CTT is defined relative to a particular test. It is not an ability that exists independently of the instrument, which is exactly the limitation the next slides develop.
The error term is assumed random and uncorrelated with true score. Systematic error — a poorly translated item, a biased context — violates that and is invisible to the framework.
ImpactMojoItem Response Theory 101www.impactmojo.in
Reliability: how much is signal, not noise?
Reliability
The share of the variation in scores that reflects true differences between people rather than measurement error. It runs from 0 (all noise) to 1 (no error).
A reliable test gives nearly the same score if a person sits it twice. Low reliability means scores wobble for reasons that have nothing to do with the trait.
Reliability estimateAnswers
Test-retestWould the same person score the same again?
Parallel formsDo two versions agree?
Split-halfDo the two halves agree?
Cronbach's alphaDo the items hang together internally?
Each answers a different question, and they can disagree. A scale can be internally consistent and unstable over time, which alpha will not tell you.
Reliability also bounds validity: a measure that is mostly noise cannot be strongly related to anything. But high reliability does not establish that you are measuring the right thing.
ImpactMojoItem Response Theory 101www.impactmojo.in
Cronbach's alpha — the workhorse statistic
Cronbach's alpha estimates reliability from a single sitting by asking how consistently the items hang together. Rules of thumb put 'acceptable' around 0.70 and up — but the number is widely misread.
Alpha rises simply by adding more items, and high alpha does not prove the scale measures one thing. It is a useful summary, not a certificate of quality.
Alpha misreadingWhat is actually true
"Alpha 0.9 means it is unidimensional"Alpha rises with item count, not dimensionality
"0.7 is the threshold"A convention, not a property
"Higher is always better"Very high alpha often means redundant items
"Alpha validates the scale"It says nothing about what is measured
Alpha is a lower bound on reliability under specific assumptions, and it increases mechanically as you add items. A 40-item scale of near-duplicates will report an excellent alpha.
An alpha above about 0.95 is usually a warning rather than a triumph: it suggests the items are asking the same question in different words, which wastes respondent time.
ImpactMojoItem Response Theory 101www.impactmojo.in
CTT statistics are sample-dependent
An item's CTT 'difficulty' is just the proportion who got it right (the p-value). Give the same item to a high-ability school and a struggling one and that proportion changes — so the item looks 'easier' or 'harder' depending on who sat it.
The item did not change. The sample did. Yet the headline statistic moved — that is the problem.
Same item, given toCTT p-valueTrue item property
A high-ability groupHigh — "easy"Unchanged
A mixed groupMiddlingUnchanged
A struggling groupLow — "hard"Unchanged
The item did not change. Only the sample did, and the statistic moved with it — which means CTT item statistics cannot be pooled across studies or carried from a pilot to the field.
The same applies to item-total correlations, the CTT discrimination statistic. Both are properties of an item-in-a-sample rather than of the item.
ImpactMojoItem Response Theory 101www.impactmojo.in
Same item, two samples, two p-values
Group% correct on Item QCTT verdict
High-ability school88%An easy item
Mixed government school61%A moderate item
Struggling school34%A hard item
Illustrative figures. One physical item, three contradictory CTT difficulties — because the p-value confounds the item with the people who answered it.
ImpactMojoItem Response Theory 101www.impactmojo.in
CTT statistics are test-dependent
A person's CTT ability is their score on this particular test. Put them on an easier form and the score climbs; a harder form and it falls. The person's standing is tangled up with the difficulty of the form they happened to take.
Two children with the same true ability can get different scores purely because they sat different forms. Comparing across forms or years becomes guesswork.
Consequence of test-dependenceWhere it bites
Scores are not comparable across formsTwo schools sitting different papers
Scores are not comparable across yearsTracking change with new items
Percentage right is not an abilityReporting "60% proficient"
Anchoring is ad hocNo principled way to link forms
This is why examination boards historically relied on the same paper for everyone, which creates its own problems: item exposure, security, and no way to adapt to the candidate.
IRT's answer — equating through anchor items — is developed in Section 10 and is the single most practically valuable thing IRT offers a monitoring programme.
ImpactMojoItem Response Theory 101www.impactmojo.in
One error figure for the whole scale
CTT reports a single standard error of measurement for everyone. But in reality a test measures a middling student far more precisely than it measures a top scorer (who finds every item easy) or a very weak one (who finds every item hard).
Precision genuinely varies across the ability range. CTT pretends it is constant — IRT will let it vary, which is closer to the truth.
Person atTest measures themBecause
Middle of the rangePreciselyMany items near their level
Top of the rangePoorlyEverything is easy for them
Bottom of the rangePoorlyEverything is too hard
A single standard error for everyone therefore overstates precision at the extremes and understates it in the middle. Confidence intervals built from it are wrong in a predictable direction.
This matters wherever the decision is at an extreme — identifying children far behind grade level, or selecting top performers. Those are exactly the cases CTT handles worst.
ImpactMojoItem Response Theory 101www.impactmojo.in
Three jobs CTT cannot do cleanly
You want to…CTT problemIRT answer (coming)
Compare items fairly across samplesp-values shift with the sampleItem params are sample-free
Compare people across forms/yearsScores depend on the formθ is on a common scale
Know precision at each levelOne SEM for allInformation varies along θ
Build an adaptive or short formHard — scores not comparableItems & people share a metric
None of this means CTT is wrong — just that it asks too much of one total score. IRT splits the job apart.
JobCTT gives youIRT gives you
Compare items across samplesp-values that shiftParameters that do not
Compare people across formsForm-dependent scoresOne theta scale
Report precisionOne SEMError as a function of theta
Shorten the instrumentDrop by item-total rDrop by information
The word "if the model fits" belongs on every row of the right-hand column. Sample-free parameters are a property of a fitting model, not a free gift.
Where the model does not fit, IRT gives you sample-dependent parameters with more decimal places. Validation, in Section 9, is what earns the claims made here.
ImpactMojoItem Response Theory 101www.impactmojo.in
03
Section Three
The Big Idea of IRT
ImpactMojoItem Response Theory 101www.impactmojo.in
Put people and items on the same ruler
The central move of IRT is deceptively simple: place each person's trait (θ) and each item's difficulty on one shared continuum. A person is somewhere on the line; so is every item.
If a person sits above an item on the scale, they are likely to get it right or endorse it. Below it, unlikely. The distance between them sets the probability.
On the shared scaleYou can now ask
Person at theta, item at bIs she above or below this item?
Two itemsWhich is harder, regardless of who sat them?
Two peopleHow far apart, in the same units?
A cut-scoreWhich items measure best right there?
None of these questions is answerable in CTT, because there the person metric (score) and the item metric (p-value) are different quantities in different units.
Putting them on one line is the whole trick, and everything else in this course is a consequence of it.
ImpactMojoItem Response Theory 101www.impactmojo.in
A trait scale centred on zero
By convention θ is scaled to have mean 0 and standard deviation 1 in the reference group, usually running from about −3 to +3. Negative means lower ability / less of the trait; positive means more.
θ ≈ −2
low trait — struggles with most items
θ ≈ 0
average for the reference group
θ ≈ +2
high trait — succeeds on most items
ThetaRoughly meansCaution
-2Low trait; bottom few per centMeasured imprecisely
0Reference-group averageOnly for that reference group
+2High trait; top few per centAlso measured imprecisely
The zero point is a convention, set by the reference group used at calibration. It carries no absolute meaning, so "theta = 0" is not a standard of proficiency.
This is why reporting raw theta to non-specialists misleads. A negative theta reads as failure and may simply mean below the average of whoever was in the calibration sample.
ImpactMojoItem Response Theory 101www.impactmojo.in
IRT predicts a probability, not a yes/no
IRT never says a person will get an item right. It gives the probability of a correct (or endorsing) response, as a function of where the person and the item sit on the scale.
P(correct | θ)
The probability that a person with trait level theta answers an item correctly (for a test) or endorses it (for an attitude / experience scale). It rises smoothly as theta rises.
IRT saysIRT does not say
P(correct) = 0.7 at this thetaThis person will get it right
This item is harder than that oneThis item is unfair
Theta is estimated at 1.2Theta is exactly 1.2
Precision is highest hereThe construct is valid
Every theta comes with a standard error, and reporting the point estimate without it invites exactly the over-reading that IRT was built to avoid.
The probabilistic framing is also what makes the model testable: predicted probabilities can be compared with observed proportions, which is what fit statistics do.
ImpactMojoItem Response Theory 101www.impactmojo.in
The S-shaped probability curve
That probability follows a smooth logistic function of θ. Far below the item, P is near 0; far above, near 1; in between it climbs through an S-shape. This curve is the heart of IRT.
Crucially the curve is monotonic: more trait always means a higher chance of success. It never dips.
Curve regionProbabilityWhat it tells you
Far below bNear 0 (or c)The item is out of reach
Just below bRising steeplyThe item is discriminating here
At b0.5 (2PL/1PL)The item's location
Far above bNear 1The item is trivial; no information
The logistic form is a modelling choice, not a discovery. The normal ogive gives an almost identical curve; logistic is used because it is mathematically convenient.
Monotonicity is the substantive assumption, not the exact functional form: more of the trait must never reduce the chance of a correct response.
ImpactMojoItem Response Theory 101www.impactmojo.in
One item's curve across the trait range
Probability of a correct response rises with θ (item at b=0)
Illustrative logistic ICC (a=1, b=0)
At θ = 0 the chance is 50%. Move up the scale and success becomes near-certain; move down and it fades to near zero.
ImpactMojoItem Response Theory 101www.impactmojo.in
Why this fixes CTT's headaches
  • Item properties are sample-independent: the curve describes the item itself, not the group that sat it (when the model fits)
  • Person estimates are test-independent: θ means the same thing whatever items you used
  • Items & people share a metric: you can match a test to a person, link forms, and build short or adaptive tests
These gains hold only when the model's assumptions hold — a promise we will scrutinise later.
PropertyHolds whenFails when
Sample-free item parametersThe model fitsMultidimensionality, DIF
Test-free person scoresItems are on one calibrated scaleForms never linked
Common metricAnchors are stableAnchor items drift
These properties are frequently stated as facts about IRT. They are conditional guarantees, and the conditions are exactly what Section 9 checks.
Invariance is also only up to a linear transformation. Two separate calibrations produce scales that must be linked before their thetas can be compared.
ImpactMojoItem Response Theory 101www.impactmojo.in
The same idea for attitudes & experiences
For a food-insecurity or empowerment scale there is no 'right' answer — the curve gives the probability of endorsing the item ('yes, we worried about food'). Severe items are endorsed only by households high on the insecurity continuum.
'Could you visit a health centre alone?' and 'did you go a whole day without eating?' are items placed at very different points on their respective scales.
Ability itemEndorsement item
Correct / incorrectEndorsed / not endorsed
b is difficultyb is severity
Guessing can inflate the floorNo guessing; social desirability instead
Higher theta = more ableHigher theta = more of the condition
The mathematics is identical; only the interpretation changes. That is why the same software fits a reading test and a food-insecurity scale.
Social desirability is the endorsement-scale analogue of guessing, and it is worse behaved: it shifts by interviewer, setting and group rather than sitting at a fixed floor.
ImpactMojoItem Response Theory 101www.impactmojo.in
What the rest of the course unpacks
Everything that follows is detail on that one S-curve: where it sits (difficulty, b), how steep it is (discrimination, a), whether it has a floor (guessing, c), how much information it carries, and whether it behaves the same for everyone (DIF).
Section aheadThe question it answers
ICCHow do I read an item's behaviour?
Difficulty (b)Where does this item work?
Discrimination (a)How sharply does it sort people?
Guessing (c)Is there a floor under the curve?
InformationWhere is my test precise?
Validation and DIFAm I entitled to any of this?
Everything is a property of one S-curve, which is why the course spends so long on a single picture. Learn to read it and the rest is bookkeeping.
The last row is not an afterthought. Sections 1 to 8 describe what IRT offers; Section 9 is what you must do to be allowed to claim it.
ImpactMojoItem Response Theory 101www.impactmojo.in
04
Section Four
The Item Characteristic Curve
ImpactMojoItem Response Theory 101www.impactmojo.in
The Item Characteristic Curve (ICC)
Item Characteristic Curve (ICC)
The graph of the probability of a correct / endorsing response (y, 0 to 1) against the latent trait theta (x). One curve per item; its shape encodes everything the model says about that item.
Also called the item response function. Read it left to right: as the person's trait increases, the chance of success increases.
Element of the ICCParameterReads as
Where it crosses 0.5bDifficulty / severity
Slope at that pointaDiscrimination
Lower asymptotecGuessing floor
Upper asymptote(d, rarely used)Careless-error ceiling
The fourth parameter exists in 4PL models and is rarely estimated in practice: it allows for high-ability respondents occasionally getting an easy item wrong.
For most development applications, b and a carry the substance and c is included only when the response format makes guessing possible.
ImpactMojoItem Response Theory 101www.impactmojo.in
Reading an ICC in three moves
  • Pick a θ on the x-axis — a person's trait level
  • Go up to the curve, then across — read the probability
  • The whole curve tells you how the item behaves for everyone
Two numbers define the basic curve: where it crosses 50% (difficulty) and how steeply it rises there (discrimination).
Read the ICC to answerBy
Is this item suitable for my group?Check whether b sits inside their theta range
Will it separate people?Look at the slope where they sit
Is it wasted?Flat curve across your whole range
Is it mis-keyed?Curve slopes downward
Reading a curve is a practical skill worth acquiring even if someone else fits the model. Most item problems are visible in the plot before any statistic flags them.
Always plot the empirical curve — observed proportions correct within theta bands — over the modelled one. Disagreement between the two is the fit check that matters.
ImpactMojoItem Response Theory 101www.impactmojo.in
The three landmarks of an ICC
Lower tail
P near 0 (or near c) for very low θ — the item is too hard for them
Inflection
the steepest point — where the item best separates people
Upper tail
P near 1 for very high θ — the item is trivially easy for them
LandmarkWhere information is
Lower tailAlmost none
Approaching bRising
At the inflectionMaximum
Upper tailAlmost none
This is the structural reason a test cannot measure everyone equally well: each item carries its information in a band, and a test is only as informative as the items you put where people are.
It also explains why adding easy items to a test for strong students changes almost nothing. They contribute near-zero information at that theta.
ImpactMojoItem Response Theory 101www.impactmojo.in
A single, well-behaved ICC
One item's ICC (a=1, b=0): monotonic, S-shaped, 0 to 1
Illustrative logistic ICC
The 50% point sits at θ = 0 here — that is this item's difficulty. The slope through that point is its discrimination.
ImpactMojoItem Response Theory 101www.impactmojo.in
An ICC only ever goes up
A valid ICC is monotonically increasing: more of the trait never lowers the chance of a correct response. If a fitted curve dips or wiggles, something is wrong — a mis-keyed item, a trick question, or a broken assumption.
A non-monotonic empirical curve is a red flag, not a finding. Investigate the item before trusting it.
Non-monotonic curve meansCheck
Mis-keyed itemThe answer key
Ambiguous wordingWhether strong respondents over-read it
Trick or double-barrelled itemThe item text
Reverse-coded and not recodedThe coding script
The last is the most common and the most embarrassing: a reverse-worded attitude item that was never recoded produces a cleanly downward curve and a negative discrimination.
Fix the coding and refit before concluding anything about the item. Only after that is a downward curve a finding about the item rather than about the data preparation.
ImpactMojoItem Response Theory 101www.impactmojo.in
Several items on one set of axes
Four items overlaid: easy/hard differ in position, steep/flat in slope
Illustrative ICCs
Overlaying ICCs is how psychometricians read an item bank at a glance — left/right tells you difficulty, steepness tells you discrimination.
ImpactMojoItem Response Theory 101www.impactmojo.in
From curves to a person's estimate
Given a person's pattern of right and wrong answers and the fitted ICCs, the software finds the θ that makes that pattern most likely. That θ — not the raw count — is the person's IRT score.
01
Item ICCs (a, b, c) estimated
02
Person's response pattern observed
03
Find θ that best fits the pattern
04
θ with a standard error = the score
Estimation stepWhat it does
CalibrationEstimate item parameters from a large sample
ScoringEstimate each person's theta given those parameters
Standard errorQuantify uncertainty at that theta
ReportingTranslate theta into something actionable
Maximum likelihood is the usual scoring method and has a known failure: a person who gets every item right or every item wrong has no finite estimate.
Bayesian estimators (EAP, MAP) handle those cases by borrowing from a prior distribution, at the cost of shrinking estimates toward the mean. Know which your software used.
ImpactMojoItem Response Theory 101www.impactmojo.in
The same raw score, different meaning
Two children each get 6 of 12 right. One passed the six easy items; the other passed six hard ones and missed easy items. CTT calls them equal. IRT, reading which items, places them at different θ.
The response pattern carries information that the raw total throws away.
Two children, 6 of 12IRT reads
Passed the six easiestTheta near the middle of the easy items
Passed six hard, missed easyAn unusual pattern — check for problems
Mixed, consistent with difficulty orderA well-behaved estimate
The second row is worth pausing on. IRT does not simply score that child higher; it flags the response pattern as improbable under the model, which is diagnostic information CTT discards.
Person-fit statistics formalise this. Aberrant patterns can indicate guessing, cheating, a mis-administered form, or an item set that does not fit that child.
ImpactMojoItem Response Theory 101www.impactmojo.in
05
Section Five
Item Difficulty — the b Parameter
ImpactMojoItem Response Theory 101www.impactmojo.in
b locates the item on the trait scale
Difficulty (b)
The point on the theta scale where the item's success probability is 50% (for a 2PL/1PL model). It is measured in the same units as theta, so an item and a person can be compared directly.
Mantra: higher b = harder item. A hard item's curve sits to the right; you need more trait to have a 50–50 chance.
b valueItem isUseful for
-2.0Very easySeparating the weakest
-1.0EasyLower half
0.0ModerateThe middle
+1.0HardUpper half
+2.0Very hardSeparating the strongest
Note the mantra carefully: higher b means harder. It reads backwards to people used to CTT p-values, where a higher number means easier.
The 50% definition holds for 1PL and 2PL. Under a 3PL with guessing, b is where the probability is halfway between c and 1, which is above 0.5.
ImpactMojoItem Response Theory 101www.impactmojo.in
Easy, medium and hard items
Three items, same a, different b: the curve slides right as b rises
Illustrative ICCs (a=1; b = −1, 0, +1)
Read off the 50% line: the easy item crosses it at θ = −1, the hard item at θ = +1. That crossing point is b.
ImpactMojoItem Response Theory 101www.impactmojo.in
b and θ share one scale
Because b is expressed in θ units, you can ask: is this person above or below this item? A child at θ = 0.5 is comfortably above an easy item (b = −1) but below a hard one (b = +1).
This shared metric is what lets you target a test — choose items whose b's surround the people you most need to measure.
Person theta vs item bExpected outcome
Well above bVery likely correct; little information
Slightly above bLikely correct; high information
At b50/50; maximum information
Well below bVery likely wrong; little information
The shared metric is what makes test targeting possible: you choose items whose b values sit where your respondents are, rather than where the item bank happens to be dense.
It is also what makes an adaptive test possible, since the machine can compare the current theta estimate with each candidate item's b directly.
ImpactMojoItem Response Theory 101www.impactmojo.in
Ordering items by difficulty
Item (early-grade reading)b (illustrative)Interpretation
Recognise a letter−2.0Very easy — most children pass
Read a familiar word−0.8Easy
Read a simple sentence0.2Moderate
Read a short paragraph1.0Hard
Answer a comprehension question1.8Very hard
Illustrative values. The ordering — letter to comprehension — is what an ASER-style ladder captures, and b puts numbers on it.
If your sample sits atThese items are
theta about -1.5Letter and word items informative; paragraph wasted
theta about 0Sentence and paragraph items informative
theta about +1Comprehension items informative; letters wasted
An assessment used across a wide age or grade range needs items spread across the range, or it will measure one grade well and the others poorly.
This is precisely the design logic behind ASER-style tools, where a child stops at the level they cannot do — a manual approximation of adaptive testing.
ImpactMojoItem Response Theory 101www.impactmojo.in
For attitude scales, b means 'severity'
On a food-insecurity scale, b is better read as severity: how much insecurity it takes before a household endorses the item. 'We worried about food' has a low b; 'a household member went a whole day without eating' has a high b.
A well-built scale spreads its items' b's from mild to severe, so it can place households all along the continuum.
FIES-style itemSeverity
Worried about foodLow
Unable to eat healthy foodLow to moderate
Ate less than you thought you shouldModerate
Ran out of foodHigh
Went a whole day without eatingHighest
The severity ordering is empirical and turns out to be remarkably stable across countries, which is what allows a common global scale — and is a genuine finding, not an assumption.
Where the ordering does not replicate in a particular country or language, that is a signal about translation or about a different lived experience of scarcity, and both are worth investigating.
ImpactMojoItem Response Theory 101www.impactmojo.in
Mapping items along the trait scale
Item-difficulty map: where five items 'bite' on the θ scale
Illustrative b values on the θ axis
Gaps in the map are blind spots: between b = −0.8 and 0.2 the test thins out, so it measures children there less well.
ImpactMojoItem Response Theory 101www.impactmojo.in
Difficulty is about the item, not the topic
b is empirical: it is set by how people actually respond, not by how hard the item looks. A question that seems advanced may be easy if everyone was taught it; a 'simple' item may be hard if the wording confuses people.
Never assign difficulty by intuition. Estimate it from data, then sanity-check the surprises — they often reveal a flawed item.
Looks hard but may be easyLooks easy but may be hard
A technical term everyone was taughtA word that is unfamiliar in the local dialect
A long question with clear structureA short question with ambiguous wording
An advanced topic recently coveredAn old topic taught years ago
Difficulty is estimated from responses, so it absorbs everything about the item as administered — the wording, the translation, the layout, the order it appeared in.
This is why b estimated in a pilot in one language does not automatically transfer to a translation. The translated item is, empirically, a different item until shown otherwise.
ImpactMojoItem Response Theory 101www.impactmojo.in
06
Section Six
Item Discrimination — the a Parameter
Coming upIn one line
a definedThe slope of the curve at its midpoint
Steep vs flatHow sharply the item separates people
Low and negative aDead weight, and broken items
a with bWhere it works, and how well it works there
Discrimination is the parameter that decides how much an item is worth, and it is the one CTT approximates poorly with the item-total correlation.
It is also the parameter Rasch deliberately fixes, which is why the choice between Rasch and 2PL is really a choice about this section.
ImpactMojoItem Response Theory 101www.impactmojo.in
a measures how sharply an item sorts people
Discrimination (a)
How steeply the ICC rises at its midpoint — how well the item distinguishes people just below its difficulty from those just above it. Higher a means a steeper curve and a sharper distinction.
Mantra: higher a = steeper ICC = better at separating low-trait from high-trait people near its b.
a valueCurveItem is
Below 0.5Nearly flatContributing little
0.5 to 1.0GentleWeak but usable
1.0 to 2.0SteepGood
Above 2.0Very steepExcellent, but narrow
NegativeDownwardBroken — investigate
These bands are conventions from the psychometric literature and depend on the scaling constant used. Software differs on whether the logistic constant 1.7 is applied, which rescales a.
Check your software's parameterisation before comparing a values with a published table, or you will be comparing numbers on two different scales.
ImpactMojoItem Response Theory 101www.impactmojo.in
A steep item and a flat item
Same difficulty (b=0), different discrimination: steep vs flat
Illustrative ICCs (b=0; a = 1.8 vs 0.5)
Both cross 50% at θ = 0 (same b). The steep item leaps from low to high probability over a narrow band — it sorts people sharply right there. The flat item barely distinguishes anyone.
ImpactMojoItem Response Theory 101www.impactmojo.in
A steep item is a sharp ruler
Near its difficulty, a high-a item changes a person's success probability a lot for a small change in θ. That sensitivity is exactly what lets it tell two nearby people apart — it carries more information (next section).
High discrimination is generally desirable — but only near that item's b. Away from b, even a steep item is flat and tells you little.
High a gives youAt the cost of
Sharp separation near bA narrow useful band
High information at one pointAlmost none elsewhere
Efficient short testsSensitivity to model misfit
Information rises with the square of a, so discrimination pays off disproportionately — which is why item selection by information favours steep items so strongly.
That is also the trap: an automated selection routine will build a test entirely of steep items clustered at one difficulty unless you constrain the b spread.
ImpactMojoItem Response Theory 101www.impactmojo.in
When an item barely sorts anyone
A low-a (flat) item gives almost the same success probability to people across a wide range of θ. Low-trait and high-trait people answer it similarly — so it adds little to distinguishing them.
Very low or negative a is a warning: the item may be ambiguous, off-topic, or mis-keyed. Items that do not discriminate are candidates for revision or removal.
Low a usually meansCheck
The item taps a different traitDimensionality
The wording is ambiguousThe item text and translation
Everyone guessesWhether c should be modelled
The item is off-targetWhether b is far from the sample
An item can show low discrimination simply because nobody in the sample sits near its difficulty. That is a sampling problem, not an item problem, and dropping the item would be wrong.
Look at the empirical curve before deleting. A flat line because there are five respondents in that band looks identical to a genuinely non-discriminating item.
ImpactMojoItem Response Theory 101www.impactmojo.in
A curve that goes the wrong way
If an item shows negative discrimination, higher-trait people do worse on it than lower-trait people — the ICC slopes downward. That should never happen for a sound item.
Usual culprits: the answer key is wrong, the item measures something else, or it is a trick question. Fix or drop it — do not leave it scoring people backwards.
Negative a causeFix
Mis-keyed answerCorrect the key, refit
Reverse-worded, not recodedRecode, refit
Item measures the opposite constructReconsider or drop
Data entry column shiftCheck the merge
Treat a negative discrimination as a data problem until proven otherwise. In practice the great majority are keying or coding errors rather than substantive findings.
If the item survives all four checks and still slopes downward, it is telling you something real about the construct — and that is worth a conversation, not a deletion.
ImpactMojoItem Response Theory 101www.impactmojo.in
Difficulty and discrimination are independent
Low a (flat)High a (steep)
Low b (easy)Easy, sorts weaklyEasy, sorts low-trait people sharply
High b (hard)Hard, sorts weaklyHard, sorts high-trait people sharply
b says where the item works; a says how well it works there. A good item bank mixes b's to cover the range and favours high a's at each level.
Low aHigh a
Low bEasy, sorts weaklyEasy, sorts sharply at the low end
High bHard, sorts weaklyHard, sorts sharply at the high end
Because the two parameters are independent, an item bank is best described as a scatter of a against b rather than by average values of either.
The plot to draw when reviewing a bank: b on the x-axis, a on the y-axis, one point per item. Gaps along x are coverage gaps; points near the bottom are dead weight.
ImpactMojoItem Response Theory 101www.impactmojo.in
Discrimination on an empowerment scale
On an empowerment scale, a high-a item is one whose endorsement cleanly separates more- from less-empowered women near its severity. A low-a item — one answered similarly regardless of empowerment — is dead weight.
Selecting high-a items is how scale developers shorten a questionnaire without losing measurement quality.
Empowerment itemLikely behaviour
"Can go to the market alone"Discriminates in restrictive settings
"Has a say in her own healthcare"Widely endorsed; low severity
"Owns land in her own name"High severity, low endorsement
"Participates in household decisions"Often low a — everyone says yes
The last row is a real and well-documented problem with agency items: broad, socially desirable statements attract near-universal agreement and therefore carry almost no information.
Specific, behavioural, recent-past items behave better than general attitudinal ones — "who decided the last major purchase" rather than "do you participate in decisions".
ImpactMojoItem Response Theory 101www.impactmojo.in
Steeper is not always better
An extremely steep item measures superbly — but only in a razor-thin band of θ. A test built only of very steep items can measure one narrow region brilliantly and everywhere else poorly.
Balance matters: you want high discrimination spread across the range you care about, not piled at one point.
Test built only ofResult
Very steep items at one bSuperb at one point, useless elsewhere
Flat items across a wide rangeBroad but imprecise everywhere
Steep items spread across bThe usual target
Which of these you want depends entirely on the decision. For a pass/fail classification at one cut-score, the first is actually correct and efficient.
For describing a population distribution, the third is right. Specify the purpose before selecting items, or the selection criterion is arbitrary.
ImpactMojoItem Response Theory 101www.impactmojo.in
07
Section Seven
Guessing & the Model Family
Coming upIn one line
The guessing floorNobody scores zero on multiple choice
Why it mattersIgnoring it biases difficulty downward
1PL, 2PL, 3PLEach adds one parameter and one data requirement
ChoosingStart simple; add only what the data supports
This section is where model choice gets decided, and the decision is usually made by sample size rather than by theory.
Most development instruments end at the 2PL. The 3PL exists for multiple-choice ability testing and adds little elsewhere.
ImpactMojoItem Response Theory 101www.impactmojo.in
On multiple-choice items, no one scores zero
On a 4-option multiple-choice item, even a child who knows nothing has roughly a 1-in-4 chance of being right. So the ICC should not fall to zero at low θ — it should flatten out at a lower asymptote.
Guessing (c)
The lower asymptote of the ICC — the success probability for someone with very low theta. It is the floor created by guessing or by partial cues.
FormatChance floorModel c?
4-option multiple choiceAbout 0.25Yes
True / falseAbout 0.5Yes, and it dominates
Open responseAbout 0No
Yes / no experience itemNot applicableNo
True/false items are a poor format for exactly this reason: half the response is chance, so the item carries little information and c is hard to separate from b.
The estimated c often comes out below the nominal chance rate, because weak respondents can usually eliminate one implausible option rather than guessing blindly.
ImpactMojoItem Response Theory 101www.impactmojo.in
A guessing floor lifts the lower tail
With guessing (c=0.25) the curve floors near 0.25, not 0
Illustrative ICCs (a=1, b=0; c=0 vs c=0.25)
Note the amber curve never drops below ~0.25. A very low-ability student still has a one-in-four chance — that is the guessing floor.
ImpactMojoItem Response Theory 101www.impactmojo.in
Ignoring guessing biases difficulty
If you fit a model with no guessing parameter to multiple-choice data, the floor created by guessing gets misread as the item being 'easier' than it is — biasing b and distorting low-ability scores.
c matters most at the bottom of the scale — precisely where many development assessments most need to measure well.
Ignoring guessing causesVisible as
b biased downwardItems look easier than they are
a biasedDiscrimination understated at the low end
Low-theta scores inflatedNobody appears to be at the bottom
Poor fit in the lower tailEmpirical curve above the model
The consequence for a development programme is specific: children who know very little are estimated as knowing something, which understates the size of the learning gap.
The lower-tail fit plot is the diagnostic. If observed proportions correct sit well above the modelled curve at low theta, guessing is unmodelled.
ImpactMojoItem Response Theory 101www.impactmojo.in
1PL, 2PL, 3PL: what each one adds
ModelFree parametersAdds…
1PL / Raschb only (a fixed equal)Difficulty differs; all items equally discriminating
2PLa and bItems also differ in discrimination
3PLa, b and cPlus a guessing floor
Each step adds realism — and demands more data to estimate the extra parameters reliably.
ModelEstimatesNeeds
1PL / Raschb onlySmallest sample; strongest assumption
2PLa and bModerate sample
3PLa, b and cLarge sample; c often fixed
Each added parameter buys flexibility and costs data and stability. The 3PL's c is notoriously hard to estimate and is frequently fixed at the reciprocal of the number of options.
More parameters will always fit the calibration sample better. The question is whether the parameters replicate in a new sample, which is what cross-validation checks.
ImpactMojoItem Response Theory 101www.impactmojo.in
The Rasch model: every item equally discriminating
The 1PL / Rasch model fixes a to be the same for every item, so items differ only in difficulty (b). The ICCs are parallel S-curves — same shape, shifted left or right.
Why people love it
  • Simple, stable, needs less data
  • Raw score is a sufficient statistic
  • Elegant measurement properties
The trade-off
It assumes equal discrimination. Where items genuinely differ in a, Rasch will misfit some of them.
Rasch is preferred becauseRasch is criticised because
Sufficient statistics: raw score maps to thetaIt assumes equal discrimination
Specific objectivity propertiesItems that misfit are dropped, not modelled
Stable with modest samplesThe assumption is often empirically false
Simple to explain and defendFit is achieved by discarding data
The disagreement is philosophical, not merely technical. Rasch practitioners treat the model as a standard the data must meet; 2PL practitioners treat the model as something to fit to the data.
Both positions are defensible and produce different workflows. Know which one you are in, because it determines whether a misfitting item is a problem with the item or with the model.
ImpactMojoItem Response Theory 101www.impactmojo.in
The 2PL model: let discrimination vary
The 2PL frees a, so each item has its own slope as well as its own difficulty. ICCs can now be steep or flat, crossing one another. It fits more datasets but needs more respondents.
Use 2PL when items plainly differ in how well they sort people — common for attitude and empowerment scales.
Use 2PL whenSigns it is needed
Items plainly differ in qualityWide spread of item-total correlations
You have the sample sizeSeveral hundred respondents
Rasch shows systematic misfitEmpirical curves steeper or flatter than modelled
The practical test is to fit both and compare. If the 2PL a values cluster tightly around one number, Rasch was adequate and is the simpler defensible choice.
If they range from 0.4 to 2.5, forcing equal discrimination is distorting the scale, and Rasch fit will have been bought by deleting the informative items.
ImpactMojoItem Response Theory 101www.impactmojo.in
The 3PL model: add a guessing floor
The 3PL adds c, the guessing asymptote — built for multiple-choice ability tests where lucky guesses are real. It is the most flexible of the three but the hungriest for data and the trickiest to estimate.
c is notoriously hard to pin down — it lives in the sparsely-populated low-θ tail. Large samples are essential, or c is often fixed to 1/(number of options).
3PL practicalitiesWhat people do
c is unstable to estimateFix it at 1/options
Needs large samplesReserve for high-stakes testing
Trades off with bConstrain c with a prior
Not needed without guessingDo not use it on yes/no scales
Fixing c is not a fudge; it is a standard and defensible choice, and it is more honest than reporting an estimated c with an enormous standard error.
Most development applications do not need the 3PL at all, because most instruments are constructed response or yes/no experience items where guessing does not arise.
ImpactMojoItem Response Theory 101www.impactmojo.in
Which model should you use?
  • Start simple. Rasch/1PL if items are similar and data is limited — many large-scale assessments use it deliberately
  • Move to 2PL when discrimination clearly varies and you have the sample size
  • Reserve 3PL for multiple-choice ability tests with guessing and very large samples
  • Let fit and theory decide — not the wish for the fanciest model
SituationModel
Yes/no experience scale, 300 respondentsRasch
Ability test, varied item quality, 8002PL
Multiple-choice, high stakes, 2,000+3PL
Ordered categories (never/sometimes/often)Graded response or partial credit
Genuinely several traitsMultidimensional, or separate scales
Start simple and add parameters only when the data demands it and supports it. A well-fitting Rasch model beats a poorly estimated 3PL every time.
The last row is a reminder that the answer is sometimes not a bigger IRT model but two scales, reported separately, because the construct was never one thing.
ImpactMojoItem Response Theory 101www.impactmojo.in
08
Section Eight
Information & Precision
Coming upIn one line
Item informationAn item informs most near its own difficulty
Test informationItem information simply adds up
Standard errorSE at any theta is 1 over the square root of information
TargetingPut the information where the decision is
Information is the concept that turns IRT from a description of items into a design tool: it tells you what a test will do before anyone sits it.
It is also the honest replacement for a single reliability figure, because it reports precision as something that varies rather than as one number.
ImpactMojoItem Response Theory 101www.impactmojo.in
Information = measurement precision
Item information
How much an item reduces uncertainty about theta at each point on the scale. An item is most informative around its own difficulty b, and more so the higher its discrimination a.
Information is the IRT replacement for one blanket reliability figure: it tells you where on the scale the item measures well.
Information rises withBecause
Discrimination a (squared)A steeper curve separates more sharply
Proximity of theta to bThe item is at its most uncertain, so most informative
Number of items nearbyInformation is additive
Absence of guessingA floor flattens the lower half
The additivity of information across items is what makes test construction tractable: you can build a target information function by choosing items, without refitting anything.
It also means one excellent item can outweigh several mediocre ones at the same location, which is the mathematical basis for shortening instruments without losing precision.
ImpactMojoItem Response Theory 101www.impactmojo.in
An item informs most around its difficulty
Item information function (item at b=0): peaks at its difficulty
Illustrative item information (a=1, b=0)
The peak sits at θ = 0 — the item's b. An item tells you most about people whose trait is near its difficulty, and little about those far from it.
ImpactMojoItem Response Theory 101www.impactmojo.in
Add items, add information
The test information function is just the sum of the item information functions. Where many items pile up, the test measures precisely; where items are sparse, it measures poorly.
Test information (3 items at b = −1, 0, +1): broad, peaks near 0
Illustrative test information function
ImpactMojoItem Response Theory 101www.impactmojo.in
More information means less error
Precision and information are two sides of one coin: the standard error of measurement at any θ is 1 / √information. High information → small standard error; low information → large error.
Standard error varies along θ — lowest where information is highest
Illustrative SEM = 1 / √(test information)
ImpactMojoItem Response Theory 101www.impactmojo.in
Precision is not constant — and IRT shows it
CTT
One standard error for everyone — pretends the test is equally precise across the whole range.
IRT
Error is lowest where information peaks (usually mid-range) and rises sharply in the tails. Honest, and actionable.
If you must measure the very weak or very strong precisely, you need items targeted there — the middle-heavy test will fail them.
CTTIRT
Error reportOne SEM for everyoneSE as a function of theta
At the extremesUnderstatedCorrectly large
In the middleOverstatedCorrectly small
ConsequenceMisleading confidence intervalsHonest ones
This matters most where the decision is at an extreme. Identifying children far below grade level using CTT confidence intervals will overstate how confidently they have been identified.
Report the standard error alongside theta wherever the number will drive a decision about an individual, and never report a proficiency classification without it.
ImpactMojoItem Response Theory 101www.impactmojo.in
Targeting a test where it matters
Two tests, same length: a 'high-θ' form moves the information peak right
Illustrative test information (items centred at θ≈0 vs θ≈+1.5)
Choosing items by their b is how you target a form — e.g. to grade top performers, or to pinpoint a pass/fail cut-score.
ImpactMojoItem Response Theory 101www.impactmojo.in
Put your information where the decision is
If a programme classifies children as 'at grade level' or not, the decision happens at one θ — the cut-score. That is exactly where you want maximum information and minimum error.
Pack items with b's near the cut-score. A test can be short yet decisive if its information is concentrated where the call is actually made.
Decision typeWhere to put information
Pass / fail at one cutConcentrate items near the cut
Three proficiency bandsPeaks at each of the two cuts
Describe the whole distributionSpread evenly across the range
Select the top 10%Concentrate high on the scale
This is the single most useful practical application of the information function, and it is the one most often skipped: instruments are usually assembled by topic coverage rather than by where the decision falls.
Coverage and targeting can conflict. Where they do, say which you chose and why, because the resulting instrument measures accurately in a particular place and not everywhere.
ImpactMojoItem Response Theory 101www.impactmojo.in
How computer-adaptive tests use information
Because items and people share a scale, a computer-adaptive test can pick each next item to be maximally informative at the test-taker's current θ estimate — honing in fast.
01
Estimate θ so far
02
Pick the most informative unused item near that θ
03
Update θ from the answer
04
Stop when the standard error is small enough
Adaptive testing needsWhich is why
A calibrated item bankItems must already be on one scale
Many items per difficulty bandExposure must be controlled
Delivery technologySelection happens between items
A stopping ruleEither fixed length or target precision
Adaptive testing is the clearest payoff from putting people and items on one scale, and it typically reaches the same precision in far fewer items than a fixed form.
The barrier in most development settings is not the algorithm but the item bank: calibrating enough items across the range is a multi-year investment.
ImpactMojoItem Response Theory 101www.impactmojo.in
09
Section Nine
Validating Scales & DIF
Coming upIn one line
Model fitDo the responses match the fitted curves?
UnidimensionalityIs it one trait, or several?
Local independenceDo items lean on each other?
DIFDoes an item behave differently by group?
This is the section that earns everything claimed in Sections 3 to 8. Without it, IRT output is a sample-dependent scale with more decimal places than CTT.
It is also the section most often skipped under deadline, which is why so many published theta scales carry properties they have not demonstrated.
ImpactMojoItem Response Theory 101www.impactmojo.in
IRT's promises hold only if assumptions do
Sample-free items, test-free scores, a clean common metric — these gifts depend on the model actually fitting the data. Validation is the work of checking that it does.
  • Does the model fit the responses?
  • Is the scale unidimensional?
  • Are responses locally independent?
  • Does each item behave the same across groups (no DIF)?
PromiseDepends on
Sample-free item parametersModel fit and no DIF
Test-free person scoresCommon calibration or linking
Meaningful information functionCorrect model choice
Comparability across groupsDIF screening
Validation is not a final quality check; it is the step that entitles you to the claims. Skipping it and reporting theta anyway is reporting a number with unearned properties.
Budget for it. In a real instrument-development cycle, validation and revision take more time than the initial item writing.
ImpactMojoItem Response Theory 101www.impactmojo.in
Checking model fit
Fit asks whether the observed responses match what the fitted ICCs predict — overall and item by item. A badly fitting item's empirical curve departs from its modelled S-curve.
Use item-fit statistics and, crucially, plot the empirical vs modelled ICC. A picture catches misfit that a single index can hide — and points to the offending item.
Fit checkWhat it catches
Item-fit statisticsItems whose responses depart from the model
Empirical vs modelled ICC plotWhere and how they depart
Person-fit statisticsAberrant response patterns
Residual analysisSystematic misfit across the bank
Item-fit chi-square statistics are sensitive to sample size: with 5,000 respondents almost every item will misfit significantly, and with 200 almost none will.
The plot is therefore more informative than the p-value. Look at the size and location of the departure, not only at whether it is statistically significant.
ImpactMojoItem Response Theory 101www.impactmojo.in
Assumption 1: one trait at a time
Unidimensionality
The assumption that a single latent trait accounts for the responses. The items should all tap the same underlying continuum — one theta, not several.
A 'reading' test that secretly also measures vocabulary and reasoning violates this. Check with factor analysis before trusting a one-dimensional θ.
Dimensionality checkWarning sign
Factor analysis of the itemsMore than one substantial factor
Principal components of residualsA structured first residual component
Content reviewItems about visibly different things
Fit across subscalesBetter fit when split
Strict unidimensionality never holds exactly. The practical question is whether the departure is large enough to distort the parameters and the scores, which is a judgement, not a test.
If a reading test also measures general knowledge, the resulting theta blends the two, and a child strong in one and weak in the other is placed somewhere meaningless.
ImpactMojoItem Response Theory 101www.impactmojo.in
Assumption 2: items don't lean on each other
Local independence
Once you account for theta, responses to different items are independent. Knowing the answer to one item should give no extra clue to another, beyond what theta already explains.
Violated by item chains — e.g. several questions about one reading passage, where missing the passage sinks them all together. Bundle or rewrite such items.
Local dependence sourceExample
Shared stimulusFive questions on one passage
Chained itemsItem 4 gives away item 5
Repeated wordingNear-duplicate items
Order effectsAn item that primes the next
The commonest case in reading assessment is the testlet: several items on one passage, which are dependent because a child who understood the passage gets them all.
The consequence is inflated reliability and information — the test looks more precise than it is, because the dependent items are counted as independent evidence.
ImpactMojoItem Response Theory 101www.impactmojo.in
Differential Item Functioning (DIF)
Differential Item Functioning (DIF)
When two people with the SAME theta but from different groups (e.g. girls vs boys, one region vs another) have different probabilities of answering an item correctly. The item behaves differently across groups.
DIF is potential item bias. Same trait, different odds — the item is reading something other than the trait for one group.
DIF isDIF is not
Different P at the same thetaA group difference in theta
A property of an itemA property of the group
Evidence of a nuisance factorProof of intent
Detectable statisticallyInterpretable without content review
The definition conditions on theta, which is the whole point: it isolates item behaviour from real differences in the trait, and that conditioning is what makes DIF meaningful.
Uniform DIF shifts the curve; non-uniform DIF changes its slope, so the item favours one group at some theta levels and the other elsewhere. The second is harder to spot and harder to fix.
ImpactMojoItem Response Theory 101www.impactmojo.in
DIF makes one item's ICC split by group
Same item, two groups: the ICCs diverge → DIF
Illustrative ICCs for one item across two groups
At any given θ, Group B's success probability is lower — the item is effectively harder for them at the same trait level. That is uniform DIF.
ImpactMojoItem Response Theory 101www.impactmojo.in
DIF is not the same as a group difference
If girls genuinely have higher reading ability than boys, they will score higher — that is a real trait difference, not DIF. DIF is when girls and boys at the same ability still differ on a specific item.
DIF conditions on θ. It isolates item bias from true group differences — a distinction CTT cannot make cleanly.
ObservationDIF?
Girls score higher overallNo — a trait difference
Girls at the same theta do better on item 7Yes
One region scores lowerNo, by itself
One region at the same theta fails item 12Yes
Conflating the two is the most common error in reading DIF output, and it runs in a damaging direction: real group differences get explained away as item bias, or genuine bias gets dismissed as a real gap.
Statistical DIF also does not establish unfairness on its own. It flags an item for content review, and the review decides whether the cause is irrelevant to the construct.
ImpactMojoItem Response Theory 101www.impactmojo.in
What to do when an item shows DIF
  • Investigate the content: a word, context or example unfamiliar to one group (an urban example, a gendered scenario)
  • Check the translation: DIF across language versions often signals a poor or unequal translation
  • Revise or remove biased items before reporting scores
  • Document the DIF review — fairness is part of validity
DIF causeAction
Unfamiliar context or exampleRewrite with a neutral context
Poor or unequal translationRetranslate and back-translate
Curriculum differs by regionMay be legitimate — document it
Gendered scenarioRewrite
No identifiable causeConsider dropping; report the decision
The third row matters: not all DIF is bias. If one state genuinely teaches a topic and another does not, the item functions differently for a reason that may be exactly what you want to detect.
Whatever you decide, record it. A dropped item and a retained flagged item are both defensible; an undocumented deletion is not.
ImpactMojoItem Response Theory 101www.impactmojo.in
10
Section Ten
Applications in Development
ImpactMojoItem Response Theory 101www.impactmojo.in
ASER- and NAS-style learning assessments
Large learning assessments — India's NAS, citizen-led ASER, and cross-national studies — lean on IRT to place children on a single proficiency scale and to compare across grades, states and years.
IRT is what lets 'reading at grade-2 level' mean the same thing whether a child sat form A or form B, this year or last.
AssessmentWhat IRT enables
NASCommon scale across grades and states
ASERA simple ordered tool with a scaling logic behind it
Cross-national studiesComparison across languages, via DIF screening
Programme endlinesComparison with a different form
Cross-national comparison is where the assumptions are stretched hardest: an item must function equivalently across languages and curricula for a common scale to mean anything.
This is why those studies publish extensive DIF and invariance analyses, and why critics focus on exactly those analyses when contesting the rankings.
ImpactMojoItem Response Theory 101www.impactmojo.in
Equating: comparing across forms and years
Equating / linking
Placing different test forms onto a common theta scale — usually via shared 'anchor' items — so scores from different forms or years are directly comparable.
Because IRT item parameters are (when the model fits) sample-independent, anchor items let you stitch separate forms into one continuous scale.
Linking designHow it works
Common items (anchors)Shared items across two forms
Common personsThe same people sit both forms
Single-groupOne group, both forms, counterbalanced
Anchor design is the usual choice in large-scale assessment because it does not require anyone to sit two tests. Its quality depends entirely on the anchor set.
Anchors should span the difficulty range and be numerous enough — a rule of thumb is at least 20% of the test or 20 items, whichever is larger.
ImpactMojoItem Response Theory 101www.impactmojo.in
Measuring change without changing the ruler
To track learning over years you must change the questions (security, age-appropriateness) without changing what the score means. Equating via anchor items keeps the ruler fixed while the items rotate.
Without equating, a 'rise in scores' could just be an easier form. Equating separates real learning gains from changes in the test.
Equating riskSymptom
Anchor driftAn anchor's b changes between years
Too few anchorsUnstable linking constants
Anchors clustered in difficultyPoor linking at the extremes
Anchor exposureAnchors become known and easy
Anchor drift is the practical hazard in a repeated survey: an item becomes familiar, is taught to, or leaks, and its difficulty falls — which then shifts the whole scale.
Check anchor stability before equating, and drop drifted anchors from the link. A trend line built on a drifted anchor will show improvement that did not happen.
ImpactMojoItem Response Theory 101www.impactmojo.in
FIES: a global IRT-based scale
The Food Insecurity Experience Scale (FIES) — eight yes/no experience items — is modelled with a Rasch/1PL approach so that severity is comparable across countries and languages, underpinning SDG indicator 2.1.2.
8 items
from 'worried' to 'a whole day without eating'
Severity = b
items ordered from mild to severe insecurity
Equated
calibrated to a global reference scale
FIES design choiceReason
Eight yes/no itemsShort enough for a household survey module
Rasch / 1PLStable with modest samples; simple to defend
Experience items, not opinionLess prone to social desirability
Global reference scaleEnables cross-country comparability
FIES underpins SDG indicator 2.1.2, which is a good example of IRT doing quiet infrastructure work: the comparability of a global indicator rests on the measurement model.
Its equating step is the same machinery as Section 10: national scales are linked to a global standard through common items.
ImpactMojoItem Response Theory 101www.impactmojo.in
Empowerment and agency scales
Women's empowerment indices and agency scales combine items on mobility, decision-making and asset control. IRT checks whether they measure one coherent trait, ranks items by severity, and flags items that work differently across regions or castes.
It turns a bag of agree/disagree items into a calibrated scale — and reveals which items actually discriminate between more- and less-empowered women.
Empowerment scaling questionWhy it is hard
Is it one trait?Mobility, decisions and assets may not co-vary
Does severity order hold across settings?Norms differ sharply
Is DIF present across groups?Almost always, by caste, region, religion
Does the scale mean the same over time?Norms shift during a programme
The last row is a genuine measurement problem for evaluation: if a programme changes what women consider normal, the same item may become easier to endorse without any change in the underlying position.
This is response shift, and it can make an effective programme look ineffective or vice versa. Anchoring vignettes are one partial remedy; honesty about the limitation is the minimum.
ImpactMojoItem Response Theory 101www.impactmojo.in
Asset and wealth indices
Asset-based wealth indices (the kind behind NFHS/DHS wealth quintiles) ask whether a household owns particular assets. IRT and related latent-trait methods place households on a wealth continuum from the pattern of ownership.
A motorcycle and a mud floor sit at very different points on the wealth scale — just as easy and hard items sit at different b's. The logic is identical.
Asset index methodNote
Principal components (DHS standard)Not a latent-trait model; weights from PCA
IRT / latent traitModels the probability of ownership
BothSensitive to which assets are included
The DHS wealth index is built with principal components rather than IRT, which is worth knowing when the two are discussed together — they answer a similar question with different machinery.
Both share the same substantive weakness: an asset list calibrated for one context measures poorly in another, because what counts as a marker of wealth is local.
ImpactMojoItem Response Theory 101www.impactmojo.in
Attitude, stigma and knowledge scales
  • Health knowledge: grade items from basic to advanced and measure understanding precisely where a campaign targets it
  • Stigma / attitude scales: order statements by how much prejudice it takes to endorse them
  • Quality-of-life & depression screeners: many are now built and validated with IRT
Scale typeWhat IRT adds
Health knowledgeGrade items and target the campaign's level
Stigma / attitudesOrder statements by how much it takes to endorse
Quality of lifeCheck whether the domains are one trait
Depression / distress screenersPrecision at the clinical cut-point
Screeners are the clearest case for targeting: the whole purpose is a classification at one threshold, so information should be concentrated there rather than spread evenly.
Where an established scale exists and is validated in your setting, use it rather than writing your own. Comparability with the literature is usually worth more than a bespoke instrument.
ImpactMojoItem Response Theory 101www.impactmojo.in
What IRT gives a development programme
  • Shorter instruments: keep the most informative items, cut respondent burden
  • Comparable numbers: across forms, years, regions and languages
  • Fairer measures: DIF screening removes biased items
  • Targeted precision: measure best exactly where decisions are made
PayoffRealistic scale
Shorter instrumentsOften 30-50% fewer items at equal precision
Comparable numbersAcross forms and years, if anchored
Fairer measuresOnly for the groups you screened
Targeted precisionAt the decision point you specify
The third row carries the caveat that matters. DIF screening protects the groups you tested for; an item biased against a group you did not analyse is undetected and unreported.
Decide your DIF groups at design, not after collection, and make sure the sample supports them. Small groups cannot be screened, which is where bias is most likely to persist.
ImpactMojoItem Response Theory 101www.impactmojo.in
11
Section Eleven
Assumptions, Limits & Tools
ImpactMojoItem Response Theory 101www.impactmojo.in
IRT is powerful, not magic
  • Unidimensionality — one trait drives the responses
  • Local independence — items don't lean on each other
  • Correct model — the chosen 1PL/2PL/3PL actually fits
  • Monotonicity — more trait, higher success probability
Break an assumption and the elegant guarantees — sample-free items, comparable scores — quietly stop holding.
AssumptionConsequence if violated
UnidimensionalityTheta blends two traits; scores uninterpretable
Local independencePrecision overstated
Correct modelBiased b and distorted low-end scores
MonotonicityThe scale is not ordered as claimed
No DIFGroup comparisons are invalid
Each violation has a distinct signature and a distinct remedy, which is why "the model did not fit" is not a diagnosis. Find out which assumption failed.
The last row is the one with the largest consequence for development work, because most of what we do with these scales is compare groups.
ImpactMojoItem Response Theory 101www.impactmojo.in
IRT is data-hungry
ModelRough sample-size guidanceNote
1PL / Rasch~200+ respondentsMost forgiving
2PL~500+ respondentsEstimating a needs more data
3PL~1,000+ respondentsc is hard to estimate; often fixed
Illustrative rules of thumb only — needs depend on test length, item quality and how spread the sample is. Small, homogeneous samples can defeat even a simple model.
Sample-size driverEffect
Number of parameters per itemMore parameters, more data
Spread of theta in the sampleConcentrated samples estimate poorly
Number of itemsMore items stabilise person estimates
Sparse response categoriesRare categories need large samples
The guidance is genuinely rough and depends on the spread of ability as much as on the count. A thousand respondents all at the same theta will not calibrate a wide-range item bank.
Where the sample is small, fixing parameters from a published calibration is often better than estimating badly — provided you check fit and DIF against your own data.
ImpactMojoItem Response Theory 101www.impactmojo.in
Where IRT can mislead
  • Garbage items in, garbage scale out — IRT cannot rescue badly written items
  • A neat θ can hide a contested concept — 'empowerment' is political, not just psychometric
  • Multidimensional traits forced onto one scale lose meaning
  • Black-box scores are harder for non-specialists to interpret than a percentage
LimitWhat it means for you
Bad items give bad scalesInvest in item writing first
A neat theta can hide a contested conceptSay what "empowerment" means here
Forced unidimensionality loses meaningReport two scales if there are two
Precision is not validityA precise measure of the wrong thing
The second row is the one that matters most in this sector. IRT can tell you a set of empowerment items forms a coherent scale; it cannot tell you the scale measures empowerment.
That question is answered by content validity, by whether the items reflect what the people concerned mean by the term, and it belongs partly to them rather than to the psychometrician.
ImpactMojoItem Response Theory 101www.impactmojo.in
Tools for fitting IRT models
ToolGood forNote
R: mirtUni- & multidimensional IRT, all common modelsFree, powerful, well documented
R: ltm1PL/2PL/3PL for dichotomous & graded itemsFree, gentle entry point
R: TAM / eRmLarge-scale & Rasch modellingFree; TAM mirrors big assessments
Stata: irt suiteIRT within a familiar stats packageBuilt-in irt commands
jMetrik / IRTPROPoint-and-click psychometricsLower coding barrier
ToolReach for it when
R: mirtYou need most models, including multidimensional
R: ltmA gentle first pass on dichotomous items
R: TAM / eRmRasch and large-scale assessment workflows
R: difR / lordifDIF screening specifically
Stata / commercialInstitutional standard, high-stakes work
Whichever you use, check the parameterisation it reports — whether the logistic constant 1.7 is applied, and whether difficulty is reported as b or as an intercept.
Mixing parameter sets from two packages without converting is a real and easy mistake, and it produces a plausible-looking scale that is wrong.
ImpactMojoItem Response Theory 101www.impactmojo.in
A sensible IRT workflow
01
CHECK dimensionality & local independence
02
FIT the simplest defensible model (start Rasch)
03
EXAMINE item fit, a, b, (c) and information
04
SCREEN for DIF across key groups
05
REVISE items, then score & report θ with its error
Workflow stepDo not skip
DimensionalityIt invalidates everything downstream
Simplest defensible modelComplexity you cannot estimate is worse
Item fit and parametersPlot, do not only test
DIF screeningDecide the groups in advance
Revise, then scoreDo not score with items you would drop
The final step is worth stating plainly: calibrate and validate first, then score. Reporting thetas from a model you are still revising produces numbers that change under you.
Keep the calibration sample and the item parameters together, versioned. A theta means nothing without the parameter set that produced it.
ImpactMojoItem Response Theory 101www.impactmojo.in
Translating θ for non-specialists
A θ of 1.2 means nothing to a programme officer or a parent. Part of doing IRT well is translating the scale back into language people act on — proficiency bands, 'can read a paragraph', percentile, or a clear cut-score.
Always report the standard error alongside the score, and describe what the cut-points mean in real-world terms. A precise number nobody understands helps no one.
Instead ofReport
"Theta = 1.2""Can read a short paragraph with comprehension"
"Mean theta rose 0.3""12% more children reached the paragraph level"
"SE = 0.35""Her level is between sentence and paragraph"
"Above the cut-score""At grade level, with the uncertainty stated"
Proficiency bands are the standard translation and require a standard-setting exercise — a structured judgement process about where on the theta scale each descriptor begins.
Standard setting is a judgement, not a calculation, and should be documented as such. Presenting a cut-score as though it were derived from the data misrepresents where it came from.
ImpactMojoItem Response Theory 101www.impactmojo.in
Where to go deeper
  • Item Response Theory for Psychologists — Embretson & Reise (the standard, readable introduction)
  • The Theory and Practice of Item Response Theory — de Ayala
  • Fundamentals of Item Response Theory — Hambleton, Swaminathan & Rogers
  • Applying the Rasch Model — Bond & Fox (Rasch-focused)
Pair this deck with ImpactMojo's Data Literacy, Survey Design and Monitoring & Evaluation 101 courses.
Read forStart with
A readable introductionEmbretson and Reise
Worked technical detailde Ayala
The classic referenceHambleton, Swaminathan and Rogers
A Rasch-specific viewBond and Fox
Applied development examplesFIES technical reports; NAS documentation
Read one general text and one applied report together. The texts explain the machinery; the reports show what a real calibration and DIF analysis look like when written up.
The FIES documentation is unusually accessible and is public, which makes it the best available worked example for anyone building an experience-based scale.
ImpactMojoItem Response Theory 101www.impactmojo.in
If you remember five things
  • The trait is hidden — IRT models the probability of each response from θ
  • Higher b = harder; higher a = steeper; c is the guessing floor
  • Information, not one reliability number — precision varies along θ, so target your test
  • DIF = same θ, different odds across groups — check for it
  • The guarantees hold only if the assumptions do — validate, don't assume
TakeawayCheck you can apply it
Model the probability, not the countExplain why a sum score misleads
Higher b harder, higher a steeperRead an ICC without the caption
Information varies along thetaSay where your test is precise
DIF is not a group gapState the difference in one sentence
Assumptions earn the benefitsName what you validated
If you take one thing into practice, take the last. IRT's advantages are conditional, and a practitioner who reports the conditions is more useful than one who reports the theta.
And if the answer to "did you validate" is no, a well-built sum score with a stated caveat is the more honest instrument.
ImpactMojoItem Response Theory 101www.impactmojo.in
Item Response Theory 101 · Complete
Now measure the
unmeasurable, well.
Continue your learning
CC BY-NC-ND 4.0·Free Forever·ImpactMojo 101 Series