| Latent trait | Observed | What can go wrong |
|---|---|---|
| Reading ability | Right / wrong on tasks | Tasks measure vocabulary instead |
| Empowerment | Agree / disagree items | The construct is contested |
| Food insecurity | Yes / no experiences | Seasonality changes the answers |
| Health knowledge | Correct answers | Recall of a campaign slogan |
| Stage | The assumption it rests on |
|---|---|
| Trait exists | The construct is coherent and one-dimensional |
| Items tap it | Content validity — the items are about the trait |
| Responses recorded | Administration is consistent and honest |
| Model estimates | The chosen model fits the data |
| Setting | Response type | Model usually used |
|---|---|---|
| ASER-style reading | Ordered tasks, pass/fail | Rasch or 2PL |
| Empowerment items | Agree / disagree | 2PL or graded response |
| Food insecurity (FIES) | Yes / no experience | Rasch / 1PL |
| Multiple-choice test | Correct / incorrect | 3PL where guessing matters |
| Sum score assumes | When it fails |
|---|---|
| Every item is worth the same | A hard item and an easy one both count 1 |
| Every gap is the same size | The jump from 4 to 5 is not the jump from 9 to 10 |
| The pattern does not matter | Two people with 6/12 are treated as identical |
| The scale is interval | It is ordinal at best |
| CTT | IRT | |
|---|---|---|
| Unit of analysis | The test as a whole | Each item |
| Person score | Sum or percentage | Estimated theta |
| Item statistics | Depend on the sample | Sample-free, if the model fits |
| Precision | One number for everyone | Varies along the scale |
| Data needed | Modest | Substantial |
| Benefit | What it enables concretely |
|---|---|
| Comparability | Track learning across years with different items |
| Fairness | Detect items biased against a group |
| Efficiency | Drop items that carry no information |
| Targeted precision | Measure accurately at the decision point |
| Honest error | Report uncertainty that varies by person |
| CTT term | Means |
|---|---|
| Observed score | What the person actually got |
| True score | Their expected score over infinite retests |
| Error | The difference, assumed random |
| Reliability | Share of variance that is true score |
| Reliability estimate | Answers |
|---|---|
| Test-retest | Would the same person score the same again? |
| Parallel forms | Do two versions agree? |
| Split-half | Do the two halves agree? |
| Cronbach's alpha | Do the items hang together internally? |
| Alpha misreading | What is actually true |
|---|---|
| "Alpha 0.9 means it is unidimensional" | Alpha rises with item count, not dimensionality |
| "0.7 is the threshold" | A convention, not a property |
| "Higher is always better" | Very high alpha often means redundant items |
| "Alpha validates the scale" | It says nothing about what is measured |
| Same item, given to | CTT p-value | True item property |
|---|---|---|
| A high-ability group | High — "easy" | Unchanged |
| A mixed group | Middling | Unchanged |
| A struggling group | Low — "hard" | Unchanged |
| Group | % correct on Item Q | CTT verdict |
|---|---|---|
| High-ability school | 88% | An easy item |
| Mixed government school | 61% | A moderate item |
| Struggling school | 34% | A hard item |
| Consequence of test-dependence | Where it bites |
|---|---|
| Scores are not comparable across forms | Two schools sitting different papers |
| Scores are not comparable across years | Tracking change with new items |
| Percentage right is not an ability | Reporting "60% proficient" |
| Anchoring is ad hoc | No principled way to link forms |
| Person at | Test measures them | Because |
|---|---|---|
| Middle of the range | Precisely | Many items near their level |
| Top of the range | Poorly | Everything is easy for them |
| Bottom of the range | Poorly | Everything is too hard |
| You want to… | CTT problem | IRT answer (coming) |
|---|---|---|
| Compare items fairly across samples | p-values shift with the sample | Item params are sample-free |
| Compare people across forms/years | Scores depend on the form | θ is on a common scale |
| Know precision at each level | One SEM for all | Information varies along θ |
| Build an adaptive or short form | Hard — scores not comparable | Items & people share a metric |
| Job | CTT gives you | IRT gives you |
|---|---|---|
| Compare items across samples | p-values that shift | Parameters that do not |
| Compare people across forms | Form-dependent scores | One theta scale |
| Report precision | One SEM | Error as a function of theta |
| Shorten the instrument | Drop by item-total r | Drop by information |
| On the shared scale | You can now ask |
|---|---|
| Person at theta, item at b | Is she above or below this item? |
| Two items | Which is harder, regardless of who sat them? |
| Two people | How far apart, in the same units? |
| A cut-score | Which items measure best right there? |
| Theta | Roughly means | Caution |
|---|---|---|
| -2 | Low trait; bottom few per cent | Measured imprecisely |
| 0 | Reference-group average | Only for that reference group |
| +2 | High trait; top few per cent | Also measured imprecisely |
| IRT says | IRT does not say |
|---|---|
| P(correct) = 0.7 at this theta | This person will get it right |
| This item is harder than that one | This item is unfair |
| Theta is estimated at 1.2 | Theta is exactly 1.2 |
| Precision is highest here | The construct is valid |
| Curve region | Probability | What it tells you |
|---|---|---|
| Far below b | Near 0 (or c) | The item is out of reach |
| Just below b | Rising steeply | The item is discriminating here |
| At b | 0.5 (2PL/1PL) | The item's location |
| Far above b | Near 1 | The item is trivial; no information |
| Property | Holds when | Fails when |
|---|---|---|
| Sample-free item parameters | The model fits | Multidimensionality, DIF |
| Test-free person scores | Items are on one calibrated scale | Forms never linked |
| Common metric | Anchors are stable | Anchor items drift |
| Ability item | Endorsement item |
|---|---|
| Correct / incorrect | Endorsed / not endorsed |
| b is difficulty | b is severity |
| Guessing can inflate the floor | No guessing; social desirability instead |
| Higher theta = more able | Higher theta = more of the condition |
| Section ahead | The question it answers |
|---|---|
| ICC | How do I read an item's behaviour? |
| Difficulty (b) | Where does this item work? |
| Discrimination (a) | How sharply does it sort people? |
| Guessing (c) | Is there a floor under the curve? |
| Information | Where is my test precise? |
| Validation and DIF | Am I entitled to any of this? |
| Element of the ICC | Parameter | Reads as |
|---|---|---|
| Where it crosses 0.5 | b | Difficulty / severity |
| Slope at that point | a | Discrimination |
| Lower asymptote | c | Guessing floor |
| Upper asymptote | (d, rarely used) | Careless-error ceiling |
| Read the ICC to answer | By |
|---|---|
| Is this item suitable for my group? | Check whether b sits inside their theta range |
| Will it separate people? | Look at the slope where they sit |
| Is it wasted? | Flat curve across your whole range |
| Is it mis-keyed? | Curve slopes downward |
| Landmark | Where information is |
|---|---|
| Lower tail | Almost none |
| Approaching b | Rising |
| At the inflection | Maximum |
| Upper tail | Almost none |
| Non-monotonic curve means | Check |
|---|---|
| Mis-keyed item | The answer key |
| Ambiguous wording | Whether strong respondents over-read it |
| Trick or double-barrelled item | The item text |
| Reverse-coded and not recoded | The coding script |
| Estimation step | What it does |
|---|---|
| Calibration | Estimate item parameters from a large sample |
| Scoring | Estimate each person's theta given those parameters |
| Standard error | Quantify uncertainty at that theta |
| Reporting | Translate theta into something actionable |
| Two children, 6 of 12 | IRT reads |
|---|---|
| Passed the six easiest | Theta near the middle of the easy items |
| Passed six hard, missed easy | An unusual pattern — check for problems |
| Mixed, consistent with difficulty order | A well-behaved estimate |
| b value | Item is | Useful for |
|---|---|---|
| -2.0 | Very easy | Separating the weakest |
| -1.0 | Easy | Lower half |
| 0.0 | Moderate | The middle |
| +1.0 | Hard | Upper half |
| +2.0 | Very hard | Separating the strongest |
| Person theta vs item b | Expected outcome |
|---|---|
| Well above b | Very likely correct; little information |
| Slightly above b | Likely correct; high information |
| At b | 50/50; maximum information |
| Well below b | Very likely wrong; little information |
| Item (early-grade reading) | b (illustrative) | Interpretation |
|---|---|---|
| Recognise a letter | −2.0 | Very easy — most children pass |
| Read a familiar word | −0.8 | Easy |
| Read a simple sentence | 0.2 | Moderate |
| Read a short paragraph | 1.0 | Hard |
| Answer a comprehension question | 1.8 | Very hard |
| If your sample sits at | These items are |
|---|---|
| theta about -1.5 | Letter and word items informative; paragraph wasted |
| theta about 0 | Sentence and paragraph items informative |
| theta about +1 | Comprehension items informative; letters wasted |
| FIES-style item | Severity |
|---|---|
| Worried about food | Low |
| Unable to eat healthy food | Low to moderate |
| Ate less than you thought you should | Moderate |
| Ran out of food | High |
| Went a whole day without eating | Highest |
| Looks hard but may be easy | Looks easy but may be hard |
|---|---|
| A technical term everyone was taught | A word that is unfamiliar in the local dialect |
| A long question with clear structure | A short question with ambiguous wording |
| An advanced topic recently covered | An old topic taught years ago |
| Coming up | In one line |
|---|---|
| a defined | The slope of the curve at its midpoint |
| Steep vs flat | How sharply the item separates people |
| Low and negative a | Dead weight, and broken items |
| a with b | Where it works, and how well it works there |
| a value | Curve | Item is |
|---|---|---|
| Below 0.5 | Nearly flat | Contributing little |
| 0.5 to 1.0 | Gentle | Weak but usable |
| 1.0 to 2.0 | Steep | Good |
| Above 2.0 | Very steep | Excellent, but narrow |
| Negative | Downward | Broken — investigate |
| High a gives you | At the cost of |
|---|---|
| Sharp separation near b | A narrow useful band |
| High information at one point | Almost none elsewhere |
| Efficient short tests | Sensitivity to model misfit |
| Low a usually means | Check |
|---|---|
| The item taps a different trait | Dimensionality |
| The wording is ambiguous | The item text and translation |
| Everyone guesses | Whether c should be modelled |
| The item is off-target | Whether b is far from the sample |
| Negative a cause | Fix |
|---|---|
| Mis-keyed answer | Correct the key, refit |
| Reverse-worded, not recoded | Recode, refit |
| Item measures the opposite construct | Reconsider or drop |
| Data entry column shift | Check the merge |
| Low a (flat) | High a (steep) | |
|---|---|---|
| Low b (easy) | Easy, sorts weakly | Easy, sorts low-trait people sharply |
| High b (hard) | Hard, sorts weakly | Hard, sorts high-trait people sharply |
| Low a | High a | |
|---|---|---|
| Low b | Easy, sorts weakly | Easy, sorts sharply at the low end |
| High b | Hard, sorts weakly | Hard, sorts sharply at the high end |
| Empowerment item | Likely behaviour |
|---|---|
| "Can go to the market alone" | Discriminates in restrictive settings |
| "Has a say in her own healthcare" | Widely endorsed; low severity |
| "Owns land in her own name" | High severity, low endorsement |
| "Participates in household decisions" | Often low a — everyone says yes |
| Test built only of | Result |
|---|---|
| Very steep items at one b | Superb at one point, useless elsewhere |
| Flat items across a wide range | Broad but imprecise everywhere |
| Steep items spread across b | The usual target |
| Coming up | In one line |
|---|---|
| The guessing floor | Nobody scores zero on multiple choice |
| Why it matters | Ignoring it biases difficulty downward |
| 1PL, 2PL, 3PL | Each adds one parameter and one data requirement |
| Choosing | Start simple; add only what the data supports |
| Format | Chance floor | Model c? |
|---|---|---|
| 4-option multiple choice | About 0.25 | Yes |
| True / false | About 0.5 | Yes, and it dominates |
| Open response | About 0 | No |
| Yes / no experience item | Not applicable | No |
| Ignoring guessing causes | Visible as |
|---|---|
| b biased downward | Items look easier than they are |
| a biased | Discrimination understated at the low end |
| Low-theta scores inflated | Nobody appears to be at the bottom |
| Poor fit in the lower tail | Empirical curve above the model |
| Model | Free parameters | Adds… |
|---|---|---|
| 1PL / Rasch | b only (a fixed equal) | Difficulty differs; all items equally discriminating |
| 2PL | a and b | Items also differ in discrimination |
| 3PL | a, b and c | Plus a guessing floor |
| Model | Estimates | Needs |
|---|---|---|
| 1PL / Rasch | b only | Smallest sample; strongest assumption |
| 2PL | a and b | Moderate sample |
| 3PL | a, b and c | Large sample; c often fixed |
| Rasch is preferred because | Rasch is criticised because |
|---|---|
| Sufficient statistics: raw score maps to theta | It assumes equal discrimination |
| Specific objectivity properties | Items that misfit are dropped, not modelled |
| Stable with modest samples | The assumption is often empirically false |
| Simple to explain and defend | Fit is achieved by discarding data |
| Use 2PL when | Signs it is needed |
|---|---|
| Items plainly differ in quality | Wide spread of item-total correlations |
| You have the sample size | Several hundred respondents |
| Rasch shows systematic misfit | Empirical curves steeper or flatter than modelled |
| 3PL practicalities | What people do |
|---|---|
| c is unstable to estimate | Fix it at 1/options |
| Needs large samples | Reserve for high-stakes testing |
| Trades off with b | Constrain c with a prior |
| Not needed without guessing | Do not use it on yes/no scales |
| Situation | Model |
|---|---|
| Yes/no experience scale, 300 respondents | Rasch |
| Ability test, varied item quality, 800 | 2PL |
| Multiple-choice, high stakes, 2,000+ | 3PL |
| Ordered categories (never/sometimes/often) | Graded response or partial credit |
| Genuinely several traits | Multidimensional, or separate scales |
| Coming up | In one line |
|---|---|
| Item information | An item informs most near its own difficulty |
| Test information | Item information simply adds up |
| Standard error | SE at any theta is 1 over the square root of information |
| Targeting | Put the information where the decision is |
| Information rises with | Because |
|---|---|
| Discrimination a (squared) | A steeper curve separates more sharply |
| Proximity of theta to b | The item is at its most uncertain, so most informative |
| Number of items nearby | Information is additive |
| Absence of guessing | A floor flattens the lower half |
| CTT | IRT | |
|---|---|---|
| Error report | One SEM for everyone | SE as a function of theta |
| At the extremes | Understated | Correctly large |
| In the middle | Overstated | Correctly small |
| Consequence | Misleading confidence intervals | Honest ones |
| Decision type | Where to put information |
|---|---|
| Pass / fail at one cut | Concentrate items near the cut |
| Three proficiency bands | Peaks at each of the two cuts |
| Describe the whole distribution | Spread evenly across the range |
| Select the top 10% | Concentrate high on the scale |
| Adaptive testing needs | Which is why |
|---|---|
| A calibrated item bank | Items must already be on one scale |
| Many items per difficulty band | Exposure must be controlled |
| Delivery technology | Selection happens between items |
| A stopping rule | Either fixed length or target precision |
| Coming up | In one line |
|---|---|
| Model fit | Do the responses match the fitted curves? |
| Unidimensionality | Is it one trait, or several? |
| Local independence | Do items lean on each other? |
| DIF | Does an item behave differently by group? |
| Promise | Depends on |
|---|---|
| Sample-free item parameters | Model fit and no DIF |
| Test-free person scores | Common calibration or linking |
| Meaningful information function | Correct model choice |
| Comparability across groups | DIF screening |
| Fit check | What it catches |
|---|---|
| Item-fit statistics | Items whose responses depart from the model |
| Empirical vs modelled ICC plot | Where and how they depart |
| Person-fit statistics | Aberrant response patterns |
| Residual analysis | Systematic misfit across the bank |
| Dimensionality check | Warning sign |
|---|---|
| Factor analysis of the items | More than one substantial factor |
| Principal components of residuals | A structured first residual component |
| Content review | Items about visibly different things |
| Fit across subscales | Better fit when split |
| Local dependence source | Example |
|---|---|
| Shared stimulus | Five questions on one passage |
| Chained items | Item 4 gives away item 5 |
| Repeated wording | Near-duplicate items |
| Order effects | An item that primes the next |
| DIF is | DIF is not |
|---|---|
| Different P at the same theta | A group difference in theta |
| A property of an item | A property of the group |
| Evidence of a nuisance factor | Proof of intent |
| Detectable statistically | Interpretable without content review |
| Observation | DIF? |
|---|---|
| Girls score higher overall | No — a trait difference |
| Girls at the same theta do better on item 7 | Yes |
| One region scores lower | No, by itself |
| One region at the same theta fails item 12 | Yes |
| DIF cause | Action |
|---|---|
| Unfamiliar context or example | Rewrite with a neutral context |
| Poor or unequal translation | Retranslate and back-translate |
| Curriculum differs by region | May be legitimate — document it |
| Gendered scenario | Rewrite |
| No identifiable cause | Consider dropping; report the decision |
| Assessment | What IRT enables |
|---|---|
| NAS | Common scale across grades and states |
| ASER | A simple ordered tool with a scaling logic behind it |
| Cross-national studies | Comparison across languages, via DIF screening |
| Programme endlines | Comparison with a different form |
| Linking design | How it works |
|---|---|
| Common items (anchors) | Shared items across two forms |
| Common persons | The same people sit both forms |
| Single-group | One group, both forms, counterbalanced |
| Equating risk | Symptom |
|---|---|
| Anchor drift | An anchor's b changes between years |
| Too few anchors | Unstable linking constants |
| Anchors clustered in difficulty | Poor linking at the extremes |
| Anchor exposure | Anchors become known and easy |
| FIES design choice | Reason |
|---|---|
| Eight yes/no items | Short enough for a household survey module |
| Rasch / 1PL | Stable with modest samples; simple to defend |
| Experience items, not opinion | Less prone to social desirability |
| Global reference scale | Enables cross-country comparability |
| Empowerment scaling question | Why it is hard |
|---|---|
| Is it one trait? | Mobility, decisions and assets may not co-vary |
| Does severity order hold across settings? | Norms differ sharply |
| Is DIF present across groups? | Almost always, by caste, region, religion |
| Does the scale mean the same over time? | Norms shift during a programme |
| Asset index method | Note |
|---|---|
| Principal components (DHS standard) | Not a latent-trait model; weights from PCA |
| IRT / latent trait | Models the probability of ownership |
| Both | Sensitive to which assets are included |
| Scale type | What IRT adds |
|---|---|
| Health knowledge | Grade items and target the campaign's level |
| Stigma / attitudes | Order statements by how much it takes to endorse |
| Quality of life | Check whether the domains are one trait |
| Depression / distress screeners | Precision at the clinical cut-point |
| Payoff | Realistic scale |
|---|---|
| Shorter instruments | Often 30-50% fewer items at equal precision |
| Comparable numbers | Across forms and years, if anchored |
| Fairer measures | Only for the groups you screened |
| Targeted precision | At the decision point you specify |
| Assumption | Consequence if violated |
|---|---|
| Unidimensionality | Theta blends two traits; scores uninterpretable |
| Local independence | Precision overstated |
| Correct model | Biased b and distorted low-end scores |
| Monotonicity | The scale is not ordered as claimed |
| No DIF | Group comparisons are invalid |
| Model | Rough sample-size guidance | Note |
|---|---|---|
| 1PL / Rasch | ~200+ respondents | Most forgiving |
| 2PL | ~500+ respondents | Estimating a needs more data |
| 3PL | ~1,000+ respondents | c is hard to estimate; often fixed |
| Sample-size driver | Effect |
|---|---|
| Number of parameters per item | More parameters, more data |
| Spread of theta in the sample | Concentrated samples estimate poorly |
| Number of items | More items stabilise person estimates |
| Sparse response categories | Rare categories need large samples |
| Limit | What it means for you |
|---|---|
| Bad items give bad scales | Invest in item writing first |
| A neat theta can hide a contested concept | Say what "empowerment" means here |
| Forced unidimensionality loses meaning | Report two scales if there are two |
| Precision is not validity | A precise measure of the wrong thing |
| Tool | Good for | Note |
|---|---|---|
| R: mirt | Uni- & multidimensional IRT, all common models | Free, powerful, well documented |
| R: ltm | 1PL/2PL/3PL for dichotomous & graded items | Free, gentle entry point |
| R: TAM / eRm | Large-scale & Rasch modelling | Free; TAM mirrors big assessments |
| Stata: irt suite | IRT within a familiar stats package | Built-in irt commands |
| jMetrik / IRTPRO | Point-and-click psychometrics | Lower coding barrier |
| Tool | Reach for it when |
|---|---|
| R: mirt | You need most models, including multidimensional |
| R: ltm | A gentle first pass on dichotomous items |
| R: TAM / eRm | Rasch and large-scale assessment workflows |
| R: difR / lordif | DIF screening specifically |
| Stata / commercial | Institutional standard, high-stakes work |
| Workflow step | Do not skip |
|---|---|
| Dimensionality | It invalidates everything downstream |
| Simplest defensible model | Complexity you cannot estimate is worse |
| Item fit and parameters | Plot, do not only test |
| DIF screening | Decide the groups in advance |
| Revise, then score | Do not score with items you would drop |
| Instead of | Report |
|---|---|
| "Theta = 1.2" | "Can read a short paragraph with comprehension" |
| "Mean theta rose 0.3" | "12% more children reached the paragraph level" |
| "SE = 0.35" | "Her level is between sentence and paragraph" |
| "Above the cut-score" | "At grade level, with the uncertainty stated" |
| Read for | Start with |
|---|---|
| A readable introduction | Embretson and Reise |
| Worked technical detail | de Ayala |
| The classic reference | Hambleton, Swaminathan and Rogers |
| A Rasch-specific view | Bond and Fox |
| Applied development examples | FIES technical reports; NAS documentation |
| Takeaway | Check you can apply it |
|---|---|
| Model the probability, not the count | Explain why a sum score misleads |
| Higher b harder, higher a steeper | Read an ICC without the caption |
| Information varies along theta | Say where your test is precise |
| DIF is not a group gap | State the difference in one sentence |
| Assumptions earn the benefits | Name what you validated |