fullscreen
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
ImpactMojo 101 Series · Free Forever
Structural
Equation
Modelling 101
Latent Variables, Confirmatory Factor Analysis, Mediation, Measurement Invariance and PLS-SEM — Measured, Fitted and Reported Honestly
Research MethodsSouth Asia Focus100 SlidesFree Access
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
What We Cover
01
What SEM Is, and Is Not
Slides 3–10
02
Latent Variables and CFA
Slides 11–19
03
Reliability and Validity
Slides 20–27
04
Specification and Identification
Slides 28–35
05
Estimation, Data and Sample Size
Slides 36–43
06
Model Fit and Modification
Slides 44–51
07
Structural Models: Mediation and Moderation
Slides 52–61
08
Measurement Invariance
Slides 62–69
09
PLS-SEM
Slides 70–80
10
Causality and Common Method Bias
Slides 81–89
11
Software, Reporting and Practice
Slides 90–99
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
01
Section One
What SEM Is, and Is Not
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
SEM is a family of models for variables you cannot observe directly
Structural equation modelling
A framework that combines a measurement model (how observed indicators relate to unobserved latent variables) with a structural model (how the latent variables relate to each other), estimated together so that measurement error is separated from the relationships of interest.
Women's empowerment, trust in institutions, food insecurity, job satisfaction, service quality: none is a column in a dataset. Each is inferred from several imperfect questions. Ordinary regression on a summed score treats the sum as the thing and its error as zero, which attenuates every coefficient and hides how well the questions measured anything. SEM makes the measurement explicit and testable.
  • Path analysis (Wright, 1921) is SEM with observed variables only.
  • Confirmatory factor analysis (Jöreskog, 1969) is SEM with a measurement model only.
  • The full model (LISREL, 1970s) puts them together: latent variables measured by indicators, related by regressions.
  • Growth curve models, multilevel SEM, latent class models and PLS-SEM are extensions or cousins; sections 07 to 09 touch them.
  • Bollen's Structural Equations with Latent Variables (Wiley, 1989) and Kline's Principles and Practice (5th ed., Guilford, 2023) are the two texts this course leans on.
SEM is a way of estimating a theory you already have. It is not a way of finding one in the data.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
The path diagram: reading the picture
SymbolMeaning
RectangleObserved (manifest) variable: a survey item, a measured quantity
OvalLatent variable (factor, construct): unobserved, inferred from indicators
Single-headed arrowA directed effect: regression coefficient (structural) or factor loading (measurement)
Double-headed arrowA covariance or correlation, with no direction claimed
Small arrow into a variable, or a circleError or disturbance: the part not explained by the model
Triangle (rare)The constant, when means and intercepts are modelled
A diagram is a set of equations. Every arrow is a parameter to estimate, every missing arrow is a restriction (a coefficient fixed at zero), and it is the missing arrows that make the model testable. A model with an arrow between every pair of variables fits any data perfectly and says nothing. The discipline of SEM is deciding which arrows to leave out, before seeing the data, for reasons you can state.
An example, in words
Membership of a self-help group (observed) affects savings behaviour (latent, three items) which affects women's decision-making (latent, five items). Membership may also affect decision-making directly. Two latent variables, eight items, three structural paths, and a mediation hypothesis. Section 07 estimates it.
Draw the diagram before collecting data. It tells you how many items each construct needs and which questions the survey must contain.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
What SEM buys, and what it costs
What it buys
Coefficients corrected for measurement error, so a true effect is not attenuated by noisy questions. A test of whether the questions measure what you say (the measurement model). Several equations estimated at once, so mediation and indirect effects come with standard errors. A test of the whole theory against the data, not one coefficient at a time. Comparison of the same model across groups, languages or waves, with a test of whether the measurement held.
What it costs
Sample sizes in the hundreds. Assumptions about distributions, linearity and the correctness of the model's omissions. A large set of decisions (indicators, estimator, fit thresholds, modifications) each of which can be made to flatter the model. And a literature in South Asian business and social-science journals where every model fits, every hypothesis is supported, and the reader learns nothing, because the decisions were made to produce that.
The method is sound and much of its use is not. This course is organised around the decisions, in the order they arise, with the defensible choice at each and the common abuse beside it. A reader who follows it will produce a model that can fail, which is the only kind worth reporting.
If a summed scale and OLS would answer the question, use them. SEM earns its place when measurement is in doubt, when mediation is the question, or when groups must be compared on a latent construct.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Covariance-based SEM and PLS-SEM: two methods with one name
Covariance-based SEM (CB-SEM)Partial least squares SEM (PLS-SEM)
What it fitsThe full covariance matrix of the indicators, by maximum likelihood or a robust variantA sequence of regressions on weighted composites of the indicators
Latent variablesCommon factors: the shared variance of the indicators, error separatedComposites: weighted sums of the indicators, error included
GoalTest a theory: does the implied covariance matrix match the observed one?Predict and explain: maximise explained variance of the endogenous composites
Global fit testYes: chi-square and the fit indicesNo exact test; SRMR and prediction-based assessment
Sample sizeLarger; hundredsSmaller samples tolerated, with caveats
Softwarelavaan (R), Mplus, AMOS, Stata sem, JASP, semopy (Python)SmartPLS, SEMinR (R), ADANCO, cSEM (R)
Where commonPsychology, sociology, public health, educationMarketing, information systems, management; South Asian business schools
Sections02 to 08, 10, 1109
They are different estimators of different things, and the choice between them is a choice about the question, not about which software the department owns. Section 09 gives the case for each; Hair and colleagues' 'silver bullet' paper (Journal of Marketing Theory and Practice 2011, 19:139) and Rönkkö and Evermann's reply (Organizational Research Methods 2013, 16:425) are the two sides.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
SEM in development and social research
FieldTypical constructsTypical model
Women's empowermentDecision-making, mobility, control over assets, attitudes to violenceCFA of a multidimensional scale; invariance across states or languages; effects of a programme on each dimension
Health behaviourKnowledge, attitudes, perceived risk, self-efficacy, intentionTheory of planned behaviour as a mediation chain; intention mediating knowledge to practice
EducationTeacher motivation, school climate, student engagement, achievementMultilevel SEM: students within schools; growth models across grades
Service delivery and governanceTrust in institutions, perceived corruption, satisfaction, willingness to payStructural model of trust to compliance; invariance across districts
Livelihoods and financeFinancial literacy, risk attitude, social capital, savings behaviourMediation of programme exposure through social capital
Management and marketing (business schools)Technology acceptance, service quality, job satisfaction, intention to adoptPLS-SEM of UTAUT or SERVQUAL-type models on convenience samples
Psychology and psychometricsWell-being, depression, resilience scalesCFA and bifactor models; validation of translated instruments
The unifying feature is a construct that a single question cannot capture and that the argument needs to treat as one thing. In development work the commonest use, and the most valuable, is validating a translated scale before an evaluation relies on it.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Vocabulary and notation used throughout
TermMeaning
Indicator, item, manifest variableAn observed variable that measures a construct
Latent variable, factor, constructThe unobserved variable the indicators measure
Loading (λ)The regression of an indicator on its factor; standardised, its correlation with the factor
ExogenousA variable with no arrows pointing into it
EndogenousA variable with at least one arrow into it
Disturbance (ζ), error (ε, δ)Unexplained variance in an endogenous latent or an indicator
Reflective measurementThe construct causes the indicators (arrows from oval to rectangles)
Formative measurementThe indicators define the construct (arrows from rectangles to oval)
TermMeaning
Free parameterEstimated from the data
Fixed parameterSet by the analyst (usually 0 or 1)
Degrees of freedomKnown covariances minus free parameters
IdentifiedEvery parameter has a unique solution
χ² (chi-square)The test of exact fit: implied versus observed covariances
Direct, indirect, total effectA path; a product of paths through a mediator; their sum
InvarianceThe same measurement model holds across groups
Modification indexThe expected drop in χ² from freeing one fixed parameter
Greek letters follow the LISREL convention where they appear; the words are what matter.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
How the course is arranged, and what it assumes
01
MEASURE: latent variables, CFA, reliability, validity (02–03)
02
SPECIFY: the model, its identification, its data and estimator (04–05)
03
FIT: judge and, carefully, modify (06)
04
RELATE: structural paths, mediation, moderation (07)
05
COMPARE: invariance across groups and languages (08)
06
ALTERNATIVE: PLS-SEM, and when it is the right tool (09)
07
INTERPRET AND REPORT: causality, method bias, software, the write-up (10–11)
Assumed: regression at the level of Econometrics 101, and the survey and scale material in Survey Design 101. Worked numbers are illustrative unless a source is named, with magnitudes chosen to be realistic so the reader learns to read the output. Code is given in lavaan syntax because it is free, readable and the reference implementation; every other package has an equivalent.
The free companion is Rosseel's lavaan tutorial at lavaan.ugent.be. For the theory, Kline (2023). For PLS-SEM, Hair and colleagues, A Primer on PLS-SEM (3rd ed., Sage, 2022), read alongside its critics.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
02
Section Two
Latent Variables and CFA
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
The common factor model: what a latent variable is, mathematically
The measurement equation
xi = τi + λiη + εi

Each indicator xi is an intercept, plus a loading times the factor η, plus an error unique to that indicator. The factor is what the indicators share; the errors are what they do not. Under the model, the covariance between two indicators is λiλj Var(η), and that is what CFA tests: does the observed covariance matrix look like one generated by a few factors?
Standardised, λ² is the share of the indicator's variance explained by the factor; 1 − λ² is error. A loading of 0.7 means about half the item is signal.
  • The factor has no natural scale. It is given one by fixing one loading to 1 (the marker indicator) or fixing the factor variance to 1. The choice changes the unstandardised numbers and nothing else.
  • Errors are assumed uncorrelated with the factor and, by default, with each other. A correlated error between two items says they share something beyond the factor (identical wording, adjacency in the questionnaire) and must be argued for.
  • Exploratory factor analysis lets every item load on every factor and asks how many factors there are. CFA fixes most loadings to zero in advance and asks whether the specified structure holds. Use EFA on a new scale in a development sample; CFA to confirm it on a fresh one.
'Latent' does not mean 'real'. A factor is a statistical summary of shared variance; whether it corresponds to empowerment is a validity argument, section 03.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Reflective or formative: which way do the arrows point?
Reflective
The construct causes the indicators. Depression causes low mood, poor sleep and loss of appetite; the items are interchangeable symptoms, they should correlate, dropping one changes little, and reliability makes sense. Most attitude and psychological scales. The common factor model applies.
Formative
The indicators define the construct. Socioeconomic status is income, education and occupation; they need not correlate, dropping one changes the meaning, and 'reliability' is the wrong question. A composite, not a factor. Estimated as a weighted sum, with weights from the outcomes it predicts (MIMIC models) or fixed by the analyst (an index).
Getting this wrong is the commonest specification error in applied SEM and it is invisible in the fit statistics. A formative construct forced into a reflective model shows low loadings and poor 'reliability' that are not defects of measurement but of the model. Jarvis, MacKenzie and Podsakoff (Journal of Consumer Research 2003, 30:199) give the decision rules; Bollen and Lennox (1991) the theory.
  • Ask: if the construct rose, would every indicator rise? (Reflective.) If one indicator rose, would the construct rise? (Formative.)
  • A wealth index from asset ownership is formative. A scale of attitudes toward girls' education is reflective. Household food insecurity (HFIAS) is argued both ways and usually treated as reflective.
  • CB-SEM handles formative constructs awkwardly and needs identification tricks; PLS-SEM handles them natively, which is one legitimate reason to use it.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Specifying a CFA: constructs, items, and the syntax
lavaan syntax
model <- '
  decide =~ d1 + d2 + d3 + d4 + d5
  mobility =~ m1 + m2 + m3
  assets =~ a1 + a2 + a3 + a4
'
fit <- cfa(model, data = df, estimator = "MLR")
summary(fit, fit.measures = TRUE, standardized = TRUE)
=~ reads 'is measured by'. Three factors, twelve items, the factors correlated by default (a double-headed arrow among the ovals), the first loading of each factor fixed to 1 by default, errors uncorrelated. Twelve lines of output later you have loadings, factor covariances, error variances and fit.
  • Three indicators per factor is the practical minimum; four or more gives the factor its own degrees of freedom and lets a bad item be dropped.
  • Each item loads on one factor (simple structure) unless there is a reason. Cross-loadings are argued, not discovered.
  • Items should be on comparable scales; a 5-point item and a 0–100 item on one factor produce loadings that are hard to read. Standardised output helps.
  • Name the factors by what the items ask, not by what you hope they measure. 'decide' for five items about who decides is honest; 'empowerment' for the same five is a claim.
Equivalents: Stata sem (Decide -> d1 d2 d3 d4 d5) ...; AMOS by drawing; JASP by dragging items into factors; Python semopy uses lavaan's syntax.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Reading CFA output: a worked example
Factor and itemUnstd. λSEStd. λComment
decide =~ d1 (who decides on food purchases)1.0000.710.50Marker; fixed
decide =~ d2 (large purchases)1.180.090.790.62
decide =~ d3 (visiting family)1.050.090.740.55
decide =~ d4 (own health care)0.970.100.680.46
decide =~ d5 (children's schooling)0.620.110.410.17Weak; shares little with the others
mobility =~ m11.0000.820.67
mobility =~ m20.940.070.770.59
mobility =~ m30.880.080.700.49
decide ~~ mobility (correlation)0.46Related, distinct
Illustrative; n = 640 women. Standardised loadings above about 0.5 (R² above 0.25) are the conventional floor; d5 fails it. The factor correlation of 0.46 says decision-making and mobility are not the same thing, which matters for section 03's discriminant validity.
What to do about d5
Not delete it on sight. Ask why: children's schooling decisions may be made jointly in most households, so the item does not discriminate; or the question was badly translated; or it belongs to a different construct. Check the item's distribution and the cognitive-interview notes. Drop it only with a reason you can write, and re-run the model on a fresh sample if the scale will be used again.
The standard errors are as important as the loadings. A loading of 0.62 with SE 0.11 is not different from 0.85 at conventional levels.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Higher-order and bifactor models: when constructs have structure
Empowerment measured as decision-making, mobility and control over assets raises the question whether there is one empowerment or three. A second-order model puts a general factor above the three, which then load on it; it is identified with three or more first-order factors and says the three are manifestations of one thing. A bifactor model lets every item load on a general factor and on its specific factor at once, and asks how much of the variance is general.
  • Compare the second-order model against the correlated-factors model: it is nested and can only fit worse; if it fits nearly as well, the general factor is supported.
  • Bifactor models fit almost anything and are over-used; Reise's omega-hierarchical and the explained common variance (ECV) say whether the general factor is worth having.
  • A general factor that explains 40% of the common variance and specific factors that explain little means 'empowerment' is one thing here; the reverse means it is three, and the programme's effects should be reported on each.
Why it matters for evaluation
A programme that raises mobility and not decision-making shows a small effect on a single empowerment score and a large one on one dimension. Kabeer's own definition (1999) treats resources, agency and achievements as related and distinct, which is a correlated-factors model rather than a second-order one. The measurement model is a theory of the construct, and it should be chosen to match the theory the evaluation is testing.
Report the correlated-factors model first. Add the higher-order structure only if the argument needs one score.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Likert items are ordinal, and the model should know
A five-point agreement item is not a continuous measurement. Treating it as one with maximum likelihood works tolerably with five or more categories and roughly symmetric distributions, and fails with skewed items or four or fewer categories: loadings are attenuated and fit is distorted. The ordinal treatment assumes a continuous latent response behind each item, estimates thresholds and polychoric correlations, and fits the model to those by diagonally weighted least squares (WLSMV in Mplus and lavaan).
  • Flora and Curran (Psychological Methods 2004, 9:466) show WLSMV recovers loadings well down to two categories; ML does not below five.
  • In lavaan: ordered = c("d1", ..., "d5") and the estimator switches automatically.
  • Binary items (yes/no decision questions, the common form in DHS-style modules) require the ordinal treatment; ML on binary items is wrong.
  • Fit indices under WLSMV are on a different footing from ML; compare within an estimator, not across.
Practical consequences
Most empowerment, attitude and knowledge items in South Asian surveys are binary or three-category. That means WLSMV, polychoric correlations, thresholds in the output, and a sample large enough for the weight matrix to be stable (several hundred). It also means the familiar Cronbach's alpha, computed on the raw items, understates reliability; the ordinal alpha or omega from the polychoric matrix is the right one.
Plot the item distributions before choosing an estimator. An item with 92% 'yes' carries little information whatever the estimator.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Exploratory factor analysis: the step before CFA on a new scale
DecisionRecommendedAvoid
ExtractionPrincipal axis factoring or ML factor analysisPrincipal components analysis (it is not factor analysis; it models total variance)
Number of factorsParallel analysis (Horn 1965); scree with judgement; theoryThe eigenvalue-greater-than-1 rule alone (over-extracts)
RotationOblique (oblimin, promax): factors are allowed to correlateVarimax by default (forces independence the constructs rarely have)
Item retentionLoading above 0.4 on one factor, cross-loadings below 0.3, with content reviewDeleting items purely on statistics
SampleA development sample, then CFA on a separate sampleEFA and CFA on the same respondents
Correlation matrixPolychoric for ordinal itemsPearson on binary items
EFA on a new or translated scale shows how the items actually cluster before you tell the CFA how they should. Split the sample at random, explore on one half, confirm on the other; or explore on the pilot and confirm on the main survey. Running CFA on the sample the EFA was tuned to is confirming what you just fitted.
Parallel analysis is in R's psych package (fa.parallel) and in JASP and jamovi. It compares eigenvalues against those from random data of the same size and is the one factor-count rule that survives scrutiny.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Items that make CFA work: what to fix at the questionnaire stage
  • Four to six items per construct, so that dropping one still leaves a testable factor.
  • Items that vary: an item everyone agrees with has no variance to share. Pilot for ceiling and floor effects.
  • One idea per item. 'I can decide about my own health care and my children's' is two items.
  • Consistent response scales within a construct, and reverse-worded items used sparingly: they often form their own method factor, especially in translation.
  • Translation with back-translation and cognitive interviewing in each language, and the same item order in every version.
  • Record which respondents answered which language version; section 08 needs it.
The reverse-item factor
Scales with half the items reversed ('I am satisfied' / 'I am not satisfied') routinely show a two-factor CFA solution that is wording, not content. In Hindi and Bangla, negatively worded items are more often misread, and the artefact is larger. Either avoid reversal or model a method factor for the reversed items and report it. Never present the wording factor as a substantive finding.
Survey Design 101 covers the questionnaire; this slide is the part of it that decides whether the CFA can succeed. The measurement model is written in the questionnaire, not in the software.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
03
Section Three
Reliability and Validity
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Reliability: alpha, omega and what each assumes
CoefficientAssumesComputed fromUse
Cronbach's α (1951)Equal loadings (tau-equivalence), uncorrelated errorsItem covariancesFamiliar; understates reliability when loadings differ; overstates it with many items or correlated errors
McDonald's ω (composite reliability)The CFA model holdsLoadings and error variancesThe default for a scale with a fitted CFA; report it instead of or beside α
Ordinal α / ωAs above, on the polychoric matrixPolychoric correlations, thresholdsFor Likert and binary items
ωh (hierarchical)Bifactor modelGeneral-factor loadingsHow much of the total score is the general factor
Test-retestStability of the constructTwo administrationsFor traits, not states; rarely feasible in field surveys
Inter-raterMultiple observersAgreement statistics (κ, ICC)Observational measures, enumerator-rated items
Reliability is the share of observed-score variance that is true-score variance: how repeatable the measurement is, not whether it measures the right thing. Values above 0.7 are conventional for research use and above 0.8 or 0.9 for individual decisions; the thresholds are conventions, not findings. Composite reliability = (Σλ)² / [(Σλ)² + Σθ] from the standardised CFA.
An α of 0.95 on a twelve-item scale is often a sign of redundancy (six items asked twice), not of excellence. Report ω and the number of items.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Convergent validity: do the items agree about something?
Convergent validity in the CFA sense is that the items of a construct share enough variance to be measuring one thing. The evidence is the loadings (all substantial, all significant) and the average variance extracted, AVE = mean of the squared standardised loadings, the share of item variance the factor explains on average. Fornell and Larcker (Journal of Marketing Research 1981, 18:39) proposed 0.5 as the floor: the factor explains more of its items than error does.
≥ 0.5
AVE floor (Fornell and Larcker 1981)
Convention
≥ 0.7
composite reliability floor
Convention
≥ 0.5
standardised loading floor, with 0.7 preferred
Hair et al.; Kline
  • In the worked CFA, 'decide' has AVE = (0.50 + 0.62 + 0.55 + 0.46 + 0.17) / 5 = 0.46 with d5, and 0.53 without it. The threshold is crossed by dropping one item, which is exactly the kind of decision that should be stated and justified.
  • AVE below 0.5 with composite reliability above 0.6 is sometimes accepted (Fornell and Larcker's own note); say so if you rely on it.
  • The thresholds are conventions from marketing research in 1981. They are useful defaults and poor gods; a scale that fails them by a little and has strong theory and invariance evidence is better than one that clears them by item deletion.
  • Convergent validity in the broader sense (the scale correlates with other measures of the same thing) is a separate and stronger test, when another measure exists.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Discriminant validity: are the constructs different from each other?
CriterionRuleNotes
Fornell-Larcker√AVE of each construct exceeds its correlation with every other constructTraditional; weak at detecting problems when loadings vary (Henseler et al. 2015)
HTMT (heterotrait-monotrait ratio)Ratio of between-construct to within-construct item correlations below 0.85 (strict) or 0.90 (lenient); bootstrap CI excludes 1Henseler, Ringle and Sarstedt, J Acad Marketing Sci 2015, 43:115; now the default in PLS-SEM and increasingly in CB-SEM
Factor correlation CIThe 95% CI of the correlation between two factors excludes 1Direct and simple; report the correlation and CI
Chi-square differenceA model with the two factors merged fits significantly worseA nested-model test; sensitive to sample size
Cross-loadingsEach item loads more on its own construct than on othersWeak; mainly a PLS-SEM habit
Two factors correlated at 0.92 are one factor with two names, and any 'effect' of one on the other is an effect of a thing on itself. Discriminant validity is the check that the constructs in a structural model are distinct enough for the paths between them to mean anything. It is where over-fitted models most often fail, and where the failure is most often hidden by reporting only the Fornell-Larcker table.
In the worked example, decide and mobility correlate at 0.46 with √AVE of 0.73 and 0.77. Distinct by every criterion.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Validity is an argument, and the CFA is one piece of it
A scale can have clean loadings, high omega and perfect discriminant validity and still not measure what it claims. Messick (1995) and the Standards for Educational and Psychological Testing (2014) frame validity as the evidence supporting a specific interpretation of scores for a specific use. The CFA is internal-structure evidence. The other kinds come from outside the model.
  • Content: do the items cover the construct as defined, judged by experts and by the people being measured? Cognitive interviews in the field are content evidence.
  • Response process: do respondents understand the items as intended? Think-aloud pilots.
  • Relations to other variables: does the score correlate with what it should (known-groups: members versus non-members) and not with what it should not (social desirability)?
  • Consequences: what happens when the score is used, and to whom?
The development-sector case
A decision-making scale validated in Bangladesh in 2005 is administered in Odisha in 2025 in Odia. The CFA fitting well says the Odia items cohere. It does not say they mean what the Bangla ones meant, that 'deciding' has the same content in a different household structure, or that a higher score is empowerment rather than a change in who is asked. Section 08's invariance tests address the first; the others need fieldwork, and a paragraph in the paper.
Write the validity argument as a paragraph with the CFA as one sentence of it. A table of AVEs is not a validity argument.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
A worked validity table for a three-construct scale
ConstructItemsωAVE√AVEr with decider with mobilityHTMT (max)
decide4 (d5 dropped)0.810.530.730.52
mobility30.810.590.770.460.52
assets40.780.480.690.380.310.44
Illustrative. 'assets' has AVE just under 0.5 with acceptable reliability; the paper says so rather than dropping a fourth item to cross the line. Every construct's √AVE exceeds its correlations, and every HTMT is below 0.85. The note gives the estimator (WLSMV), the sample, and that d5 was dropped for a stated reason before the structural model was fitted.
What the table is for
A reader can see in one glance that the three constructs are reliable enough, coherent enough and distinct enough to carry a structural model, and where the weak point is. That is the whole job of the measurement section. The structural results in section 07 are read against it: an effect of assets on decision-making is an effect of a construct with AVE 0.48, and the reader is entitled to weigh that.
Report this table for every SEM paper, before any path coefficient. Referees at good journals look for it first.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Known-groups and criterion validity: does the score behave as it should?
The cheapest external validity check is to ask whether the latent score differs between groups that theory says it must. A decision-making scale should score higher among women with their own income than without, higher among household heads than daughters-in-law, and should not differ by enumerator. Each is a latent mean comparison (section 04) or a regression of the factor on the grouping variable, and each either supports the interpretation or raises a question the CFA could not.
  • Choose the groups before the analysis, from the construct's theory. Post hoc 'known groups' are noise.
  • A difference in the wrong direction is a finding about the scale or the setting, not a nuisance.
  • Criterion validity: the score should predict an outcome measured differently (an administrative record, a partner's report, later behaviour). One such correlation is worth several fit indices.
Illustrative
Latent decision-making: women with own income 0.34 SD higher (95% CI 0.18 to 0.50); household heads 0.51 SD higher (0.29 to 0.73); no difference by enumerator (largest 0.06, CI including 0). The husband's independent report of who decides correlates 0.38 with the woman's latent score. Together with the CFA and the reliability table, that is a validity argument; the CFA alone is a coherence argument.
Plan the criterion into the survey. A partner's report or an administrative link costs a module and buys the strongest validity evidence available.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Validating a translated scale: the sequence
01
FORWARD translate by two independent translators; reconcile
02
BACK translate blind; compare with the source; resolve
03
EXPERT panel: content equivalence, cultural fit, reading level
04
COGNITIVE interviews in the field (n = 10–20 per language)
05
PILOT (n = 100–200 per language): item distributions, EFA
06
MAIN survey: CFA per language, then invariance tests across languages (section 08)
A South Asian evaluation often fields one instrument in three languages. Pooling the languages in one CFA assumes the items work the same way in each, which is exactly what has to be shown. The sequence above is the standard (Beaton and colleagues, Spine 2000, 25:3186, for health measures; the ITC guidelines for tests), and the last step is the one most often skipped.
Budget the cognitive interviews. They cost a fortnight and find the item that means something different in Bangla, which no statistic will.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
04
Section Four
Specification and Identification
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Specification: the model is the theory written as arrows
Every arrow in the structural model is a hypothesis with a direction, and every absent arrow is the hypothesis that two things are unrelated once the rest of the model is accounted for. Both come from theory, prior evidence, or the logic of the programme, and both are written down before the data are analysed. A model whose arrows were chosen to fit is not a test of anything; it is a description of one sample.
  • Start from the theory of change or the published theory (the theory of planned behaviour, Kabeer's resources-agency-achievements) and translate it construct by construct.
  • Decide the direction of each arrow from time order, mechanism or design. Cross-sectional data cannot decide it for you.
  • Decide what is exogenous. In an evaluation, programme exposure is; in a cross-sectional attitude survey, almost nothing is.
  • Write the list of omitted paths and why each is omitted. That list is the model's content.
Recursive and non-recursive
A recursive model has no feedback loops: arrows run one way and disturbances are uncorrelated. It is always identified given an identified measurement model. A non-recursive model (trust affects participation and participation affects trust) needs instruments, one excluded exogenous variable per loop, and is identified only if they exist. Most applied models are recursive and should say so; a feedback loop is a claim that needs an instrument, not an arrow.
Pre-register the diagram. OSF takes a PDF of it, and a registered model is one that referees cannot suspect of being found in the data.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Identification: can the parameters be estimated at all?
The data supply p(p + 1)/2 unique variances and covariances for p observed variables (plus p means if modelled). The model asks for some number of free parameters. If it asks for more than the data supply, it is under-identified and no unique solution exists. If exactly as many, just-identified: it fits perfectly and tests nothing. If fewer, over-identified, with degrees of freedom equal to the difference, and the fit test has something to test.
78
known moments for 12 indicators: 12 × 13 / 2
Counting rule
27
free parameters in a 3-factor CFA: 9 loadings, 12 error variances, 3 factor variances, 3 covariances
With three markers fixed
51
degrees of freedom: the model is over-identified and testable
78 − 27
  • The counting rule (df ≥ 0) is necessary, not sufficient. A model can pass it and still have unidentified parts.
  • Each latent variable needs a scale: fix one loading to 1 or the factor variance to 1.
  • The three-indicator rule: a single factor with three or more indicators, uncorrelated errors, is identified. With two indicators it is identified only if the factor correlates with something else in the model. One indicator needs its error variance fixed from a known reliability.
  • Software warns of non-identification with messages about non-positive-definite matrices, huge standard errors or failure to converge. Take the warnings literally.
Kline's chapter on identification and Bollen's chapter 4 give the rules for every case; the counting rule and the three-indicator rule cover most applied models.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Single indicators and observed variables in a latent model
Programme membership, age, household size and income are measured once and enter the model as observed variables, with arrows from rectangles. That is correct for variables measured without much error. For a single-item construct that is noisy (a one-question trust measure), the model can still separate error if the analyst fixes the error variance at (1 − reliability) × the observed variance, using a reliability from a prior study.
  • Do not build a 'latent' variable from one item with a free error variance; it is not identified and software will either fail or fix it silently.
  • A latent variable measured by a single composite score (a summed scale) is a common compromise: fix the error from the scale's omega. It loses the measurement test but keeps the error correction.
  • Categorical exogenous variables (state, caste category) enter as dummies, as in regression.
Programme exposure as the exogenous variable
In an evaluation, the structural model's first exogenous variable is treatment, observed and binary. Its paths into the latent outcomes are the programme effects, corrected for measurement error in the outcomes, and its path into a mediator and the mediator's path onward give the mechanism. Randomisation makes the treatment arrow causal; nothing makes the mediator arrow causal (section 10). The model is still worth fitting for the measurement correction alone.
Treatment as an observed exogenous variable is the one arrow in most SEMs that a reviewer will accept as causal without argument.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Covariates: where they go and what they do
A covariate that predicts an endogenous latent variable gets an arrow into it, from an observed rectangle. Every endogenous variable the covariate might affect gets one; leaving a covariate out of one equation is a restriction that must be true for the model to be right. Covariates correlate with each other and with the exogenous latents by default (double-headed arrows), and those correlations are estimated, not restricted.
  • Include the covariates the design needs (stratification variables, baseline outcome) and the confounders theory names. Not every variable in the dataset.
  • Each covariate costs parameters and degrees of freedom; with twelve covariates and three endogenous latents that is 36 paths. The model gets large and the fit indices harder to read.
  • An alternative for many covariates: residualise the indicators on the covariates first, then fit the SEM to the residuals. It is honest if stated and it keeps the model readable.
  • In a randomised evaluation, covariates in the treatment equation are not needed and their presence changes nothing except precision.
Clustered data
Respondents in villages, students in schools. The default standard errors assume independence and are too small. Options, in order of ease: cluster-robust standard errors (cluster = "village" in lavaan; vce(cluster) in Stata); a multilevel SEM if the village-level construct is itself of interest; or design-based estimation with survey weights (lavaan.survey). The first is enough for most evaluation uses and is the one referees now expect.
Survey weights change the estimates; clustering changes the standard errors. Both belong in the methods, with the level of clustering stated.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Equivalent models: the same fit, a different story
For almost any structural model there exist others, with different arrows, that imply exactly the same covariance matrix and therefore fit identically. Reverse a path between two exogenous-free variables, replace a path with a correlation, swap a mediator and an outcome: the χ² does not move. MacCallum and colleagues (Psychological Bulletin 1993, 114:185) found that published models routinely had dozens of equivalent alternatives and that authors never mentioned them.
  • Fit supports a model only relative to the alternatives that fit worse. It cannot choose among those that fit the same.
  • Stelzl's and Lee-Hershberger's replacing rules generate the equivalent models; Kline's chapter 8 walks through them.
  • The only things that break the equivalence are design (randomisation, time order) and theory strong enough to rule alternatives out.
The applied consequence
'Social capital increases savings behaviour (β = 0.34, p < .001), and the model fits well' is compatible with savings behaviour increasing social capital, and with both being driven by an unmeasured trait, at identical fit. The path coefficient is the same number under all three. A paper that presents one story as established by the fit has not understood what the fit tested.
State at least one equivalent model in the discussion and say why you prefer yours. It is a sentence, and it is the difference between a test and an assertion.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Means and intercepts: when the model needs them
The default SEM fits covariances only; means are ignored and every intercept is free. A mean structure adds the item intercepts and the latent means to the model and is needed for three purposes: comparing latent means across groups (does the treatment group have higher empowerment, on the latent scale?), measurement invariance beyond the metric level (section 08), and growth models where the trajectory's intercept and slope are the latent variables.
  • With a mean structure, the marker indicator's intercept is fixed to 0 (or the latent mean to 0 in a reference group) to identify the latent mean.
  • meanstructure = TRUE in lavaan; Stata's sem includes means by default.
  • Latent mean differences are reported in the latent metric, which has no units; standardise by the reference group's latent SD to get a Cohen's d.
Why the latent mean beats the summed score
A treatment effect on a summed score of five items weights each item equally and mixes measurement error into the estimate. The latent mean difference weights items by their loadings, separates error, and, under scalar invariance, is a comparison on the same scale in both groups. It is the right effect size for a latent outcome in an evaluation, and it comes with a standard error from the same model.
Report both. The summed-score difference is what a programme officer understands; the latent difference is what a methodologist trusts.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Before estimation: the specification checklist
QuestionAnswer written down
What are the constructs, and is each reflective or formative?A definition and a direction for each, with the source
Which items measure which construct, and why?The mapping, with the questionnaire numbers
Which structural paths exist, and in which direction?The diagram, with a one-line justification per arrow
Which paths are deliberately absent?The list, because it is what makes the model testable
Which variables are exogenous?Named; treatment if randomised, otherwise argued
Which covariates, into which equations?The list and the reason
Is the model recursive?Yes, or the instruments for each loop
Is it identified?The counting rule and the indicator rule, checked
Is the data clustered or weighted?The cluster variable and the weights
Are the items ordinal?Then the estimator is WLSMV and thresholds are in the model
Are means needed?Yes for group comparison, invariance or growth
Registered?Where, when, with the diagram
Twelve questions. Answered before the software opens, they become the methods section. Answered after, they become the things a referee asks about, and the answers tend to be whatever made the model fit.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
05
Section Five
Estimation, Data and Sample Size
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Estimators: ML, robust ML, and WLSMV
EstimatorAssumesUse whenNotes
ML (maximum likelihood)Multivariate normality of continuous indicators; complete data or FIMLContinuous, roughly normal items; the default in most softwareStandard errors and χ² wrong under non-normality, usually too small
MLR / MLM (robust ML)Continuous indicators; corrects SEs and χ² for non-normality (Satorra-Bentler)Continuous items that are skewed, which is most of themThe safe default for continuous data; lavaan estimator = "MLR"
FIMLData missing at random (MAR); used with ML/MLRAny missing data on the indicatorsUses every case; better than listwise deletion in almost every situation (Enders 2010)
WLSMV / DWLSLatent continuous response behind each ordinal itemLikert with fewer than five categories, binary itemsPairwise-present for missing data by default; needs several hundred cases
BayesianPriors on parametersSmall samples, complex models, informative prior knowledgeMplus and blavaan; report priors and sensitivity
ULS / GLSWeaker assumptions; less usedRarelyHistorical interest
PLSNone on distributions; composites, not factorsSection 09Not a CB-SEM estimator
The estimator follows from the data, not from the software's default. Five-point items with a skewed distribution and some missingness: WLSMV with pairwise or multiple imputation, or MLR with FIML if the categories are numerous enough. Report the estimator, the missing-data treatment and the reason for each.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Non-normality is the norm, and what it does
Attitude and behaviour items in field surveys are skewed: most respondents agree, or most report never. Under ordinary ML, skew and kurtosis inflate the χ² (the model is rejected too often), deflate the standard errors (paths look more significant than they are), and distort the fit indices in the direction of rejection. The parameter estimates themselves are mostly fine. The fix is a robust estimator and it costs nothing.
  • Check univariate skewness and kurtosis for every item; Mardia's multivariate kurtosis as a summary. Values well above 2 (skew) or 7 (kurtosis) on many items call for MLR or WLSMV.
  • Satorra-Bentler scaled χ² (1994; 2001 difference test) and the robust fit indices are what to report under MLR.
  • Bootstrapping the standard errors (Bollen-Stine for the χ²) is the alternative, and it is what mediation needs anyway (section 07).
  • Transforming items to reduce skew changes what they measure; do not.
Outliers and careless responses
A respondent who gave the same answer to forty items, or who finished a thirty-minute module in four, contributes a pattern no model should fit. Screen for straight-lining, implausible speed and multivariate outliers (Mahalanobis distance) before estimation, decide the rule in advance, and report how many cases were removed and why. Enumerator effects are the field version: a CFA that fails in one enumerator's interviews and fits in the rest is a data-quality finding.
Survey Design 101's high-frequency checks are the upstream fix. A clean dataset makes every slide in this section easier.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Missing data: FIML and multiple imputation, and what not to do
Listwise deletion (drop any case with any missing item) throws away information, biases estimates unless the data are missing completely at random, and can halve a sample with twenty items and 5% missingness each. Full-information maximum likelihood uses every case's available data under the weaker missing-at-random assumption and is the default in Mplus and one option away in lavaan (missing = "fiml"). Enders (Applied Missing Data Analysis, Guilford, 2010) is the reference.
  • FIML with MLR handles continuous items; for WLSMV, use multiple imputation (m = 20 or more) and pool with semTools' runMI, or accept pairwise-present with caution.
  • Auxiliary variables that predict missingness (age, enumerator, interview length) make MAR more plausible; include them via the saturated-correlates approach.
  • Report the missingness per item and its pattern. 'Data were complete' is rarely true in a field survey and is checked.
Not missing at random
Women who refuse the decision-making module may be the ones with least say. No estimator fixes that from the data alone. The honest treatments are a sensitivity analysis (pattern-mixture or selection models, or simply bounding the result under extreme assumptions) and a sentence in the limitations that says which direction the bias would run. 'Missing at random was assumed' is an assumption, and a reader is owed the reasoning.
Mean imputation and 'replace with the scale mean' shrink variances and inflate correlations. Never.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
How many respondents: the rules of thumb and the honest answer
RuleSaysStatus
n ≥ 200A floor for any SEMA folk rule; often too few for ordinal items or many parameters, sometimes more than needed
n:q ≥ 10 (Jackson 2003)Ten cases per free parameter; 20 preferredA better rule; 27 parameters need 270–540
Five to ten per indicatorFor CFACrude; ignores loadings and the number of factors
Power for RMSEA (MacCallum, Browne and Sugawara 1996)n for a given df to detect misfitAnswers the fit question, not the parameter question
Monte Carlo simulation (Muthén and Muthén 2002)Simulate the planned model with expected loadings and effects; find n for adequate power and biasThe right answer; simsem in R, Mplus's Monte Carlo
Wolf et al. 2013 (Educ Psychol Meas 73:913)Required n ranged from 30 to 460 across ordinary modelsThe evidence that no single rule works
Sample size depends on the number of indicators, the size of the loadings (strong loadings need fewer cases), the number and size of structural paths, the estimator (WLSMV wants more), missingness, and what power is wanted for which effect. A simulation with the planned model takes an afternoon and replaces every rule of thumb. For a field evaluation with three latent outcomes, four items each and binary items, expect to need 500 or more, and to say why.
Small samples do not make SEM impossible; they make convergence failures, improper solutions and wide intervals likely, and the paper must report all three rather than the one model that ran.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
When the model will not run, or runs wrong
SymptomLikely causeWhat to do
Non-convergence after many iterationsUnder-identification; poor starting values; an empirically weak factorCheck identification; simplify; supply starting values from a simpler model
Heywood case: negative error varianceOver-fitting; a factor with two indicators; an outlier; a misspecified modelDo not fix the variance to zero and move on; find the cause; consider dropping the two-indicator factor or merging
Standardised loading above 1As above, or a factor correlation near 1 (two factors are one)Discriminant validity check; merge factors
Correlation between factors above 1Two factors are oneMerge, or the model is wrong
Enormous standard errors on one parameterEmpirical under-identificationThe parameter is not estimable with these data; fix or drop it
Non-positive-definite covariance matrixLinear dependence among items; a mis-coded item; too few casesCheck the correlation matrix; find the duplicate
Different results across softwareDifferent defaults (marker, estimator, missing data, means)Match the defaults; report them
Every warning is information about the model or the data. Suppressing warnings, fixing variances at zero, or trying estimators until one converges are ways of hiding it. A paper that reports an improper solution and explains it is more credible than one that reports a clean one nobody believes.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
A worked estimation report
ItemChoiceReason
Samplen = 640 women in 32 villages; 14 cases removed for straight-lining (rule set before analysis)Design; pre-specified screening
Items11 items across three constructs, all binary or three-categoryQuestionnaire
EstimatorWLSMV on polychoric correlations, with thresholdsOrdinal items; Flora and Curran 2004
Missing data3.1% of item responses missing, at most 6% on any item; pairwise-present under WLSMV; multiple-imputation check (m = 20) gave the same loadings to two decimalsReported and checked
ClusteringVillage-clustered standard errorsSampling design
ScalingFirst loading of each factor fixed at 1; results reported standardisedConvention
SoftwareR 4.4, lavaan 0.6-19, semTools 0.5-6Reproducibility
ConvergenceNormal; no improper values; all standardised loadings between 0.41 and 0.82Checked
Illustrative. This is the estimation paragraph of the methods, as a table. Every row is a decision a reader might have made differently and can now see. It takes ten minutes to write and is absent from most SEM papers in the journals this course's readers publish in.
Software versions matter: lavaan's default handling of ordinal missing data and its robust-fit-index formulas have changed across versions.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Bootstrapping: standard errors and intervals without normality
Resample the cases with replacement, refit the model, record every parameter; repeat 1,000 to 5,000 times. The spread of the estimates is a standard error that assumes nothing about normality, and the percentiles are a confidence interval that can be asymmetric, which matters for products of coefficients (indirect effects) whose sampling distribution is skewed. Section 07 relies on it.
  • se = "bootstrap", bootstrap = 5000 in lavaan; vce(bootstrap) in Stata; a checkbox in AMOS.
  • Bias-corrected intervals were the standard (Preacher and Hayes 2008) and can be liberal; percentile intervals are now often preferred (Hayes 2022).
  • Bootstrap within clusters when the data are clustered, or the intervals are too narrow.
  • The Bollen-Stine bootstrap gives a χ² p-value robust to non-normality, as an alternative to the Satorra-Bentler correction.
Cost and reporting
Five thousand refits of an ordinal model with FIML can take an hour. Run it once, at the end, on the final model, and store the results. Report the number of resamples, the interval type, the random seed, and the interval itself rather than a p-value: 'indirect effect 0.11, 95% percentile bootstrap CI 0.04 to 0.19, 5,000 resamples' is the sentence.
The bootstrap resamples the respondents you have. It does not fix a biased sample, a wrong model or a missing confounder.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
06
Section Six
Model Fit and Modification
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
The chi-square test: the one exact test, and why nobody likes it
The model implies a covariance matrix; the data supply one; the χ² statistic, (n − 1) times the fit function, tests whether the difference is zero. A significant χ² means the model does not reproduce the data exactly. With a few hundred cases almost every model is rejected, because no model is exactly right and the test has power to see it; with fifty cases almost none is, because the test cannot see anything. The test is correct and its verdict is rarely the question.
  • Report it always: χ², df, p, and the scaling factor under MLR. Never omit it because it is significant.
  • χ²/df ratios ('below 3', 'below 5') have no theoretical basis and depend on n; retire them.
  • The difference between two nested models' χ² is a test of the restriction that separates them, and that use of the statistic is unambiguous and valuable (sections 07 and 08).
  • A non-significant χ² with n = 80 is not evidence of good fit; it is evidence of low power.
What a rejected model means
That something is missing or wrong: a cross-loading, a correlated error, a path, a non-linearity, an outlying group. The residual covariance matrix (observed minus implied, standardised) says where. Look at it before looking at the fit indices, which summarise the misfit into one number and lose the location. A large residual between two items of different factors is a finding about the questionnaire; a large residual between two items of one factor is a finding about the factor.
Standardised residuals above |2| are worth looking at; above |4| are worth explaining in the paper.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
The fit indices: what each measures and the conventional cutoffs
IndexMeasuresConventional thresholdBehaviour
RMSEA (Steiger-Lind; Browne and Cudeck 1993)Misfit per degree of freedom, in the population; with a 90% CI≤ 0.06 (Hu and Bentler 1999); ≤ 0.08 acceptable; CI upper bound < 0.10Penalises complexity; unreliable with small df; report the CI
CFI (Bentler 1990)Improvement over the null model of no correlations, 0 to 1≥ 0.95; ≥ 0.90 acceptableDepends on how bad the null is; inflated when items barely correlate
TLI / NNFIAs CFI, with a parsimony penalty; can exceed 1≥ 0.95As CFI
SRMRAverage standardised residual correlation≤ 0.08The most direct; not affected by n in the same way
AIC, BICLikelihood with a complexity penaltyLower is better; no absolute meaningFor comparing non-nested models on the same data
GFI, AGFIHistoricalDo not reportDepend on n; superseded
Hu and Bentler (Structural Equation Modeling 1999, 6:1) derived their cutoffs from simulations of particular models and warned against treating them as universal; the field ignored the warning. Marsh, Hau and Wen (2004) and Kline (2023) both argue the cutoffs are too strict for some models and too lenient for others, and that the residuals matter more. Report RMSEA with its CI, CFI, TLI and SRMR, alongside the χ², and interpret them together.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Reading fit with judgement: three cases
Caseχ² (df)RMSEA [90% CI]CFISRMRReading
A. 3-factor CFA, n = 640, WLSMV112.4 (41), p < .0010.052 [0.041, 0.064]0.960.048Rejected by χ² as expected at this n; approximate fit is acceptable on every index; check residuals for the location of misfit
B. Same model, n = 14048.7 (41), p = 0.190.037 [0.000, 0.078]0.970.071Not rejected, but the RMSEA CI runs to 0.078 and SRMR is near its limit: the data cannot distinguish good fit from mediocre
C. Structural model with 6 constructs, 24 items, n = 640612.0 (237), p < .0010.050 [0.045, 0.055]0.910.062CFI at the lenient threshold, RMSEA good: typical of large models where CFI suffers; acceptable, with the misfit located and reported
Illustrative. The indices disagree in each case, and the disagreement is information: about sample size in B, about model size in C. A paper that reports 'all fit indices met the recommended thresholds' has stopped where the reading starts.
Fit indices under WLSMV are computed differently and tend to look better than under ML for the same model. Compare within an estimator.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Modification indices: the road to a model that fits and means nothing
For every fixed parameter, the software reports the expected drop in χ² if it were freed. Free the largest, refit, repeat, and any model will reach any fit threshold. MacCallum, Roznowski and Necowitz (Psychological Bulletin 1992, 111:490) showed that such specification searches capitalise on chance: the modifications rarely replicate in a new sample, and the final model is a description of the sample's noise. This is the single most common way SEM papers in applied journals are made to 'fit'.
  • Look at modification indices to understand misfit, not to fix it.
  • Free a parameter only if it has a substantive reason you would have accepted before seeing the index: two items with near-identical wording, two items adjacent in the questionnaire, a path the theory predicts but you forgot.
  • Report every modification, its reason, and the fit before and after. A reader must be able to see the original model.
  • Cross-validate: modify on one half of the sample, test on the other. If the modified model does not fit the holdout, it was noise.
Correlated errors, in particular
The commonest modification is a correlated error between two items. It says the items share something beyond their factor. Sometimes true (a shared stem, a shared method), and then it should have been in the model from the start. Usually it is a sign that the factor is not what the items measure, and adding the correlation hides that. Five correlated errors in a twelve-item CFA is a different scale from the one specified, and the reliability and validity numbers must be recomputed on it.
A model with modifications is exploratory. Say so, and do not test hypotheses on it as if it were the pre-specified one.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Comparing nested models: the chi-square difference test
Two models are nested if one is the other with some parameters fixed. The difference in their χ² values, on the difference in their degrees of freedom, tests whether the restriction holds. This is the principled way to ask most SEM questions: is a path zero (fix it and compare)? Are two factors one (fix their correlation to 1 and compare)? Does the model hold across groups (section 08)? Is the second-order factor adequate (compare with the correlated-factors model)?
  • Under MLR or WLSMV, the scaled difference test (Satorra-Bentler 2001; lavaan's lavTestLRT) is needed; the raw difference of scaled χ² values is not χ² distributed.
  • With large n the test rejects small restrictions; report the change in CFI and RMSEA alongside (ΔCFI > 0.01 is the usual practical criterion, Cheung and Rensvold 2002).
  • For non-nested models (different indicators, different constructs), AIC and BIC, or out-of-sample prediction, not the difference test.
The model comparison as the result
A paper that fits one model and reports its indices has shown the model is not absurd. A paper that fits the theoretical model and two rivals (the mediation model against the direct-effects model; the three-factor scale against a one-factor version) and shows which the data prefer has shown something. Design the comparison before estimation; it is what turns fit into evidence.
In the worked CFA: the one-factor model has χ² = 418 on 44 df against 112 on 41. The three-factor structure is supported by a difference of 306 on 3 df, whatever the absolute indices say.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Reporting fit: the sentence and the table
The sentence
'The three-factor model fit the data acceptably: scaled χ²(41) = 112.4, p < .001; RMSEA = 0.052, 90% CI [0.041, 0.064]; CFI = 0.96; TLI = 0.95; SRMR = 0.048 (WLSMV, n = 640). The largest standardised residual (2.9) was between items d2 and a1; no modifications were made. A one-factor alternative fit substantially worse (Δχ²(3) = 306, ΔCFI = 0.19).'
Every number a reader needs, the misfit located, the alternative rejected, the modifications (none) stated. Four lines.
Modelχ² (df)RMSEA [CI]CFITLISRMRΔχ² (Δdf) vs M1
M1: three correlated factors112.4 (41)0.052 [0.041, 0.064]0.960.950.048
M2: one factor418.3 (44)0.115 [0.105, 0.126]0.770.710.104306 (3), p < .001
M3: second-order factor over the three115.0 (41)0.053 [0.042, 0.065]0.960.950.0502.6 (0), not nested with M1 in df; equivalent fit
M4: M1 with d5 dropped78.9 (32)0.048 [0.035, 0.061]0.970.960.041Different items; not comparable by χ²
Illustrative. The table is where the model comparison lives; the sentence is what the abstract carries.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
How fit is gamed, and how a reader can tell
PracticeWhat it doesTell-tale sign
Freeing correlated errors from modification indicesImproves fit by describing noiseSeveral error correlations with no substantive rationale; fit 'after modification' only
Deleting items until AVE and fit passChanges the construct to whatever the survivors measureA 12-item scale reported with 6 items and no account of the other 6
Item parcelling (averaging items into a few composites)Hides item-level misfit; inflates fitTwo or three 'indicators' per factor that are themselves averages; no item-level CFA
Reporting the model that convergedSurvivorshipNo mention of alternatives tried or of warnings
Reporting χ²/df instead of χ² and pObscures rejectionThe ratio without the components
Choosing lenient thresholds after seeing the valuesMoves the goalpostsCitations to whichever paper's cutoff the model clears
Fitting a saturated structural modelNothing to testdf of the structural part is zero; fit is the CFA's fit
Comparing WLSMV indices against ML cutoffsFlatters the modelWLSMV with CFI = 0.99 on a model that looks ordinary
None of these is fraud; each is a decision made after seeing the data and reported as if made before. The remedy is the same in every case: register the model, report the first fit and the final fit, and list what changed between them.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
07
Section Seven
Structural Models: Mediation and Moderation
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Measurement first, then structure: the two-step approach
Anderson and Gerbing (Psychological Bulletin 1988, 103:411) argued for fitting the measurement model as a CFA (all factors correlated) before imposing any structural paths, so that misfit in the measurement is found and fixed where it belongs and the structural model is tested against a measurement model that is known to hold. The structural model is nested in the CFA (it replaces free correlations with paths and zeros), so the χ² difference between them is the test of the structural restrictions alone.
01
CFA: all constructs, all correlated; assess and fix measurement
02
STRUCTURAL: replace correlations with the theorised paths
03
COMPARE: Δχ² between them tests the omitted paths
04
INTERPRET: paths, indirect effects, R² of each endogenous construct
  • If the structural model fits much worse than the CFA, the omitted paths matter; the residuals say which.
  • If the structural part is saturated (every construct connected to every other), it fits exactly as well as the CFA and tests nothing. Such models are common and their fit statistics are the CFA's.
  • Report the CFA fit and the structural fit separately, and the difference.
  • The one-step alternative (fit everything at once) is fine for a well-established scale, and hides measurement problems behind structural ones for a new one.
R² for each endogenous latent variable is reported alongside the paths; a significant path into a construct with R² = 0.04 explains little.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Mediation: direct, indirect and total effects
The model
M = a X + eM
Y = c′ X + b M + eY

Indirect effect of X on Y through M: ab. Direct effect: c′. Total: c′ + ab. In SEM the three are estimated together, with M and Y latent if they are measured by several items, and the indirect effect's standard error comes from the bootstrap (section 05).
Baron and Kenny's 1986 causal-steps procedure (test c, then a, then b, then see if c′ shrinks) is superseded: it has low power, it requires a total effect that need not exist when paths offset, and it never tests the indirect effect itself. Test ab directly with a bootstrap interval (Preacher and Hayes, Behavior Research Methods 2008, 40:879; Hayes, 2022).
  • The Sobel test assumes ab is normal; it is not, so the test is conservative. Report it only alongside the bootstrap.
  • 'Full' and 'partial' mediation are labels about whether c′ is significant, which depends on n; report the effects and their intervals and drop the labels (Zhao, Lynch and Chen, Journal of Consumer Research 2010, 37:197).
  • The proportion mediated, ab/(c′ + ab), is unstable when the total effect is small; report it with its interval or not at all.
  • Several mediators: specific indirect effects through each, and the contrast between them, all bootstrapped.
Every mediation estimate assumes no unmeasured confounding of the MY relationship. Randomising X does not deliver that. Section 10.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
A worked mediation: SHG membership, savings, decision-making
PathStd. estimateSE95% bootstrap CIReads as
a: membership → savings (latent)0.380.05[0.28, 0.48]Members score 0.38 SD higher on the savings factor
b: savings → decide (latent)0.310.06[0.19, 0.42]Savings behaviour predicts decision-making, membership held constant
c′: membership → decide (direct)0.090.05[−0.01, 0.19]Direct effect small and imprecise
ab: indirect0.120.03[0.07, 0.18]The indirect path is clear
Total (c′ + ab)0.210.05[0.11, 0.30]Overall association of membership with decision-making
R² savings0.14
R² decide0.12
Illustrative; 5,000 percentile bootstrap resamples, village-clustered; covariates age, education, household size into both endogenous constructs; WLSMV; structural fit RMSEA 0.049, CFI 0.95, SRMR 0.051.
What can be said
Membership was assigned by lottery in this design, so a and the total effect are causal. The b path and therefore the indirect effect assume that nothing unmeasured drives both savings and decision-making among members, which the design does not guarantee: a woman whose husband has migrated may save more and decide more, and migration is not in the model. The sentence in the paper: 'the association is consistent with mediation through savings; the mediator was not randomised'.
Two-thirds of the total effect runs through the savings path in this sample. That is a finding about where to look, not a proof of mechanism.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Moderation: when the effect depends on something else
A moderator changes the size or sign of an effect: membership raises decision-making more where women's baseline mobility is low, or less in households with a resident mother-in-law. Statistically, an interaction term. With observed variables it is the product X × W, mean-centred, as in regression. With a latent moderator or latent predictor the product of latents is needed, and that is where SEM has to work harder.
  • Multigroup: split by the moderator (two or three groups), fit the model in each, and test whether the path differs (equality constraint, Δχ²). Simple, and the right choice for a categorical moderator.
  • Product indicators (Marsh, Wen and Hau 2004): form products of the indicators of X and W as indicators of the latent interaction. Workable; sensitive to the pairing scheme.
  • Latent moderated structural equations (Klein and Moosbrugger, Psychometrika 2000, 65:457): the principled method; in Mplus (XWITH) and now in some R packages. No conventional fit indices.
  • Observed moderator: keep it observed and interact it with the latent predictor via the product-indicator route, or use multigroup with the moderator cut at meaningful points.
Interpreting an interaction
The coefficient on the product says how much the effect of X changes per unit of W. It is not interpretable alone. Plot simple slopes: the effect of X at low, mean and high W (one SD below and above), with confidence bands, or the Johnson-Neyman region where the effect is significant. 'The interaction was significant (β = 0.11)' is a sentence no reader can use; 'membership raised decision-making by 0.3 SD where baseline mobility was low and by 0.05 SD where it was high' is.
Interactions need more power than main effects: roughly four times the sample for the same precision. Plan for it or do not test them.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Moderated mediation: when the mechanism depends on context
The indirect effect may itself vary with a moderator: savings mediate membership's effect on decision-making where a bank is far away and not where one is near. Conditional process models (Preacher, Rucker and Hayes, Multivariate Behavioral Research 2007, 42:185; Hayes 2022) estimate the indirect effect at values of the moderator and test whether it changes, using the index of moderated mediation (Hayes, 2015): the slope of the indirect effect on the moderator, with a bootstrap interval.
  • Specify which path is moderated (a, b, or both) from theory, before estimation. Each is a different model.
  • Report the conditional indirect effects at meaningful moderator values with intervals, and the index with its interval.
  • Hayes's PROCESS macro (SPSS, SAS, R) does this with observed variables; SEM software does it with latents and is more work. Use PROCESS when the constructs are reliable summed scores and SEM when measurement is the point.
  • Every assumption of mediation applies, plus the moderator's exogeneity.
How much model is too much
A model with two mediators, two moderators and their products, on 300 respondents, has perhaps twenty structural parameters and the power to detect none of the interactions it exists to test. The result is a paper with one 'significant' conditional indirect effect out of eight tested, reported as the finding. Pre-specify one conditional hypothesis, power the study for it, and present the others as exploratory.
The diagram for a conditional process model should fit on one slide and be explicable in three sentences. If not, it is several studies.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Panel data: cross-lagged models and their trap
With the same constructs measured at two or more waves, the cross-lagged panel model regresses each construct at wave 2 on both constructs at wave 1, and the cross-lagged paths (trust at 1 to participation at 2, and the reverse) are read as evidence of direction. Hamaker, Kuiper and Grasman (Psychological Methods 2015, 20:102) showed the standard model confounds within-person change with stable between-person differences, and can find cross-lagged 'effects' that are entirely trait differences. The random-intercept cross-lagged panel model separates the two and needs three or more waves.
  • Two waves: a cross-lagged model is possible and its causal reading is weak; report it as association over time.
  • Three or more waves: RI-CLPM, with the within-person cross-lagged paths as the estimates of interest.
  • Measurement invariance across waves (section 08) is required before any of this; a change in the construct's measurement looks like change in the construct.
Latent growth models
With three or more waves of one construct, a latent growth model treats each respondent's trajectory as an intercept and a slope, both latent, with their means, variances and predictors. 'Did the programme change the slope of empowerment over three years, and for whom?' is a growth-model question. Multilevel models answer it too; the SEM version adds latent outcomes and measurement invariance tests, at the cost of needing balanced waves.
Longitudinal SEM is where measurement invariance stops being a formality. An item whose meaning shifts between baseline and endline manufactures a trajectory.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Effect sizes in SEM: what to report and in what units
QuantityReportInterpretation
Structural path, standardised (β)Estimate, SE, 95% CISD change in the outcome per SD change in the predictor, other predictors held constant
Structural path, unstandardised (B)Estimate, SE, CI, in the marker indicator's unitsNeeded when the predictor is binary (treatment): the effect in latent-outcome units
Treatment effect on a latent outcomeLatent mean difference / reference-group latent SDA Cohen's d corrected for measurement error
Indirect effectStandardised and unstandardised, bootstrap CIProduct of paths; its interval is the test
R² of each endogenous constructValueShare of the construct's variance explained; small values are common and honest
f² for a path(R²with − R²without) / (1 − R²with)Cohen's local effect size; 0.02, 0.15, 0.35 are his small, medium, large, and are too demanding for field data
Factor loadingsStandardised, with SEMeasurement quality; not effect sizes
Standardised paths are comparable within a model and not across samples with different variances. For evaluation reporting, the unstandardised treatment effect in a named unit (or the latent d) is what a programme reader needs, and it is rarely reported.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
The structural model in lavaan, with mediation and covariates
Syntax
model <- '
  # measurement
  savings =~ s1 + s2 + s3
  decide =~ d1 + d2 + d3 + d4
  # structural
  savings ~ a*member + age + educ + hhsize
  decide ~ cp*member + b*savings + age + educ + hhsize
  # effects
  ab := a*b
  total := cp + a*b
'
fit <- sem(model, data = df, ordered = c("s1",...,"d4"),
    cluster = "village", se = "bootstrap", bootstrap = 5000)
Labels (a*, b*) name parameters; := defines a new quantity from them and gives it a standard error and, with the bootstrap, an interval. The covariates enter both endogenous equations. Note the limitation: lavaan's bootstrap and its ordinal estimator do not combine in every version, and clustered bootstrapping needs care; check the current documentation and say what was done.
  • Stata: sem (Savings -> s1 s2 s3) (Decide -> d1 d2 d3 d4) (Savings <- member age educ hhsize) (Decide <- member Savings age educ hhsize), vce(cluster village), then estat teffects for indirect effects, and bootstrap around it.
  • AMOS: draw it; tick 'indirect, direct and total effects' and 'perform bootstrap'.
  • Mplus: MODEL INDIRECT; the most complete for latent interactions.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Structural-model pitfalls
PitfallConsequenceFix
Saturated structural modelNothing tested; fit is the CFA'sOmit paths the theory does not need; compare with alternatives
Baron-Kenny steps as the mediation testLow power; wrong logicBootstrap the indirect effect
'Full mediation' claimed from a non-significant direct pathSample-size artefactReport effects and intervals; drop the labels
Cross-sectional mediation read as processMaxwell and Cole (2007) show the bias can be any size and signLongitudinal design, or say 'consistent with'
Interaction without simple slopesUninterpretablePlot at ±1 SD; Johnson-Neyman
Many mediators and moderators on a small sampleOne significant result out of many; a finding by chancePre-specify one; treat others as exploratory
Standardised effects of a binary treatmentMeaningless SD of a dummyUnstandardised, or the latent d
Covariates in some equations and not others without reasonHidden restrictionsState the rule and apply it
Reversed arrows fit identicallyDirection asserted, not testedName the equivalent model; use design
The last row is the one that turns a competent SEM into an over-claimed one, and section 10 is about it.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
08
Section Eight
Measurement Invariance
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Are we measuring the same thing in both groups?
Comparing women in Bihar and Tamil Nadu on a decision-making scale, or the same women before and after a programme, or respondents who answered in Hindi and in Bangla, assumes that the items relate to the construct in the same way in each group. If a Bangla item loads more weakly, or has a different intercept (Bangla speakers answer 'yes' more often at the same true level), the group difference in scores mixes real difference with measurement difference and cannot be separated. Invariance testing checks the assumption, level by level.
Measurement invariance
The property that the measurement model (loadings, intercepts, error variances) is the same across groups or occasions, so that differences in latent means and relationships reflect the construct and not the instrument. Vandenberg and Lance, Organizational Research Methods 2000, 3:4.
  • Required before comparing latent means across groups, before pooling groups in one model, and before longitudinal models.
  • Almost never tested in South Asian multi-language surveys, and almost always partly violated when it is.
  • The test is a sequence of nested multigroup models with increasing equality constraints, compared by Δχ² and ΔCFI.
  • Partial invariance (most items invariant, a few freed) is the usual outcome and is enough for most comparisons if handled honestly (Byrne, Shavelson and Muthén, Psychological Bulletin 1989, 105:456).
Item Response Theory 101 in this series covers differential item functioning, which is the same question asked item by item.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Configural, metric, scalar, strict: the sequence
LevelConstrained equal across groupsLicensesTypical outcome
ConfiguralNothing; the same pattern of loadingsThe construct has the same structureUsually holds
Metric (weak)LoadingsComparing relationships (paths, correlations) across groupsUsually holds, or nearly
Scalar (strong)Loadings and intercepts (thresholds for ordinal items)Comparing latent means; pooling groupsOften fails on one or two items; partial scalar invariance is the common result
StrictPlus error variancesComparing observed (summed) scoresOften fails; rarely needed, since latent comparisons do not require it
Each level is nested in the one above; the Δχ² tests whether the added constraints hold. Because Δχ² rejects trivial differences at large n, the practical criteria are ΔCFI ≤ 0.01 (Cheung and Rensvold, Structural Equation Modeling 2002, 9:233) and ΔRMSEA ≤ 0.015 (Chen 2007), used together.
For ordinal items the sequence differs: thresholds and loadings are tested in a particular order (Wu and Estabrook, Psychometrika 2016, 81:1014); semTools' measEq.syntax writes the models correctly.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
A worked invariance test: Hindi and Bangla versions of a scale
Modelχ² (df)CFIRMSEAΔχ² (Δdf), pΔCFIVerdict
Configural96.2 (82)0.9810.033Same structure in both
Metric108.9 (90)0.9750.03612.7 (8), 0.12−0.006Loadings equal: holds
Scalar151.3 (98)0.9300.05842.4 (8), < .001−0.045Fails
Partial scalar (d3 and m2 thresholds freed)117.6 (94)0.9690.0408.7 (4) vs metric, 0.07−0.006Holds with two items freed
Latent mean difference (Bangla − Hindi), under partial scalar−0.18 SD, 95% CI [−0.36, 0.00]
Illustrative; n = 320 per language; WLSMV; thresholds rather than intercepts because the items are ordinal. Two items behave differently across languages: d3 ('visiting your natal family') and m2 ('going to the market alone').
What it means, and what to do
The scale measures the same construct in both languages (configural, metric), but two items are answered differently at the same true level, plausibly because the practices they name differ between the regions rather than because of translation. With those two freed, the latent means can be compared, and the Bangla group is slightly lower, imprecisely. Had the scalar failure been ignored and summed scores compared, the difference would have been larger and partly an artefact of d3 and m2.
Go back to the cognitive interviews for d3 and m2. The statistics found the items; the fieldwork explains them.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Partial invariance: how much is enough, and how to find the items
When scalar invariance fails, the aim is to find the few items responsible, free their intercepts or thresholds, and check that the rest hold. Modification indices in the constrained model point to the items; free them one at a time, largest first, retesting after each, and stop when ΔCFI against the metric model is within tolerance. Byrne, Shavelson and Muthén's rule is that at least two invariant items per factor (the marker plus one) are needed for the latent mean comparison to be identified and meaningful; more is better.
  • Freeing more than a third of a factor's items is not partial invariance; it is a different scale in each group.
  • Report which items were freed, in which direction they differ, and what the latent mean difference is with and without the freeing.
  • The freed items are findings about the instrument and the setting. They belong in the discussion, not a footnote.
  • Alternative: the alignment method (Asparouhov and Muthén, Structural Equation Modeling 2014, 21:495) for many groups (all Indian states) where sequential testing is impractical.
When invariance really fails
Sometimes the construct is different in the two groups: 'mobility' for a woman in a Bihar village and in a Chennai apartment may not be one thing. Then no amount of freeing rescues the comparison, and the honest report is that the groups cannot be compared on this scale, with the evidence. That is a result, and a more useful one for the next study than a forced comparison.
Test invariance on the pilot data if the pilot is large enough. Finding a non-invariant item before the main survey means it can be rewritten.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Invariance across waves: the same scale before and after
Longitudinal invariance is tested the same way, with waves as 'groups' in a single model where the same respondents appear at each wave and each item's error is allowed to correlate with itself across waves (the same person's idiosyncratic response to item d2 persists). Without it, a change in the latent mean between baseline and endline could be a change in how the items are understood, which a programme that talks to women about decision-making might well produce.
  • Response shift: the programme changes the respondent's frame of reference, so the same answer means something different. Invariance testing detects it as a scalar failure at endline.
  • Enumerator change between waves is confounded with time; note it.
  • With two waves and a control group, test invariance across waves within the control group first; if it holds there, a failure in the treatment group is itself evidence of response shift.
Why it matters for evaluations
An empowerment programme that raises the summed score by 0.3 SD may have raised the construct, or taught respondents what the questions are asking, or both. The treatment-group scalar failure at endline is the diagnostic, and the latent mean difference under partial invariance is the corrected estimate. Most evaluations report neither, and their effects on self-reported attitudes are correspondingly hard to read.
Behavioural items (did you go to the market alone last week?) are less prone to response shift than attitudinal ones. Mix them.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Multigroup structural models: does the effect differ by group?
Once metric invariance holds, structural paths can be compared across groups: is the effect of savings on decision-making the same for women with and without their own income? Fit the structural model in both groups at once, then constrain the path equal and test the constraint. This is moderation by a categorical variable (section 07), done with full measurement models and a proper test.
  • Constrain one path at a time and test each; an omnibus test of all paths equal is rarely the question.
  • Report the path in each group with its CI, and the Δχ² for the equality constraint.
  • Group sizes: each group needs to be large enough for its own model; 150 per group is a practical floor for a modest model with ordinal items.
  • In lavaan: group = "language", group.equal = c("loadings", "intercepts"), then equality labels on the paths.
Many groups
Twenty-eight states, or thirty districts, cannot be handled by pairwise multigroup models. Options: multilevel SEM with the group as a random level, the alignment method for invariance, or collapsing to a few theoretically meaningful groups (north and south; high and low female labour force participation). The last is usually the right one and the grouping should be stated before the data are seen.
Group differences in a structural path are among the most useful things SEM can show a programme, because they say for whom the mechanism works.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Invariance testing in practice: the code and the report
lavaan with semTools
library(semTools)
cfg <- cfa(model, df, group = "lang", ordered = items)
met <- cfa(model, df, group = "lang", ordered = items,
   group.equal = "loadings")
sca <- cfa(model, df, group = "lang", ordered = items,
   group.equal = c("loadings", "thresholds"))
compareFit(cfg, met, sca)
lavTestScore(sca)  # which constraints to release
# or, for the correct ordinal sequence:
measEq.syntax(model, df, group = "lang", ordered = items,
   ID.cat = "Wu.Estabrook.2016", group.equal = ...)
Stata: sem ..., group(lang) ginvariant(mcoef) and estat ginvariant. Mplus: MODEL = CONFIGURAL METRIC SCALAR in one line. AMOS: the multigroup wizard. All produce the same table; the judgement about partial invariance is the analyst's.
  • Report the table of nested models with χ², df, CFI, RMSEA, the differences and the criteria used.
  • Report the freed items and the reasoning.
  • Report the latent mean difference (or the group path difference) under the final model, with its interval.
  • Say which comparisons the achieved level licenses and which it does not.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
09
Section Nine
PLS-SEM
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
PLS-SEM: composites, regressions, and a different question
Partial least squares path modelling (Wold, 1982; Lohmöller, 1989) builds each construct as a weighted composite of its indicators, with weights chosen iteratively to maximise the explained variance of the endogenous composites, then estimates the structural paths by ordinary regression among the composites. There is no common factor, no separation of measurement error, and no overall fit function to test. What there is: a prediction-oriented model that runs on small samples, handles formative constructs natively, and converges when CB-SEM will not.
Composite
A weighted sum of indicators, treated as the construct. Unlike a common factor, it contains the indicators' error, and its path coefficients are attenuated accordingly unless the consistent PLS correction (PLSc) is applied.
  • Standard in marketing, information systems and management; the dominant method in South Asian business-school research, usually via SmartPLS.
  • Hair, Hult, Ringle and Sarstedt, A Primer on PLS-SEM (3rd ed., Sage, 2022) is the manual; Hair and colleagues, European Business Review 2019, 31:2, the reporting guide.
  • Its critics (Rönkkö and Evermann 2013; Rönkkö, McIntosh, Antonakis and Edwards, Journal of Operations Management 2016, 47–48:9) argue that most of its claimed advantages do not survive simulation and that it is often used because it makes weak models look supported.
  • Both sides are right about something. This section says what each is right about.
The honest statement of purpose: PLS-SEM estimates a predictive model on composites. It is not a cheaper way to do CB-SEM.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
When PLS-SEM is the right tool, and when it is the convenient one
A defensible reasonWhy
The constructs are formative composites (an index of adoption barriers, a service-quality composite)CB-SEM handles formative constructs awkwardly; PLS is built for them
The goal is prediction of a target construct, evaluated out of samplePLS maximises explained variance and PLSpredict assesses it honestly
The model is complex relative to the sample and CB-SEM will not convergePLS always converges; the estimates are still to be read with the composite caveat
Exploratory theory building in a new domainNo fit test to fail; the structure can be revised and reported as exploratory
Secondary data with single-item and archival measuresComposites of one item are fine; factors of one item are not
A convenient reasonWhy it does not hold
'The sample is small' (n = 120)Goodhue, Lewis and Thompson (MIS Quarterly 2012, 36:981): PLS has no more power than regression at small n; it just produces numbers
'The data are non-normal'Robust ML and WLSMV handle non-normality in CB-SEM; PLS's distribution-free property is about estimation, not inference
'PLS does not require a fit test'That is a cost, not a benefit; the model is not tested against the data
'The 10-times rule says 100 is enough'The rule is discredited; use the inverse-square-root method (Kock and Hadaya, Information Systems Journal 2018, 28:227) or a power analysis
'Everyone in my field uses it'True, and the field's replication record is the argument against
Choose the method for the question, state the reason in the methods, and cite both the advocates and the critics. Referees at good journals are now from both camps.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Assessing the measurement (outer) model in PLS-SEM
Construct typeCriterionThreshold (Hair et al. 2019, 2022)Notes
ReflectiveIndicator loadings≥ 0.708 (so loading², the indicator reliability, ≥ 0.5)Loadings 0.4–0.7 retained only if deletion does not raise CR or AVE above threshold
ReflectiveInternal consistency: Cronbach's α, ρA (Dijkstra-Henseler), composite reliability ρc0.70–0.95; ρA is the recommended oneAbove 0.95 signals redundant items
ReflectiveConvergent validity: AVE≥ 0.50As in CB-SEM
ReflectiveDiscriminant validity: HTMT< 0.85 (conceptually distinct) or 0.90; bootstrap CI excludes 1Fornell-Larcker and cross-loadings are no longer sufficient
FormativeConvergent validity: redundancy analysisCorrelation ≥ 0.70 with a single global item measuring the same constructRequires the global item to have been asked; plan it in the questionnaire
FormativeCollinearity: VIF of indicators< 3 (ideal), < 5 (tolerable)High VIF means indicators overlap; the weights become unstable
FormativeOuter weights significance and relevanceWeight significant by bootstrap; if not, retain if loading ≥ 0.5 and theory supportsDeleting a formative indicator changes the construct
Every row is a decision that removes or keeps an item, and the recommended practice is to report the initial and final sets with reasons. A model whose outer assessment deleted a third of the items has changed its constructs and should say so.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Assessing the structural (inner) model and predictive power
CriterionThreshold or readingNotes
Collinearity among predictor composites (VIF)< 3Above 5 the path coefficients are unreliable
Path coefficientsSign, size, bootstrap CI (10,000 subsamples, percentile)The estimate and interval, not stars
R² of endogenous constructsHair's 0.75 / 0.50 / 0.25 as substantial / moderate / weak are from marketing and far too demanding for field data; report and interpret in contextAdjusted R² when comparing models
f² effect size0.02 / 0.15 / 0.35 (Cohen)The change in R² when a predictor is removed
Predictive relevancePLSpredict (Shmueli et al., European Journal of Marketing 2019, 53:2322): out-of-sample RMSE of the PLS model versus a linear-model benchmark on the indicatorsReplaces the older blindfolding Q²; if PLS does not beat the benchmark on most indicators, the model has no predictive power
Model fit (approximate)SRMR < 0.08 (Henseler et al. 2014); dULS, dG bootstrap testsContested; report SRMR and say it is approximate
Mediation, moderationAs in CB-SEM: bootstrapped indirect effects; interaction terms with simple slopesTwo-stage approach for interactions
PLSpredict is the important addition. A PLS-SEM paper that reports R² and path significance without out-of-sample prediction has not shown what the method is for. Most do not report it.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
A worked PLS-SEM assessment: mobile-money adoption among traders
ConstructTypeItemsρAAVEHTMT maxNotes
Performance expectancyReflective40.860.640.71All loadings > 0.75
Effort expectancyReflective30.810.610.71
Social influenceReflective30.780.580.55One loading 0.66, retained
Adoption barriersFormative5VIF < 2.4; redundancy r = 0.74; two weights not significant, loadings > 0.5, retained
Intention to adoptReflective30.890.730.710.42
Use (self-reported transactions)Single item10.19PLSpredict: PLS RMSE lower than LM benchmark on the item; Q²predict = 0.11
Illustrative; n = 310 small traders in two districts, UTAUT-type model (Venkatesh and colleagues, MIS Quarterly 2012, 36:157). Paths: performance expectancy → intention 0.31 [0.19, 0.42]; barriers → intention −0.24 [−0.35, −0.13]; intention → use 0.44 [0.33, 0.54]; 10,000 bootstrap subsamples.
Reading it
The measurement is adequate on every criterion, the barriers composite is validly formative, and intention explains 42% of its own variance and predicts reported use out of sample, modestly. That is what a PLS-SEM study can show. It cannot show that raising performance expectancy would raise adoption, that the constructs are distinct common factors, or that the model fits in the CB-SEM sense; the paper should not say any of those.
R² = 0.42 is 'weak' by Hair's marketing benchmarks and respectable for a field survey. Say which yardstick you are using.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Consistent PLS and the composite-versus-factor question
Ordinary PLS path coefficients between reflective constructs are biased toward zero, because the composites carry measurement error, and loadings are biased upward. Consistent PLS (Dijkstra and Henseler, MIS Quarterly 2015, 39:297) corrects both using ρA, so that PLSc estimates for a common-factor model are consistent, and for a well-specified model come close to CB-SEM's. It is one click in SmartPLS and one argument in SEMinR and cSEM.
  • If the constructs are argued to be common factors (attitudes, traits), use PLSc or use CB-SEM.
  • If they are argued to be composites (indices, formed constructs), ordinary PLS is the right estimator and the 'bias' is not a bias.
  • Say which. The composite-factor distinction is the actual disagreement between the PLS advocates and their critics, and a paper that names its constructs' status has answered most of the objection.
  • Report both estimates when the choice is contested; if they differ substantially, the measurement error is large and that is worth knowing.
Sample size for PLS-SEM
The '10 times the largest number of arrows into a construct' rule has no statistical basis and Hair and colleagues (2019) now reject it. The inverse-square-root method gives the n needed to detect a path of a given size at 80% power: about 155 for a standardised path of 0.2, and about 620 for 0.1 (Kock and Hadaya 2018). A power analysis for the smallest path of interest is the honest answer, as in CB-SEM.
PLS on 80 respondents does produce estimates. So does regression. Neither has the power to learn anything from them.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
PLS-SEM software: SmartPLS, SEMinR, cSEM
ToolCostStrengthsLimits
SmartPLS 4 (Ringle, Wende and Becker)Licence; free student and time-limited trial licencesMenu-driven, complete: PLSc, bootstrapping, PLSpredict, HTMT with CIs, mediation, moderation, multigroup, necessary condition analysis, and a CB-SEM module since version 4Cost; point-and-click makes the analysis hard to reproduce unless the project file is shared
SEMinR (R)FreeReadable syntax that mirrors the PLS-SEM vocabulary; bootstrapping, PLSc, PLSpredict, plots; the Hair et al. companion packageR
cSEM (R)FreeComposite and factor models in one framework; tests of overall fit; the methodologically most completeLess friendly documentation
plspm (R)FreeThe older standard; still worksFewer recent methods
ADANCOLicenceHenseler's package; strong on composite models and fit testsSmaller user base
Python: semopy (CB), plspm portFreeFor pipelinesPLS coverage is thinner than R's
For a student in a South Asian business school where SmartPLS is taught: learn it, and reproduce one analysis in SEMinR so that the work can be shared as code. For anyone else: SEMinR or cSEM.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Reporting a PLS-SEM study: the Hair et al. 2019 checklist, condensed
SectionReport
JustificationWhy PLS-SEM rather than CB-SEM, in terms of the question and the constructs' status (composite or factor)
ModelThe diagram; each construct's measurement type; the theory behind each path
DataSample, sampling, n, missing data and treatment, distributional summary, power analysis or inverse-square-root n
SettingsSoftware and version; weighting scheme (path); PLSc or not; bootstrap subsamples (10,000), interval type; PLSpredict folds and repetitions; seed
Outer modelThe full table from the outer-model slide, initial and final item sets, deletions with reasons
Inner modelVIF; paths with bootstrap CIs; R², f²; PLSpredict results against the LM benchmark; SRMR
Additional analysesMediation, moderation, multigroup (with MICOM invariance test for composites), as pre-specified
LimitationsCross-sectional, self-report, composite measurement, the causal status of the paths
AvailabilityThe data and the SmartPLS project file or R script
Papers that follow this list are rare and reviewable. Papers that report loadings, AVEs, R² and starred paths are the norm and are the reason the method has the reputation it has.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
PLS-SEM pitfalls, in the order they appear in the typical thesis
PitfallFix
Chosen because CB-SEM would not fitSay so, and present the CB-SEM misfit as a finding about the model
Reflective constructs with ordinary PLS and no PLScPLSc, or argue the composite status
Fornell-Larcker only for discriminant validityHTMT with bootstrap CI
R² interpreted with marketing benchmarksContextual interpretation; compare with the literature's values
No out-of-sample predictionPLSpredict against the LM benchmark
Item deletion to pass thresholdsReport initial and final sets; do not delete formative indicators
5,000 bootstrap samples with bias-corrected intervals reported as p < 0.001 on everything10,000 percentile; report intervals
Multigroup comparison without MICOMTest measurement invariance of composites first (Henseler, Ringle and Sarstedt 2016)
Convenience sample of 150 students generalised to 'consumers in India'Describe the sample as what it is
Causal language throughoutSection 10: 'is associated with', 'predicts'
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Running both: CB-SEM and PLS-SEM on the same data as a robustness check
When the constructs are reflective and the sample allows it, fitting the model both ways is cheap and informative. If the CB-SEM fits and the PLSc paths agree with it, the result does not depend on the estimator. If CB-SEM rejects the model while PLS reports strong paths, the paths are estimated on a structure the data contradict, and the PLS result is the weaker one. Papers that do this are rare; referees notice them.
  • Same items, same constructs, same paths; CB-SEM with MLR or WLSMV, PLS with PLSc.
  • Compare the path coefficients and intervals in one table; report the CB-SEM fit.
  • SmartPLS 4 runs both; in R, lavaan and SEMinR share the data frame.
The disagreement between the two camps is partly a disagreement about what a construct is. The analyst who can say which kind theirs are, and show the result under both estimators, has stepped out of the argument.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
10
Section Ten
Causality and Common Method Bias
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
SEM does not establish causation, and never did
The arrows in a path diagram are causal hypotheses. Estimating them assumes the causal structure; it does not test it. Bollen and Pearl ('Eight myths about causality and structural equation models', in Morgan, ed., Handbook of Causal Analysis for Social Research, Springer, 2013) list the misreadings on both sides: that SEM proves causation (it cannot), and that SEM is merely correlational (it is a language for stating causal assumptions and deriving what they imply). The coefficients are causal effects if, and only if, the assumptions hold.
  • Every path XY assumes no unmeasured common cause of X and Y, no reverse effect, and no selection on the outcome.
  • Randomising X secures the assumptions for X's paths and for nothing downstream of it.
  • Good fit is compatible with the arrows reversed (section 04) and with an omitted confounder that the model absorbs into a path.
  • The language follows the design: 'effect' and 'causes' for randomised paths; 'is associated with', 'predicts', 'is consistent with' for the rest.
Antonakis's list
Antonakis, Bendahan, Jacquart and Lalive ('On making causal claims', Leadership Quarterly 2010, 21:1086) reviewed a field's SEM papers and found the large majority made causal claims their designs could not support, through omitted variables, simultaneity, measurement error in predictors treated as observed, selection, and common-method variance. The paper is written for management research and applies without change to South Asian development research using the same methods.
Draw the DAG with the unmeasured variables in it. If an unmeasured node points into both ends of a path, that path is not identified by the SEM.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Mediation's hidden assumption, and the sensitivity analysis that exposes it
Imai, Keele and Tingley (Psychological Methods 2010, 15:309) formalised what mediation needs: sequential ignorability, meaning that the treatment is as good as random given covariates, and that the mediator is as good as random given treatment and covariates. The first is delivered by randomisation; the second is delivered by nothing, because the mediator was not assigned. Their sensitivity analysis asks how strong an unmeasured confounder of the mediator-outcome path would have to be (the correlation ρ between the two error terms) to reduce the indirect effect to zero.
  • R's mediation package implements it; report the ρ at which the effect vanishes and the R² it corresponds to.
  • An indirect effect that disappears at ρ = 0.1 is fragile; at 0.4 it is not easily explained away.
  • Bullock, Green and Ha (JPSP 2010, 98:550) show how badly mediation estimates behave when the mediator is confounded, and argue for designs that manipulate the mediator directly.
  • Cross-sectional mediation adds a second problem (Maxwell and Cole 2007): the cross-sectional estimate of a longitudinal process can have the wrong sign.
What a careful paper says
'Membership was randomised; its total effect on decision-making is causal. The indirect effect through savings assumes no unmeasured confounder of savings and decision-making among members. A sensitivity analysis indicates the indirect effect would be eliminated by a confounder correlating with both error terms at ρ = 0.28. We interpret the mediation result as consistent with a savings mechanism and not as a test of it.'
Four sentences. They distinguish a paper that understands its method from one that runs it.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Common method bias: when the instrument makes the correlations
Every construct in the model measured from one respondent, in one sitting, on the same response format, shares variance that has nothing to do with the constructs: the respondent's mood, acquiescence, social desirability, the desire to be consistent, the enumerator's manner. Podsakoff, MacKenzie, Lee and Podsakoff (Journal of Applied Psychology 2003, 88:879) catalogued the sources and the remedies; their 2012 review (Annual Review of Psychology 63:539) updated them. The bias inflates correlations among self-reported constructs and therefore every path between them.
Common method variance
Variance attributable to the measurement method rather than to the constructs the measures represent. It is present in every single-source, single-occasion survey; the question is how much.
  • Procedural remedies, at design: different sources for predictor and outcome (the woman's report and the husband's; self-report and an administrative record); temporal separation; different response formats; behavioural rather than attitudinal outcomes; assurance of anonymity.
  • Statistical remedies, after: a marker variable theoretically unrelated to the constructs (Lindell and Whitney 2001; the CFA marker technique of Williams, Hartman and Cavazotte, Organizational Research Methods 2010, 13:477); an unmeasured latent method factor with equal loadings on every item.
  • Not a remedy: Harman's single-factor test (one unrotated factor explaining less than 50%). It has almost no power and detects nothing; its presence in a paper signals that the authors looked for the cheapest possible check.
A programme's effect on self-reported empowerment, measured by the programme's own enumerators, from the participant, with the mediator on the same page, has every source of method variance at once.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
A method factor in the model: how to do it and what it shows
Modelχ² (df)CFIMethod variance shareKey path (savings → decide)Reading
Substantive CFA only112.4 (41)0.960.31 [0.19, 0.42]Baseline
Plus unmeasured latent method factor (equal loadings)104.1 (40)0.978%0.27 [0.15, 0.39]Some method variance; path attenuated a little
Plus CFA marker (attitude to cricket, 3 items)128.6 (60)0.965%0.28 [0.16, 0.40]Consistent with the ULMC
Outcome from a different source (husband's report of who decides)n/an/an/a0.19 [0.06, 0.32]The procedural check: smaller, still present
Illustrative. The three statistical approaches agree that method variance is present and modest and that the path survives it. The different-source estimate is the most convincing and the smallest, which is the usual pattern and the reason procedural remedies beat statistical ones.
Reporting
State the procedural steps taken at design and the statistical check run after, with the method variance share and the key paths before and after. If no procedural step was possible, say so and treat the paths among self-reported constructs as upper bounds. A marker variable has to be planned into the questionnaire: three items on something the constructs cannot plausibly relate to, on the same response scale.
The ULMC with equal loadings is an approximation; with free loadings it is not identified. Report which was used.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Selection, reverse causality and instruments in SEM
Women who join a self-help group differ from those who do not in ways the model does not measure, and those ways affect decision-making. Membership is endogenous, and its path in an observational SEM is biased by whatever drives both. This is the same problem as in any regression, and SEM offers the same remedies: an instrument (a variable that affects membership and affects the outcome only through it), a design (randomisation, a discontinuity), or an honest limitation.
  • An instrument enters the SEM as an exogenous variable with a path into membership and no path into the outcome; the disturbances of membership and outcome are allowed to correlate. This is 2SLS inside a latent model, and its assumptions are 2SLS's.
  • A weak or invalid instrument (distance to the nearest group, a common choice, affects many things) makes it worse, not better.
  • Impact Evaluation 101 covers the designs. SEM adds latent outcomes and mediation to them; it does not replace them.
The typical South Asian SEM paper
A cross-sectional survey, all constructs self-reported, a convenience or purposive sample, a model with six constructs and twelve paths, all significant, fit indices at threshold, and a conclusion that construct A 'significantly impacts' construct F. Every problem in this section is present and none is mentioned. The same data, analysed the same way and reported with sections 10 and 11's caveats, would be a competent descriptive study. The difference is only in what is claimed.
Write the limitations paragraph before the results. It disciplines the verbs in the results.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Drawing the causal diagram before the path diagram
A directed acyclic graph is a path diagram with the unmeasured variables drawn in and no estimation implied. Drawing one for the study, including the things you cannot measure (household bargaining power, the husband's attitudes, the enumerator's effect), shows which paths are identified by adjustment, which need a design, and which are hopeless. Pearl's back-door criterion gives the rule: a path XY is identified if every back-door path from X to Y is blocked by measured covariates.
  • DAGitty (free, browser) draws the graph and lists the adjustment sets and the testable implications.
  • A collider (a variable caused by both X and Y) must not be conditioned on; conditioning on it creates a spurious association. Selecting the sample on the outcome, or controlling for a post-treatment variable, does this routinely.
  • The DAG is the pre-registration document that a path diagram alone is not.
Post-treatment bias, the common case
Controlling for savings when estimating membership's total effect on decision-making blocks the very path the programme works through and, if savings is confounded with the outcome, introduces bias. The mediation model does this deliberately and reports the direct effect as one of its outputs; a 'control for savings' in a regression does it by accident and reports the direct effect as the total. The DAG makes the difference visible.
The SEM estimates what the DAG says it can. The DAG is the argument; the SEM is the arithmetic.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Scaling the verbs to the design: a table
DesignPathVerbExample sentence
Randomised XX → latent Yraises, reduces, has an effect of'Membership raises the latent decision-making score by 0.21 SD (95% CI 0.11 to 0.30).'
Randomised X, observed MXMYis consistent with mediation through; the indirect association'The pattern is consistent with mediation through savings; the mediator was not assigned.'
Observational, well-controlled, DAG-justifiedXYis associated with; predicts; under the assumptions in Figure 1, the estimated effect is'Under the adjustment set in Figure 1, savings behaviour is associated with a 0.31 SD higher score.'
Observational, cross-sectional, self-reportAnyis associated with; correlates with'Performance expectancy is associated with adoption intention (β = 0.31).'
Panel with RI-CLPMCross-laggedprecedes; within-person increases in X are followed by'Within-person increases in trust are followed by increases in participation the next year.'
AnyGood fitthe data are consistent with the model; the model was not rejectedNever 'the model is confirmed' or 'proven'
The table is not pedantry. A ministry reading 'social influence significantly impacts adoption' will fund a social-influence campaign. The verb is the policy claim.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Robustness checks that a careful SEM paper includes
CheckWhat it testsReport
Alternative estimator (ML vs WLSMV; PLS vs PLSc)Sensitivity to distributional treatmentPaths and fit under each
Equivalent modelWhether the data can distinguish your story from a rivalName one, show the fit is the same, argue the choice
Item deletion reversedWhether the results depend on the items droppedPaths with the full item set
Method factorCommon method variancePaths with and without
Mediation sensitivityUnmeasured mediator-outcome confoundingρ at which the indirect effect vanishes
Cross-validationModifications capitalising on chanceFit on the holdout half
Subsample stabilityEnumerator, district, language effectsMultigroup or dummies; anything that changes the story
Covariate sensitivityDependence on the adjustment setPaths with the minimal and the full set
Missing-data assumptionMARFIML vs multiple imputation; a pattern-mixture bound
Not every paper needs all nine; every paper needs the ones its design makes relevant, and the appendix is where they live.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
11
Section Eleven
Software, Reporting and Practice
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Software for CB-SEM: what each does and what it costs
ToolCostStrengthsLimits
lavaan (R; Rosseel, J Stat Softw 2012, 48(2))FreeThe reference implementation: CFA, SEM, multigroup, ordinal (WLSMV), FIML, MLR, bootstrap, growth, multilevel; readable syntax; semTools for invariance and reliability, semPlot for diagramsLatent interactions need workarounds; some ordinal-plus-missing combinations are awkward
Mplus (Muthén and Muthén)LicenceThe most complete: latent interactions (XWITH), mixtures, Bayesian, complex survey designs, alignmentCost; syntax; closed
Stata sem / gsemLicenceClean integration with survey data (svy), clustering, weights; gsem for ordinal and multilevel; estat commands for fit, invariance, effectsNo WLSMV; ordinal models via gsem are slower and lack the usual fit indices
IBM SPSS AMOSLicence (SPSS add-on)Drawing interface; bootstrap; the tool most Indian universities teachNo robust ML, no ordinal estimator, no FIML with ordinal; encourages modification-index fishing
JASP and jamovi (SEMLj)Freelavaan with menus; good for learning and for departments without RFewer options exposed
semopy (Python)Freelavaan-style syntax in Python; pipelinesYounger; fewer estimators
OpenMx, blavaan (R)FreeFlexible matrix specification; Bayesian SEMSpecialist
For a South Asian researcher: lavaan, learned through JASP if menus help at first. AMOS's limitations (no robust or ordinal estimation) are real limitations for Likert data, and a thesis committee that requires AMOS is requiring a weaker analysis.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Reporting: what the standards ask for
ElementContentSource
Theory and modelThe diagram, every hypothesised path, the omitted paths, the constructs' definitions and statusKline 2023, ch. 18; APA JARS-Quant SEM table (Appelbaum et al., American Psychologist 2018, 73:3)
Sample and datan, design, missingness, screening, distributions, clustering, weightsJackson, Gillaspy and Purc-Stephenson, Psychological Methods 2009, 14:6
MeasurementItems (in an appendix, in every language), loadings with SEs, reliability (ω), AVE, discriminant validity, invariance where groups are comparedHoyle and Isherwood, Archives of Scientific Psychology 2013, 1:14
EstimationEstimator and why; missing-data method; software and version; identification; convergence; improper solutionsAs above
Fitχ² (df, p, scaling), RMSEA with CI, CFI, TLI, SRMR; residuals; nested comparisons; every modification with reasonSection 06
Structural resultsUnstandardised and standardised paths with SEs and CIs; indirect effects with bootstrap CIs; R²Section 07
Causal statusThe design, the assumptions, the equivalent models, the sensitivity analyses, the verbsSection 10
ReproducibilityData (or synthetic data with the covariance matrix), the syntax, the correlation matrix with SDs in an appendixA covariance matrix and n is enough to refit any CB-SEM
The last row is the one that costs nothing and is almost never done. A correlation matrix with standard deviations and n lets any reader refit the model, and its absence is the reason most published SEMs cannot be checked.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
An SEM paper, section by section
01
INTRODUCTION: the question, the constructs, the model as a figure, the hypotheses as paths
02
METHODS: sample, instrument and translation, measurement model, estimation, analysis plan, registration
03
MEASUREMENT RESULTS: CFA fit, loadings, reliability, validity, invariance
04
STRUCTURAL RESULTS: fit and comparison with rivals, paths, indirect effects, effect sizes
05
ROBUSTNESS: the checks the design made relevant
06
DISCUSSION: what the paths mean, in the verbs the design allows; the equivalent model; limitations with direction of bias; implications
Academic Writing 101 in this series covers the prose. The SEM-specific rule is that the measurement results come before the structural results and are read as a condition on them: a reader who does not trust the constructs has no reason to read the paths.
Two figures: the hypothesised model with paths labelled, and the estimated model with standardised coefficients and their significance. Two tables: the measurement table and the structural table. Everything else goes to the appendix.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Pre-registering an SEM: the template
ItemWhat to write
ConstructsName, definition, reflective or formative, the items (numbered) that measure each
ModelThe diagram; the list of paths with predicted signs; the omitted paths
HypothesesEach as a path or a comparison, with the effect size that would matter
SamplePlanned n and the power analysis or simulation behind it; the sampling design
EstimationEstimator; missing-data treatment; clustering; software
Fit criteriaThe indices and the thresholds that will be used, chosen now
Modification policy'No post hoc modifications' or the specific class allowed (e.g. correlated errors between items with shared stems), with the rule for reporting
InvarianceThe groups to be compared and the level required before comparison
Mediation and moderationWhich effects, which bootstrap, which sensitivity analysis
ExploratoryWhat will be reported as exploratory if run
Registered on OSF as a PDF before the data are cleaned. It turns every decision in sections 04 to 07 from something a referee suspects into something a referee can check, and it is the single step that most separates a credible SEM from the typical one.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
A worked study outline: validating and using an empowerment scale in an evaluation
StepDecisionSection
QuestionDoes an SHG programme raise women's decision-making, and does it work through savings?01
Instrument12 items adapted from a published scale; translated into Hindi and Bangla with cognitive interviews; pilot EFA on n = 180 gives three factors02, 03
DesignVillage-level randomisation; baseline and endline; n = 640 women04, 10
RegistrationDiagram, hypotheses, estimator, fit criteria, invariance plan on OSF before endline11
MeasurementCFA per wave and language; WLSMV; d5 dropped with reason; partial scalar invariance across languages (two thresholds freed) and waves (one)02, 05, 08
StructuralTreatment → savings → decide with covariates; village-clustered; 5,000 bootstrap; CFA-then-structural comparison06, 07
Finding (illustrative)Latent treatment effect on decide 0.21 SD [0.11, 0.30]; indirect through savings 0.12 [0.07, 0.18]; sensitivity ρ = 0.2807, 10
RobustnessSummed-score effect; ML vs WLSMV; method factor; different-source outcome for a subsample; equivalent model stated10
ReportMeasurement table, invariance table, structural table, two figures, correlation matrix and syntax in the appendix11
Numbers illustrative; the structure is the point. Each row names the section that governs it, and the whole is what a competent SEM contribution to an evaluation looks like: the scale validated, the effect estimated on the latent outcome, the mechanism explored with its assumptions stated.
The randomised effect is the headline. The mediation is the second paragraph, in the conditional mood.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
The dozen errors that get SEM papers rejected
ErrorSectionFix in one line
Formative construct modelled as reflective02Decide the direction from theory; PLS or MIMIC for formative
ML on binary or skewed Likert items05WLSMV or MLR
Listwise deletion05FIML or multiple imputation
Fit reached by modification indices06Report the original model; justify each change; cross-validate
χ²/df and GFI as fit evidence06χ² with df and p, RMSEA with CI, CFI, TLI, SRMR
Saturated structural model reported as tested07Omit theoretically absent paths; compare with rivals
Mediation by Baron-Kenny steps, or 'full mediation' claimed07Bootstrap the indirect effect; report intervals
Groups or languages compared without invariance08Configural, metric, scalar; partial if needed
PLS-SEM chosen for small n or non-normality09State a defensible reason or use CB-SEM
Causal verbs from cross-sectional self-report10Associated with; predicts
No common-method check in a single-source survey10Procedural remedy at design; ULMC or marker after
No correlation matrix, syntax or data11Appendix and repository
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Before you submit: the SEM checklist
MeasurementDone
Each construct defined and its reflective or formative status stated
Items listed, with translations, in an appendix
CFA fit reported in full; residuals inspected
Loadings with SEs; ω; AVE; HTMT or equivalent
Estimator matched to the items; missing data handled and reported
Invariance tested for every group and wave comparison
Item deletions and modifications listed with reasons
Structure and claimsDone
Structural model compared with the CFA and with at least one rival
Paths with SEs and CIs, unstandardised and standardised
Indirect effects bootstrapped; sensitivity analysis where the mediator was not assigned
Clustering and weights handled
Common method bias addressed procedurally or statistically
An equivalent model named and argued against
Verbs scaled to the design
Correlation matrix, SDs, n and syntax available
Fifteen lines. Invariance and the equivalent model are the two most often missing.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Where to go next
ResourceWhat it coversNotes
Kline, Principles and Practice of Structural Equation Modeling (5th ed., Guilford, 2023)Everything in sections 01 to 08 and 10, with the argumentsThe first book to buy
Rosseel, lavaan tutorial (lavaan.ugent.be) and J Stat Softw 2012, 48(2)The software, with worked examplesFree
Brown, Confirmatory Factor Analysis for Applied Research (2nd ed., Guilford, 2015)CFA in depth: ordinal items, invariance, higher-order modelsFor sections 02, 03, 08
Bollen, Structural Equations with Latent Variables (Wiley, 1989)The theoryReference
Hayes, Introduction to Mediation, Moderation, and Conditional Process Analysis (3rd ed., Guilford, 2022)Section 07 with observed variablesPROCESS
Enders, Applied Missing Data Analysis (Guilford, 2010; 2nd ed. 2022)Missing dataSection 05
Putnick and Bornstein, Developmental Review 2016, 41:71Measurement invariance: a review and guideSection 08
Hair, Hult, Ringle and Sarstedt, A Primer on PLS-SEM (3rd ed., Sage, 2022); Rönkkö et al. 2016PLS-SEM, from both sidesSection 09
Bollen and Pearl 2013; Antonakis et al. 2010; Imai, Keele and Tingley 2010Causality, claims, and mediation sensitivitySection 10
Podsakoff et al. 2003, 2012Common method biasSection 10
ImpactMojo: Survey Design 101, Item Response Theory 101, Impact Evaluation 101, Econometrics 101The instrument, the items, the designs, the regressionsimpactmojo.in/101-courses/
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
What to remember
  • A latent variable is shared variance among items. Whether it is empowerment is a validity argument the CFA does not make.
  • Decide reflective or formative from theory. It changes everything downstream.
  • Match the estimator to the items: ordinal items, WLSMV; skewed continuous, MLR; missing data, FIML or imputation.
  • Report the χ² and the residuals; read the indices with judgement; modify only with a reason stated in advance.
  • Test the structural model against the CFA and against a rival. A saturated structural model tests nothing.
  • Bootstrap the indirect effect; say what its identification assumes; run the sensitivity analysis.
  • No group, language or wave comparison without invariance testing, and partial invariance reported honestly.
  • PLS-SEM for composites and prediction, with PLSc and PLSpredict; not as an escape from fit.
  • Fit does not choose between equivalent models; design and theory do. Name one.
  • Scale the verbs to the design, address method bias, and put the correlation matrix and syntax where a reader can find them.
The method is sound. Most of its use is not. Be in the first group.
ImpactMojoStructural Equation Modelling 101www.impactmojo.in
Structural Equation Modelling 101 · Complete
Now go measure
something that can fail.
CC BY-NC-ND 4.0·Free Forever·ImpactMojo 101 Series