fullscreen
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
ImpactMojo 101 Series · Free Forever
Systematic
Reviews &
Evidence
Synthesis 101
Finding, Appraising & Combining What Is Already Known — from the Review Question to PRISMA, Meta-Analysis, GRADE and Bibliometrics
Research MethodsSouth Asia Focus100 SlidesFree Access
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
What We Cover
01
Why Synthesis, and What Kind
Slides 3–10
02
The Question and the Protocol
Slides 11–19
03
Searching
Slides 20–28
04
Screening and Selection
Slides 29–36
05
Data Extraction
Slides 37–43
06
Risk of Bias and Study Quality
Slides 44–52
07
Meta-Analysis
Slides 53–63
08
Synthesis Without Meta-Analysis
Slides 64–71
09
GRADE and Reporting
Slides 72–79
10
Bibliometric Analysis
Slides 80–91
11
Practice, Tools and Pitfalls
Slides 92–99
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
01
Section One
Why Synthesis, and What Kind
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
A review is a study of studies
A systematic review answers a defined question by finding every study that bears on it, judging each one by stated criteria, and combining what survives. The unit of observation is the study, not the person. That is the whole difference from an ordinary literature review: a systematic review is itself a piece of research, with a method that another team could repeat and get the same set of studies.
Systematic review
A review that uses explicit, pre-specified and reproducible methods to identify, select, appraise and synthesise all research relevant to a particular question, so that the conclusion depends on the evidence rather than on which papers the author happened to have read.
A literature review
Written from what the author knows and can find. Selection is invisible. Two authors on the same topic produce two different reading lists and two different conclusions, and neither can say why.
A systematic review
Written from a search that is recorded, criteria that are stated in advance, and an appraisal applied to every study the same way. A reader can check the search, re-run it, and see exactly where each study went.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Nobody can read everything
75
randomised trials published per day, health alone
Bastian, Glasziou & Chalmers, PLoS Medicine 2010
11
systematic reviews published per day, same estimate
Bastian, Glasziou & Chalmers, PLoS Medicine 2010
3rd
India's rank among countries by volume of research articles
NSF Science & Engineering Indicators 2022
Those figures are from 2010 and have only grown. A programme officer deciding whether a cash transfer should be conditional, a state official weighing a mid-day meal reform, or a doctoral student framing a thesis cannot read the primary literature on the question. Somebody has to read it for them, in a way that can be trusted. That is what a systematic review is for, and why the funders who commission evaluations (3ie, the World Bank, FCDO, the Gates Foundation) also commission reviews.
The alternative to a systematic review is not "no review". It is an unsystematic one, whose selection of studies nobody can see.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
When the pile of studies said something different
Deworming
Miguel and Kremer's 2004 Kenyan trial (Econometrica 72(1)) reported large effects of school-based deworming on attendance and became the basis of mass programmes. The Cochrane review by Taylor-Robinson and colleagues (2015) pooled the trials and found little or no effect of mass deworming on weight, haemoglobin or attendance. A 2015 reanalysis of the Kenyan data by Aiken, Davey and others opened a dispute that is still cited as the "worm wars". The point is not who was right. It is that one trial and the body of trials disagreed, and only a review could show it.
Microfinance
Through the 2000s microfinance was described as a proven route out of poverty on the strength of case studies and early evaluations. Duvendack and colleagues' 2011 systematic review for the EPPI-Centre and DFID concluded that the evidence base was weak, that the strongest-looking studies had the weakest designs, and that no robust effect on poverty could be shown. Randomised trials published after 2011 (the six studies collected in the American Economic Journal: Applied Economics 2015 symposium) found modest effects at best.
In both cases the field's belief rested on the studies people had read. The review changed the belief by changing the denominator.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Systematic is one kind; choose the right one
TypeQuestion it answersTimeTypical use
Systematic reviewDoes X work, for whom, by how much?9–18 monthsPolicy decisions, guidelines, funding
Meta-analysisWhat is the pooled effect across studies?Within a systematic reviewWhen studies are similar enough to combine
Scoping reviewWhat has been studied, how, and where are the gaps?3–9 monthsNew or diffuse fields; before a full review
Rapid reviewWhat does the evidence say, by next month?4–12 weeksA decision with a deadline
Realist synthesisWhat works for whom, in what circumstances, and why?6–18 monthsComplex programmes with mechanisms
Qualitative evidence synthesisHow do people experience or explain X?6–12 monthsAcceptability, implementation, meaning
Evidence gap mapWhere is the evidence dense and where absent?3–6 monthsSetting a research agenda
Umbrella reviewWhat do the existing reviews say?3–6 monthsMature fields with many reviews
Bibliometric analysisWho publishes what, where, citing whom?2–8 weeksMapping a field, not judging its findings
The durations are the ranges reported in methods literature and reviewers' own accounts; Borah and colleagues (BMJ Open 2017, 7:e012545) measured a median of 67 weeks from PROSPERO registration to publication across 195 reviews.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Which review, decided by the question
  • "Does conditional cash transfer raise secondary enrolment?" An effect question with a comparison: systematic review, with meta-analysis if the trials are similar.
  • "What is known about self-help groups and women's political participation?" A mapping question: scoping review or evidence gap map.
  • "Why did community health worker schemes work in Chhattisgarh and stall elsewhere?" A mechanism-and-context question: realist synthesis.
  • "How do adolescent girls experience menstrual health programmes?" An experience question: qualitative evidence synthesis.
The test
Write the question down. If it contains a verb like works, reduces, increases, it is an effect question and the systematic review machinery applies in full. If it contains what is known or how has this been studied, you are scoping. If it contains why or under what conditions, no pooled number will answer it, and pretending otherwise produces a meta-analysis of things that are not the same.
The most common error in commissioned reviews is asking a scoping question and paying for a systematic-review timeline, or the reverse.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Who does this work, and where it lives
BodyFoundedWhat it holds
Cochrane1993Health reviews; the Handbook is the method reference for every field
Campbell Collaboration2000Education, crime, social welfare and, with 3ie, international development
3ie (International Initiative for Impact Evaluation)2008Development Evidence Portal: impact evaluations, reviews and gap maps
EPPI-Centre, UCL1993Methods and reviews across social policy; the microfinance review above
JBI (Joanna Briggs Institute)1996Scoping-review guidance and appraisal tools
PROSPERO, University of York2011Prospective register of review protocols
For an Indian or South Asian question, start at the 3ie portal and Campbell's international development group, then Cochrane for anything with a health outcome. Most reviews that touch India are not Indian reviews; the search and the framing will be yours to add.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Ten sections, one workflow
01
Question
02
Protocol
03
Search
04
Screen
05
Extract
06
Appraise
07
Synthesise
08
Grade
09
Report
  • Sections 2 to 9 walk the workflow above, in the order a real review runs.
  • Section 10 covers bibliometric analysis, which answers a different question with a different toolkit and is often confused with a review.
  • Section 11 is the practical close: a 12-week plan, free tools, and the errors reviewers make.
Who this is for
Practitioners asked to "pull together the evidence" before a proposal; doctoral students whose thesis needs a literature chapter that would survive examination; evaluators reading someone else's review and needing to know whether to trust it. Nothing here needs a licence: every tool named is free or has a free equivalent.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
02
Section Two
The Question and the Protocol
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
PICO: the question has four parts
ElementAsksExample: cash transfers and schooling
PopulationWhose outcomes?Children aged 6–18 in households below a poverty threshold, low- and middle-income countries
InterventionWhat is being done?Cash transfer to the household, conditional on school attendance
ComparisonCompared with what?No transfer, or an unconditional transfer of the same value
OutcomeMeasured how?Enrolment, attendance, completion, learning (test scores)
Adding S (study design) gives PICOS, which is where you decide whether only randomised trials count or whether matched and difference-in-differences designs are in. That choice shapes everything after it.
Variants exist for other question types. PEO (population, exposure, outcome) suits observational questions; SPIDER (sample, phenomenon of interest, design, evaluation, research type) was written for qualitative synthesis.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Too broad drowns you; too narrow finds nothing
Too broad
"What is the effect of social protection on wellbeing in Asia?" Pulls in pensions, insurance, transfers, public works and school meals, across forty countries and every outcome. Screening runs to tens of thousands of records and the synthesis compares things that share only a label.
Too narrow
"What is the effect of PM-KISAN on fertiliser use in Bihar?" One scheme, one state, one outcome, launched in 2019. There may be two studies. A review of two studies is a reading, not a synthesis, and the time would be better spent on a primary study.
The workable middle is a question where you expect between roughly ten and a few hundred eligible studies. A quick scoping search in one database, with a day's reading of what it returns, tells you which side of that range you are on before the protocol is written. Adjust the population, the time window or the outcome set until the question is answerable.
Write the eligibility criteria as sentences a screener could apply without asking you. "Studies of women's empowerment" is not a criterion. "Studies reporting at least one of: decision-making index, mobility, control over earnings, for women aged 15–49" is.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Inclusion and exclusion, written before searching
DimensionIncludeExcludeWhy it matters
PopulationHouseholds in LMICs; children 6–18High-income countries; adult learnersTransferability of the answer
InterventionCash conditional on attendanceIn-kind transfers, scholarshipsDifferent mechanism, different question
ComparisonNo transfer; unconditional transferBefore–after with no comparison groupCannot separate effect from trend
OutcomesEnrolment, attendance, completion, test scoresSelf-reported "benefit"Comparable measures
DesignRCT, RDD, DiD, matchedCross-sectional correlationsAbility to support a causal claim
Time2000 onwardEarlierProgramme era; data quality
LanguageAny, with translation budgetNone excludedEnglish-only excludes Indian-language evaluations
PublicationPeer-reviewed and greyOpinion pieces, newsGrey literature holds most development evaluations
The criteria are a contract with the reader. Once the search is run, changing them to admit a study you like or drop one you do not is the review equivalent of moving the goalposts, and the protocol exists so that it shows.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
What goes in a protocol
  • Background and rationale: why this review, why now, what exists already
  • The question in PICO(S) form and the eligibility criteria as sentences
  • Information sources: every database, register, website and person you will search
  • The full search strategy for at least one database, with dates
  • Screening procedure: how many screeners, how disagreements are resolved
  • Data extraction items and the form
  • Risk-of-bias tool, named
  • Synthesis plan: meta-analysis or not, model, subgroup analyses named in advance
  • Certainty assessment (GRADE) and reporting standard (PRISMA 2020)
Why pre-specify subgroups
If you decide after seeing the results that the effect is "driven by South Asian studies", a reader cannot tell whether you found a pattern or went looking for one. Naming the subgroups in the protocol (by region, by income level, by design, by conditionality) turns the same analysis from a fishing trip into a test. PRISMA-P (Moher and colleagues, Systematic Reviews 2015, 4:1) is the checklist for protocols, with 17 items.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Register the protocol before you search
PROSPERO
The University of York's Centre for Reviews and Dissemination has run PROSPERO since 2011. It accepts reviews with a health-related outcome, broadly read: nutrition, WASH, maternal health, mental health, violence. Registration is free, public and time-stamped. Reviewers then cite the registration number in the paper.
OSF Registries
For reviews outside health (education, livelihoods, governance), the Open Science Framework accepts a protocol as a registration, free, with the same time stamp. Campbell reviews register with Campbell itself; 3ie registers the reviews it funds. The register matters less than the date.
Registration does two things. It stops duplication: a search of PROSPERO before you start will show whether a team in Dhaka registered your question last year. And it stops drift: the registered version is what your published methods will be compared against, so departures have to be explained rather than hidden.
A review that changes its outcome list after seeing the studies, and does not say so, is the review-level version of selective reporting. Registration is the cheapest protection there is.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Who a review needs
RoleDoesMinimum
Lead reviewerOwns the question, protocol and write-upOne, with time
Second screener and extractorIndependent screening and extraction of every recordOne; dual screening is not optional
Information specialistBuilds and translates the search across databasesA librarian, even for a day
Statistician or methodologistMeta-analysis, heterogeneity, bias assessmentFor any pooled analysis
Subject expertKnows the programmes, the grey literature and the people to emailAdvisory
Language readersScreen and extract studies in Hindi, Bangla, Tamil and so onAs the question demands
The two-person minimum is a method requirement, not a staffing nicety. Every reporting standard asks whether screening and extraction were done in duplicate, and a single-reviewer review is downgraded by readers and by AMSTAR 2, the tool used to appraise reviews. A doctoral student can meet it with a supervisor or a fellow student screening a sample and the disagreement rate reported.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
A review takes longer than you think
67
median weeks from registration to publication, 195 reviews
Borah et al., BMJ Open 2017
1,781
median records screened per review in the same sample
Borah et al., BMJ Open 2017
5
median number of authors
Borah et al., BMJ Open 2017
Most of the time goes into screening and extraction, which scale with the number of records, and into chasing full texts and unreported statistics from authors, which scales with how much of the literature is grey. A development review leans on grey literature and so runs at the slow end.
A rapid review buys time by restricting databases, dates and languages, screening titles once, and extracting a shorter form. Cochrane's rapid review guidance (Garritty and colleagues, Journal of Clinical Epidemiology 2021, 130:13) lists which shortcuts cost least. The restrictions must be reported as limitations.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
A protocol in one slide: self-help groups and empowerment
ItemDecision
QuestionDo economic self-help group programmes improve women's economic, social and political empowerment?
PopulationAdult women in LMICs; group-based savings or credit programmes
ComparisonNo programme, or waitlist
OutcomesEconomic (income, assets, savings), social (mobility, decision-making), political (participation)
DesignsRCT, quasi-experimental with comparison group
SearchNine databases plus 3ie, J-PAL, IPA, NGO sites; no language limit
SynthesisRandom-effects meta-analysis by outcome domain; narrative for the rest
The real one
This is the shape of Brody and colleagues' Campbell review (Campbell Systematic Reviews 2015, 11:19), which found 23 quantitative studies and reported small positive effects on economic, social and political empowerment, with an accompanying qualitative synthesis on the mechanisms. Its protocol was published two years before the review. Read it beside the review to see how the plan and the product relate.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
03
Section Three
Searching
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Databases: what each covers
SourceCoversAccessUse for
PubMed / MEDLINEBiomedicine, public health; MeSH vocabularyFreeAny health outcome
ScopusMultidisciplinary, 1970s onward; strong on Indian journalsSubscriptionBroad coverage, citation data
Web of ScienceMultidisciplinary core journalsSubscriptionCitation chasing, bibliometrics
EconLitEconomics journals and working papersSubscriptionDevelopment economics
ERICEducationFreeSchooling outcomes
OpenAlexEverything with a DOI, plus much withoutFree, APIFree multidisciplinary search; bibliometrics
Google ScholarWidest net, opaque rankingFreeGrey literature, citation chasing; not as a sole source
3ie Development Evidence PortalImpact evaluations and reviews in developmentFreeEvery development review
IDEAS/RePEc, SSRN, NBERWorking papersFreeEconomics before it is published
Shodhganga (INFLIBNET)Indian doctoral theses, full textFreeIndian evidence nobody else indexes
Bramer and colleagues (Systematic Reviews 2017, 6:245) tested database combinations against 58 published reviews and found that no single database recalled all included studies; MEDLINE, Embase, Web of Science and Google Scholar together reached 98 per cent. Plan on at least three, plus the field-specific ones.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Boolean logic, one concept at a time
  • Write one block per PICO concept. Inside a block, join synonyms with OR: ("cash transfer*" OR "conditional cash" OR CCT OR "social pension").
  • Join the blocks with AND: population AND intervention AND outcome. Comparison is rarely searched; it is applied at screening.
  • Truncation (*) catches plurals and variants: school* finds school, schools, schooling.
  • Phrase quotes keep words together: "self-help group" rather than self AND help AND group.
  • Field tags limit where the term must appear: [tiab] in PubMed searches title and abstract; TITLE-ABS-KEY() does the same in Scopus.
Do not search the outcome too tightly
A study of cash transfers that reports enrolment as a secondary outcome may never mention it in the abstract. Searching population AND intervention, and leaving outcome to screening, recalls more at the price of more records. In development, where outcomes are named inconsistently, that trade is usually worth making.
Test the string against five studies you already know should be found. If it misses one, the string is wrong, not the study.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
MeSH and thesaurus terms catch what free text misses
Indexers at MEDLINE tag every record with terms from the Medical Subject Headings (MeSH) tree, so a paper that says "undernourished" in its abstract is still tagged Malnutrition. A search that combines the MeSH term with free-text synonyms recalls both the indexed and the not-yet-indexed. Scopus and EconLit have their own thesauri; ERIC has descriptors. Google Scholar has none, which is one reason it cannot be the only source.
PubMed example
("Malnutrition"[Mesh] OR malnutrition[tiab] OR undernutrition[tiab] OR stunting[tiab] OR wasting[tiab]) AND ("Child, Preschool"[Mesh] OR "under five"[tiab] OR "under-five"[tiab]) AND (India[tiab] OR "India"[Mesh])
Translating between databases
The same concept needs a different string in each database: different field tags, different truncation symbols, different thesauri. The Polyglot Search Translator (Bond University, free) converts a PubMed string into Scopus, Web of Science and others as a first draft, which a person then corrects.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Most development evaluations are not in journals
An evaluation commissioned by a donor is a PDF on the donor's site, or on the implementing NGO's, or nowhere public at all. A review that searches only journals will find the academic subset, which is not a random sample: it overrepresents randomised trials, positive results and English. Grey literature searching is where a development review differs most from a clinical one.
  • 3ie Development Evidence Portal, J-PAL and IPA evaluation databases
  • World Bank Open Knowledge Repository; ADB, UNICEF, UNDP, WFP evaluation libraries
  • FCDO (DevTracker), USAID (DEC), GIZ, Gates Foundation research
  • NITI Aayog, NCAER, state evaluation organisations, IDR's archive
  • OpenGrey, ProQuest theses, Shodhganga
Record it like a database search
Grey searches are the ones reviewers forget to document, and PRISMA-S (Rethlefsen and colleagues, Systematic Reviews 2021, 10:39) asks for each website, the date, the terms used and the number of records. Keep a log as you go; reconstructing it later is guesswork.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Where Indian and regional evidence hides
SourceWhat it holdsNote
ShodhgangaFull-text doctoral theses from Indian universities, via INFLIBNETTheses hold primary data on programmes nobody published on
Economic and Political Weekly archiveFifty-plus years of applied Indian social scienceIndexed unevenly by Scopus; search directly
IndMED / IndMedicaIndian biomedical journals not in MEDLINEPublic-health and nutrition evaluations
National Digital Library of IndiaAggregated theses, reports, booksCoverage varies; useful for older material
NCAER, IGIDR, CDS, ISEC, IEG working papersInstitute seriesSearch each site; few are indexed
Ministry and state evaluation reportsProgramme evaluations by DMEO, NITI Aayog, state bodiesPDFs, often undated; record the retrieval date
icddr,b (Bangladesh), NIPS and PIDE (Pakistan), CBS Nepal, IPS Sri LankaNational institutes' evaluations and surveysRegional coverage that global databases miss
A regional review that skips these sources will conclude, wrongly, that the evidence from South Asia is thin. It is thinly indexed, which is different.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Use it; do not rely on it
What it is good for
Citation chasing: the "cited by" link finds later studies that built on a known one, which databases do badly. Grey literature: it indexes PDFs on institutional sites. Full text: it often links to a free copy. Recall: in Bramer's tests it added studies no database had.
What it cannot do
It shows at most 1,000 results per query and ranks them by an undisclosed relevance formula, so the same search returns different results on different days and to different users. It has no controlled vocabulary, limited Boolean support, and no export of a full result set. A search that cannot be reproduced cannot be the backbone of a systematic review.
The convention is to run the main string in the databases, then run a simplified version in Google Scholar and screen the first 200 to 300 results, reporting that cut-off. Haddaway and colleagues (PLoS ONE 2015, 10:e0138237) found that this captured most of the additional grey literature Scholar contributes.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Hand-searching, citation chasing and asking people
  • Backward citation chasing: read the reference lists of every included study and of prior reviews.
  • Forward citation chasing: find everything that cites an included study (Scopus, Web of Science, OpenAlex, Scholar).
  • Hand-searching: read the tables of contents of the three or four journals that publish most in the field, for the review period.
  • Conference proceedings: the Indian Statistical Institute, NEUDC, PacDev and CSAE programmes hold papers years before publication.
  • Ask: email the twenty people who work on the question. Unpublished and in-progress studies exist only in their inboxes.
Why this is not optional
Horsley, Dingwall and Sampson's Cochrane methodology review (2011) found that checking reference lists identified additional eligible studies in every review that tested it. In development, where working papers circulate for years, citation chasing routinely finds a fifth or more of the final included set. Report the number found by each route.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Zotero, deduplication and the search log
Export every database result set in full (RIS or BibTeX), date-stamped, into a reference manager. Zotero is free and handles tens of thousands of records; EndNote and Mendeley are the paid and semi-free alternatives. Deduplicate before screening: the same paper arrives from three databases with three slightly different titles, and screening it three times wastes the second screener's day.
Deduplication tools: Zotero's built-in merge, the free Deduklick and SRA Deduplicator, or Rayyan's detection on import. Record how many duplicates were removed; PRISMA asks.
Log entryExample
DatabaseScopus
Date2026-09-10
StringTITLE-ABS-KEY(("self-help group*" OR SHG) AND (women OR female) AND (empower* OR "decision-making"))
Limits2000–2026; no language limit
Records1,412
Exported asscopus_shg_20260910.ris
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
04
Section Four
Screening and Selection
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Titles and abstracts first, full texts second
01
Records after dedup
02
Title/abstract screen
03
Full texts sought
04
Full-text screen
05
Included studies
Stage one
Every record is read by two people against the eligibility criteria, quickly, and marked include, exclude or unsure. The rule is when in doubt, keep: a wrongly excluded study is lost for good, a wrongly included one costs a full-text read. At this stage a reviewer manages one to two hundred records an hour.
Stage two
The full text of every surviving record is obtained and read against the criteria in full. Each exclusion is given a reason from a fixed list (wrong population, wrong design, no comparison group, outcome not reported), and the reasons are counted. A study excluded here for "wrong design" should be listed by name in an appendix; readers will ask why a study they know is missing.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Two screeners, and how much they disagree
Independent screening by two people, with disagreements settled by discussion or a third reviewer, is the standard in every guidance document. Single screening misses studies: Waffenschmidt and colleagues' methods study (BMC Medical Research Methodology 2019, 19:132) found single screeners missed a median of 13 per cent of eligible records. Report the agreement, usually as Cohen's kappa, and the number of conflicts.
Cohen's kappa
Agreement between two raters corrected for the agreement expected by chance. Values above 0.6 are conventionally read as substantial; below 0.4 the criteria are probably ambiguous and need rewording before screening continues.
The pilot
Before screening in earnest, both screeners take the same 100 records, compare, and argue about every disagreement. Most disagreements turn out to be criteria that were not as clear as they seemed. Fix the wording, record the change as a protocol amendment, and then start. Kappa on the pilot is the number to report.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Rayyan, Covidence and the spreadsheet
ToolCostDoesLimit
RayyanFree tierBlinded dual screening, conflict resolution, deduplication, keyword highlightingFull-text management is basic
CovidencePaid; free for Cochrane authorsScreening, extraction forms, risk of bias, PRISMA numbersCost for a student team
EPPI-ReviewerPaid; some free accessCoding, screening, machine-learning priority screeningLearning curve
ASReviewFree, open sourceActive-learning screening: ranks records by predicted relevanceA stopping rule must be chosen and reported
Zotero + a spreadsheetFreeEverything, by handBlinding and conflicts are manual
Rayyan (Ouzzani and colleagues, Systematic Reviews 2016, 5:210) is the default for a team without a budget: import the RIS files, invite the second screener, switch on blind mode so neither sees the other's votes, and resolve conflicts at the end. It counts everything PRISMA asks for.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Prioritised screening: what it can and cannot replace
Active-learning tools such as ASReview and the classifier in EPPI-Reviewer learn from your first decisions and re-order the remaining records so the likely includes come first. On large sets they let a team find most eligible studies after screening a fraction of the pile. They do not decide; a person still reads each record that is shown.
Large language models are now used the same way, and the evidence on their recall is mixed and moving. Until a standard exists, treat an LLM as a third screener whose decisions are checked, never as a replacement for the second human.
The stopping rule
The unsolved problem is when to stop. Screening until 50 consecutive records are irrelevant, or until a statistical estimate of remaining includes falls below one, are the usual rules, and each must be stated in the methods. A review that says "we used ASReview" and nothing more has not described its screening.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
PRISMA's flow diagram: every record accounted for
StageCountWhere it goes
Records identified from databases4,212Box 1, by database
Records from other sources318Registers, websites, citation chasing, separately
Duplicates removed1,140Before screening
Records screened (title/abstract)3,390
Records excluded3,102
Reports sought for retrieval288
Reports not retrieved9Named in appendix
Reports assessed for eligibility279
Reports excluded, with reasons241Counts per reason
Studies included38 (in 41 reports)Studies, not papers
Illustrative counts. The 2020 diagram (Page and colleagues, BMJ 2021, 372:n71) separates database records from other-source records and distinguishes studies from reports: one trial may produce three papers and one paper may report two trials. Count studies for the synthesis, reports for retrieval.
The diagram is the first thing a methodologist reads. If the arithmetic does not add up, nothing after it is trusted.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Getting the papers, and what to do when you cannot
  • Institutional library access, then the free routes: Unpaywall (browser extension), Europe PMC, author manuscripts on RePEc, SSRN and institutional repositories.
  • Email the corresponding author. Response rates are low but non-zero, and the email is the same one you will send for missing statistics later.
  • Interlibrary loan through INFLIBNET's N-LIST or a university library, which reaches most Indian institutions.
  • Report the number not retrieved and list them. A reader may have access you lack.
Reports in other languages
A study in Bangla, Tamil or Nepali is eligible if it meets the criteria. Exclusion by language is a decision to be justified, not a default, and reviews that impose it should say what they lost. Machine translation is now adequate for screening; extraction should be checked by a reader of the language.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
What 5,000 records actually means
5,000
records after deduplication
Illustrative
50–80 h
of title/abstract screening, per screener
At 60–100 records an hour
300
full texts to obtain and read, at 30–60 minutes each
Illustrative
So a modest development review costs two people roughly three working weeks of screening before extraction begins. The way to cut it is upstream: a tighter search string, a narrower date window, or exclusion of designs that cannot answer the question. The way not to cut it is to skip the second screener, which trades weeks of labour for a review whose completeness cannot be defended.
Budget the screening in hours before agreeing the timeline. Commissioners who want a review in eight weeks are usually asking for a rapid review, and should be told so.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
05
Section Five
Data Extraction
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Design the extraction form before you open a paper
Extraction is the step where 40 papers become one table. The form decides what that table can answer, so it is written from the protocol, not from the first paper you happen to read. A field that is missing from the form is missing from the review, and adding it after 20 papers means going back through 20 papers.
  • One row per study, one column per item, coded where possible (country as ISO code, design as a fixed list) so the table can be sorted and counted.
  • Free-text fields for anything that will be quoted: the intervention description, the outcome definition, the authors' own caveats.
  • A 'page and table' column for every number, so a second extractor or a reader can find it in seconds.
  • A 'notes and doubts' column. Half the value of extraction is the queries it raises.
Pilot it on five studies
Two extractors take the same five papers, fill the form independently and compare. Every disagreement is either a form problem (the field was ambiguous) or a paper problem (the study reports it two ways). Fix the form, write the decision into a guidance note, then extract the rest. Cochrane's Handbook (chapter 5) asks for exactly this pilot and for the form to be included with the protocol.
Keep the raw extraction and the analysis table separate. The first records what the paper says; the second records what you did with it.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
The items every review needs, and the ones people forget
GroupItemsWhy it matters later
IdentificationStudy ID, all reports of it, year, country, funder, registration numberLinking reports to studies; funding as a bias signal
SettingRural/urban, state or district, baseline poverty or prevalence, delivery agencyApplicability to your context; subgroup analysis
PopulationEligibility, age, sex, numbers randomised and analysed per armAttrition; denominators for effect sizes
InterventionComponents, intensity, duration, who delivered, cost if reported, comparator in the same detailThe comparator explains half of all heterogeneity
DesignRCT/cluster/quasi-experimental method, unit of assignment, clustering handled?Cluster trials analysed as individual data have false precision
OutcomesEvery outcome and timepoint pre-specified in the protocol, with definition, instrument and who measured itSelective reporting; measurement bias
ResultsMeans and SDs, or counts and denominators, or coefficients with SEs, per arm and timepointEffect sizes; never extract only the p-value
AnalysisAdjusted vs unadjusted, covariates, intention-to-treat vs per-protocolWhich estimate to pool
Authors' claimsTheir stated conclusion, verbatimFor the discrepancy check between claim and data
The rows people forget are the comparator, the unit of assignment and the timepoints. A microfinance review that records 'access to credit' as the intervention and nothing about what the control group could borrow is pooling different contrasts under one name.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Getting an effect size out of what was reported
Paper reportsYou needConversion
Mean, SE, n per armSDSD = SE × √n
Mean and 95% CISDSD = √n × (upper − lower) / 3.92
Median and IQRMean and SDApproximate (Wan et al. 2014, BMC Med Res Methodol 14:135); note the approximation
Regression coefficient and SEMean differenceUse directly if unadjusted comparison; adjusted estimates pooled separately
t-statistic and nSMDd = t × √(1/n1 + 1/n2)
Percentages and n per armRisk ratio, odds ratioRebuild the 2×2 table; take logs before pooling
'Significant at 5%' onlyNothing usableEmail the authors; otherwise record as not extractable
The Cochrane Handbook chapter 6 gives every formula, and the Campbell Collaboration's online effect-size calculator (David Wilson's) implements them. Record the route taken for each number in the extraction form. Two reviews that pool the same studies can disagree because one converted from CIs and the other from t-statistics.
Economics papers report coefficients from many specifications. Extract the one the protocol named (usually the authors' preferred, fully-specified estimate) and record the others in notes.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Two extractors, because one makes errors nobody catches
+21.7%
more errors with single extraction than double (relative difference, P = 0.019)
Buscemi et al. 2006, J Clin Epidemiol 59:697
−36.1%
less time for single extraction: the saving that buys those errors
Same trial, reviewers randomised to roles
2
extractors per outcome, the Cochrane minimum for numerical data
Cochrane Handbook, ch. 5
Buscemi and colleagues randomised reviewers to extract or verify, blind to the hypothesis, and counted errors against the source papers. Single extraction produced a fifth more errors, at a third less time. Effect estimates barely moved in that pilot, but the errors were in numbers that feed a meta-analysis, and in a smaller review one wrong SD can move the pooled result. The cost of double extraction is 20–30 hours on a 40-study review.
  • Where the team is small, use one extractor and a second who verifies every number against the paper. That is cheaper than two full extractions and catches most errors.
  • Resolve disagreements by going back to the paper, not by averaging.
  • Report the method: 'Data were extracted by one reviewer and checked by a second' is honest and acceptable in most journals.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Multiple arms, multiple reports, clustered designs
One study, three arms
A trial comparing cash, cash plus training and control gives two comparisons sharing one control group. Pooling both as if independent double-counts the control. Options: combine the two treatment arms into one (Cochrane Handbook, chapter 23), split the control group between them, or pick the arm that matches the review question and record the choice.
One study, three papers
A working paper, a journal article and a follow-up. Link them under one study ID, extract from the most complete report, and note where numbers differ between versions. The published version is not always the most complete: journals cut tables that working papers keep.
Cluster-randomised trials
Most development trials randomise villages or schools. If the paper analysed at the individual level without accounting for clustering, its standard errors are too small. Inflate them by the design effect, 1 + (m − 1)ρ, where m is the average cluster size and ρ the intra-cluster correlation, using an ICC from the paper or from a similar study, and say which.
Every one of these decisions goes in the methods section. A reader who cannot reproduce your effect size from the paper and your stated rules cannot check the review.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
What the paper does not say, and asking for it
Many trials in any health review report no usable standard deviation for at least one outcome, and the share is higher in economics, where results arrive as regression tables. There are three responses, in order of preference.
01
ASK: one email to the corresponding author with a precise request and a deadline
02
DERIVE: from CIs, SEs, t-statistics, or related outcomes in the same paper
03
IMPUTE: borrow an SD from similar studies, and test whether the result changes
  • Ask for the specific number, not 'the data'. 'The standard deviation of household consumption at endline in the treatment and control arms, Table 4' gets answered; 'your dataset' does not.
  • Set a deadline of three weeks and log every request and reply in the review record.
  • If you impute, run the meta-analysis with and without the imputed studies and report both. If the conclusion depends on the imputation, say so in the abstract.
  • Never treat 'not reported' as 'no effect'. A missing outcome is a risk-of-bias signal and is recorded as one.
Selective reporting is common enough that RoB 2 has a whole domain for it. Missing numbers are evidence, not just an inconvenience.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
06
Section Six
Risk of Bias and Study Quality
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
A pooled estimate is only as good as its worst well-weighted study
Risk of bias is the likelihood that a study's design or conduct led it to over- or under-estimate the true effect. It is about internal validity, not about whether the study was well written, large, or published in a good journal. A large trial with broken randomisation is at high risk; a small one with concealed allocation and complete follow-up is at low risk.
Risk of bias
A judgement, per study and per outcome, about whether specific features of design or conduct could have produced a systematic error in the estimated effect. Assessed by domain, never by a summed score.
  • Assess per outcome, not per study. Blinding matters for a self-reported outcome and hardly at all for mortality.
  • Two assessors, independently, with disagreements resolved by discussion. Report the agreement.
  • The judgement feeds three later steps: sensitivity analysis (drop high-risk studies), GRADE (downgrade for risk of bias), and the discussion.
  • Do not sum points. Scales such as Jadad weight items arbitrarily; Cochrane advises against any scale that produces a total score (Handbook, chapter 7).
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Cochrane RoB 2 for randomised trials: five domains
DomainQuestion it asksA development example
1. Randomisation processWas the sequence random and allocation concealed? Do baselines suggest a problem?Lottery in public with sealed lists: low. NGO staff choosing 'eligible' villages after the list: high
2. Deviations from intended interventionsDid participants or staff know the assignment, and did that change what they did? Was the analysis appropriate?Control villages receiving a similar scheme from another donor mid-trial
3. Missing outcome dataHow much attrition, and was it related to the outcome?30% attrition in a migration-prone district, higher in the control arm
4. Measurement of the outcomeWas the assessor blind? Could knowing the assignment change the measurement?Enumerators from the implementing NGO measuring self-reported income
5. Selection of the reported resultWas there a pre-analysis plan, and does the paper report what it planned?Twelve outcomes measured, three reported, no registration
Sterne and colleagues, BMJ 2019, 366:l4898. Each domain gets one of three judgements: low risk, some concerns, high risk, reached through signalling questions with a written algorithm. The overall judgement is the worst domain. Assess per outcome. The tool and its guidance are free at riskofbias.info.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
ROBINS-I for non-randomised studies: the confounding domain comes first
DomainJudgement hinges on
ConfoundingWere the important confounders (listed in advance by the review team) measured and controlled? Time-varying confounding?
Selection of participantsDid entry into the study depend on characteristics observed after the intervention started?
Classification of interventionsWas exposure status defined clearly and recorded without knowledge of outcome?
Deviations from intended interventionsCo-interventions, contamination, switching
Missing dataAttrition, and whether it differs by arm and outcome
Measurement of outcomesAssessor knowledge of exposure; comparable methods across groups
Selection of the reported resultMultiple analyses, subgroups, outcomes, with no plan
Sterne and colleagues, BMJ 2016, 355:i4919. Judgements run low, moderate, serious, critical, or no information. 'Low' means comparable to a well-conducted trial, which a matching study rarely reaches; 'critical' means the study is too problematic to inform the synthesis and is dropped from it.
The confounding domain requires the review team to write down, in the protocol, which confounders a credible study must handle. For a cash-transfer review that is baseline income, household size, and programme targeting rules.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Appraising DiD, matching, IV and RDD studies
Development evidence is mostly quasi-experimental, and generic tools written for cohort studies miss the questions an economist would ask. 3ie and the Campbell Collaboration's International Development Coordinating Group developed a risk-of-bias tool for these designs (Waddington and colleagues, J Clin Epidemiol 2017, 89:43, paper 6 of the quasi-experimental series), organised around the identifying assumption of each method.
  • Difference-in-differences: is there evidence for parallel pre-trends, and are there enough pre-periods to see them?
  • Matching: was there common support, and were the matched covariates measured before treatment?
  • Instrumental variables: is the exclusion restriction argued rather than asserted, and is the first stage strong (F above 10, or the newer Lee et al. 2022 thresholds)?
  • Regression discontinuity: is the running variable manipulable? Is there a density test at the cut-off?
Two further checks for any design
Was the estimate pre-registered, or does the paper read like the specification was found after the fact? And is the reported outcome the one the programme aimed at, or a proxy chosen because it moved? Brodeur, Cook and Heyes (American Economic Review 2020, 110:3634) find bunching of test statistics just above conventional thresholds in IV and DiD papers, and much less in RCTs and RDD.
Record the identifying assumption and the evidence for it in the extraction form. It is the single item that most determines the weight a study should get.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Tools for the other designs you will meet
DesignToolNotes
Cohort, case-controlNewcastle-Ottawa Scale; ROBINS-INOS is widely used and has poor inter-rater reliability (Hartling et al. 2013, J Clin Epidemiol 66:982); prefer ROBINS-I where the question is causal
Cross-sectional prevalenceJBI critical appraisal checklist; Hoy et al. 2012 toolSampling frame, response rate and case definition are the items that matter
Diagnostic accuracyQUADAS-2Four domains: patient selection, index test, reference standard, flow and timing
QualitativeCASP qualitative checklist; JBI qualitative toolAppraises credibility and reflexivity, not bias; feeds GRADE-CERQual
Mixed methodsMMAT (Hong et al. 2018)One tool across designs, five items each, no summed score
Economic evaluationsCHEERS 2022 (reporting); Drummond checklistCosting perspective and discount rate are the usual gaps
Modelling and simulationNo standard toolJudge against ISPOR good-practice guidance; say that no validated tool exists
Whichever tool you use, present the results per study in a table or a traffic-light figure (the robvis R package and web app draw these from a spreadsheet), and use them. An appraisal that is reported and then ignored in the synthesis is decoration.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
The studies you did not find are not a random sample
Studies with significant, positive results are more likely to be written up, submitted, accepted and cited. A review that pools only what is published therefore starts biased upward. Franco, Malhotra and Simonovits (Science 2014, 345:1502) traced 221 social-science experiments run through one funded programme (TESS), so the file drawer could be seen: about two-thirds of the null results were never written up at all, and strong results were 60 percentage points more likely to be written up and 40 points more likely to be published.
10 of 48
null-result studies published
Franco et al. 2014, TESS studies
56 of 91
strong-result studies published
Same sample
  • Prevention beats detection: search registries, working-paper series, evaluation repositories and theses so unpublished work is in the review.
  • Detection: a funnel plot of effect against precision, and Egger's regression test for asymmetry (Egger et al. 1997, BMJ 315:629). Neither is meaningful with fewer than ten studies (Cochrane Handbook, chapter 13).
  • Asymmetry has other causes, including true heterogeneity where small studies target populations with larger effects. Say which explanation you favour and why.
  • In economics, the FAT-PET framework (Stanley and Doucouliagos 2012) regresses effect on standard error; the intercept is the bias-corrected estimate. Andrews and Kasy (AER 2019, 109:2766) offer a selection-model correction.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Reading a funnel plot, and what it cannot tell you
Each study is a point: effect size on the horizontal axis, standard error (inverted, so precise studies sit at the top) on the vertical. Without bias the points form a symmetrical funnel around the pooled estimate, because small studies scatter more. With publication bias the bottom-left corner, small studies with small or negative effects, is empty.
  • Trim-and-fill (Duval and Tweedie 2000, Biometrics 56:455) imputes the missing mirror-image studies and re-pools. Treat it as a sensitivity analysis, not a corrected estimate.
  • Contour-enhanced funnel plots shade the regions of significance, which helps separate publication bias from other asymmetry.
  • Do not draw one for a review of eight studies. It will look asymmetric or symmetric by chance, and readers will over-read it.
Where the asymmetry comes from in development
Small pilot studies run by the implementing organisation in a favourable site; larger, independent replications at scale with smaller effects. Vivalt (Journal of the European Economic Association 2020, 18:3045) finds across 20 intervention types that government-implemented programmes report smaller effects than NGO- or researcher-implemented ones, and that effect sizes are hard to predict from one setting to the next. That is heterogeneity, not only bias, and the review should model it rather than 'correct' it away.
Report the plot, the test, and your reading of it. The reader is entitled to disagree.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Three places the risk-of-bias judgement must show up
01
TABLE: per-study, per-domain judgements with the reason for each, in the report or appendix
02
SYNTHESIS: sensitivity analysis restricted to low-risk studies; subgroup by risk if enough studies
03
CERTAINTY: GRADE downgrade for risk of bias when high-risk studies drive the estimate
04
DISCUSSION: what the bias would do to the direction of the result, stated plainly
The most common failure in published development reviews is a careful appraisal table followed by a synthesis that treats every study alike. If the pooled estimate falls by half when high-risk studies are removed, that is the finding, and it belongs in the abstract.
Write the sentence now: 'Restricting to studies at low risk of bias (k = 6) gave a pooled effect of X, compared with Y for all studies.' If you cannot write it, the appraisal was not used.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
07
Section Seven
Meta-Analysis
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
A weighted average, where precision sets the weight
Meta-analysis
The statistical combination of effect estimates from two or more studies into one pooled estimate with a confidence interval, each study weighted by the inverse of its variance, so that precise studies count for more.
The arithmetic is short. What takes judgement is deciding whether the studies are similar enough to combine, which effect measure to combine, whether to assume one true effect or a distribution of them, and how to describe the disagreement between studies. A pooled number with no account of those choices is a number nobody should use.
  • Pool only when the protocol said you would, and only studies that ask the same question of comparable populations with a comparable comparator.
  • Two studies can be meta-analysed. The result will be fragile and the heterogeneity estimate near useless, and both should be said.
  • The pooled estimate is not the truth. It is the best summary of these studies, subject to their biases and to what was not found.
  • Never pool across effect measures. Convert everything to one metric first, and record the conversions.
Apples and oranges can be combined if the question is about fruit. The protocol defines the fruit.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Choosing what to pool
Outcome typeMeasureWhen to use itWatch for
Continuous, same scaleMean difference (MD)All studies report the outcome in the same units: rupees per month, cm, test score on one instrumentUnits and scaling (monthly vs annual) must match exactly
Continuous, different scalesStandardised mean difference (Hedges' g)Studies use different tests or indices for the same constructSensitive to the SD used; heterogeneous populations widen SDs and shrink g
BinaryRisk ratio (RR)Most interpretable for events such as enrolment, immunisation, defaultCannot be used when the control event rate is zero
BinaryOdds ratio (OR)Case-control designs; logistic regressionsOverstates RR when events are common; often misread as RR
BinaryRisk difference (RD)Absolute effect for policy: percentage points of coverage gainedVaries with baseline rate, so usually more heterogeneous than RR
RateRate ratio, hazard ratioEvents per person-time; survivalHazard ratios need the estimate and CI from the paper; rarely derivable
Regression coefficientPartial correlation, elasticity, or MD from coefficientEconomics literatures where every study is a regressionComparability across specifications; see Stanley and Doucouliagos 2012
Ratios are pooled on the log scale and back-transformed. Hedges' g corrects Cohen's d for small samples by J = 1 − 3 / (4df − 1); at 20 participants per arm the correction is about 2%. Kraft (Educational Researcher 2020, 49:241) argues that in education 0.05 SD is small, 0.05–0.20 medium and above 0.20 large, against Cohen's 0.2/0.5/0.8, which came from psychology and are far too demanding for field programmes.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Three studies, fixed-effect: the arithmetic in full
StudyEffect (SMD)SEWeight w = 1/SE²Sharew × effect
A0.200.1010019.0%20.00
B0.350.20254.8%8.75
C0.050.0540076.2%20.00
Sum525100%48.75
Pooled effect = 48.75 / 525 = 0.093. SE of the pooled effect = √(1/525) = 0.044. 95% CI = 0.093 ± 1.96 × 0.044 = 0.007 to 0.178. Illustrative numbers.
What the table shows
Study C, the most precise, carries three-quarters of the weight and pulls the pooled estimate toward its own 0.05. Study B, with the largest effect, barely registers. Precision, not effect size and not sample size directly, is what decides influence. A study's weight in a forest plot is drawn as the size of its square for this reason.
This is the fixed-effect model: it assumes A, B and C estimate the same true effect and differ only by sampling error. The next slide tests that assumption.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Q, I² and τ²: measuring how much the studies disagree
StatisticFormulaOur exampleReads as
Cochran's QΣ wii − θ̂)²3.54 on 2 df (p = 0.17)A test with low power when k is small; do not rely on its p-value
(Q − df) / Q43%Share of observed variation beyond chance; a proportion, not an amount
τ²(Q − df) / C, C = Σw − Σw²/Σw0.0077 (τ = 0.088)Between-study variance in effect-size units; the amount
Prediction intervalθ̂ ± tk−2 √(τ² + SE²)Wide, with k = 3Where a new study's effect would likely fall
Higgins and Thompson (Statistics in Medicine 2002, 21:1539) introduced I²; the Cochrane Handbook reads 0–40% as possibly unimportant, 30–60% moderate, 50–90% substantial, 75–100% considerable, with the overlaps deliberate. I² rises with study precision even when τ² does not, so large trials produce high I² for small disagreements.
Report all three, and the prediction interval. IntHout and colleagues (BMJ Open 2016, 6:e010247) found that in a fifth of Cochrane reviews with a significant pooled effect the prediction interval included harm.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
The same three studies under a random-effects model
StudyEffectSE² + τ²Weight w*ShareFixed share
A0.200.0100 + 0.007756.632.2%19.0%
B0.350.0400 + 0.007721.011.9%4.8%
C0.050.0025 + 0.007798.255.9%76.2%
Pooled0.134SE 0.075175.8FE: 0.093
95% CI: −0.014 to 0.282, against 0.007 to 0.178 under fixed effect. Adding τ² to every study's variance flattens the weights: the small study gains, the large one loses, and the interval widens to reflect the disagreement. The conclusion changed from 'significant' to 'not'. Illustrative numbers.
Which model, and which estimator
Fixed effect answers 'what is the one common effect?'; random effects answers 'what is the average of a distribution of effects?'. Development programmes differ by site, implementer and comparator, so random effects is the default and the protocol should say so before the data are seen. DerSimonian and Laird (1986) is the classic τ² estimator and is still the default in RevMan; REML is better with few studies, and the Hartung-Knapp adjustment to the interval is recommended when k is small (IntHout, Ioannidis and Borm, BMC Med Res Methodol 2014, 14:25).
Random effects is not a fix for heterogeneity. It describes it. Explaining it is the next slide.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Reading a forest plot
  • One row per study: the point estimate as a square whose area is the weight, the confidence interval as a horizontal line.
  • The vertical line of no effect at 0 (differences) or 1 (ratios). A study's line crossing it means that study alone cannot rule out no effect.
  • The diamond at the bottom is the pooled estimate; its width is the confidence interval. Some plots add a bar for the prediction interval, which is the honest one.
  • Below the diamond: k, Q, I², τ², and the test for overall effect.
  • Order the rows by something meaningful (year, risk of bias, effect size), never alphabetically. Sorted by effect, a plot shows heterogeneity at a glance.
Five questions to ask of any forest plot
Do the intervals overlap, or are there two clusters? Is one study carrying most of the weight, and is it at low risk of bias? Do the small studies sit systematically to one side? Does the diamond's interval exclude no effect, and does the prediction interval? Would a policymaker reading only the diamond be misled by what the rows show?
A diamond that excludes zero above rows that mostly cross it is a common and legitimate result. A diamond that excludes zero above two clusters of rows on opposite sides of it is a result that should not have been pooled.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Subgroups and meta-regression
When studies disagree, the interesting question is why. Subgroup analysis splits the studies by a characteristic named in the protocol (implementer, region, intensity, risk of bias) and tests whether the pooled effects differ. Meta-regression does the same with a continuous moderator, or several at once, by regressing effect size on study characteristics with weights.
  • Pre-specify the subgroups and keep them few. Ten post-hoc subgroups will produce one 'significant' difference by chance.
  • Cochrane's rule of thumb: at least ten studies per moderator in a meta-regression. With 15 studies, one moderator.
  • Test the difference between subgroups, not whether each subgroup's effect is significant on its own.
  • Study-level moderators are ecological. 'Effects are larger in studies with more women' is not 'effects are larger for women'.
Two economics examples
Card, Kluve and Weber (Journal of the European Economic Association 2018, 16:894) meta-analyse over 200 active labour market evaluations and find that programme type and time horizon explain much of the variation: training shows small short-run effects and larger ones after two years. Meager (AEJ: Applied 2019, 11:57) uses a Bayesian hierarchical model on seven microcredit RCTs and finds that most of the sites' effects are consistent with a common small average, with the heterogeneity concentrated in the upper tail of business outcomes.
Heterogeneity is often the finding. Reporting that a programme works in NGO pilots and not at government scale is more useful than one averaged number.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Many effect sizes from one study
An economics paper reports six outcomes at two follow-ups from one sample. Treating the twelve estimates as twelve independent studies gives that paper twelve votes and shrinks the confidence interval falsely. Ignoring all but one throws away evidence. Neither is acceptable in a review that will be read by methodologists.
  • Pick one effect per study per outcome domain by a rule written in the protocol (the primary outcome, the longest follow-up), and report the rule.
  • Or average the effects within a study, with a variance that accounts for their correlation (Borenstein et al. 2009, chapter 24).
  • Or use robust variance estimation, which allows every effect in and corrects the standard errors for clustering within studies (Hedges, Tipton and Johnson, Research Synthesis Methods 2010, 1:39). Implemented in robumeta and clubSandwich in R.
The economics convention
Meta-regression studies in economics routinely take every estimate from every paper, hundreds or thousands of them, and cluster standard errors by study. That is a form of robust variance estimation and is fine when the point is to model what drives estimates. It is not fine when the point is one pooled effect and half the estimates come from three papers. Say which purpose the analysis serves.
Whatever the choice, state the number of studies and the number of effect sizes separately, everywhere a k appears.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Sensitivity analyses: showing the result is not an artefact of one decision
DecisionAlternative to testReport as
Random vs fixed effectThe other modelBoth estimates, one sentence on the difference
Inclusion of high-risk studiesRestrict to low riskPooled effect with and without; the difference is a finding
Imputed SDs or ICCsExclude imputed studies; halve and double the ICCRange of pooled estimates
One influential studyLeave-one-out: re-run k times dropping each studyPlot of k estimates; name any study whose removal changes the conclusion
Effect measureRR instead of OR; MD instead of SMD where possibleDirection and significance under each
OutliersExclude studies whose CI does not overlap the pooled CIWith and without, plus a reason the outlier differs
Publication biasTrim-and-fill; PET-PEESE; selection modelAdjusted estimate labelled as a sensitivity result
Pre-specify the list and run all of it, including the ones that turn out inconvenient. A sensitivity analysis that is only reported when it supports the main result is the selective reporting the review exists to detect in others.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Tools for meta-analysis, free ones first
ToolCostBest forLimits
R: metaforFreeEverything: all models, meta-regression, RVE, plots; the reference implementation (Viechtbauer, J Stat Softw 2010, 36(3))Code, not menus; a learning curve of a few days
R: metaFreeQuick pooled analyses and publication-quality forest plots with one function callFewer model options than metafor
RevMan WebFreeCochrane-format reviews, risk-of-bias tables, summary-of-findings tablesLimited models; no meta-regression
Jamovi with the MAJOR moduleFreePoint-and-click meta-analysis for a first course; metafor underneathFewer diagnostics; export is limited
JASPFreeClassical and Bayesian meta-analysis with menusNewer, fewer worked examples online
Stata meta suiteLicenceEconomists already in Stata; meta-regression, funnel tests, forest plotsCost; RVE needs the user-written robumeta
Comprehensive Meta-AnalysisLicenceEffect-size conversion from almost any reported statisticCost; closed
For a South Asian team the choice is R or Jamovi. Both run on an old laptop offline, both are free, and a Jamovi analysis can be reproduced in R when a reviewer asks. Keep the extraction sheet and the script in a public repository so the review is reproducible by anyone.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
08
Section Eight
Synthesis Without Meta-Analysis
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Most development reviews cannot meta-analyse everything, and should not try
Outcomes measured on incompatible scales, designs too different to weight against each other, statistics that cannot be turned into effect sizes, or three studies where one is a different programme under the same name: any of these means the evidence has to be synthesised in words and structured tables. That is a method, with its own reporting standard, not a fallback for reviews that failed.
SWiM
Synthesis Without Meta-analysis: a nine-item reporting guideline (Campbell and colleagues, BMJ 2020, 368:l6890) covering how studies were grouped, the standardised metric used, the synthesis method, how heterogeneity was investigated, how certainty was assessed, and how the data were presented.
  • Group studies by a logic stated in advance: intervention type, outcome domain, population. Tables per group.
  • Where effect sizes exist but cannot be pooled, still standardise them and tabulate with confidence intervals, so a reader can see the direction and size.
  • Where they do not exist, record direction of effect and whether the study's own test was significant, separately.
  • Never write 'most studies found a positive effect' without saying how many, out of how many, and at what risk of bias.
The failure mode is the narrative review in disguise: a paragraph per study, then a conclusion the paragraphs do not support.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Counting directions, not p-values
Counting how many studies were 'significant' is the wrong count. A programme with a true small effect tested in ten underpowered studies will produce two significant results and eight 'no effect' findings, and the count will say it does not work. Counting the direction of effect, regardless of significance, is a legitimate method (Cochrane Handbook chapter 12) and comes with a sign test: under no effect, directions split 50:50.
9 of 10
studies with a positive direction: sign test p = 0.02
Binomial, two-sided
6 of 10
positive: p = 0.75, consistent with no effect
Same test
Ways to show it
Harvest plot (Ogilvie and colleagues, BMC Med Res Methodol 2008, 8:8): one bar per study, placed by direction, bar height by quality, shading by design. Effect direction plot: a table of arrows per outcome per study, sized by sample. Albatross plot: p-values against sample size, with contours for effect sizes, useful when only p and n are reported. All three are drawable in R or by hand in a spreadsheet.
Direction counting tells you whether there is an effect. It says nothing about how big. Say that in the results.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Synthesising qualitative evidence
MethodWhat it doesSourceUse when
Thematic synthesisCodes findings line by line, builds descriptive themes, then analytical themes that go beyond the studiesThomas and Harden 2008, BMC Med Res Methodol 8:45The question is about experience, acceptability, barriers
Framework synthesisStarts from an a priori framework (a theory of change, a policy's own logic) and codes into it, adding themes that do not fitCarroll et al. 2011; Booth and Carroll 2015A commissioner has a framework and wants evidence mapped to it
Meta-ethnographyTranslates concepts across studies (reciprocal, refutational), then a line of argumentNoblit and Hare 1988; eMERGe reporting guidance 2019Interpretive depth on a small set of rich studies
Realist synthesisAsks what works for whom in what circumstances; builds context-mechanism-outcome configurationsPawson et al. 2005; RAMESES standards, Wong et al. 2013, BMC Medicine 11:21Complex interventions whose effect depends on how people respond
Meta-aggregationPools findings into categories and synthesised statements with recommendationsJBIPractice guidance in health
Report to ENTREQ (Tong and colleagues, BMC Med Res Methodol 2012, 12:181) and assess confidence in each finding with GRADE-CERQual (Lewin and colleagues, PLoS Medicine 2015, 12:e1001895), which asks about methodological limitations, coherence, adequacy of data and relevance. A finding supported by two thin studies from one district is 'low confidence' however vivid the quotes.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Context, mechanism, outcome: the realist question
A cash transfer conditional on school attendance raises enrolment in one state and not in another. A conventional review averages the two. A realist review asks what the money did in each place: in one, it covered the cost of a uniform the school required; in the other, the school was 8 km away and no transfer changes that. The mechanism (relieving a cash constraint) fires only in a context where cash was the constraint.
01
Elicit candidate programme theories from documents and stakeholders
02
Search purposively, including grey literature and process evaluations, for evidence on each theory
03
Extract context-mechanism-outcome configurations, not effect sizes
04
Refine the theory; report what works for whom, where, and why
What a realist review is not
It is not an excuse to skip systematic searching or appraisal; RAMESES (Wong and colleagues 2013) sets 19 reporting items. It does not produce an effect size and should not be commissioned by someone who wants one. It is also easy to do badly: 'mechanism' gets used for any intermediate step, and the configurations become a list of things that happened. A mechanism is a change in the reasoning or resources of the people the programme reaches.
Realist and statistical synthesis are complementary. The meta-analysis says the average is 0.1 SD with high heterogeneity; the realist review says where the 0.3 came from.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Combining quantitative and qualitative streams
  • Segregated: run a quantitative synthesis and a qualitative one separately, then bring them together in a matrix. Simplest and most common.
  • Sequential: the qualitative synthesis generates hypotheses (which components matter, which barriers) that the quantitative synthesis then tests as subgroups. The EPPI-Centre's approach.
  • Integrated: convert findings to one form (usually qualitative) and synthesise together. Rare, and hard to do transparently.
  • Whatever the design, keep the two appraisals separate: risk of bias for effect studies, CASP or similar for qualitative ones.
The matrix
Rows are the intervention components or barriers found in the qualitative stream; columns are the trials. A cell records whether the trial's intervention had that component and what effect it found. Reading across, you see whether trials that addressed the barriers people named did better. It is descriptive, it is honest about being descriptive, and it is often the most useful table in the review.
Mixed-methods reviews take longer than either stream alone and need two skill sets on the team. Budget both.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Scoping reviews and evidence gap maps
A scoping review asks what evidence exists, of what kind, on what, and where the gaps are. It does not appraise risk of bias and does not synthesise effects. Arksey and O'Malley (International Journal of Social Research Methodology 2005, 8:19) set the framework; PRISMA-ScR (Tricco and colleagues, Annals of Internal Medicine 2018, 169:467) is the reporting standard. Use one before a systematic review, to decide whether one is feasible, or instead of one when the question is about the shape of a literature.
  • A scoping review is not a quick systematic review. The search is as rigorous; what is dropped is appraisal and synthesis.
  • Charting replaces extraction: a table of study characteristics, not results.
  • The output is a map and a gap list, and both should be specific enough to commission from.
Evidence gap maps
3ie's format (Snilstveit and colleagues, J Clin Epidemiol 2016, 79:120): a grid with interventions as rows and outcomes as columns, each cell holding the studies and reviews found, coloured by confidence. Blank cells are the map's point. 3ie's maps on social protection, WASH and agriculture are open at developmentevidence.3ieimpact.org and are the first place to look before proposing a review on any of those topics.
A gap map takes weeks, not months, and a funder can read it in ten minutes. It is the most cost-effective synthesis product for a South Asian research team to offer.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
When six reviews of the same question disagree
Evans and Popova (World Bank Research Observer 2016, 31:242) compared six systematic reviews of what improves learning in developing countries. The reviews reached different conclusions, and the reason was not the analysis: of 227 studies across the six, only three appeared in all six reviews. Different inclusion rules, different classifications of interventions and different dates produced different evidence bases under one question.
227
distinct studies across six reviews of learning outcomes
Evans and Popova 2016
3
studies included in all six
Same paper
  • Before starting, read the existing reviews on your question and tabulate their inclusion criteria against yours. If yours will differ, say why and what that changes.
  • Classify interventions by a published taxonomy where one exists, so your categories can be compared with others'.
  • An overview of reviews (an umbrella review) is a recognised design for exactly this situation: appraise the reviews with AMSTAR 2 (Shea et al. 2017, BMJ 358:j4008) and explain the divergence.
  • Evans and Popova found that once categories were harmonised, the reviews agreed on more than they appeared to: pedagogy-focused interventions and individualised instruction came out consistently.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
09
Section Nine
GRADE and Reporting
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
GRADE: how sure are we, outcome by outcome
A pooled estimate is a number; GRADE is the judgement about how much to trust it. The Grading of Recommendations Assessment, Development and Evaluation approach (Guyatt and colleagues, BMJ 2008, 336:924) rates the certainty of evidence for each outcome as high, moderate, low or very low. It is now the standard in Cochrane, Campbell, WHO guidelines and 3ie reviews, and journals increasingly expect it.
Certainty of evidence
The extent to which we are confident that the estimate of effect is close to the true effect, for a specific outcome. It is a property of the body of evidence, not of any one study, and it can differ between outcomes in the same review.
  • Randomised trials start at high; observational and quasi-experimental studies start at low.
  • Downgrade one or two levels for each of five reasons: risk of bias, inconsistency, indirectness, imprecision, publication bias.
  • Upgrade observational evidence for a large effect, a dose-response gradient, or when plausible confounding would reduce the observed effect.
  • The result is a sentence a minister can read: 'Cash transfers probably increase school enrolment (moderate certainty).'
Two assessors, a written reason for every downgrade, and the same rules applied to outcomes you like and outcomes you do not.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Deciding when to downgrade
DomainDowngrade whenA development example
Risk of biasMost of the weight comes from studies at high risk or with some concerns, and a sensitivity analysis moves the estimateThree of five trials unregistered with outcomes chosen after the fact
InconsistencyEffects point in different directions, intervals barely overlap, I² is high and subgroups do not explain itEnrolment effects of 0.02 and 0.25 across states with no moderator found
IndirectnessThe studies' population, intervention, comparator or outcome differ from the review questionQuestion is about Bangladesh; evidence is from Mexico and Brazil; outcome is attendance, not learning
ImprecisionThe confidence interval includes both a worthwhile benefit and no effect (or harm); total sample below the optimal information sizePooled RR 1.15 (0.92 to 1.44) from 600 households
Publication biasSmall-study effects, or many registered trials with no results, or a field where nulls are known not to be publishedFunnel asymmetry across 14 microenterprise studies; 6 registered trials unreported
Each domain can cost one level (serious) or two (very serious). Randomised evidence with two serious problems is 'low'; observational evidence with one is 'very low'. Balshem and colleagues (J Clin Epidemiol 2011, 64:401) give the definitions; GRADEpro GDT, free for non-commercial use, walks through the judgements and produces the table.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
The summary-of-findings table: one page that carries the review
OutcomeStudies (participants)Relative effect (95% CI)Absolute effectCertaintyComment
School enrolment, 1 year7 RCTs (41,200 children)RR 1.08 (1.04 to 1.12)6 more per 100 enrolled (3 to 9 more), from 75 per 100ModerateDowngraded for inconsistency
Test scores, 2 years4 RCTs (18,500)SMD 0.04 (−0.03 to 0.11)Roughly 1 percentile pointLowDowngraded for imprecision and indirectness
Child labour3 RCTs, 2 quasi (9,800)RR 0.91 (0.80 to 1.03)4 fewer per 100 (9 fewer to 1 more)LowDowngraded for risk of bias, imprecision
Household consumption9 studies (52,000)MD +7% (4 to 10)About Rs 420 per month at the baseline meanHighConsistent, precise, direct
Illustrative table for a cash-transfer review. Absolute effects are computed at a stated baseline risk, because 'RR 1.08' means something different at 20% enrolment and at 95%.
Why it comes first
Most readers of a review read the abstract and this table. It should hold the seven or so outcomes that matter to a decision, stated in absolute terms, with the certainty rating beside each. Cochrane puts it before the introduction. Campbell and 3ie reviews carry it in the plain-language summary.
If your review has no summary-of-findings table, a policymaker has to build one from your results section. They will not.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Reporting to PRISMA 2020: the 27 items
SectionItemsWhat reviewers check
Title, abstract1–2'Systematic review' in the title; the 12-item abstract checklist
Introduction3–4Rationale and explicit objectives
Methods5–15Eligibility, sources, full search strategy, selection and extraction processes, risk of bias, effect measures, synthesis methods, certainty
Results16–22Flow diagram, study characteristics, risk of bias per study, individual and synthesised results, certainty
Discussion23Interpretation, limitations of evidence and of the review, implications
Other24–27Registration and protocol, support, competing interests, availability of data and code
Page and colleagues, BMJ 2021, 372:n71, with the explanation and elaboration paper at 372:n160. The checklist is a table of where each item appears in your manuscript; most journals require it as a supplement. Filling it in honestly is the quickest audit of a draft: an item with no page number is a gap.
  • Item 7 wants the full search strategy for every database, verbatim, as run. An appendix, not a paraphrase.
  • Item 27 asks where the extraction sheet and analysis code live. A public repository answers it.
  • Extensions: PRISMA-S for searches, PRISMA-ScR for scoping reviews, PRISMA-P for protocols; MOOSE (Stroup et al. 2000, JAMA 283:2008) for observational syntheses.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
The discussion: what the evidence supports, and for whom
01
MAIN FINDINGS: each primary outcome, effect and certainty, in one paragraph
02
EVIDENCE LIMITS: bias, inconsistency, gaps, what was not studied
03
REVIEW LIMITS: what your search, criteria and resources could not do
04
APPLICABILITY: to your setting, with the reasons an effect might transfer or not
05
IMPLICATIONS: for policy, hedged to the certainty; for research, specific enough to fund
Keep the two kinds of limitation apart. 'The trials were short' is a limit of the evidence; 'we searched only English-language databases' is a limit of the review. Conflating them lets the review's shortcuts hide behind the literature's.
Applicability to South Asia
Most systematic reviews in development pool evidence from three continents. A review written for a Bihar or Sindh audience should say which studies came from comparable settings, whether the effect differed there, and what in the implementation context (front-line worker density, banking access, school distance) would change the mechanism. Vivalt's 2020 finding that effects are hard to predict across sites is the reason this paragraph exists.
'More research is needed' is not an implication. 'A trial of X in a government delivery system with a two-year follow-up on learning outcomes' is.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Where reviews are published, and what each outlet expects
OutletScopeExpectations
Campbell Systematic ReviewsSocial welfare, education, crime, international development, methods; open access, no feeRegistered title and protocol first; full Campbell and GRADE standards; editorial and methods review before acceptance
Cochrane Database of Systematic ReviewsHealth, including nutrition, WASH, maternal and child healthCochrane review group registration; RevMan format; MECIR conduct and reporting standards
3ie systematic review seriesDevelopment effectiveness; commissioned and open access3ie protocol approval; quasi-experimental risk-of-bias tool; evidence gap map alongside
Journal of Development EffectivenessImpact evaluation and synthesis in developmentPRISMA; interest in methods and policy relevance
World Development, Journal of Development EconomicsGeneral development; reviews accepted selectivelyContribution beyond summary; meta-regression rather than description
Systematic Reviews (BMC), Research Synthesis MethodsProtocols, methods, reviews across fieldsOpen access with article fee; waivers for low- and middle-income authors
Indian outlets: EPW, Indian Journal of Medical Research, IJCMPolicy audiences and health; systematic reviews acceptedShorter formats; PRISMA still expected; check the journal's indexing before submission
Publish the protocol first, wherever the review is going. A registered protocol is what separates a systematic review from a literature review with a flow diagram.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Reviews go stale: updating and living reviews
Shojania and colleagues (Annals of Internal Medicine 2007, 147:224) estimated that the median systematic review needs updating within about five and a half years, and a quarter within two. Development literatures move faster where a topic is fashionable with funders. A review should state its search date on the first page and say when it plans to be updated.
  • Keep the search strategy, extraction sheet and code so an update re-runs the search from the last date rather than starting again.
  • An update is a new review with the old studies pre-extracted: same protocol, same appraisal, new records screened.
  • Say what changed. 'Two new trials, both at low risk, moved the estimate from 0.08 to 0.11' is the finding of an update.
Living systematic reviews
A living review (Elliott and colleagues, J Clin Epidemiol 2017, 91:23) re-runs the search monthly or quarterly and incorporates new studies as they appear, with the current version always published. It needs saved searches with alerts, a screening team on retainer and a publication format that can be versioned. Cochrane ran several during COVID-19; in development, 3ie's gap maps are updated on a schedule and are the closest equivalent.
Do not promise a living review without a budget line for it. An unfunded living review is a stale one with a misleading label.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
10
Section Ten
Bibliometric Analysis
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Bibliometrics asks what a field has done, not what the evidence says
Bibliometric analysis
The quantitative study of publications and their citation and authorship links: how much a field publishes, who publishes it, which works and journals it builds on, how its topics cluster and move over time.
A systematic review reads 40 studies closely. A bibliometric analysis counts 4,000 and reads none of them; it maps structure, not findings. The two answer different questions and are often confused in South Asian journals, where 'bibliometric review' can mean a citation count dressed as a literature review. Donthu and colleagues (Journal of Business Research 2021, 133:285) set out when each is appropriate.
  • Use it when the literature is too large to read, when the question is about the field itself (who, where, what topics, what trends), or to scope a systematic review's boundaries.
  • Do not use it to claim a programme works. Citation counts measure attention, and attention is not evidence.
  • Its inputs are database exports; its outputs are tables of counts and network maps. Everything depends on the database's coverage, which is the first limitation to state.
'The most cited paper on X' is a fact about citation, and a useful one. It is not a fact about X.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
The two families: performance analysis and science mapping
TechniqueUnitQuestionOutput
Publication and citation countsAuthors, institutions, countries, journals, yearsWho produces the field, and where is it published?Ranked tables; annual output curve
Citation analysisDocumentsWhich works are the field's foundations?Most-cited list, citation half-life
Co-citationPairs of documents cited togetherWhat is the intellectual structure? Documents co-cited often belong to one schoolClusters of foundational works (Small 1973)
Bibliographic couplingPairs of documents sharing referencesWhat is the current research front? Papers citing the same sources work on the same problemClusters of recent papers (Kessler 1963)
Co-authorshipAuthors, institutions, countriesWho collaborates with whom; where are the isolated groups?Collaboration network
Co-word (keyword co-occurrence)Keywords or title termsWhat are the topics and how do they connect and shift over time?Thematic map; overlay by year
The first two are performance analysis: counting. The last four are science mapping: networks. A competent bibliometric paper does at least one of each, states the database and date, and interprets the clusters by reading the papers at their centres. A map with unlabelled clusters is a picture.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Where the records come from, and who each source leaves out
SourceAccessCoverageExport
Scopus (Elsevier)Subscription; many Indian universities via consortiaOver 27,000 active titles; curated; weaker on regional and non-English journalsCSV, RIS, BibTeX with references; 2,000 records at a time
Web of Science (Clarivate)SubscriptionNarrower and older; the source of the Journal Impact FactorTab-delimited with cited references; 500 or 1,000 at a time
OpenAlexFree, open API and web interface; launched 2022 as the successor to Microsoft Academic GraphBroadest: over 250 million works, but with noisier metadataCSV, JSON via API; no limit in practice
DimensionsFree basic search; paid for fullWide, includes grants and policy documentsLimited in the free tier
Lens.orgFree for individualsScholarly works plus patents; good for applied fieldsCSV, RIS
Google ScholarFreeWidest, including theses and reports; no quality controlNo export; Publish or Perish scrapes it in small batches
Shodhganga (INFLIBNET)FreeIndian theses in full textNo structured export; useful as a source, not a dataset
Mongeon and Paul-Hus (Scientometrics 2016, 106:213) show that Scopus and Web of Science under-represent the social sciences, humanities and non-English publishing. A bibliometric study of development research in South Asia run on Web of Science alone will systematically miss Indian, Bangladeshi and Nepali journals, and should say so.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
From query to map: the steps, and where time goes
01
DEFINE: field boundaries as a search string, with the same rigour as a review search
02
EXPORT: full records with references and keywords from one or more databases
03
CLEAN: merge duplicates, disambiguate author names, harmonise keywords and institutions
04
ANALYSE: counts, trends, top lists; then networks
05
MAP: clusters, overlays by year, density
06
INTERPRET: read the central papers of each cluster and name what it is about
Cleaning takes half the time. 'A. Banerjee', 'Banerjee, Abhijit' and 'Banerjee AV' are one author; 'Jawaharlal Nehru University' appears under a dozen spellings; 'cash transfer', 'cash transfers' and 'CCT' are one keyword. Every count and every map depends on those merges, and none of the tools does them well automatically. Keep a thesaurus file and cite it.
Report the query, the database, the date, the record count at each cleaning step, and the thesaurus. A bibliometric analysis that cannot be re-run is an opinion with a network diagram.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Software: VOSviewer, Bibliometrix, and the rest
ToolCostStrengthLimit
VOSviewerFree (Leiden); van Eck and Waltman, Scientometrics 2010, 84:523Network maps: co-authorship, co-citation, coupling, co-word; overlay and density views; reads Scopus, WoS, OpenAlex, LensMaps only; counts and trends need another tool; thesaurus by text file
Bibliometrix and biblioshiny (R)Free; Aria and Cuccurullo, Journal of Informetrics 2017, 11:959Full workflow in one package: import, descriptives, laws, networks, thematic maps, a point-and-click interfaceR installation; large networks are slow in the browser interface
CiteSpaceFree for academic use; Chen, JASIST 2006, 57:359Bursts and turning points over time; timeline viewsJava, dated interface; WoS-centred
Publish or Perish (Harzing)FreeAuthor and journal metrics from Google Scholar, Scopus, OpenAlex, CrossrefSmall batches; no networks
GephiFreeAny network, with full layout and statistics controlYou build the network file yourself
Python: pyalex, pybibxFreeScripted pulls from OpenAlex; reproducible pipelinesCode required
For a first project: export from Scopus or OpenAlex, run descriptives in biblioshiny, draw the maps in VOSviewer. Both run offline on a modest laptop. Save the VOSviewer map and network files alongside the export so a reader can reopen exactly the figure in the paper.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Reading a VOSviewer map without over-reading it
  • Node size is the item's weight: number of documents, or citations, as you chose. Distance approximates relatedness; two nodes close together are linked often. Colour is the cluster the algorithm assigned.
  • Overlay view colours nodes by average publication year, which shows where the field has moved: blue for older topics, yellow for recent.
  • Density view shows where the mass of the field sits; sparse regions are either gaps or artefacts of the query.
  • The resolution parameter sets how many clusters appear. Change it and the picture changes. Report the value.
What the picture does not show
Clusters are statistical, not conceptual, until you name them by reading their central papers. Distances in a two-dimensional layout are approximations of a high-dimensional network and can mislead at the edges. A node's size records citations, so an older paper is always bigger than a better recent one. And nothing on the map tells you whether any of the papers is sound.
Label the clusters in the caption with two or three of their central works, and say how you decided the label. That is the analysis; the map is the evidence for it.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Citation indicators and their misuse
IndicatorDefinitionUseMisuse
h-indexh papers with at least h citations each (Hirsch, PNAS 2005, 102:16569)A rough summary of an author's output and uptakeComparing across fields or career stages; it only rises, and it rewards volume
Journal Impact FactorCitations in year t to items from t−1 and t−2, divided by citable itemsComparing journals within one fieldJudging an article or an author by its journal; skewed by a few highly cited papers
CiteScore (Scopus)Four-year window, all document typesAs above, wider coverageAs above
Field-weighted citation impactCitations relative to the world average for the same field, year and document typeCross-field comparison at institutional scaleSmall numbers; a single paper's FWCI is noise
AltmetricsMentions in policy documents, news, social mediaTracing policy uptake of development researchTreating attention as quality
The Leiden Manifesto (Hicks and colleagues, Nature 2015, 520:429) gives ten principles, the first of which is that quantitative evaluation should support, not replace, expert judgement, and DORA (2012) asks institutions to stop using journal metrics to assess individuals. Both matter in India, where promotion rules have at times scored publications by journal lists and impact factors, and where that scoring fed the market for predatory journals.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
The regularities every bibliometric paper reports
  • Lotka's law: the number of authors producing n papers falls roughly as 1/n². Most authors in any field publish once; a few publish most of it.
  • Bradford's law: a core of a few journals holds a third of the papers on a topic, a second larger zone holds another third, and a long tail holds the rest.
  • Price's law and exponential growth: literatures tend to double over fixed periods until they saturate. The annual output curve is the first figure in most papers.
  • Citation ageing: the half-life of citations, which varies by field; economics cites older work than computer science.
What they are for
Bradford's core journals tell a systematic reviewer where to hand-search. Lotka's distribution tells a funder that a field with 400 authors has perhaps 20 sustained researchers. The growth curve tells everyone whether a topic is emerging, mature or declining. Reporting the laws without saying what follows from them is a common way to fill a bibliometric paper with numbers that change nothing.
Biblioshiny computes all of them in one click. The judgement is in the sentence after the number.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
A bibliometric study of cash-transfer research, step by step
StepDecisionIllustrative result
QueryTITLE-ABS-KEY("cash transfer*" AND (poverty OR welfare OR "social protection")), 2000–2025, articles and reviews3,140 Scopus records
CleanMerge 212 author variants, 96 institution variants, keyword thesaurus of 140 lines3,088 records after de-duplication
OutputAnnual curveUnder 20 a year to 2005; about 300 a year by 2022
ProducersCountries by corresponding authorUSA, UK, then Brazil, Mexico, South Africa; India seventh, Bangladesh and Pakistan in the top twenty
JournalsBradford coreWorld Development, Journal of Development Effectiveness, Social Science & Medicine, Journal of Development Economics
Co-citationDocuments, minimum 20 citationsClusters around Progresa evaluations, unconditional-transfer trials, and health and nutrition outcomes
Co-word overlayAuthor keywords, minimum 10 occurrencesRecent yellow: 'COVID-19', 'digital payments', 'universal basic income'; older blue: 'conditionality', 'Progresa'
Counts here are illustrative, constructed for teaching. The point is the shape of the writing: each step has a decision, each decision has a number, and the interpretation sits beside both. Note what the study cannot say: whether cash transfers work, whether the India-based papers are good, or anything about the hundreds of evaluations published as reports outside Scopus.
Combine it with a gap map or a systematic review and the two become useful to each other: the map shows the field's shape, the review shows what its centre found.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Bibliometrics of and for South Asian research
India is among the largest producers of science and engineering articles by count (NSF Science & Engineering Indicators 2022, as section 01 noted). A count says nothing about where the work is published or whether anyone reads it, so a bibliometric study of the region has to report uptake alongside output and say which database each number came from.
  • A bibliometric study of the region should use OpenAlex or Dimensions alongside Scopus, and say what the difference in coverage was.
  • Distinguish output from uptake: papers, and then citations per paper normalised by field.
  • Co-authorship maps show a pattern worth reporting: South Asian authors linked to North American and European institutions far more than to each other.
Predatory publishing
Journal lists used in promotion rules created a market for journals that accept anything for a fee, and bibliometric databases index some of them. Before counting a journal, check it against DOAJ, the Scopus source list, COPE membership and the Think Check Submit criteria; Beall's list, though closed in 2017, remains archived and Cabells maintains a paid successor. A bibliometric analysis that counts predatory output as research output reports the problem as achievement.
The region's evaluation reports, working papers and theses are mostly outside every index. A field map built on Scopus alone is a map of the part that faces outward.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Structure of a bibliometric paper that reviewers accept
01
QUESTION: a question about the field, stated as one that counts and maps can answer
02
DATA: source, query, date, cleaning, counts at each step
03
PERFORMANCE: output, producers, outlets, with the laws where they add something
04
STRUCTURE: at least one network, with clusters named from reading
05
SO WHAT: what the field is missing, and what a researcher or funder should do
Reviewers at Scientometrics, Journal of Informetrics and the applied journals reject the same paper repeatedly: a Scopus export, the default biblioshiny figures, one paragraph per figure describing what it shows, and a conclusion that the field is growing. The difference is the question and the reading.
Pair it with something
The most useful bibliometric work in development is a component of something else: the scoping stage of a systematic review, the baseline of a research-capacity evaluation, or the map behind a funding call. Donthu and colleagues (2021) give the standalone template; the pairing is what makes the numbers matter to anyone outside informetrics.
Do not call it a systematic review. Do not report PRISMA for it. Do not draw conclusions about what works from it. Reviewers check all three.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
11
Section Eleven
Practice, Tools and Pitfalls
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
A complete toolkit at no cost, running offline
StageToolNotes
ProtocolPROSPERO or OSF Registries; PRISMA-P checklistOSF accepts any discipline; PROSPERO is health-focused and slow to register
SearchPubMed, Google Scholar, OpenAlex, RePEc/IDEAS, 3ie repository, Campbell library; Scopus where the institution has itSave every strategy as run, with date and count
Reference managementZotero, with the Better BibTeX pluginFree, open; de-duplication and full-text retrieval built in
ScreeningRayyan (free tier); or a shared Zotero library with tagsRayyan blinds screeners to each other and logs conflicts
ExtractionGoogle Sheets or LibreOffice Calc from a piloted template; SRDR+ (AHRQ, free)One row per study; a 'source page' column for every number
Risk of biasRoB 2 and ROBINS-I Excel tools from riskofbias.info; robvis for figuresTwo assessors; keep the signalling-question answers
AnalysisR with metafor and meta; Jamovi with MAJOR for menusScript and data in a public repository
CertaintyGRADEpro GDT (free for non-commercial use)Produces the summary-of-findings table
BibliometricsOpenAlex export, biblioshiny, VOSviewerAll offline after export
ReportingPRISMA 2020 checklist and flow diagram generator (Haddaway et al. 2022, R package and web app)Fill the checklist against your own draft before submission
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Who does what, and how long it takes
PhaseWeeks (typical)PeopleOutput
Question, scoping searches, protocol4–8Lead, methodologist, subject expert, librarianRegistered protocol
Search and de-duplication2–3Librarian or trained searcherSearch log; record set
Title and abstract screening3–6Two screenersCalibrated screening; conflict log
Full-text screening3–5Two screenersExcluded-with-reasons list
Extraction and appraisal6–10Two extractors, two assessorsExtraction sheet; risk-of-bias table
Synthesis and GRADE4–8Statistician or methodologist, leadAnalyses; summary-of-findings table
Writing, PRISMA, submission4–6Lead, all authorsManuscript with checklist and appendices
Total26–46Four to six people, part timeConsistent with Borah et al.'s median of 67 weeks elapsed
The roles that get skipped are the librarian and the second screener, and those are the two that cannot be recovered later. A subject expert who does not do methods work is still essential at three points: the question, the interpretation of heterogeneity, and the applicability paragraph.
A rapid review compresses this to 8–12 weeks by single screening, limiting databases and dates, and skipping meta-regression, and says so in the title and the limitations (Garritty et al. 2021).
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Using AI tools in a review without disqualifying it
Language-model tools now offer to write search strings, screen abstracts, extract data and summarise papers. Some of that is useful. None of it removes the requirement that a named person made each decision and that the method is reported in enough detail to be reproduced. Treat an AI tool as a third screener whose accuracy you have measured, never as the second screener.
  • Search: useful for drafting synonyms and translating a strategy between database syntaxes. Check every line; models invent field tags.
  • Screening: run the model on a sample the humans have already screened and report its sensitivity. Below about 95% sensitivity it is not fit to exclude anything.
  • Extraction: acceptable as a first pass verified against the paper, number by number; it fabricates numbers with confidence.
  • Writing: never for the results. A model summarising your own extraction table will smooth over the disagreements that are the point.
Discovery tools
Elicit, Consensus, Semantic Scholar and similar search a corpus built largely from open metadata, not from Scopus or Web of Science, and rank by relevance models you cannot audit. They are good for scoping a question and finding a seed set for citation chasing, and they are not a database search. A review whose only search was an AI tool fails PRISMA item 7 on its face.
Report it: the tool, version, date, prompt, what it was used for, and how its output was checked. Most publishers' policies now require that disclosure in the methods.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
The twelve ways development reviews go wrong
PitfallWhat it looks likePrevention
No protocolCriteria change as papers arriveRegister before searching
One database'We searched Google Scholar'Three or more, plus grey literature, plus citation chasing
Single screener'Studies were selected by the first author'Two, with a conflict log
Pooling different comparatorsCash vs nothing pooled with cash vs in-kindComparator as an eligibility criterion and a subgroup
Appraisal ignoredA traffic-light figure and then equal weightsSensitivity analysis by risk of bias, in the abstract
Counting p-values'Seven of ten studies found significant effects'Direction counts or effect sizes
I² as a verdict'Heterogeneity was high (I² = 82%), so results should be interpreted with caution'τ², a prediction interval, and pre-specified moderators
Twelve estimates from one paperk = 60 from 14 studies with no adjustmentOne per study, or RVE
Funnel plot with six studies'No evidence of publication bias'Do not test below ten; search the file drawer instead
No certainty ratingResults reported as if all equally reliableGRADE per outcome
Applicability unaddressedEvidence from Latin America applied to Nepal without commentA named paragraph on transfer
Bibliometrics as evidence'The most-cited interventions are…'Keep the two products apart
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
How to read someone else's review in twenty minutes
01
REGISTRATION: is there a protocol, and does the paper match it?
02
SEARCH: which databases, what date, is the strategy in an appendix?
03
FLOW: do the numbers add up, and are exclusions listed with reasons?
04
APPRAISAL: which tool, by how many people, and was it used in the synthesis?
05
SYNTHESIS: pooled or narrative, which model, heterogeneity reported how?
06
CERTAINTY: GRADE, and does the abstract's language match it?
AMSTAR 2 (Shea and colleagues 2017) formalises this with 16 items, seven of them critical; a review failing more than one critical item is rated critically low. Most reviews published in general development journals fail at least one, usually the protocol or the appraisal.
Reading for a decision
A programme officer does not need to reproduce the review. They need to know three things: the effect on the outcomes they care about in absolute terms, how certain it is, and whether the evidence came from settings like theirs. The summary-of-findings table and the applicability paragraph answer all three. If the review does not have them, its conclusions are the authors' opinion of their own work.
The abstract of a review with low-certainty evidence should say 'may'. 'Cash transfers improve learning' above a low rating is a mismatch a reader should catch.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Where to go next
ResourceWhat it isAccess
Cochrane Handbook for Systematic Reviews of Interventions, version 6.4 (2023)The reference for every step, from question to GRADE; chapters 5, 6, 8, 10 and 13 are the coreFree online at training.cochrane.org/handbook
Gough, Oliver and Thomas, An Introduction to Systematic Reviews, 2nd ed. (Sage, 2017)The social-science and mixed-methods treatment, from the EPPI-CentreBook
Petticrew and Roberts, Systematic Reviews in the Social Sciences (Blackwell, 2006)Still the clearest account of why the method transfers beyond medicineBook
Borenstein, Hedges, Higgins and Rothstein, Introduction to Meta-Analysis, 2nd ed. (Wiley, 2021)Meta-analysis from first principles with worked arithmeticBook
Stanley and Doucouliagos, Meta-Regression Analysis in Economics and Business (Routledge, 2012)The economics approach: FAT-PET, publication bias, meta-regressionBook
Campbell Collaboration and 3ie methods guidesStandards and templates for development reviews; the quasi-experimental risk-of-bias toolFree at campbellcollaboration.org and 3ieimpact.org
Donthu et al. 2021, Journal of Business Research 133:285Guidelines for a bibliometric analysisJournal article
Cochrane Interactive Learning; Campbell's online courseStructured courses with exercisesCochrane's is paid, with free access in some low- and middle-income countries; Campbell's is free
ImpactMojoSurvey Design 101, Impact Evaluation 101 and Research Methods 101 in this series cover the primary studies these reviews synthesiseimpactmojo.in/101-courses/
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
What to remember
  • A systematic review is a study, with a protocol, a sample and a method, and it is judged as one.
  • The search decides the evidence base. Three databases, grey literature, citation chasing, and a strategy saved verbatim.
  • Two people at every judgement: screening, extraction, appraisal, GRADE.
  • Risk of bias is assessed per outcome and used in the synthesis, or it was decoration.
  • Pool when the protocol said to and the studies allow it; report Q, I², τ² and a prediction interval, and explain heterogeneity rather than lament it.
  • Where you cannot pool, synthesise to SWiM: directions, standardised tables, structured argument.
  • Rate certainty with GRADE and let the abstract's verbs match it.
  • Report to PRISMA 2020 and put the search, the extraction sheet and the code where anyone can find them.
  • Bibliometrics maps the field; it does not weigh the evidence. Keep the two products apart and pair them when it helps.
  • Write the applicability paragraph. A review for South Asia that does not say what transfers has not finished.
Every one of these is a decision you can defend in writing. That is the whole method.
ImpactMojoSystematic Reviews & Evidence Synthesis 101www.impactmojo.in
Systematic Reviews & Evidence Synthesis 101 · Complete
Now go find out
what is already known.
CC BY-NC-ND 4.0·Free Forever·ImpactMojo 101 Series