Paeds · professional-practice-and-evidence
Evidence-based medicine and critical appraisal
Also known as Evidence-based medicine in paediatrics · Critical appraisal of the medical literature · Hierarchy of evidence and levels of evidence · Risk of bias and validity appraisal · GRADE certainty of evidence
Fellowship guide to evidence-based medicine and critical appraisal in paediatrics: the five steps of evidence-based practice, PICO question formulation, hierarchy of evidence, validity and risk-of-bias appraisal (RoB 2, ROBINS-I, QUADAS-2), results appraisal (effect size, confidence intervals, number needed to treat), GRADE certainty of evidence, publication bias, and applicability to the child — with ANZ, UK, US and Canada guidance.
On this page & tools
Your progress
Saved locally on this device.
Practise this topic
Target exams
Red flags
Life stages
Care settings
Clinical exam formats
Board mappings
Appraise any paper in the right order
Is it valid — check design and risk of bias first · Results — magnitude and precision of the effect · Applicability — does it fit my patient · Versus the rest — how does it sit with the totality of evidence. Remember it as I-R-A-V: valid, results, applicable, versus. Never read the magnitude of an effect before you have judged whether the estimate can be trusted. [13] [14]
Overview & Definition
A parent hands you a printout claiming a new therapy halves the risk of their child's chronic illness, and asks whether to start it. Before you answer the clinical question, you face an evidence question: is the claim trustworthy, how large is the effect in absolute terms, and does the study describe a child anything like yours. Evidence-based medicine is the discipline that answers these. [1] [13]
Evidence-based medicine is the conscientious, explicit, and judicious integration of the best current research evidence with clinical expertise and patient or family values. It is not cookbook medicine, nor the worship of randomised trials for their own sake; the best evidence only ever informs a decision that still rests on your judgement and the family's circumstances. [1]
The work moves through five steps, often called the five A's: ask a focused question, acquire the best evidence, appraise it critically, apply it to the patient, and assess the outcome. This page owns the appraisal engine that drives them. The statistical machinery — sensitivity, specificity, likelihood ratios, confidence intervals, and effect measures — belongs to the dedicated leaves on diagnostic accuracy, clinical epidemiology, and biostatistics, and you should cross-link them rather than rebuild their content here. [1] [14]
Classification
Sort the evidence you find by the strength of the design that produced it, then match each design to the appraisal checklist it demands. [1] [13]
Levels of evidence. At the base sit expert opinion, bench research, and mechanistic reasoning, which generate hypotheses but rarely settle them. Above them sit case reports and case series, then case-control and cohort studies that observe without intervention, then randomised controlled trials that test causation directly, and at the apex sit systematic reviews and meta-analyses that pool trials appraised for bias. The pyramid ranks the resistance of each design to confounding, not its infallibility. [1] [14]
Filtered versus unfiltered evidence. Pre-appraised and filtered resources — systematic reviews, evidence summaries, and guidelines — do the appraisal work for you and should be searched first. Primary studies and raw databases are unfiltered, so you must appraise them yourself before acting. Reaching for the filtered layer first saves time and reduces error. [14]
Matching design to appraisal tool. Each design carries its own failure modes, so each needs its own tool: RoB 2 for randomised trials, ROBINS-I for non-randomised studies of interventions, QUADAS-2 for diagnostic accuracy studies, and PRISMA for the reporting of systematic reviews. Choosing the wrong tool misses the bias that matters. [7] [8] [15]

Epidemiology & Risk Factors
Children have historically been therapeutic orphans. Far more paediatric treatments are used off-label or extrapolated from adult trials than in any other age group, and the evidence base for children sits, on average, lower down the hierarchy and in smaller studies. [16] [1]
Publication and reporting bias warp the literature you can see. Studies with positive findings reach print more often than negative ones, so the published record overstates true effects, and a substantial fraction of even published conclusions are later overturned or weakened by larger and better evidence. [16] [12]
Clinician factors threaten the application of evidence at the bedside. Time pressure, knowledge gaps in appraisal skills, the cognitive pull of a recent or vivid case, and uncritical acceptance of an authoritative summary all push decisions away from the best evidence toward habit or convenience. Naming these pressures is the first defence against them. [1] [13]
The common paediatric settings that demand appraisal are familiar: choosing a therapy, weighing whether a screening test does more good than harm, deciding whether to adopt a new guideline, and counselling a family through an uncertain prognosis. Each turns on whether the evidence can be trusted and whether it applies to the child. [1] [14]
Pathophysiology
Think about the three forces that drag an estimate away from the truth, and how rigorous design resists each. [1] [13]
Bias is systematic error built into the conduct of a study. Selection bias tilts who ends up in each group, performance bias tilts the care they receive, detection bias tilts how outcomes are measured, attrition bias tilts who is lost, and reporting bias tilts which results surface. Each pushes the estimate in one direction, and a large biased study is just a precisely wrong answer. [7] [8]
Confounding mimics a treatment effect when a third variable is linked to both the exposure and the outcome. Randomisation, restriction, matching, and statistical adjustment each try to break that link, and observational studies live or die by how completely they do so, which is why ROBINS-I scrutinises confounding so closely. [8]
Chance produces spurious findings in any single study. A p-value below the conventional threshold does not rule it out, and the more comparisons a study makes, the more likely a false positive becomes. A confidence interval conveys precision more honestly than a single cut-off, because it shows the range of values compatible with the data. [13] [16]
Publication and reporting bias skew the published record toward positive results, inflating the apparent effect in any meta-analysis that pools only what reached print. Funnel-plot asymmetry and trial-registration checks are the guards against this distortion. [12] [16]

Clinical Presentation
You will meet evidence questions in several recognisable shapes. [1] [13]
The new-treatment question. A family asks whether a therapy they read about online really works and is safe for their child. The job is to find the best evidence, judge its validity, report the absolute benefit and harm, and check it applies before counselling. [1] [14]
The guideline-versus-patient mismatch. A guideline recommends an intervention, but the child differs from the trial population in age, severity, or comorbidity. You must weigh directness and decide whether to follow, adapt, or depart from the recommendation with reasons. [10] [11]
The journal-club or viva task. A supervisor hands you a recent randomised trial or meta-analysis and asks you to summarise and appraise it. Move in the fixed order — question, validity, results, applicability — and defend each judgement. [13] [14]
The promoted screening or diagnostic test. A test is offered as accurate and worthwhile, and you must judge whether its sensitivity, specificity, and likelihood ratios, weighed against its harms and the population, justify its use. [3] [15]
The conflicting-studies problem. Two studies report opposite results, and you must weigh their validity and precision to decide which to trust, and how the totality of the evidence falls. [14] [12]
Differential Diagnosis
Name the true problem before you act on a number. [13] [16]
| You see | Prefer this framing | Trap |
|---|---|---|
| A large and impressive effect | Check validity first; bias and confounding inflate estimates | Assume large means real |
| A p-value below 0.05 | Ask the magnitude and clinical importance of the effect | Treat significance as clinical importance |
| A narrow confidence interval | A large precise study — but is it unbiased and applicable? | Forget validity because it looks precise |
| A result from unlike patients | Question directness and applicability to your child | Extrapolate without checking fit |
| A glowing narrative review | Seek an independent systematic review with appraisal | Trust the summary over the synthesis |
Real effect versus artefact. A large effect may be genuine, but in observational paediatric data it is more often the fingerprint of bias or confounding, so confirm it with the risk-of-bias tool before you believe it. [8] [16]
Significance versus importance. A statistically significant result can be clinically trivial, especially in a large trial where tiny differences clear the p-value threshold. Always translate the finding into the absolute effect the family will feel. [13] [1]
Clinical & Bedside Assessment
Frame the question before you search. State it in PICO terms — population, intervention, comparator, and outcome — because a fuzzy question returns a fuzzy search and a fuzzy appraisal. The focused question is the hinge of the whole process. [2]
Judge design and validity before results. Identify the study design and score its risk of bias with the matched tool before you read a single effect size. A biased estimate is wrong no matter how large or precise, so validity always comes first. [7] [8]
Assess the magnitude and precision of the effect. Extract the absolute risk reduction and the number needed to treat alongside any relative figure, and read the confidence interval to see the range compatible with the data. An interval that crosses the null is not a positive result. [13] [1]
Assess applicability to the child. Compare the trial's population, intervention, comparator, outcomes, and setting against your patient, your resources, and the family's values. Direct evidence is always stronger than extrapolation. [14] [10]
Investigations
The investigations here are searches and extractions, not blood tests. [1] [14]
Acquire the best evidence efficiently. Search the filtered layer first — systematic reviews and evidence summaries — and only descend to primary studies when you must. A focused, reproducible search strategy prevents both drowning in results and missing the key paper. [14]
Confirm the design and choose the tool. Identify whether the study is a randomised trial, an observational study, a diagnostic accuracy study, or a systematic review, and apply the appraisal tool built for that design. The wrong tool misses the bias that matters most. [7] [8] [15]
Extract the absolute effect and its precision. Pull the absolute risk reduction, the number needed to treat, and the 95 percent confidence interval, not only the relative figure or the p-value. The family needs the magnitude they can feel, bounded by the range the data allow. [13] [1]
Hunt for publication and reporting bias in a meta-analysis. Check whether the review searched trial registers, looked for unpublished data, and examined funnel-plot asymmetry. A pooled estimate built only on published positive studies is an overestimate. [12] [6]
Document the appraisal. Record the question asked, the evidence found, its level and certainty, the decision reached, and the plan to reassess. Documentation turns a private judgement into a defensible, teachable act. [1] [14]
Management — Resuscitation
Some moments in evidence work are emergencies of a different kind, where a wrong or inapplicable result is about to drive harm. [1] [16]
Imminent treatment on a flawed study. A child is about to receive a therapy on the strength of a biased, underpowered, or inapplicable study. Correct the appraisal before the decision, state the certainty honestly, and do not let momentum carry a harmful choice. [7] [1]
Demand for an unproven intervention. A family, swayed by marketing or online claims, insists on a low-certainty or disproven treatment. Acknowledge the hope, present the best evidence in plain absolute terms, and offer a shared, evidence-based alternative. [1] [14]
Guideline that conflicts with the evidence or the patient. When a guideline outruns its evidence or clashes with the child's circumstances, depart from it with explicit reasons, document the reasoning, and seek a second opinion where the stakes are high. [10] [11]
Misinformation in the team. A colleague overclaims a result or spreads an uncritical summary. Correct it respectfully, point to the independent systematic review, and align the team on one defensible message for the family. [14] [16]
Management — Definitive & Stepwise
Work through the five steps in order, then rate the certainty of what you have found. [1] [14]
- Ask. Frame a focused PICO question that names the population, intervention, comparator, and outcome, so the search that follows is precise. [2]
- Acquire. Search the filtered layer first, then primary studies with a reproducible strategy, and record where you looked and what you found. [14]
- Appraise. Score the risk of bias with the matched tool, then extract the absolute effect and its confidence interval, and check for publication bias in any synthesis. [7] [12]
- Apply. Integrate the evidence with clinical expertise and the family's values through shared decision-making, after confirming the result applies to the child. [1] [10]
- Assess. Audit whether the decision achieved its outcome, and update your approach as better evidence arrives. [1]
Rate the certainty with GRADE. GRADE starts each body of evidence at a level set by its design and moves it down for risk of bias, inconsistency, indirectness, imprecision, and publication bias, and up for a large effect, a dose-response, or plausible residual confounding. The resulting certainty — high, moderate, low, or very low — and the balance of benefits, harms, values, and resources set the direction and strength of the recommendation. [9] [10] [11]
What the framework supports overall. Systematic reviews of appraised randomised trials provide the most trustworthy estimates of effect, and GRADE makes the move from that evidence to a transparent recommendation. A well-built question, a risk-of-bias score, an absolute effect with its interval, and a certainty rating are the four moves that defend any evidence-based decision. [14] [9]

Specific Subtypes & Scenarios
Appraising a randomised controlled trial. Use the Users' Guides and RoB 2 to score randomisation, allocation concealment, blinding, attrition, and selective reporting, then report the absolute effect and its confidence interval. CONSORT 2010 tells you whether the trial was reported completely enough to appraise at all. [5] [7]
Appraising a systematic review or meta-analysis. Check the PRISMA reporting, confirm the search was reproducible and included unpublished data, and examine heterogeneity and funnel-plot asymmetry before you trust the pooled estimate. A well-conducted review of well-conducted trials is the most reliable evidence there is. [6] [14]
Appraising a diagnostic accuracy study. Apply QUADAS-2 to patient selection, index test, reference standard, and flow and timing, then report sensitivity, specificity, and likelihood ratios in the clinically relevant subgroup. The Users' Guides walk you from validity to whether the results will help your patient. [3] [4] [15]
Appraising a clinical practice guideline. Judge whether it was developed by an accountable group, based on a systematic review, current, and applicable to your setting, and check how it rated its own evidence. A trustworthy guideline is a starting point, not a substitute for the child in front of you. [10] [11]
Rating certainty and moving to a recommendation with GRADE. Combine the certainty of the evidence with the balance of benefits and harms, values and preferences, and resource use to set a strong or weak recommendation for or against. The strength tells the clinician how confidently to follow it, not how large the effect is. [9] [11]
Resolving two conflicting studies. Weigh the validity and precision of each, look for heterogeneity and the reasons for it, and let the totality of the appraised evidence — not the single most striking paper — settle the question. [14] [12]
Complications & Pitfalls
- Equating statistical significance with clinical importance and acting on a trivial p-value. [13]
- Reading results before validity and accepting a biased estimate as truth. [7]
- Quoting a relative risk reduction without the absolute baseline or number needed to treat. [1]
- Treating a confidence interval that crosses the null as a positive result. [13]
- Extrapolating trial results to a child who differs in age, severity, or comorbidity without checking applicability. [14]
- Trusting a narrative review or an industry summary over an independent systematic review. [14] [16]
- Applying a guideline rigidly without weighing the certainty of its evidence or the patient's values. [10] [11]
- Ignoring publication bias when pooling only published studies into a meta-analysis. [12] [16]
Prognosis & Disposition
A good appraisal is measured by the decision it supports, not by how elegantly it ran. [1] [14]
Markers of success. The decision rests on the best current evidence, rated transparently for certainty, and matched to the family's values and circumstances. The clinician can state the question, the evidence, its certainty, and the reasoning in plain terms. [9] [10]
When to defer. Where the evidence is weak and higher-quality work is imminent or the decision is reversible, defer acting until better evidence arrives, or choose the reversible option and reassess. [1]
When to escalate. Sparse evidence combined with high stakes warrants a second opinion, specialist input, or ethics consultation, especially where values conflict or the child is vulnerable. [10] [11]
Disposition includes documentation. Record the question, the evidence found, its level and certainty, the shared decision, and the plan to reassess, so the reasoning survives the moment and teaches the next clinician. [14]
Special Populations
Neonates and infants. Much of the evidence is extrapolated from older children or adults, so appraise directness and dose with particular care and watch for harm that age-specific pharmacology can produce. [1] [16]
Children with rare disease. Evidence may be limited to case series and expert consensus, so rate the certainty low and lean on shared decision-making, patient registries, and networks that pool the few cases that exist. [14]
Aboriginal and Torres Strait Islander, Maori, and other Indigenous children. Appraise whether the evidence was generated with the community and is culturally and epidemiologically applicable, and privilege locally endorsed guidance and Indigenous data sovereignty. [1]
Adolescents. Ensure the trial population actually includes adolescents, and apply the evidence with the young person through shared decision-making that respects their emerging autonomy. [10]
Children with medical complexity. Weigh applicability to a heterogeneous group that trials often exclude, and prioritise patient-centred outcomes such as quality of life and family burden alongside survival. [14]
Evidence, Guidelines & Regional Differences
Core anchors are the Sackett definition of evidence-based medicine, Richardson on the well-built clinical question, the Jaeschke Users' Guides for diagnostic tests, the Schulz CONSORT statement for trials, the Sterne RoB 2 and ROBINS-I tools for risk of bias, the Moher and Liberati PRISMA statement for systematic reviews, the Balshem and Andrews GRADE guidelines for rating certainty and moving to recommendations, the Egger funnel-plot test for publication bias, the Greenhalgh how-to-read-a-paper series, the Murad Users' Guide for systematic reviews, the Whiting QUADAS-2 tool for diagnostic accuracy, and the Ioannidis argument on why most published findings are false. [1] [2] [3] [4] [5] [6] [7] [8] [9] [10] [11] [12] [13] [14] [15] [16]
The RACP curriculum names evidence-based practice and critical appraisal as core professional skills, and the Cochrane Australasia centre and national guideline organisations publish appraised syntheses. Use locally endorsed guidelines and registries, and where evidence is sparse for Aboriginal and Torres Strait Islander and Maori children, seek community-generated data and culturally safe pathways. [1] [14]
The RCPCH Progress+ curriculum frames evidence-based practice and research methodology as a core professional skill, and the National Institute for Health and Care Excellence publishes transparently graded guidelines. Use NICE and Cochrane syntheses, and follow PRISMA and GRADE in any appraisal you perform. [6] [9]
The American Academy of Pediatrics issues clinical practice guidelines built on graded evidence, and the United States Preventive Services Task Force grades recommendations on a transparent certainty-and-benefit scale. Use these alongside Cochrane reviews, and apply GRADE when moving from evidence to a recommendation. [9] [10]
The CanMEDS Scholar role maps directly onto locating, appraising, and applying evidence, and Canadian guideline bodies publish GRADE-based recommendations. Use locally endorsed guidance and appraised syntheses, and document the certainty of the evidence behind each decision. [9] [14]
Controversies: how much weight to give mechanistic reasoning and observational data against the randomised-trial hierarchy; whether a single large well-conducted trial can outweigh a meta-analysis of smaller ones; and how to apply GRADE in paediatrics when direct evidence is sparse. Exam answers show a structured appraisal, an honest certainty rating, and local humility. [1] [9] [16]
Exam Pearls
- Frame the question in PICO before you search, so the search is focused and the appraisal is fair. [2]
- Judge validity before results; a biased estimate is wrong no matter how large. [7]
- Report absolute risk reduction and number needed to treat alongside relative risk reduction. [1]
- A confidence interval that crosses the null is not a positive result. [13]
- Match the tool to the design: RoB 2 for trials, ROBINS-I for observational, QUADAS-2 for diagnostic. [7] [8] [15]
- GRADE rates certainty as high, moderate, low, or very low, and recommendations as strong or weak. [9] [11]
- Statistical significance is not clinical significance; a small p-value can accompany a trivial effect. [13]
- Applicability turns on whether the population, intervention, comparator, and outcome match your patient. [14]
Appraise and apply a paper at the bedside
Ask a focused PICO question
Acquire the best evidence, filtered layer first
Appraise validity with the matched risk-of-bias tool
Extract the absolute effect and its confidence interval
Rate the certainty of the evidence with GRADE
Apply through shared decision-making and assess the outcome
References
- [1]Sackett DL, Rosenberg WM, Gray JA, Haynes RB, Richardson WS Evidence based medicine: what it is and what it isn't BMJ, 1996.PMID 8555924
- [2]Richardson WS, Wilson MC, Nishikawa J, Hayward RS The well-built clinical question: a key to evidence-based decisions ACP J Club, 1995.PMID 7582737
- [3]Jaeschke R, Guyatt G, Sackett DL Users' guides to the medical literature. III. How to use an article about a diagnostic test. A. Are the results of the study valid? Evidence-Based Medicine Working Group JAMA, 1994.PMID 8283589
- [4]Jaeschke R, Guyatt GH, Sackett DL Users' guides to the medical literature. III. How to use an article about a diagnostic test. B. What are the results and will they help me in caring for my patients? The Evidence-Based Medicine Working Group JAMA, 1994.PMID 8309035
- [5]Schulz KF, Altman DG, Moher D CONSORT 2010 statement: Updated guidelines for reporting parallel group randomised trials J Pharmacol Pharmacother, 2010.PMID 21350618
- [6]Moher D, Liberati A, Tetzlaff J, Altman DG, PRISMA Group Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statement J Clin Epidemiol, 2009.PMID 19631508
- [7]Sterne JAC, Savovic J, Page MJ, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials BMJ, 2019.PMID 31462531
- [8]Sterne JA, Hernan MA, Reeves BC, et al. ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions BMJ, 2016.PMID 27733354
- [9]Balshem H, Helfand M, Schunemann HJ, et al. GRADE guidelines: 3. Rating the quality of evidence J Clin Epidemiol, 2011.PMID 21208779
- [10]Andrews J, Guyatt G, Oxman AD, et al. GRADE guidelines: 14. Going from evidence to recommendations: the significance and presentation of recommendations J Clin Epidemiol, 2013.PMID 23312392
- [11]Andrews JC, Schunemann HJ, Oxman AD, et al. GRADE guidelines: 15. Going from evidence to recommendation-determinants of a recommendation's direction and strength J Clin Epidemiol, 2013.PMID 23570745
- [12]Egger M, Davey Smith G, Schneider M, Minder C Bias in meta-analysis detected by a simple, graphical test BMJ, 1997.PMID 9310563
- [13]Greenhalgh T How to read a paper. Statistics for the non-statistician. II: Significant relations and their pitfalls BMJ, 1997.PMID 9277611
- [14]Murad MH, Montori VM, Ioannidis JP, et al. How to read a systematic review and meta-analysis and apply the results to patient care: users' guides to the medical literature JAMA, 2014.PMID 25005654
- [15]Whiting PF, Rutjes AW, Westwood ME, et al. QUADAS-2: a revised tool for the quality assessment of diagnostic accuracy studies Ann Intern Med, 2011.PMID 22007046
- [16]Ioannidis JP Why most published research findings are false PLoS Med, 2005.PMID 16060722