Sport and exercise medicine is unusually exposed to weak evidence. Practice spreads through conference talks, social media and what a successful team is seen to be doing, and the commercial incentive to adopt novel interventions is considerable. The defence is the ability to appraise a paper rather than to recall a conclusion. This page covers the hierarchy of evidence and why it changes with the question being asked, the main study designs and what each is good for, how to appraise a study systematically, and why a study sitting high in the hierarchy can still be worthless.
The familiar pyramid ranks designs by their vulnerability to bias, running from expert opinion at the base through case reports and series, observational studies, randomised controlled trials, and high-quality systematic reviews of appropriate studies at the apex. Meta-analysis is a statistical method rather than automatically the highest form of evidence, and a review of poor or unsuitable studies does not outrank a robust primary study simply because it is a review.
The essential caveat is that the hierarchy was constructed for questions about whether an intervention works, and it does not transfer unchanged to other question types. For intervention efficacy, randomisation is what balances confounding, so the randomised controlled trial (RCT) sits near the top. Qualitative research is not low-grade evidence; it answers different questions about experience, acceptability, implementation and context, and sits outside this hierarchy rather than beneath it. For prognosis, the question is what happens to people with a condition over time, and the appropriate design is a prospective cohort study followed from a common point early in the disease. For diagnostic accuracy, the appropriate design is a cross-sectional study comparing the index test against a reference standard in a consecutive series of patients in whom the diagnosis is genuinely uncertain, with blinded interpretation where possible and the same reference standard applied regardless of the index result. For harms, particularly rare or delayed ones, observational data and case control designs often outperform trials, which are usually too small and too short.
Modern frameworks have moved further, assessing the certainty of a body of evidence rather than the rank of a single study, taking account of risk of bias, inconsistency, indirectness, imprecision and publication bias. That is why a well-conducted observational study with a very large effect can provide more certainty than a small, poorly conducted trial.
Experimental designs allocate the intervention. The randomised controlled trial allocates at random, which makes measured and unmeasured prognostic factors balanced on average, particularly in adequately sized trials. It does not guarantee identical groups in any individual trial, but it is the only design that addresses unmeasured confounding at all. A quasi-experimental study is defined by the absence of random allocation, using methods such as allocation by clinic, by season or by clinician preference. This is often the only practicable approach in a small squad, though stronger forms such as controlled interrupted time series should not be lumped in with a simple before and after comparison.
Observational designs do not allocate anything. A cohort study follows exposed and unexposed groups forward to see who develops the outcome, and can measure incidence and relative risk. A case control study starts from people who already have the outcome and looks backwards for exposure, which makes it efficient for rare outcomes but vulnerable to recall bias and unable to give incidence. A cross-sectional study measures exposure and outcome at the same moment, which is why it establishes association rather than sequence.
Synthesis designs combine studies. A systematic review applies a pre-specified search and selection strategy to identify all relevant studies; a meta-analysis statistically pools their results. The distinction matters, because a systematic review need not include a meta-analysis, and pooling heterogeneous studies produces a precise-looking answer to a question nobody asked. Qualitative designs answer a different class of question entirely, concerning experience, meaning and barriers, and are the right tool when the question is why an intervention is not being taken up rather than whether it works.
Structured appraisal tools such as the Critical Appraisal Skills Programme (CASP) checklists exist for each major design, and all of them organise around the same three questions.
First, are the results valid. This is internal validity and it is assessed before anything else, because results that are not valid cannot become useful by being applicable. For a trial it covers whether allocation was genuinely random and concealed, whether groups were similar at baseline, whether participants and assessors were blinded, whether follow-up was adequate and whether analysis was by intention to treat.
Second, what are the results. This covers the size of the effect and the precision around it. A confidence interval is more informative than a p value because it shows the range of effects compatible with the data, and it should be read against what would count as a clinically meaningful benefit or harm. A result is uninformative when it remains compatible with materially different clinical conclusions; a narrow interval around a trivial effect may cross no effect while still excluding meaningful benefit.
Third, will the results help locally. This is external validity: whether the population, the intervention as delivered, the comparator and the outcomes resemble your setting closely enough for the finding to transfer. A trial of a rehabilitation protocol in professional male footballers may say little about recreational masters athletes.
One further check belongs to all three. Ask what outcome was actually measured. A surrogate outcome such as a change in an imaging appearance is not the same as an outcome the athlete cares about, and composite outcomes can be driven entirely by their least important component.
Position in the hierarchy describes the design, not the execution, and confusing the two is the most common appraisal error. Several failures recur.
A randomised trial can be underpowered, so that a real effect is missed and the non-significant result is misreported as evidence of no difference. Absence of evidence is not evidence of absence. Allocation may be random but not concealed, allowing recruiters to steer participants. Blinding is frequently impossible in rehabilitation and exercise trials, which matters most when the outcome is subjective. Attrition may be high or differential between arms. Analysis may be per protocol rather than by intention to treat. A naive per protocol analysis can be biased because adherence is itself influenced by prognosis and by post-randomisation factors, not simply because it removes those for whom treatment failed.
A systematic review inherits every weakness of the studies it includes. Meta-analysis can increase precision without correcting systematic bias in the results being pooled, which is what makes a tight confidence interval around a biased estimate more dangerous than no estimate at all. Heterogeneity may be substantial and glossed over. Publication bias means the studies that exist are more likely to be positive than the studies that were conducted.
Finally, statistical significance is not clinical importance. With a large enough sample a trivial difference becomes significant, which is why an effect should be judged against a threshold for meaningful change rather than against a p value, alongside the confidence interval, adverse effects, burden, cost and patient priorities. A group mean below a meaningful threshold does not establish that no individual benefited. Conflicts of interest and funding source deserve scrutiny of design, analysis, missing outcomes and sponsor involvement in a field where much research is industry-funded, though commercial funding does not by itself invalidate a finding. Registration, protocol and statistical analysis plan should be compared against the report, because absence of registration prevents verification of whether outcomes were selected after the results were known.
Sign up to get full access to 10 topics of your choice, including all sections, clinical pearls, and exam tips.
Sign up free10 free topics included with your account. Full access from £24.17/month.
Sections included with full access