New procedures arrive in sport and exercise medicine (SEM) faster than evidence does, and the clinician who can appraise one properly is more useful than the clinician who has memorised the current list of them. Procedures are harder to evaluate than drugs: there is no dose to standardise, the operator learns as they go, blinding is difficult, patient and clinician preferences are strong, and the intervention itself often changes shape while it is being studied. This page sets out a structured way to think about where a procedure sits in its evolution, what study design should be expected at that point, and what to look for when reading a paper or being asked to offer something new.
Where is the procedure in its evolution?
The most useful mental model for procedural innovation is a staged pathway, most commonly described by the framework known by the initials of its stages: idea, development, exploration, assessment and long-term study, often abbreviated to IDEAL, with pre-clinical work sitting before the first stage. At the idea stage the procedure is used in humans for the first time, and the appropriate output is a transparent report of what was done and what happened, not a claim of effectiveness. At the development stage the technique is refined case by case, the operator's learning curve is steep, and modifications are frequently made without being documented. At the exploration stage the procedure has reached a reasonably stable form and can be studied prospectively in a cohort, with attention to patient selection and quality of delivery. Only at the assessment stage is a randomised comparison against the current standard appropriate and feasible, and the long-term study stage uses registries and surveillance to detect rare and late events that trials are too small and too short to find.
A structured pathway describes how an interventional procedure normally evolves, from first use in humans through refinement and prospective cohorts to randomised comparison and long-term surveillance.
The practical value of this model is that it tells you which criticisms are fair. Complaining that a first-in-human case series was not randomised misses the point, because randomisation at that stage would be premature and probably unethical. Conversely, a procedure that has been in widespread commercial use for a decade and still has only case series behind it is not early in its evolution; it has skipped a stage, and that is a legitimate and serious criticism. The framework also explains two recurring problems in this field. First, the learning curve means early results may understate a procedure's potential while later results from enthusiasts may overstate it. Second, procedures evolve during evaluation, so a trial may test a version of the procedure that nobody performs by the time it reports.
What should you look for in the evidence?
Six questions do most of the work. The first is what the procedure was compared with, because a comparison against no treatment measures the effect of having something done rather than the effect of the procedure. The second is whether the control was credible: for an invasive procedure that usually means a sham, which is difficult to design and to obtain approval for, but without it the substantial placebo response to an impressive intervention remains unmeasured. The third is who was included and excluded, since a procedure studied in carefully selected patients who have failed everything else may perform very differently in an unselected clinic population. The fourth is what was measured, and specifically whether the outcome is something the patient experiences, such as pain and function, or a surrogate such as an imaging appearance or a Doppler signal, which may change without the patient feeling any better. The fifth is who paid for the study and who stands to benefit, since commercial funding and clinician equity in a device are not disqualifying but do change how confidently a single positive result should be read. The sixth is what happens to patients who receive nothing, because in conditions that often improve with time, an uncontrolled series will look impressive whatever is injected.
Six questions do most of the work when appraising a new procedure: the comparator, the credibility of the control, patient selection, the outcome chosen, funding and conflicts, and the natural history of the condition.
Two further points are worth carrying. Statistical significance is not clinical importance, and the question to ask of any difference is whether it exceeds the smallest change patients actually notice, conventionally described as the minimal clinically important difference (MCID); many procedural trials report differences that are statistically significant and clinically trivial. And the absence of evidence of harm in a small early study is not evidence of safety, because these studies are almost always too small to detect uncommon complications, which is precisely why registries and long-term surveillance form the final stage of the pathway. In the UK, the national body that appraises interventional procedures publishes positions that reflect exactly this reasoning, and the recurring formulations are worth recognising: use with special arrangements for clinical governance, consent and audit or research where efficacy evidence is inadequate, and use only in the context of research where it is weaker still.
Exam Tips
•Procedures are harder to evaluate than drugs because there is no dose to standardise, operators learn as they go, blinding is difficult and the intervention often changes during evaluation.
•The staged pathway runs idea, development, exploration, assessment and long-term study, with pre-clinical work beforehand; the expected study design differs at each stage.
•Criticising a first-in-human series for not being randomised misses the point, but a decade-old commercial procedure with only case series behind it has skipped a stage.
•Ask six questions: the comparator, whether the control was credible, patient selection, whether the outcome is patient-reported or a surrogate, funding and conflicts, and the natural history of the condition.
•Statistical significance is not clinical importance; ask whether the difference exceeds the minimal clinically important difference (MCID).
•Absence of harm in a small early study is not evidence of safety, which is why registries and long-term surveillance form the final stage.