Many psychological instruments used in sport are self-completed by the athlete, and that is entirely appropriate; what matters is that the purpose, administration, interpretation and response pathway are right. Clinical screening and diagnostic tools require oversight by suitably qualified professionals, and a score is a prompt for a conversation rather than an answer in itself. A screening instrument indicates that further assessment is warranted; it does not make a diagnosis, and a score below a threshold does not mean an athlete is well. In sport and exercise medicine (SEM) these tools are used mainly to detect mental health symptoms early in a population that under-reports them, to assess psychological readiness for return to sport, and to support clinical decisions. This page covers what makes an instrument trustworthy, the structured approach to athlete mental health screening and an important recent correction to how it is applied, the instruments used around injury, and the governance and ethical limits that should constrain their use.
What Makes a Psychometric Instrument Trustworthy?
Two families of property underpin everything. Reliability describes consistency, and it is not a single quantity: test-retest reliability concerns agreement on repeated administration when nothing has changed, inter-rater reliability concerns agreement between different assessors where an instrument is administered or scored by a clinician, and internal consistency concerns whether items behave coherently as a set. High internal consistency does not prove that an instrument measures a single construct and can simply reflect redundant, near-duplicate items.
A simplified analogy: reliability is consistency and validity concerns whether score interpretations are supported. The analogy is imperfect, because validity is not literally accuracy at a centre and is itself constrained by poor reliability.
Validity is best understood not as one global property an instrument either has or lacks, but as the body of evidence supporting a particular interpretation and use of its scores in a particular population. An instrument can be highly reliable while its scores do not support the interpretation being placed on them, which is why both must be demonstrated for the intended purpose. Beyond these, responsiveness describes the ability to detect genuine change over time, normative data allow an individual score to be placed in context, and sensitivity and specificity describe performance against a threshold. No screening instrument achieves perfect sensitivity and specificity, so false positives and false negatives are inevitable and must be anticipated rather than treated as failures. Crucially, these properties are established in a specific population: an instrument validated in a general adult sample may behave quite differently in elite athletes, and one validated in elite athletes may not transfer to amateur or youth populations.
How Is Athlete Mental Health Screening Structured?
The most widely used structure is the International Olympic Committee Sport Mental Health Assessment Tool 1, developed for sports medicine physicians and other licensed or registered health professionals to assess elite athletes aged sixteen and over who may be at risk of, or already experiencing, mental health symptoms and disorders.
Athlete mental health assessment moves from athlete-specific triage through symptom-domain screening to clinical assessment, but a negative triage should not be used as a gate that stops the pathway, and clinical concern can enter the pathway at any point.
As published it works in three steps. Step one is triage using the Athlete Psychological Strain Questionnaire (APSQ), a short athlete-specific instrument covering performance concerns, difficulty with self-regulation and external coping, with a defined threshold. Step two applies six disorder-specific screening instruments covering anxiety, depression, sleep, alcohol use, substance use and disordered eating. Step three is either brief intervention with reassessment, or comprehensive clinical assessment and management by a sports medicine physician or a licensed mental health professional such as a psychiatrist or clinical psychologist.
One point has become important since publication and should change how programmes are run. The developers noted that the triage instrument does not reach perfect sensitivity or specificity, and subsequent evaluation in a large national Olympic and Paralympic delegation found an overall false-negative rate of around two-thirds for the triage step in identifying athletes who screened positive on the later disorder-specific instruments, a finding since echoed in further analyses. Using a negative triage as a gate that ends the pathway therefore misses a substantial number of athletes. Programmes should consider administering the full disorder-specific battery, or using another clinically governed pathway, rather than stopping after a negative triage, and clinical concern should always be able to trigger assessment irrespective of any score.
A separate recognition tool exists for athletes themselves and their entourage, including family, teammates and coaches. This is deliberately not the same instrument: it is designed to help non-clinicians notice warning signs and encourage help-seeking, and it identifies red flags such as comments about harming oneself or others that warrant immediate help. Keeping the clinician assessment tool and the recognition tool distinct preserves the principle that screening scores are interpreted by people qualified to act on them.
Testing Around Injury and Return to Sport
Psychological readiness provides information that is distinct from physical performance testing and is associated with return-to-sport outcomes, so it belongs in the decision alongside strength and functional measures rather than being treated as superior or inferior to them. The Anterior Cruciate Ligament Return to Sport after Injury scale assesses confidence, emotion and risk appraisal in athletes recovering from cruciate reconstruction, and lower scores are associated with failure to return to previous level of sport. It is condition-specific and should not be used as a generic readiness instrument for unrelated injuries. Fear of movement and reinjury is measured with kinesiophobia scales, and general distress, coping and confidence measures are also used. No single instrument clears an athlete to return.
Two further applications need a more cautious framing. Cognitive testing is an optional adjunct in concussion management rather than an essential or independently determinative test, and computerised cognitive results must not be used in isolation to diagnose concussion or to clear an athlete to return. A pre-season baseline has genuine but limited value, because performance is influenced by effort, sleep, language, attention or learning disorders, practice effects and ordinary test-retest variability, and deliberate underperformance at baseline in order to ease later clearance is one recognised limitation among several. Personality and psychometric profiling is sometimes marketed for talent identification and selection; the evidence that such profiles predict sporting success is weak, and using clinical or quasi-clinical instruments to inform selection decisions is ethically problematic and undermines the trust on which honest disclosure depends.
Governance, Consent and Limits
Because these instruments generate sensitive personal data in an environment where athletes may feel unable to decline, governance is not an afterthought. Screening should not begin unless there is a clear, timely and confidential pathway for assessment, crisis response and treatment. Consent should be genuinely informed, explaining whether participation is optional, who sees individual and aggregate results, the limits of confidentiality, how long data are retained, how urgent risk will be handled, and whether results are separated from selection and contractual decisions. Consent must be real, which means athletes can decline without penalty to selection.
Two limitations of self-report deserve attention in interpretation. Athletes may minimise symptoms where they believe results could affect selection, and evidence indicates that responses on mental health scales differ depending on whether administration is anonymous. Cultural and language factors also affect how items are understood, and using an instrument in another language requires cultural adaptation and revalidation rather than word-for-word translation alone. A final distinction is worth making: routine wellness or readiness monitoring used for performance support is not automatically equivalent to clinical mental health screening, and it should not be labelled or interpreted as though it were. For all these reasons a score is a prompt for a conversation, never a substitute for one, and a clinician who is worried about an athlete should act on that concern whatever the questionnaire says.
Exam Tips
•Reliability covers test-retest, inter-rater and internal consistency; high internal consistency does not prove a single construct.
•Validity is evidence supporting a particular interpretation and use of scores in a particular population.
•Screening indicates the need for further assessment and never makes a diagnosis.
•A negative athlete triage has a high false-negative rate and should not be used as a gate that stops the screening pathway.
•The clinician assessment tool and the athlete and entourage recognition tool are deliberately different instruments.
•Screening requires a timely confidential pathway to assessment, crisis response and treatment before it begins.