top of page
Search

The Score Is Not the Story: What Rating Scales Can—and Cannot—Tell Us About Mental Health

  • Writer: Sarah Marchand-Lacoursière
    Sarah Marchand-Lacoursière
  • Jun 29
  • 8 min read

Mental health rating scales are everywhere. The PHQ-9, GAD-7, HAM-D, BDI, PANSS, MADRS, and PCL-5 have become part of the ordinary language of psychiatry, primary care, clinical trials, epidemiology, and digital mental health. Their appeal is obvious: they are brief, standardized, scalable, easy to score, and useful for communication across clinicians, settings, and research studies.


But their very usefulness can create a dangerous illusion. A rating scale can make distress appear more measurable than it really is. It can turn a person’s private, contextual, fluctuating experience into a number that looks precise, comparable, and clinically definitive. The central lesson from the evidence is not that rating scales are useless. It is that a score is not a diagnosis, a cut-off is not a clinical formulation, and a symptom checklist is not the same thing as understanding a person.


The promise and the trap of quantification

Rating scales were designed to operationalize mental phenomena that are otherwise difficult to observe directly. The HAM-D emerged in 1960, the BDI in 1961, and the PHQ-9 later became one of the most widely used depression screening tools in primary care (Hamilton, 1960; Beck et al., 1961; Kroenke et al., 2001). These tools helped psychiatry standardize symptom assessment, compare outcomes across studies, and monitor change over time.


Yet psychiatric symptoms are not laboratory values. When a patient answers a question such as whether they have felt “down, depressed, or hopeless,” the response reflects much more than symptom severity. It also reflects how they interpret the words, how they remember the past two weeks, what they believe is socially acceptable to disclose, how they understand response options such as “several days,” and whether the item captures the way their suffering actually appears in their life. Uher (2022) argues that rating scales often conflate the mental phenomena being studied with the verbal and numerical tools used to study them. The result is a form of “pseudoprecision”: numbers that look objective, but may rest on assumptions that have not been formally demonstrated.


This matters because summed scores are often treated as if they represent a stable, continuous quantity. But many scales rely on ordinal response options, assume that items can be added together, and impose cut-offs on experiences that may not have natural thresholds. Depression, anxiety, and psychotic phenomena are often better understood as dimensions rather than discrete categories (Brown & Barlow, 2005). A cut-off can be clinically useful, but it can also make a gradient look like a boundary.


The PHQ-9 paradox

The PHQ-9 illustrates the paradox clearly. It is one of the best-known and most validated tools in mental health, yet also one of the strongest examples of what goes wrong when screening is confused with diagnosis.


In an individual participant data meta-analysis including 44 primary studies and 9,242 participants, Levis et al. (2020) found that PHQ-9 scores of 10 or higher produced a pooled depression prevalence of 24.6%, while SCID-based major depressive disorder prevalence in the same samples was 12.1%. In other words, the standard PHQ-9 cut-off identified approximately 2.5 times as many cases as were confirmed by structured diagnostic interview.


That does not mean the PHQ-9 is a bad screener. At the commonly used threshold of 10 or higher, another meta-analysis found pooled sensitivity of 0.88 and specificity of 0.85 compared with semi-structured diagnostic interviews (Levis et al., 2019). But a screening tool is designed to identify people who need further assessment. It is not designed to replace the assessment.


The distinction is not semantic. When a tool optimized for screening is used diagnostically, false positives become clinically consequential. People may be labelled, treated, or monitored as if they have a disorder that has not been established through proper clinical evaluation. At the same time, false negatives remain possible: even at an optimal cut-off, the PHQ-9 may miss a meaningful minority of true major depressive episodes (Levis et al., 2019).


Suicide risk cannot be reduced to one item

The limitations become especially important when risk is involved. Item 9 of the PHQ-9 asks about thoughts of being better off dead or self-harm. It is clinically important, but it is not a suicide risk assessment.


Simon et al. (2018) found that PHQ-9 item 9 had a sensitivity of 87.6%, specificity of 66.1%, and positive predictive value of only 28.6% for suicide risk as defined by the Columbia Suicide Severity Rating Scale. Fewer than one in three people flagged by the item met criteria for clinically significant suicidal ideation by that standard. More concerningly, more than one-third of patients who attempted suicide or died by suicide within 30 days of PHQ-9 screening had reported no suicidal ideation on item 9 (Davis et al., 2024; Brent et al., 2024).


The lesson is not to ignore the item. The lesson is to understand what it is: a signal that may require follow-up, not a complete risk formulation. Suicide risk is dynamic, contextual, multidimensional, and often concealed. No single frequency item can carry that clinical responsibility.


Symptoms are not context-free

Many rating scales include somatic items such as sleep disturbance, fatigue, appetite change, concentration problems, and psychomotor changes. These symptoms can be part of depression. They can also reflect chronic pain, traumatic brain injury, medical illness, medications, perinatal changes, aging, anxiety, substance use, or ordinary stress.


This creates a validity problem. Williams et al. (2004) showed that depression measures emphasizing somatic symptoms may overestimate depression severity in patients with chronic pain. The BDI has been criticized for poor discriminant validity against anxiety (Steer et al., 1997). The HAM-D places disproportionate emphasis on somatic and insomnia symptoms, which can distort interpretation in antidepressant trials: a medication that improves sleep may appear more effective, while a medication with gastrointestinal side effects may appear less effective, even when the underlying antidepressant effect differs from the score (Bagby et al., 2004; Molnar et al., 2022).


Psychiatric diagnosis depends on context: onset, duration, functional impairment, meaning, developmental history, comorbidity, medical factors, trauma, culture, and longitudinal course. Rating scales capture fragments of this reality. They do not reconstruct the whole pattern.


Equity begins with measurement

A score is only fair if it means the same thing across people. That assumption often fails.

The PHQ-9 was developed in predominantly White, English-speaking North American primary care populations. Its use across cultures, languages, and demographic groups requires evidence of measurement invariance: the items, factor structure, and score interpretation should operate similarly across groups. Studies have found partial or non-invariance across American Indian and Alaska Native adults, African American and non-Latinx White patients, and other cultural and linguistic groups (Forkus et al., 2019; Tompkins et al., 2021). Evidence from low- and middle-income countries suggests that sensitivity and specificity may be lower than in high-income settings.


This is not a technical footnote. If an item about sleep, appetite, energy, or psychomotor change functions differently across groups, then the same score may not represent the same clinical reality. If emotional distress is expressed through somatic idioms in one culture and cognitive-affective language in another, a standardized questionnaire may under-detect one kind of suffering while over-detecting another.

There is also a deeper equity problem: many widely used measures were not developed with substantive involvement from people with lived experience. Patients often identify outcomes such as meaning, identity, connection, recovery, and quality of life as central to mental health, yet these domains may be absent from standard symptom scales. What we measure shapes what we value. What we fail to measure can disappear from care.


Measurement-based care needs clinical humility

Measurement-based care can improve communication, self-awareness, and monitoring. But evidence from patients and clinicians shows persistent concern that standardized measures may fail to reflect clinical complexity, introduce reporting biases, and reduce care to a numerical exercise (Rødevand, de Jong, & Bschor, 2025). Delgadillo et al. (2021) also warned that mandating a narrow set of standardized instruments may reduce methodological diversity, create artificial hierarchies between imperfect tools, and narrow the scope of mental health science.


Screening without infrastructure is another ethical risk. Khaled et al. (2025) emphasized that administering mental health questionnaires without clear protocols for positive screens can generate harm, especially when suicidal ideation or severe distress is identified but no follow-up pathway exists. Screening only improves care when it is embedded in systems capable of assessment, treatment, referral, and monitoring.


Toward better assessment, not less assessment

The evidence does not argue for abandoning rating scales. It argues for putting them in their proper place: useful signals within a broader clinical assessment, not substitutes for one. Across instruments, the same limitation appears again and again: mental health cannot be reduced to a score. Clinical meaning depends on context, trajectory, comorbidity, functioning, culture, risk, protective factors, and the patient’s own language.


The next standard in psychiatric assessment should therefore move beyond isolated screening outputs toward a more complete, structured, and contextualized picture of the person. This is the direction Aion is pursuing with Elyx: a future where clinical information is gathered with greater depth, organized with greater clarity, and used to support more informed human judgment.


A rating scale can start the conversation. The goal now is to help clinicians see more of the story before that conversation begins.


References

Bagby, R. M., Ryder, A. G., Schuller, D. R., & Marshall, M. B. (2004). The Hamilton Depression Rating Scale: Has the gold standard become a lead weight? American Journal of Psychiatry, 161(12), 2163–2177.

Beck, A. T., Ward, C. H., Mendelson, M., Mock, J., & Erbaugh, J. (1961). An inventory for measuring depression. Archives of General Psychiatry, 4(6), 561–571.

Brown, T. A., & Barlow, D. H. (2005). Dimensional versus categorical classification of mental disorders in the fifth edition of the Diagnostic and Statistical Manual of Mental Disorders and beyond: Comment on the special section. Journal of Abnormal Psychology, 114(4), 551–556.

Delgadillo, J., de Jong, K., & Lucock, M. (2021). Editorial perspective: Prescribing measures — Unintended negative consequences of mandating standardized mental health measurement. Journal of Child Psychology and Psychiatry, 62(8), 1032–1036.

Forkus, S. R., Tolin, D. F., Mathew, A. R., & DiMauro, J. (2019). Measurement invariance of the Patient Health Questionnaire-9 (PHQ-9) across U.S. sociodemographic groups. Journal of Affective Disorders, 256, 115–123.

Hamilton, M. (1960). A rating scale for depression. Journal of Neurology, Neurosurgery, and Psychiatry, 23(1), 56–62.

Khaled, K., Tsofliou, F., & Hundley, V. (2025). Ethical issues and challenges regarding the use of mental health questionnaires in public health nutrition research. Nutrients, 17(4), 715.

Kroenke, K., Spitzer, R. L., & Williams, J. B. W. (2001). The PHQ-9: Validity of a brief depression severity measure. Journal of General Internal Medicine, 16(9), 606–613.

Levis, B., Benedetti, A., Ioannidis, J. P. A., Sun, Y., Negeri, Z., He, C., Wu, Y., Krishnan, A., Bhandari, P. M., Neupane, D., Osório, F. de L., Loureiro, S. R., Chagas, M. H. N., & the DEPRESSD Collaboration. (2020). Patient Health Questionnaire-9 scores do not accurately estimate depression prevalence: Individual participant data meta-analysis. Journal of Clinical Epidemiology, 122, 115–128.

Levis, B., Benedetti, A., Thombs, B. D., & the DEPRESSD Collaboration. (2019). Accuracy of Patient Health Questionnaire-9 (PHQ-9) for screening to detect major depression: Individual participant data meta-analysis. BMJ, 365, l1476.

Molnar, G., Fonseka, T. M., & Cipriani, A. (2022). The HAM-D is not Hamilton's depression scale. BJPsych Open, 8(Suppl. 1).

Rødevand, L., de Jong, K., & Bschor, T. (2025). Quantifying care, qualifying experiences: A systematic review of measurement-based care in psychiatry from patient and provider perspectives. BMJ Mental Health, 28(1), e301663.

Simon, G. E., Rutter, C. M., Peterson, D., Oliver, M., Whiteside, U., Operskalski, B., & Ludman, E. J. (2018). Does response on the PHQ-9 Depression Questionnaire predict subsequent suicide attempt or suicide death? Psychiatric Services, 64(12).

Steer, R. A., Ball, R., Ranieri, W. F., & Beck, A. T. (1997). Further evidence for the construct validity of the Beck Depression Inventory-II with psychiatric outpatients. Psychological Reports, 80(2), 443–446.

Tompkins, S. A., Barlow, A., Jones, M., et al. (2021). Evaluating the cross-cultural measurement invariance of the PHQ-9 between American Indian/Alaska Native adults and diverse racial and ethnic groups. Journal of Affective Disorders Reports, 4, 100121.

Uher, J. (2022). Rating scales institutionalise a network of logical errors and conceptual problems in research practices: A rigorous analysis showing ways to tackle psychology's crises. Frontiers in Psychology, 13, 1009893.

Williams, L. S., Jones, W. J., Shen, J., Robinson, R. L., Weinberger, M., & Kroenke, K. (2004). Prevalence and impact of depression and pain in neurology outpatients. Journal of Neurology, Neurosurgery, and Psychiatry, 74(11), 1587–1589.


 
 
 

Comments


bottom of page