Uncategorized

AI-Powered Depression Screening Tool Outperforms Standard Questionnaire, Chinese Study Finds

A Zhengzhou Normal University study found that a ChatGPT-based depression screening tool, BDI-FS-GPT, achieved 89.3% sensitivity, outperforming the widely used PHQ-9 questionnaire, though researchers caution larger trials are still needed.

A new tool that lets a ChatGPT-based system conduct depression screening through open-ended conversation, rather than forcing patients to pick from a fixed list of answers, has shown higher accuracy than one of the most widely used screening questionnaires in medicine. The study, conducted by researchers at Zhengzhou Normal University, was published in JMIR Formative Research on April 13, 2026.

Rethinking How Screening Questions Get Asked

Standard depression screening tools like the PHQ-9 and the Beck Depression Inventory Fast Screen (BDI-FS) ask patients to rate their symptoms on a fixed numerical scale, a format that is fast and easy to score but can flatten the nuance of how someone actually describes their emotional state. The Zhengzhou team built a system called BDI-FS-GPT, which embeds the seven-item BDI-FS into a custom ChatGPT-based interface. Instead of clicking a number, patients respond in their own words, and the AI system analyzes that natural language and maps it back to the BDI-FS’s standardized scoring anchors using rule-based processing.

How the Study Was Designed

Researchers tested the tool on 115 participants, including 28 people who had already been diagnosed with depression by clinicians. Each participant’s results from BDI-FS-GPT were compared against both their clinical diagnosis and their scores on the traditional PHQ-9 questionnaire, allowing the team to measure not just whether the AI tool worked, but whether it worked better than the screening instrument already in widespread clinical use.

The Numbers Behind the Claim

The results favored the AI-driven approach. BDI-FS-GPT achieved 89.3% sensitivity in identifying depression, alongside an 11.5% false-positive rate. By comparison, the standard PHQ-9 showed only 71.4% agreement with clinician diagnoses in the same study population. When researchers compared the overall diagnostic accuracy of the two approaches using area-under-the-curve analysis, a common statistical measure of a screening tool’s discriminative power, BDI-FS-GPT also came out ahead of the PHQ-9’s benchmark score of 0.859. Participants reported slightly but significantly higher satisfaction with the conversational BDI-FS-GPT format compared to the traditional fixed-response version of the same questionnaire.

Why This Matters for Underdiagnosis

The push for better AI-assisted screening tools is partly a response to how often depression goes undetected in routine care, particularly among older adults and people who present with physical rather than emotional symptoms. Separate research this year on AI-based screening models built from routine health checkup data has highlighted the same underlying problem: depression frequently remains undiagnosed due to stigma, somatic symptom presentation, and limited access to mental health specialists, especially among people living alone. A conversational tool that can pick up on subtler linguistic cues than a checkbox questionnaire is seen by researchers as one way to close that detection gap.

The Study’s Own Limits and the Skeptical View

The Zhengzhou researchers were careful to flag the boundaries of their findings. The study was conducted at a single clinical setting, and participants with severe depression were deliberately excluded from the trial for safety reasons, meaning the tool has not yet been validated for the population that arguably needs accurate screening the most. The authors explicitly called for broader, multi-site validation before the tool could be considered for wider clinical implementation. That caution echoes a wider pattern in AI depression research, where studies built on relatively small samples and narrow clinical settings often struggle to generalize once deployed at scale, and where an AI system’s apparent statistical edge over a decades-old paper questionnaire does not automatically translate into better real-world outcomes for patients.

What Comes Next

If validated in larger and more diverse populations, tools like BDI-FS-GPT could reshape how depression screening happens in primary care and community health settings, particularly in places where specialist mental health resources are scarce. The broader research trend, which one recent scoping review found spans 87 primary studies published mostly between January 2024 and April 2026, shows the field moving away from isolated statistical models toward more sophisticated systems that combine natural language processing with clinical scoring frameworks. Whether regulators and health systems adopt these AI-driven screening tools will likely hinge on exactly the kind of larger, more representative trials the Zhengzhou team says still needs to happen. Health systems considering these tools will also need to weigh practical questions beyond raw accuracy, including how patient data from open-ended conversations is stored and protected, how clinicians are trained to act on AI-generated screening results, and how a tool trained largely on one population and language performs when deployed elsewhere. Those implementation questions, as much as the underlying statistics, are likely to determine whether conversational screening tools like BDI-FS-GPT move from a promising single-site study into something primary care providers actually use.