While lawsuits over general-purpose AI companion apps have dominated headlines, a very different kind of chatbot quietly produced the first randomized controlled trial data showing a purpose-built AI therapy tool can meaningfully reduce psychiatric symptoms. Therabot, developed by a team at Dartmouth College’s Geisel School of Medicine, was tested in a national trial of 210 adults and published its results on March 27, 2025, in NEJM AI, the American Medical Association-affiliated journal’s artificial intelligence publication.
A Deliberately Different Design Philosophy
Unlike consumer companion apps built on general-purpose large language models and optimized for open-ended engagement, Therabot was designed and trained from the outset by an expert clinical team, including a board-certified psychiatrist and a clinical psychologist, specifically to deliver evidence-based therapeutic content. That distinction, according to Dartmouth’s own announcement of the results, was central to the research question: could a chatbot built like a clinical intervention, rather than a general-purpose conversational product, produce measurable mental health benefits under controlled conditions.
The Trial Design
The study randomly assigned 210 adults with clinically significant symptoms of major depressive disorder, generalized anxiety disorder, or clinically high risk for feeding and eating disorders to either a four-week Therabot intervention, 106 participants, or a waitlist control group, 104 participants, according to the published trial and coverage from MIT Technology Review and Dartmouth News. Participants who used Therabot engaged heavily with it, logging an average of more than six hours of use over the study period and giving the tool high user ratings, according to Dartmouth’s summary of the results.
What the Numbers Showed
The results were the most substantial evidence to date that a generative AI chatbot can produce clinically meaningful improvement. Participants with major depressive disorder who used Therabot saw symptoms decrease by 51%, according to Dartmouth’s reporting on the trial. Those with generalized anxiety disorder reported an average symptom reduction of 31%, with many participants shifting from moderate to mild anxiety, or from mild anxiety to below the clinical diagnostic threshold entirely. Participants at clinically high risk for eating disorders saw a 19% reduction in symptoms tied to concerns about body image and weight.
Caveats the Researchers Themselves Flagged
Even the trial’s own authors were careful to frame the results as promising rather than conclusive. Coverage from Healio and the American Psychological Association’s practice publication noted that researchers and outside reviewers stressed that generative AI chatbots for mental health treatment still need clinical supervision and further study before being deployed at scale, given the relatively short four-week window and the comparison against a waitlist control rather than an active treatment like standard cognitive behavioral therapy delivered by a human clinician. Outside researchers who reviewed the study for Healio also pointed out that a waitlist comparison, while a standard and accepted design in early-stage digital therapeutics research, tends to produce larger apparent effect sizes than a head-to-head comparison against an active treatment would, meaning the true clinical benefit relative to existing therapies remains an open question.
A Sharp Contrast With the Unregulated Chatbot Market
Therabot’s results land amid intense scrutiny of general-purpose AI companion apps, which have been linked in lawsuits to teenage suicides and have faced Federal Trade Commission inquiries and new state laws restricting unsupervised AI therapy. Mental health researchers increasingly point to that contrast directly: a chatbot built by a clinical team, tested in a randomized trial, and published in a peer-reviewed medical journal is a fundamentally different product category from a general chatbot repurposed by users as an ad hoc therapist. Advocates for AI-assisted care argue Therabot demonstrates what responsible development looks like, while skeptics note that even Therabot’s promising numbers came from a controlled academic study with a relatively small sample, not a real-world deployment at the scale of consumer apps like Character.AI or ChatGPT. That gap in scale is precisely the point critics raise: it is far easier to maintain rigorous safety guardrails across 210 closely monitored trial participants than across the tens of millions of users a commercial mental health chatbot would eventually need to serve to be financially viable.
What’s Next
Dartmouth’s team has signaled interest in larger, longer follow-up trials and eventual FDA engagement, following a path already being walked by other clinically oriented mental health apps such as Woebot Health, which has FDA-cleared versions for adjunctive treatment of major depressive disorder, and Wysa and Sanvello, both of which hold FDA Breakthrough Device Designation. Whether Therabot or similar clinically supervised tools can scale beyond a 210-person study while preserving both their safety profile and their therapeutic effect is the question that will determine whether this NEJM AI trial marks the beginning of a genuinely new treatment category or remains a promising but isolated data point.