CLINICAL TECHNOLOGY REVIEWEDITION 01 / SEPTEMBER 2026
Explore results →
← Articles & methodology

RESULTS · Published

The best AI scribe for mental health

Hanah was the best AI scribe we tested on a low mood consultation. Heidi and Lyrebird were close behind, and all three are within the margin of error on time to a finalised note. Hanah was quickest on average and had the fewest note errors.

The test

The consultation is a GP video appointment from the PriMock57 set: a patient with two months of low energy, poor sleep and reduced appetite since starting a new job, with a family history of suicide. Six AI scribes recorded it five times each, and every note was graded blind against the 26 clinical facts said in the consultation. We rank scribes by time to a finalised note, which is the wait after pressing stop plus the time a clinician needs to review the note and fix what is wrong. The recording, transcripts, notes and grading reasons are on the test page.

Results

RankScribeTime to a finalised noteNote errors per noteCoverageHallucination flags per noteWord error rate
1Hanah2 min 26 s4.292%1.413.7%
2Heidi2 min 45 s+19 s6.8+2.680%−12%0.813.2%
3Lyrebird3 min 02 s+36 s8.0+3.884%−8%2.8+1.417.8%+4.1%
4CliniScripts4 min 40 s+2 min 15 s13.2+9.065%−27%3.2+1.813.3%

A Hanah note was ready to finalise 2 minutes 26 seconds after the consultation ended, on average. It captured 92% of the facts, with 4.2 note errors per note. Reviewing and fixing it took 77% less time than writing the note yourself.

Heidi averaged 2 minutes 45 seconds and Lyrebird 3 minutes 2 seconds. Hanah's slowest run took 2 minutes 52 seconds, while Heidi's quickest took 2 minutes 27 seconds and Lyrebird's 2 minutes 26 seconds, which is why we treat the three as close. Heidi had 6.8 note errors per note and Lyrebird 8.0. Heidi had the fewest hallucination flags at 0.8 per note, with Hanah at 1.4.

Heidi also had the lowest word error rate on this consultation at 13.2%, with CliniScripts at 13.3% and Hanah at 13.7%. Lyrebird produced its note fastest after pressing stop, though its longer review put it third overall.

Findings

Figures in this article come from the ScribeStandard benchmark. See the full results, each test and how we test.