CLINICAL TECHNOLOGY REVIEWEDITION 01 / SEPTEMBER 2026
Explore results →
← Articles & methodology

RESULTS · Published

The best AI scribe for neurology

Hanah was the best AI scribe we tested on a suspected transient ischaemic attack (TIA) consultation. Heidi was second, within the margin of error on time, and Lyrebird was third.

The test

The consultation is a GP video appointment from the PriMock57 set: a patient with episodes of pins and needles and numbness in the right hand, with weakness of the right arm and leg, that come and go over a few days. Six AI scribes recorded it five times each, and every note was graded blind against the 46 clinical facts said in the consultation. We rank scribes by time to a finalised note, which is the wait after pressing stop plus the time a clinician needs to review the note and fix what is wrong. The recording, transcripts, notes and grading reasons are on the test page.

Results

RankScribeTime to a finalised noteNote errors per noteCoverageHallucination flags per noteWord error rate
1Hanah3 min 41 s5.494%0.89.1%
2Heidi5 min 07 s+1 min 27 s11.8+6.479%−15%0.410.8%+1.7%
3Lyrebird6 min 09 s+2 min 28 s14.6+9.282%−12%3.6+2.815.1%+6.0%
4CliniScripts6 min 12 s+2 min 32 s12.2+6.880%−14%1.4+0.68.7%

A Hanah note was ready to finalise 3 minutes 41 seconds after the consultation ended, on average. It captured 94% of the facts, with 5.4 note errors per note. Reviewing and fixing it took 62% less time than writing the note yourself.

Heidi averaged 5 minutes 7 seconds. Its quickest run, at 3 minutes 50 seconds, beat Hanah's slowest at 3 minutes 51 seconds, which puts the two within the margin of error on time. Hanah was quicker on average and had less than half of Heidi's note errors, 5.4 against 11.8. Heidi had the fewest hallucination flags at 0.4 per note, with Hanah at 0.8.

Lyrebird was third at 6 minutes 9 seconds and CliniScripts fourth at 6 minutes 12 seconds. CliniScripts had the lowest word error rate on this consultation at 8.7%, with Hanah at 9.1%.

Findings

Figures in this article come from the ScribeStandard benchmark. See the full results, each test and how we test.