CLINICAL TECHNOLOGY REVIEWEDITION 01 / SEPTEMBER 2026
Explore results →
← Articles & methodology

RESULTS · Published

The best AI scribe for gastroenterology

Hanah was the best AI scribe we tested on a gastroenteritis consultation. Heidi came second, close enough that the gap is within the margin of error, and Lyrebird was third.

The test

The consultation is a GP video appointment from the PriMock57 set: three days of watery diarrhoea with cramping lower abdominal pain, after illness in the household. Six AI scribes recorded it five times each, and every note was graded blind against the 45 clinical facts said in the consultation. We rank scribes by time to a finalised note, which is the wait after pressing stop plus the time a clinician needs to review the note and fix what is wrong. The recording, transcripts, notes and grading reasons are on the test page.

Results

RankScribeTime to a finalised noteNote errors per noteCoverageHallucination flags per noteWord error rate
1Hanah2 min 16 s2.896%0.215.2%
2Heidi3 min 19 s+1 min 03 s9.6+6.883%−12%1.6+1.416.0%+0.8%
3Lyrebird3 min 51 s+1 min 36 s11.6+8.883%−12%3.4+3.220.0%+4.9%
4Freed4 min 17 s+2 min 01 s6.0+3.291%−5%0.8+0.621.4%+6.3%

A Hanah note was ready to finalise 2 minutes 16 seconds after the consultation ended, on average. It captured 96% of the facts, with 2.8 note errors and 0.2 hallucination flags per note. Reviewing and fixing it took 77% less time than writing the note yourself.

Heidi averaged 3 minutes 19 seconds. Its quickest run, at 2 minutes 9 seconds, beat Hanah's slowest at 2 minutes 30 seconds, which puts the two within the margin of error on time. Hanah was quicker on average and had less than a third of Heidi's note errors, 2.8 against 9.6.

Lyrebird was third at 3 minutes 51 seconds and Freed fourth at 4 minutes 17 seconds. Freed had the second fewest note errors at 6.0 and the second highest coverage at 91%, but took longer to produce its note.

Hanah also had the lowest word error rate on this consultation at 15.2%, with Heidi at 16.0%.

Findings

Figures in this article come from the ScribeStandard benchmark. See the full results, each test and how we test.