CLINICAL TECHNOLOGY REVIEWEDITION 01 / SEPTEMBER 2026
Explore results →
← Articles & methodology

RESULTS · Published

Hanah is the best AI scribe in our benchmark

We recorded the same 12 consultations with six AI scribes and graded every note blind against the same inventory of clinical facts. Hanah ranked first on 18 of the 19 measures we publish, including every measure of note quality and every measure of clinician time. The recordings, transcripts, notes and grading reasons for each consultation are on its test page and can be downloaded.

Errors in the note

A note error is a relevant fact the note missed or captured only partly, or a claim the consultation does not support. We count them the way Weiner et al. (2020) counted errors in doctors' notes. Averaged over at least five runs of all 12 consultations, Hanah's notes had 4.1 errors each. Freed averaged 8.3, Heidi 9.0, CliniScripts 9.7, Lyrebird 11.6 and PatientNotes 18.0.

Hanah captured 93% of the relevant facts, against 87% for Freed, 82% for Lyrebird, 81% for Heidi, 80% for CliniScripts and 61% for PatientNotes. It also had the fewest hallucination flags at 0.80 per note, with Heidi close behind at 1.00.

Weiner et al. found 5.0 errors per note in doctors' own notes checked against recordings of the visit. Our test counts more strictly than theirs, and Hanah was the only scribe below that figure.

Time to review and finalise

We estimate how long a clinician takes to read each note and fix what is wrong with it. A Hanah note took 2 minutes 2 seconds to review, against 2 minutes 58 seconds for Heidi, 3 minutes 43 seconds for Lyrebird, 3 minutes 50 seconds for CliniScripts, 4 minutes 5 seconds for Freed and 5 minutes 11 seconds for PatientNotes. Writing the note yourself takes a median 8.1 minutes (Apathy et al. 2023), which makes reviewing and fixing a Hanah note 75% quicker than writing it by hand. Heidi was 63% quicker, Lyrebird 54%, CliniScripts 53%, Freed 50% and PatientNotes 36%.

Including the wait after pressing stop, a Hanah note was ready to finalise 2 minutes 23 seconds after the consultation ended. Heidi took 3 minutes 25 seconds, CliniScripts 4 minutes 3 seconds, Lyrebird 4 minutes 6 seconds, Freed 4 minutes 40 seconds and PatientNotes 5 minutes 39 seconds.

Hanah's slowest note in any run took 4 minutes 15 seconds to review. Every other scribe had at least one note that took more than seven and a half minutes, and the slowest CliniScripts note took 10 minutes 47 seconds.

Transcription, speed and data

Hanah had the lowest word error rate at 12.8%, with CliniScripts at 13.0% and Heidi at 13.6%. On mis-heard words that change clinical meaning, Hanah averaged 8.6 per run of the 12 consultations, CliniScripts 9.0 and the other scribes between 21 and 36.

Hanah produced its note a median 19 seconds after stop and used 9.1 MB of data per consultation. The other scribes took 21 to 49 seconds and used 17 to 45 MB.

Memory use and unedited notes

In a 50 minute stress test with every scribe recording at once on one machine, CliniScripts averaged 972 MB of memory against Hanah's 1,103 MB. It is the one measure Hanah did not lead. Hanah's CPU use was 28.6%, level with CliniScripts at 28.8%.

Hanah had the highest share of notes ready to finalise as written, 8%, meaning every fact captured in full and no hallucination flags. That is below the 10% of doctors' notes Weiner et al. found error-free, and every note from every scribe still needs a clinician's review before it is finalised.

Figures in this article come from the ScribeStandard benchmark. See the full results, each test and how we test.