CLINICAL TECHNOLOGY REVIEWEDITION 01 / SEPTEMBER 2026
Explore results →
← Articles & methodology

RESULTS · Published

The best AI scribe for allergy and immunology

Hanah was the best AI scribe we tested on an allergic reaction consultation, and it led by more than the margin of error. Heidi was second and CliniScripts third.

The test

The consultation is a GP video appointment from the PriMock57 set: a young patient whose upper lip started swelling after a prawn laksa, which the doctor treats as suspected anaphylaxis and sends to hospital by ambulance. Six AI scribes recorded it five times each, and every note was graded blind against the 30 clinical facts said in the consultation. We rank scribes by time to a finalised note, which is the wait after pressing stop plus the time a clinician needs to review the note and fix what is wrong. The recording, transcripts, notes and grading reasons are on the test page.

Results

RankScribeTime to a finalised noteNote errors per noteCoverageHallucination flags per noteWord error rate
1Hanah2 min 54 s6.885%1.015.1%
2Heidi4 min 44 s+1 min 50 s14.2+7.473%−12%3.4+2.414.2%
3CliniScripts4 min 52 s+1 min 58 s12.6+5.873%−12%2.2+1.218.3%+3.3%
4Freed5 min 21 s+2 min 27 s10.6+3.885%4.4+3.417.2%+2.1%

A Hanah note was ready to finalise 2 minutes 54 seconds after the consultation ended, on average. It captured 85% of the facts, with 6.8 note errors and 1.0 hallucination flags per note, the fewest on both counts. Reviewing and fixing it took 66% less time than writing the note yourself.

Heidi was second at 4 minutes 44 seconds and CliniScripts third at 4 minutes 52 seconds, close to each other. Freed was fourth at 5 minutes 21 seconds. Freed matched Hanah on coverage at 85% and had the second fewest note errors at 10.6.

Heidi had the lowest word error rate on this consultation at 14.2%, with Hanah at 15.1%. CliniScripts had the fewest mis-heard words that change clinical meaning, at 2.4 per run.

Findings

Figures in this article come from the ScribeStandard benchmark. See the full results, each test and how we test.