CLINICAL TECHNOLOGY REVIEWEDITION 01 / SEPTEMBER 2026
Explore results →
← Articles & methodology

RESULTS · Published

The best AI scribe for dermatology

Hanah was the best AI scribe we tested on an eczema consultation. Heidi came second, within the margin of error on time, and Lyrebird was third.

The test

The consultation is a GP video appointment from the PriMock57 set: four days of an itchy, sore red rash on the chest, hands and inner elbows that was keeping the patient awake at night. Six AI scribes recorded it five times each, and every note was graded blind against the 39 clinical facts said in the consultation. We rank scribes by time to a finalised note, which is the wait after pressing stop plus the time a clinician needs to review the note and fix what is wrong. The recording, transcripts, notes and grading reasons are on the test page.

Results

RankScribeTime to a finalised noteNote errors per noteCoverageHallucination flags per noteWord error rate
1Hanah2 min 46 s4.693%1.014.3%
2Heidi3 min 17 s+32 s8.0+3.484%−9%0.416.9%+2.5%
3Lyrebird4 min 02 s+1 min 17 s12.0+7.486%−7%3.8+2.818.7%+4.4%
4Freed4 min 44 s+1 min 58 s6.6+2.089%−4%0.620.3%+5.9%
5CliniScripts4 min 58 s+2 min 12 s12.8+8.274%−18%1.014.9%+0.6%

A Hanah note was ready to finalise 2 minutes 46 seconds after the consultation ended, on average. It captured 93% of the facts, with 4.6 note errors and 1.0 hallucination flags per note. Reviewing and fixing it took 71% less time than writing the note yourself.

Heidi averaged 3 minutes 17 seconds. Its quickest run, at 2 minutes 41 seconds, was faster than Hanah's slowest at 2 minutes 59 seconds, which puts the two within the margin of error on time. Hanah was quicker on average, captured more of the consultation (93% against 84%) and had fewer note errors (4.6 against 8.0). Heidi had the fewest hallucination flags on this consultation at 0.4 per note.

Lyrebird was third at 4 minutes 2 seconds and Freed fourth at 4 minutes 44 seconds. Freed had the second fewest note errors at 6.6.

Hanah also had the lowest word error rate on this consultation at 14.3%, with CliniScripts close at 14.9%.

Findings

Figures in this article come from the ScribeStandard benchmark. See the full results, each test and how we test.