CLINICAL TECHNOLOGY REVIEWEDITION 01 / SEPTEMBER 2026
Explore results →
← Articles & methodology

RESULTS · Published

The best AI scribe for cardiology

Hanah was the best AI scribe we tested on both cardiac consultations in our benchmark. It led each one by more than the margin of error.

The tests

Both consultations are GP video appointments from the PriMock57 set. In the first, a patient calls with a couple of hours of pressure-like chest discomfort, nausea, sweating and breathlessness, and the doctor calls an ambulance for a possible heart attack. In the second, a patient with known heart failure has breathlessness that has worsened over two weeks, and the doctor arranges bloods and an echocardiogram.

Six AI scribes recorded each consultation five times, and every note was graded blind against the clinical facts said in the consultation: 33 for the chest pain consultation and 53 for heart failure. We rank scribes by time to a finalised note, which is the wait after pressing stop plus the time a clinician needs to review the note and fix what is wrong. The recordings, transcripts, notes and grading reasons are on the chest pain and heart failure test pages.

Chest pain

RankScribeTime to a finalised noteNote errors per noteCoverageHallucination flags per noteWord error rate
1Hanah2 min 09 s2.298%0.817.2%
2CliniScripts2 min 58 s+49 s5.8+3.686%−11%0.216.9%
3Heidi3 min 25 s+1 min 16 s9.4+7.277%−21%0.617.2%
4Lyrebird3 min 32 s+1 min 23 s8.8+6.684%−13%2.8+2.022.8%+5.6%

A Hanah note was ready to finalise 2 minutes 9 seconds after the consultation ended, on average. It captured 98% of the facts, with 2.2 note errors per note. Reviewing and fixing it took 78% less time than writing the note yourself. Hanah's slowest run was quicker than every other scribe's average.

CliniScripts was second at 2 minutes 58 seconds, with 86% coverage and 5.8 note errors. CliniScripts had the fewest hallucination flags on this consultation at 0.2 per note, and the lowest word error rate at 16.9%, with Hanah and Heidi level at 17.2%. Heidi was third at 3 minutes 25 seconds and Lyrebird fourth at 3 minutes 32 seconds.

Heart failure

RankScribeTime to a finalised noteNote errors per noteCoverageHallucination flags per noteWord error rate
1Hanah4 min 12 s9.888%1.216.3%
2Heidi6 min 20 s+2 min 08 s18.6+8.875%−12%2.0+0.815.3%
3Freed6 min 59 s+2 min 47 s15.4+5.681%−7%3.6+2.425.3%+9.0%
4Lyrebird7 min 05 s+2 min 53 s22.2+12.476%−12%5.2+4.025.7%+9.3%
5CliniScripts7 min 11 s+2 min 59 s19.6+9.873%−15%2.8+1.615.4%

This consultation has the most facts of any in the benchmark. A Hanah note was ready to finalise 4 minutes 12 seconds after the consultation ended, on average, and took 3 minutes 46 seconds to review. It captured 88% of the facts, with 9.8 note errors and 1.2 hallucination flags per note, the fewest on both counts.

Heidi was second at 6 minutes 20 seconds, Freed third at 6 minutes 59 seconds and Lyrebird fourth at 7 minutes 5 seconds. Freed had the second fewest note errors at 15.4. Heidi had the lowest word error rate at 15.3%, with CliniScripts at 15.4% and Hanah at 16.3%. On mis-heard words that change clinical meaning, Hanah and CliniScripts were level at 1.2 per run.

Findings

Figures in this article come from the ScribeStandard benchmark. See the full results, each test and how we test.