CLINICAL TECHNOLOGY REVIEWEDITION 01 / SEPTEMBER 2026
Explore results →
← Articles & methodology

RESULTS · Published

The best AI scribe for French

Hanah was the best AI scribe we tested on a consultation that switches between French and English. Every Hanah note captured every fact with no hallucination flags, and it was the only scribe with a note ready to finalise as written. Heidi was second and CliniScripts third.

The test

The consultation switches between French and English throughout. A patient with four days of abdominal pain that started in the upper abdomen and moved to the right lower side is assessed for appendicitis and admitted for surgical review. Five AI scribes recorded it five times each, and every note was graded blind against the 32 clinical facts said in the consultation. Lyrebird does not support this case. We rank scribes by time to a finalised note, which is the wait after pressing stop plus the time a clinician needs to review the note and fix what is wrong. The recording, transcripts, notes and grading reasons are on the test page.

Results

RankScribeTime to a finalised noteNote errors per noteCoverageHallucination flags per noteWord error rate
1Hanah1 min 09 s0.0100%0.07.3%
2Heidi1 min 45 s+35 s2.4+2.494%−6%0.4+0.413.1%+5.9%
3CliniScripts1 min 48 s+39 s2.2+2.296%−4%0.6+0.66.7%

A Hanah note was ready to finalise 1 minute 9 seconds after the consultation ended, on average. In all five runs it captured 100% of the facts, with no note errors and no hallucination flags. Every Hanah note was ready to finalise as written, and its review time is reading time alone.

Heidi was second at 1 minute 45 seconds and CliniScripts third at 1 minute 48 seconds, close to each other. CliniScripts had 2.2 note errors per note and Heidi 2.4.

CliniScripts had the lowest word error rate on this consultation at 6.7%, with Hanah at 7.3%, a difference within the margin of error. Heidi and CliniScripts had no mis-heard words that change clinical meaning; Hanah had one across its five runs.

Findings

Figures in this article come from the ScribeStandard benchmark. See the full results, each test and how we test.