CLINICAL TECHNOLOGY REVIEWEDITION 01 / SEPTEMBER 2026
Explore results →
← Articles & methodology

METHODOLOGY · Published

An industry standard test for AI medical scribes

ScribeStandard puts every AI medical scribe through transparent and rigorous tests designed to show how scribes work in the real world - beyond the accuracy claims in their marketing.

The consultations

The benchmark uses 12 recorded consultations. Eight are mock GP consultations from PriMock57, a public dataset of primary care consultations, covering cardiology, dermatology, gastroenterology, neurology, urology, mental health and allergy. The others are a consultation held entirely in Hindi, one in French, a short pain management history and a 50 minute myotherapy treatment session recorded with a television and voices from another room in the background. They run from 31 seconds to 50 minutes.

You can play and download each recording from its test page. Each has a reference transcript, and where a dataset's transcript did not match what was said we corrected it. All six scribes are scored against the same corrected text.

Testing the product clinicians use

We test each scribe through its own web app. A script drives a visible Chromium browser signed in to a normal account, presses record, plays the consultation into the microphone input and presses stop. Every scribe receives identical audio. Scribes that take a note instruction all get the same one, SOAP headings with bullet points and concise clinical shorthand. Freed and CliniScripts don't take an instruction this way and are graded on their default note.

Every scribe recorded every consultation it supports at least five times. Results change from run to run on identical audio, and a single run can make a product look better or worse than it is. Figures on the site are the average across runs, and each cell opens a chart with one dot per run.

Grading the notes

Each consultation has an inventory of every fact said that relates to the patient's conditions, symptoms or signs, 358 facts across the 12 consultations. This is the standard Weiner et al. (2020) used when they checked 105 doctors' notes against recordings of the visit. The inventory quotes the transcript line behind each fact. A diagnosis is on the inventory only when the clinician said it, because a scribe should not diagnose.

Notes for the same consultation are shuffled together with the product names removed and graded by a jury of AI reviewers. For each fact the jury marks the note full, partial or missing. A separate pass flags anything the note documents that the consultation does not support, such as a symptom nobody mentioned or a wrong dose, side or timing. These are hallucination flags. A missed or partial fact is a note error, and so is a hallucination flag. A wrong detail counts twice, as Weiner et al. counted it.

Clinician time

We estimate review time for each note from reading it at 190 words a minute, recalling and typing each missing fact, and checking and deleting each hallucination flag. Time to a finalised note adds the wait after pressing stop. Both are compared with writing the note yourself, a median 8.1 minutes per visit across 203,728 US doctors in Apathy et al. (2023).

We rank scribes in each test by time to a finalised note, because it is the time a clinician spends between the end of a consultation and a finished note.

Transcription, speed and data

Word error rate is measured on the final transcript each scribe keeps, against the reference transcript. It weighs every word the same, and a wrong drug or a flipped negation counts no more than a dropped "the". We also count significant transcript errors, the differences a blind rater judges could change diagnosis, treatment, follow-up or safety.

We time how long each scribe takes to produce its note after stop and measure all the data a recording session moves, from opening the app to closing the browser. A 50 minute stress test with every scribe recording at once on one machine measures memory and CPU use.

Doctors as a reference point

Results sit next to published figures for doctors writing their own notes: 5.0 errors per note and 10% of notes error-free in Weiner et al. (2020), and 8.1 minutes to write in Apathy et al. (2023). Our grading is stricter than Weiner et al.'s, which favours the doctors. When we graded six PriMock57 doctors' own notes blind, mixed in with the scribes' notes for the same consultations, they averaged 26 errors each against the same inventories.

Current results

Across all 12 consultations, note errors ranged from 4.1 to 18.0 per note and review time from 2 minutes 2 seconds to 5 minutes 11 seconds. Hanah ranked first on 18 of the 19 measures we publish. The results table has every scribe on every measure, and each specialty and language has its own article.

Limits

The grading is done by AI reviewers and has not been validated by independent clinicians. A hallucination flag counts once whatever its clinical consequence. Review time is an estimate from fixed reading and typing rates, and the doctor figures come from different studies of different consultations. The like-for-like comparison would have independent clinicians write notes for the 12 consultations and grade them blind alongside the scribes, and we have not done that yet.

We don't test user experience, because pressing record and getting a note works much the same way in every scribe we tested.

The data

The transcript, note, fact verdicts, hallucination flags and jury reasoning for every run are in one JSON file, with the scores alone in a CSV. The full method is on the methodology page. Anyone with accounts for these scribes can play the same recordings through them and check our numbers.

Figures in this article come from the ScribeStandard benchmark. See the full results, each test and how we test.