CLINICAL TECHNOLOGY REVIEWEDITION 01 / SEPTEMBER 2026
ScribeStandard.Explore results

SCRIBE RESULTS / 22 SEPTEMBER 2026

Hanah

Averages across 5 graded runs of 11 consultations, the same audio and instruction for every scribe.

1m 28sTime to a signable note · ranked 1 of 7
45%Notes ready to sign as written · ranked 1 of 7
2.4Significant note errors per run · ranked 1 of 7
1.6Hallucination flags per run · ranked 1 of 7
12.7%Word error rate · ranked 1 of 7

Summary

Hanah ranked 1 of 7 for time to a signable note at 1m 28s. 45% of its notes were ready to sign as written, ranked 1 of 7, and its clinical coverage was 97%, also ranked 1 of 7.

Hanah averaged 1.6 hallucination flags per run, ranked joint 1 of 7, with a range of 1.0 to 2.0 between runs. Significant note errors averaged 2.4 per run and ranged from 1.0 to 4.0 between runs.

Every measure

MeasureAverageRun rangeRank
Time to a signable noteWait from pressing stop to the graded note, plus review time: how long until the note can be signed if it is reviewed straight away.1m 28s1m 24s–1m 33s1 of 7
Notes ready to sign as writtenShare of notes that captured every checklist fact in full with no hallucination flags, so nothing needed editing.45%36%–55%1 of 7
Editing effort0% means no editing required; 100% means writing the full note by hand. Fixes are weighted by the work they take: partial fact 1, missing fact 2, hallucination flag 3, and +2 when a fact is captured wrongly.6%5%–7%1 of 7
Review time per noteEstimated clinician time to proofread and fix one note (190 wpm reading, 5 s to recall and 40 wpm to type each missing or partial fact, 12 s to verify and delete each hallucination flag, including a wrong fact before it is retyped).1m 11s1m 9s–1m 12s1 of 7
Worst note, any runThe single slowest note to review and fix across every case and run.1m 57s1 of 7
Time saved per noteEstimated time to write the note yourself (5 s to recall each checklist fact, typing it at 40 wpm) minus review time.3m 15s3m 14s–3m 17s1 of 7
Significant note errors per runMissed facts, partial facts and hallucinations that a blind AI rater judged could change diagnosis, treatment, follow-up, safety or the legal record, across 11 cases. Significance is a judgement call; the flat count is in Hallucination flags.2.41.0–4.01 of 7
Hallucination flags per runEvery claim in the note with no support in the consultation, counted flat: each one counts, however minor. No severity judgement. Across 11 cases.1.61.0–2.01 of 7
Edit actions per noteMissing items, partial items and hallucination flags a clinician would need to fix, per note.0.730.55–0.911 of 7
Clinical coverageShare of each case checklist captured by the note (partial counts as half).97%96%–97%1 of 7
Significant transcript errors per runMis-heard words that change clinical meaning, rated blind, across 11 cases.7.87.0–9.01 of 7
Word error ratePooled word error rate per run against the human reference. Every word counts equally, including fillers.12.7%12.4%–13.3%1 of 7
Stop to prompted noteMedian time from pressing stop to the note written to our instruction, per run.17s14s–22s1 of 6
Data per consultAverage network data per consultation, per run.8.3 MB8.3 MB–8.5 MB1 of 7

Definitions and weights are in the methodology. Compare Hanah head to head ↗

Consultations