SCRIBE RESULTS / 22 SEPTEMBER 2026
Heidi
Averages across 6 graded runs of 11 consultations, the same audio and instruction for every scribe. Wide variation between runs, additional run included for fine tuning.
1m 52sTime to a signable note · ranked 2 of 7
11%Notes ready to sign as written · ranked 5 of 7
8.7Significant note errors per run · ranked 5 of 7
4.3Hallucination flags per run · ranked 3 of 7
13.9%Word error rate · ranked 3 of 7
Summary
Heidi ranked 2 of 7 for time to a signable note at 1m 52s and for review time at 1m 26s. Time saved per note was 3m, also ranked 2 of 7. Its word error rate was 13.9%, ranked 3 of 7.
11% of Heidi's notes were ready to sign as written, ranked 5 of 7. Significant note errors averaged 8.7 per run, ranked 5 of 7, and ranged from 2.0 to 13.0 between runs. Clinical coverage was 88%, ranked 6 of 7.
Every measure
| Measure | Average | Run range | Rank |
|---|---|---|---|
| Time to a signable noteWait from pressing stop to the graded note, plus review time: how long until the note can be signed if it is reviewed straight away. | 1m 52s | 1m 40s–1m 57s | 2 of 7 |
| Notes ready to sign as writtenShare of notes that captured every checklist fact in full with no hallucination flags, so nothing needed editing. | 11% | 9%–18% | 5 of 7 |
| Editing effort0% means no editing required; 100% means writing the full note by hand. Fixes are weighted by the work they take: partial fact 1, missing fact 2, hallucination flag 3, and +2 when a fact is captured wrongly. | 16% | 11%–18% | 4 of 7 |
| Review time per noteEstimated clinician time to proofread and fix one note (190 wpm reading, 5 s to recall and 40 wpm to type each missing or partial fact, 12 s to verify and delete each hallucination flag, including a wrong fact before it is retyped). | 1m 26s | 1m 14s–1m 32s | 2 of 7 |
| Worst note, any runThe single slowest note to review and fix across every case and run. | 2m 48s | — | 3 of 7 |
| Time saved per noteEstimated time to write the note yourself (5 s to recall each checklist fact, typing it at 40 wpm) minus review time. | 3m | 2m 54s–3m 12s | 2 of 7 |
| Significant note errors per runMissed facts, partial facts and hallucinations that a blind AI rater judged could change diagnosis, treatment, follow-up, safety or the legal record, across 11 cases. Significance is a judgement call; the flat count is in Hallucination flags. | 8.7 | 2.0–13.0 | 5 of 7 |
| Hallucination flags per runEvery claim in the note with no support in the consultation, counted flat: each one counts, however minor. No severity judgement. Across 11 cases. | 4.3 | 3.0–6.0 | 3 of 7 |
| Edit actions per noteMissing items, partial items and hallucination flags a clinician would need to fix, per note. | 2.21 | 1.55–2.82 | 4 of 7 |
| Clinical coverageShare of each case checklist captured by the note (partial counts as half). | 88% | 84%–92% | 6 of 7 |
| Significant transcript errors per runMis-heard words that change clinical meaning, rated blind, across 11 cases. | 20.3 | 16.0–24.0 | 3 of 7 |
| Word error ratePooled word error rate per run against the human reference. Every word counts equally, including fillers. | 13.9% | 13.5%–14.2% | 3 of 7 |
| Stop to prompted noteMedian time from pressing stop to the note written to our instruction, per run. | 26s | 25s–27s | 4 of 6 |
| Data per consultAverage network data per consultation, per run. | 43.3 MB | 41.0 MB–46.1 MB | 7 of 7 |
Definitions and weights are in the methodology. Compare Heidi head to head ↗
Consultations
CONSULTATION