METHODS / VERSIONED, NOT MYSTERIOUS
Evaluation methods and limitations
Every number needs a definition. Every comparison needs the same inputs. Here is what this edition measures, and what it cannot establish.
The actual product experience
We replay the same audio through visible Chromium browsers controlled by Playwright. Each capture has an isolated copy of a signed-in profile. Audio is prepared with the shared harness as mono PCM at 44.1 kHz, with one second of leading silence and ten seconds of trailing silence. This is microphone input to the product; it does not prescribe the codec the product sends to its servers.
The harness records visible transcript changes, final transcripts, generated notes, UI events and whole-session network totals. It does not record video. Products that require a language selection are configured before recording; unsupported combinations are excluded.
Five runs, not one
Scribes vary from one run to the next, so one capture is not a result. Every scribe ran the full suite of 11 consultations at least five times. Heidi ran six: its first run was its weakest on notes, which prompted us to rerun the test, and all six runs are included. Every figure is the average across runs, and every cell opens a chart with one dot per run.
Lyrebird and Preve do not support every language case and are scored on the cases they support. Two captures failed inside the product (one PatientNotes, one Preve); those runs are averaged over the cases that completed. Four Lyrebird notes were recaptured after our harness saved them before they had finished streaming.
Final transcripts, not live drafts
Word error rate (WER) is substitutions plus deletions plus insertions, divided by reference words, pooled across every case in a run. Lower is better. We wait for the product's finalisation controls and transcript stability before capturing the text.
Equivalent numbers such as "ten" and "10" count as the same word, and the case pages don't highlight them. Contractions remain literal differences. Hindi uses a documented script-equivalence policy.
WER treats every word alike, so we also count significant transcript errors. Each transcript is compared with the reference, differences that are only fillers or function words are dropped, and a blind rater judges every remaining difference: could it change diagnosis, treatment, follow-up, safety or the record? Wrong drugs, doses, numbers, sides and flipped negations usually are. Spelling, paraphrase and dictation commands written out are not. An identical difference gets the same rating wherever it appears.
Physician dictation
The two dictation cases come with the publisher's formatted report, not a word-for-word transcript. The report adds headings and rewords phrases the doctor never said. We built a verbatim reference for each by changing the report only where the scribes agreed on what was actually said. Spoken commands such as "period", "next paragraph" and "end of dictation" are removed from every transcript before scoring, so writing them out and acting on them score the same.
Reference corrections
One reference transcript was wrong. In the day 3 food-allergy consultation the dataset marked the dish the patient names as unintelligible; on listening she says "laksa". We corrected the reference and applied it to every scribe: WER and transcript errors for that case are scored against the corrected text, and hallucination flags that only claimed the dish was laksa no longer count. Hanah, CliniScripts and Freed heard "laksa" every time; the other scribes mostly did not.
Grading the notes blind
Every note was graded by a jury of AI reviewers (Claude subagents) that could not see which product wrote it. Notes for the same consultation were shuffled into packets labelled A to G, with identical context for every note.
Each consultation has a fixed checklist of clinical facts. The jury marks each fact full, partial or missing, and flags any claim the consultation doesn't support or contradicts as a hallucination flag, with its reasoning. A wrong fact that is already counted as missing or partial is not counted again as a flag.
A separate rater then judges each missed fact, partial fact and flag for clinical significance, using the same test as the transcripts. Significance is a judgement call, and raters can disagree on borderline cases, so hallucination flags are also reported as a flat count with no severity weighting. These are AI reviews. They are not independent clinician validation and do not establish clinical safety.
Review time
Review time estimates how long a clinician would take to make each note safe to sign:
- reading the whole note at 190 words a minute;
- 5 seconds to recall each missing or partial fact, then typing it at 40 words a minute (half of it for a partial fact);
- 12 seconds to check and delete each hallucination flag. A flag that gets a checklist fact wrong also costs the 12 seconds, on top of retyping the fact.
Formatting fixes are not counted, because every scribe follows formatting instructions well. Waiting for the note is not counted either, because clinicians usually review notes later in the day. Time to a signable note, below, counts the wait.
Time saved
Time saved compares review time with writing the same note yourself. Writing it yourself is estimated from the consultation's checklist of clinical facts:
- 5 seconds to recall each fact;
- typing each fact at 40 words a minute.
Time saved is that estimate minus the note's review time. It is the same baseline for every scribe on a consultation, so a scribe saves more time only by needing less review.
Time to a signable note
Time to a signable note is the wait from pressing stop until the graded note is ready, plus its review time. It is the case where a clinician reviews each note straight after the consultation. If the gap between patients is longer than this, the note is signed before the next patient comes in. For scribes graded on their default note, the wait is timed to the default note. Like stop to prompted note, it uses production runs only.
Notes ready to sign as written
A note is ready to sign as written when the jury marked every checklist fact as fully captured and found no hallucination flags, so nothing needs editing. Formatting is not counted, as in review time. Each scribe's figure is the share of its notes that met this.
Editing effort
Editing effort shows how much of the note the clinician still has to fix, whatever its length. 0% means no editing required. 100% means writing the full note by hand.
Each fix is weighted by the work it takes:
- partial fact, 1: it only needs amending;
- missing fact, 2: it has to be recalled and written;
- hallucination flag, 3: it has to be spotted, checked and deleted;
- fact captured wrongly, 2 more: the wrong version has to be spotted and checked before it is fixed.
A note's weights are added up and divided by the score for writing it yourself, which is 2 for every fact on the consultation's checklist. Hallucinations can push a note above 100%.
Face-off
The face-off compares two scribes on every measure. Under the score, four takeaways each use a single measure, so no gap is counted twice: word error rate (transcription), hallucination flags, significant note errors (note accuracy) and time to a signable note. Each gap is how much less of it the better scribe has: (worse − better) ÷ worse.
"Weighted in a scribe's favour" is the plain average of those four gaps, equally weighted, with a tie counted as 0%. It is our own summary, not an established index: the choice of four measures and equal weights is ours, and a large gap counts for at most 100%. The takeaways themselves are the figures to quote.
The face-off opens on the two leading scribes by this weighting: every scribe is compared with every other, the two with the most head-to-head wins are shown, and the summed margin breaks a tie.
Notes and timing
Every scribe that takes a note instruction was given this one, through its own UI:
Write SOAP notes using the headings Subjective, Objective, Assessment, and Plan. Use bullet points in every section and concise clinical shorthand.
Freed, CliniScripts and Preve don't take an instruction this way, so they are graded on their default note. "Stop to prompted note" still times the prompted note where one exists.
Stop to prompted note is the median, per run, of the time from pressing stop to the requested note being ready. Completion gates include short UI stability waits, and captures share a machine and internet connection, so this is not an isolated server-latency test.
The whole bandwidth bill
Chromium NetLog totals include app loading, recording, finalisation, notes and browser shutdown, including WebSocket transport. Socket and UDP byte events are counted once; decrypted HTTP and WebSocket payloads are not added again. This is not ISP-billed traffic: packet headers and TCP retransmissions are not included. Missing or invalid logs are unavailable, never zero.
This edition
This edition covers AU production captures from 19 to 22 September 2026: seven scribes, 11 consultations, and at least five full runs each. Future changes will produce a new version with their cohort, capture dates, settings and review policy kept.
Reproducibility
Download everything as one JSON file: every run's scores, transcript, error marks, note, checklist verdicts and jury reasoning for all 11 consultations. The scores alone are also a CSV. Source recordings can be played or downloaded from each consultation page. Authentication profiles and raw network logs are not published.