RESULTS · Published
The best AI scribe for myotherapy
Hanah was the best AI scribe we tested on a 50-minute myotherapy consultation, and it led by more than the margin of error. Heidi was second and CliniScripts third.
The test
The consultation is a long treatment session for a flare of long-standing low back pain and a sore right shoulder, ending with a home stretching program and lifting advice. It is the benchmark's stress test, with a television and voices from another room playing in the background. Six AI scribes recorded it five times each, and every note was graded blind against the 40 clinical facts said in the consultation. We rank scribes by time to a finalised note, which is the wait after pressing stop plus the time a clinician needs to review the note and fix what is wrong. The recording, transcripts, notes and grading reasons are on the test page.
Results
| Rank | Scribe | Time to a finalised note | Note errors per note | Coverage | Hallucination flags per note | Word error rate |
|---|---|---|---|---|---|---|
| 1 | Hanah | 2 min 45 s | 5.2 | 92% | 1.0 | 1.3% |
| 2 | Heidi | 3 min 41 s+56 s | 9.6+4.4 | 82%−10% | 0.8 | 3.9%+2.6% |
| 3 | CliniScripts | 4 min 16 s+1 min 31 s | 9.0+3.8 | 88%−4% | 2.2+1.2 | 1.1% |
| 4 | Freed | 4 min 40 s+1 min 54 s | 6.4+1.2 | 86%−5% | 0.8 | 66.8%+65.6% |
A Hanah note was ready to finalise 2 minutes 45 seconds after the consultation ended, on average, and its five runs were within 7 seconds of each other. It captured 92% of the facts, with 5.2 note errors per note, the fewest of any scribe. Reviewing and fixing it took 71% less time than writing the note yourself.
Heidi was second at 3 minutes 41 seconds and CliniScripts third at 4 minutes 16 seconds. CliniScripts had the second highest coverage at 88%. Freed was fourth at 4 minutes 40 seconds and had the second fewest note errors at 6.4.
CliniScripts and Hanah had the lowest word error rates, 1.1% and 1.3%, a difference within the margin of error. Neither had a mis-heard word that changes clinical meaning in any run.
Findings
- Hanah used 13.5 MB of data over the 50 minutes. Heidi used 145 MB and CliniScripts 132 MB.
- Freed's transcripts included long passages that were not part of the consultation, which took its word error rate to 66.8%.
- Hanah and CliniScripts had no mis-heard words that change clinical meaning in any run, despite the background television.
- Hanah's one hallucination flag per note was the same in every run: it labelled the back pain as a "mechanical-pattern" flare, which the therapist did not say.
Figures in this article come from the ScribeStandard benchmark. See the full results, each test and how we test.
Explore results