RESULTS · Published
The best AI scribe for pain management
Hanah was the best AI scribe we tested on a pain management dictation. The leading scribes were close on accuracy, and several differences here are within the margin of error, but Hanah had the quickest note to finalise and captured every fact in every run.
The test
The test is a short dictated pain clinic summary for a 55-year-old woman with low back pain and left leg numbness, an MRI showing foraminal stenosis at L4-5, and a plan for a lumbar epidural steroid injection. Six AI scribes recorded it five times each, and every note was graded blind against the 12 clinical facts in the dictation. We rank scribes by time to a finalised note, which is the wait after pressing stop plus the time a clinician needs to review the note and fix what is wrong. The reference transcript has not yet been verified word for word against the audio, and its word error rates are provisional. The recording, transcripts, notes and grading reasons are on the test page.
Results
| Rank | Scribe | Time to a finalised note | Note errors per note | Coverage | Hallucination flags per note | Word error rate (provisional) |
|---|---|---|---|---|---|---|
| 1 | Hanah | 42 s | 1.0 | 100% | 1.0 | 7.9% |
| 2 | Lyrebird | 54 s+12 s | 1.4+0.4 | 92%−8% | 0.4 | 7.0% |
| 3 | CliniScripts | 59 s+18 s | 1.4+0.4 | 92%−8% | 0.4 | 6.7% |
| 5 | PatientNotes | 1 min 02 s+20 s | 0.8 | 99%−1% | 0.6 | 12.3%+4.4% |
A Hanah note was ready to finalise 42 seconds after the dictation ended, on average, and took 33 seconds to review. It captured 100% of the facts in all five runs. Each Hanah note carried one hallucination flag, which gave it 1.0 note errors per note.
Lyrebird was second at 54 seconds and CliniScripts third at 59 seconds. Both had 1.4 note errors per note and 92% coverage. The quickest CliniScripts run, at 39 seconds, beat Hanah's average, and one CliniScripts note in five needed no editing at all.
PatientNotes had the fewest note errors on this dictation at 0.8 per note, and two of its five notes needed no editing. The difference from Hanah's 1.0 is within the margin of error, and PatientNotes took 1 minute 2 seconds to a finalised note.
CliniScripts had the lowest provisional word error rate at 6.7%, with Lyrebird at 7.0% and Hanah at 7.9%.
Findings
- The dictation gives the patient's ethnicity. Hanah and PatientNotes recorded it in every run, CliniScripts in one of five and Lyrebird in none.
- Hanah's one hallucination flag per note was the same in every run: it linked the MRI finding to the left leg symptoms, a connection the dictation does not make.
- Hanah, CliniScripts and Lyrebird had no mis-heard words that change clinical meaning. Other scribes garbled "foraminal stenosis" and "extremity".
- Hanah used 2.5 MB of data for the dictation, the least of any scribe.
Figures in this article come from the ScribeStandard benchmark. See the full results, each test and how we test.
Explore results