RESULTS · Published
The best AI scribe for allergy and immunology
Hanah was the best AI scribe we tested on an allergic reaction consultation, and it led by more than the margin of error. Heidi was second and CliniScripts third.
The test
The consultation is a GP video appointment from the PriMock57 set: a young patient whose upper lip started swelling after a prawn laksa, which the doctor treats as suspected anaphylaxis and sends to hospital by ambulance. Six AI scribes recorded it five times each, and every note was graded blind against the 30 clinical facts said in the consultation. We rank scribes by time to a finalised note, which is the wait after pressing stop plus the time a clinician needs to review the note and fix what is wrong. The recording, transcripts, notes and grading reasons are on the test page.
Results
| Rank | Scribe | Time to a finalised note | Note errors per note | Coverage | Hallucination flags per note | Word error rate |
|---|---|---|---|---|---|---|
| 1 | Hanah | 2 min 54 s | 6.8 | 85% | 1.0 | 15.1% |
| 2 | Heidi | 4 min 44 s+1 min 50 s | 14.2+7.4 | 73%−12% | 3.4+2.4 | 14.2% |
| 3 | CliniScripts | 4 min 52 s+1 min 58 s | 12.6+5.8 | 73%−12% | 2.2+1.2 | 18.3%+3.3% |
| 4 | Freed | 5 min 21 s+2 min 27 s | 10.6+3.8 | 85% | 4.4+3.4 | 17.2%+2.1% |
A Hanah note was ready to finalise 2 minutes 54 seconds after the consultation ended, on average. It captured 85% of the facts, with 6.8 note errors and 1.0 hallucination flags per note, the fewest on both counts. Reviewing and fixing it took 66% less time than writing the note yourself.
Heidi was second at 4 minutes 44 seconds and CliniScripts third at 4 minutes 52 seconds, close to each other. Freed was fourth at 5 minutes 21 seconds. Freed matched Hanah on coverage at 85% and had the second fewest note errors at 10.6.
Heidi had the lowest word error rate on this consultation at 14.2%, with Hanah at 15.1%. CliniScripts had the fewest mis-heard words that change clinical meaning, at 2.4 per run.
Findings
- The patient usually eats the vegetarian soup at this restaurant, and the prawn laksa was new. Only one of the 30 notes from the six scribes recorded this in full.
- The patient said they felt a little dizzy during the call. Most scribes heard "dizzy" as "busy" in at least some runs, and many notes left the symptom out.
- Freed recorded the patient as living alone in four of five runs. Nothing in the consultation says that.
- The patient uses no other inhalers. Hanah recorded this in every run, and every other scribe missed it in at least three runs.
Figures in this article come from the ScribeStandard benchmark. See the full results, each test and how we test.
Explore results