RESULTS · Published
The best AI scribe for cardiology
Hanah was the best AI scribe we tested on both cardiac consultations in our benchmark. It led each one by more than the margin of error.
The tests
Both consultations are GP video appointments from the PriMock57 set. In the first, a patient calls with a couple of hours of pressure-like chest discomfort, nausea, sweating and breathlessness, and the doctor calls an ambulance for a possible heart attack. In the second, a patient with known heart failure has breathlessness that has worsened over two weeks, and the doctor arranges bloods and an echocardiogram.
Six AI scribes recorded each consultation five times, and every note was graded blind against the clinical facts said in the consultation: 33 for the chest pain consultation and 53 for heart failure. We rank scribes by time to a finalised note, which is the wait after pressing stop plus the time a clinician needs to review the note and fix what is wrong. The recordings, transcripts, notes and grading reasons are on the chest pain and heart failure test pages.
Chest pain
| Rank | Scribe | Time to a finalised note | Note errors per note | Coverage | Hallucination flags per note | Word error rate |
|---|---|---|---|---|---|---|
| 1 | Hanah | 2 min 09 s | 2.2 | 98% | 0.8 | 17.2% |
| 2 | CliniScripts | 2 min 58 s+49 s | 5.8+3.6 | 86%−11% | 0.2 | 16.9% |
| 3 | Heidi | 3 min 25 s+1 min 16 s | 9.4+7.2 | 77%−21% | 0.6 | 17.2% |
| 4 | Lyrebird | 3 min 32 s+1 min 23 s | 8.8+6.6 | 84%−13% | 2.8+2.0 | 22.8%+5.6% |
A Hanah note was ready to finalise 2 minutes 9 seconds after the consultation ended, on average. It captured 98% of the facts, with 2.2 note errors per note. Reviewing and fixing it took 78% less time than writing the note yourself. Hanah's slowest run was quicker than every other scribe's average.
CliniScripts was second at 2 minutes 58 seconds, with 86% coverage and 5.8 note errors. CliniScripts had the fewest hallucination flags on this consultation at 0.2 per note, and the lowest word error rate at 16.9%, with Hanah and Heidi level at 17.2%. Heidi was third at 3 minutes 25 seconds and Lyrebird fourth at 3 minutes 32 seconds.
Heart failure
| Rank | Scribe | Time to a finalised note | Note errors per note | Coverage | Hallucination flags per note | Word error rate |
|---|---|---|---|---|---|---|
| 1 | Hanah | 4 min 12 s | 9.8 | 88% | 1.2 | 16.3% |
| 2 | Heidi | 6 min 20 s+2 min 08 s | 18.6+8.8 | 75%−12% | 2.0+0.8 | 15.3% |
| 3 | Freed | 6 min 59 s+2 min 47 s | 15.4+5.6 | 81%−7% | 3.6+2.4 | 25.3%+9.0% |
| 4 | Lyrebird | 7 min 05 s+2 min 53 s | 22.2+12.4 | 76%−12% | 5.2+4.0 | 25.7%+9.3% |
| 5 | CliniScripts | 7 min 11 s+2 min 59 s | 19.6+9.8 | 73%−15% | 2.8+1.6 | 15.4% |
This consultation has the most facts of any in the benchmark. A Hanah note was ready to finalise 4 minutes 12 seconds after the consultation ended, on average, and took 3 minutes 46 seconds to review. It captured 88% of the facts, with 9.8 note errors and 1.2 hallucination flags per note, the fewest on both counts.
Heidi was second at 6 minutes 20 seconds, Freed third at 6 minutes 59 seconds and Lyrebird fourth at 7 minutes 5 seconds. Freed had the second fewest note errors at 15.4. Heidi had the lowest word error rate at 15.3%, with CliniScripts at 15.4% and Hanah at 16.3%. On mis-heard words that change clinical meaning, Hanah and CliniScripts were level at 1.2 per run.
Findings
- In the chest pain consultation, the patient said they had no allergies other than aspirin. Hanah recorded this in every run, and every other scribe missed it in every run.
- The patient said aspirin makes them swell up. Heidi and Lyrebird transcribed "swell up" as "swallow" in all five runs. Hanah and CliniScripts heard it correctly every time.
- In the heart failure consultation, the patient takes one furosemide tablet a day but does not know the dose. Hanah recorded this in every run, and no other scribe recorded it in full in any run.
- Heidi transcribed furosemide as "fruit" in all five heart failure runs. Lyrebird heard bisoprolol as "omeprazole" in four of five.
Figures in this article come from the ScribeStandard benchmark. See the full results, each test and how we test.
Explore results