RESULTS · Published
The best AI scribe for Hindi
Hanah was the best AI scribe we tested on a consultation held entirely in Hindi. CliniScripts was second and Heidi third. CliniScripts and Hanah were within the margin of error on note accuracy, and Hanah led on time and transcription.
The test
The consultation is a short clinic visit in Hindi: two weeks of tiredness, headaches that are worse in the evening, occasional dizziness and poor sleep from work stress, with a blood test ordered and advice on meals, water and sleep. Six AI scribes recorded it five times each, each set to Hindi where the product asks for a language, and every note was graded blind against the 17 clinical facts said in the consultation. Notes could be written in English or Hindi. We rank scribes by time to a finalised note, which is the wait after pressing stop plus the time a clinician needs to review the note and fix what is wrong. The recording, transcripts, notes and grading reasons are on the test page.
Results
| Rank | Scribe | Time to a finalised note | Note errors per note | Coverage | Hallucination flags per note | Word error rate |
|---|---|---|---|---|---|---|
| 1 | Hanah | 1 min 07 s | 2.4 | 93% | 0.0 | 2.0% |
| 2 | CliniScripts | 1 min 50 s+42 s | 2.2 | 94% | 0.6+0.6 | 3.6%+1.6% |
| 3 | Heidi | 1 min 56 s+49 s | 3.8+1.4 | 86%−7% | 0.6+0.6 | 4.3%+2.3% |
| 4 | Lyrebird | 2 min 12 s+1 min 05 s | 5.8+3.4 | 82%−11% | 1.4+1.4 | 3.5%+1.5% |
A Hanah note was ready to finalise 1 minute 7 seconds after the consultation ended, on average. It captured 93% of the facts, with 2.4 note errors per note and no hallucination flags in any run. Reviewing and fixing it took 88% less time than writing the note yourself.
CliniScripts was second at 1 minute 50 seconds. It had the fewest note errors at 2.2 per note and the highest coverage at 94%, both within the margin of error of Hanah's. Heidi was third at 1 minute 56 seconds, with 3.8 note errors per note, and Lyrebird fourth at 2 minutes 12 seconds.
Hanah transcribed Hindi most accurately, with a word error rate of 2.0%. Lyrebird was next at 3.5% and CliniScripts at 3.6%. Hanah produced its note a median 10 seconds after stop.
Findings
- The patient said they sometimes skip meals. One scribe translated this as cannabis being "sometimes skipped" and recorded cannabis use in every note.
- Hanah had no hallucination flags in any run, and it was the only scribe with none.
- The patient put their poor sleep down to work stress. CliniScripts recorded this in every run, while Hanah recorded it only in part in every run.
- Hanah used 3.3 MB of data for the consultation, the least of any scribe. Heidi used 23.3 MB.
Figures in this article come from the ScribeStandard benchmark. See the full results, each test and how we test.
Explore results