Batuhan GulerThe most accurate AI scribe for vets: our benchmark results
Lupa's AI lab benchmarked 5 AI models on real veterinary consults. Why medical AI models underperform on vet data, and why Lupa's scribe is the most accurate.
Key points
- We benchmarked 5 AI models across real veterinary consults, scored against transcripts corrected by hand by a vet.
- We weighted the scoring toward what makes a clinical note right or wrong. In the transcript: drug names, dosages, diagnoses and patient names. In the written summary: clinical accuracy, completeness, usefulness and hallucination.
- Lupa's model captured clinical details most accurately, scoring 0.83 on entity accuracy against 0.74 for a leading medical scribe, with roughly half the word error rate.
- Models built for human medicine performed worse on veterinary audio than general-purpose ones.
- A reviewing vet preferred a Lupa note in 17 of 18 consults.
I run the AI lab at Lupa. We had feedback that our AI scribe wasn't performing as well as we'd like, so we focused on making big upgrades to improve the accuracy.
That work forced a question that deserved a proper answer: which AI models are genuinely best for veterinary consults? There's no single best model. A model that excels at one job can be mediocre at another, so someone has to do the legwork for the specific job of veterinary clinical language. We did it: five models benchmarked across real veterinary consults, scored against transcripts a vet corrected by hand, measuring the things that make a clinical note right or wrong.
The model we selected for Lupa's scribe came out ahead of every alternative we tested, including the leading human-medical scribe now being sold to vets.
Every AI scribe claims to be accurate. Almost none publish what they measured, what they measured it against, or how.
Why word error rate is the wrong measure
The standard way to score a speech-to-text model is word error rate: what proportion of words did it get wrong?
For a veterinary consult, that number is close to useless. A transcript can score well on word error rate and still be dangerous. If a model transcribes ninety-eight words correctly and gets the drug name wrong, the score barely moves, but the note is now clinically wrong.
So we scored something else. We measured how accurately each model identified the details that make a clinical record right or wrong:
- Drug names
- Dosages
- Diagnoses
- Patient names
Get those right and a small mistake elsewhere in the sentence doesn't matter much. Get one of those wrong and the rest of the transcript being perfect doesn't help you.
How we ran the benchmark
We sampled real consults from Lupa, deliberately spread across different users and different accents rather than picking clean audio. Practices are noisy, vets have accents, and a benchmark on studio-quality recordings tells you nothing about a Tuesday afternoon in a busy consult room.
We built a ground truth by hand. A vet corrected the transcript of every consult manually. Those became the reference each model was scored against.
We tested five models, including general-purpose speech-to-text models and models specialized for medical audio.
We then tested the summarization step separately. An AI scribe does two jobs, not one. First it turns speech into a transcript. Then it turns that transcript into a structured clinical summary. A model can be strong at one and weak at the other, so we evaluated them separately.
For summarization, practicing vets compared outputs blind and scored them on five criteria:
- Conciseness
- Clinical usefulness
- Hallucination
- Completeness
- Clinical accuracy
Here's the breakdown, and the results:
Finding 1: medical models underperform on veterinary audio
The result we didn't expect. The speech-to-text model trained specifically on medical data performed worse on our veterinary consults than the general-purpose model from the same provider: 0.79 on entity accuracy against 0.83, and a higher word error rate.

On plain transcription accuracy the difference is wider still. Lupa's model misread 22.1% of words. The leading medical scribe misread 39.1%, close to twice as many.

Our read is that they've been tuned heavily on human medical data. Human medicine and veterinary medicine share plenty of vocabulary. They diverge on drug names, dosing conventions, species-specific terms and the shape of a consult. A model optimized hard for one is not automatically better at the other, and can be actively worse.
This is the single biggest reason the model we selected for Lupa's scribe outperforms the human-medical tools now being sold into veterinary practice. We didn't build these models; we tested them and chose the one that performs best on veterinary consults. The medical tools took the opposite path: built for human healthcare and brought across, with the specialization that helps them in human medicine working against them on a veterinary consult.
If you're evaluating any AI scribe, "it's built for medicine" is not the same claim as "it's accurate on veterinary consults."
Finding 2: Lupa's notes were the most complete
Writing the transcript is only half the job. The second half is turning it into a structured clinical note, so we scored that separately, consult by consult, in both orders to cancel out any bias from which note was shown first.

Completeness is where Lupa was strongest: its note was judged better on 67% of consults and worse on 9%.
Hallucination, meaning content appearing in a note that was never said, mostly came out level. Both products tie on 58% of consults. On fabrication, a leading medical scribe is roughly as safe as we are, and both are better than most people assume.
Hallucination is still the failure mode that should worry you most when choosing any scribe. A summary that's slightly too long wastes thirty seconds. A summary that invents a clinical finding goes into the patient record, gets relied on by the next vet who opens that file, and becomes a legal problem if anyone ever examines it.
Test any scribe for it specifically. Read the summary against what was actually said, on your own consults, and count what appears that shouldn't.
Finding 3: Lupa's notes match a vet's own note more closely than two vets match each other
The hardest test we ran. Two vets wrote their own note for a set of consults, by hand, the way they normally would. We then measured how much of each vet's note Lupa's version captured.

Against the first vet's notes, Lupa captured 94% of the clinical facts and 91% of the numbers. Against the second, 85% of both.
The line to look at is the dashed one. The two vets only agreed with each other 84% of the time. Two experienced clinicians writing up the same consult don't produce the same note, which is the realistic ceiling for this task. Lupa's notes sat at or above that line for both vets.
Finding 4: Vets are editing Lupa's notes less and less
Benchmarks are run in controlled conditions. The more interesting number comes from what vets do with the notes once they're in front of them.
Every note Lupa produces can be edited before it goes into the record, and we measure how much of it gets changed. In mid-July the median vet was rewriting 12% of the note. By early August that had fallen to 4.4%.

That drop follows a run of improvements we made to the scribe over those weeks, and it's the clearest evidence we have that they worked. Vets are correcting Lupa's notes 3 times less than they were, which is a huge timesaver across a week of consults.
What a vet made of the notes
Scores are one thing. What a vet would actually sign off is another, so a practicing vet read the notes for 18 consults and picked the best one each time, without knowing which system produced which.

A Lupa note was picked in 17 of the 18. Of the notes the vet chose, 14 of 15 needed no more than minor edits before they'd be happy to use them.
One vet on 18 consults is a small sample, and we'd rather say that than dress it up. It points the same way as everything else here, and we're expanding it.
Why Lupa's scribe is more accurate
We have labeled veterinary data that nobody else has. Practicing vets have gone through model outputs on real consults and corrected them, for both the transcripts and the written summaries, which gives us a scored dataset specific to veterinary work at each step. That's what lets us tell whether a new model is genuinely better or just newer. It also tests how we use each model. Give a model a list of veterinary drug names before it starts listening, for example, and it gets far more of them right.
We re-test every three months. New frontier models arrive at roughly that pace, and the best model for this job changes. A vendor that picked a model two years ago and never revisited it is running old technology no matter how good the marketing is. Our next accuracy review is scheduled for September.
A benchmark is a snapshot. For what that accuracy looks like after months of daily use in real practices, rather than a one-off test, see AI scribe accuracy in practice.
More and more vets are using Lupa's AI scribe
The upgrades changed how vets behave. Over the 4 weeks after they went live, the share of all appointments across Lupa practices written up with the scribe doubled, from 4.8% to 10%.

The people already using it are using it harder. Among active scribe users, the share of their own consults they choose to scribe rose from 22% to 32% over the same period.
And new people keep trying it. Every week, a fresh group of vets and nurses record their first consult, and most of the week-on-week growth comes from people who tried it and kept going.
Voice is becoming the default way to write up an appointment in Lupa practices. Adoption curves like this don't come from a feature being available. They come from it working.
What this means for your practice
Voice is becoming the normal way to write up an appointment. The vets using it aren't typing their notes any more, and are making noticeable time savings on every consult.
Dr Michelle Mooridge at Burghley Vets saves around 3 minutes on every consult using Lupa. Across a full list, that's most of an hour back every day. How much time an AI scribe saves depends on your own caseload and documentation habits, and is worth working out for your practice specifically.
The accuracy is what makes that saving real. A scribe you have to check line by line doesn't save you anything, which is why we benchmark the way we do.
AI scribing is included in Lupa's core price rather than charged as an add-on, so accuracy isn't something you pay extra for.
Lupa's AI scribe also sits inside the practice management system rather than beside it. That means the AI is embedded in your software, which is so powerful, and something to get genuinely excited about. A standalone scribe gives you a note to paste somewhere. Lupa's runs on your live practice data, so the same recording does more than write the note. It can automatically flag charges mentioned in the consult that never made it onto the invoice. It can add the medical diagnosis to the record. It can map clinical observations against the patient's history. This is a giant leap forward for veterinarians who want to get home on time. Lupa is saving you admin and driving so much efficiency.
See it on your own consults
The only benchmark that should decide it for you is your own. Book a demo and we'll run Lupa's scribe on the kind of appointments you actually see.
Frequently asked questions
What is the most accurate AI scribe for veterinary practices?
Lupa was the most accurate of the models we tested. In our benchmarking of five models against real veterinary consults, scored on drug names, dosages, diagnoses and patient names, Lupa's scribe came out top, and its notes were judged the most complete.
How accurate are AI scribes for veterinary consults?
Accuracy differs between tools, and word error rate is a poor way to judge it. A more useful measure is how reliably the model captures drug names, dosages, diagnoses and patient names. An error in any of those makes the record wrong, however much of the surrounding text is correct. When evaluating any scribe, test it on your own consults rather than relying on a published accuracy figure.
Are medical AI scribes accurate for veterinary use?
Not necessarily. In our benchmarking, speech-to-text models specialized for medical audio performed worse on veterinary consults than general-purpose models. The likely reason is that they've been tuned heavily on human medical data. A tool built for human healthcare is not automatically suited to veterinary work.
What is hallucination in an AI scribe?
Hallucination is when the AI generates content that was never said during the consult, such as a clinical finding or a symptom that didn't occur. It's the most serious failure mode for a clinical scribe, because the invented content goes into the patient record and may be relied on by whoever reads it next. Any scribe should be tested specifically for this before it's trusted with clinical notes.
How much time does an AI scribe save a vet?
It depends on how you work and how complex your consults are. Lupa clients typically say 3 to 5 minutes per consult, which adds up to hours saved across a week. The saving comes from not writing the note after the appointment, so it scales with the number of consults you see. This guide covers what actually determines the saving for a given practice.
Does Lupa charge extra for AI scribing?
No. AI scribing is included in Lupa's core subscription rather than sold as an add-on. See Lupa's pricing.

Batuhan Guler
Batuhan Guler is Head of AI at Lupa, where he runs the AI Lab, the only dedicated AI lab in the veterinary industry. He has spent the last year working closely with vets to build AI designed specifically for veterinary workflows.
Keep reading
All articles
Online booking for veterinary practices: what clients want
Client booking preferences are shifting fast. Here is what the evidence shows clients want, and where most practice booking systems still fall short.

Automated client communications for vet practices: what good looks like
How veterinary practices can make automated reminders timely, relevant and easier for clients to act on.

Staff recruitment in veterinary practice: what distinguishes the practices that keep their team
Recruitment is hard for every UK practice right now. Here is what actually separates the practices that fill a role once from the ones that keep refilling it, from the advert through the first ninety days.