News

What Perch found when it tested AI scribes for invented drug names

· Mohammed Usama

On September 2 we ran a blind test built to catch one failure: an AI scribe writing a drug into the medical record that nobody said. Today we are publishing all of it, free for anyone to run on any scribe, ours included.

The test

The set is 70 clips. Of those, 45 are ordinary veterinary dictations (SOAP visits, discharge talks, phone callbacks) with 84 drug mentions between them. The other 25 are traps, written to bait a scribe into charting a drug nobody said: ordinary words that sound like drug brands, real drugs that are not on the test's formulary, and pets and clients whose names sound like medications. An author who never saw Perch's matching code wrote every clip, and a synthetic voice read each one aloud.

We ran the set through Perch and through a commercial veterinary AI scribe, using that product's documented upload flow on a trial account with its stock note template. This is Perch's own test: one run, synthetic voices, 84 drug mentions. We publish it as a test anyone can rerun, not as an accuracy claim about either product.

What it found

On Perch's side, zero drugs were charted that nobody said, across all 70 clips and both of the speech configurations Perch runs in production.

The other product charted one. A dictation said:

"Started on Baytril, enrofloxacin, 5 mg/kg IV once."

That is a brand name and its own generic, said together the way clinicians talk. The other product's transcript heard "batril and roflexacin," and its note charted two medications: Baytril, and Roflexacin. Roflexacin does not exist. One antibiotic order became two in what would have been a patient record.

A trap clip, a phone note about "the next guard shift," came back from the same product's speech layer as "the NexGard shift." It stayed in the visit narrative rather than the medication list, but a flea and tick brand was in the note where nobody had said one.

Two misses on the record are ours. On a separate, later test set, a trap described a diabetic cat "doing well on Bexacat." The Perch build that was live until mid September charted Zycortal, a real but wrong drug, from the garbled audio. The build live now charts nothing on that clip and leaves the word for the vet. And one configuration we evaluated while building the current matcher, and never shipped, charted ProHeart from a misheard "pretreated." This test set caught it before it reached a clinic.

The other product also did something ours did not. Its repair layer recovered atenolol from a badly garbled "8 nilal," on a clip where one of Perch's two speech configurations missed it.

Why the fix and the failure are the same thing

Raw speech recognition is weak at drug names. A repair step on top of the transcript can fix what it got wrong, and that repair is useful: it is how the other product recovered atenolol. The same step that turns a mangled name into the right drug can turn a mangled sound into a drug that sounds plausible and is not real. Making a scribe good at recovering drug names and making it prone to inventing them are the same capability, pointed in two directions.

Perch is built to stop short of that guess. If a drug name in the dictation can't be matched to a known term, Perch leaves it out of the note instead of filling in a best guess. A blank gets checked. A wrong guess gets missed.

What we are publishing

All of it: the 70 clips, the written dictations, the scoring script, and a ledger of every clip and every miss on both sides, ours included. The scorer prints raw counts, never a percentage.

The clips and the ledger are CC BY 4.0, and the scorer is MIT licensed.

Check your own scribe in twenty minutes

You do not need our test set to run the same check. Say a brand name and its generic together, the way you would for an antibiotic or an NSAID: "Started on Rimadyl, carprofen, 2 mg per kg twice daily." Check whether both halves survive and whether a third drug shows up that nobody said. Then say a sentence of ordinary words that happen to sound like drug names, with no clinical content at all, and check whether a medication appears anyway. Read the finished note against your memory of the visit, one line at a time, and treat every drug name as a claim to check rather than a fact to accept.

To see how Perch handles the same check, talk to it at perch.talcoe.com/demo. No account, nothing recorded. Questions about the test set go to mo@talcoe.com.

Mohammed Usama, cofounder, Talcoe LLC

All news · Press