Engines transcribe by probability, so the rare words your session turns on lose to common words that sound like them. The dictionary is the lever you control.
A live caption engine has never seen your agenda. It transcribes by probability: given these sounds, what words does language usually make? For everyday speech, that works startlingly well. For the words your session actually turns on — the drug, the witness, the product, the case number — it fails in one specific, predictable way: the engine replaces words it doesn't know with common words it does. A custom dictionary is the lever that fixes this whole class of error, and it's the lever you control most directly.
Speech recognition is trained on enormous amounts of general language, so it carries strong expectations about which words follow which sounds. When audio is ambiguous — and live audio always is — the engine resolves the ambiguity toward the statistically likely option. "Xarelto" is not statistically likely. "Zarelto," or worse, "za relto," is what a general-language model reaches for when it hears an anticoagulant it never trained on.
This is why a caption feed can be excellent and useless at the same time. An engine can transcribe 49 of 50 words correctly and still miss the only word the sentence existed to deliver. Overall accuracy is measured across all words; your session's value is concentrated in a few of them. (For how accuracy is actually scored, see our word error rate explainer.) The words that concentrate that value are almost always the same kinds:
A custom dictionary tells the engine, before the session starts, that certain low-probability strings are high-probability *here*. It doesn't change how the engine hears — it changes how the engine decides. When the audio could plausibly be your term, the term wins instead of losing to a common near-homophone. The effect on the errors that matter is immediate:
| What was said | What a general engine hears | With the term preloaded |
|---|---|---|
| Xarelto | "zarelto" / "za relto" | Xarelto |
| Farxiga | "far sika" | Farxiga |
| voir dire | "vwar deer" | voir dire |
| Ms. Nguyen | "Miss when" | Ms. Nguyen |
| Kubernetes | "cooper netties" | Kubernetes |
Be clear-eyed about the limits, too. A dictionary can't rescue audio the engine never heard cleanly — a term mumbled under crosstalk stays lost. And it can't referee between two common words that are both plausible in context. It fixes the guessing problem, which happens to be the problem specialist sessions have most.
The good news: the terms an assignment will use are almost never a mystery. They're sitting in documents you already have. Ten minutes before the session is usually enough.
Agendas, slide decks, case files, pleadings, discharge summaries, product briefs. Skim for every capitalized word, every italicized term, every string you'd have to look up. Those are your entries.
Every participant, organization, and place likely to be spoken aloud. Names are the highest-frequency terminology errors in most sessions, and the most noticed — people always catch their own name misspelled.
Case numbers, statute citations, SKUs, acronyms. Decide the written form you want once — "24-CV-1187," not three competing renderings — and enter it that way.
A dictionary is for rare terms. Loading it with everyday vocabulary adds nothing and can nudge the engine toward your entries where ordinary words were right. Rare and specific beats long and thorough.
A glossary built for one client compounds for the next assignment with them. Interpreters have kept term lists forever — the dictionary is that habit, made executable. Our interpreter prep checklist folds this into a full pre-session routine.
Unicaption approaches this from both ends — a broad specialist base you don't have to build, and a personal layer you do:
The pattern that works is boring and repeatable: when the assignment lands, skim the documents, pull the names and terms, enter them, done. Five to ten minutes, once. Every session after that starts with an engine that already knows the words the day depends on — which is the closest thing to a free accuracy upgrade that live captioning has.
A custom dictionary is a list of terms you give a speech recognition engine before a session — names, drug names, product names, case numbers, jargon. It tells the engine those rare strings are expected here, so ambiguous audio resolves to your term instead of a common word that sounds similar. In Unicaption, it sits alongside built-in medical, legal, and government terminology.
Because engines transcribe by probability learned from general language. A rare word like a drug name or a surname has no strong statistical support, so the engine substitutes the nearest common word — "Xarelto" becomes "zarelto," "Nguyen" becomes "when." Preloading those terms in a custom dictionary is the direct fix; the engine stops guessing once it knows the word is expected.
Curate rather than dump: typically a few dozen terms per assignment — the names, codes, and jargon a general engine would actually mishear. Skip everyday words; they add nothing and can nudge the engine wrongly. Verify spellings against a written source, because a typo in the dictionary reproduces itself in every caption.
Yes. Unicaption includes built-in medical, legal, and government terminology, plus a per-user custom dictionary for the names and terms specific to your assignment. Captions run in real time, and session content is never stored or used to train models — relevant when your glossary comes from case files or patient-adjacent documents.
Built-in specialist terminology plus your own dictionary, live in 60+ languages. 30 free minutes every week — no credit card.
Start Free Trial →