Simultaneous interpreters already work with a partner who writes down numbers. A live transcript does that job tirelessly — if it is set up with booth discipline.
Conference interpreters work in one of the most cognitively demanding jobs that exists: listening in one language while speaking in another, a few seconds behind a speaker who does not slow down. That is why booths hold two interpreters trading turns of roughly half an hour, and why the off-turn partner spends much of that rest writing numbers and names on a pad for the one on mic. The live transcript does not compete with the interpreter — it competes with the notepad, and it wins on speed, stamina, and spelling. Here is the booth workflow.
Everything about booth practice is load management: the two-interpreter team, the timed turns, the console with its relay channels, the documents requested in advance. The profession discovered long ago that ear-to-voice interpreting consumes nearly all available working memory — which is why the details with no redundancy, like figures and proper names, are the first casualties when a speaker accelerates. (The research and the coping strategies are a topic of their own — see our piece on interpreter cognitive load and burnout.)
A live transcript slots into that existing system. It does not interpret, and it does not need to: it is the boothmate's number pad, running continuously, never tiring, and never on its own turn. The boothmate stays — for terminology lookups, for judgment calls, for the turn handover — but the mechanical part of their support role gets automated.
The failure points of simultaneous interpreting are famous inside the profession and invisible outside it:
Every one of these is a text problem more than a language problem. Which is precisely why a text layer helps.
Run Unicaption in a browser on a laptop and give it the cleanest audio available: the booth console's floor feed via system audio, or the event's stream in another tab. Never a microphone pointed at the booth speaker — the tool should hear what you hear in your headphones.
Speaker names, delegation names, product and org terms from the program go into the custom dictionary before the first session — the same preparation discipline as a glossary, one more artifact of it. (It belongs on your interpreter prep checklist.)
The transcript sits in peripheral vision for verification — a number checked, a spelling confirmed, a half-heard name recovered. Interpreters who try to read along while rendering split attention and do both jobs worse. The discipline is the same as with the boothmate's pad: use it when you need it.
The resting interpreter watches the transcript and the hall, flagging upcoming terminology and writing the judgment calls the machine can't make. The pad survives; it just carries less.
In relay, an interpreter works not from the floor but from a colleague's output — the Estonian speech reaches the Portuguese booth through the English pivot. Relay is unavoidable in large multilingual events, and it stacks latencies and compounds any pivot error downstream.
A live transcript of the original floor language, translated side by side, gives the relay-taking booth something it has never had: a direct line to the source while working from the pivot. Nobody interprets from the machine translation — but when a figure through relay sounds off, a glance at the floor transcript settles it. It is a cross-check on the pivot, not a bypass around the colleague providing it.
The same setup carries over to remote and hybrid events, where interpreters work from streams rather than consoles — the growing norm we examine in remote simultaneous interpreting in 2026.
Booths are shared professional space with norms older than any of this technology. The tool enters on the booth's terms:
Machine output does not deliver a diplomat's irony or repair a speaker's broken syntax in flight — interpreters do. The tool's job is to make the humans in the booth harder to rattle:
We make the full argument — including what would have to change before this stance changed — in Will AI replace interpreters? The short version: the booth keeps the judgment; the copilot keeps the digits. 30 free minutes every week, no credit card — enough to run it against one real conference session.
As a verification layer, not a script. A laptop runs Unicaption off the booth's floor feed via system audio, and the live transcript sits in peripheral vision so the interpreter can confirm numbers, names, and acronyms without breaking their rendition. The off-turn boothmate watches it too, flagging terminology for the partner on mic — the traditional support role, with better tooling.
Not in any setting where nuance, register, and accountability matter — which is most conference work. Machine translation cannot carry a speaker's irony, repair broken syntax in real time, or take professional responsibility for a rendition. AI transcription is genuinely good at holding figures and names as text, which is why tools like Unicaption position themselves as booth copilots rather than replacements.
The same feed the interpreter hears: the console's floor audio, captured via system audio on the laptop, or the event's clean stream in a browser tab. Never a microphone in the booth — it would hear the interpreter's own voice over the speaker. With a clean floor feed, Unicaption captions at state-of-the-art streaming accuracy moments behind the speaker.
Yes, in a specific way: it gives booths working from a pivot language a direct reference to the original floor. The interpreter still works from the relay, but a live transcript of the source — with translation side by side — lets them verify figures and names against the original when something through the pivot sounds off.
Live captions of the floor — figures, names, and lists held as text while you render. 30 free minutes every week, no credit card.
Start Free Trial →