One mode speaks while still listening. The other holds whole passages in memory. They are different skills, used in different rooms — and technology helps each one differently.
People outside the profession say "interpreter" as if it names one job. It names at least two. Simultaneous interpreting and consecutive interpreting differ in timing, in cognitive machinery, in the rooms that use them, and in what can go wrong — a conference interpreter and a medical interpreter are running genuinely different mental programs. Understanding the two modes is the key to almost every practical question about interpreting: what to book, what it costs, why interpreters work in pairs, and where technology actually helps. Here is the clean version, mode by mode.
In simultaneous interpreting, the interpreter speaks the target language while the source speaker keeps talking, trailing a few seconds behind. There is no pause, no turn-taking — the meeting runs at full speed and the interpretation rides alongside it.
It is the mode of conference booths, courtroom proceedings (interpreting quietly for a defendant as the hearing runs), international institutions, and multilingual broadcasts — anywhere stopping the speaker is not an option. The defining demand is split attention: listening to the next sentence while producing the last one, continuously, with an ear-voice span of a few seconds that must never collapse. It is why simultaneous interpreters work in pairs and swap every twenty to thirty minutes, and why the mode's classic failure points are precisely the items that cannot be inferred from context — numbers, names, dates, lists.
In consecutive interpreting, speaker and interpreter alternate: the speaker delivers a passage — a few sentences, sometimes several minutes — then pauses while the interpreter renders it. The conversation takes roughly twice as long, and in exchange every party hears everything in their language, completely.
It is the mode of medical appointments, depositions and witness interviews, business negotiations, asylum and immigration interviews, and community settings. The defining demand shifts from split attention to working memory and note-taking: holding a full passage — its structure, its qualifiers, its exact figures — and reproducing it faithfully. Professional consecutive interpreters use a personal shorthand of symbols and structure to pin down what memory alone would drop, and the mode's failure points are long unbroken passages, dense number sequences, and speakers who won't pause.
A third cousin appears constantly in real work: sight translation, reading a written document aloud in another language — consent forms, court orders, letters. It is technically translation performed live, and most working interpreters do it daily between turns.
| Simultaneous | Consecutive | |
|---|---|---|
| Timing | Speaks while the speaker continues | Speaks during pauses, in turns |
| Typical settings | Conferences, court proceedings, broadcasts, institutions | Medical visits, depositions, negotiations, community work |
| Core cognitive demand | Split attention; constant ear-voice span | Working memory; structured note-taking |
| Pace of the encounter | Full speed, uninterrupted | Roughly doubled in length |
| Staffing convention | Pairs, rotating every 20–30 minutes | Often solo, session-length dependent |
| Classic failure points | Numbers, names, lists, speed spikes | Long passages, dense figures, no pauses |
The staffing line explains a chunk of interpreting economics: simultaneous work bills for two professionals plus, historically, equipment — part of why conference interpreting is priced the way it is. The cognitive demands explain the profession's burnout problem, which we examine in interpreter cognitive load and burnout.
Notice that the failure points of each mode are exactly where a live, low-latency transcript earns its place — but it plays a different position in each.
In simultaneous mode, the interpreter cannot pause, cannot ask for a repeat, and cannot write much down. A real-time transcript running beside the audio acts as a safety net for the uninferable: when a speaker fires off "four hundred eighty-seven million" or a Polish surname mid-sentence, the number and the name are sitting in text a glance away, still on screen after the sound is gone. The boothmate traditionally jots exactly these items for the working interpreter — a transcript does it tirelessly, for both of them. How this is reshaping booth practice is its own story: AI in the interpreting booth, and its remote cousin in remote simultaneous interpreting.
Consecutive mode gets a different gift. The interpreter listens to a full passage, then renders it — and a live transcript changes both halves of that turn:
Unicaption was built as a copilot for working interpreters, and the design maps onto the two modes directly:
The first 30 minutes are free with no credit card — one consecutive appointment or one booth turn is enough to feel what the safety net changes. And to be clear about the boundary: the transcript assists the interpreter; it does not replace either mode — our honest treatment of that line is in Will AI replace interpreters?
Timing. A simultaneous interpreter renders speech into the target language while the speaker is still talking, trailing a few seconds behind — the mode of conference booths and court proceedings. A consecutive interpreter works in turns: the speaker delivers a passage and pauses, and the interpreter reproduces it from working memory and notes — the mode of medical appointments, depositions, and negotiations. They demand different skills: sustained split attention versus structured memory.
Whenever the encounter can afford pauses and precision matters more than pace: medical visits, legal interviews and depositions, negotiations, community and social-service settings. Consecutive roughly doubles the length of a conversation but gives every party a complete rendition. Simultaneous is chosen when the event cannot stop — conferences, live proceedings, broadcasts — and requires the interpreter to keep pace with an unbroken speaker.
Two show up constantly in real work. Whispered interpreting (chuchotage) is simultaneous without a booth — the interpreter whispers the live rendition to one or two listeners. Sight translation is reading a written document aloud in another language on the spot: consent forms, court orders, correspondence. Most working interpreters mix all four in a normal week.
Differently, and in neither case by replacing them. In simultaneous work, a live transcript like Unicaption's — in real time, with translation side by side — acts as a safety net for numbers, names, and lists that split attention drops. In consecutive work it supports note-taking by capturing the verbatim layer and lets the interpreter verify each rendition against the source text before speaking. It is a copilot for the interpreter's accuracy, not a substitute for the interpreting.
Live transcript and translation, side by side, in real time — in the booth and at the bedside. 30 free minutes every week.
Start Free Trial →