Real-Time Voice Translation Guide
Table of Contents
- Overview
- Prerequisites
- Starting Voice Translation
- Sending Audio
- Receiving Recognition and Translation Results
- Operation Controls
- Advanced Features
- Multi-Channel Speaker Diarization
- Conversation Mode
- Stopping and Summary
- Complete Flow Diagram
- Related Documents
Overview
The VAS real-time voice translation service provides low-latency speech-to-text (STT) and real-time translation over WebSocket. The complete flow is:
- The client captures audio from the microphone
- The audio stream is sent to the VAS server
- The server performs speech recognition and returns the transcript
- Multi-language translation is performed in parallel and the results are returned
- (Optional) TTS speech is synthesized to play back the translation results
Use Cases
| Scenario | Recording Type (type) |
|---|---|
| Meeting notes, interview records | transcribe |
| Bilingual real-time interpretation, cross-language conversation | conversation |
| Voice memos, quick notes | record |
| Lectures, presentations, live streaming | broadcast (see the Broadcast Guide) |
Prerequisites
1. Obtain an API Key
Make sure you have a valid API Key (format: vas_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx). For authentication details, see Authentication.
2. Obtain a Ticket
WebSocket connections are authenticated using a Ticket mechanism. First, exchange your API Key for a one-time Ticket:
curl -X POST "https://vas-poc.vurbo.ai/api/v1/auth/ticket" \
-H "X-API-Key: vas_your_api_key_here"
Response:
{
"ticket": "aBcDeFgHiJkLmNoPqRsTuVwXyZ012345",
"expires_in": 60
}
Note: A Ticket is valid for 60 seconds and can be used only once.
3. Establish a WebSocket Connection
Place the Ticket in Sec-WebSocket-Protocol using the format ticket.{TICKET_VALUE}:
const ws = new WebSocket('wss://vas-poc.vurbo.ai/ws', [`ticket.${ticket}`]);
ws.onopen = () => {
console.log('WebSocket connected');
};
4. Maintain a Heartbeat
We recommend sending a ping every 30 seconds to ensure the connection does not time out:
{
"type": "health",
"data": { "action": "ping" }
}
The server responds with pong.
Starting Voice Translation
After the connection is established, send the start action to launch a voice translation session.
Basic Request
{
"type": "voice-translation",
"data": {
"action": "start",
"transcription_languages": ["zh-TW"],
"translation_languages": ["en-US"],
"type": "transcribe",
"audio_format": "pcm",
"summary_template": "meeting"
}
}
Core Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
transcription_languages | string[] | Yes | Speech recognition languages, up to 10 (e.g., ["zh-TW"]) |
translation_languages | string[] | No | Target translation languages; multiple allowed (up to 12; an empty array or omitting it means no translation). As of v1.6.7, specifying multiple languages translates all of them in real time — see Multi-Language Translation |
type | string | Yes | Recording type: transcribe, conversation, record, broadcast |
audio_format | string | No | Audio format: pcm (default) or webm |
summary_template | string | Conditional | Summary template (required for the transcribe type, e.g., meeting, interview) |
realtime_translation | boolean | No | Real-time translation mode (default false) |
recognition_mode | string | No | single (single speaker, default) or multi_speaker (multi-speaker diarization); under multi_speaker, transcription_languages must contain exactly 1 language, otherwise a diarization_multilang_conflict error is returned and the session is refused (type=conversation is exempt as of v1.7.2) |
name | string | No | Initial default recording name (max 60 characters; the system may still override it; if not provided, a name such as Transcription #1 is generated automatically) |
Best practice: list only the transcription languages that will actually occur
When you provide multiple transcription languages, the language of each speech segment is identified automatically; the more candidate languages you provide, the more likely misidentification becomes — especially when the content contains words shared across languages (proper nouns, numbers, loanwords). Recommendations:
- List only the languages that will actually occur; don't pad the list to 10 "just in case".
- Fewer candidates are more accurate; when you are certain there is only one language, provide just one (no language identification is needed, and accuracy is highest).
- Avoid listing languages that share a script or are close relatives (e.g. multiple Latin-script European languages,
zh-CNandzh-TW, or regional variants of the same language), which are the most easily confused.
Request with TTS
To enable speech synthesis of the translation results, add the TTS-related parameters:
{
"type": "voice-translation",
"data": {
"action": "start",
"transcription_languages": ["zh-TW"],
"translation_languages": ["en-US"],
"type": "transcribe",
"audio_format": "pcm",
"summary_template": "meeting",
"tts_enabled": true,
"tts_language": "en-US",
"tts_voice": "en-US-JennyNeural",
"tts_mode": "sync"
}
}
| TTS Parameter | Description |
|---|---|
tts_enabled | Whether to enable TTS (default false) |
tts_language | TTS output language (must be in translation_languages) |
tts_voice | TTS voice name (e.g., en-US-JennyNeural) |
tts_mode | sync (automatic playback, default) or async (manual control) |
Success Response
After a successful start, the server returns a session_started event:
{
"type": "voice-translation",
"data": {
"action": "session_started",
"session_id": "550e8400-e29b-41d4-a716-446655440000",
"task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
"recording_type": "transcribe",
"recognition_mode": "single",
"message": "Speech recognition started"
}
}
Save the session_id and task_id; they will be used in subsequent API operations.
Sending Audio
Once the session has started, continuously send audio data to the server.
Audio Format Requirements
PCM format (default, recommended):
| Item | Specification |
|---|---|
| Sample rate | 16000 Hz |
| Bit depth | 16-bit |
| Channels | Mono |
| Byte order | Little-endian |
WebM/Opus format: Any sample rate and number of channels; the server converts automatically.
Sending Format
Audio data must be Base64-encoded and sent with the audio action:
{
"type": "voice-translation",
"data": {
"action": "audio",
"payload": "Base64-encoded audio data..."
}
}
Front-End Audio Capture Example
const stream = await navigator.mediaDevices.getUserMedia({ audio: true });
const audioContext = new AudioContext({ sampleRate: 16000 });
const source = audioContext.createMediaStreamSource(stream);
const processor = audioContext.createScriptProcessor(4096, 1, 1);
processor.onaudioprocess = (e) => {
const float32 = e.inputBuffer.getChannelData(0);
// Convert to 16-bit PCM
const int16 = new Int16Array(float32.length);
for (let i = 0; i < float32.length; i++) {
int16[i] = Math.max(-32768, Math.min(32767, float32[i] * 32768));
}
// Base64-encode and send
const base64 = btoa(String.fromCharCode(...new Uint8Array(int16.buffer)));
ws.send(JSON.stringify({
type: 'voice-translation',
data: { action: 'audio', payload: base64 }
}));
};
source.connect(processor);
processor.connect(audioContext.destination);
Receiving Recognition and Translation Results
The server pushes recognition and translation results via the result event.
Speech Recognition Result (Origin)
{
"type": "voice-translation",
"data": {
"action": "result",
"origin": {
"sid": 1,
"language": "zh-TW",
"text": "你好,很高興認識你",
"is_final": true,
"speaker_id": "0",
"detected_language": "zh-TW",
"start_time": "00:05"
}
}
}
| Field | Description |
|---|---|
sid | Sentence number, incrementing from 1 |
text | The recognized text |
is_final | false for intermediate results (which will be overwritten); true for final results |
speaker_id | Speaker ID (meaningful in multi-speaker mode) |
start_time | Sentence start time (format mm:ss) |
Translation Result (Translations)
{
"type": "voice-translation",
"data": {
"action": "result",
"translations": {
"en-US": {
"sid": 1,
"text": "Hello, nice to meet you",
"is_final": true
}
}
}
}
Important:
originandtranslationsmay arrive in the sameresultevent, or they may be pushed separately. The front end should match them bysid.
When a sentence mixes in other languages, everything is translated into the translation language, except proper nouns such as names, brands, and products, and all-caps abbreviations (such as AI and API). Parts of the original that are already in the translation language are left unchanged.
TTS Audio Ready (TTS Ready)
If TTS is enabled, you receive a tts_ready event after the translation completes:
{
"type": "voice-translation",
"data": {
"action": "tts_ready",
"sid": 1,
"language": "en-US",
"text": "Hello, nice to meet you",
"audio": "Base64EncodedMP3...",
"format": "mp3",
"duration_ms": 2500,
"boundaries": [...]
}
}
The boundaries array contains Word Boundary information, which can be used to implement karaoke-style synchronized highlighting.
Operation Controls
Pause
Temporarily stop speech recognition processing:
{
"type": "voice-translation",
"data": { "action": "pause" }
}
Resume
Resume paused speech recognition:
{
"type": "voice-translation",
"data": { "action": "resume" }
}
Set Recording Name
There are two ways to set the recording name:
Method 1: Specify the name parameter at start (initial default name)
{
"type": "voice-translation",
"data": {
"action": "start",
"transcription_languages": ["zh-TW"],
"type": "transcribe",
"summary_template": "meeting",
"name": "Product Planning Meeting"
}
}
This name is an initial default; when the session ends, the system may still override it based on the transcript content.
Method 2: Use set_name during recording (fixed name)
{
"type": "voice-translation",
"data": {
"action": "set_name",
"name": "Product Planning Meeting"
}
}
A name set via
set_namewill not be overridden by the system.
If no name is set, the system automatically uses a "type + sequence number" format (e.g., Transcription #1, Broadcast #3). After the session ends, the system attempts to automatically generate a more meaningful name based on the transcript content (but it will not override a name set via set_name).
Switch Translation Language
Single-language sessions: switch the target language during recording; the system automatically retranslates all previously translated sentences:
{
"type": "voice-translation",
"data": {
"action": "switch_language",
"translation_languages": ["ja-JP"]
}
}
The system returns a language_switch_start event, followed by multiple batch_retranslation events, and finally a language_switch_done event, in order.
Multi-language sessions (v1.6.7): switch_language is redefined as "add / remove a single language" and the op parameter is required:
{
"type": "voice-translation",
"data": {
"action": "switch_language",
"op": "add",
"translation_languages": ["de-DE"]
}
}
op: "add": adds the language and automatically backfills existing sentences (same response sequence as above); the limit is 12 languagesop: "remove": removes the language and returns atranslation_language_removedevent; existing translations are kept, and at least 1 language must remain- Omitting
opin a multi-language session returns aswitch_language_op_requirederror (this prevents the single-language replace semantics from accidentally corrupting the language set)
Language-set sync: op:add and single-language replace emit identical events (both go through language_switch_start/done). These events carry translation_languages (the current full set of translation languages) — overwrite your local language set with it directly; do not infer append vs replace from the single translation_language (when only 1 language exists, an op:add expanding to a 2nd would be misread as a replace and drop the existing language).
See the WebSocket reference for switch_language.
Retranslate a Specific Sentence
After correcting a recognition error, you can retranslate a single sentence:
{
"type": "voice-translation",
"data": {
"action": "retranslate",
"sid": 1,
"translation_languages": ["en-US"],
"text": "Corrected source text"
}
}
Advanced Features
Multi-Language Translation
Specify multiple target languages in translation_languages (up to 12) to translate into several languages at once. This applies to transcribe and broadcast (for two-way translation the language is set by the server and is always a single language; record does not support translation as of v1.7.0):
{
"transcription_languages": ["zh-TW"],
"translation_languages": ["en-US", "ja-JP", "ko-KR"]
}
How results arrive (v1.6.7): each language returns its own independent result event — the same sentence (same sid) receives N result events, each with a single language key inside translations. Clients must accumulate translations by "sid + language code" instead of overwriting; completion order across languages is not fixed (translations run in parallel), and if one language fails the remaining languages are still delivered.
Real-time behavior: same as single-language, governed by realtime_translation — when true, all languages are translated word-by-word while the sentence is still being recognized (interim); when false (default), all languages are translated once the sentence is finalized.
Adjusting languages mid-session: multi-language sessions can add or remove languages via switch_language (op: "add" / op: "remove"); see Switch Translation Language.
Billing reminder: translation is billed per language starting from the second language — specifying N languages is billed as N languages (see the Pricing Guide). Multi-language real-time translation consumes noticeably more credits than single-language; choose the language count based on actual needs.
Speaker Recognition (Multi Speaker)
Set recognition_mode to multi_speaker to enable speaker recognition:
{
"recognition_mode": "multi_speaker"
}
Note: In
multi_speakermode,transcription_languagesmust contain exactly 1 language. If you provide multiple languages, you will receive adiarization_multilang_conflicterror and the session will be refused. Two-way translation (type=conversation) has been exempt from this restriction since v1.7.2 — it acceptsspeaker_diarizationbut ignores it.
Tip: If every speaker has a dedicated microphone, we recommend using Multi-Channel Speaker Diarization instead — speaker identity is determined by the channel with no inference needed, and each channel can be bound to a different language.
Once enabled, the speaker_id in the recognition results automatically distinguishes different speakers (e.g., Guest-1, Guest-2). You can manage speakers with the following operations:
rename_speaker: Globally rename a speaker (e.g., changeGuest-1toManager Wang)reassign_speaker: Change the speaker identity of a single sentencemerge_speakers: Merge two speakers (assign all sentences from one to the other)
TTS Playback Control
In async mode, you can manually control TTS playback:
Play a specific sentence:
{
"type": "voice-translation",
"data": {
"action": "tts_play",
"sid": 5,
"length": 3
}
}
Stop playback:
{
"type": "voice-translation",
"data": { "action": "tts_stop" }
}
Switch playback mode:
{
"type": "voice-translation",
"data": {
"action": "tts_mode",
"tts_mode": "async"
}
}
| Mode | Behavior |
|---|---|
sync | Automatically plays the latest is_final=true translation; the next sentence plays only after the previous one finishes |
async | Manually controls playback via tts_play |
Text Processing Parameters (Config)
Before start or during recording, you can use the config action to set the terminology list, fuzzy-term correction, and the translation dictionary:
{
"type": "voice-translation",
"data": {
"action": "config",
"terminology": {
"zh-TW": [
{ "term": "語者分離" },
{ "term": "CVD製程" }
]
},
"translation_dict": {
"en-US": [{ "source": "語者分離", "target": "Speaker Diarization" }]
}
}
}
| Setting | Description |
|---|---|
terminology | Terminology list -- improves recognition accuracy for specific terms (up to 500 entries across all languages combined) |
fuzzy_correction | Fuzzy-term correction -- corrects misspellings that sound different from the term (misspellings that sound the same are already covered by terminology, so this field is usually not needed) |
translation_dict | Translation dictionary -- ensures consistent translation of proper nouns (up to 3000 entries per language) |
Recommended practice: reach for
terminologyfirst. Besides driving homophone correction, it also improves recognition itself — reducing errors at the source rather than only fixing them afterwards. Only misspellings that sound different from the term (a foreign brand name recognized as a phonetically unrelated word, say) need afuzzy_correctionrule.Beyond 500 terms:
terminologyis capped at 500 entries across all languages. Anything beyond that can go intofuzzy_correctionwithcorrectonly and noincorrect(Chinese only) — it is matched by pronunciation the same way, and the rule cap is 4000. The trade-off is that this path does not improve recognition; it only corrects afterwards.
Multi-Channel Speaker Diarization
Multi-channel speaker diarization (recognition_mode: "multi_channel") lets a single recording take in multiple physical microphones at the same time, with each microphone running its own independent speech recognition. Speaker identity is determined by the channel — who is on which channel is declared at start, with no AI inference involved. It applies to recordings whose type is transcribe or record.
Note: This feature requires activation before use. In an environment where it is not activated, sending
recognition_mode: "multi_channel"returns aninvalid_recognition_modeerror.
Comparing the Three Speaker Diarization Approaches
| Approach | Setting | Speaker attribution | Best for |
|---|---|---|---|
| Single channel (no diarization) | recognition_mode: "single" (default) | No speaker distinction | Solo dictation, single-presenter recording |
| AI speaker diarization | recognition_mode: "multi_speaker" | Inferred by the system from voice characteristics | One microphone shared by several people (e.g., a single meeting-room mic) |
| Physical channel separation | recognition_mode: "multi_channel" | Determined by the channel (physical microphone) | Every speaker has a dedicated microphone |
How to choose:
- If you can give each person their own microphone, use physical channel separation: speaker attribution is determined by the channel and cannot be misassigned, simultaneous speech is recognized fully on each channel, and each channel can be bound to a different transcription language.
- When a single microphone picks up several people, use AI speaker diarization: the system infers the speaker from voice characteristics;
transcription_languagesmust contain exactly 1 language. - Multi-channel is itself a form of speaker diarization and cannot be combined with the
speaker_diarizationparameter (returnsinvalid_parameter). type=conversationdoes not support multi-channel (returnsinvalid_parameter); neither doestype=broadcast(returnsmultichannel_broadcast_not_allowed).
Starting a Multi-Channel Recording
{
"type": "voice-translation",
"data": {
"action": "start",
"type": "transcribe",
"recognition_mode": "multi_channel",
"channel_mode": "per_channel",
"channels": [
{ "channel_id": 1, "speaker_name": "Presenter", "transcription_languages": ["zh-TW"] },
{ "channel_id": 2, "speaker_name": "Panelist A", "transcription_languages": ["en-US"] }
],
"transcription_languages": ["zh-TW", "en-US"],
"translation_languages": ["ja-JP"],
"audio_format": "pcm",
"summary_template": "meeting"
}
}
| Rule | Description |
|---|---|
channel_mode | Required: "per_channel" (each channel is recognized independently) or "shared" (channels take turns speaking and share one recognition stream; see Taking Turns: Shared Mode below); other values return invalid_channel_mode |
channels | 1–8 channels (including the main presenter, conventionally channel_id: 1); channel_id must be in the range 1–8 and must not repeat |
| Per-channel language | per_channel: transcription_languages must contain exactly 1 language per channel; the union of all channels' languages must exactly match the session-level transcription_languages (after deduplication), otherwise channel_language_mismatch is returned. shared: channels carry no language |
audio_format | Only pcm is supported (16000 Hz / 16-bit / Mono / Little-endian); other values return multichannel_requires_pcm |
| TTS | Not supported in this first release: tts_enabled: true returns multichannel_tts_not_allowed |
| Plan cap | A plan can cap the number of channels; exceeding the cap makes start / add_channel return plan_feature_not_allowed on the spot (details.field: "max_stt_streams") |
After a successful start, the data of session_started (and of resume_ok after a reconnect) carries channel_mode and channels[] (each entry contains channel_id, speaker_name, transcription_languages, and status; transcription_languages is not present in shared mode), which you can use directly to initialize or restore the UI.
Sending Multi-Channel Audio
In multi-channel mode, every audio frame must carry a channel_id identifying its source channel:
{
"type": "voice-translation",
"data": {
"action": "audio",
"channel_id": 2,
"payload": "Base64-encoded audio data..."
}
}
- An unknown or already-removed
channel_idreturns anunknown_channel_iderror. - We recommend sending one frame every 100 ms.
- Every channel must keep sending audio continuously (including silence): a single silent channel does not end the recording; the recording ends automatically only when no channel recognizes any text for a long time (see Automatic End After a Long Silence).
Receiving Multi-Channel Results
In the result event, origin carries channel_id and the channel-determined speaker fields:
{
"type": "voice-translation",
"data": {
"action": "result",
"origin": {
"sid": 3,
"language": "en-US",
"text": "Hello everyone",
"is_final": true,
"channel_id": 2,
"speaker_id": "channel_2",
"speaker_label": "Panelist A",
"detected_language": "en-US",
"start_time": "00:12"
}
}
}
- The
speaker_idformat ischannel_{N}(determined by the channel and immutable);speaker_labelis the channel'sspeaker_name. translationsdoes not carrychannel_id; match it back tooriginbysid.- In multi-channel mode each channel produces sentences independently, so
sidis not guaranteed to be monotonically increasing; each transcript sentence carries achannel_idfield (single-channel recordings do not have this field).
Managing Channels During Recording
| Action | Purpose | Key points |
|---|---|---|
add_channel | Dynamically add a channel (channels must contain exactly 1 element) | Channel numbers cannot be reused (including removed ones — returns channel_id_in_use); subject to the overall cap of 8 channels and the plan's channel cap; the new channel starts producing text within about 4 seconds (audio sent during this period is not lost — text is just delayed) |
remove_channel | Disable a channel (carries channel_id) | The last remaining channel cannot be removed (channel_remove_not_allowed); transcripts and audio already produced are kept; a wind-down window of about 3 seconds lets the final sentence come back; the billed channel count drops from the next minute |
set_channel_language | Change one channel's language (carries channel_id and exactly 1 new language) | channel_id / speaker_id stay the same and the transcript remains continuous; only one change per channel is accepted every 5 seconds (channel_rebuild_too_frequent); introducing a new language is subject to the platform cap of 10 languages and the plan's language cap (too_many_languages); text resumes about 4 seconds after the change |
set_speaking_speed | Adjust the sentence-segmentation threshold | Supported in multi-channel mode; the new threshold applies to all channels, each channel briefly pausing text output while it is applied (same 5-second minimum interval between changes, channel_rebuild_too_frequent) |
switch_language(includingop:add/op:remove) always returnsmultichannel_switch_language_not_allowedin multi-channel mode — languages are bound to channels; useset_channel_languageinstead.- Channel operations are not allowed while paused (they return
channel_action_while_paused);resumefirst. - To change a channel's language, use
set_channel_language— do not useremove_channel+add_channel: channel numbers cannot be reused, and changing the number would split the same person into two speakers in the transcript.
For each action's full parameters and error codes, see the Voice Translation Reference.
The channel_status Event and UI Status Indicators
Whenever a channel is created, has its settings changed, is removed, or runs into a problem, the server pushes a channel_status event:
{
"type": "voice-translation",
"data": {
"action": "channel_status",
"channel": {
"channel_id": 2,
"speaker_name": "Panelist A",
"transcription_languages": ["ja-JP"],
"status": "preparing"
},
"active_channels": 3,
"stt_stream_count": 3,
"reason": "language_change"
}
}
status | Meaning | Suggested UI |
|---|---|---|
preparing | Preparing (being set up or applying a new setting); speech on this channel will take a few seconds before text appears | Yellow light (preparing) |
ready | The channel has started producing text | Green light (active) |
removed | Disabled by remove_channel | Gray light (disabled) |
error | The channel failed and cannot recover automatically | Red light (error) |
reason explains why the status changed. Possible values: added (channel added), language_change (language changed), reconnect (automatic reconnection after a drop), resumed (resumed after a pause), removed (disabled), speaking_speed (speaking-speed adjustment), stt_error (recognition failure).
Implementation tip: Drive each channel's status indicator with
channel_status— duringpreparing, show a "preparing" hint so users know to wait a moment (speech during this period is not lost, text is just delayed; the one exception is a sentence already in progress at the moment of the switch, which is reported as segment_discarded); show a green light onready; onerror, prompt the user to check the channel or re-add it. After a reconnect, restore all channel statuses in one pass from thechannels[]snapshot inresume_okinstead of relying only on accumulated events.
Equipment Requirements
- Use directional / close-talking microphones with a recommended pickup distance of 5–10 cm.
- Keep enough spacing between microphones so that one person's speech is not picked up by multiple channels at once (crosstalk).
- Crosstalk caused by not meeting the equipment requirements is outside the quality guarantee: when one person's speech is picked up on multiple channels, the same sentence may appear once on each of those channels.
Client-side crosstalk countermeasures (use either or both):
- Channel selection: at any given moment, send only the channel with the strongest energy (keep the other channels alive with silent frames).
- Physical isolation: make sure each microphone picks up only its own speaker.
Taking Turns: Shared Mode
When only one person speaks at a time (for example, a host and guests taking turns), you can use channel_mode: "shared": all channels share one recognition stream, the speaker is labeled by which channel that stretch of audio came in on, and billing is always counted as 1 channel (speech recognition 1.0 + speaker diarization 0.5 = 1.5 credits per minute).
{
"type": "voice-translation",
"data": {
"action": "start",
"type": "transcribe",
"recognition_mode": "multi_channel",
"channel_mode": "shared",
"transcription_languages": ["zh-TW"],
"audio_format": "pcm",
"channels": [
{ "channel_id": 1, "speaker_name": "Host" },
{ "channel_id": 2, "speaker_name": "Guest A" },
{ "channel_id": 3, "speaker_name": "Guest B" },
{ "channel_id": 4, "speaker_name": "Guest C" }
]
}
}
The client must:
- Send only one channel at a time: send
audioonly on the current speaker's channel. If several channels are sent at once, their audio is queued into the same recognition stream in the order received and the transcript becomes garbled. - Keep sending silence: when nobody is speaking, keep sending silent frames on the current channel; do not stop sending.
- Use
pcmonly. - Share languages across the session: channels carry no
transcription_languages(providing it returnschannel_language_not_allowed); the language cannot be changed during the recording. - Never remove the first channel: the first entry in
channels[]carries the recognition for the whole session, andremove_channelreturnschannel_remove_not_allowed.
Channel status: the other channels' preparing / ready / error follow the first channel, and each channel receives its own channel_status whenever the first channel's status changes; a channel added with add_channel takes on the first channel's current status directly; removed follows each channel's own status. See channel_status.
Known limitations:
- When the gap between speakers is shorter than about 0.8 seconds, the words of the two people may be merged into one sentence labeled with only one speaker.
- Speakers are determined by channel labeling, which is highly accurate when people take turns; speakers cannot be reassigned or merged afterward.
- Retroactive transcription after resuming from a pause covers the whole last 60 seconds, for all channels together, with speakers labeled.
Known Limitations
- TTS speech synthesis is not supported in this first release (
multichannel_tts_not_allowed). - Broadcast (
multichannel_broadcast_not_allowed) and conversation mode (invalid_parameter) are not supported. - File import does not support multi-channel.
- While paused, the audio file keeps being saved but no text is produced; after resuming, speech from the paused period is transcribed retroactively (timestamps reflect when the words were actually spoken). Catch-up transcription is capped at the trailing 60 seconds in total per session (under
per_channel, split evenly across channels when there are several; undershared, the whole last 60 seconds); anything beyond that is kept only in the audio file and does not enter the transcript. A sentence cut off mid-utterance at the moment of pausing may not appear in the transcript (same as single-channel). - Speaker management:
rename_speakeris available (you can rename a channel's speaker before it speaks);reassign_speakerandmerge_speakersdo not apply to multi-channel (they returnspeaker_op_not_allowed_multi_channel). - Audio file length cap: the audio saved for one multi-channel recording has a total cap, reached sooner the more channels are open (about 70 minutes with 8 channels). Once the cap is reached, the audio file and the recording duration stop at that point, while the transcript and credit charges continue as usual.
Billing reminder: Multi-channel is a form of speaker diarization. It is billed as speech recognition at 1.0 credit/minute plus speaker diarization at 0.5 credit/minute, with
per_channeladding a surcharge of (N − 1) × 1.0 credits/minute based on the current number of channels N (the 1st channel is included in the base rate), whilesharedis always counted as 1 channel with no surcharge. Credits are deducted at the start of each minute based on the channels open at that moment: a channel added viaadd_channelis billed from the next minute, and afterremove_channelthe rate drops from the next minute. Under an unlimited plan, multi-channel requires the plan to include this feature. See the Pricing Guide for details.
Conversation Mode
Conversation mode lets two people who speak different languages hold a real-time interpreted conversation over a single WebSocket connection. The system automatically detects the language of each utterance, translates it into the other person's language, and returns the translation result as TTS audio. Language detection is fully automatic; no manual switching is required.
Start a Conversation
{
"type": "voice-translation",
"data": {
"action": "start",
"type": "conversation",
"transcription_languages": ["zh-TW", "en-US"],
"audio_format": "pcm",
"tts_config": {
"zh-TW": { "voice": "zh-TW-HsiaoChenNeural", "speaking_rate": 1.0 },
"en-US": { "voice": "en-US-JennyNeural", "speaking_rate": 1.0 }
}
}
}
transcription_languagesmust contain exactly 2 languagesactive_languageis optional and specifies the initial preferred language (language detection is still automatic)tts_configcan be omitted; the system uses default voices automaticallytts_enableddefaults totrue; set it tofalseto return text translations only- Conversation mode translates only finalized sentences;
realtime_translationhas no effect in conversation mode, and the result is the same whether or not you send it
Automatic Language Detection
The system automatically detects the language of each utterance. The origin.language of each utterance directly reflects the detected language, and the translation target is automatically the other of the two languages.
Note: You do not need to call
switch_languagemanually to switch languages; the system detects them automatically.switch_languagecan still be used, but it only updates the internal preference state.
Switching TTS Settings Mid-Conversation
During a conversation, you can use set_tts to toggle TTS on or off or to update voice settings:
{
"type": "voice-translation",
"data": {
"action": "set_tts",
"tts_enabled": true,
"tts_config": {
"en-US": { "voice": "en-US-GuyNeural", "speaking_rate": 1.2 }
}
}
}
On success, you receive a tts_updated event containing the full updated settings.
Complete Conversation Flow
1. start (conversation, zh-TW + en-US)
2. session_started
3. Send audio (Person A speaks Chinese)
4. result (origin.language: "zh-TW", translations: en-US) ← automatic detection
5. tts_ready (en-US audio → played to Person B)
6. Send audio (Person B speaks English, no switching needed!)
7. result (origin.language: "en-US", translations: zh-TW) ← automatic detection
8. tts_ready (zh-TW audio → played to Person A)
9. stop
10. task_complete
Stopping and Summary
Stop Recording
Send the stop action to end the voice translation session:
{
"type": "voice-translation",
"data": { "action": "stop" }
}
A recording that goes a certain length of time without recognizing any text (15 minutes by default) ends automatically; what follows is the same as sending
stop. See Automatic End After a Long Silence.
Event Flow
After stopping, the system performs the following steps in order and pushes events:
status-- confirms that speech recognition has stopped- (Background processing) -- uploads the audio file and saves the transcript
task_complete-- task processing is complete, including thetask_id
{
"type": "voice-translation",
"data": {
"action": "task_complete",
"task_id": "550e8400-e29b-41d4-a716-446655440000",
"message": "Task processing complete"
}
}
When no audio was received during the whole recording,
task_completecarriesnoAudio: trueand there is no transcript to read. See task_complete.
- (If a summary template was set) -- the system automatically generates a summary; when the available credits cannot cover the summary fee, no summary is generated and
summary_erroris sent instead
Save the
task_idso you can later query the results via the Tasks API or load the history via the SSE API.
Complete Flow Diagram
Prerequisites
│
┌──────────────┼──────────────┐
│ │ │
Get API Key Get Ticket Open WebSocket
│ │ │
└──────────────┼──────────────┘
│
┌───────▼───────┐
│ config (optional)│ Set terminology / correction rules
└───────┬───────┘
│
┌───────▼───────┐
│ start │ Start voice translation
└───────┬───────┘
│
session_started
│
┌────────────▼────────────┐
│ │
┌─────▼─────┐ ┌─────▼─────┐
│ audio │────────────│ result │
│ (ongoing) │ Send audio │ Results │
└─────┬─────┘ └─────┬─────┘
│ │
│ ┌───────────────────┤
│ │ │
│ origin translations
│ (source) (translation)
│ │
│ ┌────────▼────────┐
│ │ tts_ready (optional)│
│ └─────────────────┘
│
┌─────▼─────┐ ┌──────────┐
│ pause / │◄──►│ resume │ Operation controls
│ resume │ └──────────┘
└─────┬─────┘
│
┌─────▼─────┐
│ stop │ Stop translation
└─────┬─────┘
│
┌─────▼──────────┐
│ task_complete │ Task complete (with task_id)
└─────┬──────────┘
│
┌─────▼─────┐
│ summary │ Summary generation (if a template is set)
└───────────┘
Related Documents
| Document | Description |
|---|---|
| Authentication | Detailed description of API Key and Ticket authentication |
| Voice Translation Reference | Complete API specification for all actions |
| Response Events Reference | Reference for all response event formats |
| History and Playback | How to load history after stopping |
| TTS Speech Synthesis | Complete guide to the TTS feature |
| Speaker Management | Renaming, reassigning, and merging speakers |
| Pricing Guide | Billing rules for each feature (including the multi-channel surcharge) |
Version: V1.24.1 Last Updated: 2026-10-07