Guides

Real-Time Voice Translation Guide

Table of Contents

  1. Overview
  2. Prerequisites
  3. Starting Voice Translation
  4. Sending Audio
  5. Receiving Recognition and Translation Results
  6. Operation Controls
  7. Advanced Features
  8. Multi-Channel Speaker Diarization
  9. Conversation Mode
  10. Stopping and Summary
  11. Complete Flow Diagram
  12. Related Documents

Overview

The VAS real-time voice translation service provides low-latency speech-to-text (STT) and real-time translation over WebSocket. The complete flow is:

  1. The client captures audio from the microphone
  2. The audio stream is sent to the VAS server
  3. The server performs speech recognition and returns the transcript
  4. Multi-language translation is performed in parallel and the results are returned
  5. (Optional) TTS speech is synthesized to play back the translation results

Use Cases

ScenarioRecording Type (type)
Meeting notes, interview recordstranscribe
Bilingual real-time interpretation, cross-language conversationconversation
Voice memos, quick notesrecord
Lectures, presentations, live streamingbroadcast (see the Broadcast Guide)

Prerequisites

1. Obtain an API Key

Make sure you have a valid API Key (format: vas_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx). For authentication details, see Authentication.

2. Obtain a Ticket

WebSocket connections are authenticated using a Ticket mechanism. First, exchange your API Key for a one-time Ticket:

curl -X POST "https://vas-poc.vurbo.ai/api/v1/auth/ticket" \
  -H "X-API-Key: vas_your_api_key_here"

Response:

{
  "ticket": "aBcDeFgHiJkLmNoPqRsTuVwXyZ012345",
  "expires_in": 60
}

Note: A Ticket is valid for 60 seconds and can be used only once.

3. Establish a WebSocket Connection

Place the Ticket in Sec-WebSocket-Protocol using the format ticket.{TICKET_VALUE}:

const ws = new WebSocket('wss://vas-poc.vurbo.ai/ws', [`ticket.${ticket}`]);

ws.onopen = () => {
  console.log('WebSocket connected');
};

4. Maintain a Heartbeat

We recommend sending a ping every 30 seconds to ensure the connection does not time out:

{
  "type": "health",
  "data": { "action": "ping" }
}

The server responds with pong.


Starting Voice Translation

After the connection is established, send the start action to launch a voice translation session.

Basic Request

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "transcription_languages": ["zh-TW"],
    "translation_languages": ["en-US"],
    "type": "transcribe",
    "audio_format": "pcm",
    "summary_template": "meeting"
  }
}

Core Parameters

ParameterTypeRequiredDescription
transcription_languagesstring[]YesSpeech recognition languages, up to 10 (e.g., ["zh-TW"])
translation_languagesstring[]NoTarget translation languages; multiple allowed (up to 12; an empty array or omitting it means no translation). As of v1.6.7, specifying multiple languages translates all of them in real time — see Multi-Language Translation
typestringYesRecording type: transcribe, conversation, record, broadcast
audio_formatstringNoAudio format: pcm (default) or webm
summary_templatestringConditionalSummary template (required for the transcribe type, e.g., meeting, interview)
realtime_translationbooleanNoReal-time translation mode (default false)
recognition_modestringNosingle (single speaker, default) or multi_speaker (multi-speaker diarization); under multi_speaker, transcription_languages must contain exactly 1 language, otherwise a diarization_multilang_conflict error is returned and the session is refused (type=conversation is exempt as of v1.7.2)
namestringNoInitial default recording name (max 60 characters; the system may still override it; if not provided, a name such as Transcription #1 is generated automatically)

Best practice: list only the transcription languages that will actually occur

When you provide multiple transcription languages, the language of each speech segment is identified automatically; the more candidate languages you provide, the more likely misidentification becomes — especially when the content contains words shared across languages (proper nouns, numbers, loanwords). Recommendations:

  • List only the languages that will actually occur; don't pad the list to 10 "just in case".
  • Fewer candidates are more accurate; when you are certain there is only one language, provide just one (no language identification is needed, and accuracy is highest).
  • Avoid listing languages that share a script or are close relatives (e.g. multiple Latin-script European languages, zh-CN and zh-TW, or regional variants of the same language), which are the most easily confused.

Request with TTS

To enable speech synthesis of the translation results, add the TTS-related parameters:

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "transcription_languages": ["zh-TW"],
    "translation_languages": ["en-US"],
    "type": "transcribe",
    "audio_format": "pcm",
    "summary_template": "meeting",
    "tts_enabled": true,
    "tts_language": "en-US",
    "tts_voice": "en-US-JennyNeural",
    "tts_mode": "sync"
  }
}
TTS ParameterDescription
tts_enabledWhether to enable TTS (default false)
tts_languageTTS output language (must be in translation_languages)
tts_voiceTTS voice name (e.g., en-US-JennyNeural)
tts_modesync (automatic playback, default) or async (manual control)

Success Response

After a successful start, the server returns a session_started event:

{
  "type": "voice-translation",
  "data": {
    "action": "session_started",
    "session_id": "550e8400-e29b-41d4-a716-446655440000",
    "task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
    "recording_type": "transcribe",
    "recognition_mode": "single",
    "message": "Speech recognition started"
  }
}

Save the session_id and task_id; they will be used in subsequent API operations.


Sending Audio

Once the session has started, continuously send audio data to the server.

Audio Format Requirements

PCM format (default, recommended):

ItemSpecification
Sample rate16000 Hz
Bit depth16-bit
ChannelsMono
Byte orderLittle-endian

WebM/Opus format: Any sample rate and number of channels; the server converts automatically.

Sending Format

Audio data must be Base64-encoded and sent with the audio action:

{
  "type": "voice-translation",
  "data": {
    "action": "audio",
    "payload": "Base64-encoded audio data..."
  }
}

Front-End Audio Capture Example

const stream = await navigator.mediaDevices.getUserMedia({ audio: true });
const audioContext = new AudioContext({ sampleRate: 16000 });
const source = audioContext.createMediaStreamSource(stream);
const processor = audioContext.createScriptProcessor(4096, 1, 1);

processor.onaudioprocess = (e) => {
  const float32 = e.inputBuffer.getChannelData(0);
  // Convert to 16-bit PCM
  const int16 = new Int16Array(float32.length);
  for (let i = 0; i < float32.length; i++) {
    int16[i] = Math.max(-32768, Math.min(32767, float32[i] * 32768));
  }
  // Base64-encode and send
  const base64 = btoa(String.fromCharCode(...new Uint8Array(int16.buffer)));
  ws.send(JSON.stringify({
    type: 'voice-translation',
    data: { action: 'audio', payload: base64 }
  }));
};

source.connect(processor);
processor.connect(audioContext.destination);

Receiving Recognition and Translation Results

The server pushes recognition and translation results via the result event.

Speech Recognition Result (Origin)

{
  "type": "voice-translation",
  "data": {
    "action": "result",
    "origin": {
      "sid": 1,
      "language": "zh-TW",
      "text": "你好,很高興認識你",
      "is_final": true,
      "speaker_id": "0",
      "detected_language": "zh-TW",
      "start_time": "00:05"
    }
  }
}
FieldDescription
sidSentence number, incrementing from 1
textThe recognized text
is_finalfalse for intermediate results (which will be overwritten); true for final results
speaker_idSpeaker ID (meaningful in multi-speaker mode)
start_timeSentence start time (format mm:ss)

Translation Result (Translations)

{
  "type": "voice-translation",
  "data": {
    "action": "result",
    "translations": {
      "en-US": {
        "sid": 1,
        "text": "Hello, nice to meet you",
        "is_final": true
      }
    }
  }
}

Important: origin and translations may arrive in the same result event, or they may be pushed separately. The front end should match them by sid.

When a sentence mixes in other languages, everything is translated into the translation language, except proper nouns such as names, brands, and products, and all-caps abbreviations (such as AI and API). Parts of the original that are already in the translation language are left unchanged.

TTS Audio Ready (TTS Ready)

If TTS is enabled, you receive a tts_ready event after the translation completes:

{
  "type": "voice-translation",
  "data": {
    "action": "tts_ready",
    "sid": 1,
    "language": "en-US",
    "text": "Hello, nice to meet you",
    "audio": "Base64EncodedMP3...",
    "format": "mp3",
    "duration_ms": 2500,
    "boundaries": [...]
  }
}

The boundaries array contains Word Boundary information, which can be used to implement karaoke-style synchronized highlighting.


Operation Controls

Pause

Temporarily stop speech recognition processing:

{
  "type": "voice-translation",
  "data": { "action": "pause" }
}

Resume

Resume paused speech recognition:

{
  "type": "voice-translation",
  "data": { "action": "resume" }
}

Set Recording Name

There are two ways to set the recording name:

Method 1: Specify the name parameter at start (initial default name)

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "transcription_languages": ["zh-TW"],
    "type": "transcribe",
    "summary_template": "meeting",
    "name": "Product Planning Meeting"
  }
}

This name is an initial default; when the session ends, the system may still override it based on the transcript content.

Method 2: Use set_name during recording (fixed name)

{
  "type": "voice-translation",
  "data": {
    "action": "set_name",
    "name": "Product Planning Meeting"
  }
}

A name set via set_name will not be overridden by the system.

If no name is set, the system automatically uses a "type + sequence number" format (e.g., Transcription #1, Broadcast #3). After the session ends, the system attempts to automatically generate a more meaningful name based on the transcript content (but it will not override a name set via set_name).

Switch Translation Language

Single-language sessions: switch the target language during recording; the system automatically retranslates all previously translated sentences:

{
  "type": "voice-translation",
  "data": {
    "action": "switch_language",
    "translation_languages": ["ja-JP"]
  }
}

The system returns a language_switch_start event, followed by multiple batch_retranslation events, and finally a language_switch_done event, in order.

Multi-language sessions (v1.6.7): switch_language is redefined as "add / remove a single language" and the op parameter is required:

{
  "type": "voice-translation",
  "data": {
    "action": "switch_language",
    "op": "add",
    "translation_languages": ["de-DE"]
  }
}
  • op: "add": adds the language and automatically backfills existing sentences (same response sequence as above); the limit is 12 languages
  • op: "remove": removes the language and returns a translation_language_removed event; existing translations are kept, and at least 1 language must remain
  • Omitting op in a multi-language session returns a switch_language_op_required error (this prevents the single-language replace semantics from accidentally corrupting the language set)

Language-set sync: op:add and single-language replace emit identical events (both go through language_switch_start/done). These events carry translation_languages (the current full set of translation languages) — overwrite your local language set with it directly; do not infer append vs replace from the single translation_language (when only 1 language exists, an op:add expanding to a 2nd would be misread as a replace and drop the existing language).

See the WebSocket reference for switch_language.

Retranslate a Specific Sentence

After correcting a recognition error, you can retranslate a single sentence:

{
  "type": "voice-translation",
  "data": {
    "action": "retranslate",
    "sid": 1,
    "translation_languages": ["en-US"],
    "text": "Corrected source text"
  }
}

Advanced Features

Multi-Language Translation

Specify multiple target languages in translation_languages (up to 12) to translate into several languages at once. This applies to transcribe and broadcast (for two-way translation the language is set by the server and is always a single language; record does not support translation as of v1.7.0):

{
  "transcription_languages": ["zh-TW"],
  "translation_languages": ["en-US", "ja-JP", "ko-KR"]
}

How results arrive (v1.6.7): each language returns its own independent result event — the same sentence (same sid) receives N result events, each with a single language key inside translations. Clients must accumulate translations by "sid + language code" instead of overwriting; completion order across languages is not fixed (translations run in parallel), and if one language fails the remaining languages are still delivered.

Real-time behavior: same as single-language, governed by realtime_translation — when true, all languages are translated word-by-word while the sentence is still being recognized (interim); when false (default), all languages are translated once the sentence is finalized.

Adjusting languages mid-session: multi-language sessions can add or remove languages via switch_language (op: "add" / op: "remove"); see Switch Translation Language.

Billing reminder: translation is billed per language starting from the second language — specifying N languages is billed as N languages (see the Pricing Guide). Multi-language real-time translation consumes noticeably more credits than single-language; choose the language count based on actual needs.

Speaker Recognition (Multi Speaker)

Set recognition_mode to multi_speaker to enable speaker recognition:

{
  "recognition_mode": "multi_speaker"
}

Note: In multi_speaker mode, transcription_languages must contain exactly 1 language. If you provide multiple languages, you will receive a diarization_multilang_conflict error and the session will be refused. Two-way translation (type=conversation) has been exempt from this restriction since v1.7.2 — it accepts speaker_diarization but ignores it.

Tip: If every speaker has a dedicated microphone, we recommend using Multi-Channel Speaker Diarization instead — speaker identity is determined by the channel with no inference needed, and each channel can be bound to a different language.

Once enabled, the speaker_id in the recognition results automatically distinguishes different speakers (e.g., Guest-1, Guest-2). You can manage speakers with the following operations:

  • rename_speaker: Globally rename a speaker (e.g., change Guest-1 to Manager Wang)
  • reassign_speaker: Change the speaker identity of a single sentence
  • merge_speakers: Merge two speakers (assign all sentences from one to the other)

TTS Playback Control

In async mode, you can manually control TTS playback:

Play a specific sentence:

{
  "type": "voice-translation",
  "data": {
    "action": "tts_play",
    "sid": 5,
    "length": 3
  }
}

Stop playback:

{
  "type": "voice-translation",
  "data": { "action": "tts_stop" }
}

Switch playback mode:

{
  "type": "voice-translation",
  "data": {
    "action": "tts_mode",
    "tts_mode": "async"
  }
}
ModeBehavior
syncAutomatically plays the latest is_final=true translation; the next sentence plays only after the previous one finishes
asyncManually controls playback via tts_play

Text Processing Parameters (Config)

Before start or during recording, you can use the config action to set the terminology list, fuzzy-term correction, and the translation dictionary:

{
  "type": "voice-translation",
  "data": {
    "action": "config",
    "terminology": {
      "zh-TW": [
        { "term": "語者分離" },
        { "term": "CVD製程" }
      ]
    },
    "translation_dict": {
      "en-US": [{ "source": "語者分離", "target": "Speaker Diarization" }]
    }
  }
}
SettingDescription
terminologyTerminology list -- improves recognition accuracy for specific terms (up to 500 entries across all languages combined)
fuzzy_correctionFuzzy-term correction -- corrects misspellings that sound different from the term (misspellings that sound the same are already covered by terminology, so this field is usually not needed)
translation_dictTranslation dictionary -- ensures consistent translation of proper nouns (up to 3000 entries per language)

Recommended practice: reach for terminology first. Besides driving homophone correction, it also improves recognition itself — reducing errors at the source rather than only fixing them afterwards. Only misspellings that sound different from the term (a foreign brand name recognized as a phonetically unrelated word, say) need a fuzzy_correction rule.

Beyond 500 terms: terminology is capped at 500 entries across all languages. Anything beyond that can go into fuzzy_correction with correct only and no incorrect (Chinese only) — it is matched by pronunciation the same way, and the rule cap is 4000. The trade-off is that this path does not improve recognition; it only corrects afterwards.


Multi-Channel Speaker Diarization

Multi-channel speaker diarization (recognition_mode: "multi_channel") lets a single recording take in multiple physical microphones at the same time, with each microphone running its own independent speech recognition. Speaker identity is determined by the channel — who is on which channel is declared at start, with no AI inference involved. It applies to recordings whose type is transcribe or record.

Note: This feature requires activation before use. In an environment where it is not activated, sending recognition_mode: "multi_channel" returns an invalid_recognition_mode error.

Comparing the Three Speaker Diarization Approaches

ApproachSettingSpeaker attributionBest for
Single channel (no diarization)recognition_mode: "single" (default)No speaker distinctionSolo dictation, single-presenter recording
AI speaker diarizationrecognition_mode: "multi_speaker"Inferred by the system from voice characteristicsOne microphone shared by several people (e.g., a single meeting-room mic)
Physical channel separationrecognition_mode: "multi_channel"Determined by the channel (physical microphone)Every speaker has a dedicated microphone

How to choose:

  • If you can give each person their own microphone, use physical channel separation: speaker attribution is determined by the channel and cannot be misassigned, simultaneous speech is recognized fully on each channel, and each channel can be bound to a different transcription language.
  • When a single microphone picks up several people, use AI speaker diarization: the system infers the speaker from voice characteristics; transcription_languages must contain exactly 1 language.
  • Multi-channel is itself a form of speaker diarization and cannot be combined with the speaker_diarization parameter (returns invalid_parameter).
  • type=conversation does not support multi-channel (returns invalid_parameter); neither does type=broadcast (returns multichannel_broadcast_not_allowed).

Starting a Multi-Channel Recording

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "type": "transcribe",
    "recognition_mode": "multi_channel",
    "channel_mode": "per_channel",
    "channels": [
      { "channel_id": 1, "speaker_name": "Presenter", "transcription_languages": ["zh-TW"] },
      { "channel_id": 2, "speaker_name": "Panelist A", "transcription_languages": ["en-US"] }
    ],
    "transcription_languages": ["zh-TW", "en-US"],
    "translation_languages": ["ja-JP"],
    "audio_format": "pcm",
    "summary_template": "meeting"
  }
}
RuleDescription
channel_modeRequired: "per_channel" (each channel is recognized independently) or "shared" (channels take turns speaking and share one recognition stream; see Taking Turns: Shared Mode below); other values return invalid_channel_mode
channels1–8 channels (including the main presenter, conventionally channel_id: 1); channel_id must be in the range 1–8 and must not repeat
Per-channel languageper_channel: transcription_languages must contain exactly 1 language per channel; the union of all channels' languages must exactly match the session-level transcription_languages (after deduplication), otherwise channel_language_mismatch is returned. shared: channels carry no language
audio_formatOnly pcm is supported (16000 Hz / 16-bit / Mono / Little-endian); other values return multichannel_requires_pcm
TTSNot supported in this first release: tts_enabled: true returns multichannel_tts_not_allowed
Plan capA plan can cap the number of channels; exceeding the cap makes start / add_channel return plan_feature_not_allowed on the spot (details.field: "max_stt_streams")

After a successful start, the data of session_started (and of resume_ok after a reconnect) carries channel_mode and channels[] (each entry contains channel_id, speaker_name, transcription_languages, and status; transcription_languages is not present in shared mode), which you can use directly to initialize or restore the UI.

Sending Multi-Channel Audio

In multi-channel mode, every audio frame must carry a channel_id identifying its source channel:

{
  "type": "voice-translation",
  "data": {
    "action": "audio",
    "channel_id": 2,
    "payload": "Base64-encoded audio data..."
  }
}
  • An unknown or already-removed channel_id returns an unknown_channel_id error.
  • We recommend sending one frame every 100 ms.
  • Every channel must keep sending audio continuously (including silence): a single silent channel does not end the recording; the recording ends automatically only when no channel recognizes any text for a long time (see Automatic End After a Long Silence).

Receiving Multi-Channel Results

In the result event, origin carries channel_id and the channel-determined speaker fields:

{
  "type": "voice-translation",
  "data": {
    "action": "result",
    "origin": {
      "sid": 3,
      "language": "en-US",
      "text": "Hello everyone",
      "is_final": true,
      "channel_id": 2,
      "speaker_id": "channel_2",
      "speaker_label": "Panelist A",
      "detected_language": "en-US",
      "start_time": "00:12"
    }
  }
}
  • The speaker_id format is channel_{N} (determined by the channel and immutable); speaker_label is the channel's speaker_name.
  • translations does not carry channel_id; match it back to origin by sid.
  • In multi-channel mode each channel produces sentences independently, so sid is not guaranteed to be monotonically increasing; each transcript sentence carries a channel_id field (single-channel recordings do not have this field).

Managing Channels During Recording

ActionPurposeKey points
add_channelDynamically add a channel (channels must contain exactly 1 element)Channel numbers cannot be reused (including removed ones — returns channel_id_in_use); subject to the overall cap of 8 channels and the plan's channel cap; the new channel starts producing text within about 4 seconds (audio sent during this period is not lost — text is just delayed)
remove_channelDisable a channel (carries channel_id)The last remaining channel cannot be removed (channel_remove_not_allowed); transcripts and audio already produced are kept; a wind-down window of about 3 seconds lets the final sentence come back; the billed channel count drops from the next minute
set_channel_languageChange one channel's language (carries channel_id and exactly 1 new language)channel_id / speaker_id stay the same and the transcript remains continuous; only one change per channel is accepted every 5 seconds (channel_rebuild_too_frequent); introducing a new language is subject to the platform cap of 10 languages and the plan's language cap (too_many_languages); text resumes about 4 seconds after the change
set_speaking_speedAdjust the sentence-segmentation thresholdSupported in multi-channel mode; the new threshold applies to all channels, each channel briefly pausing text output while it is applied (same 5-second minimum interval between changes, channel_rebuild_too_frequent)
  • switch_language (including op:add / op:remove) always returns multichannel_switch_language_not_allowed in multi-channel mode — languages are bound to channels; use set_channel_language instead.
  • Channel operations are not allowed while paused (they return channel_action_while_paused); resume first.
  • To change a channel's language, use set_channel_language — do not use remove_channel + add_channel: channel numbers cannot be reused, and changing the number would split the same person into two speakers in the transcript.

For each action's full parameters and error codes, see the Voice Translation Reference.

The channel_status Event and UI Status Indicators

Whenever a channel is created, has its settings changed, is removed, or runs into a problem, the server pushes a channel_status event:

{
  "type": "voice-translation",
  "data": {
    "action": "channel_status",
    "channel": {
      "channel_id": 2,
      "speaker_name": "Panelist A",
      "transcription_languages": ["ja-JP"],
      "status": "preparing"
    },
    "active_channels": 3,
    "stt_stream_count": 3,
    "reason": "language_change"
  }
}
statusMeaningSuggested UI
preparingPreparing (being set up or applying a new setting); speech on this channel will take a few seconds before text appearsYellow light (preparing)
readyThe channel has started producing textGreen light (active)
removedDisabled by remove_channelGray light (disabled)
errorThe channel failed and cannot recover automaticallyRed light (error)

reason explains why the status changed. Possible values: added (channel added), language_change (language changed), reconnect (automatic reconnection after a drop), resumed (resumed after a pause), removed (disabled), speaking_speed (speaking-speed adjustment), stt_error (recognition failure).

Implementation tip: Drive each channel's status indicator with channel_status — during preparing, show a "preparing" hint so users know to wait a moment (speech during this period is not lost, text is just delayed; the one exception is a sentence already in progress at the moment of the switch, which is reported as segment_discarded); show a green light on ready; on error, prompt the user to check the channel or re-add it. After a reconnect, restore all channel statuses in one pass from the channels[] snapshot in resume_ok instead of relying only on accumulated events.

Equipment Requirements

  • Use directional / close-talking microphones with a recommended pickup distance of 5–10 cm.
  • Keep enough spacing between microphones so that one person's speech is not picked up by multiple channels at once (crosstalk).
  • Crosstalk caused by not meeting the equipment requirements is outside the quality guarantee: when one person's speech is picked up on multiple channels, the same sentence may appear once on each of those channels.

Client-side crosstalk countermeasures (use either or both):

  • Channel selection: at any given moment, send only the channel with the strongest energy (keep the other channels alive with silent frames).
  • Physical isolation: make sure each microphone picks up only its own speaker.

Taking Turns: Shared Mode

When only one person speaks at a time (for example, a host and guests taking turns), you can use channel_mode: "shared": all channels share one recognition stream, the speaker is labeled by which channel that stretch of audio came in on, and billing is always counted as 1 channel (speech recognition 1.0 + speaker diarization 0.5 = 1.5 credits per minute).

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "type": "transcribe",
    "recognition_mode": "multi_channel",
    "channel_mode": "shared",
    "transcription_languages": ["zh-TW"],
    "audio_format": "pcm",
    "channels": [
      { "channel_id": 1, "speaker_name": "Host" },
      { "channel_id": 2, "speaker_name": "Guest A" },
      { "channel_id": 3, "speaker_name": "Guest B" },
      { "channel_id": 4, "speaker_name": "Guest C" }
    ]
  }
}

The client must:

  • Send only one channel at a time: send audio only on the current speaker's channel. If several channels are sent at once, their audio is queued into the same recognition stream in the order received and the transcript becomes garbled.
  • Keep sending silence: when nobody is speaking, keep sending silent frames on the current channel; do not stop sending.
  • Use pcm only.
  • Share languages across the session: channels carry no transcription_languages (providing it returns channel_language_not_allowed); the language cannot be changed during the recording.
  • Never remove the first channel: the first entry in channels[] carries the recognition for the whole session, and remove_channel returns channel_remove_not_allowed.

Channel status: the other channels' preparing / ready / error follow the first channel, and each channel receives its own channel_status whenever the first channel's status changes; a channel added with add_channel takes on the first channel's current status directly; removed follows each channel's own status. See channel_status.

Known limitations:

  • When the gap between speakers is shorter than about 0.8 seconds, the words of the two people may be merged into one sentence labeled with only one speaker.
  • Speakers are determined by channel labeling, which is highly accurate when people take turns; speakers cannot be reassigned or merged afterward.
  • Retroactive transcription after resuming from a pause covers the whole last 60 seconds, for all channels together, with speakers labeled.

Known Limitations

  • TTS speech synthesis is not supported in this first release (multichannel_tts_not_allowed).
  • Broadcast (multichannel_broadcast_not_allowed) and conversation mode (invalid_parameter) are not supported.
  • File import does not support multi-channel.
  • While paused, the audio file keeps being saved but no text is produced; after resuming, speech from the paused period is transcribed retroactively (timestamps reflect when the words were actually spoken). Catch-up transcription is capped at the trailing 60 seconds in total per session (under per_channel, split evenly across channels when there are several; under shared, the whole last 60 seconds); anything beyond that is kept only in the audio file and does not enter the transcript. A sentence cut off mid-utterance at the moment of pausing may not appear in the transcript (same as single-channel).
  • Speaker management: rename_speaker is available (you can rename a channel's speaker before it speaks); reassign_speaker and merge_speakers do not apply to multi-channel (they return speaker_op_not_allowed_multi_channel).
  • Audio file length cap: the audio saved for one multi-channel recording has a total cap, reached sooner the more channels are open (about 70 minutes with 8 channels). Once the cap is reached, the audio file and the recording duration stop at that point, while the transcript and credit charges continue as usual.

Billing reminder: Multi-channel is a form of speaker diarization. It is billed as speech recognition at 1.0 credit/minute plus speaker diarization at 0.5 credit/minute, with per_channel adding a surcharge of (N − 1) × 1.0 credits/minute based on the current number of channels N (the 1st channel is included in the base rate), while shared is always counted as 1 channel with no surcharge. Credits are deducted at the start of each minute based on the channels open at that moment: a channel added via add_channel is billed from the next minute, and after remove_channel the rate drops from the next minute. Under an unlimited plan, multi-channel requires the plan to include this feature. See the Pricing Guide for details.


Conversation Mode

Conversation mode lets two people who speak different languages hold a real-time interpreted conversation over a single WebSocket connection. The system automatically detects the language of each utterance, translates it into the other person's language, and returns the translation result as TTS audio. Language detection is fully automatic; no manual switching is required.

Start a Conversation

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "type": "conversation",
    "transcription_languages": ["zh-TW", "en-US"],
    "audio_format": "pcm",
    "tts_config": {
      "zh-TW": { "voice": "zh-TW-HsiaoChenNeural", "speaking_rate": 1.0 },
      "en-US": { "voice": "en-US-JennyNeural", "speaking_rate": 1.0 }
    }
  }
}
  • transcription_languages must contain exactly 2 languages
  • active_language is optional and specifies the initial preferred language (language detection is still automatic)
  • tts_config can be omitted; the system uses default voices automatically
  • tts_enabled defaults to true; set it to false to return text translations only
  • Conversation mode translates only finalized sentences; realtime_translation has no effect in conversation mode, and the result is the same whether or not you send it

Automatic Language Detection

The system automatically detects the language of each utterance. The origin.language of each utterance directly reflects the detected language, and the translation target is automatically the other of the two languages.

Note: You do not need to call switch_language manually to switch languages; the system detects them automatically. switch_language can still be used, but it only updates the internal preference state.

Switching TTS Settings Mid-Conversation

During a conversation, you can use set_tts to toggle TTS on or off or to update voice settings:

{
  "type": "voice-translation",
  "data": {
    "action": "set_tts",
    "tts_enabled": true,
    "tts_config": {
      "en-US": { "voice": "en-US-GuyNeural", "speaking_rate": 1.2 }
    }
  }
}

On success, you receive a tts_updated event containing the full updated settings.

Complete Conversation Flow

1. start (conversation, zh-TW + en-US)
2. session_started
3. Send audio (Person A speaks Chinese)
4. result (origin.language: "zh-TW", translations: en-US)  ← automatic detection
5. tts_ready (en-US audio → played to Person B)
6. Send audio (Person B speaks English, no switching needed!)
7. result (origin.language: "en-US", translations: zh-TW)  ← automatic detection
8. tts_ready (zh-TW audio → played to Person A)
9. stop
10. task_complete

Stopping and Summary

Stop Recording

Send the stop action to end the voice translation session:

{
  "type": "voice-translation",
  "data": { "action": "stop" }
}

A recording that goes a certain length of time without recognizing any text (15 minutes by default) ends automatically; what follows is the same as sending stop. See Automatic End After a Long Silence.

Event Flow

After stopping, the system performs the following steps in order and pushes events:

  1. status -- confirms that speech recognition has stopped
  2. (Background processing) -- uploads the audio file and saves the transcript
  3. task_complete -- task processing is complete, including the task_id
{
  "type": "voice-translation",
  "data": {
    "action": "task_complete",
    "task_id": "550e8400-e29b-41d4-a716-446655440000",
    "message": "Task processing complete"
  }
}

When no audio was received during the whole recording, task_complete carries noAudio: true and there is no transcript to read. See task_complete.

  1. (If a summary template was set) -- the system automatically generates a summary; when the available credits cannot cover the summary fee, no summary is generated and summary_error is sent instead

Save the task_id so you can later query the results via the Tasks API or load the history via the SSE API.


Complete Flow Diagram

                    Prerequisites
                       │
        ┌──────────────┼──────────────┐
        │              │              │
   Get API Key     Get Ticket    Open WebSocket
        │              │              │
        └──────────────┼──────────────┘
                       │
               ┌───────▼───────┐
               │ config (optional)│  Set terminology / correction rules
               └───────┬───────┘
                       │
               ┌───────▼───────┐
               │     start     │  Start voice translation
               └───────┬───────┘
                       │
               session_started
                       │
          ┌────────────▼────────────┐
          │                         │
    ┌─────▼─────┐            ┌─────▼─────┐
    │   audio   │────────────│   result  │
    │ (ongoing) │  Send audio │  Results  │
    └─────┬─────┘            └─────┬─────┘
          │                        │
          │    ┌───────────────────┤
          │    │                   │
          │  origin           translations
          │  (source)          (translation)
          │                        │
          │               ┌────────▼────────┐
          │               │ tts_ready (optional)│
          │               └─────────────────┘
          │
    ┌─────▼─────┐    ┌──────────┐
    │  pause /  │◄──►│ resume   │  Operation controls
    │  resume   │    └──────────┘
    └─────┬─────┘
          │
    ┌─────▼─────┐
    │   stop    │  Stop translation
    └─────┬─────┘
          │
    ┌─────▼──────────┐
    │  task_complete  │  Task complete (with task_id)
    └─────┬──────────┘
          │
    ┌─────▼─────┐
    │  summary  │  Summary generation (if a template is set)
    └───────────┘

DocumentDescription
AuthenticationDetailed description of API Key and Ticket authentication
Voice Translation ReferenceComplete API specification for all actions
Response Events ReferenceReference for all response event formats
History and PlaybackHow to load history after stopping
TTS Speech SynthesisComplete guide to the TTS feature
Speaker ManagementRenaming, reassigning, and merging speakers
Pricing GuideBilling rules for each feature (including the multi-channel surcharge)

Version: V1.24.1 Last Updated: 2026-10-07

Copyright © 2026