Voice Translation Actions
Overview
A complete list of all actions available under the voice-translation type. For connection and authentication, see Connection and Authentication; for response event formats, see Response Events.
Table of Contents
- start - Start Voice Translation
- config - Configure Terminology / Correction Rules
- audio - Send Audio
- pause - Pause Translation
- resume - Resume Translation
- stop - Stop Translation
- retranslate - Retranslate a Single Sentence
- switch_language - Switch Language
- set_name - Set Recording Name
- rename_speaker - Globally Rename a Speaker
- reassign_speaker - Change the Speaker of a Single Sentence
- merge_speakers - Merge Speakers
- tts_play - Play TTS
- tts_stop - Stop TTS
- tts_mode - Switch TTS Mode
- set_tts - Two-Way Translation TTS Settings
- start_speaking - Start Speaking (Manual Mode)
- stop_speaking - Stop Speaking (Manual Mode)
- switch_conversation_mode - Switch Conversation Mode
- set_speaker_language - Set Speaker Language
- set_speaking_speed - Change Speaking Speed During Recording
- add_channel - Add a Channel (Multi-Channel)
- remove_channel - Disable a Channel (Multi-Channel)
- set_channel_language - Change a Channel's Language (Multi-Channel)
- set_summary - Change Summary Settings During Recording
- broadcast_go_live - Switch to the Live Phase
- broadcast_announcement - Send an Announcement
- set_standby_message - Set the Standby Phase Message
start - Start Voice Translation
Note: Do not send glossaries inside
start. If thestartpayload carriesterminology,fuzzy_correction, ortranslation_dict, those fields are ignored (startitself still succeeds) and aconfig_ignored_in_startwarning is returned, withdetails.ignored_fieldslisting what was dropped. Always send glossaries through theconfigaction instead.
Description
Start a new voice translation session and begin processing audio according to the configured parameters.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value start |
transcription_languages | string[] | Yes | Speech recognition languages (up to 10) |
translation_languages | string[] | No | Translation target languages; multiple allowed (up to 12; empty = no translation). As of v1.6.7, transcribe and broadcast translate all specified languages in real time, with one result event per language (see the "Multi-Language Translation" example below). Two-way translation (conversation) does not apply: the server overwrites this field with the counterpart language, always a single language. record no longer supports translation as of v1.7.0 and returns 400 record_translation_not_allowed |
realtime_translation | boolean | No | Real-time translation mode (default false). true: translates word-by-word while the sentence is still being recognized (interim); false: translates only when the sentence is finalized. This flag also governs the real-time behavior of multi-language translation. Always treated as true for broadcast; has no effect for conversation, which translates finalized sentences only |
recognition_mode | string | No | Recognition mode: single (single speaker, default), multi_speaker (multiple speakers), multi_channel (multi-channel, v1.10.0; must be enabled for your environment before use — see "Multi-Channel Mode Description" below); under multi_speaker, transcription_languages must contain exactly 1 language, otherwise a diarization_multilang_conflict error is returned and the session is refused (type=conversation is exempt: two-way translation forces single-speaker mode and has been exempt from this check since v1.7.2) |
type | string | Yes | Recording type: transcribe, conversation, record, broadcast |
audio_format | string | No | Audio format: pcm (default), webm |
summary_template | string | Conditional | Summary template (required for transcribe, optional for conversation/broadcast) |
options | object | No | Speech recognition options |
tts_enabled | boolean | No | Whether to enable TTS speech synthesis (default false) |
tts_language | string | No | TTS output language (must be in translation_languages) |
tts_voice | string | No | TTS voice name (e.g. en-US-JennyNeural) |
tts_mode | string | No | TTS playback mode: sync (synchronous, default), async (asynchronous). Only these two lowercase values are accepted, and an empty string is treated as not provided (sync); any other value (for example "Async") is rejected (invalid_parameter, details.field is tts_mode) and the recording does not start. Checked for every recording type, whether or not TTS is enabled |
broadcast_token | string | Conditional | Broadcast token (required for broadcast type, obtained from the REST API). Allowed only with the broadcast type: any other type that carries it is rejected (invalid_parameter, details.field is broadcast_token) and the recording does not start |
active_language | string | No | Initial active language in two-way translation mode (default transcription_languages[0]) |
tts_config | object | No | Multi-language TTS settings (broadcast / two-way translation mode) |
broadcast_phase | string | No | Initial broadcast phase: standby, live (default). Only these two lowercase values are accepted, and an empty string is treated as not provided (live); any other value (for example "Live") is rejected (invalid_parameter, details.field is broadcast_phase) and the recording does not start |
standby_message | string | No | Message viewers see during the standby phase (default: "Preparing, please wait...") |
name | string | No | Initial default recording name (max 60 characters after trimming leading and trailing whitespace; the system may still override it; if not provided, one is generated automatically, e.g. Transcription #1). A longer name is rejected (invalid_parameter, details.field is name) and the recording does not start |
summary_language | string | No | Summary output language (defaults to the recognition language when not specified; in broadcast mode it is read automatically from the channel settings). Up to 20 characters; a longer value is rejected (invalid_parameter, details.field is summary_language) and the recording does not start |
summary_mode | string | No | Summary mode enum: builtin (apply the built-in template, default) / custom (the customer prompt fully replaces the default). When omitted, builtin is inferred automatically |
summary_prompt | string | No | Required in custom mode (a value with only whitespace counts as not provided); treated as supplementary instructions in builtin mode. ≤3000 characters |
summary_prompt_slug | string | No | Required in custom mode (a value with only whitespace counts as not provided); must not be provided in builtin mode. The customer's own identifier (≤64 characters, Unicode, no control characters; passed through and stored in the backend record for historical lookup) |
summary_plain_text | boolean | No | Request plain-text summary output (default false; when enabled, the backend performs Markdown post-processing) |
speakers | object[] | No | Speaker language settings for two-way translation mode (exactly 2 entries when provided, see below). When omitted, speaker 1 uses transcription_languages[0] and speaker 2 uses transcription_languages[1] |
conversation_mode | string | No | Two-way conversation mode: auto (automatic detection, default), manual (manual PTT). Only these two lowercase values are accepted, and an empty string is treated as not provided (auto); any other value is rejected (invalid_parameter, details.field is conversation_mode) and the recording does not start. Also checked for recording types other than two-way |
channel_mode | string | Conditional | Multi-channel sub-mode (required for multi_channel, v1.10.0): per_channel (each channel is recognized independently) or shared (channels take turns speaking and share one recognition stream, v1.21.0); other values return invalid_channel_mode |
channels | object[] | Conditional | Multi-channel channel list (required for multi_channel, 1–8 channels, including the main speaker, conventionally channel_id: 1; see "Multi-Channel Mode Description" below for the fields) |
silenceTimeoutSeconds | integer | No | How many consecutive seconds without detected speech end the recording automatically: omitted or null uses the default (900 seconds); 0 means this session never ends for lack of speech; otherwise it must be an integer from 60 to 86400, and any other value is rejected. See Automatic End After a Long Silence below |
options Sub-fields
options is the speech recognition options object; all fields are optional and use their respective defaults when omitted.
| Field | Type | Default | Description |
|---|---|---|---|
speaking_speed | string | normal | Speaking speed, which affects the silence threshold for sentence segmentation: very_slow / slow / normal / fast / very_fast (see speaking_speed levels below for each threshold; the default normal is 800ms). Use a slower setting for slower speakers (longer threshold, avoids cutting on mid-sentence pauses); a faster setting segments sooner. Can be adjusted dynamically during recording via set_speaking_speed. Only these five lowercase values are accepted, and an empty string is treated as not provided (normal); any other value is rejected (invalid_parameter, details.field is options.speaking_speed) and the recording does not start |
profanity_handling | string | mask | Profanity handling: mask (mask with ***) / remove (remove) / show (show original). Only these three lowercase values are accepted, and an empty string is treated as not provided (mask); any other value is rejected (invalid_parameter, details.field is options.profanity_handling) and the recording does not start |
Note: These options apply to STT sentence segmentation; multi-speaker mode (
multi_speaker) does not currently applyspeaking_speed.
speaking_speed levels
A sentence ends once silence lasts longer than the threshold. A longer threshold tolerates pauses within a sentence (a sentence is less likely to be split in two), but each sentence appears later.
| Level | Silence threshold for ending a sentence |
|---|---|
very_fast | 300ms |
fast | 600ms |
normal (default) | 800ms |
slow | 1200ms |
very_slow | 1500ms |
Request Example (Basic)
{
"type": "voice-translation",
"data": {
"action": "start",
"transcription_languages": ["zh-TW"],
"translation_languages": ["en-US"],
"realtime_translation": false,
"type": "transcribe",
"audio_format": "pcm",
"summary_template": "meeting",
"options": {
"speaking_speed": "normal",
"profanity_handling": "mask"
}
}
}
Request Example (Multi-Language Translation, v1.6.7)
Specify multiple languages in translation_languages (up to 12) to translate into several languages at once. This applies to transcribe and broadcast (for two-way translation the language is set by the server and is always a single language; record does not support translation as of v1.7.0):
{
"type": "voice-translation",
"data": {
"action": "start",
"transcription_languages": ["zh-TW"],
"translation_languages": ["en-US", "ja-JP", "ko-KR"],
"realtime_translation": true,
"type": "transcribe",
"summary_template": "meeting"
}
}
How results arrive: each language returns its own independent result event (same sid, with a single language key inside translations). Multiple languages are never merged into one event. Clients must accumulate translations by "sid + language code" instead of overwriting. Completion order across languages is not fixed (translations run in parallel), and if one language fails the remaining languages are still delivered (the failed language additionally receives an error event whose details.translation_language identifies it).
Billing reminder: translation is billed per language starting from the second language (see the Pricing Guide). Specifying N languages is billed as N languages.
Request Example (Initial Default Name)
{
"type": "voice-translation",
"data": {
"action": "start",
"transcription_languages": ["zh-TW"],
"translation_languages": ["en-US"],
"type": "transcribe",
"audio_format": "pcm",
"summary_template": "meeting",
"name": "Product Planning Meeting"
}
}
Recording Name Rules
| Scenario | Name | name_source | Overridden by system? |
|---|---|---|---|
start with a name parameter | Initial default name | default | Yes |
start without a name | Auto-generated (e.g. Transcription #1, Broadcast #3) | default | Yes |
Set via set_name | Name explicitly set by the user | user | No |
| Auto-generated by the system after the session ends | Summary name generated from the transcript content | llm | — |
Note: The
nameinstartis an initial default name; the system may still override it when the session ends. If you need a fixed name, useset_name.
Default name formats (fixed English):
| Recording Type | Default Name Format |
|---|---|
transcribe | Transcription #N |
conversation | Conversation #N |
record | Recording #N |
broadcast | Broadcast #N |
Nis the sequential number of recordings of the same type for that user. Name priority:user>llm>default. Once the user sets a name, the system will not override it when the session ends.
Request Example (with TTS)
{
"type": "voice-translation",
"data": {
"action": "start",
"transcription_languages": ["zh-TW"],
"translation_languages": ["en-US"],
"realtime_translation": true,
"type": "transcribe",
"tts_enabled": true,
"tts_language": "en-US",
"tts_voice": "en-US-JennyNeural",
"tts_mode": "sync"
}
}
Request Example (Two-Way Translation Mode - Automatic Detection)
{
"type": "voice-translation",
"data": {
"action": "start",
"type": "conversation",
"transcription_languages": ["zh-TW", "en-US"],
"active_language": "zh-TW",
"audio_format": "pcm",
"speakers": [
{ "id": 1, "language": "zh-TW" },
{ "id": 2, "language": "en-US" }
],
"tts_config": {
"zh-TW": { "voice": "zh-TW-HsiaoChenNeural", "speaking_rate": 1.0 },
"en-US": { "voice": "en-US-JennyNeural", "speaking_rate": 1.0 }
}
}
}
Request Example (Two-Way Translation Mode - Manual Mode)
{
"type": "voice-translation",
"data": {
"action": "start",
"type": "conversation",
"transcription_languages": ["zh-TW", "en-US"],
"conversation_mode": "manual",
"audio_format": "pcm",
"speakers": [
{ "id": 1, "language": "zh-TW" },
{ "id": 2, "language": "en-US" }
],
"tts_config": {
"zh-TW": { "voice": "zh-TW-HsiaoChenNeural", "speaking_rate": 1.0 },
"en-US": { "voice": "en-US-JennyNeural", "speaking_rate": 1.0 }
}
}
}
Special rules for two-way translation mode:
| Item | Description |
|---|---|
transcription_languages | Must contain exactly 2 languages, and they must differ |
translation_languages | Overwritten by the server: ignored even if supplied; always forced to the counterpart of the active language |
realtime_translation | Has no effect in conversation mode: translations are sent only after a sentence is finalized, regardless of whether this field is true or false |
active_language | Optional, defaults to transcription_languages[0] |
recognition_mode | Forced to single; speaker_diarization is accepted but ignored (as of v1.7.2. Before that, sending speaker_diarization=true was wrongly rejected with a 400 "speaker diarization conflicts with multiple languages") |
tts_enabled | Defaults to true; set to false to return text translation only |
tts_config | Optional; configures the TTS voice for each of the two languages; leave empty to use the default voices automatically |
summary_template | Optional; when provided, a summary is generated automatically after stopping |
speakers | Optional; specifies each user's language (exactly 2 entries when provided). When omitted, users 1 and 2 map to the two transcription_languages in order |
conversation_mode | Optional, auto (automatic detection, default) or manual (manual PTT) |
speakers field description:
| Field | Type | Required | Description |
|---|---|---|---|
id | int | Yes | User number (1 or 2) |
language | string | Yes | That user's language code (must be in transcription_languages) |
conversation_mode description:
| Mode | Description |
|---|---|
auto (default) | The system automatically detects the spoken language and segments sentences automatically |
manual | The user controls the speaking interval via start_speaking / stop_speaking; audio during that interval is merged into a single sentence |
Successful Response
After a successful start, a session_started event is returned, containing the complete initial session information. For real-time recording, the first minute is deducted first, and the event is returned once that deduction completes (see the Pricing Guide).
General recording (transcribe / conversation / record):
{
"type": "voice-translation",
"data": {
"action": "session_started",
"session_id": "550e8400-e29b-41d4-a716-446655440000",
"task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
"recording_type": "transcribe",
"recognition_mode": "single",
"message": "Speech recognition started",
"resume_token": "L0VBAwIy... (43 chars)",
"resume_grace_seconds": 45,
"server_time": 1749550000000
}
}
Broadcast mode (broadcast):
{
"type": "voice-translation",
"data": {
"action": "session_started",
"session_id": "550e8400-e29b-41d4-a716-446655440000",
"task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
"recording_type": "broadcast",
"recognition_mode": "multi_speaker",
"phase": "standby",
"viewer_count": 0,
"queue_count": 0,
"peak_viewers": 0,
"total_viewers": 0,
"message": "Speech recognition started",
"resume_token": "L0VBAwIy... (43 chars)",
"resume_grace_seconds": 45,
"server_time": 1749550000000
}
}
For response field descriptions, see the session_started event.
Recording Type Descriptions
| type | Description | Use Case |
|---|---|---|
transcribe | Speech-to-text | Meeting minutes, interview records |
conversation | Conversation log | Two-way communication, customer service dialogues |
record | Plain recording | Voice memos, quick notes |
broadcast | Broadcast / live stream | Lectures, speeches, live content |
Broadcast Mode Description (type: "broadcast")
In broadcast mode, the language settings are obtained automatically from the broadcast channel settings and do not need to be sent in the WebSocket message.
Required parameters:
| Parameter | Type | Description |
|---|---|---|
type | string | Must be "broadcast" |
broadcast_token | string | Broadcast token (obtained after creating a broadcast via the REST API) |
audio_format | string | Audio format (pcm or webm) |
Optional parameters (override broadcast channel settings):
| Parameter | Type | Description |
|---|---|---|
tts_config | object | Multi-language TTS settings (override the settings used at creation) |
summary_template | string | Summary template slug (overrides the settings used at creation; if not provided, the broadcast channel default is used) |
Automatically configured parameters (can be omitted):
transcription_languages: read automatically from the broadcast settingstranslation_languages: read automatically from the broadcast settingsrealtime_translation: always enabled in broadcast mode, andfalseis treated astrue; translation is always billed at the real-time translation ratesummary_template: read automatically from the broadcast settings (the value passed via WebSocket takes precedence)summary_language: read automatically from the broadcast settings (the value passed via WebSocket takes precedence)
Both the host and viewers receive interim translations: a sentence's translation first arrives with
is_final: false, followed by the finalized version withis_final: true. Overwrite the display bysidpluslanguage; to show only finalized translations, skipis_final: false.The recording name does not reuse the channel name. If
startomitsname, a name such asBroadcast #1is generated; naming otherwise works the same as for other types, see Recording Name Rules.
Broadcast phase description:
| broadcast_phase | Description | Behavior |
|---|---|---|
live (default) | Live phase | STT/translation results are broadcast to viewers and written to the transcript |
standby | Standby phase | STT/translation results go only to the host; viewers see the standby_message |
Purpose of the standby phase: Lets the host run STT/translation warm-up tests before going live, confirming the equipment works before switching to the live phase.
The standby phase has a time limit (30 minutes by default): when the accumulated standby time reaches the limit, the session ends automatically, with a warning about 2 minutes before. See Standby Time Limit.
broadcast_phaseaccepts only lowercasestandbyandlive: an empty string is treated aslive, and any other value is rejected (invalid_parameter).
Broadcast mode request example:
{
"type": "voice-translation",
"data": {
"action": "start",
"type": "broadcast",
"broadcast_token": "a3f9",
"audio_format": "pcm"
}
}
Broadcast mode request example (standby phase + override summary template):
{
"type": "voice-translation",
"data": {
"action": "start",
"type": "broadcast",
"broadcast_token": "a3f9",
"audio_format": "pcm",
"broadcast_phase": "standby",
"standby_message": "The talk is about to begin, please wait...",
"summary_template": "lecture"
}
}
Summary template priority: the value passed in the WebSocket
start> the default set when creating the broadcast channel. If neither is set, no summary is generated automatically.
Broadcast mode TTS settings (tts_config):
Use the tts_config parameter to specify which translation languages should produce TTS audio for viewers.
| tts_config field | Type | Description |
|---|---|---|
| voice | string | TTS voice name |
| speaking_rate | number | Speaking rate (0.5–2.0, default 1.0). Values outside the range are adjusted to the nearest bound |
{
"type": "voice-translation",
"data": {
"action": "start",
"type": "broadcast",
"broadcast_token": "a3f9",
"audio_format": "pcm",
"tts_config": {
"en-US": {
"voice": "en-US-JennyNeural",
"speaking_rate": 1.0
},
"ja-JP": {
"voice": "ja-JP-NanamiNeural",
"speaking_rate": 1.0
}
}
}
}
Note:
- The TTS language must be a valid language in
translation_languages; invalid languages are ignored automatically- The host (WebSocket) does not receive TTS audio; only SSE viewers receive the
tts_readyevent- TTS is sent only during the
livephase; it is not sent during thestandbyphase
Multi-Channel Mode Description (recognition_mode: "multi_channel")
A single recording session takes input from multiple physical microphones at the same time (one per person), and speaker identity is determined by the channel — whichever microphone captured the audio identifies the speaker; the system performs no speaker inference. Suitable for meetings, round-table discussions, and other settings where every participant has a dedicated microphone (v1.10.0).
There are two sub-modes (channel_mode):
| Sub-mode | Recognition | Language | Best for |
|---|---|---|---|
per_channel | Each channel is recognized independently; people can speak at the same time | Each channel is bound to exactly 1 language; channels can differ | Several people may speak at once, each in a different language |
shared (v1.21.0) | Channels take turns speaking and share one recognition stream | The whole session shares the session-level setting | Only one person speaks at a time (a host and guests taking turns, etc.) |
Scope and hard limits:
| Item | Rule |
|---|---|
| Feature activation | Must be enabled before use; sending it in an environment where it is not enabled returns invalid_recognition_mode |
| Recording types | Only transcribe and record; conversation returns invalid_parameter and broadcast returns multichannel_broadcast_not_allowed |
| Audio format | Only pcm (16kHz / 16-bit / mono / little-endian); other values return multichannel_requires_pcm |
| TTS | Not supported; tts_enabled: true returns multichannel_tts_not_allowed |
speaker_diarization | Cannot be specified at the same time (multi-channel is itself a form of speaker diarization); providing both returns invalid_parameter |
switch_language | Disabled in multi-channel mode (languages are bound to channels); always returns multichannel_switch_language_not_allowed — under per_channel, use set_channel_language instead |
| Audio file length cap | The audio saved for one multi-channel recording has a total cap, reached sooner the more channels are open (about 70 minutes with 8 channels). Once the cap is reached, the audio file and the recording duration (duration_ms) stop at that point, while the transcript and credit charges continue as usual |
channels field description:
| Field | Type | Required | Description |
|---|---|---|---|
channel_id | int | Yes | Channel number, range 1–8, must be unique. Becomes that channel's speaker speaker_id (format channel_{N}); by convention the main speaker uses 1 |
speaker_name | string | No | Display label for that channel's speaker (max 100 characters, no control characters). If not provided, it can be set later during recording with rename_speaker |
transcription_languages | string[] | Conditional | Transcription language for that channel. Required under per_channel, exactly 1 (each channel is bound to one language); must not be provided under shared (returns channel_language_not_allowed) |
Language consistency rule (
per_channel): the union of all channel languages (deduplicated) must exactly match the session-leveltranscription_languages; otherwisechannel_language_mismatchis returned (detailslists both language sets so you can reconcile them).
Multi-channel request example:
{
"type": "voice-translation",
"data": {
"action": "start",
"type": "transcribe",
"recognition_mode": "multi_channel",
"channel_mode": "per_channel",
"transcription_languages": ["zh-TW", "ja-JP"],
"translation_languages": ["en-US"],
"audio_format": "pcm",
"summary_template": "meeting",
"channels": [
{ "channel_id": 1, "speaker_name": "Main Speaker", "transcription_languages": ["zh-TW"] },
{ "channel_id": 2, "speaker_name": "Manager Wang", "transcription_languages": ["zh-TW"] },
{ "channel_id": 3, "speaker_name": "Sato", "transcription_languages": ["ja-JP"] }
]
}
}
Successful response: the data of session_started includes channel_mode and channels[] (each entry with channel_id, speaker_name, transcription_languages, status; transcription_languages is not present in shared mode). Each channel starts with status preparing and transitions to ready via a channel_status event once that channel recognizes its first sentence.
Note: Always check that
session_startedincludeschannel_mode: this is the signal that the server actually started in multi-channel mode. If the response has nochannel_mode, the server does not support multi-channel —channels[]was ignored and the session is actually running single-channel.
Other multi-channel behavior:
- Every
audioframe must includechannel_id(see audio) - During recording you can add and disable channels dynamically via add_channel / remove_channel, and change a single channel's language via set_channel_language;
set_speaking_speedalso supports multi-channel (see that section) rename_speakerworks normally (a channel can be renamed before it is first used). The new name must not duplicate another channel's name (including the names set for each channel instart); duplicates returnspeaker_name_duplicate.reassign_speakerandmerge_speakersdo not apply to multi-channel recordings (speaker identity is determined by the channel, so there is nothing to reassign) and returnspeaker_op_not_allowed_multi_channelpause/resume: while paused, the recording file is still saved but no transcript is produced; after resuming, speech from the paused period is transcribed retroactively (timestamps reflect the actual speaking time). Retroactive transcription has a limit — a trailing 60 seconds shared across the whole session (underper_channel, split evenly across channels when there are several; undershared, the whole last 60 seconds, for all channels together, with speakers labeled); anything beyond that is kept only in the audio file. A sentence cut off at the moment of pausing may not appear in the transcript (same as single-channel recording); when it is indeed not kept, resuming also sends segment_discarded (reason: "resumed") to identify it- The
originofresultevents includeschannel_idandspeaker_id(formatchannel_{N}); every transcript sentence also carrieschannel_id(see Response Events) - Unlimited plans must include the multi-channel feature; a plan may also cap the number of channels — exceeding the cap returns
plan_feature_not_allowedon the spot atstart/add_channel(details.fieldismax_stt_streams) - Billing: speech recognition plus speaker diarization as the base;
per_channeladds a surcharge based on the number of active channels at the time, whilesharedis always counted as 1 channel — see the Pricing Guide
Shared mode (v1.21.0):
Request example:
{
"type": "voice-translation",
"data": {
"action": "start",
"type": "transcribe",
"recognition_mode": "multi_channel",
"channel_mode": "shared",
"transcription_languages": ["zh-TW", "en-US"],
"audio_format": "pcm",
"channels": [
{ "channel_id": 1, "speaker_name": "Host" },
{ "channel_id": 2, "speaker_name": "Guest A" },
{ "channel_id": 3, "speaker_name": "Guest B" }
]
}
}
Client requirements:
- Send only one channel at a time: send
audioonly on the current speaker's channel. If several channels are sent at once, their audio is queued into the same recognition stream in the order received and the transcript becomes garbled (billing is not affected). - Keep sending silence: when nobody is speaking, keep sending silence on the current channel (one frame every 100ms recommended); do not stop sending.
pcmonly (same as the common multi-channel rule).- Languages are shared across the session: channels must not specify
transcription_languages(returnschannel_language_not_allowed); the language cannot be changed during the recording (set_channel_languagereturnschannel_language_not_allowed). - The first channel cannot be removed: the first entry in
channels[]carries the recognition for the whole session, andremove_channelreturnschannel_remove_not_allowed; other channels can be removed.
Channel status:
- The
statusof the other channels follows the first channel: when the first channel turnsready, all channels turnreadytogether; when the first channel prepares recognition again (resume after pause, resume after disconnect, automatic reconnection), all channels return topreparingtogether, and each channel receives its own channel_status event with the samereasonas the first channel. After resuming from a disconnect,preparingis reported in theresume_oksnapshot rather than by a separate event. - A channel added with
add_channeltakes on the first channel's current status directly;removedfollows each channel's own status. - When the first channel turns
error, the other channels turnerroras well (with the samereason); when the first channel recovers, they recover together. - The
channels[]snapshots insession_startedandresume_okfollow the same rules; channel entries do not carrytranscription_languages.
Known limitations:
- When the gap between speakers is shorter than about 0.8 seconds, the words of the two people may be merged into one sentence labeled with only one speaker.
- Speakers are determined by channel labeling, which is highly accurate when people take turns;
reassign_speakerandmerge_speakersdo not apply, so this cannot be corrected afterward.
TTS Playback Mode Description
| Mode | Description | Behavior |
|---|---|---|
sync | Synchronous mode (default) | Automatically plays the most recent is_final=true translated sentence; if the previous sentence is still playing, it enters the queue and waits |
async | Asynchronous mode (manual control) | The user can select any translated sentence for TTS, controlled with the tts_play command |
Automatic End After a Long Silence
A recording that goes a certain length of time without recognizing any text (including interim results that are not final yet) ends automatically, so a recording someone forgot to stop does not keep being billed.
- The default threshold is 900 seconds (15 minutes) and can be changed per session with
silenceTimeoutSeconds. About 2 minutes before the end, a warning is sent once. - Any recognized text restarts the count from 0;
resumeandstart_speakingalso restart it. - The count does not run in these cases:
- While paused: the count restarts from 0 after resuming. A paused recording is still billed; to stop billing, end the recording.
- While a disconnected session waits to be resumed: after resuming, the count continues from where it was.
- Broadcasts (
type: "broadcast"): never end for lack of speech. The standby phase has its own time limit; see the Broadcast Guide.
- Multi-channel: counts only when no channel recognizes any text; a single silent channel does not end the recording.
- Conversation manual mode: speech recognized while the speak button is not pressed also restarts the count.
Values of silenceTimeoutSeconds (top level of the start data, optional)
| Value | Effect |
|---|---|
Omitted, or null | Uses the default threshold |
0 | This session never ends automatically for lack of speech; suited to sessions kept open for long periods during which nobody may speak |
| Integer from 60 to 86400 | The threshold in seconds for this session |
| Any other value | start is rejected with the error code invalid_parameter (details.field is silenceTimeoutSeconds) and the recording does not start |
- The value must be a JSON integer: strings (for example
"900"), decimal notations (for example900.0or1e3), and booleans are all rejected. - The parameter name is
silenceTimeoutSeconds; sendingsilence_timeout_secondsis rejected (details.fieldissilence_timeout_seconds), not ignored. - A valid value has no effect on broadcasts, but an invalid value is still rejected.
- With a threshold of 120 seconds or less, no warning is sent; the recording simply ends when the time is up.
- Resuming a disconnected session keeps the original setting; send it again when starting a new recording.
Events you receive
stt_silence_warning(anerrorevent withseverity: "warning"): a warning; the recording continues.details.silenceSecondsis how long the silence has lasted anddetails.remainingSecondsis how many seconds remain. Do not treat it as the end of the recording; recognized text or resuming the recording restarts the count.stt_silence_timeout(severity: "fatal"): the recording has ended automatically.details.silence_secondsis the threshold in seconds.- Then
status: "ended"andtask_complete, in that order, just as when the client sendsstop: the recording is saved and summarized as usual, and billing runs until the end.
{
"type": "error",
"data": {
"error_code": "stt_silence_warning",
"severity": "warning",
"message": "No speech detected for a while; the recording will end automatically soon",
"context": "stt",
"request_id": "req_abc123xyz789",
"timestamp": "2026-09-25T10:28:00.000Z",
"details": {
"silenceSeconds": 780,
"remainingSeconds": 120
}
}
}
After
stt_silence_timeout, stop sending audio. Each audio message that arrives afterwards gets asession_not_startedreply, which you can ignore.
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
missing_transcription_languages | 400 | No language parameter provided | Make sure the request includes transcription_languages |
invalid_transcription_language | 400 | Invalid language code | Make sure the language code format is correct (e.g. zh-TW) |
too_many_languages | 400 | Number of languages exceeds the limit | Up to 10 transcription languages and 12 translation languages |
invalid_recording_type | 400 | Invalid recording type | Use a valid type value |
audio_format_unsupported | 400 | audio_format is not supported; details.supported_formats lists the accepted values | Switch to a supported audio format |
invalid_summary_template | 400 | Invalid summary template | Make sure the template identifier is correct |
stt_init_failed | 503 | Service initialization failed | Retry later |
auth_insufficient_credit | 402 | Insufficient credit | Top up your credit balance |
auth_quota_exceeded | 402 | Available credits are insufficient; the recording did not start (less than one minute for real-time recording; the connection is not closed; details.remaining_budget is the available credit as of the most recent settlement, and details.budget_scope states whose credit it is) | Top up and start again |
daily_limit_reached | — | Usage has reached the plan's limit; the recording did not start | Available again after the plan's reset (daily limits reset the next day) |
auth_service_error | 500 | Service temporarily unavailable; the recording did not start (the connection is not closed) | start again later |
service_shutdown | — | The service is shutting down; the recording did not start, and the connection is then closed | Reconnect later and start again |
tts_init_failed | 503 | TTS service initialization failed | Retry later |
tts_invalid_language | 400 | TTS language is not in the translation languages | Make sure tts_language is in translation_languages |
broadcast_token_required | 400 | Broadcast mode requires a token | A broadcast type must provide broadcast_token |
broadcast_token_invalid | 401 | Invalid broadcast token | Make sure the token is correct and has not expired |
broadcast_not_ready | 503 | Broadcast service not yet started | Retry later |
summary_invalid_mode | 400 | summary_mode is not builtin / custom | Change to a valid mode |
summary_mode_field_mismatch | 400 | The mode and field combination do not match (a required field is missing / a forbidden field was provided) | Adjust the fields according to the mode rules |
summary_prompt_too_long | 400 | summary_prompt exceeds 3000 characters | Shorten the custom prompt |
summary_prompt_slug_too_long | 400 | summary_prompt_slug exceeds 64 characters | Shorten the identifier |
summary_prompt_slug_invalid | 400 | summary_prompt_slug contains control characters (\n / \r / \t / \0, etc.) | Remove the control characters |
invalid_recognition_mode | 400 | Invalid recognition mode, or multi-channel is not enabled for this environment (v1.10.0) | Check the recognition_mode value; multi-channel must be enabled first |
channel_mode_required | 400 | channel_mode is missing for multi-channel (v1.10.0) | Provide channel_mode (per_channel or shared) |
invalid_channel_mode | 400 | channel_mode is not per_channel or shared; also returned when shared is not enabled in this environment (v1.10.0) | Fix channel_mode; if shared is not enabled, use per_channel |
channel_language_not_allowed | 400 | A channel carries transcription_languages in shared mode (v1.21.0) | Remove transcription_languages from each channel |
channels_required | 400 | channels is missing for multi-channel (v1.10.0) | Provide 1–8 channel configurations |
too_many_channels | 400 | The channel count exceeds the limit of 8 (v1.10.0) | Reduce the number of channels |
invalid_channel_id | 400 | channel_id is out of range (1–8) or duplicated (v1.10.0) | Fix the channel_id |
channel_language_required | 400 | A channel does not specify exactly one language (v1.10.0) | Provide exactly 1 entry in each channel's transcription_languages |
channel_language_mismatch | 400 | The union of channel languages does not match transcription_languages (v1.10.0) | Reconcile the two language lists |
speaker_name_duplicate | 422 | Two or more channels use the same speaker_name (v1.10.0) | Each channel needs a distinct name (channels without a name are unaffected) |
multichannel_requires_pcm | 400 | Multi-channel supports only the pcm audio format (v1.10.0) | Set audio_format to pcm |
multichannel_tts_not_allowed | 400 | Multi-channel does not support speech synthesis (v1.10.0) | Disable tts_enabled |
multichannel_broadcast_not_allowed | 400 | Broadcast does not support multi-channel (v1.10.0) | Use multi_speaker or single for broadcast |
invalid_parameter | 400 | Multi-channel combined with conversation or speaker_diarization, or speaker_name is too long / contains control characters (v1.10.0); an invalid value or name for silenceTimeoutSeconds, an invalid broadcast_phase value, or broadcast_token sent with a type other than broadcast (v1.18.0); name longer than 60 characters, summary_language longer than 20 characters, or a value outside the accepted list for options.speaking_speed, options.profanity_handling, conversation_mode, or tts_mode (v1.24.0; details.valid_values lists the accepted values) | Fix the parameter indicated by details.field |
plan_feature_not_allowed | 403 | The plan does not include multi-channel, or the channel count exceeds the plan's limit (details.field is max_stt_streams, v1.10.0) | Reduce channels or upgrade the plan |
config - Configure Terminology / Correction Rules
Description
Send terminology, fuzzy-word correction rules, and translation dictionary settings before or during a recording. These settings can improve STT accuracy, fix homophone errors, and ensure translation consistency.
Terminology also drives homophone correction: When terminology is provided, the terms become the reference for homophone matching — any span in the transcript that sounds the same but is written differently is corrected back to the spelling of the term. Providing terminology alone is therefore enough to get correction; you do not need to list possible misspellings by hand.
For the full description of homophone correction (the reference table, the fact that it applies to Chinese only, common-word protection, and the pronunciation limits), see Terminology Guide → How Terminology Participates in Homophone Correction.
Limits
Per-block entry limits, length limits, and the matching error codes are collected in Terminology Guide → Limits.
The numbers are defaults; the limit actually in force can be tuned per environment — always treat
maxin the error response'sdetailsas authoritative rather than hard-coding the numbers.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value config |
terminology | object | No | Terminology settings |
fuzzy_correction | object | No | Fuzzy-word correction rules |
translation_dict | object | No | Translation dictionary |
Note: At least one setting item must be provided.
Note: All or nothing: all three blocks are validated before any of them is applied. If any block fails validation, an error is returned and none of the three blocks is applied — the settings stay as they were.
For example, sending a valid terminology list together with more than 3000 translation-dictionary entries returns
config_too_many_dict_entries, and the terminology does not take effect either. Fix the problem and resend the completeconfig; all three blocks replace their previous value wholesale, so resending does not stack on top of earlier settings.
Terminology Format (terminology)
Keyed by language code, with an array of terms as the value:
{
"zh-TW": [
{ "term": "語者分離" },
{ "term": "WebSocket" }
],
"en-US": [
{ "term": "diarization" }
]
}
| Field | Type | Required | Description |
|---|---|---|---|
term | string | Yes | The term (max 100 characters) |
Limit: Up to 500 terms across all languages combined in one
config— not 500 per language. Exceeding the limit returnsconfig_too_many_entries(withcountandmaxindetails).Language scope: Terms apply only to recognition in the language they are registered under. Single-language situations — each multi-channel track, speaker diarization, and file import — use only that language's terms; multi-language transcription and conversation mode use the terms for the languages declared for the session. Terms registered under a language not used in this session have no effect, but still count toward the 500 combined total above.
These numbers are defaults: the limit actually in force can be tuned per environment; always treat
maxin the error response'sdetailsas authoritative.
Fuzzy-Word Correction Format (fuzzy_correction)
Note: This field usually does not need to be set manually —
terminologyalready corrects misspellings that sound the same or nearly the same. Use it in these three cases:
- The misspelling is itself an ordinary word and is therefore blocked by common-word protection — 晶圓 heard as 金元, for example
- The misspelling sounds very different from the correct term, such as a foreign brand name recognized as a phonetically unrelated word
- Misspellings in Japanese, Korean or English, which do not participate in homophone matching
When
correctis Chinese, it also becomes a reference for homophone matching — homophone misspellings not listed inincorrectare corrected tocorrectas well.case_insensitiveapplies only to the literal matching ofincorrect; it does not affect homophone matching.
Keyed by language code, with an array of correction rules as the value:
{
"zh-TW": [
{ "correct": "語者分離", "incorrect": ["語這分離", "語者分力"] },
{ "correct": "IPEVO", "incorrect": ["ltfo"], "case_insensitive": true }
]
}
| Field | Type | Required | Description |
|---|---|---|---|
correct | string | Yes | The correct term |
incorrect | string[] | Conditional | List of incorrect variants, each up to 200 characters. Can be omitted for Chinese terms (see below); required otherwise — an empty array returns config_invalid_entry (reason: "empty") |
case_insensitive | boolean | No | Whether this rule's variants match regardless of case (defaults to false = exact-case matching) |
Supplying only the correct term: when
correctis Chinese (contains Han characters),incorrectmay be omitted entirely — the system matches by pronunciation, and spellings in the transcript that sound the same or nearly the same are corrected back tocorrect.{ "fuzzy_correction": { "zh-TW": [{ "correct": "艾思通" }] } }No misspellings need to be listed above: 愛思通, 愛時通, 愛司東 and 愛似通 are all corrected. Only spellings that sound quite different (愛自動, say) or have a different number of syllables (愛松) still need to be listed in
incorrect.Note: Both conditions must hold: the language must be Chinese (
zh-TW,zh-CN,zh-HKand so on) andcorrectmust contain Han characters. Otherwiseincorrectremains required — omitting it in those cases would have no effect at all, and accepting it would leave you believing the setting took. List misspellings explicitly for Japanese, Korean and English.
Case sensitivity:
case_insensitiveis optional and defaults tofalse(exact-case matching). When set totrue, everyincorrectvariant in that rule matches regardless of case. The flag is per rule — the samecorrectterm can be split across several rules with different settings, for example making variants that cannot collide with ordinary words case-insensitive while keeping variants that could hit a personal name exact. It has no effect on Chinese rules (Chinese has no letter case).Note: Enabling it widens the false-positive surface: if
ivois case-insensitive, the personal nameIvois replaced too.
When the same incorrect variant appears in more than one rule: collisions are resolved on
incorrect(the variant), not oncorrect.
- When several rules point at different correct terms, which one actually takes effect is not guaranteed; do not rely on any ordering (registration order included)
- The case flag resolves to strict wins — if any rule leaves
case_insensitiveoff, that variant is matched with exact case"Strict wins" is a deliberately conservative choice: it prevents a permissive rule elsewhere in your vocabulary from silently loosening a brand-name rule you explicitly set to strict.
Splitting one
correctterm across several rules is therefore safe, as long as theirincorrectvariants do not overlap. But if your data contains the same incorrect variant mapped to different correct terms, the later one is dropped with no warning — check for duplicate variants before sending.
Limit: Up to 4000 rules across all languages combined in one
config. Exceeding it returnsconfig_too_many_entries(withfield: "fuzzy_correction",countandmaxindetails).This number is a default: the limit actually in force can be tuned per environment, so always treat
maxin the error response'sdetailsas authoritative. Do not hard-code the numbers in your integration — if you want to check before sending, readmaxand fill it back in. When you do hit a limit,detailscarries bothcount(what you sent) andmax(the limit in force).Language scope: A correction rule applies only to sentences in the language it is registered under — a rule under
zh-TWwill not alter an English sentence. When the sentence language cannot be determined, all rules are applied as a fallback (better to over-apply than to skip the sentence entirely). Homophone matching driven by terminology likewise follows the language each term is registered under.
Translation Dictionary Format (translation_dict)
Group entries by language code, giving each language its own dictionary:
{
"en-US": [
{ "source": "語者分離", "target": "Speaker Diarization" },
{ "source": "晶圓", "target": "wafer", "case_sensitive": true }
],
"ja-JP": [
{ "source": "語者分離", "target": "話者分離" }
]
}
| Field | Type | Required | Description |
|---|---|---|---|
| (top-level key) | string | Yes | Target language code |
source | string | Yes | The source word (in the STT language), up to 200 characters |
target | string | Yes | The required translation for this language, up to 200 characters |
case_sensitive | boolean | No | Whether the entry applies only on an exact-case match (defaults to false = case-insensitive) |
Limit: Up to 3000 entries per language. Exceeding the limit returns
config_too_many_dict_entries; thedetailsin the response identify which language exceeded it.Note: The entry count directly affects translation workload and cost. A single translation only carries entries whose source word actually appears in that piece of text, up to 100 of them; beyond that, longer source words are kept first. Also, the more entries there are, the smaller the share that is reliably honored — an inherent limit that a higher cap does not change.
Case sensitivity:
case_sensitiveis optional and defaults tofalse(case-insensitive). When set totrue, the entry applies only where the source text matchessourceexactly, including case. The flag is per entry.Note: The translation dictionary guides the model through prompting rather than literal substitution, so it is best-effort, not deterministic — the case flag is likewise a hint and is not guaranteed to be honored. Use
fuzzy_correctionwhen you need deterministic replacement.
Case-Flag Comparison
fuzzy_correction and translation_dict each have a case switch. Their field names are opposites, and so is the behavior their default value produces:
| Block | Field | Default | Default behavior |
|---|---|---|---|
fuzzy_correction | case_insensitive | false | Strict (case-sensitive) |
translation_dict | case_sensitive | false | Permissive (case-insensitive) |
Both default to false, yet one means strict and the other means permissive. Do not share a single variable between them, and do not mirror one value onto the other — getting it wrong produces no error at all, only matching behavior opposite to what you intended.
Request Example (Recommended: terminology only)
{
"type": "voice-translation",
"data": {
"action": "config",
"terminology": {
"zh-TW": [
{ "term": "語者分離" },
{ "term": "CVD製程" },
{ "term": "wafer良率" }
]
}
}
}
Request Example (Full settings, with manual correction rules)
{
"type": "voice-translation",
"data": {
"action": "config",
"terminology": {
"zh-TW": [
{ "term": "語者分離" },
{ "term": "即時轉錄" }
]
},
"fuzzy_correction": {
"zh-TW": [
{ "correct": "語者分離", "incorrect": ["語這分離", "語者分力"] }
]
},
"translation_dict": {
"en-US": [{ "source": "語者分離", "target": "Speaker Diarization" }]
}
}
}
Successful Response
{
"type": "voice-translation",
"data": {
"action": "config_updated",
"updated": ["terminology", "fuzzy_correction", "translation_dict"],
"message": "Settings updated"
}
}
For response field descriptions, see the config_updated event.
Error Codes
Important: Your client must also listen for
type: "error"messages — do not wait only forconfig_updated.When the server rejects a configuration it sends a
type: "error"message and notconfig_updated. An integration that waits only forconfig_updatedwill hang until its own timeout and appear as "the server never responded", even though the error was delivered anddata.error_codealready states the reason.
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
config_empty | 400 | No configuration provided. Note: An empty object {} does not count as "provided" — sending {"terminology": {}, "fuzzy_correction": {}, "translation_dict": {}} triggers this error | Provide at least one setting that actually has content. To clear a language's glossary, send {"lang": []} (for example {"zh-TW": []}) |
config_term_too_long | 400 | Term exceeds 100 characters | Shorten the term |
config_too_many_entries | 400 | More than 500 terms, or more than 4000 fuzzy correction rules (both across all languages combined) | Remove terms or correction rules |
config_too_many_dict_entries | 400 | Translation dictionary exceeds 3000 entries for a single language (details.language identifies which one) | Reduce the dictionary entries for that language |
config_invalid_entry | 400 | A glossary entry has an invalid field (details carries language, index, field, and reason for locating it; depending on the case it may also carry variant_index, max_length, or count/max) | Fix the entry at the location given in details |
audio - Send Audio
Description
Send audio data to the server for speech recognition. The audio must be Base64-encoded before sending.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value audio |
payload | string | Yes | Base64-encoded audio data |
channel_id | int | Conditional | Source channel number (required on every frame in multi-channel mode, v1.10.0). Omitting it returns channel_id_required; an unknown or removed number returns unknown_channel_id. Ignored outside multi-channel mode |
Audio Format Requirements
PCM format (default):
| Item | Specification |
|---|---|
| Format | PCM (raw audio) |
| Sample rate | 16000 Hz |
| Bit depth | 16-bit |
| Channels | Mono |
| Byte order | Little-endian |
| Transfer encoding | Base64 |
WebM/Opus format:
| Item | Specification |
|---|---|
| Format | WebM container + Opus codec |
| Sample rate | Any (the server converts automatically) |
| Channels | Mono or Stereo (the server converts automatically) |
| Transfer encoding | Base64 |
Request Example
{
"type": "voice-translation",
"data": {
"action": "audio",
"payload": "Base64-encoded PCM audio data"
}
}
Request Example (Multi-Channel, v1.10.0)
In multi-channel mode every frame must include channel_id, identifying which microphone the audio came from:
{
"type": "voice-translation",
"data": {
"action": "audio",
"channel_id": 2,
"payload": "Base64-encoded PCM audio data"
}
}
Multi-channel sending tips: send frames of roughly 100ms; keep sending on every channel even while nobody is speaking (silent audio). A single silent channel does not end the recording; see Automatic End After a Long Silence.
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
session_not_started | 400 | Speech recognition has not started | Call the start action first |
audio_invalid_format | 400 | Invalid audio data format | Make sure the Base64 encoding is correct |
audio_decode_failed | 400 | Audio decoding failed | Make sure the audio format is correct. The recording continues; for WebM, send a new container (with its header) to recover. Time that cannot be decoded is not billed |
audio_process_failed | 500 | STT/diarization writes keep failing, exceeding the tolerance threshold | We recommend reconnecting |
channel_id_required | 400 | An audio frame in multi-channel mode is missing channel_id (v1.10.0) | Include channel_id on every frame |
unknown_channel_id | 400 | Unknown or removed channel_id (v1.10.0) | Make sure the channel exists and has not been removed |
pause - Pause Translation
Description
Pause speech recognition processing. Audio received while paused is cached and processing resumes afterward.
Two-way translation is an exception: audio received while paused is not kept. For a sentence in progress at the moment of pausing, the server first waits for the recognizer to finish the last part (usually about 1 second, at most about 3 seconds), sends it with is_final: true, and only then returns status: "paused"; in manual mode, if the user is speaking, speaking is ended automatically and that sentence's final result is sent.
Billing continues while paused. See Pricing.
Request Example
{
"type": "voice-translation",
"data": {
"action": "pause"
}
}
Successful Response
{
"type": "voice-translation",
"data": {
"action": "status",
"status": "paused",
"message": "Speech recognition paused"
}
}
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
session_not_started | 400 | Speech recognition has not started | Call start first |
session_already_paused | 400 | Already paused | You can ignore this error |
resume - Resume Translation
Description
Resume paused speech recognition processing.
Request Example
{
"type": "voice-translation",
"data": {
"action": "resume"
}
}
Successful Response
{
"type": "voice-translation",
"data": {
"action": "status",
"status": "live",
"message": "Speech recognition resumed"
}
}
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
session_not_started | 400 | Speech recognition has not started | Call start first |
session_not_paused | 400 | Not paused | You can ignore this error |
If a sentence is in progress when this happens, it cannot be kept and is reported separately via segment_discarded (
reason: "resumed"). In multi-channel mode each channel reports separately.
stop - Stop Translation
Description
Stop speech recognition and end the session. The server first waits for the recognizer to finish the last sentence (usually about 1 second, at most about 3 seconds) so that it is included in the transcript; if it does not arrive in time, the last interim result shown is used as that sentence. In two-way translation manual mode, if the user is speaking, that sentence's final result and translation are sent first. The system automatically uploads the audio file and transcript and generates a summary (if configured; when the available credits cannot cover the summary fee, no summary is generated and summary_error is sent instead).
Request Example
{
"type": "voice-translation",
"data": {
"action": "stop"
}
}
Successful Response
{
"type": "voice-translation",
"data": {
"action": "status",
"status": "ended",
"message": "Speech recognition stopped"
}
}
After stopping, once the audio file and transcript have finished uploading, you receive a task_complete event containing task_id (the Recording UUID).
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
session_not_started | 400 | The session has not started, or this recording has already ended (for example, stop sent twice) | Call start first if it has not started; if this was a duplicate stop, the error can be ignored |
Sending
stopa second time returns this error rather than another success response.task_completeis sent only once, after the first successful stop, when the audio and transcript have finished uploading.
retranslate - Retranslate a Single Sentence
Description
Retranslate a specified sentence. This is useful when the source text has been corrected and the translation needs to be updated.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value retranslate |
sid | int | Yes | The number of the sentence to retranslate |
translation_languages | string[] | Yes | Array of translation language codes. Only the first element is retranslated; any additional languages are ignored. In multi-language sessions, send a separate retranslate request per language |
text | string | Yes | The source text to translate (the user's corrected text) |
Request Example
{
"type": "voice-translation",
"data": {
"action": "retranslate",
"sid": 1,
"translation_languages": ["en-US"],
"text": "The user's corrected source text"
}
}
Successful Response
A translation event is returned (sharing the same schema as normal translation results), and the translation result includes is_retranslation: true:
{
"type": "voice-translation",
"data": {
"action": "translation",
"sid": 1,
"translations": {
"en-US": {
"sid": 1,
"text": "The new translation result",
"is_final": true,
"is_retranslation": true
}
}
}
}
The
dataobject also carriessid, identical to thesidinside each language oftranslations.
v1.5.6 documentation correction: Earlier documentation described retranslate returning
action: "result", but on the wire it is actuallyaction: "translation". If your client's dispatcher originally handled theresultaction, add a handler branch for thetranslationaction.
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
invalid_data | 422 | No sid provided | Include sid |
record_translation_not_allowed | 400 | Recording-only sessions do not support translation | Use the transcribe type |
retranslate_session_not_active | 400 | The session is not started or has ended | Check the session status |
retranslate_no_target_lang | 400 | No target language provided | Provide translation_languages |
retranslate_no_text | 400 | No text to translate provided | Provide the text parameter |
retranslate_llm_not_ready | 503 | The translation service is not ready | Retry later |
retranslate_llm_failed | 500 | Translation service failed | Retry later |
If the translation service returns a specific failure code (for example
llm_content_filteredwhen the content cannot be translated), that code is returned as-is instead of being wrapped inretranslate_llm_failed.
switch_language - Switch Language
Description
Switch or adjust translation languages during real-time translation. The behavior depends on the recording type and the number of translation languages:
- General mode, single language (
translation_languageshas 1 entry): replaces the translation target language and automatically batch-retranslates all already-translated sentences - General mode, multiple languages (2 or more entries, v1.6.7): redefined as "add or remove a single language". The
opparameter is required; omitting it returns aswitch_language_op_requirederror - Two-way translation mode (conversation): switches the STT source language (the spoken language); the translation target switches automatically to the other language
- Multi-channel mode (
multi_channel, v1.10.0): disabled. Languages are bound to channels, so every form ofswitch_language(includingop: "add"/op: "remove") returnsmultichannel_switch_language_not_allowed; to change a single channel's language, use set_channel_language instead
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value switch_language |
translation_languages | string[] | Conditional | Array of translation language codes (required in general mode; only the first element is used as the operation target) |
op | string | Conditional | v1.6.7 multi-language operation: add (add a language) or remove (remove a language). Required for multi-language sessions; single-language sessions may use add to expand into multi-language, or omit it to keep the existing replace semantics |
transcription_languages | string[] | Conditional | The target language to switch to (two-way translation mode; if omitted, automatically toggles to the other language) |
Request Example (General Mode, Single-Language Replace)
{
"type": "voice-translation",
"data": {
"action": "switch_language",
"translation_languages": ["ja-JP"]
}
}
Request Example (Multi-Language, Add a Language, v1.6.7)
{
"type": "voice-translation",
"data": {
"action": "switch_language",
"op": "add",
"translation_languages": ["de-DE"]
}
}
After a successful add: existing sentences are automatically backfilled with the new language (the response sequence is the same as single-language replace: language_switch_start → multiple batch_retranslation → language_switch_done), and subsequent sentences include the new language in real-time translation. The limit is 12 languages; exceeding it returns too_many_languages.
Billing reminder: after adding a language, billing reflects the new language count starting from the next billed minute.
Request Example (Multi-Language, Remove a Language, v1.6.7)
{
"type": "voice-translation",
"data": {
"action": "switch_language",
"op": "remove",
"translation_languages": ["ko-KR"]
}
}
After a successful remove, a translation_language_removed event is returned. Existing translations for that language are kept; subsequent sentences are no longer translated into it. At least 1 translation language must remain (removing the last one returns a switch_language_last_language error):
{
"type": "voice-translation",
"data": {
"action": "translation_language_removed",
"translation_language": "ko-KR",
"translation_languages": ["en-US", "ja-JP"]
}
}
translation_languagesis an authoritative snapshot of the full translation-language set after removal; overwrite your local language set with it directly (see events.md).
Request Example (Two-Way Translation Mode)
Specify the switch target:
{
"type": "voice-translation",
"data": {
"action": "switch_language",
"transcription_languages": ["en-US"]
}
}
Automatic toggle (no parameters):
{
"type": "voice-translation",
"data": {
"action": "switch_language"
}
}
Special behavior in two-way translation mode:
- Two-way translation mode uses automatic language detection, so you usually don't need to switch the language manually
switch_languageonly updates the internal preference state- After a successful switch, a language_switched event is returned (not the language_switch_start/done sequence)
- Switching to the same language returns a
conversation_same_languagewarning
Response Sequence (General Mode)
After switching the language, you receive the following events in order:
- language_switch_start: notifies that the switch has started
{
"type": "voice-translation",
"data": {
"action": "language_switch_start",
"translation_language": "ja-JP",
"translation_languages": ["en-US", "ja-JP"],
"total_segments": 15
}
}
- batch_retranslation (multiple): returns retranslation results sentence by sentence
{
"type": "voice-translation",
"data": {
"action": "batch_retranslation",
"sid": 3,
"translations": {
"ja-JP": {
"sid": 3,
"text": "今日はプロジェクトの進捗について話し合いましょう",
"is_final": true,
"is_retranslation": true
}
}
}
}
- language_switch_done: notifies that the switch is complete
{
"type": "voice-translation",
"data": {
"action": "language_switch_done",
"translation_language": "ja-JP",
"translation_languages": ["en-US", "ja-JP"],
"success_count": 15,
"failed_count": 0
}
}
Language-set sync: op:add and single-language replace emit identical events (both go through language_switch_start/done). These events carry translation_languages (the current full set of translation languages) — overwrite your local language set with it directly; do not infer append vs replace from the single translation_language (when only 1 language exists, an op:add expanding to a 2nd would be misread as a replace and drop the existing language).
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
switch_language_no_target | 400 | No target language provided | Provide translation_languages |
switch_language_in_progress | 400 | The previous switch is not yet complete | Wait for the switch to complete |
switch_language_same_target | 400 | The target language is the same as the current one | You can ignore this error |
switch_language_op_required | 400 | op is missing in a multi-language session (v1.6.7) | Provide op: "add" or op: "remove" |
switch_language_already_exists | 400 | The language to add is already in the translation list (v1.6.7) | You can ignore this warning |
switch_language_not_in_session | 400 | The language to remove is not in the translation list (v1.6.7) | Check the language code |
switch_language_last_language | 400 | At least one translation language must remain (v1.6.7) | The last language cannot be removed |
too_many_languages | 400 | Adding would exceed the limit of 12 languages (v1.6.7) | Remove an existing language first |
invalid_translation_language | 400 | The language code to add is invalid (v1.6.7) | Check the supported language list |
conversation_requires_two_languages | 400 | Two-way translation mode requires exactly two languages | Make sure transcription_languages has 2 entries |
conversation_languages_identical | 400 | The two languages in two-way translation cannot be the same | Provide two different languages |
conversation_invalid_language | 400 | Invalid two-way translation language | Make sure the language is in transcription_languages |
conversation_same_language | 400 | Already the current language | You can ignore this warning |
multichannel_switch_language_not_allowed | 400 | switch_language is disabled in multi-channel mode (v1.10.0) | Use set_channel_language to change a single channel's language |
set_name - Set Recording Name
Description
Set the name during a recording. After it is set, name_source flips to user, and the system will not override it when the recording ends (even if the LLM generates a summary name, it yields to the user-set name). For the full semantics and priority of name_source, see § Recording Name Rules above.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value set_name |
name | string | Yes | Recording name (max 60 characters) |
Request Example
{
"type": "voice-translation",
"data": {
"action": "set_name",
"name": "Product Planning Meeting"
}
}
Successful Response
{
"type": "voice-translation",
"data": {
"action": "status",
"event": "name_set",
"name": "Product Planning Meeting",
"message": "Recording name set"
}
}
Compatibility note: For backward compatibility, the
set_namesuccess response keepsaction: "status"(unchanged). New clients should identify a successfulset_nameviaevent: "name_set"(together with thenamefield). Relying onaction: "status"to detectset_namesuccess is deprecated and may be removed in a future version.
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
set_name_empty | 400 | Recording name is empty | Provide a non-empty name |
set_name_too_long | 400 | Recording name exceeds the length limit (>60 chars); the response details includes max_length | Shorten the name (≤60 characters) |
rename_speaker - Globally Rename a Speaker
Description
In multi-speaker diarization mode (multi_speaker), globally rename a speaker. All sentences that use that speaker ID are updated in sync.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value rename_speaker |
speaker_id | string | Yes | The original speaker ID (e.g. Guest-1); also accepts the current display label for consecutive renames; max 100 characters |
new_label | string | Yes | The new display label; max 100 characters, must not contain control characters (\x00-\x1F, \x7F) or line breaks |
Request Example
{
"type": "voice-translation",
"data": {
"action": "rename_speaker",
"speaker_id": "Guest-1",
"new_label": "Manager Wang"
}
}
Successful Response
Returns the speaker_renamed event:
{
"type": "voice-translation",
"data": {
"action": "speaker_renamed",
"speaker_id": "Guest-1",
"new_label": "Manager Wang",
"affected_sids": [1, 3, 5, 8]
}
}
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
speaker_not_found | 422 | The specified speaker was not found | Make sure the speaker ID or alias exists |
speaker_name_empty | 422 | The speaker name cannot be empty | Provide a valid name |
speaker_name_duplicate | 422 | The speaker name is already in use | Use another name, or first rename the conflicting speaker |
session_not_started | 400 | Speech recognition has not started | Call start first |
reassign_speaker - Change the Speaker of a Single Sentence
Description
Change the speaker identity of a specific sentence, assigning the sentence to an existing speaker.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value reassign_speaker |
sid | int | Yes | The number of the sentence to change |
target_speaker_id | string | Yes | The target speaker's original ID (taken from init_sentence.speaker_id; reassign does not accept display labels) |
Request Example
{
"type": "voice-translation",
"data": {
"action": "reassign_speaker",
"sid": 5,
"target_speaker_id": "Guest-2"
}
}
Successful Response
Returns the speaker_reassigned event:
{
"type": "voice-translation",
"data": {
"action": "speaker_reassigned",
"sid": 5,
"old_speaker_id": "Guest-1",
"new_speaker_id": "Guest-2",
"new_speaker_label": "Lisa Lee"
}
}
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
speaker_sid_not_found | 422 | The specified sentence was not found | Make sure the SID exists |
speaker_not_found | 422 | The target speaker does not exist | Use an existing speaker ID |
speaker_name_empty | 422 | The target speaker ID cannot be empty | Provide a valid speaker ID |
session_not_started | 400 | Speech recognition has not started | Call start first |
invalid_parameter | 400 | Creating a new speaker is not supported | Use an existing speaker ID |
merge_speakers - Merge Speakers
Description
Merge all sentences from one speaker into another. After merging, future recognition results from that speaker are also automatically converted to the target speaker. This applies within the current recognition pass only: after a connection recovers, speakers are re-identified and you need to merge again.
Difference from reassign_speaker
| Feature | Scope | Future Effect |
|---|---|---|
reassign_speaker | A single sentence (1 SID) | None |
merge_speakers | All sentences of that speaker | Future occurrences of the source are also automatically converted to the target |
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value merge_speakers |
source_speaker_id | string | Yes | The speaker ID to be merged (e.g. Guest-2) |
target_speaker_id | string | Yes | The target speaker ID to merge into (e.g. Guest-1) |
Request Example
{
"type": "voice-translation",
"data": {
"action": "merge_speakers",
"source_speaker_id": "Guest-2",
"target_speaker_id": "Guest-1"
}
}
Successful Response
Returns the speakers_merged event:
{
"type": "voice-translation",
"data": {
"action": "speakers_merged",
"source_speaker_id": "Guest-2",
"target_speaker_id": "Guest-1",
"affected_sids": [3, 5, 7]
}
}
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
speaker_not_found | 422 | The speaker does not exist | Make sure the speaker ID exists |
merge_speakers_same_id | 400 | The source and target speakers are the same | Use different speaker IDs |
speaker_name_empty | 422 | The speaker ID cannot be empty | Provide a valid speaker ID |
session_not_started | 400 | Speech recognition has not started | Call start first |
tts_play - Play TTS
Description
In async mode, manually play the TTS audio of a specified sentence. Repeated requests for the same sid are supported (replay).
Two-way translation mode (conversation):
tts_playautomatically synthesizes the translation in the appropriate language based on the voice settings intts_config; you don't need to specifytts_languageseparately.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value tts_play |
sid | int | Yes | The starting sentence ID |
length | int | No | The number of sentences to play (default 1, max 20) |
Request Example (Single Sentence)
{
"type": "voice-translation",
"data": {
"action": "tts_play",
"sid": 5
}
}
Request Example (Multiple Sentences)
{
"type": "voice-translation",
"data": {
"action": "tts_play",
"sid": 5,
"length": 3
}
}
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
tts_not_enabled | 400 | TTS not enabled | Make sure TTS was enabled at start |
A missing sentence or translation does not return
error: If the startingsiddoes not exist, or the sentence has no translation in the target language, noerrormessage is returned. Atts_errorevent is sent instead, witherrorset tosentence_not_foundandtranslation_not_foundrespectively. When playing multiple sentences, a failing sentence is skipped and the rest still play.
tts_stop - Stop TTS
Description
Stop the currently playing TTS audio.
Request Example
{
"type": "voice-translation",
"data": {
"action": "tts_stop"
}
}
Successful Response
{
"type": "voice-translation",
"data": {
"action": "status",
"message": "TTS stopped"
}
}
tts_mode - Switch TTS Mode
Description
Switch the TTS playback mode (synchronous/asynchronous) during a recording.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value tts_mode |
tts_mode | string | Yes | Mode: sync (synchronous) or async (asynchronous); only these two lowercase values are accepted |
Request Example
{
"type": "voice-translation",
"data": {
"action": "tts_mode",
"tts_mode": "async"
}
}
Successful Response
Returns the tts_mode_changed event:
{
"type": "voice-translation",
"data": {
"action": "tts_mode_changed",
"tts_mode": "async"
}
}
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
invalid_data | 422 | No tts_mode provided, or a value other than sync or async. For the latter, details.field is tts_mode and details.valid_values lists the accepted values; the mode does not change and no tts_mode_changed is sent | Include sync or async |
set_tts - Two-Way Translation TTS Settings
Description
During two-way translation mode (conversation), toggle TTS on/off or update the TTS voice settings mid-session. Available only under the conversation type.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value set_tts |
tts_enabled | boolean | No | Toggle TTS on/off |
tts_config | object | No | Update the TTS settings for a specific language (only the two two-way translation languages are valid) |
Request Example (Disable TTS)
{
"type": "voice-translation",
"data": {
"action": "set_tts",
"tts_enabled": false
}
}
Request Example (Update TTS Voice)
{
"type": "voice-translation",
"data": {
"action": "set_tts",
"tts_enabled": true,
"tts_config": {
"en-US": { "voice": "en-US-GuyNeural", "speaking_rate": 1.2 }
}
}
}
Successful Response
Returns the tts_updated event:
{
"type": "voice-translation",
"data": {
"action": "tts_updated",
"tts_enabled": true,
"tts_config": {
"zh-TW": { "voice": "zh-TW-HsiaoChenNeural", "speaking_rate": 1.0 },
"en-US": { "voice": "en-US-GuyNeural", "speaking_rate": 1.2 }
}
}
}
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
invalid_action | 400 | This operation is not supported outside two-way translation mode | Use only under the conversation type |
start_speaking - Start Speaking (Manual Mode)
Description
In two-way translation manual mode (conversation_mode: "manual"), notify the system that the user has started speaking. From this moment, audio is sent to STT for recognition, and all recognition results accumulate into the same sentence (no automatic segmentation). If it is called again while already speaking, the system first ends the previous sentence (waiting for the last part to finish, then sending its final result) and starts a new one.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value start_speaking |
speaker | int | Yes | User number (1 or 2) |
Request Example
{
"type": "voice-translation",
"data": {
"action": "start_speaking",
"speaker": 1
}
}
Successful Response
{
"type": "voice-translation",
"data": {
"action": "status",
"message": "Speaking started"
}
}
If it is called again while already speaking, the final result of the previous sentence is sent as usual, followed by this same status.
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
invalid_action | 400 | Not in two-way translation mode | Use only under the conversation type |
conversation_not_manual_mode | 400 | Not in manual mode | Use only in manual mode |
conversation_invalid_speaker | 400 | Invalid user number | Use 1 or 2 |
stop_speaking - Stop Speaking (Manual Mode)
Description
In two-way translation manual mode, notify the system that the user has stopped speaking. The system first waits for the recognizer to finish the last part (usually about 1 second, at most about 3 seconds), then merges the recognition results accumulated during this period into one complete sentence, and translates it and synthesizes TTS. For languages that separate words with spaces (such as English), a space is added between parts automatically.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value stop_speaking |
Request Example
{
"type": "voice-translation",
"data": {
"action": "stop_speaking"
}
}
Successful Response
After speaking stops, the system sends a complete result event (containing origin and translations):
{
"type": "voice-translation",
"data": {
"action": "result",
"origin": {
"sid": 1,
"language": "zh-TW",
"text": "The complete sentence merged from all recognition during this period",
"is_final": true,
"speaker_id": "Speaker-1",
"start_time": "00:05"
},
"translations": {
"en-US": {
"sid": 1,
"text": "The complete merged sentence from this speaking period",
"is_final": true
}
}
}
}
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
invalid_action | 400 | Not in two-way translation mode | Use only under the conversation type |
conversation_not_speaking | 400 | Not in the speaking state | Call start_speaking first |
switch_conversation_mode - Switch Conversation Mode
Description
During two-way translation mode, switch between automatic detection mode (auto) and manual mode (manual). If the user is speaking during the switch, speaking ends automatically.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value switch_conversation_mode |
conversation_mode | string | Yes | Target mode: auto or manual |
Request Example
{
"type": "voice-translation",
"data": {
"action": "switch_conversation_mode",
"conversation_mode": "manual"
}
}
Successful Response
Returns the conversation_mode_changed event:
{
"type": "voice-translation",
"data": {
"action": "conversation_mode_changed",
"conversation_mode": "manual"
}
}
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
invalid_action | 400 | Not in two-way translation mode | Use only under the conversation type |
conversation_invalid_mode | 400 | Invalid conversation mode | Use auto or manual |
set_speaker_language - Set Speaker Language
Description
During two-way translation mode, change a specified user's language in real time. The system rebuilds the STT connection to adapt to the new language, and the translation target is updated automatically. Transcript content before the change keeps its original language, and timestamps continue counting without resetting.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value set_speaker_language |
speaker | int | Yes | User number (1 or 2) |
language | string | Yes | The new language code (e.g. ja-JP) |
Request Example
{
"type": "voice-translation",
"data": {
"action": "set_speaker_language",
"speaker": 1,
"language": "ja-JP"
}
}
Successful Response
Returns the speaker_language_changed event:
{
"type": "voice-translation",
"data": {
"action": "speaker_language_changed",
"speaker_language_map": {
"1": "ja-JP",
"2": "en-US"
}
}
}
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
invalid_action | 400 | Not in two-way translation mode | Use only under the conversation type |
conversation_invalid_speaker | 400 | Invalid user number | Use 1 or 2 |
conversation_invalid_language | 400 | No language provided | Include language |
invalid_transcription_language | 400 | Invalid language code | Use a valid BCP 47 language code |
session_not_started | 400 | Recording has not started | Call start first |
conversation_same_language | 400 | Same as the current language | You can ignore this warning |
conversation_language_same_as_peer | 400 | The new language is the same as the other user's | The two users cannot have the same language |
conversation_speaking | 400 | Currently speaking, cannot change the language | End speaking first, then change |
conversation_language_change_failed | 500 | Language change failed (STT rebuild failed) | Retry later |
set_speaking_speed - Change Speaking Speed During Recording
Description
Dynamically adjust the speaking speed (which controls the silence threshold for segmentation) while recording. The system rebuilds the STT connection to apply the new setting, causing a brief interruption in recognition (same as changing a speaker's language during recording). Supported in all recognition modes except multi-speaker (multi_speaker) — including broadcast and multi-language LID; speaking_speed is not applied in multi-speaker mode.
Multi-channel mode (multi_channel, v1.10.0): supported. The new segmentation threshold applies to all channels, taking effect channel by channel; each channel sends its own channel_status event (reason: "speaking_speed" — first preparing, then ready once that channel starts producing text again), and produces no text for about 4 seconds while the change is applied. Like set_channel_language, there is a minimum 5-second interval between changes (calling too frequently returns channel_rebuild_too_frequent; a single adjustment applies to all channels, so this interval is shared across the whole session). While paused it returns channel_action_while_paused — call resume first, then adjust.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value set_speaking_speed |
speaking_speed | string | Yes | very_slow / slow / normal / fast / very_fast (see speaking_speed levels for each threshold) |
Request Example
{
"type": "voice-translation",
"data": { "action": "set_speaking_speed", "speaking_speed": "slow" }
}
Success Response
{
"type": "voice-translation",
"data": { "action": "speaking_speed_changed", "speaking_speed": "slow" }
}
Error Codes
| Error Code | HTTP | Description | Suggested Handling |
|---|---|---|---|
invalid_data | 422 | Invalid speaking_speed value | Use one of the five values |
session_not_started | 400 | Recording not started | Call start first |
invalid_action | 400 | Not supported in multi-speaker | Do not call in multi-speaker |
set_speaking_speed_failed | 400 | Failed to rebuild STT | Retry later |
channel_rebuild_too_frequent | 400 | Multi-channel: speed adjustments are too frequent (only one accepted every 5 seconds, v1.10.0) | Retry later |
channel_action_while_paused | 400 | Multi-channel: the speed cannot be adjusted while paused (v1.10.0) | Call resume first, then adjust |
If a sentence is in progress when this happens, it cannot be kept and is reported separately via segment_discarded (
reason: "speaking_speed"). You may receive it even whenset_speaking_speed_failedis returned — recognition is interrupted before the rebuild.
add_channel - Add a Channel (Multi-Channel)
Description
Dynamically add a channel while a multi-channel recording (recognition_mode: "multi_channel", v1.10.0) is in progress: when a participant joins on the spot, you can open a new channel for them without stopping the recording. Once the channel is added, you can start sending audio frames with that channel_id.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value add_channel |
channels | object[] | Yes | Exactly 1 channel configuration, with the same fields as channels[] in start (channel_id, optional speaker_name; transcription_languages with exactly 1 entry under per_channel, omitted under shared) |
Request Example
{
"type": "voice-translation",
"data": {
"action": "add_channel",
"channels": [
{ "channel_id": 4, "speaker_name": "Chen", "transcription_languages": ["zh-TW"] }
]
}
}
Successful Response
Returns a channel_status event (reason: "added", status: "preparing"):
{
"type": "voice-translation",
"data": {
"action": "channel_status",
"channel": {
"channel_id": 4,
"speaker_name": "Chen",
"transcription_languages": ["zh-TW"],
"status": "preparing"
},
"active_channels": 4,
"stt_stream_count": 4,
"reason": "added"
}
}
When that channel recognizes its first sentence, another channel_status arrives (status: "ready", with reason carried over from the triggering added).
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
not_multi_channel_session | 400 | This recording is not in multi-channel mode | Use only with recognition_mode: "multi_channel" |
channels_required | 400 | channels is missing or does not contain exactly 1 element | Add one channel at a time |
invalid_channel_id | 400 | channel_id is out of range (1–8) | Fix the channel number |
channel_id_in_use | 400 | The number is in use, or has already been removed (numbers cannot be reused) | Use a number that has never been used |
speaker_name_duplicate | 422 | speaker_name duplicates another channel's current name | Use a different name (a name freed up by a rename can be taken) |
too_many_channels | 400 | The channel count has reached the limit | Call remove_channel first, then add |
channel_language_required | 400 | Exactly one transcription language was not specified (per_channel) | Provide exactly 1 entry in transcription_languages |
channel_language_not_allowed | 400 | transcription_languages was provided in shared mode | Remove the field |
invalid_transcription_language | 400 | Invalid language code | Make sure the language code format is correct (e.g. zh-TW) |
invalid_parameter | 400 | speaker_name is too long (>100 characters) or contains control characters | Fix the name |
plan_feature_not_allowed | 403 | Exceeds the plan's channel count limit (details.field is max_stt_streams) | Remove another channel or upgrade the plan |
too_many_languages | 400 | The new channel's language pushes the number of simultaneous recognition languages over the plan limit | Use an existing language or upgrade the plan |
channel_action_while_paused | 400 | Channels cannot be added or disabled while paused | Call resume first |
session_not_started | 400 | Recording has not started | Call start first |
stt_start_failed | 500 | Speech recognition failed to start on the channel (this add did not take effect) | Retry later |
Notes
- Numbers cannot be reused: once a
channel_idhas been used (including removed ones), it is occupied permanently — speaker identity is fixed at the moment it is written into the transcript, and reusing a number would make two different people share the same speaker identity. - Startup delay: after a successful add, the channel begins producing text within about 4 seconds; audio sent during this period is not lost, only delayed.
- The new channel's timeline is automatically aligned with the session (someone joining at second 65 gets speech timestamps starting from about second 65).
- The billed channel count is taken as each minute begins, so an added channel is billed from the next minute — see the Pricing Guide.
sharedmode: the added channel shares the session's recognition and does not increase the billed channel count; its status takes on the first channel's current status directly.
remove_channel - Disable a Channel (Multi-Channel)
Description
Disable a channel while a multi-channel recording is in progress (v1.10.0): when a participant leaves early, disabling their channel stops billing for that channel from the next billed minute onward. The semantics are "disable", not "delete" — the transcript and single-track audio that channel has already produced are all preserved.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value remove_channel |
channel_id | int | Yes | The channel number to disable |
Request Example
{
"type": "voice-translation",
"data": {
"action": "remove_channel",
"channel_id": 4
}
}
Successful Response
Returns a channel_status event (reason: "removed", status: "removed"):
{
"type": "voice-translation",
"data": {
"action": "channel_status",
"channel": {
"channel_id": 4,
"speaker_name": "Chen",
"transcription_languages": ["zh-TW"],
"status": "removed"
},
"active_channels": 3,
"stt_stream_count": 3,
"reason": "removed"
}
}
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
not_multi_channel_session | 400 | This recording is not in multi-channel mode | Use only with recognition_mode: "multi_channel" |
channel_id_required | 400 | channel_id is missing | Provide the number of the channel to disable |
invalid_channel_id | 400 | channel_id is out of range (1–8) | Fix the channel number |
unknown_channel_id | 400 | Unknown channel_id (it may have been removed already) | Make sure the channel exists and has not been removed |
channel_remove_not_allowed | 400 | The last remaining channel cannot be disabled; or the first channel in shared mode | Use stop to end the recording |
channel_action_while_paused | 400 | Channels cannot be added or disabled while paused | Call resume first |
session_not_started | 400 | Recording has not started | Call start first |
Notes
- Wrap-up window: after disabling, there is a wrap-up window of about 3 seconds so that a final sentence already spoken on that channel but not yet returned can still make it into the transcript; the channel no longer accepts new audio during this period.
- Numbers cannot be reused: once disabled, that
channel_idcannot be used withadd_channelagain (it returnschannel_id_in_use). - Sending
audiowith that number after disabling returnsunknown_channel_id. - Billing: one fewer channel is billed starting from the next billed minute — see the Pricing Guide.
set_channel_language - Change a Channel's Language (Multi-Channel)
Description
Change the transcription language of a single channel while a multi-channel recording is in progress (v1.10.0). The channel_id and speaker identity (speaker_id) stay the same and the transcript remains continuous — which is exactly why this action exists: with remove_channel + add_channel, channel numbers cannot be reused, so the same person would become two different speakers in the transcript.
The system switches that channel to the new language (other channels are unaffected). The switch takes about 4 seconds to take effect, during which the channel produces no text.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value set_channel_language |
channel_id | int | Yes | Target channel number |
transcription_languages | string[] | One of the two | The new transcription language, exactly 1 (same shape as channels[] in start) |
language | string | One of the two | The new transcription language (single-value form). Providing both fields with different values returns invalid_parameter |
Request Example
{
"type": "voice-translation",
"data": {
"action": "set_channel_language",
"channel_id": 3,
"transcription_languages": ["en-US"]
}
}
Successful Response
Returns a channel_status event (reason: "language_change", status: "preparing"):
{
"type": "voice-translation",
"data": {
"action": "channel_status",
"channel": {
"channel_id": 3,
"speaker_name": "Sato",
"transcription_languages": ["en-US"],
"status": "preparing"
},
"active_channels": 3,
"stt_stream_count": 3,
"reason": "language_change"
}
}
When the channel recognizes its first sentence in the new language, another channel_status arrives (status: "ready").
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
not_multi_channel_session | 400 | This recording is not in multi-channel mode | Use only with recognition_mode: "multi_channel" |
channel_id_required | 400 | channel_id is missing | Provide the target channel number |
invalid_channel_id | 400 | channel_id is out of range (1–8) | Fix the channel number |
unknown_channel_id | 400 | Unknown channel_id (it may have been removed) | Make sure the channel exists and has not been removed |
channel_language_required | 400 | No new language provided | Provide transcription_languages or language |
invalid_parameter | 400 | More than one language provided / the two language fields disagree / same as the channel's current language | Fix the parameters according to details |
invalid_transcription_language | 400 | Invalid language code | Make sure the language code format is correct (e.g. zh-TW) |
too_many_languages | 400 | The new language pushes the number of simultaneous recognition languages over the platform limit of 10 or the plan limit | Use an existing language or upgrade the plan |
channel_rebuild_too_frequent | 400 | The same channel was changed again within 5 seconds | Retry later (details.cooldown_seconds is the number of seconds to wait) |
channel_action_while_paused | 400 | The channel language cannot be changed while paused | Call resume first |
session_not_started | 400 | Recording has not started | Call start first |
channel_language_not_allowed | 400 | shared mode does not support changing a channel's language (v1.21.0) | The session language is set at start and cannot be changed during the recording |
stt_start_failed | 500 | The change did not take effect (a channel_status event with status: "error" is also sent; the system will try to recover automatically) | Wait for automatic recovery or retry later |
Notes
- Minimum interval between changes (5 seconds): two language changes on the same channel must be at least 5 seconds apart (
channel_rebuild_too_frequent); the interval starts counting only after all validation passes — rejected requests do not consume it. - The language set only grows: when a language change introduces a language the session has not used yet, that language is merged into the session-level
transcription_languages; the old language is not removed (the channel's earlier transcript is still in the old language). Language changes are therefore governed by the platform limit on simultaneous recognition languages (10) and the plan's language limit. - Changing to the language the channel currently uses returns
invalid_parameter(the no-op change is not performed). switch_languageis always disabled in multi-channel mode (it returnsmultichannel_switch_language_not_allowed); always use this action to change languages.
If a sentence is in progress when this happens, it cannot be kept and is reported separately via segment_discarded (
reason: "language_change"). The event carrieschannel_idto identify the channel.
set_summary - Change Summary Settings During Recording
Description
Change the summary settings that will be applied when the recording stops. A live recording generates its summary exactly once, at stop time, so "takes effect immediately" here means "that automatic summary will use the latest settings" — if you send this action several times, the last one before stopping wins.
Typical use: you decide to switch to a different summary template or custom prompt after the recording has already started. Updating in place avoids having to regenerate the summary after stopping, which saves one summary generation charge.
Parameters fall into two independent groups:
- Summary source group:
summary_mode,summary_template,summary_prompt,summary_prompt_slug. Applied only whensummary_modeis present, and the four fields are replaced as a set — switching modes automatically clears the fields belonging to the other mode. - Option group:
auto_summary,summary_language,summary_plain_text. Each is independent and is updated only when present.
The two groups do not affect each other: updating the summary source does not turn auto_summary on. If automatic summaries were disabled when the recording started, you must also send auto_summary: true to resume generation.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value set_summary |
summary_mode | string | Conditional | builtin (shared template) or custom (custom prompt). Required when changing the summary source; may be omitted when updating only the option group |
summary_template | string | Conditional | Template identifier for builtin mode. Not allowed in custom mode |
summary_prompt | string | Conditional | Full prompt for custom mode, up to 3000 characters. Required in custom mode (a value with only whitespace counts as not provided) |
summary_prompt_slug | string | Conditional | Identifier for custom mode, up to 64 characters. Required in custom mode (a value with only whitespace counts as not provided); not allowed in builtin mode |
auto_summary | boolean | No | false skips summary generation at stop; true resumes it |
summary_language | string | No | Summary output language, up to 20 characters |
summary_plain_text | boolean | No | Whether to output plain text (Markdown removed) |
You must provide either
summary_modeor at least one field from the option group; otherwise the request returnsinvalid_data. Unlikestart, this action does not inferbuiltinwhensummary_modeis missing. Sending summary source fields without a mode returnssummary_mode_field_mismatch.
Summary Source Requirements
The source is validated only when the request affects the summary source or whether summaries are generated — that is, when it carries the summary source group or explicitly sends auto_summary: true:
- Changing only
summary_language/summary_plain_text→ not validated. Such a request changes neither whether a summary is produced nor the source, so recordings that never specified a template can send it too. - Sending
auto_summary: false→ not validated; you may send only the option group. - Carrying the summary source group, or explicitly sending
auto_summary: true→ evaluated in order:- This request includes
summary_templateorsummary_prompt→ the new value is used. - This request includes neither, or includes only whitespace → treated as not provided, and the previous setting is kept; if none was ever specified, the request returns
summary_mode_field_mismatchasking you to provide one.
- This request includes
Recordings without a
summary_template(conversationandbroadcastmay omit it) fall back to the system default template when generating the summary. To explicitly enable the automatic summary or change the source on such a recording, you still need to supply a template or a custom prompt.
Request Example (switch to a custom prompt)
{
"type": "voice-translation",
"data": {
"action": "set_summary",
"summary_mode": "custom",
"summary_prompt": "List the decisions and action items from this meeting, and note the owner of each item.",
"summary_prompt_slug": "meeting-actions-v2"
}
}
Request Example (switch to a shared template)
{
"type": "voice-translation",
"data": {
"action": "set_summary",
"summary_mode": "builtin",
"summary_template": "meeting"
}
}
Request Example (skip the automatic summary)
{
"type": "voice-translation",
"data": { "action": "set_summary", "auto_summary": false }
}
Success Response
Returns the settings currently in effect. For safety the response omits the full summary_prompt; use summary_prompt_slug to identify it:
{
"type": "voice-translation",
"data": {
"action": "summary_updated",
"summary_mode": "custom",
"summary_prompt_slug": "meeting-actions-v2",
"summary_language": "zh-TW",
"auto_summary": true,
"summary_plain_text": false,
"message": "Summary settings updated"
}
}
Error Codes
| Error Code | HTTP | Description | Suggested Handling |
|---|---|---|---|
invalid_data | 422 | No fields provided, or summary_language too long | Include at least one updatable field |
summary_invalid_mode | 400 | summary_mode is neither builtin nor custom | Use one of the two valid values |
summary_mode_field_mismatch | 400 | Field combination conflicts with the mode, or the source is missing | Add the field named by missing_field in details |
summary_prompt_too_long | 400 | summary_prompt exceeds 3000 characters | Shorten the prompt |
summary_prompt_slug_too_long | 400 | summary_prompt_slug exceeds 64 characters | Shorten the identifier |
summary_prompt_slug_invalid | 400 | summary_prompt_slug contains control characters | Remove the control characters |
invalid_summary_template | 400 | The template does not exist or is not enabled | Use a valid template |
session_not_started | 400 | Recording has not started, or has already stopped | Call start first; if stopped, use the regeneration API |
Notes
- After reconnecting, the
settingsobject inresume_okreturns the updated summary settings, so you can reconcile client state directly. - Calling this action after
stopreturnssession_not_started. A summary that is already being generated is unaffected. - After the recording stops you can still regenerate the summary with any template through the summary regeneration endpoint.
broadcast_go_live - Switch to the Live Phase
Description
Switch from the broadcast standby phase (standby) to the live phase (live). After the switch, STT/translation results begin broadcasting to viewers and start being written to the transcript.
Request Example
{
"type": "voice-translation",
"data": {
"action": "broadcast_go_live"
}
}
Successful Response
Returns the broadcast_phase_changed event:
{
"type": "voice-translation",
"data": {
"action": "broadcast_phase_changed",
"phase": "live",
"message": "Broadcast started"
}
}
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
broadcast_not_enabled | 400 | Not in broadcast mode | Make sure type: "broadcast" |
session_not_started | 400 | Speech recognition has not started, or the recording has ended (including while it is still being processed after ending) | Call start first if it has not started |
Note: If already in the live phase, a status message "Broadcast is already in progress" is returned; this is not treated as an error.
If a sentence is in progress when this happens, it cannot be kept and is reported separately via segment_discarded (
reason: "broadcast_go_live"). Standby-phase segments are never written to the final transcript anyway.
broadcast_announcement - Send an Announcement
Description
The host sends a custom announcement message to all viewers. Viewers receive an announcement event via SSE. The announcement message is automatically translated into all translation languages, and the SSE event viewers receive includes a translations field.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value broadcast_announcement |
message | string | Yes | The announcement message content |
Request Example
{
"type": "voice-translation",
"data": {
"action": "broadcast_announcement",
"message": "The meeting will end in 5 minutes"
}
}
Successful Response
{
"type": "voice-translation",
"data": {
"action": "status",
"message": "Announcement sent"
}
}
The SSE event received on the viewer side (with translations):
event: announcement
data: {"message":"The meeting will end in 5 minutes","translations":{"en-US":"The meeting will end in 5 minutes","ja-JP":"会議は5分後に終了します"}}
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
broadcast_not_enabled | 400 | Not in broadcast mode | Make sure type: "broadcast" |
invalid_parameter | 400 | The message is empty | Provide a valid message parameter |
set_standby_message - Set the Standby Phase Message
Description
Dynamically set the message displayed to viewers during the broadcast standby phase (standby). This lets the host enter standby mode first and set the waiting message afterward, rather than having to provide it at start.
The message is automatically translated into all translation languages, and the SSE event viewers receive includes a translations field.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value set_standby_message |
message | string | Yes | The standby phase display text (translated for viewers in each language via the translation pipeline) |
Request Example
{
"type": "voice-translation",
"data": {
"action": "set_standby_message",
"message": "The talk is about to begin, please wait..."
}
}
Successful Response
{
"type": "voice-translation",
"data": {
"action": "status",
"message": "Standby phase text updated"
}
}
The SSE event received on the viewer side (with translations):
event: standby
data: {"message":"The talk is about to begin, please wait...","translations":{"en-US":"The presentation is about to begin, please wait...","ja-JP":"プレゼンテーションがまもなく始まります。お待ちください..."}}
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
broadcast_not_enabled | 400 | Not in broadcast mode | Make sure type: "broadcast" |
broadcast_not_in_standby | 400 | Not in the standby phase | Can only be used during the standby phase |
Note: This action can only be used during the standby phase (standby). If you have already entered the live phase (live), an error is returned.
Version: V1.24.1 Last Updated: 2026-10-07