WebSocket API

Voice Translation Actions

Overview

A complete list of all actions available under the voice-translation type. For connection and authentication, see Connection and Authentication; for response event formats, see Response Events.


Table of Contents

  1. start - Start Voice Translation
  2. config - Configure Terminology / Correction Rules
  3. audio - Send Audio
  4. pause - Pause Translation
  5. resume - Resume Translation
  6. stop - Stop Translation
  7. retranslate - Retranslate a Single Sentence
  8. switch_language - Switch Language
  9. set_name - Set Recording Name
  10. rename_speaker - Globally Rename a Speaker
  11. reassign_speaker - Change the Speaker of a Single Sentence
  12. merge_speakers - Merge Speakers
  13. tts_play - Play TTS
  14. tts_stop - Stop TTS
  15. tts_mode - Switch TTS Mode
  16. set_tts - Two-Way Translation TTS Settings
  17. start_speaking - Start Speaking (Manual Mode)
  18. stop_speaking - Stop Speaking (Manual Mode)
  19. switch_conversation_mode - Switch Conversation Mode
  20. set_speaker_language - Set Speaker Language
  21. set_speaking_speed - Change Speaking Speed During Recording
  22. add_channel - Add a Channel (Multi-Channel)
  23. remove_channel - Disable a Channel (Multi-Channel)
  24. set_channel_language - Change a Channel's Language (Multi-Channel)
  25. set_summary - Change Summary Settings During Recording
  26. broadcast_go_live - Switch to the Live Phase
  27. broadcast_announcement - Send an Announcement
  28. set_standby_message - Set the Standby Phase Message

start - Start Voice Translation

Note: Do not send glossaries inside start. If the start payload carries terminology, fuzzy_correction, or translation_dict, those fields are ignored (start itself still succeeds) and a config_ignored_in_start warning is returned, with details.ignored_fields listing what was dropped. Always send glossaries through the config action instead.

Description

Start a new voice translation session and begin processing audio according to the configured parameters.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value start
transcription_languagesstring[]YesSpeech recognition languages (up to 10)
translation_languagesstring[]NoTranslation target languages; multiple allowed (up to 12; empty = no translation). As of v1.6.7, transcribe and broadcast translate all specified languages in real time, with one result event per language (see the "Multi-Language Translation" example below). Two-way translation (conversation) does not apply: the server overwrites this field with the counterpart language, always a single language. record no longer supports translation as of v1.7.0 and returns 400 record_translation_not_allowed
realtime_translationbooleanNoReal-time translation mode (default false). true: translates word-by-word while the sentence is still being recognized (interim); false: translates only when the sentence is finalized. This flag also governs the real-time behavior of multi-language translation. Always treated as true for broadcast; has no effect for conversation, which translates finalized sentences only
recognition_modestringNoRecognition mode: single (single speaker, default), multi_speaker (multiple speakers), multi_channel (multi-channel, v1.10.0; must be enabled for your environment before use — see "Multi-Channel Mode Description" below); under multi_speaker, transcription_languages must contain exactly 1 language, otherwise a diarization_multilang_conflict error is returned and the session is refused (type=conversation is exempt: two-way translation forces single-speaker mode and has been exempt from this check since v1.7.2)
typestringYesRecording type: transcribe, conversation, record, broadcast
audio_formatstringNoAudio format: pcm (default), webm
summary_templatestringConditionalSummary template (required for transcribe, optional for conversation/broadcast)
optionsobjectNoSpeech recognition options
tts_enabledbooleanNoWhether to enable TTS speech synthesis (default false)
tts_languagestringNoTTS output language (must be in translation_languages)
tts_voicestringNoTTS voice name (e.g. en-US-JennyNeural)
tts_modestringNoTTS playback mode: sync (synchronous, default), async (asynchronous). Only these two lowercase values are accepted, and an empty string is treated as not provided (sync); any other value (for example "Async") is rejected (invalid_parameter, details.field is tts_mode) and the recording does not start. Checked for every recording type, whether or not TTS is enabled
broadcast_tokenstringConditionalBroadcast token (required for broadcast type, obtained from the REST API). Allowed only with the broadcast type: any other type that carries it is rejected (invalid_parameter, details.field is broadcast_token) and the recording does not start
active_languagestringNoInitial active language in two-way translation mode (default transcription_languages[0])
tts_configobjectNoMulti-language TTS settings (broadcast / two-way translation mode)
broadcast_phasestringNoInitial broadcast phase: standby, live (default). Only these two lowercase values are accepted, and an empty string is treated as not provided (live); any other value (for example "Live") is rejected (invalid_parameter, details.field is broadcast_phase) and the recording does not start
standby_messagestringNoMessage viewers see during the standby phase (default: "Preparing, please wait...")
namestringNoInitial default recording name (max 60 characters after trimming leading and trailing whitespace; the system may still override it; if not provided, one is generated automatically, e.g. Transcription #1). A longer name is rejected (invalid_parameter, details.field is name) and the recording does not start
summary_languagestringNoSummary output language (defaults to the recognition language when not specified; in broadcast mode it is read automatically from the channel settings). Up to 20 characters; a longer value is rejected (invalid_parameter, details.field is summary_language) and the recording does not start
summary_modestringNoSummary mode enum: builtin (apply the built-in template, default) / custom (the customer prompt fully replaces the default). When omitted, builtin is inferred automatically
summary_promptstringNoRequired in custom mode (a value with only whitespace counts as not provided); treated as supplementary instructions in builtin mode. ≤3000 characters
summary_prompt_slugstringNoRequired in custom mode (a value with only whitespace counts as not provided); must not be provided in builtin mode. The customer's own identifier (≤64 characters, Unicode, no control characters; passed through and stored in the backend record for historical lookup)
summary_plain_textbooleanNoRequest plain-text summary output (default false; when enabled, the backend performs Markdown post-processing)
speakersobject[]NoSpeaker language settings for two-way translation mode (exactly 2 entries when provided, see below). When omitted, speaker 1 uses transcription_languages[0] and speaker 2 uses transcription_languages[1]
conversation_modestringNoTwo-way conversation mode: auto (automatic detection, default), manual (manual PTT). Only these two lowercase values are accepted, and an empty string is treated as not provided (auto); any other value is rejected (invalid_parameter, details.field is conversation_mode) and the recording does not start. Also checked for recording types other than two-way
channel_modestringConditionalMulti-channel sub-mode (required for multi_channel, v1.10.0): per_channel (each channel is recognized independently) or shared (channels take turns speaking and share one recognition stream, v1.21.0); other values return invalid_channel_mode
channelsobject[]ConditionalMulti-channel channel list (required for multi_channel, 1–8 channels, including the main speaker, conventionally channel_id: 1; see "Multi-Channel Mode Description" below for the fields)
silenceTimeoutSecondsintegerNoHow many consecutive seconds without detected speech end the recording automatically: omitted or null uses the default (900 seconds); 0 means this session never ends for lack of speech; otherwise it must be an integer from 60 to 86400, and any other value is rejected. See Automatic End After a Long Silence below

options Sub-fields

options is the speech recognition options object; all fields are optional and use their respective defaults when omitted.

FieldTypeDefaultDescription
speaking_speedstringnormalSpeaking speed, which affects the silence threshold for sentence segmentation: very_slow / slow / normal / fast / very_fast (see speaking_speed levels below for each threshold; the default normal is 800ms). Use a slower setting for slower speakers (longer threshold, avoids cutting on mid-sentence pauses); a faster setting segments sooner. Can be adjusted dynamically during recording via set_speaking_speed. Only these five lowercase values are accepted, and an empty string is treated as not provided (normal); any other value is rejected (invalid_parameter, details.field is options.speaking_speed) and the recording does not start
profanity_handlingstringmaskProfanity handling: mask (mask with ***) / remove (remove) / show (show original). Only these three lowercase values are accepted, and an empty string is treated as not provided (mask); any other value is rejected (invalid_parameter, details.field is options.profanity_handling) and the recording does not start

Note: These options apply to STT sentence segmentation; multi-speaker mode (multi_speaker) does not currently apply speaking_speed.

speaking_speed levels

A sentence ends once silence lasts longer than the threshold. A longer threshold tolerates pauses within a sentence (a sentence is less likely to be split in two), but each sentence appears later.

LevelSilence threshold for ending a sentence
very_fast300ms
fast600ms
normal (default)800ms
slow1200ms
very_slow1500ms

Request Example (Basic)

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "transcription_languages": ["zh-TW"],
    "translation_languages": ["en-US"],
    "realtime_translation": false,
    "type": "transcribe",
    "audio_format": "pcm",
    "summary_template": "meeting",
    "options": {
      "speaking_speed": "normal",
      "profanity_handling": "mask"
    }
  }
}

Request Example (Multi-Language Translation, v1.6.7)

Specify multiple languages in translation_languages (up to 12) to translate into several languages at once. This applies to transcribe and broadcast (for two-way translation the language is set by the server and is always a single language; record does not support translation as of v1.7.0):

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "transcription_languages": ["zh-TW"],
    "translation_languages": ["en-US", "ja-JP", "ko-KR"],
    "realtime_translation": true,
    "type": "transcribe",
    "summary_template": "meeting"
  }
}

How results arrive: each language returns its own independent result event (same sid, with a single language key inside translations). Multiple languages are never merged into one event. Clients must accumulate translations by "sid + language code" instead of overwriting. Completion order across languages is not fixed (translations run in parallel), and if one language fails the remaining languages are still delivered (the failed language additionally receives an error event whose details.translation_language identifies it).

Billing reminder: translation is billed per language starting from the second language (see the Pricing Guide). Specifying N languages is billed as N languages.

Request Example (Initial Default Name)

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "transcription_languages": ["zh-TW"],
    "translation_languages": ["en-US"],
    "type": "transcribe",
    "audio_format": "pcm",
    "summary_template": "meeting",
    "name": "Product Planning Meeting"
  }
}

Recording Name Rules

ScenarioNamename_sourceOverridden by system?
start with a name parameterInitial default namedefaultYes
start without a nameAuto-generated (e.g. Transcription #1, Broadcast #3)defaultYes
Set via set_nameName explicitly set by the useruserNo
Auto-generated by the system after the session endsSummary name generated from the transcript contentllm—

Note: The name in start is an initial default name; the system may still override it when the session ends. If you need a fixed name, use set_name.

Default name formats (fixed English):

Recording TypeDefault Name Format
transcribeTranscription #N
conversationConversation #N
recordRecording #N
broadcastBroadcast #N

N is the sequential number of recordings of the same type for that user. Name priority: user > llm > default. Once the user sets a name, the system will not override it when the session ends.

Request Example (with TTS)

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "transcription_languages": ["zh-TW"],
    "translation_languages": ["en-US"],
    "realtime_translation": true,
    "type": "transcribe",
    "tts_enabled": true,
    "tts_language": "en-US",
    "tts_voice": "en-US-JennyNeural",
    "tts_mode": "sync"
  }
}

Request Example (Two-Way Translation Mode - Automatic Detection)

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "type": "conversation",
    "transcription_languages": ["zh-TW", "en-US"],
    "active_language": "zh-TW",
    "audio_format": "pcm",
    "speakers": [
      { "id": 1, "language": "zh-TW" },
      { "id": 2, "language": "en-US" }
    ],
    "tts_config": {
      "zh-TW": { "voice": "zh-TW-HsiaoChenNeural", "speaking_rate": 1.0 },
      "en-US": { "voice": "en-US-JennyNeural", "speaking_rate": 1.0 }
    }
  }
}

Request Example (Two-Way Translation Mode - Manual Mode)

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "type": "conversation",
    "transcription_languages": ["zh-TW", "en-US"],
    "conversation_mode": "manual",
    "audio_format": "pcm",
    "speakers": [
      { "id": 1, "language": "zh-TW" },
      { "id": 2, "language": "en-US" }
    ],
    "tts_config": {
      "zh-TW": { "voice": "zh-TW-HsiaoChenNeural", "speaking_rate": 1.0 },
      "en-US": { "voice": "en-US-JennyNeural", "speaking_rate": 1.0 }
    }
  }
}

Special rules for two-way translation mode:

ItemDescription
transcription_languagesMust contain exactly 2 languages, and they must differ
translation_languagesOverwritten by the server: ignored even if supplied; always forced to the counterpart of the active language
realtime_translationHas no effect in conversation mode: translations are sent only after a sentence is finalized, regardless of whether this field is true or false
active_languageOptional, defaults to transcription_languages[0]
recognition_modeForced to single; speaker_diarization is accepted but ignored (as of v1.7.2. Before that, sending speaker_diarization=true was wrongly rejected with a 400 "speaker diarization conflicts with multiple languages")
tts_enabledDefaults to true; set to false to return text translation only
tts_configOptional; configures the TTS voice for each of the two languages; leave empty to use the default voices automatically
summary_templateOptional; when provided, a summary is generated automatically after stopping
speakersOptional; specifies each user's language (exactly 2 entries when provided). When omitted, users 1 and 2 map to the two transcription_languages in order
conversation_modeOptional, auto (automatic detection, default) or manual (manual PTT)

speakers field description:

FieldTypeRequiredDescription
idintYesUser number (1 or 2)
languagestringYesThat user's language code (must be in transcription_languages)

conversation_mode description:

ModeDescription
auto (default)The system automatically detects the spoken language and segments sentences automatically
manualThe user controls the speaking interval via start_speaking / stop_speaking; audio during that interval is merged into a single sentence

Successful Response

After a successful start, a session_started event is returned, containing the complete initial session information. For real-time recording, the first minute is deducted first, and the event is returned once that deduction completes (see the Pricing Guide).

General recording (transcribe / conversation / record):

{
  "type": "voice-translation",
  "data": {
    "action": "session_started",
    "session_id": "550e8400-e29b-41d4-a716-446655440000",
    "task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
    "recording_type": "transcribe",
    "recognition_mode": "single",
    "message": "Speech recognition started",
    "resume_token": "L0VBAwIy... (43 chars)",
    "resume_grace_seconds": 45,
    "server_time": 1749550000000
  }
}

Broadcast mode (broadcast):

{
  "type": "voice-translation",
  "data": {
    "action": "session_started",
    "session_id": "550e8400-e29b-41d4-a716-446655440000",
    "task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
    "recording_type": "broadcast",
    "recognition_mode": "multi_speaker",
    "phase": "standby",
    "viewer_count": 0,
    "queue_count": 0,
    "peak_viewers": 0,
    "total_viewers": 0,
    "message": "Speech recognition started",
    "resume_token": "L0VBAwIy... (43 chars)",
    "resume_grace_seconds": 45,
    "server_time": 1749550000000
  }
}

For response field descriptions, see the session_started event.

Recording Type Descriptions

typeDescriptionUse Case
transcribeSpeech-to-textMeeting minutes, interview records
conversationConversation logTwo-way communication, customer service dialogues
recordPlain recordingVoice memos, quick notes
broadcastBroadcast / live streamLectures, speeches, live content

Broadcast Mode Description (type: "broadcast")

In broadcast mode, the language settings are obtained automatically from the broadcast channel settings and do not need to be sent in the WebSocket message.

Required parameters:

ParameterTypeDescription
typestringMust be "broadcast"
broadcast_tokenstringBroadcast token (obtained after creating a broadcast via the REST API)
audio_formatstringAudio format (pcm or webm)

Optional parameters (override broadcast channel settings):

ParameterTypeDescription
tts_configobjectMulti-language TTS settings (override the settings used at creation)
summary_templatestringSummary template slug (overrides the settings used at creation; if not provided, the broadcast channel default is used)

Automatically configured parameters (can be omitted):

  • transcription_languages: read automatically from the broadcast settings
  • translation_languages: read automatically from the broadcast settings
  • realtime_translation: always enabled in broadcast mode, and false is treated as true; translation is always billed at the real-time translation rate
  • summary_template: read automatically from the broadcast settings (the value passed via WebSocket takes precedence)
  • summary_language: read automatically from the broadcast settings (the value passed via WebSocket takes precedence)

Both the host and viewers receive interim translations: a sentence's translation first arrives with is_final: false, followed by the finalized version with is_final: true. Overwrite the display by sid plus language; to show only finalized translations, skip is_final: false.

The recording name does not reuse the channel name. If start omits name, a name such as Broadcast #1 is generated; naming otherwise works the same as for other types, see Recording Name Rules.

Broadcast phase description:

broadcast_phaseDescriptionBehavior
live (default)Live phaseSTT/translation results are broadcast to viewers and written to the transcript
standbyStandby phaseSTT/translation results go only to the host; viewers see the standby_message

Purpose of the standby phase: Lets the host run STT/translation warm-up tests before going live, confirming the equipment works before switching to the live phase.

The standby phase has a time limit (30 minutes by default): when the accumulated standby time reaches the limit, the session ends automatically, with a warning about 2 minutes before. See Standby Time Limit.

broadcast_phase accepts only lowercase standby and live: an empty string is treated as live, and any other value is rejected (invalid_parameter).

Broadcast mode request example:

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "type": "broadcast",
    "broadcast_token": "a3f9",
    "audio_format": "pcm"
  }
}

Broadcast mode request example (standby phase + override summary template):

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "type": "broadcast",
    "broadcast_token": "a3f9",
    "audio_format": "pcm",
    "broadcast_phase": "standby",
    "standby_message": "The talk is about to begin, please wait...",
    "summary_template": "lecture"
  }
}

Summary template priority: the value passed in the WebSocket start > the default set when creating the broadcast channel. If neither is set, no summary is generated automatically.

Broadcast mode TTS settings (tts_config):

Use the tts_config parameter to specify which translation languages should produce TTS audio for viewers.

tts_config fieldTypeDescription
voicestringTTS voice name
speaking_ratenumberSpeaking rate (0.5–2.0, default 1.0). Values outside the range are adjusted to the nearest bound
{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "type": "broadcast",
    "broadcast_token": "a3f9",
    "audio_format": "pcm",
    "tts_config": {
      "en-US": {
        "voice": "en-US-JennyNeural",
        "speaking_rate": 1.0
      },
      "ja-JP": {
        "voice": "ja-JP-NanamiNeural",
        "speaking_rate": 1.0
      }
    }
  }
}

Note:

  • The TTS language must be a valid language in translation_languages; invalid languages are ignored automatically
  • The host (WebSocket) does not receive TTS audio; only SSE viewers receive the tts_ready event
  • TTS is sent only during the live phase; it is not sent during the standby phase

Multi-Channel Mode Description (recognition_mode: "multi_channel")

A single recording session takes input from multiple physical microphones at the same time (one per person), and speaker identity is determined by the channel — whichever microphone captured the audio identifies the speaker; the system performs no speaker inference. Suitable for meetings, round-table discussions, and other settings where every participant has a dedicated microphone (v1.10.0).

There are two sub-modes (channel_mode):

Sub-modeRecognitionLanguageBest for
per_channelEach channel is recognized independently; people can speak at the same timeEach channel is bound to exactly 1 language; channels can differSeveral people may speak at once, each in a different language
shared (v1.21.0)Channels take turns speaking and share one recognition streamThe whole session shares the session-level settingOnly one person speaks at a time (a host and guests taking turns, etc.)

Scope and hard limits:

ItemRule
Feature activationMust be enabled before use; sending it in an environment where it is not enabled returns invalid_recognition_mode
Recording typesOnly transcribe and record; conversation returns invalid_parameter and broadcast returns multichannel_broadcast_not_allowed
Audio formatOnly pcm (16kHz / 16-bit / mono / little-endian); other values return multichannel_requires_pcm
TTSNot supported; tts_enabled: true returns multichannel_tts_not_allowed
speaker_diarizationCannot be specified at the same time (multi-channel is itself a form of speaker diarization); providing both returns invalid_parameter
switch_languageDisabled in multi-channel mode (languages are bound to channels); always returns multichannel_switch_language_not_allowed — under per_channel, use set_channel_language instead
Audio file length capThe audio saved for one multi-channel recording has a total cap, reached sooner the more channels are open (about 70 minutes with 8 channels). Once the cap is reached, the audio file and the recording duration (duration_ms) stop at that point, while the transcript and credit charges continue as usual

channels field description:

FieldTypeRequiredDescription
channel_idintYesChannel number, range 1–8, must be unique. Becomes that channel's speaker speaker_id (format channel_{N}); by convention the main speaker uses 1
speaker_namestringNoDisplay label for that channel's speaker (max 100 characters, no control characters). If not provided, it can be set later during recording with rename_speaker
transcription_languagesstring[]ConditionalTranscription language for that channel. Required under per_channel, exactly 1 (each channel is bound to one language); must not be provided under shared (returns channel_language_not_allowed)

Language consistency rule (per_channel): the union of all channel languages (deduplicated) must exactly match the session-level transcription_languages; otherwise channel_language_mismatch is returned (details lists both language sets so you can reconcile them).

Multi-channel request example:

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "type": "transcribe",
    "recognition_mode": "multi_channel",
    "channel_mode": "per_channel",
    "transcription_languages": ["zh-TW", "ja-JP"],
    "translation_languages": ["en-US"],
    "audio_format": "pcm",
    "summary_template": "meeting",
    "channels": [
      { "channel_id": 1, "speaker_name": "Main Speaker", "transcription_languages": ["zh-TW"] },
      { "channel_id": 2, "speaker_name": "Manager Wang", "transcription_languages": ["zh-TW"] },
      { "channel_id": 3, "speaker_name": "Sato", "transcription_languages": ["ja-JP"] }
    ]
  }
}

Successful response: the data of session_started includes channel_mode and channels[] (each entry with channel_id, speaker_name, transcription_languages, status; transcription_languages is not present in shared mode). Each channel starts with status preparing and transitions to ready via a channel_status event once that channel recognizes its first sentence.

Note: Always check that session_started includes channel_mode: this is the signal that the server actually started in multi-channel mode. If the response has no channel_mode, the server does not support multi-channel — channels[] was ignored and the session is actually running single-channel.

Other multi-channel behavior:

  • Every audio frame must include channel_id (see audio)
  • During recording you can add and disable channels dynamically via add_channel / remove_channel, and change a single channel's language via set_channel_language; set_speaking_speed also supports multi-channel (see that section)
  • rename_speaker works normally (a channel can be renamed before it is first used). The new name must not duplicate another channel's name (including the names set for each channel in start); duplicates return speaker_name_duplicate. reassign_speaker and merge_speakers do not apply to multi-channel recordings (speaker identity is determined by the channel, so there is nothing to reassign) and return speaker_op_not_allowed_multi_channel
  • pause / resume: while paused, the recording file is still saved but no transcript is produced; after resuming, speech from the paused period is transcribed retroactively (timestamps reflect the actual speaking time). Retroactive transcription has a limit — a trailing 60 seconds shared across the whole session (under per_channel, split evenly across channels when there are several; under shared, the whole last 60 seconds, for all channels together, with speakers labeled); anything beyond that is kept only in the audio file. A sentence cut off at the moment of pausing may not appear in the transcript (same as single-channel recording); when it is indeed not kept, resuming also sends segment_discarded (reason: "resumed") to identify it
  • The origin of result events includes channel_id and speaker_id (format channel_{N}); every transcript sentence also carries channel_id (see Response Events)
  • Unlimited plans must include the multi-channel feature; a plan may also cap the number of channels — exceeding the cap returns plan_feature_not_allowed on the spot at start / add_channel (details.field is max_stt_streams)
  • Billing: speech recognition plus speaker diarization as the base; per_channel adds a surcharge based on the number of active channels at the time, while shared is always counted as 1 channel — see the Pricing Guide

Shared mode (v1.21.0):

Request example:

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "type": "transcribe",
    "recognition_mode": "multi_channel",
    "channel_mode": "shared",
    "transcription_languages": ["zh-TW", "en-US"],
    "audio_format": "pcm",
    "channels": [
      { "channel_id": 1, "speaker_name": "Host" },
      { "channel_id": 2, "speaker_name": "Guest A" },
      { "channel_id": 3, "speaker_name": "Guest B" }
    ]
  }
}

Client requirements:

  • Send only one channel at a time: send audio only on the current speaker's channel. If several channels are sent at once, their audio is queued into the same recognition stream in the order received and the transcript becomes garbled (billing is not affected).
  • Keep sending silence: when nobody is speaking, keep sending silence on the current channel (one frame every 100ms recommended); do not stop sending.
  • pcm only (same as the common multi-channel rule).
  • Languages are shared across the session: channels must not specify transcription_languages (returns channel_language_not_allowed); the language cannot be changed during the recording (set_channel_language returns channel_language_not_allowed).
  • The first channel cannot be removed: the first entry in channels[] carries the recognition for the whole session, and remove_channel returns channel_remove_not_allowed; other channels can be removed.

Channel status:

  • The status of the other channels follows the first channel: when the first channel turns ready, all channels turn ready together; when the first channel prepares recognition again (resume after pause, resume after disconnect, automatic reconnection), all channels return to preparing together, and each channel receives its own channel_status event with the same reason as the first channel. After resuming from a disconnect, preparing is reported in the resume_ok snapshot rather than by a separate event.
  • A channel added with add_channel takes on the first channel's current status directly; removed follows each channel's own status.
  • When the first channel turns error, the other channels turn error as well (with the same reason); when the first channel recovers, they recover together.
  • The channels[] snapshots in session_started and resume_ok follow the same rules; channel entries do not carry transcription_languages.

Known limitations:

  • When the gap between speakers is shorter than about 0.8 seconds, the words of the two people may be merged into one sentence labeled with only one speaker.
  • Speakers are determined by channel labeling, which is highly accurate when people take turns; reassign_speaker and merge_speakers do not apply, so this cannot be corrected afterward.

TTS Playback Mode Description

ModeDescriptionBehavior
syncSynchronous mode (default)Automatically plays the most recent is_final=true translated sentence; if the previous sentence is still playing, it enters the queue and waits
asyncAsynchronous mode (manual control)The user can select any translated sentence for TTS, controlled with the tts_play command

Automatic End After a Long Silence

A recording that goes a certain length of time without recognizing any text (including interim results that are not final yet) ends automatically, so a recording someone forgot to stop does not keep being billed.

  • The default threshold is 900 seconds (15 minutes) and can be changed per session with silenceTimeoutSeconds. About 2 minutes before the end, a warning is sent once.
  • Any recognized text restarts the count from 0; resume and start_speaking also restart it.
  • The count does not run in these cases:
    • While paused: the count restarts from 0 after resuming. A paused recording is still billed; to stop billing, end the recording.
    • While a disconnected session waits to be resumed: after resuming, the count continues from where it was.
    • Broadcasts (type: "broadcast"): never end for lack of speech. The standby phase has its own time limit; see the Broadcast Guide.
  • Multi-channel: counts only when no channel recognizes any text; a single silent channel does not end the recording.
  • Conversation manual mode: speech recognized while the speak button is not pressed also restarts the count.

Values of silenceTimeoutSeconds (top level of the start data, optional)

ValueEffect
Omitted, or nullUses the default threshold
0This session never ends automatically for lack of speech; suited to sessions kept open for long periods during which nobody may speak
Integer from 60 to 86400The threshold in seconds for this session
Any other valuestart is rejected with the error code invalid_parameter (details.field is silenceTimeoutSeconds) and the recording does not start
  • The value must be a JSON integer: strings (for example "900"), decimal notations (for example 900.0 or 1e3), and booleans are all rejected.
  • The parameter name is silenceTimeoutSeconds; sending silence_timeout_seconds is rejected (details.field is silence_timeout_seconds), not ignored.
  • A valid value has no effect on broadcasts, but an invalid value is still rejected.
  • With a threshold of 120 seconds or less, no warning is sent; the recording simply ends when the time is up.
  • Resuming a disconnected session keeps the original setting; send it again when starting a new recording.

Events you receive

  1. stt_silence_warning (an error event with severity: "warning"): a warning; the recording continues. details.silenceSeconds is how long the silence has lasted and details.remainingSeconds is how many seconds remain. Do not treat it as the end of the recording; recognized text or resuming the recording restarts the count.
  2. stt_silence_timeout (severity: "fatal"): the recording has ended automatically. details.silence_seconds is the threshold in seconds.
  3. Then status: "ended" and task_complete, in that order, just as when the client sends stop: the recording is saved and summarized as usual, and billing runs until the end.
{
  "type": "error",
  "data": {
    "error_code": "stt_silence_warning",
    "severity": "warning",
    "message": "No speech detected for a while; the recording will end automatically soon",
    "context": "stt",
    "request_id": "req_abc123xyz789",
    "timestamp": "2026-09-25T10:28:00.000Z",
    "details": {
      "silenceSeconds": 780,
      "remainingSeconds": 120
    }
  }
}

After stt_silence_timeout, stop sending audio. Each audio message that arrives afterwards gets a session_not_started reply, which you can ignore.

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
missing_transcription_languages400No language parameter providedMake sure the request includes transcription_languages
invalid_transcription_language400Invalid language codeMake sure the language code format is correct (e.g. zh-TW)
too_many_languages400Number of languages exceeds the limitUp to 10 transcription languages and 12 translation languages
invalid_recording_type400Invalid recording typeUse a valid type value
audio_format_unsupported400audio_format is not supported; details.supported_formats lists the accepted valuesSwitch to a supported audio format
invalid_summary_template400Invalid summary templateMake sure the template identifier is correct
stt_init_failed503Service initialization failedRetry later
auth_insufficient_credit402Insufficient creditTop up your credit balance
auth_quota_exceeded402Available credits are insufficient; the recording did not start (less than one minute for real-time recording; the connection is not closed; details.remaining_budget is the available credit as of the most recent settlement, and details.budget_scope states whose credit it is)Top up and start again
daily_limit_reached—Usage has reached the plan's limit; the recording did not startAvailable again after the plan's reset (daily limits reset the next day)
auth_service_error500Service temporarily unavailable; the recording did not start (the connection is not closed)start again later
service_shutdown—The service is shutting down; the recording did not start, and the connection is then closedReconnect later and start again
tts_init_failed503TTS service initialization failedRetry later
tts_invalid_language400TTS language is not in the translation languagesMake sure tts_language is in translation_languages
broadcast_token_required400Broadcast mode requires a tokenA broadcast type must provide broadcast_token
broadcast_token_invalid401Invalid broadcast tokenMake sure the token is correct and has not expired
broadcast_not_ready503Broadcast service not yet startedRetry later
summary_invalid_mode400summary_mode is not builtin / customChange to a valid mode
summary_mode_field_mismatch400The mode and field combination do not match (a required field is missing / a forbidden field was provided)Adjust the fields according to the mode rules
summary_prompt_too_long400summary_prompt exceeds 3000 charactersShorten the custom prompt
summary_prompt_slug_too_long400summary_prompt_slug exceeds 64 charactersShorten the identifier
summary_prompt_slug_invalid400summary_prompt_slug contains control characters (\n / \r / \t / \0, etc.)Remove the control characters
invalid_recognition_mode400Invalid recognition mode, or multi-channel is not enabled for this environment (v1.10.0)Check the recognition_mode value; multi-channel must be enabled first
channel_mode_required400channel_mode is missing for multi-channel (v1.10.0)Provide channel_mode (per_channel or shared)
invalid_channel_mode400channel_mode is not per_channel or shared; also returned when shared is not enabled in this environment (v1.10.0)Fix channel_mode; if shared is not enabled, use per_channel
channel_language_not_allowed400A channel carries transcription_languages in shared mode (v1.21.0)Remove transcription_languages from each channel
channels_required400channels is missing for multi-channel (v1.10.0)Provide 1–8 channel configurations
too_many_channels400The channel count exceeds the limit of 8 (v1.10.0)Reduce the number of channels
invalid_channel_id400channel_id is out of range (1–8) or duplicated (v1.10.0)Fix the channel_id
channel_language_required400A channel does not specify exactly one language (v1.10.0)Provide exactly 1 entry in each channel's transcription_languages
channel_language_mismatch400The union of channel languages does not match transcription_languages (v1.10.0)Reconcile the two language lists
speaker_name_duplicate422Two or more channels use the same speaker_name (v1.10.0)Each channel needs a distinct name (channels without a name are unaffected)
multichannel_requires_pcm400Multi-channel supports only the pcm audio format (v1.10.0)Set audio_format to pcm
multichannel_tts_not_allowed400Multi-channel does not support speech synthesis (v1.10.0)Disable tts_enabled
multichannel_broadcast_not_allowed400Broadcast does not support multi-channel (v1.10.0)Use multi_speaker or single for broadcast
invalid_parameter400Multi-channel combined with conversation or speaker_diarization, or speaker_name is too long / contains control characters (v1.10.0); an invalid value or name for silenceTimeoutSeconds, an invalid broadcast_phase value, or broadcast_token sent with a type other than broadcast (v1.18.0); name longer than 60 characters, summary_language longer than 20 characters, or a value outside the accepted list for options.speaking_speed, options.profanity_handling, conversation_mode, or tts_mode (v1.24.0; details.valid_values lists the accepted values)Fix the parameter indicated by details.field
plan_feature_not_allowed403The plan does not include multi-channel, or the channel count exceeds the plan's limit (details.field is max_stt_streams, v1.10.0)Reduce channels or upgrade the plan

config - Configure Terminology / Correction Rules

Description

Send terminology, fuzzy-word correction rules, and translation dictionary settings before or during a recording. These settings can improve STT accuracy, fix homophone errors, and ensure translation consistency.

Terminology also drives homophone correction: When terminology is provided, the terms become the reference for homophone matching — any span in the transcript that sounds the same but is written differently is corrected back to the spelling of the term. Providing terminology alone is therefore enough to get correction; you do not need to list possible misspellings by hand.

For the full description of homophone correction (the reference table, the fact that it applies to Chinese only, common-word protection, and the pronunciation limits), see Terminology Guide → How Terminology Participates in Homophone Correction.

Limits

Per-block entry limits, length limits, and the matching error codes are collected in Terminology Guide → Limits.

The numbers are defaults; the limit actually in force can be tuned per environment — always treat max in the error response's details as authoritative rather than hard-coding the numbers.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value config
terminologyobjectNoTerminology settings
fuzzy_correctionobjectNoFuzzy-word correction rules
translation_dictobjectNoTranslation dictionary

Note: At least one setting item must be provided.

Note: All or nothing: all three blocks are validated before any of them is applied. If any block fails validation, an error is returned and none of the three blocks is applied — the settings stay as they were.

For example, sending a valid terminology list together with more than 3000 translation-dictionary entries returns config_too_many_dict_entries, and the terminology does not take effect either. Fix the problem and resend the complete config; all three blocks replace their previous value wholesale, so resending does not stack on top of earlier settings.

Terminology Format (terminology)

Keyed by language code, with an array of terms as the value:

{
  "zh-TW": [
    { "term": "語者分離" },
    { "term": "WebSocket" }
  ],
  "en-US": [
    { "term": "diarization" }
  ]
}
FieldTypeRequiredDescription
termstringYesThe term (max 100 characters)

Limit: Up to 500 terms across all languages combined in one config — not 500 per language. Exceeding the limit returns config_too_many_entries (with count and max in details).

Language scope: Terms apply only to recognition in the language they are registered under. Single-language situations — each multi-channel track, speaker diarization, and file import — use only that language's terms; multi-language transcription and conversation mode use the terms for the languages declared for the session. Terms registered under a language not used in this session have no effect, but still count toward the 500 combined total above.

These numbers are defaults: the limit actually in force can be tuned per environment; always treat max in the error response's details as authoritative.

Fuzzy-Word Correction Format (fuzzy_correction)

Note: This field usually does not need to be set manually — terminology already corrects misspellings that sound the same or nearly the same. Use it in these three cases:

  1. The misspelling is itself an ordinary word and is therefore blocked by common-word protection — 晶圓 heard as 金元, for example
  2. The misspelling sounds very different from the correct term, such as a foreign brand name recognized as a phonetically unrelated word
  3. Misspellings in Japanese, Korean or English, which do not participate in homophone matching

When correct is Chinese, it also becomes a reference for homophone matching — homophone misspellings not listed in incorrect are corrected to correct as well. case_insensitive applies only to the literal matching of incorrect; it does not affect homophone matching.

Keyed by language code, with an array of correction rules as the value:

{
  "zh-TW": [
    { "correct": "語者分離", "incorrect": ["語這分離", "語者分力"] },
    { "correct": "IPEVO", "incorrect": ["ltfo"], "case_insensitive": true }
  ]
}
FieldTypeRequiredDescription
correctstringYesThe correct term
incorrectstring[]ConditionalList of incorrect variants, each up to 200 characters. Can be omitted for Chinese terms (see below); required otherwise — an empty array returns config_invalid_entry (reason: "empty")
case_insensitivebooleanNoWhether this rule's variants match regardless of case (defaults to false = exact-case matching)

Supplying only the correct term: when correct is Chinese (contains Han characters), incorrect may be omitted entirely — the system matches by pronunciation, and spellings in the transcript that sound the same or nearly the same are corrected back to correct.

{ "fuzzy_correction": { "zh-TW": [{ "correct": "艾思通" }] } }

No misspellings need to be listed above: 愛思通, 愛時通, 愛司東 and 愛似通 are all corrected. Only spellings that sound quite different (愛自動, say) or have a different number of syllables (愛松) still need to be listed in incorrect.

Note: Both conditions must hold: the language must be Chinese (zh-TW, zh-CN, zh-HK and so on) and correct must contain Han characters. Otherwise incorrect remains required — omitting it in those cases would have no effect at all, and accepting it would leave you believing the setting took. List misspellings explicitly for Japanese, Korean and English.

Case sensitivity: case_insensitive is optional and defaults to false (exact-case matching). When set to true, every incorrect variant in that rule matches regardless of case. The flag is per rule — the same correct term can be split across several rules with different settings, for example making variants that cannot collide with ordinary words case-insensitive while keeping variants that could hit a personal name exact. It has no effect on Chinese rules (Chinese has no letter case).

Note: Enabling it widens the false-positive surface: if ivo is case-insensitive, the personal name Ivo is replaced too.

When the same incorrect variant appears in more than one rule: collisions are resolved on incorrect (the variant), not on correct.

  • When several rules point at different correct terms, which one actually takes effect is not guaranteed; do not rely on any ordering (registration order included)
  • The case flag resolves to strict wins — if any rule leaves case_insensitive off, that variant is matched with exact case

"Strict wins" is a deliberately conservative choice: it prevents a permissive rule elsewhere in your vocabulary from silently loosening a brand-name rule you explicitly set to strict.

Splitting one correct term across several rules is therefore safe, as long as their incorrect variants do not overlap. But if your data contains the same incorrect variant mapped to different correct terms, the later one is dropped with no warning — check for duplicate variants before sending.

Limit: Up to 4000 rules across all languages combined in one config. Exceeding it returns config_too_many_entries (with field: "fuzzy_correction", count and max in details).

This number is a default: the limit actually in force can be tuned per environment, so always treat max in the error response's details as authoritative. Do not hard-code the numbers in your integration — if you want to check before sending, read max and fill it back in. When you do hit a limit, details carries both count (what you sent) and max (the limit in force).

Language scope: A correction rule applies only to sentences in the language it is registered under — a rule under zh-TW will not alter an English sentence. When the sentence language cannot be determined, all rules are applied as a fallback (better to over-apply than to skip the sentence entirely). Homophone matching driven by terminology likewise follows the language each term is registered under.

Translation Dictionary Format (translation_dict)

Group entries by language code, giving each language its own dictionary:

{
  "en-US": [
    { "source": "語者分離", "target": "Speaker Diarization" },
    { "source": "晶圓", "target": "wafer", "case_sensitive": true }
  ],
  "ja-JP": [
    { "source": "語者分離", "target": "話者分離" }
  ]
}
FieldTypeRequiredDescription
(top-level key)stringYesTarget language code
sourcestringYesThe source word (in the STT language), up to 200 characters
targetstringYesThe required translation for this language, up to 200 characters
case_sensitivebooleanNoWhether the entry applies only on an exact-case match (defaults to false = case-insensitive)

Limit: Up to 3000 entries per language. Exceeding the limit returns config_too_many_dict_entries; the details in the response identify which language exceeded it.

Note: The entry count directly affects translation workload and cost. A single translation only carries entries whose source word actually appears in that piece of text, up to 100 of them; beyond that, longer source words are kept first. Also, the more entries there are, the smaller the share that is reliably honored — an inherent limit that a higher cap does not change.

Case sensitivity: case_sensitive is optional and defaults to false (case-insensitive). When set to true, the entry applies only where the source text matches source exactly, including case. The flag is per entry.

Note: The translation dictionary guides the model through prompting rather than literal substitution, so it is best-effort, not deterministic — the case flag is likewise a hint and is not guaranteed to be honored. Use fuzzy_correction when you need deterministic replacement.

Case-Flag Comparison

fuzzy_correction and translation_dict each have a case switch. Their field names are opposites, and so is the behavior their default value produces:

BlockFieldDefaultDefault behavior
fuzzy_correctioncase_insensitivefalseStrict (case-sensitive)
translation_dictcase_sensitivefalsePermissive (case-insensitive)

Both default to false, yet one means strict and the other means permissive. Do not share a single variable between them, and do not mirror one value onto the other — getting it wrong produces no error at all, only matching behavior opposite to what you intended.

{
  "type": "voice-translation",
  "data": {
    "action": "config",
    "terminology": {
      "zh-TW": [
        { "term": "語者分離" },
        { "term": "CVD製程" },
        { "term": "wafer良率" }
      ]
    }
  }
}

Request Example (Full settings, with manual correction rules)

{
  "type": "voice-translation",
  "data": {
    "action": "config",
    "terminology": {
      "zh-TW": [
        { "term": "語者分離" },
        { "term": "即時轉錄" }
      ]
    },
    "fuzzy_correction": {
      "zh-TW": [
        { "correct": "語者分離", "incorrect": ["語這分離", "語者分力"] }
      ]
    },
    "translation_dict": {
      "en-US": [{ "source": "語者分離", "target": "Speaker Diarization" }]
    }
  }
}

Successful Response

{
  "type": "voice-translation",
  "data": {
    "action": "config_updated",
    "updated": ["terminology", "fuzzy_correction", "translation_dict"],
    "message": "Settings updated"
  }
}

For response field descriptions, see the config_updated event.

Error Codes

Important: Your client must also listen for type: "error" messages — do not wait only for config_updated.

When the server rejects a configuration it sends a type: "error" message and notconfig_updated. An integration that waits only for config_updated will hang until its own timeout and appear as "the server never responded", even though the error was delivered and data.error_code already states the reason.

Error CodeHTTP StatusDescriptionRecommended Action
config_empty400No configuration provided. Note: An empty object {} does not count as "provided" — sending {"terminology": {}, "fuzzy_correction": {}, "translation_dict": {}} triggers this errorProvide at least one setting that actually has content. To clear a language's glossary, send {"lang": []} (for example {"zh-TW": []})
config_term_too_long400Term exceeds 100 charactersShorten the term
config_too_many_entries400More than 500 terms, or more than 4000 fuzzy correction rules (both across all languages combined)Remove terms or correction rules
config_too_many_dict_entries400Translation dictionary exceeds 3000 entries for a single language (details.language identifies which one)Reduce the dictionary entries for that language
config_invalid_entry400A glossary entry has an invalid field (details carries language, index, field, and reason for locating it; depending on the case it may also carry variant_index, max_length, or count/max)Fix the entry at the location given in details

audio - Send Audio

Description

Send audio data to the server for speech recognition. The audio must be Base64-encoded before sending.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value audio
payloadstringYesBase64-encoded audio data
channel_idintConditionalSource channel number (required on every frame in multi-channel mode, v1.10.0). Omitting it returns channel_id_required; an unknown or removed number returns unknown_channel_id. Ignored outside multi-channel mode

Audio Format Requirements

PCM format (default):

ItemSpecification
FormatPCM (raw audio)
Sample rate16000 Hz
Bit depth16-bit
ChannelsMono
Byte orderLittle-endian
Transfer encodingBase64

WebM/Opus format:

ItemSpecification
FormatWebM container + Opus codec
Sample rateAny (the server converts automatically)
ChannelsMono or Stereo (the server converts automatically)
Transfer encodingBase64

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "audio",
    "payload": "Base64-encoded PCM audio data"
  }
}

Request Example (Multi-Channel, v1.10.0)

In multi-channel mode every frame must include channel_id, identifying which microphone the audio came from:

{
  "type": "voice-translation",
  "data": {
    "action": "audio",
    "channel_id": 2,
    "payload": "Base64-encoded PCM audio data"
  }
}

Multi-channel sending tips: send frames of roughly 100ms; keep sending on every channel even while nobody is speaking (silent audio). A single silent channel does not end the recording; see Automatic End After a Long Silence.

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
session_not_started400Speech recognition has not startedCall the start action first
audio_invalid_format400Invalid audio data formatMake sure the Base64 encoding is correct
audio_decode_failed400Audio decoding failedMake sure the audio format is correct. The recording continues; for WebM, send a new container (with its header) to recover. Time that cannot be decoded is not billed
audio_process_failed500STT/diarization writes keep failing, exceeding the tolerance thresholdWe recommend reconnecting
channel_id_required400An audio frame in multi-channel mode is missing channel_id (v1.10.0)Include channel_id on every frame
unknown_channel_id400Unknown or removed channel_id (v1.10.0)Make sure the channel exists and has not been removed

pause - Pause Translation

Description

Pause speech recognition processing. Audio received while paused is cached and processing resumes afterward.

Two-way translation is an exception: audio received while paused is not kept. For a sentence in progress at the moment of pausing, the server first waits for the recognizer to finish the last part (usually about 1 second, at most about 3 seconds), sends it with is_final: true, and only then returns status: "paused"; in manual mode, if the user is speaking, speaking is ended automatically and that sentence's final result is sent.

Billing continues while paused. See Pricing.

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "pause"
  }
}

Successful Response

{
  "type": "voice-translation",
  "data": {
    "action": "status",
    "status": "paused",
    "message": "Speech recognition paused"
  }
}

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
session_not_started400Speech recognition has not startedCall start first
session_already_paused400Already pausedYou can ignore this error

resume - Resume Translation

Description

Resume paused speech recognition processing.

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "resume"
  }
}

Successful Response

{
  "type": "voice-translation",
  "data": {
    "action": "status",
    "status": "live",
    "message": "Speech recognition resumed"
  }
}

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
session_not_started400Speech recognition has not startedCall start first
session_not_paused400Not pausedYou can ignore this error

If a sentence is in progress when this happens, it cannot be kept and is reported separately via segment_discarded (reason: "resumed"). In multi-channel mode each channel reports separately.


stop - Stop Translation

Description

Stop speech recognition and end the session. The server first waits for the recognizer to finish the last sentence (usually about 1 second, at most about 3 seconds) so that it is included in the transcript; if it does not arrive in time, the last interim result shown is used as that sentence. In two-way translation manual mode, if the user is speaking, that sentence's final result and translation are sent first. The system automatically uploads the audio file and transcript and generates a summary (if configured; when the available credits cannot cover the summary fee, no summary is generated and summary_error is sent instead).

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "stop"
  }
}

Successful Response

{
  "type": "voice-translation",
  "data": {
    "action": "status",
    "status": "ended",
    "message": "Speech recognition stopped"
  }
}

After stopping, once the audio file and transcript have finished uploading, you receive a task_complete event containing task_id (the Recording UUID).

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
session_not_started400The session has not started, or this recording has already ended (for example, stop sent twice)Call start first if it has not started; if this was a duplicate stop, the error can be ignored

Sending stop a second time returns this error rather than another success response. task_complete is sent only once, after the first successful stop, when the audio and transcript have finished uploading.


retranslate - Retranslate a Single Sentence

Description

Retranslate a specified sentence. This is useful when the source text has been corrected and the translation needs to be updated.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value retranslate
sidintYesThe number of the sentence to retranslate
translation_languagesstring[]YesArray of translation language codes. Only the first element is retranslated; any additional languages are ignored. In multi-language sessions, send a separate retranslate request per language
textstringYesThe source text to translate (the user's corrected text)

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "retranslate",
    "sid": 1,
    "translation_languages": ["en-US"],
    "text": "The user's corrected source text"
  }
}

Successful Response

A translation event is returned (sharing the same schema as normal translation results), and the translation result includes is_retranslation: true:

{
  "type": "voice-translation",
  "data": {
    "action": "translation",
    "sid": 1,
    "translations": {
      "en-US": {
        "sid": 1,
        "text": "The new translation result",
        "is_final": true,
        "is_retranslation": true
      }
    }
  }
}

The data object also carries sid, identical to the sid inside each language of translations.

v1.5.6 documentation correction: Earlier documentation described retranslate returning action: "result", but on the wire it is actually action: "translation". If your client's dispatcher originally handled the result action, add a handler branch for the translation action.

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
invalid_data422No sid providedInclude sid
record_translation_not_allowed400Recording-only sessions do not support translationUse the transcribe type
retranslate_session_not_active400The session is not started or has endedCheck the session status
retranslate_no_target_lang400No target language providedProvide translation_languages
retranslate_no_text400No text to translate providedProvide the text parameter
retranslate_llm_not_ready503The translation service is not readyRetry later
retranslate_llm_failed500Translation service failedRetry later

If the translation service returns a specific failure code (for example llm_content_filtered when the content cannot be translated), that code is returned as-is instead of being wrapped in retranslate_llm_failed.


switch_language - Switch Language

Description

Switch or adjust translation languages during real-time translation. The behavior depends on the recording type and the number of translation languages:

  • General mode, single language (translation_languages has 1 entry): replaces the translation target language and automatically batch-retranslates all already-translated sentences
  • General mode, multiple languages (2 or more entries, v1.6.7): redefined as "add or remove a single language". The op parameter is required; omitting it returns a switch_language_op_required error
  • Two-way translation mode (conversation): switches the STT source language (the spoken language); the translation target switches automatically to the other language
  • Multi-channel mode (multi_channel, v1.10.0): disabled. Languages are bound to channels, so every form of switch_language (including op: "add" / op: "remove") returns multichannel_switch_language_not_allowed; to change a single channel's language, use set_channel_language instead

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value switch_language
translation_languagesstring[]ConditionalArray of translation language codes (required in general mode; only the first element is used as the operation target)
opstringConditionalv1.6.7 multi-language operation: add (add a language) or remove (remove a language). Required for multi-language sessions; single-language sessions may use add to expand into multi-language, or omit it to keep the existing replace semantics
transcription_languagesstring[]ConditionalThe target language to switch to (two-way translation mode; if omitted, automatically toggles to the other language)

Request Example (General Mode, Single-Language Replace)

{
  "type": "voice-translation",
  "data": {
    "action": "switch_language",
    "translation_languages": ["ja-JP"]
  }
}

Request Example (Multi-Language, Add a Language, v1.6.7)

{
  "type": "voice-translation",
  "data": {
    "action": "switch_language",
    "op": "add",
    "translation_languages": ["de-DE"]
  }
}

After a successful add: existing sentences are automatically backfilled with the new language (the response sequence is the same as single-language replace: language_switch_start → multiple batch_retranslation → language_switch_done), and subsequent sentences include the new language in real-time translation. The limit is 12 languages; exceeding it returns too_many_languages.

Billing reminder: after adding a language, billing reflects the new language count starting from the next billed minute.

Request Example (Multi-Language, Remove a Language, v1.6.7)

{
  "type": "voice-translation",
  "data": {
    "action": "switch_language",
    "op": "remove",
    "translation_languages": ["ko-KR"]
  }
}

After a successful remove, a translation_language_removed event is returned. Existing translations for that language are kept; subsequent sentences are no longer translated into it. At least 1 translation language must remain (removing the last one returns a switch_language_last_language error):

{
  "type": "voice-translation",
  "data": {
    "action": "translation_language_removed",
    "translation_language": "ko-KR",
    "translation_languages": ["en-US", "ja-JP"]
  }
}

translation_languages is an authoritative snapshot of the full translation-language set after removal; overwrite your local language set with it directly (see events.md).

Request Example (Two-Way Translation Mode)

Specify the switch target:

{
  "type": "voice-translation",
  "data": {
    "action": "switch_language",
    "transcription_languages": ["en-US"]
  }
}

Automatic toggle (no parameters):

{
  "type": "voice-translation",
  "data": {
    "action": "switch_language"
  }
}

Special behavior in two-way translation mode:

  • Two-way translation mode uses automatic language detection, so you usually don't need to switch the language manually
  • switch_language only updates the internal preference state
  • After a successful switch, a language_switched event is returned (not the language_switch_start/done sequence)
  • Switching to the same language returns a conversation_same_language warning

Response Sequence (General Mode)

After switching the language, you receive the following events in order:

  1. language_switch_start: notifies that the switch has started
{
  "type": "voice-translation",
  "data": {
    "action": "language_switch_start",
    "translation_language": "ja-JP",
    "translation_languages": ["en-US", "ja-JP"],
    "total_segments": 15
  }
}
  1. batch_retranslation (multiple): returns retranslation results sentence by sentence
{
  "type": "voice-translation",
  "data": {
    "action": "batch_retranslation",
    "sid": 3,
    "translations": {
      "ja-JP": {
        "sid": 3,
        "text": "今日はプロジェクトの進捗について話し合いましょう",
        "is_final": true,
        "is_retranslation": true
      }
    }
  }
}
  1. language_switch_done: notifies that the switch is complete
{
  "type": "voice-translation",
  "data": {
    "action": "language_switch_done",
    "translation_language": "ja-JP",
    "translation_languages": ["en-US", "ja-JP"],
    "success_count": 15,
    "failed_count": 0
  }
}

Language-set sync: op:add and single-language replace emit identical events (both go through language_switch_start/done). These events carry translation_languages (the current full set of translation languages) — overwrite your local language set with it directly; do not infer append vs replace from the single translation_language (when only 1 language exists, an op:add expanding to a 2nd would be misread as a replace and drop the existing language).

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
switch_language_no_target400No target language providedProvide translation_languages
switch_language_in_progress400The previous switch is not yet completeWait for the switch to complete
switch_language_same_target400The target language is the same as the current oneYou can ignore this error
switch_language_op_required400op is missing in a multi-language session (v1.6.7)Provide op: "add" or op: "remove"
switch_language_already_exists400The language to add is already in the translation list (v1.6.7)You can ignore this warning
switch_language_not_in_session400The language to remove is not in the translation list (v1.6.7)Check the language code
switch_language_last_language400At least one translation language must remain (v1.6.7)The last language cannot be removed
too_many_languages400Adding would exceed the limit of 12 languages (v1.6.7)Remove an existing language first
invalid_translation_language400The language code to add is invalid (v1.6.7)Check the supported language list
conversation_requires_two_languages400Two-way translation mode requires exactly two languagesMake sure transcription_languages has 2 entries
conversation_languages_identical400The two languages in two-way translation cannot be the sameProvide two different languages
conversation_invalid_language400Invalid two-way translation languageMake sure the language is in transcription_languages
conversation_same_language400Already the current languageYou can ignore this warning
multichannel_switch_language_not_allowed400switch_language is disabled in multi-channel mode (v1.10.0)Use set_channel_language to change a single channel's language

set_name - Set Recording Name

Description

Set the name during a recording. After it is set, name_source flips to user, and the system will not override it when the recording ends (even if the LLM generates a summary name, it yields to the user-set name). For the full semantics and priority of name_source, see § Recording Name Rules above.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value set_name
namestringYesRecording name (max 60 characters)

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "set_name",
    "name": "Product Planning Meeting"
  }
}

Successful Response

{
  "type": "voice-translation",
  "data": {
    "action": "status",
    "event": "name_set",
    "name": "Product Planning Meeting",
    "message": "Recording name set"
  }
}

Compatibility note: For backward compatibility, the set_name success response keeps action: "status" (unchanged). New clients should identify a successful set_name via event: "name_set" (together with the name field). Relying on action: "status" to detect set_name success is deprecated and may be removed in a future version.

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
set_name_empty400Recording name is emptyProvide a non-empty name
set_name_too_long400Recording name exceeds the length limit (>60 chars); the response details includes max_lengthShorten the name (≤60 characters)

rename_speaker - Globally Rename a Speaker

Description

In multi-speaker diarization mode (multi_speaker), globally rename a speaker. All sentences that use that speaker ID are updated in sync.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value rename_speaker
speaker_idstringYesThe original speaker ID (e.g. Guest-1); also accepts the current display label for consecutive renames; max 100 characters
new_labelstringYesThe new display label; max 100 characters, must not contain control characters (\x00-\x1F, \x7F) or line breaks

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "rename_speaker",
    "speaker_id": "Guest-1",
    "new_label": "Manager Wang"
  }
}

Successful Response

Returns the speaker_renamed event:

{
  "type": "voice-translation",
  "data": {
    "action": "speaker_renamed",
    "speaker_id": "Guest-1",
    "new_label": "Manager Wang",
    "affected_sids": [1, 3, 5, 8]
  }
}

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
speaker_not_found422The specified speaker was not foundMake sure the speaker ID or alias exists
speaker_name_empty422The speaker name cannot be emptyProvide a valid name
speaker_name_duplicate422The speaker name is already in useUse another name, or first rename the conflicting speaker
session_not_started400Speech recognition has not startedCall start first

reassign_speaker - Change the Speaker of a Single Sentence

Description

Change the speaker identity of a specific sentence, assigning the sentence to an existing speaker.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value reassign_speaker
sidintYesThe number of the sentence to change
target_speaker_idstringYesThe target speaker's original ID (taken from init_sentence.speaker_id; reassign does not accept display labels)

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "reassign_speaker",
    "sid": 5,
    "target_speaker_id": "Guest-2"
  }
}

Successful Response

Returns the speaker_reassigned event:

{
  "type": "voice-translation",
  "data": {
    "action": "speaker_reassigned",
    "sid": 5,
    "old_speaker_id": "Guest-1",
    "new_speaker_id": "Guest-2",
    "new_speaker_label": "Lisa Lee"
  }
}

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
speaker_sid_not_found422The specified sentence was not foundMake sure the SID exists
speaker_not_found422The target speaker does not existUse an existing speaker ID
speaker_name_empty422The target speaker ID cannot be emptyProvide a valid speaker ID
session_not_started400Speech recognition has not startedCall start first
invalid_parameter400Creating a new speaker is not supportedUse an existing speaker ID

merge_speakers - Merge Speakers

Description

Merge all sentences from one speaker into another. After merging, future recognition results from that speaker are also automatically converted to the target speaker. This applies within the current recognition pass only: after a connection recovers, speakers are re-identified and you need to merge again.

Difference from reassign_speaker

FeatureScopeFuture Effect
reassign_speakerA single sentence (1 SID)None
merge_speakersAll sentences of that speakerFuture occurrences of the source are also automatically converted to the target

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value merge_speakers
source_speaker_idstringYesThe speaker ID to be merged (e.g. Guest-2)
target_speaker_idstringYesThe target speaker ID to merge into (e.g. Guest-1)

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "merge_speakers",
    "source_speaker_id": "Guest-2",
    "target_speaker_id": "Guest-1"
  }
}

Successful Response

Returns the speakers_merged event:

{
  "type": "voice-translation",
  "data": {
    "action": "speakers_merged",
    "source_speaker_id": "Guest-2",
    "target_speaker_id": "Guest-1",
    "affected_sids": [3, 5, 7]
  }
}

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
speaker_not_found422The speaker does not existMake sure the speaker ID exists
merge_speakers_same_id400The source and target speakers are the sameUse different speaker IDs
speaker_name_empty422The speaker ID cannot be emptyProvide a valid speaker ID
session_not_started400Speech recognition has not startedCall start first

tts_play - Play TTS

Description

In async mode, manually play the TTS audio of a specified sentence. Repeated requests for the same sid are supported (replay).

Two-way translation mode (conversation): tts_play automatically synthesizes the translation in the appropriate language based on the voice settings in tts_config; you don't need to specify tts_language separately.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value tts_play
sidintYesThe starting sentence ID
lengthintNoThe number of sentences to play (default 1, max 20)

Request Example (Single Sentence)

{
  "type": "voice-translation",
  "data": {
    "action": "tts_play",
    "sid": 5
  }
}

Request Example (Multiple Sentences)

{
  "type": "voice-translation",
  "data": {
    "action": "tts_play",
    "sid": 5,
    "length": 3
  }
}

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
tts_not_enabled400TTS not enabledMake sure TTS was enabled at start

A missing sentence or translation does not return error: If the starting sid does not exist, or the sentence has no translation in the target language, no error message is returned. A tts_error event is sent instead, with error set to sentence_not_found and translation_not_found respectively. When playing multiple sentences, a failing sentence is skipped and the rest still play.


tts_stop - Stop TTS

Description

Stop the currently playing TTS audio.

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "tts_stop"
  }
}

Successful Response

{
  "type": "voice-translation",
  "data": {
    "action": "status",
    "message": "TTS stopped"
  }
}

tts_mode - Switch TTS Mode

Description

Switch the TTS playback mode (synchronous/asynchronous) during a recording.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value tts_mode
tts_modestringYesMode: sync (synchronous) or async (asynchronous); only these two lowercase values are accepted

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "tts_mode",
    "tts_mode": "async"
  }
}

Successful Response

Returns the tts_mode_changed event:

{
  "type": "voice-translation",
  "data": {
    "action": "tts_mode_changed",
    "tts_mode": "async"
  }
}

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
invalid_data422No tts_mode provided, or a value other than sync or async. For the latter, details.field is tts_mode and details.valid_values lists the accepted values; the mode does not change and no tts_mode_changed is sentInclude sync or async

set_tts - Two-Way Translation TTS Settings

Description

During two-way translation mode (conversation), toggle TTS on/off or update the TTS voice settings mid-session. Available only under the conversation type.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value set_tts
tts_enabledbooleanNoToggle TTS on/off
tts_configobjectNoUpdate the TTS settings for a specific language (only the two two-way translation languages are valid)

Request Example (Disable TTS)

{
  "type": "voice-translation",
  "data": {
    "action": "set_tts",
    "tts_enabled": false
  }
}

Request Example (Update TTS Voice)

{
  "type": "voice-translation",
  "data": {
    "action": "set_tts",
    "tts_enabled": true,
    "tts_config": {
      "en-US": { "voice": "en-US-GuyNeural", "speaking_rate": 1.2 }
    }
  }
}

Successful Response

Returns the tts_updated event:

{
  "type": "voice-translation",
  "data": {
    "action": "tts_updated",
    "tts_enabled": true,
    "tts_config": {
      "zh-TW": { "voice": "zh-TW-HsiaoChenNeural", "speaking_rate": 1.0 },
      "en-US": { "voice": "en-US-GuyNeural", "speaking_rate": 1.2 }
    }
  }
}

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
invalid_action400This operation is not supported outside two-way translation modeUse only under the conversation type

start_speaking - Start Speaking (Manual Mode)

Description

In two-way translation manual mode (conversation_mode: "manual"), notify the system that the user has started speaking. From this moment, audio is sent to STT for recognition, and all recognition results accumulate into the same sentence (no automatic segmentation). If it is called again while already speaking, the system first ends the previous sentence (waiting for the last part to finish, then sending its final result) and starts a new one.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value start_speaking
speakerintYesUser number (1 or 2)

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "start_speaking",
    "speaker": 1
  }
}

Successful Response

{
  "type": "voice-translation",
  "data": {
    "action": "status",
    "message": "Speaking started"
  }
}

If it is called again while already speaking, the final result of the previous sentence is sent as usual, followed by this same status.

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
invalid_action400Not in two-way translation modeUse only under the conversation type
conversation_not_manual_mode400Not in manual modeUse only in manual mode
conversation_invalid_speaker400Invalid user numberUse 1 or 2

stop_speaking - Stop Speaking (Manual Mode)

Description

In two-way translation manual mode, notify the system that the user has stopped speaking. The system first waits for the recognizer to finish the last part (usually about 1 second, at most about 3 seconds), then merges the recognition results accumulated during this period into one complete sentence, and translates it and synthesizes TTS. For languages that separate words with spaces (such as English), a space is added between parts automatically.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value stop_speaking

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "stop_speaking"
  }
}

Successful Response

After speaking stops, the system sends a complete result event (containing origin and translations):

{
  "type": "voice-translation",
  "data": {
    "action": "result",
    "origin": {
      "sid": 1,
      "language": "zh-TW",
      "text": "The complete sentence merged from all recognition during this period",
      "is_final": true,
      "speaker_id": "Speaker-1",
      "start_time": "00:05"
    },
    "translations": {
      "en-US": {
        "sid": 1,
        "text": "The complete merged sentence from this speaking period",
        "is_final": true
      }
    }
  }
}

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
invalid_action400Not in two-way translation modeUse only under the conversation type
conversation_not_speaking400Not in the speaking stateCall start_speaking first

switch_conversation_mode - Switch Conversation Mode

Description

During two-way translation mode, switch between automatic detection mode (auto) and manual mode (manual). If the user is speaking during the switch, speaking ends automatically.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value switch_conversation_mode
conversation_modestringYesTarget mode: auto or manual

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "switch_conversation_mode",
    "conversation_mode": "manual"
  }
}

Successful Response

Returns the conversation_mode_changed event:

{
  "type": "voice-translation",
  "data": {
    "action": "conversation_mode_changed",
    "conversation_mode": "manual"
  }
}

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
invalid_action400Not in two-way translation modeUse only under the conversation type
conversation_invalid_mode400Invalid conversation modeUse auto or manual

set_speaker_language - Set Speaker Language

Description

During two-way translation mode, change a specified user's language in real time. The system rebuilds the STT connection to adapt to the new language, and the translation target is updated automatically. Transcript content before the change keeps its original language, and timestamps continue counting without resetting.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value set_speaker_language
speakerintYesUser number (1 or 2)
languagestringYesThe new language code (e.g. ja-JP)

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "set_speaker_language",
    "speaker": 1,
    "language": "ja-JP"
  }
}

Successful Response

Returns the speaker_language_changed event:

{
  "type": "voice-translation",
  "data": {
    "action": "speaker_language_changed",
    "speaker_language_map": {
      "1": "ja-JP",
      "2": "en-US"
    }
  }
}

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
invalid_action400Not in two-way translation modeUse only under the conversation type
conversation_invalid_speaker400Invalid user numberUse 1 or 2
conversation_invalid_language400No language providedInclude language
invalid_transcription_language400Invalid language codeUse a valid BCP 47 language code
session_not_started400Recording has not startedCall start first
conversation_same_language400Same as the current languageYou can ignore this warning
conversation_language_same_as_peer400The new language is the same as the other user'sThe two users cannot have the same language
conversation_speaking400Currently speaking, cannot change the languageEnd speaking first, then change
conversation_language_change_failed500Language change failed (STT rebuild failed)Retry later

set_speaking_speed - Change Speaking Speed During Recording

Description

Dynamically adjust the speaking speed (which controls the silence threshold for segmentation) while recording. The system rebuilds the STT connection to apply the new setting, causing a brief interruption in recognition (same as changing a speaker's language during recording). Supported in all recognition modes except multi-speaker (multi_speaker) — including broadcast and multi-language LID; speaking_speed is not applied in multi-speaker mode.

Multi-channel mode (multi_channel, v1.10.0): supported. The new segmentation threshold applies to all channels, taking effect channel by channel; each channel sends its own channel_status event (reason: "speaking_speed" — first preparing, then ready once that channel starts producing text again), and produces no text for about 4 seconds while the change is applied. Like set_channel_language, there is a minimum 5-second interval between changes (calling too frequently returns channel_rebuild_too_frequent; a single adjustment applies to all channels, so this interval is shared across the whole session). While paused it returns channel_action_while_paused — call resume first, then adjust.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value set_speaking_speed
speaking_speedstringYesvery_slow / slow / normal / fast / very_fast (see speaking_speed levels for each threshold)

Request Example

{
  "type": "voice-translation",
  "data": { "action": "set_speaking_speed", "speaking_speed": "slow" }
}

Success Response

{
  "type": "voice-translation",
  "data": { "action": "speaking_speed_changed", "speaking_speed": "slow" }
}

Error Codes

Error CodeHTTPDescriptionSuggested Handling
invalid_data422Invalid speaking_speed valueUse one of the five values
session_not_started400Recording not startedCall start first
invalid_action400Not supported in multi-speakerDo not call in multi-speaker
set_speaking_speed_failed400Failed to rebuild STTRetry later
channel_rebuild_too_frequent400Multi-channel: speed adjustments are too frequent (only one accepted every 5 seconds, v1.10.0)Retry later
channel_action_while_paused400Multi-channel: the speed cannot be adjusted while paused (v1.10.0)Call resume first, then adjust

If a sentence is in progress when this happens, it cannot be kept and is reported separately via segment_discarded (reason: "speaking_speed"). You may receive it even when set_speaking_speed_failed is returned — recognition is interrupted before the rebuild.


add_channel - Add a Channel (Multi-Channel)

Description

Dynamically add a channel while a multi-channel recording (recognition_mode: "multi_channel", v1.10.0) is in progress: when a participant joins on the spot, you can open a new channel for them without stopping the recording. Once the channel is added, you can start sending audio frames with that channel_id.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value add_channel
channelsobject[]YesExactly 1 channel configuration, with the same fields as channels[] in start (channel_id, optional speaker_name; transcription_languages with exactly 1 entry under per_channel, omitted under shared)

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "add_channel",
    "channels": [
      { "channel_id": 4, "speaker_name": "Chen", "transcription_languages": ["zh-TW"] }
    ]
  }
}

Successful Response

Returns a channel_status event (reason: "added", status: "preparing"):

{
  "type": "voice-translation",
  "data": {
    "action": "channel_status",
    "channel": {
      "channel_id": 4,
      "speaker_name": "Chen",
      "transcription_languages": ["zh-TW"],
      "status": "preparing"
    },
    "active_channels": 4,
    "stt_stream_count": 4,
    "reason": "added"
  }
}

When that channel recognizes its first sentence, another channel_status arrives (status: "ready", with reason carried over from the triggering added).

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
not_multi_channel_session400This recording is not in multi-channel modeUse only with recognition_mode: "multi_channel"
channels_required400channels is missing or does not contain exactly 1 elementAdd one channel at a time
invalid_channel_id400channel_id is out of range (1–8)Fix the channel number
channel_id_in_use400The number is in use, or has already been removed (numbers cannot be reused)Use a number that has never been used
speaker_name_duplicate422speaker_name duplicates another channel's current nameUse a different name (a name freed up by a rename can be taken)
too_many_channels400The channel count has reached the limitCall remove_channel first, then add
channel_language_required400Exactly one transcription language was not specified (per_channel)Provide exactly 1 entry in transcription_languages
channel_language_not_allowed400transcription_languages was provided in shared modeRemove the field
invalid_transcription_language400Invalid language codeMake sure the language code format is correct (e.g. zh-TW)
invalid_parameter400speaker_name is too long (>100 characters) or contains control charactersFix the name
plan_feature_not_allowed403Exceeds the plan's channel count limit (details.field is max_stt_streams)Remove another channel or upgrade the plan
too_many_languages400The new channel's language pushes the number of simultaneous recognition languages over the plan limitUse an existing language or upgrade the plan
channel_action_while_paused400Channels cannot be added or disabled while pausedCall resume first
session_not_started400Recording has not startedCall start first
stt_start_failed500Speech recognition failed to start on the channel (this add did not take effect)Retry later

Notes

  • Numbers cannot be reused: once a channel_id has been used (including removed ones), it is occupied permanently — speaker identity is fixed at the moment it is written into the transcript, and reusing a number would make two different people share the same speaker identity.
  • Startup delay: after a successful add, the channel begins producing text within about 4 seconds; audio sent during this period is not lost, only delayed.
  • The new channel's timeline is automatically aligned with the session (someone joining at second 65 gets speech timestamps starting from about second 65).
  • The billed channel count is taken as each minute begins, so an added channel is billed from the next minute — see the Pricing Guide.
  • shared mode: the added channel shares the session's recognition and does not increase the billed channel count; its status takes on the first channel's current status directly.

remove_channel - Disable a Channel (Multi-Channel)

Description

Disable a channel while a multi-channel recording is in progress (v1.10.0): when a participant leaves early, disabling their channel stops billing for that channel from the next billed minute onward. The semantics are "disable", not "delete" — the transcript and single-track audio that channel has already produced are all preserved.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value remove_channel
channel_idintYesThe channel number to disable

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "remove_channel",
    "channel_id": 4
  }
}

Successful Response

Returns a channel_status event (reason: "removed", status: "removed"):

{
  "type": "voice-translation",
  "data": {
    "action": "channel_status",
    "channel": {
      "channel_id": 4,
      "speaker_name": "Chen",
      "transcription_languages": ["zh-TW"],
      "status": "removed"
    },
    "active_channels": 3,
    "stt_stream_count": 3,
    "reason": "removed"
  }
}

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
not_multi_channel_session400This recording is not in multi-channel modeUse only with recognition_mode: "multi_channel"
channel_id_required400channel_id is missingProvide the number of the channel to disable
invalid_channel_id400channel_id is out of range (1–8)Fix the channel number
unknown_channel_id400Unknown channel_id (it may have been removed already)Make sure the channel exists and has not been removed
channel_remove_not_allowed400The last remaining channel cannot be disabled; or the first channel in shared modeUse stop to end the recording
channel_action_while_paused400Channels cannot be added or disabled while pausedCall resume first
session_not_started400Recording has not startedCall start first

Notes

  • Wrap-up window: after disabling, there is a wrap-up window of about 3 seconds so that a final sentence already spoken on that channel but not yet returned can still make it into the transcript; the channel no longer accepts new audio during this period.
  • Numbers cannot be reused: once disabled, that channel_id cannot be used with add_channel again (it returns channel_id_in_use).
  • Sending audio with that number after disabling returns unknown_channel_id.
  • Billing: one fewer channel is billed starting from the next billed minute — see the Pricing Guide.

set_channel_language - Change a Channel's Language (Multi-Channel)

Description

Change the transcription language of a single channel while a multi-channel recording is in progress (v1.10.0). The channel_id and speaker identity (speaker_id) stay the same and the transcript remains continuous — which is exactly why this action exists: with remove_channel + add_channel, channel numbers cannot be reused, so the same person would become two different speakers in the transcript.

The system switches that channel to the new language (other channels are unaffected). The switch takes about 4 seconds to take effect, during which the channel produces no text.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value set_channel_language
channel_idintYesTarget channel number
transcription_languagesstring[]One of the twoThe new transcription language, exactly 1 (same shape as channels[] in start)
languagestringOne of the twoThe new transcription language (single-value form). Providing both fields with different values returns invalid_parameter

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "set_channel_language",
    "channel_id": 3,
    "transcription_languages": ["en-US"]
  }
}

Successful Response

Returns a channel_status event (reason: "language_change", status: "preparing"):

{
  "type": "voice-translation",
  "data": {
    "action": "channel_status",
    "channel": {
      "channel_id": 3,
      "speaker_name": "Sato",
      "transcription_languages": ["en-US"],
      "status": "preparing"
    },
    "active_channels": 3,
    "stt_stream_count": 3,
    "reason": "language_change"
  }
}

When the channel recognizes its first sentence in the new language, another channel_status arrives (status: "ready").

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
not_multi_channel_session400This recording is not in multi-channel modeUse only with recognition_mode: "multi_channel"
channel_id_required400channel_id is missingProvide the target channel number
invalid_channel_id400channel_id is out of range (1–8)Fix the channel number
unknown_channel_id400Unknown channel_id (it may have been removed)Make sure the channel exists and has not been removed
channel_language_required400No new language providedProvide transcription_languages or language
invalid_parameter400More than one language provided / the two language fields disagree / same as the channel's current languageFix the parameters according to details
invalid_transcription_language400Invalid language codeMake sure the language code format is correct (e.g. zh-TW)
too_many_languages400The new language pushes the number of simultaneous recognition languages over the platform limit of 10 or the plan limitUse an existing language or upgrade the plan
channel_rebuild_too_frequent400The same channel was changed again within 5 secondsRetry later (details.cooldown_seconds is the number of seconds to wait)
channel_action_while_paused400The channel language cannot be changed while pausedCall resume first
session_not_started400Recording has not startedCall start first
channel_language_not_allowed400shared mode does not support changing a channel's language (v1.21.0)The session language is set at start and cannot be changed during the recording
stt_start_failed500The change did not take effect (a channel_status event with status: "error" is also sent; the system will try to recover automatically)Wait for automatic recovery or retry later

Notes

  • Minimum interval between changes (5 seconds): two language changes on the same channel must be at least 5 seconds apart (channel_rebuild_too_frequent); the interval starts counting only after all validation passes — rejected requests do not consume it.
  • The language set only grows: when a language change introduces a language the session has not used yet, that language is merged into the session-level transcription_languages; the old language is not removed (the channel's earlier transcript is still in the old language). Language changes are therefore governed by the platform limit on simultaneous recognition languages (10) and the plan's language limit.
  • Changing to the language the channel currently uses returns invalid_parameter (the no-op change is not performed).
  • switch_language is always disabled in multi-channel mode (it returns multichannel_switch_language_not_allowed); always use this action to change languages.

If a sentence is in progress when this happens, it cannot be kept and is reported separately via segment_discarded (reason: "language_change"). The event carries channel_id to identify the channel.


set_summary - Change Summary Settings During Recording

Description

Change the summary settings that will be applied when the recording stops. A live recording generates its summary exactly once, at stop time, so "takes effect immediately" here means "that automatic summary will use the latest settings" — if you send this action several times, the last one before stopping wins.

Typical use: you decide to switch to a different summary template or custom prompt after the recording has already started. Updating in place avoids having to regenerate the summary after stopping, which saves one summary generation charge.

Parameters fall into two independent groups:

  • Summary source group: summary_mode, summary_template, summary_prompt, summary_prompt_slug. Applied only when summary_mode is present, and the four fields are replaced as a set — switching modes automatically clears the fields belonging to the other mode.
  • Option group: auto_summary, summary_language, summary_plain_text. Each is independent and is updated only when present.

The two groups do not affect each other: updating the summary source does not turn auto_summary on. If automatic summaries were disabled when the recording started, you must also send auto_summary: true to resume generation.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value set_summary
summary_modestringConditionalbuiltin (shared template) or custom (custom prompt). Required when changing the summary source; may be omitted when updating only the option group
summary_templatestringConditionalTemplate identifier for builtin mode. Not allowed in custom mode
summary_promptstringConditionalFull prompt for custom mode, up to 3000 characters. Required in custom mode (a value with only whitespace counts as not provided)
summary_prompt_slugstringConditionalIdentifier for custom mode, up to 64 characters. Required in custom mode (a value with only whitespace counts as not provided); not allowed in builtin mode
auto_summarybooleanNofalse skips summary generation at stop; true resumes it
summary_languagestringNoSummary output language, up to 20 characters
summary_plain_textbooleanNoWhether to output plain text (Markdown removed)

You must provide either summary_mode or at least one field from the option group; otherwise the request returns invalid_data. Unlike start, this action does not infer builtin when summary_mode is missing. Sending summary source fields without a mode returns summary_mode_field_mismatch.

Summary Source Requirements

The source is validated only when the request affects the summary source or whether summaries are generated — that is, when it carries the summary source group or explicitly sends auto_summary: true:

  • Changing only summary_language / summary_plain_text → not validated. Such a request changes neither whether a summary is produced nor the source, so recordings that never specified a template can send it too.
  • Sending auto_summary: false → not validated; you may send only the option group.
  • Carrying the summary source group, or explicitly sending auto_summary: true → evaluated in order:
    1. This request includes summary_template or summary_prompt → the new value is used.
    2. This request includes neither, or includes only whitespace → treated as not provided, and the previous setting is kept; if none was ever specified, the request returns summary_mode_field_mismatch asking you to provide one.

Recordings without a summary_template (conversation and broadcast may omit it) fall back to the system default template when generating the summary. To explicitly enable the automatic summary or change the source on such a recording, you still need to supply a template or a custom prompt.

Request Example (switch to a custom prompt)

{
  "type": "voice-translation",
  "data": {
    "action": "set_summary",
    "summary_mode": "custom",
    "summary_prompt": "List the decisions and action items from this meeting, and note the owner of each item.",
    "summary_prompt_slug": "meeting-actions-v2"
  }
}

Request Example (switch to a shared template)

{
  "type": "voice-translation",
  "data": {
    "action": "set_summary",
    "summary_mode": "builtin",
    "summary_template": "meeting"
  }
}

Request Example (skip the automatic summary)

{
  "type": "voice-translation",
  "data": { "action": "set_summary", "auto_summary": false }
}

Success Response

Returns the settings currently in effect. For safety the response omits the full summary_prompt; use summary_prompt_slug to identify it:

{
  "type": "voice-translation",
  "data": {
    "action": "summary_updated",
    "summary_mode": "custom",
    "summary_prompt_slug": "meeting-actions-v2",
    "summary_language": "zh-TW",
    "auto_summary": true,
    "summary_plain_text": false,
    "message": "Summary settings updated"
  }
}

Error Codes

Error CodeHTTPDescriptionSuggested Handling
invalid_data422No fields provided, or summary_language too longInclude at least one updatable field
summary_invalid_mode400summary_mode is neither builtin nor customUse one of the two valid values
summary_mode_field_mismatch400Field combination conflicts with the mode, or the source is missingAdd the field named by missing_field in details
summary_prompt_too_long400summary_prompt exceeds 3000 charactersShorten the prompt
summary_prompt_slug_too_long400summary_prompt_slug exceeds 64 charactersShorten the identifier
summary_prompt_slug_invalid400summary_prompt_slug contains control charactersRemove the control characters
invalid_summary_template400The template does not exist or is not enabledUse a valid template
session_not_started400Recording has not started, or has already stoppedCall start first; if stopped, use the regeneration API

Notes

  • After reconnecting, the settings object in resume_ok returns the updated summary settings, so you can reconcile client state directly.
  • Calling this action after stop returns session_not_started. A summary that is already being generated is unaffected.
  • After the recording stops you can still regenerate the summary with any template through the summary regeneration endpoint.

broadcast_go_live - Switch to the Live Phase

Description

Switch from the broadcast standby phase (standby) to the live phase (live). After the switch, STT/translation results begin broadcasting to viewers and start being written to the transcript.

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "broadcast_go_live"
  }
}

Successful Response

Returns the broadcast_phase_changed event:

{
  "type": "voice-translation",
  "data": {
    "action": "broadcast_phase_changed",
    "phase": "live",
    "message": "Broadcast started"
  }
}

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
broadcast_not_enabled400Not in broadcast modeMake sure type: "broadcast"
session_not_started400Speech recognition has not started, or the recording has ended (including while it is still being processed after ending)Call start first if it has not started

Note: If already in the live phase, a status message "Broadcast is already in progress" is returned; this is not treated as an error.

If a sentence is in progress when this happens, it cannot be kept and is reported separately via segment_discarded (reason: "broadcast_go_live"). Standby-phase segments are never written to the final transcript anyway.


broadcast_announcement - Send an Announcement

Description

The host sends a custom announcement message to all viewers. Viewers receive an announcement event via SSE. The announcement message is automatically translated into all translation languages, and the SSE event viewers receive includes a translations field.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value broadcast_announcement
messagestringYesThe announcement message content

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "broadcast_announcement",
    "message": "The meeting will end in 5 minutes"
  }
}

Successful Response

{
  "type": "voice-translation",
  "data": {
    "action": "status",
    "message": "Announcement sent"
  }
}

The SSE event received on the viewer side (with translations):

event: announcement
data: {"message":"The meeting will end in 5 minutes","translations":{"en-US":"The meeting will end in 5 minutes","ja-JP":"会議は5分後に終了します"}}

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
broadcast_not_enabled400Not in broadcast modeMake sure type: "broadcast"
invalid_parameter400The message is emptyProvide a valid message parameter

set_standby_message - Set the Standby Phase Message

Description

Dynamically set the message displayed to viewers during the broadcast standby phase (standby). This lets the host enter standby mode first and set the waiting message afterward, rather than having to provide it at start.

The message is automatically translated into all translation languages, and the SSE event viewers receive includes a translations field.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value set_standby_message
messagestringYesThe standby phase display text (translated for viewers in each language via the translation pipeline)

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "set_standby_message",
    "message": "The talk is about to begin, please wait..."
  }
}

Successful Response

{
  "type": "voice-translation",
  "data": {
    "action": "status",
    "message": "Standby phase text updated"
  }
}

The SSE event received on the viewer side (with translations):

event: standby
data: {"message":"The talk is about to begin, please wait...","translations":{"en-US":"The presentation is about to begin, please wait...","ja-JP":"プレゼンテーションがまもなく始まります。お待ちください..."}}

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
broadcast_not_enabled400Not in broadcast modeMake sure type: "broadcast"
broadcast_not_in_standby400Not in the standby phaseCan only be used during the standby phase

Note: This action can only be used during the standby phase (standby). If you have already entered the live phase (live), an error is returned.


Version: V1.24.1 Last Updated: 2026-10-07

Copyright © 2026