WebSocket API
Note: This is a consolidated document. For detailed specifications, refer to the individual documents under reference/websocket/.
Note: The URL used in this document (
vas-poc.vurbo.ai) is the planned deployment address. A separate notice will be issued after the official launch.
Table of Contents
- Connection Info
- Authentication
- Message Format
- Health - Heartbeat Service
- Voice Translation - start
- Voice Translation - config
- Voice Translation - audio
- Voice Translation - pause
- Voice Translation - resume
- Voice Translation - stop
- Voice Translation - retranslate
- Voice Translation - switch_language
- Voice Translation - set_name
- Voice Translation - rename_speaker
- Voice Translation - reassign_speaker
- Voice Translation - merge_speakers
- Voice Translation - tts_play
- Voice Translation - tts_stop
- Voice Translation - tts_mode
- Voice Translation - set_tts
- Voice Translation - start_speaking
- Voice Translation - stop_speaking
- Voice Translation - switch_conversation_mode
- Voice Translation - set_speaker_language
- Voice Translation - set_speaking_speed
- Voice Translation - add_channel
- Voice Translation - remove_channel
- Voice Translation - set_channel_language
- Voice Translation - broadcast_go_live
- Voice Translation - broadcast_announcement
- Voice Translation - set_standby_message
- Response Events
Connection Info
| Item | Value |
|---|---|
| Endpoint | wss://vas-poc.vurbo.ai/ws |
| Protocol | WebSocket |
| Data Format | JSON |
| Auth Method | Ticket (see below) |
Authentication
The VAS WebSocket uses a Ticket mechanism for authentication, passing a one-time Ticket via Sec-WebSocket-Protocol. For details, refer to Authentication.
Step 1: Obtain a Ticket
Exchange your API Key for a one-time Ticket via the REST API:
POST /api/v1/auth/ticket
X-API-Key: vas_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
Response:
{
"ticket": "aBcDeFgHiJkLmNoPqRsTuVwXyZ012345",
"expires_in": 60
}
| Field | Type | Description |
|---|---|---|
ticket | string | One-time Ticket (32 chars) |
expires_in | int | Validity period (seconds) |
Step 2: Connect to the WebSocket using the Ticket
Place the Ticket into Sec-WebSocket-Protocol in the format ticket.{TICKET_VALUE}:
// Native browser support
const ws = new WebSocket('wss://vas-poc.vurbo.ai/ws', [`ticket.${ticket}`]);
ws.onopen = () => {
console.log('Connected! Protocol:', ws.protocol);
// Start using the WebSocket...
};
ws.onerror = (error) => {
console.error('Connection failed:', error);
};
Node.js example:
const WebSocket = require('ws');
const ws = new WebSocket('wss://vas-poc.vurbo.ai/ws', [`ticket.${ticket}`]);
Ticket Characteristics
| Characteristic | Description |
|---|---|
| Validity period | 60 seconds |
| Usage count | Can be used only once (deleted immediately after) |
| Security | The API Key is never exposed in the WebSocket connection |
| Replay protection | Uses an atomic operation to guarantee single use |
Ticket Error Codes
| Error Code | HTTP Status | Description |
|---|---|---|
ticket_invalid | 401 | Ticket invalid or expired |
ticket_expired | 401 | Ticket expired |
ticket_already_used | 401 | Ticket already used |
ticket_validation_failed | 500 | Ticket validation failed |
For the full API specification, refer to Auth Ticket API.
Message Format
All messages use a unified nested structure:
{
"type": "service type",
"data": { ... }
}
Service Types
| type | Description |
|---|---|
health | Heartbeat mechanism |
voice-translation | Voice translation service |
error | Error message |
Error Message Format
When an error occurs, the server returns a message with type: "error":
{
"type": "error",
"data": {
"error_code": "auth_invalid_api_key",
"severity": "fatal",
"message": "Invalid API key",
"context": "auth",
"request_id": "req_abc123xyz789",
"timestamp": "2026-01-15T10:30:45.123Z"
}
}
Sentence-level errors (such as a translation failure for one language of a sentence) additionally carry sid and details:
{
"type": "error",
"data": {
"error_code": "llm_content_filtered",
"severity": "warning",
"message": "Content filtered",
"context": "translation",
"sid": 5,
"request_id": "req_abc123xyz789",
"timestamp": "2026-01-15T10:30:45.123Z",
"details": {
"provider": "llm_service",
"translation_language": "ja-JP"
}
}
}
Session-level translation service errors (escalated after consecutive failures reach a threshold) do not carry sid. The frontend should display a global notice but does not need to disconnect:
{
"type": "error",
"data": {
"error_code": "translation_service_unavailable",
"severity": "error",
"message": "Translation service unavailable",
"context": "translation",
"request_id": "req_abc123xyz789",
"timestamp": "2026-01-15T10:30:45.123Z",
"details": {
"provider": "llm_service",
"last_error_code": "llm_provider_error",
"fail_count": 5
}
}
}
For the full trigger rules (consecutive failure threshold, error code classification), refer to the
translation_service_unavailablesection in Error Code Reference.
Plan Limit Errors (unlimited plans, v1.9.0)
API Keys on an "unlimited plan" may receive the following error messages (not applicable to credit-based keys). Use GET /api/v1/me/plan to look up the plan contents, current usage, and when a restriction lifts:
| Error Code | severity | Description | Client Handling |
|---|---|---|---|
plan_feature_not_allowed | fatal | The plan does not include the feature in use | Use features included in your plan or upgrade the plan; query GET /api/v1/me/plan for the plan contents |
concurrency_limit_reached | error | This API Key has reached its concurrent recording limit | The connection is not closed; start again after another recording ends |
daily_limit_disconnect | error | The plan's usage threshold was reached; the current recording was stopped | You may start a new recording immediately (when restarting on the same connection, session_started arrives after the previous recording finishes processing) |
daily_limit_reached | fatal | Usage has reached the plan's limit | Available again after the plan's reset (daily limits reset the next day) |
plan_feature_not_allowed has two occurrence points:
startrejected: when the plan does not include a feature enabled in the request, the start is rejected outright; the connection is not closed — adjust the parameters andstartagain.- Detected while recording: for example, a feature not in the plan is turned on mid-session; once the check before the next minute detects it, the current recording is stopped.
For real-time recording, usage limits are checked before each minute begins; once a limit is reached, the next minute does not start and is not counted toward usage. When the same API Key runs several recordings at once, after the periodic-stop threshold is reached, each recording stops before its own next minute begins.
In addition, details.max of too_many_languages may come from the plan's cap on simultaneously recognized transcription languages, in addition to the system-wide limit (10 transcription languages); details carries max and received. See Error Code Reference — Plan and Usage Limit Errors.
Single-Message Error (Message Failed, Connection Kept)
When the server encounters an unexpected internal error while handling a single WebSocket message (such as set_name, switch_language, tts_play, etc.), it returns internal_error. This error indicates only that the specific message failed to process; the connection is not terminated. The frontend should keep the connection open and may retry the operation:
{
"type": "error",
"data": {
"error_code": "internal_error",
"severity": "error",
"message": "Internal server error",
"context": "general",
"request_id": "req_abc123xyz789",
"timestamp": "2026-05-08T10:30:45.123Z",
"details": {
"message_type": "voice-translation",
"action": "set_name"
}
}
}
details Fields
| Field | Type | Description |
|---|---|---|
message_type | string | Service type: voice-translation / health |
action | string | (Optional) The specific operation that failed, such as set_name, switch_language, tts_play, tts_mode, retranslate, config, speaker.rename, etc. This field is absent when the message payload has no action field (such as a plain init message). |
What the Frontend Should Do
- Keep the WebSocket connection open: Do not call
ws.close(), navigate away, or return to a history page because of this error. The recording is still in progress. - Decide on follow-up handling based on
details.action:Scenario Recommended Action Idempotent operations such as set_name/switch_language/tts_mode/configSimply resend the same message. These operations use a "last write wins" approach, so retrying has no side effects. tts_play/tts_stop/retranslateUsually safe to retry directly. If the user is waiting for TTS playback, consider showing a transient toast indicating the retry is in progress. speaker.rename/speaker.mergeBefore retrying, use the REST API (speakers) to confirm the current DB state and avoid duplicate operations (for example, the rename already succeeded and only the response frame failed). details.actionis absentThe error occurred after the message payload was parsed, so the system cannot infer the specific operation. The frontend can infer it from "the most recent message the user sent," or display a generic error message such as "Operation failed, please retry." - User experience: Show a transient toast / inline error. Do not interrupt the user flow with a modal or a redirect.
- Telemetry / reporting: Report
request_id+detailsto your frontend error tracking (Sentry, Datadog, etc.) to make it easier to correlate with backend logs during troubleshooting.
What Will Not Happen (Guarantees)
- The recording will not be interrupted:
segment_uploaded,result,origin, and other messages keep arriving. - The connection will not be actively closed by the server.
- The session state will not be reset (
session_idstays the same). - State already written to the DB will not be rolled back (for example, if
set_namewas written to the DB successfully and only the response frame failed, the name still takes effect).
Client Handling Example
ws.onmessage = (event) => {
const msg = JSON.parse(event.data);
if (msg.type !== 'error') {
handleNormalMessage(msg);
return;
}
const { error_code, severity, request_id, details } = msg.data;
// Single-message failure: keep the connection, decide whether to retry based on action
if (error_code === 'internal_error') {
console.warn('[ws] message handler failed (connection kept)', {
request_id,
message_type: details?.message_type,
action: details?.action,
});
showTransientToast(`Failed to process "${details?.action ?? 'operation'}", please retry`);
// Note: do not call ws.close() and do not navigate away from the current page
return;
}
// Handle other errors with your existing logic (only fatal errors require disconnecting)
handleErrorBySeverity(severity, msg.data);
};
| Field | Type | Description |
|---|---|---|
error_code | string | Error code (for programmatic handling) |
severity | string | Severity: fatal / error / warning |
message | string | Human-readable error message |
context | string | Error source category |
sid | int | Optional. The sentence number for sentence-level errors (such as a translation failure); absent for non-sentence-level errors |
request_id | string | Request tracking ID |
timestamp | string | Time the error occurred (ISO 8601) |
details | object | Optional. Error context; common keys: provider, translation_language, source_lang, etc. |
For the full list of error codes, refer to Error Code Reference.
Health (Heartbeat Service)
Description
Used to confirm that the WebSocket connection is healthy. We recommend sending a ping every 30 seconds; if no pong is received, treat the connection as dropped and reconnect. While the previous recording is still being processed after it ends (for example, while its summary is being generated), pong can be delayed by a few seconds to a few tens of seconds; allow enough waiting time and do not treat the connection as dropped just because a pong has not arrived yet.
Use Cases
- Maintaining a long-lived connection
- Detecting connection status
- Preventing connection timeouts
Request - Ping
{
"type": "health",
"data": {
"action": "ping"
}
}
Response - Pong
{
"type": "health",
"data": {
"action": "pong"
}
}
Voice Translation - start (Start Voice Translation)
Description
Starts a new voice translation session and begins processing audio according to the configured parameters.
Use Cases
- Starting a meeting record
- Starting real-time translation
- Starting a voice memo
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value start |
transcription_languages | string[] | Yes | Speech recognition languages (up to 10) |
translation_languages | string[] | No | Translation target languages; multiple allowed (up to 12; empty = no translation). As of v1.6.7, transcribe and broadcast translate all specified languages in real time, with one result event per language (same sid, single language key inside translations) — clients must accumulate by language code instead of overwriting. Two-way translation (conversation) does not apply: the server overwrites this field with the counterpart language, always a single language. record no longer supports translation as of v1.7.0 and returns 400 record_translation_not_allowed. See the WebSocket reference |
realtime_translation | boolean | No | Real-time translation mode (default false). true: translates word-by-word while the sentence is still being recognized (interim); false: translates only when the sentence is finalized. This flag also governs multi-language translation. Always treated as true for broadcast; has no effect for conversation, which translates finalized sentences only |
recognition_mode | string | No | Recognition mode: single (single speaker, default), multi_speaker (multiple speakers), multi_channel (multi-channel, v1.10.0: one recording session takes input from multiple physical microphones, with speaker identity determined by the channel — see "Multi-Channel Mode Description" below). Under multi_speaker, transcription_languages must contain exactly 1 language; otherwise the server returns a diarization_multilang_conflict error and refuses to start (type=conversation is exempt: two-way translation forces single-speaker mode and has been exempt from this check since v1.7.2). |
type | string | Yes | Recording type: transcribe, conversation, record, broadcast |
audio_format | string | No | Audio format: pcm (default), webm |
summary_template | string | Conditional | Summary template. Required for transcribe when summary_mode=builtin; forbidden when summary_mode=custom; optional for conversation/broadcast. |
options | object | No | Speech recognition options |
tts_enabled | boolean | No | Whether to enable TTS speech synthesis (default false) |
tts_language | string | No | TTS output language (must be in translation_languages) |
tts_voice | string | No | TTS voice name (such as en-US-JennyNeural) |
tts_mode | string | No | TTS playback mode: sync (synchronous, default), async (asynchronous). Only these two lowercase values are accepted, and an empty string is treated as not provided (sync); any other value (for example "Async") is rejected (invalid_parameter, details.field is tts_mode) and the recording does not start. Checked for every recording type, whether or not TTS is enabled |
broadcast_token | string | Conditional | Broadcast token (required for the broadcast type, obtained from the REST API). Allowed only with the broadcast type: any other type that carries it is rejected (invalid_parameter, details.field is broadcast_token) and the recording does not start |
active_language | string | No | Initial active language for two-way mode (default transcription_languages[0]) |
speakers | array | No | User-to-language mapping for two-way mode (exactly 2 users when provided). When omitted, user 1 uses transcription_languages[0] and user 2 uses transcription_languages[1] |
conversation_mode | string | No | Two-way conversation mode: auto (auto-detect, default), manual (push-to-talk). Only these two lowercase values are accepted, and an empty string is treated as not provided (auto); any other value is rejected (invalid_parameter, details.field is conversation_mode) and the recording does not start. Also checked for recording types other than two-way |
speaker_diarization | boolean | No | Speaker diarization (forcibly ignored in two-way mode) |
tts_config | object | No | Multi-language TTS settings (applies to both broadcast mode and two-way mode) |
broadcast_phase | string | No | Initial broadcast phase: standby, live (default). Only these two lowercase values are accepted, and an empty string is treated as not provided (live); any other value (for example "Live") is rejected (invalid_parameter, details.field is broadcast_phase) and the recording does not start |
standby_message | string | No | The message viewers see during the standby phase (default: "Getting ready, please wait...") |
name | string | No | Initial default recording name (max 60 chars after trimming leading and trailing whitespace; the system may still override it; if not provided, auto-generated such as Transcription #1). A longer name is rejected (invalid_parameter, details.field is name) and the recording does not start |
summary_language | string | No | Summary output language (defaults to the recognition language when unspecified; in broadcast mode, read automatically from the channel settings). Up to 20 characters; a longer value is rejected (invalid_parameter, details.field is summary_language) and the recording does not start |
summary_mode | string | No | Summary mode enum: builtin (default) / custom. Inferred as builtin when omitted. |
summary_prompt | string | No | Required in custom mode (a value with only whitespace counts as not provided); supplemental instructions in builtin mode. <= 3000 characters. |
summary_prompt_slug | string | No | Required in custom mode (a value with only whitespace counts as not provided); forbidden in builtin mode. Your own identifier (<= 64 characters, Unicode, no control characters; passed through and stored in the backend record for historical lookup). |
summary_plain_text | boolean | No | Request plain-text summary output (default false; when enabled, the backend performs Markdown post-processing). |
channel_mode | string | Conditional | Multi-channel sub-mode (required when recognition_mode=multi_channel): per_channel (each channel is recognized independently) or shared (channels take turns speaking and share one recognition stream). Other values return invalid_channel_mode |
channels | array | Conditional | Multi-channel channel list (required when recognition_mode=multi_channel; includes the main speaker, conventionally channel_id: 1). See "Multi-Channel Mode Description" below |
silenceTimeoutSeconds | integer | No | How many consecutive seconds without detected speech end the recording automatically: omitted or null uses the default (900 seconds); 0 means this session never ends for lack of speech; otherwise it must be an integer from 60 to 86400, and any other value is rejected. See Automatic End After a Long Silence below |
options Sub-fields
options is the speech recognition options object; all fields are optional and use their respective defaults when omitted.
| Field | Type | Default | Description |
|---|---|---|---|
speaking_speed | string | normal | Speaking speed, which affects the silence threshold for sentence segmentation: very_slow / slow / normal / fast / very_fast. The default normal corresponds to an 800ms silence threshold (see speaking_speed levels for each level). Use a slower setting for slower speakers (longer threshold, avoids cutting on mid-sentence pauses); a faster setting segments sooner. Can be adjusted dynamically during recording via set_speaking_speed. Only these five lowercase values are accepted, and an empty string is treated as not provided (normal); any other value is rejected (invalid_parameter, details.field is options.speaking_speed) and the recording does not start |
profanity_handling | string | mask | Profanity handling: mask (mask with ***) / remove (remove) / show (show original). Only these three lowercase values are accepted, and an empty string is treated as not provided (mask); any other value is rejected (invalid_parameter, details.field is options.profanity_handling) and the recording does not start |
Note: These options apply to STT sentence segmentation; multi-speaker mode (
multi_speaker) does not currently applyspeaking_speed.
Recording Type Descriptions
| type | Description | Use Cases |
|---|---|---|
transcribe | Speech-to-text | Meeting minutes, interview notes |
conversation | Conversation record | Two-way communication, customer service conversations |
record | Plain recording | Voice memos, quick notes |
broadcast | Broadcast/live | Lectures, talks, live content |
Request Example (Basic)
{
"type": "voice-translation",
"data": {
"action": "start",
"transcription_languages": ["zh-TW"],
"translation_languages": ["en-US"],
"realtime_translation": false,
"type": "transcribe",
"audio_format": "pcm",
"summary_template": "meeting",
"options": {
"speaking_speed": "normal",
"profanity_handling": "mask"
}
}
}
Request Example (Initial Default Name)
{
"type": "voice-translation",
"data": {
"action": "start",
"transcription_languages": ["zh-TW"],
"translation_languages": ["en-US"],
"type": "transcribe",
"audio_format": "pcm",
"summary_template": "meeting",
"name": "Product Planning Meeting"
}
}
Recording Name Rules
| Scenario | Name | name_source | System Override? |
|---|---|---|---|
start with a name parameter | Initial default name | default | Yes |
start without a name | Auto-generated (such as Transcription #1, Broadcast #3) | default | Yes |
Set via set_name | The name explicitly set by the user | user | No |
| Auto-generated by the system after the session ends | A summary name generated from the transcript content | llm | — |
Note: The
nameinstartis the initial default name; the system may still override it when the session ends. If you need a fixed name, useset_name.
Default name format (fixed English):
| Recording Type | Default Name Format |
|---|---|
transcribe | Transcription #N |
conversation | Conversation #N |
record | Recording #N |
broadcast | Broadcast #N |
Nis the sequential number for that user's recordings of the same type. Name priority:user>llm>default. Once the user sets a name, the system will not override it when the session ends.
Request Example (With TTS)
{
"type": "voice-translation",
"data": {
"action": "start",
"transcription_languages": ["zh-TW"],
"translation_languages": ["en-US"],
"realtime_translation": true,
"type": "transcribe",
"tts_enabled": true,
"tts_language": "en-US",
"tts_voice": "en-US-JennyNeural",
"tts_mode": "sync"
}
}
Request Example (Two-Way Mode - Auto-Detect)
{
"type": "voice-translation",
"data": {
"action": "start",
"type": "conversation",
"transcription_languages": ["zh-TW", "en-US"],
"active_language": "zh-TW",
"audio_format": "pcm",
"speakers": [
{ "id": 1, "language": "zh-TW" },
{ "id": 2, "language": "en-US" }
],
"tts_config": {
"zh-TW": { "voice": "zh-TW-HsiaoChenNeural", "speaking_rate": 1.0 },
"en-US": { "voice": "en-US-JennyNeural", "speaking_rate": 1.0 }
}
}
}
Request Example (Two-Way Mode - Manual Mode)
{
"type": "voice-translation",
"data": {
"action": "start",
"type": "conversation",
"transcription_languages": ["zh-TW", "en-US"],
"conversation_mode": "manual",
"audio_format": "pcm",
"speakers": [
{ "id": 1, "language": "zh-TW" },
{ "id": 2, "language": "en-US" }
],
"tts_config": {
"zh-TW": { "voice": "zh-TW-HsiaoChenNeural", "speaking_rate": 1.0 },
"en-US": { "voice": "en-US-JennyNeural", "speaking_rate": 1.0 }
}
}
}
Request Example (Custom Summary Prompt - custom mode)
In
mode=custom, yoursummary_promptcontent completely replaces the built-in template rules, and the backend already adds prompt injection protection. Thesummary_prompt_slugis metadata for your own identification (stored in the backend record) and does not enter the prompt content.If you want to keep the built-in template and add your own supplemental instructions afterward, use
summary_mode=builtin+summary_template=<slug>+summary_prompt=<supplemental instructions>instead (in builtin mode,summary_promptis treated as supplemental and appended after the built-in template).
{
"type": "voice-translation",
"data": {
"action": "start",
"transcription_languages": ["zh-TW"],
"translation_languages": ["en-US"],
"type": "transcribe",
"audio_format": "pcm",
"summary_language": "zh-TW",
"summary_mode": "custom",
"summary_prompt": "You are a meeting-minutes assistant. List every amount and committed date discussed in bullet points, and note the responsible person for each.",
"summary_prompt_slug": "client_x_finance_v3",
"summary_plain_text": false
}
}
Important — How to Retrieve the Summary Result: In WebSocket mode, summaries are non-streaming by design; final_content is not pushed back via a WebSocket event (the summary_done event only signals completion and does not contain the content). The client must retrieve it afterward over HTTP:
- After receiving the
summary_doneevent, callGET /api/v1/sse/history/transcribe/{taskId}to retrieve the summary (theinit_summaryevent carries a top-levelsummaryplain string +summary_mode/summary_template/summary_plain_text/summary_prompt_snapshot+ the two content-filter fallback audit fieldssummary_fallback_level/summary_dropped_segmentsadded in v1.5.5). - Or query the
summary_mode/summary_template/summary_prompt_slugfields of the transcript record via the REST API.
v1.5.5 Content-Filter Automatic Downgrade: If your prompt or transcript content triggers the LLM service's content filter, the system automatically downgrades (standard mode → neutral mode → segment-omission mode). The
summary_fallback_levelfield of thesummary_doneevent (value2or3; omitted when standard mode succeeds directly) tells the client which path was actually taken, so the frontend can display hints such as "neutral mode in use" / "N segments omitted." See reference/websocket/events.md – summary_done and the V1.5.5 changelog.
Two-Way Mode Special Rules:
| Item | Description |
|---|---|
transcription_languages | Must contain exactly 2 languages, and they cannot be the same. |
translation_languages | Not required (automatically derived as the non-active language). |
realtime_translation | Has no effect in conversation mode: translations are sent only after a sentence is finalized, regardless of whether this field is true or false. |
active_language | Optional, defaults to transcription_languages[0]. |
recognition_mode | Forced to single (ignores speaker_diarization). |
tts_enabled | Defaults to true; set to false to return text translations only. |
tts_config | Optional; sets the TTS voice for each of the two languages; leave empty to use the default voices automatically. |
summary_template | Optional; when provided, a summary is automatically generated after stopping. |
speakers | Optional; specifies each user's language (exactly 2 users when provided). When omitted, users 1 and 2 map to the two transcription_languages in order. |
conversation_mode | Optional; auto (auto-detect, default) or manual (push-to-talk). |
speakers Field Descriptions:
| Field | Type | Required | Description |
|---|---|---|---|
id | int | Yes | User number (1 or 2) |
language | string | Yes | The user's language code (must be in transcription_languages) |
conversation_mode Descriptions:
| Mode | Description |
|---|---|
auto (default) | The system automatically detects the spoken language and segments sentences automatically. |
manual | The user controls speaking periods via start_speaking / stop_speaking, during which the audio is merged into a single sentence. |
Multi-Channel Mode Description (recognition_mode: "multi_channel")
A single recording session takes input from multiple physical microphones at the same time. Speaker identity is determined by the channel (not inferred by AI): each transcript sentence is attributed to the speaker of the channel_id it came from. There are two sub-modes:
per_channel: each microphone runs its own independent speech recognition; channels can speak at the same time and each can be bound to a different language.shared: channels take turns speaking and share one recognition stream, and the whole session uses the same set of languages. Suited to situations where only one person speaks at a time; see Shared Mode below.
Scope and limits:
| Item | Description |
|---|---|
| Recording types | Only transcribe and record; conversation returns invalid_parameter and broadcast returns multichannel_broadcast_not_allowed |
| Feature activation | Multi-channel must be enabled before it can be used; in an environment where it is not enabled, start returns invalid_recognition_mode |
channel_mode | Required; omitting it returns channel_mode_required. Allowed values are per_channel and shared; other values return invalid_channel_mode |
channels | Required (omitting it returns channels_required), 1–8 channels (including the main speaker, conventionally channel_id: 1); channel_id range 1–8, must be unique |
| Per-channel language | Under per_channel, each channel must specify exactly 1 transcription language (each channel is bound to one language); the union of all channel languages (deduplicated) must exactly match the session-level transcription_languages, otherwise channel_language_mismatch is returned. Under shared, channels must not carry transcription_languages (returns channel_language_not_allowed); the whole session uses the session-level transcription_languages |
audio_format | Only pcm is supported (16kHz / 16-bit / mono / little-endian); other values return multichannel_requires_pcm |
| TTS | Not supported in this first release: tts_enabled: true returns multichannel_tts_not_allowed |
speaker_diarization | Cannot be specified at the same time (multi-channel is itself a form of speaker diarization); providing both returns invalid_parameter |
| Plan limits | Unlimited plans must include the multi-channel feature; a plan may also cap the number of simultaneous recognition channels — exceeding the cap returns plan_feature_not_allowed on the spot at start (details.field="max_stt_streams"). shared is always counted as 1 recognition channel |
| Audio file length cap | The audio saved for one multi-channel recording has a total cap, reached sooner the more channels are open (about 70 minutes with 8 channels). Once the cap is reached, the audio file and the recording duration (duration_ms) stop at that point, while the transcript and credit charges continue as usual |
channels field descriptions:
| Field | Type | Required | Description |
|---|---|---|---|
channel_id | int | Yes | Channel number, range 1–8, must be unique; includes the main speaker, conventionally 1 for the main speaker |
speaker_name | string | No | Display name of that channel's speaker (max 100 characters, no control characters, otherwise invalid_parameter is returned); when not provided, the speaker ID (channel_{N}) is displayed |
transcription_languages | string[] | Conditional | Transcription language for that channel. Required under per_channel, exactly 1 (omitting it or providing more than one returns channel_language_required); must not be provided under shared (returns channel_language_not_allowed) |
Multi-channel mode request example:
{
"type": "voice-translation",
"data": {
"action": "start",
"type": "transcribe",
"recognition_mode": "multi_channel",
"channel_mode": "per_channel",
"transcription_languages": ["zh-TW", "en-US"],
"translation_languages": ["ja-JP"],
"audio_format": "pcm",
"channels": [
{ "channel_id": 1, "speaker_name": "Manager Wang", "transcription_languages": ["zh-TW"] },
{ "channel_id": 2, "speaker_name": "Alex", "transcription_languages": ["en-US"] },
{ "channel_id": 3, "speaker_name": "Lee", "transcription_languages": ["zh-TW"] }
]
}
}
Successful response (session_started, multi-channel):
The top level of data additionally carries channel_mode and channels[] (each entry with channel_id, speaker_name, transcription_languages, status; transcription_languages is not present in shared mode). The frontend should use this to confirm the server actually started in multi-channel mode:
{
"type": "voice-translation",
"data": {
"action": "session_started",
"session_id": "550e8400-e29b-41d4-a716-446655440000",
"task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
"recording_type": "transcribe",
"recognition_mode": "multi_channel",
"channel_mode": "per_channel",
"channels": [
{ "channel_id": 1, "speaker_name": "Manager Wang", "transcription_languages": ["zh-TW"], "status": "preparing" },
{ "channel_id": 2, "speaker_name": "Alex", "transcription_languages": ["en-US"], "status": "preparing" },
{ "channel_id": 3, "speaker_name": "Lee", "transcription_languages": ["zh-TW"], "status": "preparing" }
],
"resume_token": "L0VBAwIy... (43 characters)",
"resume_grace_seconds": 45,
"server_time": 1749550000000,
"message": "Speech recognition started"
}
}
Multi-channel mode behavior summary:
- Every
audioframe must includechannel_id(see the audio section); we recommend one frame every 100ms, and every channel should keep sending audio even while silent. A single silent channel does not end the recording; see Automatic End After a Long Silence - The
originofresultevents carrieschannel_id,speaker_id(formatchannel_{N}), andspeaker_label(=speaker_name; after arename_speaker, the new label);translationsdoes not carrychannel_id— usesidto match back toorigin - During recording you can dynamically add or disable channels and change a channel's language via
add_channel/remove_channel/set_channel_language;switch_languagealways returnsmultichannel_switch_language_not_allowed rename_speakerworks (a channel's speaker can be renamed before they say anything);reassign_speakerandmerge_speakersdo not apply to multi-channel (speaker identity is determined by the physical channel)- Channel status changes are reported via the
channel_statusevent
Shared Mode (Taking Turns)
channel_mode: "shared": the microphones take turns speaking and share one recognition stream. The speaker is labeled by which channel that stretch of audio came in on, which is highly accurate when people take turns; billing is always counted as 1 channel (see Pricing).
Request example:
{
"type": "voice-translation",
"data": {
"action": "start",
"type": "transcribe",
"recognition_mode": "multi_channel",
"channel_mode": "shared",
"transcription_languages": ["zh-TW", "en-US"],
"audio_format": "pcm",
"channels": [
{ "channel_id": 1, "speaker_name": "Host" },
{ "channel_id": 2, "speaker_name": "Guest A" },
{ "channel_id": 3, "speaker_name": "Guest B" }
]
}
}
session_started returns channel_mode: "shared"; the entries in channels[] do not carry transcription_languages.
Client requirements:
| Item | Description |
|---|---|
| Send only one channel at a time | At any moment, send audio only on the current speaker's channel. If several channels are sent at once, their audio is queued into the same recognition stream in the order received and the transcript becomes garbled (billing is not affected) |
| Keep sending silence | When nobody is speaking, keep sending silence on the current channel (one frame every 100ms recommended); do not stop sending |
| Audio format | pcm only (same as the common multi-channel rule) |
| Language | The whole session shares the session-level transcription_languages; channels cannot specify a language, and the language cannot be changed during the recording (set_channel_language returns channel_language_not_allowed) |
| The first channel cannot be removed | The first entry in channels[] carries the recognition for the whole session and cannot be removed with remove_channel (returns channel_remove_not_allowed); other channels can be removed |
Channel status:
- The
statusof the other channels follows the first channel: when the first channel turnsready, all channels turnreadytogether; when the first channel prepares recognition again (resume after pause, resume after disconnect, automatic reconnection), all channels return topreparingtogether. Each channel receives its ownchannel_statusevent, with the samereasonas the first channel. After resuming from a disconnect,preparingis reported in theresume_oksnapshot rather than by a separate event. - A channel added with
add_channelduring the recording takes on the first channel's current status directly. removedfollows each channel's own status; a removed channel receives no further events.- When the first channel turns
error, the other channels turnerroras well (with the samereason); when the first channel recovers, they recover together. - The
channels[]snapshots insession_startedandresume_okfollow the same rules.
Known limitations:
- When the gap between speakers is shorter than about 0.8 seconds, the words of the two people may be merged into one sentence labeled with only one speaker.
- Speakers are determined by channel labeling; merging and reassigning speakers through the speaker APIs do not apply (same as the common multi-channel rule), so this cannot be corrected afterward.
- Retroactive transcription after resuming from a pause covers the whole last 60 seconds, for all channels together, with speakers labeled (see pause).
Multi-channel start-specific errors:
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
invalid_recognition_mode | 400 | Multi-channel is not enabled for this environment | Contact the platform to enable multi-channel |
channel_mode_required | 400 | channel_mode is missing | Provide channel_mode (per_channel or shared) |
invalid_channel_mode | 400 | channel_mode is not per_channel or shared; also returned when shared is not enabled in this environment (explained in details.message) | Use per_channel or shared; if not enabled, use per_channel |
channel_language_not_allowed | 400 | A channel carries transcription_languages in shared mode (details carries channel_id) | Remove transcription_languages from each channel and use the session-level setting |
channels_required | 400 | channels is missing | Provide 1–8 channel configurations (including the main speaker) |
too_many_channels | 400 | The channel count exceeds the limit (details carries max and received) | Reduce the number of channels |
invalid_channel_id | 400 | channel_id is out of range (1–8) or duplicated | Use unique numbers within 1–8 |
channel_language_required | 400 | A channel does not specify exactly one transcription language (details carries channel_id) | Provide exactly 1 language in each channel's transcription_languages |
channel_language_mismatch | 400 | The union of channel languages does not match transcription_languages (details lists both sets) | Reconcile the two language lists |
multichannel_requires_pcm | 400 | audio_format is not pcm | Use pcm (16kHz / 16-bit / mono) |
multichannel_tts_not_allowed | 400 | Multi-channel does not support speech synthesis in this first release | Disable tts_enabled |
multichannel_broadcast_not_allowed | 400 | Broadcast does not support multi-channel mode | Use another recognition mode for broadcast |
invalid_parameter | 400 | Combined with speaker_diarization, type=conversation, or a malformed speaker_name (details.field indicates the field) | Fix the parameters according to details |
Broadcast Mode Description (type: "broadcast")
In broadcast mode, the language settings are automatically obtained from the broadcast channel settings and do not need to be sent in the WebSocket message.
Required parameters:
| Parameter | Type | Description |
|---|---|---|
type | string | Must be "broadcast" |
broadcast_token | string | Broadcast token (obtained after creating the broadcast via the REST API) |
audio_format | string | Audio format (pcm or webm) |
Optional parameters (override the broadcast channel settings):
| Parameter | Type | Description |
|---|---|---|
tts_config | object | Multi-language TTS settings (overrides the settings from creation time) |
summary_template | string | Summary template slug (overrides the settings from creation time; if not provided, the broadcast channel default is used) |
Auto-configured parameters (can be omitted):
transcription_languages: read automatically from the broadcast settingstranslation_languages: read automatically from the broadcast settingsrealtime_translation: always enabled in broadcast mode, andfalseis treated astrue; translation is always billed at the real-time translation ratesummary_template: read automatically from the broadcast settings (the value passed via WebSocket takes precedence)summary_language: read automatically from the broadcast settings (the value passed via WebSocket takes precedence)
Both the host and viewers receive interim translations: a sentence's translation first arrives with
is_final: false, followed by the finalized version withis_final: true. Overwrite the display bysidpluslanguage; to show only finalized translations, skipis_final: false.The recording name does not reuse the channel name. If
startomitsname, a name such asBroadcast #1is generated; naming otherwise works the same as for other types, see Recording Name Rules.
Broadcast Phase Descriptions:
| broadcast_phase | Description | Behavior |
|---|---|---|
live (default) | Live phase | STT/translation results are broadcast to viewers and written to the transcript. |
standby | Standby phase | STT/translation results go only to the host; viewers see the standby_message. |
Standby phase purpose: Lets the host warm up STT/translation before going live, confirm that equipment is working, and then switch to the live phase.
The standby phase has a time limit (30 minutes by default): when the accumulated standby time reaches the limit, the session ends automatically, with a warning about 2 minutes before. See Standby Time Limit.
broadcast_phaseaccepts only lowercasestandbyandlive: an empty string is treated aslive, and any other value is rejected (invalid_parameter).
Broadcast Mode Request Example:
{
"type": "voice-translation",
"data": {
"action": "start",
"type": "broadcast",
"broadcast_token": "a3f9",
"audio_format": "pcm"
}
}
Broadcast Mode Request Example (Standby Phase + Override Summary Template):
{
"type": "voice-translation",
"data": {
"action": "start",
"type": "broadcast",
"broadcast_token": "a3f9",
"audio_format": "pcm",
"broadcast_phase": "standby",
"standby_message": "The talk is about to begin, please wait...",
"summary_template": "lecture"
}
}
Summary template priority: The value passed in the WebSocket
start> the default set when the broadcast channel was created. If neither is set, no summary is automatically generated.
Broadcast Mode TTS Settings (tts_config):
Use the tts_config parameter to specify which translation languages should produce TTS audio for viewers.
| tts_config Field | Type | Description |
|---|---|---|
| voice | string | TTS voice name |
| speaking_rate | number | Speaking rate (0.5–2.0, default 1.0). Values outside the range are adjusted to the nearest bound |
{
"type": "voice-translation",
"data": {
"action": "start",
"type": "broadcast",
"broadcast_token": "a3f9",
"audio_format": "pcm",
"tts_config": {
"en-US": {
"voice": "en-US-JennyNeural",
"speaking_rate": 1.0
},
"ja-JP": {
"voice": "ja-JP-NanamiNeural",
"speaking_rate": 1.0
}
}
}
}
Note:
- TTS languages must be valid languages in
translation_languages; invalid languages are automatically ignored.- The host (WebSocket) does not receive TTS audio; only SSE viewers receive the
tts_readyevent.- TTS is sent only during the
livephase; nothing is sent during thestandbyphase.
TTS Playback Mode Descriptions
| Mode | Description | Behavior |
|---|---|---|
sync | Synchronous mode (default) | Automatically plays the latest is_final=true translated sentence; if the previous sentence is still playing, it enters the queue and waits. |
async | Asynchronous mode (manual control) | The user can choose any translated sentence for TTS, controlled with the tts_play command. |
Success Response
After a successful start, a session_started event is returned containing complete session initialization info. For real-time recording, the first minute is deducted first, and the event is returned once that deduction completes (see Pricing).
General recordings (transcribe / conversation / record):
{
"type": "voice-translation",
"data": {
"action": "session_started",
"session_id": "550e8400-e29b-41d4-a716-446655440000",
"task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
"recording_type": "transcribe",
"recognition_mode": "single",
"resume_token": "L0VBAwIy... (43 characters)",
"resume_grace_seconds": 45,
"server_time": 1749550000000,
"message": "Speech recognition started"
}
}
Broadcast mode (broadcast):
{
"type": "voice-translation",
"data": {
"action": "session_started",
"session_id": "550e8400-e29b-41d4-a716-446655440000",
"task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
"recording_type": "broadcast",
"recognition_mode": "multi_speaker",
"phase": "standby",
"viewer_count": 0,
"queue_count": 0,
"peak_viewers": 0,
"total_viewers": 0,
"resume_token": "L0VBAwIy... (43 characters)",
"resume_grace_seconds": 45,
"server_time": 1749550000000,
"message": "Speech recognition started"
}
}
| Field | Type | Description |
|---|---|---|
session_id | string | Session ID |
task_id | string | Task ID (can be used for subsequent API queries) |
recording_type | string | Recording type: transcribe, conversation, record, broadcast |
recognition_mode | string | Recognition mode: single, multi_speaker |
phase | string | Broadcast phase: standby or live (broadcast mode only) |
viewer_count | int | Current number of online viewers (broadcast mode only) |
queue_count | int | Number of viewers waiting in the queue (broadcast mode only) |
peak_viewers | int | Peak number of viewers for this broadcast (broadcast mode only) |
total_viewers | int | Total cumulative number of viewers who have connected (broadcast mode only) |
message | string | Status description message |
Multi-channel mode: the top level of
dataadditionally carrieschannel_modeandchannels[]; see "Multi-Channel Mode Description" above for an example and field descriptions.
Automatic End After a Long Silence
A recording that goes a certain length of time without recognizing any text (including interim results that are not final yet) ends automatically, so a recording someone forgot to stop does not keep being billed.
- The default threshold is 900 seconds (15 minutes) and can be changed per session with
silenceTimeoutSeconds. About 2 minutes before the end, a warning is sent once. - Any recognized text restarts the count from 0;
resumeandstart_speakingalso restart it. - The count does not run in these cases:
- While paused: the count restarts from 0 after resuming. A paused recording is still billed; to stop billing, end the recording.
- While a disconnected session waits to be resumed: after resuming, the count continues from where it was.
- Broadcasts (
type: "broadcast"): never end for lack of speech. The standby phase has its own time limit; see the Broadcast Guide.
- Multi-channel: counts only when no channel recognizes any text; a single silent channel does not end the recording.
- Conversation manual mode: speech recognized while the speak button is not pressed also restarts the count.
Values of silenceTimeoutSeconds (top level of the start data, optional)
| Value | Effect |
|---|---|
Omitted, or null | Uses the default threshold |
0 | This session never ends automatically for lack of speech; suited to sessions kept open for long periods during which nobody may speak |
| Integer from 60 to 86400 | The threshold in seconds for this session |
| Any other value | start is rejected with the error code invalid_parameter (details.field is silenceTimeoutSeconds) and the recording does not start |
- The value must be a JSON integer: strings (for example
"900"), decimal notations (for example900.0or1e3), and booleans are all rejected. - The parameter name is
silenceTimeoutSeconds; sendingsilence_timeout_secondsis rejected (details.fieldissilence_timeout_seconds), not ignored. - A valid value has no effect on broadcasts, but an invalid value is still rejected.
- With a threshold of 120 seconds or less, no warning is sent; the recording simply ends when the time is up.
- Resuming a disconnected session keeps the original setting; send it again when starting a new recording.
Events you receive
stt_silence_warning(anerrorevent withseverity: "warning"): a warning; the recording continues.details.silenceSecondsis how long the silence has lasted anddetails.remainingSecondsis how many seconds remain. Do not treat it as the end of the recording; recognized text or resuming the recording restarts the count.stt_silence_timeout(severity: "fatal"): the recording has ended automatically.details.silence_secondsis the threshold in seconds.- Then
status: "ended"andtask_complete, in that order, just as when the client sendsstop: the recording is saved and summarized as usual, and billing runs until the end.
{
"type": "error",
"data": {
"error_code": "stt_silence_warning",
"severity": "warning",
"message": "No speech detected for a while; the recording will end automatically soon",
"context": "stt",
"request_id": "req_abc123xyz789",
"timestamp": "2026-09-25T10:28:00.000Z",
"details": {
"silenceSeconds": 780,
"remainingSeconds": 120
}
}
}
After
stt_silence_timeout, stop sending audio. Each audio message that arrives afterwards gets asession_not_startedreply, which you can ignore.
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
missing_transcription_languages | 400 | No language parameter provided | Make sure the request includes transcription_languages |
invalid_transcription_language | 400 | Invalid language code | Confirm the language code format is correct (such as zh-TW) |
too_many_languages | 400 | Number of languages exceeds the limit (details.max may be the system limit or the plan's cap on simultaneously recognized transcription languages) | Up to 10 transcription languages and 12 translation languages; on an unlimited plan, adjust per details.max |
invalid_recording_type | 400 | Invalid recording type | Use a valid type value |
audio_format_unsupported | 400 | audio_format is not supported; details.supported_formats lists the accepted values | Switch to a supported audio format |
invalid_summary_template | 400 | Invalid summary template | Confirm the template identifier is correct |
stt_init_failed | 503 | Service initialization failed | Retry later |
auth_insufficient_credit | 402 | Insufficient credit | Top up your credit balance |
auth_quota_exceeded | 402 | Available credits are insufficient; the recording did not start (less than one minute for real-time recording; the connection is not closed; details.remaining_budget is the available credit as of the most recent settlement, and details.budget_scope states whose credit it is) | Top up and start again |
daily_limit_reached | — | Usage has reached the plan's limit; the recording did not start | Available again after the plan's reset (daily limits reset the next day) |
auth_service_error | 500 | Service temporarily unavailable; the recording did not start (the connection is not closed) | start again later |
plan_feature_not_allowed | 403 | The unlimited plan does not include a feature enabled in the request (the connection is not closed) | Use features included in your plan or upgrade the plan; query GET /api/v1/me/plan for the plan contents |
concurrency_limit_reached | — | This API Key has reached its concurrent recording limit (the connection is not closed) | start again after another recording ends |
service_shutdown | — | The service is shutting down; the recording did not start, and the connection is then closed | Reconnect later and start again |
tts_init_failed | 503 | TTS service initialization failed | Retry later |
tts_invalid_language | 400 | TTS language not in the translation languages | Confirm tts_language is in translation_languages |
broadcast_token_required | 400 | Broadcast mode requires a token | The broadcast type must provide a broadcast_token |
broadcast_token_invalid | 401 | Invalid broadcast token | Confirm the token is correct and not expired |
broadcast_not_ready | 503 | Broadcast service not yet started | Retry later |
summary_invalid_mode | 400 | summary_mode is not builtin / custom | Use a valid mode |
summary_mode_field_mismatch | 400 | The mode and field combination does not match (a required field is missing / a forbidden field was included) | Adjust fields per the mode rules |
summary_prompt_too_long | 400 | summary_prompt exceeds 3000 characters | Shorten the custom prompt |
summary_prompt_slug_too_long | 400 | summary_prompt_slug exceeds 64 characters | Shorten the identifier |
summary_prompt_slug_invalid | 400 | summary_prompt_slug contains control characters (\n / \r / \t / \0, etc.) | Remove the control characters |
invalid_parameter | 400 | Invalid parameter, for example an invalid value or name for silenceTimeoutSeconds, an invalid broadcast_phase value, broadcast_token sent with a type other than broadcast, name longer than 60 characters, summary_language longer than 20 characters, or a value outside the accepted list for options.speaking_speed, options.profanity_handling, conversation_mode, or tts_mode (details.valid_values lists the accepted values) | Fix the parameter indicated by details.field |
Voice Translation - config (Set Terminology / Correction Rules)
Description
Before or during recording, pass in terminology, fuzzy-word correction rules, and translation dictionary settings. These settings improve STT accuracy, fix homophone errors, and ensure translation consistency.
Terminology also drives homophone correction: When terminology is provided, the terms become the reference for homophone matching — any span in the transcript that sounds the same but is written differently is corrected back to the spelling of the term. Providing terminology alone is therefore enough to get correction; you do not need to list possible misspellings by hand.
Limits
Per-block entry limits, length limits, and the matching error codes are collected in Terminology Guide → Limits.
The numbers are defaults; the limit actually in force can be tuned per environment — always treat
maxin the error response'sdetailsas authoritative rather than hard-coding the numbers.
Use Cases
- Pass in professional terminology (Phrase List) before recording starts
- Set fuzzy-word correction rules (homophone correction) - optional; terminology itself already provides homophone correction
- Set a translation dictionary (ensure consistent terminology translation)
Timing
| Setting Type | Recommended Timing | Update During Recording |
|---|---|---|
| Terminology | Before or during start | Supported (from the next utterance) |
| Fuzzy-word correction | Before or during start | Supported (from the next utterance) |
| Translation dictionary | Before or during start | Supported (from the next utterance) |
Note: When you update terminology, fuzzy-word correction, or the translation dictionary during recording, the new settings take effect from the next utterance after they are sent; the utterance currently being recognized is not guaranteed to be covered. No reconnection is needed. When terminology is updated, the response includes a
terminology_effective: "next_turn"field as a hint.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value config |
terminology | object | No | Terminology settings |
fuzzy_correction | object | No | Fuzzy-word correction rules |
translation_dict | object | No | Translation dictionary |
Note: At least one setting item must be provided.
Note: All or nothing: all three blocks are validated before any of them is applied. If any block fails validation, an error is returned and none of the three blocks is applied — the settings stay as they were.
For example, sending a valid terminology list together with more than 3000 translation-dictionary entries returns
config_too_many_dict_entries, and the terminology does not take effect either. Fix the problem and resend the completeconfig; all three blocks replace their previous value wholesale, so resending does not stack on top of earlier settings.
Terminology Format (terminology)
Keyed by language code, with an array of terms as the value:
{
"zh-TW": [
{ "term": "語者分離" },
{ "term": "WebSocket" }
],
"en-US": [
{ "term": "diarization" }
]
}
| Field | Type | Required | Description |
|---|---|---|---|
term | string | Yes | The term (max 100 characters) |
Limit: Up to 500 terms across all languages combined in one
config— not 500 per language. Exceeding the limit returnsconfig_too_many_entries(withcountandmaxindetails).Language scope: Terms apply only to recognition in the language they are registered under. Single-language situations — each multi-channel track, speaker diarization, and file import — use only that language's terms; multi-language transcription and conversation mode use the terms for the languages declared for the session. Terms registered under a language not used in this session have no effect, but still count toward the 500 combined total above.
These numbers are defaults: the limit actually in force can be tuned per environment; always treat
maxin the error response'sdetailsas authoritative.
Fuzzy-Word Correction Format (fuzzy_correction)
Note: This field usually does not need to be set manually —
terminologyalready corrects misspellings that sound the same or nearly the same. Use it in these three cases:
- The misspelling is itself an ordinary word and is therefore blocked by common-word protection — 晶圓 heard as 金元, for example
- The misspelling sounds very different from the correct term, such as a foreign brand name recognized as a phonetically unrelated word
- Misspellings in Japanese, Korean or English, which do not participate in homophone matching
When
correctis Chinese, it also becomes a reference for homophone matching — homophone misspellings not listed inincorrectare corrected tocorrectas well.case_insensitiveapplies only to the literal matching ofincorrect; it does not affect homophone matching.
Keyed by language code, with an array of correction rules as the value:
{
"zh-TW": [
{ "correct": "語者分離", "incorrect": ["語這分離", "語者分力"] },
{ "correct": "IPEVO", "incorrect": ["ltfo"], "case_insensitive": true }
]
}
| Field | Type | Required | Description |
|---|---|---|---|
correct | string | Yes | The correct word |
incorrect | string[] | Conditional | List of incorrect variants, each up to 200 characters. Can be omitted for Chinese terms (see below); required otherwise — an empty array returns config_invalid_entry (reason: "empty") |
case_insensitive | boolean | No | Whether this rule's variants match regardless of case (defaults to false = exact-case matching) |
Supplying only the correct term: when
correctis Chinese (contains Han characters),incorrectmay be omitted entirely — the system matches by pronunciation, and spellings in the transcript that sound the same or nearly the same are corrected back tocorrect.{ "fuzzy_correction": { "zh-TW": [{ "correct": "艾思通" }] } }No misspellings need to be listed above: 愛思通, 愛時通, 愛司東 and 愛似通 are all corrected. Only spellings that sound quite different (愛自動, say) or have a different number of syllables (愛松) still need to be listed in
incorrect.Note: Both conditions must hold: the language must be Chinese (
zh-TW,zh-CN,zh-HKand so on) andcorrectmust contain Han characters. Otherwiseincorrectremains required — omitting it in those cases would have no effect at all, and accepting it would leave you believing the setting took. List misspellings explicitly for Japanese, Korean and English.
Case sensitivity:
case_insensitiveis optional and defaults tofalse(exact-case matching). When set totrue, everyincorrectvariant in that rule matches regardless of case. The flag is per rule — the samecorrectterm can be split across several rules with different settings, for example making variants that cannot collide with ordinary words case-insensitive while keeping variants that could hit a personal name exact. It has no effect on Chinese rules (Chinese has no letter case).Note: Enabling it widens the false-positive surface: if
ivois case-insensitive, the personal nameIvois replaced too.
When the same incorrect variant appears in more than one rule: collisions are resolved on
incorrect(the variant), not oncorrect.
- When several rules point at different correct terms, which one actually takes effect is not guaranteed; do not rely on any ordering (registration order included)
- The case flag resolves to strict wins — if any rule leaves
case_insensitiveoff, that variant is matched with exact case"Strict wins" is a deliberately conservative choice: it prevents a permissive rule elsewhere in your vocabulary from silently loosening a brand-name rule you explicitly set to strict.
Splitting one
correctterm across several rules is therefore safe, as long as theirincorrectvariants do not overlap. But if your data contains the same incorrect variant mapped to different correct terms, the later one is dropped with no warning — check for duplicate variants before sending.
How Terminology Participates in Homophone Correction
Once terminology is provided, the terms become the reference for homophone matching. Any span in the transcript that sounds the same or nearly the same but is written differently is corrected back to the spelling of the term:
| Term | Appears in transcript as | Corrected to |
|---|---|---|
| 紡拓會 | 訪拓會 | 紡拓會 |
| 語者分離 | 語這分離, 與者分離 | 語者分離 |
| 晶圓 | 晶園 | 晶圓 |
You do not need to list the possible misspellings in advance — matching is based on pronunciation, so coverage is not limited by how many variants you can think of.
Mixed Chinese-English terms: For a term like CVD製程, matching applies only to the Chinese portion; the Latin portion is left unchanged.
Applicable languages: Homophone matching applies to Chinese only (Traditional and Simplified are interchangeable, since they share pronunciation). Terms in Japanese, Korean or English do not participate in homophone matching — list their misspellings explicitly with fuzzy_correction.
Matching range: Both identical and near-identical pronunciations are covered, including accent differences such as
jinvsjing(final -n vs -ng) and retroflex vs non-retroflex initials. For the term 晶圓廠, for example, 金圓廠 in the transcript is corrected.Common-word protection: If the span in the transcript is itself an ordinary word (金元 or 反案, say), it is not changed even when it shares a pronunciation with a term — this prevents normal sentences from being altered. To force such a correction, list the misspelling explicitly with
fuzzy_correction: explicitly listed misspellings are not subject to common-word protection.
Limit: Up to 4000 rules across all languages combined in one
config. Exceeding it returnsconfig_too_many_entries(withfield: "fuzzy_correction",countandmaxindetails).This number is a default: the limit actually in force can be tuned per environment, so always treat
maxin the error response'sdetailsas authoritative. Do not hard-code the numbers in your integration — if you want to check before sending, readmaxand fill it back in. When you do hit a limit,detailscarries bothcount(what you sent) andmax(the limit in force).Language scope: A correction rule applies only to sentences in the language it is registered under — a rule under
zh-TWwill not alter an English sentence. When the sentence language cannot be determined, all rules are applied as a fallback (better to over-apply than to skip the sentence entirely). Homophone matching driven by terminology likewise follows the language each term is registered under.
Translation Dictionary Format (translation_dict)
Group entries by language code, giving each language its own dictionary:
{
"en-US": [
{ "source": "語者分離", "target": "Speaker Diarization" },
{ "source": "晶圓", "target": "wafer", "case_sensitive": true }
],
"ja-JP": [
{ "source": "語者分離", "target": "話者分離" }
]
}
| Field | Type | Required | Description |
|---|---|---|---|
| (top-level key) | string | Yes | Target language code |
source | string | Yes | The source word (in the STT language), up to 200 characters |
target | string | Yes | The required translation for this language, up to 200 characters |
case_sensitive | boolean | No | Whether the entry applies only on an exact-case match (defaults to false = case-insensitive) |
Limit: Up to 3000 entries per language. Exceeding the limit returns
config_too_many_dict_entries; thedetailsin the response identify which language exceeded it.Note: The entry count directly affects translation workload and cost. A single translation only carries entries whose source word actually appears in that piece of text, up to 100 of them; beyond that, longer source words are kept first. Also, the more entries there are, the smaller the share that is reliably honored — an inherent limit that a higher cap does not change.
The previous format is still supported: the earlier array-of-entries format (
[{ "source": ..., "translations": { "language code": ... } }]) is still accepted, with identical content and behavior. Existing integrations need no changes. The same dictionary sent in either format produces identical results.On resume, the
translation_dictreturned inresume_okmatches the format you last sent — send the old format and you get the old format back; send the new one and you get the new one.Note: Clearing the whole dictionary is not supported: an empty object is treated as not sending this setting at all. Clearing a single language is supported — send that language an empty array (e.g.
{"en-US": []}).
Case sensitivity:
case_sensitiveis optional and defaults tofalse(case-insensitive). When set totrue, the entry applies only where the source text matchessourceexactly, including case. The flag is per entry.Note: The translation dictionary guides the model through prompting rather than literal substitution, so it is best-effort, not deterministic — the case flag is likewise a hint and is not guaranteed to be honored. Use
fuzzy_correctionwhen you need deterministic replacement.
Case-Flag Comparison
fuzzy_correction and translation_dict each have a case switch. Their field names are opposites, and so is the behavior their default value produces:
| Block | Field | Default | Default behavior |
|---|---|---|---|
fuzzy_correction | case_insensitive | false | Strict (case-sensitive) |
translation_dict | case_sensitive | false | Permissive (case-insensitive) |
Both default to false, yet one means strict and the other means permissive. Do not share a single variable between them, and do not mirror one value onto the other — getting it wrong produces no error at all, only matching behavior opposite to what you intended.
Request Example (Recommended: Terminology Only)
{
"type": "voice-translation",
"data": {
"action": "config",
"terminology": {
"zh-TW": [
{ "term": "語者分離" },
{ "term": "CVD製程" },
{ "term": "wafer良率" }
]
}
}
}
Request Example (Full Settings, Including Manual Correction Rules)
{
"type": "voice-translation",
"data": {
"action": "config",
"terminology": {
"zh-TW": [
{ "term": "語者分離" },
{ "term": "即時轉錄" }
]
},
"fuzzy_correction": {
"zh-TW": [
{ "correct": "語者分離", "incorrect": ["語這分離", "語者分力"] }
]
},
"translation_dict": {
"en-US": [{ "source": "語者分離", "target": "Speaker Diarization" }]
}
}
}
Success Response
{
"type": "voice-translation",
"data": {
"action": "config_updated",
"updated": ["terminology", "fuzzy_correction", "translation_dict"],
"message": "Settings updated"
}
}
| Field | Type | Description |
|---|---|---|
updated | string[] | The setting types that were updated |
message | string | Status message |
Error Responses
Important: Your client must also listen for
type: "error"messages — do not wait only forconfig_updated.When the server rejects a configuration it sends a
type: "error"message and notconfig_updated. An integration that waits only forconfig_updatedwill hang until its own timeout and appear as "the server never responded", even though the error was delivered anddata.error_codealready states the reason.
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
config_empty | 400 | No configuration provided. Note: An empty object {} does not count as "provided" — sending {"terminology": {}, "fuzzy_correction": {}, "translation_dict": {}} triggers this error | Provide at least one setting that actually has content. To clear a language's glossary, send {"lang": []} (for example {"zh-TW": []}) |
config_term_too_long | 400 | Term exceeds 100 characters | Shorten the term length |
config_too_many_entries | 400 | More than 500 terms, or more than 4000 fuzzy correction rules (both across all languages combined) | Remove terms or correction rules |
config_too_many_dict_entries | 400 | Translation dictionary exceeds 3000 entries for a single language (details.language identifies which one) | Reduce the dictionary entries for that language |
config_invalid_entry | 400 | A glossary entry has an invalid field (details carries language, index, field, and reason for locating it; depending on the case it may also carry variant_index, max_length, or count/max) | Fix the entry at the location given in details |
Voice Translation - audio (Send Audio)
Description
Sends audio data to the server for speech recognition. The audio must be Base64-encoded before sending.
Use Cases
- Continuously sending microphone audio
- Sending recorded audio segments
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value audio |
payload | string | Yes | Base64-encoded audio data |
channel_id | int | Conditional | In multi-channel mode (recognition_mode=multi_channel), required on every frame: the source channel number of this audio frame. Omitting it returns channel_id_required; an unknown or removed number returns unknown_channel_id. Outside multi-channel mode this field is always ignored (backward compatible) |
Audio Format Requirements
PCM format (default):
| Item | Specification |
|---|---|
| Format | PCM (raw audio) |
| Sample rate | 16000 Hz |
| Bit depth | 16-bit |
| Channels | Mono |
| Byte order | Little-endian |
| Transport encoding | Base64 |
WebM/Opus format:
| Item | Specification |
|---|---|
| Format | WebM container + Opus codec |
| Sample rate | Any (the server converts automatically) |
| Channels | Mono or Stereo (the server converts automatically) |
| Transport encoding | Base64 |
Request Example
{
"type": "voice-translation",
"data": {
"action": "audio",
"payload": "Base64-encoded PCM audio data"
}
}
Request Example (Multi-Channel Mode)
In multi-channel mode every frame must include channel_id:
{
"type": "voice-translation",
"data": {
"action": "audio",
"channel_id": 2,
"payload": "Base64-encoded PCM audio data"
}
}
Multi-channel mode notes: we recommend sending one frame every 100ms; every channel should keep sending audio even while silent. Multi-channel supports only the
pcmformat (16kHz / 16-bit / mono / little-endian).
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
session_not_started | 400 | Speech recognition has not started | Call the start action first |
audio_invalid_format | 400 | Invalid audio data format | Confirm the Base64 encoding is correct |
audio_decode_failed | 400 | Audio decoding failed | Confirm the audio format is correct. The recording continues; for WebM, send a new container (with its header) to recover. Time that cannot be decoded is not billed |
channel_id_required | 400 | An audio frame in multi-channel mode is missing channel_id | Include the source channel number on every frame |
unknown_channel_id | 400 | Unknown or removed channel_id | Use a valid number declared in start or added via add_channel |
Voice Translation - pause (Pause Translation)
Description
Pauses speech recognition processing. Audio received during the pause is buffered and continues to be processed after resuming. Two-way translation is an exception: audio received while paused is not kept; for a sentence in progress at the moment of pausing, the server first waits for the recognizer to finish the last part (usually about 1 second, at most about 3 seconds), sends it with is_final: true, and only then returns status: "paused". Billing continues while paused; see Pricing.
Use Cases
- The user steps away temporarily
- You need to pause recording
Request Example
{
"type": "voice-translation",
"data": {
"action": "pause"
}
}
Success Response
{
"type": "voice-translation",
"data": {
"action": "status",
"status": "paused",
"message": "Speech recognition paused"
}
}
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
session_not_started | 400 | Speech recognition has not started | Call start first |
session_already_paused | 400 | Already paused | You can ignore this error |
Multi-channel mode: while paused, every channel's audio file is still saved as usual, but no text is produced (no transcript is generated). Adding or disabling channels, changing a channel's language, or adjusting the speaking speed while paused returns
channel_action_while_paused— callresumefirst, then operate.
Voice Translation - resume (Resume Translation)
Description
Resumes paused speech recognition processing.
Use Cases
- The user returns to continue
- You need to continue recording
Request Example
{
"type": "voice-translation",
"data": {
"action": "resume"
}
}
Success Response
{
"type": "voice-translation",
"data": {
"action": "status",
"status": "live",
"message": "Speech recognition resumed"
}
}
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
session_not_started | 400 | Speech recognition has not started | Call start first |
session_not_paused | 400 | Not paused | You can ignore this error |
Multi-channel mode: after resuming, speech from the paused period is transcribed retroactively (timestamps land at the moment the words were actually spoken, not after the resume). Retroactive transcription has a limit: a trailing 60 seconds shared across the whole session (under
per_channel, split evenly across channels when there are several; undershared, the whole last 60 seconds, for all channels together, with speakers labeled); anything earlier is kept only in the audio file and does not enter the transcript. A sentence cut off at the moment of pausing may not appear in the transcript (same as single-stream behavior); when it is indeed not kept, resuming also sends segment_discarded (reason: "resumed") to identify it. On resume, each channel prepares recognition again and sends its ownchannel_statusevent (reason: "resumed").
Voice Translation - stop (Stop Translation)
Description
Stops speech recognition and ends the session. The server first waits for the recognizer to finish the last sentence (usually about 1 second, at most about 3 seconds) so that it is included in the transcript; if it does not arrive in time, the last interim result shown is used as that sentence. The system automatically uploads the audio file and transcript, and generates a summary (if configured; when the available credits cannot cover the summary fee, no summary is generated and summary_error is sent instead).
Use Cases
- The meeting ends
- Recording is complete
Request Example
{
"type": "voice-translation",
"data": {
"action": "stop"
}
}
Success Response
{
"type": "voice-translation",
"data": {
"action": "status",
"status": "ended",
"message": "Speech recognition stopped"
}
}
Task Complete Event
This event is sent after the audio file and transcript have been uploaded:
{
"type": "voice-translation",
"data": {
"action": "task_complete",
"task_id": "550e8400-e29b-41d4-a716-446655440000",
"message": "Task processing complete"
}
}
| Field | Type | Description |
|---|---|---|
task_id | string | Recording UUID, can be used for subsequent API queries |
noAudio | boolean | Present only when no audio was received during the whole recording, and always true; see task_complete |
Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
session_not_started | 400 | The session has not started, or this recording has already ended (for example, stop sent twice) | Call start first if it has not started; if this was a duplicate stop, the error can be ignored |
Sending
stopa second time returns this error rather than another success response.task_completeis sent only once, after the first successful stop, when the audio and transcript have finished uploading.
Voice Translation - retranslate (Retranslate)
Description
Retranslates a specified sentence, useful when the original text has been corrected and the translation needs to be updated.
Use Cases
- The user edits the original text and the translation needs updating
- Correcting recognition errors
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value retranslate |
sid | int | Yes | The sentence number to retranslate |
translation_languages | string[] | Yes | Array of translation language codes |
text | string | Yes | The original text to translate (the user-corrected text) |
Request Example
{
"type": "voice-translation",
"data": {
"action": "retranslate",
"sid": 1,
"translation_languages": ["en-US"],
"text": "The user-corrected original text"
}
}
Success Response
{
"type": "voice-translation",
"data": {
"action": "result",
"translations": {
"en-US": {
"sid": 1,
"text": "The new translation result",
"is_final": true,
"is_retranslation": true
}
}
}
}
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
invalid_data | 422 | No sid provided | Include sid |
record_translation_not_allowed | 400 | Recording-only sessions do not support translation | Use the transcribe type |
retranslate_session_not_active | 400 | The session is not started or has ended | Confirm the session state |
retranslate_no_target_lang | 400 | No target language provided | Provide translation_languages |
retranslate_no_text | 400 | No text to translate provided | Provide the text parameter |
retranslate_llm_not_ready | 503 | The translation service is not ready | Retry later |
retranslate_llm_failed | 500 | Translation service failed | Retry later |
If the translation service returns a specific failure code (for example
llm_content_filteredwhen the content cannot be translated), that code is returned as-is instead of being wrapped inretranslate_llm_failed.
Voice Translation - switch_language (Switch Language)
Description
Switches or adjusts translation languages while real-time translation is in progress. The behavior varies by recording type and the number of translation languages:
- General mode, single language (
translation_languageshas 1 entry): replaces the translation target language and automatically batch-retranslates all already-translated sentences. - General mode, multiple languages (2 or more entries, v1.6.7): redefined as "add or remove a single language". The
opparameter (add/remove) is required; omitting it returns aswitch_language_op_requirederror.addautomatically backfills existing sentences (same response sequence as single-language replace);removereturns atranslation_language_removedevent and keeps existing translations. See the WebSocket reference - Two-way mode (conversation): switches the STT source language (spoken language); the translation target automatically switches to the other language.
- Multi-channel mode (multi_channel, v1.10.0): not supported — always returns
multichannel_switch_language_not_allowed(includingop: add/op: remove). Languages are bound to channels; useset_channel_languageinstead.
Use Cases
- Switching the translation target language
- A change in language needs mid-meeting
- Adding / removing a translation language mid-session in multi-language sessions (v1.6.7)
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value switch_language |
translation_languages | string[] | Conditional | Array of translation language codes (required in general mode; only the first element is used as the operation target) |
op | string | Conditional | v1.6.7 multi-language operation: add or remove. Required for multi-language sessions |
transcription_languages | string[] | Conditional | The target language to switch to (two-way mode; if omitted, automatically toggles to the other language) |
Request Example (General Mode, Single-Language Replace)
{
"type": "voice-translation",
"data": {
"action": "switch_language",
"translation_languages": ["ja-JP"]
}
}
Request Example (Multi-Language Add / Remove, v1.6.7)
{
"type": "voice-translation",
"data": {
"action": "switch_language",
"op": "add",
"translation_languages": ["de-DE"]
}
}
{
"type": "voice-translation",
"data": {
"action": "switch_language",
"op": "remove",
"translation_languages": ["ko-KR"]
}
}
Request Example (Two-Way Mode)
Specify the target to switch to:
{
"type": "voice-translation",
"data": {
"action": "switch_language",
"transcription_languages": ["en-US"]
}
}
Automatic toggle (no parameters):
{
"type": "voice-translation",
"data": {
"action": "switch_language"
}
}
Two-Way Mode Special Behavior:
- Two-way mode uses automatic language detection and usually does not require manually switching the language.
switch_languageonly updates the internal preference state.- After a successful switch, a
language_switchedevent is returned (not a language_switch_start/done sequence). - Switching to the same language returns a
conversation_same_languagewarning.
Response Sequence (General Mode)
After switching the language, you receive the following events in order:
- language_switch_start: notifies that the switch has begun
{
"type": "voice-translation",
"data": {
"action": "language_switch_start",
"translation_language": "ja-JP",
"translation_languages": ["en-US", "ja-JP"],
"total_segments": 15
}
}
- batch_retranslation (multiple): returns retranslation results sentence by sentence
{
"type": "voice-translation",
"data": {
"action": "batch_retranslation",
"sid": 3,
"translations": {
"ja-JP": {
"sid": 3,
"text": "今日はプロジェクトの進捗について話し合いましょう",
"is_final": true,
"is_retranslation": true
}
}
}
}
- language_switch_done: notifies that the switch is complete
{
"type": "voice-translation",
"data": {
"action": "language_switch_done",
"translation_language": "ja-JP",
"translation_languages": ["en-US", "ja-JP"],
"success_count": 15,
"failed_count": 0
}
}
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
switch_language_no_target | 400 | No target language provided | Provide translation_languages |
switch_language_in_progress | 400 | The previous switch is not yet complete | Wait for the switch to complete |
switch_language_same_target | 400 | The target language is the same as the current one | You can ignore this error |
conversation_requires_two_languages | 400 | Two-way mode requires exactly two languages | Confirm transcription_languages has 2 |
conversation_languages_identical | 400 | The two two-way languages cannot be the same | Provide two different languages |
conversation_invalid_language | 400 | Invalid two-way language | Confirm the language is in transcription_languages |
conversation_same_language | 400 | Already the current language | You can ignore this warning |
multichannel_switch_language_not_allowed | 400 | Multi-channel mode does not support switch_language | Use set_channel_language to change a single channel's language |
Voice Translation - set_name (Set Recording Name)
Description
Sets the name while recording is in progress. After it is set, this name is used when the recording ends and will not be auto-generated.
Tip: You can also set an initial default name via the
nameparameter atstart, but that name may still be overridden by the system when the session ends. If you need a fixed name, useset_name.
Use Cases
- Customizing the recording title after recording starts
- Overriding an auto-generated name or a previously set name
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value set_name |
name | string | Yes | Recording name (max 60 chars) |
Request Example
{
"type": "voice-translation",
"data": {
"action": "set_name",
"name": "Product Planning Meeting"
}
}
Success Response
{
"type": "voice-translation",
"data": {
"action": "status",
"event": "name_set",
"name": "Product Planning Meeting",
"message": "Recording name updated"
}
}
Compatibility note: For backward compatibility, the
set_namesuccess response keepsaction: "status"(unchanged). New clients should identify a successfulset_nameviaevent: "name_set"(together with thenamefield). Relying onaction: "status"to detectset_namesuccess is deprecated and may be removed in a future version.
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
set_name_empty | 400 | Recording name is empty | Provide a non-empty name |
set_name_too_long | 400 | Recording name exceeds the length limit (>60 chars); the response details includes max_length | Shorten the name (≤60 characters) |
Voice Translation - rename_speaker (Globally Rename a Speaker)
Description
In multi-speaker diarization mode (multi_speaker), globally renames a speaker. All sentences using that speaker ID are updated in sync.
Also available in multi-channel mode (multi_channel): speaker IDs use the format
channel_{N}(such aschannel_1); speakers are registered atstart/add_channel, so a channel's speaker can be renamed before they say anything. After renaming, thespeaker_labelinresultevents and the transcript both reflect the new label.
Use Cases
- Changing a system-assigned speaker ID (such as
Guest-1) to a meaningful name (such asManager Wang) - Naming a newly recognized speaker during a meeting
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value rename_speaker |
speaker_id | string | Yes | The original speaker ID (such as Guest-1); the current display label is also accepted for consecutive renaming; max 100 characters |
new_label | string | Yes | The new display label; max 100 characters, must not contain control characters (\x00-\x1F, \x7F) or line breaks |
Request Example
{
"type": "voice-translation",
"data": {
"action": "rename_speaker",
"speaker_id": "Guest-1",
"new_label": "Manager Wang"
}
}
Success Response
{
"type": "voice-translation",
"data": {
"action": "speaker_renamed",
"speaker_id": "Guest-1",
"new_label": "Manager Wang",
"affected_sids": [1, 3, 5, 8]
}
}
| Field | Type | Description |
|---|---|---|
speaker_id | string | The resolved original speaker ID (even if the input was a display label, the event returns the original ID) |
new_label | string | The new display label |
affected_sids | int[] | The list of affected sentence numbers |
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
speaker_not_found | 422 | The specified speaker was not found | Confirm the speaker_id or display label exists |
speaker_name_empty | 422 | new_label is empty | Provide a valid label |
speaker_name_duplicate | 422 | The display label is already in use | Use a different label, or first change the conflicting speaker |
session_not_started | 400 | Speech recognition has not started | Call start first |
Voice Translation - reassign_speaker (Change the Speaker of a Single Sentence)
Description
Changes the speaker identity of a specific sentence, assigning the sentence to an existing speaker.
Not applicable in multi-channel mode (multi_channel): speaker identity is determined by the physical channel (each sentence's
speaker_idis its source channel), so reassigning a sentence to another channel is not supported.
Use Cases
- Correcting a speaker identity that the system recognized incorrectly
- Reassigning a sentence to another known speaker
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value reassign_speaker |
sid | int | Yes | The sentence number to change |
target_speaker_id | string | Yes | The target speaker's original ID (taken from init_sentence.speaker_id; reassign does not accept display labels) |
Request Example
{
"type": "voice-translation",
"data": {
"action": "reassign_speaker",
"sid": 5,
"target_speaker_id": "Guest-2"
}
}
Success Response
{
"type": "voice-translation",
"data": {
"action": "speaker_reassigned",
"sid": 5,
"old_speaker_id": "Guest-1",
"new_speaker_id": "Guest-2",
"new_speaker_label": "Lee Hsiao-hua"
}
}
| Field | Type | Description |
|---|---|---|
sid | int | The changed sentence number |
old_speaker_id | string | The original speaker ID |
new_speaker_id | string | The new original speaker ID |
new_speaker_label | string | The new speaker display label (after applying speaker_aliases; equals new_speaker_id when no alias exists) |
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
speaker_sid_not_found | 422 | The specified sentence was not found | Confirm the SID exists |
speaker_not_found | 422 | The target speaker does not exist | Use an existing speaker ID |
speaker_name_empty | 422 | The target speaker ID cannot be empty | Provide a valid speaker ID |
session_not_started | 400 | Speech recognition has not started | Call start first |
invalid_parameter | 400 | Creating a new speaker is not supported | Use an existing speaker ID |
Voice Translation - merge_speakers (Merge Speakers)
Description
Merges all sentences of one speaker into another speaker. After the merge, future recognition results for that speaker are also automatically converted to the target speaker. This applies within the current recognition pass only: after a connection recovers, speakers are re-identified and you need to merge again.
Not applicable in multi-channel mode (multi_channel): speaker identity is determined by the physical channel, so merging speakers across channels is not supported.
Use Cases
- The speech recognition engine sometimes misidentifies the same person's voice as multiple speakers (for example, Guest-1 and Guest-2 are actually the same person)
- Use this feature to merge all of Guest-2's sentences into Guest-1
- After the merge, future Guest-2 recognition results are automatically displayed as Guest-1
- This interception applies only to the current recognition pass: if the connection drops and recovers mid-recording, speakers are identified again and receive new IDs, and the merge does not carry over — merge again once you confirm it is the same person
Difference from reassign_speaker
| Feature | Scope | Future Impact |
|---|---|---|
reassign_speaker | A single sentence (1 SID) | None |
merge_speakers | All sentences of that speaker | Future appearances of the source are also automatically converted to the target |
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value merge_speakers |
source_speaker_id | string | Yes | The speaker ID to be merged (such as Guest-2) |
target_speaker_id | string | Yes | The merge target speaker ID (such as Guest-1) |
Request Example
{
"type": "voice-translation",
"data": {
"action": "merge_speakers",
"source_speaker_id": "Guest-2",
"target_speaker_id": "Guest-1"
}
}
Success Response
{
"type": "voice-translation",
"data": {
"action": "speakers_merged",
"source_speaker_id": "Guest-2",
"target_speaker_id": "Guest-1",
"affected_sids": [3, 5, 7]
}
}
| Field | Type | Description |
|---|---|---|
source_speaker_id | string | The original ID of the merged speaker |
target_speaker_id | string | The original ID of the merge target |
affected_sids | number[] | The list of affected sentence IDs: sentences that belonged to the source speaker, plus the target speaker's existing sentences whose display name changed because of the merge (for example, when the source speaker's custom name is carried over to the target) |
To obtain the target speaker's display label, query
speaker_aliasesor the nextinit_metadataevent.
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
speaker_not_found | 422 | The speaker does not exist | Confirm the speaker ID exists |
merge_speakers_same_id | 400 | The source and target speaker are the same | Use different speaker IDs |
speaker_name_empty | 422 | The speaker ID cannot be empty | Provide a valid speaker ID |
session_not_started | 400 | Speech recognition has not started | Call start first |
Voice Translation - tts_play (Play TTS)
Description
In async mode, manually plays the TTS audio for a specified sentence.
Use Cases
- The user selects a specific sentence for TTS playback
- Playing multiple consecutive sentences
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value tts_play |
sid | int | Yes | The starting sentence ID |
length | int | No | Number of sentences to play (default 1, max 20) |
Note: The maximum value of
lengthis controlled by a server-side setting (default 20).
Request Example (Single Sentence)
{
"type": "voice-translation",
"data": {
"action": "tts_play",
"sid": 5
}
}
Request Example (Multiple Sentences)
{
"type": "voice-translation",
"data": {
"action": "tts_play",
"sid": 5,
"length": 3
}
}
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
tts_not_enabled | 400 | TTS not enabled | Confirm TTS was enabled at start |
A missing sentence or translation does not return
error: If the startingsiddoes not exist, or the sentence has no translation in the target language, noerrormessage is returned. Atts_errorevent is sent instead (see "Response Events"), witherrorset tosentence_not_foundandtranslation_not_foundrespectively. When playing multiple sentences, a failing sentence is skipped and the rest still play.
Voice Translation - tts_stop (Stop TTS)
Description
Stops the TTS audio that is currently playing.
Use Cases
- The user manually stops TTS playback
- Stopping the current playback before switching to another sentence
Request Example
{
"type": "voice-translation",
"data": {
"action": "tts_stop"
}
}
Success Response
{
"type": "voice-translation",
"data": {
"action": "status",
"message": "TTS stopped"
}
}
Voice Translation - tts_mode (Switch TTS Mode)
Description
Switches the TTS playback mode (synchronous/asynchronous) while recording is in progress.
Use Cases
- Switching from automatic playback to manual control
- Switching from manual control to automatic playback
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value tts_mode |
tts_mode | string | Yes | Mode: sync (synchronous) or async (asynchronous); only these two lowercase values are accepted |
Request Example
{
"type": "voice-translation",
"data": {
"action": "tts_mode",
"tts_mode": "async"
}
}
Success Response
{
"type": "voice-translation",
"data": {
"action": "tts_mode_changed",
"tts_mode": "async"
}
}
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
invalid_data | 422 | No tts_mode provided, or a value other than sync or async. For the latter, details.field is tts_mode and details.valid_values lists the accepted values; the mode does not change and no tts_mode_changed is sent | Include sync or async |
Voice Translation - set_tts (Two-Way TTS Settings)
Description
While a two-way mode (conversation) recording is in progress, dynamically toggles TTS on/off or updates the TTS voice settings. Available only in two-way mode.
Use Cases
- Turning the TTS audio response off/on mid-conversation in two-way mode
- Changing the TTS voice or speaking rate for a specific language
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value set_tts |
tts_enabled | boolean | No | Whether to enable two-way TTS (true / false) |
tts_config | object | No | TTS settings per language; the key is the language code, and the value is {voice, speaking_rate} |
Note: At least one of
tts_enabledandtts_configmust be provided.tts_configupdates only the settings for the specified languages; unspecified languages remain unchanged.
Request Example (Disable TTS)
{
"type": "voice-translation",
"data": {
"action": "set_tts",
"tts_enabled": false
}
}
Request Example (Update Voice Settings)
{
"type": "voice-translation",
"data": {
"action": "set_tts",
"tts_enabled": true,
"tts_config": {
"en-US": {
"voice": "en-US-GuyNeural",
"speaking_rate": 1.2
}
}
}
}
Success Response
{
"type": "voice-translation",
"data": {
"action": "tts_updated",
"tts_enabled": true,
"tts_config": {
"zh-TW": { "voice": "zh-TW-HsiaoChenNeural", "speaking_rate": 1.0 },
"en-US": { "voice": "en-US-GuyNeural", "speaking_rate": 1.2 }
}
}
}
| Field | Type | Description |
|---|---|---|
tts_enabled | boolean | The current TTS enabled state |
tts_config | object | The current complete TTS settings (all languages) |
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
invalid_action | 400 | Not two-way mode | This action is available only in conversation mode |
session_not_started | 400 | Speech recognition has not started | Call start first |
Voice Translation - start_speaking (Start Speaking / Manual Mode)
Description
In two-way manual mode (conversation_mode: "manual"), notifies the system that the user has started speaking. From this point on, audio is sent to STT for recognition, and all recognition results accumulate into a single sentence (no automatic segmentation). If it is called again while already speaking, the system first ends the previous sentence (waiting for the last part to finish, then sending its final result) and starts a new one.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value start_speaking |
speaker | int | Yes | User number (1 or 2) |
Request Example
{
"type": "voice-translation",
"data": {
"action": "start_speaking",
"speaker": 1
}
}
Success Response
{
"type": "voice-translation",
"data": {
"action": "status",
"message": "Speaking started"
}
}
If it is called again while already speaking, the final result of the previous sentence is sent as usual, followed by this same status.
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
invalid_action | 400 | Not two-way mode | Use only under the conversation type |
conversation_not_manual_mode | 400 | Not manual mode | Use only in manual mode |
conversation_invalid_speaker | 400 | Invalid user number | Use 1 or 2 |
Voice Translation - stop_speaking (Stop Speaking / Manual Mode)
Description
In two-way manual mode, notifies the system that the user has stopped speaking. The server first waits for the recognizer to finish the last part (usually about 1 second, at most about 3 seconds). The system merges the recognition results accumulated during the period into a single complete sentence and performs translation and TTS synthesis.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value stop_speaking |
Request Example
{
"type": "voice-translation",
"data": {
"action": "stop_speaking"
}
}
Success Response
After stopping speaking, the system sends a complete result event (containing origin and translations):
{
"type": "voice-translation",
"data": {
"action": "result",
"origin": {
"sid": 1,
"language": "zh-TW",
"text": "The complete sentence merged from all recognition during this period",
"is_final": true,
"speaker_id": "Speaker-1",
"start_time": "00:05"
},
"translations": {
"en-US": {
"sid": 1,
"text": "The complete merged sentence from this speaking period",
"is_final": true
}
}
}
}
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
invalid_action | 400 | Not two-way mode | Use only under the conversation type |
conversation_not_speaking | 400 | Not in a speaking state | Call start_speaking first |
Voice Translation - switch_conversation_mode (Switch Conversation Mode)
Description
While two-way mode is in progress, switches between auto-detect mode (auto) and manual mode (manual). If the user is currently speaking when the switch happens, speaking is ended automatically.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value switch_conversation_mode |
conversation_mode | string | Yes | The target mode: auto or manual |
Request Example
{
"type": "voice-translation",
"data": {
"action": "switch_conversation_mode",
"conversation_mode": "manual"
}
}
Success Response
{
"type": "voice-translation",
"data": {
"action": "conversation_mode_changed",
"conversation_mode": "manual"
}
}
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
invalid_action | 400 | Not two-way mode | Use only under the conversation type |
conversation_invalid_mode | 400 | Invalid conversation mode | Use auto or manual |
Voice Translation - set_speaker_language (Set User Language)
Description
While two-way mode is in progress, changes a specified user's language in real time. The system rebuilds the STT connection to accommodate the new language, and the translation target is also updated automatically. The transcript content before the change keeps its original language, and the timestamp continues to count without resetting.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value set_speaker_language |
speaker | int | Yes | User number (1 or 2) |
language | string | Yes | The new language code (such as ja-JP) |
Request Example
{
"type": "voice-translation",
"data": {
"action": "set_speaker_language",
"speaker": 1,
"language": "ja-JP"
}
}
Success Response
{
"type": "voice-translation",
"data": {
"action": "speaker_language_changed",
"speaker_language_map": {
"1": "ja-JP",
"2": "en-US"
}
}
}
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
invalid_action | 400 | Not two-way mode | Use only under the conversation type |
conversation_invalid_speaker | 400 | Invalid user number | Use 1 or 2 |
conversation_invalid_language | 400 | No language provided | Include language |
invalid_transcription_language | 400 | Invalid language code | Use a valid BCP 47 language code |
session_not_started | 400 | Recording has not started | Call start first |
conversation_same_language | 400 | Same as the current language | You can ignore this warning |
conversation_language_same_as_peer | 400 | The new language is the same as the other user | The two users cannot have the same language |
conversation_speaking | 400 | Currently speaking, cannot change language | End speaking before changing |
conversation_language_change_failed | 500 | Language change failed (STT rebuild failed) | Retry later |
Voice Translation - set_speaking_speed (Change Speaking Speed During Recording)
Description
Dynamically adjust the speaking speed (which controls the silence threshold for segmentation) while recording. The system rebuilds the STT connection to apply the new setting, causing a brief interruption in recognition (same as changing a speaker's language during recording). Supported in all recognition modes except multi-speaker (multi_speaker) — including broadcast and multi-language LID; speaking_speed is not applied in multi-speaker mode.
Multi-channel mode (multi_channel, v1.10.0): supported. The new segmentation threshold applies to all channels and takes effect channel by channel (each channel sends its own
channel_statusevent withreason: "speaking_speed", and briefly produces no text while the new setting is applied). The whole session shares a rate limit of one adjustment per 5 seconds — adjusting too quickly returnschannel_rebuild_too_frequent; while paused it returnschannel_action_while_paused.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value set_speaking_speed |
speaking_speed | string | Yes | very_slow / slow / normal / fast / very_fast (see speaking_speed levels for each threshold) |
Request Example
{
"type": "voice-translation",
"data": { "action": "set_speaking_speed", "speaking_speed": "slow" }
}
Success Response
{
"type": "voice-translation",
"data": { "action": "speaking_speed_changed", "speaking_speed": "slow" }
}
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
invalid_data | 422 | Invalid speaking_speed value | Use one of the five valid values |
session_not_started | 400 | Recording not started | Call start first |
invalid_action | 400 | Not supported in multi-speaker mode | Do not call in multi-speaker mode |
set_speaking_speed_failed | 400 | STT rebuild failed | Retry later |
channel_rebuild_too_frequent | 400 | Multi-channel: speaking-speed adjustments are too frequent (only one per 5 seconds is accepted; details carries cooldown_seconds) | Retry later |
channel_action_while_paused | 400 | Multi-channel: the speaking speed cannot be adjusted while paused | Call resume first, then adjust |
If a sentence is in progress when this happens, it cannot be kept and is reported separately via segment_discarded (
reason: "speaking_speed"). You may receive it even whenset_speaking_speed_failedis returned.
Voice Translation - add_channel (Add a Channel / Multi-Channel)
Multi-channel mode only (v1.10.0)
Description
Dynamically adds a channel while a multi-channel recording (recognition_mode: "multi_channel") is in progress. The new channel begins producing text after about 4 seconds; audio sent during this period is not lost, only delayed. After a successful add, the billed channel count follows the actual number of channels from the next minute onward.
Use Cases
- A new speaker joins mid-meeting and gets a new microphone
- Opening channels one by one as attendance grows
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value add_channel |
channels | array | Yes | Exactly 1 channel configuration element (same fields as channels[] in start: channel_id, speaker_name, transcription_languages; in shared mode, omit transcription_languages) |
Numbers cannot be reused: once a
channel_idhas been used (including ones already removed viaremove_channel), it can never be used again; reusing it returnschannel_id_in_use. Speaker identity in the transcript is fixed at the moment each sentence is written, and reusing a number would make one speaker ID correspond to two different people.
Shared mode: the added channel shares the session's recognition and does not increase the billed channel count; its status takes on the first channel's current status directly (see Shared Mode).
Request Example
{
"type": "voice-translation",
"data": {
"action": "add_channel",
"channels": [
{ "channel_id": 4, "speaker_name": "Director Lin", "transcription_languages": ["ja-JP"] }
]
}
}
Success Response
On success, a channel_status event is returned (reason: "added"). The new channel starts in preparing; once that channel produces its first text, another event with status: "ready" arrives:
{
"type": "voice-translation",
"data": {
"action": "channel_status",
"channel": {
"channel_id": 4,
"speaker_name": "Director Lin",
"transcription_languages": ["ja-JP"],
"status": "preparing"
},
"active_channels": 4,
"stt_stream_count": 4,
"reason": "added"
}
}
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
not_multi_channel_session | 400 | This recording is not in multi-channel mode | Available only with recognition_mode=multi_channel |
channels_required | 400 | channels is missing or does not contain exactly 1 element | Provide exactly 1 channel configuration |
invalid_channel_id | 400 | channel_id is out of range (1–8) | Use a number within 1–8 that has never been used |
channel_id_in_use | 400 | The number is in use or has already been removed; it cannot be reused | Use a number that has never been used |
too_many_channels | 400 | The channel count has reached the limit (details carries max and current) | Remove another channel first |
channel_language_required | 400 | Exactly one transcription language was not specified (per_channel) | Provide exactly 1 language in transcription_languages |
channel_language_not_allowed | 400 | transcription_languages was provided in shared mode | Remove the field |
invalid_transcription_language | 400 | Invalid language code | Confirm the language code format is correct (such as zh-TW) |
plan_feature_not_allowed | 403 | Exceeds the plan's limit on simultaneous recognition channels (details.field="max_stt_streams") | Remove another channel or upgrade the plan |
too_many_languages | 400 | The new channel's language would push the number of simultaneously recognized languages over the limit (plan cap; details carries max and received) | Use an existing language or upgrade the plan |
channel_action_while_paused | 400 | Channels cannot be added or disabled while paused | Call resume first |
speaker_name_duplicate | 422 | speaker_name duplicates another channel's current name | Use a different name |
invalid_parameter | 400 | speaker_name is too long (>100 characters) or contains control characters | Fix the parameter per details.field |
session_not_started | 400 | The recording has not started | Call start first |
stt_start_failed | 500 | Speech recognition failed to start for this channel (the addition did not take effect) | Retry later |
Voice Translation - remove_channel (Disable a Channel / Multi-Channel)
Multi-channel mode only (v1.10.0)
Description
Disables a channel while a multi-channel recording is in progress. Once disabled, the channel no longer accepts new audio and no longer counts toward the recognition channel count; the billed channel count decreases immediately from the next minute, and the transcript and audio that channel has already produced are all preserved. Disabling leaves a wrap-up window of about 3 seconds so the channel's last sentence has a chance to make it into the transcript.
Notes:
- The last remaining channel cannot be removed (returns
channel_remove_not_allowed); usestopto end the recording- In
sharedmode, the first entry inchannels[]carries the recognition for the whole session and cannot be removed (returnschannel_remove_not_allowed)- A removed
channel_idcannot be reused (seeadd_channel); to change a language, useset_channel_language— do not remove and re-add- Sending audio with that number after removal returns
unknown_channel_id
Use Cases
- A speaker leaves mid-session; release the channel to reduce billing
- Reclaiming a temporarily opened guest microphone
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value remove_channel |
channel_id | int | Yes | The channel number to disable |
Request Example
{
"type": "voice-translation",
"data": {
"action": "remove_channel",
"channel_id": 4
}
}
Success Response
On success, a channel_status event is returned (reason: "removed", status: "removed"):
{
"type": "voice-translation",
"data": {
"action": "channel_status",
"channel": {
"channel_id": 4,
"speaker_name": "Director Lin",
"transcription_languages": ["ja-JP"],
"status": "removed"
},
"active_channels": 3,
"stt_stream_count": 3,
"reason": "removed"
}
}
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
not_multi_channel_session | 400 | This recording is not in multi-channel mode | Available only with recognition_mode=multi_channel |
channel_id_required | 400 | channel_id is missing | Provide the number of the channel to disable |
invalid_channel_id | 400 | channel_id is out of range (1–8) | Use a valid channel number |
unknown_channel_id | 400 | Unknown channel_id (it may have been removed already) | Make sure the channel exists and has not been removed |
channel_remove_not_allowed | 400 | The last remaining channel cannot be removed; or the first channel in shared mode | Use stop to end the recording |
channel_action_while_paused | 400 | Channels cannot be added or disabled while paused | Call resume first |
session_not_started | 400 | The recording has not started | Call start first |
Voice Translation - set_channel_language (Change a Channel's Language / Multi-Channel)
Multi-channel mode only (v1.10.0)
Description
Changes the transcription language of a single channel while a multi-channel recording is in progress. The channel_id and speaker identity (speaker_id) stay the same and the transcript remains continuous; the system switches that channel to the new language (it takes effect in about 4 seconds, during which the channel briefly produces no text and sends a channel_status event with reason: "language_change"). When the new language is one the session has not used yet, it is governed by the platform limit (10 simultaneously recognized languages) and the plan's limit on simultaneously recognized languages.
Do not substitute
remove_channel+add_channel: channel numbers cannot be reused, and switching numbers would turn the same person into two different speakers in the transcript — with no way to fix it afterward.
Use Cases
- The same speaker switches to another language mid-session
- Correcting a channel language that was set incorrectly at
start
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value set_channel_language |
channel_id | int | Yes | Target channel number |
transcription_languages | string[] | Conditional | The new transcription language (exactly 1; more than 1 returns invalid_parameter). Provide either this or language |
language | string | Conditional | The new transcription language (single-value form). Provide either this or transcription_languages; providing both with different values returns invalid_parameter |
Request Example
{
"type": "voice-translation",
"data": {
"action": "set_channel_language",
"channel_id": 2,
"transcription_languages": ["ja-JP"]
}
}
Success Response
On success, a channel_status event is returned (reason: "language_change", status: "preparing"); once the switch completes and the channel produces its first text, another event with status: "ready" arrives:
{
"type": "voice-translation",
"data": {
"action": "channel_status",
"channel": {
"channel_id": 2,
"speaker_name": "Alex",
"transcription_languages": ["ja-JP"],
"status": "preparing"
},
"active_channels": 3,
"stt_stream_count": 3,
"reason": "language_change"
}
}
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
not_multi_channel_session | 400 | This recording is not in multi-channel mode | Available only with recognition_mode=multi_channel |
channel_id_required | 400 | channel_id is missing | Provide the target channel number |
invalid_channel_id | 400 | channel_id is out of range (1–8) | Use a valid channel number |
unknown_channel_id | 400 | Unknown channel_id (it may have been removed) | Make sure the channel exists and has not been removed |
channel_language_required | 400 | No new transcription language provided | Provide transcription_languages[0] or language |
invalid_parameter | 400 | Same as the current language, both fields provided with different values, or more than one language | Fix the parameters according to details |
invalid_transcription_language | 400 | Invalid language code | Confirm the language code format is correct (such as zh-TW) |
channel_rebuild_too_frequent | 400 | Only one switch per channel is accepted within 5 seconds (details carries cooldown_seconds) | Retry later |
too_many_languages | 400 | The new language would push the number of simultaneously recognized languages over the platform limit (10) or the plan limit | Use an existing language or upgrade the plan |
channel_action_while_paused | 400 | The channel language cannot be changed while paused | Call resume first |
session_not_started | 400 | The recording has not started | Call start first |
channel_language_not_allowed | 400 | shared mode does not support changing a channel's language | The session language is set at start and cannot be changed during the recording |
stt_start_failed | 500 | Speech recognition failed to rebuild for this channel with the new language | Retry later (the channel waits for reconnection) |
If a sentence is in progress when this happens, it cannot be kept and is reported separately via segment_discarded (
reason: "language_change"). The event carrieschannel_idto identify the channel.
Voice Translation - broadcast_go_live (Switch to the Live Phase)
Description
Switches from the broadcast standby phase (standby) to the live phase (live). After switching, STT/translation results begin broadcasting to viewers and start being written to the transcript.
Use Cases
- The host confirms the equipment is working and starts the official broadcast
- Switching from the warm-up phase to live streaming
Request Example
{
"type": "voice-translation",
"data": {
"action": "broadcast_go_live"
}
}
Success Response
{
"type": "voice-translation",
"data": {
"action": "broadcast_phase_changed",
"phase": "live",
"message": "Broadcast started"
}
}
| Field | Type | Description |
|---|---|---|
phase | string | The new phase (live) |
message | string | Status description message |
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
broadcast_not_enabled | 400 | Not broadcast mode | Confirm type: "broadcast" |
session_not_started | 400 | Speech recognition has not started, or the recording has ended (including while it is still being processed after ending) | Call start first if it has not started |
Note: If already in the live phase, a status message "Broadcast is already in progress" is returned and is not treated as an error.
If a sentence is in progress when this happens, it cannot be kept and is reported separately via segment_discarded (
reason: "broadcast_go_live"). Standby-phase segments are never written to the final transcript anyway.
Voice Translation - broadcast_announcement (Send an Announcement)
Description
The host sends a custom message announcement to all viewers. Viewers receive an announcement event via SSE. The announcement message is automatically translated into all translation languages, and the SSE event viewers receive includes a translations field.
Use Cases
- Notifying viewers that the meeting is about to end
- Sending an important reminder or announcement
- One-way communication with viewers
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value broadcast_announcement |
message | string | Yes | The announcement message content |
Request Example
{
"type": "voice-translation",
"data": {
"action": "broadcast_announcement",
"message": "The meeting will end in 5 minutes"
}
}
Success Response
{
"type": "voice-translation",
"data": {
"action": "status",
"message": "Announcement sent"
}
}
The SSE event viewers receive (with translations):
event: announcement
data: {"message":"The meeting will end in 5 minutes","translations":{"en-US":"The meeting will end in 5 minutes","ja-JP":"会議は5分後に終了します"}}
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
broadcast_not_enabled | 400 | Not broadcast mode | Confirm type: "broadcast" |
invalid_parameter | 400 | Message is empty | Provide a valid message parameter |
Voice Translation - set_standby_message (Set the Standby Phase Message)
Description
During the broadcast standby phase (standby), dynamically sets the message shown to viewers. This allows the host to enter standby mode and then set the waiting message, rather than being required to provide it at start.
The message is automatically translated into all translation languages, and the SSE event viewers receive includes a translations field.
Use Cases
- After entering standby mode, dynamically set the waiting message shown to viewers
- Update the text on the standby screen before going live
- Reduce the required fields before starting the broadcast
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Fixed value set_standby_message |
message | string | Yes | The text displayed during the standby phase (translated for viewers of each language via the existing translation pipeline) |
Request Example
{
"type": "voice-translation",
"data": {
"action": "set_standby_message",
"message": "The talk is about to begin, please wait..."
}
}
Success Response
{
"type": "voice-translation",
"data": {
"action": "status",
"message": "Standby phase text updated"
}
}
Event Viewers Receive
After a successful setting, all viewers in the standby phase receive an updated standby event via SSE:
event: standby
data: {"message":"The talk is about to begin, please wait...","translations":{"en-US":"The presentation is about to begin, please wait...","ja-JP":"プレゼンテーションがまもなく始まります。お待ちください..."}}
Note: The
translationsfield contains the translation results for all translation languages. The frontend can display the corresponding translation based on the language the viewer selects.
Error Responses
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
broadcast_not_enabled | 400 | Not broadcast mode | Confirm type: "broadcast" |
broadcast_not_in_standby | 400 | Not in the standby phase | Can be used only during the standby phase |
Note: This action can be used only during the standby phase (standby). If the broadcast has already entered the live phase (live), an error is returned.
Response Events
The following are the commonly used WebSocket response events.
Note: This is not the complete list:
translation_language_removed,summary_updated,summary_done,summary_error,upload_errorandspeakers_auto_mergedare not documented on this page. See the event reference for all 36 events.
session_started - Session Started Successfully
After a start action succeeds, the server returns an event containing complete session initialization info. The frontend can distinguish the recording type via recording_type.
General recordings (transcribe / conversation / record):
{
"type": "voice-translation",
"data": {
"action": "session_started",
"session_id": "550e8400-e29b-41d4-a716-446655440000",
"task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
"recording_type": "transcribe",
"recognition_mode": "single",
"resume_token": "L0VBAwIy... (43 characters)",
"resume_grace_seconds": 45,
"server_time": 1749550000000,
"message": "Speech recognition started"
}
}
Broadcast mode (broadcast):
{
"type": "voice-translation",
"data": {
"action": "session_started",
"session_id": "550e8400-e29b-41d4-a716-446655440000",
"task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
"recording_type": "broadcast",
"recognition_mode": "multi_speaker",
"phase": "standby",
"viewer_count": 0,
"queue_count": 0,
"peak_viewers": 0,
"total_viewers": 0,
"resume_token": "L0VBAwIy... (43 characters)",
"resume_grace_seconds": 45,
"server_time": 1749550000000,
"message": "Speech recognition started"
}
}
| Field | Type | Description |
|---|---|---|
session_id | string | Session ID |
task_id | string | Task ID (can be used for subsequent API queries) |
recording_type | string | Recording type: transcribe, conversation, record, broadcast |
recognition_mode | string | Recognition mode: single, multi_speaker |
resume_token | string | Resume token for reconnection (43 characters). After a disconnect, reconnecting with this token within the resume_grace_seconds grace window rejoins the original session. It is sent ahead of time in session_started; keep it until the session ends. |
resume_grace_seconds | int | Reconnection grace period in seconds (default 45). This is wall-clock real time and keeps counting down after a disconnect. |
server_time | int64 | Current server time in unix milliseconds. The client may use the pair (server_time, client time when received) to estimate clock skew for reference; the grace countdown is still based on the client wall-clock. |
message | string | Status description message |
phase | string | Broadcast phase: standby or live (broadcast mode only) |
viewer_count | int | Current number of online viewers (broadcast mode only) |
queue_count | int | Number of viewers waiting in the queue (broadcast mode only) |
peak_viewers | int | Peak number of viewers for this broadcast (broadcast mode only) |
total_viewers | int | Total cumulative number of viewers who have connected (broadcast mode only) |
channel_mode | string | Multi-channel sub-mode: per_channel or shared (multi-channel mode only). This is also how the frontend confirms the server actually started in multi-channel mode |
channels | array | Multi-channel channel list (multi-channel mode only). Each entry carries channel_id, speaker_name, transcription_languages, status (preparing / ready / removed / error); transcription_languages is not present in shared mode |
resume_ok - Resume Succeeded
The event returned by the server when a reconnection with resume_token succeeds within the resume_grace_seconds grace window after a disconnect. Upon receiving it, restart the audio stream just as you would after start (WebM/Opus must send a fresh container; PCM can continue directly), and align your local content using server_last_sid / server_last_offset_ms. If is_paused is true, the session was paused before the disconnect, so stay paused after resuming and send no audio. For the full resume handshake flow, see Connection and Authentication; for field details, see WebSocket Events Reference.
{
"type": "voice-translation",
"data": {
"action": "resume_ok",
"session_id": "550e8400-e29b-41d4-a716-446655440000",
"task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
"recording_type": "transcribe",
"recognition_mode": "single",
"server_last_sid": 42,
"server_last_offset_ms": 125000,
"server_recording_ms": 127500,
"is_paused": false,
"settings": { ... },
"message": "Session resumed"
}
}
| Field | Type | Description |
|---|---|---|
session_id | string | Session ID (WS connection level; invalidated when the connection ends) |
task_id | string | Task ID (can be used for subsequent API queries) |
recording_type | string | Recording type: transcribe, conversation, record, broadcast |
recognition_mode | string | Recognition mode: single, multi_speaker |
server_last_sid | int | The server's current last sentence number (sid). The client should ignore duplicate messages with sid ≤ this value. |
server_last_offset_ms | int64 | The transcript timeline position (milliseconds) of the breakpoint, based on the length of audio already processed (not wall-clock; the audio timeline is frozen during the disconnect). The client uses it to place reconnected content at the correct timeline position. |
server_recording_ms | int64 | The recording-head timestamp (milliseconds, including silence) that the transcript timeline resumes from after reconnect. The client uses it to align its recording-second header to the same timeline the transcript uses. The difference from server_last_offset_ms is the trailing audio (silence or not-yet-finalized speech) after the last finalized sentence before the disconnect. Omitted when 0. |
is_paused | bool | The paused state the server considers authoritative after resuming: true = it was paused before the disconnect, so the client should stay paused (reopen the audio stream then pause immediately and send no audio); false = recording normally. Used to align the pause UI after a network reconnect or a full-page refresh. Omitted is treated as false (backward compatible with older servers). |
settings | object | The recording settings currently held by the session (for state reconciliation, added in v1.6.4); see Connection and Authentication for details. |
message | string | Status description message (always "Session resumed") |
Note: The two times must not be mixed: the transcript timeline (
server_last_offset_ms) is based on the audio and freezes during a disconnect; the grace reconnection window (resume_grace_seconds) is wall-clock real time and keeps counting down during a disconnect. Deciding "whether reconnection is still possible" must use wall-clock, not the audio timeline position.
Multi-channel mode: the
settingsofresume_okcontainschannel_modeandchannels[](each entry withchannel_id,speaker_name,transcription_languages,status;transcription_languagesis not present insharedmode), so that after resuming the frontend can confirm the server is still in multi-channel mode and that each channel's channel-to-language binding is unchanged. After the resume, each channel first showspreparingin the snapshot, and sends its ownchannel_statusevent (ready,reason: "reconnect") when it starts producing text. Insharedmode, the other channels' status in the snapshot and events follows the first channel (see Shared Mode).
result - Recognition/Translation Result
Speech recognition and translation results. A single result event may contain origin (recognition result) and/or translations (translation results).
origin (speech recognition result):
{
"type": "voice-translation",
"data": {
"action": "result",
"origin": {
"sid": 1,
"language": "zh-TW",
"text": "Hello, nice to meet you",
"is_final": true,
"speaker_id": "0",
"detected_language": "zh-TW",
"start_time": "00:05"
}
}
}
| Field | Type | Description |
|---|---|---|
sid | int | Sentence number, starting from 1 |
language | string | Source language code. In two-way mode and multi-language transcription this is the language determined by the system — in multi-language mode it is determined per sentence (so it varies within a single recording), and when the determination is unreliable the configured value is kept, so the field is never empty. |
text | string | The recognized text |
is_final | boolean | Whether it is the final result |
speaker_id | string | Speaker ID. In multi-channel mode it is determined by the channel, format channel_{N} (such as channel_1) |
speaker_label | string | Speaker display label (optional). In multi-channel mode this is the speaker_name provided at start / add_channel (after a rename_speaker, the new label; equal to speaker_id when no name was provided) |
detected_language | string | The detected language. In two-way mode, this is determined automatically by the system. |
start_time | string | Sentence start time (mm:ss); not sent during the broadcast standby phase; after going live, counts from 00:00. |
channel_id | int | Multi-channel mode only: the source channel number of this sentence. Not present outside multi-channel mode |
Multi-channel mode: each channel segments sentences independently, so sentences from different channels arrive interleaved and
sidis not guaranteed to be sequential per channel;origin.languageis the language bound to that channel.
translations (translation results):
{
"type": "voice-translation",
"data": {
"action": "result",
"translations": {
"en-US": {
"sid": 1,
"text": "Hello, nice to meet you",
"is_final": true
}
}
}
}
Translation results are keyed by language code, and each language's translation object contains:
| Field | Type | Description |
|---|---|---|
sid | int | Sentence number |
text | string | The translated text |
is_final | boolean | Whether it is the final result |
is_retranslation | boolean | Whether it is a retranslation result (only during retranslate) |
Multi-channel mode:
translationsdoes not carrychannel_id; usesidto match back tooriginfor the source channel.
status - Generic Status Response
Used to confirm operations such as pause, resume, stop, set_name, tts_stop, and start_speaking.
{
"type": "voice-translation",
"data": {
"action": "status",
"status": "paused",
"message": "Speech recognition paused"
}
}
| Field | Type | Description |
|---|---|---|
status | string | Machine-readable recording lifecycle state: live (resumed) / paused / ended (stopped). Only pause / resume / stop carry it; set_name etc. do not. ended is always sent before task_complete. Floating subtitle consumers should act on it: paused → freeze, ended → close the window, live → resume. |
message | string | Human-readable status text (format not guaranteed; do not parse — rely on the status field) |
task_complete - Task Processing Complete
Triggered after stop when the audio file and transcript have been uploaded. task_id can be used to query task details via the REST API afterward.
- It is always sent after
status: "ended". - The transcript includes the summary, so this event waits until the title and summary have been generated. This usually takes a few seconds to tens of seconds, and longer for a long summary or a slow service, up to about 8 minutes.
- The concurrent recording slot is released when this event is sent: you can start the next recording as soon as you receive it.
- To detect completion, we recommend also supporting the Webhook or a REST query rather than relying on this event alone.
- You also receive this event when a recording ends automatically after a long silence.
{
"type": "voice-translation",
"data": {
"action": "task_complete",
"task_id": "550e8400-e29b-41d4-a716-446655440000",
"message": "Task processing complete"
}
}
Recordings that received no audio at all: whether the recording ended with stop, insufficient credit, or automatically, if no audio was received during the whole recording you still receive status: "ended" and then this event, with noAudio: true. Such a recording has no transcript or audio file and its status is failed; a recording.failed Webhook is also sent (failure_source is no_audio). Do not try to read the transcript; tell the user that no sound was received instead. Minutes already elapsed are billed as usual.
{
"type": "voice-translation",
"data": {
"action": "task_complete",
"task_id": "550e8400-e29b-41d4-a716-446655440000",
"noAudio": true,
"message": "Task processing complete"
}
}
A broadcast that ends during the standby phase (before going live) sends only
status: "ended", not this event: no recording is created during the standby phase.
| Field | Type | Description |
|---|---|---|
task_id | string | Recording UUID, can be used for subsequent API queries |
noAudio | boolean | Present only when no audio was received during the whole recording, and always true; absent when audio was recorded |
message | string | Status description |
config_updated - Settings Update Complete
Triggered after the config action succeeds.
Note: Sent only when the configuration is accepted. A rejected configuration returns
type: "error"instead, so your client must listen for that as well or the request appears to go unanswered.
{
"type": "voice-translation",
"data": {
"action": "config_updated",
"updated": ["terminology", "fuzzy_correction", "translation_dict"],
"message": "Settings updated"
}
}
| Field | Type | Description |
|---|---|---|
updated | string[] | The setting types that were updated: terminology, fuzzy_correction, translation_dict |
message | string | Status message |
terminology_effective | string | (Optional) Appears when terminology is updated during recording; a value of "next_turn" means the new terminology takes effect from the next sentence. Does not appear in the initial config |
unknown_languages | string[] | (Optional) Glossary language codes that could not be recognized (for example zh, chinese). Those entries will not take effect, but the config still succeeds |
inactive_languages | string[] | (Optional) Codes that are valid but are not used by this recording. Does not appear before the recording starts (the language list is not settled yet) |
inactive_dict_languages | string[] | (Optional) The same, for translation-dictionary target languages |
homophone_conflicts | object[] | (Optional) Groups of terms in your glossary that share a pronunciation; each entry carries languages and terms. Not present before the recording has started (the language list is not settled yet) |
Reading
homophone_conflicts: when two terms that sound alike are registered together (公事包 and 公式包, say) and the transcript contains a third spelling with the same pronunciation, the system can only correct it to one of them, and which one is not guaranteed. Both terms themselves still work; the only ambiguity is which term an unregistered homophone misspelling is attributed to.This is a warning, not an error — the glossary is still accepted and
configstill succeeds. If the pair matters to you, list the misspelling explicitly withfuzzy_correctionso it is pinned to the term you want.
languageslists every language that shares the same term index. Chinese regional codes (zh-TW,zh-CN,zh-HKand so on) are treated as one group for glossary matching, so they share a single entry rather than each reporting one."homophone_conflicts": [ { "languages": ["zh-TW"], "terms": ["公事包", "公式包"] } ]
tts_ready - TTS Audio Ready
TTS speech synthesis completion event. Contains the audio data and Word Boundary information (which can be used for a karaoke effect).
{
"type": "voice-translation",
"data": {
"action": "tts_ready",
"sid": 1,
"language": "en-US",
"transcript": "你好,很高興認識你",
"text": "Hello, nice to meet you",
"audio": "Base64EncodedMP3...",
"format": "mp3",
"duration_ms": 2500,
"boundaries": [
{"offset_ms": 0, "duration_ms": 350, "text_offset": 0, "word_length": 5, "text": "Hello"},
{"offset_ms": 350, "duration_ms": 100, "text_offset": 5, "word_length": 1, "text": ","},
{"offset_ms": 500, "duration_ms": 250, "text_offset": 7, "word_length": 4, "text": "nice"},
{"offset_ms": 750, "duration_ms": 200, "text_offset": 12, "word_length": 2, "text": "to"},
{"offset_ms": 950, "duration_ms": 350, "text_offset": 15, "word_length": 4, "text": "meet"},
{"offset_ms": 1300, "duration_ms": 300, "text_offset": 20, "word_length": 3, "text": "you"}
]
}
}
| Field | Type | Description |
|---|---|---|
sid | int | Sentence number |
language | string | TTS language |
transcript | string | The original transcript (STT recognition result) |
text | string | The translated text (TTS synthesis source) |
audio | string | Base64-encoded MP3 audio |
format | string | Audio format (fixed value mp3) |
duration_ms | int | Total audio duration (milliseconds) |
boundaries | array | Array of Word Boundaries |
Word Boundary Field Descriptions
| Field | Type | Description |
|---|---|---|
offset_ms | int | The word's start time in the audio (milliseconds) |
duration_ms | int | The word's duration (milliseconds) |
text_offset | int | Position in the original string (character index) |
word_length | int | Word length (number of characters) |
text | string | The word content |
tts_error - TTS Synthesis Failed
TTS synthesis failure event.
{
"type": "voice-translation",
"data": {
"action": "tts_error",
"sid": 1,
"language": "en-US",
"error": "translation_not_found",
"message": "No translation available for language: en-US"
}
}
| Field | Type | Description |
|---|---|---|
sid | int | Sentence number |
language | string | TTS language |
error | string | Error code |
message | string | Error message |
transcript | string | The corresponding original transcript, to help the frontend locate the point of failure. This field is always present; it is an empty string "" when the sentence is not found |
TTS Error Codes
| Error Code | Description |
|---|---|
sentence_not_found | The specified sentence was not found (the sid passed to tts_play does not exist) |
translation_not_found | No translation found for that language |
tts_invalid_language | The TTS language is not supported |
tts_invalid_voice | The voice name is invalid |
tts_connection_failed | Could not connect to the speech synthesis service |
tts_timeout | Speech synthesis timed out |
tts_synthesis_failed | Speech synthesis failed |
viewer_count - Viewer Count Update
Broadcast mode only
During a broadcast, the system monitors the viewer count and pushes this event to the host whenever it changes.
{
"type": "voice-translation",
"data": {
"action": "viewer_count",
"viewer_count": 45,
"queue_count": 8,
"peak_viewers": 50,
"total_viewers": 123
}
}
| Field | Type | Description |
|---|---|---|
viewer_count | int | Current number of online viewers |
queue_count | int | Number of viewers waiting in the queue |
peak_viewers | int | Peak number of viewers for this broadcast |
total_viewers | int | Total cumulative number of viewers who have connected |
Note: This event is pushed only when the viewer count or queue count changes, to avoid unnecessary message traffic.
viewer_joined - Viewer Joined
Broadcast mode only
When a viewer joins the broadcast, the host receives this event.
{
"type": "voice-translation",
"data": {
"action": "viewer_joined",
"viewer_count": 5,
"queue_count": 2
}
}
| Field | Type | Description |
|---|---|---|
viewer_count | number | Current number of viewers |
queue_count | number | Number waiting in the queue |
viewer_left - Viewer Left
Broadcast mode only
When a viewer leaves the broadcast, the host receives this event.
{
"type": "voice-translation",
"data": {
"action": "viewer_left",
"viewer_count": 4,
"queue_count": 1
}
}
| Field | Type | Description |
|---|---|---|
viewer_count | number | Current number of viewers |
queue_count | number | Number waiting in the queue |
broadcast_phase_changed - Broadcast Phase Changed
Triggered when the broadcast phase switches from standby to live.
{
"type": "voice-translation",
"data": {
"action": "broadcast_phase_changed",
"phase": "live",
"message": "Broadcast started"
}
}
| Field | Type | Description |
|---|---|---|
phase | string | The new phase: standby or live |
message | string | Status description message |
broadcast_recording_ready - Broadcast Recording Ready
Triggered after a broadcast goes live, returning the finalized task_id for this broadcast. The task_id in session_started is an initial value for the connection stage, not the final ID for this broadcast; you receive this event after going live regardless of whether the broadcast went through the standby phase first or started directly with broadcast_phase: "live" (the default). For subsequent operations (such as exchanging for a Floating Subtitle Feed Token), always use the task_id from this event — using the initial value returns no matching record.
- When
startgoes live directly, this event always arrives aftersession_started. - When the host disconnects and, within the grace period, sends
startagain with the samebroadcast_token(a takeover), this event on the new connection carries a newtask_id. The broadcast therefore has two recordings, one before and one after the takeover, each completed and notified separately.
{
"type": "voice-translation",
"data": {
"action": "broadcast_recording_ready",
"task_id": "3f9a1c2e-..."
}
}
| Field | Type | Description |
|---|---|---|
task_id | string | The finalized recording ID for this broadcast (valid after going live) |
speaker_renamed - Speaker Renamed
Speaker global rename completion event.
{
"type": "voice-translation",
"data": {
"action": "speaker_renamed",
"speaker_id": "Guest-1",
"new_label": "Manager Wang",
"affected_sids": [1, 3, 5, 8]
}
}
| Field | Type | Description |
|---|---|---|
speaker_id | string | The resolved original speaker ID (even if the input was a display label, the event returns the original ID) |
new_label | string | The new display label |
affected_sids | int[] | The list of affected sentence numbers |
speaker_reassigned - Speaker Identity Changed
Single-sentence speaker identity change completion event.
{
"type": "voice-translation",
"data": {
"action": "speaker_reassigned",
"sid": 5,
"old_speaker_id": "Guest-1",
"new_speaker_id": "Guest-2",
"new_speaker_label": "Lee Hsiao-hua"
}
}
| Field | Type | Description |
|---|---|---|
sid | int | The changed sentence number |
old_speaker_id | string | The original speaker ID |
new_speaker_id | string | The new original speaker ID |
new_speaker_label | string | The new speaker display label (after applying speaker_aliases; equals new_speaker_id when no alias exists) |
speakers_merged - Speakers Merged
Speaker merge completion event. After the merge, future recognition results for that source speaker are also automatically converted to the target speaker. This applies within the current recognition pass only: after a connection recovers, speakers are re-identified and you need to merge again.
{
"type": "voice-translation",
"data": {
"action": "speakers_merged",
"source_speaker_id": "Guest-2",
"target_speaker_id": "Guest-1",
"affected_sids": [3, 5, 7]
}
}
| Field | Type | Description |
|---|---|---|
source_speaker_id | string | The original ID of the merged speaker |
target_speaker_id | string | The original ID of the merge target |
affected_sids | number[] | The list of affected sentence IDs: sentences that belonged to the source speaker, plus the target speaker's existing sentences whose display name changed because of the merge (for example, when the source speaker's custom name is carried over to the target) |
To obtain the target speaker's display label, query
speaker_aliasesor the nextinit_metadataevent.
language_switch_start - Language Switch Started
Language switch start event, sent after the switch_language action is triggered.
{
"type": "voice-translation",
"data": {
"action": "language_switch_start",
"translation_language": "ja-JP",
"translation_languages": ["en-US", "ja-JP"],
"total_segments": 15
}
}
| Field | Type | Description |
|---|---|---|
translation_language | string | The single language for this operation (the new target for a replace, or the language added by op:add) |
translation_languages | string[] | Authoritative snapshot of the current full set of translation languages. For a multi-language op:add this is the complete set including existing languages. Clients should overwrite their local language set with this directly, rather than inferring "append" vs "replace" from translation_language (when only 1 language exists, an op:add expanding to a 2nd would be misread as a replace and drop the existing language). |
total_segments | int | The number of sentences that need retranslation |
batch_retranslation - Batch Retranslation Result
Batch retranslation result event, sent sentence by sentence during the language switch process.
{
"type": "voice-translation",
"data": {
"action": "batch_retranslation",
"sid": 3,
"translations": {
"ja-JP": {
"sid": 3,
"text": "今日はプロジェクトの進捗について話し合いましょう",
"is_final": true,
"is_retranslation": true
}
}
}
}
| Field | Type | Description |
|---|---|---|
sid | int | Sentence number |
translations | object | Translation results (same format as result's translations) |
language_switch_done - Language Switch Complete
Language switch completion event.
{
"type": "voice-translation",
"data": {
"action": "language_switch_done",
"translation_language": "ja-JP",
"translation_languages": ["en-US", "ja-JP"],
"success_count": 15,
"failed_count": 0
}
}
| Field | Type | Description |
|---|---|---|
translation_language | string | The single language for this operation |
translation_languages | string[] | Authoritative snapshot of the current full translation-language set (same as language_switch_start; overwrite the local set). |
success_count | int | The number of successfully translated sentences |
failed_count | int | The number of sentences that failed to translate |
tts_mode_changed - TTS Mode Changed
TTS playback mode change event.
{
"type": "voice-translation",
"data": {
"action": "tts_mode_changed",
"tts_mode": "async"
}
}
| Field | Type | Description |
|---|---|---|
tts_mode | string | The new mode: sync or async |
language_switched - Two-Way Language Switch Complete
Two-way mode (conversation) language switch completion event. Triggered after switch_language successfully switches the STT source language in two-way mode.
{
"type": "voice-translation",
"data": {
"action": "language_switched",
"language": "en-US",
"translation_language": "zh-TW",
"message": "Language switched"
}
}
| Field | Type | Description |
|---|---|---|
language | string | The new active language (STT source) |
translation_language | string | The new translation target language |
message | string | Status message |
tts_updated - Two-Way TTS Settings Updated
Two-way mode (conversation) TTS settings update event. Triggered after set_tts successfully updates the TTS toggle or voice settings.
{
"type": "voice-translation",
"data": {
"action": "tts_updated",
"tts_enabled": true,
"tts_config": {
"zh-TW": { "voice": "zh-TW-HsiaoChenNeural", "speaking_rate": 1.0 },
"en-US": { "voice": "en-US-GuyNeural", "speaking_rate": 1.2 }
}
}
}
| Field | Type | Description |
|---|---|---|
tts_enabled | boolean | Whether TTS is enabled |
tts_config | object | The TTS settings for each language (voice, speaking_rate) |
conversation_mode_changed - Conversation Mode Changed
Two-way mode (conversation) conversation mode change event. Triggered after switch_conversation_mode successfully switches between auto/manual mode.
{
"type": "voice-translation",
"data": {
"action": "conversation_mode_changed",
"conversation_mode": "manual"
}
}
| Field | Type | Description |
|---|---|---|
conversation_mode | string | The new conversation mode: auto or manual |
speaker_language_changed - User Language Changed
Two-way mode (conversation) user language change event. Triggered after set_speaker_language successfully changes a user's language, including the complete language mapping after the change.
{
"type": "voice-translation",
"data": {
"action": "speaker_language_changed",
"speaker_language_map": {
"1": "ja-JP",
"2": "en-US"
}
}
}
| Field | Type | Description |
|---|---|---|
speaker_language_map | object | The user language mapping after the change (keys are user number strings) |
speaking_speed_changed - Speaking Speed Changed
Speaking speed change event during recording. Triggered after set_speaking_speed successfully applies the new speed (STT rebuild complete), returning the applied speed level.
{
"type": "voice-translation",
"data": {
"action": "speaking_speed_changed",
"speaking_speed": "slow"
}
}
| Field | Type | Description |
|---|---|---|
speaking_speed | string | The applied speed: very_slow / slow / normal / fast / very_fast |
channel_status - Channel Status Changed
Multi-channel mode only (v1.10.0)
In multi-channel mode (recognition_mode: "multi_channel"), fired whenever a channel's status changes: the success responses of add_channel / remove_channel, setting changes from set_channel_language and set_speaking_speed, pause/resume and automatic reconnection after a disconnect, and channel failures.
{
"type": "voice-translation",
"data": {
"action": "channel_status",
"channel": {
"channel_id": 3,
"speaker_name": "Lee",
"transcription_languages": ["zh-TW"],
"status": "preparing"
},
"active_channels": 3,
"stt_stream_count": 3,
"reason": "added"
}
}
| Field | Type | Description |
|---|---|---|
channel | object | The channel whose status changed, with channel_id, speaker_name, transcription_languages, status |
active_channels | int | Current number of active channels (number of speakers) |
stt_stream_count | int | Current number of recognition channels being counted (billed on this count as each minute begins); always 1 in shared mode |
reason | string | Optional. Why this status change happened; see the table below |
status state machine
Each time a channel starts or prepares recognition again, it first enters preparing and transitions to ready once it produces its first text:
| status | Description |
|---|---|
preparing | Preparing (creating the channel or applying a new setting). Speech on this channel appears a few seconds late, but is not lost; the one exception is a sentence in progress when the setting is applied, which cannot be kept (see segment_discarded) |
ready | The channel has received its first recognition result |
removed | Disabled by remove_channel |
error | Speech recognition on this channel failed and cannot recover automatically |
Shared mode: the other channels'
preparing/ready/errorfollow the first channel, and each channel receives its own event when the first channel's status changes;removedfollows each channel's own status. See Shared Mode.
reason values
| reason | Description |
|---|---|
added | Added by add_channel |
language_change | Language switched by set_channel_language |
reconnect | Automatic reconnection after the channel failed unexpectedly |
resumed | Resumed after a pause |
removed | Disabled by remove_channel |
speaking_speed | New setting applied channel by channel via set_speaking_speed |
stt_error | Speech recognition failure |
segment_discarded - Segment Discarded
Tells you that a sid has been discarded and will receive no further events — no is_final: true original text, and no translation for that segment.
Some operations interrupt recognition. If a sentence happens to be in progress at that moment, it cannot be kept. Your client may already have received interim results (is_final: false) for it, so on receiving this event, clear that sid from any "translating" state.
Automatic conversation mode never receives this event — in that mode, a sentence in progress is finalized and sent with is_final: true, and its content is kept in the transcript.
{
"type": "voice-translation",
"data": {
"action": "segment_discarded",
"sid": 7,
"reason": "speaking_speed",
"channel_id": 1
}
}
| Field | Type | Description |
|---|---|---|
sid | int | The discarded segment number |
reason | string | speaking_speed (set_speaking_speed) / language_change (set_channel_language) / reconnect (automatic reconnection after the recognition connection failed) / resumed (session resume or resuming after a pause) / broadcast_go_live (moving from standby to live) |
channel_id | int | Optional. Source channel number; sent only in multi-channel mode |
The first four reason values come from the same set as channel_status, so you can match both events back to the same operation.
You may already have received this event even when
set_speaking_speedfails (set_speaking_speed_failed) — recognition is interrupted before the rebuild, so the segment cannot be kept whether the rebuild succeeds or not.
segment_uploaded - Audio Segment Upload Complete
Audio segment upload completion event. Triggered each time an audio segment is successfully uploaded to cloud storage; can be used to show upload progress on the frontend.
{
"type": "voice-translation",
"data": {
"action": "segment_uploaded",
"segment_index": 0,
"duration_sec": 30.5
}
}
| Field | Type | Description |
|---|---|---|
segment_index | number | Segment index (starting from 0) |
duration_sec | number | The duration of this segment (seconds) |
stt_event - STT Connection Status Event
STT connection status event. Triggered when the connection status of the speech recognition service changes; can be used to show the STT service status on the frontend.
{
"type": "voice-translation",
"data": {
"action": "stt_event",
"event": "reconnected",
"message": "STT reconnected"
}
}
| Field | Type | Description |
|---|---|---|
event | string | Event type: session_started (recognition session established), session_stopped (recognition session ended), canceled (recognition aborted; see message for the reason), reconnecting (connection lost, reconnecting automatically), reconnected (reconnected) |
message | string | Event description message. The wording is not guaranteed to be stable, so always branch on event |
error - Error Event
Triggered when an operation fails or a system anomaly occurs.
{
"type": "error",
"data": {
"error_code": "session_not_started",
"severity": "error",
"message": "Session not started",
"context": "voice-translation",
"request_id": "req_abc123xyz789",
"timestamp": "2026-01-15T10:30:45.123Z"
}
}
| Field | Type | Description |
|---|---|---|
error_code | string | Error code (for programmatic handling) |
severity | string | Severity: fatal / error / warning |
message | string | Human-readable error message |
context | string | Error source category |
request_id | string | Request tracking ID |
timestamp | string | Time the error occurred (ISO 8601) |
Severity Descriptions
| severity | Description | Recommended Action |
|---|---|---|
fatal | Fatal error | Stop the service and require reconnection |
error | Operation failed | Show an error notice and allow retry |
warning | Warning | Show a warning without blocking the operation |
A
warning-level error does not mean the recording has ended; examples arestt_silence_warning(about to end for lack of speech) andbroadcast_standby_warning(the standby phase is about to reach its time limit). The recording continues; when it actually ends you receive the correspondingfatalerror andstatus: "ended".
For the full list of error codes, refer to Error Code Reference.
Version: V1.24.1 Last Updated: 2026-10-07