WebSocket Response Events
Overview
A reference for all response event formats you may receive over WebSocket. For connection and authentication, see Connection and Authentication; for request operations, see Voice Translation Actions.
Table of Contents
- session_started - Session started successfully
- resume_ok - Resume succeeded
- result - Recognition/translation result
- status - Generic status response
- task_complete - Task processing complete
- config_updated - Configuration update complete
- tts_ready - TTS audio ready
- tts_error - TTS synthesis failed
- viewer_count - Viewer count update
- broadcast_phase_changed - Broadcast phase changed
- broadcast_recording_ready - Broadcast recording ready
- speaker_renamed - Speaker renamed
- speaker_reassigned - Speaker identity reassigned
- speakers_merged - Speakers merged
- language_switch_start - Language switch started
- batch_retranslation - Batch retranslation result
- language_switch_done - Language switch complete
- translation_language_removed - Translation language removed
- tts_mode_changed - TTS mode changed
- language_switched - Conversation language switch complete
- tts_updated - Conversation TTS settings updated
- conversation_mode_changed - Conversation mode changed
- speaker_language_changed - Speaker language changed
- speaking_speed_changed - Speaking speed changed
- summary_updated - Summary settings updated
- segment_uploaded - Audio segment upload complete
- stt_event - STT connection status event
- channel_status - Channel status changed
- segment_discarded - Segment Discarded
- viewer_joined - Viewer joined event
- viewer_left - Viewer left event
- error - Error event
- upload_error - Upload error
- speakers_auto_merged - Speakers merged automatically
- summary_done - Summary generation complete
- summary_error - Summary failed or not generated
session_started
Description
After a start action succeeds, the server returns an event containing the complete initial session information. The frontend can use recording_type to distinguish the recording type.
Standard recording (transcribe / conversation / record)
{
"type": "voice-translation",
"data": {
"action": "session_started",
"session_id": "550e8400-e29b-41d4-a716-446655440000",
"task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
"recording_type": "transcribe",
"recognition_mode": "single",
"resume_token": "L0VBAwIy... (43 characters)",
"resume_grace_seconds": 45,
"server_time": 1749550000000,
"message": "Speech recognition started"
}
}
Broadcast mode (broadcast)
{
"type": "voice-translation",
"data": {
"action": "session_started",
"session_id": "550e8400-e29b-41d4-a716-446655440000",
"task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
"recording_type": "broadcast",
"recognition_mode": "multi_speaker",
"phase": "standby",
"viewer_count": 0,
"queue_count": 0,
"peak_viewers": 0,
"total_viewers": 0,
"resume_token": "L0VBAwIy... (43 characters)",
"resume_grace_seconds": 45,
"server_time": 1749550000000,
"message": "Speech recognition started"
}
}
Multi-channel mode (multi_channel)
For a multi-channel (physical channel separation) session, session_started additionally carries channel_mode and channels[] — a snapshot of the channel configuration actually in effect on the server.
{
"type": "voice-translation",
"data": {
"action": "session_started",
"session_id": "550e8400-e29b-41d4-a716-446655440000",
"task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
"recording_type": "transcribe",
"recognition_mode": "multi_channel",
"channel_mode": "per_channel",
"channels": [
{
"channel_id": 1,
"speaker_name": "Host",
"transcription_languages": ["zh-TW"],
"status": "preparing"
},
{
"channel_id": 2,
"speaker_name": "Panelist",
"transcription_languages": ["en-US"],
"status": "preparing"
}
],
"resume_token": "L0VBAwIy... (43 characters)",
"resume_grace_seconds": 45,
"server_time": 1749550000000,
"message": "Speech recognition started"
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
session_id | string | Session ID (WS connection scope; invalid once the connection ends) |
task_id | string | Task ID (the same identifier as REST /api/v1/tasks/{taskId} and Webhook data.task_id) |
recording_type | string | Recording type: transcribe, conversation, record, broadcast |
recognition_mode | string | Recognition mode: single, multi_speaker, multi_language, multi_channel |
channel_mode | string | Multi-channel sub-mode (present only in multi_channel sessions): per_channel (each channel is recognized independently) or shared (channels take turns speaking and share one recognition stream) |
channels | object[] | Snapshot of the channel configuration (present only in multi_channel sessions); each element carries the fields below |
channels[].channel_id | int | Channel number (1–8, unique within the session) |
channels[].speaker_name | string | The speaker name bound to this channel (omitted when not set) |
channels[].transcription_languages | string[] | The transcription language bound to this channel (exactly 1) |
channels[].status | string | Channel status: preparing / ready / removed / error; see channel_status for the state machine |
resume_token | string | Resume token (43 characters). After a disconnect, reconnect with it within the resume_grace_seconds grace window to rejoin the original session; sent ahead of time in session_started, store it until the session ends |
resume_grace_seconds | int | Reconnect grace period in seconds (default 45). This is real wall-clock time and keeps counting down after a disconnect |
server_time | int64 | The server's current time in unix milliseconds. The client can estimate clock skew from (server_time, client time when received) as a reference; however, the grace countdown is still measured against the client's wall clock |
message | string | Status description message |
phase | string | Broadcast phase: standby or live (broadcast mode only) |
viewer_count | int | Current number of online viewers (broadcast mode only) |
queue_count | int | Number of viewers waiting in the queue (broadcast mode only) |
peak_viewers | int | Peak viewer count for this broadcast (broadcast mode only) |
total_viewers | int | Cumulative total of viewers that have ever connected (broadcast mode only) |
ID alignment tip: The WebSocket, REST, and Webhook interfaces all use
task_id(a UUID) as the unified identifier for a task.session_idis a WS connection-scope identifier and is a different concept from a task.
Multi-channel sessions: verify
channel_mode: when a multi-channelstartsucceeds,session_startedalways carrieschannel_modeandchannels[]. If the response is missing these two fields, this connection did not start in multi-channel mode (for example, the service is being updated and the version handling this connection does not yet support multi-channel) — stop immediately and reconnect; do not keep sending audio as if this were a multi-channel session.
resume_ok
Description
The event the server returns when, after a disconnect, you reconnect successfully with resume_token within the resume_grace_seconds grace window. After receiving it, start sending an audio stream again just like after start (for WebM/Opus send a brand-new container; for PCM simply continue sending), and align your local content using server_last_sid / server_last_offset_ms. If is_paused is true, the session was paused before the disconnect, so stay paused after resuming (reopen the audio stream then pause immediately and send no audio) instead of resuming recording on your own. For the full resume handshake flow, see Connection and Authentication.
{
"type": "voice-translation",
"data": {
"action": "resume_ok",
"session_id": "550e8400-e29b-41d4-a716-446655440000",
"task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
"recording_type": "transcribe",
"recognition_mode": "single",
"server_last_sid": 42,
"server_last_offset_ms": 125000,
"server_recording_ms": 127500,
"is_paused": false,
"settings": { ... },
"message": "Resumed the original session"
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
session_id | string | Session ID (WS connection scope; invalid once the connection ends) |
task_id | string | Task ID (same as the original session; the identifier is preserved after resuming) |
recording_type | string | Recording type: transcribe, conversation, record, broadcast |
recognition_mode | string | Recognition mode: single, multi_speaker, multi_language, multi_channel |
server_last_sid | int | The server's current last sentence id (sid). The client should ignore duplicate messages whose sid is ≤ this value |
server_last_offset_ms | int64 | The transcript timeline position (in milliseconds) of the breakpoint, based on the amount of audio already processed (not wall-clock time; the audio timeline is frozen during the disconnect). The client uses it to splice post-reconnect content at the correct timeline position |
server_recording_ms | int64 | The recording-head timestamp (in milliseconds, including silence) that the transcript timeline resumes from after reconnect. The client uses it to align its recording-second header to the same timeline the transcript uses. The difference from server_last_offset_ms is the trailing audio (silence or not-yet-finalized speech) after the last finalized sentence before the disconnect. Omitted when 0 |
is_paused | bool | The paused state the server considers authoritative after resuming: true = it was paused before the disconnect, so the client should stay paused (reopen the audio stream then pause immediately and send no audio); false = recording normally. Used to align the pause UI after a network reconnect or a full-page refresh. Omitted is treated as false (backward compatible with older servers) |
settings | object | The recording settings the session currently holds for this recording (for state reconciliation; added in v1.6.4). See Connection and Authentication for details |
message | string | Status description message (always "Resumed the original session") |
Note: Do not mix the two notions of time. The transcript timeline (
server_last_offset_ms) is based on the audio file and is frozen during a disconnect; the grace reconnect window (resume_grace_seconds) is real wall-clock time and keeps counting down during a disconnect. Deciding "whether reconnect is still possible" must use wall-clock time, never the audio-file time position.
Multi-channel mode (multi_channel):
settingscontainschannel_modeandchannels[](same fields aschannels[]in session_started). Resuming does not re-validate thestartparameters, so use this snapshot to confirm the server is still in multi-channel mode and that each channel's language binding is unchanged, and restore your per-channel status display from each channel'sstatus— channels still preparing after a resume arepreparingand begin producing text within roughly 4 seconds (speech during this window is not lost, only delayed).
result
Description
Speech recognition and translation results. A single result event may contain origin (the recognition result) and/or translations (the translation results).
origin (speech recognition result)
{
"type": "voice-translation",
"data": {
"action": "result",
"origin": {
"sid": 1,
"language": "zh-TW",
"text": "Hello, nice to meet you",
"is_final": true,
"speaker_id": "0",
"detected_language": "zh-TW",
"start_time": "00:05"
}
}
}
origin field descriptions
| Field | Type | Description |
|---|---|---|
sid | int | Sentence number, starting from 1 |
language | string | Source language code. In conversation mode and multi-language transcription this is the language determined by the system — in multi-language mode it is determined per sentence (so it varies within a single recording), and when the determination is unreliable the configured value is kept, so the field is never empty |
text | string | The recognized text |
is_final | boolean | Whether this is the final result |
speaker_id | string | Original speaker ID |
speaker_label | string | (Multi-speaker mode) Display label (after applying the alias; equals speaker_id when no alias exists) |
detected_language | string | The detected language. In conversation mode, determined automatically by the system |
start_time | string | Sentence start time (mm:ss); not sent during the broadcast standby phase, and counted from 00:00 once live begins |
channel_id | int | (Multi-channel mode) The channel number this sentence came from; present only in multi_channel sessions, on every sentence |
Multi-channel mode (multi_channel): in
origin, every sentence carrieschannel_id, andspeaker_idalways has the formatchannel_{N}(Nis the channel number, e.g.channel_3) — speaker identity is determined by the physical channel, stays stable for the whole session, and is never inferred by a model.speaker_labelis the display name: the new name after arename_speaker, otherwise thespeaker_nameconfigured for the channel, and equal tospeaker_idwhen neither is set.
translations (translation result)
{
"type": "voice-translation",
"data": {
"action": "result",
"translations": {
"en-US": {
"sid": 1,
"text": "Hello, nice to meet you",
"is_final": true
}
}
}
}
translations field descriptions
Translation results are keyed by language code; each language's translation object contains:
| Field | Type | Description |
|---|---|---|
sid | int | Sentence number |
text | string | The translated text |
is_final | boolean | Whether this is the final result |
is_retranslation | boolean | Whether this is a retranslation result (only for retranslate) |
speaker_id | string | (Multi-speaker mode) Original speaker ID (aligned with origin since v1.5.3) |
speaker_label | string | (Multi-speaker mode) Display label (after applying the alias; equals speaker_id when no alias exists) |
Multi-language translation (v1.6.7): when multiple languages are specified in
translation_languages, each language returns its own independentresultevent — the samesidreceives Nresultevents, each with a single language key insidetranslations. Multiple languages are never merged into one event. Clients must accumulate translations by "sid+ language code" instead of overwriting; arrival order across languages is not fixed (translations run in parallel).
Multi-channel mode (multi_channel):
translationsdoes not carrychannel_id. Usesidto match a translation back to the same sentence'sorigin, which carries the channel and speaker information.
Important: The success response for the
retranslateaction uses a separateaction: "translation"event (notresult); the payload structure is the same as the table above. Seevoice-translation.mdretranslate success response.
status
Description
A generic status response, used to confirm operations such as pause, resume, stop, set_name, tts_stop, and start_speaking.
{
"type": "voice-translation",
"data": {
"action": "status",
"status": "paused",
"message": "Speech recognition paused"
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
status | string | Machine-readable recording lifecycle state: live (resumed) / paused / ended (stopped). Only pause / resume / stop carry it; set_name etc. do not. ended is always sent before task_complete. |
message | string | Human-readable status text (format not guaranteed; do not parse — rely on the status field) |
Floating subtitle consumers (see Floating Subtitle SSE) should act on
status:paused→ freeze,ended→ close the window,live→ resume.
task_complete
Description
Triggered after stop, once the audio file and transcript have finished uploading. The task_id can be used for subsequent REST API queries about the task details.
- It is always sent after
status: "ended". - The transcript includes the summary, so this event waits until the title and summary have been generated. This usually takes a few seconds to tens of seconds, and longer for a long summary or a slow service, up to about 8 minutes.
- The concurrent recording slot is released when this event is sent: you can start the next recording as soon as you receive it.
- To detect completion, we recommend also supporting the Webhook or a REST query rather than relying on this event alone.
- You also receive this event when a recording ends automatically after a long silence.
{
"type": "voice-translation",
"data": {
"action": "task_complete",
"task_id": "550e8400-e29b-41d4-a716-446655440000",
"message": "Task processing complete"
}
}
Recordings that received no audio at all: whether the recording ended with stop, insufficient credit, or automatically, if no audio was received during the whole recording you still receive status: "ended" and then this event, with noAudio: true. Such a recording has no transcript or audio file and its status is failed; a recording.failed Webhook is also sent (failure_source is no_audio). Do not try to read the transcript; tell the user that no sound was received instead. Minutes already elapsed are billed as usual.
{
"type": "voice-translation",
"data": {
"action": "task_complete",
"task_id": "550e8400-e29b-41d4-a716-446655440000",
"noAudio": true,
"message": "Task processing complete"
}
}
A broadcast that ends during the standby phase (before going live) sends only
status: "ended", not this event: no recording is created during the standby phase.
Field descriptions
| Field | Type | Description |
|---|---|---|
task_id | string | Recording UUID, usable for subsequent API queries |
noAudio | boolean | Present only when no audio was received during the whole recording, and always true; absent when audio was recorded |
message | string | Status description |
config_updated
Description
Configuration update complete event, triggered after a config action succeeds.
Note: Sent only when the configuration is accepted. A rejected configuration returns
type: "error"instead, so your client must listen for that as well or the request appears to go unanswered.
{
"type": "voice-translation",
"data": {
"action": "config_updated",
"updated": ["terminology", "fuzzy_correction", "translation_dict"],
"message": "Settings updated"
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
updated | string[] | The configuration types that were updated: terminology, fuzzy_correction, translation_dict |
message | string | Status message |
terminology_effective | string | (Optional) Appears when terminology is updated during recording; a value of "next_turn" means the new terminology takes effect from the next sentence. Does not appear in the initial config |
unknown_languages | string[] | (Optional) Glossary language codes that could not be recognized (for example zh, chinese). Those entries will not take effect, but the config still succeeds |
inactive_languages | string[] | (Optional) Codes that are valid but are not used by this recording. Does not appear before the recording starts (the language list is not settled yet) |
inactive_dict_languages | string[] | (Optional) The same, for translation-dictionary target languages |
homophone_conflicts | object[] | (Optional) Groups of terms in your glossary that share a pronunciation; each entry carries languages and terms. Not present before the recording has started (the language list is not settled yet) |
Reading
homophone_conflicts: when two terms that sound alike are registered together (公事包 and 公式包, say) and the transcript contains a third spelling with the same pronunciation, the system can only correct it to one of them, and which one is not guaranteed. Both terms themselves still work; the only ambiguity is which term an unregistered homophone misspelling is attributed to.This is a warning, not an error — the glossary is still accepted and
configstill succeeds. If the pair matters to you, list the misspelling explicitly withfuzzy_correctionso it is pinned to the term you want.
languageslists every language that shares the same term index. Chinese regional codes (zh-TW,zh-CN,zh-HKand so on) are treated as one group for glossary matching, so they share a single entry rather than each reporting one."homophone_conflicts": [ { "languages": ["zh-TW"], "terms": ["公事包", "公式包"] } ]
tts_ready
Description
TTS speech synthesis complete event. It contains the audio data and Word Boundary information (which can be used for a karaoke effect).
{
"type": "voice-translation",
"data": {
"action": "tts_ready",
"sid": 1,
"language": "en-US",
"transcript": "Hello, nice to meet you",
"text": "Hello, nice to meet you",
"audio": "Base64EncodedMP3...",
"format": "mp3",
"duration_ms": 2500,
"boundaries": [
{"offset_ms": 0, "duration_ms": 350, "text_offset": 0, "word_length": 5, "text": "Hello", "boundary_type": "WordBoundary"},
{"offset_ms": 350, "duration_ms": 100, "text_offset": 5, "word_length": 1, "text": ",", "boundary_type": "PunctuationBoundary"},
{"offset_ms": 500, "duration_ms": 250, "text_offset": 7, "word_length": 4, "text": "nice", "boundary_type": "WordBoundary"},
{"offset_ms": 750, "duration_ms": 200, "text_offset": 12, "word_length": 2, "text": "to", "boundary_type": "WordBoundary"},
{"offset_ms": 950, "duration_ms": 350, "text_offset": 15, "word_length": 4, "text": "meet", "boundary_type": "WordBoundary"},
{"offset_ms": 1300, "duration_ms": 300, "text_offset": 20, "word_length": 3, "text": "you", "boundary_type": "WordBoundary"}
]
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
sid | int | Sentence number |
language | string | TTS language |
transcript | string | Original transcript (STT recognition result) |
text | string | Translated text (the source for TTS synthesis) |
audio | string | Base64-encoded MP3 audio |
format | string | Audio format (always mp3) |
duration_ms | int | Total audio duration (milliseconds) |
boundaries | array | Word Boundary array |
Word Boundary field descriptions
| Field | Type | Description |
|---|---|---|
offset_ms | int | The word's start time within the audio (milliseconds) |
duration_ms | int | The word's duration (milliseconds) |
text_offset | int | The position within the original string (character index) |
word_length | int | Word length (number of characters) |
text | string | The word content |
boundary_type | string | Boundary type; common values: WordBoundary, PunctuationBoundary, SentenceBoundary, etc. |
tts_error
Description
TTS synthesis failed event.
{
"type": "voice-translation",
"data": {
"action": "tts_error",
"sid": 1,
"language": "en-US",
"error": "translation_not_found",
"message": "No translation available for language: en-US"
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
sid | int | Sentence number |
language | string | TTS language |
error | string | Error code |
message | string | Error message |
transcript | string | The corresponding original transcript, to help the frontend locate the point of failure. This field is always present; it is an empty string "" when the sentence is not found |
TTS error codes
| Error code | Description |
|---|---|
sentence_not_found | The specified sentence was not found (the sid passed to tts_play does not exist) |
translation_not_found | No translation found for this language |
tts_invalid_language | The TTS language is not supported |
tts_invalid_voice | The voice name is invalid |
tts_connection_failed | Could not connect to the speech synthesis service |
tts_timeout | Speech synthesis timed out |
tts_synthesis_failed | Speech synthesis failed |
viewer_count
Broadcast mode only
Description
While a broadcast is in progress, the system checks the viewer count every 3 seconds and pushes this event to the host whenever it changes.
{
"type": "voice-translation",
"data": {
"action": "viewer_count",
"viewer_count": 45,
"queue_count": 8,
"peak_viewers": 50,
"total_viewers": 123
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
viewer_count | int | Current number of online viewers |
queue_count | int | Number of viewers waiting in the queue |
peak_viewers | int | Peak viewer count for this broadcast |
total_viewers | int | Cumulative total of viewers that have ever connected |
Note: This event is pushed only when the viewer count or queue count changes, to avoid unnecessary message transmission.
broadcast_phase_changed
Description
Triggered when the broadcast phase switches from standby to live.
{
"type": "voice-translation",
"data": {
"action": "broadcast_phase_changed",
"phase": "live",
"message": "Broadcast started"
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
phase | string | The new phase: standby or live |
message | string | Status description message |
broadcast_recording_ready
Description
Triggered after a broadcast goes live, returning the finalized task_id for this broadcast.
- When
startgoes live directly (broadcast_phase: "live", the default), this event always arrives aftersession_started. - When the host disconnects and, within the grace period, sends
startagain with the samebroadcast_token(a takeover), this event on the new connection carries a newtask_id. The broadcast therefore has two recordings, one before and one after the takeover, each completed and notified separately.
The
task_idinsession_startedis an initial value for the connection stage, not the final ID for this broadcast. You receive this event after going live regardless of whether the broadcast went through the standby phase first or started directly withbroadcast_phase: "live"(the default). For subsequent operations (for example, exchanging for a Floating Subtitle Feed Token), always use thetask_idfrom this event — using the initial value returns no matching record.
{
"type": "voice-translation",
"data": {
"action": "broadcast_recording_ready",
"task_id": "3f9a1c2e-..."
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
task_id | string | The finalized recording ID for this broadcast (valid after going live) |
speaker_renamed
Description
Global speaker rename complete event.
{
"type": "voice-translation",
"data": {
"action": "speaker_renamed",
"speaker_id": "Guest-1",
"new_label": "Manager Wang",
"affected_sids": [1, 3, 5, 8]
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
speaker_id | string | The resolved original speaker ID (even if the input was a display label, the event returns the original ID) |
new_label | string | The new display label |
affected_sids | int[] | List of affected sentence numbers |
speaker_reassigned
Description
Single-sentence speaker reassignment complete event.
{
"type": "voice-translation",
"data": {
"action": "speaker_reassigned",
"sid": 5,
"old_speaker_id": "Guest-1",
"new_speaker_id": "Guest-2",
"new_speaker_label": "Lisa Lee"
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
sid | int | The sentence number that was changed |
old_speaker_id | string | The original speaker ID |
new_speaker_id | string | The new original speaker ID |
new_speaker_label | string | The new speaker display label (after applying speaker_aliases; equals new_speaker_id when no alias exists) |
speakers_merged
Description
Speaker merge complete event. After the merge, future recognition results produced by the source speaker are also automatically converted to the target speaker. This applies within the current recognition pass only: after a connection recovers, speakers are re-identified and you need to merge again.
{
"type": "voice-translation",
"data": {
"action": "speakers_merged",
"source_speaker_id": "Guest-2",
"target_speaker_id": "Guest-1",
"affected_sids": [3, 5, 7]
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
source_speaker_id | string | The original ID of the merged-away speaker |
target_speaker_id | string | The original ID of the merge target speaker |
affected_sids | number[] | List of affected sentence IDs: sentences that belonged to the source speaker, plus the target speaker's existing sentences whose display name changed because of the merge (for example, when the source speaker's custom name is carried over to the target) |
To obtain the target speaker's display label, query
speaker_aliasesor the nextinit_metadataevent.
language_switch_start
Description
Language switch started event, sent after the switch_language action is triggered.
{
"type": "voice-translation",
"data": {
"action": "language_switch_start",
"translation_language": "ja-JP",
"translation_languages": ["en-US", "ja-JP"],
"total_segments": 15
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
translation_language | string | The single language for this operation (the new target for a replace, or the language added by op:add) |
translation_languages | string[] | Authoritative snapshot of the current full set of translation languages. For a multi-language op:add this is the complete set including existing languages. Clients should overwrite their local language set with this directly, rather than inferring "append" vs "replace" from translation_language (when only 1 language exists, an op:add expanding to a 2nd would be misread as a replace and drop the existing language). |
total_segments | int | The number of sentences to retranslate |
batch_retranslation
Description
Batch retranslation result event, sent sentence by sentence during the language switch process.
{
"type": "voice-translation",
"data": {
"action": "batch_retranslation",
"sid": 3,
"translations": {
"ja-JP": {
"sid": 3,
"text": "今日はプロジェクトの進捗について話し合いましょう",
"is_final": true,
"is_retranslation": true
}
}
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
sid | int | Sentence number |
translations | object | Translation result (same format as the translations in result) |
language_switch_done
Description
Language switch complete event.
{
"type": "voice-translation",
"data": {
"action": "language_switch_done",
"translation_language": "ja-JP",
"translation_languages": ["en-US", "ja-JP"],
"success_count": 15,
"failed_count": 2,
"failed_sids": [3, 7]
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
translation_language | string | The single language for this operation |
translation_languages | string[] | Authoritative snapshot of the current full translation-language set (same as language_switch_start; overwrite the local set). |
success_count | int | The number of sentences successfully translated |
failed_count | int | The number of sentences that failed to translate |
failed_sids | int[] | List of sentence numbers that failed to translate (included only when failed_count > 0) |
translation_language_removed
Description
Translation language removed event (new in v1.6.7). Returned when a multi-language session successfully sends switch_language with op: "remove". Existing translations for that language are kept; subsequent sentences are no longer translated into it.
{
"type": "voice-translation",
"data": {
"action": "translation_language_removed",
"translation_language": "ko-KR",
"translation_languages": ["en-US", "ja-JP"]
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
translation_language | string | The removed translation language |
translation_languages | string[] | Authoritative snapshot of the full translation-language set after removal; overwrite the local set directly. |
tts_mode_changed
Description
TTS playback mode changed event.
{
"type": "voice-translation",
"data": {
"action": "tts_mode_changed",
"tts_mode": "async"
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
tts_mode | string | The new mode: sync or async |
language_switched
Description
Conversation-mode (conversation) language switch complete event. Triggered when switch_language successfully switches the STT source language in conversation mode.
{
"type": "voice-translation",
"data": {
"action": "language_switched",
"language": "en-US",
"translation_language": "zh-TW",
"message": "Language switched"
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
language | string | The new active language (STT source) |
translation_language | string | The new translation target language |
message | string | Status message |
tts_updated
Description
Conversation-mode (conversation) TTS settings updated event. Triggered when set_tts successfully updates the TTS toggle or voice settings.
{
"type": "voice-translation",
"data": {
"action": "tts_updated",
"tts_enabled": true,
"tts_config": {
"zh-TW": { "voice": "zh-TW-HsiaoChenNeural", "speaking_rate": 1.0 },
"en-US": { "voice": "en-US-GuyNeural", "speaking_rate": 1.2 }
}
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
tts_enabled | boolean | Whether TTS is enabled |
tts_config | object | The TTS settings for each language (voice, speaking_rate) |
conversation_mode_changed
Description
Conversation-mode (conversation) mode changed event. Triggered when switch_conversation_mode successfully switches between auto and manual mode.
{
"type": "voice-translation",
"data": {
"action": "conversation_mode_changed",
"conversation_mode": "manual"
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
conversation_mode | string | The new conversation mode: auto or manual |
speaker_language_changed
Description
Conversation-mode (conversation) speaker language changed event. Triggered when set_speaker_language successfully changes a speaker's language; it includes the complete language map after the change.
{
"type": "voice-translation",
"data": {
"action": "speaker_language_changed",
"speaker_language_map": {
"1": "ja-JP",
"2": "en-US"
}
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
speaker_language_map | object | The speaker language map after the change (the key is the speaker number as a string) |
speaking_speed_changed
Description
Speaking speed changed event during recording. Triggered when set_speaking_speed successfully applies the new speed (STT rebuild complete); it returns the applied speed level. Supported in all recognition modes except multi-speaker (multi_speaker) — including broadcast and multi-language LID.
{
"type": "voice-translation",
"data": {
"action": "speaking_speed_changed",
"speaking_speed": "slow"
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
speaking_speed | string | The applied speed: very_slow / slow / normal / fast / very_fast |
summary_updated
Description
Summary settings updated during recording. Emitted after set_summary is applied successfully, carrying the settings currently in effect.
For safety the event omits the full summary_prompt; in custom mode, use summary_prompt_slug to identify it.
{
"type": "voice-translation",
"data": {
"action": "summary_updated",
"summary_mode": "custom",
"summary_prompt_slug": "meeting-actions-v2",
"summary_language": "zh-TW",
"auto_summary": true,
"summary_plain_text": false,
"message": "Summary settings updated"
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
summary_mode | string | Current summary mode: builtin / custom |
summary_template | string | Current shared template identifier (present in builtin mode) |
summary_prompt_slug | string | Current custom prompt identifier (present in custom mode) |
summary_language | string | Current summary output language |
auto_summary | boolean | Whether a summary is generated automatically when recording stops |
summary_plain_text | boolean | Whether the output is plain text |
message | string | Status message |
segment_uploaded
Description
Audio segment upload complete event. Triggered whenever an audio segment is successfully uploaded to cloud storage; it can be used to display upload progress in the frontend.
{
"type": "voice-translation",
"data": {
"action": "segment_uploaded",
"segment_index": 0,
"duration_sec": 30.5
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
segment_index | number | Segment index (starting from 0) |
duration_sec | number | The duration of this segment (seconds) |
stt_event
Description
STT connection status event. Triggered when the connection status of the speech recognition service changes; it can be used to display the STT service status in the frontend.
{
"type": "voice-translation",
"data": {
"action": "stt_event",
"event": "reconnected",
"message": "STT reconnected"
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
event | string | Event type: session_started (recognition session established), session_stopped (recognition session ended), canceled (recognition aborted; see message for the reason), reconnecting (connection lost, reconnecting automatically), reconnected (reconnected) |
message | string | Event description message. The wording is not guaranteed to be stable, so always branch on event |
channel_status
Multi-channel mode (multi_channel) only
Description
Channel status change event. In a multi-channel session, this fires whenever any channel's status changes: the success response of add_channel / remove_channel, settings changes triggered by set_channel_language / set_speaking_speed, resuming after a pause (resume), automatic reconnection after a failure, and when a channel fails and cannot recover automatically. The frontend can use it to display per-channel live status (for example, a loading indicator while a new channel is preparing, or a warning on a failed channel).
{
"type": "voice-translation",
"data": {
"action": "channel_status",
"channel": {
"channel_id": 3,
"speaker_name": "Manager Wang",
"transcription_languages": ["ja-JP"],
"status": "preparing"
},
"active_channels": 3,
"stt_stream_count": 3,
"reason": "added"
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
channel | object | Snapshot of the channel whose status changed; same fields as a channels[] element in session_started (channel_id, speaker_name, transcription_languages, status) |
channel.status | string | The channel's current status: preparing / ready / removed / error (see the state machine below) |
active_channels | int | Number of channels currently capturing audio (excluding removed channels) |
stt_stream_count | int | The number of recognition channels currently counted (billed on this count as each minute begins; drops immediately after remove_channel); always 1 in shared mode |
reason | string | Optional. Why the status changed: added / language_change / reconnect / resumed / removed / speaking_speed / stt_error (see the table below). When omitted, the channel became ready for the first time (see the state machine notes) |
reason values
| reason | When it fires |
|---|---|
added | add_channel added a channel (status becomes preparing) |
language_change | Language change via set_channel_language |
reconnect | Automatic reconnection after the channel failed; recovery after a session resume (resume_token) also uses this reason |
resumed | Resuming after a pause (resume); every channel prepares again |
removed | remove_channel deactivated a channel (status becomes removed) |
speaking_speed | set_speaking_speed applying the new setting channel by channel |
stt_error | The channel failed and cannot recover automatically (status becomes error) |
status state machine
preparing: the channel returns to this state every time it starts or re-prepares recognition — the initial setup atstart, joining viaadd_channel, settings changes fromset_channel_language/set_speaking_speed, automatic reconnection, and resuming after a pause. The channel needs roughly 4 seconds in this state to finish preparing; speech is not lost, only delayed. The one exception is a sentence already in progress at the moment of the switch or reconnect: it cannot be kept, and is reported separately via segment_discarded.ready: entered when the channel receives its first recognition result (starts producing text). Areadyevent triggered by a settings change reuses the samereasonas thepreparingthat started it, so the frontend can match "this ready" back to "that operation"; the first ready afterstartoradd_channelcarries noreason.removed: the channel was deactivated byremove_channel. Transcripts and audio already produced are kept; the channel number cannot be reused (including viaadd_channel).error: the channel failed and cannot recover automatically (reason: "stt_error"). If the channel later recovers on its own, it reportspreparing(reason: "reconnect") →readyagain, or switches straight toreadywhen text production resumes.
Shared mode: only the first channel (the first element of
channels[]) actually performs recognition, and the other channels'preparing/ready/errorfollow the first channel — when the first channel turnsready, all channels turnreadytogether; when the first channel prepares recognition again, all channels return topreparingtogether. Each channel receives its own event, with the samereasonas the first channel. A channel added withadd_channeltakes on the first channel's current status directly;removedfollows each channel's own status. When the first channel turnserror, the other channels turnerroras well (with the samereason); when the first channel recovers, they recover together.
These four status values are consistent with the
channels[].statussnapshot in session_started and thesettings.channels[].statussnapshot in resume_ok — after resuming a session, restore each channel's status display from the snapshot (erroralso appears in snapshots; it is never scrubbed).
Sequence example (add_channel)
→ send add_channel (channel_id: 3, language ja-JP)
← channel_status channel.status: "preparing", reason: "added",
active_channels: 3, stt_stream_count: 3
(roughly 4 seconds of preparation; keep sending channel 3's audio as usual —
nothing is lost, text is only delayed)
← channel_status channel.status: "ready" (no reason)
(the channel produced its first text; from now on this channel's result
events carry channel_id: 3 and speaker_id: "channel_3" in origin)
Sequence example (set_channel_language)
→ send set_channel_language (channel_id: 2, switch to en-US)
← channel_status channel.status: "preparing", reason: "language_change"
(the switch takes effect after roughly 4 seconds; channel_id and speaker_id
are unchanged and the transcript stays continuous)
← channel_status channel.status: "ready", reason: "language_change"
segment_discarded
Description
Segment discarded notification. It tells you that a sid has been discarded and will receive no further events — no is_final: true original text, and no translation for that segment.
Some operations interrupt recognition. If a sentence happens to be in progress at that moment, it cannot be kept. Your client may already have received interim results (is_final: false) for it, so on receiving this event, clear that sid from any "translating" state.
Two-way translation mode (automatic and manual) never receives this event — in automatic mode, a sentence in progress is finalized and sent with is_final: true; in manual mode, if recognition is rebuilt while the user is speaking (for example, after a speaking speed change or a reconnect), that part is merged into the sentence's final result. The content is kept in the transcript in both cases.
{
"type": "voice-translation",
"data": {
"action": "segment_discarded",
"sid": 7,
"reason": "speaking_speed",
"channel_id": 1
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
sid | int | The discarded segment number |
reason | string | Why it was discarded: speaking_speed / language_change / reconnect / resumed / broadcast_go_live (see below) |
channel_id | int | Optional. Source channel number; sent only in multi-channel mode |
reason values
| reason | When it fires |
|---|---|
speaking_speed | set_speaking_speed applying a new speed |
language_change | set_channel_language changing that channel's language |
reconnect | Automatic reconnection after the recognition connection failed |
resumed | Session resume (resume_token), or resuming after a pause (resume) |
broadcast_go_live | broadcast_go_live moving from standby to live; standby segments are not written to the final transcript |
The first four values come from the same set as the reason of channel_status, so you can match both events back to the same operation. The exception is a session resume (resume_token): channel_status uses reconnect, while this event uses resumed.
You may already have received this event even when
set_speaking_speedfails (set_speaking_speed_failed) — recognition is interrupted before the rebuild, so the segment cannot be kept whether the rebuild succeeds or not. A failure response does not mean "nothing happened".
viewer_joined
Description
Viewer joined event (broadcast mode only). When a viewer joins the broadcast, the host receives this event.
{
"type": "voice-translation",
"data": {
"action": "viewer_joined",
"viewer": {
"id": "viewer_abc123",
"ip": "192.168.1.100",
"language": "zh-TW"
},
"viewer_count": 5,
"queue_count": 2
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
viewer | object | Information about the viewer who joined |
viewer.id | string | Viewer ID |
viewer.ip | string | Viewer IP address |
viewer.language | string | The language the viewer selected |
viewer_count | number | Current viewer count |
queue_count | number | Number of viewers in the queue |
viewer_left
Description
Viewer left event (broadcast mode only). When a viewer leaves the broadcast, the host receives this event.
{
"type": "voice-translation",
"data": {
"action": "viewer_left",
"viewer_id": "viewer_abc123",
"viewer_count": 4,
"queue_count": 1
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
viewer_id | string | The ID of the viewer who left |
viewer_count | number | Current viewer count |
queue_count | number | Number of viewers in the queue |
error
Description
Error event. Triggered when an operation fails or a system error occurs.
{
"type": "error",
"data": {
"error_code": "session_not_started",
"severity": "error",
"message": "Session not started",
"context": "voice-translation",
"request_id": "req_abc123xyz789",
"timestamp": "2026-01-15T10:30:45.123Z"
}
}
A sentence-level error (such as a translation failure for one language of a sentence) additionally carries sid and details:
{
"type": "error",
"data": {
"error_code": "llm_content_filtered",
"severity": "warning",
"message": "Content filtered",
"context": "translation",
"sid": 5,
"request_id": "req_abc123xyz789",
"timestamp": "2026-04-26T10:30:45.123Z",
"details": {
"provider": "llm_service",
"source_lang": "zh-TW",
"translation_language": "ja-JP"
}
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
error_code | string | Error code (for programmatic handling) |
severity | string | Severity: fatal / error / warning |
message | string | Human-readable error message |
context | string | Error source category |
sid | int | Optional. The sentence number for a sentence-level error (such as a translation failure for that sentence); not included for non-sentence-level errors |
request_id | string | Request tracing ID |
timestamp | string | The time the error occurred (ISO 8601) |
details | object | Optional. Error context; common keys: provider, translation_language, source_lang. internal_error (a single message failed but the connection is kept) additionally carries message_type (always present) and action (best-effort; absent when parsing fails). See websocket-api.md Per-Message Errors |
Severity descriptions
| severity | Description | Recommended handling |
|---|---|---|
fatal | Fatal error | Stop the service and require reconnection |
error | Operation failed | Show an error prompt and allow retry |
warning | Warning | Show a warning without blocking the operation |
A
warning-level error does not mean the recording has ended; examples arestt_silence_warning(about to end for lack of speech) andbroadcast_standby_warning(the standby phase is about to reach its time limit). The recording continues; when it actually ends you receive the correspondingfatalerror andstatus: "ended".
For the complete list of error codes, see Error Code Reference.
upload_error
v1.5.6 documentation fix: Earlier documentation described a standalone event format of
type: "voice-translation"+action: "upload_error", but in practice this format was never sent on the wire. Storage upload failures always use the unifiederrorenvelope, with one of the three error codes in the table below.If your client listens for
action === "upload_error", switch to listening fortype === "error"and matching onerror_code.
Storage-layer error codes (sent via the error event)
| Error code | Description |
|---|---|
storage_connection_failed | Storage service connection failed |
storage_upload_failed | File upload failed |
storage_queue_full | Upload queue full |
speakers_auto_merged
Sent in multi-speaker mode when the system determines that two speakers are in fact the same person and merges them automatically.
{
"type": "voice-translation",
"data": {
"action": "speakers_auto_merged",
"source_speaker_id": "Guest-2",
"target_speaker_id": "Guest-1"
}
}
| Field | Type | Description |
|---|---|---|
source_speaker_id | string | The speaker that was merged away |
target_speaker_id | string | The speaker that was kept after the merge |
An automatic merge usually happens before any transcript entry has been produced, so the event carries no list of affected sentences. On receiving it, the client merges source_speaker_id into target_speaker_id in its local speaker list.
summary_done
Description
An event pushed after recording stops, once server-side non-streaming summary generation is complete. After receiving this event, the client can call GET /api/v1/sse/history/transcribe/{taskId} to retrieve the summary content (the payload does not include final_content, to avoid bloating the WebSocket message).
v1.5.5 adds two fallback audit fields, summary_fallback_level / summary_dropped_segments: when a custom prompt or transcript content triggers the LLM service content filter, the system automatically regenerates the summary in a reduced mode (standard mode -> neutral mode -> segment-omission mode) and uses these two fields to notify the client of the path actually taken.
Examples
Standard mode succeeds directly (no fallback, no filtering triggered):
{
"type": "voice-translation",
"data": {
"action": "summary_done",
"task_id": "550e8400-e29b-41d4-a716-446655440000",
"summary_id": "sum_a1b2c3d4e5f6g7h8",
"summary_mode": "custom",
"summary_template": "skin-clinic-acme-v2",
"summary_plain_text": true,
"tokens_used": { "input": 1234, "output": 567 }
}
}
Segment-omission mode triggered (summary produced after some transcript segments were omitted):
{
"type": "voice-translation",
"data": {
"action": "summary_done",
"task_id": "550e8400-e29b-41d4-a716-446655440000",
"summary_id": "sum_a1b2c3d4e5f6g7h8",
"summary_mode": "custom",
"summary_template": "skin-clinic-acme-v2",
"summary_plain_text": true,
"tokens_used": { "input": 3456, "output": 789 },
"summary_fallback_level": 3,
"summary_dropped_segments": [3, 7]
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
action | string | Always summary_done |
task_id | string | Recording UUID |
summary_id | string | The internal ID of this summary |
summary_mode | string | "builtin" or "custom" |
summary_template | string | effective slug — builtin → the built-in template slug (such as meeting); custom → the customer slug |
summary_plain_text | boolean | Whether the output is plain text |
tokens_used.input / .output | int | Token usage (a cumulative value across every generation request made for this summary when a fallback is triggered) |
summary_fallback_level | int (omit) | Present only when a fallback was triggered (2 or 3); omitted when standard mode succeeds directly. 2 = neutral mode (regenerated with a neutral instruction); 3 = segment-omission mode (regenerated after the offending segments were omitted) |
summary_dropped_segments | int[] (omit) | Present only when fallback_level=3; the indices of the trimmed transcript segments (in original order) |
Interpreting the fallback level (for frontend UI hints)
summary_fallback_level | Meaning | Suggested UI hint |
|---|---|---|
| (field omitted) | Standard mode succeeds directly, no fallback | Do not show a hint |
2 | The customer prompt triggered filtering, so a neutral fallback prompt was used instead | "Your custom instructions contained terms the content filter could not process; the summary was generated using neutral mode" |
3 | The transcript content triggered filtering; the offending segments were trimmed before producing the summary | "The transcript contained N segments that could not be processed; the summary was generated after omitting the relevant content" (N = summary_dropped_segments.length) |
If segment-omission mode also fails,
summary_doneis not sent; instead,summary_erroris sent witherror_code=llm_content_filtered(see §summary_error below).Note: The payload deliberately does not include
final_content. The client must callGET /api/v1/sse/history/transcribe/{taskId}itself to retrieve the full summary text.summary_fallback_levelandsummary_dropped_segmentsare also provided as top-level fields of theinit_summaryevent during history playback.
summary_error
Description
An event pushed when summary generation fails, or when no summary is generated because the available credits are insufficient, so the client does not need to keep polling to find out.
Example
{
"type": "voice-translation",
"data": {
"action": "summary_error",
"task_id": "550e8400-e29b-41d4-a716-446655440000",
"error_code": "summary_failed",
"message": "Summary generation failed"
}
}
Field descriptions
| Field | Type | Description |
|---|---|---|
action | string | Always summary_error |
task_id | string | Recording UUID |
error_code | string | Summary error code (such as summary_failed / summary_timeout / summary_mode_field_mismatch / summary_insufficient_credit, etc.) |
message | string | Human-readable error message (already sanitized; does not include the LLM raw error) |
When
error_codeissummary_insufficient_credit, the available credits were insufficient (the recording ended because the credits ran out, or the credits at the end could not cover the summary fee); the transcript and audio are still saved, and after topping up you can get a summary through Regenerate Summary.
Version: V1.24.1 Last Updated: 2026-10-07