WebSocket API

WebSocket Response Events

Overview

A reference for all response event formats you may receive over WebSocket. For connection and authentication, see Connection and Authentication; for request operations, see Voice Translation Actions.


Table of Contents

  1. session_started - Session started successfully
  2. resume_ok - Resume succeeded
  3. result - Recognition/translation result
  4. status - Generic status response
  5. task_complete - Task processing complete
  6. config_updated - Configuration update complete
  7. tts_ready - TTS audio ready
  8. tts_error - TTS synthesis failed
  9. viewer_count - Viewer count update
  10. broadcast_phase_changed - Broadcast phase changed
  11. broadcast_recording_ready - Broadcast recording ready
  12. speaker_renamed - Speaker renamed
  13. speaker_reassigned - Speaker identity reassigned
  14. speakers_merged - Speakers merged
  15. language_switch_start - Language switch started
  16. batch_retranslation - Batch retranslation result
  17. language_switch_done - Language switch complete
  18. translation_language_removed - Translation language removed
  19. tts_mode_changed - TTS mode changed
  20. language_switched - Conversation language switch complete
  21. tts_updated - Conversation TTS settings updated
  22. conversation_mode_changed - Conversation mode changed
  23. speaker_language_changed - Speaker language changed
  24. speaking_speed_changed - Speaking speed changed
  25. summary_updated - Summary settings updated
  26. segment_uploaded - Audio segment upload complete
  27. stt_event - STT connection status event
  28. channel_status - Channel status changed
  29. segment_discarded - Segment Discarded
  30. viewer_joined - Viewer joined event
  31. viewer_left - Viewer left event
  32. error - Error event
  33. upload_error - Upload error
  34. speakers_auto_merged - Speakers merged automatically
  35. summary_done - Summary generation complete
  36. summary_error - Summary failed or not generated

session_started

Description

After a start action succeeds, the server returns an event containing the complete initial session information. The frontend can use recording_type to distinguish the recording type.

Standard recording (transcribe / conversation / record)

{
  "type": "voice-translation",
  "data": {
    "action": "session_started",
    "session_id": "550e8400-e29b-41d4-a716-446655440000",
    "task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
    "recording_type": "transcribe",
    "recognition_mode": "single",
    "resume_token": "L0VBAwIy... (43 characters)",
    "resume_grace_seconds": 45,
    "server_time": 1749550000000,
    "message": "Speech recognition started"
  }
}

Broadcast mode (broadcast)

{
  "type": "voice-translation",
  "data": {
    "action": "session_started",
    "session_id": "550e8400-e29b-41d4-a716-446655440000",
    "task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
    "recording_type": "broadcast",
    "recognition_mode": "multi_speaker",
    "phase": "standby",
    "viewer_count": 0,
    "queue_count": 0,
    "peak_viewers": 0,
    "total_viewers": 0,
    "resume_token": "L0VBAwIy... (43 characters)",
    "resume_grace_seconds": 45,
    "server_time": 1749550000000,
    "message": "Speech recognition started"
  }
}

Multi-channel mode (multi_channel)

For a multi-channel (physical channel separation) session, session_started additionally carries channel_mode and channels[] — a snapshot of the channel configuration actually in effect on the server.

{
  "type": "voice-translation",
  "data": {
    "action": "session_started",
    "session_id": "550e8400-e29b-41d4-a716-446655440000",
    "task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
    "recording_type": "transcribe",
    "recognition_mode": "multi_channel",
    "channel_mode": "per_channel",
    "channels": [
      {
        "channel_id": 1,
        "speaker_name": "Host",
        "transcription_languages": ["zh-TW"],
        "status": "preparing"
      },
      {
        "channel_id": 2,
        "speaker_name": "Panelist",
        "transcription_languages": ["en-US"],
        "status": "preparing"
      }
    ],
    "resume_token": "L0VBAwIy... (43 characters)",
    "resume_grace_seconds": 45,
    "server_time": 1749550000000,
    "message": "Speech recognition started"
  }
}

Field descriptions

FieldTypeDescription
session_idstringSession ID (WS connection scope; invalid once the connection ends)
task_idstringTask ID (the same identifier as REST /api/v1/tasks/{taskId} and Webhook data.task_id)
recording_typestringRecording type: transcribe, conversation, record, broadcast
recognition_modestringRecognition mode: single, multi_speaker, multi_language, multi_channel
channel_modestringMulti-channel sub-mode (present only in multi_channel sessions): per_channel (each channel is recognized independently) or shared (channels take turns speaking and share one recognition stream)
channelsobject[]Snapshot of the channel configuration (present only in multi_channel sessions); each element carries the fields below
channels[].channel_idintChannel number (1–8, unique within the session)
channels[].speaker_namestringThe speaker name bound to this channel (omitted when not set)
channels[].transcription_languagesstring[]The transcription language bound to this channel (exactly 1)
channels[].statusstringChannel status: preparing / ready / removed / error; see channel_status for the state machine
resume_tokenstringResume token (43 characters). After a disconnect, reconnect with it within the resume_grace_seconds grace window to rejoin the original session; sent ahead of time in session_started, store it until the session ends
resume_grace_secondsintReconnect grace period in seconds (default 45). This is real wall-clock time and keeps counting down after a disconnect
server_timeint64The server's current time in unix milliseconds. The client can estimate clock skew from (server_time, client time when received) as a reference; however, the grace countdown is still measured against the client's wall clock
messagestringStatus description message
phasestringBroadcast phase: standby or live (broadcast mode only)
viewer_countintCurrent number of online viewers (broadcast mode only)
queue_countintNumber of viewers waiting in the queue (broadcast mode only)
peak_viewersintPeak viewer count for this broadcast (broadcast mode only)
total_viewersintCumulative total of viewers that have ever connected (broadcast mode only)

ID alignment tip: The WebSocket, REST, and Webhook interfaces all use task_id (a UUID) as the unified identifier for a task. session_id is a WS connection-scope identifier and is a different concept from a task.

Multi-channel sessions: verify channel_mode: when a multi-channel start succeeds, session_started always carries channel_mode and channels[]. If the response is missing these two fields, this connection did not start in multi-channel mode (for example, the service is being updated and the version handling this connection does not yet support multi-channel) — stop immediately and reconnect; do not keep sending audio as if this were a multi-channel session.


resume_ok

Description

The event the server returns when, after a disconnect, you reconnect successfully with resume_token within the resume_grace_seconds grace window. After receiving it, start sending an audio stream again just like after start (for WebM/Opus send a brand-new container; for PCM simply continue sending), and align your local content using server_last_sid / server_last_offset_ms. If is_paused is true, the session was paused before the disconnect, so stay paused after resuming (reopen the audio stream then pause immediately and send no audio) instead of resuming recording on your own. For the full resume handshake flow, see Connection and Authentication.

{
  "type": "voice-translation",
  "data": {
    "action": "resume_ok",
    "session_id": "550e8400-e29b-41d4-a716-446655440000",
    "task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
    "recording_type": "transcribe",
    "recognition_mode": "single",
    "server_last_sid": 42,
    "server_last_offset_ms": 125000,
    "server_recording_ms": 127500,
    "is_paused": false,
    "settings": { ... },
    "message": "Resumed the original session"
  }
}

Field descriptions

FieldTypeDescription
session_idstringSession ID (WS connection scope; invalid once the connection ends)
task_idstringTask ID (same as the original session; the identifier is preserved after resuming)
recording_typestringRecording type: transcribe, conversation, record, broadcast
recognition_modestringRecognition mode: single, multi_speaker, multi_language, multi_channel
server_last_sidintThe server's current last sentence id (sid). The client should ignore duplicate messages whose sid is ≤ this value
server_last_offset_msint64The transcript timeline position (in milliseconds) of the breakpoint, based on the amount of audio already processed (not wall-clock time; the audio timeline is frozen during the disconnect). The client uses it to splice post-reconnect content at the correct timeline position
server_recording_msint64The recording-head timestamp (in milliseconds, including silence) that the transcript timeline resumes from after reconnect. The client uses it to align its recording-second header to the same timeline the transcript uses. The difference from server_last_offset_ms is the trailing audio (silence or not-yet-finalized speech) after the last finalized sentence before the disconnect. Omitted when 0
is_pausedboolThe paused state the server considers authoritative after resuming: true = it was paused before the disconnect, so the client should stay paused (reopen the audio stream then pause immediately and send no audio); false = recording normally. Used to align the pause UI after a network reconnect or a full-page refresh. Omitted is treated as false (backward compatible with older servers)
settingsobjectThe recording settings the session currently holds for this recording (for state reconciliation; added in v1.6.4). See Connection and Authentication for details
messagestringStatus description message (always "Resumed the original session")

Note: Do not mix the two notions of time. The transcript timeline (server_last_offset_ms) is based on the audio file and is frozen during a disconnect; the grace reconnect window (resume_grace_seconds) is real wall-clock time and keeps counting down during a disconnect. Deciding "whether reconnect is still possible" must use wall-clock time, never the audio-file time position.

Multi-channel mode (multi_channel): settings contains channel_mode and channels[] (same fields as channels[] in session_started). Resuming does not re-validate the start parameters, so use this snapshot to confirm the server is still in multi-channel mode and that each channel's language binding is unchanged, and restore your per-channel status display from each channel's status — channels still preparing after a resume are preparing and begin producing text within roughly 4 seconds (speech during this window is not lost, only delayed).


result

Description

Speech recognition and translation results. A single result event may contain origin (the recognition result) and/or translations (the translation results).

origin (speech recognition result)

{
  "type": "voice-translation",
  "data": {
    "action": "result",
    "origin": {
      "sid": 1,
      "language": "zh-TW",
      "text": "Hello, nice to meet you",
      "is_final": true,
      "speaker_id": "0",
      "detected_language": "zh-TW",
      "start_time": "00:05"
    }
  }
}

origin field descriptions

FieldTypeDescription
sidintSentence number, starting from 1
languagestringSource language code. In conversation mode and multi-language transcription this is the language determined by the system — in multi-language mode it is determined per sentence (so it varies within a single recording), and when the determination is unreliable the configured value is kept, so the field is never empty
textstringThe recognized text
is_finalbooleanWhether this is the final result
speaker_idstringOriginal speaker ID
speaker_labelstring(Multi-speaker mode) Display label (after applying the alias; equals speaker_id when no alias exists)
detected_languagestringThe detected language. In conversation mode, determined automatically by the system
start_timestringSentence start time (mm:ss); not sent during the broadcast standby phase, and counted from 00:00 once live begins
channel_idint(Multi-channel mode) The channel number this sentence came from; present only in multi_channel sessions, on every sentence

Multi-channel mode (multi_channel): in origin, every sentence carries channel_id, and speaker_id always has the format channel_{N} (N is the channel number, e.g. channel_3) — speaker identity is determined by the physical channel, stays stable for the whole session, and is never inferred by a model. speaker_label is the display name: the new name after a rename_speaker, otherwise the speaker_name configured for the channel, and equal to speaker_id when neither is set.

translations (translation result)

{
  "type": "voice-translation",
  "data": {
    "action": "result",
    "translations": {
      "en-US": {
        "sid": 1,
        "text": "Hello, nice to meet you",
        "is_final": true
      }
    }
  }
}

translations field descriptions

Translation results are keyed by language code; each language's translation object contains:

FieldTypeDescription
sidintSentence number
textstringThe translated text
is_finalbooleanWhether this is the final result
is_retranslationbooleanWhether this is a retranslation result (only for retranslate)
speaker_idstring(Multi-speaker mode) Original speaker ID (aligned with origin since v1.5.3)
speaker_labelstring(Multi-speaker mode) Display label (after applying the alias; equals speaker_id when no alias exists)

Multi-language translation (v1.6.7): when multiple languages are specified in translation_languages, each language returns its own independent result event — the same sid receives N result events, each with a single language key inside translations. Multiple languages are never merged into one event. Clients must accumulate translations by "sid + language code" instead of overwriting; arrival order across languages is not fixed (translations run in parallel).

Multi-channel mode (multi_channel): translations does not carry channel_id. Use sid to match a translation back to the same sentence's origin, which carries the channel and speaker information.

Important: The success response for the retranslate action uses a separate action: "translation" event (not result); the payload structure is the same as the table above. See voice-translation.md retranslate success response.


status

Description

A generic status response, used to confirm operations such as pause, resume, stop, set_name, tts_stop, and start_speaking.

{
  "type": "voice-translation",
  "data": {
    "action": "status",
    "status": "paused",
    "message": "Speech recognition paused"
  }
}

Field descriptions

FieldTypeDescription
statusstringMachine-readable recording lifecycle state: live (resumed) / paused / ended (stopped). Only pause / resume / stop carry it; set_name etc. do not. ended is always sent before task_complete.
messagestringHuman-readable status text (format not guaranteed; do not parse — rely on the status field)

Floating subtitle consumers (see Floating Subtitle SSE) should act on status: paused → freeze, ended → close the window, live → resume.


task_complete

Description

Triggered after stop, once the audio file and transcript have finished uploading. The task_id can be used for subsequent REST API queries about the task details.

  • It is always sent after status: "ended".
  • The transcript includes the summary, so this event waits until the title and summary have been generated. This usually takes a few seconds to tens of seconds, and longer for a long summary or a slow service, up to about 8 minutes.
  • The concurrent recording slot is released when this event is sent: you can start the next recording as soon as you receive it.
  • To detect completion, we recommend also supporting the Webhook or a REST query rather than relying on this event alone.
  • You also receive this event when a recording ends automatically after a long silence.
{
  "type": "voice-translation",
  "data": {
    "action": "task_complete",
    "task_id": "550e8400-e29b-41d4-a716-446655440000",
    "message": "Task processing complete"
  }
}

Recordings that received no audio at all: whether the recording ended with stop, insufficient credit, or automatically, if no audio was received during the whole recording you still receive status: "ended" and then this event, with noAudio: true. Such a recording has no transcript or audio file and its status is failed; a recording.failed Webhook is also sent (failure_source is no_audio). Do not try to read the transcript; tell the user that no sound was received instead. Minutes already elapsed are billed as usual.

{
  "type": "voice-translation",
  "data": {
    "action": "task_complete",
    "task_id": "550e8400-e29b-41d4-a716-446655440000",
    "noAudio": true,
    "message": "Task processing complete"
  }
}

A broadcast that ends during the standby phase (before going live) sends only status: "ended", not this event: no recording is created during the standby phase.

Field descriptions

FieldTypeDescription
task_idstringRecording UUID, usable for subsequent API queries
noAudiobooleanPresent only when no audio was received during the whole recording, and always true; absent when audio was recorded
messagestringStatus description

config_updated

Description

Configuration update complete event, triggered after a config action succeeds.

Note: Sent only when the configuration is accepted. A rejected configuration returns type: "error" instead, so your client must listen for that as well or the request appears to go unanswered.

{
  "type": "voice-translation",
  "data": {
    "action": "config_updated",
    "updated": ["terminology", "fuzzy_correction", "translation_dict"],
    "message": "Settings updated"
  }
}

Field descriptions

FieldTypeDescription
updatedstring[]The configuration types that were updated: terminology, fuzzy_correction, translation_dict
messagestringStatus message
terminology_effectivestring(Optional) Appears when terminology is updated during recording; a value of "next_turn" means the new terminology takes effect from the next sentence. Does not appear in the initial config
unknown_languagesstring[](Optional) Glossary language codes that could not be recognized (for example zh, chinese). Those entries will not take effect, but the config still succeeds
inactive_languagesstring[](Optional) Codes that are valid but are not used by this recording. Does not appear before the recording starts (the language list is not settled yet)
inactive_dict_languagesstring[](Optional) The same, for translation-dictionary target languages
homophone_conflictsobject[](Optional) Groups of terms in your glossary that share a pronunciation; each entry carries languages and terms. Not present before the recording has started (the language list is not settled yet)

Reading homophone_conflicts: when two terms that sound alike are registered together (公事包 and 公式包, say) and the transcript contains a third spelling with the same pronunciation, the system can only correct it to one of them, and which one is not guaranteed. Both terms themselves still work; the only ambiguity is which term an unregistered homophone misspelling is attributed to.

This is a warning, not an error — the glossary is still accepted and config still succeeds. If the pair matters to you, list the misspelling explicitly with fuzzy_correction so it is pinned to the term you want.

languages lists every language that shares the same term index. Chinese regional codes (zh-TW, zh-CN, zh-HK and so on) are treated as one group for glossary matching, so they share a single entry rather than each reporting one.

"homophone_conflicts": [
  { "languages": ["zh-TW"], "terms": ["公事包", "公式包"] }
]

tts_ready

Description

TTS speech synthesis complete event. It contains the audio data and Word Boundary information (which can be used for a karaoke effect).

{
  "type": "voice-translation",
  "data": {
    "action": "tts_ready",
    "sid": 1,
    "language": "en-US",
    "transcript": "Hello, nice to meet you",
    "text": "Hello, nice to meet you",
    "audio": "Base64EncodedMP3...",
    "format": "mp3",
    "duration_ms": 2500,
    "boundaries": [
      {"offset_ms": 0, "duration_ms": 350, "text_offset": 0, "word_length": 5, "text": "Hello", "boundary_type": "WordBoundary"},
      {"offset_ms": 350, "duration_ms": 100, "text_offset": 5, "word_length": 1, "text": ",", "boundary_type": "PunctuationBoundary"},
      {"offset_ms": 500, "duration_ms": 250, "text_offset": 7, "word_length": 4, "text": "nice", "boundary_type": "WordBoundary"},
      {"offset_ms": 750, "duration_ms": 200, "text_offset": 12, "word_length": 2, "text": "to", "boundary_type": "WordBoundary"},
      {"offset_ms": 950, "duration_ms": 350, "text_offset": 15, "word_length": 4, "text": "meet", "boundary_type": "WordBoundary"},
      {"offset_ms": 1300, "duration_ms": 300, "text_offset": 20, "word_length": 3, "text": "you", "boundary_type": "WordBoundary"}
    ]
  }
}

Field descriptions

FieldTypeDescription
sidintSentence number
languagestringTTS language
transcriptstringOriginal transcript (STT recognition result)
textstringTranslated text (the source for TTS synthesis)
audiostringBase64-encoded MP3 audio
formatstringAudio format (always mp3)
duration_msintTotal audio duration (milliseconds)
boundariesarrayWord Boundary array

Word Boundary field descriptions

FieldTypeDescription
offset_msintThe word's start time within the audio (milliseconds)
duration_msintThe word's duration (milliseconds)
text_offsetintThe position within the original string (character index)
word_lengthintWord length (number of characters)
textstringThe word content
boundary_typestringBoundary type; common values: WordBoundary, PunctuationBoundary, SentenceBoundary, etc.

tts_error

Description

TTS synthesis failed event.

{
  "type": "voice-translation",
  "data": {
    "action": "tts_error",
    "sid": 1,
    "language": "en-US",
    "error": "translation_not_found",
    "message": "No translation available for language: en-US"
  }
}

Field descriptions

FieldTypeDescription
sidintSentence number
languagestringTTS language
errorstringError code
messagestringError message
transcriptstringThe corresponding original transcript, to help the frontend locate the point of failure. This field is always present; it is an empty string "" when the sentence is not found

TTS error codes

Error codeDescription
sentence_not_foundThe specified sentence was not found (the sid passed to tts_play does not exist)
translation_not_foundNo translation found for this language
tts_invalid_languageThe TTS language is not supported
tts_invalid_voiceThe voice name is invalid
tts_connection_failedCould not connect to the speech synthesis service
tts_timeoutSpeech synthesis timed out
tts_synthesis_failedSpeech synthesis failed

viewer_count

Broadcast mode only

Description

While a broadcast is in progress, the system checks the viewer count every 3 seconds and pushes this event to the host whenever it changes.

{
  "type": "voice-translation",
  "data": {
    "action": "viewer_count",
    "viewer_count": 45,
    "queue_count": 8,
    "peak_viewers": 50,
    "total_viewers": 123
  }
}

Field descriptions

FieldTypeDescription
viewer_countintCurrent number of online viewers
queue_countintNumber of viewers waiting in the queue
peak_viewersintPeak viewer count for this broadcast
total_viewersintCumulative total of viewers that have ever connected

Note: This event is pushed only when the viewer count or queue count changes, to avoid unnecessary message transmission.


broadcast_phase_changed

Description

Triggered when the broadcast phase switches from standby to live.

{
  "type": "voice-translation",
  "data": {
    "action": "broadcast_phase_changed",
    "phase": "live",
    "message": "Broadcast started"
  }
}

Field descriptions

FieldTypeDescription
phasestringThe new phase: standby or live
messagestringStatus description message

broadcast_recording_ready

Description

Triggered after a broadcast goes live, returning the finalized task_id for this broadcast.

  • When start goes live directly (broadcast_phase: "live", the default), this event always arrives after session_started.
  • When the host disconnects and, within the grace period, sends start again with the same broadcast_token (a takeover), this event on the new connection carries a new task_id. The broadcast therefore has two recordings, one before and one after the takeover, each completed and notified separately.

The task_id in session_started is an initial value for the connection stage, not the final ID for this broadcast. You receive this event after going live regardless of whether the broadcast went through the standby phase first or started directly with broadcast_phase: "live" (the default). For subsequent operations (for example, exchanging for a Floating Subtitle Feed Token), always use the task_id from this event — using the initial value returns no matching record.

{
  "type": "voice-translation",
  "data": {
    "action": "broadcast_recording_ready",
    "task_id": "3f9a1c2e-..."
  }
}

Field descriptions

FieldTypeDescription
task_idstringThe finalized recording ID for this broadcast (valid after going live)

speaker_renamed

Description

Global speaker rename complete event.

{
  "type": "voice-translation",
  "data": {
    "action": "speaker_renamed",
    "speaker_id": "Guest-1",
    "new_label": "Manager Wang",
    "affected_sids": [1, 3, 5, 8]
  }
}

Field descriptions

FieldTypeDescription
speaker_idstringThe resolved original speaker ID (even if the input was a display label, the event returns the original ID)
new_labelstringThe new display label
affected_sidsint[]List of affected sentence numbers

speaker_reassigned

Description

Single-sentence speaker reassignment complete event.

{
  "type": "voice-translation",
  "data": {
    "action": "speaker_reassigned",
    "sid": 5,
    "old_speaker_id": "Guest-1",
    "new_speaker_id": "Guest-2",
    "new_speaker_label": "Lisa Lee"
  }
}

Field descriptions

FieldTypeDescription
sidintThe sentence number that was changed
old_speaker_idstringThe original speaker ID
new_speaker_idstringThe new original speaker ID
new_speaker_labelstringThe new speaker display label (after applying speaker_aliases; equals new_speaker_id when no alias exists)

speakers_merged

Description

Speaker merge complete event. After the merge, future recognition results produced by the source speaker are also automatically converted to the target speaker. This applies within the current recognition pass only: after a connection recovers, speakers are re-identified and you need to merge again.

{
  "type": "voice-translation",
  "data": {
    "action": "speakers_merged",
    "source_speaker_id": "Guest-2",
    "target_speaker_id": "Guest-1",
    "affected_sids": [3, 5, 7]
  }
}

Field descriptions

FieldTypeDescription
source_speaker_idstringThe original ID of the merged-away speaker
target_speaker_idstringThe original ID of the merge target speaker
affected_sidsnumber[]List of affected sentence IDs: sentences that belonged to the source speaker, plus the target speaker's existing sentences whose display name changed because of the merge (for example, when the source speaker's custom name is carried over to the target)

To obtain the target speaker's display label, query speaker_aliases or the next init_metadata event.


language_switch_start

Description

Language switch started event, sent after the switch_language action is triggered.

{
  "type": "voice-translation",
  "data": {
    "action": "language_switch_start",
    "translation_language": "ja-JP",
    "translation_languages": ["en-US", "ja-JP"],
    "total_segments": 15
  }
}

Field descriptions

FieldTypeDescription
translation_languagestringThe single language for this operation (the new target for a replace, or the language added by op:add)
translation_languagesstring[]Authoritative snapshot of the current full set of translation languages. For a multi-language op:add this is the complete set including existing languages. Clients should overwrite their local language set with this directly, rather than inferring "append" vs "replace" from translation_language (when only 1 language exists, an op:add expanding to a 2nd would be misread as a replace and drop the existing language).
total_segmentsintThe number of sentences to retranslate

batch_retranslation

Description

Batch retranslation result event, sent sentence by sentence during the language switch process.

{
  "type": "voice-translation",
  "data": {
    "action": "batch_retranslation",
    "sid": 3,
    "translations": {
      "ja-JP": {
        "sid": 3,
        "text": "今日はプロジェクトの進捗について話し合いましょう",
        "is_final": true,
        "is_retranslation": true
      }
    }
  }
}

Field descriptions

FieldTypeDescription
sidintSentence number
translationsobjectTranslation result (same format as the translations in result)

language_switch_done

Description

Language switch complete event.

{
  "type": "voice-translation",
  "data": {
    "action": "language_switch_done",
    "translation_language": "ja-JP",
    "translation_languages": ["en-US", "ja-JP"],
    "success_count": 15,
    "failed_count": 2,
    "failed_sids": [3, 7]
  }
}

Field descriptions

FieldTypeDescription
translation_languagestringThe single language for this operation
translation_languagesstring[]Authoritative snapshot of the current full translation-language set (same as language_switch_start; overwrite the local set).
success_countintThe number of sentences successfully translated
failed_countintThe number of sentences that failed to translate
failed_sidsint[]List of sentence numbers that failed to translate (included only when failed_count > 0)

translation_language_removed

Description

Translation language removed event (new in v1.6.7). Returned when a multi-language session successfully sends switch_language with op: "remove". Existing translations for that language are kept; subsequent sentences are no longer translated into it.

{
  "type": "voice-translation",
  "data": {
    "action": "translation_language_removed",
    "translation_language": "ko-KR",
    "translation_languages": ["en-US", "ja-JP"]
  }
}

Field descriptions

FieldTypeDescription
translation_languagestringThe removed translation language
translation_languagesstring[]Authoritative snapshot of the full translation-language set after removal; overwrite the local set directly.

tts_mode_changed

Description

TTS playback mode changed event.

{
  "type": "voice-translation",
  "data": {
    "action": "tts_mode_changed",
    "tts_mode": "async"
  }
}

Field descriptions

FieldTypeDescription
tts_modestringThe new mode: sync or async

language_switched

Description

Conversation-mode (conversation) language switch complete event. Triggered when switch_language successfully switches the STT source language in conversation mode.

{
  "type": "voice-translation",
  "data": {
    "action": "language_switched",
    "language": "en-US",
    "translation_language": "zh-TW",
    "message": "Language switched"
  }
}

Field descriptions

FieldTypeDescription
languagestringThe new active language (STT source)
translation_languagestringThe new translation target language
messagestringStatus message

tts_updated

Description

Conversation-mode (conversation) TTS settings updated event. Triggered when set_tts successfully updates the TTS toggle or voice settings.

{
  "type": "voice-translation",
  "data": {
    "action": "tts_updated",
    "tts_enabled": true,
    "tts_config": {
      "zh-TW": { "voice": "zh-TW-HsiaoChenNeural", "speaking_rate": 1.0 },
      "en-US": { "voice": "en-US-GuyNeural", "speaking_rate": 1.2 }
    }
  }
}

Field descriptions

FieldTypeDescription
tts_enabledbooleanWhether TTS is enabled
tts_configobjectThe TTS settings for each language (voice, speaking_rate)

conversation_mode_changed

Description

Conversation-mode (conversation) mode changed event. Triggered when switch_conversation_mode successfully switches between auto and manual mode.

{
  "type": "voice-translation",
  "data": {
    "action": "conversation_mode_changed",
    "conversation_mode": "manual"
  }
}

Field descriptions

FieldTypeDescription
conversation_modestringThe new conversation mode: auto or manual

speaker_language_changed

Description

Conversation-mode (conversation) speaker language changed event. Triggered when set_speaker_language successfully changes a speaker's language; it includes the complete language map after the change.

{
  "type": "voice-translation",
  "data": {
    "action": "speaker_language_changed",
    "speaker_language_map": {
      "1": "ja-JP",
      "2": "en-US"
    }
  }
}

Field descriptions

FieldTypeDescription
speaker_language_mapobjectThe speaker language map after the change (the key is the speaker number as a string)

speaking_speed_changed

Description

Speaking speed changed event during recording. Triggered when set_speaking_speed successfully applies the new speed (STT rebuild complete); it returns the applied speed level. Supported in all recognition modes except multi-speaker (multi_speaker) — including broadcast and multi-language LID.

{
  "type": "voice-translation",
  "data": {
    "action": "speaking_speed_changed",
    "speaking_speed": "slow"
  }
}

Field descriptions

FieldTypeDescription
speaking_speedstringThe applied speed: very_slow / slow / normal / fast / very_fast

summary_updated

Description

Summary settings updated during recording. Emitted after set_summary is applied successfully, carrying the settings currently in effect.

For safety the event omits the full summary_prompt; in custom mode, use summary_prompt_slug to identify it.

{
  "type": "voice-translation",
  "data": {
    "action": "summary_updated",
    "summary_mode": "custom",
    "summary_prompt_slug": "meeting-actions-v2",
    "summary_language": "zh-TW",
    "auto_summary": true,
    "summary_plain_text": false,
    "message": "Summary settings updated"
  }
}

Field descriptions

FieldTypeDescription
summary_modestringCurrent summary mode: builtin / custom
summary_templatestringCurrent shared template identifier (present in builtin mode)
summary_prompt_slugstringCurrent custom prompt identifier (present in custom mode)
summary_languagestringCurrent summary output language
auto_summarybooleanWhether a summary is generated automatically when recording stops
summary_plain_textbooleanWhether the output is plain text
messagestringStatus message

segment_uploaded

Description

Audio segment upload complete event. Triggered whenever an audio segment is successfully uploaded to cloud storage; it can be used to display upload progress in the frontend.

{
  "type": "voice-translation",
  "data": {
    "action": "segment_uploaded",
    "segment_index": 0,
    "duration_sec": 30.5
  }
}

Field descriptions

FieldTypeDescription
segment_indexnumberSegment index (starting from 0)
duration_secnumberThe duration of this segment (seconds)

stt_event

Description

STT connection status event. Triggered when the connection status of the speech recognition service changes; it can be used to display the STT service status in the frontend.

{
  "type": "voice-translation",
  "data": {
    "action": "stt_event",
    "event": "reconnected",
    "message": "STT reconnected"
  }
}

Field descriptions

FieldTypeDescription
eventstringEvent type: session_started (recognition session established), session_stopped (recognition session ended), canceled (recognition aborted; see message for the reason), reconnecting (connection lost, reconnecting automatically), reconnected (reconnected)
messagestringEvent description message. The wording is not guaranteed to be stable, so always branch on event

channel_status

Multi-channel mode (multi_channel) only

Description

Channel status change event. In a multi-channel session, this fires whenever any channel's status changes: the success response of add_channel / remove_channel, settings changes triggered by set_channel_language / set_speaking_speed, resuming after a pause (resume), automatic reconnection after a failure, and when a channel fails and cannot recover automatically. The frontend can use it to display per-channel live status (for example, a loading indicator while a new channel is preparing, or a warning on a failed channel).

{
  "type": "voice-translation",
  "data": {
    "action": "channel_status",
    "channel": {
      "channel_id": 3,
      "speaker_name": "Manager Wang",
      "transcription_languages": ["ja-JP"],
      "status": "preparing"
    },
    "active_channels": 3,
    "stt_stream_count": 3,
    "reason": "added"
  }
}

Field descriptions

FieldTypeDescription
channelobjectSnapshot of the channel whose status changed; same fields as a channels[] element in session_started (channel_id, speaker_name, transcription_languages, status)
channel.statusstringThe channel's current status: preparing / ready / removed / error (see the state machine below)
active_channelsintNumber of channels currently capturing audio (excluding removed channels)
stt_stream_countintThe number of recognition channels currently counted (billed on this count as each minute begins; drops immediately after remove_channel); always 1 in shared mode
reasonstringOptional. Why the status changed: added / language_change / reconnect / resumed / removed / speaking_speed / stt_error (see the table below). When omitted, the channel became ready for the first time (see the state machine notes)

reason values

reasonWhen it fires
addedadd_channel added a channel (status becomes preparing)
language_changeLanguage change via set_channel_language
reconnectAutomatic reconnection after the channel failed; recovery after a session resume (resume_token) also uses this reason
resumedResuming after a pause (resume); every channel prepares again
removedremove_channel deactivated a channel (status becomes removed)
speaking_speedset_speaking_speed applying the new setting channel by channel
stt_errorThe channel failed and cannot recover automatically (status becomes error)

status state machine

  • preparing: the channel returns to this state every time it starts or re-prepares recognition — the initial setup at start, joining via add_channel, settings changes from set_channel_language / set_speaking_speed, automatic reconnection, and resuming after a pause. The channel needs roughly 4 seconds in this state to finish preparing; speech is not lost, only delayed. The one exception is a sentence already in progress at the moment of the switch or reconnect: it cannot be kept, and is reported separately via segment_discarded.
  • ready: entered when the channel receives its first recognition result (starts producing text). A ready event triggered by a settings change reuses the same reason as the preparing that started it, so the frontend can match "this ready" back to "that operation"; the first ready after start or add_channel carries no reason.
  • removed: the channel was deactivated by remove_channel. Transcripts and audio already produced are kept; the channel number cannot be reused (including via add_channel).
  • error: the channel failed and cannot recover automatically (reason: "stt_error"). If the channel later recovers on its own, it reports preparing (reason: "reconnect") → ready again, or switches straight to ready when text production resumes.

Shared mode: only the first channel (the first element of channels[]) actually performs recognition, and the other channels' preparing / ready / error follow the first channel — when the first channel turns ready, all channels turn ready together; when the first channel prepares recognition again, all channels return to preparing together. Each channel receives its own event, with the same reason as the first channel. A channel added with add_channel takes on the first channel's current status directly; removed follows each channel's own status. When the first channel turns error, the other channels turn error as well (with the same reason); when the first channel recovers, they recover together.

These four status values are consistent with the channels[].status snapshot in session_started and the settings.channels[].status snapshot in resume_ok — after resuming a session, restore each channel's status display from the snapshot (error also appears in snapshots; it is never scrubbed).

Sequence example (add_channel)

→ send add_channel (channel_id: 3, language ja-JP)
← channel_status  channel.status: "preparing", reason: "added",
                  active_channels: 3, stt_stream_count: 3
   (roughly 4 seconds of preparation; keep sending channel 3's audio as usual —
    nothing is lost, text is only delayed)
← channel_status  channel.status: "ready" (no reason)
   (the channel produced its first text; from now on this channel's result
    events carry channel_id: 3 and speaker_id: "channel_3" in origin)

Sequence example (set_channel_language)

→ send set_channel_language (channel_id: 2, switch to en-US)
← channel_status  channel.status: "preparing", reason: "language_change"
   (the switch takes effect after roughly 4 seconds; channel_id and speaker_id
    are unchanged and the transcript stays continuous)
← channel_status  channel.status: "ready", reason: "language_change"

segment_discarded

Description

Segment discarded notification. It tells you that a sid has been discarded and will receive no further events — no is_final: true original text, and no translation for that segment.

Some operations interrupt recognition. If a sentence happens to be in progress at that moment, it cannot be kept. Your client may already have received interim results (is_final: false) for it, so on receiving this event, clear that sid from any "translating" state.

Two-way translation mode (automatic and manual) never receives this event — in automatic mode, a sentence in progress is finalized and sent with is_final: true; in manual mode, if recognition is rebuilt while the user is speaking (for example, after a speaking speed change or a reconnect), that part is merged into the sentence's final result. The content is kept in the transcript in both cases.

{
  "type": "voice-translation",
  "data": {
    "action": "segment_discarded",
    "sid": 7,
    "reason": "speaking_speed",
    "channel_id": 1
  }
}

Field descriptions

FieldTypeDescription
sidintThe discarded segment number
reasonstringWhy it was discarded: speaking_speed / language_change / reconnect / resumed / broadcast_go_live (see below)
channel_idintOptional. Source channel number; sent only in multi-channel mode

reason values

reasonWhen it fires
speaking_speedset_speaking_speed applying a new speed
language_changeset_channel_language changing that channel's language
reconnectAutomatic reconnection after the recognition connection failed
resumedSession resume (resume_token), or resuming after a pause (resume)
broadcast_go_livebroadcast_go_live moving from standby to live; standby segments are not written to the final transcript

The first four values come from the same set as the reason of channel_status, so you can match both events back to the same operation. The exception is a session resume (resume_token): channel_status uses reconnect, while this event uses resumed.

You may already have received this event even when set_speaking_speed fails (set_speaking_speed_failed) — recognition is interrupted before the rebuild, so the segment cannot be kept whether the rebuild succeeds or not. A failure response does not mean "nothing happened".


viewer_joined

Description

Viewer joined event (broadcast mode only). When a viewer joins the broadcast, the host receives this event.

{
  "type": "voice-translation",
  "data": {
    "action": "viewer_joined",
    "viewer": {
      "id": "viewer_abc123",
      "ip": "192.168.1.100",
      "language": "zh-TW"
    },
    "viewer_count": 5,
    "queue_count": 2
  }
}

Field descriptions

FieldTypeDescription
viewerobjectInformation about the viewer who joined
viewer.idstringViewer ID
viewer.ipstringViewer IP address
viewer.languagestringThe language the viewer selected
viewer_countnumberCurrent viewer count
queue_countnumberNumber of viewers in the queue

viewer_left

Description

Viewer left event (broadcast mode only). When a viewer leaves the broadcast, the host receives this event.

{
  "type": "voice-translation",
  "data": {
    "action": "viewer_left",
    "viewer_id": "viewer_abc123",
    "viewer_count": 4,
    "queue_count": 1
  }
}

Field descriptions

FieldTypeDescription
viewer_idstringThe ID of the viewer who left
viewer_countnumberCurrent viewer count
queue_countnumberNumber of viewers in the queue

error

Description

Error event. Triggered when an operation fails or a system error occurs.

{
  "type": "error",
  "data": {
    "error_code": "session_not_started",
    "severity": "error",
    "message": "Session not started",
    "context": "voice-translation",
    "request_id": "req_abc123xyz789",
    "timestamp": "2026-01-15T10:30:45.123Z"
  }
}

A sentence-level error (such as a translation failure for one language of a sentence) additionally carries sid and details:

{
  "type": "error",
  "data": {
    "error_code": "llm_content_filtered",
    "severity": "warning",
    "message": "Content filtered",
    "context": "translation",
    "sid": 5,
    "request_id": "req_abc123xyz789",
    "timestamp": "2026-04-26T10:30:45.123Z",
    "details": {
      "provider": "llm_service",
      "source_lang": "zh-TW",
      "translation_language": "ja-JP"
    }
  }
}

Field descriptions

FieldTypeDescription
error_codestringError code (for programmatic handling)
severitystringSeverity: fatal / error / warning
messagestringHuman-readable error message
contextstringError source category
sidintOptional. The sentence number for a sentence-level error (such as a translation failure for that sentence); not included for non-sentence-level errors
request_idstringRequest tracing ID
timestampstringThe time the error occurred (ISO 8601)
detailsobjectOptional. Error context; common keys: provider, translation_language, source_lang. internal_error (a single message failed but the connection is kept) additionally carries message_type (always present) and action (best-effort; absent when parsing fails). See websocket-api.md Per-Message Errors

Severity descriptions

severityDescriptionRecommended handling
fatalFatal errorStop the service and require reconnection
errorOperation failedShow an error prompt and allow retry
warningWarningShow a warning without blocking the operation

A warning-level error does not mean the recording has ended; examples are stt_silence_warning (about to end for lack of speech) and broadcast_standby_warning (the standby phase is about to reach its time limit). The recording continues; when it actually ends you receive the corresponding fatal error and status: "ended".

For the complete list of error codes, see Error Code Reference.


upload_error

v1.5.6 documentation fix: Earlier documentation described a standalone event format of type: "voice-translation" + action: "upload_error", but in practice this format was never sent on the wire. Storage upload failures always use the unified error envelope, with one of the three error codes in the table below.

If your client listens for action === "upload_error", switch to listening for type === "error" and matching on error_code.

Storage-layer error codes (sent via the error event)

Error codeDescription
storage_connection_failedStorage service connection failed
storage_upload_failedFile upload failed
storage_queue_fullUpload queue full

speakers_auto_merged

Sent in multi-speaker mode when the system determines that two speakers are in fact the same person and merges them automatically.

{
  "type": "voice-translation",
  "data": {
    "action": "speakers_auto_merged",
    "source_speaker_id": "Guest-2",
    "target_speaker_id": "Guest-1"
  }
}
FieldTypeDescription
source_speaker_idstringThe speaker that was merged away
target_speaker_idstringThe speaker that was kept after the merge

An automatic merge usually happens before any transcript entry has been produced, so the event carries no list of affected sentences. On receiving it, the client merges source_speaker_id into target_speaker_id in its local speaker list.


summary_done

Description

An event pushed after recording stops, once server-side non-streaming summary generation is complete. After receiving this event, the client can call GET /api/v1/sse/history/transcribe/{taskId} to retrieve the summary content (the payload does not include final_content, to avoid bloating the WebSocket message).

v1.5.5 adds two fallback audit fields, summary_fallback_level / summary_dropped_segments: when a custom prompt or transcript content triggers the LLM service content filter, the system automatically regenerates the summary in a reduced mode (standard mode -> neutral mode -> segment-omission mode) and uses these two fields to notify the client of the path actually taken.

Examples

Standard mode succeeds directly (no fallback, no filtering triggered):

{
  "type": "voice-translation",
  "data": {
    "action": "summary_done",
    "task_id": "550e8400-e29b-41d4-a716-446655440000",
    "summary_id": "sum_a1b2c3d4e5f6g7h8",
    "summary_mode": "custom",
    "summary_template": "skin-clinic-acme-v2",
    "summary_plain_text": true,
    "tokens_used": { "input": 1234, "output": 567 }
  }
}

Segment-omission mode triggered (summary produced after some transcript segments were omitted):

{
  "type": "voice-translation",
  "data": {
    "action": "summary_done",
    "task_id": "550e8400-e29b-41d4-a716-446655440000",
    "summary_id": "sum_a1b2c3d4e5f6g7h8",
    "summary_mode": "custom",
    "summary_template": "skin-clinic-acme-v2",
    "summary_plain_text": true,
    "tokens_used": { "input": 3456, "output": 789 },
    "summary_fallback_level": 3,
    "summary_dropped_segments": [3, 7]
  }
}

Field descriptions

FieldTypeDescription
actionstringAlways summary_done
task_idstringRecording UUID
summary_idstringThe internal ID of this summary
summary_modestring"builtin" or "custom"
summary_templatestringeffective slug — builtin → the built-in template slug (such as meeting); custom → the customer slug
summary_plain_textbooleanWhether the output is plain text
tokens_used.input / .outputintToken usage (a cumulative value across every generation request made for this summary when a fallback is triggered)
summary_fallback_levelint (omit)Present only when a fallback was triggered (2 or 3); omitted when standard mode succeeds directly. 2 = neutral mode (regenerated with a neutral instruction); 3 = segment-omission mode (regenerated after the offending segments were omitted)
summary_dropped_segmentsint[] (omit)Present only when fallback_level=3; the indices of the trimmed transcript segments (in original order)

Interpreting the fallback level (for frontend UI hints)

summary_fallback_levelMeaningSuggested UI hint
(field omitted)Standard mode succeeds directly, no fallbackDo not show a hint
2The customer prompt triggered filtering, so a neutral fallback prompt was used instead"Your custom instructions contained terms the content filter could not process; the summary was generated using neutral mode"
3The transcript content triggered filtering; the offending segments were trimmed before producing the summary"The transcript contained N segments that could not be processed; the summary was generated after omitting the relevant content" (N = summary_dropped_segments.length)

If segment-omission mode also fails, summary_done is not sent; instead, summary_error is sent with error_code=llm_content_filtered (see §summary_error below).

Note: The payload deliberately does not include final_content. The client must call GET /api/v1/sse/history/transcribe/{taskId} itself to retrieve the full summary text. summary_fallback_level and summary_dropped_segments are also provided as top-level fields of the init_summary event during history playback.


summary_error

Description

An event pushed when summary generation fails, or when no summary is generated because the available credits are insufficient, so the client does not need to keep polling to find out.

Example

{
  "type": "voice-translation",
  "data": {
    "action": "summary_error",
    "task_id": "550e8400-e29b-41d4-a716-446655440000",
    "error_code": "summary_failed",
    "message": "Summary generation failed"
  }
}

Field descriptions

FieldTypeDescription
actionstringAlways summary_error
task_idstringRecording UUID
error_codestringSummary error code (such as summary_failed / summary_timeout / summary_mode_field_mismatch / summary_insufficient_credit, etc.)
messagestringHuman-readable error message (already sanitized; does not include the LLM raw error)

When error_code is summary_insufficient_credit, the available credits were insufficient (the recording ended because the credits ran out, or the credits at the end could not cover the summary fee); the transcript and audio are still saved, and after topping up you can get a summary through Regenerate Summary.


Version: V1.24.1 Last Updated: 2026-10-07

Copyright © 2026