Broadcast Feature Guide
Table of Contents
- Overview
- End-to-End Flow Diagram
- Host-Side Flow
- Viewer-Side Flow
- Password Protection
- Capacity Management and Queuing
- Broadcast Status and Lifecycle
- Standby Phase in Detail
- Announcements
- Error Handling
- Related Reference Documents
Overview
The VAS broadcast feature provides one-to-many real-time subtitle broadcasting. The host sends audio over WebSocket for speech recognition and translation, while viewers receive multilingual subtitles in real time over SSE.
Key Features
- Real-time subtitles: The host's speech is instantly converted into multilingual subtitles
- One-to-many architecture: One host, multiple viewers receiving simultaneously
- Multilingual translation: Supports multiple translation languages; viewers can choose their preferred language
- TTS audio: Viewers can receive translated speech playback
- Standby phase: The host can warm up and test first, confirming the equipment works before going live
- Password protection: Supports public or password-protected broadcasts
- Capacity management: Viewer limits and queuing
- Speaker identification: Supports multi-speaker diarization (Speaker Diarization)
API Types Involved
| API Type | Role | Purpose |
|---|---|---|
| REST API | Host | Create, query, update, and revoke broadcasts |
| REST API | Viewer | Query broadcast information, verify password |
| WebSocket | Host | Send audio, control the broadcast, manage viewers |
| SSE | Viewer | Receive live subtitles, translations, and TTS |
| Broadcast REST API | Viewer | Query the live broadcast status |
Authentication
- Host side: The REST API uses an API Key, and WebSocket uses Ticket authentication. See Authentication for details.
- Viewer side: Authenticated via the broadcast Token; no API Key is required. Password-protected broadcasts additionally require a
viewer_access_token.
End-to-End Flow Diagram
Host Side Viewer Side
========= ===========
[1] POST /api/v1/broadcasts
Create the broadcast, get token
|
[2] Share your viewer page URL with token ─> Receive the share link
| |
[3] POST /api/v1/auth/ticket [A] GET /api/v1/viewer/broadcasts/{token}
Get the WebSocket Ticket Query public broadcast info
| |
[4] WebSocket connection [B] (if password protected)
wss://vas-poc.vurbo.ai/ws POST /viewer/broadcasts/{token}/verify
| Get viewer_access_token
[5] Send start action |
type: "broadcast" [C] EventSource connection
broadcast_token: "xxx" GET /broadcast/{token}/text
broadcast_phase: "standby" |
| |
===== Standby Phase ===== |
| Receive connected event
[6] STT/translation results visible Receive standby event
to host only Show "Preparing, please wait..."
Host tests the equipment |
| |
===== Switch to live phase ===== |
| |
[7] broadcast_go_live ─────────────────────> Receive phase_changed event
| phase: "live"
===== Live ===== |
| |
[8] Send audio (audio action) [D] Receive live subtitles
Receive recognition results origin event (source text)
Receive translation results translation event (translation)
| tts_ready event (TTS audio)
| |
[9] Pause (pause) ─────────────────────────> Receive paused event
| |
[10] Resume (resume) ──────────────────────> Receive resumed event
| |
[11] Send announcement ────────────────────> Receive announcement event
| |
[12] Stop (stop) ──────────────────────────> Receive ended event
Close the SSE connection
Host-Side Flow
Step 1: Create a Broadcast
Create a broadcast channel via the REST API and obtain the share Token and link.
curl -X POST "https://vas-poc.vurbo.ai/api/v1/broadcasts" \
-H "X-API-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"transcription_languages": ["zh-TW", "en-US"],
"translation_languages": ["en-US", "ja-JP"],
"name": "Tech Seminar Live Subtitles",
"access_type": "public",
"max_viewers": 100,
"speaker_diarization": false,
"tts_config": {
"en-US": {"voice": "en-US-JennyNeural", "speaking_rate": 1.0},
"ja-JP": {"voice": "ja-JP-NanamiNeural", "speaking_rate": 1.0}
},
"summary_template": "meeting",
"summary_language": "zh-TW"
}'
Key request parameters:
| Parameter | Required | Description |
|---|---|---|
transcription_languages | Yes (or use the deprecated transcription_language) | Array of transcription languages (string[], up to 10, distinct; the languages the host speaks). When speaker_diarization is true, only a single transcription language is supported; providing multiple languages returns 422 |
transcription_language | No | (Deprecated, kept for backward compatibility) Equivalent to the first element of transcription_languages |
translation_languages | No | Array of translation languages |
name | No | Channel name (max 100 characters; cannot be changed after creation and is not used as the recording name) |
access_type | No | public (default) or password |
pass_code | Conditional | Password (required when type is password, 4-12 characters; letters, digits and common punctuation only (no Chinese characters, no spaces)) |
max_viewers | No | Maximum number of viewers (1 to the account viewer limit; defaults to that limit when omitted) |
speaker_diarization | No | Whether to enable speaker identification (default false) |
tts_config | No | Default TTS settings |
summary_template | No | Summary template slug (e.g., meeting, lecture) |
summary_language | No | Summary output language; must be a language code from the supported language list (defaults to the first transcription language, i.e. the first element of transcription_languages, if not specified) |
callback_url | No | Webhook callback URL (notified when recording processing completes/fails) |
Best practice: list only the transcription languages that will actually occur
When you provide multiple transcription languages, the language of each speech segment is identified automatically; the more candidate languages you provide, the more likely misidentification becomes — especially when the content contains words shared across languages (proper nouns, numbers, loanwords). Recommendations:
- List only the languages that will actually occur; don't pad the list to 10 "just in case".
- Fewer candidates are more accurate; when you are certain there is only one language, provide just one (no language identification is needed, and accuracy is highest).
- Avoid listing languages that share a script or are close relatives (e.g. multiple Latin-script European languages,
zh-CNandzh-TW, or regional variants of the same language), which are the most easily confused.
Key fields in a successful response:
| Field | Description |
|---|---|
id | Broadcast ID (used for subsequent REST API operations) |
token | Share Token (4-character short code, a-z0-9, used for viewer connections) |
share_url | Default share URL (not an openable viewer page; see Step 2) |
status | Initial status is pending |
For the full parameter list, see REST API - Broadcasts.
Webhook notifications: After adding the
callback_urlparameter, you receive arecording.completedevent notification when the broadcast recording finishes processing. See the Webhook Callback Guide for details.
Step 2: Share the Link with Viewers
The viewer page is provided by your application. Put the token from the response into your viewer page URL and share it with viewers, for example:
https://your-app.example.com/broadcast/a3f9
Your viewer page then uses the token to get the broadcast information and receive captions, as described in the Viewer-Side Flow.
The
share_urlin the response is not an openable viewer page. Do not share it with viewers directly.
Step 3: WebSocket Connection
The host needs to establish a WebSocket connection to send audio.
3.1 Get a Ticket:
curl -X POST "https://vas-poc.vurbo.ai/api/v1/auth/ticket" \
-H "X-API-Key: YOUR_API_KEY"
3.2 Establish the WebSocket connection:
const ws = new WebSocket('wss://vas-poc.vurbo.ai/ws', [`ticket.${ticket}`]);
3.3 Send the start action:
{
"type": "voice-translation",
"data": {
"action": "start",
"type": "broadcast",
"broadcast_token": "a3f9",
"audio_format": "pcm",
"broadcast_phase": "standby",
"standby_message": "The talk is about to begin, please wait..."
}
}
Key parameters for start in broadcast mode:
| Parameter | Required | Description |
|---|---|---|
type | Yes | Must be "broadcast" |
broadcast_token | Yes | The Token obtained when creating the broadcast |
audio_format | Yes | pcm or webm |
broadcast_phase | No | standby or live (default). Only these two lowercase values are accepted; an empty string is treated as live, and any other value is rejected |
standby_message | No | The message viewers see during the standby phase |
tts_config | No | Multilingual TTS settings (can override the settings from creation time) |
summary_template | No | Summary template slug (can override the settings from creation time; uses the broadcast channel default if not specified) |
Note:
broadcast_tokenis allowed only withtype: "broadcast"; any other type that carriesbroadcast_tokenis rejected (invalid_parameter) and the recording does not start.
Note: In broadcast mode, the language settings are automatically taken from the broadcast channel settings, so you do not need to send
transcription_languagesandtranslation_languagesin start.summary_templateandsummary_languageare also taken from the broadcast channel settings; pass them only when you need to override.
Note: The host also receives interim translations (
is_final: false); handle them the same way as on the viewer side, see Step 4: Receive Live Content. The recording name does not reuse the channel name; ifstartomitsname, a name such asBroadcast #1is generated.
Successful response:
{
"type": "voice-translation",
"data": {
"action": "session_started",
"session_id": "550e8400-e29b-41d4-a716-446655440000",
"task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
"recording_type": "broadcast",
"recognition_mode": "multi_speaker",
"phase": "standby",
"viewer_count": 0,
"queue_count": 0,
"peak_viewers": 0,
"total_viewers": 0,
"message": "Speech recognition started"
}
}
Step 4: Standby Phase (standby)
The standby phase lets the host test equipment and warm up STT/translation before going live.
Standby phase characteristics:
- STT/translation results are visible to the host only; viewers cannot see them
- Viewers see the waiting message set in
standby_message - The host can confirm that the microphone, recognition accuracy, and so on are working correctly
- The standby message can be updated dynamically at any time
- There is a time limit (30 minutes by default); see Standby Time Limit
Update the standby message dynamically:
{
"type": "voice-translation",
"data": {
"action": "set_standby_message",
"message": "The presenter is getting ready, expected to start in about 5 minutes..."
}
}
Viewers immediately receive the updated standby event, and the message is automatically translated into all translation languages.
For details, see Standby Phase in Detail.
Step 5: Going Live (go_live)
Once you confirm the equipment is working, switch to the live phase:
{
"type": "voice-translation",
"data": {
"action": "broadcast_go_live"
}
}
Successful response:
{
"type": "voice-translation",
"data": {
"action": "broadcast_phase_changed",
"phase": "live",
"message": "Broadcast has started"
}
}
After switching:
- STT/translation results begin broadcasting to viewers
- Results begin being written to the transcript
- TTS begins being sent to viewers
- Viewers receive the
phase_changedevent
Note: If you set
broadcast_phase: "live"(the default) at start, the standby phase is skipped and the broadcast goes live directly.
Viewers who join after the broadcast goes live do not receive what the host said during the standby phase or its translations.
When a broadcast goes from standby to live, billing starts at the moment it goes live; the standby phase is not billed.
Obtaining the task_id for this broadcast
Once the broadcast goes live, you receive a broadcast_recording_ready event carrying the finalized task_id for this broadcast:
{
"type": "voice-translation",
"data": {
"action": "broadcast_recording_ready",
"task_id": "3f9a1c2e-..."
}
}
Important: The
task_idinsession_startedis an initial value for the connection stage — it is not the final ID for this broadcast. You receive this event after going live regardless of whether you went through the standby phase first or started directly withbroadcast_phase: "live"(the default). For all subsequent operations (querying the task, exchanging for a Floating Subtitle Feed Token, and so on), always use thetask_idfrom this event — using the initial value returns no matching record.When the broadcast goes live directly, this event always arrives after
session_started. When the host disconnects and goes live again within the grace period (a takeover), the new connection receives this event with a newtask_id; see Disconnection Handling.
Step 6: Operations During the Broadcast
While live, the host can perform the following operations:
Send audio:
{
"type": "voice-translation",
"data": {
"action": "audio",
"payload": "Base64-encoded audio data"
}
}
Pause the broadcast:
{
"type": "voice-translation",
"data": {
"action": "pause"
}
}
Viewers receive the paused event and see "Broadcast paused."
Resume the broadcast:
{
"type": "voice-translation",
"data": {
"action": "resume"
}
}
Viewers receive the resumed event.
Step 7: Update Settings Dynamically
While the broadcast is running, you can adjust settings in real time via the REST API:
curl -X PATCH "https://vas-poc.vurbo.ai/api/v1/broadcasts/{id}" \
-H "X-API-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"translation_languages": ["en-US", "ja-JP", "ko-KR"],
"max_viewers": 200
}'
Settings that can be updated dynamically:
| Setting | Description |
|---|---|
access_type | Switch between public/password protected |
pass_code | Update the password |
max_viewers | Adjust the viewer limit |
transcription_languages | Change the transcription languages (string[], up to 10, distinct; when speaker_diarization is true, only a single transcription language is supported) |
transcription_language | Change the transcription language (deprecated, kept for backward compatibility; equivalent to the first element of transcription_languages) |
translation_languages | Add or remove translation languages |
speaker_diarization | Enable/disable speaker identification |
tts_config | Update TTS settings |
summary_template | Summary template (empty string clears it) |
summary_language | Summary output language; must be a supported language code (empty string clears it) |
Only broadcasts in
pending,active, orpausedstatus can be updated. The channel name cannot be changed after creation; anamefield is ignored.
Step 8: Viewer Management
The host can monitor the viewer count during the broadcast. The system checks every 3 seconds and pushes a viewer_count event whenever there is a change:
{
"type": "voice-translation",
"data": {
"action": "viewer_count",
"viewer_count": 45,
"queue_count": 8,
"peak_viewers": 50,
"total_viewers": 123
}
}
| Field | Description |
|---|---|
viewer_count | Current online viewer count |
queue_count | Number of viewers waiting in the queue |
peak_viewers | Peak viewer count for this broadcast |
total_viewers | Total number of viewers that have ever connected |
Step 9: Stop the Broadcast
{
"type": "voice-translation",
"data": {
"action": "stop"
}
}
After stopping:
- All viewers receive the
endedevent - The system automatically uploads the audio file and transcript
- The broadcast status changes to
ended - The host receives a
task_completeevent (itstask_idmatches the one returned bybroadcast_recording_ready, and can be used for subsequent queries)
Viewer-Side Flow
Your viewer page can live on your own domain: the viewer API and caption stream below can be called directly from a web page on any domain. Do not send credentials with the request (such as credentials: 'include'). See Calling from a Browser.
Step 1: Retrieve Broadcast Information
After a viewer opens the share link, they first query the broadcast information.
Method A: via the Viewer API
curl -X GET "https://vas-poc.vurbo.ai/api/v1/viewer/broadcasts/{token}"
The response includes the channel name, access type, available languages, the TTS voice list, and other information. The frontend can use this to display the broadcast information page.
Key fields in the Viewer API response:
| Field | Description |
|---|---|
name | Channel name |
access_type | public or password |
requires_password | Whether password verification is required |
is_live | Whether the broadcast is currently live |
translation_languages | Available translation languages |
tts_languages | Languages that support TTS |
tts_voices | The list of available TTS voices for each language |
Step 2: Password Verification (if required)
If the broadcast is set to password protected (requires_password: true), the viewer must verify the password first:
curl -X POST "https://vas-poc.vurbo.ai/api/v1/viewer/broadcasts/{token}/verify" \
-H "Content-Type: application/json" \
-d '{"password": "mySecret123"}'
Successful response:
{
"data": {
"viewer_access_token": "aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vWaB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW",
"expires_at": "2026-01-04T10:00:00.000Z"
}
}
The viewer_access_token you obtain must be included in the subsequent SSE connection. The token is valid for 24 hours.
Step 3: SSE Connection to Receive Subtitles
Viewers receive the subtitle stream in real time over an SSE connection.
Basic connection (public broadcast):
const eventSource = new EventSource(
'https://vas-poc.vurbo.ai/broadcast/{token}/text'
);
Filter to a specific translation language:
const eventSource = new EventSource(
'https://vas-poc.vurbo.ai/broadcast/{token}/text?lang=en-US'
);
Enable TTS:
const eventSource = new EventSource(
'https://vas-poc.vurbo.ai/broadcast/{token}/text?lang=en-US&tts=true'
);
Password-protected broadcast:
const eventSource = new EventSource(
'https://vas-poc.vurbo.ai/broadcast/{token}/text?lang=en-US&viewer_access_token=aB3dE5fG7hI9jK1l...'
);
SSE connection parameters:
| Parameter | Required | Description |
|---|---|---|
token | Yes | Broadcast share Token (path parameter) |
lang | No | Filter to a specific translation language (e.g., en-US) |
tts | No | Whether to enable TTS (true / false) |
viewer_access_token | Conditional | Required for password-protected broadcasts |
Step 4: Receive Live Content
After connecting successfully, the viewer receives the following events in order:
4.1 connected (connection confirmation)
{
"session_id": "abc123",
"source_lang": "zh-TW",
"subscribed_lang": "en-US",
"available_langs": ["en-US", "ja-JP"],
"tts_languages": ["en-US"],
"phase": "live",
"recognition_mode": "single",
"client_id": "client_xyz"
}
4.2 origin (source text)
{
"sid": 1,
"text": "Hello everyone, welcome to today's technical seminar",
"speaker_id": "0",
"start_time": "00:05",
"is_final": true
}
4.3 translation
{
"sid": 1,
"language": "en-US",
"text": "Hello everyone, welcome to today's technical seminar",
"speaker_id": "Royx",
"is_final": true
}
The translation of a sentence may first arrive several times as interim results with is_final: false, followed by the finalized version with is_final: true. Match them by sid and language and replace what is displayed.
4.4 tts_ready (TTS audio, requires tts=true)
{
"sid": 1,
"language": "en-US",
"transcript": "Hello everyone, welcome to today's technical seminar",
"text": "Hello everyone, welcome to today's technical seminar",
"audio": "Base64EncodedMP3...",
"format": "mp3",
"duration_ms": 3200,
"boundaries": [...]
}
Step 5: Select Language and TTS
Viewers can select their preferred translation language and TTS settings through the SSE connection parameters.
Switching languages: You need to close the existing SSE connection and reconnect with a new lang parameter.
// Switch to Japanese
eventSource.close();
const newEventSource = new EventSource(
`https://vas-poc.vurbo.ai/broadcast/${token}/text?lang=ja-JP&tts=true`
);
Complete viewer-side example:
function connectBroadcast(token, lang, enableTts, viewerToken) {
let url = `https://vas-poc.vurbo.ai/broadcast/${token}/text`;
const params = new URLSearchParams();
if (lang) params.set('lang', lang);
if (enableTts) params.set('tts', 'true');
if (viewerToken) params.set('viewer_access_token', viewerToken);
if (params.toString()) url += `?${params.toString()}`;
const eventSource = new EventSource(url);
// Connection confirmation
eventSource.addEventListener('connected', (e) => {
const data = JSON.parse(e.data);
console.log(`Connected, available languages: ${data.available_langs.join(', ')}`);
});
// Standby phase
eventSource.addEventListener('standby', (e) => {
const data = JSON.parse(e.data);
const msg = data.translations?.[lang] || data.message;
showWaitingScreen(msg);
});
// Phase change
eventSource.addEventListener('phase_changed', (e) => {
const data = JSON.parse(e.data);
if (data.phase === 'live') {
hideWaitingScreen();
}
});
// Source text
eventSource.addEventListener('origin', (e) => {
const data = JSON.parse(e.data);
displayOriginText(data.sid, data.text, data.speaker_id, data.start_time);
});
// Translation
eventSource.addEventListener('translation', (e) => {
const data = JSON.parse(e.data);
displayTranslation(data.sid, data.language, data.text);
});
// TTS
eventSource.addEventListener('tts_ready', (e) => {
const data = JSON.parse(e.data);
playTtsAudio(data.audio, data.boundaries, data.text);
});
// Pause/resume
eventSource.addEventListener('paused', (e) => {
showPausedOverlay();
});
eventSource.addEventListener('resumed', (e) => {
hidePausedOverlay();
});
// Announcement
eventSource.addEventListener('announcement', (e) => {
const data = JSON.parse(e.data);
const msg = data.translations?.[lang] || data.message;
showAnnouncement(msg);
});
// End
eventSource.addEventListener('ended', (e) => {
const data = JSON.parse(e.data);
showEndedScreen(data.reason);
eventSource.close();
});
// Kicked
eventSource.addEventListener('kicked', (e) => {
showKickedMessage();
eventSource.close();
});
// Queue
eventSource.addEventListener('queued', (e) => {
const data = JSON.parse(e.data);
showQueuePosition(data.position, data.estimated_wait);
});
eventSource.addEventListener('admitted', (e) => {
hideQueueScreen();
});
// Error
eventSource.addEventListener('error', (e) => {
if (e.data) {
const error = JSON.parse(e.data);
handleError(error.error_code, error.message);
}
eventSource.close();
});
return eventSource;
}
Password Protection
A broadcast can be set to password protected, requiring viewers to enter the correct password before entering.
Create a Password-Protected Broadcast
curl -X POST "https://vas-poc.vurbo.ai/api/v1/broadcasts" \
-H "X-API-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"transcription_languages": ["zh-TW", "en-US"],
"translation_languages": ["en-US"],
"access_type": "password",
"pass_code": "mySecret123"
}'
Viewer-Side Verification Flow
[1] GET /api/v1/viewer/broadcasts/{token}
↓ Confirm requires_password: true
[2] Show the password entry screen
↓ Viewer enters the password
[3] POST /api/v1/viewer/broadcasts/{token}/verify
↓ Get viewer_access_token
[4] SSE connection includes viewer_access_token
GET /broadcast/{token}/text?viewer_access_token=xxx
Switch Access Type Dynamically
While the broadcast is running, you can change a public broadcast to password protected, or vice versa:
# Change to password protected
curl -X PATCH "https://vas-poc.vurbo.ai/api/v1/broadcasts/{id}" \
-H "X-API-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"access_type": "password",
"pass_code": "newPassword"
}'
# Change back to public
curl -X PATCH "https://vas-poc.vurbo.ai/api/v1/broadcasts/{id}" \
-H "X-API-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"access_type": "public"
}'
Capacity Management and Queuing
Viewer Limit
Each broadcast can set a maximum number of viewers (max_viewers). Viewers beyond the limit enter the queue.
Queue Flow
When the viewer count reaches the limit:
- After a new viewer connects via SSE, they receive a
queuedevent - The system informs them of their queue position and estimated wait time
- As viewers leave, viewers in the queue are admitted in order
- When a viewer enters the broadcast, they receive an
admittedevent
queued event:
{
"position": 3,
"estimated_wait": "About 2 minutes"
}
admitted event:
{
"message": "You have entered the broadcast"
}
Queue Timeout
Viewers who are still waiting in the queue after 5 minutes receive an error event (error_code is broadcast_queue_timeout), and the connection then closes; to keep waiting, reconnect:
event: error
data: {"error_code":"broadcast_queue_timeout","severity":"error","message":"Queue timeout","context":"broadcast","request_id":"req_abc123xyz789","timestamp":"2026-09-26T10:30:45Z"}
Adjust the Limit Dynamically
The host can adjust the viewer limit dynamically during the broadcast:
curl -X PATCH "https://vas-poc.vurbo.ai/api/v1/broadcasts/{id}" \
-H "X-API-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"max_viewers": 200}'
After raising the limit, viewers in the queue are admitted automatically.
Broadcast Status and Lifecycle
Status Transitions
pending ──(WebSocket start)──> active ──(pause)──> paused
| |
| <──(resume)───────
|
├──(stop)──> ended
|
pending ──(DELETE)──> revoked
Status Descriptions
| Status | Description | Allowed Operations |
|---|---|---|
pending | Created, waiting for the host to start | Update settings, revoke |
active | Live | Pause, stop, update settings, send audio |
paused | Paused | Resume, stop, update settings |
ended | Ended | None (read-only) |
revoked | Revoked (only pending broadcasts can be revoked) | None (read-only) |
is_live Field
is_live is a convenience field that is true when the status is active or paused, indicating that the broadcast is in progress.
End Reasons
When a broadcast ends (the host ends it, or the host stays disconnected past the grace period), viewers receive an ended event and the connection then closes. ended carries only a notice text (data.message) and no end reason; every case sends the same event.
A queue timeout is not ended but an error event (broadcast_queue_timeout); see Queue Timeout above.
Standby Phase in Detail
The standby phase is the warm-up period before the broadcast officially starts, letting the host test equipment and confirm recognition quality before going live.
Enable the Standby Phase
Set broadcast_phase: "standby" in the WebSocket start action:
{
"type": "voice-translation",
"data": {
"action": "start",
"type": "broadcast",
"broadcast_token": "YOUR_TOKEN",
"audio_format": "pcm",
"broadcast_phase": "standby",
"standby_message": "The talk is about to begin, please wait..."
}
}
Standby Phase Behavior
| Item | Behavior |
|---|---|
| STT recognition | Operates normally; results are sent to the host only |
| Translation | Operates normally; results are sent to the host only |
| Viewer subtitles | Not sent; viewers see standby_message |
| TTS | Not sent to viewers |
| Transcript | Not written |
| Viewer history | Does not include standby content (viewers who join after going live do not see it either) |
| Billing | Not billed; billing starts at the moment the broadcast goes live |
| Time limit | 30 minutes by default; see Standby Time Limit below |
Update the Standby Message Dynamically
After entering the standby phase, you can update the message shown to viewers at any time:
{
"type": "voice-translation",
"data": {
"action": "set_standby_message",
"message": "Adjusting the equipment, expected to start in 3 minutes"
}
}
After the update, all viewers immediately receive a new standby event:
event: standby
data: {"message":"Adjusting the equipment, expected to start in 3 minutes","translations":{"en-US":"Adjusting equipment, expected to start in 3 minutes","ja-JP":"機器の調整中、約3分後に開始予定"}}
Note:
set_standby_messagecan only be used during the standby phase. It returns an error once the broadcast is already live.
Standby Time Limit
When the accumulated standby time reaches the limit (30 minutes by default), the session ends automatically.
- About 2 minutes before the limit, the host first receives a
broadcast_standby_warning(anerrorevent withseverity: "warning"); the broadcast continues.details.standbySecondsis the accumulated standby time in seconds,details.remainingSecondsis the time remaining, anddetails.limitSecondsis the limit. - At the limit, the host receives
broadcast_standby_timeout(severity: "fatal", withstandbySecondsandlimitSecondsindetails), followed bystatus: "ended". No recording exists during the standby phase, so there is notask_complete, no Webhook, and no charge. - The count stops once the broadcast goes live (
broadcast_go_live). - Time spent waiting to resume after a disconnect still counts; after resuming, the count continues from where it was rather than restarting.
- When the host disconnects and sends
startagain within the grace period (a takeover), the new connection continues the accumulated standby time if the original session was still in standby. - To start again after the limit, get a new Ticket and send
start.
{
"type": "error",
"data": {
"error_code": "broadcast_standby_warning",
"severity": "warning",
"message": "Standby is about to reach its time limit; the broadcast will end automatically",
"context": "broadcast",
"request_id": "req_abc123xyz789",
"timestamp": "2026-09-25T10:28:00.000Z",
"details": {
"standbySeconds": 1680,
"remainingSeconds": 120,
"limitSeconds": 1800
}
}
}
When the warning arrives, we recommend prompting the host: "Standby time is almost up. Go live or start again."
The Viewer Experience During the Standby Phase
- After connecting, the viewer receives a
connectedevent (phase: "standby") - Immediately afterward, they receive a
standbyevent containing the waiting message and its multilingual translations - When the host switches to the live phase, the viewer receives a
phase_changedevent (phase: "live") - They then begin receiving subtitles and translations
eventSource.addEventListener('standby', (e) => {
const data = JSON.parse(e.data);
// Display the corresponding translation based on the viewer's selected language
const displayLang = 'en-US';
const msg = data.translations?.[displayLang] || data.message;
showWaitingScreen(msg);
});
eventSource.addEventListener('phase_changed', (e) => {
const data = JSON.parse(e.data);
if (data.phase === 'live') {
hideWaitingScreen();
// Start displaying the subtitle area
}
});
Announcements
The host can send announcement messages to all viewers during the broadcast.
Send an Announcement
{
"type": "voice-translation",
"data": {
"action": "broadcast_announcement",
"message": "The meeting will end in 5 minutes"
}
}
Viewer-Side Reception
Viewers receive an announcement event over SSE, containing the source text and its multilingual translations:
{
"message": "The meeting will end in 5 minutes",
"translations": {
"en-US": "The meeting will end in 5 minutes",
"ja-JP": "会議は5分後に終了します"
}
}
The frontend can display the corresponding translation based on the viewer's selected language:
eventSource.addEventListener('announcement', (e) => {
const data = JSON.parse(e.data);
const displayLang = 'en-US';
const msg = data.translations?.[displayLang] || data.message;
showAnnouncementPopup(msg);
});
Error Handling
Common Host-Side Errors
| Error Code | Description | Suggested Action |
|---|---|---|
broadcast_token_required | Broadcast mode requires a Token | Confirm that broadcast_token is provided |
broadcast_token_invalid | Invalid Token | Confirm the Token is correct and not expired |
broadcast_not_ready | Broadcast service not yet started | Retry later |
broadcast_not_enabled | Not in broadcast mode | Confirm type: "broadcast" |
broadcast_not_in_standby | Not in the standby phase | Can only be used during the standby phase |
broadcast_cannot_revoke | Cannot revoke | Only pending broadcasts can be revoked |
broadcast_standby_warning | The standby phase is about to reach its time limit (the broadcast continues) | Prompt the host to go live or start again |
broadcast_standby_timeout | The standby phase reached its time limit and the session has ended | Get a new Ticket and send start |
invalid_parameter | Invalid broadcast_phase value, or broadcast_token sent with a type other than broadcast | Fix the parameter indicated by details.field |
Common Viewer-Side Errors
| Error Code | Description | Suggested Action |
|---|---|---|
broadcast_not_found | Broadcast not found | Confirm the Token is correct |
broadcast_session_ended | Broadcast has ended | Inform the viewer that the broadcast has ended |
broadcast_capacity_full | Viewer limit reached | Join the queue |
broadcast_token_invalid | Invalid Token | Confirm the Token is correct |
broadcast_token_revoked | Token revoked | The broadcast has been revoked |
broadcast_password_incorrect | Incorrect password | Re-enter the password |
Disconnection Handling
Host disconnection:
- Viewers receive a
host_disconnectedevent: the broadcast is frozen but not ended, so viewers do not need to disconnect; show a "host reconnecting" notice - If the host reconnects within the grace period, viewers receive a
host_reconnectedevent; whendata.resumedistrue, playback resumes at the same time (it isfalse, and the broadcast stays paused, if the host had paused it before disconnecting) - If the host fails to reconnect before timing out, viewers receive an
endedevent
Host goes live again within the grace period (takeover):
When the host disconnects and, within the grace period, sends start again with the same broadcast_token, the new session takes over the channel:
broadcast_recording_readyon the new connection carries a newtask_id; the old connection is already gone and receives nothing.- The broadcast is therefore split into two recordings: the one before the takeover completes when the old session wraps up, and the one after completes when the new session ends. Each sends its own
recording.completedWebhook (a recording with no audio at all sendsrecording.failedinstead). - The channel stays live during the takeover and viewer access is not cleared; the new session is billed as usual.
- Floating subtitle viewers: the subtitles for the old recording end; use the share link for the new recording.
- While the old connection is still connected (not disconnected), a new
startwith the samebroadcast_tokenis rejected (broadcast_token_already_used). To start a new session with the same token, sendstopon the old connection first, wait fortask_complete, then send the newstart.
Viewer disconnection:
- After the SSE connection drops, the browser reconnects automatically (the built-in EventSource mechanism)
- After reconnecting, the viewer receives a new
connectedevent - SSE uses a 15-second heartbeat interval to keep the connection alive
Related Reference Documents
- REST API - Broadcasts (Create / Query / Update / Revoke)
- REST API - Speakers (Offline Speaker Editing)
- WebSocket - Voice Translation (start / broadcast_go_live / broadcast_announcement, etc.)
- SSE - Broadcast Viewer (Viewer Live Subtitle Stream)
- TTS Speech Synthesis Guide
- Speaker Management Guide
Version: V1.24.1 Last Updated: 2026-10-07