SSE API
Note: The URL used in this document (
vas-poc.vurbo.ai) is the planned deployment URL. The official URL will be announced separately after launch.
Table of Contents
- Table of Contents
- Connection Information
- Broadcast SSE API
- Floating Subtitle SSE — see dedicated spec
- GET /api/v1/sse/history/transcribe/{taskId} (Retrieve Conversation History)
- GET /api/v1/sse/retranslate/{taskId} (Retranslate Full Transcript)
- GET /api/v1/sse/recordings/{taskId}/entries/{sid}/retranslate (Single-Sentence Retranslation, added in v1.4.0)
- GET /api/v1/sse/retranslate/summary/{taskId} (Retranslate Summary)
- Regenerate Summary (GET Preview / POST Save)
- POST /api/v1/sse/summary (Ad-hoc Summary, added in v1.9.1)
- POST /api/v1/sse/summary/translate (Summary Translation, added in v1.17.0)
- GET /api/v1/sse/audio/{taskId} (Audio Streaming Playback) — see the dedicated spec
- GET /api/v1/sse/tts/{taskId} (TTS Audio Stream)
- GET /api/v1/sse/imports/{importId}/progress (Import Progress Stream)
Connection Information
| Item | Value |
|---|---|
| Base Path | https://vas-poc.vurbo.ai/api/v1/sse |
| Protocol | HTTP + Server-Sent Events (SSE) |
| Data Format | text/event-stream |
| Authentication | Header X-API-Key: {KEY} |
Authentication
SSE APIs that require authentication accept two delivery methods for the API key (both are supported):
# Method A: HTTP Header (recommended, better security)
X-API-Key: vas_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
# Method B: Query string (native browser EventSource fallback)
?api_key=vas_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
Note: The native browser EventSource API does not support custom headers. You can use the
?api_key=query string instead, or use the fetch API with a ReadableStream / an SSE client library that supports headers. In query string mode, the API key appears in the URL, so avoid writing the full URL into server logs or leaking it in screenshots.
Floating Subtitle SSE
A read-only "floating subtitle" transcript stream for an in-progress recording (source-language original text + target-language translations), subscribed over an independent connection with cross-device/window support. The base path is https://vas-poc.vurbo.ai (the real-time service). First exchange for a feed_token via POST /api/v1/auth/tasks/{taskId}/subtitle-feed-token, then connect to GET /tasks/{task_id}/subtitle?feed_token=....
The owner can also enable sharing so that other on-site audience members connect with a share secret (read-only, not charged separately; the audience limit is server-configured, default 10, excluding the owner). While connected, the viewers event reports the current viewer count.
See Floating Subtitle SSE spec for details.
Broadcast SSE API
The Broadcast SSE API provides a live subtitle streaming feature, allowing viewers to watch real-time transcription and translation content through a share link.
Note: The base path for Broadcast SSE is
https://vas-poc.vurbo.ai/broadcast, which differs from the other SSE APIs.
GET /broadcast/{token}/text (Viewer Live Subtitle Stream)
Description
Viewers connect using a share token to receive an SSE stream of real-time transcription and translation.
Use Cases
- Viewers watching live subtitles
- Multilingual translation subtitle display
- TTS audio playback
Authentication
Token authentication (no API key required): verified through the {token} in the URL path.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
token | string | Yes | Broadcast share token (4-character short code, a-z0-9, path parameter) |
lang | string | No | Filter for a specific translation language (e.g., en-US) |
tts | boolean | No | Whether to enable TTS (true / false, default false) |
viewer_access_token | string | Conditional | Viewer access token (required for password-protected broadcasts) |
Password Protection Note: When a broadcast is set to password-protected, viewers must first obtain a
viewer_access_tokenthrough the password verification API, then include this token in the query parameters of the SSE connection.
Request Example
// Receive all languages
const eventSource = new EventSource(
'https://vas-poc.vurbo.ai/broadcast/a3f9/text'
);
// Receive only English translation
const eventSource = new EventSource(
'https://vas-poc.vurbo.ai/broadcast/a3f9/text?lang=en-US'
);
// Receive English translation and enable TTS
const eventSource = new EventSource(
'https://vas-poc.vurbo.ai/broadcast/a3f9/text?lang=en-US&tts=true'
);
Event Types
| Event | Description | Notes |
|---|---|---|
connected | Connection confirmation | - |
queued | Added to the waiting queue | Queueing mechanism |
admitted | Entered live from the queue | Queueing mechanism |
origin | Original text (STT) | - |
translation | Translation result | - |
tts_ready | TTS audio ready | - |
paused | Broadcast paused | Host paused manually (a host disconnection is reported as host_disconnected) |
resumed | Broadcast resumed | Host resumed |
ended | Broadcast ended | - |
kicked | Removed | Viewer management |
error | Error | - |
speaker_renamed | Speaker renamed | - |
speaker_reassigned | Single-sentence speaker change | - |
speakers_merged | Speakers merged | - |
recording_started | New recording started | Viewers should clear the previous subtitles on screen |
max_viewers_changed | Viewer limit changed | - |
host_disconnected | Host temporarily offline and reconnecting | The broadcast is frozen, not ended |
host_reconnected | Host has reconnected | data.resumed indicates whether playback resumed as well |
standby | Standby phase notification | - |
phase_changed | Phase change notification | - |
announcement | Host announcement | - |
Event Format
connected:
{
"available_langs": ["en-US", "ja-JP"],
"tts_languages": ["en-US"],
"phase": "standby",
"recognition_mode": "single"
}
| Field | Type | Description |
|---|---|---|
available_langs | array | List of available translation languages |
tts_languages | array | List of languages with TTS enabled (an empty array means no TTS) |
phase | string | Broadcast phase: standby (preparing) or live (active) |
recognition_mode | string | Recognition mode: single (single speaker) or multi_speaker (multi-speaker diarization) |
The
connectedevent does not carry the list of transcription languages. Retrieve the channel's public information (name,transcription_languages,translation_languages, TTS voices) withGET /api/v1/viewer/broadcasts/{token}instead - see Viewer API.
queued:
{
"position": 3,
"estimated_wait": "約 2 分鐘"
}
| Field | Type | Description |
|---|---|---|
position | number | Position in the queue (1 = next up) |
estimated_wait | string | Estimated wait time. Server-fixed value in Traditional Chinese: 少於 1 分鐘 ("less than 1 minute") or 約 N 分鐘 ("about N minutes") |
admitted:
{
"message": "已進入直播"
}
Several broadcast events carry a
messagestring that the server returns as a fixed Traditional Chinese value (已進入直播,廣播已暫停,廣播已恢復,廣播已結束,廣播已開始). These are display strings only - render them as-is or substitute your own copy, and never parse or compare against them.
origin:
{
"sid": 1,
"text": "Hello everyone",
"speaker_id": "Guest-1",
"speaker_label": "Guest-1",
"start_time": "00:05",
"is_final": true
}
| Field | Type | Description |
|---|---|---|
sid | number | Sentence ID |
text | string | Original text content |
speaker_id | string | Optional. Original speaker ID (immutable); sent only in multi-speaker diarization mode, omitted in single-speaker mode |
speaker_label | string | Display label (after applying speaker_aliases; equals speaker_id when no alias exists) |
start_time | string | Start time (mm:ss), counted from 00:00 once the broadcast goes live. Standby content is never sent to viewers |
is_final | boolean | Whether this is the final result |
translation:
{
"sid": 1,
"language": "en-US",
"text": "Hello everyone",
"speaker_id": "Guest-1",
"speaker_label": "Royx",
"is_final": true
}
| Field | Type | Description |
|---|---|---|
sid | number | The corresponding sentence ID |
language | string | Translation language |
text | string | Translated content |
speaker_id | string | Original speaker ID (multi-speaker conversation mode; immutable) |
speaker_label | string | Display label (after applying speaker_aliases) |
is_final | boolean | Whether this is the final result |
tts_ready:
{
"sid": 1,
"language": "en-US",
"transcript": "Hello, hi everyone",
"text": "Hello everyone",
"audio": "//uQxAAAAAANIAAAAAExBTUUzLjEwMFVVVV...",
"format": "mp3",
"duration_ms": 2340,
"boundaries": [
{"offset_ms": 0, "duration_ms": 320, "text": "Hello", "text_offset": 0, "word_length": 5},
{"offset_ms": 320, "duration_ms": 280, "text": "everyone", "text_offset": 6, "word_length": 8}
]
}
| Field | Type | Description |
|---|---|---|
sid | number | The corresponding sentence ID |
language | string | TTS language |
transcript | string | Original transcript (source text) |
text | string | Translated text |
audio | string | Base64-encoded MP3 audio |
format | string | Audio format, fixed as "mp3" |
duration_ms | number | Audio duration (milliseconds) |
boundaries | array | Word boundaries (optional, see table below) |
Word boundary fields (each object in the boundaries array):
| Field | Type | Description |
|---|---|---|
offset_ms | number | The word's start time in the audio (ms) |
duration_ms | number | The word's pronunciation duration (ms) |
text | string | The word text |
text_offset | number | The word's starting position in the text |
word_length | number | The word's character length |
Note:
- The host must specify which languages enable TTS via the
tts_configparameter in thestartcommand - Only viewers who subscribed to that language and enabled TTS will receive this event
- It is sent only during the
livephase; no TTS is sent during thestandbyphase
paused:
{
"message": "廣播已暫停"
}
| Field | Type | Description |
|---|---|---|
message | string | Notification message |
A host going offline and coming back are separate events (
host_disconnected/host_reconnected), not reason values onpaused.
resumed:
{
"message": "廣播已恢復"
}
ended:
{
"message": "廣播已結束"
}
| Field | Type | Description |
|---|---|---|
message | string | Notification message |
All three events carry only
message— no end reason and no timestamps. Themessagetext is for display only; determine state from the event name itself, not from this string.
kicked:
{
"message": "Kicked by host"
}
messageis for display only, never parse it. When the host removes a viewer the value is the fixed stringKicked by host; when the removal is caused by an access-type or passcode change, the value is a Traditional Chinese explanation of that change.
error:
{
"error_code": "broadcast_session_ended",
"severity": "error",
"message": "Broadcast session ended",
"context": "broadcast",
"request_id": "req_abc123xyz789",
"timestamp": "2025-12-05T10:30:45.123Z"
}
Sentence-level errors (such as a translation failure for a specific language) additionally carry sid and translation_language, making it easy for the frontend to flag which language failed for a given sentence:
{
"error_code": "llm_content_filtered",
"severity": "warning",
"message": "Content filtered",
"context": "translation",
"sid": 5,
"translation_language": "ja-JP",
"timestamp": "2026-04-26T10:30:45.123Z"
}
| Field | Type | Description |
|---|---|---|
error_code | string | Error code |
severity | string | Severity: warning / error / fatal |
message | string | Error message (a fixed English display string returned by the server; display only, never parse it) |
context | string | The context in which the error occurred (e.g., broadcast, translation) |
sid | int | Optional. The sentence number for a sentence-level error (e.g., when that sentence's translation fails) |
translation_language | string | Optional. The target language that failed to translate (viewers can use this to determine whether a specific language failed for that sentence) |
request_id | string | Optional. Request tracking ID. Present only on connection-phase errors (non-200 HTTP status); sentence-level and session-level errors omit it |
timestamp | string | Time the error occurred (ISO 8601) |
speaker_renamed:
Multi-speaker conversation mode only. Sent when the host performs a global speaker rename.
{
"speaker_id": "Guest-1",
"new_label": "Royx",
"affected_sids": [1, 3, 5, 7]
}
| Field | Type | Description |
|---|---|---|
speaker_id | string | The resolved original speaker ID (even if the input is a display label, the event returns the original ID) |
new_label | string | New display label (e.g., Royx) |
affected_sids | array | List of affected sentence IDs |
speaker_reassigned:
Multi-speaker conversation mode only. Sent when the host changes the speaker of a single sentence.
{
"sid": 3,
"old_speaker_id": "Guest-1",
"new_speaker_id": "Guest-2",
"new_speaker_label": "Amy"
}
| Field | Type | Description |
|---|---|---|
sid | number | The sentence ID that was modified |
old_speaker_id | string | Original speaker ID (e.g., Guest-1) |
new_speaker_id | string | The new original speaker ID (e.g., Guest-2) |
new_speaker_label | string | New speaker display label (after applying speaker_aliases; equals the original ID when no alias exists) |
speakers_merged:
Multi-speaker conversation mode only. Sent when the host merges speakers. After merging, all sentences belonging to that speaker are reassigned to the target speaker.
{
"source_speaker_id": "Guest-2",
"target_speaker_id": "Guest-1",
"target_speaker_label": "Manager Wang",
"affected_sids": [3, 5, 7]
}
| Field | Type | Description |
|---|---|---|
source_speaker_id | string | The original speaker ID being merged (e.g., Guest-2) |
target_speaker_id | string | The original speaker ID of the merge target (e.g., Guest-1) |
target_speaker_label | string | Target speaker display label (after applying speaker_aliases; equals the original ID when no alias exists) |
affected_sids | array | List of affected sentence IDs: sentences that belonged to the source speaker, plus the target speaker's existing sentences whose display name changed because of the merge (for example, when the source speaker's custom name is carried over to the target) |
standby:
When a viewer connects during the standby phase, this event is received immediately after the
connectedevent, indicating that the broadcast has not yet officially started. The host can dynamically update the standby message via the WebSocketset_standby_messageaction; after the update, all viewers receive a newstandbyevent.
{
"message": "The presentation is about to begin, please wait...",
"translations": {
"en-US": "The presentation is about to begin, please wait...",
"ja-JP": "プレゼンテーションがまもなく始まります。お待ちください..."
}
}
| Field | Type | Description |
|---|---|---|
message | string | The message displayed during the standby phase (original text) |
translations | object | Translation results (optional); the key is the language code and the value is the translated text |
phase_changed:
Sent when the broadcast switches from the standby phase to the active phase.
{
"phase": "live",
"message": "Broadcast has started"
}
| Field | Type | Description |
|---|---|---|
phase | string | The new phase: live (active phase) |
message | string | Phase change message |
announcement:
An announcement message sent by the host; all viewers receive it.
{
"message": "The meeting will end in 5 minutes",
"translations": {
"en-US": "The meeting will end in 5 minutes",
"ja-JP": "会議は5分後に終了します"
}
}
| Field | Type | Description |
|---|---|---|
message | string | The announcement content (original text) |
translations | object | Translation results (optional); the key is the language code and the value is the translated text |
Heartbeat Mechanism
The SSE connection uses a heartbeat to keep the connection alive:
- Interval: 15 seconds
- Format: SSE comment (starting with
:) - The frontend does not need to handle it; the browser automatically ignores it
: heartbeat
Error Responses
| Error Code | HTTP Status | Description | Recommended Handling |
|---|---|---|---|
broadcast_session_not_found | 404 | Broadcast not found | Confirm the token is correct |
broadcast_session_ended | 410 | Broadcast ended | Notify the user that the broadcast has ended |
broadcast_capacity_exceeded | 503 | Viewer capacity reached | Join the waiting queue |
Note: If an SSE endpoint encounters an unexpected internal exception, it may return
internal_error(consistent with how WebSocket handles a single-message failure); expected domain errors return the corresponding error code (e.g.,sse_translation_failed).
Frontend Example
function connectBroadcast(token, lang = null) {
let url = `https://vas-poc.vurbo.ai/broadcast/${token}/text`;
if (lang) {
url += `?lang=${lang}`;
}
const eventSource = new EventSource(url);
eventSource.addEventListener('connected', (e) => {
const data = JSON.parse(e.data);
console.log(`Connected, phase: ${data.phase}, recognition mode: ${data.recognition_mode}`);
console.log(`Available translations: ${data.available_langs.join(', ')}`);
});
eventSource.addEventListener('queued', (e) => {
const data = JSON.parse(e.data);
console.log(`In queue, position: ${data.position}, estimated wait: ${data.estimated_wait}`);
});
eventSource.addEventListener('admitted', (e) => {
console.log('Entered live');
});
eventSource.addEventListener('origin', (e) => {
const data = JSON.parse(e.data);
console.log(`[${data.start_time}] ${data.text}`);
});
eventSource.addEventListener('translation', (e) => {
const data = JSON.parse(e.data);
console.log(`Translation (${data.language}): ${data.text}`);
});
eventSource.addEventListener('tts_ready', (e) => {
const data = JSON.parse(e.data);
// Decode the Base64 audio and play it
const byteCharacters = atob(data.audio);
const byteNumbers = new Array(byteCharacters.length);
for (let i = 0; i < byteCharacters.length; i++) {
byteNumbers[i] = byteCharacters.charCodeAt(i);
}
const blob = new Blob([new Uint8Array(byteNumbers)], { type: 'audio/mpeg' });
const audio = new Audio(URL.createObjectURL(blob));
audio.play();
});
eventSource.addEventListener('paused', (e) => {
const data = JSON.parse(e.data);
console.log(`Broadcast paused: ${data.message}`);
});
eventSource.addEventListener('resumed', (e) => {
console.log('Broadcast resumed');
});
eventSource.addEventListener('ended', (e) => {
const data = JSON.parse(e.data);
console.log(`Broadcast ended: ${data.message}`);
eventSource.close();
});
eventSource.addEventListener('kicked', (e) => {
console.log('You have been removed');
eventSource.close();
});
eventSource.addEventListener('speaker_renamed', (e) => {
const data = JSON.parse(e.data);
console.log(`Speaker renamed: ${data.speaker_id} → ${data.new_label}`);
console.log(`Affected sentences: ${data.affected_sids.join(', ')}`);
// Update the speaker display name for all affected sentences
});
eventSource.addEventListener('speaker_reassigned', (e) => {
const data = JSON.parse(e.data);
console.log(`Speaker of sentence ${data.sid} changed from ${data.old_speaker_id} to: ${data.new_speaker_label}`);
// Update the speaker display name for that sentence
});
eventSource.addEventListener('standby', (e) => {
const data = JSON.parse(e.data);
// Display the translation matching the viewer's selected language
const displayLang = 'en-US'; // The language the viewer selected
const displayMessage = data.translations?.[displayLang] || data.message;
console.log(`Standby phase: ${displayMessage}`);
// Show the waiting screen
});
eventSource.addEventListener('phase_changed', (e) => {
const data = JSON.parse(e.data);
console.log(`Phase changed: ${data.phase} - ${data.message}`);
// Remove the waiting screen and start displaying subtitles
});
eventSource.addEventListener('announcement', (e) => {
const data = JSON.parse(e.data);
// Display the translation matching the viewer's selected language
const displayLang = 'en-US'; // The language the viewer selected
const displayMessage = data.translations?.[displayLang] || data.message;
console.log(`Announcement: ${displayMessage}`);
// Show the announcement message
});
eventSource.addEventListener('error', (e) => {
if (e.data) {
const error = JSON.parse(e.data);
console.error(`Error [${error.error_code}]: ${error.message}`);
}
eventSource.close();
});
return eventSource;
}
REST API also available: For the channel's public information (name, language lists, TTS voices), see Viewer API.
GET /api/v1/sse/history/transcribe/{taskId} (Retrieve Conversation History)
Description
Loads the complete conversation history for the specified task, including all sentences and the summary. Delivered one item at a time via an SSE stream.
Use Cases
- Viewing the recording details page
- Loading the historical transcript
Authentication
Header: X-API-Key: YOUR_API_KEY
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
taskId | string | Yes | Recording ID (path parameter) |
Request Example
// Use the fetch API (because EventSource does not support headers)
async function connectSSE(taskId, apiKey) {
const response = await fetch(
`https://vas-poc.vurbo.ai/api/v1/sse/history/transcribe/${taskId}`,
{
headers: {
'X-API-Key': apiKey
}
}
);
const reader = response.body.getReader();
// ... handle SSE events
}
Event Sequence
1. connected → Connection confirmation
2. init_metadata → Send task metadata
3. init_sentence → Send sentences one at a time (repeated N times)
4. init_summary → Send the summary
5. init_done → Initialization complete
Event Format
connected:
{"message": "History service connected (taskId: xxx)"}
init_metadata:
{
"task_id": "550e8400-e29b-41d4-a716-446655440000",
"title": "Meeting Notes",
"created_at": "2025-12-17T10:00:00Z",
"type": "transcribe",
"has_speaker_diarization": true,
"transcription_languages": ["zh-TW"],
"translation_languages": ["en-US"],
"summary_template": "general",
"summary_language": "zh-TW",
"speaker_aliases": {"speaker_1": "Manager Wang"}
}
speaker_aliases is a mapping of "original speaker ID → display name"; it is {} (an empty object, not an array) when there are no aliases. The frontend can use this mapping to run a duplicate-name pre-check before renaming a speaker (added in v1.3.12).
init_sentence:
{
"sid": 1,
"origin": "Hello, nice to meet you",
"translations": {
"en-US": "Hello, nice to meet you"
},
"start_time": "00:05",
"speaker_id": "speaker_1",
"speaker_label": "Manager Wang"
}
If a sentence has a translation failure, it additionally carries a translation_errors field (only present when there is a failure), so the frontend can distinguish between "that language was not scheduled for translation" (the key is missing from translations) and "translated but failed" (the key is present in translation_errors). The same language may have both an older translation and a failure record (a failed retranslation keeps the previous translation), so read both fields together:
{
"sid": 5,
"origin": "Sentence with sensitive words",
"translations": {
"en-US": "Sensitive sentence"
},
"translation_errors": {
"ja-JP": "llm_content_filtered"
},
"start_time": "00:25",
"speaker_id": "speaker_1",
"speaker_label": "Manager Wang"
}
| Field | Type | Description |
|---|---|---|
sid | int | Sentence number |
origin | string | Original text |
translations | object | Translation results (optional); the key is the language code and the value is the translated text |
translation_errors | object | Optional. Translation failure error codes; the key is the language code and the value is the error_code (e.g., llm_content_filtered) |
channel_id | number | Optional. (Multi-channel only) which physical channel this sentence came from; present only for multi-channel recordings, omitted for single-channel |
start_time | string | Start time (mm:ss format) |
speaker_id | string|null | Original speaker ID (immutable, e.g., speaker_1); the source for target_speaker_id in PATCH /speakers/reassign (flipped in v1.5.3: previously the display name) |
speaker_label | string|null | Display label (the human-readable name after applying speaker_aliases, e.g., Manager Wang); equals speaker_id when no alias exists (added in v1.5.3 to replace the original speaker_id display semantics) |
init_summary:
{
"text": "This is a summary of the meeting notes...",
"mode": "builtin",
"template": "meeting",
"plain_text": false,
"summary_language": "zh-TW"
}
mode,template,plain_textandsummary_languageare always present (summary_languageis the language of this summary and isnullwhen there is no summary);custommode also carriesprompt_snapshot, and an automatic downgrade addsfallback_level/dropped_segments. See the history endpoint for the full field list.
init_done:
{"totalSentences": 10}
Error Responses
| Error Code | HTTP Status | Description | Recommended Handling |
|---|---|---|---|
recording_not_found | 404 | Recording not found | Confirm the taskId is correct |
sse_transcript_not_found | 404 | Transcript not found | The recording may not have finished processing yet |
Frontend Example
// Use the fetch API to handle SSE (you must parse the event-stream yourself)
async function loadHistory(taskId, apiKey) {
const response = await fetch(
`https://vas-poc.vurbo.ai/api/v1/sse/history/transcribe/${taskId}`,
{
headers: {
'X-API-Key': apiKey
}
}
);
const reader = response.body.getReader();
const decoder = new TextDecoder();
while (true) {
const { done, value } = await reader.read();
if (done) break;
const text = decoder.decode(value);
// Parse the SSE format: event: xxx\ndata: {...}\n\n
const events = parseSSE(text);
for (const event of events) {
if (event.type === 'init_metadata') {
console.log('Task info:', event.data.title);
} else if (event.type === 'init_sentence') {
console.log(`[${event.data.start_time}] ${event.data.origin}`);
if (event.data.translations) {
console.log(`Translation: ${event.data.translations['en-US']}`);
}
} else if (event.type === 'init_done') {
console.log('Loading complete');
}
}
}
}
GET /api/v1/sse/retranslate/{taskId} (Retranslate Full Transcript)
Description
Retranslates all sentences of the specified task into the target language. Translation results are delivered one at a time via an SSE stream.
Use Cases
- Switching the display language
- Updating the translation content
Authentication
Header: X-API-Key: YOUR_API_KEY
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
taskId | string | Yes | Recording ID (path parameter) |
targetLang | string | Yes | Target language code |
segmented | string | No | 1 declares that the client supports segmented retranslation (v1.18.0) |
fromSid | number | No | Only together with segmented=1: start from this sentence (inclusive) (v1.18.0) |
expectedRevision | number | No | Transcript revision; on a mismatch nothing is translated or charged and transcript_revision_conflict is returned (v1.18.0) |
Use segmented retranslation for long transcripts: without
segmented=1, the whole transcript is translated at once; if it is too long to translate in one request, HTTP 422retranslate_segmentation_requiredis returned before the stream starts (no charge). Withsegmented=1, each request translates one segment;donecarriesrevision, and when the segment was cut short alsotruncated: trueandnextSid. Send the next segment withfromSid={nextSid}untildoneno longer carriestruncated. Each segment is saved and charged separately. See reference/sse/retranslate.md for the full rules.
Request Example
// Use the fetch API (because EventSource does not support headers)
async function retranslateSSE(taskId, targetLang, apiKey) {
const response = await fetch(
`https://vas-poc.vurbo.ai/api/v1/sse/retranslate/${taskId}?targetLang=${targetLang}`,
{
headers: {
'X-API-Key': apiKey
}
}
);
const reader = response.body.getReader();
// ... handle SSE events
}
Event Format
translation:
{"sid": 1, "text": "Hello, nice to meet you", "is_final": true}
done:
{"totalUpdated": 10, "characters_billed": 12700, "charged": "6.4", "billed": true}
The three billing fields appear only when the request actually incurred consumption, so the
request was billed only when billed is true — always use that field as the criterion.
charged is the points
consumed by this operation, calculated from the rate, and reflects usage — usage already
covered by an unlimited plan is still reported here. Integrators that call this API on behalf
of end users and bill them separately can use this value directly instead of deriving it.
error (per-sid sentence translation failure):
When a sentence fails to translate, instead of a translation event, an event: error is sent carrying sid + error_code, interleaved with the translation events. The failure event format matches real-time translation over WebSocket, so the frontend can share one error handler:
event: error
data: {"error_code": "sse_translation_failed", "severity": "error", "message": "SSE translation failed", "context": "sse", "sid": 5, "request_id": "req_abc123xyz789", "timestamp": "2026-04-27T10:30:45.123Z", "details": {"translation_language": "ja-JP", "original_error": "..."}}
| Field | Type | Description |
|---|---|---|
error_code | string | Error code: sse_translation_failed or llm_content_filtered |
severity | string | error for sse_translation_failed; warning for llm_content_filtered |
message | string | Human-readable message |
context | string | sse for sse_translation_failed; translation for llm_content_filtered |
sid | int | The sentence number that failed |
request_id | string | Request tracking ID |
timestamp | string | Time the error occurred (ISO 8601) |
details | object | Includes debug info such as translation_language and original_error |
Failed sentences are saved as translation error records (see the history-playback guide), and the failure markers are visible the next time the history is loaded. For the full specification, see reference/sse/retranslate.md.
Error Responses
| Error Code | HTTP Status | Description | Recommended Handling |
|---|---|---|---|
sse_translation_failed | 500 | Translation failed (per-sid) | The failed sentence is still reported via event: error; the overall flow is not interrupted |
llm_content_filtered | 400 | The sentence content cannot be translated (per-sid) | Retrying will not help; revise the original text and try again. The sentence is excluded from totalUpdated and incurs no consumption |
storage_upload_failed | 500 | Saving the transcript failed | The entire run is discarded and not billed; retry later. After this code you will not receive done |
transcript_revision_conflict | 409 | Another write to the same transcript is in progress, the transcript changed while the request was running, or expectedRevision does not match the current revision | The entire run is discarded and not billed; on a revision mismatch reload the transcript, otherwise simply retry later. After this code you will not receive done |
retranslate_segmentation_required | 422 | segmented=1 was not sent and the transcript is too long to translate in one request (details carries sentenceCount, maxSentences) (v1.18.0) | JSON response before the stream starts, no charge; send segmented=1 to retranslate in segments |
Frontend Example
// Use the fetch API to handle SSE
async function retranslate(taskId, targetLang, apiKey) {
const response = await fetch(
`https://vas-poc.vurbo.ai/api/v1/sse/retranslate/${taskId}?targetLang=${targetLang}`,
{
headers: {
'X-API-Key': apiKey
}
}
);
const reader = response.body.getReader();
const decoder = new TextDecoder();
while (true) {
const { done, value } = await reader.read();
if (done) break;
const events = parseSSE(decoder.decode(value));
for (const event of events) {
if (event.type === 'translation') {
console.log(`Sentence ${event.data.sid}: ${event.data.text}`);
} else if (event.type === 'error') {
console.warn(`Sentence ${event.data.sid} failed to translate: ${event.data.error_code}`);
} else if (event.type === 'done') {
console.log(`Complete, ${event.data.totalUpdated} sentences updated`);
}
}
}
}
GET /api/v1/sse/recordings/{taskId}/entries/{sid}/retranslate (Single-Sentence Retranslation, added in v1.4.0)
Description
Retranslates a single sentence. The most common scenario: after a user edits the original text via PATCH /api/v1/tasks/{id}/entries/{sid}, you call this endpoint to redo all translations for that sentence.
Differences from full-transcript retranslation (/retranslate/{taskId}):
- Full-transcript retranslation: all sentences are translated into a single target language
- Single-sentence retranslation: only one sentence is translated, and every language it has been translated into or failed to translate into can be retried at once; supports optimistic locking
Authentication
Query: api_key (the browser EventSource does not support headers)
Request Parameters
| Parameter | Location | Type | Required | Description |
|---|---|---|---|---|
taskId | path | string | Yes | Recording ID (UUID) |
sid | path | number | Yes | Sentence ID (1-based) |
targetLang | query | string | No | Target language code. When omitted, every language that sentence has either been translated into or failed to translate into is retried, that is the union of the translations and the translation error records |
expectedRevision | query | number | No | Optimistic lock: the current transcript revision; a mismatch returns transcript_revision_conflict |
api_key | query | string | Yes | API key |
Event Format
Event sequence: connected → progress / translated / error ×N → done
// progress (when translation begins for each language)
{ "sid": 5, "lang": "en-US", "status": "translating" }
// translated (when each language completes successfully)
{ "sid": 5, "lang": "en-US", "text": "Hello world", "tokens_used": 25 }
// done (all complete; successfully translated languages are listed in languages_translated)
{
"sid": 5,
"revision": 6,
"original_text_edited_at": "2026-05-06T10:30:00.000000Z",
"languages_translated": ["en-US"],
"languages_failed": ["ja-JP"]
}
Error Responses
| Error Code | HTTP | Description |
|---|---|---|
recording_not_found | 404 | Recording does not exist or does not belong to the user |
recording_not_completed | 422 | The recording has not finished processing |
entry_not_found | 404 | The specified sentence was not found |
entry_text_empty | 422 | The original text of that sentence is empty (whitespace-only counts as empty) |
sse_translation_failed | 500 | A target language failed to translate (per-lang) |
llm_content_filtered | 400 | A target language's content cannot be translated (per-lang); retrying will not help |
transcript_revision_conflict | 409 | Revision mismatch, or another write to the same transcript is in progress |
storage_upload_failed | 500 | Failed to save the transcript |
For the full event format and a workflow example combining optimistic locking with PATCH, see reference/sse/retranslate.md.
init_sentence Edit Marker Fields (added in v1.4.0)
For sentences edited by a user, historyTranscribe adds two fields to the init_sentence event (only present after editing):
{
"sid": 7,
"origin": "Corrected text",
"original_text_raw": "Original STT output",
"original_text_edited_at": "2026-05-06T10:30:00.000000Z",
"translations": { "en-US": "Corrected text" }
}
Frontend detection: determine this by the presence of the field ('original_text_raw' in data); do not compare origin === original_text_raw — a user may edit and then change it back to the same string, in which case the text is equal but the "edited" marker should still be shown. See reference/sse/history.md.
GET /api/v1/sse/retranslate/summary/{taskId} (Retranslate Summary)
Description
Retranslates the summary of the specified task into the target language. Translation results are delivered segment by segment via an SSE stream.
Use Cases
- Switching the summary display language
- Obtaining the summary in a different language
Authentication
Header: X-API-Key: YOUR_API_KEY
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
taskId | string | Yes | Recording ID |
targetLang | string | Yes | Target language code |
Request Example
// Use the fetch API (because EventSource does not support headers)
async function retranslateSummarySSE(taskId, targetLang, apiKey) {
const response = await fetch(
`https://vas-poc.vurbo.ai/api/v1/sse/retranslate/summary/${taskId}?targetLang=${targetLang}`,
{
headers: {
'X-API-Key': apiKey
}
}
);
const reader = response.body.getReader();
// ... handle SSE events
}
Event Format
summary_translation:
{"text": "Accumulated translation result...", "is_final": false}
done:
{"totalUpdated": 1}
The retranslated summary is not saved: the saved summary and its language are unchanged. To switch the summary to another language and keep it, use the save endpoint of Regenerate Summary (POST, billed).
When the translation is incomplete,
donealso carriestruncated: true(added in v1.17.0). Processing time and timeouts follow the same rules as Summary Translation.
This endpoint is not billed. Its
doneevent does not include thecharacters_billed/charged/billedfields.
Error Responses
| Error Code | HTTP Status | Description | Recommended Handling |
|---|---|---|---|
sse_summary_not_found | 404 | Summary not found | This recording has no summary |
sse_summary_translation_failed | 500 | Summary translation failed. details.original_error of Translation timed out means the request timed out | Retry later |
llm_content_filtered | 400 | The summary content cannot be translated | Retrying will not help; revise the summary and try again |
Regenerate Summary (GET Preview / POST Save)
Split into two endpoints + mode-aware. For the full schema, see reference/sse/regenerate-summary.md; this is a quick summary.
| Method | Endpoint | Persists Result | Saves Transcript | Billed | Purpose |
|---|---|---|---|---|---|
| GET | /api/v1/sse/regenerate/summary/{taskId} | No | No | Yes | Preview (dry run) |
| POST | /api/v1/sse/regenerate/summary/{taskId} | Yes | Yes (and increments revision) | Yes | Save (persist officially) |
Known limitation: GET is also billed — the LLM actually consumes tokens, so the GET endpoint cannot be used for free.
Shared Parameters (GET via query string, POST via JSON body)
| Parameter | Type | Required | Description |
|---|---|---|---|
taskId (path) | string | Yes | Recording UUID |
mode | string | Yes | Summary mode enum: builtin / custom |
template | string | Required for builtin / forbidden for custom | Built-in template slug |
prompt | string | Required for custom / forbidden for builtin | The customer's full prompt (replaces the built-in layered prompt, ≤3000 characters) |
promptSlug | string | Required for custom / forbidden for builtin | The customer's own identifier (≤64 Unicode characters, no control characters) |
language | string | No | Output language (defaults to the first transcription language) |
plainText | boolean | No | Whether to request plain-text output (default false) |
Mutual exclusivity rule: a violation is a parameter validation failure (HTTP 200 + an error event whose data carries only message, with no error_code).
Request Example
# Preview builtin (does not persist the result)
curl -N "https://vas-poc.vurbo.ai/api/v1/sse/regenerate/summary/550e8400-...?mode=builtin&template=meeting&language=zh-TW&plainText=true" \
-H "X-API-Key: YOUR_API_KEY"
# Save custom
curl -N -X POST "https://vas-poc.vurbo.ai/api/v1/sse/regenerate/summary/550e8400-..." \
-H "X-API-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"mode":"custom","prompt":"Please emphasize KPIs","promptSlug":"acme-v2","plainText":true}'
Event Sequence
1. connected → Connection confirmation (includes mode=builtin|custom, endpoint=preview|persist)
2. summary_regeneration → Stream summary segments (accumulating; is_final=true marks the last one)
3. done → Complete, includes final_content / mode / template(effective) / prompt_snapshot (only for custom)
or
3. error → Generation failed (sse_summary_regeneration_failed; can occur on both endpoints;
nothing is saved, no done is sent, and it is not billed)
or
3. error → Save failed, or another write is in progress (save endpoint only; the request
stops, no done is sent, and it is not billed)
done event
{
"task_id": "550e8400-...",
"tokens_used": 123,
"final_content": "This meeting...",
"mode": "custom",
"template": "acme-v2",
"plain_text": true,
"persisted": true,
"summary_language": "zh-TW",
"characters_billed": 12700,
"charged": "1.3",
"billed": true,
"prompt_snapshot": "Please emphasize KPIs"
}
mode: summary mode (builtin / custom)template: effective slug — builtin → built-in template slug; custom → customer slugpersisted: whether this summary has been officially saved (falsefor GET,truefor POST)summary_language: the language actually used for this summary. Thelanguagevalue if one was sent; otherwise the first transcription language. Present for both preview (GET) and save (POST), and always has a valuecharacters_billed/charged/billed: consumption for this request. Present only when consumption actually occurred (absent when generation fails); always usebilledto determine whether the request was billed. Both preview (GET) and save (POST) are billed.chargedreflects usage — usage already covered by an unlimited plan is still reported hereprompt_snapshot: only present in custom mode; thepromptcontent the customer passed in verbatim (a mandatory snapshot, the sole basis for reconstruction)truncated: present only when the summary is incomplete (the value is alwaystrue): the summary reached the output length limit, or generation reached the processing time limit (only the part completed so far is returned). Billed as usual, and the save endpoint also saves this summary
Error Codes
| Error Code | HTTP | Description |
|---|---|---|
recording_not_found | 404 | Recording not found |
sse_template_not_found | 404 | Summary template not found |
sse_transcript_not_found | 404 | Transcript not found |
summary_text_empty | 400 | The transcript has no content |
summary_text_too_long | 400 | The transcript exceeds the 200,000-character limit |
sse_summary_regeneration_failed | 500 | Regeneration failed (the response does not include internal error details; content filtering, or a stream that stalls or does not end normally, also falls under this code; nothing is saved or billed) |
Parameter validation failures carry no error code. An invalid
mode, a field combination that does not match the mode, apromptorpromptSlugthat is too long or contains control characters — all of these are rejected during parameter validation. The response is anerrorevent whosedatacarries onlymessage, with noerror_code. Presentmessageto the user; do not try to match on an error code.(The
set_summaryaction for realtime recording takes a different path, and it does have dedicated error codes — see the WebSocket API documentation.)
Frontend Example
async function regenerateSummary(taskId, body, apiKey, { persist = false } = {}) {
const url = `https://vas-poc.vurbo.ai/api/v1/sse/regenerate/summary/${taskId}`;
const init = persist
? { method: 'POST', headers: { 'X-API-Key': apiKey, 'Content-Type': 'application/json' }, body: JSON.stringify(body) }
: { method: 'GET', headers: { 'X-API-Key': apiKey } };
if (!persist) {
const params = new URLSearchParams(body);
return fetch(`${url}?${params}`, init);
}
return fetch(url, init);
}
POST /api/v1/sse/summary (Ad-hoc Summary, added in v1.9.1)
Generates a summary from arbitrary text supplied in the request (not tied to a recording). For the full schema, see reference/sse/adhoc-summary.md; this is a quick summary.
| Method | Endpoint | Persists Result | Billed | Purpose |
|---|---|---|---|---|
| POST | /api/v1/sse/summary | No | Yes | Generate a summary from the content supplied in the request (content the server does not have, such as merged full transcripts or edited transcripts); the result is not stored |
- Required:
content(≤200,000 characters),idempotency_key(≤64 characters,A-Z a-z 0-9 . _ -),mode(builtin / custom; the field mutual-exclusion rules are the same as Regenerate Summary) - Billing: 0.1 credits per 1,000
contentcharacters (same rate as summaries); billed only on successful generation - Duplicate-request guarantee (
idempotency_keyis only valid within a single API key): the entire request (content plus every summary parameter) determines whether a call is a retry. Retrying with the same identifier + an identical request is not billed again (but regenerates — the previous result is not replayed); the same identifier + any differing field returns 409summary_idempotency_key_conflict; failures do not claim the identifier - Error contract differs from other SSE endpoints: before the stream starts, real HTTP status codes are returned (401/403 / 422 / 404 / 400 / 402 / 409), and authentication only accepts the header —
?api_key=is not accepted; SSEerrorevents are only used after the stream starts - Event sequence:
connected→summary_regeneration×N →done(the done event has notask_id/persisted, and adds thecharacters_billed/charged/idempotency_key/billedfields;summary_languageis the same as in Regenerate Summary and iszh-TWwhenlanguageis not sent). When the summary is incomplete,donealso carriestruncated: true(billed as usual); when generation fails,error(sse_summary_regeneration_failed) is sent instead and nothing is billed - Billing fields:
characters_billedandchargedare always present;billedisfalsefor free retries and empty generation results — reconcile against that field.chargedreflects usage — usage already covered by an unlimited plan is still reported here
POST /api/v1/sse/summary/translate (Summary Translation, added in v1.17.0)
Translates summary text supplied in the request into the specified language, without tying it to a recording. See reference/sse/summary-translate.md for the full specification; this is a quick overview.
| Method | Endpoint | Result saved | Billed | Purpose |
|---|---|---|---|---|
| POST | /api/v1/sse/summary/translate | No | Yes | Translate the content supplied in the request, such as a merged or edited summary the server does not have. The result is not saved |
- Required:
content: 30,000 characters or fewertarget_languageidempotency_key: 64 characters or fewer, using onlyA-Z a-z 0-9 . _ -
- Optional:
source_language. Detected automatically when omitted. Returns 422 if it matchestarget_language. - Billing: 0.1 credits per 200 characters of
content; a partial unit counts as a full one. This is the same rate as full-transcript retranslation. Billed only on success. - Duplicate requests: same rules as Ad-hoc Summary.
- Retrying the identical request with the same identifier is not billed again.
- Reusing the identifier with any field changed returns 409
summary_idempotency_key_conflict.
- Error contract: errors before the stream starts return a real HTTP status code (401/403, 422, 402, 409, 429). Only errors after the stream starts use an SSE
errorevent. - Event sequence:
connected→summary_translation×N (cumulative full text) →done.- For long content, the stream often delivers a large block at once, and there may be a pause of a few seconds between events.
donefields:tokens_used,source_language,target_language,characters_billed,charged,idempotency_key,billed.- An incomplete translation also carries
truncated: true: it either hit the processing-time or length limit and was cut off (billed as usual), or was judged unfinished (not billed). Checkbilledto see whether credits were actually deducted.
- An incomplete translation also carries
- Processing time: a request has a limit of about 230 seconds. If translation pauses for more than 60 seconds, an
erroris sent with error codesse_summary_translation_failedanddetails.original_errorset toTranslation timed out.
GET /api/v1/sse/tts/{taskId} (TTS Audio Stream)
Description
Converts the translated content of a historical recording into TTS audio, delivered sentence by sentence via an SSE stream. The frontend can control how many sentences are returned per request.
Use Cases
- Audio playback of translations from historical recordings
- Karaoke effect (combined with word boundaries)
- Voice readout of translated content
Authentication
Header: X-API-Key: YOUR_API_KEY
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
taskId | string | Yes | Recording ID (path parameter) |
language | string | Yes | TTS output language (e.g., en-US) |
voice | string | No | Specify a voice name (e.g., en-US-JennyNeural) |
sid | int | No | Starting sentence ID (default 1, starting from the first sentence) |
length | int | No | Number of sentences to return (default 1, maximum 20) |
Note: The maximum value of
lengthis 20 by default and is configured on the server side. It is automatically truncated when it exceeds the maximum.
Request Example (Single Sentence Playback)
// Use the fetch API (because EventSource does not support headers)
async function playTTSSingle(taskId, language, sid, apiKey) {
const response = await fetch(
`https://vas-poc.vurbo.ai/api/v1/sse/tts/${taskId}?language=${language}&sid=${sid}`,
{
headers: {
'X-API-Key': apiKey
}
}
);
const reader = response.body.getReader();
// ... handle SSE events
}
Request Example (Multiple Sentence Playback)
// Play sentences 5, 6, and 7 (3 sentences total)
async function playTTSMultiple(taskId, language, sid, length, apiKey) {
const response = await fetch(
`https://vas-poc.vurbo.ai/api/v1/sse/tts/${taskId}?language=${language}&sid=${sid}&length=${length}`,
{
headers: {
'X-API-Key': apiKey
}
}
);
const reader = response.body.getReader();
// ... handle SSE events
}
Event Sequence
1. connected → Connection confirmation
2. tts_audio → Send TTS audio sentence by sentence (repeated N times, N = length)
3. tts_done → Playback complete
Event Format
connected:
{
"task_id": "550e8400-e29b-41d4-a716-446655440000",
"language": "en-US",
"voice": "en-US-JennyNeural",
"start_sid": 5,
"length": 3
}
tts_audio:
{
"sid": 5,
"transcript": "Hello, nice to meet you",
"text": "Hello, nice to meet you",
"audio": "Base64EncodedMP3...",
"duration_ms": 2500,
"boundaries": [
{"offset_ms": 0, "duration_ms": 350, "text_offset": 0, "word_length": 5, "text": "Hello"},
{"offset_ms": 350, "duration_ms": 100, "text_offset": 5, "word_length": 1, "text": ","},
{"offset_ms": 500, "duration_ms": 250, "text_offset": 7, "word_length": 4, "text": "nice"},
{"offset_ms": 750, "duration_ms": 200, "text_offset": 12, "word_length": 2, "text": "to"},
{"offset_ms": 950, "duration_ms": 350, "text_offset": 15, "word_length": 4, "text": "meet"},
{"offset_ms": 1300, "duration_ms": 300, "text_offset": 20, "word_length": 3, "text": "you"}
],
"characters_used": 23
}
| Field | Type | Description |
|---|---|---|
sid | int | Sentence ID |
transcript | string | Original transcript (STT recognition result) |
text | string | Translated text (the TTS synthesis source) |
audio | string | Base64-encoded MP3 audio |
duration_ms | int | Audio duration (milliseconds) |
boundaries | array | Word boundary array |
characters_used | int | Characters consumed synthesizing this sentence; 0 on a cache hit (tts_done.total_characters_used is the sum of this field) |
Word Boundary Field Descriptions
| Field | Type | Description |
|---|---|---|
offset_ms | int | The word's start time in the audio (milliseconds) |
duration_ms | int | The word's duration (milliseconds) |
text_offset | int | Position in the original text string (character index) |
word_length | int | Word length (number of characters) |
text | string | Word content |
tts_done:
{
"sentences_sent": 3,
"total_duration_ms": 7500,
"total_characters_used": 142
}
| Field | Type | Description |
|---|---|---|
sentences_sent | int | The number of sentences actually sent |
total_duration_ms | int | The total audio duration of all sentences (milliseconds) |
total_characters_used | int | The total number of characters synthesized in this TTS request (used for quota calculation) |
Error Responses
| Error Code | HTTP Status | Description | Recommended Handling |
|---|---|---|---|
recording_not_found | 404 | Recording not found | Confirm the taskId is correct |
tts_synthesis_failed | 500 | TTS synthesis failed | Retry later |
Frontend Example
// Use the fetch API to handle TTS SSE
async function playTTS(taskId, language, apiKey, startSid = 1, length = 1) {
const url = new URL(`https://vas-poc.vurbo.ai/api/v1/sse/tts/${taskId}`);
url.searchParams.set('language', language);
url.searchParams.set('sid', startSid);
url.searchParams.set('length', length);
const response = await fetch(url, {
headers: {
'X-API-Key': apiKey
}
});
const reader = response.body.getReader();
const decoder = new TextDecoder();
while (true) {
const { done, value } = await reader.read();
if (done) break;
const events = parseSSE(decoder.decode(value));
for (const event of events) {
if (event.type === 'connected') {
console.log(`TTS connection successful, voice: ${event.data.voice}`);
} else if (event.type === 'tts_audio') {
console.log(`Sentence ${event.data.sid}: ${event.data.text}`);
// Play the audio
const audioBlob = base64ToBlob(event.data.audio, 'audio/mp3');
const audioUrl = URL.createObjectURL(audioBlob);
const audio = new Audio(audioUrl);
// Set up the karaoke effect
setupKaraoke(audio, event.data.boundaries, event.data.text);
audio.play();
} else if (event.type === 'tts_done') {
console.log(`Playback complete, ${event.data.sentences_sent} sentences total`);
}
}
}
}
// Base64 to Blob
function base64ToBlob(base64, mimeType) {
const byteCharacters = atob(base64);
const byteNumbers = new Array(byteCharacters.length);
for (let i = 0; i < byteCharacters.length; i++) {
byteNumbers[i] = byteCharacters.charCodeAt(i);
}
const byteArray = new Uint8Array(byteNumbers);
return new Blob([byteArray], { type: mimeType });
}
// Karaoke effect
function setupKaraoke(audio, boundaries, text) {
const updateHighlight = () => {
const currentTimeMs = audio.currentTime * 1000;
const currentWord = boundaries.find((b, i) => {
const nextOffset = boundaries[i + 1]?.offset_ms ?? Infinity;
return currentTimeMs >= b.offset_ms && currentTimeMs < nextOffset;
});
if (currentWord) {
// Highlight the current word
highlightWord(text, currentWord.text_offset, currentWord.word_length);
}
};
const interval = setInterval(updateHighlight, 50);
audio.addEventListener('ended', () => clearInterval(interval));
}
GET /api/v1/sse/imports/{importId}/progress (Import Progress Stream)
Description
Tracks the processing progress of an audio file import in real time. After connecting, progress updates are continuously pushed via an SSE stream until the import completes, fails, or the connection times out.
Use Cases
- Showing a real-time processing progress bar after uploading an audio file
- Tracking the progress of each stage: audio conversion, transcription, translation, summary, etc.
Authentication
Header: X-API-Key: YOUR_API_KEY
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
importId | string | Yes | Import task ID (UUID, path parameter) |
Request Example
curl -N "https://vas-poc.vurbo.ai/api/v1/sse/imports/550e8400-e29b-41d4-a716-446655440000/progress" \
-H "X-API-Key: vas_aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW"
Event Sequence
Scenario 1: Import still in progress
1. connected → Connection confirmation
2. progress → Send the current progress
3. progress ×N → Continuously pushed when progress changes
heartbeat ×N → Sent every 15 seconds when there is no progress change
4. completed → Import succeeded, connection ends
or failed → Import failed, connection ends
or timeout → Exceeded 15 minutes, connection ends
Scenario 2: Import already complete (terminal state)
1. connected → Connection confirmation
2. progress → Send the final progress
3. completed → Send the completed event directly and end
or failed → Send the failed event directly and end
Event Format
connected:
{"message": "Import progress service connected (importId: xxx)"}
progress:
{
"import_id": "550e8400-e29b-41d4-a716-446655440000",
"status": "processing",
"stage": "transcribing",
"progress": 45,
"message": "Transcribing..."
}
| Field | Type | Description |
|---|---|---|
import_id | string | Import task ID (UUID) |
status | string | Import status: pending / processing / completed / failed |
stage | string / null | The current processing stage |
progress | integer | Progress percentage (0-100) |
message | string | Human-readable progress message |
Stage values and their corresponding progress ranges:
| Value | Description | Progress Range |
|---|---|---|
converting | Audio format conversion | 0% - 10% |
transcribing | Speech-to-text | 10% - 60% |
translating | Text translation | 60% - 85% |
summarizing | Generating the summary | 85% - 100% |
completed | Import complete | 100% |
null | Not started yet | — |
completed:
{
"import_id": "550e8400-e29b-41d4-a716-446655440000",
"status": "completed",
"task_id": "abc123-e29b-41d4-a716-446655440000",
"message": "Processing complete"
}
| Field | Type | Description |
|---|---|---|
import_id | string | Import task ID |
status | string | Fixed as completed |
task_id | string | The generated task ID, which can be used for subsequent queries |
message | string | Fixed as Processing complete |
failed:
{
"import_id": "550e8400-e29b-41d4-a716-446655440000",
"status": "failed",
"error_code": "import_invalid_format",
"error_message": "Unsupported audio format"
}
| Field | Type | Description |
|---|---|---|
import_id | string | Import task ID |
status | string | Fixed as failed |
error_code | string | Error code |
error_message | string | Human-readable error message (a general description without internal details; provide the import_id for troubleshooting) |
heartbeat:
Sent every 15 seconds when there is no progress change, used to keep the connection alive.
{"timestamp": 1708761600}
timeout:
Sent when the import has not completed after 15 minutes; the connection ends automatically.
{"message": "Connection timeout"}
Error Responses
| Error Code | HTTP Status | Description | Recommended Handling |
|---|---|---|---|
import_not_found | 404 | The specified import task was not found | Confirm the importId is correct |
Frontend Example
async function trackImportProgress(importId, apiKey) {
const response = await fetch(
`https://vas-poc.vurbo.ai/api/v1/sse/imports/${importId}/progress`,
{
headers: {
'X-API-Key': apiKey
}
}
);
const reader = response.body.getReader();
const decoder = new TextDecoder();
let buffer = '';
while (true) {
const { done, value } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const events = buffer.split('\n\n');
buffer = events.pop();
for (const eventStr of events) {
if (!eventStr.trim()) continue;
const lines = eventStr.split('\n');
let eventType = '';
let eventData = '';
for (const line of lines) {
if (line.startsWith('event: ')) eventType = line.slice(7);
else if (line.startsWith('data: ')) eventData = line.slice(6);
}
if (!eventType || !eventData) continue;
const data = JSON.parse(eventData);
switch (eventType) {
case 'connected':
console.log('Connected:', data.message);
break;
case 'progress':
console.log(`[${data.stage}] ${data.progress}% - ${data.message}`);
updateProgressBar(data.progress, data.stage, data.message);
break;
case 'completed':
console.log('Import complete! Recording ID:', data.task_id);
navigateToRecording(data.task_id);
break;
case 'failed':
console.error('Import failed:', data.error_code, data.error_message);
showError(data.error_message);
break;
case 'timeout':
console.warn('Connection timeout:', data.message);
break;
}
}
}
}
Version: V1.24.1 Last Updated: 2026-09-28