WebSocket API

WebSocket Connection and Authentication

Table of Contents

  1. Connection Information
  2. Authentication Method
  3. Message Format
  4. Heartbeat Mechanism (Health)

Connection Information

ItemValue
Endpointwss://vas-poc.vurbo.ai/ws
ProtocolWebSocket
Data FormatJSON
Auth MethodTicket (see below)

Authentication Method

VAS WebSocket uses a Ticket mechanism for authentication, passing a one-time Ticket via Sec-WebSocket-Protocol. For details, see Authentication.

Step 1: Obtain a Ticket

Use your API Key to exchange for a one-time Ticket via the REST API:

POST /api/v1/auth/ticket
X-API-Key: vas_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx

Response:

{
  "ticket": "aBcDeFgHiJkLmNoPqRsTuVwXyZ012345",
  "expires_in": 60
}
FieldTypeDescription
ticketstringOne-time Ticket (32 chars)
expires_inintValidity period (seconds)

Step 2: Connect to WebSocket Using the Ticket

Place the Ticket in Sec-WebSocket-Protocol using the format ticket.{TICKET_VALUE}:

// Native browser support
const ws = new WebSocket('wss://vas-poc.vurbo.ai/ws', [`ticket.${ticket}`]);

ws.onopen = () => {
  console.log('Connected! Protocol:', ws.protocol);
  // Start using the WebSocket...
};

ws.onerror = (error) => {
  console.error('Connection failed:', error);
};

Node.js example:

const WebSocket = require('ws');

const ws = new WebSocket('wss://vas-poc.vurbo.ai/ws', [`ticket.${ticket}`]);

Ticket Characteristics

CharacteristicDescription
Validity period60 seconds
Usage countSingle use only (deleted immediately after use)
SecurityThe API Key is never exposed in the WebSocket connection
Replay protectionAtomic operations ensure single use

Ticket Error Codes

Error CodeHTTP StatusDescription
ticket_invalid401Ticket invalid or expired
ticket_expired401Ticket expired
ticket_already_used401Ticket already used
ticket_validation_failed500Ticket validation failed

For the complete API specification, see Auth Ticket API.


Message Format

All messages use a unified nested structure:

{
  "type": "service type",
  "data": { ... }
}

Maximum Size of a Single Message

A single WebSocket message has a maximum size (1 MB by default, adjustable per environment). Exceeding it closes the connection outright (close code 1009) with no error message of any kind — the client only sees the connection drop for no apparent reason.

Ordinary use never comes close to this limit: the recommended audio frame is 100 milliseconds (about 4 KB), and even a full second at a time is only a few tens of KB.

Note: The one thing that can hit it is a large glossary. The three glossary blocks can be sent across several config messages — a block you leave out is untouched and keeps its previous value, so splitting the send is safe. If the connection drops immediately after a large config with no error message at all, this is the direction to investigate.

Service Types

typeDescription
healthHeartbeat mechanism
voice-translationVoice translation service
errorError message

Error Message Format

When an error occurs, the server returns a message with type: "error":

{
  "type": "error",
  "data": {
    "error_code": "auth_invalid_api_key",
    "severity": "fatal",
    "message": "Invalid API key",
    "context": "auth",
    "request_id": "req_abc123xyz789",
    "timestamp": "2026-01-15T10:30:45.123Z"
  }
}
FieldTypeDescription
error_codestringError code (for programmatic handling)
severitystringSeverity: fatal / error / warning
messagestringHuman-readable error message
contextstringError source category
request_idstringRequest tracking ID
timestampstringTime the error occurred (ISO 8601)

For the complete list of error codes, see Error Code Reference.


Heartbeat Mechanism (Health)

Description

Used to confirm whether the WebSocket connection is healthy. We recommend sending a ping every 30 seconds; if no pong is received, treat the connection as dropped and reconnect. While the previous recording is still being processed after it ends (for example, while its summary is being generated), pong can be delayed by a few seconds to a few tens of seconds; allow enough waiting time and do not treat the connection as dropped just because a pong has not arrived yet.

Use Cases

  • Maintain long-lived connections
  • Detect connection status
  • Prevent connection timeouts

Request - Ping

{
  "type": "health",
  "data": {
    "action": "ping"
  }
}

Response - Pong

{
  "type": "health",
  "data": {
    "action": "pong"
  }
}

Session Resume

If a WebSocket connection is unexpectedly dropped due to network instability, the client may reconnect with its resume_token within a grace period (default 45 seconds) to rejoin the original recording session — keeping the same task_id and sentence ids (sid), with the transcript timeline continuing from the breakpoint. No need to restart the whole recording.

Use cases: mobile network switching, brief outages in elevators/tunnels, Wi-Fi roaming, etc. Brief network jitter that TCP recovers from within seconds does not drop the connection; session resume targets cases where the connection is actually severed.

Flow

1. start succeeds → session_started contains resume_token + resume_grace_seconds (store them)
2. connection drops → within the grace period:
   a. obtain a NEW single-use Ticket from the backend
   b. reconnect with Sec-WebSocket-Protocol: ["ticket.<newTicket>", "resume.<resume_token>"]
3. success → server returns resume_ok (with server_last_sid) → start a fresh audio stream as after start
   failure → server returns a resume_* error → obtain a new Ticket, then fall back to a fresh start (unless the user has ended the recording)

New session_started fields

{
  "type": "voice-translation",
  "data": {
    "action": "session_started",
    "session_id": "...",
    "task_id": "...",
    "resume_token": "<43 chars; store until this session ends>",
    "resume_grace_seconds": 45,
    "server_time": 1749550000000
  }
}
  • server_time is the server's current time in unix milliseconds. The client can compute its clock skew relative to the server from (server_time, client time when received) as a reference; however, the grace-period countdown must still be measured against the client's wall clock (the instant the connection drops, the server can no longer push messages, so the client must compute the remaining time from the moment of disconnect itself).

Reconnect handshake

On reconnect, Sec-WebSocket-Protocol carries both the ticket and the resume token (ticket always first):

const ws = new WebSocket(url, [`ticket.${newTicket}`, `resume.${resumeToken}`]);
  • Every reconnect must obtain a new Ticket (tickets are single-use).
  • The server validates the Ticket first (re-authentication + quota), then verifies the resume_token ownership against the user_id / api_key_id derived from the Ticket.

Success response - resume_ok

{
  "type": "voice-translation",
  "data": {
    "action": "resume_ok",
    "session_id": "...",
    "task_id": "...",
    "recording_type": "transcribe",
    "recognition_mode": "single",
    "server_last_sid": 42,
    "server_last_offset_ms": 125000,
    "server_recording_ms": 127500,
    "is_paused": false,
    "settings": {
      "speaking_speed": "normal",
      "profanity_handling": "mask",
      "audio_format": "pcm",
      "transcription_languages": ["zh-TW"],
      "translation_languages": ["en-US"],
      "realtime_translation": true,
      "auto_summary": true,
      "summary_plain_text": false,
      "summary_template": "meeting",
      "summary_language": "zh-TW",
      "summary_mode": "builtin",
      "name": "Product Meeting"
    },
    "message": "Resumed the original session"
  }
}

After receiving resume_ok:

  1. Start sending an audio stream again, just like after start. For WebM/Opus you must send a brand-new container (restart the encoder); do not continue the old container. For PCM, simply continue sending.
  2. server_last_sid is the server's current last sentence id. If your locally received maximum sid is lower, a few sentences in flight were lost at the moment of disconnect — the live view may show a brief gap, but the complete content will appear in the final transcript when the recording ends.
  3. server_last_offset_ms is the transcript timeline position (in milliseconds) of the breakpoint, based on the amount of audio already processed (not real elapsed time; the audio timeline is frozen during the disconnect). The client can use it to splice post-reconnect content at the correct timeline position, aligned together with server_last_sid.
  4. server_recording_ms is the recording-head timestamp (in milliseconds, including silence) that the transcript timeline resumes from after reconnect. The client uses it to align its recording-second header to the same timeline the transcript uses; the difference from server_last_offset_ms is the trailing audio (silence or not-yet-finalized speech) after the last finalized sentence before the disconnect (omitted when 0).
  5. is_paused is the paused state the server considers authoritative after resuming. If true, the user was paused before the disconnect — after reopening the audio stream the client should pause immediately, send no audio, and keep the pause UI, instead of resuming recording on its own; false or omitted means recording normally. This lets both a network reconnect and a full-page refresh realign the paused state (after a refresh the local pause memory is lost, so the server is authoritative).
  6. settings is the set of recording settings currently held by the session, used to reconcile the client's local state with the server's authoritative values — see The settings object below.

Note: Do not mix the two notions of time. The transcript timeline (server_last_offset_ms) is based on the audio file and is frozen during a disconnect; the grace reconnect window (resume_grace_seconds) is real wall-clock time and keeps counting down during a disconnect. Deciding "whether reconnect is still possible" must use wall-clock time, never the audio-file time position.

The settings object (added in v1.6.4)

resume_ok returns the recording settings currently held by the session, letting the client reconcile its local state against the server's authoritative values — eliminating drift from edge cases such as pause/resume, session resume, and multiple tabs.

Three principles:

  • Current values: reflects changes made during recording via set_speaking_speed / config / set_name / set_tts, not a snapshot taken at start.
  • API format: values are returned in the public API format (e.g., speaking_speed returns the level string "normal", not an the underlying numeric value).
  • recording_type / recognition_mode are already at the top level of resume_ok and are not duplicated.
FieldTypeDescription
speaking_speedstringCurrent speaking speed (very_slow / slow / normal / fast / very_fast; returns normal if unset)
profanity_handlingstringProfanity handling (mask / remove / show; returns the default mask if unset)
audio_formatstringThe audio format from start (pcm / webm). Resumed connections must keep using the same format (the server decodes with the original format and does not renegotiate)
transcription_languagesstring[]Source languages (in conversation mode, use speaker_language_map as the source of truth after mid-recording language changes)
translation_languagesstring[]Translation target languages
realtime_translationbooleanReal-time translation switch (always present; false is meaningful)
tts_enabledbooleanTTS switch (present when conversation TTS or single-mode TTS is enabled)
tts_modestringTTS mode sync / async (same as above)
tts_configobjectTTS voice settings (conversation bilingual / broadcast multilingual / single-mode single language; { "lang": { "voice", "speaking_rate" } })
conversation_modestringConversation dialog mode auto / manual (changeable mid-recording via switch_conversation_mode; conversation mode only)
speaker_language_mapobjectConversation speaker language map { "1": "zh-TW", "2": "en-US" } (changeable mid-recording via set_speaker_language; conversation mode only, present whether or not speakers was provided at start)
terminologyarrayTerminology (current values accumulated via config; no language grouping)
fuzzy_correctionarrayFuzzy correction [{ "correct", "incorrect": [], "case_insensitive" }] (same as above; case_insensitive is omitted when false)
translation_dictobject | arrayTranslation dictionary. The format matches the one you last sent — send the language-grouped format ({ "language code": [{ "source", "target", "case_sensitive" }] }) and you get that back; send the previous array-of-entries format and you get that back. The case_sensitive field is omitted when false
auto_summarybooleanWhether to auto-generate a summary (always present; false is meaningful)
summary_plain_textbooleanWhether the summary is plain-text output (always present; false is meaningful)
summary_templatestringSummary template slug
summary_languagestringSummary output language
summary_modestringbuiltin / custom
summary_promptstringMode-aware: custom = full prompt / builtin = supplementary instructions (present only when set)
summary_prompt_slugstringPresent in custom mode only
namestringCurrent recording name (reflects set_name changes)
channel_modestringMulti-channel only (recognition_mode: "multi_channel"): channel mode (per_channel / shared)
channelsarrayMulti-channel only: current snapshot of all channels (including removed ones); see below

Empty-collection fields (e.g., no terminology configured) are omitted entirely; treat "field absent" as "not configured".

Multi-channel fields channel_mode / channels[] (added in v1.10.0)

For multi-channel recordings (recognition_mode: "multi_channel"), settings additionally returns channel_mode and channels[]. Session resume does not re-validate the start parameters — after reconnecting, this snapshot is the client's only way to confirm that the server is still in multi-channel mode and that the channel-to-language bindings are unchanged; use it to overwrite the local channel state. In shared mode, channels are not bound to a language; the session's languages are given by settings.transcription_languages.

Fields for each entry in channels[]:

FieldTypeDescription
channel_idintChannel number
speaker_namestringSpeaker name for this channel (present only when set)
transcription_languagesstring[]Language bound to this channel (exactly one; current value after set_channel_language changes). Not present in shared mode
statusstringChannel status: preparing (preparing, not yet producing text) / ready (has started producing text) / removed (disabled via remove_channel) / error (failed and cannot recover automatically). In shared mode, the other channels' preparing / ready / error follow the first channel

Note: Removed channels also appear in the snapshot (status: "removed") — this is not an omission: channel numbers are never reusable (including removed ones). After reconnecting, do not reuse a removed number in add_channel — the server returns channel_id_in_use. For the full multi-channel documentation, see Voice Translation Actions.

Failure responses

Resume failures always return severity: "error" (not fatal). Any failure code means the Ticket attached to this handshake has been consumed; obtain a new Ticket before falling back to a fresh start. If the user has already ended the recording (has sent stop), do not fall back, so that you do not start a new recording the user did not ask for.

error_codeMeaningClient action
resume_token_invalidToken invalid or not found; also returned after a service restart or update, because sessions waiting to be resumed are finalized at that pointNew Ticket → fresh start
resume_grace_expiredGrace period expired, session finalizedNew Ticket → fresh start
resume_ownership_mismatchToken ownership mismatch (user/api_key)New Ticket → fresh start
resume_unavailableTemporarily unavailable (this connection cannot resume the original session; also returned when the original session has been sent stop or has been ended and is still being processed)New Ticket → fresh start

When the service restarts or is updated, sessions waiting to be resumed are finalized right away: the recording is saved and the completion notification is sent as usual. Connections that are recording and not disconnected are unaffected; they first receive service_shutdown and can finish the recording. A start sent while the service is shutting down receives service_shutdown; no recording is started, and the connection is then closed. Reconnect later and start the recording again.

Content during the disconnect

  • The few seconds of audio during the disconnect are not recovered (handled with "pause" semantics): the transcript timeline continues seamlessly, but the disconnected interval does not appear in the recording, and no extra charge applies for the disconnect. The minute in progress when the connection dropped was already deducted when it began and is not refunded; if the next minute begins during the disconnect, it is deducted once the resume succeeds.
  • Conversation mode: a half-sentence already spoken but not yet finalized at the moment of disconnect is flushed as a provisional result.
  • Single / speaker-diarization mode: a not-yet-finalized sentence at the moment of disconnect is not guaranteed to be preserved.

Trust model and security

  • api_key is the trust boundary: the server uses a "last connection wins (takeover)" policy. Any client holding the same api_key and the corresponding resume_token can take over the session within the grace period (the existing connection is silently closed). Therefore do not share a single api_key across trust domains; a B2B team sharing a key is treated as one trust boundary.
  • Transmitted only over WSS (TLS): both resume_token and Ticket are sent via Sec-WebSocket-Protocol; any plaintext (non-TLS) ingress is treated as credential leakage.

Reconnect implementation example (JavaScript)

The following wraps the full session-resume logic: store the token, detect disconnects, reconnect within the grace period, and branch on resume_ok vs failure. Audio capture (startAudioStream, etc.) depends on your format (PCM / WebM); for WebM you must reopen with a brand-new container after resume_ok.

class ResumableVASClient {
  // getTicket: async () => string (always obtain a NEW single-use Ticket per reconnect)
  constructor({ wsUrl, getTicket, buildStartMessage }) {
    this.wsUrl = wsUrl;
    this.getTicket = getTicket;
    this.buildStartMessage = buildStartMessage; // () => start message object
    this.resumeToken = null;
    this.graceSeconds = 45;
    this.resumeDeadline = 0; // grace deadline (epoch ms); 0 = not resuming
    this.closedByUser = false;
  }

  async start() {
    this.closedByUser = false;
    await this._connect(false);
  }

  stop() {
    this.closedByUser = true;
    this.resumeToken = null;
    this.ws?.close(1000, 'client stop');
  }

  async _connect(isResume) {
    const ticket = await this.getTicket(); // single-use; a new one on every reconnect
    const protocols = isResume && this.resumeToken
      ? [`ticket.${ticket}`, `resume.${this.resumeToken}`]
      : [`ticket.${ticket}`];

    this.ws = new WebSocket(this.wsUrl, protocols);
    this.ws.onopen = () => {
      if (!isResume) this.ws.send(JSON.stringify(this.buildStartMessage())); // fresh start
      // isResume: wait for resume_ok before reopening the audio stream (see onmessage)
    };
    // staleness guard: ignore late events from a stale socket (e.target !== this.ws) to avoid
    // double connections after a resume fallback, or a stale onclose corrupting the new state.
    this.ws.onmessage = (e) => { if (e.target === this.ws) this._onMessage(JSON.parse(e.data)); };
    this.ws.onclose = (e) => { if (e.target === this.ws) this._onClose(); };
    this.ws.onerror = () => {}; // onclose handles it uniformly
  }

  _onMessage(msg) {
    const d = msg.data || {};
    switch (d.action) {
      case 'session_started':
        this.resumeToken = d.resume_token;            // store the token
        this.graceSeconds = d.resume_grace_seconds || 45;
        this.resumeDeadline = 0;                       // connection healthy; clear resume state
        this.startAudioStream();                       // start sending audio
        break;
      case 'resume_ok':
        // resume succeeded: reopen the audio stream as after start (WebM must send a new container)
        this.resumeDeadline = 0;
        this.restartAudioStream();
        break;
      default:
        if (msg.type === 'error' && String(d.error_code).startsWith('resume_')) {
          // resume failed -> clear token, fall back to a fresh start (_connect obtains a new ticket)
          this.resumeToken = null;
          this.resumeDeadline = 0;
          // the user has ended the recording (for example, stop() was called during a resume): do not start a new one
          if (this.closedByUser) return;
          this._connect(false).catch((err) => this.onFatal(err));
        }
        // other events (origin / translation, etc.) update the UI by sid
    }
  }

  _onClose() {
    if (this.closedByUser || !this.resumeToken) return this.onFatal?.();
    // First disconnect: set the grace deadline
    if (this.resumeDeadline === 0) {
      this.resumeDeadline = Date.now() + this.graceSeconds * 1000;
    }
    if (Date.now() < this.resumeDeadline) {
      // Within grace: attempt resume (retry with backoff while still within the window)
      this._connect(true).catch(() => {
        if (Date.now() < this.resumeDeadline) setTimeout(() => this._onClose(), 1000);
        else this.onResumeExhausted?.();
      });
    } else {
      this.onResumeExhausted?.(); // grace expired; give up resuming
    }
  }

  // Implement these per your audio format:
  startAudioStream() {}    // begin capturing and sending audio frames
  restartAudioStream() {}  // reopen the audio stream (WebM = new container; PCM = continue)
  onResumeExhausted() {}   // could not resume within grace (treat as session end / notify user)
  onFatal() {}             // unrecoverable
}

Key points: (1) obtain a new Ticket on every reconnect; (2) reopen the audio stream only after resume_ok; (3) on a resume_* error, fall back to a fresh start, but not if the user has ended the recording; (4) WebM must send a new container on reopen.


Version: V1.24.1 Last Updated: 2026-10-07

Copyright © 2026