Appendix

Changelog

V1.24.1

2026-10-07

Fix: Viewer API Can Be Called from Any Domain

  • When your viewer page was on your own domain, the browser blocked broadcast information lookups (GET /api/v1/viewer/broadcasts/{token}) and password verification (POST /api/v1/viewer/broadcasts/{token}/verify). Starting with this version, both endpoints can be called directly from a web page on any domain.
  • Do not send credentials with the request (such as credentials: 'include'); otherwise the browser blocks the response. When a 429 is returned, the Retry-After header can be read.
  • The viewer caption stream already accepted connections from any domain and is unaffected; put viewer_access_token in the URL query parameter.
  • See Viewer REST API — Calling from a Browser.

Documentation: Terminology Troubleshooting Example Corrected

  • The "Terminology has no effect" row under Troubleshooting used zh-CN registered with zh-TW recognition as its example, but the two are in the same language family and the terminology still takes effect. The example now uses languages from different families.
  • Added: a language code that is not recognized (e.g. zh) also keeps terminology from taking effect; inactive_languages and unknown_languages in config_updated list these languages.
  • See Terminology Guide — Troubleshooting.

V1.24.0

2026-10-07

Behavior change: Broadcasts Always Use Real-Time Translation

  • Broadcasts (broadcast) always use real-time translation; realtime_translation in start is treated as true whether it is sent as false or not sent at all. Billing is unchanged and still uses the real-time translation rate.
  • Both the host and viewers first receive interim translations of a sentence (is_final: false) and then the finalized version (is_final: true). Match them by sid and language and replace what is displayed; to show only finalized results, skip events with is_final: false. A host whose start sent realtime_translation as false or omitted it now also receives interim results.
  • A viewer who joins partway through receives only the most recent few sentences of original text and finalized translations; interim results are not resent.
  • See Broadcast Viewer SSE — Translation Result, Voice Translation Actions — Broadcast Mode Description, and Pricing — Broadcast Billing.

Behavior change: start Checks the Length of the Name and Summary Language

  • If name in start is longer than 60 characters after trimming leading and trailing whitespace, or summary_language is longer than 20 characters, invalid_parameter is returned and the recording does not start; details.field names the field.
  • Previously, start could return auth_service_error, or the recording could start but no task_complete was received after it ended.
  • When summary_mode is custom, a summary_prompt or summary_prompt_slug containing only whitespace counts as not provided: start returns summary_mode_field_mismatch; set_summary keeps the previous setting and returns summary_mode_field_mismatch only if none was ever set.
  • See Voice Translation Actions — start and Error Codes — General Errors.

Behavior change: start Checks Fields That Accept Only Fixed Values

  • If options.speaking_speed, options.profanity_handling, conversation_mode, or tts_mode in start has a value outside the accepted list, invalid_parameter is returned and the recording does not start; details.field names the field and details.valid_values lists the accepted values.
  • Values are case-sensitive; for example, "Async" for tts_mode is rejected. If a field is not sent or is an empty string, the default is used as before.
  • These fields are checked for every recording type: an invalid conversation_mode on a recording that is not a conversation, or an invalid tts_mode with TTS turned off, is also rejected.
  • Previously, an invalid value was treated as the default, while the settings in resume_ok after reconnecting showed the value originally sent.
  • When switching with tts_mode during a recording, a value other than sync or async returns invalid_data; the mode does not change and no tts_mode_changed is sent. Previously, tts_mode_changed was sent even though the mode did not actually change.
  • See Voice Translation Actions — start and Voice Translation Actions — tts_mode.

Behavior change: start_speaking Replies with status on Success

  • In manual conversation mode, a successful start_speaking now returns a status message; previously nothing was returned on success.
  • If it is called again while already speaking, the final result of the previous sentence is sent as usual, followed by the same status; no error is returned.
  • The floating subtitle stream also receives this status message, without the status field; clients that rely on the status field to detect pause, resume, and stop are unaffected.
  • See Voice Translation Actions — start_speaking and Floating Subtitle SSE — status.

Behavior change: The Recording Completion Event May Arrive Twice

  • In rare cases, recording.completed for the same recording is delivered twice, with a different delivery_id each time, so checking delivery_id alone will not catch it.
  • Also check data.task_id together with event to decide whether it has already been processed. A repeated delivery does not deduct credits twice.
  • See Webhook Callback Guide — Idempotent Processing.

Fix: Manual Speaking and Speaker Language Changes in Conversations Without speakers

  • When start for a conversation (conversation) omits speakers, users 1 and 2 map to the two transcription_languages in order, and start_speaking and set_speaker_language work normally; previously they returned conversation_invalid_speaker.
  • speakers is now marked optional and must still contain exactly 2 entries when provided; conversation_missing_speakers in the error code reference is marked as no longer returned.
  • For conversations started without speakers, the settings in resume_ok after reconnecting now also include speaker_language_map.
  • See Real-Time Voice Translation Guide — Conversation Mode and Voice Translation Actions — start.

Documentation: When Conversation Mode Translates

  • Conversation mode translates only finalized sentences; realtime_translation has no effect in conversation mode, and the result is the same whether or not you send it. The field was removed from the conversation request examples.
  • See Core Concepts — Real-Time vs. Non-Real-Time Translation.

Documentation: When Terminology Settings Take Effect

  • Updates to terminology, fuzzy_correction, and translation_dict during a recording take effect from the next utterance after they are sent; the utterance being recognized at that moment is not guaranteed to use them. The guide previously described this as the next recognition turn boundary or as immediate.
  • See Terminology Guide — When Settings Take Effect.

Documentation: Javanese and Wu Chinese Support Translation

  • jv-ID (Javanese) and wuu-CN (Wu Chinese) can be used as translation sources and targets, so all 145 languages support translation; the exception stating they did not support translation was removed.
  • See Supported Languages List — Language Overview.

Documentation: tts_stop Reply Text Corrected

Documentation: Broadcast Channel Name and Summary Language Rules

  • When creating or updating a broadcast, summary_language must be a language code from the supported language list; otherwise the request returns 422 validation_failed.
  • The channel name cannot be changed after creation; a name field in an update is ignored without an error. The recording name does not reuse the channel name either.
  • An update with no updatable field returns 200 and nothing changes; the REST API reference previously said at least one field must be provided.
  • See Broadcast Feature Guide — Step 1: Create a Broadcast and Broadcast Feature Guide — Step 7: Update Settings Dynamically.

V1.23.0

2026-10-05

Behavior change: Rate Limit for Floating Subtitle Audience Token Requests

  • Audience Token requests for floating subtitles are now limited to 30 per minute per recording, counting all viewers together, including requests with an invalid share link.
  • Requests are no longer counted separately by source, and a source that sends invalid share links repeatedly is no longer temporarily blocked from requesting Tokens for other recordings.
  • When the limit is exceeded, HTTP 429 (too_many_requests) is still returned with a Retry-After header.
  • See Subtitle Feed Token API — Audience Sharing.

Behavior change: Sending start While the Service Is Shutting Down

  • A start sent while the service is shutting down (for example, for a maintenance update) receives service_shutdown; no new recording is started, and the connection is then closed.
  • Recordings already in progress are not affected and can be finished as usual.
  • After receiving it, reconnect later and send start again.
  • See Error Codes — Session Errors.

Documentation: Webhook Guide Corrections

  • Corrected the retry count: each notification is sent at most 5 times (the first delivery plus 4 retries), at intervals of 10, 30, 90, and 270 seconds; the guide previously said 5 retries and 6 attempts in total.
  • Added that 3xx and 4xx responses are also retried and that redirects are not followed; to have a notification that was marked as failed sent again, contact us.
  • Corrected when a Webhook Secret is generated: you can choose to generate it when creating an API Key, and one is generated automatically when you set a URL for a key without a secret (secrets generated through batch settings are not shown in plaintext).
  • Corrected the verification request sent when you save a URL: the system checks only whether the receiving endpoint responds with 2xx within 8 seconds and does not check the signature; if your endpoint verifies signatures, set the secret on it first.
  • Corrected the repeat rule for credit balance notifications: each type is sent at most once every 24 hours per account and is re-evaluated after credits are added; also added that these two notifications are sent to every active API Key in the account that has a webhook_url, one copy per key.
  • Added an "API Key for Each Notification" section: real-time recordings cannot specify a callback_url and notify the API Key used to obtain the connection ticket; for broadcasts, the key that started the broadcast applies once the recording is finalized; no notifications are sent after an API Key is deleted.
  • Added credit.low and credit.exhausted to the event table in Core Concepts.
  • See the Webhook Callback Guide.

V1.22.1

2026-10-02

Fix: Numbers with Thousands Separators or Decimal Points in Real-Time Translation

  • Numbers that contain thousands separators or decimal points (for example 1,200, 12,500, 1,200,000,000, and 3.14159) now stay intact in real-time translations and broadcast announcements, instead of keeping only their last part (for example, "1,200" translated as "200").
  • These numbers keep the notation used in the original.

V1.22.0

2026-10-02

Behavior change: Translating Sentences That Mix Languages

  • Words that are not in the translation language are always translated into it, even a single word.
  • Names, companies, brands, products, places, all-caps abbreviations, model numbers, code, URLs, and email addresses are kept as in the original; words specified in the translation dictionary follow the dictionary.
  • Parts of the original that are already in the translation language are left unchanged.
  • Summary translation follows the same rule; numbers in summary translations keep the original's notation.

Behavior change: Other Language Columns in the Translation Dictionary Are Also Matched

  • Words an entry has in other language columns are also matched against the original, so some words now follow the dictionary. For example, when the recognition language is Chinese, Japanese sentences are matched against the Japanese column; English words mixed into Chinese, Japanese, and similar sentences are matched against the English column.
  • When translating into the recognition language (for example, Chinese to Chinese), sessions whose dictionary has other languages filled in also use the dictionary.
  • See Terminology Guide — Other Language Columns Are Also Matched and Real-Time Voice Translation Guide — Translation Result.

V1.21.1

2026-10-01

Documentation: Response fields in multichannel shared mode

  • Completed the description of multi-channel shared mode in responses. API behavior is unchanged:
    • settings.channel_mode in resume_ok can be per_channel or shared.
    • In shared mode, the channels[] entries in session_started and resume_ok do not carry transcription_languages; the session's languages are given by settings.transcription_languages.
    • After resuming from a disconnect, each channel's preparing status is reported in the resume_ok snapshot rather than by a separate event; when a channel starts producing text, it sends its own channel_status event (ready, reason: "reconnect").
  • See Connection and Authentication — the settings object for field details.

V1.21.0

2026-10-01

New: Multi-Channel Shared Mode (Taking Turns)

Multi-channel adds channel_mode: "shared": the microphones take turns speaking and share one recognition stream, and the speaker is labeled by which channel the audio came in on. Suited to situations where only one person speaks at a time. See WebSocket API.

  • Billing is always counted as 1 channel: 1.5 credits per minute (speech recognition 1.0 + speaker diarization 0.5); adding or removing channels does not change the rate. stt_stream_count in channel_status is always 1.
  • Client requirements: send only one channel at a time, keep sending silence on the current channel when nobody is speaking, and use pcm only. Languages are shared across the session: channels do not carry transcription_languages, and the language cannot be changed during the recording (returns channel_language_not_allowed).
  • The first channel in channels[] cannot be removed (returns channel_remove_not_allowed).
  • The preparing / ready / error status of the other channels follows the first channel.
  • Retroactive transcription after resuming from a pause covers the whole last 60 seconds, for all channels together.
  • Known limitation: when the gap between speakers is shorter than about 0.8 seconds, the words of the two people may be merged into one sentence labeled with only one speaker.
  • Environments where shared is not enabled still return invalid_channel_mode.

Documentation: Multi-Channel Audio File Length Cap

  • Clarified: the audio saved for one multi-channel recording has a total cap (about 70 minutes with 8 channels). Once the cap is reached, the audio file and the recording duration stop at that point, while the transcript and credit charges continue as usual.

V1.20.0

2026-09-30

New: API Key Self-Service Endpoints

Three new read-only endpoints return only this API Key's own data and remain available when credits run out. See API Key Self-Service.

  • GET /api/v1/me/credit-lots: the credit lots that charges draw on, with remaining credit and expiry time.
  • GET /api/v1/me/usage: charge history, one entry per recording / broadcast and per import / summary / retranslation, paginated.
  • GET /api/v1/me/key: the API Key's name, expiry date, monthly credit limit and this month's spend, concurrency limit, webhook URL, and whether a source IP restriction is configured.

V1.19.0

2026-09-30

New: Query Specific Tasks in the Task List

  • GET /api/v1/tasks adds the task_ids[] parameter (UUIDs, 1–100 entries) and returns only the listed tasks that belong to the current account. IDs that do not exist or do not belong to the current account are skipped.
  • status still applies; without it, only completed tasks are returned. To query regardless of status, add status=all.
  • The response format is unchanged.

Fix: Task List Failing for Accounts with Many Tasks

  • GET /api/v1/tasks now returns normally for accounts with a large number of tasks.

For earlier versions, see Changelog Archive.


Version: V1.24.1 Last Updated: 2026-10-07

Copyright © 2026