Changelog Archive
Release notes for V1.18 and earlier. For the latest changes, see Changelog.
V1.18.6
2026-09-29
Fix: Complete Sentences in Broadcast Subtitles
- Viewers who join a broadcast mid-way receive history subtitles as complete sentences.
- For viewers already watching, a completed subtitle no longer reverts to incomplete.
V1.18.5
2026-09-29
Fix: Import Duration and Billing for Some Transcription Languages
The following applies to file imports whose first transcription language is one of: wuu-CN, yue-CN, zh-CN-shandong, zh-CN-sichuan, ar-DZ, ar-MA, ar-TN, ar-YE, as-IN, gu-IN, kn-IN, mr-IN, or-IN, pa-IN, fr-BE, fr-CA, fr-CH, it-CH, nl-BE, bs-BA, km-KH, ne-NP, si-LK, sw-TZ.
- Duration and billing follow the actual length of the audio file. In earlier versions they could differ from the actual length, and the transcript could be incorrect.
- Short files of 1 to 5 seconds complete normally.
- Files that cannot be read end with
failed(error_codeimport_stt_failed) and are not billed.
V1.18.4
2026-09-29
Fix: Recordings Without Audio Are Not Billed
- A recording that receives no usable audio at all is not billed, and any credits already deducted are refunded automatically.
V1.18.3
2026-09-28
Fix: Handling Audio That Cannot Be Recognized
- If audio cannot be recognized during a recording, the recording continues and
audio_decode_failedis returned; send a new audio stream to recover. - After resuming a disconnected session, always send a new audio stream.
- Time during which audio cannot be recognized is not billed.
V1.18.2
2026-09-28
Documentation Updates
- The Broadcast Guide and the API reference now correctly describe
share_url: it is not an openable viewer page, so do not share it with viewers directly. The viewer page is provided by your application and connects with thetokenas described in the Viewer-Side Flow. API behavior is unchanged.
V1.18.1
2026-09-27
Fix: profanity_handling Now Takes Effect
options.profanity_handling in the WebSocket start action previously had no effect: profanity in real-time transcripts was always masked. Starting with this version, all three values take effect:
| Value | Behavior |
|---|---|
mask (default) | Profanity is masked with * |
remove | Profanity is removed from the transcript |
show | The original text is shown |
Translations are produced from the processed transcript. This setting applies to real-time recordings only; transcripts of imported files keep the original text.
Fix: Untranslatable Sentences Now Return an Error Code
Some sentences containing abuse or threats could previously receive a translation unrelated to the original. Starting with this version, such sentences return llm_content_filtered instead, the same as the existing "content cannot be translated" handling. This applies to real-time recordings, broadcasts, file imports, and retranslation; during retranslation, such sentences are not billed.
Documentation Updates
- In Summary Customization, the values of
profanity_handlingare corrected tomask/remove/show(previously written asremovedandraw).
V1.18.0
2026-09-27
Behavior Change: Recordings End Automatically After a Long Silence
A real-time recording now ends automatically after 15 minutes (by default) without detected speech, so a recording someone forgot to stop does not keep being billed.
- About 2 minutes before the end, the warning
stt_silence_warningis sent (the recording continues, and recognized text or resuming restarts the count); at the limit,stt_silence_timeoutis sent and the recording wraps up as usual withtask_complete. - The count does not run while paused, while waiting to resume after a disconnect, or for broadcasts.
- Use
silenceTimeoutSecondsinstartto adjust the threshold or turn it off (0); see Automatic End After a Long Silence.
Behavior Change: Recordings That Received No Audio
A real-time recording that received no audio at all previously did not send task_complete, and was marked as failed only about 2 hours later. Starting with this version, you receive task_complete with noAudio: true, the recording is marked failed immediately, and recording.failed is sent with the new failure_source value no_audio. The recording has no transcript or audio file.
Behavior Change: Time Limit for the Broadcast Standby Phase
When the accumulated standby time reaches the limit (30 minutes by default), the session ends automatically: the warning broadcast_standby_warning is sent about 2 minutes before, and broadcast_standby_timeout and status: "ended" at the limit. No recording exists during standby, and it is not billed. See Standby Time Limit.
Behavior Change: Broadcast Parameter Checks in start
broadcast_phase accepts only lowercase standby and live (an empty string is treated as live), and broadcast_token is allowed only with type: "broadcast". Otherwise invalid_parameter is returned and the recording does not start.
Behavior Change: Host Goes Live Again After a Disconnect (Takeover)
When the host goes live again within the grace period with the same broadcast_token, broadcast_recording_ready on the new connection carries a new task_id, and the broadcast is split into two recordings, each notified separately. Previously both parts shared one recording, and the later content overwrote the earlier content.
Behavior Change: Segmented Full Retranslation
Long transcripts can be retranslated in segments with segmented=1; without it, a transcript that is too long returns HTTP 422 retranslate_segmentation_required before the stream starts, with no charge. See Retranslate SSE.
Behavior Change: Import Failures and Failure Descriptions
import.failedis sent only once per failed import, and the status staysprocessingwhile it is being processed.failedis a final state: you never receiveimport.completedfor the same import afterward. An import that does not start processing for a long time is markedfailed(PROCESSING_TIMEOUT).error_messageof a failed import anderrorinrecording.failedare now fixed descriptions without internal messages; error codes are unchanged.
Behavior Change: Viewer-Side Rate Limits
- The broadcast viewer page and password verification are now limited per channel (about twice the channel's maximum number of viewers per minute, with a minimum of 200) instead of 10 per minute per source IP; many viewers joining at once from the same network no longer tend to receive 429.
- Too many incorrect passwords cause a temporary lock (30 from the same source within 5 minutes, or 100 for the same channel within 5 minutes), during which even the correct password returns 429; too many lookups of nonexistent channels from the same source also return 429 temporarily. See Rate Limits.
- Floating subtitle audience Token requests are now limited to 30 per minute for the same recording from the same source; too many invalid share links from the same source return 429 temporarily.
Behavior Change: Import Recognition Modes and File Formats
- An upload with
recognition_modeset tomulti_languageormulti_channelnow returns HTTP 422import_recognition_mode_unsupported(data.details.supportedModeslists the available modes). Previouslymulti_languagewas accepted and failed later during processing, andmulti_channelreturnedvalidation_failed. - When the file extension does not match the actual content (for example a
.mp3name with WAV content), the file is now processed according to its content instead of failing.
Behavior Change: Summary Generation Completeness
Applies to Ad-hoc Summary and Regenerate Summary:
- When generation takes too long, the part completed so far is returned within the processing time limit, with
truncated: trueindone, and billed as usual; previously neitherdonenorerrormight arrive. - When the stream stalls or does not end normally,
error(sse_summary_regeneration_failed) is sent instead, and nothing is saved or billed; previously an incomplete summary could be delivered as if it were complete.
Other Behavior Changes
- When speakers are merged, the target speaker's existing sentences whose display name changes because of the merge (for example, when the source speaker's custom name is carried over to the target) also switch to the merged name and are included in
affected_sids; this applies to both REST and WebSocket. - When the service restarts or is updated, sessions waiting to be resumed are finalized right away and saved as usual; resuming afterward returns
resume_token_invalid. - A broadcast's
current_recording_idhas a value only while live, and it is the recording of this go-live; it isnullduring standby or when not live.
Fixes
- Viewers who joined after a broadcast went live could see content from the standby phase.
- A broadcast that went from standby to live could be billed up to 1 extra minute; billing now starts at the moment it goes live.
- For a broadcast that went live directly,
broadcast_recording_readyoccasionally arrived beforesession_started. - The full audio file sometimes lacked a short piece at the very end of the recording.
- For recordings using
audio_format: "webm", audio after switching recording devices mid-session is no longer lost. - After a brief service disruption, recordings in progress are billed as usual; usage during the disruption is also billed once service recovers.
New Error Codes
| Error code | Type | Description |
|---|---|---|
stt_silence_warning | WebSocket, warning | About to end for lack of speech; the recording continues |
broadcast_standby_warning | WebSocket, warning | The standby phase is about to reach its time limit; the broadcast continues |
broadcast_standby_timeout | WebSocket, fatal | The standby phase reached its time limit and the session ended |
retranslate_segmentation_required | REST, HTTP 422 | The transcript is too long for a full retranslation; use segmented retranslation |
import_recognition_mode_unsupported | REST, HTTP 422 | This recognition mode is not supported for file imports; use single or multi_speaker |
Client Recommendations
- Recognize
stt_silence_warningandbroadcast_standby_warning: the recording continues; it has not ended. - For sessions during which nobody may speak for a long time (for example a meeting room), send
silenceTimeoutSeconds: 0instart. - When
task_completecarriesnoAudio: true, do not read the transcript. - After a broadcast takeover, use the recording ID from
broadcast_recording_readyon the new connection. - For full retranslation of long transcripts, send
segmented=1and continue with the next segment according totruncatedindone. - When a viewer endpoint returns 429, wait for the number of seconds in
Retry-Afterbefore retrying; do not retry immediately. - For imports, use only
singleormulti_speakerforrecognition_mode. - When a summary stream ends with
error, discard the fragments already received; whendonecarriestruncated, tell the user the summary is incomplete.
Documentation Updates
- Error Codes now lists the five error codes that can appear when an import fails, as well as
invalid_parameter. - Corrected the multi-channel note that every channel must keep sending audio even while silent: a single silent channel does not end the recording.
- Added the error code
too_many_requeststo the 429 responses of the viewer endpoints and the floating subtitle audience Token request, andinvalid_shareto the latter's 403 response.
V1.17.0
2026-09-24
New: Summary Translation Endpoint
POST /api/v1/sse/summary/translate translates summary text supplied in the request into a specified language. It is not tied to a recording and the result is not saved. It is billed at 0.1 credits per 200 characters. See Summary Translation.
Fix: Retranslate Summary
- After a summary was regenerated in another language, retranslating it could return the original text without translating it. This is fixed.
- Retranslating a long summary could return only the first part of the translation. This is fixed.
Fix: Multi-Channel Session Resume
In multi-channel mode, after a session resume, the channels stopped producing recognition results and the audio after the resume was not saved. This is fixed.
Fix: Sentences in Progress Lost in Two-Way Translation
- In manual mode, the last part that was still being recognized when the user stopped speaking could be lost. This is fixed; it is now merged into the sentence's final result.
- In manual mode, the part being spoken could be lost after a speaking speed change or a reconnect, and the whole sentence could be lost if the recording ended while the user was speaking. This is fixed.
- In automatic mode, a sentence in progress when the recording was paused could be overwritten by the next sentence after resuming. This is fixed; it is now sent with
is_final: trueat the moment of pausing. - For languages that separate words with spaces (such as English), parts merged in manual mode had no space between them. This is fixed.
- In single-speaker and two-way translation modes, the same sentence could occasionally appear twice, or a discarded sentence could reappear, right after an operation such as a speaking speed change. This is fixed.
Fix: The Last Sentence Could Be Missing When a Recording Is Stopped
When stop was sent, a sentence that had not finished being recognized was not included in the transcript. This is fixed: the server now waits for the recognizer to finish the last sentence (usually about 1 second, at most about 3 seconds) before finalizing; if it does not arrive in time, the last interim result shown is used as that sentence. pause in two-way translation and stop_speaking in manual mode wait in the same way.
Fix: task_complete Could Be Missed When Finalizing Took Longer
When finalizing after stop (title, summary, upload) took longer, the connection could be closed before task_complete was sent. This is fixed. The title and the summary are now generated at the same time, which shortens the wait between stop and task_complete.
Fix: Recordings That Had Just Ended Could Be Incomplete During Maintenance
During service maintenance, for a recording that had just ended or just disconnected, the transcript, the full audio file and the completion notification sometimes could not be processed in time; the recording stayed in processing and was later marked as failed. This is fixed: the service now finishes this work before it shuts down.
Fix: Stream Errors No Longer Include Internal System Messages
When an unexpected error occurs in Import Progress SSE or TTS SSE Streaming, details.original_error is now always Service error, consistent with the other streaming endpoints; the error code is unchanged.
Fix: Error Messages in TTS SSE Streaming
In TTS SSE Streaming, the message of tts_error events previously carried an internal identifier; it is now the standard English message (for example, TTS synthesis failed). The error code in error is unchanged. The documented error code for a sentence with no translation is corrected to tts_translation_not_found, the value actually sent; behavior is unchanged.
Behavior Change: Connections When the Service Is Shutting Down
- After
service_shutdownis sent, a connection that is not recording is closed about 2 seconds later. - A connection that is recording can finish the recording, still receives
task_completeafterstop, and is closed after that.
Behavior Change: Stop, Pause and Concurrent Recording Slots
status: "ended"is always sent beforetask_complete.- The concurrent recording slot is now released when
task_completeis sent: you can start the next recording as soon as you receive it. - Responses to
stop, topausein two-way translation, and tostop_speakingin manual mode are delayed by about 1 second (at most about 3 seconds).
Client Recommendations: Detecting Completion
task_completewaits until the title and summary have been generated, which usually takes a few seconds to tens of seconds and up to about 8 minutes. To detect completion, we recommend also supporting the Webhook or a REST query rather than relying on this event alone.
Behavior Change: Retranslate Summary
- When the translation is incomplete,
donecarriestruncated: true. - A request has a processing limit of about 230 seconds, and an
erroris sent if translation pauses for more than 60 seconds. A timeout now reportsdetails.original_errorasTranslation timed outinstead ofService error. - For a long summary, the stream may sometimes deliver a larger block at once, with a pause of a few seconds between parts.
Client Recommendations: Summary Translation
- When you receive
truncated: true, the translation is incomplete; do not save it as a complete result. - If you want to render the translation character by character, smooth it on the client side.
Documentation Updates
- The
reasondescription of segment_discarded now notes an exception: after a session resume,channel_statususesreconnect, whilesegment_discardedusesresumed. - start_speaking corrected: calling it again while already speaking ends the previous sentence and starts a new one, instead of returning an error.
- segment_discarded now notes that two-way translation manual mode does not receive this event either.
- task_complete now documents its order relative to
status: "ended", that it waits for the title and summary, when the concurrency slot is released, and how to detect completion. pause,stopandstop_speakingnow document the wait for the last sentence.- Pricing now clarifies billing during pauses and disconnects: billing continues while paused, and the time spent waiting to resume after a disconnect is not billed.
- Pricing now clarifies how multi-channel channels are counted: a channel added mid-minute and disabled before the next deduction is counted once, and a channel already deducted at the start of the minute is not billed again after it is disabled.
V1.16.5
2026-09-21
New: Summary Events Carry summary_language
The following events now include summary_language (additive):
| Event | Value |
|---|---|
init_summary of History | The language of the saved summary; null when there is no summary |
done of Regenerate Summary | The language used this time; the first transcription language when language is not sent |
done of Ad-hoc Summary | The language used this time; zh-TW when language is not sent |
To determine the language of the summary itself, use summary_language of init_summary (when there is no summary, it may differ from the field of the same name in init_metadata).
Fix: Summary Language of Imported Tasks
For imported tasks with a summary, summary_language in the task list and in init_metadata was null. This is fixed, and existing tasks have been corrected.
Documentation Updates
summary_languageofinit_metadata: when no summary language is specified for a recording, the value is usually the first transcription language (in two-way mode withactive_languagespecified, it is that language), notnull. Only tasks without a summary may havenull.- Retranslated summaries (
GET /api/v1/sse/retranslate/summary/{taskId}) are not saved. To keep a summary in another language, use the POST endpoint of Regenerate Summary (billed).
V1.16.4
2026-09-17
Import Quota Preflight Now Reports "Daily Usage Exhausted"
POST /api/v1/imports/check-quota previously only looked at credit and plan entitlement, so it still returned allowed: true when the day's usage was exhausted and only the upload was rejected. It now returns allowed: false with reason: "plan_daily_limit_reached", matching the error code the upload returns so the two map directly.
Note: When allowed is false, topping up does not solve every case — exhausted daily usage resets tomorrow, and a plan without import requires an upgrade. Tailor the message to reason; see the Import API reference.
Insufficient-Credit Errors Add budget_scope
The details of auth_quota_exceeded and stt_quota_exceeded now carry budget_scope, stating whose credit the accompanying remaining_budget refers to. Additive; existing fields are unchanged.
Important: remaining_budget is the credit available to the API key that made the request, not the end user's balance. If you call this service on behalf of other users, do not show this number directly to them. See the Error Code Reference for what each value means.
V1.16.3
2026-09-17
Billing Fix: Channels That Are Not Working Are No Longer Counted
In multi-channel recording, if a channel's audio cannot be received at the moment a minute is settled — and the channel therefore produces no transcript — that minute no longer counts this channel; previously it was still billed. See Pricing for how the count is determined, including the first minute and channels that send no audio.
Behavior Change: A Broadcast Token May Only Be Used by the Account That Owns It
Starting a broadcast with a broadcast_token that belongs to another account returns broadcast_token_invalid. Different API keys under the same account are unaffected.
New Error Code: task_already_processing (409)
While the same task is still being processed, POST /api/v1/tasks/{taskId}/retry returns 409 task_already_processing and the task status is unchanged; send the same request again later. Previously this situation returned 200 without the task being reprocessed.
Fix: Summary Language After Regenerating a Summary
When regenerating a summary without specifying language, the summary language recorded in the task data was not updated and did not match the actual summary content. The two are now consistent.
V1.16.2
2026-09-16
Documentation update: how to obtain the broadcast task_id
For a broadcast, the authoritative task_id is the one returned by the broadcast_recording_ready event;
the task_id in session_started is an initial value for the connection stage, and querying with it returns
no matching record. You receive this event either way a broadcast goes live: after the standby phase, or by
starting directly with broadcast_phase: "live" (the default). Behavior has not changed; this version only
completes the description.
V1.16.1
2026-09-16
Documentation update: how billed minutes are counted at the boundary
For real-time recording, each minute boundary has a 1-second grace period: a 60.5-second recording counts as 1 minute, and a 61.5-second recording counts as 2 minutes. Billing behavior has not changed; this version only completes the description.
V1.16.0
2026-09-15
Behavior change: real-time recording is now deducted at the start of each minute
Real-time recording (broadcasts excluded) used to be deducted after each minute was used, with any partial minute counted as a full minute when the recording ended. Starting with this version, credits are deducted at the start of each minute:
- The first minute is deducted when the recording starts, and
session_startedarrives once that deduction completes - Each following minute is deducted as it begins
- Nothing extra is deducted when the recording stops
The number of billed minutes per recording still follows "any partial minute counts as a full minute". Previously, some recordings were billed one minute less than this rule; this version always applies the rule, so a recording of the same length may be billed one minute more than before.
When features are turned on or off, or channels or translation languages are added or removed during a recording, the rate changes from the next minute. For how channels are counted in multi-channel mode, see the Pricing Guide.
Behavior change: a recording cannot start or continue when the available credits do not cover one minute
- When starting a recording: the recording does not start, and you receive
auth_quota_exceededwithdetails.remaining_budgetset to the available credit as of the most recent settlement. The connection stays open; after topping up, you can sendstartagain right away - During a recording: the recording ends before the next minute begins, and you receive
stt_quota_exceeded; everything recorded so far is still saved - The minute that cannot be covered is not deducted, and the remaining credits are not deducted either
If the available credits cannot be confirmed at the moment, start returns auth_service_error and the
recording does not start; please try again later.
New: summary_insufficient_credit error code
In the following cases no summary is generated, and summary_error is sent with the error code
summary_insufficient_credit:
- The recording ended because the available credits ran out
- At a normal end, the available credits cannot cover the summary fee
The transcript and audio are still saved; after topping up, you can get a summary by regenerating it.
Behavior change: unlimited-plan usage limits are checked before each minute begins
- The single-recording limit, the usage-hour thresholds, and the daily hard limit are all checked before each minute begins; once one is reached, the next minute does not start and is not counted toward usage
- If today's usage has already reached the limit when a recording starts,
daily_limit_reachedis returned right away and the recording does not start - Inside a restriction window, each stretch of continuous recording now lasts exactly the interval set by the plan (previously one minute longer)
- When the same API Key runs several recordings at once, after the periodic-stop threshold is reached, each recording stops before its own next minute begins (previously only one of them was stopped each time)
- The used minutes in
GET /api/v1/me/planinclude a minute as soon as it begins
Behavior change: when credit.exhausted is sent
credit.exhausted is also sent when a real-time recording cannot start or continue because the available
credits do not cover one minute and the account balance cannot cover that minute; in this case balance
may be greater than 0.0. When only an API Key's dedicated allocation is insufficient while the account
balance can still cover the minute, it is not sent, as before.
Fix: problems when recording several times on the same connection
After one recording ended and the next one started on the same connection, the following could happen:
- The next recording's summary fee also counted the transcript characters of earlier recordings, and a summary fee could be charged even when that recording produced no summary
- If an earlier recording had text-to-speech turned on, the next recording was billed at the rate including text-to-speech even when it was off
- If an earlier recording used custom summary settings, the next recording could fail to be saved when it ended
- Multi-channel recording: if an earlier recording's audio was not saved completely, the next recording could also fail to be saved
- Floating subtitles did not start after starting again
- An insufficient-credit or plan-limit notice for an earlier recording stopped the next recording
All of these are fixed in this version.
Fix: starting again right after a recording was ended could leave the previous recording unsaved
After a recording ended because of insufficient credits, a plan restriction, a usage limit, or a long period
without sound, sending start right away on the same connection could leave the previous recording
unsaved, and the new recording could end up without a transcript.
Starting with this version, the new recording starts only after the previous one finishes processing (after
status: "ended"); depending on the summary length, session_started may be delayed by a few seconds to a
few tens of seconds. Starting a recording on a new connection is not affected.
Behavior changes
- Reconnecting with a resume token after sending
stopand beforestatus: "ended"arrives (including while a recording that was ended is still being processed) returnsresume_unavailable; a recording that has already ended is not resumed - Sending
broadcast_go_livewhile the previous recording is still being processed returnssession_not_started details.remaining_budgetinauth_quota_exceededandstt_quota_exceededis now the available credit as of the most recent settlement- When a recording ends because of insufficient credits,
stt_quota_exceededis sent only once - After a recording has ended, no more insufficient-credit or plan-limit errors for that recording are sent
- When a recording is ended while its connection is interrupted, the API Key's concurrent-recording slot is released right away (previously it was released only after about 3 minutes, during which
concurrency_limit_reachedcould be returned)
Client recommendations
- Starting a recording may return
auth_quota_exceededbecause of insufficient credits: the connection stays open, so just sendstartagain after topping up - Listen for
summary_insufficient_creditinsummary_errorand tell the user that no summary was generated because of insufficient credits - When starting again on the same connection after a recording was ended,
session_startedmay be delayed; do not treat this as a timeout - While the previous recording is still being processed,
pongmay also be delayed; allow enough waiting time and do not treat the connection as dropped just because aponghas not arrived yet - On
resume_unavailable, if the user has already ended the recording, do not automatically start a new one
Documentation updates
- The WebSocket
starterror table now listsauth_quota_exceeded,daily_limit_reached, andauth_service_error; the HTTP status ofbroadcast_token_invalidis corrected to 401 to match the Error Code Reference - The
broadcast_go_liveerror table now listssession_not_started - Session resume example: when the user has ended the recording, a failed resume no longer starts a new recording automatically
- The heartbeat description now notes that
pongmay be delayed while the previous recording is being processed - The description of
stt_quota_exceededfor audio imports is corrected to insufficient available credits for the import - The event examples in the summary customization guide now use the actual message format (
typeisvoice-translation, and events are distinguished bydata.action)
V1.15.10
2026-09-12
Fix: retranslation reported success with no translation when content could not be translated
When retranslation (full text, single sentence, or summary) encountered content that could not be translated, it reported success with a blank translation; full-text retranslation was also billed and overwrote the translation that language already had.
This version reports failure instead: the previous translation is kept, and failed sentences are neither counted toward the updated total nor billed.
Fix: results could be lost or billed incorrectly when several operations ran at once
When several operations ran against the same recording at once (for example, retranslating into different languages at the same time, or retranslating while renaming a speaker), a result could be billed without being kept, or billing fields could appear even though nothing was consumed.
This version reports failure in these cases, and the request is not billed.
Behavior changes
- Retranslation may return
llm_content_filtered(severitywarning); other translation failures returnsse_translation_failed, and summary retranslation returnssse_summary_translation_failed - When the transcript cannot be saved, retranslation, saving a summary, and speaker operations all return
storage_upload_failedand stop;donewill not follow - When several changes are made to the same transcript at once, the later one receives
transcript_revision_conflict(409) donecarriescharacters_billed/charged/billedonly when the request was actually billed- When a speaker is specified by a display name that maps to more than one speaker,
speaker_name_duplicateis returned - Translations that contain only whitespace no longer appear in exports or speech synthesis
Client recommendations
- Retrying
llm_content_filteredwill not help; revise the original text instead. If you filter events byseverity, do not filter outwarning - A
transcript_revision_conflictmeans another change is in progress; simply retry - Use
billedto decide whether a request was billed; integrators that bill their own end users in particular should not infer it from other fields
V1.15.9
2026-09-10
Fix: After a connection dropped and recovered, a speaker's name could be applied to someone else
With speaker identification (multi-speaker mode), if the connection dropped and recovered mid-recording, a name assigned earlier could be applied to a different speaker, and a merge configured earlier could combine two different people into one. Neither produced any notice.
This version fixes it: speakers appearing for the first time after a recovery receive new IDs, carrying over neither the earlier names nor the earlier merges.
Behavior change: speaker IDs are no longer guaranteed to start at 1 or run consecutively
Speaker IDs are unique within a recording, but are not guaranteed to start at 1 or to be consecutive. After a connection recovers, and when a broadcast moves from standby to live, speakers appearing afterwards receive IDs that have not been used before.
rename_speaker and merge_speakers apply only to the speakers identified so far. After a recovery,
merge again once you confirm it is the same person; when the target has no name yet, merging carries
the source's name across.
When a broadcast goes live, speaker names set during standby are not kept.
Single-speaker, conversation, and multi-channel modes are unaffected.
Client recommendations
- If your application assumes IDs start at 1, run consecutively, or stay the same for one person throughout, use the
speaker_idyou receive instead - To continue with the same speaker after a recovery, use
merge_speakers; callingrename_speakerwith the earlier name returnsspeaker_name_duplicate - On the broadcast host side, clear your accumulated speaker list when the phase changes to
live
V1.15.8
2026-09-09
Fix: In conversation mode, a sentence in progress disappeared when switching modes
In conversation mode, switching between automatic and manual mode — or pressing the talk button — while a sentence was still being spoken made that sentence reappear under a new segment number. The original number then had no further events, and the sentence did not appear in the transcript; the sentence spoken next could also end up without its own translation.
This version fixes it: a sentence in progress now stays on its original number and ends normally; only the sentence spoken afterwards gets a new number.
New: segment_discarded event
When the speaking speed is changed, a channel's language is changed, a broadcast goes live, a session is resumed, or recording resumes after a pause, a sentence that happens to be in progress cannot be kept. Previously none of these produced a notification, so that sentence never received an ending signal.
From this version, segment_discarded is sent to state that the segment number will receive
no further events, with reason identifying the operation; multi-channel recordings also
carry channel_id. Floating-subtitle viewers receive it as well.
Behavior change
Re-sending the conversation mode that is already in effect no longer affects a sentence in progress; the mode change notification is still returned.
Client recommendations
- On
segment_discarded, clear the matching segment from any unfinished state - Automatic conversation mode does not receive this event; there, a sentence in progress ends normally and its content is kept in the transcript
- When
set_speaking_speed_failedis returned, that operation may still have sentsegment_discarded - See segment_discarded for fields and the full
reasonvalue set
V1.15.7
2026-09-08
Bug fix: Re-translation sometimes returned something other than a translation
When re-translating a completed transcript, a shorter sentence could come back as explanatory text about that sentence, and be stored as its translation.
Fixed in this release. Re-translating an affected segment once returns the correct translation; per-segment re-translation is not billed.
V1.15.6
2026-09-06
Fix: The 202 response when uploading audio returned null for progress
progress in the successful response from POST /api/v1/imports should have been 0,
but was actually null. Querying the same import through
GET /api/v1/imports/{importId} returned 0, so the two disagreed.
Impact: a client that validates the response structure strictly reports a
successful upload as a failure, so it never captures import_id and never receives the
completion notification. Each retry creates another import that runs to completion and
consumes credits again.
From this version, progress is always 0 in the 202 response, matching both the
documentation and the query endpoint.
Fix: The 201 response when creating a broadcast returned null for three counter fields
peak_viewers, total_viewers, and duration_ms in the successful response from
POST /api/v1/broadcasts should have been 0, but were actually null. Querying the same
broadcast through GET /api/v1/broadcasts or GET /api/v1/broadcasts/{broadcastId}
returned 0.
The symptom is the same as the previous item: a client that validates the response structure strictly reports a successful creation as a failure. Retrying leaves several usable broadcasts behind, each holding a token.
From this version these three fields are always 0 in the 201 response. Creating a
broadcast does not consume credits, and existing broadcasts and their statistics are
unaffected.
Documentation Update: task_id is present as soon as the upload succeeds
task_id was previously described as being populated only after processing completes, and
both the POST and GET response examples showed it as null. In fact task_id already
exists at the moment the upload succeeds (202); there is no need to wait for processing.
Clients can navigate to the task as soon as they receive the 202 response, rather than
polling to obtain task_id. The examples and field descriptions have been corrected.
Client Recommendations
No integration changes are required. If you relaxed your validation because progress or the
three broadcast counter fields were null, you can require an integer again; if you waited
for processing to finish before using task_id, you can now use it from the successful
upload onward.
V1.15.5
2026-09-04
Fix: Speaking a language outside the transcription list could return the untranslated original
When multiple transcription languages were specified but the speaker used a language not on the list, one of the translation languages could show the original text instead of a translation.
For example, with transcription languages zh-TW, id-ID and translation languages en-US,
zh-TW, id-ID: when English was spoken, the Indonesian translation showed the English original.
This has been fixed.
Behavior change: with multiple transcription languages, every translation is produced by translating
When multiple transcription languages are specified, every translation language is now produced by actually translating, even when the source language matches that translation language. Such translations may differ slightly in wording from the original (meaning preserved).
Sessions with a single transcription language are unaffected: when the source language matches a translation language, that translation still reuses the original text.
Billing is unchanged: translation is charged by the number of translation languages, regardless of whether a given translation was actually produced by translating.
Client recommendations
No integration changes are required. Include every language that will actually be spoken in the transcription language list.
V1.15.4
2026-09-04
Fix: Some target languages returned the untranslated original text
When multiple transcription languages were specified, one of the target languages could show the original text instead of a translation.
For example, with transcription languages en-US, zh-TW, id-ID and the same three
translation languages: when Indonesian was spoken, the English translation showed the original
Indonesian text, while the Chinese translation was correct. This has been fixed.
Fix: Language and speaker tagging in conversation mode
With certain language pairs (for example English and Indonesian), a sentence could be tagged as the other language, and the speaker label and transcript language tag were wrong as a result. This has been fixed.
Fix: Translations when an import declares multiple transcription languages
When an import declared multiple transcription languages and the translation targets included the first of those languages, that language could return the untranslated original text. This has been fixed.
Behavior change
With certain language combinations (for example English and Indonesian specified together), translations whose source and target language are the same may now differ slightly in wording from the original, while the meaning is preserved. Previously such translations reused the original text verbatim. Sessions with a single transcription language are unaffected.
Client recommendations
No integration changes are required. If your application checks whether the original and the translated text are identical in order to decide whether a sentence was translated, please note the change above.
V1.15.3
2026-09-03
Fix: Fuzzy correction could corrupt text that was already correct
When an incorrect variant in fuzzy_correction was written in Latin script, it was previously
applied to part of a longer word, corrupting text that was already correct.
For example, with the term Remote View and Emote View listed as an incorrect variant
(with case_insensitive enabled): in a transcript that already read Remote View, the
emote View inside it was treated as a match and replaced, producing RRemote View —
an extra character at the start.
From this version, incorrect variants written in Latin script are only applied to whole words. Matching for Chinese, Japanese, and Korean is unchanged.
Fix: Punctuation or spacing immediately after a correction could disappear
Punctuation or a space immediately following a variant was also consumed, with no indication:
| Transcript content | Previous result |
|---|---|
open Remote Vue, then quit | open Remote View then quit (comma lost) |
open Remote Vue. | open Remote View (period lost) |
open Remote Vue now | open Remote Viewnow (two words run together) |
Chinese punctuation was affected in the same way. Both are now preserved.
Behavior Changes
Common-word protection in homophone correction is no longer bypassed by adjacent spaces. Previously, when a misrecognized word happened to have spaces around it, common-word protection did not take effect and the word was still corrected; without those spaces it was not. The result for the same word depended on whether spaces happened to surround it.
Both cases are now consistent and common-word protection always applies. A small number of common
words that used to be corrected are therefore no longer corrected — if you do need them corrected,
list them explicitly in fuzzy_correction.
Leading and trailing spaces in glossary entries are now ignored. Previously an entry with leading or trailing spaces produced inconsistent results, and in some cases spacing in the transcript was removed. This has been fixed, which also means two entries that differ only by surrounding spaces are treated as the same entry.
Documentation Updates
- The terminology guide now states that matching in Latin script operates on whole words
- Corrected an example in the terminology guide that used an incorrect variant with a trailing space (such an entry has no effect)
Client Recommendations
No changes are required to existing integrations. Glossary configuration format, limits, error
codes, and conflict codes are all unchanged. If your glossary registers a shorter Latin-script
form and relies on it covering a longer word (such as wafer covering wafers), register both
forms separately.
V1.15.2
2026-09-03
Behavior Changes
Sending stop before the recording has started, or after it has already ended, now returns a session_not_started error.
Previously these two cases produced no response at all, leaving the client to wait for a timeout
before discovering that the action had no effect.
The most common case is sending stop twice: the second one returns this error rather than
another success response. task_complete is still sent only once, after the first successful stop.
If your client already treats a duplicate stop as harmless, simply ignore this error.
Documentation Updates
- Added the error code section for
stop - Filled in error codes that were missing from the tables for the multi-channel
add_channel,remove_channel, andset_channel_languageactions
V1.15.1
2026-09-03
Fixes
- When
configreturnsconfig_empty, the error message now states explicitly that an empty object{}is not treated as a provided setting, and shows the correct way to clear a glossary - Documentation now notes that a rejected
configreturnstype: "error"rather thanconfig_updated. Clients must listen forerroras well, or the request appears to go unanswered
Error codes and response structures are unchanged.
V1.15.0
2026-09-03
Fix: Summaries of long meetings were cut off
When generating a summary for a longer meeting, the summary could stop partway through with an
unfinished sentence — and no error was reported. The stream ended normally and the done event
was still sent, so there was no way for a client to tell that the content was incomplete.
The longer the recording, the more likely this was to happen.
From this release, meetings of an hour or more produce a complete structured summary.
This is distinct from the existing
summary_fallback_level/summary_dropped_segmentsbehavior. Those indicate that part of the input was omitted (the sentences are complete, but a section of the meeting is missing); what this release fixes is the summary itself stopping halfway.
New
Incomplete summaries are now flagged. The done event of both the ad-hoc summary endpoint
(POST /api/v1/sse/summary) and Regenerate Summary
(GET / POST /api/v1/sse/regenerate/summary/{taskId}) gains an optional truncated field: it is
true when the summary could not be produced in full, and the field is absent entirely when the
summary is complete.
This is a purely additive, backward-compatible field — existing integrations can ignore it. Even if an exceptionally long summary still cannot be produced in full, the client can now see that.
Behavior Changes
The character limit for summary input has been raised from 100,000 to 200,000. Affected endpoints:
contentonPOST /api/v1/sse/summary- The transcript for
GET/POST /api/v1/sse/regenerate/summary/{taskId} - The summary generated automatically when a recording ends
Previously, a long recording whose transcript exceeded 100,000 characters received
summary_text_too_long and got no summary at all. The new limit is aligned with the longest
audio duration file import accepts (10 hours), so "the import succeeds but there is no summary" no
longer happens.
Glossary coverage in summaries has been widened. Previously, when the transcript was long, only part of your glossary was applied to the summary, with no indication that this had happened. This release substantially relaxes that. Real-time translation behavior is unchanged.
Free regeneration for ad-hoc summaries is now capped. Re-sending the same idempotency_key
with identical content previously regenerated the summary for free with no limit on how many times;
from this release, exceeding the cap returns 429 too_many_requests.
Free retries exist for the "already charged but delivery failed" case, where one or two attempts is
normally enough. To generate again, use a new idempotency_key (charged as a new request). The
initial paid request does not consume the retry budget, and neither does a failure that produced no
content. The budget is counted over a rolling 24-hour window.
Client Recommendations
- Important: Review your client-side timeout settings. Summaries are now generated in full, so long meetings take longer than before. We recommend allowing at least five minutes for summary-related stream connections. A shorter timeout will disconnect before the summary finishes
- Long recordings that previously failed with
summary_text_too_longcan now be regenerated - The
doneevent gains an optionaltruncatedfield (see "New" above). Existing integrations can ignore it, but we recommend checking it so you know whether the summary you received is complete - Ad-hoc summary requests that reuse a fixed
idempotency_keymay now receive a 429. The normal "one new key per request" usage is unaffected
Documentation Corrections
- The "Character and Length Limits" table in
guides/summary-customization.mddid not list the limit for summary input text. It has been added
V1.14.2
2026-09-02
Fix: Viewers could not connect to the subtitle stream of password-protected broadcast channels
After entering the correct password and obtaining a viewer_access_token, the viewer's connection to the subtitle stream was still rejected (HTTP 401, broadcast_password_required). With this release, the connection succeeds once the password is verified.
- Switching networks between password verification and stream connection (for example mobile data vs. Wi-Fi, or a corporate multi-line egress) no longer causes the connection to be rejected
- The token remains valid for 24 hours; the response fields of the
verifyendpoint are unchanged
Fix: tts_languages in viewer channel info was not returned correctly
GET /api/v1/viewer/broadcasts/{token} could return an empty tts_languages array in some environments, leading clients to assume the channel offered no voice playback. It now correctly returns the voice languages enabled by the host. The field format is unchanged.
The definition of tts_languages is also clarified: it lists the languages for which the host has enabled voice playback, and is an empty array when the channel is not live. The previous wording, "otherwise from the default settings", did not correspond to any actual behavior.
Behavior change: Error responses from the viewer subtitle stream are readable cross-origin
When the viewer subtitle stream (GET /broadcast/{token}/text) rejects a connection (401 / 404 / 503, etc.), the response now carries the same cross-origin headers as a successful response. Previously, cross-origin clients received a cross-origin error in these cases and could not read the response at all.
The shape of the failure changes as well: a cross-origin rejection was previously classified as a network error, which per specification causes EventSource to keep reconnecting. It now terminates the connection outright (readyState becomes CLOSED) with no further retries. If your client wraps EventSource in its own reconnect logic, confirm that its behavior in this case is what you expect.
Client recommendation: A browser-native EventSource never exposes the response body to the page, so it still surfaces only onerror. To show viewers a specific reason (such as broadcast_password_required or broadcast_not_ready), read the stream with fetch, or call GET /api/v1/viewer/broadcasts/{token} before connecting to check the channel status and whether a password is required.
Documentation
- Viewer API: removed the network-address binding note from
viewer_access_token; corrected thetts_languagesfield description - Error Codes: corrected the HTTP status of
broadcast_password_requiredfrom 422 to 401 (the actual behavior has always been 401)
References: Viewer API, Viewer Subtitle Stream
V1.14.1
2026-09-02
Documentation: Request Size Limit for Glossary Validation
POST /api/v1/glossary/validate did not previously state the actual request size limit, leaving integrators unable to assess ahead of time whether their glossary would fit. This release documents:
- The default limit: 2 MB (2,097,152 bytes). Reading
details.max_bytesfrom the response is still the recommended approach — do not hard-code a value - How to split a glossary that exceeds the limit.
terminologyandfuzzy_correctionmust be sent in the same request;translation_dictmay be sent on its own
No API behavior changed. This release is documentation only.
Note: If you already split requests yourself, review this. When
terminologyandfuzzy_correctionare sent as two separate requests, some conflicts are not detected, and the response gives no indication of it.
Reference: Glossary Validation API
V1.14.0
2026-09-01
Added: Glossary Validation API
Adds POST /api/v1/glossary/validate, so a glossary management UI can check a glossary for format problems and internal conflicts before saving.
Most glossary problems produce no error message at all — they simply yield unexpected results, silently, during recording or translation. The worst of them is "an incorrect variant that is also the correct term of another rule": a user saying that word normally has it rewritten into something else. This endpoint surfaces situations like that at the moment the user presses Save.
Characteristics
- Free: no points are deducted, no task or recording is created, and no settings are written
- All three blocks — terminology, fuzzy correction, translation dictionary — are optional; whatever you send is what gets validated
- Every problem comes with the index of the entry (language code plus index), so a management UI can flag the item directly
- The response carries no human-readable text; integrators compose their own sentences from the conflict codes, entirely in the language and phrasing they choose
- Rate limited to 120 requests per minute per API Key, as a separate quota counted independently of the other endpoints
- Can be called directly from a browser (cross-origin requests are supported). Note: This means the API key is present in the browser, and the same key can also create recordings, import audio, and consume credits — acceptable for an internal admin tool, but for a public page route the call through your own backend instead
Note: The host is the realtime service domain (the same domain as
wss://, with the scheme swapped forhttps://), which may differ from the host of the other REST endpoints.
The eight detectable conflicts
| Code | Severity | Meaning |
|---|---|---|
variant_shadows_term | error | An incorrect variant is also another rule's correct term or a registered term; normal text gets broken |
dict_duplicate_source | error | The same source is registered more than once under one target language; only one takes effect |
variant_ambiguous | warning | The same incorrect variant maps to different correct terms |
case_flag_conflict | warning | The case_insensitive settings for one incorrect variant are inconsistent |
variant_equals_term | warning | An incorrect variant is identical to the correct term of its own rule and has no effect |
duplicate_term | warning | The same term is registered more than once within the same language family |
boost_out_of_range | warning | A term's boost value is outside the valid range and is adjusted automatically |
homophone_conflict | warning | Two correct terms or registered terms share a pronunciation |
The homophone check can be turned off
Of the eight conflicts, the homophone check is the slowest; the other seven return almost immediately. For live feedback while the user edits, send check_homophones: false; run the full check when they actually save.
Important: The homophone check applies to Chinese only, and may not always complete.
homophones_checkedin the response is what tells these two situations apart: when it isfalse, the check did not complete, which does not mean there are no conflicts. Have your save gate check this field as well.
Added: error codes
| Error code | HTTP | Description |
|---|---|---|
config_too_many_languages | 400 | There are more language codes than allowed (details.field names the field at fault) |
config_payload_too_large | 413 | The request body exceeds the size limit |
Behavior clarification: config is all-or-nothing
Earlier documentation described config as "processed in order, aborting on failure, with the blocks ahead of the failure already applied". That is not the actual behavior: if any one of the three glossary blocks fails, none of the three is applied — the settings stay exactly as they were.
For integrators this is a change for the better: when an error comes back you can be certain that nothing changed, with no half-applied state to worry about. The affected documentation has been corrected throughout.
Behavior clarification: variant collisions do not guarantee a winner
Earlier documentation said that when one incorrect variant maps to several rules, a fixed order decided the winner. Testing did not match that description, so it is now stated plainly: which rule actually takes effect is not guaranteed — do not rely on any ordering, registration order included. This is consistent with the new variant_ambiguous conflict code, which deliberately does not report which rule wins.
Behavior change: the TTS voice catalog now lists only usable languages
GET /api/v1/tts/voices previously listed all 154 locales, but 13 of them cannot actually be used — a TTS target language must be one of the translation output languages, and translation output languages are limited to the transcription language list, which those 13 are not in. Customers could look up voices for them and then have translation_languages reject the same code.
Starting with this version the voice catalog lists only the languages that are actually usable, and the published language and voice counts have been changed to the usable figures:
| Previously (including unusable) | This version (actually usable) | |
|---|---|---|
| TTS languages | 154 | 141 |
| TTS voices | 325 | 302 |
Possible impact:
- Querying
GET /api/v1/tts/voicesfor one of those 13 locales now returns an empty voice list (voices were listed before, but they could not be used for synthesis anyway) - Querying
GET /api/v1/tts/voices/{voiceName}/samplefor one of their voices now returns 404tts_voice_not_found, consistent with the catalog not listing them
All other languages are unaffected.
The 13 affected locales: bn-BD, ta-LK, ta-MY, ta-SG, ur-PK, su-ID, sr-Latn-RS, iu-Cans-CA, iu-Latn-CA, and four zh-CN dialects (zh-CN-henan, zh-CN-guangxi, zh-CN-liaoning, zh-CN-shaanxi).
Behavior change: the audio import duration limit is now actually enforced
The "File limits" table has always listed a maximum duration of 10 hours and a minimum of 1 second, but until now only the optional import pre-check endpoint verified them — a file uploaded directly was subject to no duration limit at all. In practice a 500 MB low-bitrate file can exceed 17 hours and would be processed in full and billed.
Starting with this version, a file outside the range ends the import with a failed event whose error_code is the newly added import_duration_out_of_range.
Possible impact: if you currently import files longer than 10 hours (or shorter than 1 second), those imports will start failing. Use audio within the allowed length, or split it and import in parts.
An over-long file uploads successfully first and then ends with
failed. To find out before uploading, call the import pre-check endpoint first.
Behavior change: speaking_rate is now clamped to its documented range
The TTS speaking_rate has always been documented with a valid range of 0.5 to 2.0, but until now values outside the range took effect as sent. Starting with this version, values outside the range are adjusted to the nearest bound (below 0.5 becomes 0.5; above 2.0 becomes 2.0), so the actual behavior matches the documented range.
Possible impact: if you currently send a speaking_rate above 2.0 (or between 0 and 0.5), the speaking rate becomes the boundary value. Values within the range are unaffected.
Behavior clarification: an out-of-range boost behaves differently on the two paths
A term's boost has a valid range of 0.5 to 5.0. When a value falls outside it, live recording and broadcast adjust it into range automatically with no notice, while file import rejects it outright (HTTP 422). Earlier documentation did not record that the same glossary produces different outcomes on the two paths.
Documentation corrections
This release also corrects the following discrepancies with actual behavior:
| Location | Correction |
|---|---|
| Ad-hoc summary | Authentication corrected to Header X-API-Key only (query string was documented in error); removed auth_insufficient_credit, which never occurs |
| Summary regeneration (SSE) | Removed five parameter-validation error codes that are never actually sent, replaced by a note that a parameter validation failure carries only message and no error_code; content filtering actually returns sse_summary_regeneration_failed, not llm_content_filtered |
| Retranslation (SSE) | The error event's context corrected to sse; added the request_id and timestamp fields |
| Single-sentence retranslation (SSE) | Added auth_insufficient_credit (402) |
config action | Added config_invalid_entry to the error code table (the most commonly triggered code, previously unlisted) |
| WebSocket connection | Added "Maximum Size of a Single Message" — exceeding it closes the connection outright with no error message; send a large glossary across several config messages |
| WebSocket events | Added the speakers_auto_merged event (previously documented only on the viewer side) |
| Broadcast API examples | Corrected a voice name that is not on the supported list |
| TTS audio streaming | Fixed a layout problem caused by an error-code row placed inside the wrong table |
| Documentation home page | Corrected the endpoint counts for speaker editing and summary templates |
start action | Documented that glossaries sent inside start are ignored (previously noted only in the error code reference) |
| Error code reference | auth_account_expired is marked as not currently returned by any endpoint; tts_invalid_voice is marked as returned only by the realtime voice channel |
Documentation updates
- The "Important Notes" section of the Terminology Guide now covers six previously undocumented situations (an incorrect variant shadowing a correct term, duplicate terms under one language, boost values adjusted automatically, an incorrect variant equal to its own correct term, a duplicate source word in the dictionary, terms that share a pronunciation), each marked with its matching conflict code
- The same guide gains a "Validating a Glossary Before Saving" section
- The Error Code Reference corrects the HTTP status description for
invalid_json(endpoints on the realtime service domain return 400, the rest return 422) and adds howinvalid_actionis used on REST endpoints - The Terminology Guide adds the language-family matching rule, the behavior of an empty language code, and the valid range of
boost
Behavior notes
- Only aggregate results are returned when the count is over the limit: when the glossary's total entry count exceeds the limit, the response contains the count problem alone — individual content problems are no longer listed and no conflict detection is run. Bring the count back within the limit and validate again.
- The number of reported items is capped: conflicts and format problems each have a reporting cap, and
truncatedistruewhen it is exceeded. Fix the items listed and validate again to see the rest.
Client recommendations
- Have the glossary management UI call this endpoint once before saving; when
validisfalse, block the save and show the offending entries - The recommended condition for the save gate is
valid === true && homophones_checked === true - On
auth_service_error(HTTP 500), retry later; do not swap the API Key — this has nothing to do with the key itself - Passing this endpoint's validation does not guarantee that audio import will accept the same glossary — import applies stricter rules
Reference
V1.13.1
2026-09-01
Added: completion events now report consumption
The done event of the following endpoints carries three new fields, so integrators no longer
need to derive usage themselves:
| Endpoint | Notes |
|---|---|
GET /api/v1/sse/retranslate/{taskId} | Full-text retranslation |
GET/POST /api/v1/sse/regenerate/summary/{taskId} | Summary regeneration (both preview and save are billed) |
POST /api/v1/sse/summary | Ad-hoc summary (adds charged alongside the existing characters_billed) |
{
"characters_billed": 12700,
"charged": "1.3",
"billed": true
}
characters_billed: character count used as the billing basischarged: points consumed by this operation, calculated from the rate. This value reflects usage — usage already covered by an unlimited plan is still reported herebilled: whether the request incurred consumption
Always use billed to determine whether a request was billed (billed only when billed is
true). When the fields appear differs by endpoint:
- Full-text retranslation and summary regeneration: when nothing was consumed (for example, when generation fails), all three fields are absent
- Ad-hoc summary:
characters_billedandchargedare always present, whilebilledisfalsefor free retries (same idempotency key with identical request content) and empty generation results — in those cases the request was not billed
Do not use the presence of charged as the criterion: on a free ad-hoc retry charged still has a
value, and billing on that basis would overcharge.
Endpoints that are not billed do not carry these fields: summary retranslation
(/retranslate/summary/{taskId}) and single-sentence retranslation
(/recordings/{taskId}/entries/{sid}/retranslate) are not billed, and their done events are unchanged.
Client Recommendations
These are purely additive fields; existing integrations are unaffected and require no changes.
Integrators that call this API on behalf of end users and bill them separately can use charged
directly rather than deriving it from character counts and rates, which avoids the two sides
computing from different bases. If you adopt it, use billed === true as the single criterion for
whether a request was billed — that one rule covers every endpoint listed above, with no per-endpoint
exceptions.
V1.13.0
2026-09-01
Changed: terminology now corrects homophone misspellings directly
Once terminology is set, spans in the transcript that sound the same or nearly the same but are written differently are corrected back to the spelling of the term. You no longer need to list the possible misspellings in advance.
| Term | Appears in transcript as | Corrected to |
|---|---|---|
| 紡拓會 | 訪拓會 | 紡拓會 |
| 語者分離 | 語這分離, 與者分離 | 語者分離 |
| 晶圓 | 晶園 | 晶圓 |
This is a behavior change, and it runs in both directions: transcripts and translations for existing integrations will start to differ.
- Homophone and near-homophone misspellings that were previously missed are now corrected (wider coverage)
- Conversely, a misspelling that is itself an ordinary word is no longer corrected (see "Common-word protection" below). If you relied on such corrections, list that misspelling explicitly with
fuzzy_correction
Existing recordings are not reprocessed; this applies to recordings created from this release onward.
Applicable languages: Chinese only (Traditional and Simplified are interchangeable, since they share pronunciation). Terms in Japanese, Korean or English do not participate in homophone matching — list their misspellings explicitly with fuzzy_correction.
For mixed Chinese-English terms (such as CVD製程), matching applies only to the Chinese portion; the Latin portion is left unchanged.
Matching range: Both identical and near-identical pronunciations are covered, including accent differences such as jin vs jing (final -n vs -ng) and retroflex vs non-retroflex initials. For the term 晶圓廠, for example, 金圓廠 in the transcript is corrected.
Common-word protection: If the span in the transcript is itself an ordinary word (金元 or 反案, say), it is not changed even when it shares a pronunciation with a term — this prevents normal sentences from being altered. To force such a correction, list the misspelling explicitly with fuzzy_correction: explicitly listed misspellings are not subject to common-word protection.
Changed: incorrect is now optional for Chinese terms in fuzzy_correction
When correct is Chinese (contains Han characters), incorrect may be omitted entirely — the system matches by pronunciation:
{ "fuzzy_correction": { "zh-TW": [{ "correct": "艾思通" }] } }
No misspellings are listed above, yet 愛思通, 愛時通, 愛司東 and 愛似通 are all corrected to 艾思通. Only spellings that sound quite different (愛自動) or have a different number of syllables (愛松) still need to be listed in incorrect.
This is a relaxation and existing integrations are unaffected — rules that carry incorrect behave exactly as before.
Note: Both conditions must hold: the language must be Chinese and correct must contain Han characters. Otherwise incorrect remains required — omitting it in those cases would have no effect at all, and accepting it would leave you believing the setting took. List misspellings explicitly for Japanese, Korean and English.
This applies to file import as well — both paths use exactly the same condition.
Added: config_updated reports terms that share a pronunciation
When two terms in your glossary sound alike (公事包 and 公式包, for example), config_updated carries an extra optional field, homophone_conflicts:
"homophone_conflicts": [
{ "languages": ["zh-TW"], "terms": ["公事包", "公式包"] }
]
languages lists every language that shares the same term index — Chinese regional codes (zh-TW, zh-CN, zh-HK and so on) are treated as one group for glossary matching, so they share a single entry rather than each reporting one.
When the transcript contains a third spelling with the same pronunciation, the system can only correct it to one of them, and which one is not guaranteed. Both terms themselves still work; the only ambiguity is which term an unregistered homophone misspelling is attributed to.
This is a warning, not an error: the glossary is still accepted and config still succeeds. If the pair matters to you, list the misspelling explicitly with fuzzy_correction so it is pinned to the term you want.
It is not present before the recording has started (the language list is not settled yet), the same as inactive_languages.
Fixed: the first segment_uploaded event was missing segment_index
Segment indices start at 0, and a value of 0 previously caused both segment_index
and duration_sec to disappear from the JSON — meaning the first segment_uploaded
of every recording carried no index, and very short segments (duration rounding to 0)
also lost duration_sec.
Both fields are now always present on segment_uploaded. Other events are unaffected
and do not carry these fields.
Removed: three optional fields on config_updated
auto_generated_fuzzy_correctionno longer appears inupdated[], which now carries onlyterminology,fuzzy_correctionandtranslation_dictauto_generated_fuzzy_correction_capped(valuevariant_budget_exhausted) is no longer returnedauto_generated_fuzzy_correction_skipped(valuetimeout) is no longer returned
The shape of the event itself is unchanged, as are the way terminology and fuzzy_correction are configured, their limits, and their error codes.
Client recommendations
- If your code reads any of the three fields above, remove that handling — they no longer appear
fuzzy_correctionis still required in these three cases:- The misspelling is itself an ordinary word and is blocked by common-word protection (晶圓 heard as 金元, for example)
- The misspelling sounds very different from the correct term, such as a foreign brand name recognized as a phonetically unrelated word
- Misspellings in Japanese, Korean or English — those languages do not participate in homophone matching
- If you previously listed many Chinese misspellings in
fuzzy_correctionto cover homophones, most of them can now be dropped. Keeping them is safe, and sometimes better — explicitly listed misspellings are not subject to common-word protection, so this is the only way to guarantee that a particular misspelling is corrected
Reference
V1.12.1
2026-08-27
Changed: the length limit for custom summary prompts has been raised
The custom prompt in custom mode goes from at most 2000 characters to 3000 characters.
This applies to summary_prompt in live recording, prompt in summary regeneration and ad-hoc summaries, and summary_prompt in file import.
The limit counts characters, so one Chinese character counts as one. Exceeding it returns summary_prompt_too_long.
Note: This is a relaxation and existing integrations are unaffected. Note, however, that the longer the prompt and the more instructions it carries, the smaller the share that is reliably applied — a higher cap does not translate into proportionally better results.
V1.12.0
2026-08-27
Translation dictionary is now grouped by language
The translation dictionary format changed from term-first to language-first, matching terminology and fuzzy_correction.
{
"en-US": [
{ "source": "語者分離", "target": "Speaker Diarization" }
],
"ja-JP": [
{ "source": "語者分離", "target": "話者分離" }
]
}
The previous format was a flat array in which all target languages shared one set of source terms, which imposed two limits:
- Languages could not have independent dictionaries. If one language needed 500 terms and another needed 100 entirely different ones, they all had to go into the same array with gaps.
- The entry cap was shared across all languages, so the more languages you used, the fewer terms each one could get.
The previous format is still supported, and existing integrations need no changes. The same dictionary sent in either format produces identical results.
This applies to live recording, broadcast, and file import.
Changed: the dictionary entry cap is now counted per language
Previously up to 3000 entries across all languages; now up to 3000 entries per language. This is a relaxation — no existing configuration is rejected.
Exceeding the cap returns config_too_many_dict_entries; the details in the response identify which language exceeded it.
Note: The total grows with the number of languages. In practice you hit the size limit on a single settings message first (roughly 8 languages at full capacity approaches it), rather than the entry cap itself.
Note: Clearing the whole dictionary is not supported: an empty dictionary is treated as not sending this setting at all. Clearing a single language is supported (send that language an empty array).
Changed: on resume, the dictionary is returned in the format you sent
The translation_dict in resume_ok now matches the format you last sent — send the previous format and you get it back; send the new one and you get the new one.
Changed: cap on dictionary entries carried into a single translation
At most 100 dictionary entries now apply to any single sentence (previously uncapped). Beyond that, longer source terms take priority.
You will not normally hit this: a sentence typically matches a handful of entries. The cap guards against cases where many single-character or very short source terms cause nearly every sentence to match a large set — and the more entries there are, the smaller the share that is reliably honored.
Changed: translations of the same text are now more consistent
Translating the same text twice previously could produce slightly different wording. Results are now stable. Dictionary content and behavior are unchanged, but translations may differ slightly from before.
What did not change: the target language is still matched exactly (an
en-GBtranslation does not apply toen-US); the case rules (per-entrycase_sensitive, defaulting to case-insensitive) are untouched; and the dictionary remains best-effort rather than literal substitution.
V1.11.1
2026-08-26
Fixed: the translation dictionary was not saved, so later re-translation and summary regeneration ignored it
A translation_dict set after start on a live recording did not apply to the following features:
- Re-translating afterwards (whole transcript and single sentence)
- Case-matching behavior when re-translating a single sentence
- The dictionary guidance used when regenerating a summary
Symptom: with the same dictionary, your preferred wording took effect during live translation but not when re-translating afterwards. If you noticed that gap, your observation was correct.
The two are now consistent. This is a behavior change: re-translation and summary regeneration will start applying the dictionary, matching live translation. Existing recordings are not backfilled; this affects recordings created from this release onward.
Broadcast dictionaries had the same problem and are fixed in this release as well.
Added: Terminology Guide
Terminology documentation was previously scattered across sections and presented as field specifications, making the applicable scenario for each block difficult to determine. This release adds a terminology guide covering the stage at which each of the three blocks acts, selection criteria, how to select and write terms, when settings take effect, limits and their error codes, and five important notes.
All three scenarios are covered — live recording, broadcast, and file import. They use the same terminology format; only the transport differs (the three import fields are JSON strings rather than objects).
Location: Feature Guides → Terminology Guide.
Documentation Change: Terminology Documentation Streamlined
Terminology descriptions, field tables, and examples have been streamlined to essential content.
The API contract is unchanged and no client action is required. Existing integrations continue to work, and no error or warning is returned.
To ensure a term is recognized, register it under the correct language code — see the newly added terminology guide for details.
V1.11.0
2026-08-26
Changed: Translation dictionary limit raised to 3000 entries
The translation dictionary limit is raised from 500 to 3000 entries.
Dictionary size no longer affects the cost of each translation.
This is a behavior change: an entry no longer applies unless its source term matches the wording in the sentence. For example, a source term written in the plural (wafers) will not apply to a sentence saying wafer, and a source term of IPEVO Inc. will not apply to a sentence that only says IPEVO. The reverse (shorter source term, longer sentence) still applies.
We recommend setting source terms to the shortest form that will actually be spoken.
Changed: fuzzy-correction limits raised
| Item | Previous limit | New limit |
|---|---|---|
| Rules (across all languages combined) | 3000 | 4000 |
Added: notification when the variant allowance runs out
config_updated gains an optional field, auto_generated_fuzzy_correction_capped. It appears when the terminology list would generate more variants than the remaining allowance, with the value variant_budget_exhausted.
When it appears, automatic variants were partially accepted (the closest matches first) rather than rejected wholesale — the rules that were accepted take effect as usual. To have them all accepted, send fewer variants of your own or reduce the number of terms.
Changed: glossary limits are now defaults, with max in the response as the source of truth
The glossary limits (terminology count, fuzzy-correction rule count, translation-dictionary entries) can now be tuned per environment. The numbers in this documentation are defaults.
Suggested integration change: do not hard-code the limits. The error response's details has always carried both count (what you sent) and max (the limit in force) — read max instead. Existing integrations keep working unchanged; this is a recommendation, not a breaking change.
Changed: the provider value on translation errors
For translation-related errors (llm_content_filtered, translation_service_unavailable, and similar), details.provider now carries the value llm_service.
The previous value is no longer used; llm_service is always returned.
This is a change to a value you can observe: provider has always been debug context inside details, and the documentation has never listed it as an enumerated value to branch on, so most integrations are unaffected. If your code did compare against the previous string (for example to route alerts), switch to the new value — or better, branch on error_code, which is the field designed for that.
Documentation correction: optional fields on config_updated
The field table for config_updated previously listed only updated and terminology_effective. The following have existed since V1.10.1 but appeared only in this changelog; they are now in the event reference:
unknown_languages/inactive_languages/inactive_dict_languages: warnings about glossary language codesauto_generated_fuzzy_correction_skipped(valuetimeout): automatic variant generation did not finish, so this round was not applied and the existing rules are kept; resendingconfighelpsauto_generated_fuzzy_correction, the fourth possible value inupdated[]
It is easy to confuse this with the auto_generated_fuzzy_correction_capped added in this release: capped means "finished but did not fit" (partially in effect, retrying will not help), skipped means "did not finish" (nothing changed, retrying will help).
V1.10.1
2026-08-25
Fixed: Fuzzy-word correction had no effect in live recording
fuzzy_correction configured for live recording (WebSocket) was never actually applied. Sending config returned config_updated and reconnecting returned the rules unchanged, but transcripts, translations, speech synthesis and summaries were all produced without the corrections — while the same glossary worked correctly for file import.
After this release, live recording and file import produce consistent results. If you ever observed "the same glossary works for import but not for live recording", that observation was correct.
This is a behavior change: transcripts and translations for existing integrations will start to differ (they now reflect the corrections).
Changed: Glossary entries apply per language
All three glossary blocks now use the language code to decide where they apply:
- Terminology: applies only to recognition in the language it is registered under. Single-language situations (each multi-channel track, speaker diarization, file import) use only that language's terms; multi-language transcription and conversation mode use the languages declared for the session.
- Fuzzy-word correction: a rule applies only to sentences in the language it is registered under — a rule under
zh-TWwill not alter an English sentence. When the sentence language cannot be determined, rules from every language are applied. - Matching is at language-family granularity:
zh-TW/zh-CN/zh-HKare interchangeable, as areen-US/en-GB.
This is a behavior change, and it is silent: glossary entries registered under a language not used in the session have no effect, and no error is reported. Please make sure the language keys you send match the languages actually used in that recording.
Changed: Subtitles apply corrections at end of sentence
Interim subtitle results are not corrected; corrections are applied when the sentence completes (is_final). Integrators will see the subtitle adjust once at the end of a sentence.
Changed: Glossary limits raised
| Block | Previous | New |
|---|---|---|
| Fuzzy-correction rules | 500 | 3000 |
| Translation dictionary entries | 50 | 500 |
| Terminology | 500 | 500 (unchanged) |
All three are totals across all languages combined.
Note: The number of translation-dictionary entries affects translation cost. Only entries that have a translation for the current target language are carried, so spreading translations across languages reduces the per-request load.
Added: Reporting for glossary language keys
config_updated carries two new optional fields that flag settings which may not take effect:
unknown_languages: language codes that cannot be recognized (for examplezh,chinese). Those glossary entries will not take effect.inactive_languages: valid codes that are not used in this recording.
Both are advisory and do not interrupt the update. inactive_languages is not reported before the recording starts (the language list is not settled yet).
Added: Field validation for glossary entries
Terminology and fuzzy-correction entries are now checked for required fields and length on receipt. Invalid entries return config_invalid_entry, with details identifying the language, index and field. These problems were previously ignored silently.
Removed: Error code config_terminology_locked
This code was never emitted — updating terminology while recording has always been allowed.
Fixed: Terminology not applied for some file-import languages
For some languages, file import used no terminology at all (no error, no warning). After this release these languages behave like the rest.
Fixed: Documentation that did not match actual behavior
- The terminology limit counts the number of terms;
boostdoes not consume capacity (previous documentation stated this two different ways in different sections). boostdoes not affect capacity in live recording; for a few languages in file import it does occupy extra slots.- The translation-dictionary limit was previously described as a recommendation in the guide and as a hard limit in the reference; the two are now consistent.
Fixed: The "exact case match" dictionary setting had no effect in post-processing
The case_sensitive flag on dictionary entries was ignored during post-hoc retranslation (both full-transcript and single-sentence): entries were applied case-insensitively whether or not the flag was set. Live per-sentence translation and summaries have always honored it; only post-processing did not.
After this release all three paths behave consistently for the same recording.
This is a behavior change: if you relied on entries with case_sensitive still being applied during post-hoc retranslation, they will no longer be substituted when the case does not match — which is what the setting was always meant to mean.
Changed: File-import summaries apply the translation dictionary
Summaries produced at the end of a live recording have always used the dictionary to keep proper nouns consistent; file-import summaries did not. They now behave the same.
When the summary is produced in the source language the dictionary has no matching translation, and nothing is injected (same as live recording).
Changed: Broadcast announcements and standby text apply the translation dictionary
Within a single broadcast, per-sentence subtitles were translated using the dictionary while announcements and standby text were not, so the same proper noun could appear two different ways. They are now consistent.
Behaviour is unchanged if you have not configured a translation dictionary.
Changed: Glossary settings are recorded with the recording
Terminology and fuzzy-word correction configured for live recording were not stored with that recording, so there was no way to check afterwards which settings had been in effect (file import has always stored them). The two are now consistent.
What is stored is the configuration you actually sent; rules the system derives automatically are not included.
Recommendations for integrators
- Check the format of your glossary language keys: use full codes (
zh-TW,en-US), not forms such aszhorchinese. After this release, keys in the wrong format fail silently. - Check which language your entries are filed under: if you placed English terms under a Chinese language key and relied on mixed-language speech to pick them up, they no longer apply to purely English sentences.
- Read the new
config_updatedfields:unknown_languagesandinactive_languagesare currently the only signal that surfaces configuration problems early. - Subtitle adjustment is expected: interim results are uncorrected; corrections land at end of sentence.
V1.10.0
2026-08-21
New: multi-channel speaker separation (physical channel separation mode)
Added the multi_channel recognition mode: one recording takes input from multiple physical microphones (up to 8 channels), each channel is bound to a single transcription language, and speaker identity is determined directly by the channel — no inference from voice characteristics — making it a good fit for settings where every speaker has a dedicated microphone. Available for the transcribe and record types; not available for conversation or broadcasts, and cannot be combined with text-to-speech or multi-speaker diarization. This feature must be enabled for your account before use; when not enabled, start returns invalid_recognition_mode.
- Physical channel separation mode:
startcarriesrecognition_mode: "multi_channel",channel_mode: "per_channel", andchannels[](each entry containschannel_id,speaker_name, andtranscription_languages); the audio format supportspcmonly, and everyaudioframe must carry achannel_id. - Adding and removing channels mid-recording: the new
add_channel/remove_channelactions open or deactivate channels while the recording is in progress; after deactivation, the transcript and audio already produced on that channel are kept. - Per-channel language switching: the new
set_channel_languageaction changes the transcription language of a single channel only — the channel number and speaker identity stay the same, and the transcript remains continuous.switch_languagedoes not apply under multi-channel; use this action instead. - Channel status events: the new
channel_statusevent reports each channel's state (preparing/ready/removed/error) and the current channel count, so integrators can display per-channel readiness. - Catch-up transcription after pause: during a pause the audio keeps being saved but produces no transcript; after resuming, speech from the pause is transcribed into the transcript (timestamps reflect the moment the words were actually spoken). Catch-up transcription is capped at the trailing 60 seconds combined across the whole session; anything beyond that is kept in the audio file only.
- Every sentence in
resultevents and in the transcript carrieschannel_id, andspeaker_idis fixed per channel (formatchannel_{N}).
New error codes: this release adds multi-channel error codes (the channel_* / multichannel_* series); see Error Codes for the full list and descriptions.
Billing: multi-channel adds a per-minute surcharge based on the highest number of channels active within that minute; channel 1 is included in the base rate, and removing a channel takes effect from the next minute. See Pricing.
Full specification: WebSocket - Voice Translation.
Client Recommendations
- Existing integrations are unaffected: recordings that do not use
multi_channelbehave and are billed exactly as before. - Multi-channel requires directional / close-talking microphones with sufficient spacing between them. Crosstalk from unsuitable equipment (one person picked up on several channels) is outside the quality guarantee; perform channel selection on the client side (at any given moment, send only the channel with the strongest signal) or ensure physical isolation.
V1.9.2
2026-08-20
Fix: some target languages were not translated in multi-language transcription
When transcription is configured with more than one language (transcription_languages) and the translation targets overlap with them, the overlapping language could be returned verbatim instead of translated.
Observed behavior (with both transcription and translation set to zh-TW / en-US / ja-JP / ko-KR):
| Source text | zh-TW output (before) | zh-TW output (after) |
|---|---|---|
Wait. | Wait. (verbatim) | translated |
NI hao. | NI hao. (verbatim) | translated |
Help with a. | Help with a. (verbatim) | translated |
After the fix, the source language in multi-language mode is determined per sentence: the original text is kept only when that sentence's actual language matches the target, and everything else is translated. Single-language transcription and conversation mode are unaffected.
Data already affected: the fix applies to new recordings only; transcripts already stored are not re-translated automatically. To repair them, run re-translation on the affected recordings. Re-translation is billed by actual usage.
This also affects origin.language: in multi-language mode the field now reports the language determined for that sentence rather than a fixed value for the whole recording. When no determination can be made, the configured value is kept, so the field is never empty. If your integration reads this field to decide the language, note that it now varies per sentence in multi-language mode.
Fix: re-translating historical transcripts returned the original text
When re-translating a historical recording (GET /api/v1/sse/retranslate/{taskId}), some sentences were returned verbatim instead of translated — most often in older data, or where the language had been determined incorrectly at the time. This is now fixed.
Data already affected: sentences that previously came back untranslated can simply be re-translated again; no extra configuration is required. Re-translation is billed by actual usage.
Fix: some client errors were returned as server errors (500)
Certain request errors previously returned 500 internal_error, leading integrators to treat them as server faults and retry. They now return the correct status code:
| Situation | Before | After |
|---|---|---|
| Path is correct but the HTTP method is not supported | 500 internal_error | 405 method_not_allowed (with an Allow header) |
| Uploaded content exceeds the size the server accepts | 500 internal_error | 413 http_error |
| Service temporarily unavailable (maintenance) | 500 internal_error | 503 |
Any other HTTP-level error keeps its original status code; when that status code has no dedicated error code, http_error is used.
Not found (404), forbidden (403), too many requests (429), validation failures (422 validation_failed) and unauthenticated requests (401) were already correct and are unaffected by this change.
Missing headers also fixed: throttled responses (429) previously omitted Retry-After and X-RateLimit-*, leaving clients unable to tell how long to wait. These are now preserved.
If your integration uses 500 to decide whether to retry, switch to the actual status code — a 4xx means the request itself needs to change, and retrying will not help.
Terminology limit: documentation corrected, and a combined check added for file import
The limit for live recording is unchanged, but the previous documentation did not match the actual behavior. Corrected in this release:
- The 500 limit for live-recording
configcounts term entries across all languages combined, not 500 per language boostconsumes extra capacity only in conversation mode. Single-speaker, multi-language transcription, speaker diarization and file import are unaffected byboost; in conversation mode a term occupiesclamp(round(boost), 1, 5)slots (rounded half up, so1.5becomes 2), and anything beyond the limit does not take effect
File import (behavior change): the effective terminology limit is likewise 500 entries across all languages combined. From this release, a combined total above 500 is rejected at upload time with a 422 stating the actual count; previously only the per-language limit was validated, so such an upload succeeded while the excess terms silently had no effect.
If your multi-language vocabulary exceeds 500 entries in total, trim it before uploading — the entries beyond the limit were never taking effect anyway.
Documentation: case flag for the translation dictionary
The case_sensitive field of translation_dict is now documented (the field itself was already supported; this release only adds the documentation): optional per entry, defaults to false = case-insensitive; when set to true the entry applies only on an exact-case match.
A case-flag comparison section has also been added. fuzzy_correction uses case_insensitive (defaults to false = strict) while translation_dict uses case_sensitive (defaults to false = permissive) — opposite field names, and opposite behavior from the same default value. Getting it wrong produces no error at all, only matching behavior opposite to what you intended.
Documentation: variant collision rules for fuzzy correction
When the same incorrect variant appears in more than one rule (collisions are resolved on incorrect, not correct):
- The case flag resolves to strict wins — if any rule leaves
case_insensitiveoff, that variant is matched with exact case - When several rules map the same variant to different correct terms, which one takes effect is not guaranteed; do not rely on it
Splitting one correct term across several rules with different case settings is therefore a safe and supported pattern, as long as their incorrect variants do not overlap. If the same incorrect variant needs to map to different correct terms, pick one.
Documentation: config is processed in order and aborts on the first failure
terminology, fuzzy_correction and translation_dict are processed in that order. When a block fails validation the request returns an error and aborts, so that block and everything after it is not applied — but any block processed before it is already in effect.
Do not assume the configuration is completely unchanged when you receive an error; fix the problem and resend the complete config. All three blocks replace their previous value wholesale and do not stack on top of earlier settings.
V1.9.1
2026-08-13
New: Ad-hoc Summary endpoint (POST /api/v1/sse/summary)
Added an SSE endpoint that generates a summary from text supplied in the request, not tied to any recording. It is intended for content the server does not have — for example, the full transcript of several merged recordings, or a transcript edited by the user. The result is only streamed back to the client and is never stored.
- Authentication: only the
X-API-Keyheader is accepted (no API keys in the query string); authentication failures return real 401/403. - Billing: 0.1 credits per 1,000
contentcharacters (same rate as summaries); billed only on successful generation. - Duplicate-request guarantee:
idempotency_keyis required and only valid within a single API key; the system compares the entire request (content plus every summary parameter) to decide whether a call is a retry. Retrying with the same identifier + an identical request is not billed again (but regenerates); the same identifier + any differing field returns 409; failures do not claim the identifier. - Error contract: before the stream starts, real HTTP status codes are returned (401/403 / 422 / 404 / 400 / 402 / 409), unlike the "200 +
errorevent" convention of other SSE endpoints.
Note: an earlier endpoint once existed at the same path (removed in V1.8.0). This endpoint is a brand-new contract — authentication, billing, and duplicate-request rules all differ, so do not reuse old integration code.
New error code (see Error Codes – Summary Errors):
| Error code | HTTP | Scenario |
|---|---|---|
summary_idempotency_key_conflict | 409 | The same idempotency_key was already used with different content |
Full specification: Ad-hoc Summary SSE.
Client Recommendations
- Use this endpoint when regenerating a summary from merged or edited full text; keep using Regenerate Summary for template or output-language changes.
- Use a stable identifier from your system (such as a merge-batch ID or revision ID) as the
idempotency_key; send a new key for each new piece of content.
V1.9.0
2026-07-31
New: unlimited plans
In addition to the credit model, this release introduces unlimited plans: a contract authorizes a fixed feature bundle and usage limits, and features included in the plan are not charged per minute during the authorized period. The plan's feature bundle and its limits (simultaneously recognized transcription languages, single-recording length cap, usage-hour thresholds, concurrent recording limit) are configured per contract.
- Features the plan does not include are rejected outright — they do not fall back to credit billing.
- Broadcasting is never included in unlimited plans; broadcasts are always billed in credits.
- Base features included in every plan: base speech recognition, professional vocabulary, summary, and full-text re-translation.
See Pricing — Unlimited Plans.
New error codes (full details in Error Code Reference — Plan and Usage Limit Errors):
| Error Code | Scenario |
|---|---|
plan_feature_not_allowed | The plan does not include the feature in use. Over WebSocket there are two occurrence points: start is rejected (the connection is not closed; adjust the parameters and retry), or the per-minute check while recording detects it (for example, a feature not in the plan was turned on mid-session) → the current recording is stopped. REST endpoints (creating a broadcast, floating subtitles, audio import) return HTTP 403 |
concurrency_limit_reached | This API Key has reached its concurrent recording limit; the connection is not closed — retry after another recording ends |
daily_limit_disconnect | The plan's usage threshold was reached and the current recording was stopped; you may start a new recording immediately |
daily_limit_reached | Usage has reached the plan's limit; available again after the plan's reset (daily limits reset the next day) |
plan_daily_limit_reached | REST: POST /api/v1/auth/ticket and the import upload when the plan's daily hard limit has been reached (HTTP 402) |
too_many_languages semantics extended: details.max may now come from the plan's cap on simultaneously recognized transcription languages, in addition to the system-wide limit (10 transcription languages); details carries max and received.
New: plan lookup endpoint
Added GET /api/v1/me/plan (authenticated with X-API-Key, read-only, queryable even with zero balance): reports the current billing mode (credit / unlimited), the plan's feature bundle, one-off features, each limit with current usage, and the estimated time a restriction lifts. When a request is rejected by a plan limit (403 / 402), use this endpoint to answer "what does my plan include, how far am I from a limit, and when does the restriction lift?"
See My Plan API.
Changed: audio import response fields
POST /api/v1/imports/check-quotaresponse gainsdata.reason:null(allowed) /insufficient_credit(topping up resolves it) /plan_not_allowed(the plan does not include audio import; a plan upgrade is required).- The 202 response of
POST /api/v1/importsgainsdata.downgraded_features(array): when the plan includes audio import but not some requested sub-features (such as speaker diarization or translation), those sub-features are skipped and the import proceeds; the skipped features are listed in this field. An empty array means nothing was downgraded.
Changed: remain_quota semantics (existing field)
For API Keys with a dedicated credit allotment, remain_quota now reflects the credit actually available to that key (its dedicated allotment) instead of the account's total balance; accounts without dedicated allotments see no change. This affects remain_quota in POST /api/v1/imports/check-quota and remaining_budget in WebSocket error messages.
Client recommendations
- Credit-based integrations require no changes; the error codes added in this release occur only with unlimited plans.
- For plan-bound integrations, call
GET /api/v1/me/planwhen you receiveplan_feature_not_allowed/plan_daily_limit_reachedto show the user the plan contents and recovery time, rather than treating it as a service outage. - On
concurrency_limit_reachedanddaily_limit_disconnect, both the connection and the account remain usable: for the former, retry after another recording ends; for the latter, you may start a new recording immediately. - If you display
remain_quota, note that for API Keys with a dedicated allotment its meaning has changed to the key's allotment.
V1.8.0
2026-07-29
Changed: pricing update
This release updates the billing model for several services.
Rate changes
| Item | Previous rate | New rate |
|---|---|---|
| Text-to-speech (TTS) | 0.5 credits / min | 1.0 credits / min |
| Professional vocabulary | 0.5 credits / min | Free |
| Audio import (base speech recognition) | 1.0 credits / min | 0.3 credits / min |
Interpretation mode: TTS is now billed separately
The interpretation integrated rate (1.5 credits / min) previously included text-to-speech. From this release, TTS is billed separately at an additional 1.0 credits / min when enabled. The rate for interpretation without TTS is unchanged (still 1.5 credits / min). TTS can be toggled at any time during a session, and the rate adjusts from that minute onward.
Summary and full-text re-translation: now billed by text volume
| Item | Previous billing | New billing |
|---|---|---|
| Meeting summary (first generation) | 3.0 credits per action | 0.1 credit per 1,000 transcript characters |
| Re-generate summary | 3.0 credits per action | 0.1 credit per 1,000 transcript characters |
| Re-translation (full text) | 2.0 credits + 0.3 credits per minute | 0.1 credit per 200 characters |
"Characters" refers to the actual character count of the transcript. Any partial billing unit is charged as a full unit, and each action is charged at least 0.1 credit. Single-sentence re-translation remains free.
Broadcast: cloud translation fee is now the sum of two items
The broadcast audience delivery fee previously used a single lookup of "maximum audience × translation-language tier". From this release, the "number of translation languages" and the "maximum audience size" are looked up separately and added together. Tiers for 10,000 / 20,000 / 30,000 viewers have been added (the default per-account limit remains 5,000; contact sales to raise it).
In addition, broadcasts are no longer charged the "from the 2nd translation language" surcharge — multi-language costs are now fully covered by the cloud translation fee.
Removed: POST /api/v1/summary endpoint
The REST endpoint for "generate a summary for any transcript" is discontinued as of this release.
Reason for removal: the endpoint stood outside the recording workflow and saw zero actual usage; its authentication also differed from every other REST endpoint (Authorization: Bearer vs X-API-Key), leaving it off the maintained path.
Alternatives: summary generation remains available through two recording-bound entry points:
| Scenario | How to use |
|---|---|
| Automatic generation after a recording ends | The summary_* fields of the WebSocket start action |
| Re-generate for an existing recording | GET / POST /api/v1/sse/regenerate/summary/{taskId} |
Who is affected: integrations calling POST /api/v1/summary directly. If you were using it to summarize transcripts from external sources, please contact your service representative to discuss alternatives.
Change to content-filter downgrade: the documentation previously suggested switching to this endpoint to trigger automatic downgrade when SSE summary re-generation returned llm_content_filtered. With the endpoint removed, please revise the prompt or transcript content and retry instead.
GET /api/v1/summary-templates(summary template lookup) is a different endpoint and is unaffected.
Client recommendations
The pricing changes require no code changes; if your workflow estimates costs from the rate table, please re-evaluate using the new rates. See Pricing for the full rate table and billing examples.
If your integration uses POST /api/v1/summary, please migrate per the alternatives above.
V1.7.7
2026-07-25
New: deployed-version lookup endpoint
Added GET /api/v1/version (REST service) and GET /version (realtime service). Both report the currently deployed version and build identifier so integrators can run a version gate before going live — for example, "this fix requires ≥ vX.Y.Z".
Both endpoints require no authentication: a version gate that fails on authentication would be indistinguishable from a version mismatch, defeating its purpose. The response contains only the version and build identifier.
{ "service": "vas-api", "version": "1.7.7", "build": "a1b2c3d4e5f6" }
Note: The realtime service and the REST service are deployed independently and may run different versions. Query whichever service owns the functionality you need to verify.
Recommended version-gate logic: this endpoint ships in V1.7.7, so a successful response by itself proves the version is ≥ V1.7.7. A 404 means the version predates V1.7.7 and must be confirmed with your service contact.
See REST API · GET /api/v1/version.
Client Recommendations
No changes required. If your release process needs to confirm the VAS version, consider automating it with this endpoint instead of manual confirmation.
V1.7.6
2026-07-25
Fix: broadcasts using a custom summary prompt could fail to save the recording
When a host started a broadcast in custom summary mode (summary_mode: "custom") and that broadcast also had a shared summary template configured, the two settings conflicted and the recording could fail to save after the broadcast ended. The summary itself was still generated correctly from the custom prompt, which made the problem easy to miss.
With this release, custom summary mode uses exactly what the host passes in and no longer falls back to the broadcast's shared template. Broadcasts using shared template mode are unaffected.
Who is affected: broadcasts started over WebSocket start with summary_mode: "custom" where the broadcast also has a summary_template configured. If your broadcasts always use shared templates, you are not affected.
Documentation correction: broadcast settings have two layers — channel defaults and the current session
A broadcast is a "channel", and one channel can go live many times, so its settings split into two layers:
| Layer | What it changes | Interface |
|---|---|---|
| Channel defaults | The starting configuration for every future broadcast | PATCH /api/v1/broadcasts/{id} |
| The in-progress session | What that session actually uses | Host-side WebSocket actions |
PATCH updates the channel defaults, so transcription_languages, translation_languages, speaker_diarization, tts_config, summary_template, and summary_language take effect on the next broadcast. This lets a host adjust settings for the next broadcast while the current one is still live. The exception is access_type, pass_code, and max_viewers, which are viewer access controls and are applied immediately to the in-progress broadcast.
This is a correction to how the behavior is documented; the API behavior itself has not changed. The previous wording described the endpoint as adjusting settings "in real time" without distinguishing the two layers, which made it easy to assume that changing the summary template would affect the broadcast already in progress. The description and the parameter table have both been corrected.
Fix: recordings without a summary template could not change only the summary language or output format
conversation and broadcast recordings may omit summary_template (the summary then falls back to the system default template). Previously, using set_summary on such a recording to change only summary_language or summary_plain_text was incorrectly rejected with summary_mode_field_mismatch for a missing summary source.
With this release, requests that only change the output language or format no longer require a summary source. Explicitly enabling the automatic summary (auto_summary: true) or changing the summary source still requires a template or a custom prompt — that rule is unchanged.
Client Recommendations
- If you call
PATCHduring a live broadcast to change the summary template and show it as "applied", change that to "applies to the next broadcast" — or switch to the WebSocketset_summaryaction so the current session's summary picks up the new settings (available since V1.7.5, and supported for broadcasts as well). - To use a custom summary prompt for broadcasts, pass it in the host-side WebSocket
startaction, and upgrade to this release to avoid the saving problem described above.
V1.7.5
2026-07-25
Summary settings can now be changed while recording
A new WebSocket action, set_summary, lets you change the summary settings that will be applied when the recording stops. It covers both the shared template mode (builtin) and the custom prompt mode (custom), and can also change the summary language, switch to plain-text output, or skip the automatic summary entirely for that recording.
A live recording generates its summary once, at stop time, so the change applies to that automatic summary; if you send the action several times, the last one before stopping wins. Previously the settings were fixed once recording began, and the only way to change them was to regenerate the summary after stopping — which produced an extra summary and an extra charge. That limitation is lifted in this release.
- The summary source (mode / template / prompt) is replaced as a set; switching modes automatically clears the fields belonging to the other mode.
- Changing the summary source does not turn the automatic summary on. If it was disabled when the recording started, send
auto_summary: trueas well. - After reconnecting,
resume_okreturns the updated summary settings, so you can reconcile client state directly. - See set_summary for details.
Consistent records when no shared template is specified
When a recording uses the shared template mode without naming a template, the system already fell back to the default shared template to generate the summary. Starting with this release that default template is also stored in the recording record, so the template actually used matches what is recorded. This corrects internal records only; summary content and billing are unaffected.
Client Recommendations
No changes required. set_summary is a new capability and existing integrations keep their current behavior. If your product lets users switch summary templates mid-recording, consider adopting this action to avoid an extra summary generation.
V1.7.4
2026-07-23
Retrying an audio import now preserves custom summaries and vocabulary settings
When you retry a failed audio import, the original custom summary prompt and vocabulary settings (terminology, fuzzy correction, translation dictionary) are now preserved and rebuilt on retry. As a result, the retry regenerates the summary and applies the vocabulary as expected — both the transcript and the summary are produced.
Previously, a failed custom-summary import would not regenerate the summary on retry and required resubmitting the whole import. That limitation is removed as of this release — simply retry.
The translation dictionary now also influences summaries
When a summary is produced in a "target language" that has matching entries in your translation dictionary, summary generation will try to follow the dictionary's specified renderings, so the summary wording stays consistent with the translated transcript.
- Scope: the end-of-session summary for live recording, and SSE "regenerate summary" (preview and persist).
- When a summary is produced in the "source language", the translation dictionary (source → target) does not apply.
- This is best-effort, not a guaranteed literal replacement; it is a behavioral extension and requires no changes to existing integrations.
Audio import transcripts now include summary provenance fields
Transcript records produced by audio import now also carry the summary provenance fields (summary_mode, summary_template, summary_language, summary_plain_text, summary_prompt_snapshot), consistent with live recording. These are additive, backward-compatible fields; existing integrations are unaffected.
Client recommendation
No changes required. If you previously switched to "resubmit the import" because a custom-summary import failed, you can now simply call retry to recover the summary.
V1.7.3
2026-07-22
Audio import supports custom summaries (custom)
Audio import summaries now support summary_mode=custom, consistent with live recording and SSE regeneration: you can send summary_prompt (fully replaces the built-in template) plus summary_prompt_slug (a custom identifier).
- When
summary_modeis omitted: behavior is exactly as before (usessummary_template). summary_mode=builtin:summary_templateis required.summary_mode=custom:summary_promptandsummary_prompt_slugare required, andsummary_templatemust not be present (mutually exclusive).
The custom prompt text is not persisted (consistent with live recording); therefore a retry of a failed custom import does not regenerate the summary (the transcript is still produced; only the summary is left empty). (This limitation is removed as of V1.7.4, above; from V1.7.4 onward, retry preserves the custom summary and regenerates it as expected.)
Fix: summary templates for audio import
Previously, specifying summary_template on an audio import had no actual effect — whether you passed meeting, interview or course, the resulting summary always used the same generic format.
As of this release, imports correctly apply the content of the specified template, matching the behavior of live recording.
Note: this means the summary content produced by existing import integrations will change (to match the template you specified). If you relied on a fixed summary format from imports, please re-check it. Behavior is unchanged when summary_template is not specified.
Fuzzy correction gains a case-insensitive option
Every rule in fuzzy_correction accepts a new optional field, case_insensitive. When set to true, all incorrect variants in that rule match regardless of letter case.
It defaults to false, which behaves exactly as before — existing integrations need no changes.
The flag is per rule, so the same correct term can be split across several rules with different settings:
{
"zh-TW": [
{ "correct": "IPEVO", "incorrect": ["ltfo"], "case_insensitive": true },
{ "correct": "IPEVO", "incorrect": ["ivo"] }
]
}
With the settings above, LTFO, LtFo and ltfo are all corrected to IPEVO, while only the lowercase ivo is corrected — the personal name Ivo is left untouched.
Where it applies: the config action for live recording, and audio import. The resume_ok settings snapshot also returns this field (omitted when false).
Note: enabling it widens the false-positive surface. If a variant is identical to an ordinary word or a personal name (for example ivo versus Ivo), keep the default exact-case matching. Chinese rules are unaffected (Chinese has no letter case).
Client guidance: no changes required; enable it per rule when you need it.
V1.7.2
2026-07-22
Two-way translation mode accepts speaker_diarization (behavior change)
Previously, a two-way translation request (type=conversation) carrying speaker_diarization=true was rejected. As of this release it is accepted, and the parameter is ignored — two-way translation does not perform speaker separation.
Behavior for other recording types is unchanged.
Billing: the per-minute rate is the same as a two-way translation session without this parameter; nothing extra is charged. Note, however, that these requests were previously rejected, so no recording was created and nothing was billed; from this release they create a recording and begin billing normally.
Client guidance: integrations that stripped speaker_diarization from two-way translation requests to work around this error can stay as they are; no change is required.
Two-way translation no longer reports a stale translation language after a mid-session change (fix)
After a two-way translation session changed languages mid-session, translation_languages in resume_ok, in the Webhook, and in the recording record still showed the language from before the change. As of this release it reflects the current language, as does the language information for floating subtitles.
Live translation itself was unaffected and was always correct.
For two-way translation, translation_languages holds a single language representing the current counterpart language; if the language was changed mid-session, the final record holds the last one.
Client guidance: if you relied on translation_languages from resume_ok or the Webhook to determine the translation language, that value is only trustworthy from this release onward.
Two-way translation now validates speakers language codes at start (fix)
Previously, when two-way translation specified languages via speakers, an unsupported language code still allowed the connection to start and audio to be sent, but the recording was never created — so no transcript was kept and no usage was recorded.
As of this release, start returns 400 invalid_transcription_language, with details.field set to speakers[].language and details.speaker_id identifying which speaker.
Client guidance: make sure speakers uses language codes from the supported list. If your integration contained a misspelled code, you will now receive an explicit error instead of the previous silent failure.
Documentation corrections
- Corrected several statements claiming that all recording types are translated in real time (
recordhas not supported translation since v1.7.0) - Capability matrix corrected: broadcast does support speaker separation (with a single transcription language only)
- Two-way translation rules table now lists:
translation_languagesis set by the server, andrealtime_translationis fixed totrue
V1.7.1
2026-07-22
Multi-language transcription is no longer surcharged (billing change)
Specifying multiple transcription languages (automatic language detection) is no longer surcharged as of this version — the number of source languages does not affect the per-minute rate. The former "Multi-language transcription (from the 2nd language, each +1) — 0.3 points/minute" line has been removed from the rate table.
Surcharges for translation output languages are unchanged (sentence translation +0.2, real-time translation +0.4 per additional translation language).
Interpretation (conversation) is the most clearly affected: interpretation always requires exactly 2 transcription languages by specification, so it was previously surcharged a fixed 0.3 points per minute. After this change, interpretation is billed at the integrated rate published in the rate table.
Client guidance: no changes required. Recordings that use multi-language detection (including all interpretation recordings) will see a lower per-minute rate. See Pricing.
V1.7.0
2026-07-22
record recording type made lightweight (breaking change)
Starting this version, record (plain recording) is positioned as a lightweight speech-recognition-only type, with the following behavior changes:
- Translation no longer supported: Sending
translation_languagesfor arecordreturns400 record_translation_not_allowed(consistent across thestart,switch_language, andretranslatelive paths, as well as the post-hoc transcript-retranslation / summary-translation endpoints). - TTS no longer supported: Sending
tts_enabled=truefor arecordreturns400 record_tts_not_allowed. - Summary now off by default, opt-in:
recordno longer generates a summary automatically by default; it is generated only whensummary_templateorsummary_mode=customis provided. Sendingauto_summary=truewithout a template returns400 record_summary_requires_template. - Speaker separation remains available:
recordcan still sendspeaker_diarization=truefor multi-speaker separation.
For billing, a plain record (with summary off) counts only speech-recognition usage, without translation or TTS costs.
New error codes: record_translation_not_allowed, record_tts_not_allowed, record_summary_requires_template (see Error Codes · Record Type Restriction Errors).
Client guidance:
- If you use
recordfor plain voice notes (no translation / summary / TTS needed): no changes required, and it costs less. - If you previously sent
translation_languagesortts_enabled=trueforrecord: remove these fields, or switch to thetranscribetype. - If you previously relied on the summary auto-generated by default for
record: explicitly providesummary_templateorsummary_mode=custom, otherwise summaries will no longer be generated automatically after this version. - Existing
recordrecording data is unaffected (reading, playback, existing transcripts and summaries all work normally); only re-translating them afterward is rejected.
V1.6.10
2026-07-16
Floating subtitle feed now supports broadcast hosts (during a live broadcast)
While a broadcast is live, a broadcast host can subscribe to their own live-transcript floating-subtitle feed using the same flow as a regular recording: exchange an API Key for an owner feed_token (Floating Subtitle Feed Token), then connect to the Floating Subtitle SSE (Floating Subtitle SSE).
Client guidance: For broadcasts, listen for the broadcast_recording_ready event to obtain the finalized task_id after going live (the standby session_started carries the initial ID, which returns 425 when exchanging for a token), then exchange it for a feed_token.
Endpoints, parameters, and response formats are unchanged; floating-subtitle behavior for regular recordings is unaffected.
New machine-readable status field on status events
The status events for pause / resume / stop (over WebSocket and the floating-subtitle SSE) now include a status field: live / paused / ended. See WebSocket Events · status and Floating Subtitle SSE · status.
Client guidance: The floating subtitle window (especially the host's own) should act on the status field — paused → freeze, ended → close the window, live → resume. Because the floating-subtitle SSE does not close automatically when the recording stops, failing to close on ended will leave it frozen on the last sentence. message is display text with no guaranteed format — do not parse it to determine state. This field appears only on those three lifecycle transitions; other status events such as set_name do not carry it. Backward compatible: existing fields are unchanged, and integrations that ignore this field are unaffected.
Floating Subtitle Feed Token: 425 / 410 split
When exchanging for a floating-subtitle feed_token (both owner and audience endpoints), "recording not ready" and "recording has ended" previously both returned 425. They are now split:
425Too Early: recording not ready (being created) → retry after a short delay.410Gone: recording has ended → do not retry.
See Floating Subtitle Feed Token. Client guidance: on 410, stop retrying and close the floating subtitle window.
Language-switch events now carry the full translation-language set
The language_switch_start, language_switch_done, and translation_language_removed events (over WebSocket and the floating-subtitle SSE) now include a translation_languages field: an authoritative snapshot of the current full set of translation languages. See WebSocket Events.
Background: Previously these events carried only a single translation_language, so consumers could not tell "add a language" from "replace the language" and could drop an existing language by mistake.
Client guidance: On these events, overwrite your local translation-language set directly with translation_languages instead of inferring the operation from the single translation_language. This also resolves op:add / replace / out-of-order / reconnect-gap language-set sync issues. Backward compatible: purely an added field; integrations that only read translation_language are unaffected.
Floating-subtitle reconnect consistency: speaker events in replay, connected reflects the current language set
Two behavior fixes on the Floating Subtitle SSE (Floating Subtitle SSE) that resolve "reconnecting or late-joining viewers see stale information":
- Speaker events are included in replay:
speaker_renamed/speaker_reassigned/speakers_merged/speakers_auto_mergedare now replayed in original order (after the sentences they affect). Client guidance: handle them during replay exactly as in live mode — retroactively update the speaker labels of existing sentences byaffected_sids; otherwise reconnecting viewers will see pre-rename speaker names. connectedreflects the current language set: thetranslation_languagesin theconnectedevent is now the current authoritative set (reflecting mid-recording language additions/removals), no longer the value frozen at recording start.
Backward compatible: no new fields and no format changes; replay is simply more complete and the snapshot more up to date.
V1.6.9
2026-07-14
Behavior change: translation output language limit raised to 12
The translation_languages count limit is raised from 8 to 12 (applies to live recording, import, and broadcast; effective for direct API access). This is a relaxed boundary and is fully backward compatible with existing integrations that use ≤8 languages.
- The input axis is unchanged: the transcription (source) language limit remains 10; input and output are two independent dimensions.
- The broadcast audience delivery fee adds a "9–12 languages" rate band (see the Pricing Guide).
- The
too_many_languageserror covers both axes: it is triggered when transcription > 10 or translation > 12.
V1.6.8
2026-07-13
Added: multi-language transcription input for broadcasts
When creating or updating a broadcast, the new transcription_languages field (array of strings, up to 10, distinct) is now the primary field for transcription (source) languages, letting you specify multiple languages for continuous multi-language recognition.
- The legacy
transcription_languagefield (single string) is now deprecated but still supported (kept for backward compatibility; it equals the first element oftranscription_languages). - The create/update response returns both fields:
transcription_language(= the first element, for backward compatibility) andtranscription_languages(the full array). - When
summary_languageis not specified, it now defaults to the first transcription language (the first element oftranscription_languages).
Added: source language array in audience info
Broadcast audience info (/info) now includes source_langs (array of strings) alongside the existing source_lang (= the first element), listing all transcription languages for the broadcast.
Behavior change: speaker diarization supports only a single transcription language
A broadcast with speaker_diarization enabled supports only a single transcription language. If multiple transcription languages are provided, the create/update request returns 422 (multi-language transcription and speaker diarization are mutually exclusive).
Client recommendations
- For new integrations, use
transcription_languagesto specify transcription languages;transcription_languagestill works but is deprecated and should be phased out. - When reading broadcast settings and audience info, rely on
transcription_languages/source_langs;transcription_language/source_langare retained as the first element for compatibility only. - Keep broadcasts that require speaker diarization on a single transcription language to avoid a 422.
V1.6.7
2026-07-10
Behavior change: real-time multi-language translation for live recording
When multiple languages are specified in translation_languages (up to 8), every recording type (transcribe / record / conversation / broadcast) now translates all specified languages in real time. Previously, non-broadcast types only translated the first language and silently ignored the rest; this release brings the behavior in line with the long-standing documentation (Multi-Language Translation).
How results arrive: each language returns its own independent result event (same sid, single language key inside translations); multiple languages are never merged into one event. Note for existing multi-language clients: you previously received only the first language's translation — after this upgrade you will start receiving all languages. Accumulate translations by "sid + language code" instead of overwriting.
- Word-by-word (interim) real-time translation requires
realtime_translation: true; with the defaultfalse, all languages are translated once the sentence is finalized - If one language fails, the remaining languages are still delivered; the failed language additionally receives an
errorevent whosedetails.translation_languageidentifies it
Behavior change: switch_language redefined as add / remove for multi-language sessions
Sessions with 2 or more translation languages must include the op parameter in switch_language:
op: "add": adds a single language and automatically backfills existing sentences (same response sequence as a single-language switch); the limit is 8 languagesop: "remove": removes a single language and returns the newtranslation_language_removedevent; existing translations are kept, and at least 1 language must remain- Omitting
opin a multi-language session returns the newswitch_language_op_requirederror (this prevents the legacy replace semantics from corrupting the language set)
Single-language sessions keep the existing replace-and-retranslate semantics; no client changes are needed for them.
New error codes: switch_language_op_required, switch_language_already_exists, switch_language_not_in_session, switch_language_last_language (see Error Codes).
Billing
The billing rules for multi-language translation are unchanged (billing has always been per language count, with a per-minute surcharge starting from the second language — see the Pricing Guide); this release simply brings the actual translation delivery in line with billing. Credit consumption reminder: multi-language real-time translation consumes noticeably more credits per minute than single-language (8 languages with real-time translation is roughly 2.8× or more). If credits run out, the recording stops per the existing rules — keep an eye on your balance.
Client recommendations
- In multi-language sessions, render translations keyed by the language code in
translations; do not let the last event overwrite the others - For word-by-word multi-language subtitles, include
realtime_translation: true - Use
op: "add"/op: "remove"to adjust languages in multi-language sessions; receiving aswitch_language_op_requirederror indicates the backend is on v1.6.7
Reference documentation
- WebSocket reference - start, multi-language translation
- WebSocket reference - switch_language
- WebSocket events - translation_language_removed
- Real-Time Voice Translation Guide - Multi-Language Translation
V1.6.6
2026-07-09
Added: Pricing page
Added a Pricing page with the full credit-based rate card:
- Credit pricing: 1 credit = TWD 1.5 (≈ USD 0.047).
- Base usage (real-time recording / import): per-minute rates for speech recognition, vocabulary refinement, multi-language transcription, speaker identification, translation (including a per-additional-language surcharge), and text-to-speech.
- Interpretation mode: real-time interpretation at an integrated rate.
- Value-added services (one-time): meeting summary, re-generated summary, re-translation.
- Broadcast billing: host side (content processing) + audience side (audience delivery fee by max audience size × translation language band), with the full rate table and billing examples.
Behavior change: transcript input language limit raised to 10
The maximum number of "input (source/transcription) languages" for transcription has been raised to 10. Affects real-time recording (WebSocket) and audio import.
- Speaker identification mode still supports only a single source language (multi-language transcription and speaker identification are mutually exclusive); this limit does not apply to that mode.
- The translation (output) language limit is unchanged (still 8).
Documentation clarification: broadcast maximum audience size is at least 1
When creating a broadcast, max_viewers is at least 1 (the API already enforced a minimum of 1; this release makes it explicit in the docs and admin panel as well).
V1.6.5
2026-07-06
Behavior change: silence thresholds lengthened across all five speaking-speed levels
The silence thresholds for sentence segmentation mapped to the five speaking_speed levels have been adjusted. Longer thresholds tolerate mid-sentence pauses better and reduce sentences being cut prematurely, at the cost of segmentation landing slightly later at every level.
| Level | Old threshold | New threshold |
|---|---|---|
very_fast | 150ms | 300ms |
fast | 300ms | 600ms |
normal (default) | 500ms | 800ms |
slow | 700ms | 1200ms |
very_slow | 1000ms | 1500ms |
- The default is still
normal, but its threshold changes from 500ms to 800ms. - Level names, the API surface, and
set_speaking_speedusage are unchanged; no integration changes are required after upgrading.
Client recommendation
- Expect segmentation to land slightly later and sentences to be more complete. If you relied on
very_fastfor the most immediate segmentation, its threshold changes from 150ms to 300ms (very_fastremains the fastest option).
V1.6.4
2026-07-04
Added: resume_ok returns the recording settings (state reconcile)
resume_oknow includes asettingsobject with the recording settings currently held by the session (speaking speed, audio format, languages, TTS, terminology / fuzzy correction / translation dictionary, summary settings, recording name, conversation dialog mode and speaker language map, etc.), letting clients reconcile local caches against the server's authoritative values after a resume or full page refresh.- Values are current (reflecting mid-recording changes via
set_speaking_speed/config/set_name/set_tts/switch_conversation_mode/set_speaker_language) and presented in API format (e.g.,speaking_speedreturns the level string"normal"). - See Connection and Authentication — the settings object for field details.
Resume format constraint (important)
settings.audio_formatreturns the audio format fromstart; resumed connections must keep using the same format (the server decodes with the original format and does not renegotiate).
Behavior change
- The
conversation_language_change_failederror response no longer includes internal error details indetails(consistent withset_speaking_speed_failed; branch on the error code and show a generic failure message).
Documentation fix
- Corrected the
set_speaking_speedsupport scope: supported in all recognition modes except multi-speaker (multi_speaker) — including broadcast and multi-language LID — not just single / conversation mode.
Client recommendations
- After receiving
resume_ok, treatsettingsas the authoritative source and reconcile your local settings cache (especially for refresh / multi-tab scenarios). - Older clients can ignore the
settingsfield; existing behavior is unaffected.
V1.6.3
2026-07-03
Speech Segmentation Settings Update
Added: Adjust speaking speed during recording
- Added the
set_speaking_speedaction to adjust the speaking speed (the silence threshold for segmentation) while recording. Applying the change causes a brief interruption in recognition. Success response:speaking_speed_changed(returns the applied speed). Supported in single / conversation mode only; not applied in multi-speaker (multi_speaker) mode.
Behavior change: segmentation_mode removed
- The
segmentation_modeoption (auto/by_time) instartoptions has been removed and is no longer available. Segmentation timing is now controlled solely byspeaking_speed. If this field is still sent, the server ignores it without affecting other parameters.
speaking_speed description update
- The default
normalcorresponds to a 500ms silence threshold (the same in every environment). - Five levels:
very_slow(1000ms) /slow(700ms) /normal(500ms) /fast(300ms) /very_fast(150ms).
Client recommendations
- To adjust segmentation timing during recording, use
set_speaking_speed(send after the UI control is released to avoid repeated STT rebuilds). - If you previously sent
segmentation_mode, you can remove that field (the server now ignores it).
Reference
V1.6.2
2026-07-02
set_name (Set Recording Name) fixes and identification improvements
Behavior changes
- When the recording name exceeds the length limit, the server now returns a clear
set_name_too_longerror, and the responsedetailsincludesmax_length(previously no message was returned and clients would time out). - The recording name length limit is 60 characters.
Success response identification (new event field)
- The set_name success response now includes
event: "name_set"and anamefield so clients can identify it precisely. - For backward compatibility, the success response keeps
action: "status"unchanged.
Deprecation notice
- Relying on
action: "status"to detect set_name success is now deprecated and may be removed in a future version. Identify success viaevent: "name_set"(together with thenamefield) instead.
Error codes
- The set_name error codes are
set_name_emptyandset_name_too_long.
Client recommendations
- Clients that integrate the WebSocket protocol directly and detect set_name success via
action: "status"should switch toevent: "name_set". No changes are needed for other clients.
Reference
V1.6.1
2026-06-29
Added: Floating Subtitle Audience Sharing
In addition to the recording owner, the floating subtitle now supports audience sharing: the owner can enable sharing and obtain a share secret (to embed in a share link / QR code), letting other on-site audience members view read-only, with no login, no API Key, and no separate charge.
New Endpoints
| Method | Endpoint | Description |
|---|---|---|
| POST | /api/v1/auth/tasks/{taskId}/subtitle-share | Owner enables / resets audience sharing and obtains a share secret |
| DELETE | /api/v1/auth/tasks/{taskId}/subtitle-share | Owner stops sharing |
| POST | /api/v1/public/tasks/{taskId}/subtitle-feed-token | Audience exchanges a share secret for a read-only audience Token (no authentication) |
New Events
- The Floating Subtitle SSE adds a
viewersevent reporting the current viewer count and limit (count/max). - The Floating Subtitle SSE adds a
subtitle_closedevent: when the host "closes sharing" or "stops the recording", the server proactively sends it and ends the audience connection; the audience client should stop reconnecting on receipt. The host's own connection is unaffected.
Behavior
- The number of audience members per recording is capped (server-configured, default 10, excluding the owner). Use the
maxfield of theviewersevent rather than hardcoding a value. When the audience limit is reached, the connection returns 429. - "Closing sharing" by the host immediately ends all audience connections (sending
subtitle_closed), and the share link is invalidated so reconnection is rejected; "stopping the recording" behaves the same. The audience is a passive receiver.
Client Recommendations
- Existing floating-subtitle (owner) integrations are unaffected and require no changes.
- To offer shared viewing: on the owner side, call subtitle-share to obtain a share secret and build a share link; on the audience side, exchange the secret for a Token via the public endpoint, then connect to the Floating Subtitle SSE.
- The audience client must handle the
subtitle_closedevent: close the connection and stop auto-reconnecting on receipt to avoid futile reconnection.
Reference
V1.6.0
2026-06-27
Breaking Change: Removed the legacy recording_id naming (naming unification complete)
The recording_id → task_id naming unification announced since V1.4.1 reaches its final step in this release: the legacy recording_id field and legacy paths are fully removed. All task identifiers are now unified under task_id.
The value of
task_idis exactly the same as the oldrecording_id(the UUID of the same recording). This migration only renames the field / path; the identifier itself is unchanged.
Affected Scope (changes required)
- WebSocket payload: The
session_startedandresume_okevents no longer carry therecording_idfield — readtask_idinstead.- This item was previously announced for removal in V2.0.0; it has been brought forward and completed in V1.6.0.
- REST endpoints removed: The following legacy
recordingspaths are removed; use the correspondingtaskspaths (identical behavior, same identifier value):Removed (legacy) Use instead (new) PATCH /api/v1/recordings/{recordingId}/speakers/renamePATCH /api/v1/tasks/{taskId}/speakers/renamePATCH /api/v1/recordings/{recordingId}/speakers/reassignPATCH /api/v1/tasks/{taskId}/speakers/reassignPATCH /api/v1/recordings/{recordingId}/entries/{sid}PATCH /api/v1/tasks/{taskId}/entries/{sid} - SSE connection message text: The
connectedmessage label for the history / retranslation / summary-regeneration streams changes from(recordingId: ...)to(taskId: ...)(a plain-text hint; the identifier value is unchanged).
Not Affected (no changes needed)
current_recording_id: Thecurrent_recording_idfield in the broadcast query response is kept unchanged (it means "the UUID of the currently in-progress recording" — an unambiguous meaning, not a target of this cleanup).- SSE retranslation path:
GET /api/v1/sse/recordings/{taskId}/entries/{sid}/retranslateis kept unchanged (therecordingssegment in the path is existing naming; the parameter is alreadytaskId). - Task queries, audio / transcript export, Webhook (
data.task_id), and other interfaces already centered ontask_idare completely unchanged.
Client Recommendations
- Change any code reading
recording_idto readtask_id(same value, drop-in replacement). - Change any REST calls to
/api/v1/recordings/{id}/...to/api/v1/tasks/{id}/.... - If you previously string-matched
recordingId:in the SSEconnectedmessage, matchtaskId:instead (better still, switch to event-type-based detection rather than relying on message text).
Reference
V1.5.12
2026-06-26
Documented recording options sub-fields (speaking_speed / segmentation_mode / profanity_handling)
The options sub-fields of the WebSocket start action are now documented and can be used to fine-tune STT sentence segmentation and profanity handling:
speaking_speed:very_slow/slow/normal(default) /fast/very_fast— adjusts the silence threshold for segmentation; use a slower setting for slower speakers to avoid cutting on mid-sentence pauses.segmentation_mode:auto(default) /by_time— segmentation strategy;by_timecombines withspeaking_speedto adjust the threshold.profanity_handling:mask(default) /remove/show— profanity handling.
Client Recommendations
- All are optional; when omitted, defaults apply (
normal/auto/mask), so existing integrations need no changes. - For slower speakers, or to avoid cutting sentences, try
speaking_speed: slow. - Multi-speaker mode (
multi_speaker) does not currently applyspeaking_speed/segmentation_mode.
Reference
V1.5.11
2026-06-25
New Feature: Floating Subtitle Transcript Feed
Added a read-only "floating subtitle" transcript stream: you can subscribe over an independent connection to the live transcript of an in-progress recording (source-language original text + target-language translations), suitable for desktop floating-subtitle windows, second-screen captions, and similar use cases. It is separate from the recording's own connection and can be opened independently on a different device or window.
- Exchange for a token:
POST /api/v1/auth/tasks/{taskId}/subtitle-feed-token(exchange an API Key for a short-livedfeed_tokenbound to the recording; only the recording owner can exchange). - Subscribe to the stream:
GET /tasks/{task_id}/subtitle?feed_token=...&lang=...(SSE). Receivesconnected/result(original text and translations) /status/ speaker and interpretation language-switch events. Original text is replaced in place bysid+is_final; translations inherit the speaker by matchingsidto the original line. - Supports
langfiltering of target languages, replay on connect, and automatic reconnection. Connection timing boundaries: not started 425, ended 410, invalid token 401, too many connections 429.
Client Recommendations
- Exchange for the
feed_tokenafter recording has started; if you receive 425 (recording not ready) right after starting, retry after a short delay. - The
feed_tokenis valid for 15 minutes and is extended automatically while connected; for long recordings, exchange for a new one before expiry.
Reference
V1.5.10
2026-06-20
New Feature: Per-Key Source IP Rules (Allowlist + Denylist)
You can configure source IP rules for each API Key (managed in the user portal), applied to both the REST API and live WebSocket:
- Allowlist (allow): only IPs on the list may use the key.
- Denylist (deny): a matching IP is denied — useful for blocking specific malicious/attacking IPs while letting everything else through.
Precedence (deny first): a match on any deny rule is rejected; if an allowlist exists, the IP must match one of its entries; both empty = unrestricted (backward compatible). When both are set, the effective permitted set = on the allowlist AND not on the denylist (an IP listed in both is denied). Even a leaked key cannot be used from an unauthorized (or denied) IP.
New Error Codes
auth_ip_not_allowed(403): The source IP is not permitted (not on the allowlist, or matched the denylist). Access from an authorized IP address.auth_account_blocked(403): The account has been blocked. Contact technical support.
Behavior Change
- When an account is blocked, API Key verification now returns 403
auth_account_blocked(distinct from the 401 returned for an invalid API Key); the live WebSocket handshake is rejected with the same reason.
Client Recommendations
- If you have configured IP rules for a key, make sure all callers (including WebSocket) originate from authorized IP addresses; deny takes precedence over allow.
- Handle
auth_ip_not_allowedandauth_account_blockedin your 403 handling; both arefatal— retrying will not help, so correct the source IP or contact technical support.
Reference
V1.5.9
2026-06-10
Enhancement: Time-Base Fields Added to Session Resume
Three time-base fields were added to session resume to help clients align with the server clock and the transcript timeline.
Added
session_startednow includesserver_time(the server's current unix time in milliseconds, as a clock-base reference; clients can estimate clock skew by comparing it against the client time at the moment of receipt).resume_oknow includesserver_last_offset_ms(the transcript timeline position at the breakpoint, in milliseconds, based on the length of audio already processed).resume_oknow includesserver_recording_ms(the recording-head timestamp the transcript timeline resumes from after reconnect, in milliseconds, including silence; lets the client align its recording-second header to the same timeline the transcript uses).
Note
- Do not mix the two notions of time: the grace period (
resume_grace_seconds) is measured in wall-clock time and keeps counting down during the disconnect, whereasserver_last_offset_msis based on the audio timeline and is frozen during the disconnect. Always use wall-clock time to decide whether a reconnect is still possible.
Reference
V1.5.8
2026-06-08
New Feature: WebSocket Session Resume
When a WebSocket connection is unexpectedly dropped, clients may reconnect within a grace period (default 45 seconds) with their resume_token to rejoin the original recording session — keeping the same recording_id and sentence ids (sid), with the transcript timeline continuing from the breakpoint. No need to restart the whole recording.
Added
session_startednow includesresume_tokenandresume_grace_seconds(store them).- New
resume_okevent (resume succeeded, includesserver_last_sid). - 4 new resume error codes:
resume_token_invalid,resume_grace_expired,resume_ownership_mismatch,resume_unavailable(allerrorseverity, not fatal). - Added a reconnect implementation example (JavaScript) to the Connection doc.
Client guidance
- Store
resume_tokenonsession_started. - On detecting a dropped connection, within the grace period: obtain a new Ticket → reconnect with
Sec-WebSocket-Protocol: ["ticket.<new>", "resume.<token>"]. - After
resume_ok, start a fresh audio stream as afterstart(WebM must send a new container). - On any
resume_*error → obtain a new Ticket and send a freshstart.
Unchanged
- Existing clients that do not use session resume require no changes; disconnect behavior is unchanged (always a fresh
start). - Audio during the few seconds of disconnect is not recovered and is not billed (the timeline continues seamlessly).
api_keyis the trust boundary: a "last connection wins" policy applies; do not share a singleapi_keyacross trust domains.
Reference
V1.5.7
2026-05-20
Documentation Update (No API Behavior Changes)
The public API behavior is completely unchanged. This release is a documentation supplement and wording revision.
New Usage Guide: Summary Prompt Customization
Added the Summary Prompt Customization Guide, consolidating the summary customization specs that were previously scattered across 6 reference documents into a single guide:
- The mutual-exclusion rules and use cases for the
builtinandcustomsummary modes - The corresponding fields for the three entry points: REST
POST /api/v1/summary, the WebSocketstartaction, and SSEregenerate/summary - Transcript record fields (including the
summary_prompt_snapshotaudit field and thesummary_fallback_level/summary_dropped_segmentsfallback audit fields) - A Profanity and Sensitive-Word Handling section, integrating the three paths (customer prompt -> neutral mode, transcript -> STT
profanity_handlingmasking, transcript -> summary-layer segment omission) and explicitly stating that the API layer does not proactively reject requests containing sensitive words - The built-in safety guard (content-neutralization guidance, prompt-injection protection) and character-length limits
- Complete examples for Node.js, Python, and WebSocket
The "Feature Guides" table on the documentation home page now includes an entry for this guide.
Documentation Wording Revision
Public-facing wording refined to use more general descriptions (field values such as summary_fallback_level are unchanged; wording only).
Reference
- guides/summary-customization.md (new)
- Documentation home page
V1.5.6
2026-05-19
Documentation Alignment Fixes (No API Behavior Changes)
This release is a documentation proofreading pass; the public API behavior is completely unchanged. If you previously implemented against the older documentation, please adjust to the current spec for the items below.
Token Formats
broadcast_token: a 4-character short code (character set a-z0-9)viewer_access_token: a 64-character alphanumeric string (not a JWT, no payload structure; do not attempt to parse it)
HTTP Status Codes
sse_missing_target_lang/sse_unsupported_language: 422broadcast_token_invalid(viewer verify endpoint): 401
Error Code Strings
POST /api/v1/importsinsufficient quota:stt_quota_exceeded- Broadcast not found on viewer SSE:
broadcast_session_not_found - Broadcast at capacity on viewer SSE:
broadcast_capacity_exceeded - The
contextfor thesse_translation_failederror event issse
WebSocket Event Naming
retranslatesuccess event:action: "translation"- Audio upload failures: delivered via a
type: "error"envelope (error_codeisstorage_upload_failed/storage_connection_failed/storage_queue_full); there is no separateupload_erroraction
Newly Documented Error Codes
| Endpoint / Action | Error Code | Description |
|---|---|---|
WebSocket set_name | set_name_empty / set_name_too_long / set_name_not_ready | Replaces the older name_too_long |
WebSocket audio | audio_process_failed | Audio processing fails repeatedly (HTTP 500; reconnecting is recommended) |
Reference
- reference/rest/viewer.md
- reference/websocket/events.md
- reference/websocket/voice-translation.md
- reference/sse/retranslate.md
V1.5.5
2026-05-13
Breaking Change: The Summary API Is Now Mode-Aware
The "template + custom_prompt combined" design introduced in V1.5.4 is now mutually exclusive: on each summary request you must choose either mode=builtin (apply the built-in template) or mode=custom (your prompt fully replaces the built-in template).
Clients must migrate: V1.5.4 clients that do not update their fields will receive a 422.
Unified New Fields Across the Three Entry Points: REST POST /api/v1/summary, SSE regenerate/summary, and the WebSocket start action
Old (V1.5.4) -> New (V1.5.5) mapping:
| Old Field | New Field | Notes |
|---|---|---|
template / templateSlug / summary_template | Same name (builtin mode only) | Unchanged, but must not be sent in custom mode |
custom_prompt / customPrompt / summary_custom_prompt | prompt / summary_prompt (custom mode only) | Renamed |
custom_prompt_slug / customPromptSlug / summary_custom_prompt_slug | prompt_slug / summary_prompt_slug (custom mode only) | Renamed |
persist_custom_prompt / persistCustomPrompt | (removed) | Custom mode always snapshots; no opt-in |
custom_instructions | (removed) | Legacy field, no longer supported |
| (none) | mode / summary_mode (required) | New required field, enum builtin / custom |
Mutual-exclusion rules:
mode=builtin:templateis required;prompt/prompt_slugmust not be sentmode=custom:prompt/prompt_slugis required;templatemust not be sent- Violations -> 422
summary_mode_field_mismatch
GET /api/v1/tasks/ Response Fields
Within data.tasks[]:
- Added
summary_mode(builtin/custom/null) summary_templatenow returns the effective slug (in custom mode it returns your slug, identical to theprompt_slugyou submitted)- Removed
summary_custom_prompt_slug(merged intosummary_template)
Backward compatibility: recordings without a generated summary have summary_mode set to null; existing builtin-mode recordings keep their original summary_template value.
Transcript Record Structure Changes
New top-level fields (not nested under the summary object):
| Field | Description |
|---|---|
summary_mode | builtin / custom |
summary_template | effective slug — builtin -> the built-in slug; custom -> your slug |
summary_plain_text | bool |
summary_prompt_snapshot | Present only in custom mode; the prompt content you passed in verbatim (not written in builtin mode) |
summary_fallback_level | Present only when a fallback was triggered (value 2 or 3); indicates that this summary went through an automatic content-filter fallback path. Omitted when the summary succeeds directly |
summary_dropped_segments | Present only when fallback_level=3; the indices of the transcript segments that were dropped (an array of integers in original order) |
In addition to the existing text, the init_summary event of GET /api/v1/sse/history/transcribe/{taskId} now adds mode / template / plain_text / prompt_snapshot (populated only in custom mode) for client traceability, plus fallback_level / dropped_segments (populated only when a fallback was triggered).
New Outbound WebSocket Events
summary_done: summary generation completed (includessummary_mode/summary_template(effective) /summary_plain_text/tokens_used/summary_fallback_level/summary_dropped_segments; does not includefinal_content)summary_error: summary generation failed (includeserror_code/message)
Clients no longer need to poll the transcript record to determine whether the summary is complete.
Automatic Content-Filter Fallback for Summaries
When a custom-mode prompt or transcript content is blocked by content filtering, the system handles it through an automatic multi-step fallback instead of failing outright. If some transcript segments still cannot be processed, they are omitted and reported via summary_dropped_segments. If even the fallback cannot produce a summary, a summary_error event is emitted with error_code=llm_content_filtered.
Client-side handling:
- Use
summary_fallback_levelto show a UI notice indicating the summary was produced through a content-filter fallback path - Use
summary_dropped_segmentsto inform the user which segments were actually omitted
Spec scope: In this release the fallback applies to two paths: WebSocket realtime summaries (auto-generated when a recording ends) and file-import summaries. Fallback integration for the SSE
regenerate/summaryendpoint is a follow-up; in the current version it still returnsllm_content_filteredwhen blocked.
Custom-Mode Prompt Safety Rule (New in V1.5.5)
The built-in safety guard applied to custom-mode prompts now adds a rule instructing the LLM to "summarize the intent of any colloquial, emotional, or sensitive wording in the source in neutral, objective language, avoiding verbatim quotation or repetition." This rule is enforced by the backend and is not exposed for client configuration; its purpose is to reduce the chance of triggering the content filter on the first attempt.
The guidance itself is not retained. Your original prompt is still stored via the summary_prompt_snapshot field as an audit reference, complementing summary_fallback_level:
summary_prompt_snapshot= your intent (the original prompt content)summary_fallback_level= the actual execution path taken by the automatic fallback
Prompt Safety in Custom Mode
Avoid concatenating untrusted end-user input directly into prompt.
New Error Codes
| Error Code | HTTP | Trigger Condition |
|---|---|---|
summary_invalid_mode | 422 (SSE) / 400 (others) | mode is not builtin / custom |
summary_mode_field_mismatch | 422 / 400 | The mode and field combination is inconsistent (a required field is missing, or a forbidden field was sent) |
summary_prompt_too_long | 422 / 400 | prompt exceeds 2000 characters |
summary_prompt_slug_too_long | 422 / 400 | prompt_slug exceeds 64 characters |
summary_prompt_slug_invalid | 422 / 400 | prompt_slug contains control characters (\n / \r / \t / \0, etc.) |
Client Recommendations
- Add the required
modefield — change existing calls usingtemplateSlug=meetingtomode=builtin&template=meeting - Rename fields —
customPrompt->prompt,customPromptSlug->promptSlug; these two fields are only used inmode=custom - Remove
persistCustomPrompt— custom mode preserves the prompt content automatically - Change
templateSlugtotemplate— and only use it inmode=builtin - Transcript records now use top-level fields — no longer nested under the
summaryobject - Clients can determine whether a summary was saved from the done event / summary_done event — check
persisted: true/false; you no longer need to infer it from the HTTP method
Reference
reference/rest/summary.md- reference/sse/regenerate-summary.md
- reference/rest/summary-templates.md
- reference/websocket/voice-translation.md
- reference/websocket/events.md
V1.5.4
2026-05-12
New Feature: Customer Prompt Customization for Summaries
Enterprise customers can now add their own rules to the summary API without modifying the built-in template. This release adds three orthogonal client parameters and splits the summary regeneration endpoint into "preview" and "save" verbs, avoiding the design gap of an HTTP GET with side effects.
Fully backward compatible — not sending the new fields = behavior identical to the previous version.
New Fields for POST /api/v1/summary
| Field | Type | Limit | Description |
|---|---|---|---|
custom_prompt | string | <=2000 characters | Customer custom instructions appended after the built-in template |
custom_prompt_slug | string | <=64 characters, Unicode, no control characters | A client-defined template identifier (pass-through) |
plain_text | bool | Default false | Request plain-text output |
persist_custom_prompt | bool | Default false | Opt-in: whether the done event echoes the custom_prompt content |
The SSE start / done events also add the corresponding fields (custom_prompt_slug, plain_text, final_content, custom_prompt_snapshot); see reference/rest/summary.md.
/api/v1/sse/regenerate/summary/{taskId} Split Into Two Endpoints
| Method | Purpose | Persists Result | Saves Transcript | Billed |
|---|---|---|---|---|
| GET | Preview (dry run, compare different prompt results) | No | No | Yes |
| POST | Save (official persistence) | Yes | Yes + bumps revision | Yes |
Client recommendation: If your integration previously relied on "the backend record updating automatically after a GET," switch to POST. GET is now a pure preview and no longer writes any backend state.
The done event adds a
persisted: boolfield, so clients can determine directly from the payload whether this call was saved, without inferring from the HTTP method.
Four New Fields for the WebSocket start Action
summary_custom_prompt / summary_custom_prompt_slug / summary_plain_text / summary_persist_custom_prompt, mapping one-to-one to the REST endpoint fields with the same limits.
New Endpoint: GET /api/v1/summary-templates/{slug}
Exposes the built-in template's full content so enterprise customers can reference the existing baseline when integrating and then decide what to add via custom_prompt.
GET /api/v1/summary-templates also adds a ?category=summary|medical|legal|all filter and a data[].category field in the response (default summary, backward compatible).
New Error Codes
| Error Code | HTTP | Trigger Condition |
|---|---|---|
custom_prompt_too_long | 400 | custom_prompt exceeds 2000 characters |
custom_prompt_slug_too_long | 400 | custom_prompt_slug exceeds 64 characters |
custom_prompt_slug_invalid | 400 | custom_prompt_slug contains control characters |
template_not_found | 404 | The template for the specified slug does not exist or is disabled |
invalid_category | 400 | ?category= is not in the allowlist |
Behavior Changes
summary_text_empty/summary_text_too_longHTTP status code fix: these previously fell through to 500 because they were not explicitly mapped; this release fixes them to a semantically correct 400.- The
POST /api/v1/summaryerror eventdetailsno longer includes the LLM raw error: the raw error goes only to the server log; thedetailsreturned to the client retains only theproviderindicator. - GET preview is still billed: generating a preview consumes the same resources as saving one. Repeated GET calls are billed repeatedly, but they do not change any stored content.
Path and Field Naming Conventions
customPromptSlugis a customer-defined pass-through identifier (semantically different from the existingtemplateSlug, which is validated for existence). In naming terms, the former is "for client traceability" and the latter is "for looking up the VAS built-in template."summary_custom_prompt_slugis recorded with each summary, so you can later query which customer template a summary corresponds to.custom_prompt_snapshot(opt-in) is stored with the transcript record only when the customer setspersist_custom_prompt=true.
Security Controls
- All endpoints require API Key authentication
- The VAS server log does not log
custom_promptor the full transcript (it logs only the length and slug) - LLM error messages are sanitized (the raw error is not exposed to the client)
custom_promptis fully isolated across tenants (session-scoped, no memory persistence)
Bug Fixes and Internal Improvements
- The WebSocket
startactionrecording_idfield deprecation target version is unified to V2.0.0 (events.md previously said V1.6.0, inconsistent with code comments) - The SSE
sse-api.mdbroken TOC anchor is fixed (it pointed to the audio section, but that content has been moved to the standalonereference/sse/audio.md) - Improved text sanitization so Chinese, Japanese and Korean characters and emoji are no longer mis-split or wrongly rejected
- The summary regeneration full text now has a 100,000-character upper limit
Reference
reference/rest/summary.md(new)- reference/rest/summary-templates.md (added
?category=andGET /{slug}) - reference/sse/regenerate-summary.md (rewritten as a two-endpoint GET / POST spec)
- reference/websocket/voice-translation.md (added the 4
summary_custom_prompt*fields + 3 error codes)
V1.5.3
2026-05-07
Breaking Change: speaker_id Naming Inversion
To support speaker editing, V1.3.12 added the original_speaker_id field to preserve the original ID, but it left a design gap where "the same name means different things at different stages": for WebSocket realtime recording, speaker_id is the original ID (e.g., Guest-1), but after an SSE historical audio load, speaker_id becomes the display name (e.g., Manager Wang, with the alias applied). Frontends often picked the wrong field and passed it to PATCH /speakers/reassign.
This release performs a one-time inversion that is not backward compatible:
| Old Name | New Name | Semantics |
|---|---|---|
speaker_id (display name) | speaker_label | Display label (after alias is applied; mutable, human-readable) |
original_speaker_id (original ID) | speaker_id | Original speaker ID (immutable, always stable) |
After the inversion, speaker_id consistently refers to the original ID across all interfaces (WebSocket / SSE / REST); the new speaker_label represents the display label after the alias is applied. Speaker editing (rename / reassign / merge) always uses speaker_id as the locating key.
REST API Field Changes
PATCH /api/v1/tasks/{taskId}/speakers/rename
| Location | Old Field | New Field |
|---|---|---|
| Request body | original_name | speaker_id (max 100 characters) |
| Request body | new_name | new_label (max 100 characters, no control characters \x00-\x1F / \x7F or newlines) |
| Response data | original_name | speaker_id |
| Response data | new_name | new_label |
speaker_id can still also accept a display label for chained renaming (e.g., first rename Guest-1 to "Manager Wang," then use "Manager Wang" to rename to "Director Wang"); the resolved response speaker_id is always the original ID.
PATCH /api/v1/tasks/{taskId}/speakers/reassign
| Location | Old Field | New Field |
|---|---|---|
| Request body | target_speaker_id | Unchanged (semantics already aligned to the original ID) |
| Response data | new_speaker_name | new_speaker_label |
target_speaker_id must be the original ID (taken from init_sentence.speaker_id); reassign does not accept a display label.
PATCH /api/v1/tasks/{taskId}/speakers/merge
| Location | Old Field | New Field |
|---|---|---|
| Request body | source_speaker_id / target_speaker_id | Unchanged (still accepts the original ID or the current display label) |
| Response data | target_speaker_name | target_speaker_label |
WebSocket Event Changes
| Event | Old Field | New Field |
|---|---|---|
rename_speaker action body | original_name / new_name | speaker_id / new_label |
result event origin / translations[lang] | only speaker_id (mixed with display name) | speaker_id (original ID) + speaker_label (display label) |
speaker_renamed event | original_name / new_name | speaker_id / new_label |
speaker_reassigned event | new_speaker_name | new_speaker_label |
speakers_merged event | (missing target label) | added target_speaker_label |
SSE Event Changes
| Event | Old Field | New Field |
|---|---|---|
init_sentence | speaker_id (display name) + original_speaker_id (original ID) | speaker_id (original ID) + speaker_label (display label) |
Broadcast viewer origin / translation | only speaker_id (mixed) | speaker_id + speaker_label |
Broadcast viewer speaker_renamed / speaker_reassigned / speakers_merged | same as the corresponding WebSocket events | as above |
The behavior and fields of init_metadata.speaker_aliases (the "original ID -> display label" mapping) are unchanged.
Client Recommendations
- Customers using WebSocket realtime recording: before upgrading, sync the handling of
result.origin.speaker_idand the newresult.origin.speaker_label; change the rename body to{ "speaker_id": "...", "new_label": "..." } - Customers using SSE historical audio:
init_sentence.speaker_idis now the original ID (previously the display name); switch tospeaker_labelfor display - Customers doing speaker editing (rename / reassign / merge):
- rename -> use
speaker_id(either the original ID or the current display label) +new_label - reassign ->
target_speaker_idmust be the original ID (taken frominit_sentence.speaker_id; you cannot send a display label) - merge ->
source_speaker_id/target_speaker_idcan still be the original ID or the current display label
- rename -> use
- Customers integrating TXT/SRT/CSV export:
new_labelnow has control-character/newline validation; if you previously sent labels containing newlines, you will now receive a 422, so change to single-line content - Customers who do not do speaker editing and only consume transcript text: the impact is minimal; the only behavior difference is that if old code rendered
speaker_iddirectly as the display name, it must switch tospeaker_label
Data Compatibility
- Not backward compatible: old transcript data (V1.3.12 ~ V1.5.1, containing
speaker+original_speaker_id) requires a data conversion before it can be read in the new version; there is no cross-version data retention commitment during the POC phase - New recordings are unaffected: transcript blobs created after V1.5.3 use the new fields directly
Documentation Update
- reference/rest/speakers.md: the body / response of all three endpoints (rename / reassign / merge) are fully aligned
- rest-api.md L2280–2455: the speakers summary is aligned to the new fields
- reference/websocket/voice-translation.md L975–1127: the rename_speaker / reassign_speaker / merge_speakers actions
- reference/websocket/events.md L405–495: the speaker_renamed / reassigned / merged events
- websocket-api.md L1310–2552: both rename / reassign / merge sections (actions first, events second) are aligned
- reference/sse/history.md L135–195: the
init_sentenceschema + client recommendations - reference/sse/broadcast-viewer.md L300–365, L605–630: viewer broadcast events + JS examples
- sse-api.md L240–475, L750–795: broadcast origin/translation + history init_sentence schema
- guides/speaker-management.md L140–390: examples + JS handler
- examples/curl.md, examples/python.md, examples/javascript.md: all rename / reassign / merge examples + TS interface
Reference
- REST - Speakers API
- WebSocket - Voice Translation
- WebSocket - Events
- SSE - Historical Audio Streaming
- Guide - Speaker Management
V1.5.1
2026-05-07
Bug Fix: POST /api/v1/imports Adds Length Validation for Terminology / Correction Fields
The length limits promised in several places in the documentation (e.g., a term's max of 100 characters) were previously not actually enforced on the file-import path, and overly long content was silently accepted. This release restores them, aligning behavior with the documentation's promises.
Behavior Changes (Aligning With Documented Promises)
POST /api/v1/imports adds 422 rejection conditions for the following fields (previously accepted):
| Field | Limit |
|---|---|
terminology.<lang> | Array, max 500 terms (per language) |
terminology.<lang>[].term | string, max 100 characters |
terminology.<lang>[].boost | numeric, 0.5–5.0 (optional, default 1.0) |
fuzzy_correction.<lang>[].correct | string, max 200 characters |
fuzzy_correction.<lang>[].incorrect[] | string, max 200 characters |
These limits are consistent with the WebSocket
configaction; previously only the WebSocket path enforced them, and this release completes the file-import path.
Client Recommendations
If you previously sent overly long terms (>100 characters) via POST /api/v1/imports, you will now receive a 422. The frontend should check the length before submitting and prompt the user. The WebSocket path is unchanged.
V1.5.0
2026-05-07
No public changes
This release contains no public API changes, and customers need to take no action.
The old naming (recording_id) will be fully removed in V1.6.0; for the related client migration guidance, see V1.4.1 Client Recommendations
V1.4.3
2026-05-07
No public changes
This release contains no public API changes, and customers need to take no action.
V1.4.2
2026-05-07
No public changes
This release contains no public API changes, and customers need to take no action.
V1.4.1
2026-05-06
Naming Unification: task_id as the Cross-Interface Task Identifier
Previously, the same task had different field names across interfaces (WebSocket used recording_id, Webhook used task_id, and some REST path variables mixed {recordingId} / {taskId}), forcing integrators to reconcile the three naming schemes themselves. This release starts the naming-unification cycle; new integrations should use task_id consistently.
WebSocket Changes (Backward Compatible)
- The
session_startedevent payload now carries bothtask_idandrecording_id, and their values are exactly the same (the UUID of the same recording) - The
recording_idfield is marked as Deprecated; it is still emitted normally and is scheduled for removal in V1.6.0 - Documentation enhancement:
session_idis the WS connection-level identifier (invalidated when the connection ends), which is a different level fromtask_id(the task identifier)
REST API Changes (Backward Compatible)
Added /api/v1/tasks/{taskId}/... alias paths that behave exactly the same as the existing /api/v1/recordings/{recordingId}/...:
| Recommended (from V1.4.1) | Deprecated (removed in V1.6.0) |
|---|---|
PATCH /api/v1/tasks/{taskId}/speakers/rename | PATCH /api/v1/recordings/{recordingId}/speakers/rename |
PATCH /api/v1/tasks/{taskId}/speakers/reassign | PATCH /api/v1/recordings/{recordingId}/speakers/reassign |
PATCH /api/v1/tasks/{taskId}/entries/{sid} | PATCH /api/v1/recordings/{recordingId}/entries/{sid} |
Client Recommendations
- New integrations: use the
task_idfield and the/api/v1/tasks/{taskId}/...paths consistently to avoid migrating again later - Existing integrations: no immediate change required.
recording_idand/api/v1/recordings/...remain available throughout the V1.x period; we recommend migrating on your schedule, at the latest before V1.6.0 ships - ID alignment logic: if you depend on both WS and Webhook, you can align the WS
task_id(or the old namerecording_id) directly with the Webhookdata.task_id; all three are the same UUID - Do not use
session_idfor alignment:session_idis meaningful only within the WS connection lifecycle and does not appear in Webhook or REST
Removal Timeline Announcement (V1.6.0)
V1.6.0 will remove the recording_id field from the WS payload and remove the /api/v1/recordings/{recordingId}/... paths. The detailed timeline will be announced separately before V1.6.0 ships.
Unchanged Items
- Webhook payload: the existing
data.task_idnaming is unchanged - Existing
/api/v1/tasks/{taskId}/...endpoints: unchanged
V1.4.0
2026-05-06
New Feature: Source-Text Editing for Historical Recordings + Automatic Retranslation
Users can correct STT recognition errors and regenerate translations; for the workflow, see Entries API Typical Workflow.
- New endpoint
PATCH /api/v1/recordings/{recordingId}/entries/{sid}: edit a single sentence's source text; on the first edit it automatically backs up the original STT output tooriginal_text_raw, recordsoriginal_text_edited_at, and clears the TTS cache for all languages of that sentence - New endpoint
GET /api/v1/sse/recordings/{taskId}/entries/{sid}/retranslate: retranslate a single sentence (you can specify languages or retranslate all existing languages), with optimistic locking (expectedRevision) - Editing and retranslation are decoupled: PATCH only changes the source text and does not touch the translation; the frontend can decide when to trigger retranslation
Historical Record SSE Exposes Edit Markers
The historyTranscribe init_sentence event carries original_text_raw (the STT original) and original_text_edited_at on edited sentences, so the frontend can show an "edited" marker and a "restore original" function.
Security Fixes
retranslate/retranslateSummaryadd a user filter: these two existing SSE endpoints previously had a horizontal privilege vulnerability (IDOR) that allowed reading other users' recordings. This release adds the permission check; other users' recordings now returnrecording_not_found.- Retranslation / summary regeneration requires the recording to be completed: the four endpoints
retranslate/retranslateSummary/retranslateEntry/regenerateSummaryrequireprocessing_status === completedto avoid racing with the in-progress flow. When not completed, they returnrecording_not_completed.
New Error Codes
| Error Code | HTTP | Description |
|---|---|---|
recording_not_completed | 422 | The recording has not finished processing; retranslation / editing / summary regeneration is not allowed |
entry_not_found | 404 | The specified sentence was not found |
entry_text_empty | 422 | The sentence's source text is empty |
entry_text_too_long | 422 | The sentence's source text exceeds the 2000-character limit |
transcript_revision_conflict | 409 | The transcript has been modified by another request (optimistic-lock conflict) |
See error-codes.md.
Client Recommendations
- After editing the STT source text: we recommend triggering single-sentence retranslation SSE immediately after the PATCH, passing the
revisionfrom the PATCH response asexpectedRevisionto avoid concurrent overwrites - Showing the edit marker: determine whether a sentence has been edited by the presence of the
original_text_rawfield in theinit_sentenceevent ('original_text_raw' in data); do not use text comparison (the user may edit and then change it back to the original value) - Recording status: calling retranslation / editing / summary regeneration on a recording that is not completed returns
recording_not_completed; the frontend should block these operations in the UI untilprocessing_status === completed
V1.3.13
2026-05-06
Behavior Changes (Breaking Changes)
- WebSocket
audio_formatlocked topcmandwebm: the previously accepted 5 formats (pcm/webm/mp3/wav/m4a) are narrowed to accepting onlypcmandwebm, consistent with the existing spec inreference/websocket/voice-translation.md. Customers who sendmp3/wav/m4awill now receiveaudio_format_unsupported(previously these were silently decoded, which was undocumented implicit behavior). File imports still go throughPOST /api/v1/importsand are unaffected.
Documentation Update
- Audio download Content-Type is always
audio/mp4: rest-api / SSE audio / tasks export / history playback / curl / javascript documentation in several places is unified to "all recording audio is returned in an M4A container (AAC encoding)," removing the previous circular "dynamically determined" description. - Supported file-import formats narrowed to
mp3/wav/m4a: removed mentions ofmp4andwebmfrom the documentation to align with the formats actually accepted (guides/file-import.md, reference/rest/imports.md).
Client Recommendations
- Customers using the WebSocket
startaction: be sure to explicitly specifyaudio_formataspcmorwebm; if you previously relied on the undocumented implicitmp3/wav/m4asupport (very rare scenarios), switch to the File Import API. - Customers downloading recording audio: all new recordings have Content-Type fixed to
audio/mp4with the.m4aextension. If older recordings still exist in storage, downloads may still returnaudio/webm; we recommend keeping a handling branch for the old extension to cover historical data.
Reference
- WebSocket - Voice Translation
- REST - Task Audio Export
- SSE - Historical Audio Streaming
- Guide - File Import
- Guide - History Playback
V1.3.12
2026-05-04
Note: Inverted in V1.5.3: the
original_speaker_idfield and the "speaker_idis the display name" design introduced in this version have been superseded by the naming inversion in V1.5.3. This section is kept as a historical record; new integrations should refer directly to the V1.5.3 spec and do not need to implement this version's client recommendations.
New Feature
- History SSE adds fields to align with the Transcribe speaker-editing UX: the historical record's
init_metadataandinit_sentenceevents each add a field, allowing the frontend to fully reuse the realtime recording page's speaker-editing menu (single-sentence reassignment + global rename).init_metadataaddsspeaker_aliases(object): the "original speaker ID -> display name" mapping. When there are no aliases it is{}(an empty object, not an empty array). It lets the frontend perform a name-collision precheck before sendingPATCH /speakers/rename, covering the implicit conflict of "an original ID that exists on the backend but does not appear on screen because it was renamed."init_sentenceaddsoriginal_speaker_id(string|null): the original speaker identifier without alias substitution, provided as the source for thetarget_speaker_idofPATCH /speakers/reassign.- Old-data fallback: if an older transcript record has no
original_speaker_id, the output automatically falls back tospeaker_id, preventing the new field from being null and disabling the editing entry point for old recordings.
Behavior Changes
- No breaking change. Both fields are pure additions; clients that ignore unknown fields are unaffected, so no version negotiation is needed.
Documentation Update
- sse-api.md L156-198: added the new field descriptions to the
init_metadata/init_sentenceexamples and field tables - reference/sse/history.md L103-180: added the detailed reference schema accordingly
Client Recommendations
- Customers doing speaker editing on the history detail page: get the original ID for reassign from
init_sentence.original_speaker_id(do not usespeaker_id, which is the display name with the alias applied); useinit_metadata.speaker_aliasesfor the name-collision precheck before a rename. - Customers who do not do speaker editing: you can ignore the new fields; existing parsing behavior is unaffected.
Reference
- SSE API - Historical Record Events
- Reference - Historical Record SSE
- Recording Speaker API
- Speaker Management Guide
V1.3.11
2026-05-04
Behavior Changes (Breaking Changes)
- STT rejects the bare
encode (the V1.3.10 changelog claimed it was removed, but it was not actually in effect): customers who sendenwill receive a 422invalid_transcription_language; use a full BCP 47 code such asen-US/en-GBinstead. - TTS removes 4 locales that were never usable:
it-CH,ar-IL,ar-PS,en-GH. TTS for these 4 locales never actually worked; previously, requesting their voices would fail. STT still supports these 4 locales.
New Feature
- TTS expanded to 154 languages and 325 voices
- Chinese dialects (4 added):
zh-CN-henan,zh-CN-guangxi,zh-CN-liaoning,zh-CN-shaanxi - South Asian languages (5 added):
bn-BDBengali (Bangladesh),ta-LKTamil (Sri Lanka),ta-MYTamil (Malaysia),ta-SGTamil (Singapore),ur-PKUrdu (Pakistan) - Southeast Asian languages (1 added):
su-IDSundanese (Indonesia) - Eastern European languages (1 added):
sr-Latn-RSSerbian (Latin script) - North American indigenous languages (2 added):
iu-Cans-CAInuktitut (Canadian syllabics),iu-Latn-CAInuktitut (Canadian Latin script)
- Chinese dialects (4 added):
Documentation Update
- languages.md TTS section rewritten, explicitly noting:
- Of the 145 STT locales, 141 are supported on both the STT and TTS sides; 4 (
it-CH,ar-IL,ar-PS,en-GH) are STT-only - Of the 154 TTS locales, 13 are TTS-only (4 zh-CN dialects + 9 other languages)
- Of the 145 STT locales, 141 are supported on both the STT and TTS sides; 4 (
- guides/tts.md numbers updated (142->154 languages, 304->325 voices)
- Documentation home page TTS description updated
Supported Counts
| STT | TTS locale | TTS voice | Diarization |
|---|---|---|---|
| 145 | 154 | 325 | 31 |
Client Recommendations
- Customers using the
enshort code: switch toen-USor another full BCP 47 code. - Customers using
it-CH/ar-IL/ar-PS/en-GHfor TTS: these already failed on the provider side; switch to another locale in the same language family (e.g.,it-CH->it-IT,ar-IL->ar-SA,en-GH->en-NG). STT is unaffected. - Customers who want to use the 13 new TTS-only locales: you can call
GET /api/v1/tts/voices?language=zh-CN-henanetc. directly to get the voice list.
Reference
V1.3.10
2026-04-30
Documentation Update
- languages.md number corrections
- Total speech-recognition languages
119->145 - Speech-translation support
117->143(145 minusjv-IDJavanese andwuu-CNWu Chinese)
- Total speech-recognition languages
This version has a residual issue; see V1.3.11: this version claimed "the bare
enwas removed and the language counts are fully consistent at 145," but the bare"en"was not actually removed (still 146), nor did it handle the TTS-sideit-CH/ar-IL/ar-PS/en-GH(not supported by the provider's TTS) or the 13 missing TTS-only locales. The full alignment fix was completed in V1.3.11.
Client Recommendations
- This version is a documentation-only number correction and does not affect running integrations.
Reference
V1.3.9
2026-04-29
New Feature
- Webhook Secret Bootstrap flow: resolves the contradiction where a client cannot obtain the secret on first webhook integration. The Dashboard adds a "Generate Webhook Secret" button (lazy generation), letting users obtain the secret and configure it on the receiving end first, then go back and set the webhook URL. The probe sent when setting the URL is signed with a secret that both sides agree on, so it passes on the first try.
- New endpoint:
POST /dashboard/api-keys/{id}/webhook/regenerate-secret(Dashboard only, rate limited to 10 requests/min/user) - Behavior: generates a 64-character random secret and stores it; does not send a probe and does not touch the webhook URL; the plaintext is returned once for the Dashboard to display
- Regeneration impact: after execution, the old secret is invalidated immediately; existing receivers will get webhooks with mismatched signatures until they switch to the new secret
- New endpoint:
Behavior Changes
- Clearing the Webhook URL no longer clears the Secret: when
PATCH /dashboard/api-keys/{id}/webhooksetswebhook_urlto null,webhook_secretis left unchanged. The Secret and URL now have independent lifecycles. A customer can generate the secret first and set the URL later; sending an empty URL in the meantime will not lose the secret. - The Dashboard no longer returns webhook_secret in plaintext:
GET /dashboard/api-keys/{id}now returnswebhook_secret_masked(prefix mask + last 4 characters) and ahas_webhook_secretboolean. The plaintext is shown only once right after generation.
Documentation Update
- guides/webhook.md: "Method 2: API Key-level webhook_url" rewritten as a two-step flow (generate secret -> set URL); added a Webhook Secret Lifecycle section; added a Bootstrap callout to the security-verification section.
Client Recommendations
- First integration: in the Dashboard, click "Generate Webhook Secret," copy it to the receiving end's
.env, enable HMAC verification, and restart the service, then go back to the Dashboard and enter the webhook URL. - Existing customers: fully compatible, no changes needed. Existing webhook_url and webhook_secret behavior is unchanged.
- Secret rotation: we recommend that the receiving end briefly accept both the old and new secrets; after the dashboard regeneration, remove the old secret once in-flight webhooks have finished processing.
V1.3.8
2026-04-27
New Feature
- Translation-service-unavailable detection (session-level): added the error code
translation_service_unavailable. When the LLM translation service fails consecutively up to a threshold, the backend emits a session-level error event once, so the frontend can show a global "translation temporarily unavailable" prompt instead of users seeing a page full of individual failed sentences in gray text.- Trigger conditions:
llm_timeout/llm_provider_error/llm_rate_limit/llm_request_failedescalate after 5 consecutive failuresllm_auth_failed/llm_deployment_not_found/llm_quota_exceededescalate immediately after 1 occurrence (configuration/billing issues)llm_content_filteredis not counted (a content issue, not a service issue)
- Deduplication: each session is notified only once; any successful sentence translation resets the count and can trigger it again
- payload:
type: "error",severity: "error"(not fatal — should not disconnect), does not carrysid,detailscontainsprovider,last_error_code,fail_count - Viewer notification: in broadcast mode, all viewers (regardless of language) also receive this event (via the SSE
event: errorchannel)
- Trigger conditions:
Documentation Update (Spec Sync)
Continuing the spec blind spots surfaced by frontend feedback since V1.3.7+, this pass completes:
- error-codes.md — sentence-level error rule: added a sid-rule paragraph below the "Severity Levels" table, explicitly stating that "when an error carries
sid, regardless ofseverity, it should be treated as a sentence-level error and should not disconnect." Afatal+sidcombination only means that sentence failed severely; the session as a whole can still continue. - error-codes.md —
translation_service_unavailableerror-code registration: added this error code and its full trigger-rule description to the "Translation Service Errors" section - websocket-api.md: added a session-level translation error example (no sid, severity error) to the "Error Message Format" section
- sse-api.md — retranslate section adds the per-sid error rule: explicitly lists the spec and payload format for "a failed sentence is re-emitted as
event: errorwithsid+error_code, interleaved withtranslation" (implemented in V1.3.7 but documented only in the reference subdirectory) - reference/sse/broadcast-viewer.md: added a
translation_service_unavailableexample and a specific error-code entry - reference/websocket/events.md: removed an obsolete translation_error action that the service never actually emitted; translation errors are delivered on the standard error channel (type: "error")
Client Recommendations
- Existing sentence-level error handling (
type: "error"withsid) needs no changes. - If you want to show a global "translation service unavailable" prompt, add a listener: when you receive
error_code === "translation_service_unavailable"(withoutsid), show a banner / toast; clear it once any subsequent sentence translation succeeds (you receive atranslationevent again). - Do not treat
translation_service_unavailableas a disconnect signal — STT (the source text) continues to operate.
Reference:
- Error Code Reference
- WebSocket API – Error Message Format
- SSE API – retranslate Event Format
- reference/sse/retranslate.md
- reference/sse/broadcast-viewer.md – Specific Error Codes
V1.3.7
2026-04-24
Behavior Changes
- Realtime recording: silent tasks now follow the normal completion flow: when a realtime recording (WebSocket) is silent throughout, is noise, or cannot recognize any sentence, it now still produces an empty transcript (
entries: []) and ends with atask_completeevent. This behavior aligns with the V1.3.5 file-import flow; the realtime and import sources now share the "zero recognition results is treated as a legitimate completed" semantics. - SSE historical record: no longer returns
sse_transcript_not_foundin silent scenarios:GET /api/v1/sse/history/transcribe/{taskId}no longer returns thesse_transcript_not_founderror for silent tasks; instead it sends the full event sequence (init_metadata → init_summary(text='') → init_done(totalSentences=0)). Clients should usetotalSentences === 0to detect this and show a "no speech content" empty state.
Bug Fixes
- Fixed the History page getting stuck on "processing" for silent recordings: previously, if a realtime recording was silent throughout, no transcript record was produced, but
task_completestill sent thetask_id, causing the frontend to receivesse_transcript_not_found(semantically "not finished processing") when loading the historical record, leaving the UI stuck on loading forever. After the fix, the realtime path matches the import path and always uploads the transcript (even with empty entries).
Client Recommendations
- If you previously had "retry / polling" handling logic for
sse_transcript_not_found, you may keep it as a defensive fallback (e.g., for blob upload delays), but you should no longer use it to determine "the task has no speech" — switch toinit_done.totalSentences === 0. - We recommend the UI prompt possible reasons when
totalSentences === 0(volume too low, silent throughout, recognition language does not match the audio), consistent with the V1.3.5 import-scenario wording.
Documentation Update
- Historical Record SSE adds a "Boundary scenario: no speech content" section and corrects the handling-recommendation description for
sse_transcript_not_found
Reference:
- Historical Record SSE
- Import Progress SSE
- File Import Guide – Behavior When Audio Cannot Be Recognized
V1.3.6
2026-04-23
New Feature
- Tasks API: added
POST /api/v1/tasks/{taskId}/force-fail: force-marks as failed a task stuck in a non-terminal state (recording/importing/uploading/pending/processing)- The body can optionally include
reason(max 500 characters) - Triggers the
recording.failedwebhook, withpayload.failure_sourceset touser_forced - A task already in a terminal state returns
invalid_processing_status(422)
- The body can optionally include
- Tasks API: added
POST /api/v1/tasks/{taskId}/retry: re-queues a task in thefailedstate for processing- Prerequisites:
processing_status = failedandaudio_status = successandtranscript_status = success - Not meeting the prerequisites returns
invalid_processing_status(422); thedetailsfield carriesaudio_status/transcript_statusto help with diagnosis
- Prerequisites:
Behavior Changes
- Error code
invalid_processing_status(422) expanded scope: now also used as the common response forforce-failandretry;detailscarriescurrent_status, and theretryscenario additionally carriesaudio_statusandtranscript_status
Documentation Update
- Tasks API adds documentation for the
force-failandretryendpoints - Error Code Reference adds the
invalid_processing_statusentry and a "Processing Status Mismatch" subsection
Reference:
V1.3.5
2026-04-22
Behavior Optimization
- File import: empty recognition-result filtering: audio imports now filter out empty recognition results (caused by silence, very low volume, noise, or a language mismatch), so imports no longer produce empty 00:00 placeholder segments
- Zero recognition results is a legitimate
completedstatus: in this scenario the import task still ends withstatus: completed(notfailed), thetask_idis produced normally, but the subsequently loaded transcriptentriesis an empty array andsegments_countis0 - Budget deducted by actual duration: unrecognizable audio is still deducted from the monthly budget based on the audio duration (no refund)
Client Recommendations
- After loading the transcript (SSE
/api/v1/sse/history/transcribe/{taskId}), if the cumulative sentence count is0, show a "no speech content was recognized in this audio" empty state - Do not treat zero recognition results as an error branch; follow the completion branch and judge by the sentence count
- We recommend the UI also prompt possible reasons (volume too low, silent throughout, recognition language does not match the audio)
Documentation Update
- File Import Guide adds a "Behavior When Audio Cannot Be Recognized" section
- Imports API adds a
completedboundary-scenario note under thestatustransition - Import Progress SSE adds a behavior note for zero recognition results under the
completedevent
Reference:
V1.3.4
2026-04-22
New Feature
- Tasks API: added
GET /api/v1/tasks/{taskId}/transcript/export: download a task's transcript, supporting five formats —txt,srt,sbv,vtt,csv- The output includes the source text and all translation languages
- CSV starts with a UTF-8 BOM, with columns
index,start,end,speaker,text,<one column per translation language>and times inHH:MM:SS(no milliseconds) - SRT times are
HH:MM:SS,mmm; SBV times areH:MM:SS.mmm, with the source text and translations joined into a single line with|; VTT uses theWEBVTTheader - The filename uses
{recording name}-transcript.{ext}(RFC 5987 UTF-8 encoded) - Added the error code
recording_transcript_not_ready(422)
Behavior Changes (Breaking)
- Speaker diarization and multi-language mutual exclusion is now a hard rejection: when
recognition_mode: multi_speakeris combined with multipletranscription_languages, it previously emitted a warning and automatically truncated to the first language; it now directly returns thediarization_multilang_conflicterror and refuses to start- The error severity is changed from
warningtoerror - The frontend must restrict "speaker diarization" and "multi-language" to one or the other before the user submits
start, or handle this error and guide the user to adjust the settings - Affected endpoint: WebSocket
voice-translation / start
- The error severity is changed from
Documentation Update
- Tasks API adds the full
transcript/exportspec and output examples for all five formats - Error Code Reference adds
recording_transcript_not_ready - curl, Python, JavaScript examples add a "Task Export" section
- Documentation home page API reference table: Tasks endpoint count updated from 8 to 9
Reference:
- Tasks API - transcript/export
- Error Code Reference
- WebSocket API Reference
- Speaker Diarization Guide
- Voice Translation Guide
V1.3.3
2026-04-21
New Documentation
- Tasks API: added the full documentation for the
GET /api/v1/tasks/{taskId}/audio/exportendpoint (the implementation existed but the documentation was missing), including parameters, dynamic Content-Type, error codes, and a frontend download example - Explained the difference between this endpoint and SSE
/api/v1/sse/audio/{taskId}: the former is for offline download (Content-Disposition: attachment), the latter is for playback (supports Range Requests)
Documentation Fixes
- Fixed the Voice Translation Actions translation-mode
speakersfield description table: the field name is corrected fromspeakertoid, consistent with the JSON example and the actual service behavior - Fixed the documentation home page API reference table endpoint counts: Tasks from 7 to 8 (added audio/export), Broadcasts from 9 to 6 (the original count was wrong)
Reference:
V1.3.2
2026-04-07
Documentation Structure Adjustment
- Removed 3 deprecated old documents (
error-codes.mdV0.6,languages.mdV0.1,authentication.mdV0.1) - Moved
appendix/error-codes.mdandappendix/languages.mdto the root directory, replacing the deprecated versions - Updated all cross-reference links
V1.3.1
2026-03-26
Batch Task Management
- Added
PUT /api/v1/tasks/batch/pin: batch-update pin status, max 100 per call - Added
DELETE /api/v1/tasks/batch: batch-delete tasks, max 100 per call - Both endpoints affect only tasks belonging to the current user; the response includes
affected_count
Batch Broadcast Cancellation
- Added
DELETE /api/v1/broadcasts/batch: batch-cancel broadcasts in the PENDING state, max 100 per call - IDs not in the PENDING state are ignored; the response includes
affected_count
Reference:
Version: V1.24.1 Last Updated: 2026-10-07