Appendix

Changelog Archive

Release notes for V1.18 and earlier. For the latest changes, see Changelog.

Release notes for V1.18 and earlier. For the latest changes, see Changelog.


V1.18.6

2026-09-29

Fix: Complete Sentences in Broadcast Subtitles

  • Viewers who join a broadcast mid-way receive history subtitles as complete sentences.
  • For viewers already watching, a completed subtitle no longer reverts to incomplete.

V1.18.5

2026-09-29

Fix: Import Duration and Billing for Some Transcription Languages

The following applies to file imports whose first transcription language is one of: wuu-CN, yue-CN, zh-CN-shandong, zh-CN-sichuan, ar-DZ, ar-MA, ar-TN, ar-YE, as-IN, gu-IN, kn-IN, mr-IN, or-IN, pa-IN, fr-BE, fr-CA, fr-CH, it-CH, nl-BE, bs-BA, km-KH, ne-NP, si-LK, sw-TZ.

  • Duration and billing follow the actual length of the audio file. In earlier versions they could differ from the actual length, and the transcript could be incorrect.
  • Short files of 1 to 5 seconds complete normally.
  • Files that cannot be read end with failed (error_code import_stt_failed) and are not billed.

V1.18.4

2026-09-29

Fix: Recordings Without Audio Are Not Billed

  • A recording that receives no usable audio at all is not billed, and any credits already deducted are refunded automatically.

V1.18.3

2026-09-28

Fix: Handling Audio That Cannot Be Recognized

  • If audio cannot be recognized during a recording, the recording continues and audio_decode_failed is returned; send a new audio stream to recover.
  • After resuming a disconnected session, always send a new audio stream.
  • Time during which audio cannot be recognized is not billed.

V1.18.2

2026-09-28

Documentation Updates

  • The Broadcast Guide and the API reference now correctly describe share_url: it is not an openable viewer page, so do not share it with viewers directly. The viewer page is provided by your application and connects with the token as described in the Viewer-Side Flow. API behavior is unchanged.

V1.18.1

2026-09-27

Fix: profanity_handling Now Takes Effect

options.profanity_handling in the WebSocket start action previously had no effect: profanity in real-time transcripts was always masked. Starting with this version, all three values take effect:

ValueBehavior
mask (default)Profanity is masked with *
removeProfanity is removed from the transcript
showThe original text is shown

Translations are produced from the processed transcript. This setting applies to real-time recordings only; transcripts of imported files keep the original text.

Fix: Untranslatable Sentences Now Return an Error Code

Some sentences containing abuse or threats could previously receive a translation unrelated to the original. Starting with this version, such sentences return llm_content_filtered instead, the same as the existing "content cannot be translated" handling. This applies to real-time recordings, broadcasts, file imports, and retranslation; during retranslation, such sentences are not billed.

Documentation Updates

  • In Summary Customization, the values of profanity_handling are corrected to mask / remove / show (previously written as removed and raw).

V1.18.0

2026-09-27

Behavior Change: Recordings End Automatically After a Long Silence

A real-time recording now ends automatically after 15 minutes (by default) without detected speech, so a recording someone forgot to stop does not keep being billed.

  • About 2 minutes before the end, the warning stt_silence_warning is sent (the recording continues, and recognized text or resuming restarts the count); at the limit, stt_silence_timeout is sent and the recording wraps up as usual with task_complete.
  • The count does not run while paused, while waiting to resume after a disconnect, or for broadcasts.
  • Use silenceTimeoutSeconds in start to adjust the threshold or turn it off (0); see Automatic End After a Long Silence.

Behavior Change: Recordings That Received No Audio

A real-time recording that received no audio at all previously did not send task_complete, and was marked as failed only about 2 hours later. Starting with this version, you receive task_complete with noAudio: true, the recording is marked failed immediately, and recording.failed is sent with the new failure_source value no_audio. The recording has no transcript or audio file.

Behavior Change: Time Limit for the Broadcast Standby Phase

When the accumulated standby time reaches the limit (30 minutes by default), the session ends automatically: the warning broadcast_standby_warning is sent about 2 minutes before, and broadcast_standby_timeout and status: "ended" at the limit. No recording exists during standby, and it is not billed. See Standby Time Limit.

Behavior Change: Broadcast Parameter Checks in start

broadcast_phase accepts only lowercase standby and live (an empty string is treated as live), and broadcast_token is allowed only with type: "broadcast". Otherwise invalid_parameter is returned and the recording does not start.

Behavior Change: Host Goes Live Again After a Disconnect (Takeover)

When the host goes live again within the grace period with the same broadcast_token, broadcast_recording_ready on the new connection carries a new task_id, and the broadcast is split into two recordings, each notified separately. Previously both parts shared one recording, and the later content overwrote the earlier content.

Behavior Change: Segmented Full Retranslation

Long transcripts can be retranslated in segments with segmented=1; without it, a transcript that is too long returns HTTP 422 retranslate_segmentation_required before the stream starts, with no charge. See Retranslate SSE.

Behavior Change: Import Failures and Failure Descriptions

  • import.failed is sent only once per failed import, and the status stays processing while it is being processed.
  • failed is a final state: you never receive import.completed for the same import afterward. An import that does not start processing for a long time is marked failed (PROCESSING_TIMEOUT).
  • error_message of a failed import and error in recording.failed are now fixed descriptions without internal messages; error codes are unchanged.

Behavior Change: Viewer-Side Rate Limits

  • The broadcast viewer page and password verification are now limited per channel (about twice the channel's maximum number of viewers per minute, with a minimum of 200) instead of 10 per minute per source IP; many viewers joining at once from the same network no longer tend to receive 429.
  • Too many incorrect passwords cause a temporary lock (30 from the same source within 5 minutes, or 100 for the same channel within 5 minutes), during which even the correct password returns 429; too many lookups of nonexistent channels from the same source also return 429 temporarily. See Rate Limits.
  • Floating subtitle audience Token requests are now limited to 30 per minute for the same recording from the same source; too many invalid share links from the same source return 429 temporarily.

Behavior Change: Import Recognition Modes and File Formats

  • An upload with recognition_mode set to multi_language or multi_channel now returns HTTP 422 import_recognition_mode_unsupported (data.details.supportedModes lists the available modes). Previously multi_language was accepted and failed later during processing, and multi_channel returned validation_failed.
  • When the file extension does not match the actual content (for example a .mp3 name with WAV content), the file is now processed according to its content instead of failing.

Behavior Change: Summary Generation Completeness

Applies to Ad-hoc Summary and Regenerate Summary:

  • When generation takes too long, the part completed so far is returned within the processing time limit, with truncated: true in done, and billed as usual; previously neither done nor error might arrive.
  • When the stream stalls or does not end normally, error (sse_summary_regeneration_failed) is sent instead, and nothing is saved or billed; previously an incomplete summary could be delivered as if it were complete.

Other Behavior Changes

  • When speakers are merged, the target speaker's existing sentences whose display name changes because of the merge (for example, when the source speaker's custom name is carried over to the target) also switch to the merged name and are included in affected_sids; this applies to both REST and WebSocket.
  • When the service restarts or is updated, sessions waiting to be resumed are finalized right away and saved as usual; resuming afterward returns resume_token_invalid.
  • A broadcast's current_recording_id has a value only while live, and it is the recording of this go-live; it is null during standby or when not live.

Fixes

  • Viewers who joined after a broadcast went live could see content from the standby phase.
  • A broadcast that went from standby to live could be billed up to 1 extra minute; billing now starts at the moment it goes live.
  • For a broadcast that went live directly, broadcast_recording_ready occasionally arrived before session_started.
  • The full audio file sometimes lacked a short piece at the very end of the recording.
  • For recordings using audio_format: "webm", audio after switching recording devices mid-session is no longer lost.
  • After a brief service disruption, recordings in progress are billed as usual; usage during the disruption is also billed once service recovers.

New Error Codes

Error codeTypeDescription
stt_silence_warningWebSocket, warningAbout to end for lack of speech; the recording continues
broadcast_standby_warningWebSocket, warningThe standby phase is about to reach its time limit; the broadcast continues
broadcast_standby_timeoutWebSocket, fatalThe standby phase reached its time limit and the session ended
retranslate_segmentation_requiredREST, HTTP 422The transcript is too long for a full retranslation; use segmented retranslation
import_recognition_mode_unsupportedREST, HTTP 422This recognition mode is not supported for file imports; use single or multi_speaker

Client Recommendations

  • Recognize stt_silence_warning and broadcast_standby_warning: the recording continues; it has not ended.
  • For sessions during which nobody may speak for a long time (for example a meeting room), send silenceTimeoutSeconds: 0 in start.
  • When task_complete carries noAudio: true, do not read the transcript.
  • After a broadcast takeover, use the recording ID from broadcast_recording_ready on the new connection.
  • For full retranslation of long transcripts, send segmented=1 and continue with the next segment according to truncated in done.
  • When a viewer endpoint returns 429, wait for the number of seconds in Retry-After before retrying; do not retry immediately.
  • For imports, use only single or multi_speaker for recognition_mode.
  • When a summary stream ends with error, discard the fragments already received; when done carries truncated, tell the user the summary is incomplete.

Documentation Updates

  • Error Codes now lists the five error codes that can appear when an import fails, as well as invalid_parameter.
  • Corrected the multi-channel note that every channel must keep sending audio even while silent: a single silent channel does not end the recording.
  • Added the error code too_many_requests to the 429 responses of the viewer endpoints and the floating subtitle audience Token request, and invalid_share to the latter's 403 response.

V1.17.0

2026-09-24

New: Summary Translation Endpoint

POST /api/v1/sse/summary/translate translates summary text supplied in the request into a specified language. It is not tied to a recording and the result is not saved. It is billed at 0.1 credits per 200 characters. See Summary Translation.

Fix: Retranslate Summary

  • After a summary was regenerated in another language, retranslating it could return the original text without translating it. This is fixed.
  • Retranslating a long summary could return only the first part of the translation. This is fixed.

Fix: Multi-Channel Session Resume

In multi-channel mode, after a session resume, the channels stopped producing recognition results and the audio after the resume was not saved. This is fixed.

Fix: Sentences in Progress Lost in Two-Way Translation

  • In manual mode, the last part that was still being recognized when the user stopped speaking could be lost. This is fixed; it is now merged into the sentence's final result.
  • In manual mode, the part being spoken could be lost after a speaking speed change or a reconnect, and the whole sentence could be lost if the recording ended while the user was speaking. This is fixed.
  • In automatic mode, a sentence in progress when the recording was paused could be overwritten by the next sentence after resuming. This is fixed; it is now sent with is_final: true at the moment of pausing.
  • For languages that separate words with spaces (such as English), parts merged in manual mode had no space between them. This is fixed.
  • In single-speaker and two-way translation modes, the same sentence could occasionally appear twice, or a discarded sentence could reappear, right after an operation such as a speaking speed change. This is fixed.

Fix: The Last Sentence Could Be Missing When a Recording Is Stopped

When stop was sent, a sentence that had not finished being recognized was not included in the transcript. This is fixed: the server now waits for the recognizer to finish the last sentence (usually about 1 second, at most about 3 seconds) before finalizing; if it does not arrive in time, the last interim result shown is used as that sentence. pause in two-way translation and stop_speaking in manual mode wait in the same way.

Fix: task_complete Could Be Missed When Finalizing Took Longer

When finalizing after stop (title, summary, upload) took longer, the connection could be closed before task_complete was sent. This is fixed. The title and the summary are now generated at the same time, which shortens the wait between stop and task_complete.

Fix: Recordings That Had Just Ended Could Be Incomplete During Maintenance

During service maintenance, for a recording that had just ended or just disconnected, the transcript, the full audio file and the completion notification sometimes could not be processed in time; the recording stayed in processing and was later marked as failed. This is fixed: the service now finishes this work before it shuts down.

Fix: Stream Errors No Longer Include Internal System Messages

When an unexpected error occurs in Import Progress SSE or TTS SSE Streaming, details.original_error is now always Service error, consistent with the other streaming endpoints; the error code is unchanged.

Fix: Error Messages in TTS SSE Streaming

In TTS SSE Streaming, the message of tts_error events previously carried an internal identifier; it is now the standard English message (for example, TTS synthesis failed). The error code in error is unchanged. The documented error code for a sentence with no translation is corrected to tts_translation_not_found, the value actually sent; behavior is unchanged.

Behavior Change: Connections When the Service Is Shutting Down

  • After service_shutdown is sent, a connection that is not recording is closed about 2 seconds later.
  • A connection that is recording can finish the recording, still receives task_complete after stop, and is closed after that.

Behavior Change: Stop, Pause and Concurrent Recording Slots

  • status: "ended" is always sent before task_complete.
  • The concurrent recording slot is now released when task_complete is sent: you can start the next recording as soon as you receive it.
  • Responses to stop, to pause in two-way translation, and to stop_speaking in manual mode are delayed by about 1 second (at most about 3 seconds).

Client Recommendations: Detecting Completion

  • task_complete waits until the title and summary have been generated, which usually takes a few seconds to tens of seconds and up to about 8 minutes. To detect completion, we recommend also supporting the Webhook or a REST query rather than relying on this event alone.

Behavior Change: Retranslate Summary

  • When the translation is incomplete, done carries truncated: true.
  • A request has a processing limit of about 230 seconds, and an error is sent if translation pauses for more than 60 seconds. A timeout now reports details.original_error as Translation timed out instead of Service error.
  • For a long summary, the stream may sometimes deliver a larger block at once, with a pause of a few seconds between parts.

Client Recommendations: Summary Translation

  • When you receive truncated: true, the translation is incomplete; do not save it as a complete result.
  • If you want to render the translation character by character, smooth it on the client side.

Documentation Updates

  • The reason description of segment_discarded now notes an exception: after a session resume, channel_status uses reconnect, while segment_discarded uses resumed.
  • start_speaking corrected: calling it again while already speaking ends the previous sentence and starts a new one, instead of returning an error.
  • segment_discarded now notes that two-way translation manual mode does not receive this event either.
  • task_complete now documents its order relative to status: "ended", that it waits for the title and summary, when the concurrency slot is released, and how to detect completion.
  • pause, stop and stop_speaking now document the wait for the last sentence.
  • Pricing now clarifies billing during pauses and disconnects: billing continues while paused, and the time spent waiting to resume after a disconnect is not billed.
  • Pricing now clarifies how multi-channel channels are counted: a channel added mid-minute and disabled before the next deduction is counted once, and a channel already deducted at the start of the minute is not billed again after it is disabled.

V1.16.5

2026-09-21

New: Summary Events Carry summary_language

The following events now include summary_language (additive):

EventValue
init_summary of HistoryThe language of the saved summary; null when there is no summary
done of Regenerate SummaryThe language used this time; the first transcription language when language is not sent
done of Ad-hoc SummaryThe language used this time; zh-TW when language is not sent

To determine the language of the summary itself, use summary_language of init_summary (when there is no summary, it may differ from the field of the same name in init_metadata).

Fix: Summary Language of Imported Tasks

For imported tasks with a summary, summary_language in the task list and in init_metadata was null. This is fixed, and existing tasks have been corrected.

Documentation Updates

  • summary_language of init_metadata: when no summary language is specified for a recording, the value is usually the first transcription language (in two-way mode with active_language specified, it is that language), not null. Only tasks without a summary may have null.
  • Retranslated summaries (GET /api/v1/sse/retranslate/summary/{taskId}) are not saved. To keep a summary in another language, use the POST endpoint of Regenerate Summary (billed).

V1.16.4

2026-09-17

Import Quota Preflight Now Reports "Daily Usage Exhausted"

POST /api/v1/imports/check-quota previously only looked at credit and plan entitlement, so it still returned allowed: true when the day's usage was exhausted and only the upload was rejected. It now returns allowed: false with reason: "plan_daily_limit_reached", matching the error code the upload returns so the two map directly.

Note: When allowed is false, topping up does not solve every case — exhausted daily usage resets tomorrow, and a plan without import requires an upgrade. Tailor the message to reason; see the Import API reference.

Insufficient-Credit Errors Add budget_scope

The details of auth_quota_exceeded and stt_quota_exceeded now carry budget_scope, stating whose credit the accompanying remaining_budget refers to. Additive; existing fields are unchanged.

Important: remaining_budget is the credit available to the API key that made the request, not the end user's balance. If you call this service on behalf of other users, do not show this number directly to them. See the Error Code Reference for what each value means.


V1.16.3

2026-09-17

Billing Fix: Channels That Are Not Working Are No Longer Counted

In multi-channel recording, if a channel's audio cannot be received at the moment a minute is settled — and the channel therefore produces no transcript — that minute no longer counts this channel; previously it was still billed. See Pricing for how the count is determined, including the first minute and channels that send no audio.

Behavior Change: A Broadcast Token May Only Be Used by the Account That Owns It

Starting a broadcast with a broadcast_token that belongs to another account returns broadcast_token_invalid. Different API keys under the same account are unaffected.

New Error Code: task_already_processing (409)

While the same task is still being processed, POST /api/v1/tasks/{taskId}/retry returns 409 task_already_processing and the task status is unchanged; send the same request again later. Previously this situation returned 200 without the task being reprocessed.

Fix: Summary Language After Regenerating a Summary

When regenerating a summary without specifying language, the summary language recorded in the task data was not updated and did not match the actual summary content. The two are now consistent.


V1.16.2

2026-09-16

Documentation update: how to obtain the broadcast task_id

For a broadcast, the authoritative task_id is the one returned by the broadcast_recording_ready event; the task_id in session_started is an initial value for the connection stage, and querying with it returns no matching record. You receive this event either way a broadcast goes live: after the standby phase, or by starting directly with broadcast_phase: "live" (the default). Behavior has not changed; this version only completes the description.


V1.16.1

2026-09-16

Documentation update: how billed minutes are counted at the boundary

For real-time recording, each minute boundary has a 1-second grace period: a 60.5-second recording counts as 1 minute, and a 61.5-second recording counts as 2 minutes. Billing behavior has not changed; this version only completes the description.


V1.16.0

2026-09-15

Behavior change: real-time recording is now deducted at the start of each minute

Real-time recording (broadcasts excluded) used to be deducted after each minute was used, with any partial minute counted as a full minute when the recording ended. Starting with this version, credits are deducted at the start of each minute:

  • The first minute is deducted when the recording starts, and session_started arrives once that deduction completes
  • Each following minute is deducted as it begins
  • Nothing extra is deducted when the recording stops

The number of billed minutes per recording still follows "any partial minute counts as a full minute". Previously, some recordings were billed one minute less than this rule; this version always applies the rule, so a recording of the same length may be billed one minute more than before.

When features are turned on or off, or channels or translation languages are added or removed during a recording, the rate changes from the next minute. For how channels are counted in multi-channel mode, see the Pricing Guide.

Behavior change: a recording cannot start or continue when the available credits do not cover one minute

  • When starting a recording: the recording does not start, and you receive auth_quota_exceeded with details.remaining_budget set to the available credit as of the most recent settlement. The connection stays open; after topping up, you can send start again right away
  • During a recording: the recording ends before the next minute begins, and you receive stt_quota_exceeded; everything recorded so far is still saved
  • The minute that cannot be covered is not deducted, and the remaining credits are not deducted either

If the available credits cannot be confirmed at the moment, start returns auth_service_error and the recording does not start; please try again later.

New: summary_insufficient_credit error code

In the following cases no summary is generated, and summary_error is sent with the error code summary_insufficient_credit:

  • The recording ended because the available credits ran out
  • At a normal end, the available credits cannot cover the summary fee

The transcript and audio are still saved; after topping up, you can get a summary by regenerating it.

Behavior change: unlimited-plan usage limits are checked before each minute begins

  • The single-recording limit, the usage-hour thresholds, and the daily hard limit are all checked before each minute begins; once one is reached, the next minute does not start and is not counted toward usage
  • If today's usage has already reached the limit when a recording starts, daily_limit_reached is returned right away and the recording does not start
  • Inside a restriction window, each stretch of continuous recording now lasts exactly the interval set by the plan (previously one minute longer)
  • When the same API Key runs several recordings at once, after the periodic-stop threshold is reached, each recording stops before its own next minute begins (previously only one of them was stopped each time)
  • The used minutes in GET /api/v1/me/plan include a minute as soon as it begins

Behavior change: when credit.exhausted is sent

credit.exhausted is also sent when a real-time recording cannot start or continue because the available credits do not cover one minute and the account balance cannot cover that minute; in this case balance may be greater than 0.0. When only an API Key's dedicated allocation is insufficient while the account balance can still cover the minute, it is not sent, as before.

Fix: problems when recording several times on the same connection

After one recording ended and the next one started on the same connection, the following could happen:

  • The next recording's summary fee also counted the transcript characters of earlier recordings, and a summary fee could be charged even when that recording produced no summary
  • If an earlier recording had text-to-speech turned on, the next recording was billed at the rate including text-to-speech even when it was off
  • If an earlier recording used custom summary settings, the next recording could fail to be saved when it ended
  • Multi-channel recording: if an earlier recording's audio was not saved completely, the next recording could also fail to be saved
  • Floating subtitles did not start after starting again
  • An insufficient-credit or plan-limit notice for an earlier recording stopped the next recording

All of these are fixed in this version.

Fix: starting again right after a recording was ended could leave the previous recording unsaved

After a recording ended because of insufficient credits, a plan restriction, a usage limit, or a long period without sound, sending start right away on the same connection could leave the previous recording unsaved, and the new recording could end up without a transcript.

Starting with this version, the new recording starts only after the previous one finishes processing (after status: "ended"); depending on the summary length, session_started may be delayed by a few seconds to a few tens of seconds. Starting a recording on a new connection is not affected.

Behavior changes

  • Reconnecting with a resume token after sending stop and before status: "ended" arrives (including while a recording that was ended is still being processed) returns resume_unavailable; a recording that has already ended is not resumed
  • Sending broadcast_go_live while the previous recording is still being processed returns session_not_started
  • details.remaining_budget in auth_quota_exceeded and stt_quota_exceeded is now the available credit as of the most recent settlement
  • When a recording ends because of insufficient credits, stt_quota_exceeded is sent only once
  • After a recording has ended, no more insufficient-credit or plan-limit errors for that recording are sent
  • When a recording is ended while its connection is interrupted, the API Key's concurrent-recording slot is released right away (previously it was released only after about 3 minutes, during which concurrency_limit_reached could be returned)

Client recommendations

  • Starting a recording may return auth_quota_exceeded because of insufficient credits: the connection stays open, so just send start again after topping up
  • Listen for summary_insufficient_credit in summary_error and tell the user that no summary was generated because of insufficient credits
  • When starting again on the same connection after a recording was ended, session_started may be delayed; do not treat this as a timeout
  • While the previous recording is still being processed, pong may also be delayed; allow enough waiting time and do not treat the connection as dropped just because a pong has not arrived yet
  • On resume_unavailable, if the user has already ended the recording, do not automatically start a new one

Documentation updates

  • The WebSocket start error table now lists auth_quota_exceeded, daily_limit_reached, and auth_service_error; the HTTP status of broadcast_token_invalid is corrected to 401 to match the Error Code Reference
  • The broadcast_go_live error table now lists session_not_started
  • Session resume example: when the user has ended the recording, a failed resume no longer starts a new recording automatically
  • The heartbeat description now notes that pong may be delayed while the previous recording is being processed
  • The description of stt_quota_exceeded for audio imports is corrected to insufficient available credits for the import
  • The event examples in the summary customization guide now use the actual message format (type is voice-translation, and events are distinguished by data.action)

V1.15.10

2026-09-12

Fix: retranslation reported success with no translation when content could not be translated

When retranslation (full text, single sentence, or summary) encountered content that could not be translated, it reported success with a blank translation; full-text retranslation was also billed and overwrote the translation that language already had.

This version reports failure instead: the previous translation is kept, and failed sentences are neither counted toward the updated total nor billed.

Fix: results could be lost or billed incorrectly when several operations ran at once

When several operations ran against the same recording at once (for example, retranslating into different languages at the same time, or retranslating while renaming a speaker), a result could be billed without being kept, or billing fields could appear even though nothing was consumed.

This version reports failure in these cases, and the request is not billed.

Behavior changes

  • Retranslation may return llm_content_filtered (severity warning); other translation failures return sse_translation_failed, and summary retranslation returns sse_summary_translation_failed
  • When the transcript cannot be saved, retranslation, saving a summary, and speaker operations all return storage_upload_failed and stop; done will not follow
  • When several changes are made to the same transcript at once, the later one receives transcript_revision_conflict (409)
  • done carries characters_billed / charged / billed only when the request was actually billed
  • When a speaker is specified by a display name that maps to more than one speaker, speaker_name_duplicate is returned
  • Translations that contain only whitespace no longer appear in exports or speech synthesis

Client recommendations

  • Retrying llm_content_filtered will not help; revise the original text instead. If you filter events by severity, do not filter out warning
  • A transcript_revision_conflict means another change is in progress; simply retry
  • Use billed to decide whether a request was billed; integrators that bill their own end users in particular should not infer it from other fields

V1.15.9

2026-09-10

Fix: After a connection dropped and recovered, a speaker's name could be applied to someone else

With speaker identification (multi-speaker mode), if the connection dropped and recovered mid-recording, a name assigned earlier could be applied to a different speaker, and a merge configured earlier could combine two different people into one. Neither produced any notice.

This version fixes it: speakers appearing for the first time after a recovery receive new IDs, carrying over neither the earlier names nor the earlier merges.

Behavior change: speaker IDs are no longer guaranteed to start at 1 or run consecutively

Speaker IDs are unique within a recording, but are not guaranteed to start at 1 or to be consecutive. After a connection recovers, and when a broadcast moves from standby to live, speakers appearing afterwards receive IDs that have not been used before.

rename_speaker and merge_speakers apply only to the speakers identified so far. After a recovery, merge again once you confirm it is the same person; when the target has no name yet, merging carries the source's name across.

When a broadcast goes live, speaker names set during standby are not kept.

Single-speaker, conversation, and multi-channel modes are unaffected.

Client recommendations

  • If your application assumes IDs start at 1, run consecutively, or stay the same for one person throughout, use the speaker_id you receive instead
  • To continue with the same speaker after a recovery, use merge_speakers; calling rename_speaker with the earlier name returns speaker_name_duplicate
  • On the broadcast host side, clear your accumulated speaker list when the phase changes to live

V1.15.8

2026-09-09

Fix: In conversation mode, a sentence in progress disappeared when switching modes

In conversation mode, switching between automatic and manual mode — or pressing the talk button — while a sentence was still being spoken made that sentence reappear under a new segment number. The original number then had no further events, and the sentence did not appear in the transcript; the sentence spoken next could also end up without its own translation.

This version fixes it: a sentence in progress now stays on its original number and ends normally; only the sentence spoken afterwards gets a new number.

New: segment_discarded event

When the speaking speed is changed, a channel's language is changed, a broadcast goes live, a session is resumed, or recording resumes after a pause, a sentence that happens to be in progress cannot be kept. Previously none of these produced a notification, so that sentence never received an ending signal.

From this version, segment_discarded is sent to state that the segment number will receive no further events, with reason identifying the operation; multi-channel recordings also carry channel_id. Floating-subtitle viewers receive it as well.

Behavior change

Re-sending the conversation mode that is already in effect no longer affects a sentence in progress; the mode change notification is still returned.

Client recommendations

  • On segment_discarded, clear the matching segment from any unfinished state
  • Automatic conversation mode does not receive this event; there, a sentence in progress ends normally and its content is kept in the transcript
  • When set_speaking_speed_failed is returned, that operation may still have sent segment_discarded
  • See segment_discarded for fields and the full reason value set

V1.15.7

2026-09-08

Bug fix: Re-translation sometimes returned something other than a translation

When re-translating a completed transcript, a shorter sentence could come back as explanatory text about that sentence, and be stored as its translation.

Fixed in this release. Re-translating an affected segment once returns the correct translation; per-segment re-translation is not billed.


V1.15.6

2026-09-06

Fix: The 202 response when uploading audio returned null for progress

progress in the successful response from POST /api/v1/imports should have been 0, but was actually null. Querying the same import through GET /api/v1/imports/{importId} returned 0, so the two disagreed.

Impact: a client that validates the response structure strictly reports a successful upload as a failure, so it never captures import_id and never receives the completion notification. Each retry creates another import that runs to completion and consumes credits again.

From this version, progress is always 0 in the 202 response, matching both the documentation and the query endpoint.

Fix: The 201 response when creating a broadcast returned null for three counter fields

peak_viewers, total_viewers, and duration_ms in the successful response from POST /api/v1/broadcasts should have been 0, but were actually null. Querying the same broadcast through GET /api/v1/broadcasts or GET /api/v1/broadcasts/{broadcastId} returned 0.

The symptom is the same as the previous item: a client that validates the response structure strictly reports a successful creation as a failure. Retrying leaves several usable broadcasts behind, each holding a token.

From this version these three fields are always 0 in the 201 response. Creating a broadcast does not consume credits, and existing broadcasts and their statistics are unaffected.

Documentation Update: task_id is present as soon as the upload succeeds

task_id was previously described as being populated only after processing completes, and both the POST and GET response examples showed it as null. In fact task_id already exists at the moment the upload succeeds (202); there is no need to wait for processing.

Clients can navigate to the task as soon as they receive the 202 response, rather than polling to obtain task_id. The examples and field descriptions have been corrected.

Client Recommendations

No integration changes are required. If you relaxed your validation because progress or the three broadcast counter fields were null, you can require an integer again; if you waited for processing to finish before using task_id, you can now use it from the successful upload onward.


V1.15.5

2026-09-04

Fix: Speaking a language outside the transcription list could return the untranslated original

When multiple transcription languages were specified but the speaker used a language not on the list, one of the translation languages could show the original text instead of a translation.

For example, with transcription languages zh-TW, id-ID and translation languages en-US, zh-TW, id-ID: when English was spoken, the Indonesian translation showed the English original. This has been fixed.

Behavior change: with multiple transcription languages, every translation is produced by translating

When multiple transcription languages are specified, every translation language is now produced by actually translating, even when the source language matches that translation language. Such translations may differ slightly in wording from the original (meaning preserved).

Sessions with a single transcription language are unaffected: when the source language matches a translation language, that translation still reuses the original text.

Billing is unchanged: translation is charged by the number of translation languages, regardless of whether a given translation was actually produced by translating.

Client recommendations

No integration changes are required. Include every language that will actually be spoken in the transcription language list.


V1.15.4

2026-09-04

Fix: Some target languages returned the untranslated original text

When multiple transcription languages were specified, one of the target languages could show the original text instead of a translation.

For example, with transcription languages en-US, zh-TW, id-ID and the same three translation languages: when Indonesian was spoken, the English translation showed the original Indonesian text, while the Chinese translation was correct. This has been fixed.

Fix: Language and speaker tagging in conversation mode

With certain language pairs (for example English and Indonesian), a sentence could be tagged as the other language, and the speaker label and transcript language tag were wrong as a result. This has been fixed.

Fix: Translations when an import declares multiple transcription languages

When an import declared multiple transcription languages and the translation targets included the first of those languages, that language could return the untranslated original text. This has been fixed.

Behavior change

With certain language combinations (for example English and Indonesian specified together), translations whose source and target language are the same may now differ slightly in wording from the original, while the meaning is preserved. Previously such translations reused the original text verbatim. Sessions with a single transcription language are unaffected.

Client recommendations

No integration changes are required. If your application checks whether the original and the translated text are identical in order to decide whether a sentence was translated, please note the change above.


V1.15.3

2026-09-03

Fix: Fuzzy correction could corrupt text that was already correct

When an incorrect variant in fuzzy_correction was written in Latin script, it was previously applied to part of a longer word, corrupting text that was already correct.

For example, with the term Remote View and Emote View listed as an incorrect variant (with case_insensitive enabled): in a transcript that already read Remote View, the emote View inside it was treated as a match and replaced, producing RRemote View — an extra character at the start.

From this version, incorrect variants written in Latin script are only applied to whole words. Matching for Chinese, Japanese, and Korean is unchanged.

Fix: Punctuation or spacing immediately after a correction could disappear

Punctuation or a space immediately following a variant was also consumed, with no indication:

Transcript contentPrevious result
open Remote Vue, then quitopen Remote View then quit (comma lost)
open Remote Vue.open Remote View (period lost)
open Remote Vue nowopen Remote Viewnow (two words run together)

Chinese punctuation was affected in the same way. Both are now preserved.

Behavior Changes

Common-word protection in homophone correction is no longer bypassed by adjacent spaces. Previously, when a misrecognized word happened to have spaces around it, common-word protection did not take effect and the word was still corrected; without those spaces it was not. The result for the same word depended on whether spaces happened to surround it.

Both cases are now consistent and common-word protection always applies. A small number of common words that used to be corrected are therefore no longer corrected — if you do need them corrected, list them explicitly in fuzzy_correction.

Leading and trailing spaces in glossary entries are now ignored. Previously an entry with leading or trailing spaces produced inconsistent results, and in some cases spacing in the transcript was removed. This has been fixed, which also means two entries that differ only by surrounding spaces are treated as the same entry.

Documentation Updates

  • The terminology guide now states that matching in Latin script operates on whole words
  • Corrected an example in the terminology guide that used an incorrect variant with a trailing space (such an entry has no effect)

Client Recommendations

No changes are required to existing integrations. Glossary configuration format, limits, error codes, and conflict codes are all unchanged. If your glossary registers a shorter Latin-script form and relies on it covering a longer word (such as wafer covering wafers), register both forms separately.


V1.15.2

2026-09-03

Behavior Changes

Sending stop before the recording has started, or after it has already ended, now returns a session_not_started error. Previously these two cases produced no response at all, leaving the client to wait for a timeout before discovering that the action had no effect.

The most common case is sending stop twice: the second one returns this error rather than another success response. task_complete is still sent only once, after the first successful stop. If your client already treats a duplicate stop as harmless, simply ignore this error.

Documentation Updates

  • Added the error code section for stop
  • Filled in error codes that were missing from the tables for the multi-channel add_channel, remove_channel, and set_channel_language actions

V1.15.1

2026-09-03

Fixes

  • When config returns config_empty, the error message now states explicitly that an empty object {} is not treated as a provided setting, and shows the correct way to clear a glossary
  • Documentation now notes that a rejected config returns type: "error" rather than config_updated. Clients must listen for error as well, or the request appears to go unanswered

Error codes and response structures are unchanged.


V1.15.0

2026-09-03

Fix: Summaries of long meetings were cut off

When generating a summary for a longer meeting, the summary could stop partway through with an unfinished sentence — and no error was reported. The stream ended normally and the done event was still sent, so there was no way for a client to tell that the content was incomplete. The longer the recording, the more likely this was to happen.

From this release, meetings of an hour or more produce a complete structured summary.

This is distinct from the existing summary_fallback_level / summary_dropped_segments behavior. Those indicate that part of the input was omitted (the sentences are complete, but a section of the meeting is missing); what this release fixes is the summary itself stopping halfway.

New

Incomplete summaries are now flagged. The done event of both the ad-hoc summary endpoint (POST /api/v1/sse/summary) and Regenerate Summary (GET / POST /api/v1/sse/regenerate/summary/{taskId}) gains an optional truncated field: it is true when the summary could not be produced in full, and the field is absent entirely when the summary is complete.

This is a purely additive, backward-compatible field — existing integrations can ignore it. Even if an exceptionally long summary still cannot be produced in full, the client can now see that.

Behavior Changes

The character limit for summary input has been raised from 100,000 to 200,000. Affected endpoints:

  • content on POST /api/v1/sse/summary
  • The transcript for GET / POST /api/v1/sse/regenerate/summary/{taskId}
  • The summary generated automatically when a recording ends

Previously, a long recording whose transcript exceeded 100,000 characters received summary_text_too_long and got no summary at all. The new limit is aligned with the longest audio duration file import accepts (10 hours), so "the import succeeds but there is no summary" no longer happens.

Glossary coverage in summaries has been widened. Previously, when the transcript was long, only part of your glossary was applied to the summary, with no indication that this had happened. This release substantially relaxes that. Real-time translation behavior is unchanged.

Free regeneration for ad-hoc summaries is now capped. Re-sending the same idempotency_key with identical content previously regenerated the summary for free with no limit on how many times; from this release, exceeding the cap returns 429 too_many_requests.

Free retries exist for the "already charged but delivery failed" case, where one or two attempts is normally enough. To generate again, use a new idempotency_key (charged as a new request). The initial paid request does not consume the retry budget, and neither does a failure that produced no content. The budget is counted over a rolling 24-hour window.

Client Recommendations

  • Important: Review your client-side timeout settings. Summaries are now generated in full, so long meetings take longer than before. We recommend allowing at least five minutes for summary-related stream connections. A shorter timeout will disconnect before the summary finishes
  • Long recordings that previously failed with summary_text_too_long can now be regenerated
  • The done event gains an optional truncated field (see "New" above). Existing integrations can ignore it, but we recommend checking it so you know whether the summary you received is complete
  • Ad-hoc summary requests that reuse a fixed idempotency_key may now receive a 429. The normal "one new key per request" usage is unaffected

Documentation Corrections

  • The "Character and Length Limits" table in guides/summary-customization.md did not list the limit for summary input text. It has been added

V1.14.2

2026-09-02

Fix: Viewers could not connect to the subtitle stream of password-protected broadcast channels

After entering the correct password and obtaining a viewer_access_token, the viewer's connection to the subtitle stream was still rejected (HTTP 401, broadcast_password_required). With this release, the connection succeeds once the password is verified.

  • Switching networks between password verification and stream connection (for example mobile data vs. Wi-Fi, or a corporate multi-line egress) no longer causes the connection to be rejected
  • The token remains valid for 24 hours; the response fields of the verify endpoint are unchanged

Fix: tts_languages in viewer channel info was not returned correctly

GET /api/v1/viewer/broadcasts/{token} could return an empty tts_languages array in some environments, leading clients to assume the channel offered no voice playback. It now correctly returns the voice languages enabled by the host. The field format is unchanged.

The definition of tts_languages is also clarified: it lists the languages for which the host has enabled voice playback, and is an empty array when the channel is not live. The previous wording, "otherwise from the default settings", did not correspond to any actual behavior.

Behavior change: Error responses from the viewer subtitle stream are readable cross-origin

When the viewer subtitle stream (GET /broadcast/{token}/text) rejects a connection (401 / 404 / 503, etc.), the response now carries the same cross-origin headers as a successful response. Previously, cross-origin clients received a cross-origin error in these cases and could not read the response at all.

The shape of the failure changes as well: a cross-origin rejection was previously classified as a network error, which per specification causes EventSource to keep reconnecting. It now terminates the connection outright (readyState becomes CLOSED) with no further retries. If your client wraps EventSource in its own reconnect logic, confirm that its behavior in this case is what you expect.

Client recommendation: A browser-native EventSource never exposes the response body to the page, so it still surfaces only onerror. To show viewers a specific reason (such as broadcast_password_required or broadcast_not_ready), read the stream with fetch, or call GET /api/v1/viewer/broadcasts/{token} before connecting to check the channel status and whether a password is required.

Documentation

  • Viewer API: removed the network-address binding note from viewer_access_token; corrected the tts_languages field description
  • Error Codes: corrected the HTTP status of broadcast_password_required from 422 to 401 (the actual behavior has always been 401)

References: Viewer API, Viewer Subtitle Stream

V1.14.1

2026-09-02

Documentation: Request Size Limit for Glossary Validation

POST /api/v1/glossary/validate did not previously state the actual request size limit, leaving integrators unable to assess ahead of time whether their glossary would fit. This release documents:

  • The default limit: 2 MB (2,097,152 bytes). Reading details.max_bytes from the response is still the recommended approach — do not hard-code a value
  • How to split a glossary that exceeds the limit. terminology and fuzzy_correction must be sent in the same request; translation_dict may be sent on its own

No API behavior changed. This release is documentation only.

Note: If you already split requests yourself, review this. When terminology and fuzzy_correction are sent as two separate requests, some conflicts are not detected, and the response gives no indication of it.

Reference: Glossary Validation API

V1.14.0

2026-09-01

Added: Glossary Validation API

Adds POST /api/v1/glossary/validate, so a glossary management UI can check a glossary for format problems and internal conflicts before saving.

Most glossary problems produce no error message at all — they simply yield unexpected results, silently, during recording or translation. The worst of them is "an incorrect variant that is also the correct term of another rule": a user saying that word normally has it rewritten into something else. This endpoint surfaces situations like that at the moment the user presses Save.

Characteristics

  • Free: no points are deducted, no task or recording is created, and no settings are written
  • All three blocks — terminology, fuzzy correction, translation dictionary — are optional; whatever you send is what gets validated
  • Every problem comes with the index of the entry (language code plus index), so a management UI can flag the item directly
  • The response carries no human-readable text; integrators compose their own sentences from the conflict codes, entirely in the language and phrasing they choose
  • Rate limited to 120 requests per minute per API Key, as a separate quota counted independently of the other endpoints
  • Can be called directly from a browser (cross-origin requests are supported). Note: This means the API key is present in the browser, and the same key can also create recordings, import audio, and consume credits — acceptable for an internal admin tool, but for a public page route the call through your own backend instead

Note: The host is the realtime service domain (the same domain as wss://, with the scheme swapped for https://), which may differ from the host of the other REST endpoints.

The eight detectable conflicts

CodeSeverityMeaning
variant_shadows_termerrorAn incorrect variant is also another rule's correct term or a registered term; normal text gets broken
dict_duplicate_sourceerrorThe same source is registered more than once under one target language; only one takes effect
variant_ambiguouswarningThe same incorrect variant maps to different correct terms
case_flag_conflictwarningThe case_insensitive settings for one incorrect variant are inconsistent
variant_equals_termwarningAn incorrect variant is identical to the correct term of its own rule and has no effect
duplicate_termwarningThe same term is registered more than once within the same language family
boost_out_of_rangewarningA term's boost value is outside the valid range and is adjusted automatically
homophone_conflictwarningTwo correct terms or registered terms share a pronunciation

The homophone check can be turned off

Of the eight conflicts, the homophone check is the slowest; the other seven return almost immediately. For live feedback while the user edits, send check_homophones: false; run the full check when they actually save.

Important: The homophone check applies to Chinese only, and may not always complete. homophones_checked in the response is what tells these two situations apart: when it is false, the check did not complete, which does not mean there are no conflicts. Have your save gate check this field as well.

Added: error codes

Error codeHTTPDescription
config_too_many_languages400There are more language codes than allowed (details.field names the field at fault)
config_payload_too_large413The request body exceeds the size limit

Behavior clarification: config is all-or-nothing

Earlier documentation described config as "processed in order, aborting on failure, with the blocks ahead of the failure already applied". That is not the actual behavior: if any one of the three glossary blocks fails, none of the three is applied — the settings stay exactly as they were.

For integrators this is a change for the better: when an error comes back you can be certain that nothing changed, with no half-applied state to worry about. The affected documentation has been corrected throughout.

Behavior clarification: variant collisions do not guarantee a winner

Earlier documentation said that when one incorrect variant maps to several rules, a fixed order decided the winner. Testing did not match that description, so it is now stated plainly: which rule actually takes effect is not guaranteed — do not rely on any ordering, registration order included. This is consistent with the new variant_ambiguous conflict code, which deliberately does not report which rule wins.

Behavior change: the TTS voice catalog now lists only usable languages

GET /api/v1/tts/voices previously listed all 154 locales, but 13 of them cannot actually be used — a TTS target language must be one of the translation output languages, and translation output languages are limited to the transcription language list, which those 13 are not in. Customers could look up voices for them and then have translation_languages reject the same code.

Starting with this version the voice catalog lists only the languages that are actually usable, and the published language and voice counts have been changed to the usable figures:

Previously (including unusable)This version (actually usable)
TTS languages154141
TTS voices325302

Possible impact:

  • Querying GET /api/v1/tts/voices for one of those 13 locales now returns an empty voice list (voices were listed before, but they could not be used for synthesis anyway)
  • Querying GET /api/v1/tts/voices/{voiceName}/sample for one of their voices now returns 404 tts_voice_not_found, consistent with the catalog not listing them

All other languages are unaffected.

The 13 affected locales: bn-BD, ta-LK, ta-MY, ta-SG, ur-PK, su-ID, sr-Latn-RS, iu-Cans-CA, iu-Latn-CA, and four zh-CN dialects (zh-CN-henan, zh-CN-guangxi, zh-CN-liaoning, zh-CN-shaanxi).

Behavior change: the audio import duration limit is now actually enforced

The "File limits" table has always listed a maximum duration of 10 hours and a minimum of 1 second, but until now only the optional import pre-check endpoint verified them — a file uploaded directly was subject to no duration limit at all. In practice a 500 MB low-bitrate file can exceed 17 hours and would be processed in full and billed.

Starting with this version, a file outside the range ends the import with a failed event whose error_code is the newly added import_duration_out_of_range.

Possible impact: if you currently import files longer than 10 hours (or shorter than 1 second), those imports will start failing. Use audio within the allowed length, or split it and import in parts.

An over-long file uploads successfully first and then ends with failed. To find out before uploading, call the import pre-check endpoint first.

Behavior change: speaking_rate is now clamped to its documented range

The TTS speaking_rate has always been documented with a valid range of 0.5 to 2.0, but until now values outside the range took effect as sent. Starting with this version, values outside the range are adjusted to the nearest bound (below 0.5 becomes 0.5; above 2.0 becomes 2.0), so the actual behavior matches the documented range.

Possible impact: if you currently send a speaking_rate above 2.0 (or between 0 and 0.5), the speaking rate becomes the boundary value. Values within the range are unaffected.

Behavior clarification: an out-of-range boost behaves differently on the two paths

A term's boost has a valid range of 0.5 to 5.0. When a value falls outside it, live recording and broadcast adjust it into range automatically with no notice, while file import rejects it outright (HTTP 422). Earlier documentation did not record that the same glossary produces different outcomes on the two paths.

Documentation corrections

This release also corrects the following discrepancies with actual behavior:

LocationCorrection
Ad-hoc summaryAuthentication corrected to Header X-API-Key only (query string was documented in error); removed auth_insufficient_credit, which never occurs
Summary regeneration (SSE)Removed five parameter-validation error codes that are never actually sent, replaced by a note that a parameter validation failure carries only message and no error_code; content filtering actually returns sse_summary_regeneration_failed, not llm_content_filtered
Retranslation (SSE)The error event's context corrected to sse; added the request_id and timestamp fields
Single-sentence retranslation (SSE)Added auth_insufficient_credit (402)
config actionAdded config_invalid_entry to the error code table (the most commonly triggered code, previously unlisted)
WebSocket connectionAdded "Maximum Size of a Single Message" — exceeding it closes the connection outright with no error message; send a large glossary across several config messages
WebSocket eventsAdded the speakers_auto_merged event (previously documented only on the viewer side)
Broadcast API examplesCorrected a voice name that is not on the supported list
TTS audio streamingFixed a layout problem caused by an error-code row placed inside the wrong table
Documentation home pageCorrected the endpoint counts for speaker editing and summary templates
start actionDocumented that glossaries sent inside start are ignored (previously noted only in the error code reference)
Error code referenceauth_account_expired is marked as not currently returned by any endpoint; tts_invalid_voice is marked as returned only by the realtime voice channel

Documentation updates

  • The "Important Notes" section of the Terminology Guide now covers six previously undocumented situations (an incorrect variant shadowing a correct term, duplicate terms under one language, boost values adjusted automatically, an incorrect variant equal to its own correct term, a duplicate source word in the dictionary, terms that share a pronunciation), each marked with its matching conflict code
  • The same guide gains a "Validating a Glossary Before Saving" section
  • The Error Code Reference corrects the HTTP status description for invalid_json (endpoints on the realtime service domain return 400, the rest return 422) and adds how invalid_action is used on REST endpoints
  • The Terminology Guide adds the language-family matching rule, the behavior of an empty language code, and the valid range of boost

Behavior notes

  • Only aggregate results are returned when the count is over the limit: when the glossary's total entry count exceeds the limit, the response contains the count problem alone — individual content problems are no longer listed and no conflict detection is run. Bring the count back within the limit and validate again.
  • The number of reported items is capped: conflicts and format problems each have a reporting cap, and truncated is true when it is exceeded. Fix the items listed and validate again to see the rest.

Client recommendations

  • Have the glossary management UI call this endpoint once before saving; when valid is false, block the save and show the offending entries
  • The recommended condition for the save gate is valid === true && homophones_checked === true
  • On auth_service_error (HTTP 500), retry later; do not swap the API Key — this has nothing to do with the key itself
  • Passing this endpoint's validation does not guarantee that audio import will accept the same glossary — import applies stricter rules

Reference


V1.13.1

2026-09-01

Added: completion events now report consumption

The done event of the following endpoints carries three new fields, so integrators no longer need to derive usage themselves:

EndpointNotes
GET /api/v1/sse/retranslate/{taskId}Full-text retranslation
GET/POST /api/v1/sse/regenerate/summary/{taskId}Summary regeneration (both preview and save are billed)
POST /api/v1/sse/summaryAd-hoc summary (adds charged alongside the existing characters_billed)
{
  "characters_billed": 12700,
  "charged": "1.3",
  "billed": true
}
  • characters_billed: character count used as the billing basis
  • charged: points consumed by this operation, calculated from the rate. This value reflects usage — usage already covered by an unlimited plan is still reported here
  • billed: whether the request incurred consumption

Always use billed to determine whether a request was billed (billed only when billed is true). When the fields appear differs by endpoint:

  • Full-text retranslation and summary regeneration: when nothing was consumed (for example, when generation fails), all three fields are absent
  • Ad-hoc summary: characters_billed and charged are always present, while billed is false for free retries (same idempotency key with identical request content) and empty generation results — in those cases the request was not billed

Do not use the presence of charged as the criterion: on a free ad-hoc retry charged still has a value, and billing on that basis would overcharge.

Endpoints that are not billed do not carry these fields: summary retranslation (/retranslate/summary/{taskId}) and single-sentence retranslation (/recordings/{taskId}/entries/{sid}/retranslate) are not billed, and their done events are unchanged.

Client Recommendations

These are purely additive fields; existing integrations are unaffected and require no changes.

Integrators that call this API on behalf of end users and bill them separately can use charged directly rather than deriving it from character counts and rates, which avoids the two sides computing from different bases. If you adopt it, use billed === true as the single criterion for whether a request was billed — that one rule covers every endpoint listed above, with no per-endpoint exceptions.


V1.13.0

2026-09-01

Changed: terminology now corrects homophone misspellings directly

Once terminology is set, spans in the transcript that sound the same or nearly the same but are written differently are corrected back to the spelling of the term. You no longer need to list the possible misspellings in advance.

TermAppears in transcript asCorrected to
紡拓會訪拓會紡拓會
語者分離語這分離, 與者分離語者分離
晶圓晶園晶圓

This is a behavior change, and it runs in both directions: transcripts and translations for existing integrations will start to differ.

  • Homophone and near-homophone misspellings that were previously missed are now corrected (wider coverage)
  • Conversely, a misspelling that is itself an ordinary word is no longer corrected (see "Common-word protection" below). If you relied on such corrections, list that misspelling explicitly with fuzzy_correction

Existing recordings are not reprocessed; this applies to recordings created from this release onward.

Applicable languages: Chinese only (Traditional and Simplified are interchangeable, since they share pronunciation). Terms in Japanese, Korean or English do not participate in homophone matching — list their misspellings explicitly with fuzzy_correction.

For mixed Chinese-English terms (such as CVD製程), matching applies only to the Chinese portion; the Latin portion is left unchanged.

Matching range: Both identical and near-identical pronunciations are covered, including accent differences such as jin vs jing (final -n vs -ng) and retroflex vs non-retroflex initials. For the term 晶圓廠, for example, 金圓廠 in the transcript is corrected.

Common-word protection: If the span in the transcript is itself an ordinary word (金元 or 反案, say), it is not changed even when it shares a pronunciation with a term — this prevents normal sentences from being altered. To force such a correction, list the misspelling explicitly with fuzzy_correction: explicitly listed misspellings are not subject to common-word protection.

Changed: incorrect is now optional for Chinese terms in fuzzy_correction

When correct is Chinese (contains Han characters), incorrect may be omitted entirely — the system matches by pronunciation:

{ "fuzzy_correction": { "zh-TW": [{ "correct": "艾思通" }] } }

No misspellings are listed above, yet 愛思通, 愛時通, 愛司東 and 愛似通 are all corrected to 艾思通. Only spellings that sound quite different (愛自動) or have a different number of syllables (愛松) still need to be listed in incorrect.

This is a relaxation and existing integrations are unaffected — rules that carry incorrect behave exactly as before.

Note: Both conditions must hold: the language must be Chinese and correct must contain Han characters. Otherwise incorrect remains required — omitting it in those cases would have no effect at all, and accepting it would leave you believing the setting took. List misspellings explicitly for Japanese, Korean and English.

This applies to file import as well — both paths use exactly the same condition.

Added: config_updated reports terms that share a pronunciation

When two terms in your glossary sound alike (公事包 and 公式包, for example), config_updated carries an extra optional field, homophone_conflicts:

"homophone_conflicts": [
  { "languages": ["zh-TW"], "terms": ["公事包", "公式包"] }
]

languages lists every language that shares the same term index — Chinese regional codes (zh-TW, zh-CN, zh-HK and so on) are treated as one group for glossary matching, so they share a single entry rather than each reporting one.

When the transcript contains a third spelling with the same pronunciation, the system can only correct it to one of them, and which one is not guaranteed. Both terms themselves still work; the only ambiguity is which term an unregistered homophone misspelling is attributed to.

This is a warning, not an error: the glossary is still accepted and config still succeeds. If the pair matters to you, list the misspelling explicitly with fuzzy_correction so it is pinned to the term you want.

It is not present before the recording has started (the language list is not settled yet), the same as inactive_languages.

Fixed: the first segment_uploaded event was missing segment_index

Segment indices start at 0, and a value of 0 previously caused both segment_index and duration_sec to disappear from the JSON — meaning the first segment_uploaded of every recording carried no index, and very short segments (duration rounding to 0) also lost duration_sec.

Both fields are now always present on segment_uploaded. Other events are unaffected and do not carry these fields.

Removed: three optional fields on config_updated

  • auto_generated_fuzzy_correction no longer appears in updated[], which now carries only terminology, fuzzy_correction and translation_dict
  • auto_generated_fuzzy_correction_capped (value variant_budget_exhausted) is no longer returned
  • auto_generated_fuzzy_correction_skipped (value timeout) is no longer returned

The shape of the event itself is unchanged, as are the way terminology and fuzzy_correction are configured, their limits, and their error codes.

Client recommendations

  • If your code reads any of the three fields above, remove that handling — they no longer appear
  • fuzzy_correction is still required in these three cases:
    1. The misspelling is itself an ordinary word and is blocked by common-word protection (晶圓 heard as 金元, for example)
    2. The misspelling sounds very different from the correct term, such as a foreign brand name recognized as a phonetically unrelated word
    3. Misspellings in Japanese, Korean or English — those languages do not participate in homophone matching
  • If you previously listed many Chinese misspellings in fuzzy_correction to cover homophones, most of them can now be dropped. Keeping them is safe, and sometimes better — explicitly listed misspellings are not subject to common-word protection, so this is the only way to guarantee that a particular misspelling is corrected

Reference


V1.12.1

2026-08-27

Changed: the length limit for custom summary prompts has been raised

The custom prompt in custom mode goes from at most 2000 characters to 3000 characters.

This applies to summary_prompt in live recording, prompt in summary regeneration and ad-hoc summaries, and summary_prompt in file import.

The limit counts characters, so one Chinese character counts as one. Exceeding it returns summary_prompt_too_long.

Note: This is a relaxation and existing integrations are unaffected. Note, however, that the longer the prompt and the more instructions it carries, the smaller the share that is reliably applied — a higher cap does not translate into proportionally better results.


V1.12.0

2026-08-27

Translation dictionary is now grouped by language

The translation dictionary format changed from term-first to language-first, matching terminology and fuzzy_correction.

{
  "en-US": [
    { "source": "語者分離", "target": "Speaker Diarization" }
  ],
  "ja-JP": [
    { "source": "語者分離", "target": "話者分離" }
  ]
}

The previous format was a flat array in which all target languages shared one set of source terms, which imposed two limits:

  • Languages could not have independent dictionaries. If one language needed 500 terms and another needed 100 entirely different ones, they all had to go into the same array with gaps.
  • The entry cap was shared across all languages, so the more languages you used, the fewer terms each one could get.

The previous format is still supported, and existing integrations need no changes. The same dictionary sent in either format produces identical results.

This applies to live recording, broadcast, and file import.

Changed: the dictionary entry cap is now counted per language

Previously up to 3000 entries across all languages; now up to 3000 entries per language. This is a relaxation — no existing configuration is rejected.

Exceeding the cap returns config_too_many_dict_entries; the details in the response identify which language exceeded it.

Note: The total grows with the number of languages. In practice you hit the size limit on a single settings message first (roughly 8 languages at full capacity approaches it), rather than the entry cap itself.

Note: Clearing the whole dictionary is not supported: an empty dictionary is treated as not sending this setting at all. Clearing a single language is supported (send that language an empty array).

Changed: on resume, the dictionary is returned in the format you sent

The translation_dict in resume_ok now matches the format you last sent — send the previous format and you get it back; send the new one and you get the new one.

Changed: cap on dictionary entries carried into a single translation

At most 100 dictionary entries now apply to any single sentence (previously uncapped). Beyond that, longer source terms take priority.

You will not normally hit this: a sentence typically matches a handful of entries. The cap guards against cases where many single-character or very short source terms cause nearly every sentence to match a large set — and the more entries there are, the smaller the share that is reliably honored.

Changed: translations of the same text are now more consistent

Translating the same text twice previously could produce slightly different wording. Results are now stable. Dictionary content and behavior are unchanged, but translations may differ slightly from before.

What did not change: the target language is still matched exactly (an en-GB translation does not apply to en-US); the case rules (per-entry case_sensitive, defaulting to case-insensitive) are untouched; and the dictionary remains best-effort rather than literal substitution.


V1.11.1

2026-08-26

Fixed: the translation dictionary was not saved, so later re-translation and summary regeneration ignored it

A translation_dict set after start on a live recording did not apply to the following features:

  • Re-translating afterwards (whole transcript and single sentence)
  • Case-matching behavior when re-translating a single sentence
  • The dictionary guidance used when regenerating a summary

Symptom: with the same dictionary, your preferred wording took effect during live translation but not when re-translating afterwards. If you noticed that gap, your observation was correct.

The two are now consistent. This is a behavior change: re-translation and summary regeneration will start applying the dictionary, matching live translation. Existing recordings are not backfilled; this affects recordings created from this release onward.

Broadcast dictionaries had the same problem and are fixed in this release as well.

Added: Terminology Guide

Terminology documentation was previously scattered across sections and presented as field specifications, making the applicable scenario for each block difficult to determine. This release adds a terminology guide covering the stage at which each of the three blocks acts, selection criteria, how to select and write terms, when settings take effect, limits and their error codes, and five important notes.

All three scenarios are covered — live recording, broadcast, and file import. They use the same terminology format; only the transport differs (the three import fields are JSON strings rather than objects).

Location: Feature Guides → Terminology Guide.

Documentation Change: Terminology Documentation Streamlined

Terminology descriptions, field tables, and examples have been streamlined to essential content.

The API contract is unchanged and no client action is required. Existing integrations continue to work, and no error or warning is returned.

To ensure a term is recognized, register it under the correct language code — see the newly added terminology guide for details.


V1.11.0

2026-08-26

Changed: Translation dictionary limit raised to 3000 entries

The translation dictionary limit is raised from 500 to 3000 entries.

Dictionary size no longer affects the cost of each translation.

This is a behavior change: an entry no longer applies unless its source term matches the wording in the sentence. For example, a source term written in the plural (wafers) will not apply to a sentence saying wafer, and a source term of IPEVO Inc. will not apply to a sentence that only says IPEVO. The reverse (shorter source term, longer sentence) still applies.

We recommend setting source terms to the shortest form that will actually be spoken.

Changed: fuzzy-correction limits raised

ItemPrevious limitNew limit
Rules (across all languages combined)30004000

Added: notification when the variant allowance runs out

config_updated gains an optional field, auto_generated_fuzzy_correction_capped. It appears when the terminology list would generate more variants than the remaining allowance, with the value variant_budget_exhausted.

When it appears, automatic variants were partially accepted (the closest matches first) rather than rejected wholesale — the rules that were accepted take effect as usual. To have them all accepted, send fewer variants of your own or reduce the number of terms.

Changed: glossary limits are now defaults, with max in the response as the source of truth

The glossary limits (terminology count, fuzzy-correction rule count, translation-dictionary entries) can now be tuned per environment. The numbers in this documentation are defaults.

Suggested integration change: do not hard-code the limits. The error response's details has always carried both count (what you sent) and max (the limit in force) — read max instead. Existing integrations keep working unchanged; this is a recommendation, not a breaking change.

Changed: the provider value on translation errors

For translation-related errors (llm_content_filtered, translation_service_unavailable, and similar), details.provider now carries the value llm_service.

The previous value is no longer used; llm_service is always returned.

This is a change to a value you can observe: provider has always been debug context inside details, and the documentation has never listed it as an enumerated value to branch on, so most integrations are unaffected. If your code did compare against the previous string (for example to route alerts), switch to the new value — or better, branch on error_code, which is the field designed for that.


Documentation correction: optional fields on config_updated

The field table for config_updated previously listed only updated and terminology_effective. The following have existed since V1.10.1 but appeared only in this changelog; they are now in the event reference:

  • unknown_languages / inactive_languages / inactive_dict_languages: warnings about glossary language codes
  • auto_generated_fuzzy_correction_skipped (value timeout): automatic variant generation did not finish, so this round was not applied and the existing rules are kept; resending config helps
  • auto_generated_fuzzy_correction, the fourth possible value in updated[]

It is easy to confuse this with the auto_generated_fuzzy_correction_capped added in this release: capped means "finished but did not fit" (partially in effect, retrying will not help), skipped means "did not finish" (nothing changed, retrying will help).


V1.10.1

2026-08-25

Fixed: Fuzzy-word correction had no effect in live recording

fuzzy_correction configured for live recording (WebSocket) was never actually applied. Sending config returned config_updated and reconnecting returned the rules unchanged, but transcripts, translations, speech synthesis and summaries were all produced without the corrections — while the same glossary worked correctly for file import.

After this release, live recording and file import produce consistent results. If you ever observed "the same glossary works for import but not for live recording", that observation was correct.

This is a behavior change: transcripts and translations for existing integrations will start to differ (they now reflect the corrections).

Changed: Glossary entries apply per language

All three glossary blocks now use the language code to decide where they apply:

  • Terminology: applies only to recognition in the language it is registered under. Single-language situations (each multi-channel track, speaker diarization, file import) use only that language's terms; multi-language transcription and conversation mode use the languages declared for the session.
  • Fuzzy-word correction: a rule applies only to sentences in the language it is registered under — a rule under zh-TW will not alter an English sentence. When the sentence language cannot be determined, rules from every language are applied.
  • Matching is at language-family granularity: zh-TW/zh-CN/zh-HK are interchangeable, as are en-US/en-GB.

This is a behavior change, and it is silent: glossary entries registered under a language not used in the session have no effect, and no error is reported. Please make sure the language keys you send match the languages actually used in that recording.

Changed: Subtitles apply corrections at end of sentence

Interim subtitle results are not corrected; corrections are applied when the sentence completes (is_final). Integrators will see the subtitle adjust once at the end of a sentence.

Changed: Glossary limits raised

BlockPreviousNew
Fuzzy-correction rules5003000
Translation dictionary entries50500
Terminology500500 (unchanged)

All three are totals across all languages combined.

Note: The number of translation-dictionary entries affects translation cost. Only entries that have a translation for the current target language are carried, so spreading translations across languages reduces the per-request load.

Added: Reporting for glossary language keys

config_updated carries two new optional fields that flag settings which may not take effect:

  • unknown_languages: language codes that cannot be recognized (for example zh, chinese). Those glossary entries will not take effect.
  • inactive_languages: valid codes that are not used in this recording.

Both are advisory and do not interrupt the update. inactive_languages is not reported before the recording starts (the language list is not settled yet).

Added: Field validation for glossary entries

Terminology and fuzzy-correction entries are now checked for required fields and length on receipt. Invalid entries return config_invalid_entry, with details identifying the language, index and field. These problems were previously ignored silently.

Removed: Error code config_terminology_locked

This code was never emitted — updating terminology while recording has always been allowed.

Fixed: Terminology not applied for some file-import languages

For some languages, file import used no terminology at all (no error, no warning). After this release these languages behave like the rest.

Fixed: Documentation that did not match actual behavior

  • The terminology limit counts the number of terms; boost does not consume capacity (previous documentation stated this two different ways in different sections).
  • boost does not affect capacity in live recording; for a few languages in file import it does occupy extra slots.
  • The translation-dictionary limit was previously described as a recommendation in the guide and as a hard limit in the reference; the two are now consistent.

Fixed: The "exact case match" dictionary setting had no effect in post-processing

The case_sensitive flag on dictionary entries was ignored during post-hoc retranslation (both full-transcript and single-sentence): entries were applied case-insensitively whether or not the flag was set. Live per-sentence translation and summaries have always honored it; only post-processing did not.

After this release all three paths behave consistently for the same recording.

This is a behavior change: if you relied on entries with case_sensitive still being applied during post-hoc retranslation, they will no longer be substituted when the case does not match — which is what the setting was always meant to mean.

Changed: File-import summaries apply the translation dictionary

Summaries produced at the end of a live recording have always used the dictionary to keep proper nouns consistent; file-import summaries did not. They now behave the same.

When the summary is produced in the source language the dictionary has no matching translation, and nothing is injected (same as live recording).

Changed: Broadcast announcements and standby text apply the translation dictionary

Within a single broadcast, per-sentence subtitles were translated using the dictionary while announcements and standby text were not, so the same proper noun could appear two different ways. They are now consistent.

Behaviour is unchanged if you have not configured a translation dictionary.

Changed: Glossary settings are recorded with the recording

Terminology and fuzzy-word correction configured for live recording were not stored with that recording, so there was no way to check afterwards which settings had been in effect (file import has always stored them). The two are now consistent.

What is stored is the configuration you actually sent; rules the system derives automatically are not included.

Recommendations for integrators

  1. Check the format of your glossary language keys: use full codes (zh-TW, en-US), not forms such as zh or chinese. After this release, keys in the wrong format fail silently.
  2. Check which language your entries are filed under: if you placed English terms under a Chinese language key and relied on mixed-language speech to pick them up, they no longer apply to purely English sentences.
  3. Read the new config_updated fields: unknown_languages and inactive_languages are currently the only signal that surfaces configuration problems early.
  4. Subtitle adjustment is expected: interim results are uncorrected; corrections land at end of sentence.

V1.10.0

2026-08-21

New: multi-channel speaker separation (physical channel separation mode)

Added the multi_channel recognition mode: one recording takes input from multiple physical microphones (up to 8 channels), each channel is bound to a single transcription language, and speaker identity is determined directly by the channel — no inference from voice characteristics — making it a good fit for settings where every speaker has a dedicated microphone. Available for the transcribe and record types; not available for conversation or broadcasts, and cannot be combined with text-to-speech or multi-speaker diarization. This feature must be enabled for your account before use; when not enabled, start returns invalid_recognition_mode.

  • Physical channel separation mode: start carries recognition_mode: "multi_channel", channel_mode: "per_channel", and channels[] (each entry contains channel_id, speaker_name, and transcription_languages); the audio format supports pcm only, and every audio frame must carry a channel_id.
  • Adding and removing channels mid-recording: the new add_channel / remove_channel actions open or deactivate channels while the recording is in progress; after deactivation, the transcript and audio already produced on that channel are kept.
  • Per-channel language switching: the new set_channel_language action changes the transcription language of a single channel only — the channel number and speaker identity stay the same, and the transcript remains continuous. switch_language does not apply under multi-channel; use this action instead.
  • Channel status events: the new channel_status event reports each channel's state (preparing / ready / removed / error) and the current channel count, so integrators can display per-channel readiness.
  • Catch-up transcription after pause: during a pause the audio keeps being saved but produces no transcript; after resuming, speech from the pause is transcribed into the transcript (timestamps reflect the moment the words were actually spoken). Catch-up transcription is capped at the trailing 60 seconds combined across the whole session; anything beyond that is kept in the audio file only.
  • Every sentence in result events and in the transcript carries channel_id, and speaker_id is fixed per channel (format channel_{N}).

New error codes: this release adds multi-channel error codes (the channel_* / multichannel_* series); see Error Codes for the full list and descriptions.

Billing: multi-channel adds a per-minute surcharge based on the highest number of channels active within that minute; channel 1 is included in the base rate, and removing a channel takes effect from the next minute. See Pricing.

Full specification: WebSocket - Voice Translation.

Client Recommendations

  • Existing integrations are unaffected: recordings that do not use multi_channel behave and are billed exactly as before.
  • Multi-channel requires directional / close-talking microphones with sufficient spacing between them. Crosstalk from unsuitable equipment (one person picked up on several channels) is outside the quality guarantee; perform channel selection on the client side (at any given moment, send only the channel with the strongest signal) or ensure physical isolation.

V1.9.2

2026-08-20

Fix: some target languages were not translated in multi-language transcription

When transcription is configured with more than one language (transcription_languages) and the translation targets overlap with them, the overlapping language could be returned verbatim instead of translated.

Observed behavior (with both transcription and translation set to zh-TW / en-US / ja-JP / ko-KR):

Source textzh-TW output (before)zh-TW output (after)
Wait.Wait. (verbatim)translated
NI hao.NI hao. (verbatim)translated
Help with a.Help with a. (verbatim)translated

After the fix, the source language in multi-language mode is determined per sentence: the original text is kept only when that sentence's actual language matches the target, and everything else is translated. Single-language transcription and conversation mode are unaffected.

Data already affected: the fix applies to new recordings only; transcripts already stored are not re-translated automatically. To repair them, run re-translation on the affected recordings. Re-translation is billed by actual usage.

This also affects origin.language: in multi-language mode the field now reports the language determined for that sentence rather than a fixed value for the whole recording. When no determination can be made, the configured value is kept, so the field is never empty. If your integration reads this field to decide the language, note that it now varies per sentence in multi-language mode.

Fix: re-translating historical transcripts returned the original text

When re-translating a historical recording (GET /api/v1/sse/retranslate/{taskId}), some sentences were returned verbatim instead of translated — most often in older data, or where the language had been determined incorrectly at the time. This is now fixed.

Data already affected: sentences that previously came back untranslated can simply be re-translated again; no extra configuration is required. Re-translation is billed by actual usage.

Fix: some client errors were returned as server errors (500)

Certain request errors previously returned 500 internal_error, leading integrators to treat them as server faults and retry. They now return the correct status code:

SituationBeforeAfter
Path is correct but the HTTP method is not supported500 internal_error405 method_not_allowed (with an Allow header)
Uploaded content exceeds the size the server accepts500 internal_error413 http_error
Service temporarily unavailable (maintenance)500 internal_error503

Any other HTTP-level error keeps its original status code; when that status code has no dedicated error code, http_error is used.

Not found (404), forbidden (403), too many requests (429), validation failures (422 validation_failed) and unauthenticated requests (401) were already correct and are unaffected by this change.

Missing headers also fixed: throttled responses (429) previously omitted Retry-After and X-RateLimit-*, leaving clients unable to tell how long to wait. These are now preserved.

If your integration uses 500 to decide whether to retry, switch to the actual status code — a 4xx means the request itself needs to change, and retrying will not help.

Terminology limit: documentation corrected, and a combined check added for file import

The limit for live recording is unchanged, but the previous documentation did not match the actual behavior. Corrected in this release:

  • The 500 limit for live-recording config counts term entries across all languages combined, not 500 per language
  • boost consumes extra capacity only in conversation mode. Single-speaker, multi-language transcription, speaker diarization and file import are unaffected by boost; in conversation mode a term occupies clamp(round(boost), 1, 5) slots (rounded half up, so 1.5 becomes 2), and anything beyond the limit does not take effect

File import (behavior change): the effective terminology limit is likewise 500 entries across all languages combined. From this release, a combined total above 500 is rejected at upload time with a 422 stating the actual count; previously only the per-language limit was validated, so such an upload succeeded while the excess terms silently had no effect.

If your multi-language vocabulary exceeds 500 entries in total, trim it before uploading — the entries beyond the limit were never taking effect anyway.

Documentation: case flag for the translation dictionary

The case_sensitive field of translation_dict is now documented (the field itself was already supported; this release only adds the documentation): optional per entry, defaults to false = case-insensitive; when set to true the entry applies only on an exact-case match.

A case-flag comparison section has also been added. fuzzy_correction uses case_insensitive (defaults to false = strict) while translation_dict uses case_sensitive (defaults to false = permissive) — opposite field names, and opposite behavior from the same default value. Getting it wrong produces no error at all, only matching behavior opposite to what you intended.

Documentation: variant collision rules for fuzzy correction

When the same incorrect variant appears in more than one rule (collisions are resolved on incorrect, not correct):

  • The case flag resolves to strict wins — if any rule leaves case_insensitive off, that variant is matched with exact case
  • When several rules map the same variant to different correct terms, which one takes effect is not guaranteed; do not rely on it

Splitting one correct term across several rules with different case settings is therefore a safe and supported pattern, as long as their incorrect variants do not overlap. If the same incorrect variant needs to map to different correct terms, pick one.

Documentation: config is processed in order and aborts on the first failure

terminology, fuzzy_correction and translation_dict are processed in that order. When a block fails validation the request returns an error and aborts, so that block and everything after it is not applied — but any block processed before it is already in effect.

Do not assume the configuration is completely unchanged when you receive an error; fix the problem and resend the complete config. All three blocks replace their previous value wholesale and do not stack on top of earlier settings.


V1.9.1

2026-08-13

New: Ad-hoc Summary endpoint (POST /api/v1/sse/summary)

Added an SSE endpoint that generates a summary from text supplied in the request, not tied to any recording. It is intended for content the server does not have — for example, the full transcript of several merged recordings, or a transcript edited by the user. The result is only streamed back to the client and is never stored.

  • Authentication: only the X-API-Key header is accepted (no API keys in the query string); authentication failures return real 401/403.
  • Billing: 0.1 credits per 1,000 content characters (same rate as summaries); billed only on successful generation.
  • Duplicate-request guarantee: idempotency_key is required and only valid within a single API key; the system compares the entire request (content plus every summary parameter) to decide whether a call is a retry. Retrying with the same identifier + an identical request is not billed again (but regenerates); the same identifier + any differing field returns 409; failures do not claim the identifier.
  • Error contract: before the stream starts, real HTTP status codes are returned (401/403 / 422 / 404 / 400 / 402 / 409), unlike the "200 + error event" convention of other SSE endpoints.

Note: an earlier endpoint once existed at the same path (removed in V1.8.0). This endpoint is a brand-new contract — authentication, billing, and duplicate-request rules all differ, so do not reuse old integration code.

New error code (see Error Codes – Summary Errors):

Error codeHTTPScenario
summary_idempotency_key_conflict409The same idempotency_key was already used with different content

Full specification: Ad-hoc Summary SSE.

Client Recommendations

  • Use this endpoint when regenerating a summary from merged or edited full text; keep using Regenerate Summary for template or output-language changes.
  • Use a stable identifier from your system (such as a merge-batch ID or revision ID) as the idempotency_key; send a new key for each new piece of content.

V1.9.0

2026-07-31

New: unlimited plans

In addition to the credit model, this release introduces unlimited plans: a contract authorizes a fixed feature bundle and usage limits, and features included in the plan are not charged per minute during the authorized period. The plan's feature bundle and its limits (simultaneously recognized transcription languages, single-recording length cap, usage-hour thresholds, concurrent recording limit) are configured per contract.

  • Features the plan does not include are rejected outright — they do not fall back to credit billing.
  • Broadcasting is never included in unlimited plans; broadcasts are always billed in credits.
  • Base features included in every plan: base speech recognition, professional vocabulary, summary, and full-text re-translation.

See Pricing — Unlimited Plans.

New error codes (full details in Error Code Reference — Plan and Usage Limit Errors):

Error CodeScenario
plan_feature_not_allowedThe plan does not include the feature in use. Over WebSocket there are two occurrence points: start is rejected (the connection is not closed; adjust the parameters and retry), or the per-minute check while recording detects it (for example, a feature not in the plan was turned on mid-session) → the current recording is stopped. REST endpoints (creating a broadcast, floating subtitles, audio import) return HTTP 403
concurrency_limit_reachedThis API Key has reached its concurrent recording limit; the connection is not closed — retry after another recording ends
daily_limit_disconnectThe plan's usage threshold was reached and the current recording was stopped; you may start a new recording immediately
daily_limit_reachedUsage has reached the plan's limit; available again after the plan's reset (daily limits reset the next day)
plan_daily_limit_reachedREST: POST /api/v1/auth/ticket and the import upload when the plan's daily hard limit has been reached (HTTP 402)

too_many_languages semantics extended: details.max may now come from the plan's cap on simultaneously recognized transcription languages, in addition to the system-wide limit (10 transcription languages); details carries max and received.

New: plan lookup endpoint

Added GET /api/v1/me/plan (authenticated with X-API-Key, read-only, queryable even with zero balance): reports the current billing mode (credit / unlimited), the plan's feature bundle, one-off features, each limit with current usage, and the estimated time a restriction lifts. When a request is rejected by a plan limit (403 / 402), use this endpoint to answer "what does my plan include, how far am I from a limit, and when does the restriction lift?"

See My Plan API.

Changed: audio import response fields

  • POST /api/v1/imports/check-quota response gains data.reason: null (allowed) / insufficient_credit (topping up resolves it) / plan_not_allowed (the plan does not include audio import; a plan upgrade is required).
  • The 202 response of POST /api/v1/imports gains data.downgraded_features (array): when the plan includes audio import but not some requested sub-features (such as speaker diarization or translation), those sub-features are skipped and the import proceeds; the skipped features are listed in this field. An empty array means nothing was downgraded.

Changed: remain_quota semantics (existing field)

For API Keys with a dedicated credit allotment, remain_quota now reflects the credit actually available to that key (its dedicated allotment) instead of the account's total balance; accounts without dedicated allotments see no change. This affects remain_quota in POST /api/v1/imports/check-quota and remaining_budget in WebSocket error messages.

Client recommendations

  • Credit-based integrations require no changes; the error codes added in this release occur only with unlimited plans.
  • For plan-bound integrations, call GET /api/v1/me/plan when you receive plan_feature_not_allowed / plan_daily_limit_reached to show the user the plan contents and recovery time, rather than treating it as a service outage.
  • On concurrency_limit_reached and daily_limit_disconnect, both the connection and the account remain usable: for the former, retry after another recording ends; for the latter, you may start a new recording immediately.
  • If you display remain_quota, note that for API Keys with a dedicated allotment its meaning has changed to the key's allotment.

V1.8.0

2026-07-29

Changed: pricing update

This release updates the billing model for several services.

Rate changes

ItemPrevious rateNew rate
Text-to-speech (TTS)0.5 credits / min1.0 credits / min
Professional vocabulary0.5 credits / minFree
Audio import (base speech recognition)1.0 credits / min0.3 credits / min

Interpretation mode: TTS is now billed separately

The interpretation integrated rate (1.5 credits / min) previously included text-to-speech. From this release, TTS is billed separately at an additional 1.0 credits / min when enabled. The rate for interpretation without TTS is unchanged (still 1.5 credits / min). TTS can be toggled at any time during a session, and the rate adjusts from that minute onward.

Summary and full-text re-translation: now billed by text volume

ItemPrevious billingNew billing
Meeting summary (first generation)3.0 credits per action0.1 credit per 1,000 transcript characters
Re-generate summary3.0 credits per action0.1 credit per 1,000 transcript characters
Re-translation (full text)2.0 credits + 0.3 credits per minute0.1 credit per 200 characters

"Characters" refers to the actual character count of the transcript. Any partial billing unit is charged as a full unit, and each action is charged at least 0.1 credit. Single-sentence re-translation remains free.

Broadcast: cloud translation fee is now the sum of two items

The broadcast audience delivery fee previously used a single lookup of "maximum audience × translation-language tier". From this release, the "number of translation languages" and the "maximum audience size" are looked up separately and added together. Tiers for 10,000 / 20,000 / 30,000 viewers have been added (the default per-account limit remains 5,000; contact sales to raise it).

In addition, broadcasts are no longer charged the "from the 2nd translation language" surcharge — multi-language costs are now fully covered by the cloud translation fee.

Removed: POST /api/v1/summary endpoint

The REST endpoint for "generate a summary for any transcript" is discontinued as of this release.

Reason for removal: the endpoint stood outside the recording workflow and saw zero actual usage; its authentication also differed from every other REST endpoint (Authorization: Bearer vs X-API-Key), leaving it off the maintained path.

Alternatives: summary generation remains available through two recording-bound entry points:

ScenarioHow to use
Automatic generation after a recording endsThe summary_* fields of the WebSocket start action
Re-generate for an existing recordingGET / POST /api/v1/sse/regenerate/summary/{taskId}

Who is affected: integrations calling POST /api/v1/summary directly. If you were using it to summarize transcripts from external sources, please contact your service representative to discuss alternatives.

Change to content-filter downgrade: the documentation previously suggested switching to this endpoint to trigger automatic downgrade when SSE summary re-generation returned llm_content_filtered. With the endpoint removed, please revise the prompt or transcript content and retry instead.

GET /api/v1/summary-templates (summary template lookup) is a different endpoint and is unaffected.

Client recommendations

The pricing changes require no code changes; if your workflow estimates costs from the rate table, please re-evaluate using the new rates. See Pricing for the full rate table and billing examples.

If your integration uses POST /api/v1/summary, please migrate per the alternatives above.


V1.7.7

2026-07-25

New: deployed-version lookup endpoint

Added GET /api/v1/version (REST service) and GET /version (realtime service). Both report the currently deployed version and build identifier so integrators can run a version gate before going live — for example, "this fix requires ≥ vX.Y.Z".

Both endpoints require no authentication: a version gate that fails on authentication would be indistinguishable from a version mismatch, defeating its purpose. The response contains only the version and build identifier.

{ "service": "vas-api", "version": "1.7.7", "build": "a1b2c3d4e5f6" }

Note: The realtime service and the REST service are deployed independently and may run different versions. Query whichever service owns the functionality you need to verify.

Recommended version-gate logic: this endpoint ships in V1.7.7, so a successful response by itself proves the version is ≥ V1.7.7. A 404 means the version predates V1.7.7 and must be confirmed with your service contact.

See REST API · GET /api/v1/version.

Client Recommendations

No changes required. If your release process needs to confirm the VAS version, consider automating it with this endpoint instead of manual confirmation.


V1.7.6

2026-07-25

Fix: broadcasts using a custom summary prompt could fail to save the recording

When a host started a broadcast in custom summary mode (summary_mode: "custom") and that broadcast also had a shared summary template configured, the two settings conflicted and the recording could fail to save after the broadcast ended. The summary itself was still generated correctly from the custom prompt, which made the problem easy to miss.

With this release, custom summary mode uses exactly what the host passes in and no longer falls back to the broadcast's shared template. Broadcasts using shared template mode are unaffected.

Who is affected: broadcasts started over WebSocket start with summary_mode: "custom" where the broadcast also has a summary_template configured. If your broadcasts always use shared templates, you are not affected.

Documentation correction: broadcast settings have two layers — channel defaults and the current session

A broadcast is a "channel", and one channel can go live many times, so its settings split into two layers:

LayerWhat it changesInterface
Channel defaultsThe starting configuration for every future broadcastPATCH /api/v1/broadcasts/{id}
The in-progress sessionWhat that session actually usesHost-side WebSocket actions

PATCH updates the channel defaults, so transcription_languages, translation_languages, speaker_diarization, tts_config, summary_template, and summary_language take effect on the next broadcast. This lets a host adjust settings for the next broadcast while the current one is still live. The exception is access_type, pass_code, and max_viewers, which are viewer access controls and are applied immediately to the in-progress broadcast.

This is a correction to how the behavior is documented; the API behavior itself has not changed. The previous wording described the endpoint as adjusting settings "in real time" without distinguishing the two layers, which made it easy to assume that changing the summary template would affect the broadcast already in progress. The description and the parameter table have both been corrected.

Fix: recordings without a summary template could not change only the summary language or output format

conversation and broadcast recordings may omit summary_template (the summary then falls back to the system default template). Previously, using set_summary on such a recording to change only summary_language or summary_plain_text was incorrectly rejected with summary_mode_field_mismatch for a missing summary source.

With this release, requests that only change the output language or format no longer require a summary source. Explicitly enabling the automatic summary (auto_summary: true) or changing the summary source still requires a template or a custom prompt — that rule is unchanged.

Client Recommendations

  • If you call PATCH during a live broadcast to change the summary template and show it as "applied", change that to "applies to the next broadcast" — or switch to the WebSocket set_summary action so the current session's summary picks up the new settings (available since V1.7.5, and supported for broadcasts as well).
  • To use a custom summary prompt for broadcasts, pass it in the host-side WebSocket start action, and upgrade to this release to avoid the saving problem described above.

V1.7.5

2026-07-25

Summary settings can now be changed while recording

A new WebSocket action, set_summary, lets you change the summary settings that will be applied when the recording stops. It covers both the shared template mode (builtin) and the custom prompt mode (custom), and can also change the summary language, switch to plain-text output, or skip the automatic summary entirely for that recording.

A live recording generates its summary once, at stop time, so the change applies to that automatic summary; if you send the action several times, the last one before stopping wins. Previously the settings were fixed once recording began, and the only way to change them was to regenerate the summary after stopping — which produced an extra summary and an extra charge. That limitation is lifted in this release.

  • The summary source (mode / template / prompt) is replaced as a set; switching modes automatically clears the fields belonging to the other mode.
  • Changing the summary source does not turn the automatic summary on. If it was disabled when the recording started, send auto_summary: true as well.
  • After reconnecting, resume_ok returns the updated summary settings, so you can reconcile client state directly.
  • See set_summary for details.

Consistent records when no shared template is specified

When a recording uses the shared template mode without naming a template, the system already fell back to the default shared template to generate the summary. Starting with this release that default template is also stored in the recording record, so the template actually used matches what is recorded. This corrects internal records only; summary content and billing are unaffected.

Client Recommendations

No changes required. set_summary is a new capability and existing integrations keep their current behavior. If your product lets users switch summary templates mid-recording, consider adopting this action to avoid an extra summary generation.


V1.7.4

2026-07-23

Retrying an audio import now preserves custom summaries and vocabulary settings

When you retry a failed audio import, the original custom summary prompt and vocabulary settings (terminology, fuzzy correction, translation dictionary) are now preserved and rebuilt on retry. As a result, the retry regenerates the summary and applies the vocabulary as expected — both the transcript and the summary are produced.

Previously, a failed custom-summary import would not regenerate the summary on retry and required resubmitting the whole import. That limitation is removed as of this release — simply retry.

The translation dictionary now also influences summaries

When a summary is produced in a "target language" that has matching entries in your translation dictionary, summary generation will try to follow the dictionary's specified renderings, so the summary wording stays consistent with the translated transcript.

  • Scope: the end-of-session summary for live recording, and SSE "regenerate summary" (preview and persist).
  • When a summary is produced in the "source language", the translation dictionary (source → target) does not apply.
  • This is best-effort, not a guaranteed literal replacement; it is a behavioral extension and requires no changes to existing integrations.

Audio import transcripts now include summary provenance fields

Transcript records produced by audio import now also carry the summary provenance fields (summary_mode, summary_template, summary_language, summary_plain_text, summary_prompt_snapshot), consistent with live recording. These are additive, backward-compatible fields; existing integrations are unaffected.

Client recommendation

No changes required. If you previously switched to "resubmit the import" because a custom-summary import failed, you can now simply call retry to recover the summary.


V1.7.3

2026-07-22

Audio import supports custom summaries (custom)

Audio import summaries now support summary_mode=custom, consistent with live recording and SSE regeneration: you can send summary_prompt (fully replaces the built-in template) plus summary_prompt_slug (a custom identifier).

  • When summary_mode is omitted: behavior is exactly as before (uses summary_template).
  • summary_mode=builtin: summary_template is required.
  • summary_mode=custom: summary_prompt and summary_prompt_slug are required, and summary_template must not be present (mutually exclusive).

The custom prompt text is not persisted (consistent with live recording); therefore a retry of a failed custom import does not regenerate the summary (the transcript is still produced; only the summary is left empty). (This limitation is removed as of V1.7.4, above; from V1.7.4 onward, retry preserves the custom summary and regenerates it as expected.)

Fix: summary templates for audio import

Previously, specifying summary_template on an audio import had no actual effect — whether you passed meeting, interview or course, the resulting summary always used the same generic format.

As of this release, imports correctly apply the content of the specified template, matching the behavior of live recording.

Note: this means the summary content produced by existing import integrations will change (to match the template you specified). If you relied on a fixed summary format from imports, please re-check it. Behavior is unchanged when summary_template is not specified.

Fuzzy correction gains a case-insensitive option

Every rule in fuzzy_correction accepts a new optional field, case_insensitive. When set to true, all incorrect variants in that rule match regardless of letter case.

It defaults to false, which behaves exactly as before — existing integrations need no changes.

The flag is per rule, so the same correct term can be split across several rules with different settings:

{
  "zh-TW": [
    { "correct": "IPEVO", "incorrect": ["ltfo"], "case_insensitive": true },
    { "correct": "IPEVO", "incorrect": ["ivo"] }
  ]
}

With the settings above, LTFO, LtFo and ltfo are all corrected to IPEVO, while only the lowercase ivo is corrected — the personal name Ivo is left untouched.

Where it applies: the config action for live recording, and audio import. The resume_ok settings snapshot also returns this field (omitted when false).

Note: enabling it widens the false-positive surface. If a variant is identical to an ordinary word or a personal name (for example ivo versus Ivo), keep the default exact-case matching. Chinese rules are unaffected (Chinese has no letter case).

Client guidance: no changes required; enable it per rule when you need it.


V1.7.2

2026-07-22

Two-way translation mode accepts speaker_diarization (behavior change)

Previously, a two-way translation request (type=conversation) carrying speaker_diarization=true was rejected. As of this release it is accepted, and the parameter is ignored — two-way translation does not perform speaker separation.

Behavior for other recording types is unchanged.

Billing: the per-minute rate is the same as a two-way translation session without this parameter; nothing extra is charged. Note, however, that these requests were previously rejected, so no recording was created and nothing was billed; from this release they create a recording and begin billing normally.

Client guidance: integrations that stripped speaker_diarization from two-way translation requests to work around this error can stay as they are; no change is required.

Two-way translation no longer reports a stale translation language after a mid-session change (fix)

After a two-way translation session changed languages mid-session, translation_languages in resume_ok, in the Webhook, and in the recording record still showed the language from before the change. As of this release it reflects the current language, as does the language information for floating subtitles.

Live translation itself was unaffected and was always correct.

For two-way translation, translation_languages holds a single language representing the current counterpart language; if the language was changed mid-session, the final record holds the last one.

Client guidance: if you relied on translation_languages from resume_ok or the Webhook to determine the translation language, that value is only trustworthy from this release onward.

Two-way translation now validates speakers language codes at start (fix)

Previously, when two-way translation specified languages via speakers, an unsupported language code still allowed the connection to start and audio to be sent, but the recording was never created — so no transcript was kept and no usage was recorded.

As of this release, start returns 400 invalid_transcription_language, with details.field set to speakers[].language and details.speaker_id identifying which speaker.

Client guidance: make sure speakers uses language codes from the supported list. If your integration contained a misspelled code, you will now receive an explicit error instead of the previous silent failure.

Documentation corrections

  • Corrected several statements claiming that all recording types are translated in real time (record has not supported translation since v1.7.0)
  • Capability matrix corrected: broadcast does support speaker separation (with a single transcription language only)
  • Two-way translation rules table now lists: translation_languages is set by the server, and realtime_translation is fixed to true

V1.7.1

2026-07-22

Multi-language transcription is no longer surcharged (billing change)

Specifying multiple transcription languages (automatic language detection) is no longer surcharged as of this version — the number of source languages does not affect the per-minute rate. The former "Multi-language transcription (from the 2nd language, each +1) — 0.3 points/minute" line has been removed from the rate table.

Surcharges for translation output languages are unchanged (sentence translation +0.2, real-time translation +0.4 per additional translation language).

Interpretation (conversation) is the most clearly affected: interpretation always requires exactly 2 transcription languages by specification, so it was previously surcharged a fixed 0.3 points per minute. After this change, interpretation is billed at the integrated rate published in the rate table.

Client guidance: no changes required. Recordings that use multi-language detection (including all interpretation recordings) will see a lower per-minute rate. See Pricing.


V1.7.0

2026-07-22

record recording type made lightweight (breaking change)

Starting this version, record (plain recording) is positioned as a lightweight speech-recognition-only type, with the following behavior changes:

  • Translation no longer supported: Sending translation_languages for a record returns 400 record_translation_not_allowed (consistent across the start, switch_language, and retranslate live paths, as well as the post-hoc transcript-retranslation / summary-translation endpoints).
  • TTS no longer supported: Sending tts_enabled=true for a record returns 400 record_tts_not_allowed.
  • Summary now off by default, opt-in: record no longer generates a summary automatically by default; it is generated only when summary_template or summary_mode=custom is provided. Sending auto_summary=true without a template returns 400 record_summary_requires_template.
  • Speaker separation remains available: record can still send speaker_diarization=true for multi-speaker separation.

For billing, a plain record (with summary off) counts only speech-recognition usage, without translation or TTS costs.

New error codes: record_translation_not_allowed, record_tts_not_allowed, record_summary_requires_template (see Error Codes · Record Type Restriction Errors).

Client guidance:

  • If you use record for plain voice notes (no translation / summary / TTS needed): no changes required, and it costs less.
  • If you previously sent translation_languages or tts_enabled=true for record: remove these fields, or switch to the transcribe type.
  • If you previously relied on the summary auto-generated by default for record: explicitly provide summary_template or summary_mode=custom, otherwise summaries will no longer be generated automatically after this version.
  • Existing record recording data is unaffected (reading, playback, existing transcripts and summaries all work normally); only re-translating them afterward is rejected.

V1.6.10

2026-07-16

Floating subtitle feed now supports broadcast hosts (during a live broadcast)

While a broadcast is live, a broadcast host can subscribe to their own live-transcript floating-subtitle feed using the same flow as a regular recording: exchange an API Key for an owner feed_token (Floating Subtitle Feed Token), then connect to the Floating Subtitle SSE (Floating Subtitle SSE).

Client guidance: For broadcasts, listen for the broadcast_recording_ready event to obtain the finalized task_id after going live (the standby session_started carries the initial ID, which returns 425 when exchanging for a token), then exchange it for a feed_token.

Endpoints, parameters, and response formats are unchanged; floating-subtitle behavior for regular recordings is unaffected.

New machine-readable status field on status events

The status events for pause / resume / stop (over WebSocket and the floating-subtitle SSE) now include a status field: live / paused / ended. See WebSocket Events · status and Floating Subtitle SSE · status.

Client guidance: The floating subtitle window (especially the host's own) should act on the status field — paused → freeze, ended → close the window, live → resume. Because the floating-subtitle SSE does not close automatically when the recording stops, failing to close on ended will leave it frozen on the last sentence. message is display text with no guaranteed format — do not parse it to determine state. This field appears only on those three lifecycle transitions; other status events such as set_name do not carry it. Backward compatible: existing fields are unchanged, and integrations that ignore this field are unaffected.

Floating Subtitle Feed Token: 425 / 410 split

When exchanging for a floating-subtitle feed_token (both owner and audience endpoints), "recording not ready" and "recording has ended" previously both returned 425. They are now split:

  • 425 Too Early: recording not ready (being created) → retry after a short delay.
  • 410 Gone: recording has ended → do not retry.

See Floating Subtitle Feed Token. Client guidance: on 410, stop retrying and close the floating subtitle window.

Language-switch events now carry the full translation-language set

The language_switch_start, language_switch_done, and translation_language_removed events (over WebSocket and the floating-subtitle SSE) now include a translation_languages field: an authoritative snapshot of the current full set of translation languages. See WebSocket Events.

Background: Previously these events carried only a single translation_language, so consumers could not tell "add a language" from "replace the language" and could drop an existing language by mistake.

Client guidance: On these events, overwrite your local translation-language set directly with translation_languages instead of inferring the operation from the single translation_language. This also resolves op:add / replace / out-of-order / reconnect-gap language-set sync issues. Backward compatible: purely an added field; integrations that only read translation_language are unaffected.

Floating-subtitle reconnect consistency: speaker events in replay, connected reflects the current language set

Two behavior fixes on the Floating Subtitle SSE (Floating Subtitle SSE) that resolve "reconnecting or late-joining viewers see stale information":

  • Speaker events are included in replay: speaker_renamed / speaker_reassigned / speakers_merged / speakers_auto_merged are now replayed in original order (after the sentences they affect). Client guidance: handle them during replay exactly as in live mode — retroactively update the speaker labels of existing sentences by affected_sids; otherwise reconnecting viewers will see pre-rename speaker names.
  • connected reflects the current language set: the translation_languages in the connected event is now the current authoritative set (reflecting mid-recording language additions/removals), no longer the value frozen at recording start.

Backward compatible: no new fields and no format changes; replay is simply more complete and the snapshot more up to date.


V1.6.9

2026-07-14

Behavior change: translation output language limit raised to 12

The translation_languages count limit is raised from 8 to 12 (applies to live recording, import, and broadcast; effective for direct API access). This is a relaxed boundary and is fully backward compatible with existing integrations that use ≤8 languages.

  • The input axis is unchanged: the transcription (source) language limit remains 10; input and output are two independent dimensions.
  • The broadcast audience delivery fee adds a "9–12 languages" rate band (see the Pricing Guide).
  • The too_many_languages error covers both axes: it is triggered when transcription > 10 or translation > 12.

V1.6.8

2026-07-13

Added: multi-language transcription input for broadcasts

When creating or updating a broadcast, the new transcription_languages field (array of strings, up to 10, distinct) is now the primary field for transcription (source) languages, letting you specify multiple languages for continuous multi-language recognition.

  • The legacy transcription_language field (single string) is now deprecated but still supported (kept for backward compatibility; it equals the first element of transcription_languages).
  • The create/update response returns both fields: transcription_language (= the first element, for backward compatibility) and transcription_languages (the full array).
  • When summary_language is not specified, it now defaults to the first transcription language (the first element of transcription_languages).

Added: source language array in audience info

Broadcast audience info (/info) now includes source_langs (array of strings) alongside the existing source_lang (= the first element), listing all transcription languages for the broadcast.

Behavior change: speaker diarization supports only a single transcription language

A broadcast with speaker_diarization enabled supports only a single transcription language. If multiple transcription languages are provided, the create/update request returns 422 (multi-language transcription and speaker diarization are mutually exclusive).

Client recommendations

  • For new integrations, use transcription_languages to specify transcription languages; transcription_language still works but is deprecated and should be phased out.
  • When reading broadcast settings and audience info, rely on transcription_languages / source_langs; transcription_language / source_lang are retained as the first element for compatibility only.
  • Keep broadcasts that require speaker diarization on a single transcription language to avoid a 422.

V1.6.7

2026-07-10

Behavior change: real-time multi-language translation for live recording

When multiple languages are specified in translation_languages (up to 8), every recording type (transcribe / record / conversation / broadcast) now translates all specified languages in real time. Previously, non-broadcast types only translated the first language and silently ignored the rest; this release brings the behavior in line with the long-standing documentation (Multi-Language Translation).

How results arrive: each language returns its own independent result event (same sid, single language key inside translations); multiple languages are never merged into one event. Note for existing multi-language clients: you previously received only the first language's translation — after this upgrade you will start receiving all languages. Accumulate translations by "sid + language code" instead of overwriting.

  • Word-by-word (interim) real-time translation requires realtime_translation: true; with the default false, all languages are translated once the sentence is finalized
  • If one language fails, the remaining languages are still delivered; the failed language additionally receives an error event whose details.translation_language identifies it

Behavior change: switch_language redefined as add / remove for multi-language sessions

Sessions with 2 or more translation languages must include the op parameter in switch_language:

  • op: "add": adds a single language and automatically backfills existing sentences (same response sequence as a single-language switch); the limit is 8 languages
  • op: "remove": removes a single language and returns the new translation_language_removed event; existing translations are kept, and at least 1 language must remain
  • Omitting op in a multi-language session returns the new switch_language_op_required error (this prevents the legacy replace semantics from corrupting the language set)

Single-language sessions keep the existing replace-and-retranslate semantics; no client changes are needed for them.

New error codes: switch_language_op_required, switch_language_already_exists, switch_language_not_in_session, switch_language_last_language (see Error Codes).

Billing

The billing rules for multi-language translation are unchanged (billing has always been per language count, with a per-minute surcharge starting from the second language — see the Pricing Guide); this release simply brings the actual translation delivery in line with billing. Credit consumption reminder: multi-language real-time translation consumes noticeably more credits per minute than single-language (8 languages with real-time translation is roughly 2.8× or more). If credits run out, the recording stops per the existing rules — keep an eye on your balance.

Client recommendations

  • In multi-language sessions, render translations keyed by the language code in translations; do not let the last event overwrite the others
  • For word-by-word multi-language subtitles, include realtime_translation: true
  • Use op: "add" / op: "remove" to adjust languages in multi-language sessions; receiving a switch_language_op_required error indicates the backend is on v1.6.7

Reference documentation


V1.6.6

2026-07-09

Added: Pricing page

Added a Pricing page with the full credit-based rate card:

  • Credit pricing: 1 credit = TWD 1.5 (≈ USD 0.047).
  • Base usage (real-time recording / import): per-minute rates for speech recognition, vocabulary refinement, multi-language transcription, speaker identification, translation (including a per-additional-language surcharge), and text-to-speech.
  • Interpretation mode: real-time interpretation at an integrated rate.
  • Value-added services (one-time): meeting summary, re-generated summary, re-translation.
  • Broadcast billing: host side (content processing) + audience side (audience delivery fee by max audience size × translation language band), with the full rate table and billing examples.

Behavior change: transcript input language limit raised to 10

The maximum number of "input (source/transcription) languages" for transcription has been raised to 10. Affects real-time recording (WebSocket) and audio import.

  • Speaker identification mode still supports only a single source language (multi-language transcription and speaker identification are mutually exclusive); this limit does not apply to that mode.
  • The translation (output) language limit is unchanged (still 8).

Documentation clarification: broadcast maximum audience size is at least 1

When creating a broadcast, max_viewers is at least 1 (the API already enforced a minimum of 1; this release makes it explicit in the docs and admin panel as well).


V1.6.5

2026-07-06

Behavior change: silence thresholds lengthened across all five speaking-speed levels

The silence thresholds for sentence segmentation mapped to the five speaking_speed levels have been adjusted. Longer thresholds tolerate mid-sentence pauses better and reduce sentences being cut prematurely, at the cost of segmentation landing slightly later at every level.

LevelOld thresholdNew threshold
very_fast150ms300ms
fast300ms600ms
normal (default)500ms800ms
slow700ms1200ms
very_slow1000ms1500ms
  • The default is still normal, but its threshold changes from 500ms to 800ms.
  • Level names, the API surface, and set_speaking_speed usage are unchanged; no integration changes are required after upgrading.

Client recommendation

  • Expect segmentation to land slightly later and sentences to be more complete. If you relied on very_fast for the most immediate segmentation, its threshold changes from 150ms to 300ms (very_fast remains the fastest option).

V1.6.4

2026-07-04

Added: resume_ok returns the recording settings (state reconcile)

  • resume_ok now includes a settings object with the recording settings currently held by the session (speaking speed, audio format, languages, TTS, terminology / fuzzy correction / translation dictionary, summary settings, recording name, conversation dialog mode and speaker language map, etc.), letting clients reconcile local caches against the server's authoritative values after a resume or full page refresh.
  • Values are current (reflecting mid-recording changes via set_speaking_speed / config / set_name / set_tts / switch_conversation_mode / set_speaker_language) and presented in API format (e.g., speaking_speed returns the level string "normal").
  • See Connection and Authentication — the settings object for field details.

Resume format constraint (important)

  • settings.audio_format returns the audio format from start; resumed connections must keep using the same format (the server decodes with the original format and does not renegotiate).

Behavior change

  • The conversation_language_change_failed error response no longer includes internal error details in details (consistent with set_speaking_speed_failed; branch on the error code and show a generic failure message).

Documentation fix

  • Corrected the set_speaking_speed support scope: supported in all recognition modes except multi-speaker (multi_speaker) — including broadcast and multi-language LID — not just single / conversation mode.

Client recommendations

  • After receiving resume_ok, treat settings as the authoritative source and reconcile your local settings cache (especially for refresh / multi-tab scenarios).
  • Older clients can ignore the settings field; existing behavior is unaffected.

V1.6.3

2026-07-03

Speech Segmentation Settings Update

Added: Adjust speaking speed during recording

  • Added the set_speaking_speed action to adjust the speaking speed (the silence threshold for segmentation) while recording. Applying the change causes a brief interruption in recognition. Success response: speaking_speed_changed (returns the applied speed). Supported in single / conversation mode only; not applied in multi-speaker (multi_speaker) mode.

Behavior change: segmentation_mode removed

  • The segmentation_mode option (auto / by_time) in start options has been removed and is no longer available. Segmentation timing is now controlled solely by speaking_speed. If this field is still sent, the server ignores it without affecting other parameters.

speaking_speed description update

  • The default normal corresponds to a 500ms silence threshold (the same in every environment).
  • Five levels: very_slow (1000ms) / slow (700ms) / normal (500ms) / fast (300ms) / very_fast (150ms).

Client recommendations

  • To adjust segmentation timing during recording, use set_speaking_speed (send after the UI control is released to avoid repeated STT rebuilds).
  • If you previously sent segmentation_mode, you can remove that field (the server now ignores it).

Reference


V1.6.2

2026-07-02

set_name (Set Recording Name) fixes and identification improvements

Behavior changes

  • When the recording name exceeds the length limit, the server now returns a clear set_name_too_long error, and the response details includes max_length (previously no message was returned and clients would time out).
  • The recording name length limit is 60 characters.

Success response identification (new event field)

  • The set_name success response now includes event: "name_set" and a name field so clients can identify it precisely.
  • For backward compatibility, the success response keeps action: "status" unchanged.

Deprecation notice

  • Relying on action: "status" to detect set_name success is now deprecated and may be removed in a future version. Identify success via event: "name_set" (together with the name field) instead.

Error codes

  • The set_name error codes are set_name_empty and set_name_too_long.

Client recommendations

  • Clients that integrate the WebSocket protocol directly and detect set_name success via action: "status" should switch to event: "name_set". No changes are needed for other clients.

Reference


V1.6.1

2026-06-29

Added: Floating Subtitle Audience Sharing

In addition to the recording owner, the floating subtitle now supports audience sharing: the owner can enable sharing and obtain a share secret (to embed in a share link / QR code), letting other on-site audience members view read-only, with no login, no API Key, and no separate charge.

New Endpoints

MethodEndpointDescription
POST/api/v1/auth/tasks/{taskId}/subtitle-shareOwner enables / resets audience sharing and obtains a share secret
DELETE/api/v1/auth/tasks/{taskId}/subtitle-shareOwner stops sharing
POST/api/v1/public/tasks/{taskId}/subtitle-feed-tokenAudience exchanges a share secret for a read-only audience Token (no authentication)

New Events

  • The Floating Subtitle SSE adds a viewers event reporting the current viewer count and limit (count / max).
  • The Floating Subtitle SSE adds a subtitle_closed event: when the host "closes sharing" or "stops the recording", the server proactively sends it and ends the audience connection; the audience client should stop reconnecting on receipt. The host's own connection is unaffected.

Behavior

  • The number of audience members per recording is capped (server-configured, default 10, excluding the owner). Use the max field of the viewers event rather than hardcoding a value. When the audience limit is reached, the connection returns 429.
  • "Closing sharing" by the host immediately ends all audience connections (sending subtitle_closed), and the share link is invalidated so reconnection is rejected; "stopping the recording" behaves the same. The audience is a passive receiver.

Client Recommendations

  • Existing floating-subtitle (owner) integrations are unaffected and require no changes.
  • To offer shared viewing: on the owner side, call subtitle-share to obtain a share secret and build a share link; on the audience side, exchange the secret for a Token via the public endpoint, then connect to the Floating Subtitle SSE.
  • The audience client must handle the subtitle_closed event: close the connection and stop auto-reconnecting on receipt to avoid futile reconnection.

Reference


V1.6.0

2026-06-27

Breaking Change: Removed the legacy recording_id naming (naming unification complete)

The recording_id → task_id naming unification announced since V1.4.1 reaches its final step in this release: the legacy recording_id field and legacy paths are fully removed. All task identifiers are now unified under task_id.

The value of task_id is exactly the same as the old recording_id (the UUID of the same recording). This migration only renames the field / path; the identifier itself is unchanged.

Affected Scope (changes required)

  1. WebSocket payload: The session_started and resume_ok events no longer carry the recording_id field — read task_id instead.
    • This item was previously announced for removal in V2.0.0; it has been brought forward and completed in V1.6.0.
  2. REST endpoints removed: The following legacy recordings paths are removed; use the corresponding tasks paths (identical behavior, same identifier value):
    Removed (legacy)Use instead (new)
    PATCH /api/v1/recordings/{recordingId}/speakers/renamePATCH /api/v1/tasks/{taskId}/speakers/rename
    PATCH /api/v1/recordings/{recordingId}/speakers/reassignPATCH /api/v1/tasks/{taskId}/speakers/reassign
    PATCH /api/v1/recordings/{recordingId}/entries/{sid}PATCH /api/v1/tasks/{taskId}/entries/{sid}
  3. SSE connection message text: The connected message label for the history / retranslation / summary-regeneration streams changes from (recordingId: ...) to (taskId: ...) (a plain-text hint; the identifier value is unchanged).

Not Affected (no changes needed)

  • current_recording_id: The current_recording_id field in the broadcast query response is kept unchanged (it means "the UUID of the currently in-progress recording" — an unambiguous meaning, not a target of this cleanup).
  • SSE retranslation path: GET /api/v1/sse/recordings/{taskId}/entries/{sid}/retranslate is kept unchanged (the recordings segment in the path is existing naming; the parameter is already taskId).
  • Task queries, audio / transcript export, Webhook (data.task_id), and other interfaces already centered on task_id are completely unchanged.

Client Recommendations

  • Change any code reading recording_id to read task_id (same value, drop-in replacement).
  • Change any REST calls to /api/v1/recordings/{id}/... to /api/v1/tasks/{id}/....
  • If you previously string-matched recordingId: in the SSE connected message, match taskId: instead (better still, switch to event-type-based detection rather than relying on message text).

Reference


V1.5.12

2026-06-26

Documented recording options sub-fields (speaking_speed / segmentation_mode / profanity_handling)

The options sub-fields of the WebSocket start action are now documented and can be used to fine-tune STT sentence segmentation and profanity handling:

  • speaking_speed: very_slow / slow / normal (default) / fast / very_fast — adjusts the silence threshold for segmentation; use a slower setting for slower speakers to avoid cutting on mid-sentence pauses.
  • segmentation_mode: auto (default) / by_time — segmentation strategy; by_time combines with speaking_speed to adjust the threshold.
  • profanity_handling: mask (default) / remove / show — profanity handling.

Client Recommendations

  • All are optional; when omitted, defaults apply (normal / auto / mask), so existing integrations need no changes.
  • For slower speakers, or to avoid cutting sentences, try speaking_speed: slow.
  • Multi-speaker mode (multi_speaker) does not currently apply speaking_speed / segmentation_mode.

Reference


V1.5.11

2026-06-25

New Feature: Floating Subtitle Transcript Feed

Added a read-only "floating subtitle" transcript stream: you can subscribe over an independent connection to the live transcript of an in-progress recording (source-language original text + target-language translations), suitable for desktop floating-subtitle windows, second-screen captions, and similar use cases. It is separate from the recording's own connection and can be opened independently on a different device or window.

  • Exchange for a token: POST /api/v1/auth/tasks/{taskId}/subtitle-feed-token (exchange an API Key for a short-lived feed_token bound to the recording; only the recording owner can exchange).
  • Subscribe to the stream: GET /tasks/{task_id}/subtitle?feed_token=...&lang=... (SSE). Receives connected / result (original text and translations) / status / speaker and interpretation language-switch events. Original text is replaced in place by sid+is_final; translations inherit the speaker by matching sid to the original line.
  • Supports lang filtering of target languages, replay on connect, and automatic reconnection. Connection timing boundaries: not started 425, ended 410, invalid token 401, too many connections 429.

Client Recommendations

  • Exchange for the feed_token after recording has started; if you receive 425 (recording not ready) right after starting, retry after a short delay.
  • The feed_token is valid for 15 minutes and is extended automatically while connected; for long recordings, exchange for a new one before expiry.

Reference


V1.5.10

2026-06-20

New Feature: Per-Key Source IP Rules (Allowlist + Denylist)

You can configure source IP rules for each API Key (managed in the user portal), applied to both the REST API and live WebSocket:

  • Allowlist (allow): only IPs on the list may use the key.
  • Denylist (deny): a matching IP is denied — useful for blocking specific malicious/attacking IPs while letting everything else through.

Precedence (deny first): a match on any deny rule is rejected; if an allowlist exists, the IP must match one of its entries; both empty = unrestricted (backward compatible). When both are set, the effective permitted set = on the allowlist AND not on the denylist (an IP listed in both is denied). Even a leaked key cannot be used from an unauthorized (or denied) IP.

New Error Codes

  • auth_ip_not_allowed (403): The source IP is not permitted (not on the allowlist, or matched the denylist). Access from an authorized IP address.
  • auth_account_blocked (403): The account has been blocked. Contact technical support.

Behavior Change

  • When an account is blocked, API Key verification now returns 403 auth_account_blocked (distinct from the 401 returned for an invalid API Key); the live WebSocket handshake is rejected with the same reason.

Client Recommendations

  • If you have configured IP rules for a key, make sure all callers (including WebSocket) originate from authorized IP addresses; deny takes precedence over allow.
  • Handle auth_ip_not_allowed and auth_account_blocked in your 403 handling; both are fatal — retrying will not help, so correct the source IP or contact technical support.

Reference


V1.5.9

2026-06-10

Enhancement: Time-Base Fields Added to Session Resume

Three time-base fields were added to session resume to help clients align with the server clock and the transcript timeline.

Added

  • session_started now includes server_time (the server's current unix time in milliseconds, as a clock-base reference; clients can estimate clock skew by comparing it against the client time at the moment of receipt).
  • resume_ok now includes server_last_offset_ms (the transcript timeline position at the breakpoint, in milliseconds, based on the length of audio already processed).
  • resume_ok now includes server_recording_ms (the recording-head timestamp the transcript timeline resumes from after reconnect, in milliseconds, including silence; lets the client align its recording-second header to the same timeline the transcript uses).

Note

  • Do not mix the two notions of time: the grace period (resume_grace_seconds) is measured in wall-clock time and keeps counting down during the disconnect, whereas server_last_offset_ms is based on the audio timeline and is frozen during the disconnect. Always use wall-clock time to decide whether a reconnect is still possible.

Reference


V1.5.8

2026-06-08

New Feature: WebSocket Session Resume

When a WebSocket connection is unexpectedly dropped, clients may reconnect within a grace period (default 45 seconds) with their resume_token to rejoin the original recording session — keeping the same recording_id and sentence ids (sid), with the transcript timeline continuing from the breakpoint. No need to restart the whole recording.

Added

  • session_started now includes resume_token and resume_grace_seconds (store them).
  • New resume_ok event (resume succeeded, includes server_last_sid).
  • 4 new resume error codes: resume_token_invalid, resume_grace_expired, resume_ownership_mismatch, resume_unavailable (all error severity, not fatal).
  • Added a reconnect implementation example (JavaScript) to the Connection doc.

Client guidance

  • Store resume_token on session_started.
  • On detecting a dropped connection, within the grace period: obtain a new Ticket → reconnect with Sec-WebSocket-Protocol: ["ticket.<new>", "resume.<token>"].
  • After resume_ok, start a fresh audio stream as after start (WebM must send a new container).
  • On any resume_* error → obtain a new Ticket and send a fresh start.

Unchanged

  • Existing clients that do not use session resume require no changes; disconnect behavior is unchanged (always a fresh start).
  • Audio during the few seconds of disconnect is not recovered and is not billed (the timeline continues seamlessly).
  • api_key is the trust boundary: a "last connection wins" policy applies; do not share a single api_key across trust domains.

Reference


V1.5.7

2026-05-20

Documentation Update (No API Behavior Changes)

The public API behavior is completely unchanged. This release is a documentation supplement and wording revision.

New Usage Guide: Summary Prompt Customization

Added the Summary Prompt Customization Guide, consolidating the summary customization specs that were previously scattered across 6 reference documents into a single guide:

  • The mutual-exclusion rules and use cases for the builtin and custom summary modes
  • The corresponding fields for the three entry points: REST POST /api/v1/summary, the WebSocket start action, and SSE regenerate/summary
  • Transcript record fields (including the summary_prompt_snapshot audit field and the summary_fallback_level / summary_dropped_segments fallback audit fields)
  • A Profanity and Sensitive-Word Handling section, integrating the three paths (customer prompt -> neutral mode, transcript -> STT profanity_handling masking, transcript -> summary-layer segment omission) and explicitly stating that the API layer does not proactively reject requests containing sensitive words
  • The built-in safety guard (content-neutralization guidance, prompt-injection protection) and character-length limits
  • Complete examples for Node.js, Python, and WebSocket

The "Feature Guides" table on the documentation home page now includes an entry for this guide.

Documentation Wording Revision

Public-facing wording refined to use more general descriptions (field values such as summary_fallback_level are unchanged; wording only).

Reference


V1.5.6

2026-05-19

Documentation Alignment Fixes (No API Behavior Changes)

This release is a documentation proofreading pass; the public API behavior is completely unchanged. If you previously implemented against the older documentation, please adjust to the current spec for the items below.

Token Formats

  • broadcast_token: a 4-character short code (character set a-z0-9)
  • viewer_access_token: a 64-character alphanumeric string (not a JWT, no payload structure; do not attempt to parse it)

HTTP Status Codes

  • sse_missing_target_lang / sse_unsupported_language: 422
  • broadcast_token_invalid (viewer verify endpoint): 401

Error Code Strings

  • POST /api/v1/imports insufficient quota: stt_quota_exceeded
  • Broadcast not found on viewer SSE: broadcast_session_not_found
  • Broadcast at capacity on viewer SSE: broadcast_capacity_exceeded
  • The context for the sse_translation_failed error event is sse

WebSocket Event Naming

  • retranslate success event: action: "translation"
  • Audio upload failures: delivered via a type: "error" envelope (error_code is storage_upload_failed / storage_connection_failed / storage_queue_full); there is no separate upload_error action

Newly Documented Error Codes

Endpoint / ActionError CodeDescription
WebSocket set_nameset_name_empty / set_name_too_long / set_name_not_readyReplaces the older name_too_long
WebSocket audioaudio_process_failedAudio processing fails repeatedly (HTTP 500; reconnecting is recommended)

Reference


V1.5.5

2026-05-13

Breaking Change: The Summary API Is Now Mode-Aware

The "template + custom_prompt combined" design introduced in V1.5.4 is now mutually exclusive: on each summary request you must choose either mode=builtin (apply the built-in template) or mode=custom (your prompt fully replaces the built-in template).

Clients must migrate: V1.5.4 clients that do not update their fields will receive a 422.

Unified New Fields Across the Three Entry Points: REST POST /api/v1/summary, SSE regenerate/summary, and the WebSocket start action

Old (V1.5.4) -> New (V1.5.5) mapping:

Old FieldNew FieldNotes
template / templateSlug / summary_templateSame name (builtin mode only)Unchanged, but must not be sent in custom mode
custom_prompt / customPrompt / summary_custom_promptprompt / summary_prompt (custom mode only)Renamed
custom_prompt_slug / customPromptSlug / summary_custom_prompt_slugprompt_slug / summary_prompt_slug (custom mode only)Renamed
persist_custom_prompt / persistCustomPrompt(removed)Custom mode always snapshots; no opt-in
custom_instructions(removed)Legacy field, no longer supported
(none)mode / summary_mode (required)New required field, enum builtin / custom

Mutual-exclusion rules:

  • mode=builtin: template is required; prompt / prompt_slug must not be sent
  • mode=custom: prompt / prompt_slug is required; template must not be sent
  • Violations -> 422 summary_mode_field_mismatch

GET /api/v1/tasks/ Response Fields

Within data.tasks[]:

  • Added summary_mode (builtin / custom / null)
  • summary_template now returns the effective slug (in custom mode it returns your slug, identical to the prompt_slug you submitted)
  • Removed summary_custom_prompt_slug (merged into summary_template)

Backward compatibility: recordings without a generated summary have summary_mode set to null; existing builtin-mode recordings keep their original summary_template value.

Transcript Record Structure Changes

New top-level fields (not nested under the summary object):

FieldDescription
summary_modebuiltin / custom
summary_templateeffective slug — builtin -> the built-in slug; custom -> your slug
summary_plain_textbool
summary_prompt_snapshotPresent only in custom mode; the prompt content you passed in verbatim (not written in builtin mode)
summary_fallback_levelPresent only when a fallback was triggered (value 2 or 3); indicates that this summary went through an automatic content-filter fallback path. Omitted when the summary succeeds directly
summary_dropped_segmentsPresent only when fallback_level=3; the indices of the transcript segments that were dropped (an array of integers in original order)

In addition to the existing text, the init_summary event of GET /api/v1/sse/history/transcribe/{taskId} now adds mode / template / plain_text / prompt_snapshot (populated only in custom mode) for client traceability, plus fallback_level / dropped_segments (populated only when a fallback was triggered).

New Outbound WebSocket Events

  • summary_done: summary generation completed (includes summary_mode / summary_template (effective) / summary_plain_text / tokens_used / summary_fallback_level / summary_dropped_segments; does not include final_content)
  • summary_error: summary generation failed (includes error_code / message)

Clients no longer need to poll the transcript record to determine whether the summary is complete.

Automatic Content-Filter Fallback for Summaries

When a custom-mode prompt or transcript content is blocked by content filtering, the system handles it through an automatic multi-step fallback instead of failing outright. If some transcript segments still cannot be processed, they are omitted and reported via summary_dropped_segments. If even the fallback cannot produce a summary, a summary_error event is emitted with error_code=llm_content_filtered.

Client-side handling:

  • Use summary_fallback_level to show a UI notice indicating the summary was produced through a content-filter fallback path
  • Use summary_dropped_segments to inform the user which segments were actually omitted

Spec scope: In this release the fallback applies to two paths: WebSocket realtime summaries (auto-generated when a recording ends) and file-import summaries. Fallback integration for the SSE regenerate/summary endpoint is a follow-up; in the current version it still returns llm_content_filtered when blocked.

Custom-Mode Prompt Safety Rule (New in V1.5.5)

The built-in safety guard applied to custom-mode prompts now adds a rule instructing the LLM to "summarize the intent of any colloquial, emotional, or sensitive wording in the source in neutral, objective language, avoiding verbatim quotation or repetition." This rule is enforced by the backend and is not exposed for client configuration; its purpose is to reduce the chance of triggering the content filter on the first attempt.

The guidance itself is not retained. Your original prompt is still stored via the summary_prompt_snapshot field as an audit reference, complementing summary_fallback_level:

  • summary_prompt_snapshot = your intent (the original prompt content)
  • summary_fallback_level = the actual execution path taken by the automatic fallback

Prompt Safety in Custom Mode

Avoid concatenating untrusted end-user input directly into prompt.

New Error Codes

Error CodeHTTPTrigger Condition
summary_invalid_mode422 (SSE) / 400 (others)mode is not builtin / custom
summary_mode_field_mismatch422 / 400The mode and field combination is inconsistent (a required field is missing, or a forbidden field was sent)
summary_prompt_too_long422 / 400prompt exceeds 2000 characters
summary_prompt_slug_too_long422 / 400prompt_slug exceeds 64 characters
summary_prompt_slug_invalid422 / 400prompt_slug contains control characters (\n / \r / \t / \0, etc.)

Client Recommendations

  1. Add the required mode field — change existing calls using templateSlug=meeting to mode=builtin&template=meeting
  2. Rename fields — customPrompt -> prompt, customPromptSlug -> promptSlug; these two fields are only used in mode=custom
  3. Remove persistCustomPrompt — custom mode preserves the prompt content automatically
  4. Change templateSlug to template — and only use it in mode=builtin
  5. Transcript records now use top-level fields — no longer nested under the summary object
  6. Clients can determine whether a summary was saved from the done event / summary_done event — check persisted: true/false; you no longer need to infer it from the HTTP method

Reference


V1.5.4

2026-05-12

New Feature: Customer Prompt Customization for Summaries

Enterprise customers can now add their own rules to the summary API without modifying the built-in template. This release adds three orthogonal client parameters and splits the summary regeneration endpoint into "preview" and "save" verbs, avoiding the design gap of an HTTP GET with side effects.

Fully backward compatible — not sending the new fields = behavior identical to the previous version.

New Fields for POST /api/v1/summary

FieldTypeLimitDescription
custom_promptstring<=2000 charactersCustomer custom instructions appended after the built-in template
custom_prompt_slugstring<=64 characters, Unicode, no control charactersA client-defined template identifier (pass-through)
plain_textboolDefault falseRequest plain-text output
persist_custom_promptboolDefault falseOpt-in: whether the done event echoes the custom_prompt content

The SSE start / done events also add the corresponding fields (custom_prompt_slug, plain_text, final_content, custom_prompt_snapshot); see reference/rest/summary.md.

/api/v1/sse/regenerate/summary/{taskId} Split Into Two Endpoints

MethodPurposePersists ResultSaves TranscriptBilled
GETPreview (dry run, compare different prompt results)NoNoYes
POSTSave (official persistence)YesYes + bumps revisionYes

Client recommendation: If your integration previously relied on "the backend record updating automatically after a GET," switch to POST. GET is now a pure preview and no longer writes any backend state.

The done event adds a persisted: bool field, so clients can determine directly from the payload whether this call was saved, without inferring from the HTTP method.

Four New Fields for the WebSocket start Action

summary_custom_prompt / summary_custom_prompt_slug / summary_plain_text / summary_persist_custom_prompt, mapping one-to-one to the REST endpoint fields with the same limits.

New Endpoint: GET /api/v1/summary-templates/{slug}

Exposes the built-in template's full content so enterprise customers can reference the existing baseline when integrating and then decide what to add via custom_prompt.

GET /api/v1/summary-templates also adds a ?category=summary|medical|legal|all filter and a data[].category field in the response (default summary, backward compatible).

New Error Codes

Error CodeHTTPTrigger Condition
custom_prompt_too_long400custom_prompt exceeds 2000 characters
custom_prompt_slug_too_long400custom_prompt_slug exceeds 64 characters
custom_prompt_slug_invalid400custom_prompt_slug contains control characters
template_not_found404The template for the specified slug does not exist or is disabled
invalid_category400?category= is not in the allowlist

Behavior Changes

  • summary_text_empty / summary_text_too_long HTTP status code fix: these previously fell through to 500 because they were not explicitly mapped; this release fixes them to a semantically correct 400.
  • The POST /api/v1/summary error event details no longer includes the LLM raw error: the raw error goes only to the server log; the details returned to the client retains only the provider indicator.
  • GET preview is still billed: generating a preview consumes the same resources as saving one. Repeated GET calls are billed repeatedly, but they do not change any stored content.

Path and Field Naming Conventions

  • customPromptSlug is a customer-defined pass-through identifier (semantically different from the existing templateSlug, which is validated for existence). In naming terms, the former is "for client traceability" and the latter is "for looking up the VAS built-in template."
  • summary_custom_prompt_slug is recorded with each summary, so you can later query which customer template a summary corresponds to.
  • custom_prompt_snapshot (opt-in) is stored with the transcript record only when the customer sets persist_custom_prompt=true.

Security Controls

  • All endpoints require API Key authentication
  • The VAS server log does not log custom_prompt or the full transcript (it logs only the length and slug)
  • LLM error messages are sanitized (the raw error is not exposed to the client)
  • custom_prompt is fully isolated across tenants (session-scoped, no memory persistence)

Bug Fixes and Internal Improvements

  • The WebSocket start action recording_id field deprecation target version is unified to V2.0.0 (events.md previously said V1.6.0, inconsistent with code comments)
  • The SSE sse-api.md broken TOC anchor is fixed (it pointed to the audio section, but that content has been moved to the standalone reference/sse/audio.md)
  • Improved text sanitization so Chinese, Japanese and Korean characters and emoji are no longer mis-split or wrongly rejected
  • The summary regeneration full text now has a 100,000-character upper limit

Reference


V1.5.3

2026-05-07

Breaking Change: speaker_id Naming Inversion

To support speaker editing, V1.3.12 added the original_speaker_id field to preserve the original ID, but it left a design gap where "the same name means different things at different stages": for WebSocket realtime recording, speaker_id is the original ID (e.g., Guest-1), but after an SSE historical audio load, speaker_id becomes the display name (e.g., Manager Wang, with the alias applied). Frontends often picked the wrong field and passed it to PATCH /speakers/reassign.

This release performs a one-time inversion that is not backward compatible:

Old NameNew NameSemantics
speaker_id (display name)speaker_labelDisplay label (after alias is applied; mutable, human-readable)
original_speaker_id (original ID)speaker_idOriginal speaker ID (immutable, always stable)

After the inversion, speaker_id consistently refers to the original ID across all interfaces (WebSocket / SSE / REST); the new speaker_label represents the display label after the alias is applied. Speaker editing (rename / reassign / merge) always uses speaker_id as the locating key.

REST API Field Changes

PATCH /api/v1/tasks/{taskId}/speakers/rename

LocationOld FieldNew Field
Request bodyoriginal_namespeaker_id (max 100 characters)
Request bodynew_namenew_label (max 100 characters, no control characters \x00-\x1F / \x7F or newlines)
Response dataoriginal_namespeaker_id
Response datanew_namenew_label

speaker_id can still also accept a display label for chained renaming (e.g., first rename Guest-1 to "Manager Wang," then use "Manager Wang" to rename to "Director Wang"); the resolved response speaker_id is always the original ID.

PATCH /api/v1/tasks/{taskId}/speakers/reassign

LocationOld FieldNew Field
Request bodytarget_speaker_idUnchanged (semantics already aligned to the original ID)
Response datanew_speaker_namenew_speaker_label

target_speaker_id must be the original ID (taken from init_sentence.speaker_id); reassign does not accept a display label.

PATCH /api/v1/tasks/{taskId}/speakers/merge

LocationOld FieldNew Field
Request bodysource_speaker_id / target_speaker_idUnchanged (still accepts the original ID or the current display label)
Response datatarget_speaker_nametarget_speaker_label

WebSocket Event Changes

EventOld FieldNew Field
rename_speaker action bodyoriginal_name / new_namespeaker_id / new_label
result event origin / translations[lang]only speaker_id (mixed with display name)speaker_id (original ID) + speaker_label (display label)
speaker_renamed eventoriginal_name / new_namespeaker_id / new_label
speaker_reassigned eventnew_speaker_namenew_speaker_label
speakers_merged event(missing target label)added target_speaker_label

SSE Event Changes

EventOld FieldNew Field
init_sentencespeaker_id (display name) + original_speaker_id (original ID)speaker_id (original ID) + speaker_label (display label)
Broadcast viewer origin / translationonly speaker_id (mixed)speaker_id + speaker_label
Broadcast viewer speaker_renamed / speaker_reassigned / speakers_mergedsame as the corresponding WebSocket eventsas above

The behavior and fields of init_metadata.speaker_aliases (the "original ID -> display label" mapping) are unchanged.

Client Recommendations

  • Customers using WebSocket realtime recording: before upgrading, sync the handling of result.origin.speaker_id and the new result.origin.speaker_label; change the rename body to { "speaker_id": "...", "new_label": "..." }
  • Customers using SSE historical audio: init_sentence.speaker_id is now the original ID (previously the display name); switch to speaker_label for display
  • Customers doing speaker editing (rename / reassign / merge):
    • rename -> use speaker_id (either the original ID or the current display label) + new_label
    • reassign -> target_speaker_id must be the original ID (taken from init_sentence.speaker_id; you cannot send a display label)
    • merge -> source_speaker_id / target_speaker_id can still be the original ID or the current display label
  • Customers integrating TXT/SRT/CSV export: new_label now has control-character/newline validation; if you previously sent labels containing newlines, you will now receive a 422, so change to single-line content
  • Customers who do not do speaker editing and only consume transcript text: the impact is minimal; the only behavior difference is that if old code rendered speaker_id directly as the display name, it must switch to speaker_label

Data Compatibility

  • Not backward compatible: old transcript data (V1.3.12 ~ V1.5.1, containing speaker + original_speaker_id) requires a data conversion before it can be read in the new version; there is no cross-version data retention commitment during the POC phase
  • New recordings are unaffected: transcript blobs created after V1.5.3 use the new fields directly

Documentation Update

Reference


V1.5.1

2026-05-07

Bug Fix: POST /api/v1/imports Adds Length Validation for Terminology / Correction Fields

The length limits promised in several places in the documentation (e.g., a term's max of 100 characters) were previously not actually enforced on the file-import path, and overly long content was silently accepted. This release restores them, aligning behavior with the documentation's promises.

Behavior Changes (Aligning With Documented Promises)

POST /api/v1/imports adds 422 rejection conditions for the following fields (previously accepted):

FieldLimit
terminology.<lang>Array, max 500 terms (per language)
terminology.<lang>[].termstring, max 100 characters
terminology.<lang>[].boostnumeric, 0.5–5.0 (optional, default 1.0)
fuzzy_correction.<lang>[].correctstring, max 200 characters
fuzzy_correction.<lang>[].incorrect[]string, max 200 characters

These limits are consistent with the WebSocket config action; previously only the WebSocket path enforced them, and this release completes the file-import path.

Client Recommendations

If you previously sent overly long terms (>100 characters) via POST /api/v1/imports, you will now receive a 422. The frontend should check the length before submitting and prompt the user. The WebSocket path is unchanged.


V1.5.0

2026-05-07

No public changes

This release contains no public API changes, and customers need to take no action.

The old naming (recording_id) will be fully removed in V1.6.0; for the related client migration guidance, see V1.4.1 Client Recommendations


V1.4.3

2026-05-07

No public changes

This release contains no public API changes, and customers need to take no action.


V1.4.2

2026-05-07

No public changes

This release contains no public API changes, and customers need to take no action.


V1.4.1

2026-05-06

Naming Unification: task_id as the Cross-Interface Task Identifier

Previously, the same task had different field names across interfaces (WebSocket used recording_id, Webhook used task_id, and some REST path variables mixed {recordingId} / {taskId}), forcing integrators to reconcile the three naming schemes themselves. This release starts the naming-unification cycle; new integrations should use task_id consistently.

WebSocket Changes (Backward Compatible)

  • The session_started event payload now carries both task_id and recording_id, and their values are exactly the same (the UUID of the same recording)
  • The recording_id field is marked as Deprecated; it is still emitted normally and is scheduled for removal in V1.6.0
  • Documentation enhancement: session_id is the WS connection-level identifier (invalidated when the connection ends), which is a different level from task_id (the task identifier)

REST API Changes (Backward Compatible)

Added /api/v1/tasks/{taskId}/... alias paths that behave exactly the same as the existing /api/v1/recordings/{recordingId}/...:

Recommended (from V1.4.1)Deprecated (removed in V1.6.0)
PATCH /api/v1/tasks/{taskId}/speakers/renamePATCH /api/v1/recordings/{recordingId}/speakers/rename
PATCH /api/v1/tasks/{taskId}/speakers/reassignPATCH /api/v1/recordings/{recordingId}/speakers/reassign
PATCH /api/v1/tasks/{taskId}/entries/{sid}PATCH /api/v1/recordings/{recordingId}/entries/{sid}

Client Recommendations

  • New integrations: use the task_id field and the /api/v1/tasks/{taskId}/... paths consistently to avoid migrating again later
  • Existing integrations: no immediate change required. recording_id and /api/v1/recordings/... remain available throughout the V1.x period; we recommend migrating on your schedule, at the latest before V1.6.0 ships
  • ID alignment logic: if you depend on both WS and Webhook, you can align the WS task_id (or the old name recording_id) directly with the Webhook data.task_id; all three are the same UUID
  • Do not use session_id for alignment: session_id is meaningful only within the WS connection lifecycle and does not appear in Webhook or REST

Removal Timeline Announcement (V1.6.0)

V1.6.0 will remove the recording_id field from the WS payload and remove the /api/v1/recordings/{recordingId}/... paths. The detailed timeline will be announced separately before V1.6.0 ships.

Unchanged Items

  • Webhook payload: the existing data.task_id naming is unchanged
  • Existing /api/v1/tasks/{taskId}/... endpoints: unchanged

V1.4.0

2026-05-06

New Feature: Source-Text Editing for Historical Recordings + Automatic Retranslation

Users can correct STT recognition errors and regenerate translations; for the workflow, see Entries API Typical Workflow.

  • New endpoint PATCH /api/v1/recordings/{recordingId}/entries/{sid}: edit a single sentence's source text; on the first edit it automatically backs up the original STT output to original_text_raw, records original_text_edited_at, and clears the TTS cache for all languages of that sentence
  • New endpoint GET /api/v1/sse/recordings/{taskId}/entries/{sid}/retranslate: retranslate a single sentence (you can specify languages or retranslate all existing languages), with optimistic locking (expectedRevision)
  • Editing and retranslation are decoupled: PATCH only changes the source text and does not touch the translation; the frontend can decide when to trigger retranslation

Historical Record SSE Exposes Edit Markers

The historyTranscribe init_sentence event carries original_text_raw (the STT original) and original_text_edited_at on edited sentences, so the frontend can show an "edited" marker and a "restore original" function.

Security Fixes

  • retranslate / retranslateSummary add a user filter: these two existing SSE endpoints previously had a horizontal privilege vulnerability (IDOR) that allowed reading other users' recordings. This release adds the permission check; other users' recordings now return recording_not_found.
  • Retranslation / summary regeneration requires the recording to be completed: the four endpoints retranslate / retranslateSummary / retranslateEntry / regenerateSummary require processing_status === completed to avoid racing with the in-progress flow. When not completed, they return recording_not_completed.

New Error Codes

Error CodeHTTPDescription
recording_not_completed422The recording has not finished processing; retranslation / editing / summary regeneration is not allowed
entry_not_found404The specified sentence was not found
entry_text_empty422The sentence's source text is empty
entry_text_too_long422The sentence's source text exceeds the 2000-character limit
transcript_revision_conflict409The transcript has been modified by another request (optimistic-lock conflict)

See error-codes.md.

Client Recommendations

  • After editing the STT source text: we recommend triggering single-sentence retranslation SSE immediately after the PATCH, passing the revision from the PATCH response as expectedRevision to avoid concurrent overwrites
  • Showing the edit marker: determine whether a sentence has been edited by the presence of the original_text_raw field in the init_sentence event ('original_text_raw' in data); do not use text comparison (the user may edit and then change it back to the original value)
  • Recording status: calling retranslation / editing / summary regeneration on a recording that is not completed returns recording_not_completed; the frontend should block these operations in the UI until processing_status === completed

V1.3.13

2026-05-06

Behavior Changes (Breaking Changes)

  • WebSocket audio_format locked to pcm and webm: the previously accepted 5 formats (pcm / webm / mp3 / wav / m4a) are narrowed to accepting only pcm and webm, consistent with the existing spec in reference/websocket/voice-translation.md. Customers who send mp3 / wav / m4a will now receive audio_format_unsupported (previously these were silently decoded, which was undocumented implicit behavior). File imports still go through POST /api/v1/imports and are unaffected.

Documentation Update

  • Audio download Content-Type is always audio/mp4: rest-api / SSE audio / tasks export / history playback / curl / javascript documentation in several places is unified to "all recording audio is returned in an M4A container (AAC encoding)," removing the previous circular "dynamically determined" description.
  • Supported file-import formats narrowed to mp3 / wav / m4a: removed mentions of mp4 and webm from the documentation to align with the formats actually accepted (guides/file-import.md, reference/rest/imports.md).

Client Recommendations

  • Customers using the WebSocket start action: be sure to explicitly specify audio_format as pcm or webm; if you previously relied on the undocumented implicit mp3 / wav / m4a support (very rare scenarios), switch to the File Import API.
  • Customers downloading recording audio: all new recordings have Content-Type fixed to audio/mp4 with the .m4a extension. If older recordings still exist in storage, downloads may still return audio/webm; we recommend keeping a handling branch for the old extension to cover historical data.

Reference


V1.3.12

2026-05-04

Note: Inverted in V1.5.3: the original_speaker_id field and the "speaker_id is the display name" design introduced in this version have been superseded by the naming inversion in V1.5.3. This section is kept as a historical record; new integrations should refer directly to the V1.5.3 spec and do not need to implement this version's client recommendations.

New Feature

  • History SSE adds fields to align with the Transcribe speaker-editing UX: the historical record's init_metadata and init_sentence events each add a field, allowing the frontend to fully reuse the realtime recording page's speaker-editing menu (single-sentence reassignment + global rename).
    • init_metadata adds speaker_aliases (object): the "original speaker ID -> display name" mapping. When there are no aliases it is {} (an empty object, not an empty array). It lets the frontend perform a name-collision precheck before sending PATCH /speakers/rename, covering the implicit conflict of "an original ID that exists on the backend but does not appear on screen because it was renamed."
    • init_sentence adds original_speaker_id (string|null): the original speaker identifier without alias substitution, provided as the source for the target_speaker_id of PATCH /speakers/reassign.
    • Old-data fallback: if an older transcript record has no original_speaker_id, the output automatically falls back to speaker_id, preventing the new field from being null and disabling the editing entry point for old recordings.

Behavior Changes

  • No breaking change. Both fields are pure additions; clients that ignore unknown fields are unaffected, so no version negotiation is needed.

Documentation Update

  • sse-api.md L156-198: added the new field descriptions to the init_metadata / init_sentence examples and field tables
  • reference/sse/history.md L103-180: added the detailed reference schema accordingly

Client Recommendations

  • Customers doing speaker editing on the history detail page: get the original ID for reassign from init_sentence.original_speaker_id (do not use speaker_id, which is the display name with the alias applied); use init_metadata.speaker_aliases for the name-collision precheck before a rename.
  • Customers who do not do speaker editing: you can ignore the new fields; existing parsing behavior is unaffected.

Reference


V1.3.11

2026-05-04

Behavior Changes (Breaking Changes)

  • STT rejects the bare en code (the V1.3.10 changelog claimed it was removed, but it was not actually in effect): customers who send en will receive a 422 invalid_transcription_language; use a full BCP 47 code such as en-US / en-GB instead.
  • TTS removes 4 locales that were never usable: it-CH, ar-IL, ar-PS, en-GH. TTS for these 4 locales never actually worked; previously, requesting their voices would fail. STT still supports these 4 locales.

New Feature

  • TTS expanded to 154 languages and 325 voices
    • Chinese dialects (4 added): zh-CN-henan, zh-CN-guangxi, zh-CN-liaoning, zh-CN-shaanxi
    • South Asian languages (5 added): bn-BD Bengali (Bangladesh), ta-LK Tamil (Sri Lanka), ta-MY Tamil (Malaysia), ta-SG Tamil (Singapore), ur-PK Urdu (Pakistan)
    • Southeast Asian languages (1 added): su-ID Sundanese (Indonesia)
    • Eastern European languages (1 added): sr-Latn-RS Serbian (Latin script)
    • North American indigenous languages (2 added): iu-Cans-CA Inuktitut (Canadian syllabics), iu-Latn-CA Inuktitut (Canadian Latin script)

Documentation Update

  • languages.md TTS section rewritten, explicitly noting:
    • Of the 145 STT locales, 141 are supported on both the STT and TTS sides; 4 (it-CH, ar-IL, ar-PS, en-GH) are STT-only
    • Of the 154 TTS locales, 13 are TTS-only (4 zh-CN dialects + 9 other languages)
  • guides/tts.md numbers updated (142->154 languages, 304->325 voices)
  • Documentation home page TTS description updated

Supported Counts

STTTTS localeTTS voiceDiarization
14515432531

Client Recommendations

  • Customers using the en short code: switch to en-US or another full BCP 47 code.
  • Customers using it-CH/ar-IL/ar-PS/en-GH for TTS: these already failed on the provider side; switch to another locale in the same language family (e.g., it-CH -> it-IT, ar-IL -> ar-SA, en-GH -> en-NG). STT is unaffected.
  • Customers who want to use the 13 new TTS-only locales: you can call GET /api/v1/tts/voices?language=zh-CN-henan etc. directly to get the voice list.

Reference


V1.3.10

2026-04-30

Documentation Update

  • languages.md number corrections
    • Total speech-recognition languages 119 -> 145
    • Speech-translation support 117 -> 143 (145 minus jv-ID Javanese and wuu-CN Wu Chinese)

This version has a residual issue; see V1.3.11: this version claimed "the bare en was removed and the language counts are fully consistent at 145," but the bare "en" was not actually removed (still 146), nor did it handle the TTS-side it-CH/ar-IL/ar-PS/en-GH (not supported by the provider's TTS) or the 13 missing TTS-only locales. The full alignment fix was completed in V1.3.11.

Client Recommendations

  • This version is a documentation-only number correction and does not affect running integrations.

Reference


V1.3.9

2026-04-29

New Feature

  • Webhook Secret Bootstrap flow: resolves the contradiction where a client cannot obtain the secret on first webhook integration. The Dashboard adds a "Generate Webhook Secret" button (lazy generation), letting users obtain the secret and configure it on the receiving end first, then go back and set the webhook URL. The probe sent when setting the URL is signed with a secret that both sides agree on, so it passes on the first try.
    • New endpoint: POST /dashboard/api-keys/{id}/webhook/regenerate-secret (Dashboard only, rate limited to 10 requests/min/user)
    • Behavior: generates a 64-character random secret and stores it; does not send a probe and does not touch the webhook URL; the plaintext is returned once for the Dashboard to display
    • Regeneration impact: after execution, the old secret is invalidated immediately; existing receivers will get webhooks with mismatched signatures until they switch to the new secret

Behavior Changes

  • Clearing the Webhook URL no longer clears the Secret: when PATCH /dashboard/api-keys/{id}/webhook sets webhook_url to null, webhook_secret is left unchanged. The Secret and URL now have independent lifecycles. A customer can generate the secret first and set the URL later; sending an empty URL in the meantime will not lose the secret.
  • The Dashboard no longer returns webhook_secret in plaintext: GET /dashboard/api-keys/{id} now returns webhook_secret_masked (prefix mask + last 4 characters) and a has_webhook_secret boolean. The plaintext is shown only once right after generation.

Documentation Update

  • guides/webhook.md: "Method 2: API Key-level webhook_url" rewritten as a two-step flow (generate secret -> set URL); added a Webhook Secret Lifecycle section; added a Bootstrap callout to the security-verification section.

Client Recommendations

  • First integration: in the Dashboard, click "Generate Webhook Secret," copy it to the receiving end's .env, enable HMAC verification, and restart the service, then go back to the Dashboard and enter the webhook URL.
  • Existing customers: fully compatible, no changes needed. Existing webhook_url and webhook_secret behavior is unchanged.
  • Secret rotation: we recommend that the receiving end briefly accept both the old and new secrets; after the dashboard regeneration, remove the old secret once in-flight webhooks have finished processing.

V1.3.8

2026-04-27

New Feature

  • Translation-service-unavailable detection (session-level): added the error code translation_service_unavailable. When the LLM translation service fails consecutively up to a threshold, the backend emits a session-level error event once, so the frontend can show a global "translation temporarily unavailable" prompt instead of users seeing a page full of individual failed sentences in gray text.
    • Trigger conditions:
      • llm_timeout / llm_provider_error / llm_rate_limit / llm_request_failed escalate after 5 consecutive failures
      • llm_auth_failed / llm_deployment_not_found / llm_quota_exceeded escalate immediately after 1 occurrence (configuration/billing issues)
      • llm_content_filtered is not counted (a content issue, not a service issue)
    • Deduplication: each session is notified only once; any successful sentence translation resets the count and can trigger it again
    • payload: type: "error", severity: "error" (not fatal — should not disconnect), does not carry sid, details contains provider, last_error_code, fail_count
    • Viewer notification: in broadcast mode, all viewers (regardless of language) also receive this event (via the SSE event: error channel)

Documentation Update (Spec Sync)

Continuing the spec blind spots surfaced by frontend feedback since V1.3.7+, this pass completes:

  • error-codes.md — sentence-level error rule: added a sid-rule paragraph below the "Severity Levels" table, explicitly stating that "when an error carries sid, regardless of severity, it should be treated as a sentence-level error and should not disconnect." A fatal + sid combination only means that sentence failed severely; the session as a whole can still continue.
  • error-codes.md — translation_service_unavailable error-code registration: added this error code and its full trigger-rule description to the "Translation Service Errors" section
  • websocket-api.md: added a session-level translation error example (no sid, severity error) to the "Error Message Format" section
  • sse-api.md — retranslate section adds the per-sid error rule: explicitly lists the spec and payload format for "a failed sentence is re-emitted as event: error with sid + error_code, interleaved with translation" (implemented in V1.3.7 but documented only in the reference subdirectory)
  • reference/sse/broadcast-viewer.md: added a translation_service_unavailable example and a specific error-code entry
  • reference/websocket/events.md: removed an obsolete translation_error action that the service never actually emitted; translation errors are delivered on the standard error channel (type: "error")

Client Recommendations

  • Existing sentence-level error handling (type: "error" with sid) needs no changes.
  • If you want to show a global "translation service unavailable" prompt, add a listener: when you receive error_code === "translation_service_unavailable" (without sid), show a banner / toast; clear it once any subsequent sentence translation succeeds (you receive a translation event again).
  • Do not treat translation_service_unavailable as a disconnect signal — STT (the source text) continues to operate.

Reference:


V1.3.7

2026-04-24

Behavior Changes

  • Realtime recording: silent tasks now follow the normal completion flow: when a realtime recording (WebSocket) is silent throughout, is noise, or cannot recognize any sentence, it now still produces an empty transcript (entries: []) and ends with a task_complete event. This behavior aligns with the V1.3.5 file-import flow; the realtime and import sources now share the "zero recognition results is treated as a legitimate completed" semantics.
  • SSE historical record: no longer returns sse_transcript_not_found in silent scenarios: GET /api/v1/sse/history/transcribe/{taskId} no longer returns the sse_transcript_not_found error for silent tasks; instead it sends the full event sequence (init_metadata → init_summary(text='') → init_done(totalSentences=0)). Clients should use totalSentences === 0 to detect this and show a "no speech content" empty state.

Bug Fixes

  • Fixed the History page getting stuck on "processing" for silent recordings: previously, if a realtime recording was silent throughout, no transcript record was produced, but task_complete still sent the task_id, causing the frontend to receive sse_transcript_not_found (semantically "not finished processing") when loading the historical record, leaving the UI stuck on loading forever. After the fix, the realtime path matches the import path and always uploads the transcript (even with empty entries).

Client Recommendations

  • If you previously had "retry / polling" handling logic for sse_transcript_not_found, you may keep it as a defensive fallback (e.g., for blob upload delays), but you should no longer use it to determine "the task has no speech" — switch to init_done.totalSentences === 0.
  • We recommend the UI prompt possible reasons when totalSentences === 0 (volume too low, silent throughout, recognition language does not match the audio), consistent with the V1.3.5 import-scenario wording.

Documentation Update

  • Historical Record SSE adds a "Boundary scenario: no speech content" section and corrects the handling-recommendation description for sse_transcript_not_found

Reference:


V1.3.6

2026-04-23

New Feature

  • Tasks API: added POST /api/v1/tasks/{taskId}/force-fail: force-marks as failed a task stuck in a non-terminal state (recording / importing / uploading / pending / processing)
    • The body can optionally include reason (max 500 characters)
    • Triggers the recording.failed webhook, with payload.failure_source set to user_forced
    • A task already in a terminal state returns invalid_processing_status (422)
  • Tasks API: added POST /api/v1/tasks/{taskId}/retry: re-queues a task in the failed state for processing
    • Prerequisites: processing_status = failed and audio_status = success and transcript_status = success
    • Not meeting the prerequisites returns invalid_processing_status (422); the details field carries audio_status / transcript_status to help with diagnosis

Behavior Changes

  • Error code invalid_processing_status (422) expanded scope: now also used as the common response for force-fail and retry; details carries current_status, and the retry scenario additionally carries audio_status and transcript_status

Documentation Update

  • Tasks API adds documentation for the force-fail and retry endpoints
  • Error Code Reference adds the invalid_processing_status entry and a "Processing Status Mismatch" subsection

Reference:


V1.3.5

2026-04-22

Behavior Optimization

  • File import: empty recognition-result filtering: audio imports now filter out empty recognition results (caused by silence, very low volume, noise, or a language mismatch), so imports no longer produce empty 00:00 placeholder segments
  • Zero recognition results is a legitimate completed status: in this scenario the import task still ends with status: completed (not failed), the task_id is produced normally, but the subsequently loaded transcript entries is an empty array and segments_count is 0
  • Budget deducted by actual duration: unrecognizable audio is still deducted from the monthly budget based on the audio duration (no refund)

Client Recommendations

  • After loading the transcript (SSE /api/v1/sse/history/transcribe/{taskId}), if the cumulative sentence count is 0, show a "no speech content was recognized in this audio" empty state
  • Do not treat zero recognition results as an error branch; follow the completion branch and judge by the sentence count
  • We recommend the UI also prompt possible reasons (volume too low, silent throughout, recognition language does not match the audio)

Documentation Update

  • File Import Guide adds a "Behavior When Audio Cannot Be Recognized" section
  • Imports API adds a completed boundary-scenario note under the status transition
  • Import Progress SSE adds a behavior note for zero recognition results under the completed event

Reference:


V1.3.4

2026-04-22

New Feature

  • Tasks API: added GET /api/v1/tasks/{taskId}/transcript/export: download a task's transcript, supporting five formats — txt, srt, sbv, vtt, csv
    • The output includes the source text and all translation languages
    • CSV starts with a UTF-8 BOM, with columns index,start,end,speaker,text,<one column per translation language> and times in HH:MM:SS (no milliseconds)
    • SRT times are HH:MM:SS,mmm; SBV times are H:MM:SS.mmm, with the source text and translations joined into a single line with |; VTT uses the WEBVTT header
    • The filename uses {recording name}-transcript.{ext} (RFC 5987 UTF-8 encoded)
    • Added the error code recording_transcript_not_ready (422)

Behavior Changes (Breaking)

  • Speaker diarization and multi-language mutual exclusion is now a hard rejection: when recognition_mode: multi_speaker is combined with multiple transcription_languages, it previously emitted a warning and automatically truncated to the first language; it now directly returns the diarization_multilang_conflict error and refuses to start
    • The error severity is changed from warning to error
    • The frontend must restrict "speaker diarization" and "multi-language" to one or the other before the user submits start, or handle this error and guide the user to adjust the settings
    • Affected endpoint: WebSocket voice-translation / start

Documentation Update

  • Tasks API adds the full transcript/export spec and output examples for all five formats
  • Error Code Reference adds recording_transcript_not_ready
  • curl, Python, JavaScript examples add a "Task Export" section
  • Documentation home page API reference table: Tasks endpoint count updated from 8 to 9

Reference:


V1.3.3

2026-04-21

New Documentation

  • Tasks API: added the full documentation for the GET /api/v1/tasks/{taskId}/audio/export endpoint (the implementation existed but the documentation was missing), including parameters, dynamic Content-Type, error codes, and a frontend download example
  • Explained the difference between this endpoint and SSE /api/v1/sse/audio/{taskId}: the former is for offline download (Content-Disposition: attachment), the latter is for playback (supports Range Requests)

Documentation Fixes

  • Fixed the Voice Translation Actions translation-mode speakers field description table: the field name is corrected from speaker to id, consistent with the JSON example and the actual service behavior
  • Fixed the documentation home page API reference table endpoint counts: Tasks from 7 to 8 (added audio/export), Broadcasts from 9 to 6 (the original count was wrong)

Reference:


V1.3.2

2026-04-07

Documentation Structure Adjustment

  • Removed 3 deprecated old documents (error-codes.md V0.6, languages.md V0.1, authentication.md V0.1)
  • Moved appendix/error-codes.md and appendix/languages.md to the root directory, replacing the deprecated versions
  • Updated all cross-reference links

V1.3.1

2026-03-26

Batch Task Management

  • Added PUT /api/v1/tasks/batch/pin: batch-update pin status, max 100 per call
  • Added DELETE /api/v1/tasks/batch: batch-delete tasks, max 100 per call
  • Both endpoints affect only tasks belonging to the current user; the response includes affected_count

Batch Broadcast Cancellation

  • Added DELETE /api/v1/broadcasts/batch: batch-cancel broadcasts in the PENDING state, max 100 per call
  • IDs not in the PENDING state are ignored; the response includes affected_count

Reference:


Version: V1.24.1 Last Updated: 2026-10-07

Copyright © 2026