Speaker Management Guide
Table of Contents
- Overview
- Enabling Speaker Recognition
- Receiving Speaker Information
- Renaming a Speaker
- Reassigning a Speaker
- Merging Speakers
- Speaker Management for Multi-Channel Recordings
- Realtime Mode vs. Offline Mode
- Best Practices
- Related Reference Documents
Overview
VAS's Speaker Diarization feature automatically identifies different speakers in multi-party conversations and tags each sentence with the speaker's identity. The system supports speaker recognition in 31 languages.
Core Features
| Feature | Description | API Type |
|---|---|---|
| Speaker recognition | Automatically identifies and distinguishes different speakers | WebSocket |
| Rename speaker | Changes Guest-1 to a real name | WebSocket / REST |
| Reassign | Corrects the speaker identity of a single sentence | WebSocket / REST |
| Merge speakers | Merges the same speaker that was mistakenly recognized as multiple people | WebSocket / REST |
Use Cases
- Meeting minutes: Automatically distinguish between participants' remarks
- Interview transcription: Tag the host and the interviewee
- Conversation records: Identify two-party or multi-party conversations
Authentication
All speaker management REST APIs require API Key authentication. See Authentication for details.
Enabling Speaker Recognition
To use the speaker recognition feature, set the following parameters in the WebSocket start action:
{
"type": "voice-translation",
"data": {
"action": "start",
"transcription_languages": ["zh-TW"],
"translation_languages": ["en-US"],
"type": "conversation",
"recognition_mode": "multi_speaker",
"audio_format": "pcm"
}
}
Key Parameters
| Parameter | Value | Description |
|---|---|---|
type | conversation | Use the conversation record type |
recognition_mode | multi_speaker | Enable multi-party speaker recognition |
Note:
typecan also be set totranscribeorbroadcast. Speaker recognition is enabled as long asrecognition_modeis set tomulti_speaker.Restriction: In
multi_speakermode,transcription_languagesmust contain exactly 1 language. If you provide multiple languages, you will receive adiarization_multilang_conflicterror and the session will be refused. You must switch to a single language or disable speaker diarization. Two-way translation (type=conversation) has been exempt from this restriction since v1.7.2 — it acceptsspeaker_diarizationbut ignores it.
Successful Response
After starting successfully, you will receive a session_started event confirming that the recognition mode is multi_speaker:
{
"type": "voice-translation",
"data": {
"action": "session_started",
"session_id": "550e8400-e29b-41d4-a716-446655440000",
"task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
"recording_type": "conversation",
"recognition_mode": "multi_speaker",
"message": "Speech recognition started"
}
}
Receiving Speaker Information
Once multi-party speaker recognition is enabled, every recognition result (the result event) includes speaker information.
Recognition Result Format
{
"type": "voice-translation",
"data": {
"action": "result",
"origin": {
"sid": 1,
"language": "zh-TW",
"text": "Today's meeting mainly discusses the project progress",
"is_final": true,
"speaker_id": "Guest-1",
"detected_language": "zh-TW",
"start_time": "00:05"
}
}
}
Speaker-Related Fields
| Field | Type | Description |
|---|---|---|
speaker_id | string | Speaker ID (automatically assigned by the system, e.g., Guest-1) |
sid | int | Sentence number. Each event carries one number; the transcript may contain more than one entry with the same number, so use the first |
is_final | boolean | Whether this is the final result |
Speaker ID Naming Rules
- The system automatically assigns IDs in the format
Guest-{N} - IDs are unique within a recording, but are not guaranteed to start at 1 or to be consecutive
- After renaming, subsequent recognition results use the new name
- If the connection drops and recovers mid-recording, speakers are identified again from scratch: the
same person may receive a new ID afterwards, and will not carry over the name set earlier. When
you are sure it is the same person, use
merge_speakersto merge them — when the target has no name yet, merging carries the source's name across
Renaming a Speaker
Change a system-assigned speaker ID (such as Guest-1) to a meaningful name (such as Manager Wang). Renaming is a global operation; all sentences that use that speaker ID are updated simultaneously.
Method 1: WebSocket (Realtime Mode)
For realtime renaming while a recording is in progress.
{
"type": "voice-translation",
"data": {
"action": "rename_speaker",
"speaker_id": "Guest-1",
"new_label": "Manager Wang"
}
}
Successful response:
{
"type": "voice-translation",
"data": {
"action": "speaker_renamed",
"speaker_id": "Guest-1",
"new_label": "Manager Wang",
"affected_sids": [1, 3, 5, 8]
}
}
affected_sids lists all affected sentence numbers, so the frontend can update the UI based on this information.
Method 2: REST API (Offline Mode)
For offline editing after a recording has ended.
curl -X PATCH "https://vas-poc.vurbo.ai/api/v1/tasks/{taskId}/speakers/rename" \
-H "X-API-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"speaker_id": "Guest-1",
"new_label": "Manager Wang"
}'
Successful response (HTTP 200):
{
"data": {
"speaker_id": "Guest-1",
"new_label": "Manager Wang",
"affected_sids": [1, 3, 5, 8, 12]
}
}
Renaming Restrictions
speaker_idmust be an original speaker ID currently present in the recording or its current display label; if it still cannot be resolved,speaker_not_foundis returnednew_labelcannot be empty, has a maximum of 100 characters, and must not contain control characters (\x00-\x1F,\x7F) or newlines- The new label cannot duplicate another speaker's current display name or original speaker ID (a
speaker_name_duplicateerror is returned) - When you identify the speaker by display name and that name maps to more than one speaker,
speaker_name_duplicateis also returned; pass the original speaker ID instead - The REST API applies to recordings in both
multi_speakerandmulti_channelmode (for multi-channel behavior, see Speaker Management for Multi-Channel Recordings)
Reassigning a Speaker
Change the speaker identity of a single sentence, assigning the sentence to another existing speaker. This is useful for correcting speaker recognition errors.
Method 1: WebSocket (Realtime Mode)
{
"type": "voice-translation",
"data": {
"action": "reassign_speaker",
"sid": 5,
"target_speaker_id": "Guest-2"
}
}
Successful response:
{
"type": "voice-translation",
"data": {
"action": "speaker_reassigned",
"sid": 5,
"old_speaker_id": "Guest-1",
"new_speaker_id": "Guest-2",
"new_speaker_label": "Lisa Lee"
}
}
Method 2: REST API (Offline Mode)
curl -X PATCH "https://vas-poc.vurbo.ai/api/v1/tasks/{taskId}/speakers/reassign" \
-H "X-API-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"sid": 5,
"target_speaker_id": "Guest-2"
}'
Reassignment Restrictions
target_speaker_idmust be the original ID of an existing speaker (creating a new speaker is not supported, and display labels are not accepted)- If that speaker has been renamed,
new_speaker_labelreflects the display label after applyingspeaker_aliases - Reassignment does not apply to multi-channel (
multi_channel) recordings; the REST API returns aspeaker_op_not_allowed_multi_channelerror (see Speaker Management for Multi-Channel Recordings)
Merging Speakers
Merge all sentences of one speaker into another speaker. This is useful when the system mistakenly recognizes the same person's voice as multiple speakers.
Use Case
The speech recognition engine sometimes recognizes the same person's voice at different times as different speakers (for example, Guest-1 and Guest-3 are actually the same person). After merging:
- All
Guest-3sentences are attributed toGuest-1 - In WebSocket mode: future recognition results identified as
Guest-3are also automatically converted toGuest-1(continuous interception) — but only for the current recognition pass. If the connection drops and recovers mid-recording, speakers are identified again and receive new IDs, and the merge does not carry over; merge again once you confirm it is the same person - In REST mode: historical recordings have no new sentences, so only the existing sentences are merged once
Method 1: WebSocket (Realtime Mode)
{
"type": "voice-translation",
"data": {
"action": "merge_speakers",
"source_speaker_id": "Guest-3",
"target_speaker_id": "Guest-1"
}
}
Successful response:
{
"type": "voice-translation",
"data": {
"action": "speakers_merged",
"source_speaker_id": "Guest-3",
"target_speaker_id": "Guest-1",
"affected_sids": [3, 5, 7]
}
}
Method 2: REST API (Offline Mode)
curl -X PATCH "https://vas-poc.vurbo.ai/api/v1/tasks/{taskId}/speakers/merge" \
-H "X-API-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"source_speaker_id": "Guest-3",
"target_speaker_id": "Guest-1"
}'
Successful response (HTTP 200):
{
"data": {
"source_speaker_id": "Guest-3",
"target_speaker_id": "Guest-1",
"target_speaker_label": "Manager Wang",
"affected_sids": [3, 5, 7]
}
}
Merge vs. Reassign Comparison
| Feature | Scope | Affects Future Recognition Results (WS) |
|---|---|---|
reassign_speaker | A single sentence (1 SID) | No |
merge_speakers | All sentences of the speaker | Yes (future occurrences of the source are automatically converted to the target, within the current recognition pass only) |
Merge Restrictions
source_speaker_idandtarget_speaker_idcannot be the same (amerge_speakers_same_iderror is returned)- Both speaker IDs must exist in the recording
- REST mode applies only to recordings with
recognition_mode: multi_speaker - Merging does not apply to multi-channel (
multi_channel) recordings; the REST API returns aspeaker_op_not_allowed_multi_channelerror (see Speaker Management for Multi-Channel Recordings)
Speaker Management for Multi-Channel Recordings
In multi-channel mode (recognition_mode: multi_channel), each physical microphone occupies its own audio channel, and speaker identity is determined by the channel — the "channel ↔ speaker" mapping is a fixed 1:1 correspondence. Speaker management therefore behaves differently from multi_speaker.
speaker_id Format and Immutability
- Multi-channel speaker IDs use the fixed format
channel_{N}(where N is the channel'schannel_id, e.g.,channel_1) and are immutable: the same channel uses the same speaker ID for the entire recording, even if the channel's language is changed mid-recording viaset_channel_language - The
originof theresultevent carries bothchannel_idandspeaker_id; each sentence in the transcript also carrieschannel_id(the physical channel number, which never changes as a result of any speaker operation) - Display priority of
speaker_label: the alias set byrename_speaker→ thespeaker_namespecified for that channel atstart→ thespeaker_iditself
rename_speaker: Available
Renaming works as usual in multi-channel mode (WebSocket and REST usage are the same as in the sections above), with one added convenience: multi-channel registers the speaker IDs of all channels at start, so a channel can be renamed before anyone on it has spoken. For example, before a meeting begins, you can rename channel_1 to the actual participant's name without waiting for them to speak first.
{
"type": "voice-translation",
"data": {
"action": "rename_speaker",
"speaker_id": "channel_1",
"new_label": "Manager Wang"
}
}
Tip: You can also specify each channel's initial display name directly via
speaker_namein thechannels[]ofstart, then adjust it later withrename_speakeras needed.
reassign_speaker / merge_speakers: Not Applicable
In multi-channel mode, speaker identity is determined by the physical channel rather than inferred by the system:
- Reassign: Attributing a sentence to a different channel would contradict what the physical microphone actually captured
- Merge: Merging two channels into a single speaker would break the fixed 1:1 "channel ↔ speaker" correspondence
These two operations are therefore not applicable to multi-channel recordings. Calling them through the REST API returns HTTP 422 with the error code speaker_op_not_allowed_multi_channel (details.recognition_mode carries the recording's recognition mode). Recordings in multi_speaker mode are unaffected — all three operations remain available.
Realtime Mode vs. Offline Mode
Speaker management offers two usage modes. The following is a complete comparison:
| Operation | Realtime Mode (WebSocket) | Offline Mode (REST API) |
|---|---|---|
| Rename speaker | rename_speaker action | PATCH /api/v1/tasks/{taskId}/speakers/rename |
| Reassign | reassign_speaker action | PATCH /api/v1/tasks/{taskId}/speakers/reassign |
| Merge speakers | merge_speakers action | PATCH /api/v1/tasks/{taskId}/speakers/merge |
| When to use | While recording is in progress | After recording has ended |
| Broadcast sync | Automatically pushed to SSE viewers | Not applicable |
REST vs. WebSocket merge difference: Both merge existing sentences; however, the WebSocket version additionally creates a mapping that "automatically converts future source IDs to the target" — within the current recognition pass only, since a connection recovery re-identifies speakers. This does not apply to historical recordings (which have no new sentences).
Multi-channel restriction: Recordings in
multi_channelmode support renaming only; reassignment and merging are not applicable. See Speaker Management for Multi-Channel Recordings.
Speaker Management in Broadcast Mode
In broadcast mode, speaker management operations are automatically synced to SSE viewers:
| WebSocket Operation | SSE Event Received by Viewers |
|---|---|
rename_speaker | speaker_renamed |
reassign_speaker | speaker_reassigned |
merge_speakers | speakers_merged |
Viewers can update their UI in real time based on these events:
eventSource.addEventListener('speaker_renamed', (e) => {
const data = JSON.parse(e.data);
// Update the display labels of all affected_sids
data.affected_sids.forEach(sid => {
updateSpeakerLabel(sid, data.new_label);
});
});
eventSource.addEventListener('speaker_reassigned', (e) => {
const data = JSON.parse(e.data);
// Update the speaker of a single sentence (speaker_id is the original ID, speaker_label is the display label)
updateSpeakerForSentence(data.sid, data.new_speaker_id, data.new_speaker_label);
});
eventSource.addEventListener('speakers_merged', (e) => {
const data = JSON.parse(e.data);
// Update the display labels of all affected sentences
data.affected_sids.forEach(sid => {
updateSpeakerLabel(sid, data.target_speaker_label);
});
});
Best Practices
1. Recognize First, Then Name
Let the system recognize the different speakers first (Guest-1, Guest-2, ...), and rename them only after confirming that recognition is stable.
2. Make Good Use of the Merge Feature
If you find that the same person has been recognized as multiple speakers (for example, they left and came back midway), using merge_speakers is more efficient than reassigning sentence by sentence with reassign_speaker, and it can also affect future recognition results.
3. Offline Editing for Correction
After a recording ends, perform final corrections on the transcript through the REST API to ensure that the speaker tags of all sentences are correct.
4. Error Handling
| Error Code | Description | Suggested Action |
|---|---|---|
speaker_not_found | The specified speaker was not found | Confirm that the speaker ID exists |
speaker_name_empty | The name cannot be empty | Provide a valid name |
speaker_name_duplicate | The name is already in use | Use a different name |
speaker_sid_not_found | The specified sentence was not found | Confirm that the SID exists |
speaker_diarization_required | Only diarization recordings are supported | Confirm that multi_speaker mode is used |
merge_speakers_same_id | Source and target are the same | Use different speaker IDs |
speaker_op_not_allowed_multi_channel | This speaker operation is not supported for multi-channel recordings | Multi-channel supports rename_speaker only; reassignment and merging are not available |
Related Reference Documents
- REST API - Recording Speaker Editing
- WebSocket - Voice Translation (rename_speaker / reassign_speaker / merge_speakers)
- SSE - Broadcast Viewer (speaker_renamed / speaker_reassigned / speakers_merged events)
Version: V1.24.1 Last Updated: 2026-09-28