Recording Speaker API
Overview
The recording speaker editing API is used to manage speakers in multi-speaker conversation mode. It provides three operations: global rename, single-sentence reassignment, and speaker merge.
All endpoints use
/api/v1/tasks/{taskId}/speakers/...as the sole official path.
PATCH /api/v1/tasks/{taskId}/speakers/rename
Description
Renames all sentences of a given speaker in a recording at once. This is useful for replacing system-generated speaker names (such as Guest-1) with real names.
Authentication
Header: X-API-Key (see Authentication)
Request Parameters
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
taskId | string | Yes | Task ID (UUID) |
Body Parameters (JSON)
| Parameter | Type | Required | Description |
|---|---|---|---|
speaker_id | string | Yes | Original speaker ID (such as "Guest-1"); also accepts the current display label (speaker_label) for consecutive renaming; maximum 100 characters |
new_label | string | Yes | New display label; maximum 100 characters, must not contain control characters (\x00-\x1F, \x7F) or line breaks (it will be written into the transcript, SSE events, and TXT/SRT/CSV exports) |
Request Example
curl -X PATCH "https://vas-poc.vurbo.ai/api/v1/tasks/rec_abc123/speakers/rename" \
-H "X-API-Key: vas_aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW" \
-H "Content-Type: application/json" \
-d '{
"speaker_id": "Guest-1",
"new_label": "Manager Wang"
}'
Success Response
HTTP 200
{
"data": {
"speaker_id": "Guest-1",
"new_label": "Manager Wang",
"affected_sids": [1, 3, 5]
}
}
Response Field Descriptions
| Field | Type | Description |
|---|---|---|
data.speaker_id | string | The resolved original speaker ID (even if the request sent a display label, the response is still the original ID) |
data.new_label | string | New display label |
data.affected_sids | array<int> | List of affected sentence SIDs |
Specific Error Codes
| Error Code | HTTP Status | Description | Suggested Action |
|---|---|---|---|
recording_not_found | 404 | Recording not found | Verify that taskId is correct |
validation_failed | 422 | Request validation failed | Verify that both speaker_id and new_label are provided, do not exceed 100 characters, and that new_label contains no control characters |
transcript_revision_conflict | 409 | Another write to the same transcript is in progress | Retry shortly; this change did not take effect |
storage_upload_failed | 500 | Failed to write the transcript back to storage | Retry shortly; this change did not take effect |
speaker_transcript_not_found | 404 | Transcript not found | Confirm that transcription has completed |
speaker_diarization_required | 422 | This operation is only supported on speaker-diarization recordings | Applies to multi-speaker recordings only |
speaker_name_empty | 422 | new_label is empty | Provide a valid new_label |
speaker_name_duplicate | 422 | That name is already used by another speaker (including another speaker's current display name and original speaker ID), or the display name you passed maps to more than one speaker | Use a name that is not already taken; when the input is ambiguous, pass the original speaker ID instead |
speaker_not_found | 422 | The specified speaker was not found | Confirm that speaker_id exists |
PATCH /api/v1/tasks/{taskId}/speakers/reassign
Description
Reassigns the speaker of a specific sentence to another speaker. This is useful for correcting errors produced by automatic recognition (Speaker Diarization).
Authentication
Header: X-API-Key (see Authentication)
Request Parameters
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
taskId | string | Yes | Task ID (UUID) |
Body Parameters (JSON)
| Parameter | Type | Required | Description |
|---|---|---|---|
sid | integer | Yes | Sentence ID |
target_speaker_id | string | Yes | Target speaker's original ID (taken from init_sentence.speaker_id; reassign does not accept display labels, you must send the original ID); maximum 100 characters |
Request Example
curl -X PATCH "https://vas-poc.vurbo.ai/api/v1/tasks/rec_abc123/speakers/reassign" \
-H "X-API-Key: vas_aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW" \
-H "Content-Type: application/json" \
-d '{
"sid": 3,
"target_speaker_id": "Guest-2"
}'
Success Response
HTTP 200
{
"data": {
"sid": 3,
"old_speaker_id": "Guest-1",
"new_speaker_id": "Guest-2",
"new_speaker_label": "Director Li"
}
}
Response Field Descriptions
| Field | Type | Description |
|---|---|---|
data.sid | integer | The modified sentence ID |
data.old_speaker_id | string | Original speaker ID |
data.new_speaker_id | string | Original speaker ID after reassignment |
data.new_speaker_label | string | Display label after reassignment (after applying speaker_aliases; equals new_speaker_id when there is no alias) |
Specific Error Codes
| Error Code | HTTP Status | Description | Suggested Action |
|---|---|---|---|
recording_not_found | 404 | Recording not found | Verify that taskId is correct |
validation_failed | 422 | Request validation failed | Verify that both sid and target_speaker_id are provided and correctly formatted |
transcript_revision_conflict | 409 | Another write to the same transcript is in progress | Retry shortly; this change did not take effect |
storage_upload_failed | 500 | Failed to write the transcript back to storage | Retry shortly; this change did not take effect |
speaker_transcript_not_found | 404 | Transcript not found | Confirm that transcription has completed |
speaker_op_not_allowed_multi_channel | 422 | Multi-channel recordings do not support this speaker operation (details.recognition_mode carries the mode) | Speakers in multi-channel recordings are determined by channel; adjust the channel settings instead |
speaker_diarization_required | 422 | This operation is only supported on speaker-diarization recordings | Applies to multi-speaker recordings only |
speaker_sid_not_found | 422 | The specified sentence was not found | Confirm that sid exists in this recording |
speaker_not_found | 422 | The specified speaker was not found | Confirm that target_speaker_id exists |
PATCH /api/v1/tasks/{taskId}/speakers/merge
Aligned with the WebSocket
merge_speakersaction, providing merge capability for historical recordings.
Description
Reassigns all sentences of the source speaker to the target speaker; the source's alias (if any) is transferred to the target (if the target has no alias yet). This is useful when the diarization model misidentifies the same person as two separate speakers (for example, Guest-1 and Guest-2 are actually the same person).
vs. reassign:
reassignchanges only a single sentence;mergechanges all sentences of the speaker. vs. rename:renamechanges only the display name (alias) without touching the speaker ID;mergeconsolidates multiple speakers into one.
Authentication
Header: X-API-Key (see Authentication)
Request Parameters
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
taskId | string | Yes | Task ID (UUID) |
Body Parameters (JSON)
| Parameter | Type | Required | Description |
|---|---|---|---|
source_speaker_id | string | Yes | The original speaker ID to be merged, or the current display label (such as Guest-2 or Manager Wang); maximum 100 characters |
target_speaker_id | string | Yes | The original ID or current display label of the merge target speaker (such as Guest-1); maximum 100 characters |
Request Example
curl -X PATCH "https://vas-poc.vurbo.ai/api/v1/tasks/rec_abc123/speakers/merge" \
-H "X-API-Key: vas_aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW" \
-H "Content-Type: application/json" \
-d '{
"source_speaker_id": "Guest-2",
"target_speaker_id": "Guest-1"
}'
Success Response
HTTP 200
{
"data": {
"source_speaker_id": "Guest-2",
"target_speaker_id": "Guest-1",
"target_speaker_label": "Manager Wang",
"affected_sids": [3, 5, 7]
}
}
Response Field Descriptions
| Field | Type | Description |
|---|---|---|
data.source_speaker_id | string | The original speaker ID that was merged (resolved back to the original ID, even if a display label was sent in the request) |
data.target_speaker_id | string | The original speaker ID of the merge target |
data.target_speaker_label | string | The target speaker's display label (after applying speaker_aliases; equals the original ID when there is no alias) |
data.affected_sids | array<int> | List of affected sentence SIDs: sentences that belonged to the source speaker, plus the target speaker's existing sentences whose display name changed because of the merge (for example, when the source speaker's custom name is carried over to the target) |
Specific Error Codes
| Error Code | HTTP Status | Description | Suggested Action |
|---|---|---|---|
merge_speakers_same_id | 400 | source and target resolve to the same speaker | Provide different speaker IDs |
speaker_name_empty | 422 | source or target is an empty string | Provide a valid speaker ID |
speaker_not_found | 422 | source or target does not exist in this recording | Verify that the speaker ID is correct |
recording_not_found | 404 | Recording not found | Verify that taskId is correct |
speaker_diarization_required | 422 | This recording is not in multi-speaker conversation mode | This feature is only available for recognition_mode: multi_speaker recordings |
validation_failed | 422 | Request validation failed | Verify that both source_speaker_id and target_speaker_id are provided and do not exceed 100 characters |
transcript_revision_conflict | 409 | Another write to the same transcript is in progress | Retry shortly; this change did not take effect |
storage_upload_failed | 500 | Failed to write the transcript back to storage | Retry shortly; this change did not take effect |
Notes
- Display name input supported: A renamed speaker (for example, Guest-1 renamed to "Manager Wang") can be sent as "Manager Wang" directly for the source or target, and the system resolves it back to the original ID. The display name each channel was given at the start of a multi-channel recording can also be passed directly
- Alias transfer: If the source has an alias but the target does not, after the merge the target inherits the source's alias; the target's existing sentences switch to that name as well and are included in
affected_sids - Concurrency protection: Each merge updates the
revision. This endpoint has noexpected_revisionparameter; if another change to the same transcript is in progress it returnstranscript_revision_conflict(409) — simply retry later - Irreversible: After a merge, the source ID no longer has any corresponding sentences in this recording; to revert, you must use
reassignto change them back one sentence at a time
Related Resources
- Speaker Management - Complete guide to renaming, reassigning, and merging speakers
Version: V1.24.1 Last Updated: 2026-09-28