Imports API
Endpoint Overview
| Method | Endpoint | Description |
|---|---|---|
| POST | /api/v1/imports/check-quota | Check credit |
| POST | /api/v1/imports | Upload audio file |
| GET | /api/v1/imports/{importId} | Query import status |
| GET | /api/v1/imports | Get import list |
POST /api/v1/imports/check-quota
Description
Check whether the user has enough credit to upload an audio file of a given duration. We recommend calling this API for a pre-check before uploading, so you avoid discovering that the credit is insufficient only after uploading a large file.
Authentication
Header: X-API-Key (see Authentication)
Request Parameters
| Parameter | Location | Type | Required | Description |
|---|---|---|---|---|
duration_ms | body | integer | Yes | Audio duration (milliseconds; 1 second to 10 hours by default -- the actual bounds follow the deployment setting and match the duration check applied after upload) |
Request Example
curl -X POST "https://vas-poc.vurbo.ai/api/v1/imports/check-quota" \
-H "X-API-Key: vas_aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW" \
-H "Content-Type: application/json" \
-d '{"duration_ms": 3600000}'
Success Response
HTTP 200
{
"data": {
"allowed": true,
"reason": null,
"is_unlimited": false,
"remain_quota": 480.0,
"duration_minutes": 60,
"estimated_points": 60.0
}
}
Response Field Description
| Field | Type | Description |
|---|---|---|
data.allowed | boolean | Whether the upload is allowed (true when credit is sufficient or the plan allows it) |
data.reason | string | null | Why the upload is not allowed: null (allowed) / insufficient_credit (insufficient credit; topping up resolves it) / plan_not_allowed (the plan does not include audio import; a plan upgrade is required) / plan_daily_limit_reached (today's plan usage is exhausted and resets tomorrow; topping up does not help, added in v1.16.4) |
data.is_unlimited | boolean | Whether the account is unlimited (no point limit) |
data.remain_quota | float | null | Remaining credit points; null when unlimited |
data.duration_minutes | integer | Estimated audio duration (minutes, rounded up) |
data.estimated_points | float | Estimated points to be charged (STT-based estimate; actual charge also includes translation/diarization, finalized at upload) |
remain_quotasemantics change (v1.9.0): for API Keys with a dedicated credit allotment, this returns the credit actually available to that key (its dedicated allotment, not the account's total balance); accounts without dedicated allotments see no change.
Specific Error Codes
This endpoint has no specific error codes; it may only return common authentication errors. Every rejection is returned as HTTP 200 with allowed: false and a reason.
Preflight reason | What the actual upload returns |
|---|---|
plan_not_allowed | 403 plan_feature_not_allowed |
insufficient_credit | 402 stt_quota_exceeded |
plan_daily_limit_reached | 402 plan_daily_limit_reached (deliberately the same string, so you can map them 1:1) |
Note: The preflight is a prediction, not a guarantee. Between an allowed: true response and the actual upload, other uploads or the per-minute settlement of an ongoing recording may still consume the quota, and the upload will then be rejected.
POST /api/v1/imports
Description
Upload an audio file for speech recognition and translation. After a successful upload, processing runs in the background, and you can track progress through the Query import status API.
Authentication
Header: X-API-Key (see Authentication)
Request Parameters (multipart/form-data)
| Parameter | Type | Required | Description |
|---|---|---|---|
file | file | Yes | Audio file (mp3/wav/m4a, max 500MB). The format is determined from the file's actual content; when the extension does not match the content (for example a .mp3 name with WAV content), the file is processed according to its content |
transcription_languages | string | Yes | Transcription languages (JSON array string, e.g. '["zh-TW"]'). Up to 10, no duplicates |
translation_languages | string | No | Translation languages (JSON array string, e.g. '["en-US"]'). Up to 12, no duplicates |
recognition_mode | string | Yes | Recognition mode: single (single speaker) / multi_speaker (multiple speakers). multi_language or multi_channel returns 422 import_recognition_mode_unsupported |
summary_template | string | No | Summary template identifier (max 50 characters; must be an enabled summary category slug) |
summary_mode | string | No | Summary mode: builtin (default, uses summary_template) or custom (uses summary_prompt). Omitted = uses summary_template |
summary_prompt | string | No | Full custom prompt for custom mode (max 3000 chars, fully replaces the built-in template). Required for custom, prohibited otherwise |
summary_prompt_slug | string | No | Custom identifier for custom mode (max 64 chars, pass-through, not validated). Required for custom, prohibited otherwise |
terminology | string | No | Terminology list (JSON object string; see format below) |
fuzzy_correction | string | No | Fuzzy correction rules (JSON object string; see format below) |
translation_dict | string | No | Translation dictionary (JSON object string; see format below) |
callback_url | string | No | Webhook callback URL (notifies on completion/failure, max 2048 characters) |
Webhook Notification: Once
callback_urlis set, you receive animport.completedevent when the import finishes and animport.failedevent when it fails. See the Webhook Guide.
Text Processing Parameter Formats
Terminology (terminology): Improves recognition accuracy for specific terms
{
"zh-TW": [
{ "term": "Speaker Diarization" },
{ "term": "Real-time Transcription" }
]
}
- Use the language code as the key and the term array as the value
term: Term text (required, max 100 characters)- Up to 500 terms per language, and no more than 500 across all languages combined
- Up to 4000 fuzzy-correction rules per language, and no more than 4000 across all languages combined
The numbers above are defaults: the limit actually in force can be tuned per environment; always treat the 422 response message as authoritative.
Note: Both limits are enforced; exceeding either one returns 422 with the actual count. Plan multi-language vocabularies against the combined total.
Fuzzy Correction (fuzzy_correction): Corrects misspellings that sound different from the term
Usually no manual configuration is needed — misspellings that sound the same are already covered by terminology. It is needed only when the misspelling and the correct term sound different.
{
"zh-TW": [
{ "correct": "Speaker Diarization", "incorrect": ["Speaker Diorization", "Speaker Diarizaion"] }
]
}
- Use the language code as the key and the correction rule array as the value
correct: Correct term (required, max 200 characters)incorrect: a list of incorrect variants (conditionally required, each up to 200 characters)
Supplying only the correct term: when
correctis Chinese (contains Han characters),incorrectmay be omitted entirely — the system matches by pronunciation, and spellings in the transcript that sound the same or nearly the same are corrected back tocorrect.{ "fuzzy_correction": { "zh-TW": [{ "correct": "艾思通" }] } }No misspellings need to be listed above: 愛思通, 愛時通, 愛司東 and 愛似通 are all corrected. Only spellings that sound quite different (愛自動, say) or have a different number of syllables (愛松) still need to be listed in
incorrect.Note: Both conditions must hold: the language must be Chinese (
zh-TW,zh-CN,zh-HKand so on) andcorrectmust contain Han characters. Otherwiseincorrectremains required — omitting it in those cases would have no effect at all, and accepting it would leave you believing the setting took.
case_insensitive: Whether this rule's variants match regardless of case (optional, defaults tofalse= exact-case matching)
Case sensitivity:
case_insensitiveis optional and defaults tofalse(exact-case matching). When set totrue, everyincorrectvariant in that rule matches regardless of case. The flag is per rule — the samecorrectterm can be split across several rules with different settings, for example making variants that cannot collide with ordinary words case-insensitive while keeping variants that could hit a personal name exact. It has no effect on Chinese rules (Chinese has no letter case).
When the same incorrect variant appears in more than one rule: collisions are resolved on
incorrect(the variant), not oncorrect. The case flag resolves to strict wins (if any rule leavescase_insensitiveoff, that variant is matched with exact case). Splitting onecorrectterm across several rules is therefore safe, as long as theirincorrectvariants do not overlap.Note: Enabling it widens the false-positive surface: if
ivois case-insensitive, the personal nameIvois replaced too.
Translation Dictionary (translation_dict): Specifies how proper nouns are translated
{
"en-US": [{ "source": "Speaker Diarization", "target": "Speaker Diarization" }]
}
- Top-level key: the target language code
source: Source term (required, max 200 characters)target: The required translation for this language (required, max 200 characters)case_sensitive: Whether the entry applies only on an exact-case match (optional, defaults tofalse= case-insensitive)- Up to 3000 entries per language
The previous format is still supported: the earlier array-of-entries format is still accepted, with identical content and behavior.
Case-flag comparison: the case switches in
fuzzy_correctionandtranslation_dicthave opposite field names, and their default value produces opposite behavior —
Block Field Default Default behavior fuzzy_correctioncase_insensitivefalseStrict (case-sensitive) translation_dictcase_sensitivefalsePermissive (case-insensitive) Both default to
false, yet one means strict and the other means permissive. Do not share a single variable between them or mirror one onto the other — getting it wrong produces no error at all, only matching behavior opposite to what you intended.
Request Example
Basic Request
curl -X POST "https://vas-poc.vurbo.ai/api/v1/imports" \
-H "X-API-Key: vas_aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW" \
-F "file=@meeting.mp3" \
-F 'transcription_languages=["zh-TW"]' \
-F 'translation_languages=["en-US"]' \
-F "recognition_mode=multi_speaker"
Request with Text Processing Settings
curl -X POST "https://vas-poc.vurbo.ai/api/v1/imports" \
-H "X-API-Key: vas_aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW" \
-F "file=@meeting.mp3" \
-F 'transcription_languages=["zh-TW"]' \
-F 'translation_languages=["en-US"]' \
-F "recognition_mode=multi_speaker" \
-F 'terminology={"zh-TW": [{"term": "Speaker Diarization"}]}' \
-F 'fuzzy_correction={"zh-TW": [{"correct": "Speaker Diarization", "incorrect": ["Speaker Diorization"]}]}' \
-F 'translation_dict={"en-US": [{"source": "Speaker Diarization", "target": "Speaker Diarization"}]}'
Request with Webhook Callback
curl -X POST "https://vas-poc.vurbo.ai/api/v1/imports" \
-H "X-API-Key: vas_aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW" \
-F "file=@meeting.mp3" \
-F 'transcription_languages=["zh-TW"]' \
-F 'translation_languages=["en-US"]' \
-F "recognition_mode=multi_speaker" \
-F "callback_url=https://your-server.com/webhooks/vas"
Success Response
HTTP 202
{
"data": {
"import_id": "550e8400-e29b-41d4-a716-446655440000",
"status": "pending",
"stage": null,
"progress": 0,
"message": null,
"original_filename": "meeting.mp3",
"file_size": "15.2 MB",
"task_id": "660f9500-f30c-52e5-b827-557766550000",
"error_code": null,
"error_message": null,
"created_at": "2026-02-23T10:00:00.000Z",
"updated_at": "2026-02-23T10:00:00.000Z",
"downgraded_features": []
}
}
Response Field Description
| Field | Type | Description |
|---|---|---|
data.import_id | string | Import ID (UUID) |
data.status | string | Status: pending |
data.stage | string | null | Processing stage (initially null) |
data.progress | integer | Progress percentage (initially 0) |
data.message | string | null | Processing message |
data.original_filename | string | Original file name |
data.file_size | string | File size (formatted) |
data.task_id | string | Task ID (populated as soon as the upload succeeds; you can navigate to the task immediately) |
data.error_code | string | null | Error code (populated only on failure) |
data.error_message | string | null | Error message (populated only on failure) |
data.created_at | string | Creation time (ISO 8601) |
data.updated_at | string | Last update time (ISO 8601) |
data.downgraded_features | array | Sub-features skipped due to downgrade (v1.9.0): when the unlimited plan includes audio import but not some sub-features (such as speaker diarization speaker_diarization or translation translation), those sub-features are skipped and the import proceeds as usual; empty array = nothing downgraded |
Specific Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
import_file_too_large | 413 | File size exceeds the 500MB limit | Compress or split the file |
import_invalid_format | 415 | Unsupported audio format | Use mp3/wav/m4a format |
import_recognition_mode_unsupported | 422 | This recognition mode is not supported for file imports (multi_language, multi_channel). data.details carries field: "recognition_mode" and supportedModes: ["single", "multi_speaker"] | Use single or multi_speaker |
stt_quota_exceeded | 402 | Available credits are insufficient for this import's estimated charge | Top up and upload again |
plan_feature_not_allowed | 403 | The unlimited plan does not include audio import | Upgrade the plan; query GET /api/v1/me/plan for the plan contents |
plan_daily_limit_reached | 402 | The plan's daily usage limit has been reached | Upload again after the plan's reset (the next day) |
GET /api/v1/imports/{importId}
Description
Query the processing status and progress of a specific import task.
Once the task created by an import is deleted, the corresponding import record is removed as well, and this request returns 404 import_not_found.
Authentication
Header: X-API-Key (see Authentication)
Request Parameters
| Parameter | Location | Type | Required | Description |
|---|---|---|---|---|
importId | path | string | Yes | Import ID (UUID) |
Request Example
curl -X GET "https://vas-poc.vurbo.ai/api/v1/imports/550e8400-e29b-41d4-a716-446655440000" \
-H "X-API-Key: vas_aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW"
Success Response
HTTP 200
{
"data": {
"import_id": "550e8400-e29b-41d4-a716-446655440000",
"status": "processing",
"stage": "transcribing",
"progress": 45,
"message": "Recognizing speech...",
"original_filename": "meeting.mp3",
"file_size": "15.2 MB",
"task_id": "660f9500-f30c-52e5-b827-557766550000",
"error_code": null,
"error_message": null,
"created_at": "2026-02-23T10:00:00.000Z",
"updated_at": "2026-02-23T10:05:00.000Z"
}
}
Response Field Description
| Field | Type | Description |
|---|---|---|
data.status | string | Status: pending / processing / completed / failed |
data.stage | string | null | Processing stage: converting / transcribing / translating / summarizing |
data.progress | integer | Progress percentage (0-100) |
data.task_id | string | null | Task ID (can be used with the Tasks API). Populated from a successful upload onward; you do not need to wait for processing to finish |
data.error_code | string | null | Error code on failure |
data.error_message | string | null | Error message on failure (a general description without internal details) |
The remaining fields are the same as the POST /api/v1/imports response.
status Transitions
pending → processing → completed
→ failed
completedandfailedare final states and never change afterward: an import markedfailednever becomescompletedand is not charged.- Temporary errors during processing are retried automatically, and the status stays
processingwhile retrying; only when retries are exhausted is the import markedfailed, and theimport.failedWebhook is sent only once. - An import that does not start processing for a long time (stays
pending) is markedfailedwitherror_codePROCESSING_TIMEOUT. error_messageis a general description based on the error code, without internal details; provide theimport_idwhen you need troubleshooting. See Error Codes for the list of codes.
stage Processing Stages
| Stage | Description |
|---|---|
converting | Converting audio format |
transcribing | Recognizing speech |
translating | Translating |
summarizing | Generating summary |
Edge Case for the completed Status (v1.3.5)
When an audio file produces no recognizable speech content—due to silence, low volume, noise, or a recognition language that does not match the audio—the system still finishes with completed (not failed), and task_id is generated normally, but the transcript entries for the corresponding task are an empty array. The client should load the data via GET /api/v1/sse/history/transcribe/{taskId} and then decide whether to show an empty state based on the sentence count. See File Import Guide – Behavior When Audio Cannot Be Recognized.
Specific Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
import_not_found | 404 | Import task not found | Verify that importId is correct; if the task created by the import has been deleted, the import record is removed as well |
GET /api/v1/imports
Description
Get the user's list of import tasks (paginated). Import records whose tasks have been deleted do not appear in the list.
Authentication
Header: X-API-Key (see Authentication)
Request Parameters
| Parameter | Location | Type | Required | Description |
|---|---|---|---|---|
per_page | query | integer | No | Items per page (default 20) |
Request Example
curl -X GET "https://vas-poc.vurbo.ai/api/v1/imports?per_page=20" \
-H "X-API-Key: vas_aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW"
Success Response
HTTP 200
{
"data": [
{
"import_id": "550e8400-e29b-41d4-a716-446655440000",
"status": "completed",
"stage": null,
"progress": 100,
"message": null,
"original_filename": "meeting.mp3",
"file_size": "15.2 MB",
"task_id": "660f9500-f30c-52e5-b827-557766550000",
"error_code": null,
"error_message": null,
"created_at": "2026-02-23T10:00:00.000Z",
"updated_at": "2026-02-23T10:15:00.000Z"
}
],
"meta": {
"current_page": 1,
"last_page": 3,
"per_page": 20,
"total": 55
}
}
Response Field Description
| Field | Type | Description |
|---|---|---|
data | array | List of import tasks (each field is the same as the Query import status response) |
meta.current_page | integer | Current page number |
meta.last_page | integer | Last page number |
meta.per_page | integer | Items per page |
meta.total | integer | Total number of items |
Specific Error Codes
This endpoint has no specific error codes; it may only return common authentication errors.
Version: V1.24.1 Last Updated: 2026-09-28