Transcription
Turn speech into text with the whisper engine. You need at least one Whisper model installed:
| Model | Size | Best for |
|---|---|---|
whisper-base |
145 MB + 260 MB engine | Live dictation, quick drafts |
whisper-small |
484 MB | Most files: better with accents and names |
whisper-large-v3-turbo |
1.6 GB | Hard audio, interviews, lectures |
All models understand about 99 languages. Whisper runs in its own isolated runtime; on an NVIDIA GPU it uses CUDA, and elsewhere a fast int8 CPU path. (CTranslate2, the library underneath, has no Apple-GPU backend.)
Transcript object
| Field | Type | Description |
|---|---|---|
id, title |
string | |
source |
string | file or live |
language |
string | Detected or requested ISO code, e.g. en |
duration_s |
number | |
model |
string | Which Whisper model was used |
text |
string | Full text |
segments |
array | [{start, end, text}], times in seconds |
has_audio |
boolean | Whether the original recording is kept (file transcripts) |
created_at |
number | Unix time in seconds |
List results contain a preview (the first ~160 characters) instead of text and segments.
POST /v1/transcriptions
Starts a transcription job in its own asr lane, so it runs alongside speech generation. Send multipart/form-data:
| Field | Required | Description |
|---|---|---|
file |
✓ | Any common audio or video file (wav, mp3, m4a, flac, mp4, mov, webm…), up to 2 GB |
language |
ISO code such as en. Omit to detect automatically. |
|
model |
A Whisper model id. Default: your asr_model setting, or the most accurate model installed. |
|
title |
Defaults to the file name |
Response 202: the job. On success, result.transcript_id points to the transcript. Errors: 400 engine_unavailable, 400 file_too_large
JOB=$(curl -s -X POST "$VOXD/v1/transcriptions" -H "Authorization: Bearer $VOXD_TOKEN" -F file=@interview.m4a | jq -r .id)
# …wait for the job (GET /v1/jobs/$JOB or its events), then:
curl -s "$VOXD/v1/transcriptions/<transcript_id>" -H "Authorization: Bearer $VOXD_TOKEN"
GET /v1/transcriptions
Newest first. Query: q (searches title and text), limit (1–500, default 100).
GET /v1/transcriptions/{id}
PATCH /v1/transcriptions/{id}
Rename or correct a transcript: {"title": "…"} and/or {"segments": [{start, end, text}, …]}. The full text is rebuilt from the segments.
DELETE /v1/transcriptions/{id}
Deletes the transcript and its kept recording. Response 204.
GET /v1/transcriptions/{id}/audio
The original recording, for file transcripts.
GET /v1/transcriptions/{id}/export?format=txt|srt|vtt|json
Downloads the transcript as plain text, SubRip subtitles, WebVTT subtitles or JSON.
curl -s "$VOXD/v1/transcriptions/<id>/export?format=srt" -H "Authorization: Bearer $VOXD_TOKEN" -o talk.srt
POST /v1/transcriptions/{id}/save?format=…
Writes the export to a file on this computer. Body: {"path": "/abs/path/talk.srt", "overwrite": false}. The path must end in .<format>.
WS /v1/transcribe/live
Real-time transcription from a microphone or any audio stream.
Query: token (required when auth is on), language, model (default: the quickest model installed), save (default true: keep the result as a transcript), title.
Protocol
- The server sends
{"type": "ready", "sample_rate": 16000}. - Send audio as binary frames of 16-bit little-endian mono PCM at 16 kHz. About 100 ms per frame works well.
- While you speak, the server sends
{"type": "partial", "text": "…"}about once a second: the current sentence so far, which may still change. - After a pause (or 25 s of continuous speech), it sends
{"type": "final", "segment": {"start", "end", "text"}}. - Send the text frame
{"type": "stop"}to finish. You get a lastfinal, then{"type": "done", "text": "…", "transcript_id": "…" | null}, and the socket closes.
If no Whisper model is installed, you get {"type": "error", "error": "engine_unavailable", "message": …} instead.
Example (Python, streaming a WAV file in real time):
import asyncio, json, wave, websockets
async def main():
w = wave.open("speech-16k-mono.wav")
pcm = w.readframes(w.getnframes())
async with websockets.connect("ws://127.0.0.1:4870/v1/transcribe/live?token=TOKEN&save=false") as ws:
await ws.recv() # ready
async def read():
async for msg in ws:
data = json.loads(msg)
print(data["type"], data.get("text") or data.get("segment", {}).get("text"))
if data["type"] == "done":
return
reader = asyncio.create_task(read())
for i in range(0, len(pcm), 3200): # 100 ms of 16 kHz PCM16
await ws.send(pcm[i:i + 3200])
await asyncio.sleep(0.1)
await ws.send(json.dumps({"type": "stop"}))
await reader
asyncio.run(main())
Measured on an Apple M1 Pro with Whisper Base: partial text updates about once a second, and the final text arrives about 0.7 s after you stop speaking.