OpenAI-compatible endpoints
Code written for OpenAI's audio API works against voxd. Set the base URL to http://127.0.0.1:4870/v1 and use a VoxStudio API key as the API key. Everything runs locally, and every result is saved in History or Transcribe.
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:4870/v1", api_key="vox_sk_…")
POST /v1/audio/speech
| Field | ||
|---|---|---|
input |
required | Text, up to 100,000 characters. Long text is rendered in chunks. |
voice |
alloy |
An OpenAI voice name, or any VoxStudio voice id (af_heart, cv_…, dv_…, mix:…) |
model |
tts-1 |
tts-1, tts-1-hd and gpt-4o-mini-tts use the best installed engine. kokoro, chatterbox or system pick one explicitly. |
response_format |
mp3 |
mp3, opus (Ogg), aac (ADTS), flac, wav, or pcm (raw 24 kHz 16-bit mono) |
speed |
1.0 |
0.25–4.0 accepted; applied within 0.5–2.0 |
instructions |
— | Accepted and ignored |
The response is the audio bytes. The X-Voxd-Take-Id header names the take saved in History.
OpenAI voice names map to similar Kokoro voices:
| alloy | ash | ballad | coral | echo | fable | nova | onyx | sage | shimmer | verse |
|---|---|---|---|---|---|---|---|---|---|---|
| af_alloy | am_adam | bm_george | af_heart | am_echo | bm_fable | af_nova | am_onyx | af_sarah | af_sky | am_michael |
If Kokoro isn't installed, the best available engine's first voice is used. mp3, opus, aac and flac need the Whisper runtime to encode. wav and pcm always work.
POST /v1/audio/transcriptions
Multipart form: file (required), model (default whisper-1), language, response_format, prompt and temperature. prompt and temperature are accepted and ignored.
whisper-1 and gpt-4o-transcribe use your preferred Whisper model. To choose a specific one, pass a VoxStudio model id such as whisper-large-v3-turbo.
response_format |
Returns |
|---|---|
json (default) |
{"text": "…"} |
text |
Plain text |
srt, vtt |
Subtitles |
verbose_json |
{"task", "language", "duration", "text", "segments": [{"id", "start", "end", "text", …}]} |
Errors
These endpoints return errors in OpenAI's shape, so SDKs raise their usual exceptions:
{"error": {"message": "Unknown voice: nobody", "type": "invalid_request_error", "param": "voice", "code": "voice_not_found"}}
Codes: voice_not_found, engine_unavailable, unsupported_format, invalid_value, file_too_large.