Engines

GET /v1/engines

Lists every engine voxd knows about, including ones that can't run on this machine.

Response 200: an array of:

Field Type Description
id string Pass this as engine elsewhere, e.g. "system"
name string Display name
capabilities string[] What it can do: asr (transcribes speech), tts (speaks text), clone (speaks in custom voices), emotion (honours emotion in speech requests)
license string License of the underlying model or tool
available boolean true if usable right now
unavailable_reason string | null Why not, e.g. "Install espeak-ng to enable system voices"
curl -s "$VOXD/v1/engines" -H "Authorization: Bearer $VOXD_TOKEN"
# [{"id":"system","name":"System Voices","capabilities":["tts"],"license":"Built into your OS","available":true,"unavailable_reason":null},
#  {"id":"kokoro","name":"Kokoro","capabilities":["tts"],"license":"Apache-2.0","available":false,"unavailable_reason":"Download Kokoro from Models to use it"}]

Available engines

id Platforms Notes
system macOS (built-in voices), Linux (espeak-ng) Instant, no download. Windows support is planned.
kokoro macOS, Windows, Linux (CPU) 54 natural voices in 9 languages. Needs the kokoro-v1 model (354 MB). Until it's installed, unavailable_reason says so.
chatterbox macOS, Windows, Linux (Apple GPU, NVIDIA or CPU) Voice cloning with emotion control (English). Needs the chatterbox-v1 model (3.2 GB) plus its engine runtime (about 1.3 GB), which is installed in its own isolated environment and runs as a separate process.
whisper macOS, Windows, Linux (CPU; NVIDIA GPU via CUDA) Speech recognition in ~99 languages. Needs a Whisper model and its runtime (about 260 MB). See Transcription.