VoxStudio Documentation

VoxStudio is a private voice studio that runs entirely on your computer. These docs cover the app and voxd, the local API that powers it.

Use the app

Guide What you'll do
Your first voice in 60 seconds Launch VoxStudio and hear a voice
Drop anywhere Drag in audio, video, books or text and pick what to do
Background work & Activity Long renders, progress and cancelling
Voice models Download natural voices and manage disk space
Studio Write a script, pick a voice, shape the delivery
Script markup Pauses, pace, emphasis and pronunciation inside your script
Clone a voice Make a voice from a short recording
Your voice library Browse, favorite, tag, share and import voices
Editor Trim, split, splice, fade and polish takes on a timeline
Compare voices Blind A/B tests and your favourite voices per language
History Find, replay, star, save and clean up takes
Projects Group your work; find every exported file
Install, update & automate Installing, updates, menus, voxstudio:// links, uninstalling
Settings Appearance, storage, performance, shortcuts, privacy and logs
Integrations & developer tools Connect Claude, Cursor, n8n and OpenAI SDKs; API keys; API explorer
Tools Clean up recordings, change a voice, fix pronunciation
Batch & watch folders Speak or transcribe many files; automate a folder
Design a voice Describe a voice in words and fine-tune it by ear
Dub a video Translate and re-voice a video
Stories & audiobooks Import books, cast characters, narrate, read along, export M4B
Transcribe Files and live speech to text; fix and export
Quick Speak & Speak Selection Hear anything you type or select, from any app
Dictation Type with your voice in any app

Build with the API

Page
Overview How voxd works, base URL, versioning
Quickstart Generate speech in 3 requests (curl, Python, JavaScript)
Authentication Tokens and API keys
Errors Error format and codes

Reference

Resource Endpoints
API keys & connection GET /v1/connection, create and revoke keys, finding voxd
OpenAI compatible POST /v1/audio/speech, POST /v1/audio/transcriptions
MCP server POST /mcp tools for AI agents; stdio bridge
System GET /v1/status, GET /v1/system
Engines GET /v1/engines
Models GET /v1/models, download, remove, mirrors
Voices GET /v1/voices, library, favorites and tags
Custom voices Create, rename or delete voices; export and import .voxvoice
Projects & exports Projects, items, membership, export history
Compare Blind ratings and the per-language leaderboard
Tools & pronunciations Clean audio, voice conversion, pronunciation dictionary
Batch & watch folders Batches of speech/transcription; folders processed automatically
Stories & audiobooks Import EPUB/DOCX/TXT/MD, cast, resumable render, timings, M4B/MP3
Dubbing Prepare, edit, cast, render and export dubs
Transcription File jobs, transcripts, exports, WS /v1/transcribe/live
Voice design Analyze, candidates from a description, designed voices
Settings GET/PATCH /v1/settings, storage usage and clean-up, free memory
Speech POST /v1/speech
Jobs POST /v1/jobs/speech, GET /v1/jobs, events (SSE), cancel, clear
Live events WS /v1/events
Takes Search, page, audio, star, export, delete, stats, retention, edit

A live, interactive reference is always available from a running voxd at http://127.0.0.1:<port>/docs.