| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
| Name | Name | Last commit date | ||
|---|---|---|---|---|
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model. Generate hours of high quality speech in minutes.
Community projects that use, recommend, or enable Kokoro-FastAPI as a backend:
Pre-built multi-arch images with models baked in.
:latest is available, but please pin to a release tag for stable usage.
No GPU (laptop, CPU-only server)
docker run -p 8880:8880 ghcr.io/remsky/kokoro-fastapi-cpu:latestNVIDIA (GTX 900-series through RTX 40; ships cu126)
docker run --gpus all -p 8880:8880 ghcr.io/remsky/kokoro-fastapi-gpu:latestNVIDIA RTX 50-series / Blackwell (ships cu128)
docker run --gpus all -p 8880:8880 ghcr.io/remsky/kokoro-fastapi-gpu:latest-cu128NVIDIA arm64 (Jetson, GH200; same tag, ships cu129)
docker run --gpus all -p 8880:8880 ghcr.io/remsky/kokoro-fastapi-gpu:latestAMD GPU (ROCm, experimental, x86_64 only)
docker run --device=/dev/kfd --device=/dev/dri -p 8880:8880 ghcr.io/remsky/kokoro-fastapi-rocm:latestApple Silicon (native MPS clone; the CPU image also works)
./start-gpu_mac.shgpu:latest is the same image as gpu:latest-cu126. Configuration via environment variables, see the configuration guide.
Quick Start (docker compose)git clone https://github.com/remsky/Kokoro-FastAPI.git
cd Kokoro-FastAPI
cd docker/gpu # For NVIDIA GPU support
# or cd docker/cpu # For CPU support
# or cd docker/rocm # For AMD GPU (ROCm, experimental, amd64 only)
docker compose up --build
# *Note for Apple Silicon (M1/M2/M3) users:
# The Docker GPU image is CUDA-only and won't run on Apple Silicon. With Docker, use `docker/cpu`.
# For native MPS (Apple GPU) acceleration, run directly via UV with `./start-gpu_mac.sh`.
cd ../.. # back to repo root for the paths below
# Models will auto-download, but if needed you can manually download:
python docker/scripts/download_model.py --output api/src/models/v1_0Configuration guide covers image vs build, the volume mounts, and env vars.
Direct Run (via uv)Install astral-uv
Install espeak-ng in your system if you want it available as a fallback for unknown words/sounds. The upstream libraries may attempt to handle this, but results have varied.
Clone the repository:
git clone https://github.com/remsky/Kokoro-FastAPI.git
cd Kokoro-FastAPIRun the model download script if you haven't already
Start directly via UV (with hot-reload)
Linux and macOS
./start-cpu.sh OR
./start-gpu.sh Windows
.\start-cpu.ps1 OR
.\start-gpu.ps1 Run locally as an OpenAI-Compatible Speech Endpoint
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8880/v1", api_key="not-needed"
)
with client.audio.speech.with_streaming_response.create(
model="kokoro",
voice="af_sky+af_bella", #single or multiple voicepack combo
input="Hello world!"
) as response:
response.stream_to_file("output.mp3")The API will be available at http://localhost:8880
API Documentation: http://localhost:8880/docs
Web Interface: http://localhost:8880/web
# Using OpenAI's Python library
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8880/v1", api_key="not-needed")
response = client.audio.speech.create(
model="kokoro",
voice="af_bella+af_sky", # see /api/src/core/openai_mappings.json to customize
input="Hello world!",
response_format="mp3"
)
response.stream_to_file("output.mp3")Or Via Requests:
import requests
response = requests.get("http://localhost:8880/v1/audio/voices")
voices = [v["id"] for v in response.json()["voices"]]
# Generate audio
response = requests.post(
"http://localhost:8880/v1/audio/speech",
json={
"model": "kokoro",
"input": "Hello world!",
"voice": "af_bella",
"response_format": "mp3", # Supported: mp3, wav, opus, flac, aac, pcm
"speed": 1.0
}
)
# Save audio
with open("output.mp3", "wb") as f:
f.write(response.content)Quick tests (run from another terminal):
python examples/assorted_checks/test_openai/test_openai_tts.py # Test OpenAI Compatibility
python examples/assorted_checks/test_voices/test_all_voices.py # Test all available voices# OpenAI-compatible streaming
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8880/v1", api_key="not-needed")
# Stream to file
with client.audio.speech.with_streaming_response.create(
model="kokoro",
voice="af_bella",
input="Hello world!"
) as response:
response.stream_to_file("output.mp3")
# Stream to speakers (requires PyAudio)
import pyaudio
player = pyaudio.PyAudio().open(
format=pyaudio.paInt16,
channels=1,
rate=24000,
output=True
)
with client.audio.speech.with_streaming_response.create(
model="kokoro",
voice="af_bella",
response_format="pcm",
input="Hello world!"
) as response:
for chunk in response.iter_bytes(chunk_size=1024):
player.write(chunk)Or via requests:
import requests
response = requests.post(
"http://localhost:8880/v1/audio/speech",
json={
"input": "Hello world!",
"voice": "af_bella",
"response_format": "pcm"
},
stream=True
)
for chunk in response.iter_content(chunk_size=1024):
if chunk:
# Process streaming chunks
passKey Streaming Metrics:
Note: Artifacts in intonation can increase with smaller chunks
Multiple Output Audio FormatsCombine voices and generate audio:
import requests
response = requests.get("http://localhost:8880/v1/audio/voices")
voices = [v["id"] for v in response.json()["voices"]]
# Weighted voice combination (67%/33% mix)
response = requests.post(
"http://localhost:8880/v1/audio/speech",
json={
"input": "Hello world!",
"voice": "af_bella(2)+af_sky(1)", # 2:1 ratio = 67%/33%
"response_format": "mp3"
}
)
# Download combined voice as .pt file
response = requests.post(
"http://localhost:8880/v1/audio/voices/combine",
json="af_bella(2)+af_sky(1)" # 2:1 ratio = 67%/33%
)
# Save the .pt file
with open("combined_voice.pt", "wb") as f:
f.write(response.content)
# Use the downloaded voice file
response = requests.post(
"http://localhost:8880/v1/audio/speech",
json={
"input": "Hello world!",
"voice": "combined_voice", # Use the saved voice file
"response_format": "mp3"
}
)Weighted mixes can get long fast. voice_aliases maps a short name per request, for both the voice field and [voice:...] tags:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8880/v1", api_key="not-needed")
client.audio.speech.create(
model="kokoro",
voice="narrator",
input="[voice:narrator] Once upon a time. [voice:villain] Never!",
extra_body={
"allow_voice_tags": True,
"voice_aliases": {"narrator": "af_bella(2)+af_sky", "villain": "am_michael"},
},
)or
curl -X POST http://localhost:8880/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{
"model": "kokoro",
"voice": "narrator",
"input": "[voice:narrator] Once upon a time. [voice:villain] Never!",
"allow_voice_tags": true,
"voice_aliases": {"narrator": "af_bella(2)+af_sky", "villain": "am_michael"},
"response_format": "mp3"
}' --output aliased.mp3POST /dev/tune takes a 3 to 30 s clip of one English speaker and speaks with a voice tuned toward it, via inno-kokoro. A tuner, not a cloner: expect the same neighbourhood, not a match. Only tune voices you have permission to use.
curl -s http://localhost:8880/dev/tune -F audio=@ref.wav -F 'request={"input":"Hello there."}' -o out.mp3
curl -s http://localhost:8880/dev/tune -F audio=@ref.wav -F return_voice_pack=true -o ref.pt
curl -s http://localhost:8880/dev/tune -F audio=@ref.wav -F save_voice=am_ref # saves am_ref_tuned, needs ALLOW_LOCAL_VOICE_SAVING=truecurl -X POST http://localhost:8880/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{
"model": "kokoro",
"voice": "af_heart",
"input": "The narrator opens. [voice:af_bella] Did it land? [pause:0.3s] [voice:am_michael] It did.",
"allow_voice_tags": true,
"response_format": "mp3"
}' --output dialogue.mp3With the official OpenAI client, pass the param in extra_body:
client.audio.speech.create(
model="kokoro",
voice="af_jadzia",
input="The narrator opens. [voice:af_bella] Did it land?",
extra_body={"allow_voice_tags": True},
)POST /dev/dialogue uses structured turns, and allows pause_between_turns to be controlled by param.
curl -X POST http://localhost:8880/dev/dialogue \
-H "Content-Type: application/json" \
-d '{
"turns": [
{"voice": "af_bella", "text": "Did the multi speaker support land?"},
{"voice": "am_michael", "text": "It did. Turns switch voices inline."}
],
"pause_between_turns": 0.4,
"response_format": "mp3"
}' --output dialogue.mp3Notes:
Number of voices has minimal impact on generation speed. For continuous swaps though, if each speaker gets less than about 2 sentences, chunking requirements slow generation down. Still a flat cost, not compounding as the text grows. Regenerate with examples/assorted_checks/test_dialogue/.
Inline Control TokensFour tokens can be embedded in the input text and are parsed server-side (API, WebUI, or any client):
The city of [Worcester](/wˈʊstər/) is easy. [pause:1s] See?
Send ssml: true with allow_voice_tags: true on /v1/audio/speech or /dev/captioned_speech to translate and speak in one call. Both flags are needed, since the translation emits [voice:] and [rate:] spans that would otherwise be read aloud; ssml without them is a 400.
{
"model": "kokoro",
"input": "<speak>Hi<break time=\"750ms\"/>there</speak>",
"voice": "af_bella",
"allow_voice_tags": true,
"ssml": true
}POST /dev/ssml does the translation on its own when you want the tokens back as text rather than audio, or want to inspect them before synthesis. Send text, plus voice if your speech request uses one, then pass the result back with allow_voice_tags: true. Without a voice, <voice>/<prosody> are stripped and their content kept.
curl -s http://localhost:8880/dev/ssml -H "Content-Type: application/json" \
-d '{"text": "<speak>The city of <phoneme alphabet=\"ipa\" ph=\"wˈʊstər\">Worcester</phoneme> is easy.<break time=\"1s\"/>See?</speak>"}'
# {"text": "The city of [Worcester](/wˈʊstər/) is easy. [pause:1.0s] See?"}The model takes up to 510 phonemized tokens per chunk, but running it that long tends to produce 'rushed' speech and other artifacts. The server adds its own chunking layer on top, sized by TARGET_MIN_TOKENS, TARGET_MAX_TOKENS, and ABSOLUTE_MAX_TOKENS (175, 250, 450 by default, set via environment variables).
Phoneme & Token RoutesConvert text to phonemes and/or generate audio directly from phonemes:
import requests
def get_phonemes(text: str, language: str = "a"):
"""Get phonemes and tokens for input text"""
response = requests.post(
"http://localhost:8880/dev/phonemize",
json={"text": text, "language": language} # "a" for American English
)
response.raise_for_status()
result = response.json()
return result["phonemes"], result["tokens"]
def generate_audio_from_phonemes(phonemes: str, voice: str = "af_bella"):
"""Generate audio from phonemes"""
response = requests.post(
"http://localhost:8880/dev/generate_from_phonemes",
json={"phonemes": phonemes, "voice": voice},
headers={"Accept": "audio/wav"}
)
if response.status_code != 200:
print(f"Error: {response.text}")
return None
return response.content
# Example usage
text = "Hello world!"
try:
# Convert text to phonemes
phonemes, tokens = get_phonemes(text)
print(f"Phonemes: {phonemes}") # e.g. ðɪs ɪz ˈoʊnli ɐ tˈɛst
print(f"Tokens: {tokens}") # Token IDs including start/end tokens
# Generate and save audio
if audio_bytes := generate_audio_from_phonemes(phonemes):
with open("speech.wav", "wb") as f:
f.write(audio_bytes)
print(f"Generated {len(audio_bytes)} bytes of audio")
except Exception as e:
print(f"Error: {e}")See examples/phoneme_examples/generate_phonemes.py for a sample script.
Timestamps (word level)Generate audio with word-level timestamps without streaming:
import requests
import base64
import json
response = requests.post(
"http://localhost:8880/dev/captioned_speech",
json={
"model": "kokoro",
"input": "Hello world!",
"voice": "af_bella",
"speed": 1.0,
"response_format": "mp3",
"stream": False,
},
stream=False
)
with open("output.mp3","wb") as f:
audio_json=json.loads(response.content)
chunk_audio=base64.b64decode(audio_json["audio"].encode("utf-8"))
f.write(chunk_audio)
print(audio_json["timestamps"])Generate audio with word-level timestamps with streaming:
import requests
import base64
import json
response = requests.post(
"http://localhost:8880/dev/captioned_speech",
json={
"model": "kokoro",
"input": "Hello world!",
"voice": "af_bella",
"speed": 1.0,
"response_format": "mp3",
"stream": True,
},
stream=True
)
f=open("output.mp3","wb")
for chunk in response.iter_lines(decode_unicode=True):
if chunk:
chunk_json=json.loads(chunk)
chunk_audio=base64.b64decode(chunk_json["audio"].encode("utf-8"))
f.write(chunk_audio)
print(chunk_json["timestamps"])With "allow_voice_tags": true, each timestamp also carries the voice that spoke the word, so multi-speaker captions can be labelled without re-deriving the split client side. Without it the field is absent.
Timestamps (streaming chunks)With stream, return_download_link, and return_timing set, the response carries an X-Timing-Path header pointing at a JSON sidecar of per-chunk timings. The audio body is unchanged and nothing extra is computed; this is what drives the web UI's read along.
response = requests.post(
"http://localhost:8880/v1/audio/speech",
json={
"input": "Hello world! [pause:1s] Again.",
"voice": "af_bella",
"stream": True,
"return_download_link": True,
"return_timing": True,
},
stream=True,
)
audio = b"".join(response.iter_content(1024))
timings = requests.get(f"http://localhost:8880/v1{response.headers['x-timing-path']}").json()
# {"chunks": [{"text": "Hello world!", "start": 0.0, "end": 0.64}, ...]}Generation through the local API, text lengths up to feature-length books (~1.5 hours output), measuring processing time and realtime factor. Run on:
Key Performance Metrics:
POST /dev/unload frees the model from VRAM and reloads lazily on the next request. Reclaim scales with load (the activation pool, not just weights) but plateaus: chunks cap at 450 tokens. Long-form = ~30 paragraphs. Same setup as above.
| Workload | Loaded | Floor | Reclaimed | Reload |
|---|---|---|---|---|
| Short (6s audio) | 3.11 GB | 2.37 GB | 758 MiB | +4.9s |
| Long-form (7.5m) | 3.98 GB | 2.37 GB | 1,656 MiB | +5.1s |
Floor is host + CUDA context. Reproduce with uv run --extra benchmarks assorted_checks/benchmarks/benchmark_model_unload.py from examples/.
To automatically unload the model after an idle timeout, set MODEL_AUTO_UNLOAD_TIMEOUT_SECONDS to a positive number of seconds. The default 0 disables auto-unload.
End-to-end roundtrip: synthesize with Kokoro, transcribe the result back with faster-whisper, compare to the source text. Scripts and data live under examples/assorted_checks/test_transcription/.
Long-form English (full book, A Journey to the Centre of the Earth, Project Gutenberg, voice af_heart, base.en Whisper on CUDA float16, baseline captured on cu126 GPU build):
| Run | Input chars | Audio length | Synth speedup | Transcribe speedup | WER |
|---|---|---|---|---|---|
| Short (~ch.7) | 64,996 | 66m 06s | 36.4x rt | 62.4x rt | 0.047 |
| Full book | 502,766 | 507m 52s | 45.7x rt | 65.1x rt | 0.033 |
See examples/assorted_checks/test_transcription/BASELINE.md for the full regression bands.
Per-language check (single-sentence per voice, multilingual Whisper small. WER for Latin scripts, CER for ja/zh/hi):
| Language | Voice | Metric | Score |
|---|---|---|---|
| English | af_heart | WER | 0.000 |
| English (UK) | bf_emma | WER | 0.111 |
| Spanish | ef_dora | WER | 0.000 |
| French | ff_siwis | WER | 0.000 |
| Italian | if_sara | WER | 0.000 |
| Portuguese | pf_dora | WER | 0.000 |
| Hindi | hf_alpha | CER | 0.059 |
| Japanese | jf_alpha | CER | 0.000 |
| Chinese | zf_xiaobei | CER | 0.143 |
Caveat: these are single short sentences, not a comprehensive per-language quality benchmark. They confirm each voice produces transcribable audio in its target language; deeper quality evaluation per language is open work.
To reproduce, see examples/assorted_checks/test_transcription/README.md.
Configuration VariablesEvery setting is an environment variable, or a line in a .env file at the project root. Full reference in the configuration guide.
Debug EndpointsSystem state and resource usage, for debugging exhaustion or performance issues. The /debug/* routes expose host and process internals, so they are off by default; set ENABLE_DEBUG_ENDPOINTS=true to enable.
Stability: the /v1/* OpenAI-compatible routes are the stable API. /dev/* and /debug/* are operational helpers, and may change or move behind flags between minor releases.
LoggingGlobal API loguru logging level can be set using the API_LOG_LEVEL environment variable. Defaults to DEBUG. Per run method in the configuration guide.
Missing words & Missing some timestampsThe API normalizes input text, which can incorrectly remove or change some phrases. Disable it with "normalization_options":{"normalize": false} in the request json, the full field list is in Text normalization:
import requests
response = requests.post(
"http://localhost:8880/v1/audio/speech",
json={
"input": "Hello world!",
"voice": "af_heart",
"response_format": "pcm",
"normalization_options":
{
"normalize": False
}
},
stream=True
)
for chunk in response.iter_content(chunk_size=1024):
if chunk:
# Process streaming chunks
passSee docs/troubleshooting.md#linux-gpu-permissions for container group, host group, and device permission options.
AMD GPU (ROCm) troubleshootingSee docs/troubleshooting.md#amd-gpu-rocm for HSA overrides, MIOpen warmup, hipBLAS fallback, and native Linux host requirement.
WAV duration reported as nonsense in some readersWAV responses use a streaming-sentinel (0xFFFFFFFF) for the size fields in the header. Most readers handle this fine: soundfile, pydub/ffmpeg, browsers, OS players. Python's stdlib wave does not, and reports a bogus duration. For exact length use soundfile.info(path).duration or ffprobe.
Versioning & DevelopmentBranching Strategy:
Note: This is a development focused project at its core.
If you run into trouble, you may have to roll back a version on the release tags if something comes up, or build up from source and/or troubleshoot + submit a PR.
Free and open source is a community effort, and there's only really so many hours in a day. If you'd like to support the work, feel free to open a PR, buy me a coffee, or report any bugs/features/etc you find during use.
Working on the code, or pointing an AI agent at it? AGENTS.md covers the repo layout, commands, and conventions.
This API uses the Kokoro-82M model from HuggingFace.
Visit the model page for more details about training, architecture, and capabilities. I have no affiliation with any of their work, and produced this wrapper for ease of use and personal projects.
License This project is licensed under the Apache License 2.0 - see below for details:The full Apache 2.0 license text can be found at: https://www.apache.org/licenses/LICENSE-2.0
Project Structure and ChurnMade with contrib.rocks.
| Back | FazBrowse Home | New Git URL |