# remsky/Kokoro-FastAPI

Dockerized OpenAI-compatible wrapper for Kokoro-82M text-to-speech w/multiplatform CPU, AMD, NVIDIA GPU PyTorch; multi-speaker, voice-mixing, auto-stitching, caption timestamps, SSML, readalong web UI

Repository: https://github.com/remsky/Kokoro-FastAPI
Canonical: https://ross.abutalabs.com/products/kokoro-fastapi
Language: Python
License: Apache-2.0
License Family: permissive
Topics: fastapi, tts, tts-api, kokoro, kokoro-tts, pytorch, openai-compatible-api, openwebui, sillytavern, text-to-speech, kokoro-82m, mutilingual-audio, tts-generation, tts-captions, multi-speaker-tts, read-along, ssml
Last push: 2026-08-24T08:26:57+00:00

## Health v2 (maintenance only)
Score: 88/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 99, release rhythm 99, longevity 43
- inputs: {"age_days": 611, "days_push": 9, "days_rel": 9, "gap_med": 7.5, "n_releases_24m": 23}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 5373, forks 882 (observed 2026-08-28T04:09:16.170121+00:00)

## What it is
A Dockerized FastAPI wrapper around the Kokoro-82M text-to-speech model exposing an OpenAI-compatible speech endpoint with CPU, NVIDIA, AMD, and Apple Silicon support. It adds multi-speaker generation, voice mixing, SSML, timestamped captions, and an optional read-along web UI.

## Use cases
- self-host a text-to-speech API compatible with OpenAI clients
- generate audiobooks with per-word caption timestamps
- add TTS to OpenWebUI or SillyTavern
- mix multiple voices or generate multi-speaker dialogue audio
- run TTS on NVIDIA, AMD, or Apple Silicon GPUs
- convert text or phonemes to speech audio via REST API

## When to choose
- you need an OpenAI-compatible TTS endpoint you can self-host
- you want GPU-accelerated Kokoro inference without writing PyTorch code
- you need long-form generation with read-along captions
- you want drop-in TTS for OpenWebUI or SillyTavern

## When to avoid
- you need voice cloning of arbitrary speakers rather than fixed Kokoro voices
- you want a fully managed cloud TTS service
- you need non-Kokoro TTS models
- you cannot run Docker or Python on your host

## Facets
- artifact type: service
- maturity: active
- function: tts, http-server, api-framework, audio-processing, llm-inference
- domain: speech-processing, artificial-intelligence, self-hosted, apis
- platform: self-hosted, python, windows
- tags: kokoro, openai-compatible, text-to-speech, fastapi, gpu-inference, ssml, captions, voice-mixing, multi-speaker, webui, audio, docker, web-server, linux, macos

## Member repositories
- remsky/Kokoro-FastAPI (main) score 88

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:09:16.170121+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T17:58:42.098817+00:00, confidence not recorded.
  - readme: https://github.com/remsky/Kokoro-FastAPI (fetched 2026-08-28T04:09:16.170121+00:00, sha de2049e063d6)
- Data as of 2026-08-30T08:39:29.467469+00:00.
