# MOSS-TTS

MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenarios, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS.

Repository: https://github.com/OpenMOSS/MOSS-TTS
Canonical: https://ross.abutalabs.com/products/moss-tts
Homepage: https://mosi.cn/models/moss-tts
Language: Python
License: Apache-2.0
License Family: permissive
Topics: audio, audio-tokenizer, llm, multimodal, text-to-speech, voice-cloning
Last push: 2026-07-26T11:27:15+00:00
Link (homepage): https://mosi.cn/models/moss-tts

## Health v2 (maintenance only)
Score: 57/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 94, release rhythm 35, longevity 14
- inputs: {"age_days": 207, "days_push": 38, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 4031, forks 365 (observed 2026-08-28T04:08:32.740834+00:00)

## What it is
MOSS-TTS-Nano is an open-source 0.1B-parameter multilingual speech generation (TTS) model from MOSI.AI and the OpenMOSS team, designed for realtime streaming synthesis. It runs directly on CPU without a GPU, making it easy to deploy in local demos, web serving, and lightweight product integrations.

## Use cases
- generate speech from text in chinese and english
- run text-to-speech on cpu without a gpu
- clone a voice from a short audio sample
- stream realtime tts audio in a web app
- embed a tiny multilingual tts model in a product
- synthesize speech locally for privacy-sensitive applications

## When to choose
- you need lightweight, realtime TTS that runs on CPU or edge devices
- you want an open-source, Apache-2.0 licensed speech generation model
- you need multilingual (Chinese/English) synthesis with voice cloning
- you want a simple deployment stack for local demos or web serving

## When to avoid
- you need the highest-fidelity, studio-quality speech synthesis available
- you require languages beyond the supported multilingual set
- you need a fully managed cloud TTS API rather than a self-run model
- you need heavy-duty voice conversion or audio editing features beyond TTS

## Facets
- artifact type: library
- maturity: active
- function: tts, speech-recognition, audio-processing, llm-inference, machine-learning
- domain: speech-processing, artificial-intelligence
- platform: python, cross-platform, windows
- tags: voice-clone, streaming-audio, multilingual, realtime, audio-tokenizer, tiny-model, on-device, chinese, english, audio, natural-language-processing, cpu, gpu, macos, linux

## Member repositories
- OpenMOSS/MOSS-TTS (main) score 57
- OpenMOSS/MOSS-TTS-Nano (main) score 57
- OpenMOSS/MOSS-TTSD (main) score 61

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:08:32.740834+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T18:22:12.052376+00:00, confidence not recorded.
  - readme: https://github.com/OpenMOSS/MOSS-TTS (fetched 2026-08-28T04:08:32.740834+00:00, sha 31d041299677)
  - homepage: https://mosi.cn/models/moss-tts (fetched 2026-08-29T09:12:09.040566+00:00, sha aaf084a1e7e0)
- Data as of 2026-08-30T08:39:29.467469+00:00.
