# k2-fsa/sherpa-onnx

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages

Repository: https://github.com/k2-fsa/sherpa-onnx
Canonical: https://ross.abutalabs.com/products/sherpa-onnx
Homepage: https://k2-fsa.github.io/sherpa/onnx/index.html
Language: C++
License: Apache-2.0
License Family: permissive
Topics: asr, onnx, windows, linux, macos, cpp, android, ios, raspberry-pi, aarch64, arm32, csharp, dotnet, mfc, speech-to-text, text-to-speech, vits, risc-v, lazarus, object-pascal
Last push: 2026-08-25T03:44:25+00:00

## Health v2 (maintenance only)
Score: 95/100 (v2, computed 2026-09-03T02:39:23.370411+00:00)
- activity 99, release rhythm 87, longevity 100
- inputs: {"age_days": 1462, "days_push": 8, "days_rel": 8, "gap_med": 6.0, "n_releases_24m": 91}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 14411, forks 1652 (observed 2026-08-28T04:11:06.950181+00:00)

## What it is
sherpa-onnx is an offline speech processing toolkit built on next-gen Kaldi and onnxruntime, supporting speech-to-text, text-to-speech, speaker diarization/verification, VAD, keyword spotting, speech enhancement, and source separation without Internet access. It runs on a wide range of platforms including Android, iOS, embedded boards, and NPUs, with bindings for 12 programming languages.

## Use cases
- run offline speech recognition on android
- convert text to speech locally without internet
- transcribe audio on raspberry pi
- add voice activity detection to an app
- do speaker diarization on device
- embed asr into an ios app
- run speech recognition on embedded arm boards
- build a local voice assistant

## When to choose
- you need fully offline, on-device speech recognition or TTS
- you target mobile, embedded, or NPU hardware like Android, iOS, Raspberry Pi, or RISC-V
- you need bindings across many languages (C++, Python, Java, Swift, Go, etc.)
- you want streaming and non-streaming ASR with VAD and keyword spotting in one toolkit

## When to avoid
- you need cloud-hosted, managed speech APIs with vendor-managed models
- you want training or fine-tuning of speech models rather than inference
- you need non-speech audio tasks like music generation
- you prefer PyTorch-native pipelines instead of ONNX export

## Facets
- artifact type: library
- maturity: active
- function: speech-recognition, tts, audio-processing, sdk, websocket, cli
- domain: speech-processing, embedded-systems, cross-platform, machine-learning
- platform: windows, cpp, python, wasm, embedded, cross-platform
- tags: onnx, onnxruntime, asr, vad, speaker-diarization, keyword-spotting, offline, kaldi, npu, raspberry-pi, risc-v, speech-enhancement, source-separation, harmonyos, natural-language-processing, android, ios, macos, linux, docker

## Member repositories
- k2-fsa/sherpa-onnx (main) score 95

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:11:06.950181+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T17:12:35.668049+00:00, confidence not recorded.
  - readme: https://github.com/k2-fsa/sherpa-onnx (fetched 2026-08-28T04:11:06.950181+00:00, sha 0e3ff0ce36f5)
  - homepage: https://k2-fsa.github.io/sherpa/onnx/index.html (fetched 2026-08-29T08:06:00.468083+00:00, sha 368e5002a04e)
- Data as of 2026-08-30T08:39:29.467469+00:00.
