Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: tts

498 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
LuckyHookin/edge-TTS-record
A Windows desktop tool that records Microsoft Edge's online neural text-to-speech voices (e.g., Xiaoxiao, Yunyang) and saves the output as …
231370maintenance
Jackywine/Bella
Bella is a self-hosted Node.js web application that acts as a personalized AI digital companion with voice interaction. It combines Whisper…
486379experimental
kripken/speak.js
speak.js is a port of the eSpeak C++ speech synthesizer to JavaScript via Emscripten, enabling text-to-speech in the browser using only Jav…
321335maintenance
kakaobrain/pororo
PORORO is a Python library from Kakao Brain that provides unified access to dozens of pretrained neural models for natural language process…
101305maintenance
sdkcarlos/artyom.js
Artyom.js is a JavaScript library that wraps the Web Speech APIs (webkitSpeechRecognition and speechSynthesis) to add voice control, speech…
231269maintenance
PlayVoice/vits_chinese
A Chinese text-to-speech library combining VITS with BERT-based prosody embeddings and NaturalSpeech infer-loss features, supporting ONNX c…
231227maintenance
bawangxx/XZVoice
XZVoice is a free, open-source desktop text-to-speech application built with Electron, Vue, and ElementUI. It uses Alibaba Cloud's speech s…
231168maintenance
spring-media/TransformerTTS
A TensorFlow 2 implementation of a non-autoregressive Transformer-based neural network for text-to-speech synthesis, based on FastSpeech an…
101161maintenance
Kyubyong/dc_tts
A TensorFlow implementation of DC-TTS, a text-to-speech model based on deep convolutional networks with guided attention. It includes train…
321156maintenance
openinterpreter/01
An open-source voice interface platform that lets users control computers conversationally, powered by Open Interpreter. It pairs a Python …
165156experimental
RafalWilinski/telegram-chatgpt-concierge-bot
A self-hosted Telegram bot that lets you chat with OpenAI's ChatGPT via text and voice messages. It uses LangChainJS for conversation histo…
301130maintenance
synesthesiam/opentts
OpenTTS is a text-to-speech server that unifies access to multiple open source TTS systems (Larynx, Glow-Speak, Coqui-TTS, MaryTTS, flite, …
101118maintenance
Syan-Lin/CyberWaifu
CyberWaifu is a Python chatbot application that combines LLMs (ChatGPT, Claude) with TTS (edge-tts, Azure) to create realistic conversation…
301099maintenance
JiehangXie/PaddleBoBo
PaddleBoBo is a Python project built on PaddlePaddle (with PaddleSpeech and PaddleGAN) that quickly generates a virtual streamer (VTuber) f…
321062maintenance
Yink/Amadeus
An Android app that replicates the Amadeus AI assistant app from Steins;Gate 0, built primarily for cosplay purposes. It features speech re…
231061maintenance
Edresson/YourTTS
YourTTS is a zero-shot multi-speaker text-to-speech and voice conversion model built on VITS, implemented in the Coqui TTS framework. It su…
231053maintenance
bravekingzhang/text2video
A Python web application that converts text (e.g., novel passages) into narrated videos. It splits text into sentences, generates images vi…
291049maintenance
cbh123/narrator
A Python app that watches your webcam and generates David Attenborough-style narration of what it sees, using GPT vision models and ElevenL…
694426experimental
susiai/susi_device
SUSI Device provides sources to install the SUSI AI assistant stack on a Raspberry Pi, combining microphone/speaker, a small display, a loc…
321004maintenance
NATSpeech/NATSpeech
A PyTorch framework for non-autoregressive text-to-speech (NAR-TTS), containing official implementations of PortaSpeech (NeurIPS 2021) and …
231004maintenance
enhuiz/vall-e
An unofficial PyTorch implementation of the VALL-E text-to-speech audio language model, built on the EnCodec tokenizer. It provides trainin…
312976experimental
Priler/jarvis
JARVIS is an offline, privacy-respecting voice assistant built in Rust with Tauri, using neural networks for speech-to-text, text-to-speech…
532909experimental
yakami129/VirtualWife
VirtualWife is a self-hosted virtual digital human (AI companion) application that combines VRM 3D character models with LLM-powered conver…
202889experimental
fikrikarim/parlor
Parlor is a fully on-device, real-time multimodal voice assistant similar to GPT-Live, combining speech recognition, a Gemma vision-languag…
782039experimental
Mini-Omni
Mini-Omni is an open-source multimodal large language model that performs real-time end-to-end speech-to-speech conversation with streaming…
221922experimental
collabora/WhisperFusion
WhisperFusion is a real-time voice chat application that combines WhisperLive speech-to-text, a Mistral/Phi LLM, and WhisperSpeech text-to-…
261647experimental
lucidrains/naturalspeech2-pytorch
A PyTorch implementation of NaturalSpeech 2, a zero-shot text-to-speech and singing synthesizer that combines a neural audio codec with a l…
201333experimental
elfvingralf/macOSpilot-ai-assistant
macOSpilot is a macOS desktop AI assistant built with Electron that answers spoken or typed questions about whatever application is current…
261157experimental
supertone-inc/supertonic
Supertonic is a lightning-fast, on-device multilingual text-to-speech system powered by ONNX Runtime, with a compact 99M-parameter open-wei…
5813734abandoned
idootop/mi-gpt
MiGPT is a Node.js/Docker application that connects Xiaomi's XiaoAI smart speakers to ChatGPT and Doubao, turning them into custom AI voice…
1012508abandoned
idootop/open-xiaoai
Open-XiaoAI is a Rust-based client/server project that takes over the audio input and output of Xiaomi XiaoAI smart speakers (LX06 and OH2P…
102595abandoned
C-Nedelcu/talk-to-chatgpt
A Chrome and Edge browser extension that lets users talk to ChatGPT using speech recognition and hear responses via text-to-speech, with op…
321929abandoned
wzpan/dingdang-robot
Dingdang is an open-source Chinese voice conversation robot / smart speaker project that runs on Raspberry Pi. It uses pluggable STT/TTS en…
101873abandoned
justLV/onju-voice
A hackable AI home assistant platform that replaces the internals of a Google Nest Mini with a custom ESP32-S3 PCB, paired with a server th…
631711abandoned
elevenlabs/elevenlabs-mcp
The official ElevenLabs Model Context Protocol (MCP) server, exposing ElevenLabs text-to-speech, speech-to-text, voice design, and conversa…
101532abandoned
fossasia/susi_smart_box
SUSI.AI Smart Box is an open-source smart speaker / voice assistant hardware project built around the SUSI.AI assistant. It provides the so…
101529abandoned
idootop/migpt-next
MiGPT-Next is a TypeScript library and Docker-runnable service that connects Xiaomi's XiaoAI smart speakers to OpenAI-compatible large lang…
101442abandoned
zenorocha/voice-elements
A pair of Polymer-based Web Components (<voice-player> and <voice-recognition>) that wrap the Web Speech API for speech synthesis (text to …
231348abandoned
dingdang-robot/dingdang-robot
Dingdang is an open-source Chinese voice conversation robot / smart speaker project that runs on Raspberry Pi and other Linux hosts. It use…
101330abandoned
MycroftAI/mimic3
Mimic 3 is a fast, local neural text-to-speech engine developed by Mycroft for the Mark II voice assistant, usable as a Python library, CLI…
291264abandoned
BrasD99/HeyGenClone
An open-source Python application that clones the HeyGen video translation system, translating videos into multiple languages with voice ov…
101027abandoned
Unsloth
Unsloth is a desktop application for running and fine-tuning LLMs, diffusion, embedding, and audio models locally, with support for NVIDIA,…
9474883active
calesthio/OpenMontage
OpenMontage is an open-source agentic video production system that turns AI coding assistants (Claude Code, Codex, Cursor, Copilot) into fu…
5951272active
mudler/LocalAI
LocalAI is an open-source, self-hosted AI inference runtime that runs LLMs, vision, speech, image, and video models behind OpenAI/Anthropic…
9348696active
lfnovo/open-notebook
Open Notebook is an open-source, privacy-focused alternative to Google's Notebook LM that lets users organize multi-modal sources (PDFs, vi…
8637711active
78/xiaozhi-esp32
XiaoZhi is an open-source MCP-based AI voice chatbot firmware for ESP32-family microcontrollers, connecting large language models like Qwen…
8829186active
openai/openai-agents-python
OpenAI Agents SDK is a lightweight Python framework for building multi-agent LLM workflows with agents, handoffs, tools, guardrails, sessio…
8228982active
facebookresearch/audiocraft
A PyTorch library from Meta for audio processing and generation with deep learning, featuring the EnCodec neural audio codec and generative…
6123586active
livekit/livekit
LiveKit is an open-source, scalable WebRTC SFU media server written in Go that provides realtime video, audio, and data transport for appli…
9920530stable
huggingface/transformers.js
Transformers.js is a JavaScript library that lets you run Hugging Face Transformers pretrained models directly in the browser (or Node.js) …
9116270active
alibaba/MNN
MNN is a lightweight, high-performance deep learning inference engine developed by Alibaba, supporting LLMs, vision, audio, and multimodal …
9315973active
pipecat-ai/pipecat
Pipecat is an open-source Python framework (BSD-2) for building real-time voice and multimodal conversational AI agents. It orchestrates sp…
8914769active
SubtitleEdit/subtitleedit
Subtitle Edit is a free, open-source desktop application for creating, editing, converting, and synchronizing subtitles, with video playbac…
9413971active
FujiwaraChoki/MoneyPrinter
MoneyPrinter is a Python application that automates the creation of YouTube Shorts from a user-provided video topic, using local Ollama mod…
5913891active
xszyou/Fay
Fay is an open-source Python digital human framework that connects 2.5D/3D/mobile/web digital humans or OpenAI-compatible LLMs to business …
7513455active
ace-step/ACE-Step-1.5
ACE-Step 1.5 is an open-source music generation foundation model combining a language model planner with a Diffusion Transformer to create …
7912421active
h2oai/h2ogpt
h2oGPT is an Apache-2.0 open-source application for chatting with local, private LLMs and querying/summarizing your own documents (PDFs, Wo…
1011969active
speechbrain/speechbrain
SpeechBrain is an open-source PyTorch-based speech toolkit for building conversational AI systems. It provides training recipes, pretrained…
8311785active
Baiyuetribe/paper2gui
Paper2GUI (now branded 小白兔AI / Xiaobaitu AI) is a desktop AI toolbox application that packages 50+ AI research models into install-free GUI…
2310709active
VoltAgent/voltagent
VoltAgent is an open-source TypeScript framework for building production-ready AI agents with memory, tools, RAG, guardrails, MCP, voice, a…
8010425active
lipku/LiveTalking
LiveTalking is an open-source real-time interactive streaming digital human engine that drives a talking-head avatar from text or audio wit…
899238active
fonoster/fonoster
Fonoster is an open-source programmable telecommunications stack (a Twilio alternative) for building voice and messaging applications, with…
998080active
a-ghorbani/pocketpal-ai
PocketPal AI is an open-source mobile app that runs GGUF language models entirely on-device for private, offline chat. It supports model do…
858067active
enricoros/big-AGI
Big-AGI is an open-source AI workspace web application that lets users chat with many state-of-the-art LLM providers (OpenAI, Anthropic, Ge…
937102active
PaddlePaddle/models
PaddlePaddle's officially maintained industry-grade model repository containing 600+ models across computer vision, NLP, speech, recommenda…
236932active
gaozhangmin/boxplayer
BoxPlayer is a cross-platform desktop application that unifies multiple cloud drives (Aliyun Drive, Baidu, 115, Quark, OneDrive, etc.), loc…
976878active
Dooy/chatgpt-web-midjourney-proxy
A unified web/desktop UI for ChatGPT plus AI image, music, and video generation services like Midjourney, Suno, Luma, Runway, and Flux. It …
766785active
multimodal-art-projection/YuE
YuE is a family of open-source foundation models based on the LLaMA2 architecture that generate full songs (up to five minutes) with vocals…
326403active
vllm-project/vllm-omni
vLLM-Omni is a Python framework extending vLLM for efficient inference and serving of omni-modality models, including diffusion transformer…
836369active
aidlearning/AidLearning-FrameWork
AidLux (originally AidLearning) is an AIoT development platform that runs a native Ubuntu Linux environment with GUI, deep learning tooling…
705797active
modstart-lib/aigcpanel
AIGCPanel is an open-source, all-in-one AI digital human desktop application built with TypeScript, Vue3, and Electron for Windows, macOS, …
885482active
lemonade-sdk/lemonade
Lemonade is a local AI server that runs optimized LLMs (plus image, speech, and embedding models) on your own GPU and NPU, exposing OpenAI-…
825472active
signalwire/freeswitch
FreeSWITCH is an open-source software-defined telecom stack that turns commodity servers into a full telephony platform for voice, video, a…
945118stable
MoonshotAI/Kimi-Audio
Kimi-Audio is an open-source audio foundation model (7B parameters) that unifies audio understanding, generation, and speech conversation i…
314731active
Osmantic/ODS
ODS (Osmantic Deployment System) is a self-hosted AI server installer and runtime that turns a PC, Mac, or Linux machine into a private AI …
804716active
ArcReel/ArcReel
ArcReel is an open-source, self-hosted AI video production workspace that turns novels, scripts, or product material into characters, scene…
824211active
QwenLM/Qwen2.5-Omni
Qwen2.5-Omni is an end-to-end multimodal model from Alibaba's Qwen team that understands text, images, audio, and video, and generates stre…
314074active
OpenMinis/OpenMinis
OpenMinis is a free, open-source mobile AI agent app for iOS and Android that connects leading LLM providers (Claude, GPT, Gemini, etc.) vi…
693985active
claraverse-space/ClaraVerse
ClaraVerse is a self-hosted, privacy-focused AI workspace that combines chat, multi-agent teams, a visual workflow builder, RAG document pr…
803894active
zly2006/zhihu-plus-plus
Zhihu++ is an open-source third-party Android client for the Chinese Q&A platform Zhihu, built in Kotlin, that removes ads, promotional pos…
853845active
tbxark/ChatGPT-Telegram-Workers
A Telegram ChatGPT bot deployable on Cloudflare Workers, Vercel, or Docker with a single-file, dependency-free setup. It supports multiple …
663811active
Chevey339/kelivo
Kelivo is an open-source, cross-platform LLM chat client built with Flutter that runs on Android, iOS, HarmonyOS, Windows, macOS, and Linux…
843766active
OvidijusParsiunas/deep-chat
Deep Chat is a fully customizable, framework-agnostic AI chatbot web component that can be added to any website with one line of code. It s…
913704active
sukeesh/Jarvis
Jarvis is a command-line personal assistant for Linux, macOS, and Windows written in Python. It offers 15+ task categories including weathe…
483645active
sligter/LandPPT
LandPPT is an AI-powered presentation generation platform that turns a topic or uploaded documents (PDF, Word, Markdown, Excel, PPT) into p…
813575active
n3d1117/chatgpt-telegram-bot
A self-hosted Telegram bot written in Python that integrates with OpenAI's official ChatGPT, DALL·E, and Whisper APIs to answer questions, …
333460active
Grt1228/chatgpt-java
An unofficial Java SDK for the OpenAI API covering all official endpoints including chat completions (GPT-3.5/GPT-4), DALL-E image generati…
213423active
chenyme/Chenyme-AAVT
Chenyme-AAVT is a fully automated audio/video translation application that uses Whisper (faster-whisper) for speech recognition, large lang…
103128active
Snailclimb/interview-guide
An open-source AI-powered interview platform built with Spring Boot 4.1, Java 25, Spring AI 2.0, React, PostgreSQL/pgvector, and Redis. It …
593106active
TanStack/ai
TanStack AI is a type-safe, provider-agnostic TypeScript SDK for building AI applications with streaming chat, tool calling, agents, struct…
803028active
SharpAI/DeepCamera
DeepCamera is an open-source AI camera skills platform that runs local VLM scene analysis, object detection, face recognition, and person r…
863019active
MacPaw/OpenAI
A community-maintained Swift package that wraps the OpenAI public API, supporting chat completions, responses, function calling, MCP tools,…
942935active
OpenMind/OM1
OM1 is a modular AI runtime and hardware abstraction layer for building multimodal AI agents that run on physical robots and in simulators.…
822897active
datascale-ai/opentalking
OpenTalking is an open-source Python framework for building real-time AI digital-human (talking avatar) conversation products. It orchestra…
652897active
luoluoluo22/jianying-editor-skill
A Python-based AI agent skill that automates video editing in JianYing (the Chinese desktop version of CapCut) by directly generating and i…
552878active
cdhigh/KindleEar
KindleEar is a self-hostable Python web application that aggregates RSS/ATOM/JSON feeds and web content (including Calibre recipes) into ep…
732865active
sahibzada-allahyar/YC-Killer
A collection of open-source, enterprise-grade AI agents intended as free alternatives to commercial Y Combinator startups. It includes a de…
642792active
ClassIsland/ClassIsland
ClassIsland is a cross-platform desktop application that displays class schedules and related information on classroom multimedia screens, …
922709active
Vali-98/ChatterUI
ChatterUI is a native mobile frontend for running large language models on-device via llama.cpp or connecting to commercial and open-source…
832689active
Project-N-E-K-O/N.E.K.O
Project N.E.K.O. is an open-source AI companion application — a proactive catgirl-style AI that lives on your desktop, initiates interactio…
852681active

← prev page 4 / 5 next →