jingyaogong/minimind-o resource
🎙️ A 0.1B Omni model trained from scratch, capable of listening, speaking, and seeing! observed · 2026-08-28
Health v2 · maintenance only
57/100
- Activity 96
- Release rhythm 35
- Longevity 8
Flags: no_releases young
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 124
- days_rel: n/a
- days_push: 27
- n_releases_24m: 0
Adoption not part of the score
2384 stars · 280 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
MiniMind-O is an open-source educational project that implements a tiny (~0.1B parameter) end-to-end Omni model from scratch in PyTorch, handling text, image, and audio inputs with text and streaming speech outputs. It includes full training code, model weights, mini/full datasets, and a technical report, designed so a single RTX 3090 can run the complete training pipeline in about 2 hours.
Use cases
- train a small multimodal omni model from scratch
- learn how thinker-talker speech architectures work
- build a model that listens, sees, and speaks
- experiment with streaming voice generation and barge-in interruption
- study voice cloning with reference audio codes
- run a lightweight GPT-4o-style model on a personal GPU
- understand end-to-end speech-text hidden state fusion
When to choose
- you want to read, modify, and train a complete omni model from first principles
- you have a single consumer GPU and limited time/budget
- you need a readable baseline for multimodal speech research or teaching
- you want full open access to code, weights, data, and a technical report
When to avoid
- you need production-grade speech quality or low-latency real-time deployment
- you want a plug-and-play inference library rather than a training/learning codebase
- you need large-scale model capacity or broad multilingual coverage
- you require enterprise support or long-term stability guarantees
Facets
learning-resource · maturity active
machine-learning deep-learning llm-training speech-recognition tts chatbot audio-processing computer-vision artificial-intelligence large-language-models speech-processing computer-vision education tutorials python cross-platform omni-model multimodal thinker-talker streaming-voice voice-cloning train-from-scratch pytorch mtp small-language-model natural-language-processing gpu
2 sources
- readme: https://github.com/jingyaogong/minimind-o · fetched 2026-08-28 · 2a77c6efbcbd
- homepage: https://jingyaogong.github.io/minimind-o · fetched 2026-08-29 · 2cf794458036
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| jingyaogong/minimind-o | main | 57 |
For agents
markdown · JSON · MCP: product_card(name="jingyaogong/minimind-o")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem