Ross ROSS = Recommend OSS · open-source software intelligence for agents

jingyaogong/minimind-o resource

🎙️ A 0.1B Omni model trained from scratch, capable of listening, speaking, and seeing! observed · 2026-08-28

github.com/jingyaogong/minimind-o · homepage · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

57/100

  • Activity 96
  • Release rhythm 35
  • Longevity 8

Flags: no_releases young

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 124
  • days_rel: n/a
  • days_push: 27
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

2384 stars · 280 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

MiniMind-O is an open-source educational project that implements a tiny (~0.1B parameter) end-to-end Omni model from scratch in PyTorch, handling text, image, and audio inputs with text and streaming speech outputs. It includes full training code, model weights, mini/full datasets, and a technical report, designed so a single RTX 3090 can run the complete training pipeline in about 2 hours.

Use cases

  • train a small multimodal omni model from scratch
  • learn how thinker-talker speech architectures work
  • build a model that listens, sees, and speaks
  • experiment with streaming voice generation and barge-in interruption
  • study voice cloning with reference audio codes
  • run a lightweight GPT-4o-style model on a personal GPU
  • understand end-to-end speech-text hidden state fusion

When to choose

  • you want to read, modify, and train a complete omni model from first principles
  • you have a single consumer GPU and limited time/budget
  • you need a readable baseline for multimodal speech research or teaching
  • you want full open access to code, weights, data, and a technical report

When to avoid

  • you need production-grade speech quality or low-latency real-time deployment
  • you want a plug-and-play inference library rather than a training/learning codebase
  • you need large-scale model capacity or broad multilingual coverage
  • you require enterprise support or long-term stability guarantees

Facets

learning-resource · maturity active

machine-learning deep-learning llm-training speech-recognition tts chatbot audio-processing computer-vision artificial-intelligence large-language-models speech-processing computer-vision education tutorials python cross-platform omni-model multimodal thinker-talker streaming-voice voice-cloning train-from-scratch pytorch mtp small-language-model natural-language-processing gpu

2 sources

Member repositories

RepositoryRoleHealth v2
jingyaogong/minimind-omain57

For agents

markdown · JSON · MCP: product_card(name="jingyaogong/minimind-o")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem