# HITsz-TMG/Uni-MoE

Uni-MoE: Lychee's Large Multimodal Model Family.

Repository: https://github.com/HITsz-TMG/Uni-MoE
Canonical: https://ross.abutalabs.com/products/uni-moe
Homepage: https://idealistxy.github.io/Uni-MoE-v2.github.io/
Language: Python
License Family: other
Last push: 2026-08-06T10:50:45+00:00

## Health v2 (maintenance only)
Score: 68/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 96, release rhythm 35, longevity 65
- inputs: {"age_days": 912, "days_push": 27, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1116, forks 71 (observed 2026-08-28T04:03:38.726688+00:00)

## What it is
Uni-MoE is a family of open-source Mixture-of-Experts (MoE) based omnimodal large language models that understand and generate across text, images, speech, audio, and video. The repository provides model weights, training code, and evaluation integration (e.g., LMMs-Eval) for versions including Uni-MoE-2.0-Omni built on Qwen2.5-7B.

## Use cases
- run an omnimodal LLM that understands images, speech, and video
- generate speech, images, and text from a single unified model
- fine-tune a MoE multimodal model on custom data
- evaluate a multimodal LLM with lmms-eval
- research mixture-of-experts architectures for multimodal learning
- convert speech to text and text to speech with one model

## When to choose
- you need a single open model handling cross-modal understanding and generation
- you want to study or extend MoE-based multimodal architectures
- you need audio generation unifying speech and music

## When to avoid
- you need a lightweight model for CPU-only or edge deployment
- you need a commercially licensed model (no license specified)
- you only need text-only LLM inference with minimal setup

## Facets
- artifact type: library
- maturity: active
- function: machine-learning, deep-learning, llm-training, speech-recognition, tts, image-processing, video-processing, audio-processing
- domain: large-language-models, deep-learning, artificial-intelligence, speech-processing
- platform: python
- tags: mixture-of-experts, multimodal, omnimodal, model-weights, research, natural-language-processing, gpu, linux

## Member repositories
- HITsz-TMG/Uni-MoE (main) score 68

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:03:38.726688+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T06:41:56.095490+00:00, confidence not recorded.
  - readme: https://github.com/HITsz-TMG/Uni-MoE (fetched 2026-08-28T04:03:38.726688+00:00, sha 1c7d5cfa87c4)
  - homepage: https://idealistxy.github.io/Uni-MoE-v2.github.io/ (fetched 2026-08-29T12:45:54.776598+00:00, sha 5c8872ee5448)
- Data as of 2026-08-30T08:39:29.467469+00:00.
