Ross ROSS = Recommend OSS · open-source software intelligence for agents

Alpha-VLLM/Lumina-T2X

Lumina-T2X is a unified framework for Text to Any Modality Generation observed · 2026-08-28

github.com/Alpha-VLLM/Lumina-T2X · Python · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

28/100

  • Activity 7
  • Release rhythm 35
  • Longevity 63

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 888
  • days_rel: n/a
  • days_push: 563
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

2250 stars · 98 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

Lumina-T2X is a unified framework for text-to-any-modality generation built on flow-based large diffusion transformers. It supports generating images, videos, audio, speech, and 3D shapes from text prompts at varying resolutions and durations.

Use cases

  • generate images from text prompts
  • text to video generation
  • synthesize speech or audio from text
  • generate 3D shapes from a text description
  • train or fine-tune diffusion transformer models
  • research flow-based generative models

When to choose

  • you need a single framework for multiple output modalities (image, video, audio, 3D)
  • you want state-of-the-art diffusion transformer research code with pretrained checkpoints
  • you are doing academic research on flow-based generative models

When to avoid

  • you need a production-ready, polished end-user application
  • you lack a high-end GPU, as training and inference are compute-heavy
  • you only need simple text-to-image via a hosted API

Facets

library · maturity active

machine-learning deep-learning image-processing video-processing speech-recognition llm-training artificial-intelligence deep-learning image-processing python diffusion-transformer text-to-image text-to-video text-to-audio flow-matching aigc multimodal-generation research video natural-language-processing gpu linux

1 source

Member repositories

RepositoryRoleHealth v2
Alpha-VLLM/Lumina-T2Xmain28

For agents

markdown · JSON · MCP: product_card(name="Alpha-VLLM/Lumina-T2X")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem