OFA-Sys/ONE-PEACE
A general representation model across vision, audio, language modalities. Paper: ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities observed · 2026-08-28
Health v2 · maintenance only
29/100
- Activity 0
- Release rhythm 35
- Longevity 85
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 1203
- days_rel: n/a
- days_push: 696
- n_releases_24m: 0
Adoption not part of the score
1060 stars · 71 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
ONE-PEACE is a general multimodal representation model that jointly encodes vision, audio, and language modalities without initializing from pretrained vision or language models. The repository provides pretrained checkpoints, fine-tuning and inference scripts, embedding and visual grounding APIs, and a multimodal retrieval demo.
Use cases
- extract joint embeddings for images, audio, and text
- retrieve images using audio, text, or combined audio+text+image queries
- fine-tune a multimodal model on vision-language tasks like VQA and captioning
- fine-tune on audio classification and audio-language tasks
- locate objects in images with visual grounding
- pretrain a modality-agnostic transformer from scratch
When to choose
- you need a single model embedding multiple modalities in one shared space
- you want strong zero-shot cross-modal retrieval including unpaired modality combinations
- you need a research foundation model for multimodal representation learning
When to avoid
- you need a lightweight production embedding service with minimal dependencies
- you only need unimodal text or image models
- you require active community support and frequent updates
Facets
library · maturity maintenance
machine-learning deep-learning nlp image-processing audio-processing search-engine machine-learning deep-learning artificial-intelligence computer-vision python multimodal representation-learning foundation-model vision-language contrastive-learning zero-shot-retrieval embeddings natural-language-processing audio
1 source
- readme: https://github.com/OFA-Sys/ONE-PEACE · fetched 2026-08-28 · 75487d7ea45f
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| OFA-Sys/ONE-PEACE | main | 29 |
For agents
markdown · JSON · MCP: product_card(name="OFA-Sys/ONE-PEACE")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem