apple/ml-4m
4M: Massively Multimodal Masked Modeling observed · 2026-08-28
Health v2 · maintenance only
35/100
- Activity 24
- Release rhythm 35
- Longevity 62
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 877
- days_rel: n/a
- days_push: 457
- n_releases_24m: 0
Adoption not part of the score
1808 stars · 114 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
4M is a framework from Apple and EPFL for training any-to-any multimodal foundation models using masked modeling over discrete tokens across tens of modalities. It includes official implementations, training code, tokenizers, and pre-trained models (4M-7, 4M-21) for versatile vision and multimodal generation tasks.
Use cases
- train any-to-any multimodal foundation models
- run a single vision model across tens of tasks and modalities
- generate images conditioned on text or other modalities
- perform multimodal image editing and in-painting
- fine-tune a generalist vision model for downstream tasks
- tokenize images, text, geometry, and semantics into discrete tokens
- download pretrained multimodal transformer checkpoints
When to choose
- you need one unified model handling many vision modalities and tasks
- you want to experiment with multimodal masked modeling research
- you need steerable multimodal generation or editing capabilities
- you want pretrained any-to-any vision models with open weights
When to avoid
- you need a lightweight single-task vision model
- you lack GPU resources for large transformer training or inference
- you need production-ready multimodal APIs rather than a research framework
- your use case is text-only NLP
Facets
library · maturity active
machine-learning deep-learning image-processing nlp data-science machine-learning computer-vision deep-learning artificial-intelligence image-processing python multimodal masked-modeling foundation-models transformer any-to-any tokenization pretrained-models vision-models generative-models pytorch gpu linux
2 sources
- readme: https://github.com/apple/ml-4m · fetched 2026-08-28 · cd7298de3e48
- homepage: https://4m.epfl.ch · fetched 2026-08-29 · c8386effc6c8
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| apple/ml-4m | main | 35 |
For agents
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem