marin-community/marin
Open-source framework for the research and development of foundation models. observed · 2026-08-28
Health v2 · maintenance only
78/100
- Activity 99
- Release rhythm 60
- Longevity 63
Flags: prerelease_only
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 0
- age_days: 894
- days_rel: 20
- days_push: 7
- n_releases_24m: 16
Adoption not part of the score
2428 stars · 212 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
Marin is an open-source Python framework and research program for training foundation models, covering the full pipeline from data curation, filtering, and tokenization through pretraining, posttraining, and evaluation. It emphasizes open development, documenting every experiment, dataset, and checkpoint in real time, and includes artifacts like the Delphi scaling suite and frontier mixture-of-experts training runs.
Use cases
- pretrain a large language model from scratch on TPUs
- curate and filter pretraining datasets like Nemotron-CC or StarCoderData
- run scaling law experiments across compute budgets
- run posttraining and SFT experiments on open LLMs
- evaluate and compare LLM checkpoints reproducibly
- train audio-text, DNA, or protein foundation models
- reproduce and share ML experiments with full provenance
When to choose
- you want an open, reproducible pipeline for LLM pretraining and posttraining research
- you need data curation, tokenization, training, and evaluation in one framework
- you want to run scaling suites or mixture-of-experts experiments on TPU research clouds
- you value open science with preregistered experiments and shared checkpoints
When to avoid
- you just need to fine-tune a small model quickly with a high-level API like HF Trainer or LoRA tooling
- you need production inference serving rather than training research
- you lack access to significant TPU/GPU compute for large-scale training
Facets
framework · maturity active
llm-training machine-learning etl data-science benchmarking large-language-models machine-learning deep-learning artificial-intelligence python cloud foundation-models pretraining posttraining data-curation scaling-laws tpu open-science mixture-of-experts tokenization evaluation data-engineering gpu linux
2 sources
- readme: https://github.com/marin-community/marin · fetched 2026-08-28 · 7f6637dd6bb4
- homepage: https://marin.community · fetched 2026-08-29 · 448181b97c23
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| marin-community/marin | main | 78 |
For agents
markdown · JSON · MCP: product_card(name="marin-community/marin")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem