# bytedance/SALMONN

SALMONN family: A suite of advanced multi-modal LLMs

Repository: https://github.com/bytedance/SALMONN
Canonical: https://ross.abutalabs.com/products/salmonn
Homepage: https://bytedance.github.io/SALMONN/
License: Apache-2.0
License Family: permissive
Topics: audio, audio-processing, large-language-models, multi-modal, speech, speech-recognition, bytedance, tsinghua-university, music, iclr2024, research, icml-2024, video, video-understanding, audio-visual-understanding
Last push: 2026-08-24T07:11:01+00:00

## Health v2 (maintenance only)
Score: 73/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 99, release rhythm 35, longevity 79
- inputs: {"age_days": 1118, "days_push": 9, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1513, forks 124 (observed 2026-08-28T04:04:56.213063+00:00)

## What it is
SALMONN is a family of open-source multi-modal large language models from ByteDance and Tsinghua that unify speech, audio, music, and video understanding with text. The repository hosts model checkpoints, inference code, training scripts, and related research releases such as video-SALMONN and speech quality assessment models.

## Use cases
- build an audio LLM that understands speech and music
- generate captions for videos with audio-visual understanding
- evaluate speech quality with natural language reasoning
- run inference with a multimodal speech and vision LLM
- fine-tune my own multimodal LLM on audio data
- benchmark audio LLMs on speech perception tasks

## When to choose
- you need a research-grade open model combining speech, audio, music, and video understanding
- you want pretrained checkpoints and full training code for multimodal LLM experiments
- you need speech quality assessment with natural language reasoning

## When to avoid
- you need a production-ready commercial speech API with support guarantees
- you lack GPU resources for large multimodal model inference or training
- you only need simple text-only LLM functionality

## Facets
- artifact type: library
- maturity: active
- function: machine-learning, speech-recognition, audio-processing, video-processing, llm-inference, llm-training
- domain: large-language-models, speech-processing, artificial-intelligence
- platform: python, cross-platform
- tags: multimodal-llm, audio-visual, speech-understanding, research-models, model-checkpoints, iclr, icml, audio, video, natural-language-processing, gpu

## Member repositories
- bytedance/SALMONN (main) score 73

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:04:56.213063+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T04:32:14.892759+00:00, confidence not recorded.
  - readme: https://github.com/bytedance/SALMONN (fetched 2026-08-28T04:04:56.213063+00:00, sha 0e77e160fdcf)
  - homepage: https://bytedance.github.io/SALMONN/ (fetched 2026-08-29T11:36:22.938175+00:00, sha 519b7cb97e27)
- Data as of 2026-08-30T08:39:29.467469+00:00.
