Ross ROSS = Recommend OSS · open-source software intelligence for agents

bytedance/SALMONN

SALMONN family: A suite of advanced multi-modal LLMs observed · 2026-08-28

github.com/bytedance/SALMONN · homepage · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

73/100

  • Activity 99
  • Release rhythm 35
  • Longevity 79

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1118
  • days_rel: n/a
  • days_push: 9
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1513 stars · 124 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

SALMONN is a family of open-source multi-modal large language models from ByteDance and Tsinghua that unify speech, audio, music, and video understanding with text. The repository hosts model checkpoints, inference code, training scripts, and related research releases such as video-SALMONN and speech quality assessment models.

Use cases

  • build an audio LLM that understands speech and music
  • generate captions for videos with audio-visual understanding
  • evaluate speech quality with natural language reasoning
  • run inference with a multimodal speech and vision LLM
  • fine-tune my own multimodal LLM on audio data
  • benchmark audio LLMs on speech perception tasks

When to choose

  • you need a research-grade open model combining speech, audio, music, and video understanding
  • you want pretrained checkpoints and full training code for multimodal LLM experiments
  • you need speech quality assessment with natural language reasoning

When to avoid

  • you need a production-ready commercial speech API with support guarantees
  • you lack GPU resources for large multimodal model inference or training
  • you only need simple text-only LLM functionality

Facets

library · maturity active

machine-learning speech-recognition audio-processing video-processing llm-inference llm-training large-language-models speech-processing artificial-intelligence python cross-platform multimodal-llm audio-visual speech-understanding research-models model-checkpoints iclr icml audio video natural-language-processing gpu

2 sources

Member repositories

RepositoryRoleHealth v2
bytedance/SALMONNmain73

For agents

markdown · JSON · MCP: product_card(name="bytedance/SALMONN")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem