Ross ROSS = Recommend OSS · open-source software intelligence for agents

xid32/SoundMind resource

We introduce the Audio Logical Reasoning (ALR) dataset, consisting of 6,446 text-audio annotated samples specifically designed for complex reasoning tasks. Building on this resource, we propose SoundMind, a rule-based reinforcement learning (RL) algorithm tailored to endow audio language models (ALMs) with deep bimodal reasoning abilities. observed · 2026-08-28

github.com/xid32/SoundMind · homepage · Python · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

43/100

  • Activity 54
  • Release rhythm 35
  • Longevity 31

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 446
  • days_rel: n/a
  • days_push: 280
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1113 stars · 131 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

SoundMind is an Audio Logical Reasoning (ALR) dataset of 6,446 audio-text annotated samples with chain-of-thought reasoning, paired with a rule-based reinforcement learning training framework for audio-language models built on verl. It was used to fine-tune Qwen2.5-Omni-7B, achieving strong improvements on audio reasoning benchmarks.

Use cases

  • train audio language models for logical reasoning with reinforcement learning
  • download an annotated audio-text reasoning dataset with chain-of-thought
  • benchmark audio-language models on complex reasoning tasks
  • fine-tune Qwen2.5-Omni with rule-based RL
  • research bimodal audio-text reasoning in LLMs

When to choose

  • you need reasoning-oriented audio training data with chain-of-thought annotations
  • you want to apply rule-based RL to audio-language models using verl
  • you are evaluating audio models on logical reasoning benchmarks

When to avoid

  • you lack multi-GPU H100/H800-class hardware for training
  • you need general audio transcription or speech recognition rather than reasoning
  • you want a production-ready inference service rather than a research codebase

Facets

dataset · maturity active

machine-learning llm-training audio-processing speech-recognition machine-learning artificial-intelligence python audio-language-model reinforcement-learning chain-of-thought benchmark qwen2.5-omni verl emnlp-2025 audio natural-language-processing gpu linux

6 sources

Member repositories

RepositoryRoleHealth v2
xid32/SoundMindmain43

For agents

markdown · JSON · MCP: product_card(name="xid32/SoundMind")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem