xid32/SoundMind resource
We introduce the Audio Logical Reasoning (ALR) dataset, consisting of 6,446 text-audio annotated samples specifically designed for complex reasoning tasks. Building on this resource, we propose SoundMind, a rule-based reinforcement learning (RL) algorithm tailored to endow audio language models (ALMs) with deep bimodal reasoning abilities. observed · 2026-08-28
Health v2 · maintenance only
43/100
- Activity 54
- Release rhythm 35
- Longevity 31
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 446
- days_rel: n/a
- days_push: 280
- n_releases_24m: 0
Adoption not part of the score
1113 stars · 131 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
SoundMind is an Audio Logical Reasoning (ALR) dataset of 6,446 audio-text annotated samples with chain-of-thought reasoning, paired with a rule-based reinforcement learning training framework for audio-language models built on verl. It was used to fine-tune Qwen2.5-Omni-7B, achieving strong improvements on audio reasoning benchmarks.
Use cases
- train audio language models for logical reasoning with reinforcement learning
- download an annotated audio-text reasoning dataset with chain-of-thought
- benchmark audio-language models on complex reasoning tasks
- fine-tune Qwen2.5-Omni with rule-based RL
- research bimodal audio-text reasoning in LLMs
When to choose
- you need reasoning-oriented audio training data with chain-of-thought annotations
- you want to apply rule-based RL to audio-language models using verl
- you are evaluating audio models on logical reasoning benchmarks
When to avoid
- you lack multi-GPU H100/H800-class hardware for training
- you need general audio transcription or speech recognition rather than reasoning
- you want a production-ready inference service rather than a research codebase
Facets
dataset · maturity active
machine-learning llm-training audio-processing speech-recognition machine-learning artificial-intelligence python audio-language-model reinforcement-learning chain-of-thought benchmark qwen2.5-omni verl emnlp-2025 audio natural-language-processing gpu linux
6 sources
- readme: https://github.com/xid32/SoundMind · fetched 2026-08-28 · 4591bc772939
- homepage: https://arxiv.org/abs/2506.12935 · fetched 2026-08-29 · 10a06cc3b52b
- site_page: https://info.arxiv.org/about/donate.html · fetched 2026-08-29 · cca9c3a11c56
- site_page: https://info.arxiv.org/about/ourmembers.html · fetched 2026-08-29 · 47cbc55ff1de
- site_page: https://info.arxiv.org/about · fetched 2026-08-29 · a1f16f915a9a
- site_page: https://info.arxiv.org/labs/index.html · fetched 2026-08-29 · b14a8d05a0ec
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| xid32/SoundMind | main | 43 |
For agents
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem