# om-ai-lab/VLM-R1

Solve Visual Understanding with Reinforced VLMs

Repository: https://github.com/om-ai-lab/VLM-R1
Canonical: https://ross.abutalabs.com/products/vlm-r1
Language: Python
License: Apache-2.0
License Family: permissive
Topics: deepseek-r1, grpo, llm, multimodal, vlm, qwen, vlm-r1, multimodal-r1, r1-zero, reinforcement-learning
Last push: 2026-07-07T01:50:29+00:00

## Health v2 (maintenance only)
Score: 63/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 91, release rhythm 40, longevity 40
- inputs: {"age_days": 573, "days_push": 58, "days_rel": 505, "gap_med": 14.0, "n_releases_24m": 3}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 6015, forks 383 (observed 2026-08-28T04:09:34.275204+00:00)

## What it is
VLM-R1 is a framework for training R1-style large vision-language models using reinforcement learning (GRPO) on top of Qwen2.5-VL. It provides training scripts for full and LoRA fine-tuning, multi-node training, and released checkpoints for tasks like referring expression comprehension and open-vocabulary detection.

## Use cases
- fine-tune a vision-language model with GRPO reinforcement learning
- train a VLM for referring expression comprehension
- improve out-of-domain generalization of visual reasoning models
- run LoRA fine-tuning on Qwen2.5-VL
- reproduce R1-style RL training for multimodal models
- train an open-vocabulary detection model

## When to choose
- you want to apply RL post-training (GRPO) to vision-language models
- you need better out-of-domain generalization than SFT provides
- you want proven checkpoints for REC or open-vocabulary detection

## When to avoid
- you only need to run inference with a VLM without training
- you lack multi-GPU resources for RL fine-tuning
- you need simple supervised fine-tuning only

## Facets
- artifact type: library
- maturity: active
- function: reinforcement-learning, llm-training, machine-learning, computer-vision
- domain: deep-learning, large-language-models, computer-vision, machine-learning
- platform: python
- tags: vlm, grpo, deepseek-r1, qwen, multimodal, fine-tuning, visual-reasoning, referring-expression-comprehension, gpu, linux

## Member repositories
- om-ai-lab/VLM-R1 (main) score 63

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:09:34.275204+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T17:49:45.918804+00:00, confidence not recorded.
  - readme: https://github.com/om-ai-lab/VLM-R1 (fetched 2026-08-28T04:09:34.275204+00:00, sha 02cb4ee89bb8)
- Data as of 2026-08-30T08:39:29.467469+00:00.
