# Visual-Agent/DeepEyes

Repository: https://github.com/Visual-Agent/DeepEyes
Canonical: https://ross.abutalabs.com/products/deepeyes
Language: Python
License: Apache-2.0
License Family: permissive
Last push: 2025-11-20T11:57:10+00:00

## Health v2 (maintenance only)
Score: 44/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 53, release rhythm 35, longevity 38
- inputs: {"age_days": 536, "days_push": 286, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1271, forks 78 (observed 2026-08-28T04:04:11.916317+00:00)

## What it is
DeepEyes is a research project that trains multimodal vision-language models to 'think with images' using end-to-end reinforcement learning, built on the VeRL framework with Qwen-2.5-VL as the base model. It includes training code, a 47k dataset, and released 7B model checkpoints.

## Use cases
- train a vision-language model with reinforcement learning
- reproduce thinking-with-images RL experiments
- improve visual grounding of a multimodal LLM
- fine-tune Qwen-2.5-VL with outcome-reward RL
- research tool-calling behavior emerging from RL training

## When to choose
- you need to RL-train a multimodal model to reason over images
- you want to reproduce the DeepEyes paper results
- you have a multi-node GPU cluster for 7B+ model training

## When to avoid
- you just want to run inference with an existing VLM
- you lack large-scale GPU resources (32+ GPUs recommended)
- you need a production-ready application rather than research code

## Facets
- artifact type: library
- maturity: active
- function: reinforcement-learning, machine-learning, llm-training, computer-vision, agent-framework
- domain: artificial-intelligence, machine-learning, computer-vision, large-language-models, deep-learning
- platform: python
- tags: multimodal, vision-language-model, rl-training, qwen, thinking-with-images, research, gpu, linux

## Member repositories
- Visual-Agent/DeepEyes (main) score 44

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:04:11.916317+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T05:03:21.083936+00:00, confidence not recorded.
  - readme: https://github.com/Visual-Agent/DeepEyes (fetched 2026-08-28T04:04:11.916317+00:00, sha e958097fbe88)
- Data as of 2026-08-30T08:39:29.467469+00:00.
