# SkyworkAI/Skywork-R1V

Skywork-R1V is an advanced multimodal AI model series developed by Skywork AI, specializing in vision-language reasoning.

Repository: https://github.com/SkyworkAI/Skywork-R1V
Canonical: https://ross.abutalabs.com/products/skywork-r1v
Homepage: https://arxiv.org/abs/2504.05599
Language: Python
License: MIT
License Family: permissive
Topics: deepseek-r1, llm, r1v, reasoning, skywork-r1v, multimodal-understanding, multimodal-r1, vlm, grpo, reinforcement-learning, vlm-r1
Last push: 2026-07-29T03:11:37+00:00

## Health v2 (maintenance only)
Score: 63/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 95, release rhythm 35, longevity 38
- inputs: {"age_days": 536, "days_push": 35, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 3167, forks 282 (observed 2026-08-28T04:07:47.311183+00:00)

## What it is
Skywork-R1V is a series of open-source multimodal reasoning models (up to 38B parameters) that extend R1-style LLMs to visual inputs via a lightweight projector and reinforcement finetuning with GRPO. The repository provides model weights, inference code, and evaluation scripts for reproducing benchmark results like MMMU and MathVista.

## Use cases
- run a multimodal reasoning model on images and text
- reproduce MMMU and MathVista benchmark results
- fine-tune a vision-language model with GRPO reinforcement learning
- deploy an open-source visual chain-of-thought model
- compare open multimodal reasoning models against GPT-4o
- run quantized 38B VLM inference on a single 30GB GPU

## When to choose
- you need an open-weights multimodal reasoning model with strong benchmark scores
- you want to study or extend RL post-training (GRPO) for vision-language models
- you need visual chain-of-thought reasoning with reproducible eval code

## When to avoid
- you need a lightweight model for CPU or edge devices
- you only need text-only LLM inference
- you want a turnkey hosted API rather than self-hosted weights

## Facets
- artifact type: learning-resource
- maturity: active
- function: machine-learning, deep-learning, llm-inference, reinforcement-learning, computer-vision, nlp
- domain: large-language-models, deep-learning, computer-vision, artificial-intelligence
- platform: python
- tags: multimodal, vision-language-model, chain-of-thought, grpo, model-weights, reasoning, vlm, awq-quantization, natural-language-processing, gpu, linux

## Member repositories
- SkyworkAI/Skywork-R1V (main) score 63

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:07:47.311183+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T18:45:31.636010+00:00, confidence not recorded.
  - readme: https://github.com/SkyworkAI/Skywork-R1V (fetched 2026-08-28T04:07:47.311183+00:00, sha 0be86484e3fe)
  - homepage: https://arxiv.org/abs/2504.05599 (fetched 2026-08-29T09:39:49.622781+00:00, sha 3da542eb5fc4)
  - site_page: https://info.arxiv.org/about/donate.html (fetched 2026-08-29T09:39:49.626598+00:00, sha cca9c3a11c56)
  - site_page: https://info.arxiv.org/about/ourmembers.html (fetched 2026-08-29T09:39:49.629850+00:00, sha 47cbc55ff1de)
  - site_page: https://info.arxiv.org/about (fetched 2026-08-29T09:39:49.631844+00:00, sha a1f16f915a9a)
  - site_page: https://info.arxiv.org/labs/index.html (fetched 2026-08-29T09:39:49.628297+00:00, sha b14a8d05a0ec)
- Data as of 2026-08-30T08:39:29.467469+00:00.
