# showlab/ShowUI

[CVPR 2025] Open-source, End-to-end, Vision-Language-Action model for GUI Agent & Computer Use.

Repository: https://github.com/showlab/ShowUI
Canonical: https://ross.abutalabs.com/products/showui
Homepage: https://arxiv.org/abs/2411.17465
Language: Python
License: Apache-2.0
License Family: permissive
Topics: computer-use, vision-language-action, vision-language-model, agent, gui-agent
Last push: 2026-04-24T07:16:50+00:00

## Health v2 (maintenance only)
Score: 57/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 79, release rhythm 35, longevity 47
- inputs: {"age_days": 671, "days_push": 131, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1893, forks 141 (observed 2026-08-28T04:05:50.609081+00:00)

## What it is
ShowUI is an open-source, lightweight 2B vision-language-action model for GUI agents and computer use, accepted at CVPR 2025. The repository provides training, fine-tuning, and inference code (including vLLM and Gradio API support) plus curated GUI instruction-following datasets.

## Use cases
- build a gui agent that clicks buttons from screenshots
- automate computer use tasks with a vision language model
- fine-tune a vlm on ui screenshots for grounding
- run screenshot grounding to locate ui elements
- train an agent to navigate web and mobile apps
- self-host an open-source computer use model

## When to choose
- you need an open-source, lightweight vision-language-action model for GUI automation
- you want to fine-tune or evaluate GUI agents on benchmarks like Mind2Web, AITW, or MiniWob
- you need zero-shot screenshot grounding of UI elements

## When to avoid
- you need a production-ready turnkey desktop automation product rather than a research model
- you lack GPU resources for model inference or training
- you only need text-based agents using HTML or accessibility trees

## Facets
- artifact type: library
- maturity: active
- function: machine-learning, deep-learning, computer-vision, agent-framework, llm-inference
- domain: artificial-intelligence, computer-vision, large-language-models
- platform: python, cross-platform
- tags: vision-language-action, gui-agent, computer-use, screenshot-grounding, ui-automation, qwen2-vl, fine-tuning, vllm, ai-agents, automation, gpu

## Member repositories
- showlab/ShowUI (main) score 57

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:05:50.609081+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T03:12:56.023418+00:00, confidence not recorded.
  - readme: https://github.com/showlab/ShowUI (fetched 2026-08-28T04:05:50.609081+00:00, sha 0e3096d0dc65)
  - homepage: https://arxiv.org/abs/2411.17465 (fetched 2026-08-29T10:52:00.404365+00:00, sha 380e888fdc42)
  - site_page: https://info.arxiv.org/about/donate.html (fetched 2026-08-29T10:52:00.413443+00:00, sha cca9c3a11c56)
  - site_page: https://info.arxiv.org/about/ourmembers.html (fetched 2026-08-29T10:52:00.416621+00:00, sha 47cbc55ff1de)
  - site_page: https://info.arxiv.org/about (fetched 2026-08-29T10:52:00.418555+00:00, sha a1f16f915a9a)
  - site_page: https://info.arxiv.org/labs/index.html (fetched 2026-08-29T10:52:00.415063+00:00, sha b14a8d05a0ec)
- Data as of 2026-08-30T08:39:29.467469+00:00.
