Ross ROSS = Recommend OSS · open-source software intelligence for agents

Anionex/agent-vision-toolkit

为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode observed · 2026-08-28

github.com/Anionex/agent-vision-toolkit · homepage · Python · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

79/100

  • Activity 99
  • Release rhythm 98
  • Longevity 2

Flags: young

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: 7
  • age_days: 32
  • days_rel: 19
  • days_push: 8
  • n_releases_24m: 2

Full methodology

Adoption not part of the score

1108 stars · 38 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

A Python toolkit of vision CLI tools plus an agent skill that gives text-only LLM agents (like DeepSeek) image understanding capabilities, including image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation. It optionally integrates via a local proxy and plugins with agents such as Codex, Claude Code, Pi, and OpenCode.

Use cases

  • let a text-only LLM agent see and answer questions about images
  • run OCR on long screenshots from a coding agent
  • restore frontend UI from a screenshot
  • automate GUI interactions with a text-only model
  • paste images into a non-multimodal coding agent
  • add vision tools to DeepSeek or other text-only model agents

When to choose

  • your agent runs on a text-only model and lacks multimodal support
  • you use Codex, Claude Code, Pi, or OpenCode and want vision capabilities
  • you need screenshot OCR or UI restoration inside an agent workflow

When to avoid

  • your model is already natively multimodal
  • you need a standalone OCR service rather than agent tooling
  • you don't use shell-capable coding agents

Facets

library · maturity active

computer-vision ocr agent-framework cli mcp image-processing artificial-intelligence computer-vision developer-tools large-language-models python cli cross-platform vision-tools text-only-llm agent-skills gui-automation screenshot-ocr ui-restoration claude-code codex proxy-integration ai-agents

1 source

Member repositories

RepositoryRoleHealth v2
Anionex/agent-vision-toolkitmain79

For agents

markdown · JSON · MCP: product_card(name="Anionex/agent-vision-toolkit")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem