Anionex/agent-vision-toolkit
为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode observed · 2026-08-28
Health v2 · maintenance only
79/100
- Activity 99
- Release rhythm 98
- Longevity 2
Flags: young
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: 7
- age_days: 32
- days_rel: 19
- days_push: 8
- n_releases_24m: 2
Adoption not part of the score
1108 stars · 38 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
A Python toolkit of vision CLI tools plus an agent skill that gives text-only LLM agents (like DeepSeek) image understanding capabilities, including image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation. It optionally integrates via a local proxy and plugins with agents such as Codex, Claude Code, Pi, and OpenCode.
Use cases
- let a text-only LLM agent see and answer questions about images
- run OCR on long screenshots from a coding agent
- restore frontend UI from a screenshot
- automate GUI interactions with a text-only model
- paste images into a non-multimodal coding agent
- add vision tools to DeepSeek or other text-only model agents
When to choose
- your agent runs on a text-only model and lacks multimodal support
- you use Codex, Claude Code, Pi, or OpenCode and want vision capabilities
- you need screenshot OCR or UI restoration inside an agent workflow
When to avoid
- your model is already natively multimodal
- you need a standalone OCR service rather than agent tooling
- you don't use shell-capable coding agents
Facets
library · maturity active
computer-vision ocr agent-framework cli mcp image-processing artificial-intelligence computer-vision developer-tools large-language-models python cli cross-platform vision-tools text-only-llm agent-skills gui-automation screenshot-ocr ui-restoration claude-code codex proxy-integration ai-agents
1 source
- readme: https://github.com/Anionex/agent-vision-toolkit · fetched 2026-08-28 · ba93bc91d90e
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| Anionex/agent-vision-toolkit | main | 79 |
For agents
markdown · JSON · MCP: product_card(name="Anionex/agent-vision-toolkit")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem