by Anionex · Codex Skill · ★ 1.1k
agent-vision-toolkit What it thinks is what it sees — give any text-only coding agent eyes: image Q&A, OCR, screenshot understanding, visual grounding, and image-to-SVG, as a vision toolkit plus a skill, with optional drop-in integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode. 🌐 中文 | English If your coding agent runs on a text-only model like DeepSeek V4, it can't look at images — screenshots, mockups, diagrams, and error dialogs are all dead ends.
| Stars | 1,090 |
| Forks | 38 |
| Language | Python |
| Category | Codex Skill |
| License | MIT |
| Quality Score | 73.7626382783291/100 |
| Open Issues | 6 |
| Last Updated | 2026-08-21 |
| Created | 2026-08-01 |
| Platforms | claude-code, codex, python |
| Est. Tokens | ~17k |
Explore other popular codex skill tools:
agent-vision-toolkit is 为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI. It is categorized as a Codex Skill with 1.1k GitHub stars.
agent-vision-toolkit is primarily written in Python. It covers topics such as agent, agent-skills, claude-code.
You can find installation instructions and usage details in the agent-vision-toolkit GitHub repository at github.com/Anionex/agent-vision-toolkit. The project has 1.1k stars and 38 forks, indicating an active community.
agent-vision-toolkit is released under the MIT license, making it free to use and modify according to the license terms.