agent-vision-toolkit

by Anionex · Codex Skill · ★ 1.1k

About agent-vision-toolkit

agent-vision-toolkit What it thinks is what it sees — give any text-only coding agent eyes: image Q&A, OCR, screenshot understanding, visual grounding, and image-to-SVG, as a vision toolkit plus a skill, with optional drop-in integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode. 🌐 中文 | English If your coding agent runs on a text-only model like DeepSeek V4, it can't look at images — screenshots, mockups, diagrams, and error dialogs are all dead ends.

agentagent-skillsclaude-codecodexcomputer-usedeepseekdsh-pluginglmharness-engineeringmultimodal

Quick Facts

Stars1,090
Forks38
LanguagePython
CategoryCodex Skill
LicenseMIT
Quality Score73.7626382783291/100
Open Issues6
Last Updated2026-08-21
Created2026-08-01
Platformsclaude-code, codex, python
Est. Tokens~17k

More Codex Skill Tools

Explore other popular codex skill tools:

View all Codex Skill tools →

Popular Python Agent Tools

Frequently Asked Questions

What is agent-vision-toolkit?

agent-vision-toolkit is 为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI. It is categorized as a Codex Skill with 1.1k GitHub stars.

What programming language is agent-vision-toolkit written in?

agent-vision-toolkit is primarily written in Python. It covers topics such as agent, agent-skills, claude-code.

How do I install or use agent-vision-toolkit?

You can find installation instructions and usage details in the agent-vision-toolkit GitHub repository at github.com/Anionex/agent-vision-toolkit. The project has 1.1k stars and 38 forks, indicating an active community.

What license does agent-vision-toolkit use?

agent-vision-toolkit is released under the MIT license, making it free to use and modify according to the license terms.

View on GitHub → Browse Codex Skill tools