YYTbit/dsh-plugin-vision-toolkit

Vision toolkit for DeepSeek Harness -- give text-only agents eyes

Gives a text-only agent the ability to see images by routing image understanding through a vision API and returning text the agent can reason about. The README documents four tools: glance (describe, ask a question about, or OCR an image), ground (locate one element and return its bounding box, e.g. 450,820,620,870), detect (find all instances of an element kind such as buttons) and crop (cut a region out to a file). The plugin registers a system-prompt skill describing the tools; when the agent meets an image it calls the right CLI tool, which reads the file, base64-encodes it, calls the configured vision endpoint and returns the text response — the model never receives raw pixels.

Agent Capabilities ★ 1 updated 2026-08-13 ✅ runtime-tested
View on GitHub ↗

Install

dsh plugin --profile your-profile add dsh-plugin-vision-toolkit

npm package dsh-plugin-vision-toolkit 0.1.1 (registry-verified 2026-09-12; repository field → github.com/YYTbit/dsh-plugin-vision-toolkit). README install line quoted verbatim — substitute your profile name for your-profile. Configure with environment variables: VISION_API_KEY (falls back to DEEPSEEK_API_KEY), VISION_BASE_URL (falls back to DEEPSEEK_BASE_URL) and VISION_MODEL (e.g. deepseek-vl2). Ships four CLI tools — glance, ground, detect, crop — registered as a dsh skill so the agent knows when to reach for them.

Compatibility

DeepSeek Harness (dsh) with the standard plugin manager. Any OpenAI-compatible multimodal endpoint works per the README (DeepSeek VL, GPT-4V, or a self-hosted endpoint). The image is base64-encoded and sent from the host, so the model itself can stay text-only.

Details

Recent updates

The README documents the tool surface and the provider matrix rather than a versioned changelog; there is no release history or migration section. The README notes the agent receives text descriptions only, never raw pixels. Version 0.1.1 on the public registry (verified 2026-09-12).

FAQ

How do I install dsh-plugin-vision-toolkit?
Run dsh plugin --profile <your-profile> add dsh-plugin-vision-toolkit, then set VISION_API_KEY (or DEEPSEEK_API_KEY), optionally VISION_BASE_URL and VISION_MODEL.
Which vision providers are supported?
DeepSeek VL, GPT-4V, or any OpenAI-compatible multimodal endpoint — the README lists these and the env vars (with DEEPSEEK_* fallbacks) used to point at one.
Can it return coordinates rather than prose?
Yes — ground returns a bounding box for a named element and detect finds all instances of an element kind; crop then cuts that region to a file.

Alternatives

tt-a1i/archify · foryourhealth111-pixel/Vibe-Skills · Anionex/agent-vision-toolkit

More plugins in Agent Capabilities

Browse more in Agent Capabilities

Guides for Agent Capabilities plugins