YYTbit/dsh-plugin-vision-toolkit
Vision toolkit for DeepSeek Harness -- give text-only agents eyes
Gives a text-only agent the ability to see images by routing image understanding through a vision API and returning text the agent can reason about. The README documents four tools: glance (describe, ask a question about, or OCR an image), ground (locate one element and return its bounding box, e.g. 450,820,620,870), detect (find all instances of an element kind such as buttons) and crop (cut a region out to a file). The plugin registers a system-prompt skill describing the tools; when the agent meets an image it calls the right CLI tool, which reads the file, base64-encodes it, calls the configured vision endpoint and returns the text response — the model never receives raw pixels.
Install
dsh plugin --profile your-profile add dsh-plugin-vision-toolkitnpm package dsh-plugin-vision-toolkit 0.1.1 (registry-verified 2026-09-12; repository field → github.com/YYTbit/dsh-plugin-vision-toolkit). README install line quoted verbatim — substitute your profile name for your-profile. Configure with environment variables: VISION_API_KEY (falls back to DEEPSEEK_API_KEY), VISION_BASE_URL (falls back to DEEPSEEK_BASE_URL) and VISION_MODEL (e.g. deepseek-vl2). Ships four CLI tools — glance, ground, detect, crop — registered as a dsh skill so the agent knows when to reach for them.
Compatibility
DeepSeek Harness (dsh) with the standard plugin manager. Any OpenAI-compatible multimodal endpoint works per the README (DeepSeek VL, GPT-4V, or a self-hosted endpoint). The image is base64-encoded and sent from the host, so the model itself can stay text-only.
Details
- Repo: YYTbit/dsh-plugin-vision-toolkit
- Category: Agent Capabilities
- Stars: 1
- Version: npm dsh-plugin-vision-toolkit 0.1.1 (registry-verified 2026-09-12)
- Last push: 2026-08-13
- First seen: 2026-08-13
Recent updates
The README documents the tool surface and the provider matrix rather than a versioned changelog; there is no release history or migration section. The README notes the agent receives text descriptions only, never raw pixels. Version 0.1.1 on the public registry (verified 2026-09-12).
FAQ
- How do I install dsh-plugin-vision-toolkit?
- Run dsh plugin --profile <your-profile> add dsh-plugin-vision-toolkit, then set VISION_API_KEY (or DEEPSEEK_API_KEY), optionally VISION_BASE_URL and VISION_MODEL.
- Which vision providers are supported?
- DeepSeek VL, GPT-4V, or any OpenAI-compatible multimodal endpoint — the README lists these and the env vars (with DEEPSEEK_* fallbacks) used to point at one.
- Can it return coordinates rather than prose?
- Yes — ground returns a bounding box for a named element and detect finds all instances of an element kind; crop then cuts that region to a file.
Alternatives
tt-a1i/archify · foryourhealth111-pixel/Vibe-Skills · Anionex/agent-vision-toolkit