s3yf1337/dsh-easyvision

Give text-only models vision: a describe_image tool that delegates images to a vision model from your dsh model list over the harness's own LLM runtime.

dsh-easyvision gives a text-only DeepSeek Harness conversation model 'eyes' with one command and zero extra APIs: attached images just work (drop an image in the composer and send — the message is admitted and the image is described through the vision model instead of the 'current model does not support images' refusal), and a describe_image tool lets the model inspect image files on its own by handing them to the vision-capable model and returning plain text. The vision call goes through ctx.llm — the exact same runtime the agent loop uses — so there are no external API keys, no extra plumbing, and your keys/retry policy/middleware apply. Any vision model from your dsh model list works (Settings → EasyVision); only models that declare image input are accepted, and model changes apply immediately without a restart. Multiple validated images per call use the same attachment pipeline as read_image.

Vision & Multimodal ★ 0 updated 2026-08-16
View on GitHub ↗

Install

curl -fsSL https://raw.githubusercontent.com/s3yf1337/dsh-easyvision/main/install.sh | bash

Install per README (EN primary): one command — curl -fsSL https://raw.githubusercontent.com/s3yf1337/dsh-easyvision/main/install.sh | bash — idempotent and safe to re-run; then open Settings → EasyVision and pick a vision-capable model from your list (default qwen3.7-plus on opencode-go). npm 404 verified 2026-09-04 — the install.sh route is the documented channel. Zero external API keys: the vision call goes through ctx.llm, the same runtime the agent loop uses.

Compatibility

DeepSeek Harness web chat; works when the conversation model is text-only; vision model picked from your own dsh model list; validated PNG/JPEG/WebP/GIF; live config — model changes apply immediately, no restart.

Details

Recent updates

No npm release (404 verified 2026-09-04); README shows v0.4.0.

FAQ

Do I need a separate vision API key?
No — the vision call goes through ctx.llm, the same runtime the agent loop uses: your existing keys, retry policy and middleware apply. Zero external APIs.
Which vision model is used?
Any vision-capable model from your own dsh model list, chosen in Settings → EasyVision (default qwen3.7-plus on opencode-go); text-only picks are refused with a clear message.
Does it change the conversation model?
No — the main model stays text-only; images are described through the vision model automatically on send, or on demand via the describe_image tool for image files.

Alternatives

1710782766/dsh-llm-vision · niuniuaba/dsh-subagent-vision

More plugins in Vision & Multimodal

Browse more in Vision & Multimodal

Guides for Vision & Multimodal plugins