niuniuaba/dsh-subagent-vision

Lets a text-only DeepSeek agent read images in the same session by delegating to a vision-capable subagent, with send-time image-to-path conversion.

dsh-subagent-vision lets a text-only DeepSeek Harness main agent read images in the same session — no model switching, no second session, no copy-paste. When a task needs vision, the main agent delegates to a fresh subagent routed to the vision-capable model you pick in settings, and the child's text result is merged back; the image itself never reaches the main model (the parent passes a file path or URL). Pasting or dropping an image just works: intake stays native (thumbnail rail, remove/undo), and on send the browser half uploads each draft image to a private temp file and appends the paths to the prompt, so the request never trips image admission and the text-only agent can delegate the paths to the vision subagent. A second delegation tool (pinned to the vision route) is exposed alongside the harness's native subagent/subagent_fork tools.

Vision & Multimodal ★ 0 updated 2026-09-16
View on GitHub ↗

Install

dsh plugin --profile web add dsh-subagent-vision

npm dsh-subagent-vision 0.1.5 verified 2026-09-04 (repository field → github.com/niuniuaba/dsh-subagent-vision; README EN primary with 中文). Install: dsh plugin --profile web add dsh-subagent-vision. The bundle's cordis.patch.yml inserts two rows: a second @deepseek-ai/dsh-tool-subagent instance (tool-subagent-vision) and the client UI half. Pick the vision model under Settings → 视觉处理模型 (factory default qwen3.8-max). Pasting/dropping an image just works on text-only sessions: intake stays native and images reach the vision tool as file paths without tripping image admission.

Compatibility

DeepSeek Harness Web profile with a text-only main agent; delegates to a fresh subagent routed to any vision-capable model you pick in settings; the current session's model stays text-only.

Details

Recent updates

0.1.5 is the current npm latest (verified 2026-09-04).

FAQ

Does the main model ever receive the image?
No — the parent passes a file path or URL to a fresh vision-route subagent and only the child's text reading is merged back; the session's model stays text-only.
Which vision model is used?
Whatever you pick in the plugin's settings (factory default qwen3.8-max) among your registered models; the child is routed to that provider/model.
Does pasting an image still work?
Yes — composer paste/drop stays native; on send each image is uploaded to a private temp file and its path is appended to the prompt so the request never trips the image-admission gate.

Alternatives

ruby1304/dsh-vision-subagent · kaixinbaba/dsh-vision-recognizer

More plugins in Vision & Multimodal

Browse more in Vision & Multimodal

Guides for Vision & Multimodal plugins