ruby1304/dsh-vision-subagent

Vision for text-only agents: a vision_agent tool that delegates image reading to a one-shot subagent on a configurable vision route (MiniMax/Kimi), plus a Codex-style paste bridge — images are analyzed on an isolated context and only text reaches the main session.

dsh-vision-subagent gives text-only DeepSeek Harness agents eyes by delegating image reading to a one-shot subagent running on a separately configured vision route (MiniMax, Kimi, or any OpenAI-compatible provider). Context isolation is the point: large screenshots and multi-image comparisons never occupy the main model's window, the child can call read_image on more workspace files before answering (multi-turn visual reasoning), and vision calls bill on the MiniMax/Kimi route while the main model only reasons. Pasting or dropping an image into the Web composer works Codex-style: the client uploads the image to the host endpoint, which stores it as a durable attachment and runs ONE vision-route analysis on an isolated context guided by your draft message — image bytes and intermediate context never enter the main session, only the final text answer returns.

Memory ★ 0 updated 2026-09-10
View on GitHub ↗

Install

dsh plugin --profile web add dsh-vision-subagent

npm dsh-vision-subagent 0.3.0 verified 2026-09-04 (repository field → github.com/ruby1304/dsh-vision-subagent; README EN primary). Install: dsh plugin --profile web add dsh-vision-subagent, then add the vision-route row to ~/.dsh/profiles/web/cordis.patch.yml (e.g. provider: kimi-coding / model: k3, or minimax-cn / MiniMax-VL-01), restart dsh web, and ask the agent to look at an image. Image bytes and the vision model's intermediate context never enter the main session — only the final text answer comes back.

Compatibility

DeepSeek Harness Web (0.1.x per README); delegates to a one-shot subagent on a separately configured vision route (MiniMax / Kimi / any OpenAI-compatible provider); paste/drop image intake is native.

Details

Recent updates

0.3.0 is the current npm latest (verified 2026-09-04).

FAQ

Which vision routes can it use?
A separately configured vision route — MiniMax, Kimi, or any OpenAI-compatible provider — set in cordis.patch.yml; the main session's model stays text-only.
Do image bytes enter the main session?
No — images are analyzed in an isolated subagent context and only the final text answer comes back; the main model never receives the raw image or the vision model's intermediate reasoning.
Can it do multi-turn visual reasoning?
Yes — the child subagent can call read_image on additional workspace files before answering, which is what makes comparisons and follow-up checks possible.

Alternatives

niuniuaba/dsh-subagent-vision · 1710782766/dsh-llm-vision

More plugins in Memory

Browse more in Memory

Guides for Memory plugins