MC5lan/dsh-multimodal
Install a pair of eyes and a brush for DeepSeek: directly paste screenshots/pictures in the session, the GLM visual model first accurately transcribes the picture content (error messages, codes, interfaces are retained verbatim), and then DeepSeek continues to deal with your problem - completed in the same round, the whole process is smooth; when you need to match pictures, De
dsh-multimodal gives DeepSeek Harness vision and image-generation capabilities. Paste a screenshot into the conversation and the configured vision provider transcribes it verbatim (error messages, code, UI text preserved), then the reasoning model continues from the transcription — all in the same turn. When the model needs an image, it calls generate_image and the configured image backend renders pictures directly in the conversation. Blank-slate design: no built-in models — you configure your own vision and image APIs.
Install
dsh plugin --profile web add https://github.com/MC5lan/dsh-multimodalNo npm package — install directly from GitHub URL: dsh plugin --profile web add https://github.com/MC5lan/dsh-multimodal. Or clone first: git clone https://github.com/MC5lan/dsh-multimodal.git && dsh plugin --profile web add /path/to/dsh-multimodal. After install, configure vision and image backends — the plugin ships with NO built-in models or providers. Supported vision backends: DeepSeek, Zhipu, Aliyun, SiliconFlow, ModelScope, Xfyun, Qianfan, local Ollama, or any OpenAI-compatible API. Requires DeepSeek Harness 0.1.0-rc.6+, Node.js 18+.
Compatibility
DeepSeek Harness 0.1.0-rc.6+ (Web and headless). Node.js 18+. Blank-slate: no built-in models or providers. Vision endpoints and image backends declared by the user — supports any OpenAI-compatible API, Zhipu, Aliyun, SiliconFlow, ModelScope, Xfyun, Qianfan, Ollama, or custom backend.
Details
- Repo: MC5lan/dsh-multimodal
- Category: Web UI Enhancements
- Stars: 5
- Version: GitHub source install (README-documented 2026-08-24; no npm package). Compatible with DeepSeek Harness 0.1.0-rc.6+.
- Last push: 2026-08-16
- First seen: 2026-08-13
Recent updates
Current: image transcription via configurable vision provider; generate_image tool for configured image backend; backend failover (tries next on failure; skips failover for AUTH/aborted); no built-in providers — all configured by user; custom backend slot for non-standard APIs; CHANGELOG.md tracks version history.
FAQ
- What vision providers are supported?
- Any OpenAI-compatible vision API. The README mentions DeepSeek, Zhipu, Aliyun, SiliconFlow, ModelScope, Xfyun, Qianfan, and local Ollama as tested providers. A custom backend slot lets you plug in any non-standard API.
- Does the plugin come with any built-in vision model?
- No — dsh-multimodal is blank-slate by design. No models, providers, or backends are pre-loaded. You configure your own endpoints. This ensures you control costs and privacy.
- What happens when an image-generation backend fails?
- The plugin tries the next configured backend (failover). AUTH errors and user-aborted requests skip failover to avoid wasting quota. If all backends fail, the model receives an error message.
Alternatives
oil-oil/dsh-vision · 121103qwq/dsh-vision-sidecar · Flyvhidbwo/dsh-vision-proxy