MC5lan/dsh-multimodal

Install a pair of eyes and a brush for DeepSeek: directly paste screenshots/pictures in the session, the GLM visual model first accurately transcribes the picture content (error messages, codes, interfaces are retained verbatim), and then DeepSeek continues to deal with your problem - completed in the same round, the whole process is smooth; when you need to match pictures, De

dsh-multimodal gives DeepSeek Harness vision and image-generation capabilities. Paste a screenshot into the conversation and the configured vision provider transcribes it verbatim (error messages, code, UI text preserved), then the reasoning model continues from the transcription — all in the same turn. When the model needs an image, it calls generate_image and the configured image backend renders pictures directly in the conversation. Blank-slate design: no built-in models — you configure your own vision and image APIs.

Web UI Enhancements ★ 5 updated 2026-08-16 ✅ runtime-tested
View on GitHub ↗

Install

dsh plugin --profile web add https://github.com/MC5lan/dsh-multimodal

No npm package — install directly from GitHub URL: dsh plugin --profile web add https://github.com/MC5lan/dsh-multimodal. Or clone first: git clone https://github.com/MC5lan/dsh-multimodal.git && dsh plugin --profile web add /path/to/dsh-multimodal. After install, configure vision and image backends — the plugin ships with NO built-in models or providers. Supported vision backends: DeepSeek, Zhipu, Aliyun, SiliconFlow, ModelScope, Xfyun, Qianfan, local Ollama, or any OpenAI-compatible API. Requires DeepSeek Harness 0.1.0-rc.6+, Node.js 18+.

Compatibility

DeepSeek Harness 0.1.0-rc.6+ (Web and headless). Node.js 18+. Blank-slate: no built-in models or providers. Vision endpoints and image backends declared by the user — supports any OpenAI-compatible API, Zhipu, Aliyun, SiliconFlow, ModelScope, Xfyun, Qianfan, Ollama, or custom backend.

Details

Recent updates

Current: image transcription via configurable vision provider; generate_image tool for configured image backend; backend failover (tries next on failure; skips failover for AUTH/aborted); no built-in providers — all configured by user; custom backend slot for non-standard APIs; CHANGELOG.md tracks version history.

FAQ

What vision providers are supported?
Any OpenAI-compatible vision API. The README mentions DeepSeek, Zhipu, Aliyun, SiliconFlow, ModelScope, Xfyun, Qianfan, and local Ollama as tested providers. A custom backend slot lets you plug in any non-standard API.
Does the plugin come with any built-in vision model?
No — dsh-multimodal is blank-slate by design. No models, providers, or backends are pre-loaded. You configure your own endpoints. This ensures you control costs and privacy.
What happens when an image-generation backend fails?
The plugin tries the next configured backend (failover). AUTH errors and user-aborted requests skip failover to avoid wasting quota. If all backends fail, the model receives an error message.

Alternatives

oil-oil/dsh-vision · 121103qwq/dsh-vision-sidecar · Flyvhidbwo/dsh-vision-proxy

More plugins in Web UI Enhancements

Browse more in Web UI Enhancements

Guides for Web UI Enhancements plugins