Vision & Multimodal

32 plugins

32 DeepSeek Harness plugins are listed in the Vision & Multimodal category, 19 of them reviewed in depth. 19 of those reviewed pages carry a runnable install command today. The most-starred reviewed entries here are corrinehu/dsh-chat-imagine (★ 10), 54xkeee/dsh-vision (★ 7), FuzzySoul/dsh-free-vision (★ 6). The newest upstream push among these repositories was 2026-09-19. Reviewed entries are listed first, then by GitHub stars.

What Vision & Multimodal plugins do

Vision & Multimodal plugins give text-only models eyes: image understanding and OCR, screenshot capture, and bridges to vision models or local CLIs.

Installing a plugin from this category

Every DSH plugin installs the same way: run dsh plugin --profile web add <package> in the profile you use, then restart the harness. Where this directory has a verified command for a plugin, that exact command is printed on the plugin's page. Step-by-step install guide · What a plugin can run on your machine

Related categories

Guides for Vision & Multimodal

Best DeepSeek Harness plugins · DeepSeek Harness review · DeepSeek Harness vs OpenCode

All 32 Vision & Multimodal plugins

corrinehu/dsh-chat-imagine★ 10
Automatically generates and displays images in the DSH chat via API channels or local CLIs (supports mmx / codex / agy).
Vision & Multimodal
54xkeee/dsh-vision★ 7
Vision for text-only DeepSeek via Doubao Web by default (zero-cost, no API key — drives your logged-in Chrome through a
Vision & Multimodal
FuzzySoul/dsh-free-vision★ 6
Free vision bridge for text-only models: image understanding, OCR, UI and debug analysis via free-tier providers (Qwen3-
Vision & Multimodal
TaurusWood/dsh-plugin-appshot★ 6
Codex Appshots for DSH: capture the frontmost active window via global shortcut and seamlessly mount it into the compose
Vision & Multimodal
Einskyle/dsh-llm-vision-bridge★ 4
Native LLM-provider vision bridge: images pasted in the chat are described by a vision model (Qwen3-VL via pi-ai/llama.c
Vision & Multimodal
54xkeee/dsh-youreyes★ 2
Vision toolkit for text-only DeepSeek: model-invokable vision tool, wrapper adapters for deepseek/opencode-go (v4 flash/
Vision & Multimodal
TZHR-invest/dsh-plugins#dsh-vision-tool★ 2
Agent-callable vision tool that describes local images via any OpenAI-compatible vision endpoint you configure, with an
Vision & Multimodal
Cheng-cheng9669/dsh-deepseek-vision★ 1
Reuses DeepSeek web's built-in vision mode for text-only models: the deepseek_vision tool drives the local deepseek-visi
Vision & Multimodal
MicroHEROX/dsh-koboldcpp-hands★ 1
Hands repetitive text and vision labor (OCR, image analysis, comparison) to a local KoboldCpp (llama.cpp) server through
Vision & Multimodal
MicroHEROX/dsh-unsloth-hands★ 1
Hands repetitive text and vision labor (OCR, image analysis, comparison) to a locally running Unsloth Desktop (Unsloth S
Vision & Multimodal
gloryxpnv/dsh-tool-vision★ 0
Local-first structured vision for text-only agents: images go to a local OpenAI-compatible VLM and come back as JSON evi
Vision & Multimodal
good-boy4069/dsh-vision-guard★ 0
Transparent image guard for text-only routes: paste images without the 400 session deadlock, plus a vision_analyze tool
Vision & Multimodal
kaixinbaba/dsh-vision-recognizer★ 0
Vision provider route that transcribes attached images to text through a configurable model (15+ OpenAI-compatible and A
Vision & Multimodal
maxwell-feng/dsh-tesseract-ocr★ 0
Local OCR for attached images via Tesseract: only the recognized text is sent to the model, never the image bytes; visio
Vision & Multimodal
maxwell-feng/dsh-windows-ocr★ 0
Local OCR for attached images via the built-in Windows engine (Windows.Media.Ocr): only the recognized text is sent to t
Vision & Multimodal
niuniuaba/dsh-subagent-vision★ 0
Lets a text-only DeepSeek agent read images in the same session by delegating to a vision-capable subagent, with send-ti
Vision & Multimodal
s3yf1337/dsh-easyvision★ 0
Give text-only models vision: a describe_image tool that delegates images to a vision model from your dsh model list ove
Vision & Multimodal
siegfly/dsh-deepseek-vision★ 0
A vision-language gateway provider route: pasted images are described by a configurable VL model (Qwen-VL by default) be
Vision & Multimodal
wulusai2333/mimo-vision★ 0
describe_image tool: a vision bridge that sends images to mimo-v2.5 through the opencode Zen API (credential OPENCODE_GO
Vision & Multimodal
dickpy/dsh-imagegen★ 44
AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpo
Vision & Multimodal
ConsoleSun/Gemini-Eyes★ 10
MCP bridge to gemini.google.com: vision analysis of images and videos, Imagen image and Veo video generation, and conver
Vision & Multimodal
GOU-GEE/deepseek-vision#plugins/dsh-plugin-deepseek-vision★ 5
Vision MCP and DSH bundle for text-only DeepSeek: analyze_image, analyze_clipboard, compare_images and vision_status too
Vision & Multimodal
NagasakiSoyo-ui/dsh-llm-deepseek-vision★ 2
Vision-augmented DeepSeek adapter: a vision-capable model describes image input, then a text-only DeepSeek model reasons
Vision & Multimodal
Isanti2016/dsh-quicksight★ 1
Two-tier image reading for text-only models: fast local OCR (RapidOCR, offline) first, vision-model fallback (modlens).
Vision & Multimodal
Elohia/dsh-plugin-image-input★ 0
Image-to-text input for the Web UI: paste or drag an image and it is transcribed into structured text and sent, giving t
Vision & Multimodal
haiziyao/dsh-vision-mix★ 0
Combine text, vision, and image-generation APIs into one Mix model with automatic routing: text-only requests go to the
Vision & Multimodal
jing-hy/picturereader★ 0
Image "reading" for text-only models: downscale + reduce color depth + structure/color fingerprints into text grids fed
Vision & Multimodal
ld-1101/dsh-vision-plugin★ 0
Give your text-only model eyes - chat image attachments are auto-described via a vision model (default prompt), with ite
Vision & Multimodal
shinjiyu/dsh-plugin-multimodal★ 0
Advertise image paste on text-only DeepSeek routes, describe attachments with a vision sidecar, and leave native vision
Vision & Multimodal
xiaozhe7772222/dsh-draw-router★ 0
Universal image generation for DeepSeek Harness: auto-discovers image models from any OpenAI-compatible endpoint (SenseN
Vision & Multimodal
ximengxiaolan/dsh-vision-bridge★ 0
Composer-attached images are transcribed to text by an OpenAI-compatible vision model before reaching text-only DeepSeek
Vision & Multimodal
xsoc1/dsh-image-vision★ 0
Chat image-attachment bridge with a view_image tool for any OpenAI-compatible VLM (local Ollama or cloud): pasted/droppe
Vision & Multimodal