RRRosmontis/dsh-qwen-mm

Qwen-MM-Plugins integration bundle for DeepSeek Harness (dsh) — multimodal MCP tools (vision, OCR, ASR, search

Makes DeepSeek Harness multimodal-native by integrating Qwen-MM-Plugins: core MCP tools (read_image, visualize, media_info, read_video…) read images/videos/documents/code/data with no key; api MCP tools (vision_chat, ocr, omni_*) add cloud vision/OCR/ASR/object localization via DashScope Qwen VL/Omni; search MCP tools (web_search, web_extractor, image_search) add web + reverse-image search; video-memory / video-edit / blender / freecad / edu-agent extend to long-video QA, editing, 3D and math. A three-part image-attachment bridge lets you drag an image onto a pure-text model: the dsh-attachment consumer registry allows image upload on text-only routes, the attachments bridge plugin exports each image block to <dshHome>/qwen-mm/attachments/<sha256>.<ext> and rewrites the message to a path reference, and usage guidance tells the model to read it via mcpqwen-mm-plugins-apivision_chat or mcpqwen-mm-plugins-coreread_image.

Coding & Development ★ 2 updated 2026-08-13 ✅ runtime-tested
View on GitHub ↗

Install

pnpm dsh --profile web plugin add github:RRRosmontis/dsh-qwen-mm

Drag-image support needs a small DSH core change (image-consumer registry) not yet in official releases, so the README recommends installing from the maintained fork: git clone https://github.com/RRRosmontis/deepseek-harness.git && cd deepseek-harness && pnpm install && pnpm run build, then pnpm dsh --profile web plugin add github:RRRosmontis/dsh-qwen-mm (or ./packages/bundle/qwen-mm) and pnpm dsh --profile web. Prereqs: Node.js, pnpm, and uv (for MCP python bridges). Bilingual README with full English section.

Compatibility

DSH (fork with image-consumer registry for full drag-image support); pure-text DeepSeek models; needs DashScope key for cloud vision/OCR/ASR, Serper/Tavily/Exa for search tools. MCP-based.

Details

Recent updates

Qwen-MM-Plugins integration; keyless core vision tools; DashScope cloud vision/OCR/ASR; web + reverse-image search; drag-image bridge for pure-text models.

FAQ

Which tools need a key?
Core read/visualize/media tools are keyless; cloud vision/OCR/ASR needs a DashScope key; web search needs Serper/Tavily/Exa.
Can pure-text DeepSeek models see images?
Yes — the attachment bridge lets you drag an image in; the message is rewritten to a path reference and the model reads it via the Qwen VL MCP tools.
Do I need the fork?
For full drag-image support yes (it contains the image-consumer registry patch); core MCP tools work without it.

Alternatives

Scorp1o117/dsh-tool-vision · yuz12-dsh-vision-helper · libinyam/dsh-vision-provider

More plugins in Coding & Development

Browse more in Coding & Development

Guides for Coding & Development plugins