good-boy4069/dsh-vision-guard

Transparent image guard for text-only routes: paste images without the 400 session deadlock, plus a vision_analyze tool for OCR/PDF/docx/pptx/video.

dsh-vision-guard lets text-only DeepSeek Harness models 'see' images without ever risking the 400 deadlock. It applies two gates: at agent/pre-step an image is converted to text by a vision model BEFORE it is written into the session log (so the log only ever contains text), and an llm/stream backstop rewrites any image block still present in replayed history (e.g. a session poisoned before install) to OCR text at request time — healing already-deadlocked conversations. It also registers a vision_analyze tool the model can invoke on workspace files: image OCR, PDF text layer plus embedded images, docx/pptx text plus embedded images, video-frame OCR (≤12 frames) and plain text files, with loud rejection for xlsx/doc.

Vision & Multimodal ★ 0 updated 2026-08-15
View on GitHub ↗

Install

dsh plugin --profile web add dsh-vision-guard

npm dsh-vision-guard 0.1.3 verified 2026-09-03 (repository field → github.com/good-boy4069/dsh-vision-guard; README EN primary with 中文 edition). Install: dsh plugin --profile web add dsh-vision-guard, or from git: dsh plugin --profile web add github:good-boy4069/dsh-vision-guard.

Compatibility

DSH with text-only model routes (deepseek-v4-pro etc.); vision model configured separately; routes that natively accept images pass through when whitelisted.

Details

Recent updates

0.1.3 current on npm (verified 2026-09-03).

FAQ

How does it prevent the image 400 deadlock?
Images are converted to text before they are appended to the session log, so no image block ever reaches a text-only upstream; a request-time backstop also rewrites images in already-poisoned histories.
Does the main model become multimodal?
No — the vision model acts as the eyes and the text-only main model keeps reasoning over the converted text; whitelisted native-vision routes pass through untouched.
What files can vision_analyze read?
Workspace images, PDFs (text layer plus embedded images), docx/pptx (text plus embedded images), video frames (≤12) and plain text; xlsx/doc are loudly rejected.

Alternatives

1710782766/dsh-llm-vision · mochgolf/dsh-deepseek-vision-router

More plugins in Vision & Multimodal

Browse more in Vision & Multimodal

Guides for Vision & Multimodal plugins