kaixinbaba/dsh-vision-recognizer

Vision provider route that transcribes attached images to text through a configurable model (15+ OpenAI-compatible and Anthropic vendors) while DeepSeek keeps answering.

dsh-vision-recognizer lets text-only DeepSeek Harness conversations attach images anyway: it registers an adaptive provider route that always admits image attachments, then resolves the selected model — models declaring native image input receive original image blocks directly, while text-only or unknown-capability models receive text transcribed by the vision model you configure, keeping DeepSeek as the default wrapped conversation brain. Hardened against hangs: local/anonymous endpoints get a 20s timeout cap, HTTP 429 fails fast, failed endpoints cool down 60s, and without a key or local Ollama it fails fast with guidance. fallbackModels entries are tried in order after a primary failure; a content-hash cache transcribes the same image at most once per process; autoLocalOllama (default on) probes localhost:11434 so local images never leave the machine.

Vision & Multimodal ★ 0 updated 2026-08-21
View on GitHub ↗

Install

dsh plugin --profile web add dsh-vision-recognizer

npm dsh-vision-recognizer 0.2.0 verified 2026-09-03 (repository field → github.com/kaixinbaba/dsh-vision-recognizer; README bilingual EN/简体中文). Install: dsh plugin --profile web add dsh-vision-recognizer — no build scripts, no native deps, no sharp approval. Configure provider/key in Settings → Plugins → Vision; changes apply immediately, no restart.

Compatibility

DSH web profile; wraps the configured conversation provider as an adaptive route ('DeepSeek + 智能识图'); 15+ vision providers incl. OpenAI, Anthropic, Gemini, OpenRouter, Azure, Ollama (local) and CN providers; OpenAI-compatible + native Anthropic wire protocols.

Details

Recent updates

0.2.0 current on npm (verified 2026-09-03).

FAQ

Does it change the main model?
No — DeepSeek stays the conversation brain; the route only admits images and decides whether to pass them natively or transcribe them via the configured vision provider.
Which providers are supported?
15+ — OpenAI, Claude, Gemini, OpenRouter, Azure, Ollama (local), plus Alibaba, Qwen, GLM, Baidu, iFlytek, Kimi, Hunyuan, Doubao, SiliconFlow and any OpenAI-compatible custom endpoint.
What if the vision endpoint hangs?
Local/anonymous endpoints get a hard 20s timeout, 429s fail fast, failed endpoints cool down for 60s, and without any key or local Ollama it fails fast with actionable guidance.

Alternatives

good-boy4069/dsh-vision-guard · mochgolf/dsh-deepseek-vision-router

More plugins in Vision & Multimodal

Browse more in Vision & Multimodal

Guides for Vision & Multimodal plugins