kaixinbaba/dsh-vision-recognizer
Vision provider route that transcribes attached images to text through a configurable model (15+ OpenAI-compatible and Anthropic vendors) while DeepSeek keeps answering.
dsh-vision-recognizer lets text-only DeepSeek Harness conversations attach images anyway: it registers an adaptive provider route that always admits image attachments, then resolves the selected model — models declaring native image input receive original image blocks directly, while text-only or unknown-capability models receive text transcribed by the vision model you configure, keeping DeepSeek as the default wrapped conversation brain. Hardened against hangs: local/anonymous endpoints get a 20s timeout cap, HTTP 429 fails fast, failed endpoints cool down 60s, and without a key or local Ollama it fails fast with guidance. fallbackModels entries are tried in order after a primary failure; a content-hash cache transcribes the same image at most once per process; autoLocalOllama (default on) probes localhost:11434 so local images never leave the machine.
Install
dsh plugin --profile web add dsh-vision-recognizernpm dsh-vision-recognizer 0.2.0 verified 2026-09-03 (repository field → github.com/kaixinbaba/dsh-vision-recognizer; README bilingual EN/简体中文). Install: dsh plugin --profile web add dsh-vision-recognizer — no build scripts, no native deps, no sharp approval. Configure provider/key in Settings → Plugins → Vision; changes apply immediately, no restart.
Compatibility
DSH web profile; wraps the configured conversation provider as an adaptive route ('DeepSeek + 智能识图'); 15+ vision providers incl. OpenAI, Anthropic, Gemini, OpenRouter, Azure, Ollama (local) and CN providers; OpenAI-compatible + native Anthropic wire protocols.
Details
- Repo: kaixinbaba/dsh-vision-recognizer
- Category: Vision & Multimodal
- Stars: 0
- Version: npm dsh-vision-recognizer 0.2.0
- Last push: 2026-08-21
- First seen: 2026-08-15
Recent updates
0.2.0 current on npm (verified 2026-09-03).
FAQ
- Does it change the main model?
- No — DeepSeek stays the conversation brain; the route only admits images and decides whether to pass them natively or transcribe them via the configured vision provider.
- Which providers are supported?
- 15+ — OpenAI, Claude, Gemini, OpenRouter, Azure, Ollama (local), plus Alibaba, Qwen, GLM, Baidu, iFlytek, Kimi, Hunyuan, Doubao, SiliconFlow and any OpenAI-compatible custom endpoint.
- What if the vision endpoint hangs?
- Local/anonymous endpoints get a hard 20s timeout, 429s fail fast, failed endpoints cool down for 60s, and without any key or local Ollama it fails fast with actionable guidance.
Alternatives
good-boy4069/dsh-vision-guard · mochgolf/dsh-deepseek-vision-router