Scorp1o117/dsh-tool-vision

Vision model for DeepSeek Harness | DeepSeek Harness external vision model plug-in

dsh-tool-vision gives DeepSeek Harness a vision model through any OpenAI-compatible API: configure baseURL, apiKey/apiKeyEnv, and model (gpt-4o-mini by default), and the model gains image understanding for screenshots, errors, UI analysis, OCR, and documents. bridgeTextOnly mode bridges pasted images to text hints on models that cannot see images, exporting bridged images to a temp dir, while multimodalModels lists model ids that receive image blocks directly. It is part of the DeepSeek Harness Enhancement Suite (Vision · Soul/Persona · Long-term Memory · Plugin Marketplace).

Other ★ 6 updated 2026-09-11 ✅ runtime-tested
View on GitHub ↗

Install

dsh plugin --profile web add dsh-tool-vision

npm package dsh-tool-vision 0.6.4 (registry-verified 2026-08-24). Install: dsh plugin --profile web add dsh-tool-vision, then mount it in the profile patch ($DSH_HOME/profiles/<name>/cordis.patch.yml) with an insert row (id: tool-vision, name: 'dsh-tool-vision', config: baseURL, apiKeyEnv, model). Or load from a local path without npm (name: './plugins/dsh-tool-vision/index.js'). Part of the DeepSeek Harness Enhancement Suite. Config defaults: baseURL https://api.openai.com/v1, apiKeyEnv VISION_API_KEY, model gpt-4o-mini, maxTokens 1024, timeoutMs 60000, maxImageBytes 10MB, bridgeTextOnly true, bridgeExportDir os.tmpdir()/dsh-vision-bridge.

Compatibility

DeepSeek Harness web profile. OpenAI-compatible vision API: baseURL (default https://api.openai.com/v1), apiKey (takes precedence over env) or apiKeyEnv (default VISION_API_KEY), model (default gpt-4o-mini), maxTokens (1024), timeoutMs (60000), maxImageBytes (10MB). bridgeTextOnly (default true) bridges pasted images to text hints on models that cannot see images, exporting bridged images to bridgeExportDir (os.tmpdir()/dsh-vision-bridge); multimodalModels lists model ids that receive image blocks directly. Part of the Enhancement Suite.

Details

Recent updates

The current README documents: install (dsh plugin add + profile patch mount, or local-path load), the full config table (baseURL, apiKey, apiKeyEnv, model, maxTokens, timeoutMs, maxImageBytes, description, bridgeTextOnly, bridgeExportDir, multimodalModels), and the Enhancement Suite relationship.

FAQ

Which API does it use?
Any OpenAI-compatible vision API — configure baseURL (default https://api.openai.com/v1), apiKey or apiKeyEnv (default VISION_API_KEY), and model (default gpt-4o-mini).
How does it work with text-only models?
bridgeTextOnly (default true) bridges pasted images to text hints on models that cannot see images, exporting bridged images to a temp dir; multimodalModels lists model ids that receive image blocks directly.
Do I need npm to load it?
No — it can also be loaded from a local path without npm (name: './plugins/dsh-tool-vision/index.js').

Alternatives

54xkeee/dsh-vision · FuzzySoul/dsh-free-vision · xsoc1/dsh-image-vision

More plugins in Other

Browse more in Other

Guides for Other plugins