FuzzySoul/dsh-free-vision

Free vision bridge for text-only models: image understanding, OCR, UI and debug analysis via free-tier providers (Qwen3-VL-Flash, Doubao, DeepSeek-OCR) with a settings GUI.

dsh-free-vision gives text-only models on DeepSeek Harness image understanding using free-tier vision models with zero MCP configuration: paste a screenshot, error, UI, document, or photo and the image_understand tool (a single general tool registered on ctx.tools) recognizes it via the chosen provider — Qwen3-VL-Flash by default at zero cost. It supports task modes (auto/general/ocr/ui/debug/describe), auto detail escalation with multi-cropping for large images, per-provider API base URL overrides (proxy, gateway, local service, or any OpenAI-compatible endpoint), and direct connections that strip proxy variables for domestic API access. A 1 MB screenshot is roughly 2,600 tokens, and the qwen free quota covers about 190,000 images.

Vision & Multimodal ★ 6 updated 2026-08-19
View on GitHub ↗

Install

dsh plugin --profile web add dsh-free-vision

npm package dsh-free-vision 1.0.8 (registry-verified 2026-08-24). Install: dsh plugin --profile web add dsh-free-vision. Restart dsh web — the image_understand tool becomes available (rename via config.toolName). Zero MCP configuration: the vision engine (luma-mcp) is bundled as a dependency and started in-process. After restart, open Settings → Free Vision for the config form (API keys, provider, tool name) — settings are saved to ~/.dsh/free-vision.json and take effect immediately on the next call without a restart.

Compatibility

DeepSeek Harness web profile. Free-first vision providers: qwen (default, Qwen3-VL-Flash via DASHSCOPE_API_KEY — Alibaba Cloud free tier ~500k tokens), volcengine (Doubao vision, VOLCENGINE_API_KEY, free 200k+ tokens), siliconflow (DeepSeek-OCR free, SILICONFLOW_API_KEY), plus zhipu (GLM-4.6V), hunyuan (HY-Vision), and custom OpenAI-compatible endpoints (CUSTOM_API_KEY + CUSTOM_BASE_URL + CUSTOM_MODEL_NAME). Every provider can override its API base URL (proxy/gateway/local service). Direct connection: the subprocess strips proxy environment variables for direct domestic-API access. Task modes: auto | general | ocr | ui | debug | describe; large images are auto multi-cropped for fidelity. Bilingual tool descriptions (EN/中文).

Details

Recent updates

The current bilingual README (English edition) documents: why free (provider free-tier table with quota estimates), features (zero MCP config, single image_understand tool, free-first multi-provider, per-provider base URL override, direct connection, task modes, bilingual), install (dsh plugin add + restart), settings UI (Settings → Free Vision, ~/.dsh/free-vision.json, immediate effect), and provider environment variables.

FAQ

Do I need an API key?
Yes, but the default providers are free-tier: qwen (Qwen3-VL-Flash, ~500k free tokens via DASHSCOPE_API_KEY), volcengine Doubao (free 200k+), and siliconflow DeepSeek-OCR (free).
Do I need MCP configuration?
No — the vision engine (luma-mcp) is bundled as a package dependency and started in-process; there is no cordis.patch.yml editing or runtime npx.
How does it handle large or complex images?
Task mode auto routes to general/ocr/ui/debug/describe as needed, and large images are automatically multi-cropped to preserve detail.

Alternatives

54xkeee/dsh-vision · Flyvhidbwo/dsh-vision-proxy · stardustlc666/dsh-codex-port

More plugins in Vision & Multimodal

Browse more in Vision & Multimodal

Guides for Vision & Multimodal plugins