1710782766/dsh-llm-vision

Reliable vision + OCR for text-only models on DeepSeek Harness: describe_image (normal/critical) + extract_text tools, auto-preprocessing, retries, and a persistent answer cache.

dsh-llm-vision gives text-only DeepSeek Harness models reliable image understanding and OCR, configured entirely in the GUI. It registers model-facing tools backed by any OpenAI-compatible vision endpoint: describe_image (natural description plus a 'critical' objective-inspection perspective that reports text misalignment, overlap, occlusion and wrapping anomalies and separates fact from guess — for UI bug reports and screenshot-vs-design comparisons; accepts one image or a batch of up to 8), extract_text (OCR through a dedicated OCR model with structured JSON/CSV output and verbatim extraction that never guesses missing text), and llm_vision_check (diagnostics that verify config, key resolution and endpoint auth — the key itself never appears in the report). The browser half rewrites paste/drag/drop image sends into attach references the text model can resolve, upgrades them to inline thumbnails in the transcript, and a live settings card (Settings → Plugins → llm-vision) controls endpoint, models, prompts, bounds, retries, preprocessing and cache. Inputs may be local absolute paths, http(s) URLs (redirects refused) or attachment references; the image never enters the session log — only the returned text crosses into the conversation.

uncategorized ★ 0 updated 2026-08-17
View on GitHub ↗

Install

dsh plugin --profile web add dsh-llm-vision@0.3.2

npm dsh-llm-vision 0.3.2 verified 2026-09-03 (repository field → github.com/1710782766/dsh-llm-vision; README EN primary with zh edition). The README pins the version on purpose: pnpm 11 holds back packages published in the last 24h, so a bare add dsh-llm-vision (latest) could silently install the previous release on launch day. Restart the GUI once after install — plugins load at boot, so the plugin and its settings card appear only after restart; later config changes never need one. Requires dsh >= 0.1.2-alpha.1 (the session log format it reads).

Compatibility

dsh >= 0.1.2-alpha.1; Node per dsh; OpenAI-compatible vision endpoints (free presets for Zhipu / Gemini / DashScope; any compatible endpoint configurable in the settings card).

Details

Recent updates

describe_image (normal + critical lenses, batch up to 8) / extract_text OCR / llm_vision_check diagnostics; GUI-only config; paste/drag/drop image rewrite; image bytes never enter the session log.

FAQ

How do I start?
Install the pinned release, restart the GUI, open Settings → Plugins → llm-vision, pick a Provider preset (zhipu / gemini / dashscope fill the endpoints — free routes), paste an API key (stored in the owner-only settings document, never shown again), then paste or drop an image into the composer.
Which models can use it?
Text-only models such as DeepSeek V4 or GLM text series — the plugin forwards images to a separate OpenAI-compatible vision/OCR endpoint and returns text.
Does the image enter the session log?
No. The image itself never crosses into the conversation; only the returned text description is recorded.

Alternatives

apheli0os/deepseek-harness-orchestrate · MoneShadow/dsh-plugin-vision · wdwind/dsh-vision-no-vision

More plugins in uncategorized

Browse more in uncategorized

Guides for uncategorized plugins