Einskyle/dsh-llm-vision-bridge

Native LLM-provider vision bridge: images pasted in the chat are described by a vision model (Qwen3-VL via pi-ai/llama.cpp) and the text description is fed to text-only DeepSeek for the reply — image admission, routing and compaction all run through harness-native mechanisms, with an LRU description cache and 503 retry.

Vision & Multimodal ★ 0 updated —
View on GitHub ↗

Install

dsh plugin --profile web add github:Einskyle/dsh-llm-vision-bridge

Details