jing-hy/picturereader

Image "reading" for text-only models: downscale + reduce color depth + structure/color fingerprints into text grids fed back to the conversation, letting the model zoom, sample and OCR autonomously like a multimodal model; fully local with zero external model dependency, ships an image-reading methodology skill and optional PaddleOCR.

Vision & Multimodal ★ 0 updated —
View on GitHub ↗

Install

dsh plugin --profile web add github:jing-hy/picturereader

Details