ch1bug/dsh-mimo-agent-tools
Turns the Xiaomi MiMo API into model tools for a DeepSeek Harness agent: mimo_search for web search with cited sources, mimo_think for a full reasoning chain plus answer, mimo_json for structured JSON output, and multimodal understanding through mimo_vision (images), mimo_audio (audio/transcription), mimo_video, mimo_asr speech-to-text and mimo_tts text-to-speech to a wav/mp3 file with preset voices or free-form voice design. The README also documents an audio-tools skill and a Python driver that the plugin calls.
Install
⚠️ Install command not yet confirmed — check the README on GitHub for the exact command.
Compatibility
README: a DSH Cordis plugin backed by the OpenAI-compatible endpoint https://api.xiaomimimo.com/v1; requires a Xiaomi MiMo API key and a Python driver on MIMO_DRIVER. The README notes dsh plugin add installs the bundle into the profile's node_modules so peer dependencies resolve against the running harness. Repo verified 2026-09-16 (5★, MIT, last push 2026-08-19).
Details
- Repo: ch1bug/dsh-mimo-agent-tools
- Category: Agent Capabilities
- Stars: 5
- Version: Source/GitHub install (no registry version asserted; README notes @deepseek-ai/dsh-tools is already loaded in the DSH process)
- Last push: 2026-08-19
- First seen: 2026-08-14
Recent updates
The README documents the tool table and the MiMo API notes rather than a release-by-release table.
FAQ
- How do I install it?
- Install the Python driver to the README's default path and point MIMO_DRIVER at it, then add the bundle from your checkout with dsh plugin --profile web add /path/to/dsh-mimo-agent-tools; the README's github:you/... line is an owner placeholder, so check it before using.
- What does it add?
- Per the README: web search, deep thinking, structured JSON output, and image/audio/video understanding plus speech-to-text and text-to-speech, all backed by the MiMo API.
- Which models does it use?
- The README's tool table names mimo-v2.5-pro and mimo-v2.5 for search/thinking/JSON/vision/audio/video, mimo-v2.5-asr for speech-to-text and mimo-v2.5-tts or -voicedesign for text-to-speech.