xwh-01/dsh-mediacrawler
Installable DeepSeek Harness profile bundle and bounded MCP adapter for MediaCra
An installable DSH profile bundle plus a bounded stdio MCP adapter that connects DeepSeek Harness to a separately installed MediaCrawler checkout, for sources where a logged-in collector is needed: search, post/video detail, creator feeds and explicitly enabled comments on Xiaohongshu, Douyin, Kuaishou, Bilibili, Weibo, Tieba and Zhihu. Each run is supervised, persisted and exposed through twelve MCP tools (check, collect, status, runs, result, delete_run, cleanup, stop, logs, artifacts, preview, export) surfaced in Harness as mcpmediacrawler<tool>. The bundle also mounts a packaged mediacrawler-collector Skill that guides the agent through checking the runtime, starting a small collection, polling status and exporting results.
Install
npx --yes @deepseek-ai/dsh@0.1.0-rc.6 plugin --profile web add "github:xwh-01/dsh-mediacrawler#v0.3.0"README step 3 quoted verbatim, pinned to the tested DSH 0.1.0-rc.6 release and bundle tag v0.3.0. The README's prerequisites are Python 3.11+, Node.js 22.19+ on the 22.x line or 24+ with pnpm on PATH, Google Chrome, a separate MediaCrawler checkout with its own working Python environment, and DeepSeek Harness. The Python MCP runtime installs separately into its own virtual environment: pip install "dsh-mediacrawler @ git+https://github.com/xwh-01/dsh-mediacrawler.git@v0.3.0", and DSH_MEDIACRAWLER_PYTHON must point at that interpreter. Verify with npx --yes @deepseek-ai/dsh@0.1.0-rc.6 --profile web --dump-config, which should contain a # == dsh-mediacrawler layer. Uninstall: npx --yes @deepseek-ai/dsh@0.1.0-rc.6 plugin --profile web remove dsh-mediacrawler.
Compatibility
Requires a separately installed MediaCrawler checkout, its browser dependencies, Chrome and a Python 3.11+ virtual environment, plus pnpm for the DSH profile. The README is explicit that this is an adapter, not a MediaCrawler fork: it does not copy or modify MediaCrawler source and does not change its license, and MediaCrawler plus its browser dependencies are intentionally not vendored. Adapter state defaults to ~/.dsh-mediacrawler.
Details
- Repo: xwh-01/dsh-mediacrawler
- Category: Files & Data
- Stars: 3
- Version: GitHub release v0.3.0 (README-pinned); tested against DSH 0.1.0-rc.6
- Last push: 2026-08-14
- First seen: 2026-08-14
Recent updates
The README documents install paths, the twelve tool contracts and runtime behavior rather than release-by-release notes; the pinned, tested release at the time of writing is v0.3.0.
FAQ
- Is this a MediaCrawler fork?
- No — the README states it is an adapter that neither copies nor modifies MediaCrawler source and does not change MediaCrawler's license.
- When should I use this instead of a web-search plugin?
- Per the README, use the Harness web-search providers for quick facts and already-indexed pages, and this adapter when the task needs logged-in platform records, creator feeds, comments or nested replies, or a durable reproducible export.
- What are the runtime prerequisites?
- Python 3.11+, Node 22.19+/24+ with pnpm, Google Chrome, a separate MediaCrawler checkout with its own Python environment, and DSH; the adapter's own Python runtime goes in a dedicated virtual environment.
Alternatives
Aidenwu0209/dsh-Unlimited-OCR-Skill · bill9109/dsh-101 · bwndlct/dsh-session-export