xwh-01/dsh-mediacrawler

Installable DeepSeek Harness profile bundle and bounded MCP adapter for MediaCra

An installable DSH profile bundle plus a bounded stdio MCP adapter that connects DeepSeek Harness to a separately installed MediaCrawler checkout, for sources where a logged-in collector is needed: search, post/video detail, creator feeds and explicitly enabled comments on Xiaohongshu, Douyin, Kuaishou, Bilibili, Weibo, Tieba and Zhihu. Each run is supervised, persisted and exposed through twelve MCP tools (check, collect, status, runs, result, delete_run, cleanup, stop, logs, artifacts, preview, export) surfaced in Harness as mcpmediacrawler<tool>. The bundle also mounts a packaged mediacrawler-collector Skill that guides the agent through checking the runtime, starting a small collection, polling status and exporting results.

Files & Data ★ 3 updated 2026-08-14 — untested
View on GitHub ↗

Install

npx --yes @deepseek-ai/dsh@0.1.0-rc.6 plugin --profile web add "github:xwh-01/dsh-mediacrawler#v0.3.0"

README step 3 quoted verbatim, pinned to the tested DSH 0.1.0-rc.6 release and bundle tag v0.3.0. The README's prerequisites are Python 3.11+, Node.js 22.19+ on the 22.x line or 24+ with pnpm on PATH, Google Chrome, a separate MediaCrawler checkout with its own working Python environment, and DeepSeek Harness. The Python MCP runtime installs separately into its own virtual environment: pip install "dsh-mediacrawler @ git+https://github.com/xwh-01/dsh-mediacrawler.git@v0.3.0", and DSH_MEDIACRAWLER_PYTHON must point at that interpreter. Verify with npx --yes @deepseek-ai/dsh@0.1.0-rc.6 --profile web --dump-config, which should contain a # == dsh-mediacrawler layer. Uninstall: npx --yes @deepseek-ai/dsh@0.1.0-rc.6 plugin --profile web remove dsh-mediacrawler.

Compatibility

Requires a separately installed MediaCrawler checkout, its browser dependencies, Chrome and a Python 3.11+ virtual environment, plus pnpm for the DSH profile. The README is explicit that this is an adapter, not a MediaCrawler fork: it does not copy or modify MediaCrawler source and does not change its license, and MediaCrawler plus its browser dependencies are intentionally not vendored. Adapter state defaults to ~/.dsh-mediacrawler.

Details

Recent updates

The README documents install paths, the twelve tool contracts and runtime behavior rather than release-by-release notes; the pinned, tested release at the time of writing is v0.3.0.

FAQ

Is this a MediaCrawler fork?
No — the README states it is an adapter that neither copies nor modifies MediaCrawler source and does not change MediaCrawler's license.
When should I use this instead of a web-search plugin?
Per the README, use the Harness web-search providers for quick facts and already-indexed pages, and this adapter when the task needs logged-in platform records, creator feeds, comments or nested replies, or a durable reproducible export.
What are the runtime prerequisites?
Python 3.11+, Node 22.19+/24+ with pnpm, Google Chrome, a separate MediaCrawler checkout with its own Python environment, and DSH; the adapter's own Python runtime goes in a dedicated virtual environment.

Alternatives

Aidenwu0209/dsh-Unlimited-OCR-Skill · bill9109/dsh-101 · bwndlct/dsh-session-export

More plugins in Files & Data

Browse more in Files & Data

Guides for Files & Data plugins