siruignaw-sys/dsh-tool-bandit-search
A search tool that learns to pick between fast and thorough search strategies using a contextual bandit (Thompson sampling), improving from real usage instead of a fixed heuristic.
dsh-tool-bandit-search replaces DeepSeek Harness's standard web_search with a search tool that learns which strategy to use through a contextual multi-armed bandit instead of one hardcoded approach. Two internal arms compete: quick (a single query capped at 5 results, fast, for simple factual lookups) and thorough (three query variants run in parallel and merged, capped at 10 results, for open-ended questions). On every call the plugin uses Thompson sampling — each arm keeps a Beta(α, β) reward distribution, both are sampled, and the higher sample wins — balancing exploration against exploitation with no training phase or manual tuning. After the call a continuous reward in [0,1] is computed from result quality (relative to that arm's own cap, so neither arm is structurally favored) and speed, then updates the chosen arm's distribution. The model never sees the arms: it just calls search(query).
Install
dsh plugin --profile web add github:siruignaw-sys/dsh-tool-bandit-searchGitHub install per README (EN primary): dsh plugin --profile web add github:siruignaw-sys/dsh-tool-bandit-search (or link:/absolute/path/to/dsh-tool-bandit-search). npm 404 verified 2026-09-04. The plugin replaces the standard web_search tool with a search tool that learns which strategy to use via a contextual multi-armed bandit.
Compatibility
DeepSeek Harness; replaces the standard web_search tool surface with one search tool; Thompson sampling over two internal arms (quick single-query ≤5 results vs thorough 3-variant parallel ≤10 results); continuous [0,1] reward from result quality + speed.
Details
- Repo: siruignaw-sys/dsh-tool-bandit-search
- Category: Tools & Capabilities
- Stars: 0
- Version: git github:siruignaw-sys/dsh-tool-bandit-search
- Last push: 2026-08-22
- First seen: 2026-08-16
Recent updates
No npm release (404 verified 2026-09-04); README documents the bandit design and reward model.
FAQ
- What are the two search strategies?
- quick — one query capped at 5 results, fast, for simple lookups; thorough — three parallel query variants merged/deduplicated, capped at 10 results, for open-ended or multi-perspective questions.
- How does it decide which to use?
- Thompson sampling: each arm keeps a Beta(α, β) distribution over its estimated reward; both are sampled per call and the higher sample's arm is used, naturally balancing exploration and exploitation.
- Does it need training data?
- No — after every call a continuous [0,1] reward (result quality relative to the arm's own cap, plus speed) updates the chosen arm's distribution, so it adapts from real usage with no separate training phase.
Alternatives
anweat/dsh-web-search-pro · crayonlu/dsh-web-search-tavily