tianji-qingtian/dsh-model-router

Model router & cost optimizer for DeepSeek Harness: flash direct answer to simple questions, automatic fault downgrade, session token/cache/cost real-time panel | Model router & cost optimizer for DeepSeek Harness: flash

Model Router & Cost Optimizer: answers simple questions directly on the cheap model (zero prefix, no cache tax) and degrades gracefully on transient provider failures. Cheap-model judge routing (zero-prefix flash judge call, SIMPLE/AGENTIC), ask-before-quick-answering via built-in question UI, direct quick-answer one-shot stream on the cheap model, vision-aware routing to image-capable models, model picker respected (never overrides composer selection), auto/off toggle, automatic fallback on RATE_LIMIT/SERVER/TIMEOUT/EMPTY_RESPONSE, and real per-session usage metering (tokens, cache hits, estimated cost) in a composer dock panel.

Web UI Enhancements ★ 6 updated 2026-09-03 ✅ runtime-tested
View on GitHub ↗

Install

dsh plugin --profile web add "github:tianji-qingtian/dsh-model-router#v0.9.4"

GitHub install per README (fresh fetch 2026-09-07): dsh plugin --profile web add "github:tianji-qingtian/dsh-model-router#v0.9.4" (prefer a release tag; #main tracks the latest commit; built lib/ artifacts are committed so no build script runs at install time). Prerequisite: the dsh CLI must be on PATH — npm install -g @deepseek-ai/dsh (harness 0.1.1-rc.2 or newer), or prefix commands with npx @deepseek-ai/dsh. npm package dsh-model-router 0.6.2 exists (registry-verified 2026-08-26) but the README documents the git tag install.

Compatibility

DeepSeek Harness 0.1.1-rc.2 or newer (dsh CLI on PATH). Web profile. The harness is in developer preview — compatibility-breaking changes expected.

Details

Recent updates

Pinned to release tag #v0.9.4 (README-preferred): v0.9.4 is a meta-only release — lib/ byte-identical to v0.9.3, only harness peer ranges widened (accepts 0.1.2-rc.1). v0.9.3 feature set: cheap-model judge routing; quick-answer question UI; vision-aware routing; auto/off toggle; fallback retry; projection-based usage metering with dock panel.

FAQ

How are simple questions detected?
A zero-prefix flash judge call (64-token cap, thinking off) classifies each request as SIMPLE or AGENTIC; heavy work by keywords/length goes straight to the main model with zero added latency.
Does it override my model picker?
No — the router never overrides the model chosen in the composer picker or via selectModel; quick answers always run on the session provider's own cheap model.
What happens on provider failures?
Transient failures (RATE_LIMIT, SERVER, TIMEOUT, EMPTY_RESPONSE) degrade the turn to the cheap model and retry once; anything else delegates to the provider's own retry policy.

Alternatives

SnowAmberX/dsh-role-router · BruceLanLan/dsh-tier-router · JuneLearn/dsh-reasoning-settings

More plugins in Web UI Enhancements

Browse more in Web UI Enhancements

Guides for Web UI Enhancements plugins