dsh-llm-gate
No description
- Stars
- 0
- Language
- JavaScript
- Created
- Aug 29, 2026
- Updated
- Aug 29, 2026
Introduction
dsh-llm-gate
Per-provider concurrency gate for DeepSeek Harness model requests.
If a provider can only serve a fixed number of requests at once (e.g a local llama-server with --parallel 1), every extra request is deferred by the server with nothing sent back. The client cannot tell "waiting for a slot" from "dead", and Node HTTP layer times out after 300 seconds with terminated. In practice this happens when there is overlap between a subagent and the main agent or compaction and the agent.
This plugin holds surplus requests inside dsh instead. A request waits in a FIFO queue before any HTTP request is made so no timeout is running while it waits. When a slot frees, the next request is dispatched.
Install
dsh plugin --profile web add dsh-llm-gate
Then configure the providers to gate in ~/.dsh/profiles/web/cordis.patch.yml:
- id: llm-gate
config:
providers:
llamacpp:
maxConcurrent: 1
maxQueued: 16
queueTimeoutMs: 3600000
The provider key is the route name from your llm-pi-ai.providers (or other adapter) settings. Providers not listed are not gated. Restart dsh web and open a new session.
Check the composed config with dsh --profile web --dump-config.
Settings
| Setting | Required | Meaning |
|---|---|---|
maxConcurrent | yes | Requests allowed in flight to this provider. For llama.cpp, match --parallel. |
maxQueued | no | Requests allowed to wait. Beyond this, a request fails at once with QUEUE_FULL. Default: unlimited. |
queueTimeoutMs | no | Longest a request may wait for a slot before failing with QUEUE_TIMEOUT. Default: wait indefinitely. |
Queue failures end the turn with the code shown. They are not retried by dsh-llm-retry.
What you will see
The plugin prints a line to the dsh terminal only when a request has to wait:
llm-gate: llamacpp session=a61e6e40 queued (depth 1)
llm-gate: llamacpp session=a61e6e40 dispatched after 5730ms
purpose=compaction or purpose=session-title is added for auxiliary requests. Requests that get a slot immediately print nothing.
Notes
- This gate serializes requests so it does not make a single-slot server faster. For parallelizing, give llama.cpp more slots (
--parallel 2 --kv-unified) and raisemaxConcurrentto match. - Waiting time is not counted by the adapter's
streamIdleTimeoutMsbecause the adapter is not called until the slot is acquired. You still needstreamIdleTimeoutMslarge enough for your prompt processing time (see thellm-pi-aiprovider settings). - A queued request is cancelled through its abort signal. Dropping the stream without aborting leaves the request queued until a slot frees, at which point it dispatches and is closed immediately.
- Requires the
llmservice; hooks thellm/streamwaterfall, so it covers every model request in the host: agents, subagents, compaction, and title generation.
License
MIT