dsh-orquestrator
DeepSeek Harness plugin: on task send, pick the model your subagents run on. Enforced in code on every child DSH starts, workflow agents included: the model, a reasoning-effort ceiling and an output-token cap.
- Stars
- 0
- Language
- TypeScript
- Created
- Sep 30, 2026
- Updated
- Oct 4, 2026
Introduction
dsh-orquestrator
A plugin for DeepSeek Harness (DSH). When you send a new task, a dialog in the stock DSH look asks one question:
Should subagents run on a different model than the one selected for the main agent?
If yes, you pick the model, and the plugin enforces it in code, on every child DSH
starts for that session: the subagent tools, but also the agents a workflow starts,
ralph rounds, one-shot background jobs and agent teams. The main agent cannot talk its way
around it: it is not an instruction to the model, it is the plugin standing in the doors DSH
starts children through. See Which delegations are covered.
Every governed child also gets a reasoning-effort ceiling and an output-token cap,
because DSH otherwise runs a re-routed child at its route's default (max on many setups):
the cause of subagents that burn their whole budget on one edge case. See
Reasoning effort.
Cancel, Escape and the close button send the task exactly as DSH always did. Nothing else about DSH changes.
| Dark | Light |
|---|---|
![]() | ![]() |
With a model picked from the composer's own list:

Português: README.pt-BR.md
0.5.0 removed the independent reviewer that 0.2 to 0.4 offered next to the model choice. What governs the models stayed, and it now is the plugin's only mechanism. Stored choices and patch files written for 0.4 keep working (the reviewer fields are ignored, with a warning). Why: Why the reviewer was removed.
Install
dsh plugin --profile web add github:frederico-kluser/dsh-orquestrator
dsh --profile web # (re)start it: see the note below
Restart dsh after installing or updating, not only the page. A plugin's host half is loaded when dsh
starts and its browser half when the page loads, so refreshing the page alone leaves the old host half running
(and the old guard with it). From 0.5.1 the two halves also keep talking to 0.2 to 0.4 halves while you do
(a disabled reviewer block stays on the wire for that), so a mixed state still saves; before that, it failed with
"config does not match the expected shape".
From a local clone: dsh plugin --profile web add /path/to/dsh-orquestrator.
Tested on DSH 0.1.6-alpha.2 (Node 24, pnpm 11). The committed lib/ is the
build output, so no build step is needed to install.
Use
Type a task in the composer and send it. The dialog appears once per new task:
- Subagent model: turn it on and pick a model from the same provider-grouped list the composer's model seat uses. It applies to every subagent, including the agents a workflow starts. Off means subagents keep the main agent's model. The dialog shows short, dated notes for models that need them (for example: MiMo-V2.6-Pro can take minutes per turn at high effort; GLM 5.3 is text only).
- Reasoning effort (collapsed, shown once a model is picked): how hard the model may think. It defaults to the recommended level for the model; open it to see or change it.
- There is no "do not ask again": the modal is raised for every new task and nothing can silence it. One answer never hides it from a later task or from another conversation. The last confirmed choice only pre-fills the dialog.
- Cancel / Esc / ✕: send the task with stock behavior and forget any stored choice.
/orquestrar opens the same dialog on demand (to change or clear the stored choice).
The dialog is skipped for anything that is not a new task: steering a running turn,
sub-agent conversations and / command lines.
Which delegations are covered
DSH has more ways to start a child than the two subagent tools. All of them go through two
doors, SubagentRuntime.start() and startContinuable(), and the start guard
(src/guard.ts) stands in both. For a session with a confirmed choice (stored, an ancestor's, or
defaults) it plans every child: the picked route, the effort the user asked for or the model's
ceiling, and the output-token cap, handed to DSH as the child's agentOptions.
| How DSH starts the child | Model, effort ceiling, token cap |
|---|---|
subagent and subagent_fork tools (the standard preset) | yes |
subagent as a one-shot background job (backgroundMode: one-shot, run_in_background: true) | yes |
the workflow tool: every agent() call of the script | yes |
ralph (off in the standard preset; it runs on the workflow engine) | yes |
| agent teams (experimental) | yes |
codex, claude-code and ACP providers | no: they run their own agents on their own models and take no agent options (the plugin logs a warning once per provider) |
| the DSH SDK provider (a separate DSH child runtime) | yes when a subagent model is picked (it takes the route, effort and token limit); with no model picked its child keeps the provider's own model |
What was run live, on the three target models: the subagent tool (foreground and background),
subagent_fork, and the workflow tool (with the default override, with keep and with the
enforcement off), headless and through the dialog in a real browser. The other rows follow from the
doors they use, which the contract tests pin against the DSH source (ralph runs on the workflow
engine, a one-shot job and the team call start / startContinuable, the SDK provider takes agent
options). No live run used a one-shot background job, ralph, an agent team, the SDK provider or
codex / claude-code / ACP.
Why a guard and not a tool wrapper: a wrapper never saw a workflow call's agents, because the
engine starts them through the service itself. In the session that exposed it, 34 workflow
agents ran on Claude Sonnet 5.5 at max (about 7.2 million output tokens and 1.26 billion
cache-read tokens) although DeepSeek V4.1 Flash was confirmed for subagents. Version 0.3 and
older have this hole; the validation page reproduces it on an isolated DSH and shows it closed.
What a governed child gets: the model the user picked, the effort the user picked (or the model's ceiling) and the output-token cap. The main agent is never touched, and a session with no confirmed choice is never touched.
If the model you confirmed is gone (renamed or removed from your DSH settings after you confirmed it),
the child's start is rejected with a message that says so and what to do (/orquestrar, or cancel the
dialog). The alternatives are worse: running the child on the main agent's model is the bug this
design exists to prevent, and forcing the dead route makes every workflow agent fail into a silent null.
A model the caller names itself (agent({ provider, model }) in a workflow script) loses to the
user's pick by default (children.explicitModel: override): the dialog is the user's explicit
instruction. keep lets the caller's model stand, under the same ceilings. A choice with an effort but
no model (defaults.workerEffort alone) leaves children on the main agent's model under that level,
and a model a script names itself stands.
Configuration
Everything is optional; without configuration the plugin does nothing until a user
confirms the dialog. Add overrides to your profile's cordis.patch.yml
(a patch replaces the row's whole config):
- id: orquestrator
config:
# Headless/TUI/SDK sessions have no dialog: apply this to every session.
defaults:
subagentModel: { provider: azure-opencode, model: DeepSeek-V4.1-Flash }
workerEffort: medium # optional; absent = the recommended level for the model
effort: # ceiling on the reasoning effort of every subagent; false turns it off
worker: medium
limits: # output tokens per model request, reasoning included; false = no cap
workerMaxTokens: 64000
children: # the start guard; false switches the enforcement off (the dialog still stores choices)
explicitModel: override # override | keep: a model the caller names itself, e.g. agent({ model }) in a workflow script
persist: true # remember choices across restarts
stateDir: ~/.dsh/dsh-orquestrator
maxSessions: 500 # stored sessions before the oldest are pruned
Fields of 0.4 and older that belonged to the reviewer (tools, reviewerProvider, reviewerContext,
structuredVerdict, workerHandoff, maxWorkerReportChars, retryOnTokenLimit, workspaceChecks,
sensitivePaths, defaults.reviewer, effort.reviewer, limits.reviewerMaxTokens) are ignored with a
warning in the DSH log, so an existing patch file keeps loading.
Reasoning effort
DSH resolves a child's options from its parent, and when the route changes without an
effort it clears the parent's level so the new model "resolves its own default". On setups
whose routes say reasoning: max that means every re-routed child thinks at max, with the
route's full declared output ceiling (131K to 943K tokens on the routes this was built on) as its limit.
The plugin therefore asks DSH for the model's own ladder and applies a ceiling, never a
setting: a route already at or below it is left alone, and a ladder that skips rungs (GLM 5.3
offers low, high, max) gets the highest rung not above the ceiling. off is never picked
on its own. The ceiling is, in order: the level picked in the dialog (it may be above the
ceiling), effort.worker, the model's row in src/models.ts (dated, with its
sources), and medium.
| Model | Subagent ceiling |
|---|---|
| DeepSeek V4.1 Flash | medium |
| MiMo-V2.6-Pro | low |
| GLM 5.3 / GLM 5.3 Flash | high |
| Claude Sonnet / Opus | high |
| Anything else | medium |
limits caps the output tokens of a single model request (reasoning included) and only ever
lowers a ceiling the model is known to have. Why these numbers, and which study recommendations
were left out and why: docs/estudos/.
Limits
- The plugin controls which model runs and how hard it thinks. It does not check what a subagent produces. The main agent reads the result as DSH always delivered it; if you want a check, ask the main agent to verify, or run the project's own tests. (An independent reviewer used to do this; see below.)
- Providers that cannot take agent options (
codex,claude-code, ACP) keep the models they run on; the plugin logs a warning (once per provider) and does not touch their children. A provider that runs its own default route (the SDK provider) is left alone when no subagent model is picked, because the plugin cannot know what its child runs on. - The output-token cap is not durable for
continuablechildren. When DSH later resumes a finished child (a follow-up message after it released the child, or after a restart) it rebuilds the child's options from the recorded descriptor, which holds the provider, model and effort but not the token limit. The effort ceiling survives; the cap applies to the first run only. Fixing it needs another seam (see decisoes.md, N21). - A model the main agent names in a
subagentcall (DSH's model selection, on in the standard preset) is ignored while a choice is confirmed;overrideapplies the same rule to every other caller. A token limit a caller sets (an operator's tool row, a team roster) is never raised: the smaller of the caller's and the plugin's stands. - Unknown top-level configuration fields are ignored with a warning in the DSH log (a mis-indented
explicitModel: keepwould otherwise run asoverride). - DSH has no hook around a child start (
subagent/startfires after the child exists), so the start guard installs its ownstartandstartContinuableon the service instance. The contract tests (test/contract/) pin that assumption against a DSH checkout and a realSubagentRuntime, and fail first when DSH changes it. They run on a machine that has a DSH checkout (DSH_CHECKOUT), not in CI. If the service cannot be wrapped the plugin fails to load; it never half works.children: falsetakes the guard out. - The reasoning-effort ceiling and the token cap apply to the children the plugin governs, never to the main agent. A model the LLM runtime cannot describe but can call keeps exactly the options the user picked (the log says so); one it cannot call at all is rejected, as described above.
- The model advice in the dialog (notes, ceilings) is dated data from studies of September and October 2026 and goes stale in weeks; a model it does not know gets the generic ceiling and no notes.
- Portuguese and Chinese strings ship as dictionaries. They appear when DSH (or another plugin) has registered that language; this plugin never registers a language itself, to avoid clashing with the plugin that owns it.
Security model
The plugin runs no code of the subagents and starts no process of its own. It changes only the options DSH starts a child with. What it does: serves its one route behind the DSH trust fence (Host/Origin fence and browser authentication) with strict wire validation, keeps its state in an owner-only file that holds provider and model ids and no credentials, and never reads, writes or forwards API keys. What it does not do: it provides no sandbox, filters no network and controls no process environment. What subagents may do is decided by the permission preset of the session, exactly as without the plugin.
Why the reviewer was removed
Versions 0.2 to 0.4 also offered an independent reviewer: a second model that checked each subagent's work, fixed what was broken and delivered the result to the main agent in the subagent's place. Version 0.5.0 removed it, for three reasons:
- It was an LLM judging an LLM. What the plugin enforces in code (which model runs, how hard it
thinks, how much it may write) is deterministic, and the main agent cannot bypass it. A reviewer's
APPROVEDis an opinion, and the plugin could only check its format and its coherence. - It cost about twice as much and made every delegation wait for a second model, and it covered two of the paths DSH starts children through, not the agents of a workflow.
- It needed a second mechanism. The reviewer required wrapping the
subagenttools; the start guard alone already governs every path (verified live, with the wrapper out of the way), so keeping both only kept two ways to get the same result.
The studies, the reviewer protocol and the decisions behind it stay in the repository as history: docs/estudos/, docs/pesquisa/padrao-revisor.md (Portuguese) and decision D16 in docs/estudos/decisoes.md. Version 0.4.0 is the last one that has it.
Verified
Version 0.5.0 was validated on a real DSH 0.1.6-alpha.2 with the three target models only: GLM 5.3
(main agent), DeepSeek V4.1 Flash (subagents) and MiMo-V2.6-Pro (the model a workflow script names
itself). Headless, eight scenarios read back from the session logs: the workflow agents on the
picked model at the model's ceiling with a 64 000-token cap; explicitModel: override and keep;
the enforcement switched off; the subagent tool in the background and subagent_fork governed with
no tool wrapper at all; and a confirmed model that no longer exists rejected with a message that names
it, through a workflow and through the tool. Through the dialog in a real browser: 65 checks, including
a real delegation and a real workflow whose children were read back from the logs. A stored choice
written by 0.4.0 (with its reviewer block) loaded and was rewritten in the new shape. Evidence and
findings, including what was not covered: docs/validation/README.md.
Develop
pnpm install
pnpm run check # typecheck + build + tests
pnpm run check:lib # the committed lib/ must equal a fresh build (run it after committing lib/)
DSH_CHECKOUT=/path/to/deepseek-harness pnpm test # also pins the DSH seams this plugin uses
Live validation against a real DSH uses only the three target models and fails if any other
one runs: scripts/e2e/run-workflow.sh (the workflow, subagent and subagent_fork tools and the
start guard, headless), with session-config.mjs reading back what each session was asked, and
scripts/e2e/ui-e2e.mjs (browser). scripts/e2e/setup-isolated-home.sh builds the isolated
DSH_HOME they expect and scripts/e2e/with-keys.sh runs them with only the two API keys the
three models need. See docs/validation/README.md.
Layout: src/ (host), src/client/ (browser), test/ (unit, integration,
contract), scripts/e2e/ (headless and browser runs against a real DSH),
docs/ (design, validation, and estudos/: the studies, their digest and the decision log).
License
MIT

