dsh-force-compact
Aggressive context compaction for local-first agents. Runs Qwen3.8‑27B on self‑hosted llama.cpp at low context, shrinking history so the live prompt stays small, fast, and private—delivering a big‑window experience without API cost or data egress. 面向本地的激进上下文压缩插件。自托管 llama.cpp 低上下文运行 Qwen3.8‑27B,不断收缩历史、保持常驻 prompt 小而快,兼顾隐私与大窗口体验,零 API 成本、数据不出本机。
- Stars
- 10
- Language
- JavaScript
- Created
- Aug 23, 2026
- Updated
- Oct 2, 2026
Introduction
dsh-force-compact
Aggressive, local-first context compaction for DeepSeek Harness agents.
A DSH Cordis function plugin that keeps the agent's working context lean by design: serve
Qwen3.8‑27B on a self-hosted llama.cpp with a modest context, and the plugin shrinks the
conversation itself — a large-window feel with no API cost and no data egress.
Why
- Self-hosted inference — the agent talks to a local OpenAI-compatible llama.cpp server through the standard DeepSeek adapter; no separate adapter needed.
- Low context, high signal — instead of fighting a small cap, the plugin shrinks the conversation, so the agent reasons over a tight prompt while keeping deep memory in the compressed head.
- Think off for compactions, passthrough everywhere else —
disableThinking: true(default) turns thinking off on this plugin's own compaction summarization call only; every other model request rides the machine's configuration unchanged. - Private & free — no per-token billing, no egress.
What it does
Two compaction engines coexist behind one facade (resolveCompaction), transparent to callers:
| Engine | Used when | Notes |
|---|---|---|
| Official | the compaction service resolves in the agent realm | Preferred; delegates to compaction/basic. |
| Builtin | automatic fallback (typical standard preset isolates the service) | Self-contained persistent transaction on ctx.sessions / ctx.llm.stream / ctx.tokenMeter; reuses the official compaction/* event vocabulary, so it replays safely across builds. |
No toggling — official wins when reachable, builtin takes over otherwise.
Trigger points
- Per-request guard (
agent/pre-step) — reads the session's projected context tokens (the exact number the harness renders bottom-right). AtautoThresholdTokensit rejects the outgoing request and compacts the head instead, retaining the latestretainLatestTokensverbatim. Below the threshold the request proceeds. - Turn end / idle (
agent/status→idle) — when the agent quiesces, optionally compacts viacompactNow(gate:turnEndForceCompactionEnabled). - Manual
/force-compact— immediatecompactNowwhen idle; when busy it queues a process-local flag consumed at the next model step. Loaded lazily — see "Command availability" under Install. In the/picker the row carries the same face as the first-party commands: an icon, a localized name and a localized description (the official presentation table only covers first-party commands, so the client half attaches the face itself — seeAGENTS.md, "指令行的官方外观"). session/flush— the awaited durability checkpoint.
Every path funnels into the single "compaction result landed in the session" boundary — the same point where the LiveUI signal fires.
Decision key is projectedTokens (provider-anchored, same figure as the UI corner), so the
plugin never drifts from what you see; the threshold-aware shrink gate skips summarizer calls
that provably cannot pull the session below the threshold (kills the low-threshold dead loop).
The builtin transaction bills shadowedTokenCount from the same tokenMeter.measure
per-node prices the official engine uses, so the meter's collapse protocol settles the drop
correctly — the bottom-right counter goes down after compaction.
Thinking control: scoped to compactions
Since the 2026-08 semantics revision, disableThinking controls one thing: whether this
plugin's own summarization call (engine/builtin.js → engine/summarizer.js →
ctx.llm.stream) carries reasoningEffort:'off'. Everything else is untouched:
| Call site | With disableThinking: true |
|---|---|
| Builtin-engine summarization call | Carries reasoningEffort:'off' |
| Every other model request (business, sub-agents, tools, other plugins) | Machine's LlmCallConfig unchanged |
Official compaction service calls | Not routed through any plugin seam — unaffected |
A route that cannot express off degrades instead of failing. Not every catalog model
declares an off level — pi-ai's thinkingLevelMap may map it to null, as
opencode-go/deepseek-v4.1-flash does (it accepts only low / high / max, while
deepseek/deepseek-flash accepts off too). The summarizer probes the route's capability
before the call (llm.resolveCallConfig, a detached lookup with no provider I/O) and, when
off is unavailable, uses the cheapest level the route does accept; a model with no
reasoning support at all gets no effort field. Because a thinking route bills its reasoning
trace against the same maxTokens cap, a fixed reasoning allowance is added on top of
maxSummaryTokens (documented as the cap on the summary text) so the summary still fits.
Both are logged once per route:
summarization route opencode-go/deepseek-v4.1-flash does not support reasoning effort 'off'
— using the cheapest supported level 'low' instead; compaction continues
Before this guard, such a route failed every compaction outright (the harness rejects the unsupported effort before any provider I/O).
When the target is a llama.cpp / OpenAI-compatible endpoint, the
thinking: { type: 'disabled' } field the adapter emits is silently ignored there — so the
summarizer ALSO stamps the llama.cpp-native top-level reasoning_effort: "none", gated on
the exact same condition. One options object carries BOTH fields:
| Endpoint family | Reads | Result |
|---|---|---|
| Real DeepSeek API | reasoningEffort:'off' → thinking:{type:'disabled'} | Thinking off ✅ |
| llama.cpp / OAI-compatible | reasoning_effort:"none" (top-level) | enable_thinking=false ✅ |
Each family tolerates-and-ignores the foreign key, so emitting both is harmless. The field is
stamped in src/engine/summarizer.js (immediately before llm.stream(options)), NOT in the
llm/stream waterfall — a prior draft injected there but proved ineffective structurally
(middle-layer returns are discarded; in-place seed mutation crashes the host); see the
src/hooks/wire-rewrite.js module header for the write-up. That hook now serves only the
LiveUI watermark role.
Need thinking off on business calls too? Set your provider's reasoningEffort at the
request-header level — the plugin deliberately stays out of that decision.
Observability: per-attempt audit lines
Every summarization attempt logs two lines (visible at the default debug: true) — the
durable proof of the scoping decision and its wire fields, without capturing traffic:
[force-compact] <sessionId>: compaction thinking-policy — settings.disableThinking=true → extra.reasoningEffort='off' (this summarization call carries thinking-OFF)
[force-compact] <sessionId>: summarization wire-fields → <provider>/<model>: reasoningEffort='off' + reasoning_effort="none" (llama.cpp-native wire field)
- Line 1 (
engine/builtin.js) records wheredisableThinkingis read and routed into the call options; with the setting off it records machine default. - Line 2 (
engine/summarizer.js) records both wire fields exactly as they leave the options object, plus resolved provider/model; unstamped fields are labeled(absent…).
Empirically grounded: probed against a local llama.cpp endpoint, a baseline request returned
populated reasoning_content (the model thinks by default), while the same request with
top-level reasoning_effort:"none" returned none at all — the field genuinely disables
thinking there, and business calls (which omit it) keep thinking.
LiveUI status
A tiny host→client messenger (the liveUi settings field mirrored live to the browser) replaces the
official running label's leading phrase beside the turn. Only the text changes: the harness's own
elapsed-time text, font, and colour are left exactly as shipped.
[强制压缩中>>>]— just before a compaction commits;[压缩完成!]— the instant a compaction lands; 3 s later a fresh random working line takes over;- a rotating playful one-liner — otherwise, on every model request;
- Restored at conversation end — when the agent goes
idle(the turn is fully done), an empty text (isImportant) is pushed: the client puts the official label back and drops its replacement prefix. Replaces the former conversation-START forced working override (removed 2026-09).
Harness 0.2.0 moved that line into its own RunningStatus component — the old
button[data-turn-process] > span now renders only settled turns — anchored on the
container attribute data-chat-running:
<div data-chat-running>
<span role="status" aria-live="polite">深度求索中</span> ← visually-hidden announcement
<span .runningDivider>
<span .runningContent>
<span .runningIcon>…whale animation (APNG mask + SVG fallback)…</span>
<TextShimmer data-shimmer>深度求索中,用时1分14秒 ···</TextShimmer>
</span>
</div>
So the client half substitutes just the leading phrase: it uses the announcement node's text
as an anchor, splits the text at their common prefix, and keeps the remainder — the harness clock
tail, the official trailing ··· included — verbatim. The whale animation icon and the divider
are left exactly as they are: no colours, no font changes, no layout. The announcement node is
never touched, so screen readers keep the official text.
A MutationObserver, connected only while a phase is active, writes both copies of the
sentence: TextShimmer renders it twice (a real text node plus an aria-hidden animated highlight
copy whose glyphs come from CSS ::after { content: attr(data-shimmer-text) }), so patching only
the text node would let the sweep reveal the official wording. The observer re-applies in the same
microtask in which React rewrites the label, so the clock keeps ticking and nothing polls.
Label text follows the app language: the host writes a locale-independent textId (the phase name or
working.N) plus the canonical Chinese text, and the client half resolves it through its own
ctx.locale zh/en/ja/ko dictionaries — an English UI shows English one-liners, a Chinese UI keeps the
originals.
Badge text follows the app language: the host writes a locale-independent textId
(phase name or working.N) alongside the canonical text, and the client half maps
it to zh/en via its ctx.locale dictionaries — English UI shows English one-liners,
Chinese UI shows the original Chinese.
Publishers are fail-safe: a messenger glitch can never disturb the actual compaction.
How it works
agent/request(payload, next) # every model request
return await next() # pure pass-through (thinking-off scopes
# ONLY to the plugin's own summarizer)
agent/pre-step(payload, next) # before each model step
projectedTokens >= autoThresholdTokens?
no -> next() # let the request proceed
yes -> compactRegion(head-before-retainLatestTokens, signal)
return { kind: "reject" } # no model request this step
agent/status({ agent, status }) # lifecycle transition
status === "idle" && turnEndForceCompactionEnabled?
-> compactNow(agent, freshSignal) # turn-end compaction
session/flush(session) # durability checkpoint
select region -> project messages -> preview + shrink gate
-> compaction.compactRegion(start, end, agent, signal)
Supporting modules:
src/hooks/guard.js—agent/requestpure pass-through +pre-stepthreshold gate + process-local force flag (thinkingDisabledsurvives only as a legacy predicate).src/hooks/command.js— the/force-compactcommand (lazily registered).src/hooks/idle.js— turn-end forced compaction.src/hooks/wire-rewrite.js— thellm/streamLiveUI watermark hook (no wire manipulation; historical note in the module header).src/engine/region.js— head/tail-anchored region selection (with the official pairing ledger).src/engine/summarizer.js— the one-shot LLM summarizer, fully aligned with officialcompaction-basic(target resolution, prefix-cache alignment,purpose:'compaction'tag, fail-closed finish classification, usage capture).src/engine/builtin.js— the builtin persistent transaction (officialcompaction/*vocab).src/engine/checkpoint.js— preview + shrink gate + delegation to the compaction service.src/core/projected.js— provider-anchoredprojectedTokens.src/core/ui-signal.js— the LiveUI messenger.
Install
As an installable package (recommended):
# from npm (published):
npm install @falling-ts/dsh-force-compact
# from git:
dsh plugin --profile web add github:falling-ts/dsh-force-compact
# from a local checkout:
dsh plugin --profile web add ./dsh-force-compact
Or, from a local checkout, as a --patch overlay without installing:
dsh web --patch dsh-force-compact/cordis.patch.yml
The plugin is loaded iff ~/.dsh/logs/dsh-force-compact.log gains:
[force-compact] debug logging enabled — writing [force-compact] lines to <absolute path>
Command availability — /force-compact loads lazily
The commands service arrives with the agent-presets plane, after the plugin's boot-time
apply, so registration happens at the first guarded-listener activation
(agent/request / agent/pre-step / agent/status / session/flush), settling
permanently on the first success. Practical effect: after (re)starting the instance, a
fresh session's / picker does NOT show /force-compact until that session makes its first
model request — send any one message, then the command is registered process-wide.
- Success:
[force-compact] /force-compact command registered (deferred) commandspermanently absent: one… still UNREGISTERED 10 min …warn explains the empty picker. Until registered, the rest of the plugin works — degradation, not an install failure.
Verify a compaction happened:
idle compaction (builtin) shadowed N nodes (~M tokens)
builtin compaction OK — replaced span seq[A..B] (N nodes, ~K tokens) with a P-char checkpoint
compaction thinking-policy — settings.disableThinking=true → extra.reasoningEffort='off' (…)
summarization wire-fields → <provider>/<model>: reasoningEffort='off' + reasoning_effort="none" (…)
(The last two lines are the per-attempt audit pair described under "Observability".)
Settings
Namespace falling-ts-force-compact (the profile entry's loader id); values are
written to the profile's cordis.patch.yml (harness 0.1.7 onward; formerly
$DSH_HOME/settings.yaml):
| key | type | default | meaning |
|---|---|---|---|
disableThinking | boolean | true | Only the plugin's own summarization call carries reasoningEffort:'off'; everything else unchanged. |
autoThresholdTokens | number ≥ 32000 | 32000 | Default projected-token trigger for the gate. Every session without its own override uses this value; Floor 32000 (clamps back up at read time). See Per-session threshold. |
retainLatestTokens | positive int ≥ 8000 | 8000 | Retain the latest N tokens verbatim; older history is summarized in one batch. Floor 8000. Drives both the auto gate and /force-compact. |
turnEndForceCompactionEnabled | boolean | true | Compact on the agent's idle transition. |
debug | boolean | true | Emit [force-compact] diagnostics to the plugin log. |
logFile | string | ~/.dsh/logs/dsh-force-compact.log | Diagnostics destination (~ expands to home dir). |
compactionMode | 'realm' | 'global' | 'realm' | Official-service resolution strategy (priority-1 path). |
builtinEnabled | boolean | true | Gate for the builtin engine fallback. |
maxSummaryTokens | integer (1024–200000) | 1024 | Cap on the summarizer LLM maxTokens. |
summarizationTimeoutMs | integer 5000–2147483647 (ms) | 90000 | Hard wall-clock cap for ONE summarization stream (hung-stream guard). Floor 5000 (a sub-5s cap would false-abort slow local endpoints); ceiling 2147483647 because the value is scheduled through AbortSignal.timeout, which throws on a fractional delay and silently degrades a 2^31..2^32-1 delay to 1 ms. Out-of-range values are clamped and fractions truncated. |
Example — an aggressive local profile:
falling-ts-force-compact:
disableThinking: true
autoThresholdTokens: 40000 # compact sooner ⇒ keep the live prompt small
retainLatestTokens: 8000
turnEndForceCompactionEnabled: true
Without the settings service the plugin falls back to the same defaults and still compacts —
the namespace is optional, never a hard dependency.
Every token-count field above also accepts a K / M suffix in the settings form
(32K, 1M, or a plain number); parsing is decimal (32K = 32000, 1M = 1000000) and
the floors and ceilings above still clamp. The same parser backs the per-session control.
Per-session threshold
autoThresholdTokens is only the default. A single conversation can override it
without touching the shared settings document:
- The control is the chip sitting just to the right of the context-usage percentage in the composer bottom strip (the same row that carries the built-in stats pills). It shows a fixed caption — icon plus "Force-compact threshold" — so the row reads the same in every session and the number never shifts the layout. The effective value, the owning session and the overridden state live in its tooltip (and in the panel).
- Click it to open a small panel: type a threshold and click Save, or click use global default to drop the override again. Enter submits.
- The input takes a plain number or a
K/Msuffix (123K,1M,32000). Anything unparseable is refused with an inline message and nothing is written; clearing it falls back to the default. The 32000 floor applies to overrides too. - While
threshold < contextWindowthe plugin also draws WHERE compaction will fire: a red dot on the composer context ring, placed at the ring angle forthreshold / contextWindow(clockwise from twelve oclock, the same direction the ring fills), plus a red vertical line on the expanded breakdown bar at the same ratio (exactly as tall as the bar). A threshold at or above the context window draws nothing — occupancy can never reach it, so there is no trigger point to mark. - Overrides live per session id under
sessionThresholdsin this namespace and are read only by the session they belong to. Every gate (auto compaction,/force-compact, region compaction, checkpoint, idle) resolvessessionThresholds[sessionId] ?? autoThresholdTokens. - Absence is meaningful: removing the key restores the default, so an override can never masquerade as a global change.
- The chip renders in both performance-and-usage display modes (compact and detailed). The built-in stats pills hide themselves when they have nothing to report; the threshold chip stays put.
Tuning for low-context llama.cpp
Keep autoThresholdTokens comfortably below the served context: the live prompt stays
small and latency flat, while the agent keeps deep memory through the compressed head.
Pressure is measured in projected tokens (provider-anchored), so the threshold maps
predictably onto the UI figure.
Behavior notes
- Runtime dependency: the
compactionservice (preset planeagent-presets:compaction-basic), read live viactx.get('compaction'); unreachable → the builtin engine takes over (or the request proceeds). - Optional dependencies:
settings/tokenMeter/commands/llm/agentsare read viactx.get(...)with guards — a missing one degrades gracefully. - Per-request settings read: parameters are read every model request, so edits take effect on the next request without a restart.
- Signals:
agent/*Waterfalls forward the current turn's signal; thesession/flushcheckpoint and theagent/statusidle listener each mint a freshAbortController. - Persistence: durable output is the
compaction/*bracket events + asurfaceOp:replaceuser/messagecheckpoint, replay-safe across builds. - Client half:
web/client.jsadds the settings section "Force Compact" (localized labels), live-editable without restart (uSES-safe mirror). - One intentional timer: the 3 s
publishDonefallback (presentation-only, documented deviation). Otherwise the plugin is pure listeners + a process-localMapforce flag.
Screenshots

Settings page — the Force Compact section; all nine fields above are editable live without a restart.

Conversation page — the LiveUI signal rewrites the leading phrase of each running turn's label (compressing / done / a rotating working line) while the harness clock keeps ticking after it; at conversation end the official label is restored. The screenshot predates the 2026-09 colour removal: the badge is now plain grey text in the official font.
License
MIT (see LICENSE).