← Back to home@falling-ts

dsh-force-compact

Aggressive context compaction for local-first agents. Runs Qwen3.8‑27B on self‑hosted llama.cpp at low context, shrinking history so the live prompt stays small, fast, and private—delivering a big‑window experience without API cost or data egress. 面向本地的激进上下文压缩插件。自托管 llama.cpp 低上下文运行 Qwen3.8‑27B,不断收缩历史、保持常驻 prompt 小而快,兼顾隐私与大窗口体验,零 API 成本、数据不出本机。

Stars
10
Language
JavaScript
Created
Aug 23, 2026
Updated
Oct 2, 2026
GitHub repo

Introduction

dsh-force-compact

Aggressive, local-first context compaction for DeepSeek Harness agents.

A DSH Cordis function plugin that keeps the agent's working context lean by design: serve Qwen3.8‑27B on a self-hosted llama.cpp with a modest context, and the plugin shrinks the conversation itself — a large-window feel with no API cost and no data egress.

中文


Why

  • Self-hosted inference — the agent talks to a local OpenAI-compatible llama.cpp server through the standard DeepSeek adapter; no separate adapter needed.
  • Low context, high signal — instead of fighting a small cap, the plugin shrinks the conversation, so the agent reasons over a tight prompt while keeping deep memory in the compressed head.
  • Think off for compactions, passthrough everywhere else — disableThinking: true (default) turns thinking off on this plugin's own compaction summarization call only; every other model request rides the machine's configuration unchanged.
  • Private & free — no per-token billing, no egress.

What it does

Two compaction engines coexist behind one facade (resolveCompaction), transparent to callers:

EngineUsed whenNotes
Officialthe compaction service resolves in the agent realmPreferred; delegates to compaction/basic.
Builtinautomatic fallback (typical standard preset isolates the service)Self-contained persistent transaction on ctx.sessions / ctx.llm.stream / ctx.tokenMeter; reuses the official compaction/* event vocabulary, so it replays safely across builds.

No toggling — official wins when reachable, builtin takes over otherwise.

Trigger points

  • Per-request guard (agent/pre-step) — reads the session's projected context tokens (the exact number the harness renders bottom-right). At autoThresholdTokens it rejects the outgoing request and compacts the head instead, retaining the latest retainLatestTokens verbatim. Below the threshold the request proceeds.
  • Turn end / idle (agent/status → idle) — when the agent quiesces, optionally compacts via compactNow (gate: turnEndForceCompactionEnabled).
  • Manual /force-compact — immediate compactNow when idle; when busy it queues a process-local flag consumed at the next model step. Loaded lazily — see "Command availability" under Install. In the / picker the row carries the same face as the first-party commands: an icon, a localized name and a localized description (the official presentation table only covers first-party commands, so the client half attaches the face itself — see AGENTS.md, "指令行的官方外观").
  • session/flush — the awaited durability checkpoint.

Every path funnels into the single "compaction result landed in the session" boundary — the same point where the LiveUI signal fires.

Decision key is projectedTokens (provider-anchored, same figure as the UI corner), so the plugin never drifts from what you see; the threshold-aware shrink gate skips summarizer calls that provably cannot pull the session below the threshold (kills the low-threshold dead loop).

The builtin transaction bills shadowedTokenCount from the same tokenMeter.measure per-node prices the official engine uses, so the meter's collapse protocol settles the drop correctly — the bottom-right counter goes down after compaction.

Thinking control: scoped to compactions

Since the 2026-08 semantics revision, disableThinking controls one thing: whether this plugin's own summarization call (engine/builtin.js → engine/summarizer.js → ctx.llm.stream) carries reasoningEffort:'off'. Everything else is untouched:

Call siteWith disableThinking: true
Builtin-engine summarization callCarries reasoningEffort:'off'
Every other model request (business, sub-agents, tools, other plugins)Machine's LlmCallConfig unchanged
Official compaction service callsNot routed through any plugin seam — unaffected

A route that cannot express off degrades instead of failing. Not every catalog model declares an off level — pi-ai's thinkingLevelMap may map it to null, as opencode-go/deepseek-v4.1-flash does (it accepts only low / high / max, while deepseek/deepseek-flash accepts off too). The summarizer probes the route's capability before the call (llm.resolveCallConfig, a detached lookup with no provider I/O) and, when off is unavailable, uses the cheapest level the route does accept; a model with no reasoning support at all gets no effort field. Because a thinking route bills its reasoning trace against the same maxTokens cap, a fixed reasoning allowance is added on top of maxSummaryTokens (documented as the cap on the summary text) so the summary still fits. Both are logged once per route:

summarization route opencode-go/deepseek-v4.1-flash does not support reasoning effort 'off'
  — using the cheapest supported level 'low' instead; compaction continues

Before this guard, such a route failed every compaction outright (the harness rejects the unsupported effort before any provider I/O).

When the target is a llama.cpp / OpenAI-compatible endpoint, the thinking: { type: 'disabled' } field the adapter emits is silently ignored there — so the summarizer ALSO stamps the llama.cpp-native top-level reasoning_effort: "none", gated on the exact same condition. One options object carries BOTH fields:

Endpoint familyReadsResult
Real DeepSeek APIreasoningEffort:'off' → thinking:{type:'disabled'}Thinking off ✅
llama.cpp / OAI-compatiblereasoning_effort:"none" (top-level)enable_thinking=false ✅

Each family tolerates-and-ignores the foreign key, so emitting both is harmless. The field is stamped in src/engine/summarizer.js (immediately before llm.stream(options)), NOT in the llm/stream waterfall — a prior draft injected there but proved ineffective structurally (middle-layer returns are discarded; in-place seed mutation crashes the host); see the src/hooks/wire-rewrite.js module header for the write-up. That hook now serves only the LiveUI watermark role.

Need thinking off on business calls too? Set your provider's reasoningEffort at the request-header level — the plugin deliberately stays out of that decision.

Observability: per-attempt audit lines

Every summarization attempt logs two lines (visible at the default debug: true) — the durable proof of the scoping decision and its wire fields, without capturing traffic:

[force-compact] <sessionId>: compaction thinking-policy — settings.disableThinking=true → extra.reasoningEffort='off' (this summarization call carries thinking-OFF)
[force-compact] <sessionId>: summarization wire-fields → <provider>/<model>: reasoningEffort='off' + reasoning_effort="none" (llama.cpp-native wire field)
  • Line 1 (engine/builtin.js) records where disableThinking is read and routed into the call options; with the setting off it records machine default.
  • Line 2 (engine/summarizer.js) records both wire fields exactly as they leave the options object, plus resolved provider/model; unstamped fields are labeled (absent…).

Empirically grounded: probed against a local llama.cpp endpoint, a baseline request returned populated reasoning_content (the model thinks by default), while the same request with top-level reasoning_effort:"none" returned none at all — the field genuinely disables thinking there, and business calls (which omit it) keep thinking.

LiveUI status

A tiny host→client messenger (the liveUi settings field mirrored live to the browser) replaces the official running label's leading phrase beside the turn. Only the text changes: the harness's own elapsed-time text, font, and colour are left exactly as shipped.

  • [强制压缩中>>>] — just before a compaction commits;
  • [压缩完成!] — the instant a compaction lands; 3 s later a fresh random working line takes over;
  • a rotating playful one-liner — otherwise, on every model request;
  • Restored at conversation end — when the agent goes idle (the turn is fully done), an empty text (isImportant) is pushed: the client puts the official label back and drops its replacement prefix. Replaces the former conversation-START forced working override (removed 2026-09).

Harness 0.2.0 moved that line into its own RunningStatus component — the old button[data-turn-process] > span now renders only settled turns — anchored on the container attribute data-chat-running:

<div data-chat-running>
  <span role="status" aria-live="polite">深度求索中</span>        ← visually-hidden announcement
  <span .runningDivider>
  <span .runningContent>
    <span .runningIcon>…whale animation (APNG mask + SVG fallback)…</span>
    <TextShimmer data-shimmer>深度求索中,用时1分14秒 ···</TextShimmer>
  </span>
</div>

So the client half substitutes just the leading phrase: it uses the announcement node's text as an anchor, splits the text at their common prefix, and keeps the remainder — the harness clock tail, the official trailing ··· included — verbatim. The whale animation icon and the divider are left exactly as they are: no colours, no font changes, no layout. The announcement node is never touched, so screen readers keep the official text.

A MutationObserver, connected only while a phase is active, writes both copies of the sentence: TextShimmer renders it twice (a real text node plus an aria-hidden animated highlight copy whose glyphs come from CSS ::after { content: attr(data-shimmer-text) }), so patching only the text node would let the sweep reveal the official wording. The observer re-applies in the same microtask in which React rewrites the label, so the clock keeps ticking and nothing polls.

Label text follows the app language: the host writes a locale-independent textId (the phase name or working.N) plus the canonical Chinese text, and the client half resolves it through its own ctx.locale zh/en/ja/ko dictionaries — an English UI shows English one-liners, a Chinese UI keeps the originals.

Badge text follows the app language: the host writes a locale-independent textId (phase name or working.N) alongside the canonical text, and the client half maps it to zh/en via its ctx.locale dictionaries — English UI shows English one-liners, Chinese UI shows the original Chinese.

Publishers are fail-safe: a messenger glitch can never disturb the actual compaction.


How it works

agent/request(payload, next)              # every model request
    return await next()                  # pure pass-through (thinking-off scopes
                                          # ONLY to the plugin's own summarizer)

agent/pre-step(payload, next)             # before each model step
    projectedTokens >= autoThresholdTokens?
        no  -> next()                     # let the request proceed
        yes -> compactRegion(head-before-retainLatestTokens, signal)
               return { kind: "reject" }  # no model request this step

agent/status({ agent, status })          # lifecycle transition
    status === "idle" && turnEndForceCompactionEnabled?
        -> compactNow(agent, freshSignal) # turn-end compaction

session/flush(session)                   # durability checkpoint
    select region -> project messages -> preview + shrink gate
    -> compaction.compactRegion(start, end, agent, signal)

Supporting modules:

  • src/hooks/guard.js — agent/request pure pass-through + pre-step threshold gate + process-local force flag (thinkingDisabled survives only as a legacy predicate).
  • src/hooks/command.js — the /force-compact command (lazily registered).
  • src/hooks/idle.js — turn-end forced compaction.
  • src/hooks/wire-rewrite.js — the llm/stream LiveUI watermark hook (no wire manipulation; historical note in the module header).
  • src/engine/region.js — head/tail-anchored region selection (with the official pairing ledger).
  • src/engine/summarizer.js — the one-shot LLM summarizer, fully aligned with official compaction-basic (target resolution, prefix-cache alignment, purpose:'compaction' tag, fail-closed finish classification, usage capture).
  • src/engine/builtin.js — the builtin persistent transaction (official compaction/* vocab).
  • src/engine/checkpoint.js — preview + shrink gate + delegation to the compaction service.
  • src/core/projected.js — provider-anchored projectedTokens.
  • src/core/ui-signal.js — the LiveUI messenger.

Install

As an installable package (recommended):

# from npm (published):
npm install @falling-ts/dsh-force-compact
# from git:
dsh plugin --profile web add github:falling-ts/dsh-force-compact
# from a local checkout:
dsh plugin --profile web add ./dsh-force-compact

Or, from a local checkout, as a --patch overlay without installing:

dsh web --patch dsh-force-compact/cordis.patch.yml

The plugin is loaded iff ~/.dsh/logs/dsh-force-compact.log gains:

[force-compact] debug logging enabled — writing [force-compact] lines to <absolute path>

Command availability — /force-compact loads lazily

The commands service arrives with the agent-presets plane, after the plugin's boot-time apply, so registration happens at the first guarded-listener activation (agent/request / agent/pre-step / agent/status / session/flush), settling permanently on the first success. Practical effect: after (re)starting the instance, a fresh session's / picker does NOT show /force-compact until that session makes its first model request — send any one message, then the command is registered process-wide.

  • Success: [force-compact] /force-compact command registered (deferred)
  • commands permanently absent: one … still UNREGISTERED 10 min … warn explains the empty picker. Until registered, the rest of the plugin works — degradation, not an install failure.

Verify a compaction happened:

idle compaction (builtin) shadowed N nodes (~M tokens)
builtin compaction OK — replaced span seq[A..B] (N nodes, ~K tokens) with a P-char checkpoint
compaction thinking-policy — settings.disableThinking=true → extra.reasoningEffort='off' (…)
summarization wire-fields → <provider>/<model>: reasoningEffort='off' + reasoning_effort="none" (…)

(The last two lines are the per-attempt audit pair described under "Observability".)


Settings

Namespace falling-ts-force-compact (the profile entry's loader id); values are written to the profile's cordis.patch.yml (harness 0.1.7 onward; formerly $DSH_HOME/settings.yaml):

keytypedefaultmeaning
disableThinkingbooleantrueOnly the plugin's own summarization call carries reasoningEffort:'off'; everything else unchanged.
autoThresholdTokensnumber ≥ 3200032000Default projected-token trigger for the gate. Every session without its own override uses this value; Floor 32000 (clamps back up at read time). See Per-session threshold.
retainLatestTokenspositive int ≥ 80008000Retain the latest N tokens verbatim; older history is summarized in one batch. Floor 8000. Drives both the auto gate and /force-compact.
turnEndForceCompactionEnabledbooleantrueCompact on the agent's idle transition.
debugbooleantrueEmit [force-compact] diagnostics to the plugin log.
logFilestring~/.dsh/logs/dsh-force-compact.logDiagnostics destination (~ expands to home dir).
compactionMode'realm' | 'global''realm'Official-service resolution strategy (priority-1 path).
builtinEnabledbooleantrueGate for the builtin engine fallback.
maxSummaryTokensinteger (1024–200000)1024Cap on the summarizer LLM maxTokens.
summarizationTimeoutMsinteger 5000–2147483647 (ms)90000Hard wall-clock cap for ONE summarization stream (hung-stream guard). Floor 5000 (a sub-5s cap would false-abort slow local endpoints); ceiling 2147483647 because the value is scheduled through AbortSignal.timeout, which throws on a fractional delay and silently degrades a 2^31..2^32-1 delay to 1 ms. Out-of-range values are clamped and fractions truncated.

Example — an aggressive local profile:

falling-ts-force-compact:
  disableThinking: true
  autoThresholdTokens: 40000   # compact sooner ⇒ keep the live prompt small
  retainLatestTokens: 8000
  turnEndForceCompactionEnabled: true

Without the settings service the plugin falls back to the same defaults and still compacts — the namespace is optional, never a hard dependency.

Every token-count field above also accepts a K / M suffix in the settings form (32K, 1M, or a plain number); parsing is decimal (32K = 32000, 1M = 1000000) and the floors and ceilings above still clamp. The same parser backs the per-session control.

Per-session threshold

autoThresholdTokens is only the default. A single conversation can override it without touching the shared settings document:

  • The control is the chip sitting just to the right of the context-usage percentage in the composer bottom strip (the same row that carries the built-in stats pills). It shows a fixed caption — icon plus "Force-compact threshold" — so the row reads the same in every session and the number never shifts the layout. The effective value, the owning session and the overridden state live in its tooltip (and in the panel).
  • Click it to open a small panel: type a threshold and click Save, or click use global default to drop the override again. Enter submits.
  • The input takes a plain number or a K / M suffix (123K, 1M, 32000). Anything unparseable is refused with an inline message and nothing is written; clearing it falls back to the default. The 32000 floor applies to overrides too.
  • While threshold < contextWindow the plugin also draws WHERE compaction will fire: a red dot on the composer context ring, placed at the ring angle for threshold / contextWindow (clockwise from twelve oclock, the same direction the ring fills), plus a red vertical line on the expanded breakdown bar at the same ratio (exactly as tall as the bar). A threshold at or above the context window draws nothing — occupancy can never reach it, so there is no trigger point to mark.
  • Overrides live per session id under sessionThresholds in this namespace and are read only by the session they belong to. Every gate (auto compaction, /force-compact, region compaction, checkpoint, idle) resolves sessionThresholds[sessionId] ?? autoThresholdTokens.
  • Absence is meaningful: removing the key restores the default, so an override can never masquerade as a global change.
  • The chip renders in both performance-and-usage display modes (compact and detailed). The built-in stats pills hide themselves when they have nothing to report; the threshold chip stays put.

Tuning for low-context llama.cpp

Keep autoThresholdTokens comfortably below the served context: the live prompt stays small and latency flat, while the agent keeps deep memory through the compressed head. Pressure is measured in projected tokens (provider-anchored), so the threshold maps predictably onto the UI figure.


Behavior notes

  • Runtime dependency: the compaction service (preset plane agent-presets:compaction-basic), read live via ctx.get('compaction'); unreachable → the builtin engine takes over (or the request proceeds).
  • Optional dependencies: settings / tokenMeter / commands / llm / agents are read via ctx.get(...) with guards — a missing one degrades gracefully.
  • Per-request settings read: parameters are read every model request, so edits take effect on the next request without a restart.
  • Signals: agent/* Waterfalls forward the current turn's signal; the session/flush checkpoint and the agent/status idle listener each mint a fresh AbortController.
  • Persistence: durable output is the compaction/* bracket events + a surfaceOp:replace user/message checkpoint, replay-safe across builds.
  • Client half: web/client.js adds the settings section "Force Compact" (localized labels), live-editable without restart (uSES-safe mirror).
  • One intentional timer: the 3 s publishDone fallback (presentation-only, documented deviation). Otherwise the plugin is pure listeners + a process-local Map force flag.

Screenshots

Settings panel — Force Compact section, all knobs live-editable

Settings page — the Force Compact section; all nine fields above are editable live without a restart.

Conversation page — the working line replacing the official label's leading phrase

Conversation page — the LiveUI signal rewrites the leading phrase of each running turn's label (compressing / done / a rotating working line) while the harness clock keeps ticking after it; at conversation end the official label is restored. The screenshot predates the 2026-09 colour removal: the badge is now plain grey text in the official font.


License

MIT (see LICENSE).