Back to home@stadeummwt

dsh-supreme

No description

Stars
0
Language
TypeScript
Created
Sep 6, 2026
Updated
Sep 9, 2026
GitHub repo

Introduction

🛡️ DSH SUPREME

The governance layer for DeepSeek Harness

Seven policy plugins · one bundle install · zero upstream patches · every claim executable

"ECC gives your harness breadth. Supreme gives it a conscience."

CI suite E2E upstream leaks schemas bundle license node

Install · dsh plugin --profile <your-profile> add github:stadeummwt/dsh-supreme

Quick install · Why Supreme · The seven plugins · Proof wall · v1.3 features · v1.2 features · Docs


🤔 Why Supreme

DSH's plugin ecosystem (3,421 catalog entries reviewed, 2026-09) is rich in single-domain tools — a router here, a memory store there, a verifier somewhere else. Each solves one slice of governance and asks you to trust its output.

Supreme is the opposite design. It is a full governance stack — cost policy, observability, routing, verification, memory policy, workflow limits, security audit — that treats proof as a product feature: every claim in this README maps to a command you can run, and every hard rule (deny paths, cost gates, secret scrubbing) is deterministic code, not model judgment.

The usual DSH pluginDSH Supreme
Solves one domainSeven governance domains, one install
"Trust the output"Verdict gatesCOMPLETE only when every check passes
Config verified by vibesConfig-key hygiene scan against real zod schemas (silent-strip trap closed)
Markdown evidencePublished JSON Schemas + append-only JSONL evidence stores
Touches core or monkey-patchesZero upstream patches — pinned upstream, worktree clean, verified every run
Security as a README paragraphSix-surface security audit (prompts · hooks · MCP · permissions · secrets · agent files) in CI
No ML dependencyAlso no ML — deterministic counting, globs and comparisons only. Speed is a feature: router decision ≈ 0.02–0.03 ms / 1k iterations

The core rule of this repo: bukti sebenar > klaim — real evidence over claims. If a statement here can't be re-run by you, it's marked as a claim, not a fact.


⚡ 60-second install

The repository is a dsh bundle — no build step needed (dist/ is committed):

dsh plugin --profile <your-profile> add github:stadeummwt/dsh-supreme

That mounts all seven plugins with safe production defaults:

PAID / TRIAL routes  → DENIED (hard rule, LAB-only override)
UNKNOWN cost class   → DENIED
commands / network   → OFF by default
router candidates    → you add yours in your own patch layer (last write wins)

Pick a composition in one more line if you don't need all seven:

FragmentActive pluginsUse it for
corepolicygovernance floor on any profile
standardpolicy · observability · memory · verifierdaily-driver
supremeall sevenfull stack
laball seven + LAB overridesexperiments only — never production
dsh --profile <your-profile> \
  --patch "$DSH_HOME/profiles/<your-profile>/node_modules/dsh-supreme/config/compositions/standard.patch.yml"

🧩 The seven governance plugins

Exactly seven. The scope is frozen (AGENTS.md) — no scope creep without a proven blocker.

#PluginServiceWhat it enforces
1supreme-policysupremePolicyCost-class / risk / delegation admission. UNKNOWN cost ⇒ DENY. Paid & trial overrides are LAB-only. Unicode-taint + encoding-blob detection & denial. CoT presence gate with visibility profiles + risk gating. Deny-circumvention (deny_retry) guard. Capability-class gate.
2supreme-observabilitysupremeObservabilityAppend-only JSONL metadata log over official DSH event seams. Allowlisted fields, secret-sentinel scrub, fail-open.
3supreme-benchmarksupremeBenchmarkReproducible task/run/score JSONL evidence; per-model aggregation that feeds the router; commitHash + irVersion provenance binding; evidenceBacked anti-sandbagging flag.
4supreme-routersupremeRouterDeterministic selection: 8 hard gates → weighted scoring → RM0-first cost-class rule → unscoredEvidenceWeight anti-sandbagging downweight → optional verifier-failure-driven effort pacing. Carries CapabilitySignal labels onto decisions.
5supreme-verifiersupremeVerifierDeterministic validator registry (exact-text · regex · JSON · file · command). Evidence > model self-confidence.
6supreme-memory-policysupremeMemoryPolicyMemory selection policy: confidence floor, injection cap, relevance ranking, bounded append-only note ledger (credential-bearing notes rejected at admission).
7supreme-workflow-policysupremeWorkflowPolicyWhen/how ctx.subagents / ctx.workflowEngine may run: limits, degradation ladder, glob path scoping (blocked beats allowed), verifier-gated close for HIGH-risk tasks, A2A contact graph + overreach audit.

Four support plugins (supreme-minimal-probe, supreme-boot-probe, supreme-gate-driver, supreme-fake-llm) exist only as test fixtures for the suite — they never ship in the bundle.


🏁 Proof wall — every verdict runnable

Don't trust this README. Run these:

CommandVerdict markerWhat it proves
bun run suiteCOMPLETE101/101 Level-A checks + 5/5 real-loader boots + v1.2/v1.3/v1.3.1 audit gates
bun run v131:verifyV131_COST_FIX_VERIFIED · V131_VERIFIER_FIX_VERIFIED · V131_MEMORY_FIX_VERIFIED · V131_A2A_FIX_VERIFIED · V131_EVIDENCE_BINDING_VERIFIED · V131_OUTCOME_ROUTING_VERIFIED · V131_FAILURE_INJECTION_VERIFIEDReview-hardening end-to-end: cost pre-dispatch deny, symlink-proof roots, JSON-schema strictness, memory isolation, A2A registry, evidence staleness, outcome routing, failure injection (465 probes — see docs/REVIEW-FIXES-v1.3.1.md)
bun run v13:verifyV13_POLICY_E2E_COMPLETE · V13_WORKFLOW_E2E_COMPLETE · V13_ROUTING_E2E_COMPLETEAll 7 v1.3 ASTRA features end-to-end: real engines + real pinned-cordis adapters (85 + 82 + 17 probes)
bun run bundle:verifyBUNDLE_E2E_COMPLETEReal dsh plugin add → reconciler → boot → 13 services → user-patch override wins → clean dispose
bun run composition:verifyCOMPOSITIONS_E2E_COMPLETEAll 4 fragments: service presence and absence, relative dataDir write-through
bun run v12:verifyV12_E2E_COMPLETEEvery v1.2 config key arrives at its service + functional probes (taint deny, effort escalation/recover, path scope, close gate, ledger)
bun run v3:verifyV3_CONFIG_REVIEW_EVIDENCEThe silent-strip trap, live: a wrong config loses 5/6 keys → corrected config enforces 6/6
Level A unit checks      101/101 PASS  (policy 17 · observability 7 · benchmark 11 · router 23
                                       verifier 11 · memory 13 · workflow 19)
v1.3.1 review probes     465/465      (cost 37 · verifier 43 · memory 88 · a2a 63 ·
                                       evidence 82 · outcome-routing 80 · failure-injection 72)
v1.3 E2E probes          184/184      (policy 85 · workflow 82 · routing 17 — real engines,
                                       real pinned-cordis adapters, no upstream build needed)
Real-loader boots        5/5 PASS     (supreme-minimal, core, standard, supreme, lab)
  boot times             supreme-minimal ~51–60 ms · core/standard/supreme/lab ~830–990 ms
Keyless scenario         9/9 gates PASS (real DSH session; router picks free route; PAID rejected)
Security                 sentinel leaks = 0 · paid automatic fallback = DISABLED
v1.2/v1.3 audit gates    config-key hygiene PASS · pinned-ref scan PASS ·
                         six-surface audit PASS (incl. the pinned `workflow/agent-start` seam) ·
                         schema contract PASS (3 schemas)
Upstream integrity       commit unchanged · worktree clean · patches = 0
Performance              router ≈ 0.02–0.04 ms / 1k · observability serialize ≈ 0.001–0.007 ms / 1k
VERDICT                  COMPLETE

Honesty rule: the real-loader path (real/boot.mjs) is the only real-integration evidence. The Level-A harness under src/harness/cordis-mini is a lifecycle fixture — it is never cited as DSH proof.


🔐 Security guarantees

GuaranteeMechanismProof
Secrets never leak through observabilitySecret-sentinel scrub on allowlisted fields, fail-open write pathsuite: sentinelLeaks = 0 every run
Tainted tool arguments can't dispatchUnicode class scan (zero-width / bidi / BOM / tag) + taintPolicy: DENY via upstream tools/pre-executeV12_E2E_COMPLETE functional probe
Denying a command actually stops itDeny-circumvention guard: same-shape retry of a denied call refused (deny_retry) — signature carries names/types, never valuesV13_POLICY_E2E_COMPLETE probes
Hidden payloads can't ride in tool argsEncoding-blob scan (≥256-char base64/hex runs), class names + lengths onlyV13_POLICY_E2E_COMPLETE probes
Self-declared capability labels can't buy permissioncapabilityClassGate ENFORCE/AUDIT; LAB allowlist is floor-bound; labeling only RESTRICTSV13_POLICY_E2E_COMPLETE probes
Inter-agent channels stay on the declared graphallowedContacts directed edges; out-of-graph audited (a2a_contact), DENY blocks pre-factV13_WORKFLOW_E2E_COMPLETE probes
Delegations can't quietly exceed their taskOverreach audit: risk ceiling + approval gate + path scope, value-freeV13_WORKFLOW_E2E_COMPLETE probes
Benchmark scores can't sandbag the routerevidenceBacked flag (verifier-PASS rule) + fixed unscoredEvidenceWeight downweightV13_ROUTING_E2E_COMPLETE probes
Values never echoed in audit eventsTaint/contact/overreach events carry class names, ids and levels onlycode + suite checks
Paid models never fire by accidentUNKNOWN cost ⇒ DENY; allowPaid refused outside LAB; no automatic fallbackkeyless scenario gate 9/9
Destructive delegation is scopedblockedPaths > allowedPaths glob enforcement; DENY_ALL secret policysuite checks 15 (workflow)
HIGH-risk work can't skip verificationrequireVerifierPassOnClose evidence gateV12_E2E_COMPLETE probe
Supply chain stays pinnedExternal refs scanned; upstream commit + irVersion bound into run recordspinned-ref scan PASS
Your own audit, offlineSix-surface audit: prompts · hooks · MCP · permissions · secrets · agent filessuite check PASS

🛡️ v1.3.1 review-hardening

Response to an external v1.3.0 review: 5 findings reproduced → fixed → proven (each with a failing test on the original code), plus outcome-based routing, evidence-bound verification, fast path/recovery, and a failure-injection harness. Full per-issue evidence — reproduction commands, root causes, before/ after outputs, remaining limits, and rollback steps (bash + PowerShell) — lives in docs/REVIEW-FIXES-v1.3.1.md.

#SeverityFinding (v1.3.0)Fix (v1.3.1)
AP1Paid/unknown-model LLM requests dispatched with no cost checkPre-dispatch deny at agent/request + llm/stream backstop; zero adapter calls on deny; UNKNOWN denied in production RM0; LAB exception contract kept
BP1Symlink inside allowedRoots escaped the file-hash verifierNative realpath validation of roots AND targets before any read; traversal/sibling-prefix/missing-file rejected; race reduced (not race-proof — documented)
CP1latestSelection shared across sessions/tasks (memory contamination)Selections bound to (session, task); unknown identity → empty; bounded LRU + cleanup on end/cancel/dispose
DP2copy_file {target: b.txt} misclassified as agent-to-agent contactTrusted comms-tool registry gates recipient extraction; post-fact emit stays detect-only
EP2JSON Schema additionalProperties:false silently ignored (false PASS)Deterministic validator: unsupported keywords → ERROR/UNAVAILABLE, never silent downgrade

Run it: bun run v131:verify (7 markers, 465 probes) — then read the doc before trusting this table.

🧬 v1.3 ASTRA-hardening features

Seven deterministic hardening features from the ASTRA-1 backlog (research/gpt6-astra-2026-09.md §7). No ML, no new deps — every feature is engine-checked in the keyless suite and proven end-to-end by bun run v13:verify (real engines + real pinned-cordis adapters). The shared label contract CapabilitySignal { capabilityClass?, cotVisibility? } is exported by supreme-policy and carried (never enforced) by the router.

supreme-policy — four features

Config keyDefaultMeaning
cotVisibilityProfiles{}routeId → expected CoT visibility. A route declared none never denies on cot_missing — ENFORCE downgrades to audit-only (empty-CoT models can't be coerced into producing a trace). Resolution: explicit signal > profile > verbose.
riskGatedCoTfalseENFORCE applies only to HIGH-risk tools (deterministic command/network/write name classifier); non-HIGH tools keep AUDIT.
denyCircumventionGuardtrueA same-shape retry of an already-denied call is refused with reason code deny_retry. The signature encodes argument NAMES + TYPES only — values can never enter it. First calls unaffected; resetDenyCircumvention(sessionId) is the escape hatch.
enableEncodingScanfalseAudit/deny ≥256-char base64/hex runs in tool arguments (encoding_blob; argument NAME + run LENGTH only). Extends the v1.2 taint surface: same event, same taintPolicy.
capabilityClassGate'OFF'Gate requests carrying capabilityClass: AUDIT records, ENFORCE denies unsanctioned classes. Unlabeled requests always pass untouched.
sanctionedCapabilityClasses / labCapabilityClassAllowlist[] / []Sanction lists; the LAB allowlist is additive and binds ONLY on the LAB floor. No implicit ROUTINE exemption — a self-declared label can only RESTRICT, never grant.

supreme-workflow-policy — two features

Config keyDefaultMeaning
agentContactPolicy / allowedContacts'LOG_ONLY' / []A2A contact graph: directed { from, to } edges of agent ids/roles (empty = inert). Out-of-graph spawn/message contacts are audited as a2a_contact; under 'DENY' the pre-fact tools/pre-execute waterfall refuses with a2a_contact_denied. Emit-mode seams are DETECT-only.
maxRiskLevel / approvalRequiredFor'HIGH' / []Overreach audit: delegations above the risk ceiling, listed task classes without an approval flag, or paths outside the v1.2 scope are audited as overreach_suspected (labels, levels, flags, config globs — never content).

supreme-router + supreme-benchmark — anti-sandbagging

Config keyPluginDefaultMeaning
requireEvidenceForScoresbenchmarkfalseScore claims without verifier-PASS evidence are flagged evidenceBacked: false on the score + run (flag only — scores never rewritten; re-evaluated when verification lands late).
unscoredEvidenceWeightrouter1FIXED multiplicative downweight for unevidenced benchmark claims (e.g. 0.5 halves such scores); ids + factors recorded on the decision + unscored_evidence events (ids only). 1 = off, back-compat.
routerCarries capabilityClass / cotVisibility labels from candidates onto the selected RouteDecision (carrier, not enforcer).

Composition fragments (v1.3 posture)

Fragmentv1.3 keys
coredenyCircumventionGuard: true pinned (the one default-ON); everything else inherits OFF defaults
standardenableEncodingScan: true + capabilityClassGate: AUDIT — audit-only, cannot block
supremesame audit-only policy posture + requireEvidenceForScores: true + workflow keys pinned at behavior-preserving defaults
labenforcing demo: capabilityClassGate: ENFORCE + labCapabilityClassAllowlist, cotVisibilityProfiles + riskGatedCoT, declared contact graph + maxRiskLevel: MEDIUM, unscoredEvidenceWeight: 0.5

🧬 v1.2 governance features

Deterministic. No ML. No new runtime deps. Every feature binds to a real pinned upstream seam and ships with engine checks + boot-level proof (bun run v12:verify).

supreme-policy — unicode taint denial + CoT presence gate

Upstream freezes tool arguments after logging (wrappers may change only exec.signal), so the enforceable host-side posture is detect → audit → deny through the official tools/pre-execute seam ({ kind: 'deny', reason } — upstream materializes the error result; Supreme never fabricates tool output):

Config keyDefaultMeaning
enableUnicodeSanitizationtruescan tool arguments for zero-width / bidi-isolate / bidi-override / tag codepoints (U+200B–200F, U+2060–206F, U+202A–202E, U+FEFF, U+E0000–E007F)
logTaintAttemptstruerecord taint_detected events — class names only, values are NEVER echoed
taintPolicyLOG_ONLYDENY refuses the call before dispatch
reasoningTracePolicyOFFAUDIT records cot_missing when an assistant message carried no reasoning trace; ENFORCE additionally denies that session's tool calls (ENFORCE refused on the CORE floor)

supreme-router — RM0-first + effort pacing

Config keyDefaultMeaning
costFirsttruescore only the cheapest eligible cost class — FREE_CONFIRMED beats a rate-limited peer with better history; hard-gate evidence for ALL candidates preserved
effortPacing.enabledfalsedeterministic costClass → reasoningEffort mapping over the pinned agent/request seam (pinned DeepSeek levels: off / low / high / max)
effortPacing.escalateOnVerifierFailtrueone-step escalation (low → high) driven only by verifier FAIL evidence via reportVerifierOutcome() — never model self-confidence; PASS recovers

supreme-workflow-policy — surgical path scope + verifier-gated close

Config keyDefaultMeaning
allowedPaths / blockedPaths[] / []zero-dependency glob scope for delegations (** crosses segments, */? stay in-segment); blocked always wins; empty allowlist = unrestricted
requireVerifierPassOnClosefalseHIGH-risk tasks may only close with recorded verifier PASS evidence

supreme-memory-policy — bounded ledger + instinct-style gates

Config keyDefaultMeaning
ledgerEnabledfalseopt-in bounded, append-only JSONL note ledger (ledgerDir, ledgerFileName, ledgerMaxEntries) — credential-bearing notes rejected at admission
minConfidence0.7notes below this confidence never inject (ECC instincts analogue — recorded evidence quality, not self-assessment)
maxInjected6hard cap per selection
relevanceRankingtruedeterministic task-token-overlap ranking before priority (counting, not ANN)

supreme-benchmark — provenance binding

Run records accept commitHash (40-hex sha or UNAVAILABLE) and irVersion — malformed values are rejected by validation, so routing evidence stays bound to the code that produced it.

Published schemas + audit suite


📦 Install as a dsh bundle

The repository IS the bundle: package.json declares dsh.bundle.patchcordis.patch.yml, which inserts the seven frozen plugins as profile rows. Any profile can adopt Supreme through the official plugin flow:

# from a local checkout…
dsh plugin --profile <your-profile> add /path/to/dsh-supreme
# …or straight from GitHub
dsh plugin --profile <your-profile> add github:stadeummwt/dsh-supreme

# prove an install end-to-end (real CLI install + boot + layering checks)
bun run bundle:verify
# prove the v1.2 config surface end-to-end
bun run v12:verify
# prove the v1.3 ASTRA-hardening features end-to-end (all three verifiers)
bun run v13:verify

The bundle mounts the seven plugins with safe production defaults (PAID / TRIAL denied, commands/network off, zero router candidates). Extend candidates, project knowledge, and workflow limits from YOUR profile patch layer — the composer applies last write wins per row id, so user config always beats bundle defaults. The four support/fixture plugins are NOT part of the bundle: they never ship into user profiles.

Composition fragments ship under config/compositions/ — the bundle-world analogue of manifest-driven install profiles. Each fragment UPDATE-patches the bundle rows by id (whole-config replacement, disabled: true for rows outside the composition) and carries no name restatement, so it stays install-location-independent. Prove all four end-to-end: bun run composition:verifyCOMPOSITIONS_E2E_COMPLETE.

Notes: dsh plugin add requires pnpm on PATH; installing from GitHub works without a prepare build because dist/ is committed. Fragment paths are relative to the dsh process working directory — override any row from your own patch layer.


🏗️ Architecture

┌────────────────────────────────────────────────────────────────────┐
│  Next.js dashboard (project app) — PROJECTION only, owns no state  │
│  GET/POST /api/supreme/*  (dev/LAB only)                           │
└──────────────────────────────┬─────────────────────────────────────┘
                               │ reads suite reports / triggers runs
┌──────────────────────────────▼─────────────────────────────────────┐
│  SUPREME PLUGIN LAYER (dsh-supreme/dist/plugins, project-owned)    │
│  policy · observability · benchmark · router · verifier ·          │
│  memory-policy · workflow-policy  (+ 4 support/fixture plugins)    │
│  Cordis conventions: name/inject/Config(Standard Schema)/apply     │
└──────────────────────────────┬─────────────────────────────────────┘
                               │ inject: official DSH service names
┌──────────────────────────────▼─────────────────────────────────────┐
│  DSH CORE (pinned upstream — never modified)                       │
│  ctx.llm · ctx.sessions · ctx.systemPrompt · ctx.tokenMeter ·      │
│  ctx.credentials · ctx.subagents · ctx.workflowEngine              │
│  Events: session/* · agent/request* · tools/execute ·              │
│          subagent/* · workflow/*                                   │
└────────────────────────────────────────────────────────────────────┘

Dependency direction (acyclic, enforced):

DSH core services    →  Supreme plugins        (injected seams)
supremePolicy        →  verifier, router, workflow-policy
supremeObservability →  router, verifier (optional), workflow-policy
supremeBenchmark     →  router                 (router reads history; benchmark NEVER depends on router)
supremeVerifier      →  workflow-policy        (verification evidence consulted)

Pinned upstream

ItemValue
Repositoryhttps://github.com/deepseek-ai/deepseek-harness
Pinned commitd347e703908d0406b7a7ef80e3a0e594d86b2215 (master, tag dsh-v0.1.3-alpha.1)
DSH version0.1.3-alpha.1
Vendored Cordis4.0.2 (vendor/cordis)
Upstream worktreekept pristineUPSTREAM_CORE_MODIFIED = NO, patch count 0
ToolchainNode v24 (v24.19.0), pnpm 11.7.0, Bun 1.3.14 (bundler)

The pinned upstream checkout is read-only for this project. It is resolved at runtime: DSH_UPSTREAM_ROOT env override → sibling ../deepseek-harness → in-project node_modules/.upstream/deepseek-harness. Prefer the sibling location: some upstream builds (pnpm + declaration emit) reject checkouts nested under a node_modules directory. All Supreme code lives in project-owned paths.

Compositions (profiles)

ProfileBundleMounted Supreme plugins
supreme-minimalnone (bare Loader)minimal probe only
core@deepseek-ai/dsh-baseminimal probe, boot probe, supreme-policy (CORE)
standard@deepseek-ai/dsh-base+ observability, memory-policy, verifier
supreme@deepseek-ai/dsh-baseall 7 + fake-llm + gate-driver (SUPREME policy)
lab@deepseek-ai/dsh-baseall 7 + fake-llm + gate-driver, LAB-only overrides (allowPaid: true, allowCommands: true, maxConcurrentAgents: 4)

Full layer map + verified real-API evidence table: docs/architecture/ARCHITECTURE.md.


🚀 Build & verify from source

Prerequisites: Node ≥ 24, pnpm 11.7.0 (upstream build), Bun ≥ 1.3. Commands assume the repo root (dsh-supreme/ as published; inside the companion Next.js workspace the suite auto-detects both layouts).

# 1. Install dependencies
bun install

# 2. Clone the pinned DSH upstream (default lookup: sibling ../deepseek-harness;
#    any location works via DSH_UPSTREAM_ROOT — avoid nesting it under node_modules)
git clone https://github.com/deepseek-ai/deepseek-harness.git ../deepseek-harness
git -C ../deepseek-harness checkout d347e703908d0406b7a7ef80e3a0e594d86b2215

# 3. Build the pinned upstream libraries — official tsconfig graph, memory-batched
#    (one tsc -b over the 217-ref host graph needs ~4 GB headroom; the batched
#    runner keeps each invocation under 2 GB)
npm run build:upstream

# 4. Bundle every Supreme plugin to dist/ (one ESM file per plugin; zod external)
PLUGINS="supreme-policy supreme-observability supreme-benchmark supreme-router \
supreme-verifier supreme-memory-policy supreme-workflow-policy \
supreme-minimal-probe supreme-boot-probe supreme-gate-driver supreme-fake-llm"
for p in $PLUGINS; do
  bun build src/plugins/$p/index.ts \
    --outfile dist/plugins/$p/index.mjs \
    --format esm --target node --external zod
done

Each dist bundle externalizes only zod and Node builtins; @deepseek-ai/cordis appears solely as erased type imports. This exact command was verified to reproduce the committed dist/plugins/supreme-policy/index.mjs byte-for-byte.

Real boot (the only real-integration evidence)

# Boot any composition through the REAL pinned DSH Loader and dispose cleanly.
# --setup installs the profile under $DSH_HOME/profiles/<name>/ from config/.
node real/boot.mjs --profile supreme-minimal --setup
node real/boot.mjs --profile core         --setup
node real/boot.mjs --profile standard     --setup
node real/boot.mjs --profile supreme      --setup
node real/boot.mjs --profile lab          --setup

Each run prints one JSON result (bootMs, disposeMs, services presence map, gate results) and exits non-zero on any failure. Gate markers are appended under data/real/ — see the runbooks for expected markers per profile.

Suite execution

bun run suite            # full suite incl. 5 real boots (needs the built upstream)
bun run suite:json       # machine-readable SuiteReport
bun run suite:keyless    # Level A only — runs without the upstream; verdict stays
                         # PARTIAL (REAL_BOOT_SKIPPED, UPSTREAM_CHECKOUT_UNAVAILABLE)
bun run suite:keyless:ci # keyless with CI-friendly exit code: 0 iff verdict is PARTIAL
                         # with only the documented keyless blockers — any real
                         # failure (UNIT/leaks/hygiene/audit/schema) still fails

The suite exits 0 only when every mandatory gate passes (verdict: COMPLETE). Any failure prints the exact blocking gates.

HTTP API (dashboard projection — dev/LAB only)

The Next.js app exposes a thin, read-mostly projection over the suite. It owns no runtime state; runs live in an in-memory store (latest 20 runs) and suite execution is disabled in production (NODE_ENV=production returns 403 unless SUPREME_ENABLE_SUITE=1).

EndpointMethodBehavior
/api/supreme/statusGETSuite scope (frozen 7 plugins, compositions), upstream commit/cleanliness, DSH/cordis versions, runtime info. Always safe.
/api/supreme/reportGETLast SuiteReport from memory; 404 if no run yet (POST /api/supreme/suite/run first).
/api/supreme/suite/runPOSTExecutes the full suite (including 5 real boots). dev/LAB only403 in production without SUPREME_ENABLE_SUITE=1.
/api/supreme/suite/runs/:idGETOne run record (runId, startedAt, durationMs, full report); 404 for unknown ids.

Implementation: src/app/api/supreme/** + src/lib/supreme-suite.ts (project app, outside dsh-supreme/).


📤 Distribution (manual, owner-driven)

Repo policy: no pull requests are opened on third-party repositories on the owner's behalf. Prepared submission artifacts live in distribution/:

  • awesome-dsh-entry.yml — catalog-ready entry (single file, category security, validator-conformant keys only).
  • SUBMISSION-GUIDE.md — how listing on dsh-market actually works (it auto-feeds from the awesome-dsh-plugin catalog), the pre-flight gate checklist, the exact manual submission commands, and the npm-publish note.

The GitHub repo already carries the dsh-plugin topic and a dsh.bundle manifest, so the only remaining step for listing is the manual one-file PR the owner chooses to make.


📁 Directory layout

dsh-supreme/                      (repo root as published)
├── README.md                  ← this file
├── VISION.md                  ← original v1 project vision (frozen architecture contract)
├── CHANGELOG.md
├── AGENTS.md                  ← engineering rules for future agents
├── SOURCE-OF-TRUTH.md         ← upstream integrity record (historical + current)
├── LICENSE                    ← MIT (v1.2)
├── package.json               # suite/boot/build/verify scripts (zod + yaml deps)
├── schemas/                   # published JSON Schemas (suite report, benchmark, ledger)
├── .github/workflows/ci.yml   # keyless + full suite on push/PR (v1.2)
├── config/
│   ├── supreme-minimal.cordis.yml   # bare-Loader probe gate
│   ├── core.cordis.yml              # CORE composition
│   ├── standard.cordis.yml          # STANDARD composition
│   ├── supreme.cordis.yml           # SUPREME composition (all 7)
│   ├── lab.cordis.yml               # LAB composition (LAB-only overrides)
│   ├── examples/                    # corrected config example (provenance noted)
│   └── compositions/                # 4 overlay fragments (core/standard/supreme/lab)
├── distribution/                    # manual submission artifacts (no auto-PRs)
│   ├── awesome-dsh-entry.yml        # catalog entry draft (one file)
│   └── SUBMISSION-GUIDE.md          # owner-driven listing walkthrough
├── real/
│   ├── boot.mjs               # REAL DSH boot harness (Loader + root-fiber dispose)
│   ├── bundle-verify.mjs      # E2E: real CLI install + layering (BUNDLE_E2E_COMPLETE)
│   ├── composition-verify.mjs # E2E: 4 composition fragments (COMPOSITIONS_E2E_COMPLETE)
│   ├── v3-config-verify.mjs   # E2E: silent-strip proof (V3_CONFIG_REVIEW_EVIDENCE)
│   ├── v12-config-verify.mjs  # E2E: v1.2 config surface + probes (V12_E2E_COMPLETE)
│   ├── v13-policy-verify.mjs  # E2E: v1.3 policy features, 85 probes (V13_POLICY_E2E_COMPLETE)
│   ├── v13-workflow-verify.mjs# E2E: v1.3 A2A + overreach, 82 probes (V13_WORKFLOW_E2E_COMPLETE)
│   ├── v13-routing-verify.mjs # E2E: v1.3 labels + anti-sandbagging, 17 probes (V13_ROUTING_E2E_COMPLETE)
│   └── build-batched.sh       # memory-batched official upstream build
├── dist/plugins/<name>/index.mjs    # bun-built ESM bundles loaded by the real Loader
├── data/
│   ├── observability/observability.jsonl   # runtime metadata log
│   ├── benchmark/benchmark.jsonl           # routing evidence store
│   └── real/*.markers.jsonl                # boot/gate markers (verified evidence)
├── src/
│   ├── plugins/               # engine.ts (pure logic) + index.ts (Cordis adapter) per plugin
│   │   ├── supreme-policy/  supreme-observability/  supreme-benchmark/
│   │   ├── supreme-router/  supreme-verifier/  supreme-memory-policy/
│   │   ├── supreme-workflow-policy/
│   │   └── supreme-minimal-probe/  supreme-boot-probe/  supreme-gate-driver/  supreme-fake-llm/
│   ├── suite/                 # runner.ts + cli.ts + engine-checks.ts + config-hygiene.ts
│   │                          # + surface-audit.ts + schema-contract.ts + harness.ts
│   └── harness/cordis-mini/   # Level-A lifecycle FIXTURE only (never cited as DSH proof)
└── docs/
    ├── architecture/ARCHITECTURE.md
    ├── decisions/ADR-0000 … ADR-0007
    └── runbooks/              # install, build, test, boot-*, upgrade-pinned-dsh, rollback

📖 Documentation map

DocContents
AGENTS.mdFrozen scope, upstream rules, Cordis conventions, ownership table, verification requirements
docs/architecture/ARCHITECTURE.mdLayer map, verified real-API evidence table, event seams, composition layering
docs/decisions/ADR-0000 (fixture history) + ADR-0001…0007 (one per major decision)
docs/runbooks/install, build, test, boot-core/standard/supreme/lab, upgrade-pinned-dsh, rollback
research/ECC dissection (253,948★) + v3 plan review — the data behind the v1.2 roadmap
SOURCE-OF-TRUTH.mdUpstream integrity record (historical + current)
Per-plugin READMEssrc/plugins/<name>/README.md — purpose, config tables, contracts, security boundaries
CHANGELOG.mdVersion history with evidence markers per release

❓ FAQ

Q: Does Supreme modify DeepSeek Harness? No. The pinned upstream worktree stays pristine — UPSTREAM_CORE_MODIFIED = NO, patch count 0, re-verified on every suite run. Supreme is an ordinary Cordis plugin layer that consumes official services and event seams.

Q: Is any of this AI-powered? None. Every gate is deterministic code — counting, glob matching, string comparison, zod validation. That's why the router decides in ~0.02–0.03 ms and why results are reproducible on your machine, today.

Q: Why does UNKNOWN cost deny the model? Because an unclassified route is an unaudited spend path. supreme-policy treats it as a hard DENY; paid/trial classes require an explicit LAB-only override. RM0-first routing then prefers FREE_CONFIRMED candidates deterministically.

Q: Can I use just the policy plugin? Yes — that's the core fragment. Or standard for the daily-driver four. Fragments are one-line overlays on your own profile.

Q: What if my config has a typo or an unknown key? The v1.2 suite runs a config-key hygiene scan: every shipped YAML row is validated against the plugin's real zod schema, so the "boot passes but your governance keys were silently stripped" trap (proven live in V3_CONFIG_REVIEW_EVIDENCE) stays closed.

Q: Does it work offline / air-gapped? The six-surface audit, taint scanning, ledger and all suite checks are fully offline and deterministic. Real boots need the pinned upstream checked out locally — no network calls at runtime.

Q: Why isn't Supreme listed in the dsh-market yet? Listing requires a one-file PR to the catalog, and this repo's policy is that such PRs are made by the owner, manually (see distribution/SUBMISSION-GUIDE.md). Everything else is already prepared.


📜 Honest limitations

  • The real-loader path via real/boot.mjs is the only real-integration evidence; the Level-A lifecycle harness (src/harness/cordis-mini) is a fixture and is never cited as DSH proof.
  • Keyless suite verdict is honestly PARTIAL (REAL_BOOT_SKIPPED) without a built pinned upstream — it does not fake completeness.
  • Deferred items (documented, not forgotten): HNSW-style memory indexing and Archify-style schema migration stay out of scope for the frozen seven.
  • Router candidates ship empty (zero-by-default): you add models from your own patch layer. Supreme governs choices; it does not preselect providers.

Built proof-first. Bukti sebenar > klaim.

If Supreme hardened your harness, consider starring the repo — it helps other DSH users find governance tooling.

⬆ back to top