Back to home@CatheadOwl

dsh-eval

Agent eval framework over dsh headless runs: case runner, session-trace assertions, and a scripted mock-LLM layer for plugin intent tests.

Stars
0
Language
JavaScript
Created
Sep 4, 2026
Updated
Sep 4, 2026

Introduction

@catheadowl/dsh-eval

English | 中文

A dsh-native agent evaluation layer for plugin authors: behavior cases run against real headless dsh traces, while review experiments test whether fresh models understand plugin outputs.

It evaluates the assembled agent harness (the graph a plugin + profile + patch + tool registry form inside a real dsh headless run), not isolated functions; verdicts come from dsh-native session-trace projections and matchers (contract assertions), not metric scores. It is not a general agent-eval platform (no dashboard / dataset hosting / metric catalog, no benchmark ranking) and not a DeepEval / OpenAI Evals replacement — those projects proved the problem space; this package picks the dsh-native vertical solution.

Documentation is Chinese-first; deep contracts live in docs/ (matchers / boundary contracts / review / report structure / host wiring / known issues).

Why it exists

LayerQuestionVerdictExecution
unit / shape testare deterministic fields and values correctautomaticthe plugin's own node:test
behavior realdoes natural-language intent pick the right tooltrace matcherdsh + real model
behavior mockis the tool pipeline and write round-trip stabletrace matcher + workspace inspectdsh + scripted mock LLM
comprehension reviewcan a fresh model understand the output and the next stepmanual rubric, converged over runsabstract review experiment + replaceable executor

A dsh plugin is correct when the assembled graph really wires tools, steers, prompts, and gates together — plugin unit tests cover only part of that, and "is the output understandable" is not a string regression at all. This package turns both layers into replayable evidence instead of manual trial runs.

plugin-owned experiment             shared framework
fixtures + prompt + rubric + observe ──► experiment/review.mjs
                                               │ task
                                               ▼
                                        adapters/dsh/review.mjs ──► dsh headless

behavior *.eval.mjs ───────────────────► dsh behavior runner (trace + mock)
  • src/experiment/ is the model- and runtime-agnostic experiment layer: blind review, live observation, byte-identical evidence across reviewers. It does not import dsh.
  • src/adapters/dsh/ is the landing layer: hands the abstract task to an isolated dsh headless run.
  • Your eval/ keeps only domain fixtures, projections/observe, prompts, rubrics, and cases — no runner duplication.

Install

npm i -D @catheadowl/dsh-eval

Requirements (wiring details and failure self-diagnostics in docs/host-wiring.md):

  • a built deepseek-harness checkout (apps/cli/lib/bin.js);
  • the plugin under test installed into a dsh profile;
  • the peer dependency @deepseek-ai/dsh-llm must be wired manually (npm auto-installs an incompatible antique version; replace it with a link pointing at the host checkout).

Quickstart

<plugin>/eval/behavior/mock/smoke.eval.mjs:

import { firstTool, toolCalled, toolCallStep, textStep } from '@catheadowl/dsh-eval'

export default {
  id: 'my-first-case',
  mode: 'mock',
  task: 'rename guide.md to intro.md',
  async prepare(workspace) { /* seed fixture files */ },
  script: { steps: [toolCallStep('md_rename', { oldPath: 'guide.md', newPath: 'intro.md' }), textStep('done')] },
  expect: [toolCalled('md_rename')],
}
dsh-eval run --mode mock eval/behavior/mock
dsh-review --dry-run eval/comprehension     # model-free dry run of the review layer

The command needs to know which dsh profile to use: pass --profile <name> explicitly, or drop a dsh-eval.config.mjs at the package root (see "Unified config" below).

Real runs use dsh-eval run --profile <p> --repo <harness checkout> <case path>; all flags (--mode/--keep-artifacts/--fail-on-skip/--format/--report) are documented in docs/report.md. Real cases auto-skip without credentials (dsh resolves credentials itself); mock and dry runs need no credentials.

Canonical layout

<plugin>/eval/
  .gitignore                 # .runs/ (no path prefix)
  README.md
  behavior/                  # optional
    real/*.eval.mjs
    mock/*.eval.mjs
    _fixtures/
  comprehension/             # optional
    <name>.review.mjs
    fixtures.json
    prompt.md
    rubric.md

Unified config dsh-eval.config.mjs

Drop one at the consumer package root; both CLIs walk upward from the working directory, and flags always override config:

export default {
  profile: 'headless',              // dsh profile
  repo: '../../deepseek-harness',   // relative, anchored at the config file's directory
  mode: 'mock',                     // behavior CLI's --mode default (review has none)
  failOnSkip: false,                // behavior CI gate default
  report: 'eval-report.json',       // --report default (anchored at the config dir)
  disableRows: ['gates'],           // plugin rows disabled by default; case-level declarations win
                                     // (explicit [] = all enabled, for gate-interaction cases)
}

Unknown keys fail loudly (typos never degrade silently). The disableRows semantics and the turn-close gate boundary contract are in docs/disablerows.md.

Docs

DocTopic
host-wiringpeer wiring (incl. the npm antique-peer trap), building the CLI, profiles, credentials, spawn requirements
reviewcomprehension review: experiment definition, sterile profile, artifacts, the six review rules
matchersthe full trace-matcher and mock-helper set (tool face / text face / model-visible face)
disablerowsdisableRows and the turn-close gate boundary contract
rowconfigthe rowConfig per-row config override contract (whole-segment replacement, restate needed keys)
intent-casesreal intent-case spec: when to write one, assertion face, guards, CI semantics
reportmachine-readable report structure (--format json / --report)
known-issuesknown issues and workarounds (e.g. REQUEST_EXTENSION in staged homes)
runner-apiprogrammatic runner API: runEvalCase options contract, EvalRunResult fields, crossing tiers for cliPath
experimentalexperimental subpath symbol list (escape hatch, no compatibility promise)

Runtime guarantees

The runner uses try/finally so temp directories and links are cleaned up on every path (prepare throwing, mock validation failure, spawn errors) — the real profile store is never polluted. The behavior and review CLIs share directory scanning (skipping .runs and node_modules); the behavior CLI validates case shapes and detects cross-file duplicate ids at load time, failing as early as possible.

License: MIT. The framework's own tests and release self-checks are carried by the repository CI and do not ship with the package.