← Back to home@moazzamak

dsh-train-dashboard

A right-pane training dashboard for the DeepSeek Harness: runs a command of yours that writes a JSON snapshot of a training run, and charts it, one series at a time, with the age of the numbers always on screen.

Stars
0
Language
JavaScript
Created
Oct 5, 2026
Updated
Oct 6, 2026
GitHub repo

Introduction

dsh-train-dashboard

A training dashboard for the DeepSeek Harness (DSH). It adds a Training dashboard tab to the right pane, beside Files, Terminal and Document, and charts a JSON snapshot of a training run: a KPI header, one series at a time on its own axis, the joints that make the series incomparable marked on the chart with the reason for each, and the next expected transition stated ahead of the curve.

The plugin does not parse your logs. It runs a command of your own, that command writes the snapshot JSON described below, and the tab reads it. So the plugin works with any training setup whose numbers you can get into that shape, and it never competes with your own parser for the truth about a run.

  • no build step, no runtime dependencies, plain JavaScript in two files
  • no process at all unless a tab is open and the snapshot is older than refreshSeconds
  • never writes anything itself: the snapshot is written by your command, and it is the only file the plugin causes to be written
  • the age of the numbers is always on screen, so a stale chart can never be mistaken for a live one
  • every word a break marker shows is your snapshot's own: the tab chooses where a marker sits and how it reads, never what it says

Package version 0.3.0 — snapshot contract version 1.


What you get

regioncontent
headerthe run's name, the snapshot's age with a fresh/stale badge, a Refresh button
KPI rowlatest step, the held-out reading, step wall time, throughput, memory peak, evaluation accuracy
chipsthe four metrics this tab leads with, then a N more series disclosure holding everything else the snapshot carries
chartONE series at a time on its own axis, with every break marker drawn on it, over a timeline you can select
registerunder the chart: every break marker in full — its leg, its kind, its reason, what that kind does to the curve, and where it was read from
frontierunder the register: the next expected transition, named as an expectation and not as a fact
footerthe snapshot path being read

The chips are a short list on purpose. A snapshot of this run carries sixty series, and a chip for every one of them is a wall of controls over a chart that draws one curve. The tab leads with the four it is steered by — the leg's own end-of-leg reading, throughput, VRAM peak and the exam — and the other thirty-six are one click away behind the disclosure. Nothing is dropped: a producer's own tag is still reachable, labelled by its own name.

One curve, not two. train/bpb_sealed and train/bpb_legval are the same chart. Measured on this run's snapshot: the curve carries 688 readings over legs 1–689 and the leg-end reading 730 over 1–731, and on all 688 legs they share the values are identical. The curve stops at the phase boundary, which is where the corpus transfer happens, and the leg-end reading keeps going for another 42 legs. Both are written by the producer, so the tab resolves the pair: where the leg-end reading exists the truncated copy is not offered at all, and where a producer writes only the curve the curve is the held-out reading. The KPI header names which of the two it read.

One series at a time is deliberate. Series carry different units — bits per byte, seconds, GiB, accuracy, counts — and drawing them on one axis would invite exactly the comparison the axis cannot support.

Breaks, kinds and the frontier

A long run crosses boundaries that are not learning, and a rise across one of them is not the model getting worse. The snapshot can carry them, and the tab draws each one as a vertical line carrying a shape and its leg number: the shape says which kind of joint it is, the number says which register row below reads it out, and the words stay in the register where there is room for them and where the reason, the meaning and the reading already live. Kinds are shapes and colours, because a kind encoded by hue alone is invisible to a reader who cannot separate the hues, and these kinds do different things to the curve:

kindshapewhat it means on the chart
corpuscirclethe data changed, so a rise here is expected and is not a regression
instrumentsquarethe reading changed, so the joint is a level shift measured on different terms: no slope across it
arithmetictrianglethe numbers changed, so the series is not comparable across the joint at all
shapediamondthe leg's geometry changed with its volume held constant, so the metric stays comparable and the wall clock moves
restarttriangle, point downthe process restarted, and the work of that window is not in the series
break (declared by break_leg)barthe snapshot named a leg and no kind, and this tab will not guess one
anything elseplusa kind this tab has never seen, drawn rather than hidden

A kind the tab has never seen is still drawn, with the plus shape and the producer's own name in the register; a marker whose kind the snapshot omits gets the UNKNOWN row rather than being shown as if the kind were known. Markers whose labels would collide are stacked so that no two write over each other, and every marker also gets a dim tick on the strip under the chart, so the joints are visible even when you have zoomed somewhere else. Hovering a marker gives the same text the register carries, so the detail is one gesture away without being written across the plot.

The register under the chart is where the markers read in full: the reason, the sentence your snapshot carries for that kind, the evidence the marker was read from, and any contradiction the snapshot attaches. Markers that fall outside the window on screen, or past the last point of the series, still appear there — the chart can only draw what its axis covers, and a marker that is missing from the picture without a word about it is a marker nobody finds.

The frontier is the next expected transition, drawn as a dashed line in a lane of its own and labelled next corpus change expected around leg N. It is deliberately not drawn like the markers above it: those are read from artifacts that exist, and this one is a projection. The tab says so, in the block under the register, and it also says why its dashed line may not be on the chart yet.

Both the markers and the frontier come from your snapshot, in your words. The tab does not decide that a rise is expected, that a joint is a level shift or that a transition is coming; it draws what your producer derived and repeats its sentences, so the chart and your own record cannot tell different stories about one run.

Selecting the timeline

The chart opens on the whole run, and the axis rescales to whatever window you select. That is the point of the interaction: on a long run one early spike can flatten the last few thousand steps — the part you are steering — into a straight line.

gestureeffect
drag across the chartzoom to that window
wheelzoom about the pointer
drag (or click) the strip under the chartselect on the whole run; a click moves the window there without changing its width
+ / − / Reset buttonszoom in, zoom out, back to the whole run
+ - 0 ← → with the chart focusedthe same, from the keyboard; Esc also resets
double-clickback to the whole run

The strip under the chart is always the entire run at a fixed scale, with the current window drawn on it, so you can see where you are and jump anywhere. The readout above the chart names the window in steps — steps 120–480 · 360 of 1200 shown — and the chart's accessible name carries the same numbers, because a range that exists only as a dragged rectangle is a range a screen reader cannot report.

There are two ways to open the tab: a chart glyph in the session header's action row (the conversation.session.header.actions seat, beside the shipped jobs and subagent controls at the top of the session), and the pane's "+" guide menu.

How it gets its data

your project                          plugin host half                 browser tab
────────────                          ────────────────                 ───────────
whatever your run writes
(metrics.jsonl, per-step logs,
 evaluation reports, a database)
        │
        │  YOUR command, run by the plugin
        │  at most once per refreshSeconds
        ▼
the snapshot JSON  ────────────────▶  GET /dsh-train-dashboard/state  ──▶ polled every 5 s
(the contract below)                  GET /dsh-train-dashboard/series ──▶ fetched only when
                                      (runs nothing itself)                the revision moves

Two readers, one parser: yours. A JavaScript re-implementation of your parsing would be free to disagree with it, and then the tab and your own records would tell different stories about one run with nothing on screen saying which is wrong.

The routes:

routepurpose
GET /dsh-train-dashboard/statesmall JSON: revision, age, staleness, whether a rebuild is in flight, whether this host can rebuild at all, and the exact command it would run
GET /dsh-train-dashboard/seriesthe whole snapshot body, plus ok: true; 503 while there is nothing yet
eitherGET only; anything else is 405 with allow: GET

/state is what the tab polls every 5 seconds. /series is fetched only when the revision changes, so a tab sitting open costs one small file read every few seconds and no process at all: the spawn happens only when the snapshot is older than refreshSeconds, and only one spawn can be in flight.

The host half is also usable on its own, without the browser tab, by anything that can make an HTTP request to the harness.

The snapshot contract

Your command writes one JSON file. That file is the whole interface between your project and this plugin. This is a complete example:

{
  "snapshot_version": 1,
  "generated_at": "2026-01-31T09:15:00Z",
  "revision": "1738314900",
  "arm": "run-7",
  "break_leg": 500,
  "series": [
    { "tag": "train/loss", "source": "metrics.jsonl", "points": [[1, 3.204], [2, 3.011], [3, 2.876]] },
    { "tag": "train/wall_seconds", "source": "step logs", "points": [[1, 41.2], [2, 39.8], [3, 40.5]] },
    { "tag": "memory/vram_peak_gib", "source": "step logs", "points": [[1, 11.4], [2, 11.6], [3, 11.5]] }
  ],
  "annotations": [],
  "breaks": [
    {
      "leg": 380,
      "kind": "corpus",
      "label": "corpus_a to corpus_b",
      "reason": "the corpus changed, so a rise in bits per byte here is expected.",
      "support": "the legs' own log names",
      "derived": true,
      "contradiction": "",
      "meaning": "the DATA changed, so a rise in bpb here is EXPECTED"
    }
  ],
  "frontier": {
    "phase_now": "corpus_b",
    "next_phase": "corpus_c",
    "boundary_leg": 380,
    "next_leg": 460,
    "legs_in_phase": 80,
    "legs_remaining": 61,
    "expected_to_raise_bpb": false,
    "imminent": false,
    "sentence": "FRONTIER: corpus_b is being trained; the next corpus transition is corpus_b to corpus_c at about leg 460."
  },
  "counts": { "points": 9 },
  "sources": "metrics.jsonl + step logs"
}
fieldrequiredwhat it does
snapshot_versionyesMust be the number 1. Any other value is refused as unusable and the tab says so; it does not offer a rebuild that would not help.
generated_atstrongly recommendedISO-8601 timestamp of when the numbers were read. Drives the age badge and the stale banner. Without it the tab reports the snapshot as stale, because it cannot know the age.
revisionrecommendedAny string that changes when the contents change: a timestamp, a file mtime, a hash. The tab fetches the series body only when this moves. If you omit it, generated_at is used as the revision.
armoptionalThe run's own name, shown in the header as run <arm>.
break_legoptionalOne x value to mark, for a producer that has one joint and no kind for it. It is drawn as a dashed marker labelled BREAK, and the tab says in the register that it is not claiming a kind for it.
breaksoptionalAn array of markers, each derived from your own artifacts. See below.
frontieroptionalThe next expected transition. See below.
seriesyes (may be empty)The array of series to chart.
series[].tagyesThe series name. A slash groups it: train/loss, memory/vram_peak_gib.
series[].pointsyesArray of [x, y] pairs of finite numbers. x is the step axis. A pair that is not two finite numbers is dropped; a series left with nothing is not offered at all.
series[].sourceoptionalShort string naming where the numbers came from; shown in the chip tooltip and the legend.
annotationsoptionalFree-form array, forwarded to the browser untouched.
countsoptionalFree-form object, e.g. {"points": 4123}; echoed by /state.
sourcesoptionalString naming the producer's inputs; echoed by /state as snapshotSources.

breaks[]: the markers

fieldrequiredwhat it does
legyesThe x value the marker sits at. A marker without a readable leg cannot be placed on an axis: it is counted and reported under the chart rather than dropped in silence.
kindstrongly recommendedOne of corpus, instrument, arithmetic, shape, restart — or any other word, which is drawn and named as it is. Omitted, the marker is labelled UNKNOWN.
labelrecommendedA few words naming the change (corpus_a to corpus_b, 512 to 256 steps a leg). Drawn on the chart, cut on a word boundary if it does not fit.
reasonrecommendedThe sentence that says what happened and whether a move across it is expected. Shown in full in the register.
meaningrecommendedWhat this kind of change does to the curve. Shown in the register under WHAT THIS <KIND> CHANGE DOES TO THE CURVE:.
supportoptionalWhere the marker was read from, in a sentence a reader can check. Shown as READ FROM:.
derivedoptionalfalse marks a leg that is a constant in your code rather than a reading; the register then says so. Absent means nothing is claimed either way.
contradictionoptionalSet when your own evidence contradicts the change the marker names. Shown prominently, because hiding it would leave the chart claiming a change the legs do not show.

Markers are drawn oldest first. Several markers on one leg are fine: their labels are stacked so they cannot overlap.

frontier: the next expected transition

fieldrequiredwhat it does
phase_nowrecommendedThe corpus being trained now.
next_phaserecommendedThe corpus expected next. Empty means none is scheduled, and the tab says exactly that instead of naming a leg.
next_legrecommendedThe leg the transition is expected at — a projection from your own budget, so the tab reads it as expected around leg N.
boundary_legoptionalThe first leg of the current phase. With imminent: true, this is the leg the tab draws, because nothing is left to predict.
legs_in_phase, legs_remainingoptionalContext for the frontier; shown inside sentence when your producer writes one.
expected_to_raise_bpboptionaltrue when a rise in bits per byte is expected at the transition. The tab then says so beside the expected leg.
imminentoptionaltrue when the next leg trained is already the new corpus. The wording moves from expected around leg N to changes AT LEG N.
sentencerecommendedYour own full statement about the frontier. Shown verbatim, under the tab's note that it is an expectation and not a measurement.

The tab always adds one caveat of its own to the frontier — that the leg is a projection from your budget and the rise is a direction rather than a magnitude — and it says why a dashed line that is past the end of the series is not on the chart. Everything else on screen is your text.

Tags the tab already has labels and units for (a convenience, not a requirement — any other tag is still charted, labelled by its own tag):

taglabelunit
train/bpb_sealedsealed bpbbpb
train/bpb_legvalleg-val bpbbpb
train/wall_secondsleg wall times
train/units_per_secondunits/su/s
train/gnormgrad norm—
train/vocabulary_sizevocab size—
memory/vram_peak_gibVRAM peakGiB
memory/vram_headroom_gibVRAM headroomGiB
eval/exam_accuracyexam accuracy—

Tags shaped like per-step traces (vram_trace/…, anything/trace/…, something_trace) are kept out of the chip list entirely: one chip per step would bury the series you came for. Give such a series a tag without that shape if you want it offered.

Writing a producer

Any program in any language that can write this file will do. Two rules make it pleasant to live with:

  1. Write it atomically (temp file plus rename). The tab polls the file while you write it; a torn read is handled — the tab keeps the last good numbers — but it should not happen on purpose.
  2. Make revision change exactly when the numbers change. A file mtime, a row count or a hash all work. If it never changes, the tab never refetches the series.

The whole contract, in one Python function:

# snapshot_writer.py — replace the reading below with your own.
import json, os, tempfile, time

def write_snapshot(path, leg, loss):
    body = {
        "snapshot_version": 1,
        "generated_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
        "revision": str(os.stat("metrics.jsonl").st_mtime),  # changes when the data does
        "arm": "run-7",
        "series": [
            {"tag": "train/loss", "source": "metrics.jsonl", "points": [[leg, loss]]},
        ],
    }
    directory = os.path.dirname(os.path.abspath(path)) or "."
    handle, temporary = tempfile.mkstemp(dir=directory, suffix=".tmp")
    with os.fdopen(handle, "w", encoding="utf-8") as stream:
        json.dump(body, stream)
    os.replace(temporary, path)  # atomic: a reader never sees half a file

Then point the plugin at it:

command: [python3, snapshot_writer.py, --out, "{snapshot}"]

Configuration

The plugin has one composition row. Its config is replaced wholesale by a higher patch layer and never deep-merged, so a profile that overrides the row must restate every key it cares about.

- id: train-dashboard
  name: dsh-train-dashboard
  config:
    workspace: /absolute/path/to/your/project
    command: [python3, -m, your_project.snapshot_tool, --out, "{snapshot}"]
    snapshotPath: train-dashboard.snapshot.json
    refreshSeconds: 30
    timeoutSeconds: 120
keydefaultmeaning
workspace'' — set itAbsolute path. The command's working directory, and the base that snapshotPath and any relative path resolve against. A host process can start in the profile directory, where a relative path resolves to nothing, so name it.
command[] — set itThe argv of your producer, one word per list item. Not a shell string: no shell is involved, so quoting and && do not apply.
snapshotPathtrain-dashboard.snapshot.jsonWhere your command writes and the tab reads. Relative to workspace, or absolute.
arm''Optional label shown in the header when the snapshot does not name the run itself.
refreshSeconds30Minimum spacing between regenerations, and half the staleness threshold (a snapshot older than twice this is reported stale). Raise it, do not lower it, while a training job is running on the same machine.
timeoutSeconds120A regeneration is abandoned after this and reported, rather than leaving a spinner forever.
any other key—Kept as written, and usable as a {placeholder} in command.

Placeholders in command. {snapshot} becomes snapshotPath, {workspace} becomes workspace, and any other {name} becomes the config key of that name. So a producer with its own arguments declares them as keys:

    command: [my-tool, --logs, "{logsDir}", --run, "{arm}", --out, "{snapshot}"]
    logsDir: logs/train
    arm: run-7

An unknown placeholder is left as written rather than blanked, so a typo shows up in the command the tab prints instead of silently becoming an empty argument.

What the plugin does with the command. It runs it with cwd set to workspace, captures up to 256 KB of each of stdout and stderr (the last few stderr lines are shown when the exit code is not zero), and asks it to stop after timeoutSeconds, with a short grace period. Success is judged by the snapshot, not by the exit code: a command that exits zero without writing a readable snapshot has not answered the tab's question, and one that writes a good snapshot has.

Install

Install the bundle into a profile with the plugin manager:

dsh plugin --profile <profile> add github:moazzamak/dsh-train-dashboard

Target the profile you actually boot, then restart the harness. The row arrives with placeholder values, so set workspace and command before expecting a chart; until then the tab says a command has not been configured.

To see the tab afterwards, expect to reload the page or restart DSH. The shipped dsh-client-hmr package delivers ordinary plugin enable/disable changes to an open page over /plugins/events without a page reload, but it states its own limit plainly: "Web transport only — Electron installation and backend restart handling do not use this SSE path." Nothing here needs pnpm run dev:web: there is no build step, so there is no bundle of ours to rebuild.

From a local checkout

Point the profile at this directory instead. In the profile directory (~/.dsh/profiles/<profile>):

  1. add this package as a dependency, with a file: specifier pointing at this directory;
  2. append its name to dsh.profile.bundles in the same package.json, so the bundle patch in this package is applied;
  3. run pnpm install --no-frozen-lockfile, with CI=true set: pnpm refuses to remove the modules directory without a terminal, and under CI it defaults to a frozen lockfile that the new dependency invalidates;
  4. restart the harness, and confirm the boot audit lists no pending entry.

Back up package.json first. If the install fails, restore it, otherwise the profile is left declaring a dependency that is not there and can fail to boot.

What it costs

  • No tab open, no cost. Nothing is polled and nothing is spawned.
  • Tab open, fresh snapshot. One small JSON read every 5 seconds, and no process at all.
  • Tab open, stale snapshot. The same poll, plus your command once per refreshSeconds, single-flight. The only cost the plugin itself adds is process start plus one file read; the rest is whatever your producer does, so make it cheap or raise refreshSeconds.

When things are wrong

A snapshot is written while the tab reads it, so a missing, stale or half-written one is normal, not exceptional. Every one of those states shows the last numbers it has with their true age, and says what is being done:

statewhat the tab shows
no snapshot, host can rebuild"No snapshot yet", what pressing Refresh will do, and the exact command it will run
no snapshot, no command configuredthat no command is configured and which config key to set
no snapshot, host cannot rebuild (no subprocess provider)the same, but saying so — so nobody waits for a rebuild that will never come
snapshot present but unreadable (version mismatch, not JSON)the host's own words, and that rebuilding will not fix it
snapshot older than twice the intervalthe chart still drawn, under a banner saying these are real numbers and not the newest
snapshot with no readable generated_atthe numbers charted, the badge reading "age unknown", and the banner saying the age cannot be told
a read that fails after a good onethe last good numbers are kept, with their age — never a blank chart
the host is unreachablethe transport error, in a notice, with the rest of the panel intact

Files

filerole
package.jsonthe manifest: dsh.bundle.patch, dsh.client, and the ./client export
cordis.patch.ymlone insert row, with placeholder defaults to replace
index.mjshost half: the two routes, the snapshot read, the bounded single-flight spawn
client.cjsbrowser half: the right-pane tab, the header action row trigger, and the hand-drawn SVG chart
test/host.test.mjs28 tests — the deferred registration, the read path, the spawn bounds, the pass-through
test/client.test.mjs54 tests — the slot wiring, the pure helpers, the freshness rules, the break markers and the frontier
test/live-smoke.mjsopt-in end-to-end check against a real project and a real command
npm test          # node --test test/host.test.mjs test/client.test.mjs
npm run check     # node --check index.mjs && node --check client.cjs
npm run test:live # needs DSH_TRAIN_DASHBOARD_WORKSPACE and ..._COMMAND; see the file

The live smoke test prints what it is missing and exits 0 when it is not configured, so a fresh clone never carries a red test nobody can satisfy.

Why there is no chart library here

Zero runtime dependencies, no build step, plain JavaScript, and the chart is hand-written SVG. This is not minimalism for its own sake:

  • the DSH client forbids importing Harness Client packages (@deepseek-ai/dsh-client-ui-primitives and friends) because they change without notice and a throwing component blanks the slot entry;
  • adding a charting library would add a second module the client loader has to resolve, and the loader resolves only the platform seed table, boot-graph rows, and registered factories — anything else throws;
  • d3 was not needed for one line chart, one axis and a break marker.

Zooming is a change of DOMAIN, not an SVG transform. Scaling the viewBox would also scale the strokes and the tick labels with it, so a zoomed chart would be a blurrier chart; recomputing the projection keeps every stroke one pixel and lets the ticks be the standard nice ones for the window on screen.

Design rules this follows, and the defects behind them

Each of these is a mistake some plugin in this harness has already paid for once:

  • One row in the patch. The boot audit fails on any entry left pending, so a second row that waited for the browser carrier would break the headless, SDK and ACP profiles, which never have one.
  • No ctx.shell. A desktop profile may mount no shell-executor row, and Cordis resolves an absent injected service leniently: the call registers fine and fails only on first use. This half uses ctx.fs and ctx.subprocess.
  • The browser carrier is read in ctx.inject, never in apply. Reading it during apply sees undefined, and the plugin then loads, the tab opens, and every fetch 404s.
  • No exported Config class. Cordis reads an exported Config as a schema and calls Config['~standard'].validate on it unconditionally, so the fiber dies at boot with no routes and the browser sees 404s that look like a missing snapshot. The function here is called normalizeConfig.
  • The tab body registers under the tab definition's id, not its kind. The pane dispatches with entryKey: definition.id ?? tab.kind, so a body keyed by kind renders the pane's own "Nothing here can view this kind of content yet." over a tab whose chip titles itself correctly.
  • Every slot registration names its slot. A registration without name throws while apply runs, which the desktop application treats as a failed startup.
  • An empty command is a state, not a crash. With no command configured the host spawns nothing, reports canRefresh: false, and the tab says which key to set — instead of running an empty argv and reporting a mystery.
  • The plugin never writes into your run directory. The only file it causes to be written is the snapshot, by your command, and the tab keeps the last good numbers when a read catches a partial one.
  • The axis is scaled to the samples IN the window, not to everything drawn. The line is drawn with one sample beyond each edge, so it meets the frame instead of starting in mid-air; scaling the axis to those edge samples hands the scale back to the sample just outside the window, so the spike you zoomed in to get away from keeps flattening the window. The tests pin both halves: pointsToDraw for the geometry, pointsInRange for the axis.
  • A window is clamped, and has a floor. Zooming is clamped to the run and stops at MIN_WINDOW_FRACTION of it, because a window with one or two samples in it is a chart that says nothing — and a drag past an edge slides back inside instead of compressing the window.
  • A zoom control is a button, not a gesture. Dragging is the fast path, but a range reachable only by dragging cannot be chosen by keyboard and cannot be read by a screen reader, so the chart takes focus (with +, -, 0 and the arrow keys) and every control carries a label.
  • The plugin owns its stylesheet by data-plugin. The client loader claims every style tag that lacks that attribute for whichever plugin materialises next, and deletes every tag whose data-plugin equals an id when that entry is replaced or pruned — so a privately tagged sheet is deleted with another plugin, which strips fill: none from the SVG paths and they fill black.
  • A break marker says what KIND of break it is, or says that it does not know. The kinds do different things to a curve — a corpus change explains a rise, an instrument change makes the joint a level shift, an arithmetic change makes the series incomparable across it — so a marker drawn as a bare line invites exactly the comparison it should prevent. The kind is a word on the line and a colour, an unknown kind keeps its own name, and a marker served without a kind is labelled UNKNOWN rather than coloured like something it may not be.
  • A marker's words are the snapshot's, never this plugin's. The label, the reason, the sentence for the kind and the evidence are carried through verbatim; the tab chooses where a marker sits, how much of it fits on the chart and how it is stacked, and never what it says. A paraphrase here would be a second statement of one fact, free to disagree with the record.
  • A marker that is not drawn is still reported. Markers past the end of the series or outside the window still read in the register, a marker with no readable leg is counted and reported rather than dropped, and the frontier says why its dashed line is not on the chart yet. A silent absence in a chart is read as "nothing happened here".
  • The frontier is drawn as an expectation. It is a projection from the producer's own budget, in its own dashed style and its own lane, labelled expected around leg N — and when the producer says the transition is imminent the wording moves to AT LEG N, because there is nothing left to predict. The tab adds one caveat of its own: the leg is a direction and not a magnitude, and the producer may end a phase early.
  • snapshot_version stayed at 1 while the new keys arrived. The contract version is a claim that a reader can read the body, and breaks and frontier are additive: an older reader serves them and simply draws less, while a bumped version would make every deployed reader refuse the snapshot outright. The host is a pass-through and a test pins that it does not reshape or drop a key it does not know.
  • An action goes in an action row, never in shell.overlay. That seat is a frame-wide floating layer for badges, toasts and status pills, and it is click-through on purpose, so a BUTTON registered there is in the wrong place (the window's top-left, beside the application menus and the sidebar's reopen control) and unclickable unless it opts back into pointer events. The trigger belongs to conversation.session.header.actions, the title-adjacent row the shipped jobs, subagent, agent-preset and agent-team controls use, and it copies their metrics — borderless, transparent, at least 28px, tertiary label colour at rest and secondary on hover — so the row reads as one set of controls. The test asserts the trigger is in that row and that nothing of this plugin is registered in the overlay.
  • The selected series belongs to the reader, not to the payload. The snapshot is rebuilt from logs on disk on every refresh, so a series can drop out of one revision and come back in the next: the live snapshot rotates its per-leg traces, and one refresh replaced guard_trace/…_leg_00722_… with …_leg_00725_…. The panel resolved the selection with offered.find(tag) ?? offered[0], so a series missing from a revision silently moved the reader onto the default curve, which looks exactly like a redraw and is actually the panel discarding a choice. resolveSelection reports the choice as MISSING instead, its chip stays marked with a dashed edge and its own words, and nothing else is drawn in its place. Those trace tags are filtered out of the chips too, by the same shape rule that already excluded vram_trace/.
  • The choice and the window outlive the component. The pane unmounts and mounts a tab body again on events the reader did not cause, and React state does not survive that, so both are remembered outside the component: the series per session, the zoomed window per series tag. Reset forgets the window, which is what keeps "reset" and "never zoomed" one state.
  • Every control is pressed by a check, not merely rendered by one. The zoom buttons were wired the wrong way round, + widening the window and - narrowing it, through a passing unit suite: the tests asserted that the labelled controls existed rather than what they did. The end-to-end check clicks each control and reads the window back, which is how the inversion was found.
  • The chart carries identifiers; the register carries the words. A marker was drawn with its leg, its KIND and the first words of its reason, which repeated what the register row below already said in full and cost the reader the curve: six joints inside fifty legs stacked into six lanes of prose across the plot. The plot now carries a shape and the leg number, and the register carries the prose. The shape is not decoration: a kind encoded by hue alone is invisible to a reader who cannot separate the hues, and these kinds do different things to the curve, so an unreadable kind invites the comparison the register exists to prevent. The register draws the same shape in its key, and a test reads the stylesheet and compares it against the shape table, because one mapping written twice is free to disagree with itself.
  • An expectation is not drawn like a fact. The frontier is the one unfilled swatch, in the label colour, on a dashed line, and its sentence sits in the register under a heading that calls it an expectation. A predicted leg drawn like a joint read from an artifact would be a guess wearing a measurement's clothes.
  • A default list is a claim about what matters, and it goes stale. The chips offered nine metrics in a fixed order chosen when the tab was written, and the snapshot had since grown to sixty series: the row became a wall of controls and the one at the front was a truncated duplicate of the one behind it. The list is now four, it is one constant with the reason for each entry beside it, and everything else is behind a disclosure rather than gone. A default nobody revisits is a default that decides for the reader.
  • When two series are one chart, the tab resolves them. train/bpb_sealed and train/bpb_legval are identical on all 688 legs they share; the first stops at the phase boundary and the second continues 42 legs past it. So one entry names the other as its fallbackFor: where both exist the truncated one is not offered at all, and where only the curve exists it is the reading. A reader should never be handed two chips for one picture and left to work out which is which.

Licence

MIT. It matches the sibling plugin dsh-knowledge-dag, whose licence this one follows, and DeepSeek Harness itself, which is also MIT.