system-one-benchmark-lab
No description
- Stars
- 1
- Language
- Python
- Created
- Sep 26, 2026
- Updated
- Oct 7, 2026
Introduction
Local System One
A local typed-decision control plane for Apple Silicon agents.
让本地小模型负责 Agent 的高频控制决策,让 LLM 专注真正需要生成、规划和深度推理的工作。
Local System One is not another chatbot and does not replace your LLM. It sits before or around an agent's main model and answers small but frequent questions such as:
- Does this task need current Web information?
- Can this request use a cheaper/faster reasoning tier?
- Should this event stay silent, go to a digest, or interrupt the user now?
- Which backend should handle this decision?
- Is a local ANE path healthy enough to use right now?
The project packages a validated Laya Typed Decisions 421M runtime, an optional Apple Neural Engine (ANE) backend, HTTP/MCP interfaces, a reversible Hermes plugin, and a native DeepSeek Harness adapter.
Why this exists
Large language models are good at generating and reasoning, but agents repeatedly spend those expensive calls on tiny routing decisions.
Local System One turns those decisions into a separate local control layer:
User / Agent
|
v
Local System One
|
+-- hard rules
+-- typed-decision model
+-- policy / health gates
|
v
LLM / tools / notification / routing action
The goal is not "use a small model for everything". The goal is to use a small local model where a structured decision is enough, while keeping the system reversible and fail-open.
What it can decide
| Workflow | Question | Current status |
|---|---|---|
| Search Gate | Does this task need public/current Web information? | Functional. Deterministic hard rules can act; model-probability routing remains conservative/Shadow-first. |
| Model Tier Gate | Can a bounded task use a faster reasoning tier? | Functional, but broad model-based routing remains experimental. |
| Notification Gate | silent / digest / notify_now? | Functional MVP. |
| Choice / Score / Noul | Generic typed decisions for your own control logic | Available through the HTTP API. |
This distinction matters: the project deliberately separates technical capability from what has enough evidence to activate automatically.
Quick start
Requirements
- macOS on Apple Silicon
- Python 3.11–3.13
- Internet access on first model download
Clone and start the default MLX backend:
git clone https://github.com/NMTZ-z/system-one-benchmark-lab.git
cd system-one-benchmark-lab
python3.12 -m venv .venv
. .venv/bin/activate
python -m pip install -e .
local-system-one --host 127.0.0.1 --port 8787
The default source alias is laya-typed-decisions. It resolves the validated checkpoint revision:
convaiinnovations/laya-typed-decisions
f9ab0b228f0fc0f14d873dbc99038f135c2da1b2
Check the service:
curl http://127.0.0.1:8787/health
Available endpoints include:
POST /v1/choice
POST /v1/score
POST /v1/noul
POST /v1/workflows/search-gate
POST /v1/workflows/model-tier-gate
POST /v1/workflows/notification-gate
GET /health
GET /metrics
The service binds to loopback by default and does not persist raw task text.
Optional Apple Neural Engine backend
Install the ANE dependencies:
python -m pip install -e '.[ane]'
Download the validated fixed-shape L512 Core ML package from Hugging Face:
hf download NMTZ/laya-typed-decisions-421m-coreml-ane \
--include 'L512/*' \
--local-dir artifacts/models/laya-typed421-ane
Then start Local System One with the downloaded package:
local-system-one \
--ane-package artifacts/models/laya-typed421-ane/L512/model.mlpackage
The Hugging Face release also includes the validated L192, L384 and L640 research variants. L512 remains the recommended default because it covered 1,966 / 2,000 frozen benchmark decisions while preserving 99.8% selected-decision agreement with the MLX reference.
If you prefer to reproduce the conversion locally, build the same validated fixed-shape L512 package with:
scripts/build_typed421_ane.sh
Then start Local System One with the locally built package:
local-system-one \
--ane-package artifacts/models/typed421-body512-fp16/model.mlpackage
The builder pins laya-coreml to:
4619e0483f07adf39068532e85b42ec2347edb83
The original 421M checkpoint is not duplicated by this project. The converted fixed-shape ANE artifacts are published separately at NMTZ/laya-typed-decisions-421m-coreml-ane, with provenance, SHA256 metadata and benchmark evidence. The Hugging Face packages are ANE transformer-body artifacts and still use the pinned upstream Laya checkpoint for tokenizer, embedding lookup and the host-side action head.
At runtime, the router does not blindly force ANE. It uses token length plus an ANE health gate and falls back to MLX when the accelerator path is unavailable, unhealthy, or unsuitable.
Hermes integration
The repository includes a native Hermes plugin with three modes:
off -> shadow -> canary
First install is inert by default.
For an evaluation profile:
scripts/install_hermes_system_one_plugin.sh systemoneeval
scripts/set_hermes_system_one_mode.sh systemoneeval shadow
scripts/hermes_system_one_status.sh systemoneeval
Shadow mode observes decisions without changing the provider request. Canary requires explicit acknowledgement and remains deliberately narrow.
Rollback is built in:
scripts/uninstall_hermes_system_one_plugin.sh systemoneeval
The plugin is designed to fail open: if Local System One is unavailable or a recommendation is incomplete, Hermes keeps its original request unchanged.
See Hermes plugin documentation for the full lifecycle and safety constraints.
DeepSeek Harness integration
The repository includes local-system-one-dsh, a native adapter for the official DeepSeek Harness plugin lifecycle. Phase 7A/7B upgrades it to Search Gate + Model Tier Gate + Notification Gate (Shadow) parity with the Hermes integration while preserving reversible, fail-open behavior.
Validated DSH targets are deliberately exact:
0.1.7-rc.2 @ 477b4f420553e8a52c2fbccc464d7561b239c443
0.2.1-alpha.1 @ 5badb15009ae1756c3afe0ae0cef1faafc290ccc
The adapter supports independent Search, Model Tier, and Notification observation. Search and Model Tier run at turn start; Notification observes the completed Agent event at turn end and remains Shadow-only. Search Canary retains the audited hard no-Web rules. Model Tier Canary is experimental and disabled by default; when explicitly acknowledged and enabled, only an audited deterministic hard-fast decision on an exact verified provider/model route may copy the first provider request and change reasoningEffort: high -> low. Model/probability decisions never receive active authority.
Real provider-wire validation paths include:
DeepSeek Harness -> Nova -> DeepSeek V4.1 Flash
DeepSeek Harness -> StepFun Step Plan -> step-5-preview
The StepFun route completed native high/low requests and an automatic Local System One high-to-low Canary transition, including tool-loop restoration to high after the first provider call.
The installation bundle remains inert in off mode. Connection errors, timeouts, malformed gate responses, unsupported provider/model routes, and any missing Canary authority fail open to native DSH behavior. Turn-scoped Search and Model Tier state is cleared at turn end.
Community bundle install for a supported headless profile:
dsh plugin --profile headless add github:NMTZ-z/system-one-benchmark-lab#dsh-plugin
See DeepSeek Harness adapter documentation, the Phase 0 / Phase 1 probe report, and the Phase 7A Model Tier parity report.
Architecture
Agent / Hermes / DeepSeek Harness / MCP client
|
v
Local System One API
|
v
Decision Engine
/ | \
rules policy typed model
|
v
Router
/ \
MLX L512 ANE
^ |
+-- health/fallback
Main components:
local_system_one/— service, router, health gate, runtime adapters and workflowsintegrations/hermes/— reversible Hermes native pluginintegrations/deepseek-harness/— reversible DeepSeek Harness Search Gate adapterscripts/— local service, ANE build, launchd and plugin operationsdocs/— design notes and product validationresults/reports/— frozen public benchmark reports
Verified results
The repository grew out of a reproducible JEV/Laya/ANE evaluation program. A few results matter directly to the product:
Typed Decisions quality
On the public 400-case / 2,000-decision benchmark:
| Backend / model | Accuracy |
|---|---|
| Laya Typed Decisions 421M MLX | 0.766 |
| Jev 1.13.0 | 0.737 |
These numbers describe this benchmark only; they are not a general model ranking.
L512 ANE engineering
For the validated 421M fixed L512 path:
- natural benchmark coverage: 1,966 / 2,000 decisions (98.3%)
- selected-decision agreement vs same-subset MLX: 99.8%
- representative gross system energy per decision: about 2.07× better than MLX
- L640 reaches 100% capacity coverage, but was not the best default latency/efficiency tradeoff
The key result is energy-efficient local inference with runtime health gating, not a claim that ANE is universally faster for every request.
Reproducibility
The public release path has been validated from an isolated sanitized bundle:
- clean Python installation
- pinned model acquisition
- pinned
laya-coremlcheckout - fresh L512 conversion
- Runtime startup probe
- old-vs-rebuilt output parity checks
See Public release notes.
Current limitations
This is an alpha control-plane project, not a universal autonomous router.
Important limits:
- Search model probabilities are not trusted as unrestricted final authority.
- Model Tier broad automatic downgrade is not proven to provide a stable latency benefit.
- The ANE backend uses a fixed L512 body and must fall back when requests do not fit or health checks fail.
- Notification preferences are not personalized yet.
- Hermes Canary is intentionally narrow and requires explicit acknowledgement.
- A second physical Mac would strengthen cross-machine reproducibility evidence, although the clean public-path rebuild already passes on the development M4 Mac mini.
Conservative defaults are intentional. A wrong routing decision can be more expensive than the small amount of compute it saves.
Documentation
Start here depending on what you want to do:
- Local System One product and API
- Search Gate
- Model Tier Gate
- Notification Gate
- MCP and launchd deployment
- Hermes plugin productization
- DeepSeek Harness adapter
- DeepSeek Harness validation report
- DeepSeek Harness Phase 7A Model Tier parity
- Phase 7B Notification + StepFun validation
- Public release / reproducibility
- Upstream provenance
For the underlying experiments:
- Jev Typed Decisions benchmark
- Laya 421M quality benchmark
- Core ML vs MLX parity
- Phase 4 ANE engineering
Project philosophy
System One is useful here as an agent control primitive, not as a replacement for a general-purpose LLM.
The practical pattern is:
hard rule when certainty is available
->
small local probabilistic decision when useful
->
policy / confidence / health gate
->
LLM or action
That makes the control layer cheap, local, inspectable, reversible, and easy to disable.
License
Local System One is released under the Apache License 2.0.
The project depends on and documents upstream work separately. See references/UPSTREAM.md for pinned revisions and provenance.