@hraness/sys1 0.17.0 → 0.17.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,677 +1,226 @@
1
1
  # Sys1
2
2
 
3
- Sys1 helps coding agents review changes against your repository's rules, with
4
- probability-scored answers from hosted Jev, a local model, or your own server.
5
- Review and final-message verification are experimental and advisory.
3
+ Sys1 gives coding agents tools to review code, check completion claims, and get
4
+ structured answers from Jev or a local model. Jev is TypeSafe’s hosted decision
5
+ model. Project skills work with Codex, Claude Code, and Devin; applications can
6
+ use the same tools through a CLI, Node/Bun client, or HTTP API.
6
7
 
7
- Underneath, Sys1 lets agents ask yes/no, choice, and score questions and get
8
- validated answers with probabilities. You choose who answers: TypeSafe's hosted
9
- [Jev](https://docs.typesafe.ai/models), a local model on your machine, or a
10
- compatible server you run.
8
+ Latest release: v0.17.1. Install the GitHub release with npm and run it with
9
+ Bun 1.3.14 or newer. MIT licensed.
11
10
 
12
- Call Sys1 from a small Node/Bun client, embed the router in a Bun app, or run a
13
- local daemon that serves the Jev-compatible `POST /v1/systemone` API.
11
+ [Get started](#install) · [Agent skills](https://sys1.io/skills) · [Documentation](https://sys1.io/docs) · [Releases](https://github.com/hraness/sys1/releases)
14
12
 
15
- Latest release: v0.17.0. Install it from the GitHub release with npm; it runs
16
- on Bun 1.3.14 or newer.
13
+ ## Choose what helps your workflow
17
14
 
18
- [Project site](https://sys1.io) · [Agent skills](https://sys1.io/skills) · [Protocol](#the-endpoint) · [Routing](#routing)
19
-
20
- ## Choose a workflow
21
-
22
- | Task | Start here | What you get |
15
+ | You want to | Use | What you receive |
23
16
  | --- | --- | --- |
24
- | Review an agent's code changes | [`sys1 review`](docs/review.md) | Candidates tied to selected rules and diff evidence, with feedback and reuse of unchanged reviews. |
25
- | Check a completion message | [`sys1 verify`](docs/verify.md) | A comparison of claimed commits, pushes, checks, and live changes with reachable evidence. |
26
- | Run one review without saved history | [`sys1 audit`](docs/audit.md) | A stateless diff check with a preview, request cap, and skipped-evidence report. |
27
- | Reuse a repository convention | [`sys1 rules`](docs/review.md#draft-a-repository-rule) | A draft you can inspect, test on examples, and activate for future reviews. |
28
- | Add decisions to an application | [Node/Bun client](#use-as-a-module) or [HTTP API](#the-endpoint) | Yes/no, choice, and score answers with a selected model and validated response shape. |
29
-
30
- The review and verification skills work in Git repositories with Codex,
31
- Claude Code, or Devin; other agents can use the same CLI. Their workflows do
32
- not depend on a specific framework or hosting provider. The bundled review
33
- rules cover two narrow JavaScript and TypeScript changes: newly empty catch
34
- blocks and removed test assertions. Add and evaluate your own rules for other
35
- conventions or languages.
17
+ | Keep long test logs out of the agent’s context | [System One Skills](https://sys1.io/skills#compact-checks), a separate package | A short result, the command’s exit status, and the full log saved locally. No model or API key. |
18
+ | Check a change against a repository rule | [`sys1 review`](#review-changes-with-your-agent) | Candidates tied to the rule and diff, with feedback and reuse of unchanged reviews. |
19
+ | Check a proposed completion message | [`sys1 verify`](#check-an-agents-completion-message) | A comparison with Git state, linked pull requests, and live pages. |
20
+ | Add a decision to application code | [Client and HTTP API](#use-as-a-module) | Yes/no, choice, or score answers with probabilities and validated response shapes. |
21
+
22
+ For compact checks alone, [install System One Skills](https://sys1.io/docs#compact-checks).
23
+ It needs Node.js 20+ on macOS or Linux and works independently of Sys1. Its
24
+ replay of 563 validation outputs measured 35% less text at the tool output;
25
+ whole-task token savings have not been demonstrated.
36
26
 
37
27
  ## Install
38
28
 
39
- Requires Bun 1.3.14 or newer. Install the release file from GitHub. Its SHA-256
40
- is listed on the release, and release files cannot be replaced after
41
- publishing. `--allow-scripts=node-llama-cpp` lets only the pinned native
42
- inference package run its install script. The installed `sys1` command runs
43
- with Bun.
29
+ Install [Bun](https://bun.sh/docs/installation) 1.3.14 or newer and have npm
30
+ available, then run:
44
31
 
45
32
  ```sh
46
33
  npm install --global --allow-scripts=node-llama-cpp \
47
- https://github.com/hraness/sys1/releases/download/v0.17.0/hraness-sys1-0.17.0.tgz
48
- sys1 doctor
49
- ```
50
-
51
- To build the current source instead:
52
-
53
- ```sh
54
- git clone https://github.com/hraness/sys1.git
55
- cd sys1
56
- bun install
57
- bun run build:dist
58
- ln -sf "$PWD/dist/cli.js" ~/.local/bin/sys1
34
+ https://github.com/hraness/sys1/releases/download/v0.17.1/hraness-sys1-0.17.1.tgz
35
+ sys1 --version
59
36
  ```
60
37
 
61
- ## Review changes with your agent (experimental)
38
+ The version command prints the installed release number. The release supports macOS, Linux,
39
+ and Windows. `--allow-scripts=node-llama-cpp` allows the optional local inference
40
+ runtime’s install script; installation downloads no model weights and enables
41
+ no hosted backend. The release includes a SHA-256 checksum.
62
42
 
63
- The project skill teaches an agent when to preview a review, investigate a
64
- candidate, record feedback, and check its completion message. It installs
65
- instructions in your repository, without activating a model or adding hooks.
43
+ ## Review changes with your agent
66
44
 
67
- First [enable hosted Jev](#add-hosted-jev) with `TYPESAFE_API_KEY` available in
68
- your environment, or choose another configured backend. Install the skill and
69
- preview selected files before sending source to that backend:
45
+ Run these commands inside the Git repository you are working on. They install
46
+ project instructions and preview a review without calling a model:
70
47
 
71
48
  ```sh
72
49
  sys1 review setup codex
73
- sys1 review checkpoint --staged --model typesafe/jev-1.13.0 \
50
+ sys1 review checkpoint --worktree --model typesafe/jev-1.13.0 \
74
51
  --max-requests 10 --dry-run --json -- src test
75
52
  ```
76
53
 
77
- Use `setup claude-code` for Claude Code or `setup devin` for Devin. Inspect the
78
- preview, then repeat without `--dry-run` to run the checkpoint. Investigate
79
- candidates, record feedback, and recheck the original evidence. Repeated
80
- complete batches reuse their review for up to 24 hours. Add `--rule <id>` to
81
- focus a checkpoint on a rule from `sys1 rules list`.
82
-
83
- Findings are advisory, and scores are not calibrated defect probabilities. See
84
- the [agent review guide](docs/review.md) for setup, feedback, rechecks, and rule
85
- drafts. For a stateless check, use [`sys1 audit`](docs/audit.md). Both report
86
- skipped evidence and incomplete coverage.
87
-
88
- ## Check an agent's completion message
89
-
90
- Install the standalone verification skill when you want completion checks
91
- without the review workflow:
92
-
93
- ```sh
94
- sys1 verify setup codex
95
- ```
96
-
97
- Use `setup claude-code` or `setup devin` for those agents. The installed
98
- instructions tell the agent to save its proposed final message and compare
99
- its claims with the repository, linked pull requests, and live pages. Save the
100
- draft outside the Git worktree so it does not appear as an uncommitted change.
101
- You can also run the command directly:
54
+ For Claude Code, use `setup claude-code`; for Devin, use `setup devin`. Replace
55
+ `src test` with the paths you changed. The preview lists selected files, rules,
56
+ skipped evidence, and the number of requests. If no requests are planned, inspect the skipped
57
+ evidence, selected paths, and active rules. Setup adds no automatic hooks and preserves
58
+ existing instructions.
102
59
 
103
- ```sh
104
- sys1 verify --message /tmp/final-message.txt --model typesafe/jev-1.13.0
105
- ```
106
-
107
- Use `--message -` for piped text and `--url https://example.com` to include a
108
- live page the message does not link. Without `--message`, verify reads the
109
- newest matching local Devin session, including that turn's check-command
110
- results. File and stdin input work with any agent; check-command results are
111
- available only through Devin transcript discovery. Contradicted claims exit 7;
112
- missing evidence is unverifiable. A clean report does not prove task completion.
113
- See [the verification guide](docs/verify.md) for evidence and limits.
60
+ The bundled rules cover newly empty catch blocks and removed test assertions
61
+ in JavaScript and TypeScript. Add [your repository’s rules](docs/review.md#draft-a-repository-rule)
62
+ for conventions that need judgment. For an exact syntax pattern, use a linter.
63
+ Review and completion checks are experimental and advisory: investigate
64
+ findings and keep the repository’s normal tests and review.
114
65
 
115
- ## Use as a module
66
+ ### Add hosted Jev
116
67
 
117
- For a Node 24 or Bun application that calls a running gateway, install the
118
- release package without the optional native runtime:
68
+ Get a key from [TypeSafe](https://console.typesafe.ai/) and provide it as
69
+ `TYPESAFE_API_KEY` through your shell or secret manager. Then enable Jev:
119
70
 
120
71
  ```sh
121
- npm install --omit=optional \
122
- https://github.com/hraness/sys1/releases/download/v0.17.0/hraness-sys1-0.17.0.tgz
123
- ```
124
-
125
- ```ts
126
- import { createClient } from "@hraness/sys1/client";
127
-
128
- const sys1 = createClient(); // http://127.0.0.1:13900
129
- const { response, metadata } = await sys1.evaluate({
130
- state: "The build failed after a dependency upgrade.",
131
- questions: {
132
- action: {
133
- type: "choice",
134
- criteria: { repair: "Fix the build", continue: "Continue work" },
135
- },
136
- },
137
- }, { signal: AbortSignal.timeout(5_000) });
138
-
139
- console.log(response.answers.action, metadata.backend);
140
- ```
141
-
142
- The client validates inputs and correlates every returned answer with its
143
- question. It bounds response bytes, supports cancellation, and returns stable
144
- sanitized `Sys1ClientError` codes. It never retries, reads credentials from the
145
- environment, starts a daemon, downloads weights, or imports native inference.
146
- Supply `baseUrl` and `headers` explicitly for another approved endpoint.
147
- Import schemas and request/response types from the same `/client` entry point.
148
-
149
- For a Bun application that owns routing and model lifecycle in-process:
150
-
151
- ```ts
152
- import { createRouter, DEFAULT_CONFIG } from "@hraness/sys1";
153
-
154
- const router = createRouter({
155
- config: DEFAULT_CONFIG,
156
- env: process.env,
157
- home: "/absolute/path/to/sys1-state", // previously installed models
158
- });
159
- try {
160
- const result = await router.evaluate({
161
- state: "All required checks passed.",
162
- questions: { ready: { type: "noul", instructions: "Are the checks passing?" } },
163
- });
164
- console.log(result.response.answers.ready);
165
- } finally {
166
- await router.dispose();
167
- }
168
- ```
169
-
170
- The embedded router opens no port. It uses the same routing and validation as
171
- the daemon and owns its local runner until disposal. Its runtime requires Bun;
172
- the `/client` entry point is portable to Node. Keep one router per application,
173
- not one per request. The root package also exposes lower-level routing and
174
- model-management APIs; applications should normally use `createClient` or
175
- `createRouter`.
176
-
177
- ### Adopting Sys1 in an existing Jev application
178
-
179
- Keep domain questions, deterministic fallback, action authorization, and quality
180
- thresholds in the application. Put endpoint configuration, transport, routing,
181
- response validation, and local engine lifecycle behind Sys1. Existing HTTP
182
- clients in other languages can use the same daemon without a JavaScript module.
183
-
184
- Use `model: "auto"` or omit `model` to use the configured routing policy and
185
- selected local model. A hardcoded `jev-1.13.0` remains a model pin and cannot
186
- select an unrelated local model.
187
- Local calls need no hosted API key; hosted activation stays explicit. A remote
188
- server's loopback address points to that server. A browser running on the user's
189
- machine can address local services, so Sys1's network listener rejects browser
190
- origins and Fetch Metadata site headers, requires a loopback request authority,
191
- and accepts decision POSTs only as `application/json`.
192
-
193
- Start with an opt-in, non-authoritative pilot. Compare decisions on the
194
- application's representative fixtures and record backend/adapter identity,
195
- latency, errors, abstentions, and disagreement with the current decision path.
196
- Do not reuse Jev probability thresholds for generic GGUF output
197
- without model-specific evidence. A local-only policy also constrains explicit
198
- pins; a pin never bypasses the policy. Broad production adoption requires the
199
- consumer's own quality and operational acceptance, not just wire compatibility.
200
-
201
- ## Quickstart: experimental local decisions
202
-
203
- ```sh
204
- sys1 setup # verifies the native runtime, installs and selects Qwen3 1.7B
205
- sys1 up # starts the gateway on 127.0.0.1:13900
206
- sys1 status
72
+ sys1 jev enable
73
+ sys1 jev status
207
74
  ```
208
75
 
209
- Local Qwen is experimental. Do not treat it as a drop-in replacement for Jev.
210
- The broader tests found 32/72 correct decisions for Qwen3 1.7B and 44/72 for
211
- Qwen3.5 4B on a different fresh fixture. [Read the evidence](https://sys1.io/docs/evaluations)
212
- before using local decisions to drive actions.
76
+ The key stays in the environment. Enabling Jev selects `hosted-only` routing,
77
+ so a failed hosted request cannot silently use a local model. Selected source
78
+ and diff context go to Jev; inspect the preview before sending them. Provider
79
+ usage is billed by TypeSafe at its [published rates](https://docs.typesafe.ai/models).
213
80
 
214
- `setup` is the explicit weight-download boundary. It installs Qwen3 1.7B and
215
- persists that choice as `local.model`, regardless of system memory. Inspect
216
- without changing anything:
81
+ Run the previewed check by removing `--dry-run`:
217
82
 
218
83
  ```sh
219
- sys1 setup --dry-run --json
84
+ sys1 review checkpoint --worktree --model typesafe/jev-1.13.0 \
85
+ --max-requests 10 --json -- src test
220
86
  ```
221
87
 
222
- `sys1 setup --tier compact` explicitly installs and selects the experimental
223
- Qwen3 0.6B diagnostic model. It is not an automatic low-memory fallback.
224
- `sys1 setup --tier quality` returns the selection to Qwen3 1.7B.
225
-
226
- The pinned llama.cpp runtime selects the best available backend automatically:
227
-
228
- | Package target | Runtime preference |
229
- | --- | --- |
230
- | macOS ARM64 | Metal (the pinned runtime's only automatic selection) |
231
- | macOS x64 | CPU |
232
- | Linux x64 | CUDA, Vulkan, then CPU |
233
- | Linux ARM64 | CPU |
234
- | Windows x64 | CUDA, Vulkan, then CPU |
235
- | Windows ARM64 | CPU |
88
+ No background gateway is needed for this hosted workflow. Investigate each
89
+ candidate and record it as useful, incorrect, or unverifiable. Complete,
90
+ unchanged batches reuse their review for up to 24 hours. The [review guide](docs/review.md)
91
+ covers feedback, rechecks, and rule drafts; [`sys1 audit`](docs/audit.md) runs a
92
+ check without saved review history.
236
93
 
237
- Other platform/architecture pairs fail closed before downloading a model. The
238
- release artifact and package smoke are exercised on Ubuntu, macOS, and Windows.
239
- This matrix does not prove every OS/architecture pair above; run
240
- `sys1 doctor` on the actual host before use.
94
+ ## Check an agent’s completion message
241
95
 
242
- Send a decision:
96
+ After setting up a backend, install the standalone verification instructions:
243
97
 
244
98
  ```sh
245
- sys1 eval <<'EOF'
246
- {
247
- "state": "Help! My payouts have been failing for 3 days.",
248
- "questions": {
249
- "urgent": {
250
- "type": "noul",
251
- "instructions": "Does this need immediate attention?",
252
- "criteria": {
253
- "true": "A customer-impacting incident is ongoing",
254
- "false": "This can wait for normal triage"
255
- }
256
- }
257
- }
258
- }
259
- EOF
99
+ sys1 verify setup codex
260
100
  ```
261
101
 
262
- Or point any System One client at `http://127.0.0.1:13900`.
263
-
264
- ## Add hosted Jev
265
-
266
- Hosted Jev is disabled by default, even if `TYPESAFE_API_KEY` is already set in
267
- the environment. Add it explicitly:
102
+ Use `setup claude-code` or `setup devin` for those agents. Ask the agent to save
103
+ its proposed final message outside the Git worktree, then compare the draft
104
+ with available evidence:
268
105
 
269
106
  ```sh
270
- export TYPESAFE_API_KEY=…
271
- sys1 jev enable
272
- sys1 jev status
107
+ sys1 verify --message /tmp/final-message.txt \
108
+ --model typesafe/jev-1.13.0 --dry-run --json
273
109
  ```
274
110
 
275
- `jev enable` requires the credential to be present, stores only
276
- `hosted.enabled: true`, and sets routing to `hosted-only`. The key remains in the
277
- environment and is never written to disk or printed. Restart a gateway that was
278
- started before the key was exported. To return to local-only operation:
279
-
280
- ```sh
281
- sys1 jev disable
282
- ```
111
+ Choose an equivalent external file path on Windows. The preview reports the
112
+ input type and model route without model calls or page fetches. Read the message
113
+ file and inspect its links before rerunning without `--dry-run`; the message and
114
+ any fetched page excerpts go to the selected backend. Use `--message -` for
115
+ piped text and `--url https://example.com` to include a live page the message
116
+ does not link.
283
117
 
284
- Enabling Jev selects `hosted-only`: a hosted outage or missing credential returns
285
- an error without substituting Qwen. To experiment with local fallback after
286
- measuring its quality on your application, explicitly run
287
- `sys1 config set routing.policy auto`. That policy prefers reachable Jev, then
288
- the installed model named by `local.model`. Disabling Jev returns a hosted-only
289
- configuration to `auto` for the selected local model. Installing additional
290
- models does not change the selection. If no eligible route is available, Sys1
291
- reports an error instead of silently choosing another installed model.
118
+ Contradictions exit 7; unavailable evidence is marked unverifiable. A clean
119
+ report does not prove completion. File input works with any coding agent;
120
+ check-command evidence requires local Devin transcript discovery. See the
121
+ [verification guide](docs/verify.md) for input modes, evidence, and exit codes.
292
122
 
293
- ## Local models
123
+ ## Use as a module
294
124
 
295
- `sys1 setup` and `sys1 pull` manage model artifacts under
296
- `~/.sys1/models` (or `$SYS1_HOME/models`). Downloads stream to a temporary
297
- file, enforce an 8 GiB ceiling, verify SHA-256, run bounded GGUF structural
298
- validation, and only then atomically enter the model store. Manifest filenames
299
- cannot escape the store, symbolic-link weights are not admitted, and the daemon
300
- never downloads weights implicitly.
125
+ For an application using Node 24 or Bun, install the portable client:
301
126
 
302
127
  ```sh
303
- sys1 pull --list
304
- sys1 pull qwen3-1.7b
305
- sys1 model list
306
- sys1 model verify qwen3-1.7b
128
+ npm install --omit=optional \
129
+ https://github.com/hraness/sys1/releases/download/v0.17.1/hraness-sys1-0.17.1.tgz
307
130
  ```
308
131
 
309
- The curated registry contains three experimental GGUF models:
310
-
311
- | Model | Kind | Download | Role |
312
- | --- | --- | ---: | --- |
313
- | `qwen3-1.7b` | `gguf` | 1.03 GiB | default local experiment |
314
- | `qwen3-0.6b` | `gguf` | 365 MiB | diagnostic model; explicit selection only |
315
- | `qwen3.5-4b` | `gguf` | 2.55 GiB | larger candidate; explicit selection only |
316
-
317
- All entries are pinned to the publisher's Hugging Face LFS SHA-256. Weight
318
- licenses and terms remain those of their publishers; weights are not included
319
- in the Sys1 package.
320
-
321
- `sys1 pull` defaults to Qwen3 1.7B and only installs an artifact; it does not
322
- change `local.model`. To change the local route, install the model first, then
323
- select its installed ID with `sys1 config set local.model MODEL`. Other
324
- installed models require a request-level model pin. This keeps a new download
325
- from becoming an unintended fallback.
326
-
327
- Older inventories containing removed CUA-S1 or Needle artifacts fail closed.
328
- Sys1 preserves those files and the manifest; use a new `SYS1_HOME` for the
329
- current GGUF store.
330
-
331
- ### Generic GGUF adapter
332
-
333
- For builtin GGUF models, Sys1 renders a bounded question prompt, evaluates
334
- the full first-token vocabulary distribution with llama.cpp, and sums
335
- probability mass over constrained answer labels. Choice and score use unique
336
- one-character labels to avoid ambiguous multi-token option names. Builtin
337
- inference supports up to 35 options per question; hosted and external
338
- backends retain the protocol's 255-option limit. Builtin answers use the
339
- official Jev wire shapes and disclose `generic-gguf` in the
340
- `x-sys1-local-adapter` response header:
341
-
342
- - Noul returns only `type` and probability-of-yes `noul`;
343
- - Choice returns `choice`, keyed `probabilities`, and `confidence`;
344
- - Score returns a zero-based probability-weighted fractional `score`, keyed
345
- `legend`, keyed `probabilities`, and `confidence`.
346
-
347
- Adapter input bounds fail closed: Sys1 never silently truncates state,
348
- instructions, or criteria. Generic GGUF accepts up to 6,000 state characters,
349
- 2,000 instruction characters, 96 characters per option name/criterion, and
350
- 16,000 characters for the complete rendered prompt. An input beyond the adapter's
351
- bounds returns an error without being sent to a different backend.
352
-
353
- Adapter quality signals stay outside those answer objects:
354
- `x-sys1-local-min-coverage` is the least total probability mass assigned to
355
- allowed labels, and `x-sys1-local-min-concentration` is the least
356
- distribution concentration in the batch. Low coverage means the model did not
357
- cleanly follow the decision instruction. These are useful local signals, not
358
- a calibration guarantee. Use hosted Jev or a task-qualified System One-specific
359
- backend when its behavior has been evaluated for your task.
360
-
361
- An unlisted public Hugging Face GGUF can be installed explicitly:
132
+ With a backend configured, run `sys1 up` to start the gateway, then call it:
362
133
 
363
- ```sh
364
- sys1 pull 'hf:owner/repository:path/model.gguf' --sha256 <64-hex-digest>
365
- ```
366
-
367
- ## The endpoint
134
+ ```ts
135
+ import { createClient } from "@hraness/sys1/client";
368
136
 
369
- | Route | Purpose |
370
- | --- | --- |
371
- | `POST /v1/systemone` | Evaluate `{model?, state, questions}` through the selected backend |
372
- | `GET /v1/models` | List model ids, backend names, kinds, and reachability |
373
- | `GET /healthz` | Report daemon liveness and version |
374
-
375
- Responses carry `x-sys1-backend` and `x-sys1-attempts`; builtin responses
376
- also carry the local adapter and diagnostic headers above. Any HTTP response
377
- from a remote backend, including 4xx or 5xx, is definitive. Only a transport
378
- failure may re-dispatch, at most once, and never for a pinned `backend/model`.
379
-
380
- `state`, `instructions`, and criterion descriptions accept text, JSON objects,
381
- JSON arrays, or `null` where the official Jev contract permits it. The public
382
- package exports request and response schemas for boundary validation.
383
-
384
- ### Request example
385
-
386
- ```json
387
- {
388
- "model": "auto",
389
- "state": { "tests": "failing", "branch": "main" },
390
- "questions": {
391
- "action": {
392
- "type": "choice",
393
- "instructions": "What should the agent do next?",
394
- "criteria": {
395
- "fix": "Repair the failure before continuing",
396
- "continue": "The failure is unrelated and safe to defer",
397
- "escalate": "Human judgment is required"
398
- }
137
+ const sys1 = createClient(); // http://127.0.0.1:13900
138
+ const { response } = await sys1.evaluate({
139
+ model: "typesafe/jev-1.13.0",
140
+ state: "Customers cannot complete checkout after today’s release.",
141
+ questions: {
142
+ urgent: {
143
+ type: "noul",
144
+ instructions: "Does this describe an active customer-impacting incident?",
399
145
  },
400
- "risk": {
401
- "type": "score",
402
- "instructions": "Rate merge risk",
403
- "criteria": ["low", "moderate", "high"]
404
- }
405
- }
406
- }
407
- ```
408
-
409
- ## Routing
410
-
411
- For unpinned requests (`model: "auto"` or omitted), `routing.policy` controls
412
- the order of enabled hosted Jev and the installed model named by `local.model`:
413
-
414
- | Policy | Order |
415
- | --- | --- |
416
- | `auto` (default) | hosted Jev, then the selected local model |
417
- | `prefer-local` | selected local model, then hosted Jev |
418
- | `prefer-hosted` | hosted Jev, then the selected local model |
419
- | `local-only` | selected local model only |
420
- | `hosted-only` | hosted Jev only |
421
-
422
- Other installed models and all registered HTTP services require an explicit
423
- request model or backend/model pin. They never receive unpinned fallback
424
- traffic. A missing selected model does not promote another installed model.
425
- Registered HTTP services are local only when their URL uses a
426
- loopback host; off-machine URLs count as hosted. Redirects are never followed.
427
- `local-only` decisions neither probe nor dispatch to hosted endpoints. Explicit
428
- model discovery and doctor may probe all configured backends. Backend names must
429
- be unique; `typesafe` and `local-*` are reserved for managed candidates. Requests can pin either a model id or an exact backend/model:
430
-
431
- - `"model": "jev-1.13.0"` selects a backend serving that hosted model;
432
- - `"model": "local-qwen3-1.7b/qwen3-1.7b"` pins the builtin Qwen runner;
433
- - `"model": "local-qwen3-0.6b/qwen3-0.6b"` explicitly selects the experimental model;
434
- - `"model": "openjev/openjev-latest"` pins a registered HTTP backend that advertises that alias.
435
-
436
- Selection is capability-aware. Sys1 compares each request's largest option
437
- count and total question count against the backend's published limits. A backend the request exceeds is skipped; when no configured backend
438
- can serve the request at all the gateway answers `422 request_unsupported`
439
- rather than dispatching a request that would fail downstream. Builtin
440
- backends publish their adapter limit (`generic-gguf` 35 options); remote backends
441
- are probed at `GET /v1/limits` (openjev-style `max_answers_per_question` and
442
- `max_questions`). A backend that publishes nothing has unknown capacity; Sys1 can enforce only
443
- its configured limits and the common protocol envelope. Missing limits never
444
- mean zero capability.
445
-
446
- ## External System One backends
447
-
448
- Sys1 can route to an [OpenJev](https://github.com/razorback16/openjev) server
449
- through its standard System One decision API. OpenJev's image, chat, and
450
- advanced sampling extensions are not supported. A server you register answers
451
- only requests that select it; adding one does not change the default route.
452
-
453
- Any service implementing `POST /v1/systemone` and `GET /v1/models` can join the
454
- same router. For an existing OpenJev server, first inspect its advertised model
455
- IDs. The example uses OpenJev's documented `openjev-latest` alias; replace it if
456
- your server advertises a different ID:
146
+ },
147
+ }, { signal: AbortSignal.timeout(5_000) });
457
148
 
458
- ```sh
459
- curl http://127.0.0.1:8080/v1/models
460
- sys1 backend add \
461
- --name openjev \
462
- --url http://127.0.0.1:8080 \
463
- --model openjev-latest
149
+ console.log(response.answers.urgent);
464
150
  ```
465
151
 
466
- Before routing agents to an operator backend, qualify its discovery, limits,
467
- and all three answer shapes:
468
-
469
- ```sh
470
- sys1 backend check --name openjev
471
- ```
152
+ The answer contains the model’s probability of yes. Sys1 validates requests and
153
+ responses, supports cancellation, and returns stable error codes. Your
154
+ application decides which actions are allowed and evaluates the model on its
155
+ own examples. A valid answer can still be wrong.
472
156
 
473
- Select it in a request with `"model": "openjev/openjev-latest"`. Registration
474
- and qualification do not add an HTTP service to automatic routing.
157
+ The [client and embedded router guide](docs/runtime.md#use-as-a-module) covers
158
+ custom endpoints, the in-process Bun router, and existing Jev applications.
475
159
 
476
- The check makes bounded calls to `/v1/models`, `/v1/limits`, and
477
- `/v1/systemone`; validates the official response schema, probability
478
- normalization, and Score arithmetic; and never prints or persists request or
479
- response bodies. Backends that do not publish limits receive a warning unless
480
- static caps were configured. All configured HTTP processes remain
481
- operator-owned: Sys1 probes and forwards to them but does not download their
482
- weights, mutate credentials, or own their lifecycle.
160
+ ## The endpoint
483
161
 
484
- ## Kev and decision profiles
162
+ The loopback gateway serves `POST /v1/systemone` for decisions,
163
+ `GET /v1/models` for discovery, and `GET /healthz` for liveness.
164
+ [HTTP request and response reference](docs/runtime.md#the-endpoint).
485
165
 
486
- [Kev](https://github.com/jaredpalmer/kev) servers use a dedicated adapter
487
- (`--adapter kev`) that handles Kev's two-decimal probability output and
488
- structured Score legends. Versioned decision profiles reuse task instructions
489
- with Jev, local models, or a separately trained Kev checkpoint.
490
- [Kev and tuning guide](docs/kev.md).
166
+ ## Local models
491
167
 
492
- After starting and verifying a separately owned Kev server, register it explicitly:
168
+ Local Qwen models are experimental. In the broader recorded studies, Qwen3
169
+ 1.7B answered 32/72 cases correctly; Qwen3.5 4B answered 44/72 on a different
170
+ fresh fixture. Evaluate the selected model on your task before relying on it.
171
+ [Model studies](https://sys1.io/docs/evaluations).
493
172
 
494
- ```sh
495
- sys1 backend add --adapter kev --name kev \
496
- --url http://127.0.0.1:8009 --model kev-latest
497
- sys1 backend check --name kev
498
- ```
173
+ Preview the download with `sys1 setup --dry-run --json`. The [local setup guide](docs/runtime.md#experimental-local-decisions)
174
+ covers supported platforms, downloads, and model selection. Setup explicitly
175
+ downloads and selects Qwen3 1.7B; other installed models are never automatic
176
+ substitutes.
499
177
 
500
- Kev remains outside automatic routing. A versioned profile pins the route and
501
- reuses the same instructions and criteria while each call supplies new state:
178
+ ## External System One backends
502
179
 
503
- ```ts
504
- import { createClient, createProfile } from "@hraness/sys1/client";
180
+ Connect a server that implements the System One API, including OpenJev or Kev.
181
+ Register and check the server, then select it in each request. Adding a server
182
+ does not change default routing. [Compatible server setup](docs/runtime.md#external-system-one-backends)
183
+ · [Kev and decision profiles](docs/kev.md).
505
184
 
506
- const triage = createProfile({
507
- version: 1, id: "ticket-triage", revision: "1", model: "kev/kev-latest",
508
- questions: {
509
- urgent: { type: "noul", instructions: "Does the ticket describe an active outage?" },
510
- },
511
- });
512
- const result = await createClient().evaluate(triage.request("Checkout is unavailable."));
513
- ```
185
+ ## Reference and troubleshooting
514
186
 
515
- Profiles work with the Node/Bun client and embedded router. Definitions are
516
- validated and frozen; each request is a fresh ordinary System One request.
517
- Only model, state, and questions cross the wire. No templating, hidden prompt
518
- injection, global profile registry, or automatic training is involved. Save
519
- the definition as JSON for `sys1 eval --profile triage.json`, which accepts
520
- only `{"state": ...}` on stdin or `--file`.
521
- [Start from the ticket-triage profile](examples/ticket-triage.profile.json).
522
-
523
- For a direct Kev endpoint, use `createClient({ baseUrl: "http://127.0.0.1:8009",
524
- adapter: "kev" })` with an ordinary request containing `model: "kev-latest"`.
525
- When calling the Sys1 gateway, leave the client adapter unset. Routing metadata
526
- identifies Kev and its two-decimal precision; probabilities are not renormalized.
527
-
528
- The backend/model pin identifies a route, not its weights. Kev always advertises
529
- `kev-latest`; verify the server's actual checkpoint separately. The
530
- [setup and tuning guide](docs/kev.md) covers pinned checkpoints, prompt revisions,
531
- fine-tuning, calibration, and held-out evaluation. Model downloads and training
532
- remain explicit Kev operations. Protocol tests do not establish model quality.
533
-
534
- ## Diagnostics
535
-
536
- `sys1 doctor` checks the install and prints one line per check (✓, ⚠ or ✗),
537
- a count, and the command to run next. It verifies the
538
- Bun floor, state-directory access, config, native llama.cpp runtime/backend,
539
- the installed-model list, every installed model's GGUF header, leftover or
540
- unknown files in the model folder, routing candidates, and daemon ownership. It does not hash entire model
541
- files; use `sys1 model verify MODEL` for exact SHA-256 verification.
187
+ <a id="diagnostics"></a>
188
+ <a id="routing"></a>
189
+ <a id="configuration"></a>
190
+ <a id="commands"></a>
191
+ <a id="kev-and-decision-profiles"></a>
542
192
 
543
- ```sh
544
- sys1 doctor
545
- sys1 doctor --json
546
- ```
547
-
548
- Warnings do not fail readiness. Failed checks return exit code 6. JSON is
549
- versioned (`version: 1`) and check identifiers are stable and additive.
550
-
551
- ## Commands
552
-
553
- ```text
554
- sys1 setup [--tier compact|quality] [--dry-run]
555
- sys1 jev status|enable|disable
556
- sys1 up|down|serve|status|doctor
557
- sys1 pull [MODEL]|pull --list
558
- sys1 model list|verify|remove
559
- sys1 models
560
- sys1 eval
561
- sys1 audit --staged|--worktree|--since <ref> --model <backend/model>
562
- sys1 review checkpoint|issues|feedback|recheck|setup
563
- sys1 rules list|check|draft
564
- sys1 verify setup codex|claude-code|devin
565
- sys1 verify --model <backend/model> [--message <file|->]
566
- sys1 usage [--days <n>]
567
- sys1 backend list|add|check|remove
568
- sys1 config path|get|set|unset
569
- sys1 --version|--help
570
- ```
193
+ - [`sys1 doctor` and diagnostics](docs/runtime.md#diagnostics): inspect the runtime, routing, model store, and gateway after setup.
194
+ - [Routing](docs/runtime.md#routing): select a model and control fallback.
195
+ - [Configuration](docs/runtime.md#configuration): settings, credentials, and gateway access.
196
+ - [Commands](docs/runtime.md#commands): CLI reference and JSON output.
197
+ - [Decision profiles](docs/runtime.md#kev-and-decision-profiles): reuse questions across application requests.
198
+ - [Model comparison](https://sys1.io/compare) and [evaluation reports](https://sys1.io/docs/evaluations): inspect model-specific evidence.
571
199
 
572
- Supporting commands accept `--json`. When an agent runs sys1 (Claude Code,
573
- Codex, Cursor, Gemini CLI, or `AI_AGENT` is set), JSON is the default;
574
- `HRANESS_AUDIENCE=human` or `agent` overrides the guess. Machine data goes to
575
- stdout; diagnostics, next-step hints and download progress go to stderr. Every
576
- command has its own help (`sys1 setup --help`). Errors are one sentence and the
577
- command to run next; with `--json` they are one `{"ok":false,"error":{...}}`
578
- object on stdout. `sys1 setup --dry-run` shows the model, its size and the
579
- folder it would be downloaded to.
580
-
581
- ## Configuration
582
-
583
- `~/.sys1/config.json` is created on the first write. `SYS1_HOME` overrides
584
- the state directory. Settable keys:
585
-
586
- - `routing.policy`;
587
- - `gateway.host` (loopback addresses only), `gateway.port`,
588
- `gateway.request_timeout_ms`, `gateway.probe_timeout_ms`;
589
- - `hosted.base_url`, `hosted.model`, `hosted.api_key_env` (activation uses
590
- `sys1 jev`);
591
- - `local.enabled`, `local.model`, `local.context_tokens`, `local.eval_timeout_ms`,
592
- `local.max_loaded_models`.
593
-
594
- Fresh config uses `routing.policy: auto`, `local.enabled: true`,
595
- `local.model: qwen3-1.7b`, `hosted.model: jev-1.13.0`, and
596
- `hosted.enabled: false`. The daemon reads config per request, so routing and
597
- backend changes do not need a restart. Restart after changing local runtime
598
- context, timeout, or residency settings. Environment variables are inherited when
599
- the daemon starts, so restart it after exporting a new Jev credential. Already
600
- loaded GGUFs stay resident up to `local.max_loaded_models` (default one) and are
601
- released on eviction or daemon shutdown. Local requests are serialized
602
- to keep context state isolated and residency bounded. GGUF inference lives in
603
- an owned worker process; abort, timeout, or disposal terminates and collects
604
- that worker before the next request can reuse the engine slot.
605
-
606
- The decision endpoint accepts loopback binds only and has no application
607
- authentication. Network admission blocks browser-originated decision dispatch and
608
- non-loopback Host authorities, but it does not authenticate local processes.
609
- Any process that can connect locally can dispatch decisions using the gateway's
610
- enabled backends and credentials; use it only on a trusted local machine.
611
- The in-process `createRouter`/`createFetchHandler` surface leaves admission to
612
- its owning application. Daemon shutdown uses a per-instance
613
- secret from its private pid file and an authenticated control endpoint. Sys1
614
- never signals an arbitrary PID read from that file.
200
+ If `sys1` is missing after installation, check that npm’s global binary
201
+ directory is on your `PATH` and that `bun --version` works. If a request cannot
202
+ find a model, inspect `sys1 jev status` and `sys1 doctor`. Restart an existing
203
+ gateway after changing its credential environment.
615
204
 
616
205
  ## Releases
617
206
 
618
207
  An annotated `v<version>` tag at the exact current `main` head requests a
619
- release. The tag must match `package.json`. The release workflow reruns the
620
- complete gate, creates one npm-format tarball and `SHA256SUMS`, installs and
621
- executes those exact bytes with the native dependency and `doctor` on Ubuntu,
622
- macOS, and Windows, then publishes them to a repository-enforced immutable
623
- GitHub Release. No npm registry package is claimed or required.
624
-
625
- Each release page copies its summary and changes from the version's section of
626
- [`CHANGELOG.md`](CHANGELOG.md) and adds the install command, the tarball's
627
- SHA-256, and the source commit. The workflow stops before publishing when that
628
- section is missing, and it fails when a published page no longer matches the
629
- changelog and the attached files.
208
+ release and must match `package.json`. The release workflow reruns the complete
209
+ gate, creates one npm-format tarball and `SHA256SUMS`, and installs those exact
210
+ bytes with the native dependency on Ubuntu, macOS, and Windows before
211
+ publishing an immutable GitHub Release. Release notes come from
212
+ [CHANGELOG.md](CHANGELOG.md). The optional npm mirror publishes that same
213
+ tarball after the package’s one-time registry setup.
630
214
 
631
215
  ## Development
632
216
 
633
217
  ```sh
218
+ git clone https://github.com/hraness/sys1.git
219
+ cd sys1
634
220
  bun install
635
221
  bun run check
636
222
  ```
637
223
 
638
- The check runs strict TypeScript, deterministic tests with fake inference,
639
- distribution builds, and an isolated packed-artifact import/CLI smoke test.
640
- CI then runs `bun run check:native` on Linux, macOS and Windows: it installs
641
- the packed candidate in a disposable prefix and checks real native runtime
642
- readiness before a release tag is needed. This does not rebuild the package
643
- or download model weights. Model inference and hosted calls remain excluded
644
- from ordinary CI; the release workflow still verifies its exact uploaded
645
- artifact on all three platforms.
646
-
647
- ## Comparing models
648
-
649
- Use [the model comparison](https://sys1.io/compare) for external JevBench
650
- accuracy on the hosted Jev and operator-registered OpenJev routes, route
651
- availability, and a workload cost calculator. It uses the pinned
652
- [JevBench v1.2.6 snapshot](https://github.com/fstandhartinger/jevbench/tree/v1.2.6);
653
- its methodology and limits are recorded in the [evidence appendix](docs/model-comparison.md).
654
- Built-in Qwen has no matching JevBench result: SemIf's Qwen3.5 4B uses a
655
- different adapter and BF16 checkpoint, so its score does not apply to Sys1's
656
- GGUF path. An external OpenJev GPU score likewise does not establish Mac MLX
657
- quality or qualify a Sys1 deployment.
658
-
659
- [Sys1's original evaluation studies](https://sys1.io/docs/evaluations) remain
660
- available with raw reports, failure analysis, and token-accounting details.
661
- The [opt-in Sys1 benchmarks](benchmarks/README.md) support adapter qualification
662
- and reproduction; they are not the cross-model leaderboard. No local candidate
663
- is generally qualified, and the original fixtures do not justify an automatic
664
- application migration.
665
-
666
- ## Related
667
-
668
- [System One Skills](https://github.com/hraness/system-one-skills) is a separate
669
- skill for Devin, Claude Code, and Codex. It runs a known noisy test or build
670
- once, returns a short result with the exit status, and keeps the full log on
671
- disk. It needs no model, API key, or Sys1 installation, and installing either
672
- project does not configure the other. [Skills guide](https://sys1.io/skills).
673
-
674
- [The thread through hraness](https://hraness.com/writing/the-thread-through-hraness)
675
- describes the design Sys1 shares with every Hraness project: a decision has a
676
- declared shape before a model is asked, and the answer comes back validated
677
- instead of as prose.
224
+ See [CONTRIBUTING.md](CONTRIBUTING.md) for the development and release workflow.
225
+ Report a bug in [GitHub Issues](https://github.com/hraness/sys1/issues), or follow
226
+ [SECURITY.md](SECURITY.md) for a security report.