@hraness/sys1 0.17.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md ADDED
@@ -0,0 +1,677 @@
1
+ # Sys1
2
+
3
+ Sys1 helps coding agents review changes against your repository's rules, with
4
+ probability-scored answers from hosted Jev, a local model, or your own server.
5
+ Review and final-message verification are experimental and advisory.
6
+
7
+ Underneath, Sys1 lets agents ask yes/no, choice, and score questions and get
8
+ validated answers with probabilities. You choose who answers: TypeSafe's hosted
9
+ [Jev](https://docs.typesafe.ai/models), a local model on your machine, or a
10
+ compatible server you run.
11
+
12
+ Call Sys1 from a small Node/Bun client, embed the router in a Bun app, or run a
13
+ local daemon that serves the Jev-compatible `POST /v1/systemone` API.
14
+
15
+ Latest release: v0.17.0. Install it from the GitHub release with npm; it runs
16
+ on Bun 1.3.14 or newer.
17
+
18
+ [Project site](https://sys1.io) · [Agent skills](https://sys1.io/skills) · [Protocol](#the-endpoint) · [Routing](#routing)
19
+
20
+ ## Choose a workflow
21
+
22
+ | Task | Start here | What you get |
23
+ | --- | --- | --- |
24
+ | Review an agent's code changes | [`sys1 review`](docs/review.md) | Candidates tied to selected rules and diff evidence, with feedback and reuse of unchanged reviews. |
25
+ | Check a completion message | [`sys1 verify`](docs/verify.md) | A comparison of claimed commits, pushes, checks, and live changes with reachable evidence. |
26
+ | Run one review without saved history | [`sys1 audit`](docs/audit.md) | A stateless diff check with a preview, request cap, and skipped-evidence report. |
27
+ | Reuse a repository convention | [`sys1 rules`](docs/review.md#draft-a-repository-rule) | A draft you can inspect, test on examples, and activate for future reviews. |
28
+ | Add decisions to an application | [Node/Bun client](#use-as-a-module) or [HTTP API](#the-endpoint) | Yes/no, choice, and score answers with a selected model and validated response shape. |
29
+
30
+ The review and verification skills work in Git repositories with Codex,
31
+ Claude Code, or Devin; other agents can use the same CLI. Their workflows do
32
+ not depend on a specific framework or hosting provider. The bundled review
33
+ rules cover two narrow JavaScript and TypeScript changes: newly empty catch
34
+ blocks and removed test assertions. Add and evaluate your own rules for other
35
+ conventions or languages.
36
+
37
+ ## Install
38
+
39
+ Requires Bun 1.3.14 or newer. Install the release file from GitHub. Its SHA-256
40
+ is listed on the release, and release files cannot be replaced after
41
+ publishing. `--allow-scripts=node-llama-cpp` lets only the pinned native
42
+ inference package run its install script. The installed `sys1` command runs
43
+ with Bun.
44
+
45
+ ```sh
46
+ npm install --global --allow-scripts=node-llama-cpp \
47
+ https://github.com/hraness/sys1/releases/download/v0.17.0/hraness-sys1-0.17.0.tgz
48
+ sys1 doctor
49
+ ```
50
+
51
+ To build the current source instead:
52
+
53
+ ```sh
54
+ git clone https://github.com/hraness/sys1.git
55
+ cd sys1
56
+ bun install
57
+ bun run build:dist
58
+ ln -sf "$PWD/dist/cli.js" ~/.local/bin/sys1
59
+ ```
60
+
61
+ ## Review changes with your agent (experimental)
62
+
63
+ The project skill teaches an agent when to preview a review, investigate a
64
+ candidate, record feedback, and check its completion message. It installs
65
+ instructions in your repository, without activating a model or adding hooks.
66
+
67
+ First [enable hosted Jev](#add-hosted-jev) with `TYPESAFE_API_KEY` available in
68
+ your environment, or choose another configured backend. Install the skill and
69
+ preview selected files before sending source to that backend:
70
+
71
+ ```sh
72
+ sys1 review setup codex
73
+ sys1 review checkpoint --staged --model typesafe/jev-1.13.0 \
74
+ --max-requests 10 --dry-run --json -- src test
75
+ ```
76
+
77
+ Use `setup claude-code` for Claude Code or `setup devin` for Devin. Inspect the
78
+ preview, then repeat without `--dry-run` to run the checkpoint. Investigate
79
+ candidates, record feedback, and recheck the original evidence. Repeated
80
+ complete batches reuse their review for up to 24 hours. Add `--rule <id>` to
81
+ focus a checkpoint on a rule from `sys1 rules list`.
82
+
83
+ Findings are advisory, and scores are not calibrated defect probabilities. See
84
+ the [agent review guide](docs/review.md) for setup, feedback, rechecks, and rule
85
+ drafts. For a stateless check, use [`sys1 audit`](docs/audit.md). Both report
86
+ skipped evidence and incomplete coverage.
87
+
88
+ ## Check an agent's completion message
89
+
90
+ Install the standalone verification skill when you want completion checks
91
+ without the review workflow:
92
+
93
+ ```sh
94
+ sys1 verify setup codex
95
+ ```
96
+
97
+ Use `setup claude-code` or `setup devin` for those agents. The installed
98
+ instructions tell the agent to save its proposed final message and compare
99
+ its claims with the repository, linked pull requests, and live pages. Save the
100
+ draft outside the Git worktree so it does not appear as an uncommitted change.
101
+ You can also run the command directly:
102
+
103
+ ```sh
104
+ sys1 verify --message /tmp/final-message.txt --model typesafe/jev-1.13.0
105
+ ```
106
+
107
+ Use `--message -` for piped text and `--url https://example.com` to include a
108
+ live page the message does not link. Without `--message`, verify reads the
109
+ newest matching local Devin session, including that turn's check-command
110
+ results. File and stdin input work with any agent; check-command results are
111
+ available only through Devin transcript discovery. Contradicted claims exit 7;
112
+ missing evidence is unverifiable. A clean report does not prove task completion.
113
+ See [the verification guide](docs/verify.md) for evidence and limits.
114
+
115
+ ## Use as a module
116
+
117
+ For a Node 24 or Bun application that calls a running gateway, install the
118
+ release package without the optional native runtime:
119
+
120
+ ```sh
121
+ npm install --omit=optional \
122
+ https://github.com/hraness/sys1/releases/download/v0.17.0/hraness-sys1-0.17.0.tgz
123
+ ```
124
+
125
+ ```ts
126
+ import { createClient } from "@hraness/sys1/client";
127
+
128
+ const sys1 = createClient(); // http://127.0.0.1:13900
129
+ const { response, metadata } = await sys1.evaluate({
130
+ state: "The build failed after a dependency upgrade.",
131
+ questions: {
132
+ action: {
133
+ type: "choice",
134
+ criteria: { repair: "Fix the build", continue: "Continue work" },
135
+ },
136
+ },
137
+ }, { signal: AbortSignal.timeout(5_000) });
138
+
139
+ console.log(response.answers.action, metadata.backend);
140
+ ```
141
+
142
+ The client validates inputs and correlates every returned answer with its
143
+ question. It bounds response bytes, supports cancellation, and returns stable
144
+ sanitized `Sys1ClientError` codes. It never retries, reads credentials from the
145
+ environment, starts a daemon, downloads weights, or imports native inference.
146
+ Supply `baseUrl` and `headers` explicitly for another approved endpoint.
147
+ Import schemas and request/response types from the same `/client` entry point.
148
+
149
+ For a Bun application that owns routing and model lifecycle in-process:
150
+
151
+ ```ts
152
+ import { createRouter, DEFAULT_CONFIG } from "@hraness/sys1";
153
+
154
+ const router = createRouter({
155
+ config: DEFAULT_CONFIG,
156
+ env: process.env,
157
+ home: "/absolute/path/to/sys1-state", // previously installed models
158
+ });
159
+ try {
160
+ const result = await router.evaluate({
161
+ state: "All required checks passed.",
162
+ questions: { ready: { type: "noul", instructions: "Are the checks passing?" } },
163
+ });
164
+ console.log(result.response.answers.ready);
165
+ } finally {
166
+ await router.dispose();
167
+ }
168
+ ```
169
+
170
+ The embedded router opens no port. It uses the same routing and validation as
171
+ the daemon and owns its local runner until disposal. Its runtime requires Bun;
172
+ the `/client` entry point is portable to Node. Keep one router per application,
173
+ not one per request. The root package also exposes lower-level routing and
174
+ model-management APIs; applications should normally use `createClient` or
175
+ `createRouter`.
176
+
177
+ ### Adopting Sys1 in an existing Jev application
178
+
179
+ Keep domain questions, deterministic fallback, action authorization, and quality
180
+ thresholds in the application. Put endpoint configuration, transport, routing,
181
+ response validation, and local engine lifecycle behind Sys1. Existing HTTP
182
+ clients in other languages can use the same daemon without a JavaScript module.
183
+
184
+ Use `model: "auto"` or omit `model` to use the configured routing policy and
185
+ selected local model. A hardcoded `jev-1.13.0` remains a model pin and cannot
186
+ select an unrelated local model.
187
+ Local calls need no hosted API key; hosted activation stays explicit. A remote
188
+ server's loopback address points to that server. A browser running on the user's
189
+ machine can address local services, so Sys1's network listener rejects browser
190
+ origins and Fetch Metadata site headers, requires a loopback request authority,
191
+ and accepts decision POSTs only as `application/json`.
192
+
193
+ Start with an opt-in, non-authoritative pilot. Compare decisions on the
194
+ application's representative fixtures and record backend/adapter identity,
195
+ latency, errors, abstentions, and disagreement with the current decision path.
196
+ Do not reuse Jev probability thresholds for generic GGUF output
197
+ without model-specific evidence. A local-only policy also constrains explicit
198
+ pins; a pin never bypasses the policy. Broad production adoption requires the
199
+ consumer's own quality and operational acceptance, not just wire compatibility.
200
+
201
+ ## Quickstart: experimental local decisions
202
+
203
+ ```sh
204
+ sys1 setup # verifies the native runtime, installs and selects Qwen3 1.7B
205
+ sys1 up # starts the gateway on 127.0.0.1:13900
206
+ sys1 status
207
+ ```
208
+
209
+ Local Qwen is experimental. Do not treat it as a drop-in replacement for Jev.
210
+ The broader tests found 32/72 correct decisions for Qwen3 1.7B and 44/72 for
211
+ Qwen3.5 4B on a different fresh fixture. [Read the evidence](https://sys1.io/docs/evaluations)
212
+ before using local decisions to drive actions.
213
+
214
+ `setup` is the explicit weight-download boundary. It installs Qwen3 1.7B and
215
+ persists that choice as `local.model`, regardless of system memory. Inspect
216
+ without changing anything:
217
+
218
+ ```sh
219
+ sys1 setup --dry-run --json
220
+ ```
221
+
222
+ `sys1 setup --tier compact` explicitly installs and selects the experimental
223
+ Qwen3 0.6B diagnostic model. It is not an automatic low-memory fallback.
224
+ `sys1 setup --tier quality` returns the selection to Qwen3 1.7B.
225
+
226
+ The pinned llama.cpp runtime selects the best available backend automatically:
227
+
228
+ | Package target | Runtime preference |
229
+ | --- | --- |
230
+ | macOS ARM64 | Metal (the pinned runtime's only automatic selection) |
231
+ | macOS x64 | CPU |
232
+ | Linux x64 | CUDA, Vulkan, then CPU |
233
+ | Linux ARM64 | CPU |
234
+ | Windows x64 | CUDA, Vulkan, then CPU |
235
+ | Windows ARM64 | CPU |
236
+
237
+ Other platform/architecture pairs fail closed before downloading a model. The
238
+ release artifact and package smoke are exercised on Ubuntu, macOS, and Windows.
239
+ This matrix does not prove every OS/architecture pair above; run
240
+ `sys1 doctor` on the actual host before use.
241
+
242
+ Send a decision:
243
+
244
+ ```sh
245
+ sys1 eval <<'EOF'
246
+ {
247
+ "state": "Help! My payouts have been failing for 3 days.",
248
+ "questions": {
249
+ "urgent": {
250
+ "type": "noul",
251
+ "instructions": "Does this need immediate attention?",
252
+ "criteria": {
253
+ "true": "A customer-impacting incident is ongoing",
254
+ "false": "This can wait for normal triage"
255
+ }
256
+ }
257
+ }
258
+ }
259
+ EOF
260
+ ```
261
+
262
+ Or point any System One client at `http://127.0.0.1:13900`.
263
+
264
+ ## Add hosted Jev
265
+
266
+ Hosted Jev is disabled by default, even if `TYPESAFE_API_KEY` is already set in
267
+ the environment. Add it explicitly:
268
+
269
+ ```sh
270
+ export TYPESAFE_API_KEY=…
271
+ sys1 jev enable
272
+ sys1 jev status
273
+ ```
274
+
275
+ `jev enable` requires the credential to be present, stores only
276
+ `hosted.enabled: true`, and sets routing to `hosted-only`. The key remains in the
277
+ environment and is never written to disk or printed. Restart a gateway that was
278
+ started before the key was exported. To return to local-only operation:
279
+
280
+ ```sh
281
+ sys1 jev disable
282
+ ```
283
+
284
+ Enabling Jev selects `hosted-only`: a hosted outage or missing credential returns
285
+ an error without substituting Qwen. To experiment with local fallback after
286
+ measuring its quality on your application, explicitly run
287
+ `sys1 config set routing.policy auto`. That policy prefers reachable Jev, then
288
+ the installed model named by `local.model`. Disabling Jev returns a hosted-only
289
+ configuration to `auto` for the selected local model. Installing additional
290
+ models does not change the selection. If no eligible route is available, Sys1
291
+ reports an error instead of silently choosing another installed model.
292
+
293
+ ## Local models
294
+
295
+ `sys1 setup` and `sys1 pull` manage model artifacts under
296
+ `~/.sys1/models` (or `$SYS1_HOME/models`). Downloads stream to a temporary
297
+ file, enforce an 8 GiB ceiling, verify SHA-256, run bounded GGUF structural
298
+ validation, and only then atomically enter the model store. Manifest filenames
299
+ cannot escape the store, symbolic-link weights are not admitted, and the daemon
300
+ never downloads weights implicitly.
301
+
302
+ ```sh
303
+ sys1 pull --list
304
+ sys1 pull qwen3-1.7b
305
+ sys1 model list
306
+ sys1 model verify qwen3-1.7b
307
+ ```
308
+
309
+ The curated registry contains three experimental GGUF models:
310
+
311
+ | Model | Kind | Download | Role |
312
+ | --- | --- | ---: | --- |
313
+ | `qwen3-1.7b` | `gguf` | 1.03 GiB | default local experiment |
314
+ | `qwen3-0.6b` | `gguf` | 365 MiB | diagnostic model; explicit selection only |
315
+ | `qwen3.5-4b` | `gguf` | 2.55 GiB | larger candidate; explicit selection only |
316
+
317
+ All entries are pinned to the publisher's Hugging Face LFS SHA-256. Weight
318
+ licenses and terms remain those of their publishers; weights are not included
319
+ in the Sys1 package.
320
+
321
+ `sys1 pull` defaults to Qwen3 1.7B and only installs an artifact; it does not
322
+ change `local.model`. To change the local route, install the model first, then
323
+ select its installed ID with `sys1 config set local.model MODEL`. Other
324
+ installed models require a request-level model pin. This keeps a new download
325
+ from becoming an unintended fallback.
326
+
327
+ Older inventories containing removed CUA-S1 or Needle artifacts fail closed.
328
+ Sys1 preserves those files and the manifest; use a new `SYS1_HOME` for the
329
+ current GGUF store.
330
+
331
+ ### Generic GGUF adapter
332
+
333
+ For builtin GGUF models, Sys1 renders a bounded question prompt, evaluates
334
+ the full first-token vocabulary distribution with llama.cpp, and sums
335
+ probability mass over constrained answer labels. Choice and score use unique
336
+ one-character labels to avoid ambiguous multi-token option names. Builtin
337
+ inference supports up to 35 options per question; hosted and external
338
+ backends retain the protocol's 255-option limit. Builtin answers use the
339
+ official Jev wire shapes and disclose `generic-gguf` in the
340
+ `x-sys1-local-adapter` response header:
341
+
342
+ - Noul returns only `type` and probability-of-yes `noul`;
343
+ - Choice returns `choice`, keyed `probabilities`, and `confidence`;
344
+ - Score returns a zero-based probability-weighted fractional `score`, keyed
345
+ `legend`, keyed `probabilities`, and `confidence`.
346
+
347
+ Adapter input bounds fail closed: Sys1 never silently truncates state,
348
+ instructions, or criteria. Generic GGUF accepts up to 6,000 state characters,
349
+ 2,000 instruction characters, 96 characters per option name/criterion, and
350
+ 16,000 characters for the complete rendered prompt. An input beyond the adapter's
351
+ bounds returns an error without being sent to a different backend.
352
+
353
+ Adapter quality signals stay outside those answer objects:
354
+ `x-sys1-local-min-coverage` is the least total probability mass assigned to
355
+ allowed labels, and `x-sys1-local-min-concentration` is the least
356
+ distribution concentration in the batch. Low coverage means the model did not
357
+ cleanly follow the decision instruction. These are useful local signals, not
358
+ a calibration guarantee. Use hosted Jev or a task-qualified System One-specific
359
+ backend when its behavior has been evaluated for your task.
360
+
361
+ An unlisted public Hugging Face GGUF can be installed explicitly:
362
+
363
+ ```sh
364
+ sys1 pull 'hf:owner/repository:path/model.gguf' --sha256 <64-hex-digest>
365
+ ```
366
+
367
+ ## The endpoint
368
+
369
+ | Route | Purpose |
370
+ | --- | --- |
371
+ | `POST /v1/systemone` | Evaluate `{model?, state, questions}` through the selected backend |
372
+ | `GET /v1/models` | List model ids, backend names, kinds, and reachability |
373
+ | `GET /healthz` | Report daemon liveness and version |
374
+
375
+ Responses carry `x-sys1-backend` and `x-sys1-attempts`; builtin responses
376
+ also carry the local adapter and diagnostic headers above. Any HTTP response
377
+ from a remote backend, including 4xx or 5xx, is definitive. Only a transport
378
+ failure may re-dispatch, at most once, and never for a pinned `backend/model`.
379
+
380
+ `state`, `instructions`, and criterion descriptions accept text, JSON objects,
381
+ JSON arrays, or `null` where the official Jev contract permits it. The public
382
+ package exports request and response schemas for boundary validation.
383
+
384
+ ### Request example
385
+
386
+ ```json
387
+ {
388
+ "model": "auto",
389
+ "state": { "tests": "failing", "branch": "main" },
390
+ "questions": {
391
+ "action": {
392
+ "type": "choice",
393
+ "instructions": "What should the agent do next?",
394
+ "criteria": {
395
+ "fix": "Repair the failure before continuing",
396
+ "continue": "The failure is unrelated and safe to defer",
397
+ "escalate": "Human judgment is required"
398
+ }
399
+ },
400
+ "risk": {
401
+ "type": "score",
402
+ "instructions": "Rate merge risk",
403
+ "criteria": ["low", "moderate", "high"]
404
+ }
405
+ }
406
+ }
407
+ ```
408
+
409
+ ## Routing
410
+
411
+ For unpinned requests (`model: "auto"` or omitted), `routing.policy` controls
412
+ the order of enabled hosted Jev and the installed model named by `local.model`:
413
+
414
+ | Policy | Order |
415
+ | --- | --- |
416
+ | `auto` (default) | hosted Jev, then the selected local model |
417
+ | `prefer-local` | selected local model, then hosted Jev |
418
+ | `prefer-hosted` | hosted Jev, then the selected local model |
419
+ | `local-only` | selected local model only |
420
+ | `hosted-only` | hosted Jev only |
421
+
422
+ Other installed models and all registered HTTP services require an explicit
423
+ request model or backend/model pin. They never receive unpinned fallback
424
+ traffic. A missing selected model does not promote another installed model.
425
+ Registered HTTP services are local only when their URL uses a
426
+ loopback host; off-machine URLs count as hosted. Redirects are never followed.
427
+ `local-only` decisions neither probe nor dispatch to hosted endpoints. Explicit
428
+ model discovery and doctor may probe all configured backends. Backend names must
429
+ be unique; `typesafe` and `local-*` are reserved for managed candidates. Requests can pin either a model id or an exact backend/model:
430
+
431
+ - `"model": "jev-1.13.0"` selects a backend serving that hosted model;
432
+ - `"model": "local-qwen3-1.7b/qwen3-1.7b"` pins the builtin Qwen runner;
433
+ - `"model": "local-qwen3-0.6b/qwen3-0.6b"` explicitly selects the experimental model;
434
+ - `"model": "openjev/openjev-latest"` pins a registered HTTP backend that advertises that alias.
435
+
436
+ Selection is capability-aware. Sys1 compares each request's largest option
437
+ count and total question count against the backend's published limits. A backend the request exceeds is skipped; when no configured backend
438
+ can serve the request at all the gateway answers `422 request_unsupported`
439
+ rather than dispatching a request that would fail downstream. Builtin
440
+ backends publish their adapter limit (`generic-gguf` 35 options); remote backends
441
+ are probed at `GET /v1/limits` (openjev-style `max_answers_per_question` and
442
+ `max_questions`). A backend that publishes nothing has unknown capacity; Sys1 can enforce only
443
+ its configured limits and the common protocol envelope. Missing limits never
444
+ mean zero capability.
445
+
446
+ ## External System One backends
447
+
448
+ Sys1 can route to an [OpenJev](https://github.com/razorback16/openjev) server
449
+ through its standard System One decision API. OpenJev's image, chat, and
450
+ advanced sampling extensions are not supported. A server you register answers
451
+ only requests that select it; adding one does not change the default route.
452
+
453
+ Any service implementing `POST /v1/systemone` and `GET /v1/models` can join the
454
+ same router. For an existing OpenJev server, first inspect its advertised model
455
+ IDs. The example uses OpenJev's documented `openjev-latest` alias; replace it if
456
+ your server advertises a different ID:
457
+
458
+ ```sh
459
+ curl http://127.0.0.1:8080/v1/models
460
+ sys1 backend add \
461
+ --name openjev \
462
+ --url http://127.0.0.1:8080 \
463
+ --model openjev-latest
464
+ ```
465
+
466
+ Before routing agents to an operator backend, qualify its discovery, limits,
467
+ and all three answer shapes:
468
+
469
+ ```sh
470
+ sys1 backend check --name openjev
471
+ ```
472
+
473
+ Select it in a request with `"model": "openjev/openjev-latest"`. Registration
474
+ and qualification do not add an HTTP service to automatic routing.
475
+
476
+ The check makes bounded calls to `/v1/models`, `/v1/limits`, and
477
+ `/v1/systemone`; validates the official response schema, probability
478
+ normalization, and Score arithmetic; and never prints or persists request or
479
+ response bodies. Backends that do not publish limits receive a warning unless
480
+ static caps were configured. All configured HTTP processes remain
481
+ operator-owned: Sys1 probes and forwards to them but does not download their
482
+ weights, mutate credentials, or own their lifecycle.
483
+
484
+ ## Kev and decision profiles
485
+
486
+ [Kev](https://github.com/jaredpalmer/kev) servers use a dedicated adapter
487
+ (`--adapter kev`) that handles Kev's two-decimal probability output and
488
+ structured Score legends. Versioned decision profiles reuse task instructions
489
+ with Jev, local models, or a separately trained Kev checkpoint.
490
+ [Kev and tuning guide](docs/kev.md).
491
+
492
+ After starting and verifying a separately owned Kev server, register it explicitly:
493
+
494
+ ```sh
495
+ sys1 backend add --adapter kev --name kev \
496
+ --url http://127.0.0.1:8009 --model kev-latest
497
+ sys1 backend check --name kev
498
+ ```
499
+
500
+ Kev remains outside automatic routing. A versioned profile pins the route and
501
+ reuses the same instructions and criteria while each call supplies new state:
502
+
503
+ ```ts
504
+ import { createClient, createProfile } from "@hraness/sys1/client";
505
+
506
+ const triage = createProfile({
507
+ version: 1, id: "ticket-triage", revision: "1", model: "kev/kev-latest",
508
+ questions: {
509
+ urgent: { type: "noul", instructions: "Does the ticket describe an active outage?" },
510
+ },
511
+ });
512
+ const result = await createClient().evaluate(triage.request("Checkout is unavailable."));
513
+ ```
514
+
515
+ Profiles work with the Node/Bun client and embedded router. Definitions are
516
+ validated and frozen; each request is a fresh ordinary System One request.
517
+ Only model, state, and questions cross the wire. No templating, hidden prompt
518
+ injection, global profile registry, or automatic training is involved. Save
519
+ the definition as JSON for `sys1 eval --profile triage.json`, which accepts
520
+ only `{"state": ...}` on stdin or `--file`.
521
+ [Start from the ticket-triage profile](examples/ticket-triage.profile.json).
522
+
523
+ For a direct Kev endpoint, use `createClient({ baseUrl: "http://127.0.0.1:8009",
524
+ adapter: "kev" })` with an ordinary request containing `model: "kev-latest"`.
525
+ When calling the Sys1 gateway, leave the client adapter unset. Routing metadata
526
+ identifies Kev and its two-decimal precision; probabilities are not renormalized.
527
+
528
+ The backend/model pin identifies a route, not its weights. Kev always advertises
529
+ `kev-latest`; verify the server's actual checkpoint separately. The
530
+ [setup and tuning guide](docs/kev.md) covers pinned checkpoints, prompt revisions,
531
+ fine-tuning, calibration, and held-out evaluation. Model downloads and training
532
+ remain explicit Kev operations. Protocol tests do not establish model quality.
533
+
534
+ ## Diagnostics
535
+
536
+ `sys1 doctor` checks the install and prints one line per check (✓, ⚠ or ✗),
537
+ a count, and the command to run next. It verifies the
538
+ Bun floor, state-directory access, config, native llama.cpp runtime/backend,
539
+ the installed-model list, every installed model's GGUF header, leftover or
540
+ unknown files in the model folder, routing candidates, and daemon ownership. It does not hash entire model
541
+ files; use `sys1 model verify MODEL` for exact SHA-256 verification.
542
+
543
+ ```sh
544
+ sys1 doctor
545
+ sys1 doctor --json
546
+ ```
547
+
548
+ Warnings do not fail readiness. Failed checks return exit code 6. JSON is
549
+ versioned (`version: 1`) and check identifiers are stable and additive.
550
+
551
+ ## Commands
552
+
553
+ ```text
554
+ sys1 setup [--tier compact|quality] [--dry-run]
555
+ sys1 jev status|enable|disable
556
+ sys1 up|down|serve|status|doctor
557
+ sys1 pull [MODEL]|pull --list
558
+ sys1 model list|verify|remove
559
+ sys1 models
560
+ sys1 eval
561
+ sys1 audit --staged|--worktree|--since <ref> --model <backend/model>
562
+ sys1 review checkpoint|issues|feedback|recheck|setup
563
+ sys1 rules list|check|draft
564
+ sys1 verify setup codex|claude-code|devin
565
+ sys1 verify --model <backend/model> [--message <file|->]
566
+ sys1 usage [--days <n>]
567
+ sys1 backend list|add|check|remove
568
+ sys1 config path|get|set|unset
569
+ sys1 --version|--help
570
+ ```
571
+
572
+ Supporting commands accept `--json`. When an agent runs sys1 (Claude Code,
573
+ Codex, Cursor, Gemini CLI, or `AI_AGENT` is set), JSON is the default;
574
+ `HRANESS_AUDIENCE=human` or `agent` overrides the guess. Machine data goes to
575
+ stdout; diagnostics, next-step hints and download progress go to stderr. Every
576
+ command has its own help (`sys1 setup --help`). Errors are one sentence and the
577
+ command to run next; with `--json` they are one `{"ok":false,"error":{...}}`
578
+ object on stdout. `sys1 setup --dry-run` shows the model, its size and the
579
+ folder it would be downloaded to.
580
+
581
+ ## Configuration
582
+
583
+ `~/.sys1/config.json` is created on the first write. `SYS1_HOME` overrides
584
+ the state directory. Settable keys:
585
+
586
+ - `routing.policy`;
587
+ - `gateway.host` (loopback addresses only), `gateway.port`,
588
+ `gateway.request_timeout_ms`, `gateway.probe_timeout_ms`;
589
+ - `hosted.base_url`, `hosted.model`, `hosted.api_key_env` (activation uses
590
+ `sys1 jev`);
591
+ - `local.enabled`, `local.model`, `local.context_tokens`, `local.eval_timeout_ms`,
592
+ `local.max_loaded_models`.
593
+
594
+ Fresh config uses `routing.policy: auto`, `local.enabled: true`,
595
+ `local.model: qwen3-1.7b`, `hosted.model: jev-1.13.0`, and
596
+ `hosted.enabled: false`. The daemon reads config per request, so routing and
597
+ backend changes do not need a restart. Restart after changing local runtime
598
+ context, timeout, or residency settings. Environment variables are inherited when
599
+ the daemon starts, so restart it after exporting a new Jev credential. Already
600
+ loaded GGUFs stay resident up to `local.max_loaded_models` (default one) and are
601
+ released on eviction or daemon shutdown. Local requests are serialized
602
+ to keep context state isolated and residency bounded. GGUF inference lives in
603
+ an owned worker process; abort, timeout, or disposal terminates and collects
604
+ that worker before the next request can reuse the engine slot.
605
+
606
+ The decision endpoint accepts loopback binds only and has no application
607
+ authentication. Network admission blocks browser-originated decision dispatch and
608
+ non-loopback Host authorities, but it does not authenticate local processes.
609
+ Any process that can connect locally can dispatch decisions using the gateway's
610
+ enabled backends and credentials; use it only on a trusted local machine.
611
+ The in-process `createRouter`/`createFetchHandler` surface leaves admission to
612
+ its owning application. Daemon shutdown uses a per-instance
613
+ secret from its private pid file and an authenticated control endpoint. Sys1
614
+ never signals an arbitrary PID read from that file.
615
+
616
+ ## Releases
617
+
618
+ An annotated `v<version>` tag at the exact current `main` head requests a
619
+ release. The tag must match `package.json`. The release workflow reruns the
620
+ complete gate, creates one npm-format tarball and `SHA256SUMS`, installs and
621
+ executes those exact bytes with the native dependency and `doctor` on Ubuntu,
622
+ macOS, and Windows, then publishes them to a repository-enforced immutable
623
+ GitHub Release. No npm registry package is claimed or required.
624
+
625
+ Each release page copies its summary and changes from the version's section of
626
+ [`CHANGELOG.md`](CHANGELOG.md) and adds the install command, the tarball's
627
+ SHA-256, and the source commit. The workflow stops before publishing when that
628
+ section is missing, and it fails when a published page no longer matches the
629
+ changelog and the attached files.
630
+
631
+ ## Development
632
+
633
+ ```sh
634
+ bun install
635
+ bun run check
636
+ ```
637
+
638
+ The check runs strict TypeScript, deterministic tests with fake inference,
639
+ distribution builds, and an isolated packed-artifact import/CLI smoke test.
640
+ CI then runs `bun run check:native` on Linux, macOS and Windows: it installs
641
+ the packed candidate in a disposable prefix and checks real native runtime
642
+ readiness before a release tag is needed. This does not rebuild the package
643
+ or download model weights. Model inference and hosted calls remain excluded
644
+ from ordinary CI; the release workflow still verifies its exact uploaded
645
+ artifact on all three platforms.
646
+
647
+ ## Comparing models
648
+
649
+ Use [the model comparison](https://sys1.io/compare) for external JevBench
650
+ accuracy on the hosted Jev and operator-registered OpenJev routes, route
651
+ availability, and a workload cost calculator. It uses the pinned
652
+ [JevBench v1.2.6 snapshot](https://github.com/fstandhartinger/jevbench/tree/v1.2.6);
653
+ its methodology and limits are recorded in the [evidence appendix](docs/model-comparison.md).
654
+ Built-in Qwen has no matching JevBench result: SemIf's Qwen3.5 4B uses a
655
+ different adapter and BF16 checkpoint, so its score does not apply to Sys1's
656
+ GGUF path. An external OpenJev GPU score likewise does not establish Mac MLX
657
+ quality or qualify a Sys1 deployment.
658
+
659
+ [Sys1's original evaluation studies](https://sys1.io/docs/evaluations) remain
660
+ available with raw reports, failure analysis, and token-accounting details.
661
+ The [opt-in Sys1 benchmarks](benchmarks/README.md) support adapter qualification
662
+ and reproduction; they are not the cross-model leaderboard. No local candidate
663
+ is generally qualified, and the original fixtures do not justify an automatic
664
+ application migration.
665
+
666
+ ## Related
667
+
668
+ [System One Skills](https://github.com/hraness/system-one-skills) is a separate
669
+ skill for Devin, Claude Code, and Codex. It runs a known noisy test or build
670
+ once, returns a short result with the exit status, and keeps the full log on
671
+ disk. It needs no model, API key, or Sys1 installation, and installing either
672
+ project does not configure the other. [Skills guide](https://sys1.io/skills).
673
+
674
+ [The thread through hraness](https://hraness.com/writing/the-thread-through-hraness)
675
+ describes the design Sys1 shares with every Hraness project: a decision has a
676
+ declared shape before a model is asked, and the answer comes back validated
677
+ instead of as prose.