soturail 1.1.0 → 1.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +41 -8
- package/dist/cli.js +6 -0
- package/dist/cli.js.map +1 -1
- package/dist/commands/feature.d.ts +2 -0
- package/dist/commands/feature.js +28 -0
- package/dist/commands/feature.js.map +1 -0
- package/dist/commands/handoff.d.ts +2 -0
- package/dist/commands/handoff.js +14 -0
- package/dist/commands/handoff.js.map +1 -0
- package/dist/commands/harness.js +20 -0
- package/dist/commands/harness.js.map +1 -1
- package/dist/commands/session.d.ts +2 -0
- package/dist/commands/session.js +17 -0
- package/dist/commands/session.js.map +1 -0
- package/dist/core/config.d.ts +7 -0
- package/dist/core/config.js +8 -0
- package/dist/core/config.js.map +1 -1
- package/dist/core/harness-lifecycle.d.ts +63 -0
- package/dist/core/harness-lifecycle.js +397 -0
- package/dist/core/harness-lifecycle.js.map +1 -0
- package/dist/core/harness-rail.js +1 -0
- package/dist/core/harness-rail.js.map +1 -1
- package/dist/core/schema-readiness.js +9 -1
- package/dist/core/schema-readiness.js.map +1 -1
- package/dist/core/version.d.ts +1 -1
- package/dist/core/version.js +1 -1
- package/docs/agent-governance-rail.md +89 -0
- package/docs/agent-harness-synthesis-2026.md +101 -0
- package/docs/agent-hosts.md +12 -0
- package/docs/agent-qa-rail.md +92 -0
- package/docs/agents.md +13 -3
- package/docs/conductor-mode.md +44 -0
- package/docs/context-packs.md +18 -0
- package/docs/dashboard-rail.md +6 -0
- package/docs/ecosystem-influences.md +65 -0
- package/docs/eval-datasets.md +40 -0
- package/docs/evaluation-suite.md +6 -0
- package/docs/evidence-provenance-rail.md +70 -0
- package/docs/external-projects-audit.md +29 -0
- package/docs/future-rails-index.md +70 -12
- package/docs/golden-agent-tests.md +44 -0
- package/docs/governance-cost-rail.md +10 -0
- package/docs/harness-lifecycle-rail.md +95 -0
- package/docs/harness-rail.md +27 -0
- package/docs/host-compatibility-rail.md +18 -0
- package/docs/host-router-rail.md +73 -0
- package/docs/knowledge-rail.md +76 -0
- package/docs/llm-as-judge-policy.md +45 -0
- package/docs/multi-agent-workflow-templates.md +63 -0
- package/docs/observability-rail.md +36 -0
- package/docs/rate-limit-and-fallback-policy.md +45 -0
- package/docs/releases/README.md +1 -0
- package/docs/releases/RELEASE_NOTES_v1.2.0.md +43 -0
- package/docs/repo-docs-audit-2026-06-05.md +55 -0
- package/docs/report-rail.md +12 -0
- package/docs/resilience-rail.md +48 -0
- package/docs/schema-contracts.md +5 -0
- package/docs/security-boundaries.md +34 -0
- package/docs/skill-rail-2.md +11 -0
- package/docs/stable-command-surface.md +10 -1
- package/docs/tasklet-rail.md +62 -0
- package/docs/v1-contract.md +3 -2
- package/docs/workflow-rail.md +12 -0
- package/package.json +1 -1
|
@@ -19,15 +19,19 @@ SotuRail should remain:
|
|
|
19
19
|
- independent from any single agent host;
|
|
20
20
|
- small enough to use in normal developer projects.
|
|
21
21
|
|
|
22
|
-
v1.0.0 froze the first stable local Context OS surface. v1.1.0 delivered Host Compatibility Rail 1.0. The remaining post-v1 sequence
|
|
22
|
+
v1.0.0 froze the first stable local Context OS surface. v1.1.0 delivered Host Compatibility Rail 1.0. The remaining post-v1 sequence is now staged with the 2026 agent-harness synthesis:
|
|
23
23
|
|
|
24
24
|
```txt
|
|
25
|
-
v1.
|
|
26
|
-
v1.
|
|
27
|
-
v1.
|
|
28
|
-
v1.
|
|
25
|
+
v1.1.1 Host Compatibility Polish, ecosystem docs and golden export checks
|
|
26
|
+
v1.2.0 Spec, Design, Diagram and Harness Lifecycle Rail
|
|
27
|
+
v1.3.0 Knowledge, Evidence and Evaluation Rail
|
|
28
|
+
v1.4.0 Skill Rail 2.0, Knowledge-to-Skill and Tasklet Packs
|
|
29
|
+
v1.5.0 Governance, Cost, Resilience and Host Router Rail
|
|
30
|
+
v1.6.0 Agent Governance / Evolution Rail
|
|
29
31
|
```
|
|
30
32
|
|
|
33
|
+
v1.2.0 also delivers a focused Harness Lifecycle slice: safe scaffold initialization, lifecycle audit, feature state, sessions and handoffs. Spec/design/diagram expansion remains staged and must be documented honestly when not implemented.
|
|
34
|
+
|
|
31
35
|
## 2026 Agent Runtime Update
|
|
32
36
|
|
|
33
37
|
The newest planning layer is tracked in [`docs/roadmap-agent-runtime-addendum.md`](roadmap-agent-runtime-addendum.md).
|
|
@@ -55,6 +59,7 @@ It tightens how the future rails connect around current agent-runtime patterns:
|
|
|
55
59
|
| MCP Exposure Rail | Report exposed MCP tools/resources/prompts/roots and local risk notes | v0.5.0 seed | `docs/roadmap-agent-runtime-addendum.md`, `docs/mcp.md`, `docs/policy-rail.md` |
|
|
56
60
|
| Skill Boundary Rail | Route skills by task, role, evidence and host capability instead of always loading them | v0.5.0 seed | `docs/roadmap-agent-runtime-addendum.md`, `docs/skill-rail.md` |
|
|
57
61
|
| Harness Rail | setup/plan/work/review/release discipline, evidence packs and failure ledger | v0.5.0 seeds, v0.7.0 expansion | `docs/harness-rail.md`, `docs/workflow-rail.md` |
|
|
62
|
+
| Harness Lifecycle Rail | Safe local instructions, state, sessions, features, audits and handoffs | v1.2.0 implemented | `docs/harness-lifecycle-rail.md`, `docs/security-boundaries.md` |
|
|
58
63
|
| Acceptance Harness Contracts | Require build/typecheck/lint/test/coverage/docs/policy gates before accepting work | v0.5.0 seed, v0.7.0 expansion | `docs/roadmap-agent-runtime-addendum.md`, `docs/harness-rail.md` |
|
|
59
64
|
| Policy Rail | Rules, approvals, risky action queue, auth checks and MCP exposure reports | v0.5.0+ | `docs/policy-rail.md`, `docs/security-model.md`, `docs/rules.md` |
|
|
60
65
|
| Structured Payload Rail | Target-aware Markdown/JSON/tagged/TOON-table/Mermaid context formats | v0.5.1+ | `docs/structured-payload-rail.md`, `docs/context-packs.md` |
|
|
@@ -66,6 +71,13 @@ It tightens how the future rails connect around current agent-runtime patterns:
|
|
|
66
71
|
| Reverse Specification Rail | Turn existing code/docs/config/logs into claims, rules, specs, gaps and validation tasks | v0.8.0 primary | `docs/reverse-specification-rail.md`, `docs/knowledge-to-rules.md` |
|
|
67
72
|
| Project Brain | Verified claims, decisions, bugs, gaps, rules, stale events and agent-safe briefs | v0.8.0 | `docs/project-brain.md`, `ROADMAP.md` |
|
|
68
73
|
| Host Compatibility Rail | Host-aware exports, capability matrix, host doctors and conservative prompt-only fallback for OpenCode, Antigravity, Claude, Codex, Cursor, Deep Agents-style and generic hosts | v1.1.0 implemented | `docs/host-compatibility-rail.md`, `docs/agents.md`, `docs/mcp.md`, `docs/agent-export-contract.md` |
|
|
74
|
+
| Agent QA Rail | Offline datasets, golden exports, regression reports and optional LLM-as-judge policy | v1.1.1 docs, v1.3.0 planned | `docs/agent-qa-rail.md`, `docs/eval-datasets.md`, `docs/golden-agent-tests.md`, `docs/llm-as-judge-policy.md` |
|
|
75
|
+
| Evidence And Provenance Rail | Provenance sidecars, verification statuses and evidence per run/report | v1.3.0 | `docs/evidence-provenance-rail.md`, `docs/report-rail.md`, `docs/filesystem-evidence-rail.md` |
|
|
76
|
+
| Knowledge Rail | Compile docs/specs/notes into on-demand SKILL.md/topic/glossary/pattern packs | v1.3.0 docs, v1.4.0 skill integration | `docs/knowledge-rail.md`, `docs/skill-rail-2.md`, `docs/project-brain.md` |
|
|
77
|
+
| Resilience Rail | Rate-limit, fallback and provider-risk documentation/reporting without proxying traffic | v1.5.0 | `docs/resilience-rail.md`, `docs/rate-limit-and-fallback-policy.md`, `docs/governance-cost-rail.md` |
|
|
78
|
+
| Host Router Rail | Route one local context source into host-specific export formats with safe fallback | v1.5.0 | `docs/host-router-rail.md`, `docs/host-compatibility-rail.md`, `docs/context-packs.md` |
|
|
79
|
+
| Tasklet Rail | Small local reusable task templates for agents | v1.4.0 exploration | `docs/tasklet-rail.md`, `docs/skill-rail-2.md`, `docs/workflow-rail.md` |
|
|
80
|
+
| Agent Governance/Evolution Rail | Trace, ledger, experiments, approval gates and improve/eval/apply loop | v1.6.0 | `docs/agent-governance-rail.md`, `docs/evaluation-suite.md`, `docs/policy-rail.md` |
|
|
69
81
|
| Spec Rail | PRD, requirements, design, tasks and acceptance criteria as workflow inputs | v1.2.0 | `docs/spec-driven-workflow.md`, `docs/design-rail.md`, `docs/diagram-rail.md` |
|
|
70
82
|
| Design Rail | Local `DESIGN.md`, design token lint/diff/export and agent-readable visual guidance | v1.2.0 | `docs/design-rail.md`, `docs/dashboard-rail.md` |
|
|
71
83
|
| Knowledge Graph Rail | Local graph of files, claims, decisions, tests, workflows, diagrams and releases | v1.3.0 | `docs/knowledge-graph-rail.md`, `docs/code-graph.md`, `docs/project-brain.md` |
|
|
@@ -73,6 +85,7 @@ It tightens how the future rails connect around current agent-runtime patterns:
|
|
|
73
85
|
| Governance And Cost Rail | Context budget, dynamic workflow risk and MCP/skill exposure warnings | v1.5.0 | `docs/governance-cost-rail.md`, `docs/report-rail.md`, `docs/policy-rail.md` |
|
|
74
86
|
| Redacted Evidence Sharing | Local redacted evidence bundles for reports/logs without hosting by default | v1.4+ exploration | `docs/report-redaction.md`, `docs/external-projects-audit.md` |
|
|
75
87
|
| Local Dashboard | Local reports, trace viewer, Mermaid rendering and policy/evidence views | v0.10.0 | `ROADMAP.md` |
|
|
88
|
+
| SotuRail Conductor | Optional future planner/verifier/reviewer/tasklet coordinator behind approval gates | Future proposed | `docs/conductor-mode.md`, `docs/security-boundaries.md` |
|
|
76
89
|
|
|
77
90
|
## Version Summary
|
|
78
91
|
|
|
@@ -219,10 +232,21 @@ Focus:
|
|
|
219
232
|
- per-host doctor reports;
|
|
220
233
|
- read-only MCP host manifests.
|
|
221
234
|
|
|
235
|
+
### v1.1.1
|
|
236
|
+
|
|
237
|
+
Focus:
|
|
238
|
+
|
|
239
|
+
- 2026 agent-harness synthesis docs;
|
|
240
|
+
- Agent QA, golden export and optional judge policy docs;
|
|
241
|
+
- Evidence/Provenance, Knowledge, Resilience, Host Router, Tasklet and Agent Governance planning docs;
|
|
242
|
+
- validate -> fix -> verify -> report workflow example;
|
|
243
|
+
- multi-agent researcher/analyst/writer/verifier template example.
|
|
244
|
+
|
|
222
245
|
### v1.2.0
|
|
223
246
|
|
|
224
247
|
Focus:
|
|
225
248
|
|
|
249
|
+
- Harness Lifecycle Rail implemented with `harness init`, `harness audit`, sessions, handoffs and feature state;
|
|
226
250
|
- Spec, Design And Diagram Rail;
|
|
227
251
|
- PRD/requirements/design/tasks scaffolds;
|
|
228
252
|
- local `DESIGN.md` lint/diff/export;
|
|
@@ -232,29 +256,47 @@ Focus:
|
|
|
232
256
|
|
|
233
257
|
Focus:
|
|
234
258
|
|
|
235
|
-
- Knowledge
|
|
259
|
+
- Knowledge, Evidence and Evaluation Rail;
|
|
236
260
|
- graph build/explain/impact/tour;
|
|
237
|
-
-
|
|
238
|
-
-
|
|
261
|
+
- Knowledge Rail ingest/estimate/compile/update/verify;
|
|
262
|
+
- evidence and provenance sidecars;
|
|
263
|
+
- local eval datasets, golden checks and regression reports;
|
|
264
|
+
- optional LLM-as-judge outside default release gates;
|
|
265
|
+
- stale-edge and missing-evidence detection.
|
|
239
266
|
|
|
240
267
|
### v1.4.0
|
|
241
268
|
|
|
242
269
|
Focus:
|
|
243
270
|
|
|
244
|
-
- Skill Rail 2.0;
|
|
271
|
+
- Skill Rail 2.0, Knowledge-to-Skill and Tasklet Packs;
|
|
245
272
|
- domain skill templates;
|
|
246
273
|
- skill lint/eval/report;
|
|
247
|
-
- role-aware exports and safety gates
|
|
274
|
+
- role-aware exports and safety gates;
|
|
275
|
+
- `skill build` and `skill fold-in` planning;
|
|
276
|
+
- small reusable tasklet templates.
|
|
248
277
|
|
|
249
278
|
### v1.5.0
|
|
250
279
|
|
|
251
280
|
Focus:
|
|
252
281
|
|
|
253
|
-
- Governance
|
|
254
|
-
- context budget reports;
|
|
282
|
+
- Governance, Cost, Resilience and Host Router Rail;
|
|
283
|
+
- context budget reports and compact context variants;
|
|
284
|
+
- rate-limit/fallback policy reports;
|
|
285
|
+
- host-format fallback without provider proxying;
|
|
255
286
|
- dynamic workflow guardrails;
|
|
256
287
|
- MCP/skill exposure risk summaries.
|
|
257
288
|
|
|
289
|
+
### v1.6.0
|
|
290
|
+
|
|
291
|
+
Focus:
|
|
292
|
+
|
|
293
|
+
- Agent Governance / Evolution Rail;
|
|
294
|
+
- trace and ledger artifacts;
|
|
295
|
+
- agent boundary policy;
|
|
296
|
+
- experiments/candidates/results;
|
|
297
|
+
- propose -> eval -> approve -> apply loop;
|
|
298
|
+
- approval gates before any patch application.
|
|
299
|
+
|
|
258
300
|
|
|
259
301
|
## What Should Not Happen
|
|
260
302
|
|
|
@@ -298,3 +340,19 @@ SotuRail should not become:
|
|
|
298
340
|
- `docs/ecosystem-influences.md`
|
|
299
341
|
- `docs/comparisons.md`
|
|
300
342
|
- `docs/roadmap-harness-diagram-payload-addendum.md`
|
|
343
|
+
- `docs/agent-harness-synthesis-2026.md`
|
|
344
|
+
- `docs/agent-qa-rail.md`
|
|
345
|
+
- `docs/eval-datasets.md`
|
|
346
|
+
- `docs/golden-agent-tests.md`
|
|
347
|
+
- `docs/llm-as-judge-policy.md`
|
|
348
|
+
- `docs/evidence-provenance-rail.md`
|
|
349
|
+
- `docs/agent-governance-rail.md`
|
|
350
|
+
- `docs/harness-lifecycle-rail.md`
|
|
351
|
+
- `docs/knowledge-rail.md`
|
|
352
|
+
- `docs/resilience-rail.md`
|
|
353
|
+
- `docs/rate-limit-and-fallback-policy.md`
|
|
354
|
+
- `docs/multi-agent-workflow-templates.md`
|
|
355
|
+
- `docs/host-router-rail.md`
|
|
356
|
+
- `docs/tasklet-rail.md`
|
|
357
|
+
- `docs/security-boundaries.md`
|
|
358
|
+
- `docs/conductor-mode.md`
|
|
@@ -0,0 +1,44 @@
|
|
|
1
|
+
# Golden Agent Tests
|
|
2
|
+
|
|
3
|
+
Golden agent tests are snapshot-like checks for agent-facing outputs. They are not tests of a real LLM. They protect SotuRail exports, reports and context packs from accidental regressions.
|
|
4
|
+
|
|
5
|
+
## Golden Targets
|
|
6
|
+
|
|
7
|
+
Planned golden targets:
|
|
8
|
+
|
|
9
|
+
- `AGENTS.md` guidance generated by `harness init`;
|
|
10
|
+
- Claude, Codex, Cursor, OpenCode, Gemini and generic exports;
|
|
11
|
+
- host matrix JSON;
|
|
12
|
+
- MCP host manifest JSON;
|
|
13
|
+
- report and dashboard summaries;
|
|
14
|
+
- skill export bundles;
|
|
15
|
+
- tasklet definitions;
|
|
16
|
+
- evidence/provenance sidecars.
|
|
17
|
+
|
|
18
|
+
## What A Golden Test Should Check
|
|
19
|
+
|
|
20
|
+
Golden tests should favor semantic assertions over brittle full-file string matches:
|
|
21
|
+
|
|
22
|
+
```txt
|
|
23
|
+
must include: local-first boundary
|
|
24
|
+
must include: safe next commands
|
|
25
|
+
must include: read-only MCP warning
|
|
26
|
+
must include: evidence paths when available
|
|
27
|
+
must not include: secrets
|
|
28
|
+
must not include: destructive tool promises
|
|
29
|
+
must not include: unrelated project identity
|
|
30
|
+
must parse: JSON artifacts
|
|
31
|
+
must keep: schemaVersion
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
## Update Policy
|
|
35
|
+
|
|
36
|
+
Golden outputs should be updated deliberately:
|
|
37
|
+
|
|
38
|
+
```bash
|
|
39
|
+
soturail eval golden --update
|
|
40
|
+
soturail eval report
|
|
41
|
+
soturail release check --strict --with-evals
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
A golden change should be explained in release evidence. If the change affects host behavior, update host docs too.
|
|
@@ -79,3 +79,13 @@ soturail brain stale --repair-plan
|
|
|
79
79
|
| MCP Report Resources | exposure inventory |
|
|
80
80
|
| Workflow Rail | phase depth and evidence completeness |
|
|
81
81
|
| Benchmark Rail | performance evidence and stale report warnings |
|
|
82
|
+
|
|
83
|
+
## Resilience And Host Router Expansion
|
|
84
|
+
|
|
85
|
+
Future v1.5 planning now includes Resilience Rail and Host Router Rail:
|
|
86
|
+
|
|
87
|
+
- [`resilience-rail.md`](resilience-rail.md) for retry/fallback/rate-limit documentation and risk reports;
|
|
88
|
+
- [`rate-limit-and-fallback-policy.md`](rate-limit-and-fallback-policy.md) for local policy shape;
|
|
89
|
+
- [`host-router-rail.md`](host-router-rail.md) for context-format routing across hosts.
|
|
90
|
+
|
|
91
|
+
This does not turn SotuRail into a model proxy, account manager, MITM bridge or billing gateway.
|
|
@@ -0,0 +1,95 @@
|
|
|
1
|
+
# Harness Lifecycle Rail
|
|
2
|
+
|
|
3
|
+
Harness Lifecycle Rail is the implemented v1.2.0 local state layer for carrying agent-assisted work across sessions and hosts without turning SotuRail into an autonomous agent runtime.
|
|
4
|
+
|
|
5
|
+
It organizes five subsystems:
|
|
6
|
+
|
|
7
|
+
```txt
|
|
8
|
+
Instructions
|
|
9
|
+
State
|
|
10
|
+
Verification
|
|
11
|
+
Scope
|
|
12
|
+
Session Lifecycle
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
The purpose is to keep coding agents in scope across sessions and prevent premature "done" states.
|
|
16
|
+
|
|
17
|
+
## Implemented Commands
|
|
18
|
+
|
|
19
|
+
```bash
|
|
20
|
+
soturail harness init
|
|
21
|
+
soturail harness audit
|
|
22
|
+
soturail harness audit --json
|
|
23
|
+
soturail session start "objective"
|
|
24
|
+
soturail session end --summary "completed work"
|
|
25
|
+
soturail handoff generate
|
|
26
|
+
soturail feature add "feature"
|
|
27
|
+
soturail feature start <id>
|
|
28
|
+
soturail feature done <id> --evidence tests/example.test.ts
|
|
29
|
+
soturail feature list
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
`harness init` creates local files without overwriting existing content:
|
|
33
|
+
|
|
34
|
+
```txt
|
|
35
|
+
.soturail/harness/AGENTS.md
|
|
36
|
+
.soturail/harness/instructions.md
|
|
37
|
+
.soturail/harness/verification.md
|
|
38
|
+
.soturail/harness/scope.md
|
|
39
|
+
.soturail/harness/lifecycle.md
|
|
40
|
+
.soturail/state/feature_list.json
|
|
41
|
+
.soturail/state/progress.md
|
|
42
|
+
.soturail/state/session-handoff.md
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
Use `--force` only after reviewing the scaffold files that will be replaced.
|
|
46
|
+
|
|
47
|
+
## Audit Model
|
|
48
|
+
|
|
49
|
+
`harness audit` checks Instructions, State, Verification, Scope, Session Lifecycle, Host Compatibility, Evidence/Reports and Security Boundaries. It writes:
|
|
50
|
+
|
|
51
|
+
```txt
|
|
52
|
+
.soturail/harness/audit.json
|
|
53
|
+
.soturail/harness/audit.md
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
The audit is deterministic and read-only. It does not run commands from `verification.md`.
|
|
57
|
+
|
|
58
|
+
## Feature And Session State
|
|
59
|
+
|
|
60
|
+
Feature lists prevent agents from doing many tasks and finishing none. `feature_list.json` has the stable local shape:
|
|
61
|
+
|
|
62
|
+
```json
|
|
63
|
+
{
|
|
64
|
+
"schemaVersion": "soturail.feature-list.v1",
|
|
65
|
+
"active": null,
|
|
66
|
+
"features": []
|
|
67
|
+
}
|
|
68
|
+
```
|
|
69
|
+
|
|
70
|
+
Only one feature can be active. Completion evidence is recorded as paths or local references supplied by the user.
|
|
71
|
+
|
|
72
|
+
`handoff generate` writes a bounded handoff with the current objective, completed work, changed-file names, verification status, blockers and next steps. It does not read private shell history or run verification automatically.
|
|
73
|
+
|
|
74
|
+
## Progressive Disclosure
|
|
75
|
+
|
|
76
|
+
Root agent docs should stay short and point to specialized docs.
|
|
77
|
+
|
|
78
|
+
```txt
|
|
79
|
+
Read first: .soturail/harness/AGENTS.md
|
|
80
|
+
If editing tests: .soturail/harness/verification.md
|
|
81
|
+
If continuing work: .soturail/state/session-handoff.md
|
|
82
|
+
If preparing release: docs/release-workflow.md
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
## Planned Extensions
|
|
86
|
+
|
|
87
|
+
Commands such as `harness benchmark`, `session verify` and a dedicated `session handoff` alias remain proposed. They are not implemented in v1.2.0.
|
|
88
|
+
|
|
89
|
+
## Boundaries
|
|
90
|
+
|
|
91
|
+
- SotuRail prepares lifecycle state and evidence; the agent host still reasons and edits.
|
|
92
|
+
- No cloud telemetry, mandatory server or LLM API key is required.
|
|
93
|
+
- No destructive MCP tool or arbitrary MCP shell execution is added.
|
|
94
|
+
- No Claude-only harness, giant root instruction file or acceptance without evidence.
|
|
95
|
+
- See [Security Boundaries](security-boundaries.md), [Harness Rail](harness-rail.md), [Conductor Mode](conductor-mode.md) and [Agent Harness Synthesis](agent-harness-synthesis-2026.md).
|
package/docs/harness-rail.md
CHANGED
|
@@ -15,6 +15,11 @@ Harness Rail focuses on workflow discipline, policy decisions, review evidence,
|
|
|
15
15
|
|
|
16
16
|
```bash
|
|
17
17
|
soturail harness note "agent said done before tests passed"
|
|
18
|
+
soturail harness init
|
|
19
|
+
soturail harness audit
|
|
20
|
+
soturail session start "objective"
|
|
21
|
+
soturail handoff generate
|
|
22
|
+
soturail feature list
|
|
18
23
|
soturail harness list
|
|
19
24
|
soturail harness explain <id>
|
|
20
25
|
soturail harness doctor
|
|
@@ -25,6 +30,12 @@ soturail workflow evidence <id>
|
|
|
25
30
|
|
|
26
31
|
`harness contract check` validates the local contract. It does not run build, typecheck or test commands by default.
|
|
27
32
|
|
|
33
|
+
## v1.2.0 Harness Lifecycle
|
|
34
|
+
|
|
35
|
+
Harness Lifecycle Rail adds safe project-local state for instructions, scope, verification, sessions, handoffs and features. `harness init` preserves existing files by default, and `harness audit` scores lifecycle readiness without executing verification commands.
|
|
36
|
+
|
|
37
|
+
See [Harness Lifecycle Rail](harness-lifecycle-rail.md) and [Security Boundaries](security-boundaries.md).
|
|
38
|
+
|
|
28
39
|
## v0.7.0 Workflow Connection
|
|
29
40
|
|
|
30
41
|
Harness Rail now connects to Workflow Rail 2.0:
|
|
@@ -139,3 +150,19 @@ Harness Rail is successful only if it makes SotuRail more predictable without ma
|
|
|
139
150
|
- tests cover clean workflow fixtures;
|
|
140
151
|
- reports distinguish token savings from quality preservation;
|
|
141
152
|
- no arbitrary shell execution is exposed through MCP by default.
|
|
153
|
+
|
|
154
|
+
## Harness Lifecycle Expansion
|
|
155
|
+
|
|
156
|
+
Future harness lifecycle work is tracked in [`harness-lifecycle-rail.md`](harness-lifecycle-rail.md).
|
|
157
|
+
|
|
158
|
+
The expansion adds five explicit subsystems:
|
|
159
|
+
|
|
160
|
+
```txt
|
|
161
|
+
Instructions
|
|
162
|
+
State
|
|
163
|
+
Verification
|
|
164
|
+
Scope
|
|
165
|
+
Session Lifecycle
|
|
166
|
+
```
|
|
167
|
+
|
|
168
|
+
The goal is to make `harness init`, `harness audit`, `session start/end`, feature lists and handoff files first-class local artifacts without becoming a Claude-only plugin or autonomous runtime.
|
|
@@ -2,6 +2,12 @@
|
|
|
2
2
|
|
|
3
3
|
Host Compatibility Rail is the v1.1.0 direction for making SotuRail outputs useful across coding-agent hosts without becoming an agent host itself.
|
|
4
4
|
|
|
5
|
+
## Host Fit
|
|
6
|
+
|
|
7
|
+
`soturail agents matrix`, `agents doctor --host <host>` and `agents doctor --all` form the current Host Fit Doctor surface. They report compatibility, export support, limitations and safe next commands without model serving or runtime installation.
|
|
8
|
+
|
|
9
|
+
Hermes and Odysseus are ecosystem influences, not direct stable host integrations. Hermes is an agent runtime and Odysseus is a broader workspace/runtime stack. Generic-compatible exports remain the safe fallback until a host-specific path is implemented and tested.
|
|
10
|
+
|
|
5
11
|
## Goal
|
|
6
12
|
|
|
7
13
|
```txt
|
|
@@ -76,3 +82,15 @@ Each host entry should eventually report:
|
|
|
76
82
|
| Policy Rail | config write, hook and tool exposure checks |
|
|
77
83
|
| Skill Rail | host-aware skill export without always loading every skill |
|
|
78
84
|
| Workflow Rail | phase-specific host handoff |
|
|
85
|
+
|
|
86
|
+
## Host Router Expansion
|
|
87
|
+
|
|
88
|
+
Future host work is tracked in [`host-router-rail.md`](host-router-rail.md).
|
|
89
|
+
|
|
90
|
+
The host-router idea means context-format routing, not model-request routing:
|
|
91
|
+
|
|
92
|
+
```txt
|
|
93
|
+
one SotuRail context source -> Claude/Codex/Cursor/OpenCode/Gemini/generic exports
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
SotuRail must not intercept IDE traffic, proxy provider requests, manage browser tokens or bypass quotas.
|
|
@@ -0,0 +1,73 @@
|
|
|
1
|
+
# Host Router Rail
|
|
2
|
+
|
|
3
|
+
Host Router Rail is a planned host/export layer inspired by router products, but it routes context formats, not model traffic.
|
|
4
|
+
|
|
5
|
+
```txt
|
|
6
|
+
One SotuRail context source -> many host-specific exports.
|
|
7
|
+
```
|
|
8
|
+
|
|
9
|
+
## Goal
|
|
10
|
+
|
|
11
|
+
Generate the best safe local artifact for each host:
|
|
12
|
+
|
|
13
|
+
| Host | Example export |
|
|
14
|
+
| --- | --- |
|
|
15
|
+
| Claude Code | `CLAUDE.md`, skills guidance, MCP read-only manifest |
|
|
16
|
+
| Codex | `AGENTS.md`, instructions and safe command handoff |
|
|
17
|
+
| Cursor | `.cursor/rules` style prompt/rule files |
|
|
18
|
+
| OpenCode | OpenCode-compatible instructions and role packs |
|
|
19
|
+
| Gemini | `GEMINI.md` style guidance |
|
|
20
|
+
| generic | Markdown context pack and report bundle |
|
|
21
|
+
|
|
22
|
+
## Proposed Commands
|
|
23
|
+
|
|
24
|
+
```bash
|
|
25
|
+
soturail export all
|
|
26
|
+
soturail translate --to claude
|
|
27
|
+
soturail translate --to codex
|
|
28
|
+
soturail translate --to cursor
|
|
29
|
+
soturail hosts status
|
|
30
|
+
soturail doctor --hosts
|
|
31
|
+
soturail export --fallback
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
## Context Optimization
|
|
35
|
+
|
|
36
|
+
Host Router Rail should connect to context budgeting:
|
|
37
|
+
|
|
38
|
+
```bash
|
|
39
|
+
soturail context optimize
|
|
40
|
+
soturail context budget
|
|
41
|
+
soturail context compact
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
Potential outputs:
|
|
45
|
+
|
|
46
|
+
```txt
|
|
47
|
+
.soturail/context/full.md
|
|
48
|
+
.soturail/context/compact.md
|
|
49
|
+
.soturail/context/ultra.md
|
|
50
|
+
.soturail/exports/claude/
|
|
51
|
+
.soturail/exports/codex/
|
|
52
|
+
.soturail/exports/cursor/
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
## Fallbacks
|
|
56
|
+
|
|
57
|
+
Fallback should mean safe format fallback, not provider or quota bypass:
|
|
58
|
+
|
|
59
|
+
- MCP unsupported -> static Markdown/resources;
|
|
60
|
+
- skills unsupported -> prompt-only instructions;
|
|
61
|
+
- huge context -> compact role pack;
|
|
62
|
+
- host unknown -> generic export.
|
|
63
|
+
|
|
64
|
+
## Hard Boundary
|
|
65
|
+
|
|
66
|
+
SotuRail must not:
|
|
67
|
+
|
|
68
|
+
- intercept IDE traffic;
|
|
69
|
+
- proxy model requests;
|
|
70
|
+
- manage or rotate provider accounts;
|
|
71
|
+
- use browser cookies or session tokens;
|
|
72
|
+
- bypass quotas;
|
|
73
|
+
- promise unlimited free model access.
|
|
@@ -0,0 +1,76 @@
|
|
|
1
|
+
# Knowledge Rail
|
|
2
|
+
|
|
3
|
+
Knowledge Rail is a planned future surface for compiling project documentation into small, on-demand knowledge packs for coding agents.
|
|
4
|
+
|
|
5
|
+
It is inspired by document-to-skill workflows: extract structure, not a giant summary.
|
|
6
|
+
|
|
7
|
+
## Goal
|
|
8
|
+
|
|
9
|
+
Turn docs, specs, READMEs, ADRs, notes and technical references into local agent-usable knowledge:
|
|
10
|
+
|
|
11
|
+
```txt
|
|
12
|
+
source docs -> topic index -> SKILL.md -> topics -> glossary -> patterns -> cheatsheet -> provenance
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
## Proposed Commands
|
|
16
|
+
|
|
17
|
+
```bash
|
|
18
|
+
soturail knowledge ingest ./docs ./README.md ./architecture.md
|
|
19
|
+
soturail knowledge estimate ./docs
|
|
20
|
+
soturail knowledge compile --mode on-demand
|
|
21
|
+
soturail knowledge update project-brain ./new-docs
|
|
22
|
+
soturail knowledge verify
|
|
23
|
+
soturail skill build ./docs --name project-architecture
|
|
24
|
+
soturail skill fold-in project-architecture ./ADR-004.md
|
|
25
|
+
soturail skill export --target claude
|
|
26
|
+
```
|
|
27
|
+
|
|
28
|
+
## Local Layout
|
|
29
|
+
|
|
30
|
+
```txt
|
|
31
|
+
.soturail/knowledge/<name>/
|
|
32
|
+
SKILL.md
|
|
33
|
+
topics/
|
|
34
|
+
architecture.md
|
|
35
|
+
testing.md
|
|
36
|
+
release.md
|
|
37
|
+
glossary.md
|
|
38
|
+
patterns.md
|
|
39
|
+
cheatsheet.md
|
|
40
|
+
source-map.json
|
|
41
|
+
metadata.json
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
## On-Demand Context
|
|
45
|
+
|
|
46
|
+
Instead of loading all docs into every prompt, an export can say:
|
|
47
|
+
|
|
48
|
+
```txt
|
|
49
|
+
Read .soturail/knowledge/project/SKILL.md first.
|
|
50
|
+
For release tasks, read topics/release.md.
|
|
51
|
+
For testing tasks, read topics/testing.md.
|
|
52
|
+
Do not load unrelated topics unless needed.
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
## Metadata
|
|
56
|
+
|
|
57
|
+
A knowledge pack should record:
|
|
58
|
+
|
|
59
|
+
- source paths;
|
|
60
|
+
- extraction date;
|
|
61
|
+
- token estimate;
|
|
62
|
+
- topic list;
|
|
63
|
+
- source map;
|
|
64
|
+
- verification status;
|
|
65
|
+
- update/fold-in history.
|
|
66
|
+
|
|
67
|
+
## Copyright And Sharing Boundary
|
|
68
|
+
|
|
69
|
+
Knowledge packs generated from private or copyrighted material should stay local unless the user has rights to redistribute them. SotuRail should prefer synthesized structure and references over copying large source text.
|
|
70
|
+
|
|
71
|
+
## Non-Goals
|
|
72
|
+
|
|
73
|
+
- no vector database required;
|
|
74
|
+
- no cloud embeddings required;
|
|
75
|
+
- no publishing third-party book content by default;
|
|
76
|
+
- no replacing Project Brain or Knowledge Graph Rail. Knowledge Rail feeds them with organized source material.
|
|
@@ -0,0 +1,45 @@
|
|
|
1
|
+
# LLM-As-Judge Policy
|
|
2
|
+
|
|
3
|
+
SotuRail may eventually support optional LLM-as-judge evaluations for hallucination, clarity and evidence quality. This document defines the boundary before that feature exists.
|
|
4
|
+
|
|
5
|
+
## Default Policy
|
|
6
|
+
|
|
7
|
+
```txt
|
|
8
|
+
No LLM-as-judge call is required for normal SotuRail development, testing or release.
|
|
9
|
+
```
|
|
10
|
+
|
|
11
|
+
Default gates should be deterministic:
|
|
12
|
+
|
|
13
|
+
- schema validation;
|
|
14
|
+
- JSON parsing;
|
|
15
|
+
- keyword and section checks;
|
|
16
|
+
- file existence checks;
|
|
17
|
+
- redaction checks;
|
|
18
|
+
- golden output checks;
|
|
19
|
+
- snapshot diff checks;
|
|
20
|
+
- link/path checks where local.
|
|
21
|
+
|
|
22
|
+
## Optional Judge Mode
|
|
23
|
+
|
|
24
|
+
A future optional command could look like:
|
|
25
|
+
|
|
26
|
+
```bash
|
|
27
|
+
soturail eval judge --provider openai
|
|
28
|
+
soturail eval judge --provider groq
|
|
29
|
+
soturail eval judge --provider local
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
Required behavior:
|
|
33
|
+
|
|
34
|
+
- opt-in only;
|
|
35
|
+
- never uploads repo data without clear user action;
|
|
36
|
+
- stores provider, model, prompt hash, timestamp and result;
|
|
37
|
+
- separates judge results from deterministic release evidence;
|
|
38
|
+
- marks results as subjective/heuristic, not proof.
|
|
39
|
+
|
|
40
|
+
## Non-Goals
|
|
41
|
+
|
|
42
|
+
- no mandatory provider keys;
|
|
43
|
+
- no hidden telemetry;
|
|
44
|
+
- no judge score as the only release gate;
|
|
45
|
+
- no claim that a judge result proves correctness.
|
|
@@ -0,0 +1,63 @@
|
|
|
1
|
+
# Multi-Agent Workflow Templates
|
|
2
|
+
|
|
3
|
+
SotuRail should support multi-agent ideas as templates and context packs, not as a required CrewAI/LangGraph runtime.
|
|
4
|
+
|
|
5
|
+
## Purpose
|
|
6
|
+
|
|
7
|
+
A multi-agent workflow template defines roles, inputs, outputs, evidence and verification criteria for external agents.
|
|
8
|
+
|
|
9
|
+
SotuRail can generate:
|
|
10
|
+
|
|
11
|
+
- role-specific context packs;
|
|
12
|
+
- role boundaries;
|
|
13
|
+
- definition of done per role;
|
|
14
|
+
- handoff files between roles;
|
|
15
|
+
- evidence requirements;
|
|
16
|
+
- a final report skeleton.
|
|
17
|
+
|
|
18
|
+
## Example Roles
|
|
19
|
+
|
|
20
|
+
```txt
|
|
21
|
+
Lead Agent
|
|
22
|
+
├── Researcher: collects local evidence and docs
|
|
23
|
+
├── Analyst: compares options and risks
|
|
24
|
+
├── Writer: produces user-facing docs/reports
|
|
25
|
+
└── Verifier: checks claims, tests, links and provenance
|
|
26
|
+
```
|
|
27
|
+
|
|
28
|
+
## Proposed Commands
|
|
29
|
+
|
|
30
|
+
```bash
|
|
31
|
+
soturail workflow template research-report
|
|
32
|
+
soturail workflow template validate-fix-verify-report
|
|
33
|
+
soturail context pack --role researcher
|
|
34
|
+
soturail context pack --role analyst
|
|
35
|
+
soturail context pack --role verifier
|
|
36
|
+
soturail agents export --role reviewer
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
## Template: Validate -> Fix -> Verify -> Report
|
|
40
|
+
|
|
41
|
+
Inspired by agent pipelines that validate input, repair issues, revalidate and then produce a report.
|
|
42
|
+
|
|
43
|
+
```txt
|
|
44
|
+
validate
|
|
45
|
+
input: selected files, schema, policy
|
|
46
|
+
output: findings.json
|
|
47
|
+
fix
|
|
48
|
+
input: findings.json and allowed files
|
|
49
|
+
output: patch proposal or instructions
|
|
50
|
+
verify
|
|
51
|
+
input: patch proposal and test/report commands
|
|
52
|
+
output: verification.md
|
|
53
|
+
report
|
|
54
|
+
input: all prior artifacts
|
|
55
|
+
output: report.md + provenance.md
|
|
56
|
+
```
|
|
57
|
+
|
|
58
|
+
## Non-Goals
|
|
59
|
+
|
|
60
|
+
- no mandatory CrewAI, LangGraph or LangChain dependency;
|
|
61
|
+
- no hidden subagent execution;
|
|
62
|
+
- no external web search by default;
|
|
63
|
+
- no autonomous edits without host/human approval.
|
|
@@ -36,3 +36,39 @@ The timeline is sorted deterministically, and the summary includes:
|
|
|
36
36
|
- safe next commands.
|
|
37
37
|
|
|
38
38
|
Observability still does not collect private shell history and does not upload telemetry.
|
|
39
|
+
|
|
40
|
+
## Pipeline Recorder Expansion
|
|
41
|
+
|
|
42
|
+
Future observability work can add a pipeline recorder inspired by validate/fix/verify/report agent flows:
|
|
43
|
+
|
|
44
|
+
```txt
|
|
45
|
+
agent_started
|
|
46
|
+
agent_finished
|
|
47
|
+
agent_failed
|
|
48
|
+
file_read
|
|
49
|
+
schema_validated
|
|
50
|
+
repair_applied
|
|
51
|
+
report_generated
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
The first implementation can remain local artifacts and dashboard cards. A live SSE/server mode is optional and should not become required.
|
|
55
|
+
|
|
56
|
+
## Related 2026 Harness Rails
|
|
57
|
+
|
|
58
|
+
Observability Rail is the local timeline layer that feeds several planned rails:
|
|
59
|
+
|
|
60
|
+
- [`agent-qa-rail.md`](agent-qa-rail.md) defines eval runs, scores and regression artifacts that can become observability events.
|
|
61
|
+
- [`evidence-provenance-rail.md`](evidence-provenance-rail.md) defines verification status and provenance sidecars that reports and timelines should surface.
|
|
62
|
+
- [`agent-governance-rail.md`](agent-governance-rail.md) extends observability into trace, ledger, experiment and approval records.
|
|
63
|
+
- [`resilience-rail.md`](resilience-rail.md) adds rate-limit/fallback/provider-risk warnings as local events, not cloud telemetry.
|
|
64
|
+
|
|
65
|
+
This keeps the boundary clear: SotuRail observes local artifacts and user-approved records; it does not upload traces or become a hosted LangFuse/OpenTelemetry replacement by default.
|
|
66
|
+
|
|
67
|
+
## Ecosystem Direction
|
|
68
|
+
|
|
69
|
+
Hermes-style trajectory compression and Odysseus-style visual workspaces reinforce two SotuRail rules:
|
|
70
|
+
|
|
71
|
+
- keep observability as local, recoverable evidence rather than uploaded telemetry;
|
|
72
|
+
- present Context Packs, Memory, Reports, Evidence, Workflows, Host Compatibility, Skills, lifecycle state and Handoffs as bounded dashboard/report sections.
|
|
73
|
+
|
|
74
|
+
SotuRail does not add a required server, chat workspace or central shell interface. See [Security Boundaries](security-boundaries.md).
|