soturail 1.1.0 → 1.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (64) hide show
  1. package/README.md +41 -8
  2. package/dist/cli.js +6 -0
  3. package/dist/cli.js.map +1 -1
  4. package/dist/commands/feature.d.ts +2 -0
  5. package/dist/commands/feature.js +28 -0
  6. package/dist/commands/feature.js.map +1 -0
  7. package/dist/commands/handoff.d.ts +2 -0
  8. package/dist/commands/handoff.js +14 -0
  9. package/dist/commands/handoff.js.map +1 -0
  10. package/dist/commands/harness.js +20 -0
  11. package/dist/commands/harness.js.map +1 -1
  12. package/dist/commands/session.d.ts +2 -0
  13. package/dist/commands/session.js +17 -0
  14. package/dist/commands/session.js.map +1 -0
  15. package/dist/core/config.d.ts +7 -0
  16. package/dist/core/config.js +8 -0
  17. package/dist/core/config.js.map +1 -1
  18. package/dist/core/harness-lifecycle.d.ts +63 -0
  19. package/dist/core/harness-lifecycle.js +397 -0
  20. package/dist/core/harness-lifecycle.js.map +1 -0
  21. package/dist/core/harness-rail.js +1 -0
  22. package/dist/core/harness-rail.js.map +1 -1
  23. package/dist/core/schema-readiness.js +9 -1
  24. package/dist/core/schema-readiness.js.map +1 -1
  25. package/dist/core/version.d.ts +1 -1
  26. package/dist/core/version.js +1 -1
  27. package/docs/agent-governance-rail.md +89 -0
  28. package/docs/agent-harness-synthesis-2026.md +101 -0
  29. package/docs/agent-hosts.md +12 -0
  30. package/docs/agent-qa-rail.md +92 -0
  31. package/docs/agents.md +13 -3
  32. package/docs/conductor-mode.md +44 -0
  33. package/docs/context-packs.md +18 -0
  34. package/docs/dashboard-rail.md +6 -0
  35. package/docs/ecosystem-influences.md +65 -0
  36. package/docs/eval-datasets.md +40 -0
  37. package/docs/evaluation-suite.md +6 -0
  38. package/docs/evidence-provenance-rail.md +70 -0
  39. package/docs/external-projects-audit.md +29 -0
  40. package/docs/future-rails-index.md +70 -12
  41. package/docs/golden-agent-tests.md +44 -0
  42. package/docs/governance-cost-rail.md +10 -0
  43. package/docs/harness-lifecycle-rail.md +95 -0
  44. package/docs/harness-rail.md +27 -0
  45. package/docs/host-compatibility-rail.md +18 -0
  46. package/docs/host-router-rail.md +73 -0
  47. package/docs/knowledge-rail.md +76 -0
  48. package/docs/llm-as-judge-policy.md +45 -0
  49. package/docs/multi-agent-workflow-templates.md +63 -0
  50. package/docs/observability-rail.md +36 -0
  51. package/docs/rate-limit-and-fallback-policy.md +45 -0
  52. package/docs/releases/README.md +1 -0
  53. package/docs/releases/RELEASE_NOTES_v1.2.0.md +43 -0
  54. package/docs/repo-docs-audit-2026-06-05.md +55 -0
  55. package/docs/report-rail.md +12 -0
  56. package/docs/resilience-rail.md +48 -0
  57. package/docs/schema-contracts.md +5 -0
  58. package/docs/security-boundaries.md +34 -0
  59. package/docs/skill-rail-2.md +11 -0
  60. package/docs/stable-command-surface.md +10 -1
  61. package/docs/tasklet-rail.md +62 -0
  62. package/docs/v1-contract.md +3 -2
  63. package/docs/workflow-rail.md +12 -0
  64. package/package.json +1 -1
@@ -0,0 +1,89 @@
1
+ # Agent Governance And Evolution Rail
2
+
3
+ Agent Governance And Evolution Rail is a later-stage direction for controlled improvement loops around SotuRail artifacts and agent-facing harnesses.
4
+
5
+ It absorbs ideas from self-improving harness systems while keeping SotuRail conservative:
6
+
7
+ ```txt
8
+ propose -> eval -> approve -> apply
9
+ ```
10
+
11
+ No future loop should silently edit projects or bypass human approval.
12
+
13
+ ## Planned Concepts
14
+
15
+ ### Agent Boundary
16
+
17
+ A future policy can define what an agent may change, read or append:
18
+
19
+ ```yaml
20
+ mutable:
21
+ - src/**
22
+ - tests/**
23
+ - docs/**
24
+ read_only:
25
+ - package.json
26
+ - package-lock.json
27
+ - .github/**
28
+ - .env*
29
+ append_only:
30
+ - .soturail/logs/**
31
+ - .soturail/ledger/**
32
+ - .soturail/reports/**
33
+ rules:
34
+ require_tests_before_apply: true
35
+ require_no_context_regression: true
36
+ one_patch_per_run: true
37
+ ```
38
+
39
+ ### Ledger
40
+
41
+ Append-only records can explain decisions:
42
+
43
+ ```json
44
+ {
45
+ "run_id": "run_2026_06_05_001",
46
+ "agent": "claude-code",
47
+ "action": "proposed_patch",
48
+ "files_touched": ["src/context.ts"],
49
+ "approved": false,
50
+ "tests_passed": true
51
+ }
52
+ ```
53
+
54
+ ### Experiments
55
+
56
+ ```txt
57
+ .soturail/experiments/
58
+ candidates/
59
+ results/
60
+ accepted/
61
+ rejected/
62
+ ```
63
+
64
+ Candidate changes should be compared through deterministic evals before approval.
65
+
66
+ ## Proposed Commands
67
+
68
+ ```bash
69
+ soturail trace start
70
+ soturail trace stop
71
+ soturail trace list
72
+ soturail trace show <id>
73
+ soturail ledger list
74
+ soturail experiment create "improve context ranking"
75
+ soturail experiment run
76
+ soturail experiment compare
77
+ soturail improve propose
78
+ soturail improve eval
79
+ soturail improve approve
80
+ soturail improve apply
81
+ ```
82
+
83
+ ## Non-Goals
84
+
85
+ - no autonomous self-modifying runtime by default;
86
+ - no hosted multi-tenant control plane;
87
+ - no required Anthropic/OpenAI/Groq provider;
88
+ - no destructive MCP tools;
89
+ - no applying patches without reviewable diff and approval.
@@ -0,0 +1,101 @@
1
+ # Agent And Harness Synthesis 2026
2
+
3
+ This document merges the latest external-repository review into one SotuRail planning note. It distinguishes agent runtimes from harness/context infrastructure so SotuRail can absorb useful ecosystem patterns without losing its product boundary. It is not a dependency list and it is not a claim that SotuRail vendors, wraps or outperforms those projects.
4
+
5
+ SotuRail remains:
6
+
7
+ ```txt
8
+ Local-first Context OS for AI coding agents.
9
+ It prepares context, memory, policies, skills, workflows, evidence, reports and host exports.
10
+ It is not the model, not the coding agent, not a proxy, not a cloud gateway and not a trading/finance agent.
11
+ ```
12
+
13
+ ## Classification
14
+
15
+ ```txt
16
+ Agent = objective + model + tools + state/memory + decision/execution loop.
17
+ Non-agent = skill, toolkit, toolset, compressor, router, runtime helper, context, prompt, rule or adapter.
18
+ ```
19
+
20
+ Hermes Agent is a self-improving agent runtime with a model, tools, memory, execution loop, skills, routines, subagents and multiple interfaces.
21
+
22
+ Odysseus is broader: a self-hosted AI workspace combining an agent runtime, chat UI, local services, model management, memory, skills, research and personal workspace features.
23
+
24
+ SotuRail is different:
25
+
26
+ ```txt
27
+ Hermes = self-improving personal agent runtime
28
+ Odysseus = workspace + runtime + agent + UI + local services
29
+ SotuRail = local-first context/harness OS for preparing and governing agents
30
+ ```
31
+
32
+ Toolkits, compressors and routers should not be described as agents unless they own a model-plus-tool execution loop.
33
+
34
+ ## Projects Reviewed In This Wave
35
+
36
+ | Project or product | Useful pattern | SotuRail direction |
37
+ | --- | --- | --- |
38
+ | Hermes Agent | self-improving runtime, trajectory compression, routines and role packs | context optimization, bounded sessions and safe tool profiles |
39
+ | Odysseus | workspace, runtime, UI and local-service integration | local workspace organization without a required server |
40
+ | `ijmf/qa-ai-agent` | agent QA with tests, datasets, observability and CI | Agent QA Rail, golden checks, deterministic evals |
41
+ | `duckdogersxd/Orquestrando-Agents-CrewAI` | multi-agent roles, fallback, retry and throttling lessons | role-pack templates and Resilience Rail docs |
42
+ | `CaioTakedaIA/agentesdeIA` | validate -> fix -> revalidate -> analyze pipeline with visible events | pipeline recorder, evidence per stage and dashboard timeline |
43
+ | `OpenTracy/OpenTracy` | propose -> eval -> approve -> apply loop, traces, ledger, candidates | Agent Governance / Evolution Rail |
44
+ | `walkinglabs/learn-harness-engineering` | instructions, state, verification, scope and session lifecycle | Harness Lifecycle Rail and feature/session handoff commands |
45
+ | `affaan-m/ECC` | install profiles, doctor/repair/uninstall, skills/rules/hooks, multi-host packaging | install state, harness audit, skills/rules profiles and host exports |
46
+ | `companion-inc/feynman` | provenance sidecars, verified/unverified/blocked/inferred status, file-based handoff | Evidence and Provenance Rail |
47
+ | `virgiliojr94/book-to-skill` | document/book -> skill, on-demand chapters, glossary, patterns and cheatsheet | Knowledge Rail and skill build/fold-in |
48
+ | Tasklet.ai | limited public information; reusable small tasks as a concept | Tasklet Rail as local task templates only |
49
+ | 9Router | multi-host routing metaphor, token/context optimization, local dashboard | Host Router Rail for context exports; no proxy/MITM |
50
+
51
+ ## Patterns To Absorb
52
+
53
+ - Context and trajectory compression with recoverable evidence.
54
+ - Session search and bounded handoffs.
55
+ - Toolset profiles and role packs.
56
+ - Tasklet/workflow templates.
57
+ - Host Fit Doctor and compatibility notes.
58
+ - Dashboard cards for Context Packs, Memory, Reports, Evidence, Workflows, Skills, Tasklets and Handoffs.
59
+ - Local-first privacy and explicit security boundaries.
60
+ - Doctor/audit/repair/uninstall paths for generated artifacts.
61
+ - Provenance, evidence and verification status for reports.
62
+
63
+ ## Patterns To Avoid
64
+
65
+ - Mandatory web servers or hosted workspaces.
66
+ - Model serving and GPU management.
67
+ - Central shell execution or destructive MCP tools.
68
+ - Provider-specific dependencies and hidden external services.
69
+ - Unreviewed autonomous edit loops.
70
+ - Proxy, MITM, billing, unlimited-free or cloud telemetry claims.
71
+ - Interception of IDE traffic, browser sessions, cookies or provider credentials.
72
+
73
+ ## Rail Map
74
+
75
+ | New or expanded rail | Main purpose | Inspired by |
76
+ | --- | --- | --- |
77
+ | Agent QA Rail | datasets, golden checks, regression reports and optional judges | `qa-ai-agent` |
78
+ | Resilience Rail | rate-limit, retry, fallback and provider-risk policy docs | CrewAI/LiteLLM-style repo, 9Router concept |
79
+ | Pipeline Recorder | record agent-stage events and evidence packs | `agentesdeIA` |
80
+ | Agent Governance Rail | trace, ledger, candidates, approval gate and improve loop | OpenTracy |
81
+ | Harness Lifecycle Rail | init, audit, feature list, session start/end and handoff | learn-harness-engineering, Hermes, Odysseus |
82
+ | Install/Profile Rail | profile-based install state, doctor, repair and uninstall | ECC |
83
+ | Evidence/Provenance Rail | provenance sidecars and verification status | Feynman |
84
+ | Knowledge Rail | document ingestion into on-demand skills | book-to-skill |
85
+ | Host Router Rail | one context source exported to many hosts with fallback formats | 9Router metaphor |
86
+ | Tasklet Rail | small reusable local task templates | Tasklet concept |
87
+
88
+ ## SotuRail Direction
89
+
90
+ v1.2.0 implements Harness Lifecycle Rail for local state, audits, feature tracking, sessions and handoffs. The optional [Conductor Mode](conductor-mode.md) remains proposed future work behind approval gates.
91
+
92
+ ```txt
93
+ v1.1.1 Host Compatibility Polish, docs synthesis, golden export checks
94
+ v1.2.0 Harness Lifecycle Rail plus staged Spec, Design and Diagram work
95
+ v1.3.0 Knowledge, Evidence and Evaluation Rail
96
+ v1.4.0 Skill Rail 2.0, Knowledge-to-Skill and Tasklet Packs
97
+ v1.5.0 Governance, Cost, Resilience and Host Router Rail
98
+ v1.6.0 Agent Governance / Evolution Rail
99
+ ```
100
+
101
+ The exact future command names are not frozen. Related: [Ecosystem Influences](ecosystem-influences.md), [External Projects Audit](external-projects-audit.md), [Harness Lifecycle Rail](harness-lifecycle-rail.md), [Security Boundaries](security-boundaries.md).
@@ -36,3 +36,15 @@ soturail mcp resources host-manifest --host codex
36
36
  No host row enables destructive MCP tools or arbitrary shell execution. Config writes remain dry-run or review-first. Use `soturail report agent --agent <host>`, `soturail agents doctor --host <host>` and review the output before agent handoff.
37
37
 
38
38
  See [host-matrix-schema.md](host-matrix-schema.md), [agent-export-contract.md](agent-export-contract.md) and [mcp-host-manifest.md](mcp-host-manifest.md).
39
+
40
+ ## Related Host Router And Resilience Docs
41
+
42
+ Host docs are connected to several planned rails:
43
+
44
+ - [`host-router-rail.md`](host-router-rail.md): one local context source exported into host-specific formats with safe fallback.
45
+ - [`host-compatibility-rail.md`](host-compatibility-rail.md): current host matrix, schema and export contracts.
46
+ - [`rate-limit-and-fallback-policy.md`](rate-limit-and-fallback-policy.md): local documentation shape for retry/fallback/rate-limit expectations.
47
+ - [`resilience-rail.md`](resilience-rail.md): provider and workflow-risk warnings without proxying model traffic.
48
+ - [`agent-qa-rail.md`](agent-qa-rail.md): golden checks for host exports.
49
+
50
+ SotuRail host support means context, instruction, report, skill and MCP-resource packaging. It does not mean intercepting IDE traffic, reusing browser tokens, routing model requests or bypassing provider quotas.
@@ -0,0 +1,92 @@
1
+ # Agent QA Rail
2
+
3
+ Agent QA Rail is a proposed future rail for testing SotuRail-generated agent artifacts like software outputs, not magic prompts.
4
+
5
+ It is inspired by small QA-agent repos that combine automated tests, fixed datasets, scoring, observability and CI. SotuRail should keep the useful discipline while staying local-first, host-independent and deterministic by default.
6
+
7
+ ## Goals
8
+
9
+ - Test host exports, context packs, skills, rules, workflows, evidence packs and reports.
10
+ - Catch regressions in agent-facing files before release.
11
+ - Keep default evals offline and cheap.
12
+ - Allow optional provider-backed judging only when explicitly requested.
13
+ - Generate JSON and Markdown reports that can be attached to release evidence.
14
+
15
+ ## Proposed Commands
16
+
17
+ ```bash
18
+ soturail eval dataset init
19
+ soturail eval dataset run
20
+ soturail eval golden
21
+ soturail eval regression
22
+ soturail eval report
23
+ soturail eval doctor
24
+ soturail eval judge --optional
25
+ ```
26
+
27
+ These commands can begin as report-only or fixture-only surfaces before becoming full command implementations.
28
+
29
+ ## Local Artifact Layout
30
+
31
+ ```txt
32
+ .soturail/evals/
33
+ datasets/
34
+ host-compatibility.json
35
+ context-quality.json
36
+ skill-routing.json
37
+ golden/
38
+ claude-export.md
39
+ codex-export.md
40
+ cursor-rules.md
41
+ runs/
42
+ latest.json
43
+ reports/
44
+ latest.md
45
+ ```
46
+
47
+ ## Deterministic Checks First
48
+
49
+ Default checks should not call real LLMs or external services.
50
+
51
+ Examples:
52
+
53
+ - export is not empty;
54
+ - export contains safe next commands;
55
+ - export includes required policy warnings;
56
+ - export does not contain secrets;
57
+ - export does not contain unrelated project names such as SoturAI when the target is SotuRail;
58
+ - JSON artifacts parse successfully;
59
+ - duplicate JSON keys are detected where relevant;
60
+ - host matrix fields are present;
61
+ - MCP exports keep read-only mutation boundaries;
62
+ - reports include evidence paths and verification status;
63
+ - generated docs keep local-first and no-cloud-by-default claims honest.
64
+
65
+ ## Optional LLM-As-Judge
66
+
67
+ Provider-backed judges can be useful for hallucination or answer-quality review, but they must not be the default release gate.
68
+
69
+ Policy:
70
+
71
+ ```txt
72
+ Offline fixtures are release-blocking.
73
+ LLM-as-judge is optional, explicit and non-blocking unless a project opts in.
74
+ Provider outputs must be stored as separate integration evidence.
75
+ ```
76
+
77
+ ## CI Split
78
+
79
+ Recommended tiers:
80
+
81
+ | Tier | Network | Release blocking | Purpose |
82
+ | --- | --- | --- | --- |
83
+ | unit/offline | no | yes | deterministic docs, schemas, exports and fixtures |
84
+ | integration | optional | no by default | provider/API behavior, external host smoke |
85
+ | nightly | optional | no by default | long-running or flaky judge/eval experiments |
86
+
87
+ ## Non-Goals
88
+
89
+ - no required Groq, OpenAI, Anthropic, LangFuse or other provider;
90
+ - no real model calls in default tests;
91
+ - no scoring metric based only on marketing-friendly numbers;
92
+ - no claim that passing evals proves a real agent will always behave well.
package/docs/agents.md CHANGED
@@ -263,11 +263,19 @@ Those exports should explain:
263
263
  - which payload format was chosen;
264
264
  - whether any long context was offloaded.
265
265
 
266
- ## Future Harness And Handoff Support
266
+ ## Harness And Handoff Support
267
267
 
268
- SotuRail can support disciplined handoffs without becoming the agent runtime.
268
+ SotuRail supports disciplined local handoffs without becoming the agent runtime:
269
269
 
270
- Possible future commands:
270
+ ```bash
271
+ soturail harness init
272
+ soturail harness audit
273
+ soturail session start "objective"
274
+ soturail handoff generate
275
+ soturail feature list
276
+ ```
277
+
278
+ Additional role-to-role host handoff aliases remain future possibilities:
271
279
 
272
280
  ```bash
273
281
  soturail agents handoff --from planner --to executor
@@ -285,3 +293,5 @@ A handoff should include:
285
293
  - evidence pack pointers;
286
294
  - raw/offload recovery IDs;
287
295
  - next safe command suggestions.
296
+
297
+ Current handoffs are written to `.soturail/state/session-handoff.md`. They include bounded local state and changed-file names, do not read private shell history and do not run verification commands.
@@ -0,0 +1,44 @@
1
+ # SotuRail Conductor Mode
2
+
3
+ Status: **Proposed future optional mode. Not implemented in v1.2.0.**
4
+
5
+ SotuRail Core remains the CLI-first, npm-first, local-first Context OS and harness layer. A future optional mode called **SotuRail Conductor** may coordinate planning, verification and documentation workflows without replacing agent hosts.
6
+
7
+ ```txt
8
+ SotuRail
9
+ |-- Core
10
+ | |-- context
11
+ | |-- memory
12
+ | |-- reports
13
+ | |-- workflows
14
+ | |-- evidence
15
+ | |-- host exports
16
+ | `-- dashboard
17
+ `-- Future optional Conductor mode
18
+ |-- planner
19
+ |-- verifier
20
+ |-- reviewer
21
+ |-- tasklet runner
22
+ |-- evidence collector
23
+ `-- approval gate
24
+ ```
25
+
26
+ ## Proposed Commands
27
+
28
+ These commands are documentation-only and do not currently exist:
29
+
30
+ ```bash
31
+ soturail conductor plan
32
+ soturail conductor audit
33
+ soturail conductor propose
34
+ soturail conductor verify
35
+ soturail conductor apply --approved
36
+ ```
37
+
38
+ ## Safe Capability Boundary
39
+
40
+ A future Conductor may read a repository, create plans and tasklets, generate reports, validate links, compare host exports and propose patches. Applying a patch must require explicit approval.
41
+
42
+ It must not become a chat product, unbounded fix-everything loop, central shell agent, browser agent, cloud agent or provider-specific runtime.
43
+
44
+ See [Security Boundaries](security-boundaries.md), [Harness Lifecycle Rail](harness-lifecycle-rail.md) and [Future Rails Index](future-rails-index.md).
@@ -266,6 +266,12 @@ Possible future diagram context:
266
266
 
267
267
  See [diagram-rail.md](diagram-rail.md).
268
268
 
269
+ ## Lifecycle And Context Budget
270
+
271
+ Harness Lifecycle handoffs should stay bounded. Prefer the active objective, feature state, changed-file names, verification status, blockers and next commands over copying full session history.
272
+
273
+ Hermes-style trajectory compression is useful inspiration only when raw evidence remains recoverable. Odysseus-style workspace breadth should not cause every local artifact to enter every context pack. Use `context budget`, role packs, offload pointers and host-specific exports to keep handoffs focused.
274
+
269
275
  ## Quality Rules
270
276
 
271
277
  Context selection, formatting and role packs should preserve:
@@ -280,3 +286,15 @@ Context selection, formatting and role packs should preserve:
280
286
  - raw recovery hints.
281
287
 
282
288
  If pruning or formatting would remove required evidence, SotuRail should report that pruning was not effective rather than pretending the context is sufficient.
289
+
290
+ ## Related Context Routing And Knowledge Docs
291
+
292
+ Context Packs connect directly to the newer knowledge and host-routing plans:
293
+
294
+ - [`knowledge-rail.md`](knowledge-rail.md): compile docs/specs/notes into on-demand topic packs instead of loading everything into one prompt.
295
+ - [`host-router-rail.md`](host-router-rail.md): translate one local context source into host-specific exports for Claude, Codex, Cursor, OpenCode, Gemini and generic Markdown.
296
+ - [`tasklet-rail.md`](tasklet-rail.md): attach a small task template to a minimal context pack.
297
+ - [`resilience-rail.md`](resilience-rail.md): warn when a context pack implies expensive, provider-dependent or long-running workflows.
298
+ - [`agent-qa-rail.md`](agent-qa-rail.md): test whether compact/role-specific context packs preserve required facts.
299
+
300
+ The key rule is progressive disclosure: SotuRail should send the smallest useful context first and keep large raw evidence recoverable by path, id or offload pointer.
@@ -4,6 +4,12 @@ v1.0.0 keeps the dashboard static and local. There is no server requirement, no
4
4
 
5
5
  Dashboard Rail builds a static local HTML dashboard from report and status artifacts.
6
6
 
7
+ ## Lifecycle Cards
8
+
9
+ Odysseus-style local workspace organization is useful dashboard inspiration, but SotuRail remains static and server-free by default. Future dashboard polish may surface bounded cards for Context Packs, Memory, Reports, Evidence, Workflows, Host Compatibility, Skills, Tasklets and Handoffs from existing local artifacts.
10
+
11
+ Harness Lifecycle state is currently available under `.soturail/state/`; dedicated dashboard cards remain planned.
12
+
7
13
  ```bash
8
14
  soturail dashboard build
9
15
  soturail dashboard open
@@ -46,6 +46,8 @@ SotuRail should remain the local rail layer that prepares those artifacts for an
46
46
  | `murillo-romeu/sonar-totvs` | domain-specific skill/report | Skill Rail 2.0 |
47
47
  | `Acauhi99/opencode-agent-system` | OpenCode agent system packaging | OpenCode export and role/skill packaging |
48
48
  | `ComposioHQ/composio` | tool/provider integration ecosystem | compatibility manifests, not marketplace cloning |
49
+ | Hermes Agent | self-improving agent runtime, skills, routines, subagents and trajectory compression | context optimization, role packs and tasklet inspiration without becoming the runtime |
50
+ | Odysseus | self-hosted AI workspace, runtime, UI and local services | local dashboard, host-fit and evidence-report inspiration without adding a required server |
49
51
 
50
52
  ### Product Rule
51
53
 
@@ -57,6 +59,19 @@ SotuRail should absorb the durable patterns and avoid cloning products:
57
59
  - Skill Rail 2.0 should package safe domain skills but not hide risky commands.
58
60
  - Governance And Cost Rail should warn about context/workflow risk but not claim provider billing accuracy without evidence.
59
61
 
62
+ ### Hermes Agent And Odysseus Boundary
63
+
64
+ Hermes and Odysseus reinforce why SotuRail must distinguish agents from harness components:
65
+
66
+ ```txt
67
+ Agent = model + tools + state/memory + decision/execution loop.
68
+ SotuRail = local context, harness, evidence, policy and host handoff layer.
69
+ ```
70
+
71
+ Hermes contributes useful inspiration for trajectory compression, session search, toolset profiles, skills and subagent role packs. Odysseus contributes useful inspiration for local dashboard cards, host fit checks, visual evidence reports and privacy/security documentation.
72
+
73
+ SotuRail does not copy their runtime, chat UI, model serving, local-service stack or central shell behavior. See [`agent-harness-synthesis-2026.md`](agent-harness-synthesis-2026.md), [`security-boundaries.md`](security-boundaries.md) and [`conductor-mode.md`](conductor-mode.md).
74
+
60
75
  ## 2026 Agent Runtime Update
61
76
 
62
77
  Newer agent hosts are converging around the same architecture: a coding agent surface, local or cloud workspaces, MCP or tool adapters, reusable skills/instructions, hooks/approvals, context budgeting and evidence that proves the task is actually complete.
@@ -472,3 +487,53 @@ These ideas are intentionally staged after the core rails are stable:
472
487
  15. **Auth Rail**: optional agent-readable auth docs and redaction checks.
473
488
  16. **UI/Report Rail**: local HTML dashboard and optional MCP Apps/AG-UI-style event outputs.
474
489
  17. **Gateway Lite**: local event routing only after memory, context selection, policy and reports are mature.
490
+
491
+ ## 2026 Agent Harness Synthesis Update
492
+
493
+ A later review wave added QA-agent, multi-agent orchestration, pipeline-agent, self-improving harness, harness-engineering, cross-host ECC, Feynman-style provenance, book-to-skill, Tasklet and 9Router-style context-routing ideas.
494
+
495
+ See [`agent-harness-synthesis-2026.md`](agent-harness-synthesis-2026.md) for the consolidated table.
496
+
497
+ The main product update is:
498
+
499
+ ```txt
500
+ SotuRail should evolve from context manager into a local harness manager:
501
+ context + state + verification + scope + lifecycle + evidence + host export.
502
+ ```
503
+
504
+ New patterns to absorb:
505
+
506
+ - Agent QA with offline datasets and golden export checks.
507
+ - Validate -> fix -> verify -> report pipelines with evidence per stage.
508
+ - File-based handoff and provenance sidecars.
509
+ - Knowledge packs generated from docs, loaded on demand.
510
+ - Host router behavior for context formats, not model traffic.
511
+ - Rate-limit/fallback policy docs, not provider proxying.
512
+ - Human approval gates for improve/eval/apply loops.
513
+
514
+ Hard boundaries:
515
+
516
+ - no SoturAI/trading scope in SotuRail docs or exports;
517
+ - no mandatory provider APIs;
518
+ - no MITM/proxy/account/quota bypass features;
519
+ - no autonomous patching without approval;
520
+ - no hidden cloud telemetry.
521
+
522
+ ## Coverage Checklist For The 2026 Harness Review
523
+
524
+ The reviewed materials are now mapped into SotuRail planning as follows:
525
+
526
+ | Source idea | SotuRail docs now covering it |
527
+ | --- | --- |
528
+ | QA agent with tests, datasets, CI and traces | [`agent-qa-rail.md`](agent-qa-rail.md), [`eval-datasets.md`](eval-datasets.md), [`golden-agent-tests.md`](golden-agent-tests.md), [`llm-as-judge-policy.md`](llm-as-judge-policy.md) |
529
+ | CrewAI/LiteLLM-style roles, fallback and rate limits | [`multi-agent-workflow-templates.md`](multi-agent-workflow-templates.md), [`resilience-rail.md`](resilience-rail.md), [`rate-limit-and-fallback-policy.md`](rate-limit-and-fallback-policy.md) |
530
+ | Validate -> fix -> verify -> report agent pipeline | [`observability-rail.md`](observability-rail.md), [`workflow-rail.md`](workflow-rail.md), [`examples/workflows/agent-pipeline-workflow.md`](../examples/workflows/agent-pipeline-workflow.md) |
531
+ | OpenTracy-style trace, ledger, approval and experiments | [`agent-governance-rail.md`](agent-governance-rail.md), [`evidence-provenance-rail.md`](evidence-provenance-rail.md) |
532
+ | Harness engineering lifecycle | [`harness-lifecycle-rail.md`](harness-lifecycle-rail.md), [`harness-rail.md`](harness-rail.md), [`workflow-rail.md`](workflow-rail.md) |
533
+ | ECC-style cross-host harness, doctor, audit and skills | [`host-compatibility-rail.md`](host-compatibility-rail.md), [`agent-hosts.md`](agent-hosts.md), [`skill-rail-2.md`](skill-rail-2.md), [`future-rails-index.md`](future-rails-index.md) |
534
+ | Feynman-style provenance and verifier status | [`evidence-provenance-rail.md`](evidence-provenance-rail.md), [`report-rail.md`](report-rail.md) |
535
+ | book-to-skill-style document-to-skill packs | [`knowledge-rail.md`](knowledge-rail.md), [`skill-rail-2.md`](skill-rail-2.md), [`context-packs.md`](context-packs.md) |
536
+ | Tasklet-style small reusable task blocks | [`tasklet-rail.md`](tasklet-rail.md), [`workflow-rail.md`](workflow-rail.md) |
537
+ | 9Router-style router metaphor and token/context savings | [`host-router-rail.md`](host-router-rail.md), [`context-packs.md`](context-packs.md), [`governance-cost-rail.md`](governance-cost-rail.md) |
538
+
539
+ This checklist is intentionally documentation-only. Runtime implementation should happen gradually through roadmap milestones and tests.
@@ -0,0 +1,40 @@
1
+ # Eval Datasets
2
+
3
+ Eval datasets are planned local fixtures for checking whether SotuRail outputs preserve the facts, paths, warnings and contracts that agents need.
4
+
5
+ ## Dataset Shape
6
+
7
+ A minimal dataset case can include:
8
+
9
+ ```json
10
+ {
11
+ "id": "host-export-readonly-mcp",
12
+ "input": {
13
+ "command": "soturail agents export --agent claude",
14
+ "fixture": "basic-typescript-project"
15
+ },
16
+ "expected": {
17
+ "mustContain": ["read-only", "MCP", "safe next commands"],
18
+ "mustNotContain": ["destructive MCP", "cloud telemetry required"],
19
+ "jsonPaths": ["host", "capabilities", "limitations"]
20
+ }
21
+ }
22
+ ```
23
+
24
+ ## Dataset Families
25
+
26
+ | Family | What it protects |
27
+ | --- | --- |
28
+ | host compatibility | exports for Claude, Codex, Cursor, OpenCode, Gemini and generic hosts |
29
+ | context quality | selected files, commands, errors and policy notes survive compaction |
30
+ | skill routing | only task-relevant skills are selected |
31
+ | evidence/provenance | reports include source paths, verification status and missing-proof notes |
32
+ | governance/cost | huge docs, broad MCP exposure and always-loaded skills warn correctly |
33
+ | release checks | version, changelog, pack, docs and evidence are connected |
34
+
35
+ ## Rules
36
+
37
+ - Datasets must be small enough to run during normal development.
38
+ - Datasets must be deterministic by default.
39
+ - Dataset results should be written as JSON and Markdown.
40
+ - A dataset failure should say what evidence was missing, not only that text did not match.
@@ -247,3 +247,9 @@ Evaluation Suite changes should not be promoted until:
247
247
  - token savings and quality are separated;
248
248
  - benchmark reports are refreshed only when intentionally requested;
249
249
  - public claims match reproducible local results.
250
+
251
+ ## Agent QA Rail Expansion
252
+
253
+ Future Agent QA work is tracked in [`agent-qa-rail.md`](agent-qa-rail.md), [`eval-datasets.md`](eval-datasets.md), [`golden-agent-tests.md`](golden-agent-tests.md) and [`llm-as-judge-policy.md`](llm-as-judge-policy.md).
254
+
255
+ The important rule remains: default evals must stay offline, deterministic and provider-agnostic. Provider-backed judges can exist only as explicit optional integration evidence.
@@ -0,0 +1,70 @@
1
+ # Evidence And Provenance Rail
2
+
3
+ Evidence And Provenance Rail is a planned expansion that makes SotuRail reports more auditable.
4
+
5
+ The core idea:
6
+
7
+ ```txt
8
+ Every important output should say what it used, what it changed, what was verified and what is still uncertain.
9
+ ```
10
+
11
+ ## Provenance Sidecars
12
+
13
+ Reports can have sidecar files:
14
+
15
+ ```txt
16
+ .soturail/reports/<slug>.md
17
+ .soturail/reports/<slug>.provenance.md
18
+ ```
19
+
20
+ Per-run evidence can be grouped as:
21
+
22
+ ```txt
23
+ .soturail/evidence/<run-id>/
24
+ report.md
25
+ provenance.md
26
+ tests.log
27
+ files-read.json
28
+ files-changed.json
29
+ verification.json
30
+ ```
31
+
32
+ ## Verification Status Values
33
+
34
+ Use explicit status labels:
35
+
36
+ | Status | Meaning |
37
+ | --- | --- |
38
+ | `verified` | backed by local file, command, test, schema or report evidence |
39
+ | `unverified` | plausible but not proven by available evidence |
40
+ | `blocked` | cannot be verified due to missing file, command failure or unavailable dependency |
41
+ | `inferred` | derived from available evidence but not directly stated |
42
+
43
+ ## Proposed Commands
44
+
45
+ ```bash
46
+ soturail evidence collect
47
+ soturail evidence verify
48
+ soturail evidence report
49
+ soturail sources compare
50
+ soturail review report
51
+ ```
52
+
53
+ ## Report Requirements
54
+
55
+ Agent-readable reports should include:
56
+
57
+ - source paths;
58
+ - command/eval/report ids;
59
+ - changed files;
60
+ - verification status;
61
+ - missing evidence;
62
+ - safe next commands;
63
+ - redaction status.
64
+
65
+ ## Non-Goals
66
+
67
+ - no fake citations;
68
+ - no unsupported certainty;
69
+ - no cloud evidence store by default;
70
+ - no raw secret exposure in provenance files.
@@ -27,9 +27,19 @@ It is not the model, not the agent brain, not a heavy gateway and not a clone of
27
27
  | `murillo-romeu/sonar-totvs` | High | Domain-specific AI skill/report | Add Skill Rail 2.0 domain skill templates and safety gates |
28
28
  | `Acauhi99/opencode-agent-system` | Medium | OpenCode agent system patterns | Add OpenCode host export and role/skill packaging |
29
29
  | `ComposioHQ/composio` | High | Tool/provider integration layer | Stay compatible through manifests and reports, do not become a tool marketplace |
30
+ | Hermes Agent | High | Self-improving personal agent runtime with tools, skills, routines, subagents and trajectory compression | Absorb context optimization and role-pack patterns while keeping SotuRail a harness |
31
+ | Odysseus | High | Self-hosted AI workspace combining runtime, UI, local services, memory and tools | Absorb dashboard/host-fit/report patterns without adding a required server or model manager |
30
32
 
31
33
  ## What SotuRail Should Absorb
32
34
 
35
+ ### Hermes And Odysseus Classification
36
+
37
+ Hermes is an agent runtime. Odysseus is a workspace plus runtime and local-service stack. SotuRail remains the local-first harness/context OS that prepares and governs artifacts for those kinds of hosts.
38
+
39
+ Useful patterns include trajectory compression, session search, toolset profiles, role packs, local dashboard cards, visual evidence and explicit privacy boundaries. Model serving, mandatory web UI, central shell access and bundled personal productivity services remain outside SotuRail scope.
40
+
41
+ See [Agent And Harness Synthesis 2026](agent-harness-synthesis-2026.md) and [Security Boundaries](security-boundaries.md).
42
+
33
43
  ### 1. Host compatibility without becoming a host
34
44
 
35
45
  OpenCode, Claude-compatible hosts, Antigravity-style hosts, Codex, Cursor and Deep Agents-style harnesses all need different context formats, rules, reports and safety guidance.
@@ -113,3 +123,22 @@ For security-oriented workflows, SotuRail should emphasize:
113
123
  | v1.3.0 Knowledge Graph Rail | Understand-Anything and Project Brain evolution |
114
124
  | v1.4.0 Skill Rail 2.0 | Sonar/TOTVS-style domain skill, Deep Agents skills, host exports |
115
125
  | v1.5.0 Governance And Cost Rail | dynamic workflows, long-horizon agents, context/token budget risks |
126
+
127
+ ## 2026 Harness/Eval/Provenance Review Addendum
128
+
129
+ This addendum records the later review wave that directly affects post-v1 docs and roadmap planning.
130
+
131
+ | Project/Product | Confidence | Primary pattern | SotuRail action |
132
+ | --- | --- | --- | --- |
133
+ | `ijmf/qa-ai-agent` | High | Agent QA through tests, datasets, traces and CI | Add Agent QA Rail, eval datasets, golden exports and optional judge policy |
134
+ | `duckdogersxd/Orquestrando-Agents-CrewAI` | Medium | Multi-agent role workflow plus fallback/rate-limit lessons | Add multi-agent workflow templates and Resilience Rail notes |
135
+ | `CaioTakedaIA/agentesdeIA` | Medium | Validate/fix/revalidate/analyze pipeline with visible logs | Add pipeline recorder concept and dashboard timeline ideas |
136
+ | `OpenTracy/OpenTracy` | High | Propose/eval/approve/apply loop, trace, ledger and candidates | Add Agent Governance/Evolution Rail after governance foundations |
137
+ | `walkinglabs/learn-harness-engineering` | High | Instructions, state, verification, scope, lifecycle | Add Harness Lifecycle Rail docs and feature/session handoff planning |
138
+ | `affaan-m/ECC` | High | Cross-host harness system with skills/rules/doctor/repair/install state | Add install profiles, audit score, skills/rules profiles and host export polish |
139
+ | `companion-inc/feynman` | High | Provenance sidecars and verification statuses | Add Evidence/Provenance Rail |
140
+ | `virgiliojr94/book-to-skill` | High | Document-to-skill with on-demand chapters/glossary/patterns | Add Knowledge Rail and skill build/fold-in planning |
141
+ | Tasklet.ai | Low | Public information was limited; small reusable task concept only | Add Tasklet Rail as local templates, not a vendor integration |
142
+ | 9Router | Medium | Router metaphor, multi-host compatibility, token/context savings | Add Host Router Rail for context exports; explicitly avoid proxy/MITM behavior |
143
+
144
+ The safe SotuRail interpretation is to generate local artifacts for external agents, not to become a model router, hosted agent platform, provider gateway or automation system that edits without approval.