soturail 1.1.0 → 1.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +41 -8
- package/dist/cli.js +6 -0
- package/dist/cli.js.map +1 -1
- package/dist/commands/feature.d.ts +2 -0
- package/dist/commands/feature.js +28 -0
- package/dist/commands/feature.js.map +1 -0
- package/dist/commands/handoff.d.ts +2 -0
- package/dist/commands/handoff.js +14 -0
- package/dist/commands/handoff.js.map +1 -0
- package/dist/commands/harness.js +20 -0
- package/dist/commands/harness.js.map +1 -1
- package/dist/commands/session.d.ts +2 -0
- package/dist/commands/session.js +17 -0
- package/dist/commands/session.js.map +1 -0
- package/dist/core/config.d.ts +7 -0
- package/dist/core/config.js +8 -0
- package/dist/core/config.js.map +1 -1
- package/dist/core/harness-lifecycle.d.ts +63 -0
- package/dist/core/harness-lifecycle.js +397 -0
- package/dist/core/harness-lifecycle.js.map +1 -0
- package/dist/core/harness-rail.js +1 -0
- package/dist/core/harness-rail.js.map +1 -1
- package/dist/core/schema-readiness.js +9 -1
- package/dist/core/schema-readiness.js.map +1 -1
- package/dist/core/version.d.ts +1 -1
- package/dist/core/version.js +1 -1
- package/docs/agent-governance-rail.md +89 -0
- package/docs/agent-harness-synthesis-2026.md +101 -0
- package/docs/agent-hosts.md +12 -0
- package/docs/agent-qa-rail.md +92 -0
- package/docs/agents.md +13 -3
- package/docs/conductor-mode.md +44 -0
- package/docs/context-packs.md +18 -0
- package/docs/dashboard-rail.md +6 -0
- package/docs/ecosystem-influences.md +65 -0
- package/docs/eval-datasets.md +40 -0
- package/docs/evaluation-suite.md +6 -0
- package/docs/evidence-provenance-rail.md +70 -0
- package/docs/external-projects-audit.md +29 -0
- package/docs/future-rails-index.md +70 -12
- package/docs/golden-agent-tests.md +44 -0
- package/docs/governance-cost-rail.md +10 -0
- package/docs/harness-lifecycle-rail.md +95 -0
- package/docs/harness-rail.md +27 -0
- package/docs/host-compatibility-rail.md +18 -0
- package/docs/host-router-rail.md +73 -0
- package/docs/knowledge-rail.md +76 -0
- package/docs/llm-as-judge-policy.md +45 -0
- package/docs/multi-agent-workflow-templates.md +63 -0
- package/docs/observability-rail.md +36 -0
- package/docs/rate-limit-and-fallback-policy.md +45 -0
- package/docs/releases/README.md +1 -0
- package/docs/releases/RELEASE_NOTES_v1.2.0.md +43 -0
- package/docs/repo-docs-audit-2026-06-05.md +55 -0
- package/docs/report-rail.md +12 -0
- package/docs/resilience-rail.md +48 -0
- package/docs/schema-contracts.md +5 -0
- package/docs/security-boundaries.md +34 -0
- package/docs/skill-rail-2.md +11 -0
- package/docs/stable-command-surface.md +10 -1
- package/docs/tasklet-rail.md +62 -0
- package/docs/v1-contract.md +3 -2
- package/docs/workflow-rail.md +12 -0
- package/package.json +1 -1
|
@@ -0,0 +1,89 @@
|
|
|
1
|
+
# Agent Governance And Evolution Rail
|
|
2
|
+
|
|
3
|
+
Agent Governance And Evolution Rail is a later-stage direction for controlled improvement loops around SotuRail artifacts and agent-facing harnesses.
|
|
4
|
+
|
|
5
|
+
It absorbs ideas from self-improving harness systems while keeping SotuRail conservative:
|
|
6
|
+
|
|
7
|
+
```txt
|
|
8
|
+
propose -> eval -> approve -> apply
|
|
9
|
+
```
|
|
10
|
+
|
|
11
|
+
No future loop should silently edit projects or bypass human approval.
|
|
12
|
+
|
|
13
|
+
## Planned Concepts
|
|
14
|
+
|
|
15
|
+
### Agent Boundary
|
|
16
|
+
|
|
17
|
+
A future policy can define what an agent may change, read or append:
|
|
18
|
+
|
|
19
|
+
```yaml
|
|
20
|
+
mutable:
|
|
21
|
+
- src/**
|
|
22
|
+
- tests/**
|
|
23
|
+
- docs/**
|
|
24
|
+
read_only:
|
|
25
|
+
- package.json
|
|
26
|
+
- package-lock.json
|
|
27
|
+
- .github/**
|
|
28
|
+
- .env*
|
|
29
|
+
append_only:
|
|
30
|
+
- .soturail/logs/**
|
|
31
|
+
- .soturail/ledger/**
|
|
32
|
+
- .soturail/reports/**
|
|
33
|
+
rules:
|
|
34
|
+
require_tests_before_apply: true
|
|
35
|
+
require_no_context_regression: true
|
|
36
|
+
one_patch_per_run: true
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
### Ledger
|
|
40
|
+
|
|
41
|
+
Append-only records can explain decisions:
|
|
42
|
+
|
|
43
|
+
```json
|
|
44
|
+
{
|
|
45
|
+
"run_id": "run_2026_06_05_001",
|
|
46
|
+
"agent": "claude-code",
|
|
47
|
+
"action": "proposed_patch",
|
|
48
|
+
"files_touched": ["src/context.ts"],
|
|
49
|
+
"approved": false,
|
|
50
|
+
"tests_passed": true
|
|
51
|
+
}
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
### Experiments
|
|
55
|
+
|
|
56
|
+
```txt
|
|
57
|
+
.soturail/experiments/
|
|
58
|
+
candidates/
|
|
59
|
+
results/
|
|
60
|
+
accepted/
|
|
61
|
+
rejected/
|
|
62
|
+
```
|
|
63
|
+
|
|
64
|
+
Candidate changes should be compared through deterministic evals before approval.
|
|
65
|
+
|
|
66
|
+
## Proposed Commands
|
|
67
|
+
|
|
68
|
+
```bash
|
|
69
|
+
soturail trace start
|
|
70
|
+
soturail trace stop
|
|
71
|
+
soturail trace list
|
|
72
|
+
soturail trace show <id>
|
|
73
|
+
soturail ledger list
|
|
74
|
+
soturail experiment create "improve context ranking"
|
|
75
|
+
soturail experiment run
|
|
76
|
+
soturail experiment compare
|
|
77
|
+
soturail improve propose
|
|
78
|
+
soturail improve eval
|
|
79
|
+
soturail improve approve
|
|
80
|
+
soturail improve apply
|
|
81
|
+
```
|
|
82
|
+
|
|
83
|
+
## Non-Goals
|
|
84
|
+
|
|
85
|
+
- no autonomous self-modifying runtime by default;
|
|
86
|
+
- no hosted multi-tenant control plane;
|
|
87
|
+
- no required Anthropic/OpenAI/Groq provider;
|
|
88
|
+
- no destructive MCP tools;
|
|
89
|
+
- no applying patches without reviewable diff and approval.
|
|
@@ -0,0 +1,101 @@
|
|
|
1
|
+
# Agent And Harness Synthesis 2026
|
|
2
|
+
|
|
3
|
+
This document merges the latest external-repository review into one SotuRail planning note. It distinguishes agent runtimes from harness/context infrastructure so SotuRail can absorb useful ecosystem patterns without losing its product boundary. It is not a dependency list and it is not a claim that SotuRail vendors, wraps or outperforms those projects.
|
|
4
|
+
|
|
5
|
+
SotuRail remains:
|
|
6
|
+
|
|
7
|
+
```txt
|
|
8
|
+
Local-first Context OS for AI coding agents.
|
|
9
|
+
It prepares context, memory, policies, skills, workflows, evidence, reports and host exports.
|
|
10
|
+
It is not the model, not the coding agent, not a proxy, not a cloud gateway and not a trading/finance agent.
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
## Classification
|
|
14
|
+
|
|
15
|
+
```txt
|
|
16
|
+
Agent = objective + model + tools + state/memory + decision/execution loop.
|
|
17
|
+
Non-agent = skill, toolkit, toolset, compressor, router, runtime helper, context, prompt, rule or adapter.
|
|
18
|
+
```
|
|
19
|
+
|
|
20
|
+
Hermes Agent is a self-improving agent runtime with a model, tools, memory, execution loop, skills, routines, subagents and multiple interfaces.
|
|
21
|
+
|
|
22
|
+
Odysseus is broader: a self-hosted AI workspace combining an agent runtime, chat UI, local services, model management, memory, skills, research and personal workspace features.
|
|
23
|
+
|
|
24
|
+
SotuRail is different:
|
|
25
|
+
|
|
26
|
+
```txt
|
|
27
|
+
Hermes = self-improving personal agent runtime
|
|
28
|
+
Odysseus = workspace + runtime + agent + UI + local services
|
|
29
|
+
SotuRail = local-first context/harness OS for preparing and governing agents
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
Toolkits, compressors and routers should not be described as agents unless they own a model-plus-tool execution loop.
|
|
33
|
+
|
|
34
|
+
## Projects Reviewed In This Wave
|
|
35
|
+
|
|
36
|
+
| Project or product | Useful pattern | SotuRail direction |
|
|
37
|
+
| --- | --- | --- |
|
|
38
|
+
| Hermes Agent | self-improving runtime, trajectory compression, routines and role packs | context optimization, bounded sessions and safe tool profiles |
|
|
39
|
+
| Odysseus | workspace, runtime, UI and local-service integration | local workspace organization without a required server |
|
|
40
|
+
| `ijmf/qa-ai-agent` | agent QA with tests, datasets, observability and CI | Agent QA Rail, golden checks, deterministic evals |
|
|
41
|
+
| `duckdogersxd/Orquestrando-Agents-CrewAI` | multi-agent roles, fallback, retry and throttling lessons | role-pack templates and Resilience Rail docs |
|
|
42
|
+
| `CaioTakedaIA/agentesdeIA` | validate -> fix -> revalidate -> analyze pipeline with visible events | pipeline recorder, evidence per stage and dashboard timeline |
|
|
43
|
+
| `OpenTracy/OpenTracy` | propose -> eval -> approve -> apply loop, traces, ledger, candidates | Agent Governance / Evolution Rail |
|
|
44
|
+
| `walkinglabs/learn-harness-engineering` | instructions, state, verification, scope and session lifecycle | Harness Lifecycle Rail and feature/session handoff commands |
|
|
45
|
+
| `affaan-m/ECC` | install profiles, doctor/repair/uninstall, skills/rules/hooks, multi-host packaging | install state, harness audit, skills/rules profiles and host exports |
|
|
46
|
+
| `companion-inc/feynman` | provenance sidecars, verified/unverified/blocked/inferred status, file-based handoff | Evidence and Provenance Rail |
|
|
47
|
+
| `virgiliojr94/book-to-skill` | document/book -> skill, on-demand chapters, glossary, patterns and cheatsheet | Knowledge Rail and skill build/fold-in |
|
|
48
|
+
| Tasklet.ai | limited public information; reusable small tasks as a concept | Tasklet Rail as local task templates only |
|
|
49
|
+
| 9Router | multi-host routing metaphor, token/context optimization, local dashboard | Host Router Rail for context exports; no proxy/MITM |
|
|
50
|
+
|
|
51
|
+
## Patterns To Absorb
|
|
52
|
+
|
|
53
|
+
- Context and trajectory compression with recoverable evidence.
|
|
54
|
+
- Session search and bounded handoffs.
|
|
55
|
+
- Toolset profiles and role packs.
|
|
56
|
+
- Tasklet/workflow templates.
|
|
57
|
+
- Host Fit Doctor and compatibility notes.
|
|
58
|
+
- Dashboard cards for Context Packs, Memory, Reports, Evidence, Workflows, Skills, Tasklets and Handoffs.
|
|
59
|
+
- Local-first privacy and explicit security boundaries.
|
|
60
|
+
- Doctor/audit/repair/uninstall paths for generated artifacts.
|
|
61
|
+
- Provenance, evidence and verification status for reports.
|
|
62
|
+
|
|
63
|
+
## Patterns To Avoid
|
|
64
|
+
|
|
65
|
+
- Mandatory web servers or hosted workspaces.
|
|
66
|
+
- Model serving and GPU management.
|
|
67
|
+
- Central shell execution or destructive MCP tools.
|
|
68
|
+
- Provider-specific dependencies and hidden external services.
|
|
69
|
+
- Unreviewed autonomous edit loops.
|
|
70
|
+
- Proxy, MITM, billing, unlimited-free or cloud telemetry claims.
|
|
71
|
+
- Interception of IDE traffic, browser sessions, cookies or provider credentials.
|
|
72
|
+
|
|
73
|
+
## Rail Map
|
|
74
|
+
|
|
75
|
+
| New or expanded rail | Main purpose | Inspired by |
|
|
76
|
+
| --- | --- | --- |
|
|
77
|
+
| Agent QA Rail | datasets, golden checks, regression reports and optional judges | `qa-ai-agent` |
|
|
78
|
+
| Resilience Rail | rate-limit, retry, fallback and provider-risk policy docs | CrewAI/LiteLLM-style repo, 9Router concept |
|
|
79
|
+
| Pipeline Recorder | record agent-stage events and evidence packs | `agentesdeIA` |
|
|
80
|
+
| Agent Governance Rail | trace, ledger, candidates, approval gate and improve loop | OpenTracy |
|
|
81
|
+
| Harness Lifecycle Rail | init, audit, feature list, session start/end and handoff | learn-harness-engineering, Hermes, Odysseus |
|
|
82
|
+
| Install/Profile Rail | profile-based install state, doctor, repair and uninstall | ECC |
|
|
83
|
+
| Evidence/Provenance Rail | provenance sidecars and verification status | Feynman |
|
|
84
|
+
| Knowledge Rail | document ingestion into on-demand skills | book-to-skill |
|
|
85
|
+
| Host Router Rail | one context source exported to many hosts with fallback formats | 9Router metaphor |
|
|
86
|
+
| Tasklet Rail | small reusable local task templates | Tasklet concept |
|
|
87
|
+
|
|
88
|
+
## SotuRail Direction
|
|
89
|
+
|
|
90
|
+
v1.2.0 implements Harness Lifecycle Rail for local state, audits, feature tracking, sessions and handoffs. The optional [Conductor Mode](conductor-mode.md) remains proposed future work behind approval gates.
|
|
91
|
+
|
|
92
|
+
```txt
|
|
93
|
+
v1.1.1 Host Compatibility Polish, docs synthesis, golden export checks
|
|
94
|
+
v1.2.0 Harness Lifecycle Rail plus staged Spec, Design and Diagram work
|
|
95
|
+
v1.3.0 Knowledge, Evidence and Evaluation Rail
|
|
96
|
+
v1.4.0 Skill Rail 2.0, Knowledge-to-Skill and Tasklet Packs
|
|
97
|
+
v1.5.0 Governance, Cost, Resilience and Host Router Rail
|
|
98
|
+
v1.6.0 Agent Governance / Evolution Rail
|
|
99
|
+
```
|
|
100
|
+
|
|
101
|
+
The exact future command names are not frozen. Related: [Ecosystem Influences](ecosystem-influences.md), [External Projects Audit](external-projects-audit.md), [Harness Lifecycle Rail](harness-lifecycle-rail.md), [Security Boundaries](security-boundaries.md).
|
package/docs/agent-hosts.md
CHANGED
|
@@ -36,3 +36,15 @@ soturail mcp resources host-manifest --host codex
|
|
|
36
36
|
No host row enables destructive MCP tools or arbitrary shell execution. Config writes remain dry-run or review-first. Use `soturail report agent --agent <host>`, `soturail agents doctor --host <host>` and review the output before agent handoff.
|
|
37
37
|
|
|
38
38
|
See [host-matrix-schema.md](host-matrix-schema.md), [agent-export-contract.md](agent-export-contract.md) and [mcp-host-manifest.md](mcp-host-manifest.md).
|
|
39
|
+
|
|
40
|
+
## Related Host Router And Resilience Docs
|
|
41
|
+
|
|
42
|
+
Host docs are connected to several planned rails:
|
|
43
|
+
|
|
44
|
+
- [`host-router-rail.md`](host-router-rail.md): one local context source exported into host-specific formats with safe fallback.
|
|
45
|
+
- [`host-compatibility-rail.md`](host-compatibility-rail.md): current host matrix, schema and export contracts.
|
|
46
|
+
- [`rate-limit-and-fallback-policy.md`](rate-limit-and-fallback-policy.md): local documentation shape for retry/fallback/rate-limit expectations.
|
|
47
|
+
- [`resilience-rail.md`](resilience-rail.md): provider and workflow-risk warnings without proxying model traffic.
|
|
48
|
+
- [`agent-qa-rail.md`](agent-qa-rail.md): golden checks for host exports.
|
|
49
|
+
|
|
50
|
+
SotuRail host support means context, instruction, report, skill and MCP-resource packaging. It does not mean intercepting IDE traffic, reusing browser tokens, routing model requests or bypassing provider quotas.
|
|
@@ -0,0 +1,92 @@
|
|
|
1
|
+
# Agent QA Rail
|
|
2
|
+
|
|
3
|
+
Agent QA Rail is a proposed future rail for testing SotuRail-generated agent artifacts like software outputs, not magic prompts.
|
|
4
|
+
|
|
5
|
+
It is inspired by small QA-agent repos that combine automated tests, fixed datasets, scoring, observability and CI. SotuRail should keep the useful discipline while staying local-first, host-independent and deterministic by default.
|
|
6
|
+
|
|
7
|
+
## Goals
|
|
8
|
+
|
|
9
|
+
- Test host exports, context packs, skills, rules, workflows, evidence packs and reports.
|
|
10
|
+
- Catch regressions in agent-facing files before release.
|
|
11
|
+
- Keep default evals offline and cheap.
|
|
12
|
+
- Allow optional provider-backed judging only when explicitly requested.
|
|
13
|
+
- Generate JSON and Markdown reports that can be attached to release evidence.
|
|
14
|
+
|
|
15
|
+
## Proposed Commands
|
|
16
|
+
|
|
17
|
+
```bash
|
|
18
|
+
soturail eval dataset init
|
|
19
|
+
soturail eval dataset run
|
|
20
|
+
soturail eval golden
|
|
21
|
+
soturail eval regression
|
|
22
|
+
soturail eval report
|
|
23
|
+
soturail eval doctor
|
|
24
|
+
soturail eval judge --optional
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
These commands can begin as report-only or fixture-only surfaces before becoming full command implementations.
|
|
28
|
+
|
|
29
|
+
## Local Artifact Layout
|
|
30
|
+
|
|
31
|
+
```txt
|
|
32
|
+
.soturail/evals/
|
|
33
|
+
datasets/
|
|
34
|
+
host-compatibility.json
|
|
35
|
+
context-quality.json
|
|
36
|
+
skill-routing.json
|
|
37
|
+
golden/
|
|
38
|
+
claude-export.md
|
|
39
|
+
codex-export.md
|
|
40
|
+
cursor-rules.md
|
|
41
|
+
runs/
|
|
42
|
+
latest.json
|
|
43
|
+
reports/
|
|
44
|
+
latest.md
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
## Deterministic Checks First
|
|
48
|
+
|
|
49
|
+
Default checks should not call real LLMs or external services.
|
|
50
|
+
|
|
51
|
+
Examples:
|
|
52
|
+
|
|
53
|
+
- export is not empty;
|
|
54
|
+
- export contains safe next commands;
|
|
55
|
+
- export includes required policy warnings;
|
|
56
|
+
- export does not contain secrets;
|
|
57
|
+
- export does not contain unrelated project names such as SoturAI when the target is SotuRail;
|
|
58
|
+
- JSON artifacts parse successfully;
|
|
59
|
+
- duplicate JSON keys are detected where relevant;
|
|
60
|
+
- host matrix fields are present;
|
|
61
|
+
- MCP exports keep read-only mutation boundaries;
|
|
62
|
+
- reports include evidence paths and verification status;
|
|
63
|
+
- generated docs keep local-first and no-cloud-by-default claims honest.
|
|
64
|
+
|
|
65
|
+
## Optional LLM-As-Judge
|
|
66
|
+
|
|
67
|
+
Provider-backed judges can be useful for hallucination or answer-quality review, but they must not be the default release gate.
|
|
68
|
+
|
|
69
|
+
Policy:
|
|
70
|
+
|
|
71
|
+
```txt
|
|
72
|
+
Offline fixtures are release-blocking.
|
|
73
|
+
LLM-as-judge is optional, explicit and non-blocking unless a project opts in.
|
|
74
|
+
Provider outputs must be stored as separate integration evidence.
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
## CI Split
|
|
78
|
+
|
|
79
|
+
Recommended tiers:
|
|
80
|
+
|
|
81
|
+
| Tier | Network | Release blocking | Purpose |
|
|
82
|
+
| --- | --- | --- | --- |
|
|
83
|
+
| unit/offline | no | yes | deterministic docs, schemas, exports and fixtures |
|
|
84
|
+
| integration | optional | no by default | provider/API behavior, external host smoke |
|
|
85
|
+
| nightly | optional | no by default | long-running or flaky judge/eval experiments |
|
|
86
|
+
|
|
87
|
+
## Non-Goals
|
|
88
|
+
|
|
89
|
+
- no required Groq, OpenAI, Anthropic, LangFuse or other provider;
|
|
90
|
+
- no real model calls in default tests;
|
|
91
|
+
- no scoring metric based only on marketing-friendly numbers;
|
|
92
|
+
- no claim that passing evals proves a real agent will always behave well.
|
package/docs/agents.md
CHANGED
|
@@ -263,11 +263,19 @@ Those exports should explain:
|
|
|
263
263
|
- which payload format was chosen;
|
|
264
264
|
- whether any long context was offloaded.
|
|
265
265
|
|
|
266
|
-
##
|
|
266
|
+
## Harness And Handoff Support
|
|
267
267
|
|
|
268
|
-
SotuRail
|
|
268
|
+
SotuRail supports disciplined local handoffs without becoming the agent runtime:
|
|
269
269
|
|
|
270
|
-
|
|
270
|
+
```bash
|
|
271
|
+
soturail harness init
|
|
272
|
+
soturail harness audit
|
|
273
|
+
soturail session start "objective"
|
|
274
|
+
soturail handoff generate
|
|
275
|
+
soturail feature list
|
|
276
|
+
```
|
|
277
|
+
|
|
278
|
+
Additional role-to-role host handoff aliases remain future possibilities:
|
|
271
279
|
|
|
272
280
|
```bash
|
|
273
281
|
soturail agents handoff --from planner --to executor
|
|
@@ -285,3 +293,5 @@ A handoff should include:
|
|
|
285
293
|
- evidence pack pointers;
|
|
286
294
|
- raw/offload recovery IDs;
|
|
287
295
|
- next safe command suggestions.
|
|
296
|
+
|
|
297
|
+
Current handoffs are written to `.soturail/state/session-handoff.md`. They include bounded local state and changed-file names, do not read private shell history and do not run verification commands.
|
|
@@ -0,0 +1,44 @@
|
|
|
1
|
+
# SotuRail Conductor Mode
|
|
2
|
+
|
|
3
|
+
Status: **Proposed future optional mode. Not implemented in v1.2.0.**
|
|
4
|
+
|
|
5
|
+
SotuRail Core remains the CLI-first, npm-first, local-first Context OS and harness layer. A future optional mode called **SotuRail Conductor** may coordinate planning, verification and documentation workflows without replacing agent hosts.
|
|
6
|
+
|
|
7
|
+
```txt
|
|
8
|
+
SotuRail
|
|
9
|
+
|-- Core
|
|
10
|
+
| |-- context
|
|
11
|
+
| |-- memory
|
|
12
|
+
| |-- reports
|
|
13
|
+
| |-- workflows
|
|
14
|
+
| |-- evidence
|
|
15
|
+
| |-- host exports
|
|
16
|
+
| `-- dashboard
|
|
17
|
+
`-- Future optional Conductor mode
|
|
18
|
+
|-- planner
|
|
19
|
+
|-- verifier
|
|
20
|
+
|-- reviewer
|
|
21
|
+
|-- tasklet runner
|
|
22
|
+
|-- evidence collector
|
|
23
|
+
`-- approval gate
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
## Proposed Commands
|
|
27
|
+
|
|
28
|
+
These commands are documentation-only and do not currently exist:
|
|
29
|
+
|
|
30
|
+
```bash
|
|
31
|
+
soturail conductor plan
|
|
32
|
+
soturail conductor audit
|
|
33
|
+
soturail conductor propose
|
|
34
|
+
soturail conductor verify
|
|
35
|
+
soturail conductor apply --approved
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
## Safe Capability Boundary
|
|
39
|
+
|
|
40
|
+
A future Conductor may read a repository, create plans and tasklets, generate reports, validate links, compare host exports and propose patches. Applying a patch must require explicit approval.
|
|
41
|
+
|
|
42
|
+
It must not become a chat product, unbounded fix-everything loop, central shell agent, browser agent, cloud agent or provider-specific runtime.
|
|
43
|
+
|
|
44
|
+
See [Security Boundaries](security-boundaries.md), [Harness Lifecycle Rail](harness-lifecycle-rail.md) and [Future Rails Index](future-rails-index.md).
|
package/docs/context-packs.md
CHANGED
|
@@ -266,6 +266,12 @@ Possible future diagram context:
|
|
|
266
266
|
|
|
267
267
|
See [diagram-rail.md](diagram-rail.md).
|
|
268
268
|
|
|
269
|
+
## Lifecycle And Context Budget
|
|
270
|
+
|
|
271
|
+
Harness Lifecycle handoffs should stay bounded. Prefer the active objective, feature state, changed-file names, verification status, blockers and next commands over copying full session history.
|
|
272
|
+
|
|
273
|
+
Hermes-style trajectory compression is useful inspiration only when raw evidence remains recoverable. Odysseus-style workspace breadth should not cause every local artifact to enter every context pack. Use `context budget`, role packs, offload pointers and host-specific exports to keep handoffs focused.
|
|
274
|
+
|
|
269
275
|
## Quality Rules
|
|
270
276
|
|
|
271
277
|
Context selection, formatting and role packs should preserve:
|
|
@@ -280,3 +286,15 @@ Context selection, formatting and role packs should preserve:
|
|
|
280
286
|
- raw recovery hints.
|
|
281
287
|
|
|
282
288
|
If pruning or formatting would remove required evidence, SotuRail should report that pruning was not effective rather than pretending the context is sufficient.
|
|
289
|
+
|
|
290
|
+
## Related Context Routing And Knowledge Docs
|
|
291
|
+
|
|
292
|
+
Context Packs connect directly to the newer knowledge and host-routing plans:
|
|
293
|
+
|
|
294
|
+
- [`knowledge-rail.md`](knowledge-rail.md): compile docs/specs/notes into on-demand topic packs instead of loading everything into one prompt.
|
|
295
|
+
- [`host-router-rail.md`](host-router-rail.md): translate one local context source into host-specific exports for Claude, Codex, Cursor, OpenCode, Gemini and generic Markdown.
|
|
296
|
+
- [`tasklet-rail.md`](tasklet-rail.md): attach a small task template to a minimal context pack.
|
|
297
|
+
- [`resilience-rail.md`](resilience-rail.md): warn when a context pack implies expensive, provider-dependent or long-running workflows.
|
|
298
|
+
- [`agent-qa-rail.md`](agent-qa-rail.md): test whether compact/role-specific context packs preserve required facts.
|
|
299
|
+
|
|
300
|
+
The key rule is progressive disclosure: SotuRail should send the smallest useful context first and keep large raw evidence recoverable by path, id or offload pointer.
|
package/docs/dashboard-rail.md
CHANGED
|
@@ -4,6 +4,12 @@ v1.0.0 keeps the dashboard static and local. There is no server requirement, no
|
|
|
4
4
|
|
|
5
5
|
Dashboard Rail builds a static local HTML dashboard from report and status artifacts.
|
|
6
6
|
|
|
7
|
+
## Lifecycle Cards
|
|
8
|
+
|
|
9
|
+
Odysseus-style local workspace organization is useful dashboard inspiration, but SotuRail remains static and server-free by default. Future dashboard polish may surface bounded cards for Context Packs, Memory, Reports, Evidence, Workflows, Host Compatibility, Skills, Tasklets and Handoffs from existing local artifacts.
|
|
10
|
+
|
|
11
|
+
Harness Lifecycle state is currently available under `.soturail/state/`; dedicated dashboard cards remain planned.
|
|
12
|
+
|
|
7
13
|
```bash
|
|
8
14
|
soturail dashboard build
|
|
9
15
|
soturail dashboard open
|
|
@@ -46,6 +46,8 @@ SotuRail should remain the local rail layer that prepares those artifacts for an
|
|
|
46
46
|
| `murillo-romeu/sonar-totvs` | domain-specific skill/report | Skill Rail 2.0 |
|
|
47
47
|
| `Acauhi99/opencode-agent-system` | OpenCode agent system packaging | OpenCode export and role/skill packaging |
|
|
48
48
|
| `ComposioHQ/composio` | tool/provider integration ecosystem | compatibility manifests, not marketplace cloning |
|
|
49
|
+
| Hermes Agent | self-improving agent runtime, skills, routines, subagents and trajectory compression | context optimization, role packs and tasklet inspiration without becoming the runtime |
|
|
50
|
+
| Odysseus | self-hosted AI workspace, runtime, UI and local services | local dashboard, host-fit and evidence-report inspiration without adding a required server |
|
|
49
51
|
|
|
50
52
|
### Product Rule
|
|
51
53
|
|
|
@@ -57,6 +59,19 @@ SotuRail should absorb the durable patterns and avoid cloning products:
|
|
|
57
59
|
- Skill Rail 2.0 should package safe domain skills but not hide risky commands.
|
|
58
60
|
- Governance And Cost Rail should warn about context/workflow risk but not claim provider billing accuracy without evidence.
|
|
59
61
|
|
|
62
|
+
### Hermes Agent And Odysseus Boundary
|
|
63
|
+
|
|
64
|
+
Hermes and Odysseus reinforce why SotuRail must distinguish agents from harness components:
|
|
65
|
+
|
|
66
|
+
```txt
|
|
67
|
+
Agent = model + tools + state/memory + decision/execution loop.
|
|
68
|
+
SotuRail = local context, harness, evidence, policy and host handoff layer.
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
Hermes contributes useful inspiration for trajectory compression, session search, toolset profiles, skills and subagent role packs. Odysseus contributes useful inspiration for local dashboard cards, host fit checks, visual evidence reports and privacy/security documentation.
|
|
72
|
+
|
|
73
|
+
SotuRail does not copy their runtime, chat UI, model serving, local-service stack or central shell behavior. See [`agent-harness-synthesis-2026.md`](agent-harness-synthesis-2026.md), [`security-boundaries.md`](security-boundaries.md) and [`conductor-mode.md`](conductor-mode.md).
|
|
74
|
+
|
|
60
75
|
## 2026 Agent Runtime Update
|
|
61
76
|
|
|
62
77
|
Newer agent hosts are converging around the same architecture: a coding agent surface, local or cloud workspaces, MCP or tool adapters, reusable skills/instructions, hooks/approvals, context budgeting and evidence that proves the task is actually complete.
|
|
@@ -472,3 +487,53 @@ These ideas are intentionally staged after the core rails are stable:
|
|
|
472
487
|
15. **Auth Rail**: optional agent-readable auth docs and redaction checks.
|
|
473
488
|
16. **UI/Report Rail**: local HTML dashboard and optional MCP Apps/AG-UI-style event outputs.
|
|
474
489
|
17. **Gateway Lite**: local event routing only after memory, context selection, policy and reports are mature.
|
|
490
|
+
|
|
491
|
+
## 2026 Agent Harness Synthesis Update
|
|
492
|
+
|
|
493
|
+
A later review wave added QA-agent, multi-agent orchestration, pipeline-agent, self-improving harness, harness-engineering, cross-host ECC, Feynman-style provenance, book-to-skill, Tasklet and 9Router-style context-routing ideas.
|
|
494
|
+
|
|
495
|
+
See [`agent-harness-synthesis-2026.md`](agent-harness-synthesis-2026.md) for the consolidated table.
|
|
496
|
+
|
|
497
|
+
The main product update is:
|
|
498
|
+
|
|
499
|
+
```txt
|
|
500
|
+
SotuRail should evolve from context manager into a local harness manager:
|
|
501
|
+
context + state + verification + scope + lifecycle + evidence + host export.
|
|
502
|
+
```
|
|
503
|
+
|
|
504
|
+
New patterns to absorb:
|
|
505
|
+
|
|
506
|
+
- Agent QA with offline datasets and golden export checks.
|
|
507
|
+
- Validate -> fix -> verify -> report pipelines with evidence per stage.
|
|
508
|
+
- File-based handoff and provenance sidecars.
|
|
509
|
+
- Knowledge packs generated from docs, loaded on demand.
|
|
510
|
+
- Host router behavior for context formats, not model traffic.
|
|
511
|
+
- Rate-limit/fallback policy docs, not provider proxying.
|
|
512
|
+
- Human approval gates for improve/eval/apply loops.
|
|
513
|
+
|
|
514
|
+
Hard boundaries:
|
|
515
|
+
|
|
516
|
+
- no SoturAI/trading scope in SotuRail docs or exports;
|
|
517
|
+
- no mandatory provider APIs;
|
|
518
|
+
- no MITM/proxy/account/quota bypass features;
|
|
519
|
+
- no autonomous patching without approval;
|
|
520
|
+
- no hidden cloud telemetry.
|
|
521
|
+
|
|
522
|
+
## Coverage Checklist For The 2026 Harness Review
|
|
523
|
+
|
|
524
|
+
The reviewed materials are now mapped into SotuRail planning as follows:
|
|
525
|
+
|
|
526
|
+
| Source idea | SotuRail docs now covering it |
|
|
527
|
+
| --- | --- |
|
|
528
|
+
| QA agent with tests, datasets, CI and traces | [`agent-qa-rail.md`](agent-qa-rail.md), [`eval-datasets.md`](eval-datasets.md), [`golden-agent-tests.md`](golden-agent-tests.md), [`llm-as-judge-policy.md`](llm-as-judge-policy.md) |
|
|
529
|
+
| CrewAI/LiteLLM-style roles, fallback and rate limits | [`multi-agent-workflow-templates.md`](multi-agent-workflow-templates.md), [`resilience-rail.md`](resilience-rail.md), [`rate-limit-and-fallback-policy.md`](rate-limit-and-fallback-policy.md) |
|
|
530
|
+
| Validate -> fix -> verify -> report agent pipeline | [`observability-rail.md`](observability-rail.md), [`workflow-rail.md`](workflow-rail.md), [`examples/workflows/agent-pipeline-workflow.md`](../examples/workflows/agent-pipeline-workflow.md) |
|
|
531
|
+
| OpenTracy-style trace, ledger, approval and experiments | [`agent-governance-rail.md`](agent-governance-rail.md), [`evidence-provenance-rail.md`](evidence-provenance-rail.md) |
|
|
532
|
+
| Harness engineering lifecycle | [`harness-lifecycle-rail.md`](harness-lifecycle-rail.md), [`harness-rail.md`](harness-rail.md), [`workflow-rail.md`](workflow-rail.md) |
|
|
533
|
+
| ECC-style cross-host harness, doctor, audit and skills | [`host-compatibility-rail.md`](host-compatibility-rail.md), [`agent-hosts.md`](agent-hosts.md), [`skill-rail-2.md`](skill-rail-2.md), [`future-rails-index.md`](future-rails-index.md) |
|
|
534
|
+
| Feynman-style provenance and verifier status | [`evidence-provenance-rail.md`](evidence-provenance-rail.md), [`report-rail.md`](report-rail.md) |
|
|
535
|
+
| book-to-skill-style document-to-skill packs | [`knowledge-rail.md`](knowledge-rail.md), [`skill-rail-2.md`](skill-rail-2.md), [`context-packs.md`](context-packs.md) |
|
|
536
|
+
| Tasklet-style small reusable task blocks | [`tasklet-rail.md`](tasklet-rail.md), [`workflow-rail.md`](workflow-rail.md) |
|
|
537
|
+
| 9Router-style router metaphor and token/context savings | [`host-router-rail.md`](host-router-rail.md), [`context-packs.md`](context-packs.md), [`governance-cost-rail.md`](governance-cost-rail.md) |
|
|
538
|
+
|
|
539
|
+
This checklist is intentionally documentation-only. Runtime implementation should happen gradually through roadmap milestones and tests.
|
|
@@ -0,0 +1,40 @@
|
|
|
1
|
+
# Eval Datasets
|
|
2
|
+
|
|
3
|
+
Eval datasets are planned local fixtures for checking whether SotuRail outputs preserve the facts, paths, warnings and contracts that agents need.
|
|
4
|
+
|
|
5
|
+
## Dataset Shape
|
|
6
|
+
|
|
7
|
+
A minimal dataset case can include:
|
|
8
|
+
|
|
9
|
+
```json
|
|
10
|
+
{
|
|
11
|
+
"id": "host-export-readonly-mcp",
|
|
12
|
+
"input": {
|
|
13
|
+
"command": "soturail agents export --agent claude",
|
|
14
|
+
"fixture": "basic-typescript-project"
|
|
15
|
+
},
|
|
16
|
+
"expected": {
|
|
17
|
+
"mustContain": ["read-only", "MCP", "safe next commands"],
|
|
18
|
+
"mustNotContain": ["destructive MCP", "cloud telemetry required"],
|
|
19
|
+
"jsonPaths": ["host", "capabilities", "limitations"]
|
|
20
|
+
}
|
|
21
|
+
}
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
## Dataset Families
|
|
25
|
+
|
|
26
|
+
| Family | What it protects |
|
|
27
|
+
| --- | --- |
|
|
28
|
+
| host compatibility | exports for Claude, Codex, Cursor, OpenCode, Gemini and generic hosts |
|
|
29
|
+
| context quality | selected files, commands, errors and policy notes survive compaction |
|
|
30
|
+
| skill routing | only task-relevant skills are selected |
|
|
31
|
+
| evidence/provenance | reports include source paths, verification status and missing-proof notes |
|
|
32
|
+
| governance/cost | huge docs, broad MCP exposure and always-loaded skills warn correctly |
|
|
33
|
+
| release checks | version, changelog, pack, docs and evidence are connected |
|
|
34
|
+
|
|
35
|
+
## Rules
|
|
36
|
+
|
|
37
|
+
- Datasets must be small enough to run during normal development.
|
|
38
|
+
- Datasets must be deterministic by default.
|
|
39
|
+
- Dataset results should be written as JSON and Markdown.
|
|
40
|
+
- A dataset failure should say what evidence was missing, not only that text did not match.
|
package/docs/evaluation-suite.md
CHANGED
|
@@ -247,3 +247,9 @@ Evaluation Suite changes should not be promoted until:
|
|
|
247
247
|
- token savings and quality are separated;
|
|
248
248
|
- benchmark reports are refreshed only when intentionally requested;
|
|
249
249
|
- public claims match reproducible local results.
|
|
250
|
+
|
|
251
|
+
## Agent QA Rail Expansion
|
|
252
|
+
|
|
253
|
+
Future Agent QA work is tracked in [`agent-qa-rail.md`](agent-qa-rail.md), [`eval-datasets.md`](eval-datasets.md), [`golden-agent-tests.md`](golden-agent-tests.md) and [`llm-as-judge-policy.md`](llm-as-judge-policy.md).
|
|
254
|
+
|
|
255
|
+
The important rule remains: default evals must stay offline, deterministic and provider-agnostic. Provider-backed judges can exist only as explicit optional integration evidence.
|
|
@@ -0,0 +1,70 @@
|
|
|
1
|
+
# Evidence And Provenance Rail
|
|
2
|
+
|
|
3
|
+
Evidence And Provenance Rail is a planned expansion that makes SotuRail reports more auditable.
|
|
4
|
+
|
|
5
|
+
The core idea:
|
|
6
|
+
|
|
7
|
+
```txt
|
|
8
|
+
Every important output should say what it used, what it changed, what was verified and what is still uncertain.
|
|
9
|
+
```
|
|
10
|
+
|
|
11
|
+
## Provenance Sidecars
|
|
12
|
+
|
|
13
|
+
Reports can have sidecar files:
|
|
14
|
+
|
|
15
|
+
```txt
|
|
16
|
+
.soturail/reports/<slug>.md
|
|
17
|
+
.soturail/reports/<slug>.provenance.md
|
|
18
|
+
```
|
|
19
|
+
|
|
20
|
+
Per-run evidence can be grouped as:
|
|
21
|
+
|
|
22
|
+
```txt
|
|
23
|
+
.soturail/evidence/<run-id>/
|
|
24
|
+
report.md
|
|
25
|
+
provenance.md
|
|
26
|
+
tests.log
|
|
27
|
+
files-read.json
|
|
28
|
+
files-changed.json
|
|
29
|
+
verification.json
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
## Verification Status Values
|
|
33
|
+
|
|
34
|
+
Use explicit status labels:
|
|
35
|
+
|
|
36
|
+
| Status | Meaning |
|
|
37
|
+
| --- | --- |
|
|
38
|
+
| `verified` | backed by local file, command, test, schema or report evidence |
|
|
39
|
+
| `unverified` | plausible but not proven by available evidence |
|
|
40
|
+
| `blocked` | cannot be verified due to missing file, command failure or unavailable dependency |
|
|
41
|
+
| `inferred` | derived from available evidence but not directly stated |
|
|
42
|
+
|
|
43
|
+
## Proposed Commands
|
|
44
|
+
|
|
45
|
+
```bash
|
|
46
|
+
soturail evidence collect
|
|
47
|
+
soturail evidence verify
|
|
48
|
+
soturail evidence report
|
|
49
|
+
soturail sources compare
|
|
50
|
+
soturail review report
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
## Report Requirements
|
|
54
|
+
|
|
55
|
+
Agent-readable reports should include:
|
|
56
|
+
|
|
57
|
+
- source paths;
|
|
58
|
+
- command/eval/report ids;
|
|
59
|
+
- changed files;
|
|
60
|
+
- verification status;
|
|
61
|
+
- missing evidence;
|
|
62
|
+
- safe next commands;
|
|
63
|
+
- redaction status.
|
|
64
|
+
|
|
65
|
+
## Non-Goals
|
|
66
|
+
|
|
67
|
+
- no fake citations;
|
|
68
|
+
- no unsupported certainty;
|
|
69
|
+
- no cloud evidence store by default;
|
|
70
|
+
- no raw secret exposure in provenance files.
|
|
@@ -27,9 +27,19 @@ It is not the model, not the agent brain, not a heavy gateway and not a clone of
|
|
|
27
27
|
| `murillo-romeu/sonar-totvs` | High | Domain-specific AI skill/report | Add Skill Rail 2.0 domain skill templates and safety gates |
|
|
28
28
|
| `Acauhi99/opencode-agent-system` | Medium | OpenCode agent system patterns | Add OpenCode host export and role/skill packaging |
|
|
29
29
|
| `ComposioHQ/composio` | High | Tool/provider integration layer | Stay compatible through manifests and reports, do not become a tool marketplace |
|
|
30
|
+
| Hermes Agent | High | Self-improving personal agent runtime with tools, skills, routines, subagents and trajectory compression | Absorb context optimization and role-pack patterns while keeping SotuRail a harness |
|
|
31
|
+
| Odysseus | High | Self-hosted AI workspace combining runtime, UI, local services, memory and tools | Absorb dashboard/host-fit/report patterns without adding a required server or model manager |
|
|
30
32
|
|
|
31
33
|
## What SotuRail Should Absorb
|
|
32
34
|
|
|
35
|
+
### Hermes And Odysseus Classification
|
|
36
|
+
|
|
37
|
+
Hermes is an agent runtime. Odysseus is a workspace plus runtime and local-service stack. SotuRail remains the local-first harness/context OS that prepares and governs artifacts for those kinds of hosts.
|
|
38
|
+
|
|
39
|
+
Useful patterns include trajectory compression, session search, toolset profiles, role packs, local dashboard cards, visual evidence and explicit privacy boundaries. Model serving, mandatory web UI, central shell access and bundled personal productivity services remain outside SotuRail scope.
|
|
40
|
+
|
|
41
|
+
See [Agent And Harness Synthesis 2026](agent-harness-synthesis-2026.md) and [Security Boundaries](security-boundaries.md).
|
|
42
|
+
|
|
33
43
|
### 1. Host compatibility without becoming a host
|
|
34
44
|
|
|
35
45
|
OpenCode, Claude-compatible hosts, Antigravity-style hosts, Codex, Cursor and Deep Agents-style harnesses all need different context formats, rules, reports and safety guidance.
|
|
@@ -113,3 +123,22 @@ For security-oriented workflows, SotuRail should emphasize:
|
|
|
113
123
|
| v1.3.0 Knowledge Graph Rail | Understand-Anything and Project Brain evolution |
|
|
114
124
|
| v1.4.0 Skill Rail 2.0 | Sonar/TOTVS-style domain skill, Deep Agents skills, host exports |
|
|
115
125
|
| v1.5.0 Governance And Cost Rail | dynamic workflows, long-horizon agents, context/token budget risks |
|
|
126
|
+
|
|
127
|
+
## 2026 Harness/Eval/Provenance Review Addendum
|
|
128
|
+
|
|
129
|
+
This addendum records the later review wave that directly affects post-v1 docs and roadmap planning.
|
|
130
|
+
|
|
131
|
+
| Project/Product | Confidence | Primary pattern | SotuRail action |
|
|
132
|
+
| --- | --- | --- | --- |
|
|
133
|
+
| `ijmf/qa-ai-agent` | High | Agent QA through tests, datasets, traces and CI | Add Agent QA Rail, eval datasets, golden exports and optional judge policy |
|
|
134
|
+
| `duckdogersxd/Orquestrando-Agents-CrewAI` | Medium | Multi-agent role workflow plus fallback/rate-limit lessons | Add multi-agent workflow templates and Resilience Rail notes |
|
|
135
|
+
| `CaioTakedaIA/agentesdeIA` | Medium | Validate/fix/revalidate/analyze pipeline with visible logs | Add pipeline recorder concept and dashboard timeline ideas |
|
|
136
|
+
| `OpenTracy/OpenTracy` | High | Propose/eval/approve/apply loop, trace, ledger and candidates | Add Agent Governance/Evolution Rail after governance foundations |
|
|
137
|
+
| `walkinglabs/learn-harness-engineering` | High | Instructions, state, verification, scope, lifecycle | Add Harness Lifecycle Rail docs and feature/session handoff planning |
|
|
138
|
+
| `affaan-m/ECC` | High | Cross-host harness system with skills/rules/doctor/repair/install state | Add install profiles, audit score, skills/rules profiles and host export polish |
|
|
139
|
+
| `companion-inc/feynman` | High | Provenance sidecars and verification statuses | Add Evidence/Provenance Rail |
|
|
140
|
+
| `virgiliojr94/book-to-skill` | High | Document-to-skill with on-demand chapters/glossary/patterns | Add Knowledge Rail and skill build/fold-in planning |
|
|
141
|
+
| Tasklet.ai | Low | Public information was limited; small reusable task concept only | Add Tasklet Rail as local templates, not a vendor integration |
|
|
142
|
+
| 9Router | Medium | Router metaphor, multi-host compatibility, token/context savings | Add Host Router Rail for context exports; explicitly avoid proxy/MITM behavior |
|
|
143
|
+
|
|
144
|
+
The safe SotuRail interpretation is to generate local artifacts for external agents, not to become a model router, hosted agent platform, provider gateway or automation system that edits without approval.
|