@kontourai/flow-agents 3.3.0 → 3.4.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.github/workflows/add-to-project.yml +15 -0
- package/.github/workflows/ci.yml +161 -0
- package/CHANGELOG.md +48 -0
- package/CONTEXT.md +5 -1
- package/README.md +19 -8
- package/build/src/builder-flow-run-adapter.d.ts +80 -0
- package/build/src/builder-flow-run-adapter.js +241 -0
- package/build/src/builder-flow-runtime.d.ts +16 -0
- package/build/src/builder-flow-runtime.js +290 -0
- package/build/src/cli/builder-run.d.ts +1 -0
- package/build/src/cli/builder-run.js +27 -0
- package/build/src/cli/effective-backlog-settings.js +70 -2
- package/build/src/cli/init.d.ts +34 -0
- package/build/src/cli/init.js +341 -61
- package/build/src/cli/kit.js +55 -12
- package/build/src/cli/pull-work-provider.js +346 -5
- package/build/src/cli/skill-drift-check.d.ts +1 -0
- package/build/src/cli/skill-drift-check.js +165 -0
- package/build/src/cli/telemetry-doctor.d.ts +37 -0
- package/build/src/cli/telemetry-doctor.js +53 -6
- package/build/src/cli/validate-hook-influence.js +37 -7
- package/build/src/cli/workflow-sidecar.d.ts +93 -8
- package/build/src/cli/workflow-sidecar.js +1175 -158
- package/build/src/cli.js +5 -0
- package/build/src/flow-kit/validate.d.ts +54 -34
- package/build/src/flow-kit/validate.js +237 -26
- package/build/src/index.d.ts +2 -0
- package/build/src/index.js +1 -0
- package/build/src/lib/console-connect-options.d.ts +97 -0
- package/build/src/lib/console-connect-options.js +199 -0
- package/build/src/lib/console-telemetry-validate.d.ts +49 -0
- package/build/src/lib/console-telemetry-validate.js +91 -0
- package/build/src/lib/flow-resolver.d.ts +56 -3
- package/build/src/lib/flow-resolver.js +151 -11
- package/build/src/lib/fs.d.ts +17 -0
- package/build/src/lib/fs.js +172 -0
- package/build/src/lib/local-artifact-root.d.ts +44 -1
- package/build/src/lib/local-artifact-root.js +131 -3
- package/build/src/runtime-adapters.d.ts +39 -3
- package/build/src/runtime-adapters.js +77 -31
- package/build/src/tools/build-universal-bundles.js +40 -2
- package/build/src/tools/codex-agent-routing.d.ts +2 -0
- package/build/src/tools/codex-agent-routing.js +49 -0
- package/build/src/tools/generate-context-map.js +1 -0
- package/build/src/tools/validate-source-tree.js +27 -1
- package/context/scripts/hooks/lib/kit-catalog.js +235 -0
- package/context/scripts/hooks/lib/runnable-command.js +177 -0
- package/context/scripts/hooks/stop-goal-fit.js +278 -48
- package/context/scripts/hooks/workflow-steering.js +121 -21
- package/context/scripts/package.json +3 -0
- package/context/scripts/telemetry/install-console-config.sh +25 -4
- package/context/scripts/telemetry/lib/config.sh +102 -12
- package/context/scripts/telemetry/lib/pricing.sh +50 -0
- package/context/scripts/telemetry/lib/session.sh +3 -0
- package/context/scripts/telemetry/lib/transport.sh +87 -0
- package/context/scripts/telemetry/lib/usage.sh +205 -4
- package/context/scripts/telemetry/telemetry.conf +6 -0
- package/context/scripts/telemetry/telemetry.sh +48 -0
- package/context/settings/workspace-backlog-provider-settings.example.json +48 -0
- package/docs/agent-usage-feedback-loop.md +35 -0
- package/docs/architecture-engine-and-kits.md +110 -0
- package/docs/context-map.md +2 -0
- package/docs/decisions/embeddable-engine.md +152 -0
- package/docs/decisions/index.md +3 -1
- package/docs/decisions/trust-ledger-retention.md +88 -0
- package/docs/decisions/workflow-enforcement.md +31 -9
- package/docs/fixture-ownership.md +3 -0
- package/docs/implementing-trust-reconciliation.md +129 -0
- package/docs/index.md +19 -9
- package/docs/integrations/flow-agents-console.md +167 -0
- package/docs/kit-authoring-guide.md +52 -21
- package/docs/spec/builder-flow-runtime.md +80 -0
- package/docs/spec/runtime-hook-surface.md +45 -1
- package/docs/specs/economics-record-contract.md +270 -0
- package/docs/specs/harness-capability-matrix.md +74 -0
- package/docs/specs/learning-review-proposals-contract.md +340 -0
- package/docs/specs/routing-efficiency-review.md +59 -0
- package/docs/verifiable-trust.md +74 -25
- package/docs/workflow-usage-guide.md +10 -0
- package/evals/acceptance/prove-capture-teeth.sh +132 -0
- package/evals/ci/antigaming-suite.sh +1 -0
- package/evals/ci/run-baseline.sh +72 -4
- package/evals/fixtures/economics/acceptance.json +12 -0
- package/evals/fixtures/economics/agents/tool-worker-1/events.jsonl +2 -0
- package/evals/fixtures/economics/agents/tool-worker-2/events.jsonl +2 -0
- package/evals/fixtures/economics/agents/tool-worker-3/events.jsonl +2 -0
- package/evals/fixtures/economics/agents/tool-worker-4/events.jsonl +1 -0
- package/evals/fixtures/economics/agents/tool-worker-5/events.jsonl +2 -0
- package/evals/fixtures/economics/critique.json +22 -0
- package/evals/fixtures/economics/expected-record.json +71 -0
- package/evals/fixtures/economics/session-usage-event.json +1 -0
- package/evals/fixtures/economics/state.json +11 -0
- package/evals/fixtures/economics/transcript.jsonl +3 -0
- package/evals/fixtures/hook-influence/cases.json +7 -7
- package/evals/fixtures/learning-review-proposals/balanced/economics.jsonl +6 -0
- package/evals/fixtures/learning-review-proposals/effect-follow-up/economics.jsonl +5 -0
- package/evals/fixtures/learning-review-proposals/effect-follow-up/sessions/task-lr-ef-1/trust.bundle +21 -0
- package/evals/fixtures/learning-review-proposals/effect-follow-up/sessions/task-lr-ef-2/trust.bundle +21 -0
- package/evals/fixtures/learning-review-proposals/effect-follow-up/sessions/task-lr-ef-3/trust.bundle +21 -0
- package/evals/fixtures/learning-review-proposals/effect-follow-up/sessions/task-lr-ef-4/trust.bundle +21 -0
- package/evals/fixtures/learning-review-proposals/effect-follow-up/sessions/task-lr-ef-5/trust.bundle +21 -0
- package/evals/fixtures/learning-review-proposals/pattern-present/economics.jsonl +6 -0
- package/evals/fixtures/learning-review-proposals/pattern-present/expected-aggregates.json +30 -0
- package/evals/fixtures/learning-review-proposals/pattern-present/expected-aggregates.md +66 -0
- package/evals/fixtures/learning-review-proposals/pattern-present/sessions/task-lr-pp-1/gate-review.inquiries.json +26 -0
- package/evals/fixtures/learning-review-proposals/pattern-present/sessions/task-lr-pp-1/trust.bundle +21 -0
- package/evals/fixtures/learning-review-proposals/pattern-present/sessions/task-lr-pp-2/gate-review.inquiries.json +26 -0
- package/evals/fixtures/learning-review-proposals/pattern-present/sessions/task-lr-pp-2/trust.bundle +21 -0
- package/evals/fixtures/learning-review-proposals/pattern-present/sessions/task-lr-pp-3/gate-review.inquiries.json +26 -0
- package/evals/fixtures/learning-review-proposals/pattern-present/sessions/task-lr-pp-3/trust.bundle +21 -0
- package/evals/fixtures/learning-review-proposals/pattern-present/sessions/task-lr-pp-4/gate-review.inquiries.json +26 -0
- package/evals/fixtures/learning-review-proposals/pattern-present/sessions/task-lr-pp-4/trust.bundle +21 -0
- package/evals/fixtures/learning-review-proposals/pattern-present/sessions/task-lr-pp-5/trust.bundle +21 -0
- package/evals/fixtures/learning-review-proposals/pattern-present/sessions/task-lr-pp-6/trust.bundle +21 -0
- package/evals/fixtures/learning-review-proposals/repeat-window/economics.jsonl +6 -0
- package/evals/fixtures/learning-review-proposals/under-threshold/economics.jsonl +3 -0
- package/evals/fixtures/telemetry/usage-transcript-sample.jsonl +4 -0
- package/evals/fixtures/trust-reconcile-exploits/mcp-degrade.json +42 -0
- package/evals/integration/test_builder_entry_enforcement.sh +241 -0
- package/evals/integration/test_builder_step_producers.sh +18 -10
- package/evals/integration/test_bundle_install.sh +172 -0
- package/evals/integration/test_console_tenant_isolation.sh +167 -0
- package/evals/integration/test_critique_supersession_roundtrip.sh +4 -1
- package/evals/integration/test_dual_emit_flow_step.sh +10 -4
- package/evals/integration/test_economics_record.sh +674 -0
- package/evals/integration/test_effective_backlog_settings.sh +1 -1
- package/evals/integration/test_evidence_capture_hook.sh +17 -2
- package/evals/integration/test_exemption_usage_review.sh +198 -0
- package/evals/integration/test_fixture_retirement_audit.sh +2 -2
- package/evals/integration/test_flow_kit_install_git.sh +83 -0
- package/evals/integration/test_flowdef_session_activation.sh +0 -1
- package/evals/integration/test_flowdef_session_history_preservation.sh +13 -3
- package/evals/integration/test_gate_lockdown.sh +7 -0
- package/evals/integration/test_gate_review_inquiry_records.sh +9 -1
- package/evals/integration/test_goal_fit_hook.sh +2031 -0
- package/evals/integration/test_hook_category_behaviors.sh +8 -1
- package/evals/integration/test_hook_influence_cases.sh +25 -1
- package/evals/integration/test_install_merge.sh +227 -2
- package/evals/integration/test_kit_conformance_levels.sh +6 -6
- package/evals/integration/test_learning_review_proposals.sh +329 -0
- package/evals/integration/test_liveness_conflict_injection.sh +26 -22
- package/evals/integration/test_liveness_console_relay.sh +166 -0
- package/evals/integration/test_liveness_heartbeat.sh +17 -17
- package/evals/integration/test_liveness_worktree_root.sh +575 -0
- package/evals/integration/test_phase_map_and_gate_claim.sh +6 -1
- package/evals/integration/test_publish_delivery.sh +331 -1
- package/evals/integration/test_pull_work_board.sh +200 -0
- package/evals/integration/test_pull_work_provider.sh +1 -1
- package/evals/integration/test_record_check.sh +378 -0
- package/evals/integration/test_routing_efficiency.sh +71 -0
- package/evals/integration/test_runtime_adapter_activation.sh +28 -0
- package/evals/integration/test_session_resume_roundtrip.sh +16 -19
- package/evals/integration/test_skill_drift_check.sh +870 -0
- package/evals/integration/test_telemetry.sh +445 -0
- package/evals/integration/test_telemetry_doctor.sh +66 -0
- package/evals/integration/test_telemetry_usage_pipeline.sh +228 -0
- package/evals/integration/test_trust_reconcile_negatives.sh +30 -13
- package/evals/integration/test_trust_reconcile_trailer_diagnostic.sh +247 -0
- package/evals/integration/test_usage_cost.sh +61 -0
- package/evals/integration/test_workflow_sidecar_writer.sh +1395 -0
- package/evals/integration/test_workflow_steering_hook.sh +157 -16
- package/evals/integration/test_workspace_settings.sh +176 -0
- package/evals/lib/env.sh +26 -0
- package/evals/lib/node.sh +8 -0
- package/evals/run.sh +29 -0
- package/evals/static/test_ci_integration_coverage.sh +115 -0
- package/evals/static/test_declared_scope_forms_documented.sh +114 -0
- package/evals/static/test_universal_bundles.sh +34 -0
- package/evals/static/test_validate_source_kit_asset_scope.sh +259 -0
- package/evals/static/test_workflow_skills.sh +1 -1
- package/kits/builder/flows/build.flow.json +9 -18
- package/kits/builder/flows/publish-learn.flow.json +5 -1
- package/kits/builder/kit.json +120 -0
- package/kits/builder/skills/deliver/SKILL.md +42 -0
- package/kits/builder/skills/evidence-gate/SKILL.md +12 -0
- package/kits/builder/skills/execute-plan/SKILL.md +9 -0
- package/kits/builder/skills/learning-review/SKILL.md +51 -0
- package/kits/builder/skills/plan-work/SKILL.md +17 -20
- package/kits/builder/skills/pull-work/SKILL.md +21 -0
- package/kits/builder/skills/release-readiness/SKILL.md +12 -0
- package/kits/knowledge/kit.json +9 -0
- package/kits/veritas-governance/docs/README.md +35 -7
- package/kits/veritas-governance/fixtures/exemption-review/mixed-fresh-stale.DECLARED.json +14 -0
- package/kits/veritas-governance/kit.json +14 -0
- package/kits/veritas-governance/skills/exemption-usage-review/SKILL.md +128 -0
- package/kits/veritas-governance/skills/exemption-usage-review/review-exemptions.mjs +231 -0
- package/package.json +2 -2
- package/packaging/manifest.json +29 -0
- package/schemas/backlog-provider-settings.schema.json +13 -0
- package/schemas/workflow-state.schema.json +44 -0
- package/scripts/README.md +4 -0
- package/scripts/check-content-boundary.cjs +8 -1
- package/scripts/ci/trust-reconcile.js +136 -0
- package/scripts/hooks/codex-hook-adapter.js +77 -2
- package/scripts/hooks/evidence-capture.js +38 -5
- package/scripts/hooks/lib/codex-exit-code.js +316 -0
- package/scripts/hooks/lib/kit-catalog.js +235 -0
- package/scripts/hooks/lib/liveness-write.js +28 -1
- package/scripts/hooks/lib/local-artifact-paths.js +97 -1
- package/scripts/hooks/lib/runnable-command.js +177 -0
- package/scripts/hooks/lib/skill-drift.js +350 -0
- package/scripts/hooks/stop-goal-fit.js +278 -48
- package/scripts/hooks/workflow-steering.js +121 -21
- package/scripts/install-codex-home.sh +97 -47
- package/scripts/install-merge.js +72 -14
- package/scripts/install-owned-files.js +178 -0
- package/scripts/liveness/relay.sh +84 -0
- package/scripts/telemetry/economics-record.schema.json +145 -0
- package/scripts/telemetry/economics-record.sh +331 -0
- package/scripts/telemetry/install-console-config.sh +25 -4
- package/scripts/telemetry/learning-review-decide.sh +124 -0
- package/scripts/telemetry/learning-review-proposals.schema.json +161 -0
- package/scripts/telemetry/learning-review-proposals.sh +484 -0
- package/scripts/telemetry/lib/config.sh +102 -12
- package/scripts/telemetry/lib/pricing.sh +14 -6
- package/scripts/telemetry/lib/session.sh +3 -0
- package/scripts/telemetry/lib/transport.sh +133 -15
- package/scripts/telemetry/lib/usage.sh +121 -28
- package/scripts/telemetry/routing-efficiency.sh +0 -0
- package/scripts/telemetry/telemetry.conf +6 -0
- package/scripts/telemetry/telemetry.sh +48 -0
- package/src/builder-flow-run-adapter.ts +357 -0
- package/src/builder-flow-runtime.ts +348 -0
- package/src/cli/builder-flow-run-adapter.test.mjs +495 -0
- package/src/cli/builder-flow-runtime.test.mjs +213 -0
- package/src/cli/builder-run.ts +28 -0
- package/src/cli/codex-agent-routing.test.mjs +44 -0
- package/src/cli/codex-exit-code.test.mjs +207 -0
- package/src/cli/console-connect-options.test.mjs +329 -0
- package/src/cli/console-telemetry-validate.test.mjs +157 -0
- package/src/cli/effective-backlog-settings.ts +68 -2
- package/src/cli/flow-resolver-composition.test.mjs +101 -0
- package/src/cli/init.test.mjs +161 -0
- package/src/cli/init.ts +407 -62
- package/src/cli/kit-metadata-security.test.mjs +443 -0
- package/src/cli/kit.ts +50 -12
- package/src/cli/pull-work-provider.ts +377 -3
- package/src/cli/sidecar-pure-helpers.test.mjs +64 -0
- package/src/cli/skill-drift-check.ts +196 -0
- package/src/cli/telemetry-doctor.test.mjs +53 -0
- package/src/cli/telemetry-doctor.ts +50 -7
- package/src/cli/validate-hook-influence.ts +37 -6
- package/src/cli/workflow-sidecar.ts +1150 -151
- package/src/cli.ts +5 -0
- package/src/flow-kit/validate.ts +277 -38
- package/src/index.ts +19 -0
- package/src/lib/console-connect-options.ts +261 -0
- package/src/lib/console-telemetry-validate.ts +88 -0
- package/src/lib/flow-resolver.ts +153 -10
- package/src/lib/fs.ts +160 -0
- package/src/lib/local-artifact-root.ts +129 -3
- package/src/runtime-adapters.ts +113 -33
- package/src/tools/build-universal-bundles.ts +36 -2
- package/src/tools/codex-agent-routing.ts +48 -0
- package/src/tools/generate-context-map.ts +1 -0
- package/src/tools/validate-source-tree.ts +26 -1
|
@@ -0,0 +1,152 @@
|
|
|
1
|
+
---
|
|
2
|
+
status: needs-decision
|
|
3
|
+
subject: Embeddable engine and adapter model
|
|
4
|
+
decided: 2026-07-07
|
|
5
|
+
evidence:
|
|
6
|
+
- kind: doc
|
|
7
|
+
ref: docs/spec/runtime-hook-surface.md
|
|
8
|
+
- kind: doc
|
|
9
|
+
ref: docs/decisions/flow-flow-agents-boundary.md
|
|
10
|
+
- kind: doc
|
|
11
|
+
ref: docs/decisions/trust-ledger-retention.md
|
|
12
|
+
- kind: pr
|
|
13
|
+
ref: https://github.com/kontourai/flow-agents/pull/497
|
|
14
|
+
- kind: issue
|
|
15
|
+
ref: https://github.com/kontourai/flow-agents/issues/410
|
|
16
|
+
---
|
|
17
|
+
|
|
18
|
+
# Embeddable engine and adapter model
|
|
19
|
+
|
|
20
|
+
> Status: **proposed** direction (shaped with Brian Anderson, 2026-07-07). Not yet
|
|
21
|
+
> fully built. Ratify + carry forward as the runtime is factored. This record is
|
|
22
|
+
> the north star the adapter/port refactors implement against; the linked backlog
|
|
23
|
+
> issues are the work.
|
|
24
|
+
|
|
25
|
+
**Decision.** Flow Agents is one **runtime-agnostic engine** with **thin
|
|
26
|
+
adapters** on the edges. A development harness (Claude Code, Codex, opencode, pi,
|
|
27
|
+
Kiro) and a customer agent built on a framework/SDK (AWS Strands, VoltAgent,
|
|
28
|
+
LangGraph, OpenAI Agents SDK) are the **same shape**: both are adapters that feed
|
|
29
|
+
canonical Flow events into the one engine and honor the same contracts. A Flow or
|
|
30
|
+
Flow-Agent-Kit implementation should behave **consistently** across development
|
|
31
|
+
workflows and customer agent implementations, with **explicit callouts** wherever
|
|
32
|
+
a runtime genuinely cannot offer a Flow feature — never silent divergence.
|
|
33
|
+
|
|
34
|
+
Five rules follow.
|
|
35
|
+
|
|
36
|
+
1. **One core, adapters at the edge (DRY by construction).** The engine owns
|
|
37
|
+
everything runtime-independent: the canonical event vocabulary, the redaction
|
|
38
|
+
contract, project/session/actor attribution, dual-channel transport, trust
|
|
39
|
+
bundles, gate/claim semantics, and state projection. An adapter's only job is
|
|
40
|
+
to translate a runtime's native surface into canonical events and to invoke
|
|
41
|
+
engine decisions — it holds **no policy of its own**. Today the harness
|
|
42
|
+
adapters already prove this: one shared shell core
|
|
43
|
+
(`scripts/telemetry/telemetry.sh` + `lib/*.sh`) with thin per-runtime JS
|
|
44
|
+
adapters (~130 lines each) that normalize hook input and shell into the core.
|
|
45
|
+
Duplicated logic across adapters is a bug against this record.
|
|
46
|
+
|
|
47
|
+
2. **Harness adapters and framework adapters are the same category.** A CLI
|
|
48
|
+
harness emits events from shell hooks; an in-process framework emits the same
|
|
49
|
+
canonical events from a language-native package — no shelling out, honoring the
|
|
50
|
+
same redaction contract in-process. The difference is the **transport into the
|
|
51
|
+
core**, not the **contract or the semantics**. `context.project`,
|
|
52
|
+
`context.cwd` redaction, actor identity, and dual-channel routing are derived
|
|
53
|
+
the same way everywhere (see `docs/spec/runtime-hook-surface.md`, which is now
|
|
54
|
+
the canonical adapter surface, and PR #497 which made `context.project` a
|
|
55
|
+
canonical field). Where the shell core cannot be reused verbatim in-process,
|
|
56
|
+
the engine's runtime-agnostic logic is extracted to a language port the
|
|
57
|
+
framework adapter calls; the shell adapter becomes one caller of that port, not
|
|
58
|
+
the definition of it.
|
|
59
|
+
|
|
60
|
+
3. **Flow Agents is an "SDK" engine layer, and Console is its remote backend.**
|
|
61
|
+
The same engine that instruments a dev session is the layer a customer embeds
|
|
62
|
+
in their own agent to get Flow's trust/state/economics for free. When embedded,
|
|
63
|
+
the **Console is the remote trust and state backend** for those SDK-embedded
|
|
64
|
+
agents — the same projection surface described in
|
|
65
|
+
`docs/decisions/trust-ledger-retention.md` (git/CI authoritative, console a
|
|
66
|
+
rebuildable projection), now fed by customer agents as well as dev harnesses.
|
|
67
|
+
The engine does not care whether the events came from a terminal or from a
|
|
68
|
+
long-running service.
|
|
69
|
+
|
|
70
|
+
4. **Cooperative in-process, authoritative at the boundary ("agent proposes, CI
|
|
71
|
+
disposes").** A framework/SDK adapter runs inside the customer's process, which
|
|
72
|
+
the engine does not control, so in-process enforcement is **cooperative**: the
|
|
73
|
+
embedded engine emits claims and honors gates as a good citizen. Authority
|
|
74
|
+
lives at the trust boundary the customer does not own — the **Console ingest
|
|
75
|
+
and CI trust-reconcile** re-verify what the agent claimed against independent
|
|
76
|
+
results (the fail-closed reconciliation from
|
|
77
|
+
`docs/adr/0022`). An adapter can be sloppy or hostile and the boundary still
|
|
78
|
+
holds the line. This is the same posture that already governs harness delivery;
|
|
79
|
+
it generalizes unchanged to embedded agents.
|
|
80
|
+
|
|
81
|
+
5. **Consistency is guaranteed by a conformance suite, not by hope
|
|
82
|
+
(OpenTelemetry model).** Every adapter — shell harness or language framework —
|
|
83
|
+
must pass a shared **conformance suite** that asserts the canonical contracts:
|
|
84
|
+
event shape, redaction defaults (full local path never leaves the machine),
|
|
85
|
+
attribution precedence, dual-channel routing, and gate/claim behavior. A new
|
|
86
|
+
adapter is "done" when it passes the suite, exactly as an OpenTelemetry SDK is
|
|
87
|
+
conformant when it passes the spec's tests. Feature gaps a runtime cannot close
|
|
88
|
+
are declared **explicitly** in the adapter's conformance report, not left for a
|
|
89
|
+
user to discover.
|
|
90
|
+
|
|
91
|
+
## Consequences / required refactors
|
|
92
|
+
|
|
93
|
+
The backlog issues that carry this direction:
|
|
94
|
+
|
|
95
|
+
- **Extract a runtime-agnostic core** (#500). The engine logic currently expressed
|
|
96
|
+
in shell (`scripts/telemetry/lib/*`) must be factored so its policy — event
|
|
97
|
+
canonicalization, redaction, attribution precedence, transport routing — is a
|
|
98
|
+
reusable port with at least a shell binding (today) and a language binding
|
|
99
|
+
(first framework adapter). No new policy may be added to an adapter that isn't
|
|
100
|
+
in the core.
|
|
101
|
+
- **The artifact/state store becomes a pluggable port** (#501). Trust bundles,
|
|
102
|
+
delivery records, and state projection are written today against a filesystem +
|
|
103
|
+
git + Console assumption. For embedded agents that seam must be an interface
|
|
104
|
+
(local-fs, git, Console-remote, customer-supplied) so the engine backend is
|
|
105
|
+
swappable without touching adapters. This is the storage-port dependency of the
|
|
106
|
+
Console-as-backend rule (3).
|
|
107
|
+
- **An adapter conformance suite** (#502) makes rule 5 checkable — one shared spec
|
|
108
|
+
every adapter (harness or framework) must pass, with explicit gap declarations.
|
|
109
|
+
- **A first framework adapter proves the model** (#503). AWS Strands (the
|
|
110
|
+
non-terminal runtime that motivated this) is the reference in-process adapter:
|
|
111
|
+
it must emit canonical events in-process, honor redaction without a shell, and
|
|
112
|
+
pass the conformance suite. Its explicit callouts define the template for "what
|
|
113
|
+
a framework cannot do that a harness can."
|
|
114
|
+
- **`runtime-hook-surface.md` is promoted from a harness spec to the adapter
|
|
115
|
+
contract** every adapter category conforms to.
|
|
116
|
+
|
|
117
|
+
## Rationale
|
|
118
|
+
|
|
119
|
+
The owner's model — "harness adapters and framework adapters would work in a very
|
|
120
|
+
similar if not exactly the same way, such that a Flow or Flow-Agent-Kit
|
|
121
|
+
implementation is consistent across development workflows and customer agent
|
|
122
|
+
implementations" — only holds if there is a single engine and the runtimes are
|
|
123
|
+
adapters over it. The alternative (per-runtime reimplementations that happen to
|
|
124
|
+
agree) drifts the moment two runtimes are maintained by different hands, which is
|
|
125
|
+
precisely the failure the numbered-ADR redesign already diagnosed elsewhere in
|
|
126
|
+
this portfolio. DRY here is not a style preference; it is what makes "Flow behaves
|
|
127
|
+
the same everywhere" a checkable property (rule 5) instead of a marketing claim.
|
|
128
|
+
|
|
129
|
+
Making Console the remote backend for embedded agents (rule 3) is the same
|
|
130
|
+
git-authoritative / console-projection split already ratified for trust retention;
|
|
131
|
+
it costs nothing new conceptually and turns the dogfood console into the customer
|
|
132
|
+
product surface. The cooperative-vs-authoritative split (rule 4) is the only
|
|
133
|
+
honest enforcement story for code running in a process we do not own, and it is
|
|
134
|
+
already how delivery reconciliation works — so embedding does not weaken the trust
|
|
135
|
+
model, it inherits it.
|
|
136
|
+
|
|
137
|
+
The explicit-callout requirement is the guard against the seductive failure mode:
|
|
138
|
+
quietly letting a framework adapter skip a Flow feature because it was hard,
|
|
139
|
+
leaving customers with an inconsistent product and no signal. A conformance report
|
|
140
|
+
that names the gap keeps the promise of consistency honest.
|
|
141
|
+
|
|
142
|
+
## Open questions
|
|
143
|
+
|
|
144
|
+
- **Language of the first extracted core port** (TypeScript is the source-policy
|
|
145
|
+
default; the shell core stays as a binding). Tracked in the extract-core issue.
|
|
146
|
+
- **Conformance-suite substrate** — reuse the existing `evals/integration`
|
|
147
|
+
harness vs a new adapter-conformance package.
|
|
148
|
+
- **Storage-port interface shape** and whether the Console-remote binding reuses
|
|
149
|
+
the existing telemetry ingest or a dedicated port.
|
|
150
|
+
|
|
151
|
+
These are implementation decisions for the linked backlog; the direction above is
|
|
152
|
+
the fixed part.
|
package/docs/decisions/index.md
CHANGED
|
@@ -15,6 +15,7 @@ Numbered ADRs under `docs/adr/` are frozen history and are not listed here.
|
|
|
15
15
|
| [context-lifecycle](./context-lifecycle.md) | needs-decision | Context lifecycle |
|
|
16
16
|
| [core-domain-kit-boundary](./core-domain-kit-boundary.md) | needs-decision | Core vs domain kit boundary |
|
|
17
17
|
| [decision-records](./decision-records.md) | current | Decision records |
|
|
18
|
+
| [embeddable-engine](./embeddable-engine.md) | needs-decision | Embeddable engine and adapter model |
|
|
18
19
|
| [flow-flow-agents-boundary](./flow-flow-agents-boundary.md) | needs-decision | Flow / Flow Agents boundary |
|
|
19
20
|
| [flow-kit](./flow-kit.md) | needs-decision | Flow Kit |
|
|
20
21
|
| [flow-skill-kit-tool-boundary](./flow-skill-kit-tool-boundary.md) | needs-decision | Flow / Skill / Kit / Tool boundary |
|
|
@@ -30,7 +31,8 @@ Numbered ADRs under `docs/adr/` are frozen history and are not listed here.
|
|
|
30
31
|
| [promotion-gate](./promotion-gate.md) | current | Promotion gate |
|
|
31
32
|
| [standing-directives](./standing-directives.md) | current | Standing directives |
|
|
32
33
|
| [three-hard-boundary-model](./three-hard-boundary-model.md) | needs-decision | Three-hard-boundary model |
|
|
34
|
+
| [trust-ledger-retention](./trust-ledger-retention.md) | needs-decision | Trust-ledger retention and console-as-projection |
|
|
33
35
|
| [trust-reconcile](./trust-reconcile.md) | current | Trust-reconcile and delivery reconciliation |
|
|
34
36
|
| [typescript-source-policy](./typescript-source-policy.md) | current | TypeScript-first source policy |
|
|
35
|
-
| [workflow-enforcement](./workflow-enforcement.md) |
|
|
37
|
+
| [workflow-enforcement](./workflow-enforcement.md) | current | Workflow Enforcement |
|
|
36
38
|
| [workflow-trust-state](./workflow-trust-state.md) | needs-decision | Workflow trust state |
|
|
@@ -0,0 +1,88 @@
|
|
|
1
|
+
---
|
|
2
|
+
status: needs-decision
|
|
3
|
+
subject: Trust-ledger retention and console-as-projection
|
|
4
|
+
decided: 2026-07-07
|
|
5
|
+
evidence:
|
|
6
|
+
- kind: adr
|
|
7
|
+
ref: docs/adr/0017-anti-gaming-trust-security-model.md
|
|
8
|
+
- kind: adr
|
|
9
|
+
ref: docs/adr/0020-trust-reconcile-manifest-and-claim-classification.md
|
|
10
|
+
- kind: adr
|
|
11
|
+
ref: docs/adr/0022-fail-closed-delivery-reconciliation-with-governed-exemptions.md
|
|
12
|
+
- kind: doc
|
|
13
|
+
ref: docs/implementing-trust-reconciliation.md
|
|
14
|
+
- kind: issue
|
|
15
|
+
ref: kontourai/console#118
|
|
16
|
+
- kind: issue
|
|
17
|
+
ref: kontourai/console#125
|
|
18
|
+
- kind: issue
|
|
19
|
+
ref: kontourai/flow-agents#463
|
|
20
|
+
- kind: issue
|
|
21
|
+
ref: kontourai/flow-agents#73
|
|
22
|
+
---
|
|
23
|
+
# Trust-ledger retention and console-as-projection
|
|
24
|
+
|
|
25
|
+
> Status: **needs-decision** (proposed direction (shaped with Brian Anderson, 2026-07-07). Not yet
|
|
26
|
+
> fully built. Ratify + carry forward as ADR on implementation.
|
|
27
|
+
|
|
28
|
+
**Decision.** Git is the **authoritative store** for trust bundles and delivery
|
|
29
|
+
records; the Console (and any hosted DB projection) is a **rebuildable, queryable
|
|
30
|
+
projection** over them, never the source of truth. Three rules follow:
|
|
31
|
+
|
|
32
|
+
1. **Two stores, two retention tiers.** The authoritative bundle is committed
|
|
33
|
+
(`delivery/<slug>/trust.bundle` + checkpoint) and pinned to the merge commit —
|
|
34
|
+
permanent, distributed, surviving any console outage. The console DB keeps a
|
|
35
|
+
projection: **coarse audit records** (bundles, decisions, gate/claim outcomes)
|
|
36
|
+
retained long; the **fine-grained raw telemetry firehose** kept for a short
|
|
37
|
+
window and rolled up into aggregates. Retention in the projection is a **cost**
|
|
38
|
+
decision, not a **correctness** one.
|
|
39
|
+
|
|
40
|
+
2. **"Needs attention" is a view-filter, never a delete-sweep.** An item leaves
|
|
41
|
+
the operational attention set when it reaches a terminal state or ages past a
|
|
42
|
+
recency cutoff — the underlying record is retained, it just stops surfacing as
|
|
43
|
+
actionable. Deleting audit history to quiet a dashboard is a category error.
|
|
44
|
+
|
|
45
|
+
3. **The gate fires at PR time; post-merge the bundle is an audit record.** A
|
|
46
|
+
bundle seals to the work-commit and tolerates lag (the ancestor check + fresh
|
|
47
|
+
CI re-run absorb later commits); it is not regenerated per push. After a
|
|
48
|
+
(squash) merge it rides into the merge commit as a committed file and becomes
|
|
49
|
+
provenance, not enforcement. See `docs/implementing-trust-reconciliation.md`.
|
|
50
|
+
|
|
51
|
+
**Rationale.** "Console isn't the source of truth" is the *enabling* property, not
|
|
52
|
+
a caveat: because git holds the authority, the projection is free to be pruned,
|
|
53
|
+
tiered, or rebuilt for cost/perf without losing anything. This dissolves the
|
|
54
|
+
tension between "keep an auditable delivery history" and "don't let a dashboard
|
|
55
|
+
fill with days-old noise" — they are different concerns (retention vs. attention)
|
|
56
|
+
that were previously conflated in one store. Observed data motivates the tiering:
|
|
57
|
+
in one live snapshot the coarse audit table held ~38 rows while the raw
|
|
58
|
+
telemetry-event table held ~30,000 — bundles are not what grows; the event stream
|
|
59
|
+
is.
|
|
60
|
+
|
|
61
|
+
**A "stale" bundle is a presentation problem, not a data problem.** A bundle is an
|
|
62
|
+
immutable dated receipt ("verified this way at commit X"); it never becomes
|
|
63
|
+
*wrong*, only misleading if surfaced as *current* truth. The query layer answers
|
|
64
|
+
"true now?" from latest-state-per-subject and "true at commit X?" from the bundle.
|
|
65
|
+
|
|
66
|
+
**Consequences / what to build.**
|
|
67
|
+
|
|
68
|
+
- **Reframe the console "janitor" ([console#125]) from reaper to attention
|
|
69
|
+
view-filter**: terminal-state + recency cutoff over the operating-state
|
|
70
|
+
projection, with no deletion of underlying events. (Today the gap is
|
|
71
|
+
structural: there is no terminal-close/expiry event type, so nothing ever
|
|
72
|
+
leaves the active set — a process that stops emitting without a terminal status
|
|
73
|
+
flags as "long-running / needs attention" forever.)
|
|
74
|
+
- **Make the projection genuinely rebuildable** ([flow-agents#463],
|
|
75
|
+
[flow-agents#73]): add a backfill/import that re-hydrates the console from the
|
|
76
|
+
committed `delivery/*/trust.bundle` files across repos. Until this exists, "the
|
|
77
|
+
DB is just a cache of git" is false in practice — a wipe/restart loses
|
|
78
|
+
queryability until producers re-push (observed 2026-07-07: a hosted-DB wipe
|
|
79
|
+
required a manual service restart and had no re-hydration path).
|
|
80
|
+
- **Add retention + rollup for raw telemetry** (relates to the console noise
|
|
81
|
+
cleanup, [flow-agents#7-equivalent]): keep raw events N days, keep aggregates
|
|
82
|
+
long.
|
|
83
|
+
- **Trust-bundle registry** ([console#118]) is the research home for org-scoped
|
|
84
|
+
retention, sharing, and third-party verification of the ledger.
|
|
85
|
+
|
|
86
|
+
**Non-goals.** This decision does not change the layered-defense posture of ADR
|
|
87
|
+
0017/0020/0022, the PR-time gate semantics, or claim classification. It adds the
|
|
88
|
+
retention/projection design layer on top.
|
|
@@ -1,18 +1,40 @@
|
|
|
1
1
|
---
|
|
2
|
-
status:
|
|
2
|
+
status: current
|
|
3
3
|
subject: Workflow Enforcement
|
|
4
|
-
decided: 2026-07-
|
|
4
|
+
decided: 2026-07-10
|
|
5
5
|
evidence:
|
|
6
6
|
- kind: adr
|
|
7
7
|
ref: docs/adr/0001-flow-agents-consumes-flow.md
|
|
8
|
+
- kind: issue
|
|
9
|
+
ref: https://github.com/kontourai/flow-agents/issues/438
|
|
10
|
+
- kind: session-archive
|
|
11
|
+
ref: .kontourai/flow-agents/builder-enforcement-remediation/builder-enforcement-remediation--pull-work.md
|
|
8
12
|
---
|
|
9
13
|
# Workflow Enforcement
|
|
10
14
|
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
+
**Decision.** A new Workflow Run always starts at the first step declared by its canonical
|
|
16
|
+
Flow Definition. A caller cannot choose a later starting step, and metadata such as an ad-hoc
|
|
17
|
+
reason cannot grant that authority. Flow-native accepted exceptions may satisfy a named gate in
|
|
18
|
+
an existing run; they do not authorize skipping the prefix that precedes that gate.
|
|
15
19
|
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
20
|
+
Recovery resumes persisted run state rather than creating a replacement run at an arbitrary
|
|
21
|
+
step. The recovered state must validate against the canonical definition identity and preserve
|
|
22
|
+
its transitions, gate outcomes, attached evidence, and accepted exceptions. An import surface,
|
|
23
|
+
when provided, must validate the same complete durable record before making it active; it is not
|
|
24
|
+
a step-selection escape hatch.
|
|
25
|
+
|
|
26
|
+
Direct primitives remain valid outside a product Workflow Run. For example, standalone
|
|
27
|
+
`plan-work` may execute with its own primitive contract, but it must not stamp or report itself
|
|
28
|
+
as a completed `builder.build` prefix. A request that intends the Builder Kit product flow must
|
|
29
|
+
enter through `pull-work`, then `design-probe`, before planning.
|
|
30
|
+
|
|
31
|
+
**Rationale.** The Flow Definition is the executable ordering authority. Allowing callers to
|
|
32
|
+
self-declare a later step makes every upstream gate advisory and lets the artifact being created
|
|
33
|
+
justify its own creation. Separating standalone primitives from product-flow runs preserves
|
|
34
|
+
useful composability without weakening ordered workflows.
|
|
35
|
+
|
|
36
|
+
**Consequences.** `ensure-session --flow-id builder.build --step-id plan` is invalid for a new
|
|
37
|
+
run, even with `--ad-hoc-reason`. Existing run recovery is identified by durable run identity,
|
|
38
|
+
not by a reason string. Refusal must happen before session files, assignment claims, or other
|
|
39
|
+
side effects. Builder skills must route productized work through the declared prefix and describe
|
|
40
|
+
direct primitive use as standalone.
|
|
@@ -17,13 +17,16 @@ run `npm run validate:source --` and `npm run fixture:retirement-audit --`.
|
|
|
17
17
|
| `evals/fixtures/backlog-provider-settings` | settings precedence fixtures | `evals/integration/test_effective_backlog_settings.sh` | Keep while backlog provider settings resolution supports global defaults and project overrides. |
|
|
18
18
|
| `evals/fixtures/builder-kit-workflow-state` | Builder Kit workflow-state fixtures | `evals/static/test_workflow_skills.sh` | Keep while Builder Kit state contract and resume behavior are documented in workflow skill contracts. |
|
|
19
19
|
| `evals/fixtures/console-learning-projection` | console learning projection fixtures | `evals/integration/test_console_learning_projection.sh` | Keep while learning projection supports correction and open-route examples. |
|
|
20
|
+
| `evals/fixtures/economics` | per-run kit-economics record fixtures (#349): a transcript with .message.usage blocks, state.json/acceptance.json/critique.json join sources, a session.usage event, and the golden expected kontour.console.economics record | `evals/integration/test_economics_record.sh` | Keep while the `kontour.console.economics` v0.1 record contract (docs/specs/economics-record-contract.md) is enforced; the golden proves cost-from-transcript, defects-from-critique, verdict-from-state, and the R7 co-required guard. |
|
|
20
21
|
| `evals/fixtures/flow-kit-repository` | Flow Kit repository contract fixtures | `evals/integration/test_flow_kit_repository.sh`, `evals/integration/test_local_flow_kit_install.sh`, `evals/integration/test_runtime_adapter_activation.sh`, `evals/integration/test_activate_npx_context.sh`, `evals/integration/test_flow_kit_install_git.sh`, `evals/static/test_workflow_skills.sh` | Keep valid and invalid cases paired with the Flow Kit repository contract. |
|
|
21
22
|
| `evals/fixtures/kit-conformance-levels` | K-level conformance and consumer-target derivation fixtures | `evals/integration/test_kit_conformance_levels.sh` | Keep while K-level derivation, degradation invariant, and consumer-target badge rules are tested. |
|
|
22
23
|
| `evals/fixtures/hook-influence` | hook influence behavioral cases | `evals/integration/test_hook_influence_cases.sh`, `evals/static/test_workflow_skills.sh`, `scripts/validate-hook-influence-cases.js` | Keep while hook influence cases define agent guidance behavior. |
|
|
24
|
+
| `evals/fixtures/learning-review-proposals` | learning-review kit/gate tuning proposal fixtures (#352): pattern-present (engineered cost-inflation + gate false-block-rate pattern, with sessions/ trust.bundle + gate-review.inquiries.json joins and a hand-computed expected-aggregates.json), balanced (proportional cost/findings movement -> zero proposals), under-threshold (below LR_MIN_WINDOW_SAMPLE), repeat-window (idempotency), and effect-follow-up (later-window effect-fill pass for a ratified proposal) | `evals/integration/test_learning_review_proposals.sh` | Keep while `scripts/telemetry/learning-review-proposals.sh`/`learning-review-decide.sh` (docs/specs/learning-review-proposals-contract.md) are enforced: hand-computed aggregates, evidence-cited proposals, zero-mutation, idempotency, and the ratify -> effect-fill trail. |
|
|
23
25
|
| `evals/fixtures/pull-work-provider` | work item provider normalization fixtures | `evals/integration/test_pull_work_provider.sh` | Keep while provider normalization preserves blockers, artifact refs, board membership, and freshness metadata. |
|
|
24
26
|
| `evals/fixtures/pull-work-wip-shepherding` | WIP shepherding state fixtures | `evals/static/test_workflow_skills.sh` | Keep while pull-work documents personal versus global WIP behavior. |
|
|
25
27
|
| `evals/fixtures/reconcile-preflight` | #356 reconcile-preflight shape fixtures not already covered by trust-reconcile-exploits (un-superseded disputed critique, standalone disputed session-local claim) | `evals/integration/test_reconcile_preflight.sh` | Keep while the local `reconcile-preflight` subcommand (#356) is proven against the two shapes trust-reconcile-exploits does not already fixture (un-superseded disputed critique, standalone disputed session-local claim); the other four shapes reuse trust-reconcile-exploits/trust-reconcile-mixed-bundle directly rather than forking near-duplicates. |
|
|
26
28
|
| `evals/fixtures/surface-trust` | Surface trust evidence fixtures | `evals/integration/test_workflow_sidecar_writer.sh` | Keep while sidecar writer maps Surface trust evidence into workflow records. |
|
|
29
|
+
| `evals/fixtures/telemetry` | hermetic Stop-hook usage transcript fixture (#defect-2026-07 telemetry-usage-cost-extraction): a multi-model JSONL transcript with .message.model + .message.usage token blocks proving real token/cost/model extraction end-to-end | `evals/integration/test_usage_cost.sh`, `evals/integration/test_telemetry_usage_pipeline.sh` | Keep while telemetry.sh's Stop-hook usage pipeline (usage_parse_transcript / usage_model_from_transcript_usage) is proven against a hermetic multi-model transcript rather than relying on live runtime data. |
|
|
27
30
|
| `evals/fixtures/trust-reconcile-exploits` | WS8 trust-reconcile anti-gaming exploit fixtures (frozen negative regressions); also reused by the #356 local reconcile-preflight eval (same shapes, no forked copies) | `evals/integration/test_trust_reconcile_negatives.sh`, `evals/integration/test_reconcile_preflight.sh` | Keep while trust-reconcile.js enforces the WS8 iteration-2 soundness properties (no-label test_output, unwaived-assumed, status-misassertion, waiver-on-command); each fixture is a permanent negative regression. |
|
|
28
31
|
| `evals/fixtures/trust-reconcile-mixed-bundle` | WS8 trust-reconcile mixed-evidence end-to-end proof fixture; also reused by the #356 reconcile-preflight eval as its CLEAN-BUNDLE (AC4) case | `evals/integration/test_trust_reconcile_mixed_bundle.sh`, `evals/integration/test_reconcile_preflight.sh` | Keep while the trust-reconcile manifest/classification/waiver contract (ADR 0020) is enforced; proves a mixed test_output + session-local + waived bundle passes the CI anchor. |
|
|
29
32
|
| `evals/fixtures/trust-reconcile-ws3` | WS8 AC6 backward-compat fixture: real ws3-kit-dependencies-namespacing old-style bundle | `evals/integration/test_trust_reconcile_negatives.sh` | Keep while backward compatibility with pre-classification (all-test_output) bundles is asserted; proves an old-style bundle still FAILS the same way (no silent pass). |
|
|
@@ -0,0 +1,129 @@
|
|
|
1
|
+
# Implementing Trust Reconciliation Correctly
|
|
2
|
+
|
|
3
|
+
> Status: DRAFT for review (not yet a merged doc). Companion to
|
|
4
|
+
> ADR 0017 (anti-gaming trust/security model), ADR 0020 (reconcile manifest +
|
|
5
|
+
> claim classification), and ADR 0022 (fail-closed delivery reconciliation).
|
|
6
|
+
|
|
7
|
+
This guide is the "how to build it without re-breaking it" companion to the
|
|
8
|
+
trust ADRs. Every principle below either drew blood in a real delivery or is
|
|
9
|
+
load-bearing config in `.github/workflows/trust-reconcile.yml`. If you are
|
|
10
|
+
porting this pattern to another repo/org, read this first.
|
|
11
|
+
|
|
12
|
+
## The one sentence to internalize
|
|
13
|
+
|
|
14
|
+
**The trust gate fires at PR time; after merge, the same bundle becomes an
|
|
15
|
+
audit record, not a gate.**
|
|
16
|
+
|
|
17
|
+
That single framing decides where the check lives and what the artifact is
|
|
18
|
+
_for_ at each stage:
|
|
19
|
+
|
|
20
|
+
- **At PR time** the reconcile job is a required, admin-enforced status check.
|
|
21
|
+
It is the thing that prevents a bad merge.
|
|
22
|
+
- **After merge** the committed bundle is a dated, immutable receipt of what was
|
|
23
|
+
claimed and how it was verified — provenance, not enforcement.
|
|
24
|
+
|
|
25
|
+
Get this wrong and you either gate on the post-merge `main` push (which either
|
|
26
|
+
no-ops or falsely fails — see §2) or you misjudge what the bundle is worth once
|
|
27
|
+
the work has landed.
|
|
28
|
+
|
|
29
|
+
## Gate mechanics (the non-obvious CI gotchas)
|
|
30
|
+
|
|
31
|
+
### 1. CI must re-run verification fresh; the bundle is only a divergence detector
|
|
32
|
+
|
|
33
|
+
Never let the reconcile job _believe_ the bundle's "pass." Re-run the canonical
|
|
34
|
+
verification in a clean CI environment the agent does not control, and use the
|
|
35
|
+
bundle solely to detect divergence ("claimed pass, CI fails" / "claimed pass,
|
|
36
|
+
command CI never ran" / "checkpoint-only bundle"). If your job reads the bundle
|
|
37
|
+
and trusts it, you have built theater, not a gate. This is invariant #1.
|
|
38
|
+
|
|
39
|
+
### 2. Fail closed on ambiguity — which forces full git history
|
|
40
|
+
|
|
41
|
+
Ownership is decided by: _is the bundle's checkpoint commit a git-ancestor of
|
|
42
|
+
(or equal to) the change's HEAD?_ (`git merge-base --is-ancestor`). On a shallow
|
|
43
|
+
clone that check is unresolvable (exit 128) and MUST be treated as stale/fail,
|
|
44
|
+
never pass. Therefore the reconcile job's checkout needs full history
|
|
45
|
+
(`fetch-depth: 0`); the default shallow clone would falsely stale every
|
|
46
|
+
legitimate bundle. Cause and required config travel together.
|
|
47
|
+
|
|
48
|
+
### 3. Compare against the PR head SHA, not the synthetic merge SHA
|
|
49
|
+
|
|
50
|
+
On a `pull_request` trigger, `github.sha` is GitHub's ephemeral merge commit
|
|
51
|
+
(`refs/pull/N/merge`) — a commit no locally-sealed checkpoint ever stamps. Use
|
|
52
|
+
`pull_request.head.sha` for the ownership comparison (fall back to `github.sha`
|
|
53
|
+
only on `push`/`workflow_dispatch`, where it IS the real commit). Get this wrong
|
|
54
|
+
and every bundle falsely stales. Silent footgun; call it out.
|
|
55
|
+
|
|
56
|
+
### 4. Post-merge is a deliberate no-op — don't re-gate it
|
|
57
|
+
|
|
58
|
+
A squash-merge creates a new commit with no git ancestry back to the
|
|
59
|
+
feature-branch commit the checkpoint was sealed against. The reconcile job on
|
|
60
|
+
the post-merge `main` push must be a loud no-op: gating already happened at PR
|
|
61
|
+
time. Event-scope enforcement by trigger (`pull_request` gates; `push` to `main`
|
|
62
|
+
observes). Do not "strengthen" this into a main-branch gate — it will only ever
|
|
63
|
+
falsely fail on squash ancestry.
|
|
64
|
+
|
|
65
|
+
### 5. Bundles seal to the work-commit and tolerate lag — don't regenerate per push
|
|
66
|
+
|
|
67
|
+
The ancestor check (§2) plus the fresh re-run (§1) absorb new commits landing on
|
|
68
|
+
top of a sealed checkpoint. Re-seal only when the _claim set_ materially
|
|
69
|
+
changes, not on every commit. A bundle sealed at commit A is still valid at
|
|
70
|
+
HEAD B as long as A is an ancestor of B. Building "regenerate on every push"
|
|
71
|
+
flows is brittle and unnecessary.
|
|
72
|
+
|
|
73
|
+
## Claims (and the single worst trap)
|
|
74
|
+
|
|
75
|
+
### 6. Separate executable manifest commands from human-readable evidence — never `bash -lc` the prose
|
|
76
|
+
|
|
77
|
+
If you build any re-check/backstop layer, it must reconcile ONLY against
|
|
78
|
+
declared, runnable manifest commands. Attestation summaries, check descriptions,
|
|
79
|
+
and acceptance-criteria prose are DATA, not code. Re-executing a check's
|
|
80
|
+
_summary_ as `bash -lc "<summary text>"` yields garbage exit codes (2/127) and
|
|
81
|
+
false "caught false-completion" alarms — noise that can mask a real failure
|
|
82
|
+
sitting in the same run. This is the highest-frequency failure mode observed in
|
|
83
|
+
practice; guard against it explicitly.
|
|
84
|
+
|
|
85
|
+
### 7. `not_verified`-by-design is correct, not a gap
|
|
86
|
+
|
|
87
|
+
An agent that cannot run a command-backed check should record it at
|
|
88
|
+
`not_verified` with the exact reconcile-manifest command attached; CI runs it
|
|
89
|
+
for real and reconciles. This is the auditor-session pattern. Do not let agents
|
|
90
|
+
"helpfully" self-mark those `pass` — that reintroduces gaming. Require every
|
|
91
|
+
command-backed claim to either name its exact manifest command or be typed
|
|
92
|
+
`external` (an independent-review/human attestation that CI does not re-run).
|
|
93
|
+
|
|
94
|
+
## Storage & retention (git is the authority)
|
|
95
|
+
|
|
96
|
+
### 8. Two stores, two retention tiers
|
|
97
|
+
|
|
98
|
+
Git is the authority: the bundle is committed (`delivery/<slug>/trust.bundle`)
|
|
99
|
+
and pinned to the merge commit — permanent, distributed, survives any console
|
|
100
|
+
outage. The database/console is a _rebuildable, queryable projection_ over those
|
|
101
|
+
bundles. Keep coarse audit records (bundles, decisions, gate/claim outcomes)
|
|
102
|
+
long; expire and roll up the fine-grained raw telemetry firehose. Retention here
|
|
103
|
+
is a cost decision, not a correctness one.
|
|
104
|
+
|
|
105
|
+
Rule of thumb from real data: one delivery ≈ a handful of small bundle rows; a
|
|
106
|
+
single working session ≈ tens of thousands of raw tool-event rows. The bundles
|
|
107
|
+
are not what grows — the raw event stream is. Prune the firehose; keep the
|
|
108
|
+
ledger.
|
|
109
|
+
|
|
110
|
+
### 9. Make the projection actually rebuildable from the authority
|
|
111
|
+
|
|
112
|
+
If the console cannot re-hydrate itself from the committed bundles across repos,
|
|
113
|
+
then "the DB is just a cache of git" is false in practice — a restart/wipe loses
|
|
114
|
+
queryability until producers re-push. If you claim the projection is derived,
|
|
115
|
+
build the backfill/import from `delivery/*/trust.bundle` that proves it.
|
|
116
|
+
|
|
117
|
+
### Bonus: "stale as history" is not "stale as current state"
|
|
118
|
+
|
|
119
|
+
A bundle is a dated receipt; it never becomes _wrong_, it only misleads if you
|
|
120
|
+
surface an old one as _current_ truth. The query layer answers "true now?" from
|
|
121
|
+
latest-state-per-subject and "true at commit X?" from the bundle. The same
|
|
122
|
+
principle keeps the audit ledger from leaking into the operational
|
|
123
|
+
"needs-attention" view — attention is a view-filter (terminal-state + recency),
|
|
124
|
+
never a delete-sweep of history.
|
|
125
|
+
|
|
126
|
+
## The tagline
|
|
127
|
+
|
|
128
|
+
**The agent proposes, CI disposes, git remembers, the console indexes — and
|
|
129
|
+
nothing re-executes prose.**
|
package/docs/index.md
CHANGED
|
@@ -4,24 +4,24 @@ title: Kontour Flow Agents
|
|
|
4
4
|
|
|
5
5
|
# Flow Agents
|
|
6
6
|
|
|
7
|
-
<p class="home-lede">A portable process-discipline layer for agentic work:
|
|
7
|
+
<p class="home-lede">A portable process-discipline layer for agentic work: a kit-neutral engine for FlowDefinition interpretation, gates, runtime adapters, evidence, trust, and kit validation — plus opt-in kits such as Builder, Knowledge, Release Evidence, and Veritas Governance. Flow Agents keeps work inspectable so you ask for outcomes and the selected kit supplies the path, the state, the checks, and the proof.</p>
|
|
8
8
|
|
|
9
9
|
<div class="value-grid">
|
|
10
10
|
<section>
|
|
11
|
-
<strong>
|
|
12
|
-
<span>
|
|
11
|
+
<strong>Kit-neutral engine</strong>
|
|
12
|
+
<span>FlowDefinition interpretation, gates, runtime and harness adapters, SDK/evidence/trust primitives, and kit validation. The engine is not Builder Kit; Builder is one kit on top.</span>
|
|
13
13
|
</section>
|
|
14
14
|
<section>
|
|
15
15
|
<strong>Survive context loss</strong>
|
|
16
|
-
<span>Durable sidecar state under <code>.flow-agents/</code> records acceptance criteria, evidence, critique, and handoff, so any session resumes from recorded state instead of chat memory.</span>
|
|
16
|
+
<span>Durable sidecar state under <code>.kontourai/flow-agents/</code> records acceptance criteria, evidence, critique, and handoff, so any session resumes from recorded state instead of chat memory.</span>
|
|
17
17
|
</section>
|
|
18
18
|
<section>
|
|
19
|
-
<strong>
|
|
20
|
-
<span>
|
|
19
|
+
<strong>Tamper-evident trust</strong>
|
|
20
|
+
<span>Local capture is advisory and best-effort; the controlled CI re-run reconciles manifest commands and git diff as the authoritative anchor before evidence is treated as CI-verified.</span>
|
|
21
21
|
</section>
|
|
22
22
|
<section>
|
|
23
23
|
<strong>Flow Kits — workflow + output shape</strong>
|
|
24
|
-
<span>A kit bundles
|
|
24
|
+
<span>A kit bundles flows and optional skills, docs, adapters, evals, and assets as a validated, installable unit. The catalog lists Builder, Knowledge, Release Evidence, and Veritas Governance; <a href="kit-authoring-guide.html">bring your own kit</a> through the same manifest path.</span>
|
|
25
25
|
</section>
|
|
26
26
|
</div>
|
|
27
27
|
|
|
@@ -45,7 +45,9 @@ flowchart LR
|
|
|
45
45
|
Evidence -->|not verified| Plan
|
|
46
46
|
```
|
|
47
47
|
|
|
48
|
-
Flow Agents adds the operating layer around the model:
|
|
48
|
+
Flow Agents adds the operating layer around the model: the engine preserves state, evaluates evidence, activates selected kits, renders structured kit triggers, and compiles canonical policies to host hooks. The gate semantics underneath — definitions, runs, evidence, route-back — belong to <a href="https://kontourai.github.io/flow/">Kontour Flow</a>.
|
|
49
|
+
|
|
50
|
+
Kits supply the workflow. Builder and Knowledge are examples on the engine, not the engine itself. The built-in catalog also includes Release Evidence and Veritas Governance, both useful proof points that a kit can be agentless and still run through the same manifest and gate model.
|
|
49
51
|
|
|
50
52
|
## Process-discipline layer
|
|
51
53
|
|
|
@@ -78,7 +80,11 @@ Install into your workspace in one command:
|
|
|
78
80
|
npx @kontourai/flow-agents init --runtime <your-agent> --dest .
|
|
79
81
|
```
|
|
80
82
|
|
|
81
|
-
Where `--runtime` is `claude-code`, `codex`, `kiro`, `opencode`, or `pi`.
|
|
83
|
+
Where `--runtime` is `claude-code`, `codex`, `kiro`, `opencode`, or `pi`. Kits are opt-in. Activate Builder when you want its two gated flows: `builder.shape` (idea → slices → filed work items) and `builder.build` (selected work item → design probe → plan → execute → verify → PR → learn).
|
|
84
|
+
|
|
85
|
+
```bash
|
|
86
|
+
npx @kontourai/flow-agents init --runtime <your-agent> --dest . --activate-kit builder
|
|
87
|
+
```
|
|
82
88
|
|
|
83
89
|
Ask your agent to shape an idea:
|
|
84
90
|
|
|
@@ -110,6 +116,10 @@ Use fix-bug. Reproduce the problem, diagnose root cause, implement the fix, and
|
|
|
110
116
|
<strong>Builder Kit Quick Start</strong>
|
|
111
117
|
<span>Zero to a running, gated build flow in two minutes: install, shape an idea into a work item, build it through the builder.shape and builder.build flows, and see what the evidence gates do.</span>
|
|
112
118
|
</a>
|
|
119
|
+
<a class="doc-card" href="architecture-engine-and-kits.html">
|
|
120
|
+
<strong>Engine and Kits</strong>
|
|
121
|
+
<span>The canonical split: product-neutral engine, opt-in kits, catalog + kit.json plugin model, structured workflow triggers, and marketplace metadata with no runtime privilege.</span>
|
|
122
|
+
</a>
|
|
113
123
|
<a class="doc-card" href="workflow-usage-guide.html">
|
|
114
124
|
<strong>Workflow Usage Guide</strong>
|
|
115
125
|
<span>Every stage from shaping ideas to learning review, with example prompts and expected behavior.</span>
|