gentle-pi 2.4.0 → 2.6.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +292 -25
- package/assets/agents/gentle-ai-worker.md +13 -0
- package/assets/agents/jd-fix-agent.md +18 -0
- package/assets/agents/jd-judge-a.md +1 -1
- package/assets/agents/jd-judge-b.md +1 -1
- package/assets/agents/sdd-apply.md +7 -5
- package/assets/agents/sdd-archive.md +5 -3
- package/assets/agents/sdd-design.md +4 -0
- package/assets/agents/sdd-explore.md +4 -0
- package/assets/agents/sdd-init.md +4 -0
- package/assets/agents/sdd-onboard.md +4 -0
- package/assets/agents/sdd-proposal.md +4 -0
- package/assets/agents/sdd-remediate.md +37 -0
- package/assets/agents/sdd-research.md +26 -3
- package/assets/agents/sdd-spec.md +4 -0
- package/assets/agents/sdd-status.md +9 -75
- package/assets/agents/sdd-sync.md +4 -0
- package/assets/agents/sdd-tasks.md +4 -0
- package/assets/agents/sdd-verify.md +5 -3
- package/assets/chains/sdd-full.chain.md +4 -0
- package/assets/chains/sdd-plan.chain.md +4 -0
- package/assets/chains/sdd-verify.chain.md +4 -0
- package/assets/migrations/managed-assets-v2.5.0.json +7 -0
- package/assets/orchestrator-delegation.md +39 -11
- package/assets/orchestrator.md +5 -5
- package/assets/sdd-orchestrator-workflow.md +54 -21
- package/assets/support/sdd-status-contract.md +34 -90
- package/contracts/telemetry/runtime-aggregate-v1.schema.json +65 -0
- package/docs/delegated-verification.md +25 -0
- package/docs/telemetry.md +94 -0
- package/docs/windows-startup-console-visibility.md +18 -0
- package/extensions/ask-user-choice.ts +159 -25
- package/extensions/codegraph-tools.ts +95 -5
- package/extensions/gentle-agents.ts +1337 -0
- package/extensions/gentle-ai.ts +2916 -386
- package/extensions/gentle-shell.ts +650 -0
- package/extensions/gentle-todo.ts +234 -0
- package/extensions/quiet-tools.ts +2 -1
- package/extensions/runtime-metrics.ts +130 -0
- package/extensions/sdd-init.ts +2 -2
- package/extensions/startup-banner.ts +52 -75
- package/lib/agent-profiles.ts +550 -0
- package/lib/agents-completion-delivery.ts +72 -0
- package/lib/agents-config.ts +315 -0
- package/lib/agents-history.ts +88 -0
- package/lib/agents-messaging.ts +187 -0
- package/lib/agents-protocol.ts +501 -0
- package/lib/agents-runner.ts +1012 -0
- package/lib/agents-thread-view.ts +57 -0
- package/lib/agents-transcript.ts +87 -0
- package/lib/agents-view-layout.ts +40 -0
- package/lib/agents-view.ts +914 -0
- package/lib/agents-widget.ts +241 -0
- package/lib/gentle-ai-binary.ts +3 -1
- package/lib/gentle-ai-renderer.ts +143 -25
- package/lib/native-choice-list.ts +194 -0
- package/lib/native-fullscreen-interaction.ts +47 -0
- package/lib/native-pointer-region.ts +164 -0
- package/lib/native-review-cli.ts +371 -13
- package/lib/orchestrator-presence.ts +337 -0
- package/lib/profiles-orchestrator.ts +203 -0
- package/lib/review-candidate-view-owner.ts +427 -0
- package/lib/review-candidate-view.ts +150 -48
- package/lib/review-consent-component.ts +247 -0
- package/lib/review-consent-ui.ts +110 -0
- package/lib/review-host-relay.ts +28 -0
- package/lib/review-integration-v2.ts +243 -11
- package/lib/review-last-event-controller.ts +8 -4
- package/lib/review-relay-contract.ts +11 -0
- package/lib/review-reminder-receipt.ts +74 -0
- package/lib/review-repository.ts +2 -2
- package/lib/review-risk-assessment.ts +339 -0
- package/lib/review-session-standing-permission-ipc.ts +309 -0
- package/lib/review-session-standing-permission.ts +240 -0
- package/lib/runtime-metrics-children.ts +199 -0
- package/lib/runtime-metrics-delivery.ts +68 -0
- package/lib/runtime-metrics-native.ts +166 -0
- package/lib/runtime-metrics-pi-identity.ts +113 -0
- package/lib/runtime-metrics-policy.ts +51 -0
- package/lib/runtime-metrics.ts +255 -0
- package/lib/sdd-preflight.ts +362 -81
- package/lib/sdd-research-capabilities.ts +228 -0
- package/lib/sdd-status.ts +29 -7
- package/lib/session-worktree-registry.ts +118 -0
- package/lib/shell-bar.ts +184 -0
- package/lib/shell-card.ts +133 -0
- package/lib/shell-changes-view.ts +530 -0
- package/lib/shell-changes.ts +290 -0
- package/lib/shell-gauge.ts +40 -0
- package/lib/shell-prompt.ts +115 -0
- package/lib/shell-sidebar-banner.ts +11 -0
- package/lib/shell-sidebar-layout.ts +213 -0
- package/lib/shell-sidebar.ts +41 -0
- package/lib/shell-todo.ts +297 -0
- package/lib/shell-usage-view.ts +76 -0
- package/lib/shell-usage.ts +246 -0
- package/lib/telemetry-trigger.ts +153 -0
- package/package.json +8 -5
- package/runtime/gentle-ai-binary.mjs +3 -1
- package/runtime/native-review-cli.mjs +370 -12
- package/runtime/review-integration-v2.mjs +243 -11
- package/runtime/review-relay-contract.mjs +11 -0
- package/runtime/review-risk-assessment.mjs +340 -0
- package/runtime/telemetry-trigger.mjs +154 -0
- package/scripts/build-runtime-modules.mjs +11 -1
- package/scripts/check-types.mjs +125 -0
- package/scripts/gentle-ai-installer.mjs +10 -10
- package/scripts/install-gentle-ai.mjs +12 -0
- package/scripts/install-tui-mode-setting.mjs +114 -0
- package/scripts/test-packed-runner.mjs +38 -2
- package/scripts/types-baseline.json +99 -0
- package/scripts/verify-package-files.mjs +8 -2
- package/skills/_shared/review-ledger-contract.md +20 -2
- package/skills/issue-creation/SKILL.md +3 -3
- package/skills/judgment-day/SKILL.md +17 -3
- package/skills/judgment-day/references/prompts-and-formats.md +14 -3
- package/tests/agent-profiles.test.ts +722 -0
- package/tests/agents-completion-delivery.test.ts +94 -0
- package/tests/agents-config.test.ts +205 -0
- package/tests/agents-fake-child.ts +66 -0
- package/tests/agents-grouping.test.ts +179 -0
- package/tests/agents-history.test.ts +54 -0
- package/tests/agents-integration.test.ts +100 -0
- package/tests/agents-messaging.test.ts +94 -0
- package/tests/agents-protocol.test.ts +198 -0
- package/tests/agents-queries.test.ts +190 -0
- package/tests/agents-responsive.test.ts +43 -0
- package/tests/agents-runner-process.test.ts +111 -0
- package/tests/agents-runner.test.ts +959 -0
- package/tests/agents-thread-view.test.ts +45 -0
- package/tests/agents-transcript.test.ts +30 -0
- package/tests/agents-view.test.ts +685 -0
- package/tests/agents-widget.test.ts +141 -0
- package/tests/artifact-language.test.ts +25 -2
- package/tests/ask-user-choice.test.ts +325 -5
- package/tests/asset-installation-runtime.test.ts +108 -0
- package/tests/autonomous-guard.test.ts +116 -1
- package/tests/codegraph-tools.test.ts +112 -2
- package/tests/delegated-key-learnings-contract.test.ts +1 -1
- package/tests/devbinary/native-review-parity.devtest.ts +110 -0
- package/tests/feature-request-form.test.ts +67 -0
- package/tests/fixtures/agents-messaging-child.mjs +5 -0
- package/tests/fixtures/agents-process-child.mjs +23 -0
- package/tests/fixtures/runtime-metrics-native-batches.json +6 -0
- package/tests/gentle-agents.test.ts +2168 -0
- package/tests/gentle-ai-binary.test.ts +7 -2
- package/tests/gentle-ai-installer.test.ts +47 -47
- package/tests/gentle-ai-renderer.test.ts +103 -0
- package/tests/gentle-ai.test.ts +971 -15
- package/tests/gentle-card-text.ts +35 -0
- package/tests/gentle-shell.test.ts +818 -0
- package/tests/gentle-todo.test.ts +226 -0
- package/tests/install-tui-mode-setting.test.ts +324 -0
- package/tests/issue-creation-skill.test.ts +22 -0
- package/tests/model-routing-authority.test.ts +12 -0
- package/tests/native-choice-list.test.ts +202 -0
- package/tests/native-fullscreen-interaction.test.ts +125 -0
- package/tests/native-pointer-region.test.ts +245 -0
- package/tests/native-review-capability-contract.test.ts +27 -1
- package/tests/native-review-cli.test.ts +317 -3
- package/tests/native-review-consent.test.ts +91 -0
- package/tests/native-review-parity-runtime.test.ts +8 -2
- package/tests/native-review-parity.test.ts +43 -29
- package/tests/native-sdd-attempt-authority.test.ts +7 -2
- package/tests/orchestrator-budget.test.ts +69 -0
- package/tests/orchestrator-presence.test.ts +389 -0
- package/tests/orchestrator-rdd-ownership.test.ts +9 -0
- package/tests/package-manifest.test.ts +243 -7
- package/tests/profiles-orchestrator.test.ts +208 -0
- package/tests/quiet-tool-rendering.test.ts +97 -37
- package/tests/rdd-aware-verification-contract.test.ts +226 -0
- package/tests/rdd-status-line.test.ts +286 -0
- package/tests/review-agent-end-preflight.test.ts +332 -24
- package/tests/review-candidate-view.test.ts +751 -7
- package/tests/review-consent-ui.test.ts +352 -0
- package/tests/review-contract-prompt.test.ts +17 -0
- package/tests/review-controller-native-recovery.test.ts +29 -4
- package/tests/review-controller-native-routing.test.ts +884 -7
- package/tests/review-controller-workspace-root.test.ts +45 -2
- package/tests/review-controller.test.ts +26 -1
- package/tests/review-host-relay-restart-parity.test.ts +142 -1
- package/tests/review-host-relay-routing.test.ts +384 -8
- package/tests/review-host-relay.test.ts +29 -0
- package/tests/review-integration-v2-forward.test.ts +44 -0
- package/tests/review-integration-v2.test.ts +276 -0
- package/tests/review-last-event-closure.test.ts +112 -3
- package/tests/review-ledger-contract.test.ts +61 -6
- package/tests/review-relay-contract.test.ts +26 -0
- package/tests/review-reminder-receipt.test.ts +62 -0
- package/tests/review-repository.test.ts +28 -1
- package/tests/review-risk-assessment.test.ts +626 -0
- package/tests/review-session-standing-permission-controller.test.ts +656 -0
- package/tests/review-session-standing-permission-ipc.test.ts +233 -0
- package/tests/review-session-standing-permission-runtime.test.ts +212 -0
- package/tests/review-session-standing-permission.test.ts +156 -0
- package/tests/runtime-harness.mjs +447 -39
- package/tests/runtime-metrics-children.test.ts +206 -0
- package/tests/runtime-metrics-delivery.test.ts +85 -0
- package/tests/runtime-metrics-extension.test.ts +187 -0
- package/tests/runtime-metrics-native.test.ts +209 -0
- package/tests/runtime-metrics-pi-identity.test.ts +113 -0
- package/tests/runtime-metrics-policy.test.ts +62 -0
- package/tests/runtime-metrics.test.ts +184 -0
- package/tests/sdd-agent-tools.test.ts +10 -1
- package/tests/sdd-execution-routing-contract.test.ts +28 -0
- package/tests/sdd-managed-runtime-settlement.test.ts +331 -0
- package/tests/sdd-native-managed-uptake.test.ts +253 -0
- package/tests/sdd-planning-routing-contract.test.ts +45 -0
- package/tests/sdd-preflight.test.ts +252 -8
- package/tests/sdd-research-capabilities.test.ts +256 -0
- package/tests/sdd-research-live.test.ts +241 -0
- package/tests/sdd-selection-transport.test.ts +504 -0
- package/tests/sdd-status.test.ts +51 -0
- package/tests/session-worktree-registry.test.ts +135 -0
- package/tests/shell-bar.test.ts +176 -0
- package/tests/shell-card.test.ts +139 -0
- package/tests/shell-changes-view.test.ts +609 -0
- package/tests/shell-changes.test.ts +350 -0
- package/tests/shell-prompt.test.ts +140 -0
- package/tests/shell-sidebar-banner.test.ts +23 -0
- package/tests/shell-sidebar-layout.test.ts +387 -0
- package/tests/shell-sidebar.test.ts +50 -0
- package/tests/shell-todo.test.ts +259 -0
- package/tests/shell-usage-view.test.ts +62 -0
- package/tests/shell-usage.test.ts +197 -0
- package/tests/startup-banner.test.ts +126 -0
- package/tests/telemetry-trigger.test.ts +351 -0
package/assets/orchestrator.md
CHANGED
|
@@ -41,7 +41,7 @@ Route work through the smallest harness that is safe. Three tiers:
|
|
|
41
41
|
|
|
42
42
|
1. **Inline Direct** — small, mechanical, parent has context (typo, one-file edit, read-only check of 1-3 known files, bash for state). No SDD ceremony; stop when it is no longer small.
|
|
43
43
|
2. **Simple Delegation** — generic non-SDD exploration → `gentle-ai-explore`; bounded implementation → `gentle-ai-worker`; command-running generic non-SDD verification → `gentle-ai-verify`. Try its package role; if missing/unusable, use native `Agent` under the same read-only mapping/verification constraints and report fallback. SDD roles stay inside SDD.
|
|
44
|
-
3. **SDD (optional)** — selected only by an explicit request (`/gentle-sdd-new`/`/gentle-sdd-ff`/`/gentle-sdd-continue` or a direct ask) or an accepted proposal; size, file count, or risk alone never selects
|
|
44
|
+
3. **SDD (optional)** — selected only by an explicit request (`/gentle-sdd-new`/`/gentle-sdd-ff`/`/gentle-sdd-continue` or a direct ask) or an accepted proposal; size, file count, or risk alone never selects it. Suggest it when proposal/spec/design/tasks would meaningfully reduce ambiguity. Once selected, create artifacts and gate for approval before implementing.
|
|
45
45
|
|
|
46
46
|
## Delegation Rules
|
|
47
47
|
|
|
@@ -49,7 +49,7 @@ Core question: does this inflate parent context without need?
|
|
|
49
49
|
|
|
50
50
|
Before launching bounded writer (`gentle-ai-worker` or `worker`), task/context needs nonempty `## Allowed edit surfaces`: narrow repository-relative paths/globs; never `.`, bare repo root, or absolute. Parent derives surfaces, maps unknown targets read-only, shows derived candidates only for genuine scope choices. Do not ask the human to author paths or globs.
|
|
51
51
|
|
|
52
|
-
Mandatory Delegation Triggers —
|
|
52
|
+
Mandatory Delegation Triggers — once fired, delegate through the best available runtime (prefer `subagent_run`, else native `Agent`):
|
|
53
53
|
|
|
54
54
|
1. **4-file rule** — 4+ files to understand → delegate a scout/mapping task.
|
|
55
55
|
2. **Multi-file write rule** — 2+ non-trivial files touched → delegate one writer.
|
|
@@ -59,7 +59,7 @@ Mandatory Delegation Triggers — stop rules; once fired, delegate through the b
|
|
|
59
59
|
|
|
60
60
|
{{GENTLE_PI_BACKGROUND_POLICY}}; rules: the background-subagents block in the delegation contract.
|
|
61
61
|
|
|
62
|
-
|
|
62
|
+
Per-action table, Work Routing Ladder examples, Cost and Context Balance, Canonical Workflows, and the mirrored gentle-ai canon (blocking-prompt relays, language, delegation): `orchestrator-delegation.md`.
|
|
63
63
|
|
|
64
64
|
## SDD Workflow (lazy-loaded)
|
|
65
65
|
|
|
@@ -73,7 +73,7 @@ Hard preflight invariant: `openspec/config.yaml`, existing SDD changes, installe
|
|
|
73
73
|
|
|
74
74
|
## Memory Contract
|
|
75
75
|
|
|
76
|
-
When memory is available, the parent selects context and subagents save discoveries before returning. Phase table
|
|
76
|
+
When memory is available, the parent selects context and subagents save discoveries before returning. Phase table and artifact keys: `orchestrator-memory.md`.
|
|
77
77
|
|
|
78
78
|
## Skill Registry Protocol
|
|
79
79
|
|
|
@@ -89,7 +89,7 @@ This package injects the mirrored provider-bundle review execution contract into
|
|
|
89
89
|
|
|
90
90
|
## Safety
|
|
91
91
|
|
|
92
|
-
-
|
|
92
|
+
- An eligible interactive Pi host may resolve `gentle-ai.review-integration.consent/v3` before the envelope reaches the model. Permission: host-owned. If `gentle_review` returns the envelope unresolved, it is still the original provider-owned two-choice contract. Use `ask_user_choice` exactly or relay losslessly and stop. Never add the host action to a decoded or relayed provider envelope.
|
|
93
93
|
- Never commit unless the user explicitly asks.
|
|
94
94
|
- Ask before destructive git operations, publishing, or irreversible file changes.
|
|
95
95
|
- Keep writes single-threaded unless isolated worktrees are explicitly approved.
|
|
@@ -22,22 +22,37 @@ proposal → design ┘
|
|
|
22
22
|
|
|
23
23
|
## Native SDD Dispatcher
|
|
24
24
|
|
|
25
|
-
|
|
25
|
+
`gentle-ai sdd-status --contract gentle-ai.sdd-status/v2` is the sole, read-only status authority for every store. The orchestrator carries its projection unchanged; it never reconstructs readiness, selects a replacement action, uses an Engram bypass, or launches a recommendation merely because status displayed it.
|
|
26
26
|
|
|
27
|
-
|
|
27
|
+
`/gentle-sdd-status` only inspects and renders that projection. Only explicitly authorized `/gentle-sdd-continue` may call native `sdd-continue` to prepare a missing change-instance marker; this is not a status fallback and grants no source roots. If native status is unavailable, malformed, or mismatched, stop and report the failure.
|
|
28
28
|
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
29
|
+
## Bounded Planning Routing
|
|
30
|
+
|
|
31
|
+
For authoritative native status, route only by the bounded `nextRecommended` token and dependency states; never infer a route from prose. Keep genuine blockers in `blockedReasons` and non-blocking diagnostics in `notes`, never in `nextRecommended`, and report them without discarding them to enable a route.
|
|
32
|
+
|
|
33
|
+
| `nextRecommended` | Planning route |
|
|
34
|
+
| --- | --- |
|
|
35
|
+
| `propose` | `sdd-proposal` |
|
|
36
|
+
| `spec` | `sdd-spec` |
|
|
37
|
+
| `design` | `sdd-design` |
|
|
38
|
+
| `tasks` | `sdd-tasks` |
|
|
39
|
+
|
|
40
|
+
Native unprefixed tokens are the only automatic planning routes. Prefixed or locally derived status tokens never authorize a phase.
|
|
32
41
|
|
|
33
|
-
|
|
42
|
+
These planning routes remain runnable when missing planning artifacts leave `dependencies.apply: blocked`; do not require apply readiness to produce those artifacts. This is a planning-only exception, not permission to run apply or another blocked non-planning phase.
|
|
34
43
|
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
44
|
+
Before any planning launch, stop for ambiguous change selection, unresolved session preflight, or unsafe action context. Carry `actionContext` and prove planned writes are within the authoritative workspace or allowed edit roots; workspace-planning without allowed edit roots remains read-only. Planning does not bypass the init guard, pre-proposal gate, or phase approval requirements.
|
|
45
|
+
|
|
46
|
+
## Bounded Execution Routing
|
|
47
|
+
|
|
48
|
+
| Native `nextRecommended` | Pi executor |
|
|
49
|
+
| --- | --- |
|
|
50
|
+
| `apply` | `sdd-apply` |
|
|
51
|
+
| `verify` | `sdd-verify` |
|
|
52
|
+
| `remediate` | `sdd-remediate` |
|
|
53
|
+
| `archive` | `sdd-archive` |
|
|
54
|
+
|
|
55
|
+
Execute only the selected native action when its dependency and `actionContext` permit it. Unknown, malformed, blocked, or unsupported values stop before work; prose and local routing cannot replace them. `notes` is separate from `blockedReasons` and never gates: report a non-empty `notes` value as informational and proceed when the dependency and `blockedReasons` gates allow. Manual sdd-sync deliberately retains its local resolver and is never automatic native-status dispatch.
|
|
41
56
|
|
|
42
57
|
## SDD Status Contract
|
|
43
58
|
|
|
@@ -60,11 +75,11 @@ Do not ask SDD setup questions on session start. The first time the user initiat
|
|
|
60
75
|
|
|
61
76
|
**Hard gate:** `openspec/config.yaml`, existing SDD changes, installed `.pi`/global SDD assets, or a todo named "preflight" are not session preflight. They are project context only. Do not mark SDD preflight complete, start `sdd-init`, launch SDD subagents/chains, or move to explore/proposal/spec/design/tasks until this session has an injected `## SDD Session Preflight` block or an equivalent resolution from the canonical authority order below.
|
|
62
77
|
|
|
63
|
-
|
|
78
|
+
On the first SDD invocation of EACH new interactive session, confirm the preflight choices even when valid preferences are saved. Persisted preferences and canonical defaults are preselected suggestions, not current-session consent. Offer confirmation of the grouped suggestions or changes; cancellation leaves preflight unresolved. Explicit current-session choices take precedence and, once resolved, are reused throughout that session. If `/gentle:sdd-preflight` is unavailable, perform the same confirmation inline. The parent `subagent_run` dispatch boundary resolves this gate for every shipped SDD agent, prepends the exact rendered `## SDD Session Preflight` block to the existing child `context`, and blocks launch on cancellation or failure. An RPC child consumes that transport but never originates or persists defaults; missing or malformed transport fails closed before process spawn. Only a safely distinguishable standalone headless parent may retain canonical/persisted defaults without UI. Missing Engram constrains the artifact store to `openspec` unless an incompatible explicit request needs a human decision.
|
|
64
79
|
|
|
65
80
|
Preflight canonical defaults are execution `auto`, artifact store `openspec`, delivery strategy `ask-on-risk`, and review budget `400`; capability and already-selected constraints may narrow them.
|
|
66
81
|
|
|
67
|
-
|
|
82
|
+
The grouped session confirmation includes defaulted and capability-constrained suggestions. If changes are requested, preselect saved values and omit redundant one-option selectors. Never reinitialize project context merely because a new session needs confirmation. `chain_strategy` remains deferred, and `exception-ok` requires explicit `size:exception` acceptance and is never inferred.
|
|
68
83
|
|
|
69
84
|
The exact `delivery_strategy` domain accepted by `sdd-tasks` and `sdd-apply` is `ask-on-risk`, `auto-chain`, `single-pr`, or `exception-ok`; above the review threshold, `auto-chain` resolves without asking again.
|
|
70
85
|
|
|
@@ -130,7 +145,9 @@ This gate is MANDATORY and applies in both execution modes; in interactive mode
|
|
|
130
145
|
- The proposer receives a confirmed pre-proposal handoff and MUST NOT interview the user or infer consent.
|
|
131
146
|
- Pi's native `gentle-pi.sdd-status` contract remains the sole status contract. Research and pre-proposal state are orchestrator-owned prose and artifacts (`sdd/{change}/research`, `sdd/{change}/preproposal`, `openspec/changes/{change}/research.md`) layered on top — never a native status field.
|
|
132
147
|
|
|
133
|
-
Runtime
|
|
148
|
+
Runtime mapping: use the injected `## SDD Research Capabilities` resolved from package-approved exact tool names intersected with active tools. Official documentation requires only `fetch_content`; open-web requires ALL FOUR tools: `web_search`, `source_check`, `fetch_content`, and `get_search_content`, each active and approved/reachable in the child. None is optional; inventory admission is not evidence of execution. Preserve explicit agent/source restrictions. The research child receives only reachable approved names in its CLI allowlist and rechecks child-local availability. Generic `mcp` and dynamic `mcp__context7` gateways do not imply authorization for arbitrary servers or remote methods; without a verified narrow route they grant nothing.
|
|
149
|
+
|
|
150
|
+
Selected supported research MUST run and persist source-backed claims with exact tool calls, URLs, publisher/version, retrieval times, supporting excerpts and claim-to-source IDs. Tool inventory and search snippets are not evidence. Block only genuinely unavailable classes, retain partial results without unvalidated claims, and keep proposal readiness false until every selected class is complete. Never recommend skipping research because of a fictitious blanket restriction, invent citations, or substitute bash for missing tools. SDD chains treat research as unselected.
|
|
134
151
|
|
|
135
152
|
## Delivery Strategy
|
|
136
153
|
|
|
@@ -218,10 +235,10 @@ Never persist caller-authored attempt counters, tokens, or state in OpenSpec art
|
|
|
218
235
|
After the external run completes, call the compact settle with a request ID distinct from acquire, reusing an operation's own ID only for idempotent replay of that exact operation:
|
|
219
236
|
|
|
220
237
|
```text
|
|
221
|
-
gentle-ai sdd-attempt settle --cwd <repo> --change <change> --token <token> --request-id <id> --outcome <failed|interrupted|passed> --evidence-revision <sha256:...> --diagnosis <text> --harness-disposition <reused|invalidated> --cleanup-evidence <text> --process-evidence <text>
|
|
238
|
+
gentle-ai sdd-attempt settle --cwd <repo> --change <change> --token <token> --request-id <id> --outcome <failed|interrupted|passed> [--evidence-revision <sha256:...>] --diagnosis <text> --harness-disposition <reused|invalidated> --cleanup-evidence <text> --process-evidence <text>
|
|
222
239
|
```
|
|
223
240
|
|
|
224
|
-
Every settle field is required: `cwd`, `change`, `token`, `request-id`, `outcome`, `
|
|
241
|
+
Every settle field except `evidence-revision` is required: `cwd`, `change`, `token`, `request-id`, `outcome`, `diagnosis`, `harness-disposition`, `cleanup-evidence`, and `process-evidence`. For `failed` or `passed`, include `--evidence-revision` with the `sha256:...` evidence hash. For `interrupted`, omit the entire `--evidence-revision` flag. Pass `--successor-lineage` only for a distinct approved successor; the current/bound lineage remains itself otherwise. Pass `--remediates-evidence-revision` only when repairing a specific failed evidence revision. Settle derives binding and remediation inputs; the orchestrator never invents them.
|
|
225
242
|
|
|
226
243
|
`status`, `begin`, `finish`, and `reset` are diagnostic/compatibility surfaces, not the normal runtime route. Route continuation only from the provider-returned `proceed|blocked|complete`. `reset` is never automatic and requires an explicit maintainer scope decision.
|
|
227
244
|
|
|
@@ -256,7 +273,25 @@ On Pi, phase model routing is user-owned and persisted, not prompt-passed: `/gen
|
|
|
256
273
|
| jd-judge-a | deep-reasoning | Adversarial review |
|
|
257
274
|
| jd-judge-b | deep-reasoning | Adversarial review |
|
|
258
275
|
| jd-fix-agent | balanced | Surgical confirmed fixes |
|
|
259
|
-
| default | balanced | SDD
|
|
276
|
+
| default | balanced | SDD phase fallback; never a Judgment Day role |
|
|
277
|
+
|
|
278
|
+
## Judgment Day fix routing
|
|
279
|
+
|
|
280
|
+
Judgment Day phase roles are never generic fallbacks. If the generic writer chain is unavailable, use the documented native generic fallback or stop. Judgment Day is independent: it neither enables nor replaces ordinary review; a separately requested ordinary review remains independent. A standalone Judgment Day fix requires no graph-v1 or native review lineage. Launch `jd-fix-agent` only for an explicit Judgment Day fix batch with this exact runtime-accepted Markdown shape: `## Judgment Day activation` contains only `User explicitly requested Judgment Day.`. Replace the example ID, frozen ledger hash, row data, and surface with controller-authorized values. The correction batch contains only one round (`1 of 2` or `2 of 2`) and one lowercase SHA-256. The exact frozen finding rows are one JSON object per line, use only the canonical row fields, and exactly match the authorized IDs.
|
|
281
|
+
|
|
282
|
+
```markdown
|
|
283
|
+
## Judgment Day activation
|
|
284
|
+
User explicitly requested Judgment Day.
|
|
285
|
+
## Exact authorized severe IDs
|
|
286
|
+
- `JD-A-001`
|
|
287
|
+
## Judgment Day correction batch
|
|
288
|
+
Round: 1 of 2.
|
|
289
|
+
Frozen ledger SHA-256: `aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa`
|
|
290
|
+
## Exact frozen finding rows
|
|
291
|
+
{"id":"JD-A-001","lens":"judgment-day","location":"path/to/authorized-file.ts:1","severity":"CRITICAL","status_at_freeze":"open","evidence_class":"deterministic","evidence_claim":"Concrete user-impact claim supported by the frozen location."}
|
|
292
|
+
## Allowed edit surfaces
|
|
293
|
+
path/to/authorized-file.ts
|
|
294
|
+
```
|
|
260
295
|
|
|
261
296
|
## Sub-Agent Launch Deduplication
|
|
262
297
|
|
|
@@ -310,9 +345,7 @@ Automatic mode does not override reviewer burnout protection.
|
|
|
310
345
|
|
|
311
346
|
## Recovery
|
|
312
347
|
|
|
313
|
-
|
|
314
|
-
- `openspec` → read `openspec/changes/<change>/` artifacts and re-derive readiness through the native status engine.
|
|
315
|
-
- `none` → state is not persisted; explain the limitation.
|
|
348
|
+
For every store, request a fresh native v2 status projection. Artifact reads may supply phase inputs only after native selection; they never re-derive readiness, replace status, or bypass native refusal. Manual sdd-sync keeps its separate local resolver.
|
|
316
349
|
|
|
317
350
|
## Provider Defect Handoff
|
|
318
351
|
|
|
@@ -15,94 +15,43 @@ Any phase that selects, continues, applies, verifies, syncs, or archives an SDD
|
|
|
15
15
|
|
|
16
16
|
## Native Engine
|
|
17
17
|
|
|
18
|
-
-
|
|
19
|
-
-
|
|
20
|
-
-
|
|
21
|
-
-
|
|
22
|
-
-
|
|
23
|
-
|
|
24
|
-
|
|
18
|
+
- `gentle-ai sdd-status --contract gentle-ai.sdd-status/v2` is the sole status authority for every store. It is read-only: inspect its native projection unchanged and never launch a phase, prepare consent, or grant roots while reading it.
|
|
19
|
+
- If native status is unavailable, malformed, or does not select the requested change/workspace, stop and report that failure. Do not construct a local status, infer readiness from artifacts, substitute continuation, or bypass it through Engram.
|
|
20
|
+
- `nextRecommended`, `dependencies`, `blockedReasons`, `actionContext`, and optional `phaseInstructions` are producer facts. Route only by their typed values, never by prose or a local lifecycle graph. A genuine blocker's human-readable explanation belongs in `blockedReasons`; a non-blocking diagnostic belongs in `notes`; neither belongs in `nextRecommended`.
|
|
21
|
+
- Runtime-attempt authority is separate from status: runtime-bearing work uses the provider `sdd-attempt acquire|settle` flow and its `proceed`, `blocked`, or `complete` result.
|
|
22
|
+
- Only an explicitly authorized `gentle-ai sdd-continue` may prepare a missing change-instance marker. `ensureChangeInstanceMarker` has no status caller; its sole production path is `PrepareChangeInstanceConsent` through `sdd-continue`.
|
|
23
|
+
|
|
24
|
+
## Bounded Planning Routing
|
|
25
|
+
|
|
26
|
+
For authoritative native status, route only by the bounded `nextRecommended` token and dependency states; never infer a route from prose. Keep genuine blockers in `blockedReasons` and non-blocking diagnostics in `notes`, never in `nextRecommended`, and report them without discarding them to enable a route.
|
|
27
|
+
|
|
28
|
+
| `nextRecommended` | Planning route |
|
|
29
|
+
| --- | --- |
|
|
30
|
+
| `propose` | `sdd-proposal` |
|
|
31
|
+
| `spec` | `sdd-spec` |
|
|
32
|
+
| `design` | `sdd-design` |
|
|
33
|
+
| `tasks` | `sdd-tasks` |
|
|
34
|
+
|
|
35
|
+
Native unprefixed tokens are the only automatic planning routes. Prefixed or locally derived status tokens never authorize a phase.
|
|
36
|
+
|
|
37
|
+
These planning routes remain runnable when missing planning artifacts leave `dependencies.apply: blocked`; do not require apply readiness to produce those artifacts. This is a planning-only exception, not permission to run apply or another blocked non-planning phase.
|
|
38
|
+
|
|
39
|
+
Before any planning launch, stop for ambiguous change selection, unresolved session preflight, or unsafe action context. Carry `actionContext` and prove planned writes are within the authoritative workspace or allowed edit roots; workspace-planning without allowed edit roots remains read-only. Planning does not bypass the init guard, pre-proposal gate, or phase approval requirements.
|
|
40
|
+
|
|
41
|
+
## Bounded Execution Routing
|
|
42
|
+
|
|
43
|
+
| Native `nextRecommended` | Pi executor |
|
|
44
|
+
| --- | --- |
|
|
45
|
+
| `apply` | `sdd-apply` |
|
|
46
|
+
| `verify` | `sdd-verify` |
|
|
47
|
+
| `remediate` | `sdd-remediate` |
|
|
48
|
+
| `archive` | `sdd-archive` |
|
|
49
|
+
|
|
50
|
+
Execute only a native selected action whose dependency and `actionContext` permit it. Unknown, malformed, blocked, or unsupported actions stop before work; no local route, prefixed token, or prose can replace them. `notes` is separate from `blockedReasons` and never gates: report a non-empty `notes` value as informational and proceed when the dependency and `blockedReasons` gates allow. Manual `sdd-sync` remains its intentional local resolver and is never an automatic native-status dispatch.
|
|
25
51
|
|
|
26
52
|
## Status Schema
|
|
27
53
|
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
```yaml
|
|
31
|
-
schemaName: spec-driven
|
|
32
|
-
changeName: <change-name>
|
|
33
|
-
artifactStore: openspec | engram | both | none
|
|
34
|
-
planningHome:
|
|
35
|
-
root: <project-or-openspec-root>
|
|
36
|
-
changesDir: <openspec/changes or memory topic prefix>
|
|
37
|
-
changeRoot: <openspec/changes/<change> or memory topic prefix>
|
|
38
|
-
artifactPaths:
|
|
39
|
-
proposal: [<path-or-topic>]
|
|
40
|
-
specs: [<path-or-topic>]
|
|
41
|
-
design: [<path-or-topic>]
|
|
42
|
-
tasks: [<path-or-topic>]
|
|
43
|
-
applyProgress: [<path-or-topic>]
|
|
44
|
-
verifyReport: [<path-or-topic>]
|
|
45
|
-
syncReport: [<path-or-topic>]
|
|
46
|
-
contextFiles:
|
|
47
|
-
proposal: [<concrete readable files/topics>]
|
|
48
|
-
specs: [<concrete readable files/topics>]
|
|
49
|
-
design: [<concrete readable files/topics>]
|
|
50
|
-
tasks: [<concrete readable files/topics>]
|
|
51
|
-
applyProgress: [<concrete readable files/topics>]
|
|
52
|
-
verifyReport: [<concrete readable files/topics>]
|
|
53
|
-
syncReport: [<concrete readable files/topics>]
|
|
54
|
-
artifacts:
|
|
55
|
-
proposal: missing | done | partial
|
|
56
|
-
specs: missing | done | partial
|
|
57
|
-
design: missing | done | partial
|
|
58
|
-
tasks: missing | done | partial
|
|
59
|
-
applyProgress: missing | done | partial
|
|
60
|
-
verifyReport: missing | done | partial
|
|
61
|
-
syncReport: missing | done | partial
|
|
62
|
-
taskProgress: # implementation-owned plus malformed unresolved rows
|
|
63
|
-
total: 0
|
|
64
|
-
complete: 0
|
|
65
|
-
remaining: 0
|
|
66
|
-
unchecked: []
|
|
67
|
-
deferredParentActions:
|
|
68
|
-
total: 0
|
|
69
|
-
complete: 0
|
|
70
|
-
remaining: 0
|
|
71
|
-
unchecked: []
|
|
72
|
-
taskArtifactErrors: []
|
|
73
|
-
applyState: blocked | all_done | ready | not_applicable
|
|
74
|
-
dependencies:
|
|
75
|
-
apply: blocked | ready | all_done | not_applicable
|
|
76
|
-
verify: blocked | ready | all_done | not_applicable
|
|
77
|
-
sync: blocked | ready | all_done | not_applicable
|
|
78
|
-
archive: blocked | ready | all_done | not_applicable
|
|
79
|
-
actionContext:
|
|
80
|
-
mode: repo-local | workspace-planning
|
|
81
|
-
workspaceRoot: <absolute path>
|
|
82
|
-
allowedEditRoots: [<absolute paths>]
|
|
83
|
-
warnings: []
|
|
84
|
-
nextRecommended: <bounded-machine-token>
|
|
85
|
-
isNonAuthoritative: false # boolean; true when the native engine is not authoritative for the store
|
|
86
|
-
```
|
|
87
|
-
|
|
88
|
-
## Task Ownership
|
|
89
|
-
|
|
90
|
-
New task checkboxes end with the terminal marker `<!-- sdd-owner: implementation -->`. An unmarked legacy checkbox is implementation-owned. Supported legacy non-implementation rows are informational only. Any line containing `sdd-owner` that is unsupported, duplicated, or non-terminal is malformed: add its exact line to `taskArtifactErrors` and `blockedReasons`, and count it as unresolved implementation work even when checked. `taskProgress` reports implementation work.
|
|
91
|
-
|
|
92
|
-
## Apply State
|
|
93
|
-
|
|
94
|
-
- `blocked`: required apply artifacts are missing, task selection is ambiguous, malformed ownership markers exist, or action context makes edits unsafe.
|
|
95
|
-
- `all_done`: tasks artifact exists and every implementation task is checked `[x]`.
|
|
96
|
-
- `ready`: tasks artifact exists, at least one implementation task remains unchecked, and edit scope is safe.
|
|
97
|
-
- `not_applicable`: emitted for non-authoritative stores (see Engine Authority by Store). This is NOT a blocker.
|
|
98
|
-
|
|
99
|
-
## Dependency States
|
|
100
|
-
|
|
101
|
-
- `apply` is `ready` only when specs, design, and tasks are available and task progress is not all done.
|
|
102
|
-
- `verify` is ready after implementation completion when tasks are complete or apply-progress exists. RDD authority and receipts never gate the apply -> verify -> sync -> archive route. Unchecked implementation tasks remain CRITICAL blockers for full archive readiness.
|
|
103
|
-
- `sync` is `ready` only when verify-report exists and has no unresolved `FAIL`, `BLOCKED`, `CRITICAL`, or verification blockers. `engram`/`none` modes may mark sync `not_applicable`.
|
|
104
|
-
- `archive` is `ready` only when verify-report exists, sync is complete or not applicable, and implementation tasks are complete. CRITICAL verification issues have no override. Explicit recorded exceptions are limited to non-critical partial archives or stale-checkbox reconciliation when apply-progress/verify-report prove completion.
|
|
105
|
-
- `not_applicable`: emitted for non-authoritative stores (engram, none, and both when no `openspec/` directory exists) when `nextRecommended: "resolve-via-engram"` is active. `not_applicable` is NOT a gate failure — readiness must be resolved from Engram instead of from these fields.
|
|
54
|
+
Consume the native v2 projection (`schemaName: gentle-ai.sdd-status`, `schemaVersion: 2`) losslessly. Its producer-defined selection, artifact locators, task progress, seven dependencies, `actionContext`, `blockedReasons`, optional execution instructions, remediation state, and `nextRecommended` are status facts, not a Pi schema to recreate.
|
|
106
55
|
|
|
107
56
|
## Action Context Guard
|
|
108
57
|
|
|
@@ -112,11 +61,6 @@ The orchestrator MUST carry `actionContext` into any phase launch.
|
|
|
112
61
|
- If `allowedEditRoots` is present, only edit or move files within those roots.
|
|
113
62
|
- If a phase cannot prove a file is inside the authoritative workspace or allowed edit roots, stop and ask for clarification.
|
|
114
63
|
|
|
115
|
-
## Engine Authority by Store
|
|
116
|
-
|
|
117
|
-
- `openspec` and `both` (when `openspec/` directory exists): the local SDD status engine resolves artifact state from disk and is authoritative. Phase executors must obey it.
|
|
118
|
-
- `engram`, `none`, and `both` (when `openspec/` directory does NOT exist): the local engine cannot read Engram artifacts. It returns `nextRecommended: "resolve-via-engram"` and empty `blockedReasons`. This output is **non-authoritative**. The orchestrator must resolve readiness directly from Engram using the Engram memory tools injected by the memory provider on the change topic keys (`sdd/{change-name}/proposal`, `sdd/{change-name}/spec`, etc.) instead of relying on the engine's dependency states. The `artifactStore` field still reflects the real chosen store value (e.g. `"both"`) and must not be rewritten.
|
|
119
|
-
|
|
120
64
|
## Native Runtime Attempt Authority
|
|
121
65
|
|
|
122
66
|
The compact SDD runtime attempt authority is separate from artifact dispatch and status. It is artifact-store agnostic: the same acquire/settle discipline applies to `openspec`, `engram`, `both`, and `none` stores. Its payload MUST NOT be embedded in the SDD v1 status schema above; status reports artifact state only, never attempt tokens or attempt counters. No OpenSpec or Engram attempt ledger may be created or mirrored by Pi.
|
|
@@ -0,0 +1,65 @@
|
|
|
1
|
+
{
|
|
2
|
+
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
|
3
|
+
"title": "gentle-ai.telemetry-runtime-aggregate/v1",
|
|
4
|
+
"description": "Sanitized stdin for gentle-ai telemetry runtime send --json. Maximum UTF-8 input 16384 bytes. One best-effort send attempt, no client persistence or retry. Rows are independent observations, not a session composition. No caller identity is accepted.",
|
|
5
|
+
"type": "object", "additionalProperties": false,
|
|
6
|
+
"required": ["schema", "registry", "host", "rows"],
|
|
7
|
+
"properties": {
|
|
8
|
+
"schema": {"const": "gentle-ai.telemetry-runtime-aggregate/v1"},
|
|
9
|
+
"registry": {"const": 1},
|
|
10
|
+
"host": {"enum": ["claude-code", "opencode", "codex", "pi"]},
|
|
11
|
+
"rows": {"type": "array", "minItems": 1, "maxItems": 32, "items": {"$ref": "#/$defs/row"}}
|
|
12
|
+
},
|
|
13
|
+
"$defs": {
|
|
14
|
+
"metric": {"oneOf": [{"type": "integer", "minimum": 0, "maximum": 999999999999}, {"type": "null"}, {"const": "unsupported"}]},
|
|
15
|
+
"count": {"type": "integer", "minimum": 0, "maximum": 999999999999},
|
|
16
|
+
"token": {
|
|
17
|
+
"type": "object", "additionalProperties": false,
|
|
18
|
+
"required": ["reported", "unavailable", "unsupported", "sum"],
|
|
19
|
+
"properties": {
|
|
20
|
+
"reported": {"$ref": "#/$defs/count"}, "unavailable": {"$ref": "#/$defs/count"},
|
|
21
|
+
"unsupported": {"$ref": "#/$defs/count"}, "sum": {"$ref": "#/$defs/count"}
|
|
22
|
+
},
|
|
23
|
+
"if": {"properties": {"reported": {"const": 0}}},
|
|
24
|
+
"then": {"properties": {"sum": {"const": 0}}}
|
|
25
|
+
},
|
|
26
|
+
"duration": {
|
|
27
|
+
"type": "object", "additionalProperties": false, "required": ["kind", "measured_count", "sum_ms"],
|
|
28
|
+
"properties": {"kind": {}, "measured_count": {}, "sum_ms": {}},
|
|
29
|
+
"oneOf": [
|
|
30
|
+
{"properties": {"kind": {"const": "unavailable"}, "measured_count": {"const": 0}, "sum_ms": {"type": "null"}}},
|
|
31
|
+
{"properties": {"kind": {"enum": ["request", "message"]}, "measured_count": {"type": "integer", "minimum": 1, "maximum": 999999999999}, "sum_ms": {"type": "number", "minimum": 0, "maximum": 999999999999}}}
|
|
32
|
+
]
|
|
33
|
+
},
|
|
34
|
+
"effort": {"enum": ["off", "minimal", "low", "medium", "high", "xhigh", "max", "not_selected", "unknown", "custom", "unavailable", "unsupported"]},
|
|
35
|
+
"model": {
|
|
36
|
+
"type": "object", "additionalProperties": false, "required": ["provider", "id"],
|
|
37
|
+
"properties": {"provider": {"type": "string"}, "id": {"type": "string"}},
|
|
38
|
+
"oneOf": [
|
|
39
|
+
{"properties": {"provider": {"const": "anthropic"}, "id": {"enum": ["claude-opus-5", "claude-haiku-4-5", "claude-haiku-4-5-20251001", "claude-sonnet-5"]}}},
|
|
40
|
+
{"properties": {"provider": {"enum": ["openai", "openai-codex"]}, "id": {"enum": ["gpt-6-astra", "gpt-5.6-sol", "gpt-5.6-terra", "gpt-5.6-luna", "gpt-5.3-codex-spark", "gpt-5.5", "gpt-5.4", "gpt-5.4-mini", "gpt-5.2", "gpt-5.3-codex", "gpt-5.6"]}}},
|
|
41
|
+
{"properties": {"provider": {"const": "opencode"}, "id": {"const": "custom"}}},
|
|
42
|
+
{"properties": {"provider": {"const": "unknown"}, "id": {"const": "unknown"}}},
|
|
43
|
+
{"properties": {"provider": {"const": "custom"}, "id": {"const": "custom"}}}
|
|
44
|
+
]
|
|
45
|
+
},
|
|
46
|
+
"row": {
|
|
47
|
+
"type": "object", "additionalProperties": false,
|
|
48
|
+
"required": ["model", "model_evidence", "agent_kind", "agent_class", "selected_effort", "effective_effort", "launches", "responses", "input_tokens", "output_tokens", "cache_read_tokens", "cache_creation_tokens", "reasoning_tokens", "total_tokens", "error_category", "duration"],
|
|
49
|
+
"properties": {
|
|
50
|
+
"reasoning_tokens": {"$ref": "#/$defs/token"},
|
|
51
|
+
"total_tokens": {"$ref": "#/$defs/token"},
|
|
52
|
+
"agent_class": {"enum": ["orchestrator", "worker", "explore", "verify", "unknown", "sdd-init", "sdd-explore", "sdd-research", "sdd-propose", "sdd-spec", "sdd-design", "sdd-tasks", "sdd-apply", "sdd-verify", "sdd-archive", "sdd-onboard", "sdd-status", "sdd-sync", "jd-judge-a", "jd-judge-b", "jd-fix-agent", "review-risk", "review-readability", "review-reliability", "review-resilience", "review-refuter", "review-validator"]},
|
|
53
|
+
"error_category": {"enum": ["none", "unknown", "auth", "output_length", "aborted", "api", "rate_limit", "server"]},
|
|
54
|
+
"duration": {"$ref": "#/$defs/duration"},
|
|
55
|
+
"model": {"$ref": "#/$defs/model"},
|
|
56
|
+
"model_evidence": {"enum": ["selected", "response", "unknown"]},
|
|
57
|
+
"agent_kind": {"enum": ["orchestrator", "built_in", "custom", "unknown"]},
|
|
58
|
+
"selected_effort": {"$ref": "#/$defs/effort"}, "effective_effort": {"$ref": "#/$defs/effort"},
|
|
59
|
+
"launches": {"$ref": "#/$defs/metric", "description": "Source-observed event count only; null when not observed. Never reconstructed from sessions."}, "responses": {"$ref": "#/$defs/metric", "description": "Source-observed response occurrence count, independent of token coverage and success."},
|
|
60
|
+
"input_tokens": {"$ref": "#/$defs/token"}, "output_tokens": {"$ref": "#/$defs/token"},
|
|
61
|
+
"cache_read_tokens": {"$ref": "#/$defs/token"}, "cache_creation_tokens": {"$ref": "#/$defs/token"}
|
|
62
|
+
}
|
|
63
|
+
}
|
|
64
|
+
}
|
|
65
|
+
}
|
|
@@ -0,0 +1,25 @@
|
|
|
1
|
+
# Delegated verification
|
|
2
|
+
|
|
3
|
+
How the Gentle Pi orchestrator decides who verifies a bounded writer's work. The always-on parent prompt renders a `Receipt-driven development: on|off|unknown` line; the delegation overlay (`assets/orchestrator-delegation.md`, trigger 5) keys the verification rule on it. This page is package-owned; `docs/review-integration.md` mirrors the Gentle AI contract and must stay byte-identical to it.
|
|
4
|
+
|
|
5
|
+
## Receipt-driven development on
|
|
6
|
+
|
|
7
|
+
The bounded writer runs the exact commands the parent lists under `## Verification`, in the foreground, and reports each as `<command>: <observed result>`. That report is the verification of record and the native review is the independent check. `gentle-ai-verify` is on-demand: a `partial` or `blocked` writer, an expensive or external check the parent wants on a cheaper profile, or a parent spot check.
|
|
8
|
+
|
|
9
|
+
This `on` path holds only while the native review actually reaches a terminal outcome for the current candidate (gentle-pi#668). A human decline of the consent envelope for this candidate (candidate-scoped, never the RDD kill switch), a clone-local RDD disable discovered mid-flow, or a refused START/STATUS all mean the review never ran, so the parent falls back to the exact risk-gated path below, as if RDD were `off` -- declining a review never lowers the bar below the RDD-off path. `gentle_review`'s `assess` operation accepts an optional `nativeReviewOutcome` (`closed`, `declined`, `unavailable`, or `unknown`) so the caller can state this directly. `closed` is never auto-derived: only a caller that itself just acknowledged the approved review for this exact candidate may pass it, right after that acknowledgement. When `nativeReviewOutcome` is omitted, `assess` only ever tries to auto-derive `declined`/`unavailable`, and only for the exact candidate the event was bound to -- keyed by that candidate's own target identity, never by repository alone, so one candidate's recorded outcome can never leak into a different candidate's `assess` call in the same clone. A missing or mismatched identity fails closed to `unknown`, verified exactly like `off`. The result's `outcome_source` (`explicit`, `derived`, or `unknown`) states which of these produced the value, so a stale or missing derivation is visible rather than silently indistinguishable from a real `unknown`.
|
|
10
|
+
|
|
11
|
+
## Receipt-driven development off or unknown (gentle-pi#662)
|
|
12
|
+
|
|
13
|
+
The host exposes one read-only native operation: `gentle-ai review assess --cwd <repo> [--base-ref <ref> --committed-only] --json` (gentle-ai#4295). It is decoded by `lib/review-risk-assessment.ts` and wired through `lib/native-review-cli.ts` exactly like the existing `reviewMode` STATUS reader -- a bounded subprocess with a typed decode, never a mutation. A non-zero exit, a failure envelope, or an older binary without the verb all fail closed to `high` risk.
|
|
14
|
+
|
|
15
|
+
The `gentle_review` tool's `assess` operation (`extensions/gentle-ai.ts`) combines that assessment with the rendered `Receipt-driven development:` line to decide whether a delegated writer's change needs a separate `gentle-ai-verify` run, following this tier table:
|
|
16
|
+
|
|
17
|
+
| Native risk tier | Verification when RDD is `off`/`unknown` |
|
|
18
|
+
| --- | --- |
|
|
19
|
+
| passive | structural readback by the parent; no separate verifier, no tests |
|
|
20
|
+
| medium | writer self-verification stands; a separate `gentle-ai-verify` run is added only when the writer profile is a small model (mini or low effort) |
|
|
21
|
+
| high | writer self-verification plus a separate `gentle-ai-verify` run, always |
|
|
22
|
+
| unknown / assess failed | treated as high |
|
|
23
|
+
|
|
24
|
+
When RDD is `on` and the native review closed for this candidate, the writer's own self-verification is the record and the closed native review is the independent check, except a passive-risk change, which still gets a structural readback instead; any other `nativeReviewOutcome` under `on` follows this same tier table instead (gentle-pi#668). The small-model bias raises the medium tier to high for verification purposes only; an unknown RDD line never lowers a tier below `off`. The parent's own spot check (re-running one reported command before delivery) stays required in every tier.
|
|
25
|
+
|
|
@@ -0,0 +1,94 @@
|
|
|
1
|
+
# Telemetry
|
|
2
|
+
|
|
3
|
+
Runtime usage telemetry is best effort: an available usage event gets at most one
|
|
4
|
+
asynchronous attempt through `gentle-ai telemetry runtime send --json`. Busy,
|
|
5
|
+
failed, disabled, or cancelled attempts are discarded silently. There is no
|
|
6
|
+
metrics disk storage, outbox, retry, backoff, cooldown, daemon, or session reconstruction.
|
|
7
|
+
|
|
8
|
+
The production encoder uses the byte-identical [native schema mirror](../contracts/telemetry/runtime-aggregate-v1.schema.json).
|
|
9
|
+
[The synthetic fixture](../tests/fixtures/runtime-metrics-native-batches.json) pins its
|
|
10
|
+
SHA-256 and exact one-shot stdin bytes. No old intake fallback is used. These are
|
|
11
|
+
fake-subprocess tests, not a live collector or deployment verification.
|
|
12
|
+
|
|
13
|
+
## Runtime usage flow
|
|
14
|
+
|
|
15
|
+
1. A finalized primary assistant message or child completion supplies available usage.
|
|
16
|
+
Child-host extensions do not independently consume primary usage.
|
|
17
|
+
2. Pi builds event-local sanitized rows, not cumulative session totals. An occupied
|
|
18
|
+
attempt slot discards the event; it never queues it for later.
|
|
19
|
+
3. One cancellable immediate defers binary verification and subprocess launch beyond
|
|
20
|
+
the provider callback. The native command owns fresh policy and exactly one POST.
|
|
21
|
+
4. Only `stored`, `duplicate`, `discarded`, or `disabled` results are recognized.
|
|
22
|
+
All are terminal. `stored` and `duplicate` reflect collector acknowledgement,
|
|
23
|
+
not client persistence; Pi retains nothing and never retries.
|
|
24
|
+
|
|
25
|
+
There is no policy or capability subprocess before send, and no ingest or flush call.
|
|
26
|
+
The packaged binary is verified; development overrides are refused for runtime usage.
|
|
27
|
+
Missing binaries and unsupported commands discard without installation or fallback.
|
|
28
|
+
|
|
29
|
+
Replacement and shutdown cancel an unstarted attempt or request termination of its
|
|
30
|
+
child immediately, without a wait loop or final send. The process slot remains busy
|
|
31
|
+
until actual close, preventing overlap even if a cancelled process is slow to exit.
|
|
32
|
+
A one-second process timeout requests termination; it is not a retry timer or a hard
|
|
33
|
+
bound on synchronous binary verification or event-loop stalls.
|
|
34
|
+
|
|
35
|
+
## Data and source limits
|
|
36
|
+
|
|
37
|
+
Wire fields are the schema/registry, host, public model with evidence, available
|
|
38
|
+
selected/effective effort, orchestrator or known built-in subagent class, source
|
|
39
|
+
launch/response occurrence coverage, six token coverages, explicitly reported typed
|
|
40
|
+
duration, and a sanitized error category. Occurrences never represent sessions.
|
|
41
|
+
Prompts, responses, code, paths, private names, source IDs, and raw errors never enter
|
|
42
|
+
the payload. There is no `batch_id`; native creates the remote `delivery_id`.
|
|
43
|
+
|
|
44
|
+
One event produces 1–32 rows within 16 KiB. Oversized or invalid events discard
|
|
45
|
+
whole rather than splitting into multiple sends. Child launch selection is a separate
|
|
46
|
+
row with zero response-token coverage, never substituted for per-response evidence.
|
|
47
|
+
|
|
48
|
+
- Token coverage distinguishes reported, unavailable, and unsupported values. Pi's
|
|
49
|
+
SDK-positive counters are usable; zero defaults and absence do not prove reported zero.
|
|
50
|
+
- Selected model/effort is captured at the request hook, separately from response
|
|
51
|
+
model and effective-effort evidence. Ambiguous request sequences discard selection.
|
|
52
|
+
- Pi hooks do not correlate requests/retries reliably, so this adapter does not infer duration.
|
|
53
|
+
- Child configuration classification uses packaged built-in definitions, not agent names.
|
|
54
|
+
Launch configuration is not proof of the model or selected effort of each child response.
|
|
55
|
+
- Bounded live child observations remain in RAM until completion. No completed child
|
|
56
|
+
event is retained for forwarding. A completion without usage creates no usage send.
|
|
57
|
+
- Primary object deduplication uses weak tombstones; copied objects are distinct.
|
|
58
|
+
Child completion tombstones are capped at 256 per extension session. Source IDs
|
|
59
|
+
stay local and are never exported. No session history is read or reconstructed.
|
|
60
|
+
|
|
61
|
+
## What Gentle Pi does
|
|
62
|
+
|
|
63
|
+
On activation of a primary session (never for a named agent or an SDD phase executor), Gentle Pi resolves the package-local `gentle-ai` binary (honoring a registered dev-binary override, same as every other native call) and spawns:
|
|
64
|
+
|
|
65
|
+
```text
|
|
66
|
+
gentle-ai telemetry trigger --json
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
- detached, with stdout/stderr discarded (`stdio: "ignore"`);
|
|
70
|
+
- a 3 s deadline: a runaway process is killed, but Gentle Pi never waits for it to exit;
|
|
71
|
+
- at most once per process, regardless of how many sessions or sub-agents run afterward.
|
|
72
|
+
|
|
73
|
+
Rate limiting, enrollment, and every opt-out live entirely in `gentle-ai`; calling the trigger once per session start is safe by construction. A missing binary, an older binary without the `telemetry` verb (which prints `unknown telemetry command` and exits non-zero), or a spawn failure are all treated as "nothing to do" and never affect activation or surface an error to the user.
|
|
74
|
+
|
|
75
|
+
Install counts for `gentle-pi` and `gentle-engram` come from npm download statistics; neither package emits an install event of its own.
|
|
76
|
+
|
|
77
|
+
## The trigger contract
|
|
78
|
+
|
|
79
|
+
`gentle-ai telemetry trigger --json` always exits `0` and prints one line of JSON:
|
|
80
|
+
|
|
81
|
+
```json
|
|
82
|
+
{"schema":"gentle-ai.telemetry-trigger/v1","decision":"enrolled|sent_install|sent_heartbeat|rate_limited|backoff|disabled","source":"<deciding source>"}
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
`gentle-ai telemetry status|enable|disable|preview [--json]` exist for the opt-out flow; `status --json` prints `gentle-ai.telemetry-status/v1`. Gentle Pi's `/gentle:telemetry` slash command runs these in the foreground (bounded to 5 s) through the same binary resolver and relays the result.
|
|
86
|
+
|
|
87
|
+
## Opting out
|
|
88
|
+
|
|
89
|
+
Any of the following disables the nudge or the underlying telemetry:
|
|
90
|
+
|
|
91
|
+
- `/gentle:telemetry disable` — asks the local `gentle-ai` binary to disable telemetry. `/gentle:telemetry status` and `/gentle:telemetry preview` inspect it without leaving Pi.
|
|
92
|
+
- `DO_NOT_TRACK=1` — Gentle Pi does not spawn the trigger at all; `gentle-ai` also honors this standard independently.
|
|
93
|
+
- `GENTLE_AI_TELEMETRY=0` — same effect, `gentle-ai`'s own environment switch.
|
|
94
|
+
- `CI=true` — Gentle Pi does not spawn the trigger in automated/CI runs, since they are not a real usage signal.
|
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
# Windows Startup Console Capture
|
|
2
|
+
|
|
3
|
+
Use this protocol to validate the startup Git child processes on a real Windows desktop. Linux CI can verify spawn configuration but cannot observe Windows console windows.
|
|
4
|
+
|
|
5
|
+
## Capture
|
|
6
|
+
|
|
7
|
+
1. Start a Windows window-event trace before launching Pi. Capture `EVENT_OBJECT_SHOW`, its timestamp, HWND, owner PID, and owner session.
|
|
8
|
+
2. Deliberately launch a known visible, short-lived console as a positive control. If its show event is absent, the capture is **INCONCLUSIVE**.
|
|
9
|
+
3. Build or install the candidate Gentle Pi package and start `pi` from a repository whose path contains spaces. Leave it idle for at least 10 seconds so the startup identity lookups (`git rev-parse --show-toplevel` and `git rev-parse --git-common-dir`), branch lookup, initial shell scan, and repeated poll run.
|
|
10
|
+
4. Correlate each show event with `GetWindowThreadProcessId`, then use its timestamp, host PID/session, and Procmon process-creation ancestry and command line to associate it with a Pi-launched Git invocation. The HWND can be owned by `conhost.exe`, OpenConsole, or Windows Terminal rather than `git.exe`.
|
|
11
|
+
|
|
12
|
+
## Expected Result
|
|
13
|
+
|
|
14
|
+
No `WS_VISIBLE` show event is associated with the Git invocation for the startup identity lookups, banner branch lookup, initial shell scan, or repeated shell polling. A visible show event proves window visibility, not that it was unobscured on screen. Process start alone is not visible-window evidence. If event-to-invocation association cannot be established, report **INCONCLUSIVE**, not pass.
|
|
15
|
+
|
|
16
|
+
## Scope
|
|
17
|
+
|
|
18
|
+
This protocol does not cover user-configured external editors. Their inherited-stdio launch is intentional and may open a visible window.
|