planr 1.10.0-alpha.7 → 1.10.0-alpha.8

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (35) hide show
  1. package/docs/ARCHITECTURE.md +1 -1
  2. package/docs/SWITCHLOOM_COMPATIBILITY.md +9 -41
  3. package/docs/contracts/BROWSER_HARNESS_ADAPTER_V1.md +75 -0
  4. package/docs/contracts/EVIDENCE_ADAPTER_PROTOCOL_V1.md +98 -0
  5. package/docs/contracts/EVIDENCE_CONTRACT_V1.md +96 -2
  6. package/docs/contracts/schemas/com.planr.web.dom_state.v1.schema.json +30 -0
  7. package/docs/fixtures/mcp-contract.json +4 -1
  8. package/npm/bin/planr.js +29 -5
  9. package/npm/native/darwin-arm64/planr +0 -0
  10. package/npm/native/darwin-arm64/planr-host-capability-validator +0 -0
  11. package/npm/native/darwin-x86_64/planr +0 -0
  12. package/npm/native/darwin-x86_64/planr-host-capability-validator +0 -0
  13. package/npm/native/linux-arm64/planr +0 -0
  14. package/npm/native/linux-arm64/planr-host-capability-validator +0 -0
  15. package/npm/native/linux-x86_64/planr +0 -0
  16. package/npm/native/linux-x86_64/planr-host-capability-validator +0 -0
  17. package/package.json +2 -2
  18. package/plugins/planr/.claude-plugin/plugin.json +1 -1
  19. package/plugins/planr/.codex-plugin/plugin.json +1 -1
  20. package/plugins/planr/agents/pi/planr-worker.md +3 -5
  21. package/plugins/planr/agents/planr-worker.md +4 -5
  22. package/plugins/planr/skills/planr/SKILL.md +3 -3
  23. package/plugins/planr/skills/planr-goal/SKILL.md +2 -2
  24. package/plugins/planr/skills/planr-loop/SKILL.md +119 -56
  25. package/plugins/planr/skills/planr-loop/agents/planr-worker.md +4 -5
  26. package/plugins/planr/skills/planr-loop/references/host-dispatch.md +10 -17
  27. package/plugins/planr/skills/planr-loop/references/recovery-and-verification.md +4 -5
  28. package/plugins/planr/skills/planr-plan/SKILL.md +1 -1
  29. package/plugins/planr/skills/planr-status/SKILL.md +1 -1
  30. package/plugins/planr/skills/planr-summary/SKILL.md +2 -2
  31. package/plugins/planr/skills/planr-task-graph/SKILL.md +2 -2
  32. package/plugins/planr/skills/planr-verify-web/SKILL.md +16 -31
  33. package/plugins/planr/skills/planr-work/SKILL.md +10 -9
  34. package/scripts/host-capability-experiment.mjs +11 -7
  35. package/scripts/planr-host-capability-validator +17 -11
@@ -61,7 +61,7 @@ Every live worker decision is a typed work packet:
61
61
 
62
62
  - `kind: "outcome"` is ordinary maker work; `mode: "finding_repair"` returns named findings to the same responsible maker and ReviewGate without creating a fix item.
63
63
  - `kind: "review_gate"` is independently leased checker work. Attempts and findings remain children of that gate, never graph items.
64
- - `kind: "verification"` is a fresh verifier lease over a frozen canonical source digest. Product source is read-only; trusted Evidence can commit only through the source-checked Evidence transaction.
64
+ - `kind: "verification"` is a coordinator-owned verifier lease over a frozen canonical source digest. Product source is read-only; trusted Evidence can commit only through the source-checked Evidence transaction. The coordinator calls `evidence verify` without another model.
65
65
  - Releasing verification is an application-owned transition: the item lease, FeatureRun verifier role, and verification budget reconcile atomically back to `source_frozen`. `pick release --repair <reference>` additionally invalidates the frozen candidate and routes a fresh repair batch to the recorded maker; the failed verifier lease is never resumed.
66
66
  - An accepted material checkpoint is source-bound after any product or admission repair. Repair settlement preserves the prior attempt, reopens the same gate with the active freeze revision/digest and repair-obligation lineage, and leaves verification unleaseable until a fresh independent acceptance of that exact binding.
67
67
  - `kind: "hold"` is an admission or capability stop. A driver must report its classification and next action; it cannot reinterpret the hold as permission to replace the maker, weaken verification, or open an ad hoc review.
@@ -1,46 +1,14 @@
1
- # Switchloom Compatibility
1
+ # Archived Switchloom v0.3.2 Compatibility Evidence
2
2
 
3
- This report records Planr-side compatibility evidence for the exact public `switchloom@0.3.2` package. It is intentionally narrow: Planr verifies the package identity, repository-local declaration boundary, and its own consumption/proof semantics. It does not turn Switchloom internals, generated semantic role names, routes, doctor output, or runtime telemetry into a future contract.
3
+ This document records a historical compatibility check performed for the external `switchloom@0.3.2` package. It is not a current Planr release gate or a promise about future Switchloom artifacts, routes, role names, host telemetry, or lifecycle behavior.
4
4
 
5
- ## Exact v0.3.2 Behavior Verified
5
+ The historical check established only that the pinned v0.3.2 package could generate repository-local `.planr/agents.toml` and `.planr/policy.toml` declarations that the Planr version at that time consumed successfully. The accompanying one-off cross-product oracle has been removed from Planr.
6
6
 
7
- - Public npm metadata for `switchloom@0.3.2` resolved from the npm registry:
8
- - `dist.integrity`: `sha512-g96AZIFKXpG1toAO+Gri1sjD8q0SxFxtRSLAcRcSVRGDCZ/dUtERepsOB+cSHHH7hUsT0jSKZYPfNqiLvfKk9Q==`
9
- - `dist.shasum`: `d7d72c74ac3ecd5a3e355edd8e297284cba04403`
10
- - `dist.tarball`: `https://registry.npmjs.org/switchloom/-/switchloom-0.3.2.tgz`
11
- - `npm pack switchloom@0.3.2 --pack-destination /private/tmp --json` returned a 9-entry package with the expected native binaries, wrapper, provenance file, README, license, and package manifest.
12
- - The packed tarball at `/private/tmp/switchloom-0.3.2.tgz` had SHA256 `0c04e94fc4372845edf395b3ea51139b8a4b46f34404940e06b3f4ec3ce22d20`.
13
- - The reviewed oracle contract expects the v0.3.2 Planr integration bundle `balanced-codex-openai@1.0.0+2.0.0`, bundle SHA256 `bf48f502080ff444ccb67bc4eeacc9391e77dbf5f0f8f277814e9abc2443e6c8`, 10 managed artifacts, 7 Codex-native roles, and 5 Planr work-type routes.
14
- - Repository docs now document the safe external operator flow with explicit `switchloom@0.3.2` commands, no normal-guidance `apply --yes`, and a caveat that the v0.3.2 tagged README preserves stale `0.3.1` examples.
15
- - The immutable live oracle used isolated source clone `/private/tmp/planr-switchloom-source.6gbzx2/source` and completed with `SWITCHLOOM_SOURCE_ROOT=/private/tmp/planr-switchloom-source.6gbzx2/source SWITCHLOOM_TARBALL=/private/tmp/switchloom-0.3.2.tgz PLANR_ORACLE_TEMP_PARENT=/private/tmp npm run verify:switchloom-cross-product`, exit 0. The live root is `/private/tmp/planr-switchloom-cross-product-Lbluca`, with oracle receipt `/private/tmp/planr-switchloom-cross-product-Lbluca/oracle-receipt.json`.
16
- - The retained replay of the same root and isolated source clone exited 0 and wrote `/private/tmp/planr-switchloom-cross-product-Lbluca/replay-receipt.json`. Replay first asserts the current source fingerprint matches the retained live receipt, then separately verifies the source remains unchanged during replay.
17
- - The live and replay receipts bind the same Switchloom source fingerprint before and after each mode: HEAD `8ff05776085d3175211e080143c513c5173abaa4`, clean status SHA256 `e3b0c442...`, inventory SHA256 `790de1be...`, file-hash SHA256 `97a26b12...`, and 303 paths. The maintainer's original sibling Switchloom worktree was not modified for this proof.
18
- - The successful live oracle proves exact package tarball SHA256 `0c04e94fc4372845edf395b3ea51139b8a4b46f34404940e06b3f4ec3ce22d20`, bundle SHA256 `bf48f502080ff444ccb67bc4eeacc9391e77dbf5f0f8f277814e9abc2443e6c8`, 7 profiles, 5 routes, 10 managed artifacts, Planr declaration consumption, Planr audit, uninstall and unrouted Planr behavior, no-auth fail-closed behavior, requested-only rejection, global sentinel preservation, and routing ownership.
19
- - The successful live oracle proves separate native maker and reviewer execution: maker `model_routing_terra_high` on `gpt-5.6-terra` with high effort, reviewer `model_routing_sol_high` on `gpt-5.6-sol` with high effort. Hidden spawn messages are intentionally opaque; the proof is exact successful parent `agent_type`, `task_name`, and `fork_turns` metadata correlated with matching direct child role rollouts, repo-local skill reads, and item-specific Planr evidence. It does not claim plaintext hidden-message recovery.
20
- - One malformed parent `spawn_agent` attempt in the retained live root was retained as a diagnostic and not counted as a successful spawn. The oracle still requires exactly two successful routed spawns and still rejects default, wrong-role, missing-role, or extra successful child execution.
21
- - A separate fresh root, `/private/tmp/planr-switchloom-cross-product-gHtZQI`, was correctly rejected: Codex claimed `spawn_agent` was unavailable and directly handled the work, producing 0 successful routed spawns. That run is negative evidence for host stochastic failure detection, not a passing compatibility run.
7
+ Planr's stable boundary remains:
22
8
 
23
- ## Stable Planr Boundary
9
+ - Planr works without Switchloom or routing declarations.
10
+ - Planr treats profiles, routes, model names, efforts, and fallbacks as provider-neutral repository data.
11
+ - Requested routing metadata is not effective execution proof.
12
+ - External tools own their install, compilation, generated host files, reload requirements, and uninstall lifecycle.
24
13
 
25
- - Planr Core is provider-neutral. It reads `.planr/agents.toml`, `.planr/policy.toml`, and route-audit evidence, but does not install, invoke, compile, apply, or uninstall Switchloom output.
26
- - `.planr/agents.toml` profiles, routes, model names, efforts, role names, and fallback chains are opaque data to Planr. Planr may place requested routing data in pick packets, but requested-only values are not effective execution proof.
27
- - `--profile` records the profile a worker reports for a run. Genuine runtime proof belongs in route-audit evidence that separates requested, host-resolved, and effective dimensions.
28
- - Missing effective host evidence remains explicitly unavailable. Planr must not infer it from generated declarations, policy files, or worker claims.
29
- - External lifecycle remains external: Switchloom owns compile, apply, generated host roles, host reload requirements, and uninstall. Planr validates its own consumption with `planr agents check` and records evidence through logs/reviews.
30
- - Planr works without routing declarations. After an external uninstall, unrouted Planr fallback behavior remains the stable Core contract.
31
-
32
- ## Not Guaranteed For Future Switchloom Work
33
-
34
- Future Switchloom work, including thread `019f8a71-5b6c-7c41-9850-7050516fcee4`, may change semantic role names, route names, doctor output, generated artifacts, runtime telemetry, and host-specific evidence shape. Planr must not contract those as stable.
35
-
36
- The stable compatibility promise is limited to Planr's boundary: consume provider-neutral declarations when present, keep requested values separate from effective evidence, reject requested-only metadata as proof, avoid owning external lifecycle, and continue operating when routing files are absent.
37
-
38
- Any future mismatch in package identity, generated artifacts, Planr declaration parsing, route-audit semantics, host evidence, source-worktree mutation, or security posture should become a new finding and responsible-maker repair rather than an optimistic compatibility claim.
39
-
40
- ## Verification Notes
41
-
42
- - Rust formatting, clippy, full serial cargo tests, e2e, eval-contract, and routing-ownership tests passed when run outside the sandbox where HTTP/process tests are permitted.
43
- - Docs reference generation, typecheck, reference verification, maintenance verification, and Node 22 production build passed. Node 26 production docs builds hung in the optimized Next build phase and were terminated; the Node 22 build is the relevant repository runtime.
44
- - npm package dry-run passed using a temp npm cache because the user npm cache contains root-owned files.
45
- - Privacy checks passed. The docs app Next.js runtime was upgraded from `16.2.10` to `16.2.11` in `apps/docs/package.json` and `pnpm-lock.yaml`; the lockfile change is limited to Next core/env/SWC and mechanically recalculated peer snapshots. `npm run security:check` passed with BetterLeaks reporting no leaks and Trivy reporting 0 vulnerabilities for both `Cargo.lock` and `pnpm-lock.yaml`.
46
- - The immutable source clone proof supersedes earlier retained-replay-only evidence. Earlier concurrent external mutation and missing-receipt attempts were treated as real failures and are not counted as passing evidence.
14
+ For current Switchloom behavior, consult the external project and its versioned documentation. Any future compatibility claim requires a new bounded check owned by that integration; Planr does not retain a permanent package-specific oracle.
@@ -0,0 +1,75 @@
1
+ # Browser Harness adapter v1
2
+
3
+ Status: frozen for `planr-browser-harness-adapter` v1.
4
+
5
+ This reference defines Planr's first process adapter for the external
6
+ `browser-harness` tool. The adapter implements
7
+ `EVIDENCE_ADAPTER_PROTOCOL_V1.md`. Planr Core does not depend on this adapter.
8
+
9
+ ## Capability
10
+
11
+ The adapter supports `com.planr.web.dom_state` observations with the payload
12
+ schema `schema://com.planr.web.dom_state.v1`.
13
+
14
+ Each requirement uses these fields:
15
+
16
+ - `subject`: A CSS selector for the observed element.
17
+ - `expected.text`: The exact trimmed `textContent`, if text is required.
18
+ - `expected.visible`: The required rendered visibility, if visibility is
19
+ required.
20
+ - `state_transitions`: A shared array of click actions for the batch. Each
21
+ action has `action = "click"`, an accessibility `role`, and an accessibility
22
+ `name`.
23
+
24
+ All requirements in one batch must use the same transition array and the same
25
+ optional execution method. The adapter rejects arbitrary actions, arbitrary
26
+ JavaScript, mixed methods, and unsupported observation types before it starts
27
+ the browser tool.
28
+
29
+ ## Execution
30
+
31
+ An availability probe runs `browser-harness --version`. A real Evidence run
32
+ starts `browser-harness` once and sends one generated Python program over
33
+ stdin. The program completes these steps in one browser session:
34
+
35
+ 1. Open the bound HTTP or HTTPS target.
36
+ 2. Find each click target in the accessibility tree by its role and name.
37
+ 3. Click the center of the target's CDP box model.
38
+ 4. Read each DOM postcondition until it passes or the five-second observation
39
+ window ends.
40
+ 5. Return one structured result for the full requirement batch.
41
+ 6. Close the tab.
42
+
43
+ The generated program inserts requirement data only as JSON. A requirement
44
+ cannot add executable Python or JavaScript.
45
+
46
+ ## Result
47
+
48
+ Each `actual` object contains:
49
+
50
+ ```json
51
+ {
52
+ "schema_ref": "schema://com.planr.web.dom_state.v1",
53
+ "selector": "#status",
54
+ "text": "B",
55
+ "visible": true
56
+ }
57
+ ```
58
+
59
+ Planr Core validates this object against the registered JSON Schema. Core then
60
+ evaluates the requirement's `expected` predicate.
61
+
62
+ If every requirement selects the `browser-harness` agent skill, the adapter
63
+ also returns the matching supervised invocation record. This record is adapter
64
+ output. It becomes trusted only after Planr validates the observed process and
65
+ all sealed bindings.
66
+
67
+ ## Capture policy
68
+
69
+ The v1 success path disables Browser Harness recordings. The generated program
70
+ does not call screenshot, recording, or trace helpers. The structured DOM
71
+ observation is the retained proof for a non-visual criterion.
72
+
73
+ The adapter does not advertise a visual observation type. Planr therefore
74
+ rejects it during capability matching for a visual requirement. A recording or
75
+ screenshot cannot upgrade `com.planr.web.dom_state` into visual Evidence.
@@ -0,0 +1,98 @@
1
+ # Evidence adapter protocol v1
2
+
3
+ Status: frozen for the `planr.evidence.adapter-request.v1` request and the
4
+ `planr.structured_observation_results.v2` result.
5
+
6
+ This reference defines the process protocol between Planr Core and a registered
7
+ Evidence adapter. It extends `EVIDENCE_CONTRACT_V1.md`. It does not change the
8
+ Evidence domain schema.
9
+
10
+ ## Trust ownership
11
+
12
+ Planr Core creates the adapter request, starts the registered process, and
13
+ validates the result. The adapter reports observations. It cannot assign
14
+ provenance, decide coverage, or close work.
15
+
16
+ Planr treats adapter output as untrusted input until all execution and binding
17
+ checks pass. A valid JSON result without a Planr-observed process execution
18
+ cannot create trusted Evidence.
19
+
20
+ ## Process transport
21
+
22
+ Planr starts the adapter with bounded time, stdout, and stderr limits. Planr
23
+ writes one JSON request to stdin and then closes stdin. The adapter writes one
24
+ JSON result to stdout.
25
+
26
+ An availability probe starts the same executable with closed, empty stdin. The
27
+ adapter can use that invocation to check runtime availability. A probe result
28
+ cannot satisfy an Evidence requirement.
29
+
30
+ Planr reserves the `PLANR_EVIDENCE_` environment variable namespace. An adapter
31
+ request does not use environment variables for target, environment, or contract
32
+ bindings.
33
+
34
+ ## `planr.evidence.adapter-request.v1`
35
+
36
+ The request has these fields:
37
+
38
+ - `schema_version`: `planr.evidence.adapter-request.v1`.
39
+ - `request_id`: A new Planr-assigned ID for this process execution.
40
+ - `request_digest`: The SHA-256 canonical JSON digest of the request without
41
+ `request_digest`.
42
+ - `obligation_id` and `criterion_id`: The exact Evidence obligation identity.
43
+ - `requirements`: The selected `ObservationRequirement` objects. The array
44
+ includes each subject, expected value, target, payload schema, and optional
45
+ execution method.
46
+ - `target` and `environment`: The runtime bindings selected by Planr.
47
+ - `fixture_disclosure`: The fixture and mock disclosure admitted by policy.
48
+ - `assurance_policy`: The policy that controls required capture strength.
49
+ - `result_contract`: The registered outer result schema binding.
50
+ - `execution_contract_digest`: The digest of the registered process contract.
51
+ - `execution_binding`: The sealed run-index subset and its requirement IDs.
52
+ - `retry`: The exact attempt number, maximum attempts, and predecessor IDs.
53
+
54
+ The request contains no provider-specific fields. A Browser Harness, native
55
+ browser, Playwright, mobile, desktop, game-engine, API, or CLI adapter receives
56
+ the same request shape.
57
+
58
+ ## `planr.structured_observation_results.v2`
59
+
60
+ A structured adapter result has these required fields:
61
+
62
+ - `schema_version`: `planr.structured_observation_results.v2`.
63
+ - `request_id` and `request_digest`: Exact copies from the current request.
64
+ - `target`, `observed_target`, and `environment`: The declared and observed
65
+ runtime identity.
66
+ - `execution_contract_digest`: An exact copy from the current request.
67
+ - `fixture_disclosure`: The actual fixture and mock use.
68
+ - `observations`: One result for every selected requirement and no other
69
+ results.
70
+
71
+ Each observation contains only `requirement_id`, `type`, and `actual`. The
72
+ `actual` object names its payload schema. Planr validates the object against the
73
+ registered JSON Schema and evaluates the expected predicate.
74
+
75
+ If a requirement selects an agent skill as its execution method, the result
76
+ also includes the structured `agent_skill` invocation record required by the
77
+ Evidence domain contract. The supervised adapter creates this record. Agent or
78
+ user JSON cannot submit it through another trusted path.
79
+
80
+ Adapters can include diagnostic top-level fields. Diagnostics do not affect
81
+ provenance or coverage.
82
+
83
+ ## Failure behavior
84
+
85
+ Planr rejects the result if the process fails or if any required binding does
86
+ not match. This includes a stale request ID, a stale request digest, a wrong
87
+ target, a wrong environment, a wrong execution contract, missing observations,
88
+ extra observations, a schema mismatch, and an unsatisfied expected predicate.
89
+
90
+ Planr records the failed attempt. Planr does not create a trusted passing
91
+ receipt from that result.
92
+
93
+ ## Version transition
94
+
95
+ `planr.structured_observation_results.v2` replaces
96
+ `planr.structured_observation_results.v1`. Planr does not keep a second trusted
97
+ compatibility path. Adapter manifests and policy registrations must use the v2
98
+ schema reference and artifact name.
@@ -21,10 +21,101 @@ Evidence Contract v1 is Planr's local-first contract for proving acceptance crit
21
21
  - Route Audit remains the owner of requested, resolved, and effective routing evidence. Evidence may consume a mapped provenance view, but it must not copy requested route declarations into effective execution proof.
22
22
  - Agent Profiles, model-routing capability classes, usage-policy capability classes, MCP protocol capabilities, and context tags are dispatch or protocol metadata. They are not verification capability instances and do not prove runtime availability.
23
23
  - Planr logs remain narrative and supporting records. A `kind = verification` log is a claim that can be referenced, but it never satisfies a binding observation.
24
- - A repository without a binding Evidence policy and without binding plan obligations is explicitly non-binding. A binding plan is `binding_unsatisfied` unless its authoritative active obligations match the declared build-plan criterion set exactly: zero, partial, duplicate, or undeclared bindings all fail closed. Planr returns a hold before leasing or creating a FeatureRun, persists a capability hold when an existing run reaches readiness, rejects coverage settlement, closure, final review, and stop activation, and never substitutes logs or an empty receipt lineage.
24
+ - A repository without a binding Evidence policy and without binding plan obligations is explicitly non-binding. A binding plan is `binding_unsatisfied` unless its authoritative active obligations match the declared build-plan criterion set exactly: zero, partial, duplicate, or undeclared bindings all fail closed. Planr returns a hold before leasing or creating a FeatureRun, persists a capability hold when an existing run reaches readiness, rejects coverage settlement, closure, and stop activation, and never substitutes logs or an empty receipt lineage.
25
25
  - Migration is the sole obligation-materialization path. It is explicit, plan-scoped, previewable, idempotent, and accepts only an exact declared criterion binding set before materializing ordinary immutable `ProofObligation` rows. It must not rewrite plans, logs, reviews, artifacts, or historical claims.
26
26
  - Planr artifacts remain files or references with digests. An artifact alone is not trusted evidence unless a trusted receipt binds it to the source revision, target, environment, execution identity, observation results, and policy.
27
27
 
28
+ ## Binding Execution Orchestration
29
+
30
+ The normal public success path is one plan-scoped `evidence verify` broker call. The caller identity
31
+ must differ from the responsible maker. Planr uses that identity for the verifier lease, readiness
32
+ admission, adapter execution, coverage evaluation, and FeatureRun settlement. A coordinator or host
33
+ must not spawn another model to perform this mechanical Evidence phase. The lower-level readiness
34
+ and run services remain diagnostic and application primitives; they do not define a second trusted
35
+ workflow.
36
+
37
+ Before verification admission, FeatureRun source freeze is legal only when no open ordinary
38
+ implementation outcome remains. Planned `code`, `fix`, `docs`, and `test` share the one domain-owned
39
+ maker-compatible classification. An active source-frozen run that still has open ordinary work and
40
+ has no verifier admission, Evidence attempt, or receipt for its immutable freeze is retired only by
41
+ the typed `premature-source-freeze` transition. Retirement preserves the freeze and all Evidence
42
+ history, releases active ordinary leases, and creates no successor. The repository-owned no-model
43
+ `com.planr.premature_freeze.lifecycle.v1` capability is reserved for the final frozen-source HARDEN
44
+ observation; registering it does not execute or satisfy that observation during BUILD.
45
+
46
+ An active Verification FeatureRun whose current admission is absent or unequal across the active
47
+ plan/run/revision, freeze, verifier worker/generation, optional item, or admitted/sealed run-index
48
+ digest is retired only by the typed `inconsistent-verification` transition. A current item whose
49
+ status or worker does not match the verifier lease is
50
+ `verification_item_ownership_conflict`. Exact equality rejects retirement. One immediate
51
+ optimistic transaction invalidates but preserves the freeze, ends or preserves-ended the referenced
52
+ batch, releases exact roles, Verification reservations, and the exact observed optional item state,
53
+ preserves every prior Evidence/history identity, emits one typed event, and creates no successor.
54
+ Repetition reads that event and writes nothing. The
55
+ repository-owned `com.planr.inconsistent_verification.retirement.v1` capability is registered during
56
+ BUILD but first executes only after HARDEN supplies the exact focused invariant.
57
+
58
+ `planr.evidence.run-index.v2` is the sole executable run-index shape. Readiness consumes canonical
59
+ authoritative obligation rows and seals exactly one run for every distinct canonical target within
60
+ each obligation. A run names one `obligation_id`, one target, and sorted non-empty unique
61
+ `requirement_ids`. For each obligation, admission recomputes the target partition and requires the
62
+ run subsets to be disjoint and to form the exact union of authoritative observation requirement
63
+ IDs. Requirements from different obligations never share a run.
64
+
65
+ Execution, structured and ordinary result validation, receipt observation construction, retry
66
+ predecessor validation, independence, hermetic reuse, and persisted attempt/receipt lineage bind
67
+ only the selected requirement subset plus the sealed source, policy, target, capability,
68
+ environment, execution contract, and run-index digest. Extra, missing, duplicate, foreign-target,
69
+ or cross-subset data fails before trusted receipt persistence. `evidence/coverage` remains the sole
70
+ coverage and closure owner; a run-index result is not a coverage verdict.
71
+
72
+ A multi-target obligation that would execute a `non_repeatable_one_shot` capability fails readiness
73
+ before launch or durable allowance claim. Host capture consumes the same sealed target/subset
74
+ contract as process execution and owns no first-observation or all-targets-equal policy.
75
+
76
+ Pre-receipt admission failure persists no attempt, receipt, coverage verdict, or ProductFinding.
77
+ One explicit optimistic plan/run/freeze/revision-scoped FeatureRun repair request has a closed
78
+ reason enum and conditional seal binding. Pre-seal `readiness-blocked` and
79
+ `run-index-seal-failed` require `run_index_digest` to be absent; post-seal pre-receipt
80
+ `sealed-run-rejected` and `capability-admission-failed` require the exact admitted digest. A failed
81
+ readiness lease transaction rolls back before Planr persists its durable capability hold and
82
+ diagnostic, so no unusable verifier lease is committed to create repair context. The repair
83
+ invalidates the active freeze, releases verifier ownership and any present verification-item
84
+ lease, restores the original maker at the next lease generation, starts one repair batch, emits one
85
+ repair event, and projects an optional `verification_item_id`. Post-receipt `product_failed` and
86
+ terminal one-shot exhaustion remain their existing distinct lifecycles.
87
+
88
+ Satisfied plan coverage settles the active FeatureRun by exact verifier lease generation, active
89
+ immutable freeze, and accepted trusted receipt lineage whose source binding exactly equals that
90
+ freeze. A verification map item is a zero-or-one projection. If one item is picked or running under
91
+ the verifier, settlement closes and logs it in the same transaction; if none exists, settlement
92
+ performs zero item/log mutations. A ready unleased verification item remains fail-closed. The
93
+ transition always reconciles the verification budget wall, applies `VerificationPassed` to
94
+ `Complete` when no ordinary work remains (or `Implementation` if ordinary work reopened),
95
+ persists/releases roles through the canonical repository transaction, and emits one event shape
96
+ with nullable `item_id` and `log_id`. Binding coverage settlement returns `next_action: none` on
97
+ completion and creates no final ReviewGate. `planr plan final-review` is a non-binding-plan flow.
98
+
99
+ Terminal `non_repeatable_one_shot` exhaustion is also zero-or-one item inside the trusted
100
+ attempt/receipt transaction. Attempt/receipt persistence, budget reconciliation, FeatureRun
101
+ cancellation, and verifier release always commit atomically; item failure/logging occurs only when
102
+ one active projection is present, and ready-unleased remains fail-closed.
103
+
104
+ Post-receipt ProductFinding repair remains distinct from pre-receipt admission repair. The
105
+ application resolves the verification map item as an optional current plan-path projection for
106
+ routing, maker work packets, settlement, idempotent replay, selective-replay handoff, and verifier
107
+ release. Absence performs no item mutation; only a present eligible item is updated. Durable
108
+ ProductRepair settlement persists invalidation/run/maker/selective-obligation/settlement/freeze
109
+ lineage and no item identity. The lifecycle does not reclassify the outcome or create a receipt.
110
+
111
+ This is a hard cut: run-index v1, first-observation target inference, all-targets-equal checks,
112
+ item-keyed `pick release --repair`, item-required repair, persisted repair-item identity,
113
+ settlement, final-review, and exhaustion ownership, readers, aliases, fallbacks, shims, and runtime
114
+ compatibility do not exist.
115
+ Completed trusted Evidence v1 receipts remain coverage records because coverage is independent of
116
+ run-index orchestration; unexecuted v1 run indexes and in-flight retry lineage lacking the v2
117
+ execution binding are not translated or inferred.
118
+
28
119
  ## Versioning And Compatibility
29
120
 
30
121
  - `schema_version` for all v1 objects is `evidence.contract.v1`.
@@ -92,7 +183,10 @@ Required fields:
92
183
 
93
184
  - `id`, `schema_version`, `version`, `adapter_kind`, `adapter_digest`.
94
185
  - Supported surfaces, observation type/schema/digest triples, interactions, artifacts, runtime targets, provenance path, permissions, costs, determinism, repeatability, independence, blind spots, and availability probe contract.
95
- - When a capability explicitly declares `repeatability = non_repeatable_one_shot`, Planr derives `max_attempts = 1`; any conflicting caller declaration is rejected before launch. Before the adapter can spawn, Planr atomically claims one durable allowance scoped to the active FeatureRun source freeze. That claim survives process, receipt, or settlement failure, so every later fresh initial, retry, or concurrent contender for the freeze is rejected without spawning. Its committed non-passing attempt (`attempt_index + 1 = max_attempts`), including `product_failed`, atomically exhausts that FeatureRun verification allowance with the attempt and receipt. Planr records `verification_attempts_exhausted`, releases the verifier lease, exposes no next verification action, and does not create a product-finding repair or replay path. Missing or other repeatability values never infer one-shot behavior.
186
+ - When a capability explicitly declares `repeatability = non_repeatable_one_shot`, Planr derives `max_attempts = 1`; any conflicting caller declaration is rejected before launch. Before the adapter can spawn, Planr atomically claims one durable allowance scoped to the active FeatureRun source freeze. That claim survives process, receipt, or settlement failure, so every later fresh initial, retry, or concurrent contender for the freeze is rejected without spawning. Its committed non-passing attempt (`attempt_index + 1 = max_attempts`), including `product_failed`, atomically exhausts that FeatureRun verification allowance with the attempt and receipt. Planr records `verification_attempts_exhausted`, releases the verifier lease, exposes no next verification action, and does not create a product-finding repair or replay path. A projected verification item is optional: when active it is failed/logged atomically; when absent there is no item/log mutation; when ready but unleased the transaction fails closed. Missing or other repeatability values never infer one-shot behavior.
187
+ - Every other capability also defaults to `max_attempts = 1`. A manifest may explicitly admit a larger bounded value only for a repeatable execution contract. Planr never starts a second attempt merely because a verifier or environment failure was returned; a later execution requires a fresh canonical invocation after the reported external state has materially changed.
188
+ - A fresh canonical invocation may append a superseding admission for the same exact active source freeze only when the previously admitted run-index digest already has both a durable Evidence attempt and receipt. The previous admission and execution history remain immutable. A conflicting admission with no execution receipt is pre-receipt state and remains rejected; it cannot bypass Verification Admission Repair.
189
+ - When source is stale after such a receipt, one canonical broker invocation may perform exactly one observed refresh: invalidate and preserve the old freeze and receipt bindings, release the exact verifier lease, freeze the current source, reacquire verification, and seal again. The broker does not repeat this transition if the replacement source becomes stale.
96
190
  - Process adapters declare a closed `availability_probe.kind = process` contract with executable name, arguments, optional working directory, timeout, stdout/stderr byte limits, and the payload schema binding for emitted observations.
97
191
 
98
192
  A manifest is a claim about what a method can observe. It is not proof that the method is available now.
@@ -0,0 +1,30 @@
1
+ {
2
+ "$schema": "https://json-schema.org/draft/2020-12/schema",
3
+ "$id": "schema://com.planr.web.dom_state.v1",
4
+ "type": "object",
5
+ "additionalProperties": false,
6
+ "required": [
7
+ "schema_ref",
8
+ "selector",
9
+ "text",
10
+ "visible"
11
+ ],
12
+ "properties": {
13
+ "schema_ref": {
14
+ "const": "schema://com.planr.web.dom_state.v1"
15
+ },
16
+ "selector": {
17
+ "type": "string",
18
+ "minLength": 1
19
+ },
20
+ "text": {
21
+ "type": [
22
+ "string",
23
+ "null"
24
+ ]
25
+ },
26
+ "visible": {
27
+ "type": "boolean"
28
+ }
29
+ }
30
+ }
@@ -35,6 +35,7 @@
35
35
  "planr_pick_stale",
36
36
  "planr_run_restart",
37
37
  "planr_run_resolve_budget_hold",
38
+ "planr_run_repair_verification_admission",
38
39
  "planr_recover_sweep",
39
40
  "planr_approval_request",
40
41
  "planr_approval_approve",
@@ -54,11 +55,12 @@
54
55
  "planr_evidence_capability_show",
55
56
  "planr_evidence_run",
56
57
  "planr_evidence_import",
58
+ "planr_evidence_host_capture_admit",
57
59
  "planr_evidence_host_capture_import",
58
- "planr_evidence_host_capture_run",
59
60
  "planr_evidence_attempts",
60
61
  "planr_evidence_receipts",
61
62
  "planr_evidence_coverage",
63
+ "planr_evidence_verify",
62
64
  "planr_evidence_explain",
63
65
  "planr_evidence_readiness",
64
66
  "planr_evidence_recover_settlement",
@@ -75,6 +77,7 @@
75
77
  "planr_review_ingest",
76
78
  "planr_review_evidence",
77
79
  "planr_review_gate_close",
80
+ "planr_review_gate_release",
78
81
  "planr_review_findings_resolve",
79
82
  "planr_close_item",
80
83
  "planr_context_create",
package/npm/bin/planr.js CHANGED
@@ -18,16 +18,40 @@ function platformTarget() {
18
18
  }
19
19
 
20
20
  const target = platformTarget();
21
- const candidates = [
21
+ const packagedCandidates = [
22
22
  process.env.PLANR_NATIVE_BIN,
23
23
  // Published package: per-platform binaries bundled at release time.
24
24
  target && path.join(here, "..", "native", target, "planr"),
25
- // Repository checkout: local cargo builds.
26
- path.join(packageRoot, "target", "release", "planr"),
27
- path.join(packageRoot, "target", "debug", "planr"),
28
25
  ].filter(Boolean);
29
26
 
30
- const binary = candidates.find(candidate => fs.existsSync(candidate));
27
+ function repositoryBuildCandidates() {
28
+ if (!fs.existsSync(path.join(packageRoot, "Cargo.toml"))) {
29
+ return [];
30
+ }
31
+ const metadata = spawnSync(
32
+ "cargo",
33
+ ["metadata", "--no-deps", "--format-version", "1"],
34
+ { cwd: packageRoot, encoding: "utf8" },
35
+ );
36
+ if (metadata.status !== 0) {
37
+ return [];
38
+ }
39
+ try {
40
+ const targetDirectory = JSON.parse(metadata.stdout).target_directory;
41
+ if (typeof targetDirectory !== "string" || targetDirectory.length === 0) {
42
+ return [];
43
+ }
44
+ return [
45
+ path.join(targetDirectory, "release", "planr"),
46
+ path.join(targetDirectory, "debug", "planr"),
47
+ ];
48
+ } catch {
49
+ return [];
50
+ }
51
+ }
52
+
53
+ const binary = packagedCandidates.find(candidate => fs.existsSync(candidate))
54
+ ?? repositoryBuildCandidates().find(candidate => fs.existsSync(candidate));
31
55
 
32
56
  if (!binary) {
33
57
  if (!target) {
Binary file
Binary file
Binary file
Binary file
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "planr",
3
- "version": "1.10.0-alpha.7",
3
+ "version": "1.10.0-alpha.8",
4
4
  "description": "Local-first planning and execution coordination for coding agents.",
5
5
  "license": "MIT",
6
6
  "repository": {
@@ -34,6 +34,7 @@
34
34
  "verify:release-script": "node scripts/test-release-script.mjs",
35
35
  "verify:change-classifier": "node scripts/test-verification-policy.mjs",
36
36
  "verify:runner": "node scripts/test-verification-runner.mjs",
37
+ "verify:single-session-loop": "node scripts/test-planr-risk-based-guidance.mjs --single-session-loop",
37
38
  "verify:ci-router": "node scripts/test-ci-router.mjs",
38
39
  "verify:docs-deployment": "node scripts/test-docs-deployment.mjs",
39
40
  "classify:changes": "node scripts/classify-changes.mjs",
@@ -64,7 +65,6 @@
64
65
  "docs:verify-clean-install": "pnpm --filter @planr/docs verify:clean-install",
65
66
  "docs:evidence-examples:generate": "cargo build --bin planr && pnpm --filter @planr/docs evidence-examples:generate",
66
67
  "docs:evidence-examples:check": "cargo build --bin planr && pnpm --filter @planr/docs evidence-examples:check",
67
- "verify:switchloom-cross-product": "node scripts/verify-switchloom-cross-product.mjs",
68
68
  "docs:reference:generate": "cargo build --bin planr && pnpm --filter @planr/docs reference:generate",
69
69
  "docs:verify-reference": "cargo build --bin planr && pnpm --filter @planr/docs reference:check && pnpm --filter @planr/docs verify:reference"
70
70
  },
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "planr",
3
3
  "description": "Skill-driven planning and execution loop for coding agents: one planr entry point, an autonomous planr-loop, and evidence-backed task graph skills powered by the planr CLI.",
4
- "version": "1.10.0-alpha.7",
4
+ "version": "1.10.0-alpha.8",
5
5
  "author": {
6
6
  "name": "instructa"
7
7
  },
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "planr",
3
- "version": "1.10.0-alpha.7",
3
+ "version": "1.10.0-alpha.8",
4
4
  "description": "Skill-driven planning and execution loop for coding agents: one $planr entry point, an autonomous $planr-loop, and evidence-backed task graph skills powered by the planr CLI.",
5
5
  "author": {
6
6
  "name": "instructa",
@@ -20,15 +20,13 @@ Planr rolls its internal three-outcome ExecutionBatch atomically; that boundary
20
20
  a host stop. Within that run, do not run binding Evidence readiness, collection,
21
21
  or an opportunistic live smoke for a mutable implementation item. Use ordinary
22
22
  changed-file checks, keep the same maker through stable source freeze, write a compact durable handoff only at a genuine stop, and stop.
23
- A fresh verification-only worker first leases `planr pick --plan <plan-id>
24
- --work-type verification --json`, then runs readiness under that same identity and
25
- executes only `readiness.run_index.repository_path`. Product source remains read-only
26
- by the canonical Evidence SOURCE_PATHS digest. `planr evidence run` enforces that digest before trusted
23
+ The coordinator runs the handoff packet's `commands.verify` command without another model.
24
+ Product source remains read-only by the canonical Evidence SOURCE_PATHS digest. `planr evidence verify` enforces that digest before trusted
27
25
  receipt commit; source mismatch records a failed non-covering attempt and zero
28
26
  new trusted receipts. When `mode` is `finding_repair`, repair the named findings
29
27
  for the same ReviewGate, log the changed files and commands on its scoped outcome,
30
28
  resolve those finding ids, and stop for re-review. No review or fix map item exists.
31
- Product findings require coordinator re-freeze and leased readiness before selective Evidence replay unless
29
+ Product findings require coordinator re-freeze and another brokered verify call before selective Evidence replay unless
32
30
  a material review requires exact-source Evidence earlier. Log changed files and the real verification
33
31
  commands you ran. Settle ordinary outcomes with `planr done --next` inside the
34
32
  authorized run and plain `planr done` when standalone; Planr materiality decides
@@ -22,16 +22,15 @@ empty pick, or budget. Planr rolls its internal three-outcome ExecutionBatch ato
22
22
  that boundary is not a host stop.
23
23
  Within that run, do not run binding Evidence readiness, collection, or an opportunistic live smoke for a
24
24
  mutable implementation item. Use ordinary changed-file checks, keep the same maker
25
- through stable source freeze, write a compact durable handoff only at a genuine stop, and stop. A fresh
26
- verification-only worker first leases `planr pick --plan <plan-id> --work-type verification --json`,
27
- then runs readiness under that same identity and executes only `readiness.run_index.repository_path`.
25
+ through stable source freeze, write a compact durable handoff only at a genuine stop, and stop. The
26
+ coordinator runs the handoff packet's `commands.verify` command without another model.
28
27
  Product source remains read-only by the canonical Evidence SOURCE_PATHS digest.
29
- `planr evidence run` enforces that digest before trusted receipt commit; source
28
+ `planr evidence verify` enforces that digest before trusted receipt commit; source
30
29
  mismatch records a failed non-covering attempt and zero new trusted receipts.
31
30
  When `mode` is `finding_repair`, repair the named findings for the same ReviewGate,
32
31
  log the changed files and commands on its scoped outcome, resolve those finding ids,
33
32
  and stop for re-review. No review or fix map item exists. Product findings require
34
- coordinator re-freeze and leased readiness before selective Evidence replay unless a material review requires
33
+ coordinator re-freeze and another brokered verify call before selective Evidence replay unless a material review requires
35
34
  exact-source Evidence earlier.
36
35
  Log changed files and the real verification commands you ran. Settle ordinary outcomes
37
36
  with `planr done --next` inside the authorized run and plain `planr done` when standalone;
@@ -45,7 +45,7 @@ Evaluate top to bottom; pick the first row that matches both intent and state:
45
45
  | new idea, PRD, scope, architecture | no plan, or plan needs refinement | `planr-plan` |
46
46
  | new feature, refactor, or fix on an existing project | project exists, scope has no plan yet | `planr-plan` (feature-scoped plan, then extend the map) |
47
47
  | plan exists but no map / map needs structure, dependencies, breakdown | build plan checked, map missing or stale | `planr-task-graph` |
48
- | ReviewGate pending or leased | open gate in the active FeatureRun | `planr-review` |
48
+ | ReviewGate pending or leased | explicit material gate in the active FeatureRun | `planr-review` |
49
49
  | ReviewGate has findings | active FeatureRun returns `work_packet.mode: "finding_repair"` to its responsible maker | `planr-work` |
50
50
  | implement, fix, continue work | ready items exist on the map | `planr-work` |
51
51
  | implement, but nothing is ready | all items blocked | `planr-status`, then report blockers to the user |
@@ -54,7 +54,7 @@ Evaluate top to bottom; pick the first row that matches both intent and state:
54
54
 
55
55
  - Route to exactly one skill per dispatch. If the request spans stages (idea -> running feature), route to `planr-loop` and let the loop sequence the stages.
56
56
  - Explicit delivery language such as "implement", "work autonomously until complete", or "finish this" outranks missing plan or map state and routes to `planr-loop`. Never turn a delivery request into a prep-only handoff.
57
- - Never skip a stage: no map items without a checked build plan, no FeatureRun completion without its required ReviewGates, and no review verdict without evidence.
57
+ - Never skip a stage: no map items without a checked build plan, no FeatureRun completion without Binding Evidence or any explicitly required material ReviewGates, and no review verdict without evidence.
58
58
  - When delegating to a subagent, prompt it with a skill reference plus item id (for example: `Use $planr-work on item <id>`), never with a hand-written workflow prompt.
59
- - The maker never reviews its own work. Reviews run through `planr-review` in a separate agent or subagent whenever the host supports it.
59
+ - The maker never reviews its own material ReviewGate. Reviews run through `planr-review` in a separate agent or subagent only when the plan or a protected-risk transition explicitly requires one.
60
60
  - If intent is genuinely ambiguous and state allows multiple routes, ask the user one short question instead of guessing.