project-tiny-context-harness 0.7.1 → 0.7.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +8 -6
- package/assets/README.md +8 -8
- package/assets/README.zh-CN.md +6 -6
- package/assets/agents/AGENTS_CORE.md +1 -1
- package/assets/skills/long-task-workflow/SKILL.md +7 -5
- package/assets/skills/long-task-workflow/references/authority-lifecycle.md +4 -2
- package/assets/skills/long-task-workflow/references/contract-authoring.md +1 -0
- package/assets/skills/long-task-workflow/references/evidence-design.md +7 -0
- package/dist/commands/long-task-revision.js +31 -4
- package/dist/commands/long-task.js +8 -1
- package/dist/lib/long-task-authority-revision-summary.js +33 -2
- package/dist/lib/long-task-authority-revision-types.d.ts +6 -0
- package/dist/lib/long-task-authority-revision.js +21 -1
- package/dist/lib/long-task-status-v2.d.ts +7 -0
- package/dist/lib/long-task-status-v2.js +15 -7
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -137,7 +137,7 @@ npm ci
|
|
|
137
137
|
npm run smoke:quickstart
|
|
138
138
|
npm run preview:pack
|
|
139
139
|
cd /path/to/your/test-repo
|
|
140
|
-
npm install -D /path/to/project-tiny-context-harness/tmp/ty-context/source-preview/package/project-tiny-context-harness-0.7.
|
|
140
|
+
npm install -D /path/to/project-tiny-context-harness/tmp/ty-context/source-preview/package/project-tiny-context-harness-0.7.2.tgz
|
|
141
141
|
npx --no-install ty-context init --adopt
|
|
142
142
|
make validate-context
|
|
143
143
|
```
|
|
@@ -235,7 +235,7 @@ Before product implementation, the Agent asks the user to continue with the curr
|
|
|
235
235
|
|
|
236
236
|
Harness cannot switch the host-selected model. It creates no checkpoint file, acknowledgement state, model route, model-tier scheduler or automatic model switch. The choice is a one-time execution-cost affordance enabled by locked Authority and Final Gate protection; it is not acceptance evidence.
|
|
237
237
|
|
|
238
|
-
Post-lock revisions use three fail-closed paths. Proven monotonic/mechanical strengthening auto-adopts. A candidate whose only protected reasons are owner/change/support expansion may run existing active Check identities with unchanged runner/verifier authority through stateless `diagnose-revision`; safe monotonic strengthening may coexist, but those transient results write no authority, pending decision, Progress, cache or Receipt and cannot accept. Semantic changes, proof weakening, runner or verifier-content changes, and risk increases are preview-only; risk downgrade is rejected. Related edits remain in the same `delivery-contract.yaml` until one ordinary `compile --revise` emits an exact hash-bound approval summary; `status` and `resume` project that same pending decision. Exact adoption invalidates
|
|
238
|
+
Post-lock revisions use three fail-closed paths. Proven monotonic/mechanical strengthening auto-adopts. A candidate whose only protected reasons are owner/change/support expansion may run existing active Check identities with unchanged runner/verifier authority through stateless `diagnose-revision`; safe monotonic strengthening may coexist, but those transient results write no authority, pending decision, Progress, cache or Receipt and cannot accept. Semantic changes, proof weakening, runner or verifier-content changes, and risk increases are preview-only; risk downgrade is rejected. A rolling blocker alone cannot reclassify or remove machine-verifiable scope; a real scope change first becomes marked Source. Related edits remain in the same `delivery-contract.yaml` until one ordinary `compile --revise` emits an exact hash-bound approval summary containing changed semantic fields, Source/Product Claim reductions, proof reductions and external-confirmation keys; `status` and `resume` project that same pending decision. Exact adoption reports `delivery_completed_by_this_event: false`, invalidates affected evidence, returns to rolling implementation or repair and never replaces the complete current-snapshot Final Gate.
|
|
239
239
|
|
|
240
240
|
```text
|
|
241
241
|
ty-context long-task init <workdir>
|
|
@@ -255,13 +255,13 @@ ty-context long-task close <workdir>
|
|
|
255
255
|
ty-context long-task abandon <workdir> [--force-corrupt-state]
|
|
256
256
|
```
|
|
257
257
|
|
|
258
|
-
Compact authoring omits only deterministic defaults and normalizes identically to the expanded form. `preflight` is a read-only aggregated Source/REQ/CTRL/OBL/AC and repository check that creates no authority, state, Receipt or runner execution. Compile generates Global plus Outcome Result/Requirement/Control-field/Non-completing/Technical Claims, rejects uncovered Claims and makes the first successful formal Compile the Authority Lock. The first Compile result emits `execution_model_checkpoint.required: true`; later Compile revisions emit `required: false`. Every later authority change still compares with active authority regardless of progress, Receipt/cache deletion or implementation restoration. Source/Context/Product/Acceptance/Global/verifier content, resolved runners and verification inputs are frozen in the common-dir Active Authority V3 record.
|
|
258
|
+
Compact authoring omits only deterministic defaults and normalizes identically to the expanded form. `preflight` is a read-only aggregated Source/REQ/CTRL/OBL/AC and repository check that creates no authority, state, Receipt or runner execution. Compile generates Global plus Outcome Result/Requirement/Control-field/Non-completing/Technical Claims, rejects uncovered Claims and makes the first successful formal Compile the Authority Lock. Every Compile result includes a lifecycle event, `delivery_completed_by_this_event: false`, `native_goal_effect: none` and a next action. The first Compile result emits `execution_model_checkpoint.required: true`; later Compile revisions emit `required: false`. Every later authority change still compares with active authority regardless of progress, Receipt/cache deletion or implementation restoration. Source/Context/Product/Acceptance/Global/verifier content, resolved runners and verification inputs are frozen in the common-dir Active Authority V3 record.
|
|
259
259
|
|
|
260
|
-
`diagnose-revision` performs a side-effect-free candidate Compile and only exercises existing active Check identities whose runner/verifier authority is unchanged. Its output explicitly denies acceptance, Progress and pending-state writes. Protected `compile --revise` emits `authority_revision_pending`, the exact decision id and a deterministic
|
|
260
|
+
`diagnose-revision` performs a side-effect-free candidate Compile and only exercises existing active Check identities whose runner/verifier authority is unchanged. Its output explicitly denies acceptance, Progress and pending-state writes. Protected `compile --revise` emits `authority_revision_pending`, the exact decision id and a deterministic material summary before failing closed; approving a different or stale id is rejected. Adoption emits `authority_revision_adopted` and returns to rolling execution rather than completion.
|
|
261
261
|
|
|
262
262
|
Targeted verify rechecks active task/revision/compiled/worktree identity before writing scoped Progress. Counterfactual Findings first enter the owning Check Result, invalidate an otherwise passed Check, clear Claim Proofs and remain visible in status/resume; Global Checks reuse the same Progress type without a Global Outcome state. Final Gate repeats the identity check after all Checks; Stop/close clear only the accepted identity through CAS. Commit, migration, clear and abandon share one active-state lock. `abandon --force-corrupt-state` is reserved for corrupt continuity or stale lock cleanup and preserves Contract, Source, Context and Git content.
|
|
263
263
|
|
|
264
|
-
`status` and read-only `resume` report the current fresh Final Receipt as `final_workflow_status` (or `null` after drift) plus the active Contract's complete `external_confirmations`. `progress_passing` is targeted repair evidence rather than “Outcome complete”; `progress_stale` is not a current pass, and `final_workflow_status: null` means unfinished.
|
|
264
|
+
`status` and read-only `resume` report the current fresh Final Receipt as `final_workflow_status` (or `null` after drift) plus the active Contract's complete `external_confirmations`. `progress_passing` is targeted repair evidence rather than “Outcome complete”; `progress_stale` is not a current pass, and `final_workflow_status: null` means unfinished. Every accepted Stop emits one non-blocking terminal-scope `systemMessage`; external-pending results also name every confirmation. Final/Stop/close report `acceptance_scope: declared_machine_authority` and `native_goal_effect: none`; close also reports `closed_scope: machine_authority`. Before platform-native Goal completion, the Agent performs a veto-only Goal/user-to-Source conformance review that cannot create proof. `status: closed` means only that machine Authority was cleared, not that the native Goal or external delivery completed.
|
|
265
265
|
|
|
266
266
|
New authoring uses inline Outcomes. Existing `outcome_files` remains physical compatibility only and creates no semantic or completion boundary. A Long Task requires real Source, and every declared Source file contains at least one Material Item; background-only references remain outside Source Authority. Every Material Source Item is wrapped in the original Markdown with a non-rendering, uniquely keyed `ty-source-item:start/end` marker; `control` is a first-class kind, marker keys and `source_claim` keys are set-equal, and statements are text-exact after limited whitespace normalization. Every non-decision Source item owns one same-kind, same-text canonical target and duplicate ownership fails. Outcome Source Acceptance maps to criterion-identical `<outcome>.<check>.<assertion>` with an independently Source-backed non-Result Claim; Global Source Acceptance maps to criterion-identical `GLOBAL.<check>.<assertion>`, proves no Outcome Claim and needs an independently Source-backed Global Claim. Typed dispositions keep Requirements, Controls, Acceptance, Results, Fact/Affected-Outcome Risk, Non-goals, External Confirmations and Decisions distinct; `out_of_scope` is retired. Ordinary prose remains valid after marker-only enumeration.
|
|
267
267
|
|
|
@@ -271,6 +271,8 @@ Supported runners: `package_script`, `project_binary`, `node_oracle`, `playwrigh
|
|
|
271
271
|
|
|
272
272
|
Supported proof surfaces: `ui_browser`, `runtime_behavior`, `api_contract`, `data_state`, `security_boundary`, `population_coverage`, `implementation_structure`.
|
|
273
273
|
|
|
274
|
+
After a blocker-driven semantic or proof revision, only affected weak-observability or high-risk behavioral Claims receive a causal-boundary review. Evidence must reach the furthest independently failing declared boundary; when carrier existence can diverge from the claimed capability, use a capability-disrupting Counterfactual. This adds no product taxonomy, universal runtime suite, mutation type or persistent review state.
|
|
275
|
+
|
|
274
276
|
## Risk And Evidence
|
|
275
277
|
|
|
276
278
|
L0 local work stays on the default workflow. L1 standard long work uses the Delivery Contract. L2 strict is the minimum for public API/schema, persistent data, migration, security/permission boundaries, irreversible effects, full-population operations, or a critical path with weak observability. Strict proof binds to the affected Outcome; multi-repository delivery is rejected.
|
|
@@ -316,7 +318,7 @@ make validate-harness
|
|
|
316
318
|
|
|
317
319
|
The modularity gate is `ty-context check-modularity`. Scoped waivers require `owner`, `introduced_at`, `reason`, `tracking_issue` and `expiry_condition`.
|
|
318
320
|
|
|
319
|
-
The synchronized local preview tarball is named `project-tiny-context-harness-0.7.
|
|
321
|
+
The synchronized local preview tarball is named `project-tiny-context-harness-0.7.2.tgz`.
|
|
320
322
|
|
|
321
323
|
## Community And Further Reading
|
|
322
324
|
|
package/assets/README.md
CHANGED
|
@@ -137,7 +137,7 @@ The smoke packs the local workspace, installs it into a disposable repo and vali
|
|
|
137
137
|
|
|
138
138
|
```sh
|
|
139
139
|
cd /path/to/your/test-repo
|
|
140
|
-
npm install -D /path/to/project-tiny-context-harness/tmp/ty-context/source-preview/package/project-tiny-context-harness-0.7.
|
|
140
|
+
npm install -D /path/to/project-tiny-context-harness/tmp/ty-context/source-preview/package/project-tiny-context-harness-0.7.2.tgz
|
|
141
141
|
npx --no-install ty-context init --adopt
|
|
142
142
|
make validate-context
|
|
143
143
|
```
|
|
@@ -258,15 +258,15 @@ Before the first successful formal Compile, `delivery-contract.yaml` is one non-
|
|
|
258
258
|
|
|
259
259
|
The first successful Compile creates Authority Lock and returns `execution_model_checkpoint.required: true`. Before implementation, the Agent asks the user to `continue_current_model` or switch models and then resume the active Long-Task. A task-specific model strategy already stated explicitly satisfies the checkpoint. Later Compile revisions return `required: false`; Harness does not switch models, persist acknowledgement/model-route state or repeat the pause.
|
|
260
260
|
|
|
261
|
-
Later revisions are classified into three paths. Formally monotonic evidence strengthening and other proven mechanical-safe changes auto-adopt. A candidate whose only protected reasons are owner, expected-change or allowed-support expansion may be exercised through `diagnose-revision` using existing active Check identities whose runner and verifier are unchanged; safe monotonic strengthening may coexist, and the results remain transient repair diagnostics rather than Progress or acceptance. Product/Source/Acceptance semantic changes, proof weakening, verifier-content or runner changes, and risk increases are preview-only and require the exact revision identity; risk downgrade remains rejected outright. Diagnosis never changes the active Authority or writes pending/approval state, cache, Progress or Receipt, so related edits can accumulate in the same `delivery-contract.yaml` before one `compile --revise` approval request. The pending decision contains a concise hash-bound summary and is projected by `status`/`resume
|
|
261
|
+
Later revisions are classified into three paths. Formally monotonic evidence strengthening and other proven mechanical-safe changes auto-adopt. A candidate whose only protected reasons are owner, expected-change or allowed-support expansion may be exercised through `diagnose-revision` using existing active Check identities whose runner and verifier are unchanged; safe monotonic strengthening may coexist, and the results remain transient repair diagnostics rather than Progress or acceptance. Product/Source/Acceptance semantic changes, proof weakening, verifier-content or runner changes, and risk increases are preview-only and require the exact revision identity; risk downgrade remains rejected outright. A rolling blocker is not itself an External Confirmation or permission to remove machine-verifiable scope. A real scope change first becomes marked Source. Diagnosis never changes the active Authority or writes pending/approval state, cache, Progress or Receipt, so related edits can accumulate in the same `delivery-contract.yaml` before one `compile --revise` approval request. The pending decision contains a concise hash-bound summary with exact changed semantic fields, Source/Product Claim reductions, proof reductions and external-confirmation keys and is projected by `status`/`resume`. Adoption reports `delivery_completed_by_this_event: false`, invalidates affected evidence and returns to rolling implementation or repair; the complete Final Gate remains mandatory.
|
|
262
262
|
|
|
263
263
|
The package-managed Long-Task Skill uses progressive disclosure: its main `SKILL.md` keeps the objective, boundaries and phase routing; one-level references are read only for Contract authoring, evidence design or authority lifecycle. This reduces routine instruction load without moving any rule into a second authority. When Source or controlling Context declares an architecture invariant, the Contract uses existing technical obligations/global constraints/forbidden shortcuts, owner/path/Binding boundaries and a project-owned executable Check. Functional acceptance cannot substitute when the architecture invariant can fail independently.
|
|
264
264
|
|
|
265
265
|
A Draft Outcome is simply an Outcome before Authority Lock. Outcomes split independently observable, decidable and target-verifiable results so the current Goal can keep a smaller dependency-ready working set, target verification, localize failures, resume findings and invalidate stale local results. `depends_on` expresses acceptance readiness; the Rolling Frontier is temporary. An Outcome is not a Worker, scheduler task, queue or parallelism unit. Outcome decomposes execution and diagnosis, not completion authority: targeted passes never replace the one complete Final Gate on the current final snapshot.
|
|
266
266
|
|
|
267
|
-
When a declared result can pass on a proxy surface while failing in its target runtime, the earliest owning Outcome declares a project-owned Check that exercises the target during the current Check execution. A tracked report, screenshot, binary, log or historical run cannot be the sole runtime proof. The Goal runs
|
|
267
|
+
When a declared result can pass on a proxy surface while failing in its target runtime, the earliest owning Outcome declares a project-owned Check that exercises the target during the current Check execution. A tracked report, screenshot, binary, log or historical run cannot be the sole runtime proof. After a blocker-driven semantic/proof revision, only affected weak-observability or high-risk behavioral Claims pay a causal review: evidence reaches the furthest independently failing declared boundary, and a Counterfactual disrupts the claimed capability when carrier presence alone can diverge. The Goal runs the live Check after the first runnable slice and, after coalescing related edits, before dependent work grows when declared `input_paths` or Binding carriers make Progress stale. This uses existing targeted verification and Final Gate semantics: it adds no product taxonomy, `platform_impact` flags, universal runtime suite or completion state, requires no full rebuild per Outcome/edit, never accepts early and is rerun by Final Gate.
|
|
268
268
|
|
|
269
|
-
The platform owns physical Goal/session lifecycle. A later session runs `resume` to reconstruct semantic state; Tiny Context does not recreate the prior physical Turn.
|
|
269
|
+
The platform owns physical Goal/session lifecycle. A later session runs `resume` to reconstruct semantic state; Tiny Context does not recreate the prior physical Turn. Machine acceptance covers only `declared_machine_authority` and reports `native_goal_effect: none`. Before completing the platform-native Goal, the Agent performs a veto-only comparison of current Goal/user meaning against accepted marked Source and checks for pending revisions, unresolved blockers or omissions; this guard may block and repair, but it never supplies acceptance proof.
|
|
270
270
|
|
|
271
271
|
### CLI
|
|
272
272
|
|
|
@@ -290,14 +290,14 @@ ty-context long-task abandon <workdir> [--force-corrupt-state]
|
|
|
290
290
|
|
|
291
291
|
- `init` creates one Compact inline-Outcome Contract template.
|
|
292
292
|
- `preflight` applies Compact defaults and reports all discoverable Source/REQ/CTRL/OBL/AC, Context, risk, path/binding, runner/input and proof diagnostics. Exact duplicate diagnostics are merged with `occurrences`; known problems may include stable `refs` and a safe `repair_hint` that never weakens authority or invents product semantics. It is read-only: no Authority Lock, marker, cache, progress, Receipt, pending revision, state lock or project Check.
|
|
293
|
-
- `compile` generates Global plus Outcome Result/Requirement/Control-field/Non-completing/Technical Claims, rejects uncovered Claims, preserves an immutable first baseline and makes the first successful formal Compile the Authority Lock. The first result also includes `execution_model_checkpoint.required: true`; later Compile results return `required: false`. Every revision compares against active authority regardless of progress, Receipt/cache deletion or implementation restoration. Source/Context/Product/Acceptance/Global/verifier materials, owner/binding authority, resolved runners and verification inputs are frozen in the common-dir Active Authority V3 snapshot; the model-choice result is not stored as Authority state.
|
|
293
|
+
- `compile` generates Global plus Outcome Result/Requirement/Control-field/Non-completing/Technical Claims, rejects uncovered Claims, preserves an immutable first baseline and makes the first successful formal Compile the Authority Lock. Every result includes a lifecycle event, `delivery_completed_by_this_event: false`, `native_goal_effect: none` and a next action. The first result also includes `execution_model_checkpoint.required: true`; later Compile results return `required: false`. Every revision compares against active authority regardless of progress, Receipt/cache deletion or implementation restoration. Source/Context/Product/Acceptance/Global/verifier materials, owner/binding authority, resolved runners and verification inputs are frozen in the common-dir Active Authority V3 snapshot; the model-choice result is not stored as Authority state.
|
|
294
294
|
- `diagnose-revision` performs a side-effect-free candidate Compile. Only a scope-only candidate may run existing active Check identities with unchanged runner/verifier authority; semantic changes, proof weakening, runner or verifier-content changes, and risk increases are summarized without runner execution, while risk downgrade is rejected. Output always has `acceptance_authorized: false`, `progress_written: false` and `pending_revision_written: false`.
|
|
295
|
-
- `compile --revise` auto-adopts proven-safe revisions. Protected revisions return `authority_revision_pending` on stdout plus the exact decision id and deterministic approval summary, then fail closed until `approve-authority-revision` approves that exact id. Candidate edits produce a new id and invalidate the old approval.
|
|
295
|
+
- `compile --revise` auto-adopts proven-safe revisions. Protected revisions return `authority_revision_pending` on stdout plus the exact decision id and deterministic material approval summary, then fail closed until `approve-authority-revision` approves that exact id. Candidate edits produce a new id and invalidate the old approval. Adoption emits `authority_revision_adopted` and returns to rolling execution; it never means delivery completion.
|
|
296
296
|
- `verify` writes scoped per-Check Progress Records only after rechecking active task/revision/compiled/worktree identity. A concurrent revision returns `active_authority_changed_during_verify` and writes no stale progress.
|
|
297
297
|
- `status` reports each Outcome as `unverified`, `progress_passing`, `progress_failing`, `progress_stale` or `blocked_external`. It also reports the fresh Final Receipt as `final_workflow_status` (or `null` after drift), the active Contract's complete `external_confirmations` and the single `pending_authority_revision` decision when present. `progress_passing` is targeted repair evidence rather than “Outcome complete”; `progress_stale` is not a current pass, and `final_workflow_status: null` means unfinished. It reads the common-dir authority snapshot and reports a missing or mismatched workdir cache as a repairable diagnostic.
|
|
298
298
|
- `resume` is read-only and reports task identity, risk, relevant Context, Git state, the same Final/external/pending decision surfaces, ready Outcomes, findings and the next safe action from the common-dir authority snapshot.
|
|
299
299
|
- `final-gate` requires a clean candidate commit, recompiles source authority, reruns every required Check on one Git-tree snapshot and rechecks active identity before acceptance.
|
|
300
|
-
- `stop-check` and `close` run that Live Final Gate themselves. They never trust status, progress, a Receipt or compiled cache for acceptance; success clears only the accepted identity through CAS.
|
|
300
|
+
- `stop-check` and `close` run that Live Final Gate themselves. They never trust status, progress, a Receipt or compiled cache for acceptance; success clears only the accepted identity through CAS. Every accepted Stop emits one non-blocking terminal-scope `systemMessage`; external-pending results additionally name all confirmations. Final/Stop/close report `acceptance_scope: declared_machine_authority` and `native_goal_effect: none`; close also reports `closed_scope: machine_authority`. `status: closed` means only that machine Authority was cleared, not that the native Goal or complete external delivery finished.
|
|
301
301
|
- `abandon` is explicit non-success cleanup. `--force-corrupt-state` is reserved for invalid/mismatched/legacy-unrecoverable state or a stale active lock and removes only deterministic local active state plus `<workdir>/.ty-context/**`; Contract, Source, Context and Git content are preserved.
|
|
302
302
|
|
|
303
303
|
### Delivery Contract
|
|
@@ -463,7 +463,7 @@ make validate-harness
|
|
|
463
463
|
|
|
464
464
|
The modularity gate is `ty-context check-modularity`. Scoped waivers require `owner`, `introduced_at`, `reason`, `tracking_issue` and `expiry_condition`.
|
|
465
465
|
|
|
466
|
-
`npm run preview:pack` produces a local preview named `project-tiny-context-harness-0.7.
|
|
466
|
+
`npm run preview:pack` produces a local preview named `project-tiny-context-harness-0.7.2.tgz` under the preview output directory.
|
|
467
467
|
|
|
468
468
|
## Community And Further Reading
|
|
469
469
|
|
package/assets/README.zh-CN.md
CHANGED
|
@@ -163,15 +163,15 @@ Long-Task Contract Authoring 会尽量保留 Source 中已有的稳定 Key 与 A
|
|
|
163
163
|
|
|
164
164
|
Agent 此时在实现前只暂停一次,请用户选择:继续当前模型,或切换模型后恢复同一 active Long-Task。如果用户已明确给出本任务的模型策略,则视为已完成选择。后续 `compile --revise` 返回 `required: false`,不会重复暂停。Harness 不会自动切换模型,也不持久化 acknowledgement、model route 或 checkpoint state;模型选择不是验收证据。
|
|
165
165
|
|
|
166
|
-
锁定后的修订分三类:机器可证明的单调证据增强和机械安全变化自动采用;如果唯一的受保护原因只是扩大 owner、expected-change 或 allowed-support path(可以同时带有安全的单调增强),就能用 `diagnose-revision` 在不切换 Authority 的前提下运行原 Active Authority 已有且未更换的 Check;产品/Source/Acceptance 语义变化、证明弱化、verifier 内容或 runner
|
|
166
|
+
锁定后的修订分三类:机器可证明的单调证据增强和机械安全变化自动采用;如果唯一的受保护原因只是扩大 owner、expected-change 或 allowed-support path(可以同时带有安全的单调增强),就能用 `diagnose-revision` 在不切换 Authority 的前提下运行原 Active Authority 已有且未更换的 Check;产品/Source/Acceptance 语义变化、证明弱化、verifier 内容或 runner 变化、风险上升只给摘要,不运行候选,风险降级则直接拒绝。滚动实现遇阻本身不是 External Confirmation,也不允许删除机器可验证范围;真正的范围变化必须先成为 marked Source。诊断结果不是 Progress 或 acceptance,也不会写 pending/approval、cache、Receipt 或 marker。相关修改只在同一份 `delivery-contract.yaml` 中累计,最终由一次 `compile --revise` 生成精确 hash 与包含语义字段、Source/Product Claim 缩减、proof 缩减和 external-confirmation key 的短摘要;`status`/`resume` 投影同一个待批决策。批准并原子采用后返回 `delivery_completed_by_this_event: false`,旧证据失效并回到滚动实现或修复,完整 Final Gate 仍必须重跑。
|
|
167
167
|
|
|
168
168
|
Long-Task Skill 采用渐进读取:主 `SKILL.md` 只保留目标、硬边界和阶段路由,Contract Authoring、Evidence Design 与 Authority Lifecycle 细节只在对应阶段读取一层 reference。这只是指令组织,不产生第二权威。
|
|
169
169
|
|
|
170
170
|
Draft Outcome 只是 Authority Lock 前的 Outcome。Outcome 按可独立观察、判断和定向验证的结果拆分,使当前 Goal 能缩小 dependency-ready 工作集、定向验证、定位失败、恢复 finding 并精确失效旧局部结果。`depends_on` 只表示 acceptance readiness,Rolling Frontier 只是临时工作状态;Outcome 不是 Worker、scheduler task、queue 或并行单元。Outcome 拆分执行和诊断,不拆分完成权威,因此最终仍必须在当前最终快照运行一次完整 Final Gate。
|
|
171
171
|
|
|
172
|
-
如果一个声明结果可能在代理表面通过、却在目标运行时独立失败,最早拥有可运行边界的 Outcome 必须声明项目自有的真实运行 Check,并在当前 Check 执行中启动或触达目标、从同一会话产生结构化 Observation
|
|
172
|
+
如果一个声明结果可能在代理表面通过、却在目标运行时独立失败,最早拥有可运行边界的 Outcome 必须声明项目自有的真实运行 Check,并在当前 Check 执行中启动或触达目标、从同一会话产生结构化 Observation。仓库内状态报告、截图、二进制、日志或历史运行不能单独证明目标运行时。阻塞驱动的语义或 proof 修订只让受影响的 weak-observability/high-risk 行为 Claim 支付因果审查成本:证据必须到达 Claim 声明的最远独立失败边界;当 carrier 存在不等于能力成立时,Counterfactual 应破坏被声明的能力。当前 Goal 在第一个可运行切片后执行一次;后续相关修改先合并,在声明的 `input_paths` 或 Binding carrier 使 Progress stale 后、扩大依赖工作前再运行。它复用 targeted verify 与 Final Gate,不增加产品 taxonomy、`platform_impact` 字段、全局 runtime 套件或完成状态,不要求每个 Outcome/每次编辑完整重建,也不提前取得接受权;Final Gate 仍会重跑。
|
|
173
173
|
|
|
174
|
-
平台负责物理 Goal/会话生命周期。新会话通过 `resume` 恢复语义状态;Tiny Context 不会重建此前的物理 Turn
|
|
174
|
+
平台负责物理 Goal/会话生命周期。新会话通过 `resume` 恢复语义状态;Tiny Context 不会重建此前的物理 Turn。机器接受只覆盖 `declared_machine_authority`,并报告 `native_goal_effect: none`。完成平台原生 Goal 前,Agent 只做一次否决型核对:当前 Goal/用户语义是否全部进入 accepted marked Source,且没有 pending revision、未解 blocker 或遗漏;它只能阻止并触发修复,不能增加验收证据。
|
|
175
175
|
|
|
176
176
|
### CLI
|
|
177
177
|
|
|
@@ -195,14 +195,14 @@ ty-context long-task abandon <workdir> [--force-corrupt-state]
|
|
|
195
195
|
|
|
196
196
|
- `init` 创建单文件 inline Outcome 的 Compact Contract 模板。
|
|
197
197
|
- `preflight` 应用 Compact 默认值并一次输出 Source/REQ/CTRL/OBL/AC、Context、风险、路径/Binding、Runner/Input 与 Proof 诊断;它完全只读,不创建 Authority Lock、marker、cache、progress、Receipt、pending revision、状态锁,也不运行项目 Check。
|
|
198
|
-
- `compile` 生成 Global 与 Outcome Result/Requirement/Control-field/Non-completing/Technical Claim,拒绝未覆盖 Claim,并让第一次正式成功 Compile 成为 Authority Lock。第一次结果附带 `execution_model_checkpoint.required: true`,后续 Compile 返回 `false
|
|
198
|
+
- `compile` 生成 Global 与 Outcome Result/Requirement/Control-field/Non-completing/Technical Claim,拒绝未覆盖 Claim,并让第一次正式成功 Compile 成为 Authority Lock。每次结果都包含 lifecycle event、`delivery_completed_by_this_event: false`、`native_goal_effect: none` 和 next action。第一次结果附带 `execution_model_checkpoint.required: true`,后续 Compile 返回 `false`;这些字段不进入 Authority state。
|
|
199
199
|
- `diagnose-revision` 只做无副作用候选 Compile;仅 scope-only 候选能运行 Active Authority 已有且未更换的 Check,输出固定为非验收、非 Progress、非 pending。
|
|
200
|
-
- `compile --revise` 自动采用可证明安全的修订;受保护修订在 stdout 返回 `authority_revision_pending`、精确 decision id
|
|
200
|
+
- `compile --revise` 自动采用可证明安全的修订;受保护修订在 stdout 返回 `authority_revision_pending`、精确 decision id 与确定性 material 摘要,并继续 fail closed,直到用户批准完全相同的 id。候选内容再变会生成新 id,并使旧批准失效。采用后输出 `authority_revision_adopted` 并回到滚动执行,不表示交付完成。
|
|
201
201
|
- `verify` 在重查 active task/revision/compiled/worktree identity 后写 scoped Progress;targeted verify 始终只是修复证据。
|
|
202
202
|
- `status` 输出 `unverified`、`progress_passing`、`progress_failing`、`progress_stale` 或 `blocked_external`,并报告 fresh `final_workflow_status`、完整 `external_confirmations` 与唯一的 `pending_authority_revision`。`progress_passing` 只能表述为定向修复证据,不能简称“Outcome 完成”;`progress_stale` 不是当前通过,`final_workflow_status: null` 表示 Goal 尚未完成。
|
|
203
203
|
- `resume` 完全只读,恢复 task/contract identity、风险、相关 Context、Git 状态、同一待批决策、ready Outcome、findings 和 next safe action。
|
|
204
204
|
- `final-gate` 在完整 Check 后再次验证 active identity;并发 revision 不能产生 accepted。
|
|
205
|
-
- `stop-check` 与 `close` 自己运行 Live Final Gate,并只用 accepted identity 做 CAS clear
|
|
205
|
+
- `stop-check` 与 `close` 自己运行 Live Final Gate,并只用 accepted identity 做 CAS clear。每次机器接受的 Stop 都给一个非阻塞 terminal-scope `systemMessage`;外部待确认时同时列出全部确认项。Final/Stop/close 输出 `acceptance_scope: declared_machine_authority` 与 `native_goal_effect: none`,close 另输出 `closed_scope: machine_authority`。`status: closed` 只表示机器 Authority 已清理,不表示原生 Goal 或完整外部交付完成。
|
|
206
206
|
- `abandon --force-corrupt-state` 仅用于损坏/mismatch/legacy-unrecoverable 状态或遗留锁,只删除确定性 active state 与 `<workdir>/.ty-context/**`。
|
|
207
207
|
|
|
208
208
|
### Delivery Contract
|
|
@@ -35,7 +35,7 @@ After the first Authority Lock, stop once before implementation and ask the user
|
|
|
35
35
|
|
|
36
36
|
Before authoring, proof design or authority lifecycle work, read the phase-specific references in the package-managed `long-task-workflow` Skill. Use `ty-context long-task help` for CLI syntax instead of treating this startup router as a command reference.
|
|
37
37
|
|
|
38
|
-
Final Gate, Stop and close recompile the source Contract and rerun every declared Check on one clean current snapshot. Targeted verify is repair evidence only. Status, progress, receipts and compiled cache are audit/recovery surfaces only; prose, historical tests or Agent judgment never create acceptance. External confirmations remain explicit
|
|
38
|
+
Final Gate, Stop and close recompile the source Contract and rerun every declared Check on one clean current snapshot. Targeted verify is repair evidence only. Status, progress, receipts and compiled cache are audit/recovery surfaces only; prose, historical tests or Agent judgment never create acceptance. An adopted Authority Revision returns to rolling execution and is never delivery completion. External confirmations remain explicit; machine acceptance covers declared machine Authority and cannot by itself authorize completing the platform-native Goal, CI, deployment or human acceptance.
|
|
39
39
|
|
|
40
40
|
Tiny Context does not create or restore platform Goals, invoke models, spawn agents, call an App Server, create branches/worktrees, merge, push, open PRs, deploy or manage process trees. `ty-context enable long-task` installs the Source Plan Authoring Skill, Long-Task Workflow Skill and package-owned completion Hook.
|
|
41
41
|
|
|
@@ -9,7 +9,7 @@ description: Author, preflight, execute, resume, verify, or close one complete S
|
|
|
9
9
|
|
|
10
10
|
Use one current native Goal, one repository, one selected workspace, one complete Contract and one Final Gate. Never create a scheduler, model worker, agent runtime, App Server, branch, worktree, merge, push, PR, deployment, Campaign/SFC/Packet/Wave chain, matrix, verdict or second Contract plan. Never activate from task size alone.
|
|
11
11
|
|
|
12
|
-
The host and user own model selection. The workflow has exactly one user-choice checkpoint after the first Authority Lock and before implementation; Harness neither switches the model nor persists model-routing/checkpoint state. No checkpoint file, acknowledgement state, model route, model-tier scheduler
|
|
12
|
+
The host and user own model selection and native-Goal lifecycle. The workflow has exactly one user-choice checkpoint after the first Authority Lock and before implementation; Harness neither switches the model nor persists model-routing/checkpoint state. No checkpoint file, acknowledgement state, model route, model-tier scheduler, automatic model switch, `authority_revision_in_progress` state or native-Goal completion state is created. Outside that boundary, do not pause a healthy Goal solely to change or downgrade the model. Do not create a separate approval checkpoint for a defensible recommended plan choice. A targeted pre-Authority clarification is still required when a missing user preference could materially change research or selection; genuine Source conflicts or choices the user explicitly reserves may likewise require a decision before Authority Lock. Capability-related drift is handled by targeted repair plus the Final Gate. Never proactively spawn, assign or coordinate parallel subagents. Platform-native internal delegation, if it occurs, is opaque and non-authoritative and must converge into the unified current workspace snapshot before verification can count.
|
|
13
13
|
|
|
14
14
|
`long-task-delivery-v2` is the only active Contract schema. `delivery-contract.yaml` is the root authoring file. New authoring uses inline Outcomes; existing `outcome_files` are physical compatibility only. `delivery-set` is retired and non-executing.
|
|
15
15
|
|
|
@@ -17,7 +17,7 @@ The host and user own model selection. The workflow has exactly one user-choice
|
|
|
17
17
|
|
|
18
18
|
Prevent false completion inside declared authority. Implementation may drift, fail or require rework, but every declared non-Result requirement and AC must remain traceable and every unsatisfied, unverifiable, insufficiently evidenced or stale item must block completion. Findings should localize repair through Source Item, Outcome, Claim, Assertion, Check, Proof Surface, Binding and owner boundary.
|
|
19
19
|
|
|
20
|
-
Only fresh evidence from the complete current final snapshot may create machine acceptance. Otherwise report the task as unfinished or qualified. `machine_accepted_external_pending` means machine-verifiable authority passed while named external confirmation remains; it is not full delivery completion. Never substitute prose, progress, historical tests, Receipts, one exit code or Agent judgment for the Final Gate.
|
|
20
|
+
Only fresh evidence from the complete current final snapshot may create machine acceptance. Otherwise report the task as unfinished or qualified. `machine_accepted_external_pending` means machine-verifiable authority passed while named external confirmation remains; it is not full delivery completion. Machine acceptance covers declared machine Authority and has no direct native-Goal effect. Never substitute prose, progress, historical tests, Receipts, one exit code or Agent judgment for the Final Gate.
|
|
21
21
|
|
|
22
22
|
Prefer the lowest practical Authoring, Runtime, State, Recovery and verification cost that preserves the same false-completion interception. Add no mechanism whose distinct protection does not materially exceed its total cost.
|
|
23
23
|
|
|
@@ -66,7 +66,7 @@ Use targeted `verify --outcome/--check` only to drive repair. Progress is repair
|
|
|
66
66
|
|
|
67
67
|
When the Contract declares a target-runtime Check because a proxy can pass while the target fails independently, run it at the earliest owning Outcome's first runnable boundary. After accumulated changes to its declared `input_paths` or Binding carriers make the result stale, rerun it before dependent work grows. Coalesce related edits and use the cheapest reliable target Check; do not mandate a full environment rebuild per Outcome or per edit. This is rolling feedback through existing targeted verify, not acceptance, a trigger queue, platform taxonomy or new state.
|
|
68
68
|
|
|
69
|
-
When implementation discovers missing Contract paths, first classify the revision. Proven monotonic evidence strengthening may use ordinary `compile --revise` directly. If every protected reason is only owner/expected-change/allowed-support expansion, continue editing the same `delivery-contract.yaml` and use `ty-context long-task diagnose-revision <workdir> [--outcome <key>] [--check <key>]` to exercise only existing active Check identities with unchanged runner/verifier authority; safe monotonic strengthening may coexist. Candidate diagnostics are transient: they authorize no acceptance and write no pending/approval state, Active Authority, cache, Progress or Receipt. Semantic changes, proof weakening, runner or verifier-content changes, and risk-increase candidates are preview-only and must not run; risk downgrade is rejected. When the candidate is complete, run ordinary `compile --revise` once, present its exact
|
|
69
|
+
When implementation discovers a blocker or missing Contract paths, first classify the revision. Difficulty or delay alone never reclassifies machine-verifiable scope as external and never removes Source; a real scope, Product, Acceptance or machine/external boundary change must first be explicit marked Source. Proven monotonic evidence strengthening may use ordinary `compile --revise` directly. If every protected reason is only owner/expected-change/allowed-support expansion, continue editing the same `delivery-contract.yaml` and use `ty-context long-task diagnose-revision <workdir> [--outcome <key>] [--check <key>]` to exercise only existing active Check identities with unchanged runner/verifier authority; safe monotonic strengthening may coexist. Candidate diagnostics are transient: they authorize no acceptance and write no pending/approval state, Active Authority, cache, Progress or Receipt. Semantic changes, proof weakening, runner or verifier-content changes, and risk-increase candidates are preview-only and must not run; risk downgrade is rejected. When the candidate is complete, run ordinary `compile --revise` once, present its exact material decision summary to the user, and never approve it yourself. Keep the previous Authority active until exact approval and atomic adoption. Adoption is not delivery completion: discard historical/candidate evidence, run `status` or `resume`, and return to rolling implementation or repair under the revised Authority before Final Gate.
|
|
70
70
|
|
|
71
71
|
## Live Final Authority
|
|
72
72
|
|
|
@@ -74,8 +74,10 @@ Complete Context, implementation and project tests, create a clean candidate com
|
|
|
74
74
|
|
|
75
75
|
Final Gate recompiles Source authority, validates active task/revision/compiled/worktree identity, creates one Git-tree snapshot, reruns every required Global and Outcome Check and rechecks active identity before acceptance. A target-runtime Check must exercise its target in that current Gate execution; rerunning a reader for a historical or tracked status report is not live target proof. Final Gate, Stop and close never trust historical Progress, Receipt or compiled cache.
|
|
76
76
|
|
|
77
|
-
Machine acceptance covers only declared machine authority. Preserve every pending external confirmation through `final-gate`, `status`, `resume`, `stop-check`, the package-owned Stop Hook and `close`; `
|
|
77
|
+
Machine acceptance covers only declared machine authority. Preserve every pending external confirmation through `final-gate`, `status`, `resume`, `stop-check`, the package-owned Stop Hook and `close`; accepted output identifies `acceptance_scope: declared_machine_authority` and `native_goal_effect: none`, while `closed_scope: machine_authority` means only Authority cleanup. Do not invent external-confirmation or native-Goal tracking state.
|
|
78
|
+
|
|
79
|
+
Before invoking platform-native Goal completion, perform one veto-only conformance review: compare the current Goal and user instructions with accepted marked Source, and check for pending revisions, unresolved blockers or omitted requirements. Any mismatch keeps the Goal active and returns to Source/Contract repair. A clean review does not add acceptance proof and never lets Agent judgment replace Final Gate.
|
|
78
80
|
|
|
79
81
|
## Handoff
|
|
80
82
|
|
|
81
|
-
Report implementation, effective risk, Claim Coverage, Live Gate result, every pending external confirmation, Context status and blockers. Use verifier terms exactly: `progress_passing` means targeted repair evidence, `progress_stale` is not a current pass, `final_workflow_status: null` means unfinished, and `machine_accepted_external_pending` must retain its named confirmations. Never shorten implementation or targeted progress to “Outcome complete” or invent `implementation_complete`, `platform_smoke_verified` or another persistent status. State the threat-model limits: undeclared requirements cannot be discovered, installed verifier/Git metadata are trusted, model selection belongs to the host/user, and internal platform delegation is not observed.
|
|
83
|
+
Report implementation, effective risk, Claim Coverage, Live Gate result, acceptance scope, every pending external confirmation, Context status and blockers. Use verifier terms exactly: `progress_passing` means targeted repair evidence, `progress_stale` is not a current pass, `final_workflow_status: null` means unfinished, `authority_revision_adopted` means return to rolling execution, and `machine_accepted_external_pending` must retain its named confirmations. Never shorten implementation or targeted progress to “Outcome complete” or invent `implementation_complete`, `platform_smoke_verified` or another persistent status. State the threat-model limits: undeclared requirements cannot be discovered, installed verifier/Git metadata are trusted, native-Goal/model selection belongs to the host/user, and internal platform delegation is not observed.
|
|
@@ -24,7 +24,7 @@ After Authority Lock, every revision compares against active authority and follo
|
|
|
24
24
|
|
|
25
25
|
`diagnose-revision` recompiles the same `delivery-contract.yaml` in memory, creates only a disposable workspace snapshot when class 2 is proven, and returns transient repair results with `acceptance_authorized: false`. It writes no pending/approval state, authority/marker, cache, Progress or Receipt. Repeated edits therefore accumulate only in the one existing Contract authoring file, not a pending Draft authority or candidate state plane.
|
|
26
26
|
|
|
27
|
-
Ordinary `compile --revise` is the only operation that may create the one pending decision. It binds a deterministic concise change summary into the revision identity. `status` and `resume` expose that same decision so the host can deduplicate the user prompt without a Harness-owned waiting state. The executing Agent never approves its own pending revision; earlier blanket authorization cannot approve a later exact identity. If the candidate changes, the identity changes and old approval is rejected. The previous Authority remains active until approved compare-and-swap adoption
|
|
27
|
+
Ordinary `compile --revise` is the only operation that may create the one pending decision. It binds a deterministic concise change summary into the revision identity and enumerates changed semantic fields, Source/Product Claim reductions, proof reductions and external-confirmation keys. `status` and `resume` expose that same decision so the host can deduplicate the user prompt without a Harness-owned waiting state. The executing Agent never approves its own pending revision; earlier blanket authorization cannot approve a later exact identity. If the candidate changes, the identity changes and old approval is rejected. The previous Authority remains active until approved compare-and-swap adoption. Adoption reports `delivery_completed_by_this_event: false`, invalidates affected evidence and returns to rolling implementation or repair under the revised Authority; the complete source-recompiled Final Gate remains mandatory.
|
|
28
28
|
|
|
29
29
|
Every path-bearing field uses canonical grammar. Internal `.`/`..`, control characters, empty segments, absolute/drive/UNC paths and unsupported glob syntax fail closed.
|
|
30
30
|
|
|
@@ -48,6 +48,8 @@ Report their exact meaning: `progress_passing` is current targeted repair eviden
|
|
|
48
48
|
|
|
49
49
|
Before Final Gate, complete Context/code/tests and create a clean candidate commit. Final Gate captures active identity, recompiles Source authority, reads complete current Context, validates common-dir record/marker, creates a Git-tree snapshot, reruns all Checks and sensitivity controls and rechecks identity before acceptance. A target-runtime Check must exercise its target again in that Final Gate execution; rereading historical status does not become live proof merely because the reader reran. A concurrent revision returns `active_authority_changed_during_final_gate`.
|
|
50
50
|
|
|
51
|
-
Commit, verifier migration, clear and abandon share one active-state lock. Stop/close clear only the identity actually accepted through CAS and preserve `machine_accepted_external_pending` plus every named external confirmation in output. A stale Receipt exposes no accepted workflow status.
|
|
51
|
+
Commit, verifier migration, clear and abandon share one active-state lock. Stop/close clear only the identity actually accepted through CAS and preserve `machine_accepted_external_pending` plus every named external confirmation in output. Final Gate/Stop/close identify `acceptance_scope: declared_machine_authority` and `native_goal_effect: none`; close additionally identifies `closed_scope: machine_authority`. The Stop Hook emits the same scope as one non-blocking message for either accepted machine status. A stale Receipt exposes no accepted workflow status.
|
|
52
|
+
|
|
53
|
+
Before platform-native Goal completion, compare current Goal/user meaning with accepted marked Source and check for a pending revision, unresolved blocker or omitted requirement. This review may only veto completion and direct Source/Contract repair; it is not a second acceptance Gate and cannot create proof.
|
|
52
54
|
|
|
53
55
|
For invalid, mismatched, unrecoverable or stale-lock continuity, use only `ty-context long-task abandon <workdir> --force-corrupt-state`; it preserves authored Contract, Source, Context and Git content.
|
|
@@ -13,6 +13,7 @@ Read this only while authoring or structurally revising the one `delivery-contra
|
|
|
13
13
|
- `delegated` in a Source Plan is provenance, not a Contract disposition or new Claim kind. An instruction to synthesize, refine, complete, implement or use judgment delegates plan-level authoring, but it does not invent material tradeoff preferences. Before comparative research or a material product, technical, architecture or provider selection, identify the criteria that could change the research scope, candidate set or recommendation. If such a preference is unknown or ambiguous, ask a concise targeted question before research or selection and keep the item `decision_required` until answered; do not impose a fixed questionnaire or re-ask preferences already supplied by the user, Source, Context or controlling constraints.
|
|
14
14
|
- Once the material preference envelope is clear, use current authoritative or primary evidence for external capability, price, quota, license, compatibility, region, security posture or support claims. When one defensible recommendation exists, record the authoring instruction, preference/evidence or conservative-default basis and exact added meaning in real Source, then preserve that keyed item as ordinary Source of its semantic kind. If ordinary prose is the Source, append the delegated item without rewriting the user's original text; never place the choice only in Contract YAML.
|
|
15
15
|
- A delegated plan choice is not action authorization. Payment, contracting, production deployment/publication, destructive production mutation, real permission grants, sensitive-data transmission and required legal/security/human approval remain named External Confirmations. Conflicting authority, an explicitly user-reserved choice, a missing material preference or the absence of a defensible recommendation remains `decision_required`; high impact or multiple options with known criteria alone does not.
|
|
16
|
+
- A rolling implementation blocker is not an External Confirmation merely because work is difficult, delayed or unavailable through the current implementation path. Reclassify or remove machine-verifiable scope only through an explicit marked Source change and protected exact approval; otherwise keep the requirement and revise the implementation/evidence path.
|
|
16
17
|
|
|
17
18
|
## Outcome Boundary
|
|
18
19
|
|
|
@@ -25,6 +25,13 @@ Across all Checks sharing a Raw Execution, one Claim-bearing Observation belongs
|
|
|
25
25
|
- Historical reports, screenshots, binaries and logs are review material. Current-run screenshots/logs may accompany a Check as Artifacts, but the accepting Observation must come from the live runner execution and cannot be imported from historical state.
|
|
26
26
|
- Bind every runtime-affecting implementation surface through `input_paths` and relevant Binding carriers; keep runner/helper/config files in `verification_inputs`. This lets existing Progress freshness identify when rolling feedback is stale without a new trigger registry.
|
|
27
27
|
|
|
28
|
+
## Causal Boundary Review After Revision
|
|
29
|
+
|
|
30
|
+
- When a rolling blocker causes a semantic or proof revision, review only the affected weak-observability or high-risk Outcomes before adoption. Ask whether a cheaper proxy, fixed response or self-reported success could pass while the declared result still fails at a farther independent boundary.
|
|
31
|
+
- Evidence must reach the furthest independently failing boundary named by the Claim. A proxy may prove its own result, but it cannot prove a downstream state or effect merely by reporting success.
|
|
32
|
+
- For a behavioral Claim, prefer a Counterfactual that disrupts the claimed causal capability when removing a carrier would prove only file dependence. `replace_file` may supply a declared inert/failing implementation fixture; `remove_paths` remains valid when carrier existence is itself the claimed boundary.
|
|
33
|
+
- Keep this risk-proportional and internal. Do not create an evidence matrix, product-effect taxonomy, universal restart/end-to-end suite, new mutation type or persistent review state.
|
|
34
|
+
|
|
28
35
|
## Playwright
|
|
29
36
|
|
|
30
37
|
Claim-bearing Playwright proof is only `playwright.case.<ac-key>.passed equals true`. `[ac:<assertion-key>]` binds one declared AC per Test Instance; ordinary tags are ignored and legacy `[<key>]` binds only a declared key.
|
|
@@ -63,8 +63,21 @@ async function compile(workdir, args) {
|
|
|
63
63
|
await clearFinalReceipt(compiled.repository_root, workdir);
|
|
64
64
|
}
|
|
65
65
|
await clearAuthorityRevision(workdir);
|
|
66
|
+
printCompileResult(compiled, previous, preserveProgress, revisionCapture.proposal);
|
|
67
|
+
}
|
|
68
|
+
function printCompileResult(compiled, previous, preserveProgress, revisionProposal) {
|
|
69
|
+
const firstAuthorityLock = previous === null;
|
|
70
|
+
const authorityChanged = previous !== null &&
|
|
71
|
+
previous.compiled_identity !== compiled.compiled_identity;
|
|
66
72
|
console.log(JSON.stringify({
|
|
67
73
|
status: "compiled",
|
|
74
|
+
lifecycle_event: firstAuthorityLock
|
|
75
|
+
? "authority_locked"
|
|
76
|
+
: authorityChanged
|
|
77
|
+
? "authority_revision_adopted"
|
|
78
|
+
: "authority_recompiled_unchanged",
|
|
79
|
+
delivery_completed_by_this_event: false,
|
|
80
|
+
native_goal_effect: "none",
|
|
68
81
|
task_id: compiled.task.id,
|
|
69
82
|
compiled_identity: compiled.compiled_identity,
|
|
70
83
|
authority_revision: compiled.authority_revision,
|
|
@@ -72,10 +85,15 @@ async function compile(workdir, args) {
|
|
|
72
85
|
outcomes: compiled.outcomes.map((outcome) => outcome.key),
|
|
73
86
|
claim_coverage: compiled.claim_coverage,
|
|
74
87
|
progress_preserved: preserveProgress,
|
|
75
|
-
authority_revision_change:
|
|
76
|
-
? projectAuthorityRevisionDecision(
|
|
88
|
+
authority_revision_change: revisionProposal
|
|
89
|
+
? projectAuthorityRevisionDecision(revisionProposal)
|
|
77
90
|
: null,
|
|
78
|
-
|
|
91
|
+
next_action: firstAuthorityLock
|
|
92
|
+
? "Complete the one-time model choice, then begin rolling implementation."
|
|
93
|
+
: authorityChanged
|
|
94
|
+
? "Run status or resume, then continue rolling implementation or repair under the adopted Authority Revision."
|
|
95
|
+
: "Continue rolling implementation or repair under the active Authority.",
|
|
96
|
+
execution_model_checkpoint: executionModelCheckpoint(firstAuthorityLock),
|
|
79
97
|
}));
|
|
80
98
|
}
|
|
81
99
|
async function compileForCommand(workdir, revise, previous, capture) {
|
|
@@ -101,8 +119,11 @@ async function printPendingDecision(workdir, previous) {
|
|
|
101
119
|
console.log(JSON.stringify({
|
|
102
120
|
status: "authority_revision_pending",
|
|
103
121
|
acceptance_authorized: false,
|
|
122
|
+
delivery_completed_by_this_event: false,
|
|
123
|
+
native_goal_effect: "none",
|
|
104
124
|
active_compiled_identity: previous?.compiled_identity ?? null,
|
|
105
125
|
pending_authority_revision: projectAuthorityRevisionDecision(pending),
|
|
126
|
+
next_action: "Ask the user to approve or reject this exact material revision; keep the previous Authority active.",
|
|
106
127
|
}));
|
|
107
128
|
}
|
|
108
129
|
async function diagnoseRevision(workdir, args) {
|
|
@@ -120,7 +141,13 @@ async function approveRevision(workdir, args) {
|
|
|
120
141
|
if (!revision)
|
|
121
142
|
throw new Error("--revision requires a value");
|
|
122
143
|
await approvePendingAuthorityRevision(workdir, revision);
|
|
123
|
-
console.log(JSON.stringify({
|
|
144
|
+
console.log(JSON.stringify({
|
|
145
|
+
status: "authority_revision_approved",
|
|
146
|
+
revision,
|
|
147
|
+
delivery_completed_by_this_event: false,
|
|
148
|
+
native_goal_effect: "none",
|
|
149
|
+
next_action: "Run compile --revise to atomically adopt the approved revision, then return to rolling implementation or repair.",
|
|
150
|
+
}));
|
|
124
151
|
}
|
|
125
152
|
function executionModelCheckpoint(firstAuthorityLock) {
|
|
126
153
|
if (!firstAuthorityLock)
|
|
@@ -68,6 +68,9 @@ export async function longTask(args) {
|
|
|
68
68
|
workdir,
|
|
69
69
|
workflow_status: result.workflow_status,
|
|
70
70
|
external_confirmations: result.external_confirmations,
|
|
71
|
+
acceptance_scope: result.acceptance_scope,
|
|
72
|
+
closed_scope: result.closed_scope,
|
|
73
|
+
native_goal_effect: result.native_goal_effect,
|
|
71
74
|
}));
|
|
72
75
|
return;
|
|
73
76
|
}
|
|
@@ -101,7 +104,11 @@ async function verify(workdir, args) {
|
|
|
101
104
|
async function finalGate(workdir, args) {
|
|
102
105
|
rejectUnknown(args, []);
|
|
103
106
|
const result = await runDeliveryFinalGate(workdir);
|
|
104
|
-
console.log(JSON.stringify(
|
|
107
|
+
console.log(JSON.stringify({
|
|
108
|
+
...result,
|
|
109
|
+
acceptance_scope: "declared_machine_authority",
|
|
110
|
+
native_goal_effect: "none",
|
|
111
|
+
}));
|
|
105
112
|
if (result.workflow_status !== "machine_accepted" &&
|
|
106
113
|
result.workflow_status !== "machine_accepted_external_pending")
|
|
107
114
|
process.exitCode = 1;
|
|
@@ -57,6 +57,21 @@ export function summarizeAuthorityRevision(diff, outcomeKeys) {
|
|
|
57
57
|
write_scope_expanded: diff.owner_or_path_boundary_changed,
|
|
58
58
|
risk_changed: diff.risk_changed,
|
|
59
59
|
external_confirmations_changed: diff.external_confirmations_changed,
|
|
60
|
+
semantic_fields_changed: uniqueSorted([
|
|
61
|
+
...diff.product_semantics_changed,
|
|
62
|
+
...diff.global_semantics_changed,
|
|
63
|
+
]),
|
|
64
|
+
source_claim_changes: uniqueSorted([
|
|
65
|
+
...diff.source_claims_added,
|
|
66
|
+
...diff.source_claims_removed_or_changed,
|
|
67
|
+
]),
|
|
68
|
+
product_claim_changes: uniqueSorted([
|
|
69
|
+
...diff.product_claims_added.map((claim) => `${claim}:added`),
|
|
70
|
+
...diff.product_claims_removed.map((claim) => `${claim}:removed`),
|
|
71
|
+
...diff.product_claims_changed.map((claim) => `${claim}:changed`),
|
|
72
|
+
]),
|
|
73
|
+
proof_reductions: uniqueSorted(diff.reduction_reasons.filter((reason) => PROOF_REDUCTION_REASONS.has(reason))),
|
|
74
|
+
external_confirmation_changes: uniqueSorted(diff.external_confirmation_changes),
|
|
60
75
|
added_verification_dependencies: uniqueSorted([
|
|
61
76
|
...diff.verification_inputs_added,
|
|
62
77
|
...diff.input_paths_added,
|
|
@@ -76,13 +91,29 @@ export function projectAuthorityRevisionDecision(value) {
|
|
|
76
91
|
verification_inputs_added: value.revision_diff.verification_inputs_added ?? [],
|
|
77
92
|
input_paths_added: value.revision_diff.input_paths_added ?? [],
|
|
78
93
|
external_confirmations_changed: value.revision_diff.external_confirmations_changed ?? false,
|
|
94
|
+
external_confirmation_changes: value.revision_diff.external_confirmation_changes ?? [],
|
|
79
95
|
};
|
|
96
|
+
const computedSummary = summarizeAuthorityRevision(diff, value.affected_outcomes_or_contracts);
|
|
97
|
+
const storedSummary = value.approval_summary;
|
|
80
98
|
return {
|
|
81
99
|
revision_identity: value.revision_identity,
|
|
82
100
|
change_class: value.change_class ?? classifyAuthorityRevision(diff),
|
|
83
101
|
approval_required: value.approval_required ?? true,
|
|
84
|
-
approval_summary:
|
|
85
|
-
|
|
102
|
+
approval_summary: storedSummary
|
|
103
|
+
? {
|
|
104
|
+
...computedSummary,
|
|
105
|
+
...storedSummary,
|
|
106
|
+
semantic_fields_changed: storedSummary.semantic_fields_changed ??
|
|
107
|
+
computedSummary.semantic_fields_changed,
|
|
108
|
+
source_claim_changes: storedSummary.source_claim_changes ??
|
|
109
|
+
computedSummary.source_claim_changes,
|
|
110
|
+
product_claim_changes: storedSummary.product_claim_changes ??
|
|
111
|
+
computedSummary.product_claim_changes,
|
|
112
|
+
proof_reductions: storedSummary.proof_reductions ?? computedSummary.proof_reductions,
|
|
113
|
+
external_confirmation_changes: storedSummary.external_confirmation_changes ??
|
|
114
|
+
computedSummary.external_confirmation_changes,
|
|
115
|
+
}
|
|
116
|
+
: computedSummary,
|
|
86
117
|
};
|
|
87
118
|
}
|
|
88
119
|
function scopeAffectedOutcomes(diff) {
|
|
@@ -10,6 +10,11 @@ export interface AuthorityRevisionApprovalSummaryV2 {
|
|
|
10
10
|
write_scope_expanded: boolean;
|
|
11
11
|
risk_changed: boolean;
|
|
12
12
|
external_confirmations_changed: boolean;
|
|
13
|
+
semantic_fields_changed: string[];
|
|
14
|
+
source_claim_changes: string[];
|
|
15
|
+
product_claim_changes: string[];
|
|
16
|
+
proof_reductions: string[];
|
|
17
|
+
external_confirmation_changes: string[];
|
|
13
18
|
added_verification_dependencies: string[];
|
|
14
19
|
expanded_owner_paths: string[];
|
|
15
20
|
expanded_expected_change_paths: string[];
|
|
@@ -79,6 +84,7 @@ export interface AuthorityRevisionDiffV2 {
|
|
|
79
84
|
counterfactuals_removed: string[];
|
|
80
85
|
population_weakened: string[];
|
|
81
86
|
external_confirmations_changed: boolean;
|
|
87
|
+
external_confirmation_changes: string[];
|
|
82
88
|
verifier_content_changed: boolean;
|
|
83
89
|
verifier_runtime_locator_changed: boolean;
|
|
84
90
|
verifier_files_changed: string[];
|
|
@@ -80,7 +80,8 @@ export function authorityRevisionDiff(previous, next, nextHashes, nextMaterials,
|
|
|
80
80
|
const riskChanged = previous.authority_hashes.risk_authority_hash !==
|
|
81
81
|
nextHashes.risk_authority_hash;
|
|
82
82
|
const acceptanceChanged = acceptanceSemanticsChanged(previous, next);
|
|
83
|
-
const
|
|
83
|
+
const externalConfirmationChanges = keyedAuthorityChanges(previous.global.acceptance.external_confirmations, next.global.acceptance.external_confirmations);
|
|
84
|
+
const externalConfirmationsChanged = externalConfirmationChanges.length > 0;
|
|
84
85
|
const monotonic = isMonotonicAcceptanceStrengthening(previous, next);
|
|
85
86
|
const reductionReasons = [
|
|
86
87
|
...(productClaimsAdded.length ? ["product_claim_added"] : []),
|
|
@@ -175,6 +176,7 @@ export function authorityRevisionDiff(previous, next, nextHashes, nextMaterials,
|
|
|
175
176
|
counterfactuals_removed: counterfactualsRemoved,
|
|
176
177
|
population_weakened: populationWeakened,
|
|
177
178
|
external_confirmations_changed: externalConfirmationsChanged,
|
|
179
|
+
external_confirmation_changes: externalConfirmationChanges,
|
|
178
180
|
...verifierDiff,
|
|
179
181
|
source_claims_changed: previous.authority_hashes.source_authority_hash !==
|
|
180
182
|
nextHashes.source_authority_hash,
|
|
@@ -194,3 +196,21 @@ export function authorityRevisionDiff(previous, next, nextHashes, nextMaterials,
|
|
|
194
196
|
reduction_reasons: [...new Set(reductionReasons)],
|
|
195
197
|
};
|
|
196
198
|
}
|
|
199
|
+
function keyedAuthorityChanges(before, after) {
|
|
200
|
+
const beforeByKey = new Map(before.map((item) => [item.key, item]));
|
|
201
|
+
const afterByKey = new Map(after.map((item) => [item.key, item]));
|
|
202
|
+
return [
|
|
203
|
+
...before
|
|
204
|
+
.filter((item) => !afterByKey.has(item.key))
|
|
205
|
+
.map((item) => `${item.key}:removed`),
|
|
206
|
+
...after
|
|
207
|
+
.filter((item) => !beforeByKey.has(item.key))
|
|
208
|
+
.map((item) => `${item.key}:added`),
|
|
209
|
+
...before
|
|
210
|
+
.filter((item) => {
|
|
211
|
+
const candidate = afterByKey.get(item.key);
|
|
212
|
+
return candidate !== undefined && !same(item, candidate);
|
|
213
|
+
})
|
|
214
|
+
.map((item) => `${item.key}:changed`),
|
|
215
|
+
].sort();
|
|
216
|
+
}
|
|
@@ -8,6 +8,8 @@ export interface DeliveryStatusV2 {
|
|
|
8
8
|
effective_risk: "standard" | "strict";
|
|
9
9
|
workspace_snapshot_sha256: string;
|
|
10
10
|
acceptance_authority: "live_final_gate_required";
|
|
11
|
+
acceptance_scope: "declared_machine_authority";
|
|
12
|
+
native_goal_effect: "none";
|
|
11
13
|
final_result: AuditGateStatusV2;
|
|
12
14
|
final_workflow_status: FinalReceiptV2["workflow_status"] | null;
|
|
13
15
|
external_confirmations: ExternalConfirmationV2[];
|
|
@@ -29,11 +31,16 @@ export interface StopCheckDeliveryResultV2 {
|
|
|
29
31
|
reason: string;
|
|
30
32
|
workflow_status?: FinalReceiptV2["workflow_status"];
|
|
31
33
|
external_confirmations?: ExternalConfirmationV2[];
|
|
34
|
+
acceptance_scope?: "declared_machine_authority";
|
|
35
|
+
native_goal_effect?: "none";
|
|
32
36
|
message?: string;
|
|
33
37
|
}
|
|
34
38
|
export interface CloseDeliveryResultV2 {
|
|
35
39
|
status: "closed";
|
|
36
40
|
workflow_status: Extract<FinalReceiptV2["workflow_status"], "machine_accepted" | "machine_accepted_external_pending">;
|
|
37
41
|
external_confirmations: ExternalConfirmationV2[];
|
|
42
|
+
acceptance_scope: "declared_machine_authority";
|
|
43
|
+
closed_scope: "machine_authority";
|
|
44
|
+
native_goal_effect: "none";
|
|
38
45
|
}
|
|
39
46
|
export declare function closeDeliveryTask(workdir: string): Promise<CloseDeliveryResultV2>;
|
|
@@ -48,6 +48,8 @@ async function readDeliveryStatusForAuthority(active) {
|
|
|
48
48
|
effective_risk: compiled.effective_risk,
|
|
49
49
|
workspace_snapshot_sha256: current.snapshot_sha256,
|
|
50
50
|
acceptance_authority: "live_final_gate_required",
|
|
51
|
+
acceptance_scope: "declared_machine_authority",
|
|
52
|
+
native_goal_effect: "none",
|
|
51
53
|
final_result: projection.finalResult,
|
|
52
54
|
final_workflow_status: projection.finalWorkflowStatus,
|
|
53
55
|
external_confirmations: compiled.global.acceptance.external_confirmations,
|
|
@@ -94,6 +96,8 @@ export async function resumeDeliveryTask(workdir) {
|
|
|
94
96
|
context_refs: compiled.task.context_refs,
|
|
95
97
|
git,
|
|
96
98
|
acceptance_authority: "live_final_gate_required",
|
|
99
|
+
acceptance_scope: "declared_machine_authority",
|
|
100
|
+
native_goal_effect: "none",
|
|
97
101
|
last_gate: status.final_result,
|
|
98
102
|
final_workflow_status: status.final_workflow_status,
|
|
99
103
|
external_confirmations: status.external_confirmations,
|
|
@@ -221,11 +225,9 @@ export async function stopCheckDeliveryTask(workdirInput, messageText = "") {
|
|
|
221
225
|
reason: result.workflow_status,
|
|
222
226
|
workflow_status: result.workflow_status,
|
|
223
227
|
external_confirmations: result.external_confirmations,
|
|
224
|
-
|
|
225
|
-
|
|
226
|
-
|
|
227
|
-
}
|
|
228
|
-
: {}),
|
|
228
|
+
acceptance_scope: "declared_machine_authority",
|
|
229
|
+
native_goal_effect: "none",
|
|
230
|
+
message: acceptedScopeMessage(result.workflow_status, result.external_confirmations),
|
|
229
231
|
};
|
|
230
232
|
}
|
|
231
233
|
return {
|
|
@@ -272,13 +274,19 @@ export async function closeDeliveryTask(workdir) {
|
|
|
272
274
|
status: "closed",
|
|
273
275
|
workflow_status: result.workflow_status,
|
|
274
276
|
external_confirmations: result.external_confirmations,
|
|
277
|
+
acceptance_scope: "declared_machine_authority",
|
|
278
|
+
closed_scope: "machine_authority",
|
|
279
|
+
native_goal_effect: "none",
|
|
275
280
|
};
|
|
276
281
|
}
|
|
277
|
-
function
|
|
282
|
+
function acceptedScopeMessage(workflowStatus, confirmations) {
|
|
283
|
+
const scope = "Declared machine Authority accepted and cleared. This result has no direct effect on the platform-native Goal; before completing it, confirm current Goal/user meaning is fully represented by accepted Source and no revision, blocker, or omitted requirement remains.";
|
|
284
|
+
if (workflowStatus === "machine_accepted")
|
|
285
|
+
return scope;
|
|
278
286
|
const pending = confirmations
|
|
279
287
|
.map((confirmation) => `${confirmation.key} (${confirmation.owner})`)
|
|
280
288
|
.join(", ");
|
|
281
|
-
return
|
|
289
|
+
return `${scope} Complete external delivery remains pending: ${pending}. Do not report complete external delivery.`;
|
|
282
290
|
}
|
|
283
291
|
function nextAction(status) {
|
|
284
292
|
if (status.pending_authority_revision)
|