project-tiny-context-harness 0.6.2 → 0.7.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -21
- package/README.md +342 -333
- package/assets/README.md +497 -487
- package/assets/README.zh-CN.md +266 -256
- package/assets/agents/.gitkeep +1 -1
- package/assets/agents/AGENTS_CORE.md +56 -56
- package/assets/context_templates/architecture.md +33 -33
- package/assets/context_templates/area.md +39 -39
- package/assets/context_templates/context.toml +30 -30
- package/assets/context_templates/deployment.md +35 -35
- package/assets/context_templates/global.md +55 -55
- package/assets/context_templates/product-surface-contract.md +63 -63
- package/assets/context_templates/verification.md +32 -32
- package/assets/github/.gitkeep +1 -1
- package/assets/github/harness.yml +41 -41
- package/assets/make/.gitkeep +1 -1
- package/assets/make/ty-context.mk +48 -48
- package/assets/skills/context_development_engineer/SKILL.md +90 -90
- package/assets/skills/context_full_project_export/SKILL.md +70 -70
- package/assets/skills/context_harness_upgrade/SKILL.md +60 -60
- package/assets/skills/context_product_plan/SKILL.md +77 -77
- package/assets/skills/context_surface_contract/SKILL.md +171 -171
- package/assets/skills/context_uiux_design/SKILL.md +99 -99
- package/assets/skills/long-task-workflow/SKILL.md +75 -71
- package/assets/skills/long-task-workflow/agents/openai.yaml +4 -4
- package/assets/skills/long-task-workflow/references/authority-lifecycle.md +53 -41
- package/assets/skills/long-task-workflow/references/contract-authoring.md +57 -43
- package/assets/skills/long-task-workflow/references/evidence-design.md +40 -32
- package/assets/skills/normal-long-task/SKILL.md +12 -12
- package/assets/skills/source-plan-authoring/SKILL.md +295 -295
- package/dist/commands/check-modularity.js +10 -10
- package/dist/commands/long-task-authoring.js +65 -65
- package/dist/commands/long-task-command-args.d.ts +3 -0
- package/dist/commands/long-task-command-args.js +20 -0
- package/dist/commands/long-task-revision.d.ts +1 -0
- package/dist/commands/long-task-revision.js +137 -0
- package/dist/commands/long-task.js +20 -120
- package/dist/lib/long-task-authority-revision-details.d.ts +2 -0
- package/dist/lib/long-task-authority-revision-details.js +10 -0
- package/dist/lib/long-task-authority-revision-diagnosis.d.ts +24 -0
- package/dist/lib/long-task-authority-revision-diagnosis.js +141 -0
- package/dist/lib/long-task-authority-revision-enforcement.d.ts +3 -1
- package/dist/lib/long-task-authority-revision-enforcement.js +17 -9
- package/dist/lib/long-task-authority-revision-summary.d.ts +11 -0
- package/dist/lib/long-task-authority-revision-summary.js +99 -0
- package/dist/lib/long-task-authority-revision-types.d.ts +44 -1
- package/dist/lib/long-task-authority-revision.js +9 -1
- package/dist/lib/long-task-delivery-compiler.d.ts +3 -0
- package/dist/lib/long-task-delivery-compiler.js +5 -2
- package/dist/lib/long-task-state.d.ts +3 -15
- package/dist/lib/long-task-status-v2.d.ts +2 -0
- package/dist/lib/long-task-status-v2.js +12 -2
- package/dist/lib/long-task-verifier-v2.d.ts +4 -0
- package/dist/lib/long-task-verifier-v2.js +1 -1
- package/migrations/README.md +8 -8
- package/package.json +5 -1
- package/source-mappings.yaml +25 -25
|
@@ -1,77 +1,81 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: long-task-workflow
|
|
3
|
-
description: Author, preflight, execute, resume, verify, or close one complete Single-Goal Delivery Contract in the current native Goal and workspace. Use only when explicitly invoked or a valid common-dir active authority binding exists.
|
|
4
|
-
---
|
|
5
|
-
|
|
6
|
-
# Single-Goal Long-Task Workflow
|
|
7
|
-
|
|
8
|
-
## Boundaries
|
|
9
|
-
|
|
10
|
-
Use one current native Goal, one repository, one selected workspace, one complete Contract and one Final Gate. Never create a scheduler, model worker, agent runtime, App Server, branch, worktree, merge, push, PR, deployment, Campaign/SFC/Packet/Wave chain, matrix, verdict or second Contract plan. Never activate from task size alone.
|
|
11
|
-
|
|
1
|
+
---
|
|
2
|
+
name: long-task-workflow
|
|
3
|
+
description: Author, preflight, execute, resume, verify, or close one complete Single-Goal Delivery Contract in the current native Goal and workspace. Use only when explicitly invoked or a valid common-dir active authority binding exists.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Single-Goal Long-Task Workflow
|
|
7
|
+
|
|
8
|
+
## Boundaries
|
|
9
|
+
|
|
10
|
+
Use one current native Goal, one repository, one selected workspace, one complete Contract and one Final Gate. Never create a scheduler, model worker, agent runtime, App Server, branch, worktree, merge, push, PR, deployment, Campaign/SFC/Packet/Wave chain, matrix, verdict or second Contract plan. Never activate from task size alone.
|
|
11
|
+
|
|
12
12
|
The host and user own model selection. The workflow has exactly one user-choice checkpoint after the first Authority Lock and before implementation; Harness neither switches the model nor persists model-routing/checkpoint state. No checkpoint file, acknowledgement state, model route, model-tier scheduler or automatic model switch is created. Outside that boundary, do not pause a healthy Goal solely to change or downgrade the model. Do not create a separate approval checkpoint for a defensible recommended plan choice. A targeted pre-Authority clarification is still required when a missing user preference could materially change research or selection; genuine Source conflicts or choices the user explicitly reserves may likewise require a decision before Authority Lock. Capability-related drift is handled by targeted repair plus the Final Gate. Never proactively spawn, assign or coordinate parallel subagents. Platform-native internal delegation, if it occurs, is opaque and non-authoritative and must converge into the unified current workspace snapshot before verification can count.
|
|
13
|
-
|
|
14
|
-
`long-task-delivery-v2` is the only active Contract schema. `delivery-contract.yaml` is the root authoring file. New authoring uses inline Outcomes; existing `outcome_files` are physical compatibility only. `delivery-set` is retired and non-executing.
|
|
15
|
-
|
|
16
|
-
## Controlling Objective
|
|
17
|
-
|
|
18
|
-
Prevent false completion inside declared authority. Implementation may drift, fail or require rework, but every declared non-Result requirement and AC must remain traceable and every unsatisfied, unverifiable, insufficiently evidenced or stale item must block completion. Findings should localize repair through Source Item, Outcome, Claim, Assertion, Check, Proof Surface, Binding and owner boundary.
|
|
19
|
-
|
|
20
|
-
Only fresh evidence from the complete current final snapshot may create machine acceptance. Otherwise report the task as unfinished or qualified. `machine_accepted_external_pending` means machine-verifiable authority passed while named external confirmation remains; it is not full delivery completion. Never substitute prose, progress, historical tests, Receipts, one exit code or Agent judgment for the Final Gate.
|
|
21
|
-
|
|
22
|
-
Prefer the lowest practical Authoring, Runtime, State, Recovery and verification cost that preserves the same false-completion interception. Add no mechanism whose distinct protection does not materially exceed its total cost.
|
|
23
|
-
|
|
24
|
-
## Progressive Reference Loading
|
|
25
|
-
|
|
26
|
-
Read only the reference needed for the current phase; these files are guidance, not new artifacts or authority:
|
|
27
|
-
|
|
28
|
-
- Before creating or structurally revising Source markers, Outcomes, requirements, controls, obligations, architecture boundaries, paths, Bindings, Assertions or risk, read [`references/contract-authoring.md`](references/contract-authoring.md).
|
|
29
|
-
- Before creating or repairing Checks, runners, Observations, proof surfaces, Playwright/structured evidence, Counterfactuals, Population or environment probes, read [`references/evidence-design.md`](references/evidence-design.md).
|
|
30
|
-
- Before Preflight, Compile, protected revision, resume, targeted verify, Final Gate, Stop, close or abandon, read [`references/authority-lifecycle.md`](references/authority-lifecycle.md).
|
|
31
|
-
|
|
32
|
-
Do not copy reference detail into another plan or state file. The same `delivery-contract.yaml`, active authority and current workspace remain the only lifecycle surfaces.
|
|
33
|
-
|
|
34
|
-
## Contract Draft And Outcome Decomposition
|
|
35
|
-
|
|
36
|
-
Before the first successful formal Compile, continuously revise the same non-authoritative `delivery-contract.yaml` as the Contract Draft. It need not be completed in one response; keep reading Source, repository and relevant Context and feed Preflight findings back into that same Draft. Draft authoring, Preflight, Compile, rolling execution, targeted verification and Final Gate are one `long-task-workflow` lifecycle. Do not create a standalone Contract Draft Skill, Draft Receipt, Authoring State, draft schema/CLI/runtime state or second plan.
|
|
37
|
-
|
|
38
|
-
A Draft Outcome is an Outcome in that pre-Authority-Lock Draft, not a new schema field or runtime entity. Decompose only independently observable, decidable and target-verifiable results whose dependencies and owner boundary can be stated. Use those boundaries to keep a dependency-ready working set, target verification, localize failures, resume findings/next actions and stale local results precisely.
|
|
39
|
-
|
|
40
|
-
`depends_on` means acceptance readiness. The current Goal may form a temporary Rolling Frontier from ready Outcomes and findings, but must not persist a scheduler, Worker queue, mandatory implementation DAG, model route or process tree. Never split for response/YAML/file length, implementation layer, module/file count, Agent capacity, Worker assignment or desired parallelism.
|
|
41
|
-
|
|
42
|
-
> Outcome decomposes execution and diagnosis, not completion authority.
|
|
43
|
-
|
|
44
|
-
## Entry And Authoring Loop
|
|
45
|
-
|
|
46
|
-
1. Read the user request or external proposal plus minimum controlling Context and decide `Context Delta: none|required`.
|
|
47
|
-
2. If a valid active binding exists, run `ty-context long-task resume <workdir>` and read the lifecycle reference.
|
|
48
|
-
3. Otherwise author one complete Delivery Contract for the whole selected delivery. Do not create a second Contract plan, matrix or top-level Contract split.
|
|
49
|
-
4. Preserve at least one real `source_path`. Wrap every material Source item in its original Markdown with non-rendering `ty-source-item:start/end` markers without rewriting the text; marked Source Item keys and `source_claim` keys are exactly equal.
|
|
50
|
-
5. An ordinary prose plan or optional Source Plan remains valid Source after marker-only enumeration and does not need to match the recommended Source Plan structure. Preserve stable semantic keys and Markdown anchors where practical.
|
|
13
|
+
|
|
14
|
+
`long-task-delivery-v2` is the only active Contract schema. `delivery-contract.yaml` is the root authoring file. New authoring uses inline Outcomes; existing `outcome_files` are physical compatibility only. `delivery-set` is retired and non-executing.
|
|
15
|
+
|
|
16
|
+
## Controlling Objective
|
|
17
|
+
|
|
18
|
+
Prevent false completion inside declared authority. Implementation may drift, fail or require rework, but every declared non-Result requirement and AC must remain traceable and every unsatisfied, unverifiable, insufficiently evidenced or stale item must block completion. Findings should localize repair through Source Item, Outcome, Claim, Assertion, Check, Proof Surface, Binding and owner boundary.
|
|
19
|
+
|
|
20
|
+
Only fresh evidence from the complete current final snapshot may create machine acceptance. Otherwise report the task as unfinished or qualified. `machine_accepted_external_pending` means machine-verifiable authority passed while named external confirmation remains; it is not full delivery completion. Never substitute prose, progress, historical tests, Receipts, one exit code or Agent judgment for the Final Gate.
|
|
21
|
+
|
|
22
|
+
Prefer the lowest practical Authoring, Runtime, State, Recovery and verification cost that preserves the same false-completion interception. Add no mechanism whose distinct protection does not materially exceed its total cost.
|
|
23
|
+
|
|
24
|
+
## Progressive Reference Loading
|
|
25
|
+
|
|
26
|
+
Read only the reference needed for the current phase; these files are guidance, not new artifacts or authority:
|
|
27
|
+
|
|
28
|
+
- Before creating or structurally revising Source markers, Outcomes, requirements, controls, obligations, architecture boundaries, paths, Bindings, Assertions or risk, read [`references/contract-authoring.md`](references/contract-authoring.md).
|
|
29
|
+
- Before creating or repairing Checks, runners, Observations, proof surfaces, Playwright/structured evidence, Counterfactuals, Population or environment probes, read [`references/evidence-design.md`](references/evidence-design.md).
|
|
30
|
+
- Before Preflight, Compile, protected revision, resume, targeted verify, Final Gate, Stop, close or abandon, read [`references/authority-lifecycle.md`](references/authority-lifecycle.md).
|
|
31
|
+
|
|
32
|
+
Do not copy reference detail into another plan or state file. The same `delivery-contract.yaml`, active authority and current workspace remain the only lifecycle surfaces.
|
|
33
|
+
|
|
34
|
+
## Contract Draft And Outcome Decomposition
|
|
35
|
+
|
|
36
|
+
Before the first successful formal Compile, continuously revise the same non-authoritative `delivery-contract.yaml` as the Contract Draft. It need not be completed in one response; keep reading Source, repository and relevant Context and feed Preflight findings back into that same Draft. Draft authoring, Preflight, Compile, rolling execution, targeted verification and Final Gate are one `long-task-workflow` lifecycle. Do not create a standalone Contract Draft Skill, Draft Receipt, Authoring State, draft schema/CLI/runtime state or second plan.
|
|
37
|
+
|
|
38
|
+
A Draft Outcome is an Outcome in that pre-Authority-Lock Draft, not a new schema field or runtime entity. Decompose only independently observable, decidable and target-verifiable results whose dependencies and owner boundary can be stated. Use those boundaries to keep a dependency-ready working set, target verification, localize failures, resume findings/next actions and stale local results precisely.
|
|
39
|
+
|
|
40
|
+
`depends_on` means acceptance readiness. The current Goal may form a temporary Rolling Frontier from ready Outcomes and findings, but must not persist a scheduler, Worker queue, mandatory implementation DAG, model route or process tree. Never split for response/YAML/file length, implementation layer, module/file count, Agent capacity, Worker assignment or desired parallelism.
|
|
41
|
+
|
|
42
|
+
> Outcome decomposes execution and diagnosis, not completion authority.
|
|
43
|
+
|
|
44
|
+
## Entry And Authoring Loop
|
|
45
|
+
|
|
46
|
+
1. Read the user request or external proposal plus minimum controlling Context and decide `Context Delta: none|required`.
|
|
47
|
+
2. If a valid active binding exists, run `ty-context long-task resume <workdir>` and read the lifecycle reference.
|
|
48
|
+
3. Otherwise author one complete Delivery Contract for the whole selected delivery. Do not create a second Contract plan, matrix or top-level Contract split.
|
|
49
|
+
4. Preserve at least one real `source_path`. Wrap every material Source item in its original Markdown with non-rendering `ty-source-item:start/end` markers without rewriting the text; marked Source Item keys and `source_claim` keys are exactly equal.
|
|
50
|
+
5. An ordinary prose plan or optional Source Plan remains valid Source after marker-only enumeration and does not need to match the recommended Source Plan structure. Preserve stable semantic keys and Markdown anchors where practical.
|
|
51
51
|
6. Continue reading repository, Source and Context and revise the same Draft. A request to synthesize, refine, complete, implement or use judgment delegates plan-level authoring, but it does not invent the user's tradeoff priorities. Before comparative research or a material product, technical, architecture or provider selection, identify the criteria that could change the research scope, candidate set or recommendation. Infer them only from the user's words, Source, Context or controlling constraints. If quality versus cost, speed, reliability, privacy, lock-in, operational burden or another material priority is unknown or ambiguous, stop before that research or selection and ask one concise targeted clarification. Do not impose a questionnaire, re-ask known preferences or interrupt minor reversible choices whose recommendation would not change.
|
|
52
52
|
7. Once the material preference envelope is clear, decide what research is needed. Use current authoritative or primary evidence for external capability, pricing, quota, license, compatibility, region, security posture or support claims. When one recommendation is then defensible, record it in real Source with the authoring instruction, preference/evidence basis and exact added meaning instead of pausing for approval. Append the delegated item without rewriting the user's original text when ordinary prose is the Source. Return only when authoritative requirements conflict, the user explicitly reserves the choice, a material preference remains unknown, critical semantics have no defensible recommendation or no falsifiable acceptance standard can be formed.
|
|
53
53
|
8. Contract expansion remains limited to meaning-preserving structural decomposition, evidence-backed repository binding and choices first recorded as delegated real Source. Never place a new product rule, default, threshold, recovery behavior, permission or platform/data scope only in Contract YAML. Default plan delegation authorizes meaning, not action: payment, contracting, production deployment or publication, destructive production mutation, real permission grants, sensitive-data transmission and required legal/security/human approval remain named external confirmations. Any conflicting, user-reserved, missing-preference or unsupported semantic remains `decision_required`.
|
|
54
54
|
9. Run read-only `ty-context long-task preflight <workdir>`, repair every error and `decision_required` finding in the same Draft, then formally Compile only when ready.
|
|
55
55
|
10. When the first Compile returns `execution_model_checkpoint.required: true`, stop before implementation and ask the user to choose `continue_current_model` or switch models and then resume the active Long-Task. A task-specific choice already stated explicitly satisfies the checkpoint. Later revisions return `required: false` and do not repeat it.
|
|
56
|
-
|
|
57
|
-
Architecture quality uses the existing authority model, not a new gate: when Source or controlling Context declares an architecture invariant, encode it as a Source-backed technical obligation/global constraint/forbidden shortcut plus owner/path/Binding boundaries and a project-owned executable Check. Functional acceptance cannot substitute when the architecture claim can fail independently. An unverifiable design preference remains task-local, durable Context or `decision_required`; it must not be promoted into false proof.
|
|
58
|
-
|
|
59
|
-
## Rolling Execution
|
|
60
|
-
|
|
61
|
-
After Authority Lock and the one-time execution-model checkpoint are satisfied, implement dependency-ready Outcomes in the current workspace. Small implementation plans and repair hypotheses are internal execution state and cannot silently change Product, Technical or Acceptance authority.
|
|
62
|
-
|
|
63
|
-
Re-evaluate `Context Delta` whenever implementation or repair discovers a durable fact. Controlling Context changes use protected revision; graph-derived, non-explicit `implementation-index` and `archive` are Supporting Context in referenced mode and may auto-revise when only navigation/background changed. Full snapshot mode treats every selected Context file as controlling.
|
|
64
|
-
|
|
65
|
-
Use targeted `verify --outcome/--check` only to drive repair. Progress is repair evidence only and never acceptance authority. Keep precise findings attached to the owning Source item, Claim, Assertion, Check, Binding and owner path. Do not add another model-switch pause or coordinate parallel subagents.
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
56
|
+
|
|
57
|
+
Architecture quality uses the existing authority model, not a new gate: when Source or controlling Context declares an architecture invariant, encode it as a Source-backed technical obligation/global constraint/forbidden shortcut plus owner/path/Binding boundaries and a project-owned executable Check. Functional acceptance cannot substitute when the architecture claim can fail independently. An unverifiable design preference remains task-local, durable Context or `decision_required`; it must not be promoted into false proof.
|
|
58
|
+
|
|
59
|
+
## Rolling Execution
|
|
60
|
+
|
|
61
|
+
After Authority Lock and the one-time execution-model checkpoint are satisfied, implement dependency-ready Outcomes in the current workspace. Small implementation plans and repair hypotheses are internal execution state and cannot silently change Product, Technical or Acceptance authority.
|
|
62
|
+
|
|
63
|
+
Re-evaluate `Context Delta` whenever implementation or repair discovers a durable fact. Controlling Context changes use protected revision; graph-derived, non-explicit `implementation-index` and `archive` are Supporting Context in referenced mode and may auto-revise when only navigation/background changed. Full snapshot mode treats every selected Context file as controlling.
|
|
64
|
+
|
|
65
|
+
Use targeted `verify --outcome/--check` only to drive repair. Progress is repair evidence only and never acceptance authority. Keep precise findings attached to the owning Source item, Claim, Assertion, Check, Binding and owner path. Do not add another model-switch pause or coordinate parallel subagents.
|
|
66
|
+
|
|
67
|
+
When the Contract declares a target-runtime Check because a proxy can pass while the target fails independently, run it at the earliest owning Outcome's first runnable boundary. After accumulated changes to its declared `input_paths` or Binding carriers make the result stale, rerun it before dependent work grows. Coalesce related edits and use the cheapest reliable target Check; do not mandate a full environment rebuild per Outcome or per edit. This is rolling feedback through existing targeted verify, not acceptance, a trigger queue, platform taxonomy or new state.
|
|
68
|
+
|
|
69
|
+
When implementation discovers missing Contract paths, first classify the revision. Proven monotonic evidence strengthening may use ordinary `compile --revise` directly. If every protected reason is only owner/expected-change/allowed-support expansion, continue editing the same `delivery-contract.yaml` and use `ty-context long-task diagnose-revision <workdir> [--outcome <key>] [--check <key>]` to exercise only existing active Check identities with unchanged runner/verifier authority; safe monotonic strengthening may coexist. Candidate diagnostics are transient: they authorize no acceptance and write no pending/approval state, Active Authority, cache, Progress or Receipt. Semantic changes, proof weakening, runner or verifier-content changes, and risk-increase candidates are preview-only and must not run; risk downgrade is rejected. When the candidate is complete, run ordinary `compile --revise` once, present its exact concise decision summary to the user, and never approve it yourself. Keep the previous Authority active until exact approval and atomic adoption; after adoption, discard historical/candidate evidence and require the complete Final Gate.
|
|
70
|
+
|
|
71
|
+
## Live Final Authority
|
|
72
|
+
|
|
73
|
+
Complete Context, implementation and project tests, create a clean candidate commit, then run `ty-context long-task final-gate <workdir>`.
|
|
74
|
+
|
|
75
|
+
Final Gate recompiles Source authority, validates active task/revision/compiled/worktree identity, creates one Git-tree snapshot, reruns every required Global and Outcome Check and rechecks active identity before acceptance. A target-runtime Check must exercise its target in that current Gate execution; rerunning a reader for a historical or tracked status report is not live target proof. Final Gate, Stop and close never trust historical Progress, Receipt or compiled cache.
|
|
76
|
+
|
|
77
|
+
Machine acceptance covers only declared machine authority. Preserve every pending external confirmation through `final-gate`, `status`, `resume`, `stop-check`, the package-owned Stop Hook and `close`; `closed` means only machine Authority cleanup. Do not invent external-confirmation tracking state.
|
|
78
|
+
|
|
79
|
+
## Handoff
|
|
80
|
+
|
|
81
|
+
Report implementation, effective risk, Claim Coverage, Live Gate result, every pending external confirmation, Context status and blockers. Use verifier terms exactly: `progress_passing` means targeted repair evidence, `progress_stale` is not a current pass, `final_workflow_status: null` means unfinished, and `machine_accepted_external_pending` must retain its named confirmations. Never shorten implementation or targeted progress to “Outcome complete” or invent `implementation_complete`, `platform_smoke_verified` or another persistent status. State the threat-model limits: undeclared requirements cannot be discovered, installed verifier/Git metadata are trusted, model selection belongs to the host/user, and internal platform delegation is not observed.
|
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
interface:
|
|
2
|
-
display_name: "Long-Task Workflow"
|
|
3
|
-
short_description: "Run one Delivery Contract in the current native Goal"
|
|
4
|
-
default_prompt: "Use /long-task-workflow to prepare, execute, resume, verify, or close one Canonical Delivery Contract in the current workspace."
|
|
1
|
+
interface:
|
|
2
|
+
display_name: "Long-Task Workflow"
|
|
3
|
+
short_description: "Run one Delivery Contract in the current native Goal"
|
|
4
|
+
default_prompt: "Use /long-task-workflow to prepare, execute, resume, verify, or close one Canonical Delivery Contract in the current workspace."
|
|
@@ -1,41 +1,53 @@
|
|
|
1
|
-
# Authority Lifecycle Reference
|
|
2
|
-
|
|
3
|
-
Read this before Preflight, Compile, revision, resume, targeted verify, Final Gate, Stop, close or abandon.
|
|
4
|
-
|
|
5
|
-
## Preflight And Compile
|
|
6
|
-
|
|
7
|
-
Run `ty-context long-task preflight <workdir>` before first formal Compile. Resolve every `error` and `decision_required` diagnostic and review warnings. Preflight is read-only: it creates no Active Authority, initial base, marker, cache, Progress, Receipt or pending revision, runs no project Check and persists no success record.
|
|
8
|
-
|
|
9
|
-
Preflight and Compile call the same activation-safety validator. Skipping Preflight bypasses no Source continuity, criterion, Claim/all-of-surface, adapter/Observation, risk, owner/path/Binding, runner/input, Counterfactual or sensitivity rule.
|
|
10
|
-
|
|
11
|
-
Preflight keeps every independently discovered diagnostic. When a structural duplicate makes the same Claim ambiguous or repeated, only that pair receives stable `diagnostic_id`, `repair_group`, `repair_priority` and `blocked_by` metadata so the structural blocker is repaired first. Independent findings keep their compact existing shape; no finding is hidden, reclassified or treated as resolved, and no repair state or authority is created.
|
|
12
|
-
|
|
13
|
-
The first successful `ty-context long-task compile <workdir>` is Authority Lock and freezes the immutable initial base and complete compiled authority snapshot in Git common-dir, bound to the worktree marker by task id, revision and compiled identity.
|
|
14
|
-
|
|
15
|
-
Its JSON result includes `execution_model_checkpoint.required: true`. Before product implementation, stop once and ask the user to continue with the current model or switch models and then resume the active Long-Task. A task-specific model choice already stated explicitly satisfies the checkpoint. Later Compile revisions return `required: false`; no checkpoint file, acknowledgement state, model route or automatic model switch is created.
|
|
16
|
-
|
|
17
|
-
## Protected Revision
|
|
18
|
-
|
|
19
|
-
After Authority Lock,
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
`
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
##
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
1
|
+
# Authority Lifecycle Reference
|
|
2
|
+
|
|
3
|
+
Read this before Preflight, Compile, revision, resume, targeted verify, Final Gate, Stop, close or abandon.
|
|
4
|
+
|
|
5
|
+
## Preflight And Compile
|
|
6
|
+
|
|
7
|
+
Run `ty-context long-task preflight <workdir>` before first formal Compile. Resolve every `error` and `decision_required` diagnostic and review warnings. Preflight is read-only: it creates no Active Authority, initial base, marker, cache, Progress, Receipt or pending revision, runs no project Check and persists no success record.
|
|
8
|
+
|
|
9
|
+
Preflight and Compile call the same activation-safety validator. Skipping Preflight bypasses no Source continuity, criterion, Claim/all-of-surface, adapter/Observation, risk, owner/path/Binding, runner/input, Counterfactual or sensitivity rule.
|
|
10
|
+
|
|
11
|
+
Preflight keeps every independently discovered diagnostic. When a structural duplicate makes the same Claim ambiguous or repeated, only that pair receives stable `diagnostic_id`, `repair_group`, `repair_priority` and `blocked_by` metadata so the structural blocker is repaired first. Independent findings keep their compact existing shape; no finding is hidden, reclassified or treated as resolved, and no repair state or authority is created.
|
|
12
|
+
|
|
13
|
+
The first successful `ty-context long-task compile <workdir>` is Authority Lock and freezes the immutable initial base and complete compiled authority snapshot in Git common-dir, bound to the worktree marker by task id, revision and compiled identity.
|
|
14
|
+
|
|
15
|
+
Its JSON result includes `execution_model_checkpoint.required: true`. Before product implementation, stop once and ask the user to continue with the current model or switch models and then resume the active Long-Task. A task-specific model choice already stated explicitly satisfies the checkpoint. Later Compile revisions return `required: false`; no checkpoint file, acknowledgement state, model route or automatic model switch is created.
|
|
16
|
+
|
|
17
|
+
## Protected Revision
|
|
18
|
+
|
|
19
|
+
After Authority Lock, every revision compares against active authority and follows one of three paths:
|
|
20
|
+
|
|
21
|
+
1. proven monotonic evidence strengthening, pure verifier relocation, proven tightening and supporting-only Context revision may auto-revise;
|
|
22
|
+
2. a candidate whose only protected reasons are owner, expected-change or allowed-support expansion remains inactive but may be exercised with `diagnose-revision` through existing active Check identities whose runner/verifier authority is unchanged; safe monotonic strengthening may coexist; or
|
|
23
|
+
3. every semantic change, proof weakening, runner or verifier-content change, risk change or other protected reason requires the exact revision identity and is never candidate-executed.
|
|
24
|
+
|
|
25
|
+
`diagnose-revision` recompiles the same `delivery-contract.yaml` in memory, creates only a disposable workspace snapshot when class 2 is proven, and returns transient repair results with `acceptance_authorized: false`. It writes no pending/approval state, authority/marker, cache, Progress or Receipt. Repeated edits therefore accumulate only in the one existing Contract authoring file, not a pending Draft authority or candidate state plane.
|
|
26
|
+
|
|
27
|
+
Ordinary `compile --revise` is the only operation that may create the one pending decision. It binds a deterministic concise change summary into the revision identity. `status` and `resume` expose that same decision so the host can deduplicate the user prompt without a Harness-owned waiting state. The executing Agent never approves its own pending revision; earlier blanket authorization cannot approve a later exact identity. If the candidate changes, the identity changes and old approval is rejected. The previous Authority remains active until approved compare-and-swap adoption, which invalidates derived evidence and leaves the complete source-recompiled Final Gate mandatory.
|
|
28
|
+
|
|
29
|
+
Every path-bearing field uses canonical grammar. Internal `.`/`..`, control characters, empty segments, absolute/drive/UNC paths and unsupported glob syntax fail closed.
|
|
30
|
+
|
|
31
|
+
Controlling Context includes core Context, explicit `context_refs`, verification/deployment Context and every file in full snapshot mode. In referenced mode only graph-derived, non-explicit `implementation-index` and `archive` files are Supporting Context; a supporting-only `compile --revise` may preserve otherwise-fresh targeted Progress.
|
|
32
|
+
|
|
33
|
+
`context.toml` retrieval guidance (`triggers`, `read_when`, `read_policy`, default selection and unselected nodes) is excluded from the selected delivery-authority projection. Selected area ownership, role/dependency structure and selected Context contents remain protected revision material. Retrieval-only edits may preserve scoped Progress, but a changed final Git tree still invalidates historical final acceptance and must pass the Live Final Gate again.
|
|
34
|
+
|
|
35
|
+
## Targeted Verification And Recovery
|
|
36
|
+
|
|
37
|
+
`verify --outcome/--check` runs scoped current-snapshot checks for repair and rechecks active identity before writing Progress. A Counterfactual finding is projected into the owning Main Check, changes an otherwise passed Check to `invalid_evidence`, clears Claim proofs and remains recoverable through `status`/`resume`. Global Checks use the same record without a new Global Outcome state.
|
|
38
|
+
|
|
39
|
+
For a declared target-runtime Check, run targeted verify once when its earliest owning Outcome reaches the first runnable boundary. After accumulated changes to its declared `input_paths` or Binding carriers make Progress stale, rerun before dependent work grows. Coalesce related edits and use the cheapest reliable Check; do not create a per-edit/per-Outcome rebuild rule, trigger queue or platform state. These runs remain `acceptance_authorized: false`.
|
|
40
|
+
|
|
41
|
+
Progress freshness binds Outcome authority, runner, verification inputs, Controlling Context and implementation inputs. Retry defaults to none; one retry is allowed only for explicit `transient_once`, idempotent, read-only/test-sandbox work.
|
|
42
|
+
|
|
43
|
+
Status, Progress, Receipts and workdir compiled output are audit/recovery projections only. Development-period authority state is `manual_required` and never migrated.
|
|
44
|
+
|
|
45
|
+
Report their exact meaning: `progress_passing` is current targeted repair evidence rather than “Outcome complete”; `progress_stale` is not a current pass; `final_workflow_status: null` means the Goal is unfinished. Do not invent `implementation_complete`, `platform_smoke_verified` or another persistent completion vocabulary.
|
|
46
|
+
|
|
47
|
+
## Final Gate And Terminal Paths
|
|
48
|
+
|
|
49
|
+
Before Final Gate, complete Context/code/tests and create a clean candidate commit. Final Gate captures active identity, recompiles Source authority, reads complete current Context, validates common-dir record/marker, creates a Git-tree snapshot, reruns all Checks and sensitivity controls and rechecks identity before acceptance. A target-runtime Check must exercise its target again in that Final Gate execution; rereading historical status does not become live proof merely because the reader reran. A concurrent revision returns `active_authority_changed_during_final_gate`.
|
|
50
|
+
|
|
51
|
+
Commit, verifier migration, clear and abandon share one active-state lock. Stop/close clear only the identity actually accepted through CAS and preserve `machine_accepted_external_pending` plus every named external confirmation in output. A stale Receipt exposes no accepted workflow status.
|
|
52
|
+
|
|
53
|
+
For invalid, mismatched, unrecoverable or stale-lock continuity, use only `ty-context long-task abandon <workdir> --force-corrupt-state`; it preserves authored Contract, Source, Context and Git content.
|
|
@@ -1,49 +1,63 @@
|
|
|
1
|
-
# Contract Authoring Reference
|
|
2
|
-
|
|
3
|
-
Read this only while authoring or structurally revising the one `delivery-contract.yaml` Draft.
|
|
4
|
-
|
|
5
|
-
## Source And Semantic Boundary
|
|
6
|
-
|
|
7
|
-
- Every declared Source file contains at least one Material Source Item. Mark items in original Markdown without rendering or changing meaning.
|
|
8
|
-
- Marker keys and `source_claim` keys are set-equal and globally unique. `statement` preserves marked text after line-ending normalization, surrounding blank-line removal and trailing-space cleanup.
|
|
9
|
-
- Typed dispositions keep Result, Requirement, Control, Technical Obligation, Non-completing Claim, Acceptance, Global Constraint/Non-goal, Forbidden Shortcut, Risk, External Confirmation and Decision distinct.
|
|
1
|
+
# Contract Authoring Reference
|
|
2
|
+
|
|
3
|
+
Read this only while authoring or structurally revising the one `delivery-contract.yaml` Draft.
|
|
4
|
+
|
|
5
|
+
## Source And Semantic Boundary
|
|
6
|
+
|
|
7
|
+
- Every declared Source file contains at least one Material Source Item. Mark items in original Markdown without rendering or changing meaning.
|
|
8
|
+
- Marker keys and `source_claim` keys are set-equal and globally unique. `statement` preserves marked text after line-ending normalization, surrounding blank-line removal and trailing-space cleanup.
|
|
9
|
+
- Typed dispositions keep Result, Requirement, Control, Technical Obligation, Non-completing Claim, Acceptance, Global Constraint/Non-goal, Forbidden Shortcut, Risk, External Confirmation and Decision distinct.
|
|
10
10
|
- Every non-decision Source item owns exactly one same-kind, text-identical canonical target; no target may collapse multiple Source items. `out_of_scope` is not a resolution.
|
|
11
11
|
- A Source AC maps criterion-identically to one named Assertion and proves at least one independently Source-backed non-Result Claim.
|
|
12
12
|
- Missing recommended Source Plan headings or keys never blocks authoring. Missing mandatory Material Source Item markers does.
|
|
13
13
|
- `delegated` in a Source Plan is provenance, not a Contract disposition or new Claim kind. An instruction to synthesize, refine, complete, implement or use judgment delegates plan-level authoring, but it does not invent material tradeoff preferences. Before comparative research or a material product, technical, architecture or provider selection, identify the criteria that could change the research scope, candidate set or recommendation. If such a preference is unknown or ambiguous, ask a concise targeted question before research or selection and keep the item `decision_required` until answered; do not impose a fixed questionnaire or re-ask preferences already supplied by the user, Source, Context or controlling constraints.
|
|
14
14
|
- Once the material preference envelope is clear, use current authoritative or primary evidence for external capability, price, quota, license, compatibility, region, security posture or support claims. When one defensible recommendation exists, record the authoring instruction, preference/evidence or conservative-default basis and exact added meaning in real Source, then preserve that keyed item as ordinary Source of its semantic kind. If ordinary prose is the Source, append the delegated item without rewriting the user's original text; never place the choice only in Contract YAML.
|
|
15
15
|
- A delegated plan choice is not action authorization. Payment, contracting, production deployment/publication, destructive production mutation, real permission grants, sensitive-data transmission and required legal/security/human approval remain named External Confirmations. Conflicting authority, an explicitly user-reserved choice, a missing material preference or the absence of a defensible recommendation remains `decision_required`; high impact or multiple options with known criteria alone does not.
|
|
16
|
-
|
|
17
|
-
## Outcome Boundary
|
|
18
|
-
|
|
19
|
-
Create an Outcome only when its result is independently observable, decidable, target-verifiable, dependency-expressible and localizable to its own Claims, Assertions, Checks and owner boundary. Requirement coupling, dependency-ready work, targeted verification, precise failure localization, semantic resume and stale-result invalidation are valid reasons to decompose. File count, implementation layer, context length, desired parallelism and Agent capacity are not.
|
|
20
|
-
|
|
21
|
-
For every Outcome declare:
|
|
22
|
-
|
|
23
|
-
- one complete observable result;
|
|
24
|
-
- atomic requirements and actually applicable controls/states;
|
|
25
|
-
- non-completing claims;
|
|
26
|
-
- owner Context/surfaces and expected/support/forbidden path envelopes;
|
|
27
|
-
- stable technical obligations, Bindings, forbidden shortcuts and recovery requirements;
|
|
28
|
-
- risk facts;
|
|
29
|
-
- executable Checks and named AC Assertions.
|
|
30
|
-
|
|
31
|
-
Global non-goals, constraints and forbidden shortcuts remain Global authority and use Global Checks/Assertions when machine proof is required.
|
|
32
|
-
|
|
16
|
+
|
|
17
|
+
## Outcome Boundary
|
|
18
|
+
|
|
19
|
+
Create an Outcome only when its result is independently observable, decidable, target-verifiable, dependency-expressible and localizable to its own Claims, Assertions, Checks and owner boundary. Requirement coupling, dependency-ready work, targeted verification, precise failure localization, semantic resume and stale-result invalidation are valid reasons to decompose. File count, implementation layer, context length, desired parallelism and Agent capacity are not.
|
|
20
|
+
|
|
21
|
+
For every Outcome declare:
|
|
22
|
+
|
|
23
|
+
- one complete observable result;
|
|
24
|
+
- atomic requirements and actually applicable controls/states;
|
|
25
|
+
- non-completing claims;
|
|
26
|
+
- owner Context/surfaces and expected/support/forbidden path envelopes;
|
|
27
|
+
- stable technical obligations, Bindings, forbidden shortcuts and recovery requirements;
|
|
28
|
+
- risk facts;
|
|
29
|
+
- executable Checks and named AC Assertions.
|
|
30
|
+
|
|
31
|
+
Global non-goals, constraints and forbidden shortcuts remain Global authority and use Global Checks/Assertions when machine proof is required.
|
|
32
|
+
|
|
33
33
|
## Architecture Closure
|
|
34
|
-
|
|
35
|
-
Architecture protection is risk-triggered and project-specific. Use it when the delivery declares module ownership, unique source of truth, dependency direction, API/schema/data boundary, state lifecycle, persistence/recovery, security boundary, compatibility/migration or a forbidden bypass.
|
|
36
|
-
|
|
37
|
-
Represent the invariant with existing Contract fields:
|
|
38
|
-
|
|
39
|
-
1. a Source-backed technical obligation, global constraint or forbidden shortcut;
|
|
40
|
-
2. owner Context and expected/support/forbidden paths;
|
|
41
|
-
3. a Binding to the implementation carrier when Counterfactual sensitivity is required;
|
|
42
|
-
4. a project-owned executable architecture check, such as the repository's lint, AST, dependency or contract test;
|
|
43
|
-
5. a separate Assertion when functional behavior could pass while the architecture invariant fails.
|
|
44
|
-
|
|
34
|
+
|
|
35
|
+
Architecture protection is risk-triggered and project-specific. Use it when the delivery declares module ownership, unique source of truth, dependency direction, API/schema/data boundary, state lifecycle, persistence/recovery, security boundary, compatibility/migration or a forbidden bypass.
|
|
36
|
+
|
|
37
|
+
Represent the invariant with existing Contract fields:
|
|
38
|
+
|
|
39
|
+
1. a Source-backed technical obligation, global constraint or forbidden shortcut;
|
|
40
|
+
2. owner Context and expected/support/forbidden paths;
|
|
41
|
+
3. a Binding to the implementation carrier when Counterfactual sensitivity is required;
|
|
42
|
+
4. a project-owned executable architecture check, such as the repository's lint, AST, dependency or contract test;
|
|
43
|
+
5. a separate Assertion when functional behavior could pass while the architecture invariant fails.
|
|
44
|
+
|
|
45
45
|
Do not encode subjective “clean architecture” or generic quality prose as machine authority. If no reliable observation can falsify it, keep it as durable Context/review judgment or return `decision_required`. Harness routes the repository's architecture check; it does not become a language-generic dependency analyzer.
|
|
46
46
|
|
|
47
|
+
## Proxy And Target Runtime Independence
|
|
48
|
+
|
|
49
|
+
When a declared result can pass on a proxy surface while failing in its target runtime, author independent target-runtime proof. Put the project-owned live Check in the earliest Outcome that owns the first runnable target boundary rather than postponing it to a terminal release/quality Outcome.
|
|
50
|
+
|
|
51
|
+
Use existing Contract semantics:
|
|
52
|
+
|
|
53
|
+
1. require `runtime_behavior` or the other proof surface that matches the actual Claim;
|
|
54
|
+
2. make the accepting runner exercise the target during the current Raw Execution and derive its asserted Observation from that same session;
|
|
55
|
+
3. include the runtime-affecting entrypoints, dependency manifests/lockfiles, configuration and integration carriers in `input_paths` or Bindings as appropriate;
|
|
56
|
+
4. freeze runner helpers/configuration as `verification_inputs` and declare only genuine environment requirements; and
|
|
57
|
+
5. add capability-specific probes only for Claims that actually require them.
|
|
58
|
+
|
|
59
|
+
A proxy check, static repository shape, tracked status report, prior screenshot, binary or historical run cannot be the sole proof of a Claim that can fail independently in the target. Do not add `platform_impact` flags, a platform taxonomy or per-platform completion state: Check ownership, proof surfaces, declared inputs, Bindings and risk already express the necessary boundary.
|
|
60
|
+
|
|
47
61
|
## Visual Delivery Authoring
|
|
48
62
|
|
|
49
63
|
When Source or controlling Context declares a design system, redesign, high-fidelity UI or other material visual result, author it through existing Contract semantics:
|
|
@@ -57,9 +71,9 @@ When Source or controlling Context declares a design system, redesign, high-fide
|
|
|
57
71
|
This guidance adds no visual Schema, Claim kind, risk level, lifecycle state, coverage artifact or Gate. It only makes visual meaning explicit enough for the existing Requirement/Control/Assertion and `ui_browser` mechanisms to verify what was actually declared.
|
|
58
72
|
|
|
59
73
|
## Compact Authoring
|
|
60
|
-
|
|
61
|
-
Compact V2 may omit only deterministic defaults: empty optional arrays/nulls, `context_snapshot_mode: referenced`, `requested_level: auto`, runner `argv: []`, `cwd: .`, `timeout_ms: 30000`, `retry_policy: none`, `idempotent: false`, and empty output/artifact/assertion/environment lists.
|
|
62
|
-
|
|
63
|
-
Goal, Source/Source Claims, Context, observable results, owners/paths, REQ, applicable CTRL states, OBL, proof surfaces, runner targets/effects, verification inputs, Assertions, risk, forbidden shortcuts and external confirmations remain explicit.
|
|
64
|
-
|
|
65
|
-
Compiler-generated Outcome/Check/Claim identities replace handwritten mechanical cross-entity references. This does not authorize compiler inference of product meaning, owners, architecture, proof or risk.
|
|
74
|
+
|
|
75
|
+
Compact V2 may omit only deterministic defaults: empty optional arrays/nulls, `context_snapshot_mode: referenced`, `requested_level: auto`, runner `argv: []`, `cwd: .`, `timeout_ms: 30000`, `retry_policy: none`, `idempotent: false`, and empty output/artifact/assertion/environment lists.
|
|
76
|
+
|
|
77
|
+
Goal, Source/Source Claims, Context, observable results, owners/paths, REQ, applicable CTRL states, OBL, proof surfaces, runner targets/effects, verification inputs, Assertions, risk, forbidden shortcuts and external confirmations remain explicit.
|
|
78
|
+
|
|
79
|
+
Compiler-generated Outcome/Check/Claim identities replace handwritten mechanical cross-entity references. This does not authorize compiler inference of product meaning, owners, architecture, proof or risk.
|
|
@@ -1,28 +1,36 @@
|
|
|
1
|
-
# Evidence Design Reference
|
|
2
|
-
|
|
3
|
-
Read this only while designing or repairing Contract Checks and proof.
|
|
4
|
-
|
|
5
|
-
## General Proof Rules
|
|
6
|
-
|
|
7
|
-
- Every Outcome has at least one executable Check and one non-Result atomic Claim.
|
|
8
|
-
- Required proof surfaces are non-empty, unique and all-of. Claim-bearing Assertions use explicit comparable Observations and expected values.
|
|
9
|
-
- `truthy`/`falsy` are diagnostic-only. `exists` proves only implementation-structure obligations. Missing or type-incomparable Observation never proves a Claim; negative proof uses an explicit value such as `equals: false`.
|
|
10
|
-
- Claim and Population proof is emitted only after the entire Check passes. Exit failure, missing artifact, failed population, failed Assertion or invalid Counterfactual yields no Claim proof.
|
|
11
|
-
- Verification inputs include entrypoints, helpers, fixtures/config, package scripts and lockfiles and cannot overlap implementation carriers.
|
|
12
|
-
- Runners receive the minimum environment whitelist plus only declared environment requirements. Never expose actual secret values in findings.
|
|
13
|
-
|
|
14
|
-
## Runner And Observation Identity
|
|
15
|
-
|
|
16
|
-
Evidence adapter is derived from runner kind. Only Playwright may prove `ui_browser`; structured runners prove non-browser surfaces. Raw Execution identity binds frozen runner identity and canonical declared Environment Requirements, not actual values.
|
|
17
|
-
|
|
18
|
-
Across all Checks sharing a Raw Execution, one Claim-bearing Observation belongs to one Assertion. Shared setup may execute once only when independent per-Check observations and artifacts remain unambiguous.
|
|
19
|
-
|
|
1
|
+
# Evidence Design Reference
|
|
2
|
+
|
|
3
|
+
Read this only while designing or repairing Contract Checks and proof.
|
|
4
|
+
|
|
5
|
+
## General Proof Rules
|
|
6
|
+
|
|
7
|
+
- Every Outcome has at least one executable Check and one non-Result atomic Claim.
|
|
8
|
+
- Required proof surfaces are non-empty, unique and all-of. Claim-bearing Assertions use explicit comparable Observations and expected values.
|
|
9
|
+
- `truthy`/`falsy` are diagnostic-only. `exists` proves only implementation-structure obligations. Missing or type-incomparable Observation never proves a Claim; negative proof uses an explicit value such as `equals: false`.
|
|
10
|
+
- Claim and Population proof is emitted only after the entire Check passes. Exit failure, missing artifact, failed population, failed Assertion or invalid Counterfactual yields no Claim proof.
|
|
11
|
+
- Verification inputs include entrypoints, helpers, fixtures/config, package scripts and lockfiles and cannot overlap implementation carriers.
|
|
12
|
+
- Runners receive the minimum environment whitelist plus only declared environment requirements. Never expose actual secret values in findings.
|
|
13
|
+
|
|
14
|
+
## Runner And Observation Identity
|
|
15
|
+
|
|
16
|
+
Evidence adapter is derived from runner kind. Only Playwright may prove `ui_browser`; structured runners prove non-browser surfaces. Raw Execution identity binds frozen runner identity and canonical declared Environment Requirements, not actual values.
|
|
17
|
+
|
|
18
|
+
Across all Checks sharing a Raw Execution, one Claim-bearing Observation belongs to one Assertion. Shared setup may execute once only when independent per-Check observations and artifacts remain unambiguous.
|
|
19
|
+
|
|
20
|
+
## Live Target Runtime Evidence
|
|
21
|
+
|
|
22
|
+
- For a target-runtime Claim, the accepting Check must exercise that target during the current runner invocation and derive structured Observations from the same runtime session. Rerunning a parser for a tracked or generated status report reruns the parser, not the target.
|
|
23
|
+
- A proxy surface may prove its own Claim but cannot substitute when proxy and target can fail independently. Static source/config shape proves structure only. The existence of a build, installation, started process or clean fatal-error scan proves only those exact assertions.
|
|
24
|
+
- If the declared result includes a runnable product surface or interaction, observe a stable product-owned sentinel or the declared interaction in the target session. A generic process/activity/window, development shell or absence of errors is insufficient for that broader Claim.
|
|
25
|
+
- Historical reports, screenshots, binaries and logs are review material. Current-run screenshots/logs may accompany a Check as Artifacts, but the accepting Observation must come from the live runner execution and cannot be imported from historical state.
|
|
26
|
+
- Bind every runtime-affecting implementation surface through `input_paths` and relevant Binding carriers; keep runner/helper/config files in `verification_inputs`. This lets existing Progress freshness identify when rolling feedback is stale without a new trigger registry.
|
|
27
|
+
|
|
20
28
|
## Playwright
|
|
21
|
-
|
|
22
|
-
Claim-bearing Playwright proof is only `playwright.case.<ac-key>.passed equals true`. `[ac:<assertion-key>]` binds one declared AC per Test Instance; ordinary tags are ignored and legacy `[<key>]` binds only a declared key.
|
|
23
|
-
|
|
24
|
-
Missing, skipped, flaky, unexpected, timed-out, interrupted, failed, multi-AC and duplicate-within-project cases fail closed. The same AC across distinct projects aggregates all-of. Aggregate status/count fields are diagnostic-only.
|
|
25
|
-
|
|
29
|
+
|
|
30
|
+
Claim-bearing Playwright proof is only `playwright.case.<ac-key>.passed equals true`. `[ac:<assertion-key>]` binds one declared AC per Test Instance; ordinary tags are ignored and legacy `[<key>]` binds only a declared key.
|
|
31
|
+
|
|
32
|
+
Missing, skipped, flaky, unexpected, timed-out, interrupted, failed, multi-AC and duplicate-within-project cases fail closed. The same AC across distinct projects aggregates all-of. Aggregate status/count fields are diagnostic-only.
|
|
33
|
+
|
|
26
34
|
Standard frozen Playwright verifier content is trusted. Weak-observability Outcomes require same-Check AC/Claim sensitivity. A weak Playwright Counterfactual may accept exit one only when every unexpected instance is uniquely a designated executed AC failure and there are no root, unbound, extra, missing, skipped, flaky, timeout, interruption, artifact, population, environment or other evidence failures. Ordinary Baseline Checks require exit zero.
|
|
27
35
|
|
|
28
36
|
## Visual UI Evidence
|
|
@@ -35,11 +43,11 @@ Standard frozen Playwright verifier content is trusted. Weak-observability Outco
|
|
|
35
43
|
- Keep subjective visual quality and approval external. A new visual direction or baseline that needs human judgment remains an explicit external confirmation even when all machine checks pass.
|
|
36
44
|
|
|
37
45
|
## Structured Evidence And Sensitivity
|
|
38
|
-
|
|
39
|
-
Every claim-bearing `structured_json_v2` Check needs same-Check Claim-related Counterfactual sensitivity unless the Claim is covered by that same Check's Population proof; weak observability removes that Population exemption. Artifacts and another Check never substitute for sensitivity.
|
|
40
|
-
|
|
41
|
-
Outcome Counterfactual V2 names an Outcome `binding_key`; Global Counterfactual V2 resolves an Outcome-owned `binding_ref`. A Counterfactual mutates only a proven subset of implementation carriers, never Source, Context, runners or verification inputs, and accepts only designated `assertion_value_mismatch` findings.
|
|
42
|
-
|
|
43
|
-
An `existing` mutation target must exist at Preflight/Compile. A `planned` target may be absent until implementation but must exist at Final Gate; once created, its changes stale targeted Progress. Population V2 proves exact eligible = observed + valid exclusions by entity id.
|
|
44
|
-
|
|
45
|
-
Artifacts remain review material. They do not prove Claim sensitivity by themselves.
|
|
46
|
+
|
|
47
|
+
Every claim-bearing `structured_json_v2` Check needs same-Check Claim-related Counterfactual sensitivity unless the Claim is covered by that same Check's Population proof; weak observability removes that Population exemption. Artifacts and another Check never substitute for sensitivity.
|
|
48
|
+
|
|
49
|
+
Outcome Counterfactual V2 names an Outcome `binding_key`; Global Counterfactual V2 resolves an Outcome-owned `binding_ref`. A Counterfactual mutates only a proven subset of implementation carriers, never Source, Context, runners or verification inputs, and accepts only designated `assertion_value_mismatch` findings.
|
|
50
|
+
|
|
51
|
+
An `existing` mutation target must exist at Preflight/Compile. A `planned` target may be absent until implementation but must exist at Final Gate; once created, its changes stale targeted Progress. Population V2 proves exact eligible = observed + valid exclusions by entity id.
|
|
52
|
+
|
|
53
|
+
Artifacts remain review material. They do not prove Claim sensitivity by themselves.
|
|
@@ -1,12 +1,12 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: normal-long-task
|
|
3
|
-
description: Retired compatibility pointer for users who explicitly invoke /normal-long-task. Direct them to /long-task-workflow and do not create legacy checklist, prompt, audit, matrix, verdict, or plan artifacts.
|
|
4
|
-
---
|
|
5
|
-
|
|
6
|
-
# Retired: Normal Long Task
|
|
7
|
-
|
|
8
|
-
`/normal-long-task` no longer defines a second long-task artifact workflow.
|
|
9
|
-
|
|
10
|
-
Use `/long-task-workflow` for one Canonical Delivery Contract, the current platform-native Goal, rolling implementation in the current workspace, targeted repair verification, a same-snapshot Final Gate and Stop freshness.
|
|
11
|
-
|
|
12
|
-
Do not create a preserved-source/checklist pair, target-mode prompt, Local Audit, matrix, verdict or second plan from this compatibility invocation.
|
|
1
|
+
---
|
|
2
|
+
name: normal-long-task
|
|
3
|
+
description: Retired compatibility pointer for users who explicitly invoke /normal-long-task. Direct them to /long-task-workflow and do not create legacy checklist, prompt, audit, matrix, verdict, or plan artifacts.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Retired: Normal Long Task
|
|
7
|
+
|
|
8
|
+
`/normal-long-task` no longer defines a second long-task artifact workflow.
|
|
9
|
+
|
|
10
|
+
Use `/long-task-workflow` for one Canonical Delivery Contract, the current platform-native Goal, rolling implementation in the current workspace, targeted repair verification, a same-snapshot Final Gate and Stop freshness.
|
|
11
|
+
|
|
12
|
+
Do not create a preserved-source/checklist pair, target-mode prompt, Local Audit, matrix, verdict or second plan from this compatibility invocation.
|