fdeops 4.0.4 → 4.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (139) hide show
  1. package/AGENTS.md +1 -1
  2. package/README.md +33 -6
  3. package/bin/check.js +14 -23
  4. package/bin/generate-skills.js +71 -0
  5. package/bin/install.js +2 -1
  6. package/bin/skill-catalog.js +17 -0
  7. package/mcp/fdeops-ingest/package.json +2 -2
  8. package/package.json +4 -3
  9. package/plugin.json +2 -2
  10. package/skills/fde/SKILL.md +18 -10
  11. package/skills/fde/references/board-memo.md +1 -1
  12. package/skills/fde/references/build.md +20 -0
  13. package/skills/fde/references/business-case.md +9 -7
  14. package/skills/fde/references/close.md +5 -3
  15. package/skills/fde/references/debug.md +18 -0
  16. package/skills/fde/references/encode-pattern.md +10 -8
  17. package/skills/fde/references/eval-pack.md +16 -33
  18. package/skills/fde/references/hold-scope.md +9 -7
  19. package/skills/fde/references/integrate.md +18 -0
  20. package/skills/fde/references/plan.md +4 -2
  21. package/skills/fde/references/poc.md +5 -3
  22. package/skills/fde/references/qa.md +18 -0
  23. package/skills/fde/references/review.md +22 -60
  24. package/skills/fde/references/ship.md +42 -287
  25. package/skills/fde/references/task-context.md +12 -0
  26. package/skills/fde/references/test-assumptions.md +2 -2
  27. package/skills/fde/references/three-options.md +19 -27
  28. package/skills/fde/references/verification.md +31 -0
  29. package/skills/fde-build/SKILL.md +21 -0
  30. package/skills/fde-build/references/build.md +20 -0
  31. package/skills/fde-build/references/debug.md +18 -0
  32. package/skills/fde-build/references/eval-pack.md +26 -0
  33. package/skills/fde-build/references/integrate.md +18 -0
  34. package/skills/fde-build/references/qa.md +18 -0
  35. package/skills/fde-build/references/review.md +39 -0
  36. package/skills/fde-build/references/ship.md +73 -0
  37. package/skills/fde-build/references/task-context.md +12 -0
  38. package/skills/fde-build/references/verification.md +31 -0
  39. package/skills/fde-debug/SKILL.md +21 -0
  40. package/skills/fde-debug/references/build.md +20 -0
  41. package/skills/fde-debug/references/debug.md +18 -0
  42. package/skills/fde-debug/references/eval-pack.md +26 -0
  43. package/skills/fde-debug/references/integrate.md +18 -0
  44. package/skills/fde-debug/references/qa.md +18 -0
  45. package/skills/fde-debug/references/review.md +39 -0
  46. package/skills/fde-debug/references/ship.md +73 -0
  47. package/skills/fde-debug/references/task-context.md +12 -0
  48. package/skills/fde-debug/references/verification.md +31 -0
  49. package/skills/fde-discover/SKILL.md +21 -0
  50. package/skills/fde-discover/references/audit.md +71 -0
  51. package/skills/fde-discover/references/discover.md +254 -0
  52. package/skills/fde-discover/references/task-context.md +12 -0
  53. package/skills/fde-evaluate/SKILL.md +21 -0
  54. package/skills/fde-evaluate/references/build.md +20 -0
  55. package/skills/fde-evaluate/references/debug.md +18 -0
  56. package/skills/fde-evaluate/references/eval-pack.md +26 -0
  57. package/skills/fde-evaluate/references/integrate.md +18 -0
  58. package/skills/fde-evaluate/references/qa.md +18 -0
  59. package/skills/fde-evaluate/references/review.md +39 -0
  60. package/skills/fde-evaluate/references/ship.md +73 -0
  61. package/skills/fde-evaluate/references/task-context.md +12 -0
  62. package/skills/fde-evaluate/references/verification.md +31 -0
  63. package/skills/fde-feedback/SKILL.md +21 -0
  64. package/skills/fde-feedback/references/encode-pattern.md +96 -0
  65. package/skills/fde-feedback/references/task-context.md +12 -0
  66. package/skills/fde-handoff/SKILL.md +21 -0
  67. package/skills/fde-handoff/references/close.md +66 -0
  68. package/skills/fde-handoff/references/encode-pattern.md +96 -0
  69. package/skills/fde-handoff/references/task-context.md +12 -0
  70. package/skills/fde-integrate/SKILL.md +21 -0
  71. package/skills/fde-integrate/references/build.md +20 -0
  72. package/skills/fde-integrate/references/debug.md +18 -0
  73. package/skills/fde-integrate/references/eval-pack.md +26 -0
  74. package/skills/fde-integrate/references/integrate.md +18 -0
  75. package/skills/fde-integrate/references/qa.md +18 -0
  76. package/skills/fde-integrate/references/review.md +39 -0
  77. package/skills/fde-integrate/references/ship.md +73 -0
  78. package/skills/fde-integrate/references/task-context.md +12 -0
  79. package/skills/fde-integrate/references/verification.md +31 -0
  80. package/skills/fde-options/SKILL.md +21 -0
  81. package/skills/fde-options/references/business-case.md +90 -0
  82. package/skills/fde-options/references/task-context.md +12 -0
  83. package/skills/fde-options/references/test-assumptions.md +102 -0
  84. package/skills/fde-options/references/three-options.md +90 -0
  85. package/skills/fde-poc/SKILL.md +21 -0
  86. package/skills/fde-poc/references/audit.md +71 -0
  87. package/skills/fde-poc/references/build.md +20 -0
  88. package/skills/fde-poc/references/business-case.md +90 -0
  89. package/skills/fde-poc/references/debug.md +18 -0
  90. package/skills/fde-poc/references/discover.md +254 -0
  91. package/skills/fde-poc/references/eval-pack.md +26 -0
  92. package/skills/fde-poc/references/integrate.md +18 -0
  93. package/skills/fde-poc/references/plan.md +167 -0
  94. package/skills/fde-poc/references/poc.md +55 -0
  95. package/skills/fde-poc/references/qa.md +18 -0
  96. package/skills/fde-poc/references/review.md +39 -0
  97. package/skills/fde-poc/references/ship.md +73 -0
  98. package/skills/fde-poc/references/task-context.md +12 -0
  99. package/skills/fde-poc/references/test-assumptions.md +102 -0
  100. package/skills/fde-poc/references/three-options.md +90 -0
  101. package/skills/fde-poc/references/verification.md +31 -0
  102. package/skills/fde-qa/SKILL.md +21 -0
  103. package/skills/fde-qa/references/build.md +20 -0
  104. package/skills/fde-qa/references/debug.md +18 -0
  105. package/skills/fde-qa/references/eval-pack.md +26 -0
  106. package/skills/fde-qa/references/integrate.md +18 -0
  107. package/skills/fde-qa/references/qa.md +18 -0
  108. package/skills/fde-qa/references/review.md +39 -0
  109. package/skills/fde-qa/references/ship.md +73 -0
  110. package/skills/fde-qa/references/task-context.md +12 -0
  111. package/skills/fde-qa/references/verification.md +31 -0
  112. package/skills/fde-readout/SKILL.md +21 -0
  113. package/skills/fde-readout/references/board-memo.md +108 -0
  114. package/skills/fde-readout/references/business-case.md +90 -0
  115. package/skills/fde-readout/references/readout.md +69 -0
  116. package/skills/fde-readout/references/task-context.md +12 -0
  117. package/skills/fde-review/SKILL.md +21 -0
  118. package/skills/fde-review/references/build.md +20 -0
  119. package/skills/fde-review/references/debug.md +18 -0
  120. package/skills/fde-review/references/eval-pack.md +26 -0
  121. package/skills/fde-review/references/integrate.md +18 -0
  122. package/skills/fde-review/references/qa.md +18 -0
  123. package/skills/fde-review/references/review.md +39 -0
  124. package/skills/fde-review/references/ship.md +73 -0
  125. package/skills/fde-review/references/task-context.md +12 -0
  126. package/skills/fde-review/references/verification.md +31 -0
  127. package/skills/fde-scope/SKILL.md +21 -0
  128. package/skills/fde-scope/references/hold-scope.md +83 -0
  129. package/skills/fde-scope/references/task-context.md +12 -0
  130. package/skills/fde-ship/SKILL.md +21 -0
  131. package/skills/fde-ship/references/build.md +20 -0
  132. package/skills/fde-ship/references/debug.md +18 -0
  133. package/skills/fde-ship/references/eval-pack.md +26 -0
  134. package/skills/fde-ship/references/integrate.md +18 -0
  135. package/skills/fde-ship/references/qa.md +18 -0
  136. package/skills/fde-ship/references/review.md +39 -0
  137. package/skills/fde-ship/references/ship.md +73 -0
  138. package/skills/fde-ship/references/task-context.md +12 -0
  139. package/skills/fde-ship/references/verification.md +31 -0
@@ -1,318 +1,73 @@
1
- # ship - Deliver the increment
1
+ # ship - Deliver and release with evidence
2
2
 
3
- **Enter when:** you are writing or updating on their codebase, they need to see something real, or you are going live.
3
+ **Enter when:** an implemented increment needs a delivery checkpoint, deployment, or wider rollout. Use [build](build.md) for implementation; an untested business or technical assumption needs an experiment before a release claim.
4
4
 
5
- **Read first:** `context.md`, `decisions.md`, `delivery.md`, `success.md`. Load `terrain.md` before you touch their code. Load `trust-profile.md` if the deploy touches regulated data or needs an approval chain. Load `evals.md` when the work touches AI/ML/LLM/RAG/agents.
5
+ Start from [task context](task-context.md). Standalone work uses supplied permitted context and a release receipt; it does not require `.fde/` initialization. In engagement mode, use privacy-safe views of confirmed context, decisions, terrain, success, delivery, applicable trust constraints, and AI evaluation evidence. Never load raw `<private>` blocks into a model. The `fde` CLI remains local-only; deployment uses the customer's authorized tools, never a new network capability inside `fde`.
6
6
 
7
- Do not ask them to pick a mode. Name where you are, then start at the matching section:
7
+ ## Establish the delivery contract
8
8
 
9
- - Nothing on their staging yet → **one change they can see**
10
- - On staging, the signer in `success.md` can reject it → **go-live**
11
- - Prod is the question → **go-live**. Do not start a second change.
9
+ Identify the exact outcome, acceptance check, scope, affected users/systems, target environment, recovery mechanism, and who or what is authorized to accept and release it. Reuse confirmed authority and checks for routine work. Do not invent missing signers, permissions, measurements, or acceptance.
12
10
 
13
- If going live, check the evidence for the recovery path: has rollback, restore, compensation, or roll-forward been exercised under representative conditions? If only planned, validate it before release. Reuse applicable drill evidence when the mechanism and relevant conditions are unchanged; record why it applies.
11
+ For an initialized engagement, run `fde doctor --ready` before a new delivery plan or material scope change. Missing binary success or a named customer-side signer blocks that planning progression until resolved. A passing doctor validates record structure, not connectivity, release readiness, or customer acceptance. Standalone work evaluates the supplied contract directly.
14
12
 
15
- A bounded experiment that tests an assumption is `poc`. This skill turns a validated direction into a maintainable change on a repo they will own, then production. Inspect existing prototype code and retain suitable tested parts; replace unsafe shortcuts based on evidence. A successful demo alone does not satisfy the readiness gates below.
13
+ A customer delivery checkpoint must let the agreed decision-maker replay and reject the acceptance check through an interface they operate. Prefer their staging; otherwise use an agreed representative environment and disclose its owner and limitations. Local green proves only the local run. Routine fixes may share an agreed checkpoint; no fixed number of changes forces a ceremony.
16
14
 
17
- **Before a new delivery plan or material scope change:** run `fde doctor --ready`. Missing binary success or a named customer-side signer blocks progression: review the proposed acceptance check and authority with the FDE first. Use a test/input and observable pass/fail under **Done when:** or **Acceptance check:**. A number, role, or successful demo alone is insufficient. Do not invent missing facts to pass lint. Routine reversible fixes within confirmed scope reuse the existing signer and acceptance check; record verification without restarting approval. New judgment in the record still requires confirmation.
15
+ ## Prepare a reviewable increment
18
16
 
19
- ## Field (name it once, then the same loop)
17
+ 1. Inspect repository instructions, working tree, overlapping work, and the complete intended release diff. Include working-tree changes when testing an uncommitted candidate. Preserve unrelated work; separate unintended behavior before release.
18
+ 2. Identify applicable before-state evidence and the changed outcome. Name dependencies, stop conditions, and irreversible effects. Existing applicable evidence may be reused with attribution, never represented as a fresh run.
19
+ 3. Complete [verification](verification.md) and [review](review.md), proportional to the change and repository requirements. A self-check is not an independent review. Record command, revision, environment, date, result, and unrun checks. Exercise relevant operating exceptions and fallback paths, not just the happy path.
20
+ 4. For AI behavior, obtain a scoped [eval verdict](eval-pack.md) for the candidate and applicable controls. A permitted bounded automation workflow remains permitted within its documented limits. A missing evaluation, failed critical case, or unknown action authority prevents release of that path; non-AI changes record eval as not applicable.
20
21
 
21
- | | Brownfield | Greenfield |
22
- |--|------------|------------|
23
- | What you touch | Code they already run | A new path or empty tree they will own |
24
- | First move | Characterise their tests, their runner, the workaround in `terrain.md` | First path a user can click. Not the whole product. |
25
- | Proof | Agreed representative environment and replayable acceptance check | Agreed representative environment and replayable acceptance check; record what remains untested before release |
26
- | Undo | Revert this change on its own | Name rollback or tested recovery; identify irreversible effects and required authority. |
22
+ ## Release gate
27
23
 
28
- Skip POC only when the killer assumption already lives in the repo (typical brownfield). If the bet is unproven, `poc` first.
24
+ Before deployment, establish these facts from existing evidence or a necessary check. Missing material evidence blocks the dependent release step; continue independent preparation. Do not ask again for approval already provided within the same scope.
29
25
 
30
- **Customer delivery means:** the signer in `success.md` can replay and reject the agreed acceptance check in an environment they operate. A green check on your laptop proves only what ran there. Routine fixes can share a delivery checkpoint; distinguish implementation, verification, deployment, and acceptance.
26
+ | Dimension | Required evidence |
27
+ |-----------|-------------------|
28
+ | Candidate | Exact revision/artifact, intended diff, dependencies, applicable required checks passing; no skipped failure presented as green |
29
+ | Target and access | Service/account/region, environment, authorized deployment identity and mechanism, secret provisioning without revealing values |
30
+ | Acceptance | Replayable check and agreed decision-maker/mechanism; record actual acceptance separately from readiness |
31
+ | Data and policy | Permitted data, applicable security/residency/change-window requirements, necessary approvals already recorded or obtained |
32
+ | Recovery | Applicable tested rollback, restore, compensation, or roll-forward within agreed recovery-time/data-loss limits; explicit authority for irreversible effects |
33
+ | Operations | Named release/recovery owner, runbook appropriate to risk, health and business signals, stop thresholds, observation coverage |
34
+ | AI, when applicable | Current applicable SHIP eval evidence, critical failures zero, enforced action boundary and required human review or documented bounded automation |
31
35
 
32
- If `terrain.md` **Data estate** lists a **Blocker** this change depends on (source or pipe): stop. That is discover, not ship. Do not build a path they cannot feed.
36
+ Check migration compatibility, old/new version coexistence, delayed jobs, caches, and already-emitted side effects where relevant. A code revert does not undo data loss or external writes. Reuse drill evidence only when the mechanism and relevant conditions are unchanged, explaining applicability. If recovery is only a plan, exercise it in a permitted representative environment before release.
33
37
 
34
- ## Method - one change they can see
38
+ Use the repository's existing secret scanning and security checks; avoid diagnostic commands that print credential matches. Retain sanitized references to results. Resolve material evidence gaps or obtain an explicit, authorized narrowing of the release; do not average critical blockers into a readiness score.
35
39
 
36
- One change = one coherent outcome with observable verification and a bounded recovery path. Prefer vertical slices that can be reviewed and exercised independently. A PR is how this often lands. It is not the job. The job is the change they can see.
40
+ For a coordinated engagement, also connect the release to the agreed value bucket and baseline/target, and record a dated receipt for the affected operating path. An unmeasured result remains pending with a measurement next step; do not invent realized value to pass a gate.
37
41
 
38
- ```
39
- BAD (layers):
40
- 1: all database models
41
- 2: all API endpoints
42
- 3: all UI components
43
- 4: wire everything together (and pray)
42
+ ## Deploy within authority
44
43
 
45
- GOOD (one user action each):
46
- 1: User can create a payment (schema + endpoint + minimal UI) - testable
47
- 2: User can view payment status (query + endpoint + UI) - testable
48
- 3: Payment retry on failure (logic + endpoint + UI feedback) - testable
49
- 4: Admin can void a payment (auth + logic + UI) - testable
50
- ```
44
+ Execute only when the requested workflow authorizes deployment to this target and the applicable gates are met. Otherwise leave a concrete release candidate, exact deployment/recovery instructions, evidence, and the remaining authorization for review. A permission to implement or test is not permission to publish.
51
45
 
52
- Prefer independently revertible changes. When data or external effects cannot be undone, name the dependency, containment, tested recovery, and authorized owner before release.
46
+ Use the customer's established pipeline and rollout mechanism. Select canary, staged exposure, blue/green, or direct rollout according to actual risk and platform capabilities; do not impose a universal cohort sequence. Define advance/abort thresholds and observation window before starting. If another operator must execute, record their handoff and report deployment pending until there is evidence it happened.
53
47
 
54
- **Before you start this change:**
48
+ During rollout inspect health, errors, key user behavior, and side-effect integrity. Halt expansion on breached thresholds or critical harm and apply authorized containment/recovery. Do not continue merely because the deploy command exited successfully.
55
49
 
56
- - [ ] It is in `decisions.md` with acceptance criteria (happy + unhappy path)
57
- - [ ] Blast radius declared: which files, which systems, which users affected
58
- - [ ] Rollback named: revert this change, or something more specific
59
- - [ ] No dependency on an unmerged change (if dependent, state it and land in order)
60
- - [ ] `Kill if` is written - the observation that stops this change
61
- - [ ] Before-state evidence identified: the relevant failing output, number, or behavior, with its source and date. For a routine fix within confirmed scope, reference applicable existing evidence and batch the delivery receipt; capture new evidence when the relevant behavior or conditions changed. Never imply an old check was rerun.
62
- - [ ] Open PRs and uncommitted work in the area checked (`gh pr list`, `gh pr diff <n> --name-only`); overlap goes to `decisions.md` before you start
50
+ ## Verify operation and hand off
63
51
 
64
- Your coding pack writes the function. This skill owns done. When they disagree with this repo, the repo wins.
52
+ Run permitted smoke and acceptance checks against the deployed candidate. Record deployment identity/time, observed signals, sample/window, failures, recovery actions, and remaining gaps. Define the pulse: metric, cadence, threshold, owner, and response. For AI, include permitted output sampling and drift/action-boundary monitoring.
65
53
 
66
- **The loop.** In this order:
54
+ Before wider exposure, verify expected load/cost, data pipeline behavior, ownership, support, and applicable governance for the proposed audience. Choose expansion conditions from evidence; a successful pilot does not establish readiness for an arbitrary larger scale. Measure adoption against the eligible users, expected workflow frequency, and agreed observation window; investigate misses without guessing their cause.
67
55
 
68
- ```
69
- Read existing code in the area (search before creating)
70
- → Characterise what is already there (their tests, their runner; greenfield: the empty tree)
71
- → Implement the smallest path that works
72
- → Verify the change; demonstrate at the agreed delivery checkpoint (below)
73
- → Cleanup pass (dedupe, simplify - behaviour unchanged)
74
- → Self-review against acceptance criteria
75
- → Commit with a message the client's team can read
76
- → Update decisions.md + delivery.md
77
- ```
56
+ ## Receipt and completion
78
57
 
79
- **Verify the change and prove customer delivery.** Match evidence to the reviewed revision, environment, and acceptance criteria.
58
+ Keep these claims separate: implemented, verified, deployed, measured outcome, and accepted. Include the candidate, target, command/pipeline, applicable checks and unrun checks, review source, evaluation where needed, authority source, recovery evidence, observation, and next owner/action. Attribute acceptance to its actual source and scope. A staging measurement is not production value, and a commit is not deployment.
80
59
 
81
- - Use **their** test commands, fixtures, and CI. Record the command, result, revision, environment, and run date in `delivery.md`. Reuse existing evidence only when the relevant code and conditions are unchanged, citing why it still applies; never claim it was rerun. Run affected checks for changed behavior and required release checks before deployment. Missing evidence means unproven, not an observed failure.
82
- - At the agreed delivery checkpoint, the signer in `success.md` must be able to replay and reject the acceptance check using an interface they operate (screen, API, report, or equivalent). Routine fixes can share that checkpoint; passing tests alone does not establish customer acceptance.
83
- - Prefer staging they operate. When unavailable, use an agreed, permitted representative test environment, record its owner and limitations, and resolve material release-evidence gaps before production. A local demonstration is not deployment.
84
- - **Representative data.** Use permitted sanitized or synthetic fixtures that exercise relevant volumes, edge cases, and operating paths. Before go-live, record gaps such as batch timing, distribution, or production-only dependencies and their impact on the acceptance and abort checks. Resolve material gaps or explicitly narrow the release; never load sensitive production data merely to make a demo realistic.
85
- - Model in the path: `eval-pack` until `evals.md` says SHIP. Do not skip because "it looked right in chat."
86
- - A model drafts. A named human on their side ships. No unsupervised loop on their production. If the brief demands lights-out write-access, that is `who-decides` / `hold-scope`, not ship.
87
-
88
- The proof is whatever this client already believes, plus one new receipt they can replay.
89
-
90
- **Size by reviewability and risk.** Keep one coherent intent, bounded context, and observable acceptance checks. Split unrelated behavior or work whose recovery and review cannot be understood together. Diff size and elapsed time are warning signals, not hard gates: generated changes may be large and low risk; a one-line permission change may be critical. Use the repository’s checks and add meaningful coverage for changed behavior, rather than a test-count quota.
91
-
92
- **Show it.** Every 2-3 changes, something the customer can see: an endpoint they can hit, a UI they can click, a metric that moved, a risk that was retired. Technical progress invisible to stakeholders is trust decay. `delivery.md` gets updated after every visible change.
93
-
94
- **The scope trap.** Mid-change discoveries ("this module also needs updating," "I should refactor this while I'm here"):
95
-
96
- - If it's in `decisions.md`: do it as a separate change.
97
- - If it's NOT in `decisions.md`: log it as a scope receipt (see `hold-scope.md`), don't touch it.
98
- - Ugly code outside this change stays ugly. That is discipline, not laziness.
99
-
100
- After each change: required checks pass with applicable evidence, acceptance criteria evaluated, blast radius as declared, `Kill if` still false. At the agreed delivery checkpoint: what did they see, and what is their signal? Before production, complete the go-live gates below.
101
-
102
- ---
103
-
104
- ## Deployment readiness gate (confirm the target before building the runway)
105
-
106
- Before scoring readiness, confirm WHERE this is going. State it in 2-3 lines - brief playback that invites correction:
107
-
108
- > "Deploying to: [target]. Pipeline: [how it gets there]. Rollback mechanism: [how to undo]. Any constraint I should know about?"
109
-
110
- **The checklist (confirm, don't assume):**
111
-
112
- | Dimension | Question | Status |
113
- |-----------|----------|--------|
114
- | **Target** | Cloud provider + service (ECS/Lambda/K8s/VM/on-prem)? | |
115
- | **Pipeline** | CI/CD exists? Manual? Who triggers prod deploy? | |
116
- | **Environments** | Dev → staging → prod path clear? Or deploying direct? | |
117
- | **Secrets** | Where do they live? (vault/SSM/env vars) Who provisions? | |
118
- | **Access** | Do YOU have deploy permissions, or does someone else push? | |
119
- | **Compliance** | Region constraints? Data residency? Encryption requirements? CAB/change window? | |
120
- | **Infra-as-code** | Terraform/Pulumi/CDK/manual? State file location? | |
121
-
122
- **If anything is blank:** ask now. Discovering deployment constraints after the change is where timelines slip. If the client hasn't defined these yet, that's a conversation before you write the runbook - not after.
123
-
124
- Write confirmed deployment context to `delivery.md` under a `## Deployment target` section.
125
-
126
- ---
127
-
128
- ## Method - readiness gate (score before touching the deploy button)
129
-
130
- Score each dimension green/amber/red. This is the gate, not a suggestion:
131
-
132
- | Dimension | Green | Amber | Red |
133
- |-----------|-------|-------|-----|
134
- | Tests | All pass on deploy branch | Flaky tests skipped with justification | Failures present or tests not run |
135
- | Recovery | Applicable tested rollback/restore/compensation/roll-forward meets agreed recovery and data-loss limits | Documented; drill evidence needs refresh | No viable recovery, failed drill, or irreversible effects lack explicit authority |
136
- | Sign-off | Stakeholder approval in `decisions.md` with date | Verbal approval, not logged | No approval sought |
137
- | Runbook | Exists and someone other than you has read it | Exists but unreviewed | Missing |
138
- | Monitoring | Alerts configured, owner named, dashboard live | Alerts configured, no named owner | No monitoring |
139
-
140
- ### Value + receipts gate (score with the table above)
141
-
142
- | Dimension | Green | Amber | Red |
143
- |-----------|-------|-------|-----|
144
- | **Value bucket** | `success.md` names primary bucket (`cost-save` \| `risk-mitigation` \| `revenue-uplift`) and a baseline→target metric; this change's value-ledger row has **Bucket** + **Promised** | Bucket named; **Measured** still `pending` with a pulse date | No bucket, or Promised empty / ticket-theater only |
145
- | **Audit receipt** | Dated line in `delivery.md` (`## Ship receipts` or ledger Evidence) proving exceptions/operating path were walked - cite `terrain.md` / `reality.md` / `audit.md` | Path described, not verified this ship | No audit receipt for this change |
146
- | **Eval receipt** | **n/a** (no AI on this change) **or** `evals.md` Verdict SHIP with dated golden run + HITL gate named | Eval pack exists; known fails open with owner + date | AI in scope and no eval receipt |
147
- | **AI eval pack** | `.fde/evals.md` Verdict SHIP; goldens run this change; critical fails 0; HITL filled if policy requires | Pack exists; run stale vs change log | AI-touching deploy and pack missing / NO-SHIP / HITL required but empty |
148
-
149
- **Any RED = stop. Do not deploy. Fix the red dimension first.**
150
- **2+ AMBER = sponsor conversation before deploying.** Present the ambers and get explicit "proceed" or "fix first."
151
-
152
- **AI-touching deploys (model, embeddings, RAG, agent, or inference path):**
153
- 1. Read `.fde/evals.md`. If missing → **RED. Do not deploy.** Create the pack (`eval-pack` / `ai` overlay) and re-score.
154
- 2. If Verdict is not **SHIP**, or Last run is older than the latest change-log row → **RED.**
155
- 3. If `trust-profile.md` requires human-in-the-loop and the HITL gate has no reviewer → **RED.**
156
- 4. Log in `delivery.md` → `## Ship receipts` before deploy: audit cite + eval receipt.
157
- 5. Non-AI deploys: Eval = **n/a** - do not invent an empty pack.
158
-
159
- Write the readiness score (including value + receipts) to `delivery.md` before deploying. The score is the evidence if anything goes wrong.
160
-
161
- ## Intent vs diff (before pre-blast)
162
-
163
- Ship the change you intended - not the drift that snuck in. Run this on the deploy branch against the **one-line intent** from `decisions.md` / `success.md` (the change you said you were building).
164
-
165
- ```bash
166
- git diff <base>...HEAD --stat
167
- git diff <base>...HEAD
168
- ```
169
-
170
- Score every touched path (or logical hunk):
171
-
172
- | Path / change | Verdict | Rule |
173
- |---------------|---------|------|
174
- | | **KEEP** | Directly required for the stated intent |
175
- | | **JUSTIFY** | Adjacent but load-bearing - one sentence why it must ship *now*, or split |
176
- | | **SPLIT** | Real work, wrong change - park in `decisions.md` kill/Next; do not deploy with this one |
177
- | | **DROP** | Noise (format-only, drive-by rename, unrelated tidy) - revert before ship |
178
-
179
- **Any SPLIT or DROP still in the tree = fix-first.** JUSTIFY without a written sentence = treat as SPLIT. Log a one-line receipt in `delivery.md`: `intent vs diff: KEEP n · JUSTIFY n · SPLIT n · DROP n - <intent>`.
180
-
181
- This is **code drift**, not stakeholder "also can you…" (that is `hold-scope`). Same family as review Stage 1 - ship refuses green when the diff outgrew the claim.
182
-
183
- ## Pre-blast challenge (before the deploy button)
184
-
185
- For any non-trivial go-live (shared infra, regulated data, irreversible migration, or first prod touch), run this once before canary - not as theater, as a stop-the-line check:
186
-
187
- ```
188
- CLAIM: <what you are about to ship, in one sentence>
189
- WHY IT MATTERS: <blast radius / who feels pain if wrong>
190
- CHALLENGE: <the strongest argument this is not ready - grounded in delivery.md / risks.md / trust-profile.md>
191
- VERDICT: proceed | fix-first | sponsor conversation
192
- ```
193
-
194
- Rules: no invented stakeholders; if evidence is missing, the verdict is **fix-first** or **sponsor conversation**, not "probably fine." Log the CLAIM + VERDICT as a dated line in `delivery.md`. Skip for mechanical one-line config with an already-tested rollback.
195
-
196
- ## Method - pre-flight (you verify each, confirmed not assumed)
197
-
198
- - All tests pass - state the command and result.
199
- - No hardcoded secrets/credentials (repeat `--include` per extension - brace globs silently match nothing):
200
- ```bash
201
- grep -rnE "(api[_-]?key|secret|password|token)\s*[:=]\s*['\"][^'\"]{8,}" \
202
- --include="*.js" --include="*.ts" --include="*.py" --include="*.env" \
203
- --include="*.yaml" --include="*.json" . | grep -vE "example|template|test" | head
204
- ```
205
- - DB migrations checked for compatibility, data loss, and old/new application coexistence. Prefer expand/contract for destructive changes. Irreversible steps require explicit authority and a tested restore, compensation, or roll-forward plan.
206
- - Recovery documented **and tested**, with acceptable recovery time and data loss.
207
- - Monitoring alerts configured, someone watching.
208
- - Team knows the deploy is happening.
209
- - Deploy window has staffed observation and recovery coverage appropriate to the risk; respect the client’s change calendar and business-critical periods.
210
- - **Change approval (CAB) environments:** window open, ticket approved. In banking/healthcare/gov, deploying outside an approved window is a compliance finding even when the deploy succeeds. "We didn't know there was a CAB process" is not a defence - find out before the deploy date.
211
-
212
- ## Method - the deploy
213
-
214
- **Rollout:** Choose canary, blue/green, staged cohorts, or the client’s proven release mechanism based on isolation, traffic, and failure cost. For a canary, set cohort size, exposure cap, observation duration, minimum sample, and advance/abort thresholds before starting; allow for delayed and batch effects. Watch errors, latency, and **the business metric this change affects**. Breached thresholds or critical harm → halt expansion and execute the tested recovery/containment plan; investigate after exposure is controlled. Advance only with sufficient evidence and a named operator.
215
-
216
- **Canary receipt** (write it, or the canary did not happen): what was watched, on whose dashboard, for how long, and that the next change did not start in the window. If prod is a CAB console, vendor button, or their pipeline, write the owner and the click path - the host agent does not get to pretend it shipped.
217
-
218
- **Programme-scale rollout (transformations)** - different problem from one service:
219
- 1. **Pilot** - one team, one use case; success metrics defined *before* it starts (after = fitting metrics to results).
220
- 2. **Limited release** - 3-5 teams, real load; this is where the failure modes the pilot hid show up.
221
- 3. **Broad release** - self-serve onboarding; if teams still need the FDE to start, onboarding isn't finished.
222
- 4. **Enterprise standard** - the FDE is no longer needed for this use case. That's the end state.
223
- Straight from pilot to standard = a high-profile failure at scale.
224
-
225
- ## Method - after
226
-
227
- Keep implementation, test results, deployment, measured outcome, and customer acceptance separate in the receipt. Record the environment, observation window/sample, source, and remaining gaps. A commit is not a deploy; a staging measurement is not realized production value. When the baseline is missing or incomparable, record the observed result and the measurement next step without claiming an improvement. Record acceptance only for what the named person actually accepted, with a dated source; an engineer's summary remains attributed to that summary.
228
-
229
- Smoke tests against production. Verify the business metric moved the right way. Then **define the pulse before closing the laptop** - a deploy without a pulse is one you'll hear about only when it breaks:
230
-
231
- 1. **Metric:** the number that says it's working - "p99 on payment endpoint < 800ms", not "errors low."
232
- 2. **Frequency:** daily week one, weekly after, monthly when stable.
233
- 3. **Threshold:** the exact value that triggers incident response. Nobody knows the number → nobody acts until too late.
234
-
235
- AI components: also define what *normal output* looks like and check a weekly sample of real production outputs - drift is technically-valid-but-wrong, and no exception will fire.
236
-
237
- ## Method - scale readiness (pilot proved it, now deploy enterprise-wide)
238
-
239
- A successful pilot does not establish readiness for wider use. Check organizational ownership, governance, and infrastructure alongside technical performance before expanding.
240
-
241
- **The scale-readiness gate (all must be YES before broad rollout):**
242
-
243
- | Dimension | Question | Ready? |
244
- |-----------|----------|--------|
245
- | **Infra** | Can the system handle 10× current load without architectural change? | |
246
- | **Ops** | Can someone other than the FDE operate it at 2am? (runbook exists, tested) | |
247
- | **Data** | Is the data pipeline automated, not manual? Does it handle upstream schema changes? | |
248
- | **Security** | Has infosec signed off for production data at scale? | |
249
- | **Cost** | Is the cost model viable at 10× volume? (AI inference costs scale non-linearly) | |
250
- | **Governance** | Is there an owner, a review cadence, and an escalation path? | |
251
- | **Support** | Can users get help without the FDE? (docs, training, L1 support path) | |
252
- | **Measurement** | Are success metrics automated and dashboarded, not manually calculated? | |
253
-
254
- **If any dimension is "No":** that's the work before scaling. Name it, size it, put it in the plan. Scaling without readiness = a high-profile failure that kills the entire programme.
255
-
256
- **The scale sequence:**
257
- 1. **Pilot** (1 team, controlled) → prove value, find failure modes
258
- 2. **Limited** (3-5 teams, real load) → prove operability, find scale bugs
259
- 3. **Broad** (self-serve onboarding) → prove the team doesn't need the FDE
260
- 4. **Standard** (enterprise default) → the FDE exits this workstream
261
-
262
- Never skip a step. The sponsor always wants to skip from pilot to standard - that's the conversation the FDE protects.
263
-
264
- ## Method - progressive adoption (built it, now people need to use it)
265
-
266
- Adoption isn't a handoff-stage problem - it starts while you are still writing the change. Software that launches to silence is software that gets decommissioned.
267
-
268
- **During the change:**
269
- - **Controlled exposure.** Use a feature flag or equivalent isolation when it reduces rollout risk. Choose cohorts and expansion criteria from traffic and impact; name the flag owner and removal point.
270
- - **Feedback loops built in.** A thumbs-up/down, a "was this helpful?", a usage counter. Instrument adoption, don't assume it.
271
- - **Resistance signals.** Watch for: workaround creation (they built a spreadsheet instead of using the tool), drop-off after day 3 (onboarding fails), vocal detractors (one influential skeptic can kill adoption). Address these before launch, not after.
272
-
273
- **At launch:**
274
- - **Champion network.** Identify 2-3 power users per team who adopt early. Support them intensely - they become your multiplier.
275
- - **Adoption targets agreed before launch.** Define the eligible users, expected usage frequency, observation window, baseline, and owner. A weekly workflow needs a different measure from a quarterly one. Investigate misses with users; usage alone does not establish whether onboarding, access, or value is the cause.
276
- - **The "switching cost" test.** If users can still do it the old way, they will. Adoption requires either: the old way is removed, the new way is dramatically better, or management mandates the switch. Know which lever applies.
277
-
278
- **Write adoption metrics to `delivery.md`:** active users, frequency, drop-off points, resistance signals. This is the evidence for renewal.
279
-
280
- ## Artifact
281
-
282
- **`decisions.md`** - each change: what was implemented, what was tested, what was deferred, `Kill if`.
283
-
284
- **`delivery.md`** - each visible change in business language; then the deployment record: what shipped, when, recovery procedure, pulse definition, **scale-readiness assessment, and adoption metrics**. Written for whoever inherits the system.
285
-
286
- ## Checkpoint
287
-
288
- After each change: required checks pass, acceptance criteria evaluated, blast radius as declared, `Kill if` still false. Batch routine fixes at the agreed delivery checkpoint; record staging and customer acceptance separately.
289
-
290
- Before full exposure: the chosen rollout’s advance criteria are met with sufficient observation and business-metric evidence, and the pulse is written into `delivery.md`. Also green: value bucket named, audit receipt dated, eval receipt **n/a or pass**, **intent vs diff clean** (no unresolved SPLIT/DROP). Missing any of those → not green. For enterprise-scale: scale-readiness gate passed before broad rollout.
60
+ In engagement mode, write confirmed implementation/decisions and delivery receipts under the existing record rules. Standalone work returns the same receipt or uses the repository's permitted release record. Committing, pushing, opening a PR, publishing, and notifying others are actions governed by the user's workflow, not mandatory steps imposed by this method.
291
61
 
292
62
  ## Worked example
293
63
 
294
- Acme, brownfield. Plan Now has three changes, not "the payments rewrite."
295
-
296
- Change 1 is *user sees retry status on a failed payment* - schema + endpoint + the existing ops screen, 180 lines, their `pytest -k payments` green, revert is this change. `Kill if:` the signer cannot reject it on the screen they already use. Ugly retry-queue code two files over stays ugly. `decisions.md` logs the change; `delivery.md` says ops can see a retry without opening the spreadsheet. Marco sees it on staging they operate. That is the proof. Local green was not.
297
-
298
- Then Thursday go-live of the failure-routing change. Readiness scoring catches two things the diff does not. The audit receipt is missing: the operating map says Marco's manual re-run is the fallback, and nobody has checked whether the new page fires *before* his morning run or after - if after, the alert changes nothing. That gets walked and cited before deploy. Second, the intent-vs-diff read shows the PR also touches the settlement retry that was deferred; it comes out.
299
-
300
- Pre-blast challenge: "what does this break if it fires at 3am and nobody acks?" Answer: nothing breaks, but the rota is not yet agreed - so the deploy waits on a name, not on code. That is a one-day slip that prevents a fake green.
301
-
302
- After deploy: `delivery.md` ship receipt with the audit cite, the kill test evidence, and the rollback line. Eval receipt: n/a, no AI in this path.
64
+ A freight team agrees that dispatchers can retry a failed export once without creating a duplicate shipment. The change uses the existing queue and ops screen. The developer records a failing duplicate-delivery case, implements idempotency, and passes the relevant checks on a named candidate. QA observes both the retry status and the single downstream record on permitted staging fixtures. A separate reviewer examines the queue race; the receipt names that review and the revision.
303
65
 
304
- Greenfield is the same loop with an empty tree: first path a user can click, on an environment they will operate, then this go-live. Not the whole product in one dump.
66
+ The export service has an approved staged-release workflow. Its owner reuses a recent recovery drill because the queue format and recovery mechanism are unchanged, recording that applicability. Deployment stops if duplicate records appear or the agreed error threshold is crossed. The authorized rollout completes, production smoke checks pass, and the dispatcher accepts the specified retry behavior with a dated source. The operating-cost benefit remains pending until the agreed measurement window closes. Implementation, deployment, acceptance, and measured value have different evidence. In the existing engagement, `decisions.md` records the agreed behavior and `delivery.md` holds the release receipt; standalone work returns those facts directly.
305
67
 
306
68
  ## Principles
307
69
 
308
- - One user action per change. Layers are untestable until assembled.
309
- - Prove the agreed outcome in their environment and test recovery. Local green is not customer delivery.
310
- - The ugly code outside this change stays ugly. That's discipline, not laziness.
311
- - A deployment needs tested recovery within agreed time and data-loss limits; irreversible effects require explicit authority.
312
- - Halt expansion on breached thresholds or critical harm; contain exposure with the tested recovery plan before investigating.
313
- - Verify the business metric, not just the technical one.
314
- - No value bucket, no green ship. No pulse, no done.
315
- - Diff larger than the stated intent without KEEP/JUSTIFY receipts = fix-first.
316
- - AI path without eval receipt = fix-first; non-AI ships leave eval as n/a.
317
- - Scale readiness is organizational, not just technical. Check all 8 dimensions.
318
- - Adoption is measured from day one, not hoped for at launch.
70
+ - Release the reviewed candidate with applicable evidence and documented authority.
71
+ - Test recovery against the effects that actually persist beyond a code revert.
72
+ - Keep missing evidence visible and distinguish local, staging, and production claims.
73
+ - Expansion follows observed acceptance and operating limits; fixed ceremonies cannot replace them.
@@ -0,0 +1,12 @@
1
+ # Task context and evidence
2
+
3
+ Use this contract for standalone methods and methods routed through `@fde`.
4
+
5
+ - **Standalone work:** use the supplied, permitted facts, notes, code, and artifacts. A client name, `.fde/` directory, or initialized engagement is not a prerequisite. Do not bootstrap records merely to run a method. Ask only for missing information or authority that changes the next action; mark other gaps as unknown.
6
+ - **Artifact names are destinations:** names such as `success.md`, `decisions.md`, and `delivery.md` identify relevant evidence and, when bound, record destinations. If absent, use supplied facts and return the requested draft or result in the current workspace or conversation. Do not invent files or require initialization to complete useful work.
7
+ - **Bound engagement:** honor the current client binding and constraints. Before reading records, run `fde privacy` to verify masking support. Obtain a fresh, identity-matching sanitized `fde resume` packet for this task (or reuse a fresh session-hook packet); retrieve missing evidence with targeted `fde recall <topic>`. Use bounded `fde handoff` for transfer work. Refresh after binding, masking, or record changes. Never substitute raw `.fde/` reads, private blocks, masking dictionaries, or full transcripts. If the CLI is unavailable, use only permitted supplied excerpts and report the context limitation.
8
+ - **Authority:** continue reversible work within authorized scope. Reuse prior authorization when it covers the specific action. Show consequential engagement-record judgments and uncertainties for confirmation before saving unless already explicitly confirmed. New scope, acceptance changes, production actions, exports, and external messages need the applicable authority; a method invocation alone does not supply it. Keep one customer's writes in that customer's record.
9
+ - **Evidence:** distinguish supplied facts, estimates, hypotheses, and unknowns. Cite actual sources; a log date is not attribution. Never invent a source, signer, signature, customer reaction, or acceptance. Keep outcomes **promised → measured → accepted** distinct, and implementation, verification, deployment, and customer acceptance separate. Missing evidence means unproven, not an observed failure.
10
+ - **Data boundary:** use only data permitted by the customer's AI policy; clarify unknown policy before loading their code or data. Never load `<private>` content into a model. Cross-client comparison and exporting reusable material require permission and removal of customer-identifying or confidential content; anonymization alone does not grant permission.
11
+
12
+ Apply the selected method to this context. Follow its linked supporting methods only when needed; do not restart discovery or repeat already answered questions.
@@ -34,10 +34,10 @@ Then classify blast radius:
34
34
  ```
35
35
  CRITICAL - if wrong, the engagement fails or the approach changes fundamentally
36
36
  → Must be validated before plan starts
37
-
37
+
38
38
  LOAD-BEARING - if wrong, significant rework or timeline change
39
39
  → Must be validated before build starts
40
-
40
+
41
41
  CONVENIENCE - if wrong, a task changes but the approach holds
42
42
  → Validate when you get there
43
43
  ```
@@ -1,10 +1,12 @@
1
1
  # three-options - Generate options
2
2
 
3
+ **Context:** apply [task context and evidence](task-context.md) before using the named records below.
4
+
3
5
  **Enter when:** a significant technical or strategic decision needs to be made, the FDE is asked "what should we do?", the team is stuck between approaches, or a fork in the engagement requires the sponsor's input.
4
6
 
5
7
  **Read first:** `reality.md`, `terrain.md`, `assumptions.md`, `success.md`, `context.md`. Load `business-case.md` if the decision has cost implications.
6
8
 
7
- One option is a request for trust. Two options is a false choice. Three options is a conversation between professionals. The FDE who presents three genuine options earns the decision-maker's respect - and their protection when things get hard.
9
+ Compare materially different, defensible alternatives. Three is a useful presentation shape when three viable paths exist; do not pad the set to meet a quota. Include keeping the current approach or deferring when those are credible choices.
8
10
 
9
11
  ## Method (you do this work)
10
12
 
@@ -12,26 +14,26 @@ One option is a request for trust. Two options is a false choice. Three options
12
14
 
13
15
  > "Decision: approach for the payment migration. Decided by: CTO. Needed by: Friday. Deferral cost: blocks the next sprint and delays the pilot by two weeks."
14
16
 
15
- **2. Generate three genuine options from what survived.** Read `assumptions.md` (CONFIRMED / DISPROVED) and `terrain.md` `## Parts`. The playbook is unavailable. Assemble from those blocks only.
17
+ **2. Generate genuine options from the evidence.** Use confirmed assumptions and known system parts; exclude disproved assumptions. If those records are absent, identify the supplied facts and unknowns. Use [test-assumptions](test-assumptions.md) only when a consequential assumption needs investigation.
16
18
 
17
- Not "good / medium / bad." Not three speeds of the same plan (Conservative / Pragmatic / Ambitious of one architecture). Three approaches that differ in **structure** - rearrange the same surviving parts.
19
+ Alternatives should differ materially in architecture, operating model, scope, cost, or reversibility. Do not label the same plan good / medium / bad or manufacture an unsafe option to favor your recommendation.
18
20
 
19
21
  Each option must be one the FDE would genuinely recommend under different circumstances. If you cannot defend an option, replace it - padding is visible.
20
22
 
21
23
  For each option, name:
22
24
  - which surviving blocks it is built from
23
- - which CONVENTION it refuses to obey
25
+ - which constraint or convention it changes, if any
24
26
  - its single biggest point of failure
25
27
  - any new building block, labelled as a new assumption (`UNKNOWN` in `assumptions.md`) - do not smuggle one in as a fact
26
28
 
27
- If all three are the same system at different risk levels, you wrote the playbook. Start over.
29
+ If only one viable path remains, explain what ruled out the alternatives and what evidence could reopen them.
28
30
 
29
31
  **3. Structure each option identically.** Same dimensions, same format - so comparison is instant:
30
32
 
31
33
  ```markdown
32
34
  ### Option A: <name>
33
35
  - **Blocks:** <which surviving assumptions / parts it is built from>
34
- - **Convention refused:** <the "we just…" it will not obey>
36
+ - **Constraint changed:** <constraint or convention changed, if any>
35
37
  - **What:** <the approach in one paragraph>
36
38
  - **Timeline:** <estimate with basis>
37
39
  - **Cost:** <effort, infrastructure, external>
@@ -40,19 +42,11 @@ If all three are the same system at different risk levels, you wrote the playboo
40
42
  - **Best when:** <the condition that makes this the right choice>
41
43
  ```
42
44
 
43
- **4. Add the comparison matrix.** Visually scannable:
44
-
45
- | Dimension | Option A | Option B | Option C |
46
- |-----------|----------|----------|----------|
47
- | Timeline | 6 weeks | 4 weeks | 3 weeks |
48
- | Risk | Low | Medium | High |
49
- | Reversibility | Easy rollback | Partial rollback | Difficult to reverse |
50
- | Team impact | Minimal | Moderate retraining | Significant ramp-up |
51
- | Long-term cost | Highest (tech debt) | Moderate | Lowest |
45
+ **4. Make comparison easy.** Use consistent dimensions: expected outcome, evidence, build and operating cost, time with estimate basis, reversibility, owner, and the most consequential uncertainty. A compact table helps when alternatives need comparison; do not fill it with invented numbers or label one path universally cheapest.
52
46
 
53
- **5. State your recommendation - and why.** The options are objective; the recommendation is your professional judgment:
47
+ **5. State your recommendation - and why.** Separate supplied facts and estimates from your judgment:
54
48
 
55
- > "I recommend Option B. The timeline pressure makes the conservative approach too slow, but the system fragility (see terrain.md hotspots) makes the ambitious approach reckier than the reward justifies. Option B gets us to pilot in 4 weeks with a tested rollback."
49
+ > "I recommend Option B. The limited team availability rules out a full rewrite this quarter, and the current hotspot makes keeping the job unchanged costly. Option B gets us to pilot in 4 weeks with a tested rollback."
56
50
 
57
51
  **6. Handle the override gracefully.** If the sponsor picks a different option:
58
52
 
@@ -66,9 +60,9 @@ If all three are the same system at different risk levels, you wrote the playboo
66
60
  ```markdown
67
61
  ## Decision: <name> - <date>
68
62
  Decided by: <who>
69
- Options presented: A (<name>), B (<name>), C (<name>)
63
+ Options presented: <viable alternatives>
70
64
  Recommended: B - <one line why>
71
- Chosen: <A/B/C> by <who>
65
+ Chosen: <option or pending> by <actual decision-maker or unknown>
72
66
  Trade-off accepted: <what the choice gives up>
73
67
  ```
74
68
 
@@ -76,22 +70,20 @@ The full option details in the same entry or linked to a section in `reality.md`
76
70
 
77
71
  ## Checkpoint
78
72
 
79
- Present the three options and the recommendation. One question to the FDE: "Which option matches what the sponsor can hear right now?" (A risk-averse sponsor after an incident → conservative. A founder pre-fundraise → ambitious.) If unsure: present all three and let the sponsor decide.
73
+ Present the viable alternatives, recommendation, and the evidence or constraint that would change it. Reuse known decision authority; if a decision remains pending, record it as pending.
80
74
 
81
75
  ## Worked example
82
76
 
83
- Acme: the reconciliation job needs to survive the FDE leaving. Priya asks "so what should we do?"
84
-
85
- Three real paths, not a strawman set. **Safe:** keep the job, add the rota and runbook - two weeks, no new failure modes, does nothing about the 47-commits/90d hotspot. **Pragmatic:** extract the settlement-matching step behind a tested interface - six weeks, retires the untested hotspot, needs Raj's time and he currently opposes it. **Aggressive:** rewrite the service - a quarter, fixes everything, and the same team already abandoned this once.
77
+ Fictional example: a support team needs completed requests written back to its service system. Its product can export a file today. A supported connector is expected in six weeks; the customer wants automation in two. A custom API adapter looks feasible, but nobody has accepted its maintenance.
86
78
 
87
- Same dimensions on each, so comparison is instant, and every cost carries a source: the six-week figure is churn-based, not felt.
79
+ Two defensible paths remain. Continue the approved export while checking the supported connector's fit, or investigate a bounded adapter whose delivery depends on a named owner and tested API behavior. The export is a bridge within the first path, not a third option invented for the slide. Compare manual effort, engineering effort, ongoing support, and the effect of waiting. Time released is capacity unless spending actually falls.
88
80
 
89
- Recommendation: pragmatic, conditional - *if* Raj is on the design, otherwise safe, because the aggressive path failed here before for exactly the reason it would fail again. `decisions.md` records the decision, who chose it, and the condition, so week 10's "why aren't we rewriting it" has an answer with a date on it.
81
+ Recommend the bridge while resolving native fit and the value of earlier automation. Reconsider the adapter if that value justifies full costs and an owner accepts it. `decisions.md` (or a standalone decision note) records the recommendation, evidence, unknowns, and pending decision. It does not claim that a sponsor chose it or that the adapter can meet the date.
90
82
 
91
83
  ## Principles
92
84
 
93
- - Three options, never one. One option is a request for trust; three is a real decision.
94
- - Assemble from surviving blocks. Three speeds of the same plan is the playbook - start over.
85
+ - Compare defensible alternatives; the number follows the evidence.
86
+ - Build from known facts and label new assumptions. Material differences make the comparison useful.
95
87
  - Each option must be genuinely defensible - no straw men.
96
88
  - Same structure for each option. Comparison should take 30 seconds.
97
89
  - Recommend one. State why. Accept the override gracefully.
@@ -0,0 +1,31 @@
1
+ # verification - Make a claim replayable
2
+
3
+ **Enter when:** reporting completion, evaluating an acceptance check, handing work to a reviewer, or preparing a release.
4
+
5
+ Use [task context](task-context.md). This method returns evidence directly or writes an existing permitted task/engagement record; it never requires `.fde/` initialization.
6
+
7
+ ## Method
8
+
9
+ 1. Translate each claim into the observation that would support or reject it. Reuse agreed acceptance criteria and required repository checks. Select focused checks for changed behavior before broadening to release requirements.
10
+ 2. Identify the actual repository commands, fixtures, runtime, and environment. Read command behavior before executing it, especially when it can write externally. Use authorized environments and avoid leaking secrets through logs or diagnostic commands.
11
+ 3. Run the checks and inspect results, including exit status and relevant output. A running job, test discovery, a mocked response, and a successful real request are different evidence. Record asynchronous completion before claiming success.
12
+ 4. Bind evidence to the tested revision and working tree. For uncommitted changes record the base revision plus changed paths and an available diff digest or snapshot identifier. For browser/manual checks record the steps, inputs, observed result, and inspected evidence.
13
+ 5. After a change, rerun checks whose behavior or assumptions were affected. Reuse prior evidence only when the relevant code, dependencies, data, and environment remain applicable; cite the original run and reason. Never imply reused evidence was rerun.
14
+ 6. Label every required check **passed**, **failed**, **blocked**, or **not run**. Include why blocked/not run, impact, and next step. Missing evidence is unproven; it is not an observed failure or a pass.
15
+
16
+ ## Receipt
17
+
18
+ Use one compact entry per check or a table with these fields:
19
+
20
+ - Claim / acceptance check and expected result.
21
+ - Exact command and working directory, or manual journey and inputs.
22
+ - Environment, runtime/tool versions when relevant, and fixture/data source.
23
+ - Revision plus working-tree identity; run date/time.
24
+ - Observed result and exit status where available; safe evidence location.
25
+ - Status, limitations, unrun checks, and next step.
26
+
27
+ Keep implementation, verification, deployment, measured outcome, and customer acceptance distinct. A local pass supports the tested local behavior. An acceptance claim needs an attributed source from the agreed decision-maker or agreed acceptance mechanism. Record no raw `<private>` blocks, credentials, or hidden reasoning.
28
+
29
+ ## Acceptance
30
+
31
+ A completion statement cites applicable evidence for its claims and explicitly names material gaps. If required checks fail, investigate or report the blocker; never skip them, edit expectations, or relabel the scope without authority to obtain a green result.
@@ -0,0 +1,21 @@
1
+ ---
2
+ name: fde-build
3
+ description: Implement a scoped customer-facing software change in the existing repository and verify its behavior. Use for delivery work with an understood outcome, not incident response.
4
+ ---
5
+
6
+ # fde-build
7
+
8
+ <!-- Generated by bin/generate-skills.js; edit the canonical references and catalog. -->
9
+
10
+ ## Purpose
11
+
12
+ Implement a scoped customer-facing software change in the existing repository and verify its behavior. Use for delivery work with an understood outcome, not incident response.
13
+
14
+ Read [the task context contract](references/task-context.md), then [the method](references/build.md). Load further references only when the task needs them. Everything linked is included in this skill; no other skill pack is required.
15
+
16
+ ## Principles
17
+
18
+ - Work directly from the supplied permitted context. Standalone work does not require an engagement folder or initialization. Record filenames in the method are optional persistence destinations when no engagement is bound.
19
+ - If called by @fde, reuse its current sanitized packet and scope. Do not restart setup, discovery or questions already answered.
20
+ - The task context contract controls persistence and authority in both modes. Preserve unknowns and distinguish implementation, verification, deployment and acceptance.
21
+ - Use the customer's repository instructions and available tools. Report a missing capability or unrun check honestly; do not claim that installing a skill provisions infrastructure.