fdeops 4.0.4 → 4.1.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (154) hide show
  1. package/AGENTS.md +1 -1
  2. package/README.md +33 -6
  3. package/bin/check.js +14 -23
  4. package/bin/generate-skills.js +144 -0
  5. package/bin/install.js +2 -1
  6. package/bin/skill-catalog.js +17 -0
  7. package/mcp/fdeops-ingest/package.json +2 -2
  8. package/package.json +4 -3
  9. package/plugin.json +2 -2
  10. package/skills/fde/SKILL.md +18 -10
  11. package/skills/fde/references/board-memo.md +1 -1
  12. package/skills/fde/references/build.md +20 -0
  13. package/skills/fde/references/business-case.md +9 -7
  14. package/skills/fde/references/close.md +6 -4
  15. package/skills/fde/references/debug.md +18 -0
  16. package/skills/fde/references/encode-pattern.md +10 -8
  17. package/skills/fde/references/eval-pack.md +16 -33
  18. package/skills/fde/references/hold-scope.md +9 -7
  19. package/skills/fde/references/integrate.md +18 -0
  20. package/skills/fde/references/plan.md +4 -2
  21. package/skills/fde/references/poc.md +5 -3
  22. package/skills/fde/references/qa.md +18 -0
  23. package/skills/fde/references/readout.md +6 -4
  24. package/skills/fde/references/review.md +22 -60
  25. package/skills/fde/references/ship.md +42 -287
  26. package/skills/fde/references/task-context.md +12 -0
  27. package/skills/fde/references/test-assumptions.md +2 -2
  28. package/skills/fde/references/three-options.md +19 -27
  29. package/skills/fde/references/verification.md +31 -0
  30. package/skills/fde-build/.fde-generated.json +16 -0
  31. package/skills/fde-build/SKILL.md +21 -0
  32. package/skills/fde-build/references/build.md +20 -0
  33. package/skills/fde-build/references/debug.md +18 -0
  34. package/skills/fde-build/references/eval-pack.md +26 -0
  35. package/skills/fde-build/references/integrate.md +18 -0
  36. package/skills/fde-build/references/qa.md +18 -0
  37. package/skills/fde-build/references/review.md +39 -0
  38. package/skills/fde-build/references/ship.md +73 -0
  39. package/skills/fde-build/references/task-context.md +12 -0
  40. package/skills/fde-build/references/verification.md +31 -0
  41. package/skills/fde-debug/.fde-generated.json +16 -0
  42. package/skills/fde-debug/SKILL.md +21 -0
  43. package/skills/fde-debug/references/build.md +20 -0
  44. package/skills/fde-debug/references/debug.md +18 -0
  45. package/skills/fde-debug/references/eval-pack.md +26 -0
  46. package/skills/fde-debug/references/integrate.md +18 -0
  47. package/skills/fde-debug/references/qa.md +18 -0
  48. package/skills/fde-debug/references/review.md +39 -0
  49. package/skills/fde-debug/references/ship.md +73 -0
  50. package/skills/fde-debug/references/task-context.md +12 -0
  51. package/skills/fde-debug/references/verification.md +31 -0
  52. package/skills/fde-discover/.fde-generated.json +10 -0
  53. package/skills/fde-discover/SKILL.md +21 -0
  54. package/skills/fde-discover/references/audit.md +71 -0
  55. package/skills/fde-discover/references/discover.md +254 -0
  56. package/skills/fde-discover/references/task-context.md +12 -0
  57. package/skills/fde-evaluate/.fde-generated.json +16 -0
  58. package/skills/fde-evaluate/SKILL.md +21 -0
  59. package/skills/fde-evaluate/references/build.md +20 -0
  60. package/skills/fde-evaluate/references/debug.md +18 -0
  61. package/skills/fde-evaluate/references/eval-pack.md +26 -0
  62. package/skills/fde-evaluate/references/integrate.md +18 -0
  63. package/skills/fde-evaluate/references/qa.md +18 -0
  64. package/skills/fde-evaluate/references/review.md +39 -0
  65. package/skills/fde-evaluate/references/ship.md +73 -0
  66. package/skills/fde-evaluate/references/task-context.md +12 -0
  67. package/skills/fde-evaluate/references/verification.md +31 -0
  68. package/skills/fde-feedback/.fde-generated.json +9 -0
  69. package/skills/fde-feedback/SKILL.md +21 -0
  70. package/skills/fde-feedback/references/encode-pattern.md +96 -0
  71. package/skills/fde-feedback/references/task-context.md +12 -0
  72. package/skills/fde-handoff/.fde-generated.json +10 -0
  73. package/skills/fde-handoff/SKILL.md +21 -0
  74. package/skills/fde-handoff/references/close.md +66 -0
  75. package/skills/fde-handoff/references/encode-pattern.md +96 -0
  76. package/skills/fde-handoff/references/task-context.md +12 -0
  77. package/skills/fde-integrate/.fde-generated.json +16 -0
  78. package/skills/fde-integrate/SKILL.md +21 -0
  79. package/skills/fde-integrate/references/build.md +20 -0
  80. package/skills/fde-integrate/references/debug.md +18 -0
  81. package/skills/fde-integrate/references/eval-pack.md +26 -0
  82. package/skills/fde-integrate/references/integrate.md +18 -0
  83. package/skills/fde-integrate/references/qa.md +18 -0
  84. package/skills/fde-integrate/references/review.md +39 -0
  85. package/skills/fde-integrate/references/ship.md +73 -0
  86. package/skills/fde-integrate/references/task-context.md +12 -0
  87. package/skills/fde-integrate/references/verification.md +31 -0
  88. package/skills/fde-options/.fde-generated.json +11 -0
  89. package/skills/fde-options/SKILL.md +21 -0
  90. package/skills/fde-options/references/business-case.md +90 -0
  91. package/skills/fde-options/references/task-context.md +12 -0
  92. package/skills/fde-options/references/test-assumptions.md +102 -0
  93. package/skills/fde-options/references/three-options.md +90 -0
  94. package/skills/fde-poc/.fde-generated.json +23 -0
  95. package/skills/fde-poc/SKILL.md +21 -0
  96. package/skills/fde-poc/references/audit.md +71 -0
  97. package/skills/fde-poc/references/build.md +20 -0
  98. package/skills/fde-poc/references/business-case.md +90 -0
  99. package/skills/fde-poc/references/debug.md +18 -0
  100. package/skills/fde-poc/references/discover.md +254 -0
  101. package/skills/fde-poc/references/eval-pack.md +26 -0
  102. package/skills/fde-poc/references/integrate.md +18 -0
  103. package/skills/fde-poc/references/plan.md +167 -0
  104. package/skills/fde-poc/references/poc.md +55 -0
  105. package/skills/fde-poc/references/qa.md +18 -0
  106. package/skills/fde-poc/references/review.md +39 -0
  107. package/skills/fde-poc/references/ship.md +73 -0
  108. package/skills/fde-poc/references/task-context.md +12 -0
  109. package/skills/fde-poc/references/test-assumptions.md +102 -0
  110. package/skills/fde-poc/references/three-options.md +90 -0
  111. package/skills/fde-poc/references/verification.md +31 -0
  112. package/skills/fde-qa/.fde-generated.json +16 -0
  113. package/skills/fde-qa/SKILL.md +21 -0
  114. package/skills/fde-qa/references/build.md +20 -0
  115. package/skills/fde-qa/references/debug.md +18 -0
  116. package/skills/fde-qa/references/eval-pack.md +26 -0
  117. package/skills/fde-qa/references/integrate.md +18 -0
  118. package/skills/fde-qa/references/qa.md +18 -0
  119. package/skills/fde-qa/references/review.md +39 -0
  120. package/skills/fde-qa/references/ship.md +73 -0
  121. package/skills/fde-qa/references/task-context.md +12 -0
  122. package/skills/fde-qa/references/verification.md +31 -0
  123. package/skills/fde-readout/.fde-generated.json +11 -0
  124. package/skills/fde-readout/SKILL.md +21 -0
  125. package/skills/fde-readout/references/board-memo.md +108 -0
  126. package/skills/fde-readout/references/business-case.md +90 -0
  127. package/skills/fde-readout/references/readout.md +71 -0
  128. package/skills/fde-readout/references/task-context.md +12 -0
  129. package/skills/fde-review/.fde-generated.json +16 -0
  130. package/skills/fde-review/SKILL.md +21 -0
  131. package/skills/fde-review/references/build.md +20 -0
  132. package/skills/fde-review/references/debug.md +18 -0
  133. package/skills/fde-review/references/eval-pack.md +26 -0
  134. package/skills/fde-review/references/integrate.md +18 -0
  135. package/skills/fde-review/references/qa.md +18 -0
  136. package/skills/fde-review/references/review.md +39 -0
  137. package/skills/fde-review/references/ship.md +73 -0
  138. package/skills/fde-review/references/task-context.md +12 -0
  139. package/skills/fde-review/references/verification.md +31 -0
  140. package/skills/fde-scope/.fde-generated.json +9 -0
  141. package/skills/fde-scope/SKILL.md +21 -0
  142. package/skills/fde-scope/references/hold-scope.md +83 -0
  143. package/skills/fde-scope/references/task-context.md +12 -0
  144. package/skills/fde-ship/.fde-generated.json +16 -0
  145. package/skills/fde-ship/SKILL.md +21 -0
  146. package/skills/fde-ship/references/build.md +20 -0
  147. package/skills/fde-ship/references/debug.md +18 -0
  148. package/skills/fde-ship/references/eval-pack.md +26 -0
  149. package/skills/fde-ship/references/integrate.md +18 -0
  150. package/skills/fde-ship/references/qa.md +18 -0
  151. package/skills/fde-ship/references/review.md +39 -0
  152. package/skills/fde-ship/references/ship.md +73 -0
  153. package/skills/fde-ship/references/task-context.md +12 -0
  154. package/skills/fde-ship/references/verification.md +31 -0
@@ -0,0 +1,167 @@
1
+ # plan - Sequence the work
2
+
3
+ **Context:** apply [task context and evidence](task-context.md). Standalone planning evaluates supplied facts directly; it does not require initialized engagement records.
4
+
5
+ **Enter when:** scope is understood and the work needs breaking down - a slice, a phase, or the whole delivery.
6
+
7
+ **Read first:** `reality.md`, `success.md`, `terrain.md`, `stakeholders.md`. Load `business-case.md` if poc produced one. Not the full folder.
8
+
9
+ **On an initialized engagement, before a new delivery plan or material scope change:** run `fde doctor --ready`. For standalone planning, check the supplied outcome, scope, acceptance and authority directly; do not initialize records to run this validator. Missing binary success or a named customer-side signer blocks progression: review the proposed acceptance check and authority with the FDE first. Use a test/input and observable pass/fail under **Done when:** or **Acceptance check:**. A number, role, or successful demo alone is insufficient. Do not invent missing facts to pass lint. Routine reversible fixes within confirmed scope reuse the existing signer, acceptance criteria, and engineering plan; record verification without reopening settled decisions.
10
+
11
+ ## Validation gate (confirm understanding, clarify where it elevates)
12
+
13
+ Before planning, state what you're working from in 2-3 lines:
14
+
15
+ > "Planning against: [success definition from success.md]. Scope boundary: [out-of-scope items]. Reality check: [brief aligns with reality.md / or note the delta]."
16
+
17
+ Then check - probe ONLY if it prevents a bad plan:
18
+
19
+ 1. **Success is measurable.** If "done" is vague ("make it better") → rephrase it: "I'm reading success as: [specific measurable outcome]. That the target?"
20
+ 2. **Reality matches the brief.** If discovery contradicted the brief → name it: "Discovery found [X] but the brief says [Y]. Planning against reality unless you say otherwise."
21
+ 3. **Out-of-scope exists.** If missing → one line: "Nothing's marked out-of-scope yet. That means every new request is implicitly in. Worth defining now or after the first plan draft?"
22
+
23
+ State your read, let the FDE correct, then plan.
24
+
25
+ An FDE plan is not a sprint backlog. The technical sequence is the easy part. The hard part is when to show progress, who approves the next phase, and where trust is thin enough that two silent weeks read as failure. A technically correct plan that ignores engagement politics fails on schedule.
26
+
27
+ ## Method (you do this work)
28
+
29
+ **0. Lock scope first.** Read `success.md`, `assumptions.md`, and the **Question** on `reality.md`. Make the boundary explicit using the supplied request; ask if an ambiguity changes the commitment. Investigate a critical open assumption before planning work that depends on it. If the problem itself is unclear, use discovery for that gap; absent filenames do not block a plan supported by supplied facts.
30
+
31
+ **Reuse check.** Before sequencing a build, compare the requested solution with the smallest existing capability or operating change that could satisfy the same acceptance test. Cite the relevant repo/config/workaround evidence. Record why reuse is sufficient or insufficient in `decisions.md`; include “no new code” when supported. A request for AI does not establish that a model is needed. If a host engineering pack already has an approved implementation plan, reference it from `decisions.md`; do not generate a parallel user-story backlog.
32
+
33
+ **1. Work backwards from success.** What's the last thing that must be true before done? And before that? That's the dependency chain - not a wish list.
34
+
35
+ **2. Front-load the fragile.** Check `terrain.md` hotspots. Risky modules go early - fail fast, not in week three.
36
+
37
+ **3. One user action per change.** Each task delivers something visible and testable ("user submits form, sees it saved"), never a layer ("build the database layer"). See `ship`.
38
+
39
+ **4. Size to a coherent, verifiable outcome.** Split unrelated work and tasks too complex to review or recover safely. Use bounded review sections for large cohesive changes; elapsed time and line count are signals to examine, not universal limits.
40
+
41
+ **5. AI components get explicit eval tasks.** "Output validated on 50 real production examples," "fallback tested under model unavailability," "inputs/outputs logging to <destination>" - these are pre-conditions of shipping, in the plan before build starts.
42
+
43
+ **6. Stakeholder touchpoints every 2-3 tasks.** "Show progress to <name from stakeholders.md>." Not ceremony: a customer who sees small wins stays bought in; silence gets filled with doubt.
44
+
45
+ **7. End with a kill list.** Every plan names what you will **not** do this phase. If everything is "later," you have no plan - you have a wish list. Cap **Now** at 3 PRs (same discipline as pick-three).
46
+
47
+ **Acceptance criteria gate:** no task moves to build without written happy-path AND unhappy-path criteria. Can't write them = the task isn't understood; the open question goes to the customer **before** the task starts. Vague criteria surface later as scope creep and rework.
48
+
49
+ ## Artifact
50
+
51
+ The plan goes to **`decisions.md`** - always. Build reads the plan from `decisions.md`; anywhere else and the build starts blind.
52
+
53
+ A plan is **not done** until all four blocks exist:
54
+
55
+ ```markdown
56
+ ## Plan - <date>
57
+ ### Now (max 3)
58
+ Task N: <outcome, not activity>
59
+ Delivers: <what someone can see/test>
60
+ Accepts: <happy path> / <unhappy path>
61
+ Touches: <files/systems - blast radius declared upfront>
62
+ Risk: <what could go wrong + fallback>
63
+ Kill if: <the observation that voids this slice - copy from assumptions.md How we test, or the check that means stop>
64
+ Verify: <specific check>
65
+ Value promised: <business unit change this slice claims>
66
+ Baseline: <value + source/date/window/environment, or pending + measurement owner>
67
+ Acceptance owner: <name + authority source, or unknown - ask: who can accept?>
68
+ Evidence to collect: <before/after check, sample/window, environment, and receipt location>
69
+ Reuse: <existing capability used, or evidence it cannot satisfy the criteria>
70
+
71
+ ### Next
72
+ - ...
73
+
74
+ ### Later
75
+ - ...
76
+
77
+ ### Kill list (explicitly not this phase)
78
+ | Item | Why killed / deferred | Who accepted |
79
+ |------|----------------------|--------------|
80
+ | <rewrite / nice-to-have / political ask> | <evidence> | <name, date> |
81
+ ```
82
+
83
+ In `Who accepted`, distinguish a proposed deferral from an agreement: use `pending` until a named person accepted this scope with a dated source. Sponsorship alone is not approval of every plan detail.
84
+
85
+ No kill list → not a finished plan. Reopen with the FDE until the deferrals are written.
86
+ ## Checkpoint
87
+
88
+ Walk the FDE through: sequence + why this order, where the fragile work sits, where the touchpoints land, the acceptance gate and **Kill if** on task 1, and the kill list. One question: "Which stakeholder sees the first visible slice, and when?" Second: "Who accepted what we are not doing?" Third: "What observation stops task 1 this week?"
89
+
90
+ ## Method - estimation (when the sponsor asks "how long, how much?")
91
+
92
+ Every FDE gets asked this in week one. The honest answer is a range, not a number. A single-point estimate is a promise; a range is a professional assessment.
93
+
94
+ **The 3-point method:**
95
+ 1. **Best case** - everything goes right, no surprises, team has capacity. This is what the sponsor wants to hear.
96
+ 2. **Expected case** - normal friction: one discovery changes the plan, one integration takes longer, one approval cycle stalls. This is what to plan against.
97
+ 3. **Worst case** - a major unknown surfaces, a dependency fails, a key person is unavailable. This is what to protect against.
98
+
99
+ **Present as:** "2-4 weeks expected, could stretch to 6 if [named risk]." Never give one number.
100
+
101
+ **The sizing table:**
102
+
103
+ | Slice | Complexity | Dependencies | Unknowns | Estimate (expected) |
104
+ |-------|-----------|--------------|----------|---------------------|
105
+ | _per vertical slice from the plan_ | Low/Med/High | Named | Named | X days/weeks |
106
+
107
+ **Rules:**
108
+ - Estimate in weeks for engagements > 1 month. Days for < 1 month.
109
+ - Add 30% buffer for integration work (it always takes longer).
110
+ - Add 50% buffer for AI/ML work (eval cycles are unpredictable).
111
+ - Name assumptions explicitly: "assumes API docs are accurate", "assumes staging environment exists."
112
+ - Each named assumption needs a **kill observation**: the result that voids the estimate. Copy it from `assumptions.md` → How we test. No kill observation = it is not an assumption, it is hope.
113
+ - Revisit estimates every 2 weeks. An estimate that never updates is fiction.
114
+
115
+ Write estimates to `decisions.md` under `## Sizing`. Include the assumptions - when they break, the estimate changes and the FDE has evidence for the conversation.
116
+
117
+ ## Method - migration strategy (when the engagement is "move from X to Y")
118
+
119
+ Migrations are the most common enterprise FDE engagement. The strategy precedes the plan:
120
+
121
+ **Step 1: Classify the migration type.**
122
+
123
+ | Type | What it means | Risk profile |
124
+ |------|---------------|-------------|
125
+ | **Rehost** (lift-and-shift) | Same code, different infrastructure | Low code risk, high ops risk |
126
+ | **Replatform** | Minor code changes to use new platform features | Medium risk, clear scope |
127
+ | **Refactor** | Rewrite components to fit the new architecture | High risk, scope creep magnet |
128
+ | **Replace** | Buy/build new, retire old | Highest risk, requires parallel running |
129
+ | **Retire** | Turn off, nobody uses it | Politically hard, technically easy |
130
+
131
+ **Step 2: Map the dependency graph.** What calls what. What breaks if this moves first. The migration order is the reverse of the dependency chain - leaf nodes first, core last.
132
+
133
+ **Step 3: Define the cutover strategy.**
134
+ - **Big bang** - everything moves at once. Fast but catastrophic on failure. Only for small systems.
135
+ - **Strangler fig** - new traffic to new system, old traffic drains. Safe but slow. Preferred for anything load-bearing.
136
+ - **Parallel run** - both systems run, outputs compared. Expensive but safest for data-critical systems.
137
+
138
+ **Step 4: Write the rollback before the migration starts.** "If we move service X and it fails, we route back to old within [time]." No rollback = no migration.
139
+
140
+ **Step 5: Define success metrics per phase.** Not "migration complete" - that's a project plan. "Error rate same or lower, latency within 10%, zero data loss, team can operate without FDE." Measurable, per service.
141
+
142
+ Write migration strategy to `decisions.md` under `## Migration`. Each service gets a row: type, order, cutover method, rollback, success metric.
143
+
144
+ ## When the plan changes mid-engagement
145
+
146
+ Never quietly update tasks. Name the reset: update `reality.md` and `success.md`, one paragraph in `decisions.md` - what changed, why, new sequence. An undocumented reset looks like drift; a documented one looks like the FDE caught something important.
147
+
148
+ ## Worked example
149
+
150
+ Acme, after discover: the reconciliation job is unowned, Marco's spreadsheet is the real fallback.
151
+
152
+ **Now** is three tasks, not eight. Task 1 is *failures reach a named human* - delivers a page to a rota, accepts "kill the job mid-run → the on-call is paged within 15 min", touches the job wrapper and the alert config, rollback is re-disable the route, **Kill if:** a real failure page is acked by nobody on the rota (the *finance would act* assumption, DISPROVED if Marco is the only name that answers), verify by killing it in staging. Value promised: `risk-mitigation - a silent failure becomes a 15-minute one`.
153
+
154
+ The kill list in `decisions.md` is where the plan earns its keep: the rewrite of the reconciliation service that Tom keeps proposing goes there - *deferred, the failure mode is ownership not architecture (Priya accepted, Jun 12)* - along with the finance dashboard finance asked for directly. Both stay visible so the same argument is not re-litigated in week 4 without a receipt.
155
+
156
+ First visible slice goes to Marco, not Priya: he is the one whose morning changes, and his confirmation is what makes the sponsor update true.
157
+
158
+ ## Principles
159
+
160
+ - Plan from success backwards, not from today forwards.
161
+ - Fragile zones early. Fail fast.
162
+ - Every 2-3 tasks, a stakeholder touchpoint. Trust decays without visibility.
163
+ - No written acceptance criteria, no build.
164
+ - No kill list, no finished plan.
165
+ - No **Kill if** on a Now PR, that PR is hope.
166
+ - Estimates are ranges, not promises. Name the assumptions and the observation that voids them.
167
+ - Migrations: leaf nodes first, core last. Rollback before cutover.
@@ -0,0 +1,55 @@
1
+ # poc - Validate the solution
2
+
3
+ **Context:** apply [task context and evidence](task-context.md) before using the named records below.
4
+
5
+ **Enter when:** a direction needs validating before committing real build time - POC, spike, show something, de-risk, pick between use cases. The output is something a sponsor can reject in a room this week, not a polished product.
6
+
7
+ **Read first:** `context.md`, `reality.md`. Load `terrain.md` only if the prototype touches the existing codebase. If `terrain.md` **Data estate** has a Blocker source this prototype needs, stop - that is discover, not a day's demo.
8
+
9
+ A green check on synthetic data is not a validated solution. The person who can say no has to see it on evidence they already believe.
10
+
11
+ ## Method (you do this work)
12
+
13
+ **0. Name the killer assumption.** With the FDE: "What's the belief that kills the project if it's wrong?" Prototype **that** - not the pretty demo. If [three-options](three-options.md) just ran: the cheapest test is for the recommended option first, unless they pick another.
14
+
15
+ **0b. Pass / fail before you build.** For the test you will run, write three lines in `prototype-log.md` first: what you will actually do (who you talk to, what you show, on whose screen); the result that **kills** this option; the result that keeps it alive. What you would learn either way. If every option's test would fail, name which `assumptions.md` block to reopen - do not invent a fourth playbook.
16
+
17
+ **1. Pick by score when several use cases compete.** Use the scoring model from `discover.md` - (Value × Data readiness) / Complexity. If discover or score-use-cases already produced a ranking, reuse it; never invent a third ranking.
18
+
19
+ **2. Build the minimum that tests the assumption.** Timebox the experiment with the FDE; aim for a same-day result when access and evidence permit it. Skip cosmetic polish, but keep the input validation, access controls, and failure handling needed to protect the test environment and data. Label shortcuts and simulated inputs. The POC is done when the person who can say no has seen the evidence and reacted, not when the code looks finished.
20
+
21
+ **3. AI directions - test these before anything else:**
22
+ - Data: available, clean, sufficient volume? Synthetic data can test mechanics, but does not establish production quality or real-world coverage.
23
+ - Environment: are external model calls even allowed here?
24
+ - Latency: acceptable against real user expectations, not ideal conditions?
25
+ - Is AI the right tool at all - or is this a data-quality or process problem wearing an AI costume?
26
+
27
+ **4. Decide at the agreed checkpoint.** Stop when the predeclared failure criterion is met or continued testing is unsafe. If feedback is absent or inconclusive, distinguish access or stakeholder availability from evidence against the assumption. Reassess the hypothesis, test design, and remaining timebox; extend only with a clear learning question and authorization for added scope or cost. Do not kill or continue solely because an iteration count was reached. If the customer cannot explain or trust high-stakes AI output, name the unresolved requirement and test whether it can be met. Record proceed / pivot / stop / inconclusive with evidence and what was learned.
28
+
29
+ **5. Translate to business language** once validated: problem solved, cost of inaction, success in numbers, 2-3 trade-offs. Three sentences max for the stakeholder - can't say it in three, don't understand it yet.
30
+
31
+ ## If proceeding to production
32
+
33
+ Carry the hypothesis, test evidence, customer reaction, and remaining unknowns into the existing [plan](plan.md) and [ship](ship.md) workflow. A working demo does not establish production readiness or customer acceptance.
34
+
35
+ Inspect the prototype before deciding what to reuse. Keep components whose behavior and boundaries are suitable and tested. Replace or harden shortcuts that fail production requirements; rewrite only where the evidence justifies it. Record the decision and remaining work in `decisions.md`, rather than treating all prototype code as disposable or all working code as ready to deploy.
36
+
37
+ Production work includes the actual data path, permissions, failure recovery, realistic load, observability, ownership, and required AI evaluations. Use the existing ship gates for those checks.
38
+
39
+ ## Artifact
40
+
41
+ **`prototype-log.md`** - what was built, shown, the actual reaction, what was learned (including kills - a killed prototype that saved three weeks is a win worth recording).
42
+
43
+ **`business-case.md`** - scored use case, cost of inaction, success metrics, trade-offs, the 3-sentence pitch. `plan` builds around this file.
44
+
45
+ ## Checkpoint
46
+
47
+ Tell the FDE: did the riskiest assumption hold · what the customer's reaction actually revealed · proceed / pivot / kill · the 3-sentence case if proceeding. The pitch is written for the person who can say yes or no.
48
+
49
+ ## Principles
50
+
51
+ - Optimize for a bounded learning outcome. Agree a timebox and revisit scope if access or evidence blocks it; never skip necessary safeguards to meet an arbitrary duration.
52
+ - Write pass/fail before you build. A demo with no kill line is a show.
53
+ - Show it rough. Polish misleads.
54
+ - Prototype the killer assumption, not the demo.
55
+ - Kill fast; log the learning.
@@ -0,0 +1,18 @@
1
+ # qa - Exercise the changed journey
2
+
3
+ **Enter when:** a feature or fix needs behavioral verification through its real interface, especially UI, API, and multi-step workflows.
4
+
5
+ Use [task context](task-context.md) and the customer's existing browser, API, fixtures, and test tooling. `.fde/` is not a prerequisite. Respect the permitted environment and authority for every side effect.
6
+
7
+ ## Method
8
+
9
+ 1. Identify the changed journey, user roles, acceptance checks, and risk-bearing neighboring paths. Record the revision and environment. Use synthetic or sanitized fixtures with understood cleanup; do not borrow production data without permission.
10
+ 2. Run the normal journey from its real entry point through the expected result. Verify persisted or downstream state when the requirement includes it; a success toast alone does not prove a write succeeded.
11
+ 3. Select negative and boundary cases from the change: invalid input, empty/loading/error states, refresh/back navigation, retries, duplicates, permissions, or interrupted work. For UI changes, inspect relevant viewport sizes, keyboard access, focus, labels, and errors. Use a real browser for the affected journey.
12
+ 4. Inspect relevant console and network evidence. Distinguish a UI defect from a failed API or unavailable environment. Retain only privacy-safe screenshots and logs. Do not claim visual verification from code inspection or a generated screenshot that was not viewed.
13
+ 5. Report failures with steps, expected/actual result, revision/environment, evidence, and impact. If repair is authorized, use [debug](debug.md), then rerun the failed journey and affected neighbors. Keep unrelated findings separate from the change.
14
+ 6. Produce a [verification receipt](verification.md). State which roles, devices, environments, or data conditions remain untested. Do not weaken acceptance checks to make the run pass.
15
+
16
+ ## Deliverable and acceptance
17
+
18
+ Return checked journeys and observed results, reproducible defects, limitations, and remaining blockers. Done means the agreed behavioral checks passed under the stated conditions. A browser smoke test does not establish load capacity, security assurance, accessibility conformance, deployment, or customer acceptance by itself. When coordinated, append the evidence to the existing delivery record; standalone QA can return it directly.
@@ -0,0 +1,39 @@
1
+ # review - Assess the actual change
2
+
3
+ **Enter when:** a diff, proposed merge, or review comment needs assessment against agreed behavior and constraints.
4
+
5
+ Use [task context](task-context.md). Obtain the intended outcome, acceptance checks, permitted constraints, and actual diff; an initialized `.fde/` is unnecessary. Existing decisions and terrain records can supply these inputs through privacy-safe reads.
6
+
7
+ ## Establish what was reviewed
8
+
9
+ Identify the repository, base and head revision, staged/unstaged changes, and relevant untracked files. Read applicable instructions and the full in-scope diff, then inspect callers and tests where needed. A committed-range diff alone omits working-tree edits. Record missing files or unavailable context as limitations.
10
+
11
+ State the review source: **self-check** when the author inspects their own work; **independent review** only when a separate person or agent actually examines it. A second pass by the same agent is still a self-check. Name the actual reviewer/source and reviewed revision when available. Do not fabricate a reviewer, dialogue, approval, or clean verdict. Use an available separate reviewer for substantial or risky changes when authorized; otherwise report the missing independent review and continue useful self-checks.
12
+
13
+ ## Check scope, then behavior
14
+
15
+ Compare each logical change with the agreed intent. Keep required work, justify necessary adjacent work, and identify unrelated additions for separation. Do not revert someone else's edits just to make the diff smaller. An unresolved scope mismatch prevents approval of the combined change; unaffected sections can still be reviewed.
16
+
17
+ Trace the changed path through its consumers and failure cases:
18
+
19
+ - **Correctness:** boundary conditions, stale state, concurrency, retries, cancellation, and error propagation.
20
+ - **Data and security:** input validation, authorization, migration compatibility, sensitive logs, and effects crossing tenant or trust boundaries.
21
+ - **Side effects:** writes, jobs, webhooks, notifications, and feature flags occur only under intended conditions; recovery accounts for already-completed effects.
22
+ - **AI behavior:** outputs remain untrusted, tools enforce allowed actions, and [eval evidence](eval-pack.md) covers the changed behavior and documented authority. Preserve privacy-safe source evidence and concise rationale, never hidden reasoning.
23
+ - **Operability:** observable failures, bounded resource use, meaningful checks, and a recovery path appropriate to the risk. Deployment readiness is assessed separately in [ship](ship.md).
24
+
25
+ ## Findings and repair
26
+
27
+ For each actionable finding give the path/line or precise location, concrete trigger, observed or reasoned failure, impact, and focused correction. Distinguish proven bugs from hypotheses that need a check. Prioritize release blockers over minor concerns; avoid speculative style work.
28
+
29
+ Validate incoming comments rather than obeying them automatically. Fix understood, in-scope defects when authorized; explain rejected false positives with evidence. Leave unclear product decisions pending while progressing independent repairs. Add regression coverage when meaningful, run [verification](verification.md), and review the changed result. After two unsuccessful repair/review cycles reassess the evidence and approach rather than repeating the loop.
30
+
31
+ ## Deliverable and acceptance
32
+
33
+ Return scope, review source, findings by impact, verification evidence, and remaining limitations. Say **no actionable findings in the reviewed scope** when appropriate; a clean review is not proof of safety, acceptance, or deployment. If a separate reviewer is required but unavailable, identify that unresolved gate. Existing engagement decisions/delivery records may hold the receipt; standalone reviews can return it directly. No commit, PR, or publication is required by this method.
34
+
35
+ ## Principles
36
+
37
+ - Findings need a concrete failure condition and a location in the reviewed change.
38
+ - Record the actual review source; self-check and independent review are different evidence.
39
+ - A clean reviewed diff does not grant release authority or establish customer acceptance.
@@ -0,0 +1,73 @@
1
+ # ship - Deliver and release with evidence
2
+
3
+ **Enter when:** an implemented increment needs a delivery checkpoint, deployment, or wider rollout. Use [build](build.md) for implementation; an untested business or technical assumption needs an experiment before a release claim.
4
+
5
+ Start from [task context](task-context.md). Standalone work uses supplied permitted context and a release receipt; it does not require `.fde/` initialization. In engagement mode, use privacy-safe views of confirmed context, decisions, terrain, success, delivery, applicable trust constraints, and AI evaluation evidence. Never load raw `<private>` blocks into a model. The `fde` CLI remains local-only; deployment uses the customer's authorized tools, never a new network capability inside `fde`.
6
+
7
+ ## Establish the delivery contract
8
+
9
+ Identify the exact outcome, acceptance check, scope, affected users/systems, target environment, recovery mechanism, and who or what is authorized to accept and release it. Reuse confirmed authority and checks for routine work. Do not invent missing signers, permissions, measurements, or acceptance.
10
+
11
+ For an initialized engagement, run `fde doctor --ready` before a new delivery plan or material scope change. Missing binary success or a named customer-side signer blocks that planning progression until resolved. A passing doctor validates record structure, not connectivity, release readiness, or customer acceptance. Standalone work evaluates the supplied contract directly.
12
+
13
+ A customer delivery checkpoint must let the agreed decision-maker replay and reject the acceptance check through an interface they operate. Prefer their staging; otherwise use an agreed representative environment and disclose its owner and limitations. Local green proves only the local run. Routine fixes may share an agreed checkpoint; no fixed number of changes forces a ceremony.
14
+
15
+ ## Prepare a reviewable increment
16
+
17
+ 1. Inspect repository instructions, working tree, overlapping work, and the complete intended release diff. Include working-tree changes when testing an uncommitted candidate. Preserve unrelated work; separate unintended behavior before release.
18
+ 2. Identify applicable before-state evidence and the changed outcome. Name dependencies, stop conditions, and irreversible effects. Existing applicable evidence may be reused with attribution, never represented as a fresh run.
19
+ 3. Complete [verification](verification.md) and [review](review.md), proportional to the change and repository requirements. A self-check is not an independent review. Record command, revision, environment, date, result, and unrun checks. Exercise relevant operating exceptions and fallback paths, not just the happy path.
20
+ 4. For AI behavior, obtain a scoped [eval verdict](eval-pack.md) for the candidate and applicable controls. A permitted bounded automation workflow remains permitted within its documented limits. A missing evaluation, failed critical case, or unknown action authority prevents release of that path; non-AI changes record eval as not applicable.
21
+
22
+ ## Release gate
23
+
24
+ Before deployment, establish these facts from existing evidence or a necessary check. Missing material evidence blocks the dependent release step; continue independent preparation. Do not ask again for approval already provided within the same scope.
25
+
26
+ | Dimension | Required evidence |
27
+ |-----------|-------------------|
28
+ | Candidate | Exact revision/artifact, intended diff, dependencies, applicable required checks passing; no skipped failure presented as green |
29
+ | Target and access | Service/account/region, environment, authorized deployment identity and mechanism, secret provisioning without revealing values |
30
+ | Acceptance | Replayable check and agreed decision-maker/mechanism; record actual acceptance separately from readiness |
31
+ | Data and policy | Permitted data, applicable security/residency/change-window requirements, necessary approvals already recorded or obtained |
32
+ | Recovery | Applicable tested rollback, restore, compensation, or roll-forward within agreed recovery-time/data-loss limits; explicit authority for irreversible effects |
33
+ | Operations | Named release/recovery owner, runbook appropriate to risk, health and business signals, stop thresholds, observation coverage |
34
+ | AI, when applicable | Current applicable SHIP eval evidence, critical failures zero, enforced action boundary and required human review or documented bounded automation |
35
+
36
+ Check migration compatibility, old/new version coexistence, delayed jobs, caches, and already-emitted side effects where relevant. A code revert does not undo data loss or external writes. Reuse drill evidence only when the mechanism and relevant conditions are unchanged, explaining applicability. If recovery is only a plan, exercise it in a permitted representative environment before release.
37
+
38
+ Use the repository's existing secret scanning and security checks; avoid diagnostic commands that print credential matches. Retain sanitized references to results. Resolve material evidence gaps or obtain an explicit, authorized narrowing of the release; do not average critical blockers into a readiness score.
39
+
40
+ For a coordinated engagement, also connect the release to the agreed value bucket and baseline/target, and record a dated receipt for the affected operating path. An unmeasured result remains pending with a measurement next step; do not invent realized value to pass a gate.
41
+
42
+ ## Deploy within authority
43
+
44
+ Execute only when the requested workflow authorizes deployment to this target and the applicable gates are met. Otherwise leave a concrete release candidate, exact deployment/recovery instructions, evidence, and the remaining authorization for review. A permission to implement or test is not permission to publish.
45
+
46
+ Use the customer's established pipeline and rollout mechanism. Select canary, staged exposure, blue/green, or direct rollout according to actual risk and platform capabilities; do not impose a universal cohort sequence. Define advance/abort thresholds and observation window before starting. If another operator must execute, record their handoff and report deployment pending until there is evidence it happened.
47
+
48
+ During rollout inspect health, errors, key user behavior, and side-effect integrity. Halt expansion on breached thresholds or critical harm and apply authorized containment/recovery. Do not continue merely because the deploy command exited successfully.
49
+
50
+ ## Verify operation and hand off
51
+
52
+ Run permitted smoke and acceptance checks against the deployed candidate. Record deployment identity/time, observed signals, sample/window, failures, recovery actions, and remaining gaps. Define the pulse: metric, cadence, threshold, owner, and response. For AI, include permitted output sampling and drift/action-boundary monitoring.
53
+
54
+ Before wider exposure, verify expected load/cost, data pipeline behavior, ownership, support, and applicable governance for the proposed audience. Choose expansion conditions from evidence; a successful pilot does not establish readiness for an arbitrary larger scale. Measure adoption against the eligible users, expected workflow frequency, and agreed observation window; investigate misses without guessing their cause.
55
+
56
+ ## Receipt and completion
57
+
58
+ Keep these claims separate: implemented, verified, deployed, measured outcome, and accepted. Include the candidate, target, command/pipeline, applicable checks and unrun checks, review source, evaluation where needed, authority source, recovery evidence, observation, and next owner/action. Attribute acceptance to its actual source and scope. A staging measurement is not production value, and a commit is not deployment.
59
+
60
+ In engagement mode, write confirmed implementation/decisions and delivery receipts under the existing record rules. Standalone work returns the same receipt or uses the repository's permitted release record. Committing, pushing, opening a PR, publishing, and notifying others are actions governed by the user's workflow, not mandatory steps imposed by this method.
61
+
62
+ ## Worked example
63
+
64
+ A freight team agrees that dispatchers can retry a failed export once without creating a duplicate shipment. The change uses the existing queue and ops screen. The developer records a failing duplicate-delivery case, implements idempotency, and passes the relevant checks on a named candidate. QA observes both the retry status and the single downstream record on permitted staging fixtures. A separate reviewer examines the queue race; the receipt names that review and the revision.
65
+
66
+ The export service has an approved staged-release workflow. Its owner reuses a recent recovery drill because the queue format and recovery mechanism are unchanged, recording that applicability. Deployment stops if duplicate records appear or the agreed error threshold is crossed. The authorized rollout completes, production smoke checks pass, and the dispatcher accepts the specified retry behavior with a dated source. The operating-cost benefit remains pending until the agreed measurement window closes. Implementation, deployment, acceptance, and measured value have different evidence. In the existing engagement, `decisions.md` records the agreed behavior and `delivery.md` holds the release receipt; standalone work returns those facts directly.
67
+
68
+ ## Principles
69
+
70
+ - Release the reviewed candidate with applicable evidence and documented authority.
71
+ - Test recovery against the effects that actually persist beyond a code revert.
72
+ - Keep missing evidence visible and distinguish local, staging, and production claims.
73
+ - Expansion follows observed acceptance and operating limits; fixed ceremonies cannot replace them.
@@ -0,0 +1,12 @@
1
+ # Task context and evidence
2
+
3
+ Use this contract for standalone methods and methods routed through `@fde`.
4
+
5
+ - **Standalone work:** use the supplied, permitted facts, notes, code, and artifacts. A client name, `.fde/` directory, or initialized engagement is not a prerequisite. Do not bootstrap records merely to run a method. Ask only for missing information or authority that changes the next action; mark other gaps as unknown.
6
+ - **Artifact names are destinations:** names such as `success.md`, `decisions.md`, and `delivery.md` identify relevant evidence and, when bound, record destinations. If absent, use supplied facts and return the requested draft or result in the current workspace or conversation. Do not invent files or require initialization to complete useful work.
7
+ - **Bound engagement:** honor the current client binding and constraints. Before reading records, run `fde privacy` to verify masking support. Obtain a fresh, identity-matching sanitized `fde resume` packet for this task (or reuse a fresh session-hook packet); retrieve missing evidence with targeted `fde recall <topic>`. Use bounded `fde handoff` for transfer work. Refresh after binding, masking, or record changes. Never substitute raw `.fde/` reads, private blocks, masking dictionaries, or full transcripts. If the CLI is unavailable, use only permitted supplied excerpts and report the context limitation.
8
+ - **Authority:** continue reversible work within authorized scope. Reuse prior authorization when it covers the specific action. Show consequential engagement-record judgments and uncertainties for confirmation before saving unless already explicitly confirmed. New scope, acceptance changes, production actions, exports, and external messages need the applicable authority; a method invocation alone does not supply it. Keep one customer's writes in that customer's record.
9
+ - **Evidence:** distinguish supplied facts, estimates, hypotheses, and unknowns. Cite actual sources; a log date is not attribution. Never invent a source, signer, signature, customer reaction, or acceptance. Keep outcomes **promised → measured → accepted** distinct, and implementation, verification, deployment, and customer acceptance separate. Missing evidence means unproven, not an observed failure.
10
+ - **Data boundary:** use only data permitted by the customer's AI policy; clarify unknown policy before loading their code or data. Never load `<private>` content into a model. Cross-client comparison and exporting reusable material require permission and removal of customer-identifying or confidential content; anonymization alone does not grant permission.
11
+
12
+ Apply the selected method to this context. Follow its linked supporting methods only when needed; do not restart discovery or repeat already answered questions.
@@ -0,0 +1,102 @@
1
+ # test-assumptions - Test assumptions
2
+
3
+ **Enter when:** the brief feels too neat, the customer is very confident about the solution (not the problem), someone says "we just need…" about a complex system, or discover surfaced contradictions between what was said and what the codebase shows.
4
+
5
+ **Read first:** `brief.md`, `reality.md`, `terrain.md`, `context.md`. The assumptions are hiding between what the brief says and what the code does.
6
+
7
+ Every engagement is built on assumptions. Most are invisible until they're wrong and the build is two weeks deep. The assumption audit makes them visible - and killable - before they cost time.
8
+
9
+ ## Method (you do this work)
10
+
11
+ **1. Extract the assumptions.** Read `brief.md`, `reality.md`, and `terrain.md` `## Parts` line by line. Every statement that isn't backed by evidence is an assumption. Treat every "obvious" block as a convention until a receipt proves it. Common hiding places:
12
+
13
+ | Where assumptions hide | Example | The real question |
14
+ |----------------------|---------|-------------------|
15
+ | **The problem statement** | "The API is slow" | Slow for whom? Measured how? Since when? |
16
+ | **The proposed solution** | "We need to migrate to microservices" | Is the monolith actually the bottleneck, or is it the database? |
17
+ | **The timeline** | "This should take two weeks" | Based on what? Who estimated? Have they done this before? |
18
+ | **The stakeholder claim** | "The team is on board" | Who specifically? Have they been asked? What did the resistors say? |
19
+ | **The data claim** | "We have good data for this" | Defined how? Validated when? By whom? Sample checked? |
20
+ | **The "just"** | "We just need to add a feature" | On what system? With what dependencies? What breaks? |
21
+
22
+ **2. Kind first, then blast radius.** For each row, classify:
23
+
24
+ | Kind | Meaning |
25
+ |------|---------|
26
+ | **FACT** | A dated receipt, a measurement, or the repo. You can point at it. |
27
+ | **CONVENTION** | How they have always done it. The playbook. "We just…" |
28
+ | **UNKNOWN** | No evidence either way. |
29
+
30
+ Order the list load-bearing first. For each CONVENTION or UNKNOWN, one line: what breaks if it is wrong, and what opens if you **invert** it (stop obeying it). A FACT with no receipt is UNKNOWN - do not promote it to protect the brief.
31
+
32
+ Then classify blast radius:
33
+
34
+ ```
35
+ CRITICAL - if wrong, the engagement fails or the approach changes fundamentally
36
+ → Must be validated before plan starts
37
+
38
+ LOAD-BEARING - if wrong, significant rework or timeline change
39
+ → Must be validated before build starts
40
+
41
+ CONVENIENCE - if wrong, a task changes but the approach holds
42
+ → Validate when you get there
43
+ ```
44
+
45
+ **3. Design the validation.** Each critical assumption gets one specific test - not a discussion, a test:
46
+
47
+ | Assumption | Validation method | Effort | Evidence threshold |
48
+ |-----------|-------------------|--------|-------------------|
49
+ | "The API is the bottleneck" | Instrument the three slowest endpoints, measure p95 over 24h | 2h | Latency data shows >80% of wait time in API layer |
50
+ | "The team will adopt the new tool" | Ask three team members individually: "Show me how you'd use this" | 1h | 2 of 3 can describe a use case without prompting |
51
+ | "The data is clean enough for ML" | Sample 200 records, count nulls/duplicates/format errors | 1h | <5% error rate on the fields the model needs |
52
+
53
+ **4. Run the killer test first.** The assumption with the highest blast radius AND the cheapest validation gets tested immediately. This single principle saves more engagement time than any other: if the killer assumption is wrong, you've saved weeks; if it holds, you've bought confidence. Write the kill observation in `How we test` as the result that would **stop** the plan - plan copies that line onto each Now PR as `Kill if`.
54
+
55
+ **5. Present findings as a fact base, not a challenge.**
56
+
57
+ The customer's assumptions are often wrong, but calling them wrong is a trust withdrawal. Frame as curiosity, not contradiction:
58
+
59
+ > "The brief says the API is the bottleneck. The codebase shows 80% of latency is in the database layer - here's the evidence. Should we adjust the focus?"
60
+
61
+ Evidence first, then the question. Let them reach the conclusion.
62
+
63
+ ## Artifact
64
+
65
+ **`assumptions.md`** - this IS the register (create if land did not). Keep one live table; do not only bury results in `reality.md`:
66
+
67
+ ```markdown
68
+ | # | Assumption | Kind | Blast radius | How we test | Status | Evidence |
69
+ |---|------------|------|--------------|-------------|--------|----------|
70
+ | 1 | API is the bottleneck | CONVENTION | CRITICAL | p95 instrumentation 24h | DISPROVED | 80% wait in DB layer (Day N) |
71
+ | 2 | Team will adopt new tool | UNKNOWN | LOAD-BEARING | 3 individual interviews | CONFIRMED | 2/3 describe a use case unprompted |
72
+ | 3 | Data clean enough for ML | UNKNOWN | CRITICAL | 200-record sample | PARTIAL → OPEN follow-up | 12% nulls on key field; cleaning task added |
73
+ ```
74
+
75
+ Status values: `OPEN` · `TESTING` · `CONFIRMED` · `DISPROVED` · `PARKED`. A CRITICAL row still `OPEN` blocks plan.
76
+
77
+ **`reality.md`** - short pointer only: which assumptions changed the approach and the implication for build.
78
+
79
+ **`decisions.md`** - if an assumption was disproved and the approach changed: what shifted, why, the evidence, same day.
80
+
81
+ ## Checkpoint
82
+
83
+ Tell the FDE: how many assumptions extracted, how many critical, which ones were tested, which changed the direction. If a critical assumption is disproved: recommend the next move (rescope, pivot, or the conversation with the sponsor) before the FDE asks. If any CRITICAL remains OPEN: do not route to plan.
84
+
85
+ ## Worked example
86
+
87
+ Acme's brief reads cleanly, which is the signal.
88
+
89
+ Extracted assumptions include one nobody said aloud: *finance would act on an alert*. The whole plan rests on it, and the evidence behind it is a sentence in a kickoff. Blast radius CRITICAL - if false, alerting changes nothing and the engagement delivers a page nobody answers.
90
+
91
+ Validation is a test, not a discussion, and it is cheap: send one real failure notification to the finance channel and watch what happens. It goes first because highest blast radius × cheapest test is the killer test.
92
+
93
+ Result: acked in 40 minutes, by Marco, not finance. Assumption DISPROVED, and the plan changes before six weeks are spent on it - the alert needs a rota with an owner, which is a different piece of work than the one that was funded. `assumptions.md` records the status, the evidence, and the date; the finding is presented to the FDE as a fact base, not as "the brief was wrong".
94
+
95
+ ## Principles
96
+
97
+ - Every "just" is an assumption. Every "should" is an assumption.
98
+ - Kind before blast radius. A FACT with no receipt is UNKNOWN.
99
+ - Kill the riskiest, cheapest-to-test assumption first.
100
+ - Evidence first, then the question. Let the customer reach the conclusion.
101
+ - A brief with zero disproved assumptions wasn't audited - it was accepted.
102
+ - Two weeks of building on a wrong assumption costs more than two hours of testing.
@@ -0,0 +1,90 @@
1
+ # three-options - Generate options
2
+
3
+ **Context:** apply [task context and evidence](task-context.md) before using the named records below.
4
+
5
+ **Enter when:** a significant technical or strategic decision needs to be made, the FDE is asked "what should we do?", the team is stuck between approaches, or a fork in the engagement requires the sponsor's input.
6
+
7
+ **Read first:** `reality.md`, `terrain.md`, `assumptions.md`, `success.md`, `context.md`. Load `business-case.md` if the decision has cost implications.
8
+
9
+ Compare materially different, defensible alternatives. Three is a useful presentation shape when three viable paths exist; do not pad the set to meet a quota. Include keeping the current approach or deferring when those are credible choices.
10
+
11
+ ## Method (you do this work)
12
+
13
+ **1. Name the decision.** One sentence: what needs to be decided, by whom, by when, and what happens if it's deferred.
14
+
15
+ > "Decision: approach for the payment migration. Decided by: CTO. Needed by: Friday. Deferral cost: blocks the next sprint and delays the pilot by two weeks."
16
+
17
+ **2. Generate genuine options from the evidence.** Use confirmed assumptions and known system parts; exclude disproved assumptions. If those records are absent, identify the supplied facts and unknowns. Use [test-assumptions](test-assumptions.md) only when a consequential assumption needs investigation.
18
+
19
+ Alternatives should differ materially in architecture, operating model, scope, cost, or reversibility. Do not label the same plan good / medium / bad or manufacture an unsafe option to favor your recommendation.
20
+
21
+ Each option must be one the FDE would genuinely recommend under different circumstances. If you cannot defend an option, replace it - padding is visible.
22
+
23
+ For each option, name:
24
+ - which surviving blocks it is built from
25
+ - which constraint or convention it changes, if any
26
+ - its single biggest point of failure
27
+ - any new building block, labelled as a new assumption (`UNKNOWN` in `assumptions.md`) - do not smuggle one in as a fact
28
+
29
+ If only one viable path remains, explain what ruled out the alternatives and what evidence could reopen them.
30
+
31
+ **3. Structure each option identically.** Same dimensions, same format - so comparison is instant:
32
+
33
+ ```markdown
34
+ ### Option A: <name>
35
+ - **Blocks:** <which surviving assumptions / parts it is built from>
36
+ - **Constraint changed:** <constraint or convention changed, if any>
37
+ - **What:** <the approach in one paragraph>
38
+ - **Timeline:** <estimate with basis>
39
+ - **Cost:** <effort, infrastructure, external>
40
+ - **Biggest failure:** <the single point that kills this option>
41
+ - **Trade-off:** <what you give up by choosing this>
42
+ - **Best when:** <the condition that makes this the right choice>
43
+ ```
44
+
45
+ **4. Make comparison easy.** Use consistent dimensions: expected outcome, evidence, build and operating cost, time with estimate basis, reversibility, owner, and the most consequential uncertainty. A compact table helps when alternatives need comparison; do not fill it with invented numbers or label one path universally cheapest.
46
+
47
+ **5. State your recommendation - and why.** Separate supplied facts and estimates from your judgment:
48
+
49
+ > "I recommend Option B. The limited team availability rules out a full rewrite this quarter, and the current hotspot makes keeping the job unchanged costly. Option B gets us to pilot in 4 weeks with a tested rollback."
50
+
51
+ **6. Handle the override gracefully.** If the sponsor picks a different option:
52
+
53
+ - Log it in `decisions.md`: the choice, who made it, the trade-off they accepted.
54
+ - Adjust the plan to the chosen option. Don't passive-aggressively optimise for your preference.
55
+ - If the chosen option has a specific risk you flagged: note the early-warning signal in `risks.md` so it's caught if it materialises.
56
+
57
+ ## Artifact
58
+
59
+ **`decisions.md`** - the options analysis:
60
+ ```markdown
61
+ ## Decision: <name> - <date>
62
+ Decided by: <who>
63
+ Options presented: <viable alternatives>
64
+ Recommended: B - <one line why>
65
+ Chosen: <option or pending> by <actual decision-maker or unknown>
66
+ Trade-off accepted: <what the choice gives up>
67
+ ```
68
+
69
+ The full option details in the same entry or linked to a section in `reality.md`.
70
+
71
+ ## Checkpoint
72
+
73
+ Present the viable alternatives, recommendation, and the evidence or constraint that would change it. Reuse known decision authority; if a decision remains pending, record it as pending.
74
+
75
+ ## Worked example
76
+
77
+ Fictional example: a support team needs completed requests written back to its service system. Its product can export a file today. A supported connector is expected in six weeks; the customer wants automation in two. A custom API adapter looks feasible, but nobody has accepted its maintenance.
78
+
79
+ Two defensible paths remain. Continue the approved export while checking the supported connector's fit, or investigate a bounded adapter whose delivery depends on a named owner and tested API behavior. The export is a bridge within the first path, not a third option invented for the slide. Compare manual effort, engineering effort, ongoing support, and the effect of waiting. Time released is capacity unless spending actually falls.
80
+
81
+ Recommend the bridge while resolving native fit and the value of earlier automation. Reconsider the adapter if that value justifies full costs and an owner accepts it. `decisions.md` (or a standalone decision note) records the recommendation, evidence, unknowns, and pending decision. It does not claim that a sponsor chose it or that the adapter can meet the date.
82
+
83
+ ## Principles
84
+
85
+ - Compare defensible alternatives; the number follows the evidence.
86
+ - Build from known facts and label new assumptions. Material differences make the comparison useful.
87
+ - Each option must be genuinely defensible - no straw men.
88
+ - Same structure for each option. Comparison should take 30 seconds.
89
+ - Recommend one. State why. Accept the override gracefully.
90
+ - An override logged with its trade-off protects the FDE when the risk materialises.
@@ -0,0 +1,31 @@
1
+ # verification - Make a claim replayable
2
+
3
+ **Enter when:** reporting completion, evaluating an acceptance check, handing work to a reviewer, or preparing a release.
4
+
5
+ Use [task context](task-context.md). This method returns evidence directly or writes an existing permitted task/engagement record; it never requires `.fde/` initialization.
6
+
7
+ ## Method
8
+
9
+ 1. Translate each claim into the observation that would support or reject it. Reuse agreed acceptance criteria and required repository checks. Select focused checks for changed behavior before broadening to release requirements.
10
+ 2. Identify the actual repository commands, fixtures, runtime, and environment. Read command behavior before executing it, especially when it can write externally. Use authorized environments and avoid leaking secrets through logs or diagnostic commands.
11
+ 3. Run the checks and inspect results, including exit status and relevant output. A running job, test discovery, a mocked response, and a successful real request are different evidence. Record asynchronous completion before claiming success.
12
+ 4. Bind evidence to the tested revision and working tree. For uncommitted changes record the base revision plus changed paths and an available diff digest or snapshot identifier. For browser/manual checks record the steps, inputs, observed result, and inspected evidence.
13
+ 5. After a change, rerun checks whose behavior or assumptions were affected. Reuse prior evidence only when the relevant code, dependencies, data, and environment remain applicable; cite the original run and reason. Never imply reused evidence was rerun.
14
+ 6. Label every required check **passed**, **failed**, **blocked**, or **not run**. Include why blocked/not run, impact, and next step. Missing evidence is unproven; it is not an observed failure or a pass.
15
+
16
+ ## Receipt
17
+
18
+ Use one compact entry per check or a table with these fields:
19
+
20
+ - Claim / acceptance check and expected result.
21
+ - Exact command and working directory, or manual journey and inputs.
22
+ - Environment, runtime/tool versions when relevant, and fixture/data source.
23
+ - Revision plus working-tree identity; run date/time.
24
+ - Observed result and exit status where available; safe evidence location.
25
+ - Status, limitations, unrun checks, and next step.
26
+
27
+ Keep implementation, verification, deployment, measured outcome, and customer acceptance distinct. A local pass supports the tested local behavior. An acceptance claim needs an attributed source from the agreed decision-maker or agreed acceptance mechanism. Record no raw `<private>` blocks, credentials, or hidden reasoning.
28
+
29
+ ## Acceptance
30
+
31
+ A completion statement cites applicable evidence for its claims and explicitly names material gaps. If required checks fail, investigate or report the blocker; never skip them, edit expectations, or relabel the scope without authority to obtain a green result.