fdeops 4.0.4 → 4.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (139) hide show
  1. package/AGENTS.md +1 -1
  2. package/README.md +33 -6
  3. package/bin/check.js +14 -23
  4. package/bin/generate-skills.js +71 -0
  5. package/bin/install.js +2 -1
  6. package/bin/skill-catalog.js +17 -0
  7. package/mcp/fdeops-ingest/package.json +2 -2
  8. package/package.json +4 -3
  9. package/plugin.json +2 -2
  10. package/skills/fde/SKILL.md +18 -10
  11. package/skills/fde/references/board-memo.md +1 -1
  12. package/skills/fde/references/build.md +20 -0
  13. package/skills/fde/references/business-case.md +9 -7
  14. package/skills/fde/references/close.md +5 -3
  15. package/skills/fde/references/debug.md +18 -0
  16. package/skills/fde/references/encode-pattern.md +10 -8
  17. package/skills/fde/references/eval-pack.md +16 -33
  18. package/skills/fde/references/hold-scope.md +9 -7
  19. package/skills/fde/references/integrate.md +18 -0
  20. package/skills/fde/references/plan.md +4 -2
  21. package/skills/fde/references/poc.md +5 -3
  22. package/skills/fde/references/qa.md +18 -0
  23. package/skills/fde/references/review.md +22 -60
  24. package/skills/fde/references/ship.md +42 -287
  25. package/skills/fde/references/task-context.md +12 -0
  26. package/skills/fde/references/test-assumptions.md +2 -2
  27. package/skills/fde/references/three-options.md +19 -27
  28. package/skills/fde/references/verification.md +31 -0
  29. package/skills/fde-build/SKILL.md +21 -0
  30. package/skills/fde-build/references/build.md +20 -0
  31. package/skills/fde-build/references/debug.md +18 -0
  32. package/skills/fde-build/references/eval-pack.md +26 -0
  33. package/skills/fde-build/references/integrate.md +18 -0
  34. package/skills/fde-build/references/qa.md +18 -0
  35. package/skills/fde-build/references/review.md +39 -0
  36. package/skills/fde-build/references/ship.md +73 -0
  37. package/skills/fde-build/references/task-context.md +12 -0
  38. package/skills/fde-build/references/verification.md +31 -0
  39. package/skills/fde-debug/SKILL.md +21 -0
  40. package/skills/fde-debug/references/build.md +20 -0
  41. package/skills/fde-debug/references/debug.md +18 -0
  42. package/skills/fde-debug/references/eval-pack.md +26 -0
  43. package/skills/fde-debug/references/integrate.md +18 -0
  44. package/skills/fde-debug/references/qa.md +18 -0
  45. package/skills/fde-debug/references/review.md +39 -0
  46. package/skills/fde-debug/references/ship.md +73 -0
  47. package/skills/fde-debug/references/task-context.md +12 -0
  48. package/skills/fde-debug/references/verification.md +31 -0
  49. package/skills/fde-discover/SKILL.md +21 -0
  50. package/skills/fde-discover/references/audit.md +71 -0
  51. package/skills/fde-discover/references/discover.md +254 -0
  52. package/skills/fde-discover/references/task-context.md +12 -0
  53. package/skills/fde-evaluate/SKILL.md +21 -0
  54. package/skills/fde-evaluate/references/build.md +20 -0
  55. package/skills/fde-evaluate/references/debug.md +18 -0
  56. package/skills/fde-evaluate/references/eval-pack.md +26 -0
  57. package/skills/fde-evaluate/references/integrate.md +18 -0
  58. package/skills/fde-evaluate/references/qa.md +18 -0
  59. package/skills/fde-evaluate/references/review.md +39 -0
  60. package/skills/fde-evaluate/references/ship.md +73 -0
  61. package/skills/fde-evaluate/references/task-context.md +12 -0
  62. package/skills/fde-evaluate/references/verification.md +31 -0
  63. package/skills/fde-feedback/SKILL.md +21 -0
  64. package/skills/fde-feedback/references/encode-pattern.md +96 -0
  65. package/skills/fde-feedback/references/task-context.md +12 -0
  66. package/skills/fde-handoff/SKILL.md +21 -0
  67. package/skills/fde-handoff/references/close.md +66 -0
  68. package/skills/fde-handoff/references/encode-pattern.md +96 -0
  69. package/skills/fde-handoff/references/task-context.md +12 -0
  70. package/skills/fde-integrate/SKILL.md +21 -0
  71. package/skills/fde-integrate/references/build.md +20 -0
  72. package/skills/fde-integrate/references/debug.md +18 -0
  73. package/skills/fde-integrate/references/eval-pack.md +26 -0
  74. package/skills/fde-integrate/references/integrate.md +18 -0
  75. package/skills/fde-integrate/references/qa.md +18 -0
  76. package/skills/fde-integrate/references/review.md +39 -0
  77. package/skills/fde-integrate/references/ship.md +73 -0
  78. package/skills/fde-integrate/references/task-context.md +12 -0
  79. package/skills/fde-integrate/references/verification.md +31 -0
  80. package/skills/fde-options/SKILL.md +21 -0
  81. package/skills/fde-options/references/business-case.md +90 -0
  82. package/skills/fde-options/references/task-context.md +12 -0
  83. package/skills/fde-options/references/test-assumptions.md +102 -0
  84. package/skills/fde-options/references/three-options.md +90 -0
  85. package/skills/fde-poc/SKILL.md +21 -0
  86. package/skills/fde-poc/references/audit.md +71 -0
  87. package/skills/fde-poc/references/build.md +20 -0
  88. package/skills/fde-poc/references/business-case.md +90 -0
  89. package/skills/fde-poc/references/debug.md +18 -0
  90. package/skills/fde-poc/references/discover.md +254 -0
  91. package/skills/fde-poc/references/eval-pack.md +26 -0
  92. package/skills/fde-poc/references/integrate.md +18 -0
  93. package/skills/fde-poc/references/plan.md +167 -0
  94. package/skills/fde-poc/references/poc.md +55 -0
  95. package/skills/fde-poc/references/qa.md +18 -0
  96. package/skills/fde-poc/references/review.md +39 -0
  97. package/skills/fde-poc/references/ship.md +73 -0
  98. package/skills/fde-poc/references/task-context.md +12 -0
  99. package/skills/fde-poc/references/test-assumptions.md +102 -0
  100. package/skills/fde-poc/references/three-options.md +90 -0
  101. package/skills/fde-poc/references/verification.md +31 -0
  102. package/skills/fde-qa/SKILL.md +21 -0
  103. package/skills/fde-qa/references/build.md +20 -0
  104. package/skills/fde-qa/references/debug.md +18 -0
  105. package/skills/fde-qa/references/eval-pack.md +26 -0
  106. package/skills/fde-qa/references/integrate.md +18 -0
  107. package/skills/fde-qa/references/qa.md +18 -0
  108. package/skills/fde-qa/references/review.md +39 -0
  109. package/skills/fde-qa/references/ship.md +73 -0
  110. package/skills/fde-qa/references/task-context.md +12 -0
  111. package/skills/fde-qa/references/verification.md +31 -0
  112. package/skills/fde-readout/SKILL.md +21 -0
  113. package/skills/fde-readout/references/board-memo.md +108 -0
  114. package/skills/fde-readout/references/business-case.md +90 -0
  115. package/skills/fde-readout/references/readout.md +69 -0
  116. package/skills/fde-readout/references/task-context.md +12 -0
  117. package/skills/fde-review/SKILL.md +21 -0
  118. package/skills/fde-review/references/build.md +20 -0
  119. package/skills/fde-review/references/debug.md +18 -0
  120. package/skills/fde-review/references/eval-pack.md +26 -0
  121. package/skills/fde-review/references/integrate.md +18 -0
  122. package/skills/fde-review/references/qa.md +18 -0
  123. package/skills/fde-review/references/review.md +39 -0
  124. package/skills/fde-review/references/ship.md +73 -0
  125. package/skills/fde-review/references/task-context.md +12 -0
  126. package/skills/fde-review/references/verification.md +31 -0
  127. package/skills/fde-scope/SKILL.md +21 -0
  128. package/skills/fde-scope/references/hold-scope.md +83 -0
  129. package/skills/fde-scope/references/task-context.md +12 -0
  130. package/skills/fde-ship/SKILL.md +21 -0
  131. package/skills/fde-ship/references/build.md +20 -0
  132. package/skills/fde-ship/references/debug.md +18 -0
  133. package/skills/fde-ship/references/eval-pack.md +26 -0
  134. package/skills/fde-ship/references/integrate.md +18 -0
  135. package/skills/fde-ship/references/qa.md +18 -0
  136. package/skills/fde-ship/references/review.md +39 -0
  137. package/skills/fde-ship/references/ship.md +73 -0
  138. package/skills/fde-ship/references/task-context.md +12 -0
  139. package/skills/fde-ship/references/verification.md +31 -0
@@ -1,5 +1,7 @@
1
1
  # encode-pattern - Encode the pattern
2
2
 
3
+ **Context:** apply [task context and evidence](task-context.md) before using the named records below.
4
+
3
5
  **Enter when:** the engagement is closing and reusable patterns exist, a technique worked well and will apply to future clients, the FDE notices themselves doing the same thing on a second engagement, or close identified a pattern worth preserving.
4
6
 
5
7
  **Read first:** `decisions.md`, `reality.md`, `delivery.md`, `retrospectives/`, `context.md`. Patterns live in what was *done*, not what was planned.
@@ -36,7 +38,7 @@ The difference between a 5-year FDE and a 15-year FDE is not talent - it's encod
36
38
  <The failure mode or edge case that makes the pattern not apply.>
37
39
 
38
40
  ### Evidence
39
- <Which engagement, what happened, what the result was. Specific, not generic.>
41
+ <Permitted source, what happened, measured result, and limits. Keep identifying evidence in its original customer record.>
40
42
  ```
41
43
 
42
44
  **3. The pattern quality test.** Before encoding:
@@ -63,21 +65,21 @@ The difference between a 5-year FDE and a 15-year FDE is not talent - it's encod
63
65
  **5. Version and evolve.** Patterns are living documents:
64
66
 
65
67
  - First use: **v0.1** - hypothesis based on one engagement
66
- - Second use: **v1.0** - confirmed pattern, refined from two experiences
68
+ - Later uses: record context, observed results, failures, and refinements. Repetition supplies evidence; it does not automatically validate the pattern. Promote a version when a substantive revision warrants it, not at a fixed use count.
67
69
  - After modification: increment minor version with what changed and why
68
70
  - After contradiction: note the counter-example, adjust the "watch out for" section
69
71
 
70
- **6. Cross-engagement pattern mining.** When the FDE has multiple engagements in `.fde/`:
72
+ **6. Cross-engagement pattern mining.** Only when explicitly authorized for the named engagements and permitted by each customer's data policy. Keep records separate; use sanitized CLI packets or targeted recall in each authorized context, never raw file comparisons. Otherwise extract a candidate from the current permitted context only.
71
73
 
72
- - Compare `reality.md` across engagements - do the same problems recur?
73
- - Compare `decisions.md` - are the same decisions being made?
74
- - Compare `retrospectives/` - are the same lessons being "learned" twice?
74
+ - Compare permitted problem summaries - do the same problems recur?
75
+ - Compare permitted decision summaries - are the same decisions being made?
76
+ - Compare permitted lessons - are the same lessons being learned twice?
75
77
 
76
78
  A pattern learned twice is a process failure. Encoding it prevents the third time.
77
79
 
78
80
  ## Artifact
79
81
 
80
- **`patterns.md`** - the pattern library, growing across engagements. Each pattern in the format above. Indexed by stage and situation trigger.
82
+ **`patterns.md`** - candidates and evidence for this engagement, indexed by stage and situation trigger. A shared library is a separate, authorized export: remove names, identifiers, distinctive operational details, secrets, and confidential code or data. Keep source receipts in the original record and export only permitted generalizations.
81
83
 
82
84
  **`retrospectives/YYYY-MM-DD-<engagement>.md`** - reference to which patterns were extracted from this engagement.
83
85
 
@@ -90,5 +92,5 @@ Present the extracted patterns to the FDE: "From this engagement, I've identifie
90
92
  - If you did it twice, encode it. The same lesson learned three times is a failure.
91
93
  - Patterns are steps, not principles. "Build trust" isn't a pattern; "fix a small visible bug on day one" is.
92
94
  - Every pattern needs a situation trigger - the FDE must recognise when it applies.
93
- - Version patterns. A pattern from one engagement is a hypothesis; from two, it's confirmed.
95
+ - Version substantive changes. State the evidence and limits; repeated use is not automatic confirmation.
94
96
  - The pattern library is the FDE's compound interest. It's what separates 5 years of experience from 1 year repeated 5 times.
@@ -1,43 +1,26 @@
1
- # eval-pack - Gate the model before it acts
1
+ # eval-pack - Evaluate the model's allowed behavior
2
2
 
3
- **Enter when:** the work touches AI/LLM/agents/RAG, or they need to POC a model, or ship/close is blocked because there is no evidence the non-deterministic path is safe. Activate alongside `ai.md`, `poc`, or `ship` - not instead of them.
3
+ **Enter when:** AI, LLM, RAG, or agent behavior needs evidence before an experiment, release, or material expansion. Non-AI work skips this method.
4
4
 
5
- **Read first:** `trust-profile.md` (AI policy + HITL), `terrain.md` (operating map), `delivery.md`. Create or extend `evals.md`.
5
+ Use [task context](task-context.md). Supplied permitted context and an evaluation report are sufficient without `.fde/`. In a coordinated engagement, use the existing trust/terrain context and keep the report in `evals.md`; read only privacy-safe views.
6
6
 
7
- Non-AI engagements skip this pack entirely.
7
+ ## Method
8
8
 
9
- ## Why this exists
9
+ 1. **Define the evaluated surface.** Name the model judgment, inputs, outputs, downstream actions, environment, and relevant failure impact. Separate model quality from deterministic tool authorization and application checks. Document the actual allowed action boundary and its source; missing authority remains unknown.
10
+ 2. **Choose cases by risk and coverage.** Use permitted historical examples, expert-labeled cases, or clearly marked synthetic fixtures. Cover relevant segments, boundary conditions, known failure modes, and critical harms. Record input, expected outcome/rubric, provenance, and critical-failure rule per case. Keep evaluation cases separate from tuning where possible; no fixed case count proves safety.
11
+ 3. **Agree the pass rule before the run.** Define quality thresholds, critical failures, coverage expectations, and acceptable uncertainty for this use. Use deterministic checks where possible and inspect subjective labels or judge reliability. Propose missing criteria for agreement; do not manufacture acceptance from the observed scores.
12
+ 4. **Run the actual evaluated path.** Record model/provider version, prompts/configuration, retrieval corpus or tools, application revision, environment, fixtures, and run date. Repeat where variability matters. Report totals, per-segment results, critical failures, and limitations using [verification](verification.md). A model-only run does not prove the agent's tool boundary works.
13
+ 5. **Verify action authority and controls.** Human approval is required where the user's policy or task requires it. Already agreed bounded automation may run within its documented actions, identities, environments, and limits; do not require fresh approval for every authorized action. Check enforcement outside the model, least privilege, input/output validation, cost/rate limits, stop conditions, observability, and recovery as applicable. Unknown or exceeded authority blocks those actions. Evaluation success never grants new authority.
14
+ 6. **Make a scoped verdict.** Report **SHIP** only when agreed criteria pass, critical failures are zero, applicable authority/control checks pass, and material coverage gaps are resolved or the release is explicitly narrowed by the responsible decision-maker. Otherwise report **NO-SHIP** with the smallest corrective step: fix, gather evidence, descope, or reconsider the judgment surface. A SHIP verdict is technical evidence for the stated scope, not permission to deploy.
10
15
 
11
- Intelligence without evidence is token-maxing with a nicer name. An FDE earns trust by showing: golden cases, failure modes, and a human gate before action.
16
+ ## Deliverable and acceptance
12
17
 
13
- ## Method (you do this work)
18
+ Return the suite/source, thresholds, counts and segments, top failure modes, control evidence, human-review gate or bounded automation authority, limitations, and dated verdict. Record unknown values honestly. Reevaluate after changes that affect model behavior, retrieval, tool permissions, or data conditions; cite why unchanged evidence remains applicable rather than implying a rerun.
14
19
 
15
- **1. Scope the judgement surface.** One sentence: which step uses model judgement, and what must never be autonomous.
16
-
17
- **2. Build a golden set sized to risk and coverage.** A pilot may start with 5-20 cases; neither that range nor 50-100 proves broad-scale safety. Include representative segments, boundary cases, and known critical failures. Expand based on observed failure modes and uncertainty; hold out cases from tuning and repeat runs when variability matters. For each case:
18
- - input (sanitized - no `<private>` raw values)
19
- - expected outcome or expert-approved acceptance note
20
- - pass rule (exact / contains / short rubric)
21
- - source: real historical example / expert label / staged fixture
22
-
23
- **3. Score pass/fail, not vibes.** Define the quality threshold and critical failure rules before running the suite. Run it; record count pass / fail, segment coverage, and limitations. Failures get a failure-mode tag (missing data, wrong record, format drift, hallucination, retrieval miss, unsafe action, other).
24
-
25
- **4. Human-in-the-loop gate.** Name which outcomes require human approve before side effects. Judgement that has a side effect (write, send, transfer, ticket, deploy, pay, page) is **NO-SHIP** without a named human on their side in the loop. Do not write "none - allowed under policy" to bless lights-out write-access. Staging may run a supervised loop with a kill switch, a cost cap, and a golden set from **their** failures. Production stays gated until they have a written policy, a named owner, and dated eval receipts on real traffic.
26
-
27
- **5. Ship rule.** Until `evals.md` shows Verdict **SHIP** with a dated run (critical fails = 0) and HITL filled, AI-touching ship stays **fix-first**. Eval fails do not sit in a backlog - they reopen plan (descope, move the judgement, or kill the path). Log a one-line eval receipt in `delivery.md` → `## Ship receipts`. No "probably fine."
28
-
29
- ## Artifact - `evals.md`
30
-
31
- Create on first AI-touching slice (not at `resume --init`). Use the stub in `templates/.fde/evals.md`. Every claim needs a source. Missing evidence → leave the cell `unknown - ask:`, never invent scores.
32
-
33
- ## Checkpoint
34
-
35
- Present to the FDE: suite size, pass rate, top failure mode, HITL gate, Verdict SHIP/NO-SHIP. If they want to ship without a run: say no, and offer the smallest suite that would unblock.
20
+ When coordinated, append a concise eval receipt to delivery records. For release use [ship](ship.md). For ongoing use define the drift signals, sample policy permitted by data handling rules, owner, and conditions that suspend or narrow automation. Do not store secrets, raw `<private>` data, or hidden chain-of-thought in reports.
36
21
 
37
22
  ## Principles
38
23
 
39
- - No golden set, no AI ship.
40
- - Pass/fail beats “looks good.”
41
- - Failure modes are the product - the happy path is table stakes.
42
- - HITL is a gate, not a slide. Side effects without a named human on their side are NO-SHIP.
43
- - Non-AI work does not need this file.
24
+ - Thresholds and authority come from the agreed contract, never from a convenient observed result.
25
+ - Critical failures block the evaluated release scope; disclose coverage and uncertainty.
26
+ - Bound automation with enforceable controls, and require human review where the policy requires it.
@@ -1,5 +1,7 @@
1
1
  # hold-scope - Hold scope
2
2
 
3
+ **Context:** apply [task context and evidence](task-context.md) before using the named records below.
4
+
3
5
  **Enter when:** "also can you…" mid-build, a stakeholder adds requirements without adjusting timeline, the FDE feels scope creeping but can't name it, or `success.md` no longer matches what's being asked.
4
6
 
5
7
  **Read first:** `success.md` (the agreed boundary), `decisions.md`, `context.md`. Load `stakeholders.md` to know who's asking and their signal.
@@ -11,7 +13,7 @@ Scope creep is the leading cause of FDE engagement failure - not technical compl
11
13
  **1. Detect before it compounds.** Three patterns that signal creep before it's named:
12
14
 
13
15
  | Pattern | What it sounds like | What's actually happening |
14
- |---------|--------------------|--------------------------|
16
+ |---------|--------------------|--------------------------|
15
17
  | **The friendly addition** | "While you're in there, could you also…" | Adjacent work getting absorbed without timeline adjustment |
16
18
  | **The evolved requirement** | "Oh, what I actually meant was…" | The original scope was never clear enough - `success.md` needs updating |
17
19
  | **The stakeholder swap** | A new person starts requesting features the original sponsor didn't | Power shifted; the real scope is being rewritten informally |
@@ -38,7 +40,7 @@ Log via `fde log decision "scope change: <summary> - requested by <who>, impact:
38
40
 
39
41
  **The key phrase: "Let me place it."** Not "that's out of scope" (adversarial) or "sure" (absorbed). "Let me place it" signals you're taking it seriously while buying time to assess the real cost.
40
42
 
41
- **4. The accumulation conversation.** When the scope receipts show a pattern - typically 3-5 absorbed changes - the FDE needs a conversation with the sponsor:
43
+ **4. The accumulation conversation.** When the scope receipts show a pattern - a material cumulative impact on delivery, cost, risk, or acceptance - the FDE needs a conversation with the sponsor:
42
44
 
43
45
  Frame it as **protection, not complaint:**
44
46
  > "We've absorbed five changes since the original agreement. Each one made sense individually. Together, they've added roughly two weeks. I want to make sure the timeline expectation still matches - should we adjust the delivery date, or reprioritise to keep the original date?"
@@ -59,23 +61,23 @@ Evidence-based: point to `decisions.md` scope receipts with dates and requesters
59
61
 
60
62
  ## Checkpoint
61
63
 
62
- Weekly check: count the scope receipts since last conversation. Three or more unaddressed → recommend the accumulation conversation to the FDE. Zero → "scope holding, `success.md` current."
64
+ Check cumulative impact against the agreed scope and remaining capacity. Recommend a conversation as soon as delivery, cost, risk, or acceptance changes materially; one consequential request may suffice. No logged requests alone does not prove scope is holding.
63
65
 
64
66
  ## Worked example
65
67
 
66
68
  Acme, week 5. Nothing has been formally added, and the slice is a week late.
67
69
 
68
- The pattern shows in three receipts, not one argument: a "quick" finance CSV export (Jun 20, half a day, from Denise directly), retry-logic cleanup asked for mid-build (Jun 24, one day, Tom), and a dashboard tile "while you're in there" (Jun 27, half a day). Each was individually reasonable; together they are the slip.
70
+ The pattern shows in three requests: a "quick" finance CSV export (Jun 20, half a day, from Denise directly), retry-logic cleanup asked for mid-build (Jun 24, one day, Tom), and a dashboard tile "while you're in there" (Jun 27, half a day). Each sounds reasonable; their cumulative estimates explain part of the slip and need a scope decision.
69
71
 
70
- Three-bucket response, applied at the moment of the third ask rather than in a retrospective: the CSV export goes to Next with an accepted trade (it displaces the runbook polish), the retry cleanup goes to the kill list in `decisions.md` with the what-breaks reason, and the tile is absorbed because it is genuinely twenty minutes - logged anyway, since an unlogged absorption is the one that gets forgotten in the accumulation conversation.
72
+ Three-bucket response, applied while the requests can still be placed: the CSV export fits this phase only with an accepted trade (it displaces the runbook polish), the retry cleanup goes to the kill list in `decisions.md` with the what-breaks reason, and the tile is absorbed because it is genuinely twenty minutes - logged anyway, since an unlogged absorption is the one that gets forgotten in the accumulation conversation.
71
73
 
72
- That conversation happens with Priya at three receipts, with the dates on screen: "these are the four asks, here is the two days, here is what moved." Not a complaint - a decision she gets to make, with evidence, before the deadline makes it for her.
74
+ That conversation happens with Priya when the added work threatens the date, with the receipts on screen: "here are the asks, their estimated impact, and what moved." Not a complaint - a decision she gets to make, with evidence, before the deadline makes it for her.
73
75
 
74
76
  ## Principles
75
77
 
76
78
  - "Let me place it" is the phrase. Not "no," not "sure."
77
79
  - Every scope change gets a receipt. The receipt is the evidence.
78
- - Three unaddressed scope changes → accumulation conversation.
80
+ - Escalate material impact, not an arbitrary count of requests.
79
81
  - Scope creep kills engagements that technical failure couldn't.
80
82
  - `success.md` is a contract - update it explicitly or defend it.
81
83
  - The FDE who absorbs everything is liked for three weeks and blamed for three months.
@@ -0,0 +1,18 @@
1
+ # integrate - Prove the system boundary
2
+
3
+ **Enter when:** connecting an API, data source, SDK, event stream, tool, or service, or changing its contract.
4
+
5
+ Start from [task context](task-context.md). Permitted supplied context is enough; `.fde/` is optional. Use the customer's existing clients, authentication, fixtures, and diagnostic tools. Do not create another integration platform to make one connection.
6
+
7
+ ## Method
8
+
9
+ 1. Map producer, consumer, owner, direction, and side effects. Inspect the actual installed version and local implementation; verify uncertain behavior against current official documentation. Identify the relevant schema, authentication scopes, network boundary, and permitted test environment.
10
+ 2. Write the acceptance example: an input at the real boundary and the observable downstream result. Include a rejection or failure example. Separate configuration validity, successful authentication, transport connectivity, contract compatibility, and end-to-end behavior; none proves the next.
11
+ 3. Inspect credentials by presence and required scope without printing values. Use existing secret storage. Check data classification and retention before moving data; never pass raw `<private>` blocks into a model. Prefer sanitized or synthetic cases approved for the target environment.
12
+ 4. Implement the narrow adapter using native repository patterns. Validate external inputs and model outputs, bound timeouts and retries, preserve error context without leaking payloads, and handle cancellation. For writes, establish idempotency or duplicate detection before retries; for events, check ordering, replay, and poison messages as applicable.
13
+ 5. Exercise a permitted success case and relevant failures: denied access, malformed data, rate limit, timeout, duplicate delivery, or partial completion. Trace correlation IDs or safe evidence across both sides. A mock proves client behavior only; if live access is unavailable, report that gap instead of claiming an integration works.
14
+ 6. Check cleanup and recovery for test side effects. Use [verification](verification.md) for receipts and [review](review.md) for security or data-contract changes. Route deployment through [ship](ship.md) only when authorized.
15
+
16
+ ## Deliverable and acceptance
17
+
18
+ Return the boundary contract, changed paths, environment, evidence at each tested layer, and remaining dependencies with owners when known. Done requires the agreed end-to-end result or an explicit narrower agreed scope. Do not silently replace live acceptance with a stub. In engagement mode, update the terrain/delivery record with confirmed facts; otherwise return the receipt directly.
@@ -1,10 +1,12 @@
1
1
  # plan - Sequence the work
2
2
 
3
+ **Context:** apply [task context and evidence](task-context.md). Standalone planning evaluates supplied facts directly; it does not require initialized engagement records.
4
+
3
5
  **Enter when:** scope is understood and the work needs breaking down - a slice, a phase, or the whole delivery.
4
6
 
5
7
  **Read first:** `reality.md`, `success.md`, `terrain.md`, `stakeholders.md`. Load `business-case.md` if poc produced one. Not the full folder.
6
8
 
7
- **Before a new delivery plan or material scope change:** run `fde doctor --ready`. Missing binary success or a named customer-side signer blocks progression: review the proposed acceptance check and authority with the FDE first. Use a test/input and observable pass/fail under **Done when:** or **Acceptance check:**. A number, role, or successful demo alone is insufficient. Do not invent missing facts to pass lint. Routine reversible fixes within confirmed scope reuse the existing signer, acceptance criteria, and engineering plan; record verification without reopening settled decisions.
9
+ **On an initialized engagement, before a new delivery plan or material scope change:** run `fde doctor --ready`. For standalone planning, check the supplied outcome, scope, acceptance and authority directly; do not initialize records to run this validator. Missing binary success or a named customer-side signer blocks progression: review the proposed acceptance check and authority with the FDE first. Use a test/input and observable pass/fail under **Done when:** or **Acceptance check:**. A number, role, or successful demo alone is insufficient. Do not invent missing facts to pass lint. Routine reversible fixes within confirmed scope reuse the existing signer, acceptance criteria, and engineering plan; record verification without reopening settled decisions.
8
10
 
9
11
  ## Validation gate (confirm understanding, clarify where it elevates)
10
12
 
@@ -24,7 +26,7 @@ An FDE plan is not a sprint backlog. The technical sequence is the easy part. Th
24
26
 
25
27
  ## Method (you do this work)
26
28
 
27
- **0. Lock scope first.** Read `success.md`, `assumptions.md`, and the **Question** on `reality.md`. If out-of-scope is undefined, define it now with the FDE - a plan on undefined scope accumulates silent commitments. If any CRITICAL assumption is still `OPEN`, stop and run test-assumptions / discover before sequencing work. If `reality.md` has no Question, stop and finish discover - you are sequencing trivia.
29
+ **0. Lock scope first.** Read `success.md`, `assumptions.md`, and the **Question** on `reality.md`. Make the boundary explicit using the supplied request; ask if an ambiguity changes the commitment. Investigate a critical open assumption before planning work that depends on it. If the problem itself is unclear, use discovery for that gap; absent filenames do not block a plan supported by supplied facts.
28
30
 
29
31
  **Reuse check.** Before sequencing a build, compare the requested solution with the smallest existing capability or operating change that could satisfy the same acceptance test. Cite the relevant repo/config/workaround evidence. Record why reuse is sufficient or insufficient in `decisions.md`; include “no new code” when supported. A request for AI does not establish that a model is needed. If a host engineering pack already has an approved implementation plan, reference it from `decisions.md`; do not generate a parallel user-story backlog.
30
32
 
@@ -1,5 +1,7 @@
1
1
  # poc - Validate the solution
2
2
 
3
+ **Context:** apply [task context and evidence](task-context.md) before using the named records below.
4
+
3
5
  **Enter when:** a direction needs validating before committing real build time - POC, spike, show something, de-risk, pick between use cases. The output is something a sponsor can reject in a room this week, not a polished product.
4
6
 
5
7
  **Read first:** `context.md`, `reality.md`. Load `terrain.md` only if the prototype touches the existing codebase. If `terrain.md` **Data estate** has a Blocker source this prototype needs, stop - that is discover, not a day's demo.
@@ -8,7 +10,7 @@ A green check on synthetic data is not a validated solution. The person who can
8
10
 
9
11
  ## Method (you do this work)
10
12
 
11
- **0. Name the killer assumption.** With the FDE: "What's the belief that kills the project if it's wrong?" Prototype **that** - not the pretty demo. If `three-options` just ran: the cheapest test is for the recommended option first, unless they pick another.
13
+ **0. Name the killer assumption.** With the FDE: "What's the belief that kills the project if it's wrong?" Prototype **that** - not the pretty demo. If [three-options](three-options.md) just ran: the cheapest test is for the recommended option first, unless they pick another.
12
14
 
13
15
  **0b. Pass / fail before you build.** For the test you will run, write three lines in `prototype-log.md` first: what you will actually do (who you talk to, what you show, on whose screen); the result that **kills** this option; the result that keeps it alive. What you would learn either way. If every option's test would fail, name which `assumptions.md` block to reopen - do not invent a fourth playbook.
14
16
 
@@ -22,13 +24,13 @@ A green check on synthetic data is not a validated solution. The person who can
22
24
  - Latency: acceptable against real user expectations, not ideal conditions?
23
25
  - Is AI the right tool at all - or is this a data-quality or process problem wearing an AI costume?
24
26
 
25
- **4. Kill it immediately if:** the assumption is disproven · the customer ignores it (indifference is a signal, not neutrality) · 3 iterations and feedback isn't converging · it works but the customer can't explain or trust the output (unexplainable AI in a high-stakes context is not a solution). When killed: write down what was *learned*, not what was built. The learning is the asset.
27
+ **4. Decide at the agreed checkpoint.** Stop when the predeclared failure criterion is met or continued testing is unsafe. If feedback is absent or inconclusive, distinguish access or stakeholder availability from evidence against the assumption. Reassess the hypothesis, test design, and remaining timebox; extend only with a clear learning question and authorization for added scope or cost. Do not kill or continue solely because an iteration count was reached. If the customer cannot explain or trust high-stakes AI output, name the unresolved requirement and test whether it can be met. Record proceed / pivot / stop / inconclusive with evidence and what was learned.
26
28
 
27
29
  **5. Translate to business language** once validated: problem solved, cost of inaction, success in numbers, 2-3 trade-offs. Three sentences max for the stakeholder - can't say it in three, don't understand it yet.
28
30
 
29
31
  ## If proceeding to production
30
32
 
31
- Carry the hypothesis, test evidence, customer reaction, and remaining unknowns into the existing `plan` and `ship` workflow. A working demo does not establish production readiness or customer acceptance.
33
+ Carry the hypothesis, test evidence, customer reaction, and remaining unknowns into the existing [plan](plan.md) and [ship](ship.md) workflow. A working demo does not establish production readiness or customer acceptance.
32
34
 
33
35
  Inspect the prototype before deciding what to reuse. Keep components whose behavior and boundaries are suitable and tested. Replace or harden shortcuts that fail production requirements; rewrite only where the evidence justifies it. Record the decision and remaining work in `decisions.md`, rather than treating all prototype code as disposable or all working code as ready to deploy.
34
36
 
@@ -0,0 +1,18 @@
1
+ # qa - Exercise the changed journey
2
+
3
+ **Enter when:** a feature or fix needs behavioral verification through its real interface, especially UI, API, and multi-step workflows.
4
+
5
+ Use [task context](task-context.md) and the customer's existing browser, API, fixtures, and test tooling. `.fde/` is not a prerequisite. Respect the permitted environment and authority for every side effect.
6
+
7
+ ## Method
8
+
9
+ 1. Identify the changed journey, user roles, acceptance checks, and risk-bearing neighboring paths. Record the revision and environment. Use synthetic or sanitized fixtures with understood cleanup; do not borrow production data without permission.
10
+ 2. Run the normal journey from its real entry point through the expected result. Verify persisted or downstream state when the requirement includes it; a success toast alone does not prove a write succeeded.
11
+ 3. Select negative and boundary cases from the change: invalid input, empty/loading/error states, refresh/back navigation, retries, duplicates, permissions, or interrupted work. For UI changes, inspect relevant viewport sizes, keyboard access, focus, labels, and errors. Use a real browser for the affected journey.
12
+ 4. Inspect relevant console and network evidence. Distinguish a UI defect from a failed API or unavailable environment. Retain only privacy-safe screenshots and logs. Do not claim visual verification from code inspection or a generated screenshot that was not viewed.
13
+ 5. Report failures with steps, expected/actual result, revision/environment, evidence, and impact. If repair is authorized, use [debug](debug.md), then rerun the failed journey and affected neighbors. Keep unrelated findings separate from the change.
14
+ 6. Produce a [verification receipt](verification.md). State which roles, devices, environments, or data conditions remain untested. Do not weaken acceptance checks to make the run pass.
15
+
16
+ ## Deliverable and acceptance
17
+
18
+ Return checked journeys and observed results, reproducible defects, limitations, and remaining blockers. Done means the agreed behavioral checks passed under the stated conditions. A browser smoke test does not establish load capacity, security assurance, accessibility conformance, deployment, or customer acceptance by itself. When coordinated, append the evidence to the existing delivery record; standalone QA can return it directly.
@@ -1,77 +1,39 @@
1
- # review - Review the change
1
+ # review - Assess the actual change
2
2
 
3
- **Enter when:** a change needs review before merge - "is this safe," "does it match what we agreed." Their team commented on the PR: same skill. Comments are to check, not to obey.
3
+ **Enter when:** a diff, proposed merge, or review comment needs assessment against agreed behavior and constraints.
4
4
 
5
- **Read first:** `context.md`, `decisions.md`, `trust-profile.md`, `terrain.md`. Not `reality.md`/`stakeholders.md` - irrelevant to reviewing code against agreed scope.
5
+ Use [task context](task-context.md). Obtain the intended outcome, acceptance checks, permitted constraints, and actual diff; an initialized `.fde/` is unnecessary. Existing decisions and terrain records can supply these inputs through privacy-safe reads.
6
6
 
7
- Engagement review ≠ product-company review: a codebase you don't own, systems you can't fully see, a customer who can't afford a bad release.
7
+ ## Establish what was reviewed
8
8
 
9
- ## Pre-flight: is this reviewable?
9
+ Identify the repository, base and head revision, staged/unstaged changes, and relevant untracked files. Read applicable instructions and the full in-scope diff, then inspect callers and tests where needed. A committed-range diff alone omits working-tree edits. Record missing files or unavailable context as limitations.
10
10
 
11
- Check whether the diff has one agreed intent and can be reviewed with the available evidence. Split unrelated work; for a large cohesive change, separate generated/mechanical output from behavioral changes and review in bounded sections. Size is a warning to investigate, not a universal stop threshold.
11
+ State the review source: **self-check** when the author inspects their own work; **independent review** only when a separate person or agent actually examines it. A second pass by the same agent is still a self-check. Name the actual reviewer/source and reviewed revision when available. Do not fabricate a reviewer, dialogue, approval, or clean verdict. Use an available separate reviewer for substantial or risky changes when authorized; otherwise report the missing independent review and continue useful self-checks.
12
12
 
13
- ## Stage 1 - did we build what we agreed? (you do this work)
13
+ ## Check scope, then behavior
14
14
 
15
- Check the diff against the **one-line intent** in `decisions.md` / acceptance criteria - what was *explicitly decided*, not what seems right:
15
+ Compare each logical change with the agreed intent. Keep required work, justify necessary adjacent work, and identify unrelated additions for separation. Do not revert someone else's edits just to make the diff smaller. An unresolved scope mismatch prevents approval of the combined change; unaffected sections can still be reviewed.
16
16
 
17
- ```bash
18
- git diff <base>...HEAD --stat
19
- ```
17
+ Trace the changed path through its consumers and failure cases:
20
18
 
21
- For each touched path (or logical hunk), assign one verdict:
19
+ - **Correctness:** boundary conditions, stale state, concurrency, retries, cancellation, and error propagation.
20
+ - **Data and security:** input validation, authorization, migration compatibility, sensitive logs, and effects crossing tenant or trust boundaries.
21
+ - **Side effects:** writes, jobs, webhooks, notifications, and feature flags occur only under intended conditions; recovery accounts for already-completed effects.
22
+ - **AI behavior:** outputs remain untrusted, tools enforce allowed actions, and [eval evidence](eval-pack.md) covers the changed behavior and documented authority. Preserve privacy-safe source evidence and concise rationale, never hidden reasoning.
23
+ - **Operability:** observable failures, bounded resource use, meaningful checks, and a recovery path appropriate to the risk. Deployment readiness is assessed separately in [ship](ship.md).
22
24
 
23
- | Verdict | Meaning |
24
- |---------|---------|
25
- | **KEEP** | Required for the stated intent |
26
- | **JUSTIFY** | Adjacent but must ship now - write one sentence why, or SPLIT |
27
- | **SPLIT** | Real work for another PR / Next / kill list - do not merge with this slice |
28
- | **DROP** | Noise / drive-by - revert before Pass |
25
+ ## Findings and repair
29
26
 
30
- Also check:
31
- - Any sacred system from `trust-profile.md` touched?
32
- - Any sensitive data newly in scope?
33
- - Rollback path defined before build still honoured?
27
+ For each actionable finding give the path/line or precise location, concrete trigger, observed or reasoned failure, impact, and focused correction. Distinguish proven bugs from hypotheses that need a check. Prioritize release blockers over minor concerns; avoid speculative style work.
34
28
 
35
- **Stage 1 fails → stop** if any SPLIT/DROP remains, or JUSTIFY lacks a written sentence. Quality review on out-of-scope code is wasted work. Record the mismatch (and the KEEP/JUSTIFY/SPLIT/DROP tally) in `decisions.md`.
29
+ Validate incoming comments rather than obeying them automatically. Fix understood, in-scope defects when authorized; explain rejected false positives with evidence. Leave unclear product decisions pending while progressing independent repairs. Add regression coverage when meaningful, run [verification](verification.md), and review the changed result. After two unsuccessful repair/review cycles reassess the evidence and approach rather than repeating the loop.
36
30
 
37
- Stakeholder "also can you…" mid-build is `hold-scope` - different axis. This stage is **code vs claim**.
31
+ ## Deliverable and acceptance
38
32
 
39
- ## Stage 2 - is it safe to live with?
40
-
41
- Five dimensions, line-specific ("line 47 fails under concurrent writes - no lock"), never "could be better":
42
-
43
- - **Correctness** - does what it says; edge cases; error paths traced.
44
- - **Blast radius** - what breaks at 2am; downstream systems; failure mode loud (errors surface) or silent (data corrupts over time)?
45
- - **Security** - input validation at boundaries; no secrets in logs; no new attack surface; `trust-profile.md` sensitivity classes respected.
46
- - **Recovery** - can the documented, tested rollback or recovery path meet the agreed recovery-time and data-loss limits? For irreversible changes, require explicit authority, compatibility checks, and a tested restore/compensation or roll-forward plan; a code revert alone is not proof.
47
- - **AI policy & components** - human-review requirements honoured; model output treated as untrusted until validated; fallback exists; privacy-safe execution evidence retained under the client’s data policy (no secrets, raw private data, or hidden chain-of-thought); outputs bounded so a hallucination can't cascade; check applicable explanation and human-review requirements with the client’s responsible owner; provide source evidence and concise rationale without claiming access to hidden reasoning.
48
-
49
- **Structural pass on AI-heavy or data-touching changes:** migration compatibility and tested recovery · destructive SQL guarded · PII/PCI/PHI paths match `trust-profile.md` · side effects (flags, webhooks, emails, jobs) fire only when intended · magic strings that break on rename · new behaviour has a test or an explicit reason it can't yet. One line problem, one line fix.
50
-
51
- ## The review-fix loop (until clean)
52
-
53
- 1. Read the full diff before commenting.
54
- 2. Verdicts: **Stage 1: Pass / Blocked (reason)** · **Stage 2: Pass / Concerns (line-specific)**.
55
- 3. Fix only **real** findings tied to this change - no drive-by refactors. Reject false positives with one sentence why. Their comments are to check, not to obey. Restate each against the one-line intent and `trust-profile.md`. An unclear item waits for clarification; continue independent, understood fixes. If it breaks a signed constraint, a sacred system, or nothing calls it: one-sentence pushback, then wait.
56
- 4. Add or update a test per bug found where possible.
57
- 5. Re-run tests/typechecks - state what ran.
58
- 6. Re-review. If two repair/review cycles do not converge, reassess the evidence and approach; continue independent fixes and escalate concrete scope/product decisions.
59
-
60
- ## Before the PR - thinking for the next reader
61
-
62
- Code alone loses the "why." Before you call the change reviewable, run the **session digest** from the memory contract (SKILL.md Session digest): TL;DR, key decisions & rationale, scope + how you verified, gotchas. Confirm with the FDE, then write into `.fde/` - `decisions.md` / `delivery.md` / `context.md`. Reviewers (or Monday-you) should answer "why this approach?" from the fieldbook, not from a chat transcript. Do **not** dump agent logs into the product repo.
63
-
64
- ## Artifact
65
-
66
- **`decisions.md`** - each cycle logged: what was reviewed, flagged, fixed, verified. Stage 1 failures recorded with the specific mismatch. Digest decisions (with *why*) land here too when the slice ships.
67
-
68
- **`delivery.md`** - scope + verification from the digest when a PR is opening; intent-vs-diff receipt stays the ship gate.
33
+ Return scope, review source, findings by impact, verification evidence, and remaining limitations. Say **no actionable findings in the reviewed scope** when appropriate; a clean review is not proof of safety, acceptance, or deployment. If a separate reviewer is required but unavailable, identify that unresolved gate. Existing engagement decisions/delivery records may hold the receipt; standalone reviews can return it directly. No commit, PR, or publication is required by this method.
69
34
 
70
35
  ## Principles
71
36
 
72
- - Stage 1 before Stage 2. Wrong scope reviewed well is still wrong scope.
73
- - KEEP / JUSTIFY / SPLIT / DROP - every path gets a verdict; silent extras fail Stage 1.
74
- - Specific or silent - vague concerns waste everyone's time.
75
- - No viable tested recovery path = a release blocker.
76
- - A clean review proves this diff is safe as agreed - not that the feature was right.
77
- - Judgment in `.fde/` beats transcript in git.
37
+ - Findings need a concrete failure condition and a location in the reviewed change.
38
+ - Record the actual review source; self-check and independent review are different evidence.
39
+ - A clean reviewed diff does not grant release authority or establish customer acceptance.