fdeops 4.0.4 → 4.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (139) hide show
  1. package/AGENTS.md +1 -1
  2. package/README.md +33 -6
  3. package/bin/check.js +14 -23
  4. package/bin/generate-skills.js +71 -0
  5. package/bin/install.js +2 -1
  6. package/bin/skill-catalog.js +17 -0
  7. package/mcp/fdeops-ingest/package.json +2 -2
  8. package/package.json +4 -3
  9. package/plugin.json +2 -2
  10. package/skills/fde/SKILL.md +18 -10
  11. package/skills/fde/references/board-memo.md +1 -1
  12. package/skills/fde/references/build.md +20 -0
  13. package/skills/fde/references/business-case.md +9 -7
  14. package/skills/fde/references/close.md +5 -3
  15. package/skills/fde/references/debug.md +18 -0
  16. package/skills/fde/references/encode-pattern.md +10 -8
  17. package/skills/fde/references/eval-pack.md +16 -33
  18. package/skills/fde/references/hold-scope.md +9 -7
  19. package/skills/fde/references/integrate.md +18 -0
  20. package/skills/fde/references/plan.md +4 -2
  21. package/skills/fde/references/poc.md +5 -3
  22. package/skills/fde/references/qa.md +18 -0
  23. package/skills/fde/references/review.md +22 -60
  24. package/skills/fde/references/ship.md +42 -287
  25. package/skills/fde/references/task-context.md +12 -0
  26. package/skills/fde/references/test-assumptions.md +2 -2
  27. package/skills/fde/references/three-options.md +19 -27
  28. package/skills/fde/references/verification.md +31 -0
  29. package/skills/fde-build/SKILL.md +21 -0
  30. package/skills/fde-build/references/build.md +20 -0
  31. package/skills/fde-build/references/debug.md +18 -0
  32. package/skills/fde-build/references/eval-pack.md +26 -0
  33. package/skills/fde-build/references/integrate.md +18 -0
  34. package/skills/fde-build/references/qa.md +18 -0
  35. package/skills/fde-build/references/review.md +39 -0
  36. package/skills/fde-build/references/ship.md +73 -0
  37. package/skills/fde-build/references/task-context.md +12 -0
  38. package/skills/fde-build/references/verification.md +31 -0
  39. package/skills/fde-debug/SKILL.md +21 -0
  40. package/skills/fde-debug/references/build.md +20 -0
  41. package/skills/fde-debug/references/debug.md +18 -0
  42. package/skills/fde-debug/references/eval-pack.md +26 -0
  43. package/skills/fde-debug/references/integrate.md +18 -0
  44. package/skills/fde-debug/references/qa.md +18 -0
  45. package/skills/fde-debug/references/review.md +39 -0
  46. package/skills/fde-debug/references/ship.md +73 -0
  47. package/skills/fde-debug/references/task-context.md +12 -0
  48. package/skills/fde-debug/references/verification.md +31 -0
  49. package/skills/fde-discover/SKILL.md +21 -0
  50. package/skills/fde-discover/references/audit.md +71 -0
  51. package/skills/fde-discover/references/discover.md +254 -0
  52. package/skills/fde-discover/references/task-context.md +12 -0
  53. package/skills/fde-evaluate/SKILL.md +21 -0
  54. package/skills/fde-evaluate/references/build.md +20 -0
  55. package/skills/fde-evaluate/references/debug.md +18 -0
  56. package/skills/fde-evaluate/references/eval-pack.md +26 -0
  57. package/skills/fde-evaluate/references/integrate.md +18 -0
  58. package/skills/fde-evaluate/references/qa.md +18 -0
  59. package/skills/fde-evaluate/references/review.md +39 -0
  60. package/skills/fde-evaluate/references/ship.md +73 -0
  61. package/skills/fde-evaluate/references/task-context.md +12 -0
  62. package/skills/fde-evaluate/references/verification.md +31 -0
  63. package/skills/fde-feedback/SKILL.md +21 -0
  64. package/skills/fde-feedback/references/encode-pattern.md +96 -0
  65. package/skills/fde-feedback/references/task-context.md +12 -0
  66. package/skills/fde-handoff/SKILL.md +21 -0
  67. package/skills/fde-handoff/references/close.md +66 -0
  68. package/skills/fde-handoff/references/encode-pattern.md +96 -0
  69. package/skills/fde-handoff/references/task-context.md +12 -0
  70. package/skills/fde-integrate/SKILL.md +21 -0
  71. package/skills/fde-integrate/references/build.md +20 -0
  72. package/skills/fde-integrate/references/debug.md +18 -0
  73. package/skills/fde-integrate/references/eval-pack.md +26 -0
  74. package/skills/fde-integrate/references/integrate.md +18 -0
  75. package/skills/fde-integrate/references/qa.md +18 -0
  76. package/skills/fde-integrate/references/review.md +39 -0
  77. package/skills/fde-integrate/references/ship.md +73 -0
  78. package/skills/fde-integrate/references/task-context.md +12 -0
  79. package/skills/fde-integrate/references/verification.md +31 -0
  80. package/skills/fde-options/SKILL.md +21 -0
  81. package/skills/fde-options/references/business-case.md +90 -0
  82. package/skills/fde-options/references/task-context.md +12 -0
  83. package/skills/fde-options/references/test-assumptions.md +102 -0
  84. package/skills/fde-options/references/three-options.md +90 -0
  85. package/skills/fde-poc/SKILL.md +21 -0
  86. package/skills/fde-poc/references/audit.md +71 -0
  87. package/skills/fde-poc/references/build.md +20 -0
  88. package/skills/fde-poc/references/business-case.md +90 -0
  89. package/skills/fde-poc/references/debug.md +18 -0
  90. package/skills/fde-poc/references/discover.md +254 -0
  91. package/skills/fde-poc/references/eval-pack.md +26 -0
  92. package/skills/fde-poc/references/integrate.md +18 -0
  93. package/skills/fde-poc/references/plan.md +167 -0
  94. package/skills/fde-poc/references/poc.md +55 -0
  95. package/skills/fde-poc/references/qa.md +18 -0
  96. package/skills/fde-poc/references/review.md +39 -0
  97. package/skills/fde-poc/references/ship.md +73 -0
  98. package/skills/fde-poc/references/task-context.md +12 -0
  99. package/skills/fde-poc/references/test-assumptions.md +102 -0
  100. package/skills/fde-poc/references/three-options.md +90 -0
  101. package/skills/fde-poc/references/verification.md +31 -0
  102. package/skills/fde-qa/SKILL.md +21 -0
  103. package/skills/fde-qa/references/build.md +20 -0
  104. package/skills/fde-qa/references/debug.md +18 -0
  105. package/skills/fde-qa/references/eval-pack.md +26 -0
  106. package/skills/fde-qa/references/integrate.md +18 -0
  107. package/skills/fde-qa/references/qa.md +18 -0
  108. package/skills/fde-qa/references/review.md +39 -0
  109. package/skills/fde-qa/references/ship.md +73 -0
  110. package/skills/fde-qa/references/task-context.md +12 -0
  111. package/skills/fde-qa/references/verification.md +31 -0
  112. package/skills/fde-readout/SKILL.md +21 -0
  113. package/skills/fde-readout/references/board-memo.md +108 -0
  114. package/skills/fde-readout/references/business-case.md +90 -0
  115. package/skills/fde-readout/references/readout.md +69 -0
  116. package/skills/fde-readout/references/task-context.md +12 -0
  117. package/skills/fde-review/SKILL.md +21 -0
  118. package/skills/fde-review/references/build.md +20 -0
  119. package/skills/fde-review/references/debug.md +18 -0
  120. package/skills/fde-review/references/eval-pack.md +26 -0
  121. package/skills/fde-review/references/integrate.md +18 -0
  122. package/skills/fde-review/references/qa.md +18 -0
  123. package/skills/fde-review/references/review.md +39 -0
  124. package/skills/fde-review/references/ship.md +73 -0
  125. package/skills/fde-review/references/task-context.md +12 -0
  126. package/skills/fde-review/references/verification.md +31 -0
  127. package/skills/fde-scope/SKILL.md +21 -0
  128. package/skills/fde-scope/references/hold-scope.md +83 -0
  129. package/skills/fde-scope/references/task-context.md +12 -0
  130. package/skills/fde-ship/SKILL.md +21 -0
  131. package/skills/fde-ship/references/build.md +20 -0
  132. package/skills/fde-ship/references/debug.md +18 -0
  133. package/skills/fde-ship/references/eval-pack.md +26 -0
  134. package/skills/fde-ship/references/integrate.md +18 -0
  135. package/skills/fde-ship/references/qa.md +18 -0
  136. package/skills/fde-ship/references/review.md +39 -0
  137. package/skills/fde-ship/references/ship.md +73 -0
  138. package/skills/fde-ship/references/task-context.md +12 -0
  139. package/skills/fde-ship/references/verification.md +31 -0
@@ -0,0 +1,73 @@
1
+ # ship - Deliver and release with evidence
2
+
3
+ **Enter when:** an implemented increment needs a delivery checkpoint, deployment, or wider rollout. Use [build](build.md) for implementation; an untested business or technical assumption needs an experiment before a release claim.
4
+
5
+ Start from [task context](task-context.md). Standalone work uses supplied permitted context and a release receipt; it does not require `.fde/` initialization. In engagement mode, use privacy-safe views of confirmed context, decisions, terrain, success, delivery, applicable trust constraints, and AI evaluation evidence. Never load raw `<private>` blocks into a model. The `fde` CLI remains local-only; deployment uses the customer's authorized tools, never a new network capability inside `fde`.
6
+
7
+ ## Establish the delivery contract
8
+
9
+ Identify the exact outcome, acceptance check, scope, affected users/systems, target environment, recovery mechanism, and who or what is authorized to accept and release it. Reuse confirmed authority and checks for routine work. Do not invent missing signers, permissions, measurements, or acceptance.
10
+
11
+ For an initialized engagement, run `fde doctor --ready` before a new delivery plan or material scope change. Missing binary success or a named customer-side signer blocks that planning progression until resolved. A passing doctor validates record structure, not connectivity, release readiness, or customer acceptance. Standalone work evaluates the supplied contract directly.
12
+
13
+ A customer delivery checkpoint must let the agreed decision-maker replay and reject the acceptance check through an interface they operate. Prefer their staging; otherwise use an agreed representative environment and disclose its owner and limitations. Local green proves only the local run. Routine fixes may share an agreed checkpoint; no fixed number of changes forces a ceremony.
14
+
15
+ ## Prepare a reviewable increment
16
+
17
+ 1. Inspect repository instructions, working tree, overlapping work, and the complete intended release diff. Include working-tree changes when testing an uncommitted candidate. Preserve unrelated work; separate unintended behavior before release.
18
+ 2. Identify applicable before-state evidence and the changed outcome. Name dependencies, stop conditions, and irreversible effects. Existing applicable evidence may be reused with attribution, never represented as a fresh run.
19
+ 3. Complete [verification](verification.md) and [review](review.md), proportional to the change and repository requirements. A self-check is not an independent review. Record command, revision, environment, date, result, and unrun checks. Exercise relevant operating exceptions and fallback paths, not just the happy path.
20
+ 4. For AI behavior, obtain a scoped [eval verdict](eval-pack.md) for the candidate and applicable controls. A permitted bounded automation workflow remains permitted within its documented limits. A missing evaluation, failed critical case, or unknown action authority prevents release of that path; non-AI changes record eval as not applicable.
21
+
22
+ ## Release gate
23
+
24
+ Before deployment, establish these facts from existing evidence or a necessary check. Missing material evidence blocks the dependent release step; continue independent preparation. Do not ask again for approval already provided within the same scope.
25
+
26
+ | Dimension | Required evidence |
27
+ |-----------|-------------------|
28
+ | Candidate | Exact revision/artifact, intended diff, dependencies, applicable required checks passing; no skipped failure presented as green |
29
+ | Target and access | Service/account/region, environment, authorized deployment identity and mechanism, secret provisioning without revealing values |
30
+ | Acceptance | Replayable check and agreed decision-maker/mechanism; record actual acceptance separately from readiness |
31
+ | Data and policy | Permitted data, applicable security/residency/change-window requirements, necessary approvals already recorded or obtained |
32
+ | Recovery | Applicable tested rollback, restore, compensation, or roll-forward within agreed recovery-time/data-loss limits; explicit authority for irreversible effects |
33
+ | Operations | Named release/recovery owner, runbook appropriate to risk, health and business signals, stop thresholds, observation coverage |
34
+ | AI, when applicable | Current applicable SHIP eval evidence, critical failures zero, enforced action boundary and required human review or documented bounded automation |
35
+
36
+ Check migration compatibility, old/new version coexistence, delayed jobs, caches, and already-emitted side effects where relevant. A code revert does not undo data loss or external writes. Reuse drill evidence only when the mechanism and relevant conditions are unchanged, explaining applicability. If recovery is only a plan, exercise it in a permitted representative environment before release.
37
+
38
+ Use the repository's existing secret scanning and security checks; avoid diagnostic commands that print credential matches. Retain sanitized references to results. Resolve material evidence gaps or obtain an explicit, authorized narrowing of the release; do not average critical blockers into a readiness score.
39
+
40
+ For a coordinated engagement, also connect the release to the agreed value bucket and baseline/target, and record a dated receipt for the affected operating path. An unmeasured result remains pending with a measurement next step; do not invent realized value to pass a gate.
41
+
42
+ ## Deploy within authority
43
+
44
+ Execute only when the requested workflow authorizes deployment to this target and the applicable gates are met. Otherwise leave a concrete release candidate, exact deployment/recovery instructions, evidence, and the remaining authorization for review. A permission to implement or test is not permission to publish.
45
+
46
+ Use the customer's established pipeline and rollout mechanism. Select canary, staged exposure, blue/green, or direct rollout according to actual risk and platform capabilities; do not impose a universal cohort sequence. Define advance/abort thresholds and observation window before starting. If another operator must execute, record their handoff and report deployment pending until there is evidence it happened.
47
+
48
+ During rollout inspect health, errors, key user behavior, and side-effect integrity. Halt expansion on breached thresholds or critical harm and apply authorized containment/recovery. Do not continue merely because the deploy command exited successfully.
49
+
50
+ ## Verify operation and hand off
51
+
52
+ Run permitted smoke and acceptance checks against the deployed candidate. Record deployment identity/time, observed signals, sample/window, failures, recovery actions, and remaining gaps. Define the pulse: metric, cadence, threshold, owner, and response. For AI, include permitted output sampling and drift/action-boundary monitoring.
53
+
54
+ Before wider exposure, verify expected load/cost, data pipeline behavior, ownership, support, and applicable governance for the proposed audience. Choose expansion conditions from evidence; a successful pilot does not establish readiness for an arbitrary larger scale. Measure adoption against the eligible users, expected workflow frequency, and agreed observation window; investigate misses without guessing their cause.
55
+
56
+ ## Receipt and completion
57
+
58
+ Keep these claims separate: implemented, verified, deployed, measured outcome, and accepted. Include the candidate, target, command/pipeline, applicable checks and unrun checks, review source, evaluation where needed, authority source, recovery evidence, observation, and next owner/action. Attribute acceptance to its actual source and scope. A staging measurement is not production value, and a commit is not deployment.
59
+
60
+ In engagement mode, write confirmed implementation/decisions and delivery receipts under the existing record rules. Standalone work returns the same receipt or uses the repository's permitted release record. Committing, pushing, opening a PR, publishing, and notifying others are actions governed by the user's workflow, not mandatory steps imposed by this method.
61
+
62
+ ## Worked example
63
+
64
+ A freight team agrees that dispatchers can retry a failed export once without creating a duplicate shipment. The change uses the existing queue and ops screen. The developer records a failing duplicate-delivery case, implements idempotency, and passes the relevant checks on a named candidate. QA observes both the retry status and the single downstream record on permitted staging fixtures. A separate reviewer examines the queue race; the receipt names that review and the revision.
65
+
66
+ The export service has an approved staged-release workflow. Its owner reuses a recent recovery drill because the queue format and recovery mechanism are unchanged, recording that applicability. Deployment stops if duplicate records appear or the agreed error threshold is crossed. The authorized rollout completes, production smoke checks pass, and the dispatcher accepts the specified retry behavior with a dated source. The operating-cost benefit remains pending until the agreed measurement window closes. Implementation, deployment, acceptance, and measured value have different evidence. In the existing engagement, `decisions.md` records the agreed behavior and `delivery.md` holds the release receipt; standalone work returns those facts directly.
67
+
68
+ ## Principles
69
+
70
+ - Release the reviewed candidate with applicable evidence and documented authority.
71
+ - Test recovery against the effects that actually persist beyond a code revert.
72
+ - Keep missing evidence visible and distinguish local, staging, and production claims.
73
+ - Expansion follows observed acceptance and operating limits; fixed ceremonies cannot replace them.
@@ -0,0 +1,12 @@
1
+ # Task context and evidence
2
+
3
+ Use this contract for standalone methods and methods routed through `@fde`.
4
+
5
+ - **Standalone work:** use the supplied, permitted facts, notes, code, and artifacts. A client name, `.fde/` directory, or initialized engagement is not a prerequisite. Do not bootstrap records merely to run a method. Ask only for missing information or authority that changes the next action; mark other gaps as unknown.
6
+ - **Artifact names are destinations:** names such as `success.md`, `decisions.md`, and `delivery.md` identify relevant evidence and, when bound, record destinations. If absent, use supplied facts and return the requested draft or result in the current workspace or conversation. Do not invent files or require initialization to complete useful work.
7
+ - **Bound engagement:** honor the current client binding and constraints. Before reading records, run `fde privacy` to verify masking support. Obtain a fresh, identity-matching sanitized `fde resume` packet for this task (or reuse a fresh session-hook packet); retrieve missing evidence with targeted `fde recall <topic>`. Use bounded `fde handoff` for transfer work. Refresh after binding, masking, or record changes. Never substitute raw `.fde/` reads, private blocks, masking dictionaries, or full transcripts. If the CLI is unavailable, use only permitted supplied excerpts and report the context limitation.
8
+ - **Authority:** continue reversible work within authorized scope. Reuse prior authorization when it covers the specific action. Show consequential engagement-record judgments and uncertainties for confirmation before saving unless already explicitly confirmed. New scope, acceptance changes, production actions, exports, and external messages need the applicable authority; a method invocation alone does not supply it. Keep one customer's writes in that customer's record.
9
+ - **Evidence:** distinguish supplied facts, estimates, hypotheses, and unknowns. Cite actual sources; a log date is not attribution. Never invent a source, signer, signature, customer reaction, or acceptance. Keep outcomes **promised → measured → accepted** distinct, and implementation, verification, deployment, and customer acceptance separate. Missing evidence means unproven, not an observed failure.
10
+ - **Data boundary:** use only data permitted by the customer's AI policy; clarify unknown policy before loading their code or data. Never load `<private>` content into a model. Cross-client comparison and exporting reusable material require permission and removal of customer-identifying or confidential content; anonymization alone does not grant permission.
11
+
12
+ Apply the selected method to this context. Follow its linked supporting methods only when needed; do not restart discovery or repeat already answered questions.
@@ -0,0 +1,31 @@
1
+ # verification - Make a claim replayable
2
+
3
+ **Enter when:** reporting completion, evaluating an acceptance check, handing work to a reviewer, or preparing a release.
4
+
5
+ Use [task context](task-context.md). This method returns evidence directly or writes an existing permitted task/engagement record; it never requires `.fde/` initialization.
6
+
7
+ ## Method
8
+
9
+ 1. Translate each claim into the observation that would support or reject it. Reuse agreed acceptance criteria and required repository checks. Select focused checks for changed behavior before broadening to release requirements.
10
+ 2. Identify the actual repository commands, fixtures, runtime, and environment. Read command behavior before executing it, especially when it can write externally. Use authorized environments and avoid leaking secrets through logs or diagnostic commands.
11
+ 3. Run the checks and inspect results, including exit status and relevant output. A running job, test discovery, a mocked response, and a successful real request are different evidence. Record asynchronous completion before claiming success.
12
+ 4. Bind evidence to the tested revision and working tree. For uncommitted changes record the base revision plus changed paths and an available diff digest or snapshot identifier. For browser/manual checks record the steps, inputs, observed result, and inspected evidence.
13
+ 5. After a change, rerun checks whose behavior or assumptions were affected. Reuse prior evidence only when the relevant code, dependencies, data, and environment remain applicable; cite the original run and reason. Never imply reused evidence was rerun.
14
+ 6. Label every required check **passed**, **failed**, **blocked**, or **not run**. Include why blocked/not run, impact, and next step. Missing evidence is unproven; it is not an observed failure or a pass.
15
+
16
+ ## Receipt
17
+
18
+ Use one compact entry per check or a table with these fields:
19
+
20
+ - Claim / acceptance check and expected result.
21
+ - Exact command and working directory, or manual journey and inputs.
22
+ - Environment, runtime/tool versions when relevant, and fixture/data source.
23
+ - Revision plus working-tree identity; run date/time.
24
+ - Observed result and exit status where available; safe evidence location.
25
+ - Status, limitations, unrun checks, and next step.
26
+
27
+ Keep implementation, verification, deployment, measured outcome, and customer acceptance distinct. A local pass supports the tested local behavior. An acceptance claim needs an attributed source from the agreed decision-maker or agreed acceptance mechanism. Record no raw `<private>` blocks, credentials, or hidden reasoning.
28
+
29
+ ## Acceptance
30
+
31
+ A completion statement cites applicable evidence for its claims and explicitly names material gaps. If required checks fail, investigate or report the blocker; never skip them, edit expectations, or relabel the scope without authority to obtain a green result.
@@ -0,0 +1,21 @@
1
+ ---
2
+ name: fde-feedback
3
+ description: Assess a field lesson for reuse or product feedback without exposing customer context. Use for recurring deployment lessons; distinguish a hypothesis from a validated pattern.
4
+ ---
5
+
6
+ # fde-feedback
7
+
8
+ <!-- Generated by bin/generate-skills.js; edit the canonical references and catalog. -->
9
+
10
+ ## Purpose
11
+
12
+ Assess a field lesson for reuse or product feedback without exposing customer context. Use for recurring deployment lessons; distinguish a hypothesis from a validated pattern.
13
+
14
+ Read [the task context contract](references/task-context.md), then [the method](references/encode-pattern.md). Load further references only when the task needs them. Everything linked is included in this skill; no other skill pack is required.
15
+
16
+ ## Principles
17
+
18
+ - Work directly from the supplied permitted context. Standalone work does not require an engagement folder or initialization. Record filenames in the method are optional persistence destinations when no engagement is bound.
19
+ - If called by @fde, reuse its current sanitized packet and scope. Do not restart setup, discovery or questions already answered.
20
+ - The task context contract controls persistence and authority in both modes. Preserve unknowns and distinguish implementation, verification, deployment and acceptance.
21
+ - Use the customer's repository instructions and available tools. Report a missing capability or unrun check honestly; do not claim that installing a skill provisions infrastructure.
@@ -0,0 +1,96 @@
1
+ # encode-pattern - Encode the pattern
2
+
3
+ **Context:** apply [task context and evidence](task-context.md) before using the named records below.
4
+
5
+ **Enter when:** the engagement is closing and reusable patterns exist, a technique worked well and will apply to future clients, the FDE notices themselves doing the same thing on a second engagement, or close identified a pattern worth preserving.
6
+
7
+ **Read first:** `decisions.md`, `reality.md`, `delivery.md`, `retrospectives/`, `context.md`. Patterns live in what was *done*, not what was planned.
8
+
9
+ The difference between a 5-year FDE and a 15-year FDE is not talent - it's encoded patterns. The 15-year FDE walks into a new engagement and recognises the situation in minutes because they've seen it before, named it, and know the move. Pattern extraction turns experience into reusable intelligence.
10
+
11
+ ## Method (you do this work)
12
+
13
+ **1. Identify the pattern candidates.** Scan the engagement for things that:
14
+
15
+ | Signal | Example |
16
+ |--------|---------|
17
+ | Worked well and would work again in a similar situation | The "show the workaround first" approach to earning ops team trust |
18
+ | Failed and the failure mode is predictable | The "refactor before understanding" mistake on legacy codebases |
19
+ | Was discovered late and should have been discovered early | The hidden cron job that broke the migration - always ask about cron jobs |
20
+ | Required a workaround that others would face too | The compliance dance for getting AI tools approved in regulated environments |
21
+ | Involved a political dynamic that repeats | The passed-over internal team dynamic - present in every engagement with external FDEs |
22
+
23
+ **2. Write the pattern in a transferable format.** Each pattern must be usable by a future FDE who has never heard of this engagement:
24
+
25
+ ```markdown
26
+ ## Pattern: <name>
27
+
28
+ ### Situation
29
+ <When does this pattern apply? What does the FDE see/hear that triggers recognition?>
30
+
31
+ ### The move
32
+ <What to do, specifically. Not advice - steps.>
33
+
34
+ ### Why it works
35
+ <The mechanism - why this approach succeeds where the obvious approach fails.>
36
+
37
+ ### Watch out for
38
+ <The failure mode or edge case that makes the pattern not apply.>
39
+
40
+ ### Evidence
41
+ <Permitted source, what happened, measured result, and limits. Keep identifying evidence in its original customer record.>
42
+ ```
43
+
44
+ **3. The pattern quality test.** Before encoding:
45
+
46
+ | Test | Pass | Fail |
47
+ |------|------|------|
48
+ | **Transferable?** | Another FDE could apply this without context from this engagement | Only makes sense if you know the specific client |
49
+ | **Specific enough?** | Contains concrete steps, not just principles | "Build trust" / "Communicate well" - too vague to act on |
50
+ | **Repeatable?** | Applies to a class of situations, not just this one | Only worked because of a unique circumstance |
51
+ | **Falsifiable?** | You can tell when the pattern is working or not | No way to measure whether applying it helped |
52
+ | **Monday bag?** | You would refuse the next similar embed without this in your bag - a named move plus an artifact you can drop on day one (pipe questions, CAB dance, eval golden shape, floor-drill script) | "We learned to communicate." Patterns that only live in this client's `.fde/` do not compound |
53
+
54
+ **4. Classify by stage.** Patterns sort into the same stages as the skills:
55
+
56
+ | Stage | Pattern type | Example |
57
+ |--------|-------------|---------|
58
+ | **Land** | Political / relational | "The passed-over team warm-up protocol" |
59
+ | **Discover** | Investigative / analytical | "The cron-job discovery checklist for legacy systems" |
60
+ | **Plan** | Structural / strategic | "The three-option presentation for nervous sponsors" |
61
+ | **Ship** | Technical / safety | "The Strangler Fig on financial transaction code" |
62
+ | **Outcome** | Operational / process | "The regulated-environment change-approval timeline buffer" |
63
+ | **Close** | Knowledge / handoff | "The 2am document format that actually gets used" |
64
+
65
+ **5. Version and evolve.** Patterns are living documents:
66
+
67
+ - First use: **v0.1** - hypothesis based on one engagement
68
+ - Later uses: record context, observed results, failures, and refinements. Repetition supplies evidence; it does not automatically validate the pattern. Promote a version when a substantive revision warrants it, not at a fixed use count.
69
+ - After modification: increment minor version with what changed and why
70
+ - After contradiction: note the counter-example, adjust the "watch out for" section
71
+
72
+ **6. Cross-engagement pattern mining.** Only when explicitly authorized for the named engagements and permitted by each customer's data policy. Keep records separate; use sanitized CLI packets or targeted recall in each authorized context, never raw file comparisons. Otherwise extract a candidate from the current permitted context only.
73
+
74
+ - Compare permitted problem summaries - do the same problems recur?
75
+ - Compare permitted decision summaries - are the same decisions being made?
76
+ - Compare permitted lessons - are the same lessons being learned twice?
77
+
78
+ A pattern learned twice is a process failure. Encoding it prevents the third time.
79
+
80
+ ## Artifact
81
+
82
+ **`patterns.md`** - candidates and evidence for this engagement, indexed by stage and situation trigger. A shared library is a separate, authorized export: remove names, identifiers, distinctive operational details, secrets, and confidential code or data. Keep source receipts in the original record and export only permitted generalizations.
83
+
84
+ **`retrospectives/YYYY-MM-DD-<engagement>.md`** - reference to which patterns were extracted from this engagement.
85
+
86
+ ## Checkpoint
87
+
88
+ Present the extracted patterns to the FDE: "From this engagement, I've identified N patterns worth encoding. The highest-value one is <name> because <it will apply to future engagements in these situations>." Confirm the pattern is accurate - the FDE's field judgment outranks the analysis.
89
+
90
+ ## Principles
91
+
92
+ - If you did it twice, encode it. The same lesson learned three times is a failure.
93
+ - Patterns are steps, not principles. "Build trust" isn't a pattern; "fix a small visible bug on day one" is.
94
+ - Every pattern needs a situation trigger - the FDE must recognise when it applies.
95
+ - Version substantive changes. State the evidence and limits; repeated use is not automatic confirmation.
96
+ - The pattern library is the FDE's compound interest. It's what separates 5 years of experience from 1 year repeated 5 times.
@@ -0,0 +1,12 @@
1
+ # Task context and evidence
2
+
3
+ Use this contract for standalone methods and methods routed through `@fde`.
4
+
5
+ - **Standalone work:** use the supplied, permitted facts, notes, code, and artifacts. A client name, `.fde/` directory, or initialized engagement is not a prerequisite. Do not bootstrap records merely to run a method. Ask only for missing information or authority that changes the next action; mark other gaps as unknown.
6
+ - **Artifact names are destinations:** names such as `success.md`, `decisions.md`, and `delivery.md` identify relevant evidence and, when bound, record destinations. If absent, use supplied facts and return the requested draft or result in the current workspace or conversation. Do not invent files or require initialization to complete useful work.
7
+ - **Bound engagement:** honor the current client binding and constraints. Before reading records, run `fde privacy` to verify masking support. Obtain a fresh, identity-matching sanitized `fde resume` packet for this task (or reuse a fresh session-hook packet); retrieve missing evidence with targeted `fde recall <topic>`. Use bounded `fde handoff` for transfer work. Refresh after binding, masking, or record changes. Never substitute raw `.fde/` reads, private blocks, masking dictionaries, or full transcripts. If the CLI is unavailable, use only permitted supplied excerpts and report the context limitation.
8
+ - **Authority:** continue reversible work within authorized scope. Reuse prior authorization when it covers the specific action. Show consequential engagement-record judgments and uncertainties for confirmation before saving unless already explicitly confirmed. New scope, acceptance changes, production actions, exports, and external messages need the applicable authority; a method invocation alone does not supply it. Keep one customer's writes in that customer's record.
9
+ - **Evidence:** distinguish supplied facts, estimates, hypotheses, and unknowns. Cite actual sources; a log date is not attribution. Never invent a source, signer, signature, customer reaction, or acceptance. Keep outcomes **promised → measured → accepted** distinct, and implementation, verification, deployment, and customer acceptance separate. Missing evidence means unproven, not an observed failure.
10
+ - **Data boundary:** use only data permitted by the customer's AI policy; clarify unknown policy before loading their code or data. Never load `<private>` content into a model. Cross-client comparison and exporting reusable material require permission and removal of customer-identifying or confidential content; anonymization alone does not grant permission.
11
+
12
+ Apply the selected method to this context. Follow its linked supporting methods only when needed; do not restart discovery or repeat already answered questions.
@@ -0,0 +1,21 @@
1
+ ---
2
+ name: fde-handoff
3
+ description: Transfer operation of a customer deployment with ownership, evidence and a tested support path. Use for handoff or an engineer rotation, not merely code delivery.
4
+ ---
5
+
6
+ # fde-handoff
7
+
8
+ <!-- Generated by bin/generate-skills.js; edit the canonical references and catalog. -->
9
+
10
+ ## Purpose
11
+
12
+ Transfer operation of a customer deployment with ownership, evidence and a tested support path. Use for handoff or an engineer rotation, not merely code delivery.
13
+
14
+ Read [the task context contract](references/task-context.md), then [the method](references/close.md). Load further references only when the task needs them. Everything linked is included in this skill; no other skill pack is required.
15
+
16
+ ## Principles
17
+
18
+ - Work directly from the supplied permitted context. Standalone work does not require an engagement folder or initialization. Record filenames in the method are optional persistence destinations when no engagement is bound.
19
+ - If called by @fde, reuse its current sanitized packet and scope. Do not restart setup, discovery or questions already answered.
20
+ - The task context contract controls persistence and authority in both modes. Preserve unknowns and distinguish implementation, verification, deployment and acceptance.
21
+ - Use the customer's repository instructions and available tools. Report a missing capability or unrun check honestly; do not claim that installing a skill provisions infrastructure.
@@ -0,0 +1,66 @@
1
+ # close - Transfer operations
2
+
3
+ **Context:** apply [task context and evidence](task-context.md) before using the named records below.
4
+
5
+ **Enter when:** the engagement is ending - the customer team must run this without the FDE.
6
+
7
+ **Read first:** bounded `fde handoff` or `fde resume`, then targeted `fde recall` for missing evidence. Build the full picture through relevant excerpts, not a full-directory load. Consult `terrain.md` only for the code paths needed by the successor.
8
+
9
+ The engagement doesn't end at ship. It ends when the customer can maintain what was built without calling.
10
+
11
+ ## Method (you do this work, with the FDE's answers)
12
+
13
+ **0. The opening question:** "What will bite them when you're gone?" Their answer shapes everything written below.
14
+
15
+ **1. The retrospective.** Work through, blame-free and specific:
16
+ - Did the real problem match the brief? (Compare `brief.md` vs `reality.md` - you have the receipts.)
17
+ - Which trust moments mattered?
18
+ - What did the codebase teach that `terrain.md` didn't know at the start?
19
+ - Which risk almost became real?
20
+ - AI components: did they behave in production? What failure modes did the prototype hide? Is the team equipped to maintain them?
21
+
22
+ **1b. Value + receipts close gate (refuse green close if any fail):**
23
+ - Primary value bucket in `success.md` matches what the sponsor funded; at least one ledger row has **Measured** (not forever-`pending`) with evidence **and a named customer-side owner in Accepted by** for that bucket - or the retrospective explicitly records “not measured; sponsor accepted pending.” A measured-but-unaccepted number closes as `claimed`; say so in the retrospective rather than closing green on arithmetic nobody signed.
24
+ - Audit receipt exists for the final shipped path (exceptions/operating map walked; cite file).
25
+ - Eval receipt: **n/a if no AI**, else final scoped eval result + operating owner and required human-review or bounded-automation authority recorded; kill switch / fallback named in `handoff.md`.
26
+ - One line in the retrospective: which bucket moved, by how much, vs baseline.
27
+
28
+ **2. The pattern.** Anything that happened here and will happen again - a compliance approach, a migration pattern, a stakeholder dynamic - gets encoded for reuse. Use [encode-pattern](encode-pattern.md) to distinguish candidate patterns from supported ones and protect customer data.
29
+
30
+ **3. The handoff.** Operational knowledge for the person woken at 2am, not technical documentation: the 3 things that will break and the fix for each · who holds the tribal knowledge · what each alert means · deploy and rollback in plain language. AI components additionally: model version, what normal output looks like (so drift is recognisable), fallback behaviour, who owns retraining, **how to disable the AI path without taking down the feature** - without this the team turns it off at the first misbehaviour and it stays off.
31
+
32
+ **4. Transformation engagements - four extra answers in `handoff.md`:**
33
+ - Who owns AI governance after the FDE leaves? (Who can pull a model from production?)
34
+ - The retraining trigger, exactly: "precision < 0.82 on validation for 3 consecutive weeks → <owner> retrains." A number, a condition, an owner - not "when performance drops."
35
+ - The operating model at scale: who coordinates twenty use cases across five teams?
36
+ - Decision authority for new use cases: intake, risk assessment, approver.
37
+
38
+ ## Artifact
39
+
40
+ **`retrospectives/YYYY-MM-DD-<engagement>.md`** - one file per close (separate files make cross-engagement patterns scannable). **`patterns.md`** - reusable patterns extracted. **`handoff.md`** - the 2am document.
41
+
42
+ ## Checkpoint
43
+
44
+ **Check the handoff as a lookup tool.** Give the intended operator one realistic task, such as finding the owner and recovery steps for a failed run. Can they locate the answer and its source in the permitted handoff without your explanation? A reader finding the instructions is not proof they can execute them; verify operation separately in the agreed safe environment. Correct the passage they could not use, rather than adding a longer introduction.
45
+
46
+ If the operator is unavailable, a fresh reviewer can attempt the same lookup using only the permitted draft and task. Report this as a simulated clarity check, not operator validation, customer approval, or a green close. Claim independent review only if a separate reviewer actually performed it; identify the reviewer and evidence available. If none is available, perform a labeled self-check and report independent review as unperformed. Use one focused pass for a consequential handoff; do not add a committee or a second approval ritual.
47
+
48
+ Direct assessment to the FDE: did the engagement achieve `success.md` · 2-3 lessons that matter · is the pattern worth encoding · is the handoff complete or where are the gaps. Also: value bucket + audit receipt green; eval **n/a or green**. Pending Measured without sponsor acceptance = gap, not green close. Honest - a gap named now is cheaper than a callback in six weeks.
49
+
50
+ ## Worked example
51
+
52
+ Acme, twelve weeks in, the FDE is rolling off.
53
+
54
+ Retrospective against the receipts: `brief.md` asked for monitoring, `reality.md` proved it was ownership - and the delta is the most useful paragraph in the file, because it is exactly the argument the next engagement will need.
55
+
56
+ The close gate bites in a useful way. The ledger shows detection at 12 minutes measured across two real incidents, but **Accepted by** is empty - Marco confirmed it in Slack, Denise (finance) never did, and Denise is whose escalation started the engagement. So it closes as `claimed` with a one-line retrospective note and a named next step, rather than a green close on a number nobody with budget agreed to.
57
+
58
+ `handoff.md` is written for the person woken at 2am: the three things that break, what the page means, how to re-run manually the way Marco does, and who holds the tribal knowledge (Raj, who built the original job - credited, because he protects it now). `patterns.md` gets *"unowned job" presents as "unmonitored job"* - it has now happened twice.
59
+
60
+ ## Principles
61
+
62
+ - Done = the customer operates without you.
63
+ - No named value bucket moved (or sponsor-accepted pending) = not a green close.
64
+ - The retrospective is an investment in the next engagement, not a post-mortem.
65
+ - Encode what repeated. The same lesson learned twice is a process failure.
66
+ - Write the handoff for 2am.
@@ -0,0 +1,96 @@
1
+ # encode-pattern - Encode the pattern
2
+
3
+ **Context:** apply [task context and evidence](task-context.md) before using the named records below.
4
+
5
+ **Enter when:** the engagement is closing and reusable patterns exist, a technique worked well and will apply to future clients, the FDE notices themselves doing the same thing on a second engagement, or close identified a pattern worth preserving.
6
+
7
+ **Read first:** `decisions.md`, `reality.md`, `delivery.md`, `retrospectives/`, `context.md`. Patterns live in what was *done*, not what was planned.
8
+
9
+ The difference between a 5-year FDE and a 15-year FDE is not talent - it's encoded patterns. The 15-year FDE walks into a new engagement and recognises the situation in minutes because they've seen it before, named it, and know the move. Pattern extraction turns experience into reusable intelligence.
10
+
11
+ ## Method (you do this work)
12
+
13
+ **1. Identify the pattern candidates.** Scan the engagement for things that:
14
+
15
+ | Signal | Example |
16
+ |--------|---------|
17
+ | Worked well and would work again in a similar situation | The "show the workaround first" approach to earning ops team trust |
18
+ | Failed and the failure mode is predictable | The "refactor before understanding" mistake on legacy codebases |
19
+ | Was discovered late and should have been discovered early | The hidden cron job that broke the migration - always ask about cron jobs |
20
+ | Required a workaround that others would face too | The compliance dance for getting AI tools approved in regulated environments |
21
+ | Involved a political dynamic that repeats | The passed-over internal team dynamic - present in every engagement with external FDEs |
22
+
23
+ **2. Write the pattern in a transferable format.** Each pattern must be usable by a future FDE who has never heard of this engagement:
24
+
25
+ ```markdown
26
+ ## Pattern: <name>
27
+
28
+ ### Situation
29
+ <When does this pattern apply? What does the FDE see/hear that triggers recognition?>
30
+
31
+ ### The move
32
+ <What to do, specifically. Not advice - steps.>
33
+
34
+ ### Why it works
35
+ <The mechanism - why this approach succeeds where the obvious approach fails.>
36
+
37
+ ### Watch out for
38
+ <The failure mode or edge case that makes the pattern not apply.>
39
+
40
+ ### Evidence
41
+ <Permitted source, what happened, measured result, and limits. Keep identifying evidence in its original customer record.>
42
+ ```
43
+
44
+ **3. The pattern quality test.** Before encoding:
45
+
46
+ | Test | Pass | Fail |
47
+ |------|------|------|
48
+ | **Transferable?** | Another FDE could apply this without context from this engagement | Only makes sense if you know the specific client |
49
+ | **Specific enough?** | Contains concrete steps, not just principles | "Build trust" / "Communicate well" - too vague to act on |
50
+ | **Repeatable?** | Applies to a class of situations, not just this one | Only worked because of a unique circumstance |
51
+ | **Falsifiable?** | You can tell when the pattern is working or not | No way to measure whether applying it helped |
52
+ | **Monday bag?** | You would refuse the next similar embed without this in your bag - a named move plus an artifact you can drop on day one (pipe questions, CAB dance, eval golden shape, floor-drill script) | "We learned to communicate." Patterns that only live in this client's `.fde/` do not compound |
53
+
54
+ **4. Classify by stage.** Patterns sort into the same stages as the skills:
55
+
56
+ | Stage | Pattern type | Example |
57
+ |--------|-------------|---------|
58
+ | **Land** | Political / relational | "The passed-over team warm-up protocol" |
59
+ | **Discover** | Investigative / analytical | "The cron-job discovery checklist for legacy systems" |
60
+ | **Plan** | Structural / strategic | "The three-option presentation for nervous sponsors" |
61
+ | **Ship** | Technical / safety | "The Strangler Fig on financial transaction code" |
62
+ | **Outcome** | Operational / process | "The regulated-environment change-approval timeline buffer" |
63
+ | **Close** | Knowledge / handoff | "The 2am document format that actually gets used" |
64
+
65
+ **5. Version and evolve.** Patterns are living documents:
66
+
67
+ - First use: **v0.1** - hypothesis based on one engagement
68
+ - Later uses: record context, observed results, failures, and refinements. Repetition supplies evidence; it does not automatically validate the pattern. Promote a version when a substantive revision warrants it, not at a fixed use count.
69
+ - After modification: increment minor version with what changed and why
70
+ - After contradiction: note the counter-example, adjust the "watch out for" section
71
+
72
+ **6. Cross-engagement pattern mining.** Only when explicitly authorized for the named engagements and permitted by each customer's data policy. Keep records separate; use sanitized CLI packets or targeted recall in each authorized context, never raw file comparisons. Otherwise extract a candidate from the current permitted context only.
73
+
74
+ - Compare permitted problem summaries - do the same problems recur?
75
+ - Compare permitted decision summaries - are the same decisions being made?
76
+ - Compare permitted lessons - are the same lessons being learned twice?
77
+
78
+ A pattern learned twice is a process failure. Encoding it prevents the third time.
79
+
80
+ ## Artifact
81
+
82
+ **`patterns.md`** - candidates and evidence for this engagement, indexed by stage and situation trigger. A shared library is a separate, authorized export: remove names, identifiers, distinctive operational details, secrets, and confidential code or data. Keep source receipts in the original record and export only permitted generalizations.
83
+
84
+ **`retrospectives/YYYY-MM-DD-<engagement>.md`** - reference to which patterns were extracted from this engagement.
85
+
86
+ ## Checkpoint
87
+
88
+ Present the extracted patterns to the FDE: "From this engagement, I've identified N patterns worth encoding. The highest-value one is <name> because <it will apply to future engagements in these situations>." Confirm the pattern is accurate - the FDE's field judgment outranks the analysis.
89
+
90
+ ## Principles
91
+
92
+ - If you did it twice, encode it. The same lesson learned three times is a failure.
93
+ - Patterns are steps, not principles. "Build trust" isn't a pattern; "fix a small visible bug on day one" is.
94
+ - Every pattern needs a situation trigger - the FDE must recognise when it applies.
95
+ - Version substantive changes. State the evidence and limits; repeated use is not automatic confirmation.
96
+ - The pattern library is the FDE's compound interest. It's what separates 5 years of experience from 1 year repeated 5 times.
@@ -0,0 +1,12 @@
1
+ # Task context and evidence
2
+
3
+ Use this contract for standalone methods and methods routed through `@fde`.
4
+
5
+ - **Standalone work:** use the supplied, permitted facts, notes, code, and artifacts. A client name, `.fde/` directory, or initialized engagement is not a prerequisite. Do not bootstrap records merely to run a method. Ask only for missing information or authority that changes the next action; mark other gaps as unknown.
6
+ - **Artifact names are destinations:** names such as `success.md`, `decisions.md`, and `delivery.md` identify relevant evidence and, when bound, record destinations. If absent, use supplied facts and return the requested draft or result in the current workspace or conversation. Do not invent files or require initialization to complete useful work.
7
+ - **Bound engagement:** honor the current client binding and constraints. Before reading records, run `fde privacy` to verify masking support. Obtain a fresh, identity-matching sanitized `fde resume` packet for this task (or reuse a fresh session-hook packet); retrieve missing evidence with targeted `fde recall <topic>`. Use bounded `fde handoff` for transfer work. Refresh after binding, masking, or record changes. Never substitute raw `.fde/` reads, private blocks, masking dictionaries, or full transcripts. If the CLI is unavailable, use only permitted supplied excerpts and report the context limitation.
8
+ - **Authority:** continue reversible work within authorized scope. Reuse prior authorization when it covers the specific action. Show consequential engagement-record judgments and uncertainties for confirmation before saving unless already explicitly confirmed. New scope, acceptance changes, production actions, exports, and external messages need the applicable authority; a method invocation alone does not supply it. Keep one customer's writes in that customer's record.
9
+ - **Evidence:** distinguish supplied facts, estimates, hypotheses, and unknowns. Cite actual sources; a log date is not attribution. Never invent a source, signer, signature, customer reaction, or acceptance. Keep outcomes **promised → measured → accepted** distinct, and implementation, verification, deployment, and customer acceptance separate. Missing evidence means unproven, not an observed failure.
10
+ - **Data boundary:** use only data permitted by the customer's AI policy; clarify unknown policy before loading their code or data. Never load `<private>` content into a model. Cross-client comparison and exporting reusable material require permission and removal of customer-identifying or confidential content; anonymization alone does not grant permission.
11
+
12
+ Apply the selected method to this context. Follow its linked supporting methods only when needed; do not restart discovery or repeat already answered questions.
@@ -0,0 +1,21 @@
1
+ ---
2
+ name: fde-integrate
3
+ description: Build or change a customer-system integration with explicit data mapping, permissions, retries and reconciliation. Use for connectors, imports, write-back and upstream APIs.
4
+ ---
5
+
6
+ # fde-integrate
7
+
8
+ <!-- Generated by bin/generate-skills.js; edit the canonical references and catalog. -->
9
+
10
+ ## Purpose
11
+
12
+ Build or change a customer-system integration with explicit data mapping, permissions, retries and reconciliation. Use for connectors, imports, write-back and upstream APIs.
13
+
14
+ Read [the task context contract](references/task-context.md), then [the method](references/integrate.md). Load further references only when the task needs them. Everything linked is included in this skill; no other skill pack is required.
15
+
16
+ ## Principles
17
+
18
+ - Work directly from the supplied permitted context. Standalone work does not require an engagement folder or initialization. Record filenames in the method are optional persistence destinations when no engagement is bound.
19
+ - If called by @fde, reuse its current sanitized packet and scope. Do not restart setup, discovery or questions already answered.
20
+ - The task context contract controls persistence and authority in both modes. Preserve unknowns and distinguish implementation, verification, deployment and acceptance.
21
+ - Use the customer's repository instructions and available tools. Report a missing capability or unrun check honestly; do not claim that installing a skill provisions infrastructure.
@@ -0,0 +1,20 @@
1
+ # build - Implement a verifiable increment
2
+
3
+ **Enter when:** an agreed behavior needs implementation in an existing or new repository. For a broken behavior, start with [debug](debug.md); for a system boundary, use [integrate](integrate.md).
4
+
5
+ Use the permitted context and authority in [task context](task-context.md). This method works without `.fde/`; an existing engagement record can supply the same contract. Do not initialize memory just to write code.
6
+
7
+ ## Method
8
+
9
+ 1. Identify the repository, its instructions, working tree, relevant callers, and test commands. Inspect examples before creating abstractions. Preserve unrelated edits and state which dependencies or interfaces the change touches.
10
+ 2. State the observable outcome, constraints, and acceptance checks. Reuse agreed criteria for routine fixes. If a consequential product choice is unresolved, surface that choice while continuing independent investigation; do not invent acceptance.
11
+ 3. Choose the smallest coherent path that demonstrates the outcome through the real entry point. Include the necessary storage, error handling, and interface behavior in that slice. Name the failure that stops expansion and the recovery path for stateful changes.
12
+ 4. Implement using the repository's tools and conventions. Search for existing services, fixtures, and validation before adding alternatives. Keep cleanup limited to what makes the changed path understandable; do not expand scope to repair unrelated code.
13
+ 5. Run focused checks, then required repository checks. Exercise the actual affected journey with [QA](qa.md) when appropriate. For uncertain model behavior, use [eval-pack](eval-pack.md). Record results with [verification](verification.md), including checks that could not run.
14
+ 6. Inspect the final diff against the agreed outcome. For substantial or risky work, seek [review](review.md) using an actual separate reviewer when available; identify a self-check honestly. Reverify affected behavior after fixes.
15
+
16
+ ## Deliverable and acceptance
17
+
18
+ Return the implemented behavior, relevant paths, evidence, remaining limitations, and any decision needed. Done means the agreed checks have applicable evidence and the change is reviewable; passing tests does not imply deployment or customer acceptance. Committing, opening a PR, merging, and publishing happen only when the requested workflow authorizes those actions.
19
+
20
+ When coordinated through `@fde`, record implementation and verification in the existing decisions/delivery records under their write rules. Standalone work can return the same receipt directly or use the repository's task record.
@@ -0,0 +1,18 @@
1
+ # debug - Find and repair the cause
2
+
3
+ **Enter when:** a reproducible failure, regression, incident symptom, or misleading output needs investigation.
4
+
5
+ Use [task context](task-context.md). Work from supplied permitted evidence without requiring `.fde/`. During an active incident, follow the authorized containment procedure before diagnosis; investigation authority alone does not authorize production writes.
6
+
7
+ ## Method
8
+
9
+ 1. Capture expected and observed behavior, exact input or trigger, affected revision/environment, and the last known working state. Preserve useful errors and timestamps without copying secrets or raw private data. Mark reports you have not reproduced as reports.
10
+ 2. Inspect the failing path, callers, recent relevant changes, and existing tests. Reproduce in a permitted environment with the smallest representative case. If reproduction is unavailable, identify what observation would distinguish causes and gather safe evidence; do not claim a hypothesis is proven.
11
+ 3. Keep a short hypothesis list. For each, name the predicted observation and a discriminating check. Change one relevant variable at a time. Trace values and control flow across the actual boundary instead of repeatedly changing code until the symptom disappears.
12
+ 4. Fix the cause at the appropriate layer. Check whether the proposed fix changes behavior for other callers, stale data, retries, concurrency, or permissions. Preserve evidence of the original failure and avoid unrelated cleanup.
13
+ 5. Add a regression check when it can meaningfully reproduce the bug; show that it fails before the fix and passes after when practical. If the check cannot run against the before-state, say so. Run affected adjacent and required checks using [verification](verification.md).
14
+ 6. Review the final diff and exercise the original journey. For substantial or risky fixes use [review](review.md). After two unsuccessful repair cycles, reassess the hypothesis and evidence instead of repeating the same attempt; continue useful investigation and isolate the missing decision or access.
15
+
16
+ ## Deliverable and acceptance
17
+
18
+ Report the cause with its evidence, the fix, the original reproducer's result, adjacent checks, and unresolved uncertainty. A disappearing symptom with no discriminating evidence is a mitigation, not a demonstrated root cause. In engagement mode record the incident/fix receipt in the appropriate existing record; standalone work may return it directly. Release or rollback requires the existing operational authority and [ship](ship.md) or recovery procedure.
@@ -0,0 +1,26 @@
1
+ # eval-pack - Evaluate the model's allowed behavior
2
+
3
+ **Enter when:** AI, LLM, RAG, or agent behavior needs evidence before an experiment, release, or material expansion. Non-AI work skips this method.
4
+
5
+ Use [task context](task-context.md). Supplied permitted context and an evaluation report are sufficient without `.fde/`. In a coordinated engagement, use the existing trust/terrain context and keep the report in `evals.md`; read only privacy-safe views.
6
+
7
+ ## Method
8
+
9
+ 1. **Define the evaluated surface.** Name the model judgment, inputs, outputs, downstream actions, environment, and relevant failure impact. Separate model quality from deterministic tool authorization and application checks. Document the actual allowed action boundary and its source; missing authority remains unknown.
10
+ 2. **Choose cases by risk and coverage.** Use permitted historical examples, expert-labeled cases, or clearly marked synthetic fixtures. Cover relevant segments, boundary conditions, known failure modes, and critical harms. Record input, expected outcome/rubric, provenance, and critical-failure rule per case. Keep evaluation cases separate from tuning where possible; no fixed case count proves safety.
11
+ 3. **Agree the pass rule before the run.** Define quality thresholds, critical failures, coverage expectations, and acceptable uncertainty for this use. Use deterministic checks where possible and inspect subjective labels or judge reliability. Propose missing criteria for agreement; do not manufacture acceptance from the observed scores.
12
+ 4. **Run the actual evaluated path.** Record model/provider version, prompts/configuration, retrieval corpus or tools, application revision, environment, fixtures, and run date. Repeat where variability matters. Report totals, per-segment results, critical failures, and limitations using [verification](verification.md). A model-only run does not prove the agent's tool boundary works.
13
+ 5. **Verify action authority and controls.** Human approval is required where the user's policy or task requires it. Already agreed bounded automation may run within its documented actions, identities, environments, and limits; do not require fresh approval for every authorized action. Check enforcement outside the model, least privilege, input/output validation, cost/rate limits, stop conditions, observability, and recovery as applicable. Unknown or exceeded authority blocks those actions. Evaluation success never grants new authority.
14
+ 6. **Make a scoped verdict.** Report **SHIP** only when agreed criteria pass, critical failures are zero, applicable authority/control checks pass, and material coverage gaps are resolved or the release is explicitly narrowed by the responsible decision-maker. Otherwise report **NO-SHIP** with the smallest corrective step: fix, gather evidence, descope, or reconsider the judgment surface. A SHIP verdict is technical evidence for the stated scope, not permission to deploy.
15
+
16
+ ## Deliverable and acceptance
17
+
18
+ Return the suite/source, thresholds, counts and segments, top failure modes, control evidence, human-review gate or bounded automation authority, limitations, and dated verdict. Record unknown values honestly. Reevaluate after changes that affect model behavior, retrieval, tool permissions, or data conditions; cite why unchanged evidence remains applicable rather than implying a rerun.
19
+
20
+ When coordinated, append a concise eval receipt to delivery records. For release use [ship](ship.md). For ongoing use define the drift signals, sample policy permitted by data handling rules, owner, and conditions that suspend or narrow automation. Do not store secrets, raw `<private>` data, or hidden chain-of-thought in reports.
21
+
22
+ ## Principles
23
+
24
+ - Thresholds and authority come from the agreed contract, never from a convenient observed result.
25
+ - Critical failures block the evaluated release scope; disclose coverage and uncertainty.
26
+ - Bound automation with enforceable controls, and require human review where the policy requires it.
@@ -0,0 +1,18 @@
1
+ # integrate - Prove the system boundary
2
+
3
+ **Enter when:** connecting an API, data source, SDK, event stream, tool, or service, or changing its contract.
4
+
5
+ Start from [task context](task-context.md). Permitted supplied context is enough; `.fde/` is optional. Use the customer's existing clients, authentication, fixtures, and diagnostic tools. Do not create another integration platform to make one connection.
6
+
7
+ ## Method
8
+
9
+ 1. Map producer, consumer, owner, direction, and side effects. Inspect the actual installed version and local implementation; verify uncertain behavior against current official documentation. Identify the relevant schema, authentication scopes, network boundary, and permitted test environment.
10
+ 2. Write the acceptance example: an input at the real boundary and the observable downstream result. Include a rejection or failure example. Separate configuration validity, successful authentication, transport connectivity, contract compatibility, and end-to-end behavior; none proves the next.
11
+ 3. Inspect credentials by presence and required scope without printing values. Use existing secret storage. Check data classification and retention before moving data; never pass raw `<private>` blocks into a model. Prefer sanitized or synthetic cases approved for the target environment.
12
+ 4. Implement the narrow adapter using native repository patterns. Validate external inputs and model outputs, bound timeouts and retries, preserve error context without leaking payloads, and handle cancellation. For writes, establish idempotency or duplicate detection before retries; for events, check ordering, replay, and poison messages as applicable.
13
+ 5. Exercise a permitted success case and relevant failures: denied access, malformed data, rate limit, timeout, duplicate delivery, or partial completion. Trace correlation IDs or safe evidence across both sides. A mock proves client behavior only; if live access is unavailable, report that gap instead of claiming an integration works.
14
+ 6. Check cleanup and recovery for test side effects. Use [verification](verification.md) for receipts and [review](review.md) for security or data-contract changes. Route deployment through [ship](ship.md) only when authorized.
15
+
16
+ ## Deliverable and acceptance
17
+
18
+ Return the boundary contract, changed paths, environment, evidence at each tested layer, and remaining dependencies with owners when known. Done requires the agreed end-to-end result or an explicit narrower agreed scope. Do not silently replace live acceptance with a stub. In engagement mode, update the terrain/delivery record with confirmed facts; otherwise return the receipt directly.
@@ -0,0 +1,18 @@
1
+ # qa - Exercise the changed journey
2
+
3
+ **Enter when:** a feature or fix needs behavioral verification through its real interface, especially UI, API, and multi-step workflows.
4
+
5
+ Use [task context](task-context.md) and the customer's existing browser, API, fixtures, and test tooling. `.fde/` is not a prerequisite. Respect the permitted environment and authority for every side effect.
6
+
7
+ ## Method
8
+
9
+ 1. Identify the changed journey, user roles, acceptance checks, and risk-bearing neighboring paths. Record the revision and environment. Use synthetic or sanitized fixtures with understood cleanup; do not borrow production data without permission.
10
+ 2. Run the normal journey from its real entry point through the expected result. Verify persisted or downstream state when the requirement includes it; a success toast alone does not prove a write succeeded.
11
+ 3. Select negative and boundary cases from the change: invalid input, empty/loading/error states, refresh/back navigation, retries, duplicates, permissions, or interrupted work. For UI changes, inspect relevant viewport sizes, keyboard access, focus, labels, and errors. Use a real browser for the affected journey.
12
+ 4. Inspect relevant console and network evidence. Distinguish a UI defect from a failed API or unavailable environment. Retain only privacy-safe screenshots and logs. Do not claim visual verification from code inspection or a generated screenshot that was not viewed.
13
+ 5. Report failures with steps, expected/actual result, revision/environment, evidence, and impact. If repair is authorized, use [debug](debug.md), then rerun the failed journey and affected neighbors. Keep unrelated findings separate from the change.
14
+ 6. Produce a [verification receipt](verification.md). State which roles, devices, environments, or data conditions remain untested. Do not weaken acceptance checks to make the run pass.
15
+
16
+ ## Deliverable and acceptance
17
+
18
+ Return checked journeys and observed results, reproducible defects, limitations, and remaining blockers. Done means the agreed behavioral checks passed under the stated conditions. A browser smoke test does not establish load capacity, security assurance, accessibility conformance, deployment, or customer acceptance by itself. When coordinated, append the evidence to the existing delivery record; standalone QA can return it directly.