fdeops 4.0.3 → 4.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +1 -1
- package/README.md +33 -6
- package/bin/check.js +14 -23
- package/bin/fde.js +24 -11
- package/bin/generate-skills.js +71 -0
- package/bin/install.js +2 -1
- package/bin/lib/context.js +8 -1
- package/bin/lib/provenance.js +10 -0
- package/bin/skill-catalog.js +17 -0
- package/mcp/fdeops-ingest/package.json +2 -2
- package/package.json +4 -3
- package/plugin.json +2 -2
- package/skills/fde/SKILL.md +18 -10
- package/skills/fde/references/board-memo.md +1 -1
- package/skills/fde/references/build.md +20 -0
- package/skills/fde/references/business-case.md +9 -7
- package/skills/fde/references/close.md +5 -3
- package/skills/fde/references/debug.md +18 -0
- package/skills/fde/references/encode-pattern.md +10 -8
- package/skills/fde/references/eval-pack.md +16 -33
- package/skills/fde/references/hold-scope.md +9 -7
- package/skills/fde/references/integrate.md +18 -0
- package/skills/fde/references/plan.md +4 -2
- package/skills/fde/references/poc.md +5 -3
- package/skills/fde/references/qa.md +18 -0
- package/skills/fde/references/review.md +22 -60
- package/skills/fde/references/ship.md +42 -287
- package/skills/fde/references/task-context.md +12 -0
- package/skills/fde/references/test-assumptions.md +2 -2
- package/skills/fde/references/three-options.md +19 -27
- package/skills/fde/references/verification.md +31 -0
- package/skills/fde-build/SKILL.md +21 -0
- package/skills/fde-build/references/build.md +20 -0
- package/skills/fde-build/references/debug.md +18 -0
- package/skills/fde-build/references/eval-pack.md +26 -0
- package/skills/fde-build/references/integrate.md +18 -0
- package/skills/fde-build/references/qa.md +18 -0
- package/skills/fde-build/references/review.md +39 -0
- package/skills/fde-build/references/ship.md +73 -0
- package/skills/fde-build/references/task-context.md +12 -0
- package/skills/fde-build/references/verification.md +31 -0
- package/skills/fde-debug/SKILL.md +21 -0
- package/skills/fde-debug/references/build.md +20 -0
- package/skills/fde-debug/references/debug.md +18 -0
- package/skills/fde-debug/references/eval-pack.md +26 -0
- package/skills/fde-debug/references/integrate.md +18 -0
- package/skills/fde-debug/references/qa.md +18 -0
- package/skills/fde-debug/references/review.md +39 -0
- package/skills/fde-debug/references/ship.md +73 -0
- package/skills/fde-debug/references/task-context.md +12 -0
- package/skills/fde-debug/references/verification.md +31 -0
- package/skills/fde-discover/SKILL.md +21 -0
- package/skills/fde-discover/references/audit.md +71 -0
- package/skills/fde-discover/references/discover.md +254 -0
- package/skills/fde-discover/references/task-context.md +12 -0
- package/skills/fde-evaluate/SKILL.md +21 -0
- package/skills/fde-evaluate/references/build.md +20 -0
- package/skills/fde-evaluate/references/debug.md +18 -0
- package/skills/fde-evaluate/references/eval-pack.md +26 -0
- package/skills/fde-evaluate/references/integrate.md +18 -0
- package/skills/fde-evaluate/references/qa.md +18 -0
- package/skills/fde-evaluate/references/review.md +39 -0
- package/skills/fde-evaluate/references/ship.md +73 -0
- package/skills/fde-evaluate/references/task-context.md +12 -0
- package/skills/fde-evaluate/references/verification.md +31 -0
- package/skills/fde-feedback/SKILL.md +21 -0
- package/skills/fde-feedback/references/encode-pattern.md +96 -0
- package/skills/fde-feedback/references/task-context.md +12 -0
- package/skills/fde-handoff/SKILL.md +21 -0
- package/skills/fde-handoff/references/close.md +66 -0
- package/skills/fde-handoff/references/encode-pattern.md +96 -0
- package/skills/fde-handoff/references/task-context.md +12 -0
- package/skills/fde-integrate/SKILL.md +21 -0
- package/skills/fde-integrate/references/build.md +20 -0
- package/skills/fde-integrate/references/debug.md +18 -0
- package/skills/fde-integrate/references/eval-pack.md +26 -0
- package/skills/fde-integrate/references/integrate.md +18 -0
- package/skills/fde-integrate/references/qa.md +18 -0
- package/skills/fde-integrate/references/review.md +39 -0
- package/skills/fde-integrate/references/ship.md +73 -0
- package/skills/fde-integrate/references/task-context.md +12 -0
- package/skills/fde-integrate/references/verification.md +31 -0
- package/skills/fde-options/SKILL.md +21 -0
- package/skills/fde-options/references/business-case.md +90 -0
- package/skills/fde-options/references/task-context.md +12 -0
- package/skills/fde-options/references/test-assumptions.md +102 -0
- package/skills/fde-options/references/three-options.md +90 -0
- package/skills/fde-poc/SKILL.md +21 -0
- package/skills/fde-poc/references/audit.md +71 -0
- package/skills/fde-poc/references/build.md +20 -0
- package/skills/fde-poc/references/business-case.md +90 -0
- package/skills/fde-poc/references/debug.md +18 -0
- package/skills/fde-poc/references/discover.md +254 -0
- package/skills/fde-poc/references/eval-pack.md +26 -0
- package/skills/fde-poc/references/integrate.md +18 -0
- package/skills/fde-poc/references/plan.md +167 -0
- package/skills/fde-poc/references/poc.md +55 -0
- package/skills/fde-poc/references/qa.md +18 -0
- package/skills/fde-poc/references/review.md +39 -0
- package/skills/fde-poc/references/ship.md +73 -0
- package/skills/fde-poc/references/task-context.md +12 -0
- package/skills/fde-poc/references/test-assumptions.md +102 -0
- package/skills/fde-poc/references/three-options.md +90 -0
- package/skills/fde-poc/references/verification.md +31 -0
- package/skills/fde-qa/SKILL.md +21 -0
- package/skills/fde-qa/references/build.md +20 -0
- package/skills/fde-qa/references/debug.md +18 -0
- package/skills/fde-qa/references/eval-pack.md +26 -0
- package/skills/fde-qa/references/integrate.md +18 -0
- package/skills/fde-qa/references/qa.md +18 -0
- package/skills/fde-qa/references/review.md +39 -0
- package/skills/fde-qa/references/ship.md +73 -0
- package/skills/fde-qa/references/task-context.md +12 -0
- package/skills/fde-qa/references/verification.md +31 -0
- package/skills/fde-readout/SKILL.md +21 -0
- package/skills/fde-readout/references/board-memo.md +108 -0
- package/skills/fde-readout/references/business-case.md +90 -0
- package/skills/fde-readout/references/readout.md +69 -0
- package/skills/fde-readout/references/task-context.md +12 -0
- package/skills/fde-review/SKILL.md +21 -0
- package/skills/fde-review/references/build.md +20 -0
- package/skills/fde-review/references/debug.md +18 -0
- package/skills/fde-review/references/eval-pack.md +26 -0
- package/skills/fde-review/references/integrate.md +18 -0
- package/skills/fde-review/references/qa.md +18 -0
- package/skills/fde-review/references/review.md +39 -0
- package/skills/fde-review/references/ship.md +73 -0
- package/skills/fde-review/references/task-context.md +12 -0
- package/skills/fde-review/references/verification.md +31 -0
- package/skills/fde-scope/SKILL.md +21 -0
- package/skills/fde-scope/references/hold-scope.md +83 -0
- package/skills/fde-scope/references/task-context.md +12 -0
- package/skills/fde-ship/SKILL.md +21 -0
- package/skills/fde-ship/references/build.md +20 -0
- package/skills/fde-ship/references/debug.md +18 -0
- package/skills/fde-ship/references/eval-pack.md +26 -0
- package/skills/fde-ship/references/integrate.md +18 -0
- package/skills/fde-ship/references/qa.md +18 -0
- package/skills/fde-ship/references/review.md +39 -0
- package/skills/fde-ship/references/ship.md +73 -0
- package/skills/fde-ship/references/task-context.md +12 -0
- package/skills/fde-ship/references/verification.md +31 -0
|
@@ -0,0 +1,90 @@
|
|
|
1
|
+
# business-case - Build the business case
|
|
2
|
+
|
|
3
|
+
**Context:** apply [task context and evidence](task-context.md) before using the named records below.
|
|
4
|
+
|
|
5
|
+
**Enter when:** the sponsor needs justification for the next phase, the FDE needs to defend budget or timeline, a feature decision needs cost/benefit evidence, or poc produced a direction that needs funding.
|
|
6
|
+
|
|
7
|
+
**Read first:** `reality.md`, `success.md`, `delivery.md`, `context.md`. Load `business-case.md` from poc if it exists - extend it, don't restart.
|
|
8
|
+
|
|
9
|
+
Technical FDEs lose engagements by shipping good code without business justification. The sponsor's boss doesn't ask "is the code clean?" - they ask "what did we get for the money?" A business case translates technical work into the language that keeps the engagement alive.
|
|
10
|
+
|
|
11
|
+
## Method (you do this work)
|
|
12
|
+
|
|
13
|
+
**1. Name the cost of doing nothing.** This is the anchor. Every business case starts not with what you'll build, but with what it costs them to leave the problem unsolved:
|
|
14
|
+
|
|
15
|
+
| Cost type | How to find it | Example |
|
|
16
|
+
|-----------|---------------|---------|
|
|
17
|
+
| **Labor capacity / direct spend** | Ask: "What does this problem cost per month in money?" | Manual reconciliation hours × loaded rate = capacity value; separately identify reducible spend |
|
|
18
|
+
| **Opportunity cost** | Ask: "What can't you do because of this problem?" | Can't onboard enterprise clients because the API can't handle their volume |
|
|
19
|
+
| **Risk cost** | Ask: "What happens if this breaks at the worst time?" | A payment processing outage during Black Friday = $X/hour in lost sales |
|
|
20
|
+
| **Velocity cost** | Measure: deployment frequency, lead time, change failure rate | Team ships once/month instead of once/week; each delay = N features not reaching customers |
|
|
21
|
+
|
|
22
|
+
**2. Build the driver model.** Not a spreadsheet - a logic chain the sponsor can trace:
|
|
23
|
+
|
|
24
|
+
```
|
|
25
|
+
Investment: <hours × rate, or fixed cost>
|
|
26
|
+
→ Delivers: <specific outcome from success.md>
|
|
27
|
+
→ Benefit: <capacity released, avoidable cash spend, revenue, or risk reduction>
|
|
28
|
+
→ Net cash: realizable incremental cash benefit - full costs over <time horizon>
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
Keep drivers, units, sources, and ranges explicit. For example, 3 people × 8h/week × $75/h × 52 weeks = $93.6K/year of labor capacity value. It is cash savings only if spend actually falls (for example, paid overtime or a contractor cost ends). Name who can realize the benefit and how. Include build, ongoing operation, adoption, and transition costs; avoid double-counting capacity and revenue enabled by the same hours. Do not calculate cash payback from capacity value alone.
|
|
32
|
+
|
|
33
|
+
**3. Sensitivity check - name the two drivers that swing the result:**
|
|
34
|
+
|
|
35
|
+
Every business case has 1-2 variables where a small change flips the outcome. Name them explicitly:
|
|
36
|
+
|
|
37
|
+
> "The capacity case assumes the team reclaims 6 hours/week per person. At 3 hours, that benefit halves. Cash payback remains unproven until finance identifies avoidable spend. Validate time-spent before and after the pilot with representative team members."
|
|
38
|
+
|
|
39
|
+
The sponsor who sees you've identified where the case could break trusts the case more, not less.
|
|
40
|
+
|
|
41
|
+
**4. Frame for the audience.** Different stakeholders need different lenses on the same case:
|
|
42
|
+
|
|
43
|
+
| Audience | Lead with | Avoid |
|
|
44
|
+
|----------|----------|-------|
|
|
45
|
+
| **CFO / finance** | ROI, payback period, cash flow impact | Technical architecture, feature lists |
|
|
46
|
+
| **CTO / engineering** | Technical debt retired, velocity improved, risk reduced | Revenue projections they can't verify |
|
|
47
|
+
| **CEO / founder** | Strategic enablement, competitive edge, customer impact | Detailed calculations (give the summary, offer the detail) |
|
|
48
|
+
| **Product** | User impact, adoption metrics, feature velocity | Cost structures that aren't their domain |
|
|
49
|
+
|
|
50
|
+
**5. The one-page format.** The business case fits one page or it isn't understood:
|
|
51
|
+
|
|
52
|
+
```markdown
|
|
53
|
+
## Business case: <initiative name>
|
|
54
|
+
|
|
55
|
+
**The problem costs:** <one line, quantified>
|
|
56
|
+
**The investment:** <hours and cost>
|
|
57
|
+
**The return:** <quantified, with time horizon>
|
|
58
|
+
**Payback:** <months from realizable cash benefits, or not established>
|
|
59
|
+
**Sensitivity:** <the 1-2 drivers that swing it, with thresholds>
|
|
60
|
+
**Risks:** <what must be true for this to hold>
|
|
61
|
+
**Recommendation:** <proceed / proceed-with-conditions / defer>
|
|
62
|
+
```
|
|
63
|
+
|
|
64
|
+
## Artifact
|
|
65
|
+
|
|
66
|
+
**`business-case.md`** - the one-page case. Lives alongside `success.md` and `reality.md` as a first-class engagement artifact. Referenced by plan, status, and close.
|
|
67
|
+
|
|
68
|
+
**`decisions.md`** - log the sponsor's response: approved, modified, deferred. With the date.
|
|
69
|
+
|
|
70
|
+
## Checkpoint
|
|
71
|
+
|
|
72
|
+
Walk the FDE through: the cost of doing nothing (anchor), the investment, the return, and the one sensitivity that matters most. If the FDE says "the sponsor won't buy the ROI number," inspect the disputed inputs and sources, test plausible ranges, and identify what measurement would resolve the disagreement. Never reverse-engineer assumptions to hit a desired number.
|
|
73
|
+
|
|
74
|
+
## Worked example
|
|
75
|
+
|
|
76
|
+
Acme phase 2 needs funding. The case starts with the cost of doing nothing, not the cost of building.
|
|
77
|
+
|
|
78
|
+
Anchor: two silent failures since March, each one day of finance reconciliation by hand plus a late close (`reality.md`, Marco's sheet). That is the number the sponsor already believes because her own team reported it.
|
|
79
|
+
|
|
80
|
+
Driver model the sponsor can trace: incidents/quarter × hours of manual reconciliation × loaded cost, plus the tail risk of a late regulatory close - stated separately, because mixing a certain small number with an uncertain large one is how a case loses credibility.
|
|
81
|
+
|
|
82
|
+
Sensitivity names the two drivers that swing it: incident frequency (2/quarter → 1/quarter and the case halves) and whether the manual re-run continues in parallel (if Marco keeps re-running every morning, the saving is theoretical). The second one is the honest weakness, so it is in the case rather than waiting to be found in the room - with the condition that makes it hold: the morning re-run stops after two clean cycles, agreed with Marco.
|
|
83
|
+
|
|
84
|
+
## Principles
|
|
85
|
+
|
|
86
|
+
- The cost of doing nothing is always the opening move. Anchor before proposing.
|
|
87
|
+
- Driver models with visible arithmetic beat magic spreadsheets.
|
|
88
|
+
- Name the sensitivity. The case that admits its weakness earns more trust.
|
|
89
|
+
- One page. If it doesn't fit, you don't understand it yet.
|
|
90
|
+
- A business case the FDE can't explain in 60 seconds won't survive the sponsor's boss.
|
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
# debug - Find and repair the cause
|
|
2
|
+
|
|
3
|
+
**Enter when:** a reproducible failure, regression, incident symptom, or misleading output needs investigation.
|
|
4
|
+
|
|
5
|
+
Use [task context](task-context.md). Work from supplied permitted evidence without requiring `.fde/`. During an active incident, follow the authorized containment procedure before diagnosis; investigation authority alone does not authorize production writes.
|
|
6
|
+
|
|
7
|
+
## Method
|
|
8
|
+
|
|
9
|
+
1. Capture expected and observed behavior, exact input or trigger, affected revision/environment, and the last known working state. Preserve useful errors and timestamps without copying secrets or raw private data. Mark reports you have not reproduced as reports.
|
|
10
|
+
2. Inspect the failing path, callers, recent relevant changes, and existing tests. Reproduce in a permitted environment with the smallest representative case. If reproduction is unavailable, identify what observation would distinguish causes and gather safe evidence; do not claim a hypothesis is proven.
|
|
11
|
+
3. Keep a short hypothesis list. For each, name the predicted observation and a discriminating check. Change one relevant variable at a time. Trace values and control flow across the actual boundary instead of repeatedly changing code until the symptom disappears.
|
|
12
|
+
4. Fix the cause at the appropriate layer. Check whether the proposed fix changes behavior for other callers, stale data, retries, concurrency, or permissions. Preserve evidence of the original failure and avoid unrelated cleanup.
|
|
13
|
+
5. Add a regression check when it can meaningfully reproduce the bug; show that it fails before the fix and passes after when practical. If the check cannot run against the before-state, say so. Run affected adjacent and required checks using [verification](verification.md).
|
|
14
|
+
6. Review the final diff and exercise the original journey. For substantial or risky fixes use [review](review.md). After two unsuccessful repair cycles, reassess the hypothesis and evidence instead of repeating the same attempt; continue useful investigation and isolate the missing decision or access.
|
|
15
|
+
|
|
16
|
+
## Deliverable and acceptance
|
|
17
|
+
|
|
18
|
+
Report the cause with its evidence, the fix, the original reproducer's result, adjacent checks, and unresolved uncertainty. A disappearing symptom with no discriminating evidence is a mitigation, not a demonstrated root cause. In engagement mode record the incident/fix receipt in the appropriate existing record; standalone work may return it directly. Release or rollback requires the existing operational authority and [ship](ship.md) or recovery procedure.
|
|
@@ -0,0 +1,254 @@
|
|
|
1
|
+
# discover - Frame the problem
|
|
2
|
+
|
|
3
|
+
**Enter when:** the brief feels wrong, the real problem is unclear, shadow processes are suspected, or any phase found that the map is missing.
|
|
4
|
+
|
|
5
|
+
**Read first:** `context.md`, `brief.md`. Load `terrain.md` if it exists - extend it, never regenerate from scratch.
|
|
6
|
+
|
|
7
|
+
## Validation gate (confirm understanding, clarify where it elevates)
|
|
8
|
+
|
|
9
|
+
Before discovering, state what you're investigating and why in 2-3 lines:
|
|
10
|
+
|
|
11
|
+
> "Investigating: [the hypothesis or problem area]. This informs: [the decision it feeds - descope/rescope/pick A over B]. Existing terrain: [what's already mapped vs. what's unknown]."
|
|
12
|
+
|
|
13
|
+
Then check - probe ONLY if it prevents wasted discovery:
|
|
14
|
+
|
|
15
|
+
1. **Hypothesis is testable.** If the stated problem is unfalsifiable ("the architecture is wrong") → rephrase it: "I'd narrow this to: [specific testable claim]. That closer to what you're seeing?"
|
|
16
|
+
2. **Discovery feeds a decision.** If there's no named decision → one line: "What changes depending on what we find? That keeps the discovery focused."
|
|
17
|
+
3. **Not repeating previous work.** If terrain.md already covers this area → name it: "Terrain already maps this from Day [X]. Extending it or has something shifted?"
|
|
18
|
+
|
|
19
|
+
State your read, let the FDE correct, then discover.
|
|
20
|
+
|
|
21
|
+
Before asking for facts, inspect the supplied brief and existing redacted records for the answer. Once the decision frame is confirmed and code access is authorized, use the scan below and targeted file reads to resolve technical unknowns. Phrase remaining questions around the discrepancy found: “The queue already exists, but alerts are disabled; who currently checks it?”
|
|
22
|
+
|
|
23
|
+
## Brief interrogation (when the hypothesis is still mush)
|
|
24
|
+
|
|
25
|
+
Use when the "problem" is unfalsifiable, success is undefined, or you cannot name the decision discovery informs. Skip when `reality.md` / `terrain.md` already pin a testable claim and the FDE is ready to dig.
|
|
26
|
+
|
|
27
|
+
Same format as land - one Q + GUESS, no checklist:
|
|
28
|
+
|
|
29
|
+
```
|
|
30
|
+
READ: <the real problem you think exists, in one sentence>
|
|
31
|
+
CONFIDENCE: ~NN% - missing: <what would falsify or confirm it>
|
|
32
|
+
Q: <one question that changes where you dig>
|
|
33
|
+
GUESS: <your answer, so they can correct it>
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
Stop when you can write the four lines under **Frame the decision first**. If a name, quote, or metric is still missing, write `unknown - ask:` - never invent ops folklore to make the map look complete.
|
|
37
|
+
|
|
38
|
+
## Frame the decision first
|
|
39
|
+
|
|
40
|
+
Same SCQA spine as readout (`S → C → Q → A`), aimed at the floor, not a deck. Write it **before** any scan. Confirm with the FDE, then dig.
|
|
41
|
+
|
|
42
|
+
| Line | What it is | Fail if |
|
|
43
|
+
|------|------------|---------|
|
|
44
|
+
| **Situation** | What they already treat as true - the workaround, the sheet, the owner who left | It could be copied from the RFP |
|
|
45
|
+
| **Complication** | What broke, so they cannot stay here | No tension, or three problems joined by "and" |
|
|
46
|
+
| **Question** | One decision the named signer must make | It smuggles the solution ("how do we add alerting") |
|
|
47
|
+
| **Answer-space** | Shape of a satisfying answer: confirm brief / descope / rescope / pause | A novel, or "insights" |
|
|
48
|
+
|
|
49
|
+
Tests on **Question** - rewrite until all five hold:
|
|
50
|
+
|
|
51
|
+
1. **Decision-shaped** - answering it changes what someone does.
|
|
52
|
+
2. **Single** - one thing, not three.
|
|
53
|
+
3. **Scoped** - who, where, by when.
|
|
54
|
+
4. **Answerable** - evidence could settle it in this engagement.
|
|
55
|
+
5. **Neutral** - does not assume the fix.
|
|
56
|
+
|
|
57
|
+
Cannot write the Question → keep interrogating. Do not `fde scan`. Every later output of this phase aims at that Question. Sub-questions go to the operating map or `assumptions.md`, not into the Question.
|
|
58
|
+
|
|
59
|
+
## Parts of the problem (decompose only)
|
|
60
|
+
|
|
61
|
+
After the Question is locked, and **before** `fde scan` or any option: write what the problem is made of. No advice, no playbook, no solution.
|
|
62
|
+
|
|
63
|
+
If the stated brief hides a deeper job, name that deeper job in one sentence and **wait**. Do not silently replace their problem with yours.
|
|
64
|
+
|
|
65
|
+
In `terrain.md` under `## Parts`, list the smallest useful pieces that still change what you examine next. Typical cuts: people, process step, system, data, time, cost. For each piece: what it contains, and how it connects to the Question. Stop when a further split would not change where you dig.
|
|
66
|
+
|
|
67
|
+
Do not mark pieces as facts or assumptions here. That is `test-assumptions`. Do not assemble options here. That is `three-options`.
|
|
68
|
+
|
|
69
|
+
## Method - part 1: the codebase (you do this work)
|
|
70
|
+
|
|
71
|
+
**First code move: `fde scan`** - after the Question is locked. It runs everything below deterministically in seconds (churn×tests, "temporary" archaeology, AI components, secrets redacted, previous attempts). Your job is then **interpretation**: read its output against the brief, follow the hotspots into the code, and connect the technical findings to the human signals in part 2.
|
|
72
|
+
|
|
73
|
+
If the CLI is unavailable, run the manual commands below. Either way: do not load the full codebase into context - scan wide, read deep only on hotspots.
|
|
74
|
+
|
|
75
|
+
**1. Stack and age.** Language, framework, build system, date of last major upgrade:
|
|
76
|
+
```bash
|
|
77
|
+
git log --reverse --format="%ad" --date=short | head -1 # repo birth
|
|
78
|
+
git log -1 --format="%ad" --date=short # last commit
|
|
79
|
+
```
|
|
80
|
+
|
|
81
|
+
**2. Churn heat - the modules everyone touches but fears:**
|
|
82
|
+
```bash
|
|
83
|
+
git log --since="90 days ago" --name-only --pretty=format: | sort | uniq -c | sort -rn | head -20
|
|
84
|
+
```
|
|
85
|
+
The highest-churn file in a legacy codebase is the one everyone is afraid to refactor but cannot avoid touching. Cross-reference with complexity (file size, nesting) and mark "handle with care."
|
|
86
|
+
|
|
87
|
+
**3. Test gaps - what's covered, what's a lie:**
|
|
88
|
+
```bash
|
|
89
|
+
find . -path ./node_modules -prune -o -name "*test*" -print | head -30
|
|
90
|
+
```
|
|
91
|
+
Map test files against the churn list. A high-churn module with no test neighbors is a load-bearing wall with no insurance. Spot-read the tests that do exist: tests that pass but assert nothing are worse than no tests - note them.
|
|
92
|
+
|
|
93
|
+
**4. The "temporary" archaeology** (repeat `--include` per extension - brace globs silently match nothing):
|
|
94
|
+
```bash
|
|
95
|
+
grep -rnE "HACK|FIXME|XXX|temporary|for now|remove this|workaround" \
|
|
96
|
+
--include="*.js" --include="*.ts" --include="*.py" --include="*.java" \
|
|
97
|
+
--include="*.go" --include="*.rb" --include="*.cs" --include="*.php" . | head -30
|
|
98
|
+
```
|
|
99
|
+
Temporary code in production is permanent code with an excuse. Each hit is a candidate for "what was never built properly."
|
|
100
|
+
|
|
101
|
+
**5. AI components - they fail silently:**
|
|
102
|
+
```bash
|
|
103
|
+
grep -rlnE "openai|anthropic|llm|prompt|embedding|vector|inference" \
|
|
104
|
+
--include="*.js" --include="*.ts" --include="*.py" --include="*.java" \
|
|
105
|
+
--include="*.go" . | head -20
|
|
106
|
+
```
|
|
107
|
+
Flag every one. AI components don't fail like regular code - they degrade as the world changes. Each needs: model version, fallback path (or note its absence), observability (or note its absence).
|
|
108
|
+
|
|
109
|
+
**6. Data flow.** Where data enters, how it moves, where it stops. Entry points first: routes, queues, cron, file drops.
|
|
110
|
+
|
|
111
|
+
**7. Existing capability.** Trace the requested user action through existing code, configuration, tests, and operating workarounds. In `terrain.md`, record what can already be reused and the evidence that it works or fails. Check whether a configuration, ownership, or process change could resolve the observed break. A disabled feature is a lead, not a proven root cause. Keep observations and hypotheses distinct; option selection still belongs to plan / three-options. Summarize the remaining gap in `reality.md`: what works today → what the customer needs → what is still missing, with sources. If existing capability meets the need, say so; do not manufacture a build requirement.
|
|
112
|
+
|
|
113
|
+
## Method - part 2: the humans (you coach, the FDE asks)
|
|
114
|
+
|
|
115
|
+
The real spec is what people **do** when the system fails - not what the slide deck says. Arm the FDE with these, in their own words:
|
|
116
|
+
|
|
117
|
+
- **"How is the team coping today without the fix?"** - the workaround is the honest requirements doc.
|
|
118
|
+
- **Find the spreadsheet.** Almost always there. Whoever maintains it is the best interview in the building.
|
|
119
|
+
- **The hesitation.** When someone says "well, there's also this other thing we do…" - stop them, ask them to finish. The main story is what they're comfortable explaining; the hesitation is the real problem.
|
|
120
|
+
- **"Which part of the codebase do you least want to touch?"** The answer is unanimous and it's the load-bearing wall. Check it against your churn scan - when the human answer and the churn data agree, that's your first map landmark.
|
|
121
|
+
- **Shadow AI.** Someone pasting data into ChatGPT to cope = a real unmet need + an uncontrolled data risk. Note both.
|
|
122
|
+
- **Exception-led operating map.** For each real break (not the slide-deck process): what fails, who notices first, what they do today, and which artifact is trusted in that moment. Prefer exceptions over happy-path swimlanes - the workaround is the operating system. Write rows under `terrain.md` → `## Operating map (exception-led)`. If the section is missing on an older engagement, add it; never regenerate the rest of terrain. When AI is in play, also fill `## Intelligence placement` (deterministic vs LLM judgement vs human approve). **`fde doctor` requires at least one filled exception row before plan/ship/outcome/close** - empty map after discover is a hygiene fail, not optional polish.
|
|
123
|
+
|
|
124
|
+
## Method - part 3: workshop facilitation
|
|
125
|
+
|
|
126
|
+
When discovery requires a structured session with multiple stakeholders (alignment, prioritisation, design):
|
|
127
|
+
|
|
128
|
+
**Before the room:**
|
|
129
|
+
- Define the single decision the workshop must produce - not "discuss options" but "rank the three candidates and commit to one."
|
|
130
|
+
- Cap at 8 people. Every person above 8 halves the probability of a decision.
|
|
131
|
+
- Time-box: 90 minutes max. Anything longer splits into two sessions.
|
|
132
|
+
- Pre-read: one page, sent 48 hours ahead. Nobody will read more.
|
|
133
|
+
|
|
134
|
+
**In the room (the FDE facilitates, not presents):**
|
|
135
|
+
1. **5 min - frame.** One slide: the decision, the constraint, the deadline. No history lesson.
|
|
136
|
+
2. **15 min - diverge.** Silent post-its (or digital equivalent). Everyone writes before anyone talks - prevents the loudest voice dominating.
|
|
137
|
+
3. **20 min - cluster.** Group themes, name them. The FDE does NOT label - the room labels.
|
|
138
|
+
4. **30 min - converge.** Dot-vote or forced-rank. The FDE counts, the room decides.
|
|
139
|
+
5. **10 min - lock.** State the decision back. "We're saying X. Anyone who can't live with this, speak now." Silence = consent.
|
|
140
|
+
6. **10 min - next steps.** Who does what by when. Written before people stand up.
|
|
141
|
+
|
|
142
|
+
**After the room:** Summary in `decisions.md` within 2 hours. Decisions decay - what felt clear at 3pm is debatable by 5pm if unwritten.
|
|
143
|
+
|
|
144
|
+
## Method - part 4: data estate and the pipe
|
|
145
|
+
|
|
146
|
+
Always map the estate before you score a use case - not only when someone said "AI." A path they cannot feed is a discover miss, not a ship surprise.
|
|
147
|
+
|
|
148
|
+
**Their words first.** In `terrain.md`, write the names the floor uses for the workaround, the sheet, the exception path, and the person who left. Later plan/ship/review sentences use those names. Do not translate their floor into generic product language.
|
|
149
|
+
|
|
150
|
+
**The 5 questions (ask the data owner, not the sponsor):**
|
|
151
|
+
1. **Where does data live?** - List every source: databases, warehouses, SaaS exports, spreadsheets, S3 buckets, vendor APIs. Map it.
|
|
152
|
+
2. **How fresh is it?** - Real-time, daily batch, "someone uploads a CSV on Mondays"? Freshness determines what's buildable.
|
|
153
|
+
3. **Who owns it?** - Not "IT" - the named person who can grant access and explain the schema. No named owner: access responsibility remains unverified.
|
|
154
|
+
4. **What's the quality?** - Sample 100 rows from each critical source. Check: nulls, duplicates, format consistency, semantic correctness. A 60% null rate in a key field = that source is fiction.
|
|
155
|
+
5. **What are the governance constraints?** - PII classification, retention policies, cross-border rules, consent basis. One missed constraint = a compliance stop later.
|
|
156
|
+
|
|
157
|
+
**The pipe (what talks to what).** For each source that a use case depends on, write: the system it flows from and to, the contract (object, table, file, API), whose credentials, what happens when the vendor 500s or the Monday file does not land, and whether the join the sponsor described actually exists. Their IdP, CRM, and warehouse are delivery work when the path needs them - policy questions in `trust-profile.md` are not a substitute.
|
|
158
|
+
|
|
159
|
+
**The data readiness matrix:**
|
|
160
|
+
|
|
161
|
+
| Source | Location | Freshness | Owner | Quality (sample) | Governance | Pipe | Verdict |
|
|
162
|
+
|--------|----------|-----------|-------|-----------------|------------|------|---------|
|
|
163
|
+
| _fill per source_ | | | | | | | Ready / Needs work / Blocker |
|
|
164
|
+
|
|
165
|
+
A use case that depends on a "Blocker" source **or a Blocker pipe** doesn't get scored - it gets a remediation conversation first. `what-breaks` finding an invisible integration at ship is already too late. Write this to `terrain.md` under a `## Data estate` section.
|
|
166
|
+
|
|
167
|
+
**Promised dependencies are not ready dependencies.** For consequential promises such as "data in two weeks," record or update one dependency entry in `assumptions.md` with the responsible owner, dated verification checkpoint, and evidence needed. Unknown owners or dates stay unknown; propose a checkpoint for confirmation. Link the affected work; if the checkpoint slips, identify what can proceed and what needs replanning. Missing ownership is an unresolved dependency, not proof that the project will fail. On-prem or restricted access is a constraint to investigate, not a red flag by itself.
|
|
168
|
+
|
|
169
|
+
**Verify the future operator now.** Check the proposed owner in `success.md` against who will actually monitor, recover, and support the result. Record whether they have agreed, access or training gaps, and a practical handoff check there; carry these into `handoff.md` at close. Keep unconfirmed ownership explicit. Reuse supplied evidence and ask only what changes the plan.
|
|
170
|
+
|
|
171
|
+
## When scope is a transformation, not a single problem
|
|
172
|
+
|
|
173
|
+
Score every candidate use case before anything gets prototyped:
|
|
174
|
+
|
|
175
|
+
| Dimension | Question | 1-5 |
|
|
176
|
+
|---|---|---|
|
|
177
|
+
| Business value | What does it cost them unsolved? | |
|
|
178
|
+
| Complexity | How hard to build safely? (5 = hardest) | |
|
|
179
|
+
| Data readiness | Available, clean, sufficient volume today? | |
|
|
180
|
+
|
|
181
|
+
**Score = (Value × Data readiness) / Complexity.** Highest score gets prototyped first (hand to `poc`). A 5-value/1-complexity/5-readiness case scores 25; a 5-value/5-complexity/2-readiness case scores 2 - they look identical on a whiteboard. Never let a technically interesting use case override the score.
|
|
182
|
+
|
|
183
|
+
## Artifact (this IS the memory - write it as you work)
|
|
184
|
+
|
|
185
|
+
**`reality.md`** - the readout the FDE takes into the sponsor meeting. Keep the three schema lines the dashboard reads (`Working theory` / `Evidence` / `Differs from brief how`). Then the decision frame:
|
|
186
|
+
|
|
187
|
+
```markdown
|
|
188
|
+
# Reality (actual problem)
|
|
189
|
+
**Working theory:** <the real problem, one sentence>
|
|
190
|
+
**Evidence:** <workaround/data/quote, source, day>
|
|
191
|
+
**Differs from brief how:** <delta, with evidence>
|
|
192
|
+
**Situation:** <what the floor already treats as true>
|
|
193
|
+
**Complication:** <what forces a decision now>
|
|
194
|
+
**Question:** <one decision-shaped sentence>
|
|
195
|
+
**Answer-space:** confirm brief / descope / rescope / pause - and what a yes looks like
|
|
196
|
+
**Implication for build:** <first change they can see>
|
|
197
|
+
**Validated with:** <who, when>
|
|
198
|
+
```
|
|
199
|
+
|
|
200
|
+
**`terrain.md`** - the map every later phase loads:
|
|
201
|
+
```markdown
|
|
202
|
+
# Terrain
|
|
203
|
+
**Stack:** <lang/framework/build, age>
|
|
204
|
+
**Hotspots (handle with care):** <file - churn n/90d - tests: none/weak/ok - why it matters>
|
|
205
|
+
**AI components:** <file - model - fallback? - observability?>
|
|
206
|
+
**Data flow:** <entry → transform → store → exit>
|
|
207
|
+
**Test landscape:** <covered / gaps / lies>
|
|
208
|
+
**Unknowns:** <named explicitly - an honest gap beats a confident guess>
|
|
209
|
+
|
|
210
|
+
## Operating map (exception-led)
|
|
211
|
+
| Exception / break | Who notices first | What they do today | System of record then | Blast | Evidence |
|
|
212
|
+
|-------------------|-------------------|--------------------|----------------------|-------|----------|
|
|
213
|
+
| <break> | <role> | <workaround> | <sheet/DB/person> | CRITICAL / LOAD-BEARING / CONVENIENCE | <who/day> |
|
|
214
|
+
```
|
|
215
|
+
|
|
216
|
+
Every line carries its evidence. `(churn: 47/90d)` `(ops lead, Day 5)` `(stated, unverified)`.
|
|
217
|
+
|
|
218
|
+
**`assumptions.md`** - update statuses from what discovery proved or disproved. Seed any new OPEN assumptions the brief never named. CRITICAL + OPEN must be named in the checkpoint.
|
|
219
|
+
|
|
220
|
+
## Checkpoint (before any build)
|
|
221
|
+
|
|
222
|
+
Present to the FDE, five things, one paragraph each - no padding:
|
|
223
|
+
1. The Question, then the real problem, with the two strongest pieces of evidence.
|
|
224
|
+
2. The top 3 risk areas of the codebase, one line of why each.
|
|
225
|
+
3. What must not be touched without characterisation tests.
|
|
226
|
+
4. The exception-led operating map: the two breaks that matter most, who owns the workaround, and where shadow systems live.
|
|
227
|
+
5. The Answer-space: confirm brief / descope / rescope - and the decision it puts in front of the sponsor.
|
|
228
|
+
|
|
229
|
+
If discovery revealed the problem is 3× the brief: the FDE tells the customer **before** telling themselves it's manageable. Lead with evidence, offer three paths (descope / rescope / pause-and-plan), confirm any reset in writing - update `success.md` and `brief.md` before continuing.
|
|
230
|
+
|
|
231
|
+
## If you've formed three wrong reads
|
|
232
|
+
|
|
233
|
+
Stop. Don't form a fourth hypothesis. Three disproven reads means the brief is actively misleading - usually the person who briefed doesn't know, or knows and can't say. Change method: stop analysing the system, ask three people separately "if you had to bet on what's actually wrong here, what would you say?" The thing they all hesitate before saying is the real problem.
|
|
234
|
+
|
|
235
|
+
## Worked example
|
|
236
|
+
|
|
237
|
+
Acme's brief blamed missing monitoring. Discovery goes to the workaround first.
|
|
238
|
+
|
|
239
|
+
`git log` shows the reconciliation module at 47 commits/90d with no tests, all from one author who left in February. Marco (ops lead) turns out to keep a spreadsheet: every morning he re-runs the job manually and eyeballs the totals - a habit nobody mentioned because to him it is just the job. That spreadsheet is the system of record when the job fails, which is the actual finding.
|
|
240
|
+
|
|
241
|
+
`reality.md` keeps the schema, then the frame. **Working theory:** the job has no owner, and the manual re-run masks failures for a day. **Evidence:** Marco's sheet, Day 5; two silent failures since March, finance escalation Mar 14. **Differs from brief how:** alerting existed last year and was disabled - adding it again without an owner reproduces the same outcome. **Situation:** Marco re-runs the job every morning and the spreadsheet is truth when it fails. **Complication:** two silent failures since March already hit finance, and the author of the module left in February. **Question:** should Priya fund a named owner on the failure path, or fund alerting and accept the same miss in six months? **Answer-space:** fund ownership / fund alerting-as-theatre / pause until she names who acks. `terrain.md` gets the hotspot row and an operating-map row: `job fails silently → Marco notices next morning → re-runs by hand → spreadsheet is truth → LOAD-BEARING (Marco, Day 5)`.
|
|
242
|
+
|
|
243
|
+
Checkpoint to the FDE leads with that Question, not a tour of the repo.
|
|
244
|
+
|
|
245
|
+
## Principles
|
|
246
|
+
|
|
247
|
+
- The brief is a hypothesis until evidence confirms it.
|
|
248
|
+
- No scan until the Question is one decision the signer must make.
|
|
249
|
+
- The workaround is more honest than the requirements document.
|
|
250
|
+
- Churn data + the human's "don't touch that" pointing at the same module = the map is true.
|
|
251
|
+
- Never modify code before the terrain map exists.
|
|
252
|
+
- Scan wide, read deep only on hotspots.
|
|
253
|
+
|
|
254
|
+
Before changing a surprising workaround, use the targeted history check in [audit](audit.md#before-changing-an-unfamiliar-workaround). Inspect only the implicated file or region; commit messages supply clues, not proof of current requirements.
|
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
# eval-pack - Evaluate the model's allowed behavior
|
|
2
|
+
|
|
3
|
+
**Enter when:** AI, LLM, RAG, or agent behavior needs evidence before an experiment, release, or material expansion. Non-AI work skips this method.
|
|
4
|
+
|
|
5
|
+
Use [task context](task-context.md). Supplied permitted context and an evaluation report are sufficient without `.fde/`. In a coordinated engagement, use the existing trust/terrain context and keep the report in `evals.md`; read only privacy-safe views.
|
|
6
|
+
|
|
7
|
+
## Method
|
|
8
|
+
|
|
9
|
+
1. **Define the evaluated surface.** Name the model judgment, inputs, outputs, downstream actions, environment, and relevant failure impact. Separate model quality from deterministic tool authorization and application checks. Document the actual allowed action boundary and its source; missing authority remains unknown.
|
|
10
|
+
2. **Choose cases by risk and coverage.** Use permitted historical examples, expert-labeled cases, or clearly marked synthetic fixtures. Cover relevant segments, boundary conditions, known failure modes, and critical harms. Record input, expected outcome/rubric, provenance, and critical-failure rule per case. Keep evaluation cases separate from tuning where possible; no fixed case count proves safety.
|
|
11
|
+
3. **Agree the pass rule before the run.** Define quality thresholds, critical failures, coverage expectations, and acceptable uncertainty for this use. Use deterministic checks where possible and inspect subjective labels or judge reliability. Propose missing criteria for agreement; do not manufacture acceptance from the observed scores.
|
|
12
|
+
4. **Run the actual evaluated path.** Record model/provider version, prompts/configuration, retrieval corpus or tools, application revision, environment, fixtures, and run date. Repeat where variability matters. Report totals, per-segment results, critical failures, and limitations using [verification](verification.md). A model-only run does not prove the agent's tool boundary works.
|
|
13
|
+
5. **Verify action authority and controls.** Human approval is required where the user's policy or task requires it. Already agreed bounded automation may run within its documented actions, identities, environments, and limits; do not require fresh approval for every authorized action. Check enforcement outside the model, least privilege, input/output validation, cost/rate limits, stop conditions, observability, and recovery as applicable. Unknown or exceeded authority blocks those actions. Evaluation success never grants new authority.
|
|
14
|
+
6. **Make a scoped verdict.** Report **SHIP** only when agreed criteria pass, critical failures are zero, applicable authority/control checks pass, and material coverage gaps are resolved or the release is explicitly narrowed by the responsible decision-maker. Otherwise report **NO-SHIP** with the smallest corrective step: fix, gather evidence, descope, or reconsider the judgment surface. A SHIP verdict is technical evidence for the stated scope, not permission to deploy.
|
|
15
|
+
|
|
16
|
+
## Deliverable and acceptance
|
|
17
|
+
|
|
18
|
+
Return the suite/source, thresholds, counts and segments, top failure modes, control evidence, human-review gate or bounded automation authority, limitations, and dated verdict. Record unknown values honestly. Reevaluate after changes that affect model behavior, retrieval, tool permissions, or data conditions; cite why unchanged evidence remains applicable rather than implying a rerun.
|
|
19
|
+
|
|
20
|
+
When coordinated, append a concise eval receipt to delivery records. For release use [ship](ship.md). For ongoing use define the drift signals, sample policy permitted by data handling rules, owner, and conditions that suspend or narrow automation. Do not store secrets, raw `<private>` data, or hidden chain-of-thought in reports.
|
|
21
|
+
|
|
22
|
+
## Principles
|
|
23
|
+
|
|
24
|
+
- Thresholds and authority come from the agreed contract, never from a convenient observed result.
|
|
25
|
+
- Critical failures block the evaluated release scope; disclose coverage and uncertainty.
|
|
26
|
+
- Bound automation with enforceable controls, and require human review where the policy requires it.
|
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
# integrate - Prove the system boundary
|
|
2
|
+
|
|
3
|
+
**Enter when:** connecting an API, data source, SDK, event stream, tool, or service, or changing its contract.
|
|
4
|
+
|
|
5
|
+
Start from [task context](task-context.md). Permitted supplied context is enough; `.fde/` is optional. Use the customer's existing clients, authentication, fixtures, and diagnostic tools. Do not create another integration platform to make one connection.
|
|
6
|
+
|
|
7
|
+
## Method
|
|
8
|
+
|
|
9
|
+
1. Map producer, consumer, owner, direction, and side effects. Inspect the actual installed version and local implementation; verify uncertain behavior against current official documentation. Identify the relevant schema, authentication scopes, network boundary, and permitted test environment.
|
|
10
|
+
2. Write the acceptance example: an input at the real boundary and the observable downstream result. Include a rejection or failure example. Separate configuration validity, successful authentication, transport connectivity, contract compatibility, and end-to-end behavior; none proves the next.
|
|
11
|
+
3. Inspect credentials by presence and required scope without printing values. Use existing secret storage. Check data classification and retention before moving data; never pass raw `<private>` blocks into a model. Prefer sanitized or synthetic cases approved for the target environment.
|
|
12
|
+
4. Implement the narrow adapter using native repository patterns. Validate external inputs and model outputs, bound timeouts and retries, preserve error context without leaking payloads, and handle cancellation. For writes, establish idempotency or duplicate detection before retries; for events, check ordering, replay, and poison messages as applicable.
|
|
13
|
+
5. Exercise a permitted success case and relevant failures: denied access, malformed data, rate limit, timeout, duplicate delivery, or partial completion. Trace correlation IDs or safe evidence across both sides. A mock proves client behavior only; if live access is unavailable, report that gap instead of claiming an integration works.
|
|
14
|
+
6. Check cleanup and recovery for test side effects. Use [verification](verification.md) for receipts and [review](review.md) for security or data-contract changes. Route deployment through [ship](ship.md) only when authorized.
|
|
15
|
+
|
|
16
|
+
## Deliverable and acceptance
|
|
17
|
+
|
|
18
|
+
Return the boundary contract, changed paths, environment, evidence at each tested layer, and remaining dependencies with owners when known. Done requires the agreed end-to-end result or an explicit narrower agreed scope. Do not silently replace live acceptance with a stub. In engagement mode, update the terrain/delivery record with confirmed facts; otherwise return the receipt directly.
|
|
@@ -0,0 +1,167 @@
|
|
|
1
|
+
# plan - Sequence the work
|
|
2
|
+
|
|
3
|
+
**Context:** apply [task context and evidence](task-context.md). Standalone planning evaluates supplied facts directly; it does not require initialized engagement records.
|
|
4
|
+
|
|
5
|
+
**Enter when:** scope is understood and the work needs breaking down - a slice, a phase, or the whole delivery.
|
|
6
|
+
|
|
7
|
+
**Read first:** `reality.md`, `success.md`, `terrain.md`, `stakeholders.md`. Load `business-case.md` if poc produced one. Not the full folder.
|
|
8
|
+
|
|
9
|
+
**On an initialized engagement, before a new delivery plan or material scope change:** run `fde doctor --ready`. For standalone planning, check the supplied outcome, scope, acceptance and authority directly; do not initialize records to run this validator. Missing binary success or a named customer-side signer blocks progression: review the proposed acceptance check and authority with the FDE first. Use a test/input and observable pass/fail under **Done when:** or **Acceptance check:**. A number, role, or successful demo alone is insufficient. Do not invent missing facts to pass lint. Routine reversible fixes within confirmed scope reuse the existing signer, acceptance criteria, and engineering plan; record verification without reopening settled decisions.
|
|
10
|
+
|
|
11
|
+
## Validation gate (confirm understanding, clarify where it elevates)
|
|
12
|
+
|
|
13
|
+
Before planning, state what you're working from in 2-3 lines:
|
|
14
|
+
|
|
15
|
+
> "Planning against: [success definition from success.md]. Scope boundary: [out-of-scope items]. Reality check: [brief aligns with reality.md / or note the delta]."
|
|
16
|
+
|
|
17
|
+
Then check - probe ONLY if it prevents a bad plan:
|
|
18
|
+
|
|
19
|
+
1. **Success is measurable.** If "done" is vague ("make it better") → rephrase it: "I'm reading success as: [specific measurable outcome]. That the target?"
|
|
20
|
+
2. **Reality matches the brief.** If discovery contradicted the brief → name it: "Discovery found [X] but the brief says [Y]. Planning against reality unless you say otherwise."
|
|
21
|
+
3. **Out-of-scope exists.** If missing → one line: "Nothing's marked out-of-scope yet. That means every new request is implicitly in. Worth defining now or after the first plan draft?"
|
|
22
|
+
|
|
23
|
+
State your read, let the FDE correct, then plan.
|
|
24
|
+
|
|
25
|
+
An FDE plan is not a sprint backlog. The technical sequence is the easy part. The hard part is when to show progress, who approves the next phase, and where trust is thin enough that two silent weeks read as failure. A technically correct plan that ignores engagement politics fails on schedule.
|
|
26
|
+
|
|
27
|
+
## Method (you do this work)
|
|
28
|
+
|
|
29
|
+
**0. Lock scope first.** Read `success.md`, `assumptions.md`, and the **Question** on `reality.md`. Make the boundary explicit using the supplied request; ask if an ambiguity changes the commitment. Investigate a critical open assumption before planning work that depends on it. If the problem itself is unclear, use discovery for that gap; absent filenames do not block a plan supported by supplied facts.
|
|
30
|
+
|
|
31
|
+
**Reuse check.** Before sequencing a build, compare the requested solution with the smallest existing capability or operating change that could satisfy the same acceptance test. Cite the relevant repo/config/workaround evidence. Record why reuse is sufficient or insufficient in `decisions.md`; include “no new code” when supported. A request for AI does not establish that a model is needed. If a host engineering pack already has an approved implementation plan, reference it from `decisions.md`; do not generate a parallel user-story backlog.
|
|
32
|
+
|
|
33
|
+
**1. Work backwards from success.** What's the last thing that must be true before done? And before that? That's the dependency chain - not a wish list.
|
|
34
|
+
|
|
35
|
+
**2. Front-load the fragile.** Check `terrain.md` hotspots. Risky modules go early - fail fast, not in week three.
|
|
36
|
+
|
|
37
|
+
**3. One user action per change.** Each task delivers something visible and testable ("user submits form, sees it saved"), never a layer ("build the database layer"). See `ship`.
|
|
38
|
+
|
|
39
|
+
**4. Size to a coherent, verifiable outcome.** Split unrelated work and tasks too complex to review or recover safely. Use bounded review sections for large cohesive changes; elapsed time and line count are signals to examine, not universal limits.
|
|
40
|
+
|
|
41
|
+
**5. AI components get explicit eval tasks.** "Output validated on 50 real production examples," "fallback tested under model unavailability," "inputs/outputs logging to <destination>" - these are pre-conditions of shipping, in the plan before build starts.
|
|
42
|
+
|
|
43
|
+
**6. Stakeholder touchpoints every 2-3 tasks.** "Show progress to <name from stakeholders.md>." Not ceremony: a customer who sees small wins stays bought in; silence gets filled with doubt.
|
|
44
|
+
|
|
45
|
+
**7. End with a kill list.** Every plan names what you will **not** do this phase. If everything is "later," you have no plan - you have a wish list. Cap **Now** at 3 PRs (same discipline as pick-three).
|
|
46
|
+
|
|
47
|
+
**Acceptance criteria gate:** no task moves to build without written happy-path AND unhappy-path criteria. Can't write them = the task isn't understood; the open question goes to the customer **before** the task starts. Vague criteria surface later as scope creep and rework.
|
|
48
|
+
|
|
49
|
+
## Artifact
|
|
50
|
+
|
|
51
|
+
The plan goes to **`decisions.md`** - always. Build reads the plan from `decisions.md`; anywhere else and the build starts blind.
|
|
52
|
+
|
|
53
|
+
A plan is **not done** until all four blocks exist:
|
|
54
|
+
|
|
55
|
+
```markdown
|
|
56
|
+
## Plan - <date>
|
|
57
|
+
### Now (max 3)
|
|
58
|
+
Task N: <outcome, not activity>
|
|
59
|
+
Delivers: <what someone can see/test>
|
|
60
|
+
Accepts: <happy path> / <unhappy path>
|
|
61
|
+
Touches: <files/systems - blast radius declared upfront>
|
|
62
|
+
Risk: <what could go wrong + fallback>
|
|
63
|
+
Kill if: <the observation that voids this slice - copy from assumptions.md How we test, or the check that means stop>
|
|
64
|
+
Verify: <specific check>
|
|
65
|
+
Value promised: <business unit change this slice claims>
|
|
66
|
+
Baseline: <value + source/date/window/environment, or pending + measurement owner>
|
|
67
|
+
Acceptance owner: <name + authority source, or unknown - ask: who can accept?>
|
|
68
|
+
Evidence to collect: <before/after check, sample/window, environment, and receipt location>
|
|
69
|
+
Reuse: <existing capability used, or evidence it cannot satisfy the criteria>
|
|
70
|
+
|
|
71
|
+
### Next
|
|
72
|
+
- ...
|
|
73
|
+
|
|
74
|
+
### Later
|
|
75
|
+
- ...
|
|
76
|
+
|
|
77
|
+
### Kill list (explicitly not this phase)
|
|
78
|
+
| Item | Why killed / deferred | Who accepted |
|
|
79
|
+
|------|----------------------|--------------|
|
|
80
|
+
| <rewrite / nice-to-have / political ask> | <evidence> | <name, date> |
|
|
81
|
+
```
|
|
82
|
+
|
|
83
|
+
In `Who accepted`, distinguish a proposed deferral from an agreement: use `pending` until a named person accepted this scope with a dated source. Sponsorship alone is not approval of every plan detail.
|
|
84
|
+
|
|
85
|
+
No kill list → not a finished plan. Reopen with the FDE until the deferrals are written.
|
|
86
|
+
## Checkpoint
|
|
87
|
+
|
|
88
|
+
Walk the FDE through: sequence + why this order, where the fragile work sits, where the touchpoints land, the acceptance gate and **Kill if** on task 1, and the kill list. One question: "Which stakeholder sees the first visible slice, and when?" Second: "Who accepted what we are not doing?" Third: "What observation stops task 1 this week?"
|
|
89
|
+
|
|
90
|
+
## Method - estimation (when the sponsor asks "how long, how much?")
|
|
91
|
+
|
|
92
|
+
Every FDE gets asked this in week one. The honest answer is a range, not a number. A single-point estimate is a promise; a range is a professional assessment.
|
|
93
|
+
|
|
94
|
+
**The 3-point method:**
|
|
95
|
+
1. **Best case** - everything goes right, no surprises, team has capacity. This is what the sponsor wants to hear.
|
|
96
|
+
2. **Expected case** - normal friction: one discovery changes the plan, one integration takes longer, one approval cycle stalls. This is what to plan against.
|
|
97
|
+
3. **Worst case** - a major unknown surfaces, a dependency fails, a key person is unavailable. This is what to protect against.
|
|
98
|
+
|
|
99
|
+
**Present as:** "2-4 weeks expected, could stretch to 6 if [named risk]." Never give one number.
|
|
100
|
+
|
|
101
|
+
**The sizing table:**
|
|
102
|
+
|
|
103
|
+
| Slice | Complexity | Dependencies | Unknowns | Estimate (expected) |
|
|
104
|
+
|-------|-----------|--------------|----------|---------------------|
|
|
105
|
+
| _per vertical slice from the plan_ | Low/Med/High | Named | Named | X days/weeks |
|
|
106
|
+
|
|
107
|
+
**Rules:**
|
|
108
|
+
- Estimate in weeks for engagements > 1 month. Days for < 1 month.
|
|
109
|
+
- Add 30% buffer for integration work (it always takes longer).
|
|
110
|
+
- Add 50% buffer for AI/ML work (eval cycles are unpredictable).
|
|
111
|
+
- Name assumptions explicitly: "assumes API docs are accurate", "assumes staging environment exists."
|
|
112
|
+
- Each named assumption needs a **kill observation**: the result that voids the estimate. Copy it from `assumptions.md` → How we test. No kill observation = it is not an assumption, it is hope.
|
|
113
|
+
- Revisit estimates every 2 weeks. An estimate that never updates is fiction.
|
|
114
|
+
|
|
115
|
+
Write estimates to `decisions.md` under `## Sizing`. Include the assumptions - when they break, the estimate changes and the FDE has evidence for the conversation.
|
|
116
|
+
|
|
117
|
+
## Method - migration strategy (when the engagement is "move from X to Y")
|
|
118
|
+
|
|
119
|
+
Migrations are the most common enterprise FDE engagement. The strategy precedes the plan:
|
|
120
|
+
|
|
121
|
+
**Step 1: Classify the migration type.**
|
|
122
|
+
|
|
123
|
+
| Type | What it means | Risk profile |
|
|
124
|
+
|------|---------------|-------------|
|
|
125
|
+
| **Rehost** (lift-and-shift) | Same code, different infrastructure | Low code risk, high ops risk |
|
|
126
|
+
| **Replatform** | Minor code changes to use new platform features | Medium risk, clear scope |
|
|
127
|
+
| **Refactor** | Rewrite components to fit the new architecture | High risk, scope creep magnet |
|
|
128
|
+
| **Replace** | Buy/build new, retire old | Highest risk, requires parallel running |
|
|
129
|
+
| **Retire** | Turn off, nobody uses it | Politically hard, technically easy |
|
|
130
|
+
|
|
131
|
+
**Step 2: Map the dependency graph.** What calls what. What breaks if this moves first. The migration order is the reverse of the dependency chain - leaf nodes first, core last.
|
|
132
|
+
|
|
133
|
+
**Step 3: Define the cutover strategy.**
|
|
134
|
+
- **Big bang** - everything moves at once. Fast but catastrophic on failure. Only for small systems.
|
|
135
|
+
- **Strangler fig** - new traffic to new system, old traffic drains. Safe but slow. Preferred for anything load-bearing.
|
|
136
|
+
- **Parallel run** - both systems run, outputs compared. Expensive but safest for data-critical systems.
|
|
137
|
+
|
|
138
|
+
**Step 4: Write the rollback before the migration starts.** "If we move service X and it fails, we route back to old within [time]." No rollback = no migration.
|
|
139
|
+
|
|
140
|
+
**Step 5: Define success metrics per phase.** Not "migration complete" - that's a project plan. "Error rate same or lower, latency within 10%, zero data loss, team can operate without FDE." Measurable, per service.
|
|
141
|
+
|
|
142
|
+
Write migration strategy to `decisions.md` under `## Migration`. Each service gets a row: type, order, cutover method, rollback, success metric.
|
|
143
|
+
|
|
144
|
+
## When the plan changes mid-engagement
|
|
145
|
+
|
|
146
|
+
Never quietly update tasks. Name the reset: update `reality.md` and `success.md`, one paragraph in `decisions.md` - what changed, why, new sequence. An undocumented reset looks like drift; a documented one looks like the FDE caught something important.
|
|
147
|
+
|
|
148
|
+
## Worked example
|
|
149
|
+
|
|
150
|
+
Acme, after discover: the reconciliation job is unowned, Marco's spreadsheet is the real fallback.
|
|
151
|
+
|
|
152
|
+
**Now** is three tasks, not eight. Task 1 is *failures reach a named human* - delivers a page to a rota, accepts "kill the job mid-run → the on-call is paged within 15 min", touches the job wrapper and the alert config, rollback is re-disable the route, **Kill if:** a real failure page is acked by nobody on the rota (the *finance would act* assumption, DISPROVED if Marco is the only name that answers), verify by killing it in staging. Value promised: `risk-mitigation - a silent failure becomes a 15-minute one`.
|
|
153
|
+
|
|
154
|
+
The kill list in `decisions.md` is where the plan earns its keep: the rewrite of the reconciliation service that Tom keeps proposing goes there - *deferred, the failure mode is ownership not architecture (Priya accepted, Jun 12)* - along with the finance dashboard finance asked for directly. Both stay visible so the same argument is not re-litigated in week 4 without a receipt.
|
|
155
|
+
|
|
156
|
+
First visible slice goes to Marco, not Priya: he is the one whose morning changes, and his confirmation is what makes the sponsor update true.
|
|
157
|
+
|
|
158
|
+
## Principles
|
|
159
|
+
|
|
160
|
+
- Plan from success backwards, not from today forwards.
|
|
161
|
+
- Fragile zones early. Fail fast.
|
|
162
|
+
- Every 2-3 tasks, a stakeholder touchpoint. Trust decays without visibility.
|
|
163
|
+
- No written acceptance criteria, no build.
|
|
164
|
+
- No kill list, no finished plan.
|
|
165
|
+
- No **Kill if** on a Now PR, that PR is hope.
|
|
166
|
+
- Estimates are ranges, not promises. Name the assumptions and the observation that voids them.
|
|
167
|
+
- Migrations: leaf nodes first, core last. Rollback before cutover.
|