fdeops 4.0.4 → 4.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +1 -1
- package/README.md +33 -6
- package/bin/check.js +14 -23
- package/bin/generate-skills.js +71 -0
- package/bin/install.js +2 -1
- package/bin/skill-catalog.js +17 -0
- package/mcp/fdeops-ingest/package.json +2 -2
- package/package.json +4 -3
- package/plugin.json +2 -2
- package/skills/fde/SKILL.md +18 -10
- package/skills/fde/references/board-memo.md +1 -1
- package/skills/fde/references/build.md +20 -0
- package/skills/fde/references/business-case.md +9 -7
- package/skills/fde/references/close.md +5 -3
- package/skills/fde/references/debug.md +18 -0
- package/skills/fde/references/encode-pattern.md +10 -8
- package/skills/fde/references/eval-pack.md +16 -33
- package/skills/fde/references/hold-scope.md +9 -7
- package/skills/fde/references/integrate.md +18 -0
- package/skills/fde/references/plan.md +4 -2
- package/skills/fde/references/poc.md +5 -3
- package/skills/fde/references/qa.md +18 -0
- package/skills/fde/references/review.md +22 -60
- package/skills/fde/references/ship.md +42 -287
- package/skills/fde/references/task-context.md +12 -0
- package/skills/fde/references/test-assumptions.md +2 -2
- package/skills/fde/references/three-options.md +19 -27
- package/skills/fde/references/verification.md +31 -0
- package/skills/fde-build/SKILL.md +21 -0
- package/skills/fde-build/references/build.md +20 -0
- package/skills/fde-build/references/debug.md +18 -0
- package/skills/fde-build/references/eval-pack.md +26 -0
- package/skills/fde-build/references/integrate.md +18 -0
- package/skills/fde-build/references/qa.md +18 -0
- package/skills/fde-build/references/review.md +39 -0
- package/skills/fde-build/references/ship.md +73 -0
- package/skills/fde-build/references/task-context.md +12 -0
- package/skills/fde-build/references/verification.md +31 -0
- package/skills/fde-debug/SKILL.md +21 -0
- package/skills/fde-debug/references/build.md +20 -0
- package/skills/fde-debug/references/debug.md +18 -0
- package/skills/fde-debug/references/eval-pack.md +26 -0
- package/skills/fde-debug/references/integrate.md +18 -0
- package/skills/fde-debug/references/qa.md +18 -0
- package/skills/fde-debug/references/review.md +39 -0
- package/skills/fde-debug/references/ship.md +73 -0
- package/skills/fde-debug/references/task-context.md +12 -0
- package/skills/fde-debug/references/verification.md +31 -0
- package/skills/fde-discover/SKILL.md +21 -0
- package/skills/fde-discover/references/audit.md +71 -0
- package/skills/fde-discover/references/discover.md +254 -0
- package/skills/fde-discover/references/task-context.md +12 -0
- package/skills/fde-evaluate/SKILL.md +21 -0
- package/skills/fde-evaluate/references/build.md +20 -0
- package/skills/fde-evaluate/references/debug.md +18 -0
- package/skills/fde-evaluate/references/eval-pack.md +26 -0
- package/skills/fde-evaluate/references/integrate.md +18 -0
- package/skills/fde-evaluate/references/qa.md +18 -0
- package/skills/fde-evaluate/references/review.md +39 -0
- package/skills/fde-evaluate/references/ship.md +73 -0
- package/skills/fde-evaluate/references/task-context.md +12 -0
- package/skills/fde-evaluate/references/verification.md +31 -0
- package/skills/fde-feedback/SKILL.md +21 -0
- package/skills/fde-feedback/references/encode-pattern.md +96 -0
- package/skills/fde-feedback/references/task-context.md +12 -0
- package/skills/fde-handoff/SKILL.md +21 -0
- package/skills/fde-handoff/references/close.md +66 -0
- package/skills/fde-handoff/references/encode-pattern.md +96 -0
- package/skills/fde-handoff/references/task-context.md +12 -0
- package/skills/fde-integrate/SKILL.md +21 -0
- package/skills/fde-integrate/references/build.md +20 -0
- package/skills/fde-integrate/references/debug.md +18 -0
- package/skills/fde-integrate/references/eval-pack.md +26 -0
- package/skills/fde-integrate/references/integrate.md +18 -0
- package/skills/fde-integrate/references/qa.md +18 -0
- package/skills/fde-integrate/references/review.md +39 -0
- package/skills/fde-integrate/references/ship.md +73 -0
- package/skills/fde-integrate/references/task-context.md +12 -0
- package/skills/fde-integrate/references/verification.md +31 -0
- package/skills/fde-options/SKILL.md +21 -0
- package/skills/fde-options/references/business-case.md +90 -0
- package/skills/fde-options/references/task-context.md +12 -0
- package/skills/fde-options/references/test-assumptions.md +102 -0
- package/skills/fde-options/references/three-options.md +90 -0
- package/skills/fde-poc/SKILL.md +21 -0
- package/skills/fde-poc/references/audit.md +71 -0
- package/skills/fde-poc/references/build.md +20 -0
- package/skills/fde-poc/references/business-case.md +90 -0
- package/skills/fde-poc/references/debug.md +18 -0
- package/skills/fde-poc/references/discover.md +254 -0
- package/skills/fde-poc/references/eval-pack.md +26 -0
- package/skills/fde-poc/references/integrate.md +18 -0
- package/skills/fde-poc/references/plan.md +167 -0
- package/skills/fde-poc/references/poc.md +55 -0
- package/skills/fde-poc/references/qa.md +18 -0
- package/skills/fde-poc/references/review.md +39 -0
- package/skills/fde-poc/references/ship.md +73 -0
- package/skills/fde-poc/references/task-context.md +12 -0
- package/skills/fde-poc/references/test-assumptions.md +102 -0
- package/skills/fde-poc/references/three-options.md +90 -0
- package/skills/fde-poc/references/verification.md +31 -0
- package/skills/fde-qa/SKILL.md +21 -0
- package/skills/fde-qa/references/build.md +20 -0
- package/skills/fde-qa/references/debug.md +18 -0
- package/skills/fde-qa/references/eval-pack.md +26 -0
- package/skills/fde-qa/references/integrate.md +18 -0
- package/skills/fde-qa/references/qa.md +18 -0
- package/skills/fde-qa/references/review.md +39 -0
- package/skills/fde-qa/references/ship.md +73 -0
- package/skills/fde-qa/references/task-context.md +12 -0
- package/skills/fde-qa/references/verification.md +31 -0
- package/skills/fde-readout/SKILL.md +21 -0
- package/skills/fde-readout/references/board-memo.md +108 -0
- package/skills/fde-readout/references/business-case.md +90 -0
- package/skills/fde-readout/references/readout.md +69 -0
- package/skills/fde-readout/references/task-context.md +12 -0
- package/skills/fde-review/SKILL.md +21 -0
- package/skills/fde-review/references/build.md +20 -0
- package/skills/fde-review/references/debug.md +18 -0
- package/skills/fde-review/references/eval-pack.md +26 -0
- package/skills/fde-review/references/integrate.md +18 -0
- package/skills/fde-review/references/qa.md +18 -0
- package/skills/fde-review/references/review.md +39 -0
- package/skills/fde-review/references/ship.md +73 -0
- package/skills/fde-review/references/task-context.md +12 -0
- package/skills/fde-review/references/verification.md +31 -0
- package/skills/fde-scope/SKILL.md +21 -0
- package/skills/fde-scope/references/hold-scope.md +83 -0
- package/skills/fde-scope/references/task-context.md +12 -0
- package/skills/fde-ship/SKILL.md +21 -0
- package/skills/fde-ship/references/build.md +20 -0
- package/skills/fde-ship/references/debug.md +18 -0
- package/skills/fde-ship/references/eval-pack.md +26 -0
- package/skills/fde-ship/references/integrate.md +18 -0
- package/skills/fde-ship/references/qa.md +18 -0
- package/skills/fde-ship/references/review.md +39 -0
- package/skills/fde-ship/references/ship.md +73 -0
- package/skills/fde-ship/references/task-context.md +12 -0
- package/skills/fde-ship/references/verification.md +31 -0
|
@@ -0,0 +1,39 @@
|
|
|
1
|
+
# review - Assess the actual change
|
|
2
|
+
|
|
3
|
+
**Enter when:** a diff, proposed merge, or review comment needs assessment against agreed behavior and constraints.
|
|
4
|
+
|
|
5
|
+
Use [task context](task-context.md). Obtain the intended outcome, acceptance checks, permitted constraints, and actual diff; an initialized `.fde/` is unnecessary. Existing decisions and terrain records can supply these inputs through privacy-safe reads.
|
|
6
|
+
|
|
7
|
+
## Establish what was reviewed
|
|
8
|
+
|
|
9
|
+
Identify the repository, base and head revision, staged/unstaged changes, and relevant untracked files. Read applicable instructions and the full in-scope diff, then inspect callers and tests where needed. A committed-range diff alone omits working-tree edits. Record missing files or unavailable context as limitations.
|
|
10
|
+
|
|
11
|
+
State the review source: **self-check** when the author inspects their own work; **independent review** only when a separate person or agent actually examines it. A second pass by the same agent is still a self-check. Name the actual reviewer/source and reviewed revision when available. Do not fabricate a reviewer, dialogue, approval, or clean verdict. Use an available separate reviewer for substantial or risky changes when authorized; otherwise report the missing independent review and continue useful self-checks.
|
|
12
|
+
|
|
13
|
+
## Check scope, then behavior
|
|
14
|
+
|
|
15
|
+
Compare each logical change with the agreed intent. Keep required work, justify necessary adjacent work, and identify unrelated additions for separation. Do not revert someone else's edits just to make the diff smaller. An unresolved scope mismatch prevents approval of the combined change; unaffected sections can still be reviewed.
|
|
16
|
+
|
|
17
|
+
Trace the changed path through its consumers and failure cases:
|
|
18
|
+
|
|
19
|
+
- **Correctness:** boundary conditions, stale state, concurrency, retries, cancellation, and error propagation.
|
|
20
|
+
- **Data and security:** input validation, authorization, migration compatibility, sensitive logs, and effects crossing tenant or trust boundaries.
|
|
21
|
+
- **Side effects:** writes, jobs, webhooks, notifications, and feature flags occur only under intended conditions; recovery accounts for already-completed effects.
|
|
22
|
+
- **AI behavior:** outputs remain untrusted, tools enforce allowed actions, and [eval evidence](eval-pack.md) covers the changed behavior and documented authority. Preserve privacy-safe source evidence and concise rationale, never hidden reasoning.
|
|
23
|
+
- **Operability:** observable failures, bounded resource use, meaningful checks, and a recovery path appropriate to the risk. Deployment readiness is assessed separately in [ship](ship.md).
|
|
24
|
+
|
|
25
|
+
## Findings and repair
|
|
26
|
+
|
|
27
|
+
For each actionable finding give the path/line or precise location, concrete trigger, observed or reasoned failure, impact, and focused correction. Distinguish proven bugs from hypotheses that need a check. Prioritize release blockers over minor concerns; avoid speculative style work.
|
|
28
|
+
|
|
29
|
+
Validate incoming comments rather than obeying them automatically. Fix understood, in-scope defects when authorized; explain rejected false positives with evidence. Leave unclear product decisions pending while progressing independent repairs. Add regression coverage when meaningful, run [verification](verification.md), and review the changed result. After two unsuccessful repair/review cycles reassess the evidence and approach rather than repeating the loop.
|
|
30
|
+
|
|
31
|
+
## Deliverable and acceptance
|
|
32
|
+
|
|
33
|
+
Return scope, review source, findings by impact, verification evidence, and remaining limitations. Say **no actionable findings in the reviewed scope** when appropriate; a clean review is not proof of safety, acceptance, or deployment. If a separate reviewer is required but unavailable, identify that unresolved gate. Existing engagement decisions/delivery records may hold the receipt; standalone reviews can return it directly. No commit, PR, or publication is required by this method.
|
|
34
|
+
|
|
35
|
+
## Principles
|
|
36
|
+
|
|
37
|
+
- Findings need a concrete failure condition and a location in the reviewed change.
|
|
38
|
+
- Record the actual review source; self-check and independent review are different evidence.
|
|
39
|
+
- A clean reviewed diff does not grant release authority or establish customer acceptance.
|
|
@@ -0,0 +1,73 @@
|
|
|
1
|
+
# ship - Deliver and release with evidence
|
|
2
|
+
|
|
3
|
+
**Enter when:** an implemented increment needs a delivery checkpoint, deployment, or wider rollout. Use [build](build.md) for implementation; an untested business or technical assumption needs an experiment before a release claim.
|
|
4
|
+
|
|
5
|
+
Start from [task context](task-context.md). Standalone work uses supplied permitted context and a release receipt; it does not require `.fde/` initialization. In engagement mode, use privacy-safe views of confirmed context, decisions, terrain, success, delivery, applicable trust constraints, and AI evaluation evidence. Never load raw `<private>` blocks into a model. The `fde` CLI remains local-only; deployment uses the customer's authorized tools, never a new network capability inside `fde`.
|
|
6
|
+
|
|
7
|
+
## Establish the delivery contract
|
|
8
|
+
|
|
9
|
+
Identify the exact outcome, acceptance check, scope, affected users/systems, target environment, recovery mechanism, and who or what is authorized to accept and release it. Reuse confirmed authority and checks for routine work. Do not invent missing signers, permissions, measurements, or acceptance.
|
|
10
|
+
|
|
11
|
+
For an initialized engagement, run `fde doctor --ready` before a new delivery plan or material scope change. Missing binary success or a named customer-side signer blocks that planning progression until resolved. A passing doctor validates record structure, not connectivity, release readiness, or customer acceptance. Standalone work evaluates the supplied contract directly.
|
|
12
|
+
|
|
13
|
+
A customer delivery checkpoint must let the agreed decision-maker replay and reject the acceptance check through an interface they operate. Prefer their staging; otherwise use an agreed representative environment and disclose its owner and limitations. Local green proves only the local run. Routine fixes may share an agreed checkpoint; no fixed number of changes forces a ceremony.
|
|
14
|
+
|
|
15
|
+
## Prepare a reviewable increment
|
|
16
|
+
|
|
17
|
+
1. Inspect repository instructions, working tree, overlapping work, and the complete intended release diff. Include working-tree changes when testing an uncommitted candidate. Preserve unrelated work; separate unintended behavior before release.
|
|
18
|
+
2. Identify applicable before-state evidence and the changed outcome. Name dependencies, stop conditions, and irreversible effects. Existing applicable evidence may be reused with attribution, never represented as a fresh run.
|
|
19
|
+
3. Complete [verification](verification.md) and [review](review.md), proportional to the change and repository requirements. A self-check is not an independent review. Record command, revision, environment, date, result, and unrun checks. Exercise relevant operating exceptions and fallback paths, not just the happy path.
|
|
20
|
+
4. For AI behavior, obtain a scoped [eval verdict](eval-pack.md) for the candidate and applicable controls. A permitted bounded automation workflow remains permitted within its documented limits. A missing evaluation, failed critical case, or unknown action authority prevents release of that path; non-AI changes record eval as not applicable.
|
|
21
|
+
|
|
22
|
+
## Release gate
|
|
23
|
+
|
|
24
|
+
Before deployment, establish these facts from existing evidence or a necessary check. Missing material evidence blocks the dependent release step; continue independent preparation. Do not ask again for approval already provided within the same scope.
|
|
25
|
+
|
|
26
|
+
| Dimension | Required evidence |
|
|
27
|
+
|-----------|-------------------|
|
|
28
|
+
| Candidate | Exact revision/artifact, intended diff, dependencies, applicable required checks passing; no skipped failure presented as green |
|
|
29
|
+
| Target and access | Service/account/region, environment, authorized deployment identity and mechanism, secret provisioning without revealing values |
|
|
30
|
+
| Acceptance | Replayable check and agreed decision-maker/mechanism; record actual acceptance separately from readiness |
|
|
31
|
+
| Data and policy | Permitted data, applicable security/residency/change-window requirements, necessary approvals already recorded or obtained |
|
|
32
|
+
| Recovery | Applicable tested rollback, restore, compensation, or roll-forward within agreed recovery-time/data-loss limits; explicit authority for irreversible effects |
|
|
33
|
+
| Operations | Named release/recovery owner, runbook appropriate to risk, health and business signals, stop thresholds, observation coverage |
|
|
34
|
+
| AI, when applicable | Current applicable SHIP eval evidence, critical failures zero, enforced action boundary and required human review or documented bounded automation |
|
|
35
|
+
|
|
36
|
+
Check migration compatibility, old/new version coexistence, delayed jobs, caches, and already-emitted side effects where relevant. A code revert does not undo data loss or external writes. Reuse drill evidence only when the mechanism and relevant conditions are unchanged, explaining applicability. If recovery is only a plan, exercise it in a permitted representative environment before release.
|
|
37
|
+
|
|
38
|
+
Use the repository's existing secret scanning and security checks; avoid diagnostic commands that print credential matches. Retain sanitized references to results. Resolve material evidence gaps or obtain an explicit, authorized narrowing of the release; do not average critical blockers into a readiness score.
|
|
39
|
+
|
|
40
|
+
For a coordinated engagement, also connect the release to the agreed value bucket and baseline/target, and record a dated receipt for the affected operating path. An unmeasured result remains pending with a measurement next step; do not invent realized value to pass a gate.
|
|
41
|
+
|
|
42
|
+
## Deploy within authority
|
|
43
|
+
|
|
44
|
+
Execute only when the requested workflow authorizes deployment to this target and the applicable gates are met. Otherwise leave a concrete release candidate, exact deployment/recovery instructions, evidence, and the remaining authorization for review. A permission to implement or test is not permission to publish.
|
|
45
|
+
|
|
46
|
+
Use the customer's established pipeline and rollout mechanism. Select canary, staged exposure, blue/green, or direct rollout according to actual risk and platform capabilities; do not impose a universal cohort sequence. Define advance/abort thresholds and observation window before starting. If another operator must execute, record their handoff and report deployment pending until there is evidence it happened.
|
|
47
|
+
|
|
48
|
+
During rollout inspect health, errors, key user behavior, and side-effect integrity. Halt expansion on breached thresholds or critical harm and apply authorized containment/recovery. Do not continue merely because the deploy command exited successfully.
|
|
49
|
+
|
|
50
|
+
## Verify operation and hand off
|
|
51
|
+
|
|
52
|
+
Run permitted smoke and acceptance checks against the deployed candidate. Record deployment identity/time, observed signals, sample/window, failures, recovery actions, and remaining gaps. Define the pulse: metric, cadence, threshold, owner, and response. For AI, include permitted output sampling and drift/action-boundary monitoring.
|
|
53
|
+
|
|
54
|
+
Before wider exposure, verify expected load/cost, data pipeline behavior, ownership, support, and applicable governance for the proposed audience. Choose expansion conditions from evidence; a successful pilot does not establish readiness for an arbitrary larger scale. Measure adoption against the eligible users, expected workflow frequency, and agreed observation window; investigate misses without guessing their cause.
|
|
55
|
+
|
|
56
|
+
## Receipt and completion
|
|
57
|
+
|
|
58
|
+
Keep these claims separate: implemented, verified, deployed, measured outcome, and accepted. Include the candidate, target, command/pipeline, applicable checks and unrun checks, review source, evaluation where needed, authority source, recovery evidence, observation, and next owner/action. Attribute acceptance to its actual source and scope. A staging measurement is not production value, and a commit is not deployment.
|
|
59
|
+
|
|
60
|
+
In engagement mode, write confirmed implementation/decisions and delivery receipts under the existing record rules. Standalone work returns the same receipt or uses the repository's permitted release record. Committing, pushing, opening a PR, publishing, and notifying others are actions governed by the user's workflow, not mandatory steps imposed by this method.
|
|
61
|
+
|
|
62
|
+
## Worked example
|
|
63
|
+
|
|
64
|
+
A freight team agrees that dispatchers can retry a failed export once without creating a duplicate shipment. The change uses the existing queue and ops screen. The developer records a failing duplicate-delivery case, implements idempotency, and passes the relevant checks on a named candidate. QA observes both the retry status and the single downstream record on permitted staging fixtures. A separate reviewer examines the queue race; the receipt names that review and the revision.
|
|
65
|
+
|
|
66
|
+
The export service has an approved staged-release workflow. Its owner reuses a recent recovery drill because the queue format and recovery mechanism are unchanged, recording that applicability. Deployment stops if duplicate records appear or the agreed error threshold is crossed. The authorized rollout completes, production smoke checks pass, and the dispatcher accepts the specified retry behavior with a dated source. The operating-cost benefit remains pending until the agreed measurement window closes. Implementation, deployment, acceptance, and measured value have different evidence. In the existing engagement, `decisions.md` records the agreed behavior and `delivery.md` holds the release receipt; standalone work returns those facts directly.
|
|
67
|
+
|
|
68
|
+
## Principles
|
|
69
|
+
|
|
70
|
+
- Release the reviewed candidate with applicable evidence and documented authority.
|
|
71
|
+
- Test recovery against the effects that actually persist beyond a code revert.
|
|
72
|
+
- Keep missing evidence visible and distinguish local, staging, and production claims.
|
|
73
|
+
- Expansion follows observed acceptance and operating limits; fixed ceremonies cannot replace them.
|
|
@@ -0,0 +1,12 @@
|
|
|
1
|
+
# Task context and evidence
|
|
2
|
+
|
|
3
|
+
Use this contract for standalone methods and methods routed through `@fde`.
|
|
4
|
+
|
|
5
|
+
- **Standalone work:** use the supplied, permitted facts, notes, code, and artifacts. A client name, `.fde/` directory, or initialized engagement is not a prerequisite. Do not bootstrap records merely to run a method. Ask only for missing information or authority that changes the next action; mark other gaps as unknown.
|
|
6
|
+
- **Artifact names are destinations:** names such as `success.md`, `decisions.md`, and `delivery.md` identify relevant evidence and, when bound, record destinations. If absent, use supplied facts and return the requested draft or result in the current workspace or conversation. Do not invent files or require initialization to complete useful work.
|
|
7
|
+
- **Bound engagement:** honor the current client binding and constraints. Before reading records, run `fde privacy` to verify masking support. Obtain a fresh, identity-matching sanitized `fde resume` packet for this task (or reuse a fresh session-hook packet); retrieve missing evidence with targeted `fde recall <topic>`. Use bounded `fde handoff` for transfer work. Refresh after binding, masking, or record changes. Never substitute raw `.fde/` reads, private blocks, masking dictionaries, or full transcripts. If the CLI is unavailable, use only permitted supplied excerpts and report the context limitation.
|
|
8
|
+
- **Authority:** continue reversible work within authorized scope. Reuse prior authorization when it covers the specific action. Show consequential engagement-record judgments and uncertainties for confirmation before saving unless already explicitly confirmed. New scope, acceptance changes, production actions, exports, and external messages need the applicable authority; a method invocation alone does not supply it. Keep one customer's writes in that customer's record.
|
|
9
|
+
- **Evidence:** distinguish supplied facts, estimates, hypotheses, and unknowns. Cite actual sources; a log date is not attribution. Never invent a source, signer, signature, customer reaction, or acceptance. Keep outcomes **promised → measured → accepted** distinct, and implementation, verification, deployment, and customer acceptance separate. Missing evidence means unproven, not an observed failure.
|
|
10
|
+
- **Data boundary:** use only data permitted by the customer's AI policy; clarify unknown policy before loading their code or data. Never load `<private>` content into a model. Cross-client comparison and exporting reusable material require permission and removal of customer-identifying or confidential content; anonymization alone does not grant permission.
|
|
11
|
+
|
|
12
|
+
Apply the selected method to this context. Follow its linked supporting methods only when needed; do not restart discovery or repeat already answered questions.
|
|
@@ -0,0 +1,31 @@
|
|
|
1
|
+
# verification - Make a claim replayable
|
|
2
|
+
|
|
3
|
+
**Enter when:** reporting completion, evaluating an acceptance check, handing work to a reviewer, or preparing a release.
|
|
4
|
+
|
|
5
|
+
Use [task context](task-context.md). This method returns evidence directly or writes an existing permitted task/engagement record; it never requires `.fde/` initialization.
|
|
6
|
+
|
|
7
|
+
## Method
|
|
8
|
+
|
|
9
|
+
1. Translate each claim into the observation that would support or reject it. Reuse agreed acceptance criteria and required repository checks. Select focused checks for changed behavior before broadening to release requirements.
|
|
10
|
+
2. Identify the actual repository commands, fixtures, runtime, and environment. Read command behavior before executing it, especially when it can write externally. Use authorized environments and avoid leaking secrets through logs or diagnostic commands.
|
|
11
|
+
3. Run the checks and inspect results, including exit status and relevant output. A running job, test discovery, a mocked response, and a successful real request are different evidence. Record asynchronous completion before claiming success.
|
|
12
|
+
4. Bind evidence to the tested revision and working tree. For uncommitted changes record the base revision plus changed paths and an available diff digest or snapshot identifier. For browser/manual checks record the steps, inputs, observed result, and inspected evidence.
|
|
13
|
+
5. After a change, rerun checks whose behavior or assumptions were affected. Reuse prior evidence only when the relevant code, dependencies, data, and environment remain applicable; cite the original run and reason. Never imply reused evidence was rerun.
|
|
14
|
+
6. Label every required check **passed**, **failed**, **blocked**, or **not run**. Include why blocked/not run, impact, and next step. Missing evidence is unproven; it is not an observed failure or a pass.
|
|
15
|
+
|
|
16
|
+
## Receipt
|
|
17
|
+
|
|
18
|
+
Use one compact entry per check or a table with these fields:
|
|
19
|
+
|
|
20
|
+
- Claim / acceptance check and expected result.
|
|
21
|
+
- Exact command and working directory, or manual journey and inputs.
|
|
22
|
+
- Environment, runtime/tool versions when relevant, and fixture/data source.
|
|
23
|
+
- Revision plus working-tree identity; run date/time.
|
|
24
|
+
- Observed result and exit status where available; safe evidence location.
|
|
25
|
+
- Status, limitations, unrun checks, and next step.
|
|
26
|
+
|
|
27
|
+
Keep implementation, verification, deployment, measured outcome, and customer acceptance distinct. A local pass supports the tested local behavior. An acceptance claim needs an attributed source from the agreed decision-maker or agreed acceptance mechanism. Record no raw `<private>` blocks, credentials, or hidden reasoning.
|
|
28
|
+
|
|
29
|
+
## Acceptance
|
|
30
|
+
|
|
31
|
+
A completion statement cites applicable evidence for its claims and explicitly names material gaps. If required checks fail, investigate or report the blocker; never skip them, edit expectations, or relabel the scope without authority to obtain a green result.
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: fde-readout
|
|
3
|
+
description: Prepare a sponsor update separating promised outcomes, measured results and customer acceptance. Use for progress readouts or defending a delivery claim.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# fde-readout
|
|
7
|
+
|
|
8
|
+
<!-- Generated by bin/generate-skills.js; edit the canonical references and catalog. -->
|
|
9
|
+
|
|
10
|
+
## Purpose
|
|
11
|
+
|
|
12
|
+
Prepare a sponsor update separating promised outcomes, measured results and customer acceptance. Use for progress readouts or defending a delivery claim.
|
|
13
|
+
|
|
14
|
+
Read [the task context contract](references/task-context.md), then [the method](references/readout.md). Load further references only when the task needs them. Everything linked is included in this skill; no other skill pack is required.
|
|
15
|
+
|
|
16
|
+
## Principles
|
|
17
|
+
|
|
18
|
+
- Work directly from the supplied permitted context. Standalone work does not require an engagement folder or initialization. Record filenames in the method are optional persistence destinations when no engagement is bound.
|
|
19
|
+
- If called by @fde, reuse its current sanitized packet and scope. Do not restart setup, discovery or questions already answered.
|
|
20
|
+
- The task context contract controls persistence and authority in both modes. Preserve unknowns and distinguish implementation, verification, deployment and acceptance.
|
|
21
|
+
- Use the customer's repository instructions and available tools. Report a missing capability or unrun check honestly; do not claim that installing a skill provisions infrastructure.
|
|
@@ -0,0 +1,108 @@
|
|
|
1
|
+
# board-memo - Brief the board
|
|
2
|
+
|
|
3
|
+
**Enter when:** the sponsor's boss needs a summary, a board update mentions the engagement, the FDE needs to justify continued investment, or a quarterly review is approaching.
|
|
4
|
+
|
|
5
|
+
**Read first:** `delivery.md`, `success.md`, `reality.md`, `risks.md`, `stakeholders.md`, `context.md`. The narrative is built from the engagement record, not from memory.
|
|
6
|
+
|
|
7
|
+
Technical FDEs lose renewals by presenting work instead of outcomes. The exec doesn't want to know what was built - they want to know what it changed. A good exec narrative takes 60 seconds to deliver and survives hostile questions.
|
|
8
|
+
|
|
9
|
+
## Method (you do this work)
|
|
10
|
+
|
|
11
|
+
**1. The Pyramid Principle.** One governing thought, supported by three arguments, each backed by evidence. The exec hears the conclusion first, not the journey:
|
|
12
|
+
|
|
13
|
+
```
|
|
14
|
+
GOVERNING THOUGHT: (one sentence - the conclusion)
|
|
15
|
+
"The payment processing overhaul cut manual reconciliation from
|
|
16
|
+
3 FTEs to 0.5 FTE and eliminated the $2M annual audit risk."
|
|
17
|
+
|
|
18
|
+
SUPPORT 1: What was done (one paragraph)
|
|
19
|
+
→ Evidence from delivery.md
|
|
20
|
+
|
|
21
|
+
SUPPORT 2: What it saved (quantified)
|
|
22
|
+
→ Evidence from business-case.md + delivery.md
|
|
23
|
+
|
|
24
|
+
SUPPORT 3: What's next (the ask)
|
|
25
|
+
→ Evidence from decisions.md + risks.md
|
|
26
|
+
```
|
|
27
|
+
|
|
28
|
+
**2. Four narrative lengths.** The same story, scaled for the context:
|
|
29
|
+
|
|
30
|
+
| Length | When | Format |
|
|
31
|
+
|--------|------|--------|
|
|
32
|
+
| **30 seconds** | Elevator, hallway, Slack thread | The governing thought + one number |
|
|
33
|
+
| **2 minutes** | Stand-up, exec check-in | Governing thought + 3 supports + the ask |
|
|
34
|
+
| **10 minutes** | Quarterly review, steering committee | Full pyramid + hard questions answered + visual |
|
|
35
|
+
| **60 minutes** | Board presentation, transformation review | Full pyramid + demos + deep-dive appendix |
|
|
36
|
+
|
|
37
|
+
Write all four. The FDE will need different lengths at different moments - having them pre-written means they're never caught improvising.
|
|
38
|
+
|
|
39
|
+
**3. The opening frame - SCQA.** Structure the first 30 seconds:
|
|
40
|
+
|
|
41
|
+
| Element | Purpose | Example |
|
|
42
|
+
|---------|---------|---------|
|
|
43
|
+
| **Situation** | Where we are (shared context) | "We started this engagement to fix the payment failures that were costing $200K/month in manual reconciliation." |
|
|
44
|
+
| **Complication** | What changed or what's at stake | "The problem was deeper than expected - the reconciliation failures traced to a data integrity issue in the core ledger." |
|
|
45
|
+
| **Question** | The decision the exec needs to make | "Should we extend the engagement to fix the root cause, or ship the workaround?" |
|
|
46
|
+
| **Answer** | Your recommendation | "Fix the root cause. The workaround adds $40K/year in maintenance and doesn't eliminate the audit risk." |
|
|
47
|
+
|
|
48
|
+
**4. Value in their units.** Translate every technical achievement:
|
|
49
|
+
|
|
50
|
+
| What you did (internal) | What it means (their units) |
|
|
51
|
+
|------------------------|---------------------------|
|
|
52
|
+
| Reduced p95 latency from 3s to 200ms | Customers complete checkout 15x faster |
|
|
53
|
+
| Added test coverage from 12% to 78% | Change failure rate dropped from 40% to 5% |
|
|
54
|
+
| Migrated from monolith to three services | Team can deploy independently - shipping frequency from monthly to weekly |
|
|
55
|
+
| Built ML fraud detection | $1.2M/year in fraud losses reduced to <$200K projected |
|
|
56
|
+
|
|
57
|
+
Never: "we refactored the authentication module." Always: what the refactoring *did* for them.
|
|
58
|
+
|
|
59
|
+
**5. Pre-wire the hostile questions.** Before any exec presentation, write the five toughest questions and one-line answers:
|
|
60
|
+
|
|
61
|
+
```markdown
|
|
62
|
+
## Hard questions - <presentation date>
|
|
63
|
+
1. "Why did this take longer than estimated?"
|
|
64
|
+
→ The original brief assumed API-only work; discovery revealed a database integrity issue. We surfaced it in week 2 instead of shipping a patch that would have required rework.
|
|
65
|
+
|
|
66
|
+
2. "How do we know it won't break again?"
|
|
67
|
+
→ Three guards: automated reconciliation check (runs daily), alerting on drift >0.1%, and the characterisation test suite covering the 12 failure modes we found.
|
|
68
|
+
|
|
69
|
+
3. "What happens when the FDE leaves?"
|
|
70
|
+
→ Handoff document written for the 2am scenario. The team ran the runbook independently last Thursday - no callbacks.
|
|
71
|
+
|
|
72
|
+
4. "Why should we fund phase 2?"
|
|
73
|
+
→ Phase 1 addressed the bleeding. Phase 2 eliminates the root cause. Without it: $40K/year maintenance on the workaround + the audit risk remains.
|
|
74
|
+
|
|
75
|
+
5. "Can the internal team do phase 2 without you?"
|
|
76
|
+
→ They can, with 2x the timeline. The value of an FDE in phase 2 is speed - the patterns are established and the trust with the ledger team is built.
|
|
77
|
+
```
|
|
78
|
+
|
|
79
|
+
**6. The one number.** Every exec narrative needs a single memorable quantity:
|
|
80
|
+
|
|
81
|
+
- "31 spreadsheet rows to zero"
|
|
82
|
+
- "p95 held at 180ms"
|
|
83
|
+
- "$200K monthly risk retired"
|
|
84
|
+
- "Time-to-deploy from 4 hours to 12 minutes"
|
|
85
|
+
|
|
86
|
+
The number should appear in the first 30 seconds and be the thing they repeat to *their* boss.
|
|
87
|
+
|
|
88
|
+
## Artifact
|
|
89
|
+
|
|
90
|
+
**`delivery.md`** - append under `## Exec narrative - <date>`:
|
|
91
|
+
- The four narrative lengths (30s, 2min, 10min, 60min)
|
|
92
|
+
- The SCQA frame
|
|
93
|
+
- The hard-question sheet
|
|
94
|
+
- The one number
|
|
95
|
+
|
|
96
|
+
**`context.md`** - note: exec narrative prepared, presentation date, what must be updated before delivery.
|
|
97
|
+
|
|
98
|
+
## Checkpoint
|
|
99
|
+
|
|
100
|
+
Dry-run the 2-minute version with the FDE. Confirm: the one number lands in the first 30 seconds, the SCQA frame answers "why now," and the hardest question has a prepared answer. If the FDE can't deliver the 30-second version from memory, simplify.
|
|
101
|
+
|
|
102
|
+
## Principles
|
|
103
|
+
|
|
104
|
+
- Conclusion first, evidence second. The exec decides in the first 30 seconds.
|
|
105
|
+
- Value in their units. Never present work; present outcomes.
|
|
106
|
+
- One number per narrative. The room remembers one thing - make it the right thing.
|
|
107
|
+
- Pre-wire every hostile question. Surprise in an exec meeting is a trust withdrawal.
|
|
108
|
+
- Write all four lengths. The FDE will need them at different moments.
|
|
@@ -0,0 +1,90 @@
|
|
|
1
|
+
# business-case - Build the business case
|
|
2
|
+
|
|
3
|
+
**Context:** apply [task context and evidence](task-context.md) before using the named records below.
|
|
4
|
+
|
|
5
|
+
**Enter when:** the sponsor needs justification for the next phase, the FDE needs to defend budget or timeline, a feature decision needs cost/benefit evidence, or poc produced a direction that needs funding.
|
|
6
|
+
|
|
7
|
+
**Read first:** `reality.md`, `success.md`, `delivery.md`, `context.md`. Load `business-case.md` from poc if it exists - extend it, don't restart.
|
|
8
|
+
|
|
9
|
+
Technical FDEs lose engagements by shipping good code without business justification. The sponsor's boss doesn't ask "is the code clean?" - they ask "what did we get for the money?" A business case translates technical work into the language that keeps the engagement alive.
|
|
10
|
+
|
|
11
|
+
## Method (you do this work)
|
|
12
|
+
|
|
13
|
+
**1. Name the cost of doing nothing.** This is the anchor. Every business case starts not with what you'll build, but with what it costs them to leave the problem unsolved:
|
|
14
|
+
|
|
15
|
+
| Cost type | How to find it | Example |
|
|
16
|
+
|-----------|---------------|---------|
|
|
17
|
+
| **Labor capacity / direct spend** | Ask: "What does this problem cost per month in money?" | Manual reconciliation hours × loaded rate = capacity value; separately identify reducible spend |
|
|
18
|
+
| **Opportunity cost** | Ask: "What can't you do because of this problem?" | Can't onboard enterprise clients because the API can't handle their volume |
|
|
19
|
+
| **Risk cost** | Ask: "What happens if this breaks at the worst time?" | A payment processing outage during Black Friday = $X/hour in lost sales |
|
|
20
|
+
| **Velocity cost** | Measure: deployment frequency, lead time, change failure rate | Team ships once/month instead of once/week; each delay = N features not reaching customers |
|
|
21
|
+
|
|
22
|
+
**2. Build the driver model.** Not a spreadsheet - a logic chain the sponsor can trace:
|
|
23
|
+
|
|
24
|
+
```
|
|
25
|
+
Investment: <hours × rate, or fixed cost>
|
|
26
|
+
→ Delivers: <specific outcome from success.md>
|
|
27
|
+
→ Benefit: <capacity released, avoidable cash spend, revenue, or risk reduction>
|
|
28
|
+
→ Net cash: realizable incremental cash benefit - full costs over <time horizon>
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
Keep drivers, units, sources, and ranges explicit. For example, 3 people × 8h/week × $75/h × 52 weeks = $93.6K/year of labor capacity value. It is cash savings only if spend actually falls (for example, paid overtime or a contractor cost ends). Name who can realize the benefit and how. Include build, ongoing operation, adoption, and transition costs; avoid double-counting capacity and revenue enabled by the same hours. Do not calculate cash payback from capacity value alone.
|
|
32
|
+
|
|
33
|
+
**3. Sensitivity check - name the two drivers that swing the result:**
|
|
34
|
+
|
|
35
|
+
Every business case has 1-2 variables where a small change flips the outcome. Name them explicitly:
|
|
36
|
+
|
|
37
|
+
> "The capacity case assumes the team reclaims 6 hours/week per person. At 3 hours, that benefit halves. Cash payback remains unproven until finance identifies avoidable spend. Validate time-spent before and after the pilot with representative team members."
|
|
38
|
+
|
|
39
|
+
The sponsor who sees you've identified where the case could break trusts the case more, not less.
|
|
40
|
+
|
|
41
|
+
**4. Frame for the audience.** Different stakeholders need different lenses on the same case:
|
|
42
|
+
|
|
43
|
+
| Audience | Lead with | Avoid |
|
|
44
|
+
|----------|----------|-------|
|
|
45
|
+
| **CFO / finance** | ROI, payback period, cash flow impact | Technical architecture, feature lists |
|
|
46
|
+
| **CTO / engineering** | Technical debt retired, velocity improved, risk reduced | Revenue projections they can't verify |
|
|
47
|
+
| **CEO / founder** | Strategic enablement, competitive edge, customer impact | Detailed calculations (give the summary, offer the detail) |
|
|
48
|
+
| **Product** | User impact, adoption metrics, feature velocity | Cost structures that aren't their domain |
|
|
49
|
+
|
|
50
|
+
**5. The one-page format.** The business case fits one page or it isn't understood:
|
|
51
|
+
|
|
52
|
+
```markdown
|
|
53
|
+
## Business case: <initiative name>
|
|
54
|
+
|
|
55
|
+
**The problem costs:** <one line, quantified>
|
|
56
|
+
**The investment:** <hours and cost>
|
|
57
|
+
**The return:** <quantified, with time horizon>
|
|
58
|
+
**Payback:** <months from realizable cash benefits, or not established>
|
|
59
|
+
**Sensitivity:** <the 1-2 drivers that swing it, with thresholds>
|
|
60
|
+
**Risks:** <what must be true for this to hold>
|
|
61
|
+
**Recommendation:** <proceed / proceed-with-conditions / defer>
|
|
62
|
+
```
|
|
63
|
+
|
|
64
|
+
## Artifact
|
|
65
|
+
|
|
66
|
+
**`business-case.md`** - the one-page case. Lives alongside `success.md` and `reality.md` as a first-class engagement artifact. Referenced by plan, status, and close.
|
|
67
|
+
|
|
68
|
+
**`decisions.md`** - log the sponsor's response: approved, modified, deferred. With the date.
|
|
69
|
+
|
|
70
|
+
## Checkpoint
|
|
71
|
+
|
|
72
|
+
Walk the FDE through: the cost of doing nothing (anchor), the investment, the return, and the one sensitivity that matters most. If the FDE says "the sponsor won't buy the ROI number," inspect the disputed inputs and sources, test plausible ranges, and identify what measurement would resolve the disagreement. Never reverse-engineer assumptions to hit a desired number.
|
|
73
|
+
|
|
74
|
+
## Worked example
|
|
75
|
+
|
|
76
|
+
Acme phase 2 needs funding. The case starts with the cost of doing nothing, not the cost of building.
|
|
77
|
+
|
|
78
|
+
Anchor: two silent failures since March, each one day of finance reconciliation by hand plus a late close (`reality.md`, Marco's sheet). That is the number the sponsor already believes because her own team reported it.
|
|
79
|
+
|
|
80
|
+
Driver model the sponsor can trace: incidents/quarter × hours of manual reconciliation × loaded cost, plus the tail risk of a late regulatory close - stated separately, because mixing a certain small number with an uncertain large one is how a case loses credibility.
|
|
81
|
+
|
|
82
|
+
Sensitivity names the two drivers that swing it: incident frequency (2/quarter → 1/quarter and the case halves) and whether the manual re-run continues in parallel (if Marco keeps re-running every morning, the saving is theoretical). The second one is the honest weakness, so it is in the case rather than waiting to be found in the room - with the condition that makes it hold: the morning re-run stops after two clean cycles, agreed with Marco.
|
|
83
|
+
|
|
84
|
+
## Principles
|
|
85
|
+
|
|
86
|
+
- The cost of doing nothing is always the opening move. Anchor before proposing.
|
|
87
|
+
- Driver models with visible arithmetic beat magic spreadsheets.
|
|
88
|
+
- Name the sensitivity. The case that admits its weakness earns more trust.
|
|
89
|
+
- One page. If it doesn't fit, you don't understand it yet.
|
|
90
|
+
- A business case the FDE can't explain in 60 seconds won't survive the sponsor's boss.
|
|
@@ -0,0 +1,69 @@
|
|
|
1
|
+
# readout - Report the outcome
|
|
2
|
+
|
|
3
|
+
**Enter when:** the weekly update is due, an exec asks "where are we," or the FDE says "I need to send Dana something." This artifact decides renewals; engineers underinvest in it.
|
|
4
|
+
|
|
5
|
+
**Read first:** `success.md` (the yardstick), `delivery.md` (value ledger), `decisions.md` (plan + kill list), `assumptions.md` (OPEN criticals), `risks.md`, `context.md`. Gather the week's facts: `fde receipts` for agreements, `git log --since='7 days ago' --oneline` for shipped work.
|
|
6
|
+
|
|
7
|
+
## Method (you do this work)
|
|
8
|
+
|
|
9
|
+
**First:** run `fde status`. It prints the value ledger before trust - promised → measured → accepted by, or `claimed, not yet accepted`. Those lines locate the Situation; check their cited records before making the claim. CLI output summarizes recorded text, not independently verified acceptance. Do not invent a number the CLI did not print. If the CLI is unavailable, use the redacted source records and say so.
|
|
10
|
+
|
|
11
|
+
**Qualify the evidence before drafting.** For each result, identify baseline source, measurement environment, observation window/sample, and the scope of acceptance. Report an informal baseline as reported and a staging sample as staging; neither establishes realized savings. “Looks good” without what was accepted is not outcome acceptance. Attribute an engineer's note as such; do not turn it into a direct customer receipt. If evidence conflicts, include the conflict and the next verification action rather than choosing the flattering version.
|
|
12
|
+
|
|
13
|
+
**Always draft in SCQA.** One page maximum. No other shape.
|
|
14
|
+
|
|
15
|
+
| Block | What to write | Source |
|
|
16
|
+
|-------|---------------|--------|
|
|
17
|
+
| **S - Situation** | Where we are against `success.md`, in their words - including whether the floor still uses the old path | success.md, delivery value ledger, reality.md workaround |
|
|
18
|
+
| **C - Complication** | What changed, what is at risk, or what we learned (bad news first) | risks.md, assumptions DISPROVED/OPEN, stakeholders signal |
|
|
19
|
+
| **Q - Question / Ask** | The one decision or help you need from them | decisions.md, access/sign-off needs |
|
|
20
|
+
| **A - Answer** | What you recommend / what happens next week (≤3 bullets) | plan Now lane, delivery promised→measured |
|
|
21
|
+
|
|
22
|
+
Then add, still on the same page:
|
|
23
|
+
1. **Value this week** - from the value ledger: promised → measured (or "pending") → **accepted by whom**, with evidence citation. A measured number nobody on the customer side has agreed to is written as `claimed`, and the Ask never rests on it - if the whole case for the next phase is a claimed number, the real ask this week is "who signs off that this is real?".
|
|
24
|
+
2. **Are they using it?** - Situation must say whether the workaround is still open: spreadsheet still running, shadow paste still happening, named operator completed Tuesday's job on the new path without you at the keyboard. A measured metric with the old path still live is `claimed`. That week's Ask is not "fund phase 2." It is "who on their side stops the old way, by when."
|
|
25
|
+
3. **Kill / defer reminder** - one line from the plan kill list so scope fights stay visible.
|
|
26
|
+
4. **Hostile Q prep** - three questions a skeptical sponsor will ask, with one-line answers from memory.
|
|
27
|
+
|
|
28
|
+
Exec voice: no jargon, explicit uncertainty where evidence is incomplete, every claim traceable (`(shipped Tue, delivery.md)`). Draft in the **FDE's voice, for the FDE to send** - never send anything yourself.
|
|
29
|
+
|
|
30
|
+
For board / renewal / sponsor's boss (longer pyramid): use `board-memo.md`. Do not invent a second weekly format.
|
|
31
|
+
|
|
32
|
+
## Artifact
|
|
33
|
+
|
|
34
|
+
Append the draft to `delivery.md` under `## Status - <date>` using the SCQA headings. Note in `context.md`: status drafted, awaiting FDE review/send.
|
|
35
|
+
|
|
36
|
+
```markdown
|
|
37
|
+
## Status - YYYY-MM-DD
|
|
38
|
+
**S:** ...
|
|
39
|
+
**C:** ...
|
|
40
|
+
**Q:** ...
|
|
41
|
+
**A:** ...
|
|
42
|
+
**Value ledger:** promised … / measured … / accepted by … (evidence) - or `claimed, unaccepted`
|
|
43
|
+
**Kill list reminder:** …
|
|
44
|
+
**Hostile Qs:** 1) … 2) … 3) …
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
## Checkpoint
|
|
48
|
+
|
|
49
|
+
Before presenting a consequential draft, check whether its intended reader can identify what changed, what remains unproven, and the decision being requested without extra explanation. Use the draft alone for this check; repair the unclear passage rather than adding another summary. A simulated reader can flag confusion but cannot confirm stakeholder understanding or acceptance. Keep the existing FDE review/send boundary.
|
|
50
|
+
|
|
51
|
+
Walk the FDE through the Complication and the Ask - confirm the framing matches what the sponsor can hear right now (check `stakeholders.md` signal first: a red-signal sponsor gets a different opening than a green one).
|
|
52
|
+
|
|
53
|
+
## Worked example
|
|
54
|
+
|
|
55
|
+
Acme, week 3, Priya's Friday update.
|
|
56
|
+
|
|
57
|
+
**S:** failure routing is live; detection is 12 min against the 4h baseline in `success.md`. **C** leads with the bad news, not the win: the second incident was acked 40 minutes late because the rota has one name on it, and that name was on leave. **Q:** one ask - a second name on the rota by Wednesday. **A:** three bullets, top of the Now lane.
|
|
58
|
+
|
|
59
|
+
Value ledger line: `promised 4h → 15min / measured 12min over 2 incidents / accepted by - (Marco confirmed operationally, finance not yet)` → written as `claimed, unaccepted`, which is what makes the Ask honest rather than a victory lap.
|
|
60
|
+
|
|
61
|
+
Hostile Q prep, from memory not imagination: "why did we pay for alerting we already had?" → the receipt from `decisions.md` and the disabled-alerting finding in `reality.md`. Kill list reminder: the service rewrite is still deferred, accepted by Priya on Jun 12.
|
|
62
|
+
|
|
63
|
+
## Principles
|
|
64
|
+
|
|
65
|
+
- SCQA every time. Situation → Complication → Ask → Answer.
|
|
66
|
+
- No surprises: anything the sponsor would be angry to learn later goes in Complication.
|
|
67
|
+
- Value from the ledger, not from ticket theater.
|
|
68
|
+
- An update without an ask is a missed move.
|
|
69
|
+
- You draft; the FDE sends. Their voice, their relationship.
|
|
@@ -0,0 +1,12 @@
|
|
|
1
|
+
# Task context and evidence
|
|
2
|
+
|
|
3
|
+
Use this contract for standalone methods and methods routed through `@fde`.
|
|
4
|
+
|
|
5
|
+
- **Standalone work:** use the supplied, permitted facts, notes, code, and artifacts. A client name, `.fde/` directory, or initialized engagement is not a prerequisite. Do not bootstrap records merely to run a method. Ask only for missing information or authority that changes the next action; mark other gaps as unknown.
|
|
6
|
+
- **Artifact names are destinations:** names such as `success.md`, `decisions.md`, and `delivery.md` identify relevant evidence and, when bound, record destinations. If absent, use supplied facts and return the requested draft or result in the current workspace or conversation. Do not invent files or require initialization to complete useful work.
|
|
7
|
+
- **Bound engagement:** honor the current client binding and constraints. Before reading records, run `fde privacy` to verify masking support. Obtain a fresh, identity-matching sanitized `fde resume` packet for this task (or reuse a fresh session-hook packet); retrieve missing evidence with targeted `fde recall <topic>`. Use bounded `fde handoff` for transfer work. Refresh after binding, masking, or record changes. Never substitute raw `.fde/` reads, private blocks, masking dictionaries, or full transcripts. If the CLI is unavailable, use only permitted supplied excerpts and report the context limitation.
|
|
8
|
+
- **Authority:** continue reversible work within authorized scope. Reuse prior authorization when it covers the specific action. Show consequential engagement-record judgments and uncertainties for confirmation before saving unless already explicitly confirmed. New scope, acceptance changes, production actions, exports, and external messages need the applicable authority; a method invocation alone does not supply it. Keep one customer's writes in that customer's record.
|
|
9
|
+
- **Evidence:** distinguish supplied facts, estimates, hypotheses, and unknowns. Cite actual sources; a log date is not attribution. Never invent a source, signer, signature, customer reaction, or acceptance. Keep outcomes **promised → measured → accepted** distinct, and implementation, verification, deployment, and customer acceptance separate. Missing evidence means unproven, not an observed failure.
|
|
10
|
+
- **Data boundary:** use only data permitted by the customer's AI policy; clarify unknown policy before loading their code or data. Never load `<private>` content into a model. Cross-client comparison and exporting reusable material require permission and removal of customer-identifying or confidential content; anonymization alone does not grant permission.
|
|
11
|
+
|
|
12
|
+
Apply the selected method to this context. Follow its linked supporting methods only when needed; do not restart discovery or repeat already answered questions.
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: fde-review
|
|
3
|
+
description: Review a proposed customer code change against its intended outcome and operational risks. Use for a diff or PR review; report evidence and actionable findings.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# fde-review
|
|
7
|
+
|
|
8
|
+
<!-- Generated by bin/generate-skills.js; edit the canonical references and catalog. -->
|
|
9
|
+
|
|
10
|
+
## Purpose
|
|
11
|
+
|
|
12
|
+
Review a proposed customer code change against its intended outcome and operational risks. Use for a diff or PR review; report evidence and actionable findings.
|
|
13
|
+
|
|
14
|
+
Read [the task context contract](references/task-context.md), then [the method](references/review.md). Load further references only when the task needs them. Everything linked is included in this skill; no other skill pack is required.
|
|
15
|
+
|
|
16
|
+
## Principles
|
|
17
|
+
|
|
18
|
+
- Work directly from the supplied permitted context. Standalone work does not require an engagement folder or initialization. Record filenames in the method are optional persistence destinations when no engagement is bound.
|
|
19
|
+
- If called by @fde, reuse its current sanitized packet and scope. Do not restart setup, discovery or questions already answered.
|
|
20
|
+
- The task context contract controls persistence and authority in both modes. Preserve unknowns and distinguish implementation, verification, deployment and acceptance.
|
|
21
|
+
- Use the customer's repository instructions and available tools. Report a missing capability or unrun check honestly; do not claim that installing a skill provisions infrastructure.
|
|
@@ -0,0 +1,20 @@
|
|
|
1
|
+
# build - Implement a verifiable increment
|
|
2
|
+
|
|
3
|
+
**Enter when:** an agreed behavior needs implementation in an existing or new repository. For a broken behavior, start with [debug](debug.md); for a system boundary, use [integrate](integrate.md).
|
|
4
|
+
|
|
5
|
+
Use the permitted context and authority in [task context](task-context.md). This method works without `.fde/`; an existing engagement record can supply the same contract. Do not initialize memory just to write code.
|
|
6
|
+
|
|
7
|
+
## Method
|
|
8
|
+
|
|
9
|
+
1. Identify the repository, its instructions, working tree, relevant callers, and test commands. Inspect examples before creating abstractions. Preserve unrelated edits and state which dependencies or interfaces the change touches.
|
|
10
|
+
2. State the observable outcome, constraints, and acceptance checks. Reuse agreed criteria for routine fixes. If a consequential product choice is unresolved, surface that choice while continuing independent investigation; do not invent acceptance.
|
|
11
|
+
3. Choose the smallest coherent path that demonstrates the outcome through the real entry point. Include the necessary storage, error handling, and interface behavior in that slice. Name the failure that stops expansion and the recovery path for stateful changes.
|
|
12
|
+
4. Implement using the repository's tools and conventions. Search for existing services, fixtures, and validation before adding alternatives. Keep cleanup limited to what makes the changed path understandable; do not expand scope to repair unrelated code.
|
|
13
|
+
5. Run focused checks, then required repository checks. Exercise the actual affected journey with [QA](qa.md) when appropriate. For uncertain model behavior, use [eval-pack](eval-pack.md). Record results with [verification](verification.md), including checks that could not run.
|
|
14
|
+
6. Inspect the final diff against the agreed outcome. For substantial or risky work, seek [review](review.md) using an actual separate reviewer when available; identify a self-check honestly. Reverify affected behavior after fixes.
|
|
15
|
+
|
|
16
|
+
## Deliverable and acceptance
|
|
17
|
+
|
|
18
|
+
Return the implemented behavior, relevant paths, evidence, remaining limitations, and any decision needed. Done means the agreed checks have applicable evidence and the change is reviewable; passing tests does not imply deployment or customer acceptance. Committing, opening a PR, merging, and publishing happen only when the requested workflow authorizes those actions.
|
|
19
|
+
|
|
20
|
+
When coordinated through `@fde`, record implementation and verification in the existing decisions/delivery records under their write rules. Standalone work can return the same receipt directly or use the repository's task record.
|
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
# debug - Find and repair the cause
|
|
2
|
+
|
|
3
|
+
**Enter when:** a reproducible failure, regression, incident symptom, or misleading output needs investigation.
|
|
4
|
+
|
|
5
|
+
Use [task context](task-context.md). Work from supplied permitted evidence without requiring `.fde/`. During an active incident, follow the authorized containment procedure before diagnosis; investigation authority alone does not authorize production writes.
|
|
6
|
+
|
|
7
|
+
## Method
|
|
8
|
+
|
|
9
|
+
1. Capture expected and observed behavior, exact input or trigger, affected revision/environment, and the last known working state. Preserve useful errors and timestamps without copying secrets or raw private data. Mark reports you have not reproduced as reports.
|
|
10
|
+
2. Inspect the failing path, callers, recent relevant changes, and existing tests. Reproduce in a permitted environment with the smallest representative case. If reproduction is unavailable, identify what observation would distinguish causes and gather safe evidence; do not claim a hypothesis is proven.
|
|
11
|
+
3. Keep a short hypothesis list. For each, name the predicted observation and a discriminating check. Change one relevant variable at a time. Trace values and control flow across the actual boundary instead of repeatedly changing code until the symptom disappears.
|
|
12
|
+
4. Fix the cause at the appropriate layer. Check whether the proposed fix changes behavior for other callers, stale data, retries, concurrency, or permissions. Preserve evidence of the original failure and avoid unrelated cleanup.
|
|
13
|
+
5. Add a regression check when it can meaningfully reproduce the bug; show that it fails before the fix and passes after when practical. If the check cannot run against the before-state, say so. Run affected adjacent and required checks using [verification](verification.md).
|
|
14
|
+
6. Review the final diff and exercise the original journey. For substantial or risky fixes use [review](review.md). After two unsuccessful repair cycles, reassess the hypothesis and evidence instead of repeating the same attempt; continue useful investigation and isolate the missing decision or access.
|
|
15
|
+
|
|
16
|
+
## Deliverable and acceptance
|
|
17
|
+
|
|
18
|
+
Report the cause with its evidence, the fix, the original reproducer's result, adjacent checks, and unresolved uncertainty. A disappearing symptom with no discriminating evidence is a mitigation, not a demonstrated root cause. In engagement mode record the incident/fix receipt in the appropriate existing record; standalone work may return it directly. Release or rollback requires the existing operational authority and [ship](ship.md) or recovery procedure.
|
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
# eval-pack - Evaluate the model's allowed behavior
|
|
2
|
+
|
|
3
|
+
**Enter when:** AI, LLM, RAG, or agent behavior needs evidence before an experiment, release, or material expansion. Non-AI work skips this method.
|
|
4
|
+
|
|
5
|
+
Use [task context](task-context.md). Supplied permitted context and an evaluation report are sufficient without `.fde/`. In a coordinated engagement, use the existing trust/terrain context and keep the report in `evals.md`; read only privacy-safe views.
|
|
6
|
+
|
|
7
|
+
## Method
|
|
8
|
+
|
|
9
|
+
1. **Define the evaluated surface.** Name the model judgment, inputs, outputs, downstream actions, environment, and relevant failure impact. Separate model quality from deterministic tool authorization and application checks. Document the actual allowed action boundary and its source; missing authority remains unknown.
|
|
10
|
+
2. **Choose cases by risk and coverage.** Use permitted historical examples, expert-labeled cases, or clearly marked synthetic fixtures. Cover relevant segments, boundary conditions, known failure modes, and critical harms. Record input, expected outcome/rubric, provenance, and critical-failure rule per case. Keep evaluation cases separate from tuning where possible; no fixed case count proves safety.
|
|
11
|
+
3. **Agree the pass rule before the run.** Define quality thresholds, critical failures, coverage expectations, and acceptable uncertainty for this use. Use deterministic checks where possible and inspect subjective labels or judge reliability. Propose missing criteria for agreement; do not manufacture acceptance from the observed scores.
|
|
12
|
+
4. **Run the actual evaluated path.** Record model/provider version, prompts/configuration, retrieval corpus or tools, application revision, environment, fixtures, and run date. Repeat where variability matters. Report totals, per-segment results, critical failures, and limitations using [verification](verification.md). A model-only run does not prove the agent's tool boundary works.
|
|
13
|
+
5. **Verify action authority and controls.** Human approval is required where the user's policy or task requires it. Already agreed bounded automation may run within its documented actions, identities, environments, and limits; do not require fresh approval for every authorized action. Check enforcement outside the model, least privilege, input/output validation, cost/rate limits, stop conditions, observability, and recovery as applicable. Unknown or exceeded authority blocks those actions. Evaluation success never grants new authority.
|
|
14
|
+
6. **Make a scoped verdict.** Report **SHIP** only when agreed criteria pass, critical failures are zero, applicable authority/control checks pass, and material coverage gaps are resolved or the release is explicitly narrowed by the responsible decision-maker. Otherwise report **NO-SHIP** with the smallest corrective step: fix, gather evidence, descope, or reconsider the judgment surface. A SHIP verdict is technical evidence for the stated scope, not permission to deploy.
|
|
15
|
+
|
|
16
|
+
## Deliverable and acceptance
|
|
17
|
+
|
|
18
|
+
Return the suite/source, thresholds, counts and segments, top failure modes, control evidence, human-review gate or bounded automation authority, limitations, and dated verdict. Record unknown values honestly. Reevaluate after changes that affect model behavior, retrieval, tool permissions, or data conditions; cite why unchanged evidence remains applicable rather than implying a rerun.
|
|
19
|
+
|
|
20
|
+
When coordinated, append a concise eval receipt to delivery records. For release use [ship](ship.md). For ongoing use define the drift signals, sample policy permitted by data handling rules, owner, and conditions that suspend or narrow automation. Do not store secrets, raw `<private>` data, or hidden chain-of-thought in reports.
|
|
21
|
+
|
|
22
|
+
## Principles
|
|
23
|
+
|
|
24
|
+
- Thresholds and authority come from the agreed contract, never from a convenient observed result.
|
|
25
|
+
- Critical failures block the evaluated release scope; disclose coverage and uncertainty.
|
|
26
|
+
- Bound automation with enforceable controls, and require human review where the policy requires it.
|
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
# integrate - Prove the system boundary
|
|
2
|
+
|
|
3
|
+
**Enter when:** connecting an API, data source, SDK, event stream, tool, or service, or changing its contract.
|
|
4
|
+
|
|
5
|
+
Start from [task context](task-context.md). Permitted supplied context is enough; `.fde/` is optional. Use the customer's existing clients, authentication, fixtures, and diagnostic tools. Do not create another integration platform to make one connection.
|
|
6
|
+
|
|
7
|
+
## Method
|
|
8
|
+
|
|
9
|
+
1. Map producer, consumer, owner, direction, and side effects. Inspect the actual installed version and local implementation; verify uncertain behavior against current official documentation. Identify the relevant schema, authentication scopes, network boundary, and permitted test environment.
|
|
10
|
+
2. Write the acceptance example: an input at the real boundary and the observable downstream result. Include a rejection or failure example. Separate configuration validity, successful authentication, transport connectivity, contract compatibility, and end-to-end behavior; none proves the next.
|
|
11
|
+
3. Inspect credentials by presence and required scope without printing values. Use existing secret storage. Check data classification and retention before moving data; never pass raw `<private>` blocks into a model. Prefer sanitized or synthetic cases approved for the target environment.
|
|
12
|
+
4. Implement the narrow adapter using native repository patterns. Validate external inputs and model outputs, bound timeouts and retries, preserve error context without leaking payloads, and handle cancellation. For writes, establish idempotency or duplicate detection before retries; for events, check ordering, replay, and poison messages as applicable.
|
|
13
|
+
5. Exercise a permitted success case and relevant failures: denied access, malformed data, rate limit, timeout, duplicate delivery, or partial completion. Trace correlation IDs or safe evidence across both sides. A mock proves client behavior only; if live access is unavailable, report that gap instead of claiming an integration works.
|
|
14
|
+
6. Check cleanup and recovery for test side effects. Use [verification](verification.md) for receipts and [review](review.md) for security or data-contract changes. Route deployment through [ship](ship.md) only when authorized.
|
|
15
|
+
|
|
16
|
+
## Deliverable and acceptance
|
|
17
|
+
|
|
18
|
+
Return the boundary contract, changed paths, environment, evidence at each tested layer, and remaining dependencies with owners when known. Done requires the agreed end-to-end result or an explicit narrower agreed scope. Do not silently replace live acceptance with a stub. In engagement mode, update the terrain/delivery record with confirmed facts; otherwise return the receipt directly.
|