fdeops 5.1.14 → 5.1.15

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/AGENTS.md CHANGED
@@ -1,6 +1,6 @@
1
1
  # AGENTS.md - working in the fdeops repository
2
2
 
3
- This repository **is** FDEOps: skills for Forward Deployed Engineers. The `fde` coordinator selects the relevant instructions for customer work. All 35 task skills also work individually. The local CLI manages customer records in `.fde/` files; users confirm consequential judgments.
3
+ This repository **is** FDEOps: skills that help Forward Deployed Engineers work across strategy, architecture and engineering through AI coding agents. All 35 task skills work individually; `fde` coordinates them for ongoing customer work. Local `.fde/` records support continuity between sessions; users confirm consequential judgments.
4
4
 
5
5
  ## If you are helping use fdeops in an engagement
6
6
 
package/README.md CHANGED
@@ -1,25 +1,19 @@
1
1
  # FDEOps
2
2
 
3
- **Forward deployed engineering skills for your AI coding agent.**
3
+ **Skills for forward deployed engineers, used through your AI coding agent.**
4
4
 
5
5
  <a name="why-use-it"></a>
6
6
 
7
- The codebase does not tell your agent what the customer agreed to last week, why an approach was rejected or who can approve the next release.
7
+ FDEOps helps you turn a customer problem into a working system: clarify the goal, choose an architecture, build and integrate, then verify the result and prepare for rollout.
8
8
 
9
- FDEOps brings that context into the work, from the first meeting to a system the customer can run. Use a skill for one task, or let `fde` coordinate the project and keep its record.
9
+ Use a task skill on its own, or let `fde` coordinate work across strategy, architecture and engineering. Customer memory keeps decisions, evidence and next steps available between sessions.
10
10
 
11
- **35 task skills + one coordinator, `fde`**. Use your existing tools and processes. You and the customer keep control of the decisions.
11
+ I built FDEOps around design thinking, first-principles thinking and systems thinking: understand the people doing the work, question assumptions, and examine how a change affects the whole system. I turned that approach into skills that help your AI coding agent investigate, build and verify, with customer memory to carry the work forward.
12
12
 
13
13
  [Get started](#quick-start) · [What it helps with](#three-things-it-helps-with) · [Choose a skill](#task-skills) · [Data boundaries](#your-records-your-control) · [Docs](docs/README.md)
14
14
 
15
- ![FDEOps terminal: insurance KYC workshop, build, healthcare client switch, tests and handover](media/chat-demo.gif)
16
-
17
- *Fictional customers. Illustrative enterprise conversation with real local review-logic tests. Synthetic data; no live model or customer deployment. [Read the conversation](media/chat-demo.md) · [View a still](media/chat-demo.png).*
18
-
19
15
  ## Quick start
20
16
 
21
- **Use your customer’s approved AI tools and data.** FDEOps runs locally; your AI agent’s settings determine what reaches its provider. Start with synthetic data until customer access is approved. [Safe setup](SECURITY.md#before-customer-work).
22
-
23
17
  ### Let `fde` coordinate a customer project
24
18
 
25
19
  Install in the terminal where your AI coding agent runs, then select your agent:
@@ -36,23 +30,22 @@ checks internal documents, then assigns each request to another team.
36
30
  Help me prepare for the first meeting. Here is the brief: ...
37
31
  ```
38
32
 
39
- The coordinator selects the right skill as the work changes. For an ongoing project, it keeps decisions, evidence and next actions in a local customer record. You bring the context and make the decisions.
33
+ The coordinator selects the relevant method as the work changes. Customer context guides the plan, code changes and verification; the record carries it between sessions. You do not need to learn CLI commands.
40
34
 
41
35
  ### Use one skill for one task
42
36
 
43
37
  ```bash
44
- npx skills add suboss87/fdeops --skill debrief
38
+ npx skills add suboss87/fdeops --skill build
45
39
  ```
46
40
 
47
41
  Then ask your agent:
48
42
 
49
43
  ```text
50
- Use FDEOps debrief to review these meeting notes.
51
- Separate decisions, requests and open questions. Return a draft only.
52
- [Paste notes you are permitted to share.]
44
+ Use FDEOps build to add a manual-review fallback to this routing service.
45
+ Here are the agreed behavior, repository and checks: ...
53
46
  ```
54
47
 
55
- Each task skill includes the instructions it needs. Use `debrief` on supplied notes without creating a customer record or installing the coordinator.
48
+ Each task skill includes the instructions it needs. Use `build` with supplied project context without creating a customer record or installing the coordinator.
56
49
 
57
50
  <details>
58
51
  <summary>Installation requirements and alternatives</summary>
@@ -65,69 +58,35 @@ These installation commands use Node.js and Git; the optional record CLI require
65
58
 
66
59
  ## Three things it helps with
67
60
 
68
- ### 1. Starting the next session without starting over
69
-
70
- A repository tells you where the code lives. It may not tell you why the customer rejected an approach, which access is still blocked or what the team promised on Tuesday.
71
-
72
- <a name="keep-a-customer-record"></a>
73
- <a name="how-skills-work"></a>
74
-
75
- For ongoing engagements, each customer gets a plain-Markdown record at `~/fde-engagements/<customer>/.fde/`. The coordinator loads a short summary and looks up details as needed. Before resuming implementation, it checks the saved next action against the current task and code. Saved lessons are searchable within that customer’s record. Meeting preparation brings back recorded open questions and commitments; sharing a lesson with another customer requires explicit approval.
76
-
77
- From the fictional demo’s `fde resume` output:
78
-
79
- ```text
80
- next: get the reconciliation runbook from Tom before touching anything. [source: meeting 2026-09-10]
81
- do first: Ask the acceptance owner to review the reported result and its evidence (delivery.md: 1 reported result awaiting acceptance)
82
- ```
83
-
84
- The next session can pick up the work while keeping acceptance pending.
61
+ ### 1. Strategy: decide what is worth building
85
62
 
86
- Use [debrief](skills/debrief/SKILL.md) after a meeting and [switch-clients](skills/switch-clients/SKILL.md) when changing customers. [How records work](docs/USAGE.md).
63
+ Turn the customer’s request into a problem to investigate, a measure of success and a bounded scope. Identify who decides and what evidence would change the plan.
87
64
 
88
- ### 2. Keeping a request from becoming an agreement
65
+ Use [discover](skills/discover/SKILL.md), [who-decides](skills/who-decides/SKILL.md) and [scope](skills/scope/SKILL.md).
89
66
 
90
- A stakeholder asks for more scope. A demo looks promising. Neither establishes a new commitment or an accepted result.
67
+ ### 2. Architecture: choose an approach that fits
91
68
 
92
- FDEOps keeps requests, confirmed decisions, reported results and open questions distinct. You review proposed record changes before saving them. Dates and sources keep claims traceable; customer approval still comes from the agreed owner.
69
+ Inspect the existing system, compare options against customer constraints and plan a small slice that tests the design. Make dependencies, tradeoffs and failure paths explicit.
93
70
 
94
- For example, these fictional notes:
71
+ Use [options](skills/options/SKILL.md), [plan](skills/plan/SKILL.md) and [integrate](skills/integrate/SKILL.md).
95
72
 
96
- > Mara agreed to keep CSV upload this phase. Devon asked for real-time sync; Mara has not answered. Two staging runs took 12 minutes. Production has not been measured.
73
+ ### 3. Engineering: build, verify and hand over
97
74
 
98
- The review separates them:
75
+ Implement the change, debug failures and test the agreed behavior. Report what passed on which revision and environment, what remains unproven and what the operating team needs before rollout.
99
76
 
100
- | Record | What the notes support |
101
- |---|---|
102
- | Decision | Keep CSV upload this phase; attributed to Mara in the supplied notes |
103
- | Request | Real-time sync remains unapproved |
104
- | Evidence | Two staging runs took 12 minutes; production benefit is unmeasured |
105
- | Next step | Resolve the scope request with Mara before changing the commitment |
106
-
107
- This is a draft, not a saved agreement. Use [who-decides](skills/who-decides/SKILL.md), [scope](skills/scope/SKILL.md) or [readout](skills/readout/SKILL.md) for the decision in front of you.
108
-
109
- ### 3. Knowing what is actually ready
110
-
111
- A local test, a deployed change and a customer-accepted result answer different questions.
112
-
113
- | Claim | Evidence it needs |
114
- |---|---|
115
- | Implemented | The change exists in the identified revision |
116
- | Verified | Applicable checks passed under stated conditions |
117
- | Deployed | The intended environment is running the change |
118
- | Measured | A result was observed against the agreed measure |
119
- | Accepted | The agreed owner or mechanism accepted the outcome |
120
-
121
- The skills use these distinctions when reporting progress; they are not automatic dashboard states.
77
+ Use [build](skills/build/SKILL.md), [debug](skills/debug/SKILL.md), [review](skills/review/SKILL.md), [ship](skills/ship/SKILL.md) and [handoff](skills/handoff/SKILL.md). A passing local test does not establish deployment or customer acceptance. [Verification and its limits](docs/verification.md).
122
78
 
123
- FDEOps carries agreed checks into implementation and ties test results to the revision and environment checked. Before rollout, it asks for operating limits, recovery evidence and an owner. You can see what is ready, what is blocked and what still needs verification.
79
+ <a name="keep-a-customer-record"></a>
80
+ <a name="how-skills-work"></a>
124
81
 
125
- Use [build](skills/build/SKILL.md), [integrate](skills/integrate/SKILL.md), [review](skills/review/SKILL.md), [ship](skills/ship/SKILL.md) and [handoff](skills/handoff/SKILL.md) as needed. [See the tests and their limits](docs/verification.md).
82
+ For ongoing work, a local Markdown record at `~/fde-engagements/<customer>/.fde/` carries decisions, evidence and next steps between sessions. The coordinator retrieves relevant context and prepares consequential updates for your review. Use [debrief](skills/debrief/SKILL.md) after meetings and [switch-clients](skills/switch-clients/SKILL.md) when changing customers. [How records work](docs/USAGE.md).
126
83
 
127
84
  ## Choose a skill
128
85
 
129
86
  <a name="task-skills"></a>
130
87
 
88
+ **35 task skills + one coordinator, `fde`**. Each task skill works on its own; the coordinator includes all underlying methods.
89
+
131
90
  | Work in front of you | Start with |
132
91
  |---|---|
133
92
  | An unclear customer request | `brief`, `discover` |
@@ -160,13 +119,12 @@ Copy an action into your agent to continue. Regenerate the view after record upd
160
119
  <a name="your-records-your-control"></a>
161
120
  <a name="your-data-stays-yours"></a>
162
121
 
163
- ## Local records, explicit data boundaries
164
-
165
- The CLI reads local files and Git without network calls or telemetry. Installation may download packages. Your AI host controls model connections and may transmit what it reads.
166
-
167
- CLI and hook outputs mask common identifier patterns. `<private>` blocks are redacted from those outputs and the dashboard. Local reports retain unmarked identifiers by default. These filters cover FDEOps output, not raw files or text you paste into an agent. Use only approved material, including when anonymised.
168
-
169
- You review consequential record updates. Enabled hooks can save mechanical session progress; direct CLI write commands update records when run. [Privacy](PRIVACY.md) · [Security](SECURITY.md) · [Local-model results](docs/verification.md#local-model-results).
122
+ > [!NOTE]
123
+ > **Use FDEOps with the setup that fits your work.** Try individual skills with sample data, or use a customer-approved local model or LLM provider for customer projects, including regulated and production work.
124
+ >
125
+ > Customer records stay in local files. The FDEOps CLI makes no network calls; your AI coding agent’s settings determine what it sends to a model provider. Use data approved for that setup.
126
+ >
127
+ > [Setup guidance](SECURITY.md#before-customer-work) · [Privacy and masking](PRIVACY.md) · [Local-model results](docs/verification.md#local-model-results)
170
128
 
171
129
  ## Who this is for
172
130
 
package/bin/check.js CHANGED
@@ -194,11 +194,13 @@ if (read('package.json').includes('postinstall')) {
194
194
  }
195
195
 
196
196
  const readme = read('README.md')
197
- if (!readme.includes('media/chat-demo.gif') || !readme.includes('media/chat-demo.md') || !readme.includes('Fictional customers')) {
198
- fail('README must include the chat walkthrough, text alternative and fictional-record disclosure')
199
- } else if (['chat-demo.gif', 'chat-demo.png', 'chat-demo.json', 'chat-demo.md', 'render-chat-demo.py'].some(name => !fs.existsSync(path.join(root, 'media', name)))) {
200
- fail('chat walkthrough must include rendered assets, text, source and renderer')
201
- } else ok('README chat walkthrough has accessible text and reproducible source')
197
+ if (readme.includes('media/chat-demo.gif')) {
198
+ if (!readme.includes('media/chat-demo.md') || !/fictional/i.test(readme)) {
199
+ fail('README animation must include a text alternative and fictional-data disclosure')
200
+ } else if (['chat-demo.gif', 'chat-demo.png', 'chat-demo.json', 'chat-demo.md', 'render-chat-demo.py'].some(name => !fs.existsSync(path.join(root, 'media', name)))) {
201
+ fail('chat walkthrough must include rendered assets, text, source and renderer')
202
+ } else ok('README animation has accessible text and reproducible source')
203
+ }
202
204
 
203
205
  const usage = read('docs/USAGE.md')
204
206
  if (!usage.includes('media/session.gif') || !usage.includes('media/record-session.sh')) {
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "fdeops-ingest-mcp",
3
- "version": "5.1.14",
3
+ "version": "5.1.15",
4
4
  "private": true,
5
5
  "description": "Thin stdio MCP sink for FDEOps ingest (stage \u2192 propose \u2192 apply). Zero runtime dependencies.",
6
6
  "bin": {
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "fdeops",
3
- "version": "5.1.14",
4
- "description": "Forward deployed engineering skills for AI coding agents. Use focused task skills or @fde for discovery, implementation, verification and handoff, with local engagement records.",
3
+ "version": "5.1.15",
4
+ "description": "Skills for forward deployed engineers across strategy, architecture and engineering. Use individual tasks or @fde coordination; local customer memory supports continuity.",
5
5
  "bin": {
6
6
  "fdeops": "bin/install.js",
7
7
  "fde": "bin/fde.js"
package/plugin.json CHANGED
@@ -1,8 +1,8 @@
1
1
  {
2
2
  "$schema": "https://agent-plugins.org/schemas/1.0.0/plugin.schema.json",
3
3
  "name": "fdeops",
4
- "version": "5.1.14",
5
- "description": "Forward deployed engineering skills for AI coding agents. Use focused task skills or @fde for discovery, implementation, verification and handoff, with local engagement records.",
4
+ "version": "5.1.15",
5
+ "description": "Skills for forward deployed engineers across strategy, architecture and engineering. Use individual tasks or @fde coordination; local customer memory supports continuity.",
6
6
  "author": {
7
7
  "name": "Subash Natarajan",
8
8
  "url": "https://github.com/suboss87"
@@ -16,7 +16,9 @@ A green check on synthetic data is not a validated solution. The person who can
16
16
 
17
17
  **1. Pick by score when several use cases compete.** Use the scoring model from `discover.md` - (Value × Data readiness) / Complexity. If discover or score-use-cases already produced a ranking, reuse it; never invent a third ranking.
18
18
 
19
- **2. Build the minimum that tests the assumption.** Timebox the experiment with the FDE; aim for a same-day result when access and evidence permit it. Skip cosmetic polish, but keep the input validation, access controls, and failure handling needed to protect the test environment and data. Label shortcuts and simulated inputs. The POC is done when the person who can say no has seen the evidence and reacted, not when the code looks finished.
19
+ **2. Build the minimum that tests the assumption.** Timebox the experiment with the FDE; aim for a same-day result when access and evidence permit it. Skip cosmetic polish, but keep the input validation, access controls, and failure handling needed to protect the test environment and data. Label shortcuts and simulated inputs. Complete the POC when the agreed assumption test has a recorded result and the responsible decision-maker has reviewed it; distinguish evidence from permission to proceed.
20
+
21
+ **2b. Observe use when the assumption concerns people.** For a workflow or usability claim, ask an affected user to attempt a representative task in a permitted environment. Record completion, errors, help needed and their feedback against the agreed pass/fail check. A sponsor liking the demo is not evidence that users can complete the task. If user access is unavailable, report that validation as pending and continue independent technical checks. A successful task trial supports usability under those conditions; sustained adoption needs evidence from actual use over an appropriate period. Carry findings into the next prototype or plan rather than treating feedback as automatic acceptance.
20
22
 
21
23
  **3. AI directions - test these before anything else:**
22
24
  - Data: available, clean, sufficient volume? Synthetic data can test mechanics, but does not establish production quality or real-world coverage.
@@ -47,7 +47,7 @@ CONVENIENCE - if wrong, a task changes but the approach holds
47
47
  | Assumption | Validation method | Effort | Evidence threshold |
48
48
  |-----------|-------------------|--------|-------------------|
49
49
  | "The API is the bottleneck" | Instrument the three slowest endpoints, measure p95 over 24h | 2h | Latency data shows >80% of wait time in API layer |
50
- | "The team will adopt the new tool" | Ask three team members individually: "Show me how you'd use this" | 1h | 2 of 3 can describe a use case without prompting |
50
+ | "Users can complete the target task with the prototype" | Observe affected users attempting a representative task in a permitted environment | Timebox agreed for the task | Pre-agreed completion, error and assistance criteria; report sample and limits, not adoption |
51
51
  | "The data is clean enough for ML" | Sample 200 records, count nulls/duplicates/format errors | 1h | <5% error rate on the fields the model needs |
52
52
 
53
53
  **4. Run the killer test first.** The assumption with the highest blast radius AND the cheapest validation gets tested immediately. This single principle saves more engagement time than any other: if the killer assumption is wrong, you've saved weeks; if it holds, you've bought confidence. Write the kill observation in `How we test` as the result that would **stop** the plan - plan copies that line onto each Now PR as `Kill if`.
@@ -68,7 +68,7 @@ Evidence first, then the question. Let them reach the conclusion.
68
68
  | # | Assumption | Kind | Blast radius | How we test | Status | Evidence |
69
69
  |---|------------|------|--------------|-------------|--------|----------|
70
70
  | 1 | API is the bottleneck | CONVENTION | CRITICAL | p95 instrumentation 24h | DISPROVED | 80% wait in DB layer (Day N) |
71
- | 2 | Team will adopt new tool | UNKNOWN | LOAD-BEARING | 3 individual interviews | CONFIRMED | 2/3 describe a use case unprompted |
71
+ | 2 | Team will adopt new tool | UNKNOWN | LOAD-BEARING | Observe task use, then assess sustained use over an agreed period | OPEN | 2/3 describe a use case unprompted; interest reported, use not yet observed |
72
72
  | 3 | Data clean enough for ML | UNKNOWN | CRITICAL | 200-record sample | PARTIAL → OPEN follow-up | 12% nulls on key field; cleaning task added |
73
73
  ```
74
74
 
@@ -98,5 +98,5 @@ Result: acked in 40 minutes, by Marco, not finance. Assumption DISPROVED, and th
98
98
  - Kind before blast radius. A FACT with no receipt is UNKNOWN.
99
99
  - Kill the riskiest, cheapest-to-test assumption first.
100
100
  - Evidence first, then the question. Let the customer reach the conclusion.
101
- - A brief with zero disproved assumptions wasn't audited - it was accepted.
101
+ - Design tests that could disprove consequential assumptions, and report what the evidence shows. All assumptions may survive a rigorous audit; never invent a contradiction to demonstrate skepticism.
102
102
  - Two weeks of building on a wrong assumption costs more than two hours of testing.
@@ -68,6 +68,8 @@ Time to detect: <5 minutes via error rate alert
68
68
 
69
69
  **Ask the team:** "Is anything outside this repo reading from or writing to <the thing you're changing>?" The answer is almost always "yes, and here's one we forgot about."
70
70
 
71
+ **5b. Check the consequences of success.** When a change alters throughput, workload or decision-making, trace what happens if it works as intended. Does faster intake move the queue to another team, increase review or recovery work, or reward a local metric while the overall outcome worsens? Use known capacity and observed behavior; label missing evidence rather than inventing downstream harm. Name the affected owner and an end-to-end outcome signal alongside the local improvement. Carry a relevant check into the existing plan and ship pulse. Skip this expansion when the change has no material workflow effect.
72
+
71
73
  **6. The 2am test.** For any SYSTEMIC or IRREVERSIBLE change, ask: "If this fails at 2am on Saturday, who gets woken up, what do they see, and what can they do?" If the answer is "they see nothing until Monday" - the monitoring plan needs work before the change ships.
72
74
 
73
75
  ## Artifact
@@ -6,7 +6,7 @@
6
6
  "agents/openai.yaml": "ff5da2d5215943e5c4e2a1339c11101035ed2741ded38e443af5180239e2a24b",
7
7
  "references/business-case.md": "32e000e8351cd59f9eaad8be40babb276df69948ea4f81e01a4672e47f48cb25",
8
8
  "references/task-context.md": "9066514a50043f3ad888d133d4e8b89b7132551e098cf2c80203c458a80126e5",
9
- "references/test-assumptions.md": "bf60d8bb4c0701fcffb196d78f7f6c8b1c472fc877fb2caf41058fbf8e2415a1",
9
+ "references/test-assumptions.md": "b3349a5f82e741e34fb0d17e16712a070b31d87f413e554be40374dd79152086",
10
10
  "references/three-options.md": "168fab9fb8ac8de85b0d1fa58e1db17deaa244cdef8c99623a04a0a6c70fe52c"
11
11
  }
12
12
  }
@@ -47,7 +47,7 @@ CONVENIENCE - if wrong, a task changes but the approach holds
47
47
  | Assumption | Validation method | Effort | Evidence threshold |
48
48
  |-----------|-------------------|--------|-------------------|
49
49
  | "The API is the bottleneck" | Instrument the three slowest endpoints, measure p95 over 24h | 2h | Latency data shows >80% of wait time in API layer |
50
- | "The team will adopt the new tool" | Ask three team members individually: "Show me how you'd use this" | 1h | 2 of 3 can describe a use case without prompting |
50
+ | "Users can complete the target task with the prototype" | Observe affected users attempting a representative task in a permitted environment | Timebox agreed for the task | Pre-agreed completion, error and assistance criteria; report sample and limits, not adoption |
51
51
  | "The data is clean enough for ML" | Sample 200 records, count nulls/duplicates/format errors | 1h | <5% error rate on the fields the model needs |
52
52
 
53
53
  **4. Run the killer test first.** The assumption with the highest blast radius AND the cheapest validation gets tested immediately. This single principle saves more engagement time than any other: if the killer assumption is wrong, you've saved weeks; if it holds, you've bought confidence. Write the kill observation in `How we test` as the result that would **stop** the plan - plan copies that line onto each Now PR as `Kill if`.
@@ -68,7 +68,7 @@ Evidence first, then the question. Let them reach the conclusion.
68
68
  | # | Assumption | Kind | Blast radius | How we test | Status | Evidence |
69
69
  |---|------------|------|--------------|-------------|--------|----------|
70
70
  | 1 | API is the bottleneck | CONVENTION | CRITICAL | p95 instrumentation 24h | DISPROVED | 80% wait in DB layer (Day N) |
71
- | 2 | Team will adopt new tool | UNKNOWN | LOAD-BEARING | 3 individual interviews | CONFIRMED | 2/3 describe a use case unprompted |
71
+ | 2 | Team will adopt new tool | UNKNOWN | LOAD-BEARING | Observe task use, then assess sustained use over an agreed period | OPEN | 2/3 describe a use case unprompted; interest reported, use not yet observed |
72
72
  | 3 | Data clean enough for ML | UNKNOWN | CRITICAL | 200-record sample | PARTIAL → OPEN follow-up | 12% nulls on key field; cleaning task added |
73
73
  ```
74
74
 
@@ -98,5 +98,5 @@ Result: acked in 40 minutes, by Marco, not finance. Assumption DISPROVED, and th
98
98
  - Kind before blast radius. A FACT with no receipt is UNKNOWN.
99
99
  - Kill the riskiest, cheapest-to-test assumption first.
100
100
  - Evidence first, then the question. Let the customer reach the conclusion.
101
- - A brief with zero disproved assumptions wasn't audited - it was accepted.
101
+ - Design tests that could disprove consequential assumptions, and report what the evidence shows. All assumptions may survive a rigorous audit; never invent a contradiction to demonstrate skepticism.
102
102
  - Two weeks of building on a wrong assumption costs more than two hours of testing.
@@ -12,12 +12,12 @@
12
12
  "references/eval-pack.md": "0590b85d3cae0903c6b1274540c92eaa2a4373047e8a0548d6942516ef0bb9e1",
13
13
  "references/integrate.md": "1cb7a60d7545b0bf224fce678a04ce6ccdf368c47877d9c8e4dc4272bb0d5b0c",
14
14
  "references/plan.md": "6c37976169723f93d3554092979569a9640bcf02bd179bad52b17815c1c30f86",
15
- "references/poc.md": "818dc90c2d401819735233ab9d69df171675d75abf8644c403dab7dd1dcf9199",
15
+ "references/poc.md": "65c91866037dc034a64484683ad9a329896786f1487265c9fddecd19b09a3039",
16
16
  "references/qa.md": "41de4d70827c83291efa217e97d777f62ec2849827687fbba7e4b1d17484b87e",
17
17
  "references/review.md": "55733ca868c00fb22110bc7b3ec7bb6c6451366c795073b19a2bccd30d2764c8",
18
18
  "references/ship.md": "95f51772b29de6f7d174f1f7678f15d46a3a5327dcb8facb94288e27907ae92c",
19
19
  "references/task-context.md": "9066514a50043f3ad888d133d4e8b89b7132551e098cf2c80203c458a80126e5",
20
- "references/test-assumptions.md": "bf60d8bb4c0701fcffb196d78f7f6c8b1c472fc877fb2caf41058fbf8e2415a1",
20
+ "references/test-assumptions.md": "b3349a5f82e741e34fb0d17e16712a070b31d87f413e554be40374dd79152086",
21
21
  "references/three-options.md": "168fab9fb8ac8de85b0d1fa58e1db17deaa244cdef8c99623a04a0a6c70fe52c",
22
22
  "references/verification.md": "8ee2502112a37eedfdacf929041e7b91e9e6fabb007408a1ee2e2992e3af6ff6"
23
23
  }
@@ -16,7 +16,9 @@ A green check on synthetic data is not a validated solution. The person who can
16
16
 
17
17
  **1. Pick by score when several use cases compete.** Use the scoring model from `discover.md` - (Value × Data readiness) / Complexity. If discover or score-use-cases already produced a ranking, reuse it; never invent a third ranking.
18
18
 
19
- **2. Build the minimum that tests the assumption.** Timebox the experiment with the FDE; aim for a same-day result when access and evidence permit it. Skip cosmetic polish, but keep the input validation, access controls, and failure handling needed to protect the test environment and data. Label shortcuts and simulated inputs. The POC is done when the person who can say no has seen the evidence and reacted, not when the code looks finished.
19
+ **2. Build the minimum that tests the assumption.** Timebox the experiment with the FDE; aim for a same-day result when access and evidence permit it. Skip cosmetic polish, but keep the input validation, access controls, and failure handling needed to protect the test environment and data. Label shortcuts and simulated inputs. Complete the POC when the agreed assumption test has a recorded result and the responsible decision-maker has reviewed it; distinguish evidence from permission to proceed.
20
+
21
+ **2b. Observe use when the assumption concerns people.** For a workflow or usability claim, ask an affected user to attempt a representative task in a permitted environment. Record completion, errors, help needed and their feedback against the agreed pass/fail check. A sponsor liking the demo is not evidence that users can complete the task. If user access is unavailable, report that validation as pending and continue independent technical checks. A successful task trial supports usability under those conditions; sustained adoption needs evidence from actual use over an appropriate period. Carry findings into the next prototype or plan rather than treating feedback as automatic acceptance.
20
22
 
21
23
  **3. AI directions - test these before anything else:**
22
24
  - Data: available, clean, sufficient volume? Synthetic data can test mechanics, but does not establish production quality or real-world coverage.
@@ -47,7 +47,7 @@ CONVENIENCE - if wrong, a task changes but the approach holds
47
47
  | Assumption | Validation method | Effort | Evidence threshold |
48
48
  |-----------|-------------------|--------|-------------------|
49
49
  | "The API is the bottleneck" | Instrument the three slowest endpoints, measure p95 over 24h | 2h | Latency data shows >80% of wait time in API layer |
50
- | "The team will adopt the new tool" | Ask three team members individually: "Show me how you'd use this" | 1h | 2 of 3 can describe a use case without prompting |
50
+ | "Users can complete the target task with the prototype" | Observe affected users attempting a representative task in a permitted environment | Timebox agreed for the task | Pre-agreed completion, error and assistance criteria; report sample and limits, not adoption |
51
51
  | "The data is clean enough for ML" | Sample 200 records, count nulls/duplicates/format errors | 1h | <5% error rate on the fields the model needs |
52
52
 
53
53
  **4. Run the killer test first.** The assumption with the highest blast radius AND the cheapest validation gets tested immediately. This single principle saves more engagement time than any other: if the killer assumption is wrong, you've saved weeks; if it holds, you've bought confidence. Write the kill observation in `How we test` as the result that would **stop** the plan - plan copies that line onto each Now PR as `Kill if`.
@@ -68,7 +68,7 @@ Evidence first, then the question. Let them reach the conclusion.
68
68
  | # | Assumption | Kind | Blast radius | How we test | Status | Evidence |
69
69
  |---|------------|------|--------------|-------------|--------|----------|
70
70
  | 1 | API is the bottleneck | CONVENTION | CRITICAL | p95 instrumentation 24h | DISPROVED | 80% wait in DB layer (Day N) |
71
- | 2 | Team will adopt new tool | UNKNOWN | LOAD-BEARING | 3 individual interviews | CONFIRMED | 2/3 describe a use case unprompted |
71
+ | 2 | Team will adopt new tool | UNKNOWN | LOAD-BEARING | Observe task use, then assess sustained use over an agreed period | OPEN | 2/3 describe a use case unprompted; interest reported, use not yet observed |
72
72
  | 3 | Data clean enough for ML | UNKNOWN | CRITICAL | 200-record sample | PARTIAL → OPEN follow-up | 12% nulls on key field; cleaning task added |
73
73
  ```
74
74
 
@@ -98,5 +98,5 @@ Result: acked in 40 minutes, by Marco, not finance. Assumption DISPROVED, and th
98
98
  - Kind before blast radius. A FACT with no receipt is UNKNOWN.
99
99
  - Kill the riskiest, cheapest-to-test assumption first.
100
100
  - Evidence first, then the question. Let the customer reach the conclusion.
101
- - A brief with zero disproved assumptions wasn't audited - it was accepted.
101
+ - Design tests that could disprove consequential assumptions, and report what the evidence shows. All assumptions may survive a rigorous audit; never invent a contradiction to demonstrate skepticism.
102
102
  - Two weeks of building on a wrong assumption costs more than two hours of testing.
@@ -5,6 +5,6 @@
5
5
  "SKILL.md": "acba10ec77763f1840c3cda95b26e13150683f6a7557a84a93611f2d96cdd389",
6
6
  "agents/openai.yaml": "c67e0342caba79241b1872cdfda8341f4949c2bd7672b8d288b529d8e94342b8",
7
7
  "references/task-context.md": "9066514a50043f3ad888d133d4e8b89b7132551e098cf2c80203c458a80126e5",
8
- "references/test-assumptions.md": "bf60d8bb4c0701fcffb196d78f7f6c8b1c472fc877fb2caf41058fbf8e2415a1"
8
+ "references/test-assumptions.md": "b3349a5f82e741e34fb0d17e16712a070b31d87f413e554be40374dd79152086"
9
9
  }
10
10
  }
@@ -47,7 +47,7 @@ CONVENIENCE - if wrong, a task changes but the approach holds
47
47
  | Assumption | Validation method | Effort | Evidence threshold |
48
48
  |-----------|-------------------|--------|-------------------|
49
49
  | "The API is the bottleneck" | Instrument the three slowest endpoints, measure p95 over 24h | 2h | Latency data shows >80% of wait time in API layer |
50
- | "The team will adopt the new tool" | Ask three team members individually: "Show me how you'd use this" | 1h | 2 of 3 can describe a use case without prompting |
50
+ | "Users can complete the target task with the prototype" | Observe affected users attempting a representative task in a permitted environment | Timebox agreed for the task | Pre-agreed completion, error and assistance criteria; report sample and limits, not adoption |
51
51
  | "The data is clean enough for ML" | Sample 200 records, count nulls/duplicates/format errors | 1h | <5% error rate on the fields the model needs |
52
52
 
53
53
  **4. Run the killer test first.** The assumption with the highest blast radius AND the cheapest validation gets tested immediately. This single principle saves more engagement time than any other: if the killer assumption is wrong, you've saved weeks; if it holds, you've bought confidence. Write the kill observation in `How we test` as the result that would **stop** the plan - plan copies that line onto each Now PR as `Kill if`.
@@ -68,7 +68,7 @@ Evidence first, then the question. Let them reach the conclusion.
68
68
  | # | Assumption | Kind | Blast radius | How we test | Status | Evidence |
69
69
  |---|------------|------|--------------|-------------|--------|----------|
70
70
  | 1 | API is the bottleneck | CONVENTION | CRITICAL | p95 instrumentation 24h | DISPROVED | 80% wait in DB layer (Day N) |
71
- | 2 | Team will adopt new tool | UNKNOWN | LOAD-BEARING | 3 individual interviews | CONFIRMED | 2/3 describe a use case unprompted |
71
+ | 2 | Team will adopt new tool | UNKNOWN | LOAD-BEARING | Observe task use, then assess sustained use over an agreed period | OPEN | 2/3 describe a use case unprompted; interest reported, use not yet observed |
72
72
  | 3 | Data clean enough for ML | UNKNOWN | CRITICAL | 200-record sample | PARTIAL → OPEN follow-up | 12% nulls on key field; cleaning task added |
73
73
  ```
74
74
 
@@ -98,5 +98,5 @@ Result: acked in 40 minutes, by Marco, not finance. Assumption DISPROVED, and th
98
98
  - Kind before blast radius. A FACT with no receipt is UNKNOWN.
99
99
  - Kill the riskiest, cheapest-to-test assumption first.
100
100
  - Evidence first, then the question. Let the customer reach the conclusion.
101
- - A brief with zero disproved assumptions wasn't audited - it was accepted.
101
+ - Design tests that could disprove consequential assumptions, and report what the evidence shows. All assumptions may survive a rigorous audit; never invent a contradiction to demonstrate skepticism.
102
102
  - Two weeks of building on a wrong assumption costs more than two hours of testing.
@@ -5,6 +5,6 @@
5
5
  "SKILL.md": "cacffe6b01d36cde634be5e7f8a2127d701d3dac90b3000c414a648956839c5e",
6
6
  "agents/openai.yaml": "6c4b97d3268698d1eed0b628f9872351f29b91dd47c9de371a12c584983cb0fe",
7
7
  "references/task-context.md": "9066514a50043f3ad888d133d4e8b89b7132551e098cf2c80203c458a80126e5",
8
- "references/what-breaks.md": "bc7f7b0d0dbaa4df520b263f0877474722ab223aac4cd6fd84b5112d9868715d"
8
+ "references/what-breaks.md": "fbf85539fe9c762d713c5772dbef731e96783d369610ba61b66fc220d9e50abf"
9
9
  }
10
10
  }
@@ -68,6 +68,8 @@ Time to detect: <5 minutes via error rate alert
68
68
 
69
69
  **Ask the team:** "Is anything outside this repo reading from or writing to <the thing you're changing>?" The answer is almost always "yes, and here's one we forgot about."
70
70
 
71
+ **5b. Check the consequences of success.** When a change alters throughput, workload or decision-making, trace what happens if it works as intended. Does faster intake move the queue to another team, increase review or recovery work, or reward a local metric while the overall outcome worsens? Use known capacity and observed behavior; label missing evidence rather than inventing downstream harm. Name the affected owner and an end-to-end outcome signal alongside the local improvement. Carry a relevant check into the existing plan and ship pulse. Skip this expansion when the change has no material workflow effect.
72
+
71
73
  **6. The 2am test.** For any SYSTEMIC or IRREVERSIBLE change, ask: "If this fails at 2am on Saturday, who gets woken up, what do they see, and what can they do?" If the answer is "they see nothing until Monday" - the monitoring plan needs work before the change ships.
72
74
 
73
75
  ## Artifact