fdeops 5.1.6 → 5.1.8

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (75) hide show
  1. package/README.md +73 -102
  2. package/mcp/fdeops-ingest/package.json +1 -1
  3. package/package.json +1 -1
  4. package/plugin.json +1 -1
  5. package/skills/build/.fde-generated.json +6 -6
  6. package/skills/build/references/build.md +24 -13
  7. package/skills/build/references/debug.md +26 -11
  8. package/skills/build/references/integrate.md +33 -18
  9. package/skills/build/references/qa.md +24 -11
  10. package/skills/build/references/ship.md +2 -2
  11. package/skills/build/references/verification.md +1 -1
  12. package/skills/debug/.fde-generated.json +6 -6
  13. package/skills/debug/references/build.md +24 -13
  14. package/skills/debug/references/debug.md +26 -11
  15. package/skills/debug/references/integrate.md +33 -18
  16. package/skills/debug/references/qa.md +24 -11
  17. package/skills/debug/references/ship.md +2 -2
  18. package/skills/debug/references/verification.md +1 -1
  19. package/skills/earn-trust/.fde-generated.json +1 -1
  20. package/skills/earn-trust/references/earn-trust.md +1 -1
  21. package/skills/evaluate/.fde-generated.json +6 -6
  22. package/skills/evaluate/references/build.md +24 -13
  23. package/skills/evaluate/references/debug.md +26 -11
  24. package/skills/evaluate/references/integrate.md +33 -18
  25. package/skills/evaluate/references/qa.md +24 -11
  26. package/skills/evaluate/references/ship.md +2 -2
  27. package/skills/evaluate/references/verification.md +1 -1
  28. package/skills/fde/references/build.md +24 -13
  29. package/skills/fde/references/debug.md +26 -11
  30. package/skills/fde/references/earn-trust.md +1 -1
  31. package/skills/fde/references/integrate.md +33 -18
  32. package/skills/fde/references/plan.md +5 -2
  33. package/skills/fde/references/qa.md +24 -11
  34. package/skills/fde/references/ship.md +2 -2
  35. package/skills/fde/references/verification.md +1 -1
  36. package/skills/integrate/.fde-generated.json +6 -6
  37. package/skills/integrate/references/build.md +24 -13
  38. package/skills/integrate/references/debug.md +26 -11
  39. package/skills/integrate/references/integrate.md +33 -18
  40. package/skills/integrate/references/qa.md +24 -11
  41. package/skills/integrate/references/ship.md +2 -2
  42. package/skills/integrate/references/verification.md +1 -1
  43. package/skills/plan/.fde-generated.json +1 -1
  44. package/skills/plan/references/plan.md +5 -2
  45. package/skills/poc/.fde-generated.json +7 -7
  46. package/skills/poc/references/build.md +24 -13
  47. package/skills/poc/references/debug.md +26 -11
  48. package/skills/poc/references/integrate.md +33 -18
  49. package/skills/poc/references/plan.md +5 -2
  50. package/skills/poc/references/qa.md +24 -11
  51. package/skills/poc/references/ship.md +2 -2
  52. package/skills/poc/references/verification.md +1 -1
  53. package/skills/qa/.fde-generated.json +6 -6
  54. package/skills/qa/references/build.md +24 -13
  55. package/skills/qa/references/debug.md +26 -11
  56. package/skills/qa/references/integrate.md +33 -18
  57. package/skills/qa/references/qa.md +24 -11
  58. package/skills/qa/references/ship.md +2 -2
  59. package/skills/qa/references/verification.md +1 -1
  60. package/skills/review/.fde-generated.json +6 -6
  61. package/skills/review/references/build.md +24 -13
  62. package/skills/review/references/debug.md +26 -11
  63. package/skills/review/references/integrate.md +33 -18
  64. package/skills/review/references/qa.md +24 -11
  65. package/skills/review/references/ship.md +2 -2
  66. package/skills/review/references/verification.md +1 -1
  67. package/skills/runbook/.fde-generated.json +1 -1
  68. package/skills/runbook/references/verification.md +1 -1
  69. package/skills/ship/.fde-generated.json +6 -6
  70. package/skills/ship/references/build.md +24 -13
  71. package/skills/ship/references/debug.md +26 -11
  72. package/skills/ship/references/integrate.md +33 -18
  73. package/skills/ship/references/qa.md +24 -11
  74. package/skills/ship/references/ship.md +2 -2
  75. package/skills/ship/references/verification.md +1 -1
@@ -1,22 +1,33 @@
1
1
  # build - Implement a verifiable increment
2
2
 
3
- **Enter when:** an agreed behavior needs implementation in an existing or new repository. For a broken behavior, start with [debug](debug.md); for a system boundary, use [integrate](integrate.md).
3
+ Build the smallest complete change that demonstrates the agreed customer outcome through the real entry point.
4
4
 
5
- Use the permitted context and authority in [task context](task-context.md). This method works without `.fde/`; an existing engagement record can supply the same contract. Do not initialize memory just to write code.
5
+ **Use when:** agreed behavior needs implementation in a new or existing repository. Use [debug](debug.md) for broken behavior and [integrate](integrate.md) for a system boundary.
6
6
 
7
- ## Method
7
+ Follow [task context](task-context.md). Supplied permitted context or an existing engagement record can provide the contract; do not initialize `.fde/` merely to write code.
8
8
 
9
- 1. Identify the repository, its instructions, working tree, relevant callers, and test commands. Inspect examples before creating abstractions. Follow the repository's branch policy and choose any needed checkout isolation according to that policy and overlapping work. Preserve unrelated edits and state which dependencies or interfaces the change touches. Before changing an untested legacy path, capture the undocumented behavior callers depend on with targeted characterization checks; distinguish behavior to preserve from the intended change.
10
- 2. State the observable outcome, constraints, and acceptance checks. Reuse agreed criteria for routine fixes. If a consequential product choice is unresolved, surface that choice while continuing independent investigation; do not invent acceptance.
11
- 3. Choose the smallest coherent path that demonstrates the outcome through the real entry point. Include the necessary storage, error handling, and interface behavior in that slice. Name the failure that stops expansion and the recovery path for stateful changes.
12
- 4. Implement using the repository's tools and conventions. Search for existing services, fixtures, and validation before adding alternatives. Keep cleanup limited to what makes the changed path understandable; do not expand scope to repair unrelated code. When changing dependencies, inspect the package source, requested version, lockfile changes and repository install-script policy before executing package code. Use the approved package manager and bootstrap controls; do not blanket-enable scripts or apply unrelated dependency upgrades.
13
- 5. Add or update automated coverage for changed behavior when meaningful and feasible, including the relevant failure path. Existing tests must actually exercise the change; explain manual-only coverage and its limits. Run focused checks, then required repository checks. Exercise the actual affected journey with [QA](qa.md) when appropriate. For uncertain model behavior, use [eval-pack](eval-pack.md). Record results with [verification](verification.md), including checks that could not run.
14
- 6. Inspect the final diff against the agreed outcome. If public behavior, interfaces, configuration or operating steps changed, update affected existing documentation and examples; exercise relevant commands or clearly mark checks that could not run. For substantial or risky work, seek [review](review.md) using an actual separate reviewer when available; identify a self-check honestly. Reverify affected behavior after fixes.
9
+ ## Understand the repository and outcome
15
10
 
16
- ## Deliverable and acceptance
11
+ Inspect repository instructions, working tree, relevant callers, examples, and test commands. Preserve unrelated edits. Follow its branch policy and choose checkout isolation according to overlapping work. Identify the dependencies and interfaces the change touches. Before changing an untested legacy path, use targeted characterization checks to capture undocumented behavior callers rely on; separate that behavior from the intended change.
17
12
 
18
- For substantial work, maintain the [recoverable checkpoint](verification.md#recoverable-checkpoint) in the existing task record as slices complete or work pauses.
13
+ State the observable outcome, constraints, and acceptance checks, reusing agreed criteria for routine fixes. Surface unresolved consequential product choices while continuing independent investigation; do not invent acceptance.
19
14
 
20
- Return the implemented behavior, relevant paths, evidence, remaining limitations, and any decision needed. Done means the agreed checks have applicable evidence and the change is reviewable; passing tests does not imply deployment or customer acceptance. Committing, opening a PR, merging, and publishing happen only when the requested workflow authorizes those actions.
15
+ ## Build one complete slice
21
16
 
22
- When coordinated through `@fde`, record implementation and verification in the existing decisions/delivery records under their write rules. Standalone work can return the same receipt directly or use the repository's task record.
17
+ Choose a coherent path through the real entry point, including its necessary storage, error handling, and interface behavior. Name the failure that would stop expansion and the recovery path for stateful changes.
18
+
19
+ Use existing services, fixtures, validation, and repository conventions before adding alternatives. Limit cleanup to making the changed path understandable. For dependency changes, inspect the package source, requested version, lockfile changes, and install-script policy before executing package code. Use the approved package manager and bootstrap controls; do not blanket-enable scripts or include unrelated upgrades.
20
+
21
+ ## Demonstrate the behavior
22
+
23
+ Add or update automated coverage when meaningful and feasible, including the relevant failure path. Check that existing tests actually exercise the change. Explain manual-only coverage and its limits. Run focused checks, then required repository checks; use [QA](qa.md) for the affected journey when appropriate and [eval-pack](eval-pack.md) for uncertain model behavior. Record evidence and unrun checks with [verification](verification.md).
24
+
25
+ Inspect the final diff against the agreed outcome. Update affected existing documentation and examples when public behavior, interfaces, configuration, or operating steps change. Exercise relevant commands or state what could not run. For substantial or risky work, use [review](review.md) with a separate reviewer when available; label a self-check honestly. Reverify affected behavior after repairs.
26
+
27
+ *Fictional example:* Northstar needs failed imports to be recoverable. A useful first slice takes one failed import through the existing retry action to a persisted result, including the retry's failure behavior. A new button alone does not demonstrate recovery.
28
+
29
+ ## Completion
30
+
31
+ Return implemented behavior, relevant paths, evidence, limitations, and any decision needed. The change is ready when agreed checks have applicable evidence and the work is reviewable. Passing tests does not establish deployment or customer acceptance. Commit, open a PR, merge, or publish only when the requested workflow authorizes it.
32
+
33
+ For substantial work, maintain a [recoverable checkpoint](verification.md#recoverable-checkpoint) in the existing task record as slices finish or work pauses. When coordinated through `@fde`, record implementation and verification in existing decisions/delivery records under their write rules. Standalone work can return the receipt directly or use the repository's task record.
@@ -1,18 +1,33 @@
1
1
  # debug - Find and repair the cause
2
2
 
3
- **Enter when:** a reproducible failure, regression, incident symptom, or misleading output needs investigation.
3
+ A useful repair explains the customer's failure and shows why the changed path now behaves correctly.
4
4
 
5
- Use [task context](task-context.md). Work from supplied permitted evidence without requiring `.fde/`. During an active incident, follow the authorized containment procedure before diagnosis; investigation authority alone does not authorize production writes.
5
+ **Use when:** a failure, regression, incident symptom, or misleading output needs investigation, whether or not it can yet be reproduced.
6
6
 
7
- ## Method
7
+ Follow [task context](task-context.md); permitted supplied evidence is enough without `.fde/`. During an incident, follow authorized containment procedures before diagnosis. Investigation authority does not authorize production writes.
8
8
 
9
- 1. Capture expected and observed behavior, exact input or trigger, affected revision/environment, and the last known working state. Preserve useful errors and timestamps without copying secrets or raw private data. Mark reports you have not reproduced as reports.
10
- 2. Inspect the failing path, callers, recent relevant changes, and existing tests. Reproduce in a permitted environment with the smallest representative case. If reproduction is unavailable, identify what observation would distinguish causes and gather safe evidence; do not claim a hypothesis is proven.
11
- 3. Keep a short hypothesis list. For each, name the predicted observation and a discriminating check. Change one relevant variable at a time. Trace values and control flow across the actual boundary instead of repeatedly changing code until the symptom disappears.
12
- 4. Fix the cause at the appropriate layer. Check whether the proposed fix changes behavior for other callers, stale data, retries, concurrency, or permissions. Preserve evidence of the original failure and avoid unrelated cleanup.
13
- 5. Add a regression check when it can meaningfully reproduce the bug; show that it fails before the fix and passes after when practical. If the check cannot run against the before-state, say so. Run affected adjacent and required checks using [verification](verification.md).
14
- 6. Review the final diff and exercise the original journey. For substantial or risky fixes use [review](review.md). After two unsuccessful repair cycles, reassess the hypothesis and evidence instead of repeating the same attempt; continue useful investigation and isolate the missing decision or access.
9
+ ## Establish what failed
15
10
 
16
- ## Deliverable and acceptance
11
+ Capture expected and observed behavior, the trigger or input, revision, environment, and last known working state. Keep useful errors and timestamps, without secrets or raw private data. Label unverified reports as reports.
17
12
 
18
- Report the cause with its evidence, the fix, the original reproducer's result, adjacent checks, and unresolved uncertainty. A disappearing symptom with no discriminating evidence is a mitigation, not a demonstrated root cause. In engagement mode record the incident/fix receipt in the appropriate existing record; standalone work may return it directly. Release or rollback requires the existing operational authority and [ship](ship.md) or recovery procedure.
13
+ Inspect the affected path, callers, relevant changes, and existing tests. Reproduce with the smallest representative case in a permitted environment when practical. If reproduction is unavailable, state the gap and use traces or other safe observations to distinguish causes. Keep investigating without promoting a hypothesis to a finding.
14
+
15
+ ## Test the explanation
16
+
17
+ Keep a short hypothesis list with a predicted observation and a discriminating check for each. Trace values and control flow across the actual boundary. Change one relevant variable at a time so the result tells you something.
18
+
19
+ After two unsuccessful repair cycles, reassess the evidence and approach. That is a signal to reconsider, not proof that a hypothesis is false. Continue useful investigation and identify any missing decision or access.
20
+
21
+ ## Repair and verify the affected path
22
+
23
+ Fix the cause at the appropriate layer, preserving evidence of the original failure. Consider other callers, stale data, retries, concurrency, and permissions; leave unrelated cleanup out.
24
+
25
+ Add a regression check when it can meaningfully reproduce the bug. Show failure before and success after when practical, and disclose when the before-state could not be checked. Exercise the original journey and run affected adjacent and required checks using [verification](verification.md). Review the diff; use [review](review.md) for substantial or risky fixes.
26
+
27
+ *Fictional example:* Northstar's imports sometimes duplicate orders. A lost-response trace suggests a retry after a committed write. If staging cannot reproduce it, report the supported hypothesis and missing evidence rather than calling a longer timeout a root-cause fix.
28
+
29
+ ## Completion
30
+
31
+ Return the cause and evidence, repair, original journey or reproducer result, adjacent checks, and unresolved uncertainty. A disappearing symptom without discriminating evidence establishes a mitigation, not a demonstrated root cause.
32
+
33
+ In an engagement, put the incident/fix receipt in the appropriate existing record under its write rules; standalone work can return it directly. Release or rollback needs the existing operational authority and [ship](ship.md) or recovery procedure.
@@ -1,28 +1,43 @@
1
1
  # integrate - Prove the system boundary
2
2
 
3
- **Enter when:** connecting an API, data source, SDK, event stream, tool, or service, or changing its contract.
3
+ An integration works when an input crosses the real boundary and produces the agreed downstream result.
4
4
 
5
- Start from [task context](task-context.md). Permitted supplied context is enough; `.fde/` is optional. Use the customer's existing clients, authentication, fixtures, and diagnostic tools. Do not create another integration platform to make one connection.
5
+ **Use when:** connecting or changing an API, data source, SDK, event stream, tool, or service contract.
6
6
 
7
- ## Method
7
+ Follow [task context](task-context.md); `.fde/` is optional. Use the customer's existing clients, authentication, fixtures, and diagnostics. One connection does not call for a new integration platform.
8
8
 
9
- 1. Map producer, consumer, owner, direction, and side effects. Inspect the actual installed version and local implementation; verify uncertain behavior against current official documentation. Identify the relevant schema, authentication scopes, network boundary, and permitted test environment. Before changing an untested existing boundary, characterize the mappings, ordering or other observable behavior its callers depend on.
10
- 2. Write the acceptance example: an input at the real boundary and the observable downstream result. Include a rejection or failure example. Separate configuration validity, successful authentication, transport connectivity, contract compatibility, and end-to-end behavior; none proves the next.
11
- 3. Inspect credentials by presence and required scope without printing values. Use existing secret storage. Check data classification and retention before moving data; never pass raw `<private>` blocks into a model. Prefer sanitized or synthetic cases approved for the target environment.
12
- 4. Implement the narrow adapter using native repository patterns. Validate external inputs and model outputs, bound timeouts and retries, preserve error codes and failure phase without leaking payloads, and handle cancellation. Keep explicit authentication or permission rejections distinguishable from transport uncertainty. For writes, establish idempotency or duplicate detection before retries and apply the uncertain-write rules below when outcomes can be ambiguous; for events, check ordering, replay, and poison messages as applicable.
13
- 5. Exercise a permitted success case and relevant failures: denied access, malformed data, rate limit, timeout, duplicate delivery, or partial completion. Trace correlation IDs or safe evidence across both sides. A mock proves client behavior only; if live access is unavailable, report that gap instead of claiming an integration works.
14
- 6. Check cleanup and recovery for test side effects. Use [verification](verification.md) for receipts and [review](review.md) for security or data-contract changes. Route deployment through [ship](ship.md) only when authorized.
9
+ ## Define the boundary and evidence
15
10
 
16
- ## When a write outcome is uncertain
11
+ Map producer, consumer, owner, direction, and side effects. Inspect the installed version and local implementation; check uncertain behavior against current official documentation. Identify schemas, authentication scopes, network boundaries, and the permitted test environment. Before changing an untested boundary, characterize mappings, ordering, and other behavior its callers depend on.
17
12
 
18
- Use the existing storage and worker mechanisms; do not introduce a new platform for these rules.
13
+ Write a real input and expected downstream result, plus a rejection or failure example. Keep configuration validity, authentication, connectivity, contract compatibility, and end-to-end behavior separate: evidence for one does not prove the next.
19
14
 
20
- - Before a replayable write, persist its tenant-scoped operation identity and payload identity, with enough state to recover after restart. Establish who owns an in-flight attempt so concurrent workers cannot independently replay it. Establish whether upstream deduplication is guaranteed, including its key, payload rules and retention window; sending a key alone proves nothing.
21
- - A timeout, cancellation or lost response after dispatch may leave a completed side effect. Preserve that uncertainty across restart; stopping the caller is not rollback. Do not silently turn an uncertain attempt into a fresh operation.
22
- - Reconcile against an authoritative receipt or lookup that matches the operation and payload. One verified result can confirm completion; conflicting or multiple matches require resolution. An empty stale, partial or eventually consistent lookup does not prove absence or authorize replay. Retry only under the verified deduplication contract or evidence establishing that repeating the write is safe.
23
- - Keep unresolved attempts visible with safe error context, a next action and a known resolution owner, or an explicit ownership gap. Manual resolution still needs authority for any corrective write; do not manufacture completion to clear a queue.
24
- - Test the relevant failure boundary: committed write with lost response, cancellation or restart before recording success, and stale lookup or concurrent replay where applicable. Record which were exercised and which remain unproven.
15
+ Inspect credentials by presence and scope without printing values; use existing secret storage. Check classification and retention before moving data. Never load raw `<private>` blocks into a model. Prefer sanitized or synthetic cases approved for the target environment.
25
16
 
26
- ## Deliverable and acceptance
17
+ ## Implement the narrow adapter
27
18
 
28
- Return the boundary contract, changed paths, environment, evidence at each tested layer, and remaining dependencies with owners when known. Done requires the agreed end-to-end result or an explicit narrower agreed scope. Do not silently replace live acceptance with a stub. In engagement mode, update the terrain/delivery record with confirmed facts; otherwise return the receipt directly.
19
+ Use native repository patterns. Validate external inputs and model outputs, bound timeouts and retries, and handle cancellation. Preserve error codes and the failure phase without leaking payloads. Keep explicit authentication or permission rejection distinguishable from transport uncertainty.
20
+
21
+ For writes, establish idempotency or duplicate detection before retries and apply the uncertain-write rules below. For events, address ordering, replay, and poison messages where relevant.
22
+
23
+ ## Resolve uncertain writes safely
24
+
25
+ Use existing storage and worker mechanisms for these rules:
26
+
27
+ - **Identify the attempt before dispatch.** For a replayable write, persist a tenant-scoped operation identity, payload identity, and enough state to recover after restart. Establish ownership of in-flight attempts so workers cannot independently replay them. Verify any upstream deduplication guarantee, including key, payload rules, and retention window; sending a key proves nothing by itself.
28
+ - **Preserve ambiguity.** A timeout, cancellation, or lost response after dispatch can hide a completed side effect. Keep that uncertainty across restart. Stopping the caller is not rollback; do not silently create a fresh operation from an uncertain attempt.
29
+ - **Reconcile before replay.** Use an authoritative receipt or lookup matching the operation and payload. One verified result can confirm completion; conflicting or multiple matches need resolution. An empty stale, partial, or eventually consistent lookup proves neither absence nor permission to replay. Retry only under the verified deduplication contract or evidence that repeating the write is safe.
30
+ - **Keep a resolution owner.** Leave unresolved attempts visible with safe error context, a next action, and a known owner or explicit ownership gap. Manual corrective writes still need authority. Do not invent completion to clear a queue.
31
+ - **Exercise the failure boundary.** Test a committed write with a lost response, cancellation or restart before success is recorded, and stale lookup or concurrent replay where applicable. Record what was exercised and what remains unproven.
32
+
33
+ ## Prove the result and recovery
34
+
35
+ Exercise a permitted success case and relevant failures, such as denied access, malformed data, rate limits, timeout, duplicate delivery, or partial completion. Trace correlation IDs or other safe evidence on both sides. Check cleanup and recovery for test side effects. A mock proves client behavior only; missing live access is a verification gap.
36
+
37
+ Use [verification](verification.md) for receipts and [review](review.md) for security or data-contract changes. Deploy through [ship](ship.md) only when authorized.
38
+
39
+ *Fictional example:* Northstar's ERP accepts an order but the response is lost. An empty, delayed search result does not justify resubmission. Keep the attempt unresolved until an authoritative receipt or verified deduplication contract supports the next action.
40
+
41
+ ## Completion
42
+
43
+ Return the contract, changed paths, environment, evidence at each tested layer, and remaining dependencies with owners when known. Completion requires the agreed end-to-end result or an explicitly agreed narrower scope; never silently replace live acceptance with a stub. In engagement mode, update existing terrain/delivery records with confirmed facts under their write rules; otherwise return the receipt directly.
@@ -1,18 +1,31 @@
1
1
  # qa - Exercise the changed journey
2
2
 
3
- **Enter when:** a feature or fix needs behavioral verification through its real interface, especially UI, API, and multi-step workflows.
3
+ Verify what the customer can do through the real interface, including the state the journey leaves behind.
4
4
 
5
- Use [task context](task-context.md) and the customer's existing browser, API, fixtures, and test tooling. `.fde/` is not a prerequisite. Respect the permitted environment and authority for every side effect.
5
+ **Use when:** a feature or fix needs behavioral verification, especially a UI, API, or multi-step workflow.
6
6
 
7
- ## Method
7
+ Follow [task context](task-context.md) and use existing browser, API, fixture, and test tooling. `.fde/` is optional. Every side effect must stay within the permitted environment and authority.
8
8
 
9
- 1. Identify the changed journey, user roles, acceptance checks, and risk-bearing neighboring paths. Record the revision and environment. Use synthetic or sanitized fixtures with understood cleanup; do not borrow production data without permission.
10
- 2. Run the normal journey from its real entry point through the expected result. Verify persisted or downstream state when the requirement includes it; a success toast alone does not prove a write succeeded.
11
- 3. Select negative and boundary cases from the change: invalid input, empty/loading/error states, refresh/back navigation, retries, duplicates, permissions, or interrupted work. For UI changes, inspect relevant viewport sizes, keyboard access, focus, labels, and errors. Use a real browser for the affected journey.
12
- 4. Inspect relevant console and network evidence. Distinguish a UI defect from a failed API or unavailable environment. Retain only privacy-safe screenshots and logs. Do not claim visual verification from code inspection or a generated screenshot that was not viewed.
13
- 5. Report failures with steps, expected/actual result, revision/environment, evidence, and impact. If repair is authorized, use [debug](debug.md), then rerun the failed journey and affected neighbors. Keep unrelated findings separate from the change.
14
- 6. Produce a [verification receipt](verification.md). State which roles, devices, environments, or data conditions remain untested. Do not weaken acceptance checks to make the run pass.
9
+ ## Choose the journey and conditions
15
10
 
16
- ## Deliverable and acceptance
11
+ Identify the changed journey, user roles, acceptance checks, and neighboring paths at risk. Record revision and environment. Use synthetic or sanitized fixtures with understood cleanup; do not borrow production data without permission.
17
12
 
18
- Return checked journeys and observed results, reproducible defects, limitations, and remaining blockers. Done means the agreed behavioral checks passed under the stated conditions. A browser smoke test does not establish load capacity, security assurance, accessibility conformance, deployment, or customer acceptance by itself. When coordinated, append the evidence to the existing delivery record; standalone QA can return it directly.
13
+ ## Exercise the real interface
14
+
15
+ Run the normal journey from entry point to expected result. Verify persisted or downstream state when required; a success toast alone does not prove a write succeeded.
16
+
17
+ Choose negative and boundary cases relevant to the change: invalid input, empty/loading/error states, refresh/back navigation, retries, duplicates, permissions, or interruptions. For UI changes, use a real browser and inspect relevant viewport sizes, keyboard access, focus, labels, and errors.
18
+
19
+ Inspect relevant console and network evidence to distinguish UI defects from API failures or an unavailable environment. Keep only privacy-safe screenshots and logs. Code inspection or an unviewed generated screenshot does not establish visual verification.
20
+
21
+ ## Resolve findings and repeat affected checks
22
+
23
+ Report defects with reproduction steps, expected and actual results, revision/environment, evidence, and impact. If repair is authorized, use [debug](debug.md), then rerun the failed journey and affected neighbors. Keep unrelated findings separate. Do not weaken acceptance to make a run pass.
24
+
25
+ *Fictional example:* Northstar's operator sees “Import complete,” but refreshing shows no new records. Check the downstream state and network response before reporting success or deciding whether the defect is in the page or the import service.
26
+
27
+ ## Completion
28
+
29
+ Return a [verification receipt](verification.md) with checked journeys, observed results, reproducible defects, limitations, and blockers. Name roles, devices, environments, and data conditions that remain untested. Completion means agreed behavioral checks passed under the stated conditions.
30
+
31
+ A browser smoke test alone does not establish load capacity, security assurance, accessibility conformance, deployment, or customer acceptance. When coordinated, append evidence to the existing delivery record under its write rules; standalone QA can return the receipt directly.
@@ -8,7 +8,7 @@ Start from [task context](task-context.md). Standalone work uses supplied permit
8
8
 
9
9
  Identify the exact outcome, acceptance check, scope, affected users/systems, target environment, recovery mechanism, and who or what is authorized to accept and release it. Reuse confirmed authority and checks for routine work. Do not invent missing signers, permissions, measurements, or acceptance.
10
10
 
11
- For an initialized engagement, run `fde doctor --ready` before a new delivery plan or material scope change. Missing binary success or a named customer-side signer blocks that planning progression until resolved. A passing doctor validates record structure, not connectivity, release readiness, or customer acceptance. Standalone work evaluates the supplied contract directly.
11
+ For an initialized engagement, run `fde doctor --ready` before a new delivery plan or material scope change. Missing acceptance criteria or a named customer-side signer blocks the affected implementation or release commitment, not a provisional plan or independent preparation. A passing doctor validates record structure, not connectivity, release readiness, or customer acceptance. Standalone work evaluates the supplied contract directly.
12
12
 
13
13
  A customer delivery checkpoint must let the agreed decision-maker replay and reject the acceptance check through an interface they operate. Prefer their staging; otherwise use an agreed representative environment and disclose its owner and limitations. Local green proves only the local run. Routine fixes may share an agreed checkpoint; no fixed number of changes forces a ceremony.
14
14
 
@@ -43,7 +43,7 @@ For a coordinated engagement, also connect the release to the agreed value bucke
43
43
 
44
44
  Execute only when the requested workflow authorizes deployment to this target and the applicable gates are met. Otherwise leave a concrete release candidate, exact deployment/recovery instructions, evidence, and the remaining authorization for review. A permission to implement or test is not permission to publish.
45
45
 
46
- Use the customer's established pipeline and rollout mechanism. Select canary, staged exposure, blue/green, or direct rollout according to actual risk and platform capabilities; do not impose a universal cohort sequence. Define advance/abort thresholds and observation window before starting. If another operator must execute, record their handoff and report deployment pending until there is evidence it happened.
46
+ Use the customer's established pipeline and rollout mechanism. Select canary, staged exposure, blue/green, or direct rollout according to actual risk and platform capabilities; do not impose a universal cohort sequence. Define advance/abort thresholds and observation window before starting. Capture an applicable pre-rollout operating baseline with its source, environment, load/cohort and window when assessing change. Compare like conditions; keep absolute safety limits even when the baseline is poor. If no comparable baseline exists, name the gap and measurement plan: improvement is unproven, while release depends on the agreed acceptance and safety evidence, not an invented universal baseline gate. If another operator must execute, record their handoff and report deployment pending until there is evidence it happened.
47
47
 
48
48
  During rollout inspect health, errors, key user behavior, and side-effect integrity. Halt expansion on breached thresholds or critical harm and apply authorized containment/recovery. Do not continue merely because the deploy command exited successfully.
49
49
 
@@ -6,7 +6,7 @@ Use [task context](task-context.md). This method returns evidence directly or wr
6
6
 
7
7
  ## Method
8
8
 
9
- 1. Translate each claim into the observation that would support or reject it. Reuse agreed acceptance criteria and required repository checks. Select focused checks for changed behavior before broadening to release requirements.
9
+ 1. Translate each claim into the observation that would support or reject it. Cite the agreed acceptance criteria and required repository checks. Compare the checks being run with that agreement; flag any weakened threshold or removed requirement without an attributed approval from the appropriate decision-maker. Select focused checks for changed behavior before broadening to release requirements.
10
10
  2. Identify the actual repository commands, fixtures, runtime, and environment. Read command behavior before executing it, especially when it can write externally. Use authorized environments and avoid leaking secrets through logs or diagnostic commands.
11
11
  3. Run the checks and inspect results, including exit status and relevant output. A running job, test discovery, a mocked response, and a successful real request are different evidence. Record asynchronous completion before claiming success. Describe a command as executed only when its actual invocation and result are available; an inferred result is not a run.
12
12
  4. Bind evidence to the tested revision and working tree. For uncommitted changes record the base revision plus changed paths and an available diff digest or snapshot identifier. For browser/manual checks record the steps, inputs, observed result, and inspected evidence.