@gobing-ai/spur 0.3.78 → 0.3.81
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/marketplace.json +1 -1
- package/config/config.example.yaml +29 -18
- package/config/config.global.yaml +10 -11
- package/config/pipeline-budgets.json +34 -2
- package/config/plugin-scripts.json +25 -0
- package/config/rules/boundary/config-loading-ownership.yaml +0 -3
- package/config/rules/boundary/dao-boundary.yaml +4 -17
- package/config/rules/boundary/planning-folder-hardcode.yaml +0 -1
- package/config/rules/boundary/sp-no-vendor-refs.yaml +3 -2
- package/config/rules/boundary/sp-runtime-path.yaml +3 -14
- package/config/rules/quality/coverage-gate.yaml +3 -14
- package/config/rules/quality/tsdoc-exports.yaml +4 -7
- package/config/rules/strict/http-boundaries.yaml +5 -8
- package/config/rules/strict/runtime-boundaries.yaml +1 -5
- package/config/rules/structure/protected-files.yaml +9 -3
- package/config/rules/structure/test-focus-skip.yaml +0 -2
- package/config/rules/structure/test-location.yaml +0 -5
- package/config/rules/surface/check-cli-surface.yaml +3 -2
- package/config/rules/typescript/bun-tooling.yaml +5 -7
- package/config/rules/typescript/guarded-happy-dom-register.yaml +0 -2
- package/config/rules/typescript/happy-dom-teardown.yaml +0 -2
- package/config/rules/typescript/no-biome-suppressions.yaml +0 -2
- package/config/rules/typescript/no-debugger.yaml +0 -2
- package/config/rules/typescript/no-eslint-suppressions.yaml +0 -4
- package/config/rules/typescript/no-leaky-module-mocks.yaml +6 -13
- package/config/rules/typescript/no-module-scope-import-calls.yaml +0 -2
- package/config/rules/typescript/no-syscall-emulation-in-boundary-mock.yaml +0 -3
- package/config/rules/typescript/no-unmocked-module-eval-side-effects.yaml +0 -3
- package/config/rules/typescript/output-boundaries.yaml +0 -3
- package/config/rules/typescript/prefer-accessible-role-for-button-queries.yaml +0 -3
- package/config/rules/ui/ui-import-boundary.yaml +1 -5
- package/config/templates/AGENTS.md +26 -23
- package/config/templates/docs/00_ADR.md +13 -23
- package/config/templates/docs/01_PRD.md +5 -2
- package/config/templates/docs/02_ROADMAP.md +9 -13
- package/config/templates/docs/03_ARCHITECTURE.md +2 -2
- package/config/templates/docs/04_DESIGN.md +12 -31
- package/config/templates/docs/05_FEATURES.md +6 -18
- package/config/templates/docs/99_PROJECT_CONSTITUTION.md +162 -394
- package/config/transition-shims.json +7 -7
- package/config/workflows/basic.yaml +4 -0
- package/config/workflows/docs-pipeline.yaml +13 -14
- package/config/workflows/feature-dev.yaml +20 -65
- package/config/workflows/history-anatomy.yaml +22 -1
- package/config/workflows/idea-pipeline.yaml +53 -97
- package/config/workflows/pr-review.yaml +21 -33
- package/config/workflows/task-pipeline.yaml +87 -330
- package/config/workflows/wayfinder-resolution.yaml +12 -26
- package/config/workflows/wrapup-pipeline.yaml +48 -189
- package/package.json +9 -9
- package/plugins/sp/README.md +22 -8
- package/plugins/sp/agents/expert-spur.md +41 -19
- package/plugins/sp/agents/super-reviewer.md +43 -8
- package/plugins/sp/lib/idea-handoff.generated.d.mts +17 -0
- package/plugins/sp/lib/idea-handoff.generated.mjs +1301 -0
- package/plugins/sp/plugin.json +1 -1
- package/plugins/sp/scripts/feature-dev-precheck.mjs +146 -0
- package/plugins/sp/scripts/feature-dev-precheck.ts +238 -0
- package/plugins/sp/scripts/idea-handoff.mjs +27 -0
- package/plugins/sp/scripts/idea-handoff.ts +44 -0
- package/plugins/sp/scripts/quality-gate.mjs +165 -0
- package/plugins/sp/scripts/quality-gate.ts +217 -0
- package/plugins/sp/scripts/verify-answer-lint.ts +21 -3
- package/plugins/sp/scripts/workflow-step-profile.mjs +319 -0
- package/plugins/sp/scripts/workflow-step-profile.ts +456 -0
- package/plugins/sp/scripts/wrapup-steps.mjs +350 -0
- package/plugins/sp/scripts/wrapup-steps.ts +466 -0
- package/plugins/sp/skills/conflict-finding/SKILL.md +6 -0
- package/plugins/sp/skills/daily-summary/SKILL.md +1 -1
- package/plugins/sp/skills/doc-evolve/SKILL.md +26 -40
- package/plugins/sp/skills/doc-evolve/references/operations.md +17 -30
- package/plugins/sp/skills/parallel-execution/references/dispatch-surface.md +1 -1
- package/plugins/sp/skills/spec-decomposition/references/decomposition.md +29 -0
- package/plugins/sp/skills/spur-cli/references/agent.md +56 -14
- package/plugins/sp/skills/spur-cli/references/message.md +30 -3
- package/plugins/sp/skills/spur-cli/references/projects.md +45 -1
- package/plugins/sp/skills/spur-cli/references/self.md +5 -4
- package/plugins/sp/skills/spur-cli/references/serve.md +5 -4
- package/plugins/sp/skills/spur-cli/references/tasks/verbs.md +17 -1
- package/plugins/sp/skills/spur-cli/references/tasks.md +32 -2
- package/plugins/sp/skills/spur-cli/references/team.md +21 -1
- package/plugins/sp/skills/spur-cli/references/workflows/operations.md +6 -3
- package/plugins/sp/skills/spur-cli/references/workflows/workflow-fit-and-tuning.md +57 -18
- package/plugins/sp/skills/spur-composer/SKILL.md +145 -0
- package/plugins/sp/skills/spur-dev/references/ac-style-guide.md +14 -0
- package/plugins/sp/skills/spur-dev/references/cross-cutting.md +3 -3
- package/plugins/sp/skills/spur-dev/references/done-housekeeping.md +12 -0
- package/plugins/sp/skills/spur-dev/references/inline-pipeline-driver.md +46 -4
- package/plugins/sp/skills/spur-dev/references/planning-workflow.md +24 -0
- package/plugins/sp/skills/spur-doctor/SKILL.md +138 -0
- package/plugins/sp/skills/taste-refactoring-api/README.md +43 -0
- package/plugins/sp/skills/taste-refactoring-api/SKILL.md +334 -0
- package/plugins/sp/skills/taste-refactoring-api/checklists/daily-api-review.md +71 -0
- package/plugins/sp/skills/taste-refactoring-api/examples/refactor-example.md +72 -0
- package/plugins/sp/skills/taste-refactoring-api/examples/review-template.md +93 -0
- package/plugins/sp/skills/taste-refactoring-api/references/api-refactoring-playbook.md +253 -0
- package/plugins/sp/skills/taste-refactoring-api/references/protocol-modes.md +79 -0
- package/plugins/sp/skills/taste-refactoring-api/references/research-basis.md +58 -0
- package/plugins/sp/skills/taste-refactoring-architect/README.md +26 -0
- package/plugins/sp/skills/taste-refactoring-architect/SKILL.md +471 -0
- package/plugins/sp/skills/taste-refactoring-architect/checklists/daily-architecture-review.md +48 -0
- package/plugins/sp/skills/taste-refactoring-architect/examples/refactor-example.md +55 -0
- package/plugins/sp/skills/taste-refactoring-architect/examples/review-template.md +51 -0
- package/plugins/sp/skills/taste-refactoring-architect/references/architecture-refactoring-playbook.md +173 -0
- package/plugins/sp/skills/taste-refactoring-architect/references/research-basis.md +28 -0
- package/plugins/sp/skills/taste-refactoring-tests/README.md +28 -0
- package/plugins/sp/skills/taste-refactoring-tests/SKILL.md +482 -0
- package/plugins/sp/skills/taste-refactoring-tests/checklists/daily-test-review.md +39 -0
- package/plugins/sp/skills/taste-refactoring-tests/examples/refactor-example.md +85 -0
- package/plugins/sp/skills/taste-refactoring-tests/examples/review-template.md +59 -0
- package/plugins/sp/skills/taste-refactoring-tests/references/research-basis.md +47 -0
- package/plugins/sp/skills/taste-refactoring-tests/references/test-refactoring-playbook.md +222 -0
- package/plugins/sp/skills/taste-refactoring-ui/README.md +12 -0
- package/plugins/sp/skills/taste-refactoring-ui/SKILL.md +290 -0
- package/plugins/sp/skills/taste-refactoring-ui/checklists/daily-ui-review.md +72 -0
- package/plugins/sp/skills/taste-refactoring-ui/examples/review-template.md +51 -0
- package/plugins/sp/skills/taste-refactoring-ui/references/refactoring-ui-playbook.md +170 -0
- package/plugins/sp/skills/wayfinder/SKILL.md +2 -2
- package/plugins/sp/skills/wayfinder/references/pipeline-resolution.md +30 -0
- package/schemas/spur-config.schema.json +49 -0
- package/spur.js +46936 -44198
- package/web/_astro/{BoardApp.CHQ1lycZ.js → BoardApp.B1U26g3I.js} +97 -95
- package/web/_astro/BoardApp.Csgyg-lS.js +1 -0
- package/web/_astro/{TaskDetail.GKfQJ60c.js → TaskDetail.DwPqpq7v.js} +1 -1
- package/web/_astro/{arc.DWEtA3Tx.js → arc.CweZEjN2.js} +1 -1
- package/web/_astro/{architectureDiagram-3BPJPVTR.DB42oWmP.js → architectureDiagram-3BPJPVTR.D89pbDuv.js} +1 -1
- package/web/_astro/{blockDiagram-GPEHLZMM.rhv-zNQV.js → blockDiagram-GPEHLZMM.BOuTeEpX.js} +1 -1
- package/web/_astro/{c4Diagram-AAUBKEIU.Ci4-4VvY.js → c4Diagram-AAUBKEIU.CASbkWZF.js} +1 -1
- package/web/_astro/channel.Cx6sXxhq.js +1 -0
- package/web/_astro/{chunk-2J33WTMH.Cc9veUgf.js → chunk-2J33WTMH.BKQYtOvY.js} +1 -1
- package/web/_astro/{chunk-4BX2VUAB.Bec9c4eI.js → chunk-4BX2VUAB.9sHLdMtG.js} +1 -1
- package/web/_astro/{chunk-55IACEB6.DoV8S1iB.js → chunk-55IACEB6.wOLXWlPs.js} +1 -1
- package/web/_astro/{chunk-727SXJPM.DwR-Qlyj.js → chunk-727SXJPM.DovFbwg3.js} +1 -1
- package/web/_astro/{chunk-AQP2D5EJ.ND_a81WY.js → chunk-AQP2D5EJ.B1Weod1X.js} +1 -1
- package/web/_astro/{chunk-FMBD7UC4.Wv_jwG48.js → chunk-FMBD7UC4.TEMS04st.js} +1 -1
- package/web/_astro/{chunk-ND2GUHAM.CXKXCMmp.js → chunk-ND2GUHAM.Cp8VT1wQ.js} +1 -1
- package/web/_astro/{chunk-QZHKN3VN.nkaoNYQq.js → chunk-QZHKN3VN.BzATdEcP.js} +1 -1
- package/web/_astro/{classDiagram-4FO5ZUOK.cMQcVlQu.js → classDiagram-4FO5ZUOK.C9BOCfAO.js} +1 -1
- package/web/_astro/{classDiagram-v2-Q7XG4LA2.cMQcVlQu.js → classDiagram-v2-Q7XG4LA2.C9BOCfAO.js} +1 -1
- package/web/_astro/{cose-bilkent-S5V4N54A.OaDJ7Mr2.js → cose-bilkent-S5V4N54A.DUnr4UAw.js} +1 -1
- package/web/_astro/{cynefin-OW5HDTMX.Chi8IphF.js → cynefin-OW5HDTMX.rYq5uM3D.js} +1 -1
- package/web/_astro/{cytoscape.esm.DzSz-X2X.js → cytoscape.esm.BB4DxJjf.js} +1 -1
- package/web/_astro/{dagre-BM42HDAG.CzK2t_Fp.js → dagre-BM42HDAG.CWeNKe3I.js} +1 -1
- package/web/_astro/{diagram-2AECGRRQ.DRvxlVS7.js → diagram-2AECGRRQ.DCkfls10.js} +1 -1
- package/web/_astro/{diagram-5GNKFQAL.CnYvNdwA.js → diagram-5GNKFQAL.D5U4JCka.js} +1 -1
- package/web/_astro/{diagram-KO2AKTUF.CpLpMw5R.js → diagram-KO2AKTUF.BZJgqaqG.js} +1 -1
- package/web/_astro/{diagram-LMA3HP47.JTb78qUA.js → diagram-LMA3HP47.DoMeHvPR.js} +1 -1
- package/web/_astro/{diagram-OG6HWLK6.Bk-1jDIb.js → diagram-OG6HWLK6.B50qwwWX.js} +1 -1
- package/web/_astro/{erDiagram-TEJ5UH35.D8hN9GZq.js → erDiagram-TEJ5UH35.DdGPG6LK.js} +1 -1
- package/web/_astro/{flowDiagram-I6XJVG4X.-6zQr6m5.js → flowDiagram-I6XJVG4X.QP2MJ12u.js} +1 -1
- package/web/_astro/{ganttDiagram-6RSMTGT7.DboLQ9ca.js → ganttDiagram-6RSMTGT7.BI6LgKSy.js} +1 -1
- package/web/_astro/{gitGraphDiagram-PVQCEYII.4tYvJKGR.js → gitGraphDiagram-PVQCEYII.npPZiC2G.js} +1 -1
- package/web/_astro/index.DayyIngm.css +1 -0
- package/web/_astro/{infoDiagram-5YYISTIA.Bd9rXpsB.js → infoDiagram-5YYISTIA.DCJCBVbp.js} +1 -1
- package/web/_astro/{ishikawaDiagram-YF4QCWOH.CvMoaf67.js → ishikawaDiagram-YF4QCWOH.BMLV-3I1.js} +1 -1
- package/web/_astro/{journeyDiagram-JHISSGLW.Ccy1CA7y.js → journeyDiagram-JHISSGLW.LE58crde.js} +1 -1
- package/web/_astro/{kanban-definition-UN3LZRKU.0MaMqHNS.js → kanban-definition-UN3LZRKU.BPbz8rH9.js} +1 -1
- package/web/_astro/{linear.CHXgcIbN.js → linear.DhZaBtYh.js} +1 -1
- package/web/_astro/{mermaid.core.Ca-kcelG.js → mermaid.core.BD5-jXum.js} +6 -6
- package/web/_astro/{mindmap-definition-RKZ34NQL.BUIDlHa0.js → mindmap-definition-RKZ34NQL.MTJyrQ65.js} +1 -1
- package/web/_astro/ordinal.BYWQX77i.js +1 -0
- package/web/_astro/{pieDiagram-4H26LBE5.2dX3CU1s.js → pieDiagram-4H26LBE5.BrDhDvIS.js} +1 -1
- package/web/_astro/{quadrantDiagram-W4KKPZXB.B3LBlRiv.js → quadrantDiagram-W4KKPZXB.71d73_5N.js} +1 -1
- package/web/_astro/{requirementDiagram-4Y6WPE33.X12I2uNx.js → requirementDiagram-4Y6WPE33.Bga6UF-z.js} +1 -1
- package/web/_astro/{sankeyDiagram-5OEKKPKP.BXohIHqx.js → sankeyDiagram-5OEKKPKP.BnHs4K82.js} +1 -1
- package/web/_astro/{sequenceDiagram-3UESZ5HK.C37ZIUzg.js → sequenceDiagram-3UESZ5HK.DsfY2gnj.js} +1 -1
- package/web/_astro/{stateDiagram-AJRCARHV.BRgz317z.js → stateDiagram-AJRCARHV.DvsTSc9a.js} +1 -1
- package/web/_astro/{stateDiagram-v2-BHNVJYJU.7VYSXN9-.js → stateDiagram-v2-BHNVJYJU.DxzzmHUR.js} +1 -1
- package/web/_astro/{timeline-definition-PNZ67QCA.BVNz_HiN.js → timeline-definition-PNZ67QCA.4ZuQmOTt.js} +1 -1
- package/web/_astro/{vennDiagram-CIIHVFJN.CHVDkPX4.js → vennDiagram-CIIHVFJN.Ck5Q86SG.js} +1 -1
- package/web/_astro/{wardleyDiagram-YWT4CUSO.EQQ_qT9v.js → wardleyDiagram-YWT4CUSO.BK7k2hXr.js} +1 -1
- package/web/_astro/{xychartDiagram-2RQKCTM6.DrAT9WoP.js → xychartDiagram-2RQKCTM6.DfCrgauK.js} +1 -1
- package/web/index.html +2 -2
- package/web/_astro/BoardApp.DV9kx0wo.js +0 -1
- package/web/_astro/channel.BAI6xLeV.js +0 -1
- package/web/_astro/index.Dcr_8fiK.css +0 -1
- package/web/_astro/ordinal.DBvzRdQf.js +0 -1
|
@@ -0,0 +1,482 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: taste-refactoring-tests
|
|
3
|
+
description: Refactor unit tests for failure sensitivity, deterministic confidence, and regression protection without mock-heavy brittle tests or vanity coverage. Backs test suite quality audits.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# taste-refactoring-tests
|
|
7
|
+
|
|
8
|
+
## Purpose
|
|
9
|
+
Use this skill to refactor existing unit tests so they provide **real delivery confidence**, not the illusion of safety created by a permanently green test suite.
|
|
10
|
+
|
|
11
|
+
The objective is not “more tests”, “higher coverage”, or “make CI green”. The objective is **failure-sensitive tests**: tests that are cheap enough to run continuously, deterministic enough to trust, precise enough to diagnose failures, and strong enough to fail when relevant behavior regresses.
|
|
12
|
+
|
|
13
|
+
This skill is deliberately skeptical of tests that always pass, mirror implementation details, assert trivialities, overuse mocks, or inflate coverage without protecting user-visible behavior and important engineering invariants.
|
|
14
|
+
|
|
15
|
+
## Core outcome
|
|
16
|
+
Given existing unit tests, production code, coverage reports, mutation reports, CI history, bug/incident history, or a prose description, produce:
|
|
17
|
+
|
|
18
|
+
1. A concise model of what the tests currently claim to protect.
|
|
19
|
+
2. A map of weak, redundant, brittle, misleading, and missing protections.
|
|
20
|
+
3. A `Confidence Contract`: the behaviors and delivery qualities that tests must meaningfully defend.
|
|
21
|
+
4. A prioritized refactoring plan using the intervention ladder below.
|
|
22
|
+
5. Safer replacement tests with stronger assertions and lower coupling to implementation.
|
|
23
|
+
6. Validation evidence showing the refactored suite can actually detect plausible regressions.
|
|
24
|
+
|
|
25
|
+
## Prime directive
|
|
26
|
+
**A test earns its place by detecting meaningful regressions.**
|
|
27
|
+
|
|
28
|
+
A passing test is not evidence of quality unless there is a credible change that would make it fail.
|
|
29
|
+
|
|
30
|
+
Ask of every test:
|
|
31
|
+
- What behavior or invariant does this protect?
|
|
32
|
+
- What plausible regression would make it fail?
|
|
33
|
+
- Would a broken implementation still pass it?
|
|
34
|
+
- Does it verify an observable outcome or merely implementation choreography?
|
|
35
|
+
- Is the assertion strong enough to distinguish correct from merely non-crashing behavior?
|
|
36
|
+
- Is the test deterministic and repeatable?
|
|
37
|
+
- Does its maintenance cost exceed the risk it protects?
|
|
38
|
+
|
|
39
|
+
## Non-negotiable principles
|
|
40
|
+
|
|
41
|
+
### 1. Optimize for signal, not green
|
|
42
|
+
Never weaken assertions, catch failures, add retries, overmock dependencies, or skip tests simply to stabilize CI.
|
|
43
|
+
|
|
44
|
+
A red test can be valuable information. A green test that cannot detect regressions is dangerous.
|
|
45
|
+
|
|
46
|
+
### 2. Coverage is a map, not proof
|
|
47
|
+
Line or branch coverage can reveal unexercised code, but high coverage does not prove strong tests. Prefer evidence that tests are sensitive to faults: meaningful assertions, boundary cases, negative cases, mutation testing, bug reproduction, and behavior-oriented review.
|
|
48
|
+
|
|
49
|
+
### 3. Test observable behavior
|
|
50
|
+
Prefer testing externally observable behavior at the unit boundary rather than private methods, internal call order, exact intermediate objects, incidental SQL, or framework choreography.
|
|
51
|
+
|
|
52
|
+
Implementation details may be asserted only when they are themselves contractual: e.g. security policy, protocol requirements, transactional semantics, side-effect cardinality, or performance-critical behavior.
|
|
53
|
+
|
|
54
|
+
### 4. Assertions must discriminate
|
|
55
|
+
Weak assertions create false confidence. Avoid tests that only assert:
|
|
56
|
+
- no exception was thrown,
|
|
57
|
+
- object is non-null,
|
|
58
|
+
- collection is non-empty,
|
|
59
|
+
- HTTP status is success while response semantics are unchecked,
|
|
60
|
+
- a mocked method was called without checking the outcome,
|
|
61
|
+
- a snapshot changed without validating meaning.
|
|
62
|
+
|
|
63
|
+
Assert the smallest set of outcomes that proves the behavior and important invariants.
|
|
64
|
+
|
|
65
|
+
### 5. Mocks are a cost
|
|
66
|
+
Mock boundaries that are expensive, nondeterministic, or outside the unit. Do not mock the class under test, simple value objects, domain logic, or collaborators merely to make tests easy.
|
|
67
|
+
|
|
68
|
+
Prefer fakes/stubs for stateful collaborators when interaction assertions are not the actual contract.
|
|
69
|
+
|
|
70
|
+
### 6. Determinism is mandatory
|
|
71
|
+
Tests must control time, randomness, concurrency, locale, environment, global state, and external I/O when those can affect outcomes.
|
|
72
|
+
|
|
73
|
+
Retries are not a substitute for determinism.
|
|
74
|
+
|
|
75
|
+
### 7. Bugs deserve permanent regression tests
|
|
76
|
+
When a bug escaped, first capture the smallest reproducing test that fails before the fix and passes after it. Then strengthen nearby tests if the failure reveals a broader gap.
|
|
77
|
+
|
|
78
|
+
### 8. Test structure should aid diagnosis
|
|
79
|
+
One test may contain multiple related assertions if they express one behavior. Split tests when failures would represent different behaviors, different setup, or different root causes.
|
|
80
|
+
|
|
81
|
+
### 9. Prefer risk-based depth
|
|
82
|
+
Spend testing effort where failure cost, change frequency, complexity, and uncertainty are high. Do not give every getter and every critical calculation equal test weight.
|
|
83
|
+
|
|
84
|
+
### 10. Preserve useful tests
|
|
85
|
+
Refactoring is not churn. Keep tests that are behavior-focused, deterministic, appropriately scoped, readable, and capable of detecting meaningful regressions.
|
|
86
|
+
|
|
87
|
+
## The intervention ladder
|
|
88
|
+
Classify each test or test cluster into exactly one primary action. Use the least disruptive action that materially improves confidence.
|
|
89
|
+
|
|
90
|
+
### T0 — KEEP
|
|
91
|
+
The test is valuable as-is.
|
|
92
|
+
|
|
93
|
+
Use when it:
|
|
94
|
+
- protects a meaningful behavior or invariant,
|
|
95
|
+
- has discriminating assertions,
|
|
96
|
+
- is deterministic,
|
|
97
|
+
- is not tightly coupled to incidental implementation,
|
|
98
|
+
- has reasonable maintenance cost.
|
|
99
|
+
|
|
100
|
+
Output: `KEEP — <behavior protected and why the test is credible>`
|
|
101
|
+
|
|
102
|
+
### T1 — DIRECT REMOVE
|
|
103
|
+
Delete tests that provide no meaningful protection and have low uncertainty.
|
|
104
|
+
|
|
105
|
+
Typical candidates:
|
|
106
|
+
- tests of language/framework guarantees,
|
|
107
|
+
- tests that only verify getters/setters or constant assignments,
|
|
108
|
+
- duplicates with no extra behavior coverage,
|
|
109
|
+
- tests that assert mocks return what the test configured them to return,
|
|
110
|
+
- obsolete tests for removed behavior,
|
|
111
|
+
- tautological tests whose assertion repeats setup,
|
|
112
|
+
- unreachable/skipped tests with no remaining purpose.
|
|
113
|
+
|
|
114
|
+
Required check before removal:
|
|
115
|
+
- no unique contract is being protected,
|
|
116
|
+
- no known escaped bug depends on it,
|
|
117
|
+
- no subtle side effect or invariant is hidden in the setup.
|
|
118
|
+
|
|
119
|
+
### T2 — SUGGEST REMOVE
|
|
120
|
+
The test appears low-value or harmful, but removal needs evidence.
|
|
121
|
+
|
|
122
|
+
Use when:
|
|
123
|
+
- it is brittle but may encode undocumented behavior,
|
|
124
|
+
- it is redundant but ownership is unclear,
|
|
125
|
+
- it depends on legacy behavior with unknown consumers,
|
|
126
|
+
- it is flaky and nobody knows whether the flake exposes a real race.
|
|
127
|
+
|
|
128
|
+
Output must state what evidence would justify deletion.
|
|
129
|
+
|
|
130
|
+
### T3 — STRENGTHEN
|
|
131
|
+
Keep the scenario, improve failure sensitivity.
|
|
132
|
+
|
|
133
|
+
Typical moves:
|
|
134
|
+
- replace weak assertions with semantic assertions,
|
|
135
|
+
- assert error type/code/message contract where meaningful,
|
|
136
|
+
- assert state transitions, side effects, persistence, or emitted events,
|
|
137
|
+
- add boundary and negative cases,
|
|
138
|
+
- verify idempotency or invariant preservation,
|
|
139
|
+
- introduce mutation-sensitive assertions,
|
|
140
|
+
- replace broad snapshots with targeted checks.
|
|
141
|
+
|
|
142
|
+
Use when the scenario matters but the current test could pass under a broken implementation.
|
|
143
|
+
|
|
144
|
+
### T4 — SIMPLIFY / DE-BRITTLE
|
|
145
|
+
Keep the behavior coverage while reducing implementation coupling or maintenance cost.
|
|
146
|
+
|
|
147
|
+
Typical moves:
|
|
148
|
+
- test through the public unit boundary,
|
|
149
|
+
- remove private-method tests,
|
|
150
|
+
- replace interaction-heavy mocks with fakes or state assertions,
|
|
151
|
+
- use builders/factories for irrelevant setup,
|
|
152
|
+
- extract stable fixtures without hiding important inputs,
|
|
153
|
+
- remove accidental ordering dependencies,
|
|
154
|
+
- control clocks/randomness/environment,
|
|
155
|
+
- collapse duplicate parameterized cases,
|
|
156
|
+
- use table-driven tests for coherent boundary sets.
|
|
157
|
+
|
|
158
|
+
### T5 — ADD MISSING PROTECTION
|
|
159
|
+
Existing tests leave a material behavior or risk uncovered.
|
|
160
|
+
|
|
161
|
+
Add tests for:
|
|
162
|
+
- core business rules,
|
|
163
|
+
- error and rejection paths,
|
|
164
|
+
- boundary values,
|
|
165
|
+
- previously escaped bugs,
|
|
166
|
+
- security/authorization decisions,
|
|
167
|
+
- serialization/parsing rules,
|
|
168
|
+
- state transitions,
|
|
169
|
+
- concurrency/idempotency invariants where unit-testable,
|
|
170
|
+
- important null/empty/overflow/rounding/time-zone cases,
|
|
171
|
+
- contract changes likely to break consumers.
|
|
172
|
+
|
|
173
|
+
Do not add tests merely to raise a coverage percentage.
|
|
174
|
+
|
|
175
|
+
### T6 — RE-DESIGN TEST STRATEGY
|
|
176
|
+
The unit tests are structurally incapable of providing trustworthy confidence.
|
|
177
|
+
|
|
178
|
+
Triggers:
|
|
179
|
+
- massive mock choreography,
|
|
180
|
+
- tests mirror private structure so closely that every refactor breaks them,
|
|
181
|
+
- important behavior can only be tested through dozens of collaborators,
|
|
182
|
+
- test fixtures are harder to understand than production code,
|
|
183
|
+
- unit tests duplicate integration tests without catching unique faults,
|
|
184
|
+
- suites are consistently green despite recurring production regressions,
|
|
185
|
+
- mutation testing shows large survivor clusters in critical logic,
|
|
186
|
+
- architecture prevents isolation of important domain behavior.
|
|
187
|
+
|
|
188
|
+
A redesign may include:
|
|
189
|
+
- moving pure decision logic into testable units,
|
|
190
|
+
- replacing orchestration tests with focused domain tests plus contract/integration tests,
|
|
191
|
+
- changing dependency boundaries,
|
|
192
|
+
- introducing fakes for stable protocols,
|
|
193
|
+
- using property-based or state-machine testing for rule-heavy domains,
|
|
194
|
+
- redistributing coverage across unit, integration, contract, and end-to-end levels.
|
|
195
|
+
|
|
196
|
+
Do not force all confidence into unit tests when another test level is more appropriate.
|
|
197
|
+
|
|
198
|
+
### T7 — QUARANTINE / INVESTIGATE
|
|
199
|
+
Do not normalize uncertainty.
|
|
200
|
+
|
|
201
|
+
Use when a flaky or suspicious test cannot yet be classified safely.
|
|
202
|
+
|
|
203
|
+
Required output:
|
|
204
|
+
- suspected nondeterminism source,
|
|
205
|
+
- instrumentation or reproduction plan,
|
|
206
|
+
- owner,
|
|
207
|
+
- expiration/review date if quarantine is unavoidable,
|
|
208
|
+
- explicit warning that quarantine is temporary and not equivalent to passing.
|
|
209
|
+
|
|
210
|
+
## Confidence Contract
|
|
211
|
+
Before refactoring, identify what must remain protected.
|
|
212
|
+
|
|
213
|
+
Capture:
|
|
214
|
+
- critical business rules,
|
|
215
|
+
- public behavior and compatibility contracts,
|
|
216
|
+
- authorization/security decisions,
|
|
217
|
+
- data integrity invariants,
|
|
218
|
+
- error semantics,
|
|
219
|
+
- financial/rounding calculations,
|
|
220
|
+
- state-machine transitions,
|
|
221
|
+
- idempotency and duplicate handling,
|
|
222
|
+
- concurrency-sensitive invariants,
|
|
223
|
+
- serialization/protocol rules,
|
|
224
|
+
- incident/bug regressions,
|
|
225
|
+
- latency/resource constraints only where unit-level checks are credible.
|
|
226
|
+
|
|
227
|
+
Unknown protections are risks, not permission to delete tests.
|
|
228
|
+
|
|
229
|
+
## Test review workflow
|
|
230
|
+
|
|
231
|
+
### Phase 0 — Establish scope
|
|
232
|
+
Identify:
|
|
233
|
+
- production units under test,
|
|
234
|
+
- test framework and mocking tools,
|
|
235
|
+
- CI execution model,
|
|
236
|
+
- coverage/mutation data if available,
|
|
237
|
+
- recent escaped defects,
|
|
238
|
+
- known flaky tests,
|
|
239
|
+
- ownership and change hotspots.
|
|
240
|
+
|
|
241
|
+
### Phase 1 — Inventory tests by protected behavior
|
|
242
|
+
For each test or coherent cluster, record:
|
|
243
|
+
- behavior/invariant claimed,
|
|
244
|
+
- inputs and boundary class,
|
|
245
|
+
- observed outputs/side effects,
|
|
246
|
+
- mocks/fakes used,
|
|
247
|
+
- assertions,
|
|
248
|
+
- runtime and flake history,
|
|
249
|
+
- production code touched conceptually,
|
|
250
|
+
- bug/requirement linkage if known.
|
|
251
|
+
|
|
252
|
+
Do not organize review only by filename. Organize by behavior.
|
|
253
|
+
|
|
254
|
+
### Phase 2 — Find false-confidence smells
|
|
255
|
+
Look for:
|
|
256
|
+
- assert-nothing tests,
|
|
257
|
+
- tautological assertions,
|
|
258
|
+
- excessive `isNotNull`/`isTrue` assertions,
|
|
259
|
+
- mock-verification-only tests,
|
|
260
|
+
- tests of private methods,
|
|
261
|
+
- one-to-one mirroring of implementation classes,
|
|
262
|
+
- overspecified call order,
|
|
263
|
+
- snapshots covering large unstable structures,
|
|
264
|
+
- golden files nobody reviews,
|
|
265
|
+
- broad fixtures with irrelevant setup,
|
|
266
|
+
- time/random/environment dependence,
|
|
267
|
+
- sleeps and retries,
|
|
268
|
+
- shared mutable test state,
|
|
269
|
+
- hidden network/database access in “unit” tests,
|
|
270
|
+
- duplicate happy paths with no negative cases,
|
|
271
|
+
- coverage-driven trivial tests,
|
|
272
|
+
- mocks that make impossible states look valid,
|
|
273
|
+
- swallowed exceptions,
|
|
274
|
+
- assertions after asynchronous work without proper synchronization,
|
|
275
|
+
- skipped/disabled tests that CI still reports as green.
|
|
276
|
+
|
|
277
|
+
### Phase 3 — Challenge tests with plausible faults
|
|
278
|
+
For each important test, imagine or introduce small faults such as:
|
|
279
|
+
- invert a condition,
|
|
280
|
+
- remove validation,
|
|
281
|
+
- change a boundary comparison,
|
|
282
|
+
- return a default value,
|
|
283
|
+
- omit a side effect,
|
|
284
|
+
- swap fields,
|
|
285
|
+
- alter rounding,
|
|
286
|
+
- ignore authorization,
|
|
287
|
+
- duplicate an event,
|
|
288
|
+
- drop an error mapping.
|
|
289
|
+
|
|
290
|
+
If the test would still pass, classify it for strengthening or replacement.
|
|
291
|
+
|
|
292
|
+
Where practical, use mutation testing as automated evidence. Mutation score is diagnostic, not a vanity target.
|
|
293
|
+
|
|
294
|
+
### Phase 4 — Refactor from behavior outward
|
|
295
|
+
Preferred order:
|
|
296
|
+
1. Capture escaped bugs and missing critical behavior.
|
|
297
|
+
2. Strengthen weak assertions.
|
|
298
|
+
3. Remove tautologies and duplicates.
|
|
299
|
+
4. Reduce mock/implementation coupling.
|
|
300
|
+
5. Make nondeterministic tests deterministic.
|
|
301
|
+
6. Simplify setup and fixtures.
|
|
302
|
+
7. Re-balance tests across levels if unit tests are carrying the wrong responsibilities.
|
|
303
|
+
|
|
304
|
+
### Phase 5 — Validate the refactor
|
|
305
|
+
A successful refactor should show evidence such as:
|
|
306
|
+
- known faults are detected,
|
|
307
|
+
- representative mutants are killed,
|
|
308
|
+
- critical bug reproductions fail before fixes,
|
|
309
|
+
- tests survive internal refactors when behavior is unchanged,
|
|
310
|
+
- flaky failures decrease without assertion weakening,
|
|
311
|
+
- suite runtime remains acceptable,
|
|
312
|
+
- test names and failure messages explain the broken behavior,
|
|
313
|
+
- removed tests did not reduce meaningful protection.
|
|
314
|
+
|
|
315
|
+
## Assertion quality rules
|
|
316
|
+
|
|
317
|
+
### Prefer semantic assertions
|
|
318
|
+
Good:
|
|
319
|
+
- exact domain result,
|
|
320
|
+
- relevant object fields,
|
|
321
|
+
- precise state transition,
|
|
322
|
+
- exact error category/code,
|
|
323
|
+
- expected persisted record,
|
|
324
|
+
- expected emitted event payload,
|
|
325
|
+
- absence of forbidden side effect.
|
|
326
|
+
|
|
327
|
+
Weak unless context justifies them:
|
|
328
|
+
- `not null`,
|
|
329
|
+
- `true`,
|
|
330
|
+
- `count > 0`,
|
|
331
|
+
- `status < 400`,
|
|
332
|
+
- `mock.verify(...)` alone.
|
|
333
|
+
|
|
334
|
+
### Assert invariants, not every field
|
|
335
|
+
Do not make tests brittle by asserting irrelevant representation details. Assert all fields required to prove the behavior and no more.
|
|
336
|
+
|
|
337
|
+
### Negative assertions matter
|
|
338
|
+
For sensitive operations, verify not only what happened but what must **not** happen:
|
|
339
|
+
- unauthorized request did not mutate state,
|
|
340
|
+
- rejected payment did not emit fulfillment event,
|
|
341
|
+
- duplicate command did not create a second record,
|
|
342
|
+
- failed validation did not call persistence.
|
|
343
|
+
|
|
344
|
+
## Mocking rules
|
|
345
|
+
|
|
346
|
+
Mock when:
|
|
347
|
+
- a dependency crosses process/network boundaries,
|
|
348
|
+
- the real collaborator is nondeterministic or expensive,
|
|
349
|
+
- a failure mode must be induced deliberately,
|
|
350
|
+
- interaction itself is the contract.
|
|
351
|
+
|
|
352
|
+
Prefer a fake/stub when:
|
|
353
|
+
- state-based verification is clearer,
|
|
354
|
+
- the protocol is stable and simple,
|
|
355
|
+
- interaction count/order is not the behavior.
|
|
356
|
+
|
|
357
|
+
Avoid mocking:
|
|
358
|
+
- value objects,
|
|
359
|
+
- pure functions,
|
|
360
|
+
- the unit under test,
|
|
361
|
+
- every internal collaborator by default.
|
|
362
|
+
|
|
363
|
+
## Property-based and parameterized testing
|
|
364
|
+
Use property-based tests when the behavior is best expressed as an invariant across many inputs, such as:
|
|
365
|
+
- parsers/serializers,
|
|
366
|
+
- ordering/sorting,
|
|
367
|
+
- numeric calculations,
|
|
368
|
+
- state transitions,
|
|
369
|
+
- normalization,
|
|
370
|
+
- round trips,
|
|
371
|
+
- algebraic/domain invariants.
|
|
372
|
+
|
|
373
|
+
Use parameterized/table-driven tests for a finite, meaningful set of boundary classes. Do not hide fundamentally different behaviors in one giant data table.
|
|
374
|
+
|
|
375
|
+
## Flakiness policy
|
|
376
|
+
A flaky test is a defect in either the test, product, environment, or synchronization model.
|
|
377
|
+
|
|
378
|
+
Do not permanently solve flakiness by:
|
|
379
|
+
- retries,
|
|
380
|
+
- longer sleeps,
|
|
381
|
+
- looser assertions,
|
|
382
|
+
- disabling the test,
|
|
383
|
+
- ignoring failures.
|
|
384
|
+
|
|
385
|
+
Use retries only as short-lived diagnostic containment when the underlying issue is actively tracked.
|
|
386
|
+
|
|
387
|
+
## Test naming
|
|
388
|
+
Names should describe behavior and condition, not implementation trivia.
|
|
389
|
+
|
|
390
|
+
Prefer:
|
|
391
|
+
`rejects_transfer_when_balance_is_insufficient`
|
|
392
|
+
|
|
393
|
+
Avoid:
|
|
394
|
+
`testProcessTransfer2`
|
|
395
|
+
|
|
396
|
+
For BDD-style naming, keep the same principle: condition + behavior + outcome.
|
|
397
|
+
|
|
398
|
+
## Review severity
|
|
399
|
+
Rate findings by delivery risk:
|
|
400
|
+
|
|
401
|
+
- **P0 Critical** — suite can green-light a severe security, financial, data integrity, or availability regression.
|
|
402
|
+
- **P1 High** — important production behavior is weakly protected or known bugs can recur unnoticed.
|
|
403
|
+
- **P2 Medium** — brittleness, flakiness, or overspecification significantly slows safe delivery.
|
|
404
|
+
- **P3 Low** — readability/duplication/fixture quality issue with limited confidence impact.
|
|
405
|
+
|
|
406
|
+
Priority is based on risk reduction, not cosmetic cleanliness.
|
|
407
|
+
|
|
408
|
+
## Required output format
|
|
409
|
+
When reviewing tests, return sections in this order:
|
|
410
|
+
|
|
411
|
+
### 1. Confidence summary
|
|
412
|
+
- What the suite protects well
|
|
413
|
+
- Where confidence is misleading
|
|
414
|
+
- Biggest delivery risks
|
|
415
|
+
|
|
416
|
+
### 2. Confidence Contract
|
|
417
|
+
A compact table of required behaviors/invariants and current protection strength.
|
|
418
|
+
|
|
419
|
+
### 3. Findings
|
|
420
|
+
For each finding:
|
|
421
|
+
- `ID`
|
|
422
|
+
- `Severity`
|
|
423
|
+
- `Action` (`T0`–`T7`)
|
|
424
|
+
- `Evidence`
|
|
425
|
+
- `Why it matters`
|
|
426
|
+
- `Proposed change`
|
|
427
|
+
- `Expected regression-detection improvement`
|
|
428
|
+
|
|
429
|
+
### 4. Refactoring plan
|
|
430
|
+
Order changes by risk reduction and dependency, not file order.
|
|
431
|
+
|
|
432
|
+
### 5. Replacement test examples
|
|
433
|
+
Show representative before/after tests or pseudocode when useful.
|
|
434
|
+
|
|
435
|
+
### 6. Validation plan
|
|
436
|
+
State how to prove the refactored tests are stronger: fault injection, mutation testing, bug reproduction, flake monitoring, refactor resistance, or other evidence.
|
|
437
|
+
|
|
438
|
+
## Decision rules
|
|
439
|
+
|
|
440
|
+
### Do not remove a test only because it is ugly
|
|
441
|
+
First determine what protection it provides. Preserve the protection even if the implementation is rewritten.
|
|
442
|
+
|
|
443
|
+
### Do not preserve a test only because it has existed a long time
|
|
444
|
+
Age is not evidence of value.
|
|
445
|
+
|
|
446
|
+
### Do not add assertions mechanically
|
|
447
|
+
Every assertion should make a broken implementation more likely to fail for a meaningful reason.
|
|
448
|
+
|
|
449
|
+
### Do not chase 100% coverage
|
|
450
|
+
Prefer a smaller suite with high fault sensitivity over a larger suite that asserts trivialities.
|
|
451
|
+
|
|
452
|
+
### Do not convert every unit test into an integration test
|
|
453
|
+
Use the narrowest test level that can reliably protect the behavior.
|
|
454
|
+
|
|
455
|
+
### Do not force behavior through mocks
|
|
456
|
+
If testing requires recreating the implementation in mock expectations, reconsider the unit boundary or test level.
|
|
457
|
+
|
|
458
|
+
## Daily operating heuristic
|
|
459
|
+
When time is limited, use this sequence:
|
|
460
|
+
|
|
461
|
+
1. Find tests around the highest-risk or most-changed code.
|
|
462
|
+
2. Identify what regression each test would catch.
|
|
463
|
+
3. Remove or mark tests that catch nothing meaningful.
|
|
464
|
+
4. Strengthen weak assertions on critical behavior.
|
|
465
|
+
5. Add one missing negative/boundary case where risk is highest.
|
|
466
|
+
6. Reduce one source of brittleness or nondeterminism.
|
|
467
|
+
7. Challenge the suite with a small plausible fault.
|
|
468
|
+
|
|
469
|
+
If the fault survives, the suite is not done merely because it is green.
|
|
470
|
+
|
|
471
|
+
## Definition of done
|
|
472
|
+
A test refactor is complete only when:
|
|
473
|
+
- required behavior remains protected,
|
|
474
|
+
- misleading tests are removed or repaired,
|
|
475
|
+
- important assertions are discriminating,
|
|
476
|
+
- nondeterminism is controlled,
|
|
477
|
+
- implementation coupling is reduced where possible,
|
|
478
|
+
- critical negative/boundary paths are represented,
|
|
479
|
+
- the suite detects representative plausible faults,
|
|
480
|
+
- delivery speed is not degraded without a justified risk trade-off.
|
|
481
|
+
|
|
482
|
+
The goal is **earned confidence**: green means something because the suite has demonstrated that it can turn red when the product is wrong.
|
|
@@ -0,0 +1,39 @@
|
|
|
1
|
+
# Daily Test Refactoring Checklist
|
|
2
|
+
|
|
3
|
+
Use this during code review, test cleanup, or before shipping a risky change.
|
|
4
|
+
|
|
5
|
+
## Confidence
|
|
6
|
+
- [ ] Can I name the behavior each important test protects?
|
|
7
|
+
- [ ] Can I name a plausible regression that would make it fail?
|
|
8
|
+
- [ ] Are critical rules protected by more than “no exception” / “not null” assertions?
|
|
9
|
+
- [ ] Are recent escaped bugs captured by regression tests?
|
|
10
|
+
|
|
11
|
+
## Assertions
|
|
12
|
+
- [ ] Assertions verify semantic outcomes, not just execution.
|
|
13
|
+
- [ ] Negative side effects are checked where important.
|
|
14
|
+
- [ ] Boundary values are represented where rules have thresholds.
|
|
15
|
+
- [ ] Tests are not tautological with their setup.
|
|
16
|
+
|
|
17
|
+
## Coupling
|
|
18
|
+
- [ ] Tests use public unit boundaries where practical.
|
|
19
|
+
- [ ] Private-method tests are avoided.
|
|
20
|
+
- [ ] Mock call order/count is asserted only when contractual.
|
|
21
|
+
- [ ] Fakes/state assertions are preferred when clearer than interaction choreography.
|
|
22
|
+
|
|
23
|
+
## Determinism
|
|
24
|
+
- [ ] Time is controlled.
|
|
25
|
+
- [ ] Randomness is seeded/injected.
|
|
26
|
+
- [ ] Shared state is isolated.
|
|
27
|
+
- [ ] No unexplained sleeps or permanent retries.
|
|
28
|
+
- [ ] No hidden network/database access in a unit test.
|
|
29
|
+
|
|
30
|
+
## Value
|
|
31
|
+
- [ ] Duplicate tests are consolidated.
|
|
32
|
+
- [ ] Obsolete/tautological tests are removed.
|
|
33
|
+
- [ ] Tests are not being added only to raise coverage.
|
|
34
|
+
- [ ] High-risk code gets more attention than trivial code.
|
|
35
|
+
|
|
36
|
+
## Failure sensitivity
|
|
37
|
+
- [ ] At least one representative plausible fault would be caught.
|
|
38
|
+
- [ ] Mutation survivors in critical logic are reviewed when mutation data exists.
|
|
39
|
+
- [ ] A green suite is treated as evidence only if it has demonstrated ability to fail correctly.
|
|
@@ -0,0 +1,85 @@
|
|
|
1
|
+
# Worked Example — From Green Tests to Earned Confidence
|
|
2
|
+
|
|
3
|
+
## Production behavior
|
|
4
|
+
A transfer service should:
|
|
5
|
+
- reject transfers larger than the available balance,
|
|
6
|
+
- leave balances unchanged on rejection,
|
|
7
|
+
- debit the source and credit the destination on success,
|
|
8
|
+
- never execute the same transfer twice for the same idempotency key.
|
|
9
|
+
|
|
10
|
+
## Before: misleading tests
|
|
11
|
+
|
|
12
|
+
```pseudo
|
|
13
|
+
test "transfer works" {
|
|
14
|
+
repo = mock()
|
|
15
|
+
repo.getBalance("A").returns(100)
|
|
16
|
+
repo.save(any()).returns(true)
|
|
17
|
+
|
|
18
|
+
service.transfer("A", "B", 60)
|
|
19
|
+
|
|
20
|
+
verify(repo).save(any())
|
|
21
|
+
}
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
Why this is weak:
|
|
25
|
+
- no result is asserted,
|
|
26
|
+
- balances are not asserted,
|
|
27
|
+
- wrong amounts can still pass,
|
|
28
|
+
- duplicate execution is not checked,
|
|
29
|
+
- rejection behavior is absent,
|
|
30
|
+
- the mock is configured so the interaction succeeds.
|
|
31
|
+
|
|
32
|
+
Classification: **T3 STRENGTHEN**, plus **T5 ADD MISSING PROTECTION**.
|
|
33
|
+
|
|
34
|
+
## After: behavior-focused tests
|
|
35
|
+
|
|
36
|
+
```pseudo
|
|
37
|
+
test "debits source and credits destination for a valid transfer" {
|
|
38
|
+
repo = FakeAccountRepository({A: 100, B: 20})
|
|
39
|
+
service = TransferService(repo)
|
|
40
|
+
|
|
41
|
+
result = service.transfer("tx-1", "A", "B", 60)
|
|
42
|
+
|
|
43
|
+
assert result == Success
|
|
44
|
+
assert repo.balance("A") == 40
|
|
45
|
+
assert repo.balance("B") == 80
|
|
46
|
+
}
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
```pseudo
|
|
50
|
+
test "rejects transfer when funds are insufficient and changes nothing" {
|
|
51
|
+
repo = FakeAccountRepository({A: 50, B: 20})
|
|
52
|
+
service = TransferService(repo)
|
|
53
|
+
|
|
54
|
+
result = service.transfer("tx-2", "A", "B", 60)
|
|
55
|
+
|
|
56
|
+
assert result == InsufficientFunds
|
|
57
|
+
assert repo.balance("A") == 50
|
|
58
|
+
assert repo.balance("B") == 20
|
|
59
|
+
}
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
```pseudo
|
|
63
|
+
test "same idempotency key is applied only once" {
|
|
64
|
+
repo = FakeAccountRepository({A: 100, B: 20})
|
|
65
|
+
service = TransferService(repo)
|
|
66
|
+
|
|
67
|
+
service.transfer("tx-3", "A", "B", 60)
|
|
68
|
+
service.transfer("tx-3", "A", "B", 60)
|
|
69
|
+
|
|
70
|
+
assert repo.balance("A") == 40
|
|
71
|
+
assert repo.balance("B") == 80
|
|
72
|
+
}
|
|
73
|
+
```
|
|
74
|
+
|
|
75
|
+
## Plausible fault challenge
|
|
76
|
+
Introduce these temporary faults and confirm tests fail:
|
|
77
|
+
- change insufficient-funds comparison from `amount > balance` to `amount >= balance`,
|
|
78
|
+
- debit 50 instead of requested amount,
|
|
79
|
+
- omit destination credit,
|
|
80
|
+
- ignore idempotency key.
|
|
81
|
+
|
|
82
|
+
If a fault survives, strengthen the relevant test before declaring the suite trustworthy.
|
|
83
|
+
|
|
84
|
+
## Result
|
|
85
|
+
The refactored tests are fewer in implementation assertions but stronger in behavior protection. They can remain stable if repository internals change, while still failing when transfer semantics break.
|
|
@@ -0,0 +1,59 @@
|
|
|
1
|
+
# Unit Test Refactoring Review
|
|
2
|
+
|
|
3
|
+
## 1. Confidence summary
|
|
4
|
+
**Strong protection:**
|
|
5
|
+
- ...
|
|
6
|
+
|
|
7
|
+
**Misleading confidence:**
|
|
8
|
+
- ...
|
|
9
|
+
|
|
10
|
+
**Top delivery risks:**
|
|
11
|
+
- ...
|
|
12
|
+
|
|
13
|
+
## 2. Confidence Contract
|
|
14
|
+
|
|
15
|
+
| Behavior / invariant | Risk if broken | Current protection | Target protection |
|
|
16
|
+
|---|---|---|---|
|
|
17
|
+
| ... | High/Med/Low | Strong/Weak/Missing | ... |
|
|
18
|
+
|
|
19
|
+
## 3. Findings
|
|
20
|
+
|
|
21
|
+
### TEST-001
|
|
22
|
+
- **Severity:** P1 High
|
|
23
|
+
- **Action:** T3 — STRENGTHEN
|
|
24
|
+
- **Evidence:** ...
|
|
25
|
+
- **Why it matters:** ...
|
|
26
|
+
- **Proposed change:** ...
|
|
27
|
+
- **Expected regression-detection improvement:** ...
|
|
28
|
+
|
|
29
|
+
### TEST-002
|
|
30
|
+
- **Severity:** P2 Medium
|
|
31
|
+
- **Action:** T4 — SIMPLIFY / DE-BRITTLE
|
|
32
|
+
- **Evidence:** ...
|
|
33
|
+
- **Why it matters:** ...
|
|
34
|
+
- **Proposed change:** ...
|
|
35
|
+
- **Expected regression-detection improvement:** ...
|
|
36
|
+
|
|
37
|
+
## 4. Refactoring plan
|
|
38
|
+
1. ...
|
|
39
|
+
2. ...
|
|
40
|
+
3. ...
|
|
41
|
+
|
|
42
|
+
## 5. Representative replacement tests
|
|
43
|
+
|
|
44
|
+
### Before
|
|
45
|
+
```text
|
|
46
|
+
...
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
### After
|
|
50
|
+
```text
|
|
51
|
+
...
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
## 6. Validation plan
|
|
55
|
+
- Fault injection / mutation: ...
|
|
56
|
+
- Escaped bug reproduction: ...
|
|
57
|
+
- Flake monitoring: ...
|
|
58
|
+
- Refactor-resistance check: ...
|
|
59
|
+
- Runtime budget: ...
|