@yottameta/yotta-code-quality 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +22 -0
- package/README.md +247 -0
- package/SKILL.md +130 -0
- package/assets/banner.png +0 -0
- package/bin/install.js +163 -0
- package/install.sh +132 -0
- package/package.json +32 -0
- package/references/AGENTS-template.md +44 -0
- package/references/common.md +141 -0
- package/references/decay-risks.md +252 -0
- package/references/editorial-extensions.md +63 -0
- package/references/examples.md +59 -0
- package/references/hooks.json +10 -0
- package/references/pr-review-guide.md +93 -0
- package/references/source-coverage.md +89 -0
- package/references/test-decay-risks.md +201 -0
|
@@ -0,0 +1,59 @@
|
|
|
1
|
+
# Examples β tone and score calibration
|
|
2
|
+
|
|
3
|
+
Use these to keep reviews consistent. Do **not** copy findings into real reports unless the code matches.
|
|
4
|
+
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
## Example A β Good Warning (R2)
|
|
8
|
+
|
|
9
|
+
**Context:** PR renames a domain field and also βcleans upβ unrelated formatting in 12 files.
|
|
10
|
+
|
|
11
|
+
```
|
|
12
|
+
### π‘ Warning
|
|
13
|
+
**Change Propagation β Unrelated drive-by edits**
|
|
14
|
+
Symptom: Diff touches 12 files; only 2 mention the renamed field `expires_at`; the rest are import reorder / whitespace.
|
|
15
|
+
Source: Fowler β Refactoring β Shotgun Surgery / Divergent Change
|
|
16
|
+
Consequence: Reviewers cannot see the real contract change; regressions hide in noise; blame history is polluted.
|
|
17
|
+
Remedy: Split into (1) rename + call sites, (2) optional format-only PR β or revert non-functional hunks.
|
|
18
|
+
```
|
|
19
|
+
|
|
20
|
+
**Score impact (balanced):** β5.
|
|
21
|
+
|
|
22
|
+
---
|
|
23
|
+
|
|
24
|
+
## Example B β Over-flag to avoid (R1)
|
|
25
|
+
|
|
26
|
+
**Bad:** Mark a 80-line React component Critical solely because `function > 20 lines`, when it is mostly JSX layout with one `map`.
|
|
27
|
+
|
|
28
|
+
**Better:** No finding, or Suggestion: extract presentational subcomponents *when editing nearby* β Source Editorial / Code Complete heuristics as hints.
|
|
29
|
+
|
|
30
|
+
---
|
|
31
|
+
|
|
32
|
+
## Example C β Editorial Critical (R7)
|
|
33
|
+
|
|
34
|
+
```
|
|
35
|
+
### π΄ Critical
|
|
36
|
+
**Release / Supply-chain Safety β Plaintext cloud credentials in publish script**
|
|
37
|
+
Symptom: `publish-*.ps1` embeds AccessKeyId/Secret as string literals committed to the repo.
|
|
38
|
+
Source: Editorial β Release Safety
|
|
39
|
+
Consequence: Anyone with repo (or old clone) access can abuse the cloud account; rotation is urgent.
|
|
40
|
+
Remedy: Remove secrets from VCS history of the file; load from env / local config ignored by git; rotate keys in the provider console now.
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
**Score impact (balanced):** β15.
|
|
44
|
+
|
|
45
|
+
---
|
|
46
|
+
|
|
47
|
+
## Example D β Health Score arithmetic
|
|
48
|
+
|
|
49
|
+
Findings: 1 Critical, 2 Warning, 1 Suggestion; `strictness: balanced`.
|
|
50
|
+
`100 β 15 β 5 β 5 β 1 = 74` β **Health Score: 74/100**.
|
|
51
|
+
|
|
52
|
+
Disclaimer in report: per-run deduction index, not βthe project is 74% healthy.β
|
|
53
|
+
|
|
54
|
+
---
|
|
55
|
+
|
|
56
|
+
## Example E β Quick mode depth
|
|
57
|
+
|
|
58
|
+
User pastes a 30-line function. **Do not** load all of `source-coverage.md`.
|
|
59
|
+
Use SKILL cheat-sheet; open `decay-risks.md` only if you are about to emit Critical.
|
|
@@ -0,0 +1,10 @@
|
|
|
1
|
+
{
|
|
2
|
+
"comment": "PreToolUse safety hooks for yotta-code-quality. Blocks obviously dangerous commands before they run. Compatible with Claude Code / Codex hook schema (PreToolUse on Bash). Adapt the matcher to your agent's hook format if needed.",
|
|
3
|
+
"hooks": [
|
|
4
|
+
{
|
|
5
|
+
"matcher": "Bash",
|
|
6
|
+
"type": "PreToolUse",
|
|
7
|
+
"command": "node -e \"const c=process.env.BASH_COMMAND||(process.argv[2]||'');const deny=[/rm\\s+-rf?\\s+\\//,/rm\\s+-rf?\\s+[A-Za-z]:\\\\?\\\\?/,/git\\s+push\\s+--force/,/git\\s+push\\s+-f\\b/,/git\\s+reset\\s+--hard/,/git\\s+clean\\s+-[a-z]*[fdx]/,/mkfs/,/:\\(\\)\\s*\\{\\s*:\\|:\\|/,/dd\\s+if=/,/chmod\\s+-R\\s+777\\s+\\//];if(deny.some(r=>r.test(c))){console.error('[yotta-code-quality] BLOCKED dangerous command by policy: '+c);process.exit(1);}process.exit(0);\""
|
|
8
|
+
}
|
|
9
|
+
]
|
|
10
|
+
}
|
|
@@ -0,0 +1,93 @@
|
|
|
1
|
+
# PR Review Guide β Mode 1
|
|
2
|
+
|
|
3
|
+
**Purpose:** Analyze a code diff or specific files for decay risks directly visible in the changed
|
|
4
|
+
code. Every finding must follow the Iron Law: Symptom β Source β Consequence β Remedy.
|
|
5
|
+
|
|
6
|
+
## Before You Start
|
|
7
|
+
- **Auto-generated files:** If the diff contains generated files (protobuf stubs, OpenAPI clients,
|
|
8
|
+
ORM migrations, lock files, minified bundles), skip them entirely. Note which were skipped and why.
|
|
9
|
+
- **Scope calibration:** Adjust analysis depth by PR size before starting.
|
|
10
|
+
|
|
11
|
+
| PR Size | Approach |
|
|
12
|
+
|---------|----------|
|
|
13
|
+
| < 50 lines | Steps 1β3 only; Step 6a only if imports changed; Step 6b if any class/method/variable was renamed or introduced |
|
|
14
|
+
| 50β300 lines | Full process, all steps |
|
|
15
|
+
| > 300 lines | Full process; note in Scope that review is sampled β cover highest-risk areas, not every file |
|
|
16
|
+
|
|
17
|
+
For PRs > 500 lines: flag in the Summary that a PR this size is itself a Change Propagation signal.
|
|
18
|
+
|
|
19
|
+
## Analysis Process (work through in order, do not skip)
|
|
20
|
+
|
|
21
|
+
### Step 1: Understand the scope
|
|
22
|
+
- What is the stated purpose of this change? Which files were modified?
|
|
23
|
+
- Flag immediately if the PR changes > 10 unrelated files β π‘ Warning: Change Propagation.
|
|
24
|
+
|
|
25
|
+
### Step 2: Scan for Change Propagation (R2) β first
|
|
26
|
+
- Does this change touch modules with no conceptual connection to the stated purpose?
|
|
27
|
+
- Does any modified class change for > 1 business reason?
|
|
28
|
+
- Does any method use more data from another class than its own?
|
|
29
|
+
- If no cross-module changes beyond what the feature requires β skip, no finding.
|
|
30
|
+
|
|
31
|
+
### Step 3: Scan for Cognitive Overload (R1)
|
|
32
|
+
- New/modified functions > 20 lines? Nesting > 3? > 4 params? Magic numbers? Unreadable names?
|
|
33
|
+
- Train-wreck chains (3+ calls chained)?
|
|
34
|
+
|
|
35
|
+
### Step 4: Scan for Knowledge Duplication (R3)
|
|
36
|
+
- Does this change introduce logic that already exists elsewhere?
|
|
37
|
+
- A new name for a concept that already has a name?
|
|
38
|
+
- A class added to a hierarchy that has a parallel in another module?
|
|
39
|
+
|
|
40
|
+
### Step 5: Scan for Accidental Complexity (R4)
|
|
41
|
+
- Abstraction with only one concrete use?
|
|
42
|
+
- A class that only wraps/delegates?
|
|
43
|
+
- Config options or extension points serving no current requirement?
|
|
44
|
+
|
|
45
|
+
### Step 6a: Scan for Dependency Disorder (R5)
|
|
46
|
+
- New imports from high-level module to low-level one (domain service imports DB driver/HTTP client)?
|
|
47
|
+
- New imports introducing a cycle between modules?
|
|
48
|
+
- An interface forcing callers to depend on methods they don't use?
|
|
49
|
+
- If no new imports/structural changes β skip, no finding.
|
|
50
|
+
|
|
51
|
+
### Step 6b: Scan for Domain Model Distortion (R6)
|
|
52
|
+
- New class/variable names match the business language?
|
|
53
|
+
- A new class holding only data with no behavior where behavior was expected?
|
|
54
|
+
- Logic that belongs to the domain placed in a service/utility layer?
|
|
55
|
+
|
|
56
|
+
## Severity Calibration
|
|
57
|
+
Apply the Iron Law format from `references/common.md`. Use `references/decay-risks.md` severity
|
|
58
|
+
guides as primary reference. Boundary tiebreaker:
|
|
59
|
+
- π΄ Critical β actively breaking velocity or creating production risk *today*.
|
|
60
|
+
- π‘ Warning β will if left unaddressed through the next few features.
|
|
61
|
+
- π’ Suggestion β worth fixing when nearby, not urgent.
|
|
62
|
+
If > 5 findings, add a one-line "Recommended fix order" at the end of Findings.
|
|
63
|
+
|
|
64
|
+
## Step 7: Quick Test Check (run last; three signals only)
|
|
65
|
+
Skip entirely if the diff contains only generated files, config, or docs with no production logic.
|
|
66
|
+
|
|
67
|
+
**Signal 1: Do tests exist for the changed behavior?**
|
|
68
|
+
- Diff modifies production code but no corresponding test changes included β π‘ Warning: Coverage Illusion (Feathers β Working Effectively with Legacy Code, Ch. 1).
|
|
69
|
+
- Pure refactor with existing tests covering behavior β no finding.
|
|
70
|
+
|
|
71
|
+
**Signal 2: Quick Mock Abuse sniff** (only if diff includes test changes)
|
|
72
|
+
- Mock setup obviously longer than test logic? Primary assertions `expect(mock).toHaveBeenCalledWith(...)` with no behavior verification? Production methods added only for tests?
|
|
73
|
+
- If yes β π‘ Warning: Mock Abuse (Osherove β The Art of Unit Testing).
|
|
74
|
+
|
|
75
|
+
**Signal 3: Quick Test Obscurity sniff** (only if diff includes test changes)
|
|
76
|
+
- Test names express scenario and expected outcome? Assertions have message strings?
|
|
77
|
+
- If vague/no messages β π’ Suggestion: Test Obscurity (Meszaros β xUnit Test Patterns, Assertion Roulette p.224).
|
|
78
|
+
|
|
79
|
+
**Output rule:** If all three signals clean β no Test findings, proceed to report. If findings exist β
|
|
80
|
+
add them in Iron Law format, labeled with the test risk name. Note in Summary: "Consider a full
|
|
81
|
+
test-quality review for systemic test problems."
|
|
82
|
+
|
|
83
|
+
## Step 8 (optional): Release / first-paint skim
|
|
84
|
+
|
|
85
|
+
When the user asked for a ship gate, or the diff touches updater, CSP, signing, publish scripts,
|
|
86
|
+
DevTools features, splash/loading, or license/status labels:
|
|
87
|
+
|
|
88
|
+
- Skim `editorial-extensions.md` (R7, UX1).
|
|
89
|
+
- Do not expand into a full Architecture audit unless asked.
|
|
90
|
+
|
|
91
|
+
## Output
|
|
92
|
+
Use the standard Report Template from `references/common.md`.
|
|
93
|
+
Mode: PR Review. Scope: list files reviewed (excluding skipped generated files).
|
|
@@ -0,0 +1,89 @@
|
|
|
1
|
+
---
|
|
2
|
+
books:
|
|
3
|
+
- The Mythical Man-Month
|
|
4
|
+
- Code Complete
|
|
5
|
+
- Refactoring
|
|
6
|
+
- Clean Architecture
|
|
7
|
+
- The Pragmatic Programmer
|
|
8
|
+
- Domain-Driven Design
|
|
9
|
+
- A Philosophy of Software Design
|
|
10
|
+
- Software Engineering at Google
|
|
11
|
+
- xUnit Test Patterns
|
|
12
|
+
- The Art of Unit Testing
|
|
13
|
+
- Working Effectively with Legacy Code
|
|
14
|
+
- How Google Tests Software
|
|
15
|
+
---
|
|
16
|
+
|
|
17
|
+
# Source Coverage Matrix
|
|
18
|
+
|
|
19
|
+
Use this file in **Architecture / Tech Debt** modes, or when two sources disagree on a finding.
|
|
20
|
+
Skip on Quick passes. It exists to prevent shallow "book-name citation" reviews.
|
|
21
|
+
|
|
22
|
+
## Review Discipline
|
|
23
|
+
- Cite a book only when the observed symptom actually matches that book's principle.
|
|
24
|
+
- A threshold crossing is a hint, not a verdict. Check context, intent, and blast radius.
|
|
25
|
+
- Look for justified tradeoffs before flagging a smell as debt.
|
|
26
|
+
- Prefer concrete architectural or domain consequences over abstract style complaints.
|
|
27
|
+
- If two books pull in different directions, state the tradeoff instead of pretending there is no tension.
|
|
28
|
+
|
|
29
|
+
---
|
|
30
|
+
|
|
31
|
+
## Frederick Brooks β *The Mythical Man-Month*
|
|
32
|
+
**Encoded today:** Change propagation as communication overhead; Second-System Effect; Conceptual Integrity.
|
|
33
|
+
**Do not ignore:** Whether the design shows a single coherent idea or competing local optimizations; whether cross-team coordination cost is becoming part of feature cost.
|
|
34
|
+
**Do not over-flag:** Large systems are not automatically second systems; multi-module designs are acceptable when they preserve conceptual integrity.
|
|
35
|
+
|
|
36
|
+
## Steve McConnell β *Code Complete*
|
|
37
|
+
**Encoded today:** Routine length, nesting, naming, magic numbers; construction-phase YAGNI; defensive programming and error-handling discipline.
|
|
38
|
+
**Do not ignore:** Whether low-level readability choices compound into operational risk; whether missing error handling makes failure modes invisible.
|
|
39
|
+
**Do not over-flag:** Small, explicit guard clauses are not cognitive overload; a long routine may be acceptable when linear, well-named, single-purpose.
|
|
40
|
+
|
|
41
|
+
## Martin Fowler β *Refactoring*
|
|
42
|
+
**Encoded today:** Long Method, Long Parameter List, Message Chains, Shotgun Surgery, Divergent Change, Feature Envy, Inappropriate Intimacy, Duplicate Code, Speculative Generality, Lazy Class, Middle Man, Data Class, Flag Arguments, Primitive Obsession.
|
|
43
|
+
**Do not ignore:** Whether the smell is local or systemic; whether a refactoring target has a natural home in the model.
|
|
44
|
+
**Do not over-flag:** Temporary duplication during an active extraction is not always debt; a DTO/boundary record may legitimately be data-focused.
|
|
45
|
+
|
|
46
|
+
## Robert C. Martin β *Clean Architecture*
|
|
47
|
+
**Encoded today:** DIP, ADP, SDP, SAP, layering direction; ISP; LSP; SRP and OCP.
|
|
48
|
+
**Do not ignore:** Policy vs detail boundaries; whether dependency arrows preserve replaceability and testability.
|
|
49
|
+
**Do not over-flag:** Composition roots may depend on concrete infrastructure by design; thin adapter layers can import both directions when explicitly boundary glue.
|
|
50
|
+
|
|
51
|
+
## Andrew Hunt & David Thomas β *The Pragmatic Programmer*
|
|
52
|
+
**Encoded today:** Orthogonality; DRY; Law of Demeter.
|
|
53
|
+
**Do not ignore:** Whether knowledge duplication is really duplicated decision-making; whether coupling is accidental or deliberate local simplification.
|
|
54
|
+
**Do not over-flag:** Similar code in different bounded contexts is not automatically a DRY violation; direct object access inside a cohesive aggregate is not always a Demeter problem.
|
|
55
|
+
|
|
56
|
+
## Eric Evans β *Domain-Driven Design*
|
|
57
|
+
**Encoded today:** Ubiquitous Language; Bounded Context; Anemic Domain Model; Entity vs Value Object; Aggregate Roots.
|
|
58
|
+
**Do not ignore:** Aggregate boundaries, invariant ownership, anti-corruption layers; whether names match the business language.
|
|
59
|
+
**Do not over-flag:** CRUD-heavy workflows may legitimately use transaction scripts; thin entities are acceptable when the domain is simple.
|
|
60
|
+
|
|
61
|
+
## John Ousterhout β *A Philosophy of Software Design*
|
|
62
|
+
**Encoded today:** Deep vs shallow modules; Strategic vs tactical programming; Information Leakage.
|
|
63
|
+
**Do not ignore:** Interface complexity relative to hidden complexity; whether repeated tactical patches raise long-term cognitive load; whether a "helper" exposes internal design decisions callers shouldn't know.
|
|
64
|
+
**Do not over-flag:** Internal implementation complexity is fine when the interface stays simple; a small wrapper is acceptable when it meaningfully absorbs volatility.
|
|
65
|
+
|
|
66
|
+
## Titus Winters, Tom Manshreck, Hyrum Wright β *Software Engineering at Google*
|
|
67
|
+
**Encoded today:** Hyrum's Law; dependency management and upgrade blockage; code sustainability (multi-year maintainability); backward compatibility.
|
|
68
|
+
**Do not ignore:** De facto APIs created by observable behavior; the maintenance cost of too much surface area; whether the dependency graph allows independent upgrades.
|
|
69
|
+
**Do not over-flag:** A stable public API is not a liability if intentionally supported; fan-out alone is not disorder when dependency policy is explicit and governed.
|
|
70
|
+
|
|
71
|
+
## Gerard Meszaros β *xUnit Test Patterns*
|
|
72
|
+
**Encoded today:** Assertion Roulette, Mystery Guest, General Fixture; Eager Test, Lazy Test, Test Code Duplication, Behavior Verification; Erratic Test.
|
|
73
|
+
**Do not ignore:** Whether test failures are diagnosable; whether the suite shape amplifies maintenance cost.
|
|
74
|
+
**Do not over-flag:** Multiple assertions are acceptable when they express one behavior with one failure story; shared fixtures are acceptable when every field is relevant.
|
|
75
|
+
|
|
76
|
+
## Roy Osherove β *The Art of Unit Testing*
|
|
77
|
+
**Encoded today:** Test naming discipline; test isolation; mock usage guidelines; completeness of edge-path tests.
|
|
78
|
+
**Do not ignore:** Whether tests verify behavior rather than wiring; whether seams simplify tests or contort production code for testability.
|
|
79
|
+
**Do not over-flag:** A mock is acceptable when the dependency is nondeterministic and the assertion still verifies behavior; naming conventions are guidance, clarity is the goal.
|
|
80
|
+
|
|
81
|
+
## Michael Feathers β *Working Effectively with Legacy Code*
|
|
82
|
+
**Encoded today:** Legacy code as code without tests; Sensing and Separation; Seams; Characterization Tests.
|
|
83
|
+
**Do not ignore:** Whether the team can change a risky area safely today; whether the code offers any seam for isolating behavior under change.
|
|
84
|
+
**Do not over-flag:** Untested code is not automatically legacy if stable and not under active change; characterization tests matter most before modifying unclear existing behavior.
|
|
85
|
+
|
|
86
|
+
## Google Engineering β *How Google Tests Software*
|
|
87
|
+
**Encoded today:** Change coverage vs line coverage; pyramid shape and suite portfolio economics.
|
|
88
|
+
**Do not ignore:** Whether the suite reflects business risk, not just percentages; whether expensive tests dominate feedback loops.
|
|
89
|
+
**Do not over-flag:** A non-70:20:10 ratio can be healthy when justified by platform constraints or product risk; high coverage is useful when paired with meaningful branch and change protection.
|
|
@@ -0,0 +1,201 @@
|
|
|
1
|
+
# Test Decay Risk Reference
|
|
2
|
+
|
|
3
|
+
Six patterns that cause test suites to degrade. Apply the Iron Law to each finding.
|
|
4
|
+
|
|
5
|
+
## Risk T1: Test Obscurity
|
|
6
|
+
|
|
7
|
+
**Diagnostic question:** How much effort does it take to understand what this test verifies?
|
|
8
|
+
|
|
9
|
+
Unclear test intent breeds distrust, missed failures, and duplicates β one step from an abandoned suite.
|
|
10
|
+
|
|
11
|
+
### Symptoms
|
|
12
|
+
- Assertion Roulette: multiple assertions with no message string.
|
|
13
|
+
- Mystery Guest: test depends on external state (files, DB rows, shared fixtures) invisible in the body.
|
|
14
|
+
- Test names that do not express scenario and expected outcome (`test1`, `shouldWork`, `testLogin`).
|
|
15
|
+
- General Fixture: oversized setUp/beforeEach shared by unrelated tests.
|
|
16
|
+
- Test body requires reading production code to understand what is verified.
|
|
17
|
+
|
|
18
|
+
### Sources
|
|
19
|
+
| Symptom | Book | Principle / Smell |
|
|
20
|
+
|---------|------|-------------------|
|
|
21
|
+
| Assertion Roulette | Meszaros β xUnit Test Patterns | Assertion Roulette (p.224) |
|
|
22
|
+
| Mystery Guest | Meszaros β xUnit Test Patterns | Mystery Guest (p.411) |
|
|
23
|
+
| General Fixture | Meszaros β xUnit Test Patterns | General Fixture (p.316) |
|
|
24
|
+
| Test naming | Osherove β The Art of Unit Testing | method_scenario_expected naming |
|
|
25
|
+
|
|
26
|
+
### Severity Guide
|
|
27
|
+
- π΄ Critical: no test name describes the behavior; all assertions lack messages.
|
|
28
|
+
- π‘ Warning: multiple Mystery Guests; several ambiguous test names.
|
|
29
|
+
- π’ Suggestion: minor naming issues; isolated General Fixture.
|
|
30
|
+
|
|
31
|
+
### What Not to Flag
|
|
32
|
+
- Multiple assertions are acceptable when they describe one coherent behavior and fail with a clear story.
|
|
33
|
+
- Shared setup is fine when every initialized value is relevant to nearly every test.
|
|
34
|
+
- Concise test names are acceptable if scenario and expected outcome are still obvious.
|
|
35
|
+
|
|
36
|
+
---
|
|
37
|
+
|
|
38
|
+
## Risk T2: Test Brittleness
|
|
39
|
+
|
|
40
|
+
**Diagnostic question:** Do tests break when you refactor without changing behavior?
|
|
41
|
+
|
|
42
|
+
Brittle tests punish refactoring β eventually developers stop refactoring to protect the suite.
|
|
43
|
+
|
|
44
|
+
### Symptoms
|
|
45
|
+
- Tests assert on private method results, internal state, or implementation details.
|
|
46
|
+
- Eager Test: one method verifies multiple unrelated behaviors.
|
|
47
|
+
- Over-specified: assertions enforce mock call order or exact param values irrelevant to behavior.
|
|
48
|
+
- Renaming/extracting a method causes > 5 tests to fail with no behavior change.
|
|
49
|
+
- Erratic Test: different results across runs without production change (race, time, random, shared state).
|
|
50
|
+
|
|
51
|
+
### Sources
|
|
52
|
+
| Symptom | Book | Principle / Smell |
|
|
53
|
+
|---------|------|-------------------|
|
|
54
|
+
| Eager Test | Meszaros β xUnit Test Patterns | Eager Test (p.228) |
|
|
55
|
+
| Erratic Test | Meszaros β xUnit Test Patterns | Erratic Test |
|
|
56
|
+
| Implementation coupling | Osherove β The Art of Unit Testing | Test isolation principle |
|
|
57
|
+
| Orthogonality violation | Hunt & Thomas β The Pragmatic Programmer | Ch. 2: Orthogonality |
|
|
58
|
+
|
|
59
|
+
### Severity Guide
|
|
60
|
+
- π΄ Critical: behavior-preserving refactor causes test failures; > 5 tests coupled to one impl detail.
|
|
61
|
+
- π‘ Warning: Eager Tests common; moderate implementation-detail assertions.
|
|
62
|
+
- π’ Suggestion: isolated over-specification in non-critical tests.
|
|
63
|
+
|
|
64
|
+
### What Not to Flag
|
|
65
|
+
- Verifying an externally observable event or emitted command is not implementation coupling.
|
|
66
|
+
- One test with several assertions is acceptable when all support one behavior claim.
|
|
67
|
+
- A fake/in-memory adapter is not brittleness if the test still asserts behavior, not wiring.
|
|
68
|
+
|
|
69
|
+
---
|
|
70
|
+
|
|
71
|
+
## Risk T3: Test Duplication
|
|
72
|
+
|
|
73
|
+
**Diagnostic question:** Is the same test scenario expressed in more than one place?
|
|
74
|
+
|
|
75
|
+
Duplicated tests must change in multiple places and create false confidence without testing distinct behavior.
|
|
76
|
+
|
|
77
|
+
### Symptoms
|
|
78
|
+
- Test Code Duplication: same setup/assertion logic copy-pasted without extraction.
|
|
79
|
+
- Lazy Test: multiple tests verifying identical behavior with no differentiation.
|
|
80
|
+
- Same boundary condition tested identically at unit, integration, and E2E with no layer differentiation.
|
|
81
|
+
- Test helpers/fixtures duplicated across files instead of shared.
|
|
82
|
+
|
|
83
|
+
### Sources
|
|
84
|
+
| Symptom | Book | Principle / Smell |
|
|
85
|
+
|---------|------|-------------------|
|
|
86
|
+
| Test Code Duplication | Meszaros β xUnit Test Patterns | Test Code Duplication (p.213) |
|
|
87
|
+
| Lazy Test | Meszaros β xUnit Test Patterns | Lazy Test (p.232) |
|
|
88
|
+
| DRY violation in tests | Hunt & Thomas β The Pragmatic Programmer | DRY |
|
|
89
|
+
|
|
90
|
+
### Severity Guide
|
|
91
|
+
- π΄ Critical: core scenario fully duplicated across all three test layers with no differentiation.
|
|
92
|
+
- π‘ Warning: common scenario setup repeated in 5+ tests without extraction.
|
|
93
|
+
- π’ Suggestion: minor helper duplication; isolated Lazy Tests.
|
|
94
|
+
|
|
95
|
+
### What Not to Flag
|
|
96
|
+
- The same scenario may appear at unit and integration level when each verifies a distinct risk.
|
|
97
|
+
- Small local setup duplication can be clearer than an over-abstracted fixture maze.
|
|
98
|
+
- Similar assertions against different domain rules are not Lazy Tests if business intent differs.
|
|
99
|
+
|
|
100
|
+
---
|
|
101
|
+
|
|
102
|
+
## Risk T4: Mock Abuse
|
|
103
|
+
|
|
104
|
+
**Diagnostic question:** Is the test more complex than the behavior it tests?
|
|
105
|
+
|
|
106
|
+
Mock abuse produces tests that pass while verifying nothing β production code can be fully broken as
|
|
107
|
+
long as the mocks are wired up.
|
|
108
|
+
|
|
109
|
+
### Symptoms
|
|
110
|
+
- Mock setup code longer than the test logic itself.
|
|
111
|
+
- Primary assertion is `expect(mock).toHaveBeenCalledWith(...)` β verifies a mock was called, not real behavior.
|
|
112
|
+
- Test-only methods added to production classes for lifecycle management in tests.
|
|
113
|
+
- Single unit test uses > 3 mocks.
|
|
114
|
+
- Incomplete Mock: mock missing fields downstream code will access (silent integration failures).
|
|
115
|
+
- Hard-Coded Test Data: no resemblance to real data shapes/constraints.
|
|
116
|
+
|
|
117
|
+
### Sources
|
|
118
|
+
| Symptom | Book | Principle / Smell |
|
|
119
|
+
|---------|------|-------------------|
|
|
120
|
+
| Mock count > 3 | Osherove β The Art of Unit Testing | Mock usage guidelines |
|
|
121
|
+
| Testing mock behavior | Meszaros β xUnit Test Patterns | Behavior Verification (p.544) |
|
|
122
|
+
| Test-only production methods | Feathers β Working Effectively with Legacy Code | Ch. 3: Sensing and Separation |
|
|
123
|
+
| Hard-Coded Test Data | Meszaros β xUnit Test Patterns | Hard-Coded Test Data (p.534) |
|
|
124
|
+
| Incomplete Mock | Osherove β The Art of Unit Testing | Mock completeness requirement |
|
|
125
|
+
|
|
126
|
+
### Severity Guide
|
|
127
|
+
- π΄ Critical: mock setup > 50% of test code; production class has methods only called from tests.
|
|
128
|
+
- π‘ Warning: mocks consistently > 3 per test; primary assertions are mock call verifications.
|
|
129
|
+
- π’ Suggestion: isolated Incomplete Mocks; minor Hard-Coded Test Data.
|
|
130
|
+
|
|
131
|
+
### What Not to Flag
|
|
132
|
+
- A small number of mocks around nondeterministic dependencies is acceptable when assertions still verify behavior.
|
|
133
|
+
- Fakes and spies used to observe state transitions are not mock abuse by default.
|
|
134
|
+
- One interaction assertion may be appropriate when the interaction itself is the behavior under test.
|
|
135
|
+
|
|
136
|
+
---
|
|
137
|
+
|
|
138
|
+
## Risk T5: Coverage Illusion
|
|
139
|
+
|
|
140
|
+
**Diagnostic question:** Does the test suite actually protect against the failures that matter?
|
|
141
|
+
|
|
142
|
+
Coverage measures execution, not verification. 90% line coverage can still miss every critical
|
|
143
|
+
failure mode.
|
|
144
|
+
|
|
145
|
+
### Symptoms
|
|
146
|
+
- High line coverage but error-handling branches, boundaries, exception paths untested.
|
|
147
|
+
- Happy-path only: no sad paths, no null/empty/zero inputs, no concurrency edge cases.
|
|
148
|
+
- Legacy code areas actively modified with no tests (Feathers: "legacy code is code without tests").
|
|
149
|
+
- Coverage % treated as a sign-off criterion; critical change paths remain untested.
|
|
150
|
+
- Tests assert return values but not important side effects (DB writes, event publishes, state transitions).
|
|
151
|
+
|
|
152
|
+
### Sources
|
|
153
|
+
| Symptom | Book | Principle / Smell |
|
|
154
|
+
|---------|------|-------------------|
|
|
155
|
+
| Legacy code = no tests | Feathers β Working Effectively with Legacy Code | Ch. 1 |
|
|
156
|
+
| Change coverage vs line coverage | Google β How Google Tests Software | Ch. 11: Testing at Google Scale |
|
|
157
|
+
| Happy-path only | Osherove β The Art of Unit Testing | Test completeness principle |
|
|
158
|
+
|
|
159
|
+
### Severity Guide
|
|
160
|
+
- π΄ Critical: legacy code area actively modified with no tests; error-handling paths entirely absent.
|
|
161
|
+
- π‘ Warning: coverage > 80% but edge and exception paths systematically absent.
|
|
162
|
+
- π’ Suggestion: a few non-critical paths missing sad-path tests.
|
|
163
|
+
|
|
164
|
+
### What Not to Flag
|
|
165
|
+
- High line coverage is useful when paired with branch, boundary, and change-path coverage.
|
|
166
|
+
- A new module may have limited coverage early if still private and low-risk.
|
|
167
|
+
- Side-effect assertions may live in integration tests without implying a gap.
|
|
168
|
+
|
|
169
|
+
---
|
|
170
|
+
|
|
171
|
+
## Risk T6: Architecture Mismatch
|
|
172
|
+
|
|
173
|
+
**Diagnostic question:** Does the test suite structure reflect the system's actual risk profile?
|
|
174
|
+
|
|
175
|
+
Wrong suite shape is slow and expensive β not from bad tests, but from using the wrong type at the
|
|
176
|
+
wrong layer.
|
|
177
|
+
|
|
178
|
+
### Symptoms
|
|
179
|
+
- Inverted test pyramid: E2E/integration count exceeds unit count β slow, fragile suite.
|
|
180
|
+
- Legacy code with no seam points (no interfaces, DI, or seams) β impossible to test in isolation.
|
|
181
|
+
- Legacy areas modified with no Characterization Tests to capture current behavior before changes.
|
|
182
|
+
- Full suite execution > 10 minutes (architectural problem, not performance).
|
|
183
|
+
- High-risk and low-risk paths tested at identical density; no risk-based prioritization.
|
|
184
|
+
|
|
185
|
+
### Sources
|
|
186
|
+
| Symptom | Book | Principle / Smell |
|
|
187
|
+
|---------|------|-------------------|
|
|
188
|
+
| Inverted pyramid | Google β How Google Tests Software | 70:20:10 unit:integration:E2E ratio |
|
|
189
|
+
| No seam points | Feathers β Working Effectively with Legacy Code | Ch. 4: Seam Model |
|
|
190
|
+
| Missing Characterization Tests | Feathers β Working Effectively with Legacy Code | Ch. 13: Characterization Tests |
|
|
191
|
+
| Suite execution time | Meszaros β xUnit Test Patterns | Slow Tests (p.253) |
|
|
192
|
+
|
|
193
|
+
### Severity Guide
|
|
194
|
+
- π΄ Critical: legacy code modified has no seams and no characterization tests; pyramid fully inverted.
|
|
195
|
+
- π‘ Warning: suite > 10 min; integration/E2E count exceeds unit tests.
|
|
196
|
+
- π’ Suggestion: localized pyramid deviation; a few legacy areas missing characterization tests.
|
|
197
|
+
|
|
198
|
+
### What Not to Flag
|
|
199
|
+
- Deviating from 70:20:10 can be justified by platform constraints or product risk.
|
|
200
|
+
- A suite heavy on integration tests can be healthy if feedback is fast and purposefully layered.
|
|
201
|
+
- A small number of critical-path E2E tests is desirable, not a smell.
|