@prestyj/cli 5.28.1 → 5.29.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/assets/motion/bin/contact-sheet.mjs +9 -3
- package/assets/motion/bin/cues.mjs +337 -0
- package/assets/motion/bin/library.mjs +53 -10
- package/assets/motion/bin/motion-blur.mjs +943 -0
- package/assets/motion/bin/motion-check.mjs +354 -3
- package/assets/motion/bin/music-fit.mjs +436 -0
- package/assets/motion/bin/pdf-extract.mjs +22 -10
- package/assets/motion/bin/reference-study.mjs +346 -0
- package/assets/motion/bin/score-synth.mjs +1052 -93
- package/assets/motion/library/README.md +50 -14
- package/assets/motion/library/kit/moves.js +1981 -0
- package/assets/motion/library/library.json +232 -0
- package/assets/motion/library/pieces/camera-rig/meta.json +13 -0
- package/assets/motion/library/pieces/camera-rig/piece.html +153 -0
- package/assets/motion/library/pieces/camera-rig/preview.jpg +0 -0
- package/assets/motion/library/pieces/chain-knock/meta.json +13 -0
- package/assets/motion/library/pieces/chain-knock/piece.html +195 -0
- package/assets/motion/library/pieces/chain-knock/preview.jpg +0 -0
- package/assets/motion/library/pieces/gather-to-logo/meta.json +13 -0
- package/assets/motion/library/pieces/gather-to-logo/piece.html +159 -0
- package/assets/motion/library/pieces/gather-to-logo/preview.jpg +0 -0
- package/assets/motion/library/pieces/morph-carry/meta.json +13 -0
- package/assets/motion/library/pieces/morph-carry/piece.html +173 -0
- package/assets/motion/library/pieces/morph-carry/preview.jpg +0 -0
- package/assets/motion/library/pieces/one-shape-journey/meta.json +13 -0
- package/assets/motion/library/pieces/one-shape-journey/piece.html +195 -0
- package/assets/motion/library/pieces/one-shape-journey/preview.jpg +0 -0
- package/assets/motion/library/pieces/open-from-subject/meta.json +13 -0
- package/assets/motion/library/pieces/open-from-subject/piece.html +168 -0
- package/assets/motion/library/pieces/open-from-subject/preview.jpg +0 -0
- package/assets/motion/library/pieces/request-to-result/meta.json +13 -0
- package/assets/motion/library/pieces/request-to-result/piece.html +212 -0
- package/assets/motion/library/pieces/request-to-result/preview.jpg +0 -0
- package/assets/motion/library/pieces/scale-dive/meta.json +13 -0
- package/assets/motion/library/pieces/scale-dive/piece.html +321 -0
- package/assets/motion/library/pieces/scale-dive/preview.jpg +0 -0
- package/assets/motion/library/pieces/screen-replica-steps/meta.json +13 -0
- package/assets/motion/library/pieces/screen-replica-steps/piece.html +366 -0
- package/assets/motion/library/pieces/screen-replica-steps/preview.jpg +0 -0
- package/assets/motion/library/pieces/zoom-into-card/meta.json +13 -0
- package/assets/motion/library/pieces/zoom-into-card/piece.html +179 -0
- package/assets/motion/library/pieces/zoom-into-card/preview.jpg +0 -0
- package/assets/motion/library/sheets/diagram.jpg +0 -0
- package/assets/motion/library/sheets/frame.jpg +0 -0
- package/assets/motion/library/sheets/transition.jpg +0 -0
- package/assets/motion/library/sheets/ui.jpg +0 -0
- package/assets/motion/references/build-sheet.md +206 -0
- package/assets/motion/references/runtime/determinism-rules.md +1 -1
- package/assets/motion/references/runtime/gsap-easing-and-stagger.md +29 -29
- package/assets/motion/references/runtime/inputs-and-assets.md +7 -12
- package/assets/motion/references/runtime/lint-validate-inspect.md +3 -3
- package/assets/motion/references/runtime/minimal-composition.md +1 -1
- package/assets/motion/references/runtime/preview-render.md +3 -3
- package/assets/motion/skills/app-walkthrough/SKILL.md +66 -0
- package/assets/motion/skills/before-after/SKILL.md +53 -0
- package/assets/motion/skills/brand-kit/SKILL.md +3 -3
- package/assets/motion/skills/dev-tool-video/SKILL.md +57 -0
- package/assets/motion/skills/launch-video/SKILL.md +62 -0
- package/assets/motion/skills/match-reference/SKILL.md +58 -0
- package/assets/motion/skills/motion/SKILL.md +78 -85
- package/assets/motion/skills/source-ingest/SKILL.md +17 -6
- package/assets/motion/skills/website-video/SKILL.md +59 -0
- package/assets/skills/bulletproof/SKILL.md +36 -11
- package/assets/skills/bulletproof/references/agent-surface.md +19 -9
- package/assets/skills/bulletproof/references/audit-protocol.md +20 -5
- package/assets/skills/bulletproof/references/platform-playbooks.md +5 -4
- package/assets/skills/bulletproof/references/provenance.md +26 -1
- package/assets/skills/bulletproof/references/secure-defaults.md +6 -5
- package/assets/skills/bulletproof/references/supply-chain.md +21 -17
- package/assets/skills/bulletproof/references/threat-landscape.md +28 -26
- package/assets/skills/bulletproof/references/verification.md +2 -0
- package/assets/skills/clarify/SKILL.md +25 -16
- package/assets/skills/code-review/SKILL.md +71 -13
- package/assets/skills/code-review/references/agent-diffs.md +27 -0
- package/assets/skills/code-review/references/tests.md +19 -0
- package/assets/skills/compliance-guard/SKILL.md +20 -5
- package/assets/skills/compliance-guard/references/artifacts.md +1 -1
- package/assets/skills/compliance-guard/references/eu-uk.md +16 -16
- package/assets/skills/compliance-guard/references/lawsuit-vectors.md +5 -5
- package/assets/skills/compliance-guard/references/provenance.md +41 -2
- package/assets/skills/compliance-guard/references/sector-gates.md +3 -3
- package/assets/skills/compliance-guard/references/security-baseline.md +2 -2
- package/assets/skills/compliance-guard/references/trigger-map.md +5 -5
- package/assets/skills/compliance-guard/references/us.md +27 -21
- package/assets/skills/durable/SKILL.md +87 -79
- package/assets/skills/durable/references/agent-db-safety.md +69 -0
- package/assets/skills/durable/references/backups-and-runtime.md +19 -12
- package/assets/skills/durable/references/migrations-and-schema.md +13 -6
- package/assets/skills/evidence-led-ui/SKILL.md +69 -127
- package/assets/skills/evidence-led-ui/references/anti-defaults.md +107 -208
- package/assets/skills/evidence-led-ui/references/direction.md +124 -0
- package/assets/skills/evidence-led-ui/references/production-contract.md +8 -0
- package/assets/skills/evidence-led-ui/references/provenance.md +24 -1
- package/assets/skills/lean/SKILL.md +90 -71
- package/assets/skills/lean/references/memory-and-processes.md +3 -2
- package/assets/skills/lean/references/playbooks.md +37 -12
- package/assets/skills/refactoring/SKILL.md +24 -3
- package/assets/skills/refactoring/references/agent-pitfalls.md +4 -1
- package/assets/skills/refactoring/references/legacy.md +21 -0
- package/assets/skills/root-cause/SKILL.md +20 -10
- package/assets/skills/shared-language/SKILL.md +16 -14
- package/assets/skills/tdd/SKILL.md +27 -15
- package/dist/app-sidecar.js +203 -47
- package/dist/app-sidecar.js.map +1 -1
- package/dist/cli.js +17 -26
- package/dist/cli.js.map +1 -1
- package/dist/core/acceptance-checks.d.ts +48 -0
- package/dist/core/acceptance-checks.js +144 -0
- package/dist/core/acceptance-checks.js.map +1 -0
- package/dist/core/agent-session.d.ts +106 -72
- package/dist/core/agent-session.js +539 -399
- package/dist/core/agent-session.js.map +1 -1
- package/dist/core/agents.d.ts +6 -5
- package/dist/core/agents.js.map +1 -1
- package/dist/core/ask-user.d.ts +90 -8
- package/dist/core/ask-user.js +124 -13
- package/dist/core/ask-user.js.map +1 -1
- package/dist/core/bundled-agents.js +1 -3
- package/dist/core/bundled-agents.js.map +1 -1
- package/dist/core/cache-diagnostics.d.ts +68 -0
- package/dist/core/cache-diagnostics.js +196 -0
- package/dist/core/cache-diagnostics.js.map +1 -0
- package/dist/core/cache-expiry.d.ts +87 -0
- package/dist/core/cache-expiry.js +111 -0
- package/dist/core/cache-expiry.js.map +1 -0
- package/dist/core/compaction/compactor.js +78 -48
- package/dist/core/compaction/compactor.js.map +1 -1
- package/dist/core/compaction/plan-step-policy.d.ts +46 -0
- package/dist/core/compaction/plan-step-policy.js +57 -0
- package/dist/core/compaction/plan-step-policy.js.map +1 -0
- package/dist/core/destructive-git-guard.d.ts +90 -0
- package/dist/core/destructive-git-guard.js +871 -0
- package/dist/core/destructive-git-guard.js.map +1 -0
- package/dist/core/event-bus.d.ts +3 -0
- package/dist/core/event-bus.js +5 -0
- package/dist/core/event-bus.js.map +1 -1
- package/dist/core/injection-detect.d.ts +38 -0
- package/dist/core/injection-detect.js +232 -0
- package/dist/core/injection-detect.js.map +1 -0
- package/dist/core/keep-awake.d.ts +88 -0
- package/dist/core/keep-awake.js +251 -0
- package/dist/core/keep-awake.js.map +1 -0
- package/dist/core/mcp/client.d.ts +72 -0
- package/dist/core/mcp/client.js +264 -41
- package/dist/core/mcp/client.js.map +1 -1
- package/dist/core/mcp/content.js +6 -2
- package/dist/core/mcp/content.js.map +1 -1
- package/dist/core/mcp/store.d.ts +6 -1
- package/dist/core/mcp/store.js +12 -1
- package/dist/core/mcp/store.js.map +1 -1
- package/dist/core/mcp/types.d.ts +18 -0
- package/dist/core/model-unavailable.d.ts +14 -0
- package/dist/core/model-unavailable.js +23 -0
- package/dist/core/model-unavailable.js.map +1 -0
- package/dist/core/node-debugger.d.ts +148 -0
- package/dist/core/node-debugger.js +642 -0
- package/dist/core/node-debugger.js.map +1 -0
- package/dist/core/package-threats.d.ts +18 -0
- package/dist/core/package-threats.js +168 -0
- package/dist/core/package-threats.js.map +1 -0
- package/dist/core/persistent-shell.d.ts +58 -6
- package/dist/core/persistent-shell.js +331 -49
- package/dist/core/persistent-shell.js.map +1 -1
- package/dist/core/process-manager.d.ts +14 -0
- package/dist/core/process-manager.js +61 -0
- package/dist/core/process-manager.js.map +1 -1
- package/dist/core/progress/git-xp.js +8 -14
- package/dist/core/progress/git-xp.js.map +1 -1
- package/dist/core/session-history.d.ts +12 -0
- package/dist/core/session-history.js +27 -0
- package/dist/core/session-history.js.map +1 -1
- package/dist/core/session-manager.d.ts +13 -1
- package/dist/core/session-manager.js +38 -18
- package/dist/core/session-manager.js.map +1 -1
- package/dist/core/session-summary-index.d.ts +37 -0
- package/dist/core/session-summary-index.js +172 -0
- package/dist/core/session-summary-index.js.map +1 -0
- package/dist/core/settings-manager.d.ts +2 -0
- package/dist/core/settings-manager.js +10 -0
- package/dist/core/settings-manager.js.map +1 -1
- package/dist/core/shell-threats-popular-packages.d.ts +11 -0
- package/dist/core/shell-threats-popular-packages.js +675 -0
- package/dist/core/shell-threats-popular-packages.js.map +1 -0
- package/dist/core/shell-threats.d.ts +8 -0
- package/dist/core/shell-threats.js +186 -0
- package/dist/core/shell-threats.js.map +1 -0
- package/dist/core/skills.js +3 -1
- package/dist/core/skills.js.map +1 -1
- package/dist/core/stream-rules.d.ts +30 -0
- package/dist/core/stream-rules.js +151 -0
- package/dist/core/stream-rules.js.map +1 -0
- package/dist/core/subagent-manager.d.ts +20 -5
- package/dist/core/subagent-manager.js +22 -8
- package/dist/core/subagent-manager.js.map +1 -1
- package/dist/core/subagent-receipt.d.ts +54 -0
- package/dist/core/subagent-receipt.js +276 -0
- package/dist/core/subagent-receipt.js.map +1 -0
- package/dist/core/subagent-turn-record.d.ts +2 -0
- package/dist/core/subagent-turn-record.js.map +1 -1
- package/dist/core/test-impact.d.ts +73 -0
- package/dist/core/test-impact.js +467 -0
- package/dist/core/test-impact.js.map +1 -0
- package/dist/core/thinking-level.d.ts +1 -1
- package/dist/core/thinking-level.js +1 -1
- package/dist/core/thinking-level.js.map +1 -1
- package/dist/core/verification-gate.d.ts +2 -0
- package/dist/core/verification-gate.js +4 -0
- package/dist/core/verification-gate.js.map +1 -1
- package/dist/core/verification-snapshot.js +3 -5
- package/dist/core/verification-snapshot.js.map +1 -1
- package/dist/core/workspace-guard.d.ts +18 -7
- package/dist/core/workspace-guard.js +227 -60
- package/dist/core/workspace-guard.js.map +1 -1
- package/dist/interactive.js +2 -1
- package/dist/interactive.js.map +1 -1
- package/dist/modes/subagent-worker-mode.js +34 -6
- package/dist/modes/subagent-worker-mode.js.map +1 -1
- package/dist/motion-agent/motion-agent.d.ts +6 -2
- package/dist/motion-agent/motion-agent.js +5 -7
- package/dist/motion-agent/motion-agent.js.map +1 -1
- package/dist/motion-agent/motion-prompt.d.ts +1 -1
- package/dist/motion-agent/motion-prompt.js +14 -19
- package/dist/motion-agent/motion-prompt.js.map +1 -1
- package/dist/motion-agent/motion-review.d.ts +10 -3
- package/dist/motion-agent/motion-review.js +15 -7
- package/dist/motion-agent/motion-review.js.map +1 -1
- package/dist/motion-agent/motion-studio-context.js +1 -1
- package/dist/motion-agent/motion-studio-context.js.map +1 -1
- package/dist/system-prompt.js +3 -1
- package/dist/system-prompt.js.map +1 -1
- package/dist/test-support/keep-alive.d.ts +14 -0
- package/dist/test-support/keep-alive.js +17 -0
- package/dist/test-support/keep-alive.js.map +1 -0
- package/dist/tools/ask-user.js +3 -3
- package/dist/tools/ask-user.js.map +1 -1
- package/dist/tools/bash-read-evidence.d.ts +10 -0
- package/dist/tools/bash-read-evidence.js +133 -0
- package/dist/tools/bash-read-evidence.js.map +1 -0
- package/dist/tools/bash.d.ts +10 -1
- package/dist/tools/bash.js +115 -7
- package/dist/tools/bash.js.map +1 -1
- package/dist/tools/debug.d.ts +54 -0
- package/dist/tools/debug.js +233 -0
- package/dist/tools/debug.js.map +1 -0
- package/dist/tools/edit.js +12 -4
- package/dist/tools/edit.js.map +1 -1
- package/dist/tools/goals.d.ts +1 -1
- package/dist/tools/index.d.ts +19 -2
- package/dist/tools/index.js +47 -7
- package/dist/tools/index.js.map +1 -1
- package/dist/tools/prompt-hints.js +2 -0
- package/dist/tools/prompt-hints.js.map +1 -1
- package/dist/tools/read-tracker.d.ts +5 -0
- package/dist/tools/read-tracker.js +19 -9
- package/dist/tools/read-tracker.js.map +1 -1
- package/dist/tools/read.js +3 -2
- package/dist/tools/read.js.map +1 -1
- package/dist/tools/skill.js +5 -0
- package/dist/tools/skill.js.map +1 -1
- package/dist/tools/subagent-control.js +44 -8
- package/dist/tools/subagent-control.js.map +1 -1
- package/dist/tools/subagent-shared.d.ts +29 -8
- package/dist/tools/subagent-shared.js +45 -14
- package/dist/tools/subagent-shared.js.map +1 -1
- package/dist/tools/subagent.d.ts +8 -2
- package/dist/tools/subagent.js +28 -10
- package/dist/tools/subagent.js.map +1 -1
- package/dist/tools/task-output.js +3 -2
- package/dist/tools/task-output.js.map +1 -1
- package/dist/tools/task-send.d.ts +1 -1
- package/dist/tools/task-send.js +15 -1
- package/dist/tools/task-send.js.map +1 -1
- package/dist/tools/tool-tiers.d.ts +2 -2
- package/dist/tools/tool-tiers.js +3 -2
- package/dist/tools/tool-tiers.js.map +1 -1
- package/dist/tools/truncate.d.ts +21 -0
- package/dist/tools/truncate.js +187 -0
- package/dist/tools/truncate.js.map +1 -1
- package/dist/tools/ui-adopt.js +2 -0
- package/dist/tools/ui-adopt.js.map +1 -1
- package/dist/ui/App.d.ts +0 -4
- package/dist/ui/App.js +5 -28
- package/dist/ui/App.js.map +1 -1
- package/dist/ui/components/ActivityIndicator.js +1 -0
- package/dist/ui/components/ActivityIndicator.js.map +1 -1
- package/dist/ui/hooks/useAgentLoop.d.ts +1 -8
- package/dist/ui/hooks/useAgentLoop.js +1 -119
- package/dist/ui/hooks/useAgentLoop.js.map +1 -1
- package/dist/ui/render.d.ts +0 -4
- package/dist/ui/render.js +0 -2
- package/dist/ui/render.js.map +1 -1
- package/dist/utils/git.d.ts +77 -0
- package/dist/utils/git.js +285 -21
- package/dist/utils/git.js.map +1 -1
- package/dist/utils/github-ci.js +2 -1
- package/dist/utils/github-ci.js.map +1 -1
- package/dist/utils/github.js +11 -9
- package/dist/utils/github.js.map +1 -1
- package/dist/utils/image.d.ts +14 -0
- package/dist/utils/image.js +16 -0
- package/dist/utils/image.js.map +1 -1
- package/dist/utils/process.d.ts +20 -0
- package/dist/utils/process.js +98 -0
- package/dist/utils/process.js.map +1 -1
- package/package.json +5 -5
- package/assets/motion/references/motion-language.md +0 -128
- package/assets/motion/skills/video-qa/SKILL.md +0 -89
- package/dist/core/ideal-review-subagent.d.ts +0 -56
- package/dist/core/ideal-review-subagent.js +0 -112
- package/dist/core/ideal-review-subagent.js.map +0 -1
- package/dist/core/ideal-review.d.ts +0 -82
- package/dist/core/ideal-review.js +0 -242
- package/dist/core/ideal-review.js.map +0 -1
- package/dist/motion-agent/motion-check-tool.d.ts +0 -35
- package/dist/motion-agent/motion-check-tool.js +0 -514
- package/dist/motion-agent/motion-check-tool.js.map +0 -1
|
@@ -1,110 +1,129 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: lean
|
|
3
|
-
description: Use when speed or resource
|
|
3
|
+
description: Use when speed or resource use matters — while writing features that load, render lists, fetch, poll, spawn processes or cache (inline gate); on "slow / laggy / eating RAM / fans spinning", leaks, zombie or orphan processes, bundle bloat, dead code, Core Web Vitals; for a performance pass or pre-ship "will this run smoothly" check. Any stack — web, backend/API/CLI, desktop (Electron, Tauri), mobile, native, game, ML. Do NOT use for copy/docs-only changes, design direction or visual polish (that is evidence-led-ui; this skill's styling scope is payload and consistency), data-loss or migration safety (durable), or when the user explicitly deprioritizes performance.
|
|
4
4
|
license: Performance engineering guidance, not a benchmark certification. Sources and snapshot date are recorded at the foot of each reference file.
|
|
5
|
-
compatibility: Snapshot dated
|
|
5
|
+
compatibility: Snapshot dated 3 October 2026. Thresholds, tool names, and defaults decay — re-verify with web access before asserting them as current. Claims sourced to that date carry a SNAPSHOT marker.
|
|
6
6
|
---
|
|
7
7
|
|
|
8
8
|
# Lean
|
|
9
9
|
|
|
10
|
-
|
|
10
|
+
**Pick your mode now:**
|
|
11
11
|
|
|
12
|
-
|
|
12
|
+
| You are… | Mode | Do |
|
|
13
|
+
|---|---|---|
|
|
14
|
+
| Writing or editing a feature (any size) | **Inline gate** | Apply *Binding defaults* + *AI-code traps* while you write. One line in the reply on what you did. No benchmarks, no spawning. |
|
|
15
|
+
| Asked "make it faster / why slow / eating RAM / will it run smoothly", a regression, pre-ship | **Full pass** | Run *Full-pass workflow*. Read `references/playbooks.md` sections for the stacks found; `references/memory-and-processes.md` for any memory or process finding. |
|
|
16
|
+
| About to run `kill`/`pkill`/`taskkill`, or "stop that stuck process" | **Process safety** | Obey *Never kill your own host* first. |
|
|
17
|
+
|
|
18
|
+
Users rarely ask for performance: they ask for a feature, then leave when it eats RAM or takes five seconds to open. This skill is on from the first line of code. Do not ship the heavy version intending to "optimize later".
|
|
13
19
|
|
|
14
20
|
## Governing rules
|
|
15
21
|
|
|
16
|
-
1. **Measure, then cut.** In a pass, never optimize from vibes: baseline
|
|
17
|
-
2. **Fix the shared cause once.** The N+1 belongs in the query layer, not a memo at each call site. Check every caller of the slow path
|
|
18
|
-
3. **Optimize user time, not machine time.**
|
|
19
|
-
4. **Memory should be flat.** After N cycles of the core loop, committed memory
|
|
20
|
-
5. **Nothing outlives its job.** Timers, listeners, observers, subscriptions, watchers, child processes, temp files, locks —
|
|
21
|
-
6. **Bounded by default.**
|
|
22
|
-
7. **Small is fast.** Dead code, unused
|
|
23
|
-
8. **Numbers or silence.**
|
|
24
|
-
9. **No perf theater.** Complexity must pay for itself in measured user time
|
|
25
|
-
10. **Label evidence** on every claim: `RUNTIME` (
|
|
22
|
+
1. **Measure, then cut.** In a pass, never optimize from vibes: baseline → find the bottleneck → fix → re-measure. In build mode the binding defaults are pre-paid by platform evidence: apply them without benchmarking.
|
|
23
|
+
2. **Fix the shared cause once.** The N+1 belongs in the query layer, not a memo at each call site. Check every caller of the slow path.
|
|
24
|
+
3. **Optimize user time, not machine time.** Startup, first paint, navigation, hot interactions, the nightly job. Micro-tuning code nobody waits on is last.
|
|
25
|
+
4. **Memory should be flat.** After N cycles of the core loop, committed memory ≈ after 1. A climbing staircase is a leak until proven otherwise; a sawtooth returning to baseline is GC.
|
|
26
|
+
5. **Nothing outlives its job.** Timers, listeners, observers, subscriptions, watchers, child processes, temp files, locks — each has an owner that ends it on success *and* every failure path.
|
|
27
|
+
6. **Bounded by default.** Cache, queue, buffer, retry, log, list render: cap and eviction policy at creation.
|
|
28
|
+
7. **Small is fast.** Dead code, unused deps, duplicate styles are shipped and paid for. Deleting is the cheapest optimization.
|
|
29
|
+
8. **Numbers or silence.** Before/after on the same machine and data, cold and warm. Say what moved, X → Y, and what you could not measure.
|
|
30
|
+
9. **No perf theater.** Complexity must pay for itself in measured user time, or revert it. Caching that adds staleness bugs for an unmeasured gain is a regression.
|
|
31
|
+
10. **Label evidence** on every claim: `RUNTIME` (measured), `CODE` (read in source), `DEDUCED` (inferred), `SNAPSHOT` (dated external source). Never present what you read as what you ran.
|
|
26
32
|
|
|
27
|
-
##
|
|
33
|
+
## Never kill your own host
|
|
28
34
|
|
|
29
|
-
|
|
35
|
+
Agent sessions run *inside* a host process (EZ Coder's daemon/sidecar, an editor, a terminal multiplexer). Killing it kills this session and every sibling session mid-turn. This has happened.
|
|
30
36
|
|
|
31
|
-
|
|
37
|
+
- **Never kill by pattern.** No `pkill -f`, `killall`, `kill $(pgrep …)`, `taskkill /IM`, or loops over `ps | grep` output. Kill only an **exact PID or process group you started in this task** (you hold the PID from your own spawn or background-task handle).
|
|
38
|
+
- **Before touching any PID you did not start:** walk its ancestry (`ps -o pid,ppid,command -p <pid>`, repeat up the PPID chain) and your own (`$$`, `$PPID` upward). If it is you, an ancestor, or another instance of the same host binary (e.g. another `app-sidecar`/daemon) — **do not kill. Report it.**
|
|
39
|
+
- **A process using lots of CPU/RAM is a finding, not a target.** Report PID, command, RSS, age; recommend the fix. Orphan cleanup belongs to the app's startup sweep, never to a live session.
|
|
40
|
+
- Unsure whether you own it? You don't. Ask the user.
|
|
32
41
|
|
|
33
42
|
## Binding defaults (build mode)
|
|
34
43
|
|
|
35
|
-
|
|
44
|
+
- **Teardown ships with the feature.** Whatever you start is torn down in the same module, success and error paths. One cleanup handle per owner: an `AbortController` for a component's fetches and listeners, a `dispose()`, a `finally`.
|
|
45
|
+
- **Cap every accumulation.** LRU/TTL cache, bounded queue, paginated query, virtualized long list, capped retries with backoff + jitter, rotating logs.
|
|
46
|
+
- **Never block the interactive thread.** Work that can exceed a frame (web long-task threshold: 50 ms) is chunked (`scheduler.yield()` where supported, else `setTimeout` chunking), deferred, or moved to a worker / utility process / background thread. No sync fs/crypto/CPU spikes in request handlers or UI code.
|
|
47
|
+
- **Lazy by default, eager only for the first screen.** Rare routes, heavy editors, charts, optional SDKs: dynamic `import()` / deferred `require`. Preload only what the critical path provably needs.
|
|
48
|
+
- **Right-size media.** AVIF/WebP with fallback, explicit `width`/`height`, `srcset`/`sizes`, `loading="lazy"` below the fold, `fetchpriority="high"` on the one LCP image only — never lazy-load it.
|
|
49
|
+
- **Stream, don't hoard.** Stream files and large responses; cursor-paginate queries; chunk big jobs.
|
|
50
|
+
- **Timeout everything external.** Network, subprocesses, locks, queues — and tear down on timeout.
|
|
51
|
+
- **Price a dependency before adopting it** in anything user-facing (bundle impact client-side; `node --cpu-prof -e "require('mod')"` server/desktop).
|
|
52
|
+
- **Batch I/O.** One query for N rows; DOM reads before writes; debounce/throttle expensive handlers.
|
|
53
|
+
- **Events over polling.** Prefer push (WebSocket/SSE, fs watchers, DB notifications, webhooks). If you must poll: slowest interval the UX tolerates, back off when idle, pause on `visibilitychange` hidden, stop on unmount.
|
|
36
54
|
|
|
37
|
-
-
|
|
38
|
-
- **Cap every accumulation.** LRU/TTL cache, bounded queue, paginated query, virtualized list, capped retries with backoff, rotating logs. Unbounded is a bug with a delay.
|
|
39
|
-
- **Never block the interactive thread.** Chunk, defer, or offload work that can exceed a frame (web main thread long task threshold: 50ms) — web workers, background threads, task queues, utility processes. I/O stays async on the hot path; no sync fs/crypto/CPU spikes inside request handlers or UI code.
|
|
40
|
-
- **Lazy by default, eager only for the first screen.** Below-the-fold media, rare routes, heavy editors, optional services: load on demand (dynamic `import()`, deferred `require`, on-demand plugin init). Preload/preconnect only what the critical path demonstrably needs.
|
|
41
|
-
- **Right-size media.** AVIF/WebP with JPEG fallback, explicit `width`/`height` (also kills layout shift), `srcset`/`sizes` for density, decode thumbnails instead of full images.
|
|
42
|
-
- **Stream, don't hoard.** Stream files and large responses, paginate/cursor DB queries, chunk large jobs. Loading a whole dataset into memory to process it item by item is the classic hog.
|
|
43
|
-
- **Timeout everything external.** Network calls, subprocesses, locks, queues — no unbounded waits, and teardown on timeout.
|
|
44
|
-
- **Price a dependency before adopting it** in anything user-facing: its load cost is your load cost (`node --cpu-prof --heap-prof -e "require('mod')"` for Node-side; bundle impact for client-side). The most-downloaded module is not the lightest.
|
|
45
|
-
- **Batch I/O and reads-then-writes.** One query for N rows, not N queries; group DOM reads before writes; coalesce events (debounce/throttle) when handlers are expensive.
|
|
46
|
-
|
|
47
|
-
## Full-pass workflow
|
|
55
|
+
## AI-code traps (check every generated diff)
|
|
48
56
|
|
|
49
|
-
|
|
57
|
+
The patterns agents produce most. Scan your own diff for them before replying.
|
|
50
58
|
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
59
|
+
| Trap | Detect | Fix |
|
|
60
|
+
|---|---|---|
|
|
61
|
+
| Whole-library imports | `import * as Icons`, `import _ from "lodash"`, root imports of big icon/UI/date kits | Per-icon/per-function imports or the library's documented tree-shakable entry; confirm in the bundle analyzer |
|
|
62
|
+
| Effect refetch loops | `useEffect` fetching with object/array/function deps created during render, or no deps array | Fetch on the server / in the route loader / via a query library keyed by primitives |
|
|
63
|
+
| Over-fetching | `SELECT *`, `findMany()` without `take`/`select`, API returning full rows to render two fields | Select needed columns, paginate, shape on the server |
|
|
64
|
+
| N+1 | `await` on a query inside `for` / `.map` | Join/include/batch (`IN (...)`, DataLoader) |
|
|
65
|
+
| Unvirtualized lists | `.map` over unbounded data into DOM or `ScrollView` | Paginate or virtualize (`FlatList`/`FlashList`, TanStack Virtual, `LazyColumn`) |
|
|
66
|
+
| Polling | `setInterval(fetch…)` without cleanup or visibility check | Events; else backoff + pause + cleanup |
|
|
67
|
+
| Everything client-side | `"use client"` at the top of a page tree; browser fetches data the server already had | Push the client boundary down to the interactive leaf |
|
|
68
|
+
| Memo sprinkling | `useMemo`/`useCallback`/`React.memo` on everything | React Compiler on → write plain code, keep manual memo only as an escape hatch (e.g. stable effect deps). Compiler off → memo only what the Profiler implicates |
|
|
69
|
+
| Unbounded module cache | `const cache = new Map()` at module scope, never evicted | LRU/TTL with a size cap |
|
|
56
70
|
|
|
57
|
-
|
|
71
|
+
## Full-pass workflow
|
|
58
72
|
|
|
59
|
-
|
|
73
|
+
1. **Profile the target from the code.** Shape (manifests, lockfiles, build/CI configs), surfaces users wait on, scale (data sizes, session length), platform constraints (low-end devices matter more than the dev machine). Build the coverage ledger (see Scaling).
|
|
74
|
+
2. **Baseline, or say you cannot.** Load/startup time, p50/p99 latency, bundle sizes, RSS/heap at rest and after repeated cycles. Commands: `references/playbooks.md`. If nothing runs, do a `CODE`-labeled static pass and say plainly nothing was measured.
|
|
75
|
+
3. **Sweep the six areas, in order:**
|
|
60
76
|
|
|
61
|
-
| # | Area |
|
|
77
|
+
| # | Area | Hunting |
|
|
62
78
|
|---|---|---|
|
|
63
|
-
| 1 |
|
|
64
|
-
| 2 |
|
|
65
|
-
| 3 |
|
|
66
|
-
| 4 |
|
|
67
|
-
| 5 |
|
|
68
|
-
| 6 |
|
|
79
|
+
| 1 | Loading & startup | Slow first paint/open, render-blocking resources, giant bundles, eager imports, request waterfalls, unoptimized media/fonts, deferrable cold-start work |
|
|
80
|
+
| 2 | Runtime responsiveness | Long tasks / long animation frames, layout thrash, re-render storms, N+1, sync hot paths, missing pagination/virtualization, allocation churn |
|
|
81
|
+
| 3 | Memory | Forgotten timers/listeners/observers, detached DOM, unbounded maps, large captured closures, non-GC cycles, full-size media, whole-file loads |
|
|
82
|
+
| 4 | Processes & lifecycle | Unreaped children, orphans (killed node not tree), no signal handling, leaked ports/fds/locks/temp files, no startup sweep |
|
|
83
|
+
| 5 | Payload & dead weight | Unused deps/files/exports, dead CSS, two libs for one job, debug payloads shipped, tree-shaking blockers |
|
|
84
|
+
| 6 | Styling consistency | Same visual thing built three ways, literals duplicating tokens (design decisions are evidence-led-ui's call) |
|
|
69
85
|
|
|
70
|
-
|
|
86
|
+
4. **Rank by user impact; fix the top three to five.**
|
|
71
87
|
|
|
72
88
|
| Severity | Meaning |
|
|
73
89
|
|---|---|
|
|
74
|
-
|
|
|
75
|
-
|
|
|
76
|
-
|
|
|
77
|
-
|
|
|
78
|
-
|
|
79
|
-
Fix the few that measurement actually implicates — three to five, rarely more. Each fix records its baseline number first. Prefer removing work over hiding it (delete the eager import beats code-splitting it beats deferring it), and prefer the platform primitive over a hand-rolled one.
|
|
80
|
-
|
|
81
|
-
### 5. Verify the fix
|
|
90
|
+
| Critical | Startup in tens of seconds, OOM, multi-second freezes, a leak that kills a session in minutes |
|
|
91
|
+
| High | Sluggish interactions, busy idle CPU, RAM climbing over a workday, zombies accumulating, hot paths 2–10× slower than needed |
|
|
92
|
+
| Medium | Bundle bloat, N+1 under load, unbounded caches, missing pagination, duplicate deps |
|
|
93
|
+
| Low | Dead code/styles, micro-tuning — only when adjacent to a real fix |
|
|
82
94
|
|
|
83
|
-
|
|
95
|
+
Record each baseline first. Prefer removing work over hiding it (delete the eager import > code-split it > defer it); prefer the platform primitive.
|
|
96
|
+
5. **Verify.** Re-measure exactly as baselined. Memory: cycle the path many times, force GC where possible, compare — flat wins. Number didn't move → revert and say so.
|
|
97
|
+
6. **Leave one guard.** Bundle budget (`size-limit`, Lighthouse CI assertions), slow-query threshold, repeated-cycle memory assertion, a test that fails if the eager import returns.
|
|
98
|
+
7. **Report.** Numbers first (X → Y, machine, data, date) → remaining findings ranked with file:line and fix → **not checked** list → changed vs recommended → dropped candidates (false positives). Label every claim.
|
|
84
99
|
|
|
85
|
-
|
|
100
|
+
## Scaling: one agent or several
|
|
86
101
|
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
-
|
|
95
|
-
-
|
|
102
|
+
| Situation | Do |
|
|
103
|
+
|---|---|
|
|
104
|
+
| Inline gate, small edit, single fix | Main thread only. Never spawn. |
|
|
105
|
+
| Full pass, one deployable, you can read every relevant file | Single-threaded; build the ledger yourself. |
|
|
106
|
+
| Full pass, several deployables (e.g. web + API + desktop shell + mobile), or more ledger rows than you can read fully | Fan out **static review** by surface — frontend / backend / processes & lifecycle — in ONE `spawn_agent` call (≤ 6). |
|
|
107
|
+
| A dated claim needs checking (threshold, tool status) | One `researcher` child. |
|
|
108
|
+
|
|
109
|
+
- **Ledger:** rows = six areas × in-scope units. Each row ends checked-with-findings / checked-clean / not-checked(reason).
|
|
110
|
+
- **Child brief (self-contained):** absolute skill root and which `references/*.md` sections to read; slice paths; owned ledger rows; evidence labels; output = findings (file:line, label, severity, fix) + explicit `checked` / `not checked` lists; "read-only: do not kill, start, or benchmark anything". Use `owl`, or a general-purpose child if it must load this skill.
|
|
111
|
+
- **Measurement does not fan out on one machine.** Benchmarks, load tests, and memory-cycle runs stay in the main thread, or one child per *separate environment*, so numbers stay comparable (same machine, same data, same build). Never run benchmarks in parallel on the same machine — they distort each other.
|
|
112
|
+
- **Merge:** a failed, timed-out, or silent child → its rows are `not checked`, never clean. Re-open every reported file:line before reporting it.
|
|
113
|
+
- **Large pre-ship pass:** one fresh-context verifier child tries to disprove the top findings.
|
|
114
|
+
- **Edits stay serialized** in the main thread (or `bee` children on strictly disjoint files); re-measure once after merging.
|
|
96
115
|
|
|
97
116
|
## Honesty rules
|
|
98
117
|
|
|
99
118
|
- Never state or imply the software is "fast", "optimized", or "leak-free". State what was measured, fixed, and left unchecked.
|
|
100
|
-
-
|
|
101
|
-
- Never invent a number, threshold, or version.
|
|
102
|
-
- "I could not measure this" is a
|
|
103
|
-
-
|
|
119
|
+
- A lint rule, analyzer, or profiler is never proof of absence; tools find a minority of what matters.
|
|
120
|
+
- Never invent a number, threshold, or version. Date-sensitive and unverified → `SNAPSHOT` or "verify this".
|
|
121
|
+
- "I could not measure this" is a useful output. A fabricated confirmation is not.
|
|
122
|
+
- Prototype with no users or load → say proportionality applies and stop early.
|
|
104
123
|
|
|
105
124
|
## Reference map
|
|
106
125
|
|
|
107
|
-
Resolve
|
|
126
|
+
Resolve paths from the installed skill root. Load only what the profile triggered.
|
|
108
127
|
|
|
109
|
-
- `references/playbooks.md` — per-stack
|
|
110
|
-
- `references/memory-and-processes.md` —
|
|
128
|
+
- `references/playbooks.md` — per-stack measurement and sweeps: web (Core Web Vitals, LoAF, React/Next/Vite), Node/backend, data & N+1, Electron, Tauri, mobile, native, Python, payload/dead-code tooling, styling consistency.
|
|
129
|
+
- `references/memory-and-processes.md` — leak taxonomy and detection per runtime; process lifecycle: host-safety rule, zombies, orphans, tree kills, graceful shutdown, startup sweeps, container PID 1, leaked handles.
|
|
@@ -44,7 +44,7 @@ ac.abort(); ro.disconnect(); clearInterval(id);
|
|
|
44
44
|
|
|
45
45
|
Web-platform APIs accept `AbortSignal` directly; the rest get disconnected in the same teardown function. One owner, one path — including the error path.
|
|
46
46
|
|
|
47
|
-
**Component frameworks (React and kin):** every subscription started in an effect is returned-from that effect for cleanup, no exceptions; stores and event buses get unsubscribe on unmount; beware module-level singletons accumulating per-component state (registered callbacks never removed — the leak lives on after the component dies). Long lists get virtualized.
|
|
47
|
+
**Component frameworks (React and kin):** every subscription started in an effect is returned-from that effect for cleanup, no exceptions; stores and event buses get unsubscribe on unmount; beware module-level singletons accumulating per-component state (registered callbacks never removed — the leak lives on after the component dies). Long lists get virtualized. With React Compiler enabled, don't hand-add `memo`/`useMemo`; without it, catch a real re-render in the Profiler first — memoization has its own memory cost and is frequently perf theater. Watch `key` misuse: index keys on reorder-prone lists cause remount storms (CPU) and stale-capture bugs (memory).
|
|
48
48
|
|
|
49
49
|
**Rust / C / C++ / Go / JVM / Python:** see the native/compiled and Python sections of `references/playbooks.md` for the per-runtime tools (valgrind/ASan/LSan, heaptrack, pprof, async-profiler/JFR, tracemalloc). The cycle shapes to hunt are listed in the taxonomy above.
|
|
50
50
|
|
|
@@ -64,6 +64,7 @@ Fix the lifetime, not the symptom: remove the registration (or disconnect the ob
|
|
|
64
64
|
`kill(pid)` kills one process; its spawned children (MCP servers, LSPs, shells, helpers) survive as orphans. Correct patterns:
|
|
65
65
|
|
|
66
66
|
- **Never kill the host you run in.** Before killing any long-lived daemon, node process, or "orphan", trace its ancestry (`ps -o pid,ppid,command` up the PPID chain) and check it is not your own host or an ancestor of it — agent sessions live *inside* a host daemon process, so a session killing that daemon kills itself mid-turn, and a child killing its parent kills the whole app's sessions. A process that is merely using lots of memory is a measurement finding, not a kill target; report it. Live-session agents do not kill host daemons at all — orphan cleanup belongs to the host's startup sweep (below), never to a session that might be running inside the thing it aims at.
|
|
67
|
+
- **Never select targets by name or pattern** (`pkill -f`, `killall`, `pgrep | xargs kill`, `taskkill /IM`): another instance of your own host, or a sibling session's server, matches the same pattern. Signal only an exact PID — or the negative PID of a process group — that you spawned in this task and still hold. PIDs are reused: if your handle is stale (process already exited), do not re-resolve by name; check it is still yours (same start time/command) or leave it.
|
|
67
68
|
- **Unix:** start the child in its own process group (`detached: true` in Node, `setsid`/`process_group` elsewhere), then kill the group: `process.kill(-pid, sig)` — the negative PID addresses the whole group. Escalate SIGTERM → grace period → SIGKILL.
|
|
68
69
|
- **Windows:** there are no process groups; `taskkill /PID <pid> /T /F` kills the tree (`/T` = tree, `/F` = force), or use a Job Object assigned to children so closing the handle terminates them all.
|
|
69
70
|
- **Both:** kill children *before* the parent exits, in order — parents can't clean up after they're dead; and children of an interactive app must die when the app does, normal exit *and* crash.
|
|
@@ -95,4 +96,4 @@ Anything that survives step 4 is the finding, and the ledger/sweep/group-kill pa
|
|
|
95
96
|
|
|
96
97
|
---
|
|
97
98
|
|
|
98
|
-
**Provenance:** snapshot 17 August 2026. Sources: Rust std `process::Child` documentation (zombies and reaping), Node.js `child_process` documentation (detached/process-group semantics on Unix and Windows), Docker/tini PID 1 reaping guidance, Chrome DevTools memory documentation (three-snapshot technique, allocation timeline, `queryObjects`), web.dev long-task and lifecycle guidance. OS signal semantics are stable; tool flags and library APIs decay — re-verify specifics with web access before asserting.
|
|
99
|
+
**Provenance:** snapshot 3 October 2026 (content otherwise from the 17 August 2026 review; OS semantics stable). Sources: Rust std `process::Child` documentation (zombies and reaping), Node.js `child_process` documentation (detached/process-group semantics on Unix and Windows), Docker/tini PID 1 reaping guidance, Chrome DevTools memory documentation (three-snapshot technique, allocation timeline, `queryObjects`), web.dev long-task and lifecycle guidance. OS signal semantics are stable; tool flags and library APIs decay — re-verify specifics with web access before asserting.
|
|
@@ -1,10 +1,12 @@
|
|
|
1
1
|
# Lean — per-stack playbooks
|
|
2
2
|
|
|
3
|
-
Load only the sections the profile triggered. Commands assume a POSIX shell unless noted. Claims sourced on the snapshot date (
|
|
3
|
+
Load only the sections the profile triggered. Commands assume a POSIX shell unless noted. Claims sourced on the snapshot date (3 October 2026) carry `SNAPSHOT`; tool versions and thresholds decay — re-verify with web access before asserting as current.
|
|
4
|
+
|
|
5
|
+
**Contents:** Web frontend · React / Next.js / Vite · Node.js / backend / API · Electron · Tauri · Mobile · Native / compiled · Python · Data & queries (N+1) · Payload & dead weight · Styling consistency
|
|
4
6
|
|
|
5
7
|
## Web frontend
|
|
6
8
|
|
|
7
|
-
**Targets (
|
|
9
|
+
**Targets (web.dev Core Web Vitals, field data at p75, mobile and desktop judged separately; unchanged since INP replaced FID in March 2024 — `SNAPSHOT`):**
|
|
8
10
|
|
|
9
11
|
| Metric | Good | Needs improvement | Poor |
|
|
10
12
|
|---|---|---|---|
|
|
@@ -12,9 +14,16 @@ Load only the sections the profile triggered. Commands assume a POSIX shell unle
|
|
|
12
14
|
| INP (Interaction to Next Paint) | ≤ 200ms | ≤ 500ms | > 500ms |
|
|
13
15
|
| CLS (Cumulative Layout Shift) | ≤ 0.1 | ≤ 0.25 | > 0.25 |
|
|
14
16
|
|
|
15
|
-
SEO blogs
|
|
17
|
+
SEO blogs claim 2026 "tightenings" (LCP 2.0 s, stricter INP). None appear on web.dev as of the snapshot — quote only the table above. Lab tools (Lighthouse) cannot measure INP; it is a field metric. Lab = diagnosis, field (CrUX / RUM) = verdict.
|
|
18
|
+
|
|
19
|
+
**Measure:**
|
|
20
|
+
|
|
21
|
+
- Lab: Lighthouse (DevTools, or `npx lighthouse <url> --output json`), DevTools Performance panel (long tasks, long animation frames, layout shifts), Network waterfall, Coverage tab for shipped-but-unused JS/CSS.
|
|
22
|
+
- Field: the `web-vitals` npm package; use its **attribution build** (`import { onINP } from "web-vitals/attribution"`) to get INP phase breakdown (input delay / processing / presentation) and the **Long Animation Frames (LoAF)** entries for the slow interaction — LoAF names the script and function that blocked the frame. LoAF is Chromium-only (shipped Chrome 123); script attribution misses cross-origin iframes, workers, and extensions.
|
|
23
|
+
- SPAs: route changes after the first load are not separate page views in CrUX. Chrome's **Soft Navigations API** was in a final origin trial (Chrome 147–149) with a planned 2026 ship; how CrUX will use it is undecided (`SNAPSHOT`). Until it ships, measure SPA route transitions yourself (`performance.mark` around route change → next paint).
|
|
24
|
+
- Regression gate: Lighthouse CI with assertions (budgets on LCP/CLS/TBT and resource sizes) and/or `size-limit` on bundle bytes.
|
|
16
25
|
|
|
17
|
-
**
|
|
26
|
+
**Navigation speed:** the **Speculation Rules API** (`<script type="speculationrules">`) prefetches/prerenders likely-next pages. Chromium-only by default as of snapshot (Safari had it behind a flag; Firefox not shipped) — progressive enhancement, harmless elsewhere. Start with `moderate`/`conservative` eagerness; never prerender pages with side effects (logout, cart mutation, analytics that count views) or per-user sensitive content.
|
|
18
27
|
|
|
19
28
|
**Symptom → usual cause:**
|
|
20
29
|
|
|
@@ -25,16 +34,24 @@ SEO blogs circulate claims of a March 2026 tightening of LCP to 2.0s (`SNAPSHOT`
|
|
|
25
34
|
| High CLS | Images/embeds without dimensions, late banners pushing content, font swap | Width/height or `aspect-ratio` on all media, reserve ad/banner slots (`min-height`), `content-visibility` for below-fold, avoid injecting above existing content |
|
|
26
35
|
| Slow nav | Waterfall fetches, client-side everything, no prefetch | Parallelize with `Promise.all`, prefetch likely-next routes, partial/staged rendering, keep-alive connections |
|
|
27
36
|
|
|
28
|
-
**Sweep list:** code splitting at routes (
|
|
37
|
+
**Sweep list:** code splitting at routes (dynamic `import()`), tree-shaking blockers (side-effectful modules, CJS in the graph, barrel files re-exporting everything), image audit (format, dimensions, `loading="lazy"` + `decoding="async"` below fold, `fetchpriority="high"` on the LCP image only), font count and subsetting, dependency weight (analyzer: `source-map-explorer` works with any bundler's source maps; `rollup-plugin-visualizer` for Rollup/Vite; `webpack-bundle-analyzer` for webpack; Next.js's built-in analyzer for Next), virtualized long lists, context split instead of one giant provider, `passive: true` scroll/touch listeners, debounced resize/search handlers, third-party scripts (tag managers, chat widgets) — often the top INP offender in LoAF data.
|
|
38
|
+
|
|
39
|
+
## React / Next.js / Vite
|
|
40
|
+
|
|
41
|
+
- **React Compiler 1.0** (stable Oct 2025; works for React and React Native) auto-memoizes components and hooks. With it on, write plain code; React's guidance keeps `useMemo`/`useCallback` only as an escape hatch (e.g. a value used as an effect dependency). In existing code, don't mass-delete manual memo — removing it can change compiled output; remove only with tests. Without the compiler, memo only what the React DevTools Profiler implicates. The compiler does **not** fix effects: an effect that refetches on every render is still a bug.
|
|
42
|
+
- **Next.js 16** (Oct 2025): Turbopack is the default bundler for dev and build. Caching is explicit — with `cacheComponents: true`, dynamic code runs per request and only `"use cache"` functions/components are cached (tune with `cacheLife`/`cacheTag`, invalidate with `revalidateTag`/`updateTag`). Perf checks: is expensive shared data cached? is per-user data kept out of shared cache (`"use cache: private"`)? is `"use client"` pushed down to leaves? are request waterfalls in server components parallelized? Turn on `reactCompiler: true` when the codebase passes the compiler's rules.
|
|
43
|
+
- **Vite 8** (stable 12 Mar 2026): Rolldown (Rust) replaces esbuild + Rollup as the single bundler; config key `build.rollupOptions` → `build.rolldownOptions`; needs Node 20.19+/22.12+. Vite 7 projects still use Rollup. Build-time wins are not runtime wins — measure the shipped bundle, not the build clock.
|
|
29
44
|
|
|
30
45
|
## Node.js / backend / API
|
|
31
46
|
|
|
47
|
+
**Runtime:** Node 26 becomes Active LTS on 28 Oct 2026; Node 24 moves to maintenance on 20 Oct 2026 (EOL 30 Apr 2028); Node 22 EOL 30 Apr 2027 (`SNAPSHOT`, nodejs/Release schedule). New projects target 24 now, 26 once it is LTS. Bun/Deno: same rules apply; benchmark your workload before switching runtimes for speed — vendor benchmarks are not your app.
|
|
48
|
+
|
|
32
49
|
**Measure:** `autocannon` or `k6` for load (watch p99, not the average — the average lies), `clinic doctor` / `clinic flame` / `0x` for CPU and event-loop diagnosis, `--inspect` + Chrome DevTools for heap. DB: slow-query log, `EXPLAIN ANALYZE`.
|
|
33
50
|
|
|
34
51
|
**Sweep list:**
|
|
35
52
|
|
|
36
53
|
- **Event-loop blocking:** sync fs/crypto/zlib in request handlers, `JSON.parse` of huge payloads on the hot path, regex backtracking. Move to workers or streams.
|
|
37
|
-
- **N+1 queries:** one join/include/dataloader instead of a loop of queries.
|
|
54
|
+
- **N+1 queries:** one join/include/dataloader instead of a loop of queries. Detection under *Data & queries*.
|
|
38
55
|
- **Missing indexes:** every frequent filter/sort column; verify with `EXPLAIN ANALYZE` that the plan uses them.
|
|
39
56
|
- **Unpaginated reads:** `SELECT *` on growing tables, `findMany()` without `take`. Cursor pagination for stable ordering.
|
|
40
57
|
- **Connection pools:** sized for the DB's real limit; check for pool exhaustion under load (requests queueing on a connection).
|
|
@@ -58,12 +75,12 @@ Official maintainer guidance (electronjs.org performance tutorial, `SNAPSHOT`):
|
|
|
58
75
|
- **Startup:** `Menu.setApplicationMenu(null)` when no menu is needed; splash/deferred window show to cut time-to-visible; V8 compile cache for large renderer bundles.
|
|
59
76
|
- Security config (`contextIsolation`, sandbox) is bulletproof's lane — but note sandboxing also shrinks renderer memory; do not weaken it for speed.
|
|
60
77
|
|
|
61
|
-
## Tauri
|
|
78
|
+
## Tauri (v2)
|
|
62
79
|
|
|
63
80
|
**Sweep list:**
|
|
64
81
|
|
|
65
82
|
- **Release profile:** in `Cargo.toml` `[profile.release]` — `lto = true` (or `"thin"`), `codegen-units = 1`, `strip = true`; `panic = "abort"` if acceptable. `opt-level = "s"/"z"` trades speed for size — measure which you need.
|
|
66
|
-
- **IPC cost:** every `invoke` serializes
|
|
83
|
+
- **IPC cost:** every `invoke` serializes (JSON by default) — for large binary payloads use Tauri 2's raw request/response bodies or channels rather than JSON arrays of bytes. Don't shuttle large blobs or big JSON back and forth per keystroke; chunk, delta, or move the work to the Rust side. Watch for per-frame IPC from frontend animation/monitoring loops.
|
|
67
84
|
- **State:** `tauri::State` with `Mutex` held across `.await` serializes everything behind it; scope locks tightly.
|
|
68
85
|
- **Frontend** follows the web section exactly — the webview is a browser; bundle size and long tasks hit the same.
|
|
69
86
|
- **Assets:** embed vs. fetch per asset class; large binaries should not ship inside the binary if they can be fetched/unpacked on demand.
|
|
@@ -72,9 +89,13 @@ Official maintainer guidance (electronjs.org performance tutorial, `SNAPSHOT`):
|
|
|
72
89
|
|
|
73
90
|
## Mobile
|
|
74
91
|
|
|
75
|
-
**Sweep list:** cold-start path (lazy screen registration, defer non-critical SDK init — analytics can wait), list virtualization (`FlatList`/`RecyclerView`/`LazyColumn` — never render 1000 rows), image assets per density
|
|
92
|
+
**Sweep list:** cold-start path (lazy screen registration, defer non-critical SDK init — analytics can wait), list virtualization (`FlatList`/`FlashList`/`RecyclerView`/`LazyColumn` — never render 1000 rows), image assets per density (not runtime-downscaled full images), main-thread discipline (decode/parse off the main thread), memory-warning handling that actually drops caches. Measure on a low-end Android device in a release build — debug builds and simulators lie.
|
|
93
|
+
|
|
94
|
+
**React Native:** since 0.82 (Oct 2025) the New Architecture is the only architecture; Hermes is the default engine; Hermes V1 was an experimental opt-in at that release (re-check status before recommending). Old-bridge libraries are a migration problem, not a perf knob. Avoid chatty JS↔native calls per frame; run animations on the UI thread (Reanimated worklets / native driver). React Compiler applies here too.
|
|
95
|
+
|
|
96
|
+
**Android:** add **Baseline Profiles** (Jetpack Macrobenchmark + Baseline Profile Gradle plugin) so startup and hot paths are AOT-compiled; measure cold start with Macrobenchmark `StartupTimingMetric` or `adb shell am start -W`. **iOS:** measure launch with Instruments App Launch template and Xcode Organizer launch metrics; keep work out of `application(_:didFinishLaunching…)` and static initializers.
|
|
76
97
|
|
|
77
|
-
**Measure/leak tools:** Android — Android Studio Profiler, `LeakCanary
|
|
98
|
+
**Measure/leak tools:** Android — Android Studio Profiler, `LeakCanary`, `dumpsys meminfo`; iOS — Instruments Allocations/Leaks, Xcode memory graph debugger.
|
|
78
99
|
|
|
79
100
|
## Native / compiled
|
|
80
101
|
|
|
@@ -88,13 +109,17 @@ Official maintainer guidance (electronjs.org performance tutorial, `SNAPSHOT`):
|
|
|
88
109
|
|
|
89
110
|
Generators/streams instead of list-building for large data; vectorize hot loops (NumPy/pandas — or Polars for large frames) instead of Python-level iteration; never block the asyncio event loop with sync I/O or CPU work (offload to processes); `functools.lru_cache` is bounded — `cached_property` and hand-rolled dict caches are not. Measure: `tracemalloc` for allocation deltas, `memory_profiler` line-level, `py-spy` for live flamegraphs without stopping the process. Watch for reference cycles keeping big objects alive when `gc` is disabled or timing-dependent.
|
|
90
111
|
|
|
112
|
+
**Threads vs GIL:** Python 3.14's free-threaded build (`python3.14t`) is officially supported but optional, not the default (PEP 779). Single-threaded code runs roughly 5–10% slower on it, and importing a C extension that hasn't declared free-threading support re-enables the GIL for the whole process. Recommend it only for CPU-bound threaded work where the dependency stack supports it, and measure both builds. Otherwise: `multiprocessing`/`ProcessPoolExecutor` for CPU, asyncio for I/O.
|
|
113
|
+
|
|
91
114
|
## Data & queries (any stack)
|
|
92
115
|
|
|
116
|
+
**N+1 detection:** turn on query logging in dev and count queries per request (a list page issuing 1 + N similar queries is the signature). Tools: Django `nplusone`/django-debug-toolbar, Rails Bullet or `strict_loading`, SQLAlchemy `lazy="raise"`, Prisma/Drizzle query logging, Laravel `Model::preventLazyLoading()`, GraphQL → DataLoader. Fix with eager loading / one batched query; add a test asserting the query count.
|
|
117
|
+
|
|
93
118
|
`EXPLAIN (ANALYZE)` the slow ones; indexes on frequent filters/sorts — but every index taxes writes, so measure both sides. Batch writes; avoid N single-row inserts inside a transaction per row. Cursor/keyset pagination over offset for deep pages. Slow-query log threshold low enough to catch regressions in CI-like environments. Cache expensive reads with explicit invalidation, not hope.
|
|
94
119
|
|
|
95
120
|
## Payload & dead weight (all stacks)
|
|
96
121
|
|
|
97
|
-
`SNAPSHOT` (
|
|
122
|
+
`SNAPSHOT` (3 Oct 2026): **Knip** finds unused files, exports, dependencies, and devDependencies across JS/TS monorepo workspaces; `depcheck` is an older, narrower alternative (dependencies only). Bundle bytes: `size-limit` (fails CI over budget) plus a visualizer (see Web sweep list). CSS: PurgeCSS (or the framework's built-in pruning) against real markup — beware class names built dynamically, which purgers cannot see; guard with a safelist. Shipped-code audit: DevTools Coverage for what the browser actually ran.
|
|
98
123
|
|
|
99
124
|
**Sweep list:** `knip` in CI; duplicate dependencies (`pnpm why <pkg>`, `npm ls <pkg>`) — two versions of one library is double weight; "two of the same job" audit (two icon sets, two date libs, two CSS systems, a utility lib plus hand-rolled copies of its functions); polyfills for browsers no longer supported; debug/symbol payloads shipped in release (strip; source maps to a symbol server, not the bundle); largest-files audit (`du`, bundle analyzer) — the top ten files are usually the whole story; unused assets in repos (images, fonts nobody references).
|
|
100
125
|
|
|
@@ -104,4 +129,4 @@ The rule: **one way to express one visual decision.** Hunt repeated rule blocks
|
|
|
104
129
|
|
|
105
130
|
---
|
|
106
131
|
|
|
107
|
-
**Provenance:** snapshot
|
|
132
|
+
**Provenance:** snapshot 3 October 2026 (all URLs accessed 3 Oct 2026). Sources: https://web.dev/articles/vitals (thresholds); https://web.dev/articles/find-slow-interactions-in-the-field and https://developer.chrome.com/docs/web-platform/long-animation-frames (attribution build, LoAF); https://developer.chrome.com/blog/final-soft-navigations-origin-trial (Chrome 147–149, CrUX use undecided); Speculation Rules support from secondary 2026 coverage (corewebvitals.io, uploadcare.com — Chromium default, Safari 26.2 flag; re-verify on caniuse); https://react.dev/blog/2025/10/07/react-compiler-1; https://nextjs.org/docs/app/api-reference/directives/use-cache; https://vite.dev/blog/announcing-vite8; https://github.com/nodejs/Release/blob/main/schedule.json (Node release dates); https://docs.python.org/3/whatsnew/3.14.html (PEP 779); React Native 0.82 release coverage (New Architecture only, Hermes V1 experimental). Carried from the 17 Aug 2026 snapshot, not re-verified: Electron official performance tutorial (module cost, lazy loading, main-process blocking, profiling guidance), Rust std `process::Child` docs (zombie reaping), Knip documentation and 2026 ecosystem coverage (dead-code standard claim — `SNAPSHOT`), Valgrind/LeakCanary/Instruments/pprof public docs for tool usage. Version-specific flags and thresholds decay fastest; re-verify before asserting.
|
|
@@ -1,11 +1,20 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: refactoring
|
|
3
|
-
description: Use when restructuring existing code without changing behavior — "refactor", "clean up
|
|
3
|
+
description: Use when restructuring existing code without changing behavior — "refactor", "clean up", "reduce duplication", "split this god class", "modernize", migrating legacy code or APIs (strangler fig, parallel change, codemods across many files, callback→async), planning a refactor, or mid-build when a change needs preparatory restructuring first. Two hats, test-guarded steps, revert-on-red. Do NOT use for new features, bug fixes (stubborn ones: root-cause), performance tuning (lean), schema/data migrations (durable), styling or copy, the refactor step of a user-requested TDD flow (tdd), or a from-scratch rewrite.
|
|
4
4
|
license: Behavior-preservation methodology synthesized from public sources (Fowler's Refactoring catalog, Tidy First?, and community agent skills by bienhoang, wondelai, mattpocock, jeffallan, vasilyu1983), audited 2026-09-12.
|
|
5
5
|
---
|
|
6
6
|
|
|
7
7
|
# Refactoring
|
|
8
8
|
|
|
9
|
+
**Route first:**
|
|
10
|
+
|
|
11
|
+
| Situation | Mode | Next |
|
|
12
|
+
|---|---|---|
|
|
13
|
+
| Tidy a module/function you are already touching, suite green | **1. In-the-small** | Execute loop, inline |
|
|
14
|
+
| User wants a plan, architectural change, or competing approaches | **2. Plan-first** | Modes → 2 |
|
|
15
|
+
| No/red/unrunnable tests, or migration bigger than one session | **3. Legacy** | `references/legacy.md` before any edit |
|
|
16
|
+
| Same mechanical change across >~10 files or >~500 lines | **Codemod** | `references/legacy.md` → Codemods |
|
|
17
|
+
|
|
9
18
|
Change the structure of code without changing what it does. Every observable behavior that exists before the refactoring — including the bugs — must exist after it. This skill exists because agents drift: without hard gates, "refactoring" silently becomes rewriting, and rewriting untested code is how behavior is lost.
|
|
10
19
|
|
|
11
20
|
**Without tests, you're not refactoring — you're editing.**
|
|
@@ -34,9 +43,9 @@ Change the structure of code without changing what it does. Every observable beh
|
|
|
34
43
|
## Execute loop
|
|
35
44
|
|
|
36
45
|
1. **Baseline.** Run the suite on unmodified code; record green. If red or absent → mode 3. If the project has no VCS or the working tree is dirty with unrelated changes, say so and stop for direction — a dirty baseline destroys the revert safety.
|
|
37
|
-
2. **Commit checkpoint.** Refactoring is safest with per-step commits on a dedicated branch — ask the user once, up front: "I'll commit each verified step on a branch — good?" If they decline, keep steps small and separable and report the step list for review at the end. Never commit without authorization
|
|
46
|
+
2. **Commit checkpoint.** Refactoring is safest with per-step commits on a dedicated branch — ask the user once, up front: "I'll commit each verified step on a branch — good?" If they decline, keep steps small and separable and report the step list for review at the end. Never commit without authorization.
|
|
38
47
|
3. **Pick one target.** Ranked by risk-adjusted value, not by how interesting it is: security → correctness → structure → duplication → naming. Hotspots first — files where churn (recent edit frequency) meets complexity. Smell catalog and metrics thresholds: `references/smells.md`.
|
|
39
|
-
4. **Apply one named transformation.** Full mechanics per transformation live in `references/smells.md`. Prefer language-aware tooling (IDE rename, AST codemods) over regex edits; at scale (>~10 files or >~500 lines), a codemod is the safe path and regex is the wrong one.
|
|
48
|
+
4. **Apply one named transformation.** Full mechanics per transformation live in `references/smells.md`. Prefer language-aware tooling (`code_nav` / IDE rename, AST codemods) over regex edits; at scale (>~10 files or >~500 lines), a codemod (ast-grep, jscodeshift, OpenRewrite) is the safe path and regex is the wrong one — workflow in `references/legacy.md`.
|
|
40
49
|
5. **Verify.** Smallest relevant suite → green ⇒ commit (message = transformation name) → next target. Red ⇒ rule 4 of the governing rules: revert, take a smaller step.
|
|
41
50
|
6. **Close.** Full suite + typecheck + lint, on the same commands CI runs — including every package that imports the code you touched, not just the one you edited (monorepos: respect build order). Report: transformations applied (named), before/after state, smells left and why, anything deferred. Agent-specific drift modes to check before closing: `references/agent-pitfalls.md`.
|
|
42
51
|
|
|
@@ -57,6 +66,18 @@ When unsure whether behavior could change: **do not apply — ask.**
|
|
|
57
66
|
- You cannot run or construct any safety net and the risk is medium+. Report and stop.
|
|
58
67
|
- The rewrite instinct hits ("this is all wrong, let me start fresh"). That is a different conversation with the user, not a refactor — and big-bang rewrites lose the one thing refactoring keeps: working software at every step.
|
|
59
68
|
|
|
69
|
+
## Scaling: one agent or several
|
|
70
|
+
|
|
71
|
+
| Situation | Do |
|
|
72
|
+
|---|---|
|
|
73
|
+
| In-the-small, or one package you can fully read | Main thread only. Never spawn. |
|
|
74
|
+
| Large codebase: changed symbol has many callers across packages | Impact map first: ONE `spawn_agent` call, ≤ 6 read-only `owl` children, one per package/directory. Brief: the symbol(s) and their definition file:line, the slice's paths, output = every caller/import/dynamic reference as file:line + `checked`/`not checked` lists. |
|
|
75
|
+
| Dated claim needed (framework migration, deprecated API) | One `researcher` child; label result SNAPSHOT with source URL. |
|
|
76
|
+
| Executing transformations | **Serialized in the main thread**, one named step at a time, revert-on-red. Never parallel edits on a shared tree. |
|
|
77
|
+
| Truly independent modules (no shared files, no import edges) | Optionally one `worker` per module on its own branch; each runs the full loop and its own baseline. Merge one branch at a time, full suite after each. |
|
|
78
|
+
|
|
79
|
+
Merge rule: a child that fails or omits a slice → that slice is `not checked`; treat its callers as unknown and do not change the public signature. Re-open each reported caller before relying on it.
|
|
80
|
+
|
|
60
81
|
## References
|
|
61
82
|
|
|
62
83
|
- `references/smells.md` — smell catalog, metrics thresholds, prioritization, transformation mechanics
|
|
@@ -11,6 +11,9 @@ LLM agents fail refactoring in specific, repeated ways. Human seniors watch for
|
|
|
11
11
|
- **Null/undefined semantics drift** — old code crashed on null; new code defaults it. Or falsy-checks (`x || y`) "simplified" from explicit `!== undefined` where `0`/`""` are valid values.
|
|
12
12
|
- **Scope and visibility drift** — module-level state introduced where locals were; caching added for performance that changes identity or staleness semantics.
|
|
13
13
|
- **Silent feature addition** — validation, logging, or "improvements" smuggled in because the old code "looked wrong". The old behavior was the spec.
|
|
14
|
+
- **Scope creep** — "while I'm here" edits in files outside the agreed target; reformatting mixed with structural change. Split or drop them.
|
|
15
|
+
- **Duplicate helpers** — extracting a new function that already exists elsewhere; search (`code_search`, `grep`) before creating one.
|
|
16
|
+
- **Config/lockfile drift** — a refactor diff must not touch dependency manifests, lockfiles, or lint/tsconfig rules unless that was the agreed step.
|
|
14
17
|
- **Flaky green** — a test fails, passes on rerun, and gets shrugged off. During a refactor, one flaky red invalidates the whole green: rerun it, and if it flakes, quarantine and fix the flakiness before trusting any suite result.
|
|
15
18
|
|
|
16
19
|
## Test-healing anti-pattern
|
|
@@ -25,7 +28,7 @@ Before deleting code that looks pointless: `git blame`, read the commit message,
|
|
|
25
28
|
|
|
26
29
|
Escalate only as far as the risk level demands:
|
|
27
30
|
|
|
28
|
-
1. **Typecheck** — catches signature/shape drift for
|
|
31
|
+
1. **Typecheck** — catches signature/shape drift for free.
|
|
29
32
|
2. **Lint** — catches dead code, suspicious patterns, unintended scope.
|
|
30
33
|
3. **Unit + characterization suites** — the core proof; smallest relevant suite per step, full suite at close.
|
|
31
34
|
4. **Contract test at the changed boundary** — input/output fixtures frozen before the refactor, diffed after.
|
|
@@ -13,6 +13,8 @@ For codebases or areas where the baseline gate cannot be met: no tests, an unrun
|
|
|
13
13
|
- Assert the captured behavior, including wrong-looking output. If output looks like a bug, record the bug and pin the bug.
|
|
14
14
|
- Sufficient coverage: the paths your refactor will touch and their immediate callers. 80% project coverage is not the bar; coverage of the change surface is.
|
|
15
15
|
|
|
16
|
+
Approval (snapshot) testing is the cheap way to do this for complex outputs: serialize the output, approve it once as the golden file, and any later diff fails the test (ApprovalTests exists for many languages; Jest/Vitest `toMatchSnapshot` works too). Scrub nondeterministic fields (timestamps, IDs) before comparing, or the net is flaky.
|
|
17
|
+
|
|
16
18
|
Gate: the characterization suite is **green against the unmodified system** before any refactoring. A red baseline proves nothing later.
|
|
17
19
|
|
|
18
20
|
**3. Create seams before structure.** A seam is a place where behavior can be intercepted without editing the logic:
|
|
@@ -42,8 +44,27 @@ Never skip the migration phase because tests pass — external consumers may exi
|
|
|
42
44
|
### Strangler Fig
|
|
43
45
|
For replacing a subsystem wholesale. Put a facade in front of the legacy system → route traffic through the facade → intercept one route/slice at a time, implementing it fresh behind the facade (guarded by a feature flag when the slice is risky) → when all slices are intercepted, the legacy core is unused; delete it. Each interception must keep the characterization suite green.
|
|
44
46
|
|
|
47
|
+
## Codemods (large mechanical changes)
|
|
48
|
+
|
|
49
|
+
For one transformation applied across many files (rename, API migration, import moves). Tools: **ast-grep** (polyglot structural search/rewrite), **jscodeshift** (JS/TS transforms), **OpenRewrite** (recipes for large-scale refactoring, Java-first, also JS/TS). Use what is already installed; installing one needs the user's OK.
|
|
50
|
+
|
|
51
|
+
1. **Search before rewrite.** Run the pattern read-only; count matches; spot-check ~5 by hand, including the weird ones (comments, strings, dynamic access, re-exports).
|
|
52
|
+
2. **Dry run on one directory**, review the diff, run that package's tests.
|
|
53
|
+
3. **Apply repo-wide as its own commit** — the codemod commit contains only codemod output. Hand fixes for leftovers go in a separate commit, so reviewers can trust the mechanical one.
|
|
54
|
+
4. **Grep for survivors** the pattern missed (dynamic calls, string references, docs); list them, do not silently hand-edit beyond the plan.
|
|
55
|
+
5. Full suite + typecheck + lint across every consuming package.
|
|
56
|
+
|
|
57
|
+
An agent hand-editing 50 files is a rewrite with 50 chances of drift; a codemod is one reviewable rule.
|
|
58
|
+
|
|
45
59
|
## Per-phase discipline (all strategies)
|
|
46
60
|
|
|
47
61
|
Every phase ends with a **validation checkpoint**: characterization + full suite green, and a stated **rollback trigger** — the condition under which this phase is reverted (a commit, a flag flip, never a re-merge). Keep commits bisectable: one phase-step per commit, `git bisect` must stay usable for finding which step changed behavior.
|
|
48
62
|
|
|
49
63
|
Feature flags guard risky slices, but flags are debt: each one gets a removal issue at creation, and is retired as soon as the new path is proven.
|
|
64
|
+
|
|
65
|
+
## Sources (SNAPSHOT, accessed 3 October 2026)
|
|
66
|
+
|
|
67
|
+
- Parallel Change — https://martinfowler.com/bliki/ParallelChange.html
|
|
68
|
+
- Strangler Fig — https://martinfowler.com/bliki/StranglerFigApplication.html
|
|
69
|
+
- ApprovalTests — https://approvaltests.com/
|
|
70
|
+
- ast-grep — https://ast-grep.github.io/ ; jscodeshift — https://jscodeshift.com/ ; OpenRewrite — https://docs.openrewrite.org/
|
|
@@ -1,36 +1,46 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: root-cause
|
|
3
|
-
description: Use when a bug resists the obvious fix,
|
|
3
|
+
description: Use when a bug resists the obvious fix, behaviour makes no sense, a symptom keeps coming back, a test is flaky, something "worked last week", or the user asks why something happens and the answer is not on the surface — including mid-build when a second fix attempt fails. Do NOT use for bugs with a clear repro and an obvious cause (reproduce, fix, re-run directly), security incident triage (bulletproof), or reviewing a diff (code-review).
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Root-Cause
|
|
7
7
|
|
|
8
|
-
|
|
8
|
+
Build a feedback loop, then let evidence kill hypotheses. Each phase has a gate. Keep a hypothesis log in your replies: `H# | prediction | experiment | result | verdict`.
|
|
9
9
|
|
|
10
10
|
## Phase 1 — Build the loop
|
|
11
11
|
|
|
12
|
-
|
|
12
|
+
ONE command that reproduces the failure on demand. Cheapest first: failing test → curl/script → CLI run → headless browser script → trace replay → minimal harness.
|
|
13
13
|
|
|
14
|
-
|
|
14
|
+
Gate: **red-capable**, **deterministic**, **fast** (seconds), **agent-runnable**. Run it red. No red command, no Phase 2.
|
|
15
|
+
|
|
16
|
+
- **Flaky?** Run it N times (e.g. 20–50) and record the fail rate; that rate is the signal. Vary one thing at a time (order, seed, parallelism, timing) until it is deterministic or the rate moves.
|
|
17
|
+
- **Worked before?** With a known-good commit and the loop as a script: `git bisect start <bad> <good>`, then `git bisect run <script>` (exit 0 = good, 125 = skip untestable commit, any other 1–127 = bad). Always `git bisect reset` after. Ask first if the tree has uncommitted work.
|
|
15
18
|
|
|
16
19
|
## Phase 2 — Shrink it
|
|
17
20
|
|
|
18
|
-
|
|
21
|
+
Remove input, config, and code until removing anything more makes it pass (delta debugging: halve, test, keep the failing half). What survives is implicated. Already minimal? Say so.
|
|
19
22
|
|
|
20
23
|
## Phase 3 — Hypotheses
|
|
21
24
|
|
|
22
|
-
|
|
25
|
+
3–5 falsifiable one-liners, each naming the observation that would kill it. Rank by likelihood × cheapness. Show the list (non-blocking).
|
|
23
26
|
|
|
24
27
|
## Phase 4 — One variable
|
|
25
28
|
|
|
26
|
-
Test one hypothesis at a time, cheapest first. Tag debug output `[DBG-xxxx]`
|
|
29
|
+
Test one hypothesis at a time, cheapest first. Never change two things between runs. Instrument boundaries (inputs, outputs, timing), not everything; prefer existing logging/tracing or the debugger. Tag new debug output `[DBG-xxxx]` (random suffix) so removal is one grep.
|
|
27
30
|
|
|
28
31
|
## Phase 5 — Regression test before fix
|
|
29
32
|
|
|
30
|
-
|
|
33
|
+
Write the failing regression test FIRST at a real seam. No honest seam? Record that as a finding and test at the nearest honest boundary. No suite and none requested? Keep the repro script as the check — never introduce a suite unasked.
|
|
31
34
|
|
|
32
35
|
## Phase 6 — Close out
|
|
33
36
|
|
|
34
|
-
If the ask was "why", stop at the answer
|
|
37
|
+
If the ask was "why", stop at the answer and the fix it implies — change code only when the user asks for the fix. Otherwise fix, run the loop green (flaky: N runs, zero failures), run the regression test, grep-remove every `[DBG-xxxx]`, and state the cause in one sentence: cause → mechanism → symptom. Label claims `RUNTIME` / `CODE` / `DEDUCED`.
|
|
38
|
+
|
|
39
|
+
Hard rules: three failed fixes → the hypothesis list is wrong; return to Phase 3. Redact secrets from any quoted or committed log line.
|
|
40
|
+
|
|
41
|
+
## Scaling: one agent or several
|
|
42
|
+
|
|
43
|
+
- Reproduction, shrinking, and experiments: main thread only — one variable at a time.
|
|
44
|
+
- Wide codebase (several packages could own the bug): in Phase 3 you may send one read-only `owl` per hypothesis in a single `spawn_agent` call; brief = symptom, repro command, paths, "return file:line evidence for/against; do not run or edit". Treat replies as `CODE` leads to re-open, not verdicts.
|
|
35
45
|
|
|
36
|
-
|
|
46
|
+
Sources (accessed 3 October 2026): https://git-scm.com/docs/git-bisect; https://www.debuggingbook.org/html/DeltaDebugger.html
|