@selesai/code 0.13.30-beta.0 → 0.13.31
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +7 -1
- package/README.md +8 -2
- package/dist/core/settings-manager-auto-handoff.test.js +11 -0
- package/dist/core/settings-manager.d.ts +10 -0
- package/dist/core/settings-manager.js +6 -0
- package/dist/defaults/settings.json +5 -0
- package/dist/extensions/capability-gateway/index.ts +17 -4
- package/dist/extensions/capability-gateway/integration.test.ts +26 -3
- package/dist/extensions/capability-gateway/routing.test.ts +11 -5
- package/dist/extensions/capability-gateway/routing.ts +6 -6
- package/dist/modes/interactive/components/settings-selector.d.ts +2 -0
- package/dist/modes/interactive/components/settings-selector.js +12 -0
- package/dist/modes/interactive/interactive-mode.js +18 -0
- package/dist/skills/code-review-and-quality/SKILL.md +396 -0
- package/dist/skills/code-simplification/SKILL.md +331 -0
- package/dist/skills/incremental-implementation/SKILL.md +249 -0
- package/dist/skills/planning-and-task-breakdown/SKILL.md +257 -0
- package/dist/skills/references/agent-skills-LICENSE +21 -0
- package/dist/skills/references/definition-of-done.md +67 -0
- package/dist/skills/references/performance-checklist.md +236 -0
- package/dist/skills/references/security-checklist.md +248 -0
- package/docs/settings.md +8 -7
- package/package.json +3 -3
|
@@ -0,0 +1,257 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: planning-and-task-breakdown
|
|
3
|
+
description: Breaks work into ordered tasks. Use when you have a spec or clear requirements and need to break work into implementable tasks. Use when a task feels too large to start, when you need to estimate scope, or when parallel work is possible.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Planning and Task Breakdown
|
|
7
|
+
|
|
8
|
+
## Overview
|
|
9
|
+
|
|
10
|
+
Decompose work into small, verifiable tasks with explicit acceptance criteria. Good task breakdown is the difference between an agent that completes work reliably and one that produces a tangled mess. Every task should be small enough to implement, test, and verify in a single focused session.
|
|
11
|
+
|
|
12
|
+
## When to Use
|
|
13
|
+
|
|
14
|
+
- You have a spec and need to break it into implementable units
|
|
15
|
+
- A task feels too large or vague to start
|
|
16
|
+
- Work needs to be parallelized across multiple agents or sessions
|
|
17
|
+
- You need to communicate scope to a human
|
|
18
|
+
- The implementation order isn't obvious
|
|
19
|
+
|
|
20
|
+
**When NOT to use:** Single-file changes with obvious scope, or when the spec already contains well-defined tasks.
|
|
21
|
+
|
|
22
|
+
## The Planning Process
|
|
23
|
+
|
|
24
|
+
### Step 1: Enter Plan Mode
|
|
25
|
+
|
|
26
|
+
Before writing any code, operate in read-only mode:
|
|
27
|
+
|
|
28
|
+
- Read the spec and relevant codebase sections
|
|
29
|
+
- Identify existing patterns and conventions
|
|
30
|
+
- Map dependencies between components
|
|
31
|
+
- Note risks and unknowns
|
|
32
|
+
|
|
33
|
+
**Do NOT write code during planning.** The output is a plan document saved to `tasks/plan.md` and a task list recorded in the task list target (see Output Files; default `tasks/todo.md`), not implementation.
|
|
34
|
+
|
|
35
|
+
### Step 2: Identify the Dependency Graph
|
|
36
|
+
|
|
37
|
+
Map what depends on what:
|
|
38
|
+
|
|
39
|
+
```
|
|
40
|
+
Database schema
|
|
41
|
+
│
|
|
42
|
+
├── API models/types
|
|
43
|
+
│ │
|
|
44
|
+
│ ├── API endpoints
|
|
45
|
+
│ │ │
|
|
46
|
+
│ │ └── Frontend API client
|
|
47
|
+
│ │ │
|
|
48
|
+
│ │ └── UI components
|
|
49
|
+
│ │
|
|
50
|
+
│ └── Validation logic
|
|
51
|
+
│
|
|
52
|
+
└── Seed data / migrations
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
Implementation order follows the dependency graph bottom-up: build foundations first.
|
|
56
|
+
|
|
57
|
+
### Step 3: Slice Vertically
|
|
58
|
+
|
|
59
|
+
Instead of building all the database, then all the API, then all the UI — build one complete feature path at a time:
|
|
60
|
+
|
|
61
|
+
**Bad (horizontal slicing):**
|
|
62
|
+
```
|
|
63
|
+
Task 1: Build entire database schema
|
|
64
|
+
Task 2: Build all API endpoints
|
|
65
|
+
Task 3: Build all UI components
|
|
66
|
+
Task 4: Connect everything
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
**Good (vertical slicing):**
|
|
70
|
+
```
|
|
71
|
+
Task 1: User can create an account (schema + API + UI for registration)
|
|
72
|
+
Task 2: User can log in (auth schema + API + UI for login)
|
|
73
|
+
Task 3: User can create a task (task schema + API + UI for creation)
|
|
74
|
+
Task 4: User can view task list (query + API + UI for list view)
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
Each vertical slice delivers working, testable functionality.
|
|
78
|
+
|
|
79
|
+
### Step 4: Write Tasks
|
|
80
|
+
|
|
81
|
+
Each task follows this structure, whether it lands in the markdown task list or as an item in an external tracker (see Output Files):
|
|
82
|
+
|
|
83
|
+
```markdown
|
|
84
|
+
## Task [N]: [Short descriptive title]
|
|
85
|
+
|
|
86
|
+
**Description:** One paragraph explaining what this task accomplishes.
|
|
87
|
+
|
|
88
|
+
**Acceptance criteria:**
|
|
89
|
+
- [ ] [Specific, testable condition]
|
|
90
|
+
- [ ] [Specific, testable condition]
|
|
91
|
+
|
|
92
|
+
**Verification:**
|
|
93
|
+
- [ ] Tests pass: [the repository's focused-test command]
|
|
94
|
+
- [ ] Build succeeds: [the repository's build command]
|
|
95
|
+
- [ ] Manual check: [description of what to verify]
|
|
96
|
+
|
|
97
|
+
**Dependencies:** [Task numbers this depends on, or "None"]
|
|
98
|
+
|
|
99
|
+
**Files likely touched:**
|
|
100
|
+
- `src/path/to/file.ts`
|
|
101
|
+
- `tests/path/to/test.ts`
|
|
102
|
+
|
|
103
|
+
**Estimated scope:** [Small: 1-2 files | Medium: 3-5 files | Large: 5+ files]
|
|
104
|
+
```
|
|
105
|
+
|
|
106
|
+
### Step 5: Order and Checkpoint
|
|
107
|
+
|
|
108
|
+
Arrange tasks so that:
|
|
109
|
+
|
|
110
|
+
1. Dependencies are satisfied (build foundation first)
|
|
111
|
+
2. Each task leaves the system in a working state
|
|
112
|
+
3. Verification checkpoints occur after every 2-3 tasks
|
|
113
|
+
4. High-risk tasks are early (fail fast)
|
|
114
|
+
|
|
115
|
+
Add explicit checkpoints to the task list target:
|
|
116
|
+
|
|
117
|
+
```markdown
|
|
118
|
+
## Checkpoint: After Tasks 1-3
|
|
119
|
+
- [ ] All tests pass
|
|
120
|
+
- [ ] Application builds without errors
|
|
121
|
+
- [ ] Core user flow works end-to-end
|
|
122
|
+
- [ ] Review with human before proceeding
|
|
123
|
+
```
|
|
124
|
+
|
|
125
|
+
## Task Sizing Guidelines
|
|
126
|
+
|
|
127
|
+
| Size | Files | Scope | Example |
|
|
128
|
+
|------|-------|-------|---------|
|
|
129
|
+
| **XS** | 1 | Single function or config change | Add a validation rule |
|
|
130
|
+
| **S** | 1-2 | One component or endpoint | Add a new API endpoint |
|
|
131
|
+
| **M** | 3-5 | One feature slice | User registration flow |
|
|
132
|
+
| **L** | 5-8 | Multi-component feature | Search with filtering and pagination |
|
|
133
|
+
| **XL** | 8+ | **Too large — break it down further** | — |
|
|
134
|
+
|
|
135
|
+
If a task is L or larger, it should be broken into smaller tasks. An agent performs best on S and M tasks.
|
|
136
|
+
|
|
137
|
+
**When to break a task down further:**
|
|
138
|
+
- It would take more than one focused session (roughly 2+ hours of agent work)
|
|
139
|
+
- You cannot describe the acceptance criteria in 3 or fewer bullet points
|
|
140
|
+
- It touches two or more independent subsystems (e.g., auth and billing)
|
|
141
|
+
- You find yourself writing "and" in the task title (a sign it is two tasks)
|
|
142
|
+
|
|
143
|
+
## Output Files
|
|
144
|
+
|
|
145
|
+
- **Plan document:** Save the implementation plan to `tasks/plan.md`. This is always a markdown file — design decisions, risks, and open questions don't map cleanly onto individual tracker issues.
|
|
146
|
+
- **Task list:** Record each task in the **task list target** (defined below).
|
|
147
|
+
|
|
148
|
+
Create the `tasks/` directory if it does not exist.
|
|
149
|
+
|
|
150
|
+
**Never overwrite an incomplete plan.** Before writing `tasks/plan.md` or `tasks/todo.md`, check whether they already exist and still contain unchecked tasks:
|
|
151
|
+
|
|
152
|
+
- Same work being replanned (the user asked to revise or extend this plan) → update the existing files in place.
|
|
153
|
+
- Different work → **stop and ask.** The unchecked tasks may be mid-build in another session. Do not delete, overwrite, or rename the existing files on your own; present the conflict and let the user decide (finish the old plan first, explicitly discard it, or tell you where the new plan should go).
|
|
154
|
+
|
|
155
|
+
The same rule applies to an external task list target: never bulk-close or delete another plan's open tracker items to make room for new ones.
|
|
156
|
+
|
|
157
|
+
### Task List Target
|
|
158
|
+
|
|
159
|
+
The task list target is where tasks and checkpoints are recorded. It is defined once, here; every other reference in this skill defers to it.
|
|
160
|
+
|
|
161
|
+
- **Default: a checklist-style markdown file at `tasks/todo.md`.** This is the convention the `/build` command and other downstream tooling expect. Use it unless the project says otherwise.
|
|
162
|
+
- **External tracker:** if the project's agent rules (`CLAUDE.md`, `AGENTS.md`, etc.) or the user designate an issue tracker (e.g. GitHub Issues, Jira, Linear, `bd`/beads), create one tracker item per task instead of writing `tasks/todo.md`. Map the Step 4 structure onto the tracker's fields: acceptance criteria and verification steps in the item body, dependencies via the tracker's linking mechanism (`bd dep add`, "blocked by", etc.). Record Step 5 checkpoints as tracker items too, or as a checklist in the plan document if the tracker has no natural equivalent.
|
|
163
|
+
|
|
164
|
+
When using an external tracker, note it in `tasks/plan.md` (e.g. "Tasks tracked in Linear project FOO") so downstream steps and future sessions know where to look, and keep the plan document's Task List section as an ordered index of tracker item IDs or links rather than a duplicate checklist.
|
|
165
|
+
|
|
166
|
+
## Plan Document Template
|
|
167
|
+
|
|
168
|
+
```markdown
|
|
169
|
+
# Implementation Plan: [Feature/Project Name]
|
|
170
|
+
|
|
171
|
+
## Overview
|
|
172
|
+
[One paragraph summary of what we're building]
|
|
173
|
+
|
|
174
|
+
## Architecture Decisions
|
|
175
|
+
- [Key decision 1 and rationale]
|
|
176
|
+
- [Key decision 2 and rationale]
|
|
177
|
+
|
|
178
|
+
## Task List
|
|
179
|
+
|
|
180
|
+
### Phase 1: Foundation
|
|
181
|
+
- [ ] Task 1: ...
|
|
182
|
+
- [ ] Task 2: ...
|
|
183
|
+
|
|
184
|
+
### Checkpoint: Foundation
|
|
185
|
+
- [ ] Tests pass, builds clean
|
|
186
|
+
|
|
187
|
+
### Phase 2: Core Features
|
|
188
|
+
- [ ] Task 3: ...
|
|
189
|
+
- [ ] Task 4: ...
|
|
190
|
+
|
|
191
|
+
### Checkpoint: Core Features
|
|
192
|
+
- [ ] End-to-end flow works
|
|
193
|
+
|
|
194
|
+
### Phase 3: Polish
|
|
195
|
+
- [ ] Task 5: ...
|
|
196
|
+
- [ ] Task 6: ...
|
|
197
|
+
|
|
198
|
+
### Checkpoint: Complete
|
|
199
|
+
- [ ] All acceptance criteria met
|
|
200
|
+
- [ ] Ready for review
|
|
201
|
+
|
|
202
|
+
## Risks and Mitigations
|
|
203
|
+
| Risk | Impact | Mitigation |
|
|
204
|
+
|------|--------|------------|
|
|
205
|
+
| [Risk] | [High/Med/Low] | [Strategy] |
|
|
206
|
+
|
|
207
|
+
## Open Questions
|
|
208
|
+
- [Question needing human input]
|
|
209
|
+
```
|
|
210
|
+
|
|
211
|
+
When tasks live in an external tracker, keep the Task List section above as an ordered index of tracker item IDs or links instead of a duplicate checklist.
|
|
212
|
+
|
|
213
|
+
## Parallelization Opportunities
|
|
214
|
+
|
|
215
|
+
When multiple agents or sessions are available:
|
|
216
|
+
|
|
217
|
+
- **Safe to parallelize:** Independent feature slices, tests for already-implemented features, documentation
|
|
218
|
+
- **Must be sequential:** Database migrations, shared state changes, dependency chains
|
|
219
|
+
- **Needs coordination:** Features that share an API contract (define the contract first, then parallelize)
|
|
220
|
+
|
|
221
|
+
## Common Rationalizations
|
|
222
|
+
|
|
223
|
+
| Rationalization | Reality |
|
|
224
|
+
|---|---|
|
|
225
|
+
| "I'll figure it out as I go" | That's how you end up with a tangled mess and rework. 10 minutes of planning saves hours. |
|
|
226
|
+
| "The tasks are obvious" | Write them down anyway. Explicit tasks surface hidden dependencies and forgotten edge cases. |
|
|
227
|
+
| "Planning is overhead" | Planning is the task. Implementation without a plan is just typing. |
|
|
228
|
+
| "I can hold it all in my head" | Context windows are finite. Written plans survive session boundaries and compaction. |
|
|
229
|
+
| "The old `tasks/plan.md` is stale, I'll just replace it" | Unchecked tasks may be mid-build in another session. Overwriting them destroys work state that exists nowhere else. Stop and ask. |
|
|
230
|
+
|
|
231
|
+
## Red Flags
|
|
232
|
+
|
|
233
|
+
- Starting implementation without a written task list
|
|
234
|
+
- Overwriting a `tasks/plan.md` or `tasks/todo.md` that still has unchecked tasks for different work, without asking
|
|
235
|
+
- Writing `tasks/todo.md` when the project has designated an external tracker (or scattering tasks across both)
|
|
236
|
+
- Tasks that say "implement the feature" without acceptance criteria
|
|
237
|
+
- No verification steps in the plan
|
|
238
|
+
- All tasks are XL-sized
|
|
239
|
+
- No checkpoints between tasks
|
|
240
|
+
- Dependency order isn't considered
|
|
241
|
+
|
|
242
|
+
## Verification
|
|
243
|
+
|
|
244
|
+
Before starting implementation, confirm:
|
|
245
|
+
|
|
246
|
+
- [ ] Every task has acceptance criteria
|
|
247
|
+
- [ ] Every task has a verification step
|
|
248
|
+
- [ ] Task dependencies are identified and ordered correctly
|
|
249
|
+
- [ ] Tasks are recorded in the task list target (default `tasks/todo.md`)
|
|
250
|
+
- [ ] No pre-existing incomplete plan was overwritten without explicit user confirmation
|
|
251
|
+
- [ ] No task touches more than ~5 files
|
|
252
|
+
- [ ] Checkpoints exist between major phases
|
|
253
|
+
- [ ] The human has reviewed and approved the plan
|
|
254
|
+
|
|
255
|
+
## See Also
|
|
256
|
+
|
|
257
|
+
Acceptance criteria are per-task and answer "did we build the right thing?". They sit on top of the project-wide Definition of Done, the standing bar every task clears before it counts as done. See `../references/definition-of-done.md`.
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2025 Addy Osmani
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,67 @@
|
|
|
1
|
+
# Definition of Done
|
|
2
|
+
|
|
3
|
+
A standing, project-wide bar that every change must clear before it counts as done. Unlike acceptance criteria, which vary per task and answer "did we build the right thing?", the Definition of Done is the same every time and answers "is this finished to our standard?". Use it as the final gate in `planning-and-task-breakdown`, `incremental-implementation`, and `shipping-and-launch`.
|
|
4
|
+
|
|
5
|
+
## Definition of Done vs. Acceptance Criteria
|
|
6
|
+
|
|
7
|
+
| | Acceptance Criteria | Definition of Done |
|
|
8
|
+
|---|---|---|
|
|
9
|
+
| Scope | Specific to one task or spec | Applies to every increment |
|
|
10
|
+
| Changes | Different for each item | Fixed and reused |
|
|
11
|
+
| Answers | "Did we build *this thing*?" | "Is it *ready*?" |
|
|
12
|
+
| Owner | Defined when planning the task | Defined once for the project |
|
|
13
|
+
| Example | "User can reset password via email link" | "Tests pass, no regressions, docs updated" |
|
|
14
|
+
|
|
15
|
+
The two are complementary. A task is done only when **its** acceptance criteria are met **and** the standing Definition of Done is satisfied. Skipping either leaves work that looks finished but is not.
|
|
16
|
+
|
|
17
|
+
## The Standing Checklist
|
|
18
|
+
|
|
19
|
+
Apply this to every change before declaring it done.
|
|
20
|
+
|
|
21
|
+
### Correctness
|
|
22
|
+
- [ ] All acceptance criteria for the task are met
|
|
23
|
+
- [ ] Code runs and behaves as intended, verified at runtime, not just compiled or typechecked
|
|
24
|
+
- [ ] New behavior is covered by tests that fail without the change and pass with it
|
|
25
|
+
- [ ] Existing tests still pass; no regressions introduced
|
|
26
|
+
- [ ] Edge cases and error paths are handled, not just the happy path
|
|
27
|
+
|
|
28
|
+
### Quality
|
|
29
|
+
- [ ] Code reveals intent through naming and structure; no comments needed to explain *what* it does
|
|
30
|
+
- [ ] No duplicated business logic
|
|
31
|
+
- [ ] No dead code, debug output, or commented-out blocks left behind
|
|
32
|
+
- [ ] Changes are scoped to the task; no unrelated refactors snuck in
|
|
33
|
+
- [ ] Linting and formatting pass
|
|
34
|
+
|
|
35
|
+
The depth behind these items lives in `code-review-and-quality` (the five-axis review) and `code-simplification` (reducing complexity without changing behavior).
|
|
36
|
+
|
|
37
|
+
### Integration
|
|
38
|
+
- [ ] Change works with the rest of the system, not just in isolation
|
|
39
|
+
- [ ] Database migrations, config changes, and feature flags are accounted for
|
|
40
|
+
- [ ] Backward compatibility considered for any public interface or API change
|
|
41
|
+
|
|
42
|
+
### Documentation
|
|
43
|
+
- [ ] Public interfaces, APIs, and user-facing behavior are documented
|
|
44
|
+
- [ ] Architectural decisions worth preserving are recorded (see `documentation-and-adrs`)
|
|
45
|
+
- [ ] Documentation describes the current state in timeless language, not the change history
|
|
46
|
+
|
|
47
|
+
### Ship-readiness
|
|
48
|
+
- [ ] Security implications reviewed for any untrusted input, auth, or data handling (see `security-and-hardening`)
|
|
49
|
+
- [ ] Observability in place for new critical paths (logs, metrics, traces) (see `observability-and-instrumentation`)
|
|
50
|
+
- [ ] Rollback path exists for anything risky (see `shipping-and-launch`)
|
|
51
|
+
- [ ] The human has reviewed and approved before merge or deploy
|
|
52
|
+
|
|
53
|
+
## How to Apply
|
|
54
|
+
|
|
55
|
+
- **Per task**: confirm the Correctness and Quality sections before checking the task off.
|
|
56
|
+
- **Per feature**: confirm Integration and Documentation before considering the feature complete.
|
|
57
|
+
- **Per release**: the full checklist is the floor; `shipping-and-launch` adds the deploy-specific gates on top.
|
|
58
|
+
|
|
59
|
+
Tailor the list to the project once, then reuse it unchanged. A Definition of Done that is renegotiated every sprint is not a Definition of Done.
|
|
60
|
+
|
|
61
|
+
## Red Flags
|
|
62
|
+
|
|
63
|
+
- "It's done, I just haven't run it yet": unverified work is not done.
|
|
64
|
+
- "Tests pass" used as a synonym for done while docs, regressions, or runtime verification are skipped.
|
|
65
|
+
- A different bar applied depending on deadline pressure.
|
|
66
|
+
- Acceptance criteria treated as the whole bar, with no standing quality floor.
|
|
67
|
+
- "Done" declared before human review on changes that need it.
|
|
@@ -0,0 +1,236 @@
|
|
|
1
|
+
# Performance Checklist
|
|
2
|
+
|
|
3
|
+
Quick reference checklist for web application performance. Use alongside the `performance-optimization` skill.
|
|
4
|
+
|
|
5
|
+
## Table of Contents
|
|
6
|
+
|
|
7
|
+
- [Core Web Vitals Targets](#core-web-vitals-targets)
|
|
8
|
+
- [TTFB Diagnosis](#ttfb-diagnosis)
|
|
9
|
+
- [Frontend Checklist](#frontend-checklist)
|
|
10
|
+
- [Backend Checklist](#backend-checklist)
|
|
11
|
+
- [Caching Strategies](#caching-strategies)
|
|
12
|
+
- [Measurement Commands](#measurement-commands)
|
|
13
|
+
- [Common Anti-Patterns](#common-anti-patterns)
|
|
14
|
+
|
|
15
|
+
## Core Web Vitals Targets
|
|
16
|
+
|
|
17
|
+
| Metric | Good | Needs Work | Poor |
|
|
18
|
+
|--------|------|------------|------|
|
|
19
|
+
| LCP (Largest Contentful Paint) | ≤ 2.5s | ≤ 4.0s | > 4.0s |
|
|
20
|
+
| INP (Interaction to Next Paint) | ≤ 200ms | ≤ 500ms | > 500ms |
|
|
21
|
+
| CLS (Cumulative Layout Shift) | ≤ 0.1 | ≤ 0.25 | > 0.25 |
|
|
22
|
+
|
|
23
|
+
## TTFB Diagnosis
|
|
24
|
+
|
|
25
|
+
When TTFB is slow (> 800ms), check each component in DevTools Network waterfall:
|
|
26
|
+
|
|
27
|
+
- [ ] **DNS resolution** slow → add `<link rel="dns-prefetch">` or `<link rel="preconnect">` for known origins
|
|
28
|
+
- [ ] **TCP/TLS handshake** slow → enable HTTP/2, consider edge deployment, verify keep-alive
|
|
29
|
+
- [ ] **Server processing** slow → profile backend, check slow queries, add caching
|
|
30
|
+
|
|
31
|
+
## Frontend Checklist
|
|
32
|
+
|
|
33
|
+
### Images
|
|
34
|
+
- [ ] Images use modern formats (WebP, AVIF)
|
|
35
|
+
- [ ] Images are responsively sized (`srcset` and `sizes`)
|
|
36
|
+
- [ ] Images and `<source>` elements have explicit `width` and `height` (prevents CLS in art direction)
|
|
37
|
+
- [ ] Below-the-fold images use `loading="lazy"` and `decoding="async"`
|
|
38
|
+
- [ ] Hero/LCP images use `fetchpriority="high"` and no lazy loading
|
|
39
|
+
|
|
40
|
+
### JavaScript
|
|
41
|
+
- [ ] Bundle size under 200KB gzipped (initial load)
|
|
42
|
+
- [ ] Code splitting with dynamic `import()` for routes and heavy features
|
|
43
|
+
- [ ] Tree shaking enabled (verify dependency ships ESM and marks `sideEffects: false`)
|
|
44
|
+
- [ ] No blocking JavaScript in `<head>` (use `defer` or `async`)
|
|
45
|
+
- [ ] Heavy computation offloaded to Web Workers (if applicable)
|
|
46
|
+
- [ ] `React.memo()` on expensive components that re-render with same props
|
|
47
|
+
- [ ] `useMemo()` / `useCallback()` only where profiling shows benefit
|
|
48
|
+
- [ ] Long tasks (> 50ms) broken up to keep the main thread available — main lever for INP
|
|
49
|
+
- [ ] `yieldToMain` pattern used inside long-running loops so input events can run between chunks
|
|
50
|
+
- [ ] Modern scheduling APIs used where available: `scheduler.yield()` (preferred), `scheduler.postTask()` with priorities, `isInputPending()` to yield only when needed
|
|
51
|
+
- [ ] `requestIdleCallback` for deferrable, non-urgent work (analytics flush, prefetch, warmup)
|
|
52
|
+
- [ ] Non-critical work deferred out of event handlers (e.g. analytics, logging) so the response to the interaction is not delayed
|
|
53
|
+
- [ ] Third-party scripts loaded with `async` / `defer`, audited for size, and fronted by a facade when heavy (chat widgets, embeds)
|
|
54
|
+
|
|
55
|
+
### CSS
|
|
56
|
+
- [ ] Critical CSS inlined or preloaded
|
|
57
|
+
- [ ] No render-blocking CSS for non-critical styles
|
|
58
|
+
- [ ] No CSS-in-JS runtime cost in production (use extraction)
|
|
59
|
+
|
|
60
|
+
### Fonts
|
|
61
|
+
- [ ] Limited to 2–3 font families, 2–3 weights each (every additional weight is another request)
|
|
62
|
+
- [ ] WOFF2 format only (smallest, universal support — skip WOFF/TTF/EOT)
|
|
63
|
+
- [ ] Self-hosted when possible (third-party font CDNs add DNS + TCP + TLS round-trips)
|
|
64
|
+
- [ ] LCP-critical fonts preloaded: `<link rel="preload" as="font" type="font/woff2" crossorigin>`
|
|
65
|
+
- [ ] `font-display: swap` (or `optional` for non-critical) to avoid FOIT blocking render
|
|
66
|
+
- [ ] Subsetted via `unicode-range` to ship only the glyphs each page needs
|
|
67
|
+
- [ ] Variable fonts considered when multiple weights/styles are required (one file replaces many)
|
|
68
|
+
- [ ] Fallback font metrics adjusted with `size-adjust`, `ascent-override`, `descent-override` to reduce CLS on font swap
|
|
69
|
+
- [ ] System font stack considered before any custom font
|
|
70
|
+
|
|
71
|
+
### Network
|
|
72
|
+
- [ ] Static assets cached with long `max-age` + content hashing
|
|
73
|
+
- [ ] API responses cached where appropriate (`Cache-Control`)
|
|
74
|
+
- [ ] HTTP/2 or HTTP/3 enabled
|
|
75
|
+
- [ ] Resources preconnected (`<link rel="preconnect">`) for known origins
|
|
76
|
+
- [ ] `fetchpriority` used on critical non-image resources (e.g., key `<link rel="preload">`, above-the-fold `<script>`) — not only on `<img>`
|
|
77
|
+
- [ ] No unnecessary redirects
|
|
78
|
+
|
|
79
|
+
### Rendering
|
|
80
|
+
- [ ] No layout thrashing (forced synchronous layouts)
|
|
81
|
+
- [ ] Animations use `transform` and `opacity` (GPU-accelerated)
|
|
82
|
+
- [ ] Long lists use virtualization (e.g., `react-window`)
|
|
83
|
+
- [ ] No unnecessary full-page re-renders
|
|
84
|
+
- [ ] Off-screen sections use `content-visibility: auto` with `contain-intrinsic-size` to skip layout/paint of non-visible areas
|
|
85
|
+
- [ ] No `unload` event handlers and no `Cache-Control: no-store` on HTML responses — preserves back/forward cache (bfcache) eligibility
|
|
86
|
+
|
|
87
|
+
## Backend Checklist
|
|
88
|
+
|
|
89
|
+
### Database
|
|
90
|
+
- [ ] No N+1 query patterns (use eager loading / joins)
|
|
91
|
+
- [ ] Queries have appropriate indexes
|
|
92
|
+
- [ ] List endpoints paginated (never `SELECT * FROM table`)
|
|
93
|
+
- [ ] Connection pooling configured
|
|
94
|
+
- [ ] Slow query logging enabled
|
|
95
|
+
|
|
96
|
+
#### Query plans
|
|
97
|
+
- [ ] `EXPLAIN ANALYZE` captured **before** the fix, not just after — it is the baseline
|
|
98
|
+
- [ ] `Seq Scan` on a large table understood: index missing, unusable, or genuinely not worth it
|
|
99
|
+
- [ ] Estimated vs actual `rows=` within an order of magnitude (if not, refresh statistics before touching indexes)
|
|
100
|
+
- [ ] No `Sort` node that a composite index could absorb
|
|
101
|
+
- [ ] Plan re-checked after the change — an index that did not change the plan gets reverted
|
|
102
|
+
|
|
103
|
+
#### Index strategy
|
|
104
|
+
- [ ] Composite index column order is equality first, then range/sort
|
|
105
|
+
- [ ] Index covers the query shape (filter + sort), not just one column in isolation
|
|
106
|
+
- [ ] Covering index considered for hot read paths (index-only scan avoids the heap fetch)
|
|
107
|
+
- [ ] Not indexing low-selectivity columns *for the dominant value*; a partial index still serves the rare-value query (`WHERE status = 'failed'`)
|
|
108
|
+
- [ ] Expression index used where the query applies a function (`lower(email)`)
|
|
109
|
+
- [ ] Full-text or trigram index used for leading-wildcard search, not a B-tree
|
|
110
|
+
- [ ] Write cost measured on write-heavy tables (every index taxes every `INSERT`/`UPDATE`)
|
|
111
|
+
- [ ] Unused and duplicate indexes dropped (they cost writes and buy nothing)
|
|
112
|
+
|
|
113
|
+
#### Connection pooling
|
|
114
|
+
- [ ] One pool per process, not per request or per module
|
|
115
|
+
- [ ] `instances × pool max` stays under the database's `max_connections`
|
|
116
|
+
- [ ] `connectionTimeoutMillis` set so exhaustion fails fast instead of queueing forever
|
|
117
|
+
- [ ] Exhaustion diagnosed before resizing: find what holds connections (long transactions, missing `await`, leaked clients)
|
|
118
|
+
- [ ] Serverless / autoscaling fronted by a multiplexing proxy (pgbouncer, RDS Proxy) rather than a larger pool
|
|
119
|
+
|
|
120
|
+
### API
|
|
121
|
+
- [ ] Response times < 200ms (p95)
|
|
122
|
+
- [ ] No synchronous heavy computation in request handlers
|
|
123
|
+
- [ ] Bulk operations instead of loops of individual calls
|
|
124
|
+
- [ ] Response compression (gzip/brotli)
|
|
125
|
+
- [ ] Appropriate caching (in-memory, Redis, CDN)
|
|
126
|
+
|
|
127
|
+
### Infrastructure
|
|
128
|
+
- [ ] CDN for static assets
|
|
129
|
+
- [ ] Server located close to users (or edge deployment)
|
|
130
|
+
- [ ] Horizontal scaling configured (if needed)
|
|
131
|
+
- [ ] Health check endpoint for load balancer
|
|
132
|
+
|
|
133
|
+
## Caching Strategies
|
|
134
|
+
|
|
135
|
+
The decision material (which layer, which invalidation strategy, what never to cache) lives in the `performance-optimization` skill. This section covers the read/write patterns and the checklist.
|
|
136
|
+
|
|
137
|
+
### Read and write patterns
|
|
138
|
+
|
|
139
|
+
| Pattern | How it works | Use when | Watch out for |
|
|
140
|
+
|---|---|---|---|
|
|
141
|
+
| **Cache-aside** (lazy) | App checks cache, on miss reads origin and populates | Default choice; read-heavy, tolerant of a cold first hit | Every miss hits the origin, so it needs stampede protection |
|
|
142
|
+
| **Read-through** | Cache layer itself loads on miss | You want the load path in one place, not at every call site | Hides origin latency; a slow origin looks like a slow cache |
|
|
143
|
+
| **Write-through** | Write goes to cache and origin together, synchronously | Reads must never see a stale value after a write | Adds cache latency to every write |
|
|
144
|
+
| **Write-behind** (write-back) | Write hits cache, origin updated asynchronously | Write-heavy, and the origin is the bottleneck | Data loss window if the cache dies before the flush. Needs durability you can defend |
|
|
145
|
+
|
|
146
|
+
### Negative caching
|
|
147
|
+
|
|
148
|
+
Cache the *absence* of a result too. A key that misses on every lookup (a nonexistent user ID probed in a loop, a 404 asset) sends every request to the origin, which is a cache that only protects the happy path.
|
|
149
|
+
|
|
150
|
+
- Store an explicit "not found" sentinel with a **shorter** TTL than positive entries
|
|
151
|
+
- Keep the negative TTL short enough that a newly created record appears promptly
|
|
152
|
+
- Never let an origin *error* become a negative cache entry, or one failing minute becomes many
|
|
153
|
+
|
|
154
|
+
### Request coalescing (stampede protection)
|
|
155
|
+
|
|
156
|
+
One recompute, N waiters. Prevents a hot key's expiry from delivering the full concurrent load to the origin:
|
|
157
|
+
|
|
158
|
+
```typescript
|
|
159
|
+
const inFlight = new Map<string, Promise<unknown>>();
|
|
160
|
+
|
|
161
|
+
function loadOnce<T>(key: string, fetcher: () => Promise<T>): Promise<T> {
|
|
162
|
+
const existing = inFlight.get(key) as Promise<T> | undefined;
|
|
163
|
+
if (existing) return existing;
|
|
164
|
+
const p = fetcher().finally(() => inFlight.delete(key));
|
|
165
|
+
inFlight.set(key, p);
|
|
166
|
+
return p;
|
|
167
|
+
}
|
|
168
|
+
```
|
|
169
|
+
|
|
170
|
+
For a shared cache, the same idea needs a distributed lock, or `stale-while-revalidate` so waiters serve the stale value instead of blocking.
|
|
171
|
+
|
|
172
|
+
### Cache checklist
|
|
173
|
+
- [ ] The cached call was measured as expensive first (caching a fast call adds a hop and buys nothing)
|
|
174
|
+
- [ ] Read/write ratio justifies the cache (re-read far more often than written)
|
|
175
|
+
- [ ] Cache key includes every input the response varies on: tenant, viewer, locale, permissions, feature flags
|
|
176
|
+
- [ ] No per-user data cached under a key that does not identify the user
|
|
177
|
+
- [ ] One invalidation strategy chosen (TTL, event/tag, or versioned keys), not an accidental mix
|
|
178
|
+
- [ ] Acceptable staleness window written down, not implied by whatever TTL was typed
|
|
179
|
+
- [ ] Stampede protection on hot keys (coalescing, lock, or `stale-while-revalidate`)
|
|
180
|
+
- [ ] Negative results cached with a shorter TTL; origin errors never cached
|
|
181
|
+
- [ ] Eviction policy and memory ceiling set (an unbounded cache is a memory leak)
|
|
182
|
+
- [ ] Hit rate monitored — a cache nobody measures is an assumption, and a low hit rate is pure overhead
|
|
183
|
+
- [ ] Nothing cached whose staleness is a correctness bug (balances, permissions, inventory at checkout)
|
|
184
|
+
|
|
185
|
+
## Measurement Commands
|
|
186
|
+
|
|
187
|
+
### INP field data and DevTools workflow
|
|
188
|
+
|
|
189
|
+
1. **Field data first** — check [CrUX Vis](https://developer.chrome.com/docs/crux/vis) or your RUM tool for real-user INP before optimising
|
|
190
|
+
2. **Identify slow interactions** — open DevTools → Performance panel → record while interacting; look for long tasks triggered by clicks/keystrokes
|
|
191
|
+
3. **Test on mid-range Android** — INP issues often only surface on slower hardware; use a real device or DevTools CPU throttling (4×–6× slowdown)
|
|
192
|
+
|
|
193
|
+
```bash
|
|
194
|
+
# Lighthouse CLI
|
|
195
|
+
npx lighthouse https://localhost:3000 --output json --output-path ./report.json
|
|
196
|
+
|
|
197
|
+
# Bundle analysis
|
|
198
|
+
npx webpack-bundle-analyzer stats.json
|
|
199
|
+
# or for Vite:
|
|
200
|
+
npx vite-bundle-visualizer
|
|
201
|
+
|
|
202
|
+
# Check bundle size
|
|
203
|
+
npx bundlesize
|
|
204
|
+
|
|
205
|
+
# Web Vitals in code
|
|
206
|
+
import { onLCP, onINP, onCLS } from 'web-vitals';
|
|
207
|
+
onLCP(console.log);
|
|
208
|
+
onINP(console.log);
|
|
209
|
+
onCLS(console.log);
|
|
210
|
+
|
|
211
|
+
# INP with interaction-level detail (attribution build)
|
|
212
|
+
import { onINP } from 'web-vitals/attribution';
|
|
213
|
+
onINP(({ value, attribution }) => {
|
|
214
|
+
const { interactionTarget, inputDelay, processingDuration, presentationDelay } = attribution;
|
|
215
|
+
console.log({ value, interactionTarget, inputDelay, processingDuration, presentationDelay });
|
|
216
|
+
});
|
|
217
|
+
```
|
|
218
|
+
|
|
219
|
+
## Common Anti-Patterns
|
|
220
|
+
|
|
221
|
+
| Anti-Pattern | Impact | Fix |
|
|
222
|
+
|---|---|---|
|
|
223
|
+
| N+1 queries | Linear DB load growth | Use joins, includes, or batch loading |
|
|
224
|
+
| Unbounded queries | Memory exhaustion, timeouts | Always paginate, add LIMIT |
|
|
225
|
+
| Missing indexes | Slow reads as data grows | Add indexes for filtered/sorted columns |
|
|
226
|
+
| Indexing without reading the plan | Write cost paid, read gain unproven | `EXPLAIN ANALYZE` before and after; revert if the plan is unchanged |
|
|
227
|
+
| Redundant / unused indexes | Every write pays for them | Audit usage stats, drop what nothing reads |
|
|
228
|
+
| Connection pool per request | Exhausts `max_connections` under load | One pool per process; proxy for serverless |
|
|
229
|
+
| Cache key missing the viewer | One user's data served to another | Key on tenant, viewer, locale, permissions |
|
|
230
|
+
| Unbounded cache | Memory leak wearing an optimization's clothing | Set eviction policy and a memory ceiling |
|
|
231
|
+
| Cache stampede on a hot key | Origin takes full concurrent load at expiry | Coalesce misses, or `stale-while-revalidate` |
|
|
232
|
+
| Layout thrashing | Jank, dropped frames | Batch DOM reads, then batch writes |
|
|
233
|
+
| Unoptimized images | Slow LCP, wasted bandwidth | Use WebP, responsive sizes, lazy load |
|
|
234
|
+
| Large bundles | Slow Time to Interactive | Code split, tree shake, audit deps |
|
|
235
|
+
| Blocking main thread | Poor INP, unresponsive UI | Chunk long tasks with `scheduler.yield()` / `yieldToMain`, offload to Web Workers |
|
|
236
|
+
| Memory leaks | Growing memory, eventual crash | Clean up listeners, intervals, refs |
|