@hecer/yoke 1.6.0 → 1.6.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (103) hide show
  1. package/.claude-plugin/plugin.json +13 -13
  2. package/.codex-plugin/plugin.json +7 -7
  3. package/CHANGELOG.md +294 -288
  4. package/README.md +874 -874
  5. package/TODOS.md +5 -5
  6. package/agents/docs.toml +6 -6
  7. package/agents/implementer.toml +6 -6
  8. package/agents/reviewer.toml +6 -6
  9. package/agents/security.toml +6 -6
  10. package/bench/README.md +86 -86
  11. package/bench/RESULTS.md +35 -35
  12. package/bench/output-compaction.mjs +65 -65
  13. package/bench/result-schema.mjs +12 -12
  14. package/bench/results/claude-2026-07-27T18-03-26.json +50 -50
  15. package/bench/results/codex-unavailable-1785175418318.json +15 -15
  16. package/bench/results/gemini-2026-07-27T18-03-44.json +46 -46
  17. package/bench/run-matrix.mjs +26 -26
  18. package/bench/run.mjs +106 -106
  19. package/canon/AGENTS.md +30 -30
  20. package/canon/context/DECISIONS.md +4 -4
  21. package/canon/context/GLOSSARY.md +11 -11
  22. package/canon/context/KNOWLEDGE.md +4 -4
  23. package/canon/context/PROJECT.md +15 -15
  24. package/canon/loop/loop-spec.md +65 -65
  25. package/canon/loop/prd.schema.md +43 -43
  26. package/canon/manifest.yaml +59 -59
  27. package/canon/policy/gates.md +7 -7
  28. package/canon/policy/roles.md +9 -9
  29. package/canon/skills/ATTRIBUTION.md +99 -99
  30. package/canon/skills/authoring-prd/SKILL.md +58 -58
  31. package/canon/skills/brainstorming/SKILL.md +164 -164
  32. package/canon/skills/codebase-design/DEEPENING.md +15 -15
  33. package/canon/skills/codebase-design/DESIGN-IT-TWICE.md +12 -12
  34. package/canon/skills/codebase-design/SKILL.md +39 -39
  35. package/canon/skills/dispatching-parallel-agents/SKILL.md +182 -182
  36. package/canon/skills/document-release/SKILL.md +302 -302
  37. package/canon/skills/domain-modeling/ADR-FORMAT.md +19 -19
  38. package/canon/skills/domain-modeling/CONTEXT-FORMAT.md +39 -39
  39. package/canon/skills/domain-modeling/SKILL.md +35 -35
  40. package/canon/skills/executing-plans/SKILL.md +70 -70
  41. package/canon/skills/finishing-a-development-branch/SKILL.md +200 -200
  42. package/canon/skills/health/SKILL.md +177 -177
  43. package/canon/skills/maintaining-context/SKILL.md +34 -34
  44. package/canon/skills/minimal-code/SKILL.md +21 -21
  45. package/canon/skills/no-ai-slop/SKILL.md +103 -103
  46. package/canon/skills/no-ai-slop/eval.md +43 -43
  47. package/canon/skills/plan-ceo-review/SKILL.md +541 -541
  48. package/canon/skills/plan-eng-review/SKILL.md +362 -362
  49. package/canon/skills/receiving-code-review/SKILL.md +213 -213
  50. package/canon/skills/requesting-code-review/SKILL.md +105 -105
  51. package/canon/skills/resolving-merge-conflicts/SKILL.md +18 -18
  52. package/canon/skills/retro/SKILL.md +397 -397
  53. package/canon/skills/review/SKILL.md +246 -246
  54. package/canon/skills/ship/SKILL.md +691 -691
  55. package/canon/skills/subagent-driven-development/SKILL.md +277 -277
  56. package/canon/skills/systematic-debugging/SKILL.md +296 -296
  57. package/canon/skills/tdd/SKILL.md +371 -371
  58. package/canon/skills/unslop-ui/SKILL.md +34 -34
  59. package/canon/skills/using-git-worktrees/SKILL.md +218 -218
  60. package/canon/skills/verification-before-completion/SKILL.md +139 -139
  61. package/canon/skills/visual-verification/SKILL.md +54 -54
  62. package/canon/skills/workflow/SKILL.md +22 -22
  63. package/canon/skills/writing-for-agents/SKILL-MECHANICS.md +27 -27
  64. package/canon/skills/writing-for-agents/SKILL.md +42 -42
  65. package/canon/skills/writing-plans/SKILL.md +152 -152
  66. package/canon/skills/writing-skills/SKILL.md +655 -655
  67. package/canon/skills/yoke-retrofit/SKILL.md +26 -26
  68. package/canon/skills/yoke-workflow/SKILL.md +20 -20
  69. package/canon/tools/codex-rtk-hook.mjs +35 -35
  70. package/canon/tools/graphify.md +3 -3
  71. package/canon/tools/playwright-mcp.md +3 -3
  72. package/canon/tools/rtk.md +7 -7
  73. package/canon/tools/serena.md +6 -6
  74. package/dist/agents/process.js +3 -0
  75. package/dist/loop/watchdog.js +1 -1
  76. package/dist/prd/command.js +17 -17
  77. package/dist/retrofit/planners/claude.js +14 -14
  78. package/dist/retrofit/preserve.js +2 -2
  79. package/docs/MIGRATING-TO-1.0.md +33 -33
  80. package/docs/MIGRATING-TO-1.1.md +27 -27
  81. package/docs/MIGRATING-TO-1.4.md +70 -70
  82. package/docs/PUBLISHING.md +91 -91
  83. package/docs/superpowers/plans/2026-06-28-baustein-e-context-layer.md +981 -981
  84. package/docs/superpowers/plans/2026-06-29-baustein-f-routing.md +258 -258
  85. package/docs/superpowers/plans/2026-06-29-baustein-g-loop-observability.md +1006 -1006
  86. package/docs/superpowers/plans/2026-06-29-baustein-h-loop-robustness.md +374 -374
  87. package/docs/superpowers/plans/2026-06-30-baustein-i-visual-design-verification.md +450 -450
  88. package/docs/superpowers/plans/2026-07-02-baustein-k-zero-to-100-bootstrap.md +1024 -1024
  89. package/docs/superpowers/plans/2026-07-02-baustein-m-flow-smoke-proofs.md +574 -574
  90. package/docs/superpowers/plans/2026-08-13-gauntlet-quality-loop.md +537 -537
  91. package/docs/superpowers/plans/2026-08-16-artifact-backed-output-compaction.md +329 -329
  92. package/docs/superpowers/specs/2026-06-28-baustein-e-context-layer-design.md +146 -146
  93. package/docs/superpowers/specs/2026-06-29-baustein-f-routing-design.md +106 -106
  94. package/docs/superpowers/specs/2026-06-29-baustein-g-loop-observability-design.md +186 -186
  95. package/docs/superpowers/specs/2026-06-29-baustein-h-loop-robustness-design.md +113 -113
  96. package/docs/superpowers/specs/2026-06-30-baustein-i-visual-design-verification-design.md +98 -98
  97. package/docs/superpowers/specs/2026-07-02-baustein-k-zero-to-100-bootstrap-design.md +200 -200
  98. package/docs/superpowers/specs/2026-07-02-baustein-m-flow-smoke-proofs-design.md +155 -155
  99. package/docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md +422 -422
  100. package/docs/superpowers/specs/2026-08-16-artifact-backed-output-compaction-design.md +166 -166
  101. package/gemini-extension.json +6 -6
  102. package/hooks/hooks.json +19 -19
  103. package/package.json +87 -87
@@ -1,397 +1,397 @@
1
- ---
2
- name: retro
3
- description: |
4
- Weekly engineering retrospective. Analyzes commit history, work patterns, and code quality
5
- metrics for the time window. Team-aware: identifies the user, then analyzes every
6
- contributor with per-person praise and growth opportunities.
7
- Use when asked for a "retro", "engineering retrospective", or "weekly summary".
8
- triggers:
9
- - retro
10
- - engineering retrospective
11
- - weekly retro
12
- - weekly summary
13
- ---
14
-
15
- # Engineering Retrospective
16
-
17
- You are running the `retro` skill. Generate a comprehensive engineering retrospective analyzing commit history, work patterns, and code quality metrics.
18
-
19
- ## Arguments
20
-
21
- - (none) — default: last 7 days
22
- - `24h` — last 24 hours
23
- - `14d` — last 14 days
24
- - `30d` — last 30 days
25
- - `compare` — compare current window vs prior same-length window
26
- - `compare 14d` — compare with explicit window
27
-
28
- ## Instructions
29
-
30
- Parse the argument to determine the time window. Default to 7 days if no argument given. All times should be reported in the user's **local timezone** (use the system default — do NOT set `TZ`).
31
-
32
- **Midnight-aligned windows:** For day (`d`) and week (`w`) units, compute an absolute start date at local midnight, not a relative string. For example, if today is 2026-03-18 and the window is 7 days: the start date is 2026-03-11. Use `--since="2026-03-11T00:00:00"` for git log queries — the explicit `T00:00:00` suffix ensures git starts from midnight. For week units, multiply by 7 to get days. For hour (`h`) units, use `--since="N hours ago"`.
33
-
34
- **Argument validation:** If the argument doesn't match a number followed by `d`, `h`, or `w`, or the word `compare` (optionally followed by a window), show usage and stop:
35
- ```
36
- Usage: retro [window | compare]
37
- retro — last 7 days (default)
38
- retro 24h — last 24 hours
39
- retro 14d — last 14 days
40
- retro 30d — last 30 days
41
- retro compare — compare this period vs prior period
42
- retro compare 14d — compare with explicit window
43
- ```
44
-
45
- ### Step 1: Gather Raw Data
46
-
47
- First, fetch origin and identify the current user:
48
- ```bash
49
- git fetch origin <default> --quiet
50
- git config user.name
51
- git config user.email
52
- ```
53
-
54
- The name returned by `git config user.name` is **"you"** — the person reading this retro. All other authors are teammates.
55
-
56
- Run ALL of these git commands in parallel (they are independent):
57
-
58
- ```bash
59
- # 1. All commits in window with timestamps, subject, hash, author, files changed
60
- git log origin/<default> --since="<window>" --format="%H|%aN|%ae|%ai|%s" --shortstat
61
-
62
- # 2. Per-commit test vs total LOC breakdown with author
63
- git log origin/<default> --since="<window>" --format="COMMIT:%H|%aN" --numstat
64
-
65
- # 3. Commit timestamps for session detection and hourly distribution (with author)
66
- git log origin/<default> --since="<window>" --format="%at|%aN|%ai|%s" | sort -n
67
-
68
- # 4. Files most frequently changed (hotspot analysis)
69
- git log origin/<default> --since="<window>" --format="" --name-only | grep -v '^$' | sort | uniq -c | sort -rn
70
-
71
- # 5. PR/MR numbers from commit messages (GitHub #NNN, GitLab !NNN)
72
- git log origin/<default> --since="<window>" --format="%s" | grep -oE '[#!][0-9]+' | sort | uniq
73
-
74
- # 6. Per-author file hotspots (who touches what)
75
- git log origin/<default> --since="<window>" --format="AUTHOR:%aN" --name-only
76
-
77
- # 7. Per-author commit counts (quick summary)
78
- git shortlog origin/<default> --since="<window>" -sn --no-merges
79
-
80
- # 8. TODOS.md backlog (if available)
81
- cat TODOS.md 2>/dev/null || true
82
-
83
- # 9. Test file count
84
- find . -name '*.test.*' -o -name '*.spec.*' -o -name '*_test.*' -o -name '*_spec.*' 2>/dev/null | grep -v node_modules | wc -l
85
-
86
- # 10. Regression test commits in window
87
- git log origin/<default> --since="<window>" --oneline --grep="test:" --grep="regression"
88
-
89
- # 11. Test files changed in window
90
- git log origin/<default> --since="<window>" --format="" --name-only | grep -E '\.(test|spec)\.' | sort -u | wc -l
91
- ```
92
-
93
- ### Step 2: Compute Metrics
94
-
95
- Calculate and present these metrics in a summary table:
96
-
97
- | Metric | Value |
98
- |--------|-------|
99
- | **Features shipped** (from CHANGELOG + merged PR titles) | N |
100
- | Commits to main | N |
101
- | Weighted commits (commits × avg files-touched, capped at 20 per commit) | N |
102
- | Contributors | N |
103
- | PRs merged | N |
104
- | **Logical SLOC added** (non-blank, non-comment — primary code-volume metric) | N |
105
- | Raw LOC: insertions | N |
106
- | Raw LOC: deletions | N |
107
- | Raw LOC: net | N |
108
- | Test LOC (insertions) | N |
109
- | Test LOC ratio | N% |
110
- | Version range | vX.Y.Z → vX.Y.Z |
111
- | Active days | N |
112
- | Detected sessions | N |
113
- | Avg raw LOC/session-hour | N |
114
- | Test Health | N total tests · M added this period · K regression tests |
115
-
116
- **Metric order rationale:** features shipped leads — what users got. Commits and weighted commits reflect intent-to-ship. Logical SLOC added reflects real new functionality. Raw LOC is demoted to context because AI inflates it.
117
-
118
- Then show a **per-author leaderboard** immediately below:
119
-
120
- ```
121
- Contributor Commits +/- Top area
122
- You (name) 32 +2400/-300 src/services/
123
- alice 12 +800/-150 app/api/
124
- bob 3 +120/-40 tests/
125
- ```
126
-
127
- Sort by commits descending. The current user (from `git config user.name`) always appears first, labeled "You (name)".
128
-
129
- **Backlog Health (if TODOS.md exists):** Compute:
130
- - Total open TODOs (exclude items in `## Completed` section)
131
- - P0/P1 count (critical/urgent items)
132
- - P2 count (important items)
133
- - Items completed this period (items in Completed section with dates within the retro window)
134
-
135
- Include in the metrics table:
136
- ```
137
- | Backlog Health | N open (X P0/P1, Y P2) · Z completed this period |
138
- ```
139
-
140
- ### Step 3: Commit Time Distribution
141
-
142
- Show hourly histogram in local time using bar chart:
143
-
144
- ```
145
- Hour Commits
146
- 00: 4 ████
147
- 07: 5 █████
148
- ...
149
- ```
150
-
151
- Identify and call out:
152
- - Peak hours
153
- - Dead zones
154
- - Whether pattern is bimodal (morning/evening) or continuous
155
- - Late-night coding clusters (after 10pm)
156
-
157
- ### Step 4: Work Session Detection
158
-
159
- Detect sessions using **45-minute gap** threshold between consecutive commits. For each session report:
160
- - Start/end time (local timezone)
161
- - Number of commits
162
- - Duration in minutes
163
-
164
- Classify sessions:
165
- - **Deep sessions** (50+ min)
166
- - **Medium sessions** (20-50 min)
167
- - **Micro sessions** (<20 min, typically single-commit fire-and-forget)
168
-
169
- Calculate:
170
- - Total active coding time (sum of session durations)
171
- - Average session length
172
- - LOC per hour of active time
173
-
174
- ### Step 5: Commit Type Breakdown
175
-
176
- Categorize by conventional commit prefix (feat/fix/refactor/test/chore/docs). Show as percentage bar:
177
-
178
- ```
179
- feat: 20 (40%) ████████████████████
180
- fix: 27 (54%) ███████████████████████████
181
- refactor: 2 ( 4%) ██
182
- ```
183
-
184
- Flag if fix ratio exceeds 50% — this signals a "ship fast, fix fast" pattern that may indicate review gaps.
185
-
186
- ### Step 6: Hotspot Analysis
187
-
188
- Show top 10 most-changed files. Flag:
189
- - Files changed 5+ times (churn hotspots)
190
- - Test files vs production files in the hotspot list
191
- - VERSION/CHANGELOG frequency (version discipline indicator)
192
-
193
- ### Step 7: PR Size Distribution
194
-
195
- From commit diffs, estimate PR sizes and bucket them:
196
- - **Small** (<100 LOC)
197
- - **Medium** (100-500 LOC)
198
- - **Large** (500-1500 LOC)
199
- - **XL** (1500+ LOC)
200
-
201
- ### Step 8: Focus Score + Ship of the Week
202
-
203
- **Focus score:** Calculate the percentage of commits touching the single most-changed top-level directory. Higher score = deeper focused work. Lower score = scattered context-switching. Report as: "Focus score: 62% (src/services/)"
204
-
205
- **Ship of the week:** Auto-identify the single highest-LOC PR in the window. Highlight it:
206
- - PR number and title
207
- - LOC changed
208
- - Why it matters (infer from commit messages and files touched)
209
-
210
- ### Step 9: Team Member Analysis
211
-
212
- For each contributor (including the current user), compute:
213
-
214
- 1. **Commits and LOC** — total commits, insertions, deletions, net LOC
215
- 2. **Areas of focus** — which directories/files they touched most (top 3)
216
- 3. **Commit type mix** — their personal feat/fix/refactor/test breakdown
217
- 4. **Session patterns** — when they code (their peak hours), session count
218
- 5. **Test discipline** — their personal test LOC ratio
219
- 6. **Biggest ship** — their single highest-impact commit or PR in the window
220
-
221
- **For the current user ("You"):** This section gets the deepest treatment. Include all the detail from the solo retro — session analysis, time patterns, focus score. Frame it in first person: "Your peak hours...", "Your biggest ship..."
222
-
223
- **For each teammate:** Write 2-3 sentences covering what they worked on and their pattern. Then:
224
-
225
- - **Praise** (1-2 specific things): Anchor in actual commits. Not "great work" — say exactly what was good.
226
- - **Opportunity for growth** (1 specific thing): Frame as a leveling-up suggestion, not criticism. Anchor in actual data.
227
-
228
- **If only one contributor (solo repo):** Skip the team breakdown — the retro is personal.
229
-
230
- **If there are Co-Authored-By trailers:** Parse `Co-Authored-By:` lines in commit messages. Credit those authors for the commit alongside the primary author. Note AI co-authors (e.g., `noreply@anthropic.com`) but do not include them as team members — instead, track "AI-assisted commits" as a separate metric.
231
-
232
- ### Step 10: Week-over-Week Trends (if window >= 14d)
233
-
234
- If the time window is 14 days or more, split into weekly buckets and show trends:
235
- - Commits per week (total and per-author)
236
- - LOC per week
237
- - Test ratio per week
238
- - Fix ratio per week
239
- - Session count per week
240
-
241
- ### Step 11: Streak Tracking
242
-
243
- Count consecutive days with at least 1 commit to origin/<default>, going back from today.
244
-
245
- ```bash
246
- # Team streak: all unique commit dates (local time) — no hard cutoff
247
- git log origin/<default> --format="%ad" --date=format:"%Y-%m-%d" | sort -u
248
-
249
- # Personal streak: only the current user's commits
250
- git log origin/<default> --author="<user_name>" --format="%ad" --date=format:"%Y-%m-%d" | sort -u
251
- ```
252
-
253
- Count backward from today — how many consecutive days have at least one commit? Display both:
254
- - "Team shipping streak: 47 consecutive days"
255
- - "Your shipping streak: 32 consecutive days"
256
-
257
- ### Step 12: Load History & Compare
258
-
259
- Before saving the new snapshot, check for prior retro history:
260
-
261
- ```bash
262
- ls -t .context/retros/*.json 2>/dev/null
263
- ```
264
-
265
- **If prior retros exist:** Load the most recent one using the Read tool. Calculate deltas for key metrics and include a **Trends vs Last Retro** section:
266
- ```
267
- Last Now Delta
268
- Test ratio: 22% → 41% ↑19pp
269
- Sessions: 10 → 14 ↑4
270
- LOC/hour: 200 → 350 ↑75%
271
- Fix ratio: 54% → 30% ↓24pp (improving)
272
- ```
273
-
274
- **If no prior retros exist:** Skip the comparison and append: "First retro recorded — run again next week to see trends."
275
-
276
- ### Step 13: Save Retro History
277
-
278
- After computing all metrics, save a JSON snapshot to `.context/retros/`:
279
-
280
- ```bash
281
- mkdir -p .context/retros
282
- # Filename: {today}-{sequence}.json
283
- ```
284
-
285
- Use the Write tool to save with this schema:
286
- ```json
287
- {
288
- "date": "2026-03-08",
289
- "window": "7d",
290
- "metrics": {
291
- "commits": 47,
292
- "contributors": 3,
293
- "prs_merged": 12,
294
- "insertions": 3200,
295
- "deletions": 800,
296
- "net_loc": 2400,
297
- "test_loc": 1300,
298
- "test_ratio": 0.41,
299
- "active_days": 6,
300
- "sessions": 14,
301
- "deep_sessions": 5,
302
- "avg_session_minutes": 42,
303
- "loc_per_session_hour": 350,
304
- "feat_pct": 0.40,
305
- "fix_pct": 0.30,
306
- "peak_hour": 22,
307
- "ai_assisted_commits": 32
308
- },
309
- "authors": {
310
- "Alice": { "commits": 32, "insertions": 2400, "deletions": 300, "test_ratio": 0.41, "top_area": "src/" }
311
- },
312
- "version_range": ["1.16.0", "1.16.1"],
313
- "streak_days": 47,
314
- "tweetable": "Week of Mar 1: 47 commits (3 contributors), 3.2k LOC, 38% tests, 12 PRs, peak: 10pm"
315
- }
316
- ```
317
-
318
- Only include `backlog` if TODOS.md exists. Only include `test_health` if test files were found.
319
-
320
- ### Step 14: Write the Narrative
321
-
322
- Structure the output as:
323
-
324
- ---
325
-
326
- **Tweetable summary** (first line, before everything else):
327
- ```
328
- Week of Mar 1: 47 commits (3 contributors), 3.2k LOC, 38% tests, 12 PRs, peak: 10pm | Streak: 47d
329
- ```
330
-
331
- ## Engineering Retro: [date range]
332
-
333
- ### Summary Table
334
- (from Step 2)
335
-
336
- ### Trends vs Last Retro
337
- (from Step 12, if prior retros exist — skip if first retro)
338
-
339
- ### Time & Session Patterns
340
- (from Steps 3-4)
341
-
342
- Narrative interpreting what the team-wide patterns mean:
343
- - When the most productive hours are and what drives them
344
- - Whether sessions are getting longer or shorter over time
345
- - Estimated hours per day of active coding (team aggregate)
346
-
347
- ### Shipping Velocity
348
- (from Steps 5-7)
349
-
350
- Narrative covering:
351
- - Commit type mix and what it reveals
352
- - PR size distribution and what it reveals about shipping cadence
353
- - Fix-chain detection (sequences of fix commits on the same subsystem)
354
-
355
- ### Code Quality Signals
356
- - Test LOC ratio trend
357
- - Hotspot analysis (are the same files churning?)
358
-
359
- ### Test Health
360
- - Total test files: N (from command 9)
361
- - Tests added this period: M (from command 11)
362
- - Regression test commits: commits matching test: or regression patterns
363
- - If test ratio < 20%: flag as growth area — "100% test coverage is the goal. Tests make coding safe."
364
-
365
- ### Focus & Highlights
366
- (from Step 8)
367
- - Focus score with interpretation
368
- - Ship of the week callout
369
-
370
- ### Your Week (personal deep-dive)
371
- (from Step 9, for the current user only)
372
-
373
- This is the section the user cares most about. Include:
374
- - Their personal commit count, LOC, test ratio
375
- - Their session patterns and peak hours
376
- - Their focus areas
377
- - Their biggest ship
378
- - **What you did well** (2-3 specific things anchored in commits)
379
- - **Where to level up** (1-2 specific, actionable suggestions)
380
-
381
- ### Team Breakdown
382
- (from Step 9, for each teammate — skip if solo repo)
383
-
384
- ### Top 3 Team Wins
385
- Identify the 3 highest-impact things shipped in the window across the whole team. For each:
386
- - What it was
387
- - Who shipped it
388
- - Why it matters (product/architecture impact)
389
-
390
- ### 3 Things to Improve
391
- Specific, actionable, anchored in actual commits. Mix personal and team-level suggestions. Phrase as "to get even better, the team could..."
392
-
393
- ### 3 Habits for Next Week
394
- Small, practical, realistic. Each must be something that takes <5 minutes to adopt. At least one should be team-oriented.
395
-
396
- ### Week-over-Week Trends
397
- (if applicable, from Step 10)
1
+ ---
2
+ name: retro
3
+ description: |
4
+ Weekly engineering retrospective. Analyzes commit history, work patterns, and code quality
5
+ metrics for the time window. Team-aware: identifies the user, then analyzes every
6
+ contributor with per-person praise and growth opportunities.
7
+ Use when asked for a "retro", "engineering retrospective", or "weekly summary".
8
+ triggers:
9
+ - retro
10
+ - engineering retrospective
11
+ - weekly retro
12
+ - weekly summary
13
+ ---
14
+
15
+ # Engineering Retrospective
16
+
17
+ You are running the `retro` skill. Generate a comprehensive engineering retrospective analyzing commit history, work patterns, and code quality metrics.
18
+
19
+ ## Arguments
20
+
21
+ - (none) — default: last 7 days
22
+ - `24h` — last 24 hours
23
+ - `14d` — last 14 days
24
+ - `30d` — last 30 days
25
+ - `compare` — compare current window vs prior same-length window
26
+ - `compare 14d` — compare with explicit window
27
+
28
+ ## Instructions
29
+
30
+ Parse the argument to determine the time window. Default to 7 days if no argument given. All times should be reported in the user's **local timezone** (use the system default — do NOT set `TZ`).
31
+
32
+ **Midnight-aligned windows:** For day (`d`) and week (`w`) units, compute an absolute start date at local midnight, not a relative string. For example, if today is 2026-03-18 and the window is 7 days: the start date is 2026-03-11. Use `--since="2026-03-11T00:00:00"` for git log queries — the explicit `T00:00:00` suffix ensures git starts from midnight. For week units, multiply by 7 to get days. For hour (`h`) units, use `--since="N hours ago"`.
33
+
34
+ **Argument validation:** If the argument doesn't match a number followed by `d`, `h`, or `w`, or the word `compare` (optionally followed by a window), show usage and stop:
35
+ ```
36
+ Usage: retro [window | compare]
37
+ retro — last 7 days (default)
38
+ retro 24h — last 24 hours
39
+ retro 14d — last 14 days
40
+ retro 30d — last 30 days
41
+ retro compare — compare this period vs prior period
42
+ retro compare 14d — compare with explicit window
43
+ ```
44
+
45
+ ### Step 1: Gather Raw Data
46
+
47
+ First, fetch origin and identify the current user:
48
+ ```bash
49
+ git fetch origin <default> --quiet
50
+ git config user.name
51
+ git config user.email
52
+ ```
53
+
54
+ The name returned by `git config user.name` is **"you"** — the person reading this retro. All other authors are teammates.
55
+
56
+ Run ALL of these git commands in parallel (they are independent):
57
+
58
+ ```bash
59
+ # 1. All commits in window with timestamps, subject, hash, author, files changed
60
+ git log origin/<default> --since="<window>" --format="%H|%aN|%ae|%ai|%s" --shortstat
61
+
62
+ # 2. Per-commit test vs total LOC breakdown with author
63
+ git log origin/<default> --since="<window>" --format="COMMIT:%H|%aN" --numstat
64
+
65
+ # 3. Commit timestamps for session detection and hourly distribution (with author)
66
+ git log origin/<default> --since="<window>" --format="%at|%aN|%ai|%s" | sort -n
67
+
68
+ # 4. Files most frequently changed (hotspot analysis)
69
+ git log origin/<default> --since="<window>" --format="" --name-only | grep -v '^$' | sort | uniq -c | sort -rn
70
+
71
+ # 5. PR/MR numbers from commit messages (GitHub #NNN, GitLab !NNN)
72
+ git log origin/<default> --since="<window>" --format="%s" | grep -oE '[#!][0-9]+' | sort | uniq
73
+
74
+ # 6. Per-author file hotspots (who touches what)
75
+ git log origin/<default> --since="<window>" --format="AUTHOR:%aN" --name-only
76
+
77
+ # 7. Per-author commit counts (quick summary)
78
+ git shortlog origin/<default> --since="<window>" -sn --no-merges
79
+
80
+ # 8. TODOS.md backlog (if available)
81
+ cat TODOS.md 2>/dev/null || true
82
+
83
+ # 9. Test file count
84
+ find . -name '*.test.*' -o -name '*.spec.*' -o -name '*_test.*' -o -name '*_spec.*' 2>/dev/null | grep -v node_modules | wc -l
85
+
86
+ # 10. Regression test commits in window
87
+ git log origin/<default> --since="<window>" --oneline --grep="test:" --grep="regression"
88
+
89
+ # 11. Test files changed in window
90
+ git log origin/<default> --since="<window>" --format="" --name-only | grep -E '\.(test|spec)\.' | sort -u | wc -l
91
+ ```
92
+
93
+ ### Step 2: Compute Metrics
94
+
95
+ Calculate and present these metrics in a summary table:
96
+
97
+ | Metric | Value |
98
+ |--------|-------|
99
+ | **Features shipped** (from CHANGELOG + merged PR titles) | N |
100
+ | Commits to main | N |
101
+ | Weighted commits (commits × avg files-touched, capped at 20 per commit) | N |
102
+ | Contributors | N |
103
+ | PRs merged | N |
104
+ | **Logical SLOC added** (non-blank, non-comment — primary code-volume metric) | N |
105
+ | Raw LOC: insertions | N |
106
+ | Raw LOC: deletions | N |
107
+ | Raw LOC: net | N |
108
+ | Test LOC (insertions) | N |
109
+ | Test LOC ratio | N% |
110
+ | Version range | vX.Y.Z → vX.Y.Z |
111
+ | Active days | N |
112
+ | Detected sessions | N |
113
+ | Avg raw LOC/session-hour | N |
114
+ | Test Health | N total tests · M added this period · K regression tests |
115
+
116
+ **Metric order rationale:** features shipped leads — what users got. Commits and weighted commits reflect intent-to-ship. Logical SLOC added reflects real new functionality. Raw LOC is demoted to context because AI inflates it.
117
+
118
+ Then show a **per-author leaderboard** immediately below:
119
+
120
+ ```
121
+ Contributor Commits +/- Top area
122
+ You (name) 32 +2400/-300 src/services/
123
+ alice 12 +800/-150 app/api/
124
+ bob 3 +120/-40 tests/
125
+ ```
126
+
127
+ Sort by commits descending. The current user (from `git config user.name`) always appears first, labeled "You (name)".
128
+
129
+ **Backlog Health (if TODOS.md exists):** Compute:
130
+ - Total open TODOs (exclude items in `## Completed` section)
131
+ - P0/P1 count (critical/urgent items)
132
+ - P2 count (important items)
133
+ - Items completed this period (items in Completed section with dates within the retro window)
134
+
135
+ Include in the metrics table:
136
+ ```
137
+ | Backlog Health | N open (X P0/P1, Y P2) · Z completed this period |
138
+ ```
139
+
140
+ ### Step 3: Commit Time Distribution
141
+
142
+ Show hourly histogram in local time using bar chart:
143
+
144
+ ```
145
+ Hour Commits
146
+ 00: 4 ████
147
+ 07: 5 █████
148
+ ...
149
+ ```
150
+
151
+ Identify and call out:
152
+ - Peak hours
153
+ - Dead zones
154
+ - Whether pattern is bimodal (morning/evening) or continuous
155
+ - Late-night coding clusters (after 10pm)
156
+
157
+ ### Step 4: Work Session Detection
158
+
159
+ Detect sessions using **45-minute gap** threshold between consecutive commits. For each session report:
160
+ - Start/end time (local timezone)
161
+ - Number of commits
162
+ - Duration in minutes
163
+
164
+ Classify sessions:
165
+ - **Deep sessions** (50+ min)
166
+ - **Medium sessions** (20-50 min)
167
+ - **Micro sessions** (<20 min, typically single-commit fire-and-forget)
168
+
169
+ Calculate:
170
+ - Total active coding time (sum of session durations)
171
+ - Average session length
172
+ - LOC per hour of active time
173
+
174
+ ### Step 5: Commit Type Breakdown
175
+
176
+ Categorize by conventional commit prefix (feat/fix/refactor/test/chore/docs). Show as percentage bar:
177
+
178
+ ```
179
+ feat: 20 (40%) ████████████████████
180
+ fix: 27 (54%) ███████████████████████████
181
+ refactor: 2 ( 4%) ██
182
+ ```
183
+
184
+ Flag if fix ratio exceeds 50% — this signals a "ship fast, fix fast" pattern that may indicate review gaps.
185
+
186
+ ### Step 6: Hotspot Analysis
187
+
188
+ Show top 10 most-changed files. Flag:
189
+ - Files changed 5+ times (churn hotspots)
190
+ - Test files vs production files in the hotspot list
191
+ - VERSION/CHANGELOG frequency (version discipline indicator)
192
+
193
+ ### Step 7: PR Size Distribution
194
+
195
+ From commit diffs, estimate PR sizes and bucket them:
196
+ - **Small** (<100 LOC)
197
+ - **Medium** (100-500 LOC)
198
+ - **Large** (500-1500 LOC)
199
+ - **XL** (1500+ LOC)
200
+
201
+ ### Step 8: Focus Score + Ship of the Week
202
+
203
+ **Focus score:** Calculate the percentage of commits touching the single most-changed top-level directory. Higher score = deeper focused work. Lower score = scattered context-switching. Report as: "Focus score: 62% (src/services/)"
204
+
205
+ **Ship of the week:** Auto-identify the single highest-LOC PR in the window. Highlight it:
206
+ - PR number and title
207
+ - LOC changed
208
+ - Why it matters (infer from commit messages and files touched)
209
+
210
+ ### Step 9: Team Member Analysis
211
+
212
+ For each contributor (including the current user), compute:
213
+
214
+ 1. **Commits and LOC** — total commits, insertions, deletions, net LOC
215
+ 2. **Areas of focus** — which directories/files they touched most (top 3)
216
+ 3. **Commit type mix** — their personal feat/fix/refactor/test breakdown
217
+ 4. **Session patterns** — when they code (their peak hours), session count
218
+ 5. **Test discipline** — their personal test LOC ratio
219
+ 6. **Biggest ship** — their single highest-impact commit or PR in the window
220
+
221
+ **For the current user ("You"):** This section gets the deepest treatment. Include all the detail from the solo retro — session analysis, time patterns, focus score. Frame it in first person: "Your peak hours...", "Your biggest ship..."
222
+
223
+ **For each teammate:** Write 2-3 sentences covering what they worked on and their pattern. Then:
224
+
225
+ - **Praise** (1-2 specific things): Anchor in actual commits. Not "great work" — say exactly what was good.
226
+ - **Opportunity for growth** (1 specific thing): Frame as a leveling-up suggestion, not criticism. Anchor in actual data.
227
+
228
+ **If only one contributor (solo repo):** Skip the team breakdown — the retro is personal.
229
+
230
+ **If there are Co-Authored-By trailers:** Parse `Co-Authored-By:` lines in commit messages. Credit those authors for the commit alongside the primary author. Note AI co-authors (e.g., `noreply@anthropic.com`) but do not include them as team members — instead, track "AI-assisted commits" as a separate metric.
231
+
232
+ ### Step 10: Week-over-Week Trends (if window >= 14d)
233
+
234
+ If the time window is 14 days or more, split into weekly buckets and show trends:
235
+ - Commits per week (total and per-author)
236
+ - LOC per week
237
+ - Test ratio per week
238
+ - Fix ratio per week
239
+ - Session count per week
240
+
241
+ ### Step 11: Streak Tracking
242
+
243
+ Count consecutive days with at least 1 commit to origin/<default>, going back from today.
244
+
245
+ ```bash
246
+ # Team streak: all unique commit dates (local time) — no hard cutoff
247
+ git log origin/<default> --format="%ad" --date=format:"%Y-%m-%d" | sort -u
248
+
249
+ # Personal streak: only the current user's commits
250
+ git log origin/<default> --author="<user_name>" --format="%ad" --date=format:"%Y-%m-%d" | sort -u
251
+ ```
252
+
253
+ Count backward from today — how many consecutive days have at least one commit? Display both:
254
+ - "Team shipping streak: 47 consecutive days"
255
+ - "Your shipping streak: 32 consecutive days"
256
+
257
+ ### Step 12: Load History & Compare
258
+
259
+ Before saving the new snapshot, check for prior retro history:
260
+
261
+ ```bash
262
+ ls -t .context/retros/*.json 2>/dev/null
263
+ ```
264
+
265
+ **If prior retros exist:** Load the most recent one using the Read tool. Calculate deltas for key metrics and include a **Trends vs Last Retro** section:
266
+ ```
267
+ Last Now Delta
268
+ Test ratio: 22% → 41% ↑19pp
269
+ Sessions: 10 → 14 ↑4
270
+ LOC/hour: 200 → 350 ↑75%
271
+ Fix ratio: 54% → 30% ↓24pp (improving)
272
+ ```
273
+
274
+ **If no prior retros exist:** Skip the comparison and append: "First retro recorded — run again next week to see trends."
275
+
276
+ ### Step 13: Save Retro History
277
+
278
+ After computing all metrics, save a JSON snapshot to `.context/retros/`:
279
+
280
+ ```bash
281
+ mkdir -p .context/retros
282
+ # Filename: {today}-{sequence}.json
283
+ ```
284
+
285
+ Use the Write tool to save with this schema:
286
+ ```json
287
+ {
288
+ "date": "2026-03-08",
289
+ "window": "7d",
290
+ "metrics": {
291
+ "commits": 47,
292
+ "contributors": 3,
293
+ "prs_merged": 12,
294
+ "insertions": 3200,
295
+ "deletions": 800,
296
+ "net_loc": 2400,
297
+ "test_loc": 1300,
298
+ "test_ratio": 0.41,
299
+ "active_days": 6,
300
+ "sessions": 14,
301
+ "deep_sessions": 5,
302
+ "avg_session_minutes": 42,
303
+ "loc_per_session_hour": 350,
304
+ "feat_pct": 0.40,
305
+ "fix_pct": 0.30,
306
+ "peak_hour": 22,
307
+ "ai_assisted_commits": 32
308
+ },
309
+ "authors": {
310
+ "Alice": { "commits": 32, "insertions": 2400, "deletions": 300, "test_ratio": 0.41, "top_area": "src/" }
311
+ },
312
+ "version_range": ["1.16.0", "1.16.1"],
313
+ "streak_days": 47,
314
+ "tweetable": "Week of Mar 1: 47 commits (3 contributors), 3.2k LOC, 38% tests, 12 PRs, peak: 10pm"
315
+ }
316
+ ```
317
+
318
+ Only include `backlog` if TODOS.md exists. Only include `test_health` if test files were found.
319
+
320
+ ### Step 14: Write the Narrative
321
+
322
+ Structure the output as:
323
+
324
+ ---
325
+
326
+ **Tweetable summary** (first line, before everything else):
327
+ ```
328
+ Week of Mar 1: 47 commits (3 contributors), 3.2k LOC, 38% tests, 12 PRs, peak: 10pm | Streak: 47d
329
+ ```
330
+
331
+ ## Engineering Retro: [date range]
332
+
333
+ ### Summary Table
334
+ (from Step 2)
335
+
336
+ ### Trends vs Last Retro
337
+ (from Step 12, if prior retros exist — skip if first retro)
338
+
339
+ ### Time & Session Patterns
340
+ (from Steps 3-4)
341
+
342
+ Narrative interpreting what the team-wide patterns mean:
343
+ - When the most productive hours are and what drives them
344
+ - Whether sessions are getting longer or shorter over time
345
+ - Estimated hours per day of active coding (team aggregate)
346
+
347
+ ### Shipping Velocity
348
+ (from Steps 5-7)
349
+
350
+ Narrative covering:
351
+ - Commit type mix and what it reveals
352
+ - PR size distribution and what it reveals about shipping cadence
353
+ - Fix-chain detection (sequences of fix commits on the same subsystem)
354
+
355
+ ### Code Quality Signals
356
+ - Test LOC ratio trend
357
+ - Hotspot analysis (are the same files churning?)
358
+
359
+ ### Test Health
360
+ - Total test files: N (from command 9)
361
+ - Tests added this period: M (from command 11)
362
+ - Regression test commits: commits matching test: or regression patterns
363
+ - If test ratio < 20%: flag as growth area — "100% test coverage is the goal. Tests make coding safe."
364
+
365
+ ### Focus & Highlights
366
+ (from Step 8)
367
+ - Focus score with interpretation
368
+ - Ship of the week callout
369
+
370
+ ### Your Week (personal deep-dive)
371
+ (from Step 9, for the current user only)
372
+
373
+ This is the section the user cares most about. Include:
374
+ - Their personal commit count, LOC, test ratio
375
+ - Their session patterns and peak hours
376
+ - Their focus areas
377
+ - Their biggest ship
378
+ - **What you did well** (2-3 specific things anchored in commits)
379
+ - **Where to level up** (1-2 specific, actionable suggestions)
380
+
381
+ ### Team Breakdown
382
+ (from Step 9, for each teammate — skip if solo repo)
383
+
384
+ ### Top 3 Team Wins
385
+ Identify the 3 highest-impact things shipped in the window across the whole team. For each:
386
+ - What it was
387
+ - Who shipped it
388
+ - Why it matters (product/architecture impact)
389
+
390
+ ### 3 Things to Improve
391
+ Specific, actionable, anchored in actual commits. Mix personal and team-level suggestions. Phrase as "to get even better, the team could..."
392
+
393
+ ### 3 Habits for Next Week
394
+ Small, practical, realistic. Each must be something that takes <5 minutes to adopt. At least one should be team-oriented.
395
+
396
+ ### Week-over-Week Trends
397
+ (if applicable, from Step 10)