@olegkoval/agent-skills 1.41.4 → 1.42.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (28) hide show
  1. package/.github/copilot-instructions.md +1 -0
  2. package/.github/prompts/agent-ops-retro.prompt.md +178 -0
  3. package/.kiro/steering/agent-ops-retro.md +177 -0
  4. package/.windsurf/rules/agent-ops-retro.md +176 -0
  5. package/README.md +3 -2
  6. package/adapters/claude/olko-reflection/skills/agent-ops-retro/SKILL.md +199 -0
  7. package/adapters/claude/olko-reflection/skills/agent-ops-retro/scripts/mine-transcripts.mjs +245 -0
  8. package/adapters/codex/olko-reflection/README.md +1 -0
  9. package/adapters/cursor/olko-reflection/skills/agent-ops-retro/SKILL.md +199 -0
  10. package/adapters/cursor/olko-reflection/skills/agent-ops-retro/scripts/mine-transcripts.mjs +245 -0
  11. package/adapters/grok/olko-reflection/skills/agent-ops-retro/SKILL.md +199 -0
  12. package/adapters/grok/olko-reflection/skills/agent-ops-retro/scripts/mine-transcripts.mjs +245 -0
  13. package/catalog/skills.json +24 -0
  14. package/package.json +1 -1
  15. package/plugins/olko-apple-kit/.claude-plugin/plugin.json +1 -1
  16. package/plugins/olko-creative/.claude-plugin/plugin.json +1 -1
  17. package/plugins/olko-garmin-kit/.claude-plugin/plugin.json +1 -1
  18. package/plugins/olko-git-tools/.claude-plugin/plugin.json +1 -1
  19. package/plugins/olko-github-pr/.claude-plugin/plugin.json +1 -1
  20. package/plugins/olko-obsidian/.claude-plugin/plugin.json +1 -1
  21. package/plugins/olko-product/.claude-plugin/plugin.json +1 -1
  22. package/plugins/olko-reflection/.claude-plugin/plugin.json +1 -1
  23. package/plugins/olko-reflection/skills/agent-ops-retro/SKILL.md +198 -0
  24. package/plugins/olko-reflection/skills/agent-ops-retro/scripts/mine-transcripts.mjs +245 -0
  25. package/plugins/olko-release/.claude-plugin/plugin.json +1 -1
  26. package/plugins/olko-skill-meta/.claude-plugin/plugin.json +1 -1
  27. package/plugins/olko-web-ops/.claude-plugin/plugin.json +1 -1
  28. package/scripts/lib/catalog.mjs +1 -1
@@ -0,0 +1,199 @@
1
+ ---
2
+ name: agent-ops-retro
3
+ description: >
4
+ Retrospective on how the agents themselves are being operated, from local Claude Code
5
+ transcripts plus the reports other jobs already produce. Answers what is going well, what is
6
+ going wrong, what the human keeps having to repeat, and where delivery is leaking, then carries
7
+ every unactioned finding forward with its age. USE THIS AUTOMATICALLY (no need for the user to
8
+ name it) whenever the user asks to check all conversations, review agent usage, reduce the manual
9
+ administration of agents, find repeating corrections or issues across sessions, audit skills and
10
+ the knowledge ledger together, or asks what gaps exist in delivery over a recent window. This is
11
+ the operator axis: for repository history and shipped work use retro-analysis, for a single
12
+ finished change use wrap-up.
13
+ license: MIT
14
+ allowed-tools: Bash, Read, Write, Grep, Glob, Agent
15
+ compatibility: Claude Code. Reads ~/.claude transcripts and local report files, so it does not run on hosts without them.
16
+ metadata:
17
+ author: Oleg Koval
18
+ targets: [_source-only]
19
+ tags:
20
+ - retrospective
21
+ - agent-operations
22
+ - telemetry
23
+ - skills
24
+ - knowledge-ledger
25
+ - delegation
26
+ - cost
27
+ ---
28
+ <!-- Generated by scripts/build-adapters.sh. Do not edit directly. -->
29
+
30
+ # Agent Ops Retro
31
+
32
+ A retrospective on the operating layer, not the code. It asks whether the agents are being run
33
+ well: what the human had to repeat, what the guardrails caught, what they cost, and which findings
34
+ have been reported before and never actioned.
35
+
36
+ ## Scope boundary
37
+
38
+ Three retro skills exist and they do not overlap.
39
+
40
+ | Skill | Looks at | Answers |
41
+ |---|---|---|
42
+ | `retro-analysis` | Repository history, delivery evidence, code quality | What shipped and how the code is trending |
43
+ | `wrap-up` | One finished change | Did this specific piece of work meet its goal |
44
+ | `agent-ops-retro` | Transcripts, hooks, skills, ledger, delegation logs | How the agents are being operated |
45
+
46
+ If the user asks about shipped work, hand off to `retro-analysis` and say so.
47
+
48
+ ## Invocation
49
+
50
+ Accept a window: `14d` (default), `7d`, `30d`, or an explicit ISO date range. State the absolute
51
+ range at the top of the output. Convert relative dates to absolute.
52
+
53
+ ## Step 0. Read, do not recompute
54
+
55
+ Several jobs already report on this state. Adding a fifth independent opinion is the failure mode
56
+ this skill exists to fix, so read their outputs first and only compute what none of them cover.
57
+
58
+ | Source | Path | Gives you |
59
+ |---|---|---|
60
+ | Skill audit | `~/.claude/skill-audit/latest.json` and the dated `.txt` beside it | Installed vs never-invoked skills |
61
+ | Delegation log | `~/.claude/delegation-metrics/delegations.jsonl` | Model, tier, depth, outcome, rework per spawn |
62
+ | Tier check | `~/.claude/delegation-metrics/tier-check-*.md` | Opus share and spend against a baseline |
63
+ | Usage miner | `~/obsidian/cloud-opus/Lead/Automation/Weekly Automation Report *.md` | Repeated commands, proposal fatigue, approval queue |
64
+ | Miner state | `~/.claude/usage-miner/state/last-run.json` | What was staged, queued, blocked |
65
+ | Knowledge ledger | `~/.local/share/agent-context/repo/ledger.json` | Traps already recorded, so you do not re-derive them |
66
+ | Prior run | the previous output of this skill (see Step 6) | Findings to carry forward |
67
+
68
+ Two traps when reading these:
69
+
70
+ - A skill-usage figure counted from attributed messages is not an invocation count. The audit's own
71
+ numbers run three orders of magnitude above the Skill-tool invocation count. Say which unit you
72
+ are quoting, every time.
73
+ - When two sources disagree on the same number, report both with their sources and say the figure
74
+ is unreconciled. Never silently pick one.
75
+
76
+ ## Step 1. Mine the transcripts
77
+
78
+ This is the axis nothing else covers. Run `scripts/mine-transcripts.mjs` from this skill directory,
79
+ or reimplement it, over `~/.claude/projects/**/*.jsonl`.
80
+
81
+ Non-negotiable filters, each of which has burned a previous run:
82
+
83
+ - Filter by the record's own `.timestamp`, never by file mtime. A resumed session rewrites old
84
+ files and pulls pre-window turns into the window.
85
+ - A `user` record is not a human turn. Exclude `tool_result` blocks, `<system-reminder>` wrappers,
86
+ `<command-name>` and hook preambles, task notifications and bash echoes. In one measured corpus
87
+ 45 percent of `user` text records were tool echoes, so the raw count nearly doubles the real one.
88
+ - Exclude `isSidechain` records and anything under a `subagents/` directory from session and
89
+ human-turn counts. Count them separately for the delegation view.
90
+ - Empty output is a state, not a success. If a scan returns nothing, prove the scan works before
91
+ concluding the thing is absent.
92
+
93
+ Extract: organic human turns per day and per session, turn-length distribution, tool call counts,
94
+ `tool_result` errors with `is_error`, Agent spawns with `subagent_type` and `model`, Skill
95
+ invocations, interruptions, compaction events, per-model token usage including cache reads, and
96
+ synthetic records naming a rate or spend limit.
97
+
98
+ ## Step 2. Classify what the human repeated
99
+
100
+ Group the organic turns. The strongest automation signal is not raw frequency, it is the number of
101
+ distinct sessions a correction appears in: a thing said twenty times in one session is one
102
+ argument, a thing said once in twenty sessions is a missing default.
103
+
104
+ Report each recurring pattern as: name, count, distinct sessions, two or three verbatim dated
105
+ quotes, and the one concrete change that removes it. Name the mechanism, not the mood. Write
106
+ "stops for a one-word ack on reversible steps", not "autonomy could be better".
107
+
108
+ ## Step 3. Split what holds from what is broken
109
+
110
+ Produce two explicit lists, never a narrative.
111
+
112
+ **Verified as working:** the claim and what proved it. Guardrails that fired correctly count here.
113
+
114
+ **Defects:** the mechanism and its consequence. Include the governance frictions, which usually
115
+ outnumber the external failures: permission denials, classifier blocks, path-gate refusals, agent
116
+ output failing its own schema. Give each a count and a share of total errors.
117
+
118
+ ## Step 4. Delivery leakage
119
+
120
+ Delivery here means whether agent output landed, not whether it was produced.
121
+
122
+ - Open PRs authored by the user, bucketed by age, with fan-out batches identified as batches.
123
+ - Tickets whose completion timestamps cluster in minutes, which is a status backfill rather than
124
+ shipping. State both the raw completion count and the count after removing the cluster.
125
+ - Any fan-out with no named landing mechanism. That is backlog, not delivery.
126
+ - Scheduled jobs: check `launchctl list` exit codes and each job's own log. A wrapper exit code is
127
+ not completion.
128
+
129
+ ## Step 5. Close the knowledge circle
130
+
131
+ This is the half that makes the retro compound instead of repeating.
132
+
133
+ 1. Before writing a finding, check whether the ledger already holds it:
134
+ `~/.local/share/agent-context/repo/recipes/tools/ledger-index --kind trap --scope <scope>`.
135
+ A finding the ledger already records is a compliance gap, not a discovery. Say which it is.
136
+ 2. For each genuinely new, durable, cross-cutting trap, draft a note: `kind`, `scope`, `title`,
137
+ `body`, and a `why` naming the concrete failure. A finding that is only true this week, or only
138
+ true in one repository, does not go in the ledger. It goes in the report.
139
+ 3. Appending is a shared write. Take the lease first
140
+ (`recipes/tools/task-claim acquire ledger-append <agent> 20`), append without touching any
141
+ existing note, run `node recipes/tools/validate-ledger.js`, assert the note count rose by
142
+ exactly the number added and that no prior id vanished, commit, push, then verify
143
+ `git rev-parse HEAD` against `git ls-remote origin main` and quote both. Release the lease even
144
+ on failure.
145
+ 4. Never append a credential, a personal path that is not the user's own machine, or a fact that is
146
+ only true this week.
147
+
148
+ Present the drafted notes to the user and get an explicit go before pushing. The ledger is shared
149
+ with another agent, so an unrequested push is an external side effect.
150
+
151
+ ## Step 6. Carry findings forward
152
+
153
+ The reason four correct reports changed nothing is that each one started clean. This skill does
154
+ not.
155
+
156
+ Write the output to a dated file and, on every run, read the previous one. Each finding carries:
157
+
158
+ - `first_seen`: the date it was first reported
159
+ - `runs_seen`: how many consecutive runs have reported it
160
+ - `status`: `new`, `carried`, `actioned`, or `closed-not-possible`
161
+
162
+ A finding at `runs_seen: 3` or more gets its own section at the top of the report titled with its
163
+ age, for example "Open for 3 runs". Reporting the same defect a fourth time without escalating it
164
+ is the failure this section prevents. `closed-not-possible` is a real terminal state and must
165
+ record why, so it is not rediscovered.
166
+
167
+ Use `context-repo` to resolve a durable private store for the snapshots if one is not already
168
+ configured. Do not write snapshots to a repo-local scratch directory.
169
+
170
+ ## Step 7. Output
171
+
172
+ Lead with a verdict in one line naming what is wrong, not what was analysed. Then, in this order:
173
+
174
+ 1. Open for N runs (carried findings, oldest first) - omit the section on a first run
175
+ 2. Verified as working
176
+ 3. Defects
177
+ 4. What repeated, ranked by distinct sessions
178
+ 5. Delivery leakage
179
+ 6. Skills and knowledge, including the ledger notes drafted this run
180
+ 7. The ranked change list, config edits before projects
181
+ 8. Receipts
182
+
183
+ Hard rules for the output:
184
+
185
+ - Numbers over adjectives. "1120 organic turns, 46 percent under 40 characters" beats "a lot of
186
+ short prompts".
187
+ - No em dashes anywhere. Hyphens or restructure.
188
+ - Every PR, issue or ticket line carries a clickable URL.
189
+ - Anything not run is named as not run. A step that was blocked is stated as blocked, with the
190
+ reason and what remains untrue because of it.
191
+ - Receipts are facts with values: paths, counts, SHAs, exit codes, what was deliberately not run.
192
+
193
+ ## Cost
194
+
195
+ Mining is grunt work. Run the extraction and the per-axis analysis as parallel leaf agents on
196
+ Sonnet or Haiku, and keep only the synthesis and the ledger decision on the session model. Pass the
197
+ agents the mined JSON paths rather than the transcripts, and require each to return its full report
198
+ in its final message: a subagent's transcript is not visible to the caller, and a report that says
199
+ "see above" arrives empty.
@@ -0,0 +1,245 @@
1
+ #!/usr/bin/env node
2
+ // mine-transcripts: extract operator-axis metrics from local Claude Code transcripts.
3
+ //
4
+ // Ships beside agent-ops-retro because a rule without its runnable half gets re-derived.
5
+ // Every filter here exists because a previous run got it wrong:
6
+ // - records are selected by their own .timestamp, never by file mtime (a resumed session
7
+ // rewrites old files and drags pre-window turns into the window)
8
+ // - a `user` record is not a human turn; tool_results, system-reminders, command preambles,
9
+ // task notifications and bash echoes are stripped, and they were 45% of the raw count
10
+ // - sidechain and subagent records are counted separately, never folded into session counts
11
+ //
12
+ // node mine-transcripts.mjs --since 2026-08-18 --until 2026-09-02 --out ./mine.json
13
+ //
14
+ // Writes <out> and <out minus .json>-turns.json. Prints a summary to stdout.
15
+
16
+ import fs from 'node:fs';
17
+ import path from 'node:path';
18
+ import readline from 'node:readline';
19
+
20
+ const arg = (k, d) => {
21
+ const i = process.argv.indexOf(`--${k}`);
22
+ return i > -1 && process.argv[i + 1] ? process.argv[i + 1] : d;
23
+ };
24
+
25
+ const ROOT = arg('root', path.join(process.env.HOME, '.claude', 'projects'));
26
+ const OUT = arg('out', './mine.json');
27
+ const SINCE = Date.parse(`${arg('since', '')}T00:00:00Z`);
28
+ const UNTIL = Date.parse(`${arg('until', '')}T00:00:00Z`);
29
+
30
+ if (!Number.isFinite(SINCE) || !Number.isFinite(UNTIL) || UNTIL <= SINCE) {
31
+ console.error('usage: mine-transcripts.mjs --since YYYY-MM-DD --until YYYY-MM-DD [--out path] [--root dir]');
32
+ process.exit(2);
33
+ }
34
+ if (!fs.existsSync(ROOT)) {
35
+ console.error(`transcript root not found: ${ROOT}. This is a state, not a success: stop here.`);
36
+ process.exit(1);
37
+ }
38
+
39
+ function walk(dir, out = []) {
40
+ for (const e of fs.readdirSync(dir, { withFileTypes: true })) {
41
+ const p = path.join(dir, e.name);
42
+ if (e.isDirectory()) walk(p, out);
43
+ else if (e.name.endsWith('.jsonl')) out.push(p);
44
+ }
45
+ return out;
46
+ }
47
+
48
+ // A user-role text record that is really harness output, not something a person typed.
49
+ const HARNESS = /^(<system-reminder|<command-name|<command-message|<local-command|<bash-input|<bash-stdout|<user-prompt-submit-hook|<task-notification|Caveat: The messages below|\[Request interrupted)/;
50
+
51
+ const textOf = (msg) => {
52
+ const c = msg?.content;
53
+ if (typeof c === 'string') return c;
54
+ if (Array.isArray(c)) return c.filter((b) => b.type === 'text').map((b) => b.text).join('\n');
55
+ return '';
56
+ };
57
+
58
+ const bump = (m, k, n = 1) => m.set(k, (m.get(k) || 0) + n);
59
+
60
+ const S = {
61
+ sessions: new Map(),
62
+ tools: new Map(),
63
+ agents: new Map(),
64
+ skills: new Map(),
65
+ models: new Map(),
66
+ errorsByTool: new Map(),
67
+ errorKinds: new Map(),
68
+ limits: [],
69
+ humanTurns: [],
70
+ echoTurns: 0,
71
+ interrupts: 0,
72
+ compacts: 0,
73
+ usage: { in: 0, out: 0, cacheRead: 0, cacheWrite: 0 },
74
+ };
75
+
76
+ function classifyError(text) {
77
+ const t = text.slice(0, 400);
78
+ if (/haven't granted it yet|Permission to use|requires approval/i.test(t)) return 'permission-not-granted';
79
+ if (/auto mode classifier/i.test(t)) return 'auto-mode-classifier';
80
+ if (/isolated in the worktree|stays inside worktree/i.test(t)) return 'worktree-isolation';
81
+ if (/Delegation rule/i.test(t)) return 'delegation-rule';
82
+ if (/does not match required schema/i.test(t)) return 'agent-schema-miss';
83
+ if (/File does not exist|no such file/i.test(t)) return 'file-not-found';
84
+ if (/has not been read yet/i.test(t)) return 'read-before-write';
85
+ if (/String to replace not found|not found in file/i.test(t)) return 'edit-string-not-found';
86
+ if (/timed out|timeout/i.test(t)) return 'timeout';
87
+ if (/\b(40[0-9]|41[0-9]|429)\b.*error|HTTP 4\d\d/i.test(t)) return 'api-4xx';
88
+ if (/\b(50[0-9]|529)\b.*error|HTTP 5\d\d/i.test(t)) return 'api-5xx';
89
+ if (/command not found/i.test(t)) return 'command-not-found';
90
+ if (/Exit code [1-9]/i.test(t)) return 'non-zero-exit';
91
+ return 'other';
92
+ }
93
+
94
+ const files = walk(ROOT);
95
+ let inWindow = 0;
96
+
97
+ for (const f of files) {
98
+ const isSub = f.includes(`${path.sep}subagents${path.sep}`);
99
+ const project = path.relative(ROOT, f).split(path.sep)[0];
100
+ const rl = readline.createInterface({ input: fs.createReadStream(f), crlfDelay: Infinity });
101
+
102
+ for await (const line of rl) {
103
+ if (!line.startsWith('{')) continue;
104
+ let r;
105
+ try { r = JSON.parse(line); } catch { continue; }
106
+
107
+ const ts = r.timestamp ? Date.parse(r.timestamp) : NaN;
108
+ if (!Number.isFinite(ts) || ts < SINCE || ts >= UNTIL) continue;
109
+ inWindow++;
110
+
111
+ const main = !isSub && !r.isSidechain;
112
+ const sid = r.sessionId || path.basename(f, '.jsonl');
113
+ let sess = null;
114
+ if (main) {
115
+ if (!S.sessions.has(sid)) {
116
+ S.sessions.set(sid, { project, first: ts, last: ts, human: 0, assistant: 0, tools: 0, agents: 0, errors: 0, interrupts: 0, peakContext: 0 });
117
+ }
118
+ sess = S.sessions.get(sid);
119
+ sess.first = Math.min(sess.first, ts);
120
+ sess.last = Math.max(sess.last, ts);
121
+ }
122
+
123
+ if (r.isCompactSummary || r.subtype === 'compact_boundary') S.compacts++;
124
+
125
+ // Capacity and API-failure notices are synthetic records and are NOT always type "assistant",
126
+ // so this check sits outside the assistant branch. Gating it on the type reported zero for a
127
+ // window that really held five weekly-limit hits.
128
+ if (r.message?.model === '<synthetic>') {
129
+ const t = textOf(r.message);
130
+ if (/weekly limit|spend limit|rate limit|usage limit/i.test(t)) {
131
+ S.limits.push({ ts: r.timestamp, text: t.slice(0, 90) });
132
+ }
133
+ }
134
+
135
+ if (r.type === 'assistant') {
136
+ const model = r.message?.model;
137
+ const u = r.message?.usage;
138
+ if (model) {
139
+ const stats = S.models.get(model) || { messages: 0, in: 0, out: 0, cacheRead: 0, cacheWrite: 0 };
140
+ stats.messages++;
141
+ stats.in += u?.input_tokens || 0;
142
+ stats.out += u?.output_tokens || 0;
143
+ stats.cacheRead += u?.cache_read_input_tokens || 0;
144
+ stats.cacheWrite += u?.cache_creation_input_tokens || 0;
145
+ S.models.set(model, stats);
146
+ }
147
+ if (u) {
148
+ S.usage.in += u.input_tokens || 0;
149
+ S.usage.out += u.output_tokens || 0;
150
+ S.usage.cacheRead += u.cache_read_input_tokens || 0;
151
+ S.usage.cacheWrite += u.cache_creation_input_tokens || 0;
152
+ if (sess) {
153
+ const ctx = (u.input_tokens || 0) + (u.cache_read_input_tokens || 0) + (u.cache_creation_input_tokens || 0);
154
+ sess.peakContext = Math.max(sess.peakContext, ctx);
155
+ }
156
+ }
157
+ if (sess) sess.assistant++;
158
+ for (const b of Array.isArray(r.message?.content) ? r.message.content : []) {
159
+ if (b.type !== 'tool_use') continue;
160
+ bump(S.tools, b.name);
161
+ if (sess) sess.tools++;
162
+ if (b.name === 'Agent') {
163
+ bump(S.agents, `${b.input?.subagent_type || '(none)'} | ${b.input?.model || '(inherit)'}`);
164
+ if (sess) sess.agents++;
165
+ }
166
+ if (b.name === 'Skill') bump(S.skills, b.input?.skill || '(unnamed)');
167
+ }
168
+ }
169
+
170
+ if (r.type !== 'user') continue;
171
+
172
+ const blocks = Array.isArray(r.message?.content) ? r.message.content : [];
173
+ let sawToolResult = false;
174
+ for (const b of blocks) {
175
+ if (b.type !== 'tool_result') continue;
176
+ sawToolResult = true;
177
+ if (!b.is_error) continue;
178
+ const body = typeof b.content === 'string' ? b.content : JSON.stringify(b.content || '');
179
+ bump(S.errorKinds, classifyError(body));
180
+ bump(S.errorsByTool, b.tool_use_id ? 'attributed' : 'unattributed');
181
+ if (sess) sess.errors++;
182
+ }
183
+
184
+ const raw = textOf(r.message);
185
+ if (raw.includes('[Request interrupted')) {
186
+ S.interrupts++;
187
+ if (sess) sess.interrupts++;
188
+ }
189
+ if (!main || r.userType !== 'external' || !raw || sawToolResult) continue;
190
+
191
+ const clean = raw.replace(/<system-reminder>[\s\S]*?<\/system-reminder>/g, '').trim();
192
+ if (!clean) continue;
193
+ if (HARNESS.test(clean)) { S.echoTurns++; continue; }
194
+
195
+ if (sess) sess.human++;
196
+ S.humanTurns.push({ ts: r.timestamp, session: sid, project, text: clean.slice(0, 600) });
197
+ }
198
+ }
199
+
200
+ const sessions = [...S.sessions].map(([id, v]) => ({
201
+ id: id.slice(0, 8), ...v, minutes: Math.round((v.last - v.first) / 60000),
202
+ })).sort((a, b) => b.human - a.human);
203
+
204
+ const humanCounts = sessions.map((s) => s.human).filter((n) => n > 0).sort((a, b) => a - b);
205
+ const pct = (p) => humanCounts[Math.floor((humanCounts.length - 1) * p)] ?? 0;
206
+ const byDay = {};
207
+ for (const t of S.humanTurns) {
208
+ const d = t.ts.slice(0, 10);
209
+ byDay[d] = (byDay[d] || 0) + 1;
210
+ }
211
+
212
+ const report = {
213
+ window: { since: new Date(SINCE).toISOString(), until: new Date(UNTIL).toISOString() },
214
+ filesScanned: files.length,
215
+ recordsInWindow: inWindow,
216
+ mainSessions: S.sessions.size,
217
+ humanTurnsOrganic: S.humanTurns.length,
218
+ humanTurnsEchoesExcluded: S.echoTurns,
219
+ humanTurnsPerDay: byDay,
220
+ turnsPerSession: { median: pct(0.5), p90: pct(0.9), max: humanCounts.at(-1) ?? 0 },
221
+ shortTurnsUnder40Chars: S.humanTurns.filter((t) => t.text.trim().length < 40).length,
222
+ interrupts: S.interrupts,
223
+ compactEvents: S.compacts,
224
+ capacityLimitEvents: S.limits.length,
225
+ capacityLimitSamples: S.limits.slice(0, 10),
226
+ usage: S.usage,
227
+ cacheReadToOutputRatio: S.usage.out ? +(S.usage.cacheRead / S.usage.out).toFixed(1) : null,
228
+ models: [...S.models].sort((a, b) => b[1].messages - a[1].messages),
229
+ topTools: [...S.tools].sort((a, b) => b[1] - a[1]).slice(0, 40),
230
+ agentSpawns: [...S.agents].sort((a, b) => b[1] - a[1]),
231
+ skillInvocations: [...S.skills].sort((a, b) => b[1] - a[1]),
232
+ errorKinds: [...S.errorKinds].sort((a, b) => b[1] - a[1]),
233
+ sessions,
234
+ };
235
+
236
+ fs.writeFileSync(OUT, JSON.stringify(report, null, 1));
237
+ fs.writeFileSync(OUT.replace(/\.json$/, '') + '-turns.json', JSON.stringify(S.humanTurns, null, 1));
238
+
239
+ console.log(JSON.stringify({
240
+ ...report,
241
+ capacityLimitSamples: report.capacityLimitSamples.length,
242
+ sessions: report.sessions.length,
243
+ }, null, 1));
244
+ console.error(`\nwrote ${OUT} and ${OUT.replace(/\.json$/, '')}-turns.json`);
245
+ console.error(`organic human turns ${report.humanTurnsOrganic}, harness echoes excluded ${report.humanTurnsEchoesExcluded}`);
@@ -7,5 +7,6 @@ Use the canonical skills directly:
7
7
  - `plugins/olko-reflection/skills/self-critique/SKILL.md`
8
8
  - `plugins/olko-reflection/skills/review-past-performance/SKILL.md`
9
9
  - `plugins/olko-reflection/skills/retro-analysis/SKILL.md`
10
+ - `plugins/olko-reflection/skills/agent-ops-retro/SKILL.md`
10
11
  - `plugins/olko-reflection/skills/crash-course/SKILL.md`
11
12
  - `plugins/olko-reflection/skills/wrap-up/SKILL.md`
@@ -0,0 +1,199 @@
1
+ ---
2
+ name: agent-ops-retro
3
+ description: >
4
+ Retrospective on how the agents themselves are being operated, from local Claude Code
5
+ transcripts plus the reports other jobs already produce. Answers what is going well, what is
6
+ going wrong, what the human keeps having to repeat, and where delivery is leaking, then carries
7
+ every unactioned finding forward with its age. USE THIS AUTOMATICALLY (no need for the user to
8
+ name it) whenever the user asks to check all conversations, review agent usage, reduce the manual
9
+ administration of agents, find repeating corrections or issues across sessions, audit skills and
10
+ the knowledge ledger together, or asks what gaps exist in delivery over a recent window. This is
11
+ the operator axis: for repository history and shipped work use retro-analysis, for a single
12
+ finished change use wrap-up.
13
+ license: MIT
14
+ allowed-tools: Bash, Read, Write, Grep, Glob, Agent
15
+ compatibility: Claude Code. Reads ~/.claude transcripts and local report files, so it does not run on hosts without them.
16
+ metadata:
17
+ author: Oleg Koval
18
+ targets: [_source-only]
19
+ tags:
20
+ - retrospective
21
+ - agent-operations
22
+ - telemetry
23
+ - skills
24
+ - knowledge-ledger
25
+ - delegation
26
+ - cost
27
+ ---
28
+ <!-- Generated by scripts/build-adapters.sh. Do not edit directly. -->
29
+
30
+ # Agent Ops Retro
31
+
32
+ A retrospective on the operating layer, not the code. It asks whether the agents are being run
33
+ well: what the human had to repeat, what the guardrails caught, what they cost, and which findings
34
+ have been reported before and never actioned.
35
+
36
+ ## Scope boundary
37
+
38
+ Three retro skills exist and they do not overlap.
39
+
40
+ | Skill | Looks at | Answers |
41
+ |---|---|---|
42
+ | `retro-analysis` | Repository history, delivery evidence, code quality | What shipped and how the code is trending |
43
+ | `wrap-up` | One finished change | Did this specific piece of work meet its goal |
44
+ | `agent-ops-retro` | Transcripts, hooks, skills, ledger, delegation logs | How the agents are being operated |
45
+
46
+ If the user asks about shipped work, hand off to `retro-analysis` and say so.
47
+
48
+ ## Invocation
49
+
50
+ Accept a window: `14d` (default), `7d`, `30d`, or an explicit ISO date range. State the absolute
51
+ range at the top of the output. Convert relative dates to absolute.
52
+
53
+ ## Step 0. Read, do not recompute
54
+
55
+ Several jobs already report on this state. Adding a fifth independent opinion is the failure mode
56
+ this skill exists to fix, so read their outputs first and only compute what none of them cover.
57
+
58
+ | Source | Path | Gives you |
59
+ |---|---|---|
60
+ | Skill audit | `~/.claude/skill-audit/latest.json` and the dated `.txt` beside it | Installed vs never-invoked skills |
61
+ | Delegation log | `~/.claude/delegation-metrics/delegations.jsonl` | Model, tier, depth, outcome, rework per spawn |
62
+ | Tier check | `~/.claude/delegation-metrics/tier-check-*.md` | Opus share and spend against a baseline |
63
+ | Usage miner | `~/obsidian/cloud-opus/Lead/Automation/Weekly Automation Report *.md` | Repeated commands, proposal fatigue, approval queue |
64
+ | Miner state | `~/.claude/usage-miner/state/last-run.json` | What was staged, queued, blocked |
65
+ | Knowledge ledger | `~/.local/share/agent-context/repo/ledger.json` | Traps already recorded, so you do not re-derive them |
66
+ | Prior run | the previous output of this skill (see Step 6) | Findings to carry forward |
67
+
68
+ Two traps when reading these:
69
+
70
+ - A skill-usage figure counted from attributed messages is not an invocation count. The audit's own
71
+ numbers run three orders of magnitude above the Skill-tool invocation count. Say which unit you
72
+ are quoting, every time.
73
+ - When two sources disagree on the same number, report both with their sources and say the figure
74
+ is unreconciled. Never silently pick one.
75
+
76
+ ## Step 1. Mine the transcripts
77
+
78
+ This is the axis nothing else covers. Run `scripts/mine-transcripts.mjs` from this skill directory,
79
+ or reimplement it, over `~/.claude/projects/**/*.jsonl`.
80
+
81
+ Non-negotiable filters, each of which has burned a previous run:
82
+
83
+ - Filter by the record's own `.timestamp`, never by file mtime. A resumed session rewrites old
84
+ files and pulls pre-window turns into the window.
85
+ - A `user` record is not a human turn. Exclude `tool_result` blocks, `<system-reminder>` wrappers,
86
+ `<command-name>` and hook preambles, task notifications and bash echoes. In one measured corpus
87
+ 45 percent of `user` text records were tool echoes, so the raw count nearly doubles the real one.
88
+ - Exclude `isSidechain` records and anything under a `subagents/` directory from session and
89
+ human-turn counts. Count them separately for the delegation view.
90
+ - Empty output is a state, not a success. If a scan returns nothing, prove the scan works before
91
+ concluding the thing is absent.
92
+
93
+ Extract: organic human turns per day and per session, turn-length distribution, tool call counts,
94
+ `tool_result` errors with `is_error`, Agent spawns with `subagent_type` and `model`, Skill
95
+ invocations, interruptions, compaction events, per-model token usage including cache reads, and
96
+ synthetic records naming a rate or spend limit.
97
+
98
+ ## Step 2. Classify what the human repeated
99
+
100
+ Group the organic turns. The strongest automation signal is not raw frequency, it is the number of
101
+ distinct sessions a correction appears in: a thing said twenty times in one session is one
102
+ argument, a thing said once in twenty sessions is a missing default.
103
+
104
+ Report each recurring pattern as: name, count, distinct sessions, two or three verbatim dated
105
+ quotes, and the one concrete change that removes it. Name the mechanism, not the mood. Write
106
+ "stops for a one-word ack on reversible steps", not "autonomy could be better".
107
+
108
+ ## Step 3. Split what holds from what is broken
109
+
110
+ Produce two explicit lists, never a narrative.
111
+
112
+ **Verified as working:** the claim and what proved it. Guardrails that fired correctly count here.
113
+
114
+ **Defects:** the mechanism and its consequence. Include the governance frictions, which usually
115
+ outnumber the external failures: permission denials, classifier blocks, path-gate refusals, agent
116
+ output failing its own schema. Give each a count and a share of total errors.
117
+
118
+ ## Step 4. Delivery leakage
119
+
120
+ Delivery here means whether agent output landed, not whether it was produced.
121
+
122
+ - Open PRs authored by the user, bucketed by age, with fan-out batches identified as batches.
123
+ - Tickets whose completion timestamps cluster in minutes, which is a status backfill rather than
124
+ shipping. State both the raw completion count and the count after removing the cluster.
125
+ - Any fan-out with no named landing mechanism. That is backlog, not delivery.
126
+ - Scheduled jobs: check `launchctl list` exit codes and each job's own log. A wrapper exit code is
127
+ not completion.
128
+
129
+ ## Step 5. Close the knowledge circle
130
+
131
+ This is the half that makes the retro compound instead of repeating.
132
+
133
+ 1. Before writing a finding, check whether the ledger already holds it:
134
+ `~/.local/share/agent-context/repo/recipes/tools/ledger-index --kind trap --scope <scope>`.
135
+ A finding the ledger already records is a compliance gap, not a discovery. Say which it is.
136
+ 2. For each genuinely new, durable, cross-cutting trap, draft a note: `kind`, `scope`, `title`,
137
+ `body`, and a `why` naming the concrete failure. A finding that is only true this week, or only
138
+ true in one repository, does not go in the ledger. It goes in the report.
139
+ 3. Appending is a shared write. Take the lease first
140
+ (`recipes/tools/task-claim acquire ledger-append <agent> 20`), append without touching any
141
+ existing note, run `node recipes/tools/validate-ledger.js`, assert the note count rose by
142
+ exactly the number added and that no prior id vanished, commit, push, then verify
143
+ `git rev-parse HEAD` against `git ls-remote origin main` and quote both. Release the lease even
144
+ on failure.
145
+ 4. Never append a credential, a personal path that is not the user's own machine, or a fact that is
146
+ only true this week.
147
+
148
+ Present the drafted notes to the user and get an explicit go before pushing. The ledger is shared
149
+ with another agent, so an unrequested push is an external side effect.
150
+
151
+ ## Step 6. Carry findings forward
152
+
153
+ The reason four correct reports changed nothing is that each one started clean. This skill does
154
+ not.
155
+
156
+ Write the output to a dated file and, on every run, read the previous one. Each finding carries:
157
+
158
+ - `first_seen`: the date it was first reported
159
+ - `runs_seen`: how many consecutive runs have reported it
160
+ - `status`: `new`, `carried`, `actioned`, or `closed-not-possible`
161
+
162
+ A finding at `runs_seen: 3` or more gets its own section at the top of the report titled with its
163
+ age, for example "Open for 3 runs". Reporting the same defect a fourth time without escalating it
164
+ is the failure this section prevents. `closed-not-possible` is a real terminal state and must
165
+ record why, so it is not rediscovered.
166
+
167
+ Use `context-repo` to resolve a durable private store for the snapshots if one is not already
168
+ configured. Do not write snapshots to a repo-local scratch directory.
169
+
170
+ ## Step 7. Output
171
+
172
+ Lead with a verdict in one line naming what is wrong, not what was analysed. Then, in this order:
173
+
174
+ 1. Open for N runs (carried findings, oldest first) - omit the section on a first run
175
+ 2. Verified as working
176
+ 3. Defects
177
+ 4. What repeated, ranked by distinct sessions
178
+ 5. Delivery leakage
179
+ 6. Skills and knowledge, including the ledger notes drafted this run
180
+ 7. The ranked change list, config edits before projects
181
+ 8. Receipts
182
+
183
+ Hard rules for the output:
184
+
185
+ - Numbers over adjectives. "1120 organic turns, 46 percent under 40 characters" beats "a lot of
186
+ short prompts".
187
+ - No em dashes anywhere. Hyphens or restructure.
188
+ - Every PR, issue or ticket line carries a clickable URL.
189
+ - Anything not run is named as not run. A step that was blocked is stated as blocked, with the
190
+ reason and what remains untrue because of it.
191
+ - Receipts are facts with values: paths, counts, SHAs, exit codes, what was deliberately not run.
192
+
193
+ ## Cost
194
+
195
+ Mining is grunt work. Run the extraction and the per-axis analysis as parallel leaf agents on
196
+ Sonnet or Haiku, and keep only the synthesis and the ledger decision on the session model. Pass the
197
+ agents the mined JSON paths rather than the transcripts, and require each to return its full report
198
+ in its final message: a subagent's transcript is not visible to the caller, and a report that says
199
+ "see above" arrives empty.