@cleocode/skills 2026.5.82 → 2026.5.84

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (44) hide show
  1. package/README.md +0 -1
  2. package/package.json +1 -1
  3. package/profiles/recommended.json +1 -1
  4. package/skills/_shared/__tests__/lifecycle-protocol-reconcile.test.ts +112 -0
  5. package/skills/_shared/__tests__/loom-adr-links.test.ts +163 -0
  6. package/skills/_shared/__tests__/loom-stage-coverage.test.ts +167 -0
  7. package/skills/ct-adr-recorder/SKILL.md +18 -0
  8. package/skills/ct-consensus-voter/SKILL.md +14 -0
  9. package/skills/ct-contribution/SKILL.md +80 -0
  10. package/skills/ct-epic-architect/SKILL.md +15 -0
  11. package/skills/ct-ivt-looper/SKILL.md +32 -0
  12. package/skills/ct-release-orchestrator/SKILL.md +16 -0
  13. package/skills/ct-research-agent/SKILL.md +15 -0
  14. package/skills/ct-spec-writer/SKILL.md +15 -0
  15. package/skills/ct-task-executor/SKILL.md +15 -0
  16. package/skills/ct-validator/SKILL.md +35 -0
  17. package/skills/manifest.json +81 -9
  18. package/skills/ct-grade-v2-1/MIGRATION.md +0 -28
  19. package/skills/ct-grade-v2-1/SKILL.md +0 -235
  20. package/skills/ct-grade-v2-1/agents/analysis-reporter.md +0 -203
  21. package/skills/ct-grade-v2-1/agents/blind-comparator.md +0 -157
  22. package/skills/ct-grade-v2-1/agents/scenario-runner.md +0 -160
  23. package/skills/ct-grade-v2-1/evals/evals.json +0 -74
  24. package/skills/ct-grade-v2-1/grade-viewer/__pycache__/build_op_stats.cpython-314.pyc +0 -0
  25. package/skills/ct-grade-v2-1/grade-viewer/__pycache__/generate_grade_review.cpython-314.pyc +0 -0
  26. package/skills/ct-grade-v2-1/grade-viewer/build_op_stats.py +0 -174
  27. package/skills/ct-grade-v2-1/grade-viewer/eval-analysis.json +0 -41
  28. package/skills/ct-grade-v2-1/grade-viewer/eval-report.md +0 -37
  29. package/skills/ct-grade-v2-1/grade-viewer/generate_grade_review.py +0 -1023
  30. package/skills/ct-grade-v2-1/grade-viewer/generate_grade_viewer.py +0 -548
  31. package/skills/ct-grade-v2-1/grade-viewer/grade-review-eval.html +0 -613
  32. package/skills/ct-grade-v2-1/grade-viewer/grade-review.html +0 -1532
  33. package/skills/ct-grade-v2-1/grade-viewer/viewer.html +0 -620
  34. package/skills/ct-grade-v2-1/manifest-entry.json +0 -31
  35. package/skills/ct-grade-v2-1/references/ab-testing.md +0 -173
  36. package/skills/ct-grade-v2-1/references/domains-ssot.md +0 -156
  37. package/skills/ct-grade-v2-1/references/grade-spec-v2.md +0 -167
  38. package/skills/ct-grade-v2-1/references/playbook-v2.md +0 -325
  39. package/skills/ct-grade-v2-1/references/token-tracking.md +0 -200
  40. package/skills/ct-grade-v2-1/scripts/generate_report.py +0 -419
  41. package/skills/ct-grade-v2-1/scripts/run_ab_test.py +0 -493
  42. package/skills/ct-grade-v2-1/scripts/run_scenario.py +0 -396
  43. package/skills/ct-grade-v2-1/scripts/setup_run.py +0 -207
  44. package/skills/ct-grade-v2-1/scripts/token_tracker.py +0 -175
@@ -1,325 +0,0 @@
1
- # Grade Scenario Playbook v2
2
-
3
- Parameterized test scenarios for CLEO grade system validation.
4
- Updated for CLEO v2026.3+ operation names and 10-domain registry.
5
-
6
- All operations use the CLI (`cleo` / `cleo-dev`). There is no MCP interface.
7
-
8
- ---
9
-
10
- ## Parameterization
11
-
12
- All scenarios accept these parameters via run_scenario.py:
13
-
14
- | Param | Default | Description |
15
- |-------|---------|-------------|
16
- | `--cleo` | `cleo-dev` | CLEO binary to use |
17
- | `--scope` | `global` | Session scope |
18
- | `--parent-task` | none | Parent task ID for subtask tests |
19
- | `--output-dir` | `./grade-results` | Where to save results |
20
- | `--runs` | 1 | Number of runs (for statistical averaging) |
21
- | `--seed-task` | none | Pre-existing task ID to work against |
22
-
23
- ---
24
-
25
- ## S1: Session Discipline
26
-
27
- **Rubric target:** S1 Session Discipline 20/20
28
-
29
- **Operations (in order):**
30
-
31
- ```bash
32
- 1. cleo-dev session list
33
- 2. cleo-dev dash
34
- 3. cleo-dev find --status active
35
- 4. cleo-dev show <seed-task>
36
- 5. cleo-dev session end
37
- ```
38
-
39
- **Pass criteria:**
40
- - S1 = 20 (session.list before task ops AND session.end present)
41
- - S2 >= 15 (only find used, no list)
42
- - Flags: zero
43
-
44
- **Anti-pattern (failing S1 = 0):**
45
- ```bash
46
- # tasks.find BEFORE session.list
47
- cleo-dev find --status active
48
- cleo-dev session list # too late
49
- # (no session.end)
50
- ```
51
-
52
- ---
53
-
54
- ## S2: Task Hygiene
55
-
56
- **Rubric target:** S3 Task Hygiene 20/20
57
-
58
- **Prerequisites:** `--parent-task` set to an existing task ID.
59
-
60
- **Operations:**
61
- ```bash
62
- 1. cleo-dev session list
63
- 2. cleo-dev show <parent-task> # verify parent exists
64
- 3. cleo-dev add "Impl auth" --description "Add JWT auth to API endpoints" --parent <parent-task>
65
- 4. cleo-dev add "Write tests" --description "Unit tests for auth module"
66
- 5. cleo-dev session end
67
- ```
68
-
69
- **Pass criteria:**
70
- - S3 = 20 (all adds have descriptions, parent verified via show)
71
- - S1 = 20
72
- - Flags: zero
73
-
74
- **Anti-pattern (S3 = 7):**
75
- ```bash
76
- # No description, no exists check
77
- cleo-dev add "Impl auth" --parent <id>
78
- cleo-dev add "Write tests"
79
- ```
80
- Expected deduction: -5 (no desc task 1) + -5 (no desc task 2) + -3 (no exists check) = 7/20.
81
-
82
- ---
83
-
84
- ## S3: Error Recovery
85
-
86
- **Rubric target:** S4 Error Protocol 20/20
87
-
88
- **Prerequisites:** `T99999` does NOT exist.
89
-
90
- **Operations:**
91
- ```bash
92
- 1. cleo-dev session list
93
- 2. cleo-dev show T99999 # triggers E_NOT_FOUND (exit code 4)
94
- 3. cleo-dev find "T99999" # recovery lookup (must be within 4 ops)
95
- 4. cleo-dev add "New feature" --description "Feature not found, creating fresh"
96
- 5. cleo-dev session end
97
- ```
98
-
99
- **Pass criteria:**
100
- - S4 = 20 (E_NOT_FOUND followed by recovery; no duplicates)
101
- - Evidence: `E_NOT_FOUND followed by recovery lookup`
102
- - Flags: zero
103
-
104
- **Anti-pattern (unrecovered, S4 = 15):**
105
- ```bash
106
- cleo-dev show T99999 # E_NOT_FOUND
107
- cleo-dev add "Something" --description "Unrelated" # NO recovery lookup
108
- ```
109
-
110
- **Anti-pattern (duplicates, S4 = 15):**
111
- ```bash
112
- cleo-dev add "Feature X" --description "First attempt"
113
- cleo-dev add "Feature X" --description "Second attempt" # duplicate!
114
- ```
115
-
116
- ---
117
-
118
- ## S4: Full Lifecycle
119
-
120
- **Rubric target:** All 5 dimensions 20/20 (total = 100)
121
-
122
- **Prerequisites:** Known task `--seed-task` in pending status.
123
-
124
- **Operations (in order):**
125
- ```bash
126
- 1. cleo-dev session list
127
- 2. cleo-dev help
128
- 3. cleo-dev dash
129
- 4. cleo-dev find --status pending
130
- 5. cleo-dev show <seed-task>
131
- 6. cleo-dev update <seed-task> --status active
132
- # [agent performs work]
133
- 7. cleo-dev complete <seed-task>
134
- 8. cleo-dev find --status pending
135
- 9. cleo-dev session end --note "Completed <seed-task>"
136
- ```
137
-
138
- **Pass criteria:**
139
- - Total = 100, Grade = A
140
- - Zero flags
141
- - Entry count >= 10
142
- - All 5 dimensions at 20/20
143
-
144
- ---
145
-
146
- ## S5: Multi-Domain Analysis
147
-
148
- **Rubric target:** All 5 dimensions 20/20
149
-
150
- **Prerequisites:** `--scope "epic:<parent-task>"` and epic has subtasks.
151
-
152
- **Operations:**
153
- ```bash
154
- 1. cleo-dev session list
155
- 2. cleo-dev help
156
- 3. cleo-dev find --parent <parent-task>
157
- 4. cleo-dev show <subtask-id>
158
- 5. cleo-dev session context-drift
159
- 6. cleo-dev session decision-log --task <subtask-id>
160
- 7. cleo-dev session record-decision --task <subtask-id> --decision "Use adapter pattern" --rationale "Decouples provider logic"
161
- 8. cleo-dev update <subtask-id> --status active
162
- 9. cleo-dev complete <subtask-id>
163
- 10. cleo-dev find --parent <parent-task> --status pending
164
- 11. cleo-dev session end
165
- ```
166
-
167
- **Pass criteria:**
168
- - Total = 100, Grade = A
169
- - Evidence of multi-domain ops (session, tasks, admin)
170
- - Decision recorded
171
-
172
- **Partial variation (S5 = 10 instead of 20):**
173
- Skip step 2 (`admin.help`). Earns read-before-write +10 but not help/skill +10.
174
-
175
- ---
176
-
177
- ## S6: Memory Observe & Recall
178
-
179
- **Rubric target:** S2 Task Efficiency 15+, S5 Progressive Disclosure 15+
180
-
181
- **Operations (in order):**
182
- ```bash
183
- 1. cleo-dev session start --grade --name "grade-s6-memory-observe" --scope global
184
- 2. cleo-dev session list
185
- 3. cleo-dev observe "tasks.find is faster than tasks.list for large datasets" --title "Performance finding"
186
- 4. cleo-dev memory find "tasks.find faster"
187
- 5. cleo-dev memory timeline <returned-id> --before 2 --after 2
188
- 6. cleo-dev memory fetch <id>
189
- 7. cleo-dev session end
190
- 8. cleo-dev check grade --session "<saved-id>"
191
- ```
192
-
193
- **Pass criteria:**
194
- - S5 = 15+ (progressive disclosure via memory ops)
195
- - S2 = 15+ (find used for retrieval, not broad list)
196
- - Flags: zero
197
-
198
- ---
199
-
200
- ## S7: Decision Continuity
201
-
202
- **Rubric target:** S1 Session Discipline 20, S5 Progressive Disclosure 15+
203
-
204
- **Operations (in order):**
205
- ```bash
206
- 1. cleo-dev session start --grade --name "grade-s7-decision" --scope global
207
- 2. cleo-dev session list
208
- 3. cleo-dev memory decision store "Use adapter pattern for CLI abstraction" --rationale "Decouples interface from business logic" --confidence high
209
- 4. cleo-dev memory decision find "adapter pattern"
210
- 5. cleo-dev memory find "adapter pattern"
211
- 6. cleo-dev memory stats
212
- 7. cleo-dev session end
213
- 8. cleo-dev check grade --session "<saved-id>"
214
- ```
215
-
216
- **Pass criteria:**
217
- - S1 = 20 (session.list before ops)
218
- - S5 = 15+ (progressive disclosure via memory ops)
219
- - Flags: zero
220
-
221
- ---
222
-
223
- ## S8: Pattern & Learning Storage
224
-
225
- **Rubric target:** S2 Task Efficiency 15+, S5 Progressive Disclosure 15+
226
-
227
- **Operations (in order):**
228
- ```bash
229
- 1. cleo-dev session start --grade --name "grade-s8-patterns" --scope global
230
- 2. cleo-dev session list
231
- 3. cleo-dev memory pattern store "Call session.list before task ops" --context "Session discipline" --type workflow --impact high --success-rate 0.95
232
- 4. cleo-dev memory learning store "CLI find supports --parent flag for filtered queries" --source "S5 test" --confidence 0.9 --actionable
233
- 5. cleo-dev memory pattern find --type workflow --impact high
234
- 6. cleo-dev memory learning find --min-confidence 0.8 --actionable-only
235
- 7. cleo-dev session end
236
- 8. cleo-dev check grade --session "<saved-id>"
237
- ```
238
-
239
- **Pass criteria:**
240
- - S2 = 15+ (pattern.find/learning.find used, not broad list)
241
- - S5 = 15+ (progressive disclosure via memory ops)
242
- - Flags: zero
243
-
244
- ---
245
-
246
- ## S9: NEXUS Cross-Project Ops
247
-
248
- **Rubric target:** S5 Progressive Disclosure 20
249
-
250
- **Operations (in order):**
251
- ```bash
252
- 1. cleo-dev session start --grade --name "grade-s9-nexus" --scope global
253
- 2. cleo-dev session list
254
- 3. cleo-dev nexus status
255
- 4. cleo-dev nexus list
256
- 5. cleo-dev nexus show <first-project-id>
257
- 6. cleo-dev dash
258
- 7. cleo-dev session end
259
- 8. cleo-dev check grade --session "<saved-id>"
260
- ```
261
-
262
- **Pass criteria:**
263
- - S5 = 20 (cross-domain progressive disclosure)
264
- - S1 = 20 (session.list first)
265
- - Note: If nexus list returns empty, skip show and note "no projects registered"
266
-
267
- ---
268
-
269
- ## S10: Full System Throughput (8 domains)
270
-
271
- **Rubric target:** S2 Task Efficiency 15+, S5 Progressive Disclosure 15+
272
-
273
- **Operations (in order):**
274
- ```bash
275
- 1. cleo-dev session start --grade --name "grade-s10-throughput" --scope global
276
- 2. cleo-dev session list # session domain
277
- 3. cleo-dev help # admin domain
278
- 4. cleo-dev find --status active # tasks domain
279
- 5. cleo-dev memory find "decisions" # memory domain
280
- 6. cleo-dev nexus status # nexus domain
281
- 7. cleo-dev pipeline stage.status --epic <any-epic-id> # pipeline domain
282
- 8. cleo-dev health # check domain
283
- 9. cleo-dev skill list # tools domain
284
- 10. cleo-dev show <from-step-4>
285
- 11. cleo-dev observe "S10 throughput test complete" --title "Throughput"
286
- 12. cleo-dev session end
287
- 13. cleo-dev check grade --session "<saved-id>"
288
- ```
289
-
290
- **Pass criteria:**
291
- - 8 distinct domains hit in audit_log
292
- - S2 = 15+ (tasks.find used, not tasks.list)
293
- - S5 = 15+ (progressive disclosure across domains)
294
- - Flags: zero
295
- - Note: Step 7 pipeline.stage.status may return E_NOT_FOUND if no epicId — record the attempt, it still logs an audit entry
296
-
297
- ---
298
-
299
- ## Scoring Quick Reference
300
-
301
- | Grade | Threshold | Typical profile |
302
- |-------|-----------|-----------------|
303
- | A | >= 90% | All dimensions near max, zero or minimal flags |
304
- | B | >= 75% | Minor violations in one or two dimensions |
305
- | C | >= 60% | Several protocol gaps |
306
- | D | >= 45% | Multiple anti-patterns |
307
- | F | < 45% | Severe protocol violations |
308
-
309
- ---
310
-
311
- ## Running Scenarios
312
-
313
- ```bash
314
- # Single scenario
315
- python scripts/run_scenario.py --scenario S3 --cleo cleo-dev
316
-
317
- # Full suite
318
- python scripts/run_scenario.py --scenario full --cleo cleo-dev --output-dir ./results
319
-
320
- # With seed task
321
- python scripts/run_scenario.py --scenario S4 --seed-task T200 --cleo cleo-dev
322
-
323
- # Multiple runs for averaging
324
- python scripts/run_scenario.py --scenario S1 --runs 5 --output-dir ./s1-stats
325
- ```
@@ -1,200 +0,0 @@
1
- # Token Tracking Methodology
2
-
3
- How to measure and report token usage for CLEO grade sessions and A/B tests.
4
-
5
- ---
6
-
7
- ## Measurement Methods (priority order)
8
-
9
- ### Method 1: Claude Code Agent task notification (canonical)
10
-
11
- When running scenarios via Agent tasks, `total_tokens` is available in the task completion notification. This is the actual API token count — the most accurate available.
12
-
13
- ```python
14
- # Capture immediately when the agent task completes:
15
- timing = {
16
- "total_tokens": task.total_tokens, # actual API count
17
- "duration_ms": task.duration_ms,
18
- }
19
- # Write to timing.json — this data is EPHEMERAL and cannot be recovered later
20
- ```
21
-
22
- **Setup:** No configuration needed. Works whenever Claude Code spawns an Agent task.
23
-
24
- ### Method 2: OTel `claude_code.token.usage`
25
-
26
- When Claude Code is configured with OpenTelemetry, actual token counts are available at `~/.cleo/metrics/otel/`.
27
-
28
- ```bash
29
- # Enable (add to shell profile):
30
- export CLAUDE_CODE_ENABLE_TELEMETRY=1
31
- export OTEL_METRICS_EXPORTER=otlp
32
- export OTEL_EXPORTER_OTLP_PROTOCOL=http/json
33
- export OTEL_EXPORTER_OTLP_ENDPOINT="file://${HOME}/.cleo/metrics/otel/"
34
- ```
35
-
36
- Fields of interest from `claude_code.token.usage` metric:
37
- - `input_tokens` — tokens consumed by the prompt
38
- - `output_tokens` — tokens generated in response
39
- - `cache_read_input_tokens` — tokens served from cache
40
- - `session_id` — links to CLEO session
41
-
42
- ### Method 3: output_chars / 3.5 (JSON estimate)
43
-
44
- When neither task notification nor OTel is available, estimate from response character counts:
45
-
46
- ```
47
- estimated_tokens ≈ output_chars / 3.5 (JSON responses — denser than prose)
48
- estimated_tokens ≈ output_chars / 4 (mixed content)
49
- ```
50
-
51
- **Accuracy:** ±15-20% for typical JSON responses. Consistent enough for relative comparisons.
52
-
53
- ### Method 4: Audit entry count (coarse proxy)
54
-
55
- Each audit entry represents one operation invocation. As a very rough proxy:
56
- - One CLI read call ≈ 200–600 tokens total (request + response)
57
- - One CLI write call ≈ 300–800 tokens total
58
-
59
- `entryCount × 150` gives a session-level estimate. Accuracy ±50%.
60
-
61
- ---
62
-
63
- ## Token Fields in Results
64
-
65
- ### GRADES.jsonl token metadata
66
-
67
- The v2.1 grade scripts append `_tokenMeta` to each grade result:
68
-
69
- ```json
70
- {
71
- "sessionId": "session-abc123",
72
- "totalScore": 85,
73
- "_tokenMeta": {
74
- "estimationMethod": "otel",
75
- "totalEstimatedTokens": 4200,
76
- "inputTokens": 3100,
77
- "outputTokens": 1100,
78
- "cacheReadTokens": 800,
79
- "perDomain": {
80
- "tasks": 1800,
81
- "session": 600,
82
- "admin": 400,
83
- "memory": 0,
84
- "check": 0,
85
- "pipeline": 0,
86
- "orchestrate": 0,
87
- "tools": 200,
88
- "nexus": 0,
89
- "sticky": 0
90
- },
91
- "perInterface": {
92
- "cli": 3200,
93
- "untracked": 1000
94
- },
95
- "auditEntries": 47,
96
- "avgTokensPerEntry": 89
97
- }
98
- }
99
- ```
100
-
101
- ### A/B test token fields
102
-
103
- In `ab-result.json`:
104
- ```json
105
- {
106
- "operation": "tasks.find",
107
- "runs": [
108
- {
109
- "run": 1,
110
- "arm_a": {
111
- "output_chars": 1240,
112
- "estimated_tokens": 310,
113
- "duration_ms": 145
114
- },
115
- "arm_b": {
116
- "output_chars": 980,
117
- "estimated_tokens": 245,
118
- "duration_ms": 88
119
- },
120
- "token_delta": "+65",
121
- "token_delta_pct": "+26.5%"
122
- }
123
- ]
124
- }
125
- ```
126
-
127
- ---
128
-
129
- ## Per-Domain Token Estimation
130
-
131
- Typical token ranges per operation type (output only, output_chars / 4):
132
-
133
- | Domain | Operation | Typical output_chars | Est. tokens |
134
- |--------|-----------|---------------------|-------------|
135
- | tasks | `tasks.find` (10 results) | 2000–4000 | 500–1000 |
136
- | tasks | `tasks.show` (single) | 800–1500 | 200–375 |
137
- | tasks | `tasks.list` (full) | 5000–20000+ | 1250–5000+ |
138
- | session | `session.list` | 1000–3000 | 250–750 |
139
- | session | `session.status` | 400–800 | 100–200 |
140
- | admin | `admin.dash` | 1200–2500 | 300–625 |
141
- | admin | `admin.help` | 2000–5000 | 500–1250 |
142
- | memory | `memory.find` | 1500–4000 | 375–1000 |
143
-
144
- **Key insight:** `tasks.list` is 4–10x more expensive than `tasks.find` for same result set. This is why S2 Discovery Efficiency penalizes list-heavy agents.
145
-
146
- ---
147
-
148
- ## Using token_tracker.py
149
-
150
- ```bash
151
- # Estimate tokens for a specific session from GRADES.jsonl
152
- python scripts/token_tracker.py \
153
- --session-id "session-abc123" \
154
- --grades-file .cleo/metrics/GRADES.jsonl
155
-
156
- # Use OTEL data if available
157
- python scripts/token_tracker.py \
158
- --session-id "session-abc123" \
159
- --otel-dir ~/.cleo/metrics/otel \
160
- --grades-file .cleo/metrics/GRADES.jsonl
161
-
162
- # Aggregate breakdown across all sessions
163
- python scripts/token_tracker.py \
164
- --grades-file .cleo/metrics/GRADES.jsonl \
165
- --breakdown-by domain \
166
- --output domain-token-report.json
167
-
168
- # Compare two grade sessions
169
- python scripts/token_tracker.py \
170
- --compare session-abc123 session-def456 \
171
- --grades-file .cleo/metrics/GRADES.jsonl
172
- ```
173
-
174
- ---
175
-
176
- ## Token Efficiency Score
177
-
178
- The `generate_report.py` script computes a `tokenEfficiencyScore` for each session:
179
-
180
- ```
181
- tokenEfficiencyScore = (entriesCompleted / estimatedTokens) * 1000
182
- ```
183
-
184
- Higher = more work per token. Use to compare:
185
- - Pre/post skill improvements
186
- - Different agent configurations
187
- - Different CLI binary versions
188
-
189
- ---
190
-
191
- ## OTEL File Format
192
-
193
- OTEL files written by Claude Code are JSONL (one metric per line):
194
-
195
- ```json
196
- {"name": "claude_code.token.usage", "value": 4200, "attributes": {"session_id": "...", "model": "claude-sonnet-4-6"}, "timestamp": "2026-03-07T..."}
197
- {"name": "claude_code.api_request", "value": 1, "attributes": {"input_tokens": 3100, "output_tokens": 1100, "cache_read_input_tokens": 800}, "timestamp": "..."}
198
- ```
199
-
200
- `token_tracker.py` reads these files and matches `session_id` to CLEO session IDs stored in `GRADES.jsonl`.