@cleocode/skills 2026.5.82 → 2026.5.83
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +0 -1
- package/package.json +1 -1
- package/profiles/recommended.json +1 -1
- package/skills/manifest.json +45 -1
- package/skills/ct-grade-v2-1/MIGRATION.md +0 -28
- package/skills/ct-grade-v2-1/SKILL.md +0 -235
- package/skills/ct-grade-v2-1/agents/analysis-reporter.md +0 -203
- package/skills/ct-grade-v2-1/agents/blind-comparator.md +0 -157
- package/skills/ct-grade-v2-1/agents/scenario-runner.md +0 -160
- package/skills/ct-grade-v2-1/evals/evals.json +0 -74
- package/skills/ct-grade-v2-1/grade-viewer/__pycache__/build_op_stats.cpython-314.pyc +0 -0
- package/skills/ct-grade-v2-1/grade-viewer/__pycache__/generate_grade_review.cpython-314.pyc +0 -0
- package/skills/ct-grade-v2-1/grade-viewer/build_op_stats.py +0 -174
- package/skills/ct-grade-v2-1/grade-viewer/eval-analysis.json +0 -41
- package/skills/ct-grade-v2-1/grade-viewer/eval-report.md +0 -37
- package/skills/ct-grade-v2-1/grade-viewer/generate_grade_review.py +0 -1023
- package/skills/ct-grade-v2-1/grade-viewer/generate_grade_viewer.py +0 -548
- package/skills/ct-grade-v2-1/grade-viewer/grade-review-eval.html +0 -613
- package/skills/ct-grade-v2-1/grade-viewer/grade-review.html +0 -1532
- package/skills/ct-grade-v2-1/grade-viewer/viewer.html +0 -620
- package/skills/ct-grade-v2-1/manifest-entry.json +0 -31
- package/skills/ct-grade-v2-1/references/ab-testing.md +0 -173
- package/skills/ct-grade-v2-1/references/domains-ssot.md +0 -156
- package/skills/ct-grade-v2-1/references/grade-spec-v2.md +0 -167
- package/skills/ct-grade-v2-1/references/playbook-v2.md +0 -325
- package/skills/ct-grade-v2-1/references/token-tracking.md +0 -200
- package/skills/ct-grade-v2-1/scripts/generate_report.py +0 -419
- package/skills/ct-grade-v2-1/scripts/run_ab_test.py +0 -493
- package/skills/ct-grade-v2-1/scripts/run_scenario.py +0 -396
- package/skills/ct-grade-v2-1/scripts/setup_run.py +0 -207
- package/skills/ct-grade-v2-1/scripts/token_tracker.py +0 -175
|
@@ -1,325 +0,0 @@
|
|
|
1
|
-
# Grade Scenario Playbook v2
|
|
2
|
-
|
|
3
|
-
Parameterized test scenarios for CLEO grade system validation.
|
|
4
|
-
Updated for CLEO v2026.3+ operation names and 10-domain registry.
|
|
5
|
-
|
|
6
|
-
All operations use the CLI (`cleo` / `cleo-dev`). There is no MCP interface.
|
|
7
|
-
|
|
8
|
-
---
|
|
9
|
-
|
|
10
|
-
## Parameterization
|
|
11
|
-
|
|
12
|
-
All scenarios accept these parameters via run_scenario.py:
|
|
13
|
-
|
|
14
|
-
| Param | Default | Description |
|
|
15
|
-
|-------|---------|-------------|
|
|
16
|
-
| `--cleo` | `cleo-dev` | CLEO binary to use |
|
|
17
|
-
| `--scope` | `global` | Session scope |
|
|
18
|
-
| `--parent-task` | none | Parent task ID for subtask tests |
|
|
19
|
-
| `--output-dir` | `./grade-results` | Where to save results |
|
|
20
|
-
| `--runs` | 1 | Number of runs (for statistical averaging) |
|
|
21
|
-
| `--seed-task` | none | Pre-existing task ID to work against |
|
|
22
|
-
|
|
23
|
-
---
|
|
24
|
-
|
|
25
|
-
## S1: Session Discipline
|
|
26
|
-
|
|
27
|
-
**Rubric target:** S1 Session Discipline 20/20
|
|
28
|
-
|
|
29
|
-
**Operations (in order):**
|
|
30
|
-
|
|
31
|
-
```bash
|
|
32
|
-
1. cleo-dev session list
|
|
33
|
-
2. cleo-dev dash
|
|
34
|
-
3. cleo-dev find --status active
|
|
35
|
-
4. cleo-dev show <seed-task>
|
|
36
|
-
5. cleo-dev session end
|
|
37
|
-
```
|
|
38
|
-
|
|
39
|
-
**Pass criteria:**
|
|
40
|
-
- S1 = 20 (session.list before task ops AND session.end present)
|
|
41
|
-
- S2 >= 15 (only find used, no list)
|
|
42
|
-
- Flags: zero
|
|
43
|
-
|
|
44
|
-
**Anti-pattern (failing S1 = 0):**
|
|
45
|
-
```bash
|
|
46
|
-
# tasks.find BEFORE session.list
|
|
47
|
-
cleo-dev find --status active
|
|
48
|
-
cleo-dev session list # too late
|
|
49
|
-
# (no session.end)
|
|
50
|
-
```
|
|
51
|
-
|
|
52
|
-
---
|
|
53
|
-
|
|
54
|
-
## S2: Task Hygiene
|
|
55
|
-
|
|
56
|
-
**Rubric target:** S3 Task Hygiene 20/20
|
|
57
|
-
|
|
58
|
-
**Prerequisites:** `--parent-task` set to an existing task ID.
|
|
59
|
-
|
|
60
|
-
**Operations:**
|
|
61
|
-
```bash
|
|
62
|
-
1. cleo-dev session list
|
|
63
|
-
2. cleo-dev show <parent-task> # verify parent exists
|
|
64
|
-
3. cleo-dev add "Impl auth" --description "Add JWT auth to API endpoints" --parent <parent-task>
|
|
65
|
-
4. cleo-dev add "Write tests" --description "Unit tests for auth module"
|
|
66
|
-
5. cleo-dev session end
|
|
67
|
-
```
|
|
68
|
-
|
|
69
|
-
**Pass criteria:**
|
|
70
|
-
- S3 = 20 (all adds have descriptions, parent verified via show)
|
|
71
|
-
- S1 = 20
|
|
72
|
-
- Flags: zero
|
|
73
|
-
|
|
74
|
-
**Anti-pattern (S3 = 7):**
|
|
75
|
-
```bash
|
|
76
|
-
# No description, no exists check
|
|
77
|
-
cleo-dev add "Impl auth" --parent <id>
|
|
78
|
-
cleo-dev add "Write tests"
|
|
79
|
-
```
|
|
80
|
-
Expected deduction: -5 (no desc task 1) + -5 (no desc task 2) + -3 (no exists check) = 7/20.
|
|
81
|
-
|
|
82
|
-
---
|
|
83
|
-
|
|
84
|
-
## S3: Error Recovery
|
|
85
|
-
|
|
86
|
-
**Rubric target:** S4 Error Protocol 20/20
|
|
87
|
-
|
|
88
|
-
**Prerequisites:** `T99999` does NOT exist.
|
|
89
|
-
|
|
90
|
-
**Operations:**
|
|
91
|
-
```bash
|
|
92
|
-
1. cleo-dev session list
|
|
93
|
-
2. cleo-dev show T99999 # triggers E_NOT_FOUND (exit code 4)
|
|
94
|
-
3. cleo-dev find "T99999" # recovery lookup (must be within 4 ops)
|
|
95
|
-
4. cleo-dev add "New feature" --description "Feature not found, creating fresh"
|
|
96
|
-
5. cleo-dev session end
|
|
97
|
-
```
|
|
98
|
-
|
|
99
|
-
**Pass criteria:**
|
|
100
|
-
- S4 = 20 (E_NOT_FOUND followed by recovery; no duplicates)
|
|
101
|
-
- Evidence: `E_NOT_FOUND followed by recovery lookup`
|
|
102
|
-
- Flags: zero
|
|
103
|
-
|
|
104
|
-
**Anti-pattern (unrecovered, S4 = 15):**
|
|
105
|
-
```bash
|
|
106
|
-
cleo-dev show T99999 # E_NOT_FOUND
|
|
107
|
-
cleo-dev add "Something" --description "Unrelated" # NO recovery lookup
|
|
108
|
-
```
|
|
109
|
-
|
|
110
|
-
**Anti-pattern (duplicates, S4 = 15):**
|
|
111
|
-
```bash
|
|
112
|
-
cleo-dev add "Feature X" --description "First attempt"
|
|
113
|
-
cleo-dev add "Feature X" --description "Second attempt" # duplicate!
|
|
114
|
-
```
|
|
115
|
-
|
|
116
|
-
---
|
|
117
|
-
|
|
118
|
-
## S4: Full Lifecycle
|
|
119
|
-
|
|
120
|
-
**Rubric target:** All 5 dimensions 20/20 (total = 100)
|
|
121
|
-
|
|
122
|
-
**Prerequisites:** Known task `--seed-task` in pending status.
|
|
123
|
-
|
|
124
|
-
**Operations (in order):**
|
|
125
|
-
```bash
|
|
126
|
-
1. cleo-dev session list
|
|
127
|
-
2. cleo-dev help
|
|
128
|
-
3. cleo-dev dash
|
|
129
|
-
4. cleo-dev find --status pending
|
|
130
|
-
5. cleo-dev show <seed-task>
|
|
131
|
-
6. cleo-dev update <seed-task> --status active
|
|
132
|
-
# [agent performs work]
|
|
133
|
-
7. cleo-dev complete <seed-task>
|
|
134
|
-
8. cleo-dev find --status pending
|
|
135
|
-
9. cleo-dev session end --note "Completed <seed-task>"
|
|
136
|
-
```
|
|
137
|
-
|
|
138
|
-
**Pass criteria:**
|
|
139
|
-
- Total = 100, Grade = A
|
|
140
|
-
- Zero flags
|
|
141
|
-
- Entry count >= 10
|
|
142
|
-
- All 5 dimensions at 20/20
|
|
143
|
-
|
|
144
|
-
---
|
|
145
|
-
|
|
146
|
-
## S5: Multi-Domain Analysis
|
|
147
|
-
|
|
148
|
-
**Rubric target:** All 5 dimensions 20/20
|
|
149
|
-
|
|
150
|
-
**Prerequisites:** `--scope "epic:<parent-task>"` and epic has subtasks.
|
|
151
|
-
|
|
152
|
-
**Operations:**
|
|
153
|
-
```bash
|
|
154
|
-
1. cleo-dev session list
|
|
155
|
-
2. cleo-dev help
|
|
156
|
-
3. cleo-dev find --parent <parent-task>
|
|
157
|
-
4. cleo-dev show <subtask-id>
|
|
158
|
-
5. cleo-dev session context-drift
|
|
159
|
-
6. cleo-dev session decision-log --task <subtask-id>
|
|
160
|
-
7. cleo-dev session record-decision --task <subtask-id> --decision "Use adapter pattern" --rationale "Decouples provider logic"
|
|
161
|
-
8. cleo-dev update <subtask-id> --status active
|
|
162
|
-
9. cleo-dev complete <subtask-id>
|
|
163
|
-
10. cleo-dev find --parent <parent-task> --status pending
|
|
164
|
-
11. cleo-dev session end
|
|
165
|
-
```
|
|
166
|
-
|
|
167
|
-
**Pass criteria:**
|
|
168
|
-
- Total = 100, Grade = A
|
|
169
|
-
- Evidence of multi-domain ops (session, tasks, admin)
|
|
170
|
-
- Decision recorded
|
|
171
|
-
|
|
172
|
-
**Partial variation (S5 = 10 instead of 20):**
|
|
173
|
-
Skip step 2 (`admin.help`). Earns read-before-write +10 but not help/skill +10.
|
|
174
|
-
|
|
175
|
-
---
|
|
176
|
-
|
|
177
|
-
## S6: Memory Observe & Recall
|
|
178
|
-
|
|
179
|
-
**Rubric target:** S2 Task Efficiency 15+, S5 Progressive Disclosure 15+
|
|
180
|
-
|
|
181
|
-
**Operations (in order):**
|
|
182
|
-
```bash
|
|
183
|
-
1. cleo-dev session start --grade --name "grade-s6-memory-observe" --scope global
|
|
184
|
-
2. cleo-dev session list
|
|
185
|
-
3. cleo-dev observe "tasks.find is faster than tasks.list for large datasets" --title "Performance finding"
|
|
186
|
-
4. cleo-dev memory find "tasks.find faster"
|
|
187
|
-
5. cleo-dev memory timeline <returned-id> --before 2 --after 2
|
|
188
|
-
6. cleo-dev memory fetch <id>
|
|
189
|
-
7. cleo-dev session end
|
|
190
|
-
8. cleo-dev check grade --session "<saved-id>"
|
|
191
|
-
```
|
|
192
|
-
|
|
193
|
-
**Pass criteria:**
|
|
194
|
-
- S5 = 15+ (progressive disclosure via memory ops)
|
|
195
|
-
- S2 = 15+ (find used for retrieval, not broad list)
|
|
196
|
-
- Flags: zero
|
|
197
|
-
|
|
198
|
-
---
|
|
199
|
-
|
|
200
|
-
## S7: Decision Continuity
|
|
201
|
-
|
|
202
|
-
**Rubric target:** S1 Session Discipline 20, S5 Progressive Disclosure 15+
|
|
203
|
-
|
|
204
|
-
**Operations (in order):**
|
|
205
|
-
```bash
|
|
206
|
-
1. cleo-dev session start --grade --name "grade-s7-decision" --scope global
|
|
207
|
-
2. cleo-dev session list
|
|
208
|
-
3. cleo-dev memory decision store "Use adapter pattern for CLI abstraction" --rationale "Decouples interface from business logic" --confidence high
|
|
209
|
-
4. cleo-dev memory decision find "adapter pattern"
|
|
210
|
-
5. cleo-dev memory find "adapter pattern"
|
|
211
|
-
6. cleo-dev memory stats
|
|
212
|
-
7. cleo-dev session end
|
|
213
|
-
8. cleo-dev check grade --session "<saved-id>"
|
|
214
|
-
```
|
|
215
|
-
|
|
216
|
-
**Pass criteria:**
|
|
217
|
-
- S1 = 20 (session.list before ops)
|
|
218
|
-
- S5 = 15+ (progressive disclosure via memory ops)
|
|
219
|
-
- Flags: zero
|
|
220
|
-
|
|
221
|
-
---
|
|
222
|
-
|
|
223
|
-
## S8: Pattern & Learning Storage
|
|
224
|
-
|
|
225
|
-
**Rubric target:** S2 Task Efficiency 15+, S5 Progressive Disclosure 15+
|
|
226
|
-
|
|
227
|
-
**Operations (in order):**
|
|
228
|
-
```bash
|
|
229
|
-
1. cleo-dev session start --grade --name "grade-s8-patterns" --scope global
|
|
230
|
-
2. cleo-dev session list
|
|
231
|
-
3. cleo-dev memory pattern store "Call session.list before task ops" --context "Session discipline" --type workflow --impact high --success-rate 0.95
|
|
232
|
-
4. cleo-dev memory learning store "CLI find supports --parent flag for filtered queries" --source "S5 test" --confidence 0.9 --actionable
|
|
233
|
-
5. cleo-dev memory pattern find --type workflow --impact high
|
|
234
|
-
6. cleo-dev memory learning find --min-confidence 0.8 --actionable-only
|
|
235
|
-
7. cleo-dev session end
|
|
236
|
-
8. cleo-dev check grade --session "<saved-id>"
|
|
237
|
-
```
|
|
238
|
-
|
|
239
|
-
**Pass criteria:**
|
|
240
|
-
- S2 = 15+ (pattern.find/learning.find used, not broad list)
|
|
241
|
-
- S5 = 15+ (progressive disclosure via memory ops)
|
|
242
|
-
- Flags: zero
|
|
243
|
-
|
|
244
|
-
---
|
|
245
|
-
|
|
246
|
-
## S9: NEXUS Cross-Project Ops
|
|
247
|
-
|
|
248
|
-
**Rubric target:** S5 Progressive Disclosure 20
|
|
249
|
-
|
|
250
|
-
**Operations (in order):**
|
|
251
|
-
```bash
|
|
252
|
-
1. cleo-dev session start --grade --name "grade-s9-nexus" --scope global
|
|
253
|
-
2. cleo-dev session list
|
|
254
|
-
3. cleo-dev nexus status
|
|
255
|
-
4. cleo-dev nexus list
|
|
256
|
-
5. cleo-dev nexus show <first-project-id>
|
|
257
|
-
6. cleo-dev dash
|
|
258
|
-
7. cleo-dev session end
|
|
259
|
-
8. cleo-dev check grade --session "<saved-id>"
|
|
260
|
-
```
|
|
261
|
-
|
|
262
|
-
**Pass criteria:**
|
|
263
|
-
- S5 = 20 (cross-domain progressive disclosure)
|
|
264
|
-
- S1 = 20 (session.list first)
|
|
265
|
-
- Note: If nexus list returns empty, skip show and note "no projects registered"
|
|
266
|
-
|
|
267
|
-
---
|
|
268
|
-
|
|
269
|
-
## S10: Full System Throughput (8 domains)
|
|
270
|
-
|
|
271
|
-
**Rubric target:** S2 Task Efficiency 15+, S5 Progressive Disclosure 15+
|
|
272
|
-
|
|
273
|
-
**Operations (in order):**
|
|
274
|
-
```bash
|
|
275
|
-
1. cleo-dev session start --grade --name "grade-s10-throughput" --scope global
|
|
276
|
-
2. cleo-dev session list # session domain
|
|
277
|
-
3. cleo-dev help # admin domain
|
|
278
|
-
4. cleo-dev find --status active # tasks domain
|
|
279
|
-
5. cleo-dev memory find "decisions" # memory domain
|
|
280
|
-
6. cleo-dev nexus status # nexus domain
|
|
281
|
-
7. cleo-dev pipeline stage.status --epic <any-epic-id> # pipeline domain
|
|
282
|
-
8. cleo-dev health # check domain
|
|
283
|
-
9. cleo-dev skill list # tools domain
|
|
284
|
-
10. cleo-dev show <from-step-4>
|
|
285
|
-
11. cleo-dev observe "S10 throughput test complete" --title "Throughput"
|
|
286
|
-
12. cleo-dev session end
|
|
287
|
-
13. cleo-dev check grade --session "<saved-id>"
|
|
288
|
-
```
|
|
289
|
-
|
|
290
|
-
**Pass criteria:**
|
|
291
|
-
- 8 distinct domains hit in audit_log
|
|
292
|
-
- S2 = 15+ (tasks.find used, not tasks.list)
|
|
293
|
-
- S5 = 15+ (progressive disclosure across domains)
|
|
294
|
-
- Flags: zero
|
|
295
|
-
- Note: Step 7 pipeline.stage.status may return E_NOT_FOUND if no epicId — record the attempt, it still logs an audit entry
|
|
296
|
-
|
|
297
|
-
---
|
|
298
|
-
|
|
299
|
-
## Scoring Quick Reference
|
|
300
|
-
|
|
301
|
-
| Grade | Threshold | Typical profile |
|
|
302
|
-
|-------|-----------|-----------------|
|
|
303
|
-
| A | >= 90% | All dimensions near max, zero or minimal flags |
|
|
304
|
-
| B | >= 75% | Minor violations in one or two dimensions |
|
|
305
|
-
| C | >= 60% | Several protocol gaps |
|
|
306
|
-
| D | >= 45% | Multiple anti-patterns |
|
|
307
|
-
| F | < 45% | Severe protocol violations |
|
|
308
|
-
|
|
309
|
-
---
|
|
310
|
-
|
|
311
|
-
## Running Scenarios
|
|
312
|
-
|
|
313
|
-
```bash
|
|
314
|
-
# Single scenario
|
|
315
|
-
python scripts/run_scenario.py --scenario S3 --cleo cleo-dev
|
|
316
|
-
|
|
317
|
-
# Full suite
|
|
318
|
-
python scripts/run_scenario.py --scenario full --cleo cleo-dev --output-dir ./results
|
|
319
|
-
|
|
320
|
-
# With seed task
|
|
321
|
-
python scripts/run_scenario.py --scenario S4 --seed-task T200 --cleo cleo-dev
|
|
322
|
-
|
|
323
|
-
# Multiple runs for averaging
|
|
324
|
-
python scripts/run_scenario.py --scenario S1 --runs 5 --output-dir ./s1-stats
|
|
325
|
-
```
|
|
@@ -1,200 +0,0 @@
|
|
|
1
|
-
# Token Tracking Methodology
|
|
2
|
-
|
|
3
|
-
How to measure and report token usage for CLEO grade sessions and A/B tests.
|
|
4
|
-
|
|
5
|
-
---
|
|
6
|
-
|
|
7
|
-
## Measurement Methods (priority order)
|
|
8
|
-
|
|
9
|
-
### Method 1: Claude Code Agent task notification (canonical)
|
|
10
|
-
|
|
11
|
-
When running scenarios via Agent tasks, `total_tokens` is available in the task completion notification. This is the actual API token count — the most accurate available.
|
|
12
|
-
|
|
13
|
-
```python
|
|
14
|
-
# Capture immediately when the agent task completes:
|
|
15
|
-
timing = {
|
|
16
|
-
"total_tokens": task.total_tokens, # actual API count
|
|
17
|
-
"duration_ms": task.duration_ms,
|
|
18
|
-
}
|
|
19
|
-
# Write to timing.json — this data is EPHEMERAL and cannot be recovered later
|
|
20
|
-
```
|
|
21
|
-
|
|
22
|
-
**Setup:** No configuration needed. Works whenever Claude Code spawns an Agent task.
|
|
23
|
-
|
|
24
|
-
### Method 2: OTel `claude_code.token.usage`
|
|
25
|
-
|
|
26
|
-
When Claude Code is configured with OpenTelemetry, actual token counts are available at `~/.cleo/metrics/otel/`.
|
|
27
|
-
|
|
28
|
-
```bash
|
|
29
|
-
# Enable (add to shell profile):
|
|
30
|
-
export CLAUDE_CODE_ENABLE_TELEMETRY=1
|
|
31
|
-
export OTEL_METRICS_EXPORTER=otlp
|
|
32
|
-
export OTEL_EXPORTER_OTLP_PROTOCOL=http/json
|
|
33
|
-
export OTEL_EXPORTER_OTLP_ENDPOINT="file://${HOME}/.cleo/metrics/otel/"
|
|
34
|
-
```
|
|
35
|
-
|
|
36
|
-
Fields of interest from `claude_code.token.usage` metric:
|
|
37
|
-
- `input_tokens` — tokens consumed by the prompt
|
|
38
|
-
- `output_tokens` — tokens generated in response
|
|
39
|
-
- `cache_read_input_tokens` — tokens served from cache
|
|
40
|
-
- `session_id` — links to CLEO session
|
|
41
|
-
|
|
42
|
-
### Method 3: output_chars / 3.5 (JSON estimate)
|
|
43
|
-
|
|
44
|
-
When neither task notification nor OTel is available, estimate from response character counts:
|
|
45
|
-
|
|
46
|
-
```
|
|
47
|
-
estimated_tokens ≈ output_chars / 3.5 (JSON responses — denser than prose)
|
|
48
|
-
estimated_tokens ≈ output_chars / 4 (mixed content)
|
|
49
|
-
```
|
|
50
|
-
|
|
51
|
-
**Accuracy:** ±15-20% for typical JSON responses. Consistent enough for relative comparisons.
|
|
52
|
-
|
|
53
|
-
### Method 4: Audit entry count (coarse proxy)
|
|
54
|
-
|
|
55
|
-
Each audit entry represents one operation invocation. As a very rough proxy:
|
|
56
|
-
- One CLI read call ≈ 200–600 tokens total (request + response)
|
|
57
|
-
- One CLI write call ≈ 300–800 tokens total
|
|
58
|
-
|
|
59
|
-
`entryCount × 150` gives a session-level estimate. Accuracy ±50%.
|
|
60
|
-
|
|
61
|
-
---
|
|
62
|
-
|
|
63
|
-
## Token Fields in Results
|
|
64
|
-
|
|
65
|
-
### GRADES.jsonl token metadata
|
|
66
|
-
|
|
67
|
-
The v2.1 grade scripts append `_tokenMeta` to each grade result:
|
|
68
|
-
|
|
69
|
-
```json
|
|
70
|
-
{
|
|
71
|
-
"sessionId": "session-abc123",
|
|
72
|
-
"totalScore": 85,
|
|
73
|
-
"_tokenMeta": {
|
|
74
|
-
"estimationMethod": "otel",
|
|
75
|
-
"totalEstimatedTokens": 4200,
|
|
76
|
-
"inputTokens": 3100,
|
|
77
|
-
"outputTokens": 1100,
|
|
78
|
-
"cacheReadTokens": 800,
|
|
79
|
-
"perDomain": {
|
|
80
|
-
"tasks": 1800,
|
|
81
|
-
"session": 600,
|
|
82
|
-
"admin": 400,
|
|
83
|
-
"memory": 0,
|
|
84
|
-
"check": 0,
|
|
85
|
-
"pipeline": 0,
|
|
86
|
-
"orchestrate": 0,
|
|
87
|
-
"tools": 200,
|
|
88
|
-
"nexus": 0,
|
|
89
|
-
"sticky": 0
|
|
90
|
-
},
|
|
91
|
-
"perInterface": {
|
|
92
|
-
"cli": 3200,
|
|
93
|
-
"untracked": 1000
|
|
94
|
-
},
|
|
95
|
-
"auditEntries": 47,
|
|
96
|
-
"avgTokensPerEntry": 89
|
|
97
|
-
}
|
|
98
|
-
}
|
|
99
|
-
```
|
|
100
|
-
|
|
101
|
-
### A/B test token fields
|
|
102
|
-
|
|
103
|
-
In `ab-result.json`:
|
|
104
|
-
```json
|
|
105
|
-
{
|
|
106
|
-
"operation": "tasks.find",
|
|
107
|
-
"runs": [
|
|
108
|
-
{
|
|
109
|
-
"run": 1,
|
|
110
|
-
"arm_a": {
|
|
111
|
-
"output_chars": 1240,
|
|
112
|
-
"estimated_tokens": 310,
|
|
113
|
-
"duration_ms": 145
|
|
114
|
-
},
|
|
115
|
-
"arm_b": {
|
|
116
|
-
"output_chars": 980,
|
|
117
|
-
"estimated_tokens": 245,
|
|
118
|
-
"duration_ms": 88
|
|
119
|
-
},
|
|
120
|
-
"token_delta": "+65",
|
|
121
|
-
"token_delta_pct": "+26.5%"
|
|
122
|
-
}
|
|
123
|
-
]
|
|
124
|
-
}
|
|
125
|
-
```
|
|
126
|
-
|
|
127
|
-
---
|
|
128
|
-
|
|
129
|
-
## Per-Domain Token Estimation
|
|
130
|
-
|
|
131
|
-
Typical token ranges per operation type (output only, output_chars / 4):
|
|
132
|
-
|
|
133
|
-
| Domain | Operation | Typical output_chars | Est. tokens |
|
|
134
|
-
|--------|-----------|---------------------|-------------|
|
|
135
|
-
| tasks | `tasks.find` (10 results) | 2000–4000 | 500–1000 |
|
|
136
|
-
| tasks | `tasks.show` (single) | 800–1500 | 200–375 |
|
|
137
|
-
| tasks | `tasks.list` (full) | 5000–20000+ | 1250–5000+ |
|
|
138
|
-
| session | `session.list` | 1000–3000 | 250–750 |
|
|
139
|
-
| session | `session.status` | 400–800 | 100–200 |
|
|
140
|
-
| admin | `admin.dash` | 1200–2500 | 300–625 |
|
|
141
|
-
| admin | `admin.help` | 2000–5000 | 500–1250 |
|
|
142
|
-
| memory | `memory.find` | 1500–4000 | 375–1000 |
|
|
143
|
-
|
|
144
|
-
**Key insight:** `tasks.list` is 4–10x more expensive than `tasks.find` for same result set. This is why S2 Discovery Efficiency penalizes list-heavy agents.
|
|
145
|
-
|
|
146
|
-
---
|
|
147
|
-
|
|
148
|
-
## Using token_tracker.py
|
|
149
|
-
|
|
150
|
-
```bash
|
|
151
|
-
# Estimate tokens for a specific session from GRADES.jsonl
|
|
152
|
-
python scripts/token_tracker.py \
|
|
153
|
-
--session-id "session-abc123" \
|
|
154
|
-
--grades-file .cleo/metrics/GRADES.jsonl
|
|
155
|
-
|
|
156
|
-
# Use OTEL data if available
|
|
157
|
-
python scripts/token_tracker.py \
|
|
158
|
-
--session-id "session-abc123" \
|
|
159
|
-
--otel-dir ~/.cleo/metrics/otel \
|
|
160
|
-
--grades-file .cleo/metrics/GRADES.jsonl
|
|
161
|
-
|
|
162
|
-
# Aggregate breakdown across all sessions
|
|
163
|
-
python scripts/token_tracker.py \
|
|
164
|
-
--grades-file .cleo/metrics/GRADES.jsonl \
|
|
165
|
-
--breakdown-by domain \
|
|
166
|
-
--output domain-token-report.json
|
|
167
|
-
|
|
168
|
-
# Compare two grade sessions
|
|
169
|
-
python scripts/token_tracker.py \
|
|
170
|
-
--compare session-abc123 session-def456 \
|
|
171
|
-
--grades-file .cleo/metrics/GRADES.jsonl
|
|
172
|
-
```
|
|
173
|
-
|
|
174
|
-
---
|
|
175
|
-
|
|
176
|
-
## Token Efficiency Score
|
|
177
|
-
|
|
178
|
-
The `generate_report.py` script computes a `tokenEfficiencyScore` for each session:
|
|
179
|
-
|
|
180
|
-
```
|
|
181
|
-
tokenEfficiencyScore = (entriesCompleted / estimatedTokens) * 1000
|
|
182
|
-
```
|
|
183
|
-
|
|
184
|
-
Higher = more work per token. Use to compare:
|
|
185
|
-
- Pre/post skill improvements
|
|
186
|
-
- Different agent configurations
|
|
187
|
-
- Different CLI binary versions
|
|
188
|
-
|
|
189
|
-
---
|
|
190
|
-
|
|
191
|
-
## OTEL File Format
|
|
192
|
-
|
|
193
|
-
OTEL files written by Claude Code are JSONL (one metric per line):
|
|
194
|
-
|
|
195
|
-
```json
|
|
196
|
-
{"name": "claude_code.token.usage", "value": 4200, "attributes": {"session_id": "...", "model": "claude-sonnet-4-6"}, "timestamp": "2026-03-07T..."}
|
|
197
|
-
{"name": "claude_code.api_request", "value": 1, "attributes": {"input_tokens": 3100, "output_tokens": 1100, "cache_read_input_tokens": 800}, "timestamp": "..."}
|
|
198
|
-
```
|
|
199
|
-
|
|
200
|
-
`token_tracker.py` reads these files and matches `session_id` to CLEO session IDs stored in `GRADES.jsonl`.
|