@cleocode/skills 2026.5.82 → 2026.5.84
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +0 -1
- package/package.json +1 -1
- package/profiles/recommended.json +1 -1
- package/skills/_shared/__tests__/lifecycle-protocol-reconcile.test.ts +112 -0
- package/skills/_shared/__tests__/loom-adr-links.test.ts +163 -0
- package/skills/_shared/__tests__/loom-stage-coverage.test.ts +167 -0
- package/skills/ct-adr-recorder/SKILL.md +18 -0
- package/skills/ct-consensus-voter/SKILL.md +14 -0
- package/skills/ct-contribution/SKILL.md +80 -0
- package/skills/ct-epic-architect/SKILL.md +15 -0
- package/skills/ct-ivt-looper/SKILL.md +32 -0
- package/skills/ct-release-orchestrator/SKILL.md +16 -0
- package/skills/ct-research-agent/SKILL.md +15 -0
- package/skills/ct-spec-writer/SKILL.md +15 -0
- package/skills/ct-task-executor/SKILL.md +15 -0
- package/skills/ct-validator/SKILL.md +35 -0
- package/skills/manifest.json +81 -9
- package/skills/ct-grade-v2-1/MIGRATION.md +0 -28
- package/skills/ct-grade-v2-1/SKILL.md +0 -235
- package/skills/ct-grade-v2-1/agents/analysis-reporter.md +0 -203
- package/skills/ct-grade-v2-1/agents/blind-comparator.md +0 -157
- package/skills/ct-grade-v2-1/agents/scenario-runner.md +0 -160
- package/skills/ct-grade-v2-1/evals/evals.json +0 -74
- package/skills/ct-grade-v2-1/grade-viewer/__pycache__/build_op_stats.cpython-314.pyc +0 -0
- package/skills/ct-grade-v2-1/grade-viewer/__pycache__/generate_grade_review.cpython-314.pyc +0 -0
- package/skills/ct-grade-v2-1/grade-viewer/build_op_stats.py +0 -174
- package/skills/ct-grade-v2-1/grade-viewer/eval-analysis.json +0 -41
- package/skills/ct-grade-v2-1/grade-viewer/eval-report.md +0 -37
- package/skills/ct-grade-v2-1/grade-viewer/generate_grade_review.py +0 -1023
- package/skills/ct-grade-v2-1/grade-viewer/generate_grade_viewer.py +0 -548
- package/skills/ct-grade-v2-1/grade-viewer/grade-review-eval.html +0 -613
- package/skills/ct-grade-v2-1/grade-viewer/grade-review.html +0 -1532
- package/skills/ct-grade-v2-1/grade-viewer/viewer.html +0 -620
- package/skills/ct-grade-v2-1/manifest-entry.json +0 -31
- package/skills/ct-grade-v2-1/references/ab-testing.md +0 -173
- package/skills/ct-grade-v2-1/references/domains-ssot.md +0 -156
- package/skills/ct-grade-v2-1/references/grade-spec-v2.md +0 -167
- package/skills/ct-grade-v2-1/references/playbook-v2.md +0 -325
- package/skills/ct-grade-v2-1/references/token-tracking.md +0 -200
- package/skills/ct-grade-v2-1/scripts/generate_report.py +0 -419
- package/skills/ct-grade-v2-1/scripts/run_ab_test.py +0 -493
- package/skills/ct-grade-v2-1/scripts/run_scenario.py +0 -396
- package/skills/ct-grade-v2-1/scripts/setup_run.py +0 -207
- package/skills/ct-grade-v2-1/scripts/token_tracker.py +0 -175
|
@@ -1,6 +1,11 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: ct-ivt-looper
|
|
3
3
|
description: "Runs a project-agnostic autonomous Implement-then-Validate-then-Test compliance loop on any git worktree. Detects the project's test framework (vitest, jest, mocha, pytest, unittest, go-test, cargo-test, rspec, phpunit, bats, or other) and iterates until the implementation satisfies its specification, recording convergence metrics to the manifest. Use when given an implementation task that must ship verified: the IVT loop is the autonomous compliance layer enforced before any release or PR. Triggers on phrases like 'implement and verify', 'run the IVT loop', 'ship this task', 'complete implementation with tests', 'verify against spec', or any implementation task with acceptance criteria. Works in any git worktree regardless of language or framework, never hardcoded to one project's tooling."
|
|
4
|
+
protocol: testing
|
|
5
|
+
loomStage: testing
|
|
6
|
+
adrRefs:
|
|
7
|
+
- ADR-051
|
|
8
|
+
- ADR-061
|
|
4
9
|
---
|
|
5
10
|
|
|
6
11
|
# IVT Looper
|
|
@@ -76,6 +81,24 @@ escalate_to_hitl() # IVT-007: exit code 65
|
|
|
76
81
|
|
|
77
82
|
The loop is a *single* stage from the lifecycle's point of view. Implement, Validate, and Test are not three separate tasks — they are three phases of one autonomous run that either converges or escalates.
|
|
78
83
|
|
|
84
|
+
## Out of Scope (T9675)
|
|
85
|
+
|
|
86
|
+
`ct-ivt-looper` operates on the **`testing`** LOOM lifecycle stage (stage 8). It performs the **dynamic** Implement-then-Validate-then-Test loop with framework detection and iterate-until-green convergence semantics.
|
|
87
|
+
|
|
88
|
+
This skill does NOT:
|
|
89
|
+
|
|
90
|
+
- Audit static artifacts for schema/compliance/RFC-2119 keyword usage, ADR-document structure, or JSON Schema conformance. Those belong to **`ct-validator`** at the `validation` stage (stage 7). When the question is "is this document/manifest well-formed?" rather than "does this code converge on its spec?", chain to `ct-validator` rather than expanding scope here.
|
|
91
|
+
- Promote a green loop to release. Release sequencing belongs to **`ct-release-orchestrator`** at stage 9.
|
|
92
|
+
|
|
93
|
+
### Chain handoffs
|
|
94
|
+
|
|
95
|
+
| Direction | When | Handoff |
|
|
96
|
+
|---|---|---|
|
|
97
|
+
| `ct-ivt-looper` → `ct-validator` | Loop converged; need to audit the resulting artifacts (e.g. final manifest, spec back-references) against schema/compliance | Emit the convergence manifest entry, then dispatch the `validation` stage |
|
|
98
|
+
| `ct-validator` → `ct-ivt-looper` | Spec is valid but implementation needs dynamic verification | Receive a dispatch from the `validation` stage; iterate the IVT loop on the worktree |
|
|
99
|
+
|
|
100
|
+
Governance: see **ADR-051** (programmatic gate integrity) which defines the evidence atoms (`tool:test`, `test-run:<json>`) the loop emits and that downstream `cleo verify --gate testsPassed` re-validates, and **ADR-061** (project-agnostic verify tools) which defines the canonical tool-resolution layer.
|
|
101
|
+
|
|
79
102
|
## Framework Detection
|
|
80
103
|
|
|
81
104
|
Framework detection is project-agnostic: the skill walks the worktree, inspects config files, and selects the correct test command. No language or framework is special-cased above another. The full detection table lives in [references/frameworks.md](references/frameworks.md). In summary: detection reads the project manifest (`package.json`, `pyproject.toml`, `Cargo.toml`, `go.mod`, `Gemfile`, `composer.json`, or `.cleo/project-context.json#testing.command`) and selects one of: `vitest`, `jest`, `mocha`, `pytest`, `unittest`, `go-test`, `cargo-test`, `rspec`, `phpunit`, `bats`, `other`.
|
|
@@ -179,3 +202,12 @@ cleo check protocol \
|
|
|
179
202
|
6. Record `framework`, `testsRun`, `testsPassed`, `testsFailed`, `ivtLoopConverged`, `ivtLoopIterations` in the manifest.
|
|
180
203
|
7. On non-convergence, exit 65 and leave the worktree untouched.
|
|
181
204
|
8. Validate every run via `cleo check protocol --protocolType testing`.
|
|
205
|
+
|
|
206
|
+
## See also / References
|
|
207
|
+
|
|
208
|
+
This skill binds to the **testing** LOOM lifecycle stage. Governing ADRs:
|
|
209
|
+
|
|
210
|
+
- [ADR-051 — programmatic gate integrity](../../../../.cleo/adrs/ADR-051-programmatic-gate-integrity.md) — defines the evidence atoms (`tool:test`, `test-run:<json>`) that the IVT loop emits and that downstream `cleo verify --gate testsPassed` re-validates.
|
|
211
|
+
- [ADR-061 — project-agnostic verify tools](../../../../.cleo/adrs/ADR-061-project-agnostic-verify-tools.md) — defines the canonical tool-resolution layer (`test`, `build`, `lint`, `typecheck`) that the loop walks for framework-agnostic execution.
|
|
212
|
+
|
|
213
|
+
LOOM coverage matrix: [docs/skills/loom-coverage-matrix.md](../../../../docs/skills/loom-coverage-matrix.md).
|
|
@@ -1,6 +1,12 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: ct-release-orchestrator
|
|
3
3
|
description: "Orchestrates the full release pipeline: version bump, then changelog, then commit, then tag, then conditionally forks to artifact-publish and provenance based on release config. Parent protocol that composes ct-artifact-publisher and ct-provenance-keeper as sub-protocols: not every release publishes artifacts (source-only releases skip it), and artifact publishers delegate signing and attestation to provenance. Use when shipping a new version, running cleo release ship, or promoting a completed epic to released status."
|
|
4
|
+
protocol: release
|
|
5
|
+
loomStage: release
|
|
6
|
+
adrRefs:
|
|
7
|
+
- ADR-053
|
|
8
|
+
- ADR-063
|
|
9
|
+
- ADR-065
|
|
4
10
|
---
|
|
5
11
|
|
|
6
12
|
# Release Orchestrator
|
|
@@ -132,3 +138,13 @@ For source-only releases, pass `--no-artifacts` to skip the artifact-publish han
|
|
|
132
138
|
6. `released` entries are immutable; hotfixes go into new entries.
|
|
133
139
|
7. Manifest entry MUST set `agent_type: "documentation"` and record the full chain via `record_release()`.
|
|
134
140
|
8. Always validate via `cleo check protocol --protocolType release` before declaring the release done.
|
|
141
|
+
|
|
142
|
+
## See also / References
|
|
143
|
+
|
|
144
|
+
This skill binds to the **release** LOOM lifecycle stage. Governing ADRs:
|
|
145
|
+
|
|
146
|
+
- [ADR-053 — project-agnostic release pipeline](../../../../.cleo/adrs/ADR-053-project-agnostic-release-pipeline.md) — defines the language-agnostic version bump → changelog → tag flow.
|
|
147
|
+
- [ADR-063 — release pipeline](../../../../.cleo/adrs/ADR-063-release-pipeline.md) — defines the 12-step `cleo release ship` integration with CI.
|
|
148
|
+
- [ADR-065 — PR-required release flow](../../../../.cleo/adrs/ADR-065-pr-required-release-flow.md) — defines the PR-gated path; direct pushes to `main` are prohibited.
|
|
149
|
+
|
|
150
|
+
LOOM coverage matrix: [docs/skills/loom-coverage-matrix.md](../../../../docs/skills/loom-coverage-matrix.md).
|
|
@@ -6,6 +6,10 @@ tier: 2
|
|
|
6
6
|
core: false
|
|
7
7
|
category: recommended
|
|
8
8
|
protocol: research
|
|
9
|
+
loomStage: research
|
|
10
|
+
adrRefs:
|
|
11
|
+
- ADR-023
|
|
12
|
+
- ADR-070
|
|
9
13
|
dependencies: []
|
|
10
14
|
sharedResources:
|
|
11
15
|
- subagent-protocol-base
|
|
@@ -224,3 +228,14 @@ If research cannot proceed (access denied, topic too broad, etc.):
|
|
|
224
228
|
- **Prioritized** - Most important first
|
|
225
229
|
- **Justified** - Tied to specific findings
|
|
226
230
|
- **Feasible** - Achievable within project constraints
|
|
231
|
+
|
|
232
|
+
---
|
|
233
|
+
|
|
234
|
+
## See also / References
|
|
235
|
+
|
|
236
|
+
This skill binds to the **research** LOOM lifecycle stage. Governing ADRs:
|
|
237
|
+
|
|
238
|
+
- [ADR-023 — protocol validation dispatch](../../../../.cleo/adrs/ADR-023-protocol-validation-dispatch.md) — defines how research output is validated before downstream stages consume it.
|
|
239
|
+
- [ADR-070 — three-tier orchestration](../../../../.cleo/adrs/ADR-070-three-tier-orchestration.md) — defines the Orchestrator → Phase Lead → Worker tiers; research runs as a leaf Worker.
|
|
240
|
+
|
|
241
|
+
LOOM coverage matrix: [docs/skills/loom-coverage-matrix.md](../../../../docs/skills/loom-coverage-matrix.md).
|
|
@@ -6,6 +6,10 @@ tier: 2
|
|
|
6
6
|
core: false
|
|
7
7
|
category: recommended
|
|
8
8
|
protocol: specification
|
|
9
|
+
loomStage: specification
|
|
10
|
+
adrRefs:
|
|
11
|
+
- ADR-014
|
|
12
|
+
- ADR-023
|
|
9
13
|
dependencies: []
|
|
10
14
|
sharedResources:
|
|
11
15
|
- subagent-protocol-base
|
|
@@ -187,3 +191,14 @@ Specifications go in: `docs/specs/{{SPEC_NAME}}.md`
|
|
|
187
191
|
- [ ] Manifest entry appended
|
|
188
192
|
- [ ] Task completed via `{{TASK_COMPLETE_CMD}}`
|
|
189
193
|
- [ ] Return summary message only
|
|
194
|
+
|
|
195
|
+
---
|
|
196
|
+
|
|
197
|
+
## See also / References
|
|
198
|
+
|
|
199
|
+
This skill binds to the **specification** LOOM lifecycle stage. Governing ADRs:
|
|
200
|
+
|
|
201
|
+
- [ADR-014 — RCASD rename and protocol validation](../../../../.cleo/adrs/ADR-014-rcasd-rename-and-protocol-validation.md) — defines the specification stage's role inside the RCASD-IVTR+C lifecycle.
|
|
202
|
+
- [ADR-023 — protocol validation dispatch](../../../../.cleo/adrs/ADR-023-protocol-validation-dispatch.md) — defines how specifications are validated before decomposition.
|
|
203
|
+
|
|
204
|
+
LOOM coverage matrix: [docs/skills/loom-coverage-matrix.md](../../../../docs/skills/loom-coverage-matrix.md).
|
|
@@ -6,6 +6,10 @@ tier: 2
|
|
|
6
6
|
core: true
|
|
7
7
|
category: core
|
|
8
8
|
protocol: implementation
|
|
9
|
+
loomStage: implementation
|
|
10
|
+
adrRefs:
|
|
11
|
+
- ADR-070
|
|
12
|
+
- ADR-062
|
|
9
13
|
dependencies: []
|
|
10
14
|
sharedResources:
|
|
11
15
|
- subagent-protocol-base
|
|
@@ -294,3 +298,14 @@ cleo session gc --include-active
|
|
|
294
298
|
| Partial deliverables | Missing outputs | Complete all or report partial |
|
|
295
299
|
| Undocumented changes | Lost context | Write detailed output file |
|
|
296
300
|
| Silent failures | Orchestrator unaware | Report via manifest status |
|
|
301
|
+
|
|
302
|
+
---
|
|
303
|
+
|
|
304
|
+
## See also / References
|
|
305
|
+
|
|
306
|
+
This skill binds to the **implementation** LOOM lifecycle stage. Governing ADRs:
|
|
307
|
+
|
|
308
|
+
- [ADR-070 — three-tier orchestration](../../../../.cleo/adrs/ADR-070-three-tier-orchestration.md) — defines the Worker tier that ct-task-executor occupies.
|
|
309
|
+
- [ADR-062 — worktree merge, not cherry-pick](../../../../.cleo/adrs/ADR-062-worktree-merge-not-cherry-pick.md) — defines the integration path that preserves the executor's commit SHAs end-to-end.
|
|
310
|
+
|
|
311
|
+
LOOM coverage matrix: [docs/skills/loom-coverage-matrix.md](../../../../docs/skills/loom-coverage-matrix.md).
|
|
@@ -6,6 +6,10 @@ tier: 2
|
|
|
6
6
|
core: false
|
|
7
7
|
category: recommended
|
|
8
8
|
protocol: validation
|
|
9
|
+
loomStage: validation
|
|
10
|
+
adrRefs:
|
|
11
|
+
- ADR-051
|
|
12
|
+
- ADR-023
|
|
9
13
|
dependencies: []
|
|
10
14
|
sharedResources:
|
|
11
15
|
- subagent-protocol-base
|
|
@@ -41,6 +45,26 @@ Context injection for compliance validation tasks spawned via cleo-subagent. Pro
|
|
|
41
45
|
|
|
42
46
|
---
|
|
43
47
|
|
|
48
|
+
## Out of Scope (T9675)
|
|
49
|
+
|
|
50
|
+
`ct-validator` operates on the **`validation`** LOOM lifecycle stage (stage 7). It performs **static** schema, compliance, and audit checks against artifacts that already exist on disk (specs, ADRs, JSON files, RFC 2119 keyword usage, manifest schemas).
|
|
51
|
+
|
|
52
|
+
This skill does NOT:
|
|
53
|
+
|
|
54
|
+
- Run a test suite, framework detection, or iterative IVT loop. Those belong to **`ct-ivt-looper`** at the `testing` stage (stage 8). When dynamic verification is required — e.g. "does the implementation actually pass its tests?" — chain to `ct-ivt-looper` rather than expanding scope here.
|
|
55
|
+
- Modify code or apply fixes. The validator reports; downstream skills remediate.
|
|
56
|
+
|
|
57
|
+
### Chain handoffs
|
|
58
|
+
|
|
59
|
+
| Direction | When | Handoff |
|
|
60
|
+
|---|---|---|
|
|
61
|
+
| `ct-validator` → `ct-ivt-looper` | Spec is valid but implementation needs dynamic verification | Emit a manifest entry, then dispatch the `testing` stage |
|
|
62
|
+
| `ct-ivt-looper` → `ct-validator` | IVT loop converged; need to audit the resulting artifacts against schema/compliance | Dispatch the `validation` stage after the test convergence record |
|
|
63
|
+
|
|
64
|
+
Governance: see **ADR-051** (programmatic gate integrity) which defines the evidence atoms each stage emits and that the other stage may re-validate, and **ADR-023** (protocol validation dispatch) which routes between them.
|
|
65
|
+
|
|
66
|
+
---
|
|
67
|
+
|
|
44
68
|
## Validation Methodology
|
|
45
69
|
|
|
46
70
|
### Standard Workflow
|
|
@@ -214,3 +238,14 @@ When invoked by orchestrator, expect these context tokens:
|
|
|
214
238
|
| Vague findings | Unclear remediation | Specific issue + file/line + fix |
|
|
215
239
|
| Missing severity | Can't prioritize | Always classify: critical/warning/suggestion |
|
|
216
240
|
| No remediation | Findings not actionable | Always provide fix for FAIL/PARTIAL |
|
|
241
|
+
|
|
242
|
+
---
|
|
243
|
+
|
|
244
|
+
## See also / References
|
|
245
|
+
|
|
246
|
+
This skill binds to the **validation** LOOM lifecycle stage. Governing ADRs:
|
|
247
|
+
|
|
248
|
+
- [ADR-051 — programmatic gate integrity](../../../../.cleo/adrs/ADR-051-programmatic-gate-integrity.md) — defines the evidence-atom grammar that the validator emits and re-validates.
|
|
249
|
+
- [ADR-023 — protocol validation dispatch](../../../../.cleo/adrs/ADR-023-protocol-validation-dispatch.md) — defines the protocol-validation routing layer that dispatches to this skill.
|
|
250
|
+
|
|
251
|
+
LOOM coverage matrix: [docs/skills/loom-coverage-matrix.md](../../../../docs/skills/loom-coverage-matrix.md).
|
package/skills/manifest.json
CHANGED
|
@@ -1,11 +1,13 @@
|
|
|
1
1
|
{
|
|
2
2
|
"$schema": "https://cleo-dev.com/schemas/v1/skills-manifest.schema.json",
|
|
3
3
|
"_meta": {
|
|
4
|
-
"schemaVersion": "2.
|
|
5
|
-
"lastUpdated": "2026-
|
|
6
|
-
"totalSkills":
|
|
7
|
-
"generatedFrom": "T260 — lifecycle pipeline rework: dedicated skills for ADR, IVT loop, consensus, release, artifact-publish, provenance",
|
|
8
|
-
"architectureNote": "Universal Subagent Architecture: All spawns use provider-neutral delegation with skill/protocol injection. Pipeline stages and cross-cutting protocols each have a dedicated skill — no overloading."
|
|
4
|
+
"schemaVersion": "2.7.0",
|
|
5
|
+
"lastUpdated": "2026-05-19",
|
|
6
|
+
"totalSkills": 22,
|
|
7
|
+
"generatedFrom": "T260 — lifecycle pipeline rework: dedicated skills for ADR, IVT loop, consensus, release, artifact-publish, provenance. T9664/T9665/T9672 (epic T9568) — adds loomStage + adrRefs fields on every LOOM-stage skill entry (underscored canonical form matching `cleo lifecycle` stage names; ADR refs point at .cleo/adrs/). T9672 reconciled dispatch_matrix.by_protocol key 'architecture-decision' → 'architecture_decision' matching the lifecycle CLI source of truth; the dashed form is retained as a keyword alias.",
|
|
8
|
+
"architectureNote": "Universal Subagent Architecture: All spawns use provider-neutral delegation with skill/protocol injection. Pipeline stages and cross-cutting protocols each have a dedicated skill — no overloading.",
|
|
9
|
+
"loomStageContract": "Each of the 10 LOOM lifecycle stages (research, consensus, architecture_decision, specification, decomposition, implementation, validation, testing, release, contribution) MUST have a bound skill whose manifest entry declares loomStage matching the lifecycle CLI stage name. The legacy `protocol` field is retained as an alias; dispatch_matrix.by_protocol is the authoritative routing table. See docs/skills/loom-coverage-matrix.md.",
|
|
10
|
+
"adrRefsContract": "Each LOOM-stage skill MUST declare an adrRefs[] array naming the ADR file(s) that govern the stage. Every referenced ADR-NNN MUST resolve to a file under .cleo/adrs/. The Vitest gate at packages/skills/skills/_shared/__tests__/loom-adr-links.test.ts enforces both conditions."
|
|
9
11
|
},
|
|
10
12
|
"dispatch_matrix": {
|
|
11
13
|
"_comment": "Maps task types/keywords to skill NAMES. Provider adapter decides HOW to execute.",
|
|
@@ -33,7 +35,7 @@
|
|
|
33
35
|
"spec|rfc|protocol|contract": "ct-spec-writer",
|
|
34
36
|
"validate|verify|audit|compliance": "ct-validator",
|
|
35
37
|
"consensus|vote|verdict|resolve the debate": "ct-consensus-voter",
|
|
36
|
-
"adr|architecture decision|formalize|lock in the choice": "ct-adr-recorder",
|
|
38
|
+
"adr|architecture decision|architecture-decision|formalize|lock in the choice": "ct-adr-recorder",
|
|
37
39
|
"release|version|ship|changelog|cut release": "ct-release-orchestrator",
|
|
38
40
|
"artifact|publish|registry|npm publish|docker push": "ct-artifact-publisher",
|
|
39
41
|
"provenance|attestation|sbom|sigstore|slsa": "ct-provenance-keeper"
|
|
@@ -41,7 +43,7 @@
|
|
|
41
43
|
"by_protocol": {
|
|
42
44
|
"research": "ct-research-agent",
|
|
43
45
|
"consensus": "ct-consensus-voter",
|
|
44
|
-
"
|
|
46
|
+
"architecture_decision": "ct-adr-recorder",
|
|
45
47
|
"specification": "ct-spec-writer",
|
|
46
48
|
"decomposition": "ct-epic-architect",
|
|
47
49
|
"implementation": "ct-task-executor",
|
|
@@ -120,6 +122,9 @@
|
|
|
120
122
|
"status": "active",
|
|
121
123
|
"tier": 0,
|
|
122
124
|
"token_budget": 8000,
|
|
125
|
+
"protocol": "implementation",
|
|
126
|
+
"loomStage": "implementation",
|
|
127
|
+
"adrRefs": ["ADR-070", "ADR-062"],
|
|
123
128
|
"references": [],
|
|
124
129
|
"capabilities": {
|
|
125
130
|
"inputs": ["TASK_ID", "TASK_NAME", "TASK_INSTRUCTIONS", "DELIVERABLES_LIST", "ACCEPTANCE_CRITERIA"],
|
|
@@ -148,6 +153,9 @@
|
|
|
148
153
|
"status": "active",
|
|
149
154
|
"tier": 1,
|
|
150
155
|
"token_budget": 8000,
|
|
156
|
+
"protocol": "decomposition",
|
|
157
|
+
"loomStage": "decomposition",
|
|
158
|
+
"adrRefs": ["ADR-066", "ADR-073"],
|
|
151
159
|
"references": ["skills/ct-epic-architect/references/bug-epic-example.md", "skills/ct-epic-architect/references/commands.md", "skills/ct-epic-architect/references/feature-epic-example.md"],
|
|
152
160
|
"capabilities": {
|
|
153
161
|
"inputs": ["TASK_ID", "FEATURE_NAME", "EPIC_ID", "SESSION_ID"],
|
|
@@ -176,6 +184,9 @@
|
|
|
176
184
|
"status": "active",
|
|
177
185
|
"tier": 1,
|
|
178
186
|
"token_budget": 8000,
|
|
187
|
+
"protocol": "research",
|
|
188
|
+
"loomStage": "research",
|
|
189
|
+
"adrRefs": ["ADR-023", "ADR-070"],
|
|
179
190
|
"references": [],
|
|
180
191
|
"capabilities": {
|
|
181
192
|
"inputs": ["TASK_ID", "TOPIC", "RESEARCH_QUESTIONS"],
|
|
@@ -204,6 +215,9 @@
|
|
|
204
215
|
"status": "active",
|
|
205
216
|
"tier": 1,
|
|
206
217
|
"token_budget": 8000,
|
|
218
|
+
"protocol": "specification",
|
|
219
|
+
"loomStage": "specification",
|
|
220
|
+
"adrRefs": ["ADR-014", "ADR-023"],
|
|
207
221
|
"references": [],
|
|
208
222
|
"capabilities": {
|
|
209
223
|
"inputs": ["TASK_ID", "SPEC_NAME", "spec_topic"],
|
|
@@ -232,6 +246,9 @@
|
|
|
232
246
|
"status": "active",
|
|
233
247
|
"tier": 1,
|
|
234
248
|
"token_budget": 6000,
|
|
249
|
+
"protocol": "validation",
|
|
250
|
+
"loomStage": "validation",
|
|
251
|
+
"adrRefs": ["ADR-051", "ADR-023"],
|
|
235
252
|
"references": [],
|
|
236
253
|
"capabilities": {
|
|
237
254
|
"inputs": ["TASK_ID", "VALIDATION_TARGET", "VALIDATION_CRITERIA"],
|
|
@@ -239,7 +256,7 @@
|
|
|
239
256
|
"dependencies": [],
|
|
240
257
|
"dispatch_triggers": ["validate", "verify", "check compliance", "audit"],
|
|
241
258
|
"compatible_subagent_types": ["general-purpose"],
|
|
242
|
-
"chains_to": [],
|
|
259
|
+
"chains_to": ["ct-ivt-looper"],
|
|
243
260
|
"dispatch_keywords": {
|
|
244
261
|
"primary": ["validate", "verify", "audit", "compliance"],
|
|
245
262
|
"secondary": ["check", "conformance", "standards", "requirements"]
|
|
@@ -400,6 +417,9 @@
|
|
|
400
417
|
"status": "active",
|
|
401
418
|
"tier": 2,
|
|
402
419
|
"token_budget": 6000,
|
|
420
|
+
"protocol": "contribution",
|
|
421
|
+
"loomStage": "contribution",
|
|
422
|
+
"adrRefs": ["ADR-015", "ADR-053"],
|
|
403
423
|
"references": [],
|
|
404
424
|
"capabilities": {
|
|
405
425
|
"inputs": ["TASK_ID", "contribution_type", "context"],
|
|
@@ -484,7 +504,9 @@
|
|
|
484
504
|
"status": "active",
|
|
485
505
|
"tier": 2,
|
|
486
506
|
"token_budget": 8000,
|
|
487
|
-
"protocol": "
|
|
507
|
+
"protocol": "architecture_decision",
|
|
508
|
+
"loomStage": "architecture_decision",
|
|
509
|
+
"adrRefs": ["ADR-053", "ADR-070"],
|
|
488
510
|
"references": [
|
|
489
511
|
"skills/ct-adr-recorder/references/cascade.md",
|
|
490
512
|
"skills/ct-adr-recorder/references/examples.md"
|
|
@@ -523,6 +545,8 @@
|
|
|
523
545
|
"tier": 2,
|
|
524
546
|
"token_budget": 10000,
|
|
525
547
|
"protocol": "testing",
|
|
548
|
+
"loomStage": "testing",
|
|
549
|
+
"adrRefs": ["ADR-051", "ADR-061"],
|
|
526
550
|
"references": [
|
|
527
551
|
"skills/ct-ivt-looper/references/escalation.md",
|
|
528
552
|
"skills/ct-ivt-looper/references/frameworks.md",
|
|
@@ -562,6 +586,8 @@
|
|
|
562
586
|
"tier": 2,
|
|
563
587
|
"token_budget": 6000,
|
|
564
588
|
"protocol": "consensus",
|
|
589
|
+
"loomStage": "consensus",
|
|
590
|
+
"adrRefs": ["ADR-015", "ADR-023"],
|
|
565
591
|
"references": ["skills/ct-consensus-voter/references/matrix-examples.md"],
|
|
566
592
|
"capabilities": {
|
|
567
593
|
"inputs": ["task-id", "question", "candidate-options"],
|
|
@@ -597,6 +623,8 @@
|
|
|
597
623
|
"tier": 2,
|
|
598
624
|
"token_budget": 8000,
|
|
599
625
|
"protocol": "release",
|
|
626
|
+
"loomStage": "release",
|
|
627
|
+
"adrRefs": ["ADR-053", "ADR-063", "ADR-065"],
|
|
600
628
|
"references": [
|
|
601
629
|
"skills/ct-release-orchestrator/references/composition.md",
|
|
602
630
|
"skills/ct-release-orchestrator/references/release-types.md"
|
|
@@ -700,6 +728,50 @@
|
|
|
700
728
|
"requires_session": false,
|
|
701
729
|
"requires_epic": false
|
|
702
730
|
}
|
|
731
|
+
},
|
|
732
|
+
{
|
|
733
|
+
"name": "ct-council",
|
|
734
|
+
"version": "1.0.0",
|
|
735
|
+
"description": "Convene \"The Council\" — a 5-advisor, shuffled gate-based peer-review, chairman-synthesis workflow for reviewing a plan, decision, architecture, or piece of work inside the current project. Operates on the current codebase — each advisor grounds their analysis in actual files/commits before opining. Output is validated by scripts/validate.py.",
|
|
736
|
+
"path": "skills/ct-council",
|
|
737
|
+
"tags": ["council", "review", "peer-review", "multi-perspective", "stress-test"],
|
|
738
|
+
"status": "active",
|
|
739
|
+
"tier": 2,
|
|
740
|
+
"token_budget": 8000,
|
|
741
|
+
"references": [
|
|
742
|
+
"skills/ct-council/references/chairman.md",
|
|
743
|
+
"skills/ct-council/references/contrarian.md",
|
|
744
|
+
"skills/ct-council/references/evidence-pack.md",
|
|
745
|
+
"skills/ct-council/references/examples.md",
|
|
746
|
+
"skills/ct-council/references/executor.md",
|
|
747
|
+
"skills/ct-council/references/expansionist.md",
|
|
748
|
+
"skills/ct-council/references/first-principles.md",
|
|
749
|
+
"skills/ct-council/references/outsider.md",
|
|
750
|
+
"skills/ct-council/references/peer-review.md"
|
|
751
|
+
],
|
|
752
|
+
"capabilities": {
|
|
753
|
+
"inputs": ["proposal", "plan", "architecture-decision", "task-id"],
|
|
754
|
+
"outputs": ["council-verdict", "advisor-reports", "convergence-report"],
|
|
755
|
+
"dependencies": [],
|
|
756
|
+
"dispatch_triggers": [
|
|
757
|
+
"convene the council",
|
|
758
|
+
"council review",
|
|
759
|
+
"run the five advisors",
|
|
760
|
+
"stress-test this",
|
|
761
|
+
"get multiple perspectives"
|
|
762
|
+
],
|
|
763
|
+
"compatible_subagent_types": ["general-purpose"],
|
|
764
|
+
"chains_to": [],
|
|
765
|
+
"dispatch_keywords": {
|
|
766
|
+
"primary": ["council", "advisors", "stress-test", "peer-review"],
|
|
767
|
+
"secondary": ["contrarian", "first-principles", "expansionist", "outsider", "executor", "chairman"]
|
|
768
|
+
}
|
|
769
|
+
},
|
|
770
|
+
"constraints": {
|
|
771
|
+
"max_context_tokens": 80000,
|
|
772
|
+
"requires_session": false,
|
|
773
|
+
"requires_epic": false
|
|
774
|
+
}
|
|
703
775
|
}
|
|
704
776
|
]
|
|
705
777
|
}
|
|
@@ -1,28 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
# MIGRATION NOTICE — ct-grade-v2-1
|
|
3
|
-
|
|
4
|
-
This directory is a **decommissioned staging copy** of `ct-grade` v2.1.
|
|
5
|
-
|
|
6
|
-
## Status
|
|
7
|
-
|
|
8
|
-
**Superseded.** All content has been merged into `packages/skills/skills/ct-grade/`.
|
|
9
|
-
|
|
10
|
-
## What Changed (T429 skill dedupe)
|
|
11
|
-
|
|
12
|
-
- `ct-grade-v2-1/SKILL.md` description, `argument-hint`, `allowed-tools`, and version
|
|
13
|
-
(2.1.0) were promoted into `ct-grade/SKILL.md`.
|
|
14
|
-
- `ct-grade-v2-1/manifest-entry.json` remains here as an archived snapshot; the
|
|
15
|
-
canonical manifest entry lives in `packages/skills/skills/manifest.json` under name
|
|
16
|
-
`ct-grade`.
|
|
17
|
-
- The `grade-viewer/` tooling in this directory was already reachable from `ct-grade/`
|
|
18
|
-
via its `agents/` and `evals/` directories. No content was lost.
|
|
19
|
-
|
|
20
|
-
## Migration Date
|
|
21
|
-
|
|
22
|
-
2026-04-08 — T429 hygiene wave, epic T382.
|
|
23
|
-
|
|
24
|
-
## Action Required
|
|
25
|
-
|
|
26
|
-
None. Do NOT load this skill. Use `ct-grade` instead.
|
|
27
|
-
If you need the A/B or blind-compare modes documented here, they are now part of
|
|
28
|
-
`ct-grade`'s SKILL.md description and invocation modes.
|
|
@@ -1,235 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: ct-grade
|
|
3
|
-
description: >-
|
|
4
|
-
CLEO session grading and A/B behavioral analysis with token tracking. Evaluates agent
|
|
5
|
-
session quality via a 5-dimension rubric (S1 session discipline, S2 discovery efficiency,
|
|
6
|
-
S3 task hygiene, S4 error protocol, S5 progressive disclosure). Supports three modes:
|
|
7
|
-
(1) scenario — run playbook scenarios S1-S5 via CLI; (2) ab — blind A/B
|
|
8
|
-
comparison of different CLI configurations for same domain operations with token cost
|
|
9
|
-
measurement; (3) blind — spawn two agents with different configurations, blind-comparator
|
|
10
|
-
picks winner, analyzer produces recommendation. Use when grading agent sessions, running
|
|
11
|
-
grade playbook scenarios, comparing behavioral differences, measuring token
|
|
12
|
-
usage across configurations, or performing multi-run blind A/B evaluation with statistical
|
|
13
|
-
analysis and comparative report. Triggers on: grade session, evaluate agent behavior,
|
|
14
|
-
A/B test CLEO configurations, run grade scenario, token usage analysis, behavioral rubric,
|
|
15
|
-
protocol compliance scoring.
|
|
16
|
-
argument-hint: "[mode=scenario|ab|blind] [scenario=s1-s5|all] [runs=N] [session-id=<id>]"
|
|
17
|
-
allowed-tools: ["Bash(python *)", "Bash(cleo-dev *)", "Bash(cleo *)", "Bash(kill *)", "Bash(lsof *)", "Agent", "Read", "Write", "Glob"]
|
|
18
|
-
---
|
|
19
|
-
|
|
20
|
-
# ct-grade v2.1 — CLEO Grading and A/B Testing
|
|
21
|
-
|
|
22
|
-
Session grading and A/B behavioral analysis for CLEO protocol compliance. Three operating modes cover everything from single-session scoring to multi-run blind comparisons between different CLI configurations.
|
|
23
|
-
|
|
24
|
-
## On Every /ct-grade Invocation
|
|
25
|
-
|
|
26
|
-
Before parsing arguments, start the grade viewer server:
|
|
27
|
-
|
|
28
|
-
```bash
|
|
29
|
-
# Kill any existing viewer on port 3119
|
|
30
|
-
lsof -ti :3119 | xargs kill -TERM 2>/dev/null || true
|
|
31
|
-
|
|
32
|
-
# Start grade viewer in background
|
|
33
|
-
python $CLAUDE_SKILL_DIR/grade-viewer/generate_grade_review.py . \
|
|
34
|
-
--port 3119 --no-browser &
|
|
35
|
-
echo "Grade viewer: http://localhost:3119"
|
|
36
|
-
```
|
|
37
|
-
|
|
38
|
-
When user says "end grading", "stop", "done", or "close viewer":
|
|
39
|
-
```bash
|
|
40
|
-
lsof -ti :3119 | xargs kill -TERM 2>/dev/null || true
|
|
41
|
-
echo "Grade viewer stopped."
|
|
42
|
-
```
|
|
43
|
-
|
|
44
|
-
---
|
|
45
|
-
|
|
46
|
-
## Operating Modes
|
|
47
|
-
|
|
48
|
-
| Mode | Purpose | Key Output |
|
|
49
|
-
|---|---|---|
|
|
50
|
-
| `scenario` | Run playbook scenarios S1-S5 as graded sessions | GradeResult per scenario |
|
|
51
|
-
| `ab` | Run same domain operations with two configurations, compare | comparison.json + token delta |
|
|
52
|
-
| `blind` | Two agents run same task, blind comparator picks winner | analysis.json + winner |
|
|
53
|
-
|
|
54
|
-
## Parameters
|
|
55
|
-
|
|
56
|
-
| Parameter | Values | Default | Description |
|
|
57
|
-
|---|---|---|---|
|
|
58
|
-
| `mode` | `scenario\|ab\|blind` | `scenario` | Operating mode |
|
|
59
|
-
| `scenario` | `s1\|s2\|s3\|s4\|s5\|all` | `all` | Grade playbook scenario(s) to run |
|
|
60
|
-
| `interface` | `cli` | `cli` | Interface to exercise (CLI only) |
|
|
61
|
-
| `domains` | comma list | `tasks,session` | Domains to test in `ab` mode |
|
|
62
|
-
| `runs` | integer | `3` | Runs per configuration for statistical confidence |
|
|
63
|
-
| `session-id` | string | — | Grade a specific existing session (skips execution) |
|
|
64
|
-
| `output-dir` | path | `ab_results/<ts>` | Where to write all run artifacts |
|
|
65
|
-
|
|
66
|
-
## Quick Start
|
|
67
|
-
|
|
68
|
-
**Grade an existing session:**
|
|
69
|
-
```
|
|
70
|
-
/ct-grade session-id=<id>
|
|
71
|
-
```
|
|
72
|
-
|
|
73
|
-
**Run scenario S4 (Full Lifecycle):**
|
|
74
|
-
```
|
|
75
|
-
/ct-grade mode=scenario scenario=s4
|
|
76
|
-
```
|
|
77
|
-
|
|
78
|
-
**A/B compare two configurations for tasks + session domains (3 runs each):**
|
|
79
|
-
```
|
|
80
|
-
/ct-grade mode=ab domains=tasks,session runs=3
|
|
81
|
-
```
|
|
82
|
-
|
|
83
|
-
**Full blind A/B test across all scenarios:**
|
|
84
|
-
```
|
|
85
|
-
/ct-grade mode=blind scenario=all runs=3
|
|
86
|
-
```
|
|
87
|
-
|
|
88
|
-
---
|
|
89
|
-
|
|
90
|
-
## Execution Flow
|
|
91
|
-
|
|
92
|
-
### Mode: scenario
|
|
93
|
-
|
|
94
|
-
1. Set up output dir with `python $CLAUDE_SKILL_DIR/scripts/setup_run.py --mode scenario --scenario <id> --output-dir <dir>`
|
|
95
|
-
2. For each scenario, spawn a `scenario-runner` agent:
|
|
96
|
-
- Agent start: `cleo session start --scope global --name "<scenario-id>" --grade`
|
|
97
|
-
- Agent executes the scenario operations (see [references/playbook-v2.md](references/playbook-v2.md))
|
|
98
|
-
- Agent end: `cleo session end`
|
|
99
|
-
- Agent runs: `ct grade <sessionId>`
|
|
100
|
-
- Agent saves: `GradeResult` to `<output-dir>/<scenario>/grade.json`
|
|
101
|
-
3. Capture `total_tokens` + `duration_ms` from task notification → `timing.json`
|
|
102
|
-
4. Run: `python $CLAUDE_SKILL_DIR/scripts/generate_report.py --run-dir <dir> --mode scenario`
|
|
103
|
-
|
|
104
|
-
### Mode: ab
|
|
105
|
-
|
|
106
|
-
1. Set up run dir with `python $CLAUDE_SKILL_DIR/scripts/setup_run.py --mode ab --output-dir <dir>`
|
|
107
|
-
2. For each target domain, spawn TWO agents in the SAME turn:
|
|
108
|
-
- **Arm A**: `agents/scenario-runner.md` with configuration A
|
|
109
|
-
- **Arm B**: `agents/scenario-runner.md` with configuration B
|
|
110
|
-
- Capture tokens from both task notifications immediately
|
|
111
|
-
3. Pass both outputs to `agents/blind-comparator.md` (does NOT know which configuration is which)
|
|
112
|
-
4. Comparator writes `comparison.json`
|
|
113
|
-
5. Run `python $CLAUDE_SKILL_DIR/scripts/generate_report.py --run-dir <dir> --mode ab`
|
|
114
|
-
|
|
115
|
-
### Mode: blind
|
|
116
|
-
|
|
117
|
-
Same as `ab` but configurations may differ (e.g., different session scopes, different agent prompts). The comparator is always blind to configuration identity.
|
|
118
|
-
|
|
119
|
-
---
|
|
120
|
-
|
|
121
|
-
## Token Capture — MANDATORY
|
|
122
|
-
|
|
123
|
-
After EVERY Agent task notification, immediately update `timing.json`:
|
|
124
|
-
|
|
125
|
-
```python
|
|
126
|
-
timing = {
|
|
127
|
-
"total_tokens": task.total_tokens, # from task notification — EPHEMERAL
|
|
128
|
-
"duration_ms": task.duration_ms, # from task notification
|
|
129
|
-
"arm": "arm-A",
|
|
130
|
-
"interface": "cli",
|
|
131
|
-
"scenario": "s4",
|
|
132
|
-
"run": 1,
|
|
133
|
-
"executor_start": start_iso,
|
|
134
|
-
"executor_end": end_iso,
|
|
135
|
-
}
|
|
136
|
-
# Write to: <output-dir>/<scenario>/arm-<interface>/timing.json
|
|
137
|
-
```
|
|
138
|
-
|
|
139
|
-
**`total_tokens` is EPHEMERAL** — it cannot be recovered if missed. Capture it immediately.
|
|
140
|
-
|
|
141
|
-
If running without task notifications (no total_tokens available):
|
|
142
|
-
- Fall back: `output_chars / 3.5` from operations.jsonl (JSON responses)
|
|
143
|
-
- Record `"method": "output_chars_estimate"` in timing.json
|
|
144
|
-
|
|
145
|
-
---
|
|
146
|
-
|
|
147
|
-
## Grade Rubric Summary
|
|
148
|
-
|
|
149
|
-
5 dimensions × 20 pts = 100 max. See [references/grade-spec-v2.md](references/grade-spec-v2.md) for full scoring logic.
|
|
150
|
-
|
|
151
|
-
| Dim | Points | What it measures |
|
|
152
|
-
|---|---|---|
|
|
153
|
-
| S1 Session Discipline | 20 | `session.list` before task ops (+10), `session.end` present (+10) |
|
|
154
|
-
| S2 Discovery Efficiency | 20 | `find:list` ratio ≥80% (+15), `tasks.show` used (+5) |
|
|
155
|
-
| S3 Task Hygiene | 20 | Starts 20, -5 per add without description, -3 if subtask no exists check |
|
|
156
|
-
| S4 Error Protocol | 20 | Starts 20, -5 per unrecovered E_NOT_FOUND, -5 if duplicates |
|
|
157
|
-
| S5 Progressive Disclosure | 20 | `admin.help`/skill lookup (+10), progressive disclosure used (+10) |
|
|
158
|
-
|
|
159
|
-
**Grade letters:** A>=90, B>=75, C>=60, D>=45, F<45
|
|
160
|
-
|
|
161
|
-
---
|
|
162
|
-
|
|
163
|
-
## Output Structure
|
|
164
|
-
|
|
165
|
-
```
|
|
166
|
-
<output-dir>/
|
|
167
|
-
run-manifest.json # run config, arms, timing summary
|
|
168
|
-
report.md # human-readable comparative report
|
|
169
|
-
token-summary.json # aggregated token stats across all runs
|
|
170
|
-
<scenario-or-domain>/
|
|
171
|
-
arm-A/
|
|
172
|
-
grade.json # GradeResult (from check.grade)
|
|
173
|
-
timing.json # token + duration data
|
|
174
|
-
operations.jsonl # operations executed (one per line)
|
|
175
|
-
arm-B/
|
|
176
|
-
grade.json
|
|
177
|
-
timing.json
|
|
178
|
-
operations.jsonl
|
|
179
|
-
comparison.json # blind comparator output
|
|
180
|
-
analysis.json # analyzer output
|
|
181
|
-
```
|
|
182
|
-
|
|
183
|
-
---
|
|
184
|
-
|
|
185
|
-
## Agents
|
|
186
|
-
|
|
187
|
-
| Agent | Role | Input | Output |
|
|
188
|
-
|---|---|---|---|
|
|
189
|
-
| [agents/scenario-runner.md](agents/scenario-runner.md) | Executes grade scenario | scenario, interface | grade.json, timing.json |
|
|
190
|
-
| [agents/blind-comparator.md](agents/blind-comparator.md) | Blind A/B judge | outputs A and B | comparison.json |
|
|
191
|
-
| [agents/analysis-reporter.md](agents/analysis-reporter.md) | Post-hoc synthesis | all comparison.json | analysis.json |
|
|
192
|
-
|
|
193
|
-
---
|
|
194
|
-
|
|
195
|
-
## Scripts
|
|
196
|
-
|
|
197
|
-
```bash
|
|
198
|
-
# Set up run directory and print execution plan
|
|
199
|
-
python $CLAUDE_SKILL_DIR/scripts/setup_run.py --mode <mode> --scenario <s> --output-dir <dir>
|
|
200
|
-
|
|
201
|
-
# Aggregate token data after runs complete
|
|
202
|
-
python $CLAUDE_SKILL_DIR/scripts/token_tracker.py --run-dir <dir>
|
|
203
|
-
|
|
204
|
-
# Generate final report (markdown)
|
|
205
|
-
python $CLAUDE_SKILL_DIR/scripts/generate_report.py --run-dir <dir> --mode <mode>
|
|
206
|
-
```
|
|
207
|
-
|
|
208
|
-
---
|
|
209
|
-
|
|
210
|
-
## Viewers
|
|
211
|
-
|
|
212
|
-
### Grade Results Viewer (A/B run artifacts) — port 3119
|
|
213
|
-
```bash
|
|
214
|
-
python $CLAUDE_SKILL_DIR/grade-viewer/generate_grade_viewer.py --run-dir <ab-run-dir>
|
|
215
|
-
python $CLAUDE_SKILL_DIR/grade-viewer/generate_grade_viewer.py --run-dir <ab-run-dir> --static results.html
|
|
216
|
-
```
|
|
217
|
-
Shows per-scenario grade cards with dimension bars, A/B comparison tables, token economy stats, blind comparator results, and recommendations. Refreshes on browser reload.
|
|
218
|
-
|
|
219
|
-
### General Grade Review (GRADES.jsonl browsing) — port 3119
|
|
220
|
-
```bash
|
|
221
|
-
python $CLAUDE_SKILL_DIR/grade-viewer/generate_grade_review.py <workspace>
|
|
222
|
-
python $CLAUDE_SKILL_DIR/grade-viewer/generate_grade_review.py <workspace> --static grade-report.html
|
|
223
|
-
```
|
|
224
|
-
Shows historical grades from GRADES.jsonl, A/B summaries from any workspace subdirectory.
|
|
225
|
-
|
|
226
|
-
---
|
|
227
|
-
|
|
228
|
-
## CLI Grade Operations
|
|
229
|
-
|
|
230
|
-
| Command | Description |
|
|
231
|
-
|---------|-------------|
|
|
232
|
-
| `ct grade <sessionId>` | Grade a specific session |
|
|
233
|
-
| `ct grade --list` | List past grade results |
|
|
234
|
-
| `ct session start --scope global --name "<n>" --grade` | Start graded session |
|
|
235
|
-
| `ct session end` | End session |
|