@cleocode/skills 2026.5.82 → 2026.5.84

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (44) hide show
  1. package/README.md +0 -1
  2. package/package.json +1 -1
  3. package/profiles/recommended.json +1 -1
  4. package/skills/_shared/__tests__/lifecycle-protocol-reconcile.test.ts +112 -0
  5. package/skills/_shared/__tests__/loom-adr-links.test.ts +163 -0
  6. package/skills/_shared/__tests__/loom-stage-coverage.test.ts +167 -0
  7. package/skills/ct-adr-recorder/SKILL.md +18 -0
  8. package/skills/ct-consensus-voter/SKILL.md +14 -0
  9. package/skills/ct-contribution/SKILL.md +80 -0
  10. package/skills/ct-epic-architect/SKILL.md +15 -0
  11. package/skills/ct-ivt-looper/SKILL.md +32 -0
  12. package/skills/ct-release-orchestrator/SKILL.md +16 -0
  13. package/skills/ct-research-agent/SKILL.md +15 -0
  14. package/skills/ct-spec-writer/SKILL.md +15 -0
  15. package/skills/ct-task-executor/SKILL.md +15 -0
  16. package/skills/ct-validator/SKILL.md +35 -0
  17. package/skills/manifest.json +81 -9
  18. package/skills/ct-grade-v2-1/MIGRATION.md +0 -28
  19. package/skills/ct-grade-v2-1/SKILL.md +0 -235
  20. package/skills/ct-grade-v2-1/agents/analysis-reporter.md +0 -203
  21. package/skills/ct-grade-v2-1/agents/blind-comparator.md +0 -157
  22. package/skills/ct-grade-v2-1/agents/scenario-runner.md +0 -160
  23. package/skills/ct-grade-v2-1/evals/evals.json +0 -74
  24. package/skills/ct-grade-v2-1/grade-viewer/__pycache__/build_op_stats.cpython-314.pyc +0 -0
  25. package/skills/ct-grade-v2-1/grade-viewer/__pycache__/generate_grade_review.cpython-314.pyc +0 -0
  26. package/skills/ct-grade-v2-1/grade-viewer/build_op_stats.py +0 -174
  27. package/skills/ct-grade-v2-1/grade-viewer/eval-analysis.json +0 -41
  28. package/skills/ct-grade-v2-1/grade-viewer/eval-report.md +0 -37
  29. package/skills/ct-grade-v2-1/grade-viewer/generate_grade_review.py +0 -1023
  30. package/skills/ct-grade-v2-1/grade-viewer/generate_grade_viewer.py +0 -548
  31. package/skills/ct-grade-v2-1/grade-viewer/grade-review-eval.html +0 -613
  32. package/skills/ct-grade-v2-1/grade-viewer/grade-review.html +0 -1532
  33. package/skills/ct-grade-v2-1/grade-viewer/viewer.html +0 -620
  34. package/skills/ct-grade-v2-1/manifest-entry.json +0 -31
  35. package/skills/ct-grade-v2-1/references/ab-testing.md +0 -173
  36. package/skills/ct-grade-v2-1/references/domains-ssot.md +0 -156
  37. package/skills/ct-grade-v2-1/references/grade-spec-v2.md +0 -167
  38. package/skills/ct-grade-v2-1/references/playbook-v2.md +0 -325
  39. package/skills/ct-grade-v2-1/references/token-tracking.md +0 -200
  40. package/skills/ct-grade-v2-1/scripts/generate_report.py +0 -419
  41. package/skills/ct-grade-v2-1/scripts/run_ab_test.py +0 -493
  42. package/skills/ct-grade-v2-1/scripts/run_scenario.py +0 -396
  43. package/skills/ct-grade-v2-1/scripts/setup_run.py +0 -207
  44. package/skills/ct-grade-v2-1/scripts/token_tracker.py +0 -175
@@ -1,6 +1,11 @@
1
1
  ---
2
2
  name: ct-ivt-looper
3
3
  description: "Runs a project-agnostic autonomous Implement-then-Validate-then-Test compliance loop on any git worktree. Detects the project's test framework (vitest, jest, mocha, pytest, unittest, go-test, cargo-test, rspec, phpunit, bats, or other) and iterates until the implementation satisfies its specification, recording convergence metrics to the manifest. Use when given an implementation task that must ship verified: the IVT loop is the autonomous compliance layer enforced before any release or PR. Triggers on phrases like 'implement and verify', 'run the IVT loop', 'ship this task', 'complete implementation with tests', 'verify against spec', or any implementation task with acceptance criteria. Works in any git worktree regardless of language or framework, never hardcoded to one project's tooling."
4
+ protocol: testing
5
+ loomStage: testing
6
+ adrRefs:
7
+ - ADR-051
8
+ - ADR-061
4
9
  ---
5
10
 
6
11
  # IVT Looper
@@ -76,6 +81,24 @@ escalate_to_hitl() # IVT-007: exit code 65
76
81
 
77
82
  The loop is a *single* stage from the lifecycle's point of view. Implement, Validate, and Test are not three separate tasks — they are three phases of one autonomous run that either converges or escalates.
78
83
 
84
+ ## Out of Scope (T9675)
85
+
86
+ `ct-ivt-looper` operates on the **`testing`** LOOM lifecycle stage (stage 8). It performs the **dynamic** Implement-then-Validate-then-Test loop with framework detection and iterate-until-green convergence semantics.
87
+
88
+ This skill does NOT:
89
+
90
+ - Audit static artifacts for schema/compliance/RFC-2119 keyword usage, ADR-document structure, or JSON Schema conformance. Those belong to **`ct-validator`** at the `validation` stage (stage 7). When the question is "is this document/manifest well-formed?" rather than "does this code converge on its spec?", chain to `ct-validator` rather than expanding scope here.
91
+ - Promote a green loop to release. Release sequencing belongs to **`ct-release-orchestrator`** at stage 9.
92
+
93
+ ### Chain handoffs
94
+
95
+ | Direction | When | Handoff |
96
+ |---|---|---|
97
+ | `ct-ivt-looper` → `ct-validator` | Loop converged; need to audit the resulting artifacts (e.g. final manifest, spec back-references) against schema/compliance | Emit the convergence manifest entry, then dispatch the `validation` stage |
98
+ | `ct-validator` → `ct-ivt-looper` | Spec is valid but implementation needs dynamic verification | Receive a dispatch from the `validation` stage; iterate the IVT loop on the worktree |
99
+
100
+ Governance: see **ADR-051** (programmatic gate integrity) which defines the evidence atoms (`tool:test`, `test-run:<json>`) the loop emits and that downstream `cleo verify --gate testsPassed` re-validates, and **ADR-061** (project-agnostic verify tools) which defines the canonical tool-resolution layer.
101
+
79
102
  ## Framework Detection
80
103
 
81
104
  Framework detection is project-agnostic: the skill walks the worktree, inspects config files, and selects the correct test command. No language or framework is special-cased above another. The full detection table lives in [references/frameworks.md](references/frameworks.md). In summary: detection reads the project manifest (`package.json`, `pyproject.toml`, `Cargo.toml`, `go.mod`, `Gemfile`, `composer.json`, or `.cleo/project-context.json#testing.command`) and selects one of: `vitest`, `jest`, `mocha`, `pytest`, `unittest`, `go-test`, `cargo-test`, `rspec`, `phpunit`, `bats`, `other`.
@@ -179,3 +202,12 @@ cleo check protocol \
179
202
  6. Record `framework`, `testsRun`, `testsPassed`, `testsFailed`, `ivtLoopConverged`, `ivtLoopIterations` in the manifest.
180
203
  7. On non-convergence, exit 65 and leave the worktree untouched.
181
204
  8. Validate every run via `cleo check protocol --protocolType testing`.
205
+
206
+ ## See also / References
207
+
208
+ This skill binds to the **testing** LOOM lifecycle stage. Governing ADRs:
209
+
210
+ - [ADR-051 — programmatic gate integrity](../../../../.cleo/adrs/ADR-051-programmatic-gate-integrity.md) — defines the evidence atoms (`tool:test`, `test-run:<json>`) that the IVT loop emits and that downstream `cleo verify --gate testsPassed` re-validates.
211
+ - [ADR-061 — project-agnostic verify tools](../../../../.cleo/adrs/ADR-061-project-agnostic-verify-tools.md) — defines the canonical tool-resolution layer (`test`, `build`, `lint`, `typecheck`) that the loop walks for framework-agnostic execution.
212
+
213
+ LOOM coverage matrix: [docs/skills/loom-coverage-matrix.md](../../../../docs/skills/loom-coverage-matrix.md).
@@ -1,6 +1,12 @@
1
1
  ---
2
2
  name: ct-release-orchestrator
3
3
  description: "Orchestrates the full release pipeline: version bump, then changelog, then commit, then tag, then conditionally forks to artifact-publish and provenance based on release config. Parent protocol that composes ct-artifact-publisher and ct-provenance-keeper as sub-protocols: not every release publishes artifacts (source-only releases skip it), and artifact publishers delegate signing and attestation to provenance. Use when shipping a new version, running cleo release ship, or promoting a completed epic to released status."
4
+ protocol: release
5
+ loomStage: release
6
+ adrRefs:
7
+ - ADR-053
8
+ - ADR-063
9
+ - ADR-065
4
10
  ---
5
11
 
6
12
  # Release Orchestrator
@@ -132,3 +138,13 @@ For source-only releases, pass `--no-artifacts` to skip the artifact-publish han
132
138
  6. `released` entries are immutable; hotfixes go into new entries.
133
139
  7. Manifest entry MUST set `agent_type: "documentation"` and record the full chain via `record_release()`.
134
140
  8. Always validate via `cleo check protocol --protocolType release` before declaring the release done.
141
+
142
+ ## See also / References
143
+
144
+ This skill binds to the **release** LOOM lifecycle stage. Governing ADRs:
145
+
146
+ - [ADR-053 — project-agnostic release pipeline](../../../../.cleo/adrs/ADR-053-project-agnostic-release-pipeline.md) — defines the language-agnostic version bump → changelog → tag flow.
147
+ - [ADR-063 — release pipeline](../../../../.cleo/adrs/ADR-063-release-pipeline.md) — defines the 12-step `cleo release ship` integration with CI.
148
+ - [ADR-065 — PR-required release flow](../../../../.cleo/adrs/ADR-065-pr-required-release-flow.md) — defines the PR-gated path; direct pushes to `main` are prohibited.
149
+
150
+ LOOM coverage matrix: [docs/skills/loom-coverage-matrix.md](../../../../docs/skills/loom-coverage-matrix.md).
@@ -6,6 +6,10 @@ tier: 2
6
6
  core: false
7
7
  category: recommended
8
8
  protocol: research
9
+ loomStage: research
10
+ adrRefs:
11
+ - ADR-023
12
+ - ADR-070
9
13
  dependencies: []
10
14
  sharedResources:
11
15
  - subagent-protocol-base
@@ -224,3 +228,14 @@ If research cannot proceed (access denied, topic too broad, etc.):
224
228
  - **Prioritized** - Most important first
225
229
  - **Justified** - Tied to specific findings
226
230
  - **Feasible** - Achievable within project constraints
231
+
232
+ ---
233
+
234
+ ## See also / References
235
+
236
+ This skill binds to the **research** LOOM lifecycle stage. Governing ADRs:
237
+
238
+ - [ADR-023 — protocol validation dispatch](../../../../.cleo/adrs/ADR-023-protocol-validation-dispatch.md) — defines how research output is validated before downstream stages consume it.
239
+ - [ADR-070 — three-tier orchestration](../../../../.cleo/adrs/ADR-070-three-tier-orchestration.md) — defines the Orchestrator → Phase Lead → Worker tiers; research runs as a leaf Worker.
240
+
241
+ LOOM coverage matrix: [docs/skills/loom-coverage-matrix.md](../../../../docs/skills/loom-coverage-matrix.md).
@@ -6,6 +6,10 @@ tier: 2
6
6
  core: false
7
7
  category: recommended
8
8
  protocol: specification
9
+ loomStage: specification
10
+ adrRefs:
11
+ - ADR-014
12
+ - ADR-023
9
13
  dependencies: []
10
14
  sharedResources:
11
15
  - subagent-protocol-base
@@ -187,3 +191,14 @@ Specifications go in: `docs/specs/{{SPEC_NAME}}.md`
187
191
  - [ ] Manifest entry appended
188
192
  - [ ] Task completed via `{{TASK_COMPLETE_CMD}}`
189
193
  - [ ] Return summary message only
194
+
195
+ ---
196
+
197
+ ## See also / References
198
+
199
+ This skill binds to the **specification** LOOM lifecycle stage. Governing ADRs:
200
+
201
+ - [ADR-014 — RCASD rename and protocol validation](../../../../.cleo/adrs/ADR-014-rcasd-rename-and-protocol-validation.md) — defines the specification stage's role inside the RCASD-IVTR+C lifecycle.
202
+ - [ADR-023 — protocol validation dispatch](../../../../.cleo/adrs/ADR-023-protocol-validation-dispatch.md) — defines how specifications are validated before decomposition.
203
+
204
+ LOOM coverage matrix: [docs/skills/loom-coverage-matrix.md](../../../../docs/skills/loom-coverage-matrix.md).
@@ -6,6 +6,10 @@ tier: 2
6
6
  core: true
7
7
  category: core
8
8
  protocol: implementation
9
+ loomStage: implementation
10
+ adrRefs:
11
+ - ADR-070
12
+ - ADR-062
9
13
  dependencies: []
10
14
  sharedResources:
11
15
  - subagent-protocol-base
@@ -294,3 +298,14 @@ cleo session gc --include-active
294
298
  | Partial deliverables | Missing outputs | Complete all or report partial |
295
299
  | Undocumented changes | Lost context | Write detailed output file |
296
300
  | Silent failures | Orchestrator unaware | Report via manifest status |
301
+
302
+ ---
303
+
304
+ ## See also / References
305
+
306
+ This skill binds to the **implementation** LOOM lifecycle stage. Governing ADRs:
307
+
308
+ - [ADR-070 — three-tier orchestration](../../../../.cleo/adrs/ADR-070-three-tier-orchestration.md) — defines the Worker tier that ct-task-executor occupies.
309
+ - [ADR-062 — worktree merge, not cherry-pick](../../../../.cleo/adrs/ADR-062-worktree-merge-not-cherry-pick.md) — defines the integration path that preserves the executor's commit SHAs end-to-end.
310
+
311
+ LOOM coverage matrix: [docs/skills/loom-coverage-matrix.md](../../../../docs/skills/loom-coverage-matrix.md).
@@ -6,6 +6,10 @@ tier: 2
6
6
  core: false
7
7
  category: recommended
8
8
  protocol: validation
9
+ loomStage: validation
10
+ adrRefs:
11
+ - ADR-051
12
+ - ADR-023
9
13
  dependencies: []
10
14
  sharedResources:
11
15
  - subagent-protocol-base
@@ -41,6 +45,26 @@ Context injection for compliance validation tasks spawned via cleo-subagent. Pro
41
45
 
42
46
  ---
43
47
 
48
+ ## Out of Scope (T9675)
49
+
50
+ `ct-validator` operates on the **`validation`** LOOM lifecycle stage (stage 7). It performs **static** schema, compliance, and audit checks against artifacts that already exist on disk (specs, ADRs, JSON files, RFC 2119 keyword usage, manifest schemas).
51
+
52
+ This skill does NOT:
53
+
54
+ - Run a test suite, framework detection, or iterative IVT loop. Those belong to **`ct-ivt-looper`** at the `testing` stage (stage 8). When dynamic verification is required — e.g. "does the implementation actually pass its tests?" — chain to `ct-ivt-looper` rather than expanding scope here.
55
+ - Modify code or apply fixes. The validator reports; downstream skills remediate.
56
+
57
+ ### Chain handoffs
58
+
59
+ | Direction | When | Handoff |
60
+ |---|---|---|
61
+ | `ct-validator` → `ct-ivt-looper` | Spec is valid but implementation needs dynamic verification | Emit a manifest entry, then dispatch the `testing` stage |
62
+ | `ct-ivt-looper` → `ct-validator` | IVT loop converged; need to audit the resulting artifacts against schema/compliance | Dispatch the `validation` stage after the test convergence record |
63
+
64
+ Governance: see **ADR-051** (programmatic gate integrity) which defines the evidence atoms each stage emits and that the other stage may re-validate, and **ADR-023** (protocol validation dispatch) which routes between them.
65
+
66
+ ---
67
+
44
68
  ## Validation Methodology
45
69
 
46
70
  ### Standard Workflow
@@ -214,3 +238,14 @@ When invoked by orchestrator, expect these context tokens:
214
238
  | Vague findings | Unclear remediation | Specific issue + file/line + fix |
215
239
  | Missing severity | Can't prioritize | Always classify: critical/warning/suggestion |
216
240
  | No remediation | Findings not actionable | Always provide fix for FAIL/PARTIAL |
241
+
242
+ ---
243
+
244
+ ## See also / References
245
+
246
+ This skill binds to the **validation** LOOM lifecycle stage. Governing ADRs:
247
+
248
+ - [ADR-051 — programmatic gate integrity](../../../../.cleo/adrs/ADR-051-programmatic-gate-integrity.md) — defines the evidence-atom grammar that the validator emits and re-validates.
249
+ - [ADR-023 — protocol validation dispatch](../../../../.cleo/adrs/ADR-023-protocol-validation-dispatch.md) — defines the protocol-validation routing layer that dispatches to this skill.
250
+
251
+ LOOM coverage matrix: [docs/skills/loom-coverage-matrix.md](../../../../docs/skills/loom-coverage-matrix.md).
@@ -1,11 +1,13 @@
1
1
  {
2
2
  "$schema": "https://cleo-dev.com/schemas/v1/skills-manifest.schema.json",
3
3
  "_meta": {
4
- "schemaVersion": "2.4.0",
5
- "lastUpdated": "2026-04-07",
6
- "totalSkills": 21,
7
- "generatedFrom": "T260 — lifecycle pipeline rework: dedicated skills for ADR, IVT loop, consensus, release, artifact-publish, provenance",
8
- "architectureNote": "Universal Subagent Architecture: All spawns use provider-neutral delegation with skill/protocol injection. Pipeline stages and cross-cutting protocols each have a dedicated skill — no overloading."
4
+ "schemaVersion": "2.7.0",
5
+ "lastUpdated": "2026-05-19",
6
+ "totalSkills": 22,
7
+ "generatedFrom": "T260 — lifecycle pipeline rework: dedicated skills for ADR, IVT loop, consensus, release, artifact-publish, provenance. T9664/T9665/T9672 (epic T9568) — adds loomStage + adrRefs fields on every LOOM-stage skill entry (underscored canonical form matching `cleo lifecycle` stage names; ADR refs point at .cleo/adrs/). T9672 reconciled dispatch_matrix.by_protocol key 'architecture-decision' → 'architecture_decision' matching the lifecycle CLI source of truth; the dashed form is retained as a keyword alias.",
8
+ "architectureNote": "Universal Subagent Architecture: All spawns use provider-neutral delegation with skill/protocol injection. Pipeline stages and cross-cutting protocols each have a dedicated skill — no overloading.",
9
+ "loomStageContract": "Each of the 10 LOOM lifecycle stages (research, consensus, architecture_decision, specification, decomposition, implementation, validation, testing, release, contribution) MUST have a bound skill whose manifest entry declares loomStage matching the lifecycle CLI stage name. The legacy `protocol` field is retained as an alias; dispatch_matrix.by_protocol is the authoritative routing table. See docs/skills/loom-coverage-matrix.md.",
10
+ "adrRefsContract": "Each LOOM-stage skill MUST declare an adrRefs[] array naming the ADR file(s) that govern the stage. Every referenced ADR-NNN MUST resolve to a file under .cleo/adrs/. The Vitest gate at packages/skills/skills/_shared/__tests__/loom-adr-links.test.ts enforces both conditions."
9
11
  },
10
12
  "dispatch_matrix": {
11
13
  "_comment": "Maps task types/keywords to skill NAMES. Provider adapter decides HOW to execute.",
@@ -33,7 +35,7 @@
33
35
  "spec|rfc|protocol|contract": "ct-spec-writer",
34
36
  "validate|verify|audit|compliance": "ct-validator",
35
37
  "consensus|vote|verdict|resolve the debate": "ct-consensus-voter",
36
- "adr|architecture decision|formalize|lock in the choice": "ct-adr-recorder",
38
+ "adr|architecture decision|architecture-decision|formalize|lock in the choice": "ct-adr-recorder",
37
39
  "release|version|ship|changelog|cut release": "ct-release-orchestrator",
38
40
  "artifact|publish|registry|npm publish|docker push": "ct-artifact-publisher",
39
41
  "provenance|attestation|sbom|sigstore|slsa": "ct-provenance-keeper"
@@ -41,7 +43,7 @@
41
43
  "by_protocol": {
42
44
  "research": "ct-research-agent",
43
45
  "consensus": "ct-consensus-voter",
44
- "architecture-decision": "ct-adr-recorder",
46
+ "architecture_decision": "ct-adr-recorder",
45
47
  "specification": "ct-spec-writer",
46
48
  "decomposition": "ct-epic-architect",
47
49
  "implementation": "ct-task-executor",
@@ -120,6 +122,9 @@
120
122
  "status": "active",
121
123
  "tier": 0,
122
124
  "token_budget": 8000,
125
+ "protocol": "implementation",
126
+ "loomStage": "implementation",
127
+ "adrRefs": ["ADR-070", "ADR-062"],
123
128
  "references": [],
124
129
  "capabilities": {
125
130
  "inputs": ["TASK_ID", "TASK_NAME", "TASK_INSTRUCTIONS", "DELIVERABLES_LIST", "ACCEPTANCE_CRITERIA"],
@@ -148,6 +153,9 @@
148
153
  "status": "active",
149
154
  "tier": 1,
150
155
  "token_budget": 8000,
156
+ "protocol": "decomposition",
157
+ "loomStage": "decomposition",
158
+ "adrRefs": ["ADR-066", "ADR-073"],
151
159
  "references": ["skills/ct-epic-architect/references/bug-epic-example.md", "skills/ct-epic-architect/references/commands.md", "skills/ct-epic-architect/references/feature-epic-example.md"],
152
160
  "capabilities": {
153
161
  "inputs": ["TASK_ID", "FEATURE_NAME", "EPIC_ID", "SESSION_ID"],
@@ -176,6 +184,9 @@
176
184
  "status": "active",
177
185
  "tier": 1,
178
186
  "token_budget": 8000,
187
+ "protocol": "research",
188
+ "loomStage": "research",
189
+ "adrRefs": ["ADR-023", "ADR-070"],
179
190
  "references": [],
180
191
  "capabilities": {
181
192
  "inputs": ["TASK_ID", "TOPIC", "RESEARCH_QUESTIONS"],
@@ -204,6 +215,9 @@
204
215
  "status": "active",
205
216
  "tier": 1,
206
217
  "token_budget": 8000,
218
+ "protocol": "specification",
219
+ "loomStage": "specification",
220
+ "adrRefs": ["ADR-014", "ADR-023"],
207
221
  "references": [],
208
222
  "capabilities": {
209
223
  "inputs": ["TASK_ID", "SPEC_NAME", "spec_topic"],
@@ -232,6 +246,9 @@
232
246
  "status": "active",
233
247
  "tier": 1,
234
248
  "token_budget": 6000,
249
+ "protocol": "validation",
250
+ "loomStage": "validation",
251
+ "adrRefs": ["ADR-051", "ADR-023"],
235
252
  "references": [],
236
253
  "capabilities": {
237
254
  "inputs": ["TASK_ID", "VALIDATION_TARGET", "VALIDATION_CRITERIA"],
@@ -239,7 +256,7 @@
239
256
  "dependencies": [],
240
257
  "dispatch_triggers": ["validate", "verify", "check compliance", "audit"],
241
258
  "compatible_subagent_types": ["general-purpose"],
242
- "chains_to": [],
259
+ "chains_to": ["ct-ivt-looper"],
243
260
  "dispatch_keywords": {
244
261
  "primary": ["validate", "verify", "audit", "compliance"],
245
262
  "secondary": ["check", "conformance", "standards", "requirements"]
@@ -400,6 +417,9 @@
400
417
  "status": "active",
401
418
  "tier": 2,
402
419
  "token_budget": 6000,
420
+ "protocol": "contribution",
421
+ "loomStage": "contribution",
422
+ "adrRefs": ["ADR-015", "ADR-053"],
403
423
  "references": [],
404
424
  "capabilities": {
405
425
  "inputs": ["TASK_ID", "contribution_type", "context"],
@@ -484,7 +504,9 @@
484
504
  "status": "active",
485
505
  "tier": 2,
486
506
  "token_budget": 8000,
487
- "protocol": "architecture-decision",
507
+ "protocol": "architecture_decision",
508
+ "loomStage": "architecture_decision",
509
+ "adrRefs": ["ADR-053", "ADR-070"],
488
510
  "references": [
489
511
  "skills/ct-adr-recorder/references/cascade.md",
490
512
  "skills/ct-adr-recorder/references/examples.md"
@@ -523,6 +545,8 @@
523
545
  "tier": 2,
524
546
  "token_budget": 10000,
525
547
  "protocol": "testing",
548
+ "loomStage": "testing",
549
+ "adrRefs": ["ADR-051", "ADR-061"],
526
550
  "references": [
527
551
  "skills/ct-ivt-looper/references/escalation.md",
528
552
  "skills/ct-ivt-looper/references/frameworks.md",
@@ -562,6 +586,8 @@
562
586
  "tier": 2,
563
587
  "token_budget": 6000,
564
588
  "protocol": "consensus",
589
+ "loomStage": "consensus",
590
+ "adrRefs": ["ADR-015", "ADR-023"],
565
591
  "references": ["skills/ct-consensus-voter/references/matrix-examples.md"],
566
592
  "capabilities": {
567
593
  "inputs": ["task-id", "question", "candidate-options"],
@@ -597,6 +623,8 @@
597
623
  "tier": 2,
598
624
  "token_budget": 8000,
599
625
  "protocol": "release",
626
+ "loomStage": "release",
627
+ "adrRefs": ["ADR-053", "ADR-063", "ADR-065"],
600
628
  "references": [
601
629
  "skills/ct-release-orchestrator/references/composition.md",
602
630
  "skills/ct-release-orchestrator/references/release-types.md"
@@ -700,6 +728,50 @@
700
728
  "requires_session": false,
701
729
  "requires_epic": false
702
730
  }
731
+ },
732
+ {
733
+ "name": "ct-council",
734
+ "version": "1.0.0",
735
+ "description": "Convene \"The Council\" — a 5-advisor, shuffled gate-based peer-review, chairman-synthesis workflow for reviewing a plan, decision, architecture, or piece of work inside the current project. Operates on the current codebase — each advisor grounds their analysis in actual files/commits before opining. Output is validated by scripts/validate.py.",
736
+ "path": "skills/ct-council",
737
+ "tags": ["council", "review", "peer-review", "multi-perspective", "stress-test"],
738
+ "status": "active",
739
+ "tier": 2,
740
+ "token_budget": 8000,
741
+ "references": [
742
+ "skills/ct-council/references/chairman.md",
743
+ "skills/ct-council/references/contrarian.md",
744
+ "skills/ct-council/references/evidence-pack.md",
745
+ "skills/ct-council/references/examples.md",
746
+ "skills/ct-council/references/executor.md",
747
+ "skills/ct-council/references/expansionist.md",
748
+ "skills/ct-council/references/first-principles.md",
749
+ "skills/ct-council/references/outsider.md",
750
+ "skills/ct-council/references/peer-review.md"
751
+ ],
752
+ "capabilities": {
753
+ "inputs": ["proposal", "plan", "architecture-decision", "task-id"],
754
+ "outputs": ["council-verdict", "advisor-reports", "convergence-report"],
755
+ "dependencies": [],
756
+ "dispatch_triggers": [
757
+ "convene the council",
758
+ "council review",
759
+ "run the five advisors",
760
+ "stress-test this",
761
+ "get multiple perspectives"
762
+ ],
763
+ "compatible_subagent_types": ["general-purpose"],
764
+ "chains_to": [],
765
+ "dispatch_keywords": {
766
+ "primary": ["council", "advisors", "stress-test", "peer-review"],
767
+ "secondary": ["contrarian", "first-principles", "expansionist", "outsider", "executor", "chairman"]
768
+ }
769
+ },
770
+ "constraints": {
771
+ "max_context_tokens": 80000,
772
+ "requires_session": false,
773
+ "requires_epic": false
774
+ }
703
775
  }
704
776
  ]
705
777
  }
@@ -1,28 +0,0 @@
1
- ---
2
- # MIGRATION NOTICE — ct-grade-v2-1
3
-
4
- This directory is a **decommissioned staging copy** of `ct-grade` v2.1.
5
-
6
- ## Status
7
-
8
- **Superseded.** All content has been merged into `packages/skills/skills/ct-grade/`.
9
-
10
- ## What Changed (T429 skill dedupe)
11
-
12
- - `ct-grade-v2-1/SKILL.md` description, `argument-hint`, `allowed-tools`, and version
13
- (2.1.0) were promoted into `ct-grade/SKILL.md`.
14
- - `ct-grade-v2-1/manifest-entry.json` remains here as an archived snapshot; the
15
- canonical manifest entry lives in `packages/skills/skills/manifest.json` under name
16
- `ct-grade`.
17
- - The `grade-viewer/` tooling in this directory was already reachable from `ct-grade/`
18
- via its `agents/` and `evals/` directories. No content was lost.
19
-
20
- ## Migration Date
21
-
22
- 2026-04-08 — T429 hygiene wave, epic T382.
23
-
24
- ## Action Required
25
-
26
- None. Do NOT load this skill. Use `ct-grade` instead.
27
- If you need the A/B or blind-compare modes documented here, they are now part of
28
- `ct-grade`'s SKILL.md description and invocation modes.
@@ -1,235 +0,0 @@
1
- ---
2
- name: ct-grade
3
- description: >-
4
- CLEO session grading and A/B behavioral analysis with token tracking. Evaluates agent
5
- session quality via a 5-dimension rubric (S1 session discipline, S2 discovery efficiency,
6
- S3 task hygiene, S4 error protocol, S5 progressive disclosure). Supports three modes:
7
- (1) scenario — run playbook scenarios S1-S5 via CLI; (2) ab — blind A/B
8
- comparison of different CLI configurations for same domain operations with token cost
9
- measurement; (3) blind — spawn two agents with different configurations, blind-comparator
10
- picks winner, analyzer produces recommendation. Use when grading agent sessions, running
11
- grade playbook scenarios, comparing behavioral differences, measuring token
12
- usage across configurations, or performing multi-run blind A/B evaluation with statistical
13
- analysis and comparative report. Triggers on: grade session, evaluate agent behavior,
14
- A/B test CLEO configurations, run grade scenario, token usage analysis, behavioral rubric,
15
- protocol compliance scoring.
16
- argument-hint: "[mode=scenario|ab|blind] [scenario=s1-s5|all] [runs=N] [session-id=<id>]"
17
- allowed-tools: ["Bash(python *)", "Bash(cleo-dev *)", "Bash(cleo *)", "Bash(kill *)", "Bash(lsof *)", "Agent", "Read", "Write", "Glob"]
18
- ---
19
-
20
- # ct-grade v2.1 — CLEO Grading and A/B Testing
21
-
22
- Session grading and A/B behavioral analysis for CLEO protocol compliance. Three operating modes cover everything from single-session scoring to multi-run blind comparisons between different CLI configurations.
23
-
24
- ## On Every /ct-grade Invocation
25
-
26
- Before parsing arguments, start the grade viewer server:
27
-
28
- ```bash
29
- # Kill any existing viewer on port 3119
30
- lsof -ti :3119 | xargs kill -TERM 2>/dev/null || true
31
-
32
- # Start grade viewer in background
33
- python $CLAUDE_SKILL_DIR/grade-viewer/generate_grade_review.py . \
34
- --port 3119 --no-browser &
35
- echo "Grade viewer: http://localhost:3119"
36
- ```
37
-
38
- When user says "end grading", "stop", "done", or "close viewer":
39
- ```bash
40
- lsof -ti :3119 | xargs kill -TERM 2>/dev/null || true
41
- echo "Grade viewer stopped."
42
- ```
43
-
44
- ---
45
-
46
- ## Operating Modes
47
-
48
- | Mode | Purpose | Key Output |
49
- |---|---|---|
50
- | `scenario` | Run playbook scenarios S1-S5 as graded sessions | GradeResult per scenario |
51
- | `ab` | Run same domain operations with two configurations, compare | comparison.json + token delta |
52
- | `blind` | Two agents run same task, blind comparator picks winner | analysis.json + winner |
53
-
54
- ## Parameters
55
-
56
- | Parameter | Values | Default | Description |
57
- |---|---|---|---|
58
- | `mode` | `scenario\|ab\|blind` | `scenario` | Operating mode |
59
- | `scenario` | `s1\|s2\|s3\|s4\|s5\|all` | `all` | Grade playbook scenario(s) to run |
60
- | `interface` | `cli` | `cli` | Interface to exercise (CLI only) |
61
- | `domains` | comma list | `tasks,session` | Domains to test in `ab` mode |
62
- | `runs` | integer | `3` | Runs per configuration for statistical confidence |
63
- | `session-id` | string | — | Grade a specific existing session (skips execution) |
64
- | `output-dir` | path | `ab_results/<ts>` | Where to write all run artifacts |
65
-
66
- ## Quick Start
67
-
68
- **Grade an existing session:**
69
- ```
70
- /ct-grade session-id=<id>
71
- ```
72
-
73
- **Run scenario S4 (Full Lifecycle):**
74
- ```
75
- /ct-grade mode=scenario scenario=s4
76
- ```
77
-
78
- **A/B compare two configurations for tasks + session domains (3 runs each):**
79
- ```
80
- /ct-grade mode=ab domains=tasks,session runs=3
81
- ```
82
-
83
- **Full blind A/B test across all scenarios:**
84
- ```
85
- /ct-grade mode=blind scenario=all runs=3
86
- ```
87
-
88
- ---
89
-
90
- ## Execution Flow
91
-
92
- ### Mode: scenario
93
-
94
- 1. Set up output dir with `python $CLAUDE_SKILL_DIR/scripts/setup_run.py --mode scenario --scenario <id> --output-dir <dir>`
95
- 2. For each scenario, spawn a `scenario-runner` agent:
96
- - Agent start: `cleo session start --scope global --name "<scenario-id>" --grade`
97
- - Agent executes the scenario operations (see [references/playbook-v2.md](references/playbook-v2.md))
98
- - Agent end: `cleo session end`
99
- - Agent runs: `ct grade <sessionId>`
100
- - Agent saves: `GradeResult` to `<output-dir>/<scenario>/grade.json`
101
- 3. Capture `total_tokens` + `duration_ms` from task notification → `timing.json`
102
- 4. Run: `python $CLAUDE_SKILL_DIR/scripts/generate_report.py --run-dir <dir> --mode scenario`
103
-
104
- ### Mode: ab
105
-
106
- 1. Set up run dir with `python $CLAUDE_SKILL_DIR/scripts/setup_run.py --mode ab --output-dir <dir>`
107
- 2. For each target domain, spawn TWO agents in the SAME turn:
108
- - **Arm A**: `agents/scenario-runner.md` with configuration A
109
- - **Arm B**: `agents/scenario-runner.md` with configuration B
110
- - Capture tokens from both task notifications immediately
111
- 3. Pass both outputs to `agents/blind-comparator.md` (does NOT know which configuration is which)
112
- 4. Comparator writes `comparison.json`
113
- 5. Run `python $CLAUDE_SKILL_DIR/scripts/generate_report.py --run-dir <dir> --mode ab`
114
-
115
- ### Mode: blind
116
-
117
- Same as `ab` but configurations may differ (e.g., different session scopes, different agent prompts). The comparator is always blind to configuration identity.
118
-
119
- ---
120
-
121
- ## Token Capture — MANDATORY
122
-
123
- After EVERY Agent task notification, immediately update `timing.json`:
124
-
125
- ```python
126
- timing = {
127
- "total_tokens": task.total_tokens, # from task notification — EPHEMERAL
128
- "duration_ms": task.duration_ms, # from task notification
129
- "arm": "arm-A",
130
- "interface": "cli",
131
- "scenario": "s4",
132
- "run": 1,
133
- "executor_start": start_iso,
134
- "executor_end": end_iso,
135
- }
136
- # Write to: <output-dir>/<scenario>/arm-<interface>/timing.json
137
- ```
138
-
139
- **`total_tokens` is EPHEMERAL** — it cannot be recovered if missed. Capture it immediately.
140
-
141
- If running without task notifications (no total_tokens available):
142
- - Fall back: `output_chars / 3.5` from operations.jsonl (JSON responses)
143
- - Record `"method": "output_chars_estimate"` in timing.json
144
-
145
- ---
146
-
147
- ## Grade Rubric Summary
148
-
149
- 5 dimensions × 20 pts = 100 max. See [references/grade-spec-v2.md](references/grade-spec-v2.md) for full scoring logic.
150
-
151
- | Dim | Points | What it measures |
152
- |---|---|---|
153
- | S1 Session Discipline | 20 | `session.list` before task ops (+10), `session.end` present (+10) |
154
- | S2 Discovery Efficiency | 20 | `find:list` ratio ≥80% (+15), `tasks.show` used (+5) |
155
- | S3 Task Hygiene | 20 | Starts 20, -5 per add without description, -3 if subtask no exists check |
156
- | S4 Error Protocol | 20 | Starts 20, -5 per unrecovered E_NOT_FOUND, -5 if duplicates |
157
- | S5 Progressive Disclosure | 20 | `admin.help`/skill lookup (+10), progressive disclosure used (+10) |
158
-
159
- **Grade letters:** A>=90, B>=75, C>=60, D>=45, F<45
160
-
161
- ---
162
-
163
- ## Output Structure
164
-
165
- ```
166
- <output-dir>/
167
- run-manifest.json # run config, arms, timing summary
168
- report.md # human-readable comparative report
169
- token-summary.json # aggregated token stats across all runs
170
- <scenario-or-domain>/
171
- arm-A/
172
- grade.json # GradeResult (from check.grade)
173
- timing.json # token + duration data
174
- operations.jsonl # operations executed (one per line)
175
- arm-B/
176
- grade.json
177
- timing.json
178
- operations.jsonl
179
- comparison.json # blind comparator output
180
- analysis.json # analyzer output
181
- ```
182
-
183
- ---
184
-
185
- ## Agents
186
-
187
- | Agent | Role | Input | Output |
188
- |---|---|---|---|
189
- | [agents/scenario-runner.md](agents/scenario-runner.md) | Executes grade scenario | scenario, interface | grade.json, timing.json |
190
- | [agents/blind-comparator.md](agents/blind-comparator.md) | Blind A/B judge | outputs A and B | comparison.json |
191
- | [agents/analysis-reporter.md](agents/analysis-reporter.md) | Post-hoc synthesis | all comparison.json | analysis.json |
192
-
193
- ---
194
-
195
- ## Scripts
196
-
197
- ```bash
198
- # Set up run directory and print execution plan
199
- python $CLAUDE_SKILL_DIR/scripts/setup_run.py --mode <mode> --scenario <s> --output-dir <dir>
200
-
201
- # Aggregate token data after runs complete
202
- python $CLAUDE_SKILL_DIR/scripts/token_tracker.py --run-dir <dir>
203
-
204
- # Generate final report (markdown)
205
- python $CLAUDE_SKILL_DIR/scripts/generate_report.py --run-dir <dir> --mode <mode>
206
- ```
207
-
208
- ---
209
-
210
- ## Viewers
211
-
212
- ### Grade Results Viewer (A/B run artifacts) — port 3119
213
- ```bash
214
- python $CLAUDE_SKILL_DIR/grade-viewer/generate_grade_viewer.py --run-dir <ab-run-dir>
215
- python $CLAUDE_SKILL_DIR/grade-viewer/generate_grade_viewer.py --run-dir <ab-run-dir> --static results.html
216
- ```
217
- Shows per-scenario grade cards with dimension bars, A/B comparison tables, token economy stats, blind comparator results, and recommendations. Refreshes on browser reload.
218
-
219
- ### General Grade Review (GRADES.jsonl browsing) — port 3119
220
- ```bash
221
- python $CLAUDE_SKILL_DIR/grade-viewer/generate_grade_review.py <workspace>
222
- python $CLAUDE_SKILL_DIR/grade-viewer/generate_grade_review.py <workspace> --static grade-report.html
223
- ```
224
- Shows historical grades from GRADES.jsonl, A/B summaries from any workspace subdirectory.
225
-
226
- ---
227
-
228
- ## CLI Grade Operations
229
-
230
- | Command | Description |
231
- |---------|-------------|
232
- | `ct grade <sessionId>` | Grade a specific session |
233
- | `ct grade --list` | List past grade results |
234
- | `ct session start --scope global --name "<n>" --grade` | Start graded session |
235
- | `ct session end` | End session |