@cleocode/skills 2026.5.83 → 2026.5.86

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (55) hide show
  1. package/package.json +1 -1
  2. package/skills/_shared/__tests__/lifecycle-protocol-reconcile.test.ts +112 -0
  3. package/skills/_shared/__tests__/loom-adr-links.test.ts +163 -0
  4. package/skills/_shared/__tests__/loom-stage-coverage.test.ts +167 -0
  5. package/skills/ct-adr-recorder/SKILL.md +92 -0
  6. package/skills/ct-adr-recorder/__tests__/skill-adr-recorder.test.ts +65 -0
  7. package/skills/ct-consensus-voter/SKILL.md +14 -0
  8. package/skills/ct-contribution/SKILL.md +80 -0
  9. package/skills/ct-docs-lookup/SKILL.md +116 -1
  10. package/skills/ct-docs-lookup/references/ctx7-workflow.md +198 -0
  11. package/skills/ct-docs-lookup/references/library-id-resolution.md +217 -0
  12. package/skills/ct-docs-lookup/references/version-specific-docs.md +220 -0
  13. package/skills/ct-docs-review/SKILL.md +133 -1
  14. package/skills/ct-docs-review/__tests__/skill-docs-review.test.ts +53 -0
  15. package/skills/ct-docs-review/references/inline-comment-patterns.md +268 -0
  16. package/skills/ct-docs-review/references/pr-review-mode.md +270 -0
  17. package/skills/ct-docs-review/references/style-violations.md +341 -0
  18. package/skills/ct-docs-write/SKILL.md +157 -1
  19. package/skills/ct-docs-write/__tests__/skill-docs-write.test.ts +55 -0
  20. package/skills/ct-docs-write/references/audience-targeting.md +305 -0
  21. package/skills/ct-docs-write/references/cleo-style-guide.md +234 -0
  22. package/skills/ct-docs-write/references/markdown-patterns.md +329 -0
  23. package/skills/ct-documentor/SKILL.md +11 -0
  24. package/skills/ct-documentor/references/anti-patterns.md +216 -0
  25. package/skills/ct-documentor/references/chain-orchestration.md +194 -0
  26. package/skills/ct-documentor/references/doc-types-and-templates.md +301 -0
  27. package/skills/ct-documentor/references/style-coordination.md +195 -0
  28. package/skills/ct-epic-architect/SKILL.md +15 -0
  29. package/skills/ct-ivt-looper/SKILL.md +32 -0
  30. package/skills/ct-release-orchestrator/SKILL.md +16 -0
  31. package/skills/ct-research-agent/SKILL.md +24 -0
  32. package/skills/ct-research-agent/references/anti-patterns.md +154 -0
  33. package/skills/ct-research-agent/references/citation-and-evidence.md +140 -0
  34. package/skills/ct-research-agent/references/source-strategy.md +116 -0
  35. package/skills/ct-research-agent/references/triggers-and-routing.md +93 -0
  36. package/skills/ct-skill-validator/SKILL.md +19 -0
  37. package/skills/ct-skill-validator/scripts/check_depth.py +306 -0
  38. package/skills/ct-spec-writer/SKILL.md +86 -1
  39. package/skills/ct-spec-writer/__tests__/skill-spec-writer.test.ts +60 -0
  40. package/skills/ct-spec-writer/references/anti-patterns.md +176 -0
  41. package/skills/ct-spec-writer/references/rfc2119-language.md +138 -0
  42. package/skills/ct-spec-writer/references/spec-templates.md +233 -0
  43. package/skills/ct-spec-writer/references/traceability-matrix.md +145 -0
  44. package/skills/ct-task-executor/SKILL.md +25 -0
  45. package/skills/ct-task-executor/references/acceptance-criteria-mapping.md +163 -0
  46. package/skills/ct-task-executor/references/anti-patterns.md +201 -0
  47. package/skills/ct-task-executor/references/common-failures.md +193 -0
  48. package/skills/ct-task-executor/references/evidence-and-gates.md +179 -0
  49. package/skills/ct-task-executor/references/implementation-patterns.md +160 -0
  50. package/skills/ct-validator/SKILL.md +44 -0
  51. package/skills/ct-validator/references/anti-patterns.md +194 -0
  52. package/skills/ct-validator/references/compliance-reports.md +199 -0
  53. package/skills/ct-validator/references/schema-checking.md +191 -0
  54. package/skills/ct-validator/references/validation-modes.md +185 -0
  55. package/skills/manifest.json +82 -16
@@ -0,0 +1,195 @@
1
+ # Style Coordination
2
+
3
+ The CLEO documentation style guide lives at
4
+ `packages/skills/skills/_shared/cleo-style-guide.md` and is referenced
5
+ by both `ct-docs-write` and `ct-docs-review`. As coordinator, the
6
+ documentor's role is to ensure the style guide is enforced consistently
7
+ across the chain — write applies it; review verifies it. This reference
8
+ covers the parts of the style guide most commonly violated and how the
9
+ documentor enforces them.
10
+
11
+ ## Tone Pillars
12
+
13
+ The CLEO style is **conversational, clear, user-focused**. Three pillars
14
+ the documentor must enforce across the chain.
15
+
16
+ ### 1. Conversational
17
+
18
+ Write the way you would explain to a colleague — not the way a corporate
19
+ press release would. This shows up at the word level.
20
+
21
+ | Avoid | Prefer |
22
+ |-------|--------|
23
+ | utilize | use |
24
+ | reference | see / look at |
25
+ | offerings | products / features |
26
+ | we will / we can | CLEO does / CLEO can |
27
+ | cannot, do not | can't, don't |
28
+ | people are able to | people can |
29
+
30
+ The contraction rule is real and unusual — CLEO docs USE contractions.
31
+ Many style guides ban them; CLEO does the opposite.
32
+
33
+ ### 2. Clear
34
+
35
+ Lead with the action; explain after. Headings state the point.
36
+
37
+ | Vague heading | Clear heading |
38
+ |---------------|---------------|
39
+ | "Environment variables" | "Use environment variables for configuration" |
40
+ | "Authentication" | "Authenticate with API tokens" |
41
+ | "Configuration" | "Configure release pipeline before first ship" |
42
+ | "Performance considerations" | "Set timeout to 30s on slow networks" |
43
+
44
+ Vague headings force the reader to scan the body to learn the point.
45
+ Clear headings let the reader skip the section if it's not what they
46
+ need.
47
+
48
+ ### 3. User-Focused
49
+
50
+ Use "people" or "companies", not "users". The first refers to humans;
51
+ the second to a system construct. The CLEO docs are for humans.
52
+
53
+ | Avoid | Prefer |
54
+ |-------|--------|
55
+ | users can | people can |
56
+ | our users | our customers / the companies using CLEO |
57
+ | user input | what people type |
58
+ | user experience | what people see |
59
+
60
+ This is jarring at first — many docs are full of "users". Once you
61
+ switch, the writing reads more concretely.
62
+
63
+ ## Forbidden Phrases
64
+
65
+ These never appear in CLEO docs. The review child rejects them on sight.
66
+
67
+ - "easy" / "simple" / "just" — patronizing; the reader doesn't know if
68
+ it's easy until they try
69
+ - "obviously" / "of course" — same problem
70
+ - "for free" / "out of the box" — vague
71
+ - "click here" / "read more here" — link text must be descriptive
72
+ - "and so on" / "etc." — be specific or cut the sentence
73
+ - "as mentioned above" — use a stable cross-reference
74
+
75
+ The forbidden list is the most common reason a draft fails review.
76
+ Pass the list to ct-docs-write as part of the input contract so it
77
+ knows what to avoid.
78
+
79
+ ## Link Discipline
80
+
81
+ Link text must describe the destination. Never link bare words like
82
+ "here", "this", "read more".
83
+
84
+ ```markdown
85
+ GOOD: See the [release pipeline ADR](../.cleo/adrs/ADR-065.md) for the gate set.
86
+
87
+ BAD: See ADR-065 [here](../.cleo/adrs/ADR-065.md) for the gate set.
88
+
89
+ BAD: For more info, [click here](../.cleo/adrs/ADR-065.md).
90
+ ```
91
+
92
+ The full link text MUST make sense out of context — a reader scanning
93
+ just the link list should understand each destination.
94
+
95
+ ## Code Block Discipline
96
+
97
+ Code blocks include their language tag for syntax highlighting and
98
+ their working-directory context.
99
+
100
+ ````markdown
101
+ GOOD:
102
+ ```bash
103
+ # In the project root
104
+ pnpm run test
105
+ ```
106
+
107
+ GOOD:
108
+ ```typescript
109
+ // packages/cleo/src/foo.ts
110
+ import { bar } from "@cleocode/contracts";
111
+ ```
112
+
113
+ BAD:
114
+ ```
115
+ pnpm run test
116
+ ```
117
+ (no language tag)
118
+
119
+ BAD:
120
+ ```ts
121
+ import { bar } from "@cleocode/contracts";
122
+ ```
123
+ (use `typescript`, not `ts`)
124
+ ````
125
+
126
+ The full language tags are `bash`, `typescript`, `tsx`, `python`,
127
+ `rust`, `json`, `yaml`, `markdown`. Avoid abbreviations.
128
+
129
+ ## Table Discipline
130
+
131
+ Tables compress information that would otherwise sprawl. Use them
132
+ liberally — they signal "look up data" mode.
133
+
134
+ | Use a table when... | Don't use when... |
135
+ |---------------------|-------------------|
136
+ | You have parallel structured items | You have unstructured prose |
137
+ | Each row has the same columns | Each "row" has different fields |
138
+ | The reader will scan, not read | The reader needs narrative |
139
+
140
+ Keep tables narrow — three to five columns is the sweet spot. Wider
141
+ tables overflow on mobile and look like data dumps.
142
+
143
+ ## Image Discipline
144
+
145
+ Images are scoped — they show one specific UI element with relevant
146
+ context. They never show full-window screenshots.
147
+
148
+ | Scope | When to use |
149
+ |-------|-------------|
150
+ | UI element + label | Annotating a feature |
151
+ | Specific dialog | Showing a flow step |
152
+ | Code in editor | Showing syntax highlighting |
153
+ | Whole window | Almost never |
154
+
155
+ Add alt text that describes what the image conveys, not what the image
156
+ is of. "Screenshot of the release pipeline dashboard showing three
157
+ green checks" rather than "Dashboard screenshot".
158
+
159
+ ## Cross-Skill Drift Detection
160
+
161
+ Documentor MUST check that write's output and review's checks agree.
162
+ If they disagree, the style guide reference is the tiebreaker.
163
+
164
+ ```text
165
+ write_output = ct-docs-write(input)
166
+ review_findings = ct-docs-review(write_output)
167
+
168
+ # Possible drift case:
169
+ # write produced "users" (it should have used "people")
170
+ # review didn't flag it (its rule list was outdated)
171
+ # → file a sync task; passing the style guide explicitly to write
172
+ # on next iteration
173
+
174
+ # Possible drift case:
175
+ # write produced "click here"
176
+ # review flagged "click here" with high confidence
177
+ # → no drift; write just made a mistake; iterate
178
+ ```
179
+
180
+ When drift is detected, surface it in `needs_followup` so the next
181
+ session updates the shared style guide or the child skill's rules.
182
+
183
+ ## Pre-PR Style Pass
184
+
185
+ Before opening a PR with documentation changes, the documentor MUST
186
+ run review one more time on the diff:
187
+
188
+ ```bash
189
+ # Review the diff for style violations
190
+ gh pr diff <PR_NUMBER> | ct-docs-review --mode=diff
191
+ ```
192
+
193
+ PR-mode review uses the `mcp__github__*` tools when available;
194
+ otherwise local mode is fine. The pre-PR pass catches drift that
195
+ sneaked in during integration.
@@ -6,6 +6,10 @@ tier: 1
6
6
  core: false
7
7
  category: recommended
8
8
  protocol: decomposition
9
+ loomStage: decomposition
10
+ adrRefs:
11
+ - ADR-066
12
+ - ADR-073
9
13
  dependencies: []
10
14
  sharedResources:
11
15
  - subagent-protocol-base
@@ -329,3 +333,14 @@ Recommendation: [Your recommendation]
329
333
  | 6 | Validation | Escape `$` as `\$`, check fields |
330
334
 
331
335
  **Shell Escaping**: Always `\$` in `--notes`/`--description`. See [shell-escaping.md](references/shell-escaping.md).
336
+
337
+ ---
338
+
339
+ ## See also / References
340
+
341
+ This skill binds to the **decomposition** LOOM lifecycle stage. Governing ADRs:
342
+
343
+ - [ADR-066 — task taxonomy consolidation](../../../../.cleo/adrs/ADR-066-task-taxonomy-consolidation.md) — defines the Type/Kind/Severity axes the decomposer must populate on every leaf task.
344
+ - [ADR-073 — above-epic naming](../../../../.cleo/adrs/ADR-073-above-epic-naming.md) — defines the Saga/Epic/Task/Subtask hierarchy that decomposition produces.
345
+
346
+ LOOM coverage matrix: [docs/skills/loom-coverage-matrix.md](../../../../docs/skills/loom-coverage-matrix.md).
@@ -1,6 +1,11 @@
1
1
  ---
2
2
  name: ct-ivt-looper
3
3
  description: "Runs a project-agnostic autonomous Implement-then-Validate-then-Test compliance loop on any git worktree. Detects the project's test framework (vitest, jest, mocha, pytest, unittest, go-test, cargo-test, rspec, phpunit, bats, or other) and iterates until the implementation satisfies its specification, recording convergence metrics to the manifest. Use when given an implementation task that must ship verified: the IVT loop is the autonomous compliance layer enforced before any release or PR. Triggers on phrases like 'implement and verify', 'run the IVT loop', 'ship this task', 'complete implementation with tests', 'verify against spec', or any implementation task with acceptance criteria. Works in any git worktree regardless of language or framework, never hardcoded to one project's tooling."
4
+ protocol: testing
5
+ loomStage: testing
6
+ adrRefs:
7
+ - ADR-051
8
+ - ADR-061
4
9
  ---
5
10
 
6
11
  # IVT Looper
@@ -76,6 +81,24 @@ escalate_to_hitl() # IVT-007: exit code 65
76
81
 
77
82
  The loop is a *single* stage from the lifecycle's point of view. Implement, Validate, and Test are not three separate tasks — they are three phases of one autonomous run that either converges or escalates.
78
83
 
84
+ ## Out of Scope (T9675)
85
+
86
+ `ct-ivt-looper` operates on the **`testing`** LOOM lifecycle stage (stage 8). It performs the **dynamic** Implement-then-Validate-then-Test loop with framework detection and iterate-until-green convergence semantics.
87
+
88
+ This skill does NOT:
89
+
90
+ - Audit static artifacts for schema/compliance/RFC-2119 keyword usage, ADR-document structure, or JSON Schema conformance. Those belong to **`ct-validator`** at the `validation` stage (stage 7). When the question is "is this document/manifest well-formed?" rather than "does this code converge on its spec?", chain to `ct-validator` rather than expanding scope here.
91
+ - Promote a green loop to release. Release sequencing belongs to **`ct-release-orchestrator`** at stage 9.
92
+
93
+ ### Chain handoffs
94
+
95
+ | Direction | When | Handoff |
96
+ |---|---|---|
97
+ | `ct-ivt-looper` → `ct-validator` | Loop converged; need to audit the resulting artifacts (e.g. final manifest, spec back-references) against schema/compliance | Emit the convergence manifest entry, then dispatch the `validation` stage |
98
+ | `ct-validator` → `ct-ivt-looper` | Spec is valid but implementation needs dynamic verification | Receive a dispatch from the `validation` stage; iterate the IVT loop on the worktree |
99
+
100
+ Governance: see **ADR-051** (programmatic gate integrity) which defines the evidence atoms (`tool:test`, `test-run:<json>`) the loop emits and that downstream `cleo verify --gate testsPassed` re-validates, and **ADR-061** (project-agnostic verify tools) which defines the canonical tool-resolution layer.
101
+
79
102
  ## Framework Detection
80
103
 
81
104
  Framework detection is project-agnostic: the skill walks the worktree, inspects config files, and selects the correct test command. No language or framework is special-cased above another. The full detection table lives in [references/frameworks.md](references/frameworks.md). In summary: detection reads the project manifest (`package.json`, `pyproject.toml`, `Cargo.toml`, `go.mod`, `Gemfile`, `composer.json`, or `.cleo/project-context.json#testing.command`) and selects one of: `vitest`, `jest`, `mocha`, `pytest`, `unittest`, `go-test`, `cargo-test`, `rspec`, `phpunit`, `bats`, `other`.
@@ -179,3 +202,12 @@ cleo check protocol \
179
202
  6. Record `framework`, `testsRun`, `testsPassed`, `testsFailed`, `ivtLoopConverged`, `ivtLoopIterations` in the manifest.
180
203
  7. On non-convergence, exit 65 and leave the worktree untouched.
181
204
  8. Validate every run via `cleo check protocol --protocolType testing`.
205
+
206
+ ## See also / References
207
+
208
+ This skill binds to the **testing** LOOM lifecycle stage. Governing ADRs:
209
+
210
+ - [ADR-051 — programmatic gate integrity](../../../../.cleo/adrs/ADR-051-programmatic-gate-integrity.md) — defines the evidence atoms (`tool:test`, `test-run:<json>`) that the IVT loop emits and that downstream `cleo verify --gate testsPassed` re-validates.
211
+ - [ADR-061 — project-agnostic verify tools](../../../../.cleo/adrs/ADR-061-project-agnostic-verify-tools.md) — defines the canonical tool-resolution layer (`test`, `build`, `lint`, `typecheck`) that the loop walks for framework-agnostic execution.
212
+
213
+ LOOM coverage matrix: [docs/skills/loom-coverage-matrix.md](../../../../docs/skills/loom-coverage-matrix.md).
@@ -1,6 +1,12 @@
1
1
  ---
2
2
  name: ct-release-orchestrator
3
3
  description: "Orchestrates the full release pipeline: version bump, then changelog, then commit, then tag, then conditionally forks to artifact-publish and provenance based on release config. Parent protocol that composes ct-artifact-publisher and ct-provenance-keeper as sub-protocols: not every release publishes artifacts (source-only releases skip it), and artifact publishers delegate signing and attestation to provenance. Use when shipping a new version, running cleo release ship, or promoting a completed epic to released status."
4
+ protocol: release
5
+ loomStage: release
6
+ adrRefs:
7
+ - ADR-053
8
+ - ADR-063
9
+ - ADR-065
4
10
  ---
5
11
 
6
12
  # Release Orchestrator
@@ -132,3 +138,13 @@ For source-only releases, pass `--no-artifacts` to skip the artifact-publish han
132
138
  6. `released` entries are immutable; hotfixes go into new entries.
133
139
  7. Manifest entry MUST set `agent_type: "documentation"` and record the full chain via `record_release()`.
134
140
  8. Always validate via `cleo check protocol --protocolType release` before declaring the release done.
141
+
142
+ ## See also / References
143
+
144
+ This skill binds to the **release** LOOM lifecycle stage. Governing ADRs:
145
+
146
+ - [ADR-053 — project-agnostic release pipeline](../../../../.cleo/adrs/ADR-053-project-agnostic-release-pipeline.md) — defines the language-agnostic version bump → changelog → tag flow.
147
+ - [ADR-063 — release pipeline](../../../../.cleo/adrs/ADR-063-release-pipeline.md) — defines the 12-step `cleo release ship` integration with CI.
148
+ - [ADR-065 — PR-required release flow](../../../../.cleo/adrs/ADR-065-pr-required-release-flow.md) — defines the PR-gated path; direct pushes to `main` are prohibited.
149
+
150
+ LOOM coverage matrix: [docs/skills/loom-coverage-matrix.md](../../../../docs/skills/loom-coverage-matrix.md).
@@ -6,6 +6,10 @@ tier: 2
6
6
  core: false
7
7
  category: recommended
8
8
  protocol: research
9
+ loomStage: research
10
+ adrRefs:
11
+ - ADR-023
12
+ - ADR-070
9
13
  dependencies: []
10
14
  sharedResources:
11
15
  - subagent-protocol-base
@@ -224,3 +228,23 @@ If research cannot proceed (access denied, topic too broad, etc.):
224
228
  - **Prioritized** - Most important first
225
229
  - **Justified** - Tied to specific findings
226
230
  - **Feasible** - Achievable within project constraints
231
+
232
+ ---
233
+
234
+ ## See also / References
235
+
236
+ This skill binds to the **research** LOOM lifecycle stage. Governing ADRs:
237
+
238
+ - [ADR-023 — protocol validation dispatch](../../../../.cleo/adrs/ADR-023-protocol-validation-dispatch.md) — defines how research output is validated before downstream stages consume it.
239
+ - [ADR-070 — three-tier orchestration](../../../../.cleo/adrs/ADR-070-three-tier-orchestration.md) — defines the Orchestrator → Phase Lead → Worker tiers; research runs as a leaf Worker.
240
+
241
+ LOOM coverage matrix: [docs/skills/loom-coverage-matrix.md](../../../../docs/skills/loom-coverage-matrix.md).
242
+
243
+ ## See references/
244
+
245
+ Progressive disclosure — load on demand only:
246
+
247
+ - `references/triggers-and-routing.md` — when to use ct-research-agent and what downstream skill consumes the output
248
+ - `references/source-strategy.md` — codebase + Context7 + web hierarchy, query patterns, time-boxing
249
+ - `references/citation-and-evidence.md` — evidence ladder, citation format, confidence labels
250
+ - `references/anti-patterns.md` — ten failure modes observed in past CLEO sessions and how to avoid them
@@ -0,0 +1,154 @@
1
+ # Anti-Patterns
2
+
3
+ Common failure modes for research tasks. Each pattern below has been
4
+ observed in real CLEO sessions (some recorded in BRAIN as patterns
5
+ `P-b2e59bf4`, `P-75035d53`, and others). Avoid them by following the
6
+ detection cue and applying the remediation.
7
+
8
+ ## 1. The Hallucinated Citation
9
+
10
+ **Symptom.** A finding cites a URL, function signature, or API that does
11
+ not actually exist. The agent inferred its existence from naming
12
+ conventions and a plausible domain.
13
+
14
+ **Detection cue.** Citation has not been retrieved during this session.
15
+ URL was not produced by a tool call; signature was not read from a file.
16
+
17
+ **Remediation.** Every citation MUST come from a tool call output in this
18
+ session — `WebFetch`, `WebSearch`, `Grep`, `Read`, `ctx7 docs`, or
19
+ `gitnexus_*`. If the source cannot be re-shown by a tool call, it is
20
+ hallucinated. Strip it and either re-fetch or downgrade the finding to
21
+ `hypothesis`.
22
+
23
+ ## 2. The Stale Authority
24
+
25
+ **Symptom.** A finding cites the official docs for library X — but the
26
+ project uses version Y, and the cited page describes version Z which has
27
+ different semantics.
28
+
29
+ **Detection cue.** No version qualifier on the citation. The
30
+ `package.json` or `Cargo.toml` for the project specifies a version that
31
+ differs from the doc page version.
32
+
33
+ **Remediation.** Always pin Context7 queries to the project's actual
34
+ version. Use the `/org/project/version` form when available. Read the
35
+ project's lockfile before citing version-dependent behavior.
36
+
37
+ ## 3. The Single-Blog Cascade
38
+
39
+ **Symptom.** A blog post made an interesting claim; the research output
40
+ cited it as fact; downstream the spec writer baked it into requirements;
41
+ the implementation fails because the claim was wrong.
42
+
43
+ **Detection cue.** A finding sourced to a single blog post or Medium
44
+ article is labeled `verified` or `documented` in the output.
45
+
46
+ **Remediation.** Blog posts are at most `reported` rung-3 evidence.
47
+ Promote to `documented` only after corroborating with the canonical
48
+ source (official docs, source code, or RFC). If no canonical source
49
+ exists, label `anecdotal` and put it under `## Hypotheses`.
50
+
51
+ ## 4. The Boil-the-Ocean
52
+
53
+ **Symptom.** The research task is "investigate JavaScript build tools".
54
+ The agent spends 90 minutes producing a 50-page survey of every tool from
55
+ Browserify to Bun, exhausts its token budget, and never answers the
56
+ question the orchestrator actually had.
57
+
58
+ **Detection cue.** Output exceeds 30 KB without a `## Recommendations`
59
+ section in the first 20% of the file.
60
+
61
+ **Remediation.** Before opening any source, restate the question in one
62
+ sentence and pick the success criterion. "Boil the ocean" tasks SHOULD be
63
+ split into multiple narrower research tasks via the orchestrator — return
64
+ a `needs_followup` list rather than producing one giant document.
65
+
66
+ ## 5. The Confirmation Loop
67
+
68
+ **Symptom.** The user (or the calling task) implied a preferred answer in
69
+ the question. The agent finds that preferred answer immediately, stops
70
+ searching, and reports it as if it had explored alternatives.
71
+
72
+ **Detection cue.** The output lists fewer than 2 alternatives even when
73
+ the task description used comparative language ("compare", "evaluate",
74
+ "options"). All findings support a single direction.
75
+
76
+ **Remediation.** For comparative tasks, mandate one paragraph per
77
+ alternative with stated pros and cons, even if the agent expects to
78
+ recommend one. The reader needs to see the tradeoff space.
79
+
80
+ ## 6. The Ungrounded Recommendation
81
+
82
+ **Symptom.** The recommendations section contains imperative statements
83
+ ("use library X", "adopt pattern Y") that are not tied to any specific
84
+ finding above.
85
+
86
+ **Detection cue.** Recommendations cannot be traced back to a numbered
87
+ finding or cited source.
88
+
89
+ **Remediation.** Each recommendation MUST reference at least one finding
90
+ by name or section heading. Use this format:
91
+
92
+ ```markdown
93
+ 1. Adopt `defineRelations` for the new schema work.
94
+ - Based on finding: "Drizzle v1 deprecates `relations()` in favor of
95
+ `defineRelations` (verified, 0.95)"
96
+ - Tradeoff: minor migration cost; offset by ergonomics + future-proof.
97
+ ```
98
+
99
+ ## 7. The Forgotten Codebase
100
+
101
+ **Symptom.** Research output cites web sources extensively but never
102
+ references the existing codebase, ADRs, or BRAIN memory. The
103
+ recommendation contradicts a decision already recorded in ADR-XXX.
104
+
105
+ **Detection cue.** No `.cleo/adrs/`, `packages/`, or `cleo memory find`
106
+ references in the citations.
107
+
108
+ **Remediation.** ALWAYS run a codebase + BRAIN sweep before opening the
109
+ web. Even when the topic is "external", the project has likely already
110
+ made related decisions that constrain the answer space. ADRs are
111
+ canonical — overriding them requires explicit consensus, not silent
112
+ research recommendation.
113
+
114
+ ## 8. The Manifest Stuffer
115
+
116
+ **Symptom.** The pipeline_manifest entry's `key_findings` array contains
117
+ 20+ items, half of which are minor or duplicative. The orchestrator's
118
+ briefing surface chokes on the noise.
119
+
120
+ **Detection cue.** `key_findings` length > 7 or contains items shorter
121
+ than 8 words.
122
+
123
+ **Remediation.** `key_findings` is for the orchestrator's roll-up — 3-7
124
+ sentence-length, action-oriented items only. Detail belongs in the
125
+ output file. If a finding cannot be summarized in one sentence, it is
126
+ two findings.
127
+
128
+ ## 9. The Premature Completion
129
+
130
+ **Symptom.** Task is marked complete; the output file says "research
131
+ complete" but the recommendations are vague ("further investigation
132
+ needed", "more work required") with no specific followup IDs.
133
+
134
+ **Detection cue.** Manifest status is `complete` but the file contains
135
+ phrases like "TBD", "future work", or "to be determined".
136
+
137
+ **Remediation.** If research cannot reach actionable recommendations,
138
+ the manifest status MUST be `partial` or `blocked`, NOT `complete`.
139
+ Specific gaps belong in `needs_followup` as concrete task descriptions,
140
+ not as soft prose in the file body.
141
+
142
+ ## 10. The Format Drift
143
+
144
+ **Symptom.** Output file does not match the template in SKILL.md.
145
+ Sections are renamed, the manifest entry uses old field names, or the
146
+ file location is wrong.
147
+
148
+ **Detection cue.** Validator script fails or the orchestrator reports
149
+ "could not parse manifest entry".
150
+
151
+ **Remediation.** Re-read SKILL.md's "Output File Format" and "Manifest
152
+ Entry Format" sections before writing. The downstream consumers are
153
+ brittle by design — they trust the format and will reject anything
154
+ unexpected.
@@ -0,0 +1,140 @@
1
+ # Citation and Evidence
2
+
3
+ Research is only useful if it is verifiable. This reference defines the
4
+ evidence ladder, citation format, and confidence labels that every research
5
+ finding in CLEO MUST carry. The downstream consumers — `ct-spec-writer`,
6
+ `ct-consensus-voter`, the orchestrator's HITL gates — rely on these signals
7
+ to weight findings correctly.
8
+
9
+ ## Evidence Ladder
10
+
11
+ Findings sit on a five-rung ladder. Each finding in the output file MUST be
12
+ marked with the strongest rung that applies.
13
+
14
+ | Rung | Label | Meaning | Citable? |
15
+ |------|-------|---------|----------|
16
+ | 5 | `verified` | Reproduced locally or confirmed in two independent canonical sources | Yes |
17
+ | 4 | `documented` | Confirmed in one canonical source (official docs, RFC, source code) | Yes |
18
+ | 3 | `reported` | Stated in a credible community source (named blog, conference talk, well-cited GitHub issue) | Yes, with caveat |
19
+ | 2 | `anecdotal` | Stated in an uncredentialed source (random blog, forum post) | No — needs corroboration |
20
+ | 1 | `hypothesis` | Inferred but not verified | Never |
21
+
22
+ Findings at rung 1-2 MUST live in a separate `## Hypotheses` section, never
23
+ mixed with verified findings. The spec writer downstream will silently treat
24
+ all findings as verified facts unless the rung labels are explicit.
25
+
26
+ ## Canonical Sources by Domain
27
+
28
+ | Domain | Canonical source |
29
+ |--------|------------------|
30
+ | HTTP/REST semantics | IETF RFCs (RFC 7230-7235, RFC 9110) |
31
+ | TLS/crypto | RFC + IANA registries + NIST publications |
32
+ | JavaScript language | ECMAScript spec (tc39.es) + MDN |
33
+ | TypeScript | typescript-go source or microsoft/TypeScript |
34
+ | Node.js APIs | nodejs.org/api/* + Node source code |
35
+ | Library X | github.com/<owner>/<repo> README + docs + release notes |
36
+ | CLEO architecture | `.cleo/adrs/ADR-*.md` (project-internal canon) |
37
+ | CLEO past decisions | `cleo memory find` + `.cleo/agent-outputs/` |
38
+ | Cloud platforms (AWS/GCP/Azure) | Vendor's "what's new" page + product docs |
39
+
40
+ Blog posts and Medium articles are NOT canonical — they are interpretive
41
+ layer that may have drifted from the source.
42
+
43
+ ## Citation Format Standards
44
+
45
+ Every citable finding includes at minimum: source identifier, retrieval
46
+ location, and (where relevant) retrieval date.
47
+
48
+ ### Web Source
49
+
50
+ ```markdown
51
+ According to the [Next.js 15 caching docs](https://nextjs.org/docs/app/building-your-application/caching),
52
+ the default fetch cache behavior changed from `force-cache` to `no-store`.
53
+ Retrieved 2026-05-19.
54
+ ```
55
+
56
+ ### Official Repository
57
+
58
+ ```markdown
59
+ The `defineRelations` API was introduced in
60
+ [drizzle-team/drizzle-orm@e2b9c1a](https://github.com/drizzle-team/drizzle-orm/commit/e2b9c1a)
61
+ and replaces the legacy `relations()` helper.
62
+ ```
63
+
64
+ ### Project-Internal (ADR)
65
+
66
+ ```markdown
67
+ ADR-065 §3 establishes the PR-gated release pipeline — direct pushes to
68
+ `main` are prohibited (`.cleo/adrs/ADR-065-release-pipeline.md`).
69
+ ```
70
+
71
+ ### Project-Internal (Code)
72
+
73
+ ```markdown
74
+ The current implementation in `packages/core/src/store/openCleoDb.ts:42-78`
75
+ routes all DB opens through a single chokepoint per Decision D003.
76
+ ```
77
+
78
+ ### Project-Internal (BRAIN)
79
+
80
+ ```markdown
81
+ BRAIN observation `O-mpd07uma-0` records that pub1-diagnoser correctly
82
+ refused under broken dispatcher — confirming the protocol-correct refusal
83
+ pattern is exercising at agent-level.
84
+ ```
85
+
86
+ ### Context7 Fetch
87
+
88
+ ```markdown
89
+ Context7 query `/vercel/next.js/v15` against the question "how do I set
90
+ revalidation interval" returns the `revalidate` segment-config option as
91
+ the canonical mechanism.
92
+ ```
93
+
94
+ ## Conflict Handling
95
+
96
+ When two canonical sources disagree, the research output MUST surface the
97
+ conflict rather than silently picking one. Recommended pattern:
98
+
99
+ ```markdown
100
+ ### Finding: Default cache behavior in Next.js 15
101
+
102
+ - The [official caching docs](https://nextjs.org/docs/app/.../caching)
103
+ state the default is `no-store`.
104
+ - The [release notes for 15.0](https://nextjs.org/blog/next-15) describe
105
+ the change but also mention a `staleTimes` config that softens it.
106
+ - Resolution: the default IS `no-store`, but `staleTimes` can restore
107
+ per-segment caching when explicitly configured.
108
+ ```
109
+
110
+ Conflict surfaces are the most valuable research output — they prevent
111
+ downstream consumers from making decisions on partial information.
112
+
113
+ ## Confidence Labels
114
+
115
+ In addition to the evidence rung, each `key_findings` entry in the manifest
116
+ SHOULD carry a confidence band when the rung is below 4. The orchestrator
117
+ uses this for HITL gating in `ct-consensus-voter`.
118
+
119
+ | Confidence | Meaning |
120
+ |------------|---------|
121
+ | 0.9–1.0 | Verified — multiple canonical sources agree |
122
+ | 0.7–0.9 | High — one canonical source, no conflicts |
123
+ | 0.5–0.7 | Medium — community sources agree, no canonical |
124
+ | 0.3–0.5 | Low — single source or conflicting evidence |
125
+ | 0.0–0.3 | Speculation — hypothesis only, not citable |
126
+
127
+ The manifest entry MUST reflect these in `key_findings`:
128
+
129
+ ```json
130
+ {
131
+ "key_findings": [
132
+ "Next.js 15 default cache is no-store (verified, 0.95)",
133
+ "staleTimes config can restore caching (documented, 0.85)",
134
+ "Migration path for existing apps requires per-route audit (reported, 0.6)"
135
+ ]
136
+ }
137
+ ```
138
+
139
+ Downstream skills filter by confidence — `ct-spec-writer` ignores findings
140
+ below 0.5 unless the user explicitly opts in.