@noir-ai/skills 1.9.3 → 1.9.4-beta.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (57) hide show
  1. package/builtin/noir-backend/SKILL.md +32 -4
  2. package/builtin/noir-backend/references/backend-patterns.md +55 -0
  3. package/builtin/noir-brainstorming/SKILL.md +49 -0
  4. package/builtin/noir-checkpoint/SKILL.md +30 -12
  5. package/builtin/noir-context/SKILL.md +29 -10
  6. package/builtin/noir-doctor/SKILL.md +23 -4
  7. package/builtin/noir-executing-plans/SKILL.md +59 -0
  8. package/builtin/noir-exploring/SKILL.md +37 -0
  9. package/builtin/noir-frontend/SKILL.md +32 -4
  10. package/builtin/noir-frontend/references/ui-patterns.md +48 -0
  11. package/builtin/noir-parallel/SKILL.md +25 -35
  12. package/builtin/noir-planning/SKILL.md +46 -0
  13. package/builtin/noir-prd/SKILL.md +38 -27
  14. package/builtin/noir-readme/SKILL.md +25 -4
  15. package/builtin/noir-recall/SKILL.md +31 -12
  16. package/builtin/noir-remember/SKILL.md +31 -17
  17. package/builtin/noir-rules/SKILL.md +28 -27
  18. package/builtin/noir-security/SKILL.md +33 -4
  19. package/builtin/noir-security/references/security-checklist.md +49 -0
  20. package/builtin/noir-shipping/SKILL.md +34 -0
  21. package/builtin/noir-spec/SKILL.md +48 -14
  22. package/builtin/noir-spec/references/spec-template.md +43 -0
  23. package/builtin/noir-subagent/SKILL.md +29 -30
  24. package/builtin/noir-subagent/references/dispatch-guide.md +50 -0
  25. package/builtin/noir-sync/SKILL.md +39 -12
  26. package/builtin/noir-systematic-debugging/SKILL.md +71 -0
  27. package/builtin/noir-systematic-debugging/references/tracing.md +51 -0
  28. package/builtin/noir-test-driven-development/SKILL.md +73 -0
  29. package/builtin/noir-test-driven-development/references/tdd-worked-example.md +65 -0
  30. package/builtin/noir-verifying/SKILL.md +40 -0
  31. package/builtin/noir-verifying/references/verification-checklist.md +29 -0
  32. package/builtin/noir-worktree/SKILL.md +23 -4
  33. package/builtin/noir-wrap/SKILL.md +29 -13
  34. package/builtin/noir-writing-skills/SKILL.md +42 -0
  35. package/dist/index.d.ts +159 -9
  36. package/dist/index.js +240 -7
  37. package/dist/index.js.map +1 -1
  38. package/evals/noir-systematic-debugging/evals.json +25 -0
  39. package/evals/noir-test-driven-development/evals.json +25 -0
  40. package/integrations/noir-clickup/SKILL.md +319 -71
  41. package/package.json +3 -2
  42. package/builtin/noir-brainstorm/SKILL.md +0 -17
  43. package/builtin/noir-branch/SKILL.md +0 -12
  44. package/builtin/noir-clarify/SKILL.md +0 -17
  45. package/builtin/noir-commit/SKILL.md +0 -12
  46. package/builtin/noir-debug/SKILL.md +0 -38
  47. package/builtin/noir-document/SKILL.md +0 -17
  48. package/builtin/noir-execute/SKILL.md +0 -17
  49. package/builtin/noir-explore/SKILL.md +0 -16
  50. package/builtin/noir-intake/SKILL.md +0 -17
  51. package/builtin/noir-plan/SKILL.md +0 -20
  52. package/builtin/noir-pr/SKILL.md +0 -12
  53. package/builtin/noir-review/SKILL.md +0 -28
  54. package/builtin/noir-skill-author/SKILL.md +0 -12
  55. package/builtin/noir-tdd/SKILL.md +0 -49
  56. package/builtin/noir-test/SKILL.md +0 -12
  57. package/builtin/noir-verify/SKILL.md +0 -17
@@ -1,47 +1,46 @@
1
1
  ---
2
2
  name: noir-subagent
3
- description: Use when executing an implementation plan with independent tasks — to drive a fresh subagent per task with review between.
3
+ description: Use when dispatching independent implementation tasks as fresh subagents — one task per subagent with review between. Use when the user says "use subagents" or "dispatch these". Do NOT use for sequential tasks that share state; use noir-executing-plans instead.
4
+ metadata:
5
+ category: execute
6
+ version: 1.0.0
7
+ license: MIT
8
+ compatibility: claude · agents-md · gemini · cursor · opencode
9
+ references:
10
+ - dispatch-guide.md
4
11
  ---
5
12
 
6
- Execute an implementation plan by dispatching a fresh implementer subagent per task, a task review (spec compliance + code quality) after each, and a broad whole-branch review at the end. The controller curates exactly what each subagent needs; subagents never inherit the controller's session history, and artifacts move as files rather than pasted text. This keeps each subagent focused and preserves the controller's context for coordination.
13
+ # noir-subagent
14
+
15
+ Dispatch each independent task to a fresh subagent with a brief, then review the result. Parallel execution with review gates — the fastest path through a fan-out plan.
16
+
17
+ ## When to use
18
+
19
+ - An implementation plan has independent tasks (no shared state, no ordering).
20
+ - The user says "use subagents", "dispatch these", or "run them in parallel."
21
+ - **Do NOT use:** for sequential tasks (task B reads task A's output) — use `noir-executing-plans`. For general concurrent work without per-task review, use `noir-parallel`.
7
22
 
8
23
  ## Procedure
9
24
 
10
25
  ### Pre-flight
11
- 1. **Read the plan once.** Note the global constraints (exact values, formats, cross-task relationships) and any conflicts between tasks. If you find tasks that contradict each other or the plan's own constraints, surface the whole batch to the user as one question before dispatching Task 1 — not one interrupt per conflict mid-plan.
12
- 2. **Open a progress ledger.** Create `.noir/sdd/progress.md`. Each task gets one line when its review comes back clean: `Task N: complete (commits <base7>..<head7>, review clean)`. After any compaction or resume, trust the ledger and `git log` over your recollection — controllers that lost their place have re-dispatched entire completed sequences.
26
+ 1. **Write a brief per task.** Each subagent gets: the task goal, the files it touches, the expected output, and a deadline. The brief lives at `.noir/sdd/task-N-brief.md`.
27
+ 2. **Validate independence.** If task B needs task A's output, they're not independent — order them or merge them.
13
28
 
14
29
  ### Per task
15
- 3. **Brief the implementer.** Extract the task's full text to `.noir/sdd/task-N-brief.md` (the single source of requirements — exact values, signatures, test cases live there). Compose the dispatch with: one line on where this task fits, the brief path ("read this first — it is your requirements"), interfaces and decisions from earlier tasks the brief cannot know, your resolution of any ambiguity you noticed, and the report path `.noir/sdd/task-N-report.md`. Name the model explicitly — omitted models inherit the session default and silently defeat the cost/performance choice below.
16
- 4. **Handle implementer status.**
17
- - **DONE** — generate the review package and dispatch the task reviewer.
18
- - **DONE_WITH_CONCERNS** — read the concerns; address correctness/scope concerns before review, note observations and proceed.
19
- - **NEEDS_CONTEXT** — provide the missing context, re-dispatch.
20
- - **BLOCKED** — change something before retrying: more context, a more capable model, a smaller task, or escalate to the user if the plan itself is wrong.
21
- 5. **Review the task.** Run `git diff <BASE> <HEAD>` for the task's commit range into a file (BASE = the commit you recorded before dispatching the implementer, never `HEAD~1` — multi-commit tasks get truncated) and pass the path to the reviewer along with the brief and the report. The reviewer returns two verdicts: spec compliance and code quality. Both are required.
22
- 6. **Adjudicate findings.** Dispatch fix subagents for Critical and Important findings (one fixer carrying the complete list, not one per finding — per-finding fixers each rebuild context and re-run suites). Re-review until clean. Resolve "cannot verify from diff" items yourself before marking the task complete — you hold the plan and cross-task context the reviewer lacks. If a finding conflicts with what the plan's text mandates, present the finding and the plan text to the user and ask which governs; do not dismiss it and do not silently fix past the plan.
23
- 7. **Mark complete.** Append the ledger line, move to the next task. Do not pause for check-ins between tasks — execute the plan through, stopping only on BLOCKED status you cannot resolve, genuine ambiguity, or all tasks complete.
30
+ 3. **Dispatch the subagent.** Fresh subagent, clean context, reads the brief and the spec. On Claude Code, use the `Task` tool with a specific subagent type (e.g. `Explore` for search, `Plan` for design) — or use `Agent` with `subagent_type` for custom configurations. On other hosts, dispatch whatever agent/subagent runner is available. **Issue multiple dispatches in ONE response to run them concurrently.**
31
+ 4. **Receive the report.** The subagent returns a `.noir/sdd/task-N-report.md` — what was done, test results, any issues.
32
+ 5. **Review.** Verify: tests pass, spec satisfied, commits are clean. If not, return to the subagent with the specific issues — don't rewrite their code yourself (that defeats the isolation). If a task needs rework, re-dispatch with the issues as context.
24
33
 
25
34
  ### Final
26
- 8. **Dispatch the whole-branch review.** Package the branch diff with `MERGE_BASE = git merge-base main HEAD` and dispatch the final reviewer on the most capable available model — this is the judgment step, not the session default. If it returns findings, dispatch one fix subagent carrying the complete list.
27
- 9. **Record and route.** Capture the per-task reviews and the final verdict via `noir.checkpoint` (the SDD execute/verify gates). Hand off to `noir-document`.
35
+ 6. **Integrate.** All tasks done, all tests green, all briefs archived. Hand off to `noir-verifying`.
36
+
37
+ ## Reference
28
38
 
29
- ## Constructing reviewer prompts
30
- - Do not add open-ended directives ("check all uses", "run race tests if useful") without a concrete, task-specific reason.
31
- - Do not ask a reviewer to re-run tests the implementer already ran on the same code — the implementer's report carries the evidence.
32
- - Do not pre-judge findings for the reviewer. If your prompt contains "do not flag", "don't treat X as a defect", "at most Minor", or "the plan chose" — stop. You are pre-judging, usually to spare yourself a review loop. Let the reviewer raise it and adjudicate in the loop.
33
- - The global-constraints block you hand the reviewer is its attention lens: copy the binding requirements verbatim from the plan (exact values, formats, stated relationships). Process rules (YAGNI, test hygiene) already live in the reviewer template.
34
- - Every fix dispatch carries the implementer contract: re-run the tests covering the change and report the result (test files, command, output). Re-dispatch the reviewer only once all three are present.
39
+ For brief-writing, independence validation, and review mechanics, see [dispatch-guide.md](references/dispatch-guide.md).
35
40
 
36
- ## Model selection
37
- Use the least powerful model that handles each role.
38
- - Mechanical, well-specified implementation (complete spec, 1–2 files, or plan text that already contains the code): cheap/fast tier.
39
- - Integration and judgment (multi-file coordination, pattern matching): standard tier.
40
- - Architecture, design, and the final whole-branch review: most capable tier. Always specify the model explicitly; an omitted model inherits the session default.
41
+ ## When done → next skill
41
42
 
42
- Cheapest models often take 2–3× the turns on multi-step work — use a mid-tier model as the floor for reviewers and for implementers working from prose. Turn count beats token price.
43
+ → `noir-verifying` for the full integration gate, then `noir-shipping`.
43
44
 
44
45
  ## Notes
45
- - One dispatch describes one task, not the session's history. Do not paste accumulated prior-task summaries into later dispatches — a fresh subagent needs its task, its interfaces, and the global constraints. Nothing else.
46
- - Never start implementation on `main`/`master`; never run multiple implementation subagents in parallel (they conflict on the working tree); never let a subagent self-review replace the actual task review — both are needed.
47
- - Discipline is observable, not rhetorical: the SDD engine records each per-task review and the final whole-branch review, with the diff packages and verdicts attached.
46
+ - This skill is a playbook — the host decides which tools to use.
@@ -0,0 +1,50 @@
1
+ # Subagent dispatch guide — briefs, independence, review
2
+
3
+ Deep reference for `noir-subagent`. The mechanics of dispatching fresh subagents per task and reviewing between.
4
+
5
+ ## The brief — what a subagent needs
6
+
7
+ A subagent starts cold. Its brief must be self-contained:
8
+
9
+ - **Goal** — one sentence: what must be true when done.
10
+ - **Files** — the exact files it may touch (and which it must NOT touch).
11
+ - **Constraints** — the interface contract it must satisfy (signatures, formats).
12
+ - **Acceptance** — how it proves done (tests pass, output shape).
13
+ - **Context pointer** — where to read the relevant spec/plan/background (NOT a dump of everything).
14
+
15
+ A good brief is 5-10 lines. If it needs a paragraph of caveats, the task isn't independent yet.
16
+
17
+ ## Independence validation
18
+
19
+ Two tasks are independent only if:
20
+
21
+ - B does not need A's output to start (or you ORDER them: A then B).
22
+ - They do not write the same file in conflicting ways.
23
+ - They do not share mutable state (a DB, a cache, a port).
24
+
25
+ If any dependency exists, you have two choices: run them sequentially (`noir-executing-plans`), or merge them into one task. Dispatching dependent tasks in parallel produces merge hell.
26
+
27
+ ## The review step (non-negotiable)
28
+
29
+ After each subagent returns, REVIEW before integrating:
30
+
31
+ 1. **Tests pass?** Run the suite. Don't trust the subagent's "all green."
32
+ 2. **Spec satisfied?** The output meets the brief's acceptance criteria.
33
+ 3. **No collateral?** The subagent didn't touch files outside its brief, didn't leave debug logs, didn't commit garbage.
34
+
35
+ If any item fails, return the report to the subagent with the specific issue. Do NOT rewrite their code yourself — that defeats the isolation.
36
+
37
+ ## Files as the contract
38
+
39
+ Subagents communicate through FILES, not conversation. Each task writes:
40
+ - `task-N-brief.md` — what it was asked (you write before dispatch).
41
+ - `task-N-report.md` — what it did + test results (it writes on return).
42
+
43
+ These become the integration record and the `noir-wrap` handoff input. Structured outputs (JSON, CSV, markdown tables) beat prose for machine consumption.
44
+
45
+ ## When NOT to dispatch a subagent
46
+
47
+ - The task is a single small change (a subagent adds more overhead than value).
48
+ - The task requires deep conversation context you already hold.
49
+ - The task touches the same file as a parallel task.
50
+ - The model is simpler than the task warrants — match subagent capability to task difficulty.
@@ -1,18 +1,45 @@
1
1
  ---
2
2
  name: noir-sync
3
- description: Use at the start of a session — to load project context (anchors, git state, setup) and recall prior memory before doing anything else.
3
+ description: Use when starting any session or conversation — load project context and route to the relevant noir skill. Fires on explicit signals (feature start, spec request, new task). Do NOT use mid-session for a status update; use noir-checkpoint.
4
+ metadata:
5
+ category: discovery
6
+ version: 1.0.0
7
+ license: MIT
8
+ compatibility: claude · agents-md · gemini · cursor · opencode
4
9
  ---
5
10
 
6
- Load project context at session start. Read-only — do not edit project files. Discipline comes from the SDD engine's observable gates, not from this skill.
11
+ # noir-sync — Session starter and skill router
12
+
13
+ Load the project context AND route to the right skill. This is the entry point for every session — read anchors, check state, recall memory, then hand off to the noir skill that matches the task. Stay silent on trivial edits; fire on explicit signals.
14
+
15
+ ## When to use
16
+
17
+ - Start of a session — before any other action.
18
+ - The user opens a new conversation or says "let's work on X."
19
+ - **Do NOT use:** mid-session for progress — use `noir-checkpoint`. Do NOT use for a one-line fix with no design.
7
20
 
8
21
  ## Procedure
9
- 1. **Read anchors.** Read `CLAUDE.md` / `AGENTS.md` (always-on context) to ground the session.
10
- 2. **Project setup.** Detect stack + the test/run command from manifests and `CLAUDE.md`; note project type. Informational — flag anything non-conventional, do not mutate.
11
- 3. **Git state.** `git rev-parse --abbrev-ref HEAD`; `git status --porcelain`; `git log --oneline -5`.
12
- 4. **Noir state.** Read `.noir/NOIR.md`; check `.noir/tasks/` for an in-flight task (the engine's persisted state). If one exists, surface its phase via `noir.workflow_status` and offer to resume it through the SDD lifecycle — do not auto-resume.
13
- 5. **Recall memory.** Query Noir memory (or the host's recall tooling, if any); surface 2–4 top facts relevant to the current task. If empty, say so — `.noir/` artifacts are the durable fallback.
14
- 6. **Print a brief** of essentials only (project type, stack, key commands, in-flight task) — no large dumps. Then await the user; do not auto-start work.
15
-
16
- ## Fallbacks
17
- - No memory found → say so; `.noir/NOIR.md` + `.noir/tasks/` are always the durable record.
18
- - Not initialized (no `.noir/`) → tell the user to run `noir init` before continuing.
22
+
23
+ 1. **Read anchors.** `CLAUDE.md`, `AGENTS.md`, `.noir/NOIR.md` — always-on context to ground the session. On Claude Code, use `Read`; on other hosts, use the equivalent file tool.
24
+ 2. **Check git state.** `git rev-parse --abbrev-ref HEAD`, `git status --porcelain`, `git log --oneline -5`. Note any dirty tree or in-progress work.
25
+ 3. **Check Noir state.** Read `.noir/tasks/` for an in-flight task. If one exists, surface its phase and offer to resume.
26
+ 4. **Recall memory.** Query Noir memory (or the host's recall tooling) for 2-4 top facts relevant to this session. If empty, say so — `.noir/` is the durable fallback.
27
+ 5. **Skill triage.** Map the user's intent to the right noir skill. These are the high-value skills:
28
+ - Feature/idea → `noir-brainstorming` → `noir-spec` → `noir-planning` → `noir-executing-plans`
29
+ - Bug/crash → `noir-systematic-debugging`
30
+ - Implement with TDD → `noir-test-driven-development`
31
+ - Verify/PR → `noir-verifying`
32
+ - Ship/integrate → `noir-shipping`
33
+ - Close session → `noir-wrap`
34
+ - Code exploration → `noir-exploring`
35
+ - Skill authoring → `noir-writing-skills`
36
+ 6. **Print a brief** — project type, stack, key commands, in-flight task. Then hand off to the routed skill.
37
+
38
+ ## Notes
39
+
40
+ - Don't auto-start work. Surface the state and let the user confirm direction.
41
+ - If no skill matches, ask what the user wants to do — the route is a suggestion, not a mandate.
42
+
43
+ ## When done → next skill
44
+
45
+ Route to the matched skill. Or is there something else you'd like to do first?
@@ -0,0 +1,71 @@
1
+ ---
2
+ name: noir-systematic-debugging
3
+ description: Use when encountering any bug, test failure, or unexpected behavior — find the root cause methodically before proposing any fix. Use when the user says "this is broken", "it fails intermittently", or pastes a stack trace. Do NOT use for a known fix where the root cause is already confirmed.
4
+ metadata:
5
+ category: execute
6
+ version: 1.0.0
7
+ license: MIT
8
+ compatibility: claude · agents-md · gemini · cursor · opencode
9
+ references:
10
+ - tracing.md
11
+ ---
12
+
13
+ # noir-systematic-debugging
14
+
15
+ Find root cause before attempting any fix. Symptom patches waste time and create new bugs; the path is investigate → hypothesize → test → fix, in that order. Skipping ahead is the most expensive move in debugging.
16
+
17
+ ## When to use
18
+
19
+ - A bug, test failure, or unexpected behavior appears — and you don't know the root cause yet.
20
+ - The user says "this is broken", "it fails", "something's wrong", or pastes a stack trace.
21
+ - A test passed yesterday and fails today, or it fails intermittently.
22
+ - **Do NOT use:** when the root cause is already clear and the fix is obvious (one-line patch with no risk of side effects).
23
+
24
+ ## Procedure
25
+
26
+ ### 1. Reproduce and read
27
+ - Read the error completely. Stack traces, line numbers, file paths, error codes — do not skim past the message; it usually narrows the search immediately.
28
+ - Reproduce reliably. Find the minimal, repeatable trigger. If you cannot reproduce it, gather more data; do not guess at a fix from a single trace.
29
+ - Check recent changes. `git diff` and recent commits often point straight at the cause. Note new dependencies, config edits, and environmental differences.
30
+
31
+ ### 2. Gather evidence at component boundaries
32
+ When the system has multiple layers (CLI → adapter → daemon → store, or API → service → database), instrument each boundary before proposing fixes: log what enters and exits, verify config propagation, check state at each layer. On Claude Code, use `Bash` to run the exact failing command with verbose flags; on other hosts, use whatever shell/execution tool is available. Run once, read the output, then investigate the specific layer the evidence implicates.
33
+
34
+ ### 3. Trace data flow to the source
35
+ For a bad value deep in a call stack, trace backward: where does it originate, who passed it down, who created it. Fix at the source, not at the symptom. Resist the urge to patch where the symptom surfaces.
36
+
37
+ ### 4. Find the pattern
38
+ Locate similar working code in the same codebase and list every difference from the broken path — however small. When applying a known pattern, read the reference implementation completely; partial understanding is a common cause of the bug you are now debugging.
39
+
40
+ ### 5. Form one hypothesis and test minimally
41
+ State the hypothesis explicitly ("I think X is the root cause because Y"). Make the smallest possible change that would prove or disprove it — one variable at a time, never a bundle. If the hypothesis survives the minimal test, proceed; if not, form a new hypothesis with the new information. Do not stack fixes on top of a failed one.
42
+
43
+ ### 6. Fix at the root, with a regression test
44
+ - Write the simplest failing test that reproduces the bug (see `noir-test-driven-development`).
45
+ - Implement the single fix at the root cause.
46
+ - Verify the test goes green and that no other test regressed.
47
+ - One change at a time; no "while I'm here" refactors bundled into the fix.
48
+
49
+ ### 7. Know when to stop fixing
50
+ If three or more fixes have failed, each revealing a new problem in a different place, the issue is likely architectural, not a missing patch. Stop, name the pattern that is not holding, and surface the architectural question (with evidence) instead of attempting fix number four.
51
+
52
+ ## Verification
53
+
54
+ - [ ] The bug is reproduced reliably (the trigger is documented).
55
+ - [ ] A root cause hypothesis is stated explicitly ("X because Y").
56
+ - [ ] The fix is at the source, not the symptom.
57
+ - [ ] A regression test passes and no other tests regressed.
58
+ - [ ] One commit per fix — no bundled refactors.
59
+
60
+ ## Notes
61
+
62
+ - "No root cause found" is usually an incomplete investigation. If the bug is genuinely environmental or timing-dependent after a complete pass, document what you investigated, implement appropriate handling (retry, timeout, error message), and add monitoring — but say so plainly rather than implying the root cause is unknowable.
63
+ - If the bug spans multiple services, use `noir-exploring` to fan out the evidence search first.
64
+
65
+ ## Reference
66
+
67
+ For deeper detail, see [tracing.md](references/tracing.md).
68
+
69
+ ## When done → next skill
70
+
71
+ → `noir-verifying` to confirm the fix is complete. Or do you need to investigate another bug?
@@ -0,0 +1,51 @@
1
+ # Root-cause tracing — evidence at every boundary
2
+
3
+ Deep reference for `noir-systematic-debugging`. Use this when the bug spans multiple layers or the root cause is not obvious from the first stack trace.
4
+
5
+ ## Why boundaries matter
6
+
7
+ A bug is rarely where it surfaces. "The API returns 500" could mean the route handler, the service, the DB layer, or the config. If you fix at the surface, you patch a symptom; the same root cause fails again through another path. Tracing evidence at each boundary converts a guess into a data-backed narrowing.
8
+
9
+ ## The boundary checklist
10
+
11
+ For a request that flows CLI → adapter → daemon → store (Noir's own shape), or API → service → DB, instrument each layer BEFORE proposing a fix:
12
+
13
+ 1. **Entry** — what arrived? Log the raw request: method, path, headers, body (sanitized).
14
+ 2. **Validation** — did the input pass validation? A rejected payload fails here.
15
+ 3. **Auth/context** — was identity resolved? A missing token fails here.
16
+ 4. **Service logic** — what did the service compute? Log inputs + outputs at the boundary.
17
+ 5. **Persistence** — what SQL/query ran? Log the query + params, then the result/error.
18
+ 6. **Exit** — what returned? Compare the response to the handler's intent.
19
+
20
+ ## Technique: log at the boundary, not inside
21
+
22
+ Put one log line at each boundary (enter/exit) rather than scattering logs through the middle. The boundary log answers "did X reach layer N correctly?" — the middle log answers "what happened inside layer N?" Start with boundaries; only go inside when a boundary shows the input was wrong.
23
+
24
+ ## Technique: binary search the layers
25
+
26
+ Instead of instrumenting all 6 boundaries, bisect: check the MIDDLE layer first. If the middle sees the right input and produces the wrong output, the bug is in the middle or below. If the middle sees wrong input, the bug is above. Each check halves the search space.
27
+
28
+ ## Technique: the "did it change?" check
29
+
30
+ The most common root cause is a recent change. Before tracing:
31
+ - `git log --oneline -10` — what changed recently?
32
+ - `git diff HEAD~1` — what's different in the code that touches this path?
33
+ - Check config/dependency changes: a new package version, a toggled feature flag, an env var change.
34
+
35
+ ## Recording the trace
36
+
37
+ For a debugging session, keep a compact trace log:
38
+
39
+ ```
40
+ REQUEST GET /api/v2/task/86eyeryfe?include_subtasks=true
41
+ VALIDATE ok
42
+ AUTH pk_*** (ok)
43
+ SERVICE fetched task, subtasks=[] (NOT null!) ← suspicious: empty array
44
+ FINDING subtasks excluded by default; include_subtasks=true missing
45
+ ```
46
+
47
+ The trace makes the evidence explicit and reviewable — the same honesty `noir-verifying` demands.
48
+
49
+ ## When the trace points to an architectural issue
50
+
51
+ If 3+ fixes have failed, each revealing a new problem in a different layer, stop fixing and name the pattern that is not holding (e.g. "the daemon caches process.env at spawn, so any env-var-dependent path is stale"). Surface it as an architectural question with the trace as evidence.
@@ -0,0 +1,73 @@
1
+ ---
2
+ name: noir-test-driven-development
3
+ description: Use when implementing any feature or bugfix — write the failing test first with the RED→GREEN→REFACTOR loop. Do NOT use for pure refactors where behavior doesn't change.
4
+ metadata:
5
+ category: execute
6
+ version: 1.0.0
7
+ license: MIT
8
+ compatibility: claude · agents-md · gemini · cursor · opencode
9
+ references:
10
+ - tdd-worked-example.md
11
+ ---
12
+
13
+ # noir-test-driven-development
14
+
15
+ Every feature starts with a failing test. The RED → GREEN → REFACTOR loop is the proof: if there's no failing test, there's no evidence the feature is needed. This skill absorbs `noir-test` — test design guidance lives here too.
16
+
17
+ ## When to use
18
+
19
+ - You're about to implement a feature or fix a bug.
20
+ - The user says "write a test", "test this", or "TDD this".
21
+ - The implementation is non-trivial — a behavior change, new logic, or API contract.
22
+ - **Do NOT use:** for a pure rename, formatting change, or refactor where the existing test suite already covers the behavior.
23
+
24
+ ## Procedure
25
+
26
+ ### RED — Write the failing test
27
+ 1. **Write the test FIRST.** Before any implementation code. The test must fail for the RIGHT reason (the feature is missing), not because of a syntax error or a broken import.
28
+ 2. **The test is the spec.** Write it with the behavior you want to see — the input, the expected output, the edge cases. A good test reads like a user story: "given X, when Y, then Z."
29
+ 3. **Run the test to confirm it fails.** The failure message must be clear — "feature not yet implemented", not "undefined is not a function."
30
+
31
+ ### Verify RED — confirm it fails for the right reason
32
+ - The test fails on the assertion you wrote, not on setup/import/boot.
33
+ - If it fails for the wrong reason, fix the test (not the code under test) until it fails cleanly.
34
+
35
+ ### GREEN — Implement the minimal code
36
+ - Write the smallest amount of production code that makes the test pass. Nothing more — no "future-proofing," no "while I'm here."
37
+ - Run the test. It must pass. Run the full suite to confirm no regression.
38
+
39
+ ### Verify GREEN — confirm it passes legitimately
40
+ - The test passes on the assertion you wrote — not on a coincidental edge case.
41
+ - No other test broke. If one did, fix it NOW; do not carry a broken suite.
42
+
43
+ ### REFACTOR — Clean up
44
+ - With a green suite as the safety net, clean up: remove duplication, improve naming, simplify logic. The test suite protects you — a broken refactor fails immediately.
45
+ - Run the full suite after every refactor step.
46
+
47
+ ### Repeat
48
+ - For the next behavior, go back to RED. One test → one implementation → one refactor per cycle.
49
+
50
+ ## Why order matters
51
+ Writing the test first forces you to state the desired behavior before you're influenced by the implementation. Writing the implementation to the test keeps you from building more than needed. Refactoring last, with a green suite, makes cleanup safe.
52
+
53
+ ## Verification
54
+
55
+ - [ ] A failing test was written BEFORE the implementation.
56
+ - [ ] The test failed for the right reason (feature missing, not broken import).
57
+ - [ ] The minimal implementation made the test pass.
58
+ - [ ] The full test suite is green after the change.
59
+ - [ ] Refactored code is cleaner, not just different.
60
+
61
+ ## Notes
62
+
63
+ - Don't test the framework. Test YOUR logic.
64
+ - One assertion per behavior, one test per behavior. A test that asserts five things is five tests fighting for attention.
65
+ - If TDD feels slow, you're probably fixing a bug that someone else shipped because they skipped it.
66
+
67
+ ## Reference
68
+
69
+ For a worked RED → GREEN → REFACTOR walkthrough, see [tdd-worked-example.md](references/tdd-worked-example.md).
70
+
71
+ ## When done → next skill
72
+
73
+ → `noir-verifying` to gather evidence the work is complete. Or is there another behavior to implement?
@@ -0,0 +1,65 @@
1
+ # TDD worked example — RED → GREEN → REFACTOR
2
+
3
+ Deep reference for `noir-test-driven-development`. A complete walkthrough of the loop on a real-shaped problem.
4
+
5
+ ## The problem
6
+
7
+ Add a function that counts words in a string, ignoring punctuation and case.
8
+
9
+ ## RED — write the failing test first
10
+
11
+ ```ts
12
+ import { describe, expect, it } from 'vitest';
13
+ import { countWords } from './count-words.js';
14
+
15
+ describe('countWords', () => {
16
+ it('counts simple words', () => {
17
+ expect(countWords('hello world')).toBe(2);
18
+ });
19
+ it('ignores punctuation and is case-insensitive', () => {
20
+ expect(countWords('Hello, WORLD!')).toBe(2);
21
+ });
22
+ it('returns 0 for empty/whitespace input', () => {
23
+ expect(countWords(' ')).toBe(0);
24
+ });
25
+ });
26
+ ```
27
+
28
+ Run it. Expected: FAIL with `Cannot find module './count-words.js'` or `countWords is not a function` — the failure is because the feature doesn't exist, which is the RIGHT reason.
29
+
30
+ **Verify RED:** the failure is on your assertion / missing module, NOT on a broken import in the test itself.
31
+
32
+ ## GREEN — minimal implementation
33
+
34
+ ```ts
35
+ export function countWords(text: string): number {
36
+ const words = text.toLowerCase().match(/[a-z]+/g);
37
+ return words ? words.length : 0;
38
+ }
39
+ ```
40
+
41
+ Run the test. PASS. Run the full suite — no regression.
42
+
43
+ **Verify GREEN:** it passes for the right reason (the regex matches words), not a coincidental edge case.
44
+
45
+ ## REFACTOR — clean up with the green suite as a net
46
+
47
+ ```ts
48
+ const WORD = /[a-z]+/g;
49
+ export function countWords(text: string): number {
50
+ return text.toLowerCase().match(WORD)?.length ?? 0;
51
+ }
52
+ ```
53
+
54
+ Run the suite again. Still green. The refactor is safe because the tests prove behavior.
55
+
56
+ ## Repeat
57
+
58
+ Add the next behavior → back to RED. One test → one implementation → one refactor per cycle. Never skip RED: a feature without a failing test has no evidence it was needed.
59
+
60
+ ## Anti-patterns this example guards against
61
+
62
+ - Writing the implementation before the test (no RED) — you can't prove the feature is needed.
63
+ - A test that passes before the implementation — you're not driving new behavior.
64
+ - Bundling a refactor into the GREEN step — the refactor belongs AFTER green.
65
+ - One test asserting five behaviors — split into five tests, each with one clear failure.
@@ -0,0 +1,40 @@
1
+ ---
2
+ name: noir-verifying
3
+ description: Use when about to claim work is complete — gather evidence (tests, gate, criteria) before asserting success. Use when the user says "is this done?" or "verify this". Do NOT use mid-implementation for a progress check; use noir-checkpoint.
4
+ metadata:
5
+ category: verify
6
+ version: 1.0.0
7
+ license: MIT
8
+ compatibility: claude · agents-md · gemini · cursor · opencode
9
+ references:
10
+ - verification-checklist.md
11
+ ---
12
+
13
+ # noir-verifying
14
+
15
+ Gather evidence before asserting success. Absorbs `noir-verify` and `noir-review` into one gate: run, read, verify, then claim done.
16
+
17
+ ## When to use
18
+
19
+ - About to claim a task, feature, or fix is complete.
20
+ - The user says "verify this", "is this done?", "check everything."
21
+ - **Do NOT use:** mid-implementation — use `noir-checkpoint`.
22
+
23
+ ## Procedure
24
+
25
+ 1. **IDENTIFY criteria.** Load spec/plan — each acceptance criterion is a verification item.
26
+ 2. **RUN the gate.** Execute the project's test command. The output IS the evidence.
27
+ 3. **READ results.** A passing suite is the minimum; verify each criterion explicitly.
28
+ 4. **Check side effects.** New lint warnings? Type errors? Doc links? Run the full gate.
29
+ 5. **ONLY THEN claim done.** State "verified: <evidence>" — never "verified: looks good."
30
+
31
+ ## Reference
32
+
33
+ For the full claim-done + PR-review checklist, see [verification-checklist.md](references/verification-checklist.md).
34
+
35
+ ## When done → next skill
36
+
37
+ → `noir-shipping` to commit and integrate, or `noir-wrap` to close. Or something else?
38
+
39
+ ## Notes
40
+ - This skill is a playbook — the host decides which tools to use. On Claude Code, prefer `AskUserQuestion` for choices; on other hosts, ask in text.
@@ -0,0 +1,29 @@
1
+ # Verification checklist — evidence before "done"
2
+
3
+ Deep reference for `noir-verifying`. A copyable checklist for the two moments it covers: claiming a task done, and reviewing a PR.
4
+
5
+ ## The claim-done gate
6
+
7
+ Before asserting "this is done," satisfy EVERY item with evidence:
8
+
9
+ - [ ] **Spec criteria met** — every acceptance criterion from the spec/plan is citable to an actual behavior, test, or output. Not "looks implemented" — point to the evidence.
10
+ - [ ] **Tests green** — the full suite passes. The OUTPUT was read, not assumed. Quote the result line.
11
+ - [ ] **No regression** — tests that passed before still pass. A green suite isn't enough if you deleted a test to make it green.
12
+ - [ ] **Lint/typecheck** — the project's static checks pass (e.g. `pnpm lint`, `pnpm typecheck`).
13
+ - [ ] **Side effects checked** — docs, generated files, or config that the change touches reflect the new reality (docs reflect shipped state, never stale).
14
+ - [ ] **The claim is evidence-backed** — "verified: <test output / command result>" not "verified: looks good."
15
+
16
+ ## The PR-review gate
17
+
18
+ Beyond the claim-done gate, review asks "does this make sense?"
19
+
20
+ - [ ] **Approach** — is this the right way to solve it, or a workaround?
21
+ - [ ] **Naming** — do names say what things are?
22
+ - [ ] **Scope** — does the PR do one thing, or did "while I'm here" changes creep in?
23
+ - [ ] **Security** — user input, auth, secrets: any exposure?
24
+ - [ ] **Tests** — do the tests test the behavior, not the implementation detail?
25
+ - [ ] **Diff hygiene** — no commented-out code, no debug logs, no accidental files.
26
+
27
+ ## The honesty rule
28
+
29
+ **"Seems right" is never sufficient.** State the evidence and let it stand on its own. If you cannot produce evidence for an item, the item is NOT done — say so, rather than asserting completion. The same discipline that keeps a debugging trace honest keeps a "done" claim honest.
@@ -1,12 +1,31 @@
1
1
  ---
2
2
  name: noir-worktree
3
- description: Use when starting feature work that needs isolation from the current workspace.
3
+ description: Use when creating an isolated git workspace for feature work — keeping the main checkout clean. Use when the user says "worktree" or "isolate this work".
4
+ metadata:
5
+ category: git
6
+ version: 1.0.0
7
+ license: MIT
8
+ compatibility: claude · agents-md · gemini · cursor · opencode
4
9
  ---
5
10
 
6
11
  # noir-worktree
7
12
 
8
- > **Stub:** this skill ships as a valid, loadable placeholder in S5; its full playbook is deepened in a later slice.
9
13
 
10
- **When to use:** you are starting work that should not disturb the current workspace (long-running branch, parallel feature, risky change).
14
+ ## When to use
15
+ - When the user triggers this skill.
11
16
 
12
- **For now:** create an isolated worktree on a fresh branch, do the work there, and merge or discard when done — keep the main workspace clean.
17
+ Isolate feature work from the current workspace via git worktrees. One feature, one directory, no cross-contamination.
18
+
19
+ ## Procedure
20
+
21
+ 1. **Confirm the feature needs isolation.** A single-file fix in a clean repo doesn't need a worktree. A multi-file feature with branches and dependencies does.
22
+ 2. **Create the worktree.** `git worktree add -b <feature-branch> <path> <base-branch>`. On Claude Code, if a `.claude/worktrees/` pattern exists, follow it.
23
+ 3. **Work in the isolated directory.** The worktree has its own checkout, and the original directory is untouched.
24
+ 4. **When done, clean up.** `git worktree remove <path>` + `git branch -d <feature-branch>` (after merge).
25
+
26
+ ## When done → next skill
27
+
28
+ → `noir-shipping` to integrate the finished work. Or continue in the worktree.
29
+
30
+ ## Notes
31
+ - This skill is a playbook — the host decides which tools to use. On Claude Code, prefer `AskUserQuestion` for choices; on other hosts, ask in text.
@@ -1,23 +1,39 @@
1
1
  ---
2
2
  name: noir-wrap
3
- description: Use when closing a session cleanly — run tests, curate docs, confirm commits, save memory, emit a host handoff.
3
+ description: Use when closing a work session cleanly — running final verification, updating docs and CHANGELOG, saving memory, and emitting a host handoff. Do NOT use mid-session for a checkpoint — use noir-checkpoint.
4
+ metadata:
5
+ category: document
6
+ version: 1.0.0
7
+ license: MIT
8
+ compatibility: claude · agents-md · gemini · cursor · opencode
4
9
  ---
5
10
 
6
11
  # noir-wrap
7
12
 
8
- Use when you are ending a session and want to leave the work in a clean, recoverable state — and hand off to the host CLI with a ready-to-paste prompt.
13
+ Close a session in a clean, recoverable state. This skill absorbs `noir-document` (update docs/CHANGELOG/memory) — wrap is the superset that covers the full close.
9
14
 
10
- ## Steps
15
+ ## When to use
11
16
 
12
- 1. Run the project's test command (the host detects it; if unknown, `pnpm test` / `npm test`).
13
- 2. Curate or delete ephemeral notes (scratch docs, dead branches, tmp files).
14
- 3. Confirm commits are made — and pushed or intentionally local (Noir keeps commits local + conservative by default).
15
- 4. Save durable memory before closing: observations, decisions, patterns the next session should recall. Prefer `noir memory save` (or the `noir.remember` MCP tool from the host) so cross-session recall works.
16
- 5. Advance the workflow task if a gate is satisfied: `noir task advance --to <phase>` (the verify gate prints the handoff hint automatically).
17
- 6. Emit the host handoff — run `noir handoff` (or the session-end alias `noir wrap`). This prints a ready-to-paste markdown prompt to STDOUT that names the active task, the next gate's skill, a bounded context/memory seed, and the exact host-launch directive. Pipe it straight into the host, or persist with `noir handoff --write` (the path is gitignored under `.noir/handoff/`).
17
+ - Ending a work session.
18
+ - The user says "wrap up", "I'm done", "close this out."
19
+ - **Do NOT use:** mid-session — use `noir-checkpoint`.
18
20
 
19
- ## Notes
21
+ ## Procedure
20
22
 
21
- - `noir handoff` reuses the same snapshot as `noir status` and the same phase→skill map as `noir task next`, so the handoff is always consistent with the live state.
22
- - The handoff directive is TEXT ONLY — Noir never launches the host. Paste the block into the host CLI to resume.
23
- - For a machine-readable handoff (e.g. a CI consumer), use `noir handoff --json`.
23
+ 1. **Run final verification.** Same gate as `noir-verifying`: test suite, lint, typecheck. Evidence, not assumption.
24
+ 2. **Update docs.** CHANGELOG, ADRs, decisions, reference docs — anything that should reflect the session's work. The rule: docs reflect shipped reality, never a stale plan.
25
+ 3. **Save memory.** Persist observations, decisions, patterns the next session should recall. `noir memory save` (or `noir.remember` MCP tool).
26
+ 4. **Confirm commits.** Commits are made and intentional (local or pushed). Noir defaults to local.
27
+ 5. **Advance the workflow task.** `noir task advance --to <phase>` if a gate is satisfied.
28
+ 6. **Emit the handoff.** `noir handoff` (text-only prompt; `--write` persists to `.noir/handoff/`; `--json` for CI). Names the active task, next gate's skill, and the host-launch directive.
29
+
30
+ ## Verification
31
+
32
+ - [ ] Gate is green (tests, lint, typecheck).
33
+ - [ ] Docs are synced (CHANGELOG, decisions, references).
34
+ - [ ] Memory is saved (key observations + decisions).
35
+ - [ ] Handoff emitted (the next session can resume).
36
+
37
+ ## When done → next skill
38
+
39
+ The session is closed. Until next time.