@noir-ai/skills 1.9.3 → 1.9.4-beta.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/builtin/noir-backend/SKILL.md +32 -4
- package/builtin/noir-backend/references/backend-patterns.md +55 -0
- package/builtin/noir-brainstorming/SKILL.md +49 -0
- package/builtin/noir-checkpoint/SKILL.md +30 -12
- package/builtin/noir-context/SKILL.md +29 -10
- package/builtin/noir-doctor/SKILL.md +23 -4
- package/builtin/noir-executing-plans/SKILL.md +59 -0
- package/builtin/noir-exploring/SKILL.md +37 -0
- package/builtin/noir-frontend/SKILL.md +32 -4
- package/builtin/noir-frontend/references/ui-patterns.md +48 -0
- package/builtin/noir-parallel/SKILL.md +25 -35
- package/builtin/noir-planning/SKILL.md +46 -0
- package/builtin/noir-prd/SKILL.md +38 -27
- package/builtin/noir-readme/SKILL.md +25 -4
- package/builtin/noir-recall/SKILL.md +31 -12
- package/builtin/noir-remember/SKILL.md +31 -17
- package/builtin/noir-rules/SKILL.md +28 -27
- package/builtin/noir-security/SKILL.md +33 -4
- package/builtin/noir-security/references/security-checklist.md +49 -0
- package/builtin/noir-shipping/SKILL.md +34 -0
- package/builtin/noir-spec/SKILL.md +48 -14
- package/builtin/noir-spec/references/spec-template.md +43 -0
- package/builtin/noir-subagent/SKILL.md +29 -30
- package/builtin/noir-subagent/references/dispatch-guide.md +50 -0
- package/builtin/noir-sync/SKILL.md +39 -12
- package/builtin/noir-systematic-debugging/SKILL.md +71 -0
- package/builtin/noir-systematic-debugging/references/tracing.md +51 -0
- package/builtin/noir-test-driven-development/SKILL.md +73 -0
- package/builtin/noir-test-driven-development/references/tdd-worked-example.md +65 -0
- package/builtin/noir-verifying/SKILL.md +40 -0
- package/builtin/noir-verifying/references/verification-checklist.md +29 -0
- package/builtin/noir-worktree/SKILL.md +23 -4
- package/builtin/noir-wrap/SKILL.md +29 -13
- package/builtin/noir-writing-skills/SKILL.md +42 -0
- package/dist/index.d.ts +159 -9
- package/dist/index.js +240 -7
- package/dist/index.js.map +1 -1
- package/evals/noir-systematic-debugging/evals.json +25 -0
- package/evals/noir-test-driven-development/evals.json +25 -0
- package/integrations/noir-clickup/SKILL.md +319 -71
- package/package.json +3 -2
- package/builtin/noir-brainstorm/SKILL.md +0 -17
- package/builtin/noir-branch/SKILL.md +0 -12
- package/builtin/noir-clarify/SKILL.md +0 -17
- package/builtin/noir-commit/SKILL.md +0 -12
- package/builtin/noir-debug/SKILL.md +0 -38
- package/builtin/noir-document/SKILL.md +0 -17
- package/builtin/noir-execute/SKILL.md +0 -17
- package/builtin/noir-explore/SKILL.md +0 -16
- package/builtin/noir-intake/SKILL.md +0 -17
- package/builtin/noir-plan/SKILL.md +0 -20
- package/builtin/noir-pr/SKILL.md +0 -12
- package/builtin/noir-review/SKILL.md +0 -28
- package/builtin/noir-skill-author/SKILL.md +0 -12
- package/builtin/noir-tdd/SKILL.md +0 -49
- package/builtin/noir-test/SKILL.md +0 -12
- package/builtin/noir-verify/SKILL.md +0 -17
|
@@ -1,47 +1,46 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: noir-subagent
|
|
3
|
-
description: Use when
|
|
3
|
+
description: Use when dispatching independent implementation tasks as fresh subagents — one task per subagent with review between. Use when the user says "use subagents" or "dispatch these". Do NOT use for sequential tasks that share state; use noir-executing-plans instead.
|
|
4
|
+
metadata:
|
|
5
|
+
category: execute
|
|
6
|
+
version: 1.0.0
|
|
7
|
+
license: MIT
|
|
8
|
+
compatibility: claude · agents-md · gemini · cursor · opencode
|
|
9
|
+
references:
|
|
10
|
+
- dispatch-guide.md
|
|
4
11
|
---
|
|
5
12
|
|
|
6
|
-
|
|
13
|
+
# noir-subagent
|
|
14
|
+
|
|
15
|
+
Dispatch each independent task to a fresh subagent with a brief, then review the result. Parallel execution with review gates — the fastest path through a fan-out plan.
|
|
16
|
+
|
|
17
|
+
## When to use
|
|
18
|
+
|
|
19
|
+
- An implementation plan has independent tasks (no shared state, no ordering).
|
|
20
|
+
- The user says "use subagents", "dispatch these", or "run them in parallel."
|
|
21
|
+
- **Do NOT use:** for sequential tasks (task B reads task A's output) — use `noir-executing-plans`. For general concurrent work without per-task review, use `noir-parallel`.
|
|
7
22
|
|
|
8
23
|
## Procedure
|
|
9
24
|
|
|
10
25
|
### Pre-flight
|
|
11
|
-
1. **
|
|
12
|
-
2. **
|
|
26
|
+
1. **Write a brief per task.** Each subagent gets: the task goal, the files it touches, the expected output, and a deadline. The brief lives at `.noir/sdd/task-N-brief.md`.
|
|
27
|
+
2. **Validate independence.** If task B needs task A's output, they're not independent — order them or merge them.
|
|
13
28
|
|
|
14
29
|
### Per task
|
|
15
|
-
3. **
|
|
16
|
-
4. **
|
|
17
|
-
|
|
18
|
-
- **DONE_WITH_CONCERNS** — read the concerns; address correctness/scope concerns before review, note observations and proceed.
|
|
19
|
-
- **NEEDS_CONTEXT** — provide the missing context, re-dispatch.
|
|
20
|
-
- **BLOCKED** — change something before retrying: more context, a more capable model, a smaller task, or escalate to the user if the plan itself is wrong.
|
|
21
|
-
5. **Review the task.** Run `git diff <BASE> <HEAD>` for the task's commit range into a file (BASE = the commit you recorded before dispatching the implementer, never `HEAD~1` — multi-commit tasks get truncated) and pass the path to the reviewer along with the brief and the report. The reviewer returns two verdicts: spec compliance and code quality. Both are required.
|
|
22
|
-
6. **Adjudicate findings.** Dispatch fix subagents for Critical and Important findings (one fixer carrying the complete list, not one per finding — per-finding fixers each rebuild context and re-run suites). Re-review until clean. Resolve "cannot verify from diff" items yourself before marking the task complete — you hold the plan and cross-task context the reviewer lacks. If a finding conflicts with what the plan's text mandates, present the finding and the plan text to the user and ask which governs; do not dismiss it and do not silently fix past the plan.
|
|
23
|
-
7. **Mark complete.** Append the ledger line, move to the next task. Do not pause for check-ins between tasks — execute the plan through, stopping only on BLOCKED status you cannot resolve, genuine ambiguity, or all tasks complete.
|
|
30
|
+
3. **Dispatch the subagent.** Fresh subagent, clean context, reads the brief and the spec. On Claude Code, use the `Task` tool with a specific subagent type (e.g. `Explore` for search, `Plan` for design) — or use `Agent` with `subagent_type` for custom configurations. On other hosts, dispatch whatever agent/subagent runner is available. **Issue multiple dispatches in ONE response to run them concurrently.**
|
|
31
|
+
4. **Receive the report.** The subagent returns a `.noir/sdd/task-N-report.md` — what was done, test results, any issues.
|
|
32
|
+
5. **Review.** Verify: tests pass, spec satisfied, commits are clean. If not, return to the subagent with the specific issues — don't rewrite their code yourself (that defeats the isolation). If a task needs rework, re-dispatch with the issues as context.
|
|
24
33
|
|
|
25
34
|
### Final
|
|
26
|
-
|
|
27
|
-
|
|
35
|
+
6. **Integrate.** All tasks done, all tests green, all briefs archived. Hand off to `noir-verifying`.
|
|
36
|
+
|
|
37
|
+
## Reference
|
|
28
38
|
|
|
29
|
-
|
|
30
|
-
- Do not add open-ended directives ("check all uses", "run race tests if useful") without a concrete, task-specific reason.
|
|
31
|
-
- Do not ask a reviewer to re-run tests the implementer already ran on the same code — the implementer's report carries the evidence.
|
|
32
|
-
- Do not pre-judge findings for the reviewer. If your prompt contains "do not flag", "don't treat X as a defect", "at most Minor", or "the plan chose" — stop. You are pre-judging, usually to spare yourself a review loop. Let the reviewer raise it and adjudicate in the loop.
|
|
33
|
-
- The global-constraints block you hand the reviewer is its attention lens: copy the binding requirements verbatim from the plan (exact values, formats, stated relationships). Process rules (YAGNI, test hygiene) already live in the reviewer template.
|
|
34
|
-
- Every fix dispatch carries the implementer contract: re-run the tests covering the change and report the result (test files, command, output). Re-dispatch the reviewer only once all three are present.
|
|
39
|
+
For brief-writing, independence validation, and review mechanics, see [dispatch-guide.md](references/dispatch-guide.md).
|
|
35
40
|
|
|
36
|
-
##
|
|
37
|
-
Use the least powerful model that handles each role.
|
|
38
|
-
- Mechanical, well-specified implementation (complete spec, 1–2 files, or plan text that already contains the code): cheap/fast tier.
|
|
39
|
-
- Integration and judgment (multi-file coordination, pattern matching): standard tier.
|
|
40
|
-
- Architecture, design, and the final whole-branch review: most capable tier. Always specify the model explicitly; an omitted model inherits the session default.
|
|
41
|
+
## When done → next skill
|
|
41
42
|
|
|
42
|
-
|
|
43
|
+
→ `noir-verifying` for the full integration gate, then `noir-shipping`.
|
|
43
44
|
|
|
44
45
|
## Notes
|
|
45
|
-
-
|
|
46
|
-
- Never start implementation on `main`/`master`; never run multiple implementation subagents in parallel (they conflict on the working tree); never let a subagent self-review replace the actual task review — both are needed.
|
|
47
|
-
- Discipline is observable, not rhetorical: the SDD engine records each per-task review and the final whole-branch review, with the diff packages and verdicts attached.
|
|
46
|
+
- This skill is a playbook — the host decides which tools to use.
|
|
@@ -0,0 +1,50 @@
|
|
|
1
|
+
# Subagent dispatch guide — briefs, independence, review
|
|
2
|
+
|
|
3
|
+
Deep reference for `noir-subagent`. The mechanics of dispatching fresh subagents per task and reviewing between.
|
|
4
|
+
|
|
5
|
+
## The brief — what a subagent needs
|
|
6
|
+
|
|
7
|
+
A subagent starts cold. Its brief must be self-contained:
|
|
8
|
+
|
|
9
|
+
- **Goal** — one sentence: what must be true when done.
|
|
10
|
+
- **Files** — the exact files it may touch (and which it must NOT touch).
|
|
11
|
+
- **Constraints** — the interface contract it must satisfy (signatures, formats).
|
|
12
|
+
- **Acceptance** — how it proves done (tests pass, output shape).
|
|
13
|
+
- **Context pointer** — where to read the relevant spec/plan/background (NOT a dump of everything).
|
|
14
|
+
|
|
15
|
+
A good brief is 5-10 lines. If it needs a paragraph of caveats, the task isn't independent yet.
|
|
16
|
+
|
|
17
|
+
## Independence validation
|
|
18
|
+
|
|
19
|
+
Two tasks are independent only if:
|
|
20
|
+
|
|
21
|
+
- B does not need A's output to start (or you ORDER them: A then B).
|
|
22
|
+
- They do not write the same file in conflicting ways.
|
|
23
|
+
- They do not share mutable state (a DB, a cache, a port).
|
|
24
|
+
|
|
25
|
+
If any dependency exists, you have two choices: run them sequentially (`noir-executing-plans`), or merge them into one task. Dispatching dependent tasks in parallel produces merge hell.
|
|
26
|
+
|
|
27
|
+
## The review step (non-negotiable)
|
|
28
|
+
|
|
29
|
+
After each subagent returns, REVIEW before integrating:
|
|
30
|
+
|
|
31
|
+
1. **Tests pass?** Run the suite. Don't trust the subagent's "all green."
|
|
32
|
+
2. **Spec satisfied?** The output meets the brief's acceptance criteria.
|
|
33
|
+
3. **No collateral?** The subagent didn't touch files outside its brief, didn't leave debug logs, didn't commit garbage.
|
|
34
|
+
|
|
35
|
+
If any item fails, return the report to the subagent with the specific issue. Do NOT rewrite their code yourself — that defeats the isolation.
|
|
36
|
+
|
|
37
|
+
## Files as the contract
|
|
38
|
+
|
|
39
|
+
Subagents communicate through FILES, not conversation. Each task writes:
|
|
40
|
+
- `task-N-brief.md` — what it was asked (you write before dispatch).
|
|
41
|
+
- `task-N-report.md` — what it did + test results (it writes on return).
|
|
42
|
+
|
|
43
|
+
These become the integration record and the `noir-wrap` handoff input. Structured outputs (JSON, CSV, markdown tables) beat prose for machine consumption.
|
|
44
|
+
|
|
45
|
+
## When NOT to dispatch a subagent
|
|
46
|
+
|
|
47
|
+
- The task is a single small change (a subagent adds more overhead than value).
|
|
48
|
+
- The task requires deep conversation context you already hold.
|
|
49
|
+
- The task touches the same file as a parallel task.
|
|
50
|
+
- The model is simpler than the task warrants — match subagent capability to task difficulty.
|
|
@@ -1,18 +1,45 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: noir-sync
|
|
3
|
-
description: Use
|
|
3
|
+
description: Use when starting any session or conversation — load project context and route to the relevant noir skill. Fires on explicit signals (feature start, spec request, new task). Do NOT use mid-session for a status update; use noir-checkpoint.
|
|
4
|
+
metadata:
|
|
5
|
+
category: discovery
|
|
6
|
+
version: 1.0.0
|
|
7
|
+
license: MIT
|
|
8
|
+
compatibility: claude · agents-md · gemini · cursor · opencode
|
|
4
9
|
---
|
|
5
10
|
|
|
6
|
-
|
|
11
|
+
# noir-sync — Session starter and skill router
|
|
12
|
+
|
|
13
|
+
Load the project context AND route to the right skill. This is the entry point for every session — read anchors, check state, recall memory, then hand off to the noir skill that matches the task. Stay silent on trivial edits; fire on explicit signals.
|
|
14
|
+
|
|
15
|
+
## When to use
|
|
16
|
+
|
|
17
|
+
- Start of a session — before any other action.
|
|
18
|
+
- The user opens a new conversation or says "let's work on X."
|
|
19
|
+
- **Do NOT use:** mid-session for progress — use `noir-checkpoint`. Do NOT use for a one-line fix with no design.
|
|
7
20
|
|
|
8
21
|
## Procedure
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
-
|
|
18
|
-
-
|
|
22
|
+
|
|
23
|
+
1. **Read anchors.** `CLAUDE.md`, `AGENTS.md`, `.noir/NOIR.md` — always-on context to ground the session. On Claude Code, use `Read`; on other hosts, use the equivalent file tool.
|
|
24
|
+
2. **Check git state.** `git rev-parse --abbrev-ref HEAD`, `git status --porcelain`, `git log --oneline -5`. Note any dirty tree or in-progress work.
|
|
25
|
+
3. **Check Noir state.** Read `.noir/tasks/` for an in-flight task. If one exists, surface its phase and offer to resume.
|
|
26
|
+
4. **Recall memory.** Query Noir memory (or the host's recall tooling) for 2-4 top facts relevant to this session. If empty, say so — `.noir/` is the durable fallback.
|
|
27
|
+
5. **Skill triage.** Map the user's intent to the right noir skill. These are the high-value skills:
|
|
28
|
+
- Feature/idea → `noir-brainstorming` → `noir-spec` → `noir-planning` → `noir-executing-plans`
|
|
29
|
+
- Bug/crash → `noir-systematic-debugging`
|
|
30
|
+
- Implement with TDD → `noir-test-driven-development`
|
|
31
|
+
- Verify/PR → `noir-verifying`
|
|
32
|
+
- Ship/integrate → `noir-shipping`
|
|
33
|
+
- Close session → `noir-wrap`
|
|
34
|
+
- Code exploration → `noir-exploring`
|
|
35
|
+
- Skill authoring → `noir-writing-skills`
|
|
36
|
+
6. **Print a brief** — project type, stack, key commands, in-flight task. Then hand off to the routed skill.
|
|
37
|
+
|
|
38
|
+
## Notes
|
|
39
|
+
|
|
40
|
+
- Don't auto-start work. Surface the state and let the user confirm direction.
|
|
41
|
+
- If no skill matches, ask what the user wants to do — the route is a suggestion, not a mandate.
|
|
42
|
+
|
|
43
|
+
## When done → next skill
|
|
44
|
+
|
|
45
|
+
Route to the matched skill. Or is there something else you'd like to do first?
|
|
@@ -0,0 +1,71 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: noir-systematic-debugging
|
|
3
|
+
description: Use when encountering any bug, test failure, or unexpected behavior — find the root cause methodically before proposing any fix. Use when the user says "this is broken", "it fails intermittently", or pastes a stack trace. Do NOT use for a known fix where the root cause is already confirmed.
|
|
4
|
+
metadata:
|
|
5
|
+
category: execute
|
|
6
|
+
version: 1.0.0
|
|
7
|
+
license: MIT
|
|
8
|
+
compatibility: claude · agents-md · gemini · cursor · opencode
|
|
9
|
+
references:
|
|
10
|
+
- tracing.md
|
|
11
|
+
---
|
|
12
|
+
|
|
13
|
+
# noir-systematic-debugging
|
|
14
|
+
|
|
15
|
+
Find root cause before attempting any fix. Symptom patches waste time and create new bugs; the path is investigate → hypothesize → test → fix, in that order. Skipping ahead is the most expensive move in debugging.
|
|
16
|
+
|
|
17
|
+
## When to use
|
|
18
|
+
|
|
19
|
+
- A bug, test failure, or unexpected behavior appears — and you don't know the root cause yet.
|
|
20
|
+
- The user says "this is broken", "it fails", "something's wrong", or pastes a stack trace.
|
|
21
|
+
- A test passed yesterday and fails today, or it fails intermittently.
|
|
22
|
+
- **Do NOT use:** when the root cause is already clear and the fix is obvious (one-line patch with no risk of side effects).
|
|
23
|
+
|
|
24
|
+
## Procedure
|
|
25
|
+
|
|
26
|
+
### 1. Reproduce and read
|
|
27
|
+
- Read the error completely. Stack traces, line numbers, file paths, error codes — do not skim past the message; it usually narrows the search immediately.
|
|
28
|
+
- Reproduce reliably. Find the minimal, repeatable trigger. If you cannot reproduce it, gather more data; do not guess at a fix from a single trace.
|
|
29
|
+
- Check recent changes. `git diff` and recent commits often point straight at the cause. Note new dependencies, config edits, and environmental differences.
|
|
30
|
+
|
|
31
|
+
### 2. Gather evidence at component boundaries
|
|
32
|
+
When the system has multiple layers (CLI → adapter → daemon → store, or API → service → database), instrument each boundary before proposing fixes: log what enters and exits, verify config propagation, check state at each layer. On Claude Code, use `Bash` to run the exact failing command with verbose flags; on other hosts, use whatever shell/execution tool is available. Run once, read the output, then investigate the specific layer the evidence implicates.
|
|
33
|
+
|
|
34
|
+
### 3. Trace data flow to the source
|
|
35
|
+
For a bad value deep in a call stack, trace backward: where does it originate, who passed it down, who created it. Fix at the source, not at the symptom. Resist the urge to patch where the symptom surfaces.
|
|
36
|
+
|
|
37
|
+
### 4. Find the pattern
|
|
38
|
+
Locate similar working code in the same codebase and list every difference from the broken path — however small. When applying a known pattern, read the reference implementation completely; partial understanding is a common cause of the bug you are now debugging.
|
|
39
|
+
|
|
40
|
+
### 5. Form one hypothesis and test minimally
|
|
41
|
+
State the hypothesis explicitly ("I think X is the root cause because Y"). Make the smallest possible change that would prove or disprove it — one variable at a time, never a bundle. If the hypothesis survives the minimal test, proceed; if not, form a new hypothesis with the new information. Do not stack fixes on top of a failed one.
|
|
42
|
+
|
|
43
|
+
### 6. Fix at the root, with a regression test
|
|
44
|
+
- Write the simplest failing test that reproduces the bug (see `noir-test-driven-development`).
|
|
45
|
+
- Implement the single fix at the root cause.
|
|
46
|
+
- Verify the test goes green and that no other test regressed.
|
|
47
|
+
- One change at a time; no "while I'm here" refactors bundled into the fix.
|
|
48
|
+
|
|
49
|
+
### 7. Know when to stop fixing
|
|
50
|
+
If three or more fixes have failed, each revealing a new problem in a different place, the issue is likely architectural, not a missing patch. Stop, name the pattern that is not holding, and surface the architectural question (with evidence) instead of attempting fix number four.
|
|
51
|
+
|
|
52
|
+
## Verification
|
|
53
|
+
|
|
54
|
+
- [ ] The bug is reproduced reliably (the trigger is documented).
|
|
55
|
+
- [ ] A root cause hypothesis is stated explicitly ("X because Y").
|
|
56
|
+
- [ ] The fix is at the source, not the symptom.
|
|
57
|
+
- [ ] A regression test passes and no other tests regressed.
|
|
58
|
+
- [ ] One commit per fix — no bundled refactors.
|
|
59
|
+
|
|
60
|
+
## Notes
|
|
61
|
+
|
|
62
|
+
- "No root cause found" is usually an incomplete investigation. If the bug is genuinely environmental or timing-dependent after a complete pass, document what you investigated, implement appropriate handling (retry, timeout, error message), and add monitoring — but say so plainly rather than implying the root cause is unknowable.
|
|
63
|
+
- If the bug spans multiple services, use `noir-exploring` to fan out the evidence search first.
|
|
64
|
+
|
|
65
|
+
## Reference
|
|
66
|
+
|
|
67
|
+
For deeper detail, see [tracing.md](references/tracing.md).
|
|
68
|
+
|
|
69
|
+
## When done → next skill
|
|
70
|
+
|
|
71
|
+
→ `noir-verifying` to confirm the fix is complete. Or do you need to investigate another bug?
|
|
@@ -0,0 +1,51 @@
|
|
|
1
|
+
# Root-cause tracing — evidence at every boundary
|
|
2
|
+
|
|
3
|
+
Deep reference for `noir-systematic-debugging`. Use this when the bug spans multiple layers or the root cause is not obvious from the first stack trace.
|
|
4
|
+
|
|
5
|
+
## Why boundaries matter
|
|
6
|
+
|
|
7
|
+
A bug is rarely where it surfaces. "The API returns 500" could mean the route handler, the service, the DB layer, or the config. If you fix at the surface, you patch a symptom; the same root cause fails again through another path. Tracing evidence at each boundary converts a guess into a data-backed narrowing.
|
|
8
|
+
|
|
9
|
+
## The boundary checklist
|
|
10
|
+
|
|
11
|
+
For a request that flows CLI → adapter → daemon → store (Noir's own shape), or API → service → DB, instrument each layer BEFORE proposing a fix:
|
|
12
|
+
|
|
13
|
+
1. **Entry** — what arrived? Log the raw request: method, path, headers, body (sanitized).
|
|
14
|
+
2. **Validation** — did the input pass validation? A rejected payload fails here.
|
|
15
|
+
3. **Auth/context** — was identity resolved? A missing token fails here.
|
|
16
|
+
4. **Service logic** — what did the service compute? Log inputs + outputs at the boundary.
|
|
17
|
+
5. **Persistence** — what SQL/query ran? Log the query + params, then the result/error.
|
|
18
|
+
6. **Exit** — what returned? Compare the response to the handler's intent.
|
|
19
|
+
|
|
20
|
+
## Technique: log at the boundary, not inside
|
|
21
|
+
|
|
22
|
+
Put one log line at each boundary (enter/exit) rather than scattering logs through the middle. The boundary log answers "did X reach layer N correctly?" — the middle log answers "what happened inside layer N?" Start with boundaries; only go inside when a boundary shows the input was wrong.
|
|
23
|
+
|
|
24
|
+
## Technique: binary search the layers
|
|
25
|
+
|
|
26
|
+
Instead of instrumenting all 6 boundaries, bisect: check the MIDDLE layer first. If the middle sees the right input and produces the wrong output, the bug is in the middle or below. If the middle sees wrong input, the bug is above. Each check halves the search space.
|
|
27
|
+
|
|
28
|
+
## Technique: the "did it change?" check
|
|
29
|
+
|
|
30
|
+
The most common root cause is a recent change. Before tracing:
|
|
31
|
+
- `git log --oneline -10` — what changed recently?
|
|
32
|
+
- `git diff HEAD~1` — what's different in the code that touches this path?
|
|
33
|
+
- Check config/dependency changes: a new package version, a toggled feature flag, an env var change.
|
|
34
|
+
|
|
35
|
+
## Recording the trace
|
|
36
|
+
|
|
37
|
+
For a debugging session, keep a compact trace log:
|
|
38
|
+
|
|
39
|
+
```
|
|
40
|
+
REQUEST GET /api/v2/task/86eyeryfe?include_subtasks=true
|
|
41
|
+
VALIDATE ok
|
|
42
|
+
AUTH pk_*** (ok)
|
|
43
|
+
SERVICE fetched task, subtasks=[] (NOT null!) ← suspicious: empty array
|
|
44
|
+
FINDING subtasks excluded by default; include_subtasks=true missing
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
The trace makes the evidence explicit and reviewable — the same honesty `noir-verifying` demands.
|
|
48
|
+
|
|
49
|
+
## When the trace points to an architectural issue
|
|
50
|
+
|
|
51
|
+
If 3+ fixes have failed, each revealing a new problem in a different layer, stop fixing and name the pattern that is not holding (e.g. "the daemon caches process.env at spawn, so any env-var-dependent path is stale"). Surface it as an architectural question with the trace as evidence.
|
|
@@ -0,0 +1,73 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: noir-test-driven-development
|
|
3
|
+
description: Use when implementing any feature or bugfix — write the failing test first with the RED→GREEN→REFACTOR loop. Do NOT use for pure refactors where behavior doesn't change.
|
|
4
|
+
metadata:
|
|
5
|
+
category: execute
|
|
6
|
+
version: 1.0.0
|
|
7
|
+
license: MIT
|
|
8
|
+
compatibility: claude · agents-md · gemini · cursor · opencode
|
|
9
|
+
references:
|
|
10
|
+
- tdd-worked-example.md
|
|
11
|
+
---
|
|
12
|
+
|
|
13
|
+
# noir-test-driven-development
|
|
14
|
+
|
|
15
|
+
Every feature starts with a failing test. The RED → GREEN → REFACTOR loop is the proof: if there's no failing test, there's no evidence the feature is needed. This skill absorbs `noir-test` — test design guidance lives here too.
|
|
16
|
+
|
|
17
|
+
## When to use
|
|
18
|
+
|
|
19
|
+
- You're about to implement a feature or fix a bug.
|
|
20
|
+
- The user says "write a test", "test this", or "TDD this".
|
|
21
|
+
- The implementation is non-trivial — a behavior change, new logic, or API contract.
|
|
22
|
+
- **Do NOT use:** for a pure rename, formatting change, or refactor where the existing test suite already covers the behavior.
|
|
23
|
+
|
|
24
|
+
## Procedure
|
|
25
|
+
|
|
26
|
+
### RED — Write the failing test
|
|
27
|
+
1. **Write the test FIRST.** Before any implementation code. The test must fail for the RIGHT reason (the feature is missing), not because of a syntax error or a broken import.
|
|
28
|
+
2. **The test is the spec.** Write it with the behavior you want to see — the input, the expected output, the edge cases. A good test reads like a user story: "given X, when Y, then Z."
|
|
29
|
+
3. **Run the test to confirm it fails.** The failure message must be clear — "feature not yet implemented", not "undefined is not a function."
|
|
30
|
+
|
|
31
|
+
### Verify RED — confirm it fails for the right reason
|
|
32
|
+
- The test fails on the assertion you wrote, not on setup/import/boot.
|
|
33
|
+
- If it fails for the wrong reason, fix the test (not the code under test) until it fails cleanly.
|
|
34
|
+
|
|
35
|
+
### GREEN — Implement the minimal code
|
|
36
|
+
- Write the smallest amount of production code that makes the test pass. Nothing more — no "future-proofing," no "while I'm here."
|
|
37
|
+
- Run the test. It must pass. Run the full suite to confirm no regression.
|
|
38
|
+
|
|
39
|
+
### Verify GREEN — confirm it passes legitimately
|
|
40
|
+
- The test passes on the assertion you wrote — not on a coincidental edge case.
|
|
41
|
+
- No other test broke. If one did, fix it NOW; do not carry a broken suite.
|
|
42
|
+
|
|
43
|
+
### REFACTOR — Clean up
|
|
44
|
+
- With a green suite as the safety net, clean up: remove duplication, improve naming, simplify logic. The test suite protects you — a broken refactor fails immediately.
|
|
45
|
+
- Run the full suite after every refactor step.
|
|
46
|
+
|
|
47
|
+
### Repeat
|
|
48
|
+
- For the next behavior, go back to RED. One test → one implementation → one refactor per cycle.
|
|
49
|
+
|
|
50
|
+
## Why order matters
|
|
51
|
+
Writing the test first forces you to state the desired behavior before you're influenced by the implementation. Writing the implementation to the test keeps you from building more than needed. Refactoring last, with a green suite, makes cleanup safe.
|
|
52
|
+
|
|
53
|
+
## Verification
|
|
54
|
+
|
|
55
|
+
- [ ] A failing test was written BEFORE the implementation.
|
|
56
|
+
- [ ] The test failed for the right reason (feature missing, not broken import).
|
|
57
|
+
- [ ] The minimal implementation made the test pass.
|
|
58
|
+
- [ ] The full test suite is green after the change.
|
|
59
|
+
- [ ] Refactored code is cleaner, not just different.
|
|
60
|
+
|
|
61
|
+
## Notes
|
|
62
|
+
|
|
63
|
+
- Don't test the framework. Test YOUR logic.
|
|
64
|
+
- One assertion per behavior, one test per behavior. A test that asserts five things is five tests fighting for attention.
|
|
65
|
+
- If TDD feels slow, you're probably fixing a bug that someone else shipped because they skipped it.
|
|
66
|
+
|
|
67
|
+
## Reference
|
|
68
|
+
|
|
69
|
+
For a worked RED → GREEN → REFACTOR walkthrough, see [tdd-worked-example.md](references/tdd-worked-example.md).
|
|
70
|
+
|
|
71
|
+
## When done → next skill
|
|
72
|
+
|
|
73
|
+
→ `noir-verifying` to gather evidence the work is complete. Or is there another behavior to implement?
|
|
@@ -0,0 +1,65 @@
|
|
|
1
|
+
# TDD worked example — RED → GREEN → REFACTOR
|
|
2
|
+
|
|
3
|
+
Deep reference for `noir-test-driven-development`. A complete walkthrough of the loop on a real-shaped problem.
|
|
4
|
+
|
|
5
|
+
## The problem
|
|
6
|
+
|
|
7
|
+
Add a function that counts words in a string, ignoring punctuation and case.
|
|
8
|
+
|
|
9
|
+
## RED — write the failing test first
|
|
10
|
+
|
|
11
|
+
```ts
|
|
12
|
+
import { describe, expect, it } from 'vitest';
|
|
13
|
+
import { countWords } from './count-words.js';
|
|
14
|
+
|
|
15
|
+
describe('countWords', () => {
|
|
16
|
+
it('counts simple words', () => {
|
|
17
|
+
expect(countWords('hello world')).toBe(2);
|
|
18
|
+
});
|
|
19
|
+
it('ignores punctuation and is case-insensitive', () => {
|
|
20
|
+
expect(countWords('Hello, WORLD!')).toBe(2);
|
|
21
|
+
});
|
|
22
|
+
it('returns 0 for empty/whitespace input', () => {
|
|
23
|
+
expect(countWords(' ')).toBe(0);
|
|
24
|
+
});
|
|
25
|
+
});
|
|
26
|
+
```
|
|
27
|
+
|
|
28
|
+
Run it. Expected: FAIL with `Cannot find module './count-words.js'` or `countWords is not a function` — the failure is because the feature doesn't exist, which is the RIGHT reason.
|
|
29
|
+
|
|
30
|
+
**Verify RED:** the failure is on your assertion / missing module, NOT on a broken import in the test itself.
|
|
31
|
+
|
|
32
|
+
## GREEN — minimal implementation
|
|
33
|
+
|
|
34
|
+
```ts
|
|
35
|
+
export function countWords(text: string): number {
|
|
36
|
+
const words = text.toLowerCase().match(/[a-z]+/g);
|
|
37
|
+
return words ? words.length : 0;
|
|
38
|
+
}
|
|
39
|
+
```
|
|
40
|
+
|
|
41
|
+
Run the test. PASS. Run the full suite — no regression.
|
|
42
|
+
|
|
43
|
+
**Verify GREEN:** it passes for the right reason (the regex matches words), not a coincidental edge case.
|
|
44
|
+
|
|
45
|
+
## REFACTOR — clean up with the green suite as a net
|
|
46
|
+
|
|
47
|
+
```ts
|
|
48
|
+
const WORD = /[a-z]+/g;
|
|
49
|
+
export function countWords(text: string): number {
|
|
50
|
+
return text.toLowerCase().match(WORD)?.length ?? 0;
|
|
51
|
+
}
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
Run the suite again. Still green. The refactor is safe because the tests prove behavior.
|
|
55
|
+
|
|
56
|
+
## Repeat
|
|
57
|
+
|
|
58
|
+
Add the next behavior → back to RED. One test → one implementation → one refactor per cycle. Never skip RED: a feature without a failing test has no evidence it was needed.
|
|
59
|
+
|
|
60
|
+
## Anti-patterns this example guards against
|
|
61
|
+
|
|
62
|
+
- Writing the implementation before the test (no RED) — you can't prove the feature is needed.
|
|
63
|
+
- A test that passes before the implementation — you're not driving new behavior.
|
|
64
|
+
- Bundling a refactor into the GREEN step — the refactor belongs AFTER green.
|
|
65
|
+
- One test asserting five behaviors — split into five tests, each with one clear failure.
|
|
@@ -0,0 +1,40 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: noir-verifying
|
|
3
|
+
description: Use when about to claim work is complete — gather evidence (tests, gate, criteria) before asserting success. Use when the user says "is this done?" or "verify this". Do NOT use mid-implementation for a progress check; use noir-checkpoint.
|
|
4
|
+
metadata:
|
|
5
|
+
category: verify
|
|
6
|
+
version: 1.0.0
|
|
7
|
+
license: MIT
|
|
8
|
+
compatibility: claude · agents-md · gemini · cursor · opencode
|
|
9
|
+
references:
|
|
10
|
+
- verification-checklist.md
|
|
11
|
+
---
|
|
12
|
+
|
|
13
|
+
# noir-verifying
|
|
14
|
+
|
|
15
|
+
Gather evidence before asserting success. Absorbs `noir-verify` and `noir-review` into one gate: run, read, verify, then claim done.
|
|
16
|
+
|
|
17
|
+
## When to use
|
|
18
|
+
|
|
19
|
+
- About to claim a task, feature, or fix is complete.
|
|
20
|
+
- The user says "verify this", "is this done?", "check everything."
|
|
21
|
+
- **Do NOT use:** mid-implementation — use `noir-checkpoint`.
|
|
22
|
+
|
|
23
|
+
## Procedure
|
|
24
|
+
|
|
25
|
+
1. **IDENTIFY criteria.** Load spec/plan — each acceptance criterion is a verification item.
|
|
26
|
+
2. **RUN the gate.** Execute the project's test command. The output IS the evidence.
|
|
27
|
+
3. **READ results.** A passing suite is the minimum; verify each criterion explicitly.
|
|
28
|
+
4. **Check side effects.** New lint warnings? Type errors? Doc links? Run the full gate.
|
|
29
|
+
5. **ONLY THEN claim done.** State "verified: <evidence>" — never "verified: looks good."
|
|
30
|
+
|
|
31
|
+
## Reference
|
|
32
|
+
|
|
33
|
+
For the full claim-done + PR-review checklist, see [verification-checklist.md](references/verification-checklist.md).
|
|
34
|
+
|
|
35
|
+
## When done → next skill
|
|
36
|
+
|
|
37
|
+
→ `noir-shipping` to commit and integrate, or `noir-wrap` to close. Or something else?
|
|
38
|
+
|
|
39
|
+
## Notes
|
|
40
|
+
- This skill is a playbook — the host decides which tools to use. On Claude Code, prefer `AskUserQuestion` for choices; on other hosts, ask in text.
|
|
@@ -0,0 +1,29 @@
|
|
|
1
|
+
# Verification checklist — evidence before "done"
|
|
2
|
+
|
|
3
|
+
Deep reference for `noir-verifying`. A copyable checklist for the two moments it covers: claiming a task done, and reviewing a PR.
|
|
4
|
+
|
|
5
|
+
## The claim-done gate
|
|
6
|
+
|
|
7
|
+
Before asserting "this is done," satisfy EVERY item with evidence:
|
|
8
|
+
|
|
9
|
+
- [ ] **Spec criteria met** — every acceptance criterion from the spec/plan is citable to an actual behavior, test, or output. Not "looks implemented" — point to the evidence.
|
|
10
|
+
- [ ] **Tests green** — the full suite passes. The OUTPUT was read, not assumed. Quote the result line.
|
|
11
|
+
- [ ] **No regression** — tests that passed before still pass. A green suite isn't enough if you deleted a test to make it green.
|
|
12
|
+
- [ ] **Lint/typecheck** — the project's static checks pass (e.g. `pnpm lint`, `pnpm typecheck`).
|
|
13
|
+
- [ ] **Side effects checked** — docs, generated files, or config that the change touches reflect the new reality (docs reflect shipped state, never stale).
|
|
14
|
+
- [ ] **The claim is evidence-backed** — "verified: <test output / command result>" not "verified: looks good."
|
|
15
|
+
|
|
16
|
+
## The PR-review gate
|
|
17
|
+
|
|
18
|
+
Beyond the claim-done gate, review asks "does this make sense?"
|
|
19
|
+
|
|
20
|
+
- [ ] **Approach** — is this the right way to solve it, or a workaround?
|
|
21
|
+
- [ ] **Naming** — do names say what things are?
|
|
22
|
+
- [ ] **Scope** — does the PR do one thing, or did "while I'm here" changes creep in?
|
|
23
|
+
- [ ] **Security** — user input, auth, secrets: any exposure?
|
|
24
|
+
- [ ] **Tests** — do the tests test the behavior, not the implementation detail?
|
|
25
|
+
- [ ] **Diff hygiene** — no commented-out code, no debug logs, no accidental files.
|
|
26
|
+
|
|
27
|
+
## The honesty rule
|
|
28
|
+
|
|
29
|
+
**"Seems right" is never sufficient.** State the evidence and let it stand on its own. If you cannot produce evidence for an item, the item is NOT done — say so, rather than asserting completion. The same discipline that keeps a debugging trace honest keeps a "done" claim honest.
|
|
@@ -1,12 +1,31 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: noir-worktree
|
|
3
|
-
description: Use when
|
|
3
|
+
description: Use when creating an isolated git workspace for feature work — keeping the main checkout clean. Use when the user says "worktree" or "isolate this work".
|
|
4
|
+
metadata:
|
|
5
|
+
category: git
|
|
6
|
+
version: 1.0.0
|
|
7
|
+
license: MIT
|
|
8
|
+
compatibility: claude · agents-md · gemini · cursor · opencode
|
|
4
9
|
---
|
|
5
10
|
|
|
6
11
|
# noir-worktree
|
|
7
12
|
|
|
8
|
-
> **Stub:** this skill ships as a valid, loadable placeholder in S5; its full playbook is deepened in a later slice.
|
|
9
13
|
|
|
10
|
-
|
|
14
|
+
## When to use
|
|
15
|
+
- When the user triggers this skill.
|
|
11
16
|
|
|
12
|
-
|
|
17
|
+
Isolate feature work from the current workspace via git worktrees. One feature, one directory, no cross-contamination.
|
|
18
|
+
|
|
19
|
+
## Procedure
|
|
20
|
+
|
|
21
|
+
1. **Confirm the feature needs isolation.** A single-file fix in a clean repo doesn't need a worktree. A multi-file feature with branches and dependencies does.
|
|
22
|
+
2. **Create the worktree.** `git worktree add -b <feature-branch> <path> <base-branch>`. On Claude Code, if a `.claude/worktrees/` pattern exists, follow it.
|
|
23
|
+
3. **Work in the isolated directory.** The worktree has its own checkout, and the original directory is untouched.
|
|
24
|
+
4. **When done, clean up.** `git worktree remove <path>` + `git branch -d <feature-branch>` (after merge).
|
|
25
|
+
|
|
26
|
+
## When done → next skill
|
|
27
|
+
|
|
28
|
+
→ `noir-shipping` to integrate the finished work. Or continue in the worktree.
|
|
29
|
+
|
|
30
|
+
## Notes
|
|
31
|
+
- This skill is a playbook — the host decides which tools to use. On Claude Code, prefer `AskUserQuestion` for choices; on other hosts, ask in text.
|
|
@@ -1,23 +1,39 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: noir-wrap
|
|
3
|
-
description: Use when closing a session cleanly —
|
|
3
|
+
description: Use when closing a work session cleanly — running final verification, updating docs and CHANGELOG, saving memory, and emitting a host handoff. Do NOT use mid-session for a checkpoint — use noir-checkpoint.
|
|
4
|
+
metadata:
|
|
5
|
+
category: document
|
|
6
|
+
version: 1.0.0
|
|
7
|
+
license: MIT
|
|
8
|
+
compatibility: claude · agents-md · gemini · cursor · opencode
|
|
4
9
|
---
|
|
5
10
|
|
|
6
11
|
# noir-wrap
|
|
7
12
|
|
|
8
|
-
|
|
13
|
+
Close a session in a clean, recoverable state. This skill absorbs `noir-document` (update docs/CHANGELOG/memory) — wrap is the superset that covers the full close.
|
|
9
14
|
|
|
10
|
-
##
|
|
15
|
+
## When to use
|
|
11
16
|
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
4. Save durable memory before closing: observations, decisions, patterns the next session should recall. Prefer `noir memory save` (or the `noir.remember` MCP tool from the host) so cross-session recall works.
|
|
16
|
-
5. Advance the workflow task if a gate is satisfied: `noir task advance --to <phase>` (the verify gate prints the handoff hint automatically).
|
|
17
|
-
6. Emit the host handoff — run `noir handoff` (or the session-end alias `noir wrap`). This prints a ready-to-paste markdown prompt to STDOUT that names the active task, the next gate's skill, a bounded context/memory seed, and the exact host-launch directive. Pipe it straight into the host, or persist with `noir handoff --write` (the path is gitignored under `.noir/handoff/`).
|
|
17
|
+
- Ending a work session.
|
|
18
|
+
- The user says "wrap up", "I'm done", "close this out."
|
|
19
|
+
- **Do NOT use:** mid-session — use `noir-checkpoint`.
|
|
18
20
|
|
|
19
|
-
##
|
|
21
|
+
## Procedure
|
|
20
22
|
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
23
|
+
1. **Run final verification.** Same gate as `noir-verifying`: test suite, lint, typecheck. Evidence, not assumption.
|
|
24
|
+
2. **Update docs.** CHANGELOG, ADRs, decisions, reference docs — anything that should reflect the session's work. The rule: docs reflect shipped reality, never a stale plan.
|
|
25
|
+
3. **Save memory.** Persist observations, decisions, patterns the next session should recall. `noir memory save` (or `noir.remember` MCP tool).
|
|
26
|
+
4. **Confirm commits.** Commits are made and intentional (local or pushed). Noir defaults to local.
|
|
27
|
+
5. **Advance the workflow task.** `noir task advance --to <phase>` if a gate is satisfied.
|
|
28
|
+
6. **Emit the handoff.** `noir handoff` (text-only prompt; `--write` persists to `.noir/handoff/`; `--json` for CI). Names the active task, next gate's skill, and the host-launch directive.
|
|
29
|
+
|
|
30
|
+
## Verification
|
|
31
|
+
|
|
32
|
+
- [ ] Gate is green (tests, lint, typecheck).
|
|
33
|
+
- [ ] Docs are synced (CHANGELOG, decisions, references).
|
|
34
|
+
- [ ] Memory is saved (key observations + decisions).
|
|
35
|
+
- [ ] Handoff emitted (the next session can resume).
|
|
36
|
+
|
|
37
|
+
## When done → next skill
|
|
38
|
+
|
|
39
|
+
The session is closed. Until next time.
|