pi-gauntlet 5.15.0 → 5.16.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,13 @@
1
1
  # Changelog
2
2
 
3
+ ## v5.16.1 - 2026-09-20
4
+
5
+ - gauntlet-handoff follows the shipped pi-cohort 7.1.0 `handoff` skill: a skill cannot expand `/skill:handoff`, so step 2 reads cohort's `skills/handoff/SKILL.md` from the session skill list and follows its procedure; the brief path comes from cohort's `Handoff written:` report line (`Handoff not written:` is a STOP) instead of being recomputed; the minimum is stated as pi-cohort >= 7.1.0 in README and the brief contract. (#40)
6
+
7
+ ## v5.16.0 - 2026-09-20
8
+
9
+ - Skills never name a provider or model: `subagent-driven-development` drops its `## Model Selection` tier table for a one-place `## Model` rule (a dispatch's `model:` is omitted, or carries a `gauntlet_setting` value, a user-named model, or the main loop's own string); `doc/configuration.md` gains `### Dispatch model precedence`; `scripts/model-literal-lint.mjs` bans provider/model literals in `skills/`, `agents/`, `extensions/` from `scripts/ci.mjs`. `writing-skills` widens its trigger to personas and prompt templates and adds `## Authoring rules` (imperative voice, low conditionality, minimal diff, oversized-skill extraction); AGENTS core `v6` makes it binding for every skill, persona, and prompt edit. `shape-ticket` and `conformance-check.md` name the main-loop-string provenance explicitly. (#42)
10
+
3
11
  ## v5.15.0 - 2026-09-20
4
12
 
5
13
  - New human-only `/skill:gauntlet-handoff`: invokes pi-cohort's `handoff` skill (`--out <path>` or `--key <name>`; default key = the run worktree's branch, `/` flattened to `-`, file `<tmpdir>/pi-handoff/<key>.md`) and appends the gauntlet `## Process state` section in the layout fixed by `skills/gauntlet-resume/reference/brief-contract.md` (`## Producers`); stops instead of writing when the skill is absent or the installed cohort still writes that section itself. `/skill:gauntlet-resume` with no arguments lists the briefs under `<tmpdir>/pi-handoff/` for a human pick, and a token that looks like a path is always a brief file (missing -> stop, never a scan). `scripts/ci.mjs` gains a drift lint (`scripts/brief-contract-lint.mjs`): the three process-state grammar lines may appear only in the contract file, both skills must cite it, and gauntlet-handoff may not carry cohort's core headings. Requires pi-cohort >= the release shipping [pi-cohort #18](https://github.com/jjuraszek/pi-cohort/issues/18) (exact version filled at release). (#40)
package/README.md CHANGED
@@ -79,11 +79,11 @@ pi-gauntlet is **opinionated**: every non-trivial change is *meant* to ride this
79
79
 
80
80
  ## Handoff and resume
81
81
 
82
- `/skill:gauntlet-handoff` and `/skill:gauntlet-resume` are a pair: cohort's `handoff` skill writes the six flow-agnostic headings, gauntlet-handoff appends `## Process state`, and gauntlet-resume reads the whole brief back. Both gauntlet skills read one grammar file, `skills/gauntlet-resume/reference/brief-contract.md`; the repo validator (`scripts/ci.mjs`) fails if a grammar line appears anywhere else under `skills/`. Requires the pi-cohort release that ships [pi-cohort #18](https://github.com/jjuraszek/pi-cohort/issues/18) (the `handoff` skill); on an older pi-cohort, gauntlet-handoff stops before writing.
82
+ `/skill:gauntlet-handoff` and `/skill:gauntlet-resume` are a pair: cohort's `handoff` skill writes the six flow-agnostic headings, gauntlet-handoff appends `## Process state`, and gauntlet-resume reads the whole brief back. Both gauntlet skills read one grammar file, `skills/gauntlet-resume/reference/brief-contract.md`; the repo validator (`scripts/ci.mjs`) fails if a grammar line appears anywhere else under `skills/`. Requires pi-cohort >= 7.1.0 (the `handoff` skill, [pi-cohort #18](https://github.com/jjuraszek/pi-cohort/issues/18)); on an older pi-cohort, gauntlet-handoff stops before writing.
83
83
 
84
84
  ### Smoke walkthrough (release-gated)
85
85
 
86
- Run by a human against the pi-cohort release that ships #18, before a pi-gauntlet release claims the pair works; record the outcome in the release commit body.
86
+ Run by a human against pi-cohort >= 7.1.0, before a pi-gauntlet release claims the pair works; record the outcome in the release commit body.
87
87
 
88
88
  1. Implement phase with tasks `complete`/`in_progress`/`pending` in a `.worktrees/<branch>` flow, session in the primary checkout: `/skill:gauntlet-handoff` writes `<tmpdir>/pi-handoff/<branch>.md` with the six core headings then `## Process state` last; a fresh session running `/skill:gauntlet-resume <path>` restores implement with the three statuses and ends `Gate history not restored; re-validating <task> before any stage advance`.
89
89
  2. Plan phase, `No plan active.`: the brief keeps that line and `Active task: none`; resume restores phase-only with no `plan_tracker init`.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-gauntlet",
3
- "version": "5.15.0",
3
+ "version": "5.16.1",
4
4
  "description": "Opinionated, gated workflow skills, subagent personas, and runtime extensions for the pi coding agent.",
5
5
  "author": "Jacek Juraszek",
6
6
  "type": "module",
@@ -34,7 +34,7 @@ loop: tracker state is extension state no subagent can read.
34
34
  - `--out` and `--key` both given -> STOP: they are exclusive in cohort's grammar.
35
35
  - `--out <path>` -> resolve to an absolute path against the session cwd
36
36
  (`node -p "require('path').resolve(process.argv[1])" -- "<path>"`); that is the
37
- expected path and the form passed to cohort.
37
+ form passed to cohort.
38
38
  - Otherwise the key is `--key <name>` when given, else derived from the **run
39
39
  worktree**: the path a flow skill reported as `Worktree ready at <path>` in this
40
40
  transcript (the same evidence cohort's snapshot uses):
@@ -54,19 +54,22 @@ loop: tracker state is extension state no subagent can read.
54
54
  `--out` (a default-branch key would collide across flows).
55
55
  - A derived key has every `/` replaced by `-` (`hotfix/x` -> `hotfix-x`) so the file
56
56
  lands flat in the default directory. A human-given `--key` is passed verbatim.
57
- - Expected path = `<tmpdir>/pi-handoff/<key>.md`, `<tmpdir>` from
58
- `node -p "require('os').tmpdir()"`. This is the one place this package restates
59
- cohort's default-path rule; step 3 verifies it.
60
- 2. **Invoke cohort.** `/skill:handoff --out <absolute path>` or `/skill:handoff --key <key>`,
61
- exactly as resolved in step 1. The skill must be present in the session skill list
62
- under `name: handoff`. Absent -> STOP: "cohort `handoff` skill not in the session
63
- skill list - install or upgrade pi-cohort to the release that ships pi-cohort #18".
64
- No version lookup, no fallback to the `/handoff` prompt.
65
- 3. **Verify the file.** Expected path missing, or line 1 not starting `# Handoff:` ->
66
- STOP with the expected path.
57
+ - With `--key`, cohort writes `<tmpdir>/pi-handoff/<key>.md`, `<tmpdir>` from
58
+ `node -p "require('os').tmpdir()"`; the authoritative path is the one cohort
59
+ reports in step 2.
60
+ 2. **Run cohort's handoff procedure.** A skill cannot expand `/skill:handoff` (pi expands
61
+ skill commands on typed input only). Find `handoff` in the session skill list, `Read`
62
+ its `SKILL.md` at the location listed there, and follow its procedure with the output
63
+ option resolved in step 1 - the brief's core (six headings, ending at `## Skills loaded`)
64
+ is cohort's to write, by cohort's rules. Absent from the skill list -> STOP: "cohort
65
+ `handoff` skill not in the session skill list - install or upgrade to pi-cohort >= 7.1.0".
66
+ No fallback to the `/handoff` prompt.
67
+ 3. **Take the path from the report.** Cohort's procedure ends with `Handoff written: <abs
68
+ path>` or `Handoff not written: <reason>`. `Handoff not written` -> STOP with that reason.
69
+ Otherwise that path is the brief; line 1 not starting `# Handoff:` -> STOP with the path.
67
70
  4. **Refuse a double section.** The file already has a line matching
68
71
  `^## Process state\s*$` -> STOP: "the installed pi-cohort still writes process state
69
- itself - upgrade to the release that ships #18".
72
+ itself - upgrade to pi-cohort >= 7.1.0".
70
73
  5. **Hotfix exclusion.** `skills/chase-bug/hotfix.md` is in this session's context ->
71
74
  append nothing; report a plain hotfix handoff and go to step 7. Checked before the
72
75
  trackers: chase-bug never touches tracker state, and a stale armed flow appended here
@@ -78,7 +81,7 @@ loop: tracker state is extension state no subagent can read.
78
81
  `reference/brief-contract.md` lays it out: both status outputs verbatim
79
82
  (`No plan active.` verbatim when there is no plan), the active-task line naming
80
83
  the first `→` task else `none`, the gate-history line. Append with a single
81
- `cat >> "<expected path>" <<'EOF' ... EOF` whose body is that block.
84
+ `cat >> "<brief path>" <<'EOF' ... EOF` whose body is that block.
82
85
  - Re-read the file tail and confirm it ends with the gate-history line from the
83
86
  contract and nothing after.
84
87
  7. **Report.** Print the path and `/skill:gauntlet-resume <path>`. When a section was
@@ -91,8 +94,9 @@ loop: tracker state is extension state no subagent can read.
91
94
 
92
95
  ## Red flags - STOP
93
96
 
94
- - Writing the brief yourself instead of invoking `/skill:handoff` (the core headings are
95
- cohort's; this skill appends one section).
97
+ - Writing the brief's core yourself instead of following cohort's `handoff` procedure (the
98
+ six core headings are cohort's; this skill appends one section).
99
+ - Recomputing the brief path instead of taking it from `Handoff written:`.
96
100
  - Restating the section layout here instead of reading `reference/brief-contract.md`.
97
101
  - Appending when step 4 or step 5 fired.
98
102
  - Deriving the key from the primary checkout's branch while a run worktree exists.
@@ -2,7 +2,7 @@
2
2
 
3
3
  Consumed by `gauntlet-resume/SKILL.md` (consumer) and `gauntlet-handoff/SKILL.md`
4
4
  (producer). Grammar lives only here; both skills cite this file and inline none of it
5
- (`scripts/ci.mjs` drift lint). Coupled to pi-cohort's `handoff` skill (pi-cohort #18)
5
+ (`scripts/ci.mjs` drift lint). Coupled to pi-cohort's `handoff` skill (pi-cohort >= 7.1.0, #18)
6
6
  for exactly six headings - `# Handoff:`, `## Intent`, `## Repo state`, `## Decisions`,
7
7
  `## Open questions`, `## Skills loaded` - and the `## Repo state` fields. `## Process
8
8
  state` grammar and its consumer rules are owned here. Cohort drift is reconciled here
@@ -176,7 +176,7 @@ Inline council dispatch, reusing spec-council config and personas - **not** `/sk
176
176
 
177
177
  1. Resolve `gauntlet_setting({ key: "specCouncil" })` when the tool exists. Verdict `council` -> dispatch `spec-council-member`s in parallel plus a `spec-council-synthesizer` chair. Verdict `worker` (or empty members) -> one fresh `worker` critique, its task text carrying the absolute path to `reference/ticket-wording.md` (resolved against this skill's own directory). Malformed config -> one warning line, then branch on verdict.
178
178
  2. **Dispatch shape**, mirroring `/skill:roasting-the-spec`: write the draft body and the source snapshot (original ticket + comments, or the create-mode inputs) to absolute temp files under `mktemp -d`; delimit untrusted snapshots as data. The raw ask (create mode) or the original ticket body (repair mode) is also passed as the `Human input (verbatim; off-limits for over-spec)` block. When a split is proposed, the draft artifact holds all N proposed bodies plus their three-line justification blocks (see Split rule) in one file, not a single body. Two separate calls - never fuse members and chair into one chain (a fused chain lets one member failure kill the roast before the chair runs). Call 1: one member fanout with `cwd` = repo root, absolute `output` paths per member, run-level `control: { needsAttentionAfterMs: 60000, inFlightSilenceCeilingMs: 240000, inFlightSilenceKillMs: 300000 }` (sits beside `tasks`, not inside each task; effective silence-kill max(300s, 240+60) = 300s - record all three fields verbatim so a pi-cohort default change cannot stretch the kill). Then probe the member output files on disk with item 7's usable test. Call 2: the chair, with the usable member files via `reads`, the same control block (`:low` chair turns are short), and task text that (a) forbids repository access - member disagreement on a fact is reported in the synthesis, never verified against the repo - and (b) states coverage: `Coverage: N of M members reported; <slug>: <reason>` (pi-cohort's kill diagnostic when present, else "no output produced"; omit reasons at full coverage; singular wording when one member reported). Member task text: *the draft at `<path>` is the artifact under review; this ticket brief supersedes your spec-axis template - emit the same findings format against the draft; content-only review: the temp files plus the two referenced reference paths (split-axes, ticket-wording) are the entire permitted input - do not read, search, or scan the repository; do not edit any file.* Include the absolute paths to `reference/split-axes.md` and `reference/ticket-wording.md` (both resolved against this skill's own directory) in each member's task text - members run with `cwd` = the consumer repo, where a package-relative path does not resolve.
179
- 3. **Effort: cheap by default.** Append a `:low` thinking suffix to each member's model string at dispatch (this beats the persona's frontmatter `xhigh` pin). Same for the chair: a configured chair string gets any existing suffix replaced with `:low`; an unconfigured chair is dispatched as the parent's model with `:low` appended. The `worker` fallback carries no thinking pin - it runs at the preset's default. **Full-roast escape:** the user may request a full roast, dispatching all model strings bare/as-configured, restoring the xhigh pins; a full roast reuses the spec-roast control blocks (members `{ needsAttentionAfterMs: 300000, inFlightSilenceCeilingMs: 300000, inFlightSilenceKillMs: 600000 }`, chair `{ needsAttentionAfterMs: 300000, inFlightSilenceCeilingMs: 600000, inFlightSilenceKillMs: 900000 }`) - the 5-minute figures in item 2 are `:low`-only.
179
+ 3. **Effort: cheap by default.** Append a `:low` thinking suffix to each member's model string at dispatch (this beats the persona's frontmatter `xhigh` pin). Same for the chair: a configured chair string gets any existing suffix replaced with `:low`; an unconfigured chair is dispatched with the main loop's own string, read from `$PI_PROVIDER`/`$PI_MODEL` with the bash tool, plus `:low`. The `worker` fallback carries no thinking pin - it runs at the preset's default. **Full-roast escape:** the user may request a full roast, dispatching all model strings bare/as-configured, restoring the xhigh pins; a full roast reuses the spec-roast control blocks (members `{ needsAttentionAfterMs: 300000, inFlightSilenceCeilingMs: 300000, inFlightSilenceKillMs: 600000 }`, chair `{ needsAttentionAfterMs: 300000, inFlightSilenceCeilingMs: 600000, inFlightSilenceKillMs: 900000 }`) - the 5-minute figures in item 2 are `:low`-only.
180
180
  4. **Brief covers three axes**, absorbing the fidelity-review role without a new persona: *fidelity* - compare draft against source intent (original ticket + comments in repair; prompt + answers in create), flag `lost` / `added` / `gap` (unpacking existing claims to satisfy the ticket wording contract is not `added`; contract-conformance findings on Context/Problem/Idea outrank fidelity flags that only object to extra explanation of the same claims); and *quality* - problem framing, AC integrity beyond the deterministic gate, scope, wording, and conformance to the ticket wording contract (reference path provided in every roast brief); and *split soundness* - if the draft proposes a split, test each slice against the split-axes reference (path provided in the task text); an architecture-shaped boundary is reported as a finding line containing the marker `split-axis:` (members keep their existing spec-axis findings template; the marker is a substring flag within it, not a new findings kind), e.g. `- [major] split-axis: <slice> - <why> -> merge`. Members may argue toward one ticket, never propose or endorse a split.
181
181
  5. Disposition: unambiguous concrete fixes applied to the draft (one re-pass max); ambiguous findings surfaced at the confirmation gate. Roast edits affect the body draft pre-write only, never posted as a tracker comment, and re-run the deterministic gates (pipeline step 5). Additionally, the parent scans the **usable member output files (item 7's structural test) directly** for lines containing `split-axis:` (substring match), independent of the chair synthesis; any such finding auto-applies a merge - the split is withdrawn and the draft becomes one ticket with phased AC groups, inside the same one-re-pass budget, and the pre-merge N-body draft is kept alongside: a human re-request of the split at the gate re-presents those N bodies old->new as the approval diff (see the Split rule's sticky override). The chair keeps every other axis; clearing a `split-axis:` finding is not on its path. The same directional rule - toward one ticket, never toward a split - binds the `worker` fallback and the runtime conditional (item 6).
182
182
  6. **Runtime conditional (the one allowed):** on a harness with no `gauntlet_setting`/`subagent()` (e.g. Claude Code), dispatch fresh general-purpose subagents via that harness's native facility at low effort, with the same three-axis brief (including the absolute `reference/ticket-wording.md` path) and temp-file artifacts - on such a harness this conditional IS the roast, so the contract path must ride along.
@@ -21,7 +21,7 @@ Your context window holds the full plan, prior decisions, and conversation histo
21
21
 
22
22
  - **No context pollution.** Task N's noise doesn't leak into Task N+1.
23
23
  - **Tighter focus.** Subagent reads less, makes fewer cross-task assumptions, ships smaller diffs.
24
- - **Cheaper at scale.** Smaller models can handle simple subtasks; you only spend top-tier tokens on orchestration and hard problems.
24
+ - **Cheaper at scale.** Operator pins (`subagents.agentOverrides.<agent>.model`) route simple subtasks to smaller models; you spend top-tier tokens on orchestration and hard problems.
25
25
 
26
26
  You are the **orchestrator**. You read the plan, dispatch, review the review, decide. You do **not** write code yourself.
27
27
 
@@ -111,30 +111,13 @@ Every implementer dispatch returns one of:
111
111
 
112
112
  Implementer prompt templates must instruct subagents to return one of these statuses explicitly.
113
113
 
114
- ## Model Selection
114
+ ## Model
115
115
 
116
- Pi-subagents accepts a per-task `model` override. Use it.
117
-
118
- | Task complexity | Model tier | Use cases |
119
- |---|---|---|
120
- | Trivial mechanical change | Cheap | Rename, formatter run, dependency bump, file-move with no edits |
121
- | Standard implementation | Default | Most plan tasks — feature work with tests, refactor with tests |
122
- | Hard / novel / large surface | Most capable | New subsystem, complex algorithm, cross-service contract change |
123
- | Spec review | Default | Reads diff + spec, mechanical comparison |
124
- | Code-quality review | Most capable | Judgment call on naming, design, complexity |
125
- | Conformance / closure | Most capable | Whole-deliverable-vs-origin intent gate (`conformance-reviewer`; model from `gauntlet_setting({ key: "closureReview" }).model`, injected call-site) |
126
-
127
- ```ts
128
- subagent({
129
- agent: "implementer",
130
- async: false,
131
- cwd: "<abs worktree path>",
132
- task: "...",
133
- model: "anthropic/claude-haiku-4" // cheap tier
134
- })
135
- ```
136
-
137
- When in doubt, default. Don't downgrade reviewers — false negatives are expensive.
116
+ - Omit `model:` so the operator's `subagents.agentOverrides.<agent>.model` pin applies, or else the child inherits the main loop.
117
+ - Pass `model:` only a runtime-resolved `gauntlet_setting` value (`closureReview.model`, `escalationLoop.implModel`, `specCouncil.members[]`, `specCouncil.chair`) or a model the user named for that dispatch.
118
+ - Read the main loop's own string from `$PI_PROVIDER`/`$PI_MODEL` with the bash tool only when a skill must bypass a persona's frontmatter `thinking` pin for cost or must pass an explicit fallback after a configured model is unreachable.
119
+ - Omit `model:` when `closureReview.model` or `specCouncil.chair` is `undefined`; pass an `implModel` string verbatim, and stop with a note when `implModel` is `undefined` (Fix-Loop Rounds).
120
+ - Apply suffix edits such as `:low` only to a runtime-resolved string, never to a name the skill wrote; never write a provider or model name, tier tasks by cost, or swap a pinned model for a cheaper or stronger one.
138
121
 
139
122
  ## Dispatch
140
123
 
@@ -45,7 +45,7 @@ The persona ships model-free. Get the model from `gauntlet_setting({ key: "closu
45
45
  If `gauntlet_setting` is unavailable, stop and report - never fall back to a manual bash/JSON
46
46
  settings merge. Inject the model **call-site** on the dispatch (omit `model:` when it is `undefined` to inherit
47
47
  the parent's model) — the same mechanism the spec-council chair uses. If the configured model
48
- is unreachable, retry once with the inherited model. Point it at the strongest reasoning model
48
+ is unreachable, retry once passing the main loop's own string explicitly (read from `$PI_PROVIDER`/`$PI_MODEL`; the phase-tracker guard blocks an omitted `model:` when `closureReview.model` is set and warns on the difference). Point it at the strongest reasoning model
49
49
  the resolved config can reach — this is the last correctness gate. `thinking` stays
50
50
  frontmatter-pinned at `xhigh` and is not call-site overridable, so the config supplies only
51
51
  `model`.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: writing-skills
3
- description: Use when creating new skills, editing existing skills, or verifying skills work before deployment
3
+ description: Use when creating, editing, or refactoring any SKILL.md, its reference/ files, an agent persona (`agents/*.md`, `.pi/agents/*.md`), or a prompt template - including one-line edits - or when such a file exceeds 500 lines, gains if/else branching, or drifts from imperative voice.
4
4
  ---
5
5
 
6
6
  > **Related skills:** Test new skills with `/skill:test-driven-development` discipline. Verify they work with `/skill:verification-before-completion`.
@@ -19,6 +19,15 @@ Write a pressure scenario for a subagent, watch it fail (baseline), write the sk
19
19
 
20
20
  **REQUIRED BACKGROUND:** Read `/skill:test-driven-development` first. This skill applies RED → GREEN → REFACTOR to documentation.
21
21
 
22
+ ## Authoring rules
23
+
24
+ These four rules bind the lines an edit adds or changes in any skill, persona, or prompt-template file. Leave pre-existing text alone (minimal diff). Where `reference/anthropic-best-practices.md` differs (its "Conditional workflow pattern"), this section wins.
25
+
26
+ - **Imperative voice.** Write instructions as commands. Never "should", "consider", "you may want to", "it is recommended".
27
+ - **Low conditionality.** One path per step. Branch only on a runtime fact the agent can observe - a tool result, a file's presence, a settings value - never on the reader's judgment. Two branches are the ceiling; move a third into a table or a `reference/` file.
28
+ - **Minimal diff.** A rule change touches the sentence that owns the rule, not the section. Never restate a rule in a second place; link the owner.
29
+ - **Oversized skill.** A change touching a SKILL.md over 500 lines extracts the concern it touches - or, when that concern is small, the largest self-contained `##` section - into `reference/<topic>.md` or a sibling md, in the same change, before it lands. Keep a one-line "read X now" pointer in the body at the step that needs it.
30
+
22
31
  ## Where Skills Live in Pi
23
32
 
24
33
  Pi discovers skills from multiple roots (see `docs/skills.md` in `@earendil-works/pi-coding-agent` for the full list). The common ones:
@@ -194,7 +203,7 @@ Don't invent capabilities. Don't reference Claude Code's `Task` tool, OpenCode h
194
203
  NO SKILL WITHOUT A FAILING TEST FIRST
195
204
  ```
196
205
 
197
- Applies to new skills AND edits to existing skills.
206
+ Applies to new skills AND edits to existing skills. The RED-GREEN-REFACTOR baseline binds skill bodies - new skills and behavior-changing edits to SKILL.md or reference files; a wording-only edit to a skill, persona, or prompt template is verified by reading the changed lines back against `## Authoring rules`.
198
207
 
199
208
  Wrote a skill before testing it? Delete it. Start over.
200
209
  Edited a skill without testing? Same violation.
@@ -376,7 +385,7 @@ Use `plan_tracker` to create tasks for each item:
376
385
 
377
386
  **Pi-specific**
378
387
  - [ ] If discipline skill *and* the rule is mechanically detectable at a tool boundary: consider an `extensions/*.ts` hook (high bar — runtime hooks are deliberately slim; only add ones that beat false-positive heuristics)
379
- - [ ] If skill has >500 lines: split deep content into `reference/<topic>.md` and instruct the agent to read the specific file inline
388
+ - [ ] If skill has >500 lines, or a skill this change touches does: split deep content into `reference/<topic>.md` and instruct the agent to read the specific file inline
380
389
  - [ ] Skill location: project-scoped lives under `.pi/skills/`; cross-harness skills shared with Claude Code under `.agents/skills/`; reusable workflow skills belong in a package like `pi-gauntlet`
381
390
  - [ ] Update routing: link from `AGENTS.md` if cross-cutting
382
391
 
@@ -409,11 +418,11 @@ Labels should carry semantic meaning.
409
418
  ## Red Flags — STOP
410
419
 
411
420
  - Wrote a skill without running a baseline scenario first
412
- - Edited a skill without re-running the relevant baseline
421
+ - Made a behavior-changing skill edit without re-running the relevant baseline
413
422
  - Description starts with "This skill does…" (summary instead of trigger)
414
423
  - Code-then-test ordering in the skill body
415
424
  - Force-loading other skills with `@`
416
- - SKILL.md over 500 lines with no `reference/` split
425
+ - SKILL.md over 500 lines, or a skill this change touches over 500 lines, with no `reference/` split
417
426
  - Referencing tools that don't exist in pi (Claude Code's `Task`, OpenCode hooks, etc.) for a pi-scope skill
418
427
  - About to ship multiple skills in a batch without testing each
419
428
  - "I'll test it later" — that means never