pi-gauntlet 5.15.0 → 5.16.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,9 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## v5.16.0 - 2026-09-20
|
|
4
|
+
|
|
5
|
+
- Skills never name a provider or model: `subagent-driven-development` drops its `## Model Selection` tier table for a one-place `## Model` rule (a dispatch's `model:` is omitted, or carries a `gauntlet_setting` value, a user-named model, or the main loop's own string); `doc/configuration.md` gains `### Dispatch model precedence`; `scripts/model-literal-lint.mjs` bans provider/model literals in `skills/`, `agents/`, `extensions/` from `scripts/ci.mjs`. `writing-skills` widens its trigger to personas and prompt templates and adds `## Authoring rules` (imperative voice, low conditionality, minimal diff, oversized-skill extraction); AGENTS core `v6` makes it binding for every skill, persona, and prompt edit. `shape-ticket` and `conformance-check.md` name the main-loop-string provenance explicitly. (#42)
|
|
6
|
+
|
|
3
7
|
## v5.15.0 - 2026-09-20
|
|
4
8
|
|
|
5
9
|
- New human-only `/skill:gauntlet-handoff`: invokes pi-cohort's `handoff` skill (`--out <path>` or `--key <name>`; default key = the run worktree's branch, `/` flattened to `-`, file `<tmpdir>/pi-handoff/<key>.md`) and appends the gauntlet `## Process state` section in the layout fixed by `skills/gauntlet-resume/reference/brief-contract.md` (`## Producers`); stops instead of writing when the skill is absent or the installed cohort still writes that section itself. `/skill:gauntlet-resume` with no arguments lists the briefs under `<tmpdir>/pi-handoff/` for a human pick, and a token that looks like a path is always a brief file (missing -> stop, never a scan). `scripts/ci.mjs` gains a drift lint (`scripts/brief-contract-lint.mjs`): the three process-state grammar lines may appear only in the contract file, both skills must cite it, and gauntlet-handoff may not carry cohort's core headings. Requires pi-cohort >= the release shipping [pi-cohort #18](https://github.com/jjuraszek/pi-cohort/issues/18) (exact version filled at release). (#40)
|
package/package.json
CHANGED
|
@@ -176,7 +176,7 @@ Inline council dispatch, reusing spec-council config and personas - **not** `/sk
|
|
|
176
176
|
|
|
177
177
|
1. Resolve `gauntlet_setting({ key: "specCouncil" })` when the tool exists. Verdict `council` -> dispatch `spec-council-member`s in parallel plus a `spec-council-synthesizer` chair. Verdict `worker` (or empty members) -> one fresh `worker` critique, its task text carrying the absolute path to `reference/ticket-wording.md` (resolved against this skill's own directory). Malformed config -> one warning line, then branch on verdict.
|
|
178
178
|
2. **Dispatch shape**, mirroring `/skill:roasting-the-spec`: write the draft body and the source snapshot (original ticket + comments, or the create-mode inputs) to absolute temp files under `mktemp -d`; delimit untrusted snapshots as data. The raw ask (create mode) or the original ticket body (repair mode) is also passed as the `Human input (verbatim; off-limits for over-spec)` block. When a split is proposed, the draft artifact holds all N proposed bodies plus their three-line justification blocks (see Split rule) in one file, not a single body. Two separate calls - never fuse members and chair into one chain (a fused chain lets one member failure kill the roast before the chair runs). Call 1: one member fanout with `cwd` = repo root, absolute `output` paths per member, run-level `control: { needsAttentionAfterMs: 60000, inFlightSilenceCeilingMs: 240000, inFlightSilenceKillMs: 300000 }` (sits beside `tasks`, not inside each task; effective silence-kill max(300s, 240+60) = 300s - record all three fields verbatim so a pi-cohort default change cannot stretch the kill). Then probe the member output files on disk with item 7's usable test. Call 2: the chair, with the usable member files via `reads`, the same control block (`:low` chair turns are short), and task text that (a) forbids repository access - member disagreement on a fact is reported in the synthesis, never verified against the repo - and (b) states coverage: `Coverage: N of M members reported; <slug>: <reason>` (pi-cohort's kill diagnostic when present, else "no output produced"; omit reasons at full coverage; singular wording when one member reported). Member task text: *the draft at `<path>` is the artifact under review; this ticket brief supersedes your spec-axis template - emit the same findings format against the draft; content-only review: the temp files plus the two referenced reference paths (split-axes, ticket-wording) are the entire permitted input - do not read, search, or scan the repository; do not edit any file.* Include the absolute paths to `reference/split-axes.md` and `reference/ticket-wording.md` (both resolved against this skill's own directory) in each member's task text - members run with `cwd` = the consumer repo, where a package-relative path does not resolve.
|
|
179
|
-
3. **Effort: cheap by default.** Append a `:low` thinking suffix to each member's model string at dispatch (this beats the persona's frontmatter `xhigh` pin). Same for the chair: a configured chair string gets any existing suffix replaced with `:low`; an unconfigured chair is dispatched
|
|
179
|
+
3. **Effort: cheap by default.** Append a `:low` thinking suffix to each member's model string at dispatch (this beats the persona's frontmatter `xhigh` pin). Same for the chair: a configured chair string gets any existing suffix replaced with `:low`; an unconfigured chair is dispatched with the main loop's own string, read from `$PI_PROVIDER`/`$PI_MODEL` with the bash tool, plus `:low`. The `worker` fallback carries no thinking pin - it runs at the preset's default. **Full-roast escape:** the user may request a full roast, dispatching all model strings bare/as-configured, restoring the xhigh pins; a full roast reuses the spec-roast control blocks (members `{ needsAttentionAfterMs: 300000, inFlightSilenceCeilingMs: 300000, inFlightSilenceKillMs: 600000 }`, chair `{ needsAttentionAfterMs: 300000, inFlightSilenceCeilingMs: 600000, inFlightSilenceKillMs: 900000 }`) - the 5-minute figures in item 2 are `:low`-only.
|
|
180
180
|
4. **Brief covers three axes**, absorbing the fidelity-review role without a new persona: *fidelity* - compare draft against source intent (original ticket + comments in repair; prompt + answers in create), flag `lost` / `added` / `gap` (unpacking existing claims to satisfy the ticket wording contract is not `added`; contract-conformance findings on Context/Problem/Idea outrank fidelity flags that only object to extra explanation of the same claims); and *quality* - problem framing, AC integrity beyond the deterministic gate, scope, wording, and conformance to the ticket wording contract (reference path provided in every roast brief); and *split soundness* - if the draft proposes a split, test each slice against the split-axes reference (path provided in the task text); an architecture-shaped boundary is reported as a finding line containing the marker `split-axis:` (members keep their existing spec-axis findings template; the marker is a substring flag within it, not a new findings kind), e.g. `- [major] split-axis: <slice> - <why> -> merge`. Members may argue toward one ticket, never propose or endorse a split.
|
|
181
181
|
5. Disposition: unambiguous concrete fixes applied to the draft (one re-pass max); ambiguous findings surfaced at the confirmation gate. Roast edits affect the body draft pre-write only, never posted as a tracker comment, and re-run the deterministic gates (pipeline step 5). Additionally, the parent scans the **usable member output files (item 7's structural test) directly** for lines containing `split-axis:` (substring match), independent of the chair synthesis; any such finding auto-applies a merge - the split is withdrawn and the draft becomes one ticket with phased AC groups, inside the same one-re-pass budget, and the pre-merge N-body draft is kept alongside: a human re-request of the split at the gate re-presents those N bodies old->new as the approval diff (see the Split rule's sticky override). The chair keeps every other axis; clearing a `split-axis:` finding is not on its path. The same directional rule - toward one ticket, never toward a split - binds the `worker` fallback and the runtime conditional (item 6).
|
|
182
182
|
6. **Runtime conditional (the one allowed):** on a harness with no `gauntlet_setting`/`subagent()` (e.g. Claude Code), dispatch fresh general-purpose subagents via that harness's native facility at low effort, with the same three-axis brief (including the absolute `reference/ticket-wording.md` path) and temp-file artifacts - on such a harness this conditional IS the roast, so the contract path must ride along.
|
|
@@ -21,7 +21,7 @@ Your context window holds the full plan, prior decisions, and conversation histo
|
|
|
21
21
|
|
|
22
22
|
- **No context pollution.** Task N's noise doesn't leak into Task N+1.
|
|
23
23
|
- **Tighter focus.** Subagent reads less, makes fewer cross-task assumptions, ships smaller diffs.
|
|
24
|
-
- **Cheaper at scale.**
|
|
24
|
+
- **Cheaper at scale.** Operator pins (`subagents.agentOverrides.<agent>.model`) route simple subtasks to smaller models; you spend top-tier tokens on orchestration and hard problems.
|
|
25
25
|
|
|
26
26
|
You are the **orchestrator**. You read the plan, dispatch, review the review, decide. You do **not** write code yourself.
|
|
27
27
|
|
|
@@ -111,30 +111,13 @@ Every implementer dispatch returns one of:
|
|
|
111
111
|
|
|
112
112
|
Implementer prompt templates must instruct subagents to return one of these statuses explicitly.
|
|
113
113
|
|
|
114
|
-
## Model
|
|
114
|
+
## Model
|
|
115
115
|
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
| Standard implementation | Default | Most plan tasks — feature work with tests, refactor with tests |
|
|
122
|
-
| Hard / novel / large surface | Most capable | New subsystem, complex algorithm, cross-service contract change |
|
|
123
|
-
| Spec review | Default | Reads diff + spec, mechanical comparison |
|
|
124
|
-
| Code-quality review | Most capable | Judgment call on naming, design, complexity |
|
|
125
|
-
| Conformance / closure | Most capable | Whole-deliverable-vs-origin intent gate (`conformance-reviewer`; model from `gauntlet_setting({ key: "closureReview" }).model`, injected call-site) |
|
|
126
|
-
|
|
127
|
-
```ts
|
|
128
|
-
subagent({
|
|
129
|
-
agent: "implementer",
|
|
130
|
-
async: false,
|
|
131
|
-
cwd: "<abs worktree path>",
|
|
132
|
-
task: "...",
|
|
133
|
-
model: "anthropic/claude-haiku-4" // cheap tier
|
|
134
|
-
})
|
|
135
|
-
```
|
|
136
|
-
|
|
137
|
-
When in doubt, default. Don't downgrade reviewers — false negatives are expensive.
|
|
116
|
+
- Omit `model:` so the operator's `subagents.agentOverrides.<agent>.model` pin applies, or else the child inherits the main loop.
|
|
117
|
+
- Pass `model:` only a runtime-resolved `gauntlet_setting` value (`closureReview.model`, `escalationLoop.implModel`, `specCouncil.members[]`, `specCouncil.chair`) or a model the user named for that dispatch.
|
|
118
|
+
- Read the main loop's own string from `$PI_PROVIDER`/`$PI_MODEL` with the bash tool only when a skill must bypass a persona's frontmatter `thinking` pin for cost or must pass an explicit fallback after a configured model is unreachable.
|
|
119
|
+
- Omit `model:` when `closureReview.model` or `specCouncil.chair` is `undefined`; pass an `implModel` string verbatim, and stop with a note when `implModel` is `undefined` (Fix-Loop Rounds).
|
|
120
|
+
- Apply suffix edits such as `:low` only to a runtime-resolved string, never to a name the skill wrote; never write a provider or model name, tier tasks by cost, or swap a pinned model for a cheaper or stronger one.
|
|
138
121
|
|
|
139
122
|
## Dispatch
|
|
140
123
|
|
|
@@ -45,7 +45,7 @@ The persona ships model-free. Get the model from `gauntlet_setting({ key: "closu
|
|
|
45
45
|
If `gauntlet_setting` is unavailable, stop and report - never fall back to a manual bash/JSON
|
|
46
46
|
settings merge. Inject the model **call-site** on the dispatch (omit `model:` when it is `undefined` to inherit
|
|
47
47
|
the parent's model) — the same mechanism the spec-council chair uses. If the configured model
|
|
48
|
-
is unreachable, retry once
|
|
48
|
+
is unreachable, retry once passing the main loop's own string explicitly (read from `$PI_PROVIDER`/`$PI_MODEL`; the phase-tracker guard blocks an omitted `model:` when `closureReview.model` is set and warns on the difference). Point it at the strongest reasoning model
|
|
49
49
|
the resolved config can reach — this is the last correctness gate. `thinking` stays
|
|
50
50
|
frontmatter-pinned at `xhigh` and is not call-site overridable, so the config supplies only
|
|
51
51
|
`model`.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: writing-skills
|
|
3
|
-
description: Use when creating
|
|
3
|
+
description: Use when creating, editing, or refactoring any SKILL.md, its reference/ files, an agent persona (`agents/*.md`, `.pi/agents/*.md`), or a prompt template - including one-line edits - or when such a file exceeds 500 lines, gains if/else branching, or drifts from imperative voice.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
> **Related skills:** Test new skills with `/skill:test-driven-development` discipline. Verify they work with `/skill:verification-before-completion`.
|
|
@@ -19,6 +19,15 @@ Write a pressure scenario for a subagent, watch it fail (baseline), write the sk
|
|
|
19
19
|
|
|
20
20
|
**REQUIRED BACKGROUND:** Read `/skill:test-driven-development` first. This skill applies RED → GREEN → REFACTOR to documentation.
|
|
21
21
|
|
|
22
|
+
## Authoring rules
|
|
23
|
+
|
|
24
|
+
These four rules bind the lines an edit adds or changes in any skill, persona, or prompt-template file. Leave pre-existing text alone (minimal diff). Where `reference/anthropic-best-practices.md` differs (its "Conditional workflow pattern"), this section wins.
|
|
25
|
+
|
|
26
|
+
- **Imperative voice.** Write instructions as commands. Never "should", "consider", "you may want to", "it is recommended".
|
|
27
|
+
- **Low conditionality.** One path per step. Branch only on a runtime fact the agent can observe - a tool result, a file's presence, a settings value - never on the reader's judgment. Two branches are the ceiling; move a third into a table or a `reference/` file.
|
|
28
|
+
- **Minimal diff.** A rule change touches the sentence that owns the rule, not the section. Never restate a rule in a second place; link the owner.
|
|
29
|
+
- **Oversized skill.** A change touching a SKILL.md over 500 lines extracts the concern it touches - or, when that concern is small, the largest self-contained `##` section - into `reference/<topic>.md` or a sibling md, in the same change, before it lands. Keep a one-line "read X now" pointer in the body at the step that needs it.
|
|
30
|
+
|
|
22
31
|
## Where Skills Live in Pi
|
|
23
32
|
|
|
24
33
|
Pi discovers skills from multiple roots (see `docs/skills.md` in `@earendil-works/pi-coding-agent` for the full list). The common ones:
|
|
@@ -194,7 +203,7 @@ Don't invent capabilities. Don't reference Claude Code's `Task` tool, OpenCode h
|
|
|
194
203
|
NO SKILL WITHOUT A FAILING TEST FIRST
|
|
195
204
|
```
|
|
196
205
|
|
|
197
|
-
Applies to new skills AND edits to existing skills.
|
|
206
|
+
Applies to new skills AND edits to existing skills. The RED-GREEN-REFACTOR baseline binds skill bodies - new skills and behavior-changing edits to SKILL.md or reference files; a wording-only edit to a skill, persona, or prompt template is verified by reading the changed lines back against `## Authoring rules`.
|
|
198
207
|
|
|
199
208
|
Wrote a skill before testing it? Delete it. Start over.
|
|
200
209
|
Edited a skill without testing? Same violation.
|
|
@@ -376,7 +385,7 @@ Use `plan_tracker` to create tasks for each item:
|
|
|
376
385
|
|
|
377
386
|
**Pi-specific**
|
|
378
387
|
- [ ] If discipline skill *and* the rule is mechanically detectable at a tool boundary: consider an `extensions/*.ts` hook (high bar — runtime hooks are deliberately slim; only add ones that beat false-positive heuristics)
|
|
379
|
-
- [ ] If skill has >500 lines: split deep content into `reference/<topic>.md` and instruct the agent to read the specific file inline
|
|
388
|
+
- [ ] If skill has >500 lines, or a skill this change touches does: split deep content into `reference/<topic>.md` and instruct the agent to read the specific file inline
|
|
380
389
|
- [ ] Skill location: project-scoped lives under `.pi/skills/`; cross-harness skills shared with Claude Code under `.agents/skills/`; reusable workflow skills belong in a package like `pi-gauntlet`
|
|
381
390
|
- [ ] Update routing: link from `AGENTS.md` if cross-cutting
|
|
382
391
|
|
|
@@ -409,11 +418,11 @@ Labels should carry semantic meaning.
|
|
|
409
418
|
## Red Flags — STOP
|
|
410
419
|
|
|
411
420
|
- Wrote a skill without running a baseline scenario first
|
|
412
|
-
-
|
|
421
|
+
- Made a behavior-changing skill edit without re-running the relevant baseline
|
|
413
422
|
- Description starts with "This skill does…" (summary instead of trigger)
|
|
414
423
|
- Code-then-test ordering in the skill body
|
|
415
424
|
- Force-loading other skills with `@`
|
|
416
|
-
- SKILL.md over 500 lines with no `reference/` split
|
|
425
|
+
- SKILL.md over 500 lines, or a skill this change touches over 500 lines, with no `reference/` split
|
|
417
426
|
- Referencing tools that don't exist in pi (Claude Code's `Task`, OpenCode hooks, etc.) for a pi-scope skill
|
|
418
427
|
- About to ship multiple skills in a batch without testing each
|
|
419
428
|
- "I'll test it later" — that means never
|