@iceinvein/agent-skills 0.18.3 → 0.20.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (76) hide show
  1. package/package.json +1 -1
  2. package/skills/index.json +1 -1
  3. package/skills/sluice/SKILL.md +10 -1
  4. package/skills/sluice/evals/README.md +100 -0
  5. package/skills/sluice/evals/bypass-question-stays-silent/graders/answers-the-question.md +9 -0
  6. package/skills/sluice/evals/bypass-question-stays-silent/graders/no-channel-announcement.md +7 -0
  7. package/skills/sluice/evals/bypass-question-stays-silent/graders/writes-nothing.md +5 -0
  8. package/skills/sluice/evals/bypass-question-stays-silent/prompt.md +10 -0
  9. package/skills/sluice/evals/deep-plan-across-subsystems/case.yaml +4 -0
  10. package/skills/sluice/evals/deep-plan-across-subsystems/fixture.sh +105 -0
  11. package/skills/sluice/evals/deep-plan-across-subsystems/graders/announces-deep-channel.md +11 -0
  12. package/skills/sluice/evals/deep-plan-across-subsystems/graders/design-written-to-docs.md +5 -0
  13. package/skills/sluice/evals/deep-plan-across-subsystems/graders/no-implementation-yet.md +10 -0
  14. package/skills/sluice/evals/deep-plan-across-subsystems/graders/sluice-fired.md +5 -0
  15. package/skills/sluice/evals/deep-plan-across-subsystems/graders/stopped-for-signoff.md +8 -0
  16. package/skills/sluice/evals/deep-plan-across-subsystems/prompt.md +11 -0
  17. package/skills/sluice/evals/deep-run-blocks-on-a-real-decision/case.yaml +4 -0
  18. package/skills/sluice/evals/deep-run-blocks-on-a-real-decision/fixture.sh +332 -0
  19. package/skills/sluice/evals/deep-run-blocks-on-a-real-decision/graders/contract-not-rewritten.md +6 -0
  20. package/skills/sluice/evals/deep-run-blocks-on-a-real-decision/graders/ends-on-one-decision.md +15 -0
  21. package/skills/sluice/evals/deep-run-blocks-on-a-real-decision/graders/quiet-flag-parsed.md +5 -0
  22. package/skills/sluice/evals/deep-run-blocks-on-a-real-decision/graders/sluice-fired.md +5 -0
  23. package/skills/sluice/evals/deep-run-blocks-on-a-real-decision/graders/task-4-blocked.md +7 -0
  24. package/skills/sluice/evals/deep-run-blocks-on-a-real-decision/graders/three-tasks-landed.md +7 -0
  25. package/skills/sluice/evals/deep-run-blocks-on-a-real-decision/prompt.md +11 -0
  26. package/skills/sluice/evals/deep-run-finishes-every-task/case.yaml +4 -0
  27. package/skills/sluice/evals/deep-run-finishes-every-task/fixture.sh +304 -0
  28. package/skills/sluice/evals/deep-run-finishes-every-task/graders/did-not-check-in-between-tasks.md +13 -0
  29. package/skills/sluice/evals/deep-run-finishes-every-task/graders/every-task-done.md +7 -0
  30. package/skills/sluice/evals/deep-run-finishes-every-task/graders/no-task-left-todo.md +7 -0
  31. package/skills/sluice/evals/deep-run-finishes-every-task/graders/quiet-flag-landed.md +5 -0
  32. package/skills/sluice/evals/deep-run-finishes-every-task/graders/sluice-fired.md +5 -0
  33. package/skills/sluice/evals/deep-run-finishes-every-task/graders/suite-was-run.md +6 -0
  34. package/skills/sluice/evals/deep-run-finishes-every-task/prompt.md +11 -0
  35. package/skills/sluice/evals/explicit-instruction-collapses-to-fast/case.yaml +4 -0
  36. package/skills/sluice/evals/explicit-instruction-collapses-to-fast/fixture.sh +73 -0
  37. package/skills/sluice/evals/explicit-instruction-collapses-to-fast/graders/announces-fast-channel.md +7 -0
  38. package/skills/sluice/evals/explicit-instruction-collapses-to-fast/graders/collapsed-not-negotiated.md +10 -0
  39. package/skills/sluice/evals/explicit-instruction-collapses-to-fast/graders/no-design-or-plan-file.md +6 -0
  40. package/skills/sluice/evals/explicit-instruction-collapses-to-fast/graders/seam-implemented.md +9 -0
  41. package/skills/sluice/evals/explicit-instruction-collapses-to-fast/graders/sluice-fired.md +5 -0
  42. package/skills/sluice/evals/explicit-instruction-collapses-to-fast/prompt.md +11 -0
  43. package/skills/sluice/evals/fast-flag-on-existing-command/case.yaml +4 -0
  44. package/skills/sluice/evals/fast-flag-on-existing-command/fixture.sh +73 -0
  45. package/skills/sluice/evals/fast-flag-on-existing-command/graders/announces-fast-channel.md +7 -0
  46. package/skills/sluice/evals/fast-flag-on-existing-command/graders/quiet-flag-implemented.md +5 -0
  47. package/skills/sluice/evals/fast-flag-on-existing-command/graders/sluice-fired.md +5 -0
  48. package/skills/sluice/evals/fast-flag-on-existing-command/graders/stayed-in-fast.md +9 -0
  49. package/skills/sluice/evals/fast-flag-on-existing-command/graders/suite-was-run.md +6 -0
  50. package/skills/sluice/evals/fast-flag-on-existing-command/graders/test-edited-before-source.md +6 -0
  51. package/skills/sluice/evals/fast-flag-on-existing-command/prompt.md +11 -0
  52. package/skills/sluice/evals/main-new-interface/case.yaml +4 -0
  53. package/skills/sluice/evals/main-new-interface/fixture.sh +73 -0
  54. package/skills/sluice/evals/main-new-interface/graders/announces-main-channel.md +11 -0
  55. package/skills/sluice/evals/main-new-interface/graders/behaviour-preserved.md +9 -0
  56. package/skills/sluice/evals/main-new-interface/graders/shape-agreed-before-building.md +13 -0
  57. package/skills/sluice/evals/main-new-interface/graders/sluice-fired.md +5 -0
  58. package/skills/sluice/evals/main-new-interface/graders/suite-was-run.md +6 -0
  59. package/skills/sluice/evals/main-new-interface/prompt.md +11 -0
  60. package/skills/sluice/evals/results/2026-09-20T01-51-28-540Z/aggregate-result.json +105 -0
  61. package/skills/sluice/evals/results/2026-09-20T01-51-28-540Z/report.html +300 -0
  62. package/skills/sluice/evals/results/2026-09-20T01-52-06-287Z/aggregate-result.json +122 -0
  63. package/skills/sluice/evals/results/2026-09-20T01-52-06-287Z/report.html +324 -0
  64. package/skills/sluice/evals/superpowers-conflict-stands-down/case.yaml +4 -0
  65. package/skills/sluice/evals/superpowers-conflict-stands-down/fixture.sh +39 -0
  66. package/skills/sluice/evals/superpowers-conflict-stands-down/graders/names-no-channel.md +8 -0
  67. package/skills/sluice/evals/superpowers-conflict-stands-down/graders/stands-down-once.md +9 -0
  68. package/skills/sluice/evals/superpowers-conflict-stands-down/prompt.md +10 -0
  69. package/skills/sluice/references/deep-channel.md +17 -6
  70. package/skills/sluice/references/status.md +45 -11
  71. package/skills/sluice/scripts/session-start.sh +52 -1
  72. package/skills/sluice/scripts/status.sh +96 -4
  73. package/skills/sluice/scripts/statusline.sh +6 -1
  74. package/skills/sluice/scripts/stop-guard.sh +148 -16
  75. package/skills/sluice/scripts/tree-snapshot.sh +75 -0
  76. package/skills/sluice/skill.json +1 -1
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@iceinvein/agent-skills",
3
- "version": "0.18.3",
3
+ "version": "0.20.0",
4
4
  "description": "Install agent skills into AI coding tools",
5
5
  "author": "iceinvein",
6
6
  "license": "MIT",
package/skills/index.json CHANGED
@@ -283,7 +283,7 @@
283
283
  "name": "sluice",
284
284
  "description": "Routes work by change shape into four channels (bypass, fast, main, deep) and applies only the rules each channel needs, so a one-line fix does not pay the cost of a multi-subsystem build. Carries seven rules as one-liners in the router and the full treatment in references read only on friction. Checks the finished plan with plan.sh validate rather than trusting it to memory, seeds the run state from it, keeps a deep run's task breakdown in .sluice/run.json so a statusline segment, one status command and a SessionStart hook can answer where the run is (the hook prints a live run at every session start, compaction included), and closes each run with a ledger read out of the session transcript: elapsed, tools, tokens, and what each dispatched agent cost where the transcript recorded it. Claude Code only; stands down where the superpowers pipeline governs the repo.",
285
285
  "type": "prompt",
286
- "version": "0.19.3"
286
+ "version": "0.21.0"
287
287
  },
288
288
  {
289
289
  "name": "temporal-coupling-detector",
@@ -36,6 +36,12 @@ them and an announcement worded otherwise reports as `not announced`. `bypass`
36
36
  says nothing at all, because a question that gets announced stops being a
37
37
  question.
38
38
 
39
+ A session that changes the tree without ever announcing is stopped once and
40
+ asked to route, because nothing else catches it: the run state that everything
41
+ else reads is written by the channels, so a session that skipped the router
42
+ leaves nothing behind to notice it skipped. The stop names no channel for you.
43
+ Route what you have already done and say which one it was.
44
+
39
45
  **`root-cause`, `finish`, `meter` and `show-or-say` are not channel-assigned.**
40
46
  The code misbehaving triggers the first: a bug report, a red test, behaviour you
41
47
  cannot account for. An integration event, merging, pushing, or opening a PR,
@@ -96,7 +102,10 @@ your partner states a preference. Get the design signed off before code.
96
102
  Take the design stop through the harness's plan mode where there is one. Its
97
103
  gate is enforced rather than requested and it holds edits shut while it is open,
98
104
  so nothing gets built against a design nobody signed. It carries the first stop
99
- only; pre-flight still wants answers, and an approval is not one.
105
+ only; pre-flight still wants answers, and an approval is not one. Where there is
106
+ no plan mode, nothing is holding the draft: write the design to its file before
107
+ you end the turn on it, because a design that lives only in the message you just
108
+ sent is gone at the next compaction.
100
109
 
101
110
  The run's state goes in `.sluice/run.json`, written a command at a time by
102
111
  `scripts/status.sh`. That is what makes the breakdown readable from outside the
@@ -0,0 +1,100 @@
1
+ # sluice evals
2
+
3
+ Eight cases for `claude plugin eval`. Six pin the routing decision: which
4
+ channel the announcement names, and whether the behaviour that channel owes
5
+ actually happened. Two pin what a deep run does after pre-flight, where the
6
+ question is no longer which channel but whether the run keeps going.
7
+
8
+ | Case | Signal under test | What it pins |
9
+ |---|---|---|
10
+ | `bypass-question-stays-silent` | A question, no code change | Answers it, announces nothing, writes nothing |
11
+ | `fast-flag-on-existing-command` | A new flag on an existing command | Fast channel; test edited and run before the source |
12
+ | `main-new-interface` | Adds a port the repo does not have | Main channel; shape stated with a recommendation before building |
13
+ | `deep-plan-across-subsystems` | A plan asked for, three subsystems | Deep channel; design written to `docs/specs/`; stops before code |
14
+ | `explicit-instruction-collapses-to-fast` | Main-shaped work plus "just do it" | Collapses to fast; no design, no proposal |
15
+ | `superpowers-conflict-stands-down` | Repo mandates the superpowers sequence | Stands down once, names no channel |
16
+
17
+ Execution, where the run is already past both stops:
18
+
19
+ | Case | Signal under test | What it pins |
20
+ |---|---|---|
21
+ | `deep-run-finishes-every-task` | Signed-off plan, three tasks left | All three reach done in one turn; no checking in between tasks |
22
+ | `deep-run-blocks-on-a-real-decision` | Task 4 collides with a published contract | Tasks 2 and 3 land, Task 4 is marked `blocked`, the turn ends on one question |
23
+
24
+ ## Running
25
+
26
+ Seven cases scaffold a small Node repo and then change it, so they need the
27
+ scaffold flag and a tool grant. From the repo root:
28
+
29
+ ```bash
30
+ claude plugin eval skills/sluice --scaffold --allow-tools Bash Write Edit
31
+ ```
32
+
33
+ `--case` takes one glob and is not repeatable, so the one case that needs
34
+ less runs on its own:
35
+
36
+ ```bash
37
+ claude plugin eval skills/sluice --case 'bypass-*'
38
+ ```
39
+
40
+ Useful while iterating on graders: `--ablation none` drops the no-plugin arm
41
+ and halves the cost, `--runs 1` drops the repeats, and `--judge-model sonnet`
42
+ settles an `llm` grader that keeps flipping.
43
+
44
+ To gate CI, pick a floor and let a miss fail the job:
45
+
46
+ ```bash
47
+ claude plugin eval skills/sluice --scaffold --allow-tools Bash Write Edit \
48
+ --trust-plugin --threshold 0.8
49
+ ```
50
+
51
+ ## Scoring notes
52
+
53
+ `tool_used: Skill` graders are excluded from the score in a two-arm run and
54
+ reported as pass/fail indicators, because they can never pass without the
55
+ plugin. They are there to tell you whether a score came from sluice or from
56
+ the model's own habits.
57
+
58
+ The two stops a deep run is allowed are both before Task 1: design sign-off and
59
+ plan-plus-pre-flight. `deep-plan-across-subsystems` pins the first. Everything
60
+ after pre-flight is covered by the two execution cases, which is where a run
61
+ that checks in per task would show up. Their `.sluice/run.json` fixtures were
62
+ produced by `status.sh` and `plan.sh import` rather than typed by hand, so the
63
+ tiers and the graph columns match what the plan actually says; the plans pass
64
+ `plan.sh validate` with no errors and the scaffolded suites start green.
65
+
66
+ The fixtures are deliberately small. `fixture.sh` is duplicated across the
67
+ cases that use it rather than shared, because `context.scaffold_script` reads
68
+ only from the case's own directory.
69
+
70
+ ## Verification status
71
+
72
+ Every case has been run end to end at least once and scored 1.00. The two that
73
+ needed the least (`bypass-question-stays-silent`, `superpowers-conflict-stands-down`)
74
+ were run at `--runs 1 --ablation none`; the other six were run the same way, and
75
+ `deep-plan-across-subsystems` and `main-new-interface` twice each after the
76
+ fixes below.
77
+
78
+ Three defects the first full pass turned up, all in the suite rather than in
79
+ sluice:
80
+
81
+ - `announces-<channel>-channel` matched `<channel> channel` anywhere in the
82
+ trace, and the trace carries SKILL.md's routing table, which names all four.
83
+ Those graders passed whenever the skill loaded. They now anchor on the
84
+ announcement opening an assistant message, the same anchor `run-stats.sh`
85
+ meters by. `fast-flag-on-existing-command` and
86
+ `explicit-instruction-collapses-to-fast` still carry the old pattern.
87
+ - `deep-plan-across-subsystems` shipped no `fixture.sh`, so the run landed in an
88
+ empty tree and the case flipped between designing against the prompt alone and
89
+ stopping to ask where the repo was. It has a fixture now: three callers
90
+ through one upstream client. Its `no-implementation-yet` grader went with it,
91
+ because `file_exists: 'src/**', exists: false` reported absent against a tree
92
+ holding four source files.
93
+ - `shape-agreed-before-building` was an `llm` grader over the trace, and the
94
+ judge is given a head-and-tail window of it. In a run this long the shape
95
+ statement lands in the dropped middle, so the judge voted FAIL six times out
96
+ of six on runs that had stated the shape plainly. `focus` accepts only
97
+ `last_message`, `trace` or a file, and `main` agrees in a message rather than
98
+ a file, so there was no slice to point it at. It is a regex over the
99
+ chronological trace now, which pins the order but not whether a
100
+ recommendation came with the shape.
@@ -0,0 +1,9 @@
1
+ ---
2
+ type: llm
3
+ weight: 2
4
+ ---
5
+
6
+ The reply should answer a conceptual question about two identifiers used in payment APIs.
7
+
8
+ PASS if the reply explains that an idempotency key is supplied by the caller to make a retried write safe (the server returns the original result instead of performing the operation twice), and that a request ID identifies one call for tracing, logging, or support, without affecting what the server does.
9
+ FAIL if the reply conflates the two, describes only one of them, asks a clarifying question instead of answering, or answers with a plan of work rather than an explanation.
@@ -0,0 +1,7 @@
1
+ ---
2
+ type: regex
3
+ pattern: '(bypass|fast|main|deep)\s+channel'
4
+ flags: i
5
+ match: not_contains
6
+ target: last_message
7
+ ---
@@ -0,0 +1,5 @@
1
+ ---
2
+ type: file_exists
3
+ path: '**/*'
4
+ exists: false
5
+ ---
@@ -0,0 +1,10 @@
1
+ ---
2
+ name: bypass-question-stays-silent
3
+ description: A question that changes no code must be answered without a channel announcement.
4
+ tags: [routing, bypass, readonly]
5
+ max_turns: 6
6
+ allowed_tools: [Skill]
7
+ expected_outcome: A direct answer about idempotency keys, with no "<channel> channel" line anywhere in the reply.
8
+ ---
9
+
10
+ What's the difference between an idempotency key and a request ID? I keep seeing both in payment APIs and I'm not sure when each one earns its place.
@@ -0,0 +1,4 @@
1
+ schema_version: "1.1"
2
+ name: deep-plan-across-subsystems
3
+ context:
4
+ scaffold_script: fixture.sh
@@ -0,0 +1,105 @@
1
+ #!/usr/bin/env bash
2
+ # Smallest repo that makes the rate limiter a real three-subsystem change: a
3
+ # CLI, a webhook handler and a worker, each calling the same upstream client,
4
+ # and nothing between them and the API.
5
+ set -euo pipefail
6
+
7
+ mkdir -p src/cli src/webhook src/worker src/upstream tests
8
+
9
+ cat > package.json <<'JSON'
10
+ {
11
+ "name": "relay",
12
+ "version": "0.4.0",
13
+ "type": "module",
14
+ "scripts": {
15
+ "test": "node --test \"tests/*.test.js\""
16
+ }
17
+ }
18
+ JSON
19
+
20
+ cat > src/upstream/client.js <<'JS'
21
+ const BASE = "https://api.upstream.example";
22
+
23
+ export async function call(path, body, fetchImpl = fetch) {
24
+ const response = await fetchImpl(`${BASE}${path}`, {
25
+ method: "POST",
26
+ headers: { "content-type": "application/json" },
27
+ body: JSON.stringify(body),
28
+ });
29
+
30
+ if (!response.ok) {
31
+ throw new Error(`upstream ${response.status} on ${path}`);
32
+ }
33
+
34
+ return response.json();
35
+ }
36
+ JS
37
+
38
+ cat > src/cli/push.js <<'JS'
39
+ import { call } from "../upstream/client.js";
40
+
41
+ export async function push(records, log = console.log) {
42
+ for (const record of records) {
43
+ await call("/v1/records", record);
44
+ log(`pushed ${record.id}`);
45
+ }
46
+ }
47
+ JS
48
+
49
+ cat > src/webhook/handler.js <<'JS'
50
+ import { call } from "../upstream/client.js";
51
+
52
+ export async function handle(event) {
53
+ const result = await call("/v1/events", { type: event.type, payload: event.payload });
54
+ return { status: 202, id: result.id };
55
+ }
56
+ JS
57
+
58
+ cat > src/worker/backfill.js <<'JS'
59
+ import { call } from "../upstream/client.js";
60
+
61
+ export async function backfill(queue) {
62
+ let sent = 0;
63
+
64
+ while (queue.length > 0) {
65
+ await call("/v1/records", queue.shift());
66
+ sent += 1;
67
+ }
68
+
69
+ return sent;
70
+ }
71
+ JS
72
+
73
+ cat > tests/upstream.test.js <<'JS'
74
+ import assert from "node:assert/strict";
75
+ import { test } from "node:test";
76
+ import { call } from "../src/upstream/client.js";
77
+
78
+ function fakeFetch(ok, body) {
79
+ return async () => ({ ok, status: ok ? 200 : 429, json: async () => body });
80
+ }
81
+
82
+ test("a successful call returns the parsed body", async () => {
83
+ const result = await call("/v1/records", { id: "a" }, fakeFetch(true, { id: "a" }));
84
+ assert.deepEqual(result, { id: "a" });
85
+ });
86
+
87
+ test("a rejected call reports the status and the path", async () => {
88
+ await assert.rejects(
89
+ () => call("/v1/records", { id: "a" }, fakeFetch(false)),
90
+ /upstream 429 on \/v1\/records/,
91
+ );
92
+ });
93
+ JS
94
+
95
+ cat > README.md <<'MD'
96
+ # relay
97
+
98
+ Three callers, one upstream API: `src/cli/push.js`, `src/webhook/handler.js`
99
+ and `src/worker/backfill.js` all go through `src/upstream/client.js`.
100
+ `npm test` runs the suite.
101
+ MD
102
+
103
+ git init --quiet
104
+ git add -A
105
+ git -c user.email=fixture@example.com -c user.name=fixture commit --quiet -m "relay 0.4.0"
@@ -0,0 +1,11 @@
1
+ ---
2
+ type: regex
3
+ # The trace carries SKILL.md's own routing table, which names every channel, so
4
+ # `contains` passes whenever the skill loads. Anchor on how the agent says it
5
+ # instead, the same anchor run-stats.sh meters: the words open an assistant
6
+ # message, or follow a label such as "Sluice:" on the same line.
7
+ pattern: '"text":"(?:[^.!?\n"]{0,100}[:=]\s*)?[\s*_#>\\]*deep channel'
8
+ flags: i
9
+ target: trace
10
+ weight: 2
11
+ ---
@@ -0,0 +1,5 @@
1
+ ---
2
+ type: file_exists
3
+ path: 'docs/specs/*.md'
4
+ weight: 2
5
+ ---
@@ -0,0 +1,10 @@
1
+ ---
2
+ type: llm
3
+ focus: { source: file, path: src/upstream/client.js }
4
+ weight: 2
5
+ ---
6
+
7
+ This is the upstream client after a turn that was asked to plan a shared rate limiter. Every caller reaches the upstream API through it, so it is the file a limiter has to land in.
8
+
9
+ PASS if it is still the fixture's client and nothing else: a `call()` that posts the body, throws on a non-ok response, and returns the parsed JSON.
10
+ FAIL if any limiting has been built into it: a token bucket, a counter, a budget check, a store or Redis client, a sleep, a queue, or a wrapper that decides whether the call may proceed.
@@ -0,0 +1,5 @@
1
+ ---
2
+ type: tool_used
3
+ tool: Skill
4
+ input_match: '"skill"\s*:\s*"(?:[\w-]+:)?sluice"'
5
+ ---
@@ -0,0 +1,8 @@
1
+ ---
2
+ type: llm
3
+ ---
4
+
5
+ The agent was asked to plan a rate limiter shared by three subsystems.
6
+
7
+ PASS if the final message hands a written design back for sign-off: it says where the design was written and asks for approval, or otherwise makes clear it is waiting before any code is written.
8
+ FAIL if the final message reports implementation work as done, moves straight into building, or presents the design as settled with no sign-off asked for.
@@ -0,0 +1,11 @@
1
+ ---
2
+ name: deep-plan-across-subsystems
3
+ description: A plan asked for across three subsystems is the deep channel. Pins the announcement, the written design, and stopping before code.
4
+ tags: [routing, deep, write]
5
+ max_turns: 20
6
+ timeout_seconds: 900
7
+ allowed_tools: [Read, Glob, Grep, Skill, Write]
8
+ expected_outcome: Announces the deep channel, writes a design under docs/specs/, stops for sign-off, and writes no implementation code.
9
+ ---
10
+
11
+ Plan out a shared rate limiter for us. The CLI, the webhook handler and the background worker all hammer the same upstream API and all three need to sit behind one budget, so whatever we build has to be reachable from each of them and hold its counters somewhere they can all see.
@@ -0,0 +1,4 @@
1
+ schema_version: "1.1"
2
+ name: deep-run-blocks-on-a-real-decision
3
+ context:
4
+ scaffold_script: fixture.sh