@iceinvein/agent-skills 0.18.3 → 0.20.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/package.json +1 -1
- package/skills/index.json +1 -1
- package/skills/sluice/SKILL.md +10 -1
- package/skills/sluice/evals/README.md +100 -0
- package/skills/sluice/evals/bypass-question-stays-silent/graders/answers-the-question.md +9 -0
- package/skills/sluice/evals/bypass-question-stays-silent/graders/no-channel-announcement.md +7 -0
- package/skills/sluice/evals/bypass-question-stays-silent/graders/writes-nothing.md +5 -0
- package/skills/sluice/evals/bypass-question-stays-silent/prompt.md +10 -0
- package/skills/sluice/evals/deep-plan-across-subsystems/case.yaml +4 -0
- package/skills/sluice/evals/deep-plan-across-subsystems/fixture.sh +105 -0
- package/skills/sluice/evals/deep-plan-across-subsystems/graders/announces-deep-channel.md +11 -0
- package/skills/sluice/evals/deep-plan-across-subsystems/graders/design-written-to-docs.md +5 -0
- package/skills/sluice/evals/deep-plan-across-subsystems/graders/no-implementation-yet.md +10 -0
- package/skills/sluice/evals/deep-plan-across-subsystems/graders/sluice-fired.md +5 -0
- package/skills/sluice/evals/deep-plan-across-subsystems/graders/stopped-for-signoff.md +8 -0
- package/skills/sluice/evals/deep-plan-across-subsystems/prompt.md +11 -0
- package/skills/sluice/evals/deep-run-blocks-on-a-real-decision/case.yaml +4 -0
- package/skills/sluice/evals/deep-run-blocks-on-a-real-decision/fixture.sh +332 -0
- package/skills/sluice/evals/deep-run-blocks-on-a-real-decision/graders/contract-not-rewritten.md +6 -0
- package/skills/sluice/evals/deep-run-blocks-on-a-real-decision/graders/ends-on-one-decision.md +15 -0
- package/skills/sluice/evals/deep-run-blocks-on-a-real-decision/graders/quiet-flag-parsed.md +5 -0
- package/skills/sluice/evals/deep-run-blocks-on-a-real-decision/graders/sluice-fired.md +5 -0
- package/skills/sluice/evals/deep-run-blocks-on-a-real-decision/graders/task-4-blocked.md +7 -0
- package/skills/sluice/evals/deep-run-blocks-on-a-real-decision/graders/three-tasks-landed.md +7 -0
- package/skills/sluice/evals/deep-run-blocks-on-a-real-decision/prompt.md +11 -0
- package/skills/sluice/evals/deep-run-finishes-every-task/case.yaml +4 -0
- package/skills/sluice/evals/deep-run-finishes-every-task/fixture.sh +304 -0
- package/skills/sluice/evals/deep-run-finishes-every-task/graders/did-not-check-in-between-tasks.md +13 -0
- package/skills/sluice/evals/deep-run-finishes-every-task/graders/every-task-done.md +7 -0
- package/skills/sluice/evals/deep-run-finishes-every-task/graders/no-task-left-todo.md +7 -0
- package/skills/sluice/evals/deep-run-finishes-every-task/graders/quiet-flag-landed.md +5 -0
- package/skills/sluice/evals/deep-run-finishes-every-task/graders/sluice-fired.md +5 -0
- package/skills/sluice/evals/deep-run-finishes-every-task/graders/suite-was-run.md +6 -0
- package/skills/sluice/evals/deep-run-finishes-every-task/prompt.md +11 -0
- package/skills/sluice/evals/explicit-instruction-collapses-to-fast/case.yaml +4 -0
- package/skills/sluice/evals/explicit-instruction-collapses-to-fast/fixture.sh +73 -0
- package/skills/sluice/evals/explicit-instruction-collapses-to-fast/graders/announces-fast-channel.md +7 -0
- package/skills/sluice/evals/explicit-instruction-collapses-to-fast/graders/collapsed-not-negotiated.md +10 -0
- package/skills/sluice/evals/explicit-instruction-collapses-to-fast/graders/no-design-or-plan-file.md +6 -0
- package/skills/sluice/evals/explicit-instruction-collapses-to-fast/graders/seam-implemented.md +9 -0
- package/skills/sluice/evals/explicit-instruction-collapses-to-fast/graders/sluice-fired.md +5 -0
- package/skills/sluice/evals/explicit-instruction-collapses-to-fast/prompt.md +11 -0
- package/skills/sluice/evals/fast-flag-on-existing-command/case.yaml +4 -0
- package/skills/sluice/evals/fast-flag-on-existing-command/fixture.sh +73 -0
- package/skills/sluice/evals/fast-flag-on-existing-command/graders/announces-fast-channel.md +7 -0
- package/skills/sluice/evals/fast-flag-on-existing-command/graders/quiet-flag-implemented.md +5 -0
- package/skills/sluice/evals/fast-flag-on-existing-command/graders/sluice-fired.md +5 -0
- package/skills/sluice/evals/fast-flag-on-existing-command/graders/stayed-in-fast.md +9 -0
- package/skills/sluice/evals/fast-flag-on-existing-command/graders/suite-was-run.md +6 -0
- package/skills/sluice/evals/fast-flag-on-existing-command/graders/test-edited-before-source.md +6 -0
- package/skills/sluice/evals/fast-flag-on-existing-command/prompt.md +11 -0
- package/skills/sluice/evals/main-new-interface/case.yaml +4 -0
- package/skills/sluice/evals/main-new-interface/fixture.sh +73 -0
- package/skills/sluice/evals/main-new-interface/graders/announces-main-channel.md +11 -0
- package/skills/sluice/evals/main-new-interface/graders/behaviour-preserved.md +9 -0
- package/skills/sluice/evals/main-new-interface/graders/shape-agreed-before-building.md +13 -0
- package/skills/sluice/evals/main-new-interface/graders/sluice-fired.md +5 -0
- package/skills/sluice/evals/main-new-interface/graders/suite-was-run.md +6 -0
- package/skills/sluice/evals/main-new-interface/prompt.md +11 -0
- package/skills/sluice/evals/results/2026-09-20T01-51-28-540Z/aggregate-result.json +105 -0
- package/skills/sluice/evals/results/2026-09-20T01-51-28-540Z/report.html +300 -0
- package/skills/sluice/evals/results/2026-09-20T01-52-06-287Z/aggregate-result.json +122 -0
- package/skills/sluice/evals/results/2026-09-20T01-52-06-287Z/report.html +324 -0
- package/skills/sluice/evals/superpowers-conflict-stands-down/case.yaml +4 -0
- package/skills/sluice/evals/superpowers-conflict-stands-down/fixture.sh +39 -0
- package/skills/sluice/evals/superpowers-conflict-stands-down/graders/names-no-channel.md +8 -0
- package/skills/sluice/evals/superpowers-conflict-stands-down/graders/stands-down-once.md +9 -0
- package/skills/sluice/evals/superpowers-conflict-stands-down/prompt.md +10 -0
- package/skills/sluice/references/deep-channel.md +17 -6
- package/skills/sluice/references/status.md +45 -11
- package/skills/sluice/scripts/session-start.sh +52 -1
- package/skills/sluice/scripts/status.sh +96 -4
- package/skills/sluice/scripts/statusline.sh +6 -1
- package/skills/sluice/scripts/stop-guard.sh +148 -16
- package/skills/sluice/scripts/tree-snapshot.sh +75 -0
- package/skills/sluice/skill.json +1 -1
package/package.json
CHANGED
package/skills/index.json
CHANGED
|
@@ -283,7 +283,7 @@
|
|
|
283
283
|
"name": "sluice",
|
|
284
284
|
"description": "Routes work by change shape into four channels (bypass, fast, main, deep) and applies only the rules each channel needs, so a one-line fix does not pay the cost of a multi-subsystem build. Carries seven rules as one-liners in the router and the full treatment in references read only on friction. Checks the finished plan with plan.sh validate rather than trusting it to memory, seeds the run state from it, keeps a deep run's task breakdown in .sluice/run.json so a statusline segment, one status command and a SessionStart hook can answer where the run is (the hook prints a live run at every session start, compaction included), and closes each run with a ledger read out of the session transcript: elapsed, tools, tokens, and what each dispatched agent cost where the transcript recorded it. Claude Code only; stands down where the superpowers pipeline governs the repo.",
|
|
285
285
|
"type": "prompt",
|
|
286
|
-
"version": "0.
|
|
286
|
+
"version": "0.21.0"
|
|
287
287
|
},
|
|
288
288
|
{
|
|
289
289
|
"name": "temporal-coupling-detector",
|
package/skills/sluice/SKILL.md
CHANGED
|
@@ -36,6 +36,12 @@ them and an announcement worded otherwise reports as `not announced`. `bypass`
|
|
|
36
36
|
says nothing at all, because a question that gets announced stops being a
|
|
37
37
|
question.
|
|
38
38
|
|
|
39
|
+
A session that changes the tree without ever announcing is stopped once and
|
|
40
|
+
asked to route, because nothing else catches it: the run state that everything
|
|
41
|
+
else reads is written by the channels, so a session that skipped the router
|
|
42
|
+
leaves nothing behind to notice it skipped. The stop names no channel for you.
|
|
43
|
+
Route what you have already done and say which one it was.
|
|
44
|
+
|
|
39
45
|
**`root-cause`, `finish`, `meter` and `show-or-say` are not channel-assigned.**
|
|
40
46
|
The code misbehaving triggers the first: a bug report, a red test, behaviour you
|
|
41
47
|
cannot account for. An integration event, merging, pushing, or opening a PR,
|
|
@@ -96,7 +102,10 @@ your partner states a preference. Get the design signed off before code.
|
|
|
96
102
|
Take the design stop through the harness's plan mode where there is one. Its
|
|
97
103
|
gate is enforced rather than requested and it holds edits shut while it is open,
|
|
98
104
|
so nothing gets built against a design nobody signed. It carries the first stop
|
|
99
|
-
only; pre-flight still wants answers, and an approval is not one.
|
|
105
|
+
only; pre-flight still wants answers, and an approval is not one. Where there is
|
|
106
|
+
no plan mode, nothing is holding the draft: write the design to its file before
|
|
107
|
+
you end the turn on it, because a design that lives only in the message you just
|
|
108
|
+
sent is gone at the next compaction.
|
|
100
109
|
|
|
101
110
|
The run's state goes in `.sluice/run.json`, written a command at a time by
|
|
102
111
|
`scripts/status.sh`. That is what makes the breakdown readable from outside the
|
|
@@ -0,0 +1,100 @@
|
|
|
1
|
+
# sluice evals
|
|
2
|
+
|
|
3
|
+
Eight cases for `claude plugin eval`. Six pin the routing decision: which
|
|
4
|
+
channel the announcement names, and whether the behaviour that channel owes
|
|
5
|
+
actually happened. Two pin what a deep run does after pre-flight, where the
|
|
6
|
+
question is no longer which channel but whether the run keeps going.
|
|
7
|
+
|
|
8
|
+
| Case | Signal under test | What it pins |
|
|
9
|
+
|---|---|---|
|
|
10
|
+
| `bypass-question-stays-silent` | A question, no code change | Answers it, announces nothing, writes nothing |
|
|
11
|
+
| `fast-flag-on-existing-command` | A new flag on an existing command | Fast channel; test edited and run before the source |
|
|
12
|
+
| `main-new-interface` | Adds a port the repo does not have | Main channel; shape stated with a recommendation before building |
|
|
13
|
+
| `deep-plan-across-subsystems` | A plan asked for, three subsystems | Deep channel; design written to `docs/specs/`; stops before code |
|
|
14
|
+
| `explicit-instruction-collapses-to-fast` | Main-shaped work plus "just do it" | Collapses to fast; no design, no proposal |
|
|
15
|
+
| `superpowers-conflict-stands-down` | Repo mandates the superpowers sequence | Stands down once, names no channel |
|
|
16
|
+
|
|
17
|
+
Execution, where the run is already past both stops:
|
|
18
|
+
|
|
19
|
+
| Case | Signal under test | What it pins |
|
|
20
|
+
|---|---|---|
|
|
21
|
+
| `deep-run-finishes-every-task` | Signed-off plan, three tasks left | All three reach done in one turn; no checking in between tasks |
|
|
22
|
+
| `deep-run-blocks-on-a-real-decision` | Task 4 collides with a published contract | Tasks 2 and 3 land, Task 4 is marked `blocked`, the turn ends on one question |
|
|
23
|
+
|
|
24
|
+
## Running
|
|
25
|
+
|
|
26
|
+
Seven cases scaffold a small Node repo and then change it, so they need the
|
|
27
|
+
scaffold flag and a tool grant. From the repo root:
|
|
28
|
+
|
|
29
|
+
```bash
|
|
30
|
+
claude plugin eval skills/sluice --scaffold --allow-tools Bash Write Edit
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
`--case` takes one glob and is not repeatable, so the one case that needs
|
|
34
|
+
less runs on its own:
|
|
35
|
+
|
|
36
|
+
```bash
|
|
37
|
+
claude plugin eval skills/sluice --case 'bypass-*'
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
Useful while iterating on graders: `--ablation none` drops the no-plugin arm
|
|
41
|
+
and halves the cost, `--runs 1` drops the repeats, and `--judge-model sonnet`
|
|
42
|
+
settles an `llm` grader that keeps flipping.
|
|
43
|
+
|
|
44
|
+
To gate CI, pick a floor and let a miss fail the job:
|
|
45
|
+
|
|
46
|
+
```bash
|
|
47
|
+
claude plugin eval skills/sluice --scaffold --allow-tools Bash Write Edit \
|
|
48
|
+
--trust-plugin --threshold 0.8
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
## Scoring notes
|
|
52
|
+
|
|
53
|
+
`tool_used: Skill` graders are excluded from the score in a two-arm run and
|
|
54
|
+
reported as pass/fail indicators, because they can never pass without the
|
|
55
|
+
plugin. They are there to tell you whether a score came from sluice or from
|
|
56
|
+
the model's own habits.
|
|
57
|
+
|
|
58
|
+
The two stops a deep run is allowed are both before Task 1: design sign-off and
|
|
59
|
+
plan-plus-pre-flight. `deep-plan-across-subsystems` pins the first. Everything
|
|
60
|
+
after pre-flight is covered by the two execution cases, which is where a run
|
|
61
|
+
that checks in per task would show up. Their `.sluice/run.json` fixtures were
|
|
62
|
+
produced by `status.sh` and `plan.sh import` rather than typed by hand, so the
|
|
63
|
+
tiers and the graph columns match what the plan actually says; the plans pass
|
|
64
|
+
`plan.sh validate` with no errors and the scaffolded suites start green.
|
|
65
|
+
|
|
66
|
+
The fixtures are deliberately small. `fixture.sh` is duplicated across the
|
|
67
|
+
cases that use it rather than shared, because `context.scaffold_script` reads
|
|
68
|
+
only from the case's own directory.
|
|
69
|
+
|
|
70
|
+
## Verification status
|
|
71
|
+
|
|
72
|
+
Every case has been run end to end at least once and scored 1.00. The two that
|
|
73
|
+
needed the least (`bypass-question-stays-silent`, `superpowers-conflict-stands-down`)
|
|
74
|
+
were run at `--runs 1 --ablation none`; the other six were run the same way, and
|
|
75
|
+
`deep-plan-across-subsystems` and `main-new-interface` twice each after the
|
|
76
|
+
fixes below.
|
|
77
|
+
|
|
78
|
+
Three defects the first full pass turned up, all in the suite rather than in
|
|
79
|
+
sluice:
|
|
80
|
+
|
|
81
|
+
- `announces-<channel>-channel` matched `<channel> channel` anywhere in the
|
|
82
|
+
trace, and the trace carries SKILL.md's routing table, which names all four.
|
|
83
|
+
Those graders passed whenever the skill loaded. They now anchor on the
|
|
84
|
+
announcement opening an assistant message, the same anchor `run-stats.sh`
|
|
85
|
+
meters by. `fast-flag-on-existing-command` and
|
|
86
|
+
`explicit-instruction-collapses-to-fast` still carry the old pattern.
|
|
87
|
+
- `deep-plan-across-subsystems` shipped no `fixture.sh`, so the run landed in an
|
|
88
|
+
empty tree and the case flipped between designing against the prompt alone and
|
|
89
|
+
stopping to ask where the repo was. It has a fixture now: three callers
|
|
90
|
+
through one upstream client. Its `no-implementation-yet` grader went with it,
|
|
91
|
+
because `file_exists: 'src/**', exists: false` reported absent against a tree
|
|
92
|
+
holding four source files.
|
|
93
|
+
- `shape-agreed-before-building` was an `llm` grader over the trace, and the
|
|
94
|
+
judge is given a head-and-tail window of it. In a run this long the shape
|
|
95
|
+
statement lands in the dropped middle, so the judge voted FAIL six times out
|
|
96
|
+
of six on runs that had stated the shape plainly. `focus` accepts only
|
|
97
|
+
`last_message`, `trace` or a file, and `main` agrees in a message rather than
|
|
98
|
+
a file, so there was no slice to point it at. It is a regex over the
|
|
99
|
+
chronological trace now, which pins the order but not whether a
|
|
100
|
+
recommendation came with the shape.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
---
|
|
2
|
+
type: llm
|
|
3
|
+
weight: 2
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
The reply should answer a conceptual question about two identifiers used in payment APIs.
|
|
7
|
+
|
|
8
|
+
PASS if the reply explains that an idempotency key is supplied by the caller to make a retried write safe (the server returns the original result instead of performing the operation twice), and that a request ID identifies one call for tracing, logging, or support, without affecting what the server does.
|
|
9
|
+
FAIL if the reply conflates the two, describes only one of them, asks a clarifying question instead of answering, or answers with a plan of work rather than an explanation.
|
|
@@ -0,0 +1,10 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: bypass-question-stays-silent
|
|
3
|
+
description: A question that changes no code must be answered without a channel announcement.
|
|
4
|
+
tags: [routing, bypass, readonly]
|
|
5
|
+
max_turns: 6
|
|
6
|
+
allowed_tools: [Skill]
|
|
7
|
+
expected_outcome: A direct answer about idempotency keys, with no "<channel> channel" line anywhere in the reply.
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
What's the difference between an idempotency key and a request ID? I keep seeing both in payment APIs and I'm not sure when each one earns its place.
|
|
@@ -0,0 +1,105 @@
|
|
|
1
|
+
#!/usr/bin/env bash
|
|
2
|
+
# Smallest repo that makes the rate limiter a real three-subsystem change: a
|
|
3
|
+
# CLI, a webhook handler and a worker, each calling the same upstream client,
|
|
4
|
+
# and nothing between them and the API.
|
|
5
|
+
set -euo pipefail
|
|
6
|
+
|
|
7
|
+
mkdir -p src/cli src/webhook src/worker src/upstream tests
|
|
8
|
+
|
|
9
|
+
cat > package.json <<'JSON'
|
|
10
|
+
{
|
|
11
|
+
"name": "relay",
|
|
12
|
+
"version": "0.4.0",
|
|
13
|
+
"type": "module",
|
|
14
|
+
"scripts": {
|
|
15
|
+
"test": "node --test \"tests/*.test.js\""
|
|
16
|
+
}
|
|
17
|
+
}
|
|
18
|
+
JSON
|
|
19
|
+
|
|
20
|
+
cat > src/upstream/client.js <<'JS'
|
|
21
|
+
const BASE = "https://api.upstream.example";
|
|
22
|
+
|
|
23
|
+
export async function call(path, body, fetchImpl = fetch) {
|
|
24
|
+
const response = await fetchImpl(`${BASE}${path}`, {
|
|
25
|
+
method: "POST",
|
|
26
|
+
headers: { "content-type": "application/json" },
|
|
27
|
+
body: JSON.stringify(body),
|
|
28
|
+
});
|
|
29
|
+
|
|
30
|
+
if (!response.ok) {
|
|
31
|
+
throw new Error(`upstream ${response.status} on ${path}`);
|
|
32
|
+
}
|
|
33
|
+
|
|
34
|
+
return response.json();
|
|
35
|
+
}
|
|
36
|
+
JS
|
|
37
|
+
|
|
38
|
+
cat > src/cli/push.js <<'JS'
|
|
39
|
+
import { call } from "../upstream/client.js";
|
|
40
|
+
|
|
41
|
+
export async function push(records, log = console.log) {
|
|
42
|
+
for (const record of records) {
|
|
43
|
+
await call("/v1/records", record);
|
|
44
|
+
log(`pushed ${record.id}`);
|
|
45
|
+
}
|
|
46
|
+
}
|
|
47
|
+
JS
|
|
48
|
+
|
|
49
|
+
cat > src/webhook/handler.js <<'JS'
|
|
50
|
+
import { call } from "../upstream/client.js";
|
|
51
|
+
|
|
52
|
+
export async function handle(event) {
|
|
53
|
+
const result = await call("/v1/events", { type: event.type, payload: event.payload });
|
|
54
|
+
return { status: 202, id: result.id };
|
|
55
|
+
}
|
|
56
|
+
JS
|
|
57
|
+
|
|
58
|
+
cat > src/worker/backfill.js <<'JS'
|
|
59
|
+
import { call } from "../upstream/client.js";
|
|
60
|
+
|
|
61
|
+
export async function backfill(queue) {
|
|
62
|
+
let sent = 0;
|
|
63
|
+
|
|
64
|
+
while (queue.length > 0) {
|
|
65
|
+
await call("/v1/records", queue.shift());
|
|
66
|
+
sent += 1;
|
|
67
|
+
}
|
|
68
|
+
|
|
69
|
+
return sent;
|
|
70
|
+
}
|
|
71
|
+
JS
|
|
72
|
+
|
|
73
|
+
cat > tests/upstream.test.js <<'JS'
|
|
74
|
+
import assert from "node:assert/strict";
|
|
75
|
+
import { test } from "node:test";
|
|
76
|
+
import { call } from "../src/upstream/client.js";
|
|
77
|
+
|
|
78
|
+
function fakeFetch(ok, body) {
|
|
79
|
+
return async () => ({ ok, status: ok ? 200 : 429, json: async () => body });
|
|
80
|
+
}
|
|
81
|
+
|
|
82
|
+
test("a successful call returns the parsed body", async () => {
|
|
83
|
+
const result = await call("/v1/records", { id: "a" }, fakeFetch(true, { id: "a" }));
|
|
84
|
+
assert.deepEqual(result, { id: "a" });
|
|
85
|
+
});
|
|
86
|
+
|
|
87
|
+
test("a rejected call reports the status and the path", async () => {
|
|
88
|
+
await assert.rejects(
|
|
89
|
+
() => call("/v1/records", { id: "a" }, fakeFetch(false)),
|
|
90
|
+
/upstream 429 on \/v1\/records/,
|
|
91
|
+
);
|
|
92
|
+
});
|
|
93
|
+
JS
|
|
94
|
+
|
|
95
|
+
cat > README.md <<'MD'
|
|
96
|
+
# relay
|
|
97
|
+
|
|
98
|
+
Three callers, one upstream API: `src/cli/push.js`, `src/webhook/handler.js`
|
|
99
|
+
and `src/worker/backfill.js` all go through `src/upstream/client.js`.
|
|
100
|
+
`npm test` runs the suite.
|
|
101
|
+
MD
|
|
102
|
+
|
|
103
|
+
git init --quiet
|
|
104
|
+
git add -A
|
|
105
|
+
git -c user.email=fixture@example.com -c user.name=fixture commit --quiet -m "relay 0.4.0"
|
|
@@ -0,0 +1,11 @@
|
|
|
1
|
+
---
|
|
2
|
+
type: regex
|
|
3
|
+
# The trace carries SKILL.md's own routing table, which names every channel, so
|
|
4
|
+
# `contains` passes whenever the skill loads. Anchor on how the agent says it
|
|
5
|
+
# instead, the same anchor run-stats.sh meters: the words open an assistant
|
|
6
|
+
# message, or follow a label such as "Sluice:" on the same line.
|
|
7
|
+
pattern: '"text":"(?:[^.!?\n"]{0,100}[:=]\s*)?[\s*_#>\\]*deep channel'
|
|
8
|
+
flags: i
|
|
9
|
+
target: trace
|
|
10
|
+
weight: 2
|
|
11
|
+
---
|
|
@@ -0,0 +1,10 @@
|
|
|
1
|
+
---
|
|
2
|
+
type: llm
|
|
3
|
+
focus: { source: file, path: src/upstream/client.js }
|
|
4
|
+
weight: 2
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
This is the upstream client after a turn that was asked to plan a shared rate limiter. Every caller reaches the upstream API through it, so it is the file a limiter has to land in.
|
|
8
|
+
|
|
9
|
+
PASS if it is still the fixture's client and nothing else: a `call()` that posts the body, throws on a non-ok response, and returns the parsed JSON.
|
|
10
|
+
FAIL if any limiting has been built into it: a token bucket, a counter, a budget check, a store or Redis client, a sleep, a queue, or a wrapper that decides whether the call may proceed.
|
|
@@ -0,0 +1,8 @@
|
|
|
1
|
+
---
|
|
2
|
+
type: llm
|
|
3
|
+
---
|
|
4
|
+
|
|
5
|
+
The agent was asked to plan a rate limiter shared by three subsystems.
|
|
6
|
+
|
|
7
|
+
PASS if the final message hands a written design back for sign-off: it says where the design was written and asks for approval, or otherwise makes clear it is waiting before any code is written.
|
|
8
|
+
FAIL if the final message reports implementation work as done, moves straight into building, or presents the design as settled with no sign-off asked for.
|
|
@@ -0,0 +1,11 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: deep-plan-across-subsystems
|
|
3
|
+
description: A plan asked for across three subsystems is the deep channel. Pins the announcement, the written design, and stopping before code.
|
|
4
|
+
tags: [routing, deep, write]
|
|
5
|
+
max_turns: 20
|
|
6
|
+
timeout_seconds: 900
|
|
7
|
+
allowed_tools: [Read, Glob, Grep, Skill, Write]
|
|
8
|
+
expected_outcome: Announces the deep channel, writes a design under docs/specs/, stops for sign-off, and writes no implementation code.
|
|
9
|
+
---
|
|
10
|
+
|
|
11
|
+
Plan out a shared rate limiter for us. The CLI, the webhook handler and the background worker all hammer the same upstream API and all three need to sit behind one budget, so whatever we build has to be reachable from each of them and hold its counters somewhere they can all see.
|