@forwardimpact/libharness 2.0.0 → 3.0.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +68 -65
- package/package.json +15 -13
- package/src/advisor.js +47 -41
- package/src/agent-runner.js +58 -48
- package/src/benchmark/apm-installer.js +28 -28
- package/src/benchmark/env-loader.js +24 -16
- package/src/benchmark/grade.js +44 -41
- package/src/benchmark/hidden-tests.js +25 -24
- package/src/benchmark/hook-env.js +11 -9
- package/src/benchmark/invariants.js +20 -17
- package/src/benchmark/judge.js +29 -28
- package/src/benchmark/npm-installer.js +9 -8
- package/src/benchmark/report.js +53 -50
- package/src/benchmark/result.js +24 -23
- package/src/benchmark/runner.js +75 -69
- package/src/benchmark/scheduler.js +17 -16
- package/src/benchmark/task-family.js +29 -27
- package/src/benchmark/trace-split.js +9 -8
- package/src/benchmark/workdir.js +27 -25
- package/src/claude-code-executable.js +11 -11
- package/src/commands/advisor-flags.js +8 -7
- package/src/commands/assert.js +16 -15
- package/src/commands/benchmark-definition.js +20 -20
- package/src/commands/benchmark-grade.js +13 -12
- package/src/commands/benchmark-report.js +5 -5
- package/src/commands/benchmark-run.js +31 -28
- package/src/commands/by-discussion.js +11 -11
- package/src/commands/callback.js +11 -11
- package/src/commands/discuss.js +8 -7
- package/src/commands/facilitate.js +16 -14
- package/src/commands/output.js +4 -3
- package/src/commands/run.js +15 -15
- package/src/commands/scan-logs.js +22 -20
- package/src/commands/selfedit.js +124 -0
- package/src/commands/supervise.js +13 -11
- package/src/commands/task-input.js +9 -9
- package/src/commands/tee.js +11 -10
- package/src/commands/trace.js +55 -42
- package/src/commands/work-tracker.js +4 -3
- package/src/cost.js +17 -17
- package/src/discuss-tools.js +16 -16
- package/src/discusser.js +39 -38
- package/src/events/github.js +54 -37
- package/src/facilitator.js +21 -21
- package/src/inbox-poller.js +4 -4
- package/src/judge.js +32 -30
- package/src/message-bus.js +12 -11
- package/src/orchestration-loop.js +35 -36
- package/src/orchestration-toolkit.js +58 -53
- package/src/orchestrator-helpers.js +2 -2
- package/src/profile-prompt.js +54 -53
- package/src/redaction.js +63 -57
- package/src/render/line-renderer.js +5 -5
- package/src/render/orchestrator-filter.js +3 -3
- package/src/render/palette.js +11 -9
- package/src/render/tool-hints.js +18 -15
- package/src/render/turn-renderer.js +4 -4
- package/src/reply-emitter.js +2 -2
- package/src/sequence-counter.js +4 -3
- package/src/signature-filter.js +7 -6
- package/src/supervisor.js +19 -18
- package/src/tee-writer.js +25 -25
- package/src/trace-collector.js +53 -48
- package/src/trace-github.js +53 -44
- package/src/trace-multi.js +16 -14
- package/src/trace-query.js +61 -52
- package/src/trace-render.js +19 -19
- package/src/trace-usage.js +31 -28
- package/src/transcript-recorder.js +24 -20
- package/bin/fit-benchmark.js +0 -44
- package/bin/fit-harness.js +0 -412
- package/bin/fit-selfedit.js +0 -165
- package/bin/fit-trace.js +0 -520
package/README.md
CHANGED
|
@@ -8,22 +8,23 @@ worked.
|
|
|
8
8
|
|
|
9
9
|
<!-- END:description -->
|
|
10
10
|
|
|
11
|
-
`libharness` provides the runtime and tool surface for multi-LLM coordination
|
|
12
|
-
|
|
13
|
-
an asynchronous discussion
|
|
14
|
-
traces they produce, and edits skill files under
|
|
11
|
+
`libharness` provides the runtime and tool surface for multi-LLM coordination.
|
|
12
|
+
An agent talks to a supervisor. A facilitator chairs a meeting. A lead drives
|
|
13
|
+
an asynchronous discussion. `libharness` also provides a CLI suite. The suite
|
|
14
|
+
runs evals, queries the traces they produce, and edits skill files under
|
|
15
|
+
controlled conditions.
|
|
15
16
|
|
|
16
17
|
## CLIs
|
|
17
18
|
|
|
18
19
|
| CLI | Purpose |
|
|
19
20
|
| --------------- | ---------------------------------------------------------------------- |
|
|
20
|
-
| `
|
|
21
|
-
| `
|
|
22
|
-
| `
|
|
23
|
-
| `
|
|
21
|
+
| `gemba-harness` | Run agents in `run`/`supervise`/`facilitate`/`discuss` subcommands. |
|
|
22
|
+
| `gemba-trace` | Download, query, and analyze the NDJSON traces `gemba-harness` produces. |
|
|
23
|
+
| `gemba-benchmark` | Run task families for N runs each and aggregate pass@k. |
|
|
24
|
+
| `gemba-selfedit` | Write stdin to `.claude/**` paths behind settings.json and branch gates. |
|
|
24
25
|
|
|
25
|
-
`
|
|
26
|
-
surface, below. The `judge` role is a profile
|
|
26
|
+
`gemba-harness`'s subcommands share one orchestration loop and one async tool
|
|
27
|
+
surface, below. The `judge` role is a profile you pass to `supervise`.
|
|
27
28
|
|
|
28
29
|
## Modes
|
|
29
30
|
|
|
@@ -36,8 +37,8 @@ surface, below. The `judge` role is a profile passed to `supervise`.
|
|
|
36
37
|
| `judge` | `judge` | (none) | `Conclude` |
|
|
37
38
|
|
|
38
39
|
`run` and `judge` are one-shot. The other three share `OrchestrationLoop`
|
|
39
|
-
plus an async Ask/Answer/Announce/RollCall tool surface
|
|
40
|
-
messages out over an in-memory bus
|
|
40
|
+
plus an async Ask/Answer/Announce/RollCall tool surface. The loop fans
|
|
41
|
+
messages out over an in-memory bus. It emits a `{source, seq, event}`
|
|
41
42
|
NDJSON envelope for every line.
|
|
42
43
|
|
|
43
44
|
## Async Ask / Answer / Announce
|
|
@@ -48,10 +49,10 @@ Answer({ message, askId? }) → routed to the asker
|
|
|
48
49
|
Announce({ message }) → broadcast, no reply expected
|
|
49
50
|
```
|
|
50
51
|
|
|
51
|
-
Every Ask returns immediately
|
|
52
|
+
Every Ask returns immediately. It registers a pending entry under an
|
|
52
53
|
`askId`. The reply arrives later on the asker's inbox as `[answer#N]
|
|
53
54
|
<participant>: <text>`. Broadcast: omit `to` on a multi-participant
|
|
54
|
-
lead. Answer's `askId` is optional
|
|
55
|
+
lead. Answer's `askId` is optional. The handler is forgiving:
|
|
55
56
|
|
|
56
57
|
- **Provided + matches an ask owed by the caller** → routes to that asker.
|
|
57
58
|
- **Provided but unknown or wrong addressee** → `isError` with a pointed
|
|
@@ -69,34 +70,35 @@ Inbox lines on resume:
|
|
|
69
70
|
```
|
|
70
71
|
|
|
71
72
|
Async means the lead can issue Asks, end its turn, and plan in the gap
|
|
72
|
-
while participants work in parallel
|
|
73
|
+
while participants work in parallel. Nothing blocks the LLM thread.
|
|
73
74
|
|
|
74
75
|
### Discuss-mode replies
|
|
75
76
|
|
|
76
|
-
In discussion mode, Answer calls routed to the lead
|
|
77
|
-
|
|
78
|
-
a separate reply
|
|
79
|
-
lead and agents can also call
|
|
80
|
-
directly to the thread (status
|
|
81
|
-
The message bus intercepts answers
|
|
77
|
+
In discussion mode, Answer calls routed to the lead stream to the
|
|
78
|
+
discussion thread as the agents produce them. Each agent's Answer becomes
|
|
79
|
+
a separate reply. The thread receives it immediately. The session does not
|
|
80
|
+
batch the replies at its end. The lead and agents can also call
|
|
81
|
+
`Acknowledge` to post brief messages directly to the thread (status
|
|
82
|
+
updates, human follow-up responses). The message bus intercepts answers
|
|
83
|
+
and appends them to `ctx.replies[]`.
|
|
82
84
|
|
|
83
85
|
`RequestForComment` is a separate coordination tool available on agent
|
|
84
86
|
roles (facilitated agents and discuss agents). It queues an intent to
|
|
85
87
|
open a new Discussion thread for long-horizon coordination on open
|
|
86
|
-
questions
|
|
87
|
-
thread replies in `ctx.replies[]`.
|
|
88
|
+
questions. These intents accumulate in `ctx.rfcs[]`. They stay separate
|
|
89
|
+
from the thread replies in `ctx.replies[]`.
|
|
88
90
|
|
|
89
91
|
## Orchestration loop
|
|
90
92
|
|
|
91
|
-
Each participant drains the bus
|
|
92
|
-
drained messages as tagged lines
|
|
93
|
-
one synthetic reminder
|
|
94
|
-
|
|
93
|
+
Each participant drains the bus, or waits. It then runs or resumes the LLM
|
|
94
|
+
with the drained messages as tagged lines. On an unanswered owed Ask, the
|
|
95
|
+
participant injects one synthetic reminder. It then emits
|
|
96
|
+
`protocol_violation`. It unblocks the asker with a synthetic null answer.
|
|
95
97
|
|
|
96
98
|
Termination uses two flags. `ctx.concluded` is explicit
|
|
97
|
-
`Conclude`/`Adjourn`/`Recess
|
|
98
|
-
see why
|
|
99
|
-
error, agent crash, abort path. Loops watch `stopped
|
|
99
|
+
`Conclude`/`Adjourn`/`Recess`. It also cancels in-flight Asks, so askers
|
|
100
|
+
see why nobody will answer their question. `stopped` is broader: lead
|
|
101
|
+
error, agent crash, abort path. Loops watch `stopped`. `ctx.concluded`
|
|
100
102
|
only feeds the summary's `success`/`verdict`.
|
|
101
103
|
|
|
102
104
|
## Tool surface, by role
|
|
@@ -113,7 +115,7 @@ only feeds the summary's `success`/`verdict`.
|
|
|
113
115
|
|
|
114
116
|
Ask's `to` accepts a participant name on multi-participant roles
|
|
115
117
|
(facilitator, discuss lead, all participants). The supervise pair has
|
|
116
|
-
only one possible target so `to
|
|
118
|
+
only one possible target, so it rejects `to`.
|
|
117
119
|
|
|
118
120
|
## Minimal example: two-participant facilitator
|
|
119
121
|
|
|
@@ -137,77 +139,78 @@ const result = await facilitator.run("Run a kata storyboard meeting.");
|
|
|
137
139
|
// result.success / result.turns / NDJSON trace on process.stdout
|
|
138
140
|
```
|
|
139
141
|
|
|
140
|
-
The facilitator gets `Ask`/`Answer`/`Announce`/`RollCall`/`Conclude
|
|
141
|
-
|
|
142
|
+
The facilitator gets `Ask`/`Answer`/`Announce`/`RollCall`/`Conclude`.
|
|
143
|
+
Each agent gets the same minus `Conclude`. Every tool call, bus
|
|
142
144
|
message, and orchestrator event becomes one trace line.
|
|
143
145
|
|
|
144
146
|
## Trace format and redaction
|
|
145
147
|
|
|
146
148
|
Each line is `{ "source": "<participant|orchestrator>", "seq": N, "event":
|
|
147
|
-
{…} }`. `seq` is monotonic across the whole trace
|
|
149
|
+
{…} }`. `seq` is monotonic across the whole trace. `orchestrator` emits
|
|
148
150
|
`session_start`, `agent_start`, `protocol_violation`, `lead_turn_limit`,
|
|
149
151
|
and `summary`. `event` is the SDK event verbatim or the orchestrator
|
|
150
|
-
payload. `
|
|
152
|
+
payload. `gemba-trace` consumes this format.
|
|
151
153
|
|
|
152
|
-
Redaction is on by default for `
|
|
154
|
+
Redaction is on by default for `gemba-harness run`/`supervise`/`facilitate`
|
|
153
155
|
and composes two layers:
|
|
154
156
|
|
|
155
157
|
- **Env-var allowlist** — `ANTHROPIC_API_KEY`, `GH_TOKEN`, `GITHUB_TOKEN`
|
|
156
|
-
by default
|
|
157
|
-
|
|
158
|
-
|
|
158
|
+
by default. Override the list with
|
|
159
|
+
`LIBHARNESS_REDACTION_ENV_VARS=NAME1,…`. The variable replaces the
|
|
160
|
+
default list. It does not extend the list. Runtime values become
|
|
161
|
+
`[REDACTED:env:NAME]` everywhere they appear.
|
|
159
162
|
- **Credential-shape patterns** — `sk-ant-`, `ghp_`, `ghs_`, `gho_`,
|
|
160
163
|
`github_pat_`. Hits become `[REDACTED:pattern:KIND]`.
|
|
161
164
|
|
|
162
|
-
Set `LIBHARNESS_REDACTION_DISABLED=1` to disable (one stderr warning
|
|
163
|
-
run). Never on CI for a public repo
|
|
164
|
-
downloadable through retention.
|
|
165
|
+
Set `LIBHARNESS_REDACTION_DISABLED=1` to disable it (one stderr warning
|
|
166
|
+
per run). Never disable it on CI for a public repo. Workflow artifacts
|
|
167
|
+
stay downloadable through retention.
|
|
165
168
|
|
|
166
169
|
## Module map
|
|
167
170
|
|
|
168
171
|
| Module | Purpose |
|
|
169
172
|
| ----------------------------------------------------------- | -------------------------------------------------------------------- |
|
|
170
|
-
| `agent-runner.js` | One Claude Agent SDK session
|
|
173
|
+
| `agent-runner.js` | One Claude Agent SDK session. Emits NDJSON through the redactor. |
|
|
171
174
|
| `message-bus.js` | Per-participant queues + `waitForMessages` Promise wakeup. |
|
|
172
175
|
| `orchestration-toolkit.js` | Shared Ask/Answer/Announce/Conclude/RollCall/RequestForComment handlers + builders. |
|
|
173
|
-
| `orchestration-loop.js` | Unified lead+participant loop
|
|
176
|
+
| `orchestration-loop.js` | Unified lead+participant loop. Handles reminders and violations. |
|
|
174
177
|
| `facilitator.js` / `supervisor.js` / `discusser.js` / `judge.js` | Per-mode class + factory + system prompt. |
|
|
175
178
|
| `discuss-tools.js` | Discuss-only `Recess`/`Adjourn`/`Acknowledge`. |
|
|
176
179
|
| `reply-emitter.js` | Fire-and-forget POST of reply/ack events to the callback URL. |
|
|
177
180
|
| `inbox-poller.js` | Long-poll the bridge inbox for injected human messages. |
|
|
178
|
-
| `trace-collector.js` / `trace-query.js` / `trace-github.js` | Trace ingestion /
|
|
181
|
+
| `trace-collector.js` / `trace-query.js` / `trace-github.js` | Trace ingestion / query / GitHub-attachment helpers. |
|
|
179
182
|
| `redaction.js` | Env-var allowlist + credential-shape pattern redaction. |
|
|
180
183
|
|
|
181
|
-
##
|
|
184
|
+
## gemba-selfedit
|
|
182
185
|
|
|
183
|
-
A narrow, audited bypass for sessions
|
|
184
|
-
writes)
|
|
185
|
-
|
|
186
|
-
|
|
186
|
+
A narrow, audited bypass for sessions that block `Edit`/`Write` (and bash
|
|
187
|
+
writes) against paths the project's own allowlist permits. It reads stdin.
|
|
188
|
+
It writes the target. It exits 0, 2 (safeguard violation), or 1 (I/O
|
|
189
|
+
error).
|
|
187
190
|
|
|
188
191
|
```sh
|
|
189
|
-
echo "<content>" | bunx
|
|
192
|
+
echo "<content>" | bunx gemba-selfedit <path>
|
|
190
193
|
```
|
|
191
194
|
|
|
192
|
-
|
|
195
|
+
The CLI checks two safeguards in order:
|
|
193
196
|
|
|
194
197
|
1. **Settings-allow.** Walk upward from the target with
|
|
195
198
|
[`Finder.findUpward`](../libutil/src/finder.js) to find the nearest
|
|
196
199
|
`.claude/settings.json`. The target relative to its grandparent
|
|
197
200
|
directory must match at least one `Edit(<glob>)` rule in
|
|
198
|
-
`permissions.allow[]
|
|
199
|
-
[`minimatch`](https://github.com/isaacs/minimatch)
|
|
200
|
-
Settings.json is the single source of truth
|
|
201
|
-
allowlist and the CLI follows.
|
|
202
|
-
|
|
203
|
-
|
|
201
|
+
`permissions.allow[]`. The CLI matches with
|
|
202
|
+
[`minimatch`](https://github.com/isaacs/minimatch) and `dot: true`.
|
|
203
|
+
Settings.json is the single source of truth. Widen the project
|
|
204
|
+
allowlist and the CLI follows. The CLI also rejects traversal like
|
|
205
|
+
`.claude/../README.md` as a side effect. `path.resolve` collapses
|
|
206
|
+
`..` first. The resolved path then tests against the rules.
|
|
204
207
|
|
|
205
208
|
2. **Branch scope.** `git rev-parse --abbrev-ref HEAD` must not be
|
|
206
209
|
`HEAD` (detached) or `main`. Edits ride a feature branch through
|
|
207
|
-
whatever merge gates the project
|
|
210
|
+
whatever merge gates the project configured.
|
|
208
211
|
|
|
209
|
-
Failure messages name the safeguard that rejected
|
|
210
|
-
lists the `Edit()` rules
|
|
212
|
+
Failure messages name the safeguard that rejected the write. Safeguard 1
|
|
213
|
+
also lists the `Edit()` rules it tried.
|
|
211
214
|
|
|
212
215
|
## Documentation
|
|
213
216
|
|
|
@@ -216,12 +219,12 @@ lists the `Edit()` rules that were tried.
|
|
|
216
219
|
facilitate / discuss) with Ask/Answer/Announce and a single NDJSON trace.
|
|
217
220
|
- [Run an Eval](https://www.forwardimpact.team/docs/libraries/prove-changes/run-eval/index.md)
|
|
218
221
|
— author a judge profile, run an eval locally, wire it into CI, and inspect
|
|
219
|
-
the
|
|
222
|
+
the trace it produces.
|
|
220
223
|
- [Prove Agent Changes](https://www.forwardimpact.team/docs/libraries/prove-changes/index.md)
|
|
221
|
-
— end-to-end workflow from dataset generation through evaluation to
|
|
222
|
-
analysis,
|
|
224
|
+
— the end-to-end workflow from dataset generation through evaluation to
|
|
225
|
+
trace analysis, with multi-agent collaboration sessions.
|
|
223
226
|
- [Analyze Traces](https://www.forwardimpact.team/docs/libraries/prove-changes/trace-analysis/index.md)
|
|
224
|
-
— read the NDJSON traces produced by `
|
|
227
|
+
— read the NDJSON traces produced by `gemba-harness` with `gemba-trace`.
|
|
225
228
|
- [Agent Teams](https://www.forwardimpact.team/docs/products/agent-teams/index.md)
|
|
226
229
|
— author the profiles consumed by `--agent-profile`, `--lead-profile`, and
|
|
227
230
|
`--agent-profiles`.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@forwardimpact/libharness",
|
|
3
|
-
"version": "
|
|
3
|
+
"version": "3.0.1",
|
|
4
4
|
"description": "Autonomous agent team harness — coordinate a lead and participant agents in one async session, with eval, benchmark, and trace tooling to prove the changes worked.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"orchestration",
|
|
@@ -40,20 +40,22 @@
|
|
|
40
40
|
"main": "./src/index.js",
|
|
41
41
|
"exports": {
|
|
42
42
|
".": "./src/index.js",
|
|
43
|
-
"./
|
|
44
|
-
"./
|
|
45
|
-
"./
|
|
46
|
-
"./
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
"
|
|
50
|
-
"
|
|
51
|
-
"
|
|
52
|
-
"
|
|
43
|
+
"./commands/output.js": "./src/commands/output.js",
|
|
44
|
+
"./commands/tee.js": "./src/commands/tee.js",
|
|
45
|
+
"./commands/run.js": "./src/commands/run.js",
|
|
46
|
+
"./commands/supervise.js": "./src/commands/supervise.js",
|
|
47
|
+
"./commands/facilitate.js": "./src/commands/facilitate.js",
|
|
48
|
+
"./commands/discuss.js": "./src/commands/discuss.js",
|
|
49
|
+
"./commands/callback.js": "./src/commands/callback.js",
|
|
50
|
+
"./commands/scan-logs.js": "./src/commands/scan-logs.js",
|
|
51
|
+
"./commands/trace.js": "./src/commands/trace.js",
|
|
52
|
+
"./commands/assert.js": "./src/commands/assert.js",
|
|
53
|
+
"./commands/by-discussion.js": "./src/commands/by-discussion.js",
|
|
54
|
+
"./commands/benchmark-definition.js": "./src/commands/benchmark-definition.js",
|
|
55
|
+
"./commands/selfedit.js": "./src/commands/selfedit.js"
|
|
53
56
|
},
|
|
54
57
|
"files": [
|
|
55
58
|
"src/**/*.js",
|
|
56
|
-
"bin/**/*.js",
|
|
57
59
|
"README.md"
|
|
58
60
|
],
|
|
59
61
|
"scripts": {
|
|
@@ -67,7 +69,7 @@
|
|
|
67
69
|
"@forwardimpact/libtelemetry": "^0.1.22",
|
|
68
70
|
"@forwardimpact/libutil": "^0.1.0",
|
|
69
71
|
"jmespath": "^0.16.0",
|
|
70
|
-
"minimatch": "^10.
|
|
72
|
+
"minimatch": "^10.2.6",
|
|
71
73
|
"zod": "^4.4.3"
|
|
72
74
|
},
|
|
73
75
|
"devDependencies": {
|
package/src/advisor.js
CHANGED
|
@@ -1,15 +1,15 @@
|
|
|
1
1
|
/**
|
|
2
|
-
* Advisor — the judge's mid-loop sibling
|
|
3
|
-
* `AgentRunner` session on a stronger model
|
|
4
|
-
* Each consult forwards the caller's whole recorded context (system
|
|
5
|
-
* delivered prompts, transcript so far) plus a focused question
|
|
6
|
-
* can inspect files read-only
|
|
7
|
-
* orchestration tools
|
|
2
|
+
* Advisor — the judge's mid-loop sibling. It is a solo, tool-restricted,
|
|
3
|
+
* one-shot `AgentRunner` session on a stronger model. Its final text is the
|
|
4
|
+
* advice. Each consult forwards the caller's whole recorded context (system
|
|
5
|
+
* prompt, delivered prompts, transcript so far) plus a focused question. The
|
|
6
|
+
* advisor can inspect files read-only. It holds no write, execute, subagent,
|
|
7
|
+
* or orchestration tools. It never appears on the message bus.
|
|
8
8
|
*
|
|
9
|
-
* Consults are stateless
|
|
10
|
-
* caller's context as it stands
|
|
11
|
-
* all resolve to an in-band `{unavailable}` result so
|
|
12
|
-
* never stalls or crashes on a consult.
|
|
9
|
+
* Consults are stateless. Each call runs one fresh session that reads the
|
|
10
|
+
* caller's context as it stands. Consults also fail open. A timeout, an
|
|
11
|
+
* error, and an abort all resolve to an in-band `{unavailable}` result, so
|
|
12
|
+
* the caller's session never stalls or crashes on a consult.
|
|
13
13
|
*
|
|
14
14
|
* Follows OO+DI: factory function, tests inject a fake `query`.
|
|
15
15
|
*/
|
|
@@ -20,39 +20,42 @@ import { createAgentRunner } from "./agent-runner.js";
|
|
|
20
20
|
import { composeSystemPrompt } from "./profile-prompt.js";
|
|
21
21
|
|
|
22
22
|
/**
|
|
23
|
-
* System-prompt trailer for the advisor session.
|
|
24
|
-
* contract (spec criterion "Advice is bounded")
|
|
25
|
-
* recommendation, unsolicited findings, with a stated
|
|
23
|
+
* System-prompt trailer for the advisor session. It fixes the response
|
|
24
|
+
* contract (spec criterion "Advice is bounded"). The contract is an
|
|
25
|
+
* assessment, a recommendation, and unsolicited findings, with a stated
|
|
26
|
+
* length ceiling.
|
|
26
27
|
*/
|
|
27
28
|
export const ADVISOR_SYSTEM_PROMPT =
|
|
28
|
-
"You are a consulted specialist
|
|
29
|
-
"Another agent paused its work to ask you one question
|
|
30
|
-
"You may Read, Glob, and Grep the files the transcript names to ground your advice
|
|
31
|
-
"Respond in one turn of prose
|
|
29
|
+
"You are a consulted specialist. You are not a worker. " +
|
|
30
|
+
"Another agent paused its work to ask you one question. Its full session context and the question are in the task. " +
|
|
31
|
+
"You may Read, Glob, and Grep the files the transcript names to ground your advice. You must never modify anything. " +
|
|
32
|
+
"Respond in one turn of prose. The caller receives your final text verbatim. " +
|
|
32
33
|
"Structure the response as: assessment (what you see), recommendation (what to do and why), and unsolicited findings (anything important the caller did not ask about). " +
|
|
33
34
|
"Keep the whole response to at most three short paragraphs. " +
|
|
34
|
-
"Do not ask follow-up questions
|
|
35
|
+
"Do not ask follow-up questions. The caller cannot reply.";
|
|
35
36
|
|
|
36
37
|
/**
|
|
37
38
|
* Consult-guidance fragment for caller system prompts, present only when
|
|
38
|
-
* the session runs with an advisor model.
|
|
39
|
-
* mandates nothing.
|
|
39
|
+
* the session runs with an advisor model. It steers the caller's judgment.
|
|
40
|
+
* It mandates nothing.
|
|
40
41
|
* @param {number} maxUses - The session-wide consult budget.
|
|
41
42
|
* @returns {string}
|
|
42
43
|
*/
|
|
43
44
|
export function advisorGuidance(maxUses) {
|
|
44
45
|
return (
|
|
45
|
-
"An `Advisor` tool is available
|
|
46
|
-
"A consult pays off at hard decision points
|
|
46
|
+
"An `Advisor` tool is available. It takes one focused question per call. A stronger model that sees your full session context answers it. " +
|
|
47
|
+
"A consult pays off at hard decision points, such as architectural forks, unclear root causes, and trade-offs you cannot rank. " +
|
|
48
|
+
"A consult also pays off early, before work builds on an unvalidated assumption. " +
|
|
47
49
|
"It does not pay off for routine reads, writes, or searches. " +
|
|
48
|
-
`The session-wide budget is ${maxUses} consult${maxUses === 1 ? "" : "s"},
|
|
49
|
-
"
|
|
50
|
+
`The session-wide budget is ${maxUses} consult${maxUses === 1 ? "" : "s"}, which all participants share. ` +
|
|
51
|
+
"A consult is your judgment. It is never mandatory."
|
|
50
52
|
);
|
|
51
53
|
}
|
|
52
54
|
|
|
53
55
|
/**
|
|
54
|
-
* Create the session-wide consult budget
|
|
55
|
-
*
|
|
56
|
+
* Create the session-wide consult budget. Every caller's tool handler
|
|
57
|
+
* shares it. The tool handler enforces the budget in code. The prompt does
|
|
58
|
+
* not enforce it.
|
|
56
59
|
* @param {number} maxUses
|
|
57
60
|
* @returns {{maxUses: number, used: number}}
|
|
58
61
|
*/
|
|
@@ -62,7 +65,7 @@ export function createAdvisorBudget(maxUses) {
|
|
|
62
65
|
|
|
63
66
|
/**
|
|
64
67
|
* Fold the consult guidance into an existing run-specific amendment when
|
|
65
|
-
* the advisor is enabled (budget present)
|
|
68
|
+
* the advisor is enabled (budget present). Return the amendment unchanged
|
|
66
69
|
* otherwise, so advisor-off prompts stay byte-identical.
|
|
67
70
|
* @param {string|undefined} amend - The existing amendment, if any.
|
|
68
71
|
* @param {{maxUses: number}|null} budget
|
|
@@ -73,16 +76,19 @@ export function withAdvisorGuidance(amend, budget) {
|
|
|
73
76
|
return [amend, advisorGuidance(budget.maxUses)].filter(Boolean).join("\n\n");
|
|
74
77
|
}
|
|
75
78
|
|
|
76
|
-
/**
|
|
79
|
+
/**
|
|
80
|
+
* Consult timeout. It is generous for a read-a-few-files-and-answer
|
|
81
|
+
* session. It is also the universal guard in modes with no stop path.
|
|
82
|
+
*/
|
|
77
83
|
export const DEFAULT_CONSULT_TIMEOUT_MS = 300_000;
|
|
78
84
|
|
|
79
85
|
const ADVISOR_ALLOWED_TOOLS = ["Read", "Glob", "Grep"];
|
|
80
86
|
|
|
81
87
|
// Under the harness's always-on bypassPermissions, `allowedTools` alone is
|
|
82
|
-
// not structural
|
|
83
|
-
//
|
|
84
|
-
//
|
|
85
|
-
// list also removes the write-capable and non-inspection built-ins the
|
|
88
|
+
// not structural. `disallowedTools` removes tools from the model's context.
|
|
89
|
+
// The lead runners get the same treatment. The advisor's contract is
|
|
90
|
+
// stricter than the leads' (read-only inspection, nothing else). So the
|
|
91
|
+
// list also removes the write-capable and non-inspection built-ins that the
|
|
86
92
|
// lead convention leaves in.
|
|
87
93
|
const ADVISOR_DISALLOWED_TOOLS = [
|
|
88
94
|
"Bash",
|
|
@@ -105,18 +111,18 @@ const devNull = new Writable({
|
|
|
105
111
|
});
|
|
106
112
|
|
|
107
113
|
/**
|
|
108
|
-
* Create a per-caller advisor
|
|
114
|
+
* Create a per-caller advisor that closes over that caller's transcript
|
|
109
115
|
* recorder.
|
|
110
116
|
*
|
|
111
117
|
* @param {object} deps
|
|
112
118
|
* @param {string} deps.model - Advisor model id.
|
|
113
|
-
* @param {string} deps.cwd - The caller's working directory, so read-only
|
|
114
|
-
* @param {function} deps.query - SDK query function (
|
|
119
|
+
* @param {string} deps.cwd - The caller's working directory, so the read-only tools see the caller's files.
|
|
120
|
+
* @param {function} deps.query - SDK query function (tests inject it).
|
|
115
121
|
* @param {{render: () => string}} deps.recorder - The caller's transcript recorder.
|
|
116
122
|
* @param {import("./redaction.js").Redactor} deps.redactor
|
|
117
123
|
* @param {import("@forwardimpact/libutil/runtime").Runtime} deps.runtime - Clock surface for timeout and duration.
|
|
118
|
-
* @param {function} deps.onLine - Re-emitter for the advisor session's NDJSON lines (
|
|
119
|
-
* @param {number} [deps.maxTurns] - Default 5
|
|
124
|
+
* @param {function} deps.onLine - Re-emitter for the advisor session's NDJSON lines (the caller tags them `source: "advisor"`).
|
|
125
|
+
* @param {number} [deps.maxTurns] - Default 5, which is single-digit per the spec criterion.
|
|
120
126
|
* @param {number} [deps.timeoutMs] - Default `DEFAULT_CONSULT_TIMEOUT_MS`.
|
|
121
127
|
* @returns {{consult: (question: string) => Promise<{advice?: string, unavailable?: boolean, reason?: string, durationMs: number}>, abort: () => void}}
|
|
122
128
|
*/
|
|
@@ -147,7 +153,7 @@ export function createAdvisor({
|
|
|
147
153
|
return {
|
|
148
154
|
/**
|
|
149
155
|
* Run one fresh advisor session over the caller's context as it
|
|
150
|
-
* stands plus the question.
|
|
156
|
+
* stands plus the question. It never rejects. Every failure shape
|
|
151
157
|
* resolves to `{unavailable, reason}` (fail-open).
|
|
152
158
|
* @param {string} question
|
|
153
159
|
*/
|
|
@@ -207,9 +213,9 @@ export function createAdvisor({
|
|
|
207
213
|
},
|
|
208
214
|
|
|
209
215
|
/**
|
|
210
|
-
* Abort the in-flight consult, if any. A consult is a
|
|
211
|
-
*
|
|
212
|
-
*
|
|
216
|
+
* Abort the in-flight consult, if any. A consult is a tool call that
|
|
217
|
+
* blocks, so one caller cannot overlap its own consults. Each advisor
|
|
218
|
+
* belongs to one caller, so the code tracks at most one runner.
|
|
213
219
|
*/
|
|
214
220
|
abort() {
|
|
215
221
|
currentRunner?.currentAbortController?.abort();
|