@withgauge/cli 0.14.0 → 0.15.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +189 -53
- package/dist/index.js +2327 -1636
- package/package.json +2 -1
package/README.md
CHANGED
|
@@ -49,6 +49,8 @@ never drifts from the commands it describes.
|
|
|
49
49
|
|
|
50
50
|
Use `gauge instructions evals` before creating evals: it covers developer tasks,
|
|
51
51
|
model selection, observable judging criteria, and a complete migration example.
|
|
52
|
+
Use `gauge instructions preference` to investigate product choices, source
|
|
53
|
+
exposure, and interventions, from individual traces to larger cohorts.
|
|
52
54
|
Use `gauge instructions optimization` for the full optimization workflow.
|
|
53
55
|
`gauge instructions --list` shows available topics.
|
|
54
56
|
`gauge instructions evals --install` installs the eval guidance as a local skill.
|
|
@@ -61,7 +63,7 @@ billing, usage, stats, query, and more.
|
|
|
61
63
|
## Run committed Markdown evals
|
|
62
64
|
|
|
63
65
|
File evals keep the task, run settings, and judged criteria in Git. Create
|
|
64
|
-
`evals/healthcheck.md` with YAML frontmatter and a Markdown task body:
|
|
66
|
+
`gauge-evals/healthcheck.md` with YAML frontmatter and a Markdown task body:
|
|
65
67
|
|
|
66
68
|
```markdown
|
|
67
69
|
---
|
|
@@ -74,66 +76,194 @@ config:
|
|
|
74
76
|
agents:
|
|
75
77
|
- agent: CODEX_CLI
|
|
76
78
|
---
|
|
79
|
+
|
|
77
80
|
Add a /health endpoint and run the relevant app checks.
|
|
78
81
|
```
|
|
79
82
|
|
|
80
|
-
|
|
83
|
+
Check your working-tree setup, then commit the files before planning and running:
|
|
81
84
|
|
|
82
85
|
```sh
|
|
83
|
-
gauge evals
|
|
84
|
-
gauge evals
|
|
86
|
+
gauge evals verify # offline; includes uncommitted working-tree edits
|
|
87
|
+
gauge evals verify -o json # versioned global/per-input yes/no/unknown checklist
|
|
88
|
+
gauge evals plan
|
|
89
|
+
gauge evals run --yes
|
|
85
90
|
gauge run-requests wait <request-id> -o json
|
|
86
91
|
```
|
|
87
92
|
|
|
88
|
-
`--files`
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
config
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
`
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
93
|
+
Cases default to `gauge-evals/**/*.md`. `--files` accepts quoted paths/globs and overrides
|
|
94
|
+
that convention for `plan` and `run`, which read committed Git HEAD, ignoring working-tree edits.
|
|
95
|
+
Each case owns its agents, models, consumer repository/ref, Persona and criteria.
|
|
96
|
+
The optional `config.inputs` selects named inputs; all attach by default.
|
|
97
|
+
|
|
98
|
+
`verify` reads the repository's current files, including non-ignored untracked files,
|
|
99
|
+
without login, network access, a PR, built CI artifacts, or sandbox compute. It checks
|
|
100
|
+
configuration, case input references, Git input paths, and recognizable workflow
|
|
101
|
+
checkout/upload declarations. An input-only repository with no cases is valid setup.
|
|
102
|
+
Each attempted global/per-input check reports `yes` (established locally), `no`
|
|
103
|
+
(a definite problem), or `unknown` (a specific static question could not be resolved,
|
|
104
|
+
such as an expression-based artifact name or a referenced reusable workflow).
|
|
105
|
+
Missing optional configuration and zero cases are informational findings, excluded
|
|
106
|
+
from checklist counts. Eligible workflow triggers are checked as declarations;
|
|
107
|
+
event filters and runtime conditions are not evaluated against a candidate PR.
|
|
108
|
+
Live input delivery, artifact contents, sandbox execution, App activation, and
|
|
109
|
+
account settings are outside this command's scope and do not create unknown rows.
|
|
110
|
+
Exit 1 means at least one `no`; exit 0 means no definite static problem was found.
|
|
111
|
+
No workflow code is executed. Commit files before `plan` or `run` consumes them.
|
|
112
|
+
|
|
113
|
+
`gauge evals verify -o json` writes exactly one JSON object to stdout, including on
|
|
114
|
+
validation failure. The versioned report provides:
|
|
115
|
+
|
|
116
|
+
- `version: 1` and `mode: "local-static"` identify the report contract and scope.
|
|
117
|
+
- `status` is `failed` if any check is `no`, otherwise `incomplete` if any check is
|
|
118
|
+
`unknown`, otherwise `passed`. It describes static checks only.
|
|
119
|
+
- `summary` contains `yes`, `no`, and `unknown` counts.
|
|
120
|
+
- `global` and `inputs[].checks` contain checks with stable `id`, `state`, and
|
|
121
|
+
explanatory `detail` fields; `inputs[].name` identifies each input.
|
|
122
|
+
- `facts.configFilePresent` and `facts.caseCount` expose discovered facts without
|
|
123
|
+
parsing prose. A `null` value means inspection did not reach that fact; zero cases
|
|
124
|
+
is distinct from case discovery being skipped or failing.
|
|
125
|
+
- `info` contains informational findings with stable `id` and `detail` fields.
|
|
126
|
+
- `notChecked` lists excluded capabilities as machine-readable identifiers.
|
|
127
|
+
|
|
128
|
+
For example, an agent can select unresolved checks with:
|
|
129
|
+
|
|
130
|
+
```sh
|
|
131
|
+
gauge evals verify -o json | jq '{status, facts, global: [.global[] | select(.state != "yes")], inputs: [.inputs[] | {name, checks: [.checks[] | select(.state != "yes")]}]}'
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
An optional `gauge.json` declares inputs produced by your normal CI or committed
|
|
135
|
+
alongside your product:
|
|
136
|
+
|
|
137
|
+
```json
|
|
138
|
+
{
|
|
139
|
+
"version": 2,
|
|
140
|
+
"org": "my-org",
|
|
141
|
+
"inputs": {
|
|
142
|
+
"cli": {
|
|
143
|
+
"type": "npm-package",
|
|
144
|
+
"source": { "type": "github-actions", "workflow": ".github/workflows/package.yml" },
|
|
145
|
+
"path": "*.tgz",
|
|
146
|
+
"install": "environment"
|
|
147
|
+
},
|
|
148
|
+
"docs": {
|
|
149
|
+
"type": "files",
|
|
150
|
+
"source": { "type": "git" },
|
|
151
|
+
"paths": ["docs/**/*.md"]
|
|
152
|
+
}
|
|
153
|
+
}
|
|
154
|
+
}
|
|
155
|
+
```
|
|
156
|
+
|
|
157
|
+
Gauge waits for the existing workflow at the pushed revision and consumes its
|
|
158
|
+
artifacts by ID. Omit `source.artifact` to select a sole artifact, or use a name
|
|
159
|
+
pattern when the workflow uploads several. It never triggers or performs a build. Connect the Gauge
|
|
160
|
+
GitHub App with repository access and Actions read permission. Plan previews CI
|
|
161
|
+
readiness; waiting requests create no sessions and incur no session debit.
|
|
162
|
+
Each input names its workflow. Gauge captures one successful run/attempt per
|
|
163
|
+
workflow before downloading any input. A committed org and matching GitHub App
|
|
164
|
+
installation enable automatic same-repository PR evaluations for repositories and
|
|
165
|
+
target branches enabled in Settings → GitHub (initially enabled for `main`).
|
|
166
|
+
A failed run or absent/expired artifact returns an actionable input failure.
|
|
167
|
+
Absent runs have a five-minute discovery grace period; known running or queued
|
|
168
|
+
CI can continue past an hour. Abandoned requests expire after seven days. Reruns retain captured bytes. CI metadata reports a revision; it
|
|
169
|
+
does not independently prove which checkout the workflow compiled.
|
|
170
|
+
|
|
171
|
+
To run automatic PR checks only when relevant files change, add `checks` to
|
|
172
|
+
`gauge.json`:
|
|
104
173
|
|
|
105
174
|
```json
|
|
106
|
-
|
|
175
|
+
"checks": {
|
|
176
|
+
"paths": ["agents/**", "docs.json", "snippets/**"]
|
|
177
|
+
}
|
|
107
178
|
```
|
|
108
179
|
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
|
|
132
|
-
`
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
|
|
180
|
+
Paths are repository-relative, case-sensitive globs; matching any pattern runs
|
|
181
|
+
the checks. `**` matches nested directories, including hidden files. Use 1–32
|
|
182
|
+
patterns; exclusions (`!`) are not supported. Changes to `gauge.json` or your
|
|
183
|
+
configured case paths always count. Deleted files and both names of renamed
|
|
184
|
+
files are included. Without `checks`, every eligible PR update runs as before.
|
|
185
|
+
|
|
186
|
+
The first check uses the PR diff. After a passing evaluation, later updates
|
|
187
|
+
compare against that evaluated commit, so unrelated follow-up edits can skip.
|
|
188
|
+
A failed or unfinished evaluation, a changed target branch, or an unavailable
|
|
189
|
+
or incomplete diff causes a run. Skipped checks start no sessions and reserve
|
|
190
|
+
no credits. GitHub's rerun action runs regardless of paths; explicit CLI runs
|
|
191
|
+
also ignore this automatic-check filter.
|
|
192
|
+
|
|
193
|
+
File inputs stay outside the consumer repository and their paths are provided to
|
|
194
|
+
the agent. npm packages install with ordinary npm using the VM's Node/npm runtime.
|
|
195
|
+
`install` defaults to `environment`: a separate tools directory with the package's
|
|
196
|
+
own declared commands on PATH, including login shells. `workspace` adds a consumer
|
|
197
|
+
application dependency and accepts libraries without commands. Transitive registry
|
|
198
|
+
dependency ranges remain live. Plan output shows the installation destination and,
|
|
199
|
+
when a previous matching preparation captured it, command metadata. Otherwise it
|
|
200
|
+
reports that commands will be determined during preparation; planning never downloads
|
|
201
|
+
archives or starts VMs.
|
|
202
|
+
|
|
203
|
+
Native executables use `type: "binary"`, with already unpacked files in a GitHub
|
|
204
|
+
Actions artifact (or small committed Git files):
|
|
205
|
+
|
|
206
|
+
```json
|
|
207
|
+
"gh": {
|
|
208
|
+
"type": "binary",
|
|
209
|
+
"source": {
|
|
210
|
+
"type": "github-actions",
|
|
211
|
+
"workflow": ".github/workflows/package.yml",
|
|
212
|
+
"artifact": "gh-linux"
|
|
213
|
+
},
|
|
214
|
+
"paths": ["dist/**"],
|
|
215
|
+
"platform": "linux-x64",
|
|
216
|
+
"commands": { "gh": "dist/bin/gh" }
|
|
217
|
+
}
|
|
218
|
+
```
|
|
219
|
+
|
|
220
|
+
Paths are relative to the uploaded artifact root. Build the exact candidate SHA;
|
|
221
|
+
for `actions/upload-artifact@v7`, use `archive: true`. Include companion files in
|
|
222
|
+
`paths` and explicitly map 1–32 command names to ordinary selected files. Gauge
|
|
223
|
+
validates Linux x86-64 ELF executables in its isolated artifact processor, restores
|
|
224
|
+
executable permissions, preserves the bundle outside the consumer repository, and
|
|
225
|
+
exposes the commands on PATH, including login shells. A candidate command overrides
|
|
226
|
+
an installed product command with the same name; two inputs selected by the same
|
|
227
|
+
case cannot expose the same command (including captured npm environment commands).
|
|
228
|
+
|
|
229
|
+
The first version supports ELF executables and PIE binaries compatible with the
|
|
230
|
+
consumer Linux runtime. It does not unpack nested release archives, install OS
|
|
231
|
+
packages, alter library search paths, or run arbitrary installation/version-check
|
|
232
|
+
commands. Prefer a static binary or build against a compatible Linux runtime;
|
|
233
|
+
header validation cannot prove shared-library compatibility. Scripts, macOS,
|
|
234
|
+
Windows, and ARM executables are unsupported. A missing loader/library encountered
|
|
235
|
+
when the agent invokes the command appears in session execution evidence, not in
|
|
236
|
+
static validation. Harness/runtime names (including shells, Node/npm, Python, Git,
|
|
237
|
+
archive/setup utilities, and agent commands) are reserved; use a different alias
|
|
238
|
+
for such a candidate. `GAUGE_INPUTS` gives the preserved bundle path.
|
|
239
|
+
|
|
240
|
+
A Skill input has `type: "skill"` and `path: "skills/product"`, or a path to a ZIP,
|
|
241
|
+
`.tar.gz`, or `.tgz` whose root contains `SKILL.md`. Use `path: "."` for a Skill at
|
|
242
|
+
the source root. Gauge validates and extracts selected archives in isolation and
|
|
243
|
+
stages the Skill in the coding agent's native discovery location. File inputs use
|
|
244
|
+
`paths` to select committed files or files inside an Actions artifact.
|
|
245
|
+
|
|
246
|
+
Website inputs use `type: "website"` and a canonical HTTPS origin in `url`.
|
|
247
|
+
The preview can come from `source: { "type": "github-deployment", "environment": "preview" }`
|
|
248
|
+
or a `github-pr-comment` source with an exact `author`, literal `bodyIncludes`
|
|
249
|
+
marker, and a readiness `pattern` with named `url` and `sha` captures. Gauge selects
|
|
250
|
+
candidate-specific evidence; timestamps alone cannot establish the revision.
|
|
251
|
+
Comments must belong to a unique open same-repository PR at the candidate head.
|
|
252
|
+
Previews must be public HTTPS origins, without authentication or path prefixes.
|
|
253
|
+
They are fetched live rather than snapshotted. See the [preview source examples and
|
|
254
|
+
limitations](../specs/docs-web-inputs.md), including Astro's existing bot comment.
|
|
255
|
+
|
|
256
|
+
Editor schemas are published at
|
|
257
|
+
`https://agents.withgauge.com/schemas/gauge/v2.4.0/gauge.json` and
|
|
258
|
+
`https://agents.withgauge.com/schemas/gauge/v2.4.0/case.json`. Associate them in editor
|
|
259
|
+
settings; `gauge.json` does not accept a `$schema` field. These release URLs are
|
|
260
|
+
immutable. Update earlier pilot configs by moving workflows into input sources,
|
|
261
|
+
replacing `use` with input `type` and `path`/`paths`, and moving cases into
|
|
262
|
+
`gauge-evals/` (or setting `cases` explicitly).
|
|
263
|
+
|
|
264
|
+
`--org` overrides the committed organization. Root `cases` can change discovery;
|
|
265
|
+
root `version` defaults to 2. Saved evals use `gauge evals run <id>`. Superusers
|
|
266
|
+
use platform funding by default and can add `--bill-to-org`.
|
|
137
267
|
|
|
138
268
|
## Measurements own their configuration
|
|
139
269
|
|
|
@@ -212,7 +342,6 @@ A scripted or agent-driven loop:
|
|
|
212
342
|
|
|
213
343
|
```sh
|
|
214
344
|
gauge optimizations diagnose --eval es_1 # where the sessions fail
|
|
215
|
-
gauge optimizations estimate --eval es_1 --target claude-code:claude-opus-4-8
|
|
216
345
|
gauge optimizations start --eval es_1 --target claude-code:claude-opus-4-8 --yes
|
|
217
346
|
gauge optimizations watch cl_1 # exit 4 = your turn
|
|
218
347
|
|
|
@@ -223,9 +352,10 @@ gauge optimizations trials add cl_1 \
|
|
|
223
352
|
--body-file start.md --launch --idempotency-key start-env-var --yes
|
|
224
353
|
|
|
225
354
|
gauge optimizations watch cl_1
|
|
226
|
-
gauge optimizations trial cl_1 R1 # score, verdict, sessions
|
|
355
|
+
gauge optimizations trial cl_1 R1 # score, verdict, sessions, errors
|
|
227
356
|
gauge optimizations change cl_1 R1 --body # the winning text, to paste
|
|
228
357
|
gauge optimizations adopt cl_1 R1 --yes # mark the winner; baseline stays fixed
|
|
358
|
+
gauge optimizations handoff cl_1 # instructions to land the winner in your source
|
|
229
359
|
gauge optimizations complete cl_1 --yes
|
|
230
360
|
```
|
|
231
361
|
|
|
@@ -260,8 +390,14 @@ the answer is you.
|
|
|
260
390
|
`--idempotency-key` makes `trials add` safe to retry: the same key returns the
|
|
261
391
|
trial the first call created, and never buys a second set of sessions.
|
|
262
392
|
|
|
263
|
-
|
|
264
|
-
|
|
393
|
+
There is no round limit or credit budget. Each round waits for you; `complete`
|
|
394
|
+
ends the optimization. Adoption marks the winner and leaves the baseline fixed.
|
|
395
|
+
It does not open a pull request. `handoff` prints the same instructions as the
|
|
396
|
+
web app's "Copy instructions for agent"; follow them in your own source.
|
|
397
|
+
|
|
398
|
+
A failed or timed-out session drops out of the score. `trial`, `get`, and
|
|
399
|
+
`watch` print a `WARNING` line for each one, with its failure reason; in JSON,
|
|
400
|
+
read `failedSessions` on the trial and `failureReason` on each run.
|
|
265
401
|
|
|
266
402
|
## Models, reports, and scripting
|
|
267
403
|
|