@withgauge/cli 0.14.0 → 0.15.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (3) hide show
  1. package/README.md +189 -53
  2. package/dist/index.js +2327 -1636
  3. package/package.json +2 -1
package/README.md CHANGED
@@ -49,6 +49,8 @@ never drifts from the commands it describes.
49
49
 
50
50
  Use `gauge instructions evals` before creating evals: it covers developer tasks,
51
51
  model selection, observable judging criteria, and a complete migration example.
52
+ Use `gauge instructions preference` to investigate product choices, source
53
+ exposure, and interventions, from individual traces to larger cohorts.
52
54
  Use `gauge instructions optimization` for the full optimization workflow.
53
55
  `gauge instructions --list` shows available topics.
54
56
  `gauge instructions evals --install` installs the eval guidance as a local skill.
@@ -61,7 +63,7 @@ billing, usage, stats, query, and more.
61
63
  ## Run committed Markdown evals
62
64
 
63
65
  File evals keep the task, run settings, and judged criteria in Git. Create
64
- `evals/healthcheck.md` with YAML frontmatter and a Markdown task body:
66
+ `gauge-evals/healthcheck.md` with YAML frontmatter and a Markdown task body:
65
67
 
66
68
  ```markdown
67
69
  ---
@@ -74,66 +76,194 @@ config:
74
76
  agents:
75
77
  - agent: CODEX_CLI
76
78
  ---
79
+
77
80
  Add a /health endpoint and run the relevant app checks.
78
81
  ```
79
82
 
80
- Commit the file, then preview and run its exact committed bytes:
83
+ Check your working-tree setup, then commit the files before planning and running:
81
84
 
82
85
  ```sh
83
- gauge evals plan --files evals/healthcheck.md
84
- gauge evals run --files evals/healthcheck.md --yes
86
+ gauge evals verify # offline; includes uncommitted working-tree edits
87
+ gauge evals verify -o json # versioned global/per-input yes/no/unknown checklist
88
+ gauge evals plan
89
+ gauge evals run --yes
85
90
  gauge run-requests wait <request-id> -o json
86
91
  ```
87
92
 
88
- `--files` takes explicit paths; pass several paths after the flag for a
89
- multi-case request. The CLI reads `HEAD` through Git, so uncommitted edits are
90
- ignored and untracked cases fail. It hashes the committed case bytes and root
91
- config and submits their commit, digest, and repository origin. The server
92
- currently records that provenance as **submitted, unverified**; it does not
93
- fetch Git independently. `plan` resolves the configuration and estimates
94
- sessions without launching. Each case owns a `config` object with its own
95
- agent roster; supported fields are `agents`, `sampleCount`, `repoUrl`,
96
- `repoRef`, `profileId`, `skillRefs`, `mcpRefs`, `connectionIds`, and optional
97
- `addons.browser` (`session`, positive `projectId`, optional `region: us`).
98
- Skills and MCP assets use pinned `name@label` references. At least one agent
99
- and one criterion are required for each case. Omitted models use the agent's
100
- default model; omitted `sampleCount` runs one sample. With no `repoUrl`, the
101
- existing default scaffold is used.
102
-
103
- An optional root `gauge.json` selects the organization and one product build:
93
+ Cases default to `gauge-evals/**/*.md`. `--files` accepts quoted paths/globs and overrides
94
+ that convention for `plan` and `run`, which read committed Git HEAD, ignoring working-tree edits.
95
+ Each case owns its agents, models, consumer repository/ref, Persona and criteria.
96
+ The optional `config.inputs` selects named inputs; all attach by default.
97
+
98
+ `verify` reads the repository's current files, including non-ignored untracked files,
99
+ without login, network access, a PR, built CI artifacts, or sandbox compute. It checks
100
+ configuration, case input references, Git input paths, and recognizable workflow
101
+ checkout/upload declarations. An input-only repository with no cases is valid setup.
102
+ Each attempted global/per-input check reports `yes` (established locally), `no`
103
+ (a definite problem), or `unknown` (a specific static question could not be resolved,
104
+ such as an expression-based artifact name or a referenced reusable workflow).
105
+ Missing optional configuration and zero cases are informational findings, excluded
106
+ from checklist counts. Eligible workflow triggers are checked as declarations;
107
+ event filters and runtime conditions are not evaluated against a candidate PR.
108
+ Live input delivery, artifact contents, sandbox execution, App activation, and
109
+ account settings are outside this command's scope and do not create unknown rows.
110
+ Exit 1 means at least one `no`; exit 0 means no definite static problem was found.
111
+ No workflow code is executed. Commit files before `plan` or `run` consumes them.
112
+
113
+ `gauge evals verify -o json` writes exactly one JSON object to stdout, including on
114
+ validation failure. The versioned report provides:
115
+
116
+ - `version: 1` and `mode: "local-static"` identify the report contract and scope.
117
+ - `status` is `failed` if any check is `no`, otherwise `incomplete` if any check is
118
+ `unknown`, otherwise `passed`. It describes static checks only.
119
+ - `summary` contains `yes`, `no`, and `unknown` counts.
120
+ - `global` and `inputs[].checks` contain checks with stable `id`, `state`, and
121
+ explanatory `detail` fields; `inputs[].name` identifies each input.
122
+ - `facts.configFilePresent` and `facts.caseCount` expose discovered facts without
123
+ parsing prose. A `null` value means inspection did not reach that fact; zero cases
124
+ is distinct from case discovery being skipped or failing.
125
+ - `info` contains informational findings with stable `id` and `detail` fields.
126
+ - `notChecked` lists excluded capabilities as machine-readable identifiers.
127
+
128
+ For example, an agent can select unresolved checks with:
129
+
130
+ ```sh
131
+ gauge evals verify -o json | jq '{status, facts, global: [.global[] | select(.state != "yes")], inputs: [.inputs[] | {name, checks: [.checks[] | select(.state != "yes")]}]}'
132
+ ```
133
+
134
+ An optional `gauge.json` declares inputs produced by your normal CI or committed
135
+ alongside your product:
136
+
137
+ ```json
138
+ {
139
+ "version": 2,
140
+ "org": "my-org",
141
+ "inputs": {
142
+ "cli": {
143
+ "type": "npm-package",
144
+ "source": { "type": "github-actions", "workflow": ".github/workflows/package.yml" },
145
+ "path": "*.tgz",
146
+ "install": "environment"
147
+ },
148
+ "docs": {
149
+ "type": "files",
150
+ "source": { "type": "git" },
151
+ "paths": ["docs/**/*.md"]
152
+ }
153
+ }
154
+ }
155
+ ```
156
+
157
+ Gauge waits for the existing workflow at the pushed revision and consumes its
158
+ artifacts by ID. Omit `source.artifact` to select a sole artifact, or use a name
159
+ pattern when the workflow uploads several. It never triggers or performs a build. Connect the Gauge
160
+ GitHub App with repository access and Actions read permission. Plan previews CI
161
+ readiness; waiting requests create no sessions and incur no session debit.
162
+ Each input names its workflow. Gauge captures one successful run/attempt per
163
+ workflow before downloading any input. A committed org and matching GitHub App
164
+ installation enable automatic same-repository PR evaluations for repositories and
165
+ target branches enabled in Settings → GitHub (initially enabled for `main`).
166
+ A failed run or absent/expired artifact returns an actionable input failure.
167
+ Absent runs have a five-minute discovery grace period; known running or queued
168
+ CI can continue past an hour. Abandoned requests expire after seven days. Reruns retain captured bytes. CI metadata reports a revision; it
169
+ does not independently prove which checkout the workflow compiled.
170
+
171
+ To run automatic PR checks only when relevant files change, add `checks` to
172
+ `gauge.json`:
104
173
 
105
174
  ```json
106
- {"org":"my-org","product":{"type":"npm-cli"}}
175
+ "checks": {
176
+ "paths": ["agents/**", "docs.json", "snippets/**"]
177
+ }
107
178
  ```
108
179
 
109
- For an npm CLI, Gauge uses the committed package manifest and lockfile, runs
110
- its `build` script when present, packs and installs the package with production
111
- dependencies, and exposes its declared CLI commands with the pinned Node
112
- runtime. Set `product.build` to a different build command when needed.
113
- The source install is frozen; Gauge resolves a fresh production install once per
114
- run request, archives that installed tree for every case, and records its
115
- runtime lockfile digest in the build receipt. New requests can resolve newer
116
- transitive dependencies until the product pins them.
117
- Each new CLI launch currently rebuilds, even at the same source SHA; bundles
118
- are shared across cases and reused after completed preparation within a request,
119
- but there is no cache across requests yet.
120
- The first adapter requires the pinned pnpm 10.34.5 package manager and a
121
- `pnpm-lock.yaml`.
122
- `--org` overrides the committed `org`; otherwise the CLI falls back to your
123
- configured default organization.
124
-
125
- For other products, `{"prepare":"pnpm run gauge:prepare","install":"./install.sh"}`
126
- remains available. `prepare` runs once at the exact definition commit with
127
- `GAUGE_OUTPUT_DIR` pointing to an empty directory; `install` is an optional
128
- bundle-relative script run once per consumer session. A committed
129
- `.gauge/prepare.sh` enables that preparation automatically when `prepare` is
130
- absent. `product` cannot be combined with either script form. The output is
131
- attached to every case. There is no authored resource manifest or per-case
132
- `inputs` field. Project-level run defaults are rejected. The optional
133
- `version` defaults to 1. Preparation requires a public HTTPS Git origin.
134
- Saved evals remain available with `gauge evals run <id>`.
135
- As with saved evals, superusers use platform funding by default and can add
136
- `--bill-to-org` to charge the selected organization.
180
+ Paths are repository-relative, case-sensitive globs; matching any pattern runs
181
+ the checks. `**` matches nested directories, including hidden files. Use 1–32
182
+ patterns; exclusions (`!`) are not supported. Changes to `gauge.json` or your
183
+ configured case paths always count. Deleted files and both names of renamed
184
+ files are included. Without `checks`, every eligible PR update runs as before.
185
+
186
+ The first check uses the PR diff. After a passing evaluation, later updates
187
+ compare against that evaluated commit, so unrelated follow-up edits can skip.
188
+ A failed or unfinished evaluation, a changed target branch, or an unavailable
189
+ or incomplete diff causes a run. Skipped checks start no sessions and reserve
190
+ no credits. GitHub's rerun action runs regardless of paths; explicit CLI runs
191
+ also ignore this automatic-check filter.
192
+
193
+ File inputs stay outside the consumer repository and their paths are provided to
194
+ the agent. npm packages install with ordinary npm using the VM's Node/npm runtime.
195
+ `install` defaults to `environment`: a separate tools directory with the package's
196
+ own declared commands on PATH, including login shells. `workspace` adds a consumer
197
+ application dependency and accepts libraries without commands. Transitive registry
198
+ dependency ranges remain live. Plan output shows the installation destination and,
199
+ when a previous matching preparation captured it, command metadata. Otherwise it
200
+ reports that commands will be determined during preparation; planning never downloads
201
+ archives or starts VMs.
202
+
203
+ Native executables use `type: "binary"`, with already unpacked files in a GitHub
204
+ Actions artifact (or small committed Git files):
205
+
206
+ ```json
207
+ "gh": {
208
+ "type": "binary",
209
+ "source": {
210
+ "type": "github-actions",
211
+ "workflow": ".github/workflows/package.yml",
212
+ "artifact": "gh-linux"
213
+ },
214
+ "paths": ["dist/**"],
215
+ "platform": "linux-x64",
216
+ "commands": { "gh": "dist/bin/gh" }
217
+ }
218
+ ```
219
+
220
+ Paths are relative to the uploaded artifact root. Build the exact candidate SHA;
221
+ for `actions/upload-artifact@v7`, use `archive: true`. Include companion files in
222
+ `paths` and explicitly map 1–32 command names to ordinary selected files. Gauge
223
+ validates Linux x86-64 ELF executables in its isolated artifact processor, restores
224
+ executable permissions, preserves the bundle outside the consumer repository, and
225
+ exposes the commands on PATH, including login shells. A candidate command overrides
226
+ an installed product command with the same name; two inputs selected by the same
227
+ case cannot expose the same command (including captured npm environment commands).
228
+
229
+ The first version supports ELF executables and PIE binaries compatible with the
230
+ consumer Linux runtime. It does not unpack nested release archives, install OS
231
+ packages, alter library search paths, or run arbitrary installation/version-check
232
+ commands. Prefer a static binary or build against a compatible Linux runtime;
233
+ header validation cannot prove shared-library compatibility. Scripts, macOS,
234
+ Windows, and ARM executables are unsupported. A missing loader/library encountered
235
+ when the agent invokes the command appears in session execution evidence, not in
236
+ static validation. Harness/runtime names (including shells, Node/npm, Python, Git,
237
+ archive/setup utilities, and agent commands) are reserved; use a different alias
238
+ for such a candidate. `GAUGE_INPUTS` gives the preserved bundle path.
239
+
240
+ A Skill input has `type: "skill"` and `path: "skills/product"`, or a path to a ZIP,
241
+ `.tar.gz`, or `.tgz` whose root contains `SKILL.md`. Use `path: "."` for a Skill at
242
+ the source root. Gauge validates and extracts selected archives in isolation and
243
+ stages the Skill in the coding agent's native discovery location. File inputs use
244
+ `paths` to select committed files or files inside an Actions artifact.
245
+
246
+ Website inputs use `type: "website"` and a canonical HTTPS origin in `url`.
247
+ The preview can come from `source: { "type": "github-deployment", "environment": "preview" }`
248
+ or a `github-pr-comment` source with an exact `author`, literal `bodyIncludes`
249
+ marker, and a readiness `pattern` with named `url` and `sha` captures. Gauge selects
250
+ candidate-specific evidence; timestamps alone cannot establish the revision.
251
+ Comments must belong to a unique open same-repository PR at the candidate head.
252
+ Previews must be public HTTPS origins, without authentication or path prefixes.
253
+ They are fetched live rather than snapshotted. See the [preview source examples and
254
+ limitations](../specs/docs-web-inputs.md), including Astro's existing bot comment.
255
+
256
+ Editor schemas are published at
257
+ `https://agents.withgauge.com/schemas/gauge/v2.4.0/gauge.json` and
258
+ `https://agents.withgauge.com/schemas/gauge/v2.4.0/case.json`. Associate them in editor
259
+ settings; `gauge.json` does not accept a `$schema` field. These release URLs are
260
+ immutable. Update earlier pilot configs by moving workflows into input sources,
261
+ replacing `use` with input `type` and `path`/`paths`, and moving cases into
262
+ `gauge-evals/` (or setting `cases` explicitly).
263
+
264
+ `--org` overrides the committed organization. Root `cases` can change discovery;
265
+ root `version` defaults to 2. Saved evals use `gauge evals run <id>`. Superusers
266
+ use platform funding by default and can add `--bill-to-org`.
137
267
 
138
268
  ## Measurements own their configuration
139
269
 
@@ -212,7 +342,6 @@ A scripted or agent-driven loop:
212
342
 
213
343
  ```sh
214
344
  gauge optimizations diagnose --eval es_1 # where the sessions fail
215
- gauge optimizations estimate --eval es_1 --target claude-code:claude-opus-4-8
216
345
  gauge optimizations start --eval es_1 --target claude-code:claude-opus-4-8 --yes
217
346
  gauge optimizations watch cl_1 # exit 4 = your turn
218
347
 
@@ -223,9 +352,10 @@ gauge optimizations trials add cl_1 \
223
352
  --body-file start.md --launch --idempotency-key start-env-var --yes
224
353
 
225
354
  gauge optimizations watch cl_1
226
- gauge optimizations trial cl_1 R1 # score, verdict, sessions
355
+ gauge optimizations trial cl_1 R1 # score, verdict, sessions, errors
227
356
  gauge optimizations change cl_1 R1 --body # the winning text, to paste
228
357
  gauge optimizations adopt cl_1 R1 --yes # mark the winner; baseline stays fixed
358
+ gauge optimizations handoff cl_1 # instructions to land the winner in your source
229
359
  gauge optimizations complete cl_1 --yes
230
360
  ```
231
361
 
@@ -260,8 +390,14 @@ the answer is you.
260
390
  `--idempotency-key` makes `trials add` safe to retry: the same key returns the
261
391
  trial the first call created, and never buys a second set of sessions.
262
392
 
263
- Adoption moves Gauge's measured baseline. It does not open a pull request, so
264
- take the winning text from `change --body` and apply it to your own source.
393
+ There is no round limit or credit budget. Each round waits for you; `complete`
394
+ ends the optimization. Adoption marks the winner and leaves the baseline fixed.
395
+ It does not open a pull request. `handoff` prints the same instructions as the
396
+ web app's "Copy instructions for agent"; follow them in your own source.
397
+
398
+ A failed or timed-out session drops out of the score. `trial`, `get`, and
399
+ `watch` print a `WARNING` line for each one, with its failure reason; in JSON,
400
+ read `failedSessions` on the trial and `failureReason` on each run.
265
401
 
266
402
  ## Models, reports, and scripting
267
403