@withgauge/cli 0.13.0 → 0.15.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -6,6 +6,7 @@ live from your terminal.
6
6
 
7
7
  ```sh
8
8
  npm install -g @withgauge/cli
9
+ gauge update # update the CLI to the latest published version
9
10
  gauge onboard # sign in, create a workspace, seed starter measurements
10
11
  gauge auth login # (or just sign in)
11
12
  gauge orgs use <org> # set your default organization
@@ -13,6 +14,18 @@ gauge orgs use <org> # set your default organization
13
14
 
14
15
  Requires Node.js ≥ 22.12.
15
16
 
17
+ Installing or upgrading the CLI also installs the
18
+ [`gauge-agents` skill](https://github.com/gauge-sh/gauge-skills/tree/main/gauge-agents)
19
+ globally through the `skills` installer, without prompts. It uses the installer's
20
+ agent detection for tools such as Codex and Claude Code. Skill installation
21
+ needs network access; failures leave the CLI installed and print a retry command.
22
+ It is skipped in CI or when `GAUGE_SKIP_SKILL_INSTALL=1` is set. If you install
23
+ with `--ignore-scripts`, install the skill manually:
24
+
25
+ ```sh
26
+ npx --yes skills@1.5.0 add gauge-sh/gauge-skills --skill gauge-agents --global --yes
27
+ ```
28
+
16
29
  ## Usage
17
30
 
18
31
  ```sh
@@ -27,12 +40,21 @@ gauge stats rankings -o json # brand rankings (share of voice) as JSON
27
40
  gauge query -d ecosystem -m installs,install_rate -o json # a flexible report
28
41
  ```
29
42
 
30
- `gauge instructions` prints the CLI's own agent instructions to stdout: one
31
- skill file covering every command group. `gauge instructions <topic>` reprints
32
- one group alone. `gauge --help` points agents at it, or write the file into the
33
- repo with `gauge instructions --install`. The text ships inside the CLI, so it
43
+ `gauge instructions` prints a Gauge overview, CLI setup and usage rules, and
44
+ roughly 100 words of essential guidance per module. Each summary points to
45
+ `gauge instructions <topic>` for the full page. `gauge --help` points agents at
46
+ it, or write the file into the repo with `gauge instructions --install`.
47
+ The text ships inside the CLI, so it
34
48
  never drifts from the commands it describes.
35
49
 
50
+ Use `gauge instructions evals` before creating evals: it covers developer tasks,
51
+ model selection, observable judging criteria, and a complete migration example.
52
+ Use `gauge instructions preference` to investigate product choices, source
53
+ exposure, and interventions, from individual traces to larger cohorts.
54
+ Use `gauge instructions optimization` for the full optimization workflow.
55
+ `gauge instructions --list` shows available topics.
56
+ `gauge instructions evals --install` installs the eval guidance as a local skill.
57
+
36
58
  Run `gauge --help` (or `gauge <command> --help`) for the full command set:
37
59
  evals, preference, models, personas, skills, mcp, repos, runs, optimizations,
38
60
  batches, brands, tags, actions, dashboards, members, org, provider keys,
@@ -41,7 +63,7 @@ billing, usage, stats, query, and more.
41
63
  ## Run committed Markdown evals
42
64
 
43
65
  File evals keep the task, run settings, and judged criteria in Git. Create
44
- `evals/healthcheck.md` with YAML frontmatter and a Markdown task body:
66
+ `gauge-evals/healthcheck.md` with YAML frontmatter and a Markdown task body:
45
67
 
46
68
  ```markdown
47
69
  ---
@@ -54,66 +76,194 @@ config:
54
76
  agents:
55
77
  - agent: CODEX_CLI
56
78
  ---
79
+
57
80
  Add a /health endpoint and run the relevant app checks.
58
81
  ```
59
82
 
60
- Commit the file, then preview and run its exact committed bytes:
83
+ Check your working-tree setup, then commit the files before planning and running:
61
84
 
62
85
  ```sh
63
- gauge evals plan --files evals/healthcheck.md
64
- gauge evals run --files evals/healthcheck.md --yes
86
+ gauge evals verify # offline; includes uncommitted working-tree edits
87
+ gauge evals verify -o json # versioned global/per-input yes/no/unknown checklist
88
+ gauge evals plan
89
+ gauge evals run --yes
65
90
  gauge run-requests wait <request-id> -o json
66
91
  ```
67
92
 
68
- `--files` takes explicit paths; pass several paths after the flag for a
69
- multi-case request. The CLI reads `HEAD` through Git, so uncommitted edits are
70
- ignored and untracked cases fail. It hashes the committed case bytes and root
71
- config and submits their commit, digest, and repository origin. The server
72
- currently records that provenance as **submitted, unverified**; it does not
73
- fetch Git independently. `plan` resolves the configuration and estimates
74
- sessions without launching. Each case owns a `config` object with its own
75
- agent roster; supported fields are `agents`, `sampleCount`, `repoUrl`,
76
- `repoRef`, `profileId`, `skillRefs`, `mcpRefs`, `connectionIds`, and optional
77
- `addons.browser` (`session`, positive `projectId`, optional `region: us`).
78
- Skills and MCP assets use pinned `name@label` references. At least one agent
79
- and one criterion are required for each case. Omitted models use the agent's
80
- default model; omitted `sampleCount` runs one sample. With no `repoUrl`, the
81
- existing default scaffold is used.
82
-
83
- An optional root `gauge.json` selects the organization and one product build:
93
+ Cases default to `gauge-evals/**/*.md`. `--files` accepts quoted paths/globs and overrides
94
+ that convention for `plan` and `run`, which read committed Git HEAD, ignoring working-tree edits.
95
+ Each case owns its agents, models, consumer repository/ref, Persona and criteria.
96
+ The optional `config.inputs` selects named inputs; all attach by default.
97
+
98
+ `verify` reads the repository's current files, including non-ignored untracked files,
99
+ without login, network access, a PR, built CI artifacts, or sandbox compute. It checks
100
+ configuration, case input references, Git input paths, and recognizable workflow
101
+ checkout/upload declarations. An input-only repository with no cases is valid setup.
102
+ Each attempted global/per-input check reports `yes` (established locally), `no`
103
+ (a definite problem), or `unknown` (a specific static question could not be resolved,
104
+ such as an expression-based artifact name or a referenced reusable workflow).
105
+ Missing optional configuration and zero cases are informational findings, excluded
106
+ from checklist counts. Eligible workflow triggers are checked as declarations;
107
+ event filters and runtime conditions are not evaluated against a candidate PR.
108
+ Live input delivery, artifact contents, sandbox execution, App activation, and
109
+ account settings are outside this command's scope and do not create unknown rows.
110
+ Exit 1 means at least one `no`; exit 0 means no definite static problem was found.
111
+ No workflow code is executed. Commit files before `plan` or `run` consumes them.
112
+
113
+ `gauge evals verify -o json` writes exactly one JSON object to stdout, including on
114
+ validation failure. The versioned report provides:
115
+
116
+ - `version: 1` and `mode: "local-static"` identify the report contract and scope.
117
+ - `status` is `failed` if any check is `no`, otherwise `incomplete` if any check is
118
+ `unknown`, otherwise `passed`. It describes static checks only.
119
+ - `summary` contains `yes`, `no`, and `unknown` counts.
120
+ - `global` and `inputs[].checks` contain checks with stable `id`, `state`, and
121
+ explanatory `detail` fields; `inputs[].name` identifies each input.
122
+ - `facts.configFilePresent` and `facts.caseCount` expose discovered facts without
123
+ parsing prose. A `null` value means inspection did not reach that fact; zero cases
124
+ is distinct from case discovery being skipped or failing.
125
+ - `info` contains informational findings with stable `id` and `detail` fields.
126
+ - `notChecked` lists excluded capabilities as machine-readable identifiers.
127
+
128
+ For example, an agent can select unresolved checks with:
129
+
130
+ ```sh
131
+ gauge evals verify -o json | jq '{status, facts, global: [.global[] | select(.state != "yes")], inputs: [.inputs[] | {name, checks: [.checks[] | select(.state != "yes")]}]}'
132
+ ```
133
+
134
+ An optional `gauge.json` declares inputs produced by your normal CI or committed
135
+ alongside your product:
84
136
 
85
137
  ```json
86
- {"org":"my-org","product":{"type":"npm-cli"}}
138
+ {
139
+ "version": 2,
140
+ "org": "my-org",
141
+ "inputs": {
142
+ "cli": {
143
+ "type": "npm-package",
144
+ "source": { "type": "github-actions", "workflow": ".github/workflows/package.yml" },
145
+ "path": "*.tgz",
146
+ "install": "environment"
147
+ },
148
+ "docs": {
149
+ "type": "files",
150
+ "source": { "type": "git" },
151
+ "paths": ["docs/**/*.md"]
152
+ }
153
+ }
154
+ }
87
155
  ```
88
156
 
89
- For an npm CLI, Gauge uses the committed package manifest and lockfile, runs
90
- its `build` script when present, packs and installs the package with production
91
- dependencies, and exposes its declared CLI commands with the pinned Node
92
- runtime. Set `product.build` to a different build command when needed.
93
- The source install is frozen; Gauge resolves a fresh production install once per
94
- run request, archives that installed tree for every case, and records its
95
- runtime lockfile digest in the build receipt. New requests can resolve newer
96
- transitive dependencies until the product pins them.
97
- Each new CLI launch currently rebuilds, even at the same source SHA; bundles
98
- are shared across cases and reused after completed preparation within a request,
99
- but there is no cache across requests yet.
100
- The first adapter requires the pinned pnpm 10.34.5 package manager and a
101
- `pnpm-lock.yaml`.
102
- `--org` overrides the committed `org`; otherwise the CLI falls back to your
103
- configured default organization.
104
-
105
- For other products, `{"prepare":"pnpm run gauge:prepare","install":"./install.sh"}`
106
- remains available. `prepare` runs once at the exact definition commit with
107
- `GAUGE_OUTPUT_DIR` pointing to an empty directory; `install` is an optional
108
- bundle-relative script run once per consumer session. A committed
109
- `.gauge/prepare.sh` enables that preparation automatically when `prepare` is
110
- absent. `product` cannot be combined with either script form. The output is
111
- attached to every case. There is no authored resource manifest or per-case
112
- `inputs` field. Project-level run defaults are rejected. The optional
113
- `version` defaults to 1. Preparation requires a public HTTPS Git origin.
114
- Saved evals remain available with `gauge evals run <id>`.
115
- As with saved evals, superusers use platform funding by default and can add
116
- `--bill-to-org` to charge the selected organization.
157
+ Gauge waits for the existing workflow at the pushed revision and consumes its
158
+ artifacts by ID. Omit `source.artifact` to select a sole artifact, or use a name
159
+ pattern when the workflow uploads several. It never triggers or performs a build. Connect the Gauge
160
+ GitHub App with repository access and Actions read permission. Plan previews CI
161
+ readiness; waiting requests create no sessions and incur no session debit.
162
+ Each input names its workflow. Gauge captures one successful run/attempt per
163
+ workflow before downloading any input. A committed org and matching GitHub App
164
+ installation enable automatic same-repository PR evaluations for repositories and
165
+ target branches enabled in Settings → GitHub (initially enabled for `main`).
166
+ A failed run or absent/expired artifact returns an actionable input failure.
167
+ Absent runs have a five-minute discovery grace period; known running or queued
168
+ CI can continue past an hour. Abandoned requests expire after seven days. Reruns retain captured bytes. CI metadata reports a revision; it
169
+ does not independently prove which checkout the workflow compiled.
170
+
171
+ To run automatic PR checks only when relevant files change, add `checks` to
172
+ `gauge.json`:
173
+
174
+ ```json
175
+ "checks": {
176
+ "paths": ["agents/**", "docs.json", "snippets/**"]
177
+ }
178
+ ```
179
+
180
+ Paths are repository-relative, case-sensitive globs; matching any pattern runs
181
+ the checks. `**` matches nested directories, including hidden files. Use 1–32
182
+ patterns; exclusions (`!`) are not supported. Changes to `gauge.json` or your
183
+ configured case paths always count. Deleted files and both names of renamed
184
+ files are included. Without `checks`, every eligible PR update runs as before.
185
+
186
+ The first check uses the PR diff. After a passing evaluation, later updates
187
+ compare against that evaluated commit, so unrelated follow-up edits can skip.
188
+ A failed or unfinished evaluation, a changed target branch, or an unavailable
189
+ or incomplete diff causes a run. Skipped checks start no sessions and reserve
190
+ no credits. GitHub's rerun action runs regardless of paths; explicit CLI runs
191
+ also ignore this automatic-check filter.
192
+
193
+ File inputs stay outside the consumer repository and their paths are provided to
194
+ the agent. npm packages install with ordinary npm using the VM's Node/npm runtime.
195
+ `install` defaults to `environment`: a separate tools directory with the package's
196
+ own declared commands on PATH, including login shells. `workspace` adds a consumer
197
+ application dependency and accepts libraries without commands. Transitive registry
198
+ dependency ranges remain live. Plan output shows the installation destination and,
199
+ when a previous matching preparation captured it, command metadata. Otherwise it
200
+ reports that commands will be determined during preparation; planning never downloads
201
+ archives or starts VMs.
202
+
203
+ Native executables use `type: "binary"`, with already unpacked files in a GitHub
204
+ Actions artifact (or small committed Git files):
205
+
206
+ ```json
207
+ "gh": {
208
+ "type": "binary",
209
+ "source": {
210
+ "type": "github-actions",
211
+ "workflow": ".github/workflows/package.yml",
212
+ "artifact": "gh-linux"
213
+ },
214
+ "paths": ["dist/**"],
215
+ "platform": "linux-x64",
216
+ "commands": { "gh": "dist/bin/gh" }
217
+ }
218
+ ```
219
+
220
+ Paths are relative to the uploaded artifact root. Build the exact candidate SHA;
221
+ for `actions/upload-artifact@v7`, use `archive: true`. Include companion files in
222
+ `paths` and explicitly map 1–32 command names to ordinary selected files. Gauge
223
+ validates Linux x86-64 ELF executables in its isolated artifact processor, restores
224
+ executable permissions, preserves the bundle outside the consumer repository, and
225
+ exposes the commands on PATH, including login shells. A candidate command overrides
226
+ an installed product command with the same name; two inputs selected by the same
227
+ case cannot expose the same command (including captured npm environment commands).
228
+
229
+ The first version supports ELF executables and PIE binaries compatible with the
230
+ consumer Linux runtime. It does not unpack nested release archives, install OS
231
+ packages, alter library search paths, or run arbitrary installation/version-check
232
+ commands. Prefer a static binary or build against a compatible Linux runtime;
233
+ header validation cannot prove shared-library compatibility. Scripts, macOS,
234
+ Windows, and ARM executables are unsupported. A missing loader/library encountered
235
+ when the agent invokes the command appears in session execution evidence, not in
236
+ static validation. Harness/runtime names (including shells, Node/npm, Python, Git,
237
+ archive/setup utilities, and agent commands) are reserved; use a different alias
238
+ for such a candidate. `GAUGE_INPUTS` gives the preserved bundle path.
239
+
240
+ A Skill input has `type: "skill"` and `path: "skills/product"`, or a path to a ZIP,
241
+ `.tar.gz`, or `.tgz` whose root contains `SKILL.md`. Use `path: "."` for a Skill at
242
+ the source root. Gauge validates and extracts selected archives in isolation and
243
+ stages the Skill in the coding agent's native discovery location. File inputs use
244
+ `paths` to select committed files or files inside an Actions artifact.
245
+
246
+ Website inputs use `type: "website"` and a canonical HTTPS origin in `url`.
247
+ The preview can come from `source: { "type": "github-deployment", "environment": "preview" }`
248
+ or a `github-pr-comment` source with an exact `author`, literal `bodyIncludes`
249
+ marker, and a readiness `pattern` with named `url` and `sha` captures. Gauge selects
250
+ candidate-specific evidence; timestamps alone cannot establish the revision.
251
+ Comments must belong to a unique open same-repository PR at the candidate head.
252
+ Previews must be public HTTPS origins, without authentication or path prefixes.
253
+ They are fetched live rather than snapshotted. See the [preview source examples and
254
+ limitations](../specs/docs-web-inputs.md), including Astro's existing bot comment.
255
+
256
+ Editor schemas are published at
257
+ `https://agents.withgauge.com/schemas/gauge/v2.4.0/gauge.json` and
258
+ `https://agents.withgauge.com/schemas/gauge/v2.4.0/case.json`. Associate them in editor
259
+ settings; `gauge.json` does not accept a `$schema` field. These release URLs are
260
+ immutable. Update earlier pilot configs by moving workflows into input sources,
261
+ replacing `use` with input `type` and `path`/`paths`, and moving cases into
262
+ `gauge-evals/` (or setting `cases` explicitly).
263
+
264
+ `--org` overrides the committed organization. Root `cases` can change discovery;
265
+ root `version` defaults to 2. Saved evals use `gauge evals run <id>`. Superusers
266
+ use platform funding by default and can add `--bill-to-org`.
117
267
 
118
268
  ## Measurements own their configuration
119
269
 
@@ -192,7 +342,6 @@ A scripted or agent-driven loop:
192
342
 
193
343
  ```sh
194
344
  gauge optimizations diagnose --eval es_1 # where the sessions fail
195
- gauge optimizations estimate --eval es_1 --target claude-code:claude-opus-4-8
196
345
  gauge optimizations start --eval es_1 --target claude-code:claude-opus-4-8 --yes
197
346
  gauge optimizations watch cl_1 # exit 4 = your turn
198
347
 
@@ -203,9 +352,10 @@ gauge optimizations trials add cl_1 \
203
352
  --body-file start.md --launch --idempotency-key start-env-var --yes
204
353
 
205
354
  gauge optimizations watch cl_1
206
- gauge optimizations trial cl_1 R1 # score, verdict, sessions
355
+ gauge optimizations trial cl_1 R1 # score, verdict, sessions, errors
207
356
  gauge optimizations change cl_1 R1 --body # the winning text, to paste
208
357
  gauge optimizations adopt cl_1 R1 --yes # mark the winner; baseline stays fixed
358
+ gauge optimizations handoff cl_1 # instructions to land the winner in your source
209
359
  gauge optimizations complete cl_1 --yes
210
360
  ```
211
361
 
@@ -240,8 +390,14 @@ the answer is you.
240
390
  `--idempotency-key` makes `trials add` safe to retry: the same key returns the
241
391
  trial the first call created, and never buys a second set of sessions.
242
392
 
243
- Adoption moves Gauge's measured baseline. It does not open a pull request, so
244
- take the winning text from `change --body` and apply it to your own source.
393
+ There is no round limit or credit budget. Each round waits for you; `complete`
394
+ ends the optimization. Adoption marks the winner and leaves the baseline fixed.
395
+ It does not open a pull request. `handoff` prints the same instructions as the
396
+ web app's "Copy instructions for agent"; follow them in your own source.
397
+
398
+ A failed or timed-out session drops out of the score. `trial`, `get`, and
399
+ `watch` print a `WARNING` line for each one, with its failure reason; in JSON,
400
+ read `failedSessions` on the trial and `failureReason` on each run.
245
401
 
246
402
  ## Models, reports, and scripting
247
403