@withgauge/cli 0.11.0 → 0.12.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +82 -1
- package/dist/index.js +1906 -708
- package/package.json +5 -2
package/README.md
CHANGED
|
@@ -39,6 +39,83 @@ experiments,
|
|
|
39
39
|
batches, brands, tags, actions, dashboards, members, org, provider keys,
|
|
40
40
|
billing, usage, stats, query, and more.
|
|
41
41
|
|
|
42
|
+
## Run committed Markdown evals
|
|
43
|
+
|
|
44
|
+
File evals keep the task, run settings, and judged criteria in Git. Create
|
|
45
|
+
`evals/healthcheck.md` with YAML frontmatter and a Markdown task body:
|
|
46
|
+
|
|
47
|
+
```markdown
|
|
48
|
+
---
|
|
49
|
+
criteria:
|
|
50
|
+
- name: Healthcheck responds
|
|
51
|
+
rubric: The endpoint returns HTTP 200 without authentication.
|
|
52
|
+
config:
|
|
53
|
+
repoUrl: https://github.com/acme/app
|
|
54
|
+
repoRef: main
|
|
55
|
+
agents:
|
|
56
|
+
- agent: CODEX_CLI
|
|
57
|
+
---
|
|
58
|
+
Add a /health endpoint and run the relevant app checks.
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
Commit the file, then preview and run its exact committed bytes:
|
|
62
|
+
|
|
63
|
+
```sh
|
|
64
|
+
gauge evals plan --files evals/healthcheck.md
|
|
65
|
+
gauge evals run --files evals/healthcheck.md --yes
|
|
66
|
+
gauge run-requests wait <request-id> -o json
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
`--files` takes explicit paths; pass several paths after the flag for a
|
|
70
|
+
multi-case request. The CLI reads `HEAD` through Git, so uncommitted edits are
|
|
71
|
+
ignored and untracked cases fail. It hashes the committed case bytes and root
|
|
72
|
+
config and submits their commit, digest, and repository origin. The server
|
|
73
|
+
currently records that provenance as **submitted, unverified**; it does not
|
|
74
|
+
fetch Git independently. `plan` resolves the configuration and estimates
|
|
75
|
+
sessions without launching. Each case owns a `config` object with its own
|
|
76
|
+
agent roster; supported fields are `agents`, `sampleCount`, `repoUrl`,
|
|
77
|
+
`repoRef`, `profileId`, `skillRefs`, `mcpRefs`, `connectionIds`, and optional
|
|
78
|
+
`addons.browser` (`session`, positive `projectId`, optional `region: us`).
|
|
79
|
+
Skills and MCP assets use pinned `name@label` references. At least one agent
|
|
80
|
+
and one criterion are required for each case. Omitted models use the agent's
|
|
81
|
+
default model; omitted `sampleCount` runs one sample. With no `repoUrl`, the
|
|
82
|
+
existing default scaffold is used.
|
|
83
|
+
|
|
84
|
+
An optional root `gauge.json` selects the organization and one product build:
|
|
85
|
+
|
|
86
|
+
```json
|
|
87
|
+
{"org":"my-org","product":{"type":"npm-cli"}}
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
For an npm CLI, Gauge uses the committed package manifest and lockfile, runs
|
|
91
|
+
its `build` script when present, packs and installs the package with production
|
|
92
|
+
dependencies, and exposes its declared CLI commands with the pinned Node
|
|
93
|
+
runtime. Set `product.build` to a different build command when needed.
|
|
94
|
+
The source install is frozen; Gauge resolves a fresh production install once per
|
|
95
|
+
run request, archives that installed tree for every case, and records its
|
|
96
|
+
runtime lockfile digest in the build receipt. New requests can resolve newer
|
|
97
|
+
transitive dependencies until the product pins them.
|
|
98
|
+
Each new CLI launch currently rebuilds, even at the same source SHA; bundles
|
|
99
|
+
are shared across cases and reused after completed preparation within a request,
|
|
100
|
+
but there is no cache across requests yet.
|
|
101
|
+
The first adapter requires the pinned pnpm 10.34.5 package manager and a
|
|
102
|
+
`pnpm-lock.yaml`.
|
|
103
|
+
`--org` overrides the committed `org`; otherwise the CLI falls back to your
|
|
104
|
+
configured default organization.
|
|
105
|
+
|
|
106
|
+
For other products, `{"prepare":"pnpm run gauge:prepare","install":"./install.sh"}`
|
|
107
|
+
remains available. `prepare` runs once at the exact definition commit with
|
|
108
|
+
`GAUGE_OUTPUT_DIR` pointing to an empty directory; `install` is an optional
|
|
109
|
+
bundle-relative script run once per consumer session. A committed
|
|
110
|
+
`.gauge/prepare.sh` enables that preparation automatically when `prepare` is
|
|
111
|
+
absent. `product` cannot be combined with either script form. The output is
|
|
112
|
+
attached to every case. There is no authored resource manifest or per-case
|
|
113
|
+
`inputs` field. Project-level run defaults are rejected. The optional
|
|
114
|
+
`version` defaults to 1. Preparation requires a public HTTPS Git origin.
|
|
115
|
+
Saved evals remain available with `gauge evals run <id>`.
|
|
116
|
+
As with saved evals, superusers use platform funding by default and can add
|
|
117
|
+
`--bill-to-org` to charge the selected organization.
|
|
118
|
+
|
|
42
119
|
## Measurements own their configuration
|
|
43
120
|
|
|
44
121
|
Every runnable thing carries its own complete run configuration — repository,
|
|
@@ -98,6 +175,10 @@ immediate launches. They use the stored configuration and ask before spending
|
|
|
98
175
|
(`--yes` skips the prompt for scripts).
|
|
99
176
|
`gauge preference run <id> --scenario <scenarioId>` (repeatable) runs a subset;
|
|
100
177
|
omit it and every scenario runs. Neither launch touches the cadence clock.
|
|
178
|
+
Superusers can use `--org <slug>` to launch in any org. Their runs use PLATFORM
|
|
179
|
+
funding by default, like Run Now in the web UI; `--bill-to-org` uses that org's
|
|
180
|
+
credits instead. Other users always use org credits and can launch only in orgs
|
|
181
|
+
they belong to.
|
|
101
182
|
|
|
102
183
|
## Optimizations: Gauge measures, you author
|
|
103
184
|
|
|
@@ -125,7 +206,7 @@ gauge optimizations trials add cl_1 \
|
|
|
125
206
|
gauge optimizations watch cl_1
|
|
126
207
|
gauge optimizations trial cl_1 R1 # score, verdict, sessions
|
|
127
208
|
gauge optimizations change cl_1 R1 --body # the winning text, to paste
|
|
128
|
-
gauge optimizations adopt cl_1 R1 --yes # the
|
|
209
|
+
gauge optimizations adopt cl_1 R1 --yes # mark the winner; baseline stays fixed
|
|
129
210
|
gauge optimizations complete cl_1 --yes
|
|
130
211
|
```
|
|
131
212
|
|