@withgauge/cli 0.8.0 → 0.10.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +72 -23
- package/dist/index.js +2955 -2369
- package/package.json +2 -2
package/README.md
CHANGED
|
@@ -1,8 +1,8 @@
|
|
|
1
1
|
# @withgauge/cli
|
|
2
2
|
|
|
3
3
|
`gauge` — the [Gauge](https://agents.withgauge.com) command-line interface.
|
|
4
|
-
Manage your eval sets, Agent Preference prompts,
|
|
5
|
-
|
|
4
|
+
Manage your eval sets, Agent Preference prompts, and agent runs, and watch runs
|
|
5
|
+
live from your terminal.
|
|
6
6
|
|
|
7
7
|
```sh
|
|
8
8
|
npm install -g @withgauge/cli
|
|
@@ -19,9 +19,8 @@ Requires Node.js ≥ 22.12.
|
|
|
19
19
|
gauge status # the org at a glance
|
|
20
20
|
gauge evals list # judged agent-experience measurements
|
|
21
21
|
gauge preference list # Agent Preference prompts and measurements
|
|
22
|
-
gauge presets list # run configurations (repo × persona × agents × skills × MCP)
|
|
23
22
|
gauge models list # supported model × harness × provider targets
|
|
24
|
-
gauge evals
|
|
23
|
+
gauge evals edit <id> --cadence weekly --samples 3 # when and how much it runs
|
|
25
24
|
gauge evals run <id> # launch now (ORG-funded, credit-gated)
|
|
26
25
|
gauge runs watch <run-id> # follow a run as the agent works
|
|
27
26
|
gauge stats rankings -o json # brand rankings (share of voice) as JSON
|
|
@@ -29,17 +28,72 @@ gauge query -d ecosystem -m installs,install_rate -o json # a flexible report
|
|
|
29
28
|
```
|
|
30
29
|
|
|
31
30
|
Run `gauge --help` (or `gauge <command> --help`) for the full command set:
|
|
32
|
-
evals, preference,
|
|
33
|
-
experiments,
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
31
|
+
evals, preference, models, personas, skills, mcp, repos, runs, optimizations,
|
|
32
|
+
experiments,
|
|
33
|
+
batches, brands, tags, actions, dashboards, members, org, provider keys,
|
|
34
|
+
billing, usage, stats, query, and more.
|
|
35
|
+
|
|
36
|
+
## Measurements own their configuration
|
|
37
|
+
|
|
38
|
+
Every runnable thing carries its own complete run configuration — repository,
|
|
39
|
+
persona, agent roster with model pins, skills, MCP servers, connections — plus
|
|
40
|
+
its own cadence and sample count. Nothing is shared, attached, or detached.
|
|
41
|
+
|
|
42
|
+
An **eval** is one prompt against one configuration:
|
|
43
|
+
|
|
44
|
+
```sh
|
|
45
|
+
gauge evals create --name "add a healthcheck" \
|
|
46
|
+
--text "Add a /health endpoint that returns 200 OK." \
|
|
47
|
+
--criterion "Returns 200: the endpoint responds 200 with no auth" \
|
|
48
|
+
--repo https://github.com/acme/api --ref main \
|
|
49
|
+
--agent claude-code:claude-opus-4-8 --agent codex \
|
|
50
|
+
--skill acme-cli@1.2.0 \
|
|
51
|
+
--cadence weekly --samples 3
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
`gauge evals edit <id>` changes any of it later; omitted flags keep their
|
|
55
|
+
stored value, and each `--clear-*` flag empties one field. To measure the same
|
|
56
|
+
prompt against a second repository or roster, copy the eval in the app and edit
|
|
57
|
+
the copy — one eval is always one configuration.
|
|
58
|
+
|
|
59
|
+
A **preference prompt** is a reusable question. It owns one set of run
|
|
60
|
+
settings — roster, resources, cadence, samples — and any number of
|
|
61
|
+
**scenarios**, each one repository × persona cell. Every scenario runs with
|
|
62
|
+
the prompt's settings, so one prompt measures the same question across several
|
|
63
|
+
repositories and personas at once:
|
|
64
|
+
|
|
65
|
+
```sh
|
|
66
|
+
gauge preference create --name "CLI deployment preference" \
|
|
67
|
+
--kind head-to-head --brand-a acme --brand-b globex \
|
|
68
|
+
--text "Which CLI should I use to deploy?" \
|
|
69
|
+
--agent claude-code --cadence daily \
|
|
70
|
+
--repo https://github.com/acme/api # first scenario
|
|
71
|
+
|
|
72
|
+
gauge preference edit <promptId> --agent codex --samples 3 --cadence weekly
|
|
73
|
+
gauge preference scenarios list <promptId>
|
|
74
|
+
gauge preference scenarios add <promptId> --persona <profileId>
|
|
75
|
+
gauge preference scenarios add <promptId> --repo https://github.com/acme/web --persona <profileId>
|
|
76
|
+
gauge preference scenarios rm <promptId> --scenario <scenarioId>
|
|
77
|
+
```
|
|
78
|
+
|
|
79
|
+
A prompt with zero scenarios is a valid definition — it simply never comes due.
|
|
80
|
+
Sessions per launch = scenarios × agent/model targets × samples. Omit the
|
|
81
|
+
run-settings flags on `create` and the org's default roster applies.
|
|
82
|
+
|
|
83
|
+
## Cadence is the only off switch
|
|
84
|
+
|
|
85
|
+
`--cadence none` is "manual": the measurement stops coming due and runs only
|
|
86
|
+
when you launch it. There is no separate pause or enabled flag, in the CLI or
|
|
87
|
+
the app. `--cadence daily|weekly|monthly` puts it back on a rhythm, and
|
|
88
|
+
`preference get` reads a manual prompt back as `cadence: manual`.
|
|
89
|
+
|
|
90
|
+
`gauge evals run <id>` and `gauge preference run <id>` are the separate
|
|
91
|
+
immediate launches. They use the stored configuration and ask before spending
|
|
92
|
+
(`--yes` skips the prompt for scripts).
|
|
93
|
+
`gauge preference run <id> --scenario <scenarioId>` (repeatable) runs a subset;
|
|
94
|
+
omit it and every scenario runs. Neither launch touches the cadence clock.
|
|
95
|
+
|
|
96
|
+
## Models, reports, and scripting
|
|
43
97
|
|
|
44
98
|
Use `gauge models list` before pinning models. It reads the organization's live
|
|
45
99
|
catalog and shows every selectable logical model with its harness, provider,
|
|
@@ -47,14 +101,9 @@ and effective capabilities. Open harnesses use `--agent pi:<model>` and
|
|
|
47
101
|
`--agent opencode:<model>` just like the frontier harnesses.
|
|
48
102
|
|
|
49
103
|
`gauge query` builds flexible reports over your event warehouse (a GA4
|
|
50
|
-
`run_report` analog): pick a `--dataset`, group by `--dimensions
|
|
51
|
-
|
|
52
|
-
`gauge query
|
|
53
|
-
can ask for.
|
|
54
|
-
|
|
55
|
-
To launch runs declaratively, `gauge apply -f <spec>.json`. Discover the spec
|
|
56
|
-
format without leaving the terminal: `gauge apply --example` prints a starter
|
|
57
|
-
spec and `gauge apply --schema` prints its JSON Schema.
|
|
104
|
+
`run_report` analog): pick a `--dataset`, group by `--dimensions`, and aggregate
|
|
105
|
+
`--metrics`, or pass a whole spec with `--json`. `gauge query datasets` /
|
|
106
|
+
`gauge query fields --dataset <name>` list what you can ask for.
|
|
58
107
|
|
|
59
108
|
Every read supports `-o json` for scripting. `GAUGE_API_TOKEN` overrides
|
|
60
109
|
stored credentials (useful in CI). Exit codes: `0` ok · `1` error ·
|