@withgauge/cli 0.8.0 → 0.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (3) hide show
  1. package/README.md +71 -23
  2. package/dist/index.js +1702 -2301
  3. package/package.json +2 -2
package/README.md CHANGED
@@ -1,8 +1,8 @@
1
1
  # @withgauge/cli
2
2
 
3
3
  `gauge` — the [Gauge](https://agents.withgauge.com) command-line interface.
4
- Manage your eval sets, Agent Preference prompts, run presets, and agent runs,
5
- and watch runs live from your terminal.
4
+ Manage your eval sets, Agent Preference prompts, and agent runs, and watch runs
5
+ live from your terminal.
6
6
 
7
7
  ```sh
8
8
  npm install -g @withgauge/cli
@@ -19,9 +19,8 @@ Requires Node.js ≥ 22.12.
19
19
  gauge status # the org at a glance
20
20
  gauge evals list # judged agent-experience measurements
21
21
  gauge preference list # Agent Preference prompts and measurements
22
- gauge presets list # run configurations (repo × persona × agents × skills × MCP)
23
22
  gauge models list # supported model × harness × provider targets
24
- gauge evals attach <id> --preset <presetId> # schedule a preset on an eval (prints the bill first)
23
+ gauge evals edit <id> --cadence weekly --samples 3 # when and how much it runs
25
24
  gauge evals run <id> # launch now (ORG-funded, credit-gated)
26
25
  gauge runs watch <run-id> # follow a run as the agent works
27
26
  gauge stats rankings -o json # brand rankings (share of voice) as JSON
@@ -29,17 +28,71 @@ gauge query -d ecosystem -m installs,install_rate -o json # a flexible report
29
28
  ```
30
29
 
31
30
  Run `gauge --help` (or `gauge <command> --help`) for the full command set:
32
- evals, preference, presets, models, personas, skills, mcp, cycles, repos, runs,
33
- experiments, batches, brands, tags, actions, impl, dashboards, members, org,
34
- provider keys, billing, usage, stats, query, and more.
35
-
36
- **Run presets** are the unit of configuration: an eval set or preference
37
- prompt runs whatever presets are attached to it, every cycle. `gauge presets
38
- create --name api-main --repo https://github.com/acme/api --agent
39
- claude-code:claude-opus-4-8 --skill acme-cli@1.2.0` authors one;
40
- `gauge evals attach` / `gauge preference attach` schedules it. Attaching
41
- spends money on a schedule, so `attach` and `run` print the per-cycle run
42
- arithmetic and ask before acting (`--yes` skips the prompt for scripts).
31
+ evals, preference, models, personas, skills, mcp, repos, runs, experiments,
32
+ batches, brands, tags, actions, dashboards, members, org, provider keys,
33
+ billing, usage, stats, query, and more.
34
+
35
+ ## Measurements own their configuration
36
+
37
+ Every runnable thing carries its own complete run configuration — repository,
38
+ persona, agent roster with model pins, skills, MCP servers, connections — plus
39
+ its own cadence and sample count. Nothing is shared, attached, or detached.
40
+
41
+ An **eval** is one prompt against one configuration:
42
+
43
+ ```sh
44
+ gauge evals create --name "add a healthcheck" \
45
+ --text "Add a /health endpoint that returns 200 OK." \
46
+ --criterion "Returns 200: the endpoint responds 200 with no auth" \
47
+ --repo https://github.com/acme/api --ref main \
48
+ --agent claude-code:claude-opus-4-8 --agent codex \
49
+ --skill acme-cli@1.2.0 \
50
+ --cadence weekly --samples 3
51
+ ```
52
+
53
+ `gauge evals edit <id>` changes any of it later; omitted flags keep their
54
+ stored value, and each `--clear-*` flag empties one field. To measure the same
55
+ prompt against a second repository or roster, copy the eval in the app and edit
56
+ the copy — one eval is always one configuration.
57
+
58
+ A **preference prompt** is a reusable question. It owns one set of run
59
+ settings — roster, resources, cadence, samples — and any number of
60
+ **scenarios**, each one repository × persona cell. Every scenario runs with
61
+ the prompt's settings, so one prompt measures the same question across several
62
+ repositories and personas at once:
63
+
64
+ ```sh
65
+ gauge preference create --name "CLI deployment preference" \
66
+ --kind head-to-head --brand-a acme --brand-b globex \
67
+ --text "Which CLI should I use to deploy?" \
68
+ --agent claude-code --cadence daily \
69
+ --repo https://github.com/acme/api # first scenario
70
+
71
+ gauge preference edit <promptId> --agent codex --samples 3 --cadence weekly
72
+ gauge preference scenarios list <promptId>
73
+ gauge preference scenarios add <promptId> --persona <profileId>
74
+ gauge preference scenarios add <promptId> --repo https://github.com/acme/web --persona <profileId>
75
+ gauge preference scenarios rm <promptId> --scenario <scenarioId>
76
+ ```
77
+
78
+ A prompt with zero scenarios is a valid definition — it simply never comes due.
79
+ Sessions per launch = scenarios × agent/model targets × samples. Omit the
80
+ run-settings flags on `create` and the org's default roster applies.
81
+
82
+ ## Cadence is the only off switch
83
+
84
+ `--cadence none` is "manual": the measurement stops coming due and runs only
85
+ when you launch it. There is no separate pause or enabled flag, in the CLI or
86
+ the app. `--cadence daily|weekly|monthly` puts it back on a rhythm, and
87
+ `preference get` reads a manual prompt back as `cadence: manual`.
88
+
89
+ `gauge evals run <id>` and `gauge preference run <id>` are the separate
90
+ immediate launches. They use the stored configuration and ask before spending
91
+ (`--yes` skips the prompt for scripts).
92
+ `gauge preference run <id> --scenario <scenarioId>` (repeatable) runs a subset;
93
+ omit it and every scenario runs. Neither launch touches the cadence clock.
94
+
95
+ ## Models, reports, and scripting
43
96
 
44
97
  Use `gauge models list` before pinning models. It reads the organization's live
45
98
  catalog and shows every selectable logical model with its harness, provider,
@@ -47,14 +100,9 @@ and effective capabilities. Open harnesses use `--agent pi:<model>` and
47
100
  `--agent opencode:<model>` just like the frontier harnesses.
48
101
 
49
102
  `gauge query` builds flexible reports over your event warehouse (a GA4
50
- `run_report` analog): pick a `--dataset`, group by `--dimensions` (including
51
- `run_preset`), and aggregate `--metrics`, or pass a whole spec with `--json`.
52
- `gauge query datasets` / `gauge query fields --dataset <name>` list what you
53
- can ask for.
54
-
55
- To launch runs declaratively, `gauge apply -f <spec>.json`. Discover the spec
56
- format without leaving the terminal: `gauge apply --example` prints a starter
57
- spec and `gauge apply --schema` prints its JSON Schema.
103
+ `run_report` analog): pick a `--dataset`, group by `--dimensions`, and aggregate
104
+ `--metrics`, or pass a whole spec with `--json`. `gauge query datasets` /
105
+ `gauge query fields --dataset <name>` list what you can ask for.
58
106
 
59
107
  Every read supports `-o json` for scripting. `GAUGE_API_TOKEN` overrides
60
108
  stored credentials (useful in CI). Exit codes: `0` ok · `1` error ·