@remits/remits-cli 0.1.136 → 0.1.138

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@remits/remits-cli",
3
- "version": "0.1.136",
3
+ "version": "0.1.138",
4
4
  "description": "Local CLI for auth, component sync, and live test execution against Remits",
5
5
  "license": "MIT",
6
6
  "private": false,
@@ -48,6 +48,7 @@ reference below, load it before you act, not after something surprises you.
48
48
  | register this session as an agent, or ask why a routed ticket never started | `agent-sessions.md` | `serve` vs `register`, what registration does, worker spawning per agent kind, routing order, the control center, multi-session auth |
49
49
  | investigate live behavior | `investigation.md` | which tool reads which record, correlation keys, reading a record's `content`, HTTP audits, AI activity, node/`localMode`, the production support flows |
50
50
  | call any `mcp_*` tool | `tool-reference.md` | every tool's parameters, semantics, and traps — read the tool's entry before building its input |
51
+ | run, describe, poll, or interrupt an Action from the CLI | `tool-reference.md` → `mcp_run_action` | Action execution through `remits-cli tool --name mcp_run_action`; there is no separate top-level `remits-cli action` / `run action` wrapper |
51
52
  | reason about repo discovery, auth state, or what a command actually sent | `cli-state.md` | the global `~/.remits-cli/` control plane vs per-repo `./.remits-cli/`, and which file answers which question |
52
53
  | need an exact flag, or the auth / host / async surface | `command-reference.md` | authentication, host vs data mode, the two async mechanisms, hierarchy-scoped reads, the full command list, prod banners |
53
54
  | give up on something that misbehaved | `troubleshooting.md` | the symptom→fix table, the two kinds of escalation, and the escalation bundle |
@@ -139,6 +140,12 @@ reference named after it.
139
140
  (`component-resolution.md`, `command-reference.md`)
140
141
  - **Host and data mode are two independent decisions.** `--base-url` picks the Remits host,
141
142
  `--data-mode` picks the data segment on it. Neither implies the other. (`command-reference.md`)
143
+ - **Localhost being down is not a verification blocker.** `remits-cli` can talk to any reachable Remits
144
+ host, including the deployed production host, by passing `--base-url`. If `http://localhost:8080`
145
+ refuses a connection, retry the same stage/test/tool/status command against the intended deployed host
146
+ with the same account, branch, workspace and data-mode facts; do not stop unless no reachable host or
147
+ valid session exists. A deployed host can still run test-lane verification: `--base-url` selects the
148
+ server, while `test run` still defaults to `--data-mode test`. (`command-reference.md`)
142
149
  - **`test run` ignores the stored session data lane.** It defaults to `test` even when `whoami` shows
143
150
  the session parked on prod; production Test runs require explicit prod provenance (`--data-mode prod`)
144
151
  or Test source declaring `dataMode 'prod'` / `[dataMode:'prod']`. Check returned `dataModeSource`
@@ -215,7 +222,8 @@ remits-cli components status # trunk or variant checkout, staging lane, wor
215
222
  remits-cli tools # which tools this account actually has (tools are per-account)
216
223
  ```
217
224
 
218
- plus the repo's `account-info.json` → `resolution` block for the account's shape.
225
+ plus the repo's `account-boot.json` → `resolution` block for the account's shape (`account-info.json`
226
+ only when you need full inventory detail).
219
227
  For a workflow-shaped request, rely on the automatic `remits-cli evidence` trail unless someone else needs
220
228
  a checkable verdict. Only then start `remits-cli verify start --summary "..."`, declare claims, and keep
221
229
  that envelope active through stage/test/token/sync.
@@ -227,7 +235,8 @@ non-default host).
227
235
  `~/.remits-cli/account-repos.json` (every local repo, plus the reserved `platform` entry for the core
228
236
  platform clone), `~/.remits-cli/sessions.json` (auth state and lanes), `~/.remits-cli/agents.json`
229
237
  (agents registered here), `./.remits-cli/actors/<local-agent>/tool-responses/<callId>.json` (the full
230
- tool payload), and the repo's `account-info.json`. `cli-state.md` maps every remaining question to its
238
+ tool payload), the repo's `account-boot.json`, and `account-info.json` when compact context is not enough.
239
+ `cli-state.md` maps every remaining question to its
231
240
  file, including legacy flat `.remits-cli/` fallbacks.
232
241
  Repo-local session JSONL is intentionally a bounded audit log: request payloads and ordinary response
233
242
  bodies are summarized with keys, sizes, hashes, and redaction markers. Open actor-scoped tool response
@@ -241,7 +250,7 @@ Keep support and development sessions lean:
241
250
  `mcp_component_view` / `mcp_component_grep` are fallback surfaces for agents without that checkout,
242
251
  or for confirming what the live DB has stored after you already understand the files. For a ticket
243
252
  with `implementationAccountId`, resolve that account's indexed repo first, pull the appropriate
244
- branch when the checkout is clean, and inspect `account-info.json` + `components/` there.
253
+ branch when the checkout is clean, and inspect `account-boot.json` + relevant `components/` there.
245
254
  - Do not read entire `.remits-cli/actors/<local-agent>/sessions/*.jsonl` (or legacy flat session logs) or
246
255
  large tool response files unless you first narrow to the relevant request, endpoint, tool, or ticket.
247
256
  - Prefer targeted Firestore queries: use `documentId`, tight `filters`, narrow `fields`, and low `limit`
@@ -23,13 +23,14 @@
23
23
  against, and where a fix belongs are answered by the account's structure — never by its name.
24
24
 
25
25
  The account model itself — types, component inheritance, primary vs membership edges, the three independent
26
- edge properties, the user model — is in the always-loaded `platform-overview.md` and in depth in
26
+ edge properties, the user model — is summarized in `platform-core.md` and in depth in
27
27
  `features/account-management.md` (`mcp_get_guide`). What follows is only what changes **what you type**.
28
28
 
29
29
  ### Read the shape first
30
30
 
31
- `account-info.json` (in a repo) and `mcp_account_view` (remotely) both carry a `resolution` block — the one
32
- place these facts appear. Field-by-field detail is under **`mcp_account_view`** in `tool-reference.md`. The
31
+ `account-boot.json` (in a repo), `account-info.json` (full local inventory), and `mcp_account_view`
32
+ (remotely) all carry a `resolution` block — the one place these facts appear. Field-by-field detail is
33
+ under **`mcp_account_view`** in `tool-reference.md`. The
33
34
  four that decide a CLI action:
34
35
 
35
36
  | Read | To decide |
@@ -57,7 +58,8 @@ Two more, easily confused: top-level **`componentBranches`** lists the variant b
57
58
  current before reading source or editing. Run `git fetch origin`; when the tree is clean,
58
59
  `git pull --ff-only origin <branch>`; then confirm `git log origin/<branch>..<branch>` and
59
60
  `git log <branch>..origin/<branch>` are both empty. A worktree can have an isolated staging workspace
60
- and still be based on a stale commit. Then read `account-info.json` and inspect `components/` directly.
61
+ and still be based on a stale commit. Then read `account-boot.json` and inspect relevant `components/`
62
+ directly; use `account-info.json` only for full inventory detail.
61
63
  - **Inside one repo but supporting a different account**: switch to that account's indexed repo if it
62
64
  exists. For tickets, prefer `implementationAccountId` / `implementationAccountName` over the reporting
63
65
  `accountId` when choosing that repo; a subscriber or client often reports the symptom while the
@@ -84,7 +86,7 @@ still go to the test lane. `Object.testMode` /
84
86
  `Event.testMode` / `Alert.testMode` (and `testMode` inside test-lane Audit documents) identify lifecycle
85
87
  rows in the test data lane. Agent-facing surfaces expose these fields:
86
88
 
87
- - `account-info.json`, `account-hierarchy.json`, and `mcp_account_view`: `resolution.testAccount` for the
89
+ - `account-boot.json`, `account-info.json`, `account-hierarchy.json`, and `mcp_account_view`: `resolution.testAccount` for the
88
90
  described account, and `testAccount` on returned hierarchy nodes.
89
91
  - `mcp_account_user_admin`: `testAccount` on `hierarchy`, `account`, and `account_create` results;
90
92
  `testUser` on `users` / `user` results.
@@ -408,6 +408,31 @@ from that checkout). Its first real sync (or commit) needs `--create-variant-bra
408
408
  variants or a subscriber the platform treats it as a feature branch and refuses to land it, so creating a
409
409
  new variant branch is always a stated decision, never a side effect of a feature branch's name.
410
410
 
411
+ For a nested branch cut from an existing variant branch, seed the platform overlays explicitly after the git
412
+ branch is created and pushed:
413
+
414
+ ```bash
415
+ remits-cli components branch forked --copy-to sandbox --dry-run
416
+ remits-cli components branch forked --copy-to sandbox
417
+ remits-cli components sync --safe --branch sandbox
418
+ ```
419
+
420
+ The copy is a DB overlay bootstrap, not a git write. It makes inherited overlays from `forked` already
421
+ `storedCurrent` on the first `sandbox` sync, so the safe diff gate only has to account for the new branch's
422
+ real file changes.
423
+
424
+ **Copy before you subscribe, or the copy needs `--force`.** The copy refuses a target that already has
425
+ stored overlays *or* live subscribers, and that second guard fires even when the target has no overlays at
426
+ all — a branch somebody is already resolving is code those accounts are running right now, and replacing it
427
+ has to be a stated decision. Seeding first and subscribing after keeps the plain form working. The copy is
428
+ also all-or-nothing: the delete of any replaced overlays and every inserted row share one transaction, so a
429
+ failure leaves the target exactly as it was rather than half-populated.
430
+
431
+ Run it as `--dry-run` first. The plan reports `plannedCopies` — the number the real run will write — which
432
+ is not always the source branch's row count: a source overlay that has become identical to trunk in both
433
+ content and metadata is sparse, is not stored, and is listed as skipped instead of being silently dropped
434
+ from the total.
435
+
411
436
  **Trunk moving also invalidates a branch.** Variant sparseness compares branch content against *current*
412
437
  trunk, so a trunk change can make an overlay obsolete without the branch changing at all. A trunk sync
413
438
  drops the affected branches' cached sync verdicts, so the next `components sync` on the branch really
@@ -15,6 +15,8 @@
15
15
  - [Hierarchy-scoped tool reads](#hierarchy-scoped-tool-reads)
16
16
  - [Data Mode](#data-mode)
17
17
  - [Command Reference](#command-reference)
18
+ - [Reading a failing run](#reading-a-failing-run)
19
+ - [A pass count only means something within one data lane](#a-pass-count-only-means-something-within-one-data-lane)
18
20
  - [Verification envelopes](#verification-envelopes)
19
21
  - [Staging modes: workset vs full snapshot](#staging-modes-workset-vs-full-snapshot)
20
22
  - [Prod banners and retryable failures](#prod-banners-and-retryable-failures)
@@ -40,6 +42,10 @@ Treat host selection and data mode as two separate decisions:
40
42
 
41
43
  - `--base-url` chooses the Remits host: localhost vs a deployed environment.
42
44
  - `--data-mode` chooses the data segment on that host: `test` vs `prod`.
45
+ - A connection failure to `http://localhost:8080` only says the local back-stage app is not reachable.
46
+ It does **not** mean staging, testing, tokens, or investigation are blocked. If a deployed Remits host
47
+ is the right target, pass it explicitly with `--base-url` and keep the same account/branch/workspace
48
+ and `--data-mode` facts.
43
49
 
44
50
  Examples:
45
51
 
@@ -47,11 +53,20 @@ Examples:
47
53
  # Deployed prod host, but test data segment
48
54
  remits-cli tools --base-url https://your-prod-host --data-mode test
49
55
 
56
+ # Deployed prod host, test data, staged component verification
57
+ remits-cli components stage --workset --base-url https://your-prod-host
58
+ remits-cli test run --test "Statement Reader Calculations" --base-url https://your-prod-host
59
+
50
60
  # Localhost host, but prod data segment on that localhost instance
51
61
  remits-cli tool --base-url http://localhost:8080 --name mcp_account_view --data-mode prod
52
62
  ```
53
63
 
54
- Do not assume `--data-mode prod` implies the deployed prod host, or that `--data-mode test` implies localhost. If host matters, read `~/.remits-cli/sessions.json` first and pass `--base-url` explicitly.
64
+ Do not assume `--data-mode prod` implies the deployed prod host, or that `--data-mode test` implies
65
+ localhost. Likewise, do not assume localhost is required because the current checkout is a back-stage
66
+ repo or because a previous command used localhost. If host matters, read `~/.remits-cli/sessions.json`,
67
+ `remits-cli whoami`, or the recent `remits-cli evidence` world blocks, then pass `--base-url` explicitly.
68
+ Only call verification blocked after trying the intended reachable host and finding that no authenticated
69
+ session or usable network path exists.
55
70
 
56
71
  ### Tool Execution Lifecycle
57
72
 
@@ -236,12 +251,13 @@ remits-cli components branches [--json] # branche
236
251
  remits-cli components branch <name> [--json] # one branch: owner account, overridden / added / removed, drift flags, subscribers
237
252
  remits-cli components branch <name> --diff <componentId> --component-type <kind> [--json]
238
253
  remits-cli components branch <name> --subscribers [--json]
254
+ remits-cli components branch <name> --copy-to <newBranch> [--dry-run] [--force] [--json] # seed a new variant branch with this branch's stored overlays; force if target has overlays/subscribers
239
255
  remits-cli components branch <name> --subscribe <accountId> [--parent-account <id>] [--domain <host>] [--dry-run] [--confirm-primary-edge] # make an account resolve this branch
240
256
  remits-cli components branch <name> --unsubscribe <accountId> # return that account to trunk
241
257
  remits-cli components branch <name> --retire [--force] # delete the branch's overlays
242
- remits-cli test run --test <id|name> [--branch <stagingScope>] [--names "a|b"] [--watch true|false] [--data-mode test|prod] [--as-account <ID>] [--variant-branch <name|none>] [--json]
258
+ remits-cli test run --test <id|name> [--branch <stagingScope>] [--names "a|b"] [--watch true|false] [--wait true|false] [--data-mode test|prod] [--as-account <ID>] [--variant-branch <name|none>] [--json]
243
259
  remits-cli test status --task-id <taskId> [--branch <stagingScope>] [--data-mode test|prod] [--json] # falls back to the DURABLE run record once the live status expires
244
- remits-cli test runs [--test <id|name>] [--limit 20] [--compare] [--as-account <ID>] [--json] # durable run history: pass counts, live AI cost, lane, content hash
260
+ remits-cli test runs [--test <id|name>] [--limit 20] [--compare] [--all-lanes] [--as-account <ID>] [--json] # durable run history: pass counts, live AI cost, lane, content hash
245
261
  remits-cli test compare --base <taskId> --head <taskId> [--json] # per-case improved / regressed / changed, cost and world deltas
246
262
  remits-cli corpus import --manifest corpus-manifest.json [--corpus <name>] [--as-account <ID>] [--data-mode test|prod --confirm-prod] [--json] # seed an evaluation corpus: cases + immutable artifacts; idempotent by caseKey
247
263
  remits-cli corpus cases --corpus <name> [--split S] [--tag T|--tags T,U] [--key K|--keys K,L] [--include-values] [--include-retired] [--limit N] [--json]
@@ -252,6 +268,54 @@ remits-cli corpus consistency --corpus <name> --case <caseKey> [--json]
252
268
  remits-cli corpus retire --corpus <name> --case <caseKey> [--case ...] [--restore] [--json] # drop a case from future runs; past measurements stay readable
253
269
  ```
254
270
 
271
+ `test run` starts a server-side task immediately. By default the CLI process polls until the task is
272
+ terminal; `--watch false` disables websocket progress streaming but still waits. Pass `--wait false` to
273
+ return after launch with the task id. A single run with `--names "a|b|c"` selects cases into one suite task,
274
+ and those cases execute sequentially in declaration order. To overlap independent slow cases, launch
275
+ separate `remits-cli test run --names "<case>" --wait false` commands from the same staged lane and keep the
276
+ printed task ids; complete each proof with `test status --task-id <id>`. Avoid parallel runs for cases that
277
+ share mutable fixtures, suite-level side effects, or undeclared live AI/provider calls.
278
+
279
+ ### Reading a failing run
280
+
281
+ `test run` and `test status` print a **verdict**, not the run. The whole run object used to be dumped as
282
+ pretty JSON in human mode - 220 KB for one 92-case suite, most of it identifiers repeated once per case - so
283
+ the command an agent runs most often was the one that spent its remaining room to think. Every byte is still
284
+ there behind `--json`, and behind `test status --task-id <id> --json` once the live status expires.
285
+
286
+ What you get instead, and what to do with it:
287
+
288
+ ```text
289
+ Cases: 59 passed, 33 failed, 92 total
290
+
291
+ Failure roots (33 failed case(s), 21 distinct root(s)):
292
+ 10x assert statement.data.processingStatus == 'Analyzed'
293
+ e.g. bundled UK statements (+9 more)
294
+ 2x assert statement.data.feeBreakdownChecked == true
295
+ e.g. 7851 Flat Rate (+1 more)
296
+ Largest root first: remits-cli test run --test "Statement Reader Calculations" --names "bundled UK statements"
297
+ ```
298
+
299
+ **Thirty-three failures are not thirty-three problems.** Cases are grouped by their assertion root - the
300
+ failure with the particulars (ids, numbers, Groovy's `Expression:`/`Values:` decoration) removed - largest
301
+ group first, and one pivot block is printed per root rather than per case. Fix the largest root against one
302
+ named case, then re-run the whole suite. Re-running the suite before you have a root is how a session spends
303
+ an afternoon at the same pass count.
304
+
305
+ ### A pass count only means something within one data lane
306
+
307
+ `test runs` is scoped to the data lane of the command by default, and says so:
308
+
309
+ ```text
310
+ Data lane: test (6 more run(s) in the other lane; --all-lanes to include)
311
+ ```
312
+
313
+ A run's `dataMode` **is** the lane it was written in, so `64/92` in the prod lane and `59/92` in the test lane
314
+ are two different facts, not a regression. `--compare` (and `test runs --compare`) picks the newest two runs
315
+ of the same **shape** - same data lane, branch, workspace and case count - and says how many newer runs it
316
+ skipped; it refuses rather than comparing across worlds. `test status --data-mode prod` on a run recorded in
317
+ the test lane returns the run's real lane and now says that is what happened.
318
+
255
319
  A run labelled **`AI MOCKED`** replayed every AI turn from `aiMock`: it produces the same outcomes, scores and
256
320
  metrics a measured run would, so read that label before treating green as evidence. `compare` warns when base and
257
321
  head used different AI modes; `consistency` warns before calling a case unstable when its runs mixed modes.
@@ -16,6 +16,7 @@
16
16
  - [When staged overrides apply](#when-staged-overrides-apply)
17
17
  - [Diagnosing which version is in play](#diagnosing-which-version-is-in-play)
18
18
  - [Working alongside other agents: the staging WORKSPACE](#working-alongside-other-agents-the-staging-workspace)
19
+ - [Keeping your remits-cli current](#keeping-your-remits-cli-current)
19
20
  - [A lane holds an OVERLAY; your workset is a different number](#a-lane-holds-an-overlay-your-workset-is-a-different-number)
20
21
  - [Stage / sync / clear with remits-cli](#stage--sync--clear-with-remits-cli)
21
22
  - [Stale after sync / commit (the in-memory compile cache)](#stale-after-sync--commit-the-in-memory-compile-cache)
@@ -250,6 +251,18 @@ index with `remits-cli workstream status`.
250
251
  `null` means unknown — a lane staged by an older CLI, or a branch whose last sync is not recorded —
251
252
  never "fresh".
252
253
 
254
+ ### Keeping your remits-cli current
255
+
256
+ The diagnostics in these references only exist in the CLI that prints them. An older install does not print a
257
+ worse version of them — it prints nothing, with no error, so nothing tells you what you are not being shown.
258
+ `components stage` and `test run` warn when yours is behind:
259
+
260
+ ```text
261
+ [remits-cli 0.1.120 -> 0.1.136] ... npm install -g @remits/remits-cli@latest
262
+ ```
263
+
264
+ Upgrade when you see it. Auto-update handles it for you on most commands, including `--json` ones.
265
+
253
266
  ### A lane holds an OVERLAY; your workset is a different number
254
267
 
255
268
  This is the distinction that decides whether a lane is legible to anyone but you.
@@ -20,6 +20,7 @@
20
20
  - [Stage your workset, not the whole repo](#stage-your-workset-not-the-whole-repo)
21
21
  - [Three numbers, three questions](#three-numbers-three-questions)
22
22
  - [Step 4: Verify the Change](#step-4-verify-the-change)
23
+ - [Many failures are usually few causes](#many-failures-are-usually-few-causes)
23
24
  - [Step 5: Iterate If Needed](#step-5-iterate-if-needed)
24
25
  - [Step 6: Update Documentation](#step-6-update-documentation)
25
26
  - [Temporary Experiment Workflow](#temporary-experiment-workflow)
@@ -171,7 +172,7 @@ wants to proceed rather than reporting the work as done.
171
172
  This is how every development task should flow:
172
173
 
173
174
  #### Step 1: Understand the Request
174
- Read the user's request. If you may need a repo other than the current one, read `~/.remits-cli/account-repos.json` first. Then review `account-info.json` and `README.md` to understand what components exist and how they relate. Read the source of any component you'll modify before changing it.
175
+ Read the user's request. If you may need a repo other than the current one, read `~/.remits-cli/account-repos.json` first. Then review `account-boot.json` and `README.md` to understand the account shape and component routing inventory. Read `account-info.json` only when the compact file lacks detail you need. Read the source of any component you'll modify before changing it.
175
176
 
176
177
  **Establish a steady git baseline before the first edit.** The platform syncs from the GitHub remote, not
177
178
  from your local files, and a worktree can be stale even when its staging lane is isolated. In the checkout
@@ -190,14 +191,14 @@ that the branch is behind trunk or that overlays were computed from an old SHA,
190
191
  code. A workspace prevents staged-cache collisions; it does not make a stale branch current.
191
192
 
192
193
  **Establish the account's shape too, not just its components.** Read the `resolution` block in
193
- `account-info.json` (or `mcp_account_view`): the account `type` decides whether this repo is even the right
194
+ `account-boot.json` (or `mcp_account_view`): the account `type` decides whether this repo is even the right
194
195
  place to change code, `resolution.relationships` shows whether the account has more than one parent (and
195
196
  which link carries a branch/namespace/host), and `resolvedDatabaseName` tells you where its data actually
196
197
  lands. See `account-targeting.md` and `features/account-management.md` (`mcp_get_guide`).
197
198
 
198
199
  **Also establish which world you are working in.** `remits-cli components status` reports whether the
199
200
  working tree is a **trunk** checkout or a **variant branch** checkout — which decides both what your test
200
- runs resolve and what a sync writes. If `account-info.json` carries a `componentBranches` section, branch
201
+ runs resolve and what a sync writes. If `account-boot.json` carries a `componentBranches` section, branch
201
202
  variants of these components exist: editing an origin component will drift them, so check
202
203
  `remits-cli components branches` before changing shared code. See `branch-variants.md`.
203
204
 
@@ -264,7 +265,7 @@ the file. Omit the key; the platform fills it in on sync:
264
265
  ```yaml
265
266
  # components/embeddables/new_MerchantPortal.meta.yml — no `id:` yet
266
267
  name: Merchant Portal
267
- summary: One-line statement of what this component is for. This is the compact text account-info.json uses first.
268
+ summary: One-line statement of what this component is for. This is the compact text account-boot.json uses first.
268
269
  description: |
269
270
  Longer technical description with line-number references to the key logic.
270
271
  path: /page/merchant-portal # Readers and Embeddables only
@@ -411,14 +412,54 @@ Tests run on the platform against your staged snapshot. They stream results in r
411
412
  id + content hash, platform sync vs local HEAD. If it is not the world you meant, stop: the result will be
412
413
  about a different world.
413
414
 
415
+ The platform launch is asynchronous. By default `remits-cli test run` waits for its task to finish, and any
416
+ cases selected by `--names "a|b"` execute sequentially inside that suite task. If several slow cases are
417
+ independent, stage once, then start separate `test run --names "<case>" --wait false` commands from the same
418
+ lane and keep the task ids they print. Complete each proof with
419
+ `remits-cli test status --task-id <id>`; that terminal status read records and attaches the final `test_run`
420
+ evidence. Do not parallelize cases that mutate the same fixture, rely on shared suite setup state, or make
421
+ undeclared live AI/provider calls.
422
+
414
423
  If cases fail, read the printed summary and pivots first: each case has an `outcome`
415
424
  (`failed`, `error`, `budget_exceeded`, `provider_unavailable`, …) with its reason, timing, trace id, AI usage
416
425
  (live vs mocked, cost), bounded `report(...)` diagnostics, live HTTP signals, and resolved component
417
426
  provenance. Fix the code, re-stage, and re-run only after those pivots explain the failure.
418
427
 
428
+ ##### Many failures are usually few causes
429
+
430
+ When several cases fail, the run leads with **failure roots** — the failures grouped by their assertion with
431
+ the particulars (ids, numbers, Groovy's `Expression:`/`Values:` decoration) stripped out, largest group first,
432
+ one pivot block per root rather than per case:
433
+
434
+ ```text
435
+ Failure roots (33 failed case(s), 21 distinct root(s)):
436
+ 10x assert statement.data.processingStatus == 'Analyzed'
437
+ e.g. bundled UK statements (+9 more)
438
+ Largest root first: remits-cli test run --test "Statement Reader Calculations" --names "bundled UK statements"
439
+ ```
440
+
441
+ **Work the largest root against ONE named case, then re-run the suite.** A full re-run after every edit is the
442
+ most expensive way to learn nothing: a real account spent two days and twenty-five runs holding a suite at
443
+ 56–59 of 92 while ten of its thirty-three failures were one cause. The narrow run is seconds, tells you
444
+ whether the cause moved, and leaves your context for the actual reasoning.
445
+
446
+ Two failure roots mean "stop and look elsewhere", not "iterate harder":
447
+
448
+ - **`aiMock ctx.replay(...) found no usable stored provider response`** — the message says which of three
449
+ things happened. *No rows at all* for that sessionId means the stored session this fixture borrows is not
450
+ on this platform and will not come back; the case cannot pass until the mock builds its own response with
451
+ `ctx.toolCall(...)` / `ctx.content(...)`. Do not re-run it. (Replaying a session that IS there renews it,
452
+ so a suite that runs regularly keeps its fixtures.)
453
+ - **`Method too large` / `Class too large` / a synthetic `_closureNN`** — a JVM limit on one method body, not
454
+ a bug in the line it names. Staging now warns *before* the refusal and names the closure's source line span.
455
+ Split that body; see `development-guide.md` → *Keep Component Bodies Split*.
456
+
419
457
  Every finished run is recorded durably: `remits-cli test status --task-id <id>` answers after the live status
420
- expires, `remits-cli test runs --test <name> --compare` compares the latest run with the previous one, and
421
- `remits-cli test compare --base <id> --head <id>` compares any two. For evaluation suites (a `corpus(name)` of
458
+ expires, `remits-cli test runs --test <name> --compare` compares the latest run with the newest **comparable**
459
+ one, and `remits-cli test compare --base <id> --head <id>` compares any two. Comparable means the same data
460
+ lane, branch, workspace and case count: a run's `dataMode` is the lane it was written in, so `64/92` in prod
461
+ and `59/92` in test are two facts and not a trend. `test runs` lists one lane at a time and names it
462
+ (`--all-lanes` to see both). For evaluation suites (a `corpus(name)` of
422
463
  cases seeded with `remits-cli corpus import`, one case per corpus case, intentional live AI inside
423
464
  `withAiBudget(...)`, measurements compared with `remits-cli corpus compare` / `corpus consistency`), read
424
465
  `guides/test-components.md` → *Evaluation Suites And Corpora*.
@@ -527,7 +568,7 @@ If the work is tied to a support ticket:
527
568
 
528
569
  Before committing, update metadata so the next session understands what changed:
529
570
 
530
- 1. **`.meta.yml` sidecars** — Update `summary`, `description`, and `mermaid` for each modified component. `summary` is what drives the compact component description in generated `account-info.json`; `description` is the fallback when no summary is set and is capped in that file. Preserve or explicitly revise dated decision notes; do not delete the evidence the next agent needs.
571
+ 1. **`.meta.yml` sidecars** — Update `summary`, `description`, and `mermaid` for each modified component. `summary` is what drives the compact component description in generated `account-boot.json`; `description` is the fallback when no summary is set and is capped there. Preserve or explicitly revise dated decision notes; do not delete the evidence the next agent needs.
531
572
  2. **`README.md`** — If the change affects account-level capabilities or workflows.
532
573
  3. **New components** — Always fill in `.meta.yml` immediately.
533
574
 
@@ -76,6 +76,9 @@ remits-cli tool --account-id 21 --as-account 37 --target-account 37 --name mcp_r
76
76
  # poll by run id: --account-id 21 --as-account 37 --target-account 37 --input '{"controlAction":"status","actionRunId":"my-stable-run-id"}'
77
77
  ```
78
78
 
79
+ That is the CLI Action-run surface. The command is still `remits-cli tool`; this CLI does not have a
80
+ separate top-level `remits-cli action`, `remits-cli actions`, or `remits-cli run action` wrapper.
81
+
79
82
  For long tools that lack their own async mode, use the CLI transport async (`--async true`), optionally with
80
83
  `--wait true` to poll locally, and `remits-cli tool status --call-id <callId>`. Do not stack both mechanisms
81
84
  (see `command-reference.md` → *Tool Execution Lifecycle*). Use `--timeout-ms <ms>` only to adjust the per-request client timeout; it is
@@ -403,12 +406,18 @@ Front-stage references:
403
406
  | `page` | no | 1-based page number. Default: `1` |
404
407
  | `pageSize` | no | Results per page. Default: `25`, max: `100` |
405
408
  | `summaryOnly` | no | Detail mode: return the MAP (record index + stats + timeline) with no payloads. Same as `parts:["index"]` |
406
- | `parts` | no | Detail parts: `index`, `stats`, `transcript`, `tool_calls`, and record sections `full`, `conversation_messages`, `system`, `response`, `tools` |
409
+ | `compact` | no | Detail mode: after reading the MAP, open selected records without repeating the `records` index and `timeline`; keeps `groupingSummary`, `representativeSession`, `stats`, selected ids, and requested payload sections |
410
+ | `payloadOnly` | no | Alias for `compact` |
411
+ | `parts` | no | Detail parts: `index`, `map`, `records`, `stats`, `transcript`, `tool_calls`, and record sections `full`, `conversation_messages`, `system`, `response`, `tools` |
407
412
  | `sections` | no | Alias for `parts` |
408
413
  | `recordIds` | no | Detail mode: open only these request/response record IDs |
409
414
  | `toolCallIds` | no | Detail mode: return the FULL exact input/result from `ai_tool_call` for these tool-call ids (what the tool PRODUCED — see the lens caveat) |
410
415
  | `consolidateContext` | no | When `true`, collapses repeated XML-like prompt context into a consolidated section |
411
416
 
417
+ Detail responses expose both lenses: `groupingSummary` is the authoritative grouping-wide summary
418
+ (counts, lanes, cost, status), while `representativeSession` names the concrete session used for
419
+ owner/runtime metadata. `session` remains only as a compatibility alias for older callers.
420
+
412
421
  Spend: each search row carries `liveRequestCount`, `mockedRequestCount`, `lanes`, `testTaskId` and
413
422
  `estimatedCost`/`estimatedCostUsd` = **live spend only** (a mocked turn is never priced, even when it replays a
414
423
  recording with a cost). `pageTotals` sums the page. Turns recorded before the platform stored the mock flag are
@@ -416,12 +425,16 @@ recording with a cost). `pageTotals` sums the page. Turns recorded before the pl
416
425
  is refused rather than silently widened; the applied bounds are echoed as `filters.rangeStart`/`rangeEnd`.
417
426
  To audit one Test run: `{"action":"search","testTaskId":"<taskId>","pageSize":100}`.
418
427
 
419
- Audit flow: `action:"search"` to find the grouping → `action:"detail"` + `summaryOnly:true` for the MAP → re-call detail with `recordIds`/`toolCallIds` + `parts` to open exactly what you need. Prefer the map → open flow over a full-detail dump. Remember the two-lens rule: `tool_calls`/`toolCallIds` is what the tool PRODUCED; `conversation_messages`/`transcript` is what the AI CONSUMED (after any `_offload`/`_hideResult`/`_message`/supersede/evict transform).
428
+ Audit flow: `action:"search"` to find the grouping → `action:"detail"` + `summaryOnly:true` for the MAP → re-call detail with `recordIds`/`toolCallIds` + `parts` to open exactly what you need, usually with `compact:true` once the map has chosen the row. Prefer the map → open flow over a full-detail dump. Remember the two-lens rule: `tool_calls`/`toolCallIds` is what the tool PRODUCED; `conversation_messages`/`transcript` is what the AI CONSUMED (after any `_offload`/`_hideResult`/`_message`/supersede/evict transform).
420
429
 
421
430
  ### `mcp_run_action`
422
431
  Run an Action on a target account, with explicit prod/test data mode, optional staged branch resolution, and
423
432
  staged-vs-DB provenance in the result.
424
433
 
434
+ Invoke it with `remits-cli tool --name mcp_run_action`. Despite the natural shorthand "run an Action", there
435
+ is no separate top-level `remits-cli action` / `actions` command and no `remits-cli run action` wrapper in
436
+ this CLI build.
437
+
425
438
  Describe the Action first when the input shape is not obvious. This does not execute the Action:
426
439
 
427
440
  ```bash
@@ -832,6 +845,10 @@ Common uses:
832
845
  - `remits-cli components branch <name> --diff <id> --component-type <kind>` — compare one variant against
833
846
  current trunk.
834
847
  - `remits-cli components branch <name> --subscribers` — list accounts resolving that branch.
848
+ - `remits-cli components branch <name> --copy-to <newBranch> [--dry-run] [--force]` — seed a new variant
849
+ branch with the source branch's stored overlays before the first safe sync of the new branch. Dry-run
850
+ reports the copy plan without writes; force is required when the target already has overlays or live
851
+ subscribers.
835
852
 
836
853
  ### `mcp_cache`
837
854
  Bounded read-only investigation of the platform Redis keyspace — the way to see exactly what a staged