@remits/remits-cli 0.1.135 → 0.1.137

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@remits/remits-cli",
3
- "version": "0.1.135",
3
+ "version": "0.1.137",
4
4
  "description": "Local CLI for auth, component sync, and live test execution against Remits",
5
5
  "license": "MIT",
6
6
  "private": false,
@@ -48,6 +48,7 @@ reference below, load it before you act, not after something surprises you.
48
48
  | register this session as an agent, or ask why a routed ticket never started | `agent-sessions.md` | `serve` vs `register`, what registration does, worker spawning per agent kind, routing order, the control center, multi-session auth |
49
49
  | investigate live behavior | `investigation.md` | which tool reads which record, correlation keys, reading a record's `content`, HTTP audits, AI activity, node/`localMode`, the production support flows |
50
50
  | call any `mcp_*` tool | `tool-reference.md` | every tool's parameters, semantics, and traps — read the tool's entry before building its input |
51
+ | run, describe, poll, or interrupt an Action from the CLI | `tool-reference.md` → `mcp_run_action` | Action execution through `remits-cli tool --name mcp_run_action`; there is no separate top-level `remits-cli action` / `run action` wrapper |
51
52
  | reason about repo discovery, auth state, or what a command actually sent | `cli-state.md` | the global `~/.remits-cli/` control plane vs per-repo `./.remits-cli/`, and which file answers which question |
52
53
  | need an exact flag, or the auth / host / async surface | `command-reference.md` | authentication, host vs data mode, the two async mechanisms, hierarchy-scoped reads, the full command list, prod banners |
53
54
  | give up on something that misbehaved | `troubleshooting.md` | the symptom→fix table, the two kinds of escalation, and the escalation bundle |
@@ -162,11 +163,18 @@ reference named after it.
162
163
  component (preferred, because it becomes regression protection) or a browser flow through
163
164
  `remits-cli token`. "Just do it" and "that's fine, commit it" are not evidence. If you genuinely
164
165
  cannot verify, say what you would need and ask. (`development-loop.md`)
165
- - **For concrete user workflows, start a verification envelope before you edit.** `remits-cli verify
166
- start --summary "..."` records the account/source/lane tuple and a manifest when you have one.
167
- Stage, test, token, tool, and sync commands attach evidence automatically while the envelope is
168
- active; finish from `remits-cli verify report`, which separates verified claims from missing or stale
169
- evidence. (`development-loop.md`, `component-resolution.md`)
166
+ - **`remits-cli evidence` answers "what have I actually run, and in which world?"** Every stage, test,
167
+ token, tool and sync appends one world-stamped line automatically, per actor. Nothing to start,
168
+ nothing to satisfy. Read it before you re-run something, and quote it when you report what you
169
+ proved. (`development-loop.md`)
170
+ - **Verification envelopes are OPTIONAL and exist for one job: a verdict someone else will rely on.**
171
+ Start one when a ticket, a human, or a handoff needs a checkable "these specific things are true" —
172
+ not as a routine step before editing. `remits-cli verify start --summary "..."` then
173
+ `remits-cli verify claim <id> --text "..." [--test "<suite>"]` names what must be true; add
174
+ `--claim <id>` to the command that proves it; `remits-cli verify report` gives the verdict. An
175
+ envelope with no claims is simply an evidence log, which is a complete state — not something to
176
+ chase. If you find yourself opening an envelope to answer a question about your own work, use
177
+ `remits-cli evidence` instead. (`development-loop.md`, `component-resolution.md`)
170
178
  - **Read the component's `.meta.yml` before changing behavior.** Sidecar descriptions can be dated
171
179
  decision records. Before changing a displayed value, helper, calculation, schema field, or prompt
172
180
  contract, check the sidecar and either preserve its decision or explicitly supersede it.
@@ -209,8 +217,9 @@ remits-cli tools # which tools this account actually has (tools
209
217
  ```
210
218
 
211
219
  plus the repo's `account-info.json` → `resolution` block for the account's shape.
212
- For a workflow-shaped request, also run `remits-cli verify start --summary "..."` once the target tuple
213
- is understood, then keep that envelope active through stage/test/token/sync.
220
+ For a workflow-shaped request, rely on the automatic `remits-cli evidence` trail unless someone else needs
221
+ a checkable verdict. Only then start `remits-cli verify start --summary "..."`, declare claims, and keep
222
+ that envelope active through stage/test/token/sync.
214
223
 
215
224
  If any command returns 401, run `remits-cli auth` (with the same `--base-url` if you were targeting a
216
225
  non-default host).
@@ -408,6 +408,31 @@ from that checkout). Its first real sync (or commit) needs `--create-variant-bra
408
408
  variants or a subscriber the platform treats it as a feature branch and refuses to land it, so creating a
409
409
  new variant branch is always a stated decision, never a side effect of a feature branch's name.
410
410
 
411
+ For a nested branch cut from an existing variant branch, seed the platform overlays explicitly after the git
412
+ branch is created and pushed:
413
+
414
+ ```bash
415
+ remits-cli components branch forked --copy-to sandbox --dry-run
416
+ remits-cli components branch forked --copy-to sandbox
417
+ remits-cli components sync --safe --branch sandbox
418
+ ```
419
+
420
+ The copy is a DB overlay bootstrap, not a git write. It makes inherited overlays from `forked` already
421
+ `storedCurrent` on the first `sandbox` sync, so the safe diff gate only has to account for the new branch's
422
+ real file changes.
423
+
424
+ **Copy before you subscribe, or the copy needs `--force`.** The copy refuses a target that already has
425
+ stored overlays *or* live subscribers, and that second guard fires even when the target has no overlays at
426
+ all — a branch somebody is already resolving is code those accounts are running right now, and replacing it
427
+ has to be a stated decision. Seeding first and subscribing after keeps the plain form working. The copy is
428
+ also all-or-nothing: the delete of any replaced overlays and every inserted row share one transaction, so a
429
+ failure leaves the target exactly as it was rather than half-populated.
430
+
431
+ Run it as `--dry-run` first. The plan reports `plannedCopies` — the number the real run will write — which
432
+ is not always the source branch's row count: a source overlay that has become identical to trunk in both
433
+ content and metadata is sparse, is not stored, and is listed as skipped instead of being silently dropped
434
+ from the total.
435
+
411
436
  **Trunk moving also invalidates a branch.** Variant sparseness compares branch content against *current*
412
437
  trunk, so a trunk change can make an overlay obsolete without the branch changing at all. A trunk sync
413
438
  drops the affected branches' cached sync verdicts, so the next `components sync` on the branch really
@@ -15,6 +15,8 @@
15
15
  - [Hierarchy-scoped tool reads](#hierarchy-scoped-tool-reads)
16
16
  - [Data Mode](#data-mode)
17
17
  - [Command Reference](#command-reference)
18
+ - [Reading a failing run](#reading-a-failing-run)
19
+ - [A pass count only means something within one data lane](#a-pass-count-only-means-something-within-one-data-lane)
18
20
  - [Verification envelopes](#verification-envelopes)
19
21
  - [Staging modes: workset vs full snapshot](#staging-modes-workset-vs-full-snapshot)
20
22
  - [Prod banners and retryable failures](#prod-banners-and-retryable-failures)
@@ -236,12 +238,13 @@ remits-cli components branches [--json] # branche
236
238
  remits-cli components branch <name> [--json] # one branch: owner account, overridden / added / removed, drift flags, subscribers
237
239
  remits-cli components branch <name> --diff <componentId> --component-type <kind> [--json]
238
240
  remits-cli components branch <name> --subscribers [--json]
241
+ remits-cli components branch <name> --copy-to <newBranch> [--dry-run] [--force] [--json] # seed a new variant branch with this branch's stored overlays; force if target has overlays/subscribers
239
242
  remits-cli components branch <name> --subscribe <accountId> [--parent-account <id>] [--domain <host>] [--dry-run] [--confirm-primary-edge] # make an account resolve this branch
240
243
  remits-cli components branch <name> --unsubscribe <accountId> # return that account to trunk
241
244
  remits-cli components branch <name> --retire [--force] # delete the branch's overlays
242
- remits-cli test run --test <id|name> [--branch <stagingScope>] [--names "a|b"] [--watch true|false] [--data-mode test|prod] [--as-account <ID>] [--variant-branch <name|none>] [--json]
245
+ remits-cli test run --test <id|name> [--branch <stagingScope>] [--names "a|b"] [--watch true|false] [--wait true|false] [--data-mode test|prod] [--as-account <ID>] [--variant-branch <name|none>] [--json]
243
246
  remits-cli test status --task-id <taskId> [--branch <stagingScope>] [--data-mode test|prod] [--json] # falls back to the DURABLE run record once the live status expires
244
- remits-cli test runs [--test <id|name>] [--limit 20] [--compare] [--as-account <ID>] [--json] # durable run history: pass counts, live AI cost, lane, content hash
247
+ remits-cli test runs [--test <id|name>] [--limit 20] [--compare] [--all-lanes] [--as-account <ID>] [--json] # durable run history: pass counts, live AI cost, lane, content hash
245
248
  remits-cli test compare --base <taskId> --head <taskId> [--json] # per-case improved / regressed / changed, cost and world deltas
246
249
  remits-cli corpus import --manifest corpus-manifest.json [--corpus <name>] [--as-account <ID>] [--data-mode test|prod --confirm-prod] [--json] # seed an evaluation corpus: cases + immutable artifacts; idempotent by caseKey
247
250
  remits-cli corpus cases --corpus <name> [--split S] [--tag T|--tags T,U] [--key K|--keys K,L] [--include-values] [--include-retired] [--limit N] [--json]
@@ -252,6 +255,54 @@ remits-cli corpus consistency --corpus <name> --case <caseKey> [--json]
252
255
  remits-cli corpus retire --corpus <name> --case <caseKey> [--case ...] [--restore] [--json] # drop a case from future runs; past measurements stay readable
253
256
  ```
254
257
 
258
+ `test run` starts a server-side task immediately. By default the CLI process polls until the task is
259
+ terminal; `--watch false` disables websocket progress streaming but still waits. Pass `--wait false` to
260
+ return after launch with the task id. A single run with `--names "a|b|c"` selects cases into one suite task,
261
+ and those cases execute sequentially in declaration order. To overlap independent slow cases, launch
262
+ separate `remits-cli test run --names "<case>" --wait false` commands from the same staged lane and keep the
263
+ printed task ids; complete each proof with `test status --task-id <id>`. Avoid parallel runs for cases that
264
+ share mutable fixtures, suite-level side effects, or undeclared live AI/provider calls.
265
+
266
+ ### Reading a failing run
267
+
268
+ `test run` and `test status` print a **verdict**, not the run. The whole run object used to be dumped as
269
+ pretty JSON in human mode - 220 KB for one 92-case suite, most of it identifiers repeated once per case - so
270
+ the command an agent runs most often was the one that spent its remaining room to think. Every byte is still
271
+ there behind `--json`, and behind `test status --task-id <id> --json` once the live status expires.
272
+
273
+ What you get instead, and what to do with it:
274
+
275
+ ```text
276
+ Cases: 59 passed, 33 failed, 92 total
277
+
278
+ Failure roots (33 failed case(s), 21 distinct root(s)):
279
+ 10x assert statement.data.processingStatus == 'Analyzed'
280
+ e.g. bundled UK statements (+9 more)
281
+ 2x assert statement.data.feeBreakdownChecked == true
282
+ e.g. 7851 Flat Rate (+1 more)
283
+ Largest root first: remits-cli test run --test "Statement Reader Calculations" --names "bundled UK statements"
284
+ ```
285
+
286
+ **Thirty-three failures are not thirty-three problems.** Cases are grouped by their assertion root - the
287
+ failure with the particulars (ids, numbers, Groovy's `Expression:`/`Values:` decoration) removed - largest
288
+ group first, and one pivot block is printed per root rather than per case. Fix the largest root against one
289
+ named case, then re-run the whole suite. Re-running the suite before you have a root is how a session spends
290
+ an afternoon at the same pass count.
291
+
292
+ ### A pass count only means something within one data lane
293
+
294
+ `test runs` is scoped to the data lane of the command by default, and says so:
295
+
296
+ ```text
297
+ Data lane: test (6 more run(s) in the other lane; --all-lanes to include)
298
+ ```
299
+
300
+ A run's `dataMode` **is** the lane it was written in, so `64/92` in the prod lane and `59/92` in the test lane
301
+ are two different facts, not a regression. `--compare` (and `test runs --compare`) picks the newest two runs
302
+ of the same **shape** - same data lane, branch, workspace and case count - and says how many newer runs it
303
+ skipped; it refuses rather than comparing across worlds. `test status --data-mode prod` on a run recorded in
304
+ the test lane returns the run's real lane and now says that is what happened.
305
+
255
306
  A run labelled **`AI MOCKED`** replayed every AI turn from `aiMock`: it produces the same outcomes, scores and
256
307
  metrics a measured run would, so read that label before treating green as evidence. `compare` warns when base and
257
308
  head used different AI modes; `consistency` warns before calling a case unstable when its runs mixed modes.
@@ -344,10 +395,24 @@ shown is the latest run's, with the distinct worlds listed beneath it.
344
395
 
345
396
  ### Verification envelopes
346
397
 
347
- Use a verification envelope for workflow-shaped work: concrete user journeys, browser-facing changes,
348
- branch variants, subscriber/forked accounts, production-vs-test lane questions, support tickets, and
349
- multi-agent work. Start it after you know the target account/branch/workspace/data-mode tuple and before
350
- the first edit:
398
+ **First, the thing you probably want instead.** `remits-cli evidence` prints what you have actually run
399
+ and in which world, grouped by world, with no envelope and nothing to start:
400
+
401
+ ```bash
402
+ remits-cli evidence # this actor's trail
403
+ remits-cli evidence --json # structured, for your final report
404
+ ```
405
+
406
+ Every stage, test, token, tool and sync appends to it automatically. Use it for "what have I already
407
+ run?", "did that execute in the prod lane?" and "what do I put in my final message?".
408
+
409
+ **Verification envelopes are optional, and exist for a verdict someone ELSE will rely on** — a support
410
+ ticket, a human who asked you to prove specific things, a handoff another agent will act on. They are not
411
+ a routine step before editing. An envelope you started for your own benefit is nearly always
412
+ `remits-cli evidence` in disguise, and it costs you turns that prove nothing.
413
+
414
+ When a verdict is genuinely wanted, start it after you know the target
415
+ account/branch/workspace/data-mode tuple:
351
416
 
352
417
  ```bash
353
418
  remits-cli workspace use --auto
@@ -355,6 +420,35 @@ remits-cli components status
355
420
  remits-cli verify start --summary "Hosted upload updates an existing profile" --manifest acceptance.json
356
421
  ```
357
422
 
423
+ **Name what must be true, or there is no verdict to reach.** An envelope with no `claims` and no
424
+ `requiredEvidence` reports `evidence_only` — a world-stamped log. That is a complete, final state, not a
425
+ partial one, and nothing about it is outstanding. When you DO want a verdict, a manifest file is one way to
426
+ declare the contract; `verify claim` is the one-command way, and it works on a live envelope:
427
+
428
+ ```bash
429
+ remits-cli verify claim market-filter --text "US market scoring excludes AU and GB statements"
430
+ remits-cli verify claim pilot-green --text "the pilot suite passes" --test "Acquirer Pilot"
431
+ ```
432
+
433
+ Prove a claim by naming it on the command that already proves it. `--claim <id>` is stamped onto the
434
+ evidence packet every command sends, so it works on `verify test`, `token`, `tool`, `stage`, `sync` and
435
+ `attach` alike, and takes a comma-separated list:
436
+
437
+ ```bash
438
+ remits-cli verify test --test "Acquirer Pilot" --claim pilot-green
439
+ remits-cli verify tool --name mcp_run_action --as-account 36 --claim market-filter
440
+ remits-cli verify attach --claim market-filter --note "AU statements excluded and counted"
441
+ ```
442
+
443
+ A claim with no shape is satisfied by any successful packet tagged with its id. A claim that names a
444
+ `--test` suite (or `--packet-type`) must be proven by that shape. Re-declaring an id replaces it, so
445
+ tightening a claim is also one command.
446
+
447
+ **When an envelope with a contract will not close, read the report — do not start another one.** Several
448
+ open envelopes for the same goal, or workspaces named `...-r7`/`-r8`/`-r9`, is the loop signature, and
449
+ `activity inspect` reports both as smells. Close what you are not finishing: `verify supersede --envelope
450
+ <old> --superseded-by <new>` or `verify abandon --envelope <old> --reason "false start"`.
451
+
358
452
  An active envelope is stored per command world (`baseUrl + accountId + git branch + workspace`)
359
453
  inside the active local actor's mirror; **data mode is deliberately not part of the active pointer key**:
360
454
  `./.remits-cli/actors/<local-agent>/verification/active-contexts.json`. The legacy flat
@@ -367,7 +461,7 @@ history:
367
461
 
368
462
  ```bash
369
463
  remits-cli verify stage --workset
370
- remits-cli verify test --test "Adyen Import Recovery" --names "browser upload recovery"
464
+ remits-cli verify test --test "Adyen Import Recovery" --names "browser upload recovery" --claim recovery
371
465
  remits-cli verify token --path /page/pricing-config --as-account 21 --data-mode test
372
466
  remits-cli verify sync --safe
373
467
  remits-cli verify report
@@ -435,10 +529,10 @@ proof noun: `passed` for a Test case, `measured` for corpus comparisons, `synced
435
529
 
436
530
  Final claims should come from `remits-cli verify report`. Treat `Verified` as the acceptance boundary.
437
531
  `Additional evidence` is useful handoff context, but it does not satisfy a missing required packet unless
438
- the report lists it under `Verified`. If the report says `partially_verified`, stale, or missing evidence,
439
- say that plainly instead of widening the claim. In particular, staged proof is not committed variant/trunk
440
- proof, a token is not browser proof, and a direct DOM or Alpine state mutation is not the same as a user
441
- action.
532
+ the report lists it under `Verified`. If the report says `evidence_only`, no verdict was requested; if it
533
+ says `partially_verified`, stale, or missing evidence, say that plainly instead of widening the claim. In
534
+ particular, staged proof is not committed variant/trunk proof, a token is not browser proof, and a direct
535
+ DOM or Alpine state mutation is not the same as a user action.
442
536
 
443
537
  Evidence from the wrong world is excluded before it can satisfy a requirement. When the manifest declares
444
538
  fields such as `repoAccountId`, `gitBranch`, `componentBranch`, `workspace`, or `dataMode`, the report names
@@ -16,6 +16,7 @@
16
16
  - [When staged overrides apply](#when-staged-overrides-apply)
17
17
  - [Diagnosing which version is in play](#diagnosing-which-version-is-in-play)
18
18
  - [Working alongside other agents: the staging WORKSPACE](#working-alongside-other-agents-the-staging-workspace)
19
+ - [Keeping your remits-cli current](#keeping-your-remits-cli-current)
19
20
  - [A lane holds an OVERLAY; your workset is a different number](#a-lane-holds-an-overlay-your-workset-is-a-different-number)
20
21
  - [Stage / sync / clear with remits-cli](#stage--sync--clear-with-remits-cli)
21
22
  - [Stale after sync / commit (the in-memory compile cache)](#stale-after-sync--commit-the-in-memory-compile-cache)
@@ -250,6 +251,18 @@ index with `remits-cli workstream status`.
250
251
  `null` means unknown — a lane staged by an older CLI, or a branch whose last sync is not recorded —
251
252
  never "fresh".
252
253
 
254
+ ### Keeping your remits-cli current
255
+
256
+ The diagnostics in these references only exist in the CLI that prints them. An older install does not print a
257
+ worse version of them — it prints nothing, with no error, so nothing tells you what you are not being shown.
258
+ `components stage` and `test run` warn when yours is behind:
259
+
260
+ ```text
261
+ [remits-cli 0.1.120 -> 0.1.136] ... npm install -g @remits/remits-cli@latest
262
+ ```
263
+
264
+ Upgrade when you see it. Auto-update handles it for you on most commands, including `--json` ones.
265
+
253
266
  ### A lane holds an OVERLAY; your workset is a different number
254
267
 
255
268
  This is the distinction that decides whether a lane is legible to anyone but you.
@@ -20,6 +20,7 @@
20
20
  - [Stage your workset, not the whole repo](#stage-your-workset-not-the-whole-repo)
21
21
  - [Three numbers, three questions](#three-numbers-three-questions)
22
22
  - [Step 4: Verify the Change](#step-4-verify-the-change)
23
+ - [Many failures are usually few causes](#many-failures-are-usually-few-causes)
23
24
  - [Step 5: Iterate If Needed](#step-5-iterate-if-needed)
24
25
  - [Step 6: Update Documentation](#step-6-update-documentation)
25
26
  - [Temporary Experiment Workflow](#temporary-experiment-workflow)
@@ -201,9 +202,39 @@ runs resolve and what a sync writes. If `account-info.json` carries a `component
201
202
  variants of these components exist: editing an origin component will drift them, so check
202
203
  `remits-cli components branches` before changing shared code. See `branch-variants.md`.
203
204
 
204
- If the request names a journey or acceptance behavior, start the envelope here, after the target tuple is
205
- understood and before editing. A manifest can be lightweight JSON; the point is that the required
206
- evidence is durable before the proof is collected.
205
+ **You do not need to start anything to have a record.** Every stage, test, token, tool and sync appends a
206
+ world-stamped line to your actor's evidence trail automatically. Read it with:
207
+
208
+ ```bash
209
+ remits-cli evidence
210
+ ```
211
+
212
+ That is the right tool for "what have I already run?", "did that test actually execute in the prod lane?",
213
+ and "what should I put in my final message?". It is per-actor, so agents sharing a checkout never read each
214
+ other's trail.
215
+
216
+ **Start a verification envelope only when someone else needs a verdict** — a support ticket, a human who
217
+ asked you to prove specific things, or a handoff another agent will act on. It is not a routine step before
218
+ editing, and an envelope you open for your own benefit is almost always `remits-cli evidence` in disguise.
219
+
220
+ When you do want a verdict, name what must be true. One command, no manifest file:
221
+
222
+ ```bash
223
+ remits-cli verify start --summary "Hosted upload updates an existing profile"
224
+ remits-cli verify claim fees-balance --text "statement fee totals reconcile to source within five cents"
225
+ remits-cli verify claim pilot-green --text "the pilot suite passes" --test "Acquirer Pilot"
226
+ ```
227
+
228
+ Prove a claim by naming it on the command that already proves it — `--claim <id>` works on `verify
229
+ test`, `token`, `tool`, `stage`, `sync` and `attach` alike:
230
+
231
+ ```bash
232
+ remits-cli verify test --test "Acquirer Pilot" --claim pilot-green
233
+ remits-cli verify attach --claim fees-balance --note "34025 reconciles at 0.02 variance"
234
+ ```
235
+
236
+ An envelope with no claims reports `evidence_only`. That is a complete, final state — a log, not a
237
+ half-finished exam. Nothing about it is outstanding.
207
238
 
208
239
  #### Step 2: Make the Change
209
240
  Edit component files under `components/`. This is local file editing — the platform doesn't know about your changes yet.
@@ -381,14 +412,54 @@ Tests run on the platform against your staged snapshot. They stream results in r
381
412
  id + content hash, platform sync vs local HEAD. If it is not the world you meant, stop: the result will be
382
413
  about a different world.
383
414
 
415
+ The platform launch is asynchronous. By default `remits-cli test run` waits for its task to finish, and any
416
+ cases selected by `--names "a|b"` execute sequentially inside that suite task. If several slow cases are
417
+ independent, stage once, then start separate `test run --names "<case>" --wait false` commands from the same
418
+ lane and keep the task ids they print. Complete each proof with
419
+ `remits-cli test status --task-id <id>`; that terminal status read records and attaches the final `test_run`
420
+ evidence. Do not parallelize cases that mutate the same fixture, rely on shared suite setup state, or make
421
+ undeclared live AI/provider calls.
422
+
384
423
  If cases fail, read the printed summary and pivots first: each case has an `outcome`
385
424
  (`failed`, `error`, `budget_exceeded`, `provider_unavailable`, …) with its reason, timing, trace id, AI usage
386
425
  (live vs mocked, cost), bounded `report(...)` diagnostics, live HTTP signals, and resolved component
387
426
  provenance. Fix the code, re-stage, and re-run only after those pivots explain the failure.
388
427
 
428
+ ##### Many failures are usually few causes
429
+
430
+ When several cases fail, the run leads with **failure roots** — the failures grouped by their assertion with
431
+ the particulars (ids, numbers, Groovy's `Expression:`/`Values:` decoration) stripped out, largest group first,
432
+ one pivot block per root rather than per case:
433
+
434
+ ```text
435
+ Failure roots (33 failed case(s), 21 distinct root(s)):
436
+ 10x assert statement.data.processingStatus == 'Analyzed'
437
+ e.g. bundled UK statements (+9 more)
438
+ Largest root first: remits-cli test run --test "Statement Reader Calculations" --names "bundled UK statements"
439
+ ```
440
+
441
+ **Work the largest root against ONE named case, then re-run the suite.** A full re-run after every edit is the
442
+ most expensive way to learn nothing: a real account spent two days and twenty-five runs holding a suite at
443
+ 56–59 of 92 while ten of its thirty-three failures were one cause. The narrow run is seconds, tells you
444
+ whether the cause moved, and leaves your context for the actual reasoning.
445
+
446
+ Two failure roots mean "stop and look elsewhere", not "iterate harder":
447
+
448
+ - **`aiMock ctx.replay(...) found no usable stored provider response`** — the message says which of three
449
+ things happened. *No rows at all* for that sessionId means the stored session this fixture borrows is not
450
+ on this platform and will not come back; the case cannot pass until the mock builds its own response with
451
+ `ctx.toolCall(...)` / `ctx.content(...)`. Do not re-run it. (Replaying a session that IS there renews it,
452
+ so a suite that runs regularly keeps its fixtures.)
453
+ - **`Method too large` / `Class too large` / a synthetic `_closureNN`** — a JVM limit on one method body, not
454
+ a bug in the line it names. Staging now warns *before* the refusal and names the closure's source line span.
455
+ Split that body; see `development-guide.md` → *Keep Component Bodies Split*.
456
+
389
457
  Every finished run is recorded durably: `remits-cli test status --task-id <id>` answers after the live status
390
- expires, `remits-cli test runs --test <name> --compare` compares the latest run with the previous one, and
391
- `remits-cli test compare --base <id> --head <id>` compares any two. For evaluation suites (a `corpus(name)` of
458
+ expires, `remits-cli test runs --test <name> --compare` compares the latest run with the newest **comparable**
459
+ one, and `remits-cli test compare --base <id> --head <id>` compares any two. Comparable means the same data
460
+ lane, branch, workspace and case count: a run's `dataMode` is the lane it was written in, so `64/92` in prod
461
+ and `59/92` in test are two facts and not a trend. `test runs` lists one lane at a time and names it
462
+ (`--all-lanes` to see both). For evaluation suites (a `corpus(name)` of
392
463
  cases seeded with `remits-cli corpus import`, one case per corpus case, intentional live AI inside
393
464
  `withAiBudget(...)`, measurements compared with `remits-cli corpus compare` / `corpus consistency`), read
394
465
  `guides/test-components.md` → *Evaluation Suites And Corpora*.
@@ -76,6 +76,9 @@ remits-cli tool --account-id 21 --as-account 37 --target-account 37 --name mcp_r
76
76
  # poll by run id: --account-id 21 --as-account 37 --target-account 37 --input '{"controlAction":"status","actionRunId":"my-stable-run-id"}'
77
77
  ```
78
78
 
79
+ That is the CLI Action-run surface. The command is still `remits-cli tool`; this CLI does not have a
80
+ separate top-level `remits-cli action`, `remits-cli actions`, or `remits-cli run action` wrapper.
81
+
79
82
  For long tools that lack their own async mode, use the CLI transport async (`--async true`), optionally with
80
83
  `--wait true` to poll locally, and `remits-cli tool status --call-id <callId>`. Do not stack both mechanisms
81
84
  (see `command-reference.md` → *Tool Execution Lifecycle*). Use `--timeout-ms <ms>` only to adjust the per-request client timeout; it is
@@ -403,12 +406,18 @@ Front-stage references:
403
406
  | `page` | no | 1-based page number. Default: `1` |
404
407
  | `pageSize` | no | Results per page. Default: `25`, max: `100` |
405
408
  | `summaryOnly` | no | Detail mode: return the MAP (record index + stats + timeline) with no payloads. Same as `parts:["index"]` |
406
- | `parts` | no | Detail parts: `index`, `stats`, `transcript`, `tool_calls`, and record sections `full`, `conversation_messages`, `system`, `response`, `tools` |
409
+ | `compact` | no | Detail mode: after reading the MAP, open selected records without repeating the `records` index and `timeline`; keeps `groupingSummary`, `representativeSession`, `stats`, selected ids, and requested payload sections |
410
+ | `payloadOnly` | no | Alias for `compact` |
411
+ | `parts` | no | Detail parts: `index`, `map`, `records`, `stats`, `transcript`, `tool_calls`, and record sections `full`, `conversation_messages`, `system`, `response`, `tools` |
407
412
  | `sections` | no | Alias for `parts` |
408
413
  | `recordIds` | no | Detail mode: open only these request/response record IDs |
409
414
  | `toolCallIds` | no | Detail mode: return the FULL exact input/result from `ai_tool_call` for these tool-call ids (what the tool PRODUCED — see the lens caveat) |
410
415
  | `consolidateContext` | no | When `true`, collapses repeated XML-like prompt context into a consolidated section |
411
416
 
417
+ Detail responses expose both lenses: `groupingSummary` is the authoritative grouping-wide summary
418
+ (counts, lanes, cost, status), while `representativeSession` names the concrete session used for
419
+ owner/runtime metadata. `session` remains only as a compatibility alias for older callers.
420
+
412
421
  Spend: each search row carries `liveRequestCount`, `mockedRequestCount`, `lanes`, `testTaskId` and
413
422
  `estimatedCost`/`estimatedCostUsd` = **live spend only** (a mocked turn is never priced, even when it replays a
414
423
  recording with a cost). `pageTotals` sums the page. Turns recorded before the platform stored the mock flag are
@@ -416,12 +425,16 @@ recording with a cost). `pageTotals` sums the page. Turns recorded before the pl
416
425
  is refused rather than silently widened; the applied bounds are echoed as `filters.rangeStart`/`rangeEnd`.
417
426
  To audit one Test run: `{"action":"search","testTaskId":"<taskId>","pageSize":100}`.
418
427
 
419
- Audit flow: `action:"search"` to find the grouping → `action:"detail"` + `summaryOnly:true` for the MAP → re-call detail with `recordIds`/`toolCallIds` + `parts` to open exactly what you need. Prefer the map → open flow over a full-detail dump. Remember the two-lens rule: `tool_calls`/`toolCallIds` is what the tool PRODUCED; `conversation_messages`/`transcript` is what the AI CONSUMED (after any `_offload`/`_hideResult`/`_message`/supersede/evict transform).
428
+ Audit flow: `action:"search"` to find the grouping → `action:"detail"` + `summaryOnly:true` for the MAP → re-call detail with `recordIds`/`toolCallIds` + `parts` to open exactly what you need, usually with `compact:true` once the map has chosen the row. Prefer the map → open flow over a full-detail dump. Remember the two-lens rule: `tool_calls`/`toolCallIds` is what the tool PRODUCED; `conversation_messages`/`transcript` is what the AI CONSUMED (after any `_offload`/`_hideResult`/`_message`/supersede/evict transform).
420
429
 
421
430
  ### `mcp_run_action`
422
431
  Run an Action on a target account, with explicit prod/test data mode, optional staged branch resolution, and
423
432
  staged-vs-DB provenance in the result.
424
433
 
434
+ Invoke it with `remits-cli tool --name mcp_run_action`. Despite the natural shorthand "run an Action", there
435
+ is no separate top-level `remits-cli action` / `actions` command and no `remits-cli run action` wrapper in
436
+ this CLI build.
437
+
425
438
  Describe the Action first when the input shape is not obvious. This does not execute the Action:
426
439
 
427
440
  ```bash
@@ -832,6 +845,10 @@ Common uses:
832
845
  - `remits-cli components branch <name> --diff <id> --component-type <kind>` — compare one variant against
833
846
  current trunk.
834
847
  - `remits-cli components branch <name> --subscribers` — list accounts resolving that branch.
848
+ - `remits-cli components branch <name> --copy-to <newBranch> [--dry-run] [--force]` — seed a new variant
849
+ branch with the source branch's stored overlays before the first safe sync of the new branch. Dry-run
850
+ reports the copy plan without writes; force is required when the target already has overlays or live
851
+ subscribers.
835
852
 
836
853
  ### `mcp_cache`
837
854
  Bounded read-only investigation of the platform Redis keyspace — the way to see exactly what a staged