@remits/remits-cli 0.1.135 → 0.1.137
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/index.js +931 -112
- package/package.json +1 -1
- package/skills/remits-cli/SKILL.md +16 -7
- package/skills/remits-cli/references/branch-variants.md +25 -0
- package/skills/remits-cli/references/command-reference.md +105 -11
- package/skills/remits-cli/references/component-resolution.md +13 -0
- package/skills/remits-cli/references/development-loop.md +76 -5
- package/skills/remits-cli/references/tool-reference.md +19 -2
package/package.json
CHANGED
|
@@ -48,6 +48,7 @@ reference below, load it before you act, not after something surprises you.
|
|
|
48
48
|
| register this session as an agent, or ask why a routed ticket never started | `agent-sessions.md` | `serve` vs `register`, what registration does, worker spawning per agent kind, routing order, the control center, multi-session auth |
|
|
49
49
|
| investigate live behavior | `investigation.md` | which tool reads which record, correlation keys, reading a record's `content`, HTTP audits, AI activity, node/`localMode`, the production support flows |
|
|
50
50
|
| call any `mcp_*` tool | `tool-reference.md` | every tool's parameters, semantics, and traps — read the tool's entry before building its input |
|
|
51
|
+
| run, describe, poll, or interrupt an Action from the CLI | `tool-reference.md` → `mcp_run_action` | Action execution through `remits-cli tool --name mcp_run_action`; there is no separate top-level `remits-cli action` / `run action` wrapper |
|
|
51
52
|
| reason about repo discovery, auth state, or what a command actually sent | `cli-state.md` | the global `~/.remits-cli/` control plane vs per-repo `./.remits-cli/`, and which file answers which question |
|
|
52
53
|
| need an exact flag, or the auth / host / async surface | `command-reference.md` | authentication, host vs data mode, the two async mechanisms, hierarchy-scoped reads, the full command list, prod banners |
|
|
53
54
|
| give up on something that misbehaved | `troubleshooting.md` | the symptom→fix table, the two kinds of escalation, and the escalation bundle |
|
|
@@ -162,11 +163,18 @@ reference named after it.
|
|
|
162
163
|
component (preferred, because it becomes regression protection) or a browser flow through
|
|
163
164
|
`remits-cli token`. "Just do it" and "that's fine, commit it" are not evidence. If you genuinely
|
|
164
165
|
cannot verify, say what you would need and ask. (`development-loop.md`)
|
|
165
|
-
-
|
|
166
|
-
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
|
|
166
|
+
- **`remits-cli evidence` answers "what have I actually run, and in which world?"** Every stage, test,
|
|
167
|
+
token, tool and sync appends one world-stamped line automatically, per actor. Nothing to start,
|
|
168
|
+
nothing to satisfy. Read it before you re-run something, and quote it when you report what you
|
|
169
|
+
proved. (`development-loop.md`)
|
|
170
|
+
- **Verification envelopes are OPTIONAL and exist for one job: a verdict someone else will rely on.**
|
|
171
|
+
Start one when a ticket, a human, or a handoff needs a checkable "these specific things are true" —
|
|
172
|
+
not as a routine step before editing. `remits-cli verify start --summary "..."` then
|
|
173
|
+
`remits-cli verify claim <id> --text "..." [--test "<suite>"]` names what must be true; add
|
|
174
|
+
`--claim <id>` to the command that proves it; `remits-cli verify report` gives the verdict. An
|
|
175
|
+
envelope with no claims is simply an evidence log, which is a complete state — not something to
|
|
176
|
+
chase. If you find yourself opening an envelope to answer a question about your own work, use
|
|
177
|
+
`remits-cli evidence` instead. (`development-loop.md`, `component-resolution.md`)
|
|
170
178
|
- **Read the component's `.meta.yml` before changing behavior.** Sidecar descriptions can be dated
|
|
171
179
|
decision records. Before changing a displayed value, helper, calculation, schema field, or prompt
|
|
172
180
|
contract, check the sidecar and either preserve its decision or explicitly supersede it.
|
|
@@ -209,8 +217,9 @@ remits-cli tools # which tools this account actually has (tools
|
|
|
209
217
|
```
|
|
210
218
|
|
|
211
219
|
plus the repo's `account-info.json` → `resolution` block for the account's shape.
|
|
212
|
-
For a workflow-shaped request,
|
|
213
|
-
|
|
220
|
+
For a workflow-shaped request, rely on the automatic `remits-cli evidence` trail unless someone else needs
|
|
221
|
+
a checkable verdict. Only then start `remits-cli verify start --summary "..."`, declare claims, and keep
|
|
222
|
+
that envelope active through stage/test/token/sync.
|
|
214
223
|
|
|
215
224
|
If any command returns 401, run `remits-cli auth` (with the same `--base-url` if you were targeting a
|
|
216
225
|
non-default host).
|
|
@@ -408,6 +408,31 @@ from that checkout). Its first real sync (or commit) needs `--create-variant-bra
|
|
|
408
408
|
variants or a subscriber the platform treats it as a feature branch and refuses to land it, so creating a
|
|
409
409
|
new variant branch is always a stated decision, never a side effect of a feature branch's name.
|
|
410
410
|
|
|
411
|
+
For a nested branch cut from an existing variant branch, seed the platform overlays explicitly after the git
|
|
412
|
+
branch is created and pushed:
|
|
413
|
+
|
|
414
|
+
```bash
|
|
415
|
+
remits-cli components branch forked --copy-to sandbox --dry-run
|
|
416
|
+
remits-cli components branch forked --copy-to sandbox
|
|
417
|
+
remits-cli components sync --safe --branch sandbox
|
|
418
|
+
```
|
|
419
|
+
|
|
420
|
+
The copy is a DB overlay bootstrap, not a git write. It makes inherited overlays from `forked` already
|
|
421
|
+
`storedCurrent` on the first `sandbox` sync, so the safe diff gate only has to account for the new branch's
|
|
422
|
+
real file changes.
|
|
423
|
+
|
|
424
|
+
**Copy before you subscribe, or the copy needs `--force`.** The copy refuses a target that already has
|
|
425
|
+
stored overlays *or* live subscribers, and that second guard fires even when the target has no overlays at
|
|
426
|
+
all — a branch somebody is already resolving is code those accounts are running right now, and replacing it
|
|
427
|
+
has to be a stated decision. Seeding first and subscribing after keeps the plain form working. The copy is
|
|
428
|
+
also all-or-nothing: the delete of any replaced overlays and every inserted row share one transaction, so a
|
|
429
|
+
failure leaves the target exactly as it was rather than half-populated.
|
|
430
|
+
|
|
431
|
+
Run it as `--dry-run` first. The plan reports `plannedCopies` — the number the real run will write — which
|
|
432
|
+
is not always the source branch's row count: a source overlay that has become identical to trunk in both
|
|
433
|
+
content and metadata is sparse, is not stored, and is listed as skipped instead of being silently dropped
|
|
434
|
+
from the total.
|
|
435
|
+
|
|
411
436
|
**Trunk moving also invalidates a branch.** Variant sparseness compares branch content against *current*
|
|
412
437
|
trunk, so a trunk change can make an overlay obsolete without the branch changing at all. A trunk sync
|
|
413
438
|
drops the affected branches' cached sync verdicts, so the next `components sync` on the branch really
|
|
@@ -15,6 +15,8 @@
|
|
|
15
15
|
- [Hierarchy-scoped tool reads](#hierarchy-scoped-tool-reads)
|
|
16
16
|
- [Data Mode](#data-mode)
|
|
17
17
|
- [Command Reference](#command-reference)
|
|
18
|
+
- [Reading a failing run](#reading-a-failing-run)
|
|
19
|
+
- [A pass count only means something within one data lane](#a-pass-count-only-means-something-within-one-data-lane)
|
|
18
20
|
- [Verification envelopes](#verification-envelopes)
|
|
19
21
|
- [Staging modes: workset vs full snapshot](#staging-modes-workset-vs-full-snapshot)
|
|
20
22
|
- [Prod banners and retryable failures](#prod-banners-and-retryable-failures)
|
|
@@ -236,12 +238,13 @@ remits-cli components branches [--json] # branche
|
|
|
236
238
|
remits-cli components branch <name> [--json] # one branch: owner account, overridden / added / removed, drift flags, subscribers
|
|
237
239
|
remits-cli components branch <name> --diff <componentId> --component-type <kind> [--json]
|
|
238
240
|
remits-cli components branch <name> --subscribers [--json]
|
|
241
|
+
remits-cli components branch <name> --copy-to <newBranch> [--dry-run] [--force] [--json] # seed a new variant branch with this branch's stored overlays; force if target has overlays/subscribers
|
|
239
242
|
remits-cli components branch <name> --subscribe <accountId> [--parent-account <id>] [--domain <host>] [--dry-run] [--confirm-primary-edge] # make an account resolve this branch
|
|
240
243
|
remits-cli components branch <name> --unsubscribe <accountId> # return that account to trunk
|
|
241
244
|
remits-cli components branch <name> --retire [--force] # delete the branch's overlays
|
|
242
|
-
remits-cli test run --test <id|name> [--branch <stagingScope>] [--names "a|b"] [--watch true|false] [--data-mode test|prod] [--as-account <ID>] [--variant-branch <name|none>] [--json]
|
|
245
|
+
remits-cli test run --test <id|name> [--branch <stagingScope>] [--names "a|b"] [--watch true|false] [--wait true|false] [--data-mode test|prod] [--as-account <ID>] [--variant-branch <name|none>] [--json]
|
|
243
246
|
remits-cli test status --task-id <taskId> [--branch <stagingScope>] [--data-mode test|prod] [--json] # falls back to the DURABLE run record once the live status expires
|
|
244
|
-
remits-cli test runs [--test <id|name>] [--limit 20] [--compare] [--as-account <ID>] [--json]
|
|
247
|
+
remits-cli test runs [--test <id|name>] [--limit 20] [--compare] [--all-lanes] [--as-account <ID>] [--json] # durable run history: pass counts, live AI cost, lane, content hash
|
|
245
248
|
remits-cli test compare --base <taskId> --head <taskId> [--json] # per-case improved / regressed / changed, cost and world deltas
|
|
246
249
|
remits-cli corpus import --manifest corpus-manifest.json [--corpus <name>] [--as-account <ID>] [--data-mode test|prod --confirm-prod] [--json] # seed an evaluation corpus: cases + immutable artifacts; idempotent by caseKey
|
|
247
250
|
remits-cli corpus cases --corpus <name> [--split S] [--tag T|--tags T,U] [--key K|--keys K,L] [--include-values] [--include-retired] [--limit N] [--json]
|
|
@@ -252,6 +255,54 @@ remits-cli corpus consistency --corpus <name> --case <caseKey> [--json]
|
|
|
252
255
|
remits-cli corpus retire --corpus <name> --case <caseKey> [--case ...] [--restore] [--json] # drop a case from future runs; past measurements stay readable
|
|
253
256
|
```
|
|
254
257
|
|
|
258
|
+
`test run` starts a server-side task immediately. By default the CLI process polls until the task is
|
|
259
|
+
terminal; `--watch false` disables websocket progress streaming but still waits. Pass `--wait false` to
|
|
260
|
+
return after launch with the task id. A single run with `--names "a|b|c"` selects cases into one suite task,
|
|
261
|
+
and those cases execute sequentially in declaration order. To overlap independent slow cases, launch
|
|
262
|
+
separate `remits-cli test run --names "<case>" --wait false` commands from the same staged lane and keep the
|
|
263
|
+
printed task ids; complete each proof with `test status --task-id <id>`. Avoid parallel runs for cases that
|
|
264
|
+
share mutable fixtures, suite-level side effects, or undeclared live AI/provider calls.
|
|
265
|
+
|
|
266
|
+
### Reading a failing run
|
|
267
|
+
|
|
268
|
+
`test run` and `test status` print a **verdict**, not the run. The whole run object used to be dumped as
|
|
269
|
+
pretty JSON in human mode - 220 KB for one 92-case suite, most of it identifiers repeated once per case - so
|
|
270
|
+
the command an agent runs most often was the one that spent its remaining room to think. Every byte is still
|
|
271
|
+
there behind `--json`, and behind `test status --task-id <id> --json` once the live status expires.
|
|
272
|
+
|
|
273
|
+
What you get instead, and what to do with it:
|
|
274
|
+
|
|
275
|
+
```text
|
|
276
|
+
Cases: 59 passed, 33 failed, 92 total
|
|
277
|
+
|
|
278
|
+
Failure roots (33 failed case(s), 21 distinct root(s)):
|
|
279
|
+
10x assert statement.data.processingStatus == 'Analyzed'
|
|
280
|
+
e.g. bundled UK statements (+9 more)
|
|
281
|
+
2x assert statement.data.feeBreakdownChecked == true
|
|
282
|
+
e.g. 7851 Flat Rate (+1 more)
|
|
283
|
+
Largest root first: remits-cli test run --test "Statement Reader Calculations" --names "bundled UK statements"
|
|
284
|
+
```
|
|
285
|
+
|
|
286
|
+
**Thirty-three failures are not thirty-three problems.** Cases are grouped by their assertion root - the
|
|
287
|
+
failure with the particulars (ids, numbers, Groovy's `Expression:`/`Values:` decoration) removed - largest
|
|
288
|
+
group first, and one pivot block is printed per root rather than per case. Fix the largest root against one
|
|
289
|
+
named case, then re-run the whole suite. Re-running the suite before you have a root is how a session spends
|
|
290
|
+
an afternoon at the same pass count.
|
|
291
|
+
|
|
292
|
+
### A pass count only means something within one data lane
|
|
293
|
+
|
|
294
|
+
`test runs` is scoped to the data lane of the command by default, and says so:
|
|
295
|
+
|
|
296
|
+
```text
|
|
297
|
+
Data lane: test (6 more run(s) in the other lane; --all-lanes to include)
|
|
298
|
+
```
|
|
299
|
+
|
|
300
|
+
A run's `dataMode` **is** the lane it was written in, so `64/92` in the prod lane and `59/92` in the test lane
|
|
301
|
+
are two different facts, not a regression. `--compare` (and `test runs --compare`) picks the newest two runs
|
|
302
|
+
of the same **shape** - same data lane, branch, workspace and case count - and says how many newer runs it
|
|
303
|
+
skipped; it refuses rather than comparing across worlds. `test status --data-mode prod` on a run recorded in
|
|
304
|
+
the test lane returns the run's real lane and now says that is what happened.
|
|
305
|
+
|
|
255
306
|
A run labelled **`AI MOCKED`** replayed every AI turn from `aiMock`: it produces the same outcomes, scores and
|
|
256
307
|
metrics a measured run would, so read that label before treating green as evidence. `compare` warns when base and
|
|
257
308
|
head used different AI modes; `consistency` warns before calling a case unstable when its runs mixed modes.
|
|
@@ -344,10 +395,24 @@ shown is the latest run's, with the distinct worlds listed beneath it.
|
|
|
344
395
|
|
|
345
396
|
### Verification envelopes
|
|
346
397
|
|
|
347
|
-
|
|
348
|
-
|
|
349
|
-
|
|
350
|
-
|
|
398
|
+
**First, the thing you probably want instead.** `remits-cli evidence` prints what you have actually run
|
|
399
|
+
and in which world, grouped by world, with no envelope and nothing to start:
|
|
400
|
+
|
|
401
|
+
```bash
|
|
402
|
+
remits-cli evidence # this actor's trail
|
|
403
|
+
remits-cli evidence --json # structured, for your final report
|
|
404
|
+
```
|
|
405
|
+
|
|
406
|
+
Every stage, test, token, tool and sync appends to it automatically. Use it for "what have I already
|
|
407
|
+
run?", "did that execute in the prod lane?" and "what do I put in my final message?".
|
|
408
|
+
|
|
409
|
+
**Verification envelopes are optional, and exist for a verdict someone ELSE will rely on** — a support
|
|
410
|
+
ticket, a human who asked you to prove specific things, a handoff another agent will act on. They are not
|
|
411
|
+
a routine step before editing. An envelope you started for your own benefit is nearly always
|
|
412
|
+
`remits-cli evidence` in disguise, and it costs you turns that prove nothing.
|
|
413
|
+
|
|
414
|
+
When a verdict is genuinely wanted, start it after you know the target
|
|
415
|
+
account/branch/workspace/data-mode tuple:
|
|
351
416
|
|
|
352
417
|
```bash
|
|
353
418
|
remits-cli workspace use --auto
|
|
@@ -355,6 +420,35 @@ remits-cli components status
|
|
|
355
420
|
remits-cli verify start --summary "Hosted upload updates an existing profile" --manifest acceptance.json
|
|
356
421
|
```
|
|
357
422
|
|
|
423
|
+
**Name what must be true, or there is no verdict to reach.** An envelope with no `claims` and no
|
|
424
|
+
`requiredEvidence` reports `evidence_only` — a world-stamped log. That is a complete, final state, not a
|
|
425
|
+
partial one, and nothing about it is outstanding. When you DO want a verdict, a manifest file is one way to
|
|
426
|
+
declare the contract; `verify claim` is the one-command way, and it works on a live envelope:
|
|
427
|
+
|
|
428
|
+
```bash
|
|
429
|
+
remits-cli verify claim market-filter --text "US market scoring excludes AU and GB statements"
|
|
430
|
+
remits-cli verify claim pilot-green --text "the pilot suite passes" --test "Acquirer Pilot"
|
|
431
|
+
```
|
|
432
|
+
|
|
433
|
+
Prove a claim by naming it on the command that already proves it. `--claim <id>` is stamped onto the
|
|
434
|
+
evidence packet every command sends, so it works on `verify test`, `token`, `tool`, `stage`, `sync` and
|
|
435
|
+
`attach` alike, and takes a comma-separated list:
|
|
436
|
+
|
|
437
|
+
```bash
|
|
438
|
+
remits-cli verify test --test "Acquirer Pilot" --claim pilot-green
|
|
439
|
+
remits-cli verify tool --name mcp_run_action --as-account 36 --claim market-filter
|
|
440
|
+
remits-cli verify attach --claim market-filter --note "AU statements excluded and counted"
|
|
441
|
+
```
|
|
442
|
+
|
|
443
|
+
A claim with no shape is satisfied by any successful packet tagged with its id. A claim that names a
|
|
444
|
+
`--test` suite (or `--packet-type`) must be proven by that shape. Re-declaring an id replaces it, so
|
|
445
|
+
tightening a claim is also one command.
|
|
446
|
+
|
|
447
|
+
**When an envelope with a contract will not close, read the report — do not start another one.** Several
|
|
448
|
+
open envelopes for the same goal, or workspaces named `...-r7`/`-r8`/`-r9`, is the loop signature, and
|
|
449
|
+
`activity inspect` reports both as smells. Close what you are not finishing: `verify supersede --envelope
|
|
450
|
+
<old> --superseded-by <new>` or `verify abandon --envelope <old> --reason "false start"`.
|
|
451
|
+
|
|
358
452
|
An active envelope is stored per command world (`baseUrl + accountId + git branch + workspace`)
|
|
359
453
|
inside the active local actor's mirror; **data mode is deliberately not part of the active pointer key**:
|
|
360
454
|
`./.remits-cli/actors/<local-agent>/verification/active-contexts.json`. The legacy flat
|
|
@@ -367,7 +461,7 @@ history:
|
|
|
367
461
|
|
|
368
462
|
```bash
|
|
369
463
|
remits-cli verify stage --workset
|
|
370
|
-
remits-cli verify test --test "Adyen Import Recovery" --names "browser upload recovery"
|
|
464
|
+
remits-cli verify test --test "Adyen Import Recovery" --names "browser upload recovery" --claim recovery
|
|
371
465
|
remits-cli verify token --path /page/pricing-config --as-account 21 --data-mode test
|
|
372
466
|
remits-cli verify sync --safe
|
|
373
467
|
remits-cli verify report
|
|
@@ -435,10 +529,10 @@ proof noun: `passed` for a Test case, `measured` for corpus comparisons, `synced
|
|
|
435
529
|
|
|
436
530
|
Final claims should come from `remits-cli verify report`. Treat `Verified` as the acceptance boundary.
|
|
437
531
|
`Additional evidence` is useful handoff context, but it does not satisfy a missing required packet unless
|
|
438
|
-
the report lists it under `Verified`. If the report says `
|
|
439
|
-
say that plainly instead of widening the claim. In
|
|
440
|
-
|
|
441
|
-
action.
|
|
532
|
+
the report lists it under `Verified`. If the report says `evidence_only`, no verdict was requested; if it
|
|
533
|
+
says `partially_verified`, stale, or missing evidence, say that plainly instead of widening the claim. In
|
|
534
|
+
particular, staged proof is not committed variant/trunk proof, a token is not browser proof, and a direct
|
|
535
|
+
DOM or Alpine state mutation is not the same as a user action.
|
|
442
536
|
|
|
443
537
|
Evidence from the wrong world is excluded before it can satisfy a requirement. When the manifest declares
|
|
444
538
|
fields such as `repoAccountId`, `gitBranch`, `componentBranch`, `workspace`, or `dataMode`, the report names
|
|
@@ -16,6 +16,7 @@
|
|
|
16
16
|
- [When staged overrides apply](#when-staged-overrides-apply)
|
|
17
17
|
- [Diagnosing which version is in play](#diagnosing-which-version-is-in-play)
|
|
18
18
|
- [Working alongside other agents: the staging WORKSPACE](#working-alongside-other-agents-the-staging-workspace)
|
|
19
|
+
- [Keeping your remits-cli current](#keeping-your-remits-cli-current)
|
|
19
20
|
- [A lane holds an OVERLAY; your workset is a different number](#a-lane-holds-an-overlay-your-workset-is-a-different-number)
|
|
20
21
|
- [Stage / sync / clear with remits-cli](#stage--sync--clear-with-remits-cli)
|
|
21
22
|
- [Stale after sync / commit (the in-memory compile cache)](#stale-after-sync--commit-the-in-memory-compile-cache)
|
|
@@ -250,6 +251,18 @@ index with `remits-cli workstream status`.
|
|
|
250
251
|
`null` means unknown — a lane staged by an older CLI, or a branch whose last sync is not recorded —
|
|
251
252
|
never "fresh".
|
|
252
253
|
|
|
254
|
+
### Keeping your remits-cli current
|
|
255
|
+
|
|
256
|
+
The diagnostics in these references only exist in the CLI that prints them. An older install does not print a
|
|
257
|
+
worse version of them — it prints nothing, with no error, so nothing tells you what you are not being shown.
|
|
258
|
+
`components stage` and `test run` warn when yours is behind:
|
|
259
|
+
|
|
260
|
+
```text
|
|
261
|
+
[remits-cli 0.1.120 -> 0.1.136] ... npm install -g @remits/remits-cli@latest
|
|
262
|
+
```
|
|
263
|
+
|
|
264
|
+
Upgrade when you see it. Auto-update handles it for you on most commands, including `--json` ones.
|
|
265
|
+
|
|
253
266
|
### A lane holds an OVERLAY; your workset is a different number
|
|
254
267
|
|
|
255
268
|
This is the distinction that decides whether a lane is legible to anyone but you.
|
|
@@ -20,6 +20,7 @@
|
|
|
20
20
|
- [Stage your workset, not the whole repo](#stage-your-workset-not-the-whole-repo)
|
|
21
21
|
- [Three numbers, three questions](#three-numbers-three-questions)
|
|
22
22
|
- [Step 4: Verify the Change](#step-4-verify-the-change)
|
|
23
|
+
- [Many failures are usually few causes](#many-failures-are-usually-few-causes)
|
|
23
24
|
- [Step 5: Iterate If Needed](#step-5-iterate-if-needed)
|
|
24
25
|
- [Step 6: Update Documentation](#step-6-update-documentation)
|
|
25
26
|
- [Temporary Experiment Workflow](#temporary-experiment-workflow)
|
|
@@ -201,9 +202,39 @@ runs resolve and what a sync writes. If `account-info.json` carries a `component
|
|
|
201
202
|
variants of these components exist: editing an origin component will drift them, so check
|
|
202
203
|
`remits-cli components branches` before changing shared code. See `branch-variants.md`.
|
|
203
204
|
|
|
204
|
-
|
|
205
|
-
|
|
206
|
-
|
|
205
|
+
**You do not need to start anything to have a record.** Every stage, test, token, tool and sync appends a
|
|
206
|
+
world-stamped line to your actor's evidence trail automatically. Read it with:
|
|
207
|
+
|
|
208
|
+
```bash
|
|
209
|
+
remits-cli evidence
|
|
210
|
+
```
|
|
211
|
+
|
|
212
|
+
That is the right tool for "what have I already run?", "did that test actually execute in the prod lane?",
|
|
213
|
+
and "what should I put in my final message?". It is per-actor, so agents sharing a checkout never read each
|
|
214
|
+
other's trail.
|
|
215
|
+
|
|
216
|
+
**Start a verification envelope only when someone else needs a verdict** — a support ticket, a human who
|
|
217
|
+
asked you to prove specific things, or a handoff another agent will act on. It is not a routine step before
|
|
218
|
+
editing, and an envelope you open for your own benefit is almost always `remits-cli evidence` in disguise.
|
|
219
|
+
|
|
220
|
+
When you do want a verdict, name what must be true. One command, no manifest file:
|
|
221
|
+
|
|
222
|
+
```bash
|
|
223
|
+
remits-cli verify start --summary "Hosted upload updates an existing profile"
|
|
224
|
+
remits-cli verify claim fees-balance --text "statement fee totals reconcile to source within five cents"
|
|
225
|
+
remits-cli verify claim pilot-green --text "the pilot suite passes" --test "Acquirer Pilot"
|
|
226
|
+
```
|
|
227
|
+
|
|
228
|
+
Prove a claim by naming it on the command that already proves it — `--claim <id>` works on `verify
|
|
229
|
+
test`, `token`, `tool`, `stage`, `sync` and `attach` alike:
|
|
230
|
+
|
|
231
|
+
```bash
|
|
232
|
+
remits-cli verify test --test "Acquirer Pilot" --claim pilot-green
|
|
233
|
+
remits-cli verify attach --claim fees-balance --note "34025 reconciles at 0.02 variance"
|
|
234
|
+
```
|
|
235
|
+
|
|
236
|
+
An envelope with no claims reports `evidence_only`. That is a complete, final state — a log, not a
|
|
237
|
+
half-finished exam. Nothing about it is outstanding.
|
|
207
238
|
|
|
208
239
|
#### Step 2: Make the Change
|
|
209
240
|
Edit component files under `components/`. This is local file editing — the platform doesn't know about your changes yet.
|
|
@@ -381,14 +412,54 @@ Tests run on the platform against your staged snapshot. They stream results in r
|
|
|
381
412
|
id + content hash, platform sync vs local HEAD. If it is not the world you meant, stop: the result will be
|
|
382
413
|
about a different world.
|
|
383
414
|
|
|
415
|
+
The platform launch is asynchronous. By default `remits-cli test run` waits for its task to finish, and any
|
|
416
|
+
cases selected by `--names "a|b"` execute sequentially inside that suite task. If several slow cases are
|
|
417
|
+
independent, stage once, then start separate `test run --names "<case>" --wait false` commands from the same
|
|
418
|
+
lane and keep the task ids they print. Complete each proof with
|
|
419
|
+
`remits-cli test status --task-id <id>`; that terminal status read records and attaches the final `test_run`
|
|
420
|
+
evidence. Do not parallelize cases that mutate the same fixture, rely on shared suite setup state, or make
|
|
421
|
+
undeclared live AI/provider calls.
|
|
422
|
+
|
|
384
423
|
If cases fail, read the printed summary and pivots first: each case has an `outcome`
|
|
385
424
|
(`failed`, `error`, `budget_exceeded`, `provider_unavailable`, …) with its reason, timing, trace id, AI usage
|
|
386
425
|
(live vs mocked, cost), bounded `report(...)` diagnostics, live HTTP signals, and resolved component
|
|
387
426
|
provenance. Fix the code, re-stage, and re-run only after those pivots explain the failure.
|
|
388
427
|
|
|
428
|
+
##### Many failures are usually few causes
|
|
429
|
+
|
|
430
|
+
When several cases fail, the run leads with **failure roots** — the failures grouped by their assertion with
|
|
431
|
+
the particulars (ids, numbers, Groovy's `Expression:`/`Values:` decoration) stripped out, largest group first,
|
|
432
|
+
one pivot block per root rather than per case:
|
|
433
|
+
|
|
434
|
+
```text
|
|
435
|
+
Failure roots (33 failed case(s), 21 distinct root(s)):
|
|
436
|
+
10x assert statement.data.processingStatus == 'Analyzed'
|
|
437
|
+
e.g. bundled UK statements (+9 more)
|
|
438
|
+
Largest root first: remits-cli test run --test "Statement Reader Calculations" --names "bundled UK statements"
|
|
439
|
+
```
|
|
440
|
+
|
|
441
|
+
**Work the largest root against ONE named case, then re-run the suite.** A full re-run after every edit is the
|
|
442
|
+
most expensive way to learn nothing: a real account spent two days and twenty-five runs holding a suite at
|
|
443
|
+
56–59 of 92 while ten of its thirty-three failures were one cause. The narrow run is seconds, tells you
|
|
444
|
+
whether the cause moved, and leaves your context for the actual reasoning.
|
|
445
|
+
|
|
446
|
+
Two failure roots mean "stop and look elsewhere", not "iterate harder":
|
|
447
|
+
|
|
448
|
+
- **`aiMock ctx.replay(...) found no usable stored provider response`** — the message says which of three
|
|
449
|
+
things happened. *No rows at all* for that sessionId means the stored session this fixture borrows is not
|
|
450
|
+
on this platform and will not come back; the case cannot pass until the mock builds its own response with
|
|
451
|
+
`ctx.toolCall(...)` / `ctx.content(...)`. Do not re-run it. (Replaying a session that IS there renews it,
|
|
452
|
+
so a suite that runs regularly keeps its fixtures.)
|
|
453
|
+
- **`Method too large` / `Class too large` / a synthetic `_closureNN`** — a JVM limit on one method body, not
|
|
454
|
+
a bug in the line it names. Staging now warns *before* the refusal and names the closure's source line span.
|
|
455
|
+
Split that body; see `development-guide.md` → *Keep Component Bodies Split*.
|
|
456
|
+
|
|
389
457
|
Every finished run is recorded durably: `remits-cli test status --task-id <id>` answers after the live status
|
|
390
|
-
expires, `remits-cli test runs --test <name> --compare` compares the latest run with the
|
|
391
|
-
`remits-cli test compare --base <id> --head <id>` compares any two.
|
|
458
|
+
expires, `remits-cli test runs --test <name> --compare` compares the latest run with the newest **comparable**
|
|
459
|
+
one, and `remits-cli test compare --base <id> --head <id>` compares any two. Comparable means the same data
|
|
460
|
+
lane, branch, workspace and case count: a run's `dataMode` is the lane it was written in, so `64/92` in prod
|
|
461
|
+
and `59/92` in test are two facts and not a trend. `test runs` lists one lane at a time and names it
|
|
462
|
+
(`--all-lanes` to see both). For evaluation suites (a `corpus(name)` of
|
|
392
463
|
cases seeded with `remits-cli corpus import`, one case per corpus case, intentional live AI inside
|
|
393
464
|
`withAiBudget(...)`, measurements compared with `remits-cli corpus compare` / `corpus consistency`), read
|
|
394
465
|
`guides/test-components.md` → *Evaluation Suites And Corpora*.
|
|
@@ -76,6 +76,9 @@ remits-cli tool --account-id 21 --as-account 37 --target-account 37 --name mcp_r
|
|
|
76
76
|
# poll by run id: --account-id 21 --as-account 37 --target-account 37 --input '{"controlAction":"status","actionRunId":"my-stable-run-id"}'
|
|
77
77
|
```
|
|
78
78
|
|
|
79
|
+
That is the CLI Action-run surface. The command is still `remits-cli tool`; this CLI does not have a
|
|
80
|
+
separate top-level `remits-cli action`, `remits-cli actions`, or `remits-cli run action` wrapper.
|
|
81
|
+
|
|
79
82
|
For long tools that lack their own async mode, use the CLI transport async (`--async true`), optionally with
|
|
80
83
|
`--wait true` to poll locally, and `remits-cli tool status --call-id <callId>`. Do not stack both mechanisms
|
|
81
84
|
(see `command-reference.md` → *Tool Execution Lifecycle*). Use `--timeout-ms <ms>` only to adjust the per-request client timeout; it is
|
|
@@ -403,12 +406,18 @@ Front-stage references:
|
|
|
403
406
|
| `page` | no | 1-based page number. Default: `1` |
|
|
404
407
|
| `pageSize` | no | Results per page. Default: `25`, max: `100` |
|
|
405
408
|
| `summaryOnly` | no | Detail mode: return the MAP (record index + stats + timeline) with no payloads. Same as `parts:["index"]` |
|
|
406
|
-
| `
|
|
409
|
+
| `compact` | no | Detail mode: after reading the MAP, open selected records without repeating the `records` index and `timeline`; keeps `groupingSummary`, `representativeSession`, `stats`, selected ids, and requested payload sections |
|
|
410
|
+
| `payloadOnly` | no | Alias for `compact` |
|
|
411
|
+
| `parts` | no | Detail parts: `index`, `map`, `records`, `stats`, `transcript`, `tool_calls`, and record sections `full`, `conversation_messages`, `system`, `response`, `tools` |
|
|
407
412
|
| `sections` | no | Alias for `parts` |
|
|
408
413
|
| `recordIds` | no | Detail mode: open only these request/response record IDs |
|
|
409
414
|
| `toolCallIds` | no | Detail mode: return the FULL exact input/result from `ai_tool_call` for these tool-call ids (what the tool PRODUCED — see the lens caveat) |
|
|
410
415
|
| `consolidateContext` | no | When `true`, collapses repeated XML-like prompt context into a consolidated section |
|
|
411
416
|
|
|
417
|
+
Detail responses expose both lenses: `groupingSummary` is the authoritative grouping-wide summary
|
|
418
|
+
(counts, lanes, cost, status), while `representativeSession` names the concrete session used for
|
|
419
|
+
owner/runtime metadata. `session` remains only as a compatibility alias for older callers.
|
|
420
|
+
|
|
412
421
|
Spend: each search row carries `liveRequestCount`, `mockedRequestCount`, `lanes`, `testTaskId` and
|
|
413
422
|
`estimatedCost`/`estimatedCostUsd` = **live spend only** (a mocked turn is never priced, even when it replays a
|
|
414
423
|
recording with a cost). `pageTotals` sums the page. Turns recorded before the platform stored the mock flag are
|
|
@@ -416,12 +425,16 @@ recording with a cost). `pageTotals` sums the page. Turns recorded before the pl
|
|
|
416
425
|
is refused rather than silently widened; the applied bounds are echoed as `filters.rangeStart`/`rangeEnd`.
|
|
417
426
|
To audit one Test run: `{"action":"search","testTaskId":"<taskId>","pageSize":100}`.
|
|
418
427
|
|
|
419
|
-
Audit flow: `action:"search"` to find the grouping → `action:"detail"` + `summaryOnly:true` for the MAP → re-call detail with `recordIds`/`toolCallIds` + `parts` to open exactly what you need. Prefer the map → open flow over a full-detail dump. Remember the two-lens rule: `tool_calls`/`toolCallIds` is what the tool PRODUCED; `conversation_messages`/`transcript` is what the AI CONSUMED (after any `_offload`/`_hideResult`/`_message`/supersede/evict transform).
|
|
428
|
+
Audit flow: `action:"search"` to find the grouping → `action:"detail"` + `summaryOnly:true` for the MAP → re-call detail with `recordIds`/`toolCallIds` + `parts` to open exactly what you need, usually with `compact:true` once the map has chosen the row. Prefer the map → open flow over a full-detail dump. Remember the two-lens rule: `tool_calls`/`toolCallIds` is what the tool PRODUCED; `conversation_messages`/`transcript` is what the AI CONSUMED (after any `_offload`/`_hideResult`/`_message`/supersede/evict transform).
|
|
420
429
|
|
|
421
430
|
### `mcp_run_action`
|
|
422
431
|
Run an Action on a target account, with explicit prod/test data mode, optional staged branch resolution, and
|
|
423
432
|
staged-vs-DB provenance in the result.
|
|
424
433
|
|
|
434
|
+
Invoke it with `remits-cli tool --name mcp_run_action`. Despite the natural shorthand "run an Action", there
|
|
435
|
+
is no separate top-level `remits-cli action` / `actions` command and no `remits-cli run action` wrapper in
|
|
436
|
+
this CLI build.
|
|
437
|
+
|
|
425
438
|
Describe the Action first when the input shape is not obvious. This does not execute the Action:
|
|
426
439
|
|
|
427
440
|
```bash
|
|
@@ -832,6 +845,10 @@ Common uses:
|
|
|
832
845
|
- `remits-cli components branch <name> --diff <id> --component-type <kind>` — compare one variant against
|
|
833
846
|
current trunk.
|
|
834
847
|
- `remits-cli components branch <name> --subscribers` — list accounts resolving that branch.
|
|
848
|
+
- `remits-cli components branch <name> --copy-to <newBranch> [--dry-run] [--force]` — seed a new variant
|
|
849
|
+
branch with the source branch's stored overlays before the first safe sync of the new branch. Dry-run
|
|
850
|
+
reports the copy plan without writes; force is required when the target already has overlays or live
|
|
851
|
+
subscribers.
|
|
835
852
|
|
|
836
853
|
### `mcp_cache`
|
|
837
854
|
Bounded read-only investigation of the platform Redis keyspace — the way to see exactly what a staged
|