@remits/remits-cli 0.1.113 → 0.1.115
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +3 -3
- package/index.js +933 -52
- package/package.json +3 -2
- package/skills/remits-cli/SKILL.md +163 -3549
- package/skills/remits-cli/references/account-targeting.md +174 -0
- package/skills/remits-cli/references/agent-sessions.md +222 -0
- package/skills/remits-cli/references/branch-variants.md +391 -0
- package/skills/remits-cli/references/cli-state.md +158 -0
- package/skills/remits-cli/references/command-reference.md +270 -0
- package/skills/remits-cli/references/component-integrity.md +180 -0
- package/skills/remits-cli/references/component-resolution.md +221 -0
- package/skills/remits-cli/references/development-loop.md +418 -0
- package/skills/remits-cli/references/investigation.md +254 -0
- package/skills/remits-cli/references/support-tickets.md +483 -0
- package/skills/remits-cli/references/tool-reference.md +962 -0
- package/skills/remits-cli/references/troubleshooting.md +135 -0
|
@@ -0,0 +1,254 @@
|
|
|
1
|
+
# Investigation and Production Support
|
|
2
|
+
|
|
3
|
+
> A `remits-cli` skill reference. **Load this when** you are investigating live behavior: which record trail to read, how to correlate a processing chain, and the production support flows.
|
|
4
|
+
>
|
|
5
|
+
> The table of contents below carries **real line numbers** (`- L84 Some Heading`), resolved when
|
|
6
|
+
> this file is installed, so they are never stale. Read the head, pick your sections, and offset-read
|
|
7
|
+
> only those. The entry text is the heading verbatim, so it also greps.
|
|
8
|
+
|
|
9
|
+
## Table of Contents
|
|
10
|
+
|
|
11
|
+
- [The Investigation Model](#the-investigation-model)
|
|
12
|
+
- [Which tool reads which record](#which-tool-reads-which-record)
|
|
13
|
+
- [Correlation keys](#correlation-keys)
|
|
14
|
+
- [Reading a record's `content` — persisted context, not a memory dump](#reading-a-records-content--persisted-context-not-a-memory-dump)
|
|
15
|
+
- [HTTP audits](#http-audits)
|
|
16
|
+
- [AI activity](#ai-activity)
|
|
17
|
+
- [Node Reference Table](#node-reference-table)
|
|
18
|
+
- [Runtime node and `localMode`](#runtime-node-and-localmode)
|
|
19
|
+
- [Production Support Workflow](#production-support-workflow)
|
|
20
|
+
- [Investigation Strategy](#investigation-strategy)
|
|
21
|
+
- [Presenting Findings](#presenting-findings)
|
|
22
|
+
- [Verifying a Production Issue Fix](#verifying-a-production-issue-fix)
|
|
23
|
+
|
|
24
|
+
## The Investigation Model
|
|
25
|
+
|
|
26
|
+
Front stage — Schemas, Readers, Actions, Embeddables, Rules, HtmlTemplates, Agents, Tests, and the Firestore
|
|
27
|
+
**documents** that hold business data — is described in `platform-overview.md`. This section is only about
|
|
28
|
+
the **back-stage lifecycle records** an investigation actually reads, and how to read them.
|
|
29
|
+
|
|
30
|
+
### Which tool reads which record
|
|
31
|
+
|
|
32
|
+
The model — Documents (Firestore business data) vs Objects and their record trail (MySQL), linked by
|
|
33
|
+
`object_id` — is in `platform-overview.md` → *Objects, Documents, and the Record Trail*. What matters
|
|
34
|
+
here is the tool and the filterable fields:
|
|
35
|
+
|
|
36
|
+
| Record | Query it with | Fields worth filtering on |
|
|
37
|
+
|---|---|---|
|
|
38
|
+
| document | `mcp_firestore_search` | any schema field, plus `account_id`, `object_id`, `_lastModifiedAt` |
|
|
39
|
+
| `object` | `mcp_record_listing` / `mcp_record_view` | `status`, `type`, `name`, `referenceId` |
|
|
40
|
+
| `object_log` | `mcp_record_listing` / `mcp_record_view` | `type`, `description`, `content`, `threadGroupingId` |
|
|
41
|
+
| `event` | `mcp_record_listing` / `mcp_record_view` | `action`, `status`, `eventDate`, `threadGroupingId` |
|
|
42
|
+
| `alert` | `mcp_record_listing` / `mcp_record_view` | `status`, `type`, `active`, `threadGroupingId` |
|
|
43
|
+
| user activity session | `mcp_user_activity` | `userId`, `accountId`, `sessionKey`, `focusedOnly`, `traceId` pivots |
|
|
44
|
+
|
|
45
|
+
`mcp_record_listing` finds candidates when you do not know the id; `mcp_record_view` opens an exact one;
|
|
46
|
+
`mcp_object_activity` returns one Object's whole timeline in order.
|
|
47
|
+
|
|
48
|
+
**Check the data lane before concluding a record does not exist** — `--data-mode test` and `prod` read
|
|
49
|
+
different lanes (see `platform-overview.md` → *Test and Production Data Lanes*).
|
|
50
|
+
|
|
51
|
+
### Correlation keys
|
|
52
|
+
|
|
53
|
+
- **`object_id`** — links documents to records. Pivot `mcp_firestore_search` → `mcp_object_activity`.
|
|
54
|
+
- **`threadGroupingId`** — groups every record and log line from one processing chain, and is also the
|
|
55
|
+
request's trace id. Pivot into `mcp_system_logs` and `mcp_performance_trace`.
|
|
56
|
+
- **`node`** — resolves to `serviceName` + `region` for log queries (table below).
|
|
57
|
+
- **`sessionKey`** — salted user-activity session key. Pivot `mcp_user_activity` `sessions` → `story`;
|
|
58
|
+
each beat then carries a `traceId` / `threadGroupingId` for trace and log investigation.
|
|
59
|
+
- **`sessionId`** — pivots into persisted AI activity via `mcp_ai_session_search`.
|
|
60
|
+
|
|
61
|
+
### Reading a record's `content` — persisted context, not a memory dump
|
|
62
|
+
|
|
63
|
+
`content` on an `object`, `object_log`, `event`, or `alert` is the **sanitized snapshot** the platform
|
|
64
|
+
deliberately kept of what the producing component knew at that workflow step. It is not everything that was
|
|
65
|
+
in memory.
|
|
66
|
+
|
|
67
|
+
- `Object.content` — ingestion/request/file context when the Object was created or updated.
|
|
68
|
+
- `ObjectLog.content` — the logging component's point-in-time view.
|
|
69
|
+
- `Event.content` — the scheduling component's view, plus the explicit event options.
|
|
70
|
+
- `Alert.content` — the raising component's view, plus the explicit alert body.
|
|
71
|
+
|
|
72
|
+
**Missing or `[REDACTED]` does not mean the component never had it.** Before persisting, the platform drops
|
|
73
|
+
non-serializable objects, strips `requestBody` / `params` / `token` recursively, summarizes very large
|
|
74
|
+
strings, prunes oversized maps and collections, and redacts secret-looking keys (`password`, `secret`,
|
|
75
|
+
`authorization`, `cookie`, `clientSecret`, …) plus any Account/User schema field marked `sensitive: true`.
|
|
76
|
+
When explaining a record, distinguish what the workflow *had* at runtime from what the platform *kept*.
|
|
77
|
+
|
|
78
|
+
**Provenance travels in the context itself:**
|
|
79
|
+
|
|
80
|
+
- **`source_bcd`** — the immediate component that produced this record, e.g. `[Action:18] DataPlus Invoice Posting`
|
|
81
|
+
- **`upstream_source_bcd`** — the prior workflow hop, when the context was inherited from one
|
|
82
|
+
|
|
83
|
+
**Never interpret `content` in isolation.** Pair it with the record type, `source_bcd`, any
|
|
84
|
+
`upstream_source_bcd`, the `object_id` timeline, the `threadGroupingId` chain, and the producing
|
|
85
|
+
component's source. The goal is not to find a suspicious record — it is to explain **why the context looks
|
|
86
|
+
exactly the way it does relative to the workflow step that produced it**.
|
|
87
|
+
|
|
88
|
+
### HTTP audits
|
|
89
|
+
|
|
90
|
+
Raw inbound (Reader) and outbound (`rest(...)`) HTTP is persisted to Firestore — but **only for accounts
|
|
91
|
+
with `Account.enableHttpAudits`**. When it is off, the absence of audit documents proves nothing.
|
|
92
|
+
|
|
93
|
+
Audits live in **monthly** collections, not one global collection:
|
|
94
|
+
|
|
95
|
+
```
|
|
96
|
+
http-audits/http-audits-YYYY-MM/entries
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
Start with the month the request ran in, and check the adjacent month if the run may have crossed a
|
|
100
|
+
boundary. Query them with `mcp_firestore_search` like any other collection; the filters worth reaching for
|
|
101
|
+
are `direction` (`INBOUND` / `OUTBOUND`), `success`, `request.method`, `request.path`, `component.name`, and
|
|
102
|
+
`response.statusCode`. Some paths are deliberately excluded from persistence, so a missing audit is not
|
|
103
|
+
proof a call was never made.
|
|
104
|
+
|
|
105
|
+
Reach for audits when the question is *what exact request went out, what came back, and did this component
|
|
106
|
+
actually make the call* — then use the component source plus the payload to place the fault in request
|
|
107
|
+
formation, the partner's response, or downstream processing. Full captured shape, websocket behavior, and
|
|
108
|
+
embeddable patterns: `features/http-audits.md` (`mcp_get_guide`).
|
|
109
|
+
|
|
110
|
+
### AI activity
|
|
111
|
+
|
|
112
|
+
Every `ai()` call and Agent turn is persisted. Two ways in:
|
|
113
|
+
|
|
114
|
+
- **`mcp_ai_session_search`** — search groupings and open their detail. A grouping spans every session
|
|
115
|
+
sharing one grouping id: the agent turns plus its guardrail and internal `ai()` calls. Use the
|
|
116
|
+
**map → open** flow described under its `tool-reference.md` entry, never a whole-detail dump.
|
|
117
|
+
- **`ai_request_response([sessionId: id])`** — from inside component code, loads the stored
|
|
118
|
+
request/response history for a session (`first: true` / `last: true` for one record). This is also how
|
|
119
|
+
you replay a stored provider response in a test.
|
|
120
|
+
|
|
121
|
+
If a document or agent state already carries a session id, that is a direct pivot into persisted AI
|
|
122
|
+
activity.
|
|
123
|
+
|
|
124
|
+
**Before tuning any prompt, read the session.** `features/ai-session-investigation.md` owns the method —
|
|
125
|
+
per-turn forensics, what each layer proves, and the rule that what a tool **PRODUCED** is not necessarily
|
|
126
|
+
what the model **CONSUMED**. `features/ai-strategy.md` owns what to change once you know. Changing a prompt
|
|
127
|
+
before reading the persisted request/response is guessing.
|
|
128
|
+
|
|
129
|
+
### Node Reference Table
|
|
130
|
+
|
|
131
|
+
| Node Name | Service Name | Region |
|
|
132
|
+
|---|---|---|
|
|
133
|
+
| remitsAdmin-east5 | remits | us-east5 |
|
|
134
|
+
| remitsActions | remits-actions | us-east1 |
|
|
135
|
+
| remitsAdmin | remits | us-east1 |
|
|
136
|
+
|
|
137
|
+
`mcp_system_logs` accepts `node` directly and resolves it automatically.
|
|
138
|
+
|
|
139
|
+
### Runtime node and `localMode`
|
|
140
|
+
|
|
141
|
+
The deployed `remits` service in `us-east5` (`remitsAdmin-east5`) runs with the platform setting
|
|
142
|
+
`localMode=true`. If someone says "localModel" in this context, confirm they mean this `localMode`
|
|
143
|
+
setting. Operationally, immediate async follow-on work stays on the same Cloud Run service/node instead of
|
|
144
|
+
being sharded to `remits-actions`:
|
|
145
|
+
|
|
146
|
+
- Pub/Sub-style follow-on messages are handled locally after commit.
|
|
147
|
+
- Near-immediate tasks are handled locally when `localMode` is enabled. Future scheduled tasks still use
|
|
148
|
+
Cloud Tasks.
|
|
149
|
+
- Local worker hops preserve the run context, including staged-source resolution, data mode,
|
|
150
|
+
`threadGroupingId`, and trace correlation.
|
|
151
|
+
- Durable boundaries such as async HTTP ingress and Events carry that same run context across the queue.
|
|
152
|
+
|
|
153
|
+
For investigations on the default deployed host (`https://remits-529558023549.us-east5.run.app`), do not
|
|
154
|
+
assume "async" means `remitsActions` / `us-east1`. Start with `node:"remitsAdmin-east5"` and the
|
|
155
|
+
`threadGroupingId`; pivot to `remitsActions` only when the Event delivery envelope, log line, or returned
|
|
156
|
+
node says the work actually ran there.
|
|
157
|
+
|
|
158
|
+
This does not change the data-lane rule: a non-null `TestMode` can exist only to carry branch/staged-source
|
|
159
|
+
resolution. Data isolation is decided by CLI `--data-mode`: a branch-scoped `--data-mode prod` run is still
|
|
160
|
+
prod data, while `--data-mode test` remains isolated test data.
|
|
161
|
+
|
|
162
|
+
## Production Support Workflow
|
|
163
|
+
|
|
164
|
+
Switch to prod mode for investigations:
|
|
165
|
+
|
|
166
|
+
```bash
|
|
167
|
+
remits-cli data-mode set prod
|
|
168
|
+
```
|
|
169
|
+
|
|
170
|
+
### Investigation Strategy
|
|
171
|
+
|
|
172
|
+
Before starting an investigation outside the confirmed current repo:
|
|
173
|
+
1. Read `~/.remits-cli/account-repos.json`
|
|
174
|
+
2. Switch to the best local repo candidate
|
|
175
|
+
3. Read that repo's `account-info.json`
|
|
176
|
+
4. Confirm whether you are in `CLIENT`, `PLATFORM`, or `PRODUCT` context
|
|
177
|
+
5. Then continue with the investigation flow below
|
|
178
|
+
|
|
179
|
+
**Document-First** (most common — user reports a data issue):
|
|
180
|
+
1. `mcp_account_view` — understand the account's schemas and components.
|
|
181
|
+
2. `mcp_firestore_search` — find the document, capture its `object_id`.
|
|
182
|
+
3. If the issue involves inbound or outbound HTTP behavior and the account has `enableHttpAudits`, query `http-audits/http-audits-YYYY-MM/entries` with `mcp_firestore_search`.
|
|
183
|
+
4. `mcp_object_activity` — scan the timeline for warnings, errors, unexpected events.
|
|
184
|
+
5. `mcp_record_listing` — search or filter alerts, events, object logs, or objects when you need to find the suspicious record first.
|
|
185
|
+
6. `mcp_record_view` — drill into suspicious entries for full content.
|
|
186
|
+
7. `mcp_ai_session_search` — if the workflow involves AI, inspect session groupings, prompts, tool definitions, and responses in human-readable form.
|
|
187
|
+
8. `mcp_user_activity` — for "user X is slow right now" reports, list sessions by `userId`/`accountId`, open the session story, and use the returned beat pivots.
|
|
188
|
+
9. `mcp_performance_trace` — for slow/sluggish reports, open the beat `traceId` with `action:"trace"`; use `action:"slowest"` when you only have a broad time window.
|
|
189
|
+
10. `mcp_system_logs` — correlate via `threadGroupingId` for raw log context when the trace needs supporting log lines.
|
|
190
|
+
11. If no local checkout exists for the responsible implementation account, use
|
|
191
|
+
`mcp_component_view`/`mcp_component_grep` to explain how the component works. If the repo exists on
|
|
192
|
+
this machine, inspect the branch/files there instead; the remote component tools are fallback and
|
|
193
|
+
live-DB comparison surfaces, not the starting point for source comprehension.
|
|
194
|
+
|
|
195
|
+
**Slow / sluggish user report:**
|
|
196
|
+
|
|
197
|
+
1. Resolve the reporting user/account with `mcp_account_user_admin` if you only have an email/name.
|
|
198
|
+
2. Call `mcp_user_activity` with `action:"sessions"` and `userId` or `accountId`.
|
|
199
|
+
3. Open the likely row with `action:"story"` and inspect beat labels, status, `ms`, `node`, and `traceId`.
|
|
200
|
+
4. Open slow or failed beat pivots with `mcp_performance_trace` before querying raw logs.
|
|
201
|
+
5. Use the returned `mcp_system_logs` pivot only when the trace needs surrounding log lines.
|
|
202
|
+
6. If there is no live session, call `mcp_user_activity` `action:"watch"` for the user/account, ask for reproduction, then read `sessions`/`story` again. Focused sessions retain sanitized request detail and emit archived `REMITS_ACTIVITY` log lines.
|
|
203
|
+
|
|
204
|
+
**Error or Alert Investigation:**
|
|
205
|
+
1. `mcp_record_listing` — search by alert type, content, error text, action, status, `threadGroupingId`, or other exact-match record properties when you do not yet know the record ID.
|
|
206
|
+
2. `mcp_record_view` — inspect the chosen record/event/alert/object in full once you have its ID.
|
|
207
|
+
3. Use `object_id` + `threadGroupingId` to pull full timeline and logs.
|
|
208
|
+
4. Cross-check Firestore document state.
|
|
209
|
+
5. Identify `source_bcd` and any `upstream_source_bcd`.
|
|
210
|
+
6. Explain the record in terms of the workflow step that produced it, not as a generic JSON blob.
|
|
211
|
+
7. If the context looks missing, redacted, or truncated, consider sanitization rules before concluding data was never present.
|
|
212
|
+
8. If AI behavior is part of the symptom, use `mcp_ai_session_search` and compare the persisted session content against the Agent component implementation and `features/ai-support.md`.
|
|
213
|
+
|
|
214
|
+
**Stuck / failed / recovered Event:**
|
|
215
|
+
|
|
216
|
+
Do **not** open the Action source first. The platform records each attempt's delivery envelope — which
|
|
217
|
+
queue delivered it, which delivery attempt this was, and how long it was ever allowed to run — and
|
|
218
|
+
classifies the failure for you.
|
|
219
|
+
|
|
220
|
+
```bash
|
|
221
|
+
remits-cli tool --name mcp_event_diagnostics --input '{"accountId":49,"eventId":18838}' --data-mode prod
|
|
222
|
+
```
|
|
223
|
+
|
|
224
|
+
The same classifier is available from a Test or any component as `eventDiagnostics(18838)`.
|
|
225
|
+
|
|
226
|
+
**Read `classification` before anything else** — only `APPLICATION_FAILURE` means the bug is in the
|
|
227
|
+
component. The full classification table, what each `abandonmentCause` implies, and the returned `pivots`
|
|
228
|
+
are under **`mcp_event_diagnostics`** in `tool-reference.md`. One thing to check every time:
|
|
229
|
+
`delivery.deliveryAttempt` above `1` means Cloud Tasks had **already** retried this event, so any
|
|
230
|
+
non-idempotent side effect may have run more than once — look for duplicate records before concluding the
|
|
231
|
+
component "ran twice for no reason".
|
|
232
|
+
|
|
233
|
+
Full detail: `features/observability.md` and `features/events-builder-guide.md` (`mcp_get_guide`).
|
|
234
|
+
|
|
235
|
+
### Presenting Findings
|
|
236
|
+
|
|
237
|
+
Users are not engineers. When reporting investigation results:
|
|
238
|
+
- Lead with what happened in plain language.
|
|
239
|
+
- Show the evidence (document values, timeline events, log excerpts).
|
|
240
|
+
- Explain why it happened if you can determine the cause.
|
|
241
|
+
- Recommend what to do next — in terms the user can act on.
|
|
242
|
+
|
|
243
|
+
### Verifying a Production Issue Fix
|
|
244
|
+
|
|
245
|
+
When a bug is reported from production, use this pattern:
|
|
246
|
+
|
|
247
|
+
1. Investigate the live issue in **prod mode** and identify the exact affected document IDs, collection names, account IDs, and component path.
|
|
248
|
+
2. Make the code change in the owning `PLATFORM` or `PRODUCT` repo when the defect is in shared implementation.
|
|
249
|
+
3. Verify in **test mode**, not prod.
|
|
250
|
+
4. Prefer a **Test component** when the behavior can be asserted programmatically, because that creates a durable regression suite and lets you explicitly construct the necessary data, operations, and assertions.
|
|
251
|
+
5. Use `remits-cli token` plus `playwright-cli` when the proof is visual or interaction-driven.
|
|
252
|
+
6. If useful, create or update a dedicated embeddable "playground" in test mode to reproduce the scenario in a controlled way.
|
|
253
|
+
|
|
254
|
+
Do not move production customer data into another account's test collection as a routine verification strategy. If you cannot verify with a Test component, Playwright flow, or controlled test-mode embeddable, explain the gap clearly instead of improvising with live production validation.
|