@remits/remits-cli 0.1.113 → 0.1.114

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,251 @@
1
+ # Investigation and Production Support
2
+
3
+ > A `remits-cli` skill reference. **Load this when** you are investigating live behavior: which record trail to read, how to correlate a processing chain, and the production support flows.
4
+ >
5
+ > The table of contents below carries **real line numbers** (`- L84 Some Heading`), resolved when
6
+ > this file is installed, so they are never stale. Read the head, pick your sections, and offset-read
7
+ > only those. The entry text is the heading verbatim, so it also greps.
8
+
9
+ ## Table of Contents
10
+
11
+ - [The Investigation Model](#the-investigation-model)
12
+ - [Which tool reads which record](#which-tool-reads-which-record)
13
+ - [Correlation keys](#correlation-keys)
14
+ - [Reading a record's `content` — persisted context, not a memory dump](#reading-a-records-content--persisted-context-not-a-memory-dump)
15
+ - [HTTP audits](#http-audits)
16
+ - [AI activity](#ai-activity)
17
+ - [Node Reference Table](#node-reference-table)
18
+ - [Runtime node and `localMode`](#runtime-node-and-localmode)
19
+ - [Production Support Workflow](#production-support-workflow)
20
+ - [Investigation Strategy](#investigation-strategy)
21
+ - [Presenting Findings](#presenting-findings)
22
+ - [Verifying a Production Issue Fix](#verifying-a-production-issue-fix)
23
+
24
+ ## The Investigation Model
25
+
26
+ Front stage — Schemas, Readers, Actions, Embeddables, Rules, HtmlTemplates, Agents, Tests, and the Firestore
27
+ **documents** that hold business data — is described in `platform-overview.md`. This section is only about
28
+ the **back-stage lifecycle records** an investigation actually reads, and how to read them.
29
+
30
+ ### Which tool reads which record
31
+
32
+ The model — Documents (Firestore business data) vs Objects and their record trail (MySQL), linked by
33
+ `object_id` — is in `platform-overview.md` → *Objects, Documents, and the Record Trail*. What matters
34
+ here is the tool and the filterable fields:
35
+
36
+ | Record | Query it with | Fields worth filtering on |
37
+ |---|---|---|
38
+ | document | `mcp_firestore_search` | any schema field, plus `account_id`, `object_id`, `_lastModifiedAt` |
39
+ | `object` | `mcp_record_listing` / `mcp_record_view` | `status`, `type`, `name`, `referenceId` |
40
+ | `object_log` | `mcp_record_listing` / `mcp_record_view` | `type`, `description`, `content`, `threadGroupingId` |
41
+ | `event` | `mcp_record_listing` / `mcp_record_view` | `action`, `status`, `eventDate`, `threadGroupingId` |
42
+ | `alert` | `mcp_record_listing` / `mcp_record_view` | `status`, `type`, `active`, `threadGroupingId` |
43
+ | user activity session | `mcp_user_activity` | `userId`, `accountId`, `sessionKey`, `focusedOnly`, `traceId` pivots |
44
+
45
+ `mcp_record_listing` finds candidates when you do not know the id; `mcp_record_view` opens an exact one;
46
+ `mcp_object_activity` returns one Object's whole timeline in order.
47
+
48
+ **Check the data lane before concluding a record does not exist** — `--data-mode test` and `prod` read
49
+ different lanes (see `platform-overview.md` → *Test and Production Data Lanes*).
50
+
51
+ ### Correlation keys
52
+
53
+ - **`object_id`** — links documents to records. Pivot `mcp_firestore_search` → `mcp_object_activity`.
54
+ - **`threadGroupingId`** — groups every record and log line from one processing chain, and is also the
55
+ request's trace id. Pivot into `mcp_system_logs` and `mcp_performance_trace`.
56
+ - **`node`** — resolves to `serviceName` + `region` for log queries (table below).
57
+ - **`sessionKey`** — salted user-activity session key. Pivot `mcp_user_activity` `sessions` → `story`;
58
+ each beat then carries a `traceId` / `threadGroupingId` for trace and log investigation.
59
+ - **`sessionId`** — pivots into persisted AI activity via `mcp_ai_session_search`.
60
+
61
+ ### Reading a record's `content` — persisted context, not a memory dump
62
+
63
+ `content` on an `object`, `object_log`, `event`, or `alert` is the **sanitized snapshot** the platform
64
+ deliberately kept of what the producing component knew at that workflow step. It is not everything that was
65
+ in memory.
66
+
67
+ - `Object.content` — ingestion/request/file context when the Object was created or updated.
68
+ - `ObjectLog.content` — the logging component's point-in-time view.
69
+ - `Event.content` — the scheduling component's view, plus the explicit event options.
70
+ - `Alert.content` — the raising component's view, plus the explicit alert body.
71
+
72
+ **Missing or `[REDACTED]` does not mean the component never had it.** Before persisting, the platform drops
73
+ non-serializable objects, strips `requestBody` / `params` / `token` recursively, summarizes very large
74
+ strings, prunes oversized maps and collections, and redacts secret-looking keys (`password`, `secret`,
75
+ `authorization`, `cookie`, `clientSecret`, …) plus any Account/User schema field marked `sensitive: true`.
76
+ When explaining a record, distinguish what the workflow *had* at runtime from what the platform *kept*.
77
+
78
+ **Provenance travels in the context itself:**
79
+
80
+ - **`source_bcd`** — the immediate component that produced this record, e.g. `[Action:18] DataPlus Invoice Posting`
81
+ - **`upstream_source_bcd`** — the prior workflow hop, when the context was inherited from one
82
+
83
+ **Never interpret `content` in isolation.** Pair it with the record type, `source_bcd`, any
84
+ `upstream_source_bcd`, the `object_id` timeline, the `threadGroupingId` chain, and the producing
85
+ component's source. The goal is not to find a suspicious record — it is to explain **why the context looks
86
+ exactly the way it does relative to the workflow step that produced it**.
87
+
88
+ ### HTTP audits
89
+
90
+ Raw inbound (Reader) and outbound (`rest(...)`) HTTP is persisted to Firestore — but **only for accounts
91
+ with `Account.enableHttpAudits`**. When it is off, the absence of audit documents proves nothing.
92
+
93
+ Audits live in **monthly** collections, not one global collection:
94
+
95
+ ```
96
+ http-audits/http-audits-YYYY-MM/entries
97
+ ```
98
+
99
+ Start with the month the request ran in, and check the adjacent month if the run may have crossed a
100
+ boundary. Query them with `mcp_firestore_search` like any other collection; the filters worth reaching for
101
+ are `direction` (`INBOUND` / `OUTBOUND`), `success`, `request.method`, `request.path`, `component.name`, and
102
+ `response.statusCode`. Some paths are deliberately excluded from persistence, so a missing audit is not
103
+ proof a call was never made.
104
+
105
+ Reach for audits when the question is *what exact request went out, what came back, and did this component
106
+ actually make the call* — then use the component source plus the payload to place the fault in request
107
+ formation, the partner's response, or downstream processing. Full captured shape, websocket behavior, and
108
+ embeddable patterns: `features/http-audits.md` (`mcp_get_guide`).
109
+
110
+ ### AI activity
111
+
112
+ Every `ai()` call and Agent turn is persisted. Two ways in:
113
+
114
+ - **`mcp_ai_session_search`** — search groupings and open their detail. A grouping spans every session
115
+ sharing one grouping id: the agent turns plus its guardrail and internal `ai()` calls. Use the
116
+ **map → open** flow described under its `tool-reference.md` entry, never a whole-detail dump.
117
+ - **`ai_request_response([sessionId: id])`** — from inside component code, loads the stored
118
+ request/response history for a session (`first: true` / `last: true` for one record). This is also how
119
+ you replay a stored provider response in a test.
120
+
121
+ If a document or agent state already carries a session id, that is a direct pivot into persisted AI
122
+ activity.
123
+
124
+ **Before tuning any prompt, read the session.** `features/ai-session-investigation.md` owns the method —
125
+ per-turn forensics, what each layer proves, and the rule that what a tool **PRODUCED** is not necessarily
126
+ what the model **CONSUMED**. `features/ai-strategy.md` owns what to change once you know. Changing a prompt
127
+ before reading the persisted request/response is guessing.
128
+
129
+ ### Node Reference Table
130
+
131
+ | Node Name | Service Name | Region |
132
+ |---|---|---|
133
+ | remitsAdmin-east5 | remits | us-east5 |
134
+ | remitsActions | remits-actions | us-east1 |
135
+ | remitsAdmin | remits | us-east1 |
136
+
137
+ `mcp_system_logs` accepts `node` directly and resolves it automatically.
138
+
139
+ ### Runtime node and `localMode`
140
+
141
+ The deployed `remits` service in `us-east5` (`remitsAdmin-east5`) runs with the platform setting
142
+ `localMode=true`. If someone says "localModel" in this context, confirm they mean this `localMode`
143
+ setting. Operationally, immediate async follow-on work stays on the same Cloud Run service/node instead of
144
+ being sharded to `remits-actions`:
145
+
146
+ - Pub/Sub-style follow-on messages are handled locally after commit.
147
+ - Near-immediate tasks are handled locally when `localMode` is enabled. Future scheduled tasks still use
148
+ Cloud Tasks.
149
+ - Local worker hops preserve the run context, including staged-source resolution, data mode,
150
+ `threadGroupingId`, and trace correlation.
151
+ - Durable boundaries such as async HTTP ingress and Events carry that same run context across the queue.
152
+
153
+ For investigations on the default deployed host (`https://remits-529558023549.us-east5.run.app`), do not
154
+ assume "async" means `remitsActions` / `us-east1`. Start with `node:"remitsAdmin-east5"` and the
155
+ `threadGroupingId`; pivot to `remitsActions` only when the Event delivery envelope, log line, or returned
156
+ node says the work actually ran there.
157
+
158
+ This does not change the data-lane rule: a non-null `TestMode` can exist only to carry branch/staged-source
159
+ resolution. Data isolation is decided by CLI `--data-mode`: a branch-scoped `--data-mode prod` run is still
160
+ prod data, while `--data-mode test` remains isolated test data.
161
+
162
+ ## Production Support Workflow
163
+
164
+ Switch to prod mode for investigations:
165
+
166
+ ```bash
167
+ remits-cli data-mode set prod
168
+ ```
169
+
170
+ ### Investigation Strategy
171
+
172
+ Before starting an investigation outside the confirmed current repo:
173
+ 1. Read `~/.remits-cli/account-repos.json`
174
+ 2. Switch to the best local repo candidate
175
+ 3. Read that repo's `account-info.json`
176
+ 4. Confirm whether you are in `CLIENT`, `PLATFORM`, or `PRODUCT` context
177
+ 5. Then continue with the investigation flow below
178
+
179
+ **Document-First** (most common — user reports a data issue):
180
+ 1. `mcp_account_view` — understand the account's schemas and components.
181
+ 2. `mcp_firestore_search` — find the document, capture its `object_id`.
182
+ 3. If the issue involves inbound or outbound HTTP behavior and the account has `enableHttpAudits`, query `http-audits/http-audits-YYYY-MM/entries` with `mcp_firestore_search`.
183
+ 4. `mcp_object_activity` — scan the timeline for warnings, errors, unexpected events.
184
+ 5. `mcp_record_listing` — search or filter alerts, events, object logs, or objects when you need to find the suspicious record first.
185
+ 6. `mcp_record_view` — drill into suspicious entries for full content.
186
+ 7. `mcp_ai_session_search` — if the workflow involves AI, inspect session groupings, prompts, tool definitions, and responses in human-readable form.
187
+ 8. `mcp_user_activity` — for "user X is slow right now" reports, list sessions by `userId`/`accountId`, open the session story, and use the returned beat pivots.
188
+ 9. `mcp_performance_trace` — for slow/sluggish reports, open the beat `traceId` with `action:"trace"`; use `action:"slowest"` when you only have a broad time window.
189
+ 10. `mcp_system_logs` — correlate via `threadGroupingId` for raw log context when the trace needs supporting log lines.
190
+ 11. `mcp_component_view`/`mcp_component_grep` — explain how the responsible component works.
191
+
192
+ **Slow / sluggish user report:**
193
+
194
+ 1. Resolve the reporting user/account with `mcp_account_user_admin` if you only have an email/name.
195
+ 2. Call `mcp_user_activity` with `action:"sessions"` and `userId` or `accountId`.
196
+ 3. Open the likely row with `action:"story"` and inspect beat labels, status, `ms`, `node`, and `traceId`.
197
+ 4. Open slow or failed beat pivots with `mcp_performance_trace` before querying raw logs.
198
+ 5. Use the returned `mcp_system_logs` pivot only when the trace needs surrounding log lines.
199
+ 6. If there is no live session, call `mcp_user_activity` `action:"watch"` for the user/account, ask for reproduction, then read `sessions`/`story` again. Focused sessions retain sanitized request detail and emit archived `REMITS_ACTIVITY` log lines.
200
+
201
+ **Error or Alert Investigation:**
202
+ 1. `mcp_record_listing` — search by alert type, content, error text, action, status, `threadGroupingId`, or other exact-match record properties when you do not yet know the record ID.
203
+ 2. `mcp_record_view` — inspect the chosen record/event/alert/object in full once you have its ID.
204
+ 3. Use `object_id` + `threadGroupingId` to pull full timeline and logs.
205
+ 4. Cross-check Firestore document state.
206
+ 5. Identify `source_bcd` and any `upstream_source_bcd`.
207
+ 6. Explain the record in terms of the workflow step that produced it, not as a generic JSON blob.
208
+ 7. If the context looks missing, redacted, or truncated, consider sanitization rules before concluding data was never present.
209
+ 8. If AI behavior is part of the symptom, use `mcp_ai_session_search` and compare the persisted session content against the Agent component implementation and `features/ai-support.md`.
210
+
211
+ **Stuck / failed / recovered Event:**
212
+
213
+ Do **not** open the Action source first. The platform records each attempt's delivery envelope — which
214
+ queue delivered it, which delivery attempt this was, and how long it was ever allowed to run — and
215
+ classifies the failure for you.
216
+
217
+ ```bash
218
+ remits-cli tool --name mcp_event_diagnostics --input '{"accountId":49,"eventId":18838}' --data-mode prod
219
+ ```
220
+
221
+ The same classifier is available from a Test or any component as `eventDiagnostics(18838)`.
222
+
223
+ **Read `classification` before anything else** — only `APPLICATION_FAILURE` means the bug is in the
224
+ component. The full classification table, what each `abandonmentCause` implies, and the returned `pivots`
225
+ are under **`mcp_event_diagnostics`** in `tool-reference.md`. One thing to check every time:
226
+ `delivery.deliveryAttempt` above `1` means Cloud Tasks had **already** retried this event, so any
227
+ non-idempotent side effect may have run more than once — look for duplicate records before concluding the
228
+ component "ran twice for no reason".
229
+
230
+ Full detail: `features/observability.md` and `features/events-builder-guide.md` (`mcp_get_guide`).
231
+
232
+ ### Presenting Findings
233
+
234
+ Users are not engineers. When reporting investigation results:
235
+ - Lead with what happened in plain language.
236
+ - Show the evidence (document values, timeline events, log excerpts).
237
+ - Explain why it happened if you can determine the cause.
238
+ - Recommend what to do next — in terms the user can act on.
239
+
240
+ ### Verifying a Production Issue Fix
241
+
242
+ When a bug is reported from production, use this pattern:
243
+
244
+ 1. Investigate the live issue in **prod mode** and identify the exact affected document IDs, collection names, account IDs, and component path.
245
+ 2. Make the code change in the owning `PLATFORM` or `PRODUCT` repo when the defect is in shared implementation.
246
+ 3. Verify in **test mode**, not prod.
247
+ 4. Prefer a **Test component** when the behavior can be asserted programmatically, because that creates a durable regression suite and lets you explicitly construct the necessary data, operations, and assertions.
248
+ 5. Use `remits-cli token` plus `playwright-cli` when the proof is visual or interaction-driven.
249
+ 6. If useful, create or update a dedicated embeddable "playground" in test mode to reproduce the scenario in a controlled way.
250
+
251
+ Do not move production customer data into another account's test collection as a routine verification strategy. If you cannot verify with a Test component, Playwright flow, or controlled test-mode embeddable, explain the gap clearly instead of improvising with live production validation.
@@ -0,0 +1,389 @@
1
+ # Support Tickets
2
+
3
+ > A `remits-cli` skill reference. **Load this when** a ticket is part of the request, you are registering to work tickets, or you are the worker inside an autonomous ticket run.
4
+ >
5
+ > The table of contents below carries **real line numbers** (`- L84 Some Heading`), resolved when
6
+ > this file is installed, so they are never stale. Read the head, pick your sections, and offset-read
7
+ > only those. The entry text is the heading verbatim, so it also greps.
8
+
9
+ ## Table of Contents
10
+
11
+ - [Support Ticket Mental Model](#support-ticket-mental-model)
12
+ - [You are an agent, and you register yourself](#you-are-an-agent-and-you-register-yourself)
13
+ - [Autonomous: one ticket, one process](#autonomous-one-ticket-one-process)
14
+ - [The manual loop](#the-manual-loop)
15
+ - [If you are the worker](#if-you-are-the-worker)
16
+ - [Before you edit anything: where you are, and whether you may](#before-you-edit-anything-where-you-are-and-whether-you-may)
17
+ - [Your account's process is binding, and it is already in your brief](#your-accounts-process-is-binding-and-it-is-already-in-your-brief)
18
+ - [Seeing the queue as a human does](#seeing-the-queue-as-a-human-does)
19
+ - [Agent components are workers too](#agent-components-are-workers-too)
20
+ - [Moving a ticket through its lifecycle](#moving-a-ticket-through-its-lifecycle)
21
+ - [When a worker needs a decision from a human](#when-a-worker-needs-a-decision-from-a-human)
22
+
23
+ ## Support Ticket Mental Model
24
+
25
+ Support tickets are a **first-class platform capability**, not an account convention. Every Remits
26
+ account reads and writes the `support_tickets` collection without owning a Schema, tickets are anchor
27
+ `Object`s on the account the work belongs to, and the lifecycle vocabulary is defined once in the
28
+ platform.
29
+
30
+ **The platform is deliberately not the standard for ticket workflow.** Different Remits
31
+ platforms and products integrate with different systems — Zendesk, Jira, a customer's own portal —
32
+ and each end client has its own rules for how tickets move. Remits owns the *record* and the *verbs*;
33
+ each front-stage platform builds its own workflow on top through its own Embeddables, Rules, and
34
+ Actions. So what you see through `remits-cli` is the shared record underneath every one of those
35
+ workflows, and never one product's view of it.
36
+
37
+ ### You are an agent, and you register yourself
38
+
39
+ **If the user says anything like "register to become a support agent", or "work support tickets",
40
+ run exactly this and nothing else first:**
41
+
42
+ ```bash
43
+ remits-cli agent serve
44
+ ```
45
+
46
+ No flags, no setup, **no need to be in an account repo** — run it from wherever the session started.
47
+ It registers this terminal for every account repo indexed on this machine that your user can reach,
48
+ against the production platform and the production lane, and then **starts working tickets on its
49
+ own**. It returns immediately; a background supervisor does the rest.
50
+
51
+ Use `remits-cli agent register` instead **only** when the user wants presence without autonomy — a
52
+ human, or you in this very tab, will work the tickets by hand. Registering alone starts nothing: it
53
+ makes the session routable and then waits to be asked.
54
+
55
+ Three tabs running an agent are three agents. Each is independently routable, each reports its own
56
+ activity, and each stops receiving work when its tab closes — a heartbeat anchored to the session dies
57
+ with it, so a crashed agent and a quit agent look identical to the platform.
58
+
59
+ **Nothing is pushed at you.** A terminal mid-task cannot receive a push, so delivery is you asking.
60
+ That is why the routing decision is a durable field on the ticket rather than a message: you can
61
+ restart this terminal, register again, and the work is still there.
62
+
63
+ Two facts about a ticket are separate and must stay separate:
64
+
65
+ | Fact | Field | Means |
66
+ |---|---|---|
67
+ | **Routed** | `routedAgentId` | which agent session should pick this up — delivery |
68
+ | **Claimed** | `claimedAgentId` | a worker process is running on it *right now* |
69
+ | **Owned** | `assignedTo` | who has claimed it — accountability |
70
+ | **Status** | `status` | where it is in the workflow |
71
+
72
+ A ticket routed to you is not yet yours. Claim it with `accept`, and the queue then shows it owned.
73
+
74
+ The **claim** is the supervisor's, not yours — it exists so two workers never start on one ticket,
75
+ and it expires on its own so a killed worker cannot park a ticket forever. You do not manage it.
76
+
77
+ ### Autonomous: one ticket, one process
78
+
79
+ ```bash
80
+ remits-cli agent serve # workers are whichever agent THIS session is
81
+ remits-cli agent serve --worker-agent codex --max-concurrent 2
82
+ remits-cli agent serve --mode investigate # read-only workers: no file edits
83
+ remits-cli agent workers # what is running right now
84
+ remits-cli agent release # stop serving, go offline
85
+ ```
86
+
87
+ **Run it in a plain terminal tab.** That tab becomes the agent host: `serve` returns immediately, a
88
+ detached supervisor is anchored to the tab's shell, and closing the tab stops the agent. Nothing in
89
+ that tab is an AI session — the AI only ever appears as the worker processes the supervisor spawns.
90
+
91
+ Starting it from inside an AI session works too and self-detects the worker kind, but it is the
92
+ lesser setup: it parks an interactive session as a heartbeat holder while its workers do the work.
93
+
94
+ When a ticket is routed to this session, the supervisor launches **a fresh headless agent process
95
+ for that ticket**, in that account's repo, with a brief the platform generates. That process exits
96
+ when the ticket is done.
97
+
98
+ **One ticket = one process is the point, and it is why nothing here ever needs `/clear`.** Context
99
+ grooming cannot be a discipline: `/clear` and `/new` are commands a *human types into a TUI*, and no
100
+ model can invoke them — so a long-lived session working ticket after ticket has no way to reset
101
+ itself. A process that exits has nothing to reset. Do not try to solve context growth by being tidy
102
+ inside one session; let the session end.
103
+
104
+ **`serve` returns immediately, and that is correct.** Do not follow it with a wait, a poll, or a
105
+ loop. The supervisor is detached precisely so this tab is free. Report that you are serving and
106
+ stop; you have not left the job half done.
107
+
108
+ What the supervisor handles for you, so you do not have to think about any of it: collecting routed
109
+ work, claiming each ticket so no second worker starts on it, renewing that claim while the worker
110
+ runs, reporting the worker's activity to the dashboard, dropping the claim when it exits, retrying a
111
+ failed ticket once, and releasing a ticket back to the queue when it has failed too often. Closing
112
+ the terminal stops everything and hands any in-flight work back.
113
+
114
+ ### The manual loop
115
+
116
+ Only when the session registered with `agent register` rather than `agent serve`.
117
+
118
+ ```bash
119
+ remits-cli agent work --wait 600 # returns as soon as work arrives
120
+ remits-cli agent status --state working --ticket 22454 --activity "reproducing the upload failure"
121
+ # ... investigate, fix, verify — keep `status` current as what you are doing changes ...
122
+ # ... close the ticket lifecycle through remits-cli ticket (accept -> status -> complete, ask, or release) ...
123
+ remits-cli agent status --state idle # then ask for work again
124
+ ```
125
+
126
+ **Nothing wakes this tab.** A routed ticket sits there until someone in this session asks for it, so
127
+ if you are working the manual loop you have to keep asking. If you find yourself wishing you could
128
+ be woken up, that is what `agent serve` is.
129
+
130
+ **Report what you are doing.** `agent status` is how an operator watching the dashboard, or another
131
+ agent, knows this session is alive and what it is on. It costs one command and it is the difference
132
+ between a visible queue and a silent one. Update it when you change what you are doing, not on a
133
+ timer.
134
+
135
+ **Work the ticket in the right repo.** A ticket names its `accountId`, and often an
136
+ `implementationAccountId` — the platform/product account whose repo holds the code. Resolve that to a
137
+ local directory through `~/.remits-cli/account-repos.json` and `cd` there before making changes. You
138
+ registered from anywhere; you do not fix anything from anywhere.
139
+
140
+ **Capacity is real, not advisory.** A serving session tells the platform how many workers it can run
141
+ (`--max-concurrent`, default 1), and the router will not send it more than that. So a session at
142
+ capacity is skipped in favour of one that is free, rather than accumulating tickets it will never
143
+ start.
144
+
145
+ **Lane note:** `register` and `serve` default to the production lane because a support agent works real tickets.
146
+ `--data-mode test` registers a fixture agent instead, which will never be routed a production ticket —
147
+ use it only when you are deliberately testing the routing itself.
148
+
149
+ ### If you are the worker
150
+
151
+ You know you are one when `REMITS_SUPPORT_TICKET_ID` is set in your environment. Your whole job is
152
+ that one ticket, and your brief is your prompt — follow it. Two things it says that are worth
153
+ repeating: **report progress** with `remits-cli agent status --ticket <id> --activity "..."` (a
154
+ headless run is invisible otherwise), and **end in a terminal state** — `complete` with a real
155
+ resolution, or `update_status` with what you established and what the next agent should try. Exiting
156
+ quietly leaves a ticket that looks in-flight forever.
157
+
158
+ ### Before you edit anything: where you are, and whether you may
159
+
160
+ Presence answers *who*. Two more facts answer *whether you may edit*, and with git worktrees and
161
+ workspace lanes an account id no longer identifies a working tree — so an agent that does not ask these
162
+ is assuming, and the assumption it makes when it guesses wrong is "this is my repository to edit".
163
+
164
+ ```bash
165
+ remits-cli ticket where --ticket 22454 # repo account, checkout, branch, staging lane, lease, claim, YOUR phase
166
+ remits-cli agent map # every repository, who holds each lease, and where each agent is working
167
+ ```
168
+
169
+ **Reading, reproducing and investigating are parallel-safe and unrestricted. Editing one account's
170
+ repository is exclusive**, enforced by a lease held per repository account per data lane:
171
+
172
+ ```bash
173
+ remits-cli ticket lease --ticket 22454 # take it when you started read-only and reached an actual edit
174
+ remits-cli ticket unlease --ticket 22454 # complete / ask / release already do this for you
175
+ ```
176
+
177
+ Four things are worth knowing and are not obvious:
178
+
179
+ - **A refusal is not an error.** The run continues read-only, and the message names who holds the lease
180
+ **and where they are working**, so you can tell a real conflict from a holder in a different worktree.
181
+ **Do not wait for a lease and do not poll for one** — investigation is most of the work on most
182
+ tickets, and a ticket that turns out to need an edit hands off with `ticket progress --next-step`
183
+ saying exactly what change is needed. A lock people queue on turns one stuck worker into a stalled
184
+ fleet.
185
+ - **A staging workspace does not replace the lease.** Separate lanes stop two runs *resolving* each
186
+ other's staged code; they do nothing about two processes writing the same files or pushing the same
187
+ branch, which is what actually destroys work.
188
+ - **One editing worker per repository account per lane — worktrees do not change this.** Two worktrees
189
+ of one repo push to the same branch on the same remote, so per-directory leases would trade file
190
+ conflicts for non-fast-forward push conflicts, which surface later and are worse.
191
+ - **A spawned ticket worker already has its own staging lane** (`REMITS_WORKSPACE=ticket-<id>`). You do
192
+ not set it, and your brief states it. Say which lane you staged into when you report what you verified
193
+ — somebody looking at the shared lane will not see your changes.
194
+
195
+ `unlease` returns **your own** lease and deliberately cannot touch anybody else's. Breaking a stale one
196
+ is a separate, human verb: `remits-cli ticket force-unlease --ticket ID --reason "..."`.
197
+
198
+ ### Your account's process is binding, and it is already in your brief
199
+
200
+ An account declares how its work is done as Prompts with purpose `OPERATIONS`, selected per the ticket's
201
+ `workstream` (`support`, `sdlc`, `incident`, `release`, or whatever that organization calls its
202
+ processes) via `Prompt.category`, with an uncategorised one as the catch-all. The matching text is
203
+ **inlined verbatim into the brief** of every agent that works one of that account's tickets and is
204
+ binding on it — follow it even where it differs from how you would normally proceed, and if it conflicts
205
+ with the brief, follow the process and say so on the ticket.
206
+
207
+ You do not fetch it: if the brief has a "The process that governs this ticket" section, that is it. If
208
+ it says the account documents no process, the likely cause is that the prompt is **staged but not
209
+ committed** — an autonomous worker resolves the committed one, so it is never bound by a procedure
210
+ nobody has reviewed. The repo root also carries a generated `OPERATIONS.md`; that file is generated
211
+ from the prompt, so **edit the prompt, not the file**.
212
+
213
+ `workstream` is not `type`: a type classifies the request, a workstream names the procedure, and they
214
+ cross — a `defect` handled by incident response out of hours goes through the SDLC in the morning.
215
+
216
+ ### Seeing the queue as a human does
217
+
218
+ `remits-cli start` opens a browser control center showing the same facts you are acting on: which
219
+ agents are registered and what each is doing, which tickets are open, and which agent each is routed
220
+ to. An operator can route a ticket to a specific agent from there. It is a view, not a second system —
221
+ what it shows and what `remits-cli agent work` returns come from the same records.
222
+
223
+ ### Agent components are workers too
224
+
225
+ An Agent component that works support tickets must register itself in the same presence registry as
226
+ CLI sessions. That makes it visible in `cliAgents()` and `remits-cli agent map`, and lets
227
+ `dispatch(...)` / `ticket route` target it by the same `agentId` field:
228
+
229
+ ```groovy
230
+ def me = workerId() // component:Agent:34
231
+ registerWorker([agentId: me, label: 'Support Triage Agent'])
232
+ def work = workerWork([agentId: me, runId: threadGroupingId()])
233
+ workerStatus([agentId: me, ticketId: work.tickets[0]?.id, activity: 'investigating'])
234
+ ```
235
+
236
+ `workerWork(...)` is the component-side analogue of `remits-cli agent work` — the same code runs behind
237
+ both, so a component and a terminal session get the same answer to the same question. It reads tickets
238
+ routed to that worker, claims what it may start, sweeps eligible unrouted work when it is idle (reported
239
+ separately as `sweptTicketIds`), takes the repository edit lease when available, and returns `phase`,
240
+ `brief`, `briefFacts`, and `editLease`. `phase:'editing'` means the component may edit and stage;
241
+ `phase:'investigating'` means another worker holds the repo lease and the component must stay read-only.
242
+ When a component reaches a resting point, pass `agentId` to `complete(...)`, `ask(...)`, or
243
+ `release(...)` so the ticket verb frees the edit lease and drops its claim.
244
+
245
+ Route to a component by its **id** (`component:Agent:34`) — that is what `workerId()` returns and what
246
+ the component polls on. `ticket route --agent component:Agent:MyAgent` also works: the name is resolved
247
+ and the id form is what gets stored.
248
+
249
+ ### Moving a ticket through its lifecycle
250
+
251
+ **`remits-cli ticket` is the platform surface, and it is what you should reach for.** It reads and
252
+ writes the platform's own ticket record, so the same commands work on every Remits platform.
253
+
254
+ ```bash
255
+ remits-cli ticket read --ticket 22454 # the full record
256
+ remits-cli ticket accept --ticket 22454 # claim it before changing anything
257
+ remits-cli ticket progress --ticket 22454 --summary "Traced it to the posting Action" --category investigation
258
+ remits-cli ticket status --ticket 22454 --status in_progress
259
+ remits-cli ticket complete --ticket 22454 --resolution "What you found and did"
260
+ remits-cli ticket ask --ticket 22454 --question "Re-issue or skip?" --context "412 affected"
261
+ remits-cli ticket answer --ticket 22454 --message "Skip them." # alias: reply
262
+ remits-cli ticket message --ticket 22454 --message "We reproduced it; a fix is staged."
263
+ remits-cli ticket planning --ticket 22454 --board-stage ready_for_qa --blocked-by CAB-112
264
+ remits-cli ticket release --ticket 22454 # hand it back to the queue
265
+ ```
266
+
267
+ **See the whole queue before deciding your ticket is unique.** Alert-raised tickets arrive in
268
+ clusters, and the same root cause routinely appears under several different account names — so the
269
+ first useful question is usually "how many of these are one fix?", and you cannot ask it from a single
270
+ ticket.
271
+
272
+ ```bash
273
+ remits-cli ticket queue --account-id 49 # triage order, most urgent first
274
+ remits-cli ticket queue --account-id 49 --unrouted --status open
275
+ remits-cli ticket queue --account-id 49 --search "firestore index"
276
+ ```
277
+
278
+ The `repo` column is the account the **edit lease** excludes on, so it also tells you which of these
279
+ could be worked at the same time and which will serialize behind one another.
280
+
281
+ Found a second, unrelated problem while working yours? **File it rather than widening the ticket you
282
+ were given** — a ticket that describes two things cannot be closed by either fix.
283
+
284
+ ```bash
285
+ remits-cli ticket create --account-id 49 --subject "Vendor name lost on re-normalization" \
286
+ --type defect --priority medium --reference-id vendor-name-lost
287
+ ```
288
+
289
+ `--reference-id` makes it get-or-create, so a re-run — or another worker reaching the same conclusion
290
+ — reconciles onto the same ticket instead of filing a duplicate.
291
+
292
+ Marking and evidence, so the next reader does not repeat your work:
293
+
294
+ ```bash
295
+ remits-cli ticket tag --ticket 22454 --tags firestore-index,cluster-aug
296
+ remits-cli ticket artifact --ticket 22454 --type log --label "Failing query" --content "..."
297
+ remits-cli ticket note --ticket 22454 --message "Internal: same cause as 23465"
298
+ remits-cli ticket field --ticket 22454 --key customerReference --value CR-9182
299
+ ```
300
+
301
+ `note` is internal — the next worker and a reviewer see it, the requester does not. Use `message` for
302
+ anything the requester should read. `field` writes your **organization's own** field, outside the
303
+ platform's bounded planning slots.
304
+
305
+ **`message`, `answer` and `note` are separated by who reads the result, not by tone.** They are easy to
306
+ confuse and they do different things to the ticket:
307
+
308
+ | Command | Who reads it | What it does to the ticket |
309
+ |---|---|---|
310
+ | `ticket message --message "..."` | the requester | appends a public entry. **Does not send anything** |
311
+ | `ticket note --message "..."` | the next worker, a reviewer | appends an internal entry |
312
+ | `ticket answer --message "..."` | whoever asked | answers the open question and **re-routes**, so a fresh worker resumes |
313
+
314
+ `ticket reply` is an **alias of `answer`** — kept because every existing brief says it. Do not read it as
315
+ "reply to the customer"; that is `message`. (Some account tools spell the conversational verb `reply`,
316
+ which is exactly why this surface teaches `answer`.)
317
+
318
+ **Appending a message is not emailing anyone.** The record is deliberately vendor-agnostic; the account
319
+ owns the transport. To ask for something to actually go out:
320
+
321
+ ```bash
322
+ remits-cli ticket deliver --ticket 22454 --channel email --to ops@acme.test \
323
+ --subject "Waiting on you" --idempotency-key waiting-22454
324
+ remits-cli ticket deliveries --ticket 22454 # what is queued, and what already went
325
+ ```
326
+
327
+ That records an outbox request. An account-owned Rule or scheduled Action sends it and marks the result
328
+ (`ticket delivered --delivery-id ...` / `ticket delivery-failed --delivery-id ... --reason "..."`). Use
329
+ this when a ticket is waiting on somebody who is not watching the queue.
330
+
331
+ ### When a worker needs a decision from a human
332
+
333
+ There are **three** ways a run can end, not two, and the third is what stops an agent guessing:
334
+
335
+ | Outcome | Verb | Means |
336
+ |---|---|---|
337
+ | Done | `ticket complete --resolution` | finished, here is what I did |
338
+ | **Blocked on a choice** | `ticket ask --question` | I understand the work; the DECISION is not mine |
339
+
340
+ | Could not finish | `ticket status --status pending_review` + `ticket progress` | stuck, here is the hand-off |
341
+
342
+ `ask` puts the question on the ticket as an outbound message and parks it. The queue then shows it as
343
+ **awaiting a response** — distinct from "done, review me", which `pending_review` alone cannot express.
344
+
345
+ Someone answers with `remits-cli ticket answer --message "..."` (`reply` is an alias), or through an
346
+ operator Embeddable.
347
+ That **re-routes the ticket by default**, so the supervisor's next poll starts a **fresh worker whose
348
+ brief already contains the exchange**. Resumption costs nothing because a worker is one process per
349
+ ticket — there is no session to restore.
350
+
351
+ Ask when the choice is genuinely not yours: which of two behaviours is wanted, whether to touch live
352
+ data, an ambiguity in the request. Do not ask for reassurance about something you can determine
353
+ yourself, and ask **once**, with the options laid out.
354
+
355
+ Inside an autonomous worker `--ticket` defaults to the ticket that worker was launched for.
356
+
357
+ A refusal is an **answer**: "accept it first", "already assigned to someone else", "reopen it
358
+ first". Act on the sentence — do not go looking for another tool that lets you skip the step.
359
+
360
+ **`mcp_support_ticket` is an account convention, not a platform guarantee.** It belongs to one
361
+ particular System Account: it may not exist on the platform you are on, it scopes to its own
362
+ account, and it will refuse a ticket owned by a different one. Use it only when you are working in
363
+ that account and want its product-specific workflow on top of the record. It offers:
364
+ - `read` — always start here to load current state
365
+ - `accept` — claim the ticket so other agents do not work it concurrently
366
+ - `update_status` — move to `in_progress` or `pending_review`
367
+ - `complete` — resolve the ticket with a summary of what was done
368
+ - `release` — unassign if you cannot continue
369
+ - **Pulling a ticket by number is enough — you do NOT need to know its account first.** The `ticketId` IS the ticket's globally-unique anchor id, and `mcp_support_ticket` resolves the owning account from it for every action that operates on an existing ticket (`read`, `accept`, `update_status`, `complete`, `release`, `record_progress`, `add_artifact`, `get_attachment`). So when a user says "pull ticket 19463", just call `read` with `{ "action": "read", "ticketId": "19463" }` from any authenticated **prod** session (the ticket lives in prod data) — omit `accountId` entirely. The response returns the resolved `accountId`/`accountName` (and `implementationAccountId` when set); use those for any follow-up work. **Only `create` requires an explicit `accountId`** (a brand-new ticket has no anchor to resolve from). If a bare `ticketId` returns "Support ticket not found", double-check you are in `--data-mode prod`, then fall back to passing an explicit `accountId`.
370
+ - `read` can return `documentState:'missing_or_empty'` with `ticket.mirrorOnly:true` and `ticket.canMutate:false`. That means the support-ticket anchor exists and the queue row is real, but the backing Firestore `support_tickets/{ticketId}` document is missing or metadata-only. Treat this as a degraded ticket, not as "ticket not found"; run or request the System Account action `Restore Support Ticket Documents From Anchor Mirrors` for the owning account before lifecycle mutations. The restore recovers mirrored scalar fields, but document-only arrays such as alert snapshots, email threads, attachments, worklog entries, and artifacts may be lost.
371
+ - If a ticket is part of the request, manage the lifecycle proactively. Do not wait for the human user to remind you to read, accept, update, complete, or release it.
372
+
373
+ Ticket-routing context:
374
+ - `accountId` / `accountName` identify the account that owns the ticket.
375
+ - If present, `implementationAccountId` / `implementationAccountName` identify the owning `PLATFORM` or `PRODUCT` implementation context.
376
+ - Prefer `implementationAccountId` when choosing the working directory for a ticket, falling back to `accountId` when the platform/product repo is not available locally. Routing already uses that same ladder, so the repo you were routed through is usually the right one.
377
+ - For enhancement work, prefer the platform/product implementation context when deciding where code changes belong.
378
+ - For defect investigations, start from the owning ticket account, then move to the platform/product context if the root cause is in shared components.
379
+
380
+ Sandbox note:
381
+ - Every `remits-cli` command that reaches the Remits service (`auth`, `tools`, `tool`, `components`, `test`, `token`, and all of `ticket` / `agent`) needs outbound network access.
382
+ - **Inside an autonomous ticket worker this is already arranged** — the supervisor launches the worker with network enabled, in both the editing and the investigating phase. An investigating worker is restricted from *writing the repository*, never from *talking to the platform*: reading its ticket, running a diagnostic and recording a finding are the whole of what investigating means.
383
+ - Elsewhere, `ENOTFOUND`, `EAI_AGAIN`, `ECONNREFUSED`, `EPERM` or a bare exit 6 / `HTTP:000` from a sandboxed session means the command needs escalated permissions or must run outside the sandbox. Do not read it as "the platform is down" — check a second command before concluding anything about the service.
384
+
385
+ Use the same repo/context rules as other tools:
386
+ - If the ticket targets a `CLIENT` account, do not assume that client's repo is the implementation repo. Confirm the parent `PLATFORM` / `PRODUCT` relationship first.
387
+ - Read `~/.remits-cli/account-repos.json` before choosing which local repo to open.
388
+ - If the correct repo for the relevant account type exists locally, switch there and inspect `account-info.json` and `/components`.
389
+ - If the repo is not available locally, use `mcp_account_view`, `mcp_component_view`, and `mcp_component_grep`.