@mmerterden/multi-agent-pipeline 16.25.0 → 16.27.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +77 -0
- package/README.md +1 -1
- package/README.tr.md +1 -1
- package/install/templates/claude-hooks.json +32 -1
- package/package.json +1 -1
- package/pipeline/commands/multi-agent/help/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/refactor/SKILL.md +23 -1
- package/pipeline/commands/multi-agent/search/SKILL.md +28 -0
- package/pipeline/commands/multi-agent/setup/SKILL.md +18 -44
- package/pipeline/commands/multi-agent/status/SKILL.md +9 -0
- package/pipeline/lib/credential-inventory.sh +142 -18
- package/pipeline/lib/fetch-crashlytics.sh +123 -28
- package/pipeline/multi-agent-refs/features/url-enrichment.md +1 -1
- package/pipeline/multi-agent-refs/features/visual-evidence.md +5 -0
- package/pipeline/multi-agent-refs/keychain.md +65 -20
- package/pipeline/multi-agent-refs/knowledge.md +27 -0
- package/pipeline/multi-agent-refs/phases/operations.md +7 -1
- package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +3 -1
- package/pipeline/multi-agent-refs/phases/phase-4-review.md +1 -8
- package/pipeline/multi-agent-refs/phases/phase-7-report.md +11 -21
- package/pipeline/multi-agent-refs/picker-contract.md +1 -1
- package/pipeline/multi-agent-refs/refactor/observations.md +81 -0
- package/pipeline/multi-agent-refs/setup/firebase.md +151 -0
- package/pipeline/schemas/learnings-ledger.schema.json +5 -0
- package/pipeline/schemas/prefs.schema.json +31 -3
- package/pipeline/schemas/skill-observation.schema.json +73 -0
- package/pipeline/scripts/capture-flush.sh +158 -0
- package/pipeline/scripts/capture-resume.sh +87 -0
- package/pipeline/scripts/crush-json.mjs +283 -0
- package/pipeline/scripts/firebase-app-discovery.sh +114 -0
- package/pipeline/scripts/keychain-save.sh +5 -8
- package/pipeline/scripts/keychain.py +76 -14
- package/pipeline/scripts/learn-from-transcripts.mjs +625 -0
- package/pipeline/scripts/learning-curve.mjs +22 -4
- package/pipeline/scripts/learnings-ledger.mjs +86 -12
- package/pipeline/scripts/note-session.sh +187 -0
- package/pipeline/scripts/observations.mjs +347 -0
- package/pipeline/scripts/offload-ref.sh +45 -2
- package/pipeline/scripts/pre-commit-check.sh +31 -1
- package/pipeline/scripts/scan-agent-config.sh +12 -3
- package/pipeline/scripts/skill-siblings.mjs +187 -0
- package/pipeline/scripts/triage-memory.mjs +73 -9
- package/pipeline/skills/.skill-manifest.json +1 -1
- package/pipeline/skills/shared/core/multi-agent-refactor/SKILL.md +23 -1
- package/pipeline/skills/shared/core/multi-agent-search/SKILL.md +28 -0
- package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +37 -6
- package/pipeline/skills/shared/core/multi-agent-status/SKILL.md +9 -0
|
@@ -150,7 +150,13 @@ This keeps orchestrator context lean and enables programmatic routing.
|
|
|
150
150
|
|
|
151
151
|
**Proactive compaction + phase-boundary checkpoint**: the orchestrator follows ~2,500 lines of phase prose in one session; once context fills, it starts dropping steps - the single biggest cause of "it got stuck / skipped a step." Two defenses, both required on full-pipeline runs:
|
|
152
152
|
|
|
153
|
-
- *Phase-boundary checkpoint.* At every phase transition, before loading the next phase doc, write the durable state (`agent-state.json` phase status + `files[]` + `retryCount`) and append a structured handoff block to `agent-log.md` (format below). The next phase reads state + log, not the back-conversation - so a transition is a clean re-entry point even if context is later compacted.
|
|
153
|
+
- *Phase-boundary checkpoint.* At every phase transition, before loading the next phase doc, write the durable state (`agent-state.json` phase status + `files[]` + `retryCount`) and append a structured handoff block to `agent-log.md` (format below). The next phase reads state + log, not the back-conversation - so a transition is a clean re-entry point even if context is later compacted. Then flush what the run has learned so far:
|
|
154
|
+
|
|
155
|
+
```bash
|
|
156
|
+
bash $HOME/.claude/scripts/capture-flush.sh --state "$STATE_FILE" --quiet
|
|
157
|
+
```
|
|
158
|
+
|
|
159
|
+
The durable stores used to be written only in Phase 7, which is the phase a run is LEAST likely to reach: a run killed in Phase 3 threw away every finding it had established, and the next run on the same repo paid to rediscover it. The flush is idempotent (the second call through writes 0 rows), costs no API tokens, calls no model, and never fails the transition. Phase 7 is now the LAST flush rather than the only one. The trigger is deliberately mechanical - a phase transition, not a model noticing that a moment qualifies.
|
|
154
160
|
- *Compaction trigger.* If conversation context exceeds ~50%, run `/compact` preserving "modified files, plan, open review findings, current phase + sub-step" before continuing. Don't wait for auto-compaction near the limit - it triggers exactly when context is worst and is lossy. After compaction, re-read `agent-state.json` AND the latest `## Handoff` block in `agent-log.md` to re-ground.
|
|
155
161
|
|
|
156
162
|
**Handoff block (v10.8.0)**: the structured artifact the phase-boundary checkpoint appends to `agent-log.md`. Written by the orchestrator from state it already holds - no agent dispatch, no extra LLM call. Cap at ~15 lines; the latest block is authoritative (earlier ones are history). This is the fresh-context re-entry contract: a resume or post-compaction session rebuilds working context from the latest handoff + `agent-state.json` + git log, never from conversation memory.
|
|
@@ -145,7 +145,9 @@ Gated by `prefs.global.repoMap.enabled` (default: `false`). When enabled, runs `
|
|
|
145
145
|
|
|
146
146
|
#### Step 2.6 - Code Graph Injection (advisory, opt-in)
|
|
147
147
|
|
|
148
|
-
Gated by `prefs.global.codeGraph.enabled` (default: `false`). With a rule file for `detectedStack`, Phase 1 refreshes the
|
|
148
|
+
Gated by `prefs.global.codeGraph.enabled` (default: `false`). With a rule file for `detectedStack`, Phase 1 refreshes the graph and queries it; `graph-affected` feeds `analysis.touchedAreas[]`. Zero API cost, read-only. Commands and measurements: `$HOME/.claude/multi-agent-refs/features/code-graph.md`.
|
|
149
|
+
|
|
150
|
+
**A valid result REPLACES the opening sweep** rather than sitting beside it: Explore starts from those files and walks outward, with no broad `Glob`/`Grep` pass first - running both pays twice, and the second re-derives what the graph said. No graph (missing rule, stale, or disabled) leaves the previous behaviour untouched; a task naming a symbol or path is still grepped directly.
|
|
149
151
|
|
|
150
152
|
#### Step 3 - Codebase Exploration
|
|
151
153
|
|
|
@@ -262,14 +262,7 @@ Visual-fidelity mismatches against the captured screenshot are BLOCKING findings
|
|
|
262
262
|
|
|
263
263
|
- Canonical component usage: the Code Connect-mapped component is used verbatim - a sound-alike substitute, a forked copy, or ad-hoc inline UI where a mapping exists is blocking (Phase 3 "Design fidelity contract")
|
|
264
264
|
- Inter-component spacing: gaps, paddings, and alignment BETWEEN components match the design's measured values mapped to spacing tokens - invented numeric values are blocking
|
|
265
|
-
-
|
|
266
|
-
- Field grouping (one rounded box vs two; separator vs gap)
|
|
267
|
-
- Character counter visibility
|
|
268
|
-
- Header style (size, weight, alignment)
|
|
269
|
-
- Inline error layout (icon presence, text colour, position relative to the input)
|
|
270
|
-
- Button height
|
|
271
|
-
- Indicator chip placement
|
|
272
|
-
- Placeholder copy and position
|
|
265
|
+
- Everything else: the 30-group catalog in `$HOME/.claude/multi-agent-refs/features/design-conformance.md`. Two of its rules carry into review: a height or inset is **measured**, never read from the token under test; an unrun check is a fail-to-verify, not a pass.
|
|
273
266
|
|
|
274
267
|
When `state.figmaAccess.tier === 3` (user-attached screenshot, no Code Connect snippet), the reviewer additionally sets `findings[i].severity = "blocking"` and `findings[i].tag = "review_blocking_tier3"` on every UI atom that lacks a confirmed canonical-component mapping. The triage step preserves these findings unless the user has explicitly cleared the open question.
|
|
275
268
|
|
|
@@ -219,38 +219,28 @@ This is independent of the channels-side `reportContent.costSummary` (which gate
|
|
|
219
219
|
|
|
220
220
|
**Missing-telemetry disclosure (required).** Name untracked phase ids as cost-unavailable in the report and closing summary. Mechanic: `payload-contracts.md`.
|
|
221
221
|
|
|
222
|
-
**
|
|
222
|
+
**Final capture flush (mandatory):** persist the triage findings and the durable learnings this run established.
|
|
223
223
|
|
|
224
224
|
```bash
|
|
225
|
-
|
|
226
|
-
# reader is `[ -f ]`-guarded, so a wrong path degrades SILENTLY.
|
|
227
|
-
TRIAGE_PATH="$(jq -r '.artifactsPath // empty' "$STATE_FILE" 2>/dev/null)/triage-output.json"
|
|
228
|
-
[ -f "$TRIAGE_PATH" ] || TRIAGE_PATH="$WORKTREE/triage-output.json"
|
|
229
|
-
if [ -f "$TRIAGE_PATH" ]; then
|
|
230
|
-
node $HOME/.claude/scripts/triage-memory.mjs ingest \
|
|
231
|
-
--triage "$TRIAGE_PATH" \
|
|
232
|
-
--task-id "$TASK_ID" \
|
|
233
|
-
--task-title "$TASK_TITLE" \
|
|
234
|
-
--stack "$DETECTED_STACK" >/dev/null 2>&1 || true
|
|
235
|
-
fi
|
|
225
|
+
bash $HOME/.claude/scripts/capture-flush.sh --state "$STATE_FILE"
|
|
236
226
|
```
|
|
237
227
|
|
|
238
|
-
|
|
228
|
+
Every phase boundary already made this call (`operations.md`), so by Phase 7 it usually writes 0 rows - and that is the point. These writes used to live ONLY here, in the phase a run is least likely to reach: a run killed in Phase 3 lost every finding it had established. Phase 7 is now the last flush, not the only one.
|
|
239
229
|
|
|
240
|
-
|
|
230
|
+
What it does, so this doc stays inspectable:
|
|
231
|
+
|
|
232
|
+
- **Triage ingest** (`triage-memory.mjs ingest`). accepted/deferred/rejected rows into the per-repo corpus, so Phase 1 enrichment and Phase 4 prior-art recall them later. Idempotent. Reads the salvaged copy under `artifactsPath` FIRST: Phase 6 removes the worktree, and a worktree-first reader degrades silently for exactly the runs worth rescuing. JSONL at `~/.claude/memory/multi-agent/<repo-slug>/triage-corpus.jsonl`, per-repo, never cross-leaking. Off when `prefs.global.priorArtEnrichment.ingestOnComplete = false`.
|
|
233
|
+
- **Ledger distill** (`learnings-ledger.mjs from-triage`). Rejected findings become durable `rejected-preference` entries so reviewers stop re-raising them. Idempotent. Safety: blocking-severity rejections are NOT distilled - a wrong rejection must never permanently suppress that class; `learnings-ledger.mjs forget` clears a stale one. JSONL beside the corpus; its brief replays into Phase 1 and Phase 4. On by default via `prefs.global.learningsLedger.enabled`.
|
|
234
|
+
|
|
235
|
+
No model runs in the flush - it derives everything from `triage-output.json` plus `agent-state.json`, which is what lets a hook call it. The model-dependent parts of this phase (Step 3 knowledge capture, memory synthesis) stay here, because a hook cannot think.
|
|
236
|
+
|
|
237
|
+
A durable architectural fact is a judgement call, so it stays here rather than in the flush:
|
|
241
238
|
|
|
242
239
|
```bash
|
|
243
|
-
if [ -f "$TRIAGE_PATH" ]; then
|
|
244
|
-
node $HOME/.claude/scripts/learnings-ledger.mjs from-triage \
|
|
245
|
-
--triage "$TRIAGE_PATH" --task "$TASK_ID" >/dev/null 2>&1 || true
|
|
246
|
-
fi
|
|
247
|
-
# Optionally capture a durable architectural fact the analysis established:
|
|
248
240
|
# node $HOME/.claude/scripts/learnings-ledger.mjs add --kind fact \
|
|
249
241
|
# --statement "<one-line fact>" --scope "<path-glob>" --task "$TASK_ID" >/dev/null 2>&1 || true
|
|
250
242
|
```
|
|
251
243
|
|
|
252
|
-
The ledger is JSONL at `~/.claude/memory/multi-agent/<repo-slug>/learnings-ledger.jsonl`, next to the triage corpus, per-repo isolated. Its brief is replayed into Phase 1 analysis and Phase 4 triage on future runs.
|
|
253
|
-
|
|
254
244
|
Print the closing report to the terminal. Two blocks, in this order - what the pipeline spent, then what it changed:
|
|
255
245
|
|
|
256
246
|
```bash
|
|
@@ -81,6 +81,6 @@ In autopilot, `ask_choice` resolves to `default` (or the safe first option) with
|
|
|
81
81
|
|
|
82
82
|
## Deterministic gates note
|
|
83
83
|
|
|
84
|
-
Claude Code's `PreToolUse` exit-2 hooks are the HARD blocking gates. Three ship, none needing run-specific arguments so they are naturally hookable: (1) `pre-commit-check.sh` scans the staged diff on every `git commit` and blocks on a detected secret; (2) `agent-guard.sh` runs on `git commit` + `git push` and blocks AI/assistant attribution in a commit message and force-push to a protected branch (main/master/develop); (3) `check-read-size.sh` runs on `Read` and on the shell commands that read a file whole, and routes an oversized read to a cheap worker (`bulk-read.sh`) instead of the caller's own rung. The first two inspect what a run WRITES; the third inspects what it pays to READ, and it is inert until `bulkRead.mode` is set to `observe` or `enforce`, so merging the block changes nothing until the user opts in. Its `observe` mode blocks nothing and only logs, which is how the baseline is measured before anything is routed. All three are self-contained, fail-open on internal error, and never execute the inspected command. The recommended hook block ships at `install/templates/claude-hooks.json`; `multi-agent:setup` offers to merge it into `~/.claude/settings.json`. The other deterministic gates (evidence, consensus, intent, learnings) are invoked by the pipeline phases with per-run arguments (a build-log path, the triage JSON, the free-text input), so they are phase-enforced by contract, not OS-hookable.
|
|
84
|
+
Claude Code's `PreToolUse` exit-2 hooks are the HARD blocking gates. Three ship, none needing run-specific arguments so they are naturally hookable: (1) `pre-commit-check.sh` scans the staged diff on every `git commit` and blocks on a detected secret; (2) `agent-guard.sh` runs on `git commit` + `git push` and blocks AI/assistant attribution in a commit message and force-push to a protected branch (main/master/develop); (3) `check-read-size.sh` runs on `Read` and on the shell commands that read a file whole, and routes an oversized read to a cheap worker (`bulk-read.sh`) instead of the caller's own rung. The first two inspect what a run WRITES; the third inspects what it pays to READ, and it is inert until `bulkRead.mode` is set to `observe` or `enforce`, so merging the block changes nothing until the user opts in. Its `observe` mode blocks nothing and only logs, which is how the baseline is measured before anything is routed. All three are self-contained, fail-open on internal error, and never execute the inspected command. Two capture hooks ship in the same block and block nothing: `SessionEnd` runs `capture-flush.sh --if-stale` (writing a killed run's findings into the per-repo stores, since every durable write used to live in Phase 7 - the phase a run is least likely to reach) plus `note-session.sh` (the mechanical shape of a non-pipeline session: tools used, commands that failed, calls the user refused - never an argument, never any output), and `SessionStart` runs `capture-resume.sh`, at most two lines about an unfinished run and a stale observation queue. Neither calls a model; both exit 0 on every path. The recommended hook block ships at `install/templates/claude-hooks.json`; `multi-agent:setup` offers to merge it into `~/.claude/settings.json`. The other deterministic gates (evidence, consensus, intent, learnings) are invoked by the pipeline phases with per-run arguments (a build-log path, the triage JSON, the free-text input), so they are phase-enforced by contract, not OS-hookable.
|
|
85
85
|
|
|
86
86
|
Copilot CLI has no `PreToolUse` equivalent, so the secret scan there is workflow-enforced (run as a phase step, not OS-blocked) plus a CI smoke-gate step.
|
|
@@ -0,0 +1,81 @@
|
|
|
1
|
+
# The observation backlog (refactor Step 0e / `backlog` mode)
|
|
2
|
+
|
|
3
|
+
Loaded on demand by `/multi-agent:refactor`. The SKILL.md carries the band intro; this file is the flow.
|
|
4
|
+
|
|
5
|
+
## 1. Why a queue exists at all
|
|
6
|
+
|
|
7
|
+
`/multi-agent:refactor` derives its findings from scratch on every invocation. That is the right design for a sweep, and it has one consequence nobody chose: a friction noticed on Tuesday is gone by Wednesday unless it was fixed within the hour. Most are not, because they surface mid-task, when stopping to fix them is the wrong call.
|
|
8
|
+
|
|
9
|
+
So the friction is written down the moment it is noticed, and this band decides what to do with it later. The two halves are deliberately separate: noticing is cheap and must never wait for a decision, deciding is expensive and must never happen mid-task.
|
|
10
|
+
|
|
11
|
+
Store: `$HOME/.claude/memory/multi-agent/_pipeline/observations/NNNN-slug.md`, one file per observation, frontmatter per `schemas/skill-observation.schema.json`. The directory listing is the index - there is no index file, because an index is a second copy of the truth and the second copy is the one that goes stale.
|
|
12
|
+
|
|
13
|
+
## 2. Reading the queue
|
|
14
|
+
|
|
15
|
+
```bash
|
|
16
|
+
node "$HOME/.claude/scripts/observations.mjs" scan --status open --json
|
|
17
|
+
```
|
|
18
|
+
|
|
19
|
+
`scan` reads frontmatter only and never opens a body, so a queue of several hundred costs almost nothing.
|
|
20
|
+
|
|
21
|
+
**Exit 3 is `SCAN BROKEN` and it is not an empty queue.** It means files exist on disk and none parsed - the reader is broken. Halt the band and say so. An empty scan otherwise reports two different facts with one answer ("nothing to find" and "the question was never asked"), and only the first is a finding; the guard exists so those two can never be confused again.
|
|
22
|
+
|
|
23
|
+
## 3. Splitting the queue
|
|
24
|
+
|
|
25
|
+
Every open observation lands in exactly one of three buckets:
|
|
26
|
+
|
|
27
|
+
| Bucket | Meaning | What happens |
|
|
28
|
+
|---|---|---|
|
|
29
|
+
| Actionable | a change worth making now | goes into the plan as a normal band item with its file and fix |
|
|
30
|
+
| Sibling propagation | already fixed in one copy, not the others | goes into the plan as a propagation item, with the exact surfaces named |
|
|
31
|
+
| To decline | not worth doing, or wrong | proposed for `declined` with a one-line reason |
|
|
32
|
+
|
|
33
|
+
The second bucket is the one this tree specifically needs. A command is never one file: it is authored under `commands/`, mirrored into `skills/shared/core/`, and installed again into the Claude, Copilot and Codex trees. A fix applied to whichever copy was open drifts from the rest silently - nothing errors and every gate stays green. Resolve the surfaces mechanically, never from memory:
|
|
34
|
+
|
|
35
|
+
```bash
|
|
36
|
+
node "$HOME/.claude/scripts/skill-siblings.mjs" <path> --json
|
|
37
|
+
node "$HOME/.claude/scripts/skill-siblings.mjs" --audit # every command at once
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
## 4. Deciding, and the deferral that wears a disguise
|
|
41
|
+
|
|
42
|
+
An observation leaves the queue with a status, never by being ignored:
|
|
43
|
+
|
|
44
|
+
- `actioned` - a change shipped. Record the commit in `reference`.
|
|
45
|
+
- `declined` - decided against. A decline with no `resolution` is indistinguishable from neglect, so the reason is written.
|
|
46
|
+
- `superseded` - another observation covers it. Name which.
|
|
47
|
+
- `parked` - decided, but blocked on something outside this repo.
|
|
48
|
+
|
|
49
|
+
`parked` requires `parked_until` naming the concrete event that unblocks it: a version, a release, an upstream fix. "Let us gather more data" is not an event. If no observation could change the decision and no date is nameable, the honest status is `declined` - a park with no expiry leaves the queue and never comes back, which is a silent decline wearing a friendlier word.
|
|
50
|
+
|
|
51
|
+
```bash
|
|
52
|
+
node "$HOME/.claude/scripts/observations.mjs" resolve --id 0007 --status actioned \
|
|
53
|
+
--resolution "backlog mode added" --reference "<sha>"
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
## 5. Applying
|
|
57
|
+
|
|
58
|
+
Changes go to a staged copy and are shown before anything is installed, exactly as the drift band does. The user installs; this band never edits an installed skill in place.
|
|
59
|
+
|
|
60
|
+
When any observation is resolved, stamp the review:
|
|
61
|
+
|
|
62
|
+
```bash
|
|
63
|
+
date +%Y-%m-%d > "$HOME/.claude/memory/multi-agent/_pipeline/last-review-date.txt"
|
|
64
|
+
```
|
|
65
|
+
|
|
66
|
+
`capture-resume.sh` reads that stamp at session start and offers one line when it is seven days old and the queue is non-empty. One line, never a block - the user's own work does not wait on the pipeline's housekeeping.
|
|
67
|
+
|
|
68
|
+
## 6. Writing an observation
|
|
69
|
+
|
|
70
|
+
Any session may add one, in the same turn the friction appears, and then carry on:
|
|
71
|
+
|
|
72
|
+
```bash
|
|
73
|
+
node "$HOME/.claude/scripts/observations.mjs" add \
|
|
74
|
+
--title "<the friction, not the fix>" \
|
|
75
|
+
--target <repo-relative path> [--target ...] \
|
|
76
|
+
--area <phase|gates|docs|...> \
|
|
77
|
+
--session "<what the session was doing>" \
|
|
78
|
+
--body "<detail>"
|
|
79
|
+
```
|
|
80
|
+
|
|
81
|
+
State the friction, not the remedy: "refactor re-derives its findings every run" is an observation, "add a backlog mode" is a proposal, and proposals age badly while observations do not. `siblings_checked` is filled by `skill-siblings.mjs` automatically and cannot be empty - the whole point is that the answer is computed rather than recalled, because recollection is the faculty that produced the drift.
|
|
@@ -0,0 +1,151 @@
|
|
|
1
|
+
# Firebase / Crashlytics Onboarding (setup Step 3c)
|
|
2
|
+
|
|
3
|
+
Loaded on demand by `/multi-agent:setup` Step 3c (optional, any platform). The SKILL.md carries the step intro; this file is the full flow.
|
|
4
|
+
|
|
5
|
+
Runs inside Step 3 alongside the other missing credentials. A user who already
|
|
6
|
+
holds a Firebase service-account JSON gets it mapped by Step 1 discovery like any
|
|
7
|
+
other credential; this flow covers what discovery cannot see - whether that
|
|
8
|
+
account may actually read Crashlytics, and what to do when it may not.
|
|
9
|
+
|
|
10
|
+
## 1. Why there are three ways in
|
|
11
|
+
|
|
12
|
+
Crashlytics has one API and three ways to authenticate against it. They fail
|
|
13
|
+
differently, so the pipeline measures rather than assumes.
|
|
14
|
+
|
|
15
|
+
| Tier | Path | Serves | Needs |
|
|
16
|
+
|---|---|---|---|
|
|
17
|
+
| 1 | Service account + `lib/fetch-crashlytics.sh` | headless: autopilot, cron, url-enrichment, review chains | a Crashlytics role on the service account |
|
|
18
|
+
| 2 | Firebase MCP + an interactive `firebase login` | a person at the keyboard | a browser, a real terminal, a session that expires |
|
|
19
|
+
| 3 | The user pastes the stack trace | last resort | nothing, and it produces degraded evidence |
|
|
20
|
+
|
|
21
|
+
Order is not fixed - whichever tier is ready wins, exactly as in the Figma chain.
|
|
22
|
+
But **tier 1 is preferred whenever both are ready**, and the reason is structural:
|
|
23
|
+
url-enrichment expands a Crashlytics link found in a Jira ticket while nobody is
|
|
24
|
+
watching, and autopilot runs with no terminal to log into. A tier that needs a
|
|
25
|
+
browser cannot serve those. Tier 2 fills the gap while a role request is pending;
|
|
26
|
+
it does not replace tier 1.
|
|
27
|
+
|
|
28
|
+
`credential-inventory.sh --probe` reports which tier is live:
|
|
29
|
+
`tier-1-ready` / `tier-2-ready` / `tier-1-no-grant` / `malformed` / `unreachable`.
|
|
30
|
+
The vocabulary and what to tell the user for each: `refs/keychain.md`.
|
|
31
|
+
|
|
32
|
+
## 2. Tier 1 - the service account
|
|
33
|
+
|
|
34
|
+
One credential, `keychainMapping.firebase`, holding the JSON exactly as Firebase
|
|
35
|
+
Console issued it. Firebase Console -> Project settings -> Service accounts ->
|
|
36
|
+
Generate new private key. Store it through the normal Token Save Flow; it is
|
|
37
|
+
multi-line and that is fine - the store round-trips it byte for byte.
|
|
38
|
+
|
|
39
|
+
Do **not** base64 it. Encoding is the store's internal business, and a wrapper on
|
|
40
|
+
the way in that nothing unwraps on the way out is how this path stayed broken.
|
|
41
|
+
|
|
42
|
+
Several Firebase projects per team is the normal case (legacy plus redesign,
|
|
43
|
+
staging plus prod). The single slot is the fallback; extra projects go in
|
|
44
|
+
`prefs.global.firebase.accounts[]` as `{projectId, keychainKey, label}` and the
|
|
45
|
+
URL's project id picks the key.
|
|
46
|
+
|
|
47
|
+
**The role is the part discovery cannot see.** A freshly generated service account
|
|
48
|
+
authenticates perfectly and reads nothing: `firebase-adminsdk-*` carries no
|
|
49
|
+
Crashlytics permission by default. The probe surfaces this as `tier-1-no-grant`,
|
|
50
|
+
and the fix is one line to whoever holds Owner on the project:
|
|
51
|
+
|
|
52
|
+
> Please grant `roles/firebasecrashlytics.viewer` on project `<projectId>` to the
|
|
53
|
+
> service account `<client_email>`. It is read-only: it allows reading crash
|
|
54
|
+
> issues and their stack traces, and nothing else.
|
|
55
|
+
|
|
56
|
+
`roles/firebase.viewer` also works and is broader. Ask for the narrower one first.
|
|
57
|
+
|
|
58
|
+
Verify with `bash lib/fetch-crashlytics.sh --probe`. It asks IAM what the account
|
|
59
|
+
may do rather than calling Crashlytics and reading the error, because a 403 from
|
|
60
|
+
the data endpoint means "no permission" but so does a 403 from a disabled API,
|
|
61
|
+
and those need different fixes.
|
|
62
|
+
|
|
63
|
+
## 3. Tier 2 - interactive session plus MCP
|
|
64
|
+
|
|
65
|
+
Two halves, both required. Either alone reaches nothing.
|
|
66
|
+
|
|
67
|
+
**The session.** `firebase login` opens a browser and does not work from inside an
|
|
68
|
+
agent harness - it needs a real terminal the user drives themselves. Ask them to
|
|
69
|
+
run it and say when it is done; `firebase login:list` confirms.
|
|
70
|
+
|
|
71
|
+
**The MCP server.** The Firebase CLI serves it itself - there is no package to
|
|
72
|
+
install beyond the CLI the login already needed:
|
|
73
|
+
|
|
74
|
+
```bash
|
|
75
|
+
claude mcp add firebase -- firebase mcp --dir "$PWD"
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
`--dir` is not decoration: the server resolves the project from that directory's
|
|
79
|
+
`.firebaserc` / `firebase.json`, so a server registered without it answers for
|
|
80
|
+
whatever directory it happens to start in.
|
|
81
|
+
|
|
82
|
+
Scope is a real choice with three answers, so ask rather than guess:
|
|
83
|
+
|
|
84
|
+
| Scope | Where it lands | When |
|
|
85
|
+
|---|---|---|
|
|
86
|
+
| `local` (default) | this user's entry for this repo, in `~/.claude.json` | the normal answer - the grant stays scoped to the repo that needs it |
|
|
87
|
+
| `user` | this user, every repo | only when the user works across several Firebase projects and says so; `--dir` then has to be re-pointed per repo |
|
|
88
|
+
| `project` | `.mcp.json` **committed in the repo** | only on an explicit ask - this registers the server for everyone who clones it |
|
|
89
|
+
|
|
90
|
+
Default to `local` and never reach for `project` on your own: the server can read
|
|
91
|
+
that Firebase project's data, and committing it hands that reach to the whole
|
|
92
|
+
team as a side effect of one person's setup.
|
|
93
|
+
|
|
94
|
+
Verify the registration answers before calling tier 2 ready:
|
|
95
|
+
|
|
96
|
+
```bash
|
|
97
|
+
claude mcp list | grep -i firebase
|
|
98
|
+
```
|
|
99
|
+
|
|
100
|
+
A tier-2 session is a person's credential with a refresh token that expires. Never
|
|
101
|
+
present it as the durable answer - it is the bridge while the role request moves.
|
|
102
|
+
|
|
103
|
+
## 4. appId discovery
|
|
104
|
+
|
|
105
|
+
Every Crashlytics call needs the opaque `appId` (`1:<number>:ios:<hex>`), and a
|
|
106
|
+
console URL only ever carries the bundle id. The fetcher resolves it through the
|
|
107
|
+
Firebase Management API on each run, which works but costs a round trip and needs
|
|
108
|
+
the project resolved first.
|
|
109
|
+
|
|
110
|
+
The repo already holds the answer. `GoogleService-Info*.plist` (iOS) and
|
|
111
|
+
`google-services.json` (Android) carry `PROJECT_ID` and `GOOGLE_APP_ID` for every
|
|
112
|
+
target:
|
|
113
|
+
|
|
114
|
+
```bash
|
|
115
|
+
bash "$HOME/.claude/scripts/firebase-app-discovery.sh" <repo-dir> --json
|
|
116
|
+
```
|
|
117
|
+
|
|
118
|
+
It prints entries shaped for `prefs.global.firebase.accounts[]`, each with an
|
|
119
|
+
`apps[]` of `{bundleId, appId, platform}`. Merge them into the account that
|
|
120
|
+
already carries the matching `projectId`, keeping its `keychainKey` - which
|
|
121
|
+
credential covers a project is the user's mapping to make, so the script never
|
|
122
|
+
guesses one.
|
|
123
|
+
|
|
124
|
+
A repo with several targets has several plists and they do not all point at the
|
|
125
|
+
same Firebase project, so the script reads every match rather than stopping at
|
|
126
|
+
the first, and skips build outputs where a copied plist would count one app twice.
|
|
127
|
+
|
|
128
|
+
Confirm the merge before writing: show the projectId and the app count, and write
|
|
129
|
+
nothing on a decline. Finding nothing is a normal answer - the fetcher still
|
|
130
|
+
resolves appIds live, one round trip per run.
|
|
131
|
+
|
|
132
|
+
## 5. v1alpha, and what to do when it breaks
|
|
133
|
+
|
|
134
|
+
Both tiers read `firebasecrashlytics.googleapis.com/v1alpha`. That surface is
|
|
135
|
+
undocumented and unversioned in the usual sense: Google may change or withdraw it
|
|
136
|
+
without notice, and when they do, both tiers fail at once.
|
|
137
|
+
|
|
138
|
+
This is written down so the failure is diagnosable rather than mysterious. The
|
|
139
|
+
symptom is a 404 or a changed response shape on a call that worked yesterday, with
|
|
140
|
+
credentials that still pass `--probe`. When it happens, the fetcher exits 3 with
|
|
141
|
+
`crashlytics-unreachable`, the orchestrator degrades to advisory, and the pipeline
|
|
142
|
+
keeps moving - it does not halt a run over a crash report.
|
|
143
|
+
|
|
144
|
+
## 6. Skipping
|
|
145
|
+
|
|
146
|
+
All of it is optional. Skip and nothing is written; Crashlytics links in tickets
|
|
147
|
+
are simply not enriched, and the run says so rather than pretending it looked.
|
|
148
|
+
|
|
149
|
+
Never lead with "paste the stack trace yourself" - that is tier 3, and offering
|
|
150
|
+
the last resort first trains the user to skip the durable fix (`refs/keychain.md`
|
|
151
|
+
Rule 2).
|
|
@@ -38,6 +38,11 @@
|
|
|
38
38
|
"type": ["string", "null"],
|
|
39
39
|
"description": "Task id that produced this entry, or null when added manually."
|
|
40
40
|
},
|
|
41
|
+
"source": {
|
|
42
|
+
"type": "string",
|
|
43
|
+
"enum": ["triage", "transcript-mining", "manual"],
|
|
44
|
+
"description": "v16.27+ - who found this. `triage` is the Phase 4 distill, `transcript-mining` is learn-from-transcripts.mjs correlating a failure with the retry that worked, `manual` is a human or a model writing it directly. Kept separate because the three have different error modes and learning-curve.mjs has to be able to say which kind of learning is actually accumulating: a machine-mined path correction and a model's architectural claim are not the same evidence, and a metric that pools them can rise while the useful half is flat. Absent means the row predates the field."
|
|
45
|
+
},
|
|
41
46
|
"confidence": {
|
|
42
47
|
"type": "string",
|
|
43
48
|
"enum": ["low", "medium", "high"],
|
|
@@ -68,7 +68,7 @@
|
|
|
68
68
|
},
|
|
69
69
|
"firebase": {
|
|
70
70
|
"type": "string",
|
|
71
|
-
"description": "Firebase JSON
|
|
71
|
+
"description": "Keychain key for the Firebase service-account JSON, stored as issued. project_id is read straight out of it - no separate firebase_project entry."
|
|
72
72
|
},
|
|
73
73
|
"fortify": {
|
|
74
74
|
"type": "string"
|
|
@@ -179,7 +179,7 @@
|
|
|
179
179
|
},
|
|
180
180
|
"firebase": {
|
|
181
181
|
"type": ["string", "null"],
|
|
182
|
-
"description": "Firebase JSON
|
|
182
|
+
"description": "Keychain key for the Firebase service-account JSON, stored exactly as Google issued it - no base64 wrapper. project_id and client_email are read straight out of it; fetch-crashlytics.sh signs a JWT with private_key and exchanges it for an access token."
|
|
183
183
|
},
|
|
184
184
|
"firebase_sa": {
|
|
185
185
|
"type": ["string", "null"],
|
|
@@ -594,6 +594,10 @@
|
|
|
594
594
|
"type": "string",
|
|
595
595
|
"description": "v15.15+ - Graylog TEST host without scheme. Optional; leaving it unset means fetch-graylog.sh only ever searches production. Test and production are separate instances, so a trx id minted by a tester does not exist in prod and searching prod alone answers 'no logs' for a complaint that is fully logged one host over. Resolves {GRAYLOG_TEST_HOST}."
|
|
596
596
|
},
|
|
597
|
+
"jenkins": {
|
|
598
|
+
"type": "string",
|
|
599
|
+
"description": "Jenkins host without scheme. credential-inventory.sh reads it to probe the `jenkins` token; the key was undeclarable before v16.26, so the probe could only ever answer no-host-configured for a token that was fine. Optional - unset leaves the token reported as present but unprobed."
|
|
600
|
+
},
|
|
597
601
|
"corpDomain": {
|
|
598
602
|
"type": "string",
|
|
599
603
|
"description": "Corporate email / cookie domain, e.g. example.com. Resolves {CORP_DOMAIN}."
|
|
@@ -619,11 +623,35 @@
|
|
|
619
623
|
},
|
|
620
624
|
"keychainKey": {
|
|
621
625
|
"type": "string",
|
|
622
|
-
"description": "Keychain key holding this project's service-account JSON
|
|
626
|
+
"description": "Keychain key holding this project's service-account JSON, stored as issued."
|
|
623
627
|
},
|
|
624
628
|
"label": {
|
|
625
629
|
"type": "string",
|
|
626
630
|
"description": "Human label used in error output, e.g. \"redesign prod\". Defaults to the projectId."
|
|
631
|
+
},
|
|
632
|
+
"apps": {
|
|
633
|
+
"type": "array",
|
|
634
|
+
"description": "v16.26+ - bundle id -> opaque Firebase appId, discovered from the repo's GoogleService-Info*.plist / google-services.json by firebase-app-discovery.sh. Every Crashlytics call needs the appId while a console URL carries only the bundle id, so without this the fetcher spends a Management API round trip resolving it on every run. Optional: absent means the fetcher resolves it live, which still works.",
|
|
635
|
+
"items": {
|
|
636
|
+
"type": "object",
|
|
637
|
+
"additionalProperties": false,
|
|
638
|
+
"required": ["bundleId", "appId", "platform"],
|
|
639
|
+
"properties": {
|
|
640
|
+
"bundleId": {
|
|
641
|
+
"type": "string",
|
|
642
|
+
"description": "iOS bundle identifier or Android application id, as it appears in a Crashlytics console URL after the platform prefix."
|
|
643
|
+
},
|
|
644
|
+
"appId": {
|
|
645
|
+
"type": "string",
|
|
646
|
+
"description": "Opaque Firebase app id, e.g. 1:123456789:ios:abcdef0123456789. GOOGLE_APP_ID in the plist, mobilesdk_app_id in google-services.json."
|
|
647
|
+
},
|
|
648
|
+
"platform": {
|
|
649
|
+
"type": "string",
|
|
650
|
+
"enum": ["ios", "android"],
|
|
651
|
+
"description": "Which console path the appId belongs under."
|
|
652
|
+
}
|
|
653
|
+
}
|
|
654
|
+
}
|
|
627
655
|
}
|
|
628
656
|
}
|
|
629
657
|
}
|
|
@@ -0,0 +1,73 @@
|
|
|
1
|
+
{
|
|
2
|
+
"$schema": "http://json-schema.org/draft-07/schema#",
|
|
3
|
+
"$id": "skill-observation.schema.json",
|
|
4
|
+
"title": "Skill observation",
|
|
5
|
+
"description": "One noticed piece of friction in the pipeline's own skills, written the moment it is noticed. The store is a directory of markdown files whose frontmatter conforms to this schema; the directory listing IS the index, so there is no index file to keep in sync and no scan that reads a body. This schema describes that frontmatter. The learnings ledger holds what a run learned about a REPO; this holds what a run learned about the PIPELINE, which nothing captured before: /multi-agent:refactor derived its findings from scratch on every invocation, so a friction noticed on Tuesday was gone by Wednesday unless it was fixed the same hour.",
|
|
6
|
+
"type": "object",
|
|
7
|
+
"additionalProperties": false,
|
|
8
|
+
"required": ["id", "title", "status", "target", "siblings_checked", "area", "date"],
|
|
9
|
+
"properties": {
|
|
10
|
+
"id": {
|
|
11
|
+
"type": "string",
|
|
12
|
+
"pattern": "^[0-9]{4}$",
|
|
13
|
+
"description": "Zero-padded sequence number, matching the file's NNNN- prefix."
|
|
14
|
+
},
|
|
15
|
+
"title": {
|
|
16
|
+
"type": "string",
|
|
17
|
+
"minLength": 8,
|
|
18
|
+
"description": "One line, stating the friction rather than the fix. 'refactor re-derives findings every run' is an observation; 'add a backlog mode' is a proposal, and belongs in the body."
|
|
19
|
+
},
|
|
20
|
+
"status": {
|
|
21
|
+
"type": "string",
|
|
22
|
+
"enum": ["open", "actioned", "declined", "superseded", "parked"],
|
|
23
|
+
"description": "open = in the queue. actioned = a change shipped. declined = decided against, with the reason in resolution. superseded = another observation covers it. parked = decided, but blocked on something outside this repo; it leaves the queue without being archived, and parked_until is then required. Five values rather than open/closed because 'declined' and 'parked' are the two that get silently dropped when the vocabulary is too small - a parked item with no expiry is a deferral wearing a disguise."
|
|
24
|
+
},
|
|
25
|
+
"parked_until": {
|
|
26
|
+
"type": "string",
|
|
27
|
+
"description": "Required when status is parked: the concrete event that unblocks it (a version, a release, an upstream fix). 'more data' is not an event - if no observation could change the decision and no date is nameable, the honest status is declined."
|
|
28
|
+
},
|
|
29
|
+
"target": {
|
|
30
|
+
"type": "array",
|
|
31
|
+
"minItems": 1,
|
|
32
|
+
"items": { "type": "string" },
|
|
33
|
+
"description": "Repo-relative paths the observation is about. Always a list, even for one path: a scalar here means every consumer needs two code paths, and the one that handles the scalar is the one that gets forgotten.",
|
|
34
|
+
"uniqueItems": true
|
|
35
|
+
},
|
|
36
|
+
"proposes_skill": {
|
|
37
|
+
"type": "array",
|
|
38
|
+
"items": { "type": "string" },
|
|
39
|
+
"description": "Skills or commands that do not exist yet but should, if this observation implies one."
|
|
40
|
+
},
|
|
41
|
+
"siblings_checked": {
|
|
42
|
+
"type": "string",
|
|
43
|
+
"minLength": 3,
|
|
44
|
+
"description": "What was found when the sibling surfaces were checked. This tree mirrors every command into skills/shared/core and again into the Copilot and Codex trees, so a fix applied to one copy drifts silently from the rest. 'checked, does not apply to the shared skill' is a valid answer; empty is not, which is why the field is required and why skill-siblings.mjs fills it mechanically rather than from memory."
|
|
45
|
+
},
|
|
46
|
+
"area": {
|
|
47
|
+
"type": "string",
|
|
48
|
+
"description": "Rough grouping for the weekly review: a phase name, a subsystem, 'gates', 'docs'."
|
|
49
|
+
},
|
|
50
|
+
"date": {
|
|
51
|
+
"type": "string",
|
|
52
|
+
"pattern": "^[0-9]{4}-[0-9]{2}-[0-9]{2}$",
|
|
53
|
+
"description": "ISO date the observation was written, absolute so it survives being read a year later."
|
|
54
|
+
},
|
|
55
|
+
"session_context": {
|
|
56
|
+
"type": "string",
|
|
57
|
+
"description": "What the session was doing when the friction appeared. Free text, kept short; it is what makes an old observation legible."
|
|
58
|
+
},
|
|
59
|
+
"resolved": {
|
|
60
|
+
"type": "string",
|
|
61
|
+
"pattern": "^[0-9]{4}-[0-9]{2}-[0-9]{2}$",
|
|
62
|
+
"description": "ISO date the status left 'open'."
|
|
63
|
+
},
|
|
64
|
+
"resolution": {
|
|
65
|
+
"type": "string",
|
|
66
|
+
"description": "One line saying what happened. Required in practice for declined and superseded - a decline with no reason is indistinguishable from neglect."
|
|
67
|
+
},
|
|
68
|
+
"reference": {
|
|
69
|
+
"type": "string",
|
|
70
|
+
"description": "A commit, PR, or issue that carried the change."
|
|
71
|
+
}
|
|
72
|
+
}
|
|
73
|
+
}
|