@mmerterden/multi-agent-pipeline 16.25.0 → 16.27.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (47) hide show
  1. package/CHANGELOG.md +77 -0
  2. package/README.md +1 -1
  3. package/README.tr.md +1 -1
  4. package/install/templates/claude-hooks.json +32 -1
  5. package/package.json +1 -1
  6. package/pipeline/commands/multi-agent/help/SKILL.md +2 -0
  7. package/pipeline/commands/multi-agent/refactor/SKILL.md +23 -1
  8. package/pipeline/commands/multi-agent/search/SKILL.md +28 -0
  9. package/pipeline/commands/multi-agent/setup/SKILL.md +18 -44
  10. package/pipeline/commands/multi-agent/status/SKILL.md +9 -0
  11. package/pipeline/lib/credential-inventory.sh +142 -18
  12. package/pipeline/lib/fetch-crashlytics.sh +123 -28
  13. package/pipeline/multi-agent-refs/features/url-enrichment.md +1 -1
  14. package/pipeline/multi-agent-refs/features/visual-evidence.md +5 -0
  15. package/pipeline/multi-agent-refs/keychain.md +65 -20
  16. package/pipeline/multi-agent-refs/knowledge.md +27 -0
  17. package/pipeline/multi-agent-refs/phases/operations.md +7 -1
  18. package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +3 -1
  19. package/pipeline/multi-agent-refs/phases/phase-4-review.md +1 -8
  20. package/pipeline/multi-agent-refs/phases/phase-7-report.md +11 -21
  21. package/pipeline/multi-agent-refs/picker-contract.md +1 -1
  22. package/pipeline/multi-agent-refs/refactor/observations.md +81 -0
  23. package/pipeline/multi-agent-refs/setup/firebase.md +151 -0
  24. package/pipeline/schemas/learnings-ledger.schema.json +5 -0
  25. package/pipeline/schemas/prefs.schema.json +31 -3
  26. package/pipeline/schemas/skill-observation.schema.json +73 -0
  27. package/pipeline/scripts/capture-flush.sh +158 -0
  28. package/pipeline/scripts/capture-resume.sh +87 -0
  29. package/pipeline/scripts/crush-json.mjs +283 -0
  30. package/pipeline/scripts/firebase-app-discovery.sh +114 -0
  31. package/pipeline/scripts/keychain-save.sh +5 -8
  32. package/pipeline/scripts/keychain.py +76 -14
  33. package/pipeline/scripts/learn-from-transcripts.mjs +625 -0
  34. package/pipeline/scripts/learning-curve.mjs +22 -4
  35. package/pipeline/scripts/learnings-ledger.mjs +86 -12
  36. package/pipeline/scripts/note-session.sh +187 -0
  37. package/pipeline/scripts/observations.mjs +347 -0
  38. package/pipeline/scripts/offload-ref.sh +45 -2
  39. package/pipeline/scripts/pre-commit-check.sh +31 -1
  40. package/pipeline/scripts/scan-agent-config.sh +12 -3
  41. package/pipeline/scripts/skill-siblings.mjs +187 -0
  42. package/pipeline/scripts/triage-memory.mjs +73 -9
  43. package/pipeline/skills/.skill-manifest.json +1 -1
  44. package/pipeline/skills/shared/core/multi-agent-refactor/SKILL.md +23 -1
  45. package/pipeline/skills/shared/core/multi-agent-search/SKILL.md +28 -0
  46. package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +37 -6
  47. package/pipeline/skills/shared/core/multi-agent-status/SKILL.md +9 -0
@@ -150,7 +150,13 @@ This keeps orchestrator context lean and enables programmatic routing.
150
150
 
151
151
  **Proactive compaction + phase-boundary checkpoint**: the orchestrator follows ~2,500 lines of phase prose in one session; once context fills, it starts dropping steps - the single biggest cause of "it got stuck / skipped a step." Two defenses, both required on full-pipeline runs:
152
152
 
153
- - *Phase-boundary checkpoint.* At every phase transition, before loading the next phase doc, write the durable state (`agent-state.json` phase status + `files[]` + `retryCount`) and append a structured handoff block to `agent-log.md` (format below). The next phase reads state + log, not the back-conversation - so a transition is a clean re-entry point even if context is later compacted.
153
+ - *Phase-boundary checkpoint.* At every phase transition, before loading the next phase doc, write the durable state (`agent-state.json` phase status + `files[]` + `retryCount`) and append a structured handoff block to `agent-log.md` (format below). The next phase reads state + log, not the back-conversation - so a transition is a clean re-entry point even if context is later compacted. Then flush what the run has learned so far:
154
+
155
+ ```bash
156
+ bash $HOME/.claude/scripts/capture-flush.sh --state "$STATE_FILE" --quiet
157
+ ```
158
+
159
+ The durable stores used to be written only in Phase 7, which is the phase a run is LEAST likely to reach: a run killed in Phase 3 threw away every finding it had established, and the next run on the same repo paid to rediscover it. The flush is idempotent (the second call through writes 0 rows), costs no API tokens, calls no model, and never fails the transition. Phase 7 is now the LAST flush rather than the only one. The trigger is deliberately mechanical - a phase transition, not a model noticing that a moment qualifies.
154
160
  - *Compaction trigger.* If conversation context exceeds ~50%, run `/compact` preserving "modified files, plan, open review findings, current phase + sub-step" before continuing. Don't wait for auto-compaction near the limit - it triggers exactly when context is worst and is lossy. After compaction, re-read `agent-state.json` AND the latest `## Handoff` block in `agent-log.md` to re-ground.
155
161
 
156
162
  **Handoff block (v10.8.0)**: the structured artifact the phase-boundary checkpoint appends to `agent-log.md`. Written by the orchestrator from state it already holds - no agent dispatch, no extra LLM call. Cap at ~15 lines; the latest block is authoritative (earlier ones are history). This is the fresh-context re-entry contract: a resume or post-compaction session rebuilds working context from the latest handoff + `agent-state.json` + git log, never from conversation memory.
@@ -145,7 +145,9 @@ Gated by `prefs.global.repoMap.enabled` (default: `false`). When enabled, runs `
145
145
 
146
146
  #### Step 2.6 - Code Graph Injection (advisory, opt-in)
147
147
 
148
- Gated by `prefs.global.codeGraph.enabled` (default: `false`). With a rule file for `detectedStack`, Phase 1 refreshes the code graph, queries it, and hands Explore a ranked starting set instead of a full scan; `graph-affected` feeds `analysis.touchedAreas[]`. Zero API cost, read-only. Narrows an open-ended search; it does not replace a grep for a name the task already spells out. Commands, validator contract, measurements: `$HOME/.claude/multi-agent-refs/features/code-graph.md`.
148
+ Gated by `prefs.global.codeGraph.enabled` (default: `false`). With a rule file for `detectedStack`, Phase 1 refreshes the graph and queries it; `graph-affected` feeds `analysis.touchedAreas[]`. Zero API cost, read-only. Commands and measurements: `$HOME/.claude/multi-agent-refs/features/code-graph.md`.
149
+
150
+ **A valid result REPLACES the opening sweep** rather than sitting beside it: Explore starts from those files and walks outward, with no broad `Glob`/`Grep` pass first - running both pays twice, and the second re-derives what the graph said. No graph (missing rule, stale, or disabled) leaves the previous behaviour untouched; a task naming a symbol or path is still grepped directly.
149
151
 
150
152
  #### Step 3 - Codebase Exploration
151
153
 
@@ -262,14 +262,7 @@ Visual-fidelity mismatches against the captured screenshot are BLOCKING findings
262
262
 
263
263
  - Canonical component usage: the Code Connect-mapped component is used verbatim - a sound-alike substitute, a forked copy, or ad-hoc inline UI where a mapping exists is blocking (Phase 3 "Design fidelity contract")
264
264
  - Inter-component spacing: gaps, paddings, and alignment BETWEEN components match the design's measured values mapped to spacing tokens - invented numeric values are blocking
265
- - Avatar icon (presence, glyph, shape, position)
266
- - Field grouping (one rounded box vs two; separator vs gap)
267
- - Character counter visibility
268
- - Header style (size, weight, alignment)
269
- - Inline error layout (icon presence, text colour, position relative to the input)
270
- - Button height
271
- - Indicator chip placement
272
- - Placeholder copy and position
265
+ - Everything else: the 30-group catalog in `$HOME/.claude/multi-agent-refs/features/design-conformance.md`. Two of its rules carry into review: a height or inset is **measured**, never read from the token under test; an unrun check is a fail-to-verify, not a pass.
273
266
 
274
267
  When `state.figmaAccess.tier === 3` (user-attached screenshot, no Code Connect snippet), the reviewer additionally sets `findings[i].severity = "blocking"` and `findings[i].tag = "review_blocking_tier3"` on every UI atom that lacks a confirmed canonical-component mapping. The triage step preserves these findings unless the user has explicitly cleared the open question.
275
268
 
@@ -219,38 +219,28 @@ This is independent of the channels-side `reportContent.costSummary` (which gate
219
219
 
220
220
  **Missing-telemetry disclosure (required).** Name untracked phase ids as cost-unavailable in the report and closing summary. Mechanic: `payload-contracts.md`.
221
221
 
222
- **Triage memory ingest (mandatory):** after Phase 4 produces a final triage output, persist the accepted/deferred/rejected rows into the per-repo triage corpus so Phase 1 enrichment and Phase 4 prior-art lookup can recall them on future tasks. Idempotent - re-running on the same task writes 0 rows.
222
+ **Final capture flush (mandatory):** persist the triage findings and the durable learnings this run established.
223
223
 
224
224
  ```bash
225
- # Salvaged copy first: Phase 6 removes the worktree once the PR is open, and this
226
- # reader is `[ -f ]`-guarded, so a wrong path degrades SILENTLY.
227
- TRIAGE_PATH="$(jq -r '.artifactsPath // empty' "$STATE_FILE" 2>/dev/null)/triage-output.json"
228
- [ -f "$TRIAGE_PATH" ] || TRIAGE_PATH="$WORKTREE/triage-output.json"
229
- if [ -f "$TRIAGE_PATH" ]; then
230
- node $HOME/.claude/scripts/triage-memory.mjs ingest \
231
- --triage "$TRIAGE_PATH" \
232
- --task-id "$TASK_ID" \
233
- --task-title "$TASK_TITLE" \
234
- --stack "$DETECTED_STACK" >/dev/null 2>&1 || true
235
- fi
225
+ bash $HOME/.claude/scripts/capture-flush.sh --state "$STATE_FILE"
236
226
  ```
237
227
 
238
- Best-effort. The corpus is JSONL at `~/.claude/memory/multi-agent/<repo-slug>/triage-corpus.jsonl` (per-repo isolation - never cross-leaks between projects). Disabled when `prefs.global.priorArtEnrichment.ingestOnComplete = false`.
228
+ Every phase boundary already made this call (`operations.md`), so by Phase 7 it usually writes 0 rows - and that is the point. These writes used to live ONLY here, in the phase a run is least likely to reach: a run killed in Phase 3 lost every finding it had established. Phase 7 is now the last flush, not the only one.
239
229
 
240
- **Learnings ledger distill (on by default via `prefs.global.learningsLedger.enabled`):** distill this run's rejected findings into durable rejected-preference entries so future reviewers stop re-raising them, and record any durable architectural fact the analysis surfaced. Idempotent (dedup by kind + statement). Safety: `from-triage` skips blocking-severity rejections - a wrong rejection of a blocking issue must never permanently suppress that class; use `learnings-ledger.mjs forget` to clear a bad/stale entry.
230
+ What it does, so this doc stays inspectable:
231
+
232
+ - **Triage ingest** (`triage-memory.mjs ingest`). accepted/deferred/rejected rows into the per-repo corpus, so Phase 1 enrichment and Phase 4 prior-art recall them later. Idempotent. Reads the salvaged copy under `artifactsPath` FIRST: Phase 6 removes the worktree, and a worktree-first reader degrades silently for exactly the runs worth rescuing. JSONL at `~/.claude/memory/multi-agent/<repo-slug>/triage-corpus.jsonl`, per-repo, never cross-leaking. Off when `prefs.global.priorArtEnrichment.ingestOnComplete = false`.
233
+ - **Ledger distill** (`learnings-ledger.mjs from-triage`). Rejected findings become durable `rejected-preference` entries so reviewers stop re-raising them. Idempotent. Safety: blocking-severity rejections are NOT distilled - a wrong rejection must never permanently suppress that class; `learnings-ledger.mjs forget` clears a stale one. JSONL beside the corpus; its brief replays into Phase 1 and Phase 4. On by default via `prefs.global.learningsLedger.enabled`.
234
+
235
+ No model runs in the flush - it derives everything from `triage-output.json` plus `agent-state.json`, which is what lets a hook call it. The model-dependent parts of this phase (Step 3 knowledge capture, memory synthesis) stay here, because a hook cannot think.
236
+
237
+ A durable architectural fact is a judgement call, so it stays here rather than in the flush:
241
238
 
242
239
  ```bash
243
- if [ -f "$TRIAGE_PATH" ]; then
244
- node $HOME/.claude/scripts/learnings-ledger.mjs from-triage \
245
- --triage "$TRIAGE_PATH" --task "$TASK_ID" >/dev/null 2>&1 || true
246
- fi
247
- # Optionally capture a durable architectural fact the analysis established:
248
240
  # node $HOME/.claude/scripts/learnings-ledger.mjs add --kind fact \
249
241
  # --statement "<one-line fact>" --scope "<path-glob>" --task "$TASK_ID" >/dev/null 2>&1 || true
250
242
  ```
251
243
 
252
- The ledger is JSONL at `~/.claude/memory/multi-agent/<repo-slug>/learnings-ledger.jsonl`, next to the triage corpus, per-repo isolated. Its brief is replayed into Phase 1 analysis and Phase 4 triage on future runs.
253
-
254
244
  Print the closing report to the terminal. Two blocks, in this order - what the pipeline spent, then what it changed:
255
245
 
256
246
  ```bash
@@ -81,6 +81,6 @@ In autopilot, `ask_choice` resolves to `default` (or the safe first option) with
81
81
 
82
82
  ## Deterministic gates note
83
83
 
84
- Claude Code's `PreToolUse` exit-2 hooks are the HARD blocking gates. Three ship, none needing run-specific arguments so they are naturally hookable: (1) `pre-commit-check.sh` scans the staged diff on every `git commit` and blocks on a detected secret; (2) `agent-guard.sh` runs on `git commit` + `git push` and blocks AI/assistant attribution in a commit message and force-push to a protected branch (main/master/develop); (3) `check-read-size.sh` runs on `Read` and on the shell commands that read a file whole, and routes an oversized read to a cheap worker (`bulk-read.sh`) instead of the caller's own rung. The first two inspect what a run WRITES; the third inspects what it pays to READ, and it is inert until `bulkRead.mode` is set to `observe` or `enforce`, so merging the block changes nothing until the user opts in. Its `observe` mode blocks nothing and only logs, which is how the baseline is measured before anything is routed. All three are self-contained, fail-open on internal error, and never execute the inspected command. The recommended hook block ships at `install/templates/claude-hooks.json`; `multi-agent:setup` offers to merge it into `~/.claude/settings.json`. The other deterministic gates (evidence, consensus, intent, learnings) are invoked by the pipeline phases with per-run arguments (a build-log path, the triage JSON, the free-text input), so they are phase-enforced by contract, not OS-hookable.
84
+ Claude Code's `PreToolUse` exit-2 hooks are the HARD blocking gates. Three ship, none needing run-specific arguments so they are naturally hookable: (1) `pre-commit-check.sh` scans the staged diff on every `git commit` and blocks on a detected secret; (2) `agent-guard.sh` runs on `git commit` + `git push` and blocks AI/assistant attribution in a commit message and force-push to a protected branch (main/master/develop); (3) `check-read-size.sh` runs on `Read` and on the shell commands that read a file whole, and routes an oversized read to a cheap worker (`bulk-read.sh`) instead of the caller's own rung. The first two inspect what a run WRITES; the third inspects what it pays to READ, and it is inert until `bulkRead.mode` is set to `observe` or `enforce`, so merging the block changes nothing until the user opts in. Its `observe` mode blocks nothing and only logs, which is how the baseline is measured before anything is routed. All three are self-contained, fail-open on internal error, and never execute the inspected command. Two capture hooks ship in the same block and block nothing: `SessionEnd` runs `capture-flush.sh --if-stale` (writing a killed run's findings into the per-repo stores, since every durable write used to live in Phase 7 - the phase a run is least likely to reach) plus `note-session.sh` (the mechanical shape of a non-pipeline session: tools used, commands that failed, calls the user refused - never an argument, never any output), and `SessionStart` runs `capture-resume.sh`, at most two lines about an unfinished run and a stale observation queue. Neither calls a model; both exit 0 on every path. The recommended hook block ships at `install/templates/claude-hooks.json`; `multi-agent:setup` offers to merge it into `~/.claude/settings.json`. The other deterministic gates (evidence, consensus, intent, learnings) are invoked by the pipeline phases with per-run arguments (a build-log path, the triage JSON, the free-text input), so they are phase-enforced by contract, not OS-hookable.
85
85
 
86
86
  Copilot CLI has no `PreToolUse` equivalent, so the secret scan there is workflow-enforced (run as a phase step, not OS-blocked) plus a CI smoke-gate step.
@@ -0,0 +1,81 @@
1
+ # The observation backlog (refactor Step 0e / `backlog` mode)
2
+
3
+ Loaded on demand by `/multi-agent:refactor`. The SKILL.md carries the band intro; this file is the flow.
4
+
5
+ ## 1. Why a queue exists at all
6
+
7
+ `/multi-agent:refactor` derives its findings from scratch on every invocation. That is the right design for a sweep, and it has one consequence nobody chose: a friction noticed on Tuesday is gone by Wednesday unless it was fixed within the hour. Most are not, because they surface mid-task, when stopping to fix them is the wrong call.
8
+
9
+ So the friction is written down the moment it is noticed, and this band decides what to do with it later. The two halves are deliberately separate: noticing is cheap and must never wait for a decision, deciding is expensive and must never happen mid-task.
10
+
11
+ Store: `$HOME/.claude/memory/multi-agent/_pipeline/observations/NNNN-slug.md`, one file per observation, frontmatter per `schemas/skill-observation.schema.json`. The directory listing is the index - there is no index file, because an index is a second copy of the truth and the second copy is the one that goes stale.
12
+
13
+ ## 2. Reading the queue
14
+
15
+ ```bash
16
+ node "$HOME/.claude/scripts/observations.mjs" scan --status open --json
17
+ ```
18
+
19
+ `scan` reads frontmatter only and never opens a body, so a queue of several hundred costs almost nothing.
20
+
21
+ **Exit 3 is `SCAN BROKEN` and it is not an empty queue.** It means files exist on disk and none parsed - the reader is broken. Halt the band and say so. An empty scan otherwise reports two different facts with one answer ("nothing to find" and "the question was never asked"), and only the first is a finding; the guard exists so those two can never be confused again.
22
+
23
+ ## 3. Splitting the queue
24
+
25
+ Every open observation lands in exactly one of three buckets:
26
+
27
+ | Bucket | Meaning | What happens |
28
+ |---|---|---|
29
+ | Actionable | a change worth making now | goes into the plan as a normal band item with its file and fix |
30
+ | Sibling propagation | already fixed in one copy, not the others | goes into the plan as a propagation item, with the exact surfaces named |
31
+ | To decline | not worth doing, or wrong | proposed for `declined` with a one-line reason |
32
+
33
+ The second bucket is the one this tree specifically needs. A command is never one file: it is authored under `commands/`, mirrored into `skills/shared/core/`, and installed again into the Claude, Copilot and Codex trees. A fix applied to whichever copy was open drifts from the rest silently - nothing errors and every gate stays green. Resolve the surfaces mechanically, never from memory:
34
+
35
+ ```bash
36
+ node "$HOME/.claude/scripts/skill-siblings.mjs" <path> --json
37
+ node "$HOME/.claude/scripts/skill-siblings.mjs" --audit # every command at once
38
+ ```
39
+
40
+ ## 4. Deciding, and the deferral that wears a disguise
41
+
42
+ An observation leaves the queue with a status, never by being ignored:
43
+
44
+ - `actioned` - a change shipped. Record the commit in `reference`.
45
+ - `declined` - decided against. A decline with no `resolution` is indistinguishable from neglect, so the reason is written.
46
+ - `superseded` - another observation covers it. Name which.
47
+ - `parked` - decided, but blocked on something outside this repo.
48
+
49
+ `parked` requires `parked_until` naming the concrete event that unblocks it: a version, a release, an upstream fix. "Let us gather more data" is not an event. If no observation could change the decision and no date is nameable, the honest status is `declined` - a park with no expiry leaves the queue and never comes back, which is a silent decline wearing a friendlier word.
50
+
51
+ ```bash
52
+ node "$HOME/.claude/scripts/observations.mjs" resolve --id 0007 --status actioned \
53
+ --resolution "backlog mode added" --reference "<sha>"
54
+ ```
55
+
56
+ ## 5. Applying
57
+
58
+ Changes go to a staged copy and are shown before anything is installed, exactly as the drift band does. The user installs; this band never edits an installed skill in place.
59
+
60
+ When any observation is resolved, stamp the review:
61
+
62
+ ```bash
63
+ date +%Y-%m-%d > "$HOME/.claude/memory/multi-agent/_pipeline/last-review-date.txt"
64
+ ```
65
+
66
+ `capture-resume.sh` reads that stamp at session start and offers one line when it is seven days old and the queue is non-empty. One line, never a block - the user's own work does not wait on the pipeline's housekeeping.
67
+
68
+ ## 6. Writing an observation
69
+
70
+ Any session may add one, in the same turn the friction appears, and then carry on:
71
+
72
+ ```bash
73
+ node "$HOME/.claude/scripts/observations.mjs" add \
74
+ --title "<the friction, not the fix>" \
75
+ --target <repo-relative path> [--target ...] \
76
+ --area <phase|gates|docs|...> \
77
+ --session "<what the session was doing>" \
78
+ --body "<detail>"
79
+ ```
80
+
81
+ State the friction, not the remedy: "refactor re-derives its findings every run" is an observation, "add a backlog mode" is a proposal, and proposals age badly while observations do not. `siblings_checked` is filled by `skill-siblings.mjs` automatically and cannot be empty - the whole point is that the answer is computed rather than recalled, because recollection is the faculty that produced the drift.
@@ -0,0 +1,151 @@
1
+ # Firebase / Crashlytics Onboarding (setup Step 3c)
2
+
3
+ Loaded on demand by `/multi-agent:setup` Step 3c (optional, any platform). The SKILL.md carries the step intro; this file is the full flow.
4
+
5
+ Runs inside Step 3 alongside the other missing credentials. A user who already
6
+ holds a Firebase service-account JSON gets it mapped by Step 1 discovery like any
7
+ other credential; this flow covers what discovery cannot see - whether that
8
+ account may actually read Crashlytics, and what to do when it may not.
9
+
10
+ ## 1. Why there are three ways in
11
+
12
+ Crashlytics has one API and three ways to authenticate against it. They fail
13
+ differently, so the pipeline measures rather than assumes.
14
+
15
+ | Tier | Path | Serves | Needs |
16
+ |---|---|---|---|
17
+ | 1 | Service account + `lib/fetch-crashlytics.sh` | headless: autopilot, cron, url-enrichment, review chains | a Crashlytics role on the service account |
18
+ | 2 | Firebase MCP + an interactive `firebase login` | a person at the keyboard | a browser, a real terminal, a session that expires |
19
+ | 3 | The user pastes the stack trace | last resort | nothing, and it produces degraded evidence |
20
+
21
+ Order is not fixed - whichever tier is ready wins, exactly as in the Figma chain.
22
+ But **tier 1 is preferred whenever both are ready**, and the reason is structural:
23
+ url-enrichment expands a Crashlytics link found in a Jira ticket while nobody is
24
+ watching, and autopilot runs with no terminal to log into. A tier that needs a
25
+ browser cannot serve those. Tier 2 fills the gap while a role request is pending;
26
+ it does not replace tier 1.
27
+
28
+ `credential-inventory.sh --probe` reports which tier is live:
29
+ `tier-1-ready` / `tier-2-ready` / `tier-1-no-grant` / `malformed` / `unreachable`.
30
+ The vocabulary and what to tell the user for each: `refs/keychain.md`.
31
+
32
+ ## 2. Tier 1 - the service account
33
+
34
+ One credential, `keychainMapping.firebase`, holding the JSON exactly as Firebase
35
+ Console issued it. Firebase Console -> Project settings -> Service accounts ->
36
+ Generate new private key. Store it through the normal Token Save Flow; it is
37
+ multi-line and that is fine - the store round-trips it byte for byte.
38
+
39
+ Do **not** base64 it. Encoding is the store's internal business, and a wrapper on
40
+ the way in that nothing unwraps on the way out is how this path stayed broken.
41
+
42
+ Several Firebase projects per team is the normal case (legacy plus redesign,
43
+ staging plus prod). The single slot is the fallback; extra projects go in
44
+ `prefs.global.firebase.accounts[]` as `{projectId, keychainKey, label}` and the
45
+ URL's project id picks the key.
46
+
47
+ **The role is the part discovery cannot see.** A freshly generated service account
48
+ authenticates perfectly and reads nothing: `firebase-adminsdk-*` carries no
49
+ Crashlytics permission by default. The probe surfaces this as `tier-1-no-grant`,
50
+ and the fix is one line to whoever holds Owner on the project:
51
+
52
+ > Please grant `roles/firebasecrashlytics.viewer` on project `<projectId>` to the
53
+ > service account `<client_email>`. It is read-only: it allows reading crash
54
+ > issues and their stack traces, and nothing else.
55
+
56
+ `roles/firebase.viewer` also works and is broader. Ask for the narrower one first.
57
+
58
+ Verify with `bash lib/fetch-crashlytics.sh --probe`. It asks IAM what the account
59
+ may do rather than calling Crashlytics and reading the error, because a 403 from
60
+ the data endpoint means "no permission" but so does a 403 from a disabled API,
61
+ and those need different fixes.
62
+
63
+ ## 3. Tier 2 - interactive session plus MCP
64
+
65
+ Two halves, both required. Either alone reaches nothing.
66
+
67
+ **The session.** `firebase login` opens a browser and does not work from inside an
68
+ agent harness - it needs a real terminal the user drives themselves. Ask them to
69
+ run it and say when it is done; `firebase login:list` confirms.
70
+
71
+ **The MCP server.** The Firebase CLI serves it itself - there is no package to
72
+ install beyond the CLI the login already needed:
73
+
74
+ ```bash
75
+ claude mcp add firebase -- firebase mcp --dir "$PWD"
76
+ ```
77
+
78
+ `--dir` is not decoration: the server resolves the project from that directory's
79
+ `.firebaserc` / `firebase.json`, so a server registered without it answers for
80
+ whatever directory it happens to start in.
81
+
82
+ Scope is a real choice with three answers, so ask rather than guess:
83
+
84
+ | Scope | Where it lands | When |
85
+ |---|---|---|
86
+ | `local` (default) | this user's entry for this repo, in `~/.claude.json` | the normal answer - the grant stays scoped to the repo that needs it |
87
+ | `user` | this user, every repo | only when the user works across several Firebase projects and says so; `--dir` then has to be re-pointed per repo |
88
+ | `project` | `.mcp.json` **committed in the repo** | only on an explicit ask - this registers the server for everyone who clones it |
89
+
90
+ Default to `local` and never reach for `project` on your own: the server can read
91
+ that Firebase project's data, and committing it hands that reach to the whole
92
+ team as a side effect of one person's setup.
93
+
94
+ Verify the registration answers before calling tier 2 ready:
95
+
96
+ ```bash
97
+ claude mcp list | grep -i firebase
98
+ ```
99
+
100
+ A tier-2 session is a person's credential with a refresh token that expires. Never
101
+ present it as the durable answer - it is the bridge while the role request moves.
102
+
103
+ ## 4. appId discovery
104
+
105
+ Every Crashlytics call needs the opaque `appId` (`1:<number>:ios:<hex>`), and a
106
+ console URL only ever carries the bundle id. The fetcher resolves it through the
107
+ Firebase Management API on each run, which works but costs a round trip and needs
108
+ the project resolved first.
109
+
110
+ The repo already holds the answer. `GoogleService-Info*.plist` (iOS) and
111
+ `google-services.json` (Android) carry `PROJECT_ID` and `GOOGLE_APP_ID` for every
112
+ target:
113
+
114
+ ```bash
115
+ bash "$HOME/.claude/scripts/firebase-app-discovery.sh" <repo-dir> --json
116
+ ```
117
+
118
+ It prints entries shaped for `prefs.global.firebase.accounts[]`, each with an
119
+ `apps[]` of `{bundleId, appId, platform}`. Merge them into the account that
120
+ already carries the matching `projectId`, keeping its `keychainKey` - which
121
+ credential covers a project is the user's mapping to make, so the script never
122
+ guesses one.
123
+
124
+ A repo with several targets has several plists and they do not all point at the
125
+ same Firebase project, so the script reads every match rather than stopping at
126
+ the first, and skips build outputs where a copied plist would count one app twice.
127
+
128
+ Confirm the merge before writing: show the projectId and the app count, and write
129
+ nothing on a decline. Finding nothing is a normal answer - the fetcher still
130
+ resolves appIds live, one round trip per run.
131
+
132
+ ## 5. v1alpha, and what to do when it breaks
133
+
134
+ Both tiers read `firebasecrashlytics.googleapis.com/v1alpha`. That surface is
135
+ undocumented and unversioned in the usual sense: Google may change or withdraw it
136
+ without notice, and when they do, both tiers fail at once.
137
+
138
+ This is written down so the failure is diagnosable rather than mysterious. The
139
+ symptom is a 404 or a changed response shape on a call that worked yesterday, with
140
+ credentials that still pass `--probe`. When it happens, the fetcher exits 3 with
141
+ `crashlytics-unreachable`, the orchestrator degrades to advisory, and the pipeline
142
+ keeps moving - it does not halt a run over a crash report.
143
+
144
+ ## 6. Skipping
145
+
146
+ All of it is optional. Skip and nothing is written; Crashlytics links in tickets
147
+ are simply not enriched, and the run says so rather than pretending it looked.
148
+
149
+ Never lead with "paste the stack trace yourself" - that is tier 3, and offering
150
+ the last resort first trains the user to skip the durable fix (`refs/keychain.md`
151
+ Rule 2).
@@ -38,6 +38,11 @@
38
38
  "type": ["string", "null"],
39
39
  "description": "Task id that produced this entry, or null when added manually."
40
40
  },
41
+ "source": {
42
+ "type": "string",
43
+ "enum": ["triage", "transcript-mining", "manual"],
44
+ "description": "v16.27+ - who found this. `triage` is the Phase 4 distill, `transcript-mining` is learn-from-transcripts.mjs correlating a failure with the retry that worked, `manual` is a human or a model writing it directly. Kept separate because the three have different error modes and learning-curve.mjs has to be able to say which kind of learning is actually accumulating: a machine-mined path correction and a model's architectural claim are not the same evidence, and a metric that pools them can rise while the useful half is flat. Absent means the row predates the field."
45
+ },
41
46
  "confidence": {
42
47
  "type": "string",
43
48
  "enum": ["low", "medium", "high"],
@@ -68,7 +68,7 @@
68
68
  },
69
69
  "firebase": {
70
70
  "type": "string",
71
- "description": "Firebase JSON (base64-encoded). project_id is parsed from the decoded JSON - no separate firebase_project entry."
71
+ "description": "Keychain key for the Firebase service-account JSON, stored as issued. project_id is read straight out of it - no separate firebase_project entry."
72
72
  },
73
73
  "fortify": {
74
74
  "type": "string"
@@ -179,7 +179,7 @@
179
179
  },
180
180
  "firebase": {
181
181
  "type": ["string", "null"],
182
- "description": "Firebase JSON (base64-encoded). project_id is parsed from the decoded JSON."
182
+ "description": "Keychain key for the Firebase service-account JSON, stored exactly as Google issued it - no base64 wrapper. project_id and client_email are read straight out of it; fetch-crashlytics.sh signs a JWT with private_key and exchanges it for an access token."
183
183
  },
184
184
  "firebase_sa": {
185
185
  "type": ["string", "null"],
@@ -594,6 +594,10 @@
594
594
  "type": "string",
595
595
  "description": "v15.15+ - Graylog TEST host without scheme. Optional; leaving it unset means fetch-graylog.sh only ever searches production. Test and production are separate instances, so a trx id minted by a tester does not exist in prod and searching prod alone answers 'no logs' for a complaint that is fully logged one host over. Resolves {GRAYLOG_TEST_HOST}."
596
596
  },
597
+ "jenkins": {
598
+ "type": "string",
599
+ "description": "Jenkins host without scheme. credential-inventory.sh reads it to probe the `jenkins` token; the key was undeclarable before v16.26, so the probe could only ever answer no-host-configured for a token that was fine. Optional - unset leaves the token reported as present but unprobed."
600
+ },
597
601
  "corpDomain": {
598
602
  "type": "string",
599
603
  "description": "Corporate email / cookie domain, e.g. example.com. Resolves {CORP_DOMAIN}."
@@ -619,11 +623,35 @@
619
623
  },
620
624
  "keychainKey": {
621
625
  "type": "string",
622
- "description": "Keychain key holding this project's service-account JSON (base64 or raw)."
626
+ "description": "Keychain key holding this project's service-account JSON, stored as issued."
623
627
  },
624
628
  "label": {
625
629
  "type": "string",
626
630
  "description": "Human label used in error output, e.g. \"redesign prod\". Defaults to the projectId."
631
+ },
632
+ "apps": {
633
+ "type": "array",
634
+ "description": "v16.26+ - bundle id -> opaque Firebase appId, discovered from the repo's GoogleService-Info*.plist / google-services.json by firebase-app-discovery.sh. Every Crashlytics call needs the appId while a console URL carries only the bundle id, so without this the fetcher spends a Management API round trip resolving it on every run. Optional: absent means the fetcher resolves it live, which still works.",
635
+ "items": {
636
+ "type": "object",
637
+ "additionalProperties": false,
638
+ "required": ["bundleId", "appId", "platform"],
639
+ "properties": {
640
+ "bundleId": {
641
+ "type": "string",
642
+ "description": "iOS bundle identifier or Android application id, as it appears in a Crashlytics console URL after the platform prefix."
643
+ },
644
+ "appId": {
645
+ "type": "string",
646
+ "description": "Opaque Firebase app id, e.g. 1:123456789:ios:abcdef0123456789. GOOGLE_APP_ID in the plist, mobilesdk_app_id in google-services.json."
647
+ },
648
+ "platform": {
649
+ "type": "string",
650
+ "enum": ["ios", "android"],
651
+ "description": "Which console path the appId belongs under."
652
+ }
653
+ }
654
+ }
627
655
  }
628
656
  }
629
657
  }
@@ -0,0 +1,73 @@
1
+ {
2
+ "$schema": "http://json-schema.org/draft-07/schema#",
3
+ "$id": "skill-observation.schema.json",
4
+ "title": "Skill observation",
5
+ "description": "One noticed piece of friction in the pipeline's own skills, written the moment it is noticed. The store is a directory of markdown files whose frontmatter conforms to this schema; the directory listing IS the index, so there is no index file to keep in sync and no scan that reads a body. This schema describes that frontmatter. The learnings ledger holds what a run learned about a REPO; this holds what a run learned about the PIPELINE, which nothing captured before: /multi-agent:refactor derived its findings from scratch on every invocation, so a friction noticed on Tuesday was gone by Wednesday unless it was fixed the same hour.",
6
+ "type": "object",
7
+ "additionalProperties": false,
8
+ "required": ["id", "title", "status", "target", "siblings_checked", "area", "date"],
9
+ "properties": {
10
+ "id": {
11
+ "type": "string",
12
+ "pattern": "^[0-9]{4}$",
13
+ "description": "Zero-padded sequence number, matching the file's NNNN- prefix."
14
+ },
15
+ "title": {
16
+ "type": "string",
17
+ "minLength": 8,
18
+ "description": "One line, stating the friction rather than the fix. 'refactor re-derives findings every run' is an observation; 'add a backlog mode' is a proposal, and belongs in the body."
19
+ },
20
+ "status": {
21
+ "type": "string",
22
+ "enum": ["open", "actioned", "declined", "superseded", "parked"],
23
+ "description": "open = in the queue. actioned = a change shipped. declined = decided against, with the reason in resolution. superseded = another observation covers it. parked = decided, but blocked on something outside this repo; it leaves the queue without being archived, and parked_until is then required. Five values rather than open/closed because 'declined' and 'parked' are the two that get silently dropped when the vocabulary is too small - a parked item with no expiry is a deferral wearing a disguise."
24
+ },
25
+ "parked_until": {
26
+ "type": "string",
27
+ "description": "Required when status is parked: the concrete event that unblocks it (a version, a release, an upstream fix). 'more data' is not an event - if no observation could change the decision and no date is nameable, the honest status is declined."
28
+ },
29
+ "target": {
30
+ "type": "array",
31
+ "minItems": 1,
32
+ "items": { "type": "string" },
33
+ "description": "Repo-relative paths the observation is about. Always a list, even for one path: a scalar here means every consumer needs two code paths, and the one that handles the scalar is the one that gets forgotten.",
34
+ "uniqueItems": true
35
+ },
36
+ "proposes_skill": {
37
+ "type": "array",
38
+ "items": { "type": "string" },
39
+ "description": "Skills or commands that do not exist yet but should, if this observation implies one."
40
+ },
41
+ "siblings_checked": {
42
+ "type": "string",
43
+ "minLength": 3,
44
+ "description": "What was found when the sibling surfaces were checked. This tree mirrors every command into skills/shared/core and again into the Copilot and Codex trees, so a fix applied to one copy drifts silently from the rest. 'checked, does not apply to the shared skill' is a valid answer; empty is not, which is why the field is required and why skill-siblings.mjs fills it mechanically rather than from memory."
45
+ },
46
+ "area": {
47
+ "type": "string",
48
+ "description": "Rough grouping for the weekly review: a phase name, a subsystem, 'gates', 'docs'."
49
+ },
50
+ "date": {
51
+ "type": "string",
52
+ "pattern": "^[0-9]{4}-[0-9]{2}-[0-9]{2}$",
53
+ "description": "ISO date the observation was written, absolute so it survives being read a year later."
54
+ },
55
+ "session_context": {
56
+ "type": "string",
57
+ "description": "What the session was doing when the friction appeared. Free text, kept short; it is what makes an old observation legible."
58
+ },
59
+ "resolved": {
60
+ "type": "string",
61
+ "pattern": "^[0-9]{4}-[0-9]{2}-[0-9]{2}$",
62
+ "description": "ISO date the status left 'open'."
63
+ },
64
+ "resolution": {
65
+ "type": "string",
66
+ "description": "One line saying what happened. Required in practice for declined and superseded - a decline with no reason is indistinguishable from neglect."
67
+ },
68
+ "reference": {
69
+ "type": "string",
70
+ "description": "A commit, PR, or issue that carried the change."
71
+ }
72
+ }
73
+ }