@whamp/pi-pstack 0.7.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +68 -0
- package/agents/comment-sicko.md +34 -0
- package/agents/poteto-agent.md +13 -0
- package/extensions/pstack/ask-user-question.ts +137 -0
- package/extensions/pstack/config.ts +55 -0
- package/extensions/pstack/index.ts +330 -0
- package/extensions/pstack/pstack-config-status.ts +72 -0
- package/extensions/pstack/pstack-role-config-store.ts +287 -0
- package/extensions/pstack/pstack-role-config.ts +496 -0
- package/extensions/pstack/pstack-role-prompt.ts +33 -0
- package/extensions/pstack/pstack-roles.ts +252 -0
- package/extensions/pstack/pstack-setup-plan.ts +223 -0
- package/extensions/pstack/skill-strip.ts +62 -0
- package/package.json +60 -0
- package/skills/architect/SKILL.md +83 -0
- package/skills/architect/references/design-red-flags.md +33 -0
- package/skills/architect/references/rationale-template.md +35 -0
- package/skills/architect/references/runner-prompt.md +20 -0
- package/skills/arena/SKILL.md +71 -0
- package/skills/automate-me/SKILL.md +104 -0
- package/skills/blast-radius/SKILL.md +50 -0
- package/skills/bro/SKILL.md +7 -0
- package/skills/create-verification-skill/SKILL.md +44 -0
- package/skills/create-verification-skill/references/feature-map-example/README.md +47 -0
- package/skills/create-verification-skill/references/feature-map-example/create-note.md +39 -0
- package/skills/create-verification-skill/references/feature-map-example/search.md +45 -0
- package/skills/deslop/SKILL.md +23 -0
- package/skills/figure-it-out/SKILL.md +53 -0
- package/skills/how/SKILL.md +52 -0
- package/skills/how/references/explainer-prompt.md +55 -0
- package/skills/how/references/explorer-prompt.md +52 -0
- package/skills/interrogate/SKILL.md +111 -0
- package/skills/interrogate/references/code-quality-review.md +47 -0
- package/skills/interrogate/references/lead-judgment.md +58 -0
- package/skills/interrogate/references/reviewer-prompt.md +72 -0
- package/skills/interrogate/references/rubric.md +77 -0
- package/skills/maintain-verification-skill/SKILL.md +39 -0
- package/skills/no-comments/SKILL.md +24 -0
- package/skills/poteto-mode/SKILL.md +147 -0
- package/skills/poteto-mode/playbooks/authoring-a-skill.md +12 -0
- package/skills/poteto-mode/playbooks/autonomous-run.md +13 -0
- package/skills/poteto-mode/playbooks/autopilot-full.md +13 -0
- package/skills/poteto-mode/playbooks/autopilot-stack.md +16 -0
- package/skills/poteto-mode/playbooks/babysit.md +27 -0
- package/skills/poteto-mode/playbooks/bug-fix.md +17 -0
- package/skills/poteto-mode/playbooks/eval.md +25 -0
- package/skills/poteto-mode/playbooks/feature.md +21 -0
- package/skills/poteto-mode/playbooks/hillclimb.md +21 -0
- package/skills/poteto-mode/playbooks/investigation.md +14 -0
- package/skills/poteto-mode/playbooks/multi-phase-plan.md +156 -0
- package/skills/poteto-mode/playbooks/opening-a-pr.md +33 -0
- package/skills/poteto-mode/playbooks/orchestrate.md +199 -0
- package/skills/poteto-mode/playbooks/pause-safely.md +10 -0
- package/skills/poteto-mode/playbooks/perf-issue.md +24 -0
- package/skills/poteto-mode/playbooks/prototype.md +14 -0
- package/skills/poteto-mode/playbooks/refactoring.md +16 -0
- package/skills/poteto-mode/playbooks/runtime-forensics.md +11 -0
- package/skills/poteto-mode/playbooks/session-pickup.md +11 -0
- package/skills/poteto-mode/playbooks/shipping.md +17 -0
- package/skills/poteto-mode/playbooks/trace-forensics.md +14 -0
- package/skills/poteto-mode/playbooks/visual-parity.md +11 -0
- package/skills/poteto-mode/playbooks/worktree-cleanup.md +14 -0
- package/skills/poteto-mode/references/bugbot-triage.md +142 -0
- package/skills/poteto-mode/scripts/bootstrap.ts +62 -0
- package/skills/poteto-mode/scripts/bun.lock +67 -0
- package/skills/poteto-mode/scripts/check-plan.mjs +186 -0
- package/skills/poteto-mode/scripts/check-plan.test.mjs +25 -0
- package/skills/poteto-mode/scripts/orch/orch.test.ts +634 -0
- package/skills/poteto-mode/scripts/orch/orch.ts +578 -0
- package/skills/poteto-mode/scripts/orch/store.ts +1607 -0
- package/skills/poteto-mode/scripts/package.json +16 -0
- package/skills/poteto-mode/scripts/watch-pr/cli.test.ts +224 -0
- package/skills/poteto-mode/scripts/watch-pr/cli.ts +223 -0
- package/skills/poteto-mode/scripts/watch-pr/fakes.test-helper.ts +118 -0
- package/skills/poteto-mode/scripts/watch-pr/github.test.ts +306 -0
- package/skills/poteto-mode/scripts/watch-pr/github.ts +699 -0
- package/skills/poteto-mode/scripts/watch-pr/policy.test.ts +420 -0
- package/skills/poteto-mode/scripts/watch-pr/policy.ts +832 -0
- package/skills/poteto-mode/scripts/watch-pr/render.ts +169 -0
- package/skills/poteto-mode/scripts/watch-pr/tsconfig.json +13 -0
- package/skills/poteto-mode/scripts/watch-pr/types.compile.ts +93 -0
- package/skills/poteto-mode/scripts/watch-pr/types.ts +401 -0
- package/skills/poteto-mode/scripts/watch-pr/watch-pr +6 -0
- package/skills/poteto-mode/scripts/worktree-audit.sh +85 -0
- package/skills/principle-attack-the-premise/SKILL.md +23 -0
- package/skills/principle-boundary-discipline/SKILL.md +34 -0
- package/skills/principle-build-the-lever/SKILL.md +23 -0
- package/skills/principle-encode-lessons-in-structure/SKILL.md +31 -0
- package/skills/principle-exhaust-the-design-space/SKILL.md +21 -0
- package/skills/principle-experience-first/SKILL.md +19 -0
- package/skills/principle-fix-root-causes/SKILL.md +23 -0
- package/skills/principle-foundational-thinking/SKILL.md +21 -0
- package/skills/principle-guard-the-context-window/SKILL.md +17 -0
- package/skills/principle-laziness-protocol/SKILL.md +18 -0
- package/skills/principle-make-operations-idempotent/SKILL.md +24 -0
- package/skills/principle-migrate-callers-then-delete-legacy-apis/SKILL.md +22 -0
- package/skills/principle-minimize-reader-load/SKILL.md +23 -0
- package/skills/principle-model-the-domain/SKILL.md +26 -0
- package/skills/principle-never-block-on-the-human/SKILL.md +22 -0
- package/skills/principle-outcome-oriented-execution/SKILL.md +22 -0
- package/skills/principle-prove-it-works/SKILL.md +33 -0
- package/skills/principle-redesign-from-first-principles/SKILL.md +16 -0
- package/skills/principle-separate-before-serializing-shared-state/SKILL.md +16 -0
- package/skills/principle-sequence-verifiable-units/SKILL.md +22 -0
- package/skills/principle-subtract-before-you-add/SKILL.md +21 -0
- package/skills/principle-test-behavior-not-implementation/SKILL.md +25 -0
- package/skills/principle-type-system-discipline/SKILL.md +31 -0
- package/skills/recall/SKILL.md +35 -0
- package/skills/reflect/SKILL.md +73 -0
- package/skills/reflect/references/divergent-reviewer.md +43 -0
- package/skills/reflect/references/judgment-reviewer.md +42 -0
- package/skills/reflect/references/synthesizer.md +56 -0
- package/skills/reflect/references/tooling-reviewer.md +55 -0
- package/skills/setup-pstack/SKILL.md +38 -0
- package/skills/setup-pstack/references/MODEL-ROLES.md +35 -0
- package/skills/show-me-your-work/SKILL.md +82 -0
- package/skills/show-me-your-work/references/decision-log-template.tsv +1 -0
- package/skills/show-me-your-work/scripts/log.sh +40 -0
- package/skills/swarm/SKILL.md +46 -0
- package/skills/tdd/SKILL.md +44 -0
- package/skills/teach/SKILL.md +21 -0
- package/skills/technical-writing/SKILL.md +127 -0
- package/skills/typescript-best-practices/SKILL.md +29 -0
- package/skills/typescript-best-practices/references/patterns.md +313 -0
- package/skills/unslop/SKILL.md +67 -0
- package/skills/why/SKILL.md +154 -0
- package/skills/why/references/epistemics.md +144 -0
- package/skills/why/references/investigator-prompt.md +103 -0
- package/skills/why/references/source-playbook.md +17 -0
- package/skills/why/references/sources/code-archaeology.md +88 -0
- package/skills/why/references/sources/databricks.md +70 -0
- package/skills/why/references/sources/datadog.md +99 -0
- package/skills/why/references/sources/incident-postmortem.md +15 -0
- package/skills/why/references/sources/linear.md +48 -0
- package/skills/why/references/sources/notion.md +55 -0
- package/skills/why/references/sources/sentry.md +100 -0
- package/skills/why/references/sources/slack.md +54 -0
- package/skills/why/references/synthesizer-prompt.md +135 -0
|
@@ -0,0 +1,154 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: why
|
|
3
|
+
description: "Use for 'why does X work this way', 'why we picked Y', design rationale, regressions, postmortems, or data-backed thresholds. Discovers available MCPs and queries each evidence category (source control, issue tracker, long-form docs, real-time chat, infrastructure observability, error tracking, product analytics warehouse) in parallel, then returns a cited read on decisions and tradeoffs. Use how for runtime behavior."
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Why
|
|
7
|
+
|
|
8
|
+
Investigate the motivation and intent behind code.
|
|
9
|
+
|
|
10
|
+
Companion to the `how` skill. `how` answers what the code does and how it works. `why` answers what forces led to its shape.
|
|
11
|
+
|
|
12
|
+
## Operating Posture
|
|
13
|
+
|
|
14
|
+
Operate as a **careful, cautious, and precise investigator**. Be honest about what you know vs what you're inferring. Read `references/epistemics.md` for the full confidence framework and phrasing guide. The synthesizer must follow it.
|
|
15
|
+
|
|
16
|
+
## Step 1. Understand the Target and the Question
|
|
17
|
+
|
|
18
|
+
Parse what the user is asking. The **target** is usually a chunk of code, a pattern, a feature, or a named design decision. The **question** is usually a design rationale, a tradeoff, a motivating edge case, an external constraint, dead code, or a broad history sweep.
|
|
19
|
+
|
|
20
|
+
If the target is vague ("why do we do it this way?" with no clear referent), make your best guess from conversation context (open files, recent edits, cursor location, what was just discussed). State your interpretation briefly so the user can redirect if you're off, then proceed.
|
|
21
|
+
|
|
22
|
+
## Step 2. Establish the Code Anchor
|
|
23
|
+
|
|
24
|
+
Before spawning investigators, anchor the investigation in concrete code. You need:
|
|
25
|
+
|
|
26
|
+
- The relevant file path(s) and line range(s)
|
|
27
|
+
- The key symbols (function names, class names, constants)
|
|
28
|
+
- An initial commit list. The last few commits touching the target.
|
|
29
|
+
- PR numbers from merge commits (pattern `(#1234)` in the subject line)
|
|
30
|
+
|
|
31
|
+
Build this inline.
|
|
32
|
+
|
|
33
|
+
```bash
|
|
34
|
+
# Blame target lines for last-touch commits
|
|
35
|
+
git blame -L <start>,<end> <file>
|
|
36
|
+
|
|
37
|
+
# Full file history, with patches, through renames
|
|
38
|
+
git log --follow -p -- <file>
|
|
39
|
+
|
|
40
|
+
# Last N commits touching the file, PR numbers visible
|
|
41
|
+
git log --oneline -20 -- <file>
|
|
42
|
+
|
|
43
|
+
# Extract PR numbers from a commit message
|
|
44
|
+
git log -1 --format=%B <commit>
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
Pull PR bodies and discussion via `gh` for any substantive commits:
|
|
48
|
+
|
|
49
|
+
```bash
|
|
50
|
+
gh pr view <number> --json title,body,author,createdAt,mergedAt,labels,closingIssuesReferences,comments,reviews
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
Capture this as seed context (file paths, symbols, commits, PR numbers, linked ticket IDs). Pass it to the investigators.
|
|
54
|
+
|
|
55
|
+
## Step 3. Spawn Parallel Investigators (default posture)
|
|
56
|
+
|
|
57
|
+
**Default to the full parallel investigation.**
|
|
58
|
+
|
|
59
|
+
### Discovery
|
|
60
|
+
|
|
61
|
+
Before spawning investigators, the parent lists its available MCP and extension tools. Map each available provider to one evidence category:
|
|
62
|
+
|
|
63
|
+
1. Source control history
|
|
64
|
+
2. Issue / ticket tracker
|
|
65
|
+
3. Long-form documents
|
|
66
|
+
4. Real-time team chat
|
|
67
|
+
5. Infrastructure observability
|
|
68
|
+
6. Error / exception tracking
|
|
69
|
+
7. Product analytics warehouse
|
|
70
|
+
|
|
71
|
+
Source control is always available through git and `gh`. For the other six, classify using the MCP name, server instructions, tool names, and resource descriptors. If an MCP could fit more than one category, choose the one matching its primary evidence. Record ambiguous cases in the coverage map.
|
|
72
|
+
|
|
73
|
+
Aim for a complete **coverage map**, not a minimal one. Document the null, don't skip the search. The parent queries each available MCP and builds one bounded evidence packet per category before launching children. A child does not inherit ambient MCP or extension tools. Use a custom agent for a child-side lookup only when that agent explicitly lists the tool and loads its provider through `extensions` or `subagentOnlyExtensions`.
|
|
74
|
+
|
|
75
|
+
Launch all matching investigators and the dependent synthesizer with one `subagent({ action: "execute", input: { async: true, maxSubagentSpawnsPerRun: N + 1, workflowScript } })` call. In `workflowScript`, await the investigators with `runs.all([{ key: "investigate-<category>", agent: "worker", task, model }])`, then return `runs.run("synthesize-why", { agent: "worker", task, model })` with their outputs. `N` is the number of evidence categories launched. Don't ask one agent to cover multiple categories.
|
|
76
|
+
|
|
77
|
+
Each investigator uses:
|
|
78
|
+
- agent: "worker"
|
|
79
|
+
- `model`: `why investigators` (default inherit-parent)
|
|
80
|
+
- `task`: instruct the investigator to inspect only
|
|
81
|
+
|
|
82
|
+
Each investigator gets:
|
|
83
|
+
1. The base prompt from `references/investigator-prompt.md`
|
|
84
|
+
2. The category playbook `references/sources/<source>.md` as an analysis rubric for the parent's evidence packet, not as child tool instructions
|
|
85
|
+
3. The parent's evidence packet for that category, including null results and gaps
|
|
86
|
+
4. The cross-cutting `references/sources/incident-postmortem.md` **if the target code looks defensive** (null checks, retry logic, timeout handling, rate limiting, feature flags, egress guards, OOM handlers)
|
|
87
|
+
5. The code anchor from Step 2 (file paths, symbols, commit hashes, PR numbers, ticket IDs)
|
|
88
|
+
6. The user's original question
|
|
89
|
+
|
|
90
|
+
### Investigator roster. One per available evidence category
|
|
91
|
+
|
|
92
|
+
Spawn one investigator per category with source-control evidence or a matching parent MCP. Each owns exactly one evidence packet.
|
|
93
|
+
|
|
94
|
+
Each entry names the category and the kind of "why" it uniquely surfaces. Use it to know what to expect back, how to name a gap when a category returns empty, and (only in the rare provably-irrelevant case) to justify a skip.
|
|
95
|
+
|
|
96
|
+
1. **Source control investigator**. Git history, `gh` for PRs, code comments, tests. Always spawn. The only guaranteed source. Best at surfacing *implementation-time rationale captured during review*.
|
|
97
|
+
|
|
98
|
+
2. **Issue / ticket tracker investigator** (e.g. Linear, Jira, GitHub Issues, Plane, Shortcut MCP). Best at surfacing *the product or business forcing function*. Strongest when the why is external to engineering.
|
|
99
|
+
|
|
100
|
+
3. **Long-form documents investigator** (e.g. Notion, Confluence, Google Docs, Coda MCP). Best at surfacing *long-form design rationale*. Where the why is written out before it becomes code.
|
|
101
|
+
|
|
102
|
+
4. **Real-time team chat investigator** (e.g. Slack, Discord, Microsoft Teams, Mattermost MCP). Best at surfacing *real-time deliberation that never reached a doc*. Especially important when the source control, ticket, and doc paper trail is thin.
|
|
103
|
+
|
|
104
|
+
5. **Infrastructure observability investigator** (e.g. Datadog, New Relic, Honeycomb, Grafana, Splunk MCP). Infra/runtime view. Best at surfacing *infrastructure and runtime reality that motivated the code*. Strongest when the target reacts to an infra signal (timeouts, retries, rate limits, circuit breakers).
|
|
105
|
+
|
|
106
|
+
6. **Error / exception tracking investigator** (e.g. Sentry, Rollbar, Bugsnag, Airbrake MCP). Best at surfacing *the specific exceptions and error trajectories that motivated defensive or corrective code*. Strongest for catch blocks, null guards, type checks, retries, and other defenses.
|
|
107
|
+
|
|
108
|
+
7. **Product analytics warehouse investigator** (e.g. Databricks, Snowflake, BigQuery, ClickHouse, dbt, Redshift MCP). Product/data view. Best at surfacing *product and data reality that shaped the code*. Strongest for flag-gated code, experiment-driven ships, data migrations, and "where did this number come from" questions.
|
|
109
|
+
|
|
110
|
+
### When to skip an investigator
|
|
111
|
+
|
|
112
|
+
Only skip with an **explicit, written justification** that goes in the final "Sources Consulted" section. Two valid reasons:
|
|
113
|
+
|
|
114
|
+
- **No MCP is available for that category** in this environment. Flag this as a gap, not a choice. Example: "Real-time team chat skipped. No matching MCP available, so the conversational record was not searchable."
|
|
115
|
+
- **The source is provably irrelevant**, not just "probably irrelevant." A high bar. Example: "Error / exception tracking skipped. Target is a build-time script with no runtime code path."
|
|
116
|
+
|
|
117
|
+
If your scope assessment suggests a single-commit trivial target where the PR description already contains the complete answer, you may answer inline **only after** confirming all seven available category searches would be redundant. Say so explicitly. This should be rare.
|
|
118
|
+
|
|
119
|
+
## Step 4. Synthesize
|
|
120
|
+
|
|
121
|
+
The same workflow launches `synthesize-why` after every investigator settles. It uses:
|
|
122
|
+
- agent: "worker"
|
|
123
|
+
- `model`: `why synthesizer` (default inherit-parent)
|
|
124
|
+
|
|
125
|
+
The synthesizer gets:
|
|
126
|
+
1. The investigator findings, including any null results and any categories skipped with justification
|
|
127
|
+
2. The code anchor from Step 2 (file paths, symbols, commit hashes, PR numbers, ticket IDs)
|
|
128
|
+
3. The user's original question
|
|
129
|
+
4. The epistemics framework from `references/epistemics.md`
|
|
130
|
+
5. The synthesizer prompt template from `references/synthesizer-prompt.md`
|
|
131
|
+
|
|
132
|
+
After the workflow completes, the parent spot-verifies citations with its own MCP and extension tools before presenting the result.
|
|
133
|
+
|
|
134
|
+
## Step 5. Present
|
|
135
|
+
|
|
136
|
+
Take the synthesizer's output and present it to the user. You may lightly edit for clarity or add context from the conversation, but **do not rewrite the confidence language**.
|
|
137
|
+
|
|
138
|
+
## Output Format
|
|
139
|
+
|
|
140
|
+
The output structure is the one in `references/synthesizer-prompt.md`: The Question, The Code in Question, What We Found, What We Can Reasonably Infer, Competing Hypotheses, What We Don't Know, Sources Consulted, Confidence Summary. Adapt as needed, but keep the confidence separation intact, and keep Sources Consulted as one line per investigator, including the ones that returned nothing or were skipped, with the reason.
|
|
141
|
+
|
|
142
|
+
After the Sources Consulted block, if the user's `why` question is a precursor to actually changing this code, convert the lineage findings into a Preserve / Change / Avoid / Risk constraint set suitable for planning the change.
|
|
143
|
+
|
|
144
|
+
## Common Failure Modes to Avoid
|
|
145
|
+
|
|
146
|
+
- **Recency bias**. Assuming the most recent commit is authoritative. The current shape is often the accretion of many earlier decisions. Trace back.
|
|
147
|
+
|
|
148
|
+
## Reference Files
|
|
149
|
+
|
|
150
|
+
- `references/epistemics.md`. Confidence tiers and phrasing guide. The synthesizer must follow it.
|
|
151
|
+
- `references/investigator-prompt.md`. Base prompt template for investigator subagents.
|
|
152
|
+
- `references/source-playbook.md`. Index pointing at the category playbooks below.
|
|
153
|
+
- `references/sources/*.md`. One self-contained example playbook per category, plus cross-cutting `incident-postmortem.md`. Give an investigator the single file that matches its category and adapt it to the available MCP.
|
|
154
|
+
- `references/synthesizer-prompt.md`. Prompt template for the synthesizer subagent, including the output format.
|
|
@@ -0,0 +1,144 @@
|
|
|
1
|
+
# Epistemics
|
|
2
|
+
|
|
3
|
+
How to reason about confidence when evidence is historical, fragmentary, and sometimes contradictory, and how to communicate it without flattening it into false certainty.
|
|
4
|
+
|
|
5
|
+
Code doesn't carry its own motivation. You can read what code does. You can't read *why it exists*. That lives in commits, PRs, tickets, docs, and conversations, all incomplete, biased, and sometimes missing entirely. Pretending otherwise produces confident-sounding guesses that mislead the user.
|
|
6
|
+
|
|
7
|
+
## Confidence Tiers
|
|
8
|
+
|
|
9
|
+
Every claim in the final output must sit in one of these tiers. The tier determines which output section the claim goes in and how it's phrased.
|
|
10
|
+
|
|
11
|
+
### 1. Direct
|
|
12
|
+
|
|
13
|
+
An explicit, textual citation that answers the question. Not "the code does X so the author must have wanted X." Something an author actually *wrote* that says why.
|
|
14
|
+
|
|
15
|
+
Examples:
|
|
16
|
+
- A PR description that says "this fixes the bug where users with >1000 items couldn't paginate"
|
|
17
|
+
- A ticket that says "we're adding this because customer Acme requested it in their security review"
|
|
18
|
+
- A code comment that says "// clamp to 100 because the upstream API rejects larger values"
|
|
19
|
+
- A design doc that says "we chose option A over option B because we need persistence across restarts"
|
|
20
|
+
- A chat message from the author saying "switching to this approach since the old one was flaky in tests"
|
|
21
|
+
|
|
22
|
+
Phrasing: confident, present tense. "This exists because X." Cite the source.
|
|
23
|
+
|
|
24
|
+
### 2. Supported
|
|
25
|
+
|
|
26
|
+
Multiple pieces of indirect evidence converge. No single source states it explicitly, but the pattern across sources makes it likely.
|
|
27
|
+
|
|
28
|
+
Examples:
|
|
29
|
+
- The PR title says "improve performance," the ticket is labeled "perf," and the surrounding commits all touch the same hot path
|
|
30
|
+
- Multiple tests were added alongside the change, all exercising edge cases with very large inputs
|
|
31
|
+
- The author's other PRs from the same week all mention the same incident in their descriptions
|
|
32
|
+
|
|
33
|
+
Phrasing: confident but clearly derived. "The evidence points strongly to X: [the specific pieces]." Cite multiple sources.
|
|
34
|
+
|
|
35
|
+
### 3. Inferred
|
|
36
|
+
|
|
37
|
+
A reasonable reading of the context, but nothing explicitly supports it. The reader should understand this is *your interpretation*, not a fact from the record.
|
|
38
|
+
|
|
39
|
+
Examples:
|
|
40
|
+
- The PR doesn't say why, but given the error was happening in production (per the incident channel timing) and the fix was rushed (merged the same day), it was likely a hotfix.
|
|
41
|
+
- The function name suggests retry logic. The retry count is 3. This matches the team's general convention of "3 retries" seen elsewhere in the codebase.
|
|
42
|
+
|
|
43
|
+
Phrasing: hedged. "It appears", "likely", "suggests", "is consistent with", "one reading is". Make the inference chain explicit: "Given A and B, C seems likely because D."
|
|
44
|
+
|
|
45
|
+
### 4. Speculative
|
|
46
|
+
|
|
47
|
+
A plausible hypothesis, but the evidence is thin and other explanations fit equally well. Presenting these is valuable, but mark them clearly as guesses.
|
|
48
|
+
|
|
49
|
+
Examples:
|
|
50
|
+
- "This might be a workaround for a browser bug that's since been fixed, but we found no contemporary evidence of that."
|
|
51
|
+
- "It's possible this threshold was chosen to match an SLA commitment, but no SLA doc references it."
|
|
52
|
+
|
|
53
|
+
Phrasing: explicitly speculative. "One possibility is X, but we have no direct evidence." Usually lives in the "Competing Hypotheses" section alongside other possibilities.
|
|
54
|
+
|
|
55
|
+
### 5. Unknown
|
|
56
|
+
|
|
57
|
+
You looked and couldn't find out. A valid and important outcome. Document it.
|
|
58
|
+
|
|
59
|
+
Phrasing: "We searched X, Y, and Z and found no evidence of why." Be specific about *what* you searched. "We couldn't find out" is less useful than "we searched the ticket tracker with keywords A and B, scanned the 6 PRs that touched this file since 2023, and grep'd the repo for string literals matching the threshold. None surfaced a rationale."
|
|
60
|
+
|
|
61
|
+
## Phrasing Guide
|
|
62
|
+
|
|
63
|
+
### Words that carry confidence. Use carefully
|
|
64
|
+
|
|
65
|
+
These imply **Direct** or **Supported** confidence. Don't use them for inferences.
|
|
66
|
+
|
|
67
|
+
- "because". Implies a causal claim with evidence
|
|
68
|
+
- "the reason is". Same
|
|
69
|
+
- "was designed to". Claims author intent
|
|
70
|
+
- "fixes", "addresses", "solves". Claims the change achieved its goal
|
|
71
|
+
- "the team decided". Claims a group decision happened
|
|
72
|
+
|
|
73
|
+
If you're using these, you should have a citation immediately adjacent.
|
|
74
|
+
|
|
75
|
+
### Words that hedge. Use for inferences
|
|
76
|
+
|
|
77
|
+
- "appears to"
|
|
78
|
+
- "seems to"
|
|
79
|
+
- "likely"
|
|
80
|
+
- "suggests"
|
|
81
|
+
- "is consistent with"
|
|
82
|
+
- "one reading is"
|
|
83
|
+
- "plausibly"
|
|
84
|
+
- "may have been"
|
|
85
|
+
- "the evidence points toward"
|
|
86
|
+
|
|
87
|
+
These signal that you're interpreting, not reporting. Use them liberally in the "What We Can Reasonably Infer" section.
|
|
88
|
+
|
|
89
|
+
### Words to avoid
|
|
90
|
+
|
|
91
|
+
- "obviously". If it were obvious, the user wouldn't be asking
|
|
92
|
+
- "clearly". Almost always precedes a claim that isn't clear
|
|
93
|
+
- "of course". Same
|
|
94
|
+
- "just" (as in "it's just X for performance"). Dismissive and usually hides uncertainty
|
|
95
|
+
- "I think" / "I believe". You're synthesizing evidence, not giving a personal opinion. Use "the evidence suggests" instead.
|
|
96
|
+
|
|
97
|
+
### Avoid rationalization
|
|
98
|
+
|
|
99
|
+
Code that "makes sense" today may have been written for reasons that no longer apply, or that were wrong when they were written. Don't retrofit a clean rationale onto messy history.
|
|
100
|
+
|
|
101
|
+
Resist the urge to:
|
|
102
|
+
- Assume the author did the "right" thing and work backward to justify it
|
|
103
|
+
- Assume a consistent pattern across the codebase was intentional when it might be copy-paste
|
|
104
|
+
- Turn an absence of evidence into evidence of absence ("no one mentioned security concerns, so it must not have been a concern")
|
|
105
|
+
|
|
106
|
+
## The Sycophancy Trap
|
|
107
|
+
|
|
108
|
+
Users often phrase `why` questions with an embedded hypothesis: "Why do we do it this way, I assume it's for performance?" Don't simply confirm it. Treat it as one candidate among others and check the evidence independently. If the evidence supports it, say so with citations. If not, say so and present what the evidence *does* support.
|
|
109
|
+
|
|
110
|
+
The user's guess is a prompt for investigation, not a conclusion to validate.
|
|
111
|
+
|
|
112
|
+
## When Evidence Contradicts
|
|
113
|
+
|
|
114
|
+
If two sources disagree (the PR description says one thing, the ticket says another), surface both. Don't pick the one that fits a tidier narrative. A typical pattern:
|
|
115
|
+
|
|
116
|
+
- **The ticket says** "we need this for customer X's compliance requirement"
|
|
117
|
+
- **The PR says** "cleaning up tech debt in this area"
|
|
118
|
+
|
|
119
|
+
Both may be true (the ticket motivated the work, the PR is the author's framing of it), or one may be wrong. Present both with their citations and let the user make the call.
|
|
120
|
+
|
|
121
|
+
## When Evidence Is Missing
|
|
122
|
+
|
|
123
|
+
An honest "we don't know" is one of the most valuable outputs this skill can produce. The user now knows:
|
|
124
|
+
|
|
125
|
+
- The answer isn't in the obvious places
|
|
126
|
+
- They'll need to ask a human (the original author, the product owner, the team lead) to find out
|
|
127
|
+
- Or they can decide the question isn't worth pursuing further
|
|
128
|
+
|
|
129
|
+
Failing to mark a gap and filling it with a confident guess actively harms the user. They'll act on the guess.
|
|
130
|
+
|
|
131
|
+
When you hit a gap, name it concretely:
|
|
132
|
+
- What question you were trying to answer
|
|
133
|
+
- What sources you searched
|
|
134
|
+
- What you searched for in each
|
|
135
|
+
- What you found (nothing, or only tangentially related material)
|
|
136
|
+
|
|
137
|
+
## Calibration Check Before Finalizing
|
|
138
|
+
|
|
139
|
+
Before delivering the output, the synthesizer should review every claim in "What We Found" and "What We Can Reasonably Infer" and ask:
|
|
140
|
+
|
|
141
|
+
1. Does this claim have a citation? If not, either add one or move it to "Inferred" / "Hypotheses".
|
|
142
|
+
2. Is the phrasing calibrated to the tier? (A Direct claim can use "because". An Inferred claim cannot.)
|
|
143
|
+
3. Am I treating the code itself as evidence for its own intent? If so, that's not evidence. Remove or reclassify.
|
|
144
|
+
4. Does the output include a "What We Don't Know" section? If no gaps are mentioned, that's suspicious. Either the evidence was unusually complete or something is being swept under the rug.
|
|
@@ -0,0 +1,103 @@
|
|
|
1
|
+
# Investigator Prompt Template
|
|
2
|
+
|
|
3
|
+
Build each investigator's prompt from this template. Fill in the placeholders. Append the single category playbook `sources/<source>.md` matching this investigator's evidence category (see `source-playbook.md` for the index). If the target code looks defensive (null checks, retry logic, timeout handling, rate limiting, feature flags, egress guards, OOM handlers), also append `sources/incident-postmortem.md` for the incident-flavored queries to run inside its own source.
|
|
4
|
+
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
You are investigating the historical context and motivation behind a piece of code. A separate synthesizer combines your findings with other investigators' into a final answer, so gather evidence accurately rather than writing prose.
|
|
8
|
+
|
|
9
|
+
Other investigators search different sources in parallel. Don't try to cover everything. Focus on your assigned source and go deep.
|
|
10
|
+
|
|
11
|
+
## Operating Posture
|
|
12
|
+
|
|
13
|
+
Work like a careful, cautious, precise investigator. Don't produce a narrative. Surface evidence and describe it accurately, including the parts that don't fit a tidy story. The more boring and exact your output, the more useful it is. A single verbatim quote with a precise citation beats a paragraph of plausible-sounding summary.
|
|
14
|
+
|
|
15
|
+
- **Quote, don't paraphrase** when the exact wording matters. Citations should let the reader jump to the source and confirm the claim in seconds.
|
|
16
|
+
- **Go wide before going deep.** Cast a broad first net so you don't miss related context. Only then narrow in.
|
|
17
|
+
- **Track what you searched, not just what you found.** An absence is only useful if the reader knows what was looked for. Record queries verbatim.
|
|
18
|
+
- **Resist the story.** If three pieces of evidence line up neatly and a fourth contradicts them, the contradiction is the most interesting finding. Don't file it away.
|
|
19
|
+
- **Consider the counterfactual.** Before reporting a finding as strong, ask whether you would expect to find it if your current reading were wrong, and how the evidence would differ.
|
|
20
|
+
- **Never invent.** If you're tempted to round a partial finding up into a confident statement, stop and label it partial. The synthesizer is counting on your output being accurate.
|
|
21
|
+
|
|
22
|
+
## The Question
|
|
23
|
+
|
|
24
|
+
> {QUESTION}
|
|
25
|
+
|
|
26
|
+
## The Code Anchor
|
|
27
|
+
|
|
28
|
+
**Target files:** {FILES_WITH_LINE_RANGES}
|
|
29
|
+
|
|
30
|
+
**Key symbols:** {SYMBOLS}
|
|
31
|
+
|
|
32
|
+
**Initial commits touching this code (most recent first):**
|
|
33
|
+
{COMMIT_LIST}
|
|
34
|
+
|
|
35
|
+
**PR numbers extracted from commit messages:** {PR_NUMBERS}
|
|
36
|
+
|
|
37
|
+
**Ticket IDs mentioned in commits or PR bodies (if any):** {TICKET_IDS}
|
|
38
|
+
|
|
39
|
+
## Your Assigned Source
|
|
40
|
+
|
|
41
|
+
{SOURCE_NAME}
|
|
42
|
+
|
|
43
|
+
{SOURCE_PLAYBOOK_SECTION}
|
|
44
|
+
|
|
45
|
+
## Investigation Instructions
|
|
46
|
+
|
|
47
|
+
Gather **evidence**. Don't answer the question directly. The synthesizer weighs the evidence and forms conclusions. Follow this loop:
|
|
48
|
+
|
|
49
|
+
1. **Cast a wide net first.** Start broad so you don't miss related context, then narrow in on specific items.
|
|
50
|
+
2. **Read the whole thing.** Read any PR, ticket, doc, or thread fully, not just the title or summary. The key evidence is often buried in a comment, a subtask, or a follow-up.
|
|
51
|
+
3. **Follow links within your assigned source.** If a PR references another PR or commit, pull it. If a ticket links a parent or sibling, pull it. If a doc links another doc, pull it. Stay inside your assigned source. When you spot a cross-source reference, do NOT chase it yourself. Record it under "Additional Leads" so the investigator assigned to that source can pick it up. The one-investigator-per-category design depends on this. Chasing cross-source links duplicates work and confuses scope.
|
|
52
|
+
4. **Capture quotes verbatim** with their location (PR number, ticket ID, URL, commit hash, file:line). The synthesizer needs to cite this precisely.
|
|
53
|
+
5. **Note absences.** If you searched for something and came up empty, that's also a finding. Record what you searched for and what you didn't find.
|
|
54
|
+
6. **Watch for contradictions.** If two items in your source disagree, record both. Don't suppress the inconvenient one.
|
|
55
|
+
|
|
56
|
+
Don't synthesize or form a final opinion on "the why." Collect the raw material honestly and completely. The synthesizer does the reasoning.
|
|
57
|
+
|
|
58
|
+
## Epistemic Discipline
|
|
59
|
+
|
|
60
|
+
- **Don't confuse mechanics with motivation.** A commit changing `limit = 50` to `limit = 100` shows the change, not necessarily why. Look for the explanation in the commit message, PR description, linked ticket, or review comments.
|
|
61
|
+
- **Don't infer intent from code style.** "The author chose a functional approach" is an observation about code, not evidence of intent. Claim intent only when the author stated it.
|
|
62
|
+
- **Preserve uncertainty.** If the evidence is ambiguous, say so. If one reading is more plausible but not certain, say that. Don't collapse ambiguity to look decisive.
|
|
63
|
+
- **No silent substitutions.** If the question is about feature X and you only find evidence about feature Y, don't present Y's evidence as if it answers X.
|
|
64
|
+
|
|
65
|
+
## Output Format
|
|
66
|
+
|
|
67
|
+
Return your findings in this structure. The synthesizer will read it directly.
|
|
68
|
+
|
|
69
|
+
### Source
|
|
70
|
+
Which source you investigated (source control, issue / ticket tracker, long-form documents, real-time team chat, infrastructure observability, error / exception tracking, product analytics warehouse, code comments, etc.).
|
|
71
|
+
|
|
72
|
+
### What I Searched
|
|
73
|
+
The queries you ran, the items you opened, the places you looked. Be specific. This tells the synthesizer how thorough the investigation was and what might still be unsearched.
|
|
74
|
+
|
|
75
|
+
### Direct Evidence Found
|
|
76
|
+
For each piece that explicitly addresses the question:
|
|
77
|
+
- **What it says**: verbatim quote or accurate paraphrase
|
|
78
|
+
- **Where it's from**: PR #123, ticket ID, doc URL, chat permalink, commit hash, or file:line
|
|
79
|
+
- **Author and date** (if available)
|
|
80
|
+
- **Relevance**: one sentence on how it bears on the question
|
|
81
|
+
|
|
82
|
+
### Indirect / Circumstantial Evidence
|
|
83
|
+
Items that don't explicitly answer the question but bear on it. For each:
|
|
84
|
+
- **What it is**: brief description
|
|
85
|
+
- **Where it's from**: location
|
|
86
|
+
- **What it suggests**: what a careful reader might infer, and why. Name the inference chain.
|
|
87
|
+
- **Alternative readings**: if the same evidence could support a different interpretation, note it
|
|
88
|
+
|
|
89
|
+
### Contradictions
|
|
90
|
+
Two items that disagree with each other, with both citations.
|
|
91
|
+
|
|
92
|
+
### Gaps
|
|
93
|
+
What you searched for and didn't find. Be specific: "Searched the issue tracker for [query] across [time range]. No matching issues." These absences are valuable data.
|
|
94
|
+
|
|
95
|
+
### Additional Leads
|
|
96
|
+
Anything that suggests further investigation in a different source. For example, if a PR references a chat thread that wasn't in your source, note it so the real-time team chat investigator or a follow-up pass can pursue it.
|
|
97
|
+
|
|
98
|
+
## What You're Not Doing
|
|
99
|
+
|
|
100
|
+
- Writing the final answer. The synthesizer does that.
|
|
101
|
+
- Picking sides in contradictions. Surface them.
|
|
102
|
+
- Speculating beyond what the evidence supports. A hunch with no evidence isn't evidence.
|
|
103
|
+
- Reading the code itself to figure out intent. You may read the code to understand what the target *is*, but don't confuse "what the code does" with "why."
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
# Source playbooks
|
|
2
|
+
|
|
3
|
+
The why skill spawns one investigator per available evidence category, each reading a single source-specific playbook below. The playbooks are concrete examples for common MCPs. Adapt them for a different MCP in the same category.
|
|
4
|
+
|
|
5
|
+
| Category | Playbook | Example MCP it documents |
|
|
6
|
+
|---|---|---|
|
|
7
|
+
| Source control history | [`code-archaeology.md`](./sources/code-archaeology.md) | git, `gh` |
|
|
8
|
+
| Issue / ticket tracker | [`linear.md`](./sources/linear.md) | Linear (adapt for Jira, GitHub Issues, Plane, Shortcut) |
|
|
9
|
+
| Long-form documents | [`notion.md`](./sources/notion.md) | Notion (adapt for Confluence, Google Docs, Coda) |
|
|
10
|
+
| Real-time team chat | [`slack.md`](./sources/slack.md) | Slack (adapt for Discord, Microsoft Teams, Mattermost) |
|
|
11
|
+
| Infrastructure observability | [`datadog.md`](./sources/datadog.md) | Datadog (adapt for New Relic, Honeycomb, Grafana, Splunk) |
|
|
12
|
+
| Error / exception tracking | [`sentry.md`](./sources/sentry.md) | Sentry (adapt for Rollbar, Bugsnag, Airbrake) |
|
|
13
|
+
| Product analytics warehouse | [`databricks.md`](./sources/databricks.md) | Databricks SQL (adapt for Snowflake, BigQuery, ClickHouse, dbt) |
|
|
14
|
+
|
|
15
|
+
Cross-cutting:
|
|
16
|
+
|
|
17
|
+
- [`incident-postmortem.md`](./sources/incident-postmortem.md). Add this if the target code looks defensive (null checks, retry, timeout, rate limit, feature flag, egress guard, OOM handler).
|
|
@@ -0,0 +1,88 @@
|
|
|
1
|
+
# Code Archaeology (git + in-repo)
|
|
2
|
+
|
|
3
|
+
## What this source contains
|
|
4
|
+
|
|
5
|
+
- Commit history (messages, dates, authors, diffs)
|
|
6
|
+
- PR descriptions, review comments, and discussion threads (via `gh`)
|
|
7
|
+
- Inline code comments, TODOs, FIXMEs, deprecation notes
|
|
8
|
+
- ADRs (architectural decision records) if the repo keeps them
|
|
9
|
+
- Tests. Names and assertions often encode the edge cases that motivated a change
|
|
10
|
+
- Related files modified in the same commits (co-change signal)
|
|
11
|
+
- CHANGELOG entries, release notes in the repo
|
|
12
|
+
- Issue/ticket IDs mentioned in commit messages and PR bodies
|
|
13
|
+
|
|
14
|
+
The most trustworthy source, tied directly to the code, and the most complete. Everything that went through the repo should be here.
|
|
15
|
+
|
|
16
|
+
## How to search it
|
|
17
|
+
|
|
18
|
+
Expand the seed commit list:
|
|
19
|
+
|
|
20
|
+
```bash
|
|
21
|
+
# Full history of the file through renames
|
|
22
|
+
git log --follow --oneline -- <file>
|
|
23
|
+
|
|
24
|
+
# Pickaxe: commits that added or removed this exact text
|
|
25
|
+
git log -S '<exact_string_from_code>' -- <file>
|
|
26
|
+
|
|
27
|
+
# Or for patterns:
|
|
28
|
+
git log -G '<regex>' -- <file>
|
|
29
|
+
|
|
30
|
+
# Who wrote each line and when
|
|
31
|
+
git blame -L <start>,<end> <file>
|
|
32
|
+
|
|
33
|
+
# The full diff of a specific commit
|
|
34
|
+
git show <hash>
|
|
35
|
+
|
|
36
|
+
# Commits between two points affecting this file
|
|
37
|
+
git log <old>..<new> -p -- <file>
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
For each substantive commit, pull the PR context:
|
|
41
|
+
|
|
42
|
+
```bash
|
|
43
|
+
# Find the PR number from the merge commit or branch
|
|
44
|
+
git log -1 --format=%B <hash>
|
|
45
|
+
|
|
46
|
+
# Full PR context: body, review comments, linked issues
|
|
47
|
+
gh pr view <number> --json title,body,author,createdAt,mergedAt,labels,closingIssuesReferences,comments,reviews,files
|
|
48
|
+
|
|
49
|
+
# The --json reviews and comments fields are where the real signal is
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
Look for out-of-band docs:
|
|
53
|
+
|
|
54
|
+
```bash
|
|
55
|
+
# ADRs often live in docs/adr/ or similar
|
|
56
|
+
rg -l -i 'architecture.decision' --glob '*.md'
|
|
57
|
+
|
|
58
|
+
# TODOs and FIXMEs near the target
|
|
59
|
+
rg -n -C2 '(TODO|FIXME|HACK|XXX|NOTE)' <target_file>
|
|
60
|
+
|
|
61
|
+
# Related tests. Names often encode the "why"
|
|
62
|
+
rg -l '<symbol>' --glob '*test*'
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
## What good evidence looks like here
|
|
66
|
+
|
|
67
|
+
- A PR description that explains the problem being solved, not just the change ("This fixes the pagination bug that caused X")
|
|
68
|
+
- A long review thread where alternatives were debated
|
|
69
|
+
- An inline comment near the target line that explains a non-obvious constraint
|
|
70
|
+
- A test named `test_handles_edge_case_when_X` that reveals an edge case motivating the code
|
|
71
|
+
- A commit message that references a ticket or incident ID
|
|
72
|
+
- A CHANGELOG entry that summarizes the user-visible rationale
|
|
73
|
+
|
|
74
|
+
## Common pitfalls
|
|
75
|
+
|
|
76
|
+
- **Squash-merge flatlands.** If the repo squashes PRs, individual commits in the branch history are lost. Fall back to PR body and comments.
|
|
77
|
+
- **Misleading commit messages.** "Small refactor" sometimes hides an intentional behavior change. Look at the diff, not the message.
|
|
78
|
+
- **Cargo-culted patterns.** The author may have copied a pattern without understanding why. Check if the pattern originated earlier in the codebase and investigate *that* commit.
|
|
79
|
+
- **Bot commits and auto-merges.** Dependabot, Renovate, and automated backports usually don't carry motivation. Skip them when trying to find intent.
|
|
80
|
+
- **Treating code as evidence of intent.** The code itself isn't evidence for why it exists. Evidence comes from commit messages, PRs, comments, tests, docs. Don't cite "the function is named X" as evidence of intent.
|
|
81
|
+
|
|
82
|
+
## What to return
|
|
83
|
+
|
|
84
|
+
Every commit/PR/comment that bears on the question, with:
|
|
85
|
+
- The exact text (quoted)
|
|
86
|
+
- The hash / PR number / file:line
|
|
87
|
+
- Author and date
|
|
88
|
+
- Whether it's direct (explicitly addresses the question) or circumstantial
|
|
@@ -0,0 +1,70 @@
|
|
|
1
|
+
# Databricks Analytics & System Tables
|
|
2
|
+
|
|
3
|
+
## What this source contains
|
|
4
|
+
|
|
5
|
+
Databricks is the product-analytics, data-pipeline, and warehouse-telemetry layer. It complements Datadog. Datadog is the *infra/runtime* view, Databricks is the *product/data* view (what users did, which experiments ran, how feature usage evolved, where a threshold constant came from).
|
|
6
|
+
|
|
7
|
+
- **Product analytics events.** `your_warehouse.events.analytics_track_event` (raw) and typed, deduplicated per-event dbt models in `<your_analytics_db>.<schema>.<table>`. User behavior: feature invocations, clicks, accepts/rejects, submissions, client-reported errors.
|
|
8
|
+
- **Usage & billing events.** `your_warehouse.events.usage_event` / `<your_analytics_db>.<schema>.stg_usage_events`, `your_warehouse.events.raw_model_event` / `<your_analytics_db>.<schema>.stg_raw_model_events`. For cost- or volume-driven decisions.
|
|
9
|
+
- **Experiment / feature-flag data.** Exposure and outcome tables. **Schema is company-specific.** Probe with `SHOW TABLES` before assuming names.
|
|
10
|
+
- **System tables.** `system.query.history`, `system.compute.warehouses`, `system.billing.*`, `system.access.audit`. Answer "was this query expensive?", "how often did anyone run this?", "when did warehouse load spike?"
|
|
11
|
+
- **dbt lineage.** Models in `<your_analytics_db>.<schema>` reveal what pipelines depend on a table/field. Upstream changes frequently motivate consumer-code changes.
|
|
12
|
+
- **Databricks notebooks.** Exploratory analyses engineers wrote before code changes. **Not queryable via the SQL MCP.** If you suspect the rationale lives in a notebook, name it as a gap.
|
|
13
|
+
|
|
14
|
+
## How to search it
|
|
15
|
+
|
|
16
|
+
Use the Databricks SQL MCP. Primary tool: `execute_sql_read_only`. If it returns a `statement_id`, poll with `poll_sql_result` rather than re-running.
|
|
17
|
+
|
|
18
|
+
**Orient before querying.** Schemas are company-specific. Probe before trusting a table name:
|
|
19
|
+
|
|
20
|
+
```sql
|
|
21
|
+
SHOW TABLES IN <your_analytics_db>.<schema> LIKE '*<keyword>*';
|
|
22
|
+
DESCRIBE TABLE <your_analytics_db>.<schema>.stg_<event>;
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
**Time-bound every query.** These tables are huge and unconstrained scans time out. Filter on `_timestamp` (events) or `start_time` (`system.query.history`) with a window bracketing the ship date, typically ~30 days before and after, wider only for strong reason.
|
|
26
|
+
|
|
27
|
+
**Prefer typed dbt models over the raw table.** `<your_analytics_db>.<schema>.<table>` is deduplicated, typed, and liquid-clustered. `your_warehouse.events.analytics_track_event` has duplicates and untyped `properties_json`. Model-name pattern: `stg_<source>_<event_name_with_underscores>`, where `<source>` is `app`, `backend`, `website`, or `cli`. Confirm the exact model name with `SHOW TABLES` when the pattern alone doesn't resolve it. Drop to the raw table only when there's no dbt model yet, or you need events from inside the dbt refresh lag.
|
|
28
|
+
|
|
29
|
+
**Column conventions on the typed dbt models** (knowing these avoids a `DESCRIBE` round-trip):
|
|
30
|
+
|
|
31
|
+
- `_timestamp`, `_id`, `_auth_id`, `_request_id`, `event_name`. Standard on every model
|
|
32
|
+
- `properties_<name>`. Typed, underscore-cased event properties (`properties_entrypoint`, `properties_size_bytes`, …)
|
|
33
|
+
- `context_team_id`, `context_client_version`, `context_country`, `context_client_os`. Pre-extracted client context
|
|
34
|
+
|
|
35
|
+
### Investigation patterns that tend to pay off
|
|
36
|
+
|
|
37
|
+
Pick the table + column combination that matches the target:
|
|
38
|
+
|
|
39
|
+
1. **Event usage trajectory.** Daily counts on the relevant `stg_*` model across a ±30d window around the PR merge. A step function from zero to steady volume within a day or two of the merge is strong circumstantial evidence the PR launched the feature. A decay to zero suggests a deprecation or deletion.
|
|
40
|
+
2. **Guard-rail / defensive-check origin.** Distribution (median / p99 / max) of the relevant `properties_<name>` column in the 14 days *before* the PR. A p99 that matches the target's threshold constant suggests the number was chosen from data.
|
|
41
|
+
3. **Experiment / feature-flag lookup.** `SHOW TABLES ... LIKE '*experiment*'` to find the exposure table, then pull exposure counts by variant for the relevant flag key near the PR date.
|
|
42
|
+
4. **Query-history evidence for migrations, backfills, or perf rewrites.** `system.query.history` filtered by `statement_text ILIKE '%<table_or_symbol>%'` with a tight `start_time` window surfaces the expensive queries that likely motivated the change (sort by `total_duration_ms` or aggregate `SUM(read_bytes)`, `COUNT(*)`).
|
|
43
|
+
5. **dbt lineage.** If the target reads from or writes into a `<your_analytics_db>.<schema>` model, the model's own git history (in this repo) often carries the rationale. Hand that lead back to the git investigator rather than chasing it yourself.
|
|
44
|
+
|
|
45
|
+
## What good evidence looks like here
|
|
46
|
+
|
|
47
|
+
Beyond the pattern shapes above:
|
|
48
|
+
|
|
49
|
+
- An error-classifying event's count drops to near zero in the days after a defensive-code PR. Suggests the PR resolved that error class
|
|
50
|
+
- An exposure table row names the target's feature-flag key with a "shipped" / "concluded" decision around the PR ship date
|
|
51
|
+
|
|
52
|
+
## Common pitfalls
|
|
53
|
+
|
|
54
|
+
- **Instrumented ≠ caused.** An event's existence means someone cared enough to log it, not that the target code exists *because* of it. Pair with a PR/commit citation from the git investigator before claiming causation.
|
|
55
|
+
- **Silent instrumentation changes.** A step function in event volume may mean a new event started being logged, not that user behavior changed. Check for instrumentation PRs in the same window before reading the ramp as a feature-launch signal.
|
|
56
|
+
- **Schema drift.** Event properties evolve. A column on the typed dbt model today may not have existed when the target was written. Older data may carry the property only inside raw `properties_json`.
|
|
57
|
+
- **dbt refresh lag.** `<your_analytics_db>.<schema>.*` is rebuilt on a schedule (often hourly/daily). For events from the last few hours, fall back to `your_warehouse.events.*` and deduplicate by `_id`.
|
|
58
|
+
- **Company-specific tables.** Experiment, feature-flag, billing, and usage tables vary. Reporting a result from a table whose existence you never confirmed is a classic failure mode. Probe with `SHOW TABLES` / `DESCRIBE TABLE` first.
|
|
59
|
+
- **Retention cliff.** If the relevant window predates the table's retention or the dbt model's creation date, that's a *gap*, not a null result. Name it explicitly so the synthesizer doesn't read "no results" as "no activity."
|
|
60
|
+
- **Notebooks aren't queryable.** The SQL MCP can't see Databricks notebooks. If you suspect the rationale lives in one, return a gap.
|
|
61
|
+
|
|
62
|
+
## What to return
|
|
63
|
+
|
|
64
|
+
For each relevant finding:
|
|
65
|
+
- Type (product event / experiment exposure / usage or billing event / system-table row / dbt model)
|
|
66
|
+
- Fully-qualified table name and the exact query you ran
|
|
67
|
+
- Time window queried
|
|
68
|
+
- Compact numeric summary (counts, percentiles, first/last-seen timestamps). **Don't dump raw rows.**
|
|
69
|
+
- Temporal correlation with the target's ship date (e.g., "first row 2024-08-15, PR #49074 merged 2024-08-14")
|
|
70
|
+
- Relevance + strength: direct / circumstantial / weak
|