@metabase/cli 0.1.10 → 0.1.12

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (159) hide show
  1. package/.claude-plugin/marketplace.json +1 -1
  2. package/README.md +9 -6
  3. package/dist/add-collection-BPtBMh8Y.mjs +10 -0
  4. package/dist/{add-collection-DUqTrC5T.mjs → add-collection-DEME4IOy.mjs} +4 -4
  5. package/dist/{archive-_GMNY8wH.mjs → archive-BFk-oupo.mjs} +3 -3
  6. package/dist/{archive-44EWiXud.mjs → archive-BNioNJtG.mjs} +3 -3
  7. package/dist/{archive-BZfpjMir.mjs → archive-BtDvBsr8.mjs} +3 -3
  8. package/dist/{archive-0krZxAXq.mjs → archive-CMeTr8jv.mjs} +3 -3
  9. package/dist/{archive-BEIyIsin.mjs → archive-DUlrUlok.mjs} +3 -3
  10. package/dist/{archive-DtE2H4A6.mjs → archive-Dc-FwNZm.mjs} +3 -3
  11. package/dist/{archive-B59Y7ajB.mjs → archive-MTTQv3wR.mjs} +3 -3
  12. package/dist/auth-CjFnuPe9.mjs +19 -0
  13. package/dist/{body-BdRyuvU4.mjs → body-89w3r7In.mjs} +2 -2
  14. package/dist/{branches-Jpv-FNds.mjs → branches-DuVKH8Ye.mjs} +4 -4
  15. package/dist/{cancel-ChC4lFd4.mjs → cancel-BEMQKVuR.mjs} +3 -3
  16. package/dist/{cancel-task-DyhIkNaL.mjs → cancel-task-D2mSPFXO.mjs} +4 -4
  17. package/dist/card-DoDpwR4t.mjs +20 -0
  18. package/dist/{cards-O-nKkQKP.mjs → cards-L9jksUMD.mjs} +3 -3
  19. package/dist/cli.mjs +23 -23
  20. package/dist/collection-CF-1RIPY.mjs +20 -0
  21. package/dist/{collection-namespace-7724zUMx.mjs → collection-namespace-w26SavJH.mjs} +1 -1
  22. package/dist/{create-D45uFXlo.mjs → create-8hMjPz7j.mjs} +4 -4
  23. package/dist/{create-B4f4Pldw.mjs → create-Bu9Jayck.mjs} +5 -5
  24. package/dist/{create-BxRsXQrm.mjs → create-C3b3MoIN.mjs} +4 -4
  25. package/dist/{create-DeZ2x2Db.mjs → create-Cxg_h8Kf.mjs} +4 -4
  26. package/dist/{create-DS52EhPd.mjs → create-D6QqOX6u.mjs} +4 -4
  27. package/dist/{create-BbF9zVFP.mjs → create-DCYc032c.mjs} +5 -5
  28. package/dist/{create-BnFHcnlL.mjs → create-Dew3rKsR.mjs} +4 -4
  29. package/dist/{create-CvYKJOcE.mjs → create-DoeqRVmx.mjs} +4 -4
  30. package/dist/{create-branch-KOWUIE72.mjs → create-branch-C5P2y8U5.mjs} +4 -4
  31. package/dist/{create-BuKx7kw6.mjs → create-zEq4ZJAT.mjs} +4 -4
  32. package/dist/{current-task-ClcWPPMc.mjs → current-task-uD4QO5bM.mjs} +4 -4
  33. package/dist/dashboard-C5zTX345.mjs +21 -0
  34. package/dist/db-4y5Z1CQw.mjs +22 -0
  35. package/dist/{delete-D68oS73R.mjs → delete-Bftif2R3.mjs} +3 -3
  36. package/dist/{delete-CX2VUA5R.mjs → delete-DGaNAFdQ.mjs} +3 -3
  37. package/dist/{delete-table-DLvL9mDA.mjs → delete-table-C0qxP6FE.mjs} +3 -3
  38. package/dist/{dirty-1OrXpc7E.mjs → dirty-CZZLpWt1.mjs} +4 -4
  39. package/dist/document-97l9Non_.mjs +19 -0
  40. package/dist/{eid-CzLhHZMW.mjs → eid-C8xTmR8e.mjs} +4 -4
  41. package/dist/{error-BWXBhqLW.mjs → error-BGH4gxGN.mjs} +2 -1
  42. package/dist/{export-B5z8w-xo.mjs → export-DWdjnBKo.mjs} +6 -6
  43. package/dist/field-BwUTrelF.mjs +18 -0
  44. package/dist/{fields-sC7pzmPX.mjs → fields-Cq08PQec.mjs} +3 -3
  45. package/dist/{get-CHb6J908.mjs → get-7IhxPyE_.mjs} +2 -2
  46. package/dist/{get-BYw3xS0X.mjs → get-8nAjRL2H.mjs} +3 -3
  47. package/dist/{get-ByZ4HR2T.mjs → get-B2vGW7Aw.mjs} +3 -3
  48. package/dist/{get-BC60bhel.mjs → get-B_BJ10pV.mjs} +3 -3
  49. package/dist/{get-DLwb_gUh.mjs → get-BaKKF50e.mjs} +3 -3
  50. package/dist/{get-Dl62Fy6Y.mjs → get-Bbij6RjC.mjs} +3 -3
  51. package/dist/{get-reSMTfQi.mjs → get-Bcr_5fAb.mjs} +2 -2
  52. package/dist/{get-lcX52Skc.mjs → get-Cr8oz_jr.mjs} +3 -3
  53. package/dist/{get-DTHLETau.mjs → get-DHPN1716.mjs} +3 -3
  54. package/dist/{get-C6n86-dS.mjs → get-DtKv-gl3.mjs} +3 -3
  55. package/dist/{get-ZesERdyk.mjs → get-DzLAmSKd.mjs} +2 -2
  56. package/dist/{get-Bo1FGyFs.mjs → get-Rqz7UHv1.mjs} +3 -3
  57. package/dist/{get-4GEDd9YN.mjs → get-f5ENeUNw.mjs} +3 -3
  58. package/dist/{get-run-CpCbHJad.mjs → get-run-DhbHXaZE.mjs} +3 -3
  59. package/dist/{get-Fn9WkNhS.mjs → get-ug26gcOI.mjs} +3 -3
  60. package/dist/git-sync-JIMEYW-E.mjs +28 -0
  61. package/dist/{has-remote-changes-CPz_-uxd.mjs → has-remote-changes-CTLh3WjZ.mjs} +4 -4
  62. package/dist/{import-DaWgprK6.mjs → import-XakIQfRp.mjs} +6 -6
  63. package/dist/{input-xewHccej.mjs → input-BZwm4yby.mjs} +25 -3
  64. package/dist/is-dirty-4gf8ujws.mjs +9 -0
  65. package/dist/{is-dirty-DZlI7lQx.mjs → is-dirty-DkOOyjsZ.mjs} +3 -3
  66. package/dist/{items-hgbRYsYD.mjs → items-DC985V8K.mjs} +3 -3
  67. package/dist/{list-BgHESP7b.mjs → list-B83-hnq5.mjs} +2 -2
  68. package/dist/{list-BGerkRHH.mjs → list-BtNboYpZ.mjs} +3 -3
  69. package/dist/{list-BWv5y307.mjs → list-CEbMS6g3.mjs} +2 -2
  70. package/dist/{list-CKeVS1IZ.mjs → list-CIEI8XK_.mjs} +2 -2
  71. package/dist/{list-Cd2nOCAx.mjs → list-Cnb9mQdd.mjs} +2 -2
  72. package/dist/{list-B9J3ujwn.mjs → list-CtAFlkLe.mjs} +2 -2
  73. package/dist/{list-36H-dvJZ.mjs → list-CyIgwqFB.mjs} +2 -2
  74. package/dist/{list-CGdOC9zX.mjs → list-Df4MhR2A.mjs} +2 -2
  75. package/dist/{list-DO8T-nmF.mjs → list-DfPwBQrZ.mjs} +2 -2
  76. package/dist/{list-DBlsRSpZ.mjs → list-OGkVAG6b.mjs} +2 -2
  77. package/dist/{list-DOVX3vCb.mjs → list-PYNFVJGM.mjs} +3 -3
  78. package/dist/{list-kath_2cX.mjs → list-STlBcPs8.mjs} +2 -2
  79. package/dist/{list-CCZnH2-Z.mjs → list-gy64AAWO.mjs} +2 -2
  80. package/dist/{list-D9EuxFHO.mjs → list-xfmTxHNu.mjs} +2 -2
  81. package/dist/{login-BzrAGJfu.mjs → login-C0VFlX8y.mjs} +4 -4
  82. package/dist/{logout-CKBiltoS.mjs → logout-BUgUTFol.mjs} +2 -2
  83. package/dist/measure-BmcmxRQ3.mjs +19 -0
  84. package/dist/{metadata-BcGcUEVJ.mjs → metadata-B1EmV1l5.mjs} +3 -3
  85. package/dist/{metadata-BjOrKtnv.mjs → metadata-Dwkcsa1c.mjs} +3 -3
  86. package/dist/{parse-id--iVTCKSo.mjs → parse-id-sKNP6yu5.mjs} +1 -1
  87. package/dist/{path-LgGU6Bd0.mjs → path-BZJR7YLR.mjs} +2 -2
  88. package/dist/{poll-BucRFJT-.mjs → poll-Cyw-dQi4.mjs} +1 -1
  89. package/dist/{poll-task-DTzKB3T3.mjs → poll-task-BG0oVnfZ.mjs} +2 -2
  90. package/dist/{preflight-CzqVX0PP.mjs → preflight-BjR63JnX.mjs} +1 -1
  91. package/dist/{query-BVZkK6Qk.mjs → query-C2xqnz24.mjs} +3 -3
  92. package/dist/{query-DG_jygDF.mjs → query-Dk55Wp2p.mjs} +4 -4
  93. package/dist/{remove-collection-HdeAfLyi.mjs → remove-collection-DbVNWGyb.mjs} +6 -6
  94. package/dist/{rescan-values-hWubCruZ.mjs → rescan-values-CuLQ9UHP.mjs} +3 -3
  95. package/dist/{run-C1-lDmQF.mjs → run-PDN33QjN.mjs} +5 -5
  96. package/dist/{runs-Q6DYQyqj.mjs → runs-B9T-7lei.mjs} +3 -3
  97. package/dist/{runtime-colqvhLf.mjs → runtime-DgHh4T6t.mjs} +1 -1
  98. package/dist/{schema-tables-BEastV_8.mjs → schema-tables-DjmNwrlT.mjs} +3 -3
  99. package/dist/{schemas-CgawwI_k.mjs → schemas-CWL8gv1C.mjs} +3 -3
  100. package/dist/{search-BWo7xSPP.mjs → search-D7JzmvfR.mjs} +3 -3
  101. package/dist/segment-DrZNuxsM.mjs +19 -0
  102. package/dist/{set-7Nm2ZTb_.mjs → set-C0dCLnUA.mjs} +4 -4
  103. package/dist/{setting-DSGXJehQ.mjs → setting-Licy7ZSs.mjs} +3 -3
  104. package/dist/{setup-aJLGLrIT.mjs → setup-BL4fqxoU.mjs} +4 -4
  105. package/dist/{skills-Q2AFsYvc.mjs → skills-Cn8FMCxG.mjs} +3 -3
  106. package/dist/snippet-C5_4nPGj.mjs +19 -0
  107. package/dist/{stash-DPQ0c-Cd.mjs → stash-SGod4B8-.mjs} +6 -6
  108. package/dist/{status-CvKPrV5X.mjs → status-BCb5iCsQ.mjs} +5 -5
  109. package/dist/{status-CvAATvV0.mjs → status-CvUuZVvC.mjs} +2 -2
  110. package/dist/{summary-CeOnoOq2.mjs → summary-qaTPTLWs.mjs} +3 -3
  111. package/dist/{sync-schema-aOPBc3CY.mjs → sync-schema-CjjAYTbO.mjs} +5 -5
  112. package/dist/table-DK1B8Hzq.mjs +19 -0
  113. package/dist/transform-Dh5746zH.mjs +24 -0
  114. package/dist/transform-job-D0nRGPYN.mjs +19 -0
  115. package/dist/{tree-MOQOBeAP.mjs → tree-C0MFcweC.mjs} +2 -2
  116. package/dist/{update-BWyCK8QV.mjs → update-BImaz9vj.mjs} +6 -6
  117. package/dist/{update-B6mg3AZD.mjs → update-BTzsE8GU.mjs} +5 -5
  118. package/dist/{update-BRrnfG0q.mjs → update-BUoWwmQV.mjs} +5 -5
  119. package/dist/{update-CitS-QRN.mjs → update-BqsGGrzx.mjs} +6 -6
  120. package/dist/{update-57uxZWcR.mjs → update-CHUUY-IB.mjs} +5 -5
  121. package/dist/{update-Bj9s0ri8.mjs → update-CJASCsG4.mjs} +5 -5
  122. package/dist/{update-DfNKr_vS.mjs → update-DCR0wFNl.mjs} +5 -5
  123. package/dist/{update-BkMWBzvk.mjs → update-Dty6LvU2.mjs} +5 -5
  124. package/dist/{update-dashcard-D_-ura3Y.mjs → update-dashcard-BkJaVVvf.mjs} +5 -5
  125. package/dist/{update-BD9xkglP.mjs → update-qoP7hZTg.mjs} +5 -5
  126. package/dist/{update-Dri4Zg2H.mjs → update-vk_M0eOR.mjs} +5 -5
  127. package/dist/{upgrade-CFkZ4USY.mjs → upgrade-TdaiAYbH.mjs} +2 -2
  128. package/dist/{uuid-DpinhSxA.mjs → uuid-BeUz8VpP.mjs} +2 -2
  129. package/dist/{values-BSS4DRxk.mjs → values-DR-3HcK6.mjs} +3 -3
  130. package/dist/{verify-B_A7v8TY.mjs → verify-D2uBlvmt.mjs} +1 -1
  131. package/dist/{wait-D3iSnjMM.mjs → wait-FAnqO-LT.mjs} +5 -5
  132. package/dist/{wait-flags-_LnHOeBA.mjs → wait-flags-B5BI_xob.mjs} +2 -2
  133. package/package.json +2 -1
  134. package/skill-data/core/SKILL.md +22 -23
  135. package/skill-data/data-analysis/SKILL.md +65 -0
  136. package/skill-data/data-transformation/SKILL.md +200 -0
  137. package/skill-data/document/SKILL.md +4 -4
  138. package/skill-data/mbql/SKILL.md +20 -20
  139. package/skill-data/robot-data-engineer/SKILL.md +142 -0
  140. package/skill-data/semantic-layer/SKILL.md +166 -0
  141. package/skill-data/transform/SKILL.md +46 -48
  142. package/skill-data/visualization/SKILL.md +5 -3
  143. package/skills/metabase-cli/SKILL.md +6 -0
  144. package/dist/add-collection-D9wXgmRj.mjs +0 -10
  145. package/dist/auth-cFC5m69m.mjs +0 -19
  146. package/dist/card-ClvGX6dQ.mjs +0 -20
  147. package/dist/collection-DjvowSJC.mjs +0 -20
  148. package/dist/dashboard-BFeURTOw.mjs +0 -21
  149. package/dist/db-CSH1kwQr.mjs +0 -22
  150. package/dist/document-KdT_Xj6r.mjs +0 -19
  151. package/dist/field-CTFnZI8G.mjs +0 -18
  152. package/dist/git-sync-C2vib8rx.mjs +0 -28
  153. package/dist/is-dirty-Bb0Rtj7x.mjs +0 -9
  154. package/dist/measure-Dw1QpRZa.mjs +0 -19
  155. package/dist/segment-DMuYvFjg.mjs +0 -19
  156. package/dist/snippet-Df2TrP7-.mjs +0 -19
  157. package/dist/table-pK4OkVtL.mjs +0 -19
  158. package/dist/transform-D60veFH8.mjs +0 -24
  159. package/dist/transform-job-DXt5LsrY.mjs +0 -19
@@ -0,0 +1,200 @@
1
+ ---
2
+ name: data-transformation
3
+ description: Turn a raw, normalized source database into a small set of clean, analysis-ready tables. Claude investigates the source, works out the real-world "things" the data is about (even when each one is scattered across several tables), decodes coded/JSON/translated values into readable text, and builds one wide, denormalized table per thing as Metabase transforms. Designed for a non-technical user who knows their domain. Use whenever someone wants to "clean up", "flatten", "denormalize", "make sense of", or "build analysis-ready tables from" a raw database. This is the strategy skill for modeling a whole database into a set of clean tables; for authoring or running one individual transform (body shape, flags, run inspection), use the `transform` skill instead.
4
+ allowed-tools: Read, Write, Edit, Bash, AskUserQuestion, EnterPlanMode, ExitPlanMode
5
+ ---
6
+
7
+ # Data Transformation
8
+
9
+ > **Shared contract (read first).** This skill is part of the `robot-data-engineer` family and follows its shared rules: ask before showing PII row-by-row (names, emails, phones) — default to aggregates; when asked for something the CLI can't do (alerts, dashboard filters), name the limit instead of erroring into raw SQL; honor the autonomy mode the user picked. The jargon rules are spelled out in detail below (**Who you're talking to**). Full contract and the autonomy slider live in the router — run `mb skills get robot-data-engineer` and read its **Shared Contract** if you haven't.
10
+
11
+ Your job: take a raw source database — usually normalized, often synced from some SaaS tool by a connector like Fivetran, Airbyte, or Stitch — and produce a **small set of wide, clean, analysis-ready tables**, one per real-world _thing_ the data is about, built as Metabase **transforms** the user can inspect.
12
+
13
+ Drive everything through the `mb` CLI. Load the skills you'll need:
14
+
15
+ ```bash
16
+ mb skills get core # auth, profiles, db/table/field inspection, query
17
+ mb skills get mbql # if you build transform queries in MBQL
18
+ mb skills get transform # creating/running transforms, run inspection
19
+ ```
20
+
21
+ Users authenticate. You pick the profile per `core`'s **Auth & profiles** and pass `--profile <name>` to every command. That profile's `url` is the instance's base URL. Browser links below are built from it, ensuring the links are consistent with your CLI usage.
22
+
23
+ If you are making transforms, use the transform skill.
24
+
25
+ ---
26
+
27
+ ## Who you're talking to
28
+
29
+ A **non-technical user who knows their domain well** — they understand the business (events, customers, invoices, etc.) but not databases.
30
+
31
+ - **No modeling jargon.** Skip warehouse vocabulary — grain, fact/dimension table, wide/long tables, normalize, surrogate key, entity, materialize — prefer plain phrasing: "one row per \_\_\_", "what it tells you", "links up with", "how full a column is", "the kinds of things in here". **But don't overdo it:** basic relational terms are fine — table, column, ERD, schema, key, foreign key (cardinality too, though "one-to-many" usually lands better). **Metabase's product terms are encouraged** — Question, Model, Segment, Measure, Metric, Transform — they're not database jargon.
32
+ - **Don't lean on raw SQL to communicate.** They may follow a simple `SELECT`, but don't explain work via SQL or ask them to read/write it.
33
+ - Group what you show by **the question a column answers**, never by which source table it came from.
34
+ - Be a **helpful assistant, not an engineer reporting status.** Elide machinery; ask sharp questions that matter.
35
+ - Your user may say "go" and come back later. **If you ever ask the user a question, wait for their answer.**
36
+
37
+ ---
38
+
39
+ ## Two kinds of decisions
40
+
41
+ Sort every choice into one of these.
42
+
43
+ **Hard rules — absolutes, never ask:**
44
+
45
+ 1. Never flatten multi-valued fields into opaque blobs (e.g. three options squished: `"email | phone | text"`). It destroys filterability (the whole point).
46
+ 2. Never use jargon with the user. Explain by domain and telos.
47
+ 3. Always surface **real data you're about to leave out** proactively, ranked by how much is extant.
48
+ 4. Never guess what schema mean from their name alone. Confirm against actual values, interpret them in context: the table the field belongs to and the relevant domain (e.g., a status on orders ≠ status on subscriptions).
49
+ 5. Never silently drop a whole _thing_. Dropping a column is routine; dropping a whole kind-of-thing (e.g. "suppliers") must be surfaced and confirmed.
50
+ 6. Never drop columns that link things together. Every table keeps its own id **and** the ids tying it to other tables — alongside the readable labels you copy in, not instead of. The label is for reading; the id is for joining. You're building tables about _related_ things, so they **will** be combined ("sales per region", "messages per customer") — dropped ids make that quietly impossible and the user will regret it. Keep the ids; don't force the user to stare at them.
51
+ 7. Never bake a non-obvious business rule into a table without confirming it in plain terms. When a transform encodes a judgment the user would have an opinion on — how money nets, which row is the "current" one, what "active" means — say it back in one plain sentence and get a yes/no first. You know only the columns; they know the business. Wrong rules hide insidiously in clean-looking tables. ("I'm treating each person's most recent sign-up as their current one — right?")
52
+ 8. Never sneak sensitive personal data through. Flag it on sight — addresses, phone numbers, emails, IPs, financial, etc. — and ask the user how to handle it (the prudential call below). Always surface, never silently expose it in a table others will browse.
53
+ 9. Never overwrite existing tables or other transforms' outputs. Before building, check the target name is unused (`mb transform list`, `mb table list`); if it's in use, stop and surface it — building over it silently destroys their data. Reuse names only for updating _your own_ transform (`transform update`), never for clobbering another.
54
+
55
+ **Prudential calls — contextual, multiple good answers, hinge on domain knowledge you lack. State a lean, then let the user decide.** The recurring ones:
56
+
57
+ - **Multi-valued attribute** (one response → many options; one order → many line items): keep it filterable! Structured columns for predefined lists, or simple join tables, never opaque text. Structure is the user's call. Lean: easiest filtering, probably flat.
58
+ - **Layering**: default **flat** — one self-contained table per thing, no hidden intermediate tables. Suggest a shared cleaned-up base table only for DRY, avoiding copying complex logic across many transforms. Even then, ask.
59
+ - **Out-of-scope things**: surface every domain-model you find and ask in/out, rather than inferring scope from what they happened to mention.
60
+ - **A repeating thing vs. the events it takes part in**: one table can mix a _stable_ thing (a customer, a company) with _repeating_ events (each order, each visit), copying the stable details onto every event row. If that thing genuinely recurs — same customer on many rows — consider a one-row-per-thing table too, linked by id, so "how many distinct X" and the per-X details have clean homes. Lean: split when recurrence is real, but one table when each appears once. (Phase 0's one-to-one / one-to-many check already tells you which.)
61
+ - **Handling sensitive data** (addresses, emails, phones, IPs, financial details): once you've flagged it (rule 8), _how_ to carry it is user's choice — keep as-is, mask (partial redaction), or drop. Lean: keep what is needed, mask the rest, drop the useless.
62
+
63
+ Phrase a prudential call as a lean plus a nod:
64
+
65
+ > "I'd keep these as one simple table rather than splitting into behind-the-scenes pieces — easier to look through. Good?"
66
+
67
+ ---
68
+
69
+ ## The process
70
+
71
+ ### Phase 0 — Get Oriented
72
+
73
+ **Pin down where the data lives — ask before you hunt.** A table or schema name the user mentions tells you _what_ but not _where_: an instance can hold several databases, each with several schemas. Rather than listing them all to find it, just ask — "Which database is this in, and the schema if you know it? No worries if you're not sure, I can find it." A confident answer short-circuits a lot of blind searching; "not sure" costs nothing and you fall back to locating it yourself. If you've genuinely looked and still can't find a table the user is sure is there, don't keep digging. One possible reason is that Metabase hasn't picked up that database's latest schema yet — gently raise it and ask whether the data's been synced recently, and let the user run the sync from Metabase if it's needed.
74
+
75
+ As soon as you know which database and schema you're in:
76
+
77
+ - **Show the user the map.** Open the instance's schema map for that schema so they can follow along: `<base-url>/data-studio/schema-viewer?database-id=<db-id>&schema=<schema>`. Open it in their browser if you can (e.g. `open` / `xdg-open`); else paste the URL. Don't skip this.
78
+ - **Ask for a head start.** "Do you have a picture or file showing how your data fits together, like an ERD?" If yes, read it — it shortcuts the next steps.
79
+ - **Ask for their conventions.** "Is there already cleaned-up data, or a past project, that shows how your team likes this done?" If yes, inspect it: it tells you their naming, their idea of "clean," and existing tables worth linking to.
80
+
81
+ ### Phase 1 — Investigate (in plan mode, if they choose)
82
+
83
+ Orientation done, you're about to go heads-down. First, offer two ways to work:
84
+
85
+ > Two ways I can take it from here:
86
+ >
87
+ > - **I dig through it all and bring you a complete plan** to approve before I build anything — quieter; you won't hear much until it's ready.
88
+ > - **We work it out together** — I share what I find and we make the calls as we go.
89
+
90
+ First path: **enter plan mode** (`EnterPlanMode`). Everything up to the agreed table list — investigate, present, prudential calls, naming (Phases 1–3) — happens inside it, read-only; you exit once, at the approval gate before building (Phase 4). Second path: skip it, shape it conversationally through the same phases. Either way, don't build until the design is settled and user-approved.
91
+
92
+ Plan mode is a long quiet stretch — they said "go" and walked off. So whenever you surface — a question now, the plan at the end — **carry your own context**: recap what it rests on right before you ask, never a back-reference to something said while they were away (the router's contract spells this out).
93
+
94
+ Then dig in. Don't narrate this — a single "Let me take a look at what's in here — one minute" is enough. Keep it cheap: never pull whole-warehouse rollups (they blow up); use compact column listings, `LIMIT`/sample queries, and `GROUP BY count(*)`.
95
+
96
+ 1. **Map the tables.** List them; pull each one's column names and types; note its own id.
97
+ 2. **Find the decode tables.** Normalized SaaS data hides meaning in lookups — `*_field`, `*_field_choice`, `*_question`, `*_choice`, `*_type`. A column like `doodad_4471` is meaningless until you join the lookup and find it's _"Preferred vehicular transport"_. Build that code → label map yourself by joining the lookups — never hand the user a coded column and ask what it means — before showing them anything.
98
+ 3. **Prove the connections — don't trust declared keys.** Synced databases usually have none. If that's the case, ask the user if they have ERD or relationship information (screenshot, JSON, documentation, etc.). For each `<x>_id`, guess it points at `<x>`, then check what fraction of values actually match the target's id: high = real link, low = decoy, discard. Note one-to-one vs one-to-many. **Also look outward** — does a thing you're about to build already exist as clean data elsewhere in the instance (an existing customers table your people match, a product list)? If so, plan to _link_ to it, not duplicate it.
99
+ 4. **Pin down "one row per what."** Count rows; check the id is unique; figure out what a single row is. **Watch for lies:** a stale count column, or a table that looks like "all of X" but is a filtered subset.
100
+ 5. **Reconcile across related tables.** Do child rows all link to a parent? Orphans? Is one table a trimmed snapshot while another keeps everything? These mismatches matter and the user can't see them — you must.
101
+ 6. **Profile the values.** List distinct values for coded/low-variety columns; check how full (% non-empty) any column you might drop is; spot multi-valued JSON fields. Profile with the cleaning checklist (end of file) in mind — surface the quality smells you hit, don't silently fix them.
102
+ 7. **Cluster into things.** Group tables and columns into the real-world things they describe — a thing may span several tables (one _customer_ across a main table + a loyalty table + custom-profile columns). Decide "one row per \_\_\_" for each and gather its attributes, decoded. Watch for a table that secretly mixes _two_ things — a stable thing plus its repeating events; that's the split in the prudential calls above.
103
+
104
+ **Then, still quietly, sketch the design space.** Once the things and how they connect are pinned, brainstorm the range of questions this data could answer — finance views, leaderboards, breakdowns. **Don't show it to the user or build any of it.** It only pressure-tests your design: would a reasonable pivot to a nearby question force a rewrite? When keeping a column or finer grain _cheaply_ preserves that flexibility, keep it. Serve the user's stated concern — but don't scope so tightly that the next question means starting over.
105
+
106
+ ### Phase 2 — Present what you found (plain language)
107
+
108
+ Three things, in order:
109
+
110
+ **(a) The things, in plain terms.** One short blurb each. E.g. in an online store:
111
+
112
+ > **Customers** — one row per customer. Who they are (name, company, location), how they've been in touch, what they've spent, whether they're active or churned.
113
+
114
+ **(b) The full inventory — including what you'd leave out.** Never infer scope silently:
115
+
116
+ > I found 6 kinds of things: **Customers, Orders, Products, Suppliers, Shipments, Returns.** I'd build the first four. **Shipments** and **Returns** also have real data — want those in, or leave them?
117
+
118
+ **(c) What would be set aside — proactively, ranked, two buckets:**
119
+
120
+ > Nothing important is lost. A few things set aside:
121
+ > • **Real data** — gift-message text (6 of 10 orders), delivery instructions (most), preferred carrier. Minor, but real — want any kept?
122
+ > • **Safe to drop** — duplicate product names in other languages, internal bookkeeping columns. No real loss.
123
+
124
+ If you spotted existing clean data to link to (step 3), raise it here too — and **always run a suspected match past the user before wiring it; never graft onto their existing data silently.** Then ask your prudential questions, one at a time, each a lean-plus-nod.
125
+
126
+ ### Phase 3 — Iterate
127
+
128
+ Cheap, because nothing's built. Adjust the set of things, what's kept, and the shape of any multi-valued pieces until the user's happy. **Agree on what each table will be called** — propose a clear name for each (matching any naming pattern you found in their existing data, Phase 0) and let them adjust. Confirm each name is free — not already an existing table or another transform's output (rule 9) — so building can't overwrite anyone's data. Settle the names before building: the name you agree on is the one you build and keep. Re-confirm the final picture in one short recap. **In plan mode, that recap _is_ your exit:** present it as the plan and call `ExitPlanMode` — approval here is the single go-ahead to build. (Iterating together? The recap is just your check before building.)
129
+
130
+ ### Phase 4 — Build, check, hand back
131
+
132
+ Design settled — now you build, the first step that writes; plan mode, if you used it, is behind you. Build one wide transform per agreed thing — and build for how it'll be judged: aim for output that's readable on sight, not just one that runs clean. Each table:
133
+
134
+ - **Denormalized, but the link stays.** Copy in related context so casual reading needs no lookups (a product's name and price on the orders table) — **and keep the linking id beside it** (the product's id too, per rule 6). Use the same id name everywhere a thing appears.
135
+ - **Decoded**: codes and JSON become readable text; bookkeeping columns and soft-deleted rows are gone (filter the source's soft-delete flag — Fivetran's `_fivetran_deleted`, Airbyte's `_ab_cdc_deleted_at`, or a plain `deleted_at`/`is_deleted` — so tombstones never reach clean data; not every source has one).
136
+ - **Clean, plain column names**, consistent across tables.
137
+ - **Multi-valued pieces** in the agreed filterable structure (rule 1).
138
+ - **Keep the detail; don't pre-summarize it away.** Build the detailed rows (one per order, one per payment), not pre-computed totals. A convenience count is fine _beside_ the rows, never _instead of_ them — a frozen total only ever answers the one question it was summed for.
139
+
140
+ Then make the links real, not just implied:
141
+
142
+ - **Wire foreign keys between your tables.** Mark each linking id as a foreign key pointing at the id it references (`mb field update` — set the column's type to foreign-key and its target). Now Metabase itself knows the tables connect and can traverse them.
143
+ - **Graft onto existing clean data** the user approved (step 3 / Phase 1): point the linking id at the existing table's id the same way. Link, don't duplicate.
144
+ - **Write down what you learned.** You decoded every column's real meaning while investigating — save it: set a short description on each table and its non-obvious columns (`mb table update` / `mb field update`). The cleaned data then explains itself inside Metabase — in search, in the Question editor, to Metabot — instead of the knowledge living only in this chat.
145
+
146
+ When you start refining a built transform _with_ the user, open its inspector for them so you're looking at the same thing — `<base-url>/data-studio/transforms/<transform-id>/inspect` — opening it in their browser if you can, else pasting the URL. Iterate with `transform update`, never delete-and-recreate.
147
+
148
+ **Check the output before handing back — the user can't.** Two passes, in order.
149
+
150
+ **Pass 1 — Correctness (did it run right).** After each transform runs, run quick ad-hoc tests against what Phase 0 led you to expect: row counts in the right ballpark, decoded columns readable (no stray codes), linking ids that resolve to the other tables, no column unexpectedly all-null or blown up in count. Treat surprises as bugs to chase, not noise. A table that can't combine with the others — a dropped id, or the same id named two ways — is a silent failure; catch it here.
151
+
152
+ **Pass 2 — Fitness (is it nice to use).** Correct isn't the bar; _usable_ is. `SELECT * FROM <table> LIMIT 20` and read every column left to right as if you'd never seen the source: would a non-technical person find each one readable? Smells that say not-yet, even though nothing errored:
153
+
154
+ - a multi-valued column still a raw JSON/array blob or `["Email","SMS"]` text — rule 1 never actually got resolved;
155
+ - decoded answers still carrying raw ids with no readable label, or one cryptic column per code;
156
+ - a code sitting beside its own label when only the label is wanted, or two columns saying the same thing;
157
+ - a "decoded" column that reads as a slug (`pref_contact_mthd`) rather than plain language.
158
+
159
+ A readability smell is a bug: fix it (`transform update`), re-run, look again. When the fix is really a shape choice (how a multi-select is structured) or a keep/drop call, that's the user's — surface it, don't silently decide.
160
+
161
+ Then report plainly:
162
+
163
+ > Done. Three tables:
164
+ > • **Customers** — transform #41
165
+ > • **Orders** — transform #42
166
+ > • **Products** — transform #43
167
+ >
168
+ > How they connect: each **Order** belongs to a **Customer**; each **Order** lists one or more **Products**.
169
+
170
+ End on that connection map: it's what the user reads to trust the result, and what lets whatever they build next join the tables on the right ids instead of guessing how they relate.
171
+
172
+ ---
173
+
174
+ ## A worked decode example (for your reference, not the user's)
175
+
176
+ The shape recurs across SaaS exports, whatever the domain. A coded column — say `c_4471` on a responses table — means nothing alone. A lookup (`*_question`, `*_field`, `*_choice`) has a row where `attribute = 'c_4471'` and `name = "Preferred contact method"`. Single-select answers are often already `{"id":…, "value":"Email"}` — use `value`. Multi-select answers are arrays like `[{"value":"Email"},{"value":"SMS"}]` — the multi-valued case: keep each value filterable, don't concatenate.
177
+
178
+ Always decode _before_ presenting, so the user sees "Preferred contact method", never `c_4471`. Three cautions:
179
+
180
+ - **Pull the readable name from the lookup, don't type it in.** The label (and any question text) should come _from_ the lookup's `name`, sourced in the query — not pasted as a literal. A hard-typed label goes wrong the moment the source changes.
181
+ - **Codes are usually specific to today's data.** `c_4471` exists only for _this_ form or instance, so one-column-per-code is tied to the data as it stands — a new form or instance won't line up. When that's unavoidable, say so on hand-back ("reflects the current form; new questions need a refresh"), and with many such codes prefer the companion-table shape (one row per answer, question text from the lookup): nothing hard-typed, and adding a question is a smaller change.
182
+ - **Normalize encodings once.** Turn raw representations clean in the table itself, so nothing downstream re-derives them: signed amounts → clear positive numbers by kind, 0/1 → true/false, timestamps → one consistent timezone, text → trimmed and case-consistent, and junk placeholders (`"NULL"`, `"N/A"`, `"-"`, empty string) → real null.
183
+
184
+ ---
185
+
186
+ ## Cleaning checklist (for your reference, not the user's)
187
+
188
+ A scan-list, not a pipeline — and the governing rule is **surface what you find, don't silently "fix" it.** Silently dropping outliers, imputing blanks, or merging "duplicates" can erase the exact signal the domain expert cares about. Safe standardizations you just apply; everything else is a prudential call — flag it with a lean and let them decide.
189
+
190
+ **Just apply** (safe, universal — already your default): consistent timestamps/timezone; trimmed, case-consistent text; junk placeholders (`"NULL"`, `"N/A"`, `"-"`, `""`) → real null; sane numeric precision; booleans from varied forms (Y/N, 1/0).
191
+
192
+ **Notice and surface** (the answer depends on their business):
193
+
194
+ - **Duplicates** — exact, or by business rule ("same email = same person"). Never merge silently.
195
+ - **Validation smells** — out-of-range numbers, malformed emails/phones/ids, `end_date < start_date`.
196
+ - **Outliers** — values that read as data-entry errors. Flag, don't drop.
197
+ - **Missing data** — random vs. systematic? Surface the pattern; never silently impute or default.
198
+ - **Free text / mixed encodings** — handle the safe parts, flag the rest.
199
+
200
+ Already covered by the rules above, listed so they stay on your radar: structural reshaping (decode/JSON/multi-value), orphans & key validity (Phase 0 step 5 + the post-run check), filtering soft-deletes & dropping bookkeeping columns (Phase 4's **Decoded** step), and recording meanings (the descriptions step).
@@ -145,10 +145,10 @@ Each entry in `cards` needs at least `{name, dataset_query, display, visualizati
145
145
  `update` replaces the whole `document` body, so the safe loop is **read → edit → write**. A fetched body already carries `_id`s on its id-bearing nodes, so preserve them — only mint new ones for id-bearing nodes you add:
146
146
 
147
147
  ```bash
148
- mb document get <id> --full --profile <name> --json | jq '.document' > /tmp/body.json
149
- # edit /tmp/body.json (add nodes — give each new id-bearing node a fresh `mb uuid` _id) …
150
- jq -n --slurpfile d /tmp/body.json '{document: $d[0]}' > /tmp/patch.json
151
- mb document update <id> --file /tmp/patch.json --profile <name> --json
148
+ mb document get <id> --full --profile <name> --json | jq '.document' > ./.scratch/body.json
149
+ # edit ./.scratch/body.json (add nodes — give each new id-bearing node a fresh `mb uuid` _id) …
150
+ jq -n --slurpfile d ./.scratch/body.json '{document: $d[0]}' > ./.scratch/patch.json
151
+ mb document update <id> --file ./.scratch/patch.json --profile <name> --json
152
152
  ```
153
153
 
154
154
  Don't hand-merge a partial node tree into a live document — pull the current `document`, mutate the array, and PUT the whole thing back. To rename without touching the body, patch only `name`: `mb document update <id> --body '{"name":"New title"}'`.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: mbql
3
- description: Author Metabase MBQL 5 query bodies for the `mb` CLI — the only hand-authorable query format. Covers the JSON shape (lib/type mbql/query, flat stages, numeric ids), the "options object always second" clause rule, when lib/uuid is needed (it's optional — only to reference a clause), the print-schema → dry-run → run validation loop, where MBQL 5 is consumed (mb query, card dataset_query, transform source.query, measure/segment definition), the flat-vs-legacy-envelope footgun, joins and FK traversal, multi-stage pipelines, and naming aggregation output columns. Load whenever building or fixing an MBQL query by hand — "write an MBQL query", "create a card from MBQL", "the dataset_query is wrong", "fix the validation errors", "aggregate and group by", "order by the count", "join two tables", "month-over-month", or any `--dry-run` / `mb query` work.
3
+ description: Author Metabase MBQL 5 query bodies for the `mb` CLI - the only hand-authorable query format. Covers the JSON shape (lib/type mbql/query, flat numeric-id stages), the options-object-always-second clause rule, when lib/uuid is needed (optional - only to reference a clause), the print-schema/dry-run/run loop, where MBQL 5 is consumed (mb query, card dataset_query, transform source.query, measure/segment definition), the flat-vs-legacy-envelope footgun, joins and FK traversal, multi-stage pipelines, naming aggregation columns. Load when building or fixing an MBQL query by hand - "write an MBQL query", "create a card from MBQL", "the dataset_query is wrong", "fix the validation errors", "aggregate and group by", "join two tables", "month-over-month", or any `--dry-run` / `mb query` work.
4
4
  allowed-tools: Read, Write, Edit, Bash, AskUserQuestion
5
5
  ---
6
6
 
@@ -8,13 +8,13 @@ allowed-tools: Read, Write, Edit, Bash, AskUserQuestion
8
8
 
9
9
  MBQL 5 is the **only query format you can author by hand** with confidence — it has a bundled JSON Schema, so the CLI pre-flight-validates it before sending. Legacy MBQL 4 and native SQL are accepted but **not** schema-validated (see "Other formats" below).
10
10
 
11
- Prefer MBQL over native SQL: it's portable across warehouse engines and the CLI pre-flight-validates it. Try it first, but don't force it — fall back to native SQL when MBQL can't express what you need, or when an MBQL body keeps failing server-side and you can't resolve it.
11
+ Prefer MBQL over native SQL: portable across warehouse engines and pre-flight-validated. Try it first; fall back to native SQL when MBQL can't express what you need, or when an MBQL body keeps failing server-side and you can't resolve it.
12
12
 
13
- The general flag conventions, body-input precedence, and output flags live in the `core` skill (`mb skills get core`).
13
+ General flag conventions, body-input precedence, and output flags live in the `core` skill (`mb skills get core`).
14
14
 
15
15
  ## The shape
16
16
 
17
- A query is a flat object — `lib/type`, a numeric `database` id, and an ordered `stages` array. No recursive `source-query` nesting; multi-step queries are sibling stages.
17
+ A flat object — `lib/type`, a numeric `database` id, and an ordered `stages` array. No recursive `source-query` nesting; multi-step queries are sibling stages.
18
18
 
19
19
  ```json
20
20
  {
@@ -31,7 +31,7 @@ A query is a flat object — `lib/type`, a numeric `database` id, and an ordered
31
31
  }
32
32
  ```
33
33
 
34
- - **Numeric ids only.** `database`, `source-table`, and field ids are integers from `mb database list` / `mb table get <id> --include fields`. (The portable YAML representation under git-sync uses _names_ like `[Sample Database, PUBLIC, ORDERS]`; the CLI's `/api/dataset` form uses numeric ids — don't mix them.)
34
+ - **Numeric ids only.** `database`, `source-table`, and field ids are integers from `mb database list` / `mb table get <id> --include fields`. (Git-sync YAML uses _names_ like `[Sample Database, PUBLIC, ORDERS]`; the `/api/dataset` form uses numeric ids — don't mix them.)
35
35
  - **First stage** carries `source-table` (a table id) or `source-card` (a saved card). Later stages omit both and read the previous stage's output columns by name.
36
36
  - `source-card` references a saved card by its **numeric id** (from `mb card list`), not its string entity id; downstream fields are referenced by column name (string), not a field id.
37
37
 
@@ -53,9 +53,9 @@ The same `[op, {options}, …]` rule holds for `aggregation`, `breakout` (a list
53
53
 
54
54
  ## UUIDs: optional — mint only to reference a clause
55
55
 
56
- `lib/uuid` is **optional — leave it out whenever you can.** Omit it and the server generates a unique one for every clause as the query comes in; an empty options object `{}` is the normal, preferred case. Don't add a UUID per clause: it's needless work, and the more UUIDs you hand-manage the easier it is to trip the server's "all `lib/uuid`s must be unique" check — a duplicated UUID passes pre-flight, then fails server-side.
56
+ `lib/uuid` is **optional — leave it out whenever you can.** Omit it and the server generates a unique one for every clause; an empty options object `{}` is the normal case. The more UUIDs you hand-manage the easier it is to trip the server's "all `lib/uuid`s must be unique" check — a duplicated UUID passes pre-flight, then fails server-side.
57
57
 
58
- Set an explicit `lib/uuid` only when you must **reference a clause from elsewhere in the query** — the one thing the server can't do for you, since you have to know the value to point at. The case that needs it: **ordering by (or otherwise reusing) an aggregation.** `["aggregation", {…}, "<uuid>"]`'s third arg is the **string** `lib/uuid` of the target aggregation, so give that aggregation an explicit `lib/uuid` and point the ref at the same string. A numeric position fails with `must be the target aggregation's lib/uuid (string), not a numeric position`.
58
+ Set an explicit `lib/uuid` only when you must **reference a clause from elsewhere in the query** — you have to know the value to point at. The case that needs it: **ordering by (or otherwise reusing) an aggregation.** `["aggregation", {…}, "<uuid>"]`'s third arg is the **string** `lib/uuid` of the target aggregation, so give that aggregation an explicit `lib/uuid` and point the ref at the same string. A numeric position fails with `must be the target aggregation's lib/uuid (string), not a numeric position`.
59
59
 
60
60
  ```json
61
61
  "aggregation": [["count", { "lib/uuid": "AGG_UUID" }]],
@@ -64,7 +64,7 @@ Set an explicit `lib/uuid` only when you must **reference a clause from elsewher
64
64
 
65
65
  (`AGG_UUID` is both the aggregation's own `lib/uuid` and the string the ref points at — one value, by string equality. Every other clause omits its UUID. Expression refs work the same way but key off the expression's `lib/expression-name` string, so expressions rarely need an explicit `lib/uuid`.)
66
66
 
67
- On the rare occasion you do need one, **always mint it with `mb uuid` — never write, guess, or copy a UUID yourself.** A hand-authored value is either rejected pre-flight as not-a-v4 (`"a1"`, `"uuid-1"`, `"agg-uuid-001"` → `must be a UUID v4 (RFC 4122) — run \`mb uuid\``) or, if it happens to look valid, risks colliding with another clause. Only `mb uuid`gives you genuine, unique v4s — mint just the few you reference (this also covers native template-tag ids and any other`format: "uuid"` slot):
67
+ When you do need one, **always mint it with `mb uuid` — never write, guess, or copy a UUID yourself.** A hand-authored value is rejected pre-flight as not-a-v4 (`"a1"`, `"uuid-1"`, `"agg-uuid-001"` → `must be a UUID v4 (RFC 4122) — run \`mb uuid\``), or if it looks valid risks colliding with another clause. Only `mb uuid`gives genuine, unique v4s — mint just the few you reference (also covers native template-tag ids and any other`format: "uuid"` slot):
68
68
 
69
69
  ```bash
70
70
  mb uuid --count 2 --json # mint only the clauses you actually reference
@@ -75,7 +75,7 @@ mb uuid --count 2 --json # mint only the clauses you actually reference
75
75
  `mb query` is the canonical authoring surface. Three modes:
76
76
 
77
77
  ```bash
78
- mb query --print-schema --profile <n> > /tmp/mbql-schema.json # 1. fetch the schema
78
+ mb query --print-schema --profile <n> > ./.scratch/mbql-schema.json # 1. fetch the schema
79
79
  mb query --file q.json --dry-run --profile <n> # 2. validate, no network
80
80
  mb query --file q.json --profile <n> --json # 3. validate + run
81
81
  ```
@@ -86,7 +86,7 @@ mb query --file q.json --profile <n> --json # 3. validate +
86
86
 
87
87
  `path` is a JSON Pointer into the body (`/stages/0/aggregation/0`); `message` is the validator error. Exit codes: `0` valid + ran, `2` validation failed / malformed body, `1` server-side error after a valid pre-flight.
88
88
 
89
- **Pre-flight is a lightweight shape check, not the full backend validator.** It checks JSON shape, `lib/uuid` format, and enum values — not operator names, the first-stage source rule, or whether a reference resolves. A clean `--dry-run` is necessary but not sufficient: a body can pass pre-flight and still fail on the server (exit `1`). The Metabase server is the authority — when a run fails, read its error and fix the body. The common ones and what they mean:
89
+ **Pre-flight is a lightweight shape check, not the full backend validator.** It checks JSON shape, `lib/uuid` format, and enum values — not operator names, the first-stage source rule, or whether a reference resolves. A clean `--dry-run` is necessary but not sufficient: a body can pass pre-flight and still fail on the server (exit `1`). The server is the authority — when a run fails, read its error and fix the body. Common ones:
90
90
 
91
91
  - `not a known MBQL clause` → a misspelled or unsupported **operator**. Check the vocabulary in `operators.md` (`mb skills get mbql --full`).
92
92
  - `Initial MBQL stage must have either :source-table or :source-card` → the **first stage** is missing its source (a numeric table or card id); only the first stage takes one, later stages read the previous stage's columns.
@@ -95,11 +95,11 @@ mb query --file q.json --profile <n> --json # 3. validate +
95
95
 
96
96
  A successful run emits the compact envelope by default: `data.rows` + slim `data.cols` (`name`, `display_name`, `base_type`, `semantic_type`). Pass `--full` for the raw `/api/dataset` envelope (`results_metadata`, `native_form`, per-column fingerprints/`field_ref`) only when you need that metadata; `--fields data.rows` narrows to rows alone. `mb query` also runs a **native** body — `{database, type:"native", native:{query:"SELECT …"}}` — which skips pre-flight; the quickest way to eyeball warehouse data.
97
97
 
98
- `--skip-validate` bypasses the pre-flight and sends as-is — use only when the bundled schema disagrees with what the server actually accepts (drift / false negative). Mutually exclusive with `--dry-run`. The same flag exists on `card create/update` and `transform create/update`.
98
+ `--skip-validate` bypasses pre-flight and sends as-is — use only when the bundled schema disagrees with what the server actually accepts (drift / false negative). Mutually exclusive with `--dry-run`. Same flag exists on `card create/update` and `transform create/update`.
99
99
 
100
100
  ## Where MBQL 5 is consumed
101
101
 
102
- The same body and the same pre-flight apply everywhere a query is embedded. Each pre-flights only when the value is MBQL 5 (`lib/type: "mbql/query"`); legacy shapes skip it; `--skip-validate` bypasses.
102
+ The same body and pre-flight apply everywhere a query is embedded. Each pre-flights only when the value is MBQL 5 (`lib/type: "mbql/query"`); legacy shapes skip it; `--skip-validate` bypasses.
103
103
 
104
104
  | Command | MBQL 5 lives at | Notes |
105
105
  | --------------------------------------- | ---------------------------------------------- | ------------------------------------------- |
@@ -122,16 +122,16 @@ The most common mistake. The legacy MBQL 4 shape `{ "type": "query", "database":
122
122
  }
123
123
  ```
124
124
 
125
- No `type:"query"` wrapper, no `query:` nesting. If you wrap MBQL 5 inside a legacy envelope the CLI rejects it pre-send with a `ConfigError` (no `--skip-validate` gets it past). If it ever reached the server it would store silently and fail at run time with `Initial MBQL stage must have either :source-table or :source-card`.
125
+ No `type:"query"` wrapper, no `query:` nesting. If you wrap MBQL 5 inside a legacy envelope the CLI rejects it pre-send with a `ConfigError` (no `--skip-validate` gets it past). If it reached the server it would store silently and fail at run time with `Initial MBQL stage must have either :source-table or :source-card`.
126
126
 
127
127
  ## Other formats skip pre-flight
128
128
 
129
- Anything that is not `lib/type: "mbql/query"` is sent as-is and normalized server-side:
129
+ Anything not `lib/type: "mbql/query"` is sent as-is and normalized server-side:
130
130
 
131
131
  - **Legacy MBQL 4** — `{ "type": "query", "database": N, "query": { "source-table": T, … } }`
132
132
  - **Native SQL** — `{ "type": "native", "database": N, "native": { "query": "SELECT …" } }`
133
133
 
134
- `mb query --file probe.json` runs these directly; `--dry-run` on them returns `{ ok: true, errors: [] }`. Don't author MBQL 4 by hand — if you need a legacy or complex query, build it in the Metabase UI and pull the body with `mb card get <id> --full --json` / `mb transform get <id> --full --json`.
134
+ `mb query --file probe.json` runs these directly; `--dry-run` on them returns `{ ok: true, errors: [] }`. Don't author MBQL 4 by hand — build a legacy or complex query in the Metabase UI and pull the body with `mb card get <id> --full --json` / `mb transform get <id> --full --json`.
135
135
 
136
136
  ## Joins and FK traversal
137
137
 
@@ -154,7 +154,7 @@ Two ways to read columns from a related table.
154
154
  "breakout": [["field", { "join-alias": "Customers" }, 1682]]
155
155
  ```
156
156
 
157
- The condition's left ref is a column of the stage's own source (`1711` = orders.customer_id); the right ref carries `join-alias` and points at the joined table's key (`1684` = customers.id). Every later reference to a joined column (`1682` = customers.plan) needs that same `join-alias`. Stack multiple objects in `joins` for multiple joins, each with its own `alias`.
157
+ Left ref is a column of the stage's own source (`1711` = orders.customer_id); the right ref carries `join-alias` and points at the joined table's key (`1684` = customers.id). Every later reference to a joined column (`1682` = customers.plan) needs that same `join-alias`. Stack multiple objects in `joins`, each with its own `alias`.
158
158
 
159
159
  **Implicit FK join via `source-field`.** For a single-hop FK lookup, skip the join — put the FK column's id in the target field's `source-field` option and Metabase traverses the relationship:
160
160
 
@@ -166,7 +166,7 @@ The condition's left ref is a column of the stage's own source (`1711` = orders.
166
166
 
167
167
  ## Multi-stage pipelines
168
168
 
169
- Stages run in order; each reads the **previous stage's output columns** — the breakouts and aggregations it produced — referenced by **string name + `base-type`**, not a numeric field id. Only the first stage takes a `source-table`/`source-card`. The reason to add a stage is to operate on an aggregate (you can't filter or order by an aggregation within the stage that computes it): aggregate, then filter the aggregate, then order + limit.
169
+ Stages run in order; each reads the **previous stage's output columns** — the breakouts and aggregations it produced — referenced by **string name + `base-type`**, not a numeric field id. Only the first stage takes a `source-table`/`source-card`. Add a stage to operate on an aggregate (you can't filter or order by an aggregation within the stage that computes it): aggregate, then filter the aggregate, then order + limit.
170
170
 
171
171
  ```json
172
172
  "stages": [
@@ -197,7 +197,7 @@ Later stages address the first stage's aggregation by the `name` you gave it (`"
197
197
 
198
198
  ## Naming aggregation output columns
199
199
 
200
- Default MBQL 5 aggregations materialize as `count`, `count_where`, `avg`, `avg_2`, `sum`, … — fine for an ad-hoc run, ugly when the output is a transform target table or a card column. Set `name` (becomes the warehouse column name) and `display-name` (the UI header) in the aggregation's options:
200
+ Default MBQL 5 aggregations materialize as `count`, `count_where`, `avg`, `avg_2`, `sum`, … — fine for an ad-hoc run, ugly for a transform target table or card column. Set `name` (the warehouse column name) and `display-name` (the UI header) in the aggregation's options:
201
201
 
202
202
  ```json
203
203
  ["count", { "name": "shipments_shipped", "display-name": "Shipments shipped" }]
@@ -205,7 +205,7 @@ Default MBQL 5 aggregations materialize as `count`, `count_where`, `avg`, `avg_2
205
205
 
206
206
  ## Operator reference
207
207
 
208
- The full operator vocabulary — filter operators (`=`, `!=`, `<`, `between`, `contains`, `is-null`, …), aggregation functions (`count`, `sum`, `avg`, `distinct`, `count-where`, `share`, …), expression operators (arithmetic, string, temporal), temporal-bucketing units, and binning strategies — lives in this skill's `references/operators.md`, in the CLI's numeric-id form. Load it on demand rather than dumping the schema:
208
+ The full operator vocabulary — filter operators (`=`, `!=`, `<`, `between`, `contains`, `is-null`, …), aggregation functions (`count`, `sum`, `avg`, `distinct`, `count-where`, `share`, …), expression operators (arithmetic, string, temporal), temporal-bucketing units, and binning strategies — lives in this skill's `references/operators.md`, in numeric-id form. Load it on demand rather than dumping the schema:
209
209
 
210
210
  ```bash
211
211
  mb skills get mbql --full # appends references/operators.md to this body
@@ -217,7 +217,7 @@ mb skills path mbql # → the skill dir; then Read references/operator
217
217
  ## Don't
218
218
 
219
219
  - Don't mint a `lib/uuid` for every clause — they're optional; omit them and the server fills them in. Mint (with `mb uuid`) only the clause you need to reference; never invent, hard-code, or copy a UUID (duplicates are rejected server-side).
220
- - Don't put the options object anywhere but slot 1, and don't use the legacy `["field", id, opts]` order.
220
+ - Keep the options object in slot 1 of every clause — `[op, {options}, ...args]`, id last (`["field", {}, 1779]`). The legacy `["field", id, opts]` order (id second) is rejected pre-flight.
221
221
  - Don't wrap an MBQL 5 body in `{type:"query", query:…}` — `dataset_query` / `source.query` / `definition` is the flat `mbql/query`.
222
222
  - Don't author MBQL 4 by hand — build it in the UI and pull it with `… get <id> --full --json`.
223
223
  - Don't skip the `--dry-run` loop on a non-trivial query — it's free and exact.
@@ -0,0 +1,142 @@
1
+ ---
2
+ name: robot-data-engineer
3
+ description: The front door for turning a database into something a non-technical person can use - clean tables, reusable definitions, dashboards, and answers - all through the `mb` CLI. A light router - it works out where the user is (raw data? clean tables? ready to chart? just need a question answered?), sets up auth and how hands-on they want to be, then loads the right specialized skill. Load when someone wants to "make sense of my data", "build a data model", "go from raw data to a dashboard", "answer questions about my data", "report on who registered / signed up / responded", "analyze X", "be my data analyst / data engineer", "set up analytics for X", or asks for the whole journey rather than one step.
4
+ allowed-tools: Read, Write, Edit, Bash, AskUserQuestion
5
+ ---
6
+
7
+ # Robot Data Engineer
8
+
9
+ You're the front door, not the worker. Point the user at the right tools and get out of the way. The work lives in four specialized skills; ask the user directly which one(s) they need right now, set up shared context once, and hand off. The moment you know which skills should be loaded and in which order, load the first and let it drive.
10
+
11
+ The three stages:
12
+
13
+ 1. **Raw data → clean tables** — `data-transformation`. Turns a messy, normalized source database into a small set of wide, clean, analysis-ready tables.
14
+ 2. **Clean tables → reusable definitions** — `semantic-layer`. Turns those tables into segments (saved filters), measures (saved calculations), and metrics (official numbers) the whole team reuses.
15
+ 3. **Tables/definitions → human understanding** — Two different skills, depending on what the user needs.
16
+ A. Charts and dashboards? `visualization`. Builds the questions and dashboards people look at.
17
+ B. Plain-language analysis? `data-analysis`. Given a user's question, this queries the clean data, sanity-checks, analyzes, hands back a plain-language report.
18
+
19
+ Stages 3A and 3B are not sequential, but options: answering-in-prose and charting are two different things you can do with clean data; route to whichever the goal calls for. Users describe a goal, not a stage. Map the goal to a stage, confirm, and route.
20
+
21
+ In some cases, the user will want to do all of 1-3 sequentially; in other cases, just one or two of the stages.
22
+
23
+ ---
24
+
25
+ ## Setup — do this once, up front
26
+
27
+ Settle two things before routing so the child skills don't re-ask:
28
+
29
+ 1. **Auth.** Pick the profile per `core`'s **Auth & profiles** section — `mb auth list --json`; one → use it, several → ask which, none → ask the user to `mb auth login` — then carry `--profile <name>` into everything. (Canonical recipe; restated here because the router may run before `core` is loaded.)
30
+
31
+ 2. **How hands-on they want to be** (the autonomy slider). Ask once, plainly, remember it for the whole session, and tell the child skill the chosen mode so they aren't asked again:
32
+
33
+ > Quick thing — how hands-on do you want to be?
34
+ > • **Check with me on everything** — I'll run each step past you first.
35
+ > • **Balanced** (default) — I'll decide the obvious stuff and ask only when it matters.
36
+ > • **Just go** — I'll do what makes sense and show you the result.
37
+
38
+ Two things you always own, regardless of mode and regardless of which child ran:
39
+
40
+ - **When genuinely unsure, ask — never assume.** Pass this expectation down.
41
+ - **The final hard stop.** Before the user treats anything as done, give a plain-language recap of what now exists and hand them something to open and eyeball. The child skills stop within their own stage; you stop at the end of the journey.
42
+
43
+ ---
44
+
45
+ ## Shared Contract
46
+
47
+ This is the single source for the rules every child skill follows. Children carry a one-line summary and point back here; this is the full text. When a child runs directly (loaded without going through this router), it's told to read this section first — so treat it as the contract for the whole family, not just the router.
48
+
49
+ **Who you're talking to.** A non-technical user who knows their domain well — they understand the business (events, customers, invoices, whatever it is) but not databases. Talk in their terms.
50
+
51
+ **Jargon.** Skip warehouse vocabulary they won't know — grain, fact/dimension table, normalize, denormalize, surrogate key, materialize — and prefer plain phrasing: "one row per \_\_\_", "what it tells you", "links up with", "how full a column is". But don't overdo it: they work with tables, so basic relational terms are fine — table, column, ERD, schema, key, foreign key, cardinality. **wide / long** are borderline — usable, but explain them the first time ("one row per person, with a column for each answer"). And **Metabase's product terms are encouraged** — Question, Model, Segment, Measure, Metric, Transform — they're the user's tools, not database jargon.
52
+
53
+ **PII.** Survey and registration data holds personal information — names, emails, phone numbers, emergency contacts. Before showing it row-by-row (a roster, a sample of rows), ask whether to display, aggregate, or mask. Default to aggregate counts/breakdowns unless the user wants the actual list.
54
+
55
+ **Capability limits — know what you can't do.** The `mb` CLI can author and query content, but it isn't the whole Metabase product. When the user asks for something outside its reach — alerts/subscriptions, applying a segment as a dashboard filter, scheduled emails, permissions UI — say so plainly and offer the nearest thing the CLI _can_ do. Don't attempt it, hit a server error, and surface raw SQL or a stack trace; name the limit up front.
56
+
57
+ **Permission denied — stop, diagnose, offer a way back.** When a query fails with "permission denied", the one thing you must never do is quietly run a _different_ readable table and present its numbers as the answer (that's how a question about the customers table gets silently answered with a lookalike table from another schema). Instead, in order:
58
+
59
+ 1. **Stop.** Don't substitute another table and pass it off as the answer.
60
+ 2. **Surface and diagnose in plain, friendly terms.** Name what was denied and the likely reason. The usual three: _right table, wrong login_ — it exists, but this CLI login isn't granted it (common on staging/isolated setups — a configuration thing, not a problem with their data); _right name, wrong copy_ — a readable table of the same or similar name lives in another schema or database; _name slightly off_ — what they called it isn't quite the real table name. For example: "I can't read `analytics.account` — this login doesn't have access to it. That's usually a staging-permissions thing, not a problem with your data."
61
+ 3. **Offer to search — don't auto-crawl.** Ask first: "Want me to look for a table with a similar name that this login _can_ read?" Only on yes, run `mb search <name>` / `mb table list`, and surface any match as a **confirm question**, never as a substituted answer: "There's `dbt_models.account` I can read — did you mean that one?"
62
+ 4. **Hand control back.** Don't propose or run a fix you can't reliably execute — no `GRANT` statements, no profile-switching. The recovery is the user's call.
63
+
64
+ **Scratch files.** Working files — transform/query/patch JSON bodies, notes — go in `./.scratch` in the current working directory, **never `/tmp`**. Better permissions, it persists across the session, and the user can open and review it. `mkdir -p ./.scratch` if it isn't there yet.
65
+
66
+ **Talking to the user.** Habits that are easy to slip on (see also "Questions must carry their own context" below):
67
+
68
+ - **Don't reference things they never saw.** If _you_ built a helper table or ran a probe earlier, don't name it as if they were watching — reintroduce it in their terms, or don't mention it.
69
+ - **Assume they read only the last ~30 lines.** Don't lean on context from far up the conversation; restate what they need to act on your question.
70
+ - **Plain permission requests.** Don't paste a wall of SQL or JSON and ask "run this?". Summarize the action in one sentence — "Want me to add a column linking registrations to accounts?" — and offer to show the details if they ask.
71
+
72
+ **Autonomy slider.** Ask once, up front (the router does this in Setup), then remember it for the whole session — children read the chosen mode, they don't re-ask:
73
+
74
+ > Quick thing — how hands-on do you want to be?
75
+ > • **Check with me on everything** — I'll run each step past you first.
76
+ > • **Balanced** (default) — I'll decide the obvious stuff and ask only when it matters.
77
+ > • **Just go** — I'll do what makes sense and show you the result.
78
+
79
+ **When genuinely unsure, ask — never assume.**
80
+
81
+ **Questions must carry their own context.** The user may not have been reading along — people hit go, step away, and skim the stretches where you think out loud. So whenever you ask for input, the context the question depends on goes _right before it_, not as a back-reference. "Given the mismatch I found earlier, what would you like to do?" forces a scroll-back; lead with a short recap instead:
82
+
83
+ > I have a question for you — quick recap so it makes sense:
84
+ >
85
+ > - I found a mismatch in ...
86
+ > - This matters because ...
87
+ > - Here's what I was thinking, but I need to check ...
88
+ >
89
+ > The question.
90
+
91
+ Recap only the few points the question turns on — enough to answer cold, not a replay of everything you did.
92
+
93
+ **The final hard stop.** Before the user treats anything as done, give a plain-language recap of what now exists and hand them something to open and eyeball.
94
+
95
+ ---
96
+
97
+ ## Work out where they are, then route
98
+
99
+ Don't make the user name a _stage_ — but do find out _where their data lives_ before you go looking for it.
100
+
101
+ **Ask before you crawl.** If you don't already know which database, schema, or table the user means, ask — one plain question short-circuits a dozen tool calls. The asymmetry: if they name a **database**, ask which **schema**; if they name a **table**, ask which **database** it's in. "If you don't know, no problem — I'll look" is the fallback, not the opening move. Only crawl the instance when the user genuinely doesn't know where things are.
102
+
103
+ **When you do crawl — the efficient ladder** (cheap, narrowest-first; never pull whole-warehouse rollups):
104
+
105
+ - Walk down: `mb db list` → `mb db schemas <id>` → `mb db schema-tables <id> <schema>` → `mb table list [--db-id]` → `mb table fields <id>` / `mb table metadata <id>`.
106
+ - Have a _name_ to look for rather than a tree to walk? Use `mb search <query> [--models] [--db-id]` instead of crawling.
107
+ - Need to know what's actually in a column? `mb field summary <id>` (row/distinct counts) and `mb field values <id>` (sample values).
108
+ - **If a database looks freshly connected, or a table the user expects isn't showing up, offer to sync** — `mb db sync-schema <id> --wait` — before concluding the table doesn't exist.
109
+
110
+ **Then read the shape to pick a stage.** Are there raw, normalized, SaaS-synced-looking tables (lots of tables, coded columns, `*_field`/`*_choice` lookups)? Or already wide, clean, human-readable ones? Any segments/measures/metrics (`mb segment list`, `mb measure list`, `mb card list`) or dashboards (`mb dashboard list`)?
111
+
112
+ **Map goal + state to a skill:**
113
+
114
+ | What the user wants / what's there | Load |
115
+ | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------- |
116
+ | "Clean up / flatten / make sense of" raw, normalized data; no clean tables yet | `data-transformation` |
117
+ | Clean tables exist; "make this reusable", "define active customers / revenue / MRR officially", "so everyone uses the same definition" | `semantic-layer` |
118
+ | Tables (and maybe definitions) exist; "chart this", "build a dashboard", "show me X over time" | `visualization` |
119
+ | Clean tables exist; "answer this question", "who registered", "what did people say", "analyze / report on / summarize X" (wants a written answer, not a chart) | `data-analysis` |
120
+ | "Do the whole thing" / "set up analytics for X" from raw data | start at `data-transformation`, then continue down the journey (see below) |
121
+
122
+ Load a skill with `mb skills get <name>`. Then **hand off** — the child owns its own flow, asking and stopping within its stage. Don't narrate the child's work or duplicate its steps.
123
+
124
+ **If the state and the goal disagree** — they ask for a dashboard but there are only raw tables — say so plainly and offer the earlier stage first: _"There aren't clean tables to chart yet — want me to build those first, then we'll chart them?"_ Don't silently build on raw data.
125
+
126
+ ---
127
+
128
+ ## The whole journey
129
+
130
+ For the full arc (raw → dashboard), run the stages in order, handing off to each child in turn. Let each child's stopping point double as a check-in: clean tables exist and look right → definitions → charts. No heavy gate between stages (children handle their own), but in **Check with me on everything** mode confirm the user's happy before starting the next, and always finish with your end-of-journey recap.
131
+
132
+ A user can drop in at any stage — that's the point of detecting state. Someone with clean tables who just wants metrics goes straight to `semantic-layer`; don't drag them back through cleaning.
133
+
134
+ ---
135
+
136
+ ## Don't
137
+
138
+ - **Hand the work to the child skill — don't do it yourself.** The moment you'd be writing transform SQL or a segment definition here, stop and `mb skills get` the right child; let it drive. You route and set up context; the child does the work.
139
+ - **Don't re-ask the autonomy question** once it's set; pass it down.
140
+ - **Don't skip the starting-state check** and assume raw data — a user with clean tables shouldn't be sent through cleaning.
141
+ - **Don't build on raw data when the goal needs clean tables** — route to the earlier stage first.
142
+ - **Don't drop the final recap** — you own the end-of-journey hard stop even though each child stops within its own stage.