@metabase/cli 0.1.16 → 0.1.18
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +21 -0
- package/dist/{add-collection-H4LcP-9B.mjs → add-collection-CRZAFCZy.mjs} +5 -5
- package/dist/add-collection-DbalcC3_.mjs +11 -0
- package/dist/{archive-D1mfiftv.mjs → archive-3mEhK9_H.mjs} +8 -7
- package/dist/{archive-C6MDtV1F.mjs → archive-C3RWdM-2.mjs} +8 -6
- package/dist/{archive-CQQWkC5S.mjs → archive-CFHYRF65.mjs} +7 -6
- package/dist/{archive-hN8PfvhX.mjs → archive-De4yVEV8.mjs} +7 -6
- package/dist/{archive-BCXoM1nX.mjs → archive-Z8ZM-ilB.mjs} +9 -7
- package/dist/{archive-CdUS-OG9.mjs → archive-d74nDcp4.mjs} +7 -6
- package/dist/{archive-B3cjzOIK.mjs → archive-errg-EKw.mjs} +8 -7
- package/dist/auth--PWX0oj5.mjs +22 -0
- package/dist/{body-DB2upz6a.mjs → body-BS2s2zDl.mjs} +3 -3
- package/dist/{branches-B6R60Vr1.mjs → branches-mcyT7wI-.mjs} +7 -6
- package/dist/{cancel-oPsWomYc.mjs → cancel-B2l-i2NS.mjs} +6 -5
- package/dist/{cancel-task-CleJVDNI.mjs → cancel-task-CugcIeIi.mjs} +7 -6
- package/dist/{capabilities-BX1rnVuH.mjs → capabilities-N0jo5U7S.mjs} +1 -1
- package/dist/card-D07WYCQY.mjs +26 -0
- package/dist/{card-wCPcuKSi.mjs → card-DEmcRlNO.mjs} +6 -5
- package/dist/{cards-Bw37jizL.mjs → cards-BGkz5BeZ.mjs} +8 -6
- package/dist/cli.mjs +54 -27
- package/dist/collection-Dss2qSEF.mjs +23 -0
- package/dist/{collection-namespace-CUDPh2rF.mjs → collection-namespace-BP_LJrAD.mjs} +2 -2
- package/dist/command-augment-DdZIfx1V.mjs +11 -0
- package/dist/{create-Bj6PuaAW.mjs → create-B4HAEEE0.mjs} +9 -8
- package/dist/{create-BEkuqChu.mjs → create-BOmmMSjF.mjs} +16 -11
- package/dist/{create-DSWjS09p.mjs → create-CD9Rp8qs.mjs} +21 -12
- package/dist/{create-DwoqFYb9.mjs → create-CkRFKtFG.mjs} +9 -8
- package/dist/{create-CZ0_s5ap.mjs → create-Cztp5kES.mjs} +7 -6
- package/dist/{create-CF2Zn4pT.mjs → create-D8KA4zJA.mjs} +10 -9
- package/dist/{create-B-mvVFIl.mjs → create-DhrftYre.mjs} +28 -12
- package/dist/{create-D-qE3Oq8.mjs → create-DzyMvAVy.mjs} +9 -8
- package/dist/{create-Cy2TNnJB.mjs → create-YfGzANNB.mjs} +9 -8
- package/dist/{create-branch-BaZ00MIc.mjs → create-branch-DsoMfx_2.mjs} +7 -6
- package/dist/{create-CfLaBO0V.mjs → create-kWzqvTGR.mjs} +16 -11
- package/dist/{create-B87ZQjNM.mjs → create-wpvDXSSY.mjs} +17 -12
- package/dist/{current-task-Dr5dOD4V.mjs → current-task-WuKPZpZ6.mjs} +7 -6
- package/dist/dashboard-Crh4v6Fq.mjs +35 -0
- package/dist/{dashboard-B4bn3z6t.mjs → dashboard-DOplbKyQ.mjs} +7 -5
- package/dist/{database-BJxGUXhA.mjs → database-D9fftP-i.mjs} +1 -1
- package/dist/db-BCAURQei.mjs +28 -0
- package/dist/{delete-DHrFA1SZ.mjs → delete-BCitmugk.mjs} +8 -7
- package/dist/{delete-DJti0TOA.mjs → delete-Buai0_e_.mjs} +8 -7
- package/dist/{delete-evwMw6Hk.mjs → delete-CmGh0jo1.mjs} +8 -7
- package/dist/{delete-runtime-CZMw_AGX.mjs → delete-runtime-B0ha5QR4.mjs} +3 -3
- package/dist/{delete-table-DvuuwVqu.mjs → delete-table-BOkYxdjA.mjs} +8 -7
- package/dist/{dependencies-DrV31Rj5.mjs → dependencies-XZrEvHUA.mjs} +7 -6
- package/dist/{dirty-C-rkrnVM.mjs → dirty-B-5NDbG9.mjs} +7 -6
- package/dist/document-upxL2nKH.mjs +22 -0
- package/dist/{eid-BeXI-eII.mjs → eid-BCuJRv7a.mjs} +13 -7
- package/dist/{error-CIObEXLY.mjs → error-BaBm-UrT.mjs} +2 -2
- package/dist/{export-CGa-SgEV.mjs → export-Dod2gdyE.mjs} +9 -8
- package/dist/field-CloFa1oe.mjs +21 -0
- package/dist/{fields-Recymu7n.mjs → fields-B0yttrzR.mjs} +8 -7
- package/dist/{get-C6g2C4dM.mjs → get-9p0Gt1c7.mjs} +7 -6
- package/dist/{get-CYh2DBsq.mjs → get-B0OsW6jk.mjs} +7 -6
- package/dist/{get-CHi8tU_1.mjs → get-B8tugTku.mjs} +9 -8
- package/dist/{get-BD_P_Ejc.mjs → get-BkGh9BPd.mjs} +7 -6
- package/dist/{get-BM-d3-zk.mjs → get-BpZ81hX8.mjs} +7 -6
- package/dist/{get-CARdLkmp.mjs → get-CbpHt4Hs.mjs} +7 -6
- package/dist/{get-b4xbfFbA.mjs → get-CfeYa0v5.mjs} +7 -6
- package/dist/{get-DaixPDU9.mjs → get-D-pNrQGA.mjs} +7 -6
- package/dist/{get-rFcAVIch.mjs → get-D7QMqPOK.mjs} +7 -6
- package/dist/{get-BAH_M5vj.mjs → get-DBY5esTW.mjs} +7 -6
- package/dist/{get-Q2WZ79q_.mjs → get-DISgP66L.mjs} +8 -7
- package/dist/{get-Nc5GOs6-.mjs → get-DJ8huA8y.mjs} +8 -6
- package/dist/{get-CXMv-r1p.mjs → get-Dz7fcqoG.mjs} +9 -7
- package/dist/{get-BQxLtFE-.mjs → get-k6M5nGMC.mjs} +7 -6
- package/dist/{get-run-CfQR6ZNa.mjs → get-run-DQJVpDw1.mjs} +7 -6
- package/dist/{get-7fWSU6ow.mjs → get-u5Uq9vms.mjs} +6 -5
- package/dist/git-sync-C9m-OHac.mjs +31 -0
- package/dist/group-BNE_RiH5.mjs +28 -0
- package/dist/{has-remote-changes-Bvyv4FLP.mjs → has-remote-changes-DvQBXudi.mjs} +7 -6
- package/dist/{import-c3o3OAx0.mjs → import-PGS8-DwE.mjs} +9 -8
- package/dist/{input-BXWgdKiS.mjs → input-7Sj85_K7.mjs} +1 -1
- package/dist/is-dirty-DpyKeGAJ.mjs +10 -0
- package/dist/{is-dirty-BE53XwOC.mjs → is-dirty-DyEVFQVJ.mjs} +4 -4
- package/dist/{items-1KkBMiO4.mjs → items-CkFEy2Du.mjs} +9 -8
- package/dist/{key-DwiMOWRQ.mjs → key-bltP32Pm.mjs} +1 -1
- package/dist/library-1AAVbk-K.mjs +24 -0
- package/dist/{list-C4KnM3Rq.mjs → list-82NkvRpI.mjs} +8 -6
- package/dist/{list-Drr4JiWg.mjs → list-B9wXg3qi.mjs} +6 -5
- package/dist/{list-BwdO1_gX.mjs → list-BBxRjuMn.mjs} +6 -5
- package/dist/{list-CSjFJDls.mjs → list-BCYTTCFD.mjs} +6 -5
- package/dist/{list-BqgbrpQQ.mjs → list-BVzu2RIZ.mjs} +6 -5
- package/dist/{list-DUEYX3bX.mjs → list-BWAYDSbQ.mjs} +6 -5
- package/dist/{list-BC2B02IR.mjs → list-BY4S32Lg.mjs} +6 -5
- package/dist/{list-7rwzxX6t.mjs → list-Bb8YRON_.mjs} +6 -5
- package/dist/{list-DIvOPW1g.mjs → list-CJx5q-Yn.mjs} +8 -7
- package/dist/{list-F0vkE22V.mjs → list-CK0p7vvK.mjs} +7 -6
- package/dist/{list-CO5J3SZU.mjs → list-CKzpoTgP.mjs} +8 -7
- package/dist/{list-FR8Q1SzV.mjs → list-CjF12k1G.mjs} +7 -6
- package/dist/{list-D52_BozQ.mjs → list-D-rgDFa5.mjs} +9 -7
- package/dist/{list-DPcPTqFU.mjs → list-Sgo3RfDY.mjs} +6 -5
- package/dist/{list-DHb4vQUM.mjs → list-pQ22nXhQ.mjs} +6 -5
- package/dist/{login-C0Rf2hg0.mjs → login-C6ZAnGHz.mjs} +10 -9
- package/dist/{logout-C7_UON-s.mjs → logout-CWjyY3Y8.mjs} +6 -5
- package/dist/{manifest-B2F8iL7X.mjs → manifest-BVf8P4bl.mjs} +9 -2
- package/dist/measure-Cbly1r0E.mjs +25 -0
- package/dist/{metadata-Db3Kpo-z.mjs → metadata-CW5Lfw5d.mjs} +9 -8
- package/dist/{metadata-FltZq5Ek.mjs → metadata-aQAqseCm.mjs} +8 -7
- package/dist/{command-augment-CAur0XOQ.mjs → notice-DyVl5aYB.mjs} +1 -11
- package/dist/parameter-CiJ4CwWE.mjs +118 -0
- package/dist/parameter-values-DmDOuE-j.mjs +56 -0
- package/dist/{parse-enum-BatHQ-Gs.mjs → parse-enum-BL9i_brN.mjs} +1 -1
- package/dist/{parse-id-B5adfBlS.mjs → parse-id-DlXnOcmP.mjs} +1 -1
- package/dist/{parse-ref-CZr1bYIl.mjs → parse-ref-CB_KvF9h.mjs} +1 -1
- package/dist/{path-BojuJkE4.mjs → path-D8IJ4YrW.mjs} +6 -5
- package/dist/{poll-AduuU55-.mjs → poll-hgnrHBoh.mjs} +2 -2
- package/dist/{poll-task-B00Qwd87.mjs → poll-task-NQNLT_aA.mjs} +2 -2
- package/dist/{preflight-QVPvG_Xg.mjs → preflight-DvaPQHHf.mjs} +4 -4
- package/dist/{process-DsGf7Mg5.mjs → process-j8UHMHc2.mjs} +1 -1
- package/dist/{prompt-Bc_bHSD0.mjs → prompt-C85xd9HR.mjs} +1 -1
- package/dist/{publish-h5RJh6im.mjs → publish-BntmFR6g.mjs} +9 -8
- package/dist/{query-BUkuB4bZ.mjs → query-BQfgKV68.mjs} +10 -8
- package/dist/{query-Dvi-Rksy.mjs → query-DSKQZu91.mjs} +19 -13
- package/dist/{query-result-L5_NrwQR.mjs → query-result-D6mfoVfQ.mjs} +1 -1
- package/dist/{remove-collection-D8ZfB2RN.mjs → remove-collection-BguaI3-6.mjs} +9 -8
- package/dist/{rescan-values-DOsDLrRG.mjs → rescan-values-DjD6IW3J.mjs} +9 -8
- package/dist/{resolve-BQ9vjlNJ.mjs → resolve-Dj2MTBkn.mjs} +1 -1
- package/dist/{run-CeG0KH5W.mjs → run-Bp1yxkBN.mjs} +6 -5
- package/dist/{run-CjhD-Zbr.mjs → run-RQfQj7Rk.mjs} +9 -8
- package/dist/{runs-6k8C6kXF.mjs → runs-CSsatfWb.mjs} +8 -7
- package/dist/{runtime-CmAIahm5.mjs → runtime-oxjmrYoP.mjs} +5 -3
- package/dist/{schema-tables-C45QegaY.mjs → schema-tables-5I5pCxHl.mjs} +8 -7
- package/dist/{schemas-EVwEFuTj.mjs → schemas-CfzFCfBt.mjs} +6 -5
- package/dist/{search-C_uw_D1U.mjs → search-Vo-BqliA.mjs} +11 -5
- package/dist/segment-BuN_IoM8.mjs +25 -0
- package/dist/{selectors-AktxTEMK.mjs → selectors-DBnJsKlW.mjs} +3 -3
- package/dist/{set-C3EAuyb8.mjs → set-1Vh0AF-T.mjs} +9 -8
- package/dist/{set-active-BW6LN6y0.mjs → set-active-C0mUZlNN.mjs} +6 -5
- package/dist/setting-BbdKR-lO.mjs +20 -0
- package/dist/{setup-CYrbNrlG.mjs → setup-CzYFvFK5.mjs} +8 -7
- package/dist/{skills-DJsuBguh.mjs → skills-B6gfH0iR.mjs} +1 -1
- package/dist/{skills-D6xQkmhu.mjs → skills-CR2xyV33.mjs} +3 -3
- package/dist/snippet-Bzo2U9Fv.mjs +22 -0
- package/dist/{stash-C1V2FvJR.mjs → stash-CQHXwBu_.mjs} +9 -8
- package/dist/{status-DY92F9mn.mjs → status-C_7aqTvB.mjs} +6 -5
- package/dist/{status-CxYw6zQM.mjs → status-DXQkM18v.mjs} +8 -7
- package/dist/{summary-DOxgqJoA.mjs → summary-1aNpQc8j.mjs} +7 -6
- package/dist/{sync-schema-CQPfffjU.mjs → sync-schema-BqLp5uTS.mjs} +11 -10
- package/dist/table-BwGOz97O.mjs +22 -0
- package/dist/{table-CDMG0Zi5.mjs → table-DE3i82T_.mjs} +1 -1
- package/dist/transform-B65ZD9-e.mjs +31 -0
- package/dist/transform-job-BWVKXSV6.mjs +25 -0
- package/dist/transform-tag-C4qvmicL.mjs +21 -0
- package/dist/{transforms-DzBJDydn.mjs → transforms-BelyllUL.mjs} +7 -6
- package/dist/{tree-B3f5F_dP.mjs → tree-YvmwGql7.mjs} +6 -5
- package/dist/{unpublish-B5RDeN-V.mjs → unpublish-CHjGLMu9.mjs} +7 -6
- package/dist/{update-CcvDVqNd.mjs → update-BACl_bJS.mjs} +10 -9
- package/dist/{update-PZPNx0Xd.mjs → update-BGMJTpA6.mjs} +17 -12
- package/dist/{update-DOfL_KPx.mjs → update-BKhekuqr.mjs} +22 -13
- package/dist/{update-DSueNZRw.mjs → update-BWlN3QDo.mjs} +10 -9
- package/dist/{update-C0pFSc1B.mjs → update-BnHYSu96.mjs} +18 -13
- package/dist/{update-D698CaeV.mjs → update-C9fndCqb.mjs} +10 -9
- package/dist/{update-I3TA2Tem.mjs → update-CrRz47aj.mjs} +22 -13
- package/dist/{update-BIZ9XhjS.mjs → update-D7vc8GBF.mjs} +17 -12
- package/dist/{update-DudyZ-FP.mjs → update-K-jjQ0Aw.mjs} +11 -10
- package/dist/{update-CiWPEqQ-.mjs → update-Zz01Woxj.mjs} +10 -9
- package/dist/{update-dashcard-BXZ4vS15.mjs → update-dashcard-Dvth-yLC.mjs} +11 -9
- package/dist/{update-D5gioyBa.mjs → update-txUfAJxy.mjs} +10 -9
- package/dist/{upgrade-4cWfLu90.mjs → upgrade-DsXlfPef.mjs} +7 -6
- package/dist/{uuid---pAboNQ.mjs → uuid-ByJmWV7b.mjs} +10 -5
- package/dist/{validate-VawhJ5Sc.mjs → validate-BqNW4Sk1.mjs} +2 -2
- package/dist/{validate-query-CSV-TTnd.mjs → validate-query-CcZVKYPV.mjs} +3 -3
- package/dist/{values-I503dI7K.mjs → values-CIjJI4sz.mjs} +7 -6
- package/dist/{verify-BMhTWW9s.mjs → verify-Jsuc2dat.mjs} +2 -2
- package/dist/{wait-oMSs_IdS.mjs → wait-5ltTFvSa.mjs} +8 -7
- package/dist/{wait-flags-HtCL2l1r.mjs → wait-flags-D6kd_G7c.mjs} +2 -2
- package/package.json +1 -1
- package/skill-data/core/SKILL.md +47 -56
- package/skill-data/dashboard/SKILL.md +107 -0
- package/skill-data/data-workflow/SKILL.md +116 -0
- package/skill-data/{data-analysis/SKILL.md → data-workflow/references/answering-questions.md} +7 -13
- package/skill-data/{data-transformation/SKILL.md → data-workflow/references/building-clean-tables.md} +23 -32
- package/skill-data/{semantic-layer/SKILL.md → data-workflow/references/reusable-definitions.md} +26 -56
- package/skill-data/document/SKILL.md +11 -21
- package/skill-data/git-sync/SKILL.md +39 -39
- package/skill-data/mbql/SKILL.md +19 -31
- package/skill-data/mbql/references/operators.md +13 -3
- package/skill-data/metadata/SKILL.md +78 -0
- package/skill-data/metadata/references/semantic-types.md +83 -0
- package/skill-data/native-sql/SKILL.md +118 -0
- package/skill-data/native-sql/references/template-tags.md +178 -0
- package/skill-data/transform/SKILL.md +30 -46
- package/skill-data/visualization/SKILL.md +8 -8
- package/skill-data/visualization/references/settings.md +14 -0
- package/skills/metabase-cli/SKILL.md +2 -2
- package/dist/add-collection-BqfYL4FU.mjs +0 -10
- package/dist/auth-BaCMFLTA.mjs +0 -19
- package/dist/card-DzH3aK0a.mjs +0 -20
- package/dist/collection-Wagz-ira.mjs +0 -20
- package/dist/dashboard-OUgS1Gi-.mjs +0 -21
- package/dist/db-Btfl5JMZ.mjs +0 -22
- package/dist/document-CU28GfFw.mjs +0 -19
- package/dist/field-BTbzlcyC.mjs +0 -18
- package/dist/git-sync-CBxS2urR.mjs +0 -28
- package/dist/is-dirty-Z-pqyVyB.mjs +0 -9
- package/dist/library-BlbH0xyK.mjs +0 -18
- package/dist/measure-BUedPu4K.mjs +0 -19
- package/dist/segment-CkZUZcWz.mjs +0 -19
- package/dist/setting-C50HEiGG.mjs +0 -17
- package/dist/snippet-BjaWAxCu.mjs +0 -19
- package/dist/table-B35ovbcd.mjs +0 -19
- package/dist/transform-DF79sJ0_.mjs +0 -25
- package/dist/transform-job-PmA_D8gz.mjs +0 -22
- package/dist/transform-tag-CE3cuO1K.mjs +0 -18
- package/skill-data/robot-data-engineer/SKILL.md +0 -142
- /package/dist/{body-flags-D7q87Btw.mjs → body-flags-DWTTxJpP.mjs} +0 -0
- /package/dist/{collection-Deiziuu2.mjs → collection-DrLpA1SO.mjs} +0 -0
- /package/dist/{document-qfwR0r63.mjs → document-1W7NRaO_.mjs} +0 -0
- /package/dist/{field-E0IBy4Uw.mjs → field-CMY_LWUe.mjs} +0 -0
- /package/dist/{measure-BCv5wDDN.mjs → measure-DoJvtCaA.mjs} +0 -0
- /package/dist/{paginate-BexjkjbY.mjs → paginate-FVZUxL4J.mjs} +0 -0
- /package/dist/{render-CkuFkWlQ.mjs → render-BTKnWL0d.mjs} +0 -0
- /package/dist/{revision-message-flag-CP5NFrWQ.mjs → revision-message-flag-CHrJgFFx.mjs} +0 -0
- /package/dist/{segment-BAUuELKs.mjs → segment-TXktTCfU.mjs} +0 -0
- /package/dist/{setting-DhMk0TNo.mjs → setting-m46MUtW5.mjs} +0 -0
- /package/dist/{snippet-D4SyVLKB.mjs → snippet-CtA2Pkoa.mjs} +0 -0
- /package/dist/{transform-MmqHKGU-.mjs → transform-DEF38FWe.mjs} +0 -0
- /package/dist/{transform-job-CtVziW85.mjs → transform-job-CtixL4An.mjs} +0 -0
- /package/dist/{transform-tag-wFiWmiyO.mjs → transform-tag-rsIrckCM.mjs} +0 -0
package/skill-data/{data-analysis/SKILL.md → data-workflow/references/answering-questions.md}
RENAMED
|
@@ -1,16 +1,10 @@
|
|
|
1
|
-
|
|
2
|
-
name: data-analysis
|
|
3
|
-
description: Answer real questions from clean, analysis-ready tables and hand back a plain-language report - an answer-finding task, not chart-building. Read the tables, turn the user's question into queries, run them on the live instance, sanity-check the numbers, write up findings the user can trust. Works over already-clean (wide, human-readable) data - survey/registration answers, event signups, customer lists, anything where the data holds the answer. Use when someone wants to "answer questions about my data", "report on who registered / signed up / responded", "what did people say", "analyze X", "explore this data", or "build me a report". For a non-technical user who knows their domain. Needs charts/dashboards? Use `visualization`. Tables still raw? Use `data-transformation` first.
|
|
4
|
-
allowed-tools: Read, Write, Edit, Bash, AskUserQuestion
|
|
5
|
-
---
|
|
6
|
-
|
|
7
|
-
# Data Analysis
|
|
1
|
+
# Answer questions
|
|
8
2
|
|
|
9
|
-
>
|
|
3
|
+
> Part of the **`data-workflow`** skill — the "answer questions" stage. It assumes that skill's **Shared Contract** (how to communicate, PII, autonomy, permission-denied) and final-recap rule. CLI mechanics: `mbql` / `mb query` for running queries.
|
|
10
4
|
|
|
11
5
|
The user has a question and clean data that already holds the answer. Your job: find the answer, check it's right, and hand it back in plain language. You're an analyst, not a dashboard builder — the deliverable is a **trustworthy written answer**, optionally backed by a saved question they can re-open.
|
|
12
6
|
|
|
13
|
-
This
|
|
7
|
+
This stage assumes the tables are already clean (wide, human-readable). If they're raw and normalized — lots of `*_field`/`*_choice` lookups, coded columns, JSON blobs — stop and route to the **build-clean-tables** stage first (`references/building-clean-tables.md`); don't analyze on top of a mess.
|
|
14
8
|
|
|
15
9
|
---
|
|
16
10
|
|
|
@@ -20,7 +14,7 @@ For each question the user asks:
|
|
|
20
14
|
|
|
21
15
|
1. **Find where the answer lives.** List tables (`mb table list`, `mb db schema-tables <db> <schema>`). Read the columns (`mb table fields <id>`). Clean datasets often ship the same facts two ways — a **wide** table (one row per thing, easy to read) and a **long** table (one row per attribute, easy to aggregate over many-valued answers). Pick the one that fits the question: per-person facts → wide; "which option was most popular" across a multi-select → long.
|
|
22
16
|
|
|
23
|
-
2. **Turn the question into a query.** Write it, run it (`mb query`). Start small — a `count(*)` and a couple of sample rows to confirm you're pointed at the right table and the columns mean what you think. Then write the real query.
|
|
17
|
+
2. **Turn the question into a query.** Write it, run it (`mb query`). Start small — a `count(*)` and a couple of sample rows to confirm you're pointed at the right table and the columns mean what you think. Then write the real query. Query mechanics → `mbql` / `mb query`.
|
|
24
18
|
|
|
25
19
|
3. **Sanity-check before you believe it.** A number with no cross-check is a guess. Confirm row counts against a total you trust, watch for nulls/blanks inflating or deflating a percentage, and re-read the column you grouped on — a `type/Category` column with "confirmed"/"cancelled" means your "how many registered" answer depends on which statuses you counted. State the denominator.
|
|
26
20
|
|
|
@@ -42,7 +36,7 @@ When genuinely unsure which interpretation they mean, ask — never silently pic
|
|
|
42
36
|
|
|
43
37
|
## Survey / registration data — the common shape
|
|
44
38
|
|
|
45
|
-
A lot of "analyze who registered / what did people say" work lands on event or survey data, which has a recognizable shape
|
|
39
|
+
A lot of "analyze who registered / what did people say" work lands on event or survey data, which has a recognizable shape:
|
|
46
40
|
|
|
47
41
|
- A **per-registrant wide table** — name, company, role, status, plus one column per single-answer question. Use it for "who registered", rosters, breakdowns by role/version/company, and any per-person filter.
|
|
48
42
|
- A **long answers table** — one row per (registrant, question, answer). Use it for **multi-select** questions (one person picks several options, so they can't flatten into one wide column) and for "which option was chosen most". Group by the question text, then by the answer value.
|
|
@@ -58,8 +52,8 @@ Three report families cover most asks:
|
|
|
58
52
|
|
|
59
53
|
## Don't
|
|
60
54
|
|
|
61
|
-
- **Don't analyze raw, un-cleaned tables.** If the data is normalized/coded/JSON, route to
|
|
55
|
+
- **Don't analyze raw, un-cleaned tables.** If the data is normalized/coded/JSON, route to the **build-clean-tables** stage first and analyze the clean output.
|
|
62
56
|
- **Don't report a number you didn't sanity-check.** No denominator, no null-check → no answer.
|
|
63
57
|
- **Don't silently pick a scope.** "Registered" vs "confirmed", all-time vs window — state which you used, or ask.
|
|
64
|
-
- **Don't build charts/dashboards here.** A written answer (and maybe one saved question) is the deliverable; if they want it visual, that's `visualization
|
|
58
|
+
- **Don't build charts/dashboards here.** A written answer (and maybe one saved question) is the deliverable; if they want it visual, that's the `visualization` skill.
|
|
65
59
|
- **Don't only count free-text.** Quote the real responses — the words carry the insight a count throws away.
|
|
@@ -1,36 +1,27 @@
|
|
|
1
|
-
|
|
2
|
-
name: data-transformation
|
|
3
|
-
description: Turn a raw, normalized source database into a small set of clean, analysis-ready tables. Claude investigates the source, works out the real-world "things" the data is about (even when each one is scattered across several tables), decodes coded/JSON/translated values into readable text, and builds one wide, denormalized table per thing as Metabase transforms. Designed for a non-technical user who knows their domain. Use whenever someone wants to "clean up", "flatten", "denormalize", "make sense of", or "build analysis-ready tables from" a raw database. This is the strategy skill for modeling a whole database into a set of clean tables; for authoring or running one individual transform (body shape, flags, run inspection), use the `transform` skill instead.
|
|
4
|
-
allowed-tools: Read, Write, Edit, Bash, AskUserQuestion, EnterPlanMode, ExitPlanMode
|
|
5
|
-
---
|
|
1
|
+
# Build clean tables
|
|
6
2
|
|
|
7
|
-
|
|
3
|
+
> Part of the **`data-workflow`** skill — the "build clean tables" stage. It assumes that skill's **Shared Contract** (how to communicate, PII, autonomy, permission-denied) and final-recap rule. CLI mechanics: `core` (auth, `field`/`table` verbs, library publish), `mbql` (transform query bodies), `transform` (creating/running transforms).
|
|
8
4
|
|
|
9
|
-
|
|
5
|
+
**Contents**
|
|
6
|
+
|
|
7
|
+
- [Two kinds of decisions](#two-kinds-of-decisions)
|
|
8
|
+
- [The process](#the-process) — [Phase 0 — Get Oriented](#phase-0--get-oriented), [Phase 1 — Investigate](#phase-1--investigate-in-plan-mode-if-they-choose), [Phase 2 — Present what you found](#phase-2--present-what-you-found-plain-language), [Phase 3 — Iterate](#phase-3--iterate), [Phase 4 — Build, check, hand back](#phase-4--build-check-hand-back)
|
|
9
|
+
- [A worked decode example](#a-worked-decode-example-for-your-reference-not-the-users)
|
|
10
|
+
- [Cleaning checklist](#cleaning-checklist-for-your-reference-not-the-users)
|
|
10
11
|
|
|
11
12
|
Your job: take a raw source database — usually normalized, often synced from a SaaS tool by a connector like Fivetran or Airbyte — and produce a **small set of wide, clean, analysis-ready tables**, one per real-world _thing_ the data is about, built as Metabase **transforms** the user can inspect.
|
|
12
13
|
|
|
13
14
|
Drive everything through the `mb` CLI. Load the skills you'll need:
|
|
14
15
|
|
|
15
16
|
```bash
|
|
16
|
-
mb skills get core # auth, profiles, db/table/field inspection, query
|
|
17
|
+
mb skills get core # auth, profiles, db/table/field inspection, query, library publish
|
|
17
18
|
mb skills get mbql # if you build transform queries in MBQL
|
|
18
19
|
mb skills get transform # creating/running transforms, run inspection
|
|
19
20
|
```
|
|
20
21
|
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
---
|
|
24
|
-
|
|
25
|
-
## Who you're talking to
|
|
26
|
-
|
|
27
|
-
A **non-technical user who knows their domain well** — they understand the business (events, customers, invoices, etc.) but not databases.
|
|
22
|
+
Pick the profile per `core`'s **Auth & profiles** and pass `--profile <name>` to every command. That profile's `url` is the instance's base URL; the browser links below are built from it.
|
|
28
23
|
|
|
29
|
-
|
|
30
|
-
- **Don't lean on raw SQL to communicate.** They may follow a simple `SELECT`, but don't explain work via SQL or ask them to read/write it.
|
|
31
|
-
- Group what you show by **the question a column answers**, never by which source table it came from.
|
|
32
|
-
- Be a **helpful assistant, not an engineer reporting status.** Elide machinery; ask sharp questions that matter.
|
|
33
|
-
- **If you ever ask the user a question, wait for their answer.** They may say "go" and come back later.
|
|
24
|
+
Two communication habits specific to this work, on top of the Shared Contract: **don't communicate through SQL** — they may follow a simple `SELECT`, but never explain your work via SQL or ask them to read or write it; and **group what you show by the question a column answers**, never by which source table it came from. Be a helpful assistant, not an engineer reporting status — elide the machinery, ask the sharp questions that matter.
|
|
34
25
|
|
|
35
26
|
---
|
|
36
27
|
|
|
@@ -42,7 +33,7 @@ Sort every choice into one of these.
|
|
|
42
33
|
|
|
43
34
|
1. Never flatten multi-valued fields into opaque blobs (e.g. three options squished: `"email | phone | text"`). It destroys filterability (the whole point).
|
|
44
35
|
2. Never use jargon with the user. Explain by domain and telos.
|
|
45
|
-
3. Always surface **real data you're about to leave out** proactively, ranked by how much is extant.
|
|
36
|
+
3. Always surface **real data you're about to leave out** proactively, ranked by how much is extant. (Phase 2(c) is where you present it.)
|
|
46
37
|
4. Never guess what schema mean from their name alone. Confirm against actual values, interpret them in context: the table the field belongs to and the relevant domain (e.g., a status on orders ≠ status on subscriptions).
|
|
47
38
|
5. Never silently drop a whole _thing_. Dropping a column is routine; dropping a whole kind-of-thing (e.g. "suppliers") must be surfaced and confirmed.
|
|
48
39
|
6. Never drop columns that link things together. Every table keeps its own id **and** the ids tying it to other tables — alongside the readable labels you copy in, not instead of (label for reading, id for joining). You're building tables about _related_ things, so they **will** be combined ("sales per region", "messages per customer") — dropped ids make that quietly impossible. Keep the ids.
|
|
@@ -68,7 +59,7 @@ Phrase a prudential call as a lean plus a nod:
|
|
|
68
59
|
|
|
69
60
|
### Phase 0 — Get Oriented
|
|
70
61
|
|
|
71
|
-
**Pin down where the data lives — ask before you hunt.** A table or schema name the user mentions tells you _what_ but not _where_: an instance can hold several databases, each with several schemas. Rather than listing them all to find it, just ask — "Which database is this in, and the schema if you know it? No worries if not — I can find it." A confident answer short-circuits a lot of blind searching; "not sure" costs nothing and you fall back to locating it yourself. If you've genuinely looked and still can't find a table the user is sure is there, don't keep digging — one likely reason is Metabase hasn't synced that database's latest schema; gently raise it and let the user run the sync from Metabase.
|
|
62
|
+
**Pin down where the data lives — ask before you hunt.** A table or schema name the user mentions tells you _what_ but not _where_: an instance can hold several databases, each with several schemas. Rather than listing them all to find it, just ask — "Which database is this in, and the schema if you know it? No worries if not — I can find it." A confident answer short-circuits a lot of blind searching; "not sure" costs nothing and you fall back to locating it yourself (use `core`'s narrowest-first crawl ladder). If you've genuinely looked and still can't find a table the user is sure is there, don't keep digging — one likely reason is Metabase hasn't synced that database's latest schema; gently raise it and let the user run the sync from Metabase.
|
|
72
63
|
|
|
73
64
|
As soon as you know which database and schema you're in:
|
|
74
65
|
|
|
@@ -87,7 +78,7 @@ Orientation done, you're about to go heads-down. First, offer two ways to work:
|
|
|
87
78
|
|
|
88
79
|
First path: **enter plan mode** (`EnterPlanMode`). Everything up to the agreed table list — investigate, present, prudential calls, naming (Phases 1–3) — happens inside it, read-only; you exit once, at the approval gate before building (Phase 4). Second path: skip it, shape it conversationally through the same phases. Either way, don't build until the design is settled and approved.
|
|
89
80
|
|
|
90
|
-
Plan mode is a long quiet stretch. So whenever you surface — a question now, the plan at the end — **carry your own context**: recap what it rests on right before you ask, never a back-reference to something said while they were away.
|
|
81
|
+
Plan mode is a long quiet stretch. So whenever you surface — a question now, the plan at the end — **carry your own context**: recap what it rests on right before you ask, never a back-reference to something said while they were away. And whenever you ask a question, **wait for their answer** — they may say "go" and come back later.
|
|
91
82
|
|
|
92
83
|
Then dig in. Don't narrate this — a single "Let me take a look at what's in here — one minute" is enough. Keep it cheap: never pull whole-warehouse rollups (they blow up); use compact column listings, `LIMIT`/sample queries, and `GROUP BY count(*)`.
|
|
93
84
|
|
|
@@ -109,11 +100,11 @@ Three things, in order:
|
|
|
109
100
|
|
|
110
101
|
> **Customers** — one row per customer. Who they are (name, company, location), how they've been in touch, what they've spent, whether they're active or churned.
|
|
111
102
|
|
|
112
|
-
**(b) The full inventory — including what you'd leave out.** Never infer scope silently:
|
|
103
|
+
**(b) The full inventory — including what you'd leave out.** Never infer scope silently (rule 5):
|
|
113
104
|
|
|
114
105
|
> I found 6 kinds of things: **Customers, Orders, Products, Suppliers, Shipments, Returns.** I'd build the first four. **Shipments** and **Returns** also have real data — want those in, or leave them?
|
|
115
106
|
|
|
116
|
-
**(c) What would be set aside — proactively, ranked, two buckets
|
|
107
|
+
**(c) What would be set aside.** This is rule 3 made concrete — proactively, ranked by how much is extant, in two buckets:
|
|
117
108
|
|
|
118
109
|
> Nothing important is lost. A few things set aside:
|
|
119
110
|
> • **Real data** — gift-message text (6 of 10 orders), delivery instructions (most), preferred carrier. Minor, but real — want any kept?
|
|
@@ -127,7 +118,7 @@ Cheap, because nothing's built. Adjust the set of things, what's kept, and the s
|
|
|
127
118
|
|
|
128
119
|
### Phase 4 — Build, check, hand back
|
|
129
120
|
|
|
130
|
-
Design settled — now you build, the first step that writes; plan mode, if you used it, is behind you. Build one wide transform per agreed thing, for how it'll be judged: output that's readable on sight, not just one that runs clean. Each table:
|
|
121
|
+
Design settled — now you build, the first step that writes; plan mode, if you used it, is behind you. Build one wide transform per agreed thing (transform body shape, create, run-with-wait — see `transform`), for how it'll be judged: output that's readable on sight, not just one that runs clean. Each table:
|
|
131
122
|
|
|
132
123
|
- **Denormalized, but the link stays.** Copy in related context so casual reading needs no lookups (a product's name and price on the orders table) — **and keep the linking id beside it** (the product's id too, per rule 6). Use the same id name everywhere a thing appears.
|
|
133
124
|
- **Decoded**: codes and JSON become readable text; bookkeeping columns and soft-deleted rows are gone (filter the source's soft-delete flag — Fivetran's `_fivetran_deleted`, Airbyte's `_ab_cdc_deleted_at`, or a plain `deleted_at`/`is_deleted` — so tombstones never reach clean data).
|
|
@@ -137,13 +128,13 @@ Design settled — now you build, the first step that writes; plan mode, if you
|
|
|
137
128
|
|
|
138
129
|
Then make the links real, not just implied:
|
|
139
130
|
|
|
140
|
-
- **Wire foreign keys between your tables.** Mark each linking id as a foreign key pointing at the id it references
|
|
131
|
+
- **Wire foreign keys between your tables.** Mark each linking id as a foreign key pointing at the id it references — set the column's type to foreign-key and its target so Metabase itself knows the tables connect and can traverse them.
|
|
141
132
|
- **Graft onto existing clean data** the user approved (step 3 / Phase 1): point the linking id at the existing table's id the same way. Link, don't duplicate.
|
|
142
133
|
|
|
143
|
-
**Set the metadata — a transform's output starts blank, and these tables are Library-bound.** A fresh transform table has no descriptions, raw column names, and untyped columns. You worked it out while investigating; don't leave that knowledge stranded in this chat. Set it on the table so the data explains itself inside Metabase (search, the Question editor, Metabot) and is fit to publish:
|
|
134
|
+
**Set the metadata — a transform's output starts blank, and these tables are Library-bound.** A fresh transform table has no descriptions, raw column names, and untyped columns. You worked it out while investigating; don't leave that knowledge stranded in this chat. Set it on the table so the data explains itself inside Metabase (search, the Question editor, Metabot) and is fit to publish. The mechanics — `mb field update` for semantic types / FK targets / display names, `mb table update` for table descriptions, and the feature each edit unlocks — are in the `metadata` skill; the calls about _what_ to set:
|
|
144
135
|
|
|
145
|
-
- **Semantic types — the highest-value piece.** A column's semantic type is what makes Metabase treat it right: `type/Email`, `type/Currency`/`type/Price`, `type/Category` (turns into a filter dropdown), `type/City`/`type/State`/`type/Country`, `type/CreationTimestamp`, `type/Description`. Set it on every column whose meaning you decoded
|
|
146
|
-
- **Descriptions.** A one-line description on each table and every non-obvious column
|
|
136
|
+
- **Semantic types — the highest-value piece.** A column's semantic type is what makes Metabase treat it right: `type/Email`, `type/Currency`/`type/Price`, `type/Category` (turns into a filter dropdown), `type/City`/`type/State`/`type/Country`, `type/CreationTimestamp`, `type/Description`. Set it on every column whose meaning you decoded. A typed column shows money as money, offers a filter dropdown, and lands on the right chart axis for everyone downstream; an untyped one is a guess.
|
|
137
|
+
- **Descriptions.** A one-line description on each table and every non-obvious column.
|
|
147
138
|
- **Display names.** When a cleaned-up column name still isn't plain English, set a readable `display_name`.
|
|
148
139
|
|
|
149
140
|
When the semantics are **already spelled out** — the user is porting dbt models (the `schema.yml` carries column descriptions and types), or you settled each field's meaning together here — that documentation _is_ the metadata. Carry it straight onto the tables and fields rather than letting it evaporate.
|
|
@@ -154,7 +145,7 @@ When refining a built transform _with_ the user, open its inspector so you're lo
|
|
|
154
145
|
|
|
155
146
|
**Pass 1 — Correctness (did it run right).** After each transform runs, run quick ad-hoc tests against what Phase 0 led you to expect: row counts in the right ballpark, decoded columns readable (no stray codes), linking ids that resolve to the other tables, no column unexpectedly all-null or blown up in count. Treat surprises as bugs to chase, not noise. A table that can't combine with the others — a dropped id, or the same id named two ways — is a silent failure; catch it here.
|
|
156
147
|
|
|
157
|
-
**Pass 2 — Fitness (is it nice to use).** Correct isn't the bar; _usable_ is. `SELECT * FROM <table> LIMIT 20` and read every column as if you'd never seen the source: would a
|
|
148
|
+
**Pass 2 — Fitness (is it nice to use).** Correct isn't the bar; _usable_ is. `SELECT * FROM <table> LIMIT 20` and read every column as if you'd never seen the source: would a business reader find each one readable? Smells that say not-yet, even though nothing errored:
|
|
158
149
|
|
|
159
150
|
- a multi-valued column still a raw JSON/array blob or `["Email","SMS"]` text — rule 1 never actually got resolved;
|
|
160
151
|
- decoded answers still carrying raw ids with no readable label, or one cryptic column per code;
|
|
@@ -174,7 +165,7 @@ Then report plainly:
|
|
|
174
165
|
|
|
175
166
|
End on that connection map: it's what the user reads to trust the result, and what lets whatever they build next join on the right ids instead of guessing.
|
|
176
167
|
|
|
177
|
-
These clean tables are exactly what belongs in the **Library** — published tables appear first when anyone picks a data source, so people start from your curated set, not the raw source. If the user wants that,
|
|
168
|
+
These clean tables are exactly what belongs in the **Library** — published tables appear first when anyone picks a data source, so people start from your curated set, not the raw source. If the user wants that, publish the polished tables to the Library (`mb library publish` / `mb library create` mechanics, premium feature, and permissions are in `core`). Defining reusable segments / measures / metrics on top is the **reusable-definitions** stage (`references/reusable-definitions.md` in this skill).
|
|
178
169
|
|
|
179
170
|
---
|
|
180
171
|
|
package/skill-data/{semantic-layer/SKILL.md → data-workflow/references/reusable-definitions.md}
RENAMED
|
@@ -1,73 +1,47 @@
|
|
|
1
|
-
|
|
2
|
-
name: semantic-layer
|
|
3
|
-
description: Turn clean, analysis-ready tables into a shared vocabulary the org reuses - Metabase segments (saved filters, e.g. active customers), measures (saved calculations, e.g. net revenue), and metrics (official numbers, e.g. monthly recurring revenue) - so people stop reinventing the same definition five ways. Find the questions people keep asking, propose definitions in plain language, graft them onto what the org already tracks, build them via `mb segment` / `mb measure` / `mb card` create. For a non-technical user who knows their domain. Load when someone wants to "make this reusable", "define X officially", "standardize how we calculate Y", or "create a segment / measure / metric". Strategy skill for designing reusable definitions; for raw `mb segment` / `mb measure` mechanics, use `core`.
|
|
4
|
-
allowed-tools: Read, Write, Edit, Bash, AskUserQuestion
|
|
5
|
-
---
|
|
1
|
+
# Define reusable metrics
|
|
6
2
|
|
|
7
|
-
|
|
3
|
+
> Part of the **`data-workflow`** skill — the "define reusable metrics" stage. It assumes that skill's **Shared Contract** (how to communicate, PII, autonomy, permission-denied) and final-recap rule. CLI mechanics: `core` (the `segment`/`measure` verbs, `revision_message`, library publish), `mbql` (definition bodies).
|
|
8
4
|
|
|
9
|
-
|
|
5
|
+
- [Autonomy applied here](#autonomy-applied-here)
|
|
6
|
+
- [Two kinds of decisions](#two-kinds-of-decisions)
|
|
7
|
+
- [The process](#the-process) — Phases 0–3
|
|
8
|
+
- [A worked example](#a-worked-example-for-your-reference-not-the-users)
|
|
10
9
|
|
|
11
10
|
Your job: take the clean, analysis-ready tables that already exist and turn the **questions people keep asking** into **shared, reusable definitions** — so "active customer", "net revenue", and "monthly recurring revenue" mean one thing across the whole organization, not five slightly-different things in five people's saved questions.
|
|
12
11
|
|
|
13
12
|
You build three kinds of reusable thing. These are real Metabase features with real names — **use the Metabase names** (segment, measure, metric) and teach them to the user as you go. They're product vocabulary, not jargon. Pair the name with a plain gloss the first time, then use it freely:
|
|
14
13
|
|
|
15
14
|
- **Segment** — a saved filter on a table. A reusable row-selector: "Active customers", "orders over $100", "EU shipments". People pick it from the **Filter** block in the query builder instead of re-typing the conditions. (Docs: <https://www.metabase.com/docs/latest/data-studio/segments>.)
|
|
16
|
-
- **Measure** — a saved aggregation on a table. A reusable calculation: "Net Promoter Score", "average order value". People pick it from the **Summarize** block instead of re-writing the formula.
|
|
15
|
+
- **Measure** — a saved aggregation on a table. A reusable calculation: "Net Promoter Score", "average order value". People pick it from the **Summarize** block instead of re-writing the formula. (Docs: <https://www.metabase.com/docs/latest/data-studio/measures>.)
|
|
17
16
|
- **Metric** — a reusable aggregation that lives in a **collection** (a folder), not bolted to a table. "Monthly recurring revenue", "weekly active users". It's the org's official definition of an important number, can be saved into the **Library**, and can carry a default time dimension for charting. (Docs: <https://www.metabase.com/docs/latest/data-modeling/metrics>.)
|
|
18
17
|
|
|
19
18
|
Introduce each like: _"I'll save this as a **segment** — that's Metabase's word for a reusable filter, so you can pull up active customers with one click anytime."_ After that, just say "segment".
|
|
20
19
|
|
|
21
|
-
This
|
|
20
|
+
This stage runs **after** the analysis-ready tables exist (make the table wider first — the **build-clean-tables** stage, `references/building-clean-tables.md`; the `transform` skill has the mechanics). Segments and measures only reach one table (hard rule 4 below), so a semantic layer on raw, normalized tables is nearly useless: a real answer rarely lives in a single raw table. **Wide clean tables first, segments/measures/metrics second.**
|
|
22
21
|
|
|
23
|
-
|
|
22
|
+
Load the CLI skills you'll need — `mb skills get core` (auth, profiles, inspection, the `segment`/`measure`/`library` verb mechanics) and `mb skills get mbql` (the definition bodies). Auth and scratch files follow `core`'s recipe: resolve the profile and carry `--profile <name>` into every command.
|
|
24
23
|
|
|
25
|
-
|
|
26
|
-
mb skills get core # auth, profiles, db/table/field inspection, query, search
|
|
27
|
-
mb skills get mbql # the definition bodies (filters and aggregations) are MBQL 5
|
|
28
|
-
```
|
|
24
|
+
## Autonomy applied here
|
|
29
25
|
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
---
|
|
33
|
-
|
|
34
|
-
## Who you're talking to
|
|
35
|
-
|
|
36
|
-
A **non-technical user who knows their domain well.** They know the business — who an "active" customer is, what counts as "revenue" — but not databases. So:
|
|
37
|
-
|
|
38
|
-
- **Teach the words a curious non-engineer can follow; skip the deep-internals jargon.** Two sets are fine and worth teaching: Metabase product terms (**segment, measure, metric, collection, Library, the Filter / Summarize blocks**) and common data words a domain user can reasonably learn (**table, column, foreign key, schema, join, filter, row**) — gloss them once, then use them. Avoid **deep-internals jargon** that buys nothing for this user: grain, cardinality, normalize/denormalize, surrogate key, MBQL, `table_id`, materialize. Prefer the plain effect when it's clearer ("this number needs data from two tables" reads easier than "this needs a join across two fact tables") — but you don't have to contort around "foreign key" or "schema".
|
|
39
|
-
- **Talk about the question, then name the object.** Lead with what it does for them, then attach the term: _"I'll save 'big orders' as a segment so you can pull them up with one click."_ Not a bare "I'll create a segment on `table_id` 235."
|
|
40
|
-
- **Be a helpful colleague, not an engineer reporting status.** Elide the wiring (ids, query bodies, the CLI). Ask the one question that actually matters.
|
|
41
|
-
|
|
42
|
-
---
|
|
43
|
-
|
|
44
|
-
## Autonomy — honor the mode the user set
|
|
45
|
-
|
|
46
|
-
The user already picked an autonomy mode (the router's Shared Contract asks the slider once, up front — don't re-ask). Apply it to building definitions:
|
|
26
|
+
The user already set an autonomy mode (the `data-workflow` autonomy slider — don't re-ask, don't redefine it). How it lands on building definitions:
|
|
47
27
|
|
|
48
28
|
| Mode | What you do |
|
|
49
29
|
| ----------------------- | ----------------------------------------------------------------------------------------------------------- |
|
|
50
30
|
| **Check on everything** | Confirm every single definition (name + plain description) before building it. |
|
|
51
|
-
| **Balanced** (default) | Build the obvious ones; ask only on the judgment calls (the prudential list
|
|
31
|
+
| **Balanced** (default) | Build the obvious ones; ask only on the judgment calls (the prudential list) and anything ambiguous. |
|
|
52
32
|
| **Just go** | Build the whole set, surface judgment calls as "here's what I picked and why — say the word to change any." |
|
|
53
33
|
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
1. **When you're genuinely unsure — ask. Never assume.** "Just go" means _decide the obvious_, not _guess on the unclear_. A wrong-but-confident definition of "active customer" is worse than a one-line question.
|
|
57
|
-
2. **The final gate is a hard stop (see Phase 3).** No mode auto-publishes. You always stop, recap in plain language, and hand the user something to eyeball before anything goes live.
|
|
58
|
-
|
|
59
|
-
---
|
|
34
|
+
Two things never bend in any mode: when genuinely unsure, **ask** (the Shared Contract's rule — "Just go" means decide the obvious, not guess on the unclear); and the final gate is a **hard stop** (Phase 3) — no mode auto-publishes.
|
|
60
35
|
|
|
61
36
|
## Two kinds of decisions
|
|
62
37
|
|
|
63
38
|
**Hard rules — absolutes, never ask:**
|
|
64
39
|
|
|
65
40
|
1. **Never invent what a word means — pin it to real data.** "Active customer" is not yours to define. Before you build a segment for it, find out (from the user, or from how the data actually behaves) what _they_ mean: ordered in the last 90 days? Has a live subscription? Logged in this month? Confirm against actual values, then build to that. A definition built on a guessed meaning is a silent lie everyone then trusts.
|
|
66
|
-
2. **Keep the language at the level
|
|
41
|
+
2. **Keep the language at the level the Shared Contract sets.** Metabase terms and common data words (table, column, foreign key, schema, join) are fine and worth teaching; deep-internals jargon (grain, cardinality, surrogate key, `table_id`) is not.
|
|
67
42
|
3. **Don't bury filters inside measures.** A measure should aggregate _what it's given_; let the user combine it with a segment at question time, rather than welding a filter into the measure. Welded-in filters collide and confuse when someone applies their own filter on top — and the metrics doc explicitly recommends against it. (Use conditional forms like `SumIf`/`CountIf` for "sum only the paid ones" — that's part of the measure's formula, not a hidden row filter.)
|
|
68
|
-
4. **Respect where each thing can reach.** Segments and measures work **only** on a question built _directly_ on their own table — not through a join, not on a question-built-on-a-question (the Limitations sections of both docs say so).
|
|
69
|
-
5. **
|
|
70
|
-
6. **Every definition keeps a clear, plain name and a one-line description in the user's words.** The name is what they'll see in a menu six weeks from now with no memory of this conversation. "Active customers (ordered in last 90 days)" beats "active_seg_v2".
|
|
43
|
+
4. **Respect where each thing can reach (single-table reach).** Segments and measures work **only** on a question built _directly_ on their own table — not through a join, not on a question-built-on-a-question (the Limitations sections of both docs say so). A metric is data-source-bound the same way: defined on table X, it appears only on questions built on table X, not on anything derived from it. If a definition needs more than one table's worth of data, you do **not** force a join into it — you make the table wider first (the **build-clean-tables** stage, `references/building-clean-tables.md`; the `transform` skill has the mechanics), then define on that. Quietly building a segment/measure/metric that silently won't show up where the user expects is a hard-rule violation.
|
|
44
|
+
5. **Every definition keeps a clear, plain name and a one-line description in the user's words.** The name is what they'll see in a menu six weeks from now with no memory of this conversation. "Active customers (ordered in last 90 days)" beats "active_seg_v2".
|
|
71
45
|
|
|
72
46
|
**Prudential calls — genuinely contextual, state your lean, let the user decide** (skip the ask in "Just go" mode — pick your lean, flag it):
|
|
73
47
|
|
|
@@ -76,7 +50,7 @@ The user already picked an autonomy mode (the router's Shared Contract asks the
|
|
|
76
50
|
- "Let me add up revenue the same way everywhere, on this table" → a **measure** on the table.
|
|
77
51
|
- "Revenue is an _official company number_ people pull onto dashboards" → a **metric** in a collection, with a default month-by-month view so it charts cleanly. Lean: make it a metric when it's a headline figure the org reuses across many questions/dashboards; keep it a measure when it's a table-local convenience.
|
|
78
52
|
- **Where the metric lives.** Metrics sit in a collection (folder). Lean: put the org's blessed ones in the shared **Library** so they surface prominently; keep experimental ones in a working collection until trusted.
|
|
79
|
-
- **Publish the official tables to the Library.** The clean, analysis-ready tables your definitions sit on are the org's official starting points — the **Library** is how you mark them as such. Tables published to the Library's **Data** section appear _first_ when anyone picks a data source, nudging people toward your curated tables instead of raw warehouse ones. Lean: publish the wide clean tables you built the semantic layer on; hold back raw or half-built ones.
|
|
53
|
+
- **Publish the official tables to the Library.** The clean, analysis-ready tables your definitions sit on are the org's official starting points — the **Library** is how you mark them as such. Tables published to the Library's **Data** section appear _first_ when anyone picks a data source, nudging people toward your curated tables instead of raw warehouse ones. Lean: publish the wide clean tables you built the semantic layer on; hold back raw or half-built ones. Surface which tables you'd publish and confirm. (Library is a Pro/Enterprise feature; only admins and data analysts can publish — mechanics in `core`.)
|
|
80
54
|
- **Default time dimension for a metric.** A monthly default makes it chart nicely on a dashboard, but doesn't lock anyone out of other groupings. Lean: set a sensible default (usually month) for anything headline; leave it off for raw counts that aren't inherently time-series.
|
|
81
55
|
- **How strict a segment is.** "Active" = last 30 vs 90 days is a real business call with no right answer from the data alone. Lean: surface the few reasonable thresholds with how many rows each catches, let the user pick.
|
|
82
56
|
|
|
@@ -84,8 +58,6 @@ Phrase a prudential call as a lean plus a nod:
|
|
|
84
58
|
|
|
85
59
|
> "I'd save 'revenue' as a metric — Metabase's term for an official, reusable number — rather than a table-only measure, since people pull it onto dashboards a lot. Good?"
|
|
86
60
|
|
|
87
|
-
---
|
|
88
|
-
|
|
89
61
|
## The process
|
|
90
62
|
|
|
91
63
|
### Phase 0 — Understand what's reusable (quietly)
|
|
@@ -94,9 +66,9 @@ Don't narrate. One "Let me see what's here and how people are already slicing it
|
|
|
94
66
|
|
|
95
67
|
1. **Confirm the analysis-ready tables exist.** List tables; find the wide, clean ones (a transform step's output). If the user is pointing you at raw normalized tables, say so plainly and suggest building the clean table first — don't build a hobbled semantic layer on raw data.
|
|
96
68
|
2. **Find the questions people keep asking.** Search existing saved questions and dashboards (`mb search`, `mb card list`) for repeated filters and repeated calculations — the same "status = active" written eleven times, five hand-rolled versions of revenue. Those repeats _are_ the semantic layer waiting to be named. This is the highest-signal input; mine it before proposing anything.
|
|
97
|
-
3. **Learn the real meanings.** For every candidate segment ("active", "churned", "high-value"), find what the words map to in actual values — distinct values of a status column, the spread of an amount column.
|
|
69
|
+
3. **Learn the real meanings.** For every candidate segment ("active", "churned", "high-value"), find what the words map to in actual values — distinct values of a status column, the spread of an amount column. Pin every definition to real data (hard rule 1).
|
|
98
70
|
4. **Graft onto what the org already tracks.** This is the part a model does worst and a human does best, so lean on the user: a new definition is far more useful when it lines up with the entities and language the organization _already_ uses. Before inventing "customer health score", ask whether there's already a notion of an active/at-risk customer in their world, and match it. Isolated definitions that don't connect to the existing model are low-value. Ask; don't infer the connection from column names.
|
|
99
|
-
5. **Check reach before promising.** For each candidate, confirm it can actually live where it needs to: a single-table segment/measure must sit on the table people will build questions on; a multi-table answer needs a wider table first (hard
|
|
71
|
+
5. **Check reach before promising.** For each candidate, confirm it can actually live where it needs to: a single-table segment/measure must sit on the table people will build questions on; a multi-table answer needs a wider table first (hard rule 4). Catch this now, not after building something that won't appear.
|
|
100
72
|
|
|
101
73
|
### Phase 1 — Propose the shared vocabulary (plain language)
|
|
102
74
|
|
|
@@ -120,16 +92,16 @@ Then surface what you're _not_ saving and why ("I left 'orders this week' alone
|
|
|
120
92
|
|
|
121
93
|
### Phase 2 — Iterate (cheap, nothing built yet)
|
|
122
94
|
|
|
123
|
-
Adjust names, meanings, thresholds, and which-kind-of-thing until the user is happy. Re-confirm the final list in one short recap. If a definition turns out to need more than one table, say so plainly and point back to making the table wider — don't smuggle in a join.
|
|
95
|
+
Adjust names, meanings, thresholds, and which-kind-of-thing until the user is happy. Re-confirm the final list in one short recap. If a definition turns out to need more than one table, say so plainly and point back to making the table wider (hard rule 4) — don't smuggle in a join.
|
|
124
96
|
|
|
125
97
|
### Phase 3 — Build, verify quietly, then hard-stop
|
|
126
98
|
|
|
127
|
-
Build each agreed definition.
|
|
99
|
+
Build each agreed definition. The verb mechanics (create/update flags, the `revision_message` audit-note rule on `update`, never delete-and-recreate) live in `core`; the definition bodies live in `mbql`:
|
|
128
100
|
|
|
129
|
-
- **Segment** → `mb segment create`.
|
|
130
|
-
- **Measure** → `mb measure create`.
|
|
131
|
-
- **Metric** → `mb card create` with the metric shape (`type: "metric"`) — it lives in a **collection**, carries
|
|
132
|
-
- **Publish the official tables** → `mb library create`
|
|
101
|
+
- **Segment** → `mb segment create`. A flat MBQL filter clause on a table.
|
|
102
|
+
- **Measure** → `mb measure create`. **Exactly one** aggregation on a table.
|
|
103
|
+
- **Metric** → `mb card create` with the metric shape (`type: "metric"`) — it lives in a **collection**, carries the aggregation plus an optional default time dimension. Put org-blessed ones in the Library collection.
|
|
104
|
+
- **Publish the official tables** → `mb library create` then `mb library publish` (mechanics in `core`) to move the clean tables your definitions sit on into the Library's **Data** section, so people start from your curated set, not raw warehouse tables.
|
|
133
105
|
|
|
134
106
|
Then **verify what the user can't see**, before you hand back:
|
|
135
107
|
|
|
@@ -158,14 +130,12 @@ Then **stop. Hard gate — every mode, no exceptions.** Recap in plain language
|
|
|
158
130
|
|
|
159
131
|
End on that plain-language map. It's what the user reads to trust the result — and it's what stops a wrong definition from quietly propagating into everything built next.
|
|
160
132
|
|
|
161
|
-
---
|
|
162
|
-
|
|
163
133
|
## A worked example (for your reference, not the user's)
|
|
164
134
|
|
|
165
135
|
User: _"Everyone calculates 'active users' differently — can you make it official?"_
|
|
166
136
|
|
|
167
137
|
- **Don't** create a segment from the phrase alone. **Find the real meaning first:** search existing questions — three people filter on "last seen in the last 30 days", two on "subscription status = active". That's the ambiguity to resolve. Ask: "I see two takes on 'active' — seen in the last 30 days, or has a live subscription. Which do you mean?" (hard rule 1).
|
|
168
|
-
- They say "live subscription, and seen in the last 30 days." **Check reach:** both pieces of info must live on the one table people build questions on. If subscription status and last-seen sit on two different tables, a single segment can't span them (hard rule 4) — to the user: "those two facts live in different places right now, so I'll widen your Customers table to carry both first, then save the filter on it."
|
|
138
|
+
- They say "live subscription, and seen in the last 30 days." **Check reach:** both pieces of info must live on the one table people build questions on. If subscription status and last-seen sit on two different tables, a single segment can't span them (hard rule 4) — to the user: "those two facts live in different places right now, so I'll widen your Customers table to carry both first, then save the filter on it." Make the table wider first (the **build-clean-tables** stage, `references/building-clean-tables.md`; the `transform` skill has the mechanics), then the segment on the wide table.
|
|
169
139
|
- Build it as a segment on the wide table. **Verify** the row count is plausible. **Recap** plainly and stop: "Saved **Active users** — live subscription and seen in the last 30 days — as a segment on your Customers table; it's in the Filter block there. Have a look before you build on it."
|
|
170
140
|
|
|
171
141
|
The shape recurs: a word people use loosely → pin it to real values → check it can live where they'll use it → build → verify → hard-stop with a plain recap.
|
|
@@ -6,9 +6,9 @@ allowed-tools: Read, Write, Edit, Bash, AskUserQuestion
|
|
|
6
6
|
|
|
7
7
|
# Documents
|
|
8
8
|
|
|
9
|
-
A **document** is a Metabase rich-text page (a "report" / notebook) that mixes prose with embedded saved questions and links to other Metabase entities. The body is a **TipTap** JSON tree (TipTap is the editor; the wire format is ProseMirror JSON,
|
|
9
|
+
A **document** is a Metabase rich-text page (a "report" / notebook) that mixes prose with embedded saved questions and links to other Metabase entities. The body is a **TipTap** JSON tree (TipTap is the editor; the wire format is ProseMirror JSON, stored under `content_type: "application/json+vnd.prose-mirror"`).
|
|
10
10
|
|
|
11
|
-
This skill covers authoring the body and driving the verbs.
|
|
11
|
+
This skill covers authoring the body and driving the verbs. Flag conventions, body-input precedence, `./.scratch`, and `mb uuid` live in `core` (`mb skills get core`).
|
|
12
12
|
|
|
13
13
|
## Command surface
|
|
14
14
|
|
|
@@ -20,19 +20,15 @@ mb document update <id> --file patch.json --profile <name> --json # PATCH sema
|
|
|
20
20
|
mb document archive <id> --profile <name> --json # soft-delete (PUT archived:true)
|
|
21
21
|
```
|
|
22
22
|
|
|
23
|
-
- `list` returns the standard envelope (`{data, returned, total}`). The compact item is `{id, name, collection_id, archived, creator_id, can_write}`
|
|
23
|
+
- `list` returns the standard envelope (`{data, returned, total}`). The compact item is `{id, name, collection_id, archived, creator_id, can_write}` and omits the (potentially huge) `document` body — pull the body with `get --full`.
|
|
24
24
|
- `archive` is the only delete, mirroring `card` / `dashboard`. **Unarchive** with `mb document update <id> --body '{"archived":false}'`.
|
|
25
|
-
- `update` is PATCH — send only the keys you want to change
|
|
25
|
+
- `update` is PATCH — send only the keys you want to change: `name`, `document`, `collection_id`, `collection_position`, `archived`, and `cards` (inline card creation works on update too, not just create — see below). Replacing `document` replaces the **whole** body; there is no partial-node patch.
|
|
26
26
|
|
|
27
27
|
## Node ids (`_id`)
|
|
28
28
|
|
|
29
|
-
The editor anchors only these node types with an `_id` (a UUID): `paragraph`, `heading`, `codeBlock`, `orderedList`, `bulletList`, `blockquote`, `cardEmbed`, `supportingText`. **`create`/`update` require a non-empty `_id` on every node of those types** and
|
|
29
|
+
The editor anchors only these node types with an `_id` (a UUID): `paragraph`, `heading`, `codeBlock`, `orderedList`, `bulletList`, `blockquote`, `cardEmbed`, `supportingText`. **`create`/`update` require a non-empty `_id` on every node of those types** — the CLI validates before sending and rejects a body missing any (`every … node needs a non-empty string _id (mint with mb uuid)`). Other node types (`doc`, `text`, `listItem`, `resizeNode`, `flexContainer`, …) take no `_id` and are left alone. Without them the editor backfills ids when the document opens, which makes a freshly-saved document show a spurious "unsaved changes" prompt.
|
|
30
30
|
|
|
31
|
-
|
|
32
|
-
mb uuid --count 5 --json # → ["…", …] one UUID per id-bearing node
|
|
33
|
-
```
|
|
34
|
-
|
|
35
|
-
Set each as that node's `attrs._id`. Without them the editor backfills ids when the document opens, which makes a freshly-saved document show a spurious "unsaved changes" prompt.
|
|
31
|
+
Mint the ids with `mb uuid --count <n> --json` (→ `["…", …]`), one per id-bearing node, and set each as that node's `attrs._id`.
|
|
36
32
|
|
|
37
33
|
## Body shape (create / update)
|
|
38
34
|
|
|
@@ -85,12 +81,12 @@ Every node is `{ "type": string, "attrs"?: object, "content"?: [nodes], "text"?:
|
|
|
85
81
|
- **`resizeNode`** — wraps a single `cardEmbed` or `flexContainer` to make it resizable (no `_id`). `attrs: { "height": <px>, "minHeight": <px> }`, `content` is exactly one `cardEmbed`/`flexContainer`.
|
|
86
82
|
- **`flexContainer`** — a horizontal row of 1–3 `cardEmbed` / `supportingText` cells side by side (no `_id`). `attrs.columnWidths` is an array of width percentages.
|
|
87
83
|
- **`supportingText`** — a text column that sits next to a card inside a `flexContainer` (id-bearing); `content` is the usual block nodes (`paragraph`, `heading`, lists, …).
|
|
88
|
-
- **`smartLink`** — an inline reference to a Metabase entity (renders as a live chip). Inline, atomic, no `_id`. `attrs: { "entityId": <id>, "model": <model>, "label": <string|null>, "href": <relative-path> }`. `model` ∈ `card`, `dataset`, `metric`, `dashboard`, `collection`, `table`, `database`, `document`, `transform`, `segment`, `
|
|
84
|
+
- **`smartLink`** — an inline reference to a Metabase entity (renders as a live chip). Inline, atomic, no `_id`. `attrs: { "entityId": <id>, "model": <model>, "label": <string|null>, "href": <relative-path> }`. `model` ∈ `card`, `dataset`, `metric`, `dashboard`, `collection`, `table`, `database`, `document`, `transform`, `segment`, `user`, `action`, `indexed-entity` (plus `measure` on v60+).
|
|
89
85
|
- **`metabot`** — an inline Metabot prompt block.
|
|
90
86
|
|
|
91
87
|
## Embedding an existing card
|
|
92
88
|
|
|
93
|
-
A document embedding
|
|
89
|
+
Find the id with `mb card list --profile <name> --json` (or `mb search --models card "<text>"`), then reference it in a `cardEmbed`. A document embedding existing card 114 under a heading (only the id-bearing nodes carry `_id`):
|
|
94
90
|
|
|
95
91
|
```json
|
|
96
92
|
{
|
|
@@ -111,8 +107,6 @@ A document embedding an existing card (id 114) under a heading (only the id-bear
|
|
|
111
107
|
}
|
|
112
108
|
```
|
|
113
109
|
|
|
114
|
-
To embed an existing card, find its id with `mb card list --profile <name> --json` (or `mb search --models card "<text>"`), then reference it in a `cardEmbed`.
|
|
115
|
-
|
|
116
110
|
## Creating brand-new cards inline with the document
|
|
117
111
|
|
|
118
112
|
You can create cards atomically with the document instead of pre-creating them. Reference each new card by a **negative** id in its `cardEmbed.attrs.id`, then supply the card definitions in a top-level `cards` map keyed by the same negative ids. The server creates the real cards and rewrites the negative ids to the real positive ids in the stored body.
|
|
@@ -138,11 +132,11 @@ You can create cards atomically with the document instead of pre-creating them.
|
|
|
138
132
|
}
|
|
139
133
|
```
|
|
140
134
|
|
|
141
|
-
Each entry in `cards` needs at least `{name, dataset_query, display, visualization_settings}` (these are card definitions, not TipTap nodes, so they take no `_id`). Author the `dataset_query` with
|
|
135
|
+
Each entry in `cards` needs at least `{name, dataset_query, display, visualization_settings}` (these are card definitions, not TipTap nodes, so they take no `_id`). Author the `dataset_query` with `mbql` (`mb skills get mbql`) and the `visualization_settings` with `visualization`. For most edits, prefer embedding cards that already exist (a plain positive `id` in `cardEmbed`) — inline creation is for "build the report and its questions in one shot".
|
|
142
136
|
|
|
143
137
|
## Iterating on a document
|
|
144
138
|
|
|
145
|
-
`update` replaces the whole `document` body, so the safe loop is **read → edit → write**. A fetched body already carries `_id`s on its id-bearing nodes
|
|
139
|
+
`update` replaces the whole `document` body, so the safe loop is **read → edit → write**. A fetched body already carries `_id`s on its id-bearing nodes — preserve them, and only mint new ones for id-bearing nodes you add. Don't hand-merge a partial node tree into a live document; pull the current `document`, mutate the array, and PUT the whole thing back.
|
|
146
140
|
|
|
147
141
|
```bash
|
|
148
142
|
mb document get <id> --full --profile <name> --json | jq '.document' > ./.scratch/body.json
|
|
@@ -151,12 +145,8 @@ jq -n --slurpfile d ./.scratch/body.json '{document: $d[0]}' > ./.scratch/patch.
|
|
|
151
145
|
mb document update <id> --file ./.scratch/patch.json --profile <name> --json
|
|
152
146
|
```
|
|
153
147
|
|
|
154
|
-
|
|
148
|
+
To rename without touching the body, patch only `name`: `mb document update <id> --body '{"name":"New title"}'`.
|
|
155
149
|
|
|
156
150
|
## Don't
|
|
157
151
|
|
|
158
|
-
- Don't omit `_id` on an id-bearing node (`paragraph`, `heading`, `codeBlock`, `orderedList`, `bulletList`, `blockquote`, `cardEmbed`, `supportingText`) — `create`/`update` reject the body ("did not match expected schema"). Mint ids with `mb uuid`.
|
|
159
|
-
- Don't paste a whole `document get` response into `update` — `update` only accepts `name`, `document`, `collection_id`, `collection_position`, `archived`. Send the body under the `document` key, not the full record.
|
|
160
|
-
- Don't put the full `document` body in `list` expectations — `list` is compact and omits it by design; use `get --full`.
|
|
161
152
|
- Don't invent node types. Stick to the inventory above; unknown block types render as empty/broken in the editor even though the response schema is lenient.
|
|
162
|
-
- Don't author a card's `dataset_query` or `visualization_settings` from this skill alone — load `mbql` and `viz`.
|