@metabase/cli 0.1.10 → 0.1.11
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/marketplace.json +1 -1
- package/README.md +9 -6
- package/dist/{add-collection-DUqTrC5T.mjs → add-collection-DKULSPwT.mjs} +4 -4
- package/dist/add-collection-DyEe_xRO.mjs +10 -0
- package/dist/{archive-44EWiXud.mjs → archive-BJ6uL3hh.mjs} +3 -3
- package/dist/{archive-BEIyIsin.mjs → archive-BL_0fXhf.mjs} +3 -3
- package/dist/{archive-DtE2H4A6.mjs → archive-BLbaqc-5.mjs} +3 -3
- package/dist/{archive-_GMNY8wH.mjs → archive-BLxX3J6I.mjs} +3 -3
- package/dist/{archive-B59Y7ajB.mjs → archive-DQsV2uwy.mjs} +3 -3
- package/dist/{archive-0krZxAXq.mjs → archive-ibJVYsBC.mjs} +3 -3
- package/dist/{archive-BZfpjMir.mjs → archive-y-_m3XID.mjs} +3 -3
- package/dist/auth-BKaQjcmP.mjs +19 -0
- package/dist/{body-BdRyuvU4.mjs → body-DCQg78AF.mjs} +1 -1
- package/dist/{branches-Jpv-FNds.mjs → branches-CJVscsZe.mjs} +4 -4
- package/dist/{cancel-ChC4lFd4.mjs → cancel-C8W0sPix.mjs} +3 -3
- package/dist/{cancel-task-DyhIkNaL.mjs → cancel-task-dDgQDswT.mjs} +4 -4
- package/dist/card-CGBvUmCW.mjs +20 -0
- package/dist/{cards-O-nKkQKP.mjs → cards-B7HE8BT6.mjs} +3 -3
- package/dist/cli.mjs +23 -23
- package/dist/collection-Ch7pPKHD.mjs +20 -0
- package/dist/{collection-namespace-7724zUMx.mjs → collection-namespace-CBCApZJf.mjs} +1 -1
- package/dist/{create-DS52EhPd.mjs → create-84jcfY4N.mjs} +3 -3
- package/dist/{create-BuKx7kw6.mjs → create-B6mzc8UK.mjs} +3 -3
- package/dist/{create-B4f4Pldw.mjs → create-BQA6nM5n.mjs} +4 -4
- package/dist/{create-BxRsXQrm.mjs → create-BWaBYzhC.mjs} +3 -3
- package/dist/{create-CvYKJOcE.mjs → create-DN8us8si.mjs} +3 -3
- package/dist/{create-BbF9zVFP.mjs → create-DPASPFGi.mjs} +4 -4
- package/dist/{create-D45uFXlo.mjs → create-Di9NY7kK.mjs} +3 -3
- package/dist/{create-BnFHcnlL.mjs → create-VjvMdgM5.mjs} +3 -3
- package/dist/{create-branch-KOWUIE72.mjs → create-branch-CCoDX2Qf.mjs} +4 -4
- package/dist/{create-DeZ2x2Db.mjs → create-vNcD4jfq.mjs} +3 -3
- package/dist/{current-task-ClcWPPMc.mjs → current-task-QdxA3mSX.mjs} +4 -4
- package/dist/dashboard-BDb2krMr.mjs +21 -0
- package/dist/db-CoVdkdVm.mjs +22 -0
- package/dist/{delete-D68oS73R.mjs → delete-ByJ6guaH.mjs} +3 -3
- package/dist/{delete-CX2VUA5R.mjs → delete-CE_Jfjni.mjs} +3 -3
- package/dist/{delete-table-DLvL9mDA.mjs → delete-table-D5LDUkOk.mjs} +3 -3
- package/dist/{dirty-1OrXpc7E.mjs → dirty-C6yzkIPn.mjs} +4 -4
- package/dist/document-BA2S1ev3.mjs +19 -0
- package/dist/{eid-CzLhHZMW.mjs → eid-ZOyix96k.mjs} +3 -3
- package/dist/{error-BWXBhqLW.mjs → error-C1bUYXn7.mjs} +2 -1
- package/dist/{export-B5z8w-xo.mjs → export-BhbBX-pl.mjs} +6 -6
- package/dist/field-B5bptFqL.mjs +18 -0
- package/dist/{fields-sC7pzmPX.mjs → fields-CZ6IK8Qf.mjs} +3 -3
- package/dist/{get-ByZ4HR2T.mjs → get-157SWRte.mjs} +3 -3
- package/dist/{get-Bo1FGyFs.mjs → get-9b0ZEWh1.mjs} +3 -3
- package/dist/{get-Dl62Fy6Y.mjs → get-B03gdAun.mjs} +3 -3
- package/dist/{get-CHb6J908.mjs → get-BKWZA_22.mjs} +2 -2
- package/dist/{get-DTHLETau.mjs → get-Bbo2Y-aL.mjs} +3 -3
- package/dist/{get-ZesERdyk.mjs → get-BtqQ98nb.mjs} +2 -2
- package/dist/{get-reSMTfQi.mjs → get-C0emTO1r.mjs} +2 -2
- package/dist/{get-DLwb_gUh.mjs → get-CqKttHWW.mjs} +3 -3
- package/dist/{get-BYw3xS0X.mjs → get-D49EdHgH.mjs} +3 -3
- package/dist/{get-lcX52Skc.mjs → get-DIyv2a7h.mjs} +3 -3
- package/dist/{get-BC60bhel.mjs → get-DclzgC5G.mjs} +3 -3
- package/dist/{get-Fn9WkNhS.mjs → get-DlzC0xCg.mjs} +3 -3
- package/dist/{get-4GEDd9YN.mjs → get-mR01pCNG.mjs} +3 -3
- package/dist/{get-C6n86-dS.mjs → get-nl41R-2q.mjs} +3 -3
- package/dist/{get-run-CpCbHJad.mjs → get-run-_dlMCGp5.mjs} +3 -3
- package/dist/git-sync-DqJmONiB.mjs +28 -0
- package/dist/{has-remote-changes-CPz_-uxd.mjs → has-remote-changes-CToITRyl.mjs} +4 -4
- package/dist/{import-DaWgprK6.mjs → import-xnYHb4QX.mjs} +6 -6
- package/dist/{is-dirty-DZlI7lQx.mjs → is-dirty-BHyMmUVV.mjs} +3 -3
- package/dist/is-dirty-BvA3yzRc.mjs +9 -0
- package/dist/{items-hgbRYsYD.mjs → items-D3AmbIOO.mjs} +3 -3
- package/dist/{list-DO8T-nmF.mjs → list-5ZZfaOzm.mjs} +2 -2
- package/dist/{list-DBlsRSpZ.mjs → list-BHCVZh7Q.mjs} +2 -2
- package/dist/{list-CCZnH2-Z.mjs → list-BdStkg8G.mjs} +2 -2
- package/dist/{list-CGdOC9zX.mjs → list-BhZYkbc9.mjs} +2 -2
- package/dist/{list-kath_2cX.mjs → list-BkzrFyJO.mjs} +2 -2
- package/dist/{list-Cd2nOCAx.mjs → list-BtTRJyrR.mjs} +2 -2
- package/dist/{list-DOVX3vCb.mjs → list-Byn-dRW4.mjs} +3 -3
- package/dist/{list-BgHESP7b.mjs → list-C5NywUBH.mjs} +2 -2
- package/dist/{list-B9J3ujwn.mjs → list-CA6MlByL.mjs} +2 -2
- package/dist/{list-BWv5y307.mjs → list-CIu1zfpU.mjs} +2 -2
- package/dist/{list-BGerkRHH.mjs → list-C_d7ppOJ.mjs} +3 -3
- package/dist/{list-CKeVS1IZ.mjs → list-D7Np4fwk.mjs} +2 -2
- package/dist/{list-D9EuxFHO.mjs → list-DKPhFATP.mjs} +2 -2
- package/dist/{list-36H-dvJZ.mjs → list-DSJwMHXB.mjs} +2 -2
- package/dist/{login-BzrAGJfu.mjs → login-jWvDDvEm.mjs} +3 -3
- package/dist/{logout-CKBiltoS.mjs → logout-Bltlsm6u.mjs} +2 -2
- package/dist/measure-DBX4k13t.mjs +19 -0
- package/dist/{metadata-BjOrKtnv.mjs → metadata-BKcyjsoA.mjs} +3 -3
- package/dist/{metadata-BcGcUEVJ.mjs → metadata-Bf2TfG8F.mjs} +3 -3
- package/dist/{parse-id--iVTCKSo.mjs → parse-id-CWRJlQVK.mjs} +1 -1
- package/dist/{path-LgGU6Bd0.mjs → path-Djw9IARu.mjs} +2 -2
- package/dist/{poll-BucRFJT-.mjs → poll-hJrV0cPm.mjs} +1 -1
- package/dist/{poll-task-DTzKB3T3.mjs → poll-task-DK6X9Juu.mjs} +2 -2
- package/dist/{preflight-CzqVX0PP.mjs → preflight-DjSPp3Ct.mjs} +1 -1
- package/dist/{query-BVZkK6Qk.mjs → query-Cc-9oPsW.mjs} +3 -3
- package/dist/{query-DG_jygDF.mjs → query-Do5n5SZc.mjs} +3 -3
- package/dist/{remove-collection-HdeAfLyi.mjs → remove-collection-DVvnmU06.mjs} +6 -6
- package/dist/{rescan-values-hWubCruZ.mjs → rescan-values-CHluR3JC.mjs} +3 -3
- package/dist/{run-C1-lDmQF.mjs → run-CcJbcaw8.mjs} +5 -5
- package/dist/{runs-Q6DYQyqj.mjs → runs-Dc8qZBhl.mjs} +3 -3
- package/dist/{runtime-colqvhLf.mjs → runtime-jWkzCPlL.mjs} +1 -1
- package/dist/{schema-tables-BEastV_8.mjs → schema-tables-CZ6EMHUZ.mjs} +3 -3
- package/dist/{schemas-CgawwI_k.mjs → schemas-Br1cpOZc.mjs} +3 -3
- package/dist/{search-BWo7xSPP.mjs → search-DB2joPgt.mjs} +3 -3
- package/dist/segment-BWcyA1P-.mjs +19 -0
- package/dist/{set-7Nm2ZTb_.mjs → set-BRhVzEWd.mjs} +3 -3
- package/dist/{setting-DSGXJehQ.mjs → setting-CuxO4T1n.mjs} +3 -3
- package/dist/{setup-aJLGLrIT.mjs → setup-Bn4jwxdl.mjs} +3 -3
- package/dist/{skills-Q2AFsYvc.mjs → skills-Db1r65S5.mjs} +3 -3
- package/dist/snippet-C9UM6NwC.mjs +19 -0
- package/dist/{stash-DPQ0c-Cd.mjs → stash-Dp2JFgmp.mjs} +6 -6
- package/dist/{status-CvAATvV0.mjs → status-BxGBwALE.mjs} +2 -2
- package/dist/{status-CvKPrV5X.mjs → status-j09luNai.mjs} +5 -5
- package/dist/{summary-CeOnoOq2.mjs → summary-DVrj7DSa.mjs} +3 -3
- package/dist/{sync-schema-aOPBc3CY.mjs → sync-schema-BVHvjzGX.mjs} +5 -5
- package/dist/table-DYI8K12W.mjs +19 -0
- package/dist/transform-CvbMbJtI.mjs +24 -0
- package/dist/transform-job-CtzgvccZ.mjs +19 -0
- package/dist/{tree-MOQOBeAP.mjs → tree-BEPw6xbY.mjs} +2 -2
- package/dist/{update-BWyCK8QV.mjs → update-B5VBJFU9.mjs} +5 -5
- package/dist/{update-Bj9s0ri8.mjs → update-BXaT0Faw.mjs} +4 -4
- package/dist/{update-CitS-QRN.mjs → update-Bd0ASKUG.mjs} +5 -5
- package/dist/{update-BD9xkglP.mjs → update-Bn5wUyTQ.mjs} +4 -4
- package/dist/{update-BRrnfG0q.mjs → update-C5vlBTII.mjs} +4 -4
- package/dist/{update-Dri4Zg2H.mjs → update-C875_WGf.mjs} +4 -4
- package/dist/{update-DfNKr_vS.mjs → update-CAoOJRwz.mjs} +4 -4
- package/dist/{update-57uxZWcR.mjs → update-DiG6SSGm.mjs} +4 -4
- package/dist/{update-BkMWBzvk.mjs → update-DzKPxRBE.mjs} +4 -4
- package/dist/{update-dashcard-D_-ura3Y.mjs → update-dashcard-7bZtw04E.mjs} +4 -4
- package/dist/{update-B6mg3AZD.mjs → update-fVqLtZ_t.mjs} +4 -4
- package/dist/{upgrade-CFkZ4USY.mjs → upgrade-DdCxE9Ml.mjs} +2 -2
- package/dist/{uuid-DpinhSxA.mjs → uuid-DywghRXO.mjs} +2 -2
- package/dist/{values-BSS4DRxk.mjs → values-DSul3siC.mjs} +3 -3
- package/dist/{verify-B_A7v8TY.mjs → verify-DzDyKXOa.mjs} +1 -1
- package/dist/{wait-D3iSnjMM.mjs → wait-CacFZI5Y.mjs} +5 -5
- package/dist/{wait-flags-_LnHOeBA.mjs → wait-flags-DV419apK.mjs} +2 -2
- package/package.json +2 -1
- package/skill-data/core/SKILL.md +22 -23
- package/skill-data/data-analysis/SKILL.md +65 -0
- package/skill-data/data-transformation/SKILL.md +200 -0
- package/skill-data/document/SKILL.md +4 -4
- package/skill-data/mbql/SKILL.md +20 -20
- package/skill-data/robot-data-engineer/SKILL.md +142 -0
- package/skill-data/semantic-layer/SKILL.md +166 -0
- package/skill-data/transform/SKILL.md +46 -48
- package/skill-data/visualization/SKILL.md +5 -3
- package/skills/metabase-cli/SKILL.md +6 -0
- package/dist/add-collection-D9wXgmRj.mjs +0 -10
- package/dist/auth-cFC5m69m.mjs +0 -19
- package/dist/card-ClvGX6dQ.mjs +0 -20
- package/dist/collection-DjvowSJC.mjs +0 -20
- package/dist/dashboard-BFeURTOw.mjs +0 -21
- package/dist/db-CSH1kwQr.mjs +0 -22
- package/dist/document-KdT_Xj6r.mjs +0 -19
- package/dist/field-CTFnZI8G.mjs +0 -18
- package/dist/git-sync-C2vib8rx.mjs +0 -28
- package/dist/is-dirty-Bb0Rtj7x.mjs +0 -9
- package/dist/measure-Dw1QpRZa.mjs +0 -19
- package/dist/segment-DMuYvFjg.mjs +0 -19
- package/dist/snippet-Df2TrP7-.mjs +0 -19
- package/dist/table-pK4OkVtL.mjs +0 -19
- package/dist/transform-D60veFH8.mjs +0 -24
- package/dist/transform-job-DXt5LsrY.mjs +0 -19
|
@@ -145,10 +145,10 @@ Each entry in `cards` needs at least `{name, dataset_query, display, visualizati
|
|
|
145
145
|
`update` replaces the whole `document` body, so the safe loop is **read → edit → write**. A fetched body already carries `_id`s on its id-bearing nodes, so preserve them — only mint new ones for id-bearing nodes you add:
|
|
146
146
|
|
|
147
147
|
```bash
|
|
148
|
-
mb document get <id> --full --profile <name> --json | jq '.document' > /
|
|
149
|
-
# edit /
|
|
150
|
-
jq -n --slurpfile d /
|
|
151
|
-
mb document update <id> --file /
|
|
148
|
+
mb document get <id> --full --profile <name> --json | jq '.document' > ./.scratch/body.json
|
|
149
|
+
# edit ./.scratch/body.json (add nodes — give each new id-bearing node a fresh `mb uuid` _id) …
|
|
150
|
+
jq -n --slurpfile d ./.scratch/body.json '{document: $d[0]}' > ./.scratch/patch.json
|
|
151
|
+
mb document update <id> --file ./.scratch/patch.json --profile <name> --json
|
|
152
152
|
```
|
|
153
153
|
|
|
154
154
|
Don't hand-merge a partial node tree into a live document — pull the current `document`, mutate the array, and PUT the whole thing back. To rename without touching the body, patch only `name`: `mb document update <id> --body '{"name":"New title"}'`.
|
package/skill-data/mbql/SKILL.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: mbql
|
|
3
|
-
description: Author Metabase MBQL 5 query bodies for the `mb` CLI
|
|
3
|
+
description: Author Metabase MBQL 5 query bodies for the `mb` CLI - the only hand-authorable query format. Covers the JSON shape (lib/type mbql/query, flat numeric-id stages), the options-object-always-second clause rule, when lib/uuid is needed (optional - only to reference a clause), the print-schema/dry-run/run loop, where MBQL 5 is consumed (mb query, card dataset_query, transform source.query, measure/segment definition), the flat-vs-legacy-envelope footgun, joins and FK traversal, multi-stage pipelines, naming aggregation columns. Load when building or fixing an MBQL query by hand - "write an MBQL query", "create a card from MBQL", "the dataset_query is wrong", "fix the validation errors", "aggregate and group by", "join two tables", "month-over-month", or any `--dry-run` / `mb query` work.
|
|
4
4
|
allowed-tools: Read, Write, Edit, Bash, AskUserQuestion
|
|
5
5
|
---
|
|
6
6
|
|
|
@@ -8,13 +8,13 @@ allowed-tools: Read, Write, Edit, Bash, AskUserQuestion
|
|
|
8
8
|
|
|
9
9
|
MBQL 5 is the **only query format you can author by hand** with confidence — it has a bundled JSON Schema, so the CLI pre-flight-validates it before sending. Legacy MBQL 4 and native SQL are accepted but **not** schema-validated (see "Other formats" below).
|
|
10
10
|
|
|
11
|
-
Prefer MBQL over native SQL:
|
|
11
|
+
Prefer MBQL over native SQL: portable across warehouse engines and pre-flight-validated. Try it first; fall back to native SQL when MBQL can't express what you need, or when an MBQL body keeps failing server-side and you can't resolve it.
|
|
12
12
|
|
|
13
|
-
|
|
13
|
+
General flag conventions, body-input precedence, and output flags live in the `core` skill (`mb skills get core`).
|
|
14
14
|
|
|
15
15
|
## The shape
|
|
16
16
|
|
|
17
|
-
A
|
|
17
|
+
A flat object — `lib/type`, a numeric `database` id, and an ordered `stages` array. No recursive `source-query` nesting; multi-step queries are sibling stages.
|
|
18
18
|
|
|
19
19
|
```json
|
|
20
20
|
{
|
|
@@ -31,7 +31,7 @@ A query is a flat object — `lib/type`, a numeric `database` id, and an ordered
|
|
|
31
31
|
}
|
|
32
32
|
```
|
|
33
33
|
|
|
34
|
-
- **Numeric ids only.** `database`, `source-table`, and field ids are integers from `mb database list` / `mb table get <id> --include fields`. (
|
|
34
|
+
- **Numeric ids only.** `database`, `source-table`, and field ids are integers from `mb database list` / `mb table get <id> --include fields`. (Git-sync YAML uses _names_ like `[Sample Database, PUBLIC, ORDERS]`; the `/api/dataset` form uses numeric ids — don't mix them.)
|
|
35
35
|
- **First stage** carries `source-table` (a table id) or `source-card` (a saved card). Later stages omit both and read the previous stage's output columns by name.
|
|
36
36
|
- `source-card` references a saved card by its **numeric id** (from `mb card list`), not its string entity id; downstream fields are referenced by column name (string), not a field id.
|
|
37
37
|
|
|
@@ -53,9 +53,9 @@ The same `[op, {options}, …]` rule holds for `aggregation`, `breakout` (a list
|
|
|
53
53
|
|
|
54
54
|
## UUIDs: optional — mint only to reference a clause
|
|
55
55
|
|
|
56
|
-
`lib/uuid` is **optional — leave it out whenever you can.** Omit it and the server generates a unique one for every clause
|
|
56
|
+
`lib/uuid` is **optional — leave it out whenever you can.** Omit it and the server generates a unique one for every clause; an empty options object `{}` is the normal case. The more UUIDs you hand-manage the easier it is to trip the server's "all `lib/uuid`s must be unique" check — a duplicated UUID passes pre-flight, then fails server-side.
|
|
57
57
|
|
|
58
|
-
Set an explicit `lib/uuid` only when you must **reference a clause from elsewhere in the query** —
|
|
58
|
+
Set an explicit `lib/uuid` only when you must **reference a clause from elsewhere in the query** — you have to know the value to point at. The case that needs it: **ordering by (or otherwise reusing) an aggregation.** `["aggregation", {…}, "<uuid>"]`'s third arg is the **string** `lib/uuid` of the target aggregation, so give that aggregation an explicit `lib/uuid` and point the ref at the same string. A numeric position fails with `must be the target aggregation's lib/uuid (string), not a numeric position`.
|
|
59
59
|
|
|
60
60
|
```json
|
|
61
61
|
"aggregation": [["count", { "lib/uuid": "AGG_UUID" }]],
|
|
@@ -64,7 +64,7 @@ Set an explicit `lib/uuid` only when you must **reference a clause from elsewher
|
|
|
64
64
|
|
|
65
65
|
(`AGG_UUID` is both the aggregation's own `lib/uuid` and the string the ref points at — one value, by string equality. Every other clause omits its UUID. Expression refs work the same way but key off the expression's `lib/expression-name` string, so expressions rarely need an explicit `lib/uuid`.)
|
|
66
66
|
|
|
67
|
-
|
|
67
|
+
When you do need one, **always mint it with `mb uuid` — never write, guess, or copy a UUID yourself.** A hand-authored value is rejected pre-flight as not-a-v4 (`"a1"`, `"uuid-1"`, `"agg-uuid-001"` → `must be a UUID v4 (RFC 4122) — run \`mb uuid\``), or if it looks valid risks colliding with another clause. Only `mb uuid`gives genuine, unique v4s — mint just the few you reference (also covers native template-tag ids and any other`format: "uuid"` slot):
|
|
68
68
|
|
|
69
69
|
```bash
|
|
70
70
|
mb uuid --count 2 --json # mint only the clauses you actually reference
|
|
@@ -75,7 +75,7 @@ mb uuid --count 2 --json # mint only the clauses you actually reference
|
|
|
75
75
|
`mb query` is the canonical authoring surface. Three modes:
|
|
76
76
|
|
|
77
77
|
```bash
|
|
78
|
-
mb query --print-schema --profile <n> > /
|
|
78
|
+
mb query --print-schema --profile <n> > ./.scratch/mbql-schema.json # 1. fetch the schema
|
|
79
79
|
mb query --file q.json --dry-run --profile <n> # 2. validate, no network
|
|
80
80
|
mb query --file q.json --profile <n> --json # 3. validate + run
|
|
81
81
|
```
|
|
@@ -86,7 +86,7 @@ mb query --file q.json --profile <n> --json # 3. validate +
|
|
|
86
86
|
|
|
87
87
|
`path` is a JSON Pointer into the body (`/stages/0/aggregation/0`); `message` is the validator error. Exit codes: `0` valid + ran, `2` validation failed / malformed body, `1` server-side error after a valid pre-flight.
|
|
88
88
|
|
|
89
|
-
**Pre-flight is a lightweight shape check, not the full backend validator.** It checks JSON shape, `lib/uuid` format, and enum values — not operator names, the first-stage source rule, or whether a reference resolves. A clean `--dry-run` is necessary but not sufficient: a body can pass pre-flight and still fail on the server (exit `1`). The
|
|
89
|
+
**Pre-flight is a lightweight shape check, not the full backend validator.** It checks JSON shape, `lib/uuid` format, and enum values — not operator names, the first-stage source rule, or whether a reference resolves. A clean `--dry-run` is necessary but not sufficient: a body can pass pre-flight and still fail on the server (exit `1`). The server is the authority — when a run fails, read its error and fix the body. Common ones:
|
|
90
90
|
|
|
91
91
|
- `not a known MBQL clause` → a misspelled or unsupported **operator**. Check the vocabulary in `operators.md` (`mb skills get mbql --full`).
|
|
92
92
|
- `Initial MBQL stage must have either :source-table or :source-card` → the **first stage** is missing its source (a numeric table or card id); only the first stage takes one, later stages read the previous stage's columns.
|
|
@@ -95,11 +95,11 @@ mb query --file q.json --profile <n> --json # 3. validate +
|
|
|
95
95
|
|
|
96
96
|
A successful run emits the compact envelope by default: `data.rows` + slim `data.cols` (`name`, `display_name`, `base_type`, `semantic_type`). Pass `--full` for the raw `/api/dataset` envelope (`results_metadata`, `native_form`, per-column fingerprints/`field_ref`) only when you need that metadata; `--fields data.rows` narrows to rows alone. `mb query` also runs a **native** body — `{database, type:"native", native:{query:"SELECT …"}}` — which skips pre-flight; the quickest way to eyeball warehouse data.
|
|
97
97
|
|
|
98
|
-
`--skip-validate` bypasses
|
|
98
|
+
`--skip-validate` bypasses pre-flight and sends as-is — use only when the bundled schema disagrees with what the server actually accepts (drift / false negative). Mutually exclusive with `--dry-run`. Same flag exists on `card create/update` and `transform create/update`.
|
|
99
99
|
|
|
100
100
|
## Where MBQL 5 is consumed
|
|
101
101
|
|
|
102
|
-
The same body and
|
|
102
|
+
The same body and pre-flight apply everywhere a query is embedded. Each pre-flights only when the value is MBQL 5 (`lib/type: "mbql/query"`); legacy shapes skip it; `--skip-validate` bypasses.
|
|
103
103
|
|
|
104
104
|
| Command | MBQL 5 lives at | Notes |
|
|
105
105
|
| --------------------------------------- | ---------------------------------------------- | ------------------------------------------- |
|
|
@@ -122,16 +122,16 @@ The most common mistake. The legacy MBQL 4 shape `{ "type": "query", "database":
|
|
|
122
122
|
}
|
|
123
123
|
```
|
|
124
124
|
|
|
125
|
-
No `type:"query"` wrapper, no `query:` nesting. If you wrap MBQL 5 inside a legacy envelope the CLI rejects it pre-send with a `ConfigError` (no `--skip-validate` gets it past). If it
|
|
125
|
+
No `type:"query"` wrapper, no `query:` nesting. If you wrap MBQL 5 inside a legacy envelope the CLI rejects it pre-send with a `ConfigError` (no `--skip-validate` gets it past). If it reached the server it would store silently and fail at run time with `Initial MBQL stage must have either :source-table or :source-card`.
|
|
126
126
|
|
|
127
127
|
## Other formats skip pre-flight
|
|
128
128
|
|
|
129
|
-
Anything
|
|
129
|
+
Anything not `lib/type: "mbql/query"` is sent as-is and normalized server-side:
|
|
130
130
|
|
|
131
131
|
- **Legacy MBQL 4** — `{ "type": "query", "database": N, "query": { "source-table": T, … } }`
|
|
132
132
|
- **Native SQL** — `{ "type": "native", "database": N, "native": { "query": "SELECT …" } }`
|
|
133
133
|
|
|
134
|
-
`mb query --file probe.json` runs these directly; `--dry-run` on them returns `{ ok: true, errors: [] }`. Don't author MBQL 4 by hand —
|
|
134
|
+
`mb query --file probe.json` runs these directly; `--dry-run` on them returns `{ ok: true, errors: [] }`. Don't author MBQL 4 by hand — build a legacy or complex query in the Metabase UI and pull the body with `mb card get <id> --full --json` / `mb transform get <id> --full --json`.
|
|
135
135
|
|
|
136
136
|
## Joins and FK traversal
|
|
137
137
|
|
|
@@ -154,7 +154,7 @@ Two ways to read columns from a related table.
|
|
|
154
154
|
"breakout": [["field", { "join-alias": "Customers" }, 1682]]
|
|
155
155
|
```
|
|
156
156
|
|
|
157
|
-
|
|
157
|
+
Left ref is a column of the stage's own source (`1711` = orders.customer_id); the right ref carries `join-alias` and points at the joined table's key (`1684` = customers.id). Every later reference to a joined column (`1682` = customers.plan) needs that same `join-alias`. Stack multiple objects in `joins`, each with its own `alias`.
|
|
158
158
|
|
|
159
159
|
**Implicit FK join via `source-field`.** For a single-hop FK lookup, skip the join — put the FK column's id in the target field's `source-field` option and Metabase traverses the relationship:
|
|
160
160
|
|
|
@@ -166,7 +166,7 @@ The condition's left ref is a column of the stage's own source (`1711` = orders.
|
|
|
166
166
|
|
|
167
167
|
## Multi-stage pipelines
|
|
168
168
|
|
|
169
|
-
Stages run in order; each reads the **previous stage's output columns** — the breakouts and aggregations it produced — referenced by **string name + `base-type`**, not a numeric field id. Only the first stage takes a `source-table`/`source-card`.
|
|
169
|
+
Stages run in order; each reads the **previous stage's output columns** — the breakouts and aggregations it produced — referenced by **string name + `base-type`**, not a numeric field id. Only the first stage takes a `source-table`/`source-card`. Add a stage to operate on an aggregate (you can't filter or order by an aggregation within the stage that computes it): aggregate, then filter the aggregate, then order + limit.
|
|
170
170
|
|
|
171
171
|
```json
|
|
172
172
|
"stages": [
|
|
@@ -197,7 +197,7 @@ Later stages address the first stage's aggregation by the `name` you gave it (`"
|
|
|
197
197
|
|
|
198
198
|
## Naming aggregation output columns
|
|
199
199
|
|
|
200
|
-
Default MBQL 5 aggregations materialize as `count`, `count_where`, `avg`, `avg_2`, `sum`, … — fine for an ad-hoc run, ugly
|
|
200
|
+
Default MBQL 5 aggregations materialize as `count`, `count_where`, `avg`, `avg_2`, `sum`, … — fine for an ad-hoc run, ugly for a transform target table or card column. Set `name` (the warehouse column name) and `display-name` (the UI header) in the aggregation's options:
|
|
201
201
|
|
|
202
202
|
```json
|
|
203
203
|
["count", { "name": "shipments_shipped", "display-name": "Shipments shipped" }]
|
|
@@ -205,7 +205,7 @@ Default MBQL 5 aggregations materialize as `count`, `count_where`, `avg`, `avg_2
|
|
|
205
205
|
|
|
206
206
|
## Operator reference
|
|
207
207
|
|
|
208
|
-
The full operator vocabulary — filter operators (`=`, `!=`, `<`, `between`, `contains`, `is-null`, …), aggregation functions (`count`, `sum`, `avg`, `distinct`, `count-where`, `share`, …), expression operators (arithmetic, string, temporal), temporal-bucketing units, and binning strategies — lives in this skill's `references/operators.md`, in
|
|
208
|
+
The full operator vocabulary — filter operators (`=`, `!=`, `<`, `between`, `contains`, `is-null`, …), aggregation functions (`count`, `sum`, `avg`, `distinct`, `count-where`, `share`, …), expression operators (arithmetic, string, temporal), temporal-bucketing units, and binning strategies — lives in this skill's `references/operators.md`, in numeric-id form. Load it on demand rather than dumping the schema:
|
|
209
209
|
|
|
210
210
|
```bash
|
|
211
211
|
mb skills get mbql --full # appends references/operators.md to this body
|
|
@@ -217,7 +217,7 @@ mb skills path mbql # → the skill dir; then Read references/operator
|
|
|
217
217
|
## Don't
|
|
218
218
|
|
|
219
219
|
- Don't mint a `lib/uuid` for every clause — they're optional; omit them and the server fills them in. Mint (with `mb uuid`) only the clause you need to reference; never invent, hard-code, or copy a UUID (duplicates are rejected server-side).
|
|
220
|
-
-
|
|
220
|
+
- Keep the options object in slot 1 of every clause — `[op, {options}, ...args]`, id last (`["field", {}, 1779]`). The legacy `["field", id, opts]` order (id second) is rejected pre-flight.
|
|
221
221
|
- Don't wrap an MBQL 5 body in `{type:"query", query:…}` — `dataset_query` / `source.query` / `definition` is the flat `mbql/query`.
|
|
222
222
|
- Don't author MBQL 4 by hand — build it in the UI and pull it with `… get <id> --full --json`.
|
|
223
223
|
- Don't skip the `--dry-run` loop on a non-trivial query — it's free and exact.
|
|
@@ -0,0 +1,142 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: robot-data-engineer
|
|
3
|
+
description: The front door for turning a database into something a non-technical person can use - clean tables, reusable definitions, dashboards, and answers - all through the `mb` CLI. A light router - it works out where the user is (raw data? clean tables? ready to chart? just need a question answered?), sets up auth and how hands-on they want to be, then loads the right specialized skill. Load when someone wants to "make sense of my data", "build a data model", "go from raw data to a dashboard", "answer questions about my data", "report on who registered / signed up / responded", "analyze X", "be my data analyst / data engineer", "set up analytics for X", or asks for the whole journey rather than one step.
|
|
4
|
+
allowed-tools: Read, Write, Edit, Bash, AskUserQuestion
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Robot Data Engineer
|
|
8
|
+
|
|
9
|
+
You're the front door, not the worker. Point the user at the right tools and get out of the way. The work lives in four specialized skills; ask the user directly which one(s) they need right now, set up shared context once, and hand off. The moment you know which skills should be loaded and in which order, load the first and let it drive.
|
|
10
|
+
|
|
11
|
+
The three stages:
|
|
12
|
+
|
|
13
|
+
1. **Raw data → clean tables** — `data-transformation`. Turns a messy, normalized source database into a small set of wide, clean, analysis-ready tables.
|
|
14
|
+
2. **Clean tables → reusable definitions** — `semantic-layer`. Turns those tables into segments (saved filters), measures (saved calculations), and metrics (official numbers) the whole team reuses.
|
|
15
|
+
3. **Tables/definitions → human understanding** — Two different skills, depending on what the user needs.
|
|
16
|
+
A. Charts and dashboards? `visualization`. Builds the questions and dashboards people look at.
|
|
17
|
+
B. Plain-language analysis? `data-analysis`. Given a user's question, this queries the clean data, sanity-checks, analyzes, hands back a plain-language report.
|
|
18
|
+
|
|
19
|
+
Stages 3A and 3B are not sequential, but options: answering-in-prose and charting are two different things you can do with clean data; route to whichever the goal calls for. Users describe a goal, not a stage. Map the goal to a stage, confirm, and route.
|
|
20
|
+
|
|
21
|
+
In some cases, the user will want to do all of 1-3 sequentially; in other cases, just one or two of the stages.
|
|
22
|
+
|
|
23
|
+
---
|
|
24
|
+
|
|
25
|
+
## Setup — do this once, up front
|
|
26
|
+
|
|
27
|
+
Settle two things before routing so the child skills don't re-ask:
|
|
28
|
+
|
|
29
|
+
1. **Auth.** Pick the profile per `core`'s **Auth & profiles** section — `mb auth list --json`; one → use it, several → ask which, none → ask the user to `mb auth login` — then carry `--profile <name>` into everything. (Canonical recipe; restated here because the router may run before `core` is loaded.)
|
|
30
|
+
|
|
31
|
+
2. **How hands-on they want to be** (the autonomy slider). Ask once, plainly, remember it for the whole session, and tell the child skill the chosen mode so they aren't asked again:
|
|
32
|
+
|
|
33
|
+
> Quick thing — how hands-on do you want to be?
|
|
34
|
+
> • **Check with me on everything** — I'll run each step past you first.
|
|
35
|
+
> • **Balanced** (default) — I'll decide the obvious stuff and ask only when it matters.
|
|
36
|
+
> • **Just go** — I'll do what makes sense and show you the result.
|
|
37
|
+
|
|
38
|
+
Two things you always own, regardless of mode and regardless of which child ran:
|
|
39
|
+
|
|
40
|
+
- **When genuinely unsure, ask — never assume.** Pass this expectation down.
|
|
41
|
+
- **The final hard stop.** Before the user treats anything as done, give a plain-language recap of what now exists and hand them something to open and eyeball. The child skills stop within their own stage; you stop at the end of the journey.
|
|
42
|
+
|
|
43
|
+
---
|
|
44
|
+
|
|
45
|
+
## Shared Contract
|
|
46
|
+
|
|
47
|
+
This is the single source for the rules every child skill follows. Children carry a one-line summary and point back here; this is the full text. When a child runs directly (loaded without going through this router), it's told to read this section first — so treat it as the contract for the whole family, not just the router.
|
|
48
|
+
|
|
49
|
+
**Who you're talking to.** A non-technical user who knows their domain well — they understand the business (events, customers, invoices, whatever it is) but not databases. Talk in their terms.
|
|
50
|
+
|
|
51
|
+
**Jargon.** Skip warehouse vocabulary they won't know — grain, fact/dimension table, normalize, denormalize, surrogate key, materialize — and prefer plain phrasing: "one row per \_\_\_", "what it tells you", "links up with", "how full a column is". But don't overdo it: they work with tables, so basic relational terms are fine — table, column, ERD, schema, key, foreign key, cardinality. **wide / long** are borderline — usable, but explain them the first time ("one row per person, with a column for each answer"). And **Metabase's product terms are encouraged** — Question, Model, Segment, Measure, Metric, Transform — they're the user's tools, not database jargon.
|
|
52
|
+
|
|
53
|
+
**PII.** Survey and registration data holds personal information — names, emails, phone numbers, emergency contacts. Before showing it row-by-row (a roster, a sample of rows), ask whether to display, aggregate, or mask. Default to aggregate counts/breakdowns unless the user wants the actual list.
|
|
54
|
+
|
|
55
|
+
**Capability limits — know what you can't do.** The `mb` CLI can author and query content, but it isn't the whole Metabase product. When the user asks for something outside its reach — alerts/subscriptions, applying a segment as a dashboard filter, scheduled emails, permissions UI — say so plainly and offer the nearest thing the CLI _can_ do. Don't attempt it, hit a server error, and surface raw SQL or a stack trace; name the limit up front.
|
|
56
|
+
|
|
57
|
+
**Permission denied — stop, diagnose, offer a way back.** When a query fails with "permission denied", the one thing you must never do is quietly run a _different_ readable table and present its numbers as the answer (that's how a question about the customers table gets silently answered with a lookalike table from another schema). Instead, in order:
|
|
58
|
+
|
|
59
|
+
1. **Stop.** Don't substitute another table and pass it off as the answer.
|
|
60
|
+
2. **Surface and diagnose in plain, friendly terms.** Name what was denied and the likely reason. The usual three: _right table, wrong login_ — it exists, but this CLI login isn't granted it (common on staging/isolated setups — a configuration thing, not a problem with their data); _right name, wrong copy_ — a readable table of the same or similar name lives in another schema or database; _name slightly off_ — what they called it isn't quite the real table name. For example: "I can't read `analytics.account` — this login doesn't have access to it. That's usually a staging-permissions thing, not a problem with your data."
|
|
61
|
+
3. **Offer to search — don't auto-crawl.** Ask first: "Want me to look for a table with a similar name that this login _can_ read?" Only on yes, run `mb search <name>` / `mb table list`, and surface any match as a **confirm question**, never as a substituted answer: "There's `dbt_models.account` I can read — did you mean that one?"
|
|
62
|
+
4. **Hand control back.** Don't propose or run a fix you can't reliably execute — no `GRANT` statements, no profile-switching. The recovery is the user's call.
|
|
63
|
+
|
|
64
|
+
**Scratch files.** Working files — transform/query/patch JSON bodies, notes — go in `./.scratch` in the current working directory, **never `/tmp`**. Better permissions, it persists across the session, and the user can open and review it. `mkdir -p ./.scratch` if it isn't there yet.
|
|
65
|
+
|
|
66
|
+
**Talking to the user.** Habits that are easy to slip on (see also "Questions must carry their own context" below):
|
|
67
|
+
|
|
68
|
+
- **Don't reference things they never saw.** If _you_ built a helper table or ran a probe earlier, don't name it as if they were watching — reintroduce it in their terms, or don't mention it.
|
|
69
|
+
- **Assume they read only the last ~30 lines.** Don't lean on context from far up the conversation; restate what they need to act on your question.
|
|
70
|
+
- **Plain permission requests.** Don't paste a wall of SQL or JSON and ask "run this?". Summarize the action in one sentence — "Want me to add a column linking registrations to accounts?" — and offer to show the details if they ask.
|
|
71
|
+
|
|
72
|
+
**Autonomy slider.** Ask once, up front (the router does this in Setup), then remember it for the whole session — children read the chosen mode, they don't re-ask:
|
|
73
|
+
|
|
74
|
+
> Quick thing — how hands-on do you want to be?
|
|
75
|
+
> • **Check with me on everything** — I'll run each step past you first.
|
|
76
|
+
> • **Balanced** (default) — I'll decide the obvious stuff and ask only when it matters.
|
|
77
|
+
> • **Just go** — I'll do what makes sense and show you the result.
|
|
78
|
+
|
|
79
|
+
**When genuinely unsure, ask — never assume.**
|
|
80
|
+
|
|
81
|
+
**Questions must carry their own context.** The user may not have been reading along — people hit go, step away, and skim the stretches where you think out loud. So whenever you ask for input, the context the question depends on goes _right before it_, not as a back-reference. "Given the mismatch I found earlier, what would you like to do?" forces a scroll-back; lead with a short recap instead:
|
|
82
|
+
|
|
83
|
+
> I have a question for you — quick recap so it makes sense:
|
|
84
|
+
>
|
|
85
|
+
> - I found a mismatch in ...
|
|
86
|
+
> - This matters because ...
|
|
87
|
+
> - Here's what I was thinking, but I need to check ...
|
|
88
|
+
>
|
|
89
|
+
> The question.
|
|
90
|
+
|
|
91
|
+
Recap only the few points the question turns on — enough to answer cold, not a replay of everything you did.
|
|
92
|
+
|
|
93
|
+
**The final hard stop.** Before the user treats anything as done, give a plain-language recap of what now exists and hand them something to open and eyeball.
|
|
94
|
+
|
|
95
|
+
---
|
|
96
|
+
|
|
97
|
+
## Work out where they are, then route
|
|
98
|
+
|
|
99
|
+
Don't make the user name a _stage_ — but do find out _where their data lives_ before you go looking for it.
|
|
100
|
+
|
|
101
|
+
**Ask before you crawl.** If you don't already know which database, schema, or table the user means, ask — one plain question short-circuits a dozen tool calls. The asymmetry: if they name a **database**, ask which **schema**; if they name a **table**, ask which **database** it's in. "If you don't know, no problem — I'll look" is the fallback, not the opening move. Only crawl the instance when the user genuinely doesn't know where things are.
|
|
102
|
+
|
|
103
|
+
**When you do crawl — the efficient ladder** (cheap, narrowest-first; never pull whole-warehouse rollups):
|
|
104
|
+
|
|
105
|
+
- Walk down: `mb db list` → `mb db schemas <id>` → `mb db schema-tables <id> <schema>` → `mb table list [--db-id]` → `mb table fields <id>` / `mb table metadata <id>`.
|
|
106
|
+
- Have a _name_ to look for rather than a tree to walk? Use `mb search <query> [--models] [--db-id]` instead of crawling.
|
|
107
|
+
- Need to know what's actually in a column? `mb field summary <id>` (row/distinct counts) and `mb field values <id>` (sample values).
|
|
108
|
+
- **If a database looks freshly connected, or a table the user expects isn't showing up, offer to sync** — `mb db sync-schema <id> --wait` — before concluding the table doesn't exist.
|
|
109
|
+
|
|
110
|
+
**Then read the shape to pick a stage.** Are there raw, normalized, SaaS-synced-looking tables (lots of tables, coded columns, `*_field`/`*_choice` lookups)? Or already wide, clean, human-readable ones? Any segments/measures/metrics (`mb segment list`, `mb measure list`, `mb card list`) or dashboards (`mb dashboard list`)?
|
|
111
|
+
|
|
112
|
+
**Map goal + state to a skill:**
|
|
113
|
+
|
|
114
|
+
| What the user wants / what's there | Load |
|
|
115
|
+
| -------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------- |
|
|
116
|
+
| "Clean up / flatten / make sense of" raw, normalized data; no clean tables yet | `data-transformation` |
|
|
117
|
+
| Clean tables exist; "make this reusable", "define active customers / revenue / MRR officially", "so everyone uses the same definition" | `semantic-layer` |
|
|
118
|
+
| Tables (and maybe definitions) exist; "chart this", "build a dashboard", "show me X over time" | `visualization` |
|
|
119
|
+
| Clean tables exist; "answer this question", "who registered", "what did people say", "analyze / report on / summarize X" (wants a written answer, not a chart) | `data-analysis` |
|
|
120
|
+
| "Do the whole thing" / "set up analytics for X" from raw data | start at `data-transformation`, then continue down the journey (see below) |
|
|
121
|
+
|
|
122
|
+
Load a skill with `mb skills get <name>`. Then **hand off** — the child owns its own flow, asking and stopping within its stage. Don't narrate the child's work or duplicate its steps.
|
|
123
|
+
|
|
124
|
+
**If the state and the goal disagree** — they ask for a dashboard but there are only raw tables — say so plainly and offer the earlier stage first: _"There aren't clean tables to chart yet — want me to build those first, then we'll chart them?"_ Don't silently build on raw data.
|
|
125
|
+
|
|
126
|
+
---
|
|
127
|
+
|
|
128
|
+
## The whole journey
|
|
129
|
+
|
|
130
|
+
For the full arc (raw → dashboard), run the stages in order, handing off to each child in turn. Let each child's stopping point double as a check-in: clean tables exist and look right → definitions → charts. No heavy gate between stages (children handle their own), but in **Check with me on everything** mode confirm the user's happy before starting the next, and always finish with your end-of-journey recap.
|
|
131
|
+
|
|
132
|
+
A user can drop in at any stage — that's the point of detecting state. Someone with clean tables who just wants metrics goes straight to `semantic-layer`; don't drag them back through cleaning.
|
|
133
|
+
|
|
134
|
+
---
|
|
135
|
+
|
|
136
|
+
## Don't
|
|
137
|
+
|
|
138
|
+
- **Hand the work to the child skill — don't do it yourself.** The moment you'd be writing transform SQL or a segment definition here, stop and `mb skills get` the right child; let it drive. You route and set up context; the child does the work.
|
|
139
|
+
- **Don't re-ask the autonomy question** once it's set; pass it down.
|
|
140
|
+
- **Don't skip the starting-state check** and assume raw data — a user with clean tables shouldn't be sent through cleaning.
|
|
141
|
+
- **Don't build on raw data when the goal needs clean tables** — route to the earlier stage first.
|
|
142
|
+
- **Don't drop the final recap** — you own the end-of-journey hard stop even though each child stops within its own stage.
|
|
@@ -0,0 +1,166 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: semantic-layer
|
|
3
|
+
description: Turn clean, analysis-ready tables into a shared vocabulary the org reuses - Metabase segments (saved filters, e.g. active customers), measures (saved calculations, e.g. net revenue), and metrics (official numbers, e.g. monthly recurring revenue) - so people stop reinventing the same definition five ways. Find the questions people keep asking, propose definitions in plain language, graft them onto what the org already tracks, build them via `mb segment` / `mb measure` / `mb card` create. For a non-technical user who knows their domain. Load when someone wants to "make this reusable", "define X officially", "standardize how we calculate Y", or "create a segment / measure / metric". Strategy skill for designing reusable definitions; for raw `mb segment` / `mb measure` mechanics, use `core`.
|
|
4
|
+
allowed-tools: Read, Write, Edit, Bash, AskUserQuestion
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Semantic Layer
|
|
8
|
+
|
|
9
|
+
> **Shared contract (read first).** This skill is part of the `robot-data-engineer` family and follows its shared rules: audience is a non-technical user, so no database jargon (skip "normalize"/"grain"; ERD/foreign key are fine; explain "wide"/"long" the first time you use them). Ask before showing PII row-by-row (names, emails, phones) — default to aggregates. When asked for something the CLI can't do (alerts, dashboard filters), name the limit instead of erroring into raw SQL. Honor the autonomy mode the user picked. Full text and the autonomy slider live in the router — run `mb skills get robot-data-engineer` and read its **Shared Contract** if you haven't.
|
|
10
|
+
|
|
11
|
+
Your job: take the clean, analysis-ready tables that already exist and turn the **questions people keep asking** into **shared, reusable definitions** — so "active customer", "net revenue", and "monthly recurring revenue" mean one thing across the whole organization, not five slightly-different things in five people's saved questions.
|
|
12
|
+
|
|
13
|
+
You build three kinds of reusable thing. These are real Metabase features with real names — **use the Metabase names** (segment, measure, metric) and teach them to the user as you go. They're product vocabulary, not jargon. Pair the name with a plain gloss the first time, then use it freely:
|
|
14
|
+
|
|
15
|
+
- **Segment** — a saved filter on a table. A reusable row-selector: "Active customers", "orders over $100", "EU shipments". People pick it from the **Filter** block in the query builder instead of re-typing the conditions. (Docs: <https://www.metabase.com/docs/latest/data-studio/segments>.)
|
|
16
|
+
- **Measure** — a saved aggregation on a table. A reusable calculation: "Net Promoter Score", "average order value". People pick it from the **Summarize** block instead of re-writing the formula. Only works on questions built directly on the measure's table. (Docs: <https://www.metabase.com/docs/latest/data-studio/measures>.)
|
|
17
|
+
- **Metric** — a reusable aggregation that lives in a **collection** (a folder), not bolted to a table. "Monthly recurring revenue", "weekly active users". It's the org's official definition of an important number, can be saved into the **Library**, and can carry a default time dimension for charting. (Docs: <https://www.metabase.com/docs/latest/data-modeling/metrics>.)
|
|
18
|
+
|
|
19
|
+
Introduce each like: _"I'll save this as a **segment** — that's Metabase's word for a reusable filter, so you can pull up active customers with one click anytime."_ After that, just say "segment".
|
|
20
|
+
|
|
21
|
+
This skill runs **after** the analysis-ready tables exist (build those with transforms — load `mb skills get transform`). Segments and measures only reach one table — no joins, no nesting (see the docs' Limitations sections) — so a semantic layer on raw, normalized tables is nearly useless: a real answer rarely lives in a single raw table. **Wide clean tables first, segments/measures/metrics second.**
|
|
22
|
+
|
|
23
|
+
You drive everything through the `mb` CLI. Load the CLI skills you'll need:
|
|
24
|
+
|
|
25
|
+
```bash
|
|
26
|
+
mb skills get core # auth, profiles, db/table/field inspection, query, search
|
|
27
|
+
mb skills get mbql # the definition bodies (filters and aggregations) are MBQL 5
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
Authentication is the user's job. Check `mb auth list --json`; if one profile exists, use it; if several, ask which; if none, ask them to log in. Pass `--profile <name>` to every command.
|
|
31
|
+
|
|
32
|
+
---
|
|
33
|
+
|
|
34
|
+
## Who you're talking to
|
|
35
|
+
|
|
36
|
+
A **non-technical user who knows their domain well.** They know the business — who an "active" customer is, what counts as "revenue" — but not databases. So:
|
|
37
|
+
|
|
38
|
+
- **Teach the words a curious non-engineer can follow; skip the deep-internals jargon.** Two sets are fine and worth teaching: Metabase product terms (**segment, measure, metric, collection, Library, the Filter / Summarize blocks**) and common data words a domain user can reasonably learn (**table, column, foreign key, schema, join, filter, row**) — gloss them once, then use them. Avoid **deep-internals jargon** that buys nothing for this user: grain, cardinality, normalize/denormalize, surrogate key, MBQL, `table_id`, materialize. Prefer the plain effect when it's clearer ("this number needs data from two tables" reads easier than "this needs a join across two fact tables") — but you don't have to contort around "foreign key" or "schema".
|
|
39
|
+
- **Talk about the question, then name the object.** Lead with what it does for them, then attach the term: _"I'll save 'big orders' as a segment so you can pull them up with one click."_ Not a bare "I'll create a segment on `table_id` 235."
|
|
40
|
+
- **Be a helpful colleague, not an engineer reporting status.** Elide the wiring (ids, query bodies, the CLI). Ask the one question that actually matters.
|
|
41
|
+
|
|
42
|
+
---
|
|
43
|
+
|
|
44
|
+
## Autonomy — honor the mode the user set
|
|
45
|
+
|
|
46
|
+
The user already picked an autonomy mode (the router's Shared Contract asks the slider once, up front — don't re-ask). Apply it to building definitions:
|
|
47
|
+
|
|
48
|
+
| Mode | What you do |
|
|
49
|
+
| ----------------------- | ----------------------------------------------------------------------------------------------------------- |
|
|
50
|
+
| **Check on everything** | Confirm every single definition (name + plain description) before building it. |
|
|
51
|
+
| **Balanced** (default) | Build the obvious ones; ask only on the judgment calls (the prudential list below) and anything ambiguous. |
|
|
52
|
+
| **Just go** | Build the whole set, surface judgment calls as "here's what I picked and why — say the word to change any." |
|
|
53
|
+
|
|
54
|
+
**Two things never bend, in any mode:**
|
|
55
|
+
|
|
56
|
+
1. **When you're genuinely unsure — ask. Never assume.** "Just go" means _decide the obvious_, not _guess on the unclear_. A wrong-but-confident definition of "active customer" is worse than a one-line question.
|
|
57
|
+
2. **The final gate is a hard stop (see Phase 3).** No mode auto-publishes. You always stop, recap in plain language, and hand the user something to eyeball before anything goes live.
|
|
58
|
+
|
|
59
|
+
---
|
|
60
|
+
|
|
61
|
+
## Two kinds of decisions
|
|
62
|
+
|
|
63
|
+
**Hard rules — absolutes, never ask:**
|
|
64
|
+
|
|
65
|
+
1. **Never invent what a word means — pin it to real data.** "Active customer" is not yours to define. Before you build a segment for it, find out (from the user, or from how the data actually behaves) what _they_ mean: ordered in the last 90 days? Has a live subscription? Logged in this month? Confirm against actual values, then build to that. A definition built on a guessed meaning is a silent lie everyone then trusts.
|
|
66
|
+
2. **Keep the language at the level set in "Who you're talking to."** Metabase terms and common data words (table, column, foreign key, schema, join) are fine and worth teaching; deep-internals jargon (grain, cardinality, surrogate key, `table_id`) is not.
|
|
67
|
+
3. **Don't bury filters inside measures.** A measure should aggregate _what it's given_; let the user combine it with a segment at question time, rather than welding a filter into the measure. Welded-in filters collide and confuse when someone applies their own filter on top — and the metrics doc explicitly recommends against it. (Use conditional forms like `SumIf`/`CountIf` for "sum only the paid ones" — that's part of the measure's formula, not a hidden row filter.)
|
|
68
|
+
4. **Respect where each thing can reach.** Segments and measures work **only** on a question built _directly_ on their own table — not through a join, not on a question-built-on-a-question (the Limitations sections of both docs say so). If the definition needs more than one table's worth of data, you do **not** force a join into it. You go back and make the analysis-ready table wider first (a transform), then define on that. Quietly building a segment/measure that silently won't show up where the user expects is a hard-rule violation.
|
|
69
|
+
5. **Don't strand a metric on a single data source.** A metric is data-source-bound the same way — defined on table X, it appears only on questions built on table X, not on anything derived from it. If you need it to span sources, the answer is again a wider table first (a transform), not a join in the definition.
|
|
70
|
+
6. **Every definition keeps a clear, plain name and a one-line description in the user's words.** The name is what they'll see in a menu six weeks from now with no memory of this conversation. "Active customers (ordered in last 90 days)" beats "active_seg_v2".
|
|
71
|
+
|
|
72
|
+
**Prudential calls — genuinely contextual, state your lean, let the user decide** (skip the ask in "Just go" mode — pick your lean, flag it):
|
|
73
|
+
|
|
74
|
+
- **Which kind of thing is it?** Same wish, three possible homes:
|
|
75
|
+
- "Let me filter to just the active ones" → a **segment** (saved filter).
|
|
76
|
+
- "Let me add up revenue the same way everywhere, on this table" → a **measure** on the table.
|
|
77
|
+
- "Revenue is an _official company number_ people pull onto dashboards" → a **metric** in a collection, with a default month-by-month view so it charts cleanly. Lean: make it a metric when it's a headline figure the org reuses across many questions/dashboards; keep it a measure when it's a table-local convenience.
|
|
78
|
+
- **Where the metric lives.** Metrics sit in a collection (folder). Lean: put the org's blessed ones in the shared **Library** so they surface prominently; keep experimental ones in a working collection until trusted.
|
|
79
|
+
- **Default time dimension for a metric.** A monthly default makes it chart nicely on a dashboard, but doesn't lock anyone out of other groupings. Lean: set a sensible default (usually month) for anything headline; leave it off for raw counts that aren't inherently time-series.
|
|
80
|
+
- **How strict a segment is.** "Active" = last 30 vs 90 days is a real business call with no right answer from the data alone. Lean: surface the few reasonable thresholds with how many rows each catches, let the user pick.
|
|
81
|
+
|
|
82
|
+
Phrase a prudential call as a lean plus a nod:
|
|
83
|
+
|
|
84
|
+
> "I'd save 'revenue' as a metric — Metabase's term for an official, reusable number — rather than a table-only measure, since people pull it onto dashboards a lot. Good?"
|
|
85
|
+
|
|
86
|
+
---
|
|
87
|
+
|
|
88
|
+
## The process
|
|
89
|
+
|
|
90
|
+
### Phase 0 — Understand what's reusable (quietly)
|
|
91
|
+
|
|
92
|
+
Don't narrate. One "Let me see what's here and how people are already slicing it" is plenty. Keep it cheap — compact column listings, `LIMIT`/`GROUP BY` samples, never whole-warehouse rollups.
|
|
93
|
+
|
|
94
|
+
1. **Confirm the analysis-ready tables exist.** List tables; find the wide, clean ones (a transform step's output). If the user is pointing you at raw normalized tables, say so plainly and suggest building the clean table first — don't build a hobbled semantic layer on raw data.
|
|
95
|
+
2. **Find the questions people keep asking.** Search existing saved questions and dashboards (`mb search`, `mb card list`) for repeated filters and repeated calculations — the same "status = active" written eleven times, five hand-rolled versions of revenue. Those repeats _are_ the semantic layer waiting to be named. This is the highest-signal input; mine it before proposing anything.
|
|
96
|
+
3. **Learn the real meanings.** For every candidate segment ("active", "churned", "high-value"), find what the words map to in actual values — distinct values of a status column, the spread of an amount column. Never define on a guessed meaning (hard rule 1).
|
|
97
|
+
4. **Graft onto what the org already tracks.** This is the part a model does worst and a human does best, so lean on the user: a new definition is far more useful when it lines up with the entities and language the organization _already_ uses. Before inventing "customer health score", ask whether there's already a notion of an active/at-risk customer in their world, and match it. Isolated definitions that don't connect to the existing model are low-value. Ask; don't infer the connection from column names.
|
|
98
|
+
5. **Check reach before promising.** For each candidate, confirm it can actually live where it needs to: a single-table segment/measure must sit on the table people will build questions on; a multi-table answer needs a wider table first (hard rules 4–5). Catch this now, not after building something that won't appear.
|
|
99
|
+
|
|
100
|
+
### Phase 1 — Propose the shared vocabulary (plain language)
|
|
101
|
+
|
|
102
|
+
Show, in plain terms, the definitions worth saving — lead with what each _does for the user_, and name the Metabase feature so they learn it:
|
|
103
|
+
|
|
104
|
+
**Segments — saved filters** (so people pull up the same set with one click):
|
|
105
|
+
|
|
106
|
+
> • **Active customers** — ordered in the last 90 days. ~2,400 of your 6,000 customers.
|
|
107
|
+
> • **Big orders** — over $100. About 1 in 5 orders.
|
|
108
|
+
|
|
109
|
+
**Measures — saved calculations** (so everyone adds it up the same way):
|
|
110
|
+
|
|
111
|
+
> • **Net revenue** — total paid, minus refunds.
|
|
112
|
+
> • **Average order value** — net revenue per order.
|
|
113
|
+
|
|
114
|
+
**Metrics — official numbers** (the headline figures, for dashboards):
|
|
115
|
+
|
|
116
|
+
> • **Monthly recurring revenue** — I'd save this as a metric with a month-by-month default, since it's a dashboard headline. Good?
|
|
117
|
+
|
|
118
|
+
Then surface what you're _not_ saving and why ("I left 'orders this week' alone — it's a one-off, not something you'd reuse"). Ask your prudential questions — one at a time, lean-plus-nod. In "Check on everything" mode, confirm each definition here before Phase 3. In "Balanced", ask only the judgment calls. In "Just go", state your picks and move on.
|
|
119
|
+
|
|
120
|
+
### Phase 2 — Iterate (cheap, nothing built yet)
|
|
121
|
+
|
|
122
|
+
Adjust names, meanings, thresholds, and which-kind-of-thing until the user is happy. Re-confirm the final list in one short recap. If a definition turns out to need more than one table, say so plainly and point back to making the table wider — don't smuggle in a join.
|
|
123
|
+
|
|
124
|
+
### Phase 3 — Build, verify quietly, then hard-stop
|
|
125
|
+
|
|
126
|
+
Build each agreed definition. Mechanics (load `mbql` for the definition bodies):
|
|
127
|
+
|
|
128
|
+
- **Segment** → `mb segment create`. Body: `name`, `table_id`, and a `definition` (a flat MBQL filter clause). Update later with `mb segment update <id>` — needs a `revision_message` (the audit note: _why_ it changed). Never delete-and-recreate.
|
|
129
|
+
- **Measure** → `mb measure create`. Body: `name`, `table_id`, and a `definition` holding **exactly one** aggregation. Same `revision_message` rule on update.
|
|
130
|
+
- **Metric** → `mb card create` with the metric shape (`type: "metric"`) — it lives in a **collection**, carries a `dataset_query` (the aggregation) and an optional default time dimension. Put org-blessed ones in the Library collection.
|
|
131
|
+
|
|
132
|
+
Then **verify what the user can't see**, before you hand back:
|
|
133
|
+
|
|
134
|
+
- Each segment actually narrows the rows you expect (`mb query` / preview the count — does "active customers" really return ~2,400?).
|
|
135
|
+
- Each measure and metric returns a sane number, not null or an error.
|
|
136
|
+
- Each definition shows up **where the user will look for it** — on a question built on the right table. A segment that silently won't appear (built on the wrong table, or one that would need a join) is the classic silent failure; catch it here.
|
|
137
|
+
|
|
138
|
+
Then **stop. Hard gate — every mode, no exceptions.** Recap in plain language and hand the user something to open and eyeball:
|
|
139
|
+
|
|
140
|
+
> Done. Here's the shared set you can now reuse:
|
|
141
|
+
>
|
|
142
|
+
> **Segments** (saved filters — in the **Filter** block on the Customers and Orders tables):
|
|
143
|
+
> • **Active customers** — ordered in the last 90 days
|
|
144
|
+
> • **Big orders** — over $100
|
|
145
|
+
>
|
|
146
|
+
> **Measures** (saved calculations — in the **Summarize** block):
|
|
147
|
+
> • **Net revenue** • **Average order value**
|
|
148
|
+
>
|
|
149
|
+
> **Metric** (in your **Library**, charts by month):
|
|
150
|
+
> • **Monthly recurring revenue**
|
|
151
|
+
>
|
|
152
|
+
> Open any of those tables' Filter or Summarize block in Metabase to see them in place and try one — give it a look before you start building dashboards on top.
|
|
153
|
+
|
|
154
|
+
End on that plain-language map. It's what the user reads to trust the result — and it's what stops a wrong definition from quietly propagating into everything built next.
|
|
155
|
+
|
|
156
|
+
---
|
|
157
|
+
|
|
158
|
+
## A worked example (for your reference, not the user's)
|
|
159
|
+
|
|
160
|
+
User: _"Everyone calculates 'active users' differently — can you make it official?"_
|
|
161
|
+
|
|
162
|
+
- **Don't** create a segment from the phrase alone. **Find the real meaning first:** search existing questions — three people filter on "last seen in the last 30 days", two on "subscription status = active". That's the ambiguity to resolve. Ask: "I see two takes on 'active' — seen in the last 30 days, or has a live subscription. Which do you mean?" (hard rule 1).
|
|
163
|
+
- They say "live subscription, and seen in the last 30 days." **Check reach:** both pieces of info must live on the one table people build questions on. If subscription status and last-seen sit on two different tables, a single segment can't span them (hard rule 4) — to the user: "those two facts live in different places right now, so I'll widen your Customers table to carry both first, then save the filter on it." Build the transform, then the segment on the wide table.
|
|
164
|
+
- Build it as a segment on the wide table. **Verify** the row count is plausible. **Recap** plainly and stop: "Saved **Active users** — live subscription and seen in the last 30 days — as a segment on your Customers table; it's in the Filter block there. Have a look before you build on it."
|
|
165
|
+
|
|
166
|
+
The shape recurs: a word people use loosely → pin it to real values → check it can live where they'll use it → build → verify → hard-stop with a plain recap.
|