@metabase/cli 0.1.10 → 0.1.11
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/marketplace.json +1 -1
- package/README.md +9 -6
- package/dist/{add-collection-DUqTrC5T.mjs → add-collection-DKULSPwT.mjs} +4 -4
- package/dist/add-collection-DyEe_xRO.mjs +10 -0
- package/dist/{archive-44EWiXud.mjs → archive-BJ6uL3hh.mjs} +3 -3
- package/dist/{archive-BEIyIsin.mjs → archive-BL_0fXhf.mjs} +3 -3
- package/dist/{archive-DtE2H4A6.mjs → archive-BLbaqc-5.mjs} +3 -3
- package/dist/{archive-_GMNY8wH.mjs → archive-BLxX3J6I.mjs} +3 -3
- package/dist/{archive-B59Y7ajB.mjs → archive-DQsV2uwy.mjs} +3 -3
- package/dist/{archive-0krZxAXq.mjs → archive-ibJVYsBC.mjs} +3 -3
- package/dist/{archive-BZfpjMir.mjs → archive-y-_m3XID.mjs} +3 -3
- package/dist/auth-BKaQjcmP.mjs +19 -0
- package/dist/{body-BdRyuvU4.mjs → body-DCQg78AF.mjs} +1 -1
- package/dist/{branches-Jpv-FNds.mjs → branches-CJVscsZe.mjs} +4 -4
- package/dist/{cancel-ChC4lFd4.mjs → cancel-C8W0sPix.mjs} +3 -3
- package/dist/{cancel-task-DyhIkNaL.mjs → cancel-task-dDgQDswT.mjs} +4 -4
- package/dist/card-CGBvUmCW.mjs +20 -0
- package/dist/{cards-O-nKkQKP.mjs → cards-B7HE8BT6.mjs} +3 -3
- package/dist/cli.mjs +23 -23
- package/dist/collection-Ch7pPKHD.mjs +20 -0
- package/dist/{collection-namespace-7724zUMx.mjs → collection-namespace-CBCApZJf.mjs} +1 -1
- package/dist/{create-DS52EhPd.mjs → create-84jcfY4N.mjs} +3 -3
- package/dist/{create-BuKx7kw6.mjs → create-B6mzc8UK.mjs} +3 -3
- package/dist/{create-B4f4Pldw.mjs → create-BQA6nM5n.mjs} +4 -4
- package/dist/{create-BxRsXQrm.mjs → create-BWaBYzhC.mjs} +3 -3
- package/dist/{create-CvYKJOcE.mjs → create-DN8us8si.mjs} +3 -3
- package/dist/{create-BbF9zVFP.mjs → create-DPASPFGi.mjs} +4 -4
- package/dist/{create-D45uFXlo.mjs → create-Di9NY7kK.mjs} +3 -3
- package/dist/{create-BnFHcnlL.mjs → create-VjvMdgM5.mjs} +3 -3
- package/dist/{create-branch-KOWUIE72.mjs → create-branch-CCoDX2Qf.mjs} +4 -4
- package/dist/{create-DeZ2x2Db.mjs → create-vNcD4jfq.mjs} +3 -3
- package/dist/{current-task-ClcWPPMc.mjs → current-task-QdxA3mSX.mjs} +4 -4
- package/dist/dashboard-BDb2krMr.mjs +21 -0
- package/dist/db-CoVdkdVm.mjs +22 -0
- package/dist/{delete-D68oS73R.mjs → delete-ByJ6guaH.mjs} +3 -3
- package/dist/{delete-CX2VUA5R.mjs → delete-CE_Jfjni.mjs} +3 -3
- package/dist/{delete-table-DLvL9mDA.mjs → delete-table-D5LDUkOk.mjs} +3 -3
- package/dist/{dirty-1OrXpc7E.mjs → dirty-C6yzkIPn.mjs} +4 -4
- package/dist/document-BA2S1ev3.mjs +19 -0
- package/dist/{eid-CzLhHZMW.mjs → eid-ZOyix96k.mjs} +3 -3
- package/dist/{error-BWXBhqLW.mjs → error-C1bUYXn7.mjs} +2 -1
- package/dist/{export-B5z8w-xo.mjs → export-BhbBX-pl.mjs} +6 -6
- package/dist/field-B5bptFqL.mjs +18 -0
- package/dist/{fields-sC7pzmPX.mjs → fields-CZ6IK8Qf.mjs} +3 -3
- package/dist/{get-ByZ4HR2T.mjs → get-157SWRte.mjs} +3 -3
- package/dist/{get-Bo1FGyFs.mjs → get-9b0ZEWh1.mjs} +3 -3
- package/dist/{get-Dl62Fy6Y.mjs → get-B03gdAun.mjs} +3 -3
- package/dist/{get-CHb6J908.mjs → get-BKWZA_22.mjs} +2 -2
- package/dist/{get-DTHLETau.mjs → get-Bbo2Y-aL.mjs} +3 -3
- package/dist/{get-ZesERdyk.mjs → get-BtqQ98nb.mjs} +2 -2
- package/dist/{get-reSMTfQi.mjs → get-C0emTO1r.mjs} +2 -2
- package/dist/{get-DLwb_gUh.mjs → get-CqKttHWW.mjs} +3 -3
- package/dist/{get-BYw3xS0X.mjs → get-D49EdHgH.mjs} +3 -3
- package/dist/{get-lcX52Skc.mjs → get-DIyv2a7h.mjs} +3 -3
- package/dist/{get-BC60bhel.mjs → get-DclzgC5G.mjs} +3 -3
- package/dist/{get-Fn9WkNhS.mjs → get-DlzC0xCg.mjs} +3 -3
- package/dist/{get-4GEDd9YN.mjs → get-mR01pCNG.mjs} +3 -3
- package/dist/{get-C6n86-dS.mjs → get-nl41R-2q.mjs} +3 -3
- package/dist/{get-run-CpCbHJad.mjs → get-run-_dlMCGp5.mjs} +3 -3
- package/dist/git-sync-DqJmONiB.mjs +28 -0
- package/dist/{has-remote-changes-CPz_-uxd.mjs → has-remote-changes-CToITRyl.mjs} +4 -4
- package/dist/{import-DaWgprK6.mjs → import-xnYHb4QX.mjs} +6 -6
- package/dist/{is-dirty-DZlI7lQx.mjs → is-dirty-BHyMmUVV.mjs} +3 -3
- package/dist/is-dirty-BvA3yzRc.mjs +9 -0
- package/dist/{items-hgbRYsYD.mjs → items-D3AmbIOO.mjs} +3 -3
- package/dist/{list-DO8T-nmF.mjs → list-5ZZfaOzm.mjs} +2 -2
- package/dist/{list-DBlsRSpZ.mjs → list-BHCVZh7Q.mjs} +2 -2
- package/dist/{list-CCZnH2-Z.mjs → list-BdStkg8G.mjs} +2 -2
- package/dist/{list-CGdOC9zX.mjs → list-BhZYkbc9.mjs} +2 -2
- package/dist/{list-kath_2cX.mjs → list-BkzrFyJO.mjs} +2 -2
- package/dist/{list-Cd2nOCAx.mjs → list-BtTRJyrR.mjs} +2 -2
- package/dist/{list-DOVX3vCb.mjs → list-Byn-dRW4.mjs} +3 -3
- package/dist/{list-BgHESP7b.mjs → list-C5NywUBH.mjs} +2 -2
- package/dist/{list-B9J3ujwn.mjs → list-CA6MlByL.mjs} +2 -2
- package/dist/{list-BWv5y307.mjs → list-CIu1zfpU.mjs} +2 -2
- package/dist/{list-BGerkRHH.mjs → list-C_d7ppOJ.mjs} +3 -3
- package/dist/{list-CKeVS1IZ.mjs → list-D7Np4fwk.mjs} +2 -2
- package/dist/{list-D9EuxFHO.mjs → list-DKPhFATP.mjs} +2 -2
- package/dist/{list-36H-dvJZ.mjs → list-DSJwMHXB.mjs} +2 -2
- package/dist/{login-BzrAGJfu.mjs → login-jWvDDvEm.mjs} +3 -3
- package/dist/{logout-CKBiltoS.mjs → logout-Bltlsm6u.mjs} +2 -2
- package/dist/measure-DBX4k13t.mjs +19 -0
- package/dist/{metadata-BjOrKtnv.mjs → metadata-BKcyjsoA.mjs} +3 -3
- package/dist/{metadata-BcGcUEVJ.mjs → metadata-Bf2TfG8F.mjs} +3 -3
- package/dist/{parse-id--iVTCKSo.mjs → parse-id-CWRJlQVK.mjs} +1 -1
- package/dist/{path-LgGU6Bd0.mjs → path-Djw9IARu.mjs} +2 -2
- package/dist/{poll-BucRFJT-.mjs → poll-hJrV0cPm.mjs} +1 -1
- package/dist/{poll-task-DTzKB3T3.mjs → poll-task-DK6X9Juu.mjs} +2 -2
- package/dist/{preflight-CzqVX0PP.mjs → preflight-DjSPp3Ct.mjs} +1 -1
- package/dist/{query-BVZkK6Qk.mjs → query-Cc-9oPsW.mjs} +3 -3
- package/dist/{query-DG_jygDF.mjs → query-Do5n5SZc.mjs} +3 -3
- package/dist/{remove-collection-HdeAfLyi.mjs → remove-collection-DVvnmU06.mjs} +6 -6
- package/dist/{rescan-values-hWubCruZ.mjs → rescan-values-CHluR3JC.mjs} +3 -3
- package/dist/{run-C1-lDmQF.mjs → run-CcJbcaw8.mjs} +5 -5
- package/dist/{runs-Q6DYQyqj.mjs → runs-Dc8qZBhl.mjs} +3 -3
- package/dist/{runtime-colqvhLf.mjs → runtime-jWkzCPlL.mjs} +1 -1
- package/dist/{schema-tables-BEastV_8.mjs → schema-tables-CZ6EMHUZ.mjs} +3 -3
- package/dist/{schemas-CgawwI_k.mjs → schemas-Br1cpOZc.mjs} +3 -3
- package/dist/{search-BWo7xSPP.mjs → search-DB2joPgt.mjs} +3 -3
- package/dist/segment-BWcyA1P-.mjs +19 -0
- package/dist/{set-7Nm2ZTb_.mjs → set-BRhVzEWd.mjs} +3 -3
- package/dist/{setting-DSGXJehQ.mjs → setting-CuxO4T1n.mjs} +3 -3
- package/dist/{setup-aJLGLrIT.mjs → setup-Bn4jwxdl.mjs} +3 -3
- package/dist/{skills-Q2AFsYvc.mjs → skills-Db1r65S5.mjs} +3 -3
- package/dist/snippet-C9UM6NwC.mjs +19 -0
- package/dist/{stash-DPQ0c-Cd.mjs → stash-Dp2JFgmp.mjs} +6 -6
- package/dist/{status-CvAATvV0.mjs → status-BxGBwALE.mjs} +2 -2
- package/dist/{status-CvKPrV5X.mjs → status-j09luNai.mjs} +5 -5
- package/dist/{summary-CeOnoOq2.mjs → summary-DVrj7DSa.mjs} +3 -3
- package/dist/{sync-schema-aOPBc3CY.mjs → sync-schema-BVHvjzGX.mjs} +5 -5
- package/dist/table-DYI8K12W.mjs +19 -0
- package/dist/transform-CvbMbJtI.mjs +24 -0
- package/dist/transform-job-CtzgvccZ.mjs +19 -0
- package/dist/{tree-MOQOBeAP.mjs → tree-BEPw6xbY.mjs} +2 -2
- package/dist/{update-BWyCK8QV.mjs → update-B5VBJFU9.mjs} +5 -5
- package/dist/{update-Bj9s0ri8.mjs → update-BXaT0Faw.mjs} +4 -4
- package/dist/{update-CitS-QRN.mjs → update-Bd0ASKUG.mjs} +5 -5
- package/dist/{update-BD9xkglP.mjs → update-Bn5wUyTQ.mjs} +4 -4
- package/dist/{update-BRrnfG0q.mjs → update-C5vlBTII.mjs} +4 -4
- package/dist/{update-Dri4Zg2H.mjs → update-C875_WGf.mjs} +4 -4
- package/dist/{update-DfNKr_vS.mjs → update-CAoOJRwz.mjs} +4 -4
- package/dist/{update-57uxZWcR.mjs → update-DiG6SSGm.mjs} +4 -4
- package/dist/{update-BkMWBzvk.mjs → update-DzKPxRBE.mjs} +4 -4
- package/dist/{update-dashcard-D_-ura3Y.mjs → update-dashcard-7bZtw04E.mjs} +4 -4
- package/dist/{update-B6mg3AZD.mjs → update-fVqLtZ_t.mjs} +4 -4
- package/dist/{upgrade-CFkZ4USY.mjs → upgrade-DdCxE9Ml.mjs} +2 -2
- package/dist/{uuid-DpinhSxA.mjs → uuid-DywghRXO.mjs} +2 -2
- package/dist/{values-BSS4DRxk.mjs → values-DSul3siC.mjs} +3 -3
- package/dist/{verify-B_A7v8TY.mjs → verify-DzDyKXOa.mjs} +1 -1
- package/dist/{wait-D3iSnjMM.mjs → wait-CacFZI5Y.mjs} +5 -5
- package/dist/{wait-flags-_LnHOeBA.mjs → wait-flags-DV419apK.mjs} +2 -2
- package/package.json +2 -1
- package/skill-data/core/SKILL.md +22 -23
- package/skill-data/data-analysis/SKILL.md +65 -0
- package/skill-data/data-transformation/SKILL.md +200 -0
- package/skill-data/document/SKILL.md +4 -4
- package/skill-data/mbql/SKILL.md +20 -20
- package/skill-data/robot-data-engineer/SKILL.md +142 -0
- package/skill-data/semantic-layer/SKILL.md +166 -0
- package/skill-data/transform/SKILL.md +46 -48
- package/skill-data/visualization/SKILL.md +5 -3
- package/skills/metabase-cli/SKILL.md +6 -0
- package/dist/add-collection-D9wXgmRj.mjs +0 -10
- package/dist/auth-cFC5m69m.mjs +0 -19
- package/dist/card-ClvGX6dQ.mjs +0 -20
- package/dist/collection-DjvowSJC.mjs +0 -20
- package/dist/dashboard-BFeURTOw.mjs +0 -21
- package/dist/db-CSH1kwQr.mjs +0 -22
- package/dist/document-KdT_Xj6r.mjs +0 -19
- package/dist/field-CTFnZI8G.mjs +0 -18
- package/dist/git-sync-C2vib8rx.mjs +0 -28
- package/dist/is-dirty-Bb0Rtj7x.mjs +0 -9
- package/dist/measure-Dw1QpRZa.mjs +0 -19
- package/dist/segment-DMuYvFjg.mjs +0 -19
- package/dist/snippet-Df2TrP7-.mjs +0 -19
- package/dist/table-pK4OkVtL.mjs +0 -19
- package/dist/transform-D60veFH8.mjs +0 -24
- package/dist/transform-job-DXt5LsrY.mjs +0 -19
|
@@ -1,10 +1,10 @@
|
|
|
1
1
|
import "./command-augment-BH9qgQ5u.mjs";
|
|
2
|
-
import { connectionFlags, outputFlags, profileFlag } from "./error-
|
|
3
|
-
import { defineMetabaseCommand } from "./runtime-
|
|
2
|
+
import { connectionFlags, outputFlags, profileFlag } from "./error-C1bUYXn7.mjs";
|
|
3
|
+
import { defineMetabaseCommand } from "./runtime-jWkzCPlL.mjs";
|
|
4
4
|
import { renderSummary } from "./capabilities-7L9GVMd_.mjs";
|
|
5
5
|
import "./input-xewHccej.mjs";
|
|
6
|
-
import { parseId } from "./parse-id
|
|
7
|
-
import { readBody } from "./body-
|
|
6
|
+
import { parseId } from "./parse-id-CWRJlQVK.mjs";
|
|
7
|
+
import { readBody } from "./body-DCQg78AF.mjs";
|
|
8
8
|
import { bodyInputFlags } from "./body-flags-D78h_-Ua.mjs";
|
|
9
9
|
import "./validate-B62TRDGV.mjs";
|
|
10
10
|
import { SEGMENT_DEFINITION_LABELS, preflightMbql5Query, skipValidateFlag } from "./validate-query-BpiN1CFu.mjs";
|
|
@@ -1,10 +1,10 @@
|
|
|
1
1
|
import "./command-augment-BH9qgQ5u.mjs";
|
|
2
|
-
import { connectionFlags, outputFlags, profileFlag } from "./error-
|
|
3
|
-
import { defineMetabaseCommand } from "./runtime-
|
|
2
|
+
import { connectionFlags, outputFlags, profileFlag } from "./error-C1bUYXn7.mjs";
|
|
3
|
+
import { defineMetabaseCommand } from "./runtime-jWkzCPlL.mjs";
|
|
4
4
|
import { renderSummary } from "./capabilities-7L9GVMd_.mjs";
|
|
5
5
|
import "./input-xewHccej.mjs";
|
|
6
|
-
import { parseId } from "./parse-id
|
|
7
|
-
import { readBody } from "./body-
|
|
6
|
+
import { parseId } from "./parse-id-CWRJlQVK.mjs";
|
|
7
|
+
import { readBody } from "./body-DCQg78AF.mjs";
|
|
8
8
|
import { bodyInputFlags } from "./body-flags-D78h_-Ua.mjs";
|
|
9
9
|
import { Document, DocumentUpdateInput, documentView } from "./document-CcfiiV3b.mjs";
|
|
10
10
|
|
|
@@ -1,12 +1,12 @@
|
|
|
1
1
|
import "./command-augment-BH9qgQ5u.mjs";
|
|
2
|
-
import { connectionFlags, outputFlags, profileFlag } from "./error-
|
|
3
|
-
import { defineMetabaseCommand } from "./runtime-
|
|
2
|
+
import { connectionFlags, outputFlags, profileFlag } from "./error-C1bUYXn7.mjs";
|
|
3
|
+
import { defineMetabaseCommand } from "./runtime-jWkzCPlL.mjs";
|
|
4
4
|
import { renderSummary } from "./capabilities-7L9GVMd_.mjs";
|
|
5
5
|
import "./input-xewHccej.mjs";
|
|
6
6
|
import "./field-MGxpNQUH.mjs";
|
|
7
7
|
import { Card, CardUpdateInput, cardView } from "./card-DnIeMmUn.mjs";
|
|
8
|
-
import { parseId } from "./parse-id
|
|
9
|
-
import { readBody } from "./body-
|
|
8
|
+
import { parseId } from "./parse-id-CWRJlQVK.mjs";
|
|
9
|
+
import { readBody } from "./body-DCQg78AF.mjs";
|
|
10
10
|
import { bodyInputFlags } from "./body-flags-D78h_-Ua.mjs";
|
|
11
11
|
import "./validate-B62TRDGV.mjs";
|
|
12
12
|
import { CARD_DATASET_QUERY_LABELS, preflightMbql5Query, skipValidateFlag } from "./validate-query-BpiN1CFu.mjs";
|
|
@@ -1,10 +1,10 @@
|
|
|
1
1
|
import { ConfigError } from "./command-augment-BH9qgQ5u.mjs";
|
|
2
|
-
import { connectionFlags, outputFlags, profileFlag } from "./error-
|
|
3
|
-
import { defineMetabaseCommand } from "./runtime-
|
|
2
|
+
import { connectionFlags, outputFlags, profileFlag } from "./error-C1bUYXn7.mjs";
|
|
3
|
+
import { defineMetabaseCommand } from "./runtime-jWkzCPlL.mjs";
|
|
4
4
|
import { renderSummary } from "./capabilities-7L9GVMd_.mjs";
|
|
5
5
|
import "./input-xewHccej.mjs";
|
|
6
|
-
import { parseId } from "./parse-id
|
|
7
|
-
import { readBody } from "./body-
|
|
6
|
+
import { parseId } from "./parse-id-CWRJlQVK.mjs";
|
|
7
|
+
import { readBody } from "./body-DCQg78AF.mjs";
|
|
8
8
|
import { bodyInputFlags } from "./body-flags-D78h_-Ua.mjs";
|
|
9
9
|
import { DashboardDetail, Dashcard, DashcardPatchInput, dashcardView } from "./dashboard-FY5UzJ_Z.mjs";
|
|
10
10
|
|
|
@@ -1,11 +1,11 @@
|
|
|
1
1
|
import "./command-augment-BH9qgQ5u.mjs";
|
|
2
|
-
import { connectionFlags, outputFlags, profileFlag } from "./error-
|
|
3
|
-
import { defineMetabaseCommand } from "./runtime-
|
|
2
|
+
import { connectionFlags, outputFlags, profileFlag } from "./error-C1bUYXn7.mjs";
|
|
3
|
+
import { defineMetabaseCommand } from "./runtime-jWkzCPlL.mjs";
|
|
4
4
|
import { renderSummary } from "./capabilities-7L9GVMd_.mjs";
|
|
5
5
|
import "./input-xewHccej.mjs";
|
|
6
6
|
import { Field, FieldUpdateInput, fieldView } from "./field-MGxpNQUH.mjs";
|
|
7
|
-
import { parseId } from "./parse-id
|
|
8
|
-
import { readBody } from "./body-
|
|
7
|
+
import { parseId } from "./parse-id-CWRJlQVK.mjs";
|
|
8
|
+
import { readBody } from "./body-DCQg78AF.mjs";
|
|
9
9
|
import { bodyInputFlags } from "./body-flags-D78h_-Ua.mjs";
|
|
10
10
|
|
|
11
11
|
//#region src/commands/field/update.ts
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
import { AbortError, NetworkError, TimeoutError, UnknownError, errorMessage, isNotFoundError } from "./command-augment-BH9qgQ5u.mjs";
|
|
2
|
-
import { outputFlags, package_default } from "./error-
|
|
3
|
-
import { HttpError, USER_AGENT, combineAborts, defineMetabaseCommand, parseJson, throwIfAborted } from "./runtime-
|
|
2
|
+
import { outputFlags, package_default } from "./error-C1bUYXn7.mjs";
|
|
3
|
+
import { HttpError, USER_AGENT, combineAborts, defineMetabaseCommand, parseJson, throwIfAborted } from "./runtime-jWkzCPlL.mjs";
|
|
4
4
|
import { renderItem, writeText } from "./capabilities-7L9GVMd_.mjs";
|
|
5
5
|
import { promptConfirm } from "./prompt-u4WhE4T5.mjs";
|
|
6
6
|
import { z } from "zod";
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
import { ConfigError } from "./command-augment-BH9qgQ5u.mjs";
|
|
2
|
-
import { outputFlags } from "./error-
|
|
3
|
-
import { defineMetabaseCommand, parseInteger } from "./runtime-
|
|
2
|
+
import { outputFlags } from "./error-C1bUYXn7.mjs";
|
|
3
|
+
import { defineMetabaseCommand, parseInteger } from "./runtime-jWkzCPlL.mjs";
|
|
4
4
|
import { writeJson, writeText } from "./capabilities-7L9GVMd_.mjs";
|
|
5
5
|
import { z } from "zod";
|
|
6
6
|
import { randomUUID } from "node:crypto";
|
|
@@ -1,9 +1,9 @@
|
|
|
1
1
|
import "./command-augment-BH9qgQ5u.mjs";
|
|
2
|
-
import { connectionFlags, outputFlags, profileFlag } from "./error-
|
|
3
|
-
import { defineMetabaseCommand } from "./runtime-
|
|
2
|
+
import { connectionFlags, outputFlags, profileFlag } from "./error-C1bUYXn7.mjs";
|
|
3
|
+
import { defineMetabaseCommand } from "./runtime-jWkzCPlL.mjs";
|
|
4
4
|
import { formatScalar, renderSummary } from "./capabilities-7L9GVMd_.mjs";
|
|
5
5
|
import { FieldValues, fieldValuesView } from "./field-MGxpNQUH.mjs";
|
|
6
|
-
import { parseId } from "./parse-id
|
|
6
|
+
import { parseId } from "./parse-id-CWRJlQVK.mjs";
|
|
7
7
|
|
|
8
8
|
//#region src/commands/field/values.ts
|
|
9
9
|
var values_default = defineMetabaseCommand({
|
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
import { MetabaseError, NetworkError, TimeoutError, errorMessage } from "./command-augment-BH9qgQ5u.mjs";
|
|
2
|
-
import { HttpError, createClient, probeServer } from "./runtime-
|
|
2
|
+
import { HttpError, createClient, probeServer } from "./runtime-jWkzCPlL.mjs";
|
|
3
3
|
import { z } from "zod";
|
|
4
4
|
|
|
5
5
|
//#region src/domain/user.ts
|
|
@@ -1,10 +1,10 @@
|
|
|
1
1
|
import "./command-augment-BH9qgQ5u.mjs";
|
|
2
|
-
import { connectionFlags, outputFlags, profileFlag } from "./error-
|
|
3
|
-
import { defineMetabaseCommand } from "./runtime-
|
|
2
|
+
import { connectionFlags, outputFlags, profileFlag } from "./error-C1bUYXn7.mjs";
|
|
3
|
+
import { defineMetabaseCommand } from "./runtime-jWkzCPlL.mjs";
|
|
4
4
|
import { renderSummary } from "./capabilities-7L9GVMd_.mjs";
|
|
5
|
-
import { parseId } from "./parse-id
|
|
6
|
-
import { DEFAULT_INTERVAL_MS, DEFAULT_TIMEOUT_MS } from "./poll-
|
|
7
|
-
import { SyncTaskOrIdle, formatSyncTask, pollSyncTask, syncTaskIdleView, syncTaskView, throwIfFailedTask } from "./poll-task-
|
|
5
|
+
import { parseId } from "./parse-id-CWRJlQVK.mjs";
|
|
6
|
+
import { DEFAULT_INTERVAL_MS, DEFAULT_TIMEOUT_MS } from "./poll-hJrV0cPm.mjs";
|
|
7
|
+
import { SyncTaskOrIdle, formatSyncTask, pollSyncTask, syncTaskIdleView, syncTaskView, throwIfFailedTask } from "./poll-task-DK6X9Juu.mjs";
|
|
8
8
|
|
|
9
9
|
//#region src/commands/git-sync/wait.ts
|
|
10
10
|
const WaitResult = SyncTaskOrIdle;
|
|
@@ -1,5 +1,5 @@
|
|
|
1
|
-
import { parseId } from "./parse-id
|
|
2
|
-
import { DEFAULT_INTERVAL_MS, DEFAULT_TIMEOUT_MS } from "./poll-
|
|
1
|
+
import { parseId } from "./parse-id-CWRJlQVK.mjs";
|
|
2
|
+
import { DEFAULT_INTERVAL_MS, DEFAULT_TIMEOUT_MS } from "./poll-hJrV0cPm.mjs";
|
|
3
3
|
|
|
4
4
|
//#region src/commands/wait-flags.ts
|
|
5
5
|
const waitScheduleFlags = {
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@metabase/cli",
|
|
3
|
-
"version": "0.1.
|
|
3
|
+
"version": "0.1.11",
|
|
4
4
|
"description": "Metabase CLI",
|
|
5
5
|
"license": "AGPL-3.0",
|
|
6
6
|
"repository": {
|
|
@@ -37,6 +37,7 @@
|
|
|
37
37
|
"typecheck": "tsc --noEmit",
|
|
38
38
|
"lint": "oxlint",
|
|
39
39
|
"lint:fix": "oxlint --fix",
|
|
40
|
+
"lint:skills": "uvx skillsaw lint skill-data/ --strict",
|
|
40
41
|
"format": "oxfmt",
|
|
41
42
|
"format:check": "oxfmt --check",
|
|
42
43
|
"sync:representations": "bun run scripts/sync-representations.ts",
|
package/skill-data/core/SKILL.md
CHANGED
|
@@ -6,7 +6,7 @@ allowed-tools: Read, Write, Edit, Bash, AskUserQuestion
|
|
|
6
6
|
|
|
7
7
|
# metabase-cli (core)
|
|
8
8
|
|
|
9
|
-
The official Metabase CLI (`mb`) drives a Metabase instance over its REST API
|
|
9
|
+
The official Metabase CLI (`mb`) drives a Metabase instance over its REST API: auth, list/get/create/update/delete on every resource, query and transform execution, content search, git-sync (representations ↔ instance), and entity-id translation.
|
|
10
10
|
|
|
11
11
|
Top-level command groups (run `mb <group> --help` to discover verbs):
|
|
12
12
|
|
|
@@ -15,7 +15,7 @@ auth | db | table | field | query | card | dashboard | snippet | segment | measu
|
|
|
15
15
|
document | transform | transform-job | setting | search | git-sync | setup | eid | uuid | upgrade | skills
|
|
16
16
|
```
|
|
17
17
|
|
|
18
|
-
The patterns below — auth, flag conventions, output flags, body input — apply across **every** group. Per-command flags, examples, and output schemas live in `mb __manifest` (see below). A few flows have their own specialized skills
|
|
18
|
+
The patterns below — auth, flag conventions, output flags, body input — apply across **every** group. Per-command flags, examples, and output schemas live in `mb __manifest` (see below). A few flows have their own specialized skills (see "Specialized skills"). When a card needs a query, prefer MBQL over native SQL (portable, pre-flight-validated) — load `mbql`; fall back to native SQL when MBQL can't express it.
|
|
19
19
|
|
|
20
20
|
## Auth & profiles
|
|
21
21
|
|
|
@@ -29,11 +29,11 @@ mb auth status --json # → {profile, present, url} for the d
|
|
|
29
29
|
mb auth status --profile <name> --json # → status of a specific profile
|
|
30
30
|
```
|
|
31
31
|
|
|
32
|
-
`auth list` is the primary enumeration path — one call returns every configured profile with sanitized URL, an `authenticated` flag, and a probe `status` (`ok` / `auth-failed` / `network-error` / `server-error` / `not-probed`). Use it before asking
|
|
32
|
+
`auth list` is the primary enumeration path — one call returns every configured profile with sanitized URL, an `authenticated` flag, and a probe `status` (`ok` / `auth-failed` / `network-error` / `server-error` / `not-probed`). Use it before asking which profile to pick. If it returns an empty `data: []`, ask the user to run `mb auth login` themselves (see the policy above) and tell you the profile name. `auth status` is a single-profile health probe when you already know the name.
|
|
33
33
|
|
|
34
34
|
### Pick the profile to use
|
|
35
35
|
|
|
36
|
-
If exactly one profile is configured and
|
|
36
|
+
If exactly one profile is configured and intent doesn't disambiguate, use it. If multiple exist and the user hasn't named one, ask via `AskUserQuestion`, presenting the names from `auth list`. Once a name is established, pass `--profile <name>` to **every** subsequent command. Profile names are arbitrary local labels — `prod`, `staging` — let the user pick.
|
|
37
37
|
|
|
38
38
|
## Flag conventions
|
|
39
39
|
|
|
@@ -52,14 +52,12 @@ If exactly one profile is configured and the user's intent doesn't disambiguate,
|
|
|
52
52
|
|
|
53
53
|
### Some outputs are JSON envelopes, not bare strings
|
|
54
54
|
|
|
55
|
-
A handful of "lookup" verbs return a JSON object even
|
|
55
|
+
A handful of "lookup" verbs return a JSON object even for a single field. `mb setting get <key>` returns `{"key": "...", "value": ...}`, not the bare value. Extract before reusing:
|
|
56
56
|
|
|
57
57
|
```bash
|
|
58
58
|
VALUE=$(mb setting get <key> --json | jq -r '.value')
|
|
59
59
|
```
|
|
60
60
|
|
|
61
|
-
If you find yourself piping a `--json` envelope straight into another flag and the receiving command rejects it, this is what happened.
|
|
62
|
-
|
|
63
61
|
## Output
|
|
64
62
|
|
|
65
63
|
Every list/get verb supports the same output flags:
|
|
@@ -94,26 +92,28 @@ Verbs that take a payload accept it from one of four sources, **first non-empty
|
|
|
94
92
|
3. stdin (auto-detected when piped, or explicit `--stdin` where supported)
|
|
95
93
|
4. positional argument
|
|
96
94
|
|
|
97
|
-
|
|
95
|
+
Exactly one required; passing two of `--body` + `--file` + `--stdin` is rejected with a `ConfigError`.
|
|
98
96
|
|
|
99
97
|
```bash
|
|
100
|
-
cat > /
|
|
98
|
+
cat > ./.scratch/body.json <<'EOF'
|
|
101
99
|
{ ... }
|
|
102
100
|
EOF
|
|
103
|
-
mb <noun> create --file /
|
|
101
|
+
mb <noun> create --file ./.scratch/body.json --profile <n> --json
|
|
104
102
|
```
|
|
105
103
|
|
|
106
104
|
Single-quoted `'EOF'` prevents the shell from interpolating `$vars` inside the JSON.
|
|
107
105
|
|
|
106
|
+
Write these working files to **`./.scratch`** in the current directory (`mkdir -p ./.scratch` first), never `/tmp` — better permissions, they persist across the session, and the user can review them.
|
|
107
|
+
|
|
108
108
|
## Discover the full surface: `mb __manifest`
|
|
109
109
|
|
|
110
|
-
|
|
110
|
+
The canonical, machine-readable inventory of every command — name, description, per-command `details`, examples, every flag with type and default, and the output JSON Schema:
|
|
111
111
|
|
|
112
112
|
```bash
|
|
113
113
|
mb __manifest
|
|
114
114
|
```
|
|
115
115
|
|
|
116
|
-
The leading `__` hides it from `--help`, but it's stable. Reach for it instead of `--help` per command.
|
|
116
|
+
The leading `__` hides it from `--help`, but it's stable. Reach for it instead of `--help` per command — to enumerate verbs, validate flag names before constructing a command, or read an output schema before parsing. Pairs with `jq`:
|
|
117
117
|
|
|
118
118
|
```bash
|
|
119
119
|
mb __manifest | jq -r '.commands[].command' # every command name
|
|
@@ -122,43 +122,42 @@ mb __manifest | jq '.commands[] | select(.command == "card query") | .args'
|
|
|
122
122
|
mb __manifest | jq '.commands[] | select(.command == "card list") | .outputSchema' # output schema before parsing
|
|
123
123
|
```
|
|
124
124
|
|
|
125
|
-
Use it to (a) enumerate verbs, (b) validate flag names before constructing a command, (c) read an output schema before parsing.
|
|
126
|
-
|
|
127
125
|
## Resource quirks worth memorizing
|
|
128
126
|
|
|
129
|
-
Routine verb shapes (list / get / create / update), every flag, and output JSON Schemas live in `mb __manifest` — pull
|
|
127
|
+
Routine verb shapes (list / get / create / update), every flag, and output JSON Schemas live in `mb __manifest` — pull on demand. Below is only what the manifest does _not_ tell you: footguns and non-obvious behaviors.
|
|
130
128
|
|
|
131
129
|
- **db traversal vs. rollup.** Default to granular: `database list` → `database schemas <db-id>` → `database schema-tables <db-id> <schema>` → `table get <table-id> --include fields`. The rollup endpoints (`database get --include tables.fields`, `database metadata <db-id>`) pull megabytes and blow the context window on any real warehouse — use them only on a small/dev db. `sync-schema` / `rescan-values` queue async work and return `{status:"ok"}` immediately; `sync-schema --wait` blocks until `initial_sync_status: complete`.
|
|
132
130
|
- **table fields.** `table get` never returns fields on its own — pass `--include fields` (compact) or use `table fields <id>` (list envelope). `table metadata <id>` adds FKs + dimensions (heavier). `table update` patches table-level metadata only; physical columns aren't editable here.
|
|
133
131
|
- **field has no `list`.** Fields are per-table — get them via `table get <id> --include fields`. Never enumerate fields across a whole db (context blow-up). `field summary` is live cardinality `{field_id, count, distincts}`; `field values` is the cached distinct set (`has_more_values: true` ⇒ truncated cache). `field update` patches metadata only; `base_type` isn't editable.
|
|
134
132
|
- **card.** `dataset_query` is the **flat** `mbql/query` value, not a legacy `{type:"query",query:…}` envelope (→ `mbql` skill). `--export-format csv|xlsx` streams the raw export (pipe to a file), bypassing the JSON envelope. `archive` is the only delete; unarchive with `update --body '{"archived":false}'`. `visualization_settings` keys are scoped by `display` and aren't pre-flighted — see the `viz` skill.
|
|
135
|
-
- **dashboard.** Dashcards round-trip through `PUT /api/dashboard/:id` (no per-dashcard endpoint): `update-dashcard <dash-id> <dashcard-id>` patches one safely; `update --body '{"dashcards":[…]}'` replaces the whole set (omitted ids are deleted server-side;
|
|
133
|
+
- **dashboard.** Dashcards round-trip through `PUT /api/dashboard/:id` (no per-dashcard endpoint): `update-dashcard <dash-id> <dashcard-id>` patches one safely; `update --body '{"dashcards":[…]}'` replaces the whole set (omitted ids are deleted server-side; negative ids for new cards). `create` accepts the **same** `dashcards` array in its initial body — lay out the whole dashboard in one call: negative ids for new cards, and `card_id:null` plus a `visualization_settings.virtual_card` block (`{display:"text"|"heading"|"link"|…}`) for non-question cards. `create`/`update` pre-flight every positive `card_id` against live server state and exit **2** with `{ok:false,errors:[…]}` on a bad ref — non-bypassable (no `--skip-validate`). `dashboard get <id>` (or `--full`) hydrates dashcards/tabs; `list` omits them. **Dashcard geometry: the grid is 24 columns wide.** Each dashcard's `{col, row, size_x, size_y}` is in grid units — `col` (0-indexed, left edge) and `size_x` are columns, `row`/`size_y` are rows; **full-width is `size_x: 24`** (`size_x: 12` is half a row — the usual cause of a card filling only half the width, since it's a common per-chart default). Keep `col + size_x ≤ 24`, start each card's `col` at 0 for a full-width stack, and don't overlap cards (the server stores whatever you send — it won't auto-fix collisions).
|
|
136
134
|
- **snippet `--archived` is a swap, not a union** — list returns _either_ active _or_ archived rows, never both. (Same shape for `--filter archived` on dashboard/collection.)
|
|
137
135
|
- **segment / measure** `update` and `archive` require a non-blank `revision_message` (audit-logged); the CLI does not synthesize it on `update`. `archive` defaults to `"Archived via mb CLI"` — override with `--revision-message`. `definition` is a flat MBQL clause (→ `mbql` skill): segment = a filter, measure = exactly one aggregation.
|
|
138
136
|
- **collection `<ref>`** accepts four forms only — positive int, `root`, `trash`, or a 21-char entity_id — anything else is a client-side `ConfigError`. `collection items` auto-paginates (cap with `--limit`, which then omits `total`). `collection tree` is **JSON-only** — `--format text` is rejected.
|
|
139
137
|
- **setting set** parses the value as **strict JSON**: a string is `'"value"'` (inner quotes), booleans `true`/`false`, numbers bare. Wrong quoting silently errors — confirm with `setting get <key>` after. `setting get --json` works on every value type (it wraps bare-text responses into `{key, value}`).
|
|
140
138
|
- **search vs. list.** For plain enumeration of cards/dashboards/collections use the dedicated `… list` verbs; reach for `search --models <kind>` only for ranking against a query string or a cross-resource lookup.
|
|
141
139
|
- **transform.** Iterate with `transform update <id>`, never `delete` + `create` — keeps the row, `entity_id`, materialized table, and YAML filename (avoids `_2` suffixes and noisy git history). `transform run` needs `--wait` (or `--sync`, which also waits for the run's output table to register and returns `target_table_id`) or you get only `{run_id, final:null}`. (→ `transform` skill.)
|
|
142
|
-
- **setup is one-shot.** `mb setup` walks `/api/setup` for a **fresh** instance only —
|
|
140
|
+
- **setup is one-shot.** `mb setup` walks `/api/setup` for a **fresh** instance only — errors against an already-configured one. Mostly for bootstrapping local / e2e instances.
|
|
143
141
|
- **eid** translates a string entity id → numeric id: `mb eid --model <model> <eid1,eid2> --json` (EIDs are a positional used with `--model`; or pass `--body '{"entity_ids":{"card":["…"]}}'`). Entity ids are NanoIDs that can start with `-`, which the positional form misreads as a flag (shell quotes don't help — the `-` survives into argv). For an id that may start with `-`, use `--body` — the id is a JSON string value, immune to flag parsing: `mb eid --body '{"entity_ids":{"card":["-…"]}}'`. Useful when an external system hands you an entity id and a verb needs the numeric one.
|
|
144
142
|
- **query / uuid.** `mb query` is the ad-hoc MBQL surface (`--print-schema` → `--dry-run` → run); `mb uuid --count <n>` mints the `lib/uuid` values every MBQL 5 clause needs. Both workflows live in the `mbql` skill.
|
|
145
143
|
|
|
146
144
|
## Specialized skills (load on demand)
|
|
147
145
|
|
|
148
|
-
This core file is enough for any single-command task. Load the relevant skill **proactively** when intent matches — don't wing an MBQL body, a transform body, or the git-sync workflow from this overview alone. Load
|
|
146
|
+
This core file is enough for any single-command task. Load the relevant skill **proactively** when intent matches — don't wing an MBQL body, a transform body, or the git-sync workflow from this overview alone. Load via `mb skills get <name>`.
|
|
147
|
+
|
|
148
|
+
**Start here for anything bigger than one command.** If the user wants an outcome rather than a single verb — "make sense of my data", "build a data model", "go from raw data to a dashboard", "be my data analyst", "set up analytics for X", "answer questions about my data" — load `robot-data-engineer` first and let it route. The rest of this list is the toolbox it routes into.
|
|
149
149
|
|
|
150
|
+
- **`robot-data-engineer`** — the front-door router for the whole journey (raw data → clean tables → reusable definitions → dashboards or written answers) for a non-technical user. Detects where the user is, sets up auth and autonomy once, and routes to `data-transformation` / `semantic-layer` / `visualization` / `data-analysis`. Load this when the user describes a goal, not a step.
|
|
150
151
|
- **`mbql`** — authoring or fixing any MBQL query body: `mb query`, a card `dataset_query`, a transform `source.query`, a measure/segment `definition`, "aggregate and group by", reading `--dry-run` errors. The query-body reference.
|
|
151
152
|
- **`viz`** — choosing a card's `display` and authoring `visualization_settings`: "make it a bar chart", "set the pie dimension/metric", "format this column as currency", "the card renders as a table instead of a chart". The presentation counterpart to `mbql`.
|
|
152
153
|
- **`transform`** — "create a transform", "run a transform", authoring transform body JSON, run inspection.
|
|
154
|
+
- **`data-transformation`** — the higher-level workflow: turning a raw, normalized source database into a small set of clean, wide, analysis-ready tables for a non-technical user — "clean up", "flatten", "denormalize", "make sense of this database", "build analysis-ready tables". Wraps `transform` (the mechanics) with the investigate → propose → build flow.
|
|
155
|
+
- **`semantic-layer`** — turning clean tables into reusable definitions: "make this filter reusable", "define active customers / net revenue / MRR officially", "create a segment / measure / metric", "so everyone uses the same definition". Builds on `mbql` (the definition bodies) and `transform` (widen a table first when a definition needs more than one).
|
|
153
156
|
- **`git-sync`** — "import the latest changes", "export to git", "git sync", "dirty check", "stash before pulling".
|
|
154
157
|
|
|
155
158
|
If a task spans more than one, load each. Specialized skills assume the conventions above and won't repeat them. `mb skills list` enumerates everything on the installed version.
|
|
156
159
|
|
|
157
160
|
## Don't
|
|
158
161
|
|
|
159
|
-
- **Don't run `mb auth login` for the user** — authentication is theirs (see §Auth).
|
|
160
162
|
- Don't paste credentials or warehouse passwords in chat. Have the user run the storing command.
|
|
161
|
-
- Don't
|
|
162
|
-
- Don't omit `--wait` on `transform run` / `git-sync import` for interactive flows; the next step will race the operation.
|
|
163
|
-
- Don't drop a JSON-envelope verb's output raw into another flag. Extract with `--json | jq -r '.<field>'`.
|
|
164
|
-
- Don't add a third-party HTTP library or shell into `curl` against `/api/...` when a `mb <verb>` exists — that bypasses retries, schema validation, and credential redaction.
|
|
163
|
+
- Don't shell into `curl` against `/api/...` (or add an HTTP library) when a `mb <verb>` exists — that bypasses retries, schema validation, and credential redaction.
|
|
@@ -0,0 +1,65 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: data-analysis
|
|
3
|
+
description: Answer real questions from clean, analysis-ready tables and hand back a plain-language report - an answer-finding task, not chart-building. Read the tables, turn the user's question into queries, run them on the live instance, sanity-check the numbers, write up findings the user can trust. Works over already-clean (wide, human-readable) data - survey/registration answers, event signups, customer lists, anything where the data holds the answer. Use when someone wants to "answer questions about my data", "report on who registered / signed up / responded", "what did people say", "analyze X", "explore this data", or "build me a report". For a non-technical user who knows their domain. Needs charts/dashboards? Use `visualization`. Tables still raw? Use `data-transformation` first.
|
|
4
|
+
allowed-tools: Read, Write, Edit, Bash, AskUserQuestion
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Data Analysis
|
|
8
|
+
|
|
9
|
+
> **Shared contract (read first).** This skill is part of the `robot-data-engineer` family and follows its shared rules: audience is a non-technical user, so no database jargon (skip "normalize"/"grain"; ERD/foreign key are fine; explain "wide"/"long" the first time you use them). Ask before showing PII row-by-row (names, emails, phones) — default to aggregates. When asked for something the CLI can't do (alerts, dashboard filters), name the limit instead of erroring into raw SQL. Honor the autonomy mode the user picked. Full text and the autonomy slider live in the router — run `mb skills get robot-data-engineer` and read its **Shared Contract** if you haven't.
|
|
10
|
+
|
|
11
|
+
The user has a question and clean data that already holds the answer. Your job: find the answer, check it's right, and hand it back in plain language. You're an analyst, not a dashboard builder — the deliverable is a **trustworthy written answer**, optionally backed by a saved question they can re-open.
|
|
12
|
+
|
|
13
|
+
This skill assumes the tables are already clean (wide, human-readable). If they're raw and normalized — lots of `*_field`/`*_choice` lookups, coded columns, JSON blobs — stop and route to `data-transformation` first; don't analyze on top of a mess.
|
|
14
|
+
|
|
15
|
+
---
|
|
16
|
+
|
|
17
|
+
## The loop
|
|
18
|
+
|
|
19
|
+
For each question the user asks:
|
|
20
|
+
|
|
21
|
+
1. **Find where the answer lives.** List tables (`mb table list`, `mb db schema-tables <db> <schema>`). Read the columns (`mb table fields <id>`). Clean datasets often ship the same facts two ways — a **wide** table (one row per thing, easy to read) and a **long** table (one row per attribute, easy to aggregate over many-valued answers). Pick the one that fits the question: per-person facts → wide; "which option was most popular" across a multi-select → long.
|
|
22
|
+
|
|
23
|
+
2. **Turn the question into a query.** Write it, run it (`mb query`). Start small — a `count(*)` and a couple of sample rows to confirm you're pointed at the right table and the columns mean what you think. Then write the real query.
|
|
24
|
+
|
|
25
|
+
3. **Sanity-check before you believe it.** A number with no cross-check is a guess. Confirm row counts against a total you trust, watch for nulls/blanks inflating or deflating a percentage, and re-read the column you grouped on — a `type/Category` column with "confirmed"/"cancelled" means your "how many registered" answer depends on which statuses you counted. State the denominator.
|
|
26
|
+
|
|
27
|
+
4. **Report in plain language.** Lead with the answer, then how you got it. Numbers get context ("9 of 10 confirmed"), not bare figures. For free-text answers, quote a few real responses rather than only counting them — the words are the value.
|
|
28
|
+
|
|
29
|
+
---
|
|
30
|
+
|
|
31
|
+
## What to ask the user up front
|
|
32
|
+
|
|
33
|
+
Don't over-interrogate, but settle the things that change the answer:
|
|
34
|
+
|
|
35
|
+
- **Scope.** All-time or a window? Everyone, or only confirmed/active? A "how many registered" with no status filter and a "how many _confirmed_" are different numbers — pick the one they mean, and say which you used.
|
|
36
|
+
- **Cut.** Do they want the headline number, or the number broken down (by role, by company, by version)? A breakdown is usually one `GROUP BY` away and far more useful.
|
|
37
|
+
- **Form of the answer.** A number in chat? A short written digest? A saved question they can re-open and refilter? If they want something durable or visual, that's the `visualization` skill — hand off.
|
|
38
|
+
|
|
39
|
+
When genuinely unsure which interpretation they mean, ask — never silently pick one and present it as the answer.
|
|
40
|
+
|
|
41
|
+
---
|
|
42
|
+
|
|
43
|
+
## Survey / registration data — the common shape
|
|
44
|
+
|
|
45
|
+
A lot of "analyze who registered / what did people say" work lands on event or survey data, which has a recognizable shape worth calling out:
|
|
46
|
+
|
|
47
|
+
- A **per-registrant wide table** — name, company, role, status, plus one column per single-answer question. Use it for "who registered", rosters, breakdowns by role/version/company, and any per-person filter.
|
|
48
|
+
- A **long answers table** — one row per (registrant, question, answer). Use it for **multi-select** questions (one person picks several options, so they can't flatten into one wide column) and for "which option was chosen most". Group by the question text, then by the answer value.
|
|
49
|
+
- **Question definitions** — the catalog of what was asked, the answer choices, free-text vs single vs multi. Read this first to know which questions exist and how each is typed before you start counting.
|
|
50
|
+
|
|
51
|
+
Three report families cover most asks:
|
|
52
|
+
|
|
53
|
+
1. **Roster** — who registered, with the facts that matter (company, role, status). A filtered, ordered read of the wide table.
|
|
54
|
+
2. **Distribution** — how the group splits on a single-select (role, version, customer-or-not). A `GROUP BY` with counts; the agent-facing answer is "X% picked A, Y% picked B".
|
|
55
|
+
3. **Open-ended digest** — what people said in free-text ("what do you want to learn / teach / discuss"). Small N usually — list the actual answers, don't just count them; the responses are the point.
|
|
56
|
+
|
|
57
|
+
---
|
|
58
|
+
|
|
59
|
+
## Don't
|
|
60
|
+
|
|
61
|
+
- **Don't analyze raw, un-cleaned tables.** If the data is normalized/coded/JSON, route to `data-transformation` first and analyze the clean output.
|
|
62
|
+
- **Don't report a number you didn't sanity-check.** No denominator, no null-check → no answer.
|
|
63
|
+
- **Don't silently pick a scope.** "Registered" vs "confirmed", all-time vs window — state which you used, or ask.
|
|
64
|
+
- **Don't build charts/dashboards here.** A written answer (and maybe one saved question) is the deliverable; if they want it visual, that's `visualization`.
|
|
65
|
+
- **Don't only count free-text.** Quote the real responses — the words carry the insight a count throws away.
|
|
@@ -0,0 +1,200 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: data-transformation
|
|
3
|
+
description: Turn a raw, normalized source database into a small set of clean, analysis-ready tables. Claude investigates the source, works out the real-world "things" the data is about (even when each one is scattered across several tables), decodes coded/JSON/translated values into readable text, and builds one wide, denormalized table per thing as Metabase transforms. Designed for a non-technical user who knows their domain. Use whenever someone wants to "clean up", "flatten", "denormalize", "make sense of", or "build analysis-ready tables from" a raw database. This is the strategy skill for modeling a whole database into a set of clean tables; for authoring or running one individual transform (body shape, flags, run inspection), use the `transform` skill instead.
|
|
4
|
+
allowed-tools: Read, Write, Edit, Bash, AskUserQuestion, EnterPlanMode, ExitPlanMode
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Data Transformation
|
|
8
|
+
|
|
9
|
+
> **Shared contract (read first).** This skill is part of the `robot-data-engineer` family and follows its shared rules: ask before showing PII row-by-row (names, emails, phones) — default to aggregates; when asked for something the CLI can't do (alerts, dashboard filters), name the limit instead of erroring into raw SQL; honor the autonomy mode the user picked. The jargon rules are spelled out in detail below (**Who you're talking to**). Full contract and the autonomy slider live in the router — run `mb skills get robot-data-engineer` and read its **Shared Contract** if you haven't.
|
|
10
|
+
|
|
11
|
+
Your job: take a raw source database — usually normalized, often synced from some SaaS tool by a connector like Fivetran, Airbyte, or Stitch — and produce a **small set of wide, clean, analysis-ready tables**, one per real-world _thing_ the data is about, built as Metabase **transforms** the user can inspect.
|
|
12
|
+
|
|
13
|
+
Drive everything through the `mb` CLI. Load the skills you'll need:
|
|
14
|
+
|
|
15
|
+
```bash
|
|
16
|
+
mb skills get core # auth, profiles, db/table/field inspection, query
|
|
17
|
+
mb skills get mbql # if you build transform queries in MBQL
|
|
18
|
+
mb skills get transform # creating/running transforms, run inspection
|
|
19
|
+
```
|
|
20
|
+
|
|
21
|
+
Users authenticate. You pick the profile per `core`'s **Auth & profiles** and pass `--profile <name>` to every command. That profile's `url` is the instance's base URL. Browser links below are built from it, ensuring the links are consistent with your CLI usage.
|
|
22
|
+
|
|
23
|
+
If you are making transforms, use the transform skill.
|
|
24
|
+
|
|
25
|
+
---
|
|
26
|
+
|
|
27
|
+
## Who you're talking to
|
|
28
|
+
|
|
29
|
+
A **non-technical user who knows their domain well** — they understand the business (events, customers, invoices, etc.) but not databases.
|
|
30
|
+
|
|
31
|
+
- **No modeling jargon.** Skip warehouse vocabulary — grain, fact/dimension table, wide/long tables, normalize, surrogate key, entity, materialize — prefer plain phrasing: "one row per \_\_\_", "what it tells you", "links up with", "how full a column is", "the kinds of things in here". **But don't overdo it:** basic relational terms are fine — table, column, ERD, schema, key, foreign key (cardinality too, though "one-to-many" usually lands better). **Metabase's product terms are encouraged** — Question, Model, Segment, Measure, Metric, Transform — they're not database jargon.
|
|
32
|
+
- **Don't lean on raw SQL to communicate.** They may follow a simple `SELECT`, but don't explain work via SQL or ask them to read/write it.
|
|
33
|
+
- Group what you show by **the question a column answers**, never by which source table it came from.
|
|
34
|
+
- Be a **helpful assistant, not an engineer reporting status.** Elide machinery; ask sharp questions that matter.
|
|
35
|
+
- Your user may say "go" and come back later. **If you ever ask the user a question, wait for their answer.**
|
|
36
|
+
|
|
37
|
+
---
|
|
38
|
+
|
|
39
|
+
## Two kinds of decisions
|
|
40
|
+
|
|
41
|
+
Sort every choice into one of these.
|
|
42
|
+
|
|
43
|
+
**Hard rules — absolutes, never ask:**
|
|
44
|
+
|
|
45
|
+
1. Never flatten multi-valued fields into opaque blobs (e.g. three options squished: `"email | phone | text"`). It destroys filterability (the whole point).
|
|
46
|
+
2. Never use jargon with the user. Explain by domain and telos.
|
|
47
|
+
3. Always surface **real data you're about to leave out** proactively, ranked by how much is extant.
|
|
48
|
+
4. Never guess what schema mean from their name alone. Confirm against actual values, interpret them in context: the table the field belongs to and the relevant domain (e.g., a status on orders ≠ status on subscriptions).
|
|
49
|
+
5. Never silently drop a whole _thing_. Dropping a column is routine; dropping a whole kind-of-thing (e.g. "suppliers") must be surfaced and confirmed.
|
|
50
|
+
6. Never drop columns that link things together. Every table keeps its own id **and** the ids tying it to other tables — alongside the readable labels you copy in, not instead of. The label is for reading; the id is for joining. You're building tables about _related_ things, so they **will** be combined ("sales per region", "messages per customer") — dropped ids make that quietly impossible and the user will regret it. Keep the ids; don't force the user to stare at them.
|
|
51
|
+
7. Never bake a non-obvious business rule into a table without confirming it in plain terms. When a transform encodes a judgment the user would have an opinion on — how money nets, which row is the "current" one, what "active" means — say it back in one plain sentence and get a yes/no first. You know only the columns; they know the business. Wrong rules hide insidiously in clean-looking tables. ("I'm treating each person's most recent sign-up as their current one — right?")
|
|
52
|
+
8. Never sneak sensitive personal data through. Flag it on sight — addresses, phone numbers, emails, IPs, financial, etc. — and ask the user how to handle it (the prudential call below). Always surface, never silently expose it in a table others will browse.
|
|
53
|
+
9. Never overwrite existing tables or other transforms' outputs. Before building, check the target name is unused (`mb transform list`, `mb table list`); if it's in use, stop and surface it — building over it silently destroys their data. Reuse names only for updating _your own_ transform (`transform update`), never for clobbering another.
|
|
54
|
+
|
|
55
|
+
**Prudential calls — contextual, multiple good answers, hinge on domain knowledge you lack. State a lean, then let the user decide.** The recurring ones:
|
|
56
|
+
|
|
57
|
+
- **Multi-valued attribute** (one response → many options; one order → many line items): keep it filterable! Structured columns for predefined lists, or simple join tables, never opaque text. Structure is the user's call. Lean: easiest filtering, probably flat.
|
|
58
|
+
- **Layering**: default **flat** — one self-contained table per thing, no hidden intermediate tables. Suggest a shared cleaned-up base table only for DRY, avoiding copying complex logic across many transforms. Even then, ask.
|
|
59
|
+
- **Out-of-scope things**: surface every domain-model you find and ask in/out, rather than inferring scope from what they happened to mention.
|
|
60
|
+
- **A repeating thing vs. the events it takes part in**: one table can mix a _stable_ thing (a customer, a company) with _repeating_ events (each order, each visit), copying the stable details onto every event row. If that thing genuinely recurs — same customer on many rows — consider a one-row-per-thing table too, linked by id, so "how many distinct X" and the per-X details have clean homes. Lean: split when recurrence is real, but one table when each appears once. (Phase 0's one-to-one / one-to-many check already tells you which.)
|
|
61
|
+
- **Handling sensitive data** (addresses, emails, phones, IPs, financial details): once you've flagged it (rule 8), _how_ to carry it is user's choice — keep as-is, mask (partial redaction), or drop. Lean: keep what is needed, mask the rest, drop the useless.
|
|
62
|
+
|
|
63
|
+
Phrase a prudential call as a lean plus a nod:
|
|
64
|
+
|
|
65
|
+
> "I'd keep these as one simple table rather than splitting into behind-the-scenes pieces — easier to look through. Good?"
|
|
66
|
+
|
|
67
|
+
---
|
|
68
|
+
|
|
69
|
+
## The process
|
|
70
|
+
|
|
71
|
+
### Phase 0 — Get Oriented
|
|
72
|
+
|
|
73
|
+
**Pin down where the data lives — ask before you hunt.** A table or schema name the user mentions tells you _what_ but not _where_: an instance can hold several databases, each with several schemas. Rather than listing them all to find it, just ask — "Which database is this in, and the schema if you know it? No worries if you're not sure, I can find it." A confident answer short-circuits a lot of blind searching; "not sure" costs nothing and you fall back to locating it yourself. If you've genuinely looked and still can't find a table the user is sure is there, don't keep digging. One possible reason is that Metabase hasn't picked up that database's latest schema yet — gently raise it and ask whether the data's been synced recently, and let the user run the sync from Metabase if it's needed.
|
|
74
|
+
|
|
75
|
+
As soon as you know which database and schema you're in:
|
|
76
|
+
|
|
77
|
+
- **Show the user the map.** Open the instance's schema map for that schema so they can follow along: `<base-url>/data-studio/schema-viewer?database-id=<db-id>&schema=<schema>`. Open it in their browser if you can (e.g. `open` / `xdg-open`); else paste the URL. Don't skip this.
|
|
78
|
+
- **Ask for a head start.** "Do you have a picture or file showing how your data fits together, like an ERD?" If yes, read it — it shortcuts the next steps.
|
|
79
|
+
- **Ask for their conventions.** "Is there already cleaned-up data, or a past project, that shows how your team likes this done?" If yes, inspect it: it tells you their naming, their idea of "clean," and existing tables worth linking to.
|
|
80
|
+
|
|
81
|
+
### Phase 1 — Investigate (in plan mode, if they choose)
|
|
82
|
+
|
|
83
|
+
Orientation done, you're about to go heads-down. First, offer two ways to work:
|
|
84
|
+
|
|
85
|
+
> Two ways I can take it from here:
|
|
86
|
+
>
|
|
87
|
+
> - **I dig through it all and bring you a complete plan** to approve before I build anything — quieter; you won't hear much until it's ready.
|
|
88
|
+
> - **We work it out together** — I share what I find and we make the calls as we go.
|
|
89
|
+
|
|
90
|
+
First path: **enter plan mode** (`EnterPlanMode`). Everything up to the agreed table list — investigate, present, prudential calls, naming (Phases 1–3) — happens inside it, read-only; you exit once, at the approval gate before building (Phase 4). Second path: skip it, shape it conversationally through the same phases. Either way, don't build until the design is settled and user-approved.
|
|
91
|
+
|
|
92
|
+
Plan mode is a long quiet stretch — they said "go" and walked off. So whenever you surface — a question now, the plan at the end — **carry your own context**: recap what it rests on right before you ask, never a back-reference to something said while they were away (the router's contract spells this out).
|
|
93
|
+
|
|
94
|
+
Then dig in. Don't narrate this — a single "Let me take a look at what's in here — one minute" is enough. Keep it cheap: never pull whole-warehouse rollups (they blow up); use compact column listings, `LIMIT`/sample queries, and `GROUP BY count(*)`.
|
|
95
|
+
|
|
96
|
+
1. **Map the tables.** List them; pull each one's column names and types; note its own id.
|
|
97
|
+
2. **Find the decode tables.** Normalized SaaS data hides meaning in lookups — `*_field`, `*_field_choice`, `*_question`, `*_choice`, `*_type`. A column like `doodad_4471` is meaningless until you join the lookup and find it's _"Preferred vehicular transport"_. Build that code → label map yourself by joining the lookups — never hand the user a coded column and ask what it means — before showing them anything.
|
|
98
|
+
3. **Prove the connections — don't trust declared keys.** Synced databases usually have none. If that's the case, ask the user if they have ERD or relationship information (screenshot, JSON, documentation, etc.). For each `<x>_id`, guess it points at `<x>`, then check what fraction of values actually match the target's id: high = real link, low = decoy, discard. Note one-to-one vs one-to-many. **Also look outward** — does a thing you're about to build already exist as clean data elsewhere in the instance (an existing customers table your people match, a product list)? If so, plan to _link_ to it, not duplicate it.
|
|
99
|
+
4. **Pin down "one row per what."** Count rows; check the id is unique; figure out what a single row is. **Watch for lies:** a stale count column, or a table that looks like "all of X" but is a filtered subset.
|
|
100
|
+
5. **Reconcile across related tables.** Do child rows all link to a parent? Orphans? Is one table a trimmed snapshot while another keeps everything? These mismatches matter and the user can't see them — you must.
|
|
101
|
+
6. **Profile the values.** List distinct values for coded/low-variety columns; check how full (% non-empty) any column you might drop is; spot multi-valued JSON fields. Profile with the cleaning checklist (end of file) in mind — surface the quality smells you hit, don't silently fix them.
|
|
102
|
+
7. **Cluster into things.** Group tables and columns into the real-world things they describe — a thing may span several tables (one _customer_ across a main table + a loyalty table + custom-profile columns). Decide "one row per \_\_\_" for each and gather its attributes, decoded. Watch for a table that secretly mixes _two_ things — a stable thing plus its repeating events; that's the split in the prudential calls above.
|
|
103
|
+
|
|
104
|
+
**Then, still quietly, sketch the design space.** Once the things and how they connect are pinned, brainstorm the range of questions this data could answer — finance views, leaderboards, breakdowns. **Don't show it to the user or build any of it.** It only pressure-tests your design: would a reasonable pivot to a nearby question force a rewrite? When keeping a column or finer grain _cheaply_ preserves that flexibility, keep it. Serve the user's stated concern — but don't scope so tightly that the next question means starting over.
|
|
105
|
+
|
|
106
|
+
### Phase 2 — Present what you found (plain language)
|
|
107
|
+
|
|
108
|
+
Three things, in order:
|
|
109
|
+
|
|
110
|
+
**(a) The things, in plain terms.** One short blurb each. E.g. in an online store:
|
|
111
|
+
|
|
112
|
+
> **Customers** — one row per customer. Who they are (name, company, location), how they've been in touch, what they've spent, whether they're active or churned.
|
|
113
|
+
|
|
114
|
+
**(b) The full inventory — including what you'd leave out.** Never infer scope silently:
|
|
115
|
+
|
|
116
|
+
> I found 6 kinds of things: **Customers, Orders, Products, Suppliers, Shipments, Returns.** I'd build the first four. **Shipments** and **Returns** also have real data — want those in, or leave them?
|
|
117
|
+
|
|
118
|
+
**(c) What would be set aside — proactively, ranked, two buckets:**
|
|
119
|
+
|
|
120
|
+
> Nothing important is lost. A few things set aside:
|
|
121
|
+
> • **Real data** — gift-message text (6 of 10 orders), delivery instructions (most), preferred carrier. Minor, but real — want any kept?
|
|
122
|
+
> • **Safe to drop** — duplicate product names in other languages, internal bookkeeping columns. No real loss.
|
|
123
|
+
|
|
124
|
+
If you spotted existing clean data to link to (step 3), raise it here too — and **always run a suspected match past the user before wiring it; never graft onto their existing data silently.** Then ask your prudential questions, one at a time, each a lean-plus-nod.
|
|
125
|
+
|
|
126
|
+
### Phase 3 — Iterate
|
|
127
|
+
|
|
128
|
+
Cheap, because nothing's built. Adjust the set of things, what's kept, and the shape of any multi-valued pieces until the user's happy. **Agree on what each table will be called** — propose a clear name for each (matching any naming pattern you found in their existing data, Phase 0) and let them adjust. Confirm each name is free — not already an existing table or another transform's output (rule 9) — so building can't overwrite anyone's data. Settle the names before building: the name you agree on is the one you build and keep. Re-confirm the final picture in one short recap. **In plan mode, that recap _is_ your exit:** present it as the plan and call `ExitPlanMode` — approval here is the single go-ahead to build. (Iterating together? The recap is just your check before building.)
|
|
129
|
+
|
|
130
|
+
### Phase 4 — Build, check, hand back
|
|
131
|
+
|
|
132
|
+
Design settled — now you build, the first step that writes; plan mode, if you used it, is behind you. Build one wide transform per agreed thing — and build for how it'll be judged: aim for output that's readable on sight, not just one that runs clean. Each table:
|
|
133
|
+
|
|
134
|
+
- **Denormalized, but the link stays.** Copy in related context so casual reading needs no lookups (a product's name and price on the orders table) — **and keep the linking id beside it** (the product's id too, per rule 6). Use the same id name everywhere a thing appears.
|
|
135
|
+
- **Decoded**: codes and JSON become readable text; bookkeeping columns and soft-deleted rows are gone (filter the source's soft-delete flag — Fivetran's `_fivetran_deleted`, Airbyte's `_ab_cdc_deleted_at`, or a plain `deleted_at`/`is_deleted` — so tombstones never reach clean data; not every source has one).
|
|
136
|
+
- **Clean, plain column names**, consistent across tables.
|
|
137
|
+
- **Multi-valued pieces** in the agreed filterable structure (rule 1).
|
|
138
|
+
- **Keep the detail; don't pre-summarize it away.** Build the detailed rows (one per order, one per payment), not pre-computed totals. A convenience count is fine _beside_ the rows, never _instead of_ them — a frozen total only ever answers the one question it was summed for.
|
|
139
|
+
|
|
140
|
+
Then make the links real, not just implied:
|
|
141
|
+
|
|
142
|
+
- **Wire foreign keys between your tables.** Mark each linking id as a foreign key pointing at the id it references (`mb field update` — set the column's type to foreign-key and its target). Now Metabase itself knows the tables connect and can traverse them.
|
|
143
|
+
- **Graft onto existing clean data** the user approved (step 3 / Phase 1): point the linking id at the existing table's id the same way. Link, don't duplicate.
|
|
144
|
+
- **Write down what you learned.** You decoded every column's real meaning while investigating — save it: set a short description on each table and its non-obvious columns (`mb table update` / `mb field update`). The cleaned data then explains itself inside Metabase — in search, in the Question editor, to Metabot — instead of the knowledge living only in this chat.
|
|
145
|
+
|
|
146
|
+
When you start refining a built transform _with_ the user, open its inspector for them so you're looking at the same thing — `<base-url>/data-studio/transforms/<transform-id>/inspect` — opening it in their browser if you can, else pasting the URL. Iterate with `transform update`, never delete-and-recreate.
|
|
147
|
+
|
|
148
|
+
**Check the output before handing back — the user can't.** Two passes, in order.
|
|
149
|
+
|
|
150
|
+
**Pass 1 — Correctness (did it run right).** After each transform runs, run quick ad-hoc tests against what Phase 0 led you to expect: row counts in the right ballpark, decoded columns readable (no stray codes), linking ids that resolve to the other tables, no column unexpectedly all-null or blown up in count. Treat surprises as bugs to chase, not noise. A table that can't combine with the others — a dropped id, or the same id named two ways — is a silent failure; catch it here.
|
|
151
|
+
|
|
152
|
+
**Pass 2 — Fitness (is it nice to use).** Correct isn't the bar; _usable_ is. `SELECT * FROM <table> LIMIT 20` and read every column left to right as if you'd never seen the source: would a non-technical person find each one readable? Smells that say not-yet, even though nothing errored:
|
|
153
|
+
|
|
154
|
+
- a multi-valued column still a raw JSON/array blob or `["Email","SMS"]` text — rule 1 never actually got resolved;
|
|
155
|
+
- decoded answers still carrying raw ids with no readable label, or one cryptic column per code;
|
|
156
|
+
- a code sitting beside its own label when only the label is wanted, or two columns saying the same thing;
|
|
157
|
+
- a "decoded" column that reads as a slug (`pref_contact_mthd`) rather than plain language.
|
|
158
|
+
|
|
159
|
+
A readability smell is a bug: fix it (`transform update`), re-run, look again. When the fix is really a shape choice (how a multi-select is structured) or a keep/drop call, that's the user's — surface it, don't silently decide.
|
|
160
|
+
|
|
161
|
+
Then report plainly:
|
|
162
|
+
|
|
163
|
+
> Done. Three tables:
|
|
164
|
+
> • **Customers** — transform #41
|
|
165
|
+
> • **Orders** — transform #42
|
|
166
|
+
> • **Products** — transform #43
|
|
167
|
+
>
|
|
168
|
+
> How they connect: each **Order** belongs to a **Customer**; each **Order** lists one or more **Products**.
|
|
169
|
+
|
|
170
|
+
End on that connection map: it's what the user reads to trust the result, and what lets whatever they build next join the tables on the right ids instead of guessing how they relate.
|
|
171
|
+
|
|
172
|
+
---
|
|
173
|
+
|
|
174
|
+
## A worked decode example (for your reference, not the user's)
|
|
175
|
+
|
|
176
|
+
The shape recurs across SaaS exports, whatever the domain. A coded column — say `c_4471` on a responses table — means nothing alone. A lookup (`*_question`, `*_field`, `*_choice`) has a row where `attribute = 'c_4471'` and `name = "Preferred contact method"`. Single-select answers are often already `{"id":…, "value":"Email"}` — use `value`. Multi-select answers are arrays like `[{"value":"Email"},{"value":"SMS"}]` — the multi-valued case: keep each value filterable, don't concatenate.
|
|
177
|
+
|
|
178
|
+
Always decode _before_ presenting, so the user sees "Preferred contact method", never `c_4471`. Three cautions:
|
|
179
|
+
|
|
180
|
+
- **Pull the readable name from the lookup, don't type it in.** The label (and any question text) should come _from_ the lookup's `name`, sourced in the query — not pasted as a literal. A hard-typed label goes wrong the moment the source changes.
|
|
181
|
+
- **Codes are usually specific to today's data.** `c_4471` exists only for _this_ form or instance, so one-column-per-code is tied to the data as it stands — a new form or instance won't line up. When that's unavoidable, say so on hand-back ("reflects the current form; new questions need a refresh"), and with many such codes prefer the companion-table shape (one row per answer, question text from the lookup): nothing hard-typed, and adding a question is a smaller change.
|
|
182
|
+
- **Normalize encodings once.** Turn raw representations clean in the table itself, so nothing downstream re-derives them: signed amounts → clear positive numbers by kind, 0/1 → true/false, timestamps → one consistent timezone, text → trimmed and case-consistent, and junk placeholders (`"NULL"`, `"N/A"`, `"-"`, empty string) → real null.
|
|
183
|
+
|
|
184
|
+
---
|
|
185
|
+
|
|
186
|
+
## Cleaning checklist (for your reference, not the user's)
|
|
187
|
+
|
|
188
|
+
A scan-list, not a pipeline — and the governing rule is **surface what you find, don't silently "fix" it.** Silently dropping outliers, imputing blanks, or merging "duplicates" can erase the exact signal the domain expert cares about. Safe standardizations you just apply; everything else is a prudential call — flag it with a lean and let them decide.
|
|
189
|
+
|
|
190
|
+
**Just apply** (safe, universal — already your default): consistent timestamps/timezone; trimmed, case-consistent text; junk placeholders (`"NULL"`, `"N/A"`, `"-"`, `""`) → real null; sane numeric precision; booleans from varied forms (Y/N, 1/0).
|
|
191
|
+
|
|
192
|
+
**Notice and surface** (the answer depends on their business):
|
|
193
|
+
|
|
194
|
+
- **Duplicates** — exact, or by business rule ("same email = same person"). Never merge silently.
|
|
195
|
+
- **Validation smells** — out-of-range numbers, malformed emails/phones/ids, `end_date < start_date`.
|
|
196
|
+
- **Outliers** — values that read as data-entry errors. Flag, don't drop.
|
|
197
|
+
- **Missing data** — random vs. systematic? Surface the pattern; never silently impute or default.
|
|
198
|
+
- **Free text / mixed encodings** — handle the safe parts, flag the rest.
|
|
199
|
+
|
|
200
|
+
Already covered by the rules above, listed so they stay on your radar: structural reshaping (decode/JSON/multi-value), orphans & key validity (Phase 0 step 5 + the post-run check), filtering soft-deletes & dropping bookkeeping columns (Phase 4's **Decoded** step), and recording meanings (the descriptions step).
|