@graphit/cli 0.2.321 → 0.2.323

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (57) hide show
  1. package/.claude-plugin/marketplace.json +3 -3
  2. package/.claude-plugin/plugin.json +1 -1
  3. package/.codex-plugin/plugin.json +1 -1
  4. package/bin/graphit +1 -1
  5. package/bin/graphit.ps1 +1 -1
  6. package/dist/api/client.js +15 -0
  7. package/dist/api/client.js.map +1 -1
  8. package/dist/commands/ds-config.js +2 -2
  9. package/dist/commands/ds-config.js.map +1 -1
  10. package/dist/commands/ds.js +2 -2
  11. package/dist/commands/ds.js.map +1 -1
  12. package/dist/commands/kb.js +347 -15
  13. package/dist/commands/kb.js.map +1 -1
  14. package/dist/commands/query.js +3 -3
  15. package/dist/commands/query.js.map +1 -1
  16. package/dist/skill-guard.js +2 -1
  17. package/dist/skill-guard.js.map +1 -1
  18. package/package.json +1 -1
  19. package/scripts/verb-policy-source.json +30 -102
  20. package/skills/graphit/SKILL.md +31 -39
  21. package/skills/graphit/VERSION.json +1 -1
  22. package/skills/graphit/references/data-source-refresh.md +27 -0
  23. package/skills/graphit/references/data-sources.md +27 -122
  24. package/skills/graphit/references/filters-advanced.md +2 -2
  25. package/skills/graphit/references/governance-explained.md +18 -29
  26. package/skills/graphit/references/governance.md +20 -83
  27. package/skills/graphit/references/kb-actions.md +28 -81
  28. package/skills/graphit/references/kb-discovery.md +35 -64
  29. package/skills/graphit/references/kb-scope.md +17 -15
  30. package/skills/graphit/references/kb-structure.md +28 -54
  31. package/skills/graphit/references/kb-traversal.md +24 -96
  32. package/skills/graphit/references/metric-families.md +19 -0
  33. package/skills/graphit/references/migration.md +2 -2
  34. package/skills/graphit/references/onboarding.md +5 -4
  35. package/skills/graphit/references/presentations.md +1 -1
  36. package/skills/graphit/references/runtime.md +6 -6
  37. package/skills/graphit/references/semantic-authoring.md +65 -0
  38. package/skills/graphit/references/sql-reference.md +10 -20
  39. package/dist/commands/kb-constraints.d.ts +0 -14
  40. package/dist/commands/kb-constraints.js +0 -53
  41. package/dist/commands/kb-constraints.js.map +0 -1
  42. package/dist/commands/kb-create.d.ts +0 -2
  43. package/dist/commands/kb-create.js +0 -296
  44. package/dist/commands/kb-create.js.map +0 -1
  45. package/dist/commands/kb-delete.d.ts +0 -2
  46. package/dist/commands/kb-delete.js +0 -37
  47. package/dist/commands/kb-delete.js.map +0 -1
  48. package/dist/commands/kb-read.d.ts +0 -2
  49. package/dist/commands/kb-read.js +0 -223
  50. package/dist/commands/kb-read.js.map +0 -1
  51. package/dist/commands/kb-shared.d.ts +0 -43
  52. package/dist/commands/kb-shared.js +0 -81
  53. package/dist/commands/kb-shared.js.map +0 -1
  54. package/dist/commands/kb-update.d.ts +0 -2
  55. package/dist/commands/kb-update.js +0 -240
  56. package/dist/commands/kb-update.js.map +0 -1
  57. package/skills/graphit/references/parameterized-metrics.md +0 -77
@@ -0,0 +1,27 @@
1
+ # Data-Source Refresh
2
+
3
+ Load when configuring refresh mode, incremental behavior, history, or reconciliation.
4
+
5
+ ## Modes
6
+
7
+ - **full:** replace the cached result from the complete source query.
8
+ - **incremental:** append/merge new source rows using a watermark and stable merge key.
9
+
10
+ Choose incremental only when the source exposes a reliable monotonic watermark and the merge key is unique. Otherwise use full refresh.
11
+
12
+ ## Incremental contract
13
+
14
+ - Filter the source early on the watermark column.
15
+ - Preserve lookback for late-arriving updates.
16
+ - Provide the correct watermark type.
17
+ - Verify merge-key uniqueness before serving the new version.
18
+ - Reconcile periodically with a full rebuild.
19
+ - Treat schema drift or grain change as a semantic review, not a blind refresh.
20
+
21
+ ## Operations
22
+
23
+ A refresh request may be asynchronous. Poll the job/source status and report counts, duration, version, and failure truthfully. Fire-and-forget is appropriate only when the user did not ask to wait.
24
+
25
+ Inspect refresh history before retrying. A failed status may follow a partially applied external action; use the receipt/status rather than assuming nothing happened.
26
+
27
+ Refresh settings and connector lifecycle require data-source write authority in the server's policy key. Never expose credentials, source rows, or concealed schema in evidence.
@@ -1,134 +1,39 @@
1
- # Data Sources: Routing and Building for Speed
1
+ # Data Sources
2
2
 
3
- A data source's shape - set by its source SQL at creation - decides how fast every dashboard built on it will be, and whether filter changes feel instant. Get this right when you create the source; it can't be fixed later in the dashboard SQL.
3
+ Load when selecting or creating the cached source a semantic model uses.
4
4
 
5
- ## Routing: which source to query
5
+ ## Routing
6
6
 
7
- Always prefer a cached data source over the live warehouse. Check what exists with `graphit ds list` before writing any query.
7
+ 1. Read the semantic model's declared data-source binding.
8
+ 2. Prefer that cached source for speed, governance, and repeatability.
9
+ 3. Use metadata discovery when physical columns are unknown.
10
+ 4. Query live warehouse only when no cached source covers the question and the user approves.
11
+ 5. Never infer a source from a similarly named model.
8
12
 
9
- | Situation | Command | Speed |
10
- |---|---|---|
11
- | The table has a cached data source | `graphit query "SQL" --ds <NAME>` | roughly 100ms, DuckDB |
12
- | No data source covers the table | `graphit query "SQL" --warehouse --connection <id>` | roughly 10s, the connected warehouse |
13
+ A group is semantic placement. Data-source creation still accepts `--domain`; pass the uppercase policy key returned by status or the group's `domain_keys`.
13
14
 
14
- `--ds` takes the data source **name** (the same name you SELECT FROM, e.g. `... FROM MARKETING_UA_DS ... --ds MARKETING_UA_DS`); a full id or unique id-prefix also resolves.
15
+ ## Source SQL
15
16
 
16
- If no data source covers the table the user needs, propose creating one for future speed rather than defaulting to repeated warehouse queries. Dialect differs by route: DuckDB for `--ds`, and the connected warehouse's own dialect (Snowflake or BigQuery) for `--warehouse`; see `sql-reference.md`.
17
+ - Select only needed columns and rows.
18
+ - Filter early using base-table columns.
19
+ - Avoid wrapping filter columns when a direct predicate works.
20
+ - Keep complete executable SQL; no ellipses, fake tables, or embedded data.
21
+ - Preserve warehouse dialect.
22
+ - Make grain and refresh mode explicit.
23
+ - Use merge key and watermark only when the source supports them.
17
24
 
18
- ## Build it right (in the source SQL)
25
+ ## Creation
19
26
 
20
- Filter changes answer from a cached, pre-aggregated result only when the source is small and aggregated to the grain you query; pull in raw rows, hundreds of columns, or thousands-of-values dimensions and it is too big to cache, so every filter change re-scans the whole source and is slow. Set the shape at creation on these four levers:
27
+ Confirm connector, relation/query, policy key, grain, refresh mode, and cost. Read columns through metadata rather than probing with ad-hoc SQL.
21
28
 
22
- | Lever | Build it right | Anti-pattern |
23
- |---|---|---|
24
- | Grain | `GROUP BY` to the grain you chart (the single biggest lever) | One row per raw event |
25
- | Columns | Only the columns dashboards use (each is downloaded + cached per query) | 400+ columns "just in case" |
26
- | Cardinality | Low-card dimensions in the base; ad/campaign names in a separate drill-down | Thousands-of-values dimensions in the base grain |
27
- | Size | A few-thousand-row typical aggregation | Sitting at the 100M-row / 5GB ceiling |
29
+ Create with automatic scan unless there is a specific reason not to. Creation may be asynchronous; report `creating` honestly and poll status rather than claiming readiness.
28
30
 
29
- ## Slow-shape signals to watch for
31
+ ## Access and safety
30
32
 
31
- Before creating a source, check your own SQL for these. Each is a reason to offer the user a faster shape, never to refuse:
33
+ - Changing a source requires data-source write capability in its policy key.
34
+ - Reading does not imply authority over connector, SQL, or refresh settings.
35
+ - Visibility and masking cover agent, canvas, render, export, and report paths.
36
+ - Private names and columns remain concealed.
37
+ - Delete/move stay in the Sources Hub where cascades are visible.
32
38
 
33
- | Signal | What it looks like | Faster option |
34
- |---|---|---|
35
- | Raw passthrough | No `GROUP BY` / no aggregate - one row per raw event | Pre-aggregate to the grain the dashboards chart |
36
- | Very wide | Far more columns than dashboards use (e.g. `SELECT *`) | Select only the columns dashboards need |
37
- | High-cardinality grain | A dimension with thousands of distinct values (ad / campaign / user ids) | Keep it out of the base; build a separate drill-down source |
38
- | Large + monolithic | A big source whose typical query still scans most rows | Pre-aggregate and narrow so typical queries touch a few thousand rows |
39
-
40
- A wide or raw source is sometimes the right call - row-level drill-down/export, columns genuinely all used, or a staging source to reshape later. Note the trade-off, then build whichever the user chooses.
41
-
42
- ## Creating data sources
43
-
44
- `graphit ds create` auto-chains: create -> poll until ready -> scan schema -> print verification link. Activating it for KB use is the `ds verify` step below.
45
-
46
- ```bash
47
- # Create with auto-scan (recommended)
48
- graphit ds create --name "MY_DS" --domain <DOMAIN> --sql "SELECT ..." --connection <id>
49
-
50
- # Create without auto-scan (for special cases)
51
- graphit ds create --name "MY_DS" --domain <DOMAIN> --sql "SELECT ..." --skip-scan
52
- ```
53
-
54
- **`--domain` is REQUIRED, in both modes.** It files the scanned table under an existing KB domain and decides who can see the source; there is no uncategorized fallback. `graphit kb list domains`, confirm the choice with the user, and `graphit kb create domain --name <NAME>` if none fits.
55
-
56
- **From a local file (Excel/CSV):** `graphit ds create --file <path> --domain <NAME>` uploads the file and creates one data source. Optional: `--name` (defaults to the file name), `--sheet <name>` (multi-sheet workbooks). `--file` and `--sql` are mutually exclusive; same flow as above.
57
-
58
- **Warehouse connection.** `--connection` names the warehouse a `--sql` source reads from. Add BigQuery with `graphit connector add bigquery-serviceaccount --key-file <path> [--project --dataset --location]` (org admin; project defaults from the key). The pipeline routes by connection type - the same `ds create` works for either warehouse.
59
-
60
- For existing unverified sources, `graphit ds verify <id>` scans and shows the schema; add `--accept-schema` to accept the AI schema and activate a warehouse/SQL source from the CLI. A file upload needs `ds verify` too - it activates without `--accept-schema`, but never at create time, so it stays unqueryable until you run it.
61
-
62
- ## Refreshing data sources
63
-
64
- Data sources cache a snapshot of the warehouse query result. Refresh when you need current data. **File-upload sources can't be refreshed - update them by re-uploading with `graphit ds create --file <path> --domain <NAME>`.**
65
-
66
- On BigQuery a refresh scans billed bytes, so keep the shape tight and prefer incremental/partition-pruned refresh over full re-scans; a per-connection scan cap (max bytes billed) fails an oversized query fast rather than running up a bill.
67
-
68
- ```bash
69
- # Refresh all data sources and wait for completion (live status table)
70
- graphit ds refresh --all
71
-
72
- # Fire-and-forget (trigger refreshes, don't wait)
73
- graphit ds refresh --all --no-wait
74
-
75
- # Refresh specific sources by ID
76
- graphit ds refresh <id1> <id2>
77
- ```
78
-
79
- `graphit ds refresh` only runs an **incremental** refresh (new rows since last update); it never re-exports the whole source. A full rebuild is UI-only (Sources -> Refresh -> Full rebuild); hand off to the UI if one is needed.
80
-
81
- Refreshes fire in parallel; polls to completion (large sources 30-60s), or returns at once with `--no-wait` (check status via `graphit ds list`). Governed per org: manual refreshes have an hourly budget and a limited number run at once (a reserved slot keeps manual ones unblocked by scheduled). A limit returns a 429 with a clear reason (reset time, or "wait for running operations to finish"); wait it out, don't loop - not source errors. Review past runs with `graphit ds refresh-history <id>`.
82
-
83
- ## Incremental refresh and early-filtering (advanced)
84
-
85
- Incremental mode fetches only rows past a watermark and merges them in. Three windows govern it: the **watermark column** (which output rows are new), the **merge window** (`--merge-window` - how far back each run re-fetches and upserts, healing late data; API responses call it `lookback_periods`), and per-table **lookback windows** (`--table-lookback` - how far back each source table is *read*). Set on a scanned source; each call sets the COMPLETE config - omitted flags reset to defaults (no `--table-lookback` = windows cleared).
86
-
87
- When a source aggregates over a wide internal window (e.g. a multi-year rollup), incremental refresh is nearly as slow as full: the outer watermark filter can't prune the inner scan. Early-filtering fixes that - get the contract right first, or older periods silently corrupt on merge:
88
-
89
- - Size each lookback to cover the merge window plus the longest rolling calculation in the query.
90
- - Set `--merge-key` (upsert) - overlapping rows double-count without one.
91
- - Rolling-window metrics (WAU / MAU / stickiness): prefer a layered base daily source instead - recent rows alone can't compute a rolling window.
92
- - `--reconciliation` is the periodic full-refresh drift backstop (default off; keep it on with the bind or in append mode).
93
-
94
- Then pick ONE early-filter mode (mutually exclusive, validated):
95
-
96
- **Per-table lookback windows - preferred; no SQL edit.** Declare how far back each table is read; the engine prunes delta scans and the SQL stays exactly as written. Day-based (date/timestamp watermark required).
97
-
98
- ```bash
99
- graphit ds refresh-config <id> --mode incremental \
100
- --watermark-column EVENT_DATE --watermark-type date --merge-key ID \
101
- --merge-window 3 --table-lookback ANALYTICS.EVENTS:EVENT_DATE:30 \
102
- --table-lookback USERS:CREATED_AT:90
103
- ```
104
-
105
- **`:graphit_watermark` bind - for SQL owners.** Place the token where the filter belongs; Graphit substitutes the last watermark on deltas, full history on reconciliation. Output-filter to the current fully-covered period.
106
-
107
- Only early-filter when an incremental source is slow for this reason; the default refresh is correct and simpler otherwise.
108
-
109
- ## What needs write access
110
-
111
- Querying a source, listing sources, reading schema or refresh history, and an ordinary `graphit ds refresh` are reads - any member who can read that source's domain can run them, and a source in a domain they cannot read returns the same uniform 404 as one that does not exist. These need `data_source_write` in the source's domain: `ds create`, editing its SQL, `ds refresh-config`, a `--force` refresh or accepting a schema, `ds verify`, scanning, per-source governance settings, and deletion. Moving a source to another domain needs write in both the old and the new domain.
112
-
113
- Check `graphit status` for those domains before proposing a create or a config change. It is advisory - the server authorizes each operation when it runs, and a denial with `retryable: false` is a stop, not a retry (`operations.md`).
114
-
115
- ## Deleting data sources
116
-
117
- `ds delete` is not available on the CLI. Deleting a data source cascades to storage and the KB table, removing all metrics, dimensions, and rules on it. Direct the user to the platform UI (Sources Hub), whose confirmation flow shows what will be affected.
118
-
119
- ## Presenting data source results
120
-
121
- The user cannot see raw CLI output - you are the rendering layer. After `graphit ds list`, present a markdown table and end with a recommendation of which source to use (or note none covers the needed table):
122
-
123
- ~~~
124
- **2 data sources:**
125
-
126
- | Name | ID | Rows | Status | Governed |
127
- |---|---|---:|---|---|
128
- | **MARKETING_UA_DS** | ds_abc123 | 1,247,832 | active | yes |
129
- | **REVENUE_EVENTS** | ds_ghi789 | 3,412,006 | stale | yes |
130
-
131
- Using **MARKETING_UA_DS** (ds_abc123), which covers spend, installs, and ROAS columns.
132
- ~~~
133
-
134
- Bold every data source name. If a source is stale, say so and offer to refresh it before querying. If `truncated` is true, raise `--limit` before recommending.
39
+ For refresh modes, history, incremental tuning, and reconciliation, load `data-source-refresh.md`.
@@ -51,13 +51,13 @@ Top values by a measure, shaped to drop straight into an array filter or an `IN
51
51
  ```js
52
52
  const top = await graphit.rank({
53
53
  column: 'COUNTRY', source: 'sales', dataSourceId: 'SALES_DS',
54
- by: '{{metric:REVENUE}}', // governed measure, or SUM(REVENUE) / COUNT(*)
54
+ by: "{{ Metric('revenue') }}", // governed metric, or a bare aggregate
55
55
  limit: 10,
56
56
  filters: { REGION: region.get() }, // optional, same contract as cascade
57
57
  })
58
58
  ```
59
59
 
60
- - Returns a plain array of values. Prefer a `{{metric:NAME}}` reference so the ranking uses the org's definition; the bare aggregate form accepts SUM, COUNT, AVG, MIN, MAX over one column.
60
+ - Returns a plain array of values. Prefer `{{ Metric('name') }}` so ranking uses the org's definition; a bare aggregate accepts SUM, COUNT, AVG, MIN, or MAX over one column.
61
61
 
62
62
  ## graphit.dateRange(id, options) - Date Presets
63
63
 
@@ -1,40 +1,29 @@
1
- # Explaining governance to the user
1
+ # Explaining Governance
2
2
 
3
- Load this when the user is confused about governance itself - asks what "governed" means, why a query was blocked or wanted a reason, what a trust tier or badge is, or how the whole system works. This is for EXPLAINING the system to a person in plain language; writing governed queries is governance.md. Relay it in your own words, tied to what the user actually hit - never paste it verbatim.
3
+ Load when the user asks what governed means, why a query changed, or why it was blocked.
4
4
 
5
- ## The one idea
5
+ ## The idea
6
6
 
7
- Governance means every number is computed the team's agreed way and carries an honest label of how trustworthy it is. It rests entirely on the knowledge base: metrics, dimensions, and rules are shared KB assets - one definition of each calculation and each guardrail, owned by the team. "Governed" is not a lock for its own sake; it is what lets one person's number be reused by everyone else without re-deriving it or second-guessing it.
7
+ The semantic layer defines a business concept. Governance decides how that definition may be queried for this user and context.
8
8
 
9
- ## The two halves
9
+ ## What happens
10
10
 
11
- **1. Enforcement, at query time.** When any query runs, the server checks it against the KB before executing. Rules attach constraints to a table (inject a required WHERE, mask a column, restrict values, require a filter or a GROUP BY) and apply automatically - the user never has to remember them. Every result is stamped with a trust tier, shown as a badge:
11
+ 1. Graphit resolves `Metric`, qualified `Dimension`, and its `Measure` extension.
12
+ 2. It finds rules through model, entity, dimension, metric, and group targets.
13
+ 3. Verified constraints are injected before execution.
14
+ 4. Resolved SQL is verified fail-closed.
15
+ 5. The transparency receipt records what changed and why.
12
16
 
13
- | Tier | Means | Badge |
14
- |------|-------|-------|
15
- | governed | Used `{{metric:X}}` / `{{dim:X}}` KB references | Teal |
16
- | verified | Raw SQL whose math matches a KB definition | Amber |
17
- | ad_hoc | Inline formula with no KB match | Gray |
17
+ A user may not see concealed targets, but still receives a generic safe refusal when access cannot be proven.
18
18
 
19
- Enforcement is server-side and identical on every channel (agent, CLI, dashboard, warehouse); it cannot be weakened from the CLI. An ad-hoc business measure requires EXPLORE access across every table it touches. A rule override additionally requires EXPLORE, the rule policy, and the user's role to allow it. If a query was blocked or asked for a justification, that is the ad-hoc gate: it used no KB reference and matched no definition, so Graphit wants you to either rewrite it with a metric or dimension (creating one if it is missing) or state honestly why a raw run is warranted. The reason is recorded in the audit log; it never bypasses missing access or a disallowed rule override.
19
+ ## Rule states
20
20
 
21
- **2. Auditing, after the fact (Proactive Insights).** Separately, Graphit continuously scans everything the team built - KB definitions, dashboard graphs, the queries that actually ran - and turns each governance gap into a ranked Fix card routed to an owner. Three principles: a card shows only when Graphit is confident it is real (right or silent), one root cause is one card even if it spans many surfaces, and each is routed to an owner (the resource's owner, else its domain owner, else org admins). The ten card types:
22
-
23
- | Card | What it flags |
24
- |------|---------------|
25
- | Govern next | The same ungoverned calculation runs in 2+ places |
26
- | Unsealed dashboard | A dashboard uses definitions the KB does not govern |
27
- | Verified, not wired | A verified definition is re-implemented as raw SQL |
28
- | Unverified in use | Live graphs depend on a definition nobody verified |
29
- | Stale reference | A graph is frozen on an old version of its definition (the only "Breaking" card) |
30
- | Deprecated in use | A deprecated definition is still used by live graphs |
31
- | Conflicting definitions | Two near-identical definitions disagree |
32
- | Rule conflict | Two governance rules contradict each other |
33
- | Unenforced rule | A rule that structurally cannot act on anything |
34
- | Missing owner | Assets with nobody responsible for them |
35
-
36
- Fix is a guided one-click resolution: it drafts and previews the exact change, you apply it (same validation as a manual edit), then Graphit re-checks and clears the card. This queue lives only in the **web app's Governance page** - the CLI cannot list or fix Insights cards (`governance status` and `governance audit` report conformance only). When a user asks about a finding, explain what it means and point them to the Governance page to Fix it.
21
+ - Draft: no effect.
22
+ - Verified body-only: guidance.
23
+ - Verified constrained: enforceable according to mode.
24
+ - Protected-column masks: never overridable.
25
+ - EXPLORE: bounded override where policy permits.
37
26
 
38
27
  ## Relaying it
39
28
 
40
- Answer the person's real question first, in plain language ("your query was blocked because it invented a revenue formula instead of using the governed one"), then give the concrete next step: the governed rewrite, creating the missing asset, or the web Governance page for an Insights card. No jargon dumps, no internal names.
29
+ State the result first, then tier and receipt facts. Name only visible rules. Never invent hidden names or quote raw compiler errors. If blocked, give the server-provided next step. If ad hoc, say so and offer a governed rewrite or supported definition.
@@ -1,100 +1,37 @@
1
1
  # Query Governance
2
2
 
3
- Load this when writing or validating a governed query, working trust tiers, declaring a conditionally-enforced rule, or working the ad-hoc frontier.
3
+ Load when writing a governed query, explaining a refusal, or reporting provenance.
4
4
 
5
- Governance is enforced server-side in the QueryGateway, the same across every channel. You CANNOT weaken it from the CLI: an ad-hoc business measure requires EXPLORE access across every queried scope, and rule overrides must also satisfy EXPLORE.
5
+ ## References
6
6
 
7
- ## Reference syntax
7
+ | Reference | Meaning |
8
+ |---|---|
9
+ | `{{ Metric('revenue') }}` | Reusable metric |
10
+ | `{{ Dimension('order__channel') }}` | Qualified grouping/filter field |
11
+ | `{{ Measure('order_total') }}` | Graphit's model-owned measure extension |
8
12
 
9
- Write governed queries with KB references, not inline formulas. The server compiles each to its KB expression, injects rule constraints, stamps the trust tier, and caches the result. The syntax is identical in `graphit query` and `graphit.resolve()`.
13
+ Legacy token grammar is refused. Keep references inside complete executable SQL and canvas `data-graphit-sql`.
10
14
 
11
- | Syntax | Expands to | Example |
12
- |--------|-----------|---------|
13
- | `{{metric:NAME}}` | Metric calculation (aggregation) | `{{metric:CPI}}` |
14
- | `{{metric:NAME(K=V)}}` | Parameterized metric | `{{metric:ARPU(DAY=7)}}` |
15
- | `{{metric_raw:NAME}}` | Raw expression, no outer aggregate | `{{metric_raw:REVENUE}}` |
16
- | `{{dim:NAME}}` | Dimension expression | `{{dim:INSTALL_MONTH}}` |
17
-
18
- ```bash
19
- graphit query "SELECT {{dim:INSTALL_MONTH}}, {{metric:CPI}} AS cpi FROM MARKETING_UA_DS GROUP BY 1" --ds MARKETING_UA_DS --verbose
20
- ```
21
-
22
- `--verbose` prints the expanded SQL and trust tier, so you can confirm the reference resolved before presenting the result.
23
-
24
- ## Parameterized metrics
25
-
26
- Some metrics (for example ARPU, ROAS, RETENTION) carry required parameters and cannot resolve without a value. Run `graphit kb list metric` and read the `params` column for the names a metric requires, then supply them inline as `{{metric:ARPU(DAY=7)}}`. Pre-baked variants such as `ARPU_D7` or `ROAS_D30` have the value fixed and need none. Omitting a required parameter returns a clear error naming the exact syntax, so read `params` first.
15
+ The governed fragment path serves simple, ratio, and derived metrics. Cumulative, conversion, shifted, time-spine, and null-fill shapes remain unavailable until Project #289. Use a supported decomposition or explicitly labeled free SQL.
27
16
 
28
17
  ## Trust tiers
29
18
 
30
- Every result is stamped with a tier, shown to the user as a badge on graphs and canvas entities - your honest signal of how trustworthy the number is.
31
-
32
- | Tier | Meaning | Badge |
33
- |------|---------|-------|
34
- | `governed` | Query used `{{metric:X}}` / `{{dim:X}}` references | Teal dot |
35
- | `verified` | Raw SQL whose expressions match KB definitions | Amber dot |
36
- | `ad_hoc` | Inline formulas with no KB match | Gray dot |
37
-
38
- Prefer the governed tier. The server may upgrade matching raw SQL to verified, but reach for references first so the result is governed by intent.
39
-
40
- ## The ad-hoc gate
41
-
42
- This is the hard frontier - the rules below are what the QueryGateway does.
43
-
44
- A query lands at the `ad_hoc` tier when it uses no `{{metric:NAME}}` / `{{dim:NAME}}` reference and its raw expressions match no KB definition. On the CLI (`graphit query`, cached `--ds` or `--warehouse`), **every** ad-hoc query must be justified - business measures (an aggregate or `GROUP BY`) and plain exploration alike (`SELECT *`, `COUNT(*)`, `DISTINCT` peeks). Governed and verified results are exempt.
45
-
46
- - **EXPLORE access is a hard prerequisite for ad-hoc business measures.** If the user lacks EXPLORE access to any queried table scope, the server blocks that measure. An ad-hoc reason cannot bypass that denial.
47
- - **Justification floor.** The ad-hoc query is withheld and asks for a justification. Preferred path first: rewrite with `{{metric:NAME}}` / `{{dim:NAME}}` references - genuinely search the KB (`graphit kb explore`, `graphit kb list metric`, `graphit kb list dimension`), and if the metric or dimension you need does not exist, CREATE it first. A filter's value list belongs in a `{{dim:NAME}}`, not a raw `SELECT DISTINCT` peek. Only if nothing fits and the user needs the raw run, pass `--adhoc-reason "<text>"` stating what you searched, what you found, and why it does not fit - pass it on the first call when you already know the query is ad-hoc. A trivial or empty reason is rejected server-side and recorded in the audit log, so it must be honest.
48
-
49
- When the gate fires, do not narrate around it or pretend the query ran. Report truthfully: it was blocked or needs approval, name the governed rewrite, and let the user decide.
50
-
51
- ## Enforceable rules and overrides
52
-
53
- Rules with typed constraints are enforced automatically, rewriting the SQL before it runs:
54
-
55
- | Type | What it does |
56
- |------|-------------|
57
- | `required_where` | Injects a WHERE predicate |
58
- | `forbidden_column` | NULLifies a column in SELECT |
59
- | `value_restriction` | Restricts a column to allowed values |
60
- | `required_filter` | Validates a column appears in WHERE |
61
- | `required_aggregation` | Validates GROUP BY includes a column |
62
-
63
- User-context variables (`${user.team_id}`, `${user.email}`) resolve server-side for row-level security. Override a rule only when the user explicitly asks; the server honors it only if the user holds EXPLORE on every queried scope. A rule that masks a column (a `forbidden_column` constraint) can never be overridden. Every override is logged. Pass several names to override more than one.
64
-
65
- ```bash
66
- graphit query "SELECT * FROM EVENTS" --ds EVENTS --override-rules EXCLUDE_RETARGETING
67
- ```
68
-
69
- ## Conditionally-enforced rules
70
-
71
- A rule's mode is Advisory (guidance), Always (every query), or **Conditional** - fires only when its plain-language body (the condition) holds for the query you wrote. No server classifier decides that; you do. Read a table's rules first (`graphit kb explore table <name>`), judge your query against each conditional rule's body, and declare it up front:
72
-
73
- - `--apply-conditional RULE` - enforce it for this query.
74
- - `--skip-conditional RULE:"reason"` - skip it; a reason is required and audit-logged.
75
-
76
- Declare up front so a clean query never stalls. An undeclared conditional returns a retryable prompt naming each unresolved rule and its condition - read it, decide, re-run. A saved tile stores your decision and replays it every refresh.
77
-
78
- ## Data-source row caps
79
-
80
- An admin can set a `max_rows` cap per data source. A cap limits the result set; it does not change authorization, trust-tier classification, or the ad-hoc justification floor.
19
+ - **governed:** verified semantic references compiled through the gateway.
20
+ - **verified:** known safe stored query without semantic references.
21
+ - **ad hoc:** raw SQL at the frontier.
81
22
 
82
- ## Presenting governance results
23
+ Prefer governed. Never present ad-hoc SQL as the team's definition.
83
24
 
84
- The user CANNOT see raw CLI output. Render every result as markdown.
25
+ ## Rules
85
26
 
86
- **After a governed query**, append a provenance footer so the trust signal travels with the number: tier, KB references used, row cap if applied. Because a governed query may be rewritten before it runs, read `provenance.injection_summary` and report which rules changed it and how (each rule, its outcome, and why), not just a count. An ungoverned query reports no rules applied. For an ad-hoc result, state the tier honestly and offer the governed `{{metric:NAME}}` / `{{dim:NAME}}` rewrite.
27
+ Rules target model, entity, dimension, metric, or group identities. Verified constraints enforce; verified body-only rules guide; drafts do nothing. Modes and EXPLORE behavior remain server-owned.
87
28
 
88
- ~~~
89
- **Trust tier:** governed - 2 KB refs, 1 rule enforced (**EXCLUDE_INTERNAL**), max rows 10000
90
- ~~~
29
+ The gateway runs before caches, injects constraints, verifies resolved SQL, and returns a transparency receipt. Do not claim a rule applied merely because it exists.
91
30
 
92
- **After `graphit governance status`**, show the 7-day conformance counts (governed / verified / ad-hoc / total) as a small markdown table.
31
+ ## Ad-hoc gate
93
32
 
94
- **When a gate or rule blocks a query**, explain it and the path forward, no raw dump:
33
+ Search the KB genuinely, explain why visible definitions do not fit, and prefer an approved reusable supported definition. Use a truthful ad-hoc reason only for a real one-off. Never use it to bypass a rule.
95
34
 
96
- ~~~
97
- **Blocked by governance.**
35
+ ## Reporting
98
36
 
99
- Rule **EXCLUDE_ORGANIC** is enforced; overriding it requires EXPLORE access on every queried scope. Rewrite the query to include the required filter, or ask an admin for EXPLORE access.
100
- ~~~
37
+ Report tier, semantic references, row cap, visible rules that changed the query, and any refusal or override. Read the receipt rather than inferring. A blocked or partial result is not success.
@@ -1,97 +1,44 @@
1
- # KB Actions (Execute)
1
+ # KB Actions
2
2
 
3
- The execute side of KB work: run approved create / update / delete through the `graphit kb` commands. The plan side - what a domain or topic is, why an asset sits where it does - lives in `kb-structure.md`.
3
+ Load when an approved gap must be authored or an existing semantic asset changed.
4
4
 
5
- Create only after the user approves the gap plan below. Names are stored UPPER_SNAKE_CASE. Run `graphit kb create <type> --help` for the exact flag spelling - this file teaches the recipes and policy, the CLI owns the syntax.
5
+ ## Approval gate
6
6
 
7
- ## Gap Plan (present before creating anything)
7
+ Before writing, present the missing concept, proposed root, exact definition, group/access scope, and verification state. Do not write until the user approves.
8
8
 
9
- When the KB-readiness gate finds the dashboard needs assets that do not exist, present a compact plan and get one approval before any write. Show, per missing asset: name, type, formula or expression, the table and topics it lands on, and any rule that applies. Then ask once.
9
+ ## Authoring contract
10
10
 
11
- ~~~
12
- **KB gap - 3 assets missing for this Marketing dashboard:**
11
+ - Create semantic models, metrics, groups, and retained rules from JSON.
12
+ - Use exactly one of `--file` or `--json`.
13
+ - Names are lowercase snake case; reserve `__` for qualified dimensions.
14
+ - Entities, dimensions, and measures mutate only through semantic-model update.
15
+ - A supplied nested list replaces the stored list whole. Read first and include every sibling that must remain.
16
+ - Explicit `meta` replaces author metadata whole. Preserve family, axes, topics, and other author fields.
17
+ - Use dedicated verify/unverify actions. Never patch metadata merely to change verification.
13
18
 
14
- | Asset | Type | Definition | Table | Topics |
15
- |---|---|---|---|---|
16
- | **ROAS_D7** | metric | `SUM(revenue_d7) / NULLIF(SUM(cost), 0)` | MARKETING_UA | ATTRIBUTION |
17
- | **CPI** | metric | `SUM(cost) / NULLIF(SUM(installs), 0)` | MARKETING_UA | ACQUISITION |
18
- | **MEDIA_SOURCE** | dimension | `media_source` | MARKETING_UA | ACQUISITION |
19
+ Read `semantic-authoring.md` for model/metric shapes and `metric-families.md` for concrete variants.
19
20
 
20
- Rule to apply: **EXCLUDE_ORGANIC** already exists and will filter these.
21
- Create these so the dashboard runs on governed references? (Approve / adjust.)
22
- ~~~
21
+ ## Rules
23
22
 
24
- Present the plan, then stop. Do not create until the user approves.
23
+ Rules remain Graphit objects. Create them from JSON with body/constraints plus `apply_on` targets. Final targets are model, entity, dimension, metric, or group identities. A rule without targets is refused.
25
24
 
26
- ## Create
25
+ Constraints keep their five semantics: required predicate, forbidden column, required filter, required aggregation, and value restriction. Use declared semantic identities and typed values.
27
26
 
28
- | Asset | Command shape |
29
- |---|---|
30
- | Metric | `graphit kb create metric --name X --sql "<expr>" --table T` (optional `--topics "A,B"`, `--default-dimensions "D1,D2"`, `--parameters`/`--parameters-file` for templates, `--skip-validate`) |
31
- | Dimension | `graphit kb create dimension --name X --expr "<expr>" --table T` (type auto-inferred; override with `--type` / `--output-type`; `--skip-validate`) |
32
- | Rule | `graphit kb create rule --name X --sql "<text>" --table T` (optional `--constraint`, `--apply-on`, `--topics`, `--skip-validate`) |
33
- | Synonym | `graphit kb create synonym --term X --canonical Y --type metric` |
34
- | Domain | `graphit kb create domain --name X` (optional `--color "#4DB6AC"`) |
35
- | Topic | `graphit kb create topic --name X` |
36
- | Relationship | `graphit kb create relationship --name X --primary-table T --primary-column C --related-table T2 --related-column C2` |
27
+ ## Update
37
28
 
38
- ## Enforceable Rule Flags
29
+ 1. Read the target through the current principal.
30
+ 2. Preserve complete nested and metadata structures.
31
+ 3. Apply the smallest patch.
32
+ 4. Re-read immediately.
33
+ 5. Verify/unverify separately when intended.
34
+ 6. Inspect receipts; a degraded write may have landed and must not be retried blindly.
39
35
 
40
- A plain rule is documentation. To make it enforced server-side at query time, pass typed constraints on `graphit kb create rule` / `graphit kb update rule`:
36
+ ## Delete
41
37
 
42
- - `--constraint <spec...>` - one or more typed constraints, each written `type:value`. Types: `required_where:"<predicate>"`, `forbidden_column:<col>`, `required_filter:<col>`, `required_aggregation:<col>`, `value_restriction:<col>:<in|not_in>:<v1,v2>`. On update the supplied list REPLACES the rule's existing constraints.
38
+ Confirm with the user and inspect usage first. The server checks known definition dependencies, not every canvas reference. A green guard is not exhaustive impact proof.
43
39
 
44
- What each constraint type does at query time, plus the override flow, lives in `governance.md`. Example: a rule that always scopes verified purchases -
40
+ ## Permissions
45
41
 
46
- ```bash
47
- graphit kb create rule --name FILTER_VERIFIED_PURCHASES --sql "Only count verified purchases" --table ORDERS --constraint required_where:"is_verified = true"
48
- ```
42
+ Read access is the ceiling for writes. `kb_write` comes from the effective policy key. Group lifecycle is admin-only. Hidden and missing targets return the same absence.
49
43
 
50
- ## Rule Targeting - what a rule governs
51
-
52
- Every rule must apply to at least one asset; a targetless rule is rejected. `--table` is required, and without `--apply-on` the rule governs that **whole table** - which cascades to every metric and dimension on it (it fires on any query touching the table). To narrow instead, pass `--apply-on metric:NAME` / `--apply-on dimension:NAME` so the rule applies only when that asset is used. The two are mutually exclusive: never mix a table target with metric/dimension targets on one rule. To govern several tables, list each: `--apply-on table:A table:B`.
53
-
54
- ```bash
55
- # Whole table - cascades to every metric/dimension on ORDERS
56
- graphit kb create rule --name ORDERS_ACTIVE_ONLY --sql "Exclude cancelled orders" --table ORDERS
57
- # Narrowed - applies only when the ARPU metric is used
58
- graphit kb create rule --name ARPU_TRIM --sql "Cap ARPU outliers" --table USERS --apply-on metric:ARPU
59
- ```
60
-
61
- ## Pre-Creation Validation
62
-
63
- Metric, dimension, and rule creates validate the formula against real data before writing; the response carries a `validation` object (status `pass` / `skipped` / `fail`). Surface it to the user - per-type result templates are in `kb-traversal.md`. A `fail` returns HTTP 422 and the asset is NOT created: show the error and failing SQL, fix the formula, retry. Validation is skipped (asset still created) for `${PARAM:X}` templates, constraint-based rules, and tables with no ready data source. Pass `--skip-validate` to bypass it on bulk creates.
64
-
65
- ## Update and Delete
66
-
67
- - Update a field: `graphit kb update <type> NAME --<field> value` (e.g. `--sql`, `--expr`, `--description`, `--table`).
68
- - Delete: `graphit kb delete <type> NAME --yes`. Deleting a parameterized parent cascades to all its children - confirm the blast radius first (see `parameterized-metrics.md`).
69
-
70
- ## Lists REPLACE - read before you write
71
-
72
- `--topics`, `--secondary-tables`, `--default-dimensions`, and `--constraint` REPLACE the existing list, they do not append. To change one value, read the current list first (`graphit kb get metric NAME`), then write the full intended list: add a topic with `--topics "EXISTING1,EXISTING2,NEW"`; remove one by writing the list minus that value.
73
-
74
- Reference a **metric or dimension** onto another table with `graphit kb update metric NAME --secondary-tables "OTHER_TABLE"` (also dimension). This is a read-only pointer marked `*` in the tree; every referenced column must exist on the target. For **rules**, govern several tables by listing them in `--apply-on` (`--apply-on table:A table:B`), not `--secondary-tables`. Topics are horizontal - one topic can tag assets across many domains (see `kb-structure.md`).
75
-
76
- ## Domain home (set on the table, cascades)
77
-
78
- Domain is set on the TABLE, never per asset, and cascades to every asset on it (model in `kb-structure.md`). To re-home a whole table at once: `graphit kb update table NAME --domain MARKETING`. Change it once on the table, never asset by asset.
79
-
80
- ## Who can write what
81
-
82
- Reads are open to every member; writes are scoped by the caller's data access profile. Check `graphit status` before presenting a gap plan, so the plan is one they can execute.
83
-
84
- | Write | Needs |
85
- |---|---|
86
- | Metric, dimension, rule, synonym, relationship, table | `kb_write` in the asset's domain |
87
- | Moving an asset or table to another domain | `kb_write` in BOTH domains - the one it leaves and the one it enters |
88
- | Domain and topic create / update / delete | Org admin; a profile never grants it |
89
- | Template create / update / delete | Org admin, OR `kb_write` in any one domain - never a per-template or per-domain grant |
90
-
91
- Every member can read and use templates. Status is advisory; a denial with `retryable: false` is a stop, not a retry (`operations.md`).
92
-
93
- To find what exists and how it connects, use the read recipes in `kb-traversal.md`.
94
-
95
- ## UI-Only (no CLI command)
96
-
97
- Platform-UI state, not KB data the CLI changes: view mode (Tree / By Topic / By Table / Flat), filter dropdowns, expand / collapse, drag-drop onto a topic, and a synonym's cross-cutting domains (extra relevance beyond its home domain has no CLI flag yet).
44
+ Group `access` is stored dbt metadata; it does not grant Graphit visibility.