@graphit/cli 0.2.321 → 0.2.323
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/marketplace.json +3 -3
- package/.claude-plugin/plugin.json +1 -1
- package/.codex-plugin/plugin.json +1 -1
- package/bin/graphit +1 -1
- package/bin/graphit.ps1 +1 -1
- package/dist/api/client.js +15 -0
- package/dist/api/client.js.map +1 -1
- package/dist/commands/ds-config.js +2 -2
- package/dist/commands/ds-config.js.map +1 -1
- package/dist/commands/ds.js +2 -2
- package/dist/commands/ds.js.map +1 -1
- package/dist/commands/kb.js +347 -15
- package/dist/commands/kb.js.map +1 -1
- package/dist/commands/query.js +3 -3
- package/dist/commands/query.js.map +1 -1
- package/dist/skill-guard.js +2 -1
- package/dist/skill-guard.js.map +1 -1
- package/package.json +1 -1
- package/scripts/verb-policy-source.json +30 -102
- package/skills/graphit/SKILL.md +31 -39
- package/skills/graphit/VERSION.json +1 -1
- package/skills/graphit/references/data-source-refresh.md +27 -0
- package/skills/graphit/references/data-sources.md +27 -122
- package/skills/graphit/references/filters-advanced.md +2 -2
- package/skills/graphit/references/governance-explained.md +18 -29
- package/skills/graphit/references/governance.md +20 -83
- package/skills/graphit/references/kb-actions.md +28 -81
- package/skills/graphit/references/kb-discovery.md +35 -64
- package/skills/graphit/references/kb-scope.md +17 -15
- package/skills/graphit/references/kb-structure.md +28 -54
- package/skills/graphit/references/kb-traversal.md +24 -96
- package/skills/graphit/references/metric-families.md +19 -0
- package/skills/graphit/references/migration.md +2 -2
- package/skills/graphit/references/onboarding.md +5 -4
- package/skills/graphit/references/presentations.md +1 -1
- package/skills/graphit/references/runtime.md +6 -6
- package/skills/graphit/references/semantic-authoring.md +65 -0
- package/skills/graphit/references/sql-reference.md +10 -20
- package/dist/commands/kb-constraints.d.ts +0 -14
- package/dist/commands/kb-constraints.js +0 -53
- package/dist/commands/kb-constraints.js.map +0 -1
- package/dist/commands/kb-create.d.ts +0 -2
- package/dist/commands/kb-create.js +0 -296
- package/dist/commands/kb-create.js.map +0 -1
- package/dist/commands/kb-delete.d.ts +0 -2
- package/dist/commands/kb-delete.js +0 -37
- package/dist/commands/kb-delete.js.map +0 -1
- package/dist/commands/kb-read.d.ts +0 -2
- package/dist/commands/kb-read.js +0 -223
- package/dist/commands/kb-read.js.map +0 -1
- package/dist/commands/kb-shared.d.ts +0 -43
- package/dist/commands/kb-shared.js +0 -81
- package/dist/commands/kb-shared.js.map +0 -1
- package/dist/commands/kb-update.d.ts +0 -2
- package/dist/commands/kb-update.js +0 -240
- package/dist/commands/kb-update.js.map +0 -1
- package/skills/graphit/references/parameterized-metrics.md +0 -77
|
@@ -0,0 +1,27 @@
|
|
|
1
|
+
# Data-Source Refresh
|
|
2
|
+
|
|
3
|
+
Load when configuring refresh mode, incremental behavior, history, or reconciliation.
|
|
4
|
+
|
|
5
|
+
## Modes
|
|
6
|
+
|
|
7
|
+
- **full:** replace the cached result from the complete source query.
|
|
8
|
+
- **incremental:** append/merge new source rows using a watermark and stable merge key.
|
|
9
|
+
|
|
10
|
+
Choose incremental only when the source exposes a reliable monotonic watermark and the merge key is unique. Otherwise use full refresh.
|
|
11
|
+
|
|
12
|
+
## Incremental contract
|
|
13
|
+
|
|
14
|
+
- Filter the source early on the watermark column.
|
|
15
|
+
- Preserve lookback for late-arriving updates.
|
|
16
|
+
- Provide the correct watermark type.
|
|
17
|
+
- Verify merge-key uniqueness before serving the new version.
|
|
18
|
+
- Reconcile periodically with a full rebuild.
|
|
19
|
+
- Treat schema drift or grain change as a semantic review, not a blind refresh.
|
|
20
|
+
|
|
21
|
+
## Operations
|
|
22
|
+
|
|
23
|
+
A refresh request may be asynchronous. Poll the job/source status and report counts, duration, version, and failure truthfully. Fire-and-forget is appropriate only when the user did not ask to wait.
|
|
24
|
+
|
|
25
|
+
Inspect refresh history before retrying. A failed status may follow a partially applied external action; use the receipt/status rather than assuming nothing happened.
|
|
26
|
+
|
|
27
|
+
Refresh settings and connector lifecycle require data-source write authority in the server's policy key. Never expose credentials, source rows, or concealed schema in evidence.
|
|
@@ -1,134 +1,39 @@
|
|
|
1
|
-
# Data Sources
|
|
1
|
+
# Data Sources
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
Load when selecting or creating the cached source a semantic model uses.
|
|
4
4
|
|
|
5
|
-
## Routing
|
|
5
|
+
## Routing
|
|
6
6
|
|
|
7
|
-
|
|
7
|
+
1. Read the semantic model's declared data-source binding.
|
|
8
|
+
2. Prefer that cached source for speed, governance, and repeatability.
|
|
9
|
+
3. Use metadata discovery when physical columns are unknown.
|
|
10
|
+
4. Query live warehouse only when no cached source covers the question and the user approves.
|
|
11
|
+
5. Never infer a source from a similarly named model.
|
|
8
12
|
|
|
9
|
-
|
|
10
|
-
|---|---|---|
|
|
11
|
-
| The table has a cached data source | `graphit query "SQL" --ds <NAME>` | roughly 100ms, DuckDB |
|
|
12
|
-
| No data source covers the table | `graphit query "SQL" --warehouse --connection <id>` | roughly 10s, the connected warehouse |
|
|
13
|
+
A group is semantic placement. Data-source creation still accepts `--domain`; pass the uppercase policy key returned by status or the group's `domain_keys`.
|
|
13
14
|
|
|
14
|
-
|
|
15
|
+
## Source SQL
|
|
15
16
|
|
|
16
|
-
|
|
17
|
+
- Select only needed columns and rows.
|
|
18
|
+
- Filter early using base-table columns.
|
|
19
|
+
- Avoid wrapping filter columns when a direct predicate works.
|
|
20
|
+
- Keep complete executable SQL; no ellipses, fake tables, or embedded data.
|
|
21
|
+
- Preserve warehouse dialect.
|
|
22
|
+
- Make grain and refresh mode explicit.
|
|
23
|
+
- Use merge key and watermark only when the source supports them.
|
|
17
24
|
|
|
18
|
-
##
|
|
25
|
+
## Creation
|
|
19
26
|
|
|
20
|
-
|
|
27
|
+
Confirm connector, relation/query, policy key, grain, refresh mode, and cost. Read columns through metadata rather than probing with ad-hoc SQL.
|
|
21
28
|
|
|
22
|
-
|
|
23
|
-
|---|---|---|
|
|
24
|
-
| Grain | `GROUP BY` to the grain you chart (the single biggest lever) | One row per raw event |
|
|
25
|
-
| Columns | Only the columns dashboards use (each is downloaded + cached per query) | 400+ columns "just in case" |
|
|
26
|
-
| Cardinality | Low-card dimensions in the base; ad/campaign names in a separate drill-down | Thousands-of-values dimensions in the base grain |
|
|
27
|
-
| Size | A few-thousand-row typical aggregation | Sitting at the 100M-row / 5GB ceiling |
|
|
29
|
+
Create with automatic scan unless there is a specific reason not to. Creation may be asynchronous; report `creating` honestly and poll status rather than claiming readiness.
|
|
28
30
|
|
|
29
|
-
##
|
|
31
|
+
## Access and safety
|
|
30
32
|
|
|
31
|
-
|
|
33
|
+
- Changing a source requires data-source write capability in its policy key.
|
|
34
|
+
- Reading does not imply authority over connector, SQL, or refresh settings.
|
|
35
|
+
- Visibility and masking cover agent, canvas, render, export, and report paths.
|
|
36
|
+
- Private names and columns remain concealed.
|
|
37
|
+
- Delete/move stay in the Sources Hub where cascades are visible.
|
|
32
38
|
|
|
33
|
-
|
|
34
|
-
|---|---|---|
|
|
35
|
-
| Raw passthrough | No `GROUP BY` / no aggregate - one row per raw event | Pre-aggregate to the grain the dashboards chart |
|
|
36
|
-
| Very wide | Far more columns than dashboards use (e.g. `SELECT *`) | Select only the columns dashboards need |
|
|
37
|
-
| High-cardinality grain | A dimension with thousands of distinct values (ad / campaign / user ids) | Keep it out of the base; build a separate drill-down source |
|
|
38
|
-
| Large + monolithic | A big source whose typical query still scans most rows | Pre-aggregate and narrow so typical queries touch a few thousand rows |
|
|
39
|
-
|
|
40
|
-
A wide or raw source is sometimes the right call - row-level drill-down/export, columns genuinely all used, or a staging source to reshape later. Note the trade-off, then build whichever the user chooses.
|
|
41
|
-
|
|
42
|
-
## Creating data sources
|
|
43
|
-
|
|
44
|
-
`graphit ds create` auto-chains: create -> poll until ready -> scan schema -> print verification link. Activating it for KB use is the `ds verify` step below.
|
|
45
|
-
|
|
46
|
-
```bash
|
|
47
|
-
# Create with auto-scan (recommended)
|
|
48
|
-
graphit ds create --name "MY_DS" --domain <DOMAIN> --sql "SELECT ..." --connection <id>
|
|
49
|
-
|
|
50
|
-
# Create without auto-scan (for special cases)
|
|
51
|
-
graphit ds create --name "MY_DS" --domain <DOMAIN> --sql "SELECT ..." --skip-scan
|
|
52
|
-
```
|
|
53
|
-
|
|
54
|
-
**`--domain` is REQUIRED, in both modes.** It files the scanned table under an existing KB domain and decides who can see the source; there is no uncategorized fallback. `graphit kb list domains`, confirm the choice with the user, and `graphit kb create domain --name <NAME>` if none fits.
|
|
55
|
-
|
|
56
|
-
**From a local file (Excel/CSV):** `graphit ds create --file <path> --domain <NAME>` uploads the file and creates one data source. Optional: `--name` (defaults to the file name), `--sheet <name>` (multi-sheet workbooks). `--file` and `--sql` are mutually exclusive; same flow as above.
|
|
57
|
-
|
|
58
|
-
**Warehouse connection.** `--connection` names the warehouse a `--sql` source reads from. Add BigQuery with `graphit connector add bigquery-serviceaccount --key-file <path> [--project --dataset --location]` (org admin; project defaults from the key). The pipeline routes by connection type - the same `ds create` works for either warehouse.
|
|
59
|
-
|
|
60
|
-
For existing unverified sources, `graphit ds verify <id>` scans and shows the schema; add `--accept-schema` to accept the AI schema and activate a warehouse/SQL source from the CLI. A file upload needs `ds verify` too - it activates without `--accept-schema`, but never at create time, so it stays unqueryable until you run it.
|
|
61
|
-
|
|
62
|
-
## Refreshing data sources
|
|
63
|
-
|
|
64
|
-
Data sources cache a snapshot of the warehouse query result. Refresh when you need current data. **File-upload sources can't be refreshed - update them by re-uploading with `graphit ds create --file <path> --domain <NAME>`.**
|
|
65
|
-
|
|
66
|
-
On BigQuery a refresh scans billed bytes, so keep the shape tight and prefer incremental/partition-pruned refresh over full re-scans; a per-connection scan cap (max bytes billed) fails an oversized query fast rather than running up a bill.
|
|
67
|
-
|
|
68
|
-
```bash
|
|
69
|
-
# Refresh all data sources and wait for completion (live status table)
|
|
70
|
-
graphit ds refresh --all
|
|
71
|
-
|
|
72
|
-
# Fire-and-forget (trigger refreshes, don't wait)
|
|
73
|
-
graphit ds refresh --all --no-wait
|
|
74
|
-
|
|
75
|
-
# Refresh specific sources by ID
|
|
76
|
-
graphit ds refresh <id1> <id2>
|
|
77
|
-
```
|
|
78
|
-
|
|
79
|
-
`graphit ds refresh` only runs an **incremental** refresh (new rows since last update); it never re-exports the whole source. A full rebuild is UI-only (Sources -> Refresh -> Full rebuild); hand off to the UI if one is needed.
|
|
80
|
-
|
|
81
|
-
Refreshes fire in parallel; polls to completion (large sources 30-60s), or returns at once with `--no-wait` (check status via `graphit ds list`). Governed per org: manual refreshes have an hourly budget and a limited number run at once (a reserved slot keeps manual ones unblocked by scheduled). A limit returns a 429 with a clear reason (reset time, or "wait for running operations to finish"); wait it out, don't loop - not source errors. Review past runs with `graphit ds refresh-history <id>`.
|
|
82
|
-
|
|
83
|
-
## Incremental refresh and early-filtering (advanced)
|
|
84
|
-
|
|
85
|
-
Incremental mode fetches only rows past a watermark and merges them in. Three windows govern it: the **watermark column** (which output rows are new), the **merge window** (`--merge-window` - how far back each run re-fetches and upserts, healing late data; API responses call it `lookback_periods`), and per-table **lookback windows** (`--table-lookback` - how far back each source table is *read*). Set on a scanned source; each call sets the COMPLETE config - omitted flags reset to defaults (no `--table-lookback` = windows cleared).
|
|
86
|
-
|
|
87
|
-
When a source aggregates over a wide internal window (e.g. a multi-year rollup), incremental refresh is nearly as slow as full: the outer watermark filter can't prune the inner scan. Early-filtering fixes that - get the contract right first, or older periods silently corrupt on merge:
|
|
88
|
-
|
|
89
|
-
- Size each lookback to cover the merge window plus the longest rolling calculation in the query.
|
|
90
|
-
- Set `--merge-key` (upsert) - overlapping rows double-count without one.
|
|
91
|
-
- Rolling-window metrics (WAU / MAU / stickiness): prefer a layered base daily source instead - recent rows alone can't compute a rolling window.
|
|
92
|
-
- `--reconciliation` is the periodic full-refresh drift backstop (default off; keep it on with the bind or in append mode).
|
|
93
|
-
|
|
94
|
-
Then pick ONE early-filter mode (mutually exclusive, validated):
|
|
95
|
-
|
|
96
|
-
**Per-table lookback windows - preferred; no SQL edit.** Declare how far back each table is read; the engine prunes delta scans and the SQL stays exactly as written. Day-based (date/timestamp watermark required).
|
|
97
|
-
|
|
98
|
-
```bash
|
|
99
|
-
graphit ds refresh-config <id> --mode incremental \
|
|
100
|
-
--watermark-column EVENT_DATE --watermark-type date --merge-key ID \
|
|
101
|
-
--merge-window 3 --table-lookback ANALYTICS.EVENTS:EVENT_DATE:30 \
|
|
102
|
-
--table-lookback USERS:CREATED_AT:90
|
|
103
|
-
```
|
|
104
|
-
|
|
105
|
-
**`:graphit_watermark` bind - for SQL owners.** Place the token where the filter belongs; Graphit substitutes the last watermark on deltas, full history on reconciliation. Output-filter to the current fully-covered period.
|
|
106
|
-
|
|
107
|
-
Only early-filter when an incremental source is slow for this reason; the default refresh is correct and simpler otherwise.
|
|
108
|
-
|
|
109
|
-
## What needs write access
|
|
110
|
-
|
|
111
|
-
Querying a source, listing sources, reading schema or refresh history, and an ordinary `graphit ds refresh` are reads - any member who can read that source's domain can run them, and a source in a domain they cannot read returns the same uniform 404 as one that does not exist. These need `data_source_write` in the source's domain: `ds create`, editing its SQL, `ds refresh-config`, a `--force` refresh or accepting a schema, `ds verify`, scanning, per-source governance settings, and deletion. Moving a source to another domain needs write in both the old and the new domain.
|
|
112
|
-
|
|
113
|
-
Check `graphit status` for those domains before proposing a create or a config change. It is advisory - the server authorizes each operation when it runs, and a denial with `retryable: false` is a stop, not a retry (`operations.md`).
|
|
114
|
-
|
|
115
|
-
## Deleting data sources
|
|
116
|
-
|
|
117
|
-
`ds delete` is not available on the CLI. Deleting a data source cascades to storage and the KB table, removing all metrics, dimensions, and rules on it. Direct the user to the platform UI (Sources Hub), whose confirmation flow shows what will be affected.
|
|
118
|
-
|
|
119
|
-
## Presenting data source results
|
|
120
|
-
|
|
121
|
-
The user cannot see raw CLI output - you are the rendering layer. After `graphit ds list`, present a markdown table and end with a recommendation of which source to use (or note none covers the needed table):
|
|
122
|
-
|
|
123
|
-
~~~
|
|
124
|
-
**2 data sources:**
|
|
125
|
-
|
|
126
|
-
| Name | ID | Rows | Status | Governed |
|
|
127
|
-
|---|---|---:|---|---|
|
|
128
|
-
| **MARKETING_UA_DS** | ds_abc123 | 1,247,832 | active | yes |
|
|
129
|
-
| **REVENUE_EVENTS** | ds_ghi789 | 3,412,006 | stale | yes |
|
|
130
|
-
|
|
131
|
-
Using **MARKETING_UA_DS** (ds_abc123), which covers spend, installs, and ROAS columns.
|
|
132
|
-
~~~
|
|
133
|
-
|
|
134
|
-
Bold every data source name. If a source is stale, say so and offer to refresh it before querying. If `truncated` is true, raise `--limit` before recommending.
|
|
39
|
+
For refresh modes, history, incremental tuning, and reconciliation, load `data-source-refresh.md`.
|
|
@@ -51,13 +51,13 @@ Top values by a measure, shaped to drop straight into an array filter or an `IN
|
|
|
51
51
|
```js
|
|
52
52
|
const top = await graphit.rank({
|
|
53
53
|
column: 'COUNTRY', source: 'sales', dataSourceId: 'SALES_DS',
|
|
54
|
-
by:
|
|
54
|
+
by: "{{ Metric('revenue') }}", // governed metric, or a bare aggregate
|
|
55
55
|
limit: 10,
|
|
56
56
|
filters: { REGION: region.get() }, // optional, same contract as cascade
|
|
57
57
|
})
|
|
58
58
|
```
|
|
59
59
|
|
|
60
|
-
- Returns a plain array of values. Prefer
|
|
60
|
+
- Returns a plain array of values. Prefer `{{ Metric('name') }}` so ranking uses the org's definition; a bare aggregate accepts SUM, COUNT, AVG, MIN, or MAX over one column.
|
|
61
61
|
|
|
62
62
|
## graphit.dateRange(id, options) - Date Presets
|
|
63
63
|
|
|
@@ -1,40 +1,29 @@
|
|
|
1
|
-
# Explaining
|
|
1
|
+
# Explaining Governance
|
|
2
2
|
|
|
3
|
-
Load
|
|
3
|
+
Load when the user asks what governed means, why a query changed, or why it was blocked.
|
|
4
4
|
|
|
5
|
-
## The
|
|
5
|
+
## The idea
|
|
6
6
|
|
|
7
|
-
|
|
7
|
+
The semantic layer defines a business concept. Governance decides how that definition may be queried for this user and context.
|
|
8
8
|
|
|
9
|
-
##
|
|
9
|
+
## What happens
|
|
10
10
|
|
|
11
|
-
|
|
11
|
+
1. Graphit resolves `Metric`, qualified `Dimension`, and its `Measure` extension.
|
|
12
|
+
2. It finds rules through model, entity, dimension, metric, and group targets.
|
|
13
|
+
3. Verified constraints are injected before execution.
|
|
14
|
+
4. Resolved SQL is verified fail-closed.
|
|
15
|
+
5. The transparency receipt records what changed and why.
|
|
12
16
|
|
|
13
|
-
|
|
14
|
-
|------|-------|-------|
|
|
15
|
-
| governed | Used `{{metric:X}}` / `{{dim:X}}` KB references | Teal |
|
|
16
|
-
| verified | Raw SQL whose math matches a KB definition | Amber |
|
|
17
|
-
| ad_hoc | Inline formula with no KB match | Gray |
|
|
17
|
+
A user may not see concealed targets, but still receives a generic safe refusal when access cannot be proven.
|
|
18
18
|
|
|
19
|
-
|
|
19
|
+
## Rule states
|
|
20
20
|
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
| Unsealed dashboard | A dashboard uses definitions the KB does not govern |
|
|
27
|
-
| Verified, not wired | A verified definition is re-implemented as raw SQL |
|
|
28
|
-
| Unverified in use | Live graphs depend on a definition nobody verified |
|
|
29
|
-
| Stale reference | A graph is frozen on an old version of its definition (the only "Breaking" card) |
|
|
30
|
-
| Deprecated in use | A deprecated definition is still used by live graphs |
|
|
31
|
-
| Conflicting definitions | Two near-identical definitions disagree |
|
|
32
|
-
| Rule conflict | Two governance rules contradict each other |
|
|
33
|
-
| Unenforced rule | A rule that structurally cannot act on anything |
|
|
34
|
-
| Missing owner | Assets with nobody responsible for them |
|
|
35
|
-
|
|
36
|
-
Fix is a guided one-click resolution: it drafts and previews the exact change, you apply it (same validation as a manual edit), then Graphit re-checks and clears the card. This queue lives only in the **web app's Governance page** - the CLI cannot list or fix Insights cards (`governance status` and `governance audit` report conformance only). When a user asks about a finding, explain what it means and point them to the Governance page to Fix it.
|
|
21
|
+
- Draft: no effect.
|
|
22
|
+
- Verified body-only: guidance.
|
|
23
|
+
- Verified constrained: enforceable according to mode.
|
|
24
|
+
- Protected-column masks: never overridable.
|
|
25
|
+
- EXPLORE: bounded override where policy permits.
|
|
37
26
|
|
|
38
27
|
## Relaying it
|
|
39
28
|
|
|
40
|
-
|
|
29
|
+
State the result first, then tier and receipt facts. Name only visible rules. Never invent hidden names or quote raw compiler errors. If blocked, give the server-provided next step. If ad hoc, say so and offer a governed rewrite or supported definition.
|
|
@@ -1,100 +1,37 @@
|
|
|
1
1
|
# Query Governance
|
|
2
2
|
|
|
3
|
-
Load
|
|
3
|
+
Load when writing a governed query, explaining a refusal, or reporting provenance.
|
|
4
4
|
|
|
5
|
-
|
|
5
|
+
## References
|
|
6
6
|
|
|
7
|
-
|
|
7
|
+
| Reference | Meaning |
|
|
8
|
+
|---|---|
|
|
9
|
+
| `{{ Metric('revenue') }}` | Reusable metric |
|
|
10
|
+
| `{{ Dimension('order__channel') }}` | Qualified grouping/filter field |
|
|
11
|
+
| `{{ Measure('order_total') }}` | Graphit's model-owned measure extension |
|
|
8
12
|
|
|
9
|
-
|
|
13
|
+
Legacy token grammar is refused. Keep references inside complete executable SQL and canvas `data-graphit-sql`.
|
|
10
14
|
|
|
11
|
-
|
|
12
|
-
|--------|-----------|---------|
|
|
13
|
-
| `{{metric:NAME}}` | Metric calculation (aggregation) | `{{metric:CPI}}` |
|
|
14
|
-
| `{{metric:NAME(K=V)}}` | Parameterized metric | `{{metric:ARPU(DAY=7)}}` |
|
|
15
|
-
| `{{metric_raw:NAME}}` | Raw expression, no outer aggregate | `{{metric_raw:REVENUE}}` |
|
|
16
|
-
| `{{dim:NAME}}` | Dimension expression | `{{dim:INSTALL_MONTH}}` |
|
|
17
|
-
|
|
18
|
-
```bash
|
|
19
|
-
graphit query "SELECT {{dim:INSTALL_MONTH}}, {{metric:CPI}} AS cpi FROM MARKETING_UA_DS GROUP BY 1" --ds MARKETING_UA_DS --verbose
|
|
20
|
-
```
|
|
21
|
-
|
|
22
|
-
`--verbose` prints the expanded SQL and trust tier, so you can confirm the reference resolved before presenting the result.
|
|
23
|
-
|
|
24
|
-
## Parameterized metrics
|
|
25
|
-
|
|
26
|
-
Some metrics (for example ARPU, ROAS, RETENTION) carry required parameters and cannot resolve without a value. Run `graphit kb list metric` and read the `params` column for the names a metric requires, then supply them inline as `{{metric:ARPU(DAY=7)}}`. Pre-baked variants such as `ARPU_D7` or `ROAS_D30` have the value fixed and need none. Omitting a required parameter returns a clear error naming the exact syntax, so read `params` first.
|
|
15
|
+
The governed fragment path serves simple, ratio, and derived metrics. Cumulative, conversion, shifted, time-spine, and null-fill shapes remain unavailable until Project #289. Use a supported decomposition or explicitly labeled free SQL.
|
|
27
16
|
|
|
28
17
|
## Trust tiers
|
|
29
18
|
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|------|---------|-------|
|
|
34
|
-
| `governed` | Query used `{{metric:X}}` / `{{dim:X}}` references | Teal dot |
|
|
35
|
-
| `verified` | Raw SQL whose expressions match KB definitions | Amber dot |
|
|
36
|
-
| `ad_hoc` | Inline formulas with no KB match | Gray dot |
|
|
37
|
-
|
|
38
|
-
Prefer the governed tier. The server may upgrade matching raw SQL to verified, but reach for references first so the result is governed by intent.
|
|
39
|
-
|
|
40
|
-
## The ad-hoc gate
|
|
41
|
-
|
|
42
|
-
This is the hard frontier - the rules below are what the QueryGateway does.
|
|
43
|
-
|
|
44
|
-
A query lands at the `ad_hoc` tier when it uses no `{{metric:NAME}}` / `{{dim:NAME}}` reference and its raw expressions match no KB definition. On the CLI (`graphit query`, cached `--ds` or `--warehouse`), **every** ad-hoc query must be justified - business measures (an aggregate or `GROUP BY`) and plain exploration alike (`SELECT *`, `COUNT(*)`, `DISTINCT` peeks). Governed and verified results are exempt.
|
|
45
|
-
|
|
46
|
-
- **EXPLORE access is a hard prerequisite for ad-hoc business measures.** If the user lacks EXPLORE access to any queried table scope, the server blocks that measure. An ad-hoc reason cannot bypass that denial.
|
|
47
|
-
- **Justification floor.** The ad-hoc query is withheld and asks for a justification. Preferred path first: rewrite with `{{metric:NAME}}` / `{{dim:NAME}}` references - genuinely search the KB (`graphit kb explore`, `graphit kb list metric`, `graphit kb list dimension`), and if the metric or dimension you need does not exist, CREATE it first. A filter's value list belongs in a `{{dim:NAME}}`, not a raw `SELECT DISTINCT` peek. Only if nothing fits and the user needs the raw run, pass `--adhoc-reason "<text>"` stating what you searched, what you found, and why it does not fit - pass it on the first call when you already know the query is ad-hoc. A trivial or empty reason is rejected server-side and recorded in the audit log, so it must be honest.
|
|
48
|
-
|
|
49
|
-
When the gate fires, do not narrate around it or pretend the query ran. Report truthfully: it was blocked or needs approval, name the governed rewrite, and let the user decide.
|
|
50
|
-
|
|
51
|
-
## Enforceable rules and overrides
|
|
52
|
-
|
|
53
|
-
Rules with typed constraints are enforced automatically, rewriting the SQL before it runs:
|
|
54
|
-
|
|
55
|
-
| Type | What it does |
|
|
56
|
-
|------|-------------|
|
|
57
|
-
| `required_where` | Injects a WHERE predicate |
|
|
58
|
-
| `forbidden_column` | NULLifies a column in SELECT |
|
|
59
|
-
| `value_restriction` | Restricts a column to allowed values |
|
|
60
|
-
| `required_filter` | Validates a column appears in WHERE |
|
|
61
|
-
| `required_aggregation` | Validates GROUP BY includes a column |
|
|
62
|
-
|
|
63
|
-
User-context variables (`${user.team_id}`, `${user.email}`) resolve server-side for row-level security. Override a rule only when the user explicitly asks; the server honors it only if the user holds EXPLORE on every queried scope. A rule that masks a column (a `forbidden_column` constraint) can never be overridden. Every override is logged. Pass several names to override more than one.
|
|
64
|
-
|
|
65
|
-
```bash
|
|
66
|
-
graphit query "SELECT * FROM EVENTS" --ds EVENTS --override-rules EXCLUDE_RETARGETING
|
|
67
|
-
```
|
|
68
|
-
|
|
69
|
-
## Conditionally-enforced rules
|
|
70
|
-
|
|
71
|
-
A rule's mode is Advisory (guidance), Always (every query), or **Conditional** - fires only when its plain-language body (the condition) holds for the query you wrote. No server classifier decides that; you do. Read a table's rules first (`graphit kb explore table <name>`), judge your query against each conditional rule's body, and declare it up front:
|
|
72
|
-
|
|
73
|
-
- `--apply-conditional RULE` - enforce it for this query.
|
|
74
|
-
- `--skip-conditional RULE:"reason"` - skip it; a reason is required and audit-logged.
|
|
75
|
-
|
|
76
|
-
Declare up front so a clean query never stalls. An undeclared conditional returns a retryable prompt naming each unresolved rule and its condition - read it, decide, re-run. A saved tile stores your decision and replays it every refresh.
|
|
77
|
-
|
|
78
|
-
## Data-source row caps
|
|
79
|
-
|
|
80
|
-
An admin can set a `max_rows` cap per data source. A cap limits the result set; it does not change authorization, trust-tier classification, or the ad-hoc justification floor.
|
|
19
|
+
- **governed:** verified semantic references compiled through the gateway.
|
|
20
|
+
- **verified:** known safe stored query without semantic references.
|
|
21
|
+
- **ad hoc:** raw SQL at the frontier.
|
|
81
22
|
|
|
82
|
-
|
|
23
|
+
Prefer governed. Never present ad-hoc SQL as the team's definition.
|
|
83
24
|
|
|
84
|
-
|
|
25
|
+
## Rules
|
|
85
26
|
|
|
86
|
-
|
|
27
|
+
Rules target model, entity, dimension, metric, or group identities. Verified constraints enforce; verified body-only rules guide; drafts do nothing. Modes and EXPLORE behavior remain server-owned.
|
|
87
28
|
|
|
88
|
-
|
|
89
|
-
**Trust tier:** governed - 2 KB refs, 1 rule enforced (**EXCLUDE_INTERNAL**), max rows 10000
|
|
90
|
-
~~~
|
|
29
|
+
The gateway runs before caches, injects constraints, verifies resolved SQL, and returns a transparency receipt. Do not claim a rule applied merely because it exists.
|
|
91
30
|
|
|
92
|
-
|
|
31
|
+
## Ad-hoc gate
|
|
93
32
|
|
|
94
|
-
|
|
33
|
+
Search the KB genuinely, explain why visible definitions do not fit, and prefer an approved reusable supported definition. Use a truthful ad-hoc reason only for a real one-off. Never use it to bypass a rule.
|
|
95
34
|
|
|
96
|
-
|
|
97
|
-
**Blocked by governance.**
|
|
35
|
+
## Reporting
|
|
98
36
|
|
|
99
|
-
|
|
100
|
-
~~~
|
|
37
|
+
Report tier, semantic references, row cap, visible rules that changed the query, and any refusal or override. Read the receipt rather than inferring. A blocked or partial result is not success.
|
|
@@ -1,97 +1,44 @@
|
|
|
1
|
-
# KB Actions
|
|
1
|
+
# KB Actions
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
Load when an approved gap must be authored or an existing semantic asset changed.
|
|
4
4
|
|
|
5
|
-
|
|
5
|
+
## Approval gate
|
|
6
6
|
|
|
7
|
-
|
|
7
|
+
Before writing, present the missing concept, proposed root, exact definition, group/access scope, and verification state. Do not write until the user approves.
|
|
8
8
|
|
|
9
|
-
|
|
9
|
+
## Authoring contract
|
|
10
10
|
|
|
11
|
-
|
|
12
|
-
|
|
11
|
+
- Create semantic models, metrics, groups, and retained rules from JSON.
|
|
12
|
+
- Use exactly one of `--file` or `--json`.
|
|
13
|
+
- Names are lowercase snake case; reserve `__` for qualified dimensions.
|
|
14
|
+
- Entities, dimensions, and measures mutate only through semantic-model update.
|
|
15
|
+
- A supplied nested list replaces the stored list whole. Read first and include every sibling that must remain.
|
|
16
|
+
- Explicit `meta` replaces author metadata whole. Preserve family, axes, topics, and other author fields.
|
|
17
|
+
- Use dedicated verify/unverify actions. Never patch metadata merely to change verification.
|
|
13
18
|
|
|
14
|
-
|
|
15
|
-
|---|---|---|---|---|
|
|
16
|
-
| **ROAS_D7** | metric | `SUM(revenue_d7) / NULLIF(SUM(cost), 0)` | MARKETING_UA | ATTRIBUTION |
|
|
17
|
-
| **CPI** | metric | `SUM(cost) / NULLIF(SUM(installs), 0)` | MARKETING_UA | ACQUISITION |
|
|
18
|
-
| **MEDIA_SOURCE** | dimension | `media_source` | MARKETING_UA | ACQUISITION |
|
|
19
|
+
Read `semantic-authoring.md` for model/metric shapes and `metric-families.md` for concrete variants.
|
|
19
20
|
|
|
20
|
-
|
|
21
|
-
Create these so the dashboard runs on governed references? (Approve / adjust.)
|
|
22
|
-
~~~
|
|
21
|
+
## Rules
|
|
23
22
|
|
|
24
|
-
|
|
23
|
+
Rules remain Graphit objects. Create them from JSON with body/constraints plus `apply_on` targets. Final targets are model, entity, dimension, metric, or group identities. A rule without targets is refused.
|
|
25
24
|
|
|
26
|
-
|
|
25
|
+
Constraints keep their five semantics: required predicate, forbidden column, required filter, required aggregation, and value restriction. Use declared semantic identities and typed values.
|
|
27
26
|
|
|
28
|
-
|
|
29
|
-
|---|---|
|
|
30
|
-
| Metric | `graphit kb create metric --name X --sql "<expr>" --table T` (optional `--topics "A,B"`, `--default-dimensions "D1,D2"`, `--parameters`/`--parameters-file` for templates, `--skip-validate`) |
|
|
31
|
-
| Dimension | `graphit kb create dimension --name X --expr "<expr>" --table T` (type auto-inferred; override with `--type` / `--output-type`; `--skip-validate`) |
|
|
32
|
-
| Rule | `graphit kb create rule --name X --sql "<text>" --table T` (optional `--constraint`, `--apply-on`, `--topics`, `--skip-validate`) |
|
|
33
|
-
| Synonym | `graphit kb create synonym --term X --canonical Y --type metric` |
|
|
34
|
-
| Domain | `graphit kb create domain --name X` (optional `--color "#4DB6AC"`) |
|
|
35
|
-
| Topic | `graphit kb create topic --name X` |
|
|
36
|
-
| Relationship | `graphit kb create relationship --name X --primary-table T --primary-column C --related-table T2 --related-column C2` |
|
|
27
|
+
## Update
|
|
37
28
|
|
|
38
|
-
|
|
29
|
+
1. Read the target through the current principal.
|
|
30
|
+
2. Preserve complete nested and metadata structures.
|
|
31
|
+
3. Apply the smallest patch.
|
|
32
|
+
4. Re-read immediately.
|
|
33
|
+
5. Verify/unverify separately when intended.
|
|
34
|
+
6. Inspect receipts; a degraded write may have landed and must not be retried blindly.
|
|
39
35
|
|
|
40
|
-
|
|
36
|
+
## Delete
|
|
41
37
|
|
|
42
|
-
|
|
38
|
+
Confirm with the user and inspect usage first. The server checks known definition dependencies, not every canvas reference. A green guard is not exhaustive impact proof.
|
|
43
39
|
|
|
44
|
-
|
|
40
|
+
## Permissions
|
|
45
41
|
|
|
46
|
-
|
|
47
|
-
graphit kb create rule --name FILTER_VERIFIED_PURCHASES --sql "Only count verified purchases" --table ORDERS --constraint required_where:"is_verified = true"
|
|
48
|
-
```
|
|
42
|
+
Read access is the ceiling for writes. `kb_write` comes from the effective policy key. Group lifecycle is admin-only. Hidden and missing targets return the same absence.
|
|
49
43
|
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
Every rule must apply to at least one asset; a targetless rule is rejected. `--table` is required, and without `--apply-on` the rule governs that **whole table** - which cascades to every metric and dimension on it (it fires on any query touching the table). To narrow instead, pass `--apply-on metric:NAME` / `--apply-on dimension:NAME` so the rule applies only when that asset is used. The two are mutually exclusive: never mix a table target with metric/dimension targets on one rule. To govern several tables, list each: `--apply-on table:A table:B`.
|
|
53
|
-
|
|
54
|
-
```bash
|
|
55
|
-
# Whole table - cascades to every metric/dimension on ORDERS
|
|
56
|
-
graphit kb create rule --name ORDERS_ACTIVE_ONLY --sql "Exclude cancelled orders" --table ORDERS
|
|
57
|
-
# Narrowed - applies only when the ARPU metric is used
|
|
58
|
-
graphit kb create rule --name ARPU_TRIM --sql "Cap ARPU outliers" --table USERS --apply-on metric:ARPU
|
|
59
|
-
```
|
|
60
|
-
|
|
61
|
-
## Pre-Creation Validation
|
|
62
|
-
|
|
63
|
-
Metric, dimension, and rule creates validate the formula against real data before writing; the response carries a `validation` object (status `pass` / `skipped` / `fail`). Surface it to the user - per-type result templates are in `kb-traversal.md`. A `fail` returns HTTP 422 and the asset is NOT created: show the error and failing SQL, fix the formula, retry. Validation is skipped (asset still created) for `${PARAM:X}` templates, constraint-based rules, and tables with no ready data source. Pass `--skip-validate` to bypass it on bulk creates.
|
|
64
|
-
|
|
65
|
-
## Update and Delete
|
|
66
|
-
|
|
67
|
-
- Update a field: `graphit kb update <type> NAME --<field> value` (e.g. `--sql`, `--expr`, `--description`, `--table`).
|
|
68
|
-
- Delete: `graphit kb delete <type> NAME --yes`. Deleting a parameterized parent cascades to all its children - confirm the blast radius first (see `parameterized-metrics.md`).
|
|
69
|
-
|
|
70
|
-
## Lists REPLACE - read before you write
|
|
71
|
-
|
|
72
|
-
`--topics`, `--secondary-tables`, `--default-dimensions`, and `--constraint` REPLACE the existing list, they do not append. To change one value, read the current list first (`graphit kb get metric NAME`), then write the full intended list: add a topic with `--topics "EXISTING1,EXISTING2,NEW"`; remove one by writing the list minus that value.
|
|
73
|
-
|
|
74
|
-
Reference a **metric or dimension** onto another table with `graphit kb update metric NAME --secondary-tables "OTHER_TABLE"` (also dimension). This is a read-only pointer marked `*` in the tree; every referenced column must exist on the target. For **rules**, govern several tables by listing them in `--apply-on` (`--apply-on table:A table:B`), not `--secondary-tables`. Topics are horizontal - one topic can tag assets across many domains (see `kb-structure.md`).
|
|
75
|
-
|
|
76
|
-
## Domain home (set on the table, cascades)
|
|
77
|
-
|
|
78
|
-
Domain is set on the TABLE, never per asset, and cascades to every asset on it (model in `kb-structure.md`). To re-home a whole table at once: `graphit kb update table NAME --domain MARKETING`. Change it once on the table, never asset by asset.
|
|
79
|
-
|
|
80
|
-
## Who can write what
|
|
81
|
-
|
|
82
|
-
Reads are open to every member; writes are scoped by the caller's data access profile. Check `graphit status` before presenting a gap plan, so the plan is one they can execute.
|
|
83
|
-
|
|
84
|
-
| Write | Needs |
|
|
85
|
-
|---|---|
|
|
86
|
-
| Metric, dimension, rule, synonym, relationship, table | `kb_write` in the asset's domain |
|
|
87
|
-
| Moving an asset or table to another domain | `kb_write` in BOTH domains - the one it leaves and the one it enters |
|
|
88
|
-
| Domain and topic create / update / delete | Org admin; a profile never grants it |
|
|
89
|
-
| Template create / update / delete | Org admin, OR `kb_write` in any one domain - never a per-template or per-domain grant |
|
|
90
|
-
|
|
91
|
-
Every member can read and use templates. Status is advisory; a denial with `retryable: false` is a stop, not a retry (`operations.md`).
|
|
92
|
-
|
|
93
|
-
To find what exists and how it connects, use the read recipes in `kb-traversal.md`.
|
|
94
|
-
|
|
95
|
-
## UI-Only (no CLI command)
|
|
96
|
-
|
|
97
|
-
Platform-UI state, not KB data the CLI changes: view mode (Tree / By Topic / By Table / Flat), filter dropdowns, expand / collapse, drag-drop onto a topic, and a synonym's cross-cutting domains (extra relevance beyond its home domain has no CLI flag yet).
|
|
44
|
+
Group `access` is stored dbt metadata; it does not grant Graphit visibility.
|