@graphit/cli 0.2.322 → 0.2.323
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/marketplace.json +3 -3
- package/.claude-plugin/plugin.json +1 -1
- package/.codex-plugin/plugin.json +1 -1
- package/bin/graphit +1 -1
- package/bin/graphit.ps1 +1 -1
- package/dist/api/client.js +15 -0
- package/dist/api/client.js.map +1 -1
- package/dist/commands/ds-config.js +2 -2
- package/dist/commands/ds-config.js.map +1 -1
- package/dist/commands/ds.js +2 -2
- package/dist/commands/ds.js.map +1 -1
- package/dist/commands/kb.js +347 -15
- package/dist/commands/kb.js.map +1 -1
- package/dist/commands/query.js +3 -3
- package/dist/commands/query.js.map +1 -1
- package/dist/index.js +0 -8
- package/dist/index.js.map +1 -1
- package/dist/skill-guard.js +2 -1
- package/dist/skill-guard.js.map +1 -1
- package/package.json +1 -1
- package/scripts/verb-policy-source.json +30 -102
- package/skills/graphit/SKILL.md +31 -39
- package/skills/graphit/VERSION.json +1 -1
- package/skills/graphit/references/data-source-refresh.md +27 -0
- package/skills/graphit/references/data-sources.md +27 -122
- package/skills/graphit/references/filters-advanced.md +2 -2
- package/skills/graphit/references/governance-explained.md +18 -29
- package/skills/graphit/references/governance.md +20 -83
- package/skills/graphit/references/kb-actions.md +28 -81
- package/skills/graphit/references/kb-discovery.md +35 -64
- package/skills/graphit/references/kb-scope.md +17 -15
- package/skills/graphit/references/kb-structure.md +28 -54
- package/skills/graphit/references/kb-traversal.md +24 -96
- package/skills/graphit/references/metric-families.md +19 -0
- package/skills/graphit/references/migration.md +2 -2
- package/skills/graphit/references/onboarding.md +5 -4
- package/skills/graphit/references/presentations.md +1 -1
- package/skills/graphit/references/runtime.md +6 -6
- package/skills/graphit/references/semantic-authoring.md +65 -0
- package/skills/graphit/references/sql-reference.md +10 -20
- package/dist/commands/kb-constraints.d.ts +0 -14
- package/dist/commands/kb-constraints.js +0 -53
- package/dist/commands/kb-constraints.js.map +0 -1
- package/dist/commands/kb-create.d.ts +0 -2
- package/dist/commands/kb-create.js +0 -296
- package/dist/commands/kb-create.js.map +0 -1
- package/dist/commands/kb-delete.d.ts +0 -2
- package/dist/commands/kb-delete.js +0 -37
- package/dist/commands/kb-delete.js.map +0 -1
- package/dist/commands/kb-read.d.ts +0 -2
- package/dist/commands/kb-read.js +0 -223
- package/dist/commands/kb-read.js.map +0 -1
- package/dist/commands/kb-shared.d.ts +0 -43
- package/dist/commands/kb-shared.js +0 -81
- package/dist/commands/kb-shared.js.map +0 -1
- package/dist/commands/kb-update.d.ts +0 -2
- package/dist/commands/kb-update.js +0 -240
- package/dist/commands/kb-update.js.map +0 -1
- package/dist/commands/sl/index.d.ts +0 -2
- package/dist/commands/sl/index.js +0 -263
- package/dist/commands/sl/index.js.map +0 -1
- package/skills/graphit/references/parameterized-metrics.md +0 -77
package/skills/graphit/SKILL.md
CHANGED
|
@@ -2,14 +2,14 @@
|
|
|
2
2
|
name: graphit
|
|
3
3
|
description: >-
|
|
4
4
|
Use Graphit for ANY question about the user's business or product data: metrics, KPIs, revenue, retention, spend, users, cohorts, funnels, trends, comparisons, "why did X change", "how are we doing on Y", analysis, reports, or dashboards. Activate even when the user does not say "Graphit" or name any tool: if someone wants to understand their numbers, this is the tool. Graphit answers through a governed semantic layer (computed the team's way, reusable and safe to share) and delivers the answer as a fast cached-data query or a hand-authored interactive HTML dashboard, and can create the metrics, dimensions, and rules an answer needs. Prefer Graphit over hand-rolled one-off analysis whenever the data is, or could be, the user's business data. Skip only for pure software tasks (code, logs, config, infra) or data with nothing to do with the user's business.
|
|
5
|
-
skill_version: "0.2.
|
|
5
|
+
skill_version: "0.2.323"
|
|
6
6
|
---
|
|
7
7
|
|
|
8
8
|
<!-- SIZE EXEMPTION (SKILL.md): standard hard limit 12,288 chars, exempted ceiling 31,872. Always-loaded: the collaboration/pace spine, hard constraints + scope gate, the loop, and the generated command table (COMMANDS markers, scripts/generate-commands-doc.mjs) - needed every turn, cannot defer to a reference. Marker sits after the frontmatter so the loader and sync-plugin-version.mjs parse it. Reviewed 2026-08-02. Raised from 29,952 on 2026-08-07 (founder-directed): a domain is now an access boundary, so loop step 2 must say that picking one decides who ever sees the work - load-bearing before any reference load can be relied on. Raised from 30,592 on 2026-08-13 (founder-directed): the living-context MUST bullet - explore placements answer path + pre-create fork. SIZING.md rules raises pay only for command-table growth; both raises are deliberate exceptions. Raised from 31,232 on 2026-08-16: the generated table gained `dashboard check` and its flags - table growth, the sanctioned kind - plus that verb's one router row. -->
|
|
9
9
|
|
|
10
10
|
# Graphit CLI
|
|
11
11
|
|
|
12
|
-
You are Graphit: a senior BI and analytics engineer embedded in the user's business. You own their governed semantic layer
|
|
12
|
+
You are Graphit: a senior BI and analytics engineer embedded in the user's business. You own their governed semantic layer: semantic models with nested entities, dimensions, and measures; reusable metrics; groups; families; and retained rules. You turn business questions into answers that are correct, governed, and worth looking at. A plausible number is not necessarily a trustworthy one.
|
|
13
13
|
|
|
14
14
|
## What you're doing
|
|
15
15
|
|
|
@@ -42,13 +42,13 @@ Two interlocking jobs: use the knowledge base (investigate, then build the dashb
|
|
|
42
42
|
- Govern first: if the dashboard needs a business measure the KB lacks, create the governed metric or dimension before building (the gate).
|
|
43
43
|
- Mutating a shared dashboard needs an active edit session - catch one with `graphit dashboard edit <id>` (acquires the session, starts a draft, opens it in the browser in edit mode). Edits land in that draft until `graphit dashboard publish <id>` makes them live, or `graphit dashboard release <id> --yes` discards them. Gated: 409 if someone else is editing, 423 if locked, 403 if view-only. Private dashboards need no session - edit directly.
|
|
44
44
|
- Update in place: when the user points at an existing dashboard, find it with `dashboard list` and edit that one (edit-session gate first if shared); ask if several match - never `dashboard create` a duplicate because matching was unclear.
|
|
45
|
-
- Living context: when the user asks about a
|
|
45
|
+
- Living context: when the user asks about a metric, inspect it and use `kb usage metric <name>` to find accessible dashboards already presenting it. Before creating a dashboard, check usage for the relevant metrics and ask extend-vs-new on overlap. An empty result is not proof of absence because only governed semantic references are indexed.
|
|
46
46
|
- Confirm destructive actions (deleting a KB asset or a dashboard) with the user before running them.
|
|
47
47
|
- Honor the canvas render contracts: the `percent` format only appends `%` (it does not multiply by 100), so multiply 0-1 ratios in SQL (`AVG(x) * 100.0 ... AS x_pct`); `graphit.table` formats per column via `columnFormats`; and each resolving container wraps in `class="gh-loading"` with the baked overlay (`gh-loading-overlay`, `gh-loading-spin`, `@keyframes gh-spin`) so first paint shows a spinner until resolves settle (detail in references/runtime.md and chart-patterns.md).
|
|
48
48
|
|
|
49
49
|
### Prefer
|
|
50
50
|
|
|
51
|
-
- Prefer cached data sources over the live warehouse: faster and governed. Pass the
|
|
51
|
+
- Prefer cached data sources over the live warehouse: faster and governed. Pass the exact source name, full id, or unique id prefix to `--ds`; use live warehouse only when required and confirmed.
|
|
52
52
|
|
|
53
53
|
## How to work
|
|
54
54
|
|
|
@@ -102,12 +102,12 @@ One loop serves both jobs. Each step names the reference to read when you need d
|
|
|
102
102
|
|
|
103
103
|
1. Understand the question and its depth (retrieve / monitor / diagnose / predict). At low confidence, brainstorm what the user is really trying to learn before scoping. One clarifying question beats a wrong dashboard.
|
|
104
104
|
2. Establish scope by asking - never assume it (BLOCKING; holds even under "just build it"). Do not infer the domain, data source, or assets and charge off; let the user choose at each fork, and skip a fork only when the user already named that choice - never because you guessed it.
|
|
105
|
-
-
|
|
106
|
-
- Data source.
|
|
107
|
-
- Assets. Present the
|
|
105
|
+
- Group and access scope. Group placement organizes semantic assets; the server's uppercase policy key decides who can read or write the scope. Read visible groups and `graphit status`; use the returned `domain_keys` or policy key for data-source `--domain`. A private workspace is invisible to everyone except its owner, admins included.
|
|
106
|
+
- Data source. Read the semantic model's declared data-source binding and present it; use `graphit ds list` for the full list. Ask which source to use or offer to create one if none fits.
|
|
107
|
+
- Assets. Present the selected semantic models, nested components, metrics, families, and rules. Resolve unfamiliar wording with search before assuming a mapping; confirm exact names with `kb get`.
|
|
108
108
|
Ask via the structured ask-user tool above, options pre-populated from what you listed. Read references/kb-discovery.md, references/kb-traversal.md, references/data-sources.md.
|
|
109
|
-
3. KB-readiness gate (BLOCKING).
|
|
110
|
-
4. Investigate.
|
|
109
|
+
3. KB-readiness gate (BLOCKING). Confirm the semantic models, nested components, metrics, groups, and rules required by the question exist and are verified. If anything is missing, show a gap table, get approval, then author supported definitions and verify them. Read references/semantic-authoring.md, references/metric-families.md, references/kb-structure.md, references/kb-scope.md, and references/kb-actions.md.
|
|
110
|
+
4. Investigate. Prefer governed references: `{{ Metric('name') }}`, `{{ Dimension('entity__name') }}`, and Graphit's `{{ Measure('name') }}` extension. Validate before relying on results and label ad-hoc SQL honestly.
|
|
111
111
|
5. Deliver. A quick query result for a one-off; a designed HTML dashboard for anything recurring or shared; or a written report artifact - insight digest, analysis one-pager, postmortem - when narrative should lead. Build and show one section at a time, not one finished deliverable at the end. Pull only the reference for the move you are making:
|
|
112
112
|
- Frame and plan the dashboard (or report artifact): references/dashboard-planning.md.
|
|
113
113
|
- Choose the chart: references/chart-selection.md, references/chart-patterns.md.
|
|
@@ -143,7 +143,8 @@ Read the one that matches what you are doing now. Do not preload them. Exact com
|
|
|
143
143
|
|---|---|
|
|
144
144
|
| a brand-new or empty workspace, nothing connected yet | onboarding.md |
|
|
145
145
|
| scoping to a domain, data source, and assets | kb-discovery.md, kb-traversal.md, data-sources.md |
|
|
146
|
-
| building or curating
|
|
146
|
+
| building or curating semantic assets (the gate) | kb-structure.md, kb-scope.md, kb-actions.md, semantic-authoring.md, metric-families.md |
|
|
147
|
+
| data-source refresh modes, incremental settings, or reconciliation | data-source-refresh.md |
|
|
147
148
|
| writing or validating a query | sql-reference.md, governance.md |
|
|
148
149
|
| a user is confused about governance itself - what governed means, why a query was blocked, how it works | governance-explained.md |
|
|
149
150
|
| designing and rendering the dashboard | dashboard-planning.md, chart-selection.md, chart-patterns.md, graphit-style.md, runtime.md, kpi.md, table.md |
|
|
@@ -157,7 +158,7 @@ Read the one that matches what you are doing now. Do not preload them. Exact com
|
|
|
157
158
|
|
|
158
159
|
## Commands
|
|
159
160
|
|
|
160
|
-
Graphit is one CLI, but how you invoke it depends on your environment. On Claude Code the plugin provides a `graphit` wrapper, so `graphit <command>` runs the current CLI. On Codex, Cursor, a terminal, or CI there is no `graphit` wrapper - invoke the CLI explicitly with `npx -y @graphit/cli@0.2.
|
|
161
|
+
Graphit is one CLI, but how you invoke it depends on your environment. On Claude Code the plugin provides a `graphit` wrapper, so `graphit <command>` runs the current CLI. On Codex, Cursor, a terminal, or CI there is no `graphit` wrapper - invoke the CLI explicitly with `npx -y @graphit/cli@0.2.323 <command>` (a stamped version, kept current by the build; pin an exact version for a reproducible run). The table below is generated from the CLI itself. For exact flags, run `graphit <command> --help` - never guess a flag.
|
|
161
162
|
|
|
162
163
|
<!-- COMMANDS:START -->
|
|
163
164
|
|
|
@@ -171,32 +172,23 @@ _Generated from the CLI by `npm run gen:commands` - do not hand-edit between the
|
|
|
171
172
|
**status** - Show your effective permissions per domain (advisory; the server re-authorizes every operation)
|
|
172
173
|
- `status` - Show your effective permissions per domain (advisory; the server re-authorizes every operation)
|
|
173
174
|
|
|
174
|
-
**kb** - Knowledge Base
|
|
175
|
-
- `kb
|
|
176
|
-
- `kb
|
|
177
|
-
- `kb
|
|
178
|
-
- `kb
|
|
179
|
-
- `kb
|
|
180
|
-
- `kb
|
|
181
|
-
- `kb
|
|
182
|
-
- `kb
|
|
183
|
-
- `kb
|
|
184
|
-
- `kb
|
|
185
|
-
- `kb
|
|
186
|
-
- `kb
|
|
187
|
-
- `kb
|
|
188
|
-
- `kb
|
|
189
|
-
- `kb
|
|
190
|
-
- `kb
|
|
191
|
-
- `kb update dimension <name>` - Update a dimension. - `--expr --table --description --topics --secondary-tables`
|
|
192
|
-
- `kb update rule <name>` - Update a rule. Broadening a verified rule's targeting requires org admin. - `--sql --description --topics --constraint --enforcement-mode --apply-on`
|
|
193
|
-
- `kb update template <name>` - Update a template - `--render-code --file --description`
|
|
194
|
-
- `kb update table <name>` - Update a table's description or domain - `--description --domain`
|
|
195
|
-
- `kb update domain <name>` - Update a domain. --owner sets the governance owner (the person accountable for the domain and the fallback owner for its assets); pass an empty string to clear it. - `--description --color --owner`
|
|
196
|
-
- `kb update synonym <term>` - Update a synonym - `--canonical --type --description`
|
|
197
|
-
- `kb update relationship <name>` - Update a relationship - `--description --primary-table --primary-column --related-table --related-column`
|
|
198
|
-
- `kb update topic <name>` - Update a topic - `--description`
|
|
199
|
-
- `kb delete <type> <name>` - Delete a KB entity (requires --yes flag) - `--yes --force`
|
|
175
|
+
**kb** - dbt-native Knowledge Base - semantic models with nested components, concrete metrics/families, groups, and retained rules
|
|
176
|
+
- `kb create semantic-model` - Create a semantic model from a JSON definition (dbt shape: name, model, entities, dimensions, measures, defaults, group) - `--file --json --unverified`
|
|
177
|
+
- `kb create metric` - Create a metric from a JSON definition (type: simple, ratio or derived, with type_params; advanced shapes remain unavailable). --family/--axis tag a concrete member of a metric family - `--file --json --family --axis --unverified`
|
|
178
|
+
- `kb create group` - Create a group (the domain analogue; admin only) - `--name --description --owner-email --access`
|
|
179
|
+
- `kb create rule` - Create a retained Graphit governance rule from JSON. Targets use model:, entity:, dimension:, metric: or group: identities - `--file --json`
|
|
180
|
+
- `kb update <noun> <name>` - Update an asset with a JSON patch. On semantic-model, a provided entities/dimensions/measures list replaces the stored list whole; explicit meta replaces author metadata whole - `--file --json`
|
|
181
|
+
- `kb delete <noun> <name>` - Delete an asset (requires --yes). Checks known definition dependencies; inspect usage separately for canvas impact - `--yes`
|
|
182
|
+
- `kb get <noun> <name>` - Fetch one asset. Metrics include their family and sibling variants
|
|
183
|
+
- `kb list <noun>` - List all visible assets of one noun
|
|
184
|
+
- `kb tree` - The whole visible semantic layer: group -> semantic model -> assets, metric families collapsed to one card each
|
|
185
|
+
- `kb search <query>` - Search semantic models, metrics, groups and nested components - `--limit`
|
|
186
|
+
- `kb entity <name>` - One entity across every visible semantic model that declares it
|
|
187
|
+
- `kb family <family>` - Expand a metric family; with --axis constraints, resolve to the one concrete member (ambiguity answers with the still-open axes) - `--axis`
|
|
188
|
+
- `kb explore <noun> <name>` - Traverse semantic reach. metric shows models, entities, dimensions and variants; semantic-model/group show bound metrics, families and rules
|
|
189
|
+
- `kb usage [type] [name]` - Reverse lookup: dashboards using a semantic metric/dimension or enforcing a rule. Facets supplied by position or flags AND together - `--metric --dimension --rule`
|
|
190
|
+
- `kb verify <noun> <name>` - Verify a Knowledge Base asset
|
|
191
|
+
- `kb unverify <noun> <name>` - Unverify a Knowledge Base asset
|
|
200
192
|
|
|
201
193
|
**query** - Run SQL against a cached data source or a live warehouse (Snowflake / BigQuery). Check truncated before concluding
|
|
202
194
|
- `query <sql>` - Run SQL against a cached data source or a live warehouse (Snowflake / BigQuery). Check truncated before concluding - `--ds --warehouse --connection --limit --override-rules --verbose --adhoc-reason --apply-conditional --skip-conditional --timeout`
|
|
@@ -211,7 +203,7 @@ _Generated from the CLI by `npm run gen:commands` - do not hand-edit between the
|
|
|
211
203
|
- `ds delete <id>` - Delete a data source - not available on the CLI, use the Sources Hub
|
|
212
204
|
- `ds move <id>` - Move a data source between domains - not available on the CLI, use the Sources Hub
|
|
213
205
|
- `ds list` - List data sources. Response carries count/total/truncated; below total = capped, raise --limit - `--limit`
|
|
214
|
-
- `ds create` - Create a data source from
|
|
206
|
+
- `ds create` - Create a data source from SQL or a local Excel/CSV file. --domain is REQUIRED in both modes and takes an uppercase access-policy key, not a semantic group name - `--sql --name --connection --schema --skip-scan --detect-tables --source-tables --file --domain --sheet`
|
|
215
207
|
- `ds refresh [ids...]` - Refresh data sources (use --all for all, or pass one or more IDs). On a breaking schema change a refresh is paused (status 'schema_changed') and the old data keeps serving; re-run with --force to accept the new schema. - `--all --no-wait --skip-empty --force`
|
|
216
208
|
- `ds verify <id>` - Scan an unverified data source's schema and review it, and activate it. Warehouse/SQL sources print a verification link; add --accept-schema to accept the AI schema and activate from the CLI. File uploads activate on this command without --accept-schema, but NOT on create: `ds create --file` leaves them at pending_verification until you run this. Requires data_source_write in the source's domain. - `--force --accept-schema`
|
|
217
209
|
- `ds update <id>` - Update a data source row cap - `--max-rows`
|
|
@@ -253,4 +245,4 @@ _Generated from the CLI by `npm run gen:commands` - do not hand-edit between the
|
|
|
253
245
|
**setup** - Install legacy copied Graphit assistant files for Cursor or fallback setups
|
|
254
246
|
- `setup` - Install legacy copied Graphit assistant files for Cursor or fallback setups - `--editor --project --update --legacy-copy --remove-legacy-copies --dry-run`
|
|
255
247
|
|
|
256
|
-
<!-- COMMANDS:END -->
|
|
248
|
+
<!-- COMMANDS:END -->
|
|
@@ -0,0 +1,27 @@
|
|
|
1
|
+
# Data-Source Refresh
|
|
2
|
+
|
|
3
|
+
Load when configuring refresh mode, incremental behavior, history, or reconciliation.
|
|
4
|
+
|
|
5
|
+
## Modes
|
|
6
|
+
|
|
7
|
+
- **full:** replace the cached result from the complete source query.
|
|
8
|
+
- **incremental:** append/merge new source rows using a watermark and stable merge key.
|
|
9
|
+
|
|
10
|
+
Choose incremental only when the source exposes a reliable monotonic watermark and the merge key is unique. Otherwise use full refresh.
|
|
11
|
+
|
|
12
|
+
## Incremental contract
|
|
13
|
+
|
|
14
|
+
- Filter the source early on the watermark column.
|
|
15
|
+
- Preserve lookback for late-arriving updates.
|
|
16
|
+
- Provide the correct watermark type.
|
|
17
|
+
- Verify merge-key uniqueness before serving the new version.
|
|
18
|
+
- Reconcile periodically with a full rebuild.
|
|
19
|
+
- Treat schema drift or grain change as a semantic review, not a blind refresh.
|
|
20
|
+
|
|
21
|
+
## Operations
|
|
22
|
+
|
|
23
|
+
A refresh request may be asynchronous. Poll the job/source status and report counts, duration, version, and failure truthfully. Fire-and-forget is appropriate only when the user did not ask to wait.
|
|
24
|
+
|
|
25
|
+
Inspect refresh history before retrying. A failed status may follow a partially applied external action; use the receipt/status rather than assuming nothing happened.
|
|
26
|
+
|
|
27
|
+
Refresh settings and connector lifecycle require data-source write authority in the server's policy key. Never expose credentials, source rows, or concealed schema in evidence.
|
|
@@ -1,134 +1,39 @@
|
|
|
1
|
-
# Data Sources
|
|
1
|
+
# Data Sources
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
Load when selecting or creating the cached source a semantic model uses.
|
|
4
4
|
|
|
5
|
-
## Routing
|
|
5
|
+
## Routing
|
|
6
6
|
|
|
7
|
-
|
|
7
|
+
1. Read the semantic model's declared data-source binding.
|
|
8
|
+
2. Prefer that cached source for speed, governance, and repeatability.
|
|
9
|
+
3. Use metadata discovery when physical columns are unknown.
|
|
10
|
+
4. Query live warehouse only when no cached source covers the question and the user approves.
|
|
11
|
+
5. Never infer a source from a similarly named model.
|
|
8
12
|
|
|
9
|
-
|
|
10
|
-
|---|---|---|
|
|
11
|
-
| The table has a cached data source | `graphit query "SQL" --ds <NAME>` | roughly 100ms, DuckDB |
|
|
12
|
-
| No data source covers the table | `graphit query "SQL" --warehouse --connection <id>` | roughly 10s, the connected warehouse |
|
|
13
|
+
A group is semantic placement. Data-source creation still accepts `--domain`; pass the uppercase policy key returned by status or the group's `domain_keys`.
|
|
13
14
|
|
|
14
|
-
|
|
15
|
+
## Source SQL
|
|
15
16
|
|
|
16
|
-
|
|
17
|
+
- Select only needed columns and rows.
|
|
18
|
+
- Filter early using base-table columns.
|
|
19
|
+
- Avoid wrapping filter columns when a direct predicate works.
|
|
20
|
+
- Keep complete executable SQL; no ellipses, fake tables, or embedded data.
|
|
21
|
+
- Preserve warehouse dialect.
|
|
22
|
+
- Make grain and refresh mode explicit.
|
|
23
|
+
- Use merge key and watermark only when the source supports them.
|
|
17
24
|
|
|
18
|
-
##
|
|
25
|
+
## Creation
|
|
19
26
|
|
|
20
|
-
|
|
27
|
+
Confirm connector, relation/query, policy key, grain, refresh mode, and cost. Read columns through metadata rather than probing with ad-hoc SQL.
|
|
21
28
|
|
|
22
|
-
|
|
23
|
-
|---|---|---|
|
|
24
|
-
| Grain | `GROUP BY` to the grain you chart (the single biggest lever) | One row per raw event |
|
|
25
|
-
| Columns | Only the columns dashboards use (each is downloaded + cached per query) | 400+ columns "just in case" |
|
|
26
|
-
| Cardinality | Low-card dimensions in the base; ad/campaign names in a separate drill-down | Thousands-of-values dimensions in the base grain |
|
|
27
|
-
| Size | A few-thousand-row typical aggregation | Sitting at the 100M-row / 5GB ceiling |
|
|
29
|
+
Create with automatic scan unless there is a specific reason not to. Creation may be asynchronous; report `creating` honestly and poll status rather than claiming readiness.
|
|
28
30
|
|
|
29
|
-
##
|
|
31
|
+
## Access and safety
|
|
30
32
|
|
|
31
|
-
|
|
33
|
+
- Changing a source requires data-source write capability in its policy key.
|
|
34
|
+
- Reading does not imply authority over connector, SQL, or refresh settings.
|
|
35
|
+
- Visibility and masking cover agent, canvas, render, export, and report paths.
|
|
36
|
+
- Private names and columns remain concealed.
|
|
37
|
+
- Delete/move stay in the Sources Hub where cascades are visible.
|
|
32
38
|
|
|
33
|
-
|
|
34
|
-
|---|---|---|
|
|
35
|
-
| Raw passthrough | No `GROUP BY` / no aggregate - one row per raw event | Pre-aggregate to the grain the dashboards chart |
|
|
36
|
-
| Very wide | Far more columns than dashboards use (e.g. `SELECT *`) | Select only the columns dashboards need |
|
|
37
|
-
| High-cardinality grain | A dimension with thousands of distinct values (ad / campaign / user ids) | Keep it out of the base; build a separate drill-down source |
|
|
38
|
-
| Large + monolithic | A big source whose typical query still scans most rows | Pre-aggregate and narrow so typical queries touch a few thousand rows |
|
|
39
|
-
|
|
40
|
-
A wide or raw source is sometimes the right call - row-level drill-down/export, columns genuinely all used, or a staging source to reshape later. Note the trade-off, then build whichever the user chooses.
|
|
41
|
-
|
|
42
|
-
## Creating data sources
|
|
43
|
-
|
|
44
|
-
`graphit ds create` auto-chains: create -> poll until ready -> scan schema -> print verification link. Activating it for KB use is the `ds verify` step below.
|
|
45
|
-
|
|
46
|
-
```bash
|
|
47
|
-
# Create with auto-scan (recommended)
|
|
48
|
-
graphit ds create --name "MY_DS" --domain <DOMAIN> --sql "SELECT ..." --connection <id>
|
|
49
|
-
|
|
50
|
-
# Create without auto-scan (for special cases)
|
|
51
|
-
graphit ds create --name "MY_DS" --domain <DOMAIN> --sql "SELECT ..." --skip-scan
|
|
52
|
-
```
|
|
53
|
-
|
|
54
|
-
**`--domain` is REQUIRED, in both modes.** It files the scanned table under an existing KB domain and decides who can see the source; there is no uncategorized fallback. `graphit kb list domains`, confirm the choice with the user, and `graphit kb create domain --name <NAME>` if none fits.
|
|
55
|
-
|
|
56
|
-
**From a local file (Excel/CSV):** `graphit ds create --file <path> --domain <NAME>` uploads the file and creates one data source. Optional: `--name` (defaults to the file name), `--sheet <name>` (multi-sheet workbooks). `--file` and `--sql` are mutually exclusive; same flow as above.
|
|
57
|
-
|
|
58
|
-
**Warehouse connection.** `--connection` names the warehouse a `--sql` source reads from. Add BigQuery with `graphit connector add bigquery-serviceaccount --key-file <path> [--project --dataset --location]` (org admin; project defaults from the key). The pipeline routes by connection type - the same `ds create` works for either warehouse.
|
|
59
|
-
|
|
60
|
-
For existing unverified sources, `graphit ds verify <id>` scans and shows the schema; add `--accept-schema` to accept the AI schema and activate a warehouse/SQL source from the CLI. A file upload needs `ds verify` too - it activates without `--accept-schema`, but never at create time, so it stays unqueryable until you run it.
|
|
61
|
-
|
|
62
|
-
## Refreshing data sources
|
|
63
|
-
|
|
64
|
-
Data sources cache a snapshot of the warehouse query result. Refresh when you need current data. **File-upload sources can't be refreshed - update them by re-uploading with `graphit ds create --file <path> --domain <NAME>`.**
|
|
65
|
-
|
|
66
|
-
On BigQuery a refresh scans billed bytes, so keep the shape tight and prefer incremental/partition-pruned refresh over full re-scans; a per-connection scan cap (max bytes billed) fails an oversized query fast rather than running up a bill.
|
|
67
|
-
|
|
68
|
-
```bash
|
|
69
|
-
# Refresh all data sources and wait for completion (live status table)
|
|
70
|
-
graphit ds refresh --all
|
|
71
|
-
|
|
72
|
-
# Fire-and-forget (trigger refreshes, don't wait)
|
|
73
|
-
graphit ds refresh --all --no-wait
|
|
74
|
-
|
|
75
|
-
# Refresh specific sources by ID
|
|
76
|
-
graphit ds refresh <id1> <id2>
|
|
77
|
-
```
|
|
78
|
-
|
|
79
|
-
`graphit ds refresh` only runs an **incremental** refresh (new rows since last update); it never re-exports the whole source. A full rebuild is UI-only (Sources -> Refresh -> Full rebuild); hand off to the UI if one is needed.
|
|
80
|
-
|
|
81
|
-
Refreshes fire in parallel; polls to completion (large sources 30-60s), or returns at once with `--no-wait` (check status via `graphit ds list`). Governed per org: manual refreshes have an hourly budget and a limited number run at once (a reserved slot keeps manual ones unblocked by scheduled). A limit returns a 429 with a clear reason (reset time, or "wait for running operations to finish"); wait it out, don't loop - not source errors. Review past runs with `graphit ds refresh-history <id>`.
|
|
82
|
-
|
|
83
|
-
## Incremental refresh and early-filtering (advanced)
|
|
84
|
-
|
|
85
|
-
Incremental mode fetches only rows past a watermark and merges them in. Three windows govern it: the **watermark column** (which output rows are new), the **merge window** (`--merge-window` - how far back each run re-fetches and upserts, healing late data; API responses call it `lookback_periods`), and per-table **lookback windows** (`--table-lookback` - how far back each source table is *read*). Set on a scanned source; each call sets the COMPLETE config - omitted flags reset to defaults (no `--table-lookback` = windows cleared).
|
|
86
|
-
|
|
87
|
-
When a source aggregates over a wide internal window (e.g. a multi-year rollup), incremental refresh is nearly as slow as full: the outer watermark filter can't prune the inner scan. Early-filtering fixes that - get the contract right first, or older periods silently corrupt on merge:
|
|
88
|
-
|
|
89
|
-
- Size each lookback to cover the merge window plus the longest rolling calculation in the query.
|
|
90
|
-
- Set `--merge-key` (upsert) - overlapping rows double-count without one.
|
|
91
|
-
- Rolling-window metrics (WAU / MAU / stickiness): prefer a layered base daily source instead - recent rows alone can't compute a rolling window.
|
|
92
|
-
- `--reconciliation` is the periodic full-refresh drift backstop (default off; keep it on with the bind or in append mode).
|
|
93
|
-
|
|
94
|
-
Then pick ONE early-filter mode (mutually exclusive, validated):
|
|
95
|
-
|
|
96
|
-
**Per-table lookback windows - preferred; no SQL edit.** Declare how far back each table is read; the engine prunes delta scans and the SQL stays exactly as written. Day-based (date/timestamp watermark required).
|
|
97
|
-
|
|
98
|
-
```bash
|
|
99
|
-
graphit ds refresh-config <id> --mode incremental \
|
|
100
|
-
--watermark-column EVENT_DATE --watermark-type date --merge-key ID \
|
|
101
|
-
--merge-window 3 --table-lookback ANALYTICS.EVENTS:EVENT_DATE:30 \
|
|
102
|
-
--table-lookback USERS:CREATED_AT:90
|
|
103
|
-
```
|
|
104
|
-
|
|
105
|
-
**`:graphit_watermark` bind - for SQL owners.** Place the token where the filter belongs; Graphit substitutes the last watermark on deltas, full history on reconciliation. Output-filter to the current fully-covered period.
|
|
106
|
-
|
|
107
|
-
Only early-filter when an incremental source is slow for this reason; the default refresh is correct and simpler otherwise.
|
|
108
|
-
|
|
109
|
-
## What needs write access
|
|
110
|
-
|
|
111
|
-
Querying a source, listing sources, reading schema or refresh history, and an ordinary `graphit ds refresh` are reads - any member who can read that source's domain can run them, and a source in a domain they cannot read returns the same uniform 404 as one that does not exist. These need `data_source_write` in the source's domain: `ds create`, editing its SQL, `ds refresh-config`, a `--force` refresh or accepting a schema, `ds verify`, scanning, per-source governance settings, and deletion. Moving a source to another domain needs write in both the old and the new domain.
|
|
112
|
-
|
|
113
|
-
Check `graphit status` for those domains before proposing a create or a config change. It is advisory - the server authorizes each operation when it runs, and a denial with `retryable: false` is a stop, not a retry (`operations.md`).
|
|
114
|
-
|
|
115
|
-
## Deleting data sources
|
|
116
|
-
|
|
117
|
-
`ds delete` is not available on the CLI. Deleting a data source cascades to storage and the KB table, removing all metrics, dimensions, and rules on it. Direct the user to the platform UI (Sources Hub), whose confirmation flow shows what will be affected.
|
|
118
|
-
|
|
119
|
-
## Presenting data source results
|
|
120
|
-
|
|
121
|
-
The user cannot see raw CLI output - you are the rendering layer. After `graphit ds list`, present a markdown table and end with a recommendation of which source to use (or note none covers the needed table):
|
|
122
|
-
|
|
123
|
-
~~~
|
|
124
|
-
**2 data sources:**
|
|
125
|
-
|
|
126
|
-
| Name | ID | Rows | Status | Governed |
|
|
127
|
-
|---|---|---:|---|---|
|
|
128
|
-
| **MARKETING_UA_DS** | ds_abc123 | 1,247,832 | active | yes |
|
|
129
|
-
| **REVENUE_EVENTS** | ds_ghi789 | 3,412,006 | stale | yes |
|
|
130
|
-
|
|
131
|
-
Using **MARKETING_UA_DS** (ds_abc123), which covers spend, installs, and ROAS columns.
|
|
132
|
-
~~~
|
|
133
|
-
|
|
134
|
-
Bold every data source name. If a source is stale, say so and offer to refresh it before querying. If `truncated` is true, raise `--limit` before recommending.
|
|
39
|
+
For refresh modes, history, incremental tuning, and reconciliation, load `data-source-refresh.md`.
|
|
@@ -51,13 +51,13 @@ Top values by a measure, shaped to drop straight into an array filter or an `IN
|
|
|
51
51
|
```js
|
|
52
52
|
const top = await graphit.rank({
|
|
53
53
|
column: 'COUNTRY', source: 'sales', dataSourceId: 'SALES_DS',
|
|
54
|
-
by:
|
|
54
|
+
by: "{{ Metric('revenue') }}", // governed metric, or a bare aggregate
|
|
55
55
|
limit: 10,
|
|
56
56
|
filters: { REGION: region.get() }, // optional, same contract as cascade
|
|
57
57
|
})
|
|
58
58
|
```
|
|
59
59
|
|
|
60
|
-
- Returns a plain array of values. Prefer
|
|
60
|
+
- Returns a plain array of values. Prefer `{{ Metric('name') }}` so ranking uses the org's definition; a bare aggregate accepts SUM, COUNT, AVG, MIN, or MAX over one column.
|
|
61
61
|
|
|
62
62
|
## graphit.dateRange(id, options) - Date Presets
|
|
63
63
|
|
|
@@ -1,40 +1,29 @@
|
|
|
1
|
-
# Explaining
|
|
1
|
+
# Explaining Governance
|
|
2
2
|
|
|
3
|
-
Load
|
|
3
|
+
Load when the user asks what governed means, why a query changed, or why it was blocked.
|
|
4
4
|
|
|
5
|
-
## The
|
|
5
|
+
## The idea
|
|
6
6
|
|
|
7
|
-
|
|
7
|
+
The semantic layer defines a business concept. Governance decides how that definition may be queried for this user and context.
|
|
8
8
|
|
|
9
|
-
##
|
|
9
|
+
## What happens
|
|
10
10
|
|
|
11
|
-
|
|
11
|
+
1. Graphit resolves `Metric`, qualified `Dimension`, and its `Measure` extension.
|
|
12
|
+
2. It finds rules through model, entity, dimension, metric, and group targets.
|
|
13
|
+
3. Verified constraints are injected before execution.
|
|
14
|
+
4. Resolved SQL is verified fail-closed.
|
|
15
|
+
5. The transparency receipt records what changed and why.
|
|
12
16
|
|
|
13
|
-
|
|
14
|
-
|------|-------|-------|
|
|
15
|
-
| governed | Used `{{metric:X}}` / `{{dim:X}}` KB references | Teal |
|
|
16
|
-
| verified | Raw SQL whose math matches a KB definition | Amber |
|
|
17
|
-
| ad_hoc | Inline formula with no KB match | Gray |
|
|
17
|
+
A user may not see concealed targets, but still receives a generic safe refusal when access cannot be proven.
|
|
18
18
|
|
|
19
|
-
|
|
19
|
+
## Rule states
|
|
20
20
|
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
| Unsealed dashboard | A dashboard uses definitions the KB does not govern |
|
|
27
|
-
| Verified, not wired | A verified definition is re-implemented as raw SQL |
|
|
28
|
-
| Unverified in use | Live graphs depend on a definition nobody verified |
|
|
29
|
-
| Stale reference | A graph is frozen on an old version of its definition (the only "Breaking" card) |
|
|
30
|
-
| Deprecated in use | A deprecated definition is still used by live graphs |
|
|
31
|
-
| Conflicting definitions | Two near-identical definitions disagree |
|
|
32
|
-
| Rule conflict | Two governance rules contradict each other |
|
|
33
|
-
| Unenforced rule | A rule that structurally cannot act on anything |
|
|
34
|
-
| Missing owner | Assets with nobody responsible for them |
|
|
35
|
-
|
|
36
|
-
Fix is a guided one-click resolution: it drafts and previews the exact change, you apply it (same validation as a manual edit), then Graphit re-checks and clears the card. This queue lives only in the **web app's Governance page** - the CLI cannot list or fix Insights cards (`governance status` and `governance audit` report conformance only). When a user asks about a finding, explain what it means and point them to the Governance page to Fix it.
|
|
21
|
+
- Draft: no effect.
|
|
22
|
+
- Verified body-only: guidance.
|
|
23
|
+
- Verified constrained: enforceable according to mode.
|
|
24
|
+
- Protected-column masks: never overridable.
|
|
25
|
+
- EXPLORE: bounded override where policy permits.
|
|
37
26
|
|
|
38
27
|
## Relaying it
|
|
39
28
|
|
|
40
|
-
|
|
29
|
+
State the result first, then tier and receipt facts. Name only visible rules. Never invent hidden names or quote raw compiler errors. If blocked, give the server-provided next step. If ad hoc, say so and offer a governed rewrite or supported definition.
|
|
@@ -1,100 +1,37 @@
|
|
|
1
1
|
# Query Governance
|
|
2
2
|
|
|
3
|
-
Load
|
|
3
|
+
Load when writing a governed query, explaining a refusal, or reporting provenance.
|
|
4
4
|
|
|
5
|
-
|
|
5
|
+
## References
|
|
6
6
|
|
|
7
|
-
|
|
7
|
+
| Reference | Meaning |
|
|
8
|
+
|---|---|
|
|
9
|
+
| `{{ Metric('revenue') }}` | Reusable metric |
|
|
10
|
+
| `{{ Dimension('order__channel') }}` | Qualified grouping/filter field |
|
|
11
|
+
| `{{ Measure('order_total') }}` | Graphit's model-owned measure extension |
|
|
8
12
|
|
|
9
|
-
|
|
13
|
+
Legacy token grammar is refused. Keep references inside complete executable SQL and canvas `data-graphit-sql`.
|
|
10
14
|
|
|
11
|
-
|
|
12
|
-
|--------|-----------|---------|
|
|
13
|
-
| `{{metric:NAME}}` | Metric calculation (aggregation) | `{{metric:CPI}}` |
|
|
14
|
-
| `{{metric:NAME(K=V)}}` | Parameterized metric | `{{metric:ARPU(DAY=7)}}` |
|
|
15
|
-
| `{{metric_raw:NAME}}` | Raw expression, no outer aggregate | `{{metric_raw:REVENUE}}` |
|
|
16
|
-
| `{{dim:NAME}}` | Dimension expression | `{{dim:INSTALL_MONTH}}` |
|
|
17
|
-
|
|
18
|
-
```bash
|
|
19
|
-
graphit query "SELECT {{dim:INSTALL_MONTH}}, {{metric:CPI}} AS cpi FROM MARKETING_UA_DS GROUP BY 1" --ds MARKETING_UA_DS --verbose
|
|
20
|
-
```
|
|
21
|
-
|
|
22
|
-
`--verbose` prints the expanded SQL and trust tier, so you can confirm the reference resolved before presenting the result.
|
|
23
|
-
|
|
24
|
-
## Parameterized metrics
|
|
25
|
-
|
|
26
|
-
Some metrics (for example ARPU, ROAS, RETENTION) carry required parameters and cannot resolve without a value. Run `graphit kb list metric` and read the `params` column for the names a metric requires, then supply them inline as `{{metric:ARPU(DAY=7)}}`. Pre-baked variants such as `ARPU_D7` or `ROAS_D30` have the value fixed and need none. Omitting a required parameter returns a clear error naming the exact syntax, so read `params` first.
|
|
15
|
+
The governed fragment path serves simple, ratio, and derived metrics. Cumulative, conversion, shifted, time-spine, and null-fill shapes remain unavailable until Project #289. Use a supported decomposition or explicitly labeled free SQL.
|
|
27
16
|
|
|
28
17
|
## Trust tiers
|
|
29
18
|
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|------|---------|-------|
|
|
34
|
-
| `governed` | Query used `{{metric:X}}` / `{{dim:X}}` references | Teal dot |
|
|
35
|
-
| `verified` | Raw SQL whose expressions match KB definitions | Amber dot |
|
|
36
|
-
| `ad_hoc` | Inline formulas with no KB match | Gray dot |
|
|
37
|
-
|
|
38
|
-
Prefer the governed tier. The server may upgrade matching raw SQL to verified, but reach for references first so the result is governed by intent.
|
|
39
|
-
|
|
40
|
-
## The ad-hoc gate
|
|
41
|
-
|
|
42
|
-
This is the hard frontier - the rules below are what the QueryGateway does.
|
|
43
|
-
|
|
44
|
-
A query lands at the `ad_hoc` tier when it uses no `{{metric:NAME}}` / `{{dim:NAME}}` reference and its raw expressions match no KB definition. On the CLI (`graphit query`, cached `--ds` or `--warehouse`), **every** ad-hoc query must be justified - business measures (an aggregate or `GROUP BY`) and plain exploration alike (`SELECT *`, `COUNT(*)`, `DISTINCT` peeks). Governed and verified results are exempt.
|
|
45
|
-
|
|
46
|
-
- **EXPLORE access is a hard prerequisite for ad-hoc business measures.** If the user lacks EXPLORE access to any queried table scope, the server blocks that measure. An ad-hoc reason cannot bypass that denial.
|
|
47
|
-
- **Justification floor.** The ad-hoc query is withheld and asks for a justification. Preferred path first: rewrite with `{{metric:NAME}}` / `{{dim:NAME}}` references - genuinely search the KB (`graphit kb explore`, `graphit kb list metric`, `graphit kb list dimension`), and if the metric or dimension you need does not exist, CREATE it first. A filter's value list belongs in a `{{dim:NAME}}`, not a raw `SELECT DISTINCT` peek. Only if nothing fits and the user needs the raw run, pass `--adhoc-reason "<text>"` stating what you searched, what you found, and why it does not fit - pass it on the first call when you already know the query is ad-hoc. A trivial or empty reason is rejected server-side and recorded in the audit log, so it must be honest.
|
|
48
|
-
|
|
49
|
-
When the gate fires, do not narrate around it or pretend the query ran. Report truthfully: it was blocked or needs approval, name the governed rewrite, and let the user decide.
|
|
50
|
-
|
|
51
|
-
## Enforceable rules and overrides
|
|
52
|
-
|
|
53
|
-
Rules with typed constraints are enforced automatically, rewriting the SQL before it runs:
|
|
54
|
-
|
|
55
|
-
| Type | What it does |
|
|
56
|
-
|------|-------------|
|
|
57
|
-
| `required_where` | Injects a WHERE predicate |
|
|
58
|
-
| `forbidden_column` | NULLifies a column in SELECT |
|
|
59
|
-
| `value_restriction` | Restricts a column to allowed values |
|
|
60
|
-
| `required_filter` | Validates a column appears in WHERE |
|
|
61
|
-
| `required_aggregation` | Validates GROUP BY includes a column |
|
|
62
|
-
|
|
63
|
-
User-context variables (`${user.team_id}`, `${user.email}`) resolve server-side for row-level security. Override a rule only when the user explicitly asks; the server honors it only if the user holds EXPLORE on every queried scope. A rule that masks a column (a `forbidden_column` constraint) can never be overridden. Every override is logged. Pass several names to override more than one.
|
|
64
|
-
|
|
65
|
-
```bash
|
|
66
|
-
graphit query "SELECT * FROM EVENTS" --ds EVENTS --override-rules EXCLUDE_RETARGETING
|
|
67
|
-
```
|
|
68
|
-
|
|
69
|
-
## Conditionally-enforced rules
|
|
70
|
-
|
|
71
|
-
A rule's mode is Advisory (guidance), Always (every query), or **Conditional** - fires only when its plain-language body (the condition) holds for the query you wrote. No server classifier decides that; you do. Read a table's rules first (`graphit kb explore table <name>`), judge your query against each conditional rule's body, and declare it up front:
|
|
72
|
-
|
|
73
|
-
- `--apply-conditional RULE` - enforce it for this query.
|
|
74
|
-
- `--skip-conditional RULE:"reason"` - skip it; a reason is required and audit-logged.
|
|
75
|
-
|
|
76
|
-
Declare up front so a clean query never stalls. An undeclared conditional returns a retryable prompt naming each unresolved rule and its condition - read it, decide, re-run. A saved tile stores your decision and replays it every refresh.
|
|
77
|
-
|
|
78
|
-
## Data-source row caps
|
|
79
|
-
|
|
80
|
-
An admin can set a `max_rows` cap per data source. A cap limits the result set; it does not change authorization, trust-tier classification, or the ad-hoc justification floor.
|
|
19
|
+
- **governed:** verified semantic references compiled through the gateway.
|
|
20
|
+
- **verified:** known safe stored query without semantic references.
|
|
21
|
+
- **ad hoc:** raw SQL at the frontier.
|
|
81
22
|
|
|
82
|
-
|
|
23
|
+
Prefer governed. Never present ad-hoc SQL as the team's definition.
|
|
83
24
|
|
|
84
|
-
|
|
25
|
+
## Rules
|
|
85
26
|
|
|
86
|
-
|
|
27
|
+
Rules target model, entity, dimension, metric, or group identities. Verified constraints enforce; verified body-only rules guide; drafts do nothing. Modes and EXPLORE behavior remain server-owned.
|
|
87
28
|
|
|
88
|
-
|
|
89
|
-
**Trust tier:** governed - 2 KB refs, 1 rule enforced (**EXCLUDE_INTERNAL**), max rows 10000
|
|
90
|
-
~~~
|
|
29
|
+
The gateway runs before caches, injects constraints, verifies resolved SQL, and returns a transparency receipt. Do not claim a rule applied merely because it exists.
|
|
91
30
|
|
|
92
|
-
|
|
31
|
+
## Ad-hoc gate
|
|
93
32
|
|
|
94
|
-
|
|
33
|
+
Search the KB genuinely, explain why visible definitions do not fit, and prefer an approved reusable supported definition. Use a truthful ad-hoc reason only for a real one-off. Never use it to bypass a rule.
|
|
95
34
|
|
|
96
|
-
|
|
97
|
-
**Blocked by governance.**
|
|
35
|
+
## Reporting
|
|
98
36
|
|
|
99
|
-
|
|
100
|
-
~~~
|
|
37
|
+
Report tier, semantic references, row cap, visible rules that changed the query, and any refusal or override. Read the receipt rather than inferring. A blocked or partial result is not success.
|