@malloy-publisher/server 0.0.243 → 0.0.245

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (26) hide show
  1. package/README.md +47 -5
  2. package/dist/app/api-doc.yaml +445 -21
  3. package/dist/app/assets/{EnvironmentPage-zu0NeDu-.js → EnvironmentPage-EjXlK1AR.js} +1 -1
  4. package/dist/app/assets/{HomePage-kuBM3jiQ.js → HomePage-BDg82oe3.js} +1 -1
  5. package/dist/app/assets/{LightMode-DGBAdVPb.js → LightMode-CM8J0nPB.js} +1 -1
  6. package/dist/app/assets/{MainPage-CCVx9yPs.js → MainPage-CkkdimtC.js} +1 -1
  7. package/dist/app/assets/{MaterializationsPage-M-bWIx8L.js → MaterializationsPage-DvVS2HMi.js} +1 -1
  8. package/dist/app/assets/{ModelPage-D-L2pqIy.js → ModelPage-Bz2je31A.js} +1 -1
  9. package/dist/app/assets/{PackagePage-BZglxDlj.js → PackagePage-BUJU3qB3.js} +1 -1
  10. package/dist/app/assets/{RouteError-DGoCQGam.js → RouteError-oqZTfMol.js} +1 -1
  11. package/dist/app/assets/{ThemeEditorPage-e8MLf7SV.js → ThemeEditorPage-W-V8hngE.js} +1 -1
  12. package/dist/app/assets/{WorkbookPage-C8ufpN5i.js → WorkbookPage-IQFxsQGR.js} +1 -1
  13. package/dist/app/assets/{core-KFsKleo0.es-D6cYhbct.js → core-BCwj9ZvB.es-C610H3gu.js} +1 -1
  14. package/dist/app/assets/{index-CDNkCta4.js → index-1xafiaYB.js} +1 -1
  15. package/dist/app/assets/{index-Dnef41hg.js → index-BUhTsFh9.js} +1 -1
  16. package/dist/app/assets/{index-B-jw9EkA.js → index-DJ-3RN5W.js} +4 -4
  17. package/dist/app/assets/{index-CLwzlAx6.js → index-Dxotbpnc.js} +1 -1
  18. package/dist/app/index.html +1 -1
  19. package/dist/instrumentation.mjs +10 -9
  20. package/dist/package_load_worker.mjs +209 -188
  21. package/dist/server.mjs +107148 -105496
  22. package/dist/sshcrypto-vd2k5hq9.node +0 -0
  23. package/package.json +4 -2
  24. package/publisher.config.example.bigquery.json +33 -0
  25. package/publisher.config.example.duckdb.json +23 -0
  26. package/dist/sshcrypto-8m50vnmb.node +0 -0
package/README.md CHANGED
@@ -2,11 +2,53 @@
2
2
 
3
3
  The Malloy Publisher Server is an Express.js server that provides an API for managing and accessing Malloy data models, packages, and queries
4
4
 
5
+ ## Quick start
6
+
7
+ This section is self-contained: this README ships in the npm tarball, where most of the repository does not, so nothing here needs a file you do not have. Later sections are written for a clone and do point into the repository.
8
+
9
+ **New workspace?** One command scaffolds a package, the server config, the MCP wiring, and the agent skills:
10
+
11
+ ```bash
12
+ npm create @malloy-publisher/malloy-package@latest sales
13
+ npm start
14
+ ```
15
+
16
+ Run bare like that, the package comes with a small sample dataset, so there is something to query immediately. To start from a file of your own instead, add `-- --data ./orders.csv`, naming a delimited or Parquet/Excel file you actually have (the `--` is required, and the path is relative to where you run the command). Keep the `@latest`: `npm create` resolves through the npx cache, and an unversioned name silently reuses whatever old copy is on the machine.
17
+
18
+ **Existing directory of models?** Write a `publisher.config.json` beside a package directory that holds a `publisher.json`, then run the server pointing at it:
19
+
20
+ ```bash
21
+ npx @malloy-publisher/server --server_root . --watch-env local
22
+ ```
23
+
24
+ ```json
25
+ {
26
+ "frozenConfig": false,
27
+ "environments": [
28
+ {
29
+ "name": "local",
30
+ "packages": [{ "name": "sales", "location": "./sales" }],
31
+ "connections": []
32
+ }
33
+ ]
34
+ }
35
+ ```
36
+
37
+ The config is read from `<server_root>/publisher.config.json` by default; `--config <path>` overrides it.
38
+
39
+ Three facts that are easy to get wrong:
40
+
41
+ - **A flat-file (CSV/Parquet/XLSX) package needs no `connections` entry at all.** Every loaded package automatically gets its own DuckDB sandbox connection named `duckdb`, which is what `duckdb.table('data/file.csv')` resolves against. That name is reserved: an environment-level connection named `duckdb` fails the whole environment at init (call an env-level DuckDB connection `shared_duckdb` or similar).
42
+ - **A package `location` is local only when it starts with `./`, `../`, `~/`, or `/`.** Write `./sales`, not `sales`: a bare name is read as a remote URI, and the package is skipped with `Invalid package path: "sales". Must be an absolute mounted path or a GCS/S3/GitHub URI.` in `loadErrors`. A `./` or `../` path resolves against the config file's directory; `~/` against your home directory.
43
+ - **Local authoring means `--watch-env <env>`.** Without it the server copies each local package into `publisher_data/` at boot and serves the copy, so edits to your source directory are never read — a reload recompiles the copy and still answers 200. Adding the flag to a later boot does not undo a copy already made: pass `--watch-env <env> --init` once to re-mount (it wipes `publisher_data/` and re-syncs from config), then plain `--watch-env` boots keep watching. `publisher_data/<env>/<pkg>` tells you which you got — a symlink is mounted, a real directory is a copy.
44
+
45
+ Poll `curl -s http://localhost:4000/api/v0/status` until `operationalState` is `"serving"`, then check `loadErrors` (absent when everything loaded). The MCP endpoint for agents is `http://localhost:4040/mcp`.
46
+
5
47
  ## Configuration
6
48
 
7
- `publisher.config.json` lives in this directory. The repository root also contains a symlink (`/publisher.config.json` → `./packages/server/publisher.config.json`) so that running the server from either location picks up the same config. Edit one and you've edited both.
49
+ Two example configs ship in the npm package, beside this README: [`publisher.config.example.duckdb.json`](./publisher.config.example.duckdb.json) (the GitHub-hosted sample packages, no connection block) and [`publisher.config.example.bigquery.json`](./publisher.config.example.bigquery.json) (adds a BigQuery connection). Copy either one to `publisher.config.json` and point `--server_root` at its directory.
8
50
 
9
- For the BigQuery-enabled variant, see [`publisher.config.example.bigquery.json`](./publisher.config.example.bigquery.json) and the [Quick Start in the repo root README](../../README.md#quick-start).
51
+ In a clone, the live `publisher.config.json` lives in this directory, and the repository root contains a symlink (`/publisher.config.json` → `./packages/server/publisher.config.json`) so that running the server from either location picks up the same config. Edit one and you've edited both.
10
52
 
11
53
  ### Remaining deprecation warnings
12
54
 
@@ -15,11 +57,11 @@ Removing the unused `trino` CLI direct dep (it pulled in `@google-cloud/translat
15
57
  - **npm CLI tooling**: `npmlog`, `gauge`, `are-we-there-yet`, `glob@7/8/10`, `rimraf@3`, `tar@6.2.1`, `inflight`, `@npmcli/move-file`, `node-domexception`, `querystring` — pulled in by npm itself and by `node-pre-gyp`/`node-gyp`. Not actionable from this repo.
16
58
  - **`uuid@8.x` / `uuid@9.x`**: surfaced across multiple transitives (Malloy, AWS SDKs, others). Resolves when each upstream bumps to `uuid@11`.
17
59
  - **`q@1.5.1`**: pulled in via `thrift` → `@databricks/sql` → `@malloydata/db-databricks`. Resolves when Databricks upgrades `@databricks/sql` past the thrift dep, or when we replace the Databricks driver.
18
- - **`aws-sdk@2.1693.0`**: listed as a direct dep in `packages/server/package.json` but not imported anywhere in source — leftover, candidate for removal in a follow-up PR. The actual consumer is `@aws-sdk/client-s3` v3.
60
+ - **`aws-sdk@2.1693.0`**: no longer a direct dep (removed from `packages/server/package.json`); anything still surfacing it is transitive. The actual S3 consumer is `@aws-sdk/client-s3` v3.
19
61
 
20
62
  ## K6 Test Presets
21
63
 
22
- The Malloy Publisher Server includes several K6 test presets to help you test its performance and stability.
64
+ The Malloy Publisher Server includes several K6 test presets to help you test its performance and stability. These live in the repository, not the npm package, so run them from a clone.
23
65
 
24
66
  Below is a list of the available test presets:
25
67
 
@@ -124,7 +166,7 @@ For more information on how to configure OpenTelemetry collectors, please refer
124
166
 
125
167
  ## MCP Prompt Capability
126
168
 
127
- Publisher's MCP interface exposes the bundled agent **skills** (under [`skills/`](../../skills/)) as **LLM-ready prompts**, so hosts that ingest MCP but do not load skill files can pull the same guidance. For authoring or contributing skills, see [`docs/agent-skills/`](../../docs/agent-skills/).
169
+ Publisher's MCP interface exposes the bundled agent **skills** as **LLM-ready prompts**, so hosts that ingest MCP but do not load skill files can pull the same guidance. The skills are published separately as [`@malloy-publisher/skills`](https://www.npmjs.com/package/@malloy-publisher/skills); their source is [`skills/`](https://github.com/malloydata/publisher/tree/main/skills), and authoring guidance is in [`docs/agent-skills/`](https://github.com/malloydata/publisher/tree/main/docs/agent-skills).
128
170
 
129
171
  List prompts:
130
172
 
@@ -2530,10 +2530,12 @@ components:
2530
2530
  description:
2531
2531
  Configured packages and environments that did not load, and why. A
2532
2532
  load failure is not fatal to the server, so operationalState stays
2533
- serving and the affected packages are simply absent from
2534
- environments; this is the only place that difference is reported, so
2535
- check it before treating a serving server as complete. Present only
2536
- when something failed to load, so a healthy server omits the field.
2533
+ serving and the affected packages are absent from environments,
2534
+ unless the entry carries stale = true, in which case the package is
2535
+ still listed and still serving its last successful compile. This is
2536
+ the only place either difference is reported, so check it before
2537
+ treating a serving server as complete or current. Present only when
2538
+ something failed, so a healthy server omits the field.
2537
2539
  items:
2538
2540
  type: object
2539
2541
  properties:
@@ -2551,6 +2553,24 @@ components:
2551
2553
  message:
2552
2554
  type: string
2553
2555
  description: Why it did not load.
2556
+ stale:
2557
+ type: boolean
2558
+ description:
2559
+ When true, the package did load earlier and is still serving
2560
+ that last successful compile; the message is why its most
2561
+ recent reload (a watch-mode save, malloy_reloadPackage, or
2562
+ reload=true) failed. The served model is older than the files
2563
+ on disk. Absent on ordinary load failures. This array is the
2564
+ only place staleness is reported, because the package's entry
2565
+ under environments, and its own package resource, both read
2566
+ SERVING; join on the package name here to tell a current
2567
+ package from a stale one.
2568
+ failedAt:
2569
+ type: string
2570
+ format: date-time
2571
+ description:
2572
+ When the failed reload happened. Only present alongside
2573
+ stale = true.
2554
2574
  frozenConfig:
2555
2575
  type: boolean
2556
2576
  description:
@@ -2753,8 +2773,19 @@ components:
2753
2773
  location:
2754
2774
  type: string
2755
2775
  description:
2756
- Package location, can be an absolute path or URI (e.g. github, s3,
2757
- gcs, etc.)
2776
+ Package location. A LOCAL path must start with `./`, `../`, `~/`,
2777
+ or `/`; a `./` or `../` path resolves against the config file's
2778
+ directory, and `~/` against the home directory. Anything else is
2779
+ read as a remote URI (github, s3, gcs, https), so a bare name like
2780
+ `sales` is not local and the package is skipped with
2781
+ "Invalid package path" in the server's loadErrors. At boot a local
2782
+ package is COPIED into publisher_data/ and the copy is served, so
2783
+ edits to the source directory are not read; starting the server
2784
+ with `--watch-env <env>` mounts local packages in place (symlink)
2785
+ and live-reloads them instead. The mount is chosen when the
2786
+ environment is first loaded from config, so adding `--watch-env`
2787
+ to a later boot needs `--init` to re-mount a package already
2788
+ copied.
2758
2789
  scope:
2759
2790
  type: string
2760
2791
  enum:
@@ -2812,9 +2843,55 @@ components:
2812
2843
  findings: a `storage=` source not served from storage, a `#@ persist`
2813
2844
  source that was not recognized as materializable, a persist-target
2814
2845
  collision, and materialization-config problems that degrade
2815
- attribution. Server-computed and read-only: ignored on create/update
2816
- requests and only returned in responses. Present only when there are
2817
- such findings.
2846
+ attribution. Discovery findings: a model listed in `explores` whose
2847
+ export closure surfaces no sources or named queries, so it lists as
2848
+ empty and renders a blank page. Server-computed and read-only:
2849
+ ignored on create/update requests and only returned in responses.
2850
+ Present only when there are such findings.
2851
+
2852
+
2853
+ **Pre-aggregation contributes nothing to this array. A
2854
+ `#@ preaggregate` that cannot be honored is an ERROR, never a
2855
+ warning** — a 400 at publish or PATCH, and a load failure thereafter.
2856
+ Rejected are a `#@ preaggregate` anywhere but on a measure;
2857
+ a `#@ preaggregate` with no `grain=`;
2858
+ a grain naming anything but a dimension the base source itself
2859
+ declares, which rules out an inline time truncation such as
2860
+ `order_time.day`; a measure that cannot be correctly
2861
+ re-aggregated to a coarser grain, whether because its aggregate has no
2862
+ exact merge function (`count(distinct)`, `median`, percentiles, `mode`,
2863
+ and `avg` until it is decomposed) or because it carries its own filter,
2864
+ an ungrouping (`all`, `exclude`), or a required group-by; a measure
2865
+ whose expression cannot be POSITIVELY classified as re-aggregatable,
2866
+ which fails closed rather than being guessed at; a base source with a
2867
+ FAN-OUT join (`join_many` or `join_cross`), which can multiply rows,
2868
+ while a `join_one` is permitted and a measure aggregating through one
2869
+ is served normally; a base source that already fails persist eligibility (unbound
2870
+ parameters, `given` references, `#(authorize)` gates); `sharing=` or
2871
+ `schedule=` at the resource level; and a synthesized rollup name that
2872
+ collides with an existing model object. Each names the measure and the
2873
+ fix.
2874
+
2875
+
2876
+ Refusing rather than silently dropping the measure is the deliberate
2877
+ choice. A dropped measure is not served from the rollup at ANY grain,
2878
+ so the declaration would have no effect whatsoever while still reading
2879
+ as applied — and the build cost of the rollup would still be paid. An
2880
+ author who writes `#@ preaggregate` on a measure that cannot be rolled
2881
+ up has made a modeling mistake, and an error is where a mistake
2882
+ belongs.
2883
+
2884
+
2885
+ The same verdict is enforced at LOAD, not only at publish, so a
2886
+ package whose declarations stopped being honorable does not load at
2887
+ all: it is absent from its environment and reported in
2888
+ `ServerStatus.loadErrors`. Worth knowing because whether a measure can
2889
+ be re-aggregated is derived from the compiled model, so a Malloy
2890
+ version change can in principle reclassify a measure that published
2891
+ cleanly long before. That surfaces as a package that stops loading
2892
+ rather than as a package quietly serving without its rollups — louder,
2893
+ and deliberately so, since the alternative is paying for rollups that
2894
+ answer nothing.
2818
2895
  items:
2819
2896
  type: object
2820
2897
  properties:
@@ -2850,15 +2927,28 @@ components:
2850
2927
  description: >-
2851
2928
  Controls whether the discovery surface is also a query boundary.
2852
2929
  `"declared"` (the default) makes queryable == discoverable: when
2853
- `explores` is declared, only `explores` model files and within
2854
- them only the `export {}` closure are valid top-level query
2855
- targets; every other source still compiles, imports, joins, and
2856
- extends but is not directly queryable (denied with 404). `"all"`
2857
- decouples them: `explores`/`export {}` gate discovery only and every
2858
- compiled source stays directly queryable. When `explores` is absent
2859
- there is no curated surface, so both modes are equivalent
2860
- (everything queryable). Invalid values fall back to `"declared"`.
2861
- Identity-based access is a separate concern see `#(authorize)`.
2930
+ `explores` is declared, only `explores` model files are valid query
2931
+ entry points, and the queryable sources are the union of those
2932
+ files' `export {}` closures — so a source exported by any listed
2933
+ file stays queryable whichever listed model path the request
2934
+ addresses it through. Admission is by declaration, not by name: a
2935
+ request clears only when the model it names resolves the name to the
2936
+ very source a listed file exported, so a same-named source declared
2937
+ in a hidden file is not admitted by the coincidence. Every other
2938
+ source still compiles, imports, joins, and extends but is not
2939
+ directly queryable (denied with 404).
2940
+ The `/compile` endpoint is exempt (compile is the authoring loop;
2941
+ the boundary is discovery curation, not access control — for
2942
+ confidentiality use `#(authorize)`, which gates compile too). A
2943
+ source that is both hidden and `#(authorize)`-gated still answers
2944
+ `/compile` with the generic 404, so the exemption cannot be used to
2945
+ enumerate gated names.
2946
+ `"all"` decouples the axes: `explores`/`export {}` gate discovery
2947
+ only and every compiled source stays directly queryable. When
2948
+ `explores` is absent there is no curated surface, so both modes are
2949
+ equivalent (everything queryable). Invalid values fall back to
2950
+ `"declared"`. Identity-based access is a separate concern — see
2951
+ `#(authorize)`.
2862
2952
  manifestLocation:
2863
2953
  type: ["string", "null"]
2864
2954
  description: |
@@ -3566,7 +3656,10 @@ components:
3566
3656
  statements a query issues, for the backend's own reporting — cost
3567
3657
  attribution, workload classification, tracing. Each backend applies it
3568
3658
  through a per-query mechanism: Snowflake a per-statement `QUERY_TAG`
3569
- (the bag as JSON, so its properties are queryable in `QUERY_HISTORY`),
3659
+ (the bag as JSON, so its properties are queryable in `QUERY_HISTORY`)
3660
+ except on a `storage=` build, whose read goes through DuckDB's
3661
+ query-passthrough and so sets the tag for the SESSION, which means the
3662
+ driver's own connection probes carry the build's bag too,
3570
3663
  BigQuery per-job `labels` (queryable in `INFORMATION_SCHEMA.JOBS`), and
3571
3664
  every other backend a leading SQL comment (`-- team="finance"`, visible
3572
3665
  in query history / `pg_stat_activity`). It never affects results, and it
@@ -3645,6 +3738,74 @@ components:
3645
3738
  Which of the two queryRowLimit came from. Only server_default means
3646
3739
  rows were probably left behind; a deliberate limit or top returning
3647
3740
  exactly that many rows is a complete answer.
3741
+ servedFrom:
3742
+ type: ["string", "null"]
3743
+ enum: [storage, live_fallback, null]
3744
+ description: >-
3745
+ How this answer was produced, when the query reached the storage
3746
+ routing decision at all.
3747
+
3748
+ `storage` — served from a materialized table in a `storage=`
3749
+ destination, via the virtual-source transform. The source's own
3750
+ warehouse was not queried.
3751
+
3752
+ `live_fallback` — the query DID route to storage and the store then
3753
+ failed underneath it, and the source's freshness fallback allowed it
3754
+ to degrade to the live warehouse. **The rows are correct and the
3755
+ request succeeded**, which is why this needs its own value: it is
3756
+ indistinguishable from a storage hit at the call site, and treating it
3757
+ as one would report a healthy hit rate for a broken store.
3758
+
3759
+ Null when the query never had a storage binding to consider — every
3760
+ query on a deployment with no `storage=` sources, which is the default.
3761
+
3762
+ A storage-served result is byte-identical to a live one by design, so
3763
+ this field is the only way a caller can tell that materialization did
3764
+ anything, or show a user which of their sources it applies to.
3765
+ executionTimeMs:
3766
+ type: integer
3767
+ description: >-
3768
+ Wall-clock milliseconds from the start of query handling through the
3769
+ end of execution.
3770
+
3771
+ Covers compile, authorize, storage routing, prepare and the warehouse
3772
+ round trip. Narrower than the HTTP round trip — it excludes
3773
+ serialization and transport, so it reads lower than a caller's own
3774
+ stopwatch — and it is the same span the `malloy_model_query_duration`
3775
+ histogram records, so the two are directly comparable. Use it to
3776
+ attribute ONE query, which an aggregate histogram cannot do; use the
3777
+ histogram for rates and percentiles.
3778
+ queryCostBytes:
3779
+ type: ["integer", "null"]
3780
+ description: >-
3781
+ Warehouse bytes this query SCANNED, as reported by the backend that
3782
+ ran it. Not what it is billed — see below.
3783
+
3784
+ **Null is not zero.** It means the backend did not report a figure,
3785
+ and there are two very different reasons for that. Either the backend
3786
+ has no inline equivalent — BigQuery returns `totalBytesProcessed` in
3787
+ the job metadata the driver already fetches, while Snowflake exposes
3788
+ bytes only in `QUERY_HISTORY` keyed by query id (join it later on the
3789
+ `query_id` query-metadata property rather than paying a round trip per
3790
+ query) and Postgres has none at all. Or the query touched no warehouse,
3791
+ which is the case for every `servedFrom: storage` answer. Check
3792
+ `servedFrom` before reading a null as "free".
3793
+
3794
+ Where it is reported it is the executed job's ACTUAL bytes
3795
+ processed, i.e. bytes SCANNED — not an estimate, and not what the
3796
+ query is billed. BigQuery bills a 10MB minimum per query, so a small
3797
+ read is billed well above this figure. Good for comparing queries
3798
+ against each other; not a spend number.
3799
+
3800
+ Named `cost` rather than `scanned` because it carries Malloy's own
3801
+ `runStats.queryCostBytes` through verbatim — the name follows the
3802
+ figure back to its source.
3803
+
3804
+ One asymmetry worth knowing on BigQuery: the connector reports a
3805
+ genuinely-zero scan as absent rather than as 0, so a fully cached
3806
+ query is indistinguishable here from a backend that cannot say. Free
3807
+ and unknown are the same value, and a savings calculation built on
3808
+ this will treat free queries as unmeasured rather than as free.
3648
3809
 
3649
3810
  LogMessage:
3650
3811
  type: object
@@ -3728,7 +3889,14 @@ components:
3728
3889
  description: Resource path to the connection
3729
3890
  name:
3730
3891
  type: string
3731
- description: Name of the connection
3892
+ description:
3893
+ Name of the connection. The name `duckdb` is RESERVED for the
3894
+ per-package DuckDB sandbox every loaded package gets automatically;
3895
+ declaring an environment-level connection with that name fails the
3896
+ whole environment at init. A flat-file (CSV/Parquet/XLSX) package
3897
+ needs no connection entry at all — its models reach the sandbox as
3898
+ `duckdb.table('data/…')`. Name an environment-level DuckDB
3899
+ connection something else (e.g. `shared_duckdb`).
3732
3900
  type:
3733
3901
  type: string
3734
3902
  description: Type of database connection
@@ -4702,6 +4870,45 @@ components:
4702
4870
  type: string
4703
4871
  dialect:
4704
4872
  type: string
4873
+ origin:
4874
+ type: string
4875
+ enum: [persist, preaggregate]
4876
+ default: persist
4877
+ description: >-
4878
+ Which authoring annotation produced this plan entry. `persist` — the
4879
+ modeler wrote `#@ persist` on the source; every entry on a deployment
4880
+ that uses no pre-aggregation, and the meaning of an absent value.
4881
+ `preaggregate` — the modeler wrote `#@ preaggregate` on measures of
4882
+ `preaggregate.baseSourceName`, and the publisher SYNTHESIZED this
4883
+ rollup source to serve them.
4884
+
4885
+
4886
+ Note that BOTH kinds carry a `#@ persist` annotation by the time they
4887
+ reach this plan, because synthesis emits one (see `annotationFields`).
4888
+ So this reports which surface the AUTHOR used, not which annotation
4889
+ sits on the source — which is the question worth answering, since a
4890
+ `preaggregate` entry's `name` appears in no model file and without
4891
+ this field an operator or orchestrator sees a source nobody wrote with
4892
+ no way to tell it from one they did.
4893
+
4894
+
4895
+ Everything downstream is deliberately indifferent to it: a synthesized
4896
+ entry is built, identified, manifested, shared and garbage-collected
4897
+ exactly like a hand-written one.
4898
+ preaggregate:
4899
+ oneOf:
4900
+ - $ref: "#/components/schemas/PreaggregatePlan"
4901
+ - type: "null"
4902
+ description: >-
4903
+ What the synthesized rollup covers: the base source it was derived
4904
+ from, the one grain it is built at, and the measures served there.
4905
+ Present only when `origin` is `preaggregate`; null otherwise.
4906
+
4907
+
4908
+ Coverage is complete by construction — a measure that cannot be
4909
+ correctly re-aggregated is refused at publish rather than dropped
4910
+ from the rollup — so there is no per-measure omission list to
4911
+ reconcile against the model.
4705
4912
  sourceEntityId:
4706
4913
  type: string
4707
4914
  description: >-
@@ -4740,6 +4947,14 @@ components:
4740
4947
  scope modes. (Per-source `sharing`/`schedule` were retired: scope is
4741
4948
  a single package-level mode and a schedule is package-root-only —
4742
4949
  declaring either on a source is a publish-time manifest error.)
4950
+
4951
+
4952
+ A synthesized rollup (`origin: preaggregate`) declares nothing of its
4953
+ own and gets no freshness policy of its own: it INHERITS its base's,
4954
+ so what is reported here is the base source's effective objective
4955
+ resolved by the rule above. A rollup and its base therefore go stale
4956
+ together, which is what makes "serve the base when the rollup is
4957
+ stale" a safe degradation rather than a change of answer.
4743
4958
  queryMetadata:
4744
4959
  oneOf:
4745
4960
  - $ref: "#/components/schemas/QueryMetadata"
@@ -4794,15 +5009,142 @@ components:
4794
5009
  a dangling or stale name is rejected rather than deferred to a build.
4795
5010
  Neither participates in `sourceEntityId`: adding a watermark to an
4796
5011
  existing source must not re-address its table.
5012
+
5013
+
5014
+ For a synthesized rollup (`origin: preaggregate`) the annotation was
5015
+ generated by the publisher rather than authored, so this reports what
5016
+ synthesis emitted. It is still the annotation that governs the build,
5017
+ which is why it is reported the same way: a rollup is an ordinary
5018
+ persist source whose `#@ persist` nobody typed.
4797
5019
  modelPath:
4798
5020
  type: string
4799
- description:
5021
+ description: >-
4800
5022
  Package-relative path of the `.malloy` model that declares this
4801
5023
  source (e.g. `order_rollup.malloy`). The source's sourceID embeds an
4802
5024
  absolute `file://` modelURL with no package boundary, so this is the
4803
5025
  only place the relative path is exposed; callers use it to deep-link
4804
5026
  the source back to its model.
4805
5027
 
5028
+
5029
+ A synthesized rollup (`origin: preaggregate`) is declared by no model,
5030
+ so this is the model of its base — the file holding the
5031
+ `#@ preaggregate` annotations it was synthesized from, which is where
5032
+ an author would go to change it.
5033
+
5034
+ PreaggregatePlan:
5035
+ type: object
5036
+ description: >-
5037
+ Provenance for a rollup source synthesized from `#@ preaggregate`
5038
+ annotations. A query never names a rollup, never changes shape because one
5039
+ exists, and gets the same answer whether one is used or not, so this is
5040
+ the only place the set of rollups and the coverage of each is visible at
5041
+ all.
5042
+
5043
+
5044
+ **This object describes ONE rollup table, at ONE grain.** A rollup is a
5045
+ single `GROUP BY`, so a grain is a property of the table, not of the
5046
+ individual measures in it: `grainDimensions` lists the dimensions of that
5047
+ one grain, and `measures` are the measures served at it.
5048
+
5049
+
5050
+ That inverts how the grain was authored, which is worth knowing before
5051
+ reading a plan. The modeler declares per measure — `#@ preaggregate
5052
+ grain="order_day, category"` — and may declare a measure at SEVERAL grains
5053
+ by writing several annotation lines on it, each of which is its own rollup.
5054
+ The publisher then packs every measure declaring the same (grain,
5055
+ connection, base source) into a single synthesized source, because ten
5056
+ measures at one grain should be one table and one CTAS rather than ten. So
5057
+ the plan has one entry per DISTINCT grain declared anywhere on the source,
5058
+ which is neither the number of annotated measures nor the number of
5059
+ annotations: a measure declared at two grains appears in two entries, and
5060
+ two measures sharing a grain appear in one. To recover the per-measure view
5061
+ that was written, iterate `buildPlan.sources`, keep the `origin:
5062
+ preaggregate` entries, and invert each entry's grain-to-measures pairing.
5063
+
5064
+
5065
+ Choosing between two rollups and one at the combined grain is a size
5066
+ judgement, not a correctness one — see `grainDimensions`.
5067
+ required: [baseSourceName, grainDimensions, measures]
5068
+ properties:
5069
+ baseSourceName:
5070
+ type: string
5071
+ description: >-
5072
+ The source whose measures carry the `#@ preaggregate` annotations this
5073
+ rollup was synthesized from. Queries name THIS source; the rollup is
5074
+ selected behind it, or bypassed, per grain coverage.
5075
+ grainDimensions:
5076
+ type: array
5077
+ items:
5078
+ type: string
5079
+ description: >-
5080
+ The dimensions of the ONE grain this rollup is built at — the parsed
5081
+ form of a single `grain=` annotation value, so `grain="order_day,
5082
+ category"` reports as two entries here and not as two grains.
5083
+
5084
+
5085
+ **Every entry is a dimension the base source itself declares**, never
5086
+ a truncation written inline. `grain="order_time.day"` is refused at
5087
+ publish: a stored column of truncated values either shadows the base's
5088
+ own `order_time` and answers untruncated queries from truncated rows,
5089
+ or is named something no query references and serves nothing. So time
5090
+ grains are reached by declaring `dimension: order_day is
5091
+ order_time.day` on the source and naming `order_day` here.
5092
+
5093
+
5094
+ **Sorted, and canonically so.** Synthesis derives the rollup's name
5095
+ and identity from the sorted grain, so two authors who write the same
5096
+ dimensions in different orders get one table rather than two. Compare
5097
+ grains as sets using this list; never by the annotation's original
5098
+ string, which is not canonical.
5099
+
5100
+
5101
+ A query can be answered here only if this grain covers it: every
5102
+ dimension the query groups by is in this list, or is a coarser
5103
+ truncation of one (a `month` of a stored `order_day` is derivable and
5104
+ does route; a `day` of a stored `order_month` is not). Which makes the
5105
+ set of grains ACROSS a package the whole performance story, and the
5106
+ reason adding a rollup per query shape is the documented way to spend
5107
+ more on builds than you save.
5108
+
5109
+
5110
+ Note that covering does not require an exact match: a query grouping
5111
+ by a SUBSET of these dimensions — or by none of them — is served here
5112
+ too, by re-aggregating across the ones it omits. So one rollup at the
5113
+ finest grain an author needs does subsume every coarser one, and a
5114
+ combined grain is often the economical choice.
5115
+
5116
+
5117
+ That subsumption is about coverage, though, not cost, and it is why a
5118
+ measure may still be declared at several grains. A combined grain has
5119
+ roughly the PRODUCT of its dimensions' cardinalities, so
5120
+ `customer_id, order_day` can approach the base table's row count and
5121
+ save almost nothing, while `order_day` alone is tiny. Two rollups then
5122
+ beat one, each query reading the smallest that covers it, at the price
5123
+ of two tables to build and refresh. Correctness is identical either
5124
+ way; the choice is size against build cost.
5125
+ measures:
5126
+ type: array
5127
+ items:
5128
+ type: string
5129
+ description: >-
5130
+ The measures served at this rollup's grain, under their original names
5131
+ — the names a query uses, not the names of the partial-aggregate
5132
+ columns stored to serve them. A query using only these, at a grain
5133
+ this rollup covers, is eligible to be served from its table.
5134
+
5135
+
5136
+ These are the measures whose `#@ preaggregate` named THIS grain. A
5137
+ measure declared at a different grain is not missing: it is in the
5138
+ plan entry for that grain. Do not read this list as the base source's
5139
+ full set of pre-aggregated measures — that is the union of these lists
5140
+ across every entry sharing a `baseSourceName`.
5141
+
5142
+
5143
+ **Every measure that declared this grain is here.** There is no
5144
+ partial rollup and no per-measure omission to check: a measure that
5145
+ cannot be correctly re-aggregated is refused at publish, so a rollup
5146
+ that exists carries everything that asked to be in it.
5147
+
4806
5148
  Realization:
4807
5149
  type: string
4808
5150
  enum: [SNAPSHOT, COPY]
@@ -5046,8 +5388,40 @@ components:
5046
5388
  properties:
5047
5389
  sourceEntityId:
5048
5390
  type: string
5391
+ error:
5392
+ type: [string, "null"]
5393
+ description: >-
5394
+ Why this source failed to materialize, in the words the warehouse or
5395
+ compiler reported. Present only on a source that failed; absent on one
5396
+ that built, whose table the remaining fields describe. A build that
5397
+ loses some of its sources still returns the ones that succeeded, so an
5398
+ entry carrying this is how a consumer tells a failed source from a
5399
+ healthy one.
5049
5400
  sourceName:
5050
5401
  type: string
5402
+ description: >-
5403
+ The persist source this table was built for. Convenience only:
5404
+ `sourceEntityId` is the identity, and this is the human-readable label
5405
+ beside it.
5406
+
5407
+
5408
+ **It may name no source in any model file.** A pre-aggregation rollup
5409
+ is a persist source the publisher synthesized, so its name is
5410
+ generated (see `PersistSourcePlan.origin`). Match on
5411
+ `sourceEntityId`; do not resolve this against a package's models and
5412
+ treat a miss as a corrupt manifest.
5413
+ origin:
5414
+ type: string
5415
+ enum: [persist, preaggregate]
5416
+ default: persist
5417
+ description: >-
5418
+ Mirrors `PersistSourcePlan.origin` for the source this entry was built
5419
+ from: `persist` for a source the modeler wrote `#@ persist` on, which
5420
+ is the meaning of an absent value, and `preaggregate` for a rollup
5421
+ synthesized from `#@ preaggregate` measures. Repeated here because a
5422
+ manifest travels without its build plan, and a consumer holding only
5423
+ the manifest otherwise cannot tell why it holds a table for a source
5424
+ that does not appear in the package.
5051
5425
  materializedTableId:
5052
5426
  type: string
5053
5427
  description: Echoes the caller-assigned id from the BuildInstruction.
@@ -5088,6 +5462,56 @@ components:
5088
5462
  $ref: "#/components/schemas/Realization"
5089
5463
  rowCount:
5090
5464
  type: ["integer", "null"]
5465
+ buildDurationMs:
5466
+ type: ["integer", "null"]
5467
+ description: >-
5468
+ Wall-clock milliseconds spent building this one source's table —
5469
+ measured around the build itself, so it excludes planning, manifest
5470
+ distribution, and anything the orchestrator does around it.
5471
+
5472
+ Null when no build ran, which is the normal case for an incremental
5473
+ source whose refresh SKIPPED because its boundary already covered the
5474
+ requested range. Null rather than 0 on purpose: a zero would average
5475
+ into build-duration series as an implausibly fast build.
5476
+
5477
+ Per SOURCE, where a run-level duration is per instruction. A run that
5478
+ builds five sources reports five of these, and their sum is less than
5479
+ the run's duration by whatever the orchestration cost.
5480
+ queryCostBytes:
5481
+ type: ["integer", "null"]
5482
+ description: >-
5483
+ Warehouse bytes SCANNED by the read that produced this table — the
5484
+ recurring cost of keeping the source materialized, and the debit
5485
+ against whatever materializing it saves on the read side.
5486
+
5487
+ Bytes scanned, never bytes billed. The two differ by up to BigQuery's
5488
+ 10MB-per-query floor, and materialization refreshes are mostly small
5489
+ reads, so a field carrying one on some paths and the other elsewhere
5490
+ would be unsummable across a package's sources.
5491
+
5492
+ **Null is not zero**, and which paths report a figure differs.
5493
+
5494
+ A COLOCATED build reports one, taken from the Malloy connection's own
5495
+ statistics.
5496
+
5497
+ A plain `storage=` build reports one when its warehouse read carried
5498
+ query metadata, and null otherwise. That read goes through DuckDB's
5499
+ query-passthrough rather than a Malloy connector, so the figure is
5500
+ read back from the warehouse's own accounting afterwards — which needs
5501
+ an identifier for the job, and obtaining one is part of how a TAGGED
5502
+ read is issued. An untagged read is issued the older way, which yields
5503
+ no such identifier, and reports null. A read whose connection cannot
5504
+ reach its own job records also reports null; that case is counted by
5505
+ `publisher_storage_build_attribution_skipped_total`.
5506
+
5507
+ An incremental delta reports null because its statements do not run
5508
+ through a single call whose result reaches here. A chained `storage=`
5509
+ build reports null because it read its parent's already-materialized
5510
+ table and touched no warehouse at all. A POSTGRES source reports null
5511
+ on every path: the label that makes a read reportable is BigQuery's,
5512
+ and Postgres has no per-statement tag to carry one. And a bag whose
5513
+ every property is unusable as a BigQuery label key renders no label,
5514
+ so such a read is issued the untagged way and reports null too.
5091
5515
  dataAsOf:
5092
5516
  type: string
5093
5517
  format: date-time
@@ -1 +1 @@
1
- import{$ as r,H as t,j as e,a1 as i,aa as o}from"./index-B-jw9EkA.js";function m(){const s=r(),{environmentName:n}=t();if(n){const a=i({environmentName:n});return e.jsx(o,{onSelectPackage:s,resourceUri:a})}else return e.jsx("div",{children:e.jsx("h2",{children:"Missing environment name"})})}export{m as default};
1
+ import{$ as r,H as t,j as e,a1 as i,aa as o}from"./index-DJ-3RN5W.js";function m(){const s=r(),{environmentName:n}=t();if(n){const a=i({environmentName:n});return e.jsx(o,{onSelectPackage:s,resourceUri:a})}else return e.jsx("div",{children:e.jsx("h2",{children:"Missing environment name"})})}export{m as default};