@gscdump/analysis 3.4.4 → 3.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -4,7 +4,8 @@
4
4
  [![npm downloads](https://img.shields.io/npm/dm/@gscdump/analysis?color=yellow)](https://npm.chart.dev/@gscdump/analysis)
5
5
  [![license](https://img.shields.io/github/license/harlan-zw/gscdump?color=yellow)](https://github.com/harlan-zw/gscdump/blob/main/LICENSE)
6
6
 
7
- > SEO analyzers for Google Search Console data. Row-based, DuckDB-native, D1-ready.
7
+ SEO Analyzers and Reports for Google Search Console data.
8
+ Use pure functions for rows you already have, or run an Analyzer against a Source.
8
9
 
9
10
  ## Install
10
11
 
@@ -12,181 +13,159 @@
12
13
  npm install @gscdump/analysis
13
14
  ```
14
15
 
15
- ## When to use which subpath
16
+ Node.js 22 or newer is required for Node consumers.
16
17
 
17
- | Subpath | Use when |
18
- |---|---|
19
- | `@gscdump/analysis` | Pure analyzers, `analyzeInBrowser`, and shared analyzer contracts. |
20
- | `@gscdump/analysis/registry` | Pre-built `defaultAnalyzerRegistry` (rows + sql). Convenience for callers who don't care about bundle size. |
21
- | `@gscdump/analysis/errors` | Typed analysis failures and rendering helpers. |
22
- | `@gscdump/analysis/report` | Report registry and runtime. |
23
- | `@gscdump/analysis/source` | Portable source factories. |
18
+ ## Choose an export
24
19
 
25
- The contract layer (`Analyzer`, `Plan`, `Capability`, `AnalysisParams`, `AnalysisResult`, `AnalysisQuerySource`, `runAnalyzerFromSource`, `createAnalyzerRegistry`, `defineAnalyzer`, period helpers, `createEngineQuerySource`) lives in `@gscdump/engine` under the `/analyzer`, `/analysis-types`, `/period`, `/source`, and `/resolver` subpaths. Most are re-exported from `@gscdump/analysis` for convenience.
20
+ | Subpath | Purpose |
21
+ | --- | --- |
22
+ | `@gscdump/analysis` | Pure Analyzers, browser dispatch, and shared Analyzer contracts |
23
+ | `@gscdump/analysis/registry` | All 29 registered Analyzers, including SQL implementations |
24
+ | `@gscdump/analysis/report` | Report registry, `runReport`, and formatting |
25
+ | `@gscdump/analysis/source` | Composite and in-memory Source factories |
26
+ | `@gscdump/analysis/errors` | Typed analysis errors and rendering helpers |
26
27
 
27
- ## Row-based analyzers
28
+ `@gscdump/engine` owns the Analyzer, Source, and period contracts.
29
+ The analysis root re-exports the common contracts.
30
+ The default registry imports every Analyzer; avoid it when you need a smaller browser bundle.
28
31
 
29
- Pure functions. Take typed arrays in, return typed results out.
32
+ ## Analyze rows
30
33
 
31
34
  ```ts
32
- import {
33
- analyzeBrandSegmentation,
34
- analyzeClustering,
35
- analyzeConcentration,
36
- analyzeDecay,
37
- analyzeMovers,
38
- analyzeOpportunity,
39
- analyzeSeasonality,
40
- padTimeseries,
41
- } from '@gscdump/analysis'
42
-
43
- const movers = analyzeMovers(currentRows, previousRows)
44
- const decay = analyzeDecay(currentRows, previousRows)
35
+ import { analyzeDecay, analyzeMovers } from '@gscdump/analysis'
36
+
37
+ const current = [{
38
+ query: 'example query',
39
+ page: 'https://example.com/docs',
40
+ clicks: 40,
41
+ impressions: 1000,
42
+ ctr: 0.04,
43
+ position: 8,
44
+ }]
45
+ const previous = [{
46
+ query: 'example query',
47
+ page: 'https://example.com/docs',
48
+ clicks: 100,
49
+ impressions: 1000,
50
+ ctr: 0.1,
51
+ position: 4,
52
+ }]
53
+
54
+ const movers = analyzeMovers({ current, previous })
55
+ const decay = analyzeDecay({ current, previous })
56
+ console.log(movers.declining, decay)
45
57
  ```
46
58
 
47
- Meta-analyses live in the report layer. Run a composed evidence report via `runReport`:
59
+ Comparison functions take one `{ current, previous }` input object.
60
+ Options are a separate second argument.
61
+ See the [SEO guide](../../docs/guides/seo-analysis.md) for a complete striking-distance example.
48
62
 
49
- ```ts
50
- import { defaultAnalyzerRegistry } from '@gscdump/analysis/registry'
51
- import {
52
- defaultReportRegistry,
53
- runReport,
54
- } from '@gscdump/analysis/report'
55
- import { resolveWindow } from '@gscdump/engine/period'
56
-
57
- const report = defaultReportRegistry.getReport('health')!
58
- const window = resolveWindow({ preset: 'last-28d', comparison: 'none' })
59
- const result = await runReport(report, {
60
- source,
61
- analyzers: defaultAnalyzerRegistry,
62
- ctx: { site: siteUrl, window, params: {}, registryVersion: defaultReportRegistry.version },
63
- })
64
- // result.sections[0].findings — bounded evidence, page+query keyed.
65
- ```
63
+ ## Sources
66
64
 
67
- Source adapters compose a GSC client + analyzer in one call:
65
+ Use `runAnalyzerFromSource` to choose an Analyzer's row or SQL plan:
68
66
 
69
67
  ```ts
70
- import { analyzeMoversFromSource } from '@gscdump/analysis'
68
+ import { defaultAnalyzerRegistry } from '@gscdump/analysis/registry'
71
69
  import { createGscApiQuerySource } from '@gscdump/engine-gsc-api'
72
-
73
- const source = createGscApiQuerySource({ client, siteUrl })
74
- const movers = await analyzeMoversFromSource(source, { current, previous })
75
- ```
76
-
77
- ## DuckDB (Node)
78
-
79
- SQL-native path. `SQL_ANALYZERS` dispatch through `runAnalyzerFromSource` against an engine-backed source.
80
-
81
- ```ts
82
- import { ROW_ANALYZERS, SQL_ANALYZERS } from '@gscdump/analysis/registry'
83
- import { createAnalyzerRegistry, runAnalyzerFromSource } from '@gscdump/engine/analyzer'
84
- import { createEngineQuerySource } from '@gscdump/engine/source'
85
-
86
- const source = createEngineQuerySource({ engine, ctx })
87
- const registry = createAnalyzerRegistry({ rows: ROW_ANALYZERS, sql: SQL_ANALYZERS })
88
- const result = await runAnalyzerFromSource(source, { type: 'striking-distance', minImpressions: 100 }, registry)
70
+ import { runAnalyzerFromSource } from '@gscdump/engine/analyzer'
71
+ import { googleSearchConsole } from 'gscdump'
72
+
73
+ const client = googleSearchConsole({ accessToken: process.env.GSC_ACCESS_TOKEN! })
74
+ const source = createGscApiQuerySource({ client, siteUrl: 'sc-domain:example.com' })
75
+ const result = await runAnalyzerFromSource(source, {
76
+ type: 'striking-distance',
77
+ startDate: '2026-08-01',
78
+ endDate: '2026-08-28',
79
+ minImpressions: 100,
80
+ }, defaultAnalyzerRegistry)
81
+
82
+ console.log(result.results)
89
83
  ```
90
84
 
91
- `attachParquetIndex` and `attachSnapshotIndex` from `@gscdump/engine/node`
92
- wire parquet files (per-day, per-month, or pre-baked `.duckdb` snapshots) into
93
- a Node DuckDB session.
85
+ Install `@gscdump/engine`, `@gscdump/engine-gsc-api`, and `gscdump` for this example.
94
86
 
95
- ## Browser (DuckDB-WASM)
87
+ Twelve Analyzers have row plans:
96
88
 
97
- ```ts
98
- import { analyzeInBrowser } from '@gscdump/analysis'
99
- // Compose your own narrow registry instead of pulling the kitchen-sink default
100
- // (which statically imports every SQL analyzer). For demo only:
101
- import { defaultAnalyzerRegistry } from '@gscdump/analysis/registry'
89
+ - `brand`, `cannibalization`, `clustering`, `concentration`
90
+ - `data-detail`, `data-query`, `decay`, `movers`
91
+ - `opportunity`, `seasonality`, `striking-distance`, `zero-click`
102
92
 
103
- const result = await analyzeInBrowser(
104
- runner,
105
- { schema: 'gsc' },
106
- { type: 'striking-distance' },
107
- defaultAnalyzerRegistry,
108
- )
109
- ```
93
+ All 29 Analyzers have SQL plans.
94
+ SQL-only Analyzers need a Source that supports their required capabilities.
110
95
 
111
- `analyzeInBrowser` wraps any runner with `query(sql, params, signal?)` in an `AnalysisQuerySource` with the `attachedTables` capability and dispatches via `runAnalyzerFromSource`.
96
+ | Factory | Import path | Input |
97
+ | --- | --- | --- |
98
+ | `createGscApiQuerySource` | `@gscdump/engine-gsc-api` | `{ client, siteUrl }` |
99
+ | `createLiveGscSource` | `@gscdump/engine-gsc-api` | `{ siteUrl, getAccessToken }` |
100
+ | `createEngineQuerySource` | `@gscdump/engine/source` | `{ engine, ctx }` |
101
+ | `createSqliteQuerySource` | `@gscdump/engine-sqlite` | `{ executor, siteId, regex? }` |
102
+ | `createInMemoryQuerySource` | `@gscdump/analysis/source` | `{ queryRows }` |
103
+ | `createCompositeSource` | `@gscdump/analysis/source` | `{ engine, live, site }` |
112
104
 
113
- `@gscdump/engine-duckdb-wasm` exports `bootDuckDBWasm`,
114
- `attachParquetUrlTables`, `createBrowserAnalysisRuntime`, and `resolveWindow`
115
- (re-exported from `@gscdump/engine/period`). Browser analysis uses attached
116
- tables rather than a canonical-schema `createEngine`; see ADR-0001.
105
+ For a composite Source, `site` contains `oldestDateSynced`, `newestDateSynced`, and optional `coveredSpans`.
106
+ It routes supported queries to Google when stored coverage is missing or stored dimensions cannot answer the query.
107
+ SQL execution always uses the Engine.
117
108
 
118
- ## SQLite (D1 / Cloudflare Workers)
109
+ ## Reports
119
110
 
120
- Mirror of the DuckDB path, dialect-targeted at sqlite-core. `@gscdump/engine-sqlite` exports `createSqliteQuerySource` (an `AnalysisQuerySource` over `executor + siteId`), `compileSqlite`, drizzle helpers (`gsc_keywords`, etc.), and `resolveWindow` (re-export from `@gscdump/engine/period`).
121
-
122
- ## Query composers (dialect-neutral)
111
+ Reports combine Analyzers into bounded Sections.
112
+ After creating `source` above, run a Report supported by that Source:
123
113
 
124
114
  ```ts
125
- import { sqliteResolverAdapter } from '@gscdump/engine-sqlite'
126
- import { pgResolverAdapter, resolveToSQL } from '@gscdump/engine/resolver'
127
-
128
- const resolved = resolveToSQL(builderState, { adapter: sqliteResolverAdapter, siteId })
129
- ```
130
-
131
- Pass `sqliteResolverAdapter` from `@gscdump/engine-sqlite` (D1, `site_id`-scoped) or `pgResolverAdapter` from `@gscdump/engine/resolver` (parquet via DuckDB, single tenant). Composers stay identical; only the column bindings + dialect compilation differ.
115
+ import { defaultReportRegistry, runReport } from '@gscdump/analysis/report'
116
+ import { resolveWindow } from '@gscdump/engine/period'
132
117
 
133
- ## Sources (portable)
118
+ const report = defaultReportRegistry.getReport('movers')!
119
+ const window = resolveWindow({ preset: 'last-28d', comparison: 'prev-period' })
120
+ const result = await runReport(report, {
121
+ source,
122
+ analyzers: defaultAnalyzerRegistry,
123
+ ctx: {
124
+ site: 'sc-domain:example.com',
125
+ window,
126
+ params: {},
127
+ registryVersion: defaultReportRegistry.version,
128
+ },
129
+ })
134
130
 
135
- `/source` is the cross-implementation seam:
131
+ console.log(result.sections, result.meta.degraded)
132
+ ```
136
133
 
137
- ```ts
138
- import { analyzeMoversFromSource } from '@gscdump/analysis'
139
- import { createEngineQuerySource, queryRows } from '@gscdump/engine/source'
134
+ See the [Report list](../../README.md#reports) for inputs and defaults.
135
+ If an optional step fails, `meta.degraded` is `true`.
136
+ A required step failure rejects the Report.
140
137
 
141
- const source = createEngineQuerySource({ engine, ctx: { userId, siteId } })
138
+ Use `defineReport` from `@gscdump/engine/report` to define your own Report.
142
139
 
143
- const rows = await queryRows(source, builderState)
144
- const movers = await analyzeMoversFromSource(source, periods)
145
- ```
140
+ ## DuckDB and browser use
146
141
 
147
- Available source factories:
142
+ For Node, create an Engine Source with `createEngineQuerySource({ engine, ctx })` from `@gscdump/engine/source`.
143
+ The context supplies `userId` and `siteId`.
144
+ Node attachment helpers live at `@gscdump/engine/node`.
148
145
 
149
- - `createGscApiQuerySource({ client, siteUrl })` — `@gscdump/engine-gsc-api`
150
- - `createLiveGscSource({ accessToken, siteUrl })` — `@gscdump/engine-gsc-api`
151
- - `createCompositeSource({ engine, gsc })` — `@gscdump/analysis/source`; engine first, GSC fallback
152
- - `createInMemoryQuerySource({ queryRows })` — `@gscdump/analysis/source`
153
- - `createEngineQuerySource({ engine, ctx })` — `@gscdump/engine/source`
154
- - `createSqliteQuerySource({ ... })` — `@gscdump/engine-sqlite`
146
+ For browsers, use [`@gscdump/engine-duckdb-wasm`](../engine-duckdb-wasm/README.md) to attach Parquet tables.
147
+ `analyzeInBrowser` accepts a runner with `query(sql, params, signal?)`, options, analysis parameters, and an Analyzer registry.
148
+ Browser analysis uses attached tables as described in [ADR-0001](../../docs/adr/0001-browser-engine-uses-attached-tables.md).
155
149
 
156
- Portable analyzers currently cover the row-based tools:
157
- `striking-distance`, `opportunity`, `brand`, `clustering`, `concentration`,
158
- `seasonality`, `movers`, and `decay`.
150
+ For SQLite and D1, use [`@gscdump/engine-sqlite`](../engine-sqlite/README.md).
151
+ Analyzer support depends on the Source's SQL dialect and capabilities.
159
152
 
160
- ## Window resolution
153
+ ## Date windows
161
154
 
162
155
  ```ts
163
156
  import { resolveWindow } from '@gscdump/analysis'
164
157
 
165
- const w = resolveWindow({ preset: 'last-30d', comparison: 'yoy' })
166
- // { start: '...', end: '...', days: 30, comparison: { start, end } }
158
+ const window = resolveWindow({ preset: 'last-30d', comparison: 'yoy' })
159
+ console.log(window.start, window.end, window.comparison)
167
160
  ```
168
161
 
169
- Presets: `last-7d`, `last-28d`, `last-30d`, `last-90d`, `last-180d`, `last-365d`, `mtd`, `ytd`, `custom`. Comparison modes: `none`, `prev-period`, `yoy`.
170
-
171
- ## Stability
172
-
173
- | Surface | Stability |
174
- |---|---|
175
- | Row analyzers (`analyzeMovers`, `analyzeDecay`, ...) | Public |
176
- | Source factories + `analyzeFromSource` | Public |
177
- | `Analyzer<P, R>` contract + `createAnalyzerRegistry` (re-exported from `@gscdump/engine/analyzer`) | Public |
178
- | Source factories under `@gscdump/analysis/source` | Public |
179
- | Per-analyzer modules under `analysis/src/analyzers/<name>` | Private |
162
+ Presets: `last-7d`, `last-28d`, `last-30d`, `last-90d`, `last-180d`, `last-365d`, `mtd`, `ytd`, and `custom`.
163
+ Comparisons: `none`, `prev-period`, and `yoy`.
180
164
 
181
- ## Related
165
+ ## Public API
182
166
 
183
- - [`gscdump`](../gscdump) — REST client + query builder (edge-safe).
184
- - [`@gscdump/engine`](../engine) — Parquet/DuckDB storage engine + analyzer/source/period contracts.
185
- - [`@gscdump/engine/node`](../engine) — Node DuckDB handle + parquet/snapshot attach helpers.
186
- - [`@gscdump/engine-duckdb-wasm`](../engine-duckdb-wasm) — DuckDB-WASM browser runtime + drizzle adapter.
187
- - [`@gscdump/engine-sqlite`](../engine-sqlite) — SQLite / D1 dialect adapter.
188
- - [`@gscdump/engine-gsc-api`](../engine-gsc-api) — GSC live-API engine adapter.
189
- - [`@gscdump/cli`](../cli) — CLI wrapping `gscdump` + `@gscdump/engine` + `@gscdump/analysis`.
167
+ Use the package exports listed above.
168
+ Files under `src/` are private and may change without a public migration path.
190
169
 
191
170
  ## License
192
171
 
@@ -1,3 +1,2 @@
1
1
  import { Analyzer } from "@gscdump/engine/analyzer";
2
- declare const ROW_ANALYZERS: readonly Analyzer[];
3
- export { ROW_ANALYZERS };
2
+ export declare const ROW_ANALYZERS: readonly Analyzer[];
@@ -1,4 +1,4 @@
1
- import { parseJsonRows, rowString } from "../analyzer/row-values.mjs";
1
+ import { rowString } from "../analyzer/row-values.mjs";
2
2
  import { num } from "@gscdump/engine/analysis-types";
3
3
  import { defineAnalyzer } from "@gscdump/engine/analyzer";
4
4
  import { periodOf } from "@gscdump/engine/period";
@@ -13,37 +13,6 @@ const bipartitePagerankAnalyzer = defineAnalyzer({
13
13
  const topQueries = 1e3;
14
14
  const topUrls = 500;
15
15
  const limit = params.limit ?? 50;
16
- const bridgingEdgeThreshold = .05;
17
- const anchoringEdgeThreshold = .05;
18
- const iterations = BIPARTITE_PAGERANK_ITERATIONS;
19
- const d = BIPARTITE_PAGERANK_DAMPING;
20
- const iterCtes = [];
21
- for (let i = 1; i <= iterations; i++) iterCtes.push(`
22
- ranks_${i} AS (
23
- SELECT
24
- 'q' AS kind,
25
- e.qid AS id,
26
- (1.0 - ${d}) / (SELECT n FROM query_count)
27
- + ${d} * SUM(e.w_u_to_q * r.rank) AS rank
28
- FROM u_to_q_weights e
29
- JOIN ranks_${i - 1} r ON r.kind = 'u' AND r.id = e.uid
30
- GROUP BY e.qid
31
- UNION ALL
32
- SELECT
33
- 'u' AS kind,
34
- e.uid AS id,
35
- (1.0 - ${d}) / (SELECT n FROM url_count)
36
- + ${d} * SUM(e.w_q_to_u * r.rank) AS rank
37
- FROM q_to_u_weights e
38
- JOIN ranks_${i - 1} r ON r.kind = 'q' AND r.id = e.qid
39
- GROUP BY e.uid
40
- )`);
41
- const deltaParts = [];
42
- for (let i = 1; i <= iterations; i++) deltaParts.push(`
43
- SELECT ${i} AS step,
44
- (SELECT COALESCE(SUM(ABS(a.rank - b.rank)), 0.0)
45
- FROM ranks_${i} a
46
- JOIN ranks_${i - 1} b USING (kind, id)) AS l1`);
47
16
  return {
48
17
  sql: `
49
18
  WITH edges0 AS (
@@ -81,118 +50,23 @@ const bipartitePagerankAnalyzer = defineAnalyzer({
81
50
  JOIN top_queries tq USING (qid)
82
51
  JOIN top_urls tu USING (uid)
83
52
  ),
84
- query_nodes AS (SELECT DISTINCT qid FROM edges),
85
- url_nodes AS (SELECT DISTINCT uid FROM edges),
86
- query_count AS (SELECT GREATEST(COUNT(*), 1) AS n FROM query_nodes),
87
- url_count AS (SELECT GREATEST(COUNT(*), 1) AS n FROM url_nodes),
88
- -- Row-stochastic transition weights in each direction. For q->u the
89
- -- weights out of a query sum to 1; symmetric for u->q.
90
- q_out AS (SELECT qid, SUM(impressions) AS s FROM edges GROUP BY qid),
91
- u_out AS (SELECT uid, SUM(impressions) AS s FROM edges GROUP BY uid),
92
- q_to_u_weights AS (
93
- SELECT e.qid, e.uid,
94
- e.impressions / NULLIF(q.s, 0) AS w_q_to_u
95
- FROM edges e JOIN q_out q USING (qid)
96
- ),
97
- u_to_q_weights AS (
98
- SELECT e.qid, e.uid,
99
- e.impressions / NULLIF(u.s, 0) AS w_u_to_q
100
- FROM edges e JOIN u_out u USING (uid)
101
- ),
102
- -- Seed: uniform distribution per side. Total mass = 2 (one unit per side).
103
- ranks_0 AS (
104
- SELECT 'q' AS kind, q.qid AS id, 1.0 / (SELECT n FROM query_count) AS rank
105
- FROM query_nodes q
106
- UNION ALL
107
- SELECT 'u' AS kind, u.uid AS id, 1.0 / (SELECT n FROM url_count) AS rank
108
- FROM url_nodes u
109
- ),
110
- ${iterCtes.join(",\n")},
111
- final_ranks AS (SELECT * FROM ranks_${iterations}),
112
- -- Hub/anchor diagnostics computed from raw edge mass (not rank). A
113
- -- query "bridges" URLs it sends >= ${bridgingEdgeThreshold} of its mass
114
- -- to; a URL "anchors" queries that contribute >= ${anchoringEdgeThreshold}
115
- -- of its incoming mass.
116
- q_bridging AS (
117
- SELECT qid, COUNT(*) AS bridging
118
- FROM q_to_u_weights
119
- WHERE w_q_to_u >= ${bridgingEdgeThreshold}
120
- GROUP BY qid
121
- ),
122
- u_anchoring AS (
123
- SELECT uid, COUNT(*) AS anchoring
124
- FROM u_to_q_weights
125
- WHERE w_u_to_q >= ${anchoringEdgeThreshold}
126
- GROUP BY uid
53
+ query_nodes AS (
54
+ SELECT qid, ROW_NUMBER() OVER (ORDER BY qid) - 1 AS idx
55
+ FROM (SELECT DISTINCT qid FROM edges)
127
56
  ),
128
- q_degree AS (
129
- SELECT qid, COUNT(*) AS degree, SUM(impressions) AS impressions
130
- FROM edges GROUP BY qid
131
- ),
132
- u_degree AS (
133
- SELECT uid, COUNT(*) AS degree, SUM(impressions) AS impressions
134
- FROM edges GROUP BY uid
135
- ),
136
- deltas AS (
137
- ${deltaParts.join("\n UNION ALL\n")}
138
- ),
139
- query_rows AS (
140
- SELECT
141
- 'query' AS kind, f.id, f.rank,
142
- COALESCE(b.bridging, 0) AS bridging,
143
- 0 AS anchoring,
144
- COALESCE(qd.degree, 0) AS degree,
145
- COALESCE(qd.impressions, 0) AS impressions
146
- FROM final_ranks f
147
- LEFT JOIN q_bridging b ON b.qid = f.id
148
- LEFT JOIN q_degree qd ON qd.qid = f.id
149
- WHERE f.kind = 'q'
150
- ORDER BY f.rank DESC
151
- LIMIT ${Number(limit)}
152
- ),
153
- url_rows AS (
154
- SELECT
155
- 'url' AS kind, f.id, f.rank,
156
- 0 AS bridging,
157
- COALESCE(a.anchoring, 0) AS anchoring,
158
- COALESCE(ud.degree, 0) AS degree,
159
- COALESCE(ud.impressions, 0) AS impressions
160
- FROM final_ranks f
161
- LEFT JOIN u_anchoring a ON a.uid = f.id
162
- LEFT JOIN u_degree ud ON ud.uid = f.id
163
- WHERE f.kind = 'u'
164
- ORDER BY f.rank DESC
165
- LIMIT ${Number(limit)}
166
- ),
167
- nodes AS (
168
- SELECT * FROM query_rows
169
- UNION ALL
170
- SELECT * FROM url_rows
171
- ),
172
- counts AS (
173
- SELECT
174
- (SELECT n FROM query_count) AS q_count,
175
- (SELECT n FROM url_count) AS u_count
176
- ),
177
- deltas_json AS (
178
- SELECT to_json(list({ 'step': step, 'l1': l1 } ORDER BY step)) AS dj
179
- FROM deltas
57
+ url_nodes AS (
58
+ SELECT uid, ROW_NUMBER() OVER (ORDER BY uid) - 1 AS idx
59
+ FROM (SELECT DISTINCT uid FROM edges)
180
60
  )
181
61
  SELECT
182
- n.kind,
183
- n.id,
184
- n.rank,
185
- n.bridging,
186
- n.anchoring,
187
- n.degree,
188
- n.impressions,
189
- c.q_count AS queryCount,
190
- c.u_count AS urlCount,
191
- dj.dj AS deltasJson
192
- FROM nodes n
193
- CROSS JOIN counts c
194
- CROSS JOIN deltas_json dj
195
- ORDER BY n.kind, n.rank DESC
62
+ (SELECT to_json(list(qid ORDER BY idx)) FROM query_nodes) AS queriesJson,
63
+ (SELECT to_json(list(uid ORDER BY idx)) FROM url_nodes) AS urlsJson,
64
+ (SELECT to_json(list([q.idx, u.idx, e.impressions]))
65
+ FROM edges e
66
+ JOIN query_nodes q USING (qid)
67
+ JOIN url_nodes u USING (uid)) AS edgesJson,
68
+ CAST(${Number(limit)} AS BIGINT) AS resultLimit
69
+ LIMIT ${Number(limit)}
196
70
  `,
197
71
  params: [
198
72
  startDate,
@@ -206,33 +80,111 @@ const bipartitePagerankAnalyzer = defineAnalyzer({
206
80
  };
207
81
  },
208
82
  reduceSql(rows) {
209
- const arr = Array.isArray(rows) ? rows : [];
83
+ const first = (Array.isArray(rows) ? rows[0] : void 0) ?? {};
84
+ const queries = JSON.parse(rowString(first.queriesJson) || "[]");
85
+ const urls = JSON.parse(rowString(first.urlsJson) || "[]");
86
+ const edges = JSON.parse(rowString(first.edgesJson) || "[]");
210
87
  const iterations = BIPARTITE_PAGERANK_ITERATIONS;
211
- const d = BIPARTITE_PAGERANK_DAMPING;
212
- const results = arr.map((r) => ({
213
- kind: rowString(r.kind),
214
- id: rowString(r.id),
215
- rank: num(r.rank),
216
- bridging: num(r.bridging),
217
- anchoring: num(r.anchoring),
218
- degree: num(r.degree),
219
- impressions: num(r.impressions)
220
- }));
221
- const first = arr[0] ?? {};
222
- const queryCount = num(first.queryCount);
223
- const urlCount = num(first.urlCount);
224
- const deltas = parseJsonRows(first.deltasJson).map((e) => ({
225
- step: num(e.step),
226
- l1: num(e.l1)
88
+ const damping = BIPARTITE_PAGERANK_DAMPING;
89
+ if (!queries?.length || !urls?.length || !edges?.length) return {
90
+ results: [],
91
+ meta: {
92
+ total: 0,
93
+ convergenceDelta: 0,
94
+ iterations,
95
+ damping,
96
+ queryCount: 0,
97
+ urlCount: 0,
98
+ deltas: []
99
+ }
100
+ };
101
+ const queryCount = queries.length;
102
+ const urlCount = urls.length;
103
+ const queryTotals = new Float64Array(queryCount);
104
+ const urlTotals = new Float64Array(urlCount);
105
+ const queryDegrees = new Uint32Array(queryCount);
106
+ const urlDegrees = new Uint32Array(urlCount);
107
+ for (const [q, u, impressions] of edges) {
108
+ queryTotals[q] += impressions;
109
+ urlTotals[u] += impressions;
110
+ queryDegrees[q]++;
111
+ urlDegrees[u]++;
112
+ }
113
+ const qToU = new Float64Array(edges.length);
114
+ const uToQ = new Float64Array(edges.length);
115
+ const bridging = new Uint32Array(queryCount);
116
+ const anchoring = new Uint32Array(urlCount);
117
+ for (let i = 0; i < edges.length; i++) {
118
+ const [q, u, impressions] = edges[i];
119
+ qToU[i] = queryTotals[q] === 0 ? NaN : impressions / queryTotals[q];
120
+ uToQ[i] = urlTotals[u] === 0 ? NaN : impressions / urlTotals[u];
121
+ if (qToU[i] >= .05) bridging[q]++;
122
+ if (uToQ[i] >= .05) anchoring[u]++;
123
+ }
124
+ let queryRanks = new Float64Array(queryCount).fill(1 / queryCount);
125
+ let urlRanks = new Float64Array(urlCount).fill(1 / urlCount);
126
+ let nextQueries = new Float64Array(queryCount);
127
+ let nextUrls = new Float64Array(urlCount);
128
+ const deltas = [];
129
+ for (let step = 1; step <= iterations; step++) {
130
+ nextQueries.fill(NaN);
131
+ nextUrls.fill(NaN);
132
+ for (let i = 0; i < edges.length; i++) {
133
+ const [q, u] = edges[i];
134
+ const toQuery = uToQ[i] * urlRanks[u];
135
+ const toUrl = qToU[i] * queryRanks[q];
136
+ if (!Number.isNaN(toQuery)) nextQueries[q] = (Number.isNaN(nextQueries[q]) ? 0 : nextQueries[q]) + toQuery;
137
+ if (!Number.isNaN(toUrl)) nextUrls[u] = (Number.isNaN(nextUrls[u]) ? 0 : nextUrls[u]) + toUrl;
138
+ }
139
+ let l1 = 0;
140
+ for (let q = 0; q < queryCount; q++) {
141
+ nextQueries[q] = .15000000000000002 / queryCount + damping * nextQueries[q];
142
+ const delta = Math.abs(nextQueries[q] - queryRanks[q]);
143
+ if (!Number.isNaN(delta)) l1 += delta;
144
+ }
145
+ for (let u = 0; u < urlCount; u++) {
146
+ nextUrls[u] = .15000000000000002 / urlCount + damping * nextUrls[u];
147
+ const delta = Math.abs(nextUrls[u] - urlRanks[u]);
148
+ if (!Number.isNaN(delta)) l1 += delta;
149
+ }
150
+ deltas.push({
151
+ step,
152
+ l1
153
+ });
154
+ [queryRanks, nextQueries] = [nextQueries, queryRanks];
155
+ [urlRanks, nextUrls] = [nextUrls, urlRanks];
156
+ }
157
+ const limit = num(first.resultLimit);
158
+ const byRank = (a, b) => Number.isNaN(a.rank) ? Number.isNaN(b.rank) ? 0 : 1 : Number.isNaN(b.rank) ? -1 : b.rank - a.rank;
159
+ const queryNodes = queries.map((id, i) => ({
160
+ kind: "query",
161
+ id,
162
+ rank: queryRanks[i],
163
+ bridging: bridging[i],
164
+ anchoring: 0,
165
+ degree: queryDegrees[i],
166
+ impressions: queryTotals[i]
167
+ })).sort(byRank).slice(0, limit);
168
+ const urlNodes = urls.map((id, i) => ({
169
+ kind: "url",
170
+ id,
171
+ rank: urlRanks[i],
172
+ bridging: 0,
173
+ anchoring: anchoring[i],
174
+ degree: urlDegrees[i],
175
+ impressions: urlTotals[i]
176
+ })).sort(byRank).slice(0, limit);
177
+ const results = [...queryNodes, ...urlNodes].map((node) => ({
178
+ ...node,
179
+ rank: Number.isNaN(node.rank) ? 0 : node.rank
227
180
  }));
228
- const convergenceDelta = deltas.length > 0 ? deltas[deltas.length - 1].l1 : 0;
229
181
  return {
230
182
  results,
231
183
  meta: {
232
184
  total: results.length,
233
- convergenceDelta,
185
+ convergenceDelta: deltas[24].l1,
234
186
  iterations,
235
- damping: d,
187
+ damping,
236
188
  queryCount,
237
189
  urlCount,
238
190
  deltas
@@ -1,25 +1,25 @@
1
1
  import { QueriesRow } from "../types.mjs";
2
- interface BrandSegmentationOptions {
2
+ export interface BrandSegmentationOptions {
3
3
  /** Brand terms to match against keywords (case-insensitive) */
4
4
  brandTerms: string[];
5
5
  /** Minimum impressions for a keyword to be included. Default: 10 */
6
6
  minImpressions?: number;
7
7
  }
8
- interface BrandSegmentationRow {
8
+ export interface BrandSegmentationRow {
9
9
  query?: string;
10
10
  keyword?: string;
11
11
  variants?: readonly string[];
12
12
  clicks?: number | null;
13
13
  impressions?: number | null;
14
14
  }
15
- interface BrandSummary {
15
+ export interface BrandSummary {
16
16
  brandClicks: number;
17
17
  nonBrandClicks: number;
18
18
  brandShare: number;
19
19
  brandImpressions: number;
20
20
  nonBrandImpressions: number;
21
21
  }
22
- interface BrandSegmentationResult<T extends BrandSegmentationRow = QueriesRow> {
22
+ export interface BrandSegmentationResult<T extends BrandSegmentationRow = QueriesRow> {
23
23
  brand: T[];
24
24
  nonBrand: T[];
25
25
  summary: BrandSummary;
@@ -28,6 +28,5 @@ interface BrandSegmentationResult<T extends BrandSegmentationRow = QueriesRow> {
28
28
  * Pure helper: segment keywords into brand and non-brand based on provided
29
29
  * brand terms. Re-exported from `@gscdump/analysis` for portable callers.
30
30
  */
31
- declare function analyzeBrandSegmentation(keywords: QueriesRow[], options: BrandSegmentationOptions): BrandSegmentationResult<QueriesRow>;
32
- declare function analyzeBrandSegmentation<T extends BrandSegmentationRow>(keywords: T[], options: BrandSegmentationOptions): BrandSegmentationResult<T>;
33
- export { BrandSegmentationOptions, BrandSegmentationResult, BrandSegmentationRow, BrandSummary, analyzeBrandSegmentation };
31
+ export declare function analyzeBrandSegmentation(keywords: QueriesRow[], options: BrandSegmentationOptions): BrandSegmentationResult<QueriesRow>;
32
+ export declare function analyzeBrandSegmentation<T extends BrandSegmentationRow>(keywords: T[], options: BrandSegmentationOptions): BrandSegmentationResult<T>;