@titan-design/session-graph 0.6.0 → 0.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +53 -3
- package/dist/index.d.ts +184 -6
- package/dist/index.js +336 -10
- package/dist/index.js.map +1 -1
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -41,12 +41,14 @@ passes 1002.
|
|
|
41
41
|
3. `rollupSessions` recomputes turn aggregates (index, end, duration, tool calls, thinking
|
|
42
42
|
time) for the sessions that changed, including those the backfill touched. Recompute, never accumulate, so incremental and
|
|
43
43
|
full passes converge.
|
|
44
|
-
4. `
|
|
44
|
+
4. `resolveOrigins` runs if the caller passed a `resolveOrigins` resolver. See migration 5.
|
|
45
|
+
5. `reconcile` folds cross-transcript observations: `gh pr merge` sightings onto PRs,
|
|
45
46
|
complete `gh pr create` sightings into new PR rows, subagent end times and parentage
|
|
46
47
|
from child sessions.
|
|
47
|
-
|
|
48
|
+
6. `enrichPrs` runs if the caller passed a `resolvePrs` resolver. See below.
|
|
49
|
+
7. `enrichTasks` runs if the caller passed a `resolveTasks` resolver, once over the whole
|
|
48
50
|
task table. See below.
|
|
49
|
-
|
|
51
|
+
8. Rows whose source file has vanished are marked `missing`. Their facts stay: surviving
|
|
50
52
|
Claude Code's own pruning is much of the point.
|
|
51
53
|
|
|
52
54
|
`resetIndex` clears every derived table and rewinds watermarks; the next refresh rebuilds
|
|
@@ -68,12 +70,39 @@ await refreshCorpus(graph, transcripts, {
|
|
|
68
70
|
- **Precedence.** Every field the resolver states wins; every field it omits or nulls keeps
|
|
69
71
|
what the transcripts derived. The store is the system of record for a task's present
|
|
70
72
|
status, while a transcript only witnesses a command that was observed to run.
|
|
73
|
+
- **Estimate.** `estimate` fills `task.estimate`, in whatever unit the store keeps.
|
|
71
74
|
- **Batching.** One call per refresh, holding every task id in the graph, because a real
|
|
72
75
|
resolver reads a database. Returning an id no transcript mentioned inserts that task.
|
|
73
76
|
- **Failure.** A resolver that throws costs that pass its enrichment and nothing else; the
|
|
74
77
|
rows stand as the transcripts left them and `summary.tasks` carries `failed` plus the
|
|
75
78
|
`error` message for the caller to log.
|
|
76
79
|
|
|
80
|
+
## The PR outcome resolver
|
|
81
|
+
|
|
82
|
+
A transcript sees a merge only when that session ran `gh pr merge`. A merge done by another
|
|
83
|
+
session or in the browser, a close, and the review history exist nowhere in the corpus. Pass
|
|
84
|
+
`resolvePrs` and a product fills `state`, `merged_at`, `closed_at` and `review_rounds`.
|
|
85
|
+
|
|
86
|
+
```ts
|
|
87
|
+
await refreshCorpus(graph, transcripts, {
|
|
88
|
+
resolvePrs: async (prs) => new Map(prs.map((pr) => [pr.prRef, myForge.outcome(pr.repo, pr.number)])),
|
|
89
|
+
});
|
|
90
|
+
```
|
|
91
|
+
|
|
92
|
+
- **Order.** It runs after `reconcile`, so it sees every merge the transcripts witnessed.
|
|
93
|
+
- **Merged is sticky.** A resolver's `open` or `closed` never replaces a `merged` state, so a
|
|
94
|
+
stale forge cache cannot reopen a PR. States are stored lower-case.
|
|
95
|
+
- **Batching.** One call per refresh with every PR that has a repo and number and whose
|
|
96
|
+
outcome may still change: never checked, or not yet merged. A merged PR is asked about
|
|
97
|
+
once. The resolver only updates rows; a PR enters the graph from a transcript.
|
|
98
|
+
- **`review_rounds`** is stored as the resolver counts it. What counts as a round is an open
|
|
99
|
+
question in the TP-256 design (Q6).
|
|
100
|
+
- **Failure.** Same as the task resolver: the pass completes, the rows stand, and
|
|
101
|
+
`summary.prs` carries `failed` plus `error`.
|
|
102
|
+
|
|
103
|
+
`reconcile` rewrites `merged_at` from merge sightings on every pass, so for a PR a transcript
|
|
104
|
+
saw merged, `merged_at` holds the sighting's time rather than the forge's.
|
|
105
|
+
|
|
77
106
|
## Tables
|
|
78
107
|
|
|
79
108
|
Kit tables: `transcript` (watermark), `edge` (bi-temporal, `session:… touched file:…`),
|
|
@@ -119,6 +148,27 @@ the several assistant lines of one API response collapse to the first line's off
|
|
|
119
148
|
the largest value of each token column. `session.account` comes from discovery, not the
|
|
120
149
|
transcript. `purgeTranscript` and `resetIndex` clear all nine tables.
|
|
121
150
|
|
|
151
|
+
Migration 5, `origin, episodes, prices`, adds four tables and three views. It creates only
|
|
152
|
+
empty tables, so existing rows are untouched.
|
|
153
|
+
|
|
154
|
+
- `session_origin` and `session_external_event` hold who launched a session and the
|
|
155
|
+
lifecycle events the launcher saw outside the transcript. `refreshCorpus` fills them
|
|
156
|
+
through the `resolveOrigins` option, once per pass, for sessions with no origin row or one
|
|
157
|
+
older than the session's last line. Each origin with a parent projects into a `spawned`
|
|
158
|
+
edge and a `subagent` row. `resetIndex` clears them and the next pass refills them.
|
|
159
|
+
- `episode` holds each segmentation heuristic's cut of a session.
|
|
160
|
+
`replaceEpisodes(graph, sessionId, heuristic, rows)` is its only writer. It replaces one
|
|
161
|
+
heuristic's rows for one session in a transaction and leaves every other heuristic's rows
|
|
162
|
+
alone, so `worker-v1` and `coordinator-v1` coexist. The rule itself lives in
|
|
163
|
+
session-analytics. A rewritten transcript purges its sessions' episodes.
|
|
164
|
+
- `price` holds USD per million tokens by model prefix and effective date.
|
|
165
|
+
`syncPrices(graph, rows, { tableVersion, source })` replaces the whole table in one
|
|
166
|
+
transaction.
|
|
167
|
+
- `request_dedup` collapses fan-out copies of a request to the earliest one. Every cost
|
|
168
|
+
query reads it, never `request`. `request_cost` prices each row by longest model prefix
|
|
169
|
+
and latest `effective_from`; an unmatched model reads `priced = 0` and costs 0.
|
|
170
|
+
`context_contribution` attributes each request's context growth to the blocks before it.
|
|
171
|
+
|
|
122
172
|
For snapshot-only usage across multiple physical sources, queries select one source
|
|
123
173
|
by latest native usage timestamp, then greatest usage-record coverage and stable
|
|
124
174
|
source ID. Source-local reset epochs cannot safely be summed across copies. This is
|
package/dist/index.d.ts
CHANGED
|
@@ -39,7 +39,7 @@ declare const KIT: {
|
|
|
39
39
|
*/
|
|
40
40
|
declare const DOMAIN_DDL = "\n CREATE TABLE IF NOT EXISTS fact (\n fact_id INTEGER PRIMARY KEY,\n transcript_id INTEGER NOT NULL,\n byte_offset INTEGER NOT NULL,\n byte_length INTEGER NOT NULL,\n event_type TEXT NOT NULL,\n ts TEXT NOT NULL,\n seq INTEGER NOT NULL,\n session_id TEXT NOT NULL,\n prompt_id TEXT,\n tool_use_id TEXT,\n t_indexed TEXT NOT NULL DEFAULT (strftime('%Y-%m-%dT%H:%M:%fZ','now')),\n UNIQUE (transcript_id, byte_offset)\n );\n CREATE INDEX IF NOT EXISTS idx_fact_session_ts ON fact(session_id, ts);\n CREATE INDEX IF NOT EXISTS idx_fact_prompt ON fact(prompt_id);\n CREATE INDEX IF NOT EXISTS idx_fact_tool_use ON fact(tool_use_id);\n\n CREATE TABLE IF NOT EXISTS session (\n session_id TEXT PRIMARY KEY,\n transcript_id INTEGER,\n started_at TEXT,\n ended_at TEXT,\n start_type TEXT,\n cwd TEXT,\n git_branch TEXT,\n ai_title TEXT,\n seed_prompt TEXT,\n cli_version TEXT,\n turn_count INTEGER NOT NULL DEFAULT 0,\n commit_count INTEGER NOT NULL DEFAULT 0,\n push_count INTEGER NOT NULL DEFAULT 0\n );\n CREATE INDEX IF NOT EXISTS idx_session_started ON session(started_at);\n\n CREATE TABLE IF NOT EXISTS session_model_usage (\n session_id TEXT NOT NULL,\n model TEXT NOT NULL,\n input_tokens INTEGER NOT NULL DEFAULT 0,\n output_tokens INTEGER NOT NULL DEFAULT 0,\n cache_read_tokens INTEGER NOT NULL DEFAULT 0,\n cache_creation_tokens INTEGER NOT NULL DEFAULT 0,\n thinking_tokens INTEGER NOT NULL DEFAULT 0,\n request_count INTEGER NOT NULL DEFAULT 0,\n PRIMARY KEY (session_id, model)\n );\n\n CREATE TABLE IF NOT EXISTS turn (\n prompt_id TEXT PRIMARY KEY,\n session_id TEXT NOT NULL,\n turn_index INTEGER NOT NULL,\n started_at TEXT NOT NULL,\n ended_at TEXT,\n duration_ms INTEGER,\n tool_call_count INTEGER NOT NULL DEFAULT 0,\n thinking_ms INTEGER NOT NULL DEFAULT 0,\n fact_id_start INTEGER\n );\n CREATE INDEX IF NOT EXISTS idx_turn_session ON turn(session_id, turn_index);\n\n CREATE TABLE IF NOT EXISTS permission_phase (\n phase_id INTEGER PRIMARY KEY,\n session_id TEXT NOT NULL,\n from_mode TEXT,\n to_mode TEXT NOT NULL,\n trigger TEXT NOT NULL,\n t_valid TEXT NOT NULL,\n t_invalid TEXT,\n fact_id INTEGER\n );\n CREATE INDEX IF NOT EXISTS idx_phase_session ON permission_phase(session_id, t_valid);\n\n CREATE TABLE IF NOT EXISTS human_edit (\n edit_id INTEGER PRIMARY KEY,\n session_id TEXT NOT NULL,\n file_path TEXT NOT NULL,\n ts TEXT NOT NULL,\n fact_id INTEGER,\n UNIQUE (session_id, file_path, ts)\n );\n\n CREATE TABLE IF NOT EXISTS file_checkpoint (\n checkpoint_id INTEGER PRIMARY KEY,\n session_id TEXT NOT NULL,\n file_path TEXT NOT NULL,\n backup_file_name TEXT NOT NULL,\n version INTEGER NOT NULL,\n backup_time TEXT NOT NULL,\n fact_id INTEGER,\n UNIQUE (session_id, file_path, backup_file_name)\n );\n\n CREATE TABLE IF NOT EXISTS pr (\n pr_ref TEXT PRIMARY KEY, number INTEGER, repo TEXT, title TEXT, state TEXT, url TEXT, merged_at TEXT\n );\n CREATE TABLE IF NOT EXISTS branch (\n branch_ref TEXT PRIMARY KEY, repo TEXT, name TEXT NOT NULL, base TEXT, created_at TEXT, deleted_at TEXT\n );\n CREATE TABLE IF NOT EXISTS file (\n file_ref TEXT PRIMARY KEY, repo TEXT, path TEXT NOT NULL\n );\n CREATE TABLE IF NOT EXISTS task (\n task_ref TEXT PRIMARY KEY, task_id TEXT NOT NULL, initiative TEXT, title TEXT, status TEXT\n );\n CREATE TABLE IF NOT EXISTS subagent (\n agent_ref TEXT PRIMARY KEY,\n session_id TEXT,\n child_session_id TEXT,\n parent_agent_ref TEXT,\n agent_type TEXT,\n label TEXT,\n started_at TEXT,\n ended_at TEXT,\n fact_id INTEGER\n );\n CREATE INDEX IF NOT EXISTS idx_subagent_child ON subagent(child_session_id);\n CREATE TABLE IF NOT EXISTS artifact (\n artifact_ref TEXT PRIMARY KEY, kind TEXT, title TEXT, url TEXT, path TEXT, created_at TEXT\n );\n CREATE TABLE IF NOT EXISTS pr_merge_observation (\n number INTEGER NOT NULL, repo_hint TEXT, merged_at TEXT NOT NULL,\n PRIMARY KEY (number, repo_hint, merged_at)\n );\n CREATE TABLE IF NOT EXISTS pr_create_observation (\n tool_use_id TEXT PRIMARY KEY, title TEXT, number INTEGER, repo TEXT, url TEXT\n );\n";
|
|
41
41
|
/** Every derived table, in an order safe to clear. The watermark table is not derived. */
|
|
42
|
-
declare const DERIVED_TABLES: readonly ["normalized_span", "normalized_event", "normalized_source", "search_span", "edge", "turn", "permission_phase", "human_edit", "file_checkpoint", "subagent", "request", "tool_call", "inbound", "context_block", "compaction", "queue_op", "session_signal", "cost_state_observation", "transcript_facet", "session_model_usage", "session", "fact", "pr", "pr_merge_observation", "pr_create_observation", "branch", "file", "task", "artifact"];
|
|
42
|
+
declare const DERIVED_TABLES: readonly ["normalized_span", "normalized_event", "normalized_source", "search_span", "edge", "turn", "permission_phase", "human_edit", "file_checkpoint", "subagent", "request", "tool_call", "inbound", "context_block", "compaction", "queue_op", "session_signal", "cost_state_observation", "transcript_facet", "session_origin", "session_external_event", "episode", "session_model_usage", "session", "fact", "pr", "pr_merge_observation", "pr_create_observation", "branch", "file", "task", "artifact"];
|
|
43
43
|
declare const MIGRATIONS: Migration[];
|
|
44
44
|
|
|
45
45
|
/** Recorded in `_migration`; store-sqlite refuses a database whose applied name differs, so never rename it. */
|
|
@@ -50,6 +50,15 @@ declare const AUDIT_TABLES: readonly ["request", "tool_call", "inbound", "contex
|
|
|
50
50
|
declare const FACET_TABLE = "transcript_facet";
|
|
51
51
|
declare const AUDIT_DDL = "\n CREATE TABLE IF NOT EXISTS request (\n transcript_id INTEGER NOT NULL,\n request_id TEXT NOT NULL,\n byte_offset INTEGER NOT NULL, -- first line seen for this request\n session_id TEXT NOT NULL,\n message_id TEXT,\n ts TEXT NOT NULL,\n model TEXT NOT NULL,\n input_tokens INTEGER NOT NULL DEFAULT 0,\n cache_read_tokens INTEGER NOT NULL DEFAULT 0,\n cache_creation_tokens INTEGER NOT NULL DEFAULT 0,\n cache_creation_5m INTEGER NOT NULL DEFAULT 0,\n cache_creation_1h INTEGER NOT NULL DEFAULT 0,\n output_tokens INTEGER NOT NULL DEFAULT 0,\n thinking_tokens INTEGER NOT NULL DEFAULT 0,\n context_tokens INTEGER NOT NULL DEFAULT 0, -- input + cache_read + cache_creation\n service_tier TEXT,\n is_sidechain INTEGER NOT NULL DEFAULT 0,\n -- rollup-owned, recomputed for touched sessions:\n seq_in_session INTEGER,\n gap_ms INTEGER, -- since previous request in the session\n ctx_delta INTEGER, -- context_tokens - (prev.context_tokens + prev.output_tokens)\n wake_cause TEXT,\n wake_delivery TEXT,\n wake_detail TEXT, -- tool name, channel sender, or NULL\n wake_tool_family TEXT,\n wake_mcp_server TEXT,\n wake_offset INTEGER, -- byte_offset of the inbound row\n PRIMARY KEY (transcript_id, request_id)\n ) WITHOUT ROWID;\n CREATE INDEX IF NOT EXISTS idx_request_session_ts ON request(session_id, ts);\n CREATE INDEX IF NOT EXISTS idx_request_ts ON request(ts);\n CREATE INDEX IF NOT EXISTS idx_request_id ON request(request_id);\n\n CREATE TABLE IF NOT EXISTS tool_call (\n transcript_id INTEGER NOT NULL, byte_offset INTEGER NOT NULL, block_index INTEGER NOT NULL,\n session_id TEXT NOT NULL, ts TEXT NOT NULL, tool_use_id TEXT NOT NULL,\n name TEXT NOT NULL, family TEXT NOT NULL, mcp_server TEXT, input_chars INTEGER NOT NULL,\n PRIMARY KEY (transcript_id, byte_offset, block_index)\n ) WITHOUT ROWID;\n CREATE INDEX IF NOT EXISTS idx_tool_call_use ON tool_call(tool_use_id);\n\n CREATE TABLE IF NOT EXISTS inbound (\n transcript_id INTEGER NOT NULL, byte_offset INTEGER NOT NULL, block_index INTEGER NOT NULL,\n session_id TEXT NOT NULL, ts TEXT NOT NULL,\n cause TEXT NOT NULL, delivery TEXT NOT NULL, detail TEXT,\n origin_server TEXT, from_name TEXT, msg_id TEXT, tool_use_id TEXT,\n is_error INTEGER NOT NULL DEFAULT 0, content_hash TEXT NOT NULL, chars INTEGER NOT NULL,\n queued_ms INTEGER, -- rollup-owned\n PRIMARY KEY (transcript_id, byte_offset, block_index)\n ) WITHOUT ROWID;\n CREATE INDEX IF NOT EXISTS idx_inbound_session ON inbound(session_id, byte_offset);\n\n CREATE TABLE IF NOT EXISTS context_block (\n transcript_id INTEGER NOT NULL, byte_offset INTEGER NOT NULL, block_index INTEGER NOT NULL,\n session_id TEXT NOT NULL, ts TEXT NOT NULL, source TEXT NOT NULL,\n tool_use_id TEXT, attachment_type TEXT, chars INTEGER NOT NULL, is_media INTEGER NOT NULL DEFAULT 0,\n PRIMARY KEY (transcript_id, byte_offset, block_index)\n ) WITHOUT ROWID;\n CREATE INDEX IF NOT EXISTS idx_context_block_session ON context_block(session_id, source);\n\n CREATE TABLE IF NOT EXISTS compaction (\n transcript_id INTEGER NOT NULL, byte_offset INTEGER NOT NULL,\n session_id TEXT NOT NULL, ts TEXT NOT NULL, trigger TEXT,\n pre_tokens INTEGER, post_tokens INTEGER, dropped_tokens INTEGER, duration_ms INTEGER,\n mid_loop INTEGER, -- rollup-owned: previous inbound was a tool_result\n PRIMARY KEY (transcript_id, byte_offset)\n ) WITHOUT ROWID;\n\n CREATE TABLE IF NOT EXISTS queue_op (\n transcript_id INTEGER NOT NULL, byte_offset INTEGER NOT NULL,\n session_id TEXT NOT NULL, ts TEXT NOT NULL, operation TEXT NOT NULL,\n content_hash TEXT, origin_server TEXT,\n PRIMARY KEY (transcript_id, byte_offset)\n ) WITHOUT ROWID;\n CREATE INDEX IF NOT EXISTS idx_queue_op_hash ON queue_op(session_id, content_hash);\n\n CREATE TABLE IF NOT EXISTS session_signal (\n transcript_id INTEGER NOT NULL, byte_offset INTEGER NOT NULL, block_index INTEGER NOT NULL,\n session_id TEXT NOT NULL, ts TEXT NOT NULL, signal TEXT NOT NULL, detail TEXT, tool_use_id TEXT,\n -- one Bash block can emit up to three signals (commit, push, pr_create); the signal name disambiguates them\n PRIMARY KEY (transcript_id, byte_offset, block_index, signal)\n ) WITHOUT ROWID;\n CREATE INDEX IF NOT EXISTS idx_signal_session ON session_signal(session_id, ts);\n\n CREATE TABLE IF NOT EXISTS cost_state_observation (\n transcript_id INTEGER NOT NULL, byte_offset INTEGER NOT NULL,\n session_id TEXT NOT NULL, ts TEXT, total_cost_usd REAL NOT NULL, model_usage TEXT NOT NULL,\n PRIMARY KEY (transcript_id, byte_offset)\n ) WITHOUT ROWID;\n\n CREATE TABLE IF NOT EXISTS transcript_facet (\n transcript_id INTEGER NOT NULL, facet TEXT NOT NULL,\n version INTEGER NOT NULL, indexed_to INTEGER NOT NULL,\n PRIMARY KEY (transcript_id, facet)\n ) WITHOUT ROWID;\n";
|
|
52
52
|
|
|
53
|
+
/** Recorded in `_migration`; store-sqlite refuses a database whose applied name differs, so never rename it. */
|
|
54
|
+
declare const ORIGIN_MIGRATION_NAME = "origin, episodes, prices";
|
|
55
|
+
/** Re-derivable from their own sources rather than from transcripts, so `resetIndex` clears them and the next pass refills them. */
|
|
56
|
+
declare const ORIGIN_TABLES: readonly ["session_origin", "session_external_event"];
|
|
57
|
+
/** Written from transcript offsets by `replaceEpisodes`; a rewritten transcript invalidates them, so purge drops them by session. */
|
|
58
|
+
declare const EPISODE_TABLE = "episode";
|
|
59
|
+
declare const ORIGIN_DDL = "\n CREATE TABLE IF NOT EXISTS session_origin ( -- not derived from transcripts\n session_id TEXT PRIMARY KEY,\n origin_system TEXT NOT NULL, -- 'agent-chat'\n agent_id TEXT, agent_name TEXT, parent_name TEXT, parent_session_id TEXT,\n profile TEXT, model_alias TEXT, surface TEXT, isolation TEXT, depth INTEGER,\n origin_kind TEXT, -- spawned | adopted | inherited | human\n config_dir TEXT, spawn_cwd TEXT, spawned_at TEXT,\n brief_chars INTEGER, brief_excerpt TEXT, brief_path TEXT, launch_args TEXT,\n resolved_at TEXT NOT NULL\n );\n CREATE INDEX IF NOT EXISTS idx_origin_parent ON session_origin(parent_session_id);\n\n CREATE TABLE IF NOT EXISTS session_external_event ( -- teleport, handoff, retire, exit\n session_id TEXT NOT NULL, ts TEXT NOT NULL, origin_system TEXT NOT NULL,\n kind TEXT NOT NULL, detail TEXT,\n PRIMARY KEY (session_id, ts, kind)\n ) WITHOUT ROWID;\n\n CREATE TABLE IF NOT EXISTS episode ( -- provisional, see section 9\n session_id TEXT NOT NULL, episode_index INTEGER NOT NULL,\n heuristic TEXT NOT NULL, heuristic_version INTEGER NOT NULL,\n started_at TEXT NOT NULL, ended_at TEXT NOT NULL,\n start_offset INTEGER NOT NULL, end_offset INTEGER NOT NULL,\n opened_by TEXT NOT NULL, -- brief | channel_followup | idle_gap | compaction\n assignment_offset INTEGER, first_deliverable_offset INTEGER, first_deliverable_signal TEXT,\n first_status_offset INTEGER,\n PRIMARY KEY (session_id, heuristic, episode_index)\n ) WITHOUT ROWID;\n\n CREATE TABLE IF NOT EXISTS price (\n model TEXT NOT NULL, effective_from TEXT NOT NULL, table_version INTEGER NOT NULL,\n input_usd_mtok REAL NOT NULL, cache_read_usd_mtok REAL NOT NULL,\n cache_write_5m_usd_mtok REAL NOT NULL, cache_write_1h_usd_mtok REAL NOT NULL,\n output_usd_mtok REAL NOT NULL, source TEXT,\n PRIMARY KEY (model, effective_from)\n ) WITHOUT ROWID;\n";
|
|
60
|
+
declare const ORIGIN_VIEWS: readonly ["\n CREATE VIEW IF NOT EXISTS request_dedup AS\n SELECT transcript_id, request_id, byte_offset, session_id, message_id, ts, model,\n input_tokens, cache_read_tokens, cache_creation_tokens, cache_creation_5m, cache_creation_1h,\n output_tokens, thinking_tokens, context_tokens, service_tier, is_sidechain,\n seq_in_session, gap_ms, ctx_delta, wake_cause, wake_delivery, wake_detail,\n wake_tool_family, wake_mcp_server, wake_offset FROM (\n SELECT r.*, ROW_NUMBER() OVER (PARTITION BY r.request_id ORDER BY r.ts, r.transcript_id) AS copy_rank\n FROM request r\n ) WHERE copy_rank = 1;\n", "\n CREATE VIEW IF NOT EXISTS request_cost AS\n WITH candidate AS (\n SELECT d.request_id, p.model AS price_model, p.effective_from AS price_effective_from,\n p.input_usd_mtok, p.cache_read_usd_mtok, p.cache_write_5m_usd_mtok, p.cache_write_1h_usd_mtok, p.output_usd_mtok,\n ROW_NUMBER() OVER (PARTITION BY d.request_id ORDER BY length(p.model) DESC, p.effective_from DESC) AS match_rank\n FROM request_dedup d\n JOIN price p ON substr(d.model, 1, length(p.model)) = p.model AND p.effective_from <= d.ts\n ), component AS (\n SELECT d.*, c.price_model, c.price_effective_from, c.price_model IS NOT NULL AS priced,\n COALESCE(d.input_tokens * c.input_usd_mtok, 0) / 1e6 AS input_cost_usd,\n COALESCE(d.cache_read_tokens * c.cache_read_usd_mtok, 0) / 1e6 AS cache_read_cost_usd,\n COALESCE(CASE WHEN d.cache_creation_5m + d.cache_creation_1h = 0 THEN d.cache_creation_tokens ELSE d.cache_creation_5m END\n * c.cache_write_5m_usd_mtok, 0) / 1e6 AS cache_write_5m_cost_usd,\n COALESCE(d.cache_creation_1h * c.cache_write_1h_usd_mtok, 0) / 1e6 AS cache_write_1h_cost_usd,\n COALESCE(d.output_tokens * c.output_usd_mtok, 0) / 1e6 AS output_cost_usd\n FROM request_dedup d\n LEFT JOIN candidate c ON c.request_id = d.request_id AND c.match_rank = 1\n )\n SELECT *,\n input_cost_usd + cache_read_cost_usd + cache_write_5m_cost_usd + cache_write_1h_cost_usd + output_cost_usd AS cost_usd,\n (cache_read_tokens < 0.2 * context_tokens AND cache_creation_tokens >= 20000) AS is_cold,\n CASE WHEN context_tokens < 50000 THEN '<50k' WHEN context_tokens < 100000 THEN '50-100k'\n WHEN context_tokens < 200000 THEN '100-200k' ELSE '200k+' END AS context_band,\n CASE WHEN gap_ms IS NULL THEN NULL WHEN gap_ms < 300000 THEN '<5m'\n WHEN gap_ms < 3600000 THEN '5-60m' ELSE '>60m' END AS gap_band\n FROM component;\n", "\n CREATE VIEW IF NOT EXISTS context_contribution AS\n WITH stream AS (\n SELECT transcript_id, byte_offset, 0 AS lane, block_index, NULL AS request_id FROM context_block\n UNION ALL\n SELECT transcript_id, byte_offset, -1 AS lane, 0, request_id FROM request\n ), ordered AS (\n SELECT *, COUNT(request_id) OVER (PARTITION BY transcript_id ORDER BY byte_offset, lane, block_index ROWS UNBOUNDED PRECEDING) AS requests_before\n FROM stream\n ), owner AS (\n SELECT b.transcript_id, b.byte_offset, b.block_index, r.request_id\n FROM ordered b JOIN ordered r ON r.transcript_id = b.transcript_id AND r.lane = -1 AND r.requests_before = b.requests_before + 1\n WHERE b.lane = 0\n ), tool AS (\n SELECT tool_use_id, name, family, mcp_server,\n ROW_NUMBER() OVER (PARTITION BY tool_use_id ORDER BY transcript_id, byte_offset, block_index) AS copy_rank\n FROM tool_call\n )\n SELECT cb.transcript_id, cb.byte_offset, cb.block_index, cb.session_id, cb.ts, cb.source,\n cb.tool_use_id, cb.attachment_type, cb.chars, cb.is_media, o.request_id, d.ctx_delta,\n CASE WHEN cb.source IN ('assistant_text', 'assistant_thinking', 'assistant_tool_input') THEN NULL\n ELSE d.ctx_delta * cb.chars * 1.0 / NULLIF(SUM(CASE WHEN cb.source IN ('assistant_text', 'assistant_thinking', 'assistant_tool_input') THEN 0 ELSE cb.chars END)\n OVER (PARTITION BY o.transcript_id, o.request_id), 0) END AS est_tokens,\n t.name AS tool_name, t.family AS tool_family, t.mcp_server\n FROM context_block cb\n JOIN owner o USING (transcript_id, byte_offset, block_index)\n JOIN request_dedup d ON d.transcript_id = o.transcript_id AND d.request_id = o.request_id\n LEFT JOIN tool t ON t.tool_use_id = cb.tool_use_id AND t.copy_rank = 1;\n"];
|
|
61
|
+
|
|
53
62
|
/** Write one delta's audit events. Runs inside `applyDelta`'s transaction, so it opens none of its own. */
|
|
54
63
|
declare function applyAudit(db: Db, transcriptId: number, delta: TranscriptDelta): void;
|
|
55
64
|
|
|
@@ -125,6 +134,121 @@ interface ReconcileCounts {
|
|
|
125
134
|
/** Whole-table reconciliations, run once at the end of a refresh pass after the rollup. */
|
|
126
135
|
declare function reconcile(graph: SessionGraph): ReconcileCounts;
|
|
127
136
|
|
|
137
|
+
/** A PR the graph holds, as the resolver needs it to ask a forge. */
|
|
138
|
+
interface PrKey {
|
|
139
|
+
prRef: string;
|
|
140
|
+
repo: string;
|
|
141
|
+
number: number;
|
|
142
|
+
}
|
|
143
|
+
/**
|
|
144
|
+
* What a forge knows and a transcript cannot state: a merge done by another
|
|
145
|
+
* session or in the browser, a close, and the review history. Every field is
|
|
146
|
+
* optional; an omitted or null field leaves whatever the transcripts derived.
|
|
147
|
+
*/
|
|
148
|
+
interface ResolvedPr {
|
|
149
|
+
/** `open`, `closed` or `merged`, in any case. */
|
|
150
|
+
state?: string | null;
|
|
151
|
+
mergedAt?: string | null;
|
|
152
|
+
closedAt?: string | null;
|
|
153
|
+
/** Stored as the resolver counts it; the round definition is open question Q6 in the TP-256 design. */
|
|
154
|
+
reviewRounds?: number | null;
|
|
155
|
+
}
|
|
156
|
+
/** Resolutions keyed by `pr_ref` (`pr:acme/demo#7`). */
|
|
157
|
+
type PrResolution = ReadonlyMap<string, ResolvedPr | null | undefined>;
|
|
158
|
+
/**
|
|
159
|
+
* Supplied by the caller, never by this package: session-graph is tier 2 and
|
|
160
|
+
* must not learn which forge a product talks to. Called once per pass with every
|
|
161
|
+
* PR whose outcome may still change.
|
|
162
|
+
*/
|
|
163
|
+
type PrResolver = (prs: readonly PrKey[]) => PromiseLike<PrResolution> | PrResolution;
|
|
164
|
+
interface PrEnrichment {
|
|
165
|
+
/** PRs handed to the resolver. */
|
|
166
|
+
requested: number;
|
|
167
|
+
/** PR rows the resolver updated. */
|
|
168
|
+
applied: number;
|
|
169
|
+
/** The resolver threw; rows stand as the transcripts left them. */
|
|
170
|
+
failed: boolean;
|
|
171
|
+
/** Why it failed, for the caller to log. Absent unless `failed`. */
|
|
172
|
+
error?: string;
|
|
173
|
+
}
|
|
174
|
+
declare const NO_PR_OUTCOMES: PrEnrichment;
|
|
175
|
+
declare function prsNeedingOutcome(graph: SessionGraph): PrKey[];
|
|
176
|
+
/**
|
|
177
|
+
* Fill PR outcomes from a caller-supplied resolver. Runs after `reconcile`, so
|
|
178
|
+
* it sees the merges the transcripts witnessed and only adds what they missed.
|
|
179
|
+
* Updates rows and never inserts one: a PR enters the graph from a transcript.
|
|
180
|
+
* A resolver that throws costs this pass its outcomes and nothing else.
|
|
181
|
+
*/
|
|
182
|
+
declare function enrichPrs(graph: SessionGraph, resolver: PrResolver | undefined): Promise<PrEnrichment>;
|
|
183
|
+
|
|
184
|
+
/**
|
|
185
|
+
* Who started a session and why, as a launcher such as agent-chat recorded it.
|
|
186
|
+
* A transcript cannot state this: a `claude -p` worker and a headless miner look
|
|
187
|
+
* the same from inside. Every field but `originSystem` is optional.
|
|
188
|
+
*/
|
|
189
|
+
interface ResolvedOrigin {
|
|
190
|
+
originSystem: string;
|
|
191
|
+
agentId?: string | null;
|
|
192
|
+
agentName?: string | null;
|
|
193
|
+
parentName?: string | null;
|
|
194
|
+
parentSessionId?: string | null;
|
|
195
|
+
profile?: string | null;
|
|
196
|
+
modelAlias?: string | null;
|
|
197
|
+
surface?: string | null;
|
|
198
|
+
isolation?: string | null;
|
|
199
|
+
depth?: number | null;
|
|
200
|
+
/** `spawned`, `adopted`, `inherited` or `human`. */
|
|
201
|
+
originKind?: string | null;
|
|
202
|
+
configDir?: string | null;
|
|
203
|
+
spawnCwd?: string | null;
|
|
204
|
+
spawnedAt?: string | null;
|
|
205
|
+
briefChars?: number | null;
|
|
206
|
+
briefExcerpt?: string | null;
|
|
207
|
+
briefPath?: string | null;
|
|
208
|
+
launchArgs?: string | null;
|
|
209
|
+
}
|
|
210
|
+
/** A lifecycle event the launcher saw outside the transcript: `teleport`, `handoff`, `retired`, `exited`. */
|
|
211
|
+
interface ExternalEvent {
|
|
212
|
+
sessionId: string;
|
|
213
|
+
ts: string;
|
|
214
|
+
kind: string;
|
|
215
|
+
detail?: string | null;
|
|
216
|
+
/** Defaults to the session's resolved origin system. */
|
|
217
|
+
originSystem?: string | null;
|
|
218
|
+
}
|
|
219
|
+
interface OriginResolution {
|
|
220
|
+
/** Keyed by session id. */
|
|
221
|
+
origins: Record<string, ResolvedOrigin>;
|
|
222
|
+
externalEvents?: readonly ExternalEvent[];
|
|
223
|
+
}
|
|
224
|
+
/**
|
|
225
|
+
* Supplied by the caller, never by this package: session-graph is tier 2 and
|
|
226
|
+
* must not learn where a launcher keeps its records. Called once per pass with
|
|
227
|
+
* every session that has no origin row or a stale one.
|
|
228
|
+
*/
|
|
229
|
+
type OriginResolver = (sessionIds: readonly string[]) => PromiseLike<OriginResolution> | OriginResolution;
|
|
230
|
+
interface OriginEnrichment {
|
|
231
|
+
/** Session ids handed to the resolver. */
|
|
232
|
+
requested: number;
|
|
233
|
+
/** `session_origin` rows the resolver wrote. */
|
|
234
|
+
applied: number;
|
|
235
|
+
/** `session_external_event` rows the resolver wrote. */
|
|
236
|
+
events: number;
|
|
237
|
+
/** The resolver threw; origin rows stand as the last pass left them. */
|
|
238
|
+
failed: boolean;
|
|
239
|
+
/** Why it failed, for the caller to log. Absent unless `failed`. */
|
|
240
|
+
error?: string;
|
|
241
|
+
}
|
|
242
|
+
declare const NO_ORIGINS: OriginEnrichment;
|
|
243
|
+
declare function sessionsNeedingOrigin(graph: SessionGraph): string[];
|
|
244
|
+
/**
|
|
245
|
+
* Resolve origins for the sessions that lack a current one, then project every
|
|
246
|
+
* stored origin into `spawned` edges and `subagent` rows. Projecting from the
|
|
247
|
+
* table, not from this pass's answer, restores rows a parent transcript's purge
|
|
248
|
+
* removed. A resolver that throws costs this pass its origins and nothing else.
|
|
249
|
+
*/
|
|
250
|
+
declare function resolveOrigins(graph: SessionGraph, resolver: OriginResolver | undefined): Promise<OriginEnrichment>;
|
|
251
|
+
|
|
128
252
|
/**
|
|
129
253
|
* What a product's task store knows and a transcript cannot state: the task's
|
|
130
254
|
* present title, its initiative, and the status it holds right now rather than
|
|
@@ -135,6 +259,8 @@ interface ResolvedTask {
|
|
|
135
259
|
initiative?: string | null;
|
|
136
260
|
title?: string | null;
|
|
137
261
|
status?: string | null;
|
|
262
|
+
/** The store's size estimate for the task, in whatever unit it keeps. */
|
|
263
|
+
estimate?: number | null;
|
|
138
264
|
}
|
|
139
265
|
/** Resolutions keyed by bare task id (`AW-23`, not `task:AW-23`). */
|
|
140
266
|
type TaskResolution = ReadonlyMap<string, ResolvedTask | null | undefined>;
|
|
@@ -173,6 +299,10 @@ interface IndexOptions {
|
|
|
173
299
|
withContentHash?: boolean;
|
|
174
300
|
/** Fill task title, initiative and present status from the caller's store. Absent means transcripts alone. */
|
|
175
301
|
resolveTasks?: TaskResolver;
|
|
302
|
+
/** Record who launched each session. Runs once per `refreshCorpus` pass, never per transcript. Absent means no origin rows. */
|
|
303
|
+
resolveOrigins?: OriginResolver;
|
|
304
|
+
/** Fill PR state, merge, close and review rounds from the caller's forge. Runs once per `refreshCorpus` pass. Absent means transcripts alone. */
|
|
305
|
+
resolvePrs?: PrResolver;
|
|
176
306
|
}
|
|
177
307
|
interface RefreshOptions extends IndexOptions {
|
|
178
308
|
/** Roll up every session, not just the ones this pass touched. */
|
|
@@ -211,6 +341,8 @@ interface RefreshSummary {
|
|
|
211
341
|
turnsRolledUp: number;
|
|
212
342
|
reconciled: ReconcileCounts;
|
|
213
343
|
tasks: TaskEnrichment;
|
|
344
|
+
origins: OriginEnrichment;
|
|
345
|
+
prs: PrEnrichment;
|
|
214
346
|
markedMissing: number;
|
|
215
347
|
facetsBackfilled: number;
|
|
216
348
|
/** Transcripts whose audit facet is still stale after this pass. */
|
|
@@ -218,10 +350,10 @@ interface RefreshSummary {
|
|
|
218
350
|
}
|
|
219
351
|
/**
|
|
220
352
|
* One pass over a corpus: index every transcript, re-extract a bounded batch of
|
|
221
|
-
* stale audit facets, roll up the sessions that changed,
|
|
222
|
-
* cross-transcript observations,
|
|
223
|
-
* is gone. Idempotent: a second pass over unchanged
|
|
224
|
-
* changes nothing.
|
|
353
|
+
* stale audit facets, roll up the sessions that changed, resolve session
|
|
354
|
+
* origins, reconcile cross-transcript observations, resolve PR outcomes,
|
|
355
|
+
* enrich tasks, and mark rows whose source file is gone. Idempotent: a second pass over unchanged
|
|
356
|
+
* files changes nothing.
|
|
225
357
|
*
|
|
226
358
|
* The resolver is hoisted out of the per-transcript loop and run once over the
|
|
227
359
|
* whole task table, so a corpus of N transcripts costs one resolver call rather
|
|
@@ -230,6 +362,52 @@ interface RefreshSummary {
|
|
|
230
362
|
*/
|
|
231
363
|
declare function refreshCorpus(graph: SessionGraph, transcripts: readonly DiscoveredTranscript[], options?: RefreshOptions): Promise<RefreshSummary>;
|
|
232
364
|
|
|
365
|
+
/**
|
|
366
|
+
* One stretch of a session as some segmentation heuristic cut it. The rule
|
|
367
|
+
* lives above storage (session-analytics); this package only keeps its output.
|
|
368
|
+
* Offsets are byte offsets into the session's transcript.
|
|
369
|
+
*/
|
|
370
|
+
interface EpisodeRow {
|
|
371
|
+
episodeIndex: number;
|
|
372
|
+
heuristicVersion: number;
|
|
373
|
+
startedAt: string;
|
|
374
|
+
endedAt: string;
|
|
375
|
+
startOffset: number;
|
|
376
|
+
endOffset: number;
|
|
377
|
+
/** `session_start`, `brief`, `channel_followup`, `idle_gap`, `pr_merge`, `spawn_wave_complete`, `task_wrap`, `context_reset`, or a newer heuristic's term. */
|
|
378
|
+
openedBy: string;
|
|
379
|
+
assignmentOffset?: number | null;
|
|
380
|
+
firstDeliverableOffset?: number | null;
|
|
381
|
+
firstDeliverableSignal?: string | null;
|
|
382
|
+
firstStatusOffset?: number | null;
|
|
383
|
+
}
|
|
384
|
+
/**
|
|
385
|
+
* The only writer of `episode`. Replaces one heuristic's segmentation of one
|
|
386
|
+
* session atomically and leaves every other heuristic's rows alone, so two rule
|
|
387
|
+
* sets can coexist while one is compared against the other. Returns rows written.
|
|
388
|
+
*/
|
|
389
|
+
declare function replaceEpisodes(graph: SessionGraph, sessionId: string, heuristic: string, rows: readonly EpisodeRow[]): number;
|
|
390
|
+
|
|
391
|
+
/** USD per million tokens for one model prefix from one date. Shaped so session-analytics' `PRICE_TABLE` passes straight through. */
|
|
392
|
+
interface PriceInput {
|
|
393
|
+
/** Matched against `request.model` by longest prefix in `request_cost`. */
|
|
394
|
+
modelPrefix: string;
|
|
395
|
+
/** ISO date or timestamp; compared as text against `request.ts`. */
|
|
396
|
+
effectiveFrom: string;
|
|
397
|
+
input: number;
|
|
398
|
+
cacheRead: number;
|
|
399
|
+
cacheWrite5m: number;
|
|
400
|
+
cacheWrite1h: number;
|
|
401
|
+
output: number;
|
|
402
|
+
}
|
|
403
|
+
interface SyncPricesOptions {
|
|
404
|
+
/** Stamped on every row so a report can name the table it priced with. */
|
|
405
|
+
tableVersion: number;
|
|
406
|
+
source?: string;
|
|
407
|
+
}
|
|
408
|
+
/** Replace every `price` row with `rows` in one transaction, so `request_cost` never reads a half-written table. */
|
|
409
|
+
declare function syncPrices(graph: SessionGraph, rows: readonly PriceInput[], options: SyncPricesOptions): number;
|
|
410
|
+
|
|
233
411
|
interface NormalizedIndexResult {
|
|
234
412
|
status: "indexed" | "unchanged" | "missing" | "quarantined";
|
|
235
413
|
conversationRef: string;
|
|
@@ -274,4 +452,4 @@ declare function normalizedUsage(graph: SessionGraph, ref: string): NormalizedUs
|
|
|
274
452
|
/** Ambiguity is explicit; workspace session bodies remain in their original namespace. */
|
|
275
453
|
declare function resolveConversationAlias(db: Db, legacyRef: string): string | null;
|
|
276
454
|
|
|
277
|
-
export { AUDIT_DDL, AUDIT_FACET, AUDIT_MIGRATION_NAME, AUDIT_TABLES, type BackfillOptions, type BackfillSummary, type ConversationSummary, DEFAULT_FACET_LIMIT, DERIVED_TABLES, DOMAIN_DDL, type DeltaSource, FACET_TABLE, type IndexOptions, type IndexedSpan, KIT, MIGRATIONS, NO_ENRICHMENT, type NormalizedIndexResult, type NormalizedUsageSummary, type OpenSessionGraphOptions, type ReconcileCounts, type RefreshOptions, type RefreshSummary, type ResolvedTask, type SessionGraph, type TaskEnrichment, type TaskResolution, type TaskResolver, type TranscriptOutcome, allSessionIds, allTaskIds, applyAudit, applyDelta, backfillFacets, enrichTasks, indexCodexSource, indexTranscript, normalizedSessions, normalizedUsage, openSessionGraph, purgeTranscript, readIndexedText, reconcile, refreshCorpus, resetIndex, resolveConversationAlias, rollupSessions };
|
|
455
|
+
export { AUDIT_DDL, AUDIT_FACET, AUDIT_MIGRATION_NAME, AUDIT_TABLES, type BackfillOptions, type BackfillSummary, type ConversationSummary, DEFAULT_FACET_LIMIT, DERIVED_TABLES, DOMAIN_DDL, type DeltaSource, EPISODE_TABLE, type EpisodeRow, type ExternalEvent, FACET_TABLE, type IndexOptions, type IndexedSpan, KIT, MIGRATIONS, NO_ENRICHMENT, NO_ORIGINS, NO_PR_OUTCOMES, type NormalizedIndexResult, type NormalizedUsageSummary, ORIGIN_DDL, ORIGIN_MIGRATION_NAME, ORIGIN_TABLES, ORIGIN_VIEWS, type OpenSessionGraphOptions, type OriginEnrichment, type OriginResolution, type OriginResolver, type PrEnrichment, type PrKey, type PrResolution, type PrResolver, type PriceInput, type ReconcileCounts, type RefreshOptions, type RefreshSummary, type ResolvedOrigin, type ResolvedPr, type ResolvedTask, type SessionGraph, type SyncPricesOptions, type TaskEnrichment, type TaskResolution, type TaskResolver, type TranscriptOutcome, allSessionIds, allTaskIds, applyAudit, applyDelta, backfillFacets, enrichPrs, enrichTasks, indexCodexSource, indexTranscript, normalizedSessions, normalizedUsage, openSessionGraph, prsNeedingOutcome, purgeTranscript, readIndexedText, reconcile, refreshCorpus, replaceEpisodes, resetIndex, resolveConversationAlias, resolveOrigins, rollupSessions, sessionsNeedingOrigin, syncPrices };
|