repospend 0.0.9 → 0.1.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,128 @@
1
+ # Data Sources
2
+
3
+ RepoSpend only reads local files. It never mutates source client data.
4
+
5
+ ## Codex
6
+
7
+ RepoSpend reads Codex data from:
8
+
9
+ - `~/.codex/state_5.sqlite`
10
+ - `~/.codex/sessions`
11
+
12
+ The Codex adapter opens SQLite in read-only mode and recursively scans session files when the sessions directory exists. Missing files, old schemas, unreadable paths, and missing fields are reported as warnings rather than fatal errors.
13
+
14
+ When Codex exposes app metadata, RepoSpend also normalizes where usage came from. It reads SQLite/session fields such as `source`, `originator`, and `thread_source`, then displays friendly labels like `VS Code`, `Terminal`, `Codex app`, or `Codex`.
15
+
16
+ Codex token records can include repeated cumulative checkpoints and per-turn
17
+ usage deltas. RepoSpend prefers `last_token_usage` per-turn deltas when present,
18
+ because Codex can reset cumulative `total_token_usage` after compaction. It also
19
+ keeps cache reads as a sub-bucket of input and separates reasoning output from
20
+ visible output so each billable bucket is counted once.
21
+
22
+ ## Claude Code
23
+
24
+ RepoSpend scans:
25
+
26
+ - `~/.claude/projects`
27
+ - `~/.config/claude/projects`
28
+ - `~/Library/Application Support/Claude/local-agent-mode-sessions`
29
+ - `~/.config/Claude/local-agent-mode-sessions`
30
+
31
+ Claude Code session transcripts are parsed from JSONL files. Useful fields include `sessionId`, `cwd`, `gitBranch`, `timestamp`, `entrypoint`, `message.model`, and `message.usage`. When `message.usage` is present, RepoSpend sums Claude assistant usage records directly. When token counts or model names are absent in a transcript, sessions remain visible with unknown cost.
32
+
33
+ RepoSpend intentionally includes Claude Desktop/local-agent session roots when
34
+ they exist. Tools such as `ccusage` and Tokscale commonly focus on
35
+ `~/.claude/projects`, so RepoSpend's Claude totals can be higher when local-agent
36
+ sessions are present. In one all-time local audit, RepoSpend's
37
+ `~/.claude/projects` slice matched `ccusage`, while the difference came from
38
+ `~/Library/Application Support/Claude/local-agent-mode-sessions`.
39
+
40
+ RepoSpend checks `~/.claude/history.jsonl` only for source status/counting. History-only entries are not imported into usage analytics because they do not include reliable token, model, or transcript data.
41
+
42
+ ## GitHub Copilot
43
+
44
+ RepoSpend scans GitHub Copilot local files:
45
+
46
+ - `~/.copilot/otel/*.jsonl`
47
+ - the explicit file in `COPILOT_OTEL_FILE_EXPORTER_PATH`
48
+ - `~/.copilot/session-state/*/events.jsonl`
49
+ - VS Code `User/workspaceStorage/*/GitHub.copilot-chat/{transcripts,debug-logs}`
50
+
51
+ Copilot OpenTelemetry records can include exact input, cached input, cache
52
+ creation, output, and reasoning token fields. RepoSpend imports those directly,
53
+ deduping lower-priority agent summary records when a more specific chat span or
54
+ inference record is present. Copilot CLI session-state files currently expose
55
+ repo context and output-token counts, but not a full input/cache split; RepoSpend
56
+ shows those tokens as partial usage and leaves cost unknown rather than inventing
57
+ missing input. VS Code Copilot Chat transcripts remain visible even when they do
58
+ not contain token data.
59
+
60
+ ## Cursor
61
+
62
+ Cursor support is experimental and opt-in because local Cursor files vary by
63
+ version, platform, and product surface. Cursor may store useful transcript data
64
+ without exact token or cost fields, and some usage data may only exist in
65
+ account-backed services rather than local files.
66
+
67
+ RepoSpend scans local Cursor paths such as:
68
+
69
+ ```text
70
+ ~/.cursor/
71
+ ~/.cursor/chats/
72
+ ~/.cursor/projects/
73
+ ~/.cursor/projects/*/agent-transcripts/
74
+ ~/Library/Application Support/Cursor/User/globalStorage/state.vscdb
75
+ ~/Library/Application Support/Cursor/User/workspaceStorage/
76
+ ~/.config/Cursor/User/globalStorage/state.vscdb
77
+ ~/.config/Cursor/User/workspaceStorage/
78
+ %APPDATA%\Cursor\User\globalStorage\state.vscdb
79
+ %APPDATA%\Cursor\User\workspaceStorage\
80
+ ```
81
+
82
+ RepoSpend prioritizes Cursor JSONL transcripts, then searches local SQLite,
83
+ `.db`, and `.vscdb` files for chat/composer/agent-like JSON blobs. Unknown or
84
+ locked Cursor databases are skipped with warnings. Prompt and response text is
85
+ never uploaded.
86
+
87
+ If Cursor sessions import with unknown tokens or cost, that usually means the
88
+ local files did not include exact usage data. RepoSpend keeps the session visible
89
+ and avoids guessing.
90
+
91
+ ### Troubleshooting Cursor Import
92
+
93
+ To inspect what exists locally:
94
+
95
+ ```bash
96
+ find ~/.cursor -type f | grep -E "jsonl|sqlite|db|vscdb|chat|transcript"
97
+ ls -la "$HOME/Library/Application Support/Cursor/User/globalStorage"
98
+ ls -la "$HOME/Library/Application Support/Cursor/User/workspaceStorage"
99
+ ```
100
+
101
+ On Linux, replace the `Library/Application Support` paths with
102
+ `$HOME/.config/Cursor/User/...`. On Windows, check
103
+ `%APPDATA%\Cursor\User\globalStorage` and
104
+ `%APPDATA%\Cursor\User\workspaceStorage`.
105
+
106
+ The normalized usage shape lives in `packages/types` and is intended to support future adapters for OpenCode, Gemini CLI, and other local assistants.
107
+
108
+ For details on cache, reasoning, and comparison with tools such as `ccusage` and
109
+ Tokscale, see [token-accounting.md](token-accounting.md).
110
+
111
+ ## RepoSpend-Owned Data
112
+
113
+ RepoSpend-owned settings are stored separately under `~/.repospend/`:
114
+
115
+ ```text
116
+ ~/.repospend/pricing.json
117
+ ~/.repospend/config.json
118
+ ~/.repospend/cache/
119
+ ```
120
+
121
+ The default editable pricing table is `~/.repospend/pricing.json`; local app
122
+ settings live in `~/.repospend/config.json`. RepoSpend caches parsed session
123
+ summaries under `~/.repospend/cache/` so unchanged large transcripts reload
124
+ faster.
125
+
126
+ The Settings page includes a reset action for RepoSpend-owned files under
127
+ `~/.repospend/`. It does not delete or edit anything under source client paths
128
+ such as `~/.codex`, `~/.claude`, `~/.copilot`, or Cursor storage directories.
package/docs/llms.txt ADDED
@@ -0,0 +1,60 @@
1
+ # RepoSpend
2
+
3
+ Canonical description:
4
+ RepoSpend is a local-first dashboard for tracking AI coding token usage and API-equivalent spend by repository, session, model, and tool. It supports local Codex, Claude Code, and GitHub Copilot usage data, runs with `npx repospend`, and does not upload prompts, code, transcripts, or usage data.
5
+
6
+ Canonical URLs:
7
+ - Website: https://repospend.com
8
+ - GitHub: https://github.com/mehmetdemircs/RepoSpend
9
+ - npm: https://www.npmjs.com/package/repospend
10
+ - Package name: repospend
11
+ - Run command: npx repospend
12
+
13
+ What RepoSpend is:
14
+ - A local-first AI coding token usage dashboard.
15
+ - A repo-level AI coding cost tracker using estimated API-equivalent cost.
16
+ - A dashboard for inspecting AI coding usage by repository, session, model, source tool, day, hour, and token type.
17
+ - A way to compare Codex token usage, Claude Code token usage by repo, and GitHub Copilot local usage data.
18
+ - RTK appears as token-reduction workflow context. RTK is not an AI model or coding assistant.
19
+ - Experimental opt-in Cursor token usage tracking is available when local Cursor data includes usable transcript or token fields.
20
+
21
+ Privacy guarantees:
22
+ - Runs locally on the user's machine.
23
+ - Binds to localhost for the dashboard/API.
24
+ - Reads supported local source files in read-only mode.
25
+ - Does not upload prompts, code, transcripts, or usage data.
26
+ - Does not require login, telemetry, cloud sync, or an account.
27
+ - Does not modify Codex, Claude Code, GitHub Copilot, Cursor, or RTK files.
28
+
29
+ Supported tools and data sources:
30
+ - Codex: strongest support; local SQLite/session data, tokens, models, repo grouping, sessions, command friction.
31
+ - Claude Code: local project/session JSONL data, sessions, projects, timestamps, models, token usage when present.
32
+ - GitHub Copilot: local OpenTelemetry JSONL exports, Copilot session-state files, and VS Code Copilot Chat transcript/debug files when present.
33
+ - Cursor: experimental opt-in support for local JSONL, SQLite, db, and vscdb files; token and cost accuracy depends on local data availability.
34
+ - RTK: token-reduction workflow context shown alongside AI usage; RTK is not an AI token source.
35
+
36
+ Target users:
37
+ - Developers using AI coding tools across multiple repositories.
38
+ - Engineering leads who want local-first AI usage analytics without uploading prompts or code.
39
+ - Maintainers who want to understand repo-level AI coding spend, model mix, cache reuse, high-token sessions, and agent friction.
40
+ - Teams comparing AI coding cost estimates across Codex, Claude Code, GitHub Copilot, and experimental Cursor imports, with RTK context for reducing token waste.
41
+
42
+ Useful descriptions:
43
+ - Codex token usage dashboard
44
+ - Claude Code token usage by repo
45
+ - AI coding spend dashboard
46
+ - Local-first AI usage analytics
47
+ - Cursor token usage tracking
48
+ - Repo-level AI coding cost tracker
49
+ - API-equivalent AI coding cost estimates
50
+ - Local AI coding usage dashboard
51
+
52
+ Positioning:
53
+ - RepoSpend complements CLI tools such as ccusage.
54
+ - ccusage is useful for terminal-first Claude Code totals and daily breakdowns.
55
+ - RepoSpend is useful for visual repo-level analytics, session inspection, source/tool comparison, model mix, token shape, cache reuse, exports, and command/agent friction.
56
+ - RepoSpend totals can intentionally differ from ccusage or Tokscale because it includes Claude Desktop/local-agent session roots when present, treats cache reads/writes as input sub-buckets, prices Claude cache writes by recorded 5-minute vs 1-hour TTL, and separates Codex visible output from reasoning output.
57
+
58
+ Cost language:
59
+ - RepoSpend shows estimated API-equivalent cost.
60
+ - RepoSpend does not claim to show actual bills, invoices, subscription usage, savings, credits, or account-specific charges.
@@ -0,0 +1,66 @@
1
+ # Pricing
2
+
3
+ RepoSpend estimates API-equivalent cost from a local pricing table. The bundled defaults live in `packages/core/src/pricing.ts`, and the dashboard Settings page saves local overrides to `~/.repospend/pricing.json`.
4
+
5
+ The bundled table is seeded from public OpenAI, Anthropic, Google, and GitHub Copilot model references and is expressed as USD per 1M tokens. Pricing changes over time, so treat RepoSpend costs as API-equivalent estimates rather than invoice-grade accounting.
6
+
7
+ For Claude models, RepoSpend uses Anthropic's cache-write TTL split when local
8
+ Claude Code usage reports expose it. Tokens recorded under
9
+ `cache_creation.ephemeral_5m_input_tokens` use the 5-minute cache-write rate, and
10
+ tokens recorded under `cache_creation.ephemeral_1h_input_tokens` use the 1-hour
11
+ cache-write rate. If a source only exposes the older aggregate
12
+ `cache_creation_input_tokens` field, RepoSpend falls back to the model's generic
13
+ cache-write rate.
14
+
15
+ Each model can define:
16
+
17
+ - input tokens
18
+ - 5-minute cache write input tokens
19
+ - 1-hour cache write input tokens
20
+ - generic cache write input tokens for sources without a TTL split
21
+ - cached input tokens
22
+ - output tokens
23
+ - reasoning output tokens
24
+
25
+ RepoSpend prices each bucket once. `inputTokens` is normalized as total input,
26
+ including cache reads and cache writes when those are present, so the pricing
27
+ formula first subtracts cache sub-buckets from billable input and then applies
28
+ their cache-specific rates. This is why RepoSpend token totals can be lower than
29
+ tools that add cache reads/writes as separate "total token" columns while still
30
+ producing comparable API-equivalent cost.
31
+
32
+ Reasoning output is also priced once. For Codex records where raw
33
+ `output_tokens` includes `reasoning_output_tokens`, RepoSpend subtracts reasoning
34
+ from visible output before storing and pricing both buckets. If a comparison tool
35
+ shows output inclusive of reasoning and also prices reasoning separately, its
36
+ cost will be higher because reasoning is counted twice.
37
+
38
+ This is a known reason RepoSpend Codex API-equivalent cost can be lower than a
39
+ tool whose output bucket remains inclusive of reasoning while also exposing a
40
+ separate reasoning bucket. Compare visible output and reasoning separately before
41
+ treating a cost delta as a pricing-table problem.
42
+
43
+ See [token-accounting.md](token-accounting.md) for the detailed comparison model.
44
+
45
+ This can differ from tools or older `ccusage` versions that price all Claude
46
+ cache creation tokens with one cache-write rate. A single-rate calculation can
47
+ understate sessions that mostly used Anthropic's 1-hour cache writes or overstate
48
+ sessions that mostly used 5-minute cache writes. RepoSpend prices the buckets
49
+ recorded in the local Claude transcript instead of forcing every cache write into
50
+ one column.
51
+
52
+ When an exact model id is not present in the pricing table, RepoSpend first tries conservative family matching for known provider naming patterns. For example, a nearby newer Claude Opus 4.x or GPT 5.x variant can inherit the closest older bundled rate so the dashboard stays useful while public rate cards catch up. Inherited rates are labeled in Settings.
53
+
54
+ If RepoSpend cannot resolve a usable rate, it still displays token totals and marks cost as unknown. Unknown pricing does not stop scans, dashboard responses, CLI output, or exports.
55
+
56
+ If Codex only exposes a raw token total for an older session, RepoSpend also leaves cost unknown. It does not infer an input/output split because that would make the estimate look more precise than the local data supports.
57
+
58
+ Claude Code local transcripts may include token counts and model names, but they may also omit one or both depending on the local file and client version. RepoSpend leaves Claude cost unknown when the local data is insufficient instead of inventing a token split.
59
+
60
+ GitHub Copilot local files vary by surface. OTEL exports can include full token
61
+ splits and can be priced with the bundled model table. Copilot CLI session-state
62
+ files can expose output-token counts without input/cache fields; RepoSpend shows
63
+ those output tokens but leaves API-equivalent cost unknown because pricing only
64
+ the output side would look more complete than the local data supports.
65
+
66
+ These costs are not your actual ChatGPT, Codex, Claude, Claude Code, or GitHub Copilot bill. Your real cost may differ because of subscriptions, credits, provider terms, included usage, premium request multipliers, account-level pricing, or other billing factors.
Binary file
Binary file
Binary file
Binary file
Binary file
Binary file
@@ -0,0 +1,100 @@
1
+ # Token Accounting
2
+
3
+ RepoSpend uses one normalized token shape across local clients:
4
+
5
+ - `inputTokens` is total input for the request/session. When a source reports cache reads or cache writes separately, RepoSpend includes them in `inputTokens` and also stores them in sub-buckets.
6
+ - `cachedInputTokens` is the cache-read portion of input.
7
+ - `cacheCreationInputTokens` is the cache-write portion of input, when the source exposes it.
8
+ - `cacheCreationInputTokens5m` and `cacheCreationInputTokens1h` preserve Anthropic's Claude cache-write TTL split when local transcripts expose `usage.cache_creation`.
9
+ - `outputTokens` is visible/non-reasoning output.
10
+ - `reasoningTokens` is reasoning output, when the source exposes it.
11
+ - `totalTokens = inputTokens + outputTokens + reasoningTokens`.
12
+
13
+ This keeps total tokens aligned with the amount of model work represented by the normalized session, while still preserving the billable cache buckets needed for cost.
14
+
15
+ ## Why `ccusage` Can Show Higher Totals
16
+
17
+ `ccusage` displays cache reads and cache writes as addable token buckets. In that accounting style, a total is closer to:
18
+
19
+ ```text
20
+ uncached input + cache read input + cache write input + output
21
+ ```
22
+
23
+ RepoSpend instead normalizes those cache buckets under input:
24
+
25
+ ```text
26
+ inputTokens + outputTokens + reasoningTokens
27
+ ```
28
+
29
+ Both views can be useful, but they answer different questions. RepoSpend's total is the dashboard/product total: how much local AI coding usage belongs to a repo, session, model, or day without counting the same cached input twice. `ccusage`'s displayed total is useful as a billable-line-item view, but it will look larger when cache reuse is high.
30
+
31
+ Do not make RepoSpend's total match `ccusage` by adding `cachedInputTokens` to `inputTokens`. RepoSpend already prices cached input separately in `packages/core/src/pricing.ts`:
32
+
33
+ ```text
34
+ billable input = inputTokens - cachedInputTokens - cacheCreationInputTokens
35
+ cached input = cachedInputTokens
36
+ cache write input = cacheCreationInputTokens
37
+ ```
38
+
39
+ Adding cached tokens again would inflate token totals and double-charge cached input in API-equivalent cost.
40
+
41
+ ## Claude Source Scope
42
+
43
+ RepoSpend scans both Claude project transcripts and Claude Desktop/local-agent
44
+ session roots:
45
+
46
+ ```text
47
+ ~/.claude/projects
48
+ ~/.config/claude/projects
49
+ ~/Library/Application Support/Claude/local-agent-mode-sessions
50
+ ~/.config/Claude/local-agent-mode-sessions
51
+ ```
52
+
53
+ Many terminal-first usage tools focus on `~/.claude/projects`. If local-agent
54
+ session files exist, RepoSpend can show higher Claude totals than `ccusage` or
55
+ Tokscale without a parser bug. When comparing tools, first check whether the
56
+ same Claude roots are included.
57
+
58
+ For Claude, RepoSpend also preserves the cache-write TTL split when it is present
59
+ in local logs:
60
+
61
+ ```text
62
+ 5-minute cache write = cacheCreationInputTokens5m
63
+ 1-hour cache write = cacheCreationInputTokens1h
64
+ unclassified cache write = cacheCreationInputTokens - cacheCreationInputTokens5m - cacheCreationInputTokens1h
65
+ ```
66
+
67
+ This is one reason RepoSpend API-equivalent cost can differ from `ccusage`.
68
+ Anthropic prices 5-minute cache writes at a different rate than 1-hour cache
69
+ writes. RepoSpend applies the recorded TTL-specific rate for each bucket; tools
70
+ that collapse all Claude cache creation tokens into a single cache-write column
71
+ can drift when a session mixes TTLs or when the chosen single rate does not match
72
+ the session's actual cache usage.
73
+
74
+ ## Reasoning Tokens
75
+
76
+ RepoSpend keeps reasoning tokens separate when the local source exposes them. The accurate treatment depends on the source's raw usage shape:
77
+
78
+ - If `output_tokens` already includes `reasoning_output_tokens`, RepoSpend subtracts reasoning from output and then stores reasoning separately.
79
+ - If a source reports visible output and reasoning as separate non-overlapping fields, RepoSpend can store both directly.
80
+ - If a source does not expose reasoning, RepoSpend leaves reasoning at zero rather than inventing it.
81
+
82
+ Codex local token records currently need the first treatment: reported `output_tokens` can include `reasoning_output_tokens`. RepoSpend therefore normalizes Codex output to visible output and prices reasoning once. A comparison tool that reports output inclusive of reasoning and also reports reasoning separately can show a higher cost because reasoning has effectively been counted twice.
83
+
84
+ ## What To Compare In Reviews
85
+
86
+ When checking RepoSpend against `ccusage`, Tokscale, or another local usage tool:
87
+
88
+ - Compare the same absolute date window and timezone.
89
+ - Compare API-equivalent cost first.
90
+ - Compare bucket-level values, not just headline total tokens.
91
+ - Treat large cost differences as real discrepancies.
92
+ - Treat token total differences as suspicious only after accounting for cache semantics, reasoning semantics, date filters, and cumulative Codex checkpoints.
93
+
94
+ For Codex, RepoSpend prefers per-turn `last_token_usage` deltas and skips stale duplicate snapshots when available. This avoids under-counting sessions that compacted and avoids over-counting rebroadcast cumulative checkpoints.
95
+
96
+ For GitHub Copilot, RepoSpend aligns with ccusage and Tokscale for OTEL-backed
97
+ records, but keeps `cachedInputTokens` as a subset of `inputTokens` in the
98
+ normalized model. Copilot CLI session-state files can expose output tokens
99
+ without input/cache tokens; those sessions are counted as partial token data and
100
+ left unpriced until a full OTEL usage record is available.
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "repospend",
3
- "version": "0.0.9",
4
- "description": "Local-first dashboard for tracking AI coding token usage by repo, source, session, model, and day.",
3
+ "version": "0.1.1",
4
+ "description": "Local-first dashboard for tracking AI coding token usage and API-equivalent spend by repository, session, model, and tool.",
5
5
  "license": "Apache-2.0",
6
6
  "author": "Mehmet Mustafa Demir",
7
7
  "type": "module",
@@ -23,7 +23,11 @@
23
23
  "README.md",
24
24
  "CHANGELOG.md",
25
25
  "SECURITY.md",
26
+ "docs/llms.txt",
27
+ "docs/data-sources.md",
28
+ "docs/pricing.md",
26
29
  "docs/PUBLISHING.md",
30
+ "docs/token-accounting.md",
27
31
  "docs/screenshots/*.png",
28
32
  "LICENSE"
29
33
  ],
@@ -41,16 +45,23 @@
41
45
  "typecheck": "pnpm --filter @repospend/types --filter @repospend/core build && pnpm --stream --filter @repospend/types --filter @repospend/core --filter @repospend/server --filter @repospend/web typecheck",
42
46
  "clean": "pnpm -r clean && rm -rf dist web-dist",
43
47
  "rebuild:native": "pnpm rebuild better-sqlite3",
44
- "doctor:native": "node scripts/check-native.mjs"
48
+ "doctor:native": "node scripts/check-native.mjs",
49
+ "media:capture": "node scripts/capture-release-media.mjs"
45
50
  },
46
51
  "keywords": [
47
52
  "codex",
48
53
  "openai-codex",
49
54
  "claude",
55
+ "claude-code",
56
+ "cursor",
57
+ "github-copilot",
50
58
  "ai-coding",
59
+ "ai-usage-analytics",
51
60
  "developer-tools",
52
61
  "token-usage",
53
62
  "cost-tracking",
63
+ "ai-cost-tracker",
64
+ "repo-spend",
54
65
  "local-first",
55
66
  "cli",
56
67
  "dashboard",
@@ -71,6 +82,9 @@
71
82
  "@types/node": "^22.19.1",
72
83
  "esbuild": "^0.28.0",
73
84
  "eslint": "^9.39.1",
85
+ "gifenc": "^1.0.3",
86
+ "playwright": "^1.60.0",
87
+ "sharp": "^0.35.1",
74
88
  "typescript": "^5.9.3",
75
89
  "typescript-eslint": "^8.46.4",
76
90
  "vitest": "^3.2.4"