transcripto 0.1.0__tar.gz → 0.1.2__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,340 @@
1
+ Metadata-Version: 2.4
2
+ Name: transcripto
3
+ Version: 0.1.2
4
+ Summary: Search everything your coding agents ever did, grade your own prompts, and price your decisions. Local, stdlib-only, your data never leaves the machine.
5
+ Author: Oscar Morke
6
+ License: MIT
7
+ Project-URL: Homepage, https://github.com/Morkeeth/transcripto
8
+ Project-URL: Source, https://github.com/Morkeeth/transcripto
9
+ Keywords: claude-code,coding-agents,transcripts,local-first,analytics
10
+ Classifier: Environment :: Console
11
+ Classifier: Programming Language :: Python :: 3
12
+ Classifier: License :: OSI Approved :: MIT License
13
+ Classifier: Topic :: Utilities
14
+ Requires-Python: >=3.9
15
+ Description-Content-Type: text/markdown
16
+ License-File: LICENSE
17
+ Dynamic: license-file
18
+
19
+ # transcripto
20
+
21
+ **your coding agents keep a transcript of every session. it is the most valuable
22
+ dataset you own, and you cannot scroll back far enough to read it. transcripto
23
+ indexes it, keeps only the turns you actually typed, and grades them.**
24
+
25
+ one command, no account, no signup, no cloud. it reads files that are already on
26
+ your disk and never opens a socket.
27
+
28
+ ```
29
+ uvx transcripto coach
30
+ ```
31
+
32
+ > **check which build you got.** `trace`, `--harness cursor`, and the "run `index`
33
+ > first" errors below all arrive in **0.1.2**. On anything older, `trace` is not a
34
+ > command, cursor is rejected, and the index-gated commands fail with a raw sqlite
35
+ > traceback instead of an instruction. One line settles which one you are holding:
36
+ >
37
+ > ```
38
+ > uvx transcripto --version # 0.1.2 or newer = this README is accurate
39
+ > ```
40
+ >
41
+ > *Read 2026-08-31: PyPI was serving 0.1.1 while this README described 0.1.2, so
42
+ > `uvx transcripto` gave the older build. If `--version` is not even a recognised
43
+ > flag, you have 0.1.1 — it was added in 0.1.2 precisely because there was no way
44
+ > to tell.* The repo always matches this README and needs nothing installed:
45
+ >
46
+ > ```
47
+ > git clone https://github.com/Morkeeth/transcripto && cd transcripto
48
+ > python3 transcripto.py coach
49
+ > ```
50
+
51
+ ## three harnesses, one instrument
52
+
53
+ ```
54
+ transcripto coach # Claude Code, ~/.claude/projects
55
+ transcripto coach --harness codex # Codex, ~/.codex
56
+ transcripto coach --harness cursor # Cursor, ~/.cursor/projects/*/agent-transcripts
57
+ ```
58
+
59
+ **Authorship is not the same gate in all three, and the tool says so rather than pooling them.**
60
+ Claude Code stamps `promptSource: typed`, which is the measured-reliable signal: about 95% of raw
61
+ `type: user` records are not the operator at all. Cursor has no such field. Its one honest
62
+ equivalent is the `<user_query>` wrapper it puts around a submitted prompt, which injected and
63
+ tool-result records do not carry. That is a weaker signal and it is labelled weaker.
64
+
65
+ ## `trace` — what actually happened after you asked
66
+
67
+ `ask` shows what you typed. `find` shows what a file went through. Neither answers the
68
+ question that matters after the fact: **you asked for X, did anything durable happen?**
69
+
70
+ ```
71
+ transcripto trace "the gate"
72
+ ```
73
+
74
+ It walks each of your matching prompts forward inside its own session and lists the writes
75
+ and edits that followed, stopping at your next prompt so one turn cannot claim the next
76
+ turn's work. Green dot = something durable landed. Red = nothing was touched.
77
+
78
+ **Honest limit:** a write following a prompt in the same session is CO-OCCURRENCE, not proof
79
+ the write was caused by that prompt or that it was correct. Same proxy `coach` uses, labelled
80
+ the same way.
81
+
82
+ ## what you get back
83
+
84
+ this is a real run on one machine, pasted unedited, 2026-08-28:
85
+
86
+ ```
87
+
88
+ YOUR PROMPT HABITS, GRADED (offline, your machine only)
89
+
90
+ harness: claude
91
+ corpus : 2720 transcript(s), 381,804 records
92
+ kept : 3678 prompts you actually typed (0.96% of records)
93
+ episodes: 1946 ranked, 910 survived (47%)
94
+ tiers : commit 373 | write/edit 537 | reverted 2 | nothing durable 1034
95
+
96
+ SURVIVAL IS A PROXY: survival = a durable Write/Edit or an un-reverted git commit in-episode. A PROXY, not proof the work was correct or shipped.
97
+
98
+ SURVIVES MOST do more of these:
99
+ 64% (167/260) detailed (>40 words)
100
+ 63% (66/104) states-a-check-or-done-condition
101
+ 58% (414/708) intent:CHANGE
102
+ 58% (21/36) no-object (pronoun/vague)
103
+ 57% (99/175) cites-a-file-or-path
104
+
105
+ SURVIVES LEAST these tend to loop:
106
+ 33% (4/12) intent:REVERT
107
+ 39% (411/1043) intent:none
108
+ 40% (4/10) intent:TEST
109
+ 42% (298/715) terse (<8 words)
110
+ 45% (77/173) intent:DESCRIBE
111
+
112
+ + your best landed prompt, with its witness:
113
+ "mTERMINAL 8 — Mountain of Helicon · ~/CODE/mountain-of-helicon Read ~/CODE/mountain-of-helicon. Two…"
114
+ COMMIT-WITNESSED: git commit · corrections: 0
115
+
116
+ - your worst looped prompt, with its witness:
117
+ "ok, and lets see they might solve it in the future so i can go back to my beloeved routine :) Befor…"
118
+ NO-DURABLE-RECORD: read-only Bash only, no file change · corrections: 15 · assistant turns: 221
119
+ ```
120
+
121
+ those are my numbers on that date, and they move every session i run, so treat
122
+ them as a snapshot rather than a constant. yours will be different, which is the
123
+ whole point. the last two lines are the ones that sting: it hands you back your
124
+ own best and worst prompt, verbatim, with the receipt for why it scored each one.
125
+
126
+ on that machine, on that date, prompts that wrote down what done looks like
127
+ survived **63% of the time (66 of 104)**. prompts with no stated intent survived
128
+ **39% (411 of 1043)**. i had spent a year blaming the model.
129
+
130
+ one day later, 2026-08-29, the same command on the same machine read 63% (67 of
131
+ 107) and 40% (424 of 1072) over 2,874 transcripts. the percentages held and the
132
+ denominators moved, which is what a snapshot is supposed to do.
133
+
134
+ ## the proxy caveat, which travels with every number
135
+
136
+ an episode "survived" if a Write or Edit landed, or a git commit ran and nothing
137
+ reverted it inside the same transcript.
138
+
139
+ that is a durable **keystroke**, not a durable **outcome**. a commit is not proof
140
+ the code was right. a revert in a later session is invisible to it. a prompt
141
+ whose payoff was a decision rather than an edit reads as dead. it is a coaching
142
+ signal, not a verdict. if it ever prints something that flatters you, distrust it.
143
+
144
+ the caveat is printed in the output itself, every run, on purpose.
145
+
146
+ ## why your own gate matters here
147
+
148
+ at fleet scale roughly 95% of the `type: user` records in a transcript are not
149
+ you. they are tool results, injected skill bodies, sub-agent prompts, and
150
+ messages from other terminals, all wearing your role. transcripto gates on
151
+ `promptSource` (typed/queued, no meta, no sidechain) so it grades what you typed.
152
+
153
+ you can watch the gate do work: in the run above, 3678 of 381,804 records
154
+ survived it. that is 0.96%.
155
+
156
+ the same gate is what makes `cost` produce a number a spend tracker cannot:
157
+
158
+ ```
159
+ cost per human decision last 30 days 2026-07-28 → 2026-08-27
160
+
161
+ API-equivalent spend $8,892.49
162
+ your decisions 2934 turns you actually typed (promptSource typed/queued)
163
+ ────────────────────────────────────────────────────
164
+ cost per human decision $3.03
165
+
166
+ 46.4k agent messages · 16 per decision · 11.5B tokens · 149 sessions
167
+ 57.7k raw `type: user` records in the same window. dividing by those instead
168
+ would read $0.15, 19.7x too cheap.
169
+ ```
170
+
171
+ it prints both, so the gate's effect is something you can check rather than
172
+ something i am asserting. these are API-equivalent dollars at list rates,
173
+ because a transcript has no cost field, only token counts. on a subscription you
174
+ did not pay this.
175
+
176
+ ## honest limits
177
+
178
+ read these before you quote a number at anyone.
179
+
180
+ - **survival is a proxy**, described above. durable keystroke, not durable outcome.
181
+ - **one operator's corpus.** every figure in this README comes from one machine.
182
+ it is an existence proof that the measurement runs, not a finding about how
183
+ people prompt. run it on yours and you get yours.
184
+ - **three harnesses today: Claude Code, Codex, Cursor.** nothing else is supported.
185
+ aider and the rest are not read. and the three are not equal: Claude Code has a
186
+ measured-reliable authorship field, Cursor has only the `<user_query>` wrapper,
187
+ which is weaker and is labelled weaker wherever it is used.
188
+ - **the habit labels are heuristics.** "states-a-check-or-done-condition" is a
189
+ pattern match over your text, not comprehension. it will misfile some prompts.
190
+ - **correlation, not instruction.** detailed prompts surviving more often does not
191
+ prove that padding a prompt causes survival.
192
+
193
+ ## privacy
194
+
195
+ it runs locally and never touches the network. there is no socket, no urllib, no
196
+ requests, no subprocess, no telemetry, no analytics, and no account. that is a
197
+ claim, so here is the grep that settles it against the single file it ships as:
198
+
199
+ ```
200
+ $ grep -nE '^[[:space:]]*(import|from) ' transcripto.py
201
+ 8:import sys, os, json, glob, re, sqlite3, argparse
202
+ 9:from datetime import datetime, timezone
203
+ 260: import time
204
+ 1051: import datetime
205
+ 1061: from the separator), so the result is checked on disk and dropped if it is
206
+ ```
207
+
208
+ five lines, four of which are imports and all four are stdlib. `time` and
209
+ `datetime` sit inside functions, which is why the pattern allows for indentation —
210
+ anchor it at `^import` and you would miss two, so do not take my word for the
211
+ anchor either. line 1061 is the pattern catching a docstring that happens to begin
212
+ with the word `from`; it is prose, not an import, and it is left in rather than
213
+ tuned out, because a grep you tuned until it agreed with you proves nothing.
214
+
215
+ what the list does NOT contain is the actual claim: no `socket`, no `urllib`, no
216
+ `requests`, no `http.client`, no `subprocess`. that one is checkable too, and the
217
+ right answer is no output at all:
218
+
219
+ ```
220
+ $ grep -nE '\b(socket|urllib|requests|http\.client|subprocess)\b' transcripto.py
221
+ $
222
+ ```
223
+
224
+ your transcripts stay in `~/.claude`, `~/.codex` and `~/.cursor`. the index it
225
+ builds stays in `~/.trace`.
226
+
227
+ ## the rest of it
228
+
229
+ `coach` and `cost` read your transcript files directly and need nothing set up.
230
+ **the other six read a local index, so run this once first:**
231
+
232
+ ```
233
+ transcripto index # a few minutes on a large corpus, incremental after that
234
+ ```
235
+
236
+ on a 2,874-file corpus that was 164 seconds, measured 2026-08-29. if you skip it,
237
+ the six say so and exit 2.
238
+
239
+ ```
240
+ transcripto index build / refresh (incremental)
241
+ transcripto watch live, new sessions get picked up as your agents work
242
+ transcripto ask YOUR OWN messages about a topic, newest first + a rollup
243
+ transcripto search full-text across everything (you + agents + tool logs)
244
+ transcripto find every session that wrote / edited / read a file
245
+ transcripto trace what durably happened after each prompt you typed (0.1.2+)
246
+ transcripto sessions recent sessions + the first prompt YOU typed in each
247
+ transcripto stats what you actually work on
248
+ transcripto cost what ONE of your decisions costs
249
+ transcripto coach which of YOUR prompt habits survive (a proxy)
250
+ ```
251
+
252
+ `ask` is the one that kills "wait, did i lose something?". it answers "what was i
253
+ thinking about X across ALL my sessions", in your own words only.
254
+
255
+ ```
256
+ $ transcripto find USER-JOURNEY.md # run 2026-08-29
257
+ USER-JOURNEY.md 4 touches across sessions (3 were writes/edits)
258
+
259
+ 2026-08-20 WROTE ~/CODE/mountain-of-helicon-main/USER-JOURNEY.md abd9e871
260
+ 2026-08-21 WROTE ~/…/Obsidian LIFE/00 Dashboard/suite-user-journey.md 0f845ede
261
+ 2026-08-27 read ~/CODE/hack-fleet-ata/docs/USER-JOURNEY.md cddfde29
262
+ 2026-08-27 WROTE ~/CODE/hack-fleet-ata/docs/USER-JOURNEY.md cddfde29
263
+ ```
264
+
265
+ the file you lost, found across every session you ever ran, with the session id
266
+ that touched it. `find` needs `transcripto index` first.
267
+
268
+ ## Codex
269
+
270
+ ```
271
+ transcripto coach --harness codex
272
+ ```
273
+
274
+ reads `~/.codex` (sessions + archived_sessions), normalises it into the same rows,
275
+ and applies the identical survival proxy. it also ingests `history.jsonl` purely
276
+ as a control on the gate: it reports how many of its input lines also show up as
277
+ typed rollout turns, so you can see the gate agreeing with a second source.
278
+
279
+ ## install
280
+
281
+ ```
282
+ uvx transcripto coach
283
+ ```
284
+
285
+ no install, nothing to set up. or put it on your PATH:
286
+
287
+ ```
288
+ pipx install transcripto
289
+ ```
290
+
291
+ or run the single file with no packaging at all:
292
+
293
+ ```
294
+ git clone https://github.com/Morkeeth/transcripto
295
+ cd transcripto
296
+ python3 transcripto.py coach
297
+ ```
298
+
299
+ no dependencies, stdlib only, one file. the packaging adds nothing at runtime, it
300
+ just gives the file a name on your PATH.
301
+
302
+ ## tests
303
+
304
+ ```
305
+ ./test_coach.sh 15 assertions
306
+ ./test_codex.sh 14 assertions
307
+ ./test_cost.sh 12 assertions
308
+ ./test_small_n.sh 7 assertions
309
+ ./test_label_bands.sh 13 assertions
310
+ ```
311
+
312
+ 61 assertions, all green, re-run 2026-08-31.
313
+
314
+ offline, no keys, on fixtures that inherit the real transcript shape including
315
+ all four ways a non-human record disguises itself as `type: user`.
316
+
317
+ the load-bearing one in `test_coach.sh` is `REVERTED IS NOT SURVIVED`: a commit
318
+ that got `reset --hard` in the same session left no durable record. flip that one
319
+ line and the suite goes red, which is the point. a generous proxy is a broken one.
320
+
321
+ `test_label_bands.sh` is the other one, and it exists because 0.1.1 shipped the
322
+ defect it pins. `SURVIVES MOST` took the top five habits and `SURVIVES LEAST` took
323
+ the bottom five, which overlap whenever you have fewer than ten rankable habits —
324
+ so a new user, who necessarily has few, read the same habit at the same percentage
325
+ under both "do more of these" and "these tend to loop". on a 3-habit corpus 0.1.1
326
+ reprinted all three, all at 65% (22/34). the suite is red on the published 0.1.1
327
+ file and green on this one, and its `wide` band asserts the fix leaves a large
328
+ corpus byte-identical.
329
+
330
+ ## why though
331
+
332
+ your agent history is proof. every "yeah it's done" has a real trace sitting
333
+ behind it. transcripto is the index that makes it checkable.
334
+
335
+ it is the fuel layer. on top of it you check what your agents *claim* against what
336
+ the trace *shows*, which is [mountain of helicon](https://github.com/Morkeeth/mountain-of-helicon).
337
+ the pitch was never "search your history". it is *prove your agent did what it
338
+ said, from your own local traces.*
339
+
340
+ local, MIT, no telemetry. star it if it finds you something you'd lost ™
@@ -0,0 +1,322 @@
1
+ # transcripto
2
+
3
+ **your coding agents keep a transcript of every session. it is the most valuable
4
+ dataset you own, and you cannot scroll back far enough to read it. transcripto
5
+ indexes it, keeps only the turns you actually typed, and grades them.**
6
+
7
+ one command, no account, no signup, no cloud. it reads files that are already on
8
+ your disk and never opens a socket.
9
+
10
+ ```
11
+ uvx transcripto coach
12
+ ```
13
+
14
+ > **check which build you got.** `trace`, `--harness cursor`, and the "run `index`
15
+ > first" errors below all arrive in **0.1.2**. On anything older, `trace` is not a
16
+ > command, cursor is rejected, and the index-gated commands fail with a raw sqlite
17
+ > traceback instead of an instruction. One line settles which one you are holding:
18
+ >
19
+ > ```
20
+ > uvx transcripto --version # 0.1.2 or newer = this README is accurate
21
+ > ```
22
+ >
23
+ > *Read 2026-08-31: PyPI was serving 0.1.1 while this README described 0.1.2, so
24
+ > `uvx transcripto` gave the older build. If `--version` is not even a recognised
25
+ > flag, you have 0.1.1 — it was added in 0.1.2 precisely because there was no way
26
+ > to tell.* The repo always matches this README and needs nothing installed:
27
+ >
28
+ > ```
29
+ > git clone https://github.com/Morkeeth/transcripto && cd transcripto
30
+ > python3 transcripto.py coach
31
+ > ```
32
+
33
+ ## three harnesses, one instrument
34
+
35
+ ```
36
+ transcripto coach # Claude Code, ~/.claude/projects
37
+ transcripto coach --harness codex # Codex, ~/.codex
38
+ transcripto coach --harness cursor # Cursor, ~/.cursor/projects/*/agent-transcripts
39
+ ```
40
+
41
+ **Authorship is not the same gate in all three, and the tool says so rather than pooling them.**
42
+ Claude Code stamps `promptSource: typed`, which is the measured-reliable signal: about 95% of raw
43
+ `type: user` records are not the operator at all. Cursor has no such field. Its one honest
44
+ equivalent is the `<user_query>` wrapper it puts around a submitted prompt, which injected and
45
+ tool-result records do not carry. That is a weaker signal and it is labelled weaker.
46
+
47
+ ## `trace` — what actually happened after you asked
48
+
49
+ `ask` shows what you typed. `find` shows what a file went through. Neither answers the
50
+ question that matters after the fact: **you asked for X, did anything durable happen?**
51
+
52
+ ```
53
+ transcripto trace "the gate"
54
+ ```
55
+
56
+ It walks each of your matching prompts forward inside its own session and lists the writes
57
+ and edits that followed, stopping at your next prompt so one turn cannot claim the next
58
+ turn's work. Green dot = something durable landed. Red = nothing was touched.
59
+
60
+ **Honest limit:** a write following a prompt in the same session is CO-OCCURRENCE, not proof
61
+ the write was caused by that prompt or that it was correct. Same proxy `coach` uses, labelled
62
+ the same way.
63
+
64
+ ## what you get back
65
+
66
+ this is a real run on one machine, pasted unedited, 2026-08-28:
67
+
68
+ ```
69
+
70
+ YOUR PROMPT HABITS, GRADED (offline, your machine only)
71
+
72
+ harness: claude
73
+ corpus : 2720 transcript(s), 381,804 records
74
+ kept : 3678 prompts you actually typed (0.96% of records)
75
+ episodes: 1946 ranked, 910 survived (47%)
76
+ tiers : commit 373 | write/edit 537 | reverted 2 | nothing durable 1034
77
+
78
+ SURVIVAL IS A PROXY: survival = a durable Write/Edit or an un-reverted git commit in-episode. A PROXY, not proof the work was correct or shipped.
79
+
80
+ SURVIVES MOST do more of these:
81
+ 64% (167/260) detailed (>40 words)
82
+ 63% (66/104) states-a-check-or-done-condition
83
+ 58% (414/708) intent:CHANGE
84
+ 58% (21/36) no-object (pronoun/vague)
85
+ 57% (99/175) cites-a-file-or-path
86
+
87
+ SURVIVES LEAST these tend to loop:
88
+ 33% (4/12) intent:REVERT
89
+ 39% (411/1043) intent:none
90
+ 40% (4/10) intent:TEST
91
+ 42% (298/715) terse (<8 words)
92
+ 45% (77/173) intent:DESCRIBE
93
+
94
+ + your best landed prompt, with its witness:
95
+ "mTERMINAL 8 — Mountain of Helicon · ~/CODE/mountain-of-helicon Read ~/CODE/mountain-of-helicon. Two…"
96
+ COMMIT-WITNESSED: git commit · corrections: 0
97
+
98
+ - your worst looped prompt, with its witness:
99
+ "ok, and lets see they might solve it in the future so i can go back to my beloeved routine :) Befor…"
100
+ NO-DURABLE-RECORD: read-only Bash only, no file change · corrections: 15 · assistant turns: 221
101
+ ```
102
+
103
+ those are my numbers on that date, and they move every session i run, so treat
104
+ them as a snapshot rather than a constant. yours will be different, which is the
105
+ whole point. the last two lines are the ones that sting: it hands you back your
106
+ own best and worst prompt, verbatim, with the receipt for why it scored each one.
107
+
108
+ on that machine, on that date, prompts that wrote down what done looks like
109
+ survived **63% of the time (66 of 104)**. prompts with no stated intent survived
110
+ **39% (411 of 1043)**. i had spent a year blaming the model.
111
+
112
+ one day later, 2026-08-29, the same command on the same machine read 63% (67 of
113
+ 107) and 40% (424 of 1072) over 2,874 transcripts. the percentages held and the
114
+ denominators moved, which is what a snapshot is supposed to do.
115
+
116
+ ## the proxy caveat, which travels with every number
117
+
118
+ an episode "survived" if a Write or Edit landed, or a git commit ran and nothing
119
+ reverted it inside the same transcript.
120
+
121
+ that is a durable **keystroke**, not a durable **outcome**. a commit is not proof
122
+ the code was right. a revert in a later session is invisible to it. a prompt
123
+ whose payoff was a decision rather than an edit reads as dead. it is a coaching
124
+ signal, not a verdict. if it ever prints something that flatters you, distrust it.
125
+
126
+ the caveat is printed in the output itself, every run, on purpose.
127
+
128
+ ## why your own gate matters here
129
+
130
+ at fleet scale roughly 95% of the `type: user` records in a transcript are not
131
+ you. they are tool results, injected skill bodies, sub-agent prompts, and
132
+ messages from other terminals, all wearing your role. transcripto gates on
133
+ `promptSource` (typed/queued, no meta, no sidechain) so it grades what you typed.
134
+
135
+ you can watch the gate do work: in the run above, 3678 of 381,804 records
136
+ survived it. that is 0.96%.
137
+
138
+ the same gate is what makes `cost` produce a number a spend tracker cannot:
139
+
140
+ ```
141
+ cost per human decision last 30 days 2026-07-28 → 2026-08-27
142
+
143
+ API-equivalent spend $8,892.49
144
+ your decisions 2934 turns you actually typed (promptSource typed/queued)
145
+ ────────────────────────────────────────────────────
146
+ cost per human decision $3.03
147
+
148
+ 46.4k agent messages · 16 per decision · 11.5B tokens · 149 sessions
149
+ 57.7k raw `type: user` records in the same window. dividing by those instead
150
+ would read $0.15, 19.7x too cheap.
151
+ ```
152
+
153
+ it prints both, so the gate's effect is something you can check rather than
154
+ something i am asserting. these are API-equivalent dollars at list rates,
155
+ because a transcript has no cost field, only token counts. on a subscription you
156
+ did not pay this.
157
+
158
+ ## honest limits
159
+
160
+ read these before you quote a number at anyone.
161
+
162
+ - **survival is a proxy**, described above. durable keystroke, not durable outcome.
163
+ - **one operator's corpus.** every figure in this README comes from one machine.
164
+ it is an existence proof that the measurement runs, not a finding about how
165
+ people prompt. run it on yours and you get yours.
166
+ - **three harnesses today: Claude Code, Codex, Cursor.** nothing else is supported.
167
+ aider and the rest are not read. and the three are not equal: Claude Code has a
168
+ measured-reliable authorship field, Cursor has only the `<user_query>` wrapper,
169
+ which is weaker and is labelled weaker wherever it is used.
170
+ - **the habit labels are heuristics.** "states-a-check-or-done-condition" is a
171
+ pattern match over your text, not comprehension. it will misfile some prompts.
172
+ - **correlation, not instruction.** detailed prompts surviving more often does not
173
+ prove that padding a prompt causes survival.
174
+
175
+ ## privacy
176
+
177
+ it runs locally and never touches the network. there is no socket, no urllib, no
178
+ requests, no subprocess, no telemetry, no analytics, and no account. that is a
179
+ claim, so here is the grep that settles it against the single file it ships as:
180
+
181
+ ```
182
+ $ grep -nE '^[[:space:]]*(import|from) ' transcripto.py
183
+ 8:import sys, os, json, glob, re, sqlite3, argparse
184
+ 9:from datetime import datetime, timezone
185
+ 260: import time
186
+ 1051: import datetime
187
+ 1061: from the separator), so the result is checked on disk and dropped if it is
188
+ ```
189
+
190
+ five lines, four of which are imports and all four are stdlib. `time` and
191
+ `datetime` sit inside functions, which is why the pattern allows for indentation —
192
+ anchor it at `^import` and you would miss two, so do not take my word for the
193
+ anchor either. line 1061 is the pattern catching a docstring that happens to begin
194
+ with the word `from`; it is prose, not an import, and it is left in rather than
195
+ tuned out, because a grep you tuned until it agreed with you proves nothing.
196
+
197
+ what the list does NOT contain is the actual claim: no `socket`, no `urllib`, no
198
+ `requests`, no `http.client`, no `subprocess`. that one is checkable too, and the
199
+ right answer is no output at all:
200
+
201
+ ```
202
+ $ grep -nE '\b(socket|urllib|requests|http\.client|subprocess)\b' transcripto.py
203
+ $
204
+ ```
205
+
206
+ your transcripts stay in `~/.claude`, `~/.codex` and `~/.cursor`. the index it
207
+ builds stays in `~/.trace`.
208
+
209
+ ## the rest of it
210
+
211
+ `coach` and `cost` read your transcript files directly and need nothing set up.
212
+ **the other six read a local index, so run this once first:**
213
+
214
+ ```
215
+ transcripto index # a few minutes on a large corpus, incremental after that
216
+ ```
217
+
218
+ on a 2,874-file corpus that was 164 seconds, measured 2026-08-29. if you skip it,
219
+ the six say so and exit 2.
220
+
221
+ ```
222
+ transcripto index build / refresh (incremental)
223
+ transcripto watch live, new sessions get picked up as your agents work
224
+ transcripto ask YOUR OWN messages about a topic, newest first + a rollup
225
+ transcripto search full-text across everything (you + agents + tool logs)
226
+ transcripto find every session that wrote / edited / read a file
227
+ transcripto trace what durably happened after each prompt you typed (0.1.2+)
228
+ transcripto sessions recent sessions + the first prompt YOU typed in each
229
+ transcripto stats what you actually work on
230
+ transcripto cost what ONE of your decisions costs
231
+ transcripto coach which of YOUR prompt habits survive (a proxy)
232
+ ```
233
+
234
+ `ask` is the one that kills "wait, did i lose something?". it answers "what was i
235
+ thinking about X across ALL my sessions", in your own words only.
236
+
237
+ ```
238
+ $ transcripto find USER-JOURNEY.md # run 2026-08-29
239
+ USER-JOURNEY.md 4 touches across sessions (3 were writes/edits)
240
+
241
+ 2026-08-20 WROTE ~/CODE/mountain-of-helicon-main/USER-JOURNEY.md abd9e871
242
+ 2026-08-21 WROTE ~/…/Obsidian LIFE/00 Dashboard/suite-user-journey.md 0f845ede
243
+ 2026-08-27 read ~/CODE/hack-fleet-ata/docs/USER-JOURNEY.md cddfde29
244
+ 2026-08-27 WROTE ~/CODE/hack-fleet-ata/docs/USER-JOURNEY.md cddfde29
245
+ ```
246
+
247
+ the file you lost, found across every session you ever ran, with the session id
248
+ that touched it. `find` needs `transcripto index` first.
249
+
250
+ ## Codex
251
+
252
+ ```
253
+ transcripto coach --harness codex
254
+ ```
255
+
256
+ reads `~/.codex` (sessions + archived_sessions), normalises it into the same rows,
257
+ and applies the identical survival proxy. it also ingests `history.jsonl` purely
258
+ as a control on the gate: it reports how many of its input lines also show up as
259
+ typed rollout turns, so you can see the gate agreeing with a second source.
260
+
261
+ ## install
262
+
263
+ ```
264
+ uvx transcripto coach
265
+ ```
266
+
267
+ no install, nothing to set up. or put it on your PATH:
268
+
269
+ ```
270
+ pipx install transcripto
271
+ ```
272
+
273
+ or run the single file with no packaging at all:
274
+
275
+ ```
276
+ git clone https://github.com/Morkeeth/transcripto
277
+ cd transcripto
278
+ python3 transcripto.py coach
279
+ ```
280
+
281
+ no dependencies, stdlib only, one file. the packaging adds nothing at runtime, it
282
+ just gives the file a name on your PATH.
283
+
284
+ ## tests
285
+
286
+ ```
287
+ ./test_coach.sh 15 assertions
288
+ ./test_codex.sh 14 assertions
289
+ ./test_cost.sh 12 assertions
290
+ ./test_small_n.sh 7 assertions
291
+ ./test_label_bands.sh 13 assertions
292
+ ```
293
+
294
+ 61 assertions, all green, re-run 2026-08-31.
295
+
296
+ offline, no keys, on fixtures that inherit the real transcript shape including
297
+ all four ways a non-human record disguises itself as `type: user`.
298
+
299
+ the load-bearing one in `test_coach.sh` is `REVERTED IS NOT SURVIVED`: a commit
300
+ that got `reset --hard` in the same session left no durable record. flip that one
301
+ line and the suite goes red, which is the point. a generous proxy is a broken one.
302
+
303
+ `test_label_bands.sh` is the other one, and it exists because 0.1.1 shipped the
304
+ defect it pins. `SURVIVES MOST` took the top five habits and `SURVIVES LEAST` took
305
+ the bottom five, which overlap whenever you have fewer than ten rankable habits —
306
+ so a new user, who necessarily has few, read the same habit at the same percentage
307
+ under both "do more of these" and "these tend to loop". on a 3-habit corpus 0.1.1
308
+ reprinted all three, all at 65% (22/34). the suite is red on the published 0.1.1
309
+ file and green on this one, and its `wide` band asserts the fix leaves a large
310
+ corpus byte-identical.
311
+
312
+ ## why though
313
+
314
+ your agent history is proof. every "yeah it's done" has a real trace sitting
315
+ behind it. transcripto is the index that makes it checkable.
316
+
317
+ it is the fuel layer. on top of it you check what your agents *claim* against what
318
+ the trace *shows*, which is [mountain of helicon](https://github.com/Morkeeth/mountain-of-helicon).
319
+ the pitch was never "search your history". it is *prove your agent did what it
320
+ said, from your own local traces.*
321
+
322
+ local, MIT, no telemetry. star it if it finds you something you'd lost ™
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
4
4
 
5
5
  [project]
6
6
  name = "transcripto"
7
- version = "0.1.0"
7
+ version = "0.1.2"
8
8
  description = "Search everything your coding agents ever did, grade your own prompts, and price your decisions. Local, stdlib-only, your data never leaves the machine."
9
9
  readme = "README.md"
10
10
  requires-python = ">=3.9"