transcripto 0.1.3__tar.gz → 0.1.5__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: transcripto
3
- Version: 0.1.3
3
+ Version: 0.1.5
4
4
  Summary: Search everything your coding agents ever did, grade your own prompts, and price your decisions. Local, stdlib-only, your data never leaves the machine.
5
5
  Author: Oscar Morke
6
6
  License: MIT
@@ -29,7 +29,7 @@ your disk and never opens a socket.
29
29
  uvx transcripto coach
30
30
  ```
31
31
 
32
- > **which build you got.** `uvx transcripto --version` should print **0.1.3**. If `--version`
32
+ > **which build you got.** `uvx transcripto --version` should print **0.1.4**. If `--version`
33
33
  > is not a recognised flag at all you are on 0.1.1, which predates `trace`, Cursor support and
34
34
  > this README. `uvx --refresh transcripto` forces a fresh resolve past uv's cache.
35
35
  >
@@ -73,18 +73,18 @@ the same way.
73
73
 
74
74
  ## what you get back
75
75
 
76
- this is a real run on one machine, pasted unedited, 2026-08-31:
76
+ this is a real run on one machine, 2026-08-31. The numbers are unedited; the two quoted prompts are synthetic stand-ins of the same shape, because your prompts never leave your machine and neither do mine:
77
77
 
78
78
  ```
79
79
 
80
80
  YOUR PROMPT HABITS, GRADED (offline, your machine only)
81
81
 
82
82
  - your worst looped prompt, with its witness:
83
- "ok, and lets see they might solve it in the future so i can go back to my beloeved routine :) Befor…"
83
+ "hmm ok lets just try again and see if it works this time, same thing as before but…"
84
84
  NO-DURABLE-RECORD: read-only Bash only, no file change · corrections: 15 · assistant turns: 221
85
85
 
86
86
  + your best landed prompt, with its witness:
87
- "mTERMINAL 8 — Mountain of Helicon · ~/CODE/mountain-of-helicon Read ~/CODE/mountain-of-helicon. Two…"
87
+ "TERMINAL 3 — parser · ~/CODE/demo Read hack.md. Fix the counter, run the suite, commit if green…"
88
88
  COMMIT-WITNESSED: git commit · corrections: 0
89
89
 
90
90
  SURVIVAL IS A PROXY: survival = a durable Write/Edit or an un-reverted git commit in-episode. A PROXY, not proof the work was correct or shipped.
@@ -142,6 +142,74 @@ signal, not a verdict. if it ever prints something that flatters you, distrust i
142
142
 
143
143
  the caveat is printed in the output itself, every run, on purpose.
144
144
 
145
+ ## correction rate
146
+
147
+ > in the repo since 2026-09-02, not yet on PyPI. `python3 transcripto.py coach` prints it.
148
+
149
+ one more line under the coach footer:
150
+
151
+ ```
152
+ correction rate: 6% measured (227 of 4061 typed turns) · v1 catches ~1 in 6, so the real rate is ~26-37%
153
+ ```
154
+
155
+ **correction rate = typed turns that correct the agent ÷ typed turns.** the denominator
156
+ is the same authorship gate as every other number here (`typed` or `queued`, never
157
+ `isMeta`, `isSidechain` or a tool result), so a tool result that happens to say "wrong"
158
+ cannot move it. a typed turn counts as a correction when a **marker** fires in its head:
159
+ the first 80 words after URLs and file paths are stripped, case-insensitive, whole-word:
160
+ a leading or bare `no`, `again`, `wrong`, `not that`, `I meant`, `revert`, `undo`, `stop`,
161
+ `instead`, plus a few measured additions. whole-word, so "against" is not "again" and
162
+ "now" is not "no". the v0 "nudge" rule (a short turn naming the agent's last file) is gone:
163
+ measured over 100 flagged rows it fired alone five times and was wrong five times.
164
+ `TRANSCRIPTO_CORRECTION=v0` brings the old classifier back for comparison.
165
+
166
+ it is one pure function, `is_correction(text)`, and it is a **floor**, now measured rather
167
+ than asserted: on a 200-row labelled sample the v1 classifier has precision 0.81 and recall
168
+ 0.16 (`docs/CORRECTION-PRECISION-2026-09-03.md`), which is why the printed line carries the
169
+ "~1 in 6" correction and a range. read it as a trend on your own history.
170
+ `test_correction.sh` pins the rule and the gate on `fixtures-correction/`.
171
+
172
+ ## export-run
173
+
174
+ > in the repo since 2026-09-02, not yet on PyPI.
175
+
176
+ ```
177
+ transcripto export-run latest # the newest session on this machine
178
+ transcripto export-run 3f9c1a2b # a session id, or a prefix of one
179
+ transcripto export-run path/to/session.jsonl # a transcript file
180
+ transcripto export-run latest --harness codex # --root / --harness as for coach
181
+ ```
182
+
183
+ one run's numbers as JSON, read straight from the transcript file (no index needed).
184
+ this is the contract other tools read (Agent Grinder's card, ZUP's board); the keys are
185
+ frozen under `schema`, and a new key is an addition, never a rename.
186
+
187
+ | key | meaning |
188
+ |---|---|
189
+ | `schema` | `transcripto.export-run/1` |
190
+ | `session_id` | the harness's session id (Claude Code: the file name; Codex: `session_meta.id`; Cursor: the file name) |
191
+ | `project` | the run's `cwd` |
192
+ | `harness` | `claude` · `codex` · `cursor` |
193
+ | `transcript` | absolute path of the file read |
194
+ | `started` · `ended` | first and last record timestamp, UTC, `…Z`; `null` if the file carries none |
195
+ | `duration_s` | `ended − started`, whole seconds |
196
+ | `records` | every record in the file, before any gate |
197
+ | `typed_turns` | records that pass the authorship gate: `promptSource` typed or queued, never meta, sidechain or tool result. the same count coach prints as "typed by you" |
198
+ | `corrections` | typed turns `is_correction()` flags (see above) |
199
+ | `correction_rate` | `corrections / typed_turns`, 3 decimals; `null` when nothing was typed |
200
+ | `tool_calls` | every `tool_use` block the agent emitted |
201
+ | `files_touched` | sorted set of `file_path` (or `notebook_path`) from Edit / Write / Read / MultiEdit / NotebookEdit calls |
202
+ | `commits_in_window` | commits stamped inside `[started, ended]` in the project's git reflog; `null` when `project` is not inside a git repo |
203
+ | `commits` | those commits as `{sha, ts, subject}`, oldest first; `null` when not a repo |
204
+ | `proxy` | the caveat, in the JSON so it travels with the numbers |
205
+
206
+ `commits_in_window` reads `.git/logs/HEAD` directly, not `git log`, because this file
207
+ does not shell out (see privacy). the reflog is the record of what **that working tree**
208
+ did: a commit made there in the window is in it, a commit pulled in from elsewhere is not.
209
+ git expires the reflog after 90 days by default, so a run older than that can read 0 here
210
+ while `git log` would still show its commits. `commit (amend)` counts; a rebase's `pick`
211
+ lines do not. a `.git` file (a worktree) is followed to its gitdir.
212
+
145
213
  ## why your own gate matters here
146
214
 
147
215
  at fleet scale roughly 95% of the `type: user` records in a transcript are not
@@ -199,15 +267,15 @@ claim, so here is the grep that settles it against the single file it ships as:
199
267
  $ grep -nE '^[[:space:]]*(import|from) ' transcripto.py
200
268
  8:import sys, os, json, glob, re, sqlite3, argparse
201
269
  9:from datetime import datetime, timezone
202
- 260: import time
203
- 1051: import datetime
204
- 1061: from the separator), so the result is checked on disk and dropped if it is
270
+ 261: import time
271
+ 1160: import datetime
272
+ 1170: from the separator), so the result is checked on disk and dropped if it is
205
273
  ```
206
274
 
207
275
  five lines, four of which are imports and all four are stdlib. `time` and
208
276
  `datetime` sit inside functions, which is why the pattern allows for indentation —
209
277
  anchor it at `^import` and you would miss two, so do not take my word for the
210
- anchor either. line 1061 is the pattern catching a docstring that happens to begin
278
+ anchor either. line 1170 is the pattern catching a docstring that happens to begin
211
279
  with the word `from`; it is prose, not an import, and it is left in rather than
212
280
  tuned out, because a grep you tuned until it agreed with you proves nothing.
213
281
 
@@ -246,6 +314,7 @@ transcripto sessions recent sessions + the first prompt YOU typed in each
246
314
  transcripto stats what you actually work on
247
315
  transcripto cost what ONE of your decisions costs
248
316
  transcripto coach which of YOUR prompt habits survive (a proxy)
317
+ transcripto export-run one run's numbers as JSON (typed turns, correction rate, commits)
249
318
  ```
250
319
 
251
320
  `ask` is the one that kills "wait, did i lose something?". it answers "what was i
@@ -255,10 +324,10 @@ thinking about X across ALL my sessions", in your own words only.
255
324
  $ transcripto find USER-JOURNEY.md # run 2026-08-29
256
325
  USER-JOURNEY.md 4 touches across sessions (3 were writes/edits)
257
326
 
258
- 2026-08-20 WROTE ~/CODE/mountain-of-helicon-main/USER-JOURNEY.md abd9e871
259
- 2026-08-21 WROTE ~/…/Obsidian LIFE/00 Dashboard/suite-user-journey.md 0f845ede
260
- 2026-08-27 read ~/CODE/hack-fleet-ata/docs/USER-JOURNEY.md cddfde29
261
- 2026-08-27 WROTE ~/CODE/hack-fleet-ata/docs/USER-JOURNEY.md cddfde29
327
+ 2026-08-20 WROTE ~/CODE/demo/USER-JOURNEY.md abd9e871
328
+ 2026-08-21 WROTE ~/CODE/demo/docs/onboarding-notes.md 0f845ede
329
+ 2026-08-27 read ~/CODE/demo-api/docs/USER-JOURNEY.md cddfde29
330
+ 2026-08-27 WROTE ~/CODE/demo-api/docs/USER-JOURNEY.md cddfde29
262
331
  ```
263
332
 
264
333
  the file you lost, found across every session you ever ran, with the session id
@@ -306,9 +375,12 @@ just gives the file a name on your PATH.
306
375
  ./test_cost.sh 12 assertions
307
376
  ./test_small_n.sh 7 assertions
308
377
  ./test_label_bands.sh 13 assertions
378
+ ./test_correction.sh 32 assertions (correction rate + export-run, 2026-09-02)
379
+ ./test_cursor_partial.sh 7 assertions
380
+ ./test_version.sh 1 assertion (VERSION matches in transcripto.py and pyproject.toml)
309
381
  ```
310
382
 
311
- 61 assertions, all green, re-run 2026-08-31.
383
+ 91 assertions, all green, re-run 2026-09-02.
312
384
 
313
385
  offline, no keys, on fixtures that inherit the real transcript shape including
314
386
  all four ways a non-human record disguises itself as `type: user`.
@@ -11,7 +11,7 @@ your disk and never opens a socket.
11
11
  uvx transcripto coach
12
12
  ```
13
13
 
14
- > **which build you got.** `uvx transcripto --version` should print **0.1.3**. If `--version`
14
+ > **which build you got.** `uvx transcripto --version` should print **0.1.4**. If `--version`
15
15
  > is not a recognised flag at all you are on 0.1.1, which predates `trace`, Cursor support and
16
16
  > this README. `uvx --refresh transcripto` forces a fresh resolve past uv's cache.
17
17
  >
@@ -55,18 +55,18 @@ the same way.
55
55
 
56
56
  ## what you get back
57
57
 
58
- this is a real run on one machine, pasted unedited, 2026-08-31:
58
+ this is a real run on one machine, 2026-08-31. The numbers are unedited; the two quoted prompts are synthetic stand-ins of the same shape, because your prompts never leave your machine and neither do mine:
59
59
 
60
60
  ```
61
61
 
62
62
  YOUR PROMPT HABITS, GRADED (offline, your machine only)
63
63
 
64
64
  - your worst looped prompt, with its witness:
65
- "ok, and lets see they might solve it in the future so i can go back to my beloeved routine :) Befor…"
65
+ "hmm ok lets just try again and see if it works this time, same thing as before but…"
66
66
  NO-DURABLE-RECORD: read-only Bash only, no file change · corrections: 15 · assistant turns: 221
67
67
 
68
68
  + your best landed prompt, with its witness:
69
- "mTERMINAL 8 — Mountain of Helicon · ~/CODE/mountain-of-helicon Read ~/CODE/mountain-of-helicon. Two…"
69
+ "TERMINAL 3 — parser · ~/CODE/demo Read hack.md. Fix the counter, run the suite, commit if green…"
70
70
  COMMIT-WITNESSED: git commit · corrections: 0
71
71
 
72
72
  SURVIVAL IS A PROXY: survival = a durable Write/Edit or an un-reverted git commit in-episode. A PROXY, not proof the work was correct or shipped.
@@ -124,6 +124,74 @@ signal, not a verdict. if it ever prints something that flatters you, distrust i
124
124
 
125
125
  the caveat is printed in the output itself, every run, on purpose.
126
126
 
127
+ ## correction rate
128
+
129
+ > in the repo since 2026-09-02, not yet on PyPI. `python3 transcripto.py coach` prints it.
130
+
131
+ one more line under the coach footer:
132
+
133
+ ```
134
+ correction rate: 6% measured (227 of 4061 typed turns) · v1 catches ~1 in 6, so the real rate is ~26-37%
135
+ ```
136
+
137
+ **correction rate = typed turns that correct the agent ÷ typed turns.** the denominator
138
+ is the same authorship gate as every other number here (`typed` or `queued`, never
139
+ `isMeta`, `isSidechain` or a tool result), so a tool result that happens to say "wrong"
140
+ cannot move it. a typed turn counts as a correction when a **marker** fires in its head:
141
+ the first 80 words after URLs and file paths are stripped, case-insensitive, whole-word:
142
+ a leading or bare `no`, `again`, `wrong`, `not that`, `I meant`, `revert`, `undo`, `stop`,
143
+ `instead`, plus a few measured additions. whole-word, so "against" is not "again" and
144
+ "now" is not "no". the v0 "nudge" rule (a short turn naming the agent's last file) is gone:
145
+ measured over 100 flagged rows it fired alone five times and was wrong five times.
146
+ `TRANSCRIPTO_CORRECTION=v0` brings the old classifier back for comparison.
147
+
148
+ it is one pure function, `is_correction(text)`, and it is a **floor**, now measured rather
149
+ than asserted: on a 200-row labelled sample the v1 classifier has precision 0.81 and recall
150
+ 0.16 (`docs/CORRECTION-PRECISION-2026-09-03.md`), which is why the printed line carries the
151
+ "~1 in 6" correction and a range. read it as a trend on your own history.
152
+ `test_correction.sh` pins the rule and the gate on `fixtures-correction/`.
153
+
154
+ ## export-run
155
+
156
+ > in the repo since 2026-09-02, not yet on PyPI.
157
+
158
+ ```
159
+ transcripto export-run latest # the newest session on this machine
160
+ transcripto export-run 3f9c1a2b # a session id, or a prefix of one
161
+ transcripto export-run path/to/session.jsonl # a transcript file
162
+ transcripto export-run latest --harness codex # --root / --harness as for coach
163
+ ```
164
+
165
+ one run's numbers as JSON, read straight from the transcript file (no index needed).
166
+ this is the contract other tools read (Agent Grinder's card, ZUP's board); the keys are
167
+ frozen under `schema`, and a new key is an addition, never a rename.
168
+
169
+ | key | meaning |
170
+ |---|---|
171
+ | `schema` | `transcripto.export-run/1` |
172
+ | `session_id` | the harness's session id (Claude Code: the file name; Codex: `session_meta.id`; Cursor: the file name) |
173
+ | `project` | the run's `cwd` |
174
+ | `harness` | `claude` · `codex` · `cursor` |
175
+ | `transcript` | absolute path of the file read |
176
+ | `started` · `ended` | first and last record timestamp, UTC, `…Z`; `null` if the file carries none |
177
+ | `duration_s` | `ended − started`, whole seconds |
178
+ | `records` | every record in the file, before any gate |
179
+ | `typed_turns` | records that pass the authorship gate: `promptSource` typed or queued, never meta, sidechain or tool result. the same count coach prints as "typed by you" |
180
+ | `corrections` | typed turns `is_correction()` flags (see above) |
181
+ | `correction_rate` | `corrections / typed_turns`, 3 decimals; `null` when nothing was typed |
182
+ | `tool_calls` | every `tool_use` block the agent emitted |
183
+ | `files_touched` | sorted set of `file_path` (or `notebook_path`) from Edit / Write / Read / MultiEdit / NotebookEdit calls |
184
+ | `commits_in_window` | commits stamped inside `[started, ended]` in the project's git reflog; `null` when `project` is not inside a git repo |
185
+ | `commits` | those commits as `{sha, ts, subject}`, oldest first; `null` when not a repo |
186
+ | `proxy` | the caveat, in the JSON so it travels with the numbers |
187
+
188
+ `commits_in_window` reads `.git/logs/HEAD` directly, not `git log`, because this file
189
+ does not shell out (see privacy). the reflog is the record of what **that working tree**
190
+ did: a commit made there in the window is in it, a commit pulled in from elsewhere is not.
191
+ git expires the reflog after 90 days by default, so a run older than that can read 0 here
192
+ while `git log` would still show its commits. `commit (amend)` counts; a rebase's `pick`
193
+ lines do not. a `.git` file (a worktree) is followed to its gitdir.
194
+
127
195
  ## why your own gate matters here
128
196
 
129
197
  at fleet scale roughly 95% of the `type: user` records in a transcript are not
@@ -181,15 +249,15 @@ claim, so here is the grep that settles it against the single file it ships as:
181
249
  $ grep -nE '^[[:space:]]*(import|from) ' transcripto.py
182
250
  8:import sys, os, json, glob, re, sqlite3, argparse
183
251
  9:from datetime import datetime, timezone
184
- 260: import time
185
- 1051: import datetime
186
- 1061: from the separator), so the result is checked on disk and dropped if it is
252
+ 261: import time
253
+ 1160: import datetime
254
+ 1170: from the separator), so the result is checked on disk and dropped if it is
187
255
  ```
188
256
 
189
257
  five lines, four of which are imports and all four are stdlib. `time` and
190
258
  `datetime` sit inside functions, which is why the pattern allows for indentation —
191
259
  anchor it at `^import` and you would miss two, so do not take my word for the
192
- anchor either. line 1061 is the pattern catching a docstring that happens to begin
260
+ anchor either. line 1170 is the pattern catching a docstring that happens to begin
193
261
  with the word `from`; it is prose, not an import, and it is left in rather than
194
262
  tuned out, because a grep you tuned until it agreed with you proves nothing.
195
263
 
@@ -228,6 +296,7 @@ transcripto sessions recent sessions + the first prompt YOU typed in each
228
296
  transcripto stats what you actually work on
229
297
  transcripto cost what ONE of your decisions costs
230
298
  transcripto coach which of YOUR prompt habits survive (a proxy)
299
+ transcripto export-run one run's numbers as JSON (typed turns, correction rate, commits)
231
300
  ```
232
301
 
233
302
  `ask` is the one that kills "wait, did i lose something?". it answers "what was i
@@ -237,10 +306,10 @@ thinking about X across ALL my sessions", in your own words only.
237
306
  $ transcripto find USER-JOURNEY.md # run 2026-08-29
238
307
  USER-JOURNEY.md 4 touches across sessions (3 were writes/edits)
239
308
 
240
- 2026-08-20 WROTE ~/CODE/mountain-of-helicon-main/USER-JOURNEY.md abd9e871
241
- 2026-08-21 WROTE ~/…/Obsidian LIFE/00 Dashboard/suite-user-journey.md 0f845ede
242
- 2026-08-27 read ~/CODE/hack-fleet-ata/docs/USER-JOURNEY.md cddfde29
243
- 2026-08-27 WROTE ~/CODE/hack-fleet-ata/docs/USER-JOURNEY.md cddfde29
309
+ 2026-08-20 WROTE ~/CODE/demo/USER-JOURNEY.md abd9e871
310
+ 2026-08-21 WROTE ~/CODE/demo/docs/onboarding-notes.md 0f845ede
311
+ 2026-08-27 read ~/CODE/demo-api/docs/USER-JOURNEY.md cddfde29
312
+ 2026-08-27 WROTE ~/CODE/demo-api/docs/USER-JOURNEY.md cddfde29
244
313
  ```
245
314
 
246
315
  the file you lost, found across every session you ever ran, with the session id
@@ -288,9 +357,12 @@ just gives the file a name on your PATH.
288
357
  ./test_cost.sh 12 assertions
289
358
  ./test_small_n.sh 7 assertions
290
359
  ./test_label_bands.sh 13 assertions
360
+ ./test_correction.sh 32 assertions (correction rate + export-run, 2026-09-02)
361
+ ./test_cursor_partial.sh 7 assertions
362
+ ./test_version.sh 1 assertion (VERSION matches in transcripto.py and pyproject.toml)
291
363
  ```
292
364
 
293
- 61 assertions, all green, re-run 2026-08-31.
365
+ 91 assertions, all green, re-run 2026-09-02.
294
366
 
295
367
  offline, no keys, on fixtures that inherit the real transcript shape including
296
368
  all four ways a non-human record disguises itself as `type: user`.
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
4
4
 
5
5
  [project]
6
6
  name = "transcripto"
7
- version = "0.1.3"
7
+ version = "0.1.5"
8
8
  description = "Search everything your coding agents ever did, grade your own prompts, and price your decisions. Local, stdlib-only, your data never leaves the machine."
9
9
  readme = "README.md"
10
10
  requires-python = ">=3.9"
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: transcripto
3
- Version: 0.1.3
3
+ Version: 0.1.5
4
4
  Summary: Search everything your coding agents ever did, grade your own prompts, and price your decisions. Local, stdlib-only, your data never leaves the machine.
5
5
  Author: Oscar Morke
6
6
  License: MIT
@@ -29,7 +29,7 @@ your disk and never opens a socket.
29
29
  uvx transcripto coach
30
30
  ```
31
31
 
32
- > **which build you got.** `uvx transcripto --version` should print **0.1.3**. If `--version`
32
+ > **which build you got.** `uvx transcripto --version` should print **0.1.4**. If `--version`
33
33
  > is not a recognised flag at all you are on 0.1.1, which predates `trace`, Cursor support and
34
34
  > this README. `uvx --refresh transcripto` forces a fresh resolve past uv's cache.
35
35
  >
@@ -73,18 +73,18 @@ the same way.
73
73
 
74
74
  ## what you get back
75
75
 
76
- this is a real run on one machine, pasted unedited, 2026-08-31:
76
+ this is a real run on one machine, 2026-08-31. The numbers are unedited; the two quoted prompts are synthetic stand-ins of the same shape, because your prompts never leave your machine and neither do mine:
77
77
 
78
78
  ```
79
79
 
80
80
  YOUR PROMPT HABITS, GRADED (offline, your machine only)
81
81
 
82
82
  - your worst looped prompt, with its witness:
83
- "ok, and lets see they might solve it in the future so i can go back to my beloeved routine :) Befor…"
83
+ "hmm ok lets just try again and see if it works this time, same thing as before but…"
84
84
  NO-DURABLE-RECORD: read-only Bash only, no file change · corrections: 15 · assistant turns: 221
85
85
 
86
86
  + your best landed prompt, with its witness:
87
- "mTERMINAL 8 — Mountain of Helicon · ~/CODE/mountain-of-helicon Read ~/CODE/mountain-of-helicon. Two…"
87
+ "TERMINAL 3 — parser · ~/CODE/demo Read hack.md. Fix the counter, run the suite, commit if green…"
88
88
  COMMIT-WITNESSED: git commit · corrections: 0
89
89
 
90
90
  SURVIVAL IS A PROXY: survival = a durable Write/Edit or an un-reverted git commit in-episode. A PROXY, not proof the work was correct or shipped.
@@ -142,6 +142,74 @@ signal, not a verdict. if it ever prints something that flatters you, distrust i
142
142
 
143
143
  the caveat is printed in the output itself, every run, on purpose.
144
144
 
145
+ ## correction rate
146
+
147
+ > in the repo since 2026-09-02, not yet on PyPI. `python3 transcripto.py coach` prints it.
148
+
149
+ one more line under the coach footer:
150
+
151
+ ```
152
+ correction rate: 6% measured (227 of 4061 typed turns) · v1 catches ~1 in 6, so the real rate is ~26-37%
153
+ ```
154
+
155
+ **correction rate = typed turns that correct the agent ÷ typed turns.** the denominator
156
+ is the same authorship gate as every other number here (`typed` or `queued`, never
157
+ `isMeta`, `isSidechain` or a tool result), so a tool result that happens to say "wrong"
158
+ cannot move it. a typed turn counts as a correction when a **marker** fires in its head:
159
+ the first 80 words after URLs and file paths are stripped, case-insensitive, whole-word:
160
+ a leading or bare `no`, `again`, `wrong`, `not that`, `I meant`, `revert`, `undo`, `stop`,
161
+ `instead`, plus a few measured additions. whole-word, so "against" is not "again" and
162
+ "now" is not "no". the v0 "nudge" rule (a short turn naming the agent's last file) is gone:
163
+ measured over 100 flagged rows it fired alone five times and was wrong five times.
164
+ `TRANSCRIPTO_CORRECTION=v0` brings the old classifier back for comparison.
165
+
166
+ it is one pure function, `is_correction(text)`, and it is a **floor**, now measured rather
167
+ than asserted: on a 200-row labelled sample the v1 classifier has precision 0.81 and recall
168
+ 0.16 (`docs/CORRECTION-PRECISION-2026-09-03.md`), which is why the printed line carries the
169
+ "~1 in 6" correction and a range. read it as a trend on your own history.
170
+ `test_correction.sh` pins the rule and the gate on `fixtures-correction/`.
171
+
172
+ ## export-run
173
+
174
+ > in the repo since 2026-09-02, not yet on PyPI.
175
+
176
+ ```
177
+ transcripto export-run latest # the newest session on this machine
178
+ transcripto export-run 3f9c1a2b # a session id, or a prefix of one
179
+ transcripto export-run path/to/session.jsonl # a transcript file
180
+ transcripto export-run latest --harness codex # --root / --harness as for coach
181
+ ```
182
+
183
+ one run's numbers as JSON, read straight from the transcript file (no index needed).
184
+ this is the contract other tools read (Agent Grinder's card, ZUP's board); the keys are
185
+ frozen under `schema`, and a new key is an addition, never a rename.
186
+
187
+ | key | meaning |
188
+ |---|---|
189
+ | `schema` | `transcripto.export-run/1` |
190
+ | `session_id` | the harness's session id (Claude Code: the file name; Codex: `session_meta.id`; Cursor: the file name) |
191
+ | `project` | the run's `cwd` |
192
+ | `harness` | `claude` · `codex` · `cursor` |
193
+ | `transcript` | absolute path of the file read |
194
+ | `started` · `ended` | first and last record timestamp, UTC, `…Z`; `null` if the file carries none |
195
+ | `duration_s` | `ended − started`, whole seconds |
196
+ | `records` | every record in the file, before any gate |
197
+ | `typed_turns` | records that pass the authorship gate: `promptSource` typed or queued, never meta, sidechain or tool result. the same count coach prints as "typed by you" |
198
+ | `corrections` | typed turns `is_correction()` flags (see above) |
199
+ | `correction_rate` | `corrections / typed_turns`, 3 decimals; `null` when nothing was typed |
200
+ | `tool_calls` | every `tool_use` block the agent emitted |
201
+ | `files_touched` | sorted set of `file_path` (or `notebook_path`) from Edit / Write / Read / MultiEdit / NotebookEdit calls |
202
+ | `commits_in_window` | commits stamped inside `[started, ended]` in the project's git reflog; `null` when `project` is not inside a git repo |
203
+ | `commits` | those commits as `{sha, ts, subject}`, oldest first; `null` when not a repo |
204
+ | `proxy` | the caveat, in the JSON so it travels with the numbers |
205
+
206
+ `commits_in_window` reads `.git/logs/HEAD` directly, not `git log`, because this file
207
+ does not shell out (see privacy). the reflog is the record of what **that working tree**
208
+ did: a commit made there in the window is in it, a commit pulled in from elsewhere is not.
209
+ git expires the reflog after 90 days by default, so a run older than that can read 0 here
210
+ while `git log` would still show its commits. `commit (amend)` counts; a rebase's `pick`
211
+ lines do not. a `.git` file (a worktree) is followed to its gitdir.
212
+
145
213
  ## why your own gate matters here
146
214
 
147
215
  at fleet scale roughly 95% of the `type: user` records in a transcript are not
@@ -199,15 +267,15 @@ claim, so here is the grep that settles it against the single file it ships as:
199
267
  $ grep -nE '^[[:space:]]*(import|from) ' transcripto.py
200
268
  8:import sys, os, json, glob, re, sqlite3, argparse
201
269
  9:from datetime import datetime, timezone
202
- 260: import time
203
- 1051: import datetime
204
- 1061: from the separator), so the result is checked on disk and dropped if it is
270
+ 261: import time
271
+ 1160: import datetime
272
+ 1170: from the separator), so the result is checked on disk and dropped if it is
205
273
  ```
206
274
 
207
275
  five lines, four of which are imports and all four are stdlib. `time` and
208
276
  `datetime` sit inside functions, which is why the pattern allows for indentation —
209
277
  anchor it at `^import` and you would miss two, so do not take my word for the
210
- anchor either. line 1061 is the pattern catching a docstring that happens to begin
278
+ anchor either. line 1170 is the pattern catching a docstring that happens to begin
211
279
  with the word `from`; it is prose, not an import, and it is left in rather than
212
280
  tuned out, because a grep you tuned until it agreed with you proves nothing.
213
281
 
@@ -246,6 +314,7 @@ transcripto sessions recent sessions + the first prompt YOU typed in each
246
314
  transcripto stats what you actually work on
247
315
  transcripto cost what ONE of your decisions costs
248
316
  transcripto coach which of YOUR prompt habits survive (a proxy)
317
+ transcripto export-run one run's numbers as JSON (typed turns, correction rate, commits)
249
318
  ```
250
319
 
251
320
  `ask` is the one that kills "wait, did i lose something?". it answers "what was i
@@ -255,10 +324,10 @@ thinking about X across ALL my sessions", in your own words only.
255
324
  $ transcripto find USER-JOURNEY.md # run 2026-08-29
256
325
  USER-JOURNEY.md 4 touches across sessions (3 were writes/edits)
257
326
 
258
- 2026-08-20 WROTE ~/CODE/mountain-of-helicon-main/USER-JOURNEY.md abd9e871
259
- 2026-08-21 WROTE ~/…/Obsidian LIFE/00 Dashboard/suite-user-journey.md 0f845ede
260
- 2026-08-27 read ~/CODE/hack-fleet-ata/docs/USER-JOURNEY.md cddfde29
261
- 2026-08-27 WROTE ~/CODE/hack-fleet-ata/docs/USER-JOURNEY.md cddfde29
327
+ 2026-08-20 WROTE ~/CODE/demo/USER-JOURNEY.md abd9e871
328
+ 2026-08-21 WROTE ~/CODE/demo/docs/onboarding-notes.md 0f845ede
329
+ 2026-08-27 read ~/CODE/demo-api/docs/USER-JOURNEY.md cddfde29
330
+ 2026-08-27 WROTE ~/CODE/demo-api/docs/USER-JOURNEY.md cddfde29
262
331
  ```
263
332
 
264
333
  the file you lost, found across every session you ever ran, with the session id
@@ -306,9 +375,12 @@ just gives the file a name on your PATH.
306
375
  ./test_cost.sh 12 assertions
307
376
  ./test_small_n.sh 7 assertions
308
377
  ./test_label_bands.sh 13 assertions
378
+ ./test_correction.sh 32 assertions (correction rate + export-run, 2026-09-02)
379
+ ./test_cursor_partial.sh 7 assertions
380
+ ./test_version.sh 1 assertion (VERSION matches in transcripto.py and pyproject.toml)
309
381
  ```
310
382
 
311
- 61 assertions, all green, re-run 2026-08-31.
383
+ 91 assertions, all green, re-run 2026-09-02.
312
384
 
313
385
  offline, no keys, on fixtures that inherit the real transcript shape including
314
386
  all four ways a non-human record disguises itself as `type: user`.