transcripto 0.1.3__tar.gz → 0.1.4__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {transcripto-0.1.3/transcripto.egg-info → transcripto-0.1.4}/PKG-INFO +81 -9
- {transcripto-0.1.3 → transcripto-0.1.4}/README.md +80 -8
- {transcripto-0.1.3 → transcripto-0.1.4}/pyproject.toml +1 -1
- {transcripto-0.1.3 → transcripto-0.1.4/transcripto.egg-info}/PKG-INFO +81 -9
- {transcripto-0.1.3 → transcripto-0.1.4}/transcripto.py +407 -15
- {transcripto-0.1.3 → transcripto-0.1.4}/LICENSE +0 -0
- {transcripto-0.1.3 → transcripto-0.1.4}/setup.cfg +0 -0
- {transcripto-0.1.3 → transcripto-0.1.4}/transcripto.egg-info/SOURCES.txt +0 -0
- {transcripto-0.1.3 → transcripto-0.1.4}/transcripto.egg-info/dependency_links.txt +0 -0
- {transcripto-0.1.3 → transcripto-0.1.4}/transcripto.egg-info/entry_points.txt +0 -0
- {transcripto-0.1.3 → transcripto-0.1.4}/transcripto.egg-info/top_level.txt +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: transcripto
|
|
3
|
-
Version: 0.1.
|
|
3
|
+
Version: 0.1.4
|
|
4
4
|
Summary: Search everything your coding agents ever did, grade your own prompts, and price your decisions. Local, stdlib-only, your data never leaves the machine.
|
|
5
5
|
Author: Oscar Morke
|
|
6
6
|
License: MIT
|
|
@@ -73,18 +73,18 @@ the same way.
|
|
|
73
73
|
|
|
74
74
|
## what you get back
|
|
75
75
|
|
|
76
|
-
this is a real run on one machine,
|
|
76
|
+
this is a real run on one machine, 2026-08-31. The numbers are unedited; the two quoted prompts are synthetic stand-ins of the same shape, because your prompts never leave your machine and neither do mine:
|
|
77
77
|
|
|
78
78
|
```
|
|
79
79
|
|
|
80
80
|
YOUR PROMPT HABITS, GRADED (offline, your machine only)
|
|
81
81
|
|
|
82
82
|
- your worst looped prompt, with its witness:
|
|
83
|
-
"ok
|
|
83
|
+
"hmm ok lets just try again and see if it works this time, same thing as before but…"
|
|
84
84
|
NO-DURABLE-RECORD: read-only Bash only, no file change · corrections: 15 · assistant turns: 221
|
|
85
85
|
|
|
86
86
|
+ your best landed prompt, with its witness:
|
|
87
|
-
"
|
|
87
|
+
"TERMINAL 3 — parser · ~/CODE/demo Read hack.md. Fix the counter, run the suite, commit if green…"
|
|
88
88
|
COMMIT-WITNESSED: git commit · corrections: 0
|
|
89
89
|
|
|
90
90
|
SURVIVAL IS A PROXY: survival = a durable Write/Edit or an un-reverted git commit in-episode. A PROXY, not proof the work was correct or shipped.
|
|
@@ -142,6 +142,74 @@ signal, not a verdict. if it ever prints something that flatters you, distrust i
|
|
|
142
142
|
|
|
143
143
|
the caveat is printed in the output itself, every run, on purpose.
|
|
144
144
|
|
|
145
|
+
## correction rate
|
|
146
|
+
|
|
147
|
+
> in the repo since 2026-09-02, not yet on PyPI. `python3 transcripto.py coach` prints it.
|
|
148
|
+
|
|
149
|
+
one more line under the coach footer:
|
|
150
|
+
|
|
151
|
+
```
|
|
152
|
+
correction rate: 6% measured (227 of 4061 typed turns) · v1 catches ~1 in 6, so the real rate is ~26-37%
|
|
153
|
+
```
|
|
154
|
+
|
|
155
|
+
**correction rate = typed turns that correct the agent ÷ typed turns.** the denominator
|
|
156
|
+
is the same authorship gate as every other number here (`typed` or `queued`, never
|
|
157
|
+
`isMeta`, `isSidechain` or a tool result), so a tool result that happens to say "wrong"
|
|
158
|
+
cannot move it. a typed turn counts as a correction when a **marker** fires in its head:
|
|
159
|
+
the first 80 words after URLs and file paths are stripped, case-insensitive, whole-word:
|
|
160
|
+
a leading or bare `no`, `again`, `wrong`, `not that`, `I meant`, `revert`, `undo`, `stop`,
|
|
161
|
+
`instead`, plus a few measured additions. whole-word, so "against" is not "again" and
|
|
162
|
+
"now" is not "no". the v0 "nudge" rule (a short turn naming the agent's last file) is gone:
|
|
163
|
+
measured over 100 flagged rows it fired alone five times and was wrong five times.
|
|
164
|
+
`TRANSCRIPTO_CORRECTION=v0` brings the old classifier back for comparison.
|
|
165
|
+
|
|
166
|
+
it is one pure function, `is_correction(text)`, and it is a **floor**, now measured rather
|
|
167
|
+
than asserted: on a 200-row labelled sample the v1 classifier has precision 0.81 and recall
|
|
168
|
+
0.16 (`docs/CORRECTION-PRECISION-2026-09-03.md`), which is why the printed line carries the
|
|
169
|
+
"~1 in 6" correction and a range. read it as a trend on your own history.
|
|
170
|
+
`test_correction.sh` pins the rule and the gate on `fixtures-correction/`.
|
|
171
|
+
|
|
172
|
+
## export-run
|
|
173
|
+
|
|
174
|
+
> in the repo since 2026-09-02, not yet on PyPI.
|
|
175
|
+
|
|
176
|
+
```
|
|
177
|
+
transcripto export-run latest # the newest session on this machine
|
|
178
|
+
transcripto export-run 3f9c1a2b # a session id, or a prefix of one
|
|
179
|
+
transcripto export-run path/to/session.jsonl # a transcript file
|
|
180
|
+
transcripto export-run latest --harness codex # --root / --harness as for coach
|
|
181
|
+
```
|
|
182
|
+
|
|
183
|
+
one run's numbers as JSON, read straight from the transcript file (no index needed).
|
|
184
|
+
this is the contract other tools read (Agent Grinder's card, ZUP's board); the keys are
|
|
185
|
+
frozen under `schema`, and a new key is an addition, never a rename.
|
|
186
|
+
|
|
187
|
+
| key | meaning |
|
|
188
|
+
|---|---|
|
|
189
|
+
| `schema` | `transcripto.export-run/1` |
|
|
190
|
+
| `session_id` | the harness's session id (Claude Code: the file name; Codex: `session_meta.id`; Cursor: the file name) |
|
|
191
|
+
| `project` | the run's `cwd` |
|
|
192
|
+
| `harness` | `claude` · `codex` · `cursor` |
|
|
193
|
+
| `transcript` | absolute path of the file read |
|
|
194
|
+
| `started` · `ended` | first and last record timestamp, UTC, `…Z`; `null` if the file carries none |
|
|
195
|
+
| `duration_s` | `ended − started`, whole seconds |
|
|
196
|
+
| `records` | every record in the file, before any gate |
|
|
197
|
+
| `typed_turns` | records that pass the authorship gate: `promptSource` typed or queued, never meta, sidechain or tool result. the same count coach prints as "typed by you" |
|
|
198
|
+
| `corrections` | typed turns `is_correction()` flags (see above) |
|
|
199
|
+
| `correction_rate` | `corrections / typed_turns`, 3 decimals; `null` when nothing was typed |
|
|
200
|
+
| `tool_calls` | every `tool_use` block the agent emitted |
|
|
201
|
+
| `files_touched` | sorted set of `file_path` (or `notebook_path`) from Edit / Write / Read / MultiEdit / NotebookEdit calls |
|
|
202
|
+
| `commits_in_window` | commits stamped inside `[started, ended]` in the project's git reflog; `null` when `project` is not inside a git repo |
|
|
203
|
+
| `commits` | those commits as `{sha, ts, subject}`, oldest first; `null` when not a repo |
|
|
204
|
+
| `proxy` | the caveat, in the JSON so it travels with the numbers |
|
|
205
|
+
|
|
206
|
+
`commits_in_window` reads `.git/logs/HEAD` directly, not `git log`, because this file
|
|
207
|
+
does not shell out (see privacy). the reflog is the record of what **that working tree**
|
|
208
|
+
did: a commit made there in the window is in it, a commit pulled in from elsewhere is not.
|
|
209
|
+
git expires the reflog after 90 days by default, so a run older than that can read 0 here
|
|
210
|
+
while `git log` would still show its commits. `commit (amend)` counts; a rebase's `pick`
|
|
211
|
+
lines do not. a `.git` file (a worktree) is followed to its gitdir.
|
|
212
|
+
|
|
145
213
|
## why your own gate matters here
|
|
146
214
|
|
|
147
215
|
at fleet scale roughly 95% of the `type: user` records in a transcript are not
|
|
@@ -199,15 +267,15 @@ claim, so here is the grep that settles it against the single file it ships as:
|
|
|
199
267
|
$ grep -nE '^[[:space:]]*(import|from) ' transcripto.py
|
|
200
268
|
8:import sys, os, json, glob, re, sqlite3, argparse
|
|
201
269
|
9:from datetime import datetime, timezone
|
|
202
|
-
|
|
203
|
-
|
|
204
|
-
|
|
270
|
+
261: import time
|
|
271
|
+
1160: import datetime
|
|
272
|
+
1170: from the separator), so the result is checked on disk and dropped if it is
|
|
205
273
|
```
|
|
206
274
|
|
|
207
275
|
five lines, four of which are imports and all four are stdlib. `time` and
|
|
208
276
|
`datetime` sit inside functions, which is why the pattern allows for indentation —
|
|
209
277
|
anchor it at `^import` and you would miss two, so do not take my word for the
|
|
210
|
-
anchor either. line
|
|
278
|
+
anchor either. line 1170 is the pattern catching a docstring that happens to begin
|
|
211
279
|
with the word `from`; it is prose, not an import, and it is left in rather than
|
|
212
280
|
tuned out, because a grep you tuned until it agreed with you proves nothing.
|
|
213
281
|
|
|
@@ -246,6 +314,7 @@ transcripto sessions recent sessions + the first prompt YOU typed in each
|
|
|
246
314
|
transcripto stats what you actually work on
|
|
247
315
|
transcripto cost what ONE of your decisions costs
|
|
248
316
|
transcripto coach which of YOUR prompt habits survive (a proxy)
|
|
317
|
+
transcripto export-run one run's numbers as JSON (typed turns, correction rate, commits)
|
|
249
318
|
```
|
|
250
319
|
|
|
251
320
|
`ask` is the one that kills "wait, did i lose something?". it answers "what was i
|
|
@@ -306,9 +375,12 @@ just gives the file a name on your PATH.
|
|
|
306
375
|
./test_cost.sh 12 assertions
|
|
307
376
|
./test_small_n.sh 7 assertions
|
|
308
377
|
./test_label_bands.sh 13 assertions
|
|
378
|
+
./test_correction.sh 32 assertions (correction rate + export-run, 2026-09-02)
|
|
379
|
+
./test_cursor_partial.sh 7 assertions
|
|
380
|
+
./test_version.sh 1 assertion (VERSION matches in transcripto.py and pyproject.toml)
|
|
309
381
|
```
|
|
310
382
|
|
|
311
|
-
|
|
383
|
+
91 assertions, all green, re-run 2026-09-02.
|
|
312
384
|
|
|
313
385
|
offline, no keys, on fixtures that inherit the real transcript shape including
|
|
314
386
|
all four ways a non-human record disguises itself as `type: user`.
|
|
@@ -55,18 +55,18 @@ the same way.
|
|
|
55
55
|
|
|
56
56
|
## what you get back
|
|
57
57
|
|
|
58
|
-
this is a real run on one machine,
|
|
58
|
+
this is a real run on one machine, 2026-08-31. The numbers are unedited; the two quoted prompts are synthetic stand-ins of the same shape, because your prompts never leave your machine and neither do mine:
|
|
59
59
|
|
|
60
60
|
```
|
|
61
61
|
|
|
62
62
|
YOUR PROMPT HABITS, GRADED (offline, your machine only)
|
|
63
63
|
|
|
64
64
|
- your worst looped prompt, with its witness:
|
|
65
|
-
"ok
|
|
65
|
+
"hmm ok lets just try again and see if it works this time, same thing as before but…"
|
|
66
66
|
NO-DURABLE-RECORD: read-only Bash only, no file change · corrections: 15 · assistant turns: 221
|
|
67
67
|
|
|
68
68
|
+ your best landed prompt, with its witness:
|
|
69
|
-
"
|
|
69
|
+
"TERMINAL 3 — parser · ~/CODE/demo Read hack.md. Fix the counter, run the suite, commit if green…"
|
|
70
70
|
COMMIT-WITNESSED: git commit · corrections: 0
|
|
71
71
|
|
|
72
72
|
SURVIVAL IS A PROXY: survival = a durable Write/Edit or an un-reverted git commit in-episode. A PROXY, not proof the work was correct or shipped.
|
|
@@ -124,6 +124,74 @@ signal, not a verdict. if it ever prints something that flatters you, distrust i
|
|
|
124
124
|
|
|
125
125
|
the caveat is printed in the output itself, every run, on purpose.
|
|
126
126
|
|
|
127
|
+
## correction rate
|
|
128
|
+
|
|
129
|
+
> in the repo since 2026-09-02, not yet on PyPI. `python3 transcripto.py coach` prints it.
|
|
130
|
+
|
|
131
|
+
one more line under the coach footer:
|
|
132
|
+
|
|
133
|
+
```
|
|
134
|
+
correction rate: 6% measured (227 of 4061 typed turns) · v1 catches ~1 in 6, so the real rate is ~26-37%
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
**correction rate = typed turns that correct the agent ÷ typed turns.** the denominator
|
|
138
|
+
is the same authorship gate as every other number here (`typed` or `queued`, never
|
|
139
|
+
`isMeta`, `isSidechain` or a tool result), so a tool result that happens to say "wrong"
|
|
140
|
+
cannot move it. a typed turn counts as a correction when a **marker** fires in its head:
|
|
141
|
+
the first 80 words after URLs and file paths are stripped, case-insensitive, whole-word:
|
|
142
|
+
a leading or bare `no`, `again`, `wrong`, `not that`, `I meant`, `revert`, `undo`, `stop`,
|
|
143
|
+
`instead`, plus a few measured additions. whole-word, so "against" is not "again" and
|
|
144
|
+
"now" is not "no". the v0 "nudge" rule (a short turn naming the agent's last file) is gone:
|
|
145
|
+
measured over 100 flagged rows it fired alone five times and was wrong five times.
|
|
146
|
+
`TRANSCRIPTO_CORRECTION=v0` brings the old classifier back for comparison.
|
|
147
|
+
|
|
148
|
+
it is one pure function, `is_correction(text)`, and it is a **floor**, now measured rather
|
|
149
|
+
than asserted: on a 200-row labelled sample the v1 classifier has precision 0.81 and recall
|
|
150
|
+
0.16 (`docs/CORRECTION-PRECISION-2026-09-03.md`), which is why the printed line carries the
|
|
151
|
+
"~1 in 6" correction and a range. read it as a trend on your own history.
|
|
152
|
+
`test_correction.sh` pins the rule and the gate on `fixtures-correction/`.
|
|
153
|
+
|
|
154
|
+
## export-run
|
|
155
|
+
|
|
156
|
+
> in the repo since 2026-09-02, not yet on PyPI.
|
|
157
|
+
|
|
158
|
+
```
|
|
159
|
+
transcripto export-run latest # the newest session on this machine
|
|
160
|
+
transcripto export-run 3f9c1a2b # a session id, or a prefix of one
|
|
161
|
+
transcripto export-run path/to/session.jsonl # a transcript file
|
|
162
|
+
transcripto export-run latest --harness codex # --root / --harness as for coach
|
|
163
|
+
```
|
|
164
|
+
|
|
165
|
+
one run's numbers as JSON, read straight from the transcript file (no index needed).
|
|
166
|
+
this is the contract other tools read (Agent Grinder's card, ZUP's board); the keys are
|
|
167
|
+
frozen under `schema`, and a new key is an addition, never a rename.
|
|
168
|
+
|
|
169
|
+
| key | meaning |
|
|
170
|
+
|---|---|
|
|
171
|
+
| `schema` | `transcripto.export-run/1` |
|
|
172
|
+
| `session_id` | the harness's session id (Claude Code: the file name; Codex: `session_meta.id`; Cursor: the file name) |
|
|
173
|
+
| `project` | the run's `cwd` |
|
|
174
|
+
| `harness` | `claude` · `codex` · `cursor` |
|
|
175
|
+
| `transcript` | absolute path of the file read |
|
|
176
|
+
| `started` · `ended` | first and last record timestamp, UTC, `…Z`; `null` if the file carries none |
|
|
177
|
+
| `duration_s` | `ended − started`, whole seconds |
|
|
178
|
+
| `records` | every record in the file, before any gate |
|
|
179
|
+
| `typed_turns` | records that pass the authorship gate: `promptSource` typed or queued, never meta, sidechain or tool result. the same count coach prints as "typed by you" |
|
|
180
|
+
| `corrections` | typed turns `is_correction()` flags (see above) |
|
|
181
|
+
| `correction_rate` | `corrections / typed_turns`, 3 decimals; `null` when nothing was typed |
|
|
182
|
+
| `tool_calls` | every `tool_use` block the agent emitted |
|
|
183
|
+
| `files_touched` | sorted set of `file_path` (or `notebook_path`) from Edit / Write / Read / MultiEdit / NotebookEdit calls |
|
|
184
|
+
| `commits_in_window` | commits stamped inside `[started, ended]` in the project's git reflog; `null` when `project` is not inside a git repo |
|
|
185
|
+
| `commits` | those commits as `{sha, ts, subject}`, oldest first; `null` when not a repo |
|
|
186
|
+
| `proxy` | the caveat, in the JSON so it travels with the numbers |
|
|
187
|
+
|
|
188
|
+
`commits_in_window` reads `.git/logs/HEAD` directly, not `git log`, because this file
|
|
189
|
+
does not shell out (see privacy). the reflog is the record of what **that working tree**
|
|
190
|
+
did: a commit made there in the window is in it, a commit pulled in from elsewhere is not.
|
|
191
|
+
git expires the reflog after 90 days by default, so a run older than that can read 0 here
|
|
192
|
+
while `git log` would still show its commits. `commit (amend)` counts; a rebase's `pick`
|
|
193
|
+
lines do not. a `.git` file (a worktree) is followed to its gitdir.
|
|
194
|
+
|
|
127
195
|
## why your own gate matters here
|
|
128
196
|
|
|
129
197
|
at fleet scale roughly 95% of the `type: user` records in a transcript are not
|
|
@@ -181,15 +249,15 @@ claim, so here is the grep that settles it against the single file it ships as:
|
|
|
181
249
|
$ grep -nE '^[[:space:]]*(import|from) ' transcripto.py
|
|
182
250
|
8:import sys, os, json, glob, re, sqlite3, argparse
|
|
183
251
|
9:from datetime import datetime, timezone
|
|
184
|
-
|
|
185
|
-
|
|
186
|
-
|
|
252
|
+
261: import time
|
|
253
|
+
1160: import datetime
|
|
254
|
+
1170: from the separator), so the result is checked on disk and dropped if it is
|
|
187
255
|
```
|
|
188
256
|
|
|
189
257
|
five lines, four of which are imports and all four are stdlib. `time` and
|
|
190
258
|
`datetime` sit inside functions, which is why the pattern allows for indentation —
|
|
191
259
|
anchor it at `^import` and you would miss two, so do not take my word for the
|
|
192
|
-
anchor either. line
|
|
260
|
+
anchor either. line 1170 is the pattern catching a docstring that happens to begin
|
|
193
261
|
with the word `from`; it is prose, not an import, and it is left in rather than
|
|
194
262
|
tuned out, because a grep you tuned until it agreed with you proves nothing.
|
|
195
263
|
|
|
@@ -228,6 +296,7 @@ transcripto sessions recent sessions + the first prompt YOU typed in each
|
|
|
228
296
|
transcripto stats what you actually work on
|
|
229
297
|
transcripto cost what ONE of your decisions costs
|
|
230
298
|
transcripto coach which of YOUR prompt habits survive (a proxy)
|
|
299
|
+
transcripto export-run one run's numbers as JSON (typed turns, correction rate, commits)
|
|
231
300
|
```
|
|
232
301
|
|
|
233
302
|
`ask` is the one that kills "wait, did i lose something?". it answers "what was i
|
|
@@ -288,9 +357,12 @@ just gives the file a name on your PATH.
|
|
|
288
357
|
./test_cost.sh 12 assertions
|
|
289
358
|
./test_small_n.sh 7 assertions
|
|
290
359
|
./test_label_bands.sh 13 assertions
|
|
360
|
+
./test_correction.sh 32 assertions (correction rate + export-run, 2026-09-02)
|
|
361
|
+
./test_cursor_partial.sh 7 assertions
|
|
362
|
+
./test_version.sh 1 assertion (VERSION matches in transcripto.py and pyproject.toml)
|
|
291
363
|
```
|
|
292
364
|
|
|
293
|
-
|
|
365
|
+
91 assertions, all green, re-run 2026-09-02.
|
|
294
366
|
|
|
295
367
|
offline, no keys, on fixtures that inherit the real transcript shape including
|
|
296
368
|
all four ways a non-human record disguises itself as `type: user`.
|
|
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
|
|
|
4
4
|
|
|
5
5
|
[project]
|
|
6
6
|
name = "transcripto"
|
|
7
|
-
version = "0.1.
|
|
7
|
+
version = "0.1.4"
|
|
8
8
|
description = "Search everything your coding agents ever did, grade your own prompts, and price your decisions. Local, stdlib-only, your data never leaves the machine."
|
|
9
9
|
readme = "README.md"
|
|
10
10
|
requires-python = ">=3.9"
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: transcripto
|
|
3
|
-
Version: 0.1.
|
|
3
|
+
Version: 0.1.4
|
|
4
4
|
Summary: Search everything your coding agents ever did, grade your own prompts, and price your decisions. Local, stdlib-only, your data never leaves the machine.
|
|
5
5
|
Author: Oscar Morke
|
|
6
6
|
License: MIT
|
|
@@ -73,18 +73,18 @@ the same way.
|
|
|
73
73
|
|
|
74
74
|
## what you get back
|
|
75
75
|
|
|
76
|
-
this is a real run on one machine,
|
|
76
|
+
this is a real run on one machine, 2026-08-31. The numbers are unedited; the two quoted prompts are synthetic stand-ins of the same shape, because your prompts never leave your machine and neither do mine:
|
|
77
77
|
|
|
78
78
|
```
|
|
79
79
|
|
|
80
80
|
YOUR PROMPT HABITS, GRADED (offline, your machine only)
|
|
81
81
|
|
|
82
82
|
- your worst looped prompt, with its witness:
|
|
83
|
-
"ok
|
|
83
|
+
"hmm ok lets just try again and see if it works this time, same thing as before but…"
|
|
84
84
|
NO-DURABLE-RECORD: read-only Bash only, no file change · corrections: 15 · assistant turns: 221
|
|
85
85
|
|
|
86
86
|
+ your best landed prompt, with its witness:
|
|
87
|
-
"
|
|
87
|
+
"TERMINAL 3 — parser · ~/CODE/demo Read hack.md. Fix the counter, run the suite, commit if green…"
|
|
88
88
|
COMMIT-WITNESSED: git commit · corrections: 0
|
|
89
89
|
|
|
90
90
|
SURVIVAL IS A PROXY: survival = a durable Write/Edit or an un-reverted git commit in-episode. A PROXY, not proof the work was correct or shipped.
|
|
@@ -142,6 +142,74 @@ signal, not a verdict. if it ever prints something that flatters you, distrust i
|
|
|
142
142
|
|
|
143
143
|
the caveat is printed in the output itself, every run, on purpose.
|
|
144
144
|
|
|
145
|
+
## correction rate
|
|
146
|
+
|
|
147
|
+
> in the repo since 2026-09-02, not yet on PyPI. `python3 transcripto.py coach` prints it.
|
|
148
|
+
|
|
149
|
+
one more line under the coach footer:
|
|
150
|
+
|
|
151
|
+
```
|
|
152
|
+
correction rate: 6% measured (227 of 4061 typed turns) · v1 catches ~1 in 6, so the real rate is ~26-37%
|
|
153
|
+
```
|
|
154
|
+
|
|
155
|
+
**correction rate = typed turns that correct the agent ÷ typed turns.** the denominator
|
|
156
|
+
is the same authorship gate as every other number here (`typed` or `queued`, never
|
|
157
|
+
`isMeta`, `isSidechain` or a tool result), so a tool result that happens to say "wrong"
|
|
158
|
+
cannot move it. a typed turn counts as a correction when a **marker** fires in its head:
|
|
159
|
+
the first 80 words after URLs and file paths are stripped, case-insensitive, whole-word:
|
|
160
|
+
a leading or bare `no`, `again`, `wrong`, `not that`, `I meant`, `revert`, `undo`, `stop`,
|
|
161
|
+
`instead`, plus a few measured additions. whole-word, so "against" is not "again" and
|
|
162
|
+
"now" is not "no". the v0 "nudge" rule (a short turn naming the agent's last file) is gone:
|
|
163
|
+
measured over 100 flagged rows it fired alone five times and was wrong five times.
|
|
164
|
+
`TRANSCRIPTO_CORRECTION=v0` brings the old classifier back for comparison.
|
|
165
|
+
|
|
166
|
+
it is one pure function, `is_correction(text)`, and it is a **floor**, now measured rather
|
|
167
|
+
than asserted: on a 200-row labelled sample the v1 classifier has precision 0.81 and recall
|
|
168
|
+
0.16 (`docs/CORRECTION-PRECISION-2026-09-03.md`), which is why the printed line carries the
|
|
169
|
+
"~1 in 6" correction and a range. read it as a trend on your own history.
|
|
170
|
+
`test_correction.sh` pins the rule and the gate on `fixtures-correction/`.
|
|
171
|
+
|
|
172
|
+
## export-run
|
|
173
|
+
|
|
174
|
+
> in the repo since 2026-09-02, not yet on PyPI.
|
|
175
|
+
|
|
176
|
+
```
|
|
177
|
+
transcripto export-run latest # the newest session on this machine
|
|
178
|
+
transcripto export-run 3f9c1a2b # a session id, or a prefix of one
|
|
179
|
+
transcripto export-run path/to/session.jsonl # a transcript file
|
|
180
|
+
transcripto export-run latest --harness codex # --root / --harness as for coach
|
|
181
|
+
```
|
|
182
|
+
|
|
183
|
+
one run's numbers as JSON, read straight from the transcript file (no index needed).
|
|
184
|
+
this is the contract other tools read (Agent Grinder's card, ZUP's board); the keys are
|
|
185
|
+
frozen under `schema`, and a new key is an addition, never a rename.
|
|
186
|
+
|
|
187
|
+
| key | meaning |
|
|
188
|
+
|---|---|
|
|
189
|
+
| `schema` | `transcripto.export-run/1` |
|
|
190
|
+
| `session_id` | the harness's session id (Claude Code: the file name; Codex: `session_meta.id`; Cursor: the file name) |
|
|
191
|
+
| `project` | the run's `cwd` |
|
|
192
|
+
| `harness` | `claude` · `codex` · `cursor` |
|
|
193
|
+
| `transcript` | absolute path of the file read |
|
|
194
|
+
| `started` · `ended` | first and last record timestamp, UTC, `…Z`; `null` if the file carries none |
|
|
195
|
+
| `duration_s` | `ended − started`, whole seconds |
|
|
196
|
+
| `records` | every record in the file, before any gate |
|
|
197
|
+
| `typed_turns` | records that pass the authorship gate: `promptSource` typed or queued, never meta, sidechain or tool result. the same count coach prints as "typed by you" |
|
|
198
|
+
| `corrections` | typed turns `is_correction()` flags (see above) |
|
|
199
|
+
| `correction_rate` | `corrections / typed_turns`, 3 decimals; `null` when nothing was typed |
|
|
200
|
+
| `tool_calls` | every `tool_use` block the agent emitted |
|
|
201
|
+
| `files_touched` | sorted set of `file_path` (or `notebook_path`) from Edit / Write / Read / MultiEdit / NotebookEdit calls |
|
|
202
|
+
| `commits_in_window` | commits stamped inside `[started, ended]` in the project's git reflog; `null` when `project` is not inside a git repo |
|
|
203
|
+
| `commits` | those commits as `{sha, ts, subject}`, oldest first; `null` when not a repo |
|
|
204
|
+
| `proxy` | the caveat, in the JSON so it travels with the numbers |
|
|
205
|
+
|
|
206
|
+
`commits_in_window` reads `.git/logs/HEAD` directly, not `git log`, because this file
|
|
207
|
+
does not shell out (see privacy). the reflog is the record of what **that working tree**
|
|
208
|
+
did: a commit made there in the window is in it, a commit pulled in from elsewhere is not.
|
|
209
|
+
git expires the reflog after 90 days by default, so a run older than that can read 0 here
|
|
210
|
+
while `git log` would still show its commits. `commit (amend)` counts; a rebase's `pick`
|
|
211
|
+
lines do not. a `.git` file (a worktree) is followed to its gitdir.
|
|
212
|
+
|
|
145
213
|
## why your own gate matters here
|
|
146
214
|
|
|
147
215
|
at fleet scale roughly 95% of the `type: user` records in a transcript are not
|
|
@@ -199,15 +267,15 @@ claim, so here is the grep that settles it against the single file it ships as:
|
|
|
199
267
|
$ grep -nE '^[[:space:]]*(import|from) ' transcripto.py
|
|
200
268
|
8:import sys, os, json, glob, re, sqlite3, argparse
|
|
201
269
|
9:from datetime import datetime, timezone
|
|
202
|
-
|
|
203
|
-
|
|
204
|
-
|
|
270
|
+
261: import time
|
|
271
|
+
1160: import datetime
|
|
272
|
+
1170: from the separator), so the result is checked on disk and dropped if it is
|
|
205
273
|
```
|
|
206
274
|
|
|
207
275
|
five lines, four of which are imports and all four are stdlib. `time` and
|
|
208
276
|
`datetime` sit inside functions, which is why the pattern allows for indentation —
|
|
209
277
|
anchor it at `^import` and you would miss two, so do not take my word for the
|
|
210
|
-
anchor either. line
|
|
278
|
+
anchor either. line 1170 is the pattern catching a docstring that happens to begin
|
|
211
279
|
with the word `from`; it is prose, not an import, and it is left in rather than
|
|
212
280
|
tuned out, because a grep you tuned until it agreed with you proves nothing.
|
|
213
281
|
|
|
@@ -246,6 +314,7 @@ transcripto sessions recent sessions + the first prompt YOU typed in each
|
|
|
246
314
|
transcripto stats what you actually work on
|
|
247
315
|
transcripto cost what ONE of your decisions costs
|
|
248
316
|
transcripto coach which of YOUR prompt habits survive (a proxy)
|
|
317
|
+
transcripto export-run one run's numbers as JSON (typed turns, correction rate, commits)
|
|
249
318
|
```
|
|
250
319
|
|
|
251
320
|
`ask` is the one that kills "wait, did i lose something?". it answers "what was i
|
|
@@ -306,9 +375,12 @@ just gives the file a name on your PATH.
|
|
|
306
375
|
./test_cost.sh 12 assertions
|
|
307
376
|
./test_small_n.sh 7 assertions
|
|
308
377
|
./test_label_bands.sh 13 assertions
|
|
378
|
+
./test_correction.sh 32 assertions (correction rate + export-run, 2026-09-02)
|
|
379
|
+
./test_cursor_partial.sh 7 assertions
|
|
380
|
+
./test_version.sh 1 assertion (VERSION matches in transcripto.py and pyproject.toml)
|
|
309
381
|
```
|
|
310
382
|
|
|
311
|
-
|
|
383
|
+
91 assertions, all green, re-run 2026-09-02.
|
|
312
384
|
|
|
313
385
|
offline, no keys, on fixtures that inherit the real transcript shape including
|
|
314
386
|
all four ways a non-human record disguises itself as `type: user`.
|
|
@@ -25,7 +25,7 @@ PROG = _prog()
|
|
|
25
25
|
# packaging. A stranger who reads the README on GitHub and installs from PyPI can be
|
|
26
26
|
# holding a different build than the one the README describes, and until this flag
|
|
27
27
|
# existed there was no way for them to tell which.
|
|
28
|
-
VERSION = "0.1.
|
|
28
|
+
VERSION = "0.1.4"
|
|
29
29
|
|
|
30
30
|
USAGE = """
|
|
31
31
|
%(p)s index build / refresh the index (incremental)
|
|
@@ -38,6 +38,7 @@ USAGE = """
|
|
|
38
38
|
%(p)s coach which of YOUR prompt habits actually survive (a proxy)
|
|
39
39
|
%(p)s coach --harness codex grade your Codex (~/.codex) transcripts instead\n %(p)s coach --harness cursor grade your Cursor (~/.cursor) transcripts instead
|
|
40
40
|
%(p)s coach --verified-human subtract likely-PASTED turns (echoes of agent output)
|
|
41
|
+
%(p)s export-run <session|latest> one run's numbers as JSON: typed turns, correction rate, commits
|
|
41
42
|
""" % {"p": PROG}
|
|
42
43
|
|
|
43
44
|
HOME = os.path.expanduser("~")
|
|
@@ -147,7 +148,7 @@ def extract(d):
|
|
|
147
148
|
if bt == "text":
|
|
148
149
|
parts.append(b.get("text", ""))
|
|
149
150
|
elif bt == "tool_use":
|
|
150
|
-
name, inp = b.get("name", ""), (b
|
|
151
|
+
name, inp = b.get("name", ""), _tool_input(b)
|
|
151
152
|
fp = inp.get("file_path") or inp.get("path") or inp.get("notebook_path")
|
|
152
153
|
if fp:
|
|
153
154
|
files.append((FILE_TOOLS.get(name, name.lower()), fp))
|
|
@@ -880,6 +881,181 @@ def _looks_corrective(text):
|
|
|
880
881
|
return any(low.startswith(" " + m) or (" " + m) in low for m in _CORRECTIVE_MARKERS)
|
|
881
882
|
|
|
882
883
|
|
|
884
|
+
# ============================================================================
|
|
885
|
+
# correction rate — the inverse of a landed prompt
|
|
886
|
+
# ============================================================================
|
|
887
|
+
# Of the turns you TYPED, how many were spent telling the agent it got the last
|
|
888
|
+
# one wrong. The denominator is the same authorship gate every other number here
|
|
889
|
+
# uses (is_human_turn: typed OR queued, never isMeta / isSidechain / tool_result),
|
|
890
|
+
# so a tool result that happens to contain the word "wrong" cannot inflate it.
|
|
891
|
+
# The classifier is lexical and it is a FLOOR, stated so it can be argued with:
|
|
892
|
+
# a correction phrased without a marker ("the header one") is missed, and a
|
|
893
|
+
# genuine "no" inside a fresh request ("no tests needed") is counted. Read it as
|
|
894
|
+
# a trend on your own history, never as a verdict on one turn.
|
|
895
|
+
|
|
896
|
+
_CORRECTION_MARKERS = ("no,", "no ", "not that", "wrong", "i meant", "i said",
|
|
897
|
+
"again", "that's not", "you didn't", "revert", "undo",
|
|
898
|
+
"stop", "instead")
|
|
899
|
+
_CORRECTION_HEAD = 80 # markers are looked for in the first N characters only
|
|
900
|
+
_CORRECTION_SHORT = 12 # a turn under this many words, right after an agent
|
|
901
|
+
# turn that names the same file, is a nudge = correction
|
|
902
|
+
_CORRECTION_RE = re.compile(
|
|
903
|
+
"(?<![a-z0-9_])(?:" + "|".join(
|
|
904
|
+
re.escape(m).replace("'", "['\u2019]")
|
|
905
|
+
+ ("" if m[-1] in ", " else "(?![a-z0-9_])")
|
|
906
|
+
for m in _CORRECTION_MARKERS) + ")")
|
|
907
|
+
|
|
908
|
+
# ---------------------------------------------------------------------------
|
|
909
|
+
# v1 — the same question, rebuilt against a measured failure set.
|
|
910
|
+
#
|
|
911
|
+
# v0 was measured on 2026-09-03 over 200 agent-labelled turns (see
|
|
912
|
+
# docs/CORRECTION-PRECISION-2026-09-03.md): precision 0.54, recall 0.12, and a
|
|
913
|
+
# printed rate of 6% against a measured ~26%. The two biggest markers were the
|
|
914
|
+
# two worst: bare "no " ran at 56% precision because it is the commonest way to
|
|
915
|
+
# say something about the WORLD ("no worries", "no clue"), and bare "again" ran
|
|
916
|
+
# at 54% because resuming work ("lets pick this up again") is lexically identical
|
|
917
|
+
# to rejecting it. The NUDGE rule went 0-for-5.
|
|
918
|
+
#
|
|
919
|
+
# So v1 keeps the two words only where the grammar makes them about the agent.
|
|
920
|
+
#
|
|
921
|
+
# WHAT THIS STILL IS NOT: a classifier that understands the sentence. It is a
|
|
922
|
+
# better lexicon, tested on a held-out sample, and it will still miss a
|
|
923
|
+
# correction phrased without any of these words.
|
|
924
|
+
|
|
925
|
+
# "no" only where a pronoun, an imperative, or a comma makes it a rejection of
|
|
926
|
+
# what the agent just did, rather than a description of the world.
|
|
927
|
+
_V1_NO = (r"(?:^|[\s,;:])no[,!]"
|
|
928
|
+
r"|(?:^|[\s,;:])no\s+(?:you|we|i|it|that|thats|dont|don't|need\s+to|"
|
|
929
|
+
r"please|wait|new\s+work|not)"
|
|
930
|
+
r"|(?:^|[\s,;:])(?:thats|that's|its|it's)\s+not\b"
|
|
931
|
+
r"|(?:^|[\s,;:])no\b\s*$")
|
|
932
|
+
|
|
933
|
+
# "again" only next to a retry imperative or a complaint. "once again" alone is
|
|
934
|
+
# continuation; "try again" and "once again <negative>" are not.
|
|
935
|
+
_V1_AGAIN = (r"\btry(?:ing)?\s+again\b"
|
|
936
|
+
r"|\bagain\?"
|
|
937
|
+
r"|\b(?:once\s+)?again\b[^.!?]{0,40}?\b(?:wrong|not|no|never|"
|
|
938
|
+
r"broke|broken|fail|failed|mix(?:ed)?\s+up|missing|stuck|horrible|"
|
|
939
|
+
r"bad|lost|misunderstood|conflict)"
|
|
940
|
+
r"|\b(?:wrong|not|never|broke|broken|fail|failed|missing|stuck|"
|
|
941
|
+
r"horrible|misunderstood|conflict|mix(?:ed)?\s+up|lo[os]t|lose|"
|
|
942
|
+
r"loose|swamp(?:ed|ing)?)\b[^.!?]{0,40}?\b(?:once\s+)?again\b")
|
|
943
|
+
|
|
944
|
+
# The register v0 had no words for at all, measured from its 24 misses.
|
|
945
|
+
_V1_EXTRA = (r"\bfix\s+(?:it|this|that)\b"
|
|
946
|
+
r"|\bi\s+mean\b"
|
|
947
|
+
r"|\bdon'?t\s+like\b"
|
|
948
|
+
r"|\b(?:doesn'?t|does\s+not|isn'?t|is\s+not)\s+work"
|
|
949
|
+
r"|\bnot\s+working\b"
|
|
950
|
+
r"|\bnonsense\b|\bgibberish\b"
|
|
951
|
+
r"|\bshould\s?n'?t\b|\bshouldnt\b"
|
|
952
|
+
r"|\bmisunderstood\b|\bmisunderstanding\b"
|
|
953
|
+
r"|\byou\s+did\s?n'?t\b|\byou'?re\s+(?:wrong|lost|on\s+the\s+wrong)\b"
|
|
954
|
+
r"|\bwhere'?s?\s+the\b"
|
|
955
|
+
r"|\bis\s+this\s+(?:accurate|right|correct)\b"
|
|
956
|
+
r"|\bmakes\s+no\s+sense\b"
|
|
957
|
+
r"|\bmeans\s+nothing\b"
|
|
958
|
+
r"|\b(?:change|redo|rewrite|remove)\s+(?:it|this|that)\b")
|
|
959
|
+
|
|
960
|
+
_V1_RE = re.compile("(?:%s|%s|%s|%s)" % (
|
|
961
|
+
_V1_NO, _V1_AGAIN, _V1_EXTRA,
|
|
962
|
+
# the v0 markers that measured well enough to keep, unchanged
|
|
963
|
+
r"\bwrong\b|\bnot\s+that\b|\bi\s+meant\b|\brevert\b|\bundo\b"
|
|
964
|
+
r"|\bstop\b|\binstead\b"), re.I)
|
|
965
|
+
|
|
966
|
+
# A head taken in WORDS after URLs and paths are stripped. v0 read the first 80
|
|
967
|
+
# CHARACTERS raw, so a turn opening with a pasted URL spent its whole window on
|
|
968
|
+
# the URL and was unclassifiable — measured, not hypothesised: one sampled row's
|
|
969
|
+
# "once again ... change it" sat at roughly character 100.
|
|
970
|
+
_V1_HEAD_WORDS = 80
|
|
971
|
+
_URLISH = re.compile(r"(?:https?://|file://|/(?:Users|tmp|var|home)/)\S+|\S+\.(?:png|jpe?g|gif|pdf|mp4|mov)\b", re.I)
|
|
972
|
+
|
|
973
|
+
|
|
974
|
+
def _v1_head(text):
|
|
975
|
+
"""The first _V1_HEAD_WORDS words, with URLs and absolute paths removed."""
|
|
976
|
+
return " ".join(_URLISH.sub(" ", text).split()[:_V1_HEAD_WORDS]).lower()
|
|
977
|
+
|
|
978
|
+
|
|
979
|
+
# Which classifier runs. v1 is the default; v0 stays reachable so the two can be
|
|
980
|
+
# compared on one corpus. An explicit selector rather than a bare boolean, so a
|
|
981
|
+
# caller comparing versions cannot accidentally compare v1 to itself.
|
|
982
|
+
_CORRECTION_VERSION = os.environ.get("TRANSCRIPTO_CORRECTION", "v1")
|
|
983
|
+
if _CORRECTION_VERSION not in ("v0", "v1"):
|
|
984
|
+
_CORRECTION_VERSION = "v1"
|
|
985
|
+
|
|
986
|
+
|
|
987
|
+
def _quoted_things(text):
|
|
988
|
+
"""The file basenames and `backticked` spans a text names, lower-cased.
|
|
989
|
+
A path is reduced to its basename so "docs/caching.md" in your turn and
|
|
990
|
+
"/repo/docs/caching.md" in the agent's tool call read as the same file."""
|
|
991
|
+
out = set()
|
|
992
|
+
for m in _FILE_RE.finditer(text):
|
|
993
|
+
out.add(os.path.basename(m.group(0).rstrip(".,;:")).lower())
|
|
994
|
+
for m in re.finditer(r"`([^`\n]{2,80})`", text):
|
|
995
|
+
out.add(m.group(1).strip().lower())
|
|
996
|
+
return {t for t in out if len(t) >= 3}
|
|
997
|
+
|
|
998
|
+
|
|
999
|
+
def is_correction(text, prev_agent="", version=None):
|
|
1000
|
+
"""True iff a typed turn reads as a correction of the agent's previous turn.
|
|
1001
|
+
|
|
1002
|
+
PURE: text in, bool out, no state. Two rules, both lexical:
|
|
1003
|
+
|
|
1004
|
+
1. MARKER - the first 80 characters (case-insensitive) start with or
|
|
1005
|
+
contain one of _CORRECTION_MARKERS as a whole word: "no," "no " "not
|
|
1006
|
+
that" "wrong" "I meant" "I said" "again" "that's not" "you didn't"
|
|
1007
|
+
"revert" "undo" "stop" "instead". Whole-word, so "against" is not
|
|
1008
|
+
"again", "undone" is not "undo", "now" and "know" are not "no ". A
|
|
1009
|
+
bare "no" as the entire turn counts (the head is padded with a space).
|
|
1010
|
+
2. NUDGE - the turn is short (< 12 words) and the agent turn immediately
|
|
1011
|
+
before it, `prev_agent`, names the same file (matched by basename) or
|
|
1012
|
+
the same `backticked` span. "docs/caching.md, shorter please" right
|
|
1013
|
+
after the agent wrote docs/caching.md is a correction with no marker.
|
|
1014
|
+
|
|
1015
|
+
`prev_agent` is the extract()-rendered text of everything the assistant did
|
|
1016
|
+
since the previous typed turn (its prose plus its tool commands and file
|
|
1017
|
+
paths), or "" when this turn opened the session. The caller owns the
|
|
1018
|
+
authorship gate; this function never sees a record, only text.
|
|
1019
|
+
"""
|
|
1020
|
+
if (version or _CORRECTION_VERSION) == "v0":
|
|
1021
|
+
head = text[:_CORRECTION_HEAD].lower().strip() + " "
|
|
1022
|
+
if _CORRECTION_RE.search(head):
|
|
1023
|
+
return True
|
|
1024
|
+
# NUDGE. Kept in v0 exactly as it shipped; v1 drops it, because measured
|
|
1025
|
+
# over 100 flagged rows it fired alone five times and was wrong five
|
|
1026
|
+
# times. A rule with no true positive is not a loose rule, it is noise
|
|
1027
|
+
# with a cost — it renders every assistant record since the last turn.
|
|
1028
|
+
if prev_agent and len(text.split()) < _CORRECTION_SHORT:
|
|
1029
|
+
return bool(_quoted_things(text) & _quoted_things(prev_agent))
|
|
1030
|
+
return False
|
|
1031
|
+
return bool(_V1_RE.search(_v1_head(text)))
|
|
1032
|
+
|
|
1033
|
+
|
|
1034
|
+
def count_corrections(rows, pasted=None, version=None):
|
|
1035
|
+
"""(typed_turns, corrections) over one transcript, walked in order so each
|
|
1036
|
+
typed turn is classified against the agent turn it answers. Same gate and
|
|
1037
|
+
same paste subtraction as the episode grader, so the denominator returned
|
|
1038
|
+
here IS the `typed by you` count coach prints."""
|
|
1039
|
+
pasted = pasted or set()
|
|
1040
|
+
typed = corrections = 0
|
|
1041
|
+
agent = [] # assistant records since the last typed turn
|
|
1042
|
+
for row in rows:
|
|
1043
|
+
t = _human_prompt(row)
|
|
1044
|
+
if t:
|
|
1045
|
+
if t in pasted:
|
|
1046
|
+
continue # an echo of agent output, not a turn of yours
|
|
1047
|
+
typed += 1
|
|
1048
|
+
prev = ""
|
|
1049
|
+
if agent and len(t.split()) < _CORRECTION_SHORT:
|
|
1050
|
+
prev = "\n".join(extract(r)[1] for r in agent)
|
|
1051
|
+
if is_correction(t, prev, version=version):
|
|
1052
|
+
corrections += 1
|
|
1053
|
+
agent = []
|
|
1054
|
+
elif row.get("type") == "assistant":
|
|
1055
|
+
agent.append(row)
|
|
1056
|
+
return typed, corrections
|
|
1057
|
+
|
|
1058
|
+
|
|
883
1059
|
def _human_prompt(d):
|
|
884
1060
|
"""The text of a genuine human turn, or '' — reuses the measured gate."""
|
|
885
1061
|
if not is_human_turn(d):
|
|
@@ -894,6 +1070,17 @@ def _human_prompt(d):
|
|
|
894
1070
|
return ""
|
|
895
1071
|
|
|
896
1072
|
|
|
1073
|
+
def _tool_input(b):
|
|
1074
|
+
"""A tool_use block's arguments as a dict, or {} when it is not one.
|
|
1075
|
+
|
|
1076
|
+
Cursor persists an IN-FLIGHT tool call with its arguments still a raw JSON
|
|
1077
|
+
string prefix (`{"contents": "`), and one such record in a corpus was enough
|
|
1078
|
+
to abort `coach --harness cursor` with an AttributeError. A partial call has
|
|
1079
|
+
no arguments to read yet, so it reads as no arguments rather than as a crash."""
|
|
1080
|
+
inp = b.get("input")
|
|
1081
|
+
return inp if isinstance(inp, dict) else {}
|
|
1082
|
+
|
|
1083
|
+
|
|
897
1084
|
def _tool_uses(rec):
|
|
898
1085
|
if rec.get("type") != "assistant":
|
|
899
1086
|
return []
|
|
@@ -1004,26 +1191,32 @@ _CODEX_TOOL_TYPES = ("function_call", "custom_tool_call", "local_shell_call")
|
|
|
1004
1191
|
def _codex_rows(path):
|
|
1005
1192
|
"""Normalise one Codex rollout .jsonl into ordered Claude-shaped rows that
|
|
1006
1193
|
extract_episodes / _human_prompt / _tool_uses consume unchanged."""
|
|
1007
|
-
rows = []
|
|
1194
|
+
rows, sid, cwd = [], "", ""
|
|
1008
1195
|
for d in _iter_json(path):
|
|
1196
|
+
p = d.get("payload") or {}
|
|
1197
|
+
if d.get("type") == "session_meta":
|
|
1198
|
+
sid, cwd = p.get("id") or sid, p.get("cwd") or cwd
|
|
1199
|
+
continue
|
|
1009
1200
|
if d.get("type") != "response_item":
|
|
1010
1201
|
continue
|
|
1011
|
-
|
|
1202
|
+
# the envelope export-run reads; coach ignores it
|
|
1203
|
+
meta = {"timestamp": d.get("timestamp") or "", "cwd": cwd, "sessionId": sid}
|
|
1012
1204
|
pt = p.get("type")
|
|
1013
1205
|
if pt == "message":
|
|
1014
1206
|
role = p.get("role")
|
|
1015
1207
|
if role == "user":
|
|
1016
1208
|
txt = _codex_user_text(p.get("content"))
|
|
1017
1209
|
if txt:
|
|
1018
|
-
rows.append(
|
|
1019
|
-
|
|
1210
|
+
rows.append(dict(meta, type="user", promptSource="typed",
|
|
1211
|
+
message={"role": "user", "content": txt}))
|
|
1020
1212
|
elif role == "assistant":
|
|
1021
1213
|
txt = _codex_out_text(p.get("content"))
|
|
1022
|
-
rows.append(
|
|
1023
|
-
|
|
1214
|
+
rows.append(dict(meta, type="assistant", message={
|
|
1215
|
+
"role": "assistant",
|
|
1216
|
+
"content": [{"type": "text", "text": txt}] if txt else []}))
|
|
1024
1217
|
elif pt in _CODEX_TOOL_TYPES:
|
|
1025
|
-
rows.append(
|
|
1026
|
-
|
|
1218
|
+
rows.append(dict(meta, type="assistant", message={
|
|
1219
|
+
"role": "assistant", "content": [_codex_tool_block(p)]}))
|
|
1027
1220
|
return rows
|
|
1028
1221
|
|
|
1029
1222
|
|
|
@@ -1056,7 +1249,7 @@ def _cursor_ts(text):
|
|
|
1056
1249
|
|
|
1057
1250
|
|
|
1058
1251
|
def _cursor_cwd(path):
|
|
1059
|
-
"""~/.cursor/projects/Users-
|
|
1252
|
+
"""~/.cursor/projects/Users-dev-CODE-proj/... -> /Users/dev/CODE/proj.
|
|
1060
1253
|
The slug is lossy (a real hyphen in a directory name is indistinguishable
|
|
1061
1254
|
from the separator), so the result is checked on disk and dropped if it is
|
|
1062
1255
|
not a real directory. A wrong cwd is worse than no cwd."""
|
|
@@ -1316,11 +1509,11 @@ def extract_episodes(rows, source="", pasted=None):
|
|
|
1316
1509
|
if name in _WRITE_TOOLS:
|
|
1317
1510
|
wrote = True
|
|
1318
1511
|
if not witness:
|
|
1319
|
-
inp = b
|
|
1512
|
+
inp = _tool_input(b)
|
|
1320
1513
|
tgt = inp.get("file_path") or inp.get("notebook_path") or "?"
|
|
1321
1514
|
witness = "%s %s" % (name, os.path.basename(tgt))
|
|
1322
1515
|
elif name == "Bash":
|
|
1323
|
-
cmd = (b
|
|
1516
|
+
cmd = _tool_input(b).get("command", "") or ""
|
|
1324
1517
|
if _COMMIT_RE.search(cmd):
|
|
1325
1518
|
committed, witness = True, "git commit"
|
|
1326
1519
|
elif _REVERT_RE.search(cmd):
|
|
@@ -1435,7 +1628,7 @@ def coach(roots=None, harness=None, verified_human=False):
|
|
|
1435
1628
|
if not roots:
|
|
1436
1629
|
roots = _coach_roots(None, harness)
|
|
1437
1630
|
paths = _coach_files(roots, harness)
|
|
1438
|
-
episodes, records, humans, pastes = [], 0, 0, 0
|
|
1631
|
+
episodes, records, humans, pastes, corrections = [], 0, 0, 0, 0
|
|
1439
1632
|
human_texts, harnesses = [], set()
|
|
1440
1633
|
for p in paths:
|
|
1441
1634
|
rows, fh = _rows_for_file(p)
|
|
@@ -1450,8 +1643,10 @@ def coach(roots=None, harness=None, verified_human=False):
|
|
|
1450
1643
|
for row in rows:
|
|
1451
1644
|
t = _human_prompt(row)
|
|
1452
1645
|
if t and t not in pasted:
|
|
1453
|
-
humans += 1
|
|
1454
1646
|
human_texts.append(t)
|
|
1647
|
+
typed, corr = count_corrections(rows, pasted) # same gate as the line above
|
|
1648
|
+
humans += typed
|
|
1649
|
+
corrections += corr
|
|
1455
1650
|
episodes += extract_episodes(rows, source=p, pasted=pasted)
|
|
1456
1651
|
|
|
1457
1652
|
patterns = rank_patterns(episodes)
|
|
@@ -1471,6 +1666,10 @@ def coach(roots=None, harness=None, verified_human=False):
|
|
|
1471
1666
|
"history_lines": hist_lines, "history_matched": hist_matched,
|
|
1472
1667
|
"files": len(paths), "total_records": records, "human_turns": humans,
|
|
1473
1668
|
"human_pct": round(100 * humans / records, 2) if records else 0.0,
|
|
1669
|
+
# correction rate = typed turns that correct the agent / typed turns.
|
|
1670
|
+
# See is_correction() for the two rules; it is a lexical floor.
|
|
1671
|
+
"corrections": corrections,
|
|
1672
|
+
"correction_rate": round(corrections / humans, 3) if humans else 0.0,
|
|
1474
1673
|
"episodes": len(episodes), "durable": durable,
|
|
1475
1674
|
"durable_rate": round(durable / len(episodes), 3) if episodes else 0.0,
|
|
1476
1675
|
"tiers": tiers,
|
|
@@ -1616,6 +1815,27 @@ def cmd_coach(args):
|
|
|
1616
1815
|
t = r["tiers"]
|
|
1617
1816
|
print(" \033[2mcommit %s · write/edit %s · reverted %s · nothing durable %s\033[0m"
|
|
1618
1817
|
% (t["commit"], t["artifact"], t["reverted"], t["none"]))
|
|
1818
|
+
# correction rate: the inverse of a landed prompt. n travels with it because
|
|
1819
|
+
# a percentage without its denominator is the number this tool distrusts.
|
|
1820
|
+
#
|
|
1821
|
+
# It used to say "a floor" and stop. That was true and useless — it named a
|
|
1822
|
+
# direction without a size, so nobody could tell whether the real figure was
|
|
1823
|
+
# 7% or 40%. Both were measured on 2026-09-03 (200 rows each, agent-labelled,
|
|
1824
|
+
# one tuning sample and one held-out): v1 runs at precision 0.80 and recall
|
|
1825
|
+
# 0.16, so the printed count is roughly a FIFTH of the real one, and the two
|
|
1826
|
+
# samples put the true rate between 26% and 37%.
|
|
1827
|
+
#
|
|
1828
|
+
# The interval is printed rather than a point estimate because the two
|
|
1829
|
+
# samples disagree by more than sampling error, and hiding that behind one
|
|
1830
|
+
# confident number is the failure this tool exists to catch.
|
|
1831
|
+
# Each version states ITS OWN measured recall. Printing one figure for both
|
|
1832
|
+
# would attach v1's recall to v0's count — a true number about the wrong
|
|
1833
|
+
# object, which is the exact error this tool was built to surface.
|
|
1834
|
+
caught = {"v0": "~1 in 18", "v1": "~1 in 6"}.get(_CORRECTION_VERSION, "?")
|
|
1835
|
+
print(" \033[2mcorrection rate: %s%% measured (%s of %s typed turns) · "
|
|
1836
|
+
"%s catches %s, so the real rate is ~26-37%%\033[0m"
|
|
1837
|
+
% (int(round(r["correction_rate"] * 100)), r["corrections"],
|
|
1838
|
+
r["human_turns"], _CORRECTION_VERSION, caught))
|
|
1619
1839
|
if r["verified_human"]:
|
|
1620
1840
|
print(" \033[2m%s typed turn(s) flagged likely-PASTED and subtracted\033[0m"
|
|
1621
1841
|
% r["pastes_flagged"])
|
|
@@ -1628,6 +1848,171 @@ def cmd_coach(args):
|
|
|
1628
1848
|
print()
|
|
1629
1849
|
|
|
1630
1850
|
|
|
1851
|
+
# ============================================================================
|
|
1852
|
+
# export-run — one run's numbers as JSON, the contract Agent Grinder and ZUP read
|
|
1853
|
+
# ============================================================================
|
|
1854
|
+
# Reads the transcript file directly (no index needed), applies the same
|
|
1855
|
+
# authorship gate and the same correction classifier coach uses, and adds the
|
|
1856
|
+
# run envelope: when it started and ended, what the agent touched, and what got
|
|
1857
|
+
# committed in that window. Every key is documented in README.md > export-run.
|
|
1858
|
+
|
|
1859
|
+
|
|
1860
|
+
def _epoch(ts):
|
|
1861
|
+
"""ISO-8601 transcript stamp -> unix seconds, or None. Accepts the 'Z' Claude
|
|
1862
|
+
Code and Codex write and the naive local stamp the Cursor connector yields."""
|
|
1863
|
+
if not ts:
|
|
1864
|
+
return None
|
|
1865
|
+
s = ts.strip()
|
|
1866
|
+
if s.endswith("Z"):
|
|
1867
|
+
s = s[:-1] + "+00:00"
|
|
1868
|
+
try:
|
|
1869
|
+
d = datetime.fromisoformat(s)
|
|
1870
|
+
except ValueError:
|
|
1871
|
+
return None
|
|
1872
|
+
if d.tzinfo is None:
|
|
1873
|
+
d = d.astimezone() # naive = local wall clock
|
|
1874
|
+
return d.timestamp()
|
|
1875
|
+
|
|
1876
|
+
|
|
1877
|
+
def _iso(epoch):
|
|
1878
|
+
return datetime.fromtimestamp(epoch, timezone.utc).isoformat(
|
|
1879
|
+
timespec="seconds").replace("+00:00", "Z")
|
|
1880
|
+
|
|
1881
|
+
|
|
1882
|
+
def _git_dir(cwd):
|
|
1883
|
+
"""The .git directory governing `cwd` (walking up), resolving a `.git` FILE
|
|
1884
|
+
(worktree / submodule) to the gitdir it points at. None if not in a repo."""
|
|
1885
|
+
d = os.path.abspath(os.path.expanduser(cwd or ""))
|
|
1886
|
+
while True:
|
|
1887
|
+
g = os.path.join(d, ".git")
|
|
1888
|
+
if os.path.isdir(g):
|
|
1889
|
+
return g
|
|
1890
|
+
if os.path.isfile(g):
|
|
1891
|
+
try:
|
|
1892
|
+
first = open(g).readline().strip()
|
|
1893
|
+
except OSError:
|
|
1894
|
+
return None
|
|
1895
|
+
if first.startswith("gitdir:"):
|
|
1896
|
+
target = first[len("gitdir:"):].strip()
|
|
1897
|
+
return os.path.normpath(os.path.join(d, target))
|
|
1898
|
+
return None
|
|
1899
|
+
parent = os.path.dirname(d)
|
|
1900
|
+
if parent == d:
|
|
1901
|
+
return None
|
|
1902
|
+
d = parent
|
|
1903
|
+
|
|
1904
|
+
|
|
1905
|
+
def _reflog_commits(cwd, start=None, end=None):
|
|
1906
|
+
"""Commits recorded in cwd's reflog (.git/logs/HEAD) stamped inside
|
|
1907
|
+
[start, end] (unix seconds, either side open when None), as a list of
|
|
1908
|
+
{sha, ts, subject}; None when cwd is not inside a git repo.
|
|
1909
|
+
|
|
1910
|
+
Read straight from the file, not from `git log`: the README's privacy claim
|
|
1911
|
+
is that nothing here shells out, and this is the one place that would have
|
|
1912
|
+
crept in. The reflog is the record of what THIS working tree did: a commit made
|
|
1913
|
+
here lands in it, a commit pulled in from elsewhere does not, which is the
|
|
1914
|
+
right side of the line for a run receipt. It is local and git expires it
|
|
1915
|
+
(90 days by default), so a run older than that can read 0 here while
|
|
1916
|
+
`git log` would still show its commits. `commit (amend)` counts as a
|
|
1917
|
+
commit; a rebase's `pick` lines do not."""
|
|
1918
|
+
g = _git_dir(cwd)
|
|
1919
|
+
if not g:
|
|
1920
|
+
return None
|
|
1921
|
+
out = []
|
|
1922
|
+
try:
|
|
1923
|
+
f = open(os.path.join(g, "logs", "HEAD"), "r", errors="replace")
|
|
1924
|
+
except OSError:
|
|
1925
|
+
return out
|
|
1926
|
+
with f:
|
|
1927
|
+
for line in f:
|
|
1928
|
+
head, _, msg = line.rstrip("\n").partition("\t")
|
|
1929
|
+
parts = head.split()
|
|
1930
|
+
if len(parts) < 4 or not msg.startswith("commit"):
|
|
1931
|
+
continue
|
|
1932
|
+
try:
|
|
1933
|
+
stamp = int(parts[-2])
|
|
1934
|
+
except ValueError:
|
|
1935
|
+
continue
|
|
1936
|
+
if (start is not None and stamp < start) or (end is not None and stamp > end):
|
|
1937
|
+
continue
|
|
1938
|
+
out.append({"sha": parts[1][:7], "ts": _iso(stamp),
|
|
1939
|
+
"subject": msg.partition(": ")[2]})
|
|
1940
|
+
return out
|
|
1941
|
+
|
|
1942
|
+
|
|
1943
|
+
def _resolve_run(target, roots, harness):
|
|
1944
|
+
"""The transcript file `target` names: 'latest' (newest by mtime; a spawned
|
|
1945
|
+
sub-agent's file is not a run of yours and is skipped), a session id or a
|
|
1946
|
+
prefix of one, or a path to a .jsonl. None when nothing matches."""
|
|
1947
|
+
if target != "latest" and os.path.isfile(target):
|
|
1948
|
+
return target
|
|
1949
|
+
paths = [p for p in _coach_files(roots, harness)
|
|
1950
|
+
if os.sep + "subagents" + os.sep not in p
|
|
1951
|
+
and os.path.basename(p) not in ("history.jsonl", "session_index.jsonl")]
|
|
1952
|
+
if not paths:
|
|
1953
|
+
return None
|
|
1954
|
+
if target == "latest":
|
|
1955
|
+
return max(paths, key=os.path.getmtime)
|
|
1956
|
+
hits = [p for p in paths if os.path.basename(p).startswith(target)]
|
|
1957
|
+
if not hits: # codex names rollouts rollout-<date>-<id>.jsonl
|
|
1958
|
+
hits = [p for p in paths if target in os.path.basename(p)]
|
|
1959
|
+
return max(hits, key=os.path.getmtime) if hits else None
|
|
1960
|
+
|
|
1961
|
+
|
|
1962
|
+
def export_run(path):
|
|
1963
|
+
"""The run-level numbers of one transcript file, as a JSON-ready dict."""
|
|
1964
|
+
rows, harness = _rows_for_file(path)
|
|
1965
|
+
sid = next((r.get("sessionId") for r in rows if r.get("sessionId")), "") \
|
|
1966
|
+
or os.path.splitext(os.path.basename(path))[0]
|
|
1967
|
+
cwd = next((r.get("cwd") for r in reversed(rows) if r.get("cwd")), "")
|
|
1968
|
+
stamps = [e for e in (_epoch(r.get("timestamp")) for r in rows) if e is not None]
|
|
1969
|
+
start, end = (min(stamps), max(stamps)) if stamps else (None, None)
|
|
1970
|
+
typed, corrections = count_corrections(rows)
|
|
1971
|
+
tool_calls, files = 0, set()
|
|
1972
|
+
for r in rows:
|
|
1973
|
+
for b in _tool_uses(r):
|
|
1974
|
+
tool_calls += 1
|
|
1975
|
+
if b.get("name") in FILE_TOOLS:
|
|
1976
|
+
fp = _tool_input(b).get("file_path") or _tool_input(b).get("notebook_path")
|
|
1977
|
+
if fp:
|
|
1978
|
+
files.add(fp)
|
|
1979
|
+
commits = _reflog_commits(cwd, start, end) if cwd else None
|
|
1980
|
+
return {
|
|
1981
|
+
"schema": "transcripto.export-run/1",
|
|
1982
|
+
"session_id": sid,
|
|
1983
|
+
"project": cwd,
|
|
1984
|
+
"harness": harness,
|
|
1985
|
+
"transcript": path,
|
|
1986
|
+
"started": _iso(start) if start is not None else None,
|
|
1987
|
+
"ended": _iso(end) if end is not None else None,
|
|
1988
|
+
"duration_s": int(round(end - start)) if stamps else None,
|
|
1989
|
+
"records": len(rows),
|
|
1990
|
+
"typed_turns": typed,
|
|
1991
|
+
"corrections": corrections,
|
|
1992
|
+
"correction_rate": round(corrections / typed, 3) if typed else None,
|
|
1993
|
+
"tool_calls": tool_calls,
|
|
1994
|
+
"files_touched": sorted(files),
|
|
1995
|
+
"commits_in_window": len(commits) if commits is not None else None,
|
|
1996
|
+
"commits": commits,
|
|
1997
|
+
"proxy": ("typed_turns is the authorship gate (typed/queued, never meta, "
|
|
1998
|
+
"sidechain or tool output); correction_rate is a lexical floor; "
|
|
1999
|
+
"commits_in_window is this tree's reflog inside the run window, "
|
|
2000
|
+
"null when the project is not a git repo."),
|
|
2001
|
+
}
|
|
2002
|
+
|
|
2003
|
+
|
|
2004
|
+
def cmd_export_run(args):
|
|
2005
|
+
roots = _coach_roots(args.root, args.harness)
|
|
2006
|
+
path = _resolve_run(args.target, roots, args.harness)
|
|
2007
|
+
if not path:
|
|
2008
|
+
sys.stderr.write("\n no transcript matches '%s'.\n" % args.target)
|
|
2009
|
+
sys.stderr.write(" looked in: %s\n" % ", ".join(roots))
|
|
2010
|
+
sys.stderr.write(" pass a session id (or a prefix of one), 'latest', or a path "
|
|
2011
|
+
"to a .jsonl; --root / --harness as for coach.\n\n")
|
|
2012
|
+
sys.exit(2)
|
|
2013
|
+
print(json.dumps(export_run(path), indent=2))
|
|
2014
|
+
|
|
2015
|
+
|
|
1631
2016
|
def main():
|
|
1632
2017
|
p = argparse.ArgumentParser(prog=PROG, description=__doc__ + USAGE,
|
|
1633
2018
|
formatter_class=argparse.RawDescriptionHelpFormatter)
|
|
@@ -1660,6 +2045,13 @@ def main():
|
|
|
1660
2045
|
"in the same session, from the human signal before grading")
|
|
1661
2046
|
s.add_argument("--json", action="store_true", help="machine-readable, for other tools")
|
|
1662
2047
|
s.set_defaults(fn=cmd_coach)
|
|
2048
|
+
s = sub.add_parser("export-run")
|
|
2049
|
+
s.add_argument("target", help="a session id (or a prefix of one), 'latest', or a "
|
|
2050
|
+
"path to a .jsonl transcript")
|
|
2051
|
+
s.add_argument("--root", help="look in this transcript dir instead of ~/.claude/projects")
|
|
2052
|
+
s.add_argument("--harness", choices=["claude", "codex", "cursor"],
|
|
2053
|
+
help="which agent's transcripts to look in. default: claude")
|
|
2054
|
+
s.set_defaults(fn=cmd_export_run)
|
|
1663
2055
|
a = p.parse_args()
|
|
1664
2056
|
if not getattr(a, "fn", None):
|
|
1665
2057
|
p.print_help(); return
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|