transcripto 0.2.0__tar.gz → 0.3.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- transcripto-0.3.0/PKG-INFO +490 -0
- transcripto-0.3.0/README.md +472 -0
- {transcripto-0.2.0 → transcripto-0.3.0}/pyproject.toml +3 -3
- transcripto-0.3.0/tests/test_harness_backfill.py +203 -0
- transcripto-0.3.0/tests/test_jev.py +423 -0
- transcripto-0.3.0/tests/test_jev_findings.py +119 -0
- transcripto-0.3.0/tests/test_jev_invalid_responses.py +55 -0
- transcripto-0.3.0/tests/test_jev_preview.py +77 -0
- transcripto-0.3.0/tests/test_public_flow.py +351 -0
- {transcripto-0.2.0 → transcripto-0.3.0}/tests/test_replay.py +117 -0
- transcripto-0.3.0/tests/test_selected_context.py +62 -0
- transcripto-0.3.0/transcripto.egg-info/PKG-INFO +490 -0
- transcripto-0.3.0/transcripto.egg-info/SOURCES.txt +22 -0
- transcripto-0.3.0/transcripto.egg-info/top_level.txt +6 -0
- {transcripto-0.2.0 → transcripto-0.3.0}/transcripto.py +1094 -87
- {transcripto-0.2.0 → transcripto-0.3.0}/transcripto_core.py +5 -1
- transcripto-0.3.0/transcripto_findings.py +87 -0
- transcripto-0.3.0/transcripto_jev.py +347 -0
- {transcripto-0.2.0 → transcripto-0.3.0}/transcripto_replay.py +53 -8
- transcripto-0.3.0/transcripto_selected.py +118 -0
- transcripto-0.2.0/PKG-INFO +0 -222
- transcripto-0.2.0/README.md +0 -204
- transcripto-0.2.0/transcripto.egg-info/PKG-INFO +0 -222
- transcripto-0.2.0/transcripto.egg-info/SOURCES.txt +0 -12
- transcripto-0.2.0/transcripto.egg-info/top_level.txt +0 -3
- {transcripto-0.2.0 → transcripto-0.3.0}/LICENSE +0 -0
- {transcripto-0.2.0 → transcripto-0.3.0}/setup.cfg +0 -0
- {transcripto-0.2.0 → transcripto-0.3.0}/transcripto.egg-info/dependency_links.txt +0 -0
- {transcripto-0.2.0 → transcripto-0.3.0}/transcripto.egg-info/entry_points.txt +0 -0
|
@@ -0,0 +1,490 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: transcripto
|
|
3
|
+
Version: 0.3.0
|
|
4
|
+
Summary: Instant replay for coding agents. Inspect requests, tool calls, and recorded results across Claude Code, Codex, and Cursor. Local by default, stdlib-only.
|
|
5
|
+
Author: Oscar Morke
|
|
6
|
+
License: MIT
|
|
7
|
+
Project-URL: Homepage, https://github.com/Morkeeth/transcripto
|
|
8
|
+
Project-URL: Source, https://github.com/Morkeeth/transcripto
|
|
9
|
+
Keywords: claude-code,coding-agents,transcripts,local-first,analytics
|
|
10
|
+
Classifier: Environment :: Console
|
|
11
|
+
Classifier: Programming Language :: Python :: 3
|
|
12
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
13
|
+
Classifier: Topic :: Utilities
|
|
14
|
+
Requires-Python: >=3.9
|
|
15
|
+
Description-Content-Type: text/markdown
|
|
16
|
+
License-File: LICENSE
|
|
17
|
+
Dynamic: license-file
|
|
18
|
+
|
|
19
|
+
# Transcripto
|
|
20
|
+
|
|
21
|
+
**You might already be keeping a journal. Read your side of it.**
|
|
22
|
+
|
|
23
|
+
Your agent transcripts contain what you asked for, what you changed your mind
|
|
24
|
+
about, and what you kept coming back to. Transcripto helps you find those words
|
|
25
|
+
and read the recorded work around them.
|
|
26
|
+
|
|
27
|
+
Claude Code · Codex · Cursor. Local files. No account. No runtime dependencies.
|
|
28
|
+
|
|
29
|
+
## Start with something you remember saying
|
|
30
|
+
|
|
31
|
+
```sh
|
|
32
|
+
uvx --from transcripto==0.3.0 transcripto ask "retry"
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
Replace `retry` with a word you remember using. `ask` searches messages identified
|
|
36
|
+
as yours and shows dated snippets, newest first. It refreshes the local index
|
|
37
|
+
automatically. The first search indexes the selected history; a large archive
|
|
38
|
+
can take minutes. Add `--harness claude`, `--harness codex`, or `--harness cursor`
|
|
39
|
+
to limit that scan. It does not generate a diary or interpret your personality.
|
|
40
|
+
|
|
41
|
+
Each hit prints an `Open:` command. Run that command to open the exact request
|
|
42
|
+
and its recorded work. This also works when search matches a word variant
|
|
43
|
+
(such as `retry` matching `retried`) or several requests share the same words.
|
|
44
|
+
|
|
45
|
+
You can also search replay directly:
|
|
46
|
+
|
|
47
|
+
```sh
|
|
48
|
+
uvx --from transcripto==0.3.0 transcripto replay "retry"
|
|
49
|
+
|
|
50
|
+
# Or open your latest human session:
|
|
51
|
+
uvx --from transcripto==0.3.0 transcripto
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
Replay puts your request, tool calls and recorded results in order. Failed edits
|
|
55
|
+
stay failed. Missing results stay unknown. Status describes tool execution,
|
|
56
|
+
not whether the task was done correctly.
|
|
57
|
+
|
|
58
|
+
Or install with `python3 -m pip install transcripto==0.3.0`, then run
|
|
59
|
+
`transcripto ask "retry"`. Requires Python 3.9 or newer.
|
|
60
|
+
|
|
61
|
+
## Try the stranger flow without your transcripts
|
|
62
|
+
|
|
63
|
+
The bundled public example is synthetic. It works in an isolated home and does
|
|
64
|
+
not depend on agent dotfiles:
|
|
65
|
+
|
|
66
|
+
```sh
|
|
67
|
+
INSTALL="$(mktemp -d)"
|
|
68
|
+
python3 -m pip install --no-deps --no-build-isolation --target "$INSTALL" .
|
|
69
|
+
export HOME="$(mktemp -d)"
|
|
70
|
+
transcripto() { PYTHONPATH="$INSTALL" python3 -m transcripto "$@"; }
|
|
71
|
+
|
|
72
|
+
transcripto import-example
|
|
73
|
+
transcripto ask "What changed about the forecast cache?"
|
|
74
|
+
transcripto changes
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
`ask` cites the imported JSONL line for every hit. `changes` is a focused view
|
|
78
|
+
of the request that was revised, the correction, and its recorded follow-up.
|
|
79
|
+
It labels missing results rather than turning a change of mind into a score.
|
|
80
|
+
|
|
81
|
+
To carry that correction to a different receiver:
|
|
82
|
+
|
|
83
|
+
```sh
|
|
84
|
+
transcripto handoff "30 seconds" \
|
|
85
|
+
--to-harness codex --output "$HOME/codex-inbox/correction.json"
|
|
86
|
+
transcripto receive-handoff \
|
|
87
|
+
"$HOME/codex-inbox/correction.json" --as-harness codex \
|
|
88
|
+
--output "$HOME/codex-work/receiver-brief.md"
|
|
89
|
+
cat "$HOME/codex-work/receiver-brief.md"
|
|
90
|
+
```
|
|
91
|
+
|
|
92
|
+
The brief includes the cited correction, recorded follow-up statuses
|
|
93
|
+
(failed / succeeded / unknown), and an `Open:` command for the exact request.
|
|
94
|
+
If the source moved, disappeared, or no longer holds the cited request, the brief
|
|
95
|
+
marks evidence uncertain and shows only the packet's own statuses as provisional. It does not invoke a receiver agent or prove
|
|
96
|
+
adoption. Synthetic provenance stays visible in search, changes, and handoffs.
|
|
97
|
+
Handoff files are local and mode `0600`; they can contain transcript text and
|
|
98
|
+
paths, so review them before sharing.
|
|
99
|
+
|
|
100
|
+
### Cross-harness lab (failed / succeeded / unknown)
|
|
101
|
+
|
|
102
|
+
For a receiving agent that needs to find a prior episode and reopen exact
|
|
103
|
+
evidence without private history:
|
|
104
|
+
|
|
105
|
+
```sh
|
|
106
|
+
transcripto import-lab
|
|
107
|
+
transcripto ask "retry"
|
|
108
|
+
# run each printed Open: command
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
All lab records are labelled synthetic. Claude shows a failed edit, Codex a
|
|
112
|
+
succeeded check, Cursor an unknown missing result.
|
|
113
|
+
|
|
114
|
+
### Offline flight card
|
|
115
|
+
|
|
116
|
+
```sh
|
|
117
|
+
transcripto quickstart --wheel /absolute/path/to/transcripto-0.3.0-py3-none-any.whl
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
Prints install, `import-lab`, search, and reopen commands for a built wheel
|
|
121
|
+
without PyPI. See also `docs/OFFLINE-QUICKSTART.md`.
|
|
122
|
+
|
|
123
|
+
**Your files remain yours.** Transcripto does not upload transcript content or
|
|
124
|
+
execute commands found in it. Search output, replay and JSON can contain private
|
|
125
|
+
words and paths; review anything you choose to share. It reads existing files,
|
|
126
|
+
not deleted history. Check your agent's retention settings and keep your own
|
|
127
|
+
backup if you want a lasting record. Authorship detection differs by harness;
|
|
128
|
+
[see the limits below](#what-each-harness-supports).
|
|
129
|
+
|
|
130
|
+
## Find the thing you remember
|
|
131
|
+
|
|
132
|
+
Search automatically refreshes a local index. No setup command is required.
|
|
133
|
+
|
|
134
|
+
```sh
|
|
135
|
+
transcripto ask "retry" # your submitted words
|
|
136
|
+
transcripto search "retry" # prompts, replies, and tool text
|
|
137
|
+
transcripto find parser.py # recorded file operations and attempts
|
|
138
|
+
transcripto trace "retry" # an alias into result-aware replay
|
|
139
|
+
transcripto sessions # sessions with submitted prompts
|
|
140
|
+
transcripto stats # activity counts
|
|
141
|
+
```
|
|
142
|
+
|
|
143
|
+
A failed or unconfirmed file change is labelled an **attempt**, never `WROTE`.
|
|
144
|
+
Queries with `--harness` or `--root` are scoped to that selection even if the
|
|
145
|
+
index already contains another corpus. `index` and `watch` remain available for
|
|
146
|
+
explicit refresh and background polling.
|
|
147
|
+
|
|
148
|
+
## The replay
|
|
149
|
+
|
|
150
|
+
This is output from `replay --demo`. **All prompts and results in this example
|
|
151
|
+
are invented.** The demo goes through the same parser as a real transcript.
|
|
152
|
+
|
|
153
|
+
```text
|
|
154
|
+
THE COMEBACK · claude · request 1
|
|
155
|
+
You asked: "Fix the login redirect and run its tests."
|
|
156
|
+
|
|
157
|
+
1 FAIL edit src/login.py
|
|
158
|
+
Tool error: Error: text not found [call L2 → result L3]
|
|
159
|
+
2 OK edit src/login.py
|
|
160
|
+
Tool reported success. [call L4 → result L5]
|
|
161
|
+
3 FAIL check pytest tests/test_login.py
|
|
162
|
+
Tool error: Process exited with code 1 [call L6 → result L7]
|
|
163
|
+
4 OK edit src/login.py
|
|
164
|
+
Tool reported success. [call L8 → result L9]
|
|
165
|
+
5 OK check pytest tests/test_login.py
|
|
166
|
+
Tool reported success. [call L10 → result L11]
|
|
167
|
+
|
|
168
|
+
Agent said: "The redirect is fixed and the tests pass."
|
|
169
|
+
|
|
170
|
+
Recorded: 3 succeeded · 2 failed · 0 unknown
|
|
171
|
+
Status describes a tool result, not task correctness. Missing results stay unknown.
|
|
172
|
+
```
|
|
173
|
+
|
|
174
|
+
The headings have rules. **The comeback** means a recorded failure was followed
|
|
175
|
+
by success for the same operation and target, without a later failed or unknown
|
|
176
|
+
attempt on that target. **The snag** means a failure is
|
|
177
|
+
present. **The missing receipt** means a result is unknown. **The answer** means
|
|
178
|
+
there were no recorded tool calls; an explanation may have been the whole task.
|
|
179
|
+
These are descriptions of the sequence, not grades for you or your agent.
|
|
180
|
+
|
|
181
|
+
On your own history, each replay names its source file and line numbers, plus a
|
|
182
|
+
command that reopens that exact request. Long sequences open around the first
|
|
183
|
+
failure or change and tell you what was omitted. `--all` shows the full sequence.
|
|
184
|
+
|
|
185
|
+
```sh
|
|
186
|
+
transcripto # latest session you submitted a request in
|
|
187
|
+
transcripto replay --failures # most recent request with a recorded failure
|
|
188
|
+
transcripto replay "login redirect" # find requests containing these words
|
|
189
|
+
transcripto replay path/to/session.jsonl # inspect one transcript
|
|
190
|
+
transcripto replay --session 3f9c1a2b # explicitly select a session prefix
|
|
191
|
+
transcripto replay path/to/session.jsonl --episode 3 --all
|
|
192
|
+
transcripto replay path/to/session.jsonl --line 42 # exact request from an ask hit
|
|
193
|
+
transcripto replay latest --json # structured events, evidence, source lines
|
|
194
|
+
transcripto replay latest --share # counts + caveat; no prompts or paths
|
|
195
|
+
```
|
|
196
|
+
|
|
197
|
+
`--share` is intentionally small. Full replay output and JSON contain your own
|
|
198
|
+
words and local paths. Replay does not upload either.
|
|
199
|
+
|
|
200
|
+
## What each harness supports
|
|
201
|
+
|
|
202
|
+
| Feature | Claude Code | Codex | Cursor |
|
|
203
|
+
|---|---|---|---|
|
|
204
|
+
| Replay, search, ask, find, trace, sessions, stats | Yes | Yes | Yes |
|
|
205
|
+
| Tool attempts | Native tool calls | Direct calls and supported static wrappers | `StrReplace`, `Shell`, `Write`, and other calls |
|
|
206
|
+
| Execution status | Matched results | Matched results; ambiguous wrappers stay unknown | Unknown when the export omits results or call IDs |
|
|
207
|
+
| Authorship | `promptSource` typed/queued, excluding injected/tool records | User messages with known injected context excluded | `<user_query>` wrapper; a weaker signal |
|
|
208
|
+
| Coach, export-run | Yes | Yes | Yes, with missing evidence preserved |
|
|
209
|
+
| API-equivalent cost | Yes | Not supported | Not supported |
|
|
210
|
+
|
|
211
|
+
```sh
|
|
212
|
+
transcripto replay --harness claude
|
|
213
|
+
transcripto replay --harness codex
|
|
214
|
+
transcripto replay --harness cursor
|
|
215
|
+
transcripto search "retry" --harness codex
|
|
216
|
+
transcripto replay --root /path/to/transcripts
|
|
217
|
+
```
|
|
218
|
+
|
|
219
|
+
Default roots are `~/.claude/projects`, `~/.codex`, and `~/.cursor`.
|
|
220
|
+
Codex reads sessions and archived sessions. Cursor reads the per-session files
|
|
221
|
+
under `projects/*/agent-transcripts/*/`. Latest-session replay skips subagents
|
|
222
|
+
and files without a submitted human request.
|
|
223
|
+
|
|
224
|
+
Cursor exports often contain calls without results. That is useful evidence
|
|
225
|
+
of an attempt, but not enough to claim success. Transcripto does not substitute
|
|
226
|
+
an assistant's closing message or a `turn_ended` record for the missing result.
|
|
227
|
+
|
|
228
|
+
## The evidence contract
|
|
229
|
+
|
|
230
|
+
1. A tool call is an **attempt**.
|
|
231
|
+
2. A matching result may establish **succeeded** or **failed** execution.
|
|
232
|
+
3. A missing result, a running command, or an ambiguous result is **unknown**.
|
|
233
|
+
4. An exit code of zero is not proof that the requested task is correct.
|
|
234
|
+
5. A later human request opens a new episode. Its work is never absorbed into
|
|
235
|
+
the previous request because the words happen to overlap.
|
|
236
|
+
6. A command mentioning `git commit` is not necessarily a commit. Quoted text,
|
|
237
|
+
dry runs, and compound shell commands are not promoted to commit evidence.
|
|
238
|
+
|
|
239
|
+
The tool never executes transcript commands. It parses a limited set of static
|
|
240
|
+
Codex wrapper forms; arbitrary JavaScript and multiple nested child calls are
|
|
241
|
+
not reconstructed. A long-running call can remain unknown when completion is
|
|
242
|
+
only present in a later polling call. Cross-session durability, semantic task
|
|
243
|
+
completion, and live repository state are not inferred from transcript text.
|
|
244
|
+
|
|
245
|
+
## Coach without invented grades
|
|
246
|
+
|
|
247
|
+
`transcripto coach` shows descriptive request history. It no longer recommends
|
|
248
|
+
prompt habits, labels a no-edit answer a bad prompt, or applies one person's
|
|
249
|
+
correction-rate calibration to someone else's data.
|
|
250
|
+
|
|
251
|
+
Habit proportions include **change attempts with known outcomes**. Unknown
|
|
252
|
+
outcomes and read-only tasks are excluded. The groups overlap and the requests
|
|
253
|
+
can be correlated, so these proportions are not significance tests or causal
|
|
254
|
+
advice. No best/worst ranking is printed. Correction markers are a lexical
|
|
255
|
+
estimate with false positives and misses, not a guaranteed lower bound.
|
|
256
|
+
|
|
257
|
+
Coach JSON is marked `transcripto.coach/2`. Legacy `durable`/`survived` fields
|
|
258
|
+
refer only to observed successful change results, not lasting work. Unknown
|
|
259
|
+
request outcomes have `survived: null`. `durable_rate` uses only known change
|
|
260
|
+
requests as its denominator and is null when there are none. `best_prompt` and `worst_prompt` are
|
|
261
|
+
retained as null compatibility fields. Use `successful_request`,
|
|
262
|
+
`failed_request`, and replay's event status to inspect evidence.
|
|
263
|
+
|
|
264
|
+
`export-run latest` always prints JSON. Its
|
|
265
|
+
existing `transcripto.export-run/1` keys remain available. `records` counts
|
|
266
|
+
normalized message records. `files_touched` lists attempted file targets,
|
|
267
|
+
including reads; it is not a successful-change count. Reflog commits are local
|
|
268
|
+
working-tree events inside the available timestamp window, not proof that this
|
|
269
|
+
agent caused them. Without a usable window, the commit fields are null.
|
|
270
|
+
|
|
271
|
+
## Privacy and limits
|
|
272
|
+
|
|
273
|
+
The default commands process transcripts locally, without telemetry or an
|
|
274
|
+
account flow. The optional Jev detector below sends filtered typed text only
|
|
275
|
+
when explicitly selected. Package installation (`pip` or `uvx`) is a separate
|
|
276
|
+
operation that may contact a package registry and write a package cache.
|
|
277
|
+
|
|
278
|
+
### Optional network detector: `--detector jev`
|
|
279
|
+
|
|
280
|
+
`coach` and `export-run` can count corrections with TypeSafe Jev instead of the
|
|
281
|
+
local regex. This is the one path that sends text off the machine, and it runs
|
|
282
|
+
only when you pass the flag on that run. No environment variable or config file
|
|
283
|
+
turns it on. The code lives in its own module, `transcripto_jev.py`, which the
|
|
284
|
+
default path never imports.
|
|
285
|
+
|
|
286
|
+
```sh
|
|
287
|
+
OPENROUTER_API_KEY=... transcripto coach --detector jev
|
|
288
|
+
```
|
|
289
|
+
|
|
290
|
+
Before the first request it prints one line to stderr: how many typed turns it
|
|
291
|
+
may send, how many the privacy filter excluded, and the URL
|
|
292
|
+
(`https://openrouter.ai/api/alpha/decisions`, model `typesafe/jev-1.13`).
|
|
293
|
+
Only your typed turns are sent, at most 2,000 characters each, with the fixed
|
|
294
|
+
question. No separate path/session metadata, agent output or tool results are
|
|
295
|
+
sent. Typed text can still contain paths and private details the filter misses.
|
|
296
|
+
What OpenRouter and the model provider keep, and for how long, is set by their
|
|
297
|
+
terms. Transcripto does not verify it. Read those terms before the first send.
|
|
298
|
+
|
|
299
|
+
The privacy filter runs before any request is built:
|
|
300
|
+
|
|
301
|
+
- **Excluded, never sent:** a turn that names your account, cites a numbered
|
|
302
|
+
notes-folder path (two digits, a space, a folder name, a `.md` file), mentions a private topic (money,
|
|
303
|
+
finance, wallet, seed, key, password, token, salary, bank, journal, health,
|
|
304
|
+
family, whole words), or holds an email address or phone-like number.
|
|
305
|
+
- **Redacted, then sent:** API keys and tokens, AWS key ids, private-key blocks,
|
|
306
|
+
40-hex `0x` addresses, and home directory paths.
|
|
307
|
+
|
|
308
|
+
Excluded turns and failed requests get no verdict. They are reported, never
|
|
309
|
+
filled in with the regex. The correction rate then uses the scored turns as its
|
|
310
|
+
denominator (`correction_rate_denominator: "jev.scored"`). JSON gains a `jev`
|
|
311
|
+
block with `sent`, `excluded`, `excluded_reasons`, `scored`, `errors`,
|
|
312
|
+
`cost_usd` and the served model.
|
|
313
|
+
|
|
314
|
+
Preview the privacy counts before choosing to send anything:
|
|
315
|
+
|
|
316
|
+
```sh
|
|
317
|
+
transcripto coach --detector jev --jev-dry-run --json
|
|
318
|
+
transcripto export-run latest --detector jev --jev-dry-run
|
|
319
|
+
```
|
|
320
|
+
|
|
321
|
+
This needs no API key and makes no requests, even with a key in the environment.
|
|
322
|
+
Both commands return the dedicated `transcripto.jev-privacy-preview/1` JSON
|
|
323
|
+
schema in dry-run mode (`coach` needs `--json`). The `jev` count block and null
|
|
324
|
+
correction fields stay available to existing count consumers. Ordinary coach
|
|
325
|
+
episodes and export session, file, tool and commit details are omitted; dry runs
|
|
326
|
+
do not inspect the project reflog or build an episode report.
|
|
327
|
+
|
|
328
|
+
It reports eligible turns, exclusions by reason, and redaction counts. It does
|
|
329
|
+
not print turn text, estimate cost, or produce correction verdicts. Eligibility
|
|
330
|
+
means the current filter permits a turn; it is not a guarantee that the text
|
|
331
|
+
contains no private information.
|
|
332
|
+
|
|
333
|
+
Options: `--jev-threshold` (default 0.30 on P(correction)), `--jev-max-usd`
|
|
334
|
+
(default 1.00, stops sending once reached), `--jev-batch` (default 1; larger
|
|
335
|
+
batches are cheaper but change the answers), `--jev-fallback-regex` (with no
|
|
336
|
+
key set, use the regex instead of exiting). A refused key (HTTP 401, 402, 403)
|
|
337
|
+
on the first request stops before any later batch. If a later request is refused,
|
|
338
|
+
completed verdicts are retained and no further wave starts.
|
|
339
|
+
|
|
340
|
+
The `eligible` count describes turns allowed by the filter; `sent` counts turns
|
|
341
|
+
submitted to transport, excluding later turns skipped by the spending stop.
|
|
342
|
+
Neither count proves that the remote service received a request successfully.
|
|
343
|
+
|
|
344
|
+
The spend limit must be finite and positive. Costs are reported after requests,
|
|
345
|
+
so requests already in flight can exceed the limit; it is not a provider-side
|
|
346
|
+
hard cap. If any request cost is missing or invalid, no further batch is sent
|
|
347
|
+
and the displayed cost is labelled an incomplete subtotal. Invalid probabilities
|
|
348
|
+
produce no verdict rather than a guessed correction label.
|
|
349
|
+
|
|
350
|
+
The default 0.30 comes from a local experiment on 185 turns, labelled by a
|
|
351
|
+
single model rater: agreement F1 about 0.77 to 0.83 against that rater, versus
|
|
352
|
+
0.68 to 0.70 for the regex. That is agreement with a model, not accuracy.
|
|
353
|
+
|
|
354
|
+
Replay and coach read transcripts without making an index. Search writes text
|
|
355
|
+
and file metadata to `~/.trace/trace.db`. A new index directory is private;
|
|
356
|
+
database and WAL files use mode `0600`. The index stays after the command exits.
|
|
357
|
+
Schema upgrades rebuild it locally. `replay --demo` briefly writes an invented
|
|
358
|
+
transcript to a temporary directory and removes it afterward.
|
|
359
|
+
|
|
360
|
+
Malformed records and unreadable files produce diagnostics. Search indexes the
|
|
361
|
+
valid records of partially malformed files and repeats the warning on later
|
|
362
|
+
queries until the source is repaired. A wholly unreadable file keeps any prior
|
|
363
|
+
indexed copy, with an explicit warning; replay always reads the source. Files larger than
|
|
364
|
+
128 MiB are skipped before parsing; lines larger than 8 MiB are discarded as
|
|
365
|
+
whole records. Split larger files into smaller JSONL files to inspect them.
|
|
366
|
+
Individual displayed text fields are bounded at 16,000 characters. Replay's
|
|
367
|
+
source references let you inspect the original. Terminal control sequences are
|
|
368
|
+
removed from rendered transcript content.
|
|
369
|
+
|
|
370
|
+
## Development
|
|
371
|
+
|
|
372
|
+
Python 3.9+, standard library only. The CLI remains in `transcripto.py`;
|
|
373
|
+
`transcripto_core.py` owns normalization and evidence; `transcripto_replay.py`
|
|
374
|
+
owns replay selection and presentation. All fixtures committed here are synthetic.
|
|
375
|
+
|
|
376
|
+
```sh
|
|
377
|
+
python3 -m unittest discover -s tests -v
|
|
378
|
+
for test in test_*.sh; do bash "$test" || exit; done
|
|
379
|
+
```
|
|
380
|
+
|
|
381
|
+
`test_distribution.sh` requires the development-only `build` package. It builds
|
|
382
|
+
an sdist, builds the wheel from that archive, installs without dependencies in a
|
|
383
|
+
fresh virtual environment, and exercises discovery, search and exact replay
|
|
384
|
+
across all three harnesses in an isolated synthetic HOME.
|
|
385
|
+
|
|
386
|
+
The regression cases include failed edits and commits, missing/mismatched
|
|
387
|
+
results, Cursor call shapes, Codex wrappers, result attribution across prompts,
|
|
388
|
+
rollback order, malformed JSON, a sparse 2 GiB file, terminal controls, private
|
|
389
|
+
index permissions, incremental search, and cross-harness retrieval.
|
|
390
|
+
|
|
391
|
+
MIT. Open an issue with the **record shape** that fails, or a synthetic
|
|
392
|
+
reproduction. Your real prompt text is not needed.
|
|
393
|
+
|
|
394
|
+
### Inspect one session's Jev findings, then carry one candidate
|
|
395
|
+
|
|
396
|
+
Added in 0.3.0: a selected-session path.
|
|
397
|
+
Preview remains counts-only, offline and keyless:
|
|
398
|
+
|
|
399
|
+
```sh
|
|
400
|
+
transcripto jev-findings /path/session.jsonl --detector jev --jev-dry-run
|
|
401
|
+
```
|
|
402
|
+
|
|
403
|
+
Only when you choose to send that session's privacy-filtered typed turns:
|
|
404
|
+
|
|
405
|
+
```sh
|
|
406
|
+
transcripto jev-findings /path/session.jsonl --detector jev \
|
|
407
|
+
--jev-max-usd 0.05 --output /your/private/findings.json
|
|
408
|
+
```
|
|
409
|
+
|
|
410
|
+
The local report contains references, source/request hashes, exact lines,
|
|
411
|
+
probabilities, threshold, model metadata and observation time, without transcript
|
|
412
|
+
text. It marks candidate, not-candidate, excluded and unknown separately. Each
|
|
413
|
+
row prints its exact replay command. A candidate is a recorded model suggestion,
|
|
414
|
+
not a confirmed human correction. The serving-model hint is not a per-turn
|
|
415
|
+
model guarantee. The cap and provider/privacy limits above still apply.
|
|
416
|
+
|
|
417
|
+
After inspecting a candidate's request and recorded work, select its exact line:
|
|
418
|
+
|
|
419
|
+
```sh
|
|
420
|
+
transcripto replay --findings /your/private/findings.json --line 3
|
|
421
|
+
transcripto handoff --findings /your/private/findings.json --line 3 \
|
|
422
|
+
--to-harness codex --output /your/private/candidate.json
|
|
423
|
+
transcripto receive-handoff /your/private/candidate.json --as-harness codex \
|
|
424
|
+
--output /your/private/receiver-brief.md
|
|
425
|
+
```
|
|
426
|
+
|
|
427
|
+
Use the actual line shown by your report and choose a receiver different from
|
|
428
|
+
the source harness. Replay and handoff refuse a changed source, including changed
|
|
429
|
+
follow-up records around an unchanged request. Excluded, unknown and negative
|
|
430
|
+
findings cannot become candidate handoffs. A previously prepared packet whose
|
|
431
|
+
source changes remains historical; its receiver brief marks outcomes provisional.
|
|
432
|
+
|
|
433
|
+
The packet and receiver brief are private local files and contain selected
|
|
434
|
+
transcript text. They retain detector provenance, synthetic/test labels, and
|
|
435
|
+
pending human confirmation and receiver acknowledgement. They do not invoke an
|
|
436
|
+
agent or send a message. Reports, packets and briefs use mode `0600`; inspect
|
|
437
|
+
before sharing. `replay --share` remains counts-only.
|
|
438
|
+
|
|
439
|
+
### Choose a local replay and author a handoff
|
|
440
|
+
|
|
441
|
+
A model report is optional. List metadata from an existing index, choose a source,
|
|
442
|
+
then explicitly permit local viewing of its requests and recorded tool outcomes:
|
|
443
|
+
|
|
444
|
+
```sh
|
|
445
|
+
transcripto selected-context runs --cwd /your/repo
|
|
446
|
+
transcripto selected-context describe --source /your/session.jsonl
|
|
447
|
+
transcripto selected-context episodes --source /your/session.jsonl \
|
|
448
|
+
--accept-sha SHA256_FROM_DESCRIBE --consent
|
|
449
|
+
transcripto handoff --source /your/session.jsonl \
|
|
450
|
+
--accept-sha SHA256_FROM_DESCRIBE --line 3 \
|
|
451
|
+
--instruction 'Repair the selected output and verify the stated condition.' \
|
|
452
|
+
--consent --to-harness claude --output /your/private/packet.json
|
|
453
|
+
transcripto receive-handoff /your/private/packet.json --as-harness claude \
|
|
454
|
+
--output /your/private/brief.md
|
|
455
|
+
```
|
|
456
|
+
|
|
457
|
+
`runs` accepts `--index /your/existing.sqlite`; it does not create or refresh an
|
|
458
|
+
index. It reads only metadata from at most the most recent 20,000 indexed records,
|
|
459
|
+
returning up to 20 runs by default (maximum 30). Suggestions are unbound: sharing a
|
|
460
|
+
working directory does not prove that a run produced your artifact. No title or
|
|
461
|
+
body is read by this query. SQLite may use locking sidecars. Choose a file explicitly
|
|
462
|
+
when the bounded index window has no suitable run.
|
|
463
|
+
|
|
464
|
+
`describe` reads bytes to compute identity without returning transcript text. The
|
|
465
|
+
consented replay returns at most the first 100 requests from one file of at most
|
|
466
|
+
16 MiB; changed bytes or parsing warnings refuse replay. Authored handoffs retain
|
|
467
|
+
the original request, source hash, exact line and recorded outcomes. The new
|
|
468
|
+
instruction is explicit authorship for this handoff, not a detector verdict,
|
|
469
|
+
inferred human REDO or research label. Same-harness refusal and synthetic labels
|
|
470
|
+
remain. These commands prepare private local files; they do not invoke a receiver
|
|
471
|
+
or send a message. Review the full brief before giving it to another process.
|
|
472
|
+
|
|
473
|
+
### Explicit authored continuation
|
|
474
|
+
|
|
475
|
+
`handoff` still requires a different receiver harness. For a separately chosen new
|
|
476
|
+
Claude session, the local producer can describe exactly one authored instruction:
|
|
477
|
+
|
|
478
|
+
```sh
|
|
479
|
+
transcripto selected-context describe --source /path/to/session.jsonl
|
|
480
|
+
transcripto selected-context authored-continuation \
|
|
481
|
+
--source /path/to/session.jsonl --accept-sha SHA256_FROM_DESCRIBE \
|
|
482
|
+
--line 1 --instruction 'Add the missing label.' --consent
|
|
483
|
+
```
|
|
484
|
+
|
|
485
|
+
This returns a typed local selection, not a handoff exception, detector verdict or
|
|
486
|
+
receiver invocation. Its source session UUID is derived from the consented bytes;
|
|
487
|
+
absent, malformed or mixed identities refuse. No filename/index fallback is used.
|
|
488
|
+
ZUP's separate new-session confirmation binds a fresh receiver UUID and actual
|
|
489
|
+
launch contract. New-session work is same-harness authored work, not independent
|
|
490
|
+
judgement or measured model improvement. Only explicitly selected context is carried.
|