transcripto 0.2.0__tar.gz → 0.2.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: transcripto
3
- Version: 0.2.0
3
+ Version: 0.2.1
4
4
  Summary: Instant replay for coding agents. Inspect requests, tool calls, and recorded results across Claude Code, Codex, and Cursor. Local, stdlib-only.
5
5
  Author: Oscar Morke
6
6
  License: MIT
@@ -18,27 +18,132 @@ Dynamic: license-file
18
18
 
19
19
  # Transcripto
20
20
 
21
- **Your agent said “done.” Replay what happened.**
21
+ **You might already be keeping a journal. Read your side of it.**
22
22
 
23
- Transcripto reads the transcripts already on your machine and puts your request,
24
- the tool calls, and their recorded results in order. Failed edits stay failed.
25
- Missing results stay unknown. A confident closing message changes neither.
23
+ Your agent transcripts contain what you asked for, what you changed your mind
24
+ about, and what you kept coming back to. Transcripto helps you find those words
25
+ and read the recorded work around them.
26
26
 
27
- Claude Code · Codex · Cursor. Local. No accounts. No runtime dependencies.
27
+ Claude Code · Codex · Cursor. Local files. No account. No runtime dependencies.
28
+
29
+ ## Start with something you remember saying
30
+
31
+ ```sh
32
+ uvx --from transcripto==0.2.1 transcripto ask "retry"
33
+ ```
34
+
35
+ Replace `retry` with a word you remember using. `ask` searches messages identified
36
+ as yours and shows dated snippets, newest first. It refreshes the local index
37
+ automatically. The first search indexes the selected history; a large archive
38
+ can take minutes. Add `--harness claude`, `--harness codex`, or `--harness cursor`
39
+ to limit that scan. It does not generate a diary or interpret your personality.
40
+
41
+ Each hit prints an `Open:` command. Run that command to open the exact request
42
+ and its recorded work. This also works when search matches a word variant
43
+ (such as `retry` matching `retried`) or several requests share the same words.
44
+
45
+ You can also search replay directly:
46
+
47
+ ```sh
48
+ uvx --from transcripto==0.2.1 transcripto replay "retry"
49
+
50
+ # Or open your latest human session:
51
+ uvx --from transcripto==0.2.1 transcripto
52
+ ```
53
+
54
+ Replay puts your request, tool calls and recorded results in order. Failed edits
55
+ stay failed. Missing results stay unknown. Status describes tool execution,
56
+ not whether the task was done correctly.
57
+
58
+ Or install with `python3 -m pip install transcripto==0.2.1`, then run
59
+ `transcripto ask "retry"`. Requires Python 3.9 or newer.
60
+
61
+ ## Try the stranger flow without your transcripts
62
+
63
+ The bundled public example is synthetic. It works in an isolated home and does
64
+ not depend on agent dotfiles:
65
+
66
+ ```sh
67
+ INSTALL="$(mktemp -d)"
68
+ python3 -m pip install --no-deps --no-build-isolation --target "$INSTALL" .
69
+ export HOME="$(mktemp -d)"
70
+ transcripto() { PYTHONPATH="$INSTALL" python3 -m transcripto "$@"; }
71
+
72
+ transcripto import-example
73
+ transcripto ask "What changed about the forecast cache?"
74
+ transcripto changes
75
+ ```
76
+
77
+ `ask` cites the imported JSONL line for every hit. `changes` is a focused view
78
+ of the request that was revised, the correction, and its recorded follow-up.
79
+ It labels missing results rather than turning a change of mind into a score.
80
+
81
+ To carry that correction to a different receiver:
28
82
 
29
83
  ```sh
30
- uvx --from transcripto==0.2.0 transcripto
84
+ transcripto handoff "30 seconds" \
85
+ --to-harness codex --output "$HOME/codex-inbox/correction.json"
86
+ transcripto receive-handoff \
87
+ "$HOME/codex-inbox/correction.json" --as-harness codex \
88
+ --output "$HOME/codex-work/receiver-brief.md"
89
+ cat "$HOME/codex-work/receiver-brief.md"
90
+ ```
91
+
92
+ The brief includes the cited correction, recorded follow-up statuses
93
+ (failed / succeeded / unknown), and an `Open:` command for the exact request.
94
+ If the source moved, disappeared, or no longer holds the cited request, the brief
95
+ marks evidence uncertain and shows only the packet's own statuses as provisional. It does not invoke a receiver agent or prove
96
+ adoption. Synthetic provenance stays visible in search, changes, and handoffs.
97
+ Handoff files are local and mode `0600`; they can contain transcript text and
98
+ paths, so review them before sharing.
31
99
 
32
- # Try a synthetic session without opening your own history:
33
- uvx --from transcripto==0.2.0 transcripto replay --demo
100
+ ### Cross-harness lab (failed / succeeded / unknown)
101
+
102
+ For a receiving agent that needs to find a prior episode and reopen exact
103
+ evidence without private history:
104
+
105
+ ```sh
106
+ transcripto import-lab
107
+ transcripto ask "retry"
108
+ # run each printed Open: command
34
109
  ```
35
110
 
36
- Or install with `python3 -m pip install transcripto==0.2.0`, then run
37
- `transcripto`. Requires Python 3.9 or newer.
111
+ All lab records are labelled synthetic. Claude shows a failed edit, Codex a
112
+ succeeded check, Cursor an unknown missing result.
113
+
114
+ ### Offline flight card
38
115
 
39
- Version **0.2.0** adds replay. The default command opens your latest human
40
- session. Pinning the version makes the reviewed build and the build you run
41
- the same object.
116
+ ```sh
117
+ transcripto quickstart --wheel /absolute/path/to/transcripto-0.2.1-py3-none-any.whl
118
+ ```
119
+
120
+ Prints install, `import-lab`, search, and reopen commands for a built wheel
121
+ without PyPI. See also `docs/OFFLINE-QUICKSTART.md`.
122
+
123
+ **Your files remain yours.** Transcripto does not upload transcript content or
124
+ execute commands found in it. Search output, replay and JSON can contain private
125
+ words and paths; review anything you choose to share. It reads existing files,
126
+ not deleted history. Check your agent's retention settings and keep your own
127
+ backup if you want a lasting record. Authorship detection differs by harness;
128
+ [see the limits below](#what-each-harness-supports).
129
+
130
+ ## Find the thing you remember
131
+
132
+ Search automatically refreshes a local index. No setup command is required.
133
+
134
+ ```sh
135
+ transcripto ask "retry" # your submitted words
136
+ transcripto search "retry" # prompts, replies, and tool text
137
+ transcripto find parser.py # recorded file operations and attempts
138
+ transcripto trace "retry" # an alias into result-aware replay
139
+ transcripto sessions # sessions with submitted prompts
140
+ transcripto stats # activity counts
141
+ ```
142
+
143
+ A failed or unconfirmed file change is labelled an **attempt**, never `WROTE`.
144
+ Queries with `--harness` or `--root` are scoped to that selection even if the
145
+ index already contains another corpus. `index` and `watch` remain available for
146
+ explicit refresh and background polling.
42
147
 
43
148
  ## The replay
44
149
 
@@ -84,6 +189,7 @@ transcripto replay "login redirect" # find requests containing these wor
84
189
  transcripto replay path/to/session.jsonl # inspect one transcript
85
190
  transcripto replay --session 3f9c1a2b # explicitly select a session prefix
86
191
  transcripto replay path/to/session.jsonl --episode 3 --all
192
+ transcripto replay path/to/session.jsonl --line 42 # exact request from an ask hit
87
193
  transcripto replay latest --json # structured events, evidence, source lines
88
194
  transcripto replay latest --share # counts + caveat; no prompts or paths
89
195
  ```
@@ -91,24 +197,6 @@ transcripto replay latest --share # counts + caveat; no prompts or pat
91
197
  `--share` is intentionally small. Full replay output and JSON contain your own
92
198
  words and local paths. The tool does not upload either.
93
199
 
94
- ## Find the thing you remember
95
-
96
- Search automatically refreshes a local index. No setup command is required.
97
-
98
- ```sh
99
- transcripto ask "retry" # your submitted words
100
- transcripto search "retry" # prompts, replies, and tool text
101
- transcripto find parser.py # recorded file operations and attempts
102
- transcripto trace "retry" # an alias into result-aware replay
103
- transcripto sessions # sessions with submitted prompts
104
- transcripto stats # activity counts
105
- ```
106
-
107
- A failed or unconfirmed file change is labelled an **attempt**, never `WROTE`.
108
- Queries with `--harness` or `--root` are scoped to that selection even if the
109
- index already contains another corpus. `index` and `watch` remain available for
110
- explicit refresh and background polling.
111
-
112
200
  ## What each harness supports
113
201
 
114
202
  | Feature | Claude Code | Codex | Cursor |
@@ -213,6 +301,11 @@ python3 -m unittest discover -s tests -v
213
301
  for test in test_*.sh; do bash "$test" || exit; done
214
302
  ```
215
303
 
304
+ `test_distribution.sh` requires the development-only `build` package. It builds
305
+ an sdist, builds the wheel from that archive, installs without dependencies in a
306
+ fresh virtual environment, and exercises discovery, search and exact replay
307
+ across all three harnesses in an isolated synthetic HOME.
308
+
216
309
  The regression cases include failed edits and commits, missing/mismatched
217
310
  results, Cursor call shapes, Codex wrappers, result attribution across prompts,
218
311
  rollback order, malformed JSON, a sparse 2 GiB file, terminal controls, private
@@ -1,26 +1,131 @@
1
1
  # Transcripto
2
2
 
3
- **Your agent said “done.” Replay what happened.**
3
+ **You might already be keeping a journal. Read your side of it.**
4
4
 
5
- Transcripto reads the transcripts already on your machine and puts your request,
6
- the tool calls, and their recorded results in order. Failed edits stay failed.
7
- Missing results stay unknown. A confident closing message changes neither.
5
+ Your agent transcripts contain what you asked for, what you changed your mind
6
+ about, and what you kept coming back to. Transcripto helps you find those words
7
+ and read the recorded work around them.
8
8
 
9
- Claude Code · Codex · Cursor. Local. No accounts. No runtime dependencies.
9
+ Claude Code · Codex · Cursor. Local files. No account. No runtime dependencies.
10
+
11
+ ## Start with something you remember saying
12
+
13
+ ```sh
14
+ uvx --from transcripto==0.2.1 transcripto ask "retry"
15
+ ```
16
+
17
+ Replace `retry` with a word you remember using. `ask` searches messages identified
18
+ as yours and shows dated snippets, newest first. It refreshes the local index
19
+ automatically. The first search indexes the selected history; a large archive
20
+ can take minutes. Add `--harness claude`, `--harness codex`, or `--harness cursor`
21
+ to limit that scan. It does not generate a diary or interpret your personality.
22
+
23
+ Each hit prints an `Open:` command. Run that command to open the exact request
24
+ and its recorded work. This also works when search matches a word variant
25
+ (such as `retry` matching `retried`) or several requests share the same words.
26
+
27
+ You can also search replay directly:
28
+
29
+ ```sh
30
+ uvx --from transcripto==0.2.1 transcripto replay "retry"
31
+
32
+ # Or open your latest human session:
33
+ uvx --from transcripto==0.2.1 transcripto
34
+ ```
35
+
36
+ Replay puts your request, tool calls and recorded results in order. Failed edits
37
+ stay failed. Missing results stay unknown. Status describes tool execution,
38
+ not whether the task was done correctly.
39
+
40
+ Or install with `python3 -m pip install transcripto==0.2.1`, then run
41
+ `transcripto ask "retry"`. Requires Python 3.9 or newer.
42
+
43
+ ## Try the stranger flow without your transcripts
44
+
45
+ The bundled public example is synthetic. It works in an isolated home and does
46
+ not depend on agent dotfiles:
47
+
48
+ ```sh
49
+ INSTALL="$(mktemp -d)"
50
+ python3 -m pip install --no-deps --no-build-isolation --target "$INSTALL" .
51
+ export HOME="$(mktemp -d)"
52
+ transcripto() { PYTHONPATH="$INSTALL" python3 -m transcripto "$@"; }
53
+
54
+ transcripto import-example
55
+ transcripto ask "What changed about the forecast cache?"
56
+ transcripto changes
57
+ ```
58
+
59
+ `ask` cites the imported JSONL line for every hit. `changes` is a focused view
60
+ of the request that was revised, the correction, and its recorded follow-up.
61
+ It labels missing results rather than turning a change of mind into a score.
62
+
63
+ To carry that correction to a different receiver:
10
64
 
11
65
  ```sh
12
- uvx --from transcripto==0.2.0 transcripto
66
+ transcripto handoff "30 seconds" \
67
+ --to-harness codex --output "$HOME/codex-inbox/correction.json"
68
+ transcripto receive-handoff \
69
+ "$HOME/codex-inbox/correction.json" --as-harness codex \
70
+ --output "$HOME/codex-work/receiver-brief.md"
71
+ cat "$HOME/codex-work/receiver-brief.md"
72
+ ```
73
+
74
+ The brief includes the cited correction, recorded follow-up statuses
75
+ (failed / succeeded / unknown), and an `Open:` command for the exact request.
76
+ If the source moved, disappeared, or no longer holds the cited request, the brief
77
+ marks evidence uncertain and shows only the packet's own statuses as provisional. It does not invoke a receiver agent or prove
78
+ adoption. Synthetic provenance stays visible in search, changes, and handoffs.
79
+ Handoff files are local and mode `0600`; they can contain transcript text and
80
+ paths, so review them before sharing.
13
81
 
14
- # Try a synthetic session without opening your own history:
15
- uvx --from transcripto==0.2.0 transcripto replay --demo
82
+ ### Cross-harness lab (failed / succeeded / unknown)
83
+
84
+ For a receiving agent that needs to find a prior episode and reopen exact
85
+ evidence without private history:
86
+
87
+ ```sh
88
+ transcripto import-lab
89
+ transcripto ask "retry"
90
+ # run each printed Open: command
16
91
  ```
17
92
 
18
- Or install with `python3 -m pip install transcripto==0.2.0`, then run
19
- `transcripto`. Requires Python 3.9 or newer.
93
+ All lab records are labelled synthetic. Claude shows a failed edit, Codex a
94
+ succeeded check, Cursor an unknown missing result.
95
+
96
+ ### Offline flight card
20
97
 
21
- Version **0.2.0** adds replay. The default command opens your latest human
22
- session. Pinning the version makes the reviewed build and the build you run
23
- the same object.
98
+ ```sh
99
+ transcripto quickstart --wheel /absolute/path/to/transcripto-0.2.1-py3-none-any.whl
100
+ ```
101
+
102
+ Prints install, `import-lab`, search, and reopen commands for a built wheel
103
+ without PyPI. See also `docs/OFFLINE-QUICKSTART.md`.
104
+
105
+ **Your files remain yours.** Transcripto does not upload transcript content or
106
+ execute commands found in it. Search output, replay and JSON can contain private
107
+ words and paths; review anything you choose to share. It reads existing files,
108
+ not deleted history. Check your agent's retention settings and keep your own
109
+ backup if you want a lasting record. Authorship detection differs by harness;
110
+ [see the limits below](#what-each-harness-supports).
111
+
112
+ ## Find the thing you remember
113
+
114
+ Search automatically refreshes a local index. No setup command is required.
115
+
116
+ ```sh
117
+ transcripto ask "retry" # your submitted words
118
+ transcripto search "retry" # prompts, replies, and tool text
119
+ transcripto find parser.py # recorded file operations and attempts
120
+ transcripto trace "retry" # an alias into result-aware replay
121
+ transcripto sessions # sessions with submitted prompts
122
+ transcripto stats # activity counts
123
+ ```
124
+
125
+ A failed or unconfirmed file change is labelled an **attempt**, never `WROTE`.
126
+ Queries with `--harness` or `--root` are scoped to that selection even if the
127
+ index already contains another corpus. `index` and `watch` remain available for
128
+ explicit refresh and background polling.
24
129
 
25
130
  ## The replay
26
131
 
@@ -66,6 +171,7 @@ transcripto replay "login redirect" # find requests containing these wor
66
171
  transcripto replay path/to/session.jsonl # inspect one transcript
67
172
  transcripto replay --session 3f9c1a2b # explicitly select a session prefix
68
173
  transcripto replay path/to/session.jsonl --episode 3 --all
174
+ transcripto replay path/to/session.jsonl --line 42 # exact request from an ask hit
69
175
  transcripto replay latest --json # structured events, evidence, source lines
70
176
  transcripto replay latest --share # counts + caveat; no prompts or paths
71
177
  ```
@@ -73,24 +179,6 @@ transcripto replay latest --share # counts + caveat; no prompts or pat
73
179
  `--share` is intentionally small. Full replay output and JSON contain your own
74
180
  words and local paths. The tool does not upload either.
75
181
 
76
- ## Find the thing you remember
77
-
78
- Search automatically refreshes a local index. No setup command is required.
79
-
80
- ```sh
81
- transcripto ask "retry" # your submitted words
82
- transcripto search "retry" # prompts, replies, and tool text
83
- transcripto find parser.py # recorded file operations and attempts
84
- transcripto trace "retry" # an alias into result-aware replay
85
- transcripto sessions # sessions with submitted prompts
86
- transcripto stats # activity counts
87
- ```
88
-
89
- A failed or unconfirmed file change is labelled an **attempt**, never `WROTE`.
90
- Queries with `--harness` or `--root` are scoped to that selection even if the
91
- index already contains another corpus. `index` and `watch` remain available for
92
- explicit refresh and background polling.
93
-
94
182
  ## What each harness supports
95
183
 
96
184
  | Feature | Claude Code | Codex | Cursor |
@@ -195,6 +283,11 @@ python3 -m unittest discover -s tests -v
195
283
  for test in test_*.sh; do bash "$test" || exit; done
196
284
  ```
197
285
 
286
+ `test_distribution.sh` requires the development-only `build` package. It builds
287
+ an sdist, builds the wheel from that archive, installs without dependencies in a
288
+ fresh virtual environment, and exercises discovery, search and exact replay
289
+ across all three harnesses in an isolated synthetic HOME.
290
+
198
291
  The regression cases include failed edits and commits, missing/mismatched
199
292
  results, Cursor call shapes, Codex wrappers, result attribution across prompts,
200
293
  rollback order, malformed JSON, a sparse 2 GiB file, terminal controls, private
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
4
4
 
5
5
  [project]
6
6
  name = "transcripto"
7
- version = "0.2.0"
7
+ version = "0.2.1"
8
8
  description = "Instant replay for coding agents. Inspect requests, tool calls, and recorded results across Claude Code, Codex, and Cursor. Local, stdlib-only."
9
9
  readme = "README.md"
10
10
  requires-python = ">=3.9"
@@ -0,0 +1,203 @@
1
+ """Unlabelled harness rows: detection, in-place migration, and path backfill.
2
+
3
+ Every record here is invented. The stores are built in a fresh HOME.
4
+
5
+ The defect this guards: a writer whose INSERT names fewer columns than the table
6
+ has leaves `harness` NULL with no error, and a schema check that only knew how to
7
+ drop and rebuild never repaired it. These tests fail on code that has no
8
+ `backfill-harness` command or that drops a version-3 store instead of migrating it.
9
+ """
10
+ import hashlib
11
+ import json
12
+ import os
13
+ from pathlib import Path
14
+ import sqlite3
15
+ import subprocess
16
+ import sys
17
+ import tempfile
18
+ import unittest
19
+
20
+ ROOT = Path(__file__).resolve().parents[1]
21
+ sys.path.insert(0, str(ROOT))
22
+
23
+ # The version-3 schema exactly as a v3 writer created it: harness present, no
24
+ # source_line, no synthetic, user_version 3.
25
+ V3_SCHEMA = """
26
+ CREATE TABLE messages(
27
+ id INTEGER PRIMARY KEY, session_id TEXT, session_file TEXT, project TEXT,
28
+ ts TEXT, role TEXT, cwd TEXT, git_branch TEXT, text TEXT,
29
+ is_human INTEGER DEFAULT 0, prompt_source TEXT, harness TEXT);
30
+ CREATE VIRTUAL TABLE messages_fts USING fts5(
31
+ text, content='messages', content_rowid='id', tokenize="porter unicode61");
32
+ CREATE TABLE files(
33
+ id INTEGER PRIMARY KEY, path TEXT, name TEXT, action TEXT,
34
+ session_id TEXT, session_file TEXT, ts TEXT, cwd TEXT, harness TEXT);
35
+ CREATE TABLE indexed(session_file TEXT PRIMARY KEY, mtime REAL, warnings TEXT);
36
+ PRAGMA user_version=3;
37
+ """
38
+
39
+ # The legacy writer: ten named columns into a table that also has harness.
40
+ LEGACY_INSERT = ("INSERT INTO messages(session_id,session_file,project,ts,role,cwd,git_branch,text,"
41
+ "is_human,prompt_source) VALUES(?,?,?,?,?,?,?,?,?,?)")
42
+ LEGACY_FILE_INSERT = ("INSERT INTO files(path,name,action,session_id,session_file,ts,cwd)"
43
+ " VALUES(?,?,?,?,?,?,?)")
44
+
45
+
46
+ def legacy_row(sid, path, text):
47
+ return (sid, str(path), "demo", "2026-09-12T10:00:00Z", "user", "/tmp/demo", "", text, 1, "typed")
48
+
49
+
50
+ class Home(unittest.TestCase):
51
+ def setUp(self):
52
+ self.temp = tempfile.TemporaryDirectory(prefix="transcripto-harness-")
53
+ self.addCleanup(self.temp.cleanup)
54
+ self.home = Path(self.temp.name)
55
+ self.db = self.home / ".trace" / "trace.db"
56
+ self.env = dict(os.environ, HOME=str(self.home), PYTHONIOENCODING="utf-8")
57
+ self.env.pop("PYTHONPATH", None)
58
+
59
+ def cli(self, *args):
60
+ return subprocess.run([sys.executable, str(ROOT / "transcripto.py")] + list(args),
61
+ cwd=self.home, env=self.env, capture_output=True, text=True, timeout=60)
62
+
63
+ def write(self, relative, records):
64
+ path = self.home / relative
65
+ path.parent.mkdir(parents=True, exist_ok=True)
66
+ path.write_text("".join(json.dumps(r) + "\n" for r in records))
67
+ return path
68
+
69
+ def claude_session(self, relative, text):
70
+ return self.write(relative, [
71
+ {"type": "user", "promptSource": "typed", "sessionId": "c-" + text[:4],
72
+ "timestamp": "2026-09-12T09:00:00Z", "message": {"content": text}},
73
+ {"type": "assistant", "timestamp": "2026-09-12T09:00:01Z", "message": {"content": [
74
+ {"type": "tool_use", "name": "Edit", "id": "e1", "input": {"file_path": "/tmp/demo/a.py"}}]}},
75
+ ])
76
+
77
+ def query(self, sql, args=()):
78
+ con = sqlite3.connect(str(self.db))
79
+ try:
80
+ return con.execute(sql, args).fetchall()
81
+ finally:
82
+ con.close()
83
+
84
+
85
+ class PathLabelTests(unittest.TestCase):
86
+ def test_paths_name_their_harness(self):
87
+ import transcripto
88
+ cases = {
89
+ "/h/.claude/projects/-tmp-demo/s.jsonl": "claude",
90
+ "/h/.codex/sessions/2026/09/12/rollout.jsonl": "codex",
91
+ "/h/.codex/archived_sessions/a.jsonl": "codex",
92
+ "/h/.codex/history.jsonl": "codex-history",
93
+ "/h/.cursor/projects/p/agent-transcripts/x/x.jsonl": "cursor",
94
+ "/h/.transcripto/imports/cursor/lab.jsonl": "cursor",
95
+ "/h/.transcripto/imports/unknown/lab.jsonl": None,
96
+ "/h/archive/.codex/sessions/old.jsonl": "codex",
97
+ "/h/notes/session.jsonl": None,
98
+ }
99
+ for path, expected in cases.items():
100
+ self.assertEqual(transcripto.harness_from_path(path), expected, path)
101
+
102
+
103
+ class LegacyInsertTests(Home):
104
+ def test_backfill_labels_rows_a_ten_column_insert_left_null(self):
105
+ self.claude_session(".claude/projects/-tmp-demo/one.jsonl", "Rename the amber helper.")
106
+ done = self.cli("index")
107
+ self.assertEqual(done.returncode, 0, done.stderr)
108
+ self.assertEqual(self.query("SELECT COUNT(*) FROM messages WHERE harness IS NULL")[0][0], 0,
109
+ "a current index pass must never write an unlabelled row")
110
+
111
+ codex = self.home / ".codex/sessions/2026/09/12/rollout-legacy.jsonl"
112
+ cursor = self.home / ".cursor/projects/p/agent-transcripts/x/x.jsonl"
113
+ claude = self.home / ".claude/projects/-tmp-demo/legacy.jsonl"
114
+ stray = self.home / "notes/unknown.jsonl"
115
+ con = sqlite3.connect(str(self.db))
116
+ for i, path in enumerate((codex, codex, cursor, claude, stray)):
117
+ con.execute(LEGACY_INSERT, legacy_row("s%d" % i, path, "invented cobalt request %d" % i))
118
+ con.execute(LEGACY_FILE_INSERT, ("/tmp/demo/b.py", "b.py", "edit", "s0", str(codex),
119
+ "2026-09-12T10:00:00Z", "/tmp/demo"))
120
+ con.commit(); con.close()
121
+ self.assertEqual(self.query("SELECT COUNT(*) FROM messages WHERE harness IS NULL")[0][0], 5)
122
+
123
+ dry = self.cli("backfill-harness", "--json")
124
+ self.assertEqual(dry.returncode, 0, dry.stderr)
125
+ report = json.loads(dry.stdout)
126
+ self.assertFalse(report["applied"])
127
+ self.assertEqual(report["unlabelled_before"], {"messages": 5, "files": 1})
128
+ self.assertEqual(report["unlabelled_after"], {"messages": 5, "files": 1}, "dry run wrote")
129
+ self.assertEqual(report["labels"]["codex"], {"files": 1, "messages": 2, "file_rows": 1})
130
+ self.assertEqual(report["unresolved_rows"], 1)
131
+
132
+ applied = self.cli("backfill-harness", "--apply", "--json")
133
+ self.assertEqual(applied.returncode, 0, applied.stderr)
134
+ self.assertEqual(json.loads(applied.stdout)["unlabelled_after"], {"messages": 1, "files": 0})
135
+ got = dict(self.query("SELECT session_file, harness FROM messages WHERE text LIKE 'invented cobalt%'"))
136
+ self.assertEqual(got[str(codex)], "codex")
137
+ self.assertEqual(got[str(cursor)], "cursor")
138
+ self.assertEqual(got[str(claude)], "claude")
139
+ self.assertIsNone(got[str(stray)], "a path that names no harness must stay unlabelled")
140
+ self.assertEqual(self.query("SELECT harness FROM files WHERE session_file=?", (str(codex),)),
141
+ [("codex",)])
142
+
143
+ def test_open_detects_and_labels_unlabelled_rows(self):
144
+ self.claude_session(".claude/projects/-tmp-demo/one.jsonl", "Rename the amber helper.")
145
+ self.assertEqual(self.cli("index").returncode, 0)
146
+ con = sqlite3.connect(str(self.db))
147
+ con.execute(LEGACY_INSERT, legacy_row("s9", self.home / ".codex/sessions/r.jsonl", "invented"))
148
+ con.commit(); con.close()
149
+ again = self.cli("index")
150
+ self.assertEqual(again.returncode, 0, again.stderr)
151
+ self.assertIn("labelled 1 unlabelled message row", again.stderr)
152
+ self.assertEqual(self.query("SELECT COUNT(*) FROM messages WHERE harness IS NULL")[0][0], 0)
153
+
154
+
155
+ class VersionThreeStoreTests(Home):
156
+ def v3_store(self):
157
+ self.db.parent.mkdir(parents=True, exist_ok=True)
158
+ con = sqlite3.connect(str(self.db))
159
+ con.executescript(V3_SCHEMA)
160
+ # A transcript indexed earlier from a --root outside the default roots. It is
161
+ # still on disk, so an index pass keeps it; a rebuild would lose it.
162
+ kept = self.claude_session("archive/.codex/sessions/old.jsonl", "Kept archive request.")
163
+ cur = con.execute(LEGACY_INSERT, legacy_row("old", kept, "invented archive request"))
164
+ con.execute("INSERT INTO messages_fts(rowid,text) VALUES(?,?)", (cur.lastrowid, "invented archive request"))
165
+ con.execute("INSERT INTO indexed(session_file,mtime,warnings) VALUES(?,?,?)",
166
+ (str(kept), os.path.getmtime(kept), "[]"))
167
+ con.commit(); con.close()
168
+ return kept
169
+
170
+ def test_index_migrates_a_version_three_store_in_place(self):
171
+ kept = self.v3_store()
172
+ self.claude_session(".claude/projects/-tmp-demo/new.jsonl", "Rename the amber helper.")
173
+ done = self.cli("index")
174
+ self.assertEqual(done.returncode, 0, done.stderr)
175
+ self.assertIn("migrated index schema: added messages.source_line, messages.synthetic", done.stderr)
176
+ self.assertEqual(self.query("PRAGMA user_version")[0][0], 5)
177
+ cols = {r[1] for r in self.query("PRAGMA table_info(messages)")}
178
+ self.assertTrue({"harness", "source_line", "synthetic"} <= cols)
179
+ self.assertEqual(self.query("SELECT harness FROM messages WHERE session_file=?", (str(kept),)),
180
+ [("codex",)], "the version-3 row was dropped or left unlabelled")
181
+ self.assertEqual(self.query("SELECT COUNT(*) FROM messages WHERE harness IS NULL")[0][0], 0)
182
+ self.assertEqual(len(self.query("SELECT rowid FROM messages_fts WHERE messages_fts MATCH 'archive'")), 1,
183
+ "full-text index lost the migrated row")
184
+
185
+ def test_backfill_on_a_version_three_store_leaves_the_schema_alone(self):
186
+ kept = self.v3_store()
187
+ before = hashlib.sha256(self.db.read_bytes()).hexdigest()
188
+ dry = self.cli("backfill-harness", "--db", str(self.db))
189
+ self.assertEqual(dry.returncode, 0, dry.stderr)
190
+ self.assertIn("DRY RUN", dry.stdout)
191
+ self.assertEqual(hashlib.sha256(self.db.read_bytes()).hexdigest(), before, "dry run changed the file")
192
+
193
+ applied = self.cli("backfill-harness", "--db", str(self.db), "--apply")
194
+ self.assertEqual(applied.returncode, 0, applied.stderr)
195
+ # A v3 writer rebuilds on any other user_version, so backfill must not bump it.
196
+ self.assertEqual(self.query("PRAGMA user_version")[0][0], 3)
197
+ self.assertEqual(len(self.query("PRAGMA table_info(messages)")), 12)
198
+ self.assertEqual(self.query("SELECT harness FROM messages WHERE session_file=?", (str(kept),)),
199
+ [("codex",)])
200
+
201
+
202
+ if __name__ == "__main__":
203
+ unittest.main()