openqodex 0.4.0 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,176 @@
1
+ # Reviewer drivers
2
+
3
+ This page is for contributors. It records how `openqodex review` starts its own reviewer process, and what was observed with the real binary before a driver was enabled. A driver is enabled only when its isolation was shown by a real run. A limit that no setting removes is stated below and in the user docs.
4
+
5
+ The reviewer is a separate coding-agent process that openqodex starts, without a window, for one review. It reads a frozen copy of the change (the snapshot) and answers with one JSON object. The trace is the agent's own event stream: every tool call, its input and whether it succeeded.
6
+
7
+ ## Claude Code
8
+
9
+ Tested with Claude Code 2.1.289 (`claude --version`) on macOS, 2026-10-03, on a throwaway folder and on the demo repo.
10
+
11
+ ### The command line
12
+
13
+ The driver starts this command without a shell, with the snapshot as the working directory, and writes the brief to standard input:
14
+
15
+ ```
16
+ claude -p --output-format stream-json --verbose --input-format stream-json
17
+ --tools Read,Grep,Glob
18
+ --permission-mode dontAsk
19
+ --setting-sources ""
20
+ --settings {"autoMemoryEnabled":false,"hooks":{},"disableAllHooks":true}
21
+ --strict-mcp-config --mcp-config {"mcpServers":{}}
22
+ --disable-slash-commands
23
+ --no-session-persistence
24
+ ```
25
+
26
+ The child gets an environment built from an allowlist (`reviewerEnv` in `packages/cli/src/reviewers/claude.ts`): `PATH`, `HOME`, `USER`, `LOGNAME`, `SHELL`, `TMPDIR`, locale, `TERM`, `TZ`, `CLAUDE_CONFIG_DIR`, proxy and CA settings, the `ANTHROPIC_*` key, URL and model variables, and the Bedrock, Vertex or Foundry variables only when the matching `CLAUDE_CODE_USE_*` flag is set; plus `OPENQODEX_REVIEW_DEPTH=1`. No other variable is copied, so no developer token and nothing that ties the child to a running Claude Code session (`CLAUDECODE`, `CLAUDE_CODE_SESSION_ID`, messaging sockets) reaches it. A run from inside a Claude Code session with this environment worked as the runs below did.
27
+
28
+ ### What each flag was observed to do
29
+
30
+ | Flag | Observed |
31
+ |---|---|
32
+ | `-p` with `--input-format stream-json` | A fresh session that reads user messages as JSON lines on standard input. After each answer it prints a `result` event and waits for the next message, so a correction round goes to the same session. Closing standard input ends the process with exit 0. |
33
+ | `--output-format stream-json --verbose` | One JSON event per line: an `init` event (tools, MCP servers, plugins, permission mode, memory paths, version), every `tool_use` with its input, every `tool_result` with `is_error`, a `permission_denied` event for each refusal, and a `result` event with the final text, turns, usage and cost. For a Read, `tool_use_result.file` gives the path, `startLine` and `numLines` delivered. |
34
+ | `--tools Read,Grep,Glob` | The `init` event lists exactly `Glob`, `Grep`, `Read`. Asked to run `ls /` and to write a file, the model answered that it has no Bash or Write tool; no such tool call appears in the trace. The `Agent` tool is absent, so no subagent can start. |
35
+ | `--permission-mode dontAsk` | Reads inside the working directory succeed. A Read of an absolute path outside it (`/tmp/.../outside/secret.txt`, `/etc/hosts`, a decoy ssh config), a relative path that leaves it (`../outside/secret.txt`), a Read through a link inside the folder that points outside, and a Grep or Glob rooted outside (`/tmp/...`, `/`) were each refused with a `permission_denied` event and an error result. A recursive Grep and a `**/*` Glob in the folder did not follow the link out. |
36
+ | `--tools Read,Grep,Glob,WebSearch,WebFetch --allowedTools WebSearch,WebFetch` (the default; dropped with `reviewer_web: off`) | The `init` event lists the five tools. Without `--allowedTools`, `dontAsk` refused both web tools ("Permission to use WebFetch has been denied because Claude Code is running in don't ask mode"); with it, a WebFetch of example.com and a WebSearch both returned results (2026-10-03). |
37
+ | `--setting-sources ""` | No user, project or local settings file is read. In the same folder, a run without this flag loaded the project `CLAUDE.md` canary (the answer ended with the canary word) and the user's global instructions (the answer quoted them, about 155,000 input tokens); with it, the input was about 4,500 tokens and neither canary nor any sentence of the global file appeared anywhere in the event stream. With `--include-hook-events`, a run reading user settings showed 11 hook events; this run showed none. |
38
+ | `--settings {"autoMemoryEnabled":false,"hooks":{},"disableAllHooks":true}` | The `init` event has no `memory_paths`: auto memory is off. With `disableAllHooks`, a SessionStart hook that a terminal wrapper (cmux) added to every `claude` it starts no longer ran: 2 hook events without it, 0 with it (2026-10-04, Claude Code 2.1.289). The driver also stops the run if any hook event appears in the stream. A repository `AGENTS.md` with a canary instruction was not followed. |
39
+ | `--strict-mcp-config --mcp-config {"mcpServers":{}}` | The `init` event lists no MCP server. |
40
+ | `--disable-slash-commands` | The `init` event lists no skill and no slash command. |
41
+ | `--no-session-persistence` | Nothing is saved for a later `--resume`; the correction rounds use the open process instead. |
42
+
43
+ `AGENTS.md`: a canary there was not loaded in any run, with or without settings. The built-in plugins (`cc-plugin-agents-md`, `cc-plugin-telemetry`, `cc-plugin-plugin-authoring`) stay listed; they are part of Claude Code and read no repository instruction file in this configuration.
44
+
45
+ ### Where it works
46
+
47
+ - Started from a Bash tool inside a running Claude Code session: works, with the same tools and the same refusals.
48
+ - Started from a plain environment (`PATH`, `HOME`, `USER`, `LOGNAME`, `TMPDIR` and the developer's own `CLAUDE_CONFIG_DIR`): works. Without `USER` the login in the macOS keychain is not found and the run ends with "Not logged in".
49
+ - A temporary `CLAUDE_CONFIG_DIR` loses the login, so the driver keeps the developer's own configuration folder and excludes its contents with the flags above.
50
+
51
+ ### Detecting it
52
+
53
+ - `claude --version` prints the version.
54
+ - `claude auth status` prints JSON with `loggedIn`. Exit 1 and `"loggedIn": false` mean the reviewer cannot start.
55
+
56
+ ### Usage
57
+
58
+ Each `result` event carries `num_turns` and `usage` for that turn (input, output and cache tokens), and `total_cost_usd` and `modelUsage` for the session so far. The driver adds up the turns and takes tokens and cost from the last `result` event.
59
+
60
+ ### The boundary and the alarm
61
+
62
+ The boundary is Claude Code's own permission rules: `--tools Read,Grep,Glob` and `--permission-mode dontAsk` with no settings source, which refused every read outside the working folder in the runs above. The alarm is the tool's own check of the event stream (`packages/cli/src/reviewers/trace.ts`), which does not trust the boundary and fails closed:
63
+
64
+ - the run fails when the `init` event lists any tool beyond Read, Grep and Glob, any MCP server or a memory path; the `Agent` tool is never listed, so no subagent or nested turn can make a call the stream does not show, and every `tool_use` in the stream is checked whichever turn it came from;
65
+ - every tool call counts from the moment the agent asks for it, with or without a result; a tool name other than the three makes the review incomplete;
66
+ - every path-bearing input (`file_path`, `path`, `notebook_path`, `cwd`, `directory`, and a `pattern` or `glob` that starts at `/`, `~`, a drive or `..`) is resolved against the snapshot, then through the real path of its deepest existing folder, and compared case-insensitively on macOS and Windows; a path with `$`, `%` or a NUL is refused, `~` is the home folder;
67
+ - an input that is not an object, or a path field that is not text, makes the review incomplete;
68
+ - any attempt outside the snapshot makes the review incomplete, even one the agent refused.
69
+
70
+ The snapshot holds no links (they are written as plain files) and secrets the scanners found are redacted in every file of it before the reviewer starts; a file too large to check is removed from it.
71
+
72
+ ### What the agent stores
73
+
74
+ With `--no-session-persistence` and auto memory off, real runs with Claude Code 2.1.289 left no transcript, no `history.jsonl` line and no project entry for a snapshot folder in the configuration folder (searched for the brief's text and the snapshot paths after the runs). An earlier run without `autoMemoryEnabled: false` left one empty `projects/<folder>/memory` folder; with the flag, none. The driver keeps the developer's configuration folder because a temporary one loses the login.
75
+
76
+ ### Not covered
77
+
78
+ - Managed (policy) settings set by an organisation still apply; they can add hooks or permission rules. The trace check above still fails a run that reads outside the snapshot.
79
+ - Each new Claude Code version can change these flags. Re-run these checks before raising the tested version.
80
+
81
+ ## Codex
82
+
83
+ Enabled since 0.6.0, with two stated limits. Tested with codex-cli 0.160.0 (`/opt/homebrew/bin/codex --version`) on macOS, 2026-10-03 and 2026-10-04, logged in with a ChatGPT account, on throwaway folders under `~/.openqodex/` with canaries. The driver is `packages/cli/src/reviewers/codex.ts`. It refuses a Codex older than 0.160.0, and a Codex whose `--version` prints no version number.
84
+
85
+ ### The command line
86
+
87
+ The driver starts this command without a shell, with the snapshot as the working directory, and writes the prompt to standard input. Its environment comes from an allowlist (`codexEnv`): `PATH`, `HOME`, `USER`, `LOGNAME`, `SHELL`, `TMPDIR`, locale, `TERM`, `TZ`, `CODEX_HOME`, `CODEX_CA_CERTIFICATE`, `SSL_CERT_FILE` and proxy settings, plus `OPENQODEX_REVIEW_DEPTH=1`. No `OPENAI_API_KEY`, no `CODEX_THREAD_ID` and no `CODEX_SANDBOX` reach it. Only absolute `PATH` entries are kept: an empty or relative one would resolve inside the snapshot.
88
+
89
+ ```
90
+ codex exec --json --color never --ephemeral --skip-git-repo-check
91
+ --ignore-user-config --ignore-rules
92
+ -C <snapshot>
93
+ -c approval_policy="never"
94
+ -c default_permissions="openqodex_review"
95
+ -c permissions.openqodex_review.filesystem={":minimal"="read",":project_roots"="read","/tmp"="deny"}
96
+ -c web_search="disabled" (web_search="cached" unless reviewer_web: off)
97
+ -c project_doc_max_bytes=0
98
+ -c allow_login_shell=false
99
+ -c shell_environment_policy.inherit="core"
100
+ -c skills.include_instructions=false -c skills.bundled.enabled=false
101
+ --disable plugins --disable apps --disable hooks --disable multi_agent --disable memories
102
+ --disable browser_use --disable computer_use --disable image_generation --disable skill_search
103
+ --disable tool_suggest --disable goals --disable in_app_browser --disable view_image
104
+ -
105
+ ```
106
+
107
+ `codex sandbox -c ... -- <command>` runs one command under the same sandbox without a model, and `codex debug prompt-input -c ...` prints what the model would be given; both cost nothing and were used for most checks below.
108
+
109
+ ### What each part was observed to do
110
+
111
+ | Part | Observed |
112
+ |---|---|
113
+ | `exec --json` | One JSON event per line: `thread.started`, `turn.started`, `item.started` and `item.completed` for messages, commands and web searches, `turn.completed` with `usage` (`input_tokens`, `cached_input_tokens`, `cache_write_input_tokens`, `output_tokens`, `reasoning_output_tokens`). `cached_input_tokens` is part of `input_tokens`. A run can print more than one `agent_message`; the last one is the answer. The run ends after one answer. |
114
+ | `--ephemeral` | No rollout file was written for the test folder. A follow-up in the same session is not possible, so each correction round is a new run that carries the brief, every earlier answer and every earlier correction, in order. A subagent cannot start: `spawn_agent` failed with "no rollout found for thread id". |
115
+ | `-s read-only` alone | Writes and network were refused, but every read was allowed: `cat` of a file in another `/tmp` folder, of `/etc/hosts` and `ls ~/.ssh` all succeeded. So the driver does not use it. |
116
+ | the `openqodex_review` permission profile | Reads outside the folder were refused (`Operation not permitted`): a file in a home folder outside it, `~/.ssh`, `~/.codex`, `~/Projects`, `$TMPDIR`, and `/tmp` with the `/tmp` deny entry. Reads inside it worked. `/etc` and the system folders `:minimal` names stay readable; without `:minimal` no command could start. `touch` was refused. `curl` could not resolve a host. The patch tool was refused ("writing is blocked by read-only sandbox"). Programs outside those folders do not start: in a real review `rg` was "command not found", and the model used `ls`, `find` and `sed`. |
117
+ | `web_search="disabled"` | The model reported no web search tool. |
118
+ | `web_search="cached"` | Asked to search, the model ran one search; the stream showed a `web_search` item with the query and its results (2026-10-04). Shell commands still had no network: the profile has no network entry. |
119
+ | `project_doc_max_bytes=0` | A canary `AGENTS.md` in the folder did not appear in the prompt input or the answer. |
120
+ | `skills.include_instructions=false` | Without it, a canary skill in the folder's `.agents/skills/` was listed to the model, which then followed it. With it, the skills block is gone from the prompt input. |
121
+ | `--ignore-user-config` | The developer's `config.toml` (MCP servers, plugins, model, trusted projects) is not read. A canary `developer_instructions` in the folder's `.codex/config.toml` did not appear either way: the folder is not a trusted project. |
122
+
123
+ ### The per-run probe of the boundary
124
+
125
+ The read confinement rests on two `-c` keys (`default_permissions` and `permissions.openqodex_review.filesystem`). A newer Codex could rename or ignore them, and the event stream would not show it. So every review proves the boundary before the first model run (`probeSandbox`). It runs one command, with no model, under `codex sandbox` with the same two keys, in the snapshot folder:
126
+
127
+ - it reads a canary file that openqodex writes in its home folder (`~/.openqodex/.openqodex-probe-<random>`, mode 0600, random content), outside the snapshot;
128
+ - it reads a file with random content that openqodex writes inside the snapshot;
129
+ - it tries to create a file inside the snapshot;
130
+ - it prints a random marker as its last act.
131
+
132
+ The script runs every program by absolute path (`/bin/cat`) with `PATH=/usr/bin:/bin`, so a program committed in the snapshot cannot stand in for one. It handles each expected refusal itself, so the marker prints only when every step ran. The review starts only when the probe exits 0 with no signal, the marker came back, the inside read worked, the canary's content did not come back and no file was created. Any other result, a probe that cannot start, or one that runs past 30 seconds ends the run as "Full review unavailable" with "Codex's sandbox did not confine reads to the review copy; the review did not start" and the `review --agent` fallback. The canary and both probe files are removed whatever happened, before the snapshot is hashed. A Ctrl-C or a kill during the probe ends the probe's process group and removes the files before `review` exits. `codex sandbox` passes on the command's exit status (3 for `exit 3`, 137 for a killed shell).
133
+
134
+ Observed with codex-cli 0.160.0 (2026-10-04):
135
+
136
+ - With the review profile, `cat` of the canary printed "Operation not permitted", `cat` of the inside file printed its content, and the write printed "Operation not permitted".
137
+ - With a profile that adds `"/"="read"`, the canary's content came back, so the probe refuses it. `packages/cli/test/codex-stream.test.ts` runs both against the real binary when Codex is installed (skipped in CI).
138
+ - `codex sandbox -c default_permissions="nope"` stops with "default_permissions refers to undefined profile `nope`". With the key misspelled it stops with "config defines `[permissions]` profiles but does not set `default_permissions`". Both print no canary content, so the probe refuses them.
139
+ - `codex sandbox` takes the same `-c` keys but has no `--ignore-user-config`: the probe reads the developer's `config.toml`, while `codex exec` does not. The `-c` keys override the same keys in that file. The probe proves that this Codex binary applies these keys to a sandboxed command; it runs through `codex sandbox`, not through `codex exec` itself, which cannot run a command without a model.
140
+ - Each probe took well under a second.
141
+
142
+ ### The two limits
143
+
144
+ 1. The developer's global instructions are loaded. `~/.codex/AGENTS.md` (or `$CODEX_HOME/AGENTS.md`) appeared in the prompt input with every flag above, and the model quoted its first sentence. No configuration key removed it (`instructions`, `user_instructions`, `agents_md.enabled`, `include_agents_md`, `features.agents_md` were tried). Only a different `CODEX_HOME` leaves it out, and that moves the login: a copy of `auth.json` would refresh its token on its own and can leave the developer's real login with a used refresh token. The driver keeps the developer's `CODEX_HOME`. In one real run the model looked for `CLAUDE.md` and `AGENT.md` files in the snapshot because the global file told it to.
145
+ 2. The event stream does not show every command. Every current model in the catalog (`codex debug models`) has `tool_mode: code_mode_only` except gpt-5.5: the shell is a nested tool inside a code tool. In one run on 2026-10-03, two shell commands ran (their output came back in the answer) and no `command_execution` event appeared in the stream. In the runs on 2026-10-04 every command did appear. Nothing guarantees it, so the driver says `traced: false`.
146
+
147
+ What `traced: false` changes in the run (`packages/cli/src/review-run.ts`, `packages/core/src/completion.ts`):
148
+
149
+ - No read in the stream counts as coverage. A changed range counts only when its diff is in the brief or the run sent it in a correction round. A range still not sent after two rounds makes the review incomplete ("not given to the reviewer").
150
+ - The commands and searches the stream shows are kept in `trace.json` with `inside: null` and their input under `detail`. They never pass or fail a review: there is no "read outside the snapshot" alarm and no "tool it was not given" check. The sandbox is the boundary.
151
+ - The completion record holds `trace_complete: false`, and empty `files_read` and `files_not_read`. The report prints "Files read: not recorded by Codex" and "Reads outside the snapshot: not recorded by Codex".
152
+
153
+ Read confinement held in every test: the permission profile is a real boundary, stronger than `-s read-only`. The code tool's own JavaScript runtime has no file or network access (`require`, `import("node:fs")` and `fetch` were all undefined or refused).
154
+
155
+ ### Detecting it
156
+
157
+ - `codex --version` prints `codex-cli 0.160.0`.
158
+ - `codex login status` prints "Logged in using ChatGPT" and exits 0 when logged in. With an empty `CODEX_HOME` it prints "Not logged in" and exits 1. The driver takes exit 0 as logged in.
159
+ - A Codex session sets these variables for the commands it runs (seen in a `codex exec` run, 2026-10-04): `CODEX_THREAD_ID` and `CODEX_SESSION_ID` (the thread id), `CODEX_VERSION` and `CODEX_CI=1`. Under a sandbox it also sets `CODEX_SANDBOX=seatbelt`, and `CODEX_SANDBOX_NETWORK_DISABLED=1` when network is off. `hostAgent` takes `CODEX_THREAD_ID` as "running inside Codex". The interactive Codex was not run for this; the binary holds the same name.
160
+
161
+ ### Where it works
162
+
163
+ - Started from a Bash tool inside a running Claude Code session, with the allowlist environment: it works.
164
+ - Started from inside a Codex sandbox (`codex sandbox -- codex exec ...`), with network off, and again with `workspace-write` and network on: `codex exec` exits 1 at once with "Error: failed to initialize in-process app-server client: Operation not permitted (os error 1)" and prints no event. `codex --version` and `codex login status` still work there. So `detect` reports Codex as unavailable whenever `CODEX_SANDBOX` is set, and `review` prints "Full review unavailable" with the `review --agent` fallback instead of starting a run that cannot answer.
165
+
166
+ ### Usage
167
+
168
+ Each `turn.completed` event carries `usage`. The driver adds `input_tokens` and `output_tokens` over all runs of a review and counts one turn per run. A ChatGPT login has no price per run, so the report shows no cost. The test runs on 2026-10-03 used 51,000 to 92,000 input tokens (most of them cached) and 600 to 950 output tokens per run, in 45 seconds or less. The web search run on 2026-10-04 used 52,911 in and 207 out. A real review of a six-line change with one planted SQL injection took 31 seconds, one run, 50,384 tokens in and 641 out, and reported the injection (2026-10-04).
169
+
170
+ ### What would remove the limits
171
+
172
+ A switch that leaves out `$CODEX_HOME/AGENTS.md` without moving the login, and an event for every command the code tool runs. Re-run the checks above on each new Codex version before raising the tested version.
173
+
174
+ ## Cursor
175
+
176
+ Not enabled. `cursor-agent` 2025.09.18-7ae6800 was on this Mac and not logged in. Its help shows `-p` ("Has access to all tools, including write and bash"), `--output-format stream-json`, `--model`, `--force` and `--resume`: no option to limit its tools, no read-only sandbox, no switch to skip the repository's rules or the developer's settings. None of the checks above can pass with those options, so no invocation was built. Cursor users get the full review through Claude Code or Codex when one is installed.
package/docs/llms.txt CHANGED
@@ -7,6 +7,7 @@
7
7
  - [FAQ](faq.md)
8
8
  - [GitHub Action](github-action.md)
9
9
  - [OpenQodex docs](index.md)
10
+ - [Reviewer drivers](internal-reviewer-drivers.md)
10
11
  - [Plumbing commands](plumbing.md)
11
12
  - [Quickstart](quickstart.md)
12
13
  - [Scanners](scanners.md)
package/docs/plumbing.md CHANGED
@@ -4,13 +4,29 @@
4
4
 
5
5
  ## scan
6
6
 
7
- Plain `openqodex review` does the same. Kept under this name for the git pre-push hook of earlier releases, the pre-commit hook and the GitHub Action, which call it.
7
+ For machines. Kept for the pre-commit hook and the GitHub Action, which call it, and for the git pre-push hook of earlier releases.
8
8
 
9
9
  ```
10
- openqodex scan [--base <ref>] [--uncommitted] [--only <list>] [--skip <list>]
10
+ openqodex scan [--base <ref>] [--uncommitted] [--only <list>] [--skip <list>] [--block-on-severity <severity>]
11
11
  ```
12
12
 
13
- Runs the scanners on the change and prints the report. No model is involved. The git hook, the pre-commit hook and the GitHub Action run this command.
13
+ Runs the scanners on the change and prints their findings, labelled as scanner data. No model is involved and nothing is checked: it is not a review. The pre-commit hook and the GitHub Action run this command. `--block-on-severity` sets the severity that makes it exit 1, and wins over `review.block_on_severity` in the config.
14
+
15
+ ## review --agent and review --finalize
16
+
17
+ The two-step protocol of earlier versions, kept so a skill installed before the one-command review keeps working, and the fallback `review` names when no reviewer can start (only Codex or only Cursor installed, or Claude Code logged out). The brief `review --agent` prints carries the whole procedure. New skills, rules and permission rules no longer name it.
18
+
19
+ ```
20
+ openqodex review --agent [--all | <target>] [...]
21
+ openqodex review --finalize [--run <id> | path]
22
+ ```
23
+
24
+ - `review --agent`: run the scanners, write the brief and print it for the agent running the command.
25
+ - `review --finalize [path]`: check the agent's findings file and write the report. Without a path it reads `agent-findings.json` in the newest report folder. With a path it finds the run by the `change_id` in that file. `--run <id>` names the run folder instead; a review of a branch or a pull request is finalized only that way.
26
+
27
+ `--finalize` exits 2 when the findings file breaks the shape (naming the first wrong field), the change or the config moved since the brief, a finding cites a scanner rule or candidate that is not in this scan, the brief was written by another openqodex version that is not installed in `~/.openqodex/runtime/`, or, for a branch or a pull request, the temporary checkout moved from the reviewed commit or is gone. When the launcher started the review, the brief's finalize command is the plain line `<launcher> review --finalize`, with `--all` and `--offline` as the review had them. When the version that runs `--finalize` is not the one that wrote the brief, and that one is installed, it hands the run to that version. It never repairs a finding.
28
+
29
+ A review finished this way is a legacy review: the agent that ran it reviewed the change itself. Its report says so on the first line after the verdict, in every format ("Reviewed by the coding agent you are using."; `reviewed_by` in `report.json`, a run property in `report.sarif`). Finalize writes a legacy record to `~/.openqodex/receipts/`, and the push hooks accept it as reviewed, with a line naming who reviewed; it never counts as a complete record.
14
30
 
15
31
  ## doctor
16
32
 
@@ -43,12 +59,12 @@ openqodex hook install [--force]
43
59
  openqodex hook uninstall
44
60
  ```
45
61
 
46
- - `hook check`: the push gate. The Claude Code and Codex hooks call it before a shell command. It reads the hook's JSON on stdin. It always exits 0.
47
- - `hook install`: add a git pre-push hook to this repository. It also sets up the launcher in `~/.openqodex/`, which the hook calls. The hook runs `hook pre-push`, which scans each commit the push sends against the remote's tip of its branch (`agents` has the details). It stops the push only when the scan exits 1. A scan that fails for its own reasons never stops the push.
62
+ - `hook check`: the push gate. The Claude Code and Codex hooks call it before a shell command. It reads the hook's JSON on stdin and looks up the review of exactly the change being pushed. It always exits 0.
63
+ - `hook install`: add a git pre-push hook to this repository. It also sets up the launcher in `~/.openqodex/`, which the hook calls. The hook runs `hook pre-push`, which does the same lookup for each commit the push sends (`agents` has the details). It prints no scanner findings and never starts a review. It stops the push (exit 1) only when the config sets `block_on_severity` and the review is missing or blocked. A lookup that fails for its own reasons never stops the push.
48
64
  - `hook install` refuses to replace a hook it did not write. `--force` replaces it and keeps the old hook as `pre-push.openqodex.bak`.
49
65
  - `hook uninstall`: remove that hook and put back the one it replaced. A hook you edited after install is left in place.
50
66
 
51
- When the repository uses husky or lefthook, `hook install` writes nothing. It prints the line to add to their pre-push hook: `npx -y openqodex@<version> hook pre-push || [ $? -ne 1 ]`. The part after `||` makes the line stop the push only on exit 1, as the hook `hook install` writes does: a scan that fails for its own reasons (exit 2) never stops the push.
67
+ When the repository uses husky or lefthook, `hook install` writes nothing. It prints the line to add to their pre-push hook. For husky it is `npx -y openqodex@<version> hook pre-push "$@" || [ $? -ne 1 ]`, so the hook gets the remote's name, which picks the base for a new branch. For lefthook the line has no `"$@"`: lefthook puts git's hook arguments into its command line as raw text, so a remote URL could carry shell code into it. Without the arguments, the hook looks the push up against `origin`, so a push to another remote is checked as if it went to `origin`. The part after `||` makes the line stop the push only on exit 1, as the hook `hook install` writes does: a lookup that fails for its own reasons (exit 2) never stops the push.
52
68
 
53
69
  `init` asks whether to install the git hook. `agents` explains the push gate.
54
70
 
@@ -12,6 +12,7 @@ The agent installs the skill, runs the review and tells you the result. The step
12
12
  ## Before you start
13
13
 
14
14
  - Node 22 or newer, and git.
15
+ - Claude Code or Codex, installed and logged in. It is the reviewer OpenQodex starts. Without either, `review` runs the scanners, says "Full review unavailable", and names the command with which the agent you are in reviews the change itself.
15
16
  - macOS or Linux. On Windows, use WSL.
16
17
  - A git repository with a change in it.
17
18
 
@@ -27,12 +28,14 @@ npx openqodex init
27
28
 
28
29
  Inside a repository, `init` also:
29
30
 
30
- - asks whether to add the git pre-push hook, so every push from that repository gets a scan, from an agent or by hand. The default is yes.
31
- - adds a short section to each agent's instruction file, such as `~/.claude/CLAUDE.md` for Claude Code: when a feature or fix is done, review it with openqodex in a separate subagent, so the agent that wrote the code does not judge its own work. It prints the section before writing it.
31
+ - asks whether to add the git pre-push hook, so every push from that repository is checked for a review, from an agent or by hand. The default is yes.
32
+ - adds a short section to each agent's instruction file, such as `~/.claude/CLAUDE.md` for Claude Code: when a feature or fix is done, review it with openqodex. It prints the section before writing it.
32
33
  - creates `.openqodex/config.yaml` and `.openqodex/custom-instructions.md`. Commit both. Write in `custom-instructions.md` what a reviewer of your repository must know: conventions, what never to flag, what always to check. The review brief carries it word for word.
33
34
 
34
35
  `init` also starts the scanner downloads that your repo needs, in the background. Running it outside the agent matters: some agents run commands in a sandbox that cannot download.
35
36
 
37
+ Last, `init` reviews: when the repository has a change, it runs `openqodex review` and prints the report. When it has none, it asks what to review: the whole repository, a pull request, a branch, or not now. With `--yes` or without a terminal it prints the three commands instead of asking. `--no-review` skips this step. The review uses the scanners already installed and never fails `init`.
38
+
36
39
  Codex only: open Codex, run `/hooks` and trust the OpenQodex hook. Codex runs a new hook only after you trust it.
37
40
 
38
41
  ## 2. Ask for a review
@@ -43,7 +46,7 @@ Say to your agent:
43
46
  review my change with openqodex
44
47
  ```
45
48
 
46
- The agent hands the review to a separate subagent where it can, and tells you when it cannot. The reviewer runs `openqodex review --agent`. That command works out the change, runs the scanners and prints a brief. The agent verifies each scanner finding, reviews the change itself, and writes its findings to a file. Then it runs `openqodex review --finalize`, which checks those findings without a model and writes the report.
49
+ The agent runs `openqodex review` and shows you the report it prints. That one command works out the change, copies it to a temporary folder, runs the scanners and the code graph, and starts its own reviewer: a separate Claude Code or Codex process that reads that copy. The reviewer checks every scanner finding and is given every changed line; a script checks its answer, and OpenQodex prints the report. It takes one to three minutes and uses your Claude Code or Codex plan. You can run the same command in your terminal.
47
50
 
48
51
  To review the whole repository instead of one change, say:
49
52
 
@@ -51,7 +54,7 @@ To review the whole repository instead of one change, say:
51
54
  review my whole repo with openqodex
52
55
  ```
53
56
 
54
- The agent runs `openqodex review --all --agent`. The scanners check every file, and the brief tells the agent where to start: the most-called functions and the files with the most scanner hits. See `docs/cli.md` for the details.
57
+ The agent runs `openqodex review --all`. The scanners check every file, and the brief tells the reviewer where to start: the most-called functions and the files with the most scanner hits. See `docs/cli.md` for the details.
55
58
 
56
59
  To review a teammate's branch or a pull request before it merges, without leaving your own work, say:
57
60
 
@@ -60,11 +63,11 @@ review the branch feature/login with openqodex
60
63
  review pull request #42 with openqodex
61
64
  ```
62
65
 
63
- The agent runs `openqodex review --agent feature/login` or `openqodex review --agent '#42'`. OpenQodex fetches the branch or the pull request, checks it out in a temporary folder and reviews what it added since it left its base. Your working folder is not touched. See "Reviewing a branch or a pull request" in `docs/cli.md`.
66
+ The agent runs `openqodex review feature/login` or `openqodex review '#42'`. OpenQodex fetches the branch or the pull request, checks it out in a temporary folder and reviews what it added since it left its base. Your working folder is not touched. See "Reviewing a branch or a pull request" in `docs/cli.md`.
64
67
 
65
68
  ## 3. Read the report
66
69
 
67
- The agent tells you the verdict and the most serious findings. The full report is in `.openqodex/reviews/<time>-<id>/report.md` in your repo. `.openqodex/.gitignore` keeps the reports out of git; `git status` shows only the two files above and that `.gitignore`, the first time.
70
+ The agent shows you the report as OpenQodex printed it: the verdict, the counts, and for each finding where it is, the problem, why it matters and the fix. A complete review means every stage ran, every scanner finding was checked and every changed line was put in front of the reviewer; anything not covered is named. It does not mean nothing was missed. The same report is in `.openqodex/reviews/<time>-<id>/report.md` in your repo. `.openqodex/.gitignore` keeps the reports out of git; `git status` shows only the two files above and that `.gitignore`, the first time.
68
71
 
69
72
  The verdict is `passed` unless `.openqodex/config.yaml` sets `review.block_on_severity` and a finding meets it. With no config, OpenQodex warns and never blocks.
70
73
 
@@ -74,17 +77,17 @@ The verdict is `passed` unless `.openqodex/config.yaml` sets `review.block_on_se
74
77
  npx openqodex demo /tmp/openqodex-demo
75
78
  ```
76
79
 
77
- `demo` builds a small repo with planted bugs: a secret, a SQL injection, a bad Dockerfile, a vulnerable lockfile, a shell bug and a workflow injection. It scans the change and prints the report. Then open the folder in your agent and ask for a review.
80
+ `demo` builds a small repo with planted bugs: a secret, a SQL injection, a bad Dockerfile, a vulnerable lockfile, a shell bug and a workflow injection. It scans the change and prints the scanner report. Then run `openqodex review` in that folder, or open it in your agent and ask for a review.
78
81
 
79
82
  ## Without an agent
80
83
 
81
- `openqodex scan` runs the scanners on the change and prints the report:
84
+ Run the review yourself:
82
85
 
83
86
  ```
84
- npx openqodex scan
87
+ npx openqodex review
85
88
  ```
86
89
 
87
- It is the same check the git hook, the pre-commit hook and the GitHub Action run.
90
+ `openqodex scan` runs the scanners only and prints their findings unchecked. It is the check the pre-commit hook and the GitHub Action run; it is not a review.
88
91
 
89
92
  ## First run
90
93
 
package/docs/security.md CHANGED
@@ -25,11 +25,11 @@ npx openqodex trust
25
25
 
26
26
  The stored sha256 is checked against the project's checksum file when the project publishes one. Otherwise it is the hash of your first download. `custom-scanners` explains the difference.
27
27
 
28
- Agents that follow the OpenQodex skill are told never to run `openqodex trust` without asking you. In user scope, `init` adds rules so Claude Code runs exactly `review --agent`, `review --finalize`, `review --agent --all` and `review --finalize --all` (each also with ` --offline`), `guide` and `guide <topic>` through the launcher without asking. An `ask` or `deny` rule in your own or your organisation's managed Claude Code settings still wins over these. Any other flag, any other command (`scan`, `doctor`, `trust`, `update`, `init`, `report`) and `init --project` grant nothing.
28
+ Agents that follow the OpenQodex skill are told never to run `openqodex trust` without asking you. In user scope, `init` adds rules so Claude Code runs exactly `review` and `review --all` (each also with ` --offline`), `guide` and `guide <topic>` through the launcher without asking. It removes the rules for the older two-step lines (`review --agent`, `review --finalize`) that an earlier `init` added. A review of a branch or a pull request names its target, so Claude Code asks before each one. An `ask` or `deny` rule in your own or your organisation's managed Claude Code settings still wins over these. Any other flag, any other command (`scan`, `doctor`, `trust`, `update`, `init`, `report`) and `init --project` grant nothing.
29
29
 
30
30
  ## What is sent where
31
31
 
32
- OpenQodex and the built-in scanners send no code anywhere. The review runs on the model your agent already uses, which sees what the agent reads. A custom scanner you approved does whatever its own command does.
32
+ OpenQodex and the built-in scanners send no code anywhere. The review runs on the model your Claude Code or Codex login uses: the reviewer process sends it the brief and what the reviewer reads (see "The reviewer process"). By default the reviewer can also search the web, and Claude Code can open web pages; `reviewer_web: off` in `~/.openqodex/config.yaml` removes that. A custom scanner you approved does whatever its own command does.
33
33
 
34
34
  OpenQodex and the built-in scanners use the network for these things only:
35
35
 
@@ -37,6 +37,7 @@ OpenQodex and the built-in scanners use the network for these things only:
37
37
  - Semgrep rule packs. semgrep fetches `p/default`, `p/security-audit` and `p/secrets` from the Semgrep registry on each run. Its metrics are off. The rules are never bundled in the package.
38
38
  - The dependency check. When the change holds a lockfile, osv-scanner sends the names and versions of the dependencies in it to osv.dev. It never sends code.
39
39
  - Custom scanners. `openqodex trust` reads the release from the GitHub API and downloads the asset. After approval, a custom scanner does whatever its own command does.
40
+ - The problem report, only when you choose it. When OpenQodex fails, a scanner breaks, or you run `openqodex report`, it prints the GitHub issue it would create and two choices. Nothing is sent unless you press 1 or run `openqodex report --send-last`. Then, when the GitHub CLI `gh` is installed and signed in, `gh` creates the issue in `openqodex/openqodex` with your GitHub account. Otherwise OpenQodex opens GitHub's new-issue page in your browser, or prints its link, with the title and body filled in, and you submit it there. The issue holds the OpenQodex version, the command and its arguments with paths and secrets taken out, the part of OpenQodex that failed, a scrubbed error line, the scanner statuses and your platform (OS, CPU type, Node major version). It never holds code, file names, paths, repository names, config or secrets.
40
41
  - The daily version check, for an install made with `init`. See "Updates" below.
41
42
  - A review of a branch or a pull request (`review <branch>`, `review '#<number>'`). git fetches the branch or `pull/<number>/head` from your remote with its own credentials, and `gh`, when it is installed and signed in, is asked for the pull request's base. OpenQodex reads no token. The target is checked out in `~/.openqodex/checkouts/`, a folder only you can open, with every git hook and filter switched off, so checking it out runs nothing from it, and a link in it becomes a small plain file. The scanners you approved for this repository do run on the target's files, with this repository's settings; one named only in the target's config never runs. If you review pull requests from people you do not trust, approve only custom scanners that do not execute the code they scan.
42
43
 
@@ -69,6 +70,38 @@ Updates are off with `openqodex update --off`, `update: off` in `~/.openqodex/co
69
70
 
70
71
  OpenQodex sends no telemetry. See `telemetry`.
71
72
 
73
+ ## The reviewer process
74
+
75
+ `openqodex review` starts Claude Code (`claude -p`) or Codex (`codex exec`) as its reviewer. Cursor is not used as a reviewer: with the version tested, it cannot be limited to reading. `docs/internal-reviewer-drivers.md` in the repository records the tests.
76
+
77
+ ### Claude Code
78
+
79
+ Claude Code sends the review brief and the files the reviewer reads to the model your Claude Code login uses, as any Claude Code session does. The reviewer:
80
+
81
+ - reads a snapshot of the change in `~/.openqodex/checkouts/`, never your folder. Secrets the scanners found are redacted in every file of the snapshot first, and a file too large to check is left out of it.
82
+ - has the read, search and list tools, plus Claude Code's WebSearch and WebFetch: no shell, no edits, no MCP server, no subagent. So the reviewer reads your code and can open web pages. A reviewer that reads private code and untrusted text from the change and can open web addresses can be talked into putting that code into a web address. `reviewer_web: off` in `~/.openqodex/config.yaml` removes the web tools; set it when you do not accept that risk. Claude Code's own permission rules refuse a read outside the snapshot; that is the boundary. OpenQodex also checks every tool call in the agent's event stream and marks the review incomplete when one names a path outside the snapshot, an unknown tool or an input it cannot read; that is the alarm.
83
+ - loads none of your Claude Code settings, hooks, plugins, memory or `CLAUDE.md` files, and none of the repository's.
84
+ - gets an environment built from a short allowlist: the variables Claude Code needs to run and find its login (`PATH`, `HOME`, `USER`, `CLAUDE_CONFIG_DIR`, proxy settings, `ANTHROPIC_*` keys, and cloud provider variables only when Claude Code is set to that provider). Other tokens in your shell, such as `GITHUB_TOKEN` or `NPM_TOKEN`, never reach it.
85
+
86
+ The reviewer runs with session saving off (`--no-session-persistence`). After real runs with Claude Code 2.1.289, no transcript, history line or project entry for a snapshot was found in the Claude Code configuration folder. Claude Code's own logs and telemetry follow its own settings.
87
+
88
+ ### Codex
89
+
90
+ Codex sends the conversation to the model your Codex login uses, as any Codex session does. The conversation holds the review brief, your global `~/.codex/AGENTS.md` and the output of each command the reviewer runs. Each correction round is a new Codex run that carries the whole conversation so far. The reviewer:
91
+
92
+ - reads the same redacted snapshot in `~/.openqodex/checkouts/`, never your folder.
93
+ - runs under a Codex permission profile: its commands can read the snapshot and the system folders Codex's `:minimal` set names (such as `/usr` and `/etc`), and nothing else. `/tmp`, your home folder, `~/.ssh` and `~/.codex` are refused. Writes and network are refused. That sandbox is the boundary. Before each review, OpenQodex proves it with a command run under the same profile without a model: a read of a file outside the snapshot and a write inside it must both be refused, or the review does not start.
94
+ - loads your global `~/.codex/AGENTS.md` (or `$CODEX_HOME/AGENTS.md`). No Codex setting leaves it out. If you keep instructions there, the reviewer sees them, including the section `init` adds for Codex.
95
+ - loads none of your `config.toml`, rules, MCP servers, plugins, hooks, memories or skills, and none of the repository's `AGENTS.md` or skills.
96
+ - has Codex's web search by default, as the cached search, which answers from OpenAI's search index and opens no address the model names. `reviewer_web: off` removes it.
97
+ - gets an environment built from a short allowlist: `PATH`, `HOME`, `USER`, `CODEX_HOME`, proxy and certificate settings. Other tokens in your shell, including `OPENAI_API_KEY` and `GITHUB_TOKEN`, never reach it. Its commands see a smaller set still (Codex's `core` environment).
98
+
99
+ There is no alarm for Codex. Its event stream does not show every command it runs, so OpenQodex cannot check from it which files were read. The commands it does show are kept in the run folder as a list for you to read; they never pass or fail a review. Coverage counts only the changed lines in the brief and those OpenQodex sent in a correction round, and the report says file reads were not recorded by Codex.
100
+
101
+ Codex runs with `--ephemeral`: after real runs with codex-cli 0.160.0, no session file was written for the snapshot folder. Codex's own logs follow its own settings.
102
+
103
+ The run folder of a review holds the brief, the scan, the reviewer's answer and the list of its tool calls: paths and line ranges for Claude Code, and the command lines Codex showed for Codex, never their output. Each file is created readable by you only, and secrets are redacted in all of them.
104
+
72
105
  ## Secrets
73
106
 
74
107
  When gitleaks finds a secret in the change, OpenQodex removes it from the brief, every report file and the terminal. It keeps the length and sha256 of each secret, to redact any text the agent quotes.
@@ -87,16 +120,18 @@ In your home folder, under `~/.openqodex/` (`OPENQODEX_HOME` moves it):
87
120
  - `runtime/<version>/` and `bin/openqodex`: the copy of the package and the launcher that the hooks call, written by `init`. Updates add copies beside it; a copy is never changed after it is written. `init` and `openqodex update` remove copies older than 7 days, except the one `init` installed, the current one and the previous one.
88
121
  - `runtime/current`: the version the launcher runs, and on a second line the version a rollback goes back to.
89
122
  - `update.json`: the state of the version check, private to you.
90
- - `config.yaml`: your own settings; today only `update`.
123
+ - `config.yaml`: your own settings: `update`, `reviewer` (which agent reviews) and `reviewer_web` (the reviewer's web tools, on by default; `off` removes them).
91
124
  - `install.json`: what `init` and `hook install` wrote, so an uninstall removes only that.
125
+ - `receipts/<repo id>/`: one small record per reviewed change, readable by you only, written by `review` at the end of a run (and by `review --finalize` for the older two-step protocol, only for a run whose scan this machine ran). The push hooks decide from these records only. The files under the repository's `.openqodex/` are the readable report, never the proof: a branch can carry those files, so a record found only there counts as no review. The check inside your agent is a reminder about your current work: it does not know what a push sends. For a plain `git push` it asks whether your current work has a passing review; any other push command it cannot tell, and says so (a deny when `block_on_severity` is set). The git pre-push hook that `init` offers is the check that sees the exact commits a push sends, and `git push --no-verify` skips it. `init` and `openqodex update` remove records older than 30 days.
126
+ - `runs/<repo id>/`: one record per `review --agent` run, readable by you only: the change and the hashes of the run files it wrote, so `review --finalize` can tell a run this machine scanned from one a branch carries. Removed with the receipts.
92
127
  - `trust.json`: your approvals of custom scanners.
93
128
 
94
129
  In the repository, under `.openqodex/` only:
95
130
 
96
131
  - `config.yaml` and `custom-instructions.md`: the team's config and instructions for the reviewer, created once and never touched after. They are meant to be committed.
97
132
  - `.gitignore`: keeps the run state below out of git, so after the first run `git status` shows only the two files above and the `.gitignore`.
98
- - `reviews/<time>-<id>/`: one folder per run, holding the brief, the scan result, the agent's findings and the reports. OpenQodex keeps the newest 20.
99
- - `latest.json`: points at the newest review; the push gate reads only this. `latest-scan.json` points at the newest scan.
133
+ - `reviews/<time>-<id>/`: one folder per run, holding the brief, the scan result, the reviewer's answer, the list of its tool calls and the reports. OpenQodex keeps the newest 20.
134
+ - `latest.json`: points at the newest review, for you and older tools; the push gate does not trust it (see `receipts/` above). `latest-scan.json` points at the newest scan.
100
135
 
101
136
  OpenQodex never reads or writes `.openqodex/` or the root `.openqodex.yaml` through a symbolic link, at the file or at any folder above it inside the repository. A link there stops the command with one line naming it, or, for a run file such as `latest.json`, counts as no file. Only regular files are read there, each within a size limit, so a link or a device in their place cannot hang a run.
102
137
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "openqodex",
3
- "version": "0.4.0",
3
+ "version": "0.6.0",
4
4
  "description": "Open source code review that runs inside your coding agent, before you push.",
5
5
  "type": "module",
6
6
  "license": "Apache-2.0",