diffninja 0.1.1 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,259 +1,115 @@
1
1
  # diffninja
2
2
 
3
- PR reviews for humans, not prose from a chatbot, run from inside your coding
4
- agent. Ask Claude Code, Codex, pi, or another MCP-capable agent CLI to review a
5
- GitHub PR link and diffninja returns a review workspace where you read the diff,
6
- write inline comments, and submit the review yourself. Give the agent a diff or
7
- a git range and diffninja returns an evidence-backed reading agenda alongside
8
- the complete diff, as structured data for the agent and as a page for you.
9
-
10
- diffninja is an MCP server (`review_diff`) plus a `diffninja setup` command that
11
- registers it. It has no terminal review mode.
12
-
13
- Analysis is local and deterministic: no model is called, no source leaves your
14
- machine, and the same input always gives the same report. Each hunk gets change
15
- facts — did a comparison, a limit, or an input check change; is a failure handed
16
- to the caller, deferred, or discarded — each pointing at the changed line it
17
- rests on. Interpreting what the change means is left to you and, if you ask it,
18
- to the agent you already use. diffninja writes no review prose. You stay the
19
- reviewer.
3
+ diffninja helps you review pull requests faster without handing the review to
4
+ a bot. You ask your coding agent (Claude Code, Codex, and others) to review a
5
+ PR. diffninja gives you a review page that shows the changes in the order worth
6
+ reading them, with the agent's notes next to the code. You read, comment, and
7
+ submit the review to GitHub yourself.
8
+
9
+ ## Why use it
10
+
11
+ - **Read the important changes first.** Big PRs are hard to follow file by
12
+ file. diffninja numbers each change, most important first, and lets you step
13
+ through them with `j` and `k`.
14
+ - **Your agent does the first pass.** For each change, the agent answers simple
15
+ questions: does it change behavior, is it tested, does it match the PR's
16
+ goal. The answers sit above the code.
17
+ - **Suggested comments, never posted for you.** The agent can suggest short
18
+ line comments. You add them with one click, edit them, or dismiss them.
19
+ Nothing reaches GitHub until you press Submit.
20
+ - **Facts you can check.** diffninja points at the exact lines that changed a
21
+ comparison, a limit, an input check, or error handling. It also shows call
22
+ flows: which functions call the changed code.
23
+ - **Private.** The analysis runs on your machine. diffninja calls no AI model,
24
+ needs no API key, and sends your code nowhere. It uses your existing GitHub
25
+ CLI login to read the PR and post your review.
20
26
 
21
27
  ## What you need
22
28
 
23
- - **Node.js 22.18 or newer.**
24
- - **An MCP-capable agent CLI:** Claude Code, Codex, OMP, pi, or any client that
25
- can launch a stdio MCP server.
26
- - **GitHub CLI (`gh`) 2.45.0+, authenticated** (`gh auth login`) — only for
27
- reviewing pull requests. diffninja never asks for a token; it reuses your
28
- `gh` session.
29
- - No API key. Git ranges may install missing parsing grammars through npm the
30
- first time a language is seen.
29
+ - **Node.js 22.18 or newer**
30
+ - **An agent CLI:** Claude Code, Codex, OMP, or pi (or any tool that supports
31
+ MCP servers)
32
+ - **GitHub CLI (`gh`) 2.45.0 or newer, logged in:** run `gh auth login` once.
33
+ Only needed for pull requests.
31
34
 
32
35
  ## Install
33
36
 
34
- Register the MCP server on every agent CLI you use (Claude Code, Codex, OMP,
35
- pi) with one command:
37
+ Run this once:
36
38
 
37
39
  ```bash
38
- npx -y diffninja setup
40
+ npx -y diffninja@latest setup
39
41
  ```
40
42
 
41
- It installs the package globally first, then registers the server in each
42
- detected CLI. `diffninja setup --help` lists the options (`--cli`,
43
- `--uninstall`, `--dry-run`, `--no-install`); `docs/mcp-setup.md` has the manual
44
- entries.
45
-
46
- The package ships `diffninja` (setup only) and `diffninja-mcp` (the MCP server).
47
- `npx` fetches the latest published version on first run; if a cached copy feels
48
- stale, pin it explicitly (`npx -y diffninja@latest setup`).
49
-
50
- On npm 12 and later, which block dependency install scripts by default, setup
51
- names the ones diffninja needs (`--allow-scripts=diffninja,tree-sitter,...`), so
52
- nothing extra is required. If you install by hand, pass the same flag:
53
- `npm install -g --allow-scripts=diffninja,tree-sitter,tree-sitter-javascript,tree-sitter-typescript diffninja`.
54
-
55
- ## Review a pull request
56
-
57
- Ask your agent to review `https://github.com/OWNER/REPO/pull/123`. It calls
58
- `review_diff` with the link, reads the pull request, sends its reading with
59
- `finish_review`, and then gives you a loopback review page, loaded from the
60
- canonical GitHub patch through your `gh` authentication. The link arrives
61
- once the agent has read the change, so the page opens complete: the agent's
62
- reading order, its answers, and its suggested comments. Select diff lines,
63
- write single-line inline comments and a review body, pick Comment, Approve, or
64
- Request changes, preview the exact payload, and submit. Comments your agent
65
- suggests wait under their lines until you add them; every word you submit is one
66
- you chose, and diffninja only carries it to GitHub. It never approves, blocks, or merges
67
- anything on its own. The page belongs to the agent's MCP connection and closes
68
- when the agent exits.
69
-
70
- The diff opens in the **reading order**: each hunk is a numbered change,
71
- most important first, with your agent's answers (does this change behavior,
72
- does a test exercise it, does it serve the stated goal) and the facts to look
73
- at above its code. Hunks outside the order follow under **Other changes**.
74
- Press `j` and `k` to step through the changes. A rail beside the diff lists
75
- them, marks the one you are reading, and fills in as you read. **By file**
76
- switches to one block per file, with each change's number where it starts,
77
- and keeps you on the change you were reading. When the agent works inside your local
78
- clone it passes it as `repo`; once the clone has the pull request's commits,
79
- the analysis adds definitions and call flows, and the page opens a file's call
80
- flow in a drawer beside the diff. diffninja never fetches or writes in the
81
- clone; when it lacks the commits, the page says so and names the `git fetch`
82
- that would add them.
83
-
84
- ## Review a diff or a git range
43
+ It installs diffninja and adds it to every agent CLI it finds on your machine.
44
+ Then restart your agent CLI.
85
45
 
86
- Ask your agent to review a patch, the working tree, or a range such as
87
- `main..HEAD` in a repository. It calls `review_diff` with `diff` text or with
88
- `repo`, `from`, and `to`, and receives the report as the tool result: the exact
89
- expected outcome when supplied (`expectedOutcome`), a short reading agenda,
90
- bounded automatic findings, explicit check coverage, and every hunk, ranked as
91
- **attention**, **uncertain**, **low**, or **passed**; git ranges add call flows
92
- and snapshot-bound source. Claims in a description are not proof that the code
93
- fulfills them. Once the agent has sent its reading (below), it gets `reportUrl`:
94
- a read-only page on `127.0.0.1` with the same report for you to read (the
95
- agenda, call-flow graphs, and every hunk), served from memory for as long as
96
- the agent's session lasts.
97
- The tool writes no report files. Arguments and examples:
98
- [docs/mcp-setup.md](docs/mcp-setup.md). Trying it on real reviews:
99
- [docs/pilot.md](docs/pilot.md).
46
+ To update later, run the same command again. To remove diffninja from your
47
+ agent CLIs, run `npx -y diffninja setup --uninstall`. Other options:
48
+ `npx -y diffninja setup --help`.
100
49
 
101
- ## How static analysis works
50
+ ## How to use it
102
51
 
103
- 1. **Evidence first.** No-op hunks and blank-only document changes pass.
104
- Repository snapshots support bounded checks for duplicate function bodies and
105
- unread `errors` response fields, plus caller and contract source cards.
106
- Optional TypeScript reference checking compares before/after diagnostics,
107
- including unchanged consumers, using an explicitly trusted installed compiler.
108
- Unsupported or incomplete checks say so.
109
- 2. **Change facts per hunk.** The added and removed lines are read lexically —
110
- strings and comments set aside, moved lines cancelling out. Code gets six facts:
111
- comparison changed, limit changed (a numeric bound, or `<` turned into `<=`),
112
- input check changed (type/shape checks, or a changed guard in front of a
113
- raise), failure handed to the caller, failure deferred or retried, failure
114
- discarded (an empty or defaulting `catch`, `except: pass`, …). Documentation
115
- (`.md`, `.rst`, `.txt`, …) is asked whether an instruction to readers changed
116
- (must, never, only, at most, …), a link target changed, or a numeric limit
117
- changed. Configuration (`.yml`, `.json`, `.toml`, Dockerfiles, `.env`, …) is
118
- asked whether a CI gate was weakened (`continue-on-error`, `|| true`, a
119
- failure turned into a warning, a check step removed), permissions or secret
120
- access changed, a version pin changed, or a limit changed. Each `yes` cites
121
- its changed line. `no` speaks only about the lines the hunk shows. A change
122
- that only touches formatting, comments, or line breaks is recognized as such.
123
- JS/TS, Java, C#, Go, Rust, C/C++, Kotlin, Swift, PHP, Python, Ruby,
124
- documentation and configuration are read; any other file type says that no
125
- facts were established instead of claiming none.
126
- 3. **Status and order.** A code or configuration change outside a test file
127
- reads **attention**; documentation reads **attention** when it changes an
128
- instruction, a link, or a limit, **low** otherwise; a test-file change reads
129
- **attention** only when it changes a limit, discards a failure, or weakens a
130
- gate, **low** otherwise; a formatting-only change **passed**;
131
- an unread file type **uncertain**, for a person to read. Priority orders hunks:
132
- a fixed base, plus 10 for a real change, plus the heaviest fact of the
133
- boundary group (what the change says or bounds) and of the failure group
134
- (failures, CI gates, permissions) — each group counts once, never summed. The report lists
135
- manual work first (binary and other metadata-only units), then the read hunks
136
- by priority, then passes; among equal priorities the hunk that changes more
137
- lines comes first. Hunks in test files (by path convention: `test/`,
138
- `*.test.ts`, `test_*.py`, `*_test.go`, …, including snapshots diffninja does
139
- not read) come after the other hunks, still ordered by their own priority: a
140
- regression test changes as much as its fix.
141
- Documentation is not demoted, because prose can be normative. Status is a
142
- label for filtering and never reorders the report.
143
- 4. **Questions for your agent.** Where a judgment needs meaning rather than
144
- syntax, the report asks the agent that requested it — does this hunk change
145
- what callers observe, does a test exercise it, does a test change weaken it,
146
- do the docs match the code, does the hunk serve the stated goal. Questions
147
- are fixed templates with closed options (always including `cannot-tell`).
148
- diffninja itself still calls no model.
149
- The agent sends its whole reading in one `finish_review` call: an answer to
150
- every question, the reading order of every hunk (most important first),
151
- and the line comments it would leave (or none). diffninja checks all of it
152
- and only then hands out the page link, so every page you open already
153
- carries the agent's answers, its order, and its comment decision; an agent
154
- cannot give you a half-read page. The pages list every hunk in the agent's
155
- order, attributed to it; diffninja's own order stays available one click
156
- away, and statuses stay diffninja's. `record_answers`, `record_order`, and
157
- `suggest_comments` update a review afterwards. On 159 held-out open-source
158
- pull requests, weighted by the severity of maintainers' actual review
159
- comments, a host model that read diffninja's report put the serious
160
- comments earlier than diffninja's deterministic order did.
161
- On a pull request, the suggested line comments are short, in the reviewer's
162
- own voice, with no "Finding 1:" scaffolding. The page shows each under its line; you add one
163
- or all of them to your draft with a click, edit or dismiss them, and submit
164
- the review yourself. Nothing is posted without you.
165
- 5. **Project context (git ranges only).** What a diff does not show is often
166
- the project around it. From the local clone alone — nothing is fetched —
167
- the report names the commits that last changed each hunk's removed lines
168
- (`git blame` at the base), earlier revert commits that touched a changed
169
- file or share a rare word with the goal or the changed file names,
170
- contributor guidelines that apply (`CONTRIBUTING`, `AGENTS.md`, `.github/`,
171
- docs policy pages such as versioning or preview rules), and, for a new
172
- file, identifiers most of its same-named siblings use and it does not
173
- (`components/*/select.py`). They add questions for your agent: does a hunk
174
- undo a fix it removes, does the change reintroduce something reverted, does
175
- it follow the guidelines and the sibling pattern. These are pointers, not
176
- verdicts, and never change status or order. A shallow clone says so: lines
177
- whose origin lies past its boundary count as unknown.
52
+ Inside your agent CLI, ask in plain words:
178
53
 
179
- Intent cross-checks keep author claims and generated summaries separate.
180
- Source matches are navigation evidence, not proof of fulfillment. Broad goals,
181
- missing metadata, and behavior not established by the available code remain
182
- explicitly unestablished. Findings likewise state their scope: duplicated syntax
183
- is not necessarily duplicated responsibility, and an unread field alone does not
184
- prove that a real failure was mishandled.
54
+ ```text
55
+ Review https://github.com/OWNER/REPO/pull/123 with diffninja
56
+ ```
185
57
 
186
- Git-range analysis supplies selected call-site blocks from both snapshots,
187
- including written arguments, declared parameters, locations, and explicit target
188
- and binding uncertainty. Unambiguous same-file JS/TS and Python calls support
189
- positional binding; Python also supports named arguments. Imports, member
190
- dispatch, dynamic targets, and other grammars do not get guessed mappings.
191
- These are static source expressions, not runtime values or data-flow analysis.
192
- TypeScript/TSX extraction includes methods of decorated exported classes,
193
- including stacked and custom decorators; method source spans retain their
194
- decorators so decorator-only changes still select the method body. Simple explicit class-field and
195
- constructor-property types identify candidate dependency methods; unsupported
196
- receiver types stay unresolved rather than borrowing the containing class's
197
- method. These are not type-checked or proven runtime bindings.
58
+ The agent reads the PR, then gives you a link to the review page (it runs on
59
+ your machine at `127.0.0.1`). On the page:
60
+
61
+ 1. Start with **Goal**: the reviewing agent's short, plain-English explanation
62
+ of what the PR is meant to do and its important limits. The full
63
+ **Original PR description** stays one click away, with Markdown formatting.
64
+ The goal is stated intent, not proof the code fulfills it.
65
+ `j` and `k` move to the next and previous change, and the list on the left
66
+ shows where you are. The agent sets the reading order; the connected page
67
+ does not label changes “Attention”.
68
+ 2. Hover a line and press **+** to write a comment, or add the agent's
69
+ suggestions.
70
+ 3. Write a summary, choose **Comment**, **Approve**, or **Request changes**.
71
+ 4. Press **Check the review** to see exactly what will be sent, then
72
+ **Submit review**.
73
+
74
+ Tip: if your agent is running inside a local clone of the repository, it can
75
+ pass the clone to diffninja. You then also get call-flow diagrams for the
76
+ changed files.
77
+
78
+ You can also review changes that aren't a PR yet:
79
+
80
+ ```text
81
+ Review my changes on this branch against main with diffninja
82
+ ```
198
83
 
199
- Review context also follows candidate event and queue relations: static
200
- `emit`/`emitAsync` keys match `@On*Event` handlers; injected queue `.add` keys
201
- match `@Processor` consumers' `job.name` cases or `@Process` methods only on a
202
- matching queue channel. String enum values can connect member keys to literal
203
- cases. These are source-derived relations, not runtime calls or proof of
204
- delivery; dynamic keys, aliases, and unrecognized framework syntax can be absent.
205
- Constant resolution is snapshot-local, including when extraction is cached.
84
+ The agent gets a report of every changed piece of code, ranked by importance,
85
+ and gives you a link to read the same report in your browser.
206
86
 
207
- Module-level TypeScript/TSX interfaces, type aliases, and enums are addressable
208
- non-callable context nodes. Signature, body, and generic type references connect
209
- them to reviewed methods and other contracts. Relative import bindings take
210
- precedence; unique module-path suffixes and unimported names are candidate
211
- matches, not type checking. Unresolved imports do not borrow same-named types.
212
- Barrel re-exports, nested declarations, and class-field contracts may be absent.
87
+ ## Good to know
213
88
 
214
- Context prioritizes calls adjacent to the hunk, with depth limited to four,
215
- at most eight arguments per call, and 120 characters per argument excerpt.
216
- Snapshot-bound changed definitions, callers, callees, dispatch endpoints, and
217
- type contracts carry complete source plus selected relation evidence. Bodies not
218
- already shown in the hunk take precedence over duplicate source. Unseen type
219
- declarations and dispatch endpoints precede ordinary caller chains, nearest
220
- first, with resulting-snapshot evidence ahead of prior-snapshot duplicates.
221
- Unresolved own-call expressions remain verbatim in a whole source body rather
222
- than repeating unknown-binding boilerplate. At most eight distinct definition
223
- nodes are kept per hunk, each carried whole, never shortened.
89
+ - The review page closes when you exit your agent CLI.
90
+ - The page and the agent's session contain source code. Treat them like the
91
+ code itself.
92
+ - diffninja never approves, blocks, or merges anything on its own.
224
93
 
225
- Automatic response/caller evidence is narrower than the candidate call graph:
226
- it requires a supported lexical or typed-constructor binding. Relative modules
227
- and simple single-target `paths` aliases from snapshot-local, standalone JSON
228
- `tsconfig.json` files are supported. JSONC, inherited configurations, package or
229
- barrel resolution, and complex receivers remain unproven rather than borrowing
230
- an unrelated same-named function. `referenceProject: "path/to/tsconfig.json"`
231
- opts into the separate TypeScript diagnostic comparison; absent dependencies,
232
- unsupported project layouts, and exceeded bounds are reported as not checked.
94
+ ## More
233
95
 
234
- ## Good to know
96
+ - [How the analysis works](docs/how-it-works.md): what diffninja checks and
97
+ how it ranks changes
98
+ - [Manual setup and tool reference](docs/mcp-setup.md): for agent CLIs that
99
+ `setup` doesn't configure
100
+ - [Releasing](docs/npm-release.md): how new versions are published
235
101
 
236
- - Tool results embed source code, including unchanged code, and stay in your
237
- agent's session. Keep transcripts that contain them private.
238
- - Connected PR reviews need authenticated `gh`. Nothing else leaves your
239
- machine: no model is called and no API key is needed.
240
-
241
- ## Dev
102
+ ## Development
242
103
 
243
104
  ```bash
244
- npm run build # tsc -> dist/
245
- npm run lint # oxlint
246
- npm test # vitest run
105
+ npm install
106
+ npm run build # compile to dist/
107
+ npm run lint
108
+ npm test
247
109
  ```
248
110
 
249
- Run the built server directly with `node dist/review/mcp-cli.js` (it speaks MCP
250
- over stdio and prints nothing else to stdout).
251
-
252
- Releases and the npm publishing setup: [docs/npm-release.md](docs/npm-release.md).
253
-
254
111
  ## Credits
255
112
 
256
113
  The call-flow engine is a fork of
257
114
  [calldiff](https://github.com/tanishqkancharla/calldiff) by Tanishq Kancharla
258
- (MIT, see LICENSE). diffninja uses its call graphs to show which flows each
259
- hunk touches.
115
+ (MIT, see LICENSE).
@@ -11,7 +11,7 @@
11
11
  * MCP client that recorded it.
12
12
  */
13
13
  import { type QuestionKind, type Verdict } from "./questions.js";
14
- import type { ReviewReport, ReviewStatus, SuggestedComment } from "./types.js";
14
+ import type { AgentSummary, ReviewReport, ReviewStatus, SuggestedComment } from "./types.js";
15
15
  /** Most agenda entries the page lists; the full report has the rest. */
16
16
  export declare const CONNECTED_AGENDA_LIMIT = 5;
17
17
  export interface ConnectedFact {
@@ -70,13 +70,21 @@ export interface ConnectedAnalysis {
70
70
  readonly total: number;
71
71
  readonly answered: number;
72
72
  };
73
- /** Line comments the reviewing agent suggested, for the human to add to their review or not. */
74
73
  /** Changed files that have call-flow diagrams, for the page's "Call flow" buttons; empty for a patch-only analysis. */
75
74
  readonly callFlowFiles: readonly string[];
75
+ /** Line comments the reviewing agent suggested, for the human to add to their review or not. */
76
76
  suggestions?: {
77
77
  readonly suggestedBy: string;
78
78
  readonly comments: readonly SuggestedComment[];
79
79
  };
80
+ /**
81
+ * The reviewing agent's own short paragraph on what the pull request does and
82
+ * why, attributed to the client that wrote it. Absent when no finish_review
83
+ * accepted one, which the page must say instead of showing a goal of its own:
84
+ * diffninja generates no summary, and this one is the agent's reading of the
85
+ * author's stated intent, not a claim that the changes achieve it.
86
+ */
87
+ summary?: AgentSummary;
80
88
  }
81
89
  export type ConnectedOrder = {
82
90
  readonly source: "agent";
@@ -159,5 +159,10 @@ export function connectedAnalysisOf(report, snapshotId, reviewId, reportUrl, sco
159
159
  if (report.agentComments !== undefined) {
160
160
  analysis.suggestions = { suggestedBy: report.agentComments.suggestedBy, comments: report.agentComments.comments };
161
161
  }
162
+ // Only finish_review stores a summary, so this is present exactly when the
163
+ // agent's whole reading was accepted for this snapshot's report; it is copied
164
+ // verbatim, attributed, and never synthesized here.
165
+ if (report.agentSummary !== undefined)
166
+ analysis.summary = report.agentSummary;
162
167
  return analysis;
163
168
  }