@ataraxy-labs/sem 0.25.0 → 0.27.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +61 -0
- package/README.md +284 -141
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,67 @@ All notable changes to sem are documented in this file.
|
|
|
4
4
|
|
|
5
5
|
## [Unreleased]
|
|
6
6
|
|
|
7
|
+
## [0.27.0] - 2026-10-04
|
|
8
|
+
|
|
9
|
+
### Changed
|
|
10
|
+
|
|
11
|
+
- **A smaller command line: eight verbs, variations as flags.** `sem --help` now lists `find`, `grep`, `impact`, `check`, `certify`, `diff`, `graph` and `history`, the `cloud` and `config` groups, and `mcp`, one line each saying which question the verb answers, with a QUICKSTART for the four agent questions (where is it, what does my change touch, is it correct, what should a human review).
|
|
12
|
+
- `sem find X --callers | --refs | --context` and `sem find --in PATH` replace `callers`, `refs`, `context` and `entities`.
|
|
13
|
+
- `sem impact --diff <range>` takes a whole change; with `--tests` in a JS/TS workspace it uses the module graph's affected-test selection.
|
|
14
|
+
- `sem certify <range> --arch` (with `--html`, `--view`, `--md`) replaces `arch-diff`.
|
|
15
|
+
- `sem graph --modules | --dataflow [--witness] | --system` replaces `topology`, `dataflow` and `system`.
|
|
16
|
+
- `sem history X` and `sem history --blame FILE` replace `log` and `blame`.
|
|
17
|
+
- `sem cloud login|logout|whoami|review|xref|repos|enable|disable` and `sem config setup|unsetup|telemetry|completions|update|stats` group the account and setup commands.
|
|
18
|
+
- Every old command name and flag still works with identical output, JSON included. At a terminal it prints a one-line note on stderr naming the new spelling; in `--json` mode or when stdout is not a terminal it prints nothing extra.
|
|
19
|
+
- **`sem mcp` lists the core verbs only:** `sem_find` (with `mode` callers, refs or context, `in`, `text` and `intent`), `sem_grep`, `sem_impact`, `sem_check`, `sem_certify`, `sem_diff`, `sem_graph` and `sem_history`. The earlier tools (`sem_entities`, `sem_context`, `sem_callers`, `sem_log`, `sem_blame`) still answer when called by name. The review-listener tools are listed for `sem mcp --review`, which the review-listener plugin now passes.
|
|
20
|
+
|
|
21
|
+
### Added
|
|
22
|
+
|
|
23
|
+
- **`sem check`, an exact incremental verifier.** Runs the project's TypeScript, lint, test, Go, Cargo and configured checkers and prints one verdict: exit 0 pass, 1 fail, 2 could not decide (nothing to check is 2, never a pass). A checker rechecks only what a change can affect when that is provably the verdict of the full tool, and otherwise runs in full and says why; `--json` adds a verification certificate. `--promises` also proves every promise in `.sem/promises` can fail.
|
|
24
|
+
- **`sem dataflow --witness` and an experimental execution-witness runner (`witness/`).** `--witness` emits one instrumentation task per static source -> sink flow (Python, TS/JS). The runner has a model propose a harness, runs it in a network-less docker sandbox with a runtime that injects a secret canary at the source, and marks a flow CONFIRMED only when the canary reaches the sink along the claimed path in 3 of 3 runs. See `witness/DESIGN.md`.
|
|
25
|
+
- **More Python data-flow sources:** typer/click command parameters, FastAPI route parameters, framework base-class handlers, functions run through `asyncio.to_thread` and executors, typed `*args`/`**kwargs`, and `getattr(self, name)` dispatch.
|
|
26
|
+
- **`sem arch-diff --view`, `--html` and `--from-json`.** A ranked, collapsed view of an arch-diff report (at most 10 items that need a human decision, each with what changed and why it matters), a self-contained HTML page with the module graph around the change, and rendering of a saved `--json` report.
|
|
27
|
+
- **`sem system deps|fetch|build` (experimental).** A layered whole-system graph (repo code, locked dependencies and stdlib, database schema, config and routes, service contracts, a runtime trace) that gives every call and boundary site one outcome per layer and reports how much of the system is knowable. Optional TypeScript front-end and runtime tracers live in `scripts/system/`.
|
|
28
|
+
- **`sem dataflow`, `sem arch-diff <base>..<head>` and `sem certify <base>..<head>`.** `sem dataflow` lowers Python, JS/TS, Go and Rust functions into one IR and reports source -> sink paths with witness paths, matched against declarative models of sources, sinks, sanitizers and callbacks; a call with no known target is an escape, never dropped. `sem arch-diff` reports the architecture delta of a change: new and removed data paths, dependency and cycle changes, signature changes with their unresolved possible callers, broken Python imports and complexity deltas, ranked by severity (text, `--json`, `--md`), in bounded time (`--budget`) and memory (`--max-memory`; above that it analyzes the diff's region and says so). `sem certify` writes a review certificate: touched entities, callers a signature change left unmodified, laws kept or newly broken with their shortest witness paths, affected tests.
|
|
29
|
+
- **An exact-or-unknown call graph for Rust, Go and Python.** Each call site resolves to a target set, external, or unknown; a base or trait method call gets dispatch edges to its overrides. It replaces the scope resolver for those languages (persisted facts are invalidated and rebuild on upgrade).
|
|
30
|
+
- **`sem topology`** for JS/TS workspaces: the module reference graph (value vs type-only imports, workspace imports resolved to source without a build, asset nodes, dynamic-import patterns), graph metrics (cycles, betweenness, propagation cost, blast radius, affected tests) and laws (`forbid`, `only`, `acyclic`, `layers`, `forbidImport`, `forbidPattern` tree-sitter queries in any sem language). `sem promises check|status|verify` evaluates laws that carry a human promise, and `verify` requires each law to break under its declared mutation. `callsFit` reports Python calls that no longer fit their target's signature.
|
|
31
|
+
|
|
32
|
+
### Changed
|
|
33
|
+
|
|
34
|
+
- **Caller answers say whether they are complete.** `sem callers`, `sem impact` and `sem certify` attach a completeness verdict: when an alias, a dispatch registration or an untyped receiver may hide callers, the answer says so and lists the possible callers, nearest first, instead of printing "callers: none" or a green "No tests found". `find`/`callers`/`refs` accept qualified names (`Class.method`, `module.func`, `name@line`) and suggest near matches on a miss; `--file` takes a directory.
|
|
35
|
+
- `sem grep` accepts rg-style trailing paths, `-l` and `-n`; `sem context` labels packed entities with `file:start-end`; members of decorated Python classes are extracted; a cold `sem find` builds through the shared graph cache.
|
|
36
|
+
|
|
37
|
+
- Experimental simple agent sessions now support explicitly acknowledged exact-source reuse and content-only diff reviews since a captured review. Cold file parses run with bounded concurrency, and batch-edit preflight avoids repeated reads of the same file. Validation result reuse is opt-in and requires an operator-owned complete-input fingerprint provider; ordinary checks continue to execute by default.
|
|
38
|
+
- **Signature changes and deletions now report the callers a batch left behind.** Before `weave_transaction` deletes an entity or applies an edit with `allow_signature_change`, it asks sem's graph for that entity's dependents and returns the ones the batch did not edit as `caller_review.unedited_dependents`, so a missed caller shows up before the build. The list comes from the syntax-level graph and is a review aid, not a compile check. Out-of-range integer arguments such as `max_entities: 0` are now clamped instead of failing the call, and the pi package README documents the transaction mode as it runs today.
|
|
39
|
+
|
|
40
|
+
## [0.26.0] - 2026-10-02
|
|
41
|
+
|
|
42
|
+
### Removed
|
|
43
|
+
|
|
44
|
+
- **The strict and adaptive transaction policies are gone from pi, leaving the simple structural session policy as the only one.** `pi/config/transaction.mjs` and `pi/config/adaptive-transaction.mjs` are removed, and `config/simple-transaction.mjs` is now the documented configuration for `PI_SEM_MODE=transaction`.
|
|
45
|
+
|
|
46
|
+
### Added
|
|
47
|
+
|
|
48
|
+
- **Opt-in simple structural session policy for agents.** `pi/config/simple-transaction.mjs` packages batched exact reads, acknowledged context reuse, scoped edit composition and snapshot-checked edits without fixed discovery/transaction call quotas. Includes regression tests and setup instructions. This policy remains experimental, not a guarantee of faster sessions or semantic completeness.
|
|
49
|
+
|
|
50
|
+
### Fixed
|
|
51
|
+
|
|
52
|
+
- **Portable transaction startup and guarded batches.** The simple server starts without JeV credentials; external ranking requires explicit `SEM_JEV_ENABLED=1` plus credentials and task. Same-file exact-edit batches enforce intermediate snapshot guards and roll back earlier writes on a later edit error.
|
|
53
|
+
|
|
54
|
+
- Experimental structural-session adapter snapshot adds explicitly scoped reads and edit programs, batched discovery, bounded validation evidence, and timeout recovery. Retained as a research branch: benchmark timings vary with agent-selected validation workloads and do not establish a universal speedup.
|
|
55
|
+
- Text search includes eligible files without a structural parser (such as Objective-C++ `.mm` and custom build files), including alongside a warm structural index. Search responses preserve coverage limits separately from result pagination; binary content remains excluded.
|
|
56
|
+
|
|
57
|
+
- Simple transaction edits preserve overloaded entity selectors. Timed-out public checks retain partial diagnostics and restart their disposable checker so subsequent validation can proceed.
|
|
58
|
+
|
|
59
|
+
- Simple transaction read batches preserve valid results when another selector is malformed or a requested path is rejected as a symlink. Partial captures report their errors explicitly; strict capture and edit checks remain unchanged.
|
|
60
|
+
|
|
61
|
+
- **Copilot CLI can connect to the MCP server again.** Unsupported discovery probes return `Method not found` without closing the connection, allowing clients to fall back to `initialize` in both standalone and shared modes. Fixes #497.
|
|
62
|
+
- **Indexed name lookup sees renames and added definitions in edited files.** `sem find` checks indexed file freshness and reparses changed files on demand, without requiring a whole dependency-graph refresh. Includes TypeScript, Python and Rust regression coverage.
|
|
63
|
+
- **Dart dependency graphs now resolve ordinary calls, constructor-bound receivers and typed parameters.** Callers and refs no longer select a same-named Dart method for a TypeScript receiver (or vice versa); imported class owners take precedence. Persisted graph/query caches are invalidated so upgrades rebuild the affected edges. Fixes #491.
|
|
64
|
+
- **Shared MCP clients keep independent context history and stay bound to their repository.** Concurrent daemon startup is serialized with an OS lock, stale sockets recover after crashes, and handshakes are bounded. Adds `sem mcp --status` for health checks and reproducible lifecycle coverage on macOS and Linux.
|
|
65
|
+
- **Telemetry uploads are no longer rejected by the server.** 0.21.0 removed the install id from the upload payload while the ingest endpoint still required one, so every batch uploaded since then was refused and active-install counts only ever reflected 0.20.0 and older. Uploads now carry `hash(local seed + day number)`, where the seed is generated once, stays on the machine and is never sent, so a batch groups with the rest of that machine's day and with nothing before or after it. Telemetry is still local-by-default and opt-in, so this only changes the contents of an upload that someone enabled with `sem telemetry on`.
|
|
66
|
+
- **`sem setup` no longer reports success when part of it failed.** It writes the global `diff.external` config first, then installs the Claude Code hook and the pre-commit hook, and those two steps used to treat a permissions or JSON failure as a warning before falling through to a closing message that listed all three features as live and returned success. An unparseable `~/.claude/settings.json` therefore left changed git config, no hook, and a final line reading "sem is wired in". The closing summary is now built from the steps that actually ran, and a failed step reports what did and did not apply before exiting 2. The successful path is unchanged. Thanks to kantorcodes1 on Reddit for the report.
|
|
67
|
+
|
|
7
68
|
## [0.25.0] - 2026-09-13
|
|
8
69
|
|
|
9
70
|
### Added
|
package/README.md
CHANGED
|
@@ -17,6 +17,7 @@
|
|
|
17
17
|
|
|
18
18
|
<p align="center">
|
|
19
19
|
<a href="https://ataraxy-labs.com/blogs/code-is-not-text">Why sem?</a> ·
|
|
20
|
+
<a href="#using-sem-with-agents">Using sem with agents</a> ·
|
|
20
21
|
<a href="#install">Install</a> ·
|
|
21
22
|
<a href="#commands">Commands</a> ·
|
|
22
23
|
<a href="#use-with-ai-agents-mcp">Agents (MCP)</a> ·
|
|
@@ -42,6 +43,83 @@ Cloud-backed queries are opt-in per repo: logging in does not upload a repo or s
|
|
|
42
43
|
<img src="assets/terminal.svg" alt="sem diff" width="800" />
|
|
43
44
|
</p>
|
|
44
45
|
|
|
46
|
+
## Using sem with agents
|
|
47
|
+
|
|
48
|
+
Four questions, four verbs. Every verb takes `--json`.
|
|
49
|
+
|
|
50
|
+
| Question | Verb |
|
|
51
|
+
|---|---|
|
|
52
|
+
| Where is it? | `sem find`, `sem grep` |
|
|
53
|
+
| What does my change touch? | `sem impact` |
|
|
54
|
+
| Is it correct? | `sem check` |
|
|
55
|
+
| What should a human review? | `sem certify` |
|
|
56
|
+
|
|
57
|
+
**Where is it?** Find a function, class or method by name:
|
|
58
|
+
|
|
59
|
+
```bash
|
|
60
|
+
sem find parse_config
|
|
61
|
+
```
|
|
62
|
+
```
|
|
63
|
+
function parse_config src/config.py:1
|
|
64
|
+
```
|
|
65
|
+
|
|
66
|
+
Who calls it (`--callers`), what it uses (`--refs`), or its code with its callers and callees (`--context`):
|
|
67
|
+
|
|
68
|
+
```bash
|
|
69
|
+
sem find parse_config --callers
|
|
70
|
+
```
|
|
71
|
+
```
|
|
72
|
+
function parse_config src/config.py:1
|
|
73
|
+
function load src/config.py:6
|
|
74
|
+
function test_parse_config tests/test_config.py:4
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
Search text, like `rg`:
|
|
78
|
+
|
|
79
|
+
```bash
|
|
80
|
+
sem grep splitlines
|
|
81
|
+
```
|
|
82
|
+
```
|
|
83
|
+
src/config.py:2: pairs = [line.split("=", 1) for line in text.splitlines() if line]
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
**What does my change touch?** Everything that depends on an entity, and the tests to run:
|
|
87
|
+
|
|
88
|
+
```bash
|
|
89
|
+
sem impact parse_config --tests
|
|
90
|
+
```
|
|
91
|
+
```
|
|
92
|
+
⚡ 1 tests affected:
|
|
93
|
+
tests/test_config.py
|
|
94
|
+
function test_parse_config (L4–5)
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
`sem impact --diff HEAD --tests` answers the same for your whole uncommitted change.
|
|
98
|
+
|
|
99
|
+
**Is it correct?** Run the project's own compiler, type checker, linter and tests, rechecking only what the change can affect when that gives the same answer. Exit 0 means pass, 1 fail, 2 could not decide:
|
|
100
|
+
|
|
101
|
+
```bash
|
|
102
|
+
sem check
|
|
103
|
+
```
|
|
104
|
+
```
|
|
105
|
+
cmd FAIL full sh (0 rechecked, 279 ms)
|
|
106
|
+
`python3 -m pytest -q tests` exited 1:
|
|
107
|
+
E AssertionError: assert {'a': '1'} == {'a': '2'}
|
|
108
|
+
```
|
|
109
|
+
|
|
110
|
+
**What should a human review?** A review certificate for a commit range: what changed, which callers the change leaves behind, which tests reach it:
|
|
111
|
+
|
|
112
|
+
```bash
|
|
113
|
+
sem certify main..HEAD
|
|
114
|
+
```
|
|
115
|
+
```
|
|
116
|
+
# sem certificate 94f4bd74e26e..d73d4d092b43
|
|
117
|
+
1 files, 2 entities changed (1 added, 1 modified, 0 deleted, 0 renamed, 0 moved).
|
|
118
|
+
- modified `parse_config` (src/config.py:1): 2 static caller(s) in 2 file(s) that this change does not modify
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
For agents that speak MCP, `sem mcp` serves the same verbs as tools (see [Use with AI agents](#use-with-ai-agents-mcp)).
|
|
122
|
+
|
|
45
123
|
## Install
|
|
46
124
|
|
|
47
125
|
```bash
|
|
@@ -129,8 +207,97 @@ If you installed via npm/bun, the binary lives in `node_modules/.bin/sem` and is
|
|
|
129
207
|
|
|
130
208
|
Works in any Git repo. No setup required. Also works outside Git for arbitrary file comparison.
|
|
131
209
|
|
|
210
|
+
| Verb | The question it answers |
|
|
211
|
+
|---|---|
|
|
212
|
+
| `sem find` | Where is it defined? Who calls it, what does it use, what is around it? |
|
|
213
|
+
| `sem grep` | Where does this text appear? |
|
|
214
|
+
| `sem impact` | What does changing this entity, or this diff, touch? |
|
|
215
|
+
| `sem check` | Is my change correct? |
|
|
216
|
+
| `sem certify` | What should a human review in this commit range? |
|
|
217
|
+
| `sem diff` | Which functions and classes changed? |
|
|
218
|
+
| `sem graph` | How is the code connected? |
|
|
219
|
+
| `sem history` | How did this entity change over time? Who last changed it? |
|
|
220
|
+
| `sem cloud` | Account, cloud reviews, cross-repo queries |
|
|
221
|
+
| `sem config` | git diff integration, telemetry, completions, updates, stats |
|
|
222
|
+
| `sem mcp` | The same verbs as MCP tools for agents |
|
|
223
|
+
|
|
224
|
+
`sem <verb> --help` shows every flag with an example.
|
|
225
|
+
|
|
132
226
|
sem stores its SQLite entity cache outside the repository, under the OS cache directory by default. Set `SEM_CACHE_DIR=/path/to/cache` to override the cache root; repo-local overrides are ignored so cache files do not dirty the working tree.
|
|
133
227
|
|
|
228
|
+
### sem find
|
|
229
|
+
|
|
230
|
+
Find entity definitions by name, then ask about one: who calls it, what it uses, or its code with the code around it. Answers come from an on-disk query index (`index.sem`, next to the entity cache); the first call in a repo builds it, later calls read it directly, with no daemon.
|
|
231
|
+
|
|
232
|
+
```bash
|
|
233
|
+
sem find "function diff_command" # where it is defined ("type name" narrows by kind)
|
|
234
|
+
sem find diff_command load_config # several names in one call
|
|
235
|
+
sem find diff_command --callers # who calls it; says when the caller set may be incomplete
|
|
236
|
+
sem find diff_command --refs # what it calls and references
|
|
237
|
+
sem find diff_command --context # its code plus callers and callees, within a token budget
|
|
238
|
+
sem find diff_command --context --budget 4000 --hops 1
|
|
239
|
+
sem find --in src/auth.ts # every entity in a file or directory
|
|
240
|
+
sem find diff_command --in src/ # only definitions under src/
|
|
241
|
+
sem find --text "retry budget" # entity bodies containing a string, named by entity
|
|
242
|
+
sem find diff_command --json # JSON on any of the above
|
|
243
|
+
```
|
|
244
|
+
|
|
245
|
+
With `--context`, when the target signature itself does not fit, JSON output reports `target_omitted: true`.
|
|
246
|
+
|
|
247
|
+
### sem grep
|
|
248
|
+
|
|
249
|
+
Text search across source files, rg-compatible `file:line:text` output, served from the index's trigram postings when one exists.
|
|
250
|
+
|
|
251
|
+
```bash
|
|
252
|
+
sem grep "TODO"
|
|
253
|
+
sem grep -i todo src/ # case-insensitive, under src/
|
|
254
|
+
sem grep -e foo -e bar --json # several patterns, hits kept apart
|
|
255
|
+
```
|
|
256
|
+
|
|
257
|
+
### sem impact
|
|
258
|
+
|
|
259
|
+
What breaks if an entity changes: its dependencies, its dependents (transitively), and the tests that reach it.
|
|
260
|
+
|
|
261
|
+
```bash
|
|
262
|
+
sem impact authenticateUser # full impact
|
|
263
|
+
sem impact authenticateUser --deps # direct dependencies only
|
|
264
|
+
sem impact authenticateUser --dependents
|
|
265
|
+
sem impact authenticateUser --tests # the tests to run
|
|
266
|
+
sem impact authenticateUser --file src/auth.ts # disambiguate by file
|
|
267
|
+
sem impact --diff HEAD --tests # tests for the uncommitted change
|
|
268
|
+
sem impact --diff main..HEAD --json # one impact report per changed entity
|
|
269
|
+
```
|
|
270
|
+
|
|
271
|
+
With `--diff` and `--tests` in a JS/TS workspace, the answer is the module graph's affected-test selection over the changed files (`sem graph --modules affected-tests`); elsewhere it is the entity graph's. `--no-default-excludes` includes generated, fixture, vendor, benchmark and build trees.
|
|
272
|
+
|
|
273
|
+
### sem check
|
|
274
|
+
|
|
275
|
+
Runs the project's compiler, type checker, linter and tests on the working tree and prints one verdict: exit 0 pass, 1 fail, 2 could not decide. Each checker gives the verdict the real tool would give on the whole project; it rechecks only what the change can affect when that is provably the same answer, and otherwise runs the tool in full and says why. Nothing to check is 2, never a pass.
|
|
276
|
+
|
|
277
|
+
```bash
|
|
278
|
+
sem check # every checker the project has (TypeScript, lint, tests, Go, Cargo)
|
|
279
|
+
sem check --checkers ts,lint,tests # only these
|
|
280
|
+
sem check --base origin/main # against origin/main instead of HEAD
|
|
281
|
+
sem check --promises # also prove every promise in .sem/promises can fail
|
|
282
|
+
sem check --json # one JSON object with a verification certificate
|
|
283
|
+
```
|
|
284
|
+
|
|
285
|
+
Commands of your own go in `.sem/check.json`, e.g. `{"commands": ["python3 -m pytest -q"]}`; they always run in full.
|
|
286
|
+
|
|
287
|
+
### sem certify
|
|
288
|
+
|
|
289
|
+
A review certificate for a commit range: entities touched, signature changes and the callers they leave behind, callee deltas, promises kept or broken (with witnesses), module reachability deltas (JS/TS), affected tests, and the static reference cone. `--arch` gives the architecture view instead: new or removed data paths, side effects, dependencies, cycles, and what did not change, ranked.
|
|
290
|
+
|
|
291
|
+
```bash
|
|
292
|
+
sem certify main..HEAD # markdown certificate
|
|
293
|
+
sem certify main..HEAD --json # the full certificate
|
|
294
|
+
sem certify main..HEAD --arch # architecture delta
|
|
295
|
+
sem certify main..HEAD --arch --view # at most 10 ranked items
|
|
296
|
+
sem certify main..HEAD --html > view.html # one self-contained page
|
|
297
|
+
```
|
|
298
|
+
|
|
299
|
+
`<base>...<head>` uses the merge base; one ref means `<ref>..HEAD`.
|
|
300
|
+
|
|
134
301
|
### sem diff
|
|
135
302
|
|
|
136
303
|
Entity-level diff with rename detection, structural hashing, and word-level inline highlights.
|
|
@@ -171,167 +338,125 @@ echo '[{"filePath":"src/main.rs","status":"modified","beforeContent":"...","afte
|
|
|
171
338
|
sem diff --file-exts .py .rs
|
|
172
339
|
```
|
|
173
340
|
|
|
174
|
-
### sem
|
|
175
|
-
|
|
176
|
-
Cross-file dependency graph shows what breaks if an entity changes.
|
|
177
|
-
|
|
178
|
-
```bash
|
|
179
|
-
# Full impact analysis
|
|
180
|
-
sem impact authenticateUser
|
|
181
|
-
|
|
182
|
-
# Direct dependencies only
|
|
183
|
-
sem impact authenticateUser --deps
|
|
184
|
-
|
|
185
|
-
# Direct dependents only
|
|
186
|
-
sem impact authenticateUser --dependents
|
|
187
|
-
|
|
188
|
-
# Affected tests only
|
|
189
|
-
sem impact authenticateUser --tests
|
|
190
|
-
|
|
191
|
-
# JSON output
|
|
192
|
-
sem impact authenticateUser --json
|
|
193
|
-
|
|
194
|
-
# Disambiguate by file
|
|
195
|
-
sem impact authenticateUser --file src/auth.ts
|
|
196
|
-
|
|
197
|
-
# Include default-excluded paths such as generated, fixture, vendor, benchmark, and build trees
|
|
198
|
-
sem impact authenticateUser --no-default-excludes
|
|
199
|
-
```
|
|
200
|
-
|
|
201
|
-
### sem blame
|
|
202
|
-
|
|
203
|
-
Entity-level blame showing who last modified each function, class, or method.
|
|
204
|
-
|
|
205
|
-
```bash
|
|
206
|
-
sem blame src/auth.ts
|
|
207
|
-
|
|
208
|
-
# JSON output
|
|
209
|
-
sem blame src/auth.ts --json
|
|
210
|
-
```
|
|
211
|
-
|
|
212
|
-
### sem log
|
|
213
|
-
|
|
214
|
-
Track how a single entity evolved through git history.
|
|
215
|
-
|
|
216
|
-
```bash
|
|
217
|
-
sem log authenticateUser
|
|
218
|
-
|
|
219
|
-
# Verbose mode (show content diff between versions)
|
|
220
|
-
sem log authenticateUser -v
|
|
221
|
-
|
|
222
|
-
# Limit commits scanned
|
|
223
|
-
sem log authenticateUser --limit 20
|
|
224
|
-
|
|
225
|
-
# JSON output
|
|
226
|
-
sem log authenticateUser --json
|
|
227
|
-
```
|
|
341
|
+
### sem graph
|
|
228
342
|
|
|
229
|
-
With no
|
|
230
|
-
**hotspots** (most-changed functions/classes, with author counts) and
|
|
231
|
-
**co-change pairs** (entities that repeatedly change in the same commits:
|
|
232
|
-
"if you touch one, don't forget the other"):
|
|
343
|
+
How the code is connected. With no flag, the entity dependency graph (the graph `sem impact` and `sem find --context` are built on); the flags select other layers.
|
|
233
344
|
|
|
234
345
|
```bash
|
|
235
|
-
sem
|
|
236
|
-
sem
|
|
237
|
-
sem
|
|
238
|
-
sem
|
|
346
|
+
sem graph --json # entities and the calls/references between them
|
|
347
|
+
sem graph --modules metrics # JS/TS module graph: density, cycles, depth, centrality
|
|
348
|
+
sem graph --modules blast-radius pkg-a # who breaks if pkg-a changes
|
|
349
|
+
sem graph --modules path src/a.ts src/b.ts
|
|
350
|
+
sem graph --dataflow # reads, writes and source -> sink paths
|
|
351
|
+
sem graph --dataflow --witness # experimental: witness tasks per static flow
|
|
352
|
+
sem graph --system # exact dependency versions from every lockfile
|
|
353
|
+
sem graph --system build --out sys/ # the layered whole-system graph
|
|
239
354
|
```
|
|
240
355
|
|
|
241
|
-
### sem
|
|
356
|
+
### sem history
|
|
242
357
|
|
|
243
|
-
|
|
358
|
+
How an entity changed through git history, logic changes told apart from cosmetic ones.
|
|
244
359
|
|
|
245
360
|
```bash
|
|
246
|
-
sem
|
|
247
|
-
|
|
248
|
-
sem
|
|
249
|
-
|
|
250
|
-
sem
|
|
251
|
-
|
|
252
|
-
# JSON output
|
|
253
|
-
sem entities --json
|
|
254
|
-
sem entities src/auth.ts --json
|
|
255
|
-
|
|
256
|
-
# Include default-excluded paths such as generated, fixture, vendor, benchmark, and build trees
|
|
257
|
-
sem entities --no-default-excludes
|
|
361
|
+
sem history authenticateUser
|
|
362
|
+
sem history authenticateUser -v # with the content diff of each version
|
|
363
|
+
sem history authenticateUser --limit 20
|
|
364
|
+
sem history # repo hotspots and co-change pairs
|
|
365
|
+
sem history --blame src/auth.ts # who last changed each entity in the file
|
|
258
366
|
```
|
|
259
367
|
|
|
260
|
-
|
|
368
|
+
With no entity, `sem history` analyzes recent repo history at the entity level: **hotspots** (most-changed functions/classes, with author counts) and **co-change pairs** (entities that repeatedly change in the same commits: "if you touch one, don't forget the other").
|
|
261
369
|
|
|
262
|
-
|
|
263
|
-
When the target signature itself does not fit, JSON output reports `target_omitted: true`.
|
|
370
|
+
### sem cloud, sem config
|
|
264
371
|
|
|
265
372
|
```bash
|
|
266
|
-
sem
|
|
267
|
-
|
|
268
|
-
#
|
|
269
|
-
sem
|
|
270
|
-
|
|
271
|
-
#
|
|
272
|
-
|
|
373
|
+
sem cloud login # API key, or GitHub when omitted
|
|
374
|
+
sem cloud enable # cloud queries for this public repo (shows what is sent, asks first)
|
|
375
|
+
sem cloud disable # stop cloud for this repo
|
|
376
|
+
sem cloud review listen <diff-id> # attach an agent to a cloud code review
|
|
377
|
+
sem cloud xref --json # cross-repo dependencies
|
|
378
|
+
sem cloud repos # where your code is stored
|
|
379
|
+
|
|
380
|
+
sem config setup # make `git diff` show sem's entity diff
|
|
381
|
+
sem config telemetry off
|
|
382
|
+
sem config completions zsh # shell completions
|
|
383
|
+
sem config update
|
|
384
|
+
sem config stats # local diff counters; nothing leaves your machine
|
|
385
|
+
```
|
|
386
|
+
|
|
387
|
+
### Old command names
|
|
388
|
+
|
|
389
|
+
Every earlier command name and flag still works, with the same output. At a terminal, sem prints a one-line note on stderr with the new spelling; in `--json` mode or when output is piped, it prints nothing extra.
|
|
390
|
+
|
|
391
|
+
| Old | New |
|
|
392
|
+
|---|---|
|
|
393
|
+
| `sem callers X` | `sem find X --callers` |
|
|
394
|
+
| `sem refs X` | `sem find X --refs` |
|
|
395
|
+
| `sem context X` | `sem find X --context` |
|
|
396
|
+
| `sem entities PATH` | `sem find --in PATH` |
|
|
397
|
+
| `sem log X` | `sem history X` |
|
|
398
|
+
| `sem blame FILE` | `sem history --blame FILE` |
|
|
399
|
+
| `sem arch-diff RANGE` | `sem certify RANGE --arch` |
|
|
400
|
+
| `sem topology OP` | `sem graph --modules OP` |
|
|
401
|
+
| `sem dataflow` | `sem graph --dataflow` |
|
|
402
|
+
| `sem system OP` | `sem graph --system OP` |
|
|
403
|
+
| `sem promises verify` | `sem check --promises` |
|
|
404
|
+
| `sem login`, `logout`, `whoami`, `review`, `xref`, `repos` | `sem cloud ...` |
|
|
405
|
+
| `sem cloud never` | `sem cloud disable` |
|
|
406
|
+
| `sem setup`, `unsetup`, `telemetry`, `completions`, `update`, `stats` | `sem config ...` |
|
|
407
|
+
|
|
408
|
+
### Promises
|
|
409
|
+
|
|
410
|
+
A promise is something an agent (or a person) claims about the codebase, written down as a check that either passes or fails. "Done" means the check passes. Promises live in `.sem/promises/*.json`; each law may carry a human `"promise"`. Code-shape laws are tree-sitter queries that work in any language sem parses; every capture whose name does not start with `_` counts as a violation, and nested captures are reported once, at the outermost node. For example, `.sem/promises/jsx-no-logic.json`:
|
|
273
411
|
|
|
274
|
-
|
|
275
|
-
|
|
412
|
+
```json
|
|
413
|
+
{ "laws": [
|
|
414
|
+
{ "id": "jsx-no-logic/conditionals",
|
|
415
|
+
"promise": "JSX contains no conditional rendering",
|
|
416
|
+
"forbidPattern": { "from": ["**/*.tsx", "**/*.jsx"], "within": ["jsx_expression"],
|
|
417
|
+
"query": "(ternary_expression) @conditional" } },
|
|
418
|
+
{ "id": "jsx-no-logic/inline-handler-bodies",
|
|
419
|
+
"promise": "JSX event handlers are references, not block bodies",
|
|
420
|
+
"forbidPattern": { "from": ["**/*.tsx", "**/*.jsx"], "within": ["jsx_attribute"],
|
|
421
|
+
"query": "(arrow_function body: (statement_block (_))) @handler_body" } }
|
|
422
|
+
] }
|
|
276
423
|
```
|
|
277
424
|
|
|
278
|
-
### sem find / callers / refs / grep
|
|
279
|
-
|
|
280
|
-
Cold-start lookups backed by an on-disk, mmap-able query index (`index.sem`, stored next to the SQLite entity cache). The first call in a repo builds the index; every call after that reads it directly, no daemon or background process involved:
|
|
281
|
-
|
|
282
425
|
```bash
|
|
283
|
-
#
|
|
284
|
-
sem
|
|
285
|
-
|
|
286
|
-
#
|
|
287
|
-
sem
|
|
288
|
-
|
|
289
|
-
# What it calls
|
|
290
|
-
sem refs diff_command
|
|
291
|
-
|
|
292
|
-
# Text search across source files (rg-compatible file:line:text output,
|
|
293
|
-
# served from the index's trigram postings when one exists)
|
|
294
|
-
sem grep "TODO"
|
|
295
|
-
|
|
296
|
-
# JSON output on any of the above
|
|
297
|
-
sem find diff_command --json
|
|
426
|
+
sem promises check # KEPT / BROKEN n per promise, then file:line:col capture snippet; exit 1 if any broken
|
|
427
|
+
sem promises check --changed src/App.tsx # only violations in these files (what an edit hook runs)
|
|
428
|
+
sem promises check --since main # only files changed since a ref, uncommitted and untracked included
|
|
429
|
+
sem promises status --json # id, kept, violation count per promise
|
|
430
|
+
sem check --promises # prove every promise can fail (each needs a mutation)
|
|
298
431
|
```
|
|
299
432
|
|
|
300
|
-
|
|
433
|
+
`sem promises` is not listed in `sem --help`; `sem certify` reports promises kept or broken for a range. The same file also works with `sem graph --modules check --laws <file>`, which adds graph and import laws (`forbid`, `only`, `acyclic`, `layers`, `forbidImport`, `allowImports`) for JS/TS workspaces. With `--changed` or `--since`, those laws still run over the whole graph, but only violations that involve a changed file (or the package containing it) are reported.
|
|
301
434
|
|
|
302
|
-
|
|
303
|
-
|
|
304
|
-
Prints the full entity dependency graph for the current repo, or `--json` for the underlying edge list (the same graph `sem impact` and `sem context` are built on top of):
|
|
435
|
+
To have an agent see a broken promise right after it makes an edit, run [`scripts/promises-hook.sh`](scripts/promises-hook.sh) after each edit. The script exits 2 and prints the violations to stderr when a promise breaks, and stays silent otherwise. In Claude Code, add it to your settings as a `PostToolUse` hook:
|
|
305
436
|
|
|
306
|
-
```
|
|
307
|
-
|
|
308
|
-
sem
|
|
437
|
+
```json
|
|
438
|
+
{ "hooks": { "PostToolUse": [ { "matcher": "Edit|Write|MultiEdit",
|
|
439
|
+
"hooks": [ { "type": "command", "command": "sh /path/to/sem/scripts/promises-hook.sh" } ] } ] } }
|
|
309
440
|
```
|
|
310
441
|
|
|
311
|
-
|
|
312
|
-
|
|
313
|
-
Local, cumulative counters: how many diffs `sem` has run in this environment and how much of that was noise filtered out. Nothing here leaves your machine (see [Telemetry](#telemetry)):
|
|
314
|
-
|
|
315
|
-
```bash
|
|
316
|
-
sem stats
|
|
317
|
-
```
|
|
442
|
+
The script reads `tool_input.file_path` from the hook's stdin. Other harnesses can pass the edited paths as arguments instead. In pi, for example, an extension can do this from its `tool_result` handler for the edit and write tools: run `sh promises-hook.sh <path>` and append the stderr to the tool result when the script exits 2.
|
|
318
443
|
|
|
319
444
|
## Use as default Git diff
|
|
320
445
|
|
|
321
446
|
Replace `git diff` output with entity-level diffs. Agents and humans get sem output automatically without changing any commands.
|
|
322
447
|
|
|
323
448
|
```bash
|
|
324
|
-
sem setup
|
|
449
|
+
sem config setup
|
|
325
450
|
```
|
|
326
451
|
|
|
327
452
|
Now `git diff` shows entity-level changes instead of line-level. No prompts, no agent configuration needed. Everything that calls `git diff` gets sem output automatically. Also installs a pre-commit hook that shows entity-level blast radius of staged changes.
|
|
328
453
|
|
|
329
|
-
On macOS and Linux, `sem setup` also registers a Claude Code `UserPromptSubmit` hook (`sem hook prompt-submit`) for prompt-time context injection. It edits `~/.claude/settings.json` idempotently, backs it up first, and leaves any hooks you already have untouched.
|
|
454
|
+
On macOS and Linux, `sem config setup` also registers a Claude Code `UserPromptSubmit` hook (`sem hook prompt-submit`) for prompt-time context injection. It edits `~/.claude/settings.json` idempotently, backs it up first, and leaves any hooks you already have untouched.
|
|
330
455
|
|
|
331
456
|
To disable and go back to normal git diff (also removes the session hooks):
|
|
332
457
|
|
|
333
458
|
```bash
|
|
334
|
-
sem unsetup
|
|
459
|
+
sem config unsetup
|
|
335
460
|
```
|
|
336
461
|
|
|
337
462
|
## Entity-level diffs on every pull request
|
|
@@ -361,10 +486,10 @@ No config, no API keys, never fails your build. See [action/](action/) for detai
|
|
|
361
486
|
|
|
362
487
|
Local is always free and always fast: the on-disk index answers day-to-day queries in single-digit milliseconds even from a cold process, so there's nothing to keep warm and no login required. You do not pay to make your laptop fast.
|
|
363
488
|
|
|
364
|
-
Cloud is for what a laptop can't do. On a very large monorepo the first local graph build can take a few seconds; a shared team graph shouldn't be rebuilt per developer; and CI wants the graph without checking anything out. `sem login` connects those cases to sem cloud, which keeps a warm, pre-built graph for your registered repos and serves the heavy queries from it instead of rebuilding locally.
|
|
489
|
+
Cloud is for what a laptop can't do. On a very large monorepo the first local graph build can take a few seconds; a shared team graph shouldn't be rebuilt per developer; and CI wants the graph without checking anything out. `sem cloud login` connects those cases to sem cloud, which keeps a warm, pre-built graph for your registered repos and serves the heavy queries from it instead of rebuilding locally.
|
|
365
490
|
|
|
366
491
|
```bash
|
|
367
|
-
sem login
|
|
492
|
+
sem cloud login # GitHub device flow, one time
|
|
368
493
|
sem impact myFunc --file src/foo.rs # served from the cloud's warm graph
|
|
369
494
|
```
|
|
370
495
|
|
|
@@ -377,19 +502,20 @@ It is fully optional and transparent:
|
|
|
377
502
|
Related commands, all cloud-account scoped:
|
|
378
503
|
|
|
379
504
|
```bash
|
|
380
|
-
sem logout # log out
|
|
381
|
-
sem whoami # show current cloud identity
|
|
382
|
-
sem cloud
|
|
383
|
-
sem cloud
|
|
384
|
-
sem cloud
|
|
385
|
-
sem cloud
|
|
386
|
-
sem
|
|
387
|
-
sem
|
|
505
|
+
sem cloud logout # log out
|
|
506
|
+
sem cloud whoami # show current cloud identity
|
|
507
|
+
sem cloud enable # turn on cloud queries for a public repo (shows what's sent first)
|
|
508
|
+
sem cloud disable # stop cloud for this repo
|
|
509
|
+
sem cloud xref --json # cross-repo dependencies across your indexed repos
|
|
510
|
+
sem cloud repos # where your code is stored: cloud-indexed repos + local caches
|
|
511
|
+
sem cloud status # cloud + telemetry state for this repo (offline; sends nothing)
|
|
512
|
+
sem cloud share # cloud queries for a private repo, with extra confirmation
|
|
513
|
+
sem cloud forget # delete this repo's cloud index and unregister it
|
|
388
514
|
```
|
|
389
515
|
|
|
390
|
-
`
|
|
516
|
+
`status`, `share`, `forget`, `list`, `preview` and `log` are not shown in `sem cloud --help` but work as listed; each one is read-only or requires explicit confirmation before it sends anything.
|
|
391
517
|
|
|
392
|
-
If your team runs code review through sem cloud, `sem review listen <diff-id-or-url>` execs a coding agent pre-configured to join that review as a live listener that answers reviewer questions anchored to specific lines of the diff.
|
|
518
|
+
If your team runs code review through sem cloud, `sem cloud review listen <diff-id-or-url>` execs a coding agent pre-configured to join that review as a live listener that answers reviewer questions anchored to specific lines of the diff.
|
|
393
519
|
|
|
394
520
|
## What it parses
|
|
395
521
|
|
|
@@ -472,9 +598,26 @@ This means sem detects renames and moves, not just additions and deletions. Stru
|
|
|
472
598
|
|
|
473
599
|
## Use with AI agents (MCP)
|
|
474
600
|
|
|
475
|
-
|
|
601
|
+
On macOS and Linux, clients in the same checkout share a warm repository daemon.
|
|
602
|
+
Run `sem mcp --status` to check it. See the [shared runtime contract and reproducible benchmark](docs/shared-mcp.md)
|
|
603
|
+
for session isolation, fallback behavior, and current platform limits.
|
|
604
|
+
|
|
605
|
+
`sem mcp` starts a [Model Context Protocol](https://modelcontextprotocol.io) server over stdin/stdout. It's not a command you run and read yourself: it's a server your coding agent launches in the background so it can ask sem questions while it works. That's the reason `mcp` lives alongside the normal commands. The agent gets the core verbs as tools:
|
|
606
|
+
|
|
607
|
+
| Tool | Answers |
|
|
608
|
+
|---|---|
|
|
609
|
+
| `sem_find` | Where is it defined? `mode`: `callers`, `refs` or `context`; `in` lists a file or directory |
|
|
610
|
+
| `sem_grep` | Where does this text appear? |
|
|
611
|
+
| `sem_impact` | What does changing this entity touch? Dependents, deps, tests |
|
|
612
|
+
| `sem_check` | Is my change correct? A verdict: pass, fail, could not decide |
|
|
613
|
+
| `sem_certify` | What should a human review in this commit range? |
|
|
614
|
+
| `sem_diff` | Which entities changed between two refs? |
|
|
615
|
+
| `sem_graph` | The entity, module, data-flow or system graph |
|
|
616
|
+
| `sem_history` | How did an entity change? `blame` for a file |
|
|
617
|
+
|
|
618
|
+
Clients that call the earlier tool names (`sem_entities`, `sem_context`, `sem_callers`, `sem_log`, `sem_blame`) still get answers; those names are just not listed. A review-listener session (`sem mcp --review`, which `sem cloud review listen` sets up) also lists `join_review`, `wait_for_branch`, `reply_to_branch` and `list_open_branches`.
|
|
476
619
|
|
|
477
|
-
Why an agent wants these: instead of reading whole files and burning tokens, it can ask "what breaks if I change `submitOrder`" (`sem_impact`) or "give me just the context to refactor this function" (`
|
|
620
|
+
Why an agent wants these: instead of reading whole files and burning tokens, it can ask "what breaks if I change `submitOrder`" (`sem_impact`) or "give me just the context to refactor this function" (`sem_find` with mode `context`, which returns the function's source plus its callers and callees) and get a precise, deterministic answer from the dependency graph instead of a grep result that might miss a caller.
|
|
478
621
|
|
|
479
622
|
Add it once, then talk to your agent normally. It calls the tools on its own.
|
|
480
623
|
|
|
@@ -579,9 +722,9 @@ Used by [weave](https://github.com/Ataraxy-Labs/weave) (semantic merge driver) a
|
|
|
579
722
|
Local by default: sem counts command names (e.g. `diff`, `impact`) on your own machine only, and in that mode nothing is ever uploaded. No code, file paths, repo names, or user identity is recorded, and no network call is made.
|
|
580
723
|
|
|
581
724
|
```bash
|
|
582
|
-
sem telemetry preview # see current mode and exactly what would be sent
|
|
583
|
-
sem telemetry on # opt in: also upload counts to help improve sem
|
|
584
|
-
sem telemetry off # record nothing at all
|
|
725
|
+
sem config telemetry preview # see current mode and exactly what would be sent
|
|
726
|
+
sem config telemetry on # opt in: also upload counts to help improve sem
|
|
727
|
+
sem config telemetry off # record nothing at all
|
|
585
728
|
```
|
|
586
729
|
|
|
587
730
|
`SEM_NO_TELEMETRY=1` or `DO_NOT_TRACK=1` force the record-nothing behavior regardless of mode. Development builds (anything run out of a `cargo build` `target/` directory) never record, so working on sem itself doesn't pollute the numbers.
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@ataraxy-labs/sem",
|
|
3
3
|
"mcpName": "io.github.Ataraxy-Labs/sem",
|
|
4
|
-
"version": "0.
|
|
4
|
+
"version": "0.27.0",
|
|
5
5
|
"description": "npm wrapper for the sem CLI. Downloads the matching release binary and exposes the sem command in node_modules/.bin.",
|
|
6
6
|
"license": "MIT OR Apache-2.0",
|
|
7
7
|
"type": "module",
|