@allansantos-dev/smart-tool 0.9.5 → 0.9.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,72 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.9.7 - beta
4
+
5
+ The hook decides code searches with a fixed rule instead of the router model.
6
+
7
+ - Deterministic redirect (`redirect_rule.py`): only the part of a Bash command that searches content is measured, with
8
+ its own paths (`cd` carried, `$VAR` set in the command or the environment, `~` and globs expanded). A content search
9
+ over a project folder (`grep -r`, `rg`, `git grep`, `git -C dir grep`, the Grep tool on a folder) or over more than
10
+ 20 listed files goes to `smart_search`; reading or searching known files, listings, Glob, commands that change files
11
+ and paths it cannot resolve run. Replayed on 2,789 real decisions of the router model: the model redirected 315,
12
+ the rule 463 (134 shared); in a sample of 40 redirected by the rule alone, 40 were content searches over a folder or
13
+ many files; in a sample of 40 redirected by the model alone, ~36 were wrong. A decision takes 0.2 ms (median)
14
+ instead of 1.5 s: the router model cost 119 min of agent waiting in 84 h. The same call always gets the same answer.
15
+ - Block patterns proposed by a model (`pattern_proposals`, on by default): bulk reads the rule let run (globs, xargs,
16
+ -exec, loops, scripts walking folders) and paths it could not resolve are reviewed in the background by the scope
17
+ model (else the router model); the agent never waits. A proposed regex is checked (it compiles, matches the call,
18
+ does not match calls allowed on purpose, blocks at most 25% of the logged calls of that tool) and waits in the setup
19
+ screen with how many logged calls it would have blocked; it redirects only after you accept it. With gpt-4o-mini, 17
20
+ of 20 reviews proposed a pattern and none was usable; with gpt-5-mini, 2 of 20, one a real gap (grep over
21
+ `$(git ls-files)` held in a variable).
22
+ - A new branch or worktree starts from the project's latest view: the scope is inherited while the folder structure
23
+ and root docs still match (no scope or profile model call), and the index starts as a copy that only reprocesses the
24
+ files whose content differs. Measured on a branch at the same commit: the first search took 52 s (scope 11 s,
25
+ profile 20.5 s, rebuilding the index 17 s), seen in real use as the p90 of 64 s over 219 searches.
26
+ - `smart_search_result` waits up to 15 s for the job before answering pending: answering at once had an agent poll the
27
+ same job 7 times in 16 s and give up on Smart Tool.
28
+ - Code redirects no longer need a router model: `smart_search` answers with the lexical index when no embedding model
29
+ is configured.
30
+
31
+ ## 0.9.6 - beta
32
+
33
+ Reported by the claude-code-boss session validating 0.9.5.
34
+
35
+ - `project_manage action=duplicates` reads the files the index does not reflect yet (working tree, files changed on
36
+ disk, commits after the last indexing) and says how many it read: it kept reporting a function removed in a commit
37
+ made after the last indexing, with no sign that the index was behind.
38
+ - In `graph` impact results, a test or caller that is an anonymous callback registered with `test('title', fn)`,
39
+ `it()` or `describe()` shows its title (`test "embedder (Q47): ..."`) instead of `callback`.
40
+ - `graph` with `symbol` accepts a partial path before `::` (`brain-embedder.js::loadConfig`), matching the files whose
41
+ path ends with it; it needed the full path from the project root.
42
+
43
+ Automatic updates.
44
+
45
+ - When the tray starts (at logon, before agent sessions use the daemon) and npm has a newer version, it runs the
46
+ install command in a visible console that closes on success and stays open on failure. Each version is installed
47
+ automatically once: a failed update restores the previous version and restarts the tray, which then only offers it.
48
+ - The tray menu has "Update to X" (or "Check for updates", which says when the installed version is the latest) and an
49
+ "Update automatically" switch (`auto_update`, on by default). Later daily checks only notify.
50
+
51
+ Agents that ignored the hook's redirect: only 44 of 368 blocked calls (12%) were followed by a `smart_search` within
52
+ 30 s; the rest rewrote the search in Bash, node or python.
53
+
54
+ - The block message names the exact tool and arguments (`mcp__smart-tool__smart_search` with the project root filled
55
+ in, `web_fetch` with the URL), says it is a routing rule and not a failure, that redoing the search through Bash,
56
+ python, node or PowerShell bypasses it, and, in Claude Code, how to load the tool with ToolSearch.
57
+ - The MCP server sends instructions on what to do when a call is blocked, and `smart_search`, `web_fetch`,
58
+ `web_search` and their `_result` tools load upfront (`anthropic/alwaysLoad`): Claude Code defers MCP tools, so a
59
+ blocked agent often did not have the replacement in its tool list. In real `claude -p` sessions, an agent asked to
60
+ read a page went straight to `web_fetch`, and one asked to run a broad `grep -rn` switched to `smart_search` after
61
+ the block.
62
+ - `cd dir; grep -n x file.py` was measured as a search over the whole tree (the `;` stuck to the directory name hid
63
+ the file): 45 of the 92 logged Bash redirects with `cd dir;` were reads of one file. A command with a recursive
64
+ search after a file read (`sed -n 1,9p a.py; grep -rn x src`) is still measured by the directory it searches.
65
+ - A subagent whose tool list has no Smart Tool tool (claude-code-guide, statusline-setup, or an agent in
66
+ `.claude/agents` with a `tools` list without it) is not redirected: it had nothing to switch to and gave up.
67
+ - Bash commands that change files, the repository or dependencies (`git checkout`, `rm`, `sed -i`, `writeFileSync`,
68
+ `npm run`, ...) are never redirected and skip the router model: 12 of 283 Bash redirects were such commands.
69
+
3
70
  ## 0.9.5 - beta
4
71
 
5
72
  Reported by the claude-code-boss session on its first day using 0.9.4.
package/README.md CHANGED
@@ -29,7 +29,9 @@ are cheaper, or just tells the agent what Smart Tool would do.
29
29
  npx @allansantos-dev/smart-tool@latest install
30
30
  ```
31
31
 
32
- Run the same command to update; `~/.smart-tool` is kept. It downloads [uv](https://docs.astral.sh/uv/) (checksum
32
+ Updates install themselves: when the tray starts (at logon) and npm has a newer version, it runs this same command in
33
+ a visible console, once per version; the tray menu has "Update to X" (or "Check for updates") on demand and an
34
+ "Update automatically" switch. Running the command by hand also updates; `~/.smart-tool` is kept. It downloads [uv](https://docs.astral.sh/uv/) (checksum
33
35
  verified) and a private Python 3.12 into `~/.smart-tool/runtime`, with nothing added to `PATH` or the registry, then
34
36
  copies Smart Tool to `%LOCALAPPDATA%\Programs\SmartTool` with its own virtual environment,
35
37
  installs the web runtime, registers a per-user Scheduled Task that starts the tray at logon and starts the daemon on
@@ -46,11 +48,16 @@ Open the setup screen from the tray icon ("Open settings") or at `http://127.0.0
46
48
  URL ending where the API starts, like the OpenAI SDK `base_url` (`https://api.openai.com/v1`,
47
49
  `https://openrouter.ai/api/v1`, `http://localhost:11434/v1`), and an optional API key. Saving tests `/models` first.
48
50
  2. **Models**: embedding, rerank (optional; without it results use the hybrid order), the router model (small, no
49
- reasoning: it decides before tool calls), the index classification model and the research model. Any id the gateway
51
+ reasoning: web search tier order and result checks), the index classification model (also reviews block patterns) and the research model. Any id the gateway
50
52
  accepts works; the catalog is shown as suggestions.
51
53
  3. **Agents**: register the MCP server in Claude Code or Codex (runs their official `mcp add` command), install the
52
54
  hook and choose its behavior:
53
- - **Redirect** (default): the native tool is denied with the reason and the Smart Tool tool to use.
55
+ - **Redirect** (default): the native tool is denied with the reason and the Smart Tool tool to use. The decision
56
+ is a fixed rule, under a millisecond and the same every time: a content search over a project folder (`grep -r`,
57
+ `rg`, `git grep`, the Grep tool on a folder) or over more than 20 listed files goes to `smart_search`; reading or
58
+ searching known files, listings, Glob and commands that change files run. With **Block patterns proposed by a
59
+ model** on (default), the scope model reviews, in the background, bulk reads the rule could not measure and
60
+ proposes regex patterns, each checked against the logged calls; a pattern blocks only after you accept it.
54
61
  - **Advise only**: the native tool runs; the agent receives the same advice next to the result and decides.
55
62
  - **Off**: no routing.
56
63
 
@@ -0,0 +1,266 @@
1
+ """Block patterns learned during operation: when the deterministic redirect rule lets a search or read run, the router
2
+ model (review_model) reviews the call in the background (the agent never waits for it) and may propose a regex for calls that should
3
+ go to Smart Tool. A proposal is checked (compiles, matches the call, how many logged calls it would have blocked; one
4
+ that blocks more than BROAD_SHARE of them is discarded) and waits for the user in the setup screen; an accepted pattern
5
+ redirects from then on, deterministically, after the rule's own exceptions (single files, commands changing files).
6
+ """
7
+ import concurrent.futures
8
+ import hashlib
9
+ import json
10
+ import os
11
+ import re
12
+ import threading
13
+ import time
14
+
15
+ import atomic_io
16
+ import config
17
+ import model_client
18
+ import paths
19
+
20
+ STORE_PATH = os.path.join(paths.DATA_DIR, "block-patterns.json")
21
+ REVIEW_TIMEOUT_S = 60
22
+ MAX_REVIEWS_PER_HOUR = 20
23
+ MAX_PENDING = 20
24
+ BROAD_SHARE = 0.25
25
+ PATTERN_MAX_CHARS = 300
26
+ CALL_MAX_CHARS = 1500
27
+ # Only calls whose reach the rule cannot measure are reviewed: bulk reads (globs, xargs, -exec, loops, scripts walking
28
+ # folders) and paths it could not resolve. Reading a few listed files is allowed on purpose and never reviewed.
29
+ REVIEWED_REASONS = ("The searched path uses a variable the hook cannot resolve.", "No content search over a folder.")
30
+ _BULK_READ = re.compile(r"\*|\bxargs\b|-exec\b|\bfor\s+\w+\s+in\b|os\.walk|rglob|glob\(|readdirSync|readdir\(|"
31
+ r"walkSync|-Recurse\b|Get-ChildItem", re.IGNORECASE)
32
+ # Calls the rule lets run on purpose: a pattern matching any of them would block what must run.
33
+ ALLOWED_ON_PURPOSE = ("grep -n handler src/app.py", "grep -n \"def load\" -A20 src/app.py | head -40",
34
+ "sed -n 1,80p src/app.py", "cat README.md", "head -50 src/app.py", "tail -20 logs/app.log",
35
+ "awk '/def /{print NR\": \"$0}' src/app.py", "git status --short", "git log --oneline -5",
36
+ "git diff --stat", "npm test", "pytest -q tests/test_app.py", "ls src", "find src -name '*.py'",
37
+ "Grep pattern=handler path=src/app.py", "cd src && grep -n x app.py",
38
+ "ls -la src docs | head -20", "wc -l src/app.py src/util.py",
39
+ "for f in src/app.py src/util.py; do grep -n handler $f; done")
40
+
41
+ _LOCK = threading.RLock()
42
+ _EXECUTOR = concurrent.futures.ThreadPoolExecutor(max_workers=1, thread_name_prefix="block-patterns")
43
+ _state = {"pending": 0, "recent": [], "seen": set(), "compiled": (None, [])}
44
+
45
+ _SYSTEM_PROMPT = (
46
+ "Smart Tool saves the context of coding agents: a PreToolUse hook sends raw searches of project code to an indexed "
47
+ "semantic search (smart_search), whose short answer costs far fewer tokens than grep output, and the index gets "
48
+ "better the more it is used. A deterministic rule already redirects content searches over a folder (grep -r, rg, "
49
+ "git grep, the Grep tool on a folder) and searches over more than 20 listed files. It lets run, on purpose: reading "
50
+ "or searching inside one or a few known files, listings (ls, find -name, Glob), commands that change files, run "
51
+ "builds, tests or git operations other than grep, and paths it cannot resolve.\n\n"
52
+ "You get one call the rule let run because it could not measure its reach: a bulk read (glob, xargs, -exec, a "
53
+ "loop, a script walking folders) or a path held in a variable. Answer block=true only when calls like it read the "
54
+ "content of many project source files at once, which smart_search answers in far fewer tokens. Answer block=false "
55
+ "for anything else, which is most calls: reading one or a few named files (grep, sed -n, cat, head, tail, awk on a "
56
+ "file), logs, test or build runs, git commands, registry or system commands, scripts that do not read code. When "
57
+ "block=true, write a Python regex matched with re.search against the call text that catches the same kind of "
58
+ "bulk read (not this exact path) and none of the calls allowed on purpose. Patterns already accepted or rejected "
59
+ "are listed; do not propose them again. Call submit_proposal."
60
+ )
61
+ _PROPOSAL_TOOL = {
62
+ "type": "function",
63
+ "function": {
64
+ "name": "submit_proposal",
65
+ "description": "Whether calls like this one should be redirected, and the pattern when they should.",
66
+ "parameters": {
67
+ "type": "object",
68
+ "properties": {
69
+ "block": {"type": "boolean"},
70
+ "pattern": {"type": "string", "description": "Python regex over the call text; empty when block is false"},
71
+ "reason": {"type": "string", "description": "One sentence on what the pattern catches and why"},
72
+ },
73
+ "required": ["block", "pattern", "reason"],
74
+ },
75
+ },
76
+ }
77
+
78
+
79
+ def call_text(tool_name, tool_input):
80
+ """The text a pattern is matched against: the command for Bash, the arguments for Grep and Glob."""
81
+ if tool_name == "Bash":
82
+ return str(tool_input.get("command") or "")
83
+ fields = ("pattern", "path", "glob", "type", "output_mode")
84
+ return f"{tool_name} " + " ".join(f"{f}={tool_input[f]}" for f in fields if tool_input.get(f) not in (None, ""))
85
+
86
+
87
+ def _empty():
88
+ return {"accepted": [], "proposals": [], "rejected": []}
89
+
90
+
91
+ def load():
92
+ """The stored patterns: {"accepted": [...], "proposals": [...], "rejected": [...]}."""
93
+ try:
94
+ with open(STORE_PATH, encoding="utf-8") as stream:
95
+ data = json.load(stream)
96
+ except FileNotFoundError:
97
+ return _empty()
98
+ if not isinstance(data, dict) or any(not isinstance(data.get(k), list) for k in _empty()):
99
+ raise ValueError(f"{STORE_PATH} is not a block patterns file: fix or delete it.")
100
+ return data
101
+
102
+
103
+ def _save(data):
104
+ os.makedirs(os.path.dirname(STORE_PATH), exist_ok=True)
105
+ atomic_io.write_secret_text(STORE_PATH, json.dumps(data, indent=1, ensure_ascii=False))
106
+
107
+
108
+ def matching(tool_name, tool_input):
109
+ """The accepted pattern matching this call, or None. Compiled patterns are cached by the file's mtime."""
110
+ try:
111
+ mtime = os.path.getmtime(STORE_PATH)
112
+ except OSError:
113
+ return None
114
+ with _LOCK:
115
+ cached_mtime, compiled = _state["compiled"]
116
+ if cached_mtime != mtime:
117
+ compiled = [(entry, re.compile(entry["pattern"])) for entry in load()["accepted"]]
118
+ _state["compiled"] = (mtime, compiled)
119
+ text = call_text(tool_name, tool_input)
120
+ return next((entry for entry, regex in compiled if entry["tool"] == tool_name and regex.search(text)), None)
121
+
122
+
123
+ def _history(tool_name, regex, metrics_path):
124
+ """(matches, total) of the logged calls of this tool the pattern would have blocked."""
125
+ matches = total = 0
126
+ for name in (metrics_path + ".1", metrics_path):
127
+ try:
128
+ with open(name, encoding="utf-8") as stream:
129
+ for line in stream:
130
+ try:
131
+ row = json.loads(line)
132
+ except ValueError:
133
+ continue
134
+ if row.get("tool_name") != tool_name or not isinstance(row.get("tool_input"), dict):
135
+ continue
136
+ total += 1
137
+ matches += bool(regex.search(call_text(tool_name, row["tool_input"])))
138
+ except FileNotFoundError:
139
+ continue
140
+ return matches, total
141
+
142
+
143
+ def _ask(tool_name, text, rule_reason, known, model):
144
+ token = model_client.get_token()
145
+ response = model_client.fetch("/v1/chat/completions", token, method="POST", timeout=REVIEW_TIMEOUT_S, body={
146
+ "model": model,
147
+ "messages": [{"role": "system", "content": _SYSTEM_PROMPT},
148
+ {"role": "user", "content": json.dumps({"tool": tool_name, "call": text[:CALL_MAX_CHARS],
149
+ "rule_allowed_because": rule_reason,
150
+ "known_patterns": known}, ensure_ascii=False)}],
151
+ "tools": [_PROPOSAL_TOOL],
152
+ "tool_choice": {"type": "function", "function": {"name": "submit_proposal"}},
153
+ "temperature": 0,
154
+ })
155
+ return json.loads(response["choices"][0]["message"]["tool_calls"][0]["function"]["arguments"])
156
+
157
+
158
+ def propose(tool_name, tool_input, rule_reason, model, metrics_path, redact):
159
+ """Asks the model about one allowed call and stores what it proposes. Returns the stored proposal, or None when the
160
+ model did not propose one or the proposal failed a check (the discarded ones are kept as rejected, with why)."""
161
+ text = call_text(tool_name, tool_input)
162
+ with _LOCK:
163
+ data = load()
164
+ known = [e["pattern"] for group in ("accepted", "proposals", "rejected") for e in data[group]
165
+ if e["tool"] == tool_name][-30:]
166
+ answer = _ask(tool_name, call_text(tool_name, redact(tool_input)), rule_reason, known, model)
167
+ pattern = str(answer.get("pattern") or "").strip()
168
+ if not answer.get("block") or not pattern:
169
+ return None
170
+ entry = {"id": hashlib.sha1(f"{tool_name}\0{pattern}".encode()).hexdigest()[:12], "tool": tool_name,
171
+ "pattern": pattern, "reason": " ".join(str(answer.get("reason") or "").split())[:300],
172
+ "example": call_text(tool_name, redact(tool_input))[:300], "proposed_at": time.time()}
173
+ problem = None
174
+ try:
175
+ regex = re.compile(pattern)
176
+ except re.error as exc:
177
+ problem, regex = f"invalid regex: {exc}", None
178
+ if regex and len(pattern) > PATTERN_MAX_CHARS:
179
+ problem = f"longer than {PATTERN_MAX_CHARS} characters"
180
+ elif regex and not regex.search(text):
181
+ problem = "does not match the call it came from"
182
+ elif regex and any(regex.search(sample) for sample in ALLOWED_ON_PURPOSE):
183
+ problem = "matches calls the rule lets run on purpose: " + next(x for x in ALLOWED_ON_PURPOSE if regex.search(x))
184
+ if regex and not problem:
185
+ matches, total = _history(tool_name, regex, metrics_path)
186
+ entry.update(history_matches=matches, history_total=total)
187
+ if total and matches / total > BROAD_SHARE:
188
+ problem = f"too broad: would block {matches} of {total} logged {tool_name} calls"
189
+ with _LOCK:
190
+ data = load()
191
+ if any(e["id"] == entry["id"] for group in data.values() for e in group):
192
+ return None
193
+ if problem:
194
+ data["rejected"].append({**entry, "rejected_by": "check", "why": problem})
195
+ _save(data)
196
+ return None
197
+ data["proposals"].append(entry)
198
+ _save(data)
199
+ return entry
200
+
201
+
202
+ def review_model(cfg):
203
+ """The scope model (a reasoning model by default), else the router model. Measured on 2026-10-09 over 20 real bulk
204
+ reads the rule let run: gpt-4o-mini proposed a pattern for 17 (14 failed the checks, the other 3 copied one session's
205
+ command); gpt-5-mini proposed 2, one of them a real gap (grep over `$(git ls-files)` held in a variable). The review
206
+ runs in the background, so the slower model costs the agent nothing."""
207
+ return cfg.get("scope_model") or cfg.get("router_model")
208
+
209
+
210
+ def reviewable(tool_name, tool_input, rule_reason):
211
+ """Whether an allowed call is worth the model's review: a path the rule could not resolve, or a bulk read."""
212
+ if rule_reason == REVIEWED_REASONS[0]:
213
+ return True
214
+ return rule_reason == REVIEWED_REASONS[1] and bool(_BULK_READ.search(call_text(tool_name, tool_input)))
215
+
216
+
217
+ def review_later(tool_name, tool_input, rule_reason, metrics_path, redact, log):
218
+ """Queues a background review of an allowed call whose reach the rule could not measure (REVIEWED_REASONS, bulk
219
+ reads only), when pattern_proposals is on and a review model is set. The same kind of call is reviewed once per
220
+ daemon run, at most MAX_REVIEWS_PER_HOUR an hour; failures go to log."""
221
+ if not reviewable(tool_name, tool_input, rule_reason):
222
+ return False
223
+ cfg = config.load_config()
224
+ model = review_model(cfg)
225
+ if not model or config.pattern_proposals(cfg) != "on":
226
+ return False
227
+ signature = (tool_name, re.sub(r"\s+", " ", call_text(tool_name, tool_input))[:120])
228
+ now = time.time()
229
+ with _LOCK:
230
+ _state["recent"] = [t for t in _state["recent"] if now - t < 3600]
231
+ if signature in _state["seen"] or len(_state["recent"]) >= MAX_REVIEWS_PER_HOUR or \
232
+ _state["pending"] >= MAX_PENDING:
233
+ return False
234
+ _state["seen"].add(signature)
235
+ _state["recent"].append(now)
236
+ _state["pending"] += 1
237
+
238
+ def run():
239
+ try:
240
+ propose(tool_name, tool_input, rule_reason, model, metrics_path, redact)
241
+ except Exception as exc:
242
+ log(f"block pattern review failed: {type(exc).__name__}: {exc}")
243
+ finally:
244
+ with _LOCK:
245
+ _state["pending"] -= 1
246
+ _EXECUTOR.submit(run)
247
+ return True
248
+
249
+
250
+ def decide_proposal(entry_id, action):
251
+ """accept moves a proposal to accepted; reject moves it to rejected; remove deletes an accepted pattern."""
252
+ with _LOCK:
253
+ data = load()
254
+ source = "accepted" if action == "remove" else "proposals"
255
+ entry = next((e for e in data[source] if e["id"] == entry_id), None)
256
+ if entry is None:
257
+ raise ValueError(f"No {source[:-1]} with id {entry_id}.")
258
+ data[source] = [e for e in data[source] if e["id"] != entry_id]
259
+ if action == "accept":
260
+ data["accepted"].append({**entry, "accepted_at": time.time()})
261
+ elif action == "reject":
262
+ data["rejected"].append({**entry, "rejected_by": "user"})
263
+ elif action != "remove":
264
+ raise ValueError("Invalid action: use accept, reject or remove.")
265
+ _save(data)
266
+ return data
package/code_impact.py CHANGED
@@ -2,6 +2,7 @@
2
2
  which tests reach it through static calls and which files import its module, each with the first line of its
3
3
  docstring. Answers the question an agent asks before an edit without a chain of searches and file reads."""
4
4
  import os
5
+ import re
5
6
  from collections import defaultdict
6
7
 
7
8
  import code_graph
@@ -17,7 +18,7 @@ NOTE = ("Static analysis of the indexed snapshot plus the files changed in the w
17
18
  def _matches(symbols, name):
18
19
  path, _sep, wanted = name.strip().rpartition("::")
19
20
  path = path.replace("\\", "/").removeprefix("./")
20
- pool = [s for s in symbols if s["path"] == path] if path else symbols
21
+ pool = [s for s in symbols if s["path"] == path or s["path"].endswith("/" + path)] if path else symbols
21
22
  exact = [s for s in pool if s["id"] == wanted or s["name"] == wanted]
22
23
  return exact or [s for s in pool if s["name"].split(".")[-1] == wanted]
23
24
 
@@ -49,6 +50,31 @@ def summaries(root, symbols, wanted_ids):
49
50
  if s["id"] in wanted_ids and (s["path"], s["start_line"]) in docs}
50
51
 
51
52
 
53
+ _TEST_TITLE = re.compile(r"""\b(test|it|describe|suite)(?:\.\w+)?\s*\(\s*(['"`])(.+?)\2""")
54
+
55
+
56
+ def _name_test_callbacks(root, by_id, sites):
57
+ """Anonymous callbacks registered with test('title', fn), it(), describe() are named after their title, so a test
58
+ reads 'test "loads the config"' instead of 'callback'."""
59
+ lines_of = {}
60
+ for site in sites:
61
+ symbol = by_id.get(site.get("id")) or {}
62
+ if symbol.get("name", "").split(".")[-1] not in ("callback", "anonymous"):
63
+ continue
64
+ path = symbol["path"]
65
+ if path not in lines_of:
66
+ try:
67
+ with open(os.path.join(root, path), encoding="utf-8") as stream:
68
+ lines_of[path] = stream.read().splitlines()
69
+ except (OSError, UnicodeDecodeError):
70
+ lines_of[path] = []
71
+ start = symbol["start_line"]
72
+ text = " ".join(lines_of[path][max(start - 2, 0):start])
73
+ match = _TEST_TITLE.search(text)
74
+ if match:
75
+ site["function"] = f'{match.group(1)} "{match.group(3)[:80]}"'
76
+
77
+
52
78
  def impact(root, symbol, view_id=None, depth=3, limit=30):
53
79
  if not isinstance(symbol, str) or not symbol.strip():
54
80
  raise ValueError("Pass symbol: a function, method (Class.method) or class name.")
@@ -92,6 +118,7 @@ def impact(root, symbol, view_id=None, depth=3, limit=30):
92
118
  importers = sorted({d["source"] for d in data["dependencies"] if d.get("target") in files and d["source"] not in files})
93
119
  ordered_tests = sorted(tests.values(), key=lambda t: (t["hops"], t["path"], t["line"]))
94
120
  shown = direct[:limit] + outgoing[:limit] + ordered_tests[:limit]
121
+ _name_test_callbacks(root, by_id, shown)
95
122
  docs = summaries(root, data["symbols"], targets | {site["id"] for site in shown})
96
123
  for site in shown:
97
124
  doc = docs.get(site.pop("id"))
package/config.py CHANGED
@@ -33,6 +33,8 @@ DEFAULT_CONFIG = {
33
33
  "hook_mode": "redirect",
34
34
  "doc_mode": "remind",
35
35
  "duplicate_mode": "warn",
36
+ "auto_update": True,
37
+ "pattern_proposals": "on",
36
38
  "contact": "",
37
39
  "model_adapter": "",
38
40
  "model_adapter_options": {},
@@ -50,6 +52,9 @@ DOC_MODES = ("require", "remind", "off")
50
52
  # Edit hook on functions that copy one already indexed (exact or near-identical body): warn = let the edit run and
51
53
  # point to the existing function; off = say nothing. Never blocks: near matches are leads, not defects.
52
54
  DUPLICATE_MODES = ("warn", "off")
55
+ # Router model reviewing, in the background, searches the deterministic redirect rule let run: on = propose block
56
+ # patterns for the user to accept in the setup screen; off = no review.
57
+ PATTERN_PROPOSAL_MODES = ("on", "off")
53
58
 
54
59
 
55
60
  def doc_mode(cfg=None):
@@ -66,6 +71,23 @@ def duplicate_mode(cfg=None):
66
71
  return value
67
72
 
68
73
 
74
+ def pattern_proposals(cfg=None):
75
+ """on: the router model reviews searches the redirect rule let run and proposes block patterns for the user to
76
+ accept; off: no review."""
77
+ value = (cfg if cfg is not None else load_config()).get("pattern_proposals") or DEFAULT_CONFIG["pattern_proposals"]
78
+ if value not in PATTERN_PROPOSAL_MODES:
79
+ raise ValueError(f"Invalid pattern_proposals ({value!r}) in {CONFIG_PATH}: use on or off.")
80
+ return value
81
+
82
+
83
+ def auto_update(cfg=None):
84
+ """Whether the tray installs a newer release by itself when it starts (once per version)."""
85
+ value = (cfg if cfg is not None else load_config()).get("auto_update", DEFAULT_CONFIG["auto_update"])
86
+ if not isinstance(value, bool):
87
+ raise ValueError(f"Invalid auto_update ({value!r}) in {CONFIG_PATH}: use true or false.")
88
+ return value
89
+
90
+
69
91
  def hook_mode(cfg=None):
70
92
  value = (cfg if cfg is not None else load_config()).get("hook_mode") or DEFAULT_CONFIG["hook_mode"]
71
93
  if value not in HOOK_MODES:
package/duplicates.py CHANGED
@@ -199,13 +199,16 @@ def _entry(path, name, start, end, file_lines, ext):
199
199
  "hash": hashlib.sha1(normalized.encode("utf-8")).hexdigest(), **_features(body, ext, short)}
200
200
 
201
201
 
202
- def _working_tree_functions(root, changes, profile):
203
- """Production functions of the files changed in the working tree, read from disk, so a copy of a function written
204
- earlier in the same session is caught before the index catches up."""
202
+ def _working_tree_functions(root, changes, profile, include_tests=False):
203
+ """Functions of the files the index does not reflect yet, read from disk (production code, plus tests when
204
+ include_tests), with the same identity as indexed ones, so a copy written earlier in the session is caught and
205
+ a function removed since the last indexing is not reported."""
205
206
  found = []
207
+ kinds = ("code", "test") if include_tests else ("code",)
206
208
  for rel, status in changes.items():
207
209
  ext = rel.rsplit(".", 1)[-1].lower() if "." in rel else ""
208
- if status == "deleted" or ext not in _LANG or index_profile.kind(rel, profile) != "code":
210
+ kind = index_profile.kind(rel, profile)
211
+ if status == "deleted" or ext not in _LANG or kind not in kinds:
209
212
  continue
210
213
  try:
211
214
  with open(os.path.join(root, rel), encoding="utf-8") as stream:
@@ -215,12 +218,18 @@ def _working_tree_functions(root, changes, profile):
215
218
  if _GENERATED.search(source[:4000]):
216
219
  continue
217
220
  lines = source.splitlines()
218
- for function in doc_check.functions(rel, source):
219
- if function["end"] - function["start"] + 1 < MIN_LINES:
221
+ functions = doc_check.functions(rel, source)
222
+ occurrences = {}
223
+ for function in functions:
224
+ if function["end"] - function["start"] + 1 < MIN_LINES or any(
225
+ other is not function and other["start"] <= function["start"] and function["end"] <= other["end"]
226
+ and (other["start"], other["end"]) != (function["start"], function["end"]) for other in functions):
220
227
  continue
221
228
  entry = _entry(rel, function["name"], function["start"], function["end"], lines, ext)
222
229
  if entry:
223
- found.append({**entry, "kind": "code"})
230
+ occurrence = occurrences[entry["hash"]] = occurrences.get(entry["hash"], 0) + 1
231
+ found.append({**entry, "kind": kind, "fp": hashlib.sha1(
232
+ f"{rel}\0{entry['hash']}\0{occurrence}".encode("utf-8")).hexdigest()})
224
233
  return found
225
234
 
226
235
 
@@ -359,6 +368,13 @@ def find(root, embed, configured_model, min_similarity=DEFAULT_MIN_SIMILARITY, i
359
368
  started = time.monotonic()
360
369
  dismissed = dismissed or {}
361
370
  functions, notes, view = _functions(root, include_tests)
371
+ if view.get("current"):
372
+ pending = code_graph.pending_changes(root, view["path"])
373
+ if pending:
374
+ profile = index_profile.current((index_scope.load_scope(root) or {}).get("profile"))
375
+ functions = [f for f in functions if f["path"] not in pending] + _working_tree_functions(
376
+ root, pending, profile, include_tests)
377
+ notes.append(f"{len(pending)} file(s) changed since the last indexing were read from disk.")
362
378
  groups = {}
363
379
  for f in functions:
364
380
  groups.setdefault(f["hash"], []).append(f)