@allansantos-dev/smart-tool 0.9.6 → 0.9.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +28 -0
- package/README.md +7 -2
- package/block_patterns.py +266 -0
- package/config.py +13 -0
- package/hook_decision.py +12 -18
- package/index_scope.py +26 -0
- package/indexer.py +17 -1
- package/package.json +1 -1
- package/redirect_rule.py +232 -0
- package/router.py +29 -357
- package/setup_ui.py +84 -3
- package/smart_tool_daemon.py +15 -2
- package/version.py +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,33 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.9.7 - beta
|
|
4
|
+
|
|
5
|
+
The hook decides code searches with a fixed rule instead of the router model.
|
|
6
|
+
|
|
7
|
+
- Deterministic redirect (`redirect_rule.py`): only the part of a Bash command that searches content is measured, with
|
|
8
|
+
its own paths (`cd` carried, `$VAR` set in the command or the environment, `~` and globs expanded). A content search
|
|
9
|
+
over a project folder (`grep -r`, `rg`, `git grep`, `git -C dir grep`, the Grep tool on a folder) or over more than
|
|
10
|
+
20 listed files goes to `smart_search`; reading or searching known files, listings, Glob, commands that change files
|
|
11
|
+
and paths it cannot resolve run. Replayed on 2,789 real decisions of the router model: the model redirected 315,
|
|
12
|
+
the rule 463 (134 shared); in a sample of 40 redirected by the rule alone, 40 were content searches over a folder or
|
|
13
|
+
many files; in a sample of 40 redirected by the model alone, ~36 were wrong. A decision takes 0.2 ms (median)
|
|
14
|
+
instead of 1.5 s: the router model cost 119 min of agent waiting in 84 h. The same call always gets the same answer.
|
|
15
|
+
- Block patterns proposed by a model (`pattern_proposals`, on by default): bulk reads the rule let run (globs, xargs,
|
|
16
|
+
-exec, loops, scripts walking folders) and paths it could not resolve are reviewed in the background by the scope
|
|
17
|
+
model (else the router model); the agent never waits. A proposed regex is checked (it compiles, matches the call,
|
|
18
|
+
does not match calls allowed on purpose, blocks at most 25% of the logged calls of that tool) and waits in the setup
|
|
19
|
+
screen with how many logged calls it would have blocked; it redirects only after you accept it. With gpt-4o-mini, 17
|
|
20
|
+
of 20 reviews proposed a pattern and none was usable; with gpt-5-mini, 2 of 20, one a real gap (grep over
|
|
21
|
+
`$(git ls-files)` held in a variable).
|
|
22
|
+
- A new branch or worktree starts from the project's latest view: the scope is inherited while the folder structure
|
|
23
|
+
and root docs still match (no scope or profile model call), and the index starts as a copy that only reprocesses the
|
|
24
|
+
files whose content differs. Measured on a branch at the same commit: the first search took 52 s (scope 11 s,
|
|
25
|
+
profile 20.5 s, rebuilding the index 17 s), seen in real use as the p90 of 64 s over 219 searches.
|
|
26
|
+
- `smart_search_result` waits up to 15 s for the job before answering pending: answering at once had an agent poll the
|
|
27
|
+
same job 7 times in 16 s and give up on Smart Tool.
|
|
28
|
+
- Code redirects no longer need a router model: `smart_search` answers with the lexical index when no embedding model
|
|
29
|
+
is configured.
|
|
30
|
+
|
|
3
31
|
## 0.9.6 - beta
|
|
4
32
|
|
|
5
33
|
Reported by the claude-code-boss session validating 0.9.5.
|
package/README.md
CHANGED
|
@@ -48,11 +48,16 @@ Open the setup screen from the tray icon ("Open settings") or at `http://127.0.0
|
|
|
48
48
|
URL ending where the API starts, like the OpenAI SDK `base_url` (`https://api.openai.com/v1`,
|
|
49
49
|
`https://openrouter.ai/api/v1`, `http://localhost:11434/v1`), and an optional API key. Saving tests `/models` first.
|
|
50
50
|
2. **Models**: embedding, rerank (optional; without it results use the hybrid order), the router model (small, no
|
|
51
|
-
reasoning:
|
|
51
|
+
reasoning: web search tier order and result checks), the index classification model (also reviews block patterns) and the research model. Any id the gateway
|
|
52
52
|
accepts works; the catalog is shown as suggestions.
|
|
53
53
|
3. **Agents**: register the MCP server in Claude Code or Codex (runs their official `mcp add` command), install the
|
|
54
54
|
hook and choose its behavior:
|
|
55
|
-
- **Redirect** (default): the native tool is denied with the reason and the Smart Tool tool to use.
|
|
55
|
+
- **Redirect** (default): the native tool is denied with the reason and the Smart Tool tool to use. The decision
|
|
56
|
+
is a fixed rule, under a millisecond and the same every time: a content search over a project folder (`grep -r`,
|
|
57
|
+
`rg`, `git grep`, the Grep tool on a folder) or over more than 20 listed files goes to `smart_search`; reading or
|
|
58
|
+
searching known files, listings, Glob and commands that change files run. With **Block patterns proposed by a
|
|
59
|
+
model** on (default), the scope model reviews, in the background, bulk reads the rule could not measure and
|
|
60
|
+
proposes regex patterns, each checked against the logged calls; a pattern blocks only after you accept it.
|
|
56
61
|
- **Advise only**: the native tool runs; the agent receives the same advice next to the result and decides.
|
|
57
62
|
- **Off**: no routing.
|
|
58
63
|
|
|
@@ -0,0 +1,266 @@
|
|
|
1
|
+
"""Block patterns learned during operation: when the deterministic redirect rule lets a search or read run, the router
|
|
2
|
+
model (review_model) reviews the call in the background (the agent never waits for it) and may propose a regex for calls that should
|
|
3
|
+
go to Smart Tool. A proposal is checked (compiles, matches the call, how many logged calls it would have blocked; one
|
|
4
|
+
that blocks more than BROAD_SHARE of them is discarded) and waits for the user in the setup screen; an accepted pattern
|
|
5
|
+
redirects from then on, deterministically, after the rule's own exceptions (single files, commands changing files).
|
|
6
|
+
"""
|
|
7
|
+
import concurrent.futures
|
|
8
|
+
import hashlib
|
|
9
|
+
import json
|
|
10
|
+
import os
|
|
11
|
+
import re
|
|
12
|
+
import threading
|
|
13
|
+
import time
|
|
14
|
+
|
|
15
|
+
import atomic_io
|
|
16
|
+
import config
|
|
17
|
+
import model_client
|
|
18
|
+
import paths
|
|
19
|
+
|
|
20
|
+
STORE_PATH = os.path.join(paths.DATA_DIR, "block-patterns.json")
|
|
21
|
+
REVIEW_TIMEOUT_S = 60
|
|
22
|
+
MAX_REVIEWS_PER_HOUR = 20
|
|
23
|
+
MAX_PENDING = 20
|
|
24
|
+
BROAD_SHARE = 0.25
|
|
25
|
+
PATTERN_MAX_CHARS = 300
|
|
26
|
+
CALL_MAX_CHARS = 1500
|
|
27
|
+
# Only calls whose reach the rule cannot measure are reviewed: bulk reads (globs, xargs, -exec, loops, scripts walking
|
|
28
|
+
# folders) and paths it could not resolve. Reading a few listed files is allowed on purpose and never reviewed.
|
|
29
|
+
REVIEWED_REASONS = ("The searched path uses a variable the hook cannot resolve.", "No content search over a folder.")
|
|
30
|
+
_BULK_READ = re.compile(r"\*|\bxargs\b|-exec\b|\bfor\s+\w+\s+in\b|os\.walk|rglob|glob\(|readdirSync|readdir\(|"
|
|
31
|
+
r"walkSync|-Recurse\b|Get-ChildItem", re.IGNORECASE)
|
|
32
|
+
# Calls the rule lets run on purpose: a pattern matching any of them would block what must run.
|
|
33
|
+
ALLOWED_ON_PURPOSE = ("grep -n handler src/app.py", "grep -n \"def load\" -A20 src/app.py | head -40",
|
|
34
|
+
"sed -n 1,80p src/app.py", "cat README.md", "head -50 src/app.py", "tail -20 logs/app.log",
|
|
35
|
+
"awk '/def /{print NR\": \"$0}' src/app.py", "git status --short", "git log --oneline -5",
|
|
36
|
+
"git diff --stat", "npm test", "pytest -q tests/test_app.py", "ls src", "find src -name '*.py'",
|
|
37
|
+
"Grep pattern=handler path=src/app.py", "cd src && grep -n x app.py",
|
|
38
|
+
"ls -la src docs | head -20", "wc -l src/app.py src/util.py",
|
|
39
|
+
"for f in src/app.py src/util.py; do grep -n handler $f; done")
|
|
40
|
+
|
|
41
|
+
_LOCK = threading.RLock()
|
|
42
|
+
_EXECUTOR = concurrent.futures.ThreadPoolExecutor(max_workers=1, thread_name_prefix="block-patterns")
|
|
43
|
+
_state = {"pending": 0, "recent": [], "seen": set(), "compiled": (None, [])}
|
|
44
|
+
|
|
45
|
+
_SYSTEM_PROMPT = (
|
|
46
|
+
"Smart Tool saves the context of coding agents: a PreToolUse hook sends raw searches of project code to an indexed "
|
|
47
|
+
"semantic search (smart_search), whose short answer costs far fewer tokens than grep output, and the index gets "
|
|
48
|
+
"better the more it is used. A deterministic rule already redirects content searches over a folder (grep -r, rg, "
|
|
49
|
+
"git grep, the Grep tool on a folder) and searches over more than 20 listed files. It lets run, on purpose: reading "
|
|
50
|
+
"or searching inside one or a few known files, listings (ls, find -name, Glob), commands that change files, run "
|
|
51
|
+
"builds, tests or git operations other than grep, and paths it cannot resolve.\n\n"
|
|
52
|
+
"You get one call the rule let run because it could not measure its reach: a bulk read (glob, xargs, -exec, a "
|
|
53
|
+
"loop, a script walking folders) or a path held in a variable. Answer block=true only when calls like it read the "
|
|
54
|
+
"content of many project source files at once, which smart_search answers in far fewer tokens. Answer block=false "
|
|
55
|
+
"for anything else, which is most calls: reading one or a few named files (grep, sed -n, cat, head, tail, awk on a "
|
|
56
|
+
"file), logs, test or build runs, git commands, registry or system commands, scripts that do not read code. When "
|
|
57
|
+
"block=true, write a Python regex matched with re.search against the call text that catches the same kind of "
|
|
58
|
+
"bulk read (not this exact path) and none of the calls allowed on purpose. Patterns already accepted or rejected "
|
|
59
|
+
"are listed; do not propose them again. Call submit_proposal."
|
|
60
|
+
)
|
|
61
|
+
_PROPOSAL_TOOL = {
|
|
62
|
+
"type": "function",
|
|
63
|
+
"function": {
|
|
64
|
+
"name": "submit_proposal",
|
|
65
|
+
"description": "Whether calls like this one should be redirected, and the pattern when they should.",
|
|
66
|
+
"parameters": {
|
|
67
|
+
"type": "object",
|
|
68
|
+
"properties": {
|
|
69
|
+
"block": {"type": "boolean"},
|
|
70
|
+
"pattern": {"type": "string", "description": "Python regex over the call text; empty when block is false"},
|
|
71
|
+
"reason": {"type": "string", "description": "One sentence on what the pattern catches and why"},
|
|
72
|
+
},
|
|
73
|
+
"required": ["block", "pattern", "reason"],
|
|
74
|
+
},
|
|
75
|
+
},
|
|
76
|
+
}
|
|
77
|
+
|
|
78
|
+
|
|
79
|
+
def call_text(tool_name, tool_input):
|
|
80
|
+
"""The text a pattern is matched against: the command for Bash, the arguments for Grep and Glob."""
|
|
81
|
+
if tool_name == "Bash":
|
|
82
|
+
return str(tool_input.get("command") or "")
|
|
83
|
+
fields = ("pattern", "path", "glob", "type", "output_mode")
|
|
84
|
+
return f"{tool_name} " + " ".join(f"{f}={tool_input[f]}" for f in fields if tool_input.get(f) not in (None, ""))
|
|
85
|
+
|
|
86
|
+
|
|
87
|
+
def _empty():
|
|
88
|
+
return {"accepted": [], "proposals": [], "rejected": []}
|
|
89
|
+
|
|
90
|
+
|
|
91
|
+
def load():
|
|
92
|
+
"""The stored patterns: {"accepted": [...], "proposals": [...], "rejected": [...]}."""
|
|
93
|
+
try:
|
|
94
|
+
with open(STORE_PATH, encoding="utf-8") as stream:
|
|
95
|
+
data = json.load(stream)
|
|
96
|
+
except FileNotFoundError:
|
|
97
|
+
return _empty()
|
|
98
|
+
if not isinstance(data, dict) or any(not isinstance(data.get(k), list) for k in _empty()):
|
|
99
|
+
raise ValueError(f"{STORE_PATH} is not a block patterns file: fix or delete it.")
|
|
100
|
+
return data
|
|
101
|
+
|
|
102
|
+
|
|
103
|
+
def _save(data):
|
|
104
|
+
os.makedirs(os.path.dirname(STORE_PATH), exist_ok=True)
|
|
105
|
+
atomic_io.write_secret_text(STORE_PATH, json.dumps(data, indent=1, ensure_ascii=False))
|
|
106
|
+
|
|
107
|
+
|
|
108
|
+
def matching(tool_name, tool_input):
|
|
109
|
+
"""The accepted pattern matching this call, or None. Compiled patterns are cached by the file's mtime."""
|
|
110
|
+
try:
|
|
111
|
+
mtime = os.path.getmtime(STORE_PATH)
|
|
112
|
+
except OSError:
|
|
113
|
+
return None
|
|
114
|
+
with _LOCK:
|
|
115
|
+
cached_mtime, compiled = _state["compiled"]
|
|
116
|
+
if cached_mtime != mtime:
|
|
117
|
+
compiled = [(entry, re.compile(entry["pattern"])) for entry in load()["accepted"]]
|
|
118
|
+
_state["compiled"] = (mtime, compiled)
|
|
119
|
+
text = call_text(tool_name, tool_input)
|
|
120
|
+
return next((entry for entry, regex in compiled if entry["tool"] == tool_name and regex.search(text)), None)
|
|
121
|
+
|
|
122
|
+
|
|
123
|
+
def _history(tool_name, regex, metrics_path):
|
|
124
|
+
"""(matches, total) of the logged calls of this tool the pattern would have blocked."""
|
|
125
|
+
matches = total = 0
|
|
126
|
+
for name in (metrics_path + ".1", metrics_path):
|
|
127
|
+
try:
|
|
128
|
+
with open(name, encoding="utf-8") as stream:
|
|
129
|
+
for line in stream:
|
|
130
|
+
try:
|
|
131
|
+
row = json.loads(line)
|
|
132
|
+
except ValueError:
|
|
133
|
+
continue
|
|
134
|
+
if row.get("tool_name") != tool_name or not isinstance(row.get("tool_input"), dict):
|
|
135
|
+
continue
|
|
136
|
+
total += 1
|
|
137
|
+
matches += bool(regex.search(call_text(tool_name, row["tool_input"])))
|
|
138
|
+
except FileNotFoundError:
|
|
139
|
+
continue
|
|
140
|
+
return matches, total
|
|
141
|
+
|
|
142
|
+
|
|
143
|
+
def _ask(tool_name, text, rule_reason, known, model):
|
|
144
|
+
token = model_client.get_token()
|
|
145
|
+
response = model_client.fetch("/v1/chat/completions", token, method="POST", timeout=REVIEW_TIMEOUT_S, body={
|
|
146
|
+
"model": model,
|
|
147
|
+
"messages": [{"role": "system", "content": _SYSTEM_PROMPT},
|
|
148
|
+
{"role": "user", "content": json.dumps({"tool": tool_name, "call": text[:CALL_MAX_CHARS],
|
|
149
|
+
"rule_allowed_because": rule_reason,
|
|
150
|
+
"known_patterns": known}, ensure_ascii=False)}],
|
|
151
|
+
"tools": [_PROPOSAL_TOOL],
|
|
152
|
+
"tool_choice": {"type": "function", "function": {"name": "submit_proposal"}},
|
|
153
|
+
"temperature": 0,
|
|
154
|
+
})
|
|
155
|
+
return json.loads(response["choices"][0]["message"]["tool_calls"][0]["function"]["arguments"])
|
|
156
|
+
|
|
157
|
+
|
|
158
|
+
def propose(tool_name, tool_input, rule_reason, model, metrics_path, redact):
|
|
159
|
+
"""Asks the model about one allowed call and stores what it proposes. Returns the stored proposal, or None when the
|
|
160
|
+
model did not propose one or the proposal failed a check (the discarded ones are kept as rejected, with why)."""
|
|
161
|
+
text = call_text(tool_name, tool_input)
|
|
162
|
+
with _LOCK:
|
|
163
|
+
data = load()
|
|
164
|
+
known = [e["pattern"] for group in ("accepted", "proposals", "rejected") for e in data[group]
|
|
165
|
+
if e["tool"] == tool_name][-30:]
|
|
166
|
+
answer = _ask(tool_name, call_text(tool_name, redact(tool_input)), rule_reason, known, model)
|
|
167
|
+
pattern = str(answer.get("pattern") or "").strip()
|
|
168
|
+
if not answer.get("block") or not pattern:
|
|
169
|
+
return None
|
|
170
|
+
entry = {"id": hashlib.sha1(f"{tool_name}\0{pattern}".encode()).hexdigest()[:12], "tool": tool_name,
|
|
171
|
+
"pattern": pattern, "reason": " ".join(str(answer.get("reason") or "").split())[:300],
|
|
172
|
+
"example": call_text(tool_name, redact(tool_input))[:300], "proposed_at": time.time()}
|
|
173
|
+
problem = None
|
|
174
|
+
try:
|
|
175
|
+
regex = re.compile(pattern)
|
|
176
|
+
except re.error as exc:
|
|
177
|
+
problem, regex = f"invalid regex: {exc}", None
|
|
178
|
+
if regex and len(pattern) > PATTERN_MAX_CHARS:
|
|
179
|
+
problem = f"longer than {PATTERN_MAX_CHARS} characters"
|
|
180
|
+
elif regex and not regex.search(text):
|
|
181
|
+
problem = "does not match the call it came from"
|
|
182
|
+
elif regex and any(regex.search(sample) for sample in ALLOWED_ON_PURPOSE):
|
|
183
|
+
problem = "matches calls the rule lets run on purpose: " + next(x for x in ALLOWED_ON_PURPOSE if regex.search(x))
|
|
184
|
+
if regex and not problem:
|
|
185
|
+
matches, total = _history(tool_name, regex, metrics_path)
|
|
186
|
+
entry.update(history_matches=matches, history_total=total)
|
|
187
|
+
if total and matches / total > BROAD_SHARE:
|
|
188
|
+
problem = f"too broad: would block {matches} of {total} logged {tool_name} calls"
|
|
189
|
+
with _LOCK:
|
|
190
|
+
data = load()
|
|
191
|
+
if any(e["id"] == entry["id"] for group in data.values() for e in group):
|
|
192
|
+
return None
|
|
193
|
+
if problem:
|
|
194
|
+
data["rejected"].append({**entry, "rejected_by": "check", "why": problem})
|
|
195
|
+
_save(data)
|
|
196
|
+
return None
|
|
197
|
+
data["proposals"].append(entry)
|
|
198
|
+
_save(data)
|
|
199
|
+
return entry
|
|
200
|
+
|
|
201
|
+
|
|
202
|
+
def review_model(cfg):
|
|
203
|
+
"""The scope model (a reasoning model by default), else the router model. Measured on 2026-10-09 over 20 real bulk
|
|
204
|
+
reads the rule let run: gpt-4o-mini proposed a pattern for 17 (14 failed the checks, the other 3 copied one session's
|
|
205
|
+
command); gpt-5-mini proposed 2, one of them a real gap (grep over `$(git ls-files)` held in a variable). The review
|
|
206
|
+
runs in the background, so the slower model costs the agent nothing."""
|
|
207
|
+
return cfg.get("scope_model") or cfg.get("router_model")
|
|
208
|
+
|
|
209
|
+
|
|
210
|
+
def reviewable(tool_name, tool_input, rule_reason):
|
|
211
|
+
"""Whether an allowed call is worth the model's review: a path the rule could not resolve, or a bulk read."""
|
|
212
|
+
if rule_reason == REVIEWED_REASONS[0]:
|
|
213
|
+
return True
|
|
214
|
+
return rule_reason == REVIEWED_REASONS[1] and bool(_BULK_READ.search(call_text(tool_name, tool_input)))
|
|
215
|
+
|
|
216
|
+
|
|
217
|
+
def review_later(tool_name, tool_input, rule_reason, metrics_path, redact, log):
|
|
218
|
+
"""Queues a background review of an allowed call whose reach the rule could not measure (REVIEWED_REASONS, bulk
|
|
219
|
+
reads only), when pattern_proposals is on and a review model is set. The same kind of call is reviewed once per
|
|
220
|
+
daemon run, at most MAX_REVIEWS_PER_HOUR an hour; failures go to log."""
|
|
221
|
+
if not reviewable(tool_name, tool_input, rule_reason):
|
|
222
|
+
return False
|
|
223
|
+
cfg = config.load_config()
|
|
224
|
+
model = review_model(cfg)
|
|
225
|
+
if not model or config.pattern_proposals(cfg) != "on":
|
|
226
|
+
return False
|
|
227
|
+
signature = (tool_name, re.sub(r"\s+", " ", call_text(tool_name, tool_input))[:120])
|
|
228
|
+
now = time.time()
|
|
229
|
+
with _LOCK:
|
|
230
|
+
_state["recent"] = [t for t in _state["recent"] if now - t < 3600]
|
|
231
|
+
if signature in _state["seen"] or len(_state["recent"]) >= MAX_REVIEWS_PER_HOUR or \
|
|
232
|
+
_state["pending"] >= MAX_PENDING:
|
|
233
|
+
return False
|
|
234
|
+
_state["seen"].add(signature)
|
|
235
|
+
_state["recent"].append(now)
|
|
236
|
+
_state["pending"] += 1
|
|
237
|
+
|
|
238
|
+
def run():
|
|
239
|
+
try:
|
|
240
|
+
propose(tool_name, tool_input, rule_reason, model, metrics_path, redact)
|
|
241
|
+
except Exception as exc:
|
|
242
|
+
log(f"block pattern review failed: {type(exc).__name__}: {exc}")
|
|
243
|
+
finally:
|
|
244
|
+
with _LOCK:
|
|
245
|
+
_state["pending"] -= 1
|
|
246
|
+
_EXECUTOR.submit(run)
|
|
247
|
+
return True
|
|
248
|
+
|
|
249
|
+
|
|
250
|
+
def decide_proposal(entry_id, action):
|
|
251
|
+
"""accept moves a proposal to accepted; reject moves it to rejected; remove deletes an accepted pattern."""
|
|
252
|
+
with _LOCK:
|
|
253
|
+
data = load()
|
|
254
|
+
source = "accepted" if action == "remove" else "proposals"
|
|
255
|
+
entry = next((e for e in data[source] if e["id"] == entry_id), None)
|
|
256
|
+
if entry is None:
|
|
257
|
+
raise ValueError(f"No {source[:-1]} with id {entry_id}.")
|
|
258
|
+
data[source] = [e for e in data[source] if e["id"] != entry_id]
|
|
259
|
+
if action == "accept":
|
|
260
|
+
data["accepted"].append({**entry, "accepted_at": time.time()})
|
|
261
|
+
elif action == "reject":
|
|
262
|
+
data["rejected"].append({**entry, "rejected_by": "user"})
|
|
263
|
+
elif action != "remove":
|
|
264
|
+
raise ValueError("Invalid action: use accept, reject or remove.")
|
|
265
|
+
_save(data)
|
|
266
|
+
return data
|
package/config.py
CHANGED
|
@@ -34,6 +34,7 @@ DEFAULT_CONFIG = {
|
|
|
34
34
|
"doc_mode": "remind",
|
|
35
35
|
"duplicate_mode": "warn",
|
|
36
36
|
"auto_update": True,
|
|
37
|
+
"pattern_proposals": "on",
|
|
37
38
|
"contact": "",
|
|
38
39
|
"model_adapter": "",
|
|
39
40
|
"model_adapter_options": {},
|
|
@@ -51,6 +52,9 @@ DOC_MODES = ("require", "remind", "off")
|
|
|
51
52
|
# Edit hook on functions that copy one already indexed (exact or near-identical body): warn = let the edit run and
|
|
52
53
|
# point to the existing function; off = say nothing. Never blocks: near matches are leads, not defects.
|
|
53
54
|
DUPLICATE_MODES = ("warn", "off")
|
|
55
|
+
# Router model reviewing, in the background, searches the deterministic redirect rule let run: on = propose block
|
|
56
|
+
# patterns for the user to accept in the setup screen; off = no review.
|
|
57
|
+
PATTERN_PROPOSAL_MODES = ("on", "off")
|
|
54
58
|
|
|
55
59
|
|
|
56
60
|
def doc_mode(cfg=None):
|
|
@@ -67,6 +71,15 @@ def duplicate_mode(cfg=None):
|
|
|
67
71
|
return value
|
|
68
72
|
|
|
69
73
|
|
|
74
|
+
def pattern_proposals(cfg=None):
|
|
75
|
+
"""on: the router model reviews searches the redirect rule let run and proposes block patterns for the user to
|
|
76
|
+
accept; off: no review."""
|
|
77
|
+
value = (cfg if cfg is not None else load_config()).get("pattern_proposals") or DEFAULT_CONFIG["pattern_proposals"]
|
|
78
|
+
if value not in PATTERN_PROPOSAL_MODES:
|
|
79
|
+
raise ValueError(f"Invalid pattern_proposals ({value!r}) in {CONFIG_PATH}: use on or off.")
|
|
80
|
+
return value
|
|
81
|
+
|
|
82
|
+
|
|
70
83
|
def auto_update(cfg=None):
|
|
71
84
|
"""Whether the tray installs a newer release by itself when it starts (once per version)."""
|
|
72
85
|
value = (cfg if cfg is not None else load_config()).get("auto_update", DEFAULT_CONFIG["auto_update"])
|
package/hook_decision.py
CHANGED
|
@@ -100,15 +100,10 @@ def agent_has_smart_tool(agent_type, cwd):
|
|
|
100
100
|
return True
|
|
101
101
|
|
|
102
102
|
|
|
103
|
-
def _search_root(
|
|
104
|
-
"""project_root for smart_search: the registered project holding
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
raw = router._target_path(tool_input, cwd)
|
|
108
|
-
paths = [router._resolve_path(raw, cwd)] if raw else router._command_paths(tool_input.get("command"), cwd)
|
|
109
|
-
target = os.path.abspath(next((p for p in paths if p), None) or cwd or ".")
|
|
110
|
-
if os.path.isfile(target):
|
|
111
|
-
target = os.path.dirname(target)
|
|
103
|
+
def _search_root(target, cwd):
|
|
104
|
+
"""project_root for smart_search: the registered project holding the folder the redirected search reads, else that
|
|
105
|
+
folder; the session folder when the rule names none."""
|
|
106
|
+
target = os.path.abspath(target or cwd or ".")
|
|
112
107
|
key = os.path.normcase(target) + os.sep
|
|
113
108
|
roots = [p["root"] for p in project_store.all_projects()
|
|
114
109
|
if key.startswith(os.path.normcase(os.path.abspath(p["root"])).rstrip(os.sep) + os.sep)]
|
|
@@ -126,10 +121,10 @@ def _load_hint(client, tool):
|
|
|
126
121
|
return f' If it is not in your tool list yet, load it first with ToolSearch, query "select:mcp__{SMART_TOOL_SERVER}__{tool}".'
|
|
127
122
|
|
|
128
123
|
|
|
129
|
-
def redirect_message(client, tool_name,
|
|
130
|
-
"""Deny text for a broad raw search: the exact tool and arguments to call
|
|
131
|
-
Bash, python, node or PowerShell is the bypass the rule
|
|
132
|
-
root = _search_root(
|
|
124
|
+
def redirect_message(client, tool_name, target, reason, cwd):
|
|
125
|
+
"""Deny text for a broad raw search: the exact tool and arguments to call (project_root from the folder the search
|
|
126
|
+
reads), and that rewriting the same search in Bash, python, node or PowerShell is the bypass the rule stops."""
|
|
127
|
+
root = _search_root(target, cwd).replace("\\", "/")
|
|
133
128
|
return (f"Blocked by Smart Tool, a routing rule set by the user, not a failure ({tool_name}: {clean_reason(reason)}). "
|
|
134
129
|
f'Run this search with {_tool_name(client, "smart_search")}: project_root="{root}", query_identifiers = what '
|
|
135
130
|
f"you are looking for in identifier terms (plus query_comments when the project has two languages)."
|
|
@@ -280,10 +275,9 @@ def _route(payload, client, web_route, cfg):
|
|
|
280
275
|
tool_name = payload.get("tool_name", "")
|
|
281
276
|
tool_input = payload.get("tool_input") if isinstance(payload.get("tool_input"), dict) else {}
|
|
282
277
|
cwd = payload.get("cwd")
|
|
283
|
-
router_model = cfg.get("router_model")
|
|
284
|
-
if not router_model:
|
|
285
|
-
return {}
|
|
286
278
|
if tool_name in WEB_TOOLS:
|
|
279
|
+
if not cfg.get("router_model"):
|
|
280
|
+
return {}
|
|
287
281
|
try:
|
|
288
282
|
present = smart_tool_in_session(cwd)
|
|
289
283
|
except (OSError, ValueError):
|
|
@@ -295,7 +289,7 @@ def _route(payload, client, web_route, cfg):
|
|
|
295
289
|
return deny(web_message(client, tool_name, url, clean_reason(decision.get("reason")))) if decision.get("redirect") else {}
|
|
296
290
|
router.CLIENT.set(client)
|
|
297
291
|
try:
|
|
298
|
-
decision, reason = router.decide(tool_name, tool_input,
|
|
292
|
+
decision, reason, target = router.decide(tool_name, tool_input, cwd=cwd)
|
|
299
293
|
except Exception as exc:
|
|
300
294
|
try:
|
|
301
295
|
router.log_unavailable(tool_name, tool_input, exc)
|
|
@@ -303,5 +297,5 @@ def _route(payload, client, web_route, cfg):
|
|
|
303
297
|
pass
|
|
304
298
|
return {"systemMessage": f"Smart Tool unavailable, routing skipped: {type(exc).__name__}"}
|
|
305
299
|
if decision == "redirect":
|
|
306
|
-
return deny(redirect_message(client, tool_name,
|
|
300
|
+
return deny(redirect_message(client, tool_name, target, reason, cwd))
|
|
307
301
|
return {}
|
package/index_scope.py
CHANGED
|
@@ -138,6 +138,32 @@ def load_scope(root):
|
|
|
138
138
|
return None
|
|
139
139
|
|
|
140
140
|
|
|
141
|
+
def sibling_scope(root):
|
|
142
|
+
"""Scope of the most recently saved other Git view of the same project, or None. A new branch or a worktree starts
|
|
143
|
+
without a scope; asking the scope and profile models again cost ~28 s of a 52 s first search (measured 2026-10-09
|
|
144
|
+
on a branch at the same commit). The caller still checks it with needs_rescan before using it."""
|
|
145
|
+
if not index_views.describe(root).get("git"):
|
|
146
|
+
return None
|
|
147
|
+
own = scope_path(root)
|
|
148
|
+
prefix = project_identity.project_id(root) + ".v-"
|
|
149
|
+
try:
|
|
150
|
+
names = [n for n in os.listdir(SCOPE_DIR) if n.startswith(prefix) and n.endswith(".scope.json")]
|
|
151
|
+
except OSError:
|
|
152
|
+
return None
|
|
153
|
+
paths = sorted((os.path.join(SCOPE_DIR, n) for n in names), key=os.path.getmtime, reverse=True)
|
|
154
|
+
for path in paths:
|
|
155
|
+
if os.path.normcase(path) == os.path.normcase(own):
|
|
156
|
+
continue
|
|
157
|
+
try:
|
|
158
|
+
with open(path, "r", encoding="utf-8") as f:
|
|
159
|
+
data = json.load(f)
|
|
160
|
+
except (OSError, ValueError, UnicodeError):
|
|
161
|
+
continue
|
|
162
|
+
if valid_scope(data):
|
|
163
|
+
return data
|
|
164
|
+
return None
|
|
165
|
+
|
|
166
|
+
|
|
141
167
|
def save_scope(root, scope):
|
|
142
168
|
if not valid_scope(scope):
|
|
143
169
|
raise ValueError("Scope must contain lists of paths in include/exclude.")
|
package/indexer.py
CHANGED
|
@@ -230,12 +230,28 @@ def migrate_legacy(root):
|
|
|
230
230
|
if not old and index_views.describe(root).get('git'):
|
|
231
231
|
old = next((os.path.join(INDEX_DIR, key + '.sqlite3') for key in
|
|
232
232
|
[project_identity.project_id(root), *project_identity.legacy_ids(root)]
|
|
233
|
-
if os.path.isfile(os.path.join(INDEX_DIR, key + '.sqlite3'))), None)
|
|
233
|
+
if os.path.isfile(os.path.join(INDEX_DIR, key + '.sqlite3'))), None) or \
|
|
234
|
+
_sibling_view_db(root, target)
|
|
234
235
|
if old:
|
|
235
236
|
_backup(old, target)
|
|
236
237
|
return target
|
|
237
238
|
|
|
238
239
|
|
|
240
|
+
def _sibling_view_db(root, target):
|
|
241
|
+
"""Most recently written index of another Git view of the same project, or None. A new branch or worktree starts
|
|
242
|
+
from its copy: the full content check a view change forces then reprocesses only the files that differ, instead
|
|
243
|
+
of chunking every file again (17 s of a 52 s first search on a branch at the same commit, measured 2026-10-09)."""
|
|
244
|
+
prefix = project_identity.project_id(root) + ".v-"
|
|
245
|
+
try:
|
|
246
|
+
names = os.listdir(INDEX_DIR)
|
|
247
|
+
except OSError:
|
|
248
|
+
return None
|
|
249
|
+
candidates = [os.path.join(INDEX_DIR, n) for n in names
|
|
250
|
+
if n.startswith(prefix) and n.endswith(".sqlite3") and ".building-" not in n]
|
|
251
|
+
candidates = [p for p in candidates if os.path.normcase(p) != os.path.normcase(target)]
|
|
252
|
+
return max(candidates, key=os.path.getmtime, default=None)
|
|
253
|
+
|
|
254
|
+
|
|
239
255
|
def _open_path(path):
|
|
240
256
|
conn = sqlite3.connect(path, timeout=10)
|
|
241
257
|
# `chunks.text` é conteúdo de arquivo em texto puro: sem isto, um DELETE só desliga a
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@allansantos-dev/smart-tool",
|
|
3
|
-
"version": "0.9.
|
|
3
|
+
"version": "0.9.7",
|
|
4
4
|
"description": "Local MCP server that gives coding agents (Claude Code, Codex) cheaper, sharper tools than their built-in search. Windows.",
|
|
5
5
|
"license": "Apache-2.0",
|
|
6
6
|
"author": "Allan Santos",
|