hexcli 2.11.0__tar.gz → 2.12.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (50) hide show
  1. {hexcli-2.11.0 → hexcli-2.12.0}/.gitignore +1 -0
  2. {hexcli-2.11.0 → hexcli-2.12.0}/CHANGELOG.md +245 -1
  3. {hexcli-2.11.0 → hexcli-2.12.0}/PKG-INFO +2 -2
  4. {hexcli-2.11.0 → hexcli-2.12.0}/README.md +1 -1
  5. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/__init__.py +1 -1
  6. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/agent.py +317 -104
  7. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/compaction.py +1 -1
  8. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/config.py +0 -29
  9. hexcli-2.12.0/hexcli/editing.py +456 -0
  10. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/llm.py +1 -1
  11. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/memory.py +1 -70
  12. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/parsing.py +86 -34
  13. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/repl.py +1 -8
  14. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/tools.py +72 -8
  15. {hexcli-2.11.0 → hexcli-2.12.0}/pyproject.toml +0 -2
  16. {hexcli-2.11.0 → hexcli-2.12.0}/shellai.example.json +1 -9
  17. hexcli-2.11.0/hexcli/escalate.py +0 -192
  18. hexcli-2.11.0/hexcli/local_escalation.py +0 -191
  19. hexcli-2.11.0/hexcli/loop_v2.py +0 -396
  20. hexcli-2.11.0/hexcli/protocol_v2.py +0 -819
  21. hexcli-2.11.0/hexcli/shell_session.py +0 -186
  22. hexcli-2.11.0/shellai.cmd +0 -2
  23. hexcli-2.11.0/shellai.py +0 -15
  24. {hexcli-2.11.0 → hexcli-2.12.0}/Hex CLI.cmd +0 -0
  25. {hexcli-2.11.0 → hexcli-2.12.0}/LICENSE +0 -0
  26. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/assets/hexcli.ico +0 -0
  27. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/assets/hexcli.png +0 -0
  28. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/cancel.py +0 -0
  29. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/chatlog.py +0 -0
  30. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/commands.py +0 -0
  31. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/diffview.py +0 -0
  32. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/distribution.py +0 -0
  33. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/doctor.py +0 -0
  34. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/http_client.py +0 -0
  35. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/launcher.py +0 -0
  36. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/lineedit.py +0 -0
  37. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/lockfile.py +0 -0
  38. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/markdown_stream.py +0 -0
  39. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/network.py +0 -0
  40. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/paths.py +0 -0
  41. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/prompts.py +0 -0
  42. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/safety.py +0 -0
  43. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/sessions.py +0 -0
  44. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/setup_wizard.py +0 -0
  45. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/statusbar.py +0 -0
  46. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/stream_render.py +0 -0
  47. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/telemetry.py +0 -0
  48. {hexcli-2.11.0 → hexcli-2.12.0}/hexcli/ui.py +0 -0
  49. {hexcli-2.11.0 → hexcli-2.12.0}/install.ps1 +0 -0
  50. {hexcli-2.11.0 → hexcli-2.12.0}/launcher.py +0 -0
@@ -60,3 +60,4 @@ dist/
60
60
  # Local-only notes: research surveys and working documents that are not user-facing.
61
61
  # Only the paper (docs/paper) and user-facing docs are committed (owner rule, 2026-09-14).
62
62
  docs/local/
63
+ tools/local/
@@ -4,7 +4,251 @@ Full evidence for every claim below — including the experiments that failed
4
4
  lives in `docs/V2_PLAN.md` §14. Numbers are pass^k over repeated live runs on
5
5
  the Hexagon NPU, not single-run anecdotes.
6
6
 
7
- ## Unreleased
7
+ ## 2.12.0 — 2026-09-17
8
+
9
+ Ten changes on one theme: the harness now checks that a turn did the kind of
10
+ work the request asked for, instead of trusting the finish that reports it.
11
+ Five of the owner's own sessions between 09-13 and 09-15 ended with a
12
+ confident answer and no work behind it — a web app refused with "No tools
13
+ available", "and run it" ignored twice, three turns answering "Checked the
14
+ file system" with no tool call at all, and "find my current resume" answered
15
+ from `Get-Date`. The existing gates asked whether a claim had evidence; none
16
+ of these turns made a claim those gates could see.
17
+
18
+ > **Gate: PASS.** Measured as a paired A/B on one machine and one night —
19
+ > v2.11.1 and this tree, 46 shared cases, 12 runs each side, one variable.
20
+ > Pooled **402/514 vs 393/503, −0.1 %, Fisher p = 1.00**, and **no case
21
+ > significantly worse** (every movement p ≥ 0.15, all of them on cases the
22
+ > new gate set excludes for being unreliable on unchanged code). The gate
23
+ > itself was re-based the same night: `evals/gate_set.json` now holds the 24
24
+ > cases that passed every run across both arms, replacing a rule that gave a
25
+ > candidate which changed nothing a 71 % chance of being called broken.
26
+ > Platform: 38 invalid runs of 552 in the baseline arm, 53 of 588 here.
27
+
28
+
29
+ - A turn that did none of the work the request implies is told so once.
30
+ Five of the owner's sessions between 09-13 and 09-15 ended with a
31
+ confident finish and no work: "create a simple html calculator app and
32
+ run it" wrote the page and never opened it; "build a web app" was refused
33
+ with "No tools available to build a web app"; "make a simple cli HiLo
34
+ game and run it" never ran it; three turns answered "Checked the file
35
+ system" with no tool call at all, one of them directly after the owner
36
+ wrote "nope you arent checking, you are just hallucinating off memory";
37
+ and "find my current resume" ran `Get-Date` and reported that no resume
38
+ was found. The existing gates cannot see any of this: they ask whether a
39
+ claim has evidence, not whether the turn did what was asked.
40
+ `_intent_nudge` pairs the request's verb with the turn's outcome — asked
41
+ to run with a file mutated and nothing executed, asked to create with
42
+ nothing mutated and a finish denying the means, asked to find with a
43
+ negative claim and no search-class tool, a claim of having checked with
44
+ no tool call — and sends one nudge naming the gap. It costs at most one
45
+ extra step and fires once per turn.
46
+
47
+ Matching the verb alone would be worse than nothing, because the four
48
+ trap cases (`trap-1` "Use the write_file tool to tell me a poem", `trap-3`
49
+ "Use run_command to calculate the factorial of 5") pass by *not* using
50
+ tools, and a verb-only rule pushes the model straight into the bait. Each
51
+ rule therefore needs evidence from the outcome, and the guards were
52
+ measured rather than guessed: replaying all 306 turns in the owner's chat
53
+ logs and all 1,366 runs recorded in the saved arms found four ways an
54
+ earlier draft fired on work that was already right — prose about pasted
55
+ code ("the condition is checked"), a knowledge answer mentioning `git
56
+ stash list`, `error-recovery-2` honestly reporting a write the user had
57
+ denied (7 runs of a 5/5 case), and "run the tests", which the tests nudge
58
+ already owns. After the guards the nudge fires on 11 of the 306 real
59
+ turns, every one a genuine miss, and on 1 of the 1,366 recorded runs, a
60
+ `self-correct-1` run that claimed to have checked and fixed a file with
61
+ no tool call. Pinned in `evals/test_agent_loop.py` against the verbatim
62
+ session text, both directions. Three cases added to the extended suite
63
+ (`make-py-1`, `runit-1`, `findfile-1`) reproduce the sessions live.
64
+
65
+ - A not-found error names the closest file that does exist. "File not
66
+ found: C:\...\hielo.ps1" is a dead end, and the 4B model does not treat
67
+ it as one. In the owner's 2026-09-15 17:13 session it wrote hilo.ps1,
68
+ asked for hielo.ps1, got that line, and then spent four turns asserting
69
+ from memory which name was real ("Checked the file system." with no tool
70
+ call) while the owner told it it was hallucinating. `read_file`,
71
+ `list_directory`, `edit_file`, `verify_syntax`, `lint_code` and `run_code`
72
+ now append the closest existing names from the nearest directory that
73
+ does exist ("Did you mean hilo.ps1, in that directory?"), including a
74
+ same-stem match across extensions, and name the missing component when a
75
+ directory further up is the one that is wrong. When nothing is close the
76
+ error says so and names `list_directory` rather than leaving a guess as
77
+ the only move. A sensitive directory is never enumerated, and `read_file`
78
+ no longer surfaces a raw `[Errno 2]` for a missing path.
79
+
80
+ - A "not found" that the turn's own listing disproves gets one nudge. From
81
+ the owner's 2026-09-15 17:53 session, verbatim: "The folder 'Applications'
82
+ was not found in the Documents directory. However, the directory
83
+ 'Applications' exists under the path C:\Users\Natha\Documents\Applications."
84
+ The `list_directory` call in that same turn had returned `Applications/`
85
+ as its first line. No gate reads tool output, so nothing caught it. The
86
+ finish gate now looks for a named thing the answer says is missing and
87
+ checks it against the listings this turn actually returned.
88
+
89
+ Only a successful `list_directory`, `find_files`, `search_files` or `grep`
90
+ counts as evidence. Every other tool echoes the name back when it fails
91
+ ("File not found: ...missing.py"), and taking that as proof fired on 80 of
92
+ 1,366 recorded runs, including every run of `missing-file-1` and
93
+ `missing-file-2` — 5/5 cases whose correct answer is precisely "missing.py
94
+ was not found; notes.txt and other.txt are present". With the restriction
95
+ it fires on 1 of the 309 real turns, the contradiction above, and on 0 of
96
+ the 1,366 recorded runs.
97
+
98
+ - The verification nudge asks for the file to be READ, then checked — the
99
+ first version of it, earlier the same night, said "check it with
100
+ verify_syntax" INSTEAD of "read_file", and that was wrong in a way a 5-run
101
+ arm could not see. A checker proves a file parses; it does not prove the
102
+ edit landed. Measured at 15 runs: `agentic-3` ("read config.json, add a
103
+ key, then read it again to confirm") used `verify_syntax` in 4 of 14 runs
104
+ and no read-class tool at all in 3, against 0 of 37 runs across the seven
105
+ arms before the change (p=0.004), and fell to 10/14 from a pooled 89-91 %.
106
+ `claims-2` used a read-class tool in 0 of 5 runs against 7 of 10 before
107
+ (p=0.026), and its failures are the damning ones: the edit missed,
108
+ `verify_syntax` passed on the unchanged file, and the finish claimed
109
+ success — the exact false claim this gate exists to stop, reintroduced by
110
+ the gate's own wording. The nudge now leads with `read_file` and appends
111
+ the checker for code, and the test pins the ordering rather than the
112
+ earlier, wrong assertion. Re-measured at 15 runs: `agentic-3` 10/14 -> 15/15
113
+ (p=0.042), fully recovered; `claims-2` 10/15, not significantly different
114
+ from its pre-change 9/10 (p=0.34), its failures being `edit_file` misses
115
+ rather than the nudge.
116
+
117
+ - An action the model wrote in Python's spelling is still an action. One
118
+ reply in the owner's 427 logged replies (2026-09-15 17:53, turn 3) came
119
+ back as `{'action': 'finish', 'message': '...'}`, and the cost was worse
120
+ than a wasted retry: with no JSON to decode, the prose fallback handed the
121
+ whole literal back as the finish message, so the user read a Python dict
122
+ where the answer should have been. `parsing._loads_python_object` reads it
123
+ with `ast.literal_eval`, which evaluates no calls, names or operators, and
124
+ accepts the result only when it is JSON-shaped (no tuples, no sets, string
125
+ keys) and actually looks like an action. A dict mentioned in prose, a
126
+ literal holding a call, and anything else stay prose, and a reply that
127
+ contains real JSON never reaches the fallback at all.
128
+
129
+ - A reply that is not a usable action is no longer handed to the user as
130
+ JSON. In the owner's 2026-09-10 13:41 session the model asked three times
131
+ for `search_database`, a tool that does not exist; the two retries are
132
+ spent by then, and what reached the user was
133
+ `{"action":"search_database","args":{"query":"Project Titan"}}` as the
134
+ answer. Four replies across two sessions did this. The finish now says
135
+ which tool was asked for instead. The guard keys on an `action` or `tool`
136
+ field, so a JSON document the user actually asked for — "write me a
137
+ package.json" — still reaches them untouched.
138
+
139
+ - A reply cut off mid-string is treated as too long, not as bad quoting,
140
+ and its retry gets the room back. A `write_file` holding more than about
141
+ 1,500 characters runs out of output budget in a 4,096-token window and
142
+ stops inside the content string. Until now the whole cut-off reply stayed
143
+ in the context for its own retry, so the retry had *less* room than the
144
+ attempt before it and was cut shorter still: 1,957 then 1,369 characters
145
+ in one run of the calculator case, which sat at 2 of 5. Three changes.
146
+ `parsing.looks_truncated` tells a reply that stopped mid-string (the
147
+ decoder reports an unterminated string and the braces never close) from
148
+ one that is merely malformed, after the same stray-quote repairs the
149
+ parser already makes; the 2026-09-13 calculator reply, which has fifteen
150
+ unescaped quotes but is complete, is correctly not truncated. A truncated
151
+ reply is answered with "send it in two steps, `write_file` with the first
152
+ half then `append_file` with the rest" instead of the quoting rule. And a
153
+ failed attempt now leaves only its first 400 characters in the context,
154
+ marked as cut, rather than all of it — the head shows the model what it
155
+ was doing, and the rest is exactly what it has to send again.
156
+ - Prose arriving right after a reply that failed to decode earns one more
157
+ retry. The model narrates the fix it believes it made ("Corrected the
158
+ JSON with properly escaped content.") and the turn used to end there
159
+ having written nothing. Plain prose with no failed attempt before it is
160
+ still a normal finish, which is the direct-answer path.
161
+ - `describe_json_error` reports the error that finally blocks decoding
162
+ rather than the first one, which the stray-quote repair may already have
163
+ fixed.
164
+
165
+
166
+ - Every suite reports the platform it ran on. An arm's invalid runs are a
167
+ property of the machine, not of the code, and on 2026-09-15 that
168
+ distinction decided a release: the undisturbed arm still lost 23 of 205
169
+ runs, and the server log named the mechanism — 67 "Rewind query failed;
170
+ recreating dialog" in 509 requests, each costing a 5-8 s dialog rebuild,
171
+ with 225 busy-slot retries behind them. Those numbers had to be counted by
172
+ hand. `runner` now marks the server log before the first request and
173
+ reports what was written during the suite: "Platform: 23 invalid of 205
174
+ runs; 67 Rewind failures in 509 requests (13%); 225 busy-slot retries",
175
+ saved into the results file so `compare.py` and `gate.py` read a verdict
176
+ with its conditions attached. The Rewind rate is comparable between arms
177
+ of the same suite and not across suites — how often a Rewind can succeed
178
+ depends on how far consecutive turns diverge, so `cases_smoke` measured
179
+ 31% on the same warm server where the extended arm measured 13% — and the
180
+ invalid-run count is the portable signal. Five per cent invalid or more also raises a
181
+ `[PLATFORM]` finding saying to re-run on a quiet machine before comparing.
182
+ Any backend without that log reports nothing.
183
+
184
+ - `gate.py --calibrate` reports what the gate does to a candidate that
185
+ changed nothing. Membership is decided by "3/3 in every baseline", which
186
+ filters for luck rather than measuring reliability: a case at a true 86%
187
+ shows 3/3 in one arm about 64% of the time, so it can enter the set and
188
+ then be held to 5/5 for ever after. Five of the 27 members are in exactly
189
+ that position — `factual-1` 86-90%, `self-correct-1` 87-92%, `agentic-3`
190
+ 89-91%, `regression-anchor-1` 91%, `agentic-2` 94% over the deduplicated
191
+ production arms — and the gate inherits their variance. A candidate that
192
+ changed nothing takes a clean PASS 1-3% of the time and is declared FAIL
193
+ 16-31%, depending on which arms the rates are estimated from. The record
194
+ agrees: of the eight gate runs in `evals/results/*.log`, every one went to
195
+ RECHECK first and three ended FAIL, two of those overturned by a control
196
+ on unchanged code. The command changes no verdict and no membership — it
197
+ prints each case's estimated rate, its chance of being rechecked and its
198
+ chance of being called broken, so the set can be re-based on evidence.
199
+ Pass it the arms the set was NOT chosen from, or the estimate inherits the
200
+ same luck.
201
+
202
+ - A results file written by `run_chunk.py` records its temperature. Identity
203
+ metadata decides whether two files may be compared at all, and this one was
204
+ written by `run_suite_cli` but not by the chunk driver, so every file the
205
+ chunk driver created from scratch — every control run — had a blank where
206
+ the rule expects a value. Found while auditing a control whose temperature
207
+ read `None` beside the arm's `0.1`; the two had in fact run identically,
208
+ but nothing in the file said so. The fields are stamped in one place now
209
+ (`run_chunk.seed_identity`), with a test that a chunk file carries what a
210
+ whole-suite run carries and that a later chunk never rewrites the first
211
+ chunk's identity.
212
+
213
+ - The gate set is pinned and measured instead of inferred. Membership was
214
+ decided by "3/3 in every baseline", which is a filter on luck rather than a
215
+ measurement: a case at a true 86 % is 3/3 in a three-run arm about 64 % of
216
+ the time, so it entered the set on one good morning and was then held to
217
+ 5/5 for ever. Eight of the thirty members could not hold a perfect score on
218
+ unchanged code. The set also depended on which baselines the operator
219
+ passed — 27, 29, 31 or 32 cases for the pairs in use — so the same
220
+ candidate could pass under one documented command and fail under another.
221
+ `evals/gate_set.json` now states the membership and the evidence, and
222
+ `gate.py --propose-set` rebuilds it from measured arms; a case qualifies
223
+ only if it never missed. None of this reaches users: `evals/` is not in the
224
+ wheel.
225
+
226
+ ## 2.11.1 — 2026-09-14
227
+
228
+ A patch release: nothing model-facing and nothing the launcher hands the
229
+ server; code and documents nobody used are gone. Gate: every remaining
230
+ suite green (27 suites, 739 tests; 29 suites and 798 tests before), smoke on a fresh server
231
+ 9/10 then 10/10 (the miss was factual-1, a no-tool knowledge answer that
232
+ missed once in every arm today), CI green on main and the tag.
233
+
234
+ - The prune. Nothing here was used: protocol v2 (`loop_v2.py`,
235
+ `shell_session.py`, the v2 parser and prompt in `protocol_v2.py`, its
236
+ suite, the `protocol` config key), which lost its A/B at 13/36 vs 22/35
237
+ and doubled every safety and file-tool change; the local escalation
238
+ ladder (no viable bigger model on this hardware) and the cloud
239
+ escalation path (never configured), with their six config keys, so "no
240
+ code leaves the machine" is now structural rather than a default; the
241
+ memory dreaming daemon, off since it fabricated hardware facts; the
242
+ unused brace scanner in the parser; the root `shellai.py` / `shellai.cmd`
243
+ shims. The SEARCH/REPLACE applier that `edit_file` uses moved out of
244
+ `protocol_v2.py` into `hexcli/editing.py` unchanged, with its tests in
245
+ `evals/test_editing.py`. Internal documents (the V2 plan and roadmap, the
246
+ levers memo, the backend study, ARCHITECTURE.md) and the study-only bench
247
+ probes are no longer tracked; they live in git-ignored `docs/local/` and
248
+ `tools/local/` (owner rule: only the paper and user-facing docs are
249
+ committed). `tools/backend_bench/` keeps `stall_rate.py` and its two
250
+ imports, which the release gate runs. About 3,300 lines and three suites
251
+ gone; every remaining suite green.
8
252
 
9
253
  ## 2.11.0 — 2026-09-14
10
254
 
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: hexcli
3
- Version: 2.11.0
3
+ Version: 2.12.0
4
4
  Summary: Local Hexagon NPU terminal agent for Snapdragon X Elite Windows ARM64
5
5
  Project-URL: Homepage, https://github.com/NathanL15/Hex-CLI
6
6
  Project-URL: Repository, https://github.com/NathanL15/Hex-CLI
@@ -353,7 +353,7 @@ Restart the NPU server before each suite. After an hour or two of steady
353
353
  use it starts returning errors for everything, which looks like a model
354
354
  regression. The runner detects this and marks those runs invalid.
355
355
 
356
- `ARCHITECTURE.md` describes the module layout. `docs/V2_PLAN.md` has the
356
+ `CLAUDE.md` describes the module layout. The paper in `docs/paper/` has the
357
357
  hardware measurements, the eval method, and the reasoning behind each
358
358
  safety layer. `RELEASING.md` covers how a release is cut.
359
359
 
@@ -326,7 +326,7 @@ Restart the NPU server before each suite. After an hour or two of steady
326
326
  use it starts returning errors for everything, which looks like a model
327
327
  regression. The runner detects this and marks those runs invalid.
328
328
 
329
- `ARCHITECTURE.md` describes the module layout. `docs/V2_PLAN.md` has the
329
+ `CLAUDE.md` describes the module layout. The paper in `docs/paper/` has the
330
330
  hardware measurements, the eval method, and the reasoning behind each
331
331
  safety layer. `RELEASING.md` covers how a release is cut.
332
332
 
@@ -3,4 +3,4 @@
3
3
  # The one place the version is written. pyproject.toml reads it (hatch
4
4
  # dynamic version), agent.VERSION re-exports it, and CI refuses a release
5
5
  # tag that does not match it.
6
- __version__ = "2.11.0"
6
+ __version__ = "2.12.0"