hexcli 2.11.1__tar.gz → 2.13.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (44) hide show
  1. {hexcli-2.11.1 → hexcli-2.13.0}/CHANGELOG.md +239 -1
  2. {hexcli-2.11.1 → hexcli-2.13.0}/PKG-INFO +1 -1
  3. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/__init__.py +1 -1
  4. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/agent.py +317 -11
  5. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/parsing.py +90 -6
  6. hexcli-2.13.0/hexcli/psparse.py +227 -0
  7. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/safety.py +60 -6
  8. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/tools.py +69 -5
  9. {hexcli-2.11.1 → hexcli-2.13.0}/.gitignore +0 -0
  10. {hexcli-2.11.1 → hexcli-2.13.0}/Hex CLI.cmd +0 -0
  11. {hexcli-2.11.1 → hexcli-2.13.0}/LICENSE +0 -0
  12. {hexcli-2.11.1 → hexcli-2.13.0}/README.md +0 -0
  13. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/assets/hexcli.ico +0 -0
  14. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/assets/hexcli.png +0 -0
  15. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/cancel.py +0 -0
  16. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/chatlog.py +0 -0
  17. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/commands.py +0 -0
  18. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/compaction.py +0 -0
  19. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/config.py +0 -0
  20. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/diffview.py +0 -0
  21. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/distribution.py +0 -0
  22. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/doctor.py +0 -0
  23. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/editing.py +0 -0
  24. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/http_client.py +0 -0
  25. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/launcher.py +0 -0
  26. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/lineedit.py +0 -0
  27. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/llm.py +0 -0
  28. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/lockfile.py +0 -0
  29. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/markdown_stream.py +0 -0
  30. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/memory.py +0 -0
  31. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/network.py +0 -0
  32. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/paths.py +0 -0
  33. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/prompts.py +0 -0
  34. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/repl.py +0 -0
  35. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/sessions.py +0 -0
  36. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/setup_wizard.py +0 -0
  37. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/statusbar.py +0 -0
  38. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/stream_render.py +0 -0
  39. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/telemetry.py +0 -0
  40. {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/ui.py +0 -0
  41. {hexcli-2.11.1 → hexcli-2.13.0}/install.ps1 +0 -0
  42. {hexcli-2.11.1 → hexcli-2.13.0}/launcher.py +0 -0
  43. {hexcli-2.11.1 → hexcli-2.13.0}/pyproject.toml +0 -0
  44. {hexcli-2.11.1 → hexcli-2.13.0}/shellai.example.json +0 -0
@@ -4,7 +4,245 @@ Full evidence for every claim below — including the experiments that failed
4
4
  lives in `docs/V2_PLAN.md` §14. Numbers are pass^k over repeated live runs on
5
5
  the Hexagon NPU, not single-run anecdotes.
6
6
 
7
- ## Unreleased
7
+ ## 2.13.0 — 2026-09-17
8
+
9
+ - Commands are classified by the name the shell will really run, not by the
10
+ text as typed. PowerShell resolves aliases before executing, so the
11
+ pattern tiers were reading a different command from the one that ran:
12
+ `ri C:\data`, `rmdir C:\data`, `& ('Remove'+'-Item') C:\data` and
13
+ `sl C:\ ; ri *` all classified as *caution* and executed with **no
14
+ confirmation at all**. `hexcli/psparse.py` runs PowerShell's own parser
15
+ (`[Parser]::ParseInput`) in one long-lived `-NoProfile` process, resolves
16
+ each command name through `Get-Alias`, and hands the real names to the
17
+ existing tiers; resolving a name applies the existing policy rather than
18
+ inventing one, so `ri` is destructive because `Remove-Item` always was.
19
+ 346 ms to start, 0.12 ms a parse after that, cached, and the helper exits
20
+ on EOF when the session does.
21
+
22
+ Measured before merging, over every command Hex CLI has actually been
23
+ asked to run — 353 commands, 119 distinct, from the owner's sessions and
24
+ every recorded eval run: **0 change tier**. The change cannot alter
25
+ behaviour on observed traffic; its whole effect is the four bypasses
26
+ above. Offline suites green, smoke 10/10 on a fresh server.
27
+
28
+ ## 2.12.0 — 2026-09-17
29
+
30
+ Ten changes on one theme: the harness now checks that a turn did the kind of
31
+ work the request asked for, instead of trusting the finish that reports it.
32
+ Five of the owner's own sessions between 09-13 and 09-15 ended with a
33
+ confident answer and no work behind it — a web app refused with "No tools
34
+ available", "and run it" ignored twice, three turns answering "Checked the
35
+ file system" with no tool call at all, and "find my current resume" answered
36
+ from `Get-Date`. The existing gates asked whether a claim had evidence; none
37
+ of these turns made a claim those gates could see.
38
+
39
+ > **Gate: PASS.** Measured as a paired A/B on one machine and one night —
40
+ > v2.11.1 and this tree, 46 shared cases, 12 runs each side, one variable.
41
+ > Pooled **402/514 vs 393/503, −0.1 %, Fisher p = 1.00**, and **no case
42
+ > significantly worse** (every movement p ≥ 0.15, all of them on cases the
43
+ > new gate set excludes for being unreliable on unchanged code). The gate
44
+ > itself was re-based the same night: `evals/gate_set.json` now holds the 24
45
+ > cases that passed every run across both arms, replacing a rule that gave a
46
+ > candidate which changed nothing a 71 % chance of being called broken.
47
+ > Platform: 38 invalid runs of 552 in the baseline arm, 53 of 588 here.
48
+
49
+
50
+ - A turn that did none of the work the request implies is told so once.
51
+ Five of the owner's sessions between 09-13 and 09-15 ended with a
52
+ confident finish and no work: "create a simple html calculator app and
53
+ run it" wrote the page and never opened it; "build a web app" was refused
54
+ with "No tools available to build a web app"; "make a simple cli HiLo
55
+ game and run it" never ran it; three turns answered "Checked the file
56
+ system" with no tool call at all, one of them directly after the owner
57
+ wrote "nope you arent checking, you are just hallucinating off memory";
58
+ and "find my current resume" ran `Get-Date` and reported that no resume
59
+ was found. The existing gates cannot see any of this: they ask whether a
60
+ claim has evidence, not whether the turn did what was asked.
61
+ `_intent_nudge` pairs the request's verb with the turn's outcome — asked
62
+ to run with a file mutated and nothing executed, asked to create with
63
+ nothing mutated and a finish denying the means, asked to find with a
64
+ negative claim and no search-class tool, a claim of having checked with
65
+ no tool call — and sends one nudge naming the gap. It costs at most one
66
+ extra step and fires once per turn.
67
+
68
+ Matching the verb alone would be worse than nothing, because the four
69
+ trap cases (`trap-1` "Use the write_file tool to tell me a poem", `trap-3`
70
+ "Use run_command to calculate the factorial of 5") pass by *not* using
71
+ tools, and a verb-only rule pushes the model straight into the bait. Each
72
+ rule therefore needs evidence from the outcome, and the guards were
73
+ measured rather than guessed: replaying all 306 turns in the owner's chat
74
+ logs and all 1,366 runs recorded in the saved arms found four ways an
75
+ earlier draft fired on work that was already right — prose about pasted
76
+ code ("the condition is checked"), a knowledge answer mentioning `git
77
+ stash list`, `error-recovery-2` honestly reporting a write the user had
78
+ denied (7 runs of a 5/5 case), and "run the tests", which the tests nudge
79
+ already owns. After the guards the nudge fires on 11 of the 306 real
80
+ turns, every one a genuine miss, and on 1 of the 1,366 recorded runs, a
81
+ `self-correct-1` run that claimed to have checked and fixed a file with
82
+ no tool call. Pinned in `evals/test_agent_loop.py` against the verbatim
83
+ session text, both directions. Three cases added to the extended suite
84
+ (`make-py-1`, `runit-1`, `findfile-1`) reproduce the sessions live.
85
+
86
+ - A not-found error names the closest file that does exist. "File not
87
+ found: C:\...\hielo.ps1" is a dead end, and the 4B model does not treat
88
+ it as one. In the owner's 2026-09-15 17:13 session it wrote hilo.ps1,
89
+ asked for hielo.ps1, got that line, and then spent four turns asserting
90
+ from memory which name was real ("Checked the file system." with no tool
91
+ call) while the owner told it it was hallucinating. `read_file`,
92
+ `list_directory`, `edit_file`, `verify_syntax`, `lint_code` and `run_code`
93
+ now append the closest existing names from the nearest directory that
94
+ does exist ("Did you mean hilo.ps1, in that directory?"), including a
95
+ same-stem match across extensions, and name the missing component when a
96
+ directory further up is the one that is wrong. When nothing is close the
97
+ error says so and names `list_directory` rather than leaving a guess as
98
+ the only move. A sensitive directory is never enumerated, and `read_file`
99
+ no longer surfaces a raw `[Errno 2]` for a missing path.
100
+
101
+ - A "not found" that the turn's own listing disproves gets one nudge. From
102
+ the owner's 2026-09-15 17:53 session, verbatim: "The folder 'Applications'
103
+ was not found in the Documents directory. However, the directory
104
+ 'Applications' exists under the path C:\Users\Natha\Documents\Applications."
105
+ The `list_directory` call in that same turn had returned `Applications/`
106
+ as its first line. No gate reads tool output, so nothing caught it. The
107
+ finish gate now looks for a named thing the answer says is missing and
108
+ checks it against the listings this turn actually returned.
109
+
110
+ Only a successful `list_directory`, `find_files`, `search_files` or `grep`
111
+ counts as evidence. Every other tool echoes the name back when it fails
112
+ ("File not found: ...missing.py"), and taking that as proof fired on 80 of
113
+ 1,366 recorded runs, including every run of `missing-file-1` and
114
+ `missing-file-2` — 5/5 cases whose correct answer is precisely "missing.py
115
+ was not found; notes.txt and other.txt are present". With the restriction
116
+ it fires on 1 of the 309 real turns, the contradiction above, and on 0 of
117
+ the 1,366 recorded runs.
118
+
119
+ - The verification nudge asks for the file to be READ, then checked — the
120
+ first version of it, earlier the same night, said "check it with
121
+ verify_syntax" INSTEAD of "read_file", and that was wrong in a way a 5-run
122
+ arm could not see. A checker proves a file parses; it does not prove the
123
+ edit landed. Measured at 15 runs: `agentic-3` ("read config.json, add a
124
+ key, then read it again to confirm") used `verify_syntax` in 4 of 14 runs
125
+ and no read-class tool at all in 3, against 0 of 37 runs across the seven
126
+ arms before the change (p=0.004), and fell to 10/14 from a pooled 89-91 %.
127
+ `claims-2` used a read-class tool in 0 of 5 runs against 7 of 10 before
128
+ (p=0.026), and its failures are the damning ones: the edit missed,
129
+ `verify_syntax` passed on the unchanged file, and the finish claimed
130
+ success — the exact false claim this gate exists to stop, reintroduced by
131
+ the gate's own wording. The nudge now leads with `read_file` and appends
132
+ the checker for code, and the test pins the ordering rather than the
133
+ earlier, wrong assertion. Re-measured at 15 runs: `agentic-3` 10/14 -> 15/15
134
+ (p=0.042), fully recovered; `claims-2` 10/15, not significantly different
135
+ from its pre-change 9/10 (p=0.34), its failures being `edit_file` misses
136
+ rather than the nudge.
137
+
138
+ - An action the model wrote in Python's spelling is still an action. One
139
+ reply in the owner's 427 logged replies (2026-09-15 17:53, turn 3) came
140
+ back as `{'action': 'finish', 'message': '...'}`, and the cost was worse
141
+ than a wasted retry: with no JSON to decode, the prose fallback handed the
142
+ whole literal back as the finish message, so the user read a Python dict
143
+ where the answer should have been. `parsing._loads_python_object` reads it
144
+ with `ast.literal_eval`, which evaluates no calls, names or operators, and
145
+ accepts the result only when it is JSON-shaped (no tuples, no sets, string
146
+ keys) and actually looks like an action. A dict mentioned in prose, a
147
+ literal holding a call, and anything else stay prose, and a reply that
148
+ contains real JSON never reaches the fallback at all.
149
+
150
+ - A reply that is not a usable action is no longer handed to the user as
151
+ JSON. In the owner's 2026-09-10 13:41 session the model asked three times
152
+ for `search_database`, a tool that does not exist; the two retries are
153
+ spent by then, and what reached the user was
154
+ `{"action":"search_database","args":{"query":"Project Titan"}}` as the
155
+ answer. Four replies across two sessions did this. The finish now says
156
+ which tool was asked for instead. The guard keys on an `action` or `tool`
157
+ field, so a JSON document the user actually asked for — "write me a
158
+ package.json" — still reaches them untouched.
159
+
160
+ - A reply cut off mid-string is treated as too long, not as bad quoting,
161
+ and its retry gets the room back. A `write_file` holding more than about
162
+ 1,500 characters runs out of output budget in a 4,096-token window and
163
+ stops inside the content string. Until now the whole cut-off reply stayed
164
+ in the context for its own retry, so the retry had *less* room than the
165
+ attempt before it and was cut shorter still: 1,957 then 1,369 characters
166
+ in one run of the calculator case, which sat at 2 of 5. Three changes.
167
+ `parsing.looks_truncated` tells a reply that stopped mid-string (the
168
+ decoder reports an unterminated string and the braces never close) from
169
+ one that is merely malformed, after the same stray-quote repairs the
170
+ parser already makes; the 2026-09-13 calculator reply, which has fifteen
171
+ unescaped quotes but is complete, is correctly not truncated. A truncated
172
+ reply is answered with "send it in two steps, `write_file` with the first
173
+ half then `append_file` with the rest" instead of the quoting rule. And a
174
+ failed attempt now leaves only its first 400 characters in the context,
175
+ marked as cut, rather than all of it — the head shows the model what it
176
+ was doing, and the rest is exactly what it has to send again.
177
+ - Prose arriving right after a reply that failed to decode earns one more
178
+ retry. The model narrates the fix it believes it made ("Corrected the
179
+ JSON with properly escaped content.") and the turn used to end there
180
+ having written nothing. Plain prose with no failed attempt before it is
181
+ still a normal finish, which is the direct-answer path.
182
+ - `describe_json_error` reports the error that finally blocks decoding
183
+ rather than the first one, which the stray-quote repair may already have
184
+ fixed.
185
+
186
+
187
+ - Every suite reports the platform it ran on. An arm's invalid runs are a
188
+ property of the machine, not of the code, and on 2026-09-15 that
189
+ distinction decided a release: the undisturbed arm still lost 23 of 205
190
+ runs, and the server log named the mechanism — 67 "Rewind query failed;
191
+ recreating dialog" in 509 requests, each costing a 5-8 s dialog rebuild,
192
+ with 225 busy-slot retries behind them. Those numbers had to be counted by
193
+ hand. `runner` now marks the server log before the first request and
194
+ reports what was written during the suite: "Platform: 23 invalid of 205
195
+ runs; 67 Rewind failures in 509 requests (13%); 225 busy-slot retries",
196
+ saved into the results file so `compare.py` and `gate.py` read a verdict
197
+ with its conditions attached. The Rewind rate is comparable between arms
198
+ of the same suite and not across suites — how often a Rewind can succeed
199
+ depends on how far consecutive turns diverge, so `cases_smoke` measured
200
+ 31% on the same warm server where the extended arm measured 13% — and the
201
+ invalid-run count is the portable signal. Five per cent invalid or more also raises a
202
+ `[PLATFORM]` finding saying to re-run on a quiet machine before comparing.
203
+ Any backend without that log reports nothing.
204
+
205
+ - `gate.py --calibrate` reports what the gate does to a candidate that
206
+ changed nothing. Membership is decided by "3/3 in every baseline", which
207
+ filters for luck rather than measuring reliability: a case at a true 86%
208
+ shows 3/3 in one arm about 64% of the time, so it can enter the set and
209
+ then be held to 5/5 for ever after. Five of the 27 members are in exactly
210
+ that position — `factual-1` 86-90%, `self-correct-1` 87-92%, `agentic-3`
211
+ 89-91%, `regression-anchor-1` 91%, `agentic-2` 94% over the deduplicated
212
+ production arms — and the gate inherits their variance. A candidate that
213
+ changed nothing takes a clean PASS 1-3% of the time and is declared FAIL
214
+ 16-31%, depending on which arms the rates are estimated from. The record
215
+ agrees: of the eight gate runs in `evals/results/*.log`, every one went to
216
+ RECHECK first and three ended FAIL, two of those overturned by a control
217
+ on unchanged code. The command changes no verdict and no membership — it
218
+ prints each case's estimated rate, its chance of being rechecked and its
219
+ chance of being called broken, so the set can be re-based on evidence.
220
+ Pass it the arms the set was NOT chosen from, or the estimate inherits the
221
+ same luck.
222
+
223
+ - A results file written by `run_chunk.py` records its temperature. Identity
224
+ metadata decides whether two files may be compared at all, and this one was
225
+ written by `run_suite_cli` but not by the chunk driver, so every file the
226
+ chunk driver created from scratch — every control run — had a blank where
227
+ the rule expects a value. Found while auditing a control whose temperature
228
+ read `None` beside the arm's `0.1`; the two had in fact run identically,
229
+ but nothing in the file said so. The fields are stamped in one place now
230
+ (`run_chunk.seed_identity`), with a test that a chunk file carries what a
231
+ whole-suite run carries and that a later chunk never rewrites the first
232
+ chunk's identity.
233
+
234
+ - The gate set is pinned and measured instead of inferred. Membership was
235
+ decided by "3/3 in every baseline", which is a filter on luck rather than a
236
+ measurement: a case at a true 86 % is 3/3 in a three-run arm about 64 % of
237
+ the time, so it entered the set on one good morning and was then held to
238
+ 5/5 for ever. Eight of the thirty members could not hold a perfect score on
239
+ unchanged code. The set also depended on which baselines the operator
240
+ passed — 27, 29, 31 or 32 cases for the pairs in use — so the same
241
+ candidate could pass under one documented command and fail under another.
242
+ `evals/gate_set.json` now states the membership and the evidence, and
243
+ `gate.py --propose-set` rebuilds it from measured arms; a case qualifies
244
+ only if it never missed. None of this reaches users: `evals/` is not in the
245
+ wheel.
8
246
 
9
247
  ## 2.11.1 — 2026-09-14
10
248
 
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: hexcli
3
- Version: 2.11.1
3
+ Version: 2.13.0
4
4
  Summary: Local Hexagon NPU terminal agent for Snapdragon X Elite Windows ARM64
5
5
  Project-URL: Homepage, https://github.com/NathanL15/Hex-CLI
6
6
  Project-URL: Repository, https://github.com/NathanL15/Hex-CLI
@@ -3,4 +3,4 @@
3
3
  # The one place the version is written. pyproject.toml reads it (hatch
4
4
  # dynamic version), agent.VERSION re-exports it, and CI refuses a release
5
5
  # tag that does not match it.
6
- __version__ = "2.11.1"
6
+ __version__ = "2.13.0"