hexcli 2.11.1__tar.gz → 2.13.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {hexcli-2.11.1 → hexcli-2.13.0}/CHANGELOG.md +239 -1
- {hexcli-2.11.1 → hexcli-2.13.0}/PKG-INFO +1 -1
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/__init__.py +1 -1
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/agent.py +317 -11
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/parsing.py +90 -6
- hexcli-2.13.0/hexcli/psparse.py +227 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/safety.py +60 -6
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/tools.py +69 -5
- {hexcli-2.11.1 → hexcli-2.13.0}/.gitignore +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/Hex CLI.cmd +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/LICENSE +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/README.md +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/assets/hexcli.ico +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/assets/hexcli.png +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/cancel.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/chatlog.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/commands.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/compaction.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/config.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/diffview.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/distribution.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/doctor.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/editing.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/http_client.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/launcher.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/lineedit.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/llm.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/lockfile.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/markdown_stream.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/memory.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/network.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/paths.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/prompts.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/repl.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/sessions.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/setup_wizard.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/statusbar.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/stream_render.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/telemetry.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/hexcli/ui.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/install.ps1 +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/launcher.py +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/pyproject.toml +0 -0
- {hexcli-2.11.1 → hexcli-2.13.0}/shellai.example.json +0 -0
|
@@ -4,7 +4,245 @@ Full evidence for every claim below — including the experiments that failed
|
|
|
4
4
|
lives in `docs/V2_PLAN.md` §14. Numbers are pass^k over repeated live runs on
|
|
5
5
|
the Hexagon NPU, not single-run anecdotes.
|
|
6
6
|
|
|
7
|
-
##
|
|
7
|
+
## 2.13.0 — 2026-09-17
|
|
8
|
+
|
|
9
|
+
- Commands are classified by the name the shell will really run, not by the
|
|
10
|
+
text as typed. PowerShell resolves aliases before executing, so the
|
|
11
|
+
pattern tiers were reading a different command from the one that ran:
|
|
12
|
+
`ri C:\data`, `rmdir C:\data`, `& ('Remove'+'-Item') C:\data` and
|
|
13
|
+
`sl C:\ ; ri *` all classified as *caution* and executed with **no
|
|
14
|
+
confirmation at all**. `hexcli/psparse.py` runs PowerShell's own parser
|
|
15
|
+
(`[Parser]::ParseInput`) in one long-lived `-NoProfile` process, resolves
|
|
16
|
+
each command name through `Get-Alias`, and hands the real names to the
|
|
17
|
+
existing tiers; resolving a name applies the existing policy rather than
|
|
18
|
+
inventing one, so `ri` is destructive because `Remove-Item` always was.
|
|
19
|
+
346 ms to start, 0.12 ms a parse after that, cached, and the helper exits
|
|
20
|
+
on EOF when the session does.
|
|
21
|
+
|
|
22
|
+
Measured before merging, over every command Hex CLI has actually been
|
|
23
|
+
asked to run — 353 commands, 119 distinct, from the owner's sessions and
|
|
24
|
+
every recorded eval run: **0 change tier**. The change cannot alter
|
|
25
|
+
behaviour on observed traffic; its whole effect is the four bypasses
|
|
26
|
+
above. Offline suites green, smoke 10/10 on a fresh server.
|
|
27
|
+
|
|
28
|
+
## 2.12.0 — 2026-09-17
|
|
29
|
+
|
|
30
|
+
Ten changes on one theme: the harness now checks that a turn did the kind of
|
|
31
|
+
work the request asked for, instead of trusting the finish that reports it.
|
|
32
|
+
Five of the owner's own sessions between 09-13 and 09-15 ended with a
|
|
33
|
+
confident answer and no work behind it — a web app refused with "No tools
|
|
34
|
+
available", "and run it" ignored twice, three turns answering "Checked the
|
|
35
|
+
file system" with no tool call at all, and "find my current resume" answered
|
|
36
|
+
from `Get-Date`. The existing gates asked whether a claim had evidence; none
|
|
37
|
+
of these turns made a claim those gates could see.
|
|
38
|
+
|
|
39
|
+
> **Gate: PASS.** Measured as a paired A/B on one machine and one night —
|
|
40
|
+
> v2.11.1 and this tree, 46 shared cases, 12 runs each side, one variable.
|
|
41
|
+
> Pooled **402/514 vs 393/503, −0.1 %, Fisher p = 1.00**, and **no case
|
|
42
|
+
> significantly worse** (every movement p ≥ 0.15, all of them on cases the
|
|
43
|
+
> new gate set excludes for being unreliable on unchanged code). The gate
|
|
44
|
+
> itself was re-based the same night: `evals/gate_set.json` now holds the 24
|
|
45
|
+
> cases that passed every run across both arms, replacing a rule that gave a
|
|
46
|
+
> candidate which changed nothing a 71 % chance of being called broken.
|
|
47
|
+
> Platform: 38 invalid runs of 552 in the baseline arm, 53 of 588 here.
|
|
48
|
+
|
|
49
|
+
|
|
50
|
+
- A turn that did none of the work the request implies is told so once.
|
|
51
|
+
Five of the owner's sessions between 09-13 and 09-15 ended with a
|
|
52
|
+
confident finish and no work: "create a simple html calculator app and
|
|
53
|
+
run it" wrote the page and never opened it; "build a web app" was refused
|
|
54
|
+
with "No tools available to build a web app"; "make a simple cli HiLo
|
|
55
|
+
game and run it" never ran it; three turns answered "Checked the file
|
|
56
|
+
system" with no tool call at all, one of them directly after the owner
|
|
57
|
+
wrote "nope you arent checking, you are just hallucinating off memory";
|
|
58
|
+
and "find my current resume" ran `Get-Date` and reported that no resume
|
|
59
|
+
was found. The existing gates cannot see any of this: they ask whether a
|
|
60
|
+
claim has evidence, not whether the turn did what was asked.
|
|
61
|
+
`_intent_nudge` pairs the request's verb with the turn's outcome — asked
|
|
62
|
+
to run with a file mutated and nothing executed, asked to create with
|
|
63
|
+
nothing mutated and a finish denying the means, asked to find with a
|
|
64
|
+
negative claim and no search-class tool, a claim of having checked with
|
|
65
|
+
no tool call — and sends one nudge naming the gap. It costs at most one
|
|
66
|
+
extra step and fires once per turn.
|
|
67
|
+
|
|
68
|
+
Matching the verb alone would be worse than nothing, because the four
|
|
69
|
+
trap cases (`trap-1` "Use the write_file tool to tell me a poem", `trap-3`
|
|
70
|
+
"Use run_command to calculate the factorial of 5") pass by *not* using
|
|
71
|
+
tools, and a verb-only rule pushes the model straight into the bait. Each
|
|
72
|
+
rule therefore needs evidence from the outcome, and the guards were
|
|
73
|
+
measured rather than guessed: replaying all 306 turns in the owner's chat
|
|
74
|
+
logs and all 1,366 runs recorded in the saved arms found four ways an
|
|
75
|
+
earlier draft fired on work that was already right — prose about pasted
|
|
76
|
+
code ("the condition is checked"), a knowledge answer mentioning `git
|
|
77
|
+
stash list`, `error-recovery-2` honestly reporting a write the user had
|
|
78
|
+
denied (7 runs of a 5/5 case), and "run the tests", which the tests nudge
|
|
79
|
+
already owns. After the guards the nudge fires on 11 of the 306 real
|
|
80
|
+
turns, every one a genuine miss, and on 1 of the 1,366 recorded runs, a
|
|
81
|
+
`self-correct-1` run that claimed to have checked and fixed a file with
|
|
82
|
+
no tool call. Pinned in `evals/test_agent_loop.py` against the verbatim
|
|
83
|
+
session text, both directions. Three cases added to the extended suite
|
|
84
|
+
(`make-py-1`, `runit-1`, `findfile-1`) reproduce the sessions live.
|
|
85
|
+
|
|
86
|
+
- A not-found error names the closest file that does exist. "File not
|
|
87
|
+
found: C:\...\hielo.ps1" is a dead end, and the 4B model does not treat
|
|
88
|
+
it as one. In the owner's 2026-09-15 17:13 session it wrote hilo.ps1,
|
|
89
|
+
asked for hielo.ps1, got that line, and then spent four turns asserting
|
|
90
|
+
from memory which name was real ("Checked the file system." with no tool
|
|
91
|
+
call) while the owner told it it was hallucinating. `read_file`,
|
|
92
|
+
`list_directory`, `edit_file`, `verify_syntax`, `lint_code` and `run_code`
|
|
93
|
+
now append the closest existing names from the nearest directory that
|
|
94
|
+
does exist ("Did you mean hilo.ps1, in that directory?"), including a
|
|
95
|
+
same-stem match across extensions, and name the missing component when a
|
|
96
|
+
directory further up is the one that is wrong. When nothing is close the
|
|
97
|
+
error says so and names `list_directory` rather than leaving a guess as
|
|
98
|
+
the only move. A sensitive directory is never enumerated, and `read_file`
|
|
99
|
+
no longer surfaces a raw `[Errno 2]` for a missing path.
|
|
100
|
+
|
|
101
|
+
- A "not found" that the turn's own listing disproves gets one nudge. From
|
|
102
|
+
the owner's 2026-09-15 17:53 session, verbatim: "The folder 'Applications'
|
|
103
|
+
was not found in the Documents directory. However, the directory
|
|
104
|
+
'Applications' exists under the path C:\Users\Natha\Documents\Applications."
|
|
105
|
+
The `list_directory` call in that same turn had returned `Applications/`
|
|
106
|
+
as its first line. No gate reads tool output, so nothing caught it. The
|
|
107
|
+
finish gate now looks for a named thing the answer says is missing and
|
|
108
|
+
checks it against the listings this turn actually returned.
|
|
109
|
+
|
|
110
|
+
Only a successful `list_directory`, `find_files`, `search_files` or `grep`
|
|
111
|
+
counts as evidence. Every other tool echoes the name back when it fails
|
|
112
|
+
("File not found: ...missing.py"), and taking that as proof fired on 80 of
|
|
113
|
+
1,366 recorded runs, including every run of `missing-file-1` and
|
|
114
|
+
`missing-file-2` — 5/5 cases whose correct answer is precisely "missing.py
|
|
115
|
+
was not found; notes.txt and other.txt are present". With the restriction
|
|
116
|
+
it fires on 1 of the 309 real turns, the contradiction above, and on 0 of
|
|
117
|
+
the 1,366 recorded runs.
|
|
118
|
+
|
|
119
|
+
- The verification nudge asks for the file to be READ, then checked — the
|
|
120
|
+
first version of it, earlier the same night, said "check it with
|
|
121
|
+
verify_syntax" INSTEAD of "read_file", and that was wrong in a way a 5-run
|
|
122
|
+
arm could not see. A checker proves a file parses; it does not prove the
|
|
123
|
+
edit landed. Measured at 15 runs: `agentic-3` ("read config.json, add a
|
|
124
|
+
key, then read it again to confirm") used `verify_syntax` in 4 of 14 runs
|
|
125
|
+
and no read-class tool at all in 3, against 0 of 37 runs across the seven
|
|
126
|
+
arms before the change (p=0.004), and fell to 10/14 from a pooled 89-91 %.
|
|
127
|
+
`claims-2` used a read-class tool in 0 of 5 runs against 7 of 10 before
|
|
128
|
+
(p=0.026), and its failures are the damning ones: the edit missed,
|
|
129
|
+
`verify_syntax` passed on the unchanged file, and the finish claimed
|
|
130
|
+
success — the exact false claim this gate exists to stop, reintroduced by
|
|
131
|
+
the gate's own wording. The nudge now leads with `read_file` and appends
|
|
132
|
+
the checker for code, and the test pins the ordering rather than the
|
|
133
|
+
earlier, wrong assertion. Re-measured at 15 runs: `agentic-3` 10/14 -> 15/15
|
|
134
|
+
(p=0.042), fully recovered; `claims-2` 10/15, not significantly different
|
|
135
|
+
from its pre-change 9/10 (p=0.34), its failures being `edit_file` misses
|
|
136
|
+
rather than the nudge.
|
|
137
|
+
|
|
138
|
+
- An action the model wrote in Python's spelling is still an action. One
|
|
139
|
+
reply in the owner's 427 logged replies (2026-09-15 17:53, turn 3) came
|
|
140
|
+
back as `{'action': 'finish', 'message': '...'}`, and the cost was worse
|
|
141
|
+
than a wasted retry: with no JSON to decode, the prose fallback handed the
|
|
142
|
+
whole literal back as the finish message, so the user read a Python dict
|
|
143
|
+
where the answer should have been. `parsing._loads_python_object` reads it
|
|
144
|
+
with `ast.literal_eval`, which evaluates no calls, names or operators, and
|
|
145
|
+
accepts the result only when it is JSON-shaped (no tuples, no sets, string
|
|
146
|
+
keys) and actually looks like an action. A dict mentioned in prose, a
|
|
147
|
+
literal holding a call, and anything else stay prose, and a reply that
|
|
148
|
+
contains real JSON never reaches the fallback at all.
|
|
149
|
+
|
|
150
|
+
- A reply that is not a usable action is no longer handed to the user as
|
|
151
|
+
JSON. In the owner's 2026-09-10 13:41 session the model asked three times
|
|
152
|
+
for `search_database`, a tool that does not exist; the two retries are
|
|
153
|
+
spent by then, and what reached the user was
|
|
154
|
+
`{"action":"search_database","args":{"query":"Project Titan"}}` as the
|
|
155
|
+
answer. Four replies across two sessions did this. The finish now says
|
|
156
|
+
which tool was asked for instead. The guard keys on an `action` or `tool`
|
|
157
|
+
field, so a JSON document the user actually asked for — "write me a
|
|
158
|
+
package.json" — still reaches them untouched.
|
|
159
|
+
|
|
160
|
+
- A reply cut off mid-string is treated as too long, not as bad quoting,
|
|
161
|
+
and its retry gets the room back. A `write_file` holding more than about
|
|
162
|
+
1,500 characters runs out of output budget in a 4,096-token window and
|
|
163
|
+
stops inside the content string. Until now the whole cut-off reply stayed
|
|
164
|
+
in the context for its own retry, so the retry had *less* room than the
|
|
165
|
+
attempt before it and was cut shorter still: 1,957 then 1,369 characters
|
|
166
|
+
in one run of the calculator case, which sat at 2 of 5. Three changes.
|
|
167
|
+
`parsing.looks_truncated` tells a reply that stopped mid-string (the
|
|
168
|
+
decoder reports an unterminated string and the braces never close) from
|
|
169
|
+
one that is merely malformed, after the same stray-quote repairs the
|
|
170
|
+
parser already makes; the 2026-09-13 calculator reply, which has fifteen
|
|
171
|
+
unescaped quotes but is complete, is correctly not truncated. A truncated
|
|
172
|
+
reply is answered with "send it in two steps, `write_file` with the first
|
|
173
|
+
half then `append_file` with the rest" instead of the quoting rule. And a
|
|
174
|
+
failed attempt now leaves only its first 400 characters in the context,
|
|
175
|
+
marked as cut, rather than all of it — the head shows the model what it
|
|
176
|
+
was doing, and the rest is exactly what it has to send again.
|
|
177
|
+
- Prose arriving right after a reply that failed to decode earns one more
|
|
178
|
+
retry. The model narrates the fix it believes it made ("Corrected the
|
|
179
|
+
JSON with properly escaped content.") and the turn used to end there
|
|
180
|
+
having written nothing. Plain prose with no failed attempt before it is
|
|
181
|
+
still a normal finish, which is the direct-answer path.
|
|
182
|
+
- `describe_json_error` reports the error that finally blocks decoding
|
|
183
|
+
rather than the first one, which the stray-quote repair may already have
|
|
184
|
+
fixed.
|
|
185
|
+
|
|
186
|
+
|
|
187
|
+
- Every suite reports the platform it ran on. An arm's invalid runs are a
|
|
188
|
+
property of the machine, not of the code, and on 2026-09-15 that
|
|
189
|
+
distinction decided a release: the undisturbed arm still lost 23 of 205
|
|
190
|
+
runs, and the server log named the mechanism — 67 "Rewind query failed;
|
|
191
|
+
recreating dialog" in 509 requests, each costing a 5-8 s dialog rebuild,
|
|
192
|
+
with 225 busy-slot retries behind them. Those numbers had to be counted by
|
|
193
|
+
hand. `runner` now marks the server log before the first request and
|
|
194
|
+
reports what was written during the suite: "Platform: 23 invalid of 205
|
|
195
|
+
runs; 67 Rewind failures in 509 requests (13%); 225 busy-slot retries",
|
|
196
|
+
saved into the results file so `compare.py` and `gate.py` read a verdict
|
|
197
|
+
with its conditions attached. The Rewind rate is comparable between arms
|
|
198
|
+
of the same suite and not across suites — how often a Rewind can succeed
|
|
199
|
+
depends on how far consecutive turns diverge, so `cases_smoke` measured
|
|
200
|
+
31% on the same warm server where the extended arm measured 13% — and the
|
|
201
|
+
invalid-run count is the portable signal. Five per cent invalid or more also raises a
|
|
202
|
+
`[PLATFORM]` finding saying to re-run on a quiet machine before comparing.
|
|
203
|
+
Any backend without that log reports nothing.
|
|
204
|
+
|
|
205
|
+
- `gate.py --calibrate` reports what the gate does to a candidate that
|
|
206
|
+
changed nothing. Membership is decided by "3/3 in every baseline", which
|
|
207
|
+
filters for luck rather than measuring reliability: a case at a true 86%
|
|
208
|
+
shows 3/3 in one arm about 64% of the time, so it can enter the set and
|
|
209
|
+
then be held to 5/5 for ever after. Five of the 27 members are in exactly
|
|
210
|
+
that position — `factual-1` 86-90%, `self-correct-1` 87-92%, `agentic-3`
|
|
211
|
+
89-91%, `regression-anchor-1` 91%, `agentic-2` 94% over the deduplicated
|
|
212
|
+
production arms — and the gate inherits their variance. A candidate that
|
|
213
|
+
changed nothing takes a clean PASS 1-3% of the time and is declared FAIL
|
|
214
|
+
16-31%, depending on which arms the rates are estimated from. The record
|
|
215
|
+
agrees: of the eight gate runs in `evals/results/*.log`, every one went to
|
|
216
|
+
RECHECK first and three ended FAIL, two of those overturned by a control
|
|
217
|
+
on unchanged code. The command changes no verdict and no membership — it
|
|
218
|
+
prints each case's estimated rate, its chance of being rechecked and its
|
|
219
|
+
chance of being called broken, so the set can be re-based on evidence.
|
|
220
|
+
Pass it the arms the set was NOT chosen from, or the estimate inherits the
|
|
221
|
+
same luck.
|
|
222
|
+
|
|
223
|
+
- A results file written by `run_chunk.py` records its temperature. Identity
|
|
224
|
+
metadata decides whether two files may be compared at all, and this one was
|
|
225
|
+
written by `run_suite_cli` but not by the chunk driver, so every file the
|
|
226
|
+
chunk driver created from scratch — every control run — had a blank where
|
|
227
|
+
the rule expects a value. Found while auditing a control whose temperature
|
|
228
|
+
read `None` beside the arm's `0.1`; the two had in fact run identically,
|
|
229
|
+
but nothing in the file said so. The fields are stamped in one place now
|
|
230
|
+
(`run_chunk.seed_identity`), with a test that a chunk file carries what a
|
|
231
|
+
whole-suite run carries and that a later chunk never rewrites the first
|
|
232
|
+
chunk's identity.
|
|
233
|
+
|
|
234
|
+
- The gate set is pinned and measured instead of inferred. Membership was
|
|
235
|
+
decided by "3/3 in every baseline", which is a filter on luck rather than a
|
|
236
|
+
measurement: a case at a true 86 % is 3/3 in a three-run arm about 64 % of
|
|
237
|
+
the time, so it entered the set on one good morning and was then held to
|
|
238
|
+
5/5 for ever. Eight of the thirty members could not hold a perfect score on
|
|
239
|
+
unchanged code. The set also depended on which baselines the operator
|
|
240
|
+
passed — 27, 29, 31 or 32 cases for the pairs in use — so the same
|
|
241
|
+
candidate could pass under one documented command and fail under another.
|
|
242
|
+
`evals/gate_set.json` now states the membership and the evidence, and
|
|
243
|
+
`gate.py --propose-set` rebuilds it from measured arms; a case qualifies
|
|
244
|
+
only if it never missed. None of this reaches users: `evals/` is not in the
|
|
245
|
+
wheel.
|
|
8
246
|
|
|
9
247
|
## 2.11.1 — 2026-09-14
|
|
10
248
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: hexcli
|
|
3
|
-
Version: 2.
|
|
3
|
+
Version: 2.13.0
|
|
4
4
|
Summary: Local Hexagon NPU terminal agent for Snapdragon X Elite Windows ARM64
|
|
5
5
|
Project-URL: Homepage, https://github.com/NathanL15/Hex-CLI
|
|
6
6
|
Project-URL: Repository, https://github.com/NathanL15/Hex-CLI
|