luciazero 2.3.0 → 2.4.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md DELETED
@@ -1,712 +0,0 @@
1
- # Changelog
2
-
3
- All notable changes to this project are documented in this file.
4
-
5
- Format: [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
6
- Versioning: [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
7
-
8
- ## [Unreleased]
9
-
10
- ## [2.3.0] - 2026-08-16
11
-
12
- ### Fixed
13
-
14
- - `test.sh` clears ambient `LUCIAZERO_*` variables before running. An exported
15
- `LUCIAZERO_VERIFY_CMD` flipped the hook fixtures into exact-match mode, so
16
- the suite went red on exactly the machines that dogfood the pack
17
- (`FAIL: stop hook nudged despite verify after edit`) while CI stayed green.
18
- A new self-test re-runs the fast tier in a child poisoned with every knob the
19
- hooks read, and quotes the child's own failing line.
20
- - The verify hook parses under bash 3.2 again — the `/bin/bash` on stock macOS.
21
- A here-document inside a command substitution (with a quoted expansion and a
22
- trailing redirection on the same line) breaks that parser, and it rejects the
23
- **whole file** at load time while pointing at an unrelated later line, so the
24
- enforcement pack silently did nothing there. The scanner program now lives in
25
- a variable. `test.sh` rejects the construct in the hooks and, with
26
- `LZ_BASH32=/path/to/bash-3.2`, parses every script with the real thing.
27
- - Both hooks call `hashlib.md5(..., usedforsecurity=False)` for their state
28
- directory name. On a FIPS-enforcing python3 the bare call raised and the
29
- tracker failed open — silently doing nothing. The digest is unchanged, so
30
- existing state keys still resolve.
31
- - CI now parses every repository shell script with Bash 3.2.
32
-
33
- ### Security
34
-
35
- - A repository's **committed** `.claude/settings.json` can no longer configure
36
- Luciazero at all: every `LUCIAZERO_*` key declared there is dropped and the
37
- hook falls back to its own defaults. `LUCIAZERO_VERIFY_REGEX` and
38
- `LUCIAZERO_VERIFY_CMD` could make any command count as a verify run,
39
- `LUCIAZERO_DOC_REGEX='.*'` made every edit look like documentation so nothing
40
- was ever unverified, and `LUCIAZERO_STRICT_VERIFY_CMD` was a command the stop
41
- hook would run. A committed `CLAUDE_CONFIG_DIR` is refused for the same
42
- reason: it could point at a repository-controlled "wired classic install" and
43
- make every hook copy stand down. Only the default `~/.claude` is treated as
44
- the user's config directory during the search — honouring `CLAUDE_CONFIG_DIR`
45
- there let a repository point it at its own `.claude` so the scanner skipped
46
- the file declaring the key. The search covers the session directory and
47
- its ancestors — Claude Code merges project settings from the repository root
48
- and a session's cwd is often a subdirectory — but it is **project scope
49
- only**: it stops at the repository root, at `CLAUDE_PROJECT_DIR`, and at
50
- `$HOME`, and never reads the user's own config directory, so a global
51
- `~/.claude/settings.json` keeps configuring the hook.
52
- `SessionStart` names the refused keys once. The
53
- personal, gitignored `.claude/settings.local.json` is untouched, a parse error
54
- leaves values alone (still fail-open), the lookup runs only in the modes that
55
- consume a knob, skips a non-regular file (a planted fifo would hang the hook),
56
- and refuses everything outright on an absurdly large settings file instead of
57
- parsing it.
58
- - Channel dedupe is decided from the running copy's own path instead of
59
- `LUCIAZERO_CHANNEL`. A committed `env` block could set that variable, hand the
60
- **classic** hook a plugin label, and make it stand itself down — disabling
61
- enforcement with one line.
62
- - `install.sh --with-hooks` refuses a python3 older than 3.9 (where hashlib
63
- gained `usedforsecurity=`) instead of installing hooks that fail open, and
64
- `--status` reports the version. README states the requirement.
65
- - CI runs with `permissions: contents: read`, and both workflows pin every
66
- action to a commit SHA.
67
-
68
- ### Changed
69
-
70
- - `shellcheck` is required, not silently skipped, when `CI` or
71
- `LZ_REQUIRE_LINT` is set — a local green must not disagree with the CI that
72
- gates the release.
73
- - README and README.th document the committed-settings refusal and the
74
- Windows/WSL requirement.
75
- - Retired the `/luciazero-bootstrap` compatibility alias after the `/ready`
76
- migration window. Installers remove only untouched Luciazero-owned copies and
77
- preserve customized user directories.
78
- - Capped skill discovery descriptions at 40 words and tightened path/offline
79
- guidance for the closeout, debug, ready, retro, and discipline skills.
80
-
81
- ## [2.2.0] - 2026-08-15
82
-
83
- ### Added
84
-
85
- - `/imouto-mode` adds Lucia's optional warm, lightly tsundere younger-sister
86
- coding voice. It is explicit-only, off by default, invocation-scoped, and
87
- keeps technical work, safety, and verification ahead of persona.
88
-
89
- ## [2.1.0] - 2026-08-15
90
-
91
- ### Added
92
-
93
- - `/show` turns code relationships, structural changes, and verification
94
- evidence into the smallest useful traceable visual.
95
- - `./test.sh --fast` provides a measured intermediate tier while the default
96
- and `--full` preserve complete CI/closeout coverage.
97
- - Opt-in Claude hooks use private per-session scratch state to record
98
- privacy-preserving aggregate turn/merged-Bash wall time and Bash, failed or
99
- successful verify, and model/user skill counts; `luciazero discipline`
100
- summarizes them.
101
-
102
- ### Changed
103
-
104
- - `/luciazero-bootstrap` is now `/ready`. The old command remains as a
105
- deprecated compatibility alias for one release.
106
- - The doctrine reserves full verification for closeout, and `/plan` plus
107
- `/debug` no longer auto-trigger for routine edits or first obvious failures.
108
-
109
- ## [2.0.3] - 2026-08-13
110
-
111
- ### Fixed
112
-
113
- - The Claude plugin now exposes its reviewer from the default root `agents/`
114
- directory. Claude Code 2.1.227 validated the former custom manifest path but
115
- reported and loaded zero agents; the default layout reports one. CI requires
116
- the plugin mirror to match the classic install source byte-for-byte and
117
- prevents the custom path from returning.
118
-
119
- ## [2.0.2] - 2026-08-13
120
-
121
- ### Fixed
122
-
123
- - `/lucia-relay` now records whether its recipient is on the same machine or a
124
- different one before producing pointers. Schema 2 permits full paths only
125
- for same-machine delivery; cross-machine validation requires a clean HEAD
126
- reachable from a locally known remote branch, rejects machine-only paths,
127
- and provides `knowledge.inline` for otherwise-local context. Schema 1 relays
128
- remain readable as same-machine-only legacy artifacts, and legacy `draft`
129
- callers without `--recipient` safely default to same-machine delivery.
130
-
131
- ## [2.0.1] - 2026-08-13
132
-
133
- ### Added
134
-
135
- - **Explicit, channel-aware updates.** `npx luciazero@latest check-update`
136
- performs a read-only, five-second npm registry check only when invoked;
137
- `npx luciazero@latest update` detects classic Claude, its hook mode, and
138
- Codex, then refreshes every detected install through the existing audited
139
- installers. It refuses to create a fresh install, downgrade a recognized
140
- newer one, or trust a malformed version sidecar. Legacy installs without the
141
- sidecar remain updatable. Classic doctrine customization now receives the
142
- same managed-snapshot backup protection as skills and agents.
143
- Plugin and skills-only update commands, Claude plugin auto-update, and GitHub
144
- release notifications are documented separately in both READMEs.
145
-
146
- ### Changed
147
-
148
- - The release workflow uses the runner's GitHub CLI instead of a Node 20-based
149
- release action. Re-runs replace the existing zip without recreating the
150
- release, and GitHub Actions no longer emits the deprecated-runtime warning.
151
-
152
- ## [2.0.0] - 2026-08-13
153
-
154
- ### Changed
155
-
156
- - The canonical Sonnet result is explicitly preliminary (`+31pp`, n=4–5).
157
- The historical `+37pp` statement is retired because its eight replacement
158
- raw rows could not be recovered.
159
- - `eval/report.sh` rejects mixed campaigns/commits/seeds, changed fixture
160
- hashes, duplicate invocation IDs, and inconsistent pair order. Published
161
- evidence also enforces registered task/arm/row/invalid/model expectations,
162
- and discloses Haiku's incomplete per-row model provenance.
163
- - Eval tasks may provide deterministic offline setup before either arm. Provider
164
- transcripts now live outside worked trees so they cannot alter Git status,
165
- repository fingerprints, or final-tree grading.
166
- - Relay fingerprints encode untracked special files without opening them, so a
167
- FIFO, socket, or device cannot block inspection or trigger device I/O.
168
- - GitHub workflows use the Node 24-based `actions/checkout@v5` and
169
- `actions/setup-node@v5` runtimes.
170
- - **Breaking: `/handoff` is now `/lucia-relay`.** The branded name avoids
171
- collisions with generic handoff skills. Installs remove an untouched v1.5
172
- copy but preserve and warn about customized copies. Relay state is now a
173
- validated `LUCIA_RELAY.json` manifest plus a generated human view, with a
174
- repository fingerprint, verification evidence, negative knowledge,
175
- cross-session/cross-agent routing, drift inspection, and explicit consume.
176
- - **Risk-routed review.** The single portable reviewer now accepts `general`,
177
- `security`, and `contract` focus modes, reads callers/consumers, and uses one
178
- blocker/major/minor policy. `/done` requests separate focused passes when a
179
- diff crosses both security and contract boundaries.
180
- - **Smart verification is repo-owned.** Monorepos create `verify-changed` from
181
- their native task graph and keep `verify-full` for closeout. The global hook
182
- never guesses dependency impact from path prefixes.
183
- - **Classic and Codex installs track component ownership.** Exact hidden
184
- snapshots distinguish Luciazero-managed skills/agents from same-name user or
185
- third-party components. Updates back up collisions/customizations, and
186
- uninstall removes only an unchanged managed copy.
187
-
188
- ### Added
189
-
190
- - Auditable benchmark evidence: canonical Claude raw JSONL, a SHA-256 campaign
191
- registry, generated README/benchmark tables, and a CI drift check.
192
- - Result schema 2 records campaign, pair, invocation, repository, fixture,
193
- prompt, platform, and arm-order metadata. Seeded arm randomization reduces
194
- fixed-order bias without making campaigns irreproducible.
195
- - Strict shared result validation rejects unsupported schemas and mistyped
196
- booleans, criteria, metrics, timestamps, platform, and campaign metadata.
197
- Output-aware `--resume` fills interrupted pairs without rerunning completed
198
- invocation IDs; `--run-offset` extends completed batches.
199
- - Three zero-quota candidate eval tasks cover archive extraction security,
200
- lossless atomic schema migration, and multi-page cursor integration. Each
201
- grader proves reference/project/anti-gamed behavior offline.
202
- - **`relay-transfer` protocol eval** grades portable state, an exact next edit,
203
- verification evidence, negative knowledge, scope preservation, and a current
204
- repository fingerprint. CI proves its 6/6 reference and rejects generic
205
- prose plus a content-complete stale relay without spending model quota.
206
- - **Lucia Relay demo** drives the shipped producer/receiver implementation in
207
- a temporary Git repository: render, validate, detect drift, re-run evidence,
208
- and explicitly consume. The checked-in GIF is generated from the same script
209
- exercised by CI.
210
- - **Central component catalogs** drive classic/Codex install, status,
211
- uninstall, and inventory tests, so a new skill or agent cannot silently ship
212
- through only one channel.
213
- - **`/plan`** defines falsifiable acceptance signals and reversible steps,
214
- while pausing for approval only on ambiguity, high stakes, destructive work,
215
- public-contract choices, or scope changes.
216
- - **`/bisect` + `safe-bisect.sh`** locate the first bad commit in a detached
217
- temporary worktree, repeat endpoints to catch flakes, preserve exit 125
218
- skips, distinguish missing commands, and clean every exit path.
219
- - **`npx luciazero discipline` + `/discipline-report`** analyze schema-v2
220
- local JSONL outcomes with day/project filters and JSON output. The hook logs
221
- a privacy-preserving project hash and verify mode; legacy records remain
222
- readable and recommendations distinguish observations from likely causes.
223
- - **`/lucia-relay` carries memory pointers** — the `Read first`
224
- section quotes the `docs/lessons.md` entries relevant to the unfinished
225
- work (a selection, never a copy — the ledger travels with the repo) and
226
- copies applicable machine-local `luciazero-heuristics.md` entries
227
- verbatim, since the relay is the only way those cross machines. The
228
- consume protocol tells the reader to follow the pointers before touching
229
- code and to adopt carried heuristics that earn their keep.
230
-
231
- - **`eval/run.sh --use-login`** — run the real eval on an existing Claude
232
- subscription (Pro/Max) instead of API dollars: seeds each per-run sandbox
233
- config dir with this machine's login state — `~/.claude.json`, plus
234
- OAuth tokens from `.credentials.json` (Linux) or a Keychain export
235
- (macOS). The copy lives only inside the mktemp sandbox and is deleted
236
- with it. Fail-soft by
237
- design: if the seed does not authenticate, `check-result.sh` marks the
238
- arm INVALID and nothing is spent. Plumbing (seed per arm, warn on missing
239
- login state) is proven offline in `test.sh`.
240
-
241
- - **`eval/check-result.sh`** — a zero exit code no longer proves the agent
242
- ran: the CLI has wrapped a `Not logged in` error in subtype `"success"`
243
- (observed 2026-08-11, caught free by a 1-run smoke). The guard inspects
244
- the result payload (`is_error`, `terminal_reason: api_error`, login
245
- errors) and `run.sh` books a refuted arm as INVALID with the reason
246
- quoted; every accept/reject path is fixture-proven in `test.sh`.
247
- - **`eval/run.sh --offline`** — synthetic smoke mode: no `claude` CLI, no
248
- API key, zero cost. Doctrine-style arms get the task's `reference/` tree,
249
- bare keeps the planted bug, and the whole copy → grade → JSONL → report
250
- loop runs in seconds. Rows are branded `"offline": true` and `report.sh`
251
- prints a SYNTHETIC banner so the numbers can never pass as behavioral
252
- results; end-to-end proven in `test.sh` plus a frozen fixture pair.
253
- - **README narrative reorder** (EN + TH): the demo GIF and the
254
- "What it prevents" table now sit directly under the intro, before
255
- Install — a newcomer sees what the pack does in 30 seconds before being
256
- asked to install anything.
257
-
258
- - **README demo GIF** — 15 seconds of the enforcement pack's real behavior:
259
- edit → `✎ unverified`, stop attempt → the rule-1 nudge, red verify →
260
- `❌ verify RED`, fix → `✅ verify`. Recorded from the checked-in
261
- `docs/assets/statusline-demo.sh`, which drives the shipped hooks in a
262
- sandbox (so the GIF cannot drift from what the scripts actually print),
263
- via the checked-in `docs/assets/demo.tape` (`vhs`). The driver script is
264
- under `test.sh`'s shellcheck net.
265
-
266
- - **`false-green` eval task** (sixth): the false-done trap — the shipped
267
- suite is green from the start while the CSV escaping bug lives outside
268
- its coverage. The untouched tree *passes its own tests* and still fails
269
- the grader (symptom probed on unseen data; bug-restored suite must go
270
- red), which is doctrine rule 1 stated as a fixture. `gamed/` (comma-only
271
- half fix) and `gamed-notest/` (correct fix, no test added) are rejected.
272
- - **`--with-lessons` eval arm** — `eval/run.sh --with-lessons` runs a third
273
- arm for tasks that ship a `lessons.md` (currently `pipeline` and
274
- `false-green`): doctrine install plus the task's ledger pre-seeded as
275
- `docs/lessons.md`, the A/B/C comparison that measures whether the
276
- learning layer pays. `report.sh` discovers arm columns from the data and
277
- renders per-arm deltas; frozen three-arm fixture added to `test.sh`.
278
- - **Per-run resource accounting** — `run.sh --out` now records duration,
279
- token usage, and cost per run (parsed fail-open from the CLI's
280
- `--output-format json` result), and `report.sh` appends per-arm resource
281
- means whenever the data is present — a pass-rate delta is only a win if
282
- the cost next to it says so.
283
- - **Community eval issue templates** — `Evaluation result` (report.sh
284
- output required, null results explicitly welcome) and `New eval task`
285
- (asks for the doctrine rule probed and the gamed tree that would cheat
286
- the grader).
287
-
288
- - **`merge-conflict` eval task** (fifth): an unresolved merge where main's
289
- bulk discount and the branch's member discount must both survive. The
290
- grader probes each feature on data the shipped tests never mention, and
291
- swaps in one-sided feature mutants to prove the worked tests actually
292
- cover both sides — `gamed/` (HEAD-only resolution, suite green) and
293
- `gamed-notests/` (correct merge, no tests added) are both rejected.
294
- - **Machine-readable closeout evidence** — `/done` step 6 now mirrors the
295
- report as a JSON block (status, verify command + exit code + decisive
296
- line, not-covered, left-out) when the result feeds CI, a PR comment, or a
297
- dashboard.
298
- - **README "What it prevents" section** (EN + TH): failure modes mapped to
299
- the shipped mechanism that catches each — no promise without a mechanism.
300
- - **SECURITY.md** — private reporting channel plus the enforced design
301
- guarantees (no network, no npm lifecycle scripts, fail-open hooks,
302
- config-dir-only writes) and the documented strict-mode env sharp edge.
303
- - **GitHub issue templates** — bug report (channel + decisive-output
304
- evidence required), feature request (doctrine-fit question), security
305
- contact link.
306
-
307
- ## [1.5.0] - 2026-08-10
308
-
309
- First version published to npm (`luciazero`); 1.4.x and below were
310
- development versions.
311
-
312
- ### Added
313
-
314
- - **Learning layer** — the pack now compounds experience across sessions,
315
- three stores, all mechanized and all pruned:
316
- - `/retro` records debugged failures to a per-repo lesson ledger
317
- (`docs/lessons.md`, fixed greppable shape: symptom → cause → proven-by →
318
- fix) and repo-independent lessons to `luciazero-heuristics.md` in the
319
- harness config dir (one line each, hard 100-line cap); stale entries are
320
- corrected or deleted, since a wrong lesson mis-seeds every future debug.
321
- - `/debug` seeds its hypothesis ledger from both files before inventing
322
- hypotheses — a match becomes H1 but is still verified.
323
- - The stop hook appends one line per stop outcome (`stop-clean` / `nudge` /
324
- `strict-block`) to `luciazero-stats.log` in the config dir — local only,
325
- fail-open, rotated at 500→250 lines — and `/retro` reads it to turn
326
- recurring discipline gaps into recorded lessons. This is the one
327
- documented exception to "state never leaves $TMPDIR"; the hook header
328
- says so.
329
- - Both uninstallers keep (and mention) the learned-data files.
330
- - test.sh: stats logging proven for all three outcomes + rotation +
331
- uninstall survival, learning-layer wiring greps on both skills, and the
332
- whole run now exports a sandbox `CLAUDE_CONFIG_DIR` so no test can ever
333
- write to the real `~/.claude`. New checks red-proven by mutation.
334
-
335
- ## [1.4.1] - 2026-08-10
336
-
337
- ### Added
338
-
339
- - **Trusted publishing (OIDC)**: `release.yml` gained an `npm-publish` job —
340
- every `v*` tag now publishes to npm from GitHub Actions with provenance
341
- attestations and no token anywhere. Guards: the tag must equal the
342
- `package.json` version, and already-live versions are skipped so re-runs
343
- cannot fail. This release exists to exercise that pipeline end to end.
344
-
345
- ### Fixed
346
-
347
- - The README shipped inside the npm tarball no longer carries the
348
- "npm publish is in flight" sentence that 1.4.0 froze in.
349
-
350
- ## [1.4.0] - 2026-08-10
351
-
352
- ### Changed
353
-
354
- - **Project renamed to Luciazero** (from "agentic-engineering"). Every brand
355
- identifier moved with it: doctrine file `claude/luciazero.md` (imported as
356
- `@luciazero.md`), hooks `luciazero-verify.sh` / `luciazero-statusline.sh`,
357
- skill `/luciazero-bootstrap`, env vars `LUCIAZERO_*` (was `AGENTIC_*`),
358
- version sidecars `.luciazero-version`, CI example
359
- `examples/luciazero-ci.example.yml`, hook state dir `luciazero-verify-state`,
360
- and the settings-cleanup matchers in both uninstallers. Nothing was
361
- published under the old name, so there is no migration path to keep.
362
- Prose still uses "agentic engineer(ing)" where it names the discipline,
363
- not the project.
364
- - **Skills moved to the repo root** (`skills/`, was `claude/skills/`) so
365
- `npx skills add <owner>/luciazero` (vercel-labs/skills) discovers them with
366
- zero registration. Installers, tests, and docs all read the new path.
367
-
368
- ### Added
369
-
370
- - **Claude Code plugin packaging**: `.claude-plugin/plugin.json` +
371
- `.claude-plugin/marketplace.json` make the repo installable as a plugin
372
- from its own single-plugin marketplace (`/plugin marketplace add
373
- <owner>/luciazero` → `/plugin install luciazero@luciazero`); `claude plugin
374
- validate` passes. `claude/hooks/hooks.json` wires the verify hooks via
375
- `${CLAUDE_PLUGIN_ROOT}`, and a new `doctrine` subcommand of
376
- `luciazero-verify.sh` loads the doctrine as SessionStart context — plugins
377
- cannot add a CLAUDE.md import line — with a guard that stays silent when a
378
- classic install exists, so the doctrine never loads twice. Honest limits
379
- documented: no statusline via plugins; pick one channel so hooks are not
380
- wired twice.
381
- - **npm wrapper** (`package.json` + `bin/luciazero.js`): `npx luciazero`
382
- routes to the bundled installers (`codex`, `uninstall`, `uninstall-codex`
383
- subcommands; flags pass through). No lifecycle scripts, ever — test.sh
384
- fails if one appears, matching npm v12's default block.
385
- - **Lucia mascot** in both READMEs, cropped from the project's character
386
- sheet (`docs/assets/lucia*.png`): plushie-hug under the title, laptop pose
387
- at the eval paragraph, and the fist-up pose celebrating
388
- `PASS all checks green` in Development.
389
- - **README rewrite (both languages)**: install-channels-first — plugin
390
- (recommended) and `npx skills add` lead, classic `git clone` demoted to the
391
- reference channel; sections condensed; the stale "Why not a Claude Code
392
- plugin" design note replaced with "How the plugin squares with this"; new
393
- "Lucia family & support" section (Lucia Discord bot + donate link).
394
- A 2-lens verification pass (bilingual fidelity + truth-to-code) confirmed
395
- the new claims and caught 4 wording issues, all fixed.
396
- - **docs/publishing.md**: dependency-ordered release checklist (GitHub →
397
- plugin directory submission → npm trusted publishing → awesome-claude-code),
398
- with the channel-honesty note that only the classic installer carries the
399
- statusline and CLAUDE.md import.
400
- - test.sh grew five gates for the above: manifest validity + version sync
401
- across CHANGELOG/plugin.json/package.json, doctrine-mode behavior (emits
402
- once, never twice), npm payload completeness + lifecycle-script ban
403
- (with a live `--status` routing probe when node is present), the plugin
404
- channel dedupe, and the installers' unknown-option rejection. All five
405
- proven red by mutation before being trusted.
406
-
407
- ### Fixed (post-review of the rename/packaging wave; 11 confirmed findings)
408
-
409
- - `install-codex.sh`, `uninstall.sh`, and `uninstall-codex.sh` now reject
410
- unknown options — previously `npx luciazero codex --status` silently
411
- performed a FULL install instead of a status check, and stray flags to the
412
- uninstallers were swallowed.
413
- - Plugin doctrine mode no longer needs python3 or stdin: it is handled before
414
- the script's shared setup, so machines without python3 (where every other
415
- mode fails open to doing nothing) still load the doctrine. It also survives
416
- an unset `HOME` (was an `set -u` abort violating the fail-open contract).
417
- - Running any hook mode by hand from a terminal no longer hangs waiting for
418
- stdin EOF.
419
- - Plugin + `install.sh --with-hooks` double-install: the plugin's hooks.json
420
- now invokes every mode with `LUCIAZERO_CHANNEL=plugin`, and the hook stands
421
- down when classic wiring exists in settings.json — the stop nudge can no
422
- longer double-fire, and a strict verify command can no longer run twice
423
- concurrently against the same repo.
424
- - The debug and bootstrap skills no longer hardcode classic-install paths
425
- (`~/.claude/...`) that do not exist under the plugin / `npx skills`
426
- channels.
427
- - Release procedure docs caught up with the version-sync gate: CONTRIBUTING's
428
- Releasing step and docs/publishing.md now both say to bump plugin.json +
429
- package.json together with the CHANGELOG heading (following the old steps
430
- verbatim would have produced a red release workflow), and publishing.md no
431
- longer hardcodes tagging the already-released v1.3.0.
432
- - docs/comparison.md no longer lists "plugin marketplace" as something
433
- superpowers has and we don't (this repo is now its own single-plugin
434
- marketplace; theirs remains a multi-plugin ecosystem).
435
-
436
- - **Strict verify gate** (opt-in on top of the opt-in enforcement pack): set
437
- `LUCIAZERO_STRICT_VERIFY_CMD` in your *personal* settings and the Stop hook
438
- actually runs that command (fast-pathing when the tracked state is already
439
- green after the last edit) and blocks a red stop with the failing output
440
- quoted. Hard timeout (`LUCIAZERO_STRICT_TIMEOUT`, default 120s); every
441
- internal error degrades to the ordinary fail-open nudge. The variable
442
- belongs in personal settings only; the hook cannot verify which settings
443
- scope set it (a committed `.claude/settings.json` env block reaches it
444
- too), and the docs say so plainly — never commit it, and treat a repo
445
- that ships it as hostile. Documented honestly as a speed bump, not a
446
- wall: a blocked stop's continuation is never re-blocked.
447
- - **Exact-match verify tracking**: `LUCIAZERO_VERIFY_CMD` switches the Bash
448
- tracker from the broad regex to prefix matching, closing a real
449
- false-green — `cat test.sh` or `grep pytest README` no longer count as a
450
- verify run. `/luciazero-bootstrap` Phase 2 now offers (ask-first) to record
451
- the established command in the repo's `.claude/settings.local.json`.
452
- - **SessionStart handoff pointer**: a `session` hook subcommand emits one
453
- context line when the project has a `HANDOFF.md` capsule — age included,
454
- stale warning past `LUCIAZERO_HANDOFF_STALE_DAYS` (default 7), silent and
455
- zero-cost when there is none, pointer only (never the contents).
456
- - **`revert-probe.sh`** (ships inside the done skill, works on both
457
- harnesses): the mechanical form of "would the new tests fail if the change
458
- were reverted?" — checks the pre-change code into a throwaway git
459
- worktree, overlays only the changed test files, runs the verify command
460
- there and inverts the result. Exit 0 tests bite / 1 vacuous or no test
461
- changes / 2 unassessable; never touches the caller's tree. `/done` and
462
- `/debug` reference it.
463
- - **Three new eval tasks**, each probing a different doctrine rule:
464
- `red-suite` (correct-but-red suite; the lazy fix is bending the tests to
465
- the bug — caught by replaying the fixture's pristine tests against the
466
- worked code), `flaky-report` (hash-seed-dependent output; graded
467
- deterministically via a `PYTHONHASHSEED` 0–9 sweep), `pipeline` (bug in
468
- the parser, symptom two modules away; graded by diff *locality* — the
469
- untouched modules must stay AST-identical).
470
- - **`gamed*/` cheat fixtures + grader auto-discovery**: every task now ships
471
- one or more hand-built cheat trees its grader must reject — including
472
- hardcoded-lookup (`red-suite/gamed-hardcode/`) and hardcoded-output
473
- (`flaky-report/gamed-hardcode/`) variants killed by unseen-data criteria —
474
- and `test.sh` auto-discovers `eval/tasks/*/` so no task can ship without
475
- proving its grader goes red, green, and anti-gamed (a missing `gamed/` is
476
- itself a red build) and speaks the new machine-readable
477
- `CRIT <id> pass|fail` / `SCORE n/m` output contract.
478
- - **`eval/run.sh --runs N --out results.jsonl` + `eval/report.sh`**: repeat
479
- runs, record per-criterion results as JSONL, and render the doctrine-vs-
480
- bare pass-*rate* table the honesty box has always prescribed — with n and
481
- a low-n warning printed unconditionally. `report.sh` is byte-compared
482
- against a frozen fixture in CI and rejects malformed input.
483
- - **`install.sh --status`**: read-only health check of an existing install —
484
- every piece listed, hook wiring verified in `settings.json` (the hooks
485
- fail open, so a broken install was previously silent), version compared,
486
- non-zero exit when a core piece is missing. Plus a version sidecar
487
- (`.luciazero-version`, both harnesses) and a documented update
488
- path in the README.
489
- - **`demo.sh`**: scaffolds the slugify planted-bug fixture into a throwaway
490
- git repo, prints the bug report and the exact commands — fix it in your
491
- own Claude session, then score the tree with the offline grader. Never
492
- invokes `claude` itself; refuses to scaffold inside the repo.
493
- - **`docs/comparison.md`**: dated, sourced, deliberately two-sided
494
- comparison against superpowers, SuperClaude, proof-loop, orchestrator
495
- runtimes, template catalogs, and the harness built-ins.
496
- - README: 60-second quickstart, a "What it looks like" section showing the
497
- actual statusline/nudge/strict-gate output (captured, not composed), and
498
- an Updating section.
499
- - **`README.th.md`** — full Thai translation of the README, replacing the
500
- abridged Thai section; English stays the default, both files cross-link,
501
- and `test.sh` trips when the section structures drift apart.
502
-
503
- - `/done` skill — closeout ritual before declaring a non-trivial task
504
- complete: full-tier verify with the decisive line quoted, a skeptic pass
505
- over the final diff, an independent adversarial review when the diff
506
- earns it, an explicit scope check naming anything left out, and a fixed
507
- report format. Doctrine rule 1 now points to it.
508
- - `/handoff` skill — transient state capsule (`HANDOFF.md`) for resuming
509
- unfinished work across sessions, machines, or harnesses: goal, verified
510
- state, one literal next command, open and refuted hypotheses, landmines.
511
- Consumed and deleted by the reader; `/retro` stays the home of permanent
512
- lessons.
513
- - `/experiment` skill — measured-change protocol for optimization work:
514
- metric and win threshold defined before any change, multi-run baseline,
515
- one variable per experiment, verdicts (including null results) recorded
516
- to `docs/experiments.md`, losers reverted immediately.
517
- - Enforcement pack (`./install.sh --with-hooks`, Claude Code only,
518
- requires python3): a verify-tracking hook pair plus statusline —
519
- PostToolUse hooks record edits and verify-ish Bash runs per project, a
520
- Stop hook nudges once (fails open, never loops) when a session ends
521
- with unverified edits, and the statusline shows `model | branch |
522
- ✅ verify 3m` / `❌ verify RED` / `✎ unverified` at a glance. The
523
- settings.json merge is additive, idempotent, backed up, respects an
524
- existing custom statusLine, and `uninstall.sh` removes exactly our
525
- entries while preserving everything else.
526
- - `examples/luciazero-ci.example.yml` — inert, REPLACE-ME-gated GitHub
527
- Actions template: on CI failure, an agent diagnoses the root cause from
528
- the failing logs (hypothesis + evidence line, logs treated as untrusted
529
- input) and posts a size-capped PR comment. Diagnosis only — it cannot
530
- push or edit code (`contents: read`, tool allowlist without Bash,
531
- `persist-credentials: false`); its single write scope is
532
- `pull-requests: write` for the comment. Fork-guarded secrets, no
533
- auto-fix.
534
- - `eval/` — A/B harness measuring whether the doctrine changes agent
535
- behavior: planted-bug task fixtures graded offline by behavioral
536
- criteria (bug actually fixed, a regression test that goes red when the
537
- buggy implementation is restored, no weakened checks). `eval/run.sh`
538
- runs both arms (doctrine vs bare config) and costs API money, so it is
539
- manual; `test.sh` verifies the graders themselves can go both red and
540
- green, offline.
541
- - `test.sh` now also exercises the enforcement-pack hook state machine,
542
- the `--with-hooks` install/uninstall cycle against a settings.json with
543
- pre-existing user content, the eval graders' red/green behavior, and the
544
- inertness of the luciazero-ci example.
545
-
546
- - `/debug` skill — hypothesis-driven debugging procedure, the on-demand
547
- expansion of the doctrine's hypothesis rule: deterministic reproduction
548
- first, minimized repro, a visible hypothesis ledger (run the refuting
549
- observation, not the edit), one variable per iteration with failed fixes
550
- reverted, close-out via a regression test red before the fix and green
551
- after. Installed by both harness installers.
552
- - `luciazero-bootstrap` now bundles `scripts/detect.sh` — a read-only,
553
- dependency-free evidence scan (bash plus standard tools; python3 to
554
- parse `package.json` when available, `sed` fallback) covering docs,
555
- manifests, script/target names, CI `run:` lines, test dirs, monorepo
556
- markers, and git status, replacing a dozen manual reads in Phase 1. It surfaces candidates only; the agent still decides.
557
- Ships to both harnesses via the existing skill copy.
558
- - Bootstrap Phase 2 hardening: verify must run unattended (no watch or
559
- interactive modes), the suite is timed once so the measurement (not a
560
- guess) decides one tier or two, fast-tier output should be near-silent,
561
- and monorepos scope the fast tier to the package being changed.
562
- - Bootstrap Phase 6 rewrite: run the fast tier twice to catch flakes,
563
- break a line a smoke test actually *covers* (breaking an uncovered line
564
- proves nothing), restore via `git checkout`/`git stash`; Phase 1 now
565
- reports git-repo status and proposes `git init` (ask first) for
566
- unversioned dirs.
567
- - `/retro` routing gate: lessons true for anyone who clones the repo go to
568
- committed notes; machine-local or personal facts go to the harness's
569
- memory system (Claude Code's per-project `memory/` dir + `MEMORY.md`
570
- index) and are never committed; on Codex (no memory system) only the
571
- machine-independent generalization is kept — an honest gap beats a note
572
- no harness loads.
573
- - `test.sh` now enforces the doctrine's word-count budget (≤420 words) and
574
- platform-neutral vocabulary, smoke-runs `detect.sh` against the repo
575
- itself, and covers the new skill and script in both sandbox cycles.
576
- - Example settings: inert check-suppression guard hook (blocks edits that
577
- add `noqa`/`ts-ignore`/`.skip(`-style markers, mechanizing "never weaken
578
- a check"), and the derived-files hook example now explains scoping by
579
- `tool_input.file_path` so it does not run on every edit.
580
- - `reviewer` agent: `model: inherit` frontmatter — an adversarial reviewer
581
- on a weaker model than the author defeats its purpose. The Codex
582
- transform drops `model:` alongside `tools:`.
583
-
584
- ### Changed
585
-
586
- - **Doctrine cut from 15 rules to 9** (568 → ~415 words). Removed
587
- outright — each relies on a behavior 2026 harnesses enforce by default;
588
- if a harness regresses, restore from here: style matching (old R7
589
- tail), read-before-overwrite (old R8 tail — mechanically enforced by
590
- Write tools), wide-read delegation (old R10 — also impossible on
591
- Codex). Folded into surviving rules, not cut: faithful run reporting
592
- (old R3 → one clause in new R1), the loop itself (old R5 → preamble),
593
- read-project-notes-first (old R12 → new R8), routine-work-without-
594
- blocking (old R13 → new R9), whole-scope completion and honest handback
595
- (old R15 / R5 tail → closing clause of new R9). The stop-and-ask list
596
- survives as new R9.
597
- - Doctrine rule 7 routes risky/wide diffs through the harness's built-in
598
- review command when one exists (Claude Code: `/code-review`), with the
599
- shipped `reviewer` agent as the portable fallback and the only reviewer
600
- on Codex.
601
- - Doctrine and skills are now platform-neutral: "notes file
602
- (`CLAUDE.md` / `AGENTS.md`)" instead of assuming one harness; bootstrap
603
- Phase 4 marks hooks/settings/`/fewer-permission-prompts` as Claude-only
604
- and tells Codex sessions to encode the same guardrails in `AGENTS.md`;
605
- Phase 5 retitled "Project notes file".
606
- - Bootstrap Phase 1 no longer early-stops at possibly-stale docs: sources
607
- are reordered CI-first and doc-claimed commands are cross-checked
608
- against CI; a docs/CI mismatch is itself a finding.
609
- - README: removed the free-floating version line (CHANGELOG is the version
610
- source of truth), corrected the `SessionStart` design note (Claude Code
611
- re-injects `CLAUDE.md` after compaction; the real anti-drift lever is
612
- doctrine size), and recorded the decision against plugin packaging.
613
-
614
- ### Fixed
615
-
616
- - The slugify grader could be gamed for a perfect score: keeping the four
617
- original test *names* with `pass` bodies plus one real regression test
618
- passed every criterion. A new contract-mutant criterion (the worked suite
619
- must go red against an implementation that breaks the original
620
- leading/trailing-separator contract while handling unicode correctly)
621
- closes it; the exact cheat is checked in as `gamed/` and CI proves the
622
- grader rejects it.
623
- - The verify-tracking hook counted *reading* the test file (`cat test.sh`,
624
- `grep pytest README`) as a verify run, flipping the statusline green and
625
- disarming the stop nudge with no test run — fixed opt-in via
626
- `LUCIAZERO_VERIFY_CMD` exact matching (the broad regex remains the default).
627
- - The example check-suppression guard blocked edits *near* a pre-existing
628
- suppression marker (it regexed the whole tool input, so an untouched
629
- `noqa` in `old_string` triggered it). Now diff-aware: it compares marker
630
- counts between new and old text (`old_string` for Edit, the on-disk file
631
- for Write) and blocks only when an edit *adds* a marker.
632
- - `uninstall.sh` no longer aborts halfway (unguarded `grep` no-match under
633
- `set -e`) when `CLAUDE.md` contains only the import line — exactly the
634
- state `install.sh` creates for a user with no prior `CLAUDE.md`. It now
635
- also removes an empty `CLAUDE.md` instead of leaving a zero-line file.
636
- Regression-tested by a fresh-user install→uninstall cycle in `test.sh`.
637
- - `install-codex.sh` reinstalls no longer grow `AGENTS.md` by one blank
638
- line per run; a reinstall is now byte-identical (asserted in `test.sh`).
639
- - Backup names (`*.bak.<timestamp>`) are collision-proof in all four
640
- install/uninstall scripts — two runs in the same second previously
641
- overwrote the earlier backup, destroying the pristine pre-install copy.
642
-
643
- ## [1.3.0] - 2026-08-07
644
-
645
- ### Added
646
-
647
- - OpenAI Codex CLI support: `install-codex.sh` / `uninstall-codex.sh`.
648
- Doctrine lands as a marker-delimited block in `~/.codex/AGENTS.md`;
649
- `luciazero-bootstrap` and `retro` copy as-is (Codex reads the same
650
- `SKILL.md` format); the `reviewer` agent ships as a Codex skill with the
651
- Claude-only `tools:` frontmatter line dropped. Honors `CODEX_HOME`,
652
- backs up `AGENTS.md`, idempotent, converts from the `claude/` sources at
653
- install time so nothing is duplicated in the repo.
654
- - `test.sh` covers the Codex cycle in its own sandbox (marker-block
655
- idempotency, skill installs, `tools:` line stripped, uninstall restores
656
- pre-existing `AGENTS.md` content).
657
-
658
- ## [1.2.0] - 2026-08-07
659
-
660
- ### Added
661
-
662
- - Doctrine rule 6 — debugging starts with a hypothesis and the command
663
- that would confirm or refute it, before any edit.
664
- - Doctrine rule 9 — review the final diff as a skeptic before declaring
665
- done; risky diffs get an independent reviewer agent. (15 rules total.)
666
- - `/retro` skill — harvest session lessons (null results, footguns,
667
- environment quirks) into the project's `CLAUDE.md`/`docs/`, deduping
668
- against and correcting existing notes.
669
- - `agents/reviewer.md` — read-only adversarial reviewer subagent:
670
- severity-tagged findings, verified against source, `No findings.` over
671
- invented ones.
672
- - Bootstrap Phase 2 — two-tier verify guidance: fast `verify` every loop
673
- iteration, `verify-full` before declaring done; single tier for small
674
- repos.
675
-
676
- ### Changed
677
-
678
- - `install.sh`/`uninstall.sh`/`test.sh` cover the new skill and agent;
679
- installer backs up a pre-existing customized `agents/reviewer.md`
680
- before overwriting.
681
-
682
- ## [1.1.0] - 2026-08-07
683
-
684
- ### Added
685
-
686
- - Doctrine rule 12 — **stop and ask before high-stakes moves, and ask
687
- clearly**: deleting data, deploying or touching production,
688
- force-pushing, changing a public API/contract, spending real money, or
689
- leaving the agreed scope require a decidable question (what, why,
690
- options, recommendation) before proceeding. "Finish the whole scope"
691
- renumbered to 13.
692
-
693
- ## [1.0.0] - 2026-08-07
694
-
695
- ### Added
696
-
697
- - 12-rule luciazero doctrine (`claude/luciazero.md`),
698
- loaded in every session via an `@luciazero.md` import in the
699
- global `~/.claude/CLAUDE.md`.
700
- - `/luciazero-bootstrap` skill — six-phase, language-agnostic procedure that
701
- makes a repository agent-ready (detect commands, establish a verify
702
- command, smoke tests, guardrails, project `CLAUDE.md`, prove verify can
703
- go red).
704
- - `install.sh` / `uninstall.sh` — idempotent, back up `CLAUDE.md` before
705
- editing, never write outside the Claude config dir, honor
706
- `CLAUDE_CONFIG_DIR`.
707
- - Inert per-repo settings example
708
- (`examples/project-settings.example.json`): permission allowlist,
709
- secret-read denies, disabled hook templates.
710
- - `test.sh` verify command; GitHub Actions CI on every push and
711
- zip-building release workflow on version tags.
712
- - MIT license.