ai-push-hooks 0.3.2 → 0.3.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,26 @@ All notable changes to this project are documented here. The format is based on
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [0.3.3] - 2026-09-14
8
+
9
+ This is a beta patch release. Python and npm both use `0.3.3`, the canonical
10
+ Git tag is `v0.3.3`, npm uses the `beta` dist-tag, and GitHub marks the release
11
+ as a prerelease. PyPI does not provide a separate beta channel, so Python users
12
+ must select the exact `0.3.3` version.
13
+
14
+ This is a documentation and validation patch only; it adds no runtime features.
15
+
16
+ ### Documentation
17
+
18
+ - Added a tested deterministic postcondition recipe that distinguishes successful
19
+ runner execution from a verified fix, while preserving legitimate no-op applies.
20
+ - Recorded live runner validation results for OpenCode, Claude Code, and Codex.
21
+ - Clarified the focused pre-push use case and how the tool complements hook
22
+ managers, background reviewers, and broader agent workflow frameworks.
23
+ - Documented custom-agent CLI support through the custom command runner.
24
+ - Updated the Codex conformance check to 0.152.0 and verified live project reads
25
+ and allowlisted edits with `gpt-5.6-luna`.
26
+
7
27
  ## [0.3.2] - 2026-09-12
8
28
 
9
29
  This is a beta patch release. Python and npm both use `0.3.2`, the canonical
@@ -44,8 +64,7 @@ must select the exact `0.3.2` version.
44
64
 
45
65
  ### Documentation
46
66
 
47
- - Added an agent setup skill, CI/npm beta badges, and a reproducible,
48
- deterministic review/fail/fix/commit/pass demo, explicitly not live AI footage.
67
+ - Added an agent setup skill and CI/npm beta badges.
49
68
  - Explained the manual-commit assertion and why applied edits are not part of
50
69
  the commit already being pushed.
51
70
  - Included the configuration guide in the npm package and removed its link to
@@ -196,7 +215,8 @@ the exact `0.2.0` version rather than a PyPI beta channel.
196
215
  - Preserved Git porcelain paths and protected pre-existing dirty allowlisted files during apply steps.
197
216
  - Clarified module-local artifact references and standardized the repository hook integration.
198
217
 
199
- [Unreleased]: https://github.com/shanebishop1/ai-push-hooks/compare/v0.3.2...HEAD
218
+ [Unreleased]: https://github.com/shanebishop1/ai-push-hooks/compare/v0.3.3...HEAD
219
+ [0.3.3]: https://github.com/shanebishop1/ai-push-hooks/compare/v0.3.2...v0.3.3
200
220
  [0.3.2]: https://github.com/shanebishop1/ai-push-hooks/compare/v0.3.1...v0.3.2
201
221
  [0.3.1]: https://github.com/shanebishop1/ai-push-hooks/compare/v0.3.0...v0.3.1
202
222
  [0.3.0]: https://github.com/shanebishop1/ai-push-hooks/compare/v0.2.1...v0.3.0
package/README.md CHANGED
@@ -6,7 +6,7 @@
6
6
 
7
7
  `AGENTS.md` tells an agent how to work. **ai-push-hooks checks whether it followed through.** Think of it as the inverse of `AGENTS.md`: a second pass over outgoing changes before `git push`, catching guidelines the agent forgot or neglected.
8
8
 
9
- Use **OpenCode, Codex, or Claude Code** to check rules that need judgment, not just a regex. Report violations, apply scoped fixes, and block pushes with explicit checks. It's a hedge against missed instructions, not a guarantee that AI catches everything.
9
+ Use **OpenCode, Codex, Claude Code, or your own agentic CLI** to check rules that need judgment, not just a regex. Report violations, apply scoped fixes, and block pushes with explicit checks. It's a hedge against missed instructions, not a guarantee that AI catches everything.
10
10
 
11
11
  ## Rules Worth Checking
12
12
 
@@ -19,9 +19,11 @@ Use **OpenCode, Codex, or Claude Code** to check rules that need judgment, not j
19
19
 
20
20
  Write your own rules in prompts or supply a rules file as context. These are examples of checks you configure, not built-in guarantees.
21
21
 
22
- ## Why Now?
22
+ ## Why This Tool?
23
23
 
24
- With lower-cost models such as GPT 5.6 Luna, GLM 5.3 Flash, Muse Spark 1.3, and Gemini Flash, running several focused agentic checks on each push can be practical, rather than reserving AI review for special occasions. Keep context narrow and measure your workflow's cost and latency. GPT 5.6 Luna is the default.
24
+ Keep your coding workflow; add a check before Git sends the commits. ai-push-hooks combines repository-specific AI checks, scripts, and explicit pass/fail gates around the outgoing diff, using your existing coding CLIs. No daemon or hosted review service is required.
25
+
26
+ Use it alongside [Lefthook](https://github.com/evilmartians/lefthook) for hook management and [roborev](https://github.com/kenn-io/roborev) for background reviews. [Archon](https://github.com/coleam00/Archon) and [TAKT](https://github.com/nrslib/takt) orchestrate development tasks; this tool deliberately stays narrower. Its job is to check your rules before a push, not take over how you build the project.
25
27
 
26
28
  ## Quick Start
27
29
 
@@ -37,12 +39,6 @@ Install and authenticate your chosen AI CLI. The current `init` starter checks d
37
39
 
38
40
  [Other installation options](docs/configuration.md#installation) | [Existing hook managers](docs/configuration.md#hook-managers)
39
41
 
40
- ## Deterministic Demo
41
-
42
- ![Review, fail, fix, commit, pass: deterministic demo](https://raw.githubusercontent.com/shanebishop1/ai-push-hooks/main/docs/demo/review-fix.gif)
43
-
44
- This generated GIF summarizes actual hook results from a disposable repository with a local deterministic command runner; it is not a live screen recording or live AI. It makes no provider calls, remote pushes, or Beads changes. Regenerate it with `uv run --no-project --with pillow==11.3.0 python docs/demo/generate.py`.
45
-
46
42
  ## Example: Check Your Rules
47
43
 
48
44
  This workflow supplies the outgoing diff and your root `AGENTS.md` explicitly, asks for violations, and blocks if any are reported. It does not rely on the runner automatically loading agent instructions.
@@ -103,7 +99,7 @@ inputs = ["review/issues.json"]
103
99
 
104
100
  `docs_issue_array` is the existing schema name for `{file, description}` findings; it works for code rules too. A missing rules file or a failed check blocks the push by default. Findings live in the run artifacts under `.git/ai-push-hooks/`.
105
101
 
106
- **Want fixes too?** Add an `apply` step with the findings, your rules, and an explicit `allow_paths` list, then recheck and run tests. Applied edits are not auto-committed: review the diff, commit approved changes, and retry the push.
102
+ **Want fixes too?** Add an `apply` step with the findings, your rules, and an explicit `allow_paths` list, then recheck and run tests. [Runner completion is not proof that a fix landed](docs/configuration.md#deterministic-postconditions-after-apply): verify the result. Applied edits are not auto-committed: review the diff, commit approved changes, and retry the push.
107
103
 
108
104
  ## Build Your Workflow
109
105
 
@@ -121,7 +117,7 @@ Each module combines the steps it needs:
121
117
 
122
118
  ## Choose Your AI
123
119
 
124
- OpenCode defaults to **`openai/gpt-5.6-luna`**. Explicit runner profiles use their own model settings. To use Codex or Claude Code, add its profile and change `[llm].runner`:
120
+ OpenCode, Codex, and Claude Code have first-class adapters. OpenCode defaults to **`openai/gpt-5.6-luna`**. Explicit runner profiles use their own model settings. To use Codex or Claude Code, add its profile and change `[llm].runner`:
125
121
 
126
122
  ```toml
127
123
  [runners.codex]
@@ -131,13 +127,14 @@ type = "codex"
131
127
  type = "claude"
132
128
  ```
133
129
 
134
- Each `ask` or `apply` step can override the runner, so one tool can review and another can fix. Use model identifiers available to your provider. Other tools, including Pi, can use a [custom command runner](docs/configuration.md#custom-runners).
130
+ Each `ask` or `apply` step can override the runner, so one tool can review and another can fix. **Any other agentic CLI, including Pi, can integrate through a [custom command runner](docs/configuration.md#custom-runners)**: supply a noninteractive command or a thin wrapper. Use model identifiers available to your provider.
135
131
 
136
132
  ## Control And Safety
137
133
 
138
134
  - Errors block pushes by default. AI judgments can still miss violations or report false positives.
139
135
  - `apply` validates file and Git state and limits propagated edits to `allow_paths`. It does not auto-commit.
140
136
  - Runners, scripts, and callbacks are trusted local programs, not an OS sandbox.
137
+ - Local hooks can be bypassed; retain CI for required enforcement.
141
138
  - Repository content may be sent to your model provider. Review its privacy and billing terms.
142
139
 
143
140
  Logs and run summaries live under `.git/ai-push-hooks/`. OpenCode transcripts are captured there by default. To intentionally skip one push: `AI_PUSH_HOOKS_SKIP=1 git push`.
package/SECURITY.md CHANGED
@@ -117,7 +117,12 @@ The [0.3.0 release record](CHANGELOG.md#030---2026-09-09) documented a pinned
117
117
  Lefthook suite reporting **407 tests with no skips**, an OpenCode **1.18.29**
118
118
  contract smoke test with an in-process loopback mock provider and no external
119
119
  model call, and version/help-only checks for Codex **0.148.0** and Claude
120
- **2.1.220**. This is historical evidence, not a current suite result or proof
121
- for every provider, model, authentication mode, platform, or live `apply` path.
120
+ **2.1.220**. The bounded validation snapshot also records Codex **0.152.0** on
121
+ 2026-09-14: its synthetic fixture verified a nonce read and an allowlisted marker
122
+ apply followed by a deterministic postcondition. See the [runner validation
123
+ snapshot](docs/configuration.md#live-validation-snapshot). Together, these records
124
+ are historical or bounded evidence, not a current suite result or proof for every
125
+ provider, model, authentication mode, platform, live `apply` path, or a general
126
+ security guarantee.
122
127
  Treat generated-hook path checks and the Lefthook runner as integration
123
128
  safeguards, not isolation boundaries.
@@ -64,6 +64,8 @@ Upgrading from 0.2.1? Rename `type = "llm"` steps to `type = "ask"`.
64
64
 
65
65
  `[llm].runner` selects the default; a step's `runner` overrides it. Referenced profiles must exist under `[runners.<name>]`, except for the implicit OpenCode default.
66
66
 
67
+ OpenCode, Codex, and Claude are first-class adapters. The generic `command` type is for other agentic CLIs and follows the custom runner contract below.
68
+
67
69
  | Profile field | Values / behavior |
68
70
  | --- | --- |
69
71
  | `type` | Required: `opencode`, `codex`, `claude`, or `command`. |
@@ -81,6 +83,27 @@ OpenCode runs with isolated configuration and permissions; project/global config
81
83
 
82
84
  Apply requires a single pushed branch whose local commit is the checked-out `HEAD`. Staging excludes Git metadata, `AGENTS.md`, ignored files, symlinks, and special files. These controls are not an OS sandbox or an automatic rollback system. See [Security](../SECURITY.md).
83
85
 
86
+ ### Live validation snapshot
87
+
88
+ Live read/review and apply checks on Linux:
89
+
90
+ | Date | Runner (CLI; model) | Review | Apply |
91
+ | --- | --- | --- | --- |
92
+ | 2026-09-12 | OpenCode (1.18.29; `openai/gpt-5.6-luna`) | Passed | Passed |
93
+ | 2026-09-12 | Claude Code (2.1.220; `sonnet`) | Passed | Passed |
94
+ | 2026-09-14 | Codex (0.152.0; `gpt-5.6-luna`) | Passed | Passed |
95
+
96
+ The Codex check read a random value from a disposable repository, then verified
97
+ that only the allowlisted README changed. These checks establish the tested
98
+ operations, not compatibility with every model or platform.
99
+
100
+ #### Local troubleshooting: Codex sandbox
101
+
102
+ For Codex 0.152, the standalone no-AI sandbox syntax is
103
+ `codex sandbox -- /usr/bin/true`. Resolve host restrictions rather than disabling
104
+ protections, and use a deterministic postcondition to verify the desired checkout
105
+ outcome after an apply.
106
+
84
107
  ### Apply and manual commits
85
108
 
86
109
  `apply` is generic: it can edit any eligible checkout file matching `allow_paths`; it is not limited to Markdown. The runner edits a temporary staging copy, and only validated changes propagate back to the checkout. Those edits do not enter the commit already being pushed, and `apply` never creates a Git commit.
@@ -97,9 +120,62 @@ inputs = ["apply/result.json"]
97
120
 
98
121
  The assertion checks `apply/result.json`'s `changed_files` and intentionally blocks when edits were propagated. Review `git diff`, run the relevant checks, commit the approved changes, and retry the push. On the retry, the assertion passes when the apply step reports no changes.
99
122
 
123
+ ### Deterministic postconditions after apply
124
+
125
+ An apply process succeeding, or reporting `changed_files = []`, is not proof that
126
+ the requested result is present. The latter can simply mean that the apply was a
127
+ legitimate no-op because the checkout was already correct. Add a deterministic
128
+ postcondition after `apply` and before the manual-commit gate when the desired
129
+ file content has a precise representation:
130
+
131
+ ```toml
132
+ [general]
133
+ require_clean_worktree = true
134
+
135
+ [workflow]
136
+ modules = ["docs"]
137
+
138
+ [modules.docs]
139
+ enabled = true
140
+
141
+ [[modules.docs.steps]]
142
+ id = "apply"
143
+ type = "apply"
144
+ prompt = "In the existing README.md, replace 'Release note: DRAFT.' with 'Release note: READY.' and make no other changes."
145
+ allow_paths = ["README.md"]
146
+
147
+ [[modules.docs.steps]]
148
+ id = "postcondition"
149
+ type = "assert"
150
+ command = [
151
+ "{python}",
152
+ "-c",
153
+ "import pathlib, sys; sys.exit(0 if pathlib.Path('README.md').read_text(encoding='utf-8') == 'Release note: READY.\\n' else 1)",
154
+ ]
155
+ inputs = ["apply/result.json"]
156
+
157
+ [[modules.docs.steps]]
158
+ id = "manual-commit"
159
+ type = "assert"
160
+ assertion = "docs_apply_requires_manual_commit"
161
+ inputs = ["apply/result.json"]
162
+ ```
163
+
164
+ This is an intentional synthetic example, not a recommendation to overwrite a
165
+ real README: its fixture starts with exactly `Release note: DRAFT.\n`, and its
166
+ desired full content is exactly `Release note: READY.\n`.
167
+ The command is a direct argv vector rather than a Python `assert`; Python
168
+ optimization must not be able to remove the check. The postcondition reads the
169
+ checkout after apply, not the commit being pushed. Keep the clean-worktree
170
+ workflow setting for the hook's starting state and the manual-commit gate for
171
+ the intentionally dirty post-apply checkout: review the resulting diff, run
172
+ relevant checks, commit it, and retry the push. This gate and postcondition
173
+ still do not prove human review, semantic correctness, or that an agent
174
+ complied with every instruction.
175
+
100
176
  ## Custom Runners
101
177
 
102
- Use a command profile for another CLI, including Pi, or your own wrapper. This example expects a local `scripts/review-agent` program that reads the prompt from stdin and writes its final response to stdout:
178
+ Use a `command` profile to invoke any other agentic CLI, including Pi, directly or through a thin wrapper. A wrapper can normalize JSONL events into a final response. The runner must be noninteractive: read the prompt from stdin or argv, write only the final response to stdout, and use exit codes to report success or failure.
103
179
 
104
180
  ```toml
105
181
  [runners.custom]
@@ -111,7 +187,7 @@ project_access = "project"
111
187
 
112
188
  Select it with `runner = "custom"` on an `ask` or `apply` step. The runner receives the full instruction and artifact packet. For apply, its working directory is the staging copy.
113
189
 
114
- Commands are argv arrays, not shell strings. Whole-argument placeholders are `{model}`, `{cwd}`, `{stage}`, and `{prompt}`. The `argv` transport requires exactly one `{prompt}` argument; stdin avoids exposing prompts in process listings. Custom programs inherit the user environment and own their permissions and session lifecycle.
190
+ Commands are argv arrays, not shell strings. Whole-argument placeholders are `{model}`, `{cwd}`, `{stage}`, and `{prompt}`. The `argv` transport requires exactly one `{prompt}` argument; stdin avoids exposing prompts in process listings. Custom programs inherit the user's environment, including authentication variables, and manage their own permissions and session setup/cleanup. For apply, the working directory is the staging copy.
115
191
 
116
192
  ## Commands And Callbacks
117
193
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "ai-push-hooks",
3
- "version": "0.3.2",
3
+ "version": "0.3.3",
4
4
  "description": "Run structured AI-assisted checks and allowlisted maintenance before git push",
5
5
  "license": "MIT",
6
6
  "author": {
package/pyproject.toml CHANGED
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
4
4
 
5
5
  [project]
6
6
  name = "ai-push-hooks"
7
- version = "0.3.2"
7
+ version = "0.3.3"
8
8
  description = "Run structured AI-assisted checks and allowlisted maintenance before git push"
9
9
  readme = "README.md"
10
10
  requires-python = ">=3.10"