arcaeon-once 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,9 @@
1
+ __pycache__/
2
+ *.pyc
3
+ .venv/
4
+ dist/
5
+ build/
6
+ *.egg-info/
7
+ *.jsonl
8
+ *.sqlite3
9
+ !server.json
@@ -0,0 +1,128 @@
1
+ # Changelog
2
+
3
+ ## 0.1.0 — the WAL-init race, found and fixed BEFORE first publish (2026-08-14)
4
+
5
+ The hold below did its job. The race is fixed, the fix is signed off, the gate
6
+ was run and passed **20/20 consecutive full-suite runs with zero flakes**
7
+ (`test_once.py` + `test_concurrency.py` + `selftest`, 283s), `dist/` was
8
+ rebuilt from the fixed source, and `ArcaeonOncePyPIRetry` is re-enabled for
9
+ 14:10 on 8/15. Nothing was ever published from the raced code — PyPI returned
10
+ 404 for the project at fix time — so this folds into 0.1.0 rather than
11
+ minting a 0.1.1 whose predecessor never existed.
12
+
13
+ **What changed:**
14
+
15
+ - **First-touch index setup is serialized by a cross-process file lock**
16
+ (`msvcrt` byte-range on Windows, `fcntl.flock` on POSIX). The one-time
17
+ `delete`→`wal` journal-mode switch now happens ONCE, by one process, with
18
+ everybody else queued *outside the database* — instead of N processes
19
+ simultaneously fighting for an exclusive lock that ignores `busy_timeout`.
20
+ - **The hot path takes no exclusive lock at all.** `_init_index()` now starts
21
+ with a lock-free, read-only check (file exists + table present + journal
22
+ mode settled) and returns immediately once the index is up. Only genuine
23
+ first touch reaches the locked setup path, double-checked on the way in.
24
+ - **WAL is treated as what it actually is — an optimisation, not the safety
25
+ mechanism.** The real race serialization is `BEGIN IMMEDIATE`, which *does*
26
+ honour `busy_timeout`. So the WAL switch is bounded-retry and best-effort:
27
+ an index stuck in `delete` mode is slower under load, never less correct,
28
+ and must not take down a guarded side effect.
29
+ - **New typed outcome: `IndexUnavailable`.** Every SQLite path in the index
30
+ (`_init_index`, `_claim`, `_mark_executed`, `_reclaim_after_indeterminate`,
31
+ `rebuild_index`) now converts residual `sqlite3.Error` into this documented
32
+ exception after a bounded retry. Nothing untyped escapes `guard()` any
33
+ more. It fails safe — raised before any claim row and any ledger row.
34
+ The single documented nuance rides on the instance as `.ledger_committed`:
35
+ raised from `done()`/`complete()`, the executed row is already durable in
36
+ the ledger and only the accelerator is stale (duplicate refusal still
37
+ works, `rebuild_index()` resyncs). Don't re-run the effect on that one.
38
+ - **`rebuild_index()` now removes orphaned `-wal`/`-shm` sidecars** before
39
+ recreating the index, so a stale WAL can't outlive the database it belonged
40
+ to.
41
+
42
+ **New regression test** —
43
+ `test_first_touch_index_race_is_typed_and_exactly_once`, 5 trials × 12
44
+ processes. Workers import the library, signal ready, and wait on a `go` file;
45
+ the parent writes it — carrying an absolute release timestamp — only once
46
+ every worker has armed, and workers then BUSY-SPIN to that instant, so all
47
+ twelve land on the journal-mode switch inside the same millisecond.
48
+
49
+ The release time is chosen AFTER everyone arms on purpose. The first version
50
+ of this test guessed a fixed 3-second lead up front, and on a loaded machine
51
+ that guess was wrong: a gate run caught it arming 9/10 and then 0/10 workers.
52
+ That test was itself flaky, which is precisely the disease under treatment —
53
+ so the barrier was rebuilt rather than papered over with a longer sleep.
54
+
55
+ Reproduction measured against the pre-fix code (the check that the test isn't
56
+ vacuous — a regression test that can't fail on the buggy code proves nothing):
57
+ **19 of 30 workers crashed** with exactly the traceback below under the first
58
+ barrier, and 2–8 of 30–80 workers under the final one, with run-level
59
+ reproduction — at least one trial red — in every attempt. Five trials, not
60
+ one, because the crash window is the sub-millisecond in which the first
61
+ process creates the database and flips its journal mode, and a single trial
62
+ can miss it by luck.
63
+
64
+ On the fixed code: 0 crashes, every outcome typed, exactly one execution per
65
+ trial, across 20 consecutive full-suite runs including three where the
66
+ machine was loaded enough to double the runtime.
67
+
68
+ ## Superseded — the hold that produced the fix (2026-08-14 QA sweep)
69
+
70
+ _Kept as the audit trail. Every claim below was true when written; the
71
+ "Candidate fix" is what got implemented, plus the file lock and the typed
72
+ outcome._
73
+
74
+ **The 0.1.0 PyPI publish was held back, and the `ArcaeonOncePyPIRetry` scheduled
75
+ task that would have shipped it unattended at 14:10 on 8/15 is DISABLED.**
76
+ Re-arm with `schtasks /change /tn "ArcaeonOncePyPIRetry" /enable` once the item
77
+ below is signed off. The task publishes unconditionally — it has no test gate —
78
+ and 14:10 falls inside a work shift, so nobody would have been watching.
79
+
80
+ **Why:** `test_concurrency.py` fails roughly 30% of runs, always the same way:
81
+
82
+ ```
83
+ File "arcaeon_once\__init__.py", line 438, in _enter_for_key
84
+ _init_index(self.state_db)
85
+ File "arcaeon_once\__init__.py", line 119, in _init_index
86
+ con.execute("PRAGMA journal_mode=WAL")
87
+ sqlite3.OperationalError: database is locked
88
+ ```
89
+
90
+ `_init_index()` runs on every `guard()` entry and unconditionally re-issues
91
+ `PRAGMA journal_mode=WAL`. The first-time `delete`→`wal` switch needs an
92
+ EXCLUSIVE lock and does **not** honour `busy_timeout` — measured: the pragma
93
+ gave up after 0.003s against a 15s timeout. Two processes first-touching the
94
+ same ledger at once collide there, and one crashes out of
95
+ `guard().__enter__()` with an untyped `OperationalError` instead of one of this
96
+ library's typed outcomes.
97
+
98
+ **What is NOT broken:** the exactly-once guarantee holds. The crash lands before
99
+ any intent or claim row is written, so it fails safe — no double-fire. 120
100
+ concurrent claims and 40 barrier-synchronised races each produced exactly one
101
+ winner. This is an availability defect, not a correctness one.
102
+
103
+ **Why it still blocks the release:** `README.md` sells this exact test as the
104
+ proof — "Verified with two real OS processes hammering the same key" — and that
105
+ proof is red a third of the time. Shipping a reliability library whose own
106
+ advertised evidence is flaky is the wrong first impression, and PyPI version
107
+ numbers cannot be reused.
108
+
109
+ **Candidate fix** (shape verified in a scratch copy, deliberately NOT applied):
110
+ read `PRAGMA journal_mode` first and only switch when it is not already `wal`,
111
+ tolerating `OperationalError` on the switch — WAL is a concurrency optimisation,
112
+ and the real serialization comes from `BEGIN IMMEDIATE`. It touches the
113
+ concurrency core of an exactly-once library, so it wants a sign-off plus 20+
114
+ consecutive green suite runs before the publish is re-armed.
115
+
116
+ ## 0.1.0 — initial release contents (2026-08-14)
117
+
118
+ - `guard(key, *, ledger_path=None, on_duplicate="raise", allow_retry_after_indeterminate=False, verify_integrity=False, store_outcome=False)` — context manager / decorator around a non-idempotent side effect, backed by an `arcaeon-ledger` hash chain.
119
+ - Two-phase `once.intent` / `once.executed` ledger rows: a crash between them leaves the key `Indeterminate` (typed, refuse-by-default) instead of silently assumed either way.
120
+ - `AlreadyExecuted`, `Indeterminate`, `TamperDetected` — typed, structured refusals, never a silent pass.
121
+ - `receipt(key, ledger_path=...)` — the tamper-evident state of a key, read straight from the hash chain.
122
+ - `complete(key, outcome, ledger_path=...)` — crash-recovery path: mark a key executed after manually confirming the effect actually ran.
123
+ - `rebuild_index(ledger_path)` — replay the ledger into a fresh SQLite concurrency index if the index file is lost.
124
+ - SQLite `BEGIN IMMEDIATE` race serialization on brand-new keys (same pattern as `arcaeon-meter`'s usage counter); every other decision is re-derived fresh from the ledger, so an out-of-band ledger edit is reflected immediately.
125
+ - MCP server (`arcaeon_once.mcp_server`) — one tool, `guard_side_effect`, with `claim`/`complete` actions mirroring the library's own two-phase design across the MCP boundary.
126
+ - CLI (`arcaeon_once.cli`) — `receipt` and `rebuild-index` commands.
127
+ - Tests: duplicate refusal, crash-window indeterminacy (real subprocess hard-exit via `os._exit`), chain tamper detection (in-place edit and tail-drop), and a real two-OS-process concurrency race (`test_concurrency.py`).
128
+ - `python -m arcaeon_once.selftest` — golden outcome-digest vector + planted-tamper fixture, runnable on any machine.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Arcaeon
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,234 @@
1
+ Metadata-Version: 2.5
2
+ Name: arcaeon-once
3
+ Version: 0.1.0
4
+ Summary: Executed-once receipts for non-idempotent agent side effects. Refuses silent double-fires, flags crash-window uncertainty honestly.
5
+ Project-URL: Homepage, https://arcaeon.io
6
+ Author: Arcaeon
7
+ License: MIT
8
+ License-File: LICENSE
9
+ Keywords: agents,ai,audit,hash-chain,idempotency,mcp,reliability,tamper-evident
10
+ Classifier: Intended Audience :: Developers
11
+ Classifier: License :: OSI Approved :: MIT License
12
+ Classifier: Programming Language :: Python :: 3
13
+ Classifier: Topic :: Software Development :: Libraries
14
+ Requires-Python: >=3.9
15
+ Requires-Dist: arcaeon-ledger>=0.5.2
16
+ Description-Content-Type: text/markdown
17
+
18
+ # arcaeon-once
19
+
20
+ **Kybernis-shaped enforcement stops the double-fire. `arcaeon-once` is the
21
+ evidence layer: a tamper-evident receipt proving a side effect ran exactly
22
+ once — or an honest flag when it didn't know.**
23
+
24
+ Agents retry. Refunds, deploys, and outbound emails do not want to be
25
+ retried. `arcaeon-once` wraps a non-idempotent side effect with an
26
+ idempotency key: it refuses to re-run a key that already executed and hands
27
+ back the original receipt instead, hash-chained via
28
+ [`arcaeon-ledger`](https://pypi.org/project/arcaeon-ledger/) so nobody can
29
+ quietly delete the record to enable a re-fire.
30
+
31
+ ```
32
+ pip install arcaeon-once # then: from arcaeon_once import guard
33
+ ```
34
+
35
+ ```python
36
+ from arcaeon_once import guard
37
+
38
+ with guard(f"refund:{charge_id}", ledger_path="ops.log.jsonl") as g:
39
+ result = stripe.Refund.create(charge=charge_id)
40
+ g.done(result)
41
+ ```
42
+
43
+ Call it again with the same key and it raises `AlreadyExecuted` — carrying
44
+ the original receipt — instead of refunding twice.
45
+
46
+ ## The non-proof, stated before any feature, because it is the point
47
+
48
+ **This library gives you at-most-once-or-flagged. Not exactly-once.**
49
+
50
+ Exactly-once over a real, non-transactional side effect — an HTTP call to a
51
+ payment processor, a `kubectl apply`, an SMTP send — is not achievable by any
52
+ wrapper running in the same process as the effect. If the process dies
53
+ between the effect executing and the record being written, nobody, this
54
+ library included, can know from the outside whether the effect happened.
55
+ Anyone telling you their idempotency library gives you exactly-once across
56
+ that boundary is telling you a story, not an engineering fact — say so
57
+ plainly, because the crowded "AI agent reliability" space is not short on
58
+ confident claims that don't survive a `kill -9` at the wrong instant.
59
+
60
+ What you actually get:
61
+
62
+ - **At-most-once**, when nothing crashes. A second call with the same key,
63
+ while the ledger is intact, is refused. Period.
64
+ - **Or-flagged**, when something crashes mid-effect. The key comes back
65
+ `Indeterminate` — a typed, refuse-by-default outcome — instead of a silent
66
+ double-fire or a silent skip. You check the real system (did the refund
67
+ post? did the deploy land?) and either `complete()` it (it did happen) or
68
+ retry with `allow_retry_after_indeterminate=True` (it didn't).
69
+
70
+ ## The crash window, designed, not hidden
71
+
72
+ Every guarded call is two-phase in the ledger: an `once.intent` row is
73
+ appended **before** the effect runs, an `once.executed` row **after** (via
74
+ `g.done(outcome)` or the module-level `complete()`). A key with an `intent`
75
+ row and no matching `executed` row means: something started and this library
76
+ does not know if it finished.
77
+
78
+ ```python
79
+ from arcaeon_once import guard, receipt, complete, Indeterminate, AlreadyExecuted
80
+
81
+ try:
82
+ with guard("deploy:build-4471", ledger_path="ops.log.jsonl") as g:
83
+ run_deploy() # process dies here -> intent, no executed
84
+ g.done({"status": "ok"})
85
+ except AlreadyExecuted as e:
86
+ print("already ran:", e.receipt.executed_ts, e.receipt.executed_chain)
87
+
88
+ # next run, same key:
89
+ r = receipt("deploy:build-4471", ledger_path="ops.log.jsonl")
90
+ r.state # "intent" -- the crash window, exactly as it happened, not glossed
91
+ try:
92
+ with guard("deploy:build-4471", ledger_path="ops.log.jsonl"):
93
+ ...
94
+ except Indeterminate as e:
95
+ # go check the actual deploy target by hand, THEN:
96
+ complete("deploy:build-4471", {"status": "ok"}, ledger_path="ops.log.jsonl")
97
+ # -- or, if it truly didn't land --
98
+ # guard("deploy:build-4471", ledger_path="ops.log.jsonl",
99
+ # allow_retry_after_indeterminate=True)
100
+ ```
101
+
102
+ That's honest exactly-once-**or-tell-you** semantics. `Indeterminate` is
103
+ refused by default — never silently treated as "safe to retry," never
104
+ silently treated as "must have worked." Resolving it is a manual step on
105
+ purpose: only you (or your ops tooling) can look at the real system and know
106
+ which way it actually went.
107
+
108
+ ## What proves the "once" — the hash chain, not a promise
109
+
110
+ Every `intent`/`executed` row is appended to an `arcaeon-ledger` hash chain:
111
+ `chain = sha256(prev_chain + canonical_json(row_without_chain))[:32]`. Delete
112
+ or edit an inconvenient `executed` row to re-enable a re-fire, and every
113
+ later link in the chain breaks — `receipt()` reports `ledger_ok=False` with a
114
+ `ledger_first_break` naming the row. Deletion doesn't erase the fact that a
115
+ deletion happened.
116
+
117
+ `guard()` does **not** re-verify the whole chain on every call by default —
118
+ that's an `O(rows)` scan of the file, and paying it on every single side
119
+ effect would not scale to a high-volume tool. Pass `verify_integrity=True`
120
+ for that stronger (and slower) guarantee inline, or call `receipt()` /
121
+ `arcaeon_ledger.verify_file()` on your own cadence (a pre-ship gate, a
122
+ nightly job). Tamper caught late is still tamper caught. Tamper never
123
+ checked is a receipt you shouldn't have trusted in the first place — this
124
+ library will not pretend otherwise to look faster in a benchmark.
125
+
126
+ ## Concurrency: exactly one process wins the claim
127
+
128
+ Two processes racing the **same brand-new key** resolve through a single
129
+ SQLite `BEGIN IMMEDIATE` transaction against a small index file next to the
130
+ ledger — the same WAL + immediate-transaction pattern
131
+ [`arcaeon-meter`](https://pypi.org/project/arcaeon-meter/) uses for its usage
132
+ counter. Exactly one caller, ever, gets back the execution claim for a given
133
+ key; the loser gets `AlreadyExecuted` or `Indeterminate` depending on timing,
134
+ never a green light to also run the effect. Verified with two real OS
135
+ processes hammering the same key, and with ten processes released onto a
136
+ brand-new ledger by a wall-clock start barrier (see `test_concurrency.py`),
137
+ not simulated with threads and a comforting mock.
138
+
139
+ The index is created lazily, and that first-touch setup is serialized by a
140
+ cross-process **file lock** — switching a brand-new SQLite file into WAL
141
+ journal mode needs a momentary EXCLUSIVE database lock that does *not* honour
142
+ `busy_timeout`, so without that serialization, N processes first-using the
143
+ same fresh ledger collide on it. If contention still can't be absorbed you get
144
+ `IndexUnavailable`: a typed, documented outcome raised **before** any claim or
145
+ ledger row, never a raw `sqlite3.OperationalError` leaking out of `guard()`.
146
+ (One documented nuance, carried on the exception as `.ledger_committed`: if it
147
+ comes from `done()`/`complete()`, the executed row is already durable in the
148
+ ledger and only the index is stale — duplicate refusal still works, and
149
+ `rebuild_index()` resyncs it. Don't re-run the effect on that one.)
150
+
151
+ That index is consulted **only** to serialize the race at the "nobody has
152
+ claimed this key yet" boundary — every other decision (already executed?
153
+ still an unresolved intent?) is re-derived fresh from the ledger itself on
154
+ every `guard()` call, on purpose: an out-of-band edit to the ledger file
155
+ (a dropped row, a tampered byte) must be reflected immediately even though
156
+ the index file wasn't touched, so a truncation attack degrades to a safe
157
+ refusal (`Indeterminate`) rather than the index quietly vouching for a row
158
+ that's no longer there. The honest cost of that choice: `guard()` scans the
159
+ ledger for the key on every call — `O(rows)` in the ledger's total size, not
160
+ `O(1)`. Fine for a day's or a service's worth of idempotency keys; if you're
161
+ guarding millions of distinct keys against one ledger file, shard the ledger
162
+ (one file per key prefix / tenant / day) rather than expecting this to stay
163
+ `O(1)` — that sharding is on you for now, stated rather than hidden. Lose the
164
+ index file entirely and you lose only the race-serialization fast path, not
165
+ correctness: `rebuild_index()` replays the ledger into a fresh one.
166
+
167
+ ## API surface
168
+
169
+ ```python
170
+ guard(key, *, ledger_path=None, state_db=None, on_duplicate="raise",
171
+ allow_retry_after_indeterminate=False, verify_integrity=False,
172
+ store_outcome=False) -> GuardContext
173
+ ```
174
+ Context manager (`with guard(key) as g: ...; g.done(outcome)`) or decorator
175
+ (`@guard(key)` for a static key, or `@guard(lambda *a, **kw: f"job:{a[0]}")`
176
+ for a key resolved per call). `on_duplicate="raise"` (default) raises
177
+ `AlreadyExecuted`; `"return_receipt"` returns without executing — check
178
+ `g.already_executed` / `g.receipt`, or for the decorator, the call returns
179
+ the `Receipt` directly instead of the wrapped function's result.
180
+
181
+ ```python
182
+ receipt(key, *, ledger_path) -> Receipt
183
+ ```
184
+ The tamper-evident state of a key, read straight from the hash chain (never
185
+ from the SQLite index). `Receipt.state` is `"never"`, `"intent"`, or
186
+ `"executed"`. `bool(receipt)` is `True` only for a clean, verified,
187
+ `"executed"` record — a tampered ledger or an unresolved intent never reads
188
+ as success.
189
+
190
+ ```python
191
+ complete(key, outcome=None, *, ledger_path, store_outcome=False) -> Receipt
192
+ ```
193
+ Mark a key executed directly, without an open `guard()` context — the
194
+ crash-recovery path once you've manually confirmed the effect actually ran.
195
+
196
+ ```python
197
+ rebuild_index(ledger_path, state_db=None) -> int
198
+ ```
199
+ Replay the ledger into a fresh SQLite concurrency index. Restores
200
+ correctness after the index file is lost; not needed for normal operation.
201
+
202
+ ## Drop it into any MCP agent
203
+
204
+ ```json
205
+ {
206
+ "mcpServers": {
207
+ "once": {
208
+ "command": "python",
209
+ "args": ["-m", "arcaeon_once.mcp_server"]
210
+ }
211
+ }
212
+ }
213
+ ```
214
+
215
+ One tool, two actions mirroring the library's own two-phase design so the
216
+ crash window is real even across the MCP boundary: `guard_side_effect(action=
217
+ "claim", key, ...)` before the agent performs the effect (skip it if
218
+ `claimed: false`), `guard_side_effect(action="complete", key, outcome, ...)`
219
+ after it succeeds. If the agent session dies between the two calls, the key
220
+ is left `intent`-only — indeterminate on the next claim, not silently
221
+ resolved by the wrapping.
222
+
223
+ ## Status
224
+
225
+ Core library, CLI, MCP server, all tested: duplicate execution refused with
226
+ the original receipt returned, the crash window (INTENT with no EXECUTED)
227
+ resolving to a typed `Indeterminate` refusal, chain tampering on an executed
228
+ row detected by `receipt()`, a real two-OS-process race resolving to exactly
229
+ one execution, and a ten-process barrier-released first-touch race that lands
230
+ zero untyped exceptions. `python -m arcaeon_once.selftest` ships in the
231
+ package so you can verify the golden digest vector and the planted-tamper
232
+ case on your own machine.
233
+
234
+ MIT.
@@ -0,0 +1,217 @@
1
+ # arcaeon-once
2
+
3
+ **Kybernis-shaped enforcement stops the double-fire. `arcaeon-once` is the
4
+ evidence layer: a tamper-evident receipt proving a side effect ran exactly
5
+ once — or an honest flag when it didn't know.**
6
+
7
+ Agents retry. Refunds, deploys, and outbound emails do not want to be
8
+ retried. `arcaeon-once` wraps a non-idempotent side effect with an
9
+ idempotency key: it refuses to re-run a key that already executed and hands
10
+ back the original receipt instead, hash-chained via
11
+ [`arcaeon-ledger`](https://pypi.org/project/arcaeon-ledger/) so nobody can
12
+ quietly delete the record to enable a re-fire.
13
+
14
+ ```
15
+ pip install arcaeon-once # then: from arcaeon_once import guard
16
+ ```
17
+
18
+ ```python
19
+ from arcaeon_once import guard
20
+
21
+ with guard(f"refund:{charge_id}", ledger_path="ops.log.jsonl") as g:
22
+ result = stripe.Refund.create(charge=charge_id)
23
+ g.done(result)
24
+ ```
25
+
26
+ Call it again with the same key and it raises `AlreadyExecuted` — carrying
27
+ the original receipt — instead of refunding twice.
28
+
29
+ ## The non-proof, stated before any feature, because it is the point
30
+
31
+ **This library gives you at-most-once-or-flagged. Not exactly-once.**
32
+
33
+ Exactly-once over a real, non-transactional side effect — an HTTP call to a
34
+ payment processor, a `kubectl apply`, an SMTP send — is not achievable by any
35
+ wrapper running in the same process as the effect. If the process dies
36
+ between the effect executing and the record being written, nobody, this
37
+ library included, can know from the outside whether the effect happened.
38
+ Anyone telling you their idempotency library gives you exactly-once across
39
+ that boundary is telling you a story, not an engineering fact — say so
40
+ plainly, because the crowded "AI agent reliability" space is not short on
41
+ confident claims that don't survive a `kill -9` at the wrong instant.
42
+
43
+ What you actually get:
44
+
45
+ - **At-most-once**, when nothing crashes. A second call with the same key,
46
+ while the ledger is intact, is refused. Period.
47
+ - **Or-flagged**, when something crashes mid-effect. The key comes back
48
+ `Indeterminate` — a typed, refuse-by-default outcome — instead of a silent
49
+ double-fire or a silent skip. You check the real system (did the refund
50
+ post? did the deploy land?) and either `complete()` it (it did happen) or
51
+ retry with `allow_retry_after_indeterminate=True` (it didn't).
52
+
53
+ ## The crash window, designed, not hidden
54
+
55
+ Every guarded call is two-phase in the ledger: an `once.intent` row is
56
+ appended **before** the effect runs, an `once.executed` row **after** (via
57
+ `g.done(outcome)` or the module-level `complete()`). A key with an `intent`
58
+ row and no matching `executed` row means: something started and this library
59
+ does not know if it finished.
60
+
61
+ ```python
62
+ from arcaeon_once import guard, receipt, complete, Indeterminate, AlreadyExecuted
63
+
64
+ try:
65
+ with guard("deploy:build-4471", ledger_path="ops.log.jsonl") as g:
66
+ run_deploy() # process dies here -> intent, no executed
67
+ g.done({"status": "ok"})
68
+ except AlreadyExecuted as e:
69
+ print("already ran:", e.receipt.executed_ts, e.receipt.executed_chain)
70
+
71
+ # next run, same key:
72
+ r = receipt("deploy:build-4471", ledger_path="ops.log.jsonl")
73
+ r.state # "intent" -- the crash window, exactly as it happened, not glossed
74
+ try:
75
+ with guard("deploy:build-4471", ledger_path="ops.log.jsonl"):
76
+ ...
77
+ except Indeterminate as e:
78
+ # go check the actual deploy target by hand, THEN:
79
+ complete("deploy:build-4471", {"status": "ok"}, ledger_path="ops.log.jsonl")
80
+ # -- or, if it truly didn't land --
81
+ # guard("deploy:build-4471", ledger_path="ops.log.jsonl",
82
+ # allow_retry_after_indeterminate=True)
83
+ ```
84
+
85
+ That's honest exactly-once-**or-tell-you** semantics. `Indeterminate` is
86
+ refused by default — never silently treated as "safe to retry," never
87
+ silently treated as "must have worked." Resolving it is a manual step on
88
+ purpose: only you (or your ops tooling) can look at the real system and know
89
+ which way it actually went.
90
+
91
+ ## What proves the "once" — the hash chain, not a promise
92
+
93
+ Every `intent`/`executed` row is appended to an `arcaeon-ledger` hash chain:
94
+ `chain = sha256(prev_chain + canonical_json(row_without_chain))[:32]`. Delete
95
+ or edit an inconvenient `executed` row to re-enable a re-fire, and every
96
+ later link in the chain breaks — `receipt()` reports `ledger_ok=False` with a
97
+ `ledger_first_break` naming the row. Deletion doesn't erase the fact that a
98
+ deletion happened.
99
+
100
+ `guard()` does **not** re-verify the whole chain on every call by default —
101
+ that's an `O(rows)` scan of the file, and paying it on every single side
102
+ effect would not scale to a high-volume tool. Pass `verify_integrity=True`
103
+ for that stronger (and slower) guarantee inline, or call `receipt()` /
104
+ `arcaeon_ledger.verify_file()` on your own cadence (a pre-ship gate, a
105
+ nightly job). Tamper caught late is still tamper caught. Tamper never
106
+ checked is a receipt you shouldn't have trusted in the first place — this
107
+ library will not pretend otherwise to look faster in a benchmark.
108
+
109
+ ## Concurrency: exactly one process wins the claim
110
+
111
+ Two processes racing the **same brand-new key** resolve through a single
112
+ SQLite `BEGIN IMMEDIATE` transaction against a small index file next to the
113
+ ledger — the same WAL + immediate-transaction pattern
114
+ [`arcaeon-meter`](https://pypi.org/project/arcaeon-meter/) uses for its usage
115
+ counter. Exactly one caller, ever, gets back the execution claim for a given
116
+ key; the loser gets `AlreadyExecuted` or `Indeterminate` depending on timing,
117
+ never a green light to also run the effect. Verified with two real OS
118
+ processes hammering the same key, and with ten processes released onto a
119
+ brand-new ledger by a wall-clock start barrier (see `test_concurrency.py`),
120
+ not simulated with threads and a comforting mock.
121
+
122
+ The index is created lazily, and that first-touch setup is serialized by a
123
+ cross-process **file lock** — switching a brand-new SQLite file into WAL
124
+ journal mode needs a momentary EXCLUSIVE database lock that does *not* honour
125
+ `busy_timeout`, so without that serialization, N processes first-using the
126
+ same fresh ledger collide on it. If contention still can't be absorbed you get
127
+ `IndexUnavailable`: a typed, documented outcome raised **before** any claim or
128
+ ledger row, never a raw `sqlite3.OperationalError` leaking out of `guard()`.
129
+ (One documented nuance, carried on the exception as `.ledger_committed`: if it
130
+ comes from `done()`/`complete()`, the executed row is already durable in the
131
+ ledger and only the index is stale — duplicate refusal still works, and
132
+ `rebuild_index()` resyncs it. Don't re-run the effect on that one.)
133
+
134
+ That index is consulted **only** to serialize the race at the "nobody has
135
+ claimed this key yet" boundary — every other decision (already executed?
136
+ still an unresolved intent?) is re-derived fresh from the ledger itself on
137
+ every `guard()` call, on purpose: an out-of-band edit to the ledger file
138
+ (a dropped row, a tampered byte) must be reflected immediately even though
139
+ the index file wasn't touched, so a truncation attack degrades to a safe
140
+ refusal (`Indeterminate`) rather than the index quietly vouching for a row
141
+ that's no longer there. The honest cost of that choice: `guard()` scans the
142
+ ledger for the key on every call — `O(rows)` in the ledger's total size, not
143
+ `O(1)`. Fine for a day's or a service's worth of idempotency keys; if you're
144
+ guarding millions of distinct keys against one ledger file, shard the ledger
145
+ (one file per key prefix / tenant / day) rather than expecting this to stay
146
+ `O(1)` — that sharding is on you for now, stated rather than hidden. Lose the
147
+ index file entirely and you lose only the race-serialization fast path, not
148
+ correctness: `rebuild_index()` replays the ledger into a fresh one.
149
+
150
+ ## API surface
151
+
152
+ ```python
153
+ guard(key, *, ledger_path=None, state_db=None, on_duplicate="raise",
154
+ allow_retry_after_indeterminate=False, verify_integrity=False,
155
+ store_outcome=False) -> GuardContext
156
+ ```
157
+ Context manager (`with guard(key) as g: ...; g.done(outcome)`) or decorator
158
+ (`@guard(key)` for a static key, or `@guard(lambda *a, **kw: f"job:{a[0]}")`
159
+ for a key resolved per call). `on_duplicate="raise"` (default) raises
160
+ `AlreadyExecuted`; `"return_receipt"` returns without executing — check
161
+ `g.already_executed` / `g.receipt`, or for the decorator, the call returns
162
+ the `Receipt` directly instead of the wrapped function's result.
163
+
164
+ ```python
165
+ receipt(key, *, ledger_path) -> Receipt
166
+ ```
167
+ The tamper-evident state of a key, read straight from the hash chain (never
168
+ from the SQLite index). `Receipt.state` is `"never"`, `"intent"`, or
169
+ `"executed"`. `bool(receipt)` is `True` only for a clean, verified,
170
+ `"executed"` record — a tampered ledger or an unresolved intent never reads
171
+ as success.
172
+
173
+ ```python
174
+ complete(key, outcome=None, *, ledger_path, store_outcome=False) -> Receipt
175
+ ```
176
+ Mark a key executed directly, without an open `guard()` context — the
177
+ crash-recovery path once you've manually confirmed the effect actually ran.
178
+
179
+ ```python
180
+ rebuild_index(ledger_path, state_db=None) -> int
181
+ ```
182
+ Replay the ledger into a fresh SQLite concurrency index. Restores
183
+ correctness after the index file is lost; not needed for normal operation.
184
+
185
+ ## Drop it into any MCP agent
186
+
187
+ ```json
188
+ {
189
+ "mcpServers": {
190
+ "once": {
191
+ "command": "python",
192
+ "args": ["-m", "arcaeon_once.mcp_server"]
193
+ }
194
+ }
195
+ }
196
+ ```
197
+
198
+ One tool, two actions mirroring the library's own two-phase design so the
199
+ crash window is real even across the MCP boundary: `guard_side_effect(action=
200
+ "claim", key, ...)` before the agent performs the effect (skip it if
201
+ `claimed: false`), `guard_side_effect(action="complete", key, outcome, ...)`
202
+ after it succeeds. If the agent session dies between the two calls, the key
203
+ is left `intent`-only — indeterminate on the next claim, not silently
204
+ resolved by the wrapping.
205
+
206
+ ## Status
207
+
208
+ Core library, CLI, MCP server, all tested: duplicate execution refused with
209
+ the original receipt returned, the crash window (INTENT with no EXECUTED)
210
+ resolving to a typed `Indeterminate` refusal, chain tampering on an executed
211
+ row detected by `receipt()`, a real two-OS-process race resolving to exactly
212
+ one execution, and a ten-process barrier-released first-touch race that lands
213
+ zero untyped exceptions. `python -m arcaeon_once.selftest` ships in the
214
+ package so you can verify the golden digest vector and the planted-tamper
215
+ case on your own machine.
216
+
217
+ MIT.