standard_audit 0.12.1 → 0.13.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: a6da3bdedaecde5d44a57afe47372bf419b9290ace08a2312d1c522df5a67561
4
- data.tar.gz: 283e8c2839ddce46fa2ec6468423c75380256f51f05a0a52844763d5454b7da3
3
+ metadata.gz: 7ca3b859a3eb6993caa5c57a9a43621f67a44871c5ee3f05b2ea5d5ca79860b4
4
+ data.tar.gz: 7a85141df78963fa16c7b223da7a7d48270c368538815f26d8e097a09487c0ff
5
5
  SHA512:
6
- metadata.gz: 96cd2995bbaaab665bc485770604f37b0b4188cf052a30639b804882a70cc276fad7e10b7ee21c65fa0294018687e06cea921ae7e387f4666a3f7298d29bbd09
7
- data.tar.gz: 810d2becd35d3e0a5325786b86415c48318b4c23d7c8df46742916053cff6eaf5f97e3cba1e94b8ab479437751ba94177c28b05bc4c12837208a975698788aa0
6
+ metadata.gz: b74d1bbbd38c3580fb36f4db24f40f7fd01c9f6f807035e3aee8e0a3979692f0c4dc8175cdb53799291d578fdf0e8ca80add04e6df1b09b819425717088abea1
7
+ data.tar.gz: 32dd1f53798d592fe517966dd4677abf42ce5ab7dd650176f56dfcd873025a43a3a81a048e276da37fdf90954cc68d3746fb9355e9c1e40559d123c87ab605af
data/CHANGELOG.md CHANGED
@@ -7,6 +7,213 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [0.13.1] - 2026-09-25
11
+
12
+ Fixes the checksum so it survives PostgreSQL `jsonb`
13
+ ([fundbright/delivery-ops#689](https://github.com/fundbright/delivery-ops/issues/689)).
14
+ Through 0.13.0, the row digest hashed `metadata.to_json` in the order the
15
+ Ruby hash was built. `jsonb` stores keys in its own order (shortest first,
16
+ then bytewise). So any row whose metadata had more than one key failed
17
+ `verify_chain` with `digest_mismatch`, unless its keys happened to be
18
+ written in that order. In fundbright production that is 17,448 of 25,910
19
+ rows (67%). It, not concurrency, was the main cause behind
20
+ fundbright/delivery-ops#433; the 0.8.0 recovery search rescues 24 of those
21
+ rows. Every host on `jsonb` is affected: fundbright, sidekick, jumpdrive,
22
+ luminality and nutripod. The gem's suite runs on SQLite, which keeps JSON as
23
+ text in insertion order, so it never saw the bug.
24
+
25
+ ### ⚠️ Deploy before 2026-10-01T00:00:00Z
26
+
27
+ The new checksum switches on **by the clock, not by deploy**. Rows whose
28
+ `created_at` is at or after `StandardAudit::CANONICAL_CHECKSUM_CUTOVER`
29
+ (**2026-10-01T00:00:00Z**) are signed with the canonical checksum and
30
+ verified strictly with it. **Every process that writes audit rows must run
31
+ 0.13.1 before then.** A web worker or job runner still on an older gem after
32
+ the cutover writes legacy-hashed rows with post-cutover timestamps. Those
33
+ rows fail verification as `:digest_mismatch`, and the only honest fix
34
+ afterwards is to explain them.
35
+
36
+ If your rollout will slip, set `config.canonical_checksum_since` to a later
37
+ time **before 2026-10-01**, in a `configure(baseline: true)` block. Never move
38
+ it once the time has passed: verification recomputes the same decision from
39
+ each row's stored `created_at`, so moving it re-judges rows under the other
40
+ algorithm.
41
+
42
+ ### Upgrade steps
43
+
44
+ 1. Bump to 0.13.1 and deploy **before 2026-10-01T00:00:00Z** (or move
45
+ `config.canonical_checksum_since`, as described above). **No migration or
46
+ new column is needed.**
47
+ 2. Regenerate Sorbet RBIs where the app uses Tapioca. This release adds
48
+ `StandardAudit::Checksum` and new `verify_chain` keywords.
49
+ 3. After the cutover, run `verify_chain` (or `rake standard_audit:verify`)
50
+ and read the new counts; see "What `verify_chain` reports" below. Rows
51
+ created after the cutover should all verify. Record the
52
+ `legacy_unverifiable` count and alert if it ever grows.
53
+ 4. **Decide the policy for legacy rows that can't be reconstructed.** This
54
+ is a human decision (#689 options a to c). The gem reports these rows; it
55
+ does not make the call.
56
+
57
+ **Requires Rails 8.1** (`activerecord`, `activejob`, `activesupport` `>= 8.1`,
58
+ was `>= 8.0`). Every consumer app runs 8.1; 8.0 was never exercised in CI.
59
+
60
+ ### Changed
61
+
62
+ - **Canonical checksum for rows created at or after the cutover.** It is
63
+ the SHA-256 of canonical JSON
64
+ `{"fields": {…}, "previous_checksum": …, "v": 2}`:
65
+ - JSON values (metadata, or any Hash/Array field) are round-tripped exactly
66
+ as the column stores them.
67
+ - Object keys are sorted bytewise at every depth.
68
+ - Integral floats hash as integers, and times as UTC ISO 8601 with
69
+ microseconds.
70
+ - Strings are escaped at the byte level.
71
+ - nil is distinct from `""`.
72
+
73
+ Every input to the digest was audited. `CHECKSUM_FIELDS` is unchanged, and
74
+ `metadata` is the only JSON field in the shipped schema. The parent is
75
+ covered. The canonical form also closes a field-shifting ambiguity in the
76
+ legacy form, where `|` inside a value could move content between adjacent
77
+ fields with the same digest.
78
+ - **Rows created before the cutover keep the legacy digest byte for byte**
79
+ (`StandardAudit::Checksum.legacy_digest`). They are written and verified
80
+ exactly as before.
81
+ - The algorithm is chosen from `created_at` on every write path: `create`,
82
+ the batched `insert_all!` path, and `backfill_checksums!`. Verification and
83
+ `relink_checksums!` make the same choice. `created_at` is fixed before the
84
+ checksum is computed, so the writer and the verifier read the same stored
85
+ value. `created_at` is not itself hashed; see "Alert if it grows" below.
86
+ - `compute_checksum_value` (instance and class) picks the algorithm from
87
+ `created_at`. `version:` forces one (`Checksum::LEGACY` / `CANONICAL`).
88
+ - **`verify_chain` returns `reordered:`, `legacy_unverifiable:` and
89
+ `unverifiable:`**, and takes `key_order_search_limit:` (default 720) and
90
+ `fail_on_legacy_unverifiable:` (default false). Existing keys and reasons
91
+ are unchanged.
92
+ - `rake standard_audit:verify` prints the cutover, the new counts and a tally
93
+ per reason. It takes `FAIL_ON_LEGACY_UNVERIFIABLE=1` and
94
+ `KEY_ORDER_SEARCH_LIMIT=<n>`.
95
+
96
+ ### What `verify_chain` reports
97
+
98
+ **Rows created at or after the cutover** are checked with the canonical
99
+ digest only, against the declared parent, else the preceding row, else the
100
+ recovery search. A mismatch is `:digest_mismatch`. These rows are never
101
+ classified as legacy. Chain linkage holds across the cutover: the first
102
+ canonical row's parent is the last legacy row's checksum.
103
+
104
+ **Rows created before the cutover** are first checked with the legacy digest
105
+ exactly as in 0.13.0: stored order, against the declared parent, else the
106
+ preceding row, else the recovery search. A row that still doesn't reproduce
107
+ goes through these steps:
108
+
109
+ 1. **Key-order reconstruction.** The insertion order can't be read back,
110
+ because `jsonb`'s order depends only on the key set. It can be searched,
111
+ though: an ordering of the stored keys (at every depth) that reproduces the
112
+ digest the row has held since it was written is a witness, in the same
113
+ sense as the parent search. A row whose values were edited reproduces no
114
+ ordering.
115
+ - Such rows count in `reordered` and are valid.
116
+ - The search is bounded by `key_order_search_limit` orderings per row.
117
+ - Orders learned from earlier rows with the same keys are tried first, so
118
+ most rows cost one hash.
119
+ - Parents tried: the declared parent, or else the preceding row and "no
120
+ parent".
121
+ 2. **`:digest_mismatch`** if the search covered every ordering against the
122
+ parent the row declares, or if the row has no JSON object with more than
123
+ one key (key order can't be why it fails).
124
+ 3. **`:missing_parent`** if the declared parent is gone.
125
+ 4. Otherwise **`:legacy_key_order_unverifiable`**: the row has a multi-key
126
+ object and every other check passed. It is listed in `unverifiable` and
127
+ counted in `legacy_unverifiable`.
128
+
129
+ **`valid` semantics.** `valid` is still `failures.empty?`. A
130
+ `:legacy_key_order_unverifiable` row **does not make the chain invalid on its
131
+ own**, because it means "cannot be proven either way": an edited legacy row
132
+ looks exactly like one whose key order was lost. It is never silent. It is
133
+ reported separately with a count, and `fail_on_legacy_unverifiable: true`
134
+ turns these rows into failures.
135
+
136
+ **Alert if it grows.** After the cutover, no new legacy row can be written,
137
+ so under honest operation `legacy_unverifiable` never grows. Growth means a
138
+ pre-cutover row was edited, or a row's `created_at` was moved back across
139
+ the cutover.
140
+
141
+ How many fundbright rows reconstruction rescues depends on their key counts
142
+ and on whether they carry `previous_checksum`. That hasn't been measured on
143
+ production data yet.
144
+
145
+ ### Added
146
+
147
+ - `StandardAudit::CANONICAL_CHECKSUM_CUTOVER` (frozen,
148
+ `Time.utc(2026, 10, 1)`) and `config.canonical_checksum_since`, which
149
+ defaults to it.
150
+ - `StandardAudit::Checksum` (`algorithm_for`, `digest`, `legacy_digest`,
151
+ `canonical_digest`, `canonical_json`) and
152
+ `StandardAudit::Checksum::KeyOrderSearch`.
153
+ - **A PostgreSQL CI leg** (`test (postgres)`: `postgres:16-alpine`, `jsonb`
154
+ metadata) runs the full suite. The SQLite matrix stays. The dummy app uses
155
+ `DATABASE_URL` when it is set. New specs cover:
156
+ - both sides of the cutover, a row just before and just after it, the
157
+ config override, and the batch and backfill paths;
158
+ - linkage across the cutover;
159
+ - the `jsonb` regression. On Postgres the specs also assert that the stored
160
+ order really changed and that the legacy digest fails on it.
161
+
162
+ ### Not included: re-sealing legacy rows
163
+
164
+ No tool re-signs legacy rows with the canonical digest. This is deliberate,
165
+ and a follow-up if you want one. Re-sealing replaces an attestation made at
166
+ write time with one made today. It needs an operator to attest to the rows
167
+ first, and it has to keep the chain intact, because each row's successor
168
+ hashed the row's *old* checksum as its parent. A safe version would need to:
169
+
170
+ 1. keep the original checksum (for example in a `legacy_checksum` column) so
171
+ the successor link can still be checked;
172
+ 2. record who re-sealed which rows and when (a `resealed_at` stamp plus an
173
+ audit event);
174
+ 3. re-seal only rows classified `:legacy_key_order_unverifiable`, never a
175
+ `:digest_mismatch`;
176
+ 4. have `verify_chain` report re-sealed rows separately, so a re-seal never
177
+ reads as original evidence.
178
+
179
+ That is a schema change plus a policy decision, so it is left for a separate
180
+ release. `backfill_checksums!` is still only for rows that never had a
181
+ checksum; don't use it for this.
182
+
183
+ ## [0.13.0] - 2026-09-24
184
+
185
+ The Phase 4 release. It removes what 0.12 deprecated, the empty engine
186
+ routing, and the gaps the five app adoptions of 0.12 ran into.
187
+
188
+ ### Removed (breaking)
189
+
190
+ - **`standard_audit:add_checksums` generator** (deprecated in 0.12.0). It was the 0.2 → 0.3 upgrade path. Every install since 0.3 creates `checksum`; hosts older than 0.8 run `standard_audit:add_previous_checksum`.
191
+ - **`isolate_namespace StandardAudit` and the empty `config/routes.rb`.** The engine has no routes, controllers or views, and no host mounts it. `AuditLog` sets its own `table_name`, so table naming is unchanged. Side effects: `StandardAudit::Engine.routes` no longer holds an (empty) isolated route set, and `StandardAudit.table_name_prefix` / `railtie_namespace` are no longer set by isolation. The gem uses neither.
192
+
193
+ ### Added
194
+
195
+ - **`config.error_reporter = ->(error, context) { … }`**, the destination for every error the gem swallows: `record(raise: false)`, subscriber writes, raising `before_checksum` hooks, and `audit!` write failures under the default policy. nil (the default) keeps `Rails.error.report(error, handled: true, context:)`, so nothing changes unless you set it. Apps that don't forward `Rails.error` to their tracker (jumpdrive-web, nutripod-web) can report straight to Sentry instead of wrapping calls in their own rescue. A raising reporter is logged and ignored. `audit_write_error_handler` still takes precedence for `audit!`.
196
+ - **`entry[:via]` in `before_write`**: `:direct` (`record` without a block, `record_audit`, `audit!`), `:notification` (the ActiveSupport::Notifications subscriber, including `record` with a block), or `:rails_event`. A hook can now tell a direct write from a subscriber write. Not persisted. `StandardAudit::VIA` lists the values.
197
+ - **`hooks:` option on the `"a standard_audit baseline"` shared example**: an Integer (exact `before_checksum` hook count) or an Array of Symbol hook names. The mutation example clears the hooks before the reset, so it fails when a hook lives outside the baseline block.
198
+
199
+ ### Fixed
200
+
201
+ - **Upgrade generators number the migration after the host's newest migration.** `add_anonymized_at`, `add_previous_checksum` and `install` stamped `Time.now`, which sorts before future-dated host migrations. They now use the later of now and one second after the newest migration in the target directory (a real timestamp, unlike ActiveRecord's `+1`, which can produce `…235960`).
202
+ - **README `before_write` example.** It showed a PII guard in `before_write` next to a masking `metadata_builder`. `before_write` runs after the builder, so the guard only ever saw masked values. The README now documents the order (resolvers → `metadata_builder` → `before_write` → dereference → redact → persist) and shows guard-then-mask inside `before_write`.
203
+
204
+ ### Upgrade notes (0.12.x → 0.13.0)
205
+
206
+ Grepped `origin/main` of sidekick-web, jumpdrive-web (control-plane), fundbright-web, luminality-web and nutripod-web on 2026-09-24.
207
+
208
+ **Required host changes: none.** No app runs `add_checksums`, mounts `StandardAudit::Engine`, or uses its route helpers. Regenerate Sorbet RBIs (`bin/tapioca gem standard_audit`, `bin/tapioca dsl`) where the app uses Tapioca.
209
+
210
+ **Optional cleanups this release enables:**
211
+ - sidekick-web `spec/initializers/standard_audit_baseline_spec.rb:49-53` ("still carries the two classification hooks after a reset"): replace with `hooks: 2` on the `it_behaves_like` call.
212
+ - jumpdrive-web `control-plane/app/services/mcp/server.rb:98-104` (`audit_tool` rescue → `ErrorReporting.notify`): replace with `StandardAudit.record(..., raise: false)` plus a `config.error_reporter` that calls `ErrorReporting.notify`. Or keep it, since it also adds `component:` context.
213
+ - nutripod-web `app/controllers/concerns/audit_auth_failure.rb:92-96`: the rescue reports through `Rails.error` itself, so it can become `record(..., raise: false)` with `config.audit_error_context_key = :audit_event`, like the other apps. Add a `config.error_reporter` if nutripod-web does not forward `Rails.error` to Sentry.
214
+ - fundbright-web `AuditWritePolicy` (`before_write`) can use `entry[:via]` if its PII guard should skip gem-published subscriber payloads.
215
+ - Any `before_write` that iterates every `entry` key now also sees `:via`.
216
+
10
217
  ## [0.12.1] - 2026-09-24
11
218
 
12
219
  ### Upgrade steps
data/README.md CHANGED
@@ -102,8 +102,8 @@ When `actor` is omitted, it falls back to the configured `current_actor_resolver
102
102
 
103
103
  Where a missing audit row must never break the request — logging an
104
104
  authentication failure, say — pass `raise: false`. A failed write is logged,
105
- reported to `Rails.error` as handled (context
106
- `{ <audit_error_context_key> => event_type, source: "StandardAudit.record" }`),
105
+ reported through `config.error_reporter` (default: `Rails.error`, as handled;
106
+ context `{ <audit_error_context_key> => event_type, source: "StandardAudit.record" }`),
107
107
  and `record` returns nil:
108
108
 
109
109
  ```ruby
@@ -116,11 +116,30 @@ StandardAudit.record("auth.token_invalid",
116
116
  The default is `raise: true` (unchanged). In block form the option only
117
117
  governs the audit write; errors from your block always propagate.
118
118
 
119
- `raise: false` does not talk to Sentry (or any other tracker) itself. It calls
120
- `Rails.error.report`, so a swallowed failure reaches your error tracker only if
121
- something subscribes to `Rails.error`. `sentry-rails` registers that subscriber
122
- for you. With a hand-rolled Sentry setup, register one
123
- (`Rails.error.subscribe(...)`), or these failures show up only in the log.
119
+ #### Where swallowed failures go: `config.error_reporter`
120
+
121
+ By default the gem reports every error it swallows with
122
+ `Rails.error.report(error, handled: true, context:)`. That covers a failed
123
+ `record(raise: false)`, a failed subscriber write, a raising `before_checksum`
124
+ hook, and a failed `audit!` write under the default policy. It reaches your
125
+ error tracker only if something subscribes to `Rails.error`. `sentry-rails`
126
+ registers that subscriber for you. If your app never forwards `Rails.error`
127
+ (a hand-rolled Sentry setup), point the gem straight at the tracker (0.13.0+):
128
+
129
+ ```ruby
130
+ config.error_reporter = ->(error, context) { Sentry.capture_exception(error, extra: context) }
131
+ ```
132
+
133
+ It receives the error and the context Hash (keyed by
134
+ `audit_error_context_key`) and **replaces** the `Rails.error` call. If you want
135
+ both, call `Rails.error.report` from it too. A reporter that raises is logged
136
+ and ignored, so it can never turn a swallowed audit failure into a raised one.
137
+ `audit_write_error_handler`, when set, still takes precedence for `audit!`
138
+ write failures.
139
+
140
+ **Replace your host code with** `config.error_reporter`. It supersedes any
141
+ rescue-and-report wrapper kept around `record(raise: false)` or `audit!` only
142
+ because the app does not forward `Rails.error` to its tracker.
124
143
 
125
144
  **Replace your host code with** `raise: false`. It supersedes the
126
145
  `AuditAuthFailure#record_auth_failure` rescue-and-report wrapper
@@ -172,18 +191,40 @@ never ran on direct `record` calls.)
172
191
  ```ruby
173
192
  config.before_write = ->(entry) {
174
193
  # entry: { event_type:, actor:, target:, scope:, metadata:,
175
- # request_id:, ip_address:, user_agent:, session_id: }
176
- AuditMetadataPii.verify!(entry[:metadata]) if StandardAudit::Operation::Audit.verify?
177
- if (surface = Current.audit_surface).present?
178
- entry[:metadata] = entry[:metadata].merge(surface: surface)
179
- end
194
+ # request_id:, ip_address:, user_agent:, session_id:, via: }
195
+ metadata = entry[:metadata]
196
+ AuditMetadataPii.verify!(metadata) if StandardAudit::Operation::Audit.verify? && entry[:via] == :direct
197
+ metadata = AuditMetadataPii.mask(metadata)
198
+ metadata = metadata.merge(surface: Current.audit_surface) if Current.audit_surface.present?
199
+ entry[:metadata] = metadata
180
200
  }
181
201
  ```
182
202
 
183
- It runs after the `Current` resolvers and `metadata_builder`, and **before**
184
- dereferencing and `sensitive_keys` redaction, so anything it injects is still
185
- filtered. Mutate `entry` in place; the return value is ignored. Raising aborts
186
- the write: direct callers see the error, the subscribers rescue and report it.
203
+ **Order, per write:** `Current` resolvers → `metadata_builder` → `before_write`
204
+ → record dereferencing → `sensitive_keys` redaction → persist (or buffer, or
205
+ enqueue). So:
206
+
207
+ - `before_write` **sees the builder's output, not the raw metadata.** If your
208
+ `metadata_builder` masks or rewrites values, a guard in `before_write` checks
209
+ the masked values and can never fire. Keep the builder to idempotent
210
+ injection (like `engine_scope`), and do guard-then-mask in `before_write`, in
211
+ that order, as above. Before 0.13 this README showed a guard in
212
+ `before_write` next to a masking builder, which does not work.
213
+ - Anything `before_write` injects is still dereferenced and redacted.
214
+
215
+ Mutate `entry` in place; the return value is ignored. Raising aborts the write:
216
+ direct callers see the error, the subscribers rescue and report it.
217
+
218
+ **`entry[:via]`** (0.13.0+) names the entry point, so a hook can treat direct
219
+ writes differently from subscriber writes. For example, it can apply a guard
220
+ only to rows your own code writes and not to payloads a gem publishes. It is
221
+ not persisted.
222
+
223
+ | `entry[:via]` | Written by |
224
+ |---|---|
225
+ | `:direct` | `StandardAudit.record` (no block), `Auditable#record_audit`, `Operation#audit!` |
226
+ | `:notification` | the ActiveSupport::Notifications subscriber (`subscribe_to` patterns), including `StandardAudit.record` **with a block** |
227
+ | `:rails_event` | the `Rails.event` subscriber (Rails 8.1+) |
187
228
 
188
229
  **Hooks run once per row, batched writes included.** `before_write` and
189
230
  `before_checksum` run for every row inside `StandardAudit.batch` too (at flush
@@ -424,7 +465,8 @@ RSpec.describe "StandardAudit configuration baseline" do
424
465
  catalogue: -> { AuditCatalogue::ACTIONS },
425
466
  sensitive_keys: %i[source_payload],
426
467
  sensitive_key_patterns: [/secret/i],
427
- present: %i[metadata_builder before_write current_scope_resolver]
468
+ present: %i[metadata_builder before_write current_scope_resolver],
469
+ hooks: 2 # before_checksum hooks: a count, or %i[backfill_scope] by name
428
470
  end
429
471
  ```
430
472
 
@@ -433,6 +475,13 @@ restored after a mutation plus `reset_configuration!`. Behaviour held in
433
475
  lambdas can only be checked for presence; keep an app-specific example for
434
476
  anything whose *result* matters.
435
477
 
478
+ `hooks:` (0.13.0+) covers `before_checksum` hooks, which a reset drops unless
479
+ the baseline re-adds them. Pass an Integer to require exactly that many, or an
480
+ Array of Symbol hook names (`config.before_checksum :name`) to require each one.
481
+ The mutation example clears the hooks before the reset, so a hook registered
482
+ outside the baseline block fails it. This replaces hand-written "still carries
483
+ the N hooks after a reset" examples.
484
+
436
485
  **Replace your host code with** the shared example. It supersedes the bulk of
437
486
  each app's `spec/initializers/standard_audit_baseline_spec.rb` (or
438
487
  `spec/config/…`). The `Current.account` / `Current.session&.id` resolver
@@ -527,9 +576,14 @@ StandardAudit.configure(baseline: true) do |config|
527
576
  config.current_scope_resolver = -> { Current.organisation }
528
577
 
529
578
  # -- before_write --
530
- # Runs on every write path, before redaction. See "One write path".
579
+ # Runs on every write path, after metadata_builder and before redaction.
580
+ # entry[:via] is :direct, :notification or :rails_event. See "One write path".
531
581
  config.before_write = ->(entry) { entry[:metadata] = entry[:metadata].merge("surface" => Current.surface) }
532
582
 
583
+ # -- Error reporting --
584
+ # Where swallowed audit failures go. nil (default) = Rails.error.report.
585
+ # config.error_reporter = ->(error, context) { Sentry.capture_exception(error, extra: context) }
586
+
533
587
  # -- Async Processing --
534
588
  # Offload audit log creation to ActiveJob.
535
589
  config.async = false
@@ -546,6 +600,13 @@ StandardAudit.configure(baseline: true) do |config|
546
600
  # Defaults from STANDARD_AUDIT_RETENTION_DAYS (see Retention below); set here
547
601
  # to override per app. Leave unset for infinite retention.
548
602
  config.retention_days = 90
603
+
604
+ # -- Checksum cutover --
605
+ # Rows created at/after this time get the canonical (key-order-independent)
606
+ # checksum. Default StandardAudit::CANONICAL_CHECKSUM_CUTOVER (2026-10-01Z).
607
+ # Only move it if your rollout slips, only to a future time, and never after
608
+ # it has passed. See "Checksum algorithm versions".
609
+ # config.canonical_checksum_since = Time.utc(2026, 10, 15)
549
610
  end
550
611
  ```
551
612
 
@@ -736,7 +797,8 @@ reproduce its own digest. Since 0.12.0 its stored `checksum` is left untouched
736
797
 
737
798
  ```ruby
738
799
  StandardAudit::AuditLog.verify_chain
739
- # => { valid: true, verified: 5576, recovered: 0, redacted: 1, failures: [] }
800
+ # => { valid: true, verified: 5576, recovered: 0, reordered: 0, redacted: 1,
801
+ # legacy_unverifiable: 0, unverifiable: [], failures: [] }
740
802
  ```
741
803
 
742
804
  A redacted row's declared parent is still checked, so deleting the row before
@@ -832,7 +894,8 @@ own digest, so editing it invalidates the row.
832
894
 
833
895
  ```ruby
834
896
  result = StandardAudit::AuditLog.verify_chain
835
- # => { valid: true, verified: 5577, recovered: 0, redacted: 0, failures: [] }
897
+ # => { valid: true, verified: 5577, recovered: 0, reordered: 0, redacted: 0,
898
+ # legacy_unverifiable: 0, unverifiable: [], failures: [] }
836
899
  ```
837
900
 
838
901
  `redacted` counts rows anonymized by `anonymize_actor!` — see "Anonymization
@@ -848,8 +911,13 @@ and the checksum chain".
848
911
  `created_at`, so rows whose two timestamps disagree can leave a hole rather
849
912
  than a pruned start, and a hole is reported — truthfully, since rows really
850
913
  are missing.
914
+ Pre-cutover rows whose metadata key order is lost are listed separately in
915
+ `unverifiable` (`reason: :legacy_key_order_unverifiable`) — see "Checksum
916
+ algorithm versions".
851
917
  - `recovered` counts rows with no `previous_checksum` whose parent had to be
852
918
  found by searching back through recent digests — see below.
919
+ - `reordered`, `legacy_unverifiable` and `unverifiable` concern pre-cutover
920
+ rows only — see "Checksum algorithm versions".
853
921
  - `verify_chain(scope: org)` skips the missing-parent check, because the log is
854
922
  global and a scoped row's parent usually belongs to another scope.
855
923
 
@@ -860,6 +928,105 @@ previous strict-line reading did not actually detect insertions either — it
860
928
  reported every concurrent append as tampering, which on one production log meant
861
929
  67% of rows red and any real signal lost in the noise.
862
930
 
931
+ ### Checksum algorithm versions
932
+
933
+ Up to 0.13.0 the digest hashed `metadata.to_json` in the order the Ruby hash
934
+ was built. PostgreSQL `jsonb` (and MySQL `JSON`) store object keys in their own
935
+ order — shortest first, then bytewise — so a row whose metadata had more than
936
+ one key could only be verified if its keys happened to be written in that
937
+ order. On one production log that was 67% of rows
938
+ ([fundbright/delivery-ops#689](https://github.com/fundbright/delivery-ops/issues/689)).
939
+ SQLite keeps JSON as text in insertion order, so the gem's own suite never saw
940
+ it; CI now runs the suite on PostgreSQL too.
941
+
942
+ There are two algorithms, and **which one a row uses is decided by its
943
+ `created_at`** — nothing extra is stored and no migration is needed:
944
+
945
+ | Rows created | Algorithm | Digest |
946
+ |--------------|-----------|--------|
947
+ | before `config.canonical_checksum_since` | legacy | `SHA256("<parent>\|field=value\|…")`, Hash values via `to_json` in Ruby key order. Kept byte-for-byte so these rows verify as they were signed. |
948
+ | at or after it | canonical | `SHA256` of canonical JSON `{"fields": {…}, "previous_checksum": …, "v": 2}`: JSON values round-tripped exactly as the column stores them, object keys sorted bytewise at every depth, integral floats as integers, times as UTC ISO 8601 (µs), nil distinct from `""`, no separator ambiguity. |
949
+
950
+ `config.canonical_checksum_since` defaults to
951
+ `StandardAudit::CANONICAL_CHECKSUM_CUTOVER`, **2026-10-01T00:00:00Z**. The
952
+ switch happens by the clock, not by deploy:
953
+
954
+ > **Deploy 0.13.1 or later before the cutover.** A process still running an
955
+ > older gem after it writes legacy-hashed rows with post-cutover
956
+ > `created_at`s, and those fail strict canonical verification as
957
+ > `:digest_mismatch`. If your rollout will slip, set
958
+ > `config.canonical_checksum_since` to a later time **before** the default
959
+ > passes — and never move it once it has passed, because verification
960
+ > recomputes the same decision from each row's stored `created_at`.
961
+
962
+ ```ruby
963
+ # config/initializers/standard_audit.rb — only if the rollout slips
964
+ StandardAudit.configure(baseline: true) do |config|
965
+ config.canonical_checksum_since = Time.utc(2026, 10, 15)
966
+ end
967
+ ```
968
+
969
+ Rows written between upgrading and the cutover are still legacy-hashed, so a
970
+ multi-key row written then can still end up unverifiable (below). Deploying
971
+ early does not change that by itself. To switch sooner, set
972
+ `canonical_checksum_since` to a time that is still in the future and after
973
+ every process runs 0.13.1. Never set it to a time that has already passed,
974
+ because rows written since then would be re-judged under the other algorithm.
975
+
976
+ **How `verify_chain` judges each row:**
977
+
978
+ - **At or after the cutover:** the canonical digest only — declared parent,
979
+ else the preceding row, else the recovery search. A mismatch is
980
+ `:digest_mismatch`; these rows are never classified as legacy.
981
+ - **Before the cutover:** the legacy digest, exactly as before (declared
982
+ parent, preceding row, recovery search). A row that still does not
983
+ reproduce is then:
984
+ 1. **Reconstructed.** The insertion order the legacy digest hashed cannot be
985
+ read back — jsonb's order depends only on the key set — but it can be
986
+ *searched*. If some ordering of the stored keys (at every depth), hashed
987
+ with the row's parent, reproduces the digest the row has held since it
988
+ was written, that is a witness in the same sense as the parent recovery
989
+ search: a row whose values were edited reproduces no ordering. Such rows
990
+ count in `reordered` and are valid. At most `key_order_search_limit:`
991
+ orderings per row (default 720, six keys in one object); orders learned
992
+ from earlier rows with the same keys are tried first, so most rows cost
993
+ one hash. Parents tried: the declared one, or else the preceding row and
994
+ "no parent" (not the 256-row window).
995
+ 2. **`:digest_mismatch`** if that search was exhaustive against the parent
996
+ the row *declares* (no key order explains it), or if the row has no JSON
997
+ object with more than one key (key order cannot be why it fails).
998
+ 3. **`:missing_parent`** if its declared parent is gone.
999
+ 4. Otherwise **`:legacy_key_order_unverifiable`**.
1000
+
1001
+ **What `valid` means.** `valid` is `failures.empty?`. A
1002
+ `:legacy_key_order_unverifiable` row does **not** make the chain invalid on its
1003
+ own: it means "cannot be proven either way" — an edited legacy row looks
1004
+ exactly like one whose key order was lost — and the policy for such rows is
1005
+ the host's. It is never silent, though: `legacy_unverifiable` counts them and
1006
+ `unverifiable` lists them (same shape as a failure).
1007
+
1008
+ ```ruby
1009
+ StandardAudit::AuditLog.verify_chain
1010
+ # => { valid: true, verified: 25910, recovered: 24, reordered: <n>, redacted: 0,
1011
+ # legacy_unverifiable: <m>, unverifiable: [...], failures: [] }
1012
+
1013
+ StandardAudit::AuditLog.verify_chain(fail_on_legacy_unverifiable: true)
1014
+ # => the same rows reported in failures, valid: false while any exist
1015
+ ```
1016
+
1017
+ After the cutover no new legacy row can be written, so under honest operation
1018
+ `legacy_unverifiable` never grows. **Alert if it does:** it means a pre-cutover
1019
+ row was edited, or a row's `created_at` was moved back across the cutover
1020
+ (`created_at` is not itself hashed).
1021
+
1022
+ **Cost.** A canonical row costs one JSON round trip and one hash. A legacy row
1023
+ that needs the search costs up to `limit × 2` hashes the first time a key set
1024
+ is seen and usually one after that. Raise the limit (e.g. 5040 for seven keys)
1025
+ if your events carry wider metadata.
1026
+
1027
+ **Re-sealing legacy rows is not provided.** See the 0.13.1 CHANGELOG for why
1028
+ and for what a safe version would need.
1029
+
863
1030
  ### Rows written before 0.8.0
864
1031
 
865
1032
  They have no `previous_checksum`. Verification falls back to the preceding row
@@ -896,13 +1063,20 @@ It is for rows that never had a checksum at all (pre-feature data).
896
1063
  | `standard_audit:install` | `audit_logs` migration + initializer (new installs) |
897
1064
  | `standard_audit:add_previous_checksum` | Adds `previous_checksum` (upgrading from < 0.8) |
898
1065
  | `standard_audit:add_anonymized_at` | Adds `anonymized_at` (upgrading from < 0.12) |
899
- | `standard_audit:add_checksums` | **Deprecated** (0.12.0; to be removed). The 0.2 → 0.3 upgrade path; warns when run |
1066
+
1067
+ The upgrade generators number their migration one second after the newest
1068
+ migration already in `db/migrate` when that is later than now (0.13.0+), so the
1069
+ new migration sorts after future-dated host migrations instead of before them.
1070
+ Before 0.13.0 they stamped the current time. (`standard_audit:add_checksums`,
1071
+ the 0.2 → 0.3 upgrade path, was removed in 0.13.0.)
900
1072
 
901
1073
  ## Rake Tasks
902
1074
 
903
1075
  ```bash
904
1076
  # Verify chain integrity (exits non-zero on failures)
905
1077
  rake standard_audit:verify
1078
+ # ...treating unreconstructable pre-cutover rows as failures too
1079
+ FAIL_ON_LEGACY_UNVERIFIABLE=1 rake standard_audit:verify
906
1080
 
907
1081
  # Record the parent digest each existing row was signed against
908
1082
  rake standard_audit:relink_checksums
@@ -950,6 +1124,10 @@ For PostgreSQL, edit the generated migration to use `jsonb` instead of `json`:
950
1124
  t.jsonb :metadata, default: {}
951
1125
  ```
952
1126
 
1127
+ `jsonb` and MySQL `JSON` reorder object keys. Rows created since the
1128
+ canonical-checksum cutover do not depend on key order; earlier rows may — see
1129
+ "Checksum algorithm versions".
1130
+
953
1131
  ## Best Practices
954
1132
 
955
1133
  **What to audit**: Authentication events, data mutations, permission changes, financial transactions, admin actions, data exports, and API access from external services.