smart-data-engine-sdk 0.1.0.dev0__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
sde/migration.py ADDED
@@ -0,0 +1,820 @@
1
+ """Copying a group into its second engine, and proving the copy is complete.
2
+
3
+ The two halves of a migration that touch data, and therefore the two halves that cannot be ours.
4
+ Copying a row means reading a client's row; comparing two copies means holding both engines open at
5
+ once. We have credentials to neither and never will, which is why this file is in the public
6
+ library and the state machine that gates on it is not.
7
+
8
+ So the division is: **this module produces numbers, and the control plane decides.**
9
+ :class:`VerifyReport` is designed around that boundary - ``as_record()`` carries seven counts and
10
+ nothing else, while the detail an operator needs to *fix* a mismatch stays here, on their machine,
11
+ in a field the record does not have.
12
+
13
+ Three decisions in the backfill are worth reading before the code, because each one is the reason a
14
+ simpler version would lose rows.
15
+
16
+ **The marker is a row count, not a key.** The obvious marker is "the last key copied", which
17
+ resumes exactly. It also needs a codec: a key value has to survive a round trip through whatever
18
+ column the marker table has, for every type a key can be, in every language that will later grow an
19
+ adapter. A codec whose failure mode is a resume point *past* rows that were never copied is a codec
20
+ whose failure mode is silent data loss. A row count cannot fail that way, and the reason is worth
21
+ stating precisely because the loose version of it is false.
22
+
23
+ Resuming means one ``OFFSET`` query - "give me the key of row N". If rows have been inserted below
24
+ that point since, row N is now an *earlier* row, so the resume key moves down and the backfill
25
+ recopies. It does **not** follow that no row is ever stepped over: rows inserted below the new
26
+ resume key are stepped over, and the loose claim that "every error points at recopying" is wrong
27
+ about them. What holds is the statement that matters. Call the rows that existed when this backfill
28
+ began S. Marker N was written when the first N rows in key order were copied, so every row of S at
29
+ or below that boundary key B is copied. The new resume key A is at or below B, because inserting
30
+ rows can only push a given row later in the order - so the rows of S at or below A are a subset of
31
+ those at or below B, and therefore copied. **No row of S is ever skipped**, which is the entire job
32
+ of the backfill. The rows that are stepped over are rows that arrived after it began, and those
33
+ belong to the fan-out - the same premise the absence of a ceiling rests on, below.
34
+
35
+ If the fan-out failed for one of them, it is missing from the copy and it now sits *below* the final
36
+ marker, so :func:`verify` counts it against the chunks rather than against the tail. The attribution
37
+ is off by one mechanism in that case and the row is still caught, which is the right way round.
38
+
39
+ **The chunk is written before the marker moves, and the copy is idempotent.** A crash between the
40
+ two leaves a chunk that will be copied again, which is why the target's key semantics have to
41
+ absorb a duplicate: ``ON CONFLICT DO NOTHING`` in PostgreSQL, ``ReplacingMergeTree`` collapsing
42
+ under ``FINAL`` in ClickHouse. Order the two writes the other way round and the crash window loses
43
+ a chunk permanently. These two decisions hold each other up: the ordering is only safe because the
44
+ copy is idempotent, and the copy only needs to be idempotent because of the ordering.
45
+
46
+ **There is no ceiling, and the backfill does not chase its own tail.** The source keeps growing
47
+ while the copy runs - that is what a live migration is - so a naive "copy until the source stops
48
+ growing" never terminates. It does not need to. ``DUAL_WRITE`` precedes ``BACKFILL``, so every row
49
+ written from then on reaches the copy through the fan-out; the backfill's job is only the rows that
50
+ were there before it started. Reaching the end of the table **once** is therefore enough, and a
51
+ chunk that comes back short is what says so. Every row in the source is then in one of two regions:
52
+ below the final marker and copied here, or written after dual-write began and copied by
53
+ :meth:`sde.Session.save`.
54
+
55
+ That argument has a visible failure mode, and it is the one the gate exists for. If dual-write was
56
+ not actually running everywhere - a deployment half-rolled-out, one process still on the previous
57
+ map - then rows written by the stragglers land above the marker and nothing copies them.
58
+ :func:`verify` reads the tail above the marker and finds them missing, and the migration stops with
59
+ reads still on the source. The design defends itself rather than trusting the operator to have
60
+ sequenced the phases correctly.
61
+
62
+ **On what ``verify`` compares, which is a correction to an earlier design of it.** That design said
63
+ chunks below the marker are immutable and so compare exactly, while the live tail above it is
64
+ checked for containment. The premise is false: "below the marker" is a position in *key* order, and
65
+ a key is not required to increase with insertion time, so a row written during the migration can
66
+ land anywhere - including below the marker, where it is perfectly mutable. What the two regions
67
+ actually distinguish is **which mechanism failed**: a source row missing below the marker means the
68
+ backfill did not copy it, and one missing above means the fan-out did not. Different cause,
69
+ different fix, and worth two counters. The rule applied is the same in both, and it is containment
70
+ with equality of content: every row the source has, the target has, byte for byte in every column.
71
+ Extra rows in the target are not gated on, for the reason the row counts are not - the two reads
72
+ happen at different instants and a write between them is not a defect.
73
+ """
74
+
75
+ from __future__ import annotations
76
+
77
+ from collections.abc import Mapping, Sequence
78
+ from dataclasses import dataclass
79
+ from datetime import UTC, datetime
80
+ from typing import TYPE_CHECKING, Any, Protocol, cast, runtime_checkable
81
+
82
+ from .capabilities import satisfies
83
+ from .errors import EngineError, MigrationRefused
84
+ from .groups import Group, colocation_groups
85
+ from .logging import log
86
+ from .placement import BACKFILL_TABLE, Materialization
87
+
88
+ if TYPE_CHECKING: # pragma: no cover - typing only
89
+ from .session import Session
90
+
91
+ __all__ = [
92
+ "BACKFILL_TABLE",
93
+ "CHUNK_ROWS",
94
+ "DIALECT_PRECISION",
95
+ "PRECISION_INDEPENDENT",
96
+ "BackfillProgress",
97
+ "Difference",
98
+ "EntityProgress",
99
+ "Migratable",
100
+ "VerifyReport",
101
+ "backfill",
102
+ "precision_refusal",
103
+ "verify",
104
+ ]
105
+
106
+ CHUNK_ROWS = 1000
107
+ """Rows per chunk, by default.
108
+
109
+ A thousand rather than a round ten thousand: a chunk is held in memory twice during
110
+ :func:`verify` - the source's rows and the target's - and the number that matters is not throughput
111
+ but how much work a crash discards, which is one chunk.
112
+ """
113
+
114
+ DIALECT_PRECISION: Mapping[tuple[str, str], int] = {
115
+ ("timestamp", "postgres"): 6,
116
+ ("timestamptz", "postgres"): 6,
117
+ ("timestamp", "clickhouse"): 6,
118
+ ("timestamptz", "clickhouse"): 6,
119
+ }
120
+ """Sub-second digits each dialect keeps, for the neutral types where dialects differ.
121
+
122
+ PostgreSQL's ``timestamptz`` is microsecond-resolution and ClickHouse's ``DateTime64(3)`` is
123
+ millisecond-resolution, which is a *stated* choice in :mod:`sde.layout` and a fine one for storage.
124
+ It is not fine for a copy: ``datetime.now()`` has microseconds, so the truncation affects
125
+ essentially every row, and it is silent - the insert succeeds and the value comes back changed. A
126
+ migration in that direction is refused before it copies anything rather than after, because
127
+ :func:`verify` would otherwise find every row mismatched at the end of a copy that took hours.
128
+ """
129
+
130
+ PRECISION_INDEPENDENT: frozenset[str] = frozenset(
131
+ {
132
+ "bool",
133
+ "int32",
134
+ "int64",
135
+ "float32",
136
+ "float64",
137
+ "string",
138
+ "bytes",
139
+ "uuid",
140
+ "date",
141
+ "json",
142
+ }
143
+ )
144
+ """Neutral types a copy between dialects does not silently change.
145
+
146
+ Not the same claim as "every value survives". PostgreSQL's ``date`` has a wider range than
147
+ ClickHouse's ``Date32``, so a date in the year 1800 does not survive that move - but it fails
148
+ *loudly*, on the insert or as a mismatch in :func:`verify`, and no ordinary business date is
149
+ anywhere near the boundary. The line drawn here is silent-and-universal loss, which is a much
150
+ smaller set than lossy, and it is the set worth a refusal that arrives before the work.
151
+
152
+ A neutral type in neither this set nor :data:`DIALECT_PRECISION` refuses the migration, so adding
153
+ one to the vocabulary forces a decision here instead of inheriting an answer nobody made.
154
+ """
155
+
156
+
157
+ @runtime_checkable
158
+ class Migratable(Protocol):
159
+ """What an engine adapter needs to offer for a group to be migrated into or out of it.
160
+
161
+ A separate optional protocol, exactly like :class:`sde.watermark.WatermarkStore` and for the
162
+ same reason: putting these on :class:`sde.session.Engine` would break every adapter anybody has
163
+ written against it, including fakes in someone else's test suite, for a capability our own
164
+ orderbook engine cannot provide. Its schema is fixed in its own source, so it has nowhere to
165
+ keep a marker, and its write path is an update of N price levels rather than a row - neither of
166
+ which can be papered over. Non-participation is therefore a **named refusal** rather than a
167
+ silent skip, because a migration that quietly copies nothing is the worst outcome available
168
+ here.
169
+ """
170
+
171
+ dialect: str
172
+
173
+ def key_range(
174
+ self,
175
+ table: str,
176
+ order: Sequence[str],
177
+ *,
178
+ after: Sequence[Any] | None = None,
179
+ upto: Sequence[Any] | None = None,
180
+ limit: int | None = None,
181
+ ) -> list[dict[str, Any]]: ...
182
+
183
+ def nth_key(
184
+ self, table: str, order: Sequence[str], *, position: int
185
+ ) -> tuple[Any, ...] | None: ...
186
+
187
+ def copy_in(self, table: str, rows: Sequence[Mapping[str, Any]]) -> None: ...
188
+
189
+ def count(self, table: str) -> int: ...
190
+
191
+ def get(self, table: str, key: Mapping[str, Any]) -> dict[str, Any] | None: ...
192
+
193
+ def backfill_marker(self, *, materialization: str, entity: str) -> int: ...
194
+
195
+ def record_backfill_marker(
196
+ self, *, materialization: str, entity: str, rows: int
197
+ ) -> None: ...
198
+
199
+
200
+ def key_columns(order: Sequence[str], table: str) -> tuple[str, ...]:
201
+ """The ordering columns for a keyset scan, refusing an empty one.
202
+
203
+ Here rather than in each adapter so that the two cannot disagree about it, and public because
204
+ an adapter written outside this repository has the same argument to check. An empty order is
205
+ not a scan of everything in an unspecified order - it is a paginated scan with no pagination,
206
+ which returns the same first page forever.
207
+ """
208
+ cols = tuple(str(c) for c in order)
209
+ if not cols:
210
+ raise EngineError(
211
+ f"a keyset scan of {table} needs at least one ordering column. With none, every page "
212
+ f"is the first page and a backfill would copy the same chunk until it was stopped."
213
+ )
214
+ return cols
215
+
216
+
217
+ def same_width(bound: Sequence[Any], cols: Sequence[str], name: str) -> None:
218
+ """A bound has one value per ordering column, or the comparison is not the one intended."""
219
+ if len(bound) != len(cols):
220
+ raise EngineError(
221
+ f"{name} has {len(bound)} values and the order has {len(cols)} columns {list(cols)}. A "
222
+ f"row-value comparison of different widths is not a narrower comparison, it is a "
223
+ f"different one."
224
+ )
225
+
226
+
227
+ @dataclass(frozen=True)
228
+ class _Copy:
229
+ """One entity, from one materialisation to one fan-out target. The unit both passes work in."""
230
+
231
+ entity: str
232
+ key: tuple[str, ...]
233
+ source: Migratable
234
+ source_engine: str
235
+ source_table: str
236
+ target: Migratable
237
+ target_engine: str
238
+ target_id: str
239
+ target_table: str
240
+
241
+
242
+ @dataclass(frozen=True)
243
+ class EntityProgress:
244
+ """How far one entity's copy into one target has got."""
245
+
246
+ entity: str
247
+ engine: str
248
+ table: str
249
+ rows_copied: int
250
+ """The marker: rows of this entity copied into this target, across every run."""
251
+ rows_this_run: int
252
+ chunks: int
253
+ complete: bool
254
+ """Whether the last chunk came back short, which is what "the tail is the fan-out's now" means.
255
+ """
256
+
257
+ def as_record(self) -> dict[str, Any]:
258
+ return {
259
+ "entity": self.entity,
260
+ "engine": self.engine,
261
+ "table": self.table,
262
+ "rows_copied": self.rows_copied,
263
+ "rows_this_run": self.rows_this_run,
264
+ "chunks": self.chunks,
265
+ "complete": self.complete,
266
+ }
267
+
268
+
269
+ @dataclass(frozen=True)
270
+ class BackfillProgress:
271
+ """What one call to :func:`backfill` did, per entity and per target."""
272
+
273
+ group: str
274
+ entities: tuple[EntityProgress, ...]
275
+
276
+ @property
277
+ def complete(self) -> bool:
278
+ """Every entity of every target has reached the end of its table at least once."""
279
+ return all(entity.complete for entity in self.entities)
280
+
281
+ @property
282
+ def rows_this_run(self) -> int:
283
+ return sum(entity.rows_this_run for entity in self.entities)
284
+
285
+ def as_record(self) -> dict[str, Any]:
286
+ return {
287
+ "group": self.group,
288
+ "complete": self.complete,
289
+ "rows_this_run": self.rows_this_run,
290
+ "entities": [entity.as_record() for entity in self.entities],
291
+ }
292
+
293
+ def for_a_human(self) -> str:
294
+ lines = [
295
+ f"backfill of {self.group}: "
296
+ f"{'complete' if self.complete else 'more to do'}, "
297
+ f"{self.rows_this_run} rows this run"
298
+ ]
299
+ for entity in self.entities:
300
+ lines.append(
301
+ f" {entity.entity} -> {entity.engine}.{entity.table}: "
302
+ f"{entity.rows_copied} rows copied "
303
+ f"({entity.rows_this_run} this run, {entity.chunks} chunks)"
304
+ f"{'' if entity.complete else ', more to do'}"
305
+ )
306
+ return "\n".join(lines)
307
+
308
+
309
+ @dataclass(frozen=True)
310
+ class Difference:
311
+ """One source row the target does not have, or has differently.
312
+
313
+ **This holds the client's own data and it is the reason ``as_record()`` does not.** The key is
314
+ here because "which row" is the first thing an operator needs and a count cannot say it; it
315
+ stays on their machine because a row is the one thing that must never travel to us. The two
316
+ facts are the same decision seen from two sides.
317
+ """
318
+
319
+ entity: str
320
+ table: str
321
+ key: Mapping[str, Any]
322
+ columns: tuple[str, ...]
323
+ """Columns whose values differ, or empty when the row is absent from the target altogether."""
324
+
325
+ @property
326
+ def absent(self) -> bool:
327
+ return not self.columns
328
+
329
+ def for_a_human(self) -> str:
330
+ where = "absent" if self.absent else f"differs in {list(self.columns)}"
331
+ return f"{self.entity} {dict(self.key)} -> {where}"
332
+
333
+
334
+ _DIFFERENCES_KEPT = 20
335
+
336
+ _ABSENT = object()
337
+ """Sentinel for "the target row has no such column", so that a stored ``None`` is not that."""
338
+
339
+
340
+ @dataclass(frozen=True)
341
+ class VerifyReport:
342
+ """What the comparison found, in the shape the gate needs and nothing wider.
343
+
344
+ ``as_record()`` is the boundary. It carries seven counts, and the control plane's
345
+ ``VerifyResult`` reads exactly those - so there is no field in which a value of the client's
346
+ could travel, and adding one would be a visible change to this method rather than an accident
347
+ somewhere in a call chain. :attr:`differences` is the other half of that: the detail that makes
348
+ a mismatch fixable, kept here and deliberately absent from the record.
349
+ """
350
+
351
+ at: str
352
+ group: str
353
+ chunks_compared: int
354
+ chunks_mismatched: int
355
+ tail_rows_read: int
356
+ tail_rows_missing_in_target: int
357
+ rows_source: int
358
+ rows_target: int
359
+ differences: tuple[Difference, ...] = ()
360
+ differences_suppressed: int = 0
361
+
362
+ @property
363
+ def matched(self) -> bool:
364
+ """Whether the target holds everything the source holds. Zero tolerance, both terms.
365
+
366
+ Zero rather than a threshold, and the reason is arithmetic rather than principled: any
367
+ non-zero threshold is an answer to "how many of your rows may we lose", and there is no
368
+ number to say out loud there.
369
+ """
370
+ return self.chunks_mismatched == 0 and self.tail_rows_missing_in_target == 0
371
+
372
+ def as_record(self) -> dict[str, Any]:
373
+ """The seven counts the gate reads. **Numbers, never rows** - see the class docstring."""
374
+ return {
375
+ "at": self.at,
376
+ "chunks_compared": self.chunks_compared,
377
+ "chunks_mismatched": self.chunks_mismatched,
378
+ "tail_rows_read": self.tail_rows_read,
379
+ "tail_rows_missing_in_target": self.tail_rows_missing_in_target,
380
+ "rows_source": self.rows_source,
381
+ "rows_target": self.rows_target,
382
+ }
383
+
384
+ def for_a_human(self) -> str:
385
+ lines = [
386
+ f"verify of {self.group} at {self.at}: "
387
+ f"{'matched' if self.matched else 'DID NOT MATCH'}",
388
+ f" below the marker: {self.chunks_compared} chunks compared, "
389
+ f"{self.chunks_mismatched} mismatched"
390
+ + ("" if self.chunks_mismatched == 0 else " <- the backfill did not copy these"),
391
+ f" above the marker: {self.tail_rows_read} rows read, "
392
+ f"{self.tail_rows_missing_in_target} missing in the copy"
393
+ + (
394
+ ""
395
+ if self.tail_rows_missing_in_target == 0
396
+ else " <- the dual-write fan-out did not reach these"
397
+ ),
398
+ f" rows: {self.rows_source} in the source, {self.rows_target} in the copy "
399
+ f"(reported, not gated on: two counts of live tables are taken at different instants)",
400
+ ]
401
+ if self.differences:
402
+ lines.append(
403
+ " the rows below are your own data. They are not part of what is reported to "
404
+ "Smart Data Engines:"
405
+ )
406
+ lines.extend(f" {d.for_a_human()}" for d in self.differences)
407
+ if self.differences_suppressed:
408
+ lines.append(f" ... and {self.differences_suppressed} more")
409
+ return "\n".join(lines)
410
+
411
+
412
+ def _group(session: Session, group: str) -> Group:
413
+ for candidate in colocation_groups(session.model):
414
+ if candidate.name == group:
415
+ return candidate
416
+ raise MigrationRefused(
417
+ f"{group!r} is not a colocation group of this model. It has "
418
+ f"{sorted(g.name for g in colocation_groups(session.model))}."
419
+ )
420
+
421
+
422
+ def _migratable(session: Session, engine_name: str, role: str, group: str) -> Migratable:
423
+ engine = session.engines[engine_name]
424
+ # `satisfies` rather than `isinstance`: a runtime_checkable protocol ignores `__getattr__`, so
425
+ # a client's wrapper around one of our adapters would be refused here for a property of their
426
+ # wrapper. See `sde.capabilities`.
427
+ if not satisfies(engine, Migratable):
428
+ raise MigrationRefused(
429
+ f"{engine_name!r} cannot act as the {role} of a migration of {group!r}: its adapter "
430
+ f"does not offer the row-level operations a copy needs. An engine whose schema is "
431
+ f"fixed in its own source has nowhere to keep a progress marker and no table to scan "
432
+ f"in key order, so this is a property of the engine rather than a missing feature. "
433
+ f"Refused here rather than skipped, because a migration that copies nothing and says "
434
+ f"nothing is the worst thing this module could do."
435
+ )
436
+ # `satisfies` cannot be a `TypeGuard`, because the protocol it checks against is a runtime
437
+ # argument. The narrowing is asserted here and the line above is what makes it true.
438
+ return cast("Migratable", engine)
439
+
440
+
441
+ def precision_refusal(
442
+ *,
443
+ group: str,
444
+ entity: str,
445
+ columns: Mapping[str, str],
446
+ source_dialect: str,
447
+ target_dialect: str,
448
+ ) -> str | None:
449
+ """Why a copy of these columns between these two dialects would change values, or ``None``.
450
+
451
+ A string rather than an exception, and dialect names rather than engines, because **two doors
452
+ ask this question and only one of them used to.** :func:`backfill` refuses a copy that would
453
+ truncate; the write fan-out in :class:`sde.Session` was doing the same truncation, one row at a
454
+ time, for as long as a `also_write` map was in force. Measured on live servers before it was
455
+ fixed: a `timestamptz` written as ``09:30:15.123456`` came back from PostgreSQL unchanged and
456
+ from ClickHouse as ``09:30:15.123``, with no error on either side.
457
+
458
+ That is worse than the case this rule was written for. A migration that truncates is caught by
459
+ `backfill` before it copies anything; a fan-out that truncates is a difference between two live
460
+ copies of one row, and the product's whole claim about a copy is that it is the same data.
461
+ """
462
+ for column, neutral in sorted(columns.items()):
463
+ if neutral.startswith("decimal(") or neutral in PRECISION_INDEPENDENT:
464
+ continue
465
+ here = DIALECT_PRECISION.get((neutral, source_dialect))
466
+ there = DIALECT_PRECISION.get((neutral, target_dialect))
467
+ if here is None or there is None:
468
+ return (
469
+ f"{group}.{entity}.{column} has neutral type {neutral!r}, and this library does "
470
+ f"not know whether {source_dialect} and {target_dialect} store it to the same "
471
+ f"precision. Refused rather than attempted: a type nobody classified is a type "
472
+ f"nobody checked, and the failure mode of guessing here is a value that comes back "
473
+ f"changed with no error anywhere."
474
+ )
475
+ if there < here:
476
+ return (
477
+ f"{group}.{entity}.{column} is {neutral!r}, which {source_dialect} stores to "
478
+ f"{here} sub-second digits and {target_dialect} to {there}. Copying it would "
479
+ f"truncate every value with more precision than that - silently, because the "
480
+ f"insert succeeds and the value comes back changed - and `verify` would then find "
481
+ f"every such row mismatched at the end of the copy rather than before it. Your "
482
+ f"rows may all happen to be aligned to {there} digits, in which case this refusal "
483
+ f"costs you a migration that would have worked; we cannot tell without reading "
484
+ f"your data, and a copy that is faithful only for the values that happen to be "
485
+ f"present is not something to build a gate on."
486
+ )
487
+ return None
488
+
489
+
490
+ def _plan(session: Session, group: str) -> tuple[_Copy, ...]:
491
+ """Every refusal, before a single row moves.
492
+
493
+ A migration is the operation with the least tolerance for a late discovery in this whole
494
+ library: the cost of finding a problem at chunk four thousand is four thousand chunks of the
495
+ client's I/O and an operator who now has to decide whether what has been copied is safe to
496
+ leave. So the shapes, the engines and the types are all settled here.
497
+ """
498
+ members = _group(session, group)
499
+ placement = session.placement.placement_of(group)
500
+ if not placement.also_write:
501
+ raise MigrationRefused(
502
+ f"{group!r} has no fan-out target in this map, so there is nothing to backfill. A "
503
+ f"migration reaches this library as a placement map with 'also_write' - there is no "
504
+ f"phase name in the document and no second channel - so a map without that key is one "
505
+ f"that says this group is not being migrated."
506
+ )
507
+ source_engine = _migratable(session, placement.source.engine, "source", group)
508
+
509
+ copies: list[_Copy] = []
510
+ for copy in placement.also_write:
511
+ target_engine = _migratable(session, copy.engine, "target", group)
512
+ for entity in members.members:
513
+ key = tuple(session.model.entity(entity).key)
514
+ if not key:
515
+ raise MigrationRefused(
516
+ f"{group}.{entity} has no key, so its rows cannot be scanned in a stable order "
517
+ f"and a chunk boundary would not mean anything."
518
+ )
519
+ _shapes_agree(group, entity, placement.source, copy)
520
+ # No precision check here, and its absence is deliberate. `Session.__init__` refuses a
521
+ # map whose fan-out would truncate, over every group, and `_plan` needs a session - so
522
+ # a copy that reaches this line has already been through that door. Asking twice would
523
+ # leave a branch no mutation can reach on its own, which is the shape this repository
524
+ # treats as worse than a missing guard: it reads as coverage.
525
+ copies.append(
526
+ _Copy(
527
+ entity=entity,
528
+ key=key,
529
+ source=source_engine,
530
+ source_engine=placement.source.engine,
531
+ source_table=placement.source.layout.table_for(entity),
532
+ target=target_engine,
533
+ target_engine=copy.engine,
534
+ target_id=copy.id,
535
+ target_table=copy.layout.table_for(entity),
536
+ )
537
+ )
538
+ return tuple(copies)
539
+
540
+
541
+ def _shapes_agree(
542
+ group: str, entity: str, source: Materialization, target: Materialization
543
+ ) -> None:
544
+ """The target's table has the same columns as the source's, or this is not a move.
545
+
546
+ A fan-out target is allowed to be any derived materialisation, and a derived materialisation is
547
+ allowed to be a denormalised wide table - which is a useful thing and not a migration target.
548
+ Filling one means reading the group's relations and assembling rows that exist in no single
549
+ table, and this module copies rows. Refused by name, because the alternative is a copy that
550
+ leaves the extra columns null and looks like it worked.
551
+ """
552
+ here = source.layout.columns.get(entity)
553
+ there = target.layout.columns.get(entity)
554
+ for label, cols, mat in (("source", here, source), ("target", there, target)):
555
+ if not cols:
556
+ raise MigrationRefused(
557
+ f"the {label} materialisation {mat.id!r} of {group!r} does not describe the "
558
+ f"columns of {entity}, so a copy cannot be checked for shape before it starts. "
559
+ f"`ensure_schema` needs them too; a layout with tables and no columns is not one "
560
+ f"this library can apply."
561
+ )
562
+ assert here is not None and there is not None # for mypy; both proved non-empty above
563
+ if set(here) != set(there):
564
+ only_source = sorted(set(here) - set(there))
565
+ only_target = sorted(set(there) - set(here))
566
+ raise MigrationRefused(
567
+ f"{group}.{entity} has different columns in {source.id!r} and {target.id!r} "
568
+ f"(only in the source: {only_source}; only in the target: {only_target}). That is a "
569
+ f"reshape rather than a move: filling a wide table means reading the group's relations "
570
+ f"and assembling rows that exist in no single table, and this module copies rows. A "
571
+ f"copy would leave the extra columns null and look like it had worked."
572
+ )
573
+
574
+
575
+ def backfill(
576
+ session: Session,
577
+ group: str,
578
+ *,
579
+ chunk_rows: int = CHUNK_ROWS,
580
+ stop_after: int | None = None,
581
+ ) -> BackfillProgress:
582
+ """Copy a group's existing rows into every fan-out target the map names. Resumable.
583
+
584
+ Called again after an interruption it picks up from the marker, and called again after
585
+ completion it does nothing - both because the marker is durable and lives in the target engine,
586
+ next to the rows it describes. That is the correct coupling: a target dropped and recreated
587
+ loses its marker with its data, and a marker kept anywhere else would claim work that no longer
588
+ exists.
589
+
590
+ ``stop_after`` bounds the work to that many chunks per entity, for an operator who wants to copy
591
+ for a while and stop, and for a test that needs to interrupt at a known point. ``None`` runs
592
+ each entity to the end of its table.
593
+ """
594
+ if chunk_rows < 1:
595
+ raise MigrationRefused(f"a chunk of {chunk_rows} rows is not a chunk")
596
+ progress: list[EntityProgress] = []
597
+ for copy in _plan(session, group):
598
+ progress.append(
599
+ _backfill_one(copy, group=group, chunk_rows=chunk_rows, stop_after=stop_after)
600
+ )
601
+ return BackfillProgress(group=group, entities=tuple(progress))
602
+
603
+
604
+ def _backfill_one(
605
+ copy: _Copy, *, group: str, chunk_rows: int, stop_after: int | None
606
+ ) -> EntityProgress:
607
+ marker = copy.target.backfill_marker(materialization=copy.target_id, entity=copy.entity)
608
+ after = _resume_point(copy, marker)
609
+ rows_this_run = 0
610
+ chunks = 0
611
+ complete = False
612
+ while stop_after is None or chunks < stop_after:
613
+ rows = copy.source.key_range(
614
+ copy.source_table, copy.key, after=after, limit=chunk_rows
615
+ )
616
+ if not rows:
617
+ complete = True
618
+ break
619
+ # The chunk first, then the marker. A crash between them costs a recopy, which the target's
620
+ # key semantics absorb; the other order costs the chunk, permanently. See the module
621
+ # docstring - this ordering is why the copy has to be idempotent, and the idempotence is
622
+ # why this ordering is free.
623
+ copy.target.copy_in(copy.target_table, rows)
624
+ marker += len(rows)
625
+ copy.target.record_backfill_marker(
626
+ materialization=copy.target_id, entity=copy.entity, rows=marker
627
+ )
628
+ after = tuple(rows[-1][column] for column in copy.key)
629
+ rows_this_run += len(rows)
630
+ chunks += 1
631
+ log(
632
+ "sde.migration.backfill_progress",
633
+ group=group,
634
+ entity=copy.entity,
635
+ engine=copy.target_engine,
636
+ table=copy.target_table,
637
+ chunk=chunks,
638
+ rows=len(rows),
639
+ rows_copied=marker,
640
+ )
641
+ if len(rows) < chunk_rows:
642
+ # The end of the table, once. Rows arriving above this point from here on are the
643
+ # fan-out's, which is why there is no ceiling and no second pass.
644
+ complete = True
645
+ break
646
+ return EntityProgress(
647
+ entity=copy.entity,
648
+ engine=copy.target_engine,
649
+ table=copy.target_table,
650
+ rows_copied=marker,
651
+ rows_this_run=rows_this_run,
652
+ chunks=chunks,
653
+ complete=complete,
654
+ )
655
+
656
+
657
+ def _resume_point(copy: _Copy, marker: int) -> tuple[Any, ...] | None:
658
+ """Turn a row count back into a key, or refuse if the source has lost rows.
659
+
660
+ One ``OFFSET`` scan, paid once per resume rather than once per chunk, which is the trade that
661
+ makes a row count an acceptable marker. If rows have been inserted below this point since the
662
+ marker was written, row N is now an earlier row, so this key moves down and the backfill
663
+ recopies. Rows inserted below the new key are stepped over - see the module docstring for why
664
+ that is safe, and for why "nothing is ever stepped over" is the wrong thing to claim.
665
+
666
+ A source with fewer rows than the marker claims were copied is the one case that refuses. It
667
+ means rows left the source outside this library, and a backfill cannot resume against a table
668
+ that has shrunk: the marker would be describing a table that no longer exists.
669
+ """
670
+ if marker <= 0:
671
+ return None
672
+ position = copy.source.nth_key(copy.source_table, copy.key, position=marker)
673
+ if position is None:
674
+ raise MigrationRefused(
675
+ f"the marker for {copy.entity} in {copy.target_engine} says {marker} rows have been "
676
+ f"copied, and {copy.source_engine}.{copy.source_table} does not have that many. Rows "
677
+ f"have left the source outside this library, so the marker describes a table that no "
678
+ f"longer exists and resuming from it would be guessing. Nothing has been copied by "
679
+ f"this call."
680
+ )
681
+ return position
682
+
683
+
684
+ def verify(
685
+ session: Session, group: str, *, chunk_rows: int = CHUNK_ROWS
686
+ ) -> VerifyReport:
687
+ """Compare both copies of a group and report counts. The gate reads the counts, not the rows.
688
+
689
+ Reads the **source first and the target second**, always, and the order is load-bearing. A write
690
+ is in the source before it is in the copy, so a row read from the source may not have reached
691
+ the copy yet - but the copy is read after the whole source window, so the window has already
692
+ elapsed. Anything still missing is then looked up once more by point read, on what is normally
693
+ an empty set. Reversing the two reads would make this flaky in the direction that stops a
694
+ healthy migration.
695
+ """
696
+ if chunk_rows < 1:
697
+ raise MigrationRefused(f"a chunk of {chunk_rows} rows is not a chunk")
698
+ chunks_compared = 0
699
+ chunks_mismatched = 0
700
+ tail_rows_read = 0
701
+ tail_missing = 0
702
+ rows_source = 0
703
+ rows_target = 0
704
+ differences: list[Difference] = []
705
+ suppressed = 0
706
+
707
+ for copy in _plan(session, group):
708
+ marker = copy.target.backfill_marker(
709
+ materialization=copy.target_id, entity=copy.entity
710
+ )
711
+ rows_source += copy.source.count(copy.source_table)
712
+ rows_target += copy.target.count(copy.target_table)
713
+ after: tuple[Any, ...] | None = None
714
+ seen = 0
715
+ while True:
716
+ below = seen < marker
717
+ want = min(chunk_rows, marker - seen) if below else chunk_rows
718
+ rows = copy.source.key_range(
719
+ copy.source_table, copy.key, after=after, limit=want
720
+ )
721
+ if not rows:
722
+ break
723
+ high = tuple(rows[-1][column] for column in copy.key)
724
+ missing = _missing_in_target(copy, rows, low=after, high=high)
725
+ if below:
726
+ chunks_compared += 1
727
+ if missing:
728
+ chunks_mismatched += 1
729
+ else:
730
+ tail_rows_read += len(rows)
731
+ tail_missing += len(missing)
732
+ for difference in missing:
733
+ if len(differences) < _DIFFERENCES_KEPT:
734
+ differences.append(difference)
735
+ else:
736
+ suppressed += 1
737
+ after = high
738
+ seen += len(rows)
739
+
740
+ return VerifyReport(
741
+ at=datetime.now(UTC).isoformat(),
742
+ group=group,
743
+ chunks_compared=chunks_compared,
744
+ chunks_mismatched=chunks_mismatched,
745
+ tail_rows_read=tail_rows_read,
746
+ tail_rows_missing_in_target=tail_missing,
747
+ rows_source=rows_source,
748
+ rows_target=rows_target,
749
+ differences=tuple(differences),
750
+ differences_suppressed=suppressed,
751
+ )
752
+
753
+
754
+ def _missing_in_target(
755
+ copy: _Copy,
756
+ rows: Sequence[Mapping[str, Any]],
757
+ *,
758
+ low: tuple[Any, ...] | None,
759
+ high: tuple[Any, ...],
760
+ ) -> tuple[Difference, ...]:
761
+ """Which of these source rows the target does not have, or has differently.
762
+
763
+ One windowed read of the target instead of one point read per row, and then point reads only
764
+ for what the window says is missing - which is normally nothing. The second look is what
765
+ absorbs a fan-out that was in flight during the first: it happens after the whole window has
766
+ been read, so the write has had that long to land.
767
+ """
768
+ mirror = copy.target.key_range(copy.target_table, copy.key, after=low, upto=high)
769
+ index = {tuple(row[column] for column in copy.key): row for row in mirror}
770
+ out: list[Difference] = []
771
+ for row in rows:
772
+ key = tuple(row[column] for column in copy.key)
773
+ there = index.get(key)
774
+ if there is not None and not _differing_columns(row, there):
775
+ continue
776
+ # Not there, or there and different. Look once more, directly, before calling it a loss.
777
+ named = {column: row[column] for column in copy.key}
778
+ again = copy.target.get(copy.target_table, named)
779
+ if again is not None:
780
+ differs = _differing_columns(row, again)
781
+ if not differs:
782
+ continue
783
+ out.append(
784
+ Difference(
785
+ entity=copy.entity, table=copy.target_table, key=named, columns=differs
786
+ )
787
+ )
788
+ continue
789
+ out.append(
790
+ Difference(entity=copy.entity, table=copy.target_table, key=named, columns=())
791
+ )
792
+ return tuple(out)
793
+
794
+
795
+ def _differing_columns(
796
+ source: Mapping[str, Any], target: Mapping[str, Any]
797
+ ) -> tuple[str, ...]:
798
+ """Columns of the **source** row whose values the target does not match, by name.
799
+
800
+ Compared value by value rather than by digest, and the reason is diagnostics. Both copies are
801
+ read by one process on one machine, so a checksum would compress a comparison that costs
802
+ nothing to do exactly - and an exact comparison can say *which column* differs, which is the
803
+ difference between an operator who can fix a migration and one who can only stop it.
804
+ Requirement 9.3 asks for checksums to agree; equal values are a stronger statement than equal
805
+ checksums, so this satisfies it rather than departing from it.
806
+
807
+ Over the source's columns and not over the union, which is a decision rather than an oversight.
808
+ ``ensure_schema`` allows a table to have columns the map does not name - it logs
809
+ ``sde.schema.extra_columns`` and permits the write, because a client may have added one outside
810
+ SDE - so a copy table with an extra column is a supported state. Comparing the union would
811
+ then report *every* row as differing, on a column that has nothing to do with the copy, and
812
+ stop a healthy migration. What the map describes is what the copy is about.
813
+ """
814
+ return tuple(
815
+ sorted(
816
+ column
817
+ for column in source
818
+ if source[column] != target.get(column, _ABSENT)
819
+ )
820
+ )