weft-kernel 0.1.0__py3-none-any.whl
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- weft_kernel/__init__.py +178 -0
- weft_kernel/blocking.py +370 -0
- weft_kernel/context.py +250 -0
- weft_kernel/discovery.py +1181 -0
- weft_kernel/errors.py +73 -0
- weft_kernel/fallback.py +159 -0
- weft_kernel/payload/__init__.py +43 -0
- weft_kernel/payload/applicability.py +298 -0
- weft_kernel/payload/ext.py +172 -0
- weft_kernel/payload/ids.py +22 -0
- weft_kernel/payload/lineage.py +90 -0
- weft_kernel/payload/media_type.py +18 -0
- weft_kernel/payload/node.py +260 -0
- weft_kernel/payload/outcome.py +40 -0
- weft_kernel/payload/property.py +60 -0
- weft_kernel/payload/vector.py +28 -0
- weft_kernel/pipeline.py +748 -0
- weft_kernel/py.typed +0 -0
- weft_kernel/registry.py +692 -0
- weft_kernel/resolution.py +1582 -0
- weft_kernel/runner.py +1436 -0
- weft_kernel/seam.py +725 -0
- weft_kernel-0.1.0.dist-info/METADATA +88 -0
- weft_kernel-0.1.0.dist-info/RECORD +27 -0
- weft_kernel-0.1.0.dist-info/WHEEL +4 -0
- weft_kernel-0.1.0.dist-info/licenses/LICENSE +21 -0
- weft_kernel-0.1.0.dist-info/licenses/NOTICE +77 -0
weft_kernel/runner.py
ADDED
|
@@ -0,0 +1,1436 @@
|
|
|
1
|
+
"""The linear runner — an explicit, ordered `StageSpec` list, resolved once and run batch by batch.
|
|
2
|
+
|
|
3
|
+
Specified in `docs/06-phase-0-build.md` step 6 and `docs/02-extension-model.md`
|
|
4
|
+
section 1 ("What a plugin receives", "Composition is typed and checked at
|
|
5
|
+
load"). This is the second of `06`'s three places Phase 0 could accidentally
|
|
6
|
+
settle **G2** (pipeline derivation semantics, open): a pipeline here is a
|
|
7
|
+
`Sequence[StageSpec]` a caller writes out in full. No `extends`, `insert`,
|
|
8
|
+
`replace`, `remove` or `set`, and no derivation of any kind — that machinery
|
|
9
|
+
is Phase 1's, after G2 closes. `06`: *"a plan with no derivation operators
|
|
10
|
+
cannot silently place a stage between two cleaning stages, which is exactly
|
|
11
|
+
why the choice is shaped this way."*
|
|
12
|
+
|
|
13
|
+
**`Lifetime` and `Stage[In, Out]` are built here**, not in an earlier step,
|
|
14
|
+
because nothing before this one needed them: step 3 (the seam) wraps a bare
|
|
15
|
+
async callable with no opinion on what it is a method of; step 4 (the
|
|
16
|
+
passport) hands `Context` to whatever calls it; discovery (step 5) registers
|
|
17
|
+
factories, never instances. The runner is the first thing that has to
|
|
18
|
+
*resolve and run a chain of them*, which is what forces the shape a plugin
|
|
19
|
+
class must have.
|
|
20
|
+
|
|
21
|
+
**Five checks happen before a single batch runs — "at resolution", never
|
|
22
|
+
discovered later as a runtime `KeyError`** (`06` step 6; the fourth is G2's,
|
|
23
|
+
task 1.2, and the fifth is Phase 2's, task 2.28):
|
|
24
|
+
|
|
25
|
+
1. **The plugin exists.** `Registry.entry` already raises `UnknownPluginError`
|
|
26
|
+
naming the contract, the name that was wanted, and every name that *is*
|
|
27
|
+
registered — reused here unchanged, not re-implemented.
|
|
28
|
+
2. **`requires` is produced by an earlier stage.** Checked model by model
|
|
29
|
+
against the `provides` every earlier stage in the same list declared,
|
|
30
|
+
accumulated as resolution walks the list in order. A miss raises
|
|
31
|
+
`UnmetRequiresError` naming the stage, the missing `ExtModel` and its
|
|
32
|
+
`__namespace__` — which doubles as the pack that owns it, the same
|
|
33
|
+
convention `docs/02-extension-model.md` section 1 uses throughout.
|
|
34
|
+
3. **Consecutive stages compose by type.** `Stage[In, Out]` — see below for
|
|
35
|
+
where `In`/`Out` actually come from.
|
|
36
|
+
4. **`intact` was not destroyed by an earlier stage.** `docs/02-extension-model.md`
|
|
37
|
+
§3 → *Ordering constraints*: G5's `requires`/`provides` solve ordering by
|
|
38
|
+
data dependency, and "it cannot solve the cleaning chain, because
|
|
39
|
+
`WhitespaceNormalizer` must run last for being *destructive*, not because
|
|
40
|
+
anyone reads its output." `intact` and `destroys` are the mirror,
|
|
41
|
+
accumulated the same way `provides` is — a running set of every
|
|
42
|
+
`weft_kernel.payload.Property` an earlier stage's `destroys` named, with
|
|
43
|
+
which stage named it — so a later stage needing one `intact` fails,
|
|
44
|
+
naming the stage, the property, the stage that destroyed it, and that the
|
|
45
|
+
only legal positions for the needing stage are *before* the one that
|
|
46
|
+
destroys it. Unlike `requires`/`provides`, `destroys` is **mandatory** on
|
|
47
|
+
every plugin registered for a contract that opts into this — see
|
|
48
|
+
`weft_kernel.registry`'s own docstring for that half, which happens at
|
|
49
|
+
registration, before this module ever sees the plugin; `intact` stays an
|
|
50
|
+
optional convention read defensively here exactly as `requires` already
|
|
51
|
+
is.
|
|
52
|
+
|
|
53
|
+
5. **Every `fallback:` name exists, and stands in for the primary**, task
|
|
54
|
+
2.28. A stage may name other plugins to try when its own refuses the
|
|
55
|
+
position, and the chain that walks them is
|
|
56
|
+
`weft_kernel.fallback.try_in_order` — a combinator over any contract,
|
|
57
|
+
which is why it lives in its own module and not in this one. What
|
|
58
|
+
happens *here* is the lookup: each name is resolved against the
|
|
59
|
+
registry at resolution and refused as `UnknownFallbackError` if
|
|
60
|
+
nothing registered it, so a chain that cannot run fails before a batch
|
|
61
|
+
does rather than on the first document the primary could not read.
|
|
62
|
+
Each is then compared with the primary it would replace — checks 2 and
|
|
63
|
+
4 above were answered by the primary's declarations alone, so a
|
|
64
|
+
fallback demanding more or promising less is refused as
|
|
65
|
+
`FallbackNotSubstitutableError` rather than corrupting a run on
|
|
66
|
+
exactly the documents nobody tested. Both are class-level reads, so
|
|
67
|
+
the candidates themselves are still **built lazily, at try time** and
|
|
68
|
+
an uninstantiable fallback costs nothing while the primary is working
|
|
69
|
+
— see `_chain`, `_attempt` and `_built_of`, the last of which is why a
|
|
70
|
+
fallback the run actually reached is still flushed.
|
|
71
|
+
|
|
72
|
+
**Where a pipeline's types actually come from — a narrowing worth stating
|
|
73
|
+
plainly.** `docs/02-extension-model.md` → *Composition is typed and checked
|
|
74
|
+
at load* writes `Stage[In, Out]` once per *contract*, not once per plugin —
|
|
75
|
+
see that section for the per-capability examples the kernel itself does not
|
|
76
|
+
restate. That reading is what this module implements: a `StageSpec.contract`
|
|
77
|
+
is expected to declare `Stage[In, Out]` as one of its own bases, and the
|
|
78
|
+
composition check reads `In`/`Out` off the *contract* via `__orig_bases__` —
|
|
79
|
+
never off the plugin, which may satisfy a contract structurally with no
|
|
80
|
+
declared base at all. This is a deliberate, documented choice, not an
|
|
81
|
+
oversight: contracts are few and already generic, one plugin implementing
|
|
82
|
+
several contracts would otherwise have to restate the same pair of types
|
|
83
|
+
redundantly, and Phase 0 has no contract yet to prove the alternative against
|
|
84
|
+
— step 7 is expected to follow this convention, and `docs/02-extension-model.md`
|
|
85
|
+
§1 carries the same note.
|
|
86
|
+
|
|
87
|
+
**`requires`, `provides` and `lifetime` are read off the constructed
|
|
88
|
+
instance, defensively, and `Stage` declares none of them.** `run` is the
|
|
89
|
+
only member `Stage`'s body carries. `typing.Protocol` computes
|
|
90
|
+
`__protocol_attrs__` — the set `isinstance` checks against on a
|
|
91
|
+
`@runtime_checkable` Protocol — by walking every base class's `__dict__`,
|
|
92
|
+
`Stage` included; a `ClassVar` declared on `Stage` would therefore become a
|
|
93
|
+
required `isinstance` member of every contract built on `Stage[In, Out]`
|
|
94
|
+
(`Chunker`, `Extractor`, `NodeStore`, …), not merely an inheritable
|
|
95
|
+
convenience. `getattr(instance, name, default)` already supplies
|
|
96
|
+
`Lifetime.RUN` / `()` / `()` when a plugin never sets them — whether or not
|
|
97
|
+
that plugin inherits `Stage` at all — so nothing about correctness depends
|
|
98
|
+
on `Stage` declaring these, only on the runner reading them defensively.
|
|
99
|
+
See `docs/02-extension-model.md`'s Phase 0 step 7 narrowing note.
|
|
100
|
+
|
|
101
|
+
**The instance cache honours `Lifetime`.** Keyed `(tenant_id, contract, name,
|
|
102
|
+
config_hash)`, per `docs/02-extension-model.md` section 1 — `config_hash` is
|
|
103
|
+
a content hash rather than the config object itself, so a config holding an
|
|
104
|
+
unhashable field (a plain `dict`, say) does not make caching impossible.
|
|
105
|
+
`Lifetime.RUN` — the default — is never written to the cache: a fresh
|
|
106
|
+
instance is built every time `resolve()` runs, reused only for the stages of
|
|
107
|
+
*that* `RunnablePipeline`, exactly as "a fresh instance per pipeline run, no
|
|
108
|
+
thread-safety obligation on the author" describes. `Lifetime.PROCESS` is
|
|
109
|
+
written once and read back by every later `resolve()` call sharing the same
|
|
110
|
+
key, on the same `Runner` — which is therefore expected to live for the
|
|
111
|
+
process, not be rebuilt per run, or the cache buys nothing.
|
|
112
|
+
|
|
113
|
+
**`flush()` is the runner's, never a stage's.** `docs/02-extension-model.md`:
|
|
114
|
+
"there is no `persist()`... `add()` may buffer; **the kernel runner calls
|
|
115
|
+
`flush`** at the end of a run and on cancellation, so no plugin can forget
|
|
116
|
+
it." The kernel names no capability, so it cannot know which stages own
|
|
117
|
+
persistence — instead every resolved instance that happens to expose an
|
|
118
|
+
async, callable `flush()` is flushed, once, whether or not it needs one.
|
|
119
|
+
`flush` is documented as idempotent, so calling it on something that has
|
|
120
|
+
nothing to flush is a legitimate no-op, never a hazard. Every stage gets its
|
|
121
|
+
chance regardless of an earlier one's failure — see `_flush_all` — and each
|
|
122
|
+
call carries the same span and attribution `weft_kernel.seam.wrap` gives
|
|
123
|
+
`run()`, via `weft_kernel.seam.wrap_flush`.
|
|
124
|
+
|
|
125
|
+
**`run` returns a `RunSummary`, sized by outcome count, never by payload.**
|
|
126
|
+
`docs/02-extension-model.md` → *Composition is typed and checked at load*:
|
|
127
|
+
"the kernel runner owns batching, so memory is bounded by batch size rather
|
|
128
|
+
than corpus size." Retaining every batch's whole `Outcome` — in particular a
|
|
129
|
+
`Produced` batch's produced value — for the life of a run would make that
|
|
130
|
+
false for the one component that owns the promise: peak memory would grow
|
|
131
|
+
with the corpus, not the batch. `RunSummary` carries only counts of
|
|
132
|
+
`Produced` / `NothingToProduce` / `Failed`, plus the reasons the latter two
|
|
133
|
+
gave, so a batch's payload is free to be collected the moment it has been
|
|
134
|
+
counted.
|
|
135
|
+
|
|
136
|
+
**Applicability is routed here, task 1.6 — `docs/02-extension-model.md` §3 →
|
|
137
|
+
*Applicability*.** `weft_kernel.payload.applicability.Applies` is what a
|
|
138
|
+
stage's `applies_to` declares; this module is "the seam" that section keeps
|
|
139
|
+
promising will do the evaluating, on the identical `getattr(instance,
|
|
140
|
+
"applies_to", ())` footing `_requires_of`/`_intact_of` already read. An
|
|
141
|
+
empty tuple — no plugin written before this task declares one — means every
|
|
142
|
+
node satisfies it vacuously, so `_run_one_batch` below skips straight to
|
|
143
|
+
the single whole-batch call it always made; nothing about a stage that never
|
|
144
|
+
opts in changes. A non-empty tuple, against a payload that is a `Sequence`
|
|
145
|
+
of `Node` (never a bare `str`/`bytes`, which are sequences of the wrong
|
|
146
|
+
thing entirely), splits the batch into its **maximal contiguous runs** of
|
|
147
|
+
"every `Applies` matches" and "at least one does not" — `_segment_by_
|
|
148
|
+
applicability` — runs the stage once per matching run, and threads every
|
|
149
|
+
non-matching run through unchanged, in the position it already held.
|
|
150
|
+
Contiguous *runs*, not a single filtered call over every matching node,
|
|
151
|
+
because a chunker's `Sequence[Node] -> Sequence[Node]` is not shape-
|
|
152
|
+
preserving: calling it once over the whole matching subset would return one
|
|
153
|
+
undifferentiated sequence with no record of where an untouched node
|
|
154
|
+
originally separated two matching ones, and there would be nothing left to
|
|
155
|
+
recombine correctly. This is `docs/11-multimodal.md` §2's own claim made
|
|
156
|
+
literal: "An atomic node passes the chunker unsplit, and the chunker does
|
|
157
|
+
not have to know that" — the chunker's `run` carries no branch on any
|
|
158
|
+
node's content; the branch lives here, once, for every stage that ever
|
|
159
|
+
declares `applies_to`. A payload that is not a routable `Sequence[Node]` at
|
|
160
|
+
all — a `Query`, say, for a stage at the query-path end of `Stage[In, Out]`
|
|
161
|
+
— runs exactly as an empty `applies_to` would: applicability is a node-
|
|
162
|
+
level mechanism, and nothing about the kernel could recognise "this payload
|
|
163
|
+
is nodes" without naming the capability that produces them.
|
|
164
|
+
|
|
165
|
+
**One batch in flight**, per `01` → *Colour*: `run` walks its batches with a
|
|
166
|
+
plain `async for`, awaiting one batch's whole chain of stages before the next
|
|
167
|
+
batch's `__anext__` is even requested. No `asyncio.gather`, no per-batch task
|
|
168
|
+
— parallelism, if a stage wants it, is the stage's own concern over the one
|
|
169
|
+
batch it was handed.
|
|
170
|
+
|
|
171
|
+
**`CancelledError` propagates untouched — including when its own cleanup
|
|
172
|
+
fails.** `run` catches `BaseException`, not `Exception`, around the batch
|
|
173
|
+
loop, so a `CancelledError` raised mid-run still reaches `_flush_all` before
|
|
174
|
+
`raise` sends it on. Every stage still gets its chance to flush. The
|
|
175
|
+
narrower hazard is what `_flush_all` normally does on a flush failure:
|
|
176
|
+
it raises `FlushError`, and raising anything from inside an `except` block
|
|
177
|
+
handling `exc` replaces `exc` on the way out — a `CancelledError` would
|
|
178
|
+
surface only as `FlushError.__context__`, silently converting a cancelled
|
|
179
|
+
task into one that merely raised. So `_flush_all` takes the exception
|
|
180
|
+
already in flight and, when one is present, never raises on top of it: a
|
|
181
|
+
flush failure is attached to it as a note instead. The result is `01` →
|
|
182
|
+
*Colour*'s requirement exactly: never caught, never rewritten, whether or
|
|
183
|
+
not cleanup itself succeeds.
|
|
184
|
+
"""
|
|
185
|
+
|
|
186
|
+
from __future__ import annotations
|
|
187
|
+
|
|
188
|
+
import hashlib
|
|
189
|
+
import typing
|
|
190
|
+
from collections.abc import AsyncIterator, Awaitable, Callable, Sequence
|
|
191
|
+
from dataclasses import dataclass
|
|
192
|
+
from enum import StrEnum
|
|
193
|
+
from typing import Protocol, cast
|
|
194
|
+
|
|
195
|
+
from pydantic import BaseModel, ConfigDict
|
|
196
|
+
|
|
197
|
+
from weft_kernel.context import Context
|
|
198
|
+
from weft_kernel.errors import UnresolvedNameError, WeftError
|
|
199
|
+
from weft_kernel.fallback import Attempt, try_in_order
|
|
200
|
+
from weft_kernel.payload import (
|
|
201
|
+
Applies,
|
|
202
|
+
ExtModel,
|
|
203
|
+
Node,
|
|
204
|
+
NothingToProduce,
|
|
205
|
+
Outcome,
|
|
206
|
+
Produced,
|
|
207
|
+
Property,
|
|
208
|
+
)
|
|
209
|
+
from weft_kernel.registry import Registry, RegistryEntry, UnknownPluginError, unwrap_factory
|
|
210
|
+
from weft_kernel.seam import wrap, wrap_flush
|
|
211
|
+
|
|
212
|
+
|
|
213
|
+
class Lifetime(StrEnum):
|
|
214
|
+
"""How long one resolved stage instance may be reused. `02` §1 → *What a plugin receives*."""
|
|
215
|
+
|
|
216
|
+
RUN = "run"
|
|
217
|
+
PROCESS = "process"
|
|
218
|
+
|
|
219
|
+
|
|
220
|
+
class Stage[In, Out](Protocol):
|
|
221
|
+
"""The shape every registered plugin class satisfies. `02` §1 → *What a plugin receives*.
|
|
222
|
+
|
|
223
|
+
A contract publishes this specialised — see `docs/02-extension-model.md`
|
|
224
|
+
→ *Composition is typed and checked at load* for the per-capability
|
|
225
|
+
examples — which is what the composition check in this module reads back
|
|
226
|
+
via `__orig_bases__`. A plugin implementing that contract does not need
|
|
227
|
+
to name `Stage` itself.
|
|
228
|
+
|
|
229
|
+
**`run` is the only member declared here, deliberately.** `Lifetime`,
|
|
230
|
+
`requires` and `provides` are conventions a plugin *may* set — read
|
|
231
|
+
defensively off the constructed instance by `_lifetime_of`,
|
|
232
|
+
`_requires_of` and `_provides_of` below — never declared as `ClassVar`s
|
|
233
|
+
on this Protocol. See the module docstring, *"`requires`, `provides` and
|
|
234
|
+
`lifetime` are read off the constructed instance, defensively, and
|
|
235
|
+
`Stage` declares none of them."*
|
|
236
|
+
"""
|
|
237
|
+
|
|
238
|
+
async def run(self, payload: In, ctx: Context) -> Outcome[Out]: ...
|
|
239
|
+
|
|
240
|
+
|
|
241
|
+
class PipelineResolutionError(WeftError):
|
|
242
|
+
"""The family base for every way a pipeline can fail to resolve. `02` §3 → *When
|
|
243
|
+
resolution fails*: "each failure is its own `WeftError` subclass under a
|
|
244
|
+
`PipelineResolutionError` family base, all carrying the same required fields — the
|
|
245
|
+
pipeline, the stage ids, the distributions in conflict, and the remedy."
|
|
246
|
+
|
|
247
|
+
**Never raised directly, task 1.13.** It used to be — this class alone covered three
|
|
248
|
+
unrelated checks (an unmet `requires`, two stages that do not compose, an `intact`
|
|
249
|
+
property already destroyed), told apart only by reading the message. `02` §3 names
|
|
250
|
+
exactly the failure this is: "a fat class would present one already-documented name
|
|
251
|
+
and let a dozen new failure modes ship undocumented" to the 0.14 coverage ratchet,
|
|
252
|
+
which derives its required set from subclass *names* — and it is not hypothetical
|
|
253
|
+
here, it is what `weft_kernel.resolution.resolve()` already proved wrong for the
|
|
254
|
+
identical three checks, giving each its own name (`UnmetRequiresError`,
|
|
255
|
+
`StageCompositionError`, `IntactViolationError`, below) for a pipeline *document*.
|
|
256
|
+
Since a `StageSpec` list and a resolved document fail these three checks for the
|
|
257
|
+
literal same reason — `weft_kernel.resolution`'s own docstring already says as much,
|
|
258
|
+
"the same check `weft_kernel.runner.PipelineResolutionError` performs for an explicit
|
|
259
|
+
`StageSpec` list" — the fix is not three *new* names, it is reusing these three,
|
|
260
|
+
defined here because `Runner.resolve` was the first to need them and `weft_kernel.
|
|
261
|
+
resolution` already imports this module's base.
|
|
262
|
+
|
|
263
|
+
**The four fields are real attributes, not prose.** Before task 1.13 the pipeline, the
|
|
264
|
+
failing stage ids, the distributions in conflict and the remedy lived only inside each
|
|
265
|
+
subclass's own formatted message — readable by a person, unreachable by a caller (the
|
|
266
|
+
CLI's exit-code mapping, a future report) without parsing English. `pipeline` is
|
|
267
|
+
`None`, never a placeholder, where there genuinely is none to name — an anonymous
|
|
268
|
+
`StageSpec` list has no name at all, the identical honesty `02` §3 already asks of
|
|
269
|
+
`UnknownParentPipelineError`'s "no stage ids and no distribution to name" for a
|
|
270
|
+
missing parent. `stages` and `distributions` default to `()`, `remedy` to `""`, for
|
|
271
|
+
the same reason: an empty tuple is a fact ("nothing here"), never a lie about data
|
|
272
|
+
that was never computed.
|
|
273
|
+
"""
|
|
274
|
+
|
|
275
|
+
def __init__(
|
|
276
|
+
self,
|
|
277
|
+
message: str,
|
|
278
|
+
*,
|
|
279
|
+
pipeline: str | None = None,
|
|
280
|
+
stages: tuple[str, ...] = (),
|
|
281
|
+
distributions: tuple[str, ...] = (),
|
|
282
|
+
remedy: str = "",
|
|
283
|
+
) -> None:
|
|
284
|
+
super().__init__(message)
|
|
285
|
+
self.pipeline = pipeline
|
|
286
|
+
self.stages = stages
|
|
287
|
+
self.distributions = distributions
|
|
288
|
+
self.remedy = remedy
|
|
289
|
+
|
|
290
|
+
|
|
291
|
+
class UnresolvedNameInPipelineResolutionError:
|
|
292
|
+
"""The shared `__init__` for every `PipelineResolutionError` subclass that is *also*
|
|
293
|
+
`UnresolvedNameError` — task 2.36's own repair.
|
|
294
|
+
|
|
295
|
+
Task 2.36 gave `UnknownParentPipelineError`, `UndefinedVarError` and
|
|
296
|
+
`StaleOperatorTargetError` (`weft_kernel.resolution`) and `UnknownFallbackError`
|
|
297
|
+
(below) an identical 19-line `__init__`: forward the family base's four fields, then
|
|
298
|
+
set `self.valid_options`. Four copies of the same body is exactly the shape `01`'s own
|
|
299
|
+
rule against a fat class exists to prevent one level down — not four failure *kinds*
|
|
300
|
+
sharing a name, but one failure kind (a name that did not resolve against an
|
|
301
|
+
enumerable set of alternatives, inside pipeline resolution specifically) written out
|
|
302
|
+
four times. This class is that body, written once; each of the four now declares it
|
|
303
|
+
as their *first* base — so it is the one Python's MRO finds `__init__` on — and
|
|
304
|
+
declares no `__init__` of its own.
|
|
305
|
+
|
|
306
|
+
**The guarantee this exists to serve is unchanged, not merely preserved by accident.**
|
|
307
|
+
`valid_options` stays a required, keyword-only parameter with no default — the same
|
|
308
|
+
signature the four used to each declare by hand — so a raise site that forgets it
|
|
309
|
+
still fails to construct the exception at all, `TypeError` before a `WeftError` even
|
|
310
|
+
exists to catch. Nothing about collapsing four bodies into one changes what the one
|
|
311
|
+
body requires.
|
|
312
|
+
|
|
313
|
+
**Not itself a `WeftError`, on the identical footing `weft_kernel.errors.
|
|
314
|
+
UnresolvedNameError`'s own docstring already states for the marker it is mixed in
|
|
315
|
+
alongside.** Three separate reasons converge on this one shape:
|
|
316
|
+
|
|
317
|
+
- *A `PipelineResolutionError` base would make it a `WeftError` subclass, and
|
|
318
|
+
`tests/docs/test_troubleshooting_coverage.py`'s own coverage ratchet — task 0.14,
|
|
319
|
+
`08` §3 clause (d) — requires a `manual/troubleshooting.md` entry for every
|
|
320
|
+
`WeftError` subclass the first-party tree defines, pinned-empty waiver only by a
|
|
321
|
+
dated decision-log entry.* This class is never raised on its own — only its four
|
|
322
|
+
concrete subclasses ever are — so it names no failure mode a user ever meets in a
|
|
323
|
+
traceback; writing a troubleshooting entry for a class nobody encounters would be
|
|
324
|
+
documentation invented to satisfy a check rather than to help a reader, exactly
|
|
325
|
+
what `08`'s own rule exists to prevent in the other direction. Not inheriting
|
|
326
|
+
`WeftError` at all is the honest way to keep it out of that requirement, mirroring
|
|
327
|
+
the precedent already sitting one class over: `UnresolvedNameError` itself is "not
|
|
328
|
+
itself a `WeftError`, and defines no `__init__` of its own" for the same reason —
|
|
329
|
+
it is a marker/mixin, not a failure mode with its own identity.
|
|
330
|
+
- *Calling `PipelineResolutionError.__init__(self, ...)` explicitly, never via
|
|
331
|
+
`super()`, is what makes the base unnecessary.* Every concrete subclass below still
|
|
332
|
+
lists `PipelineResolutionError` as an actual base (so `isinstance` and
|
|
333
|
+
`PipelineResolutionError.__subclasses__()` see it exactly as before) — this class
|
|
334
|
+
only has to supply the four raise sites' shared `__init__` *body*, not sit in the
|
|
335
|
+
inheritance chain itself. Listed first among a subclass's bases, so Python's MRO
|
|
336
|
+
finds `__init__` here before it reaches `PipelineResolutionError`'s own. The `cast`
|
|
337
|
+
below is what that split costs statically: this class alone has no way to promise
|
|
338
|
+
pyright that whatever mixes it in also mixes in `PipelineResolutionError`, so the
|
|
339
|
+
call is annotated true rather than left for strict mode to refuse. Every real
|
|
340
|
+
subclass keeps the promise the cast makes.
|
|
341
|
+
- *Not `_`-prefixed, because pyright's strict mode makes that choice unavailable for a
|
|
342
|
+
class actually reused across modules.* A leading underscore was tried first, on
|
|
343
|
+
`weft_kernel.pipeline._QUALIFIER`'s own precedent — but that precedent is for a
|
|
344
|
+
private *constant*, duplicated on purpose rather than imported, exactly because
|
|
345
|
+
duplicating one character is cheaper than coupling two modules over it. A shared
|
|
346
|
+
`__init__` is not that: the entire point of writing it once is that both this
|
|
347
|
+
module (`UnknownFallbackError`) and `weft_kernel.resolution` (the other three)
|
|
348
|
+
construct the *same* class, so duplicating it would silently reproduce the very
|
|
349
|
+
duplication task 2.36's review found. `reportPrivateUsage` under
|
|
350
|
+
`[tool.pyright] typeCheckingMode = "strict"` refuses a leading-underscore name
|
|
351
|
+
imported into another module's source regardless of package boundary, which makes
|
|
352
|
+
the `_QUALIFIER` shape structurally unavailable here — the concrete fact that
|
|
353
|
+
decided the question, not merely a style preference. It stays out of
|
|
354
|
+
`weft_kernel/__init__.py`'s own export list all the same: no contract in `02` gives
|
|
355
|
+
a pack a reason to raise `PipelineResolutionError`'s specific four-field shape —
|
|
356
|
+
that family is raised only by `weft_kernel.resolution.resolve` and `Runner.resolve`
|
|
357
|
+
themselves — so a pack wanting `01` requirement 5's guarantee for its own error
|
|
358
|
+
hierarchy mixes in `UnresolvedNameError` directly instead, exactly as
|
|
359
|
+
`weft_llm.models.UnknownModelError` and every other non-`PipelineResolutionError`
|
|
360
|
+
family member already do.
|
|
361
|
+
|
|
362
|
+
**Deliberately does not mix in `UnresolvedNameError` itself.** If it did, it would be
|
|
363
|
+
a direct subclass of the marker, and fitness function 12's own family walk
|
|
364
|
+
(`test_ff12_unresolvable_name_carries_options.py`'s `_all_unresolved_name_subclasses`,
|
|
365
|
+
which starts from `UnresolvedNameError.__subclasses__()`) would discover it as a 21st
|
|
366
|
+
member alongside the 20 pinned in `NAME_RESOLUTION_FAMILY` — an accounting artifact of
|
|
367
|
+
this refactor, not a new failure kind, and exactly the kind of thing that check exists
|
|
368
|
+
to catch rather than silently absorb. Each of the four concrete subclasses below still
|
|
369
|
+
writes `UnresolvedNameError` as its own base, so `issubclass(cls, UnresolvedNameError)`
|
|
370
|
+
is unchanged for every one of them and the family the fitness function counts stays at
|
|
371
|
+
20. Neither that walk nor `weft_kernel.errors.UnresolvedNameError.__subclasses__()`
|
|
372
|
+
needed to become more recursive for this to hold — both already are, and a class that
|
|
373
|
+
is this one's *user* rather than its subclass is never reached walking downward from
|
|
374
|
+
either the marker or `WeftError`, since this class is a subclass of neither.
|
|
375
|
+
"""
|
|
376
|
+
|
|
377
|
+
valid_options: tuple[str, ...]
|
|
378
|
+
|
|
379
|
+
def __init__(
|
|
380
|
+
self,
|
|
381
|
+
message: str,
|
|
382
|
+
*,
|
|
383
|
+
valid_options: tuple[str, ...],
|
|
384
|
+
pipeline: str | None = None,
|
|
385
|
+
stages: tuple[str, ...] = (),
|
|
386
|
+
distributions: tuple[str, ...] = (),
|
|
387
|
+
remedy: str = "",
|
|
388
|
+
) -> None:
|
|
389
|
+
PipelineResolutionError.__init__(
|
|
390
|
+
cast("PipelineResolutionError", self),
|
|
391
|
+
message,
|
|
392
|
+
pipeline=pipeline,
|
|
393
|
+
stages=stages,
|
|
394
|
+
distributions=distributions,
|
|
395
|
+
remedy=remedy,
|
|
396
|
+
)
|
|
397
|
+
self.valid_options = valid_options
|
|
398
|
+
|
|
399
|
+
|
|
400
|
+
class UnmetRequiresError(PipelineResolutionError):
|
|
401
|
+
"""A stage's `requires` names an `ExtModel` no earlier stage in this list provides.
|
|
402
|
+
|
|
403
|
+
Moved here at task 1.13 from being one of three checks `PipelineResolutionError`
|
|
404
|
+
itself used to raise bare — see that class's own docstring. `weft_kernel.resolution`
|
|
405
|
+
imports this rather than declaring a second class of the same name for the identical
|
|
406
|
+
check against a pipeline *document*: two classes for one kind is exactly what `02` §3
|
|
407
|
+
→ *When resolution fails* rules out ("one class per kind rather than one class with a
|
|
408
|
+
kind field"), and `Runner.resolve` and `weft_kernel.resolution.resolve` disagreeing
|
|
409
|
+
about what to call the same failure would be that rule broken across two modules
|
|
410
|
+
instead of inside one.
|
|
411
|
+
"""
|
|
412
|
+
|
|
413
|
+
|
|
414
|
+
class StageCompositionError(PipelineResolutionError):
|
|
415
|
+
"""Two consecutive stages do not compose: one's `Out` is not the next one's `In` — or a
|
|
416
|
+
contract in the list does not declare `Stage[In, Out]` as a base at all, so there is no
|
|
417
|
+
type pair to compare in the first place. See `UnmetRequiresError`'s own docstring for
|
|
418
|
+
why this is imported by `weft_kernel.resolution` rather than redefined there.
|
|
419
|
+
"""
|
|
420
|
+
|
|
421
|
+
|
|
422
|
+
class IntactViolationError(PipelineResolutionError):
|
|
423
|
+
"""A stage needs a `Property` `intact` that an earlier stage's `destroys` already named.
|
|
424
|
+
|
|
425
|
+
Task 1.2, `02` §3 → *Ordering constraints*. See `UnmetRequiresError`'s own docstring
|
|
426
|
+
for why this is imported by `weft_kernel.resolution` rather than redefined there.
|
|
427
|
+
"""
|
|
428
|
+
|
|
429
|
+
|
|
430
|
+
class UnknownFallbackError(
|
|
431
|
+
UnresolvedNameInPipelineResolutionError, PipelineResolutionError, UnresolvedNameError
|
|
432
|
+
):
|
|
433
|
+
"""A stage's `fallback:` list names a plugin no distribution registered — task 2.28.
|
|
434
|
+
|
|
435
|
+
No `__init__` of its own — task 2.36's repair collapsed this class's own 19-line
|
|
436
|
+
forwarding body into `UnresolvedNameInPipelineResolutionError` above, which it now
|
|
437
|
+
inherits unmodified; see that class's own docstring for why the shared body lives
|
|
438
|
+
here rather than in `weft_kernel.resolution`, and why `valid_options` staying
|
|
439
|
+
required and keyword-only, with no default, is unaffected by the collapse.
|
|
440
|
+
|
|
441
|
+
**Refused here, and deliberately not in `weft_kernel.resolution`.** That module carries
|
|
442
|
+
a document's `fallback:` names through unchecked, on purpose and at length (see
|
|
443
|
+
`ResolvedStage.fallback`'s own docstring): a pipeline document may legitimately name
|
|
444
|
+
`ocr` as a fallback before any pack ships one, and it must stay authorable, storable,
|
|
445
|
+
diffable and derivable while that is true. What this class refuses is a different
|
|
446
|
+
claim — making that document **runnable**. The two are separable because
|
|
447
|
+
`Runner.resolve` is a later step than `resolve()`, and keeping them separate is what
|
|
448
|
+
lets the late-binding promise and the no-silent-fallback rule both hold.
|
|
449
|
+
|
|
450
|
+
**Why not at try time.** Refusing when the chain is first *reached* would make the
|
|
451
|
+
failure depend on encountering a document the primary cannot read, so a pipeline could
|
|
452
|
+
be green for a year and fail in production the first time it met a scanned page.
|
|
453
|
+
Skipping the unknown name instead is `01` requirement 5's silent fallback with extra
|
|
454
|
+
steps — it degrades quality precisely on the inputs the fallback existed for.
|
|
455
|
+
|
|
456
|
+
Fitness function 12's family: `valid_options` is every name registered for the
|
|
457
|
+
stage's own contract.
|
|
458
|
+
"""
|
|
459
|
+
|
|
460
|
+
|
|
461
|
+
class FallbackNotSubstitutableError(PipelineResolutionError):
|
|
462
|
+
"""A `fallback:` entry demands more or promises less than the primary it stands in for.
|
|
463
|
+
|
|
464
|
+
A repair to task 2.28, which checked a fallback only for existence. Every
|
|
465
|
+
`requires`, `intact` and downstream `IntactViolationError` check `resolve` performs
|
|
466
|
+
was answered by the **primary**'s declarations, because the primary is the only
|
|
467
|
+
candidate resolution knows will run. A fallback that `destroys` a `Property` the
|
|
468
|
+
primary does not therefore resolves clean and then corrupts a later stage's input —
|
|
469
|
+
and it does so only on the documents the primary could not read, which is precisely
|
|
470
|
+
where a test never looks. `02` §3 names that asymmetry as the one the kernel exists
|
|
471
|
+
to close: "forgetting `destroys` silently corrupts a stranger's, and the pack that
|
|
472
|
+
caused it never sees a failure." `provides` is the mirror: a downstream
|
|
473
|
+
`UnmetRequiresError` satisfied by the primary's declaration leaves a later stage
|
|
474
|
+
running against an `ExtModel` that is not there.
|
|
475
|
+
|
|
476
|
+
So the rule is one sentence — **a fallback may demand no more and promise no less
|
|
477
|
+
than the primary** — and its four halves are `requires`, `intact`, `provides` and
|
|
478
|
+
`destroys`. Declaring *more* than the primary in the safe direction (providing an
|
|
479
|
+
extra model, destroying one property fewer) is not refused: it is the difference
|
|
480
|
+
that would invalidate a check already made that is.
|
|
481
|
+
|
|
482
|
+
**Read off the class, not off an instance** — `weft_kernel.registry.unwrap_factory`,
|
|
483
|
+
exactly as `weft_kernel.resolution.resolve` and `Registry`'s own `destroys`
|
|
484
|
+
mandatoriness already read the same four declarations. That is what lets this check
|
|
485
|
+
cost a `getattr` while candidates stay built lazily, at try time.
|
|
486
|
+
"""
|
|
487
|
+
|
|
488
|
+
|
|
489
|
+
class TenantMismatchError(WeftError):
|
|
490
|
+
"""`run()` was given a `Context` for a tenant this pipeline was not resolved for.
|
|
491
|
+
|
|
492
|
+
`docs/02-extension-model.md`'s cache note is explicit: the instance
|
|
493
|
+
cache is keyed `(tenant_id, contract, name, config_hash)`, and "the
|
|
494
|
+
tenant is in the key or multi-tenancy is broken on day one." A
|
|
495
|
+
`RunnablePipeline` remembers the tenant `resolve()` built it for; `run()`
|
|
496
|
+
refuses rather than silently executing it against a different tenant's
|
|
497
|
+
`Context`, naming both.
|
|
498
|
+
"""
|
|
499
|
+
|
|
500
|
+
|
|
501
|
+
class FlushError(WeftError):
|
|
502
|
+
"""One or more resolved stages failed to flush. Raised once, after every stage was tried.
|
|
503
|
+
|
|
504
|
+
`Runner._flush_all` gives every stage a chance to flush regardless of an
|
|
505
|
+
earlier one's failure — a completed run must stay durable even when an
|
|
506
|
+
unrelated stage's flush raised, and the store is normally last in the
|
|
507
|
+
list. `__cause__` is the first failure encountered, so its traceback is
|
|
508
|
+
never lost; the message names every stage that failed.
|
|
509
|
+
"""
|
|
510
|
+
|
|
511
|
+
|
|
512
|
+
@dataclass(frozen=True, slots=True)
|
|
513
|
+
class StageSpec:
|
|
514
|
+
"""One named position in an explicit, fully-written-out pipeline.
|
|
515
|
+
|
|
516
|
+
`06` step 6's half of G2's minimal, reversible choice: a pipeline is a
|
|
517
|
+
plain `Sequence[StageSpec]`, never a derivation. `id` is what resolution
|
|
518
|
+
errors and `weft_kernel.seam.wrap`'s span both use to say which stage
|
|
519
|
+
failed or ran.
|
|
520
|
+
|
|
521
|
+
**Not `weft_kernel.pipeline.StageDeclaration`, and 1.3 reconciles the two.**
|
|
522
|
+
That one is a stage as an author *wrote* it — a bare `use:` name and an
|
|
523
|
+
unvalidated `with:` block; this one is a stage already resolved far enough
|
|
524
|
+
to run, carrying the contract type and the plugin's own config object.
|
|
525
|
+
They are the two ends of resolution, which neither of them performs.
|
|
526
|
+
"""
|
|
527
|
+
|
|
528
|
+
id: str
|
|
529
|
+
contract: type[object]
|
|
530
|
+
name: str
|
|
531
|
+
config: object = None
|
|
532
|
+
fallback: tuple[str, ...] = ()
|
|
533
|
+
"""Plugin names to try, in order, when `name` refuses this position — task 2.28.
|
|
534
|
+
|
|
535
|
+
The same list `StageDeclaration.fallback` holds in a document, reaching the runner
|
|
536
|
+
at last. Every entry names a plugin under the *same* `contract`, so composition needs
|
|
537
|
+
no second check: a fallback is substitutable for the primary by construction.
|
|
538
|
+
|
|
539
|
+
**A fallback carries no `with:` block, and that is a stated narrowing rather than an
|
|
540
|
+
oversight.** The document grammar has nowhere to put one — `fallback:` is a list of
|
|
541
|
+
bare names — so a fallback runs on its plugin's own defaults (`_build` calls
|
|
542
|
+
`factory(None)`). Widening the grammar is not this task's, and inventing a place for
|
|
543
|
+
the configuration here would put a second spelling of `with:` in the kernel instead of
|
|
544
|
+
in the document that owns it. The cost is real and is not to be papered over
|
|
545
|
+
elsewhere: one plugin in two configurations — the same parser in two modes — is the
|
|
546
|
+
most natural chain there is, and it cannot be expressed at all until the grammar
|
|
547
|
+
carries a configuration. Nothing routes a failed batch into another pipeline either,
|
|
548
|
+
so there is no mechanism standing in for it today.
|
|
549
|
+
"""
|
|
550
|
+
|
|
551
|
+
|
|
552
|
+
@dataclass(slots=True)
|
|
553
|
+
class _Candidate:
|
|
554
|
+
"""One entry in a stage's chain: how to build a plugin, and the instance once built.
|
|
555
|
+
|
|
556
|
+
Not frozen, unlike everything else resolution produces, and the mutability is the
|
|
557
|
+
point: `instance` is filled in **at try time**, so a fallback that is never reached is
|
|
558
|
+
never constructed and an uninstantiable one costs nothing while the primary works.
|
|
559
|
+
It is also what `_built_of` reads, so a fallback the chain did reach is flushed like
|
|
560
|
+
any other stage rather than losing whatever it buffered.
|
|
561
|
+
"""
|
|
562
|
+
|
|
563
|
+
name: str
|
|
564
|
+
distribution: str
|
|
565
|
+
factory: Callable[..., object]
|
|
566
|
+
instance: object | None = None
|
|
567
|
+
|
|
568
|
+
|
|
569
|
+
@dataclass(frozen=True, slots=True)
|
|
570
|
+
class _ResolvedStage:
|
|
571
|
+
"""One `StageSpec`, checked and instantiated. Built only by `Runner.resolve`."""
|
|
572
|
+
|
|
573
|
+
id: str
|
|
574
|
+
contract_name: str
|
|
575
|
+
plugin_name: str
|
|
576
|
+
distribution: str
|
|
577
|
+
instance: object
|
|
578
|
+
chain: tuple[_Candidate, ...] = ()
|
|
579
|
+
"""`[primary, *fallbacks]` when this stage declared any, and empty when it did not.
|
|
580
|
+
|
|
581
|
+
Empty rather than a one-element chain for a stage with no `fallback:`, so the
|
|
582
|
+
ordinary path through `_invoke_stage` is the one it always was — a combinator
|
|
583
|
+
silently interposed on every stage in the tree would be a change to every pipeline
|
|
584
|
+
in exchange for a uniformity nothing reads.
|
|
585
|
+
"""
|
|
586
|
+
|
|
587
|
+
|
|
588
|
+
@dataclass(frozen=True, slots=True)
|
|
589
|
+
class RunnablePipeline:
|
|
590
|
+
"""A `StageSpec` list after every resolution check in `06` step 6 has passed.
|
|
591
|
+
|
|
592
|
+
What `Runner.run` executes. Never constructed directly — see
|
|
593
|
+
`Runner.resolve`. `tenant_id` is the tenant `resolve()` built every
|
|
594
|
+
cached instance for; `run()` checks a `Context` against it — see
|
|
595
|
+
`TenantMismatchError`.
|
|
596
|
+
|
|
597
|
+
**Renamed from `ResolvedPipeline` at task 1.3.** `docs/02-extension-model.md`
|
|
598
|
+
§3 → *Derivation* uses "the resolved form" for something else entirely: a
|
|
599
|
+
frozen, printable, diffable **Pydantic model** — no live plugin instance
|
|
600
|
+
anywhere in it — that a pipeline document derives into, built by
|
|
601
|
+
`weft_kernel.resolution.resolve`. This dataclass is the opposite of that
|
|
602
|
+
on purpose: `instance` on every `_ResolvedStage` it carries is a
|
|
603
|
+
constructed plugin object, which is exactly what makes it runnable and
|
|
604
|
+
exactly what makes it unfit to log, diff or compare across two runs — an
|
|
605
|
+
instance carries no `__eq__` a comparison could trust and holds open
|
|
606
|
+
resources a frozen data value must not. Keeping the old name on this
|
|
607
|
+
class once the data-shaped one existed would have made "the resolved
|
|
608
|
+
pipeline" ambiguous between the two in every docstring and every log
|
|
609
|
+
line that used it, so this one takes the name that says what it is:
|
|
610
|
+
the runnable — built by a resolver, held by a `Runner`, run once.
|
|
611
|
+
"""
|
|
612
|
+
|
|
613
|
+
tenant_id: str
|
|
614
|
+
stages: tuple[_ResolvedStage, ...]
|
|
615
|
+
|
|
616
|
+
|
|
617
|
+
class RunSummary(BaseModel):
|
|
618
|
+
"""What a run produced, sized by outcome count — never by a batch's payload.
|
|
619
|
+
|
|
620
|
+
See the module docstring, *"`run` returns a `RunSummary`..."*. Every
|
|
621
|
+
field is a count or a tuple of short reason strings, never a produced
|
|
622
|
+
value, so this object's size tracks the number of batches a run saw, not
|
|
623
|
+
the volume of data any one of them carried.
|
|
624
|
+
"""
|
|
625
|
+
|
|
626
|
+
model_config = ConfigDict(frozen=True, extra="forbid")
|
|
627
|
+
|
|
628
|
+
produced: int = 0
|
|
629
|
+
nothing_to_produce: int = 0
|
|
630
|
+
failed: int = 0
|
|
631
|
+
nothing_to_produce_reasons: tuple[str, ...] = ()
|
|
632
|
+
failed_reasons: tuple[str, ...] = ()
|
|
633
|
+
|
|
634
|
+
|
|
635
|
+
_FlushFn = Callable[[], Awaitable[None]]
|
|
636
|
+
|
|
637
|
+
|
|
638
|
+
class Runner:
|
|
639
|
+
"""Resolves an explicit `StageSpec` list once, then runs it batch by batch.
|
|
640
|
+
|
|
641
|
+
One `Runner` is expected to live for the process — its instance cache is
|
|
642
|
+
where `Lifetime.PROCESS` stages are actually reused across separate
|
|
643
|
+
`resolve()` calls; a `Runner` rebuilt per call defeats that half of the
|
|
644
|
+
contract.
|
|
645
|
+
"""
|
|
646
|
+
|
|
647
|
+
def __init__(self, registry: Registry) -> None:
|
|
648
|
+
self._registry = registry
|
|
649
|
+
self._process_cache: dict[tuple[str, type[object], str, str], object] = {}
|
|
650
|
+
|
|
651
|
+
def resolve(
|
|
652
|
+
self,
|
|
653
|
+
specs: Sequence[StageSpec],
|
|
654
|
+
*,
|
|
655
|
+
tenant_id: str,
|
|
656
|
+
entry_type: type[object] | None = None,
|
|
657
|
+
) -> RunnablePipeline:
|
|
658
|
+
"""Check plugin existence, `requires`/`provides`, `intact`/`destroys` and composition.
|
|
659
|
+
|
|
660
|
+
Raises `UnknownPluginError` (via `Registry.entry`) if a plugin name
|
|
661
|
+
was never registered; `UnmetRequiresError` if a `requires` goes
|
|
662
|
+
unmet; `IntactViolationError` if an `intact` property was already
|
|
663
|
+
destroyed by an earlier stage; or `StageCompositionError` if two
|
|
664
|
+
consecutive stages do not compose, or if the first stage cannot
|
|
665
|
+
accept `entry_type` — task 1.13: three distinct
|
|
666
|
+
`PipelineResolutionError` subclasses, never one bare class covering
|
|
667
|
+
all three (see that class's own docstring). Nothing here runs a
|
|
668
|
+
stage — see `run`.
|
|
669
|
+
|
|
670
|
+
`entry_type`, defaulted to `None`, is what the caller is about to hand the
|
|
671
|
+
resolved pipeline's first stage. Leaving it `None` makes no claim about that —
|
|
672
|
+
the honest default for a `Runner` that does not know its own caller — and the
|
|
673
|
+
first stage's declared `In` goes unchecked exactly as before this parameter
|
|
674
|
+
existed.
|
|
675
|
+
"""
|
|
676
|
+
if not specs:
|
|
677
|
+
return RunnablePipeline(tenant_id=tenant_id, stages=())
|
|
678
|
+
|
|
679
|
+
_check_composition(specs, entry_type=entry_type)
|
|
680
|
+
|
|
681
|
+
resolved: list[_ResolvedStage] = []
|
|
682
|
+
provided_models: set[type[ExtModel]] = set()
|
|
683
|
+
destroyed_by: dict[type[Property], str] = {}
|
|
684
|
+
for spec in specs:
|
|
685
|
+
entry = self._registry.entry(spec.contract, spec.name)
|
|
686
|
+
key = (tenant_id, spec.contract, spec.name, _config_hash(spec.config))
|
|
687
|
+
|
|
688
|
+
instance = self._process_cache.get(key)
|
|
689
|
+
if instance is None:
|
|
690
|
+
instance = entry.factory(spec.config)
|
|
691
|
+
if _lifetime_of(instance) is Lifetime.PROCESS:
|
|
692
|
+
self._process_cache[key] = instance
|
|
693
|
+
|
|
694
|
+
for required in _requires_of(instance):
|
|
695
|
+
if required not in provided_models:
|
|
696
|
+
raise UnmetRequiresError(
|
|
697
|
+
f"stage '{spec.id}' ({spec.contract.__name__}:{spec.name}) requires "
|
|
698
|
+
f"'{required.__name__}' — namespace '{required.__namespace__}', "
|
|
699
|
+
f"published by the pack of that name — but no earlier stage in this "
|
|
700
|
+
f"pipeline provides it.",
|
|
701
|
+
stages=(spec.id,),
|
|
702
|
+
distributions=(required.__namespace__,),
|
|
703
|
+
remedy=(
|
|
704
|
+
f"add an earlier stage that provides '{required.__name__}', or "
|
|
705
|
+
f"reorder this StageSpec list so one already does."
|
|
706
|
+
),
|
|
707
|
+
)
|
|
708
|
+
for needed_intact in _intact_of(instance):
|
|
709
|
+
destroyer = destroyed_by.get(needed_intact)
|
|
710
|
+
if destroyer is not None:
|
|
711
|
+
raise IntactViolationError(
|
|
712
|
+
f"stage '{spec.id}' ({spec.contract.__name__}:{spec.name}) needs "
|
|
713
|
+
f"'{needed_intact.__name__}' intact — namespace "
|
|
714
|
+
f"'{needed_intact.__namespace__}' — but stage '{destroyer}' earlier in "
|
|
715
|
+
f"this pipeline already destroys it. The only legal positions for "
|
|
716
|
+
f"'{spec.id}' are before '{destroyer}', never after.",
|
|
717
|
+
stages=(spec.id, destroyer),
|
|
718
|
+
remedy=f"move '{spec.id}' to before '{destroyer}', never after.",
|
|
719
|
+
)
|
|
720
|
+
provided_models.update(_provides_of(instance))
|
|
721
|
+
for destroyed in _destroys_of(instance):
|
|
722
|
+
destroyed_by.setdefault(destroyed, spec.id)
|
|
723
|
+
|
|
724
|
+
resolved.append(
|
|
725
|
+
_ResolvedStage(
|
|
726
|
+
id=spec.id,
|
|
727
|
+
contract_name=spec.contract.__name__,
|
|
728
|
+
plugin_name=spec.name,
|
|
729
|
+
distribution=entry.distribution,
|
|
730
|
+
instance=instance,
|
|
731
|
+
chain=self._chain(spec, primary=entry, instance=instance),
|
|
732
|
+
)
|
|
733
|
+
)
|
|
734
|
+
|
|
735
|
+
return RunnablePipeline(tenant_id=tenant_id, stages=tuple(resolved))
|
|
736
|
+
|
|
737
|
+
def _chain(
|
|
738
|
+
self, spec: StageSpec, *, primary: RegistryEntry, instance: object
|
|
739
|
+
) -> tuple[_Candidate, ...]:
|
|
740
|
+
"""`spec`'s whole chain — `[primary, *fallbacks]` — or empty when it declares none.
|
|
741
|
+
|
|
742
|
+
Every fallback name is looked up **now**, so an unregistered one is refused before
|
|
743
|
+
a single batch runs (`UnknownFallbackError`), and none of them is *built* now, so
|
|
744
|
+
the lookup costs a dict read and nothing more. The primary is already constructed
|
|
745
|
+
— `resolve` did it above, under the full `requires`/`intact`/`Lifetime` treatment
|
|
746
|
+
— and is carried in as candidate zero rather than rebuilt.
|
|
747
|
+
|
|
748
|
+
**And every fallback is checked to be substitutable, which costs no instance.**
|
|
749
|
+
Task 2.28 shipped this checking existence only, on the stated ground that the four
|
|
750
|
+
declarations are read off a constructed instance; that ground was wrong.
|
|
751
|
+
`weft_kernel.resolution.resolve` reads all four off the *unconstructed* class
|
|
752
|
+
through `unwrap_factory`, and `Registry` reads `destroys` the same way at
|
|
753
|
+
registration — so the check is available for a `getattr` with laziness intact, and
|
|
754
|
+
`FallbackNotSubstitutableError` is what it raises. It has to be made here rather
|
|
755
|
+
than left to the author's claim: every `requires`/`intact` check `resolve` performs
|
|
756
|
+
was answered by the *primary*'s declarations, so a fallback that destroys more or
|
|
757
|
+
provides less runs against checks nobody made — and only on the documents the
|
|
758
|
+
primary could not read.
|
|
759
|
+
|
|
760
|
+
What is still not checked is a declaration a plugin computes in `__init__` and
|
|
761
|
+
never states on its class, which reads as empty here. That direction is safe by
|
|
762
|
+
construction: an unstated declaration narrows nothing and fails nothing, it only
|
|
763
|
+
leaves this check with less to compare.
|
|
764
|
+
"""
|
|
765
|
+
if not spec.fallback:
|
|
766
|
+
return ()
|
|
767
|
+
expected = _declared_by(unwrap_factory(primary.factory))
|
|
768
|
+
candidates = [
|
|
769
|
+
_Candidate(
|
|
770
|
+
name=spec.name,
|
|
771
|
+
distribution=primary.distribution,
|
|
772
|
+
factory=primary.factory,
|
|
773
|
+
instance=instance,
|
|
774
|
+
)
|
|
775
|
+
]
|
|
776
|
+
for position, name in enumerate(spec.fallback, start=1):
|
|
777
|
+
try:
|
|
778
|
+
entry = self._registry.entry(spec.contract, name)
|
|
779
|
+
except UnknownPluginError as exc:
|
|
780
|
+
registered = tuple(sorted(self._registry.names_for(spec.contract)))
|
|
781
|
+
options = (
|
|
782
|
+
", ".join(f"'{option}'" for option in registered) if registered else "none"
|
|
783
|
+
)
|
|
784
|
+
raise UnknownFallbackError(
|
|
785
|
+
f"stage '{spec.id}' names '{name}' as fallback {position} of "
|
|
786
|
+
f"{len(spec.fallback)}, but no distribution registered that name for "
|
|
787
|
+
f"{spec.contract.__name__}, so this pipeline cannot be run. Names "
|
|
788
|
+
f"registered for {spec.contract.__name__}: {options}.",
|
|
789
|
+
valid_options=registered,
|
|
790
|
+
stages=(spec.id,),
|
|
791
|
+
remedy=(
|
|
792
|
+
f"install a distribution that registers '{name}' for "
|
|
793
|
+
f"{spec.contract.__name__}, or remove '{name}' from stage "
|
|
794
|
+
f"'{spec.id}'s fallback list."
|
|
795
|
+
),
|
|
796
|
+
) from exc
|
|
797
|
+
differences = _substitutability_differences(
|
|
798
|
+
expected, _declared_by(unwrap_factory(entry.factory))
|
|
799
|
+
)
|
|
800
|
+
if differences:
|
|
801
|
+
raise FallbackNotSubstitutableError(
|
|
802
|
+
f"stage '{spec.id}' names '{name}' as fallback {position} of "
|
|
803
|
+
f"{len(spec.fallback)}, but it cannot stand in for '{spec.name}' at that "
|
|
804
|
+
f"position: it {'; and it '.join(differences)}. Every requires, intact and "
|
|
805
|
+
f"ordering check this pipeline passed was answered by '{spec.name}'s "
|
|
806
|
+
f"declarations, and a chain reaching '{name}' would run against checks "
|
|
807
|
+
f"nobody made.",
|
|
808
|
+
stages=(spec.id,),
|
|
809
|
+
distributions=(primary.distribution, entry.distribution),
|
|
810
|
+
remedy=(
|
|
811
|
+
f"declare on '{name}' what '{spec.name}' declares — a fallback may "
|
|
812
|
+
f"demand no more and promise no less than the plugin it replaces — or "
|
|
813
|
+
f"give '{name}' a stage of its own instead of '{spec.id}'s fallback list."
|
|
814
|
+
),
|
|
815
|
+
)
|
|
816
|
+
candidates.append(
|
|
817
|
+
_Candidate(name=name, distribution=entry.distribution, factory=entry.factory)
|
|
818
|
+
)
|
|
819
|
+
return tuple(candidates)
|
|
820
|
+
|
|
821
|
+
async def run(
|
|
822
|
+
self, pipeline: RunnablePipeline, batches: AsyncIterator[object], ctx: Context
|
|
823
|
+
) -> RunSummary:
|
|
824
|
+
"""Run every batch through the whole stage list, one batch in flight, `flush()` at the end.
|
|
825
|
+
|
|
826
|
+
Raises `TenantMismatchError`, naming both tenants, if `ctx.tenant_id`
|
|
827
|
+
does not match the tenant `pipeline` was resolved for — the instance
|
|
828
|
+
cache is keyed by tenant, so running it for a different one would
|
|
829
|
+
silently reach across that boundary. A batch that any stage answers
|
|
830
|
+
with `NothingToProduce` or `Failed` stops there — the remaining
|
|
831
|
+
stages never see it, and that outcome counts toward the returned
|
|
832
|
+
`RunSummary` without its payload (there is none) or reason being
|
|
833
|
+
retained past that count. `flush()` is called on every resolved
|
|
834
|
+
instance that has one, once, whether this loop finishes normally or
|
|
835
|
+
an exception — `CancelledError` above all — cuts it off mid-batch; a
|
|
836
|
+
flush failure in that second case never displaces the exception
|
|
837
|
+
already propagating, which continues on unmodified. See `_flush_all`.
|
|
838
|
+
"""
|
|
839
|
+
_check_tenant(pipeline, ctx, called="run()")
|
|
840
|
+
|
|
841
|
+
produced = 0
|
|
842
|
+
nothing_to_produce = 0
|
|
843
|
+
failed = 0
|
|
844
|
+
nothing_to_produce_reasons: list[str] = []
|
|
845
|
+
failed_reasons: list[str] = []
|
|
846
|
+
try:
|
|
847
|
+
async for batch in batches:
|
|
848
|
+
outcome = await self._run_one_batch(pipeline, batch, ctx)
|
|
849
|
+
if isinstance(outcome, Produced):
|
|
850
|
+
produced += 1
|
|
851
|
+
elif isinstance(outcome, NothingToProduce):
|
|
852
|
+
nothing_to_produce += 1
|
|
853
|
+
nothing_to_produce_reasons.append(outcome.reason)
|
|
854
|
+
else:
|
|
855
|
+
failed += 1
|
|
856
|
+
failed_reasons.append(outcome.reason)
|
|
857
|
+
except BaseException as exc:
|
|
858
|
+
# An exception already in flight — `CancelledError` above all — must reach the
|
|
859
|
+
# caller untouched. `_flush_all` still runs and every stage still gets its chance,
|
|
860
|
+
# but a flush failure is attached to `exc` rather than raised as `FlushError`,
|
|
861
|
+
# which would otherwise replace `exc` in Python's exception chain. See the module
|
|
862
|
+
# docstring, *"`CancelledError` propagates untouched."*
|
|
863
|
+
await self._flush_all(pipeline, in_flight=exc)
|
|
864
|
+
raise
|
|
865
|
+
else:
|
|
866
|
+
await self._flush_all(pipeline)
|
|
867
|
+
return RunSummary(
|
|
868
|
+
produced=produced,
|
|
869
|
+
nothing_to_produce=nothing_to_produce,
|
|
870
|
+
failed=failed,
|
|
871
|
+
nothing_to_produce_reasons=tuple(nothing_to_produce_reasons),
|
|
872
|
+
failed_reasons=tuple(failed_reasons),
|
|
873
|
+
)
|
|
874
|
+
|
|
875
|
+
async def run_once(
|
|
876
|
+
self, pipeline: RunnablePipeline, payload: object, ctx: Context
|
|
877
|
+
) -> Outcome[object]:
|
|
878
|
+
"""Run one payload through the whole stage list and give the caller its outcome back.
|
|
879
|
+
|
|
880
|
+
Task **2.4**. `run` above returns a `RunSummary` — counts and reason
|
|
881
|
+
strings, no payload — which is right for ingest, where the result is
|
|
882
|
+
in the store and a batch's payload is nobody's business once it has
|
|
883
|
+
landed. A query path is the opposite shape: exactly one payload, and
|
|
884
|
+
the payload *is* the point. There is no way to get it out of `run`,
|
|
885
|
+
and the alternatives are worse than fifteen lines here — a pack
|
|
886
|
+
re-implementing applicability routing, seam wrapping, tenant checking
|
|
887
|
+
and flush is exactly the four things this module's own docstring says
|
|
888
|
+
not to hand-write, and a mutable box passed in through a `Context`
|
|
889
|
+
field would make the answer a side effect.
|
|
890
|
+
|
|
891
|
+
Everything else is `run`'s behaviour unchanged, because it is
|
|
892
|
+
literally the same code: the same tenant check, the same
|
|
893
|
+
`_run_one_batch` walk (so `applies_to`, the seam and fallback chains
|
|
894
|
+
all still apply), and the same `_flush_all` discipline — flushed once
|
|
895
|
+
on the way out, flushed again with the exception in flight when one
|
|
896
|
+
cuts the walk off, `CancelledError` above all, which continues on
|
|
897
|
+
unmodified.
|
|
898
|
+
"""
|
|
899
|
+
_check_tenant(pipeline, ctx, called="run_once()")
|
|
900
|
+
try:
|
|
901
|
+
outcome = await self._run_one_batch(pipeline, payload, ctx)
|
|
902
|
+
except BaseException as exc:
|
|
903
|
+
await self._flush_all(pipeline, in_flight=exc)
|
|
904
|
+
raise
|
|
905
|
+
await self._flush_all(pipeline)
|
|
906
|
+
return outcome
|
|
907
|
+
|
|
908
|
+
async def _run_one_batch(
|
|
909
|
+
self, pipeline: RunnablePipeline, batch: object, ctx: Context
|
|
910
|
+
) -> Outcome[object]:
|
|
911
|
+
payload = batch
|
|
912
|
+
for stage in pipeline.stages:
|
|
913
|
+
outcome = await self._run_stage(stage, payload, ctx)
|
|
914
|
+
if not isinstance(outcome, Produced):
|
|
915
|
+
return outcome
|
|
916
|
+
payload = outcome.value
|
|
917
|
+
return Produced(value=payload)
|
|
918
|
+
|
|
919
|
+
async def _run_stage(
|
|
920
|
+
self, stage: _ResolvedStage, payload: object, ctx: Context
|
|
921
|
+
) -> Outcome[object]:
|
|
922
|
+
"""One stage's whole contribution to a batch — routed by `applies_to`, task 1.6.
|
|
923
|
+
|
|
924
|
+
See the module docstring, *"Applicability is routed here"*. An empty
|
|
925
|
+
`applies_to` (every stage before this task, and most after it) or an
|
|
926
|
+
unroutable payload calls `stage.instance.run` once, exactly as
|
|
927
|
+
before this task existed. A non-empty `applies_to` over a
|
|
928
|
+
`Sequence[Node]` instead calls it once per maximal contiguous run of
|
|
929
|
+
matching nodes — `_segment_by_applicability` — threading every
|
|
930
|
+
non-matching run through untouched, and stops at the first outcome
|
|
931
|
+
that is not `Produced`, exactly the short-circuit `_run_one_batch`
|
|
932
|
+
itself already applies stage to stage.
|
|
933
|
+
"""
|
|
934
|
+
applies_to = _applies_to_of(stage.instance)
|
|
935
|
+
if not applies_to or not _is_routable(payload):
|
|
936
|
+
return await self._invoke_stage(stage, payload, ctx)
|
|
937
|
+
|
|
938
|
+
items = cast("Sequence[object]", payload)
|
|
939
|
+
produced: list[object] = []
|
|
940
|
+
for matches, segment in _segment_by_applicability(items, applies_to):
|
|
941
|
+
if not matches:
|
|
942
|
+
produced.extend(segment)
|
|
943
|
+
continue
|
|
944
|
+
outcome = await self._invoke_stage(stage, segment, ctx)
|
|
945
|
+
if not isinstance(outcome, Produced):
|
|
946
|
+
return outcome
|
|
947
|
+
produced.extend(cast("Sequence[object]", outcome.value))
|
|
948
|
+
return Produced(value=tuple(produced))
|
|
949
|
+
|
|
950
|
+
async def _invoke_stage(
|
|
951
|
+
self, stage: _ResolvedStage, payload: object, ctx: Context
|
|
952
|
+
) -> Outcome[object]:
|
|
953
|
+
"""One call into this stage's position — through its whole chain when it has one.
|
|
954
|
+
|
|
955
|
+
A stage that declared no `fallback:` takes the path it always took, unchanged: one
|
|
956
|
+
wrapped call, one span named for the stage id. A stage that declared one hands
|
|
957
|
+
`weft_kernel.fallback.try_in_order` the candidates and lets it decide what
|
|
958
|
+
continues a chain — the runner itself has no opinion on `Outcome` types here, which
|
|
959
|
+
is what keeps the three-outcome rule in one testable place instead of two.
|
|
960
|
+
"""
|
|
961
|
+
if not stage.chain:
|
|
962
|
+
return await _wrapped_run(
|
|
963
|
+
stage.instance,
|
|
964
|
+
distribution=stage.distribution,
|
|
965
|
+
contract=stage.contract_name,
|
|
966
|
+
plugin=stage.plugin_name,
|
|
967
|
+
stage=stage.id,
|
|
968
|
+
position=stage.id,
|
|
969
|
+
)(payload, ctx)
|
|
970
|
+
return await try_in_order(
|
|
971
|
+
payload,
|
|
972
|
+
ctx,
|
|
973
|
+
stage=stage.id,
|
|
974
|
+
attempts=[_attempt(stage, candidate) for candidate in stage.chain],
|
|
975
|
+
)
|
|
976
|
+
|
|
977
|
+
async def _flush_all(
|
|
978
|
+
self, pipeline: RunnablePipeline, *, in_flight: BaseException | None = None
|
|
979
|
+
) -> None:
|
|
980
|
+
"""Flush every instance a run built, once, regardless of an earlier one's failure.
|
|
981
|
+
|
|
982
|
+
Every instance, not every resolved stage: a fallback is constructed at try time,
|
|
983
|
+
so a chain that reached one has an instance resolution never saw. Skipping it here
|
|
984
|
+
would lose whatever it buffered, silently — the exact shape of failure `flush`
|
|
985
|
+
being the runner's rather than a plugin's exists to make impossible. `_built_of`
|
|
986
|
+
is what asks the question, and a fallback that was never reached was never built
|
|
987
|
+
and so has nothing to answer for.
|
|
988
|
+
|
|
989
|
+
Each flush runs through `weft_kernel.seam.wrap_flush`, so a failure
|
|
990
|
+
carries the same span and pack/contract/plugin/stage attribution a
|
|
991
|
+
stage's `run()` gets. Every stage is tried — a completed run must
|
|
992
|
+
stay durable even when one stage's flush fails and another (normally
|
|
993
|
+
the store, last in the list) would otherwise never get the chance.
|
|
994
|
+
|
|
995
|
+
`in_flight` is the exception already propagating out of `run()`'s
|
|
996
|
+
batch loop, if any. When it is `None` (the ordinary end of a
|
|
997
|
+
successful run), flush failures are collected and raised together,
|
|
998
|
+
once, as a single `FlushError` whose `__cause__` is the first one
|
|
999
|
+
encountered — the caller has nothing else it is already about to
|
|
1000
|
+
raise, so `FlushError` is the loudest correct thing to raise. When
|
|
1001
|
+
`in_flight` is not `None`, raising here would replace it — Python
|
|
1002
|
+
lets an exception raised inside an `except` block's handling
|
|
1003
|
+
displace the one being handled — so a flush failure is instead
|
|
1004
|
+
attached to `in_flight` as a note and `in_flight` is left to
|
|
1005
|
+
propagate exactly as it arrived. This is what keeps a `CancelledError`
|
|
1006
|
+
a `CancelledError` even when the cleanup it triggers itself fails; see
|
|
1007
|
+
the module docstring, *"`CancelledError` propagates untouched."*
|
|
1008
|
+
"""
|
|
1009
|
+
failures: list[WeftError] = []
|
|
1010
|
+
for stage in pipeline.stages:
|
|
1011
|
+
for plugin_name, distribution, instance in _built_of(stage):
|
|
1012
|
+
flush = _flush_of(instance)
|
|
1013
|
+
if flush is None:
|
|
1014
|
+
continue
|
|
1015
|
+
wrapped_flush = wrap_flush(
|
|
1016
|
+
flush,
|
|
1017
|
+
distribution=distribution,
|
|
1018
|
+
contract=stage.contract_name,
|
|
1019
|
+
plugin=plugin_name,
|
|
1020
|
+
stage=stage.id,
|
|
1021
|
+
)
|
|
1022
|
+
try:
|
|
1023
|
+
await wrapped_flush()
|
|
1024
|
+
except WeftError as exc:
|
|
1025
|
+
failures.append(exc)
|
|
1026
|
+
|
|
1027
|
+
if not failures:
|
|
1028
|
+
return
|
|
1029
|
+
|
|
1030
|
+
failed_stages = ", ".join(f"'{exc.stage}'" for exc in failures)
|
|
1031
|
+
message = (
|
|
1032
|
+
f"{len(failures)} of {len(pipeline.stages)} stage(s) failed to flush: {failed_stages}."
|
|
1033
|
+
)
|
|
1034
|
+
if in_flight is not None:
|
|
1035
|
+
in_flight.add_note(f"FlushError: {message}")
|
|
1036
|
+
return
|
|
1037
|
+
raise FlushError(message) from failures[0]
|
|
1038
|
+
|
|
1039
|
+
|
|
1040
|
+
def _wrapped_run(
|
|
1041
|
+
instance: object,
|
|
1042
|
+
*,
|
|
1043
|
+
distribution: str,
|
|
1044
|
+
contract: str,
|
|
1045
|
+
plugin: str,
|
|
1046
|
+
stage: str,
|
|
1047
|
+
position: str,
|
|
1048
|
+
) -> Callable[[object, Context], Awaitable[Outcome[object]]]:
|
|
1049
|
+
"""`instance.run`, through the registration seam. The only place this module calls `wrap`.
|
|
1050
|
+
|
|
1051
|
+
`position` is `stage.id` and `stage` is the span's label — the same string on the ordinary
|
|
1052
|
+
path and deliberately different on the fallback path, where the label carries the backend
|
|
1053
|
+
that answered (`f"{stage.id}:{candidate.name}"`) so a trace says *which position, filled by
|
|
1054
|
+
whom*. What travels on a `TokenChunk` is the **position**, because a reader asking "is this
|
|
1055
|
+
the answer?" is asking about the pipeline, not about which candidate won. Carried repair
|
|
1056
|
+
**R10.1**.
|
|
1057
|
+
"""
|
|
1058
|
+
return wrap(
|
|
1059
|
+
cast("Stage[object, object]", instance).run,
|
|
1060
|
+
distribution=distribution,
|
|
1061
|
+
contract=contract,
|
|
1062
|
+
plugin=plugin,
|
|
1063
|
+
stage=stage,
|
|
1064
|
+
position=position,
|
|
1065
|
+
)
|
|
1066
|
+
|
|
1067
|
+
|
|
1068
|
+
def _attempt(stage: _ResolvedStage, candidate: _Candidate) -> Attempt[object, object]:
|
|
1069
|
+
"""One chain candidate as an `Attempt` — built the moment the chain reaches it, never before.
|
|
1070
|
+
|
|
1071
|
+
The span is named `f"{stage.id}:{candidate.name}"` rather than `stage.id`, which is
|
|
1072
|
+
what makes *which backend answered* readable off a trace with nothing else to consult:
|
|
1073
|
+
a chain of three produces three spans, and their names say which position was being
|
|
1074
|
+
filled and by whom.
|
|
1075
|
+
"""
|
|
1076
|
+
|
|
1077
|
+
async def _run(payload: object, ctx: Context) -> Outcome[object]:
|
|
1078
|
+
if candidate.instance is None:
|
|
1079
|
+
candidate.instance = _build(stage, candidate)
|
|
1080
|
+
return await _wrapped_run(
|
|
1081
|
+
candidate.instance,
|
|
1082
|
+
distribution=candidate.distribution,
|
|
1083
|
+
contract=stage.contract_name,
|
|
1084
|
+
plugin=candidate.name,
|
|
1085
|
+
stage=f"{stage.id}:{candidate.name}",
|
|
1086
|
+
position=stage.id,
|
|
1087
|
+
)(payload, ctx)
|
|
1088
|
+
|
|
1089
|
+
return Attempt(name=candidate.name, run=_run)
|
|
1090
|
+
|
|
1091
|
+
|
|
1092
|
+
def _build(stage: _ResolvedStage, candidate: _Candidate) -> object:
|
|
1093
|
+
"""Construct `candidate`, turning a construction failure into an attributed `WeftError`.
|
|
1094
|
+
|
|
1095
|
+
`weft_kernel.seam.wrap` covers a plugin's `run`, and construction happens outside it —
|
|
1096
|
+
so without this, a fallback whose factory raises escapes as a bare `TypeError` naming
|
|
1097
|
+
neither the pack nor the stage, and it does so only on the inputs the primary could not
|
|
1098
|
+
read. Raising `WeftError` instead puts it where `try_in_order` records it, which means
|
|
1099
|
+
a chain that ends with an unbuildable last candidate reports *"could not be built"*
|
|
1100
|
+
beside every other candidate's reason rather than replacing them with a traceback.
|
|
1101
|
+
|
|
1102
|
+
`except Exception` is the same deliberate breadth `seam.wrap` declares for the same
|
|
1103
|
+
reason: a factory is third-party code, nothing is swallowed, and `__cause__` keeps the
|
|
1104
|
+
traceback. `CancelledError` is a `BaseException` and passes through untouched.
|
|
1105
|
+
"""
|
|
1106
|
+
label = f"{stage.id}:{candidate.name}"
|
|
1107
|
+
try:
|
|
1108
|
+
# `fallback:` carries no `with:` block — see `StageSpec.fallback` — so a fallback
|
|
1109
|
+
# is built on its plugin's own defaults, the same call an unconfigured stage gets.
|
|
1110
|
+
return candidate.factory(None)
|
|
1111
|
+
except Exception as exc:
|
|
1112
|
+
raise WeftError(
|
|
1113
|
+
f"'{label}' could not be built: {exc}",
|
|
1114
|
+
pack=candidate.distribution,
|
|
1115
|
+
contract=stage.contract_name,
|
|
1116
|
+
plugin=candidate.name,
|
|
1117
|
+
stage=label,
|
|
1118
|
+
) from exc
|
|
1119
|
+
|
|
1120
|
+
|
|
1121
|
+
def _built_of(stage: _ResolvedStage) -> tuple[tuple[str, str, object], ...]:
|
|
1122
|
+
"""Every `(plugin_name, distribution, instance)` this stage actually built.
|
|
1123
|
+
|
|
1124
|
+
The primary always; a fallback only once a chain reached it, which is precisely when
|
|
1125
|
+
it can have buffered anything worth flushing. `stage.chain[1:]` rather than the whole
|
|
1126
|
+
chain because candidate zero *is* `stage.instance` — including it twice would flush
|
|
1127
|
+
the primary twice, and `flush` is documented idempotent rather than free.
|
|
1128
|
+
"""
|
|
1129
|
+
return (
|
|
1130
|
+
(stage.plugin_name, stage.distribution, stage.instance),
|
|
1131
|
+
*(
|
|
1132
|
+
(candidate.name, candidate.distribution, candidate.instance)
|
|
1133
|
+
for candidate in stage.chain[1:]
|
|
1134
|
+
if candidate.instance is not None
|
|
1135
|
+
),
|
|
1136
|
+
)
|
|
1137
|
+
|
|
1138
|
+
|
|
1139
|
+
def _check_tenant(pipeline: RunnablePipeline, ctx: Context, *, called: str) -> None:
|
|
1140
|
+
"""Refuse a `Context` for a tenant `pipeline` was not resolved for. One place, two callers.
|
|
1141
|
+
|
|
1142
|
+
Extracted from `run` when `run_once` arrived (task 2.4). The instance cache is keyed
|
|
1143
|
+
by tenant, so an entry point that forgot this check would silently reach across the
|
|
1144
|
+
boundary the key exists to draw — and the failure would look like a data leak rather
|
|
1145
|
+
than like a missing guard. `called` names the entry point in the message because
|
|
1146
|
+
which one was used is the first thing a reader wants and the traceback is the last
|
|
1147
|
+
place they should have to find it.
|
|
1148
|
+
"""
|
|
1149
|
+
if ctx.tenant_id != pipeline.tenant_id:
|
|
1150
|
+
raise TenantMismatchError(
|
|
1151
|
+
f"this pipeline was resolved for tenant '{pipeline.tenant_id}', but {called} "
|
|
1152
|
+
f"was given a Context for tenant '{ctx.tenant_id}'. An instance cached for "
|
|
1153
|
+
f"one tenant must never run for another."
|
|
1154
|
+
)
|
|
1155
|
+
|
|
1156
|
+
|
|
1157
|
+
def _stage_signature(contract: type[object]) -> tuple[object, object]:
|
|
1158
|
+
"""The `(In, Out)` `contract` declared via `Stage[In, Out]` as one of its own bases."""
|
|
1159
|
+
for base in getattr(contract, "__orig_bases__", ()):
|
|
1160
|
+
if typing.get_origin(base) is Stage:
|
|
1161
|
+
args = typing.get_args(base)
|
|
1162
|
+
if len(args) == 2: # noqa: PLR2004 - Stage is fixed at two type parameters
|
|
1163
|
+
return args[0], args[1]
|
|
1164
|
+
raise StageCompositionError(
|
|
1165
|
+
f"'{contract.__name__}' does not declare Stage[In, Out] as a base — every contract "
|
|
1166
|
+
f"used in a pipeline states what it consumes and produces, e.g. "
|
|
1167
|
+
f"class YourContract(Stage[list[In], list[Out]], Protocol).",
|
|
1168
|
+
remedy=(
|
|
1169
|
+
f"declare `class {contract.__name__}(Stage[In, Out], Protocol)` on the contract itself."
|
|
1170
|
+
),
|
|
1171
|
+
)
|
|
1172
|
+
|
|
1173
|
+
|
|
1174
|
+
def _type_name(value: object) -> str:
|
|
1175
|
+
"""A type as a reader recognises it — `QuerySet`, never `<class '...payload.QuerySet'>`.
|
|
1176
|
+
|
|
1177
|
+
`repr` on a class is the noisy form, and this message's whole job is to name two
|
|
1178
|
+
types clearly enough that somebody can see which document is wrong.
|
|
1179
|
+
`_stage_signature` may hand back something that is not a class at all, so the
|
|
1180
|
+
fallback is `repr` rather than an attribute access that raises inside an error path.
|
|
1181
|
+
"""
|
|
1182
|
+
return getattr(value, "__name__", None) or repr(value)
|
|
1183
|
+
|
|
1184
|
+
|
|
1185
|
+
def _check_composition(
|
|
1186
|
+
specs: Sequence[StageSpec], *, entry_type: type[object] | None = None
|
|
1187
|
+
) -> None:
|
|
1188
|
+
"""Every consecutive pair of `specs` composes: one stage's `Out` is the next stage's `In`.
|
|
1189
|
+
|
|
1190
|
+
`entry_type`, when given, is compared against the *first* spec's own expected `In` —
|
|
1191
|
+
the value `_stage_signature` already computes for it below, previously discarded
|
|
1192
|
+
because the `if previous is not None` guard skips the first iteration entirely. A
|
|
1193
|
+
caller that knows what it is about to hand the pipeline can name it here; a caller
|
|
1194
|
+
that does not (the default, `None`) is making no claim, and the first stage's own
|
|
1195
|
+
`In` goes unchecked exactly as it always has.
|
|
1196
|
+
"""
|
|
1197
|
+
previous: tuple[str, object] | None = None
|
|
1198
|
+
for spec in specs:
|
|
1199
|
+
payload_type, produced_type = _stage_signature(spec.contract)
|
|
1200
|
+
# `expected` is `entry_type` widened to `object` purely for the comparison:
|
|
1201
|
+
# `_stage_signature` returns the `In`/`Out` pair as `object`, so comparing a
|
|
1202
|
+
# `type[object]` against it directly reads to a type checker as two things that
|
|
1203
|
+
# can never be equal. The parameter stays `type[object]` because that is what a
|
|
1204
|
+
# caller actually has.
|
|
1205
|
+
expected: object = entry_type
|
|
1206
|
+
if previous is None and entry_type is not None and expected != payload_type:
|
|
1207
|
+
raise StageCompositionError(
|
|
1208
|
+
f"stage '{spec.id}' ({spec.contract.__name__}:{spec.name}) expects "
|
|
1209
|
+
f"{_type_name(payload_type)}, but this pipeline will be handed "
|
|
1210
|
+
f"{_type_name(entry_type)}.",
|
|
1211
|
+
stages=(spec.id,),
|
|
1212
|
+
remedy=(
|
|
1213
|
+
f"add a stage ahead of '{spec.id}' that produces {payload_type!r}, or "
|
|
1214
|
+
f"use a pipeline whose first stage expects {entry_type!r}."
|
|
1215
|
+
),
|
|
1216
|
+
)
|
|
1217
|
+
if previous is not None:
|
|
1218
|
+
previous_id, previous_produced = previous
|
|
1219
|
+
if previous_produced != payload_type:
|
|
1220
|
+
raise StageCompositionError(
|
|
1221
|
+
f"stage '{spec.id}' ({spec.contract.__name__}:{spec.name}) expects "
|
|
1222
|
+
f"{payload_type!r}, but the previous stage '{previous_id}' produces "
|
|
1223
|
+
f"{previous_produced!r}. Consecutive stages must compose by type.",
|
|
1224
|
+
stages=(previous_id, spec.id),
|
|
1225
|
+
remedy=(
|
|
1226
|
+
f"reorder this StageSpec list so '{previous_id}' precedes a stage "
|
|
1227
|
+
f"expecting {previous_produced!r}, or so '{spec.id}' follows one "
|
|
1228
|
+
f"producing {payload_type!r}."
|
|
1229
|
+
),
|
|
1230
|
+
)
|
|
1231
|
+
previous = (spec.id, produced_type)
|
|
1232
|
+
|
|
1233
|
+
|
|
1234
|
+
def _config_hash(config: object) -> str:
|
|
1235
|
+
"""A stable digest of `config`, for the instance cache key.
|
|
1236
|
+
|
|
1237
|
+
A content hash rather than `config` itself: the cache key must be
|
|
1238
|
+
hashable even when `config` is a Pydantic model carrying an unhashable
|
|
1239
|
+
field (a plain `dict`, say), which `frozen=True` alone does not fix.
|
|
1240
|
+
"""
|
|
1241
|
+
payload = config.model_dump_json() if isinstance(config, BaseModel) else repr(config)
|
|
1242
|
+
return hashlib.sha256(payload.encode("utf-8")).hexdigest()
|
|
1243
|
+
|
|
1244
|
+
|
|
1245
|
+
def _lifetime_of(instance: object) -> Lifetime:
|
|
1246
|
+
return cast(Lifetime, getattr(instance, "lifetime", Lifetime.RUN))
|
|
1247
|
+
|
|
1248
|
+
|
|
1249
|
+
def _requires_of(instance: object) -> tuple[type[ExtModel], ...]:
|
|
1250
|
+
return cast("tuple[type[ExtModel], ...]", getattr(instance, "requires", ()))
|
|
1251
|
+
|
|
1252
|
+
|
|
1253
|
+
def _provides_of(instance: object) -> tuple[type[ExtModel], ...]:
|
|
1254
|
+
return cast("tuple[type[ExtModel], ...]", getattr(instance, "provides", ()))
|
|
1255
|
+
|
|
1256
|
+
|
|
1257
|
+
def _applies_to_of(instance: object) -> tuple[Applies, ...]:
|
|
1258
|
+
"""`instance.applies_to`, read defensively — task 1.6, on `_requires_of`'s own footing.
|
|
1259
|
+
|
|
1260
|
+
An empty default means every node satisfies it vacuously — `02` §3: "A
|
|
1261
|
+
stage that declares no `applies_to` applies to everything." Nothing
|
|
1262
|
+
refuses a plugin that never mentions it, unlike `destroys`: applicability
|
|
1263
|
+
narrows a stage's own reach, so forgetting it costs nothing to a
|
|
1264
|
+
stranger's pipeline the way forgetting `destroys` would.
|
|
1265
|
+
"""
|
|
1266
|
+
return cast("tuple[Applies, ...]", getattr(instance, "applies_to", ()))
|
|
1267
|
+
|
|
1268
|
+
|
|
1269
|
+
def _is_routable(payload: object) -> bool:
|
|
1270
|
+
"""Whether `payload` is shaped like a batch of nodes `applies_to` could filter.
|
|
1271
|
+
|
|
1272
|
+
A repair, not part of the original task 1.6 lift: this used to read
|
|
1273
|
+
`isinstance(payload, Sequence) and not isinstance(payload, str | bytes)`,
|
|
1274
|
+
which checks only the *container*. A `Sequence[SourceDoc]` — an
|
|
1275
|
+
extractor's own payload type — satisfies that check exactly as well as a
|
|
1276
|
+
`Sequence[Node]` does, so a stage that declared `applies_to` over a
|
|
1277
|
+
non-node payload was not falling back to the unfiltered call this
|
|
1278
|
+
docstring promises; it was falling into `_segment_by_applicability`
|
|
1279
|
+
instead, where `isinstance(item, Node)` marks every one of its items
|
|
1280
|
+
non-matching. The whole payload came back unchanged — the stage's `run`
|
|
1281
|
+
was never invoked at all — which is silent by construction: nothing
|
|
1282
|
+
records that `applies_to` was ignored, and the first symptom shows up as
|
|
1283
|
+
an `AttributeError` inside whatever stage runs next.
|
|
1284
|
+
|
|
1285
|
+
So this now also confirms the *elements*, not only the container: a
|
|
1286
|
+
`Sequence`, excluding `str`/`bytes` (both satisfy `Sequence[object]`
|
|
1287
|
+
structurally, but iterating either one apart is never what a pipeline
|
|
1288
|
+
author meant by "a batch of nodes"), that is both non-empty and holds
|
|
1289
|
+
only `Node`s. An **empty** sequence is deliberately excluded too — with
|
|
1290
|
+
it included, a non-empty `applies_to` over zero items would segment into
|
|
1291
|
+
nothing and hand back `Produced(value=())` without ever calling the
|
|
1292
|
+
stage, the exact silently-empty `Produced` every contract docstring in
|
|
1293
|
+
this tree forbids; excluded, an empty batch instead takes the unfiltered
|
|
1294
|
+
path below and lets the stage answer for itself, the same
|
|
1295
|
+
`NothingToProduce` guard every shipped `Cleaner`/`Chunker` already opens
|
|
1296
|
+
its own `run` with. Anything else that fails this check (a `Query`, an
|
|
1297
|
+
`Answer`, a `Sequence[SourceDoc]`, an ordinary `str` a `str`-shaped stage
|
|
1298
|
+
happens to receive) is not a routable node-level payload, so `_run_stage`
|
|
1299
|
+
calls the stage once, unfiltered, exactly as an empty `applies_to`
|
|
1300
|
+
already does — applicability has nothing to route without nodes to
|
|
1301
|
+
route.
|
|
1302
|
+
"""
|
|
1303
|
+
if not isinstance(payload, Sequence) or isinstance(payload, str | bytes):
|
|
1304
|
+
return False
|
|
1305
|
+
items = cast("Sequence[object]", payload)
|
|
1306
|
+
return len(items) > 0 and all(isinstance(item, Node) for item in items)
|
|
1307
|
+
|
|
1308
|
+
|
|
1309
|
+
def _segment_by_applicability(
|
|
1310
|
+
items: Sequence[object], applies_to: tuple[Applies, ...]
|
|
1311
|
+
) -> list[tuple[bool, list[object]]]:
|
|
1312
|
+
"""`items` split into its maximal contiguous runs of "every `Applies` matches".
|
|
1313
|
+
|
|
1314
|
+
See the module docstring, *"Applicability is routed here"*, for why a
|
|
1315
|
+
run rather than one filtered call: a chunker's output is not shape-
|
|
1316
|
+
preserving, so the only way to know where an untouched node belongs once
|
|
1317
|
+
the matching nodes around it have been transformed is to never let them
|
|
1318
|
+
separate from their neighbours before the call happens. A node not
|
|
1319
|
+
matching *every* `Applies` in the tuple — the tuple is a conjunction,
|
|
1320
|
+
the same reading `requires` already gives a stage's own tuple of models
|
|
1321
|
+
— starts (or extends) a non-matching run instead; an item that is not
|
|
1322
|
+
even a `Node` (a payload `_is_routable` let through structurally, that
|
|
1323
|
+
nonetheless holds something with no `ext` to carry a fact at all) is
|
|
1324
|
+
treated the identical way — not matching, failing to the safe side
|
|
1325
|
+
exactly as an absent fact does.
|
|
1326
|
+
"""
|
|
1327
|
+
segments: list[tuple[bool, list[object]]] = []
|
|
1328
|
+
for item in items:
|
|
1329
|
+
matches = isinstance(item, Node) and all(applies.matches(item) for applies in applies_to)
|
|
1330
|
+
if segments and segments[-1][0] == matches:
|
|
1331
|
+
segments[-1][1].append(item)
|
|
1332
|
+
else:
|
|
1333
|
+
segments.append((matches, [item]))
|
|
1334
|
+
return segments
|
|
1335
|
+
|
|
1336
|
+
|
|
1337
|
+
@dataclass(frozen=True, slots=True)
|
|
1338
|
+
class _Declarations:
|
|
1339
|
+
"""The four facts one plugin states about its place in a pipeline, as sets.
|
|
1340
|
+
|
|
1341
|
+
Sets rather than the declared tuples because the only question asked of them is
|
|
1342
|
+
membership — see `_substitutability_differences` — and order carries no meaning in
|
|
1343
|
+
any of the four.
|
|
1344
|
+
"""
|
|
1345
|
+
|
|
1346
|
+
requires: frozenset[type[ExtModel]]
|
|
1347
|
+
provides: frozenset[type[ExtModel]]
|
|
1348
|
+
intact: frozenset[type[Property]]
|
|
1349
|
+
destroys: frozenset[type[Property]]
|
|
1350
|
+
|
|
1351
|
+
|
|
1352
|
+
def _declared_by(target: object) -> _Declarations:
|
|
1353
|
+
"""The four declarations, read off an unconstructed plugin class.
|
|
1354
|
+
|
|
1355
|
+
`target` is what `weft_kernel.registry.unwrap_factory` returns — the class itself,
|
|
1356
|
+
with any `functools.partial` binding pack settings peeled away, since `partial` does
|
|
1357
|
+
not proxy attribute access. Reading a class is what keeps a fallback unbuilt until
|
|
1358
|
+
the chain reaches it; it is also exactly what `weft_kernel.resolution.resolve` does
|
|
1359
|
+
for a pipeline document, so the two resolvers judge the same declarations.
|
|
1360
|
+
"""
|
|
1361
|
+
return _Declarations(
|
|
1362
|
+
requires=frozenset(_requires_of(target)),
|
|
1363
|
+
provides=frozenset(_provides_of(target)),
|
|
1364
|
+
intact=frozenset(_intact_of(target)),
|
|
1365
|
+
destroys=frozenset(_destroys_of(target)),
|
|
1366
|
+
)
|
|
1367
|
+
|
|
1368
|
+
|
|
1369
|
+
def _substitutability_differences(
|
|
1370
|
+
primary: _Declarations, fallback: _Declarations
|
|
1371
|
+
) -> tuple[str, ...]:
|
|
1372
|
+
"""Every way `fallback` demands more or promises less than `primary`. Empty means it fits.
|
|
1373
|
+
|
|
1374
|
+
One direction only, and that asymmetry is the point: a fallback providing an extra
|
|
1375
|
+
model or destroying one property fewer invalidates no check `Runner.resolve` already
|
|
1376
|
+
made against the primary, so it is not refused. See
|
|
1377
|
+
`FallbackNotSubstitutableError`.
|
|
1378
|
+
"""
|
|
1379
|
+
differences: list[str] = []
|
|
1380
|
+
if extra_requires := fallback.requires - primary.requires:
|
|
1381
|
+
differences.append(f"requires {_named(extra_requires)}, which the primary does not")
|
|
1382
|
+
if extra_intact := fallback.intact - primary.intact:
|
|
1383
|
+
differences.append(f"needs {_named(extra_intact)} intact, which the primary does not")
|
|
1384
|
+
if missing_provides := primary.provides - fallback.provides:
|
|
1385
|
+
differences.append(
|
|
1386
|
+
f"declares no {_named(missing_provides)}, which the primary provides to every "
|
|
1387
|
+
f"stage after it"
|
|
1388
|
+
)
|
|
1389
|
+
if extra_destroys := fallback.destroys - primary.destroys:
|
|
1390
|
+
differences.append(f"destroys {_named(extra_destroys)}, which the primary does not")
|
|
1391
|
+
return tuple(differences)
|
|
1392
|
+
|
|
1393
|
+
|
|
1394
|
+
def _named(types: frozenset[type[ExtModel]] | frozenset[type[Property]]) -> str:
|
|
1395
|
+
"""A stable, readable list of type names for a refusal's message."""
|
|
1396
|
+
return ", ".join(
|
|
1397
|
+
f"'{declared.__name__}'" for declared in sorted(types, key=lambda t: t.__name__)
|
|
1398
|
+
)
|
|
1399
|
+
|
|
1400
|
+
|
|
1401
|
+
def _intact_of(instance: object) -> tuple[type[Property], ...]:
|
|
1402
|
+
"""`instance.intact`, read defensively — an optional convention, never enforced at registration.
|
|
1403
|
+
|
|
1404
|
+
`02` §3: "`intact` stays an optional convention." Unlike `destroys`,
|
|
1405
|
+
nothing refuses a plugin that never mentions it — the cost of forgetting
|
|
1406
|
+
`intact` lands on the forgetful stage itself, the moment it needs a
|
|
1407
|
+
property some earlier stage already destroyed, so there is nothing here
|
|
1408
|
+
to guard against beyond the check `resolve` already performs with what
|
|
1409
|
+
this returns.
|
|
1410
|
+
"""
|
|
1411
|
+
return cast("tuple[type[Property], ...]", getattr(instance, "intact", ()))
|
|
1412
|
+
|
|
1413
|
+
|
|
1414
|
+
def _destroys_of(instance: object) -> tuple[type[Property], ...]:
|
|
1415
|
+
"""`instance.destroys`, read defensively — mandatory *at registration* for a governed contract.
|
|
1416
|
+
|
|
1417
|
+
Reading it the same defensive way `_requires_of`/`_provides_of` do is
|
|
1418
|
+
still correct here even though `weft_kernel.registry` already refused an
|
|
1419
|
+
ungoverned-contract plugin that omits `destroys`: an ungoverned
|
|
1420
|
+
contract's plugin is never required to declare it at all, and this
|
|
1421
|
+
function is the one place both cases have to be read uniformly.
|
|
1422
|
+
"""
|
|
1423
|
+
return cast("tuple[type[Property], ...]", getattr(instance, "destroys", ()))
|
|
1424
|
+
|
|
1425
|
+
|
|
1426
|
+
def _flush_of(instance: object) -> _FlushFn | None:
|
|
1427
|
+
"""`instance.flush`, if it has one and it is callable — never a bare `TypeError` later.
|
|
1428
|
+
|
|
1429
|
+
An attribute merely named `flush` that is not callable (a plain value, an
|
|
1430
|
+
author's typo) is treated the same as no `flush` at all: `flush` is
|
|
1431
|
+
documented as an optional method, not a required, checked-elsewhere one.
|
|
1432
|
+
"""
|
|
1433
|
+
found = getattr(instance, "flush", None)
|
|
1434
|
+
if found is None or not callable(found):
|
|
1435
|
+
return None
|
|
1436
|
+
return cast(_FlushFn, found)
|