weft-kernel 0.1.0__py3-none-any.whl
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- weft_kernel/__init__.py +178 -0
- weft_kernel/blocking.py +370 -0
- weft_kernel/context.py +250 -0
- weft_kernel/discovery.py +1181 -0
- weft_kernel/errors.py +73 -0
- weft_kernel/fallback.py +159 -0
- weft_kernel/payload/__init__.py +43 -0
- weft_kernel/payload/applicability.py +298 -0
- weft_kernel/payload/ext.py +172 -0
- weft_kernel/payload/ids.py +22 -0
- weft_kernel/payload/lineage.py +90 -0
- weft_kernel/payload/media_type.py +18 -0
- weft_kernel/payload/node.py +260 -0
- weft_kernel/payload/outcome.py +40 -0
- weft_kernel/payload/property.py +60 -0
- weft_kernel/payload/vector.py +28 -0
- weft_kernel/pipeline.py +748 -0
- weft_kernel/py.typed +0 -0
- weft_kernel/registry.py +692 -0
- weft_kernel/resolution.py +1582 -0
- weft_kernel/runner.py +1436 -0
- weft_kernel/seam.py +725 -0
- weft_kernel-0.1.0.dist-info/METADATA +88 -0
- weft_kernel-0.1.0.dist-info/RECORD +27 -0
- weft_kernel-0.1.0.dist-info/WHEEL +4 -0
- weft_kernel-0.1.0.dist-info/licenses/LICENSE +21 -0
- weft_kernel-0.1.0.dist-info/licenses/NOTICE +77 -0
weft_kernel/seam.py
ADDED
|
@@ -0,0 +1,725 @@
|
|
|
1
|
+
"""`wrap` — the registration seam every stage's execution passes through.
|
|
2
|
+
|
|
3
|
+
Specified in `docs/06-phase-0-build.md` step 3 and `docs/01-high-level-plan.md`
|
|
4
|
+
→ *Fitness functions*. Four cross-cutting concerns attach here, applied
|
|
5
|
+
without the author asking, because measurement shows what the alternative
|
|
6
|
+
costs: every concern applied automatically held perfectly, and every concern
|
|
7
|
+
left for an author to remember decayed — spans missing or off-convention at
|
|
8
|
+
38 of 54 hand-written call sites out of 58; observability lost entirely on an
|
|
9
|
+
untraced ingest stage. `wrap` is built *before* there is a `Stage` protocol,
|
|
10
|
+
a `Context`, or a single published capability contract (those are steps 4, 6
|
|
11
|
+
and 7) — deliberately, so nothing downstream ever has the chance to wrap a
|
|
12
|
+
call by hand. It is generic over what it wraps: a plain async callable and
|
|
13
|
+
three identifying strings (`distribution`, `contract`, `plugin`) are the
|
|
14
|
+
whole surface, supplied by whatever calls `wrap` — this module never chooses,
|
|
15
|
+
enumerates or hard-codes any of them, so it names no capability of its own
|
|
16
|
+
and assumes no pipeline concept.
|
|
17
|
+
|
|
18
|
+
The four concerns:
|
|
19
|
+
|
|
20
|
+
1. **Span wrapping**, via `opentelemetry-api`. `docs/02-extension-model.md`
|
|
21
|
+
→ *What a plugin receives*: "span name and `span_kind` are derived at the
|
|
22
|
+
registration seam from contract and plugin name" — never written by a
|
|
23
|
+
stage. Every span uses `SpanKind.INTERNAL`: the seam cannot know whether a
|
|
24
|
+
given contract is a store, a client, or neither, so a per-capability kind
|
|
25
|
+
would mean the kernel naming a capability. `SpanKind` beyond "internal
|
|
26
|
+
pipeline step" is a pack's own concern, added inside its stage if it wants
|
|
27
|
+
one.
|
|
28
|
+
2. **Error attribution**, naming the pack, contract, plugin and stage — see
|
|
29
|
+
`errors.py`. An exception a pack already raised as `WeftError` is not
|
|
30
|
+
replaced; only the four fields it left as `None` are filled in, because a
|
|
31
|
+
library re-raising with its own attribution knows more than the seam
|
|
32
|
+
does. Anything else escaping is wrapped fresh, `__cause__` preserved, so
|
|
33
|
+
no traceback is hidden. `CancelledError` is a `BaseException`, not an
|
|
34
|
+
`Exception` — it is never caught here, so it is never at risk of being
|
|
35
|
+
swallowed or rewrapped, by construction rather than by an added clause.
|
|
36
|
+
3. **`__transient__` stripping** — see `payload/ext.py` and
|
|
37
|
+
`payload/node.py`. A produced `Node`, or a list/tuple of them, has every
|
|
38
|
+
transient namespace stripped before the result leaves this function. This
|
|
39
|
+
is the type-level fact `Node.without_transient` exists to apply.
|
|
40
|
+
4. **The categorical blocking-call detector**, fitness function 7(b) — see
|
|
41
|
+
`blocking.py`. Scoped to exactly the `await` below, so a blocking call
|
|
42
|
+
made anywhere else — a fixture, an import, a factory building an instance
|
|
43
|
+
— is out of scope by construction, not by an exclusion list.
|
|
44
|
+
5. **The NUL-byte sanitiser** — see `_sanitize_control_bytes` below, riding
|
|
45
|
+
the same `Produced` → `Node` / `tuple` / `list` walk `_strip_transient`
|
|
46
|
+
already performs, immediately after it. `docs/build-ledger.md` → **2.34**
|
|
47
|
+
settles where this lives, against two alternatives, with evidence:
|
|
48
|
+
|
|
49
|
+
- **Not in an extractor pack.** Eight sites across `packages/` build a
|
|
50
|
+
`Node` from text that came from outside the process — `weft_extract/
|
|
51
|
+
text.py:80 'Node.syn'`, `weft_pdf/document.py:205 'rows: tu'`,
|
|
52
|
+
`weft_chunk/fixed_size.py:117 'destroys'
|
|
53
|
+
'destroys'`,
|
|
54
|
+
`weft_clean/dictionary_spacing.py:108-109 'intact: '`,
|
|
55
|
+
`weft_clean/hyphenation.py:71-72 'intact: '
|
|
56
|
+
'intact:'`,
|
|
57
|
+
`weft_clean/whitespace.py:64-65 'intact: t'`, `weft_clean/table_linearizer.py:79 'destroys:'`,
|
|
58
|
+
`weft_index/raptor.py:254 'owns retrying the same'`. A fix in one extractor covers two of the
|
|
59
|
+
eight — the same fragility shape at a smaller scale: a new construction
|
|
60
|
+
path that forgets the call reaches storage uncleaned.
|
|
61
|
+
- **Not in a store.** `weft_store/pgvector_store.py`'s `weft_nodes.content`
|
|
62
|
+
column is `TEXT NOT NULL` (`:138`) and Postgres refuses a NUL byte in it;
|
|
63
|
+
a second store backend sending the same payload over its own wire
|
|
64
|
+
protocol would not refuse it. Fixing this at a store means the same
|
|
65
|
+
corpus indexes under one backend and fails under another — the exact
|
|
66
|
+
shape this task line refuses ("refused by whichever backend happens to
|
|
67
|
+
notice").
|
|
68
|
+
- **The seam already owns this class of concern.** Spans, error
|
|
69
|
+
attribution and transient stripping all attach here rather than at a
|
|
70
|
+
rule an author must remember (`CLAUDE.md` → *The rules that are already
|
|
71
|
+
settled*), this function already knows what a `Node` is, and `wrap`'s
|
|
72
|
+
own signature already carries `distribution`, `contract` and `plugin` —
|
|
73
|
+
so a diagnostic triple — source, chunk index, extractor name — is
|
|
74
|
+
structural here, read off the span's own attributes, rather than four
|
|
75
|
+
keyword arguments an author has to remember to pass at every call site.
|
|
76
|
+
|
|
77
|
+
**NUL becomes a space, never a deletion**, because `weft_chunk.payload.ChunkOffset` records a
|
|
78
|
+
character offset into a parent's content, so deleting a byte would
|
|
79
|
+
silently shift every offset recorded downstream of the node being
|
|
80
|
+
cleaned. A space is one character for one character; every offset already
|
|
81
|
+
recorded against this content stays correct.
|
|
82
|
+
|
|
83
|
+
**Scope is `Node.content` and the `str`-typed fields of whatever
|
|
84
|
+
`ExtModel`s `Node.ext` carries** — `weft_store/pgvector_store.py`'s
|
|
85
|
+
`ext` column is `JSONB NOT NULL` (`:141`), and Postgres JSONB refuses a
|
|
86
|
+
NUL byte in a string value exactly as `TEXT` does, so an extension model
|
|
87
|
+
that ever carries verbatim extractor output is `content`'s twin, not a
|
|
88
|
+
narrower case. No current first-party `ExtModel` does — `weft_pdf.
|
|
89
|
+
PdfPages` (`weft_pdf/document.py:94-131 'ExtModel,'`) is the one built directly from
|
|
90
|
+
what a PDF backend reads, and its two fields are `backend: str` (a
|
|
91
|
+
plugin name, never extractor output) and `starts: tuple[int, ...]`
|
|
92
|
+
(offsets, not text) — so today's corpus exercises `content` only. The
|
|
93
|
+
walk still covers `ext` because the JSONB fact above is about the column,
|
|
94
|
+
not about any one model's current fields, and because covering it costs
|
|
95
|
+
one `isinstance` check per field on a namespace that already changed,
|
|
96
|
+
never a maintained list of which namespaces to check — the invariant is
|
|
97
|
+
"ext is safe to store," not the name of whichever field happens to hold
|
|
98
|
+
it.
|
|
99
|
+
6. **Never with a name list.** Every `ExtModel` namespace `Node.ext` carries
|
|
100
|
+
is walked by `type(model).model_fields`, the same field-introspection
|
|
101
|
+
idiom `weft_kernel.pipeline` and `weft_kernel.resolution` already use
|
|
102
|
+
elsewhere in this distribution — never a maintained tuple of which
|
|
103
|
+
fields are known to carry text.
|
|
104
|
+
|
|
105
|
+
**Where `distribution`, `contract` and `plugin` come from.** `wrap` does not
|
|
106
|
+
derive them: they are supplied by whatever calls it, exactly as
|
|
107
|
+
`Registry.add` takes `distribution` as a parameter rather than discovering it
|
|
108
|
+
(see `registry.py`). Discovery (step 5) is what will thread a pack's own
|
|
109
|
+
distribution name through; this module has no opinion on where that string
|
|
110
|
+
originates, only that it is attributed once it is known.
|
|
111
|
+
|
|
112
|
+
**`stage` is an optional parameter, defaulted for the caller that has no
|
|
113
|
+
pipeline concept.** `docs/06-phase-0-build.md` step 3 asks for "error
|
|
114
|
+
attribution naming the stage and the distribution." A pipeline *position* —
|
|
115
|
+
which slot in an ordered list a plugin fills — was not knowable at
|
|
116
|
+
registration, before step 6's runner existed to resolve a pipeline; `wrap`
|
|
117
|
+
therefore falls back to `f"{contract}:{plugin}"`, the one identifying label
|
|
118
|
+
available at registration time, whenever a caller does not supply `stage`
|
|
119
|
+
itself. **Step 6's runner is that caller**: it passes the resolved
|
|
120
|
+
`StageSpec.id`, so two positions in one pipeline that happen to name the same
|
|
121
|
+
plugin now produce distinguishable spans and error attribution — closing the
|
|
122
|
+
pipeline-position gap this paragraph used to describe as open.
|
|
123
|
+
|
|
124
|
+
**`wrap_flush`, below, is the same seam for a plugin's `flush()`.** `flush`
|
|
125
|
+
returns nothing an `Outcome` could decide, so it cannot share `wrap`'s
|
|
126
|
+
signature — there is nothing for `_strip_transient` to strip — but the other
|
|
127
|
+
three concerns still apply: a span, the blocking-call guard, and, for a bare
|
|
128
|
+
`WeftError` a pack raises itself, the same four-field attribution `wrap`
|
|
129
|
+
gives `run()`.
|
|
130
|
+
|
|
131
|
+
**`guard_blocking_calls`, added by `weft-cli` task 3.4 for a caller outside
|
|
132
|
+
`Runner`'s own reach.** `docs/build-ledger.md` 3.2 tried running a
|
|
133
|
+
`weft_command.contract.Command` invocation through this function unchanged
|
|
134
|
+
and reverted: `weft index`'s synchronous filesystem walk tripped concern 4,
|
|
135
|
+
which exists because a blocking `Stage` starves an event loop *other stages
|
|
136
|
+
share* — a `Command` is CLI orchestration invoked once per invocation or
|
|
137
|
+
REPL turn, with nothing else scheduled on that loop to starve, so the guard
|
|
138
|
+
was a false positive for it rather than a caught defect. 3.4's own analysis
|
|
139
|
+
(recorded in full in its `docs/build-ledger.md` entry) rejected two other
|
|
140
|
+
shapes for this fact: a second, hand-written span-and-attribution wrapper in
|
|
141
|
+
`weft_cli` (a second implementation of a concern this module already owns,
|
|
142
|
+
free to drift from it) and applying the guard unconditionally (the reverted
|
|
143
|
+
status quo, which left every `Command` invocation with no attribution at
|
|
144
|
+
all — a rule an author had to remember, which CLAUDE.md refuses). A keyword
|
|
145
|
+
defaulting to `True` keeps every existing caller's behaviour identical by
|
|
146
|
+
construction; only `weft_cli.cli.run_command` passes `False`, and only
|
|
147
|
+
concern 4 is affected — spans, attribution and `Produced` post-processing
|
|
148
|
+
run exactly as they do when the guard is armed.
|
|
149
|
+
|
|
150
|
+
**`Deprecation` and `warn_deprecated`, added at task 5.2e.** `docs/09-release.md` §3: "a
|
|
151
|
+
deprecated plugin, contract or config key is marked at registration, and the warning is
|
|
152
|
+
emitted by the registration wrapper — the same wrapper that applies spans, error
|
|
153
|
+
attribution and blocking-call detection." `weft_kernel.discovery.PackRegistrar.deprecate`
|
|
154
|
+
buffers a `Deprecation` per surface a pack names, exactly as `add_pipeline_resource`
|
|
155
|
+
buffers a `PipelineResource` — nothing is warned about until `register()` returns without
|
|
156
|
+
raising, so a pack that raises partway through warns about nothing it only half-marked.
|
|
157
|
+
`discovery._activate` is the one caller: once a pack's buffer has committed, it hands
|
|
158
|
+
whatever `PackRegistrar.deprecations` collected to `warn_deprecated` here, so a pack author
|
|
159
|
+
states the fact once, at registration, and never writes the warning by hand — the same
|
|
160
|
+
measured argument the module docstring opens with, applied to a fifth concern rather
|
|
161
|
+
than the original four. `docs/02-extension-model.md` §2's status vocabulary gains no member
|
|
162
|
+
for this: `weft_kernel.discovery.PackReport.deprecations` is read by `weft plugins doctor`
|
|
163
|
+
as a flag beside a pack's existing status, exactly as `ambient` already is one.
|
|
164
|
+
"""
|
|
165
|
+
|
|
166
|
+
import contextlib
|
|
167
|
+
import warnings
|
|
168
|
+
from collections.abc import Awaitable, Callable, Iterable, Sequence
|
|
169
|
+
from contextvars import ContextVar
|
|
170
|
+
from dataclasses import dataclass
|
|
171
|
+
from enum import Enum
|
|
172
|
+
from importlib import metadata
|
|
173
|
+
from typing import cast
|
|
174
|
+
|
|
175
|
+
from opentelemetry import trace
|
|
176
|
+
from opentelemetry.trace import SpanKind
|
|
177
|
+
from pydantic import ValidationError
|
|
178
|
+
|
|
179
|
+
from weft_kernel import blocking
|
|
180
|
+
from weft_kernel.errors import WeftError
|
|
181
|
+
from weft_kernel.payload import ExtModel, Node, Outcome, Produced
|
|
182
|
+
|
|
183
|
+
#: The span attribute a NUL count is recorded under — see `_sanitize_control_bytes`.
|
|
184
|
+
_NUL_BYTES_ATTRIBUTE = "weft.nul_bytes_removed"
|
|
185
|
+
|
|
186
|
+
_tracer = trace.get_tracer("weft_kernel")
|
|
187
|
+
|
|
188
|
+
#: The pipeline position currently executing, published by `wrap` for the length of each call.
|
|
189
|
+
#:
|
|
190
|
+
#: **Carried repair `R10.1`.** A `TokenChunk` carried a `role`, which names a *model mapping*,
|
|
191
|
+
#: and nothing finer reached a sink — so a stage making several concurrent calls on the
|
|
192
|
+
#: answering role interleaved its intermediate output with the answer, word by word, on a run
|
|
193
|
+
#: whose own configuration was correct. Whether a chunk is the answer is a fact about the
|
|
194
|
+
#: **stage**, and this is how that fact reaches code the seam wraps without every author having
|
|
195
|
+
#: to thread it: the concern attaches at the registration seam, which is `CLAUDE.md`'s rule and
|
|
196
|
+
#: the reason spans, error attribution and blocking detection all already live here.
|
|
197
|
+
#:
|
|
198
|
+
#: `""` outside any wrapped call, and readers are required to treat that as *unknown* rather
|
|
199
|
+
#: than as a stage name — `weft_llm.payload.TokenChunk.stage` carries the same convention.
|
|
200
|
+
_current_stage: ContextVar[str] = ContextVar("weft_current_stage", default="")
|
|
201
|
+
|
|
202
|
+
|
|
203
|
+
def current_stage() -> str:
|
|
204
|
+
"""The pipeline position executing on this task, or `""` outside any wrapped call."""
|
|
205
|
+
return _current_stage.get()
|
|
206
|
+
|
|
207
|
+
|
|
208
|
+
_FlushFn = Callable[[], Awaitable[None]]
|
|
209
|
+
|
|
210
|
+
|
|
211
|
+
class RemovalClock(Enum):
|
|
212
|
+
"""When a deprecated surface may be removed — the three honest answers, task **6.5**.
|
|
213
|
+
|
|
214
|
+
G9 settled the unit (`docs/09-release.md` §2.3, dependency 3): "**Releases, not months,
|
|
215
|
+
and the unit is one major of the publishing distribution.** A calendar window needs a
|
|
216
|
+
cadence promise this project does not make... A deprecated surface keeps working,
|
|
217
|
+
warning at registration, until its publisher's next major."
|
|
218
|
+
|
|
219
|
+
`UNPROMISED_BEFORE_1_0` is the member that matters and the one it would be easy not to
|
|
220
|
+
have. G9 also settled that "inside 0.x a contract may move without a deprecation period
|
|
221
|
+
but never silently", so a 0.x publisher's answer is **not** "removed in 1.0.0" — that
|
|
222
|
+
would promise a window 0.x explicitly reserves the right not to give. It is that there
|
|
223
|
+
is no window, said out loud, which is what makes the clock observable rather than
|
|
224
|
+
invented. Six distributions read `0.1.0` today (`09` §2.2), so this is the common case
|
|
225
|
+
rather than a corner.
|
|
226
|
+
"""
|
|
227
|
+
|
|
228
|
+
NEXT_MAJOR = "next-major"
|
|
229
|
+
UNPROMISED_BEFORE_1_0 = "unpromised-before-1.0"
|
|
230
|
+
VERSION_UNREADABLE = "version-unreadable"
|
|
231
|
+
|
|
232
|
+
|
|
233
|
+
@dataclass(frozen=True, slots=True)
|
|
234
|
+
class Removal:
|
|
235
|
+
"""The removal clock for one deprecated surface, **derived and never declared**.
|
|
236
|
+
|
|
237
|
+
Task **6.5**. A `removed_in` a pack author types is a number that goes stale on that
|
|
238
|
+
pack's next release with nothing to notice — `CLAUDE.md`'s measured rule: every concern
|
|
239
|
+
applied automatically held, and every concern an author had to remember decayed.
|
|
240
|
+
G9's unit makes the answer a pure function of the publishing distribution's own installed
|
|
241
|
+
version, so it is computed at the registration seam, once, and read by both consumers:
|
|
242
|
+
the `DeprecationWarning` below and `weft plugins doctor`.
|
|
243
|
+
|
|
244
|
+
`installed_version` is carried even when it could not be turned into a clock, because
|
|
245
|
+
"there is a version and it is not a number I can read" and "there is no version at all"
|
|
246
|
+
are different problems for whoever has to fix them.
|
|
247
|
+
"""
|
|
248
|
+
|
|
249
|
+
clock: RemovalClock
|
|
250
|
+
distribution: str
|
|
251
|
+
installed_version: str | None
|
|
252
|
+
release: str | None
|
|
253
|
+
|
|
254
|
+
def describe(self) -> str:
|
|
255
|
+
"""One phrase, owned here, so the warning and `doctor` cannot drift apart."""
|
|
256
|
+
if self.clock is RemovalClock.NEXT_MAJOR:
|
|
257
|
+
return f"removed in {self.release}"
|
|
258
|
+
if self.clock is RemovalClock.UNPROMISED_BEFORE_1_0:
|
|
259
|
+
return (
|
|
260
|
+
f"'{self.distribution}' is 0.x ({self.installed_version}), which promises no "
|
|
261
|
+
f"deprecation period — this surface may be removed in any release"
|
|
262
|
+
)
|
|
263
|
+
return (
|
|
264
|
+
f"removal release unknown — no readable version is recorded for '{self.distribution}'"
|
|
265
|
+
)
|
|
266
|
+
|
|
267
|
+
|
|
268
|
+
def removal_for(distribution: str, version_of: Callable[[str], str] = metadata.version) -> Removal:
|
|
269
|
+
"""G9's clock for `distribution`, read off its own installed version.
|
|
270
|
+
|
|
271
|
+
`version_of` is a parameter rather than a hard call so the derivation can be exercised
|
|
272
|
+
against every version shape without installing four distributions to do it; production
|
|
273
|
+
passes nothing and gets `importlib.metadata.version`.
|
|
274
|
+
|
|
275
|
+
The major is taken as the leading run of digits, so `2.0.0rc1` reads as major 2 —
|
|
276
|
+
`packaging` is not available here and never will be (G1 fixes this distribution's
|
|
277
|
+
dependencies at `pydantic` and `opentelemetry-api`), and a bare
|
|
278
|
+
`int(version.split(".")[0])` gets a pre-release wrong by raising.
|
|
279
|
+
"""
|
|
280
|
+
try:
|
|
281
|
+
version = version_of(distribution)
|
|
282
|
+
except Exception: # noqa: BLE001 — any lookup failure is the same answer to the caller
|
|
283
|
+
return Removal(
|
|
284
|
+
clock=RemovalClock.VERSION_UNREADABLE,
|
|
285
|
+
distribution=distribution,
|
|
286
|
+
installed_version=None,
|
|
287
|
+
release=None,
|
|
288
|
+
)
|
|
289
|
+
|
|
290
|
+
digits = ""
|
|
291
|
+
for character in version:
|
|
292
|
+
if not character.isdigit():
|
|
293
|
+
break
|
|
294
|
+
digits += character
|
|
295
|
+
|
|
296
|
+
if not digits:
|
|
297
|
+
return Removal(
|
|
298
|
+
clock=RemovalClock.VERSION_UNREADABLE,
|
|
299
|
+
distribution=distribution,
|
|
300
|
+
installed_version=version,
|
|
301
|
+
release=None,
|
|
302
|
+
)
|
|
303
|
+
|
|
304
|
+
major = int(digits)
|
|
305
|
+
if major == 0:
|
|
306
|
+
return Removal(
|
|
307
|
+
clock=RemovalClock.UNPROMISED_BEFORE_1_0,
|
|
308
|
+
distribution=distribution,
|
|
309
|
+
installed_version=version,
|
|
310
|
+
release=None,
|
|
311
|
+
)
|
|
312
|
+
|
|
313
|
+
return Removal(
|
|
314
|
+
clock=RemovalClock.NEXT_MAJOR,
|
|
315
|
+
distribution=distribution,
|
|
316
|
+
installed_version=version,
|
|
317
|
+
release=f"{distribution} {major + 1}.0.0",
|
|
318
|
+
)
|
|
319
|
+
|
|
320
|
+
|
|
321
|
+
@dataclass(frozen=True, slots=True)
|
|
322
|
+
class Unavailable:
|
|
323
|
+
"""One surface a pack declared it could not provide, and why — ledger task **6.29**.
|
|
324
|
+
|
|
325
|
+
`01` → *Fitness functions* 5: "Every capability a plugin declares must resolve to a live
|
|
326
|
+
implementation **at discovery time**, or the plugin must declare it unavailable and say why."
|
|
327
|
+
This is the second half. `PackStatus.PARTIAL` has been in `02` §2's status vocabulary since
|
|
328
|
+
Phase 0 and nothing could produce it — `weft_kernel.discovery`'s own docstring deferred the
|
|
329
|
+
mechanism to *"a later step's job"* and no later step took it — so a plugin that could not run
|
|
330
|
+
said so when a **run** failed rather than when `weft plugins doctor` asked.
|
|
331
|
+
|
|
332
|
+
`surface` is a short label the pack chooses, exactly as `Deprecation.surface` is: a plugin
|
|
333
|
+
name, a `"Contract:name"` pair, or the pack itself. The kernel names no capability and never
|
|
334
|
+
interprets it. `reason` is what a human reads, and it is required — an unavailability with no
|
|
335
|
+
reason is the silent drop this exists to end.
|
|
336
|
+
"""
|
|
337
|
+
|
|
338
|
+
distribution: str
|
|
339
|
+
surface: str
|
|
340
|
+
reason: str
|
|
341
|
+
|
|
342
|
+
|
|
343
|
+
@dataclass(frozen=True, slots=True)
|
|
344
|
+
class Deprecation:
|
|
345
|
+
"""One surface a pack marked deprecated at its own registration.
|
|
346
|
+
|
|
347
|
+
Task 5.2e. `surface` is a short label the pack author chooses — a
|
|
348
|
+
plugin name, a `"Contract:name"` pair, or the pack itself — never
|
|
349
|
+
interpreted by the kernel, which names no capability; `reason` is the
|
|
350
|
+
free text a human reads in `weft plugins doctor`. Buffered by
|
|
351
|
+
`weft_kernel.discovery.PackRegistrar.deprecate` exactly as
|
|
352
|
+
`PipelineResource` is buffered by `add_pipeline_resource`, for the
|
|
353
|
+
identical reason: a pack whose `register()` raises after calling this
|
|
354
|
+
must not leave a warning standing about a mark that never actually
|
|
355
|
+
committed.
|
|
356
|
+
|
|
357
|
+
`removal` — task **6.5**, `09` §3 — is when the surface may go, derived by
|
|
358
|
+
`removal_for` from the publishing distribution's own version rather than declared. It
|
|
359
|
+
is required rather than defaulted: a deprecation that does not say when it ends is the
|
|
360
|
+
thing this task exists to abolish, and a default would let one back in silently.
|
|
361
|
+
"""
|
|
362
|
+
|
|
363
|
+
distribution: str
|
|
364
|
+
surface: str
|
|
365
|
+
reason: str
|
|
366
|
+
removal: Removal
|
|
367
|
+
|
|
368
|
+
|
|
369
|
+
def warn_deprecated(deprecations: Iterable[Deprecation]) -> None:
|
|
370
|
+
"""Emit one `DeprecationWarning` per surface a pack marked deprecated at registration.
|
|
371
|
+
|
|
372
|
+
The registration wrapper this module's docstring names — see its own
|
|
373
|
+
paragraph on `Deprecation` above. Called by `weft_kernel.discovery._activate`,
|
|
374
|
+
once, after a pack's `register()` has returned without raising and every
|
|
375
|
+
buffered `PackRegistrar.deprecate` call is therefore known to have
|
|
376
|
+
committed; never called by a pack itself. Uses the stdlib `warnings`
|
|
377
|
+
machinery rather than `WeftError` or a printed line: a deprecation is
|
|
378
|
+
not a failure — the pack is still `ACTIVE` — and `warnings.warn` is the
|
|
379
|
+
one channel every Python tool already knows how to filter, capture or
|
|
380
|
+
promote to an error, without this module inventing a second one.
|
|
381
|
+
"""
|
|
382
|
+
for notice in deprecations:
|
|
383
|
+
warnings.warn(
|
|
384
|
+
f"'{notice.distribution}' marks '{notice.surface}' deprecated: {notice.reason}"
|
|
385
|
+
f" — {notice.removal.describe()}",
|
|
386
|
+
DeprecationWarning,
|
|
387
|
+
stacklevel=2,
|
|
388
|
+
)
|
|
389
|
+
|
|
390
|
+
|
|
391
|
+
def wrap[**P, T](
|
|
392
|
+
run: Callable[P, Awaitable[Outcome[T]]],
|
|
393
|
+
*,
|
|
394
|
+
distribution: str,
|
|
395
|
+
contract: str,
|
|
396
|
+
plugin: str,
|
|
397
|
+
stage: str | None = None,
|
|
398
|
+
position: str | None = None,
|
|
399
|
+
guard_blocking_calls: bool = True,
|
|
400
|
+
) -> Callable[P, Awaitable[Outcome[T]]]:
|
|
401
|
+
"""Wrap `run` so every call through it carries spans, attribution, stripping and the guard.
|
|
402
|
+
|
|
403
|
+
`run` is any async callable returning an `Outcome` — a stage's `run`
|
|
404
|
+
method, bound, is the intended shape, but this function never asserts
|
|
405
|
+
that: it calls what it is given and reacts only to what comes back.
|
|
406
|
+
`position` is the pipeline position this call *is*, and only `weft_kernel.runner` passes
|
|
407
|
+
one — see the comment on `_current_stage` above for why it is separate from `stage`.
|
|
408
|
+
|
|
409
|
+
`stage` names the pipeline position this call fills, when the caller has
|
|
410
|
+
one to give — the runner (`06` step 6) always does. A caller with no
|
|
411
|
+
pipeline concept (registration, step 3) may omit it and falls back to
|
|
412
|
+
`f"{contract}:{plugin}"`, the label available at registration time.
|
|
413
|
+
|
|
414
|
+
`guard_blocking_calls` defaults to `True` for every existing caller — see
|
|
415
|
+
the module docstring's own paragraph on why and who passes `False`.
|
|
416
|
+
"""
|
|
417
|
+
stage_label = stage if stage is not None else f"{contract}:{plugin}"
|
|
418
|
+
|
|
419
|
+
async def _wrapped(*args: P.args, **kwargs: P.kwargs) -> Outcome[T]:
|
|
420
|
+
# Carried repair **R10.1**. The stage this call is running, published for the length
|
|
421
|
+
# of the call so anything underneath can ask which pipeline position it is inside —
|
|
422
|
+
# `weft_llm.client.LLMClient.complete` stamps it onto every `TokenChunk`, so a sink
|
|
423
|
+
# can tell the answer from an intermediate call. Set here rather than passed down
|
|
424
|
+
# because that is `CLAUDE.md`'s own rule: a cross-cutting concern attaches at the
|
|
425
|
+
# registration seam, never in something an author has to remember. `reset` on the
|
|
426
|
+
# token rather than to a constant, so nested stages restore their parent's value.
|
|
427
|
+
# **`position`, never `stage_label`, and that distinction is `R10.1`'s second defect.**
|
|
428
|
+
# `wrap` is called for services and providers as well as stages, and those calls nest
|
|
429
|
+
# *inside* a stage: a stage asks an `LLM`, which asks a provider, each wrapped in turn.
|
|
430
|
+
# Keying on `stage_label` — or even on "was `stage` passed" — meant the innermost call
|
|
431
|
+
# won, and `weft_llm.client` wraps its own with `stage=f"llm:{role}"`, so every
|
|
432
|
+
# `TokenChunk` was stamped `llm:generate` rather than with the pipeline position that
|
|
433
|
+
# asked. Measured from the shipped binary: the sink's filter then matched nothing and
|
|
434
|
+
# the repair was inert. Only `weft_kernel.runner` passes `position`, and a call without
|
|
435
|
+
# one leaves its caller's in place — which is what makes this mean *which pipeline
|
|
436
|
+
# position am I inside*, rather than *what is the nearest wrapped call*.
|
|
437
|
+
token = _current_stage.set(position) if position is not None else None
|
|
438
|
+
try:
|
|
439
|
+
with _tracer.start_as_current_span(stage_label, kind=SpanKind.INTERNAL) as span:
|
|
440
|
+
span.set_attribute("weft.pack", distribution)
|
|
441
|
+
span.set_attribute("weft.contract", contract)
|
|
442
|
+
span.set_attribute("weft.plugin", plugin)
|
|
443
|
+
# A fresh context manager per call, never hoisted above `_wrapped`: a
|
|
444
|
+
# `@contextmanager`-built one (`blocking.guard`) can only be entered once, and
|
|
445
|
+
# `_wrapped` itself is reusable across many invocations.
|
|
446
|
+
guard_cm = (
|
|
447
|
+
blocking.guard(stage_label)
|
|
448
|
+
if guard_blocking_calls
|
|
449
|
+
else contextlib.nullcontext()
|
|
450
|
+
)
|
|
451
|
+
with guard_cm:
|
|
452
|
+
try:
|
|
453
|
+
outcome = await run(*args, **kwargs)
|
|
454
|
+
except WeftError as exc:
|
|
455
|
+
_attribute(
|
|
456
|
+
exc,
|
|
457
|
+
distribution=distribution,
|
|
458
|
+
contract=contract,
|
|
459
|
+
plugin=plugin,
|
|
460
|
+
stage=stage_label,
|
|
461
|
+
)
|
|
462
|
+
raise
|
|
463
|
+
except Exception as exc:
|
|
464
|
+
raise WeftError(
|
|
465
|
+
f"'{stage_label}' failed: {exc}",
|
|
466
|
+
pack=distribution,
|
|
467
|
+
contract=contract,
|
|
468
|
+
plugin=plugin,
|
|
469
|
+
stage=stage_label,
|
|
470
|
+
) from exc
|
|
471
|
+
outcome, nul_count = _sanitize_control_bytes(_strip_transient(outcome))
|
|
472
|
+
span.set_attribute(_NUL_BYTES_ATTRIBUTE, nul_count)
|
|
473
|
+
return outcome
|
|
474
|
+
|
|
475
|
+
finally:
|
|
476
|
+
if token is not None:
|
|
477
|
+
_current_stage.reset(token)
|
|
478
|
+
|
|
479
|
+
return _wrapped
|
|
480
|
+
|
|
481
|
+
|
|
482
|
+
def wrap_flush(
|
|
483
|
+
flush: _FlushFn,
|
|
484
|
+
*,
|
|
485
|
+
distribution: str,
|
|
486
|
+
contract: str,
|
|
487
|
+
plugin: str,
|
|
488
|
+
stage: str,
|
|
489
|
+
) -> _FlushFn:
|
|
490
|
+
"""Wrap a plugin's `flush()` with the same span, guard and attribution `wrap` gives `run()`.
|
|
491
|
+
|
|
492
|
+
No `_strip_transient` — `flush` returns nothing an `Outcome` could
|
|
493
|
+
decide, so there is nothing to strip. `stage` is required, not optional,
|
|
494
|
+
because the only caller (`06` step 6's `Runner._flush_all`) always has a
|
|
495
|
+
resolved `StageSpec.id` to give it.
|
|
496
|
+
"""
|
|
497
|
+
|
|
498
|
+
async def _wrapped() -> None:
|
|
499
|
+
with _tracer.start_as_current_span(f"{stage}:flush", kind=SpanKind.INTERNAL) as span:
|
|
500
|
+
span.set_attribute("weft.pack", distribution)
|
|
501
|
+
span.set_attribute("weft.contract", contract)
|
|
502
|
+
span.set_attribute("weft.plugin", plugin)
|
|
503
|
+
with blocking.guard(f"{stage}:flush"):
|
|
504
|
+
try:
|
|
505
|
+
await flush()
|
|
506
|
+
except WeftError as exc:
|
|
507
|
+
_attribute(
|
|
508
|
+
exc,
|
|
509
|
+
distribution=distribution,
|
|
510
|
+
contract=contract,
|
|
511
|
+
plugin=plugin,
|
|
512
|
+
stage=stage,
|
|
513
|
+
)
|
|
514
|
+
raise
|
|
515
|
+
except Exception as exc:
|
|
516
|
+
raise WeftError(
|
|
517
|
+
f"'{stage}' flush failed: {exc}",
|
|
518
|
+
pack=distribution,
|
|
519
|
+
contract=contract,
|
|
520
|
+
plugin=plugin,
|
|
521
|
+
stage=stage,
|
|
522
|
+
) from exc
|
|
523
|
+
|
|
524
|
+
return _wrapped
|
|
525
|
+
|
|
526
|
+
|
|
527
|
+
def _attribute(
|
|
528
|
+
exc: WeftError, *, distribution: str, contract: str, plugin: str, stage: str
|
|
529
|
+
) -> None:
|
|
530
|
+
"""Fill in whichever of `exc`'s four attribution fields a pack's own raise left `None`.
|
|
531
|
+
|
|
532
|
+
`errors.py`: a plain `WeftError` a pack raises directly has no reason to
|
|
533
|
+
know its own attribution. A field the pack *did* set — a library
|
|
534
|
+
re-raising with its own `plugin=`, say — is left exactly as it was.
|
|
535
|
+
"""
|
|
536
|
+
if exc.pack is None:
|
|
537
|
+
exc.pack = distribution
|
|
538
|
+
if exc.contract is None:
|
|
539
|
+
exc.contract = contract
|
|
540
|
+
if exc.plugin is None:
|
|
541
|
+
exc.plugin = plugin
|
|
542
|
+
if exc.stage is None:
|
|
543
|
+
exc.stage = stage
|
|
544
|
+
|
|
545
|
+
|
|
546
|
+
def _strip_transient[T](outcome: Outcome[T]) -> Outcome[T]:
|
|
547
|
+
"""Strip every `__transient__` namespace from a produced `Node`, or a list/tuple of them.
|
|
548
|
+
|
|
549
|
+
Anything else a stage produces — a scalar, a future `Answer` — passes
|
|
550
|
+
through untouched: transience is a fact about `Node.ext`, so this looks
|
|
551
|
+
for `Node` and nothing wider. A list or tuple is walked item by item
|
|
552
|
+
rather than gated on every item being a `Node`: a container mixing
|
|
553
|
+
`Node`s with other values must not carry a `Node`'s transients through
|
|
554
|
+
just because it as a whole failed an `all(...)` check — that would be a
|
|
555
|
+
success path and a do-nothing path indistinguishable to the caller, the
|
|
556
|
+
exact silent-fallback shape `payload/outcome.py` exists to avoid. `list`
|
|
557
|
+
is walked alongside `tuple`, not just `tuple`, because a `Produced[list[Node]]`
|
|
558
|
+
— the natural shape for a stage such as a chunker that emits many chunks
|
|
559
|
+
— is exactly as much "a fact about `Node.ext`, in a different container"
|
|
560
|
+
as a tuple is; nothing in Phase 0 restricts a stage's return shape to
|
|
561
|
+
forbid it. The container type is preserved: a `list` in, a `list` out.
|
|
562
|
+
"""
|
|
563
|
+
if not isinstance(outcome, Produced):
|
|
564
|
+
return outcome
|
|
565
|
+
|
|
566
|
+
value = outcome.value
|
|
567
|
+
if isinstance(value, Node):
|
|
568
|
+
return Produced(value=cast(T, value.without_transient()))
|
|
569
|
+
if isinstance(value, tuple):
|
|
570
|
+
items = cast("tuple[object, ...]", value)
|
|
571
|
+
stripped_tuple = tuple(
|
|
572
|
+
item.without_transient() if isinstance(item, Node) else item for item in items
|
|
573
|
+
)
|
|
574
|
+
return Produced(value=cast(T, stripped_tuple))
|
|
575
|
+
if isinstance(value, list):
|
|
576
|
+
entries = cast("list[object]", value)
|
|
577
|
+
stripped_list = [
|
|
578
|
+
item.without_transient() if isinstance(item, Node) else item for item in entries
|
|
579
|
+
]
|
|
580
|
+
return Produced(value=cast(T, stripped_list))
|
|
581
|
+
return outcome
|
|
582
|
+
|
|
583
|
+
|
|
584
|
+
def _sanitize_control_bytes[T](outcome: Outcome[T]) -> tuple[Outcome[T], int]:
|
|
585
|
+
"""Replace every NUL byte a produced `Node`'s `content` or ext `str` fields carry.
|
|
586
|
+
|
|
587
|
+
Returns the (possibly unchanged) outcome and how many bytes were found.
|
|
588
|
+
|
|
589
|
+
Same shape as `_strip_transient` immediately above — `Produced` → `Node`
|
|
590
|
+
/ `tuple` / `list`, anything else untouched — because it walks the exact
|
|
591
|
+
same values, one step later. See the module docstring, concern 5, for
|
|
592
|
+
why this lives here rather than in an extractor or a store, why the
|
|
593
|
+
replacement is a space rather than a deletion, and why `ext` is in scope
|
|
594
|
+
beside `content`. A container or a `Node` that needed no change is
|
|
595
|
+
returned as the same object the caller passed in — a corpus with no NUL
|
|
596
|
+
bytes, the overwhelming case, costs one `"\\x00" in s` scan per string and
|
|
597
|
+
not one rebuild.
|
|
598
|
+
"""
|
|
599
|
+
if not isinstance(outcome, Produced):
|
|
600
|
+
return outcome, 0
|
|
601
|
+
|
|
602
|
+
value = outcome.value
|
|
603
|
+
if isinstance(value, Node):
|
|
604
|
+
cleaned, count = _sanitize_node(value)
|
|
605
|
+
return (outcome if count == 0 else Produced(value=cast(T, cleaned))), count
|
|
606
|
+
if isinstance(value, tuple):
|
|
607
|
+
items = cast("tuple[object, ...]", value)
|
|
608
|
+
cleaned_items, count = _sanitize_items(items)
|
|
609
|
+
return (outcome if count == 0 else Produced(value=cast(T, tuple(cleaned_items)))), count
|
|
610
|
+
if isinstance(value, list):
|
|
611
|
+
entries = cast("list[object]", value)
|
|
612
|
+
cleaned_items, count = _sanitize_items(entries)
|
|
613
|
+
return (outcome if count == 0 else Produced(value=cast(T, cleaned_items))), count
|
|
614
|
+
return outcome, 0
|
|
615
|
+
|
|
616
|
+
|
|
617
|
+
def _sanitize_items(items: Sequence[object]) -> tuple[list[object], int]:
|
|
618
|
+
"""`_sanitize_node` applied to every `Node` in `items`; anything else passes through."""
|
|
619
|
+
total = 0
|
|
620
|
+
cleaned: list[object] = []
|
|
621
|
+
for item in items:
|
|
622
|
+
if isinstance(item, Node):
|
|
623
|
+
node, count = _sanitize_node(item)
|
|
624
|
+
cleaned.append(node)
|
|
625
|
+
total += count
|
|
626
|
+
else:
|
|
627
|
+
cleaned.append(item)
|
|
628
|
+
return cleaned, total
|
|
629
|
+
|
|
630
|
+
|
|
631
|
+
def _sanitize_node(node: Node) -> tuple[Node, int]:
|
|
632
|
+
"""`node` with every NUL in `content` and every ext model's `str` fields turned to a space.
|
|
633
|
+
|
|
634
|
+
Rebuilds through `Node._replace` — the same revalidating path
|
|
635
|
+
`without_transient` uses — only when something actually changed, and
|
|
636
|
+
only the fields that did: a node with a clean `content` but a dirty ext
|
|
637
|
+
field does not get its content re-validated for nothing, and the reverse.
|
|
638
|
+
|
|
639
|
+
The walk is `node.ext` specifically — the same "`Node` and nothing wider"
|
|
640
|
+
reach `_strip_transient` already states above. `weft_retrieve.payload`'s
|
|
641
|
+
`QuerySet` and `Candidates` carry their own `ext: ExtMap` and are out of
|
|
642
|
+
scope here, unchanged from that pre-existing rule.
|
|
643
|
+
"""
|
|
644
|
+
cleaned_content, content_count = _clean_str(node.content)
|
|
645
|
+
|
|
646
|
+
ext_updates: dict[str, ExtModel] = {}
|
|
647
|
+
ext_count = 0
|
|
648
|
+
for namespace, model in node.ext.items():
|
|
649
|
+
cleaned_model, model_count = _sanitize_ext_model(model)
|
|
650
|
+
if model_count:
|
|
651
|
+
ext_updates[namespace] = cleaned_model
|
|
652
|
+
ext_count += model_count
|
|
653
|
+
|
|
654
|
+
total = content_count + ext_count
|
|
655
|
+
if total == 0:
|
|
656
|
+
return node, 0
|
|
657
|
+
|
|
658
|
+
updates: dict[str, object] = {}
|
|
659
|
+
if content_count:
|
|
660
|
+
updates["content"] = cleaned_content
|
|
661
|
+
if ext_updates:
|
|
662
|
+
updates["ext"] = {**node.ext, **ext_updates}
|
|
663
|
+
# `Node._replace` is kernel-internal, not a third-party API — this module is the one
|
|
664
|
+
# other place in the same distribution `node.py` names as the sanctioned revalidating
|
|
665
|
+
# rebuild path, alongside `with_ext`/`without_transient`, which are both defined on
|
|
666
|
+
# `Node` itself and so need no such suppression.
|
|
667
|
+
return node._replace(**updates), total # pyright: ignore[reportPrivateUsage]
|
|
668
|
+
|
|
669
|
+
|
|
670
|
+
def _sanitize_ext_model(model: ExtModel) -> tuple[ExtModel, int]:
|
|
671
|
+
"""`model` with every `str`-typed field's NUL bytes turned to a space, by field introspection.
|
|
672
|
+
|
|
673
|
+
`type(model).model_fields` is walked rather than a maintained list of
|
|
674
|
+
which fields carry text — the module docstring's concern 6 — so a pack's
|
|
675
|
+
own `ExtModel` subclass needs no registration here to be covered; it is
|
|
676
|
+
covered the moment it declares a `str` field.
|
|
677
|
+
|
|
678
|
+
Scoped to a field whose *runtime value* is a `str` — not `tuple[str, ...]`
|
|
679
|
+
or `list[str]`, checked by `isinstance` rather than by the field's
|
|
680
|
+
annotation. `docs/build-ledger.md` → 2.34 scopes this task to "`Node.content`
|
|
681
|
+
and the `str`-typed fields of the `ExtModel`s in `Node.ext`" literally; no
|
|
682
|
+
shipped `ExtModel` as of this task holds a string collection built from
|
|
683
|
+
verbatim extractor text, so widening to collections has no motivating case
|
|
684
|
+
yet and is left for whichever future field needs it, named at that point
|
|
685
|
+
rather than guessed at here.
|
|
686
|
+
|
|
687
|
+
Rebuilt through `model_validate`, never `model_copy(update=...)` — the same
|
|
688
|
+
choice `node.py`'s `Node._replace` makes, and for its own stated reason:
|
|
689
|
+
`model_copy(update=...)` skips validation entirely, so it would let this
|
|
690
|
+
substitution smuggle a NUL-cleaned value past a pack's own `Field` or
|
|
691
|
+
`field_validator` guard. Unlike `_replace`, a validation failure here is
|
|
692
|
+
not the caller's mistake to see raw: nothing upstream asked for this
|
|
693
|
+
rebuild, so `ValidationError` is caught and re-raised naming the model,
|
|
694
|
+
the cause, and carrying pydantic's own field-level detail, rather than
|
|
695
|
+
surfacing as an unexplained validation error out of a sanitiser no caller
|
|
696
|
+
invoked directly.
|
|
697
|
+
"""
|
|
698
|
+
updates: dict[str, str] = {}
|
|
699
|
+
total = 0
|
|
700
|
+
for field_name in type(model).model_fields:
|
|
701
|
+
if not isinstance(value := getattr(model, field_name), str):
|
|
702
|
+
continue
|
|
703
|
+
cleaned, count = _clean_str(value)
|
|
704
|
+
if count:
|
|
705
|
+
updates[field_name] = cleaned
|
|
706
|
+
total += count
|
|
707
|
+
if total == 0:
|
|
708
|
+
return model, 0
|
|
709
|
+
try:
|
|
710
|
+
return type(model).model_validate({**model.__dict__, **updates}), total
|
|
711
|
+
except ValidationError as exc:
|
|
712
|
+
raise WeftError(f"NUL sanitisation left {type(model).__name__} invalid: {exc}") from exc
|
|
713
|
+
|
|
714
|
+
|
|
715
|
+
def _clean_str(value: str) -> tuple[str, int]:
|
|
716
|
+
"""`value` with every NUL byte replaced by a space, and how many there were.
|
|
717
|
+
|
|
718
|
+
The early return is the cost discipline the module docstring promises: a
|
|
719
|
+
clean string — the overwhelming majority — costs one C-level containment
|
|
720
|
+
check and nothing else, never a `.replace` call that would return an
|
|
721
|
+
identical string.
|
|
722
|
+
"""
|
|
723
|
+
if "\x00" not in value:
|
|
724
|
+
return value, 0
|
|
725
|
+
return value.replace("\x00", " "), value.count("\x00")
|