@xlaunch/llm 0.2.0-beta.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +22 -0
- package/README.md +174 -0
- package/lib/index.js +2300 -0
- package/lib/invariant.js +84 -0
- package/lib/typert.host.d.ts +3 -0
- package/lib/typert.host.js +547 -0
- package/lib/typert.remote-client.d.ts +26 -0
- package/lib/typert.remote-client.js +102 -0
- package/lib/types/adapter-failure.d.ts +14 -0
- package/lib/types/adapter-failure.js +105 -0
- package/lib/types/api-key.d.ts +28 -0
- package/lib/types/api-key.js +34 -0
- package/lib/types/assembler.d.ts +75 -0
- package/lib/types/assembler.js +191 -0
- package/lib/types/assistant-stream.d.ts +166 -0
- package/lib/types/assistant-stream.js +458 -0
- package/lib/types/attribution.d.ts +47 -0
- package/lib/types/attribution.js +46 -0
- package/lib/types/brand.d.ts +56 -0
- package/lib/types/brand.js +53 -0
- package/lib/types/call-config.d.ts +53 -0
- package/lib/types/call-config.js +46 -0
- package/lib/types/content.d.ts +130 -0
- package/lib/types/content.js +284 -0
- package/lib/types/error.d.ts +73 -0
- package/lib/types/error.js +145 -0
- package/lib/types/index.d.ts +408 -0
- package/lib/types/index.js +920 -0
- package/lib/types/invariant.d.ts +13 -0
- package/lib/types/invariant.js +100 -0
- package/lib/types/message.d.ts +197 -0
- package/lib/types/message.js +82 -0
- package/lib/types/retry-policy.d.ts +66 -0
- package/lib/types/retry-policy.js +127 -0
- package/lib/types/types.d.ts +430 -0
- package/lib/types/types.js +7 -0
- package/package.json +81 -0
package/LICENSE
ADDED
|
@@ -0,0 +1,22 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Portions Copyright (c) 2026 Northlatch Labs LLC
|
|
4
|
+
Copyright (c) 2026 DeepSeek
|
|
5
|
+
|
|
6
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
7
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
8
|
+
in the Software without restriction, including without limitation the rights
|
|
9
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
10
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
11
|
+
furnished to do so, subject to the following conditions:
|
|
12
|
+
|
|
13
|
+
The above copyright notice and this permission notice shall be included in all
|
|
14
|
+
copies or substantial portions of the Software.
|
|
15
|
+
|
|
16
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
17
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
18
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
19
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
20
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
21
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
22
|
+
SOFTWARE.
|
package/README.md
ADDED
|
@@ -0,0 +1,174 @@
|
|
|
1
|
+
---
|
|
2
|
+
description: "The provider-neutral model-call service for users and maintainers streaming requests, registering provider adapters, or resolving model metadata."
|
|
3
|
+
kind: "package-reference"
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# @xlaunch/llm
|
|
7
|
+
|
|
8
|
+
## Summary
|
|
9
|
+
|
|
10
|
+
`@xlaunch/llm` is the provider-neutral model-call service at the center of the harness's LLM capability. Every composition that streams a request to a model provider goes through it, and it owns the shared vocabulary — messages, content blocks, raw stream chunks, and compact Assistant stream records — that the agent loop, session log, and every plugin speak. With it you can register provider adapters, stream one model call, list and discover models, resolve exact-model metadata and call defaults, and capture each provider's retry policy; every request is logged so it stays reconstructable from the session log. It executes no retries and owns no provider wire logic: adapters translate their provider's format, and the optional `xlaunch-llm-retry` package re-runs failed requests at durable step boundaries. Requests are deep-frozen before dispatch, so middleware and adapters can read them but never rewrite them.
|
|
11
|
+
|
|
12
|
+
## Table of Contents
|
|
13
|
+
|
|
14
|
+
- [Use this package](#use-this-package)
|
|
15
|
+
- [Understand the implementation](#understand-the-implementation)
|
|
16
|
+
- [Further Exploration](#further-exploration)
|
|
17
|
+
- [Model Experience](#model-experience)
|
|
18
|
+
- [Known Limitations and Deferred Work](#known-limitations-and-deferred-work)
|
|
19
|
+
- [Dev Note](#dev-note)
|
|
20
|
+
|
|
21
|
+
-----
|
|
22
|
+
|
|
23
|
+
<a id="use-this-package"></a>
|
|
24
|
+
## Use this package
|
|
25
|
+
|
|
26
|
+
Any composition that calls a model provider — an agent loop, a session-title generator, a compaction summarizer — streams its requests through this service. Mount it together with at least one provider adapter; the service itself has no configuration and no provider wire code.
|
|
27
|
+
|
|
28
|
+
### When to choose it
|
|
29
|
+
|
|
30
|
+
Choose this package whenever a plugin or composition needs to call a model: it is the only supported path into provider adapters, and it keeps one vocabulary across the loop, the session log, and every consumer. Do not reach for it when you need provider-specific wire behavior (that belongs in the `xlaunch-llm-gateway` adapter) or retry execution (that belongs in `xlaunch-llm-retry`).
|
|
31
|
+
|
|
32
|
+
### Minimal composition
|
|
33
|
+
|
|
34
|
+
Mount the service and at least one adapter, then select the provider by name in every request:
|
|
35
|
+
|
|
36
|
+
```yaml
|
|
37
|
+
- name: '@xlaunch/llm'
|
|
38
|
+
- name: '@xlaunch/llm-gateway'
|
|
39
|
+
config:
|
|
40
|
+
providers:
|
|
41
|
+
xlaunch:
|
|
42
|
+
apiKeyEnv: XLAUNCH_API_KEY
|
|
43
|
+
api: openai-completions
|
|
44
|
+
baseURL: https://gateway.xlaunch.work/v1
|
|
45
|
+
models:
|
|
46
|
+
- id: auto
|
|
47
|
+
- id: fusion
|
|
48
|
+
```
|
|
49
|
+
|
|
50
|
+
A stream returns token-level chunks and always ends with one terminal `finish` chunk. `BlockAssembler` turns the chunks into content blocks and messages; `AssistantStreamAccumulator` preserves their exact timestamps and token boundaries in a compact representation that the loop embeds in one durable attempt settlement:
|
|
51
|
+
|
|
52
|
+
```text
|
|
53
|
+
for await (const chunk of ctx.llm.stream({
|
|
54
|
+
provider: 'xlaunch',
|
|
55
|
+
model: 'auto',
|
|
56
|
+
messages: [createUserMessage({ content: [{ type: 'text', text: 'Hello' }] })],
|
|
57
|
+
})) {
|
|
58
|
+
// chunks: block-start, text-delta, ..., usage, finish
|
|
59
|
+
}
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
After a successful mount, `ctx.llm.listProviders()` reports the registered routes in registration order.
|
|
63
|
+
|
|
64
|
+
### What you can do
|
|
65
|
+
|
|
66
|
+
- **Stream one model call** — `ctx.llm.stream(options)` yields raw chunks (token-level deltas) for any registered provider and model; consumers assemble them with `BlockAssembler`.
|
|
67
|
+
- **Register provider adapters** — an adapter owns one or more provider routes, and its registration captures that route's retry policy; registering the same route twice fails with `DUPLICATE_ADAPTER`.
|
|
68
|
+
- **Expose and activate providers through configuration** — adapters declare configurable-provider routes plus a settings namespace, so configuration surfaces can activate dormant providers and edit connection facts without a restart.
|
|
69
|
+
- **Discover and resolve models** — list the models an adapter advertises, interrogate an endpoint for the models it serves, and resolve one exact model's context window, output default, reasoning efforts, and input modalities.
|
|
70
|
+
- **Validate call config** — an explicit or configured reasoning effort is checked against the exact model before any provider I/O, and an adapter-configured output cap is materialized when the request omits one.
|
|
71
|
+
- **Read an embedded Assistant stream without expanding it** — `assistantStreamFirstTokenTime` (first token), `assistantStreamHasVisibleContent` (any visible content), and `assistantStreamHasVisibleText` (any visible text) answer their questions from the compact records with early exit; `lastAssistantStreamChunk` scans backward to the last raw chunk of one type, `assistantStreamChunks` and `joinAssistantStreamText` scan the whole stream, and `assembleAssistantStream` feeds a `BlockAssembler` one joined delta per run with the same blocks, usage, and replay state as the per-member expansion. `runFirstTokenTime` and `runFirstVisibleTime` do the early-exit scan for one packed run, and `isTokenDelta`, `isVisibleChunk`, and `chunkHasVisibleText` define the token and visibility rules for a single chunk. `expandAssistantStream` remains the validating path for records read at a durable boundary; it is not memoized, because a retained expansion costs roughly ten times the compact stream for as long as the event lives.
|
|
72
|
+
|
|
73
|
+
### Failures and recovery
|
|
74
|
+
|
|
75
|
+
Every stream ends in exactly one terminal `finish` chunk: `{ kind: 'error', failure }` on failure, `{ kind: 'aborted', failure }` on cancellation. Failures carry stable codes such as `NO_ADAPTER`, `MISSING_CREDENTIAL`, `AUTH`, `RATE_LIMIT`, and `CONTEXT_WINDOW_EXCEEDED`; consumers route on the code, never on message text. A request naming an unregistered provider fails with `NO_ADAPTER`, and a malformed credential fails with `INVALID_CREDENTIAL` instead of surfacing as an opaque fetch error. This service never re-runs a request: retrying is the job of `xlaunch-llm-retry` at the agent's failed-step extension point.
|
|
76
|
+
|
|
77
|
+
-----
|
|
78
|
+
|
|
79
|
+
<a id="understand-the-implementation"></a>
|
|
80
|
+
## Understand the implementation
|
|
81
|
+
|
|
82
|
+
<details>
|
|
83
|
+
<summary>Implementation internals — click to expand</summary>
|
|
84
|
+
|
|
85
|
+
This section explains the design behind the service; the observable behavior is fully covered in [Use this package](#use-this-package).
|
|
86
|
+
|
|
87
|
+
### Design philosophy
|
|
88
|
+
|
|
89
|
+
The service is built on one separation: **the logical contract is provider-neutral, adapters own the wire.** It defines the canonical message, content-block, and stream-chunk vocabulary once, and every provider adapter translates only its own wire format into that vocabulary. The registry is the topology owner — adapter routes, configurable-provider entries, and discovery offers all register here and are disposed with their fiber — while a request stays a pure function of the session log: loop-built requests arrive deep-frozen, so listeners and adapters read them and never rewrite them.
|
|
90
|
+
|
|
91
|
+
### Source map
|
|
92
|
+
|
|
93
|
+
| File | Role |
|
|
94
|
+
|---|---|
|
|
95
|
+
| [`src/index.ts`](src/index.ts) | The `LlmRuntime` service: adapter registry, configurable-provider directory, model discovery, call preparation, and the streaming boundary |
|
|
96
|
+
| [`src/types.ts`](src/types.ts) | The `StreamChunk` protocol, content-block map, finish reasons, and shared vocabulary |
|
|
97
|
+
| [`src/message.ts`](src/message.ts) | Immutable message constructors shared by delivery, history, and requests |
|
|
98
|
+
| [`src/assembler.ts`](src/assembler.ts) | `BlockAssembler`: incremental chunk-to-block assembly |
|
|
99
|
+
| [`src/assistant-stream.ts`](src/assistant-stream.ts) | Compact timed Assistant stream accumulation, strict validation, exact expansion, and record-level readers |
|
|
100
|
+
| [`src/call-config.ts`](src/call-config.ts) | Call-config validation, adapter-default materialization, and request freezing |
|
|
101
|
+
| [`src/retry-policy.ts`](src/retry-policy.ts) | Provider-owned retry policy resolution (normal and always modes) |
|
|
102
|
+
| [`src/error.ts`](src/error.ts) | `HarnessError`/`LlmError` taxonomy and provider-neutral failure codes |
|
|
103
|
+
| [`src/content.ts`](src/content.ts) | Shared image-content helpers, including request-image offloading |
|
|
104
|
+
| [`src/api-key.ts`](src/api-key.ts) | Credential format check shared by every adapter |
|
|
105
|
+
| [`src/adapter-failure.ts`](src/adapter-failure.ts) | Failure normalization into terminal finish chunks |
|
|
106
|
+
|
|
107
|
+
### Main flow
|
|
108
|
+
|
|
109
|
+
A request is validated against its exact model's capability — context window, output default, reasoning efforts, and input modalities — and any adapter-configured defaults are materialized, then the whole request is deep-frozen. `prepareCall()` binds those facts, detached context, and retry policy to the exact adapter generation that performs terminal dispatch, so HMR or dynamic settings cannot combine one generation's image capability with another generation's endpoint. An image-capable adapter projects durable references into route-specific request versions; `resolveImageAttachmentAccess()` separately maps an attachment provider's optional host object into the current tool execution world without changing the request image or its `variantId`. A text-only route receives deterministic per-image placeholders, including nested tool-result images, without rewriting append-only session history. Durable `FileBlock` references never reach any adapter: request assembly replaces each one, nested tool-result occurrences included, with deterministic handle text naming the file and its saved read-only path, resolved through the mounted attachment and filesystem providers. `ctx.llm.fileRequestText(ref)` exposes that exact synchronous projection to request measurement. `offloadRequestImagesWithPolicy()` removes oldest images deterministically by raw or base64 size and count or byte quanta; the pure `offloadedImagePrefixCount()` exposes that decision so route-owned request pricing can reproduce it without building the projection. Adapters that charge visual tokens declare per-route `imageRequestPricing`, which `ctx.llm.imageRequestPricing(provider, model)` resolves synchronously for the token meter. Dispatch goes through the `llm/stream` waterfall, then chunks return as token-level deltas and every adapter outcome reaches the consumer as one terminal `finish` chunk.
|
|
110
|
+
|
|
111
|
+
### Invariants
|
|
112
|
+
|
|
113
|
+
- **Model-visible ⟺ logged** — anything that reaches a provider request is reconstructable from the session log; loop-built requests are deep-frozen and never rewritten.
|
|
114
|
+
- **Replay state travels only within one adapter** — assistant replay state rides along only when the same adapter instance owns the historical and target routes; otherwise it is dropped before dispatch.
|
|
115
|
+
- **Prepared calls are one-shot** — a prepared call can be dispatched exactly once, and its call-config fields must match the prepared config.
|
|
116
|
+
- **Image projection follows the captured route** — durable `ImageBlock` references become route-specific request versions only for image-capable models; text-only models receive stable placeholders.
|
|
117
|
+
- **File projection is unconditional** — no provider receives file bytes; every route gets one deterministic handle line per `FileBlock`, and the model reads the saved copy with its file tools on demand.
|
|
118
|
+
- **Protocol ordering** — `usage` precedes `finish`, tool arguments stay raw JSON strings, and nothing follows the terminal `finish`.
|
|
119
|
+
- **Registry mutations are atomic** — route and directory registration validates the whole candidate set before anything moves, so a refused change leaves the previous state serving.
|
|
120
|
+
|
|
121
|
+
</details>
|
|
122
|
+
|
|
123
|
+
-----
|
|
124
|
+
|
|
125
|
+
<a id="further-exploration"></a>
|
|
126
|
+
## Further Exploration
|
|
127
|
+
|
|
128
|
+
Read these pages when the package-level contract is not enough. They move from the shared types to the concrete adapters, the retry executor, and the measurement service.
|
|
129
|
+
|
|
130
|
+
- [LLM streaming subsystem](../../../docs/subsystems/llm-streaming.md) — the message and block types, compact Assistant stream records, the `StreamChunk` protocol, and the adapter contract.
|
|
131
|
+
- [llm-gateway adapter](../llm-gateway/README.md) — the gateway-backed multi-provider implementation.
|
|
132
|
+
- [llm-retry](../llm-retry/README.md) — the retry executor that re-runs failed model requests.
|
|
133
|
+
- [Token meter](../token-meter/README.md) — replay-aware request and context pressure measurement.
|
|
134
|
+
- [Terminal LLM stream failures](../../../.agents/notes/implemented/architecture/2026-07-29-terminal-llm-stream-failures.md) — the service boundary between model-request outcomes and plugin failures.
|
|
135
|
+
|
|
136
|
+
-----
|
|
137
|
+
|
|
138
|
+
<a id="model-experience"></a>
|
|
139
|
+
## Model Experience
|
|
140
|
+
|
|
141
|
+
None, as the LLM service adds no content; adapters choose when to add the shared image descriptors and per-image placeholders exported by this package.
|
|
142
|
+
|
|
143
|
+
#### KV Cache effect
|
|
144
|
+
|
|
145
|
+
Reasoning-effort materialization preserves the assembled request prefix. Image identity and request-preview text are deterministic, while an optional execution-world path is resolved for each request; a changed path or image-offload boundary can prevent reuse from that image.
|
|
146
|
+
|
|
147
|
+
## Known Limitations and Deferred Work
|
|
148
|
+
|
|
149
|
+
<a id="known-limitations-and-deferred-work"></a>
|
|
150
|
+
|
|
151
|
+
|
|
152
|
+
These limits define where this service stops and other packages or future work begin. They are current package constraints, not a task backlog.
|
|
153
|
+
|
|
154
|
+
- **No retry execution, caching, or rate limiting ships in this service** — provider registration stores the retry policy, but a stream remains a single provider attempt; `@xlaunch/llm-retry` executes the policy at durable agent-step boundaries.
|
|
155
|
+
- **`GenerateOptions` sampling is `temperature`/`maxTokens`/`stop` only** — no `tool_choice`, `top_p`, or penalty fields; the vocabulary grows when a producer lands ([dropped inert knobs](../../../.agents/notes/archived/simplification/2026-07-04-drop-inert-request-knobs.md)).
|
|
156
|
+
- **Producer-gated variants stay out until produced** — `prefill`, per-tool `strict`, block `cache` hints, and the `agent` message-source variant have no producer ([Agent Note](../../../.agents/notes/archived/simplification/2026-07-04-prune-producerless-vocabulary-variants.md)).
|
|
157
|
+
- **`BlockAssembler` handles core block kinds only** — a plugin-added block type whose stream is never closed by `block-end` makes `blocks()` throw.
|
|
158
|
+
- **`GenerateOptions.sessionId` is a locally-declared brand** — importing xlaunch-session's `SessionId` would create a dependency cycle.
|
|
159
|
+
|
|
160
|
+
<a id="dev-note"></a>
|
|
161
|
+
### Dev Note
|
|
162
|
+
|
|
163
|
+
<details>
|
|
164
|
+
<summary>Working context for maintainers — click to expand</summary>
|
|
165
|
+
|
|
166
|
+
This Dev Note is non-authoritative working context: open questions and undecided directions. Shipped behavior and accepted rationale live in the sections above, the package code, and the linked Agent Notes.
|
|
167
|
+
|
|
168
|
+
#### Open items
|
|
169
|
+
|
|
170
|
+
- `GenerateOptions.sessionId` is a locally-declared brand because importing xlaunch-session's `SessionId` would create a dependency cycle; a future ids-owning package could dissolve the workaround.
|
|
171
|
+
- Reasoning-effort identifiers are adapter-owned opaque strings resolved only against each adapter's advertised set; a shared cross-adapter effort vocabulary is not decided.
|
|
172
|
+
- The `llm/adapters-updated` event is payload-free by design; consumers re-read the registries instead of receiving the new topology in the event.
|
|
173
|
+
|
|
174
|
+
</details>
|