@latimer-woods-tech/llm 0.2.0 → 0.3.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,16 +1,81 @@
1
- ## [Unreleased]
1
+ # Changelog
2
2
 
3
- ## [0.2.0] - 2026-04-29
3
+ ## 0.3.1 — 2026-05-02 PM
4
+
5
+ ### Added (no breaking changes)
6
+
7
+ - **Grok opt-in provider.** `LLMProvider` now includes `'grok'`. Call with `{ model: 'grok-4-fast' }` or `{ model: 'grok-3-mini-latest' }` to route to xAI. Not in default tier routing — explicit model override only.
8
+ - `LLMEnv.GROK_API_KEY` is **optional**; required only when a caller opts in via a `grok-*` model override.
9
+ - `MODELS.grok.{fast, mini}` added to the catalogue.
10
+
11
+ ### Rationale
12
+
13
+ Reversal of the 0.3.0 partial decision (D1 in `docs/supervisor/DECISIONS.md`): Grok stays available for workloads that benefit from its quirks — cheap experimental prompting in the xico-city user economy, artist-platform surface, etc. It does not compete with Anthropic / Gemini / Groq for the tier slots; it sits alongside as a per-call opt-in.
14
+
15
+ ### Consumers
16
+
17
+ - Downstream packages do not need to declare `GROK_API_KEY` unless a route explicitly sends Grok model overrides. Existing code that dropped the key on 0.3.0 migration continues to work.
18
+
19
+
20
+ ## 0.3.0 — 2026-05-02
21
+
22
+ ### Breaking
23
+
24
+ - **Grok removed.** `LLMProvider` no longer includes `grok`; `LLMEnv.GROK_API_KEY` removed.
25
+ - **AI Gateway mandatory.** `LLMEnv.AI_GATEWAY_BASE_URL` is required. All provider traffic flows through
26
+ Cloudflare AI Gateway for unified logging, rate limiting, and cost telemetry.
27
+ - **Tier-based routing.** `LLMOptions.tier` (`fast | balanced | smart | verifier`) replaces the
28
+ flat Anthropic → Grok → Groq failover chain. Routing is workload-split:
29
+ - `fast` → Anthropic Haiku (latency-sensitive, short completions)
30
+ - `balanced` → Anthropic Sonnet (default)
31
+ - `smart` → Anthropic Opus; Gemini 2.5 Pro for long-context (>150k tokens est.)
32
+ - `verifier` → Groq Llama 3.3 70B (cheap second opinion, no fallback)
33
+ - **Dependency declarations fixed.** `@latimer-woods-tech/errors` and `@latimer-woods-tech/logger`
34
+ now declared as `^0.2.0` instead of `file:../*` so external consumers can install the package.
4
35
 
5
36
  ### Added
6
- - Multi-provider LLM completion chain with Anthropic primary, Grok fallback, and Groq tertiary fallback.
7
- - Streaming support through the Anthropic-compatible provider path.
8
- - `withSystem()` helper for consistently applying system prompts at call sites.
9
- - Quality-gated package build with lint, typecheck, coverage, and ESM output.
10
37
 
11
- ### Fixed
12
- - Restored `@latimer-woods-tech/errors` and `@latimer-woods-tech/logger` as runtime dependencies so published consumers resolve package imports correctly.
38
+ - Gemini 2.5 Pro via Vertex AI as long-context fallback (`VERTEX_ACCESS_TOKEN`, `VERTEX_PROJECT`, `VERTEX_LOCATION`).
39
+ - Anthropic prompt caching (auto-enabled for system prompts ≥ 4096 chars; override via `LLMOptions.promptCache`).
40
+ - Per-call cancellation via `LLMOptions.signal: AbortSignal`.
41
+ - 3-attempt exponential backoff (250ms · 2^n + jitter) on retryable status (408, 425, 429, 5xx).
42
+ - Ledger stamping: `LLMOptions.{runId, project, actor}` propagated to logger for downstream
43
+ consumption by `@latimer-woods-tech/llm-meter`.
44
+ - `LLMResult.{model, tier, attempts, gatewayRequestId, tokens.cacheRead, tokens.cacheWrite}` for
45
+ observability and budget accounting.
46
+
47
+ ### Changed
48
+
49
+ - Default model constants renamed and grouped under exported `MODELS` catalogue. Kept in sync with
50
+ `docs/architecture/FACTORY_V1.md § LLM substrate`.
51
+ - Empty provider responses now surface as leg failures and trigger fallback rather than returning
52
+ an empty `LLMResult`.
53
+
54
+ ### Migration
55
+
56
+ Replace:
57
+
58
+ ```ts
59
+ const res = await complete(msgs, { ANTHROPIC_API_KEY, GROK_API_KEY, GROQ_API_KEY }, { model: 'claude-sonnet-4-...' });
60
+ ```
61
+
62
+ With:
63
+
64
+ ```ts
65
+ const res = await complete(
66
+ msgs,
67
+ {
68
+ AI_GATEWAY_BASE_URL: env.AI_GATEWAY_BASE_URL,
69
+ ANTHROPIC_API_KEY: env.ANTHROPIC_API_KEY,
70
+ GROQ_API_KEY: env.GROQ_API_KEY,
71
+ VERTEX_ACCESS_TOKEN: env.VERTEX_ACCESS_TOKEN, // minted via JWT-bearer flow
72
+ VERTEX_PROJECT: env.VERTEX_PROJECT,
73
+ VERTEX_LOCATION: env.VERTEX_LOCATION,
74
+ },
75
+ { tier: 'balanced', runId, project: 'prime-self', actor: 'worker' },
76
+ );
77
+ ```
78
+
79
+ ## 0.2.0
13
80
 
14
- ### Verification
15
- - `npm install` regenerated the package lock and ran the package prepublish gate on Apr 29, 2026.
16
- - `npm run lint`, `npm run typecheck`, `npm test -- --coverage`, and `npm run build` passed during the prepublish gate.
81
+ Initial public release of the failover orchestrator.
package/README.md CHANGED
@@ -1,3 +1,60 @@
1
- # @factory/llm
1
+ # @latimer-woods-tech/llm
2
2
 
3
- LLM failover orchestration across Factory AI providers.
3
+ Tier-routed LLM orchestration for the Factory platform, with Cloudflare AI Gateway, Anthropic
4
+ primary, Gemini 2.5 Pro long-context fallback, and Groq verifier.
5
+
6
+ ## Routing (0.3.0)
7
+
8
+ | Tier | Primary | Fallback | Notes |
9
+ |---|---|---|---|
10
+ | `fast` | Claude Haiku 4 | — | latency-sensitive, short completions |
11
+ | `balanced` *(default)* | Claude Sonnet 4 | Gemini 2.5 Pro | swaps to Gemini when est. tokens ≥ 150k |
12
+ | `smart` | Claude Opus 4 | Gemini 2.5 Pro | ditto; tools, long reasoning |
13
+ | `verifier` | Groq Llama 3.3 70B | — | cheap second opinion; no fallback |
14
+
15
+ All traffic flows through `AI_GATEWAY_BASE_URL`. The gateway handles caching, rate-limit shedding,
16
+ and per-project cost telemetry.
17
+
18
+ ## Usage
19
+
20
+ ```ts
21
+ import { complete } from '@latimer-woods-tech/llm';
22
+
23
+ const ctl = new AbortController();
24
+ setTimeout(() => ctl.abort(), 30_000);
25
+
26
+ const res = await complete(
27
+ [{ role: 'user', content: 'Draft a release note for llm@0.3.0.' }],
28
+ {
29
+ AI_GATEWAY_BASE_URL: env.AI_GATEWAY_BASE_URL,
30
+ ANTHROPIC_API_KEY: env.ANTHROPIC_API_KEY,
31
+ GROQ_API_KEY: env.GROQ_API_KEY,
32
+ VERTEX_ACCESS_TOKEN: env.VERTEX_ACCESS_TOKEN,
33
+ VERTEX_PROJECT: env.VERTEX_PROJECT,
34
+ VERTEX_LOCATION: env.VERTEX_LOCATION,
35
+ },
36
+ {
37
+ tier: 'balanced',
38
+ signal: ctl.signal,
39
+ runId: 'sup-run-0042',
40
+ project: 'prime-self',
41
+ actor: 'supervisor',
42
+ },
43
+ );
44
+
45
+ if (res.ok) {
46
+ console.log(res.data.content, res.data.provider, res.data.tokens);
47
+ }
48
+ ```
49
+
50
+ ## Vertex access token
51
+
52
+ Minted via the JWT-bearer flow from a GCP service account. See
53
+ `docs/runbooks/rotate-gcp-sa.md` for the mint procedure. Tokens are short-lived (1h); callers
54
+ should refresh before every cold start or every 50 minutes, whichever comes first.
55
+
56
+ ## Budget
57
+
58
+ Every call emits a `llm.complete` log line with `{ provider, model, tier, tokens, runId, project,
59
+ actor }`. `@latimer-woods-tech/llm-meter` (0.1.0+) consumes these lines to enforce the per-run
60
+ $5 hard cap and the per-project steady-state budget.
package/dist/index.d.mts CHANGED
@@ -8,38 +8,83 @@ interface LLMMessage {
8
8
  role: 'user' | 'assistant' | 'system';
9
9
  content: string;
10
10
  }
11
+ /**
12
+ * Quality tier selected by the caller. Routing is workload-split:
13
+ * - `fast` → Anthropic Haiku (short, latency-sensitive)
14
+ * - `balanced` → Anthropic Sonnet (default)
15
+ * - `smart` → Anthropic Opus OR Gemini 2.5 Pro if input is long-context (>150k tokens estimated)
16
+ * - `verifier` → Groq Llama (cheap second opinion; only used from verifier code path)
17
+ */
18
+ type LLMTier = 'fast' | 'balanced' | 'smart' | 'verifier';
11
19
  /**
12
20
  * Options that influence LLM completion behaviour.
13
21
  */
14
22
  interface LLMOptions {
23
+ /** Quality tier; see {@link LLMTier}. Defaults to `balanced`. */
24
+ tier?: LLMTier;
25
+ /** Explicit model override. Takes precedence over tier. */
15
26
  model?: string;
16
27
  maxTokens?: number;
17
28
  temperature?: number;
18
29
  system?: string;
30
+ /** Token budget above which we force long-context routing (Gemini). */
31
+ longContextThreshold?: number;
32
+ /** Per-call cancellation signal. Aborts the in-flight provider request. */
33
+ signal?: AbortSignal;
34
+ /** Optional run identifier stamped on ledger rows + logs. */
35
+ runId?: string;
36
+ /** Optional project identifier stamped on ledger rows + logs. */
37
+ project?: string;
38
+ /** Optional actor identifier (supervisor / worker / human). */
39
+ actor?: string;
40
+ /** Anthropic prompt-cache control. Defaults to `true` for `system` prompts ≥ 1024 tokens. */
41
+ promptCache?: boolean;
19
42
  }
20
43
  /**
21
- * Provider that produced an LLM response.
44
+ * Provider that produced an LLM response. `grok` removed in 0.3.0.
22
45
  */
23
- type LLMProvider = 'anthropic' | 'grok' | 'groq';
46
+ type LLMProvider = 'anthropic' | 'gemini' | 'groq' | 'grok';
24
47
  /**
25
48
  * Result returned by a successful completion.
26
49
  */
27
50
  interface LLMResult {
28
51
  content: string;
29
52
  provider: LLMProvider;
53
+ model: string;
54
+ tier: LLMTier;
30
55
  tokens: {
31
56
  input: number;
32
57
  output: number;
58
+ cacheRead?: number;
59
+ cacheWrite?: number;
33
60
  };
34
61
  latency: number;
62
+ /** Number of attempts before success (1 = primary succeeded). */
63
+ attempts: number;
64
+ /** Monotonic request id from AI Gateway, if present in headers. */
65
+ gatewayRequestId?: string;
35
66
  }
36
67
  /**
37
68
  * Environment bindings required by {@link complete}.
69
+ *
70
+ * `AI_GATEWAY_BASE_URL` is REQUIRED in 0.3.0. All provider calls flow through the
71
+ * Cloudflare AI Gateway for unified logging, rate limiting, and cost telemetry.
72
+ * In test/dev the caller may pass a custom fetch impl that short-circuits this.
38
73
  */
39
74
  interface LLMEnv {
75
+ AI_GATEWAY_BASE_URL: string;
40
76
  ANTHROPIC_API_KEY: string;
41
- GROK_API_KEY: string;
42
77
  GROQ_API_KEY: string;
78
+ /** Optional — only required when caller passes `{ model: 'grok-*' }` override. */
79
+ GROK_API_KEY?: string;
80
+ /**
81
+ * Google Cloud short-lived access token with `aiplatform.endpoints.predict`.
82
+ * Callers mint this via the JWT-bearer flow (service account → token exchange);
83
+ * see `docs/runbooks/rotate-gcp-sa.md`. Token must be valid for ≥ 5 minutes.
84
+ */
85
+ VERTEX_ACCESS_TOKEN: string;
86
+ VERTEX_PROJECT: string;
87
+ VERTEX_LOCATION: string;
43
88
  }
44
89
  /**
45
90
  * Optional dependencies for {@link complete}.
@@ -49,39 +94,113 @@ interface LLMDeps {
49
94
  logger?: Logger;
50
95
  now?: () => number;
51
96
  }
97
+ declare const MODELS: {
98
+ readonly anthropic: {
99
+ readonly fast: "claude-haiku-4-20250514";
100
+ readonly balanced: "claude-sonnet-4-6";
101
+ readonly smart: "claude-opus-4-7";
102
+ };
103
+ readonly gemini: {
104
+ readonly smart: "gemini-2.5-pro";
105
+ };
106
+ readonly groq: {
107
+ readonly verifier: "llama-3.3-70b-versatile";
108
+ };
109
+ readonly grok: {
110
+ /** Opt-in only via `{ model: 'grok-*' }`. Not in default tier routing. */
111
+ readonly fast: "grok-4-fast";
112
+ readonly mini: "grok-3-mini-latest";
113
+ };
114
+ };
115
+ /** Cooldown duration in ms after a provider exhausts all retries. */
116
+ declare const PROVIDER_COOLDOWN_MS = 30000;
117
+ /**
118
+ * Returns `true` if the provider is currently in its cooldown window.
119
+ * Uses the injected `now` function (or `Date.now`) for testability.
120
+ */
121
+ declare function isProviderCoolingDown(provider: LLMProvider, now?: () => number): boolean;
52
122
  /**
53
- * Runs a completion through the Anthropic → Grok → Groq failover chain.
123
+ * Marks a provider as cooling down for {@link PROVIDER_COOLDOWN_MS} milliseconds.
124
+ */
125
+ declare function markProviderCoolingDown(provider: LLMProvider, now?: () => number): void;
126
+ /**
127
+ * Clears the cooldown state for a provider after a successful call.
128
+ */
129
+ declare function clearProviderCooldown(provider: LLMProvider): void;
130
+ declare const BASE_BACKOFF_MS = 250;
131
+ /**
132
+ * Run a completion through the routing plan for the requested tier.
133
+ *
134
+ * Routing summary (0.3.0):
135
+ * - `fast` → Anthropic Haiku
136
+ * - `balanced` → Anthropic Sonnet; Gemini 2.5 Pro if `longContextThreshold` exceeded
137
+ * - `smart` → Anthropic Opus; Gemini 2.5 Pro if long-context
138
+ * - `verifier` → Groq Llama 3.3 70B (no fallback — verifier is inherently cheap/best-effort)
139
+ *
140
+ * All provider traffic flows through Cloudflare AI Gateway at `AI_GATEWAY_BASE_URL`.
141
+ *
142
+ * Per-provider reliability guarantees (0.4.0):
143
+ * - Exponential backoff with jitter on 429 / 5xx (base 500ms, cap 8s, up to 2 retries).
144
+ * - Provider cooldown: after exhausting retries the provider is marked cooling down
145
+ * for 30 seconds; subsequent calls skip it and go straight to the fallback leg.
54
146
  *
55
147
  * @param messages - Ordered chat history.
56
- * @param env - API key bindings.
57
- * @param opts - Optional model/parameters override.
148
+ * @param env - API key + gateway bindings.
149
+ * @param opts - Optional tier/model/parameters override.
58
150
  * @param deps - Optional fetch/logger/clock injection (for testing).
59
151
  * @returns A {@link FactoryResponse} carrying either an {@link LLMResult} or
60
- * an `LLM_ALL_PROVIDERS_FAILED` error.
152
+ * an error (`LLM_ALL_PROVIDERS_FAILED`, `LLM_RATE_LIMITED`, or `INTERNAL_ERROR`).
61
153
  */
62
154
  declare function complete(messages: LLMMessage[], env: LLMEnv, opts?: LLMOptions, deps?: LLMDeps): Promise<FactoryResponse<LLMResult>>;
63
155
  /**
64
- * Streams a completion from Anthropic. No failover is performed for streaming
65
- * responses; callers should fall back to {@link complete} on failure.
156
+ * Streams a completion from the primary Anthropic provider, yielding text chunks
157
+ * as they arrive. Falls back to the non-streaming {@link complete} function when
158
+ * the provider does not support streaming (i.e. a non-Anthropic primary is selected).
159
+ *
160
+ * The generator's **return value** (accessible via `gen.return()` or by consuming
161
+ * the full iteration) is an {@link LLMResult} with the same shape as {@link complete}.
162
+ *
163
+ * Usage pattern:
164
+ * ```ts
165
+ * const gen = completionStream(messages, env, opts);
166
+ * for await (const chunk of gen) {
167
+ * // stream chunk to client
168
+ * }
169
+ * const result = (await gen.return(undefined)).value; // LLMResult
170
+ * ```
66
171
  *
67
172
  * @param messages - Ordered chat history.
68
- * @param env - Anthropic API key binding.
69
- * @param opts - Optional model/parameters override.
70
- * @param deps - Optional fetch override (for testing).
71
- * @returns The raw Anthropic streaming response body.
173
+ * @param env - API key + gateway bindings.
174
+ * @param opts - Optional tier/model/parameters override. Accepts `deps` as nested field.
175
+ * @returns An async generator that yields `string` chunks and returns an {@link LLMResult}.
72
176
  */
73
- declare function stream(messages: LLMMessage[], env: {
74
- ANTHROPIC_API_KEY: string;
75
- }, opts?: LLMOptions, deps?: {
76
- fetch?: typeof fetch;
77
- }): Promise<ReadableStream<Uint8Array>>;
177
+ declare function completionStream(messages: LLMMessage[], env: LLMEnv, opts?: LLMOptions & {
178
+ deps?: LLMDeps;
179
+ }): AsyncGenerator<string, LLMResult, unknown>;
78
180
  /**
79
- * Returns a {@link complete}-compatible function with a system prompt
80
- * pre-bound, so callers can hand around a domain-specific shortcut.
181
+ * Returns `true` if `response` contains at least one verbatim phrase of at
182
+ * least 5 consecutive whitespace-delimited tokens that also appears in one of
183
+ * the `sources` strings.
184
+ *
185
+ * Returns `true` unconditionally when `sources` is empty (no grounding
186
+ * documents means grounding cannot be violated).
187
+ *
188
+ * This is a lightweight guard for RAG pipelines — it detects obvious
189
+ * hallucinations where the model generates content not present in any
190
+ * retrieved source. It is NOT a semantic similarity check.
191
+ *
192
+ * @param response - The LLM-generated text to inspect.
193
+ * @param sources - Retrieved source documents to check against.
194
+ * @returns `true` if the response is grounded, `false` if hallucination detected.
81
195
  *
82
- * @param system - System prompt to prepend to every call.
83
- * @returns A function that invokes {@link complete} with `system` injected.
196
+ * @example
197
+ * ```ts
198
+ * const grounded = assertGrounding(llmAnswer, retrievedDocs);
199
+ * if (!grounded) {
200
+ * // flag or re-rank the response
201
+ * }
202
+ * ```
84
203
  */
85
- declare function withSystem(system: string): (messages: LLMMessage[], env: LLMEnv, opts?: LLMOptions, deps?: LLMDeps) => Promise<FactoryResponse<LLMResult>>;
204
+ declare function assertGrounding(response: string, sources: string[]): boolean;
86
205
 
87
- export { type LLMDeps, type LLMEnv, type LLMMessage, type LLMOptions, type LLMProvider, type LLMResult, complete, stream, withSystem };
206
+ export { BASE_BACKOFF_MS, type LLMDeps, type LLMEnv, type LLMMessage, type LLMOptions, type LLMProvider, type LLMResult, type LLMTier, MODELS, PROVIDER_COOLDOWN_MS, assertGrounding, clearProviderCooldown, complete, completionStream, isProviderCoolingDown, markProviderCoolingDown };