@latimer-woods-tech/llm 0.2.0 → 0.4.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +159 -11
- package/README.md +66 -2
- package/dist/index.d.mts +248 -23
- package/dist/index.mjs +693 -168
- package/dist/index.mjs.map +1 -1
- package/package.json +13 -5
package/CHANGELOG.md
CHANGED
|
@@ -1,16 +1,164 @@
|
|
|
1
|
-
|
|
1
|
+
# Changelog
|
|
2
2
|
|
|
3
|
-
##
|
|
3
|
+
## 0.4.1 — 2026-06-03
|
|
4
|
+
|
|
5
|
+
### Added (no breaking changes)
|
|
6
|
+
|
|
7
|
+
- Export `MODEL_PRICE_PER_1M` — the canonical USD-per-1M-tokens rate table. This
|
|
8
|
+
makes it the single source of truth for pricing across the platform;
|
|
9
|
+
`@latimer-woods-tech/llm-meter` now derives its cents table from it and enforces
|
|
10
|
+
parity with a drift-guard test. Make all rate changes here.
|
|
11
|
+
|
|
12
|
+
---
|
|
13
|
+
|
|
14
|
+
## 0.3.4 — 2026-05-28
|
|
15
|
+
|
|
16
|
+
### Added (no breaking changes)
|
|
17
|
+
|
|
18
|
+
- **`workbench` tier** routes to `deepseek-chat` with Groq fallback for boring,
|
|
19
|
+
reviewable, non-sensitive internal batch work.
|
|
20
|
+
- **`DEEPSEEK_API_KEY`** added as an optional `LLMEnv` binding. It is required only
|
|
21
|
+
for `tier: 'workbench'` or explicit `deepseek-*` model overrides.
|
|
22
|
+
- **DeepSeek pricing entries** for `deepseek-chat` and `deepseek-reasoner` so cost
|
|
23
|
+
caps and ledger rows use known rates instead of the conservative Opus fallback.
|
|
24
|
+
|
|
25
|
+
### Guardrail
|
|
26
|
+
|
|
27
|
+
- `workbench` is for docs summaries, changelog drafts, issue triage, and classification.
|
|
28
|
+
Do not route secrets, customer PII, billing data, production ops, or final
|
|
29
|
+
customer-facing answers through DeepSeek.
|
|
30
|
+
|
|
31
|
+
---
|
|
32
|
+
|
|
33
|
+
## 0.3.3 — 2026-05-27
|
|
34
|
+
|
|
35
|
+
### Changed (no breaking changes)
|
|
36
|
+
|
|
37
|
+
- **`fast` tier now routes to Grok 4.3 (primary) → Anthropic Haiku (fallback).**
|
|
38
|
+
When `GROK_API_KEY` is present in `LLMEnv`, the `fast` tier sends completions to
|
|
39
|
+
`grok-4.3` first and falls back to `claude-haiku-4-20250514` if Grok is unavailable
|
|
40
|
+
or returns an error. When `GROK_API_KEY` is absent, the request goes directly to
|
|
41
|
+
Anthropic Haiku (same behaviour as before). Callers that already set `tier: 'fast'`
|
|
42
|
+
pick up Grok routing automatically; no code change needed.
|
|
4
43
|
|
|
5
44
|
### Added
|
|
6
|
-
- Multi-provider LLM completion chain with Anthropic primary, Grok fallback, and Groq tertiary fallback.
|
|
7
|
-
- Streaming support through the Anthropic-compatible provider path.
|
|
8
|
-
- `withSystem()` helper for consistently applying system prompts at call sites.
|
|
9
|
-
- Quality-gated package build with lint, typecheck, coverage, and ESM output.
|
|
10
45
|
|
|
11
|
-
|
|
12
|
-
|
|
46
|
+
- **`LLMOptions.reasoningEffort`** — `'none' | 'low' | 'medium' | 'high'`
|
|
47
|
+
(optional, defaults to `'none'`). Forwarded to the Grok `reasoning_effort` parameter
|
|
48
|
+
for `grok-4.3`; ignored for all other providers.
|
|
49
|
+
- **`MODELS.grok.fast`** updated from `'grok-4-fast'` to `'grok-4.3'`. Old alias
|
|
50
|
+
`'grok-4-fast'` retained as deprecated with updated pricing so historical ledger
|
|
51
|
+
rows remain accurate.
|
|
52
|
+
- **Grok pricing update**: `grok-4.3` at $1.25/$2.50 per MTok (in/out).
|
|
53
|
+
Old `grok-4-fast` and `grok-3-mini-latest` re-priced to the same $1.25/$2.50 rate.
|
|
54
|
+
- **`for (const [legIndex, leg] of routeLegs.entries())`** — renamed loop variable so
|
|
55
|
+
`legIndex` is accessible for the last-leg 429 rate-limit short-circuit.
|
|
56
|
+
|
|
57
|
+
### Consumers
|
|
58
|
+
|
|
59
|
+
- To opt-in to Grok 4.3: add `GROK_API_KEY` to your `LLMEnv` binding. No other code
|
|
60
|
+
change required for `tier: 'fast'` callers.
|
|
61
|
+
- `GROK_API_KEY` is optional. If absent, fast-tier routing is identical to 0.3.2
|
|
62
|
+
(Anthropic Haiku only).
|
|
63
|
+
|
|
64
|
+
---
|
|
65
|
+
|
|
66
|
+
## 0.3.2 — 2026-05-27
|
|
67
|
+
|
|
68
|
+
### Added (no breaking changes)
|
|
69
|
+
|
|
70
|
+
- **`LLMOptions.workload`** — optional string label (`'insights'`, `'copy'`, `'lead-qualification'`, …)
|
|
71
|
+
forwarded to cost-recording calls for per-workload cost attribution in dashboards.
|
|
72
|
+
- **Missing model pricing entries** in `MODEL_PRICE_PER_1M`:
|
|
73
|
+
`claude-haiku-4-5-20251001`, `claude-sonnet-4-20250514`, `claude-opus-4-20250514`
|
|
74
|
+
(aliases for variants that share pricing with their shorthand names; prevents
|
|
75
|
+
`estimateCostUsd` from silently returning `$0` for these model IDs).
|
|
76
|
+
- **`workload` forwarded** to `recordOrgCostUsage` so per-workload breakdowns appear
|
|
77
|
+
in the cost-tracking KV store.
|
|
78
|
+
|
|
79
|
+
### Consumers
|
|
80
|
+
|
|
81
|
+
- Existing callers do not need to set `workload`; it is optional and defaults to `undefined`.
|
|
82
|
+
- The `recordOrgCostUsage` signature is unchanged; `workload` is an additive internal field.
|
|
83
|
+
|
|
84
|
+
---
|
|
85
|
+
|
|
86
|
+
## 0.3.1 — 2026-05-02 PM
|
|
87
|
+
|
|
88
|
+
### Added (no breaking changes)
|
|
89
|
+
|
|
90
|
+
- **Grok opt-in provider.** `LLMProvider` now includes `'grok'`. Call with `{ model: 'grok-4-fast' }` or `{ model: 'grok-3-mini-latest' }` to route to xAI. Not in default tier routing — explicit model override only.
|
|
91
|
+
- `LLMEnv.GROK_API_KEY` is **optional**; required only when a caller opts in via a `grok-*` model override.
|
|
92
|
+
- `MODELS.grok.{fast, mini}` added to the catalogue.
|
|
93
|
+
|
|
94
|
+
### Rationale
|
|
95
|
+
|
|
96
|
+
Reversal of the 0.3.0 partial decision (D1 in `docs/supervisor/DECISIONS.md`): Grok stays available for workloads that benefit from its quirks — cheap experimental prompting in the xico-city user economy, artist-platform surface, etc. It does not compete with Anthropic / Gemini / Groq for the tier slots; it sits alongside as a per-call opt-in.
|
|
97
|
+
|
|
98
|
+
### Consumers
|
|
99
|
+
|
|
100
|
+
- Downstream packages do not need to declare `GROK_API_KEY` unless a route explicitly sends Grok model overrides. Existing code that dropped the key on 0.3.0 migration continues to work.
|
|
101
|
+
|
|
102
|
+
|
|
103
|
+
## 0.3.0 — 2026-05-02
|
|
104
|
+
|
|
105
|
+
### Breaking
|
|
106
|
+
|
|
107
|
+
- **Grok removed.** `LLMProvider` no longer includes `grok`; `LLMEnv.GROK_API_KEY` removed.
|
|
108
|
+
- **AI Gateway mandatory.** `LLMEnv.AI_GATEWAY_BASE_URL` is required. All provider traffic flows through
|
|
109
|
+
Cloudflare AI Gateway for unified logging, rate limiting, and cost telemetry.
|
|
110
|
+
- **Tier-based routing.** `LLMOptions.tier` (`fast | balanced | smart | verifier`) replaces the
|
|
111
|
+
flat Anthropic → Grok → Groq failover chain. Routing is workload-split:
|
|
112
|
+
- `fast` → Anthropic Haiku (latency-sensitive, short completions)
|
|
113
|
+
- `balanced` → Anthropic Sonnet (default)
|
|
114
|
+
- `smart` → Anthropic Opus; Gemini 2.5 Pro for long-context (>150k tokens est.)
|
|
115
|
+
- `verifier` → Groq Llama 3.3 70B (cheap second opinion, no fallback)
|
|
116
|
+
- **Dependency declarations fixed.** `@latimer-woods-tech/errors` and `@latimer-woods-tech/logger`
|
|
117
|
+
now declared as `^0.2.0` instead of `file:../*` so external consumers can install the package.
|
|
118
|
+
|
|
119
|
+
### Added
|
|
120
|
+
|
|
121
|
+
- Gemini 2.5 Pro via Vertex AI as long-context fallback (`VERTEX_ACCESS_TOKEN`, `VERTEX_PROJECT`, `VERTEX_LOCATION`).
|
|
122
|
+
- Anthropic prompt caching (auto-enabled for system prompts ≥ 4096 chars; override via `LLMOptions.promptCache`).
|
|
123
|
+
- Per-call cancellation via `LLMOptions.signal: AbortSignal`.
|
|
124
|
+
- 3-attempt exponential backoff (250ms · 2^n + jitter) on retryable status (408, 425, 429, 5xx).
|
|
125
|
+
- Ledger stamping: `LLMOptions.{runId, project, actor}` propagated to logger for downstream
|
|
126
|
+
consumption by `@latimer-woods-tech/llm-meter`.
|
|
127
|
+
- `LLMResult.{model, tier, attempts, gatewayRequestId, tokens.cacheRead, tokens.cacheWrite}` for
|
|
128
|
+
observability and budget accounting.
|
|
129
|
+
|
|
130
|
+
### Changed
|
|
131
|
+
|
|
132
|
+
- Default model constants renamed and grouped under exported `MODELS` catalogue. Kept in sync with
|
|
133
|
+
`docs/architecture/FACTORY_V1.md § LLM substrate`.
|
|
134
|
+
- Empty provider responses now surface as leg failures and trigger fallback rather than returning
|
|
135
|
+
an empty `LLMResult`.
|
|
136
|
+
|
|
137
|
+
### Migration
|
|
138
|
+
|
|
139
|
+
Replace:
|
|
140
|
+
|
|
141
|
+
```ts
|
|
142
|
+
const res = await complete(msgs, { ANTHROPIC_API_KEY, GROK_API_KEY, GROQ_API_KEY }, { model: 'claude-sonnet-4-...' });
|
|
143
|
+
```
|
|
144
|
+
|
|
145
|
+
With:
|
|
146
|
+
|
|
147
|
+
```ts
|
|
148
|
+
const res = await complete(
|
|
149
|
+
msgs,
|
|
150
|
+
{
|
|
151
|
+
AI_GATEWAY_BASE_URL: env.AI_GATEWAY_BASE_URL,
|
|
152
|
+
ANTHROPIC_API_KEY: env.ANTHROPIC_API_KEY,
|
|
153
|
+
GROQ_API_KEY: env.GROQ_API_KEY,
|
|
154
|
+
VERTEX_ACCESS_TOKEN: env.VERTEX_ACCESS_TOKEN, // minted via JWT-bearer flow
|
|
155
|
+
VERTEX_PROJECT: env.VERTEX_PROJECT,
|
|
156
|
+
VERTEX_LOCATION: env.VERTEX_LOCATION,
|
|
157
|
+
},
|
|
158
|
+
{ tier: 'balanced', runId, project: 'prime-self', actor: 'worker' },
|
|
159
|
+
);
|
|
160
|
+
```
|
|
161
|
+
|
|
162
|
+
## 0.2.0
|
|
13
163
|
|
|
14
|
-
|
|
15
|
-
- `npm install` regenerated the package lock and ran the package prepublish gate on Apr 29, 2026.
|
|
16
|
-
- `npm run lint`, `npm run typecheck`, `npm test -- --coverage`, and `npm run build` passed during the prepublish gate.
|
|
164
|
+
Initial public release of the failover orchestrator.
|
package/README.md
CHANGED
|
@@ -1,3 +1,67 @@
|
|
|
1
|
-
# @
|
|
1
|
+
# @latimer-woods-tech/llm
|
|
2
2
|
|
|
3
|
-
LLM
|
|
3
|
+
Tier-routed LLM orchestration for the Factory platform, with Cloudflare AI Gateway, Anthropic
|
|
4
|
+
primary, Gemini 2.5 Pro long-context fallback, Groq verifier, and DeepSeek workbench routing.
|
|
5
|
+
|
|
6
|
+
## Routing (0.3.0)
|
|
7
|
+
|
|
8
|
+
| Tier | Primary | Fallback | Notes |
|
|
9
|
+
|---|---|---|---|
|
|
10
|
+
| `fast` | Claude Haiku 4 | — | latency-sensitive, short completions |
|
|
11
|
+
| `balanced` *(default)* | Claude Sonnet 4 | Gemini 2.5 Pro | swaps to Gemini when est. tokens ≥ 150k |
|
|
12
|
+
| `smart` | Claude Opus 4 | Gemini 2.5 Pro | ditto; tools, long reasoning |
|
|
13
|
+
| `verifier` | Groq Llama 3.3 70B | — | cheap second opinion; no fallback |
|
|
14
|
+
| `workbench` | DeepSeek Chat | Groq Llama | boring, reviewable, non-sensitive batch work |
|
|
15
|
+
|
|
16
|
+
All traffic flows through `AI_GATEWAY_BASE_URL`. The gateway handles caching, rate-limit shedding,
|
|
17
|
+
and per-project cost telemetry.
|
|
18
|
+
|
|
19
|
+
## Usage
|
|
20
|
+
|
|
21
|
+
```ts
|
|
22
|
+
import { complete } from '@latimer-woods-tech/llm';
|
|
23
|
+
|
|
24
|
+
const ctl = new AbortController();
|
|
25
|
+
setTimeout(() => ctl.abort(), 30_000);
|
|
26
|
+
|
|
27
|
+
const res = await complete(
|
|
28
|
+
[{ role: 'user', content: 'Draft a release note for llm@0.3.0.' }],
|
|
29
|
+
{
|
|
30
|
+
AI_GATEWAY_BASE_URL: env.AI_GATEWAY_BASE_URL,
|
|
31
|
+
ANTHROPIC_API_KEY: env.ANTHROPIC_API_KEY,
|
|
32
|
+
GROQ_API_KEY: env.GROQ_API_KEY,
|
|
33
|
+
DEEPSEEK_API_KEY: env.DEEPSEEK_API_KEY,
|
|
34
|
+
VERTEX_ACCESS_TOKEN: env.VERTEX_ACCESS_TOKEN,
|
|
35
|
+
VERTEX_PROJECT: env.VERTEX_PROJECT,
|
|
36
|
+
VERTEX_LOCATION: env.VERTEX_LOCATION,
|
|
37
|
+
},
|
|
38
|
+
{
|
|
39
|
+
tier: 'balanced',
|
|
40
|
+
signal: ctl.signal,
|
|
41
|
+
runId: 'sup-run-0042',
|
|
42
|
+
project: 'prime-self',
|
|
43
|
+
actor: 'supervisor',
|
|
44
|
+
},
|
|
45
|
+
);
|
|
46
|
+
|
|
47
|
+
if (res.ok) {
|
|
48
|
+
console.log(res.data.content, res.data.provider, res.data.tokens);
|
|
49
|
+
}
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
Use `tier: 'workbench'` only for low-risk internal jobs: docs summaries, changelog drafts,
|
|
53
|
+
classification, issue triage, and other outputs a human or higher-trust model can review. Do not
|
|
54
|
+
send secrets, customer PII, production ops requests, billing data, or final customer-facing answers
|
|
55
|
+
through this lane.
|
|
56
|
+
|
|
57
|
+
## Vertex access token
|
|
58
|
+
|
|
59
|
+
Minted via the JWT-bearer flow from a GCP service account. See
|
|
60
|
+
`docs/runbooks/rotate-gcp-sa.md` for the mint procedure. Tokens are short-lived (1h); callers
|
|
61
|
+
should refresh before every cold start or every 50 minutes, whichever comes first.
|
|
62
|
+
|
|
63
|
+
## Budget
|
|
64
|
+
|
|
65
|
+
Every call emits a `llm.complete` log line with `{ provider, model, tier, tokens, runId, project,
|
|
66
|
+
actor }`. `@latimer-woods-tech/llm-meter` (0.1.0+) consumes these lines to enforce the per-run
|
|
67
|
+
$5 hard cap and the per-project steady-state budget.
|
package/dist/index.d.mts
CHANGED
|
@@ -8,38 +8,165 @@ interface LLMMessage {
|
|
|
8
8
|
role: 'user' | 'assistant' | 'system';
|
|
9
9
|
content: string;
|
|
10
10
|
}
|
|
11
|
+
/**
|
|
12
|
+
* Quality tier selected by the caller. Routing is workload-split:
|
|
13
|
+
* - `fast` → Grok 4.3 with Anthropic Haiku fallback (routine drafts/small jobs)
|
|
14
|
+
* - `balanced` → Anthropic Sonnet (default)
|
|
15
|
+
* - `smart` → Anthropic Opus OR Gemini 2.5 Pro if input is long-context (>150k tokens estimated)
|
|
16
|
+
* - `verifier` → Groq Llama (cheap second opinion; only used from verifier code path)
|
|
17
|
+
* - `workbench` → DeepSeek Chat with Groq fallback (boring, reviewable, non-sensitive batch work)
|
|
18
|
+
*/
|
|
19
|
+
type LLMTier = 'fast' | 'balanced' | 'smart' | 'verifier' | 'workbench';
|
|
11
20
|
/**
|
|
12
21
|
* Options that influence LLM completion behaviour.
|
|
13
22
|
*/
|
|
14
23
|
interface LLMOptions {
|
|
24
|
+
/** Quality tier; see {@link LLMTier}. Defaults to `balanced`. */
|
|
25
|
+
tier?: LLMTier;
|
|
26
|
+
/** Explicit model override. Takes precedence over tier. */
|
|
15
27
|
model?: string;
|
|
16
28
|
maxTokens?: number;
|
|
17
29
|
temperature?: number;
|
|
18
30
|
system?: string;
|
|
31
|
+
/** Token budget above which we force long-context routing (Gemini). */
|
|
32
|
+
longContextThreshold?: number;
|
|
33
|
+
/** Per-call cancellation signal. Aborts the in-flight provider request. */
|
|
34
|
+
signal?: AbortSignal;
|
|
35
|
+
/** Optional run identifier stamped on ledger rows + logs. */
|
|
36
|
+
runId?: string;
|
|
37
|
+
/** Optional project identifier stamped on ledger rows + logs. */
|
|
38
|
+
project?: string;
|
|
39
|
+
/** Optional actor identifier (supervisor / worker / human). */
|
|
40
|
+
actor?: string;
|
|
41
|
+
/** Optional workload label used in logs and cost-policy call sites. */
|
|
42
|
+
workload?: string;
|
|
43
|
+
/** Grok reasoning effort. Defaults to `none` for cost-controlled fast/draft calls. */
|
|
44
|
+
reasoningEffort?: 'none' | 'low' | 'medium' | 'high';
|
|
45
|
+
/** Anthropic prompt-cache control. Defaults to `true` for `system` prompts ≥ 1024 tokens. */
|
|
46
|
+
promptCache?: boolean;
|
|
47
|
+
/**
|
|
48
|
+
* Maximum estimated cost in USD for this completion.
|
|
49
|
+
* This cap is enforced after the provider returns because it uses actual
|
|
50
|
+
* response token counts to compute the final cost.
|
|
51
|
+
* If the post-call estimated cost exceeds this cap, `complete` returns a
|
|
52
|
+
* {@link RateLimitError} with code `LLM_COST_CAP_EXCEEDED` and
|
|
53
|
+
* `completionStream` throws the same error.
|
|
54
|
+
* Pricing is based on {@link MODEL_PRICE_PER_1M}; unknown models default to
|
|
55
|
+
* Opus rates (conservative upper bound).
|
|
56
|
+
*/
|
|
57
|
+
maxCostUsd?: number;
|
|
58
|
+
/**
|
|
59
|
+
* Org-level daily cost cap in USD. Requires `env.LLM_COST_KV` to be set.
|
|
60
|
+
* When today's cumulative spend read from KV is >= this value, `complete`
|
|
61
|
+
* returns a {@link RateLimitError} with code `LLM_DAILY_CAP_EXCEEDED`
|
|
62
|
+
* without making any provider call. After a successful call the daily
|
|
63
|
+
* accumulator is updated in KV (TTL: 48 h).
|
|
64
|
+
*/
|
|
65
|
+
dailyCapUsd?: number;
|
|
66
|
+
/**
|
|
67
|
+
* Org-level monthly cost cap in USD. Requires `env.LLM_COST_KV` to be set.
|
|
68
|
+
* Same enforcement pattern as {@link dailyCapUsd} but keyed by YYYY-MM.
|
|
69
|
+
* KV TTL: 40 days.
|
|
70
|
+
*/
|
|
71
|
+
monthlyCapUsd?: number;
|
|
72
|
+
/**
|
|
73
|
+
* Metering context. When supplied and `deps.onRecord` is set, a {@link LLMRecordRow}
|
|
74
|
+
* is emitted after every successful completion. Errors are swallowed.
|
|
75
|
+
*/
|
|
76
|
+
ledger?: LLMRecordContext;
|
|
19
77
|
}
|
|
20
78
|
/**
|
|
21
79
|
* Provider that produced an LLM response.
|
|
22
80
|
*/
|
|
23
|
-
type LLMProvider = 'anthropic' | '
|
|
81
|
+
type LLMProvider = 'anthropic' | 'gemini' | 'groq' | 'grok' | 'deepseek';
|
|
24
82
|
/**
|
|
25
83
|
* Result returned by a successful completion.
|
|
26
84
|
*/
|
|
27
85
|
interface LLMResult {
|
|
28
86
|
content: string;
|
|
29
87
|
provider: LLMProvider;
|
|
88
|
+
model: string;
|
|
89
|
+
tier: LLMTier;
|
|
30
90
|
tokens: {
|
|
31
91
|
input: number;
|
|
32
92
|
output: number;
|
|
93
|
+
cacheRead?: number;
|
|
94
|
+
cacheWrite?: number;
|
|
33
95
|
};
|
|
34
96
|
latency: number;
|
|
97
|
+
/** Number of attempts before success (1 = primary succeeded). */
|
|
98
|
+
attempts: number;
|
|
99
|
+
/** Monotonic request id from AI Gateway, if present in headers. */
|
|
100
|
+
gatewayRequestId?: string;
|
|
35
101
|
}
|
|
36
102
|
/**
|
|
37
103
|
* Environment bindings required by {@link complete}.
|
|
104
|
+
*
|
|
105
|
+
* `AI_GATEWAY_BASE_URL` is REQUIRED in 0.3.0. All provider calls flow through the
|
|
106
|
+
* Cloudflare AI Gateway for unified logging, rate limiting, and cost telemetry.
|
|
107
|
+
* In test/dev the caller may pass a custom fetch impl that short-circuits this.
|
|
38
108
|
*/
|
|
39
109
|
interface LLMEnv {
|
|
110
|
+
AI_GATEWAY_BASE_URL: string;
|
|
40
111
|
ANTHROPIC_API_KEY: string;
|
|
41
|
-
GROK_API_KEY: string;
|
|
42
112
|
GROQ_API_KEY: string;
|
|
113
|
+
/** Optional — only required for `{ tier: 'workbench' }` or `deepseek-*` model overrides. */
|
|
114
|
+
DEEPSEEK_API_KEY?: string;
|
|
115
|
+
/** Optional — only required when caller passes `{ model: 'grok-*' }` override. */
|
|
116
|
+
GROK_API_KEY?: string;
|
|
117
|
+
/**
|
|
118
|
+
* Google Cloud short-lived access token with `aiplatform.endpoints.predict`.
|
|
119
|
+
* Callers mint this via the JWT-bearer flow (service account → token exchange);
|
|
120
|
+
* see `docs/runbooks/rotate-gcp-sa.md`. Token must be valid for ≥ 5 minutes.
|
|
121
|
+
*/
|
|
122
|
+
VERTEX_ACCESS_TOKEN: string;
|
|
123
|
+
VERTEX_PROJECT: string;
|
|
124
|
+
VERTEX_LOCATION: string;
|
|
125
|
+
/**
|
|
126
|
+
* Optional KV store for org-level daily/monthly cost tracking and enforcement.
|
|
127
|
+
* When provided alongside {@link LLMOptions.dailyCapUsd} or {@link LLMOptions.monthlyCapUsd},
|
|
128
|
+
* `complete` will block calls that would exceed the declared cap.
|
|
129
|
+
* Any KV-like store satisfying `get`/`put` works (e.g. Cloudflare KV, in-memory stub).
|
|
130
|
+
*/
|
|
131
|
+
LLM_COST_KV?: CostKvStore;
|
|
132
|
+
}
|
|
133
|
+
/**
|
|
134
|
+
* Minimal KV store interface for org-level LLM cost tracking.
|
|
135
|
+
* Cloudflare KV satisfies this. An in-memory stub is sufficient for tests.
|
|
136
|
+
*/
|
|
137
|
+
interface CostKvStore {
|
|
138
|
+
get(key: string): Promise<string | null>;
|
|
139
|
+
put(key: string, value: string, options?: {
|
|
140
|
+
expirationTtl?: number;
|
|
141
|
+
}): Promise<void>;
|
|
142
|
+
}
|
|
143
|
+
/**
|
|
144
|
+
* Caller-supplied context stamped on every metering row.
|
|
145
|
+
* Mirrors the `LLMRecordContext` in `@latimer-woods-tech/llm-meter`; kept inline
|
|
146
|
+
* to avoid a circular dependency (llm-meter imports llm).
|
|
147
|
+
*/
|
|
148
|
+
interface LLMRecordContext {
|
|
149
|
+
project: string;
|
|
150
|
+
actor: string;
|
|
151
|
+
runId?: string;
|
|
152
|
+
workload?: string;
|
|
153
|
+
tenantId?: string;
|
|
154
|
+
}
|
|
155
|
+
/**
|
|
156
|
+
* Row shape passed to the optional {@link LLMDeps.onRecord} callback.
|
|
157
|
+
* Callers can wire this directly to `recordCall` from `@latimer-woods-tech/llm-meter`.
|
|
158
|
+
*/
|
|
159
|
+
interface LLMRecordRow extends LLMRecordContext {
|
|
160
|
+
model: string;
|
|
161
|
+
provider: LLMProvider;
|
|
162
|
+
tier: LLMTier;
|
|
163
|
+
inputTokens: number;
|
|
164
|
+
outputTokens: number;
|
|
165
|
+
cacheReadTokens: number;
|
|
166
|
+
cacheWriteTokens: number;
|
|
167
|
+
latencyMs: number;
|
|
168
|
+
costUsd: number;
|
|
169
|
+
yyyyMm: string;
|
|
43
170
|
}
|
|
44
171
|
/**
|
|
45
172
|
* Optional dependencies for {@link complete}.
|
|
@@ -48,40 +175,138 @@ interface LLMDeps {
|
|
|
48
175
|
fetch?: typeof fetch;
|
|
49
176
|
logger?: Logger;
|
|
50
177
|
now?: () => number;
|
|
178
|
+
/**
|
|
179
|
+
* Optional metering callback. Called after every successful completion.
|
|
180
|
+
* Errors are swallowed so metering never blocks the caller.
|
|
181
|
+
* Wire to `recordCall` from `@latimer-woods-tech/llm-meter`.
|
|
182
|
+
*/
|
|
183
|
+
onRecord?: (row: LLMRecordRow) => Promise<void>;
|
|
51
184
|
}
|
|
185
|
+
declare const MODELS: {
|
|
186
|
+
readonly anthropic: {
|
|
187
|
+
readonly fast: "claude-haiku-4-20250514";
|
|
188
|
+
readonly balanced: "claude-sonnet-4-6";
|
|
189
|
+
readonly smart: "claude-opus-4-7";
|
|
190
|
+
};
|
|
191
|
+
readonly gemini: {
|
|
192
|
+
readonly smart: "gemini-2.5-pro";
|
|
193
|
+
};
|
|
194
|
+
readonly groq: {
|
|
195
|
+
readonly verifier: "llama-4-maverick";
|
|
196
|
+
};
|
|
197
|
+
readonly grok: {
|
|
198
|
+
readonly fast: "grok-4.3";
|
|
199
|
+
};
|
|
200
|
+
readonly deepseek: {
|
|
201
|
+
readonly workbench: "deepseek-chat";
|
|
202
|
+
};
|
|
203
|
+
};
|
|
204
|
+
/** Cooldown duration in ms after a provider exhausts all retries. */
|
|
205
|
+
declare const PROVIDER_COOLDOWN_MS = 30000;
|
|
52
206
|
/**
|
|
53
|
-
*
|
|
207
|
+
* Returns `true` if the provider is currently in its cooldown window.
|
|
208
|
+
* Uses the injected `now` function (or `Date.now`) for testability.
|
|
209
|
+
*/
|
|
210
|
+
declare function isProviderCoolingDown(provider: LLMProvider, now?: () => number): boolean;
|
|
211
|
+
/**
|
|
212
|
+
* USD cost per 1 million tokens for each model.
|
|
213
|
+
* Source: Anthropic / Google / xAI pricing pages as of 2026-05.
|
|
214
|
+
* Keep these model names in sync with the default routing constants in
|
|
215
|
+
* {@link MODELS}; unknown models fall back to Opus rates (conservative upper bound).
|
|
216
|
+
*
|
|
217
|
+
* CANONICAL pricing source for the platform. `@latimer-woods-tech/llm-meter`
|
|
218
|
+
* derives its cents-denominated rates from this table and a drift-guard test
|
|
219
|
+
* there fails CI if they diverge — make all rate changes here.
|
|
220
|
+
*/
|
|
221
|
+
declare const MODEL_PRICE_PER_1M: Record<string, {
|
|
222
|
+
input: number;
|
|
223
|
+
output: number;
|
|
224
|
+
cacheRead: number;
|
|
225
|
+
cacheWrite: number;
|
|
226
|
+
}>;
|
|
227
|
+
/**
|
|
228
|
+
* Marks a provider as cooling down for {@link PROVIDER_COOLDOWN_MS} milliseconds.
|
|
229
|
+
*/
|
|
230
|
+
declare function markProviderCoolingDown(provider: LLMProvider, now?: () => number): void;
|
|
231
|
+
/**
|
|
232
|
+
* Clears the cooldown state for a provider after a successful call.
|
|
233
|
+
*/
|
|
234
|
+
declare function clearProviderCooldown(provider: LLMProvider): void;
|
|
235
|
+
declare const BASE_BACKOFF_MS = 250;
|
|
236
|
+
/**
|
|
237
|
+
* Run a completion through the routing plan for the requested tier.
|
|
238
|
+
*
|
|
239
|
+
* Routing summary (0.3.0):
|
|
240
|
+
* - `fast` → Grok 4.3; Anthropic Haiku fallback when Grok is unavailable
|
|
241
|
+
* - `balanced` → Anthropic Sonnet; Gemini 2.5 Pro if `longContextThreshold` exceeded
|
|
242
|
+
* - `smart` → Anthropic Opus; Gemini 2.5 Pro if long-context
|
|
243
|
+
* - `verifier` → Groq Llama 3.3 70B (no fallback — verifier is inherently cheap/best-effort)
|
|
244
|
+
* - `workbench` → DeepSeek Chat; Groq fallback for boring/reviewable internal batch jobs
|
|
245
|
+
*
|
|
246
|
+
* All provider traffic flows through Cloudflare AI Gateway at `AI_GATEWAY_BASE_URL`.
|
|
247
|
+
*
|
|
248
|
+
* Per-provider reliability guarantees (0.4.0):
|
|
249
|
+
* - Exponential backoff with jitter on 429 / 5xx (base 500ms, cap 8s, up to 2 retries).
|
|
250
|
+
* - Provider cooldown: after exhausting retries the provider is marked cooling down
|
|
251
|
+
* for 30 seconds; subsequent calls skip it and go straight to the fallback leg.
|
|
54
252
|
*
|
|
55
253
|
* @param messages - Ordered chat history.
|
|
56
|
-
* @param env - API key bindings.
|
|
57
|
-
* @param opts - Optional model/parameters override.
|
|
254
|
+
* @param env - API key + gateway bindings.
|
|
255
|
+
* @param opts - Optional tier/model/parameters override.
|
|
58
256
|
* @param deps - Optional fetch/logger/clock injection (for testing).
|
|
59
257
|
* @returns A {@link FactoryResponse} carrying either an {@link LLMResult} or
|
|
60
|
-
* an `LLM_ALL_PROVIDERS_FAILED`
|
|
258
|
+
* an error (`LLM_ALL_PROVIDERS_FAILED`, `LLM_RATE_LIMITED`, or `INTERNAL_ERROR`).
|
|
61
259
|
*/
|
|
62
260
|
declare function complete(messages: LLMMessage[], env: LLMEnv, opts?: LLMOptions, deps?: LLMDeps): Promise<FactoryResponse<LLMResult>>;
|
|
63
261
|
/**
|
|
64
|
-
* Streams a completion from
|
|
65
|
-
*
|
|
262
|
+
* Streams a completion from the primary Anthropic provider, yielding text chunks
|
|
263
|
+
* as they arrive. Falls back to the non-streaming {@link complete} function when
|
|
264
|
+
* the provider does not support streaming (i.e. a non-Anthropic primary is selected).
|
|
265
|
+
*
|
|
266
|
+
* The generator's **return value** (accessible via `gen.return()` or by consuming
|
|
267
|
+
* the full iteration) is an {@link LLMResult} with the same shape as {@link complete}.
|
|
268
|
+
*
|
|
269
|
+
* Usage pattern:
|
|
270
|
+
* ```ts
|
|
271
|
+
* const gen = completionStream(messages, env, opts);
|
|
272
|
+
* for await (const chunk of gen) {
|
|
273
|
+
* // stream chunk to client
|
|
274
|
+
* }
|
|
275
|
+
* const result = (await gen.return(undefined)).value; // LLMResult
|
|
276
|
+
* ```
|
|
66
277
|
*
|
|
67
278
|
* @param messages - Ordered chat history.
|
|
68
|
-
* @param env -
|
|
69
|
-
* @param opts - Optional model/parameters override.
|
|
70
|
-
* @
|
|
71
|
-
* @returns The raw Anthropic streaming response body.
|
|
279
|
+
* @param env - API key + gateway bindings.
|
|
280
|
+
* @param opts - Optional tier/model/parameters override. Accepts `deps` as nested field.
|
|
281
|
+
* @returns An async generator that yields `string` chunks and returns an {@link LLMResult}.
|
|
72
282
|
*/
|
|
73
|
-
declare function
|
|
74
|
-
|
|
75
|
-
},
|
|
76
|
-
fetch?: typeof fetch;
|
|
77
|
-
}): Promise<ReadableStream<Uint8Array>>;
|
|
283
|
+
declare function completionStream(messages: LLMMessage[], env: LLMEnv, opts?: LLMOptions & {
|
|
284
|
+
deps?: LLMDeps;
|
|
285
|
+
}): AsyncGenerator<string, LLMResult, unknown>;
|
|
78
286
|
/**
|
|
79
|
-
* Returns
|
|
80
|
-
*
|
|
287
|
+
* Returns `true` if `response` contains at least one verbatim phrase of at
|
|
288
|
+
* least 5 consecutive whitespace-delimited tokens that also appears in one of
|
|
289
|
+
* the `sources` strings.
|
|
290
|
+
*
|
|
291
|
+
* Returns `true` unconditionally when `sources` is empty (no grounding
|
|
292
|
+
* documents means grounding cannot be violated).
|
|
293
|
+
*
|
|
294
|
+
* This is a lightweight guard for RAG pipelines — it detects obvious
|
|
295
|
+
* hallucinations where the model generates content not present in any
|
|
296
|
+
* retrieved source. It is NOT a semantic similarity check.
|
|
297
|
+
*
|
|
298
|
+
* @param response - The LLM-generated text to inspect.
|
|
299
|
+
* @param sources - Retrieved source documents to check against.
|
|
300
|
+
* @returns `true` if the response is grounded, `false` if hallucination detected.
|
|
81
301
|
*
|
|
82
|
-
* @
|
|
83
|
-
*
|
|
302
|
+
* @example
|
|
303
|
+
* ```ts
|
|
304
|
+
* const grounded = assertGrounding(llmAnswer, retrievedDocs);
|
|
305
|
+
* if (!grounded) {
|
|
306
|
+
* // flag or re-rank the response
|
|
307
|
+
* }
|
|
308
|
+
* ```
|
|
84
309
|
*/
|
|
85
|
-
declare function
|
|
310
|
+
declare function assertGrounding(response: string, sources: string[]): boolean;
|
|
86
311
|
|
|
87
|
-
export { type LLMDeps, type LLMEnv, type LLMMessage, type LLMOptions, type LLMProvider, type LLMResult, complete,
|
|
312
|
+
export { BASE_BACKOFF_MS, type CostKvStore, type LLMDeps, type LLMEnv, type LLMMessage, type LLMOptions, type LLMProvider, type LLMRecordContext, type LLMRecordRow, type LLMResult, type LLMTier, MODELS, MODEL_PRICE_PER_1M, PROVIDER_COOLDOWN_MS, assertGrounding, clearProviderCooldown, complete, completionStream, isProviderCoolingDown, markProviderCoolingDown };
|