@juspay/neurolink 12.46.1 → 12.47.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,8 +1,8 @@
1
- ## [12.46.1](https://github.com/juspay/neurolink/compare/v12.46.0...v12.46.1) (2026-10-04)
1
+ ## [12.47.0](https://github.com/juspay/neurolink/compare/v12.46.1...v12.47.0) (2026-10-05)
2
2
 
3
- ### Bug Fixes
3
+ ### Features
4
4
 
5
- - **(providers):** fix the provider defects the reviewers found ([6dac3ed](https://github.com/juspay/neurolink/commit/6dac3ed724353637c375a7ee732c3922e71c3d7d))
5
+ - **(decide):** add Cloudflare Clef as a decide provider with image input ([9d14900](https://github.com/juspay/neurolink/commit/9d14900509f3552d6647f4909bf8a4f227b3ca83))
6
6
 
7
7
  ## [11.2.3](https://github.com/juspay/neurolink/compare/v11.2.2...v11.2.3) (2026-08-19)
8
8
 
package/README.md CHANGED
@@ -47,7 +47,7 @@ const decision = await pipe.tryDecide({
47
47
 
48
48
  **NeuroLink is the pipe layer of an AI nervous system.** Providers — OpenAI, Anthropic, Google, AWS, Azure, Mistral, local runtimes like Ollama, and dozens more — are the neurons: each generates a different kind of intelligence, at a different cost and latency. NeuroLink is the vascular layer that carries that intelligence, as a stream, to the applications — the organs — that consume it, across three inference types: `generate` and `stream` produce text, `decide` produces a calibrated `boolean`/`choice`/`score` judgment instead. A curated model registry (64 models, 132 aliases) backs metadata, routing, and context-window checks out of the box, and hundreds more models are reachable through aggregator providers — 100+ via LiteLLM, 300+ via OpenRouter.
49
49
 
50
- Extracted from production systems at Juspay, NeuroLink provides a practical, TypeScript-first way to plug any application into that nervous system. Switch which neuron answers a request with a single parameter change — OpenAI, Anthropic, Google, AWS Bedrock, Azure, a local runtime, or any provider you add. `decide` is the third inference type — a typed, calibrated judgment instead of text — for the model-routing and gating decisions `generate`/`stream` were never meant to make, powered by a purpose-built decision model (TypeSafe Jev, Perplexity's hosted Decisions API, or the open-weights Laya or XOR) rather than a general-purpose LLM: with Jev, routing decisions land in ~400ms for about $0.00002, instead of a full generation call.
50
+ Extracted from production systems at Juspay, NeuroLink provides a practical, TypeScript-first way to plug any application into that nervous system. Switch which neuron answers a request with a single parameter change — OpenAI, Anthropic, Google, AWS Bedrock, Azure, a local runtime, or any provider you add. `decide` is the third inference type — a typed, calibrated judgment instead of text — for the model-routing and gating decisions `generate`/`stream` were never meant to make, powered by a purpose-built decision model (TypeSafe Jev, Perplexity's hosted Decisions API, Cloudflare's hosted Clef, or the open-weights Laya or XOR) rather than a general-purpose LLM: with Jev, routing decisions land in ~400ms for about $0.00002, instead of a full generation call.
51
51
 
52
52
  **Why NeuroLink?** Three genuine inference types, not one dressed up three ways — `generate` and `stream` produce text; `decide` produces a calibrated `boolean`/`choice`/`score` judgment, and which types a provider serves is declared per-provider via `inferenceKinds` rather than inferred from behavior. Every neuron plugs into the same pipe, including 3 fully local runtimes (Ollama, LM Studio, llama.cpp) with per-request credential overrides, and MCP support covers all 4 transports (stdio, HTTP, SSE, WebSocket). Every AI-driven optimization the pipe performs — model routing, context compaction, tool selection — fails open: no key configured behaves exactly like NeuroLink without it, and routing uses asymmetric confidence thresholds (upgrade at 0.3, downgrade at 0.6) rather than a single cutoff, because a wrong downgrade costs more than a wrong upgrade. Switch providers with a single parameter change, leverage built-in tools plus any MCP-compliant tool server, deploy with confidence using enterprise features like Redis memory and multi-provider failover, and optimize costs automatically with intelligent routing. Use it via our professional CLI or TypeScript SDK—whichever fits your workflow.
53
53
 
@@ -59,37 +59,37 @@ Extracted from production systems at Juspay, NeuroLink provides a practical, Typ
59
59
 
60
60
  ## What's New
61
61
 
62
- | Feature | Version | Description | Guide |
63
- | -------------------------------------------------------------------- | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- |
64
- | **`decide` Inference Type + TypeSafe Jev + Laya + XOR + Perplexity** | next | A third inference type alongside `generate`/`stream`: typed, calibrated judgments (`boolean`, `choice`, `score`) via `neurolink.decide()` / `tryDecide()`, one parallel pass (~400ms and ~$0.00002/decision on Jev). Providers: TypeSafe Jev (`TYPESAFE_API_KEY`, also reachable via the Vercel AI Gateway); [Laya](docs/getting-started/providers/laya.md) (`LAYA_API_KEY` + `LAYA_BASE_URL`), an open-weights model you run yourself; [XOR](docs/getting-started/providers/xor.md) (`XOR_API_KEY` + `XOR_BASE_URL`), Juspay's open-weights model; and [Perplexity](docs/getting-started/providers/perplexity-decider.md) (`PERPLEXITY_API_KEY`, the same key as its text provider), a hosted API that also reads images — the first one configured, in that order, runs. Used internally for model routing, context budgeting, relevance compaction and tool routing — fail-open and a no-op when no decision provider is configured. Per-query RAG planning is opt-in via `RAGPipeline`. | [Decide Guide](docs/features/decide-inference-type.md) |
65
- | **7 More Catalog Providers** | v12.11.0–v12.16.0 | Baseten, GMI Cloud, Inception Labs, io.net Intelligence, Mancer, Upstage and API Route onboarded as Tier-2 catalog entries — one JSON file each, roster live-verified against the provider's own `/v1/models`. | [Tier 2 Onboarding](docs/provider-integration/tiers/tier-2-catalog-entry.md) |
66
- | **Claude-on-Vertex Proxy Fallback** | v12.18.0 | The Anthropic proxy pool can fall back to Claude served on Google Vertex, so an agentic turn survives losing its primary backend mid-conversation instead of failing the turn. | [Claude Proxy](docs/features/claude-proxy.md) |
67
- | **Native-Loop V3 Conversation Reclaim** | v12.17.0 | Reclaims V3 conversations without splitting tool-call/tool-result pairs — the pairing a provider rejects the whole request over. | [Claude Proxy Architecture](docs/features/claude-proxy-architecture.md) |
68
- | **Multi-Modal Embeddings** | v12.15.0 | `embed()` / `embedMany()` accept images alongside text on providers whose embedding models are multi-modal, for cross-modal retrieval in RAG and custom vector search. | [Embeddings Guide](docs/features/embeddings.md) |
69
- | **Grok Build Auto-Configuration** | v12.14.0 | The proxy configures Grok Build automatically, deriving context windows and backends from the model catalog rather than hardcoded values. | [Proxy CLI Onboarding](docs/features/proxy-cli-onboarding.md) |
70
- | **Anthropic Execution-Control Contract** | v12.13.0 | Truthful stream termination plus an opt-in execution-control contract, so a stream that stopped early reports why instead of looking like a clean finish. | [Claude Proxy](docs/features/claude-proxy.md) |
71
- | **Catalog Tool Declarations Honoured at Runtime** | v12.12.0 | A Tier-2 catalog entry declaring `tools: false` (e.g. Mancer) no longer has tools offered to it at runtime — the JSON declaration is enforced, not just documented. | [Tier 2 Onboarding](docs/provider-integration/tiers/tier-2-catalog-entry.md) |
72
- | **Artifact Stores: Redis, Custom, Range Reads, Search** | v12.10.0 | Artifacts can be backed by Redis or a custom store, read by byte range, and searched — instead of being held only in process memory. | [Claude Proxy](docs/features/claude-proxy.md) |
73
- | **Local CLI Spend Reading** | v12.6.0–v12.9.0 | Reads token usage directly from other coding CLIs' own local stores — Cursor, Grok Build, Hermes Agent and three more — and names them in proxy traffic, so spend is attributed per client. | [Proxy CLI Onboarding](docs/features/proxy-cli-onboarding.md) |
74
- | **Native OpenAI Audio Streaming** | v12.7.0 | OpenAI TTS audio streams natively rather than being buffered to completion first. | [TTS Guide](docs/features/tts.md) |
75
- | **HITL Pending-Confirmation State** | v12.5.0 | Exposes whether a human-in-the-loop confirmation is still outstanding, so a caller can distinguish 'waiting on a human' from 'finished'. | [Task Manager](docs/features/task-manager.md) |
76
- | **OpenCode + Gemini CLI Proxy Clients** | v12.4.0 | OpenCode's generated config is actually loadable, and Gemini CLI is onboarded as a proxy client. | [OpenCode Proxy](docs/features/opencode-proxy-support.md) \| [Proxy CLI Onboarding](docs/features/proxy-cli-onboarding.md) |
77
- | **SambaNova Provider** | v12.3.0 | RDU-accelerated open-weight flagships: Llama 3.3 70B (default), GPT-OSS 120B, DeepSeek V3.x, MiniMax, Gemma 4 (vision) — OpenAI-compatible Tier 2 catalog entry. Note: new SambaNova accounts require purchased credits. | [SambaNova Guide](docs/getting-started/providers/sambanova.md) |
78
- | **Cerebras Provider** | v12.1.0 | Wafer-scale inference at ~3000 tok/s: GPT-OSS 120B (default) + Gemma 4 31B, OpenAI-compatible Tier 2 catalog entry, live-verified end to end (generate, stream, tools, structured output). | [Cerebras Guide](docs/getting-started/providers/cerebras.md) |
79
- | **Avatar / Music Modalities + 12 Providers** | v9.65.0 | New `output: { mode: "avatar" \| "music" }` dispatch with handlers for D-ID, HeyGen, Replicate-MuseTalk (avatar) and Beatoven, ElevenLabs Music, Lyria, Replicate-MusicGen (music). Plus Fish Audio TTS, Kling/Runway/Replicate video, xAI/Groq/Cohere/Together/Fireworks/Perplexity/Cloudflare LLMs, Voyage/Jina embeddings, Stability/Ideogram/Recraft/Replicate image-gen. | [Provider Integration](docs/provider-integration/) |
80
- | **Multi-Provider Voice (TTS/STT)** | v9.62.0 | 6 TTS providers (OpenAI TTS, ElevenLabs, Google TTS, Azure TTS, Fish Audio, Cartesia) + 4 STT providers (Whisper, Deepgram, Azure STT, Google STT) + 2 realtime APIs (OpenAI Realtime, Gemini Live). | [TTS Guide](docs/features/tts.md) \| [STT Guide](docs/features/audio-input.md) \| [Realtime Guide](docs/features/real-time-services.md) |
81
- | **4 New Providers** | v9.60.0 | DeepSeek (V3/R1), NVIDIA NIM (400+ catalog), LM Studio (local), llama.cpp (GGUF local). | [Provider Setup](docs/getting-started/provider-setup.md) |
82
- | **ModelAccessDeniedError** | v9.59.0 | Typed `ModelAccessDeniedError` + `sdk.checkCredentials()` API for proactive credential validation before first call. | [Error Reference](docs/reference/troubleshooting.md) |
83
- | **Provider Fallback Policy** | v9.58.0 | `providerFallback` callback + `modelChain` config for centralized multi-provider fallback logic. | [Advanced Guide](docs/advanced/index.md) |
84
- | **Per-Request Credentials** | v9.52.0 | Pass credentials per-call or per-instance for all providers. Per-call overrides instance; instance overrides env vars. | [Credentials Guide](docs/features/per-request-credentials.md) |
85
- | **AutoResearch** | v9.53.0 | Autonomous AI experiment engine: proposes code changes, runs experiments, evaluates metrics — unattended for hours. | [AutoResearch Guide](docs/features/autoresearch.md) |
86
- | **Gemini 3 Multi-turn Tool Fix** | v9.49.0 | Fixed multi-step agentic tool calling on Vertex AI Gemini 3. Correct `thoughtSignature` replay, `stepIndex` grouping, `executionId` session isolation, 5-min timeout. | [Vertex AI Guide](docs/getting-started/providers/google-vertex.md) |
87
- | **MCP Enhancements** | v9.16.0 | Tool routing (6 strategies), result caching (LRU/FIFO/LFU), request batching, annotations, elicitation protocol, multi-server management. | [MCP Enhancements Guide](docs/features/mcp-enhancements.md) |
88
- | **Memory** | v9.12.0 | Per-user condensed memory across conversations. LLM-powered condensation with S3, Redis, or SQLite. | [Memory Guide](docs/features/memory.md) |
89
- | **Context Window Management** | v9.2.0 | 5-stage compaction pipeline with budget gate at 80% usage, per-provider token estimation. | [Context Compaction Guide](docs/features/context-compaction.md) |
90
- | **Tool Execution Control** | v9.3.0 | `prepareStep` and `toolChoice` for per-step tool enforcement in multi-step agentic loops. | [API Reference](docs/api/type-aliases/GenerateOptions.md#preparestep) |
91
- | **File Processor System** | v9.1.0 | 17+ file type processors with ProcessorRegistry, security sanitization, SVG text injection. | [File Processors Guide](docs/features/file-processors.md) |
92
- | **RAG with generate()/stream()** | v9.2.0 | Pass `rag: { files }` for automatic document chunking, embedding, and AI-powered search. 10 chunking strategies, hybrid search, reranking, and a choice of 4 vector stores (in-memory, Chroma, PgVector, Pinecone). | [RAG Guide](docs/features/rag.md) |
62
+ | Feature | Version | Description | Guide |
63
+ | -------------------------------------------------------------------------------------- | ----------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- |
64
+ | **`decide` Inference Type + TypeSafe Jev + Laya + XOR + Perplexity + Cloudflare Clef** | next | A third inference type alongside `generate`/`stream`: typed, calibrated judgments (`boolean`, `choice`, `score`) via `neurolink.decide()` / `tryDecide()`, one parallel pass (~400ms and ~$0.00002/decision on Jev). Providers: TypeSafe Jev (`TYPESAFE_API_KEY`, also reachable via the Vercel AI Gateway); [Laya](docs/getting-started/providers/laya.md) (`LAYA_API_KEY` + `LAYA_BASE_URL`), an open-weights model you run yourself; [XOR](docs/getting-started/providers/xor.md) (`XOR_API_KEY` + `XOR_BASE_URL`), Juspay's open-weights model; [Perplexity](docs/getting-started/providers/perplexity-decider.md) (`PERPLEXITY_API_KEY`, the same key as its text provider), a hosted API that also reads images; and [Cloudflare Clef](docs/getting-started/providers/cloudflare-clef.md) (`CLOUDFLARE_API_KEY` + `CLOUDFLARE_ACCOUNT_ID`, the same two as the Workers AI text provider), hosted on Workers AI, which also reads images — the first one configured, in that order, runs. Used internally for model routing, context budgeting, relevance compaction and tool routing — fail-open and a no-op when no decision provider is configured. Per-query RAG planning is opt-in via `RAGPipeline`. | [Decide Guide](docs/features/decide-inference-type.md) |
65
+ | **7 More Catalog Providers** | v12.11.0–v12.16.0 | Baseten, GMI Cloud, Inception Labs, io.net Intelligence, Mancer, Upstage and API Route onboarded as Tier-2 catalog entries — one JSON file each, roster live-verified against the provider's own `/v1/models`. | [Tier 2 Onboarding](docs/provider-integration/tiers/tier-2-catalog-entry.md) |
66
+ | **Claude-on-Vertex Proxy Fallback** | v12.18.0 | The Anthropic proxy pool can fall back to Claude served on Google Vertex, so an agentic turn survives losing its primary backend mid-conversation instead of failing the turn. | [Claude Proxy](docs/features/claude-proxy.md) |
67
+ | **Native-Loop V3 Conversation Reclaim** | v12.17.0 | Reclaims V3 conversations without splitting tool-call/tool-result pairs — the pairing a provider rejects the whole request over. | [Claude Proxy Architecture](docs/features/claude-proxy-architecture.md) |
68
+ | **Multi-Modal Embeddings** | v12.15.0 | `embed()` / `embedMany()` accept images alongside text on providers whose embedding models are multi-modal, for cross-modal retrieval in RAG and custom vector search. | [Embeddings Guide](docs/features/embeddings.md) |
69
+ | **Grok Build Auto-Configuration** | v12.14.0 | The proxy configures Grok Build automatically, deriving context windows and backends from the model catalog rather than hardcoded values. | [Proxy CLI Onboarding](docs/features/proxy-cli-onboarding.md) |
70
+ | **Anthropic Execution-Control Contract** | v12.13.0 | Truthful stream termination plus an opt-in execution-control contract, so a stream that stopped early reports why instead of looking like a clean finish. | [Claude Proxy](docs/features/claude-proxy.md) |
71
+ | **Catalog Tool Declarations Honoured at Runtime** | v12.12.0 | A Tier-2 catalog entry declaring `tools: false` (e.g. Mancer) no longer has tools offered to it at runtime — the JSON declaration is enforced, not just documented. | [Tier 2 Onboarding](docs/provider-integration/tiers/tier-2-catalog-entry.md) |
72
+ | **Artifact Stores: Redis, Custom, Range Reads, Search** | v12.10.0 | Artifacts can be backed by Redis or a custom store, read by byte range, and searched — instead of being held only in process memory. | [Claude Proxy](docs/features/claude-proxy.md) |
73
+ | **Local CLI Spend Reading** | v12.6.0–v12.9.0 | Reads token usage directly from other coding CLIs' own local stores — Cursor, Grok Build, Hermes Agent and three more — and names them in proxy traffic, so spend is attributed per client. | [Proxy CLI Onboarding](docs/features/proxy-cli-onboarding.md) |
74
+ | **Native OpenAI Audio Streaming** | v12.7.0 | OpenAI TTS audio streams natively rather than being buffered to completion first. | [TTS Guide](docs/features/tts.md) |
75
+ | **HITL Pending-Confirmation State** | v12.5.0 | Exposes whether a human-in-the-loop confirmation is still outstanding, so a caller can distinguish 'waiting on a human' from 'finished'. | [Task Manager](docs/features/task-manager.md) |
76
+ | **OpenCode + Gemini CLI Proxy Clients** | v12.4.0 | OpenCode's generated config is actually loadable, and Gemini CLI is onboarded as a proxy client. | [OpenCode Proxy](docs/features/opencode-proxy-support.md) \| [Proxy CLI Onboarding](docs/features/proxy-cli-onboarding.md) |
77
+ | **SambaNova Provider** | v12.3.0 | RDU-accelerated open-weight flagships: Llama 3.3 70B (default), GPT-OSS 120B, DeepSeek V3.x, MiniMax, Gemma 4 (vision) — OpenAI-compatible Tier 2 catalog entry. Note: new SambaNova accounts require purchased credits. | [SambaNova Guide](docs/getting-started/providers/sambanova.md) |
78
+ | **Cerebras Provider** | v12.1.0 | Wafer-scale inference at ~3000 tok/s: GPT-OSS 120B (default) + Gemma 4 31B, OpenAI-compatible Tier 2 catalog entry, live-verified end to end (generate, stream, tools, structured output). | [Cerebras Guide](docs/getting-started/providers/cerebras.md) |
79
+ | **Avatar / Music Modalities + 12 Providers** | v9.65.0 | New `output: { mode: "avatar" \| "music" }` dispatch with handlers for D-ID, HeyGen, Replicate-MuseTalk (avatar) and Beatoven, ElevenLabs Music, Lyria, Replicate-MusicGen (music). Plus Fish Audio TTS, Kling/Runway/Replicate video, xAI/Groq/Cohere/Together/Fireworks/Perplexity/Cloudflare LLMs, Voyage/Jina embeddings, Stability/Ideogram/Recraft/Replicate image-gen. | [Provider Integration](docs/provider-integration/) |
80
+ | **Multi-Provider Voice (TTS/STT)** | v9.62.0 | 6 TTS providers (OpenAI TTS, ElevenLabs, Google TTS, Azure TTS, Fish Audio, Cartesia) + 4 STT providers (Whisper, Deepgram, Azure STT, Google STT) + 2 realtime APIs (OpenAI Realtime, Gemini Live). | [TTS Guide](docs/features/tts.md) \| [STT Guide](docs/features/audio-input.md) \| [Realtime Guide](docs/features/real-time-services.md) |
81
+ | **4 New Providers** | v9.60.0 | DeepSeek (V3/R1), NVIDIA NIM (400+ catalog), LM Studio (local), llama.cpp (GGUF local). | [Provider Setup](docs/getting-started/provider-setup.md) |
82
+ | **ModelAccessDeniedError** | v9.59.0 | Typed `ModelAccessDeniedError` + `sdk.checkCredentials()` API for proactive credential validation before first call. | [Error Reference](docs/reference/troubleshooting.md) |
83
+ | **Provider Fallback Policy** | v9.58.0 | `providerFallback` callback + `modelChain` config for centralized multi-provider fallback logic. | [Advanced Guide](docs/advanced/index.md) |
84
+ | **Per-Request Credentials** | v9.52.0 | Pass credentials per-call or per-instance for all providers. Per-call overrides instance; instance overrides env vars. | [Credentials Guide](docs/features/per-request-credentials.md) |
85
+ | **AutoResearch** | v9.53.0 | Autonomous AI experiment engine: proposes code changes, runs experiments, evaluates metrics — unattended for hours. | [AutoResearch Guide](docs/features/autoresearch.md) |
86
+ | **Gemini 3 Multi-turn Tool Fix** | v9.49.0 | Fixed multi-step agentic tool calling on Vertex AI Gemini 3. Correct `thoughtSignature` replay, `stepIndex` grouping, `executionId` session isolation, 5-min timeout. | [Vertex AI Guide](docs/getting-started/providers/google-vertex.md) |
87
+ | **MCP Enhancements** | v9.16.0 | Tool routing (6 strategies), result caching (LRU/FIFO/LFU), request batching, annotations, elicitation protocol, multi-server management. | [MCP Enhancements Guide](docs/features/mcp-enhancements.md) |
88
+ | **Memory** | v9.12.0 | Per-user condensed memory across conversations. LLM-powered condensation with S3, Redis, or SQLite. | [Memory Guide](docs/features/memory.md) |
89
+ | **Context Window Management** | v9.2.0 | 5-stage compaction pipeline with budget gate at 80% usage, per-provider token estimation. | [Context Compaction Guide](docs/features/context-compaction.md) |
90
+ | **Tool Execution Control** | v9.3.0 | `prepareStep` and `toolChoice` for per-step tool enforcement in multi-step agentic loops. | [API Reference](docs/api/type-aliases/GenerateOptions.md#preparestep) |
91
+ | **File Processor System** | v9.1.0 | 17+ file type processors with ProcessorRegistry, security sanitization, SVG text injection. | [File Processors Guide](docs/features/file-processors.md) |
92
+ | **RAG with generate()/stream()** | v9.2.0 | Pass `rag: { files }` for automatic document chunking, embedding, and AI-powered search. 10 chunking strategies, hybrid search, reranking, and a choice of 4 vector stores (in-memory, Chroma, PgVector, Pinecone). | [RAG Guide](docs/features/rag.md) |
93
93
 
94
94
  ```typescript
95
95
  // decide() — a third inference type: calibrated judgments, not text (next)
@@ -281,7 +281,7 @@ shorter state than TypeSafe's — about 768 tokens on the default
281
281
  `typed-decisions` checkpoint (320 on `english`/`auto`), against TypeSafe's
282
282
  ~33,000 — so it fits short, structured decisions rather than long context.
283
283
  Every built-in consumer below asks for the first configured decision provider,
284
- in the order TypeSafe, Laya, XOR, Perplexity: with TypeSafe and Laya both
284
+ in the order TypeSafe, Laya, XOR, Perplexity, Cloudflare Clef: with TypeSafe and Laya both
285
285
  configured, TypeSafe runs; with only Laya's key and base URL set, Laya runs.
286
286
 
287
287
  [**XOR**](docs/getting-started/providers/xor.md) is Juspay's Apache-2.0,
@@ -308,6 +308,23 @@ configure one of TypeSafe, Laya and XOR; the guide lists
308
308
  [what each consumer sends](docs/getting-started/providers/perplexity-decider.md#what-is-sent-to-perplexity)
309
309
  and [the exact switches](docs/getting-started/providers/perplexity-decider.md#one-key-two-providers).
310
310
 
311
+ [**Cloudflare Clef**](docs/getting-started/providers/cloudflare-clef.md) is
312
+ Cloudflare's hosted decision models on Workers AI (`clef`, 27B, the default, and
313
+ `clef-flash`, 9B; provider id `cloudflare-clef`) and also reads images, up to 4,
314
+ but no video. It has a public endpoint, so credentials alone configure it:
315
+ `CLOUDFLARE_API_KEY` (Workers AI permission) **and** `CLOUDFLARE_ACCOUNT_ID`, or
316
+ `credentials.cloudflareClef`. **Those are the same two variables the Workers AI
317
+ text provider (`cloudflare`) reads**, so setting them for it also configures
318
+ `decide`: built-in features then use Clef whenever none of TypeSafe, Laya, XOR or
319
+ Perplexity is configured, and it never displaces one that is. No switch turns
320
+ that off while the two variables are set, and `credentials.cloudflare` does not
321
+ configure `decide`. **The Workers AI endpoint ignores text past about 2,048
322
+ tokens without an error (hosted service or model: unknown)**; NeuroLink refuses a state it estimates at more
323
+ than 1,500 tokens with `max_tokens_exceeded`, which each built-in consumer treats
324
+ as "carry on as before". The guide covers
325
+ [the limits](docs/getting-started/providers/cloudflare-clef.md#limits) and
326
+ [when NeuroLink uses it](docs/getting-started/providers/cloudflare-clef.md#when-neurolink-uses-it).
327
+
311
328
  ```typescript
312
329
  import {
313
330
  NeuroLink,
@@ -397,7 +414,15 @@ The figures above are TypeSafe's. Perplexity's Decisions API, measured on a real
397
414
 
398
415
  The [guide's limits section](docs/getting-started/providers/perplexity-decider.md#limits) has the full table, including tokens per character and the image-size rule.
399
416
 
400
- **[Decide Guide](docs/features/decide-inference-type.md)** · **[TypeSafe Provider Guide](docs/getting-started/providers/typesafe.md)** · **[Laya Provider Guide](docs/getting-started/providers/laya.md)** · **[XOR Provider Guide](docs/getting-started/providers/xor.md)** · **[Perplexity Provider Guide](docs/getting-started/providers/perplexity-decider.md)**
417
+ Cloudflare Clef, measured on a real account in October 2026: the follow-up of 133 probe calls checked structured-text cuts and four-image/JPG cases on both models; ids, options, score levels, formats and natural scripts were checked on `clef` only. The guide records each model and date.
418
+
419
+ - The Workers AI endpoint ignores state text past about 2,048 tokens without an error (hosted service or model: unknown; Cloudflare documents 64K). NeuroLink refuses a state it estimates over 1,500 tokens (digits count 1 token each, punctuation 0.75, emoji 3, other non-ASCII text 1.5 per character, other text about 4 characters a token) with `max_tokens_exceeded`.
420
+ - 64 questions and 4 images per request (both Cloudflare's documented limits; the 65th question and a fifth image were refused; `tryDecide()` splits a larger question map into batches of 64). NeuroLink caps the encoded request at 256,000 bytes. On 2026-10-03, `clef-flash` accepted 262,000 text characters and refused 270,000. On 2026-10-04 both models accepted 520,000 text characters and refused 525,000 with 413/code 5021; the smaller local cap is retained. Exactly four images and the `image/jpg` alias were also accepted on both models.
421
+ - On 2026-10-03, latency was 0.3 to 1.0 s for a small request and 1.1 s (`clef-flash`) / 1.3 s (`clef`) for 64 questions. A 64-question request with the same questions and a shorter state took 1.5 s / 2.3 s on 2026-10-04. The price is $0.24 per million input tokens for `clef` and $0.09 for `clef-flash`; no output price is listed on Cloudflare's Workers AI pricing page.
422
+
423
+ The [guide's limits section](docs/getting-started/providers/cloudflare-clef.md#limits) has the full table, including tokens per character for each kind of text.
424
+
425
+ **[Decide Guide](docs/features/decide-inference-type.md)** · **[TypeSafe Provider Guide](docs/getting-started/providers/typesafe.md)** · **[Laya Provider Guide](docs/getting-started/providers/laya.md)** · **[XOR Provider Guide](docs/getting-started/providers/xor.md)** · **[Perplexity Provider Guide](docs/getting-started/providers/perplexity-decider.md)** · **[Cloudflare Clef Provider Guide](docs/getting-started/providers/cloudflare-clef.md)**
401
426
 
402
427
  ## Enterprise Security: Human-in-the-Loop (HITL)
403
428
 
@@ -676,7 +701,7 @@ NeuroLink is a comprehensive AI development platform. Every feature below is shi
676
701
 
677
702
  ### 🤖 AI Provider Integration
678
703
 
679
- **Provider neurons behind one API** - Switch providers with a single parameter change. Nearly all serve `generate`/`stream`; TypeSafe Jev, Laya, XOR and Perplexity Decisions serve `decide` instead. Tool support varies by provider and model (see the provider catalog); embedding-, media- and decision-only providers serve no tools. 3 are fully local runtimes (Ollama, LM Studio, llama.cpp) and need no cloud account or API key. LiteLLM needs no NeuroLink env vars (it defaults to `localhost:4000`) but requires a running LiteLLM proxy, which holds the upstream provider keys for hosted models. 9 providers (OpenAI, Google AI Studio, Google Vertex, Amazon Bedrock, Cohere, Ollama, LiteLLM, Voyage, Jina) expose `embed()`/`embedMany()` natively for RAG and custom vector search.
704
+ **Provider neurons behind one API** - Switch providers with a single parameter change. Nearly all serve `generate`/`stream`; TypeSafe Jev, Laya, XOR, Perplexity Decisions and Cloudflare Clef serve `decide` instead. Tool support varies by provider and model (see the provider catalog); embedding-, media- and decision-only providers serve no tools. 3 are fully local runtimes (Ollama, LM Studio, llama.cpp) and need no cloud account or API key. LiteLLM needs no NeuroLink env vars (it defaults to `localhost:4000`) but requires a running LiteLLM proxy, which holds the upstream provider keys for hosted models. 9 providers (OpenAI, Google AI Studio, Google Vertex, Amazon Bedrock, Cohere, Ollama, LiteLLM, Voyage, Jina) expose `embed()`/`embedMany()` natively for RAG and custom vector search.
680
705
 
681
706
  | Provider | Models | Free Tier | Tool Support | Status | Documentation |
682
707
  | --------------------- | -------------------------------------------------------------------------- | --------------- | ------------ | ------------- | ----------------------------------------------------------------------------------------------------------------------------- |
@@ -710,9 +735,9 @@ NeuroLink is a comprehensive AI development platform. Every feature below is shi
710
735
 
711
736
  **Media generation** — [Replicate](docs/getting-started/providers/replicate.md) (`REPLICATE_API_TOKEN`) · [Stability AI](docs/getting-started/providers/stability.md) (`STABILITY_API_KEY`) · [Ideogram](docs/getting-started/providers/ideogram.md) (`IDEOGRAM_API_KEY`) · [Recraft](docs/getting-started/providers/recraft.md) (`RECRAFT_API_KEY`)
712
737
 
713
- **Decision** — [TypeSafe Jev](docs/getting-started/providers/typesafe.md) (`TYPESAFE_API_KEY`, or `AI_GATEWAY_API_KEY` via the Vercel AI Gateway) · [Laya](docs/getting-started/providers/laya.md) (`LAYA_API_KEY` + `LAYA_BASE_URL`, open-weights, self-hosted) · [XOR](docs/getting-started/providers/xor.md) (`XOR_API_KEY` + `XOR_BASE_URL`, open-weights) · [Perplexity Decisions](docs/getting-started/providers/perplexity-decider.md) (`PERPLEXITY_API_KEY`, hosted; the same key as the Perplexity text provider) — the providers serving `decide` rather than `generate`/`stream`.
738
+ **Decision** — [TypeSafe Jev](docs/getting-started/providers/typesafe.md) (`TYPESAFE_API_KEY`, or `AI_GATEWAY_API_KEY` via the Vercel AI Gateway) · [Laya](docs/getting-started/providers/laya.md) (`LAYA_API_KEY` + `LAYA_BASE_URL`, open-weights, self-hosted) · [XOR](docs/getting-started/providers/xor.md) (`XOR_API_KEY` + `XOR_BASE_URL`, open-weights) · [Perplexity Decisions](docs/getting-started/providers/perplexity-decider.md) (`PERPLEXITY_API_KEY`, hosted; the same key as the Perplexity text provider) · [Cloudflare Clef](docs/getting-started/providers/cloudflare-clef.md) (`CLOUDFLARE_API_KEY` + `CLOUDFLARE_ACCOUNT_ID`, hosted on Workers AI; the same two variables as the Cloudflare text provider) — the providers serving `decide` rather than `generate`/`stream`.
714
739
 
715
- **Decision-only providers:** **TypeSafe Jev** (`TYPESAFE_API_KEY`), **Laya** (`LAYA_API_KEY` + `LAYA_BASE_URL`), **XOR** (`XOR_API_KEY` + `XOR_BASE_URL`) and **Perplexity Decisions** (`PERPLEXITY_API_KEY`) do not appear in the table above because none of them serves `generate`/`stream` — they are the providers for the `decide` inference type, tried in the order TypeSafe, Laya, XOR, Perplexity when several are configured. Perplexity's key is shared with the Perplexity text provider, which does appear in the table. See [Decide: Calibrated Judgments, Not Text](#decide-calibrated-judgments-not-text).
740
+ **Decision-only providers:** **TypeSafe Jev** (`TYPESAFE_API_KEY`), **Laya** (`LAYA_API_KEY` + `LAYA_BASE_URL`), **XOR** (`XOR_API_KEY` + `XOR_BASE_URL`), **Perplexity Decisions** (`PERPLEXITY_API_KEY`) and **Cloudflare Clef** (`CLOUDFLARE_API_KEY` + `CLOUDFLARE_ACCOUNT_ID`) do not appear in the table above because none of them serves `generate`/`stream` — they are the providers for the `decide` inference type, tried in the order TypeSafe, Laya, XOR, Perplexity, Cloudflare Clef when several are configured. Perplexity's key is shared with the Perplexity text provider, and Clef's two variables with the Cloudflare Workers AI text provider; both text providers appear in the table. See [Decide: Calibrated Judgments, Not Text](#decide-calibrated-judgments-not-text).
716
741
 
717
742
  **[📖 Provider Comparison Guide](docs/reference/provider-comparison.md)** - Detailed feature matrix and selection criteria
718
743
  **[🔬 Provider Feature Compatibility](docs/reference/provider-feature-compatibility.md)** - Test-based compatibility reference for 19 features (dated snapshot covering a subset of the full provider list)
@@ -1234,18 +1259,18 @@ Full command and API breakdown lives in [`docs/cli/commands.md`](docs/cli/comman
1234
1259
 
1235
1260
  ## Platform Capabilities at a Glance
1236
1261
 
1237
- | Capability | Highlights |
1238
- | ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
1239
- | **Provider unification** | Provider neurons behind one API, with automatic fallback, cost-aware routing, `providerFallback` policy, `modelChain` config. |
1240
- | **Decision inference** | Third inference type (`decide`) alongside generate/stream: calibrated `boolean`/`choice`/`score` judgments via TypeSafe Jev (~400ms flat, ~$0.00002/decision), Laya, a self-hosted open-weights alternative, XOR, Juspay's open-weights model, or Perplexity's hosted Decisions API. Used internally for model routing, context budgeting, relevance compaction and tool routing; per-query RAG planning is opt-in via `RAGPipeline`. |
1241
- | **Multimodal pipeline** | Stream images + CSV data + PDF documents across providers with local/remote assets. Auto-detection for mixed file types. |
1242
- | **Voice pipeline** | TTS (6 providers: Google, OpenAI, ElevenLabs, Azure, Fish Audio, Cartesia) + STT (4 providers) + realtime voice APIs (OpenAI Realtime, Gemini Live). |
1243
- | **Quality & governance** | Auto-evaluation engine (14 scorers), guardrails middleware, HITL workflows, audit logging. |
1244
- | **Memory & context** | Per-user condensed memory (S3/Redis/SQLite), Redis session export, 5-stage context compaction. |
1245
- | **CLI tooling** | 34 commands: loop sessions, setup wizard, config validation, Redis auto-detect, JSON output, TTS/STT flags. |
1246
- | **Enterprise ops** | Claude proxy, OTLP observability, OpenObserve dashboard, regional routing, credential management. |
1247
- | **Tool ecosystem** | MCP auto discovery, HTTP/stdio/SSE/WebSocket transports, LiteLLM hub access, SageMaker custom deployment, web search. |
1248
- | **Engineering rigor** | 129 end-to-end test suites (every suite drives the public `generate`/`stream`/`decide`/CLI surface, never internals), 13 custom ESLint rules enforcing the architecture (no `interface`, unique type names, barrel-only type imports) — all AST-based, no regex heuristics. |
1262
+ | Capability | Highlights |
1263
+ | ------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
1264
+ | **Provider unification** | Provider neurons behind one API, with automatic fallback, cost-aware routing, `providerFallback` policy, `modelChain` config. |
1265
+ | **Decision inference** | Third inference type (`decide`) alongside generate/stream: calibrated `boolean`/`choice`/`score` judgments via TypeSafe Jev (~400ms flat, ~$0.00002/decision), Laya, a self-hosted open-weights alternative, XOR, Juspay's open-weights model, Perplexity's hosted Decisions API, or Cloudflare Clef on Workers AI. Used internally for model routing, context budgeting, relevance compaction and tool routing; per-query RAG planning is opt-in via `RAGPipeline`. |
1266
+ | **Multimodal pipeline** | Stream images + CSV data + PDF documents across providers with local/remote assets. Auto-detection for mixed file types. |
1267
+ | **Voice pipeline** | TTS (6 providers: Google, OpenAI, ElevenLabs, Azure, Fish Audio, Cartesia) + STT (4 providers) + realtime voice APIs (OpenAI Realtime, Gemini Live). |
1268
+ | **Quality & governance** | Auto-evaluation engine (14 scorers), guardrails middleware, HITL workflows, audit logging. |
1269
+ | **Memory & context** | Per-user condensed memory (S3/Redis/SQLite), Redis session export, 5-stage context compaction. |
1270
+ | **CLI tooling** | 34 commands: loop sessions, setup wizard, config validation, Redis auto-detect, JSON output, TTS/STT flags. |
1271
+ | **Enterprise ops** | Claude proxy, OTLP observability, OpenObserve dashboard, regional routing, credential management. |
1272
+ | **Tool ecosystem** | MCP auto discovery, HTTP/stdio/SSE/WebSocket transports, LiteLLM hub access, SageMaker custom deployment, web search. |
1273
+ | **Engineering rigor** | 129 end-to-end test suites (every suite drives the public `generate`/`stream`/`decide`/CLI surface, never internals), 13 custom ESLint rules enforcing the architecture (no `interface`, unique type names, barrel-only type imports) — all AST-based, no regex heuristics. |
1249
1274
 
1250
1275
  ## Documentation Map
1251
1276
 
@@ -1273,7 +1298,7 @@ Full command and API breakdown lives in [`docs/cli/commands.md`](docs/cli/comman
1273
1298
 
1274
1299
  **Decision Inference:**
1275
1300
 
1276
- - [Decide Guide](docs/features/decide-inference-type.md) - The `decide` inference type: boolean/choice/score primitives, TypeSafe Jev / Laya / XOR / Perplexity setup, measured latency/cost
1301
+ - [Decide Guide](docs/features/decide-inference-type.md) - The `decide` inference type: boolean/choice/score primitives, TypeSafe Jev / Laya / XOR / Perplexity / Cloudflare Clef setup, measured latency/cost
1277
1302
 
1278
1303
  **Provider Intelligence:**
1279
1304