@juspay/neurolink 12.46.1 → 12.47.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (37) hide show
  1. package/CHANGELOG.md +2 -2
  2. package/README.md +75 -50
  3. package/dist/adapters/video/ffmpegAdapter.d.ts +3 -1
  4. package/dist/adapters/video/ffmpegAdapter.js +13 -2
  5. package/dist/browser/neurolink.min.js +421 -421
  6. package/dist/cli/commands/decide.js +2 -2
  7. package/dist/cli/commands/setup.js +8 -6
  8. package/dist/cli/factories/commandFactory.js +1 -1
  9. package/dist/cli/proxy-clients/copilot.js +55 -1
  10. package/dist/constants/contextWindows.js +12 -0
  11. package/dist/constants/enums.d.ts +16 -0
  12. package/dist/constants/enums.js +17 -0
  13. package/dist/core/toolExecutionGuards.d.ts +1 -9
  14. package/dist/core/toolExecutionGuards.js +1 -9
  15. package/dist/factories/providerDescriptors.js +84 -5
  16. package/dist/factories/providerFactory.js +16 -1
  17. package/dist/factories/providerRegistry.js +10 -1
  18. package/dist/files/fileReferenceRegistry.js +3 -3
  19. package/dist/models/manifestRegistry.js +2 -0
  20. package/dist/models/manifests/cloudflareClef.d.ts +16 -0
  21. package/dist/models/manifests/cloudflareClef.js +42 -0
  22. package/dist/neurolink.js +13 -2
  23. package/dist/processors/media/VideoProcessor.js +16 -4
  24. package/dist/providers/cloudflareClef.d.ts +52 -0
  25. package/dist/providers/cloudflareClef.js +331 -0
  26. package/dist/providers/googleNativeGemini3/utils.d.ts +9 -0
  27. package/dist/providers/googleNativeGemini3/utils.js +9 -0
  28. package/dist/providers/systemOneDecision.d.ts +12 -1
  29. package/dist/providers/systemOneDecision.js +70 -18
  30. package/dist/types/decision.d.ts +22 -0
  31. package/dist/types/providers.d.ts +15 -0
  32. package/dist/utils/modelChoices.js +13 -1
  33. package/dist/utils/pricing.js +12 -0
  34. package/dist/utils/providerConfig.d.ts +7 -0
  35. package/dist/utils/providerConfig.js +19 -0
  36. package/docs-site/static/search-index.json +78 -56
  37. package/package.json +2 -1
@@ -114,9 +114,9 @@
114
114
  {"objectID":"b38e417fcd8ea5d3b983875efebb5b5a309883fdfe8bfdc80300e473bea5e081","title":"Gradual Adoption","url":"/docs/WORKFLOW-ENGINE-LLD#gradual-adoption","content":"Phase 1: Users can try workflows alongside existing methods\nPhase 2: Workflows become recommended for high-stakes queries\nPhase 3: Workflows are default with single-model as fallback","hierarchy":{"lvl0":"WORKFLOW ENGINE LLD","lvl1":"Workflow Engine - Low-Level Design","lvl2":"Gradual Adoption","lvl3":""}},
115
115
  {"objectID":"2127ec4796813559375c700e37b4944f77c72f22a0410fe93bd6f5f7cff8a4c0","title":"17. Performance Benchmarks (Expected)","url":"/docs/WORKFLOW-ENGINE-LLD#17-performance-benchmarks-expected","content":"| Workflow | Models | Judge | Latency (p50) | Latency (p95) | Cost Multiplier |\n| ------------- | ------ | ----- | ------------- | ------------- | --------------- |\n| consensus-3 | 3 | 1 | 3.2s | 5.1s | 4.2x |\n| fast-fallback | 1-2 | 0 | 1.1s | 2.8s | 1.3x |\n| quality-max | 2 | 1 | 3.5s | 4.9s | 3.1x |\n| multi-judge-5 | 3 | 2 | 4.8s | 6.7s | 5.3x |","hierarchy":{"lvl0":"WORKFLOW ENGINE LLD","lvl1":"Workflow Engine - Low-Level Design","lvl2":"17. Performance Benchmarks (Expected)","lvl3":""}},
116
116
  {"objectID":"2f54303255393f1b5c8695c6342964ce83cb786fd3ae00f535bf17c44cae16cd","title":"šŸ“ Implementation Checklist","url":"/docs/WORKFLOW-ENGINE-LLD#-implementation-checklist","content":"[ ] Create src/workflow/ directory structure\n[ ] Implement types.ts with all interfaces\n[ ] Implement config.ts with Zod schemas\n[ ] Implement ensembleExecutor.ts\n[ ] Implement judgeScorer.ts\n[ ] Implement responseConditioner.ts\n[ ] Implement workflowRegistry.ts\n[ ] Implement workflowRunner.ts\n[ ] Create built-in workflows (consensus, fallback, quality-max)\n[ ] Add methods to NeuroLink class\n[ ] Export types from src/lib/index.ts\n[ ] Write unit tests (80% coverage target)\n[ ] Write integration tests\n[ ] Add JSDoc documentation\n[ ] Create user guide with examples\n[ ] Add CLI support (optional Phase 2)\n\nDocument Status: āœ… Ready for Implementation \nNext Step: Code generation upon approval","hierarchy":{"lvl0":"WORKFLOW ENGINE LLD","lvl1":"Workflow Engine - Low-Level Design","lvl2":"šŸ“ Implementation Checklist","lvl3":""}},
117
- {"objectID":"de72b620d608319199e3abf25b54298afc8fe5f294c73c538956a2feb63fe9b1","title":"The Nervous System Model","url":"/docs/about/nervous-system-model","content":"The Nervous System Model\n\nNeuroLink is built around a biological metaphor — not as decoration, but as a structural model that governs every architectural decision.\n\nThe Three Components\n\nNeurons — LLM Providers\n\nNeurons are where intelligence is generated. In NeuroLink, neurons are the AI providers, including: Anthropic, OpenAI, Google (AI Studio + Vertex), AWS (Bedrock + SageMaker), Azure, Mistral, LiteLLM, OpenRouter, Ollama, Hugging Face, DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, OpenAI-compatible endpoints, TypeSafe Jev, Laya, XOR and Perplexity Decisions (decision-only) — plus voice neurons (OpenAI TTS, ElevenLabs, Google TTS, Azure TTS, Whisper, Deepgram, Azure STT, Google STT), realtime neurons (OpenAI Realtime, Gemini Live), and media-generation neurons (image, video, music, avatar).\n\nEach provider is a different type of neuron — different capabilities, different costs, different latency profiles. NeuroLink's ProviderRegistry gives you access to all of them through one interface, switchable with a single line.\n\nThe Pipe — NeuroLink\n\nThe pipe is the vascular layer that carries streams between neurons and organs. This is NeuroLink itself.\n\nWhat the pipe does every time you call generate() or stream():\nContext Building — RAG retrieval, memory lookup, file processing merge into the prompt\nBudget Check — BudgetChecker validates the assembled context fits the model's window\nProvider Dispatch — ProviderRegistry routes to the correct neuron\nStream Emission — Tokens flow as an async iterable\nTool Interception — When the model calls a tool, the stream pauses, MCP tool executes, result injects, stream continues\nObservability — Every stage emits OpenTelemetry spans\n\nOrgans — Connectors\n\nOrgans are the applications that consume the pipe. They connect to the vascular layer and open a gateway — a specific way for people or systems to interact with AI.\n\nEvery application built on NeuroLink is an organ. Production organs today:\nAutomatic — Shopify operations hub: address intelligence, RTO risk scoring\nTara — Slack engineering assistant: conversational AI with MCP tool access\nYama — Code review judge: automated PR analysis and governance\n\nWhy This Model Works\n\nThe metaphor enforces good architecture:\n\nSeparation of concerns: Neurons (generation) and organs (consumption) are completely decoupled. Changing AI provider doesn't touch the application. Changing the application doesn't touch the provider.\n\nSingle flow direction: Intelligence flows one way — neuron → pipe → organ. There's no confusion about where logic lives.\n\nObservable by default: A vascular system you can't monitor is dangerous. Every stage of the pipe emits telemetry by design.\n\nComposable: Multiple organs can share the same pipe. One NeuroLink instance serves many connectors.\n\nExtending the System\n\nThe nervous system model scales in three directions:\nAdd neurons — New AI provider? Register it in ProviderRegistry.\nExtend the pipe — New capability (chunking strategy, reranker, compaction stage)? Add it to the pipeline.\nBuild organs — New application? Import NeuroLink, connect to the pipe, open your gateway.\n\nSee Pipe Architecture → for the technical implementation.","hierarchy":{"lvl0":"About","lvl1":"The Nervous System Model","lvl2":"","lvl3":""}},
117
+ {"objectID":"de72b620d608319199e3abf25b54298afc8fe5f294c73c538956a2feb63fe9b1","title":"The Nervous System Model","url":"/docs/about/nervous-system-model","content":"The Nervous System Model\n\nNeuroLink is built around a biological metaphor — not as decoration, but as a structural model that governs every architectural decision.\n\nThe Three Components\n\nNeurons — LLM Providers\n\nNeurons are where intelligence is generated. In NeuroLink, neurons are the AI providers, including: Anthropic, OpenAI, Google (AI Studio + Vertex), AWS (Bedrock + SageMaker), Azure, Mistral, LiteLLM, OpenRouter, Ollama, Hugging Face, DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, OpenAI-compatible endpoints, TypeSafe Jev, Laya, XOR, Perplexity Decisions and Cloudflare Clef (decision-only) — plus voice neurons (OpenAI TTS, ElevenLabs, Google TTS, Azure TTS, Whisper, Deepgram, Azure STT, Google STT), realtime neurons (OpenAI Realtime, Gemini Live), and media-generation neurons (image, video, music, avatar).\n\nEach provider is a different type of neuron — different capabilities, different costs, different latency profiles. NeuroLink's ProviderRegistry gives you access to all of them through one interface, switchable with a single line.\n\nThe Pipe — NeuroLink\n\nThe pipe is the vascular layer that carries streams between neurons and organs. This is NeuroLink itself.\n\nWhat the pipe does every time you call generate() or stream():\nContext Building — RAG retrieval, memory lookup, file processing merge into the prompt\nBudget Check — BudgetChecker validates the assembled context fits the model's window\nProvider Dispatch — ProviderRegistry routes to the correct neuron\nStream Emission — Tokens flow as an async iterable\nTool Interception — When the model calls a tool, the stream pauses, MCP tool executes, result injects, stream continues\nObservability — Every stage emits OpenTelemetry spans\n\nOrgans — Connectors\n\nOrgans are the applications that consume the pipe. They connect to the vascular layer and open a gateway — a specific way for people or systems to interact with AI.\n\nEvery application built on NeuroLink is an organ. Production organs today:\nAutomatic — Shopify operations hub: address intelligence, RTO risk scoring\nTara — Slack engineering assistant: conversational AI with MCP tool access\nYama — Code review judge: automated PR analysis and governance\n\nWhy This Model Works\n\nThe metaphor enforces good architecture:\n\nSeparation of concerns: Neurons (generation) and organs (consumption) are completely decoupled. Changing AI provider doesn't touch the application. Changing the application doesn't touch the provider.\n\nSingle flow direction: Intelligence flows one way — neuron → pipe → organ. There's no confusion about where logic lives.\n\nObservable by default: A vascular system you can't monitor is dangerous. Every stage of the pipe emits telemetry by design.\n\nComposable: Multiple organs can share the same pipe. One NeuroLink instance serves many connectors.\n\nExtending the System\n\nThe nervous system model scales in three directions:\nAdd neurons — New AI provider? Register it in ProviderRegistry.\nExtend the pipe — New capability (chunking strategy, reranker, compaction stage)? Add it to the pipeline.\nBuild organs — New application? Import NeuroLink, connect to the pipe, open your gateway.\n\nSee Pipe Architecture → for the technical implementation.","hierarchy":{"lvl0":"About","lvl1":"The Nervous System Model","lvl2":"","lvl3":""}},
118
118
  {"objectID":"27ef114b7e0cecc6ca45dbf3e9285980b691cd27d21e1912f2a27a7f309f68fe","title":"The Nervous System Model","url":"/docs/about/nervous-system-model#the-nervous-system-model","content":"NeuroLink is built around a biological metaphor — not as decoration, but as a structural model that governs every architectural decision.","hierarchy":{"lvl0":"About","lvl1":"The Nervous System Model","lvl2":"The Nervous System Model","lvl3":""}},
119
- {"objectID":"58e5ab0de0df2a0949ae12798cd53ac3133b2c253880bcbf0c573938376fd84f","title":"Neurons — LLM Providers","url":"/docs/about/nervous-system-model#neurons-llm-providers","content":"Neurons are where intelligence is generated. In NeuroLink, neurons are the AI providers, including: Anthropic, OpenAI, Google (AI Studio + Vertex), AWS (Bedrock + SageMaker), Azure, Mistral, LiteLLM, OpenRouter, Ollama, Hugging Face, DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, OpenAI-compatible endpoints, TypeSafe Jev, Laya, XOR and Perplexity Decisions (decision-only) — plus voice neurons (OpenAI TTS, ElevenLabs, Google TTS, Azure TTS, Whisper, Deepgram, Azure STT, Google STT), realtime neurons (OpenAI Realtime, Gemini Live), and media-generation neurons (image, video, music, avatar).\n\nEach provider is a different type of neuron — different capabilities, different costs, different latency profiles. NeuroLink's ProviderRegistry gives you access to all of them through one interface, switchable with a single line.","hierarchy":{"lvl0":"About","lvl1":"The Nervous System Model","lvl2":"Neurons — LLM Providers","lvl3":""}},
119
+ {"objectID":"58e5ab0de0df2a0949ae12798cd53ac3133b2c253880bcbf0c573938376fd84f","title":"Neurons — LLM Providers","url":"/docs/about/nervous-system-model#neurons-llm-providers","content":"Neurons are where intelligence is generated. In NeuroLink, neurons are the AI providers, including: Anthropic, OpenAI, Google (AI Studio + Vertex), AWS (Bedrock + SageMaker), Azure, Mistral, LiteLLM, OpenRouter, Ollama, Hugging Face, DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, OpenAI-compatible endpoints, TypeSafe Jev, Laya, XOR, Perplexity Decisions and Cloudflare Clef (decision-only) — plus voice neurons (OpenAI TTS, ElevenLabs, Google TTS, Azure TTS, Whisper, Deepgram, Azure STT, Google STT), realtime neurons (OpenAI Realtime, Gemini Live), and media-generation neurons (image, video, music, avatar).\n\nEach provider is a different type of neuron — different capabilities, different costs, different latency profiles. NeuroLink's ProviderRegistry gives you access to all of them through one interface, switchable with a single line.","hierarchy":{"lvl0":"About","lvl1":"The Nervous System Model","lvl2":"Neurons — LLM Providers","lvl3":""}},
120
120
  {"objectID":"96f94c64476fe5567b61e5374da6d1b5c5020e22654c6c082b8936ae3f72eb51","title":"The Pipe — NeuroLink","url":"/docs/about/nervous-system-model#the-pipe-neurolink","content":"The pipe is the vascular layer that carries streams between neurons and organs. This is NeuroLink itself.\n\nWhat the pipe does every time you call generate() or stream():\nContext Building — RAG retrieval, memory lookup, file processing merge into the prompt\nBudget Check — BudgetChecker validates the assembled context fits the model's window\nProvider Dispatch — ProviderRegistry routes to the correct neuron\nStream Emission — Tokens flow as an async iterable\nTool Interception — When the model calls a tool, the stream pauses, MCP tool executes, result injects, stream continues\nObservability — Every stage emits OpenTelemetry spans","hierarchy":{"lvl0":"About","lvl1":"The Nervous System Model","lvl2":"The Pipe — NeuroLink","lvl3":""}},
121
121
  {"objectID":"1c74200e8d4572e84dca981b50bb162912b4c6700ea7495d60fe286cfb16347f","title":"Organs — Connectors","url":"/docs/about/nervous-system-model#organs-connectors","content":"Organs are the applications that consume the pipe. They connect to the vascular layer and open a gateway — a specific way for people or systems to interact with AI.\n\nEvery application built on NeuroLink is an organ. Production organs today:\nAutomatic — Shopify operations hub: address intelligence, RTO risk scoring\nTara — Slack engineering assistant: conversational AI with MCP tool access\nYama — Code review judge: automated PR analysis and governance","hierarchy":{"lvl0":"About","lvl1":"The Nervous System Model","lvl2":"Organs — Connectors","lvl3":""}},
122
122
  {"objectID":"6c97c4e15e0de12e62fbb1994cc7b18328c9f2b0afa2b721d8f40f01e70d32f7","title":"Why This Model Works","url":"/docs/about/nervous-system-model#why-this-model-works","content":"The metaphor enforces good architecture:\n\nSeparation of concerns: Neurons (generation) and organs (consumption) are completely decoupled. Changing AI provider doesn't touch the application. Changing the application doesn't touch the provider.\n\nSingle flow direction: Intelligence flows one way — neuron → pipe → organ. There's no confusion about where logic lives.\n\nObservable by default: A vascular system you can't monitor is dangerous. Every stage of the pipe emits telemetry by design.\n\nComposable: Multiple organs can share the same pipe. One NeuroLink instance serves many connectors.","hierarchy":{"lvl0":"About","lvl1":"The Nervous System Model","lvl2":"Why This Model Works","lvl3":""}},
@@ -784,7 +784,7 @@
784
784
  {"objectID":"6df70184b50cfe765cb953c7f99da86bcc5085d7c1bd9c009ad1d936cd313d54","title":"decide [state]","url":"/docs/cli/commands#decide","content":"Get typed, calibrated judgements from a decision model — the decide inference type, alongside generate and stream. Takes one state plus a map of named typed questions and prints one answer per question; there is no free text anywhere in the response.\n\n`bash","hierarchy":{"lvl0":"Cli","lvl1":"CLI Command Reference","lvl2":"decide [state]","lvl3":""}},
785
785
  {"objectID":"0076ed0bb97c3c91b6019b1168d3de94687353bc588fc474812ff9a55bf325d3","title":"State as a positional argument","url":"/docs/cli/commands#state-as-a-positional-argument","content":"npx @juspay/neurolink decide \"Refund request for a damaged item\" \\\n --questions '{\"urgent\":{\"type\":\"boolean\",\"instructions\":\"Is this urgent?\"}}'","hierarchy":{"lvl0":"Cli","lvl1":"CLI Command Reference","lvl2":"State as a positional argument","lvl3":""}},
786
786
  {"objectID":"e7605e932a3b0b4ce109cdb2cbb0d1d5012cb18891ee886ed14d4873544bd69e","title":"State and questions from files, raw JSON output","url":"/docs/cli/commands#state-and-questions-from-files-raw-json-output","content":"npx @juspay/neurolink decide --state-file ticket.json \\\n --questions-file questions.json --format json","hierarchy":{"lvl0":"Cli","lvl1":"CLI Command Reference","lvl2":"State and questions from files, raw JSON output","lvl3":""}},
787
- {"objectID":"756842dd6795d63598c25343fd2ca0b535c5880ab387395525d1fa15dad33f5d","title":"Ask about an image (XOR and Perplexity read images); repeat --image for several, or use --video (XOR)","url":"/docs/cli/commands#ask-about-an-image-xor-and-perplexity-read-images-repeat---image-for-several-or-use---video-xor","content":"npx @juspay/neurolink decide \"What color is this?\" --provider xor \\\n --image ./photo.png \\\n --questions '{\"color\":{\"type\":\"choice\",\"instructions\":\"What color is the image?\",\"criteria\":{\"red\":\"red\",\"blue\":\"blue\"}}}'\n\n\n| Option | Description |\n| --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- |\n| state | The content to judge, as a positional argument (or use --state-file). |\n| --state-file | Path to a file holding the state (plain text or JSON). |\n| --questions | Inline JSON map of questions. Exactly one of this or --questions-file. |\n| --questions-file | Path to a JSON file holding the questions map. |\n| --provider | Decision provider to use: typesafe, laya, xor or perplexity-decider. Defaults to the first configured, in that order (TypeSafe, Laya, XOR, Perplexity). |\n| --model | Overrides the provider's configured model for this call. |\n| --image | An image for the model to read, as a file path or a data: URL. Repeat the flag for several (XOR and Perplexity take up to 8).","hierarchy":{"lvl0":"Cli","lvl1":"CLI Command Reference","lvl2":"Ask about an image (XOR and Perplexity read images); repeat --image for several, or use --video (XOR)","lvl3":""}},
787
+ {"objectID":"bbd1f018ee4492d0ec7f492e6a84291f0eaa12b81d211cf285fe6d5ac8ed273d","title":"Ask about an image (XOR, Perplexity and Cloudflare Clef read images); repeat --image for several, or use --video (XOR)","url":"/docs/cli/commands#ask-about-an-image-xor-perplexity-and-cloudflare-clef-read-images-repeat---image-for-several-or-use---video-xor","content":"npx @juspay/neurolink decide \"What color is this?\" --provider xor \\\n --image ./photo.png \\\n --questions '{\"color\":{\"type\":\"choice\",\"instructions\":\"What color is the image?\",\"criteria\":{\"red\":\"red\",\"blue\":\"blue\"}}}'\n\n\n| Option | Description |\n| --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |\n| state | The content to judge, as a positional argument (or use --state-file). |\n| --state-file | Path to a file holding the state (plain text or JSON). |\n| --questions | Inline JSON map of questions. Exactly one of this or --questions-file. |\n| --questions-file | Path to a JSON file holding the questions map. |\n| --provider | Decision provider to use: typesafe, laya, xor, perplexity-decider or cloudflare-clef. Defaults to the first configured, in that order (TypeSafe, Laya, XOR, Perplexity, Cloudflare Clef). |\n| --model | Overrides the provider's configured model for this call.","hierarchy":{"lvl0":"Cli","lvl1":"CLI Command Reference","lvl2":"Ask about an image (XOR, Perplexity and Cloudflare Clef read images); repeat --image for several, or use --video (XOR)","lvl3":""}},
788
788
  {"objectID":"c8d0c91cc9a6c43b1ffaf7778b1b847d0082ea6ac5b07bef70f901690fab3f2d","title":"batch <file>","url":"/docs/cli/commands#batch","content":"Process multiple prompts from a file in sequence.\n\n`bash","hierarchy":{"lvl0":"Cli","lvl1":"CLI Command Reference","lvl2":"batch <file>","lvl3":""}},
789
789
  {"objectID":"7295bf29c7231cdfbc08cfe63fa54bdd96d158a2f3d4ffb093a8d86826c42abf","title":"Process prompts from a file","url":"/docs/cli/commands#process-prompts-from-a-file","content":"npx @juspay/neurolink batch prompts.txt","hierarchy":{"lvl0":"Cli","lvl1":"CLI Command Reference","lvl2":"Process prompts from a file","lvl3":""}},
790
790
  {"objectID":"f5a81a31f8f017423e190eed5892ad06dbeb9bc3d55c3224a9e34909c9332aad","title":"Export results as JSON","url":"/docs/cli/commands#export-results-as-json","content":"npx @juspay/neurolink batch questions.txt --format json","hierarchy":{"lvl0":"Cli","lvl1":"CLI Command Reference","lvl2":"Export results as JSON","lvl3":""}},
@@ -2623,19 +2623,19 @@
2623
2623
  {"objectID":"b0a5b78dc4cfcabcd52b58ea94936ee81e97b59b5907f058ccba74881a7aeacd","title":"Context-window filtering","url":"/docs/features/classifier-router-catalog#context-window-filtering","content":"Nothing in routing read maxContextTokens before this. Two independent checks\nnow do:\nrankCatalogue() filters candidates whose contextWindow is below the\n request's estimated input tokens, applied only when it would leave something\n behind (an empty result falls back to the unfiltered pool rather than\n routing nowhere).\nThe classifier's own pick is separately vetoed. Even when jev picks a\n model directly, fitsRequest() checks whether that specific model's window\n can hold the request. If it can't, the pick is dropped regardless of\n confidence — this is a hard provider error, not a degraded answer, because\n an oversized request against a real model puts that model into ModelPool's\n permanent (10-year) cooldown. A wrong guess here is not \"less accurate,\" it\n is unrecoverable for the life of the process.","hierarchy":{"lvl0":"Features","lvl1":"The model catalogue","lvl2":"Context-window filtering","lvl3":""}},
2624
2624
  {"objectID":"e34966644db91aef39b46ef5c2f2fa259e9c828895197c9b2260a07c6e640bce","title":"What this is bad at","url":"/docs/features/classifier-router-catalog#what-this-is-bad-at","content":"Registry quality is coarse, and the catalogue path never shows its own\n work. Auto-enriched quality is a 3-bucket scale (high/medium/low)\n before buildModelCatalog() turns it into 1/2/3; two \"high\" models\n cannot be separated on capability alone. Worse for this page's topic: because\n the catalogue always populates cost/quality, its rendered line never\n shows the registry's own price, speed bucket, or \"strong at …\" scores — only\n the terse capability N / relative cost N form. Declare tiers/quality\n on a hand-declared pool member for finer control over the number itself; there\n is no way to get the richer rendering for a catalogue-sourced model.\nA catalogue candidate's \"relative cost\" is real pricing wearing a relative\n label. buildModelCatalog() seeds it from inputCostPer1K + outputCostPer1K\n — a small decimal like 0.00074 — not a small integer a host would typically\n pick for a hand-declared cost. The number is still correct and still never\n rendered as currency, but it does not compare cleanly against a hand-declared\n member's cost: 1 in the same pool, since one is a real price and the other\n is an arbitrary scale.\nUnknown-to-the-registry models rank on relative numbers, not real prices.\n A model the registry doesn't know (self-hosted, brand-new) keeps whatever\n cost/quality the host declared, compared only against other declared\n values on the same relative scale — never rendered as a dollar figure it\n isn't.\nA wide catalogue is still one choice question, and the cap is hard, not\n smart, when it does bind. Every model adds tokens to that single question;\n maxModels is the safeguard, and a model past it would simply never be\n offered rather than offered with lower priority. Today this is theoretical —\n the registry is 64 models against a default cap of 120, so nothing is\n dropped — but the mechanism has no ranking behavior for the day a registry\n does exceed it.\nPermissive reachability means occasional dead candidates.","hierarchy":{"lvl0":"Features","lvl1":"The model catalogue","lvl2":"What this is bad at","lvl3":""}},
2625
2625
  {"objectID":"2d7354001109ce1051977450cf365d1f2a8e13b345ad206ee23b1f5466228454","title":"See also","url":"/docs/features/classifier-router-catalog#see-also","content":"Classifier Router\nModel routing with a decision model\nThe decide inference type","hierarchy":{"lvl0":"Features","lvl1":"The model catalogue","lvl2":"See also","lvl3":""}},
2626
- {"objectID":"57703ab4e28ce3ad7f4c83d5d3e43f1b1d5541cbb2ac1d6f10f7a057046f0e1b","title":"Model routing with a decision model","url":"/docs/features/classifier-router-jev-strategy","content":"Model routing with a decision model\n\nThe classifier router has a jev strategy: one\ndecision-model round trip answers difficulty,\nrequired capabilities, risk and the model pick simultaneously, with a\ncalibrated confidence on each. This page is the strategy's own mechanics —\nclassifier-router.md covers the router as a whole.\n\nThe degradation contract. With no decision provider configured,\nresolveStrategy() resolves auto to heuristic, exactly as it always did.\nclassifyJev() itself never throws: any failure, timeout, or malformed answer\nfalls back to classifyHeuristic(). Configuring a decision provider (for example TYPESAFE_API_KEY or\nAI_GATEWAY_API_KEY) upgrades routing; it cannot make routing worse than before\nthe key existed. The same holds for PERPLEXITY_API_KEY, which is also the\nPerplexity text provider's key.\n\nWhat goes into the state\n\nclassifyJev() sends the request truncated to 8000 characters, plus four\nsignals the caller already has on hand:\n\nAlongside it, one batch asks: difficulty (a choice over the five tiers),\nneeds_vision / needs_tools / needs_reasoning (boolean), risky\n(boolean), context (a score — see\nper-request context budget), and, only when the\npool has more than one member, model (a choice over the pool, rendered by\nthe catalogue). All of this rides in\none request. On TypeSafe that takes ~400ms, because latency is flat in question\ncount; on Perplexity each further question adds about 65 ms, up to 128 (see\nits guide). See\nthe batching rule.\n\nThe difficulty rubric\n\nFive tiers, ordered easiest to hardest, worded about the shape of the work —\nthe model is never told a provider or model name:\n\n| Tier | Criterion |\n| ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- |\n| trivial | Mechanical and local: rename a symbol, fix a typo, add an import, run one named command, or answer something already stated. |\n| simple | A small localised change or a direct factual answer. One file, one obvious approach. |\n| moderate | Ordinary engineering: implement a well-specified change across a few files, write tests, fix a clearly described bug, review a small diff. |\n| hard | Deep reasoning: architecture and design, debugging a failure whose cause is unknown, security analysis, concurrency, cross-system refactors. |\n| expert | Frontier-level work: ambiguous requirements, novel design with no established pattern, or analysis where a wrong answer is costly and hard to detect. |\n\nAsymmetric confidence bars, and why they differ\n\nA verdict harder than the neutral tier (moderate) spends more if wrong; a\nverdict easier than neutral spends less if wrong. Those are not the same\nmistake:\nRouting a simple task to an expensive model wastes money. Cheap to be wrong\n about, so the bar to route up is low: minUpgradeConfidence defaults to\n 0.3.\nRouting a hard task to a weak model produces a wrong answer. Expensive to be\n wrong about, so the bar to route down is high: minDowngradeConfidence\n defaults to 0.6.\n\nBelow the applicable bar, the difficulty verdict is discarded and the heuristic\nclassifier's tier stands instead — not a downgraded jev answer, the ordinary\nzero-cost fallback.\n\nOne case skips both bars: a risky reading above 0.7 forces the tier to at\nleast hard, unconditionally. Risk can only ever raise the tier, never lower\none, and it is not itself gated by confidence.\n\nAsk about the act, not the subject\n\nThe risk question is not \"does this touch production, money, or credentials\":\n\nCarrying out this request would itself change production, move real money,\nexpose credentials, or alter data that cannot be restored. Writing or testing\ncode that deals with such things, without running it against the real system,\ndoes not count.\n\nThe naive phrasing was tried first and scored high on ordinary code that merely\nconcerns those things — \"add a refund endpoint that calls Stripe\" — which\nwould have escalated every such request to the most expensive tier. The second\nsentence is what separates writing the code from running it against something\nreal.\n\nThe model pick faces a bar too — but a different one\n\nWhen the pool has more than one member, jev is also asked to choose directly:\n_\"Which of these models is the cheapest one that can still complete this\nrequest correctly?\"_ — over\nthe rendered candidate lines.\nThat pick is reported with its own confidence, separate from the difficulty\nconfidence, and the router decides which bar applies:\n\nPicking something costlier than what the difficulty tier would have picked\non its own risks only spending more than necessary, so it clears the low\nupgrade bar. Picking something cheaper risks handing the task to a model\ntha","hierarchy":{"lvl0":"Features","lvl1":"Model routing with a decision model","lvl2":"","lvl3":""}},
2627
- {"objectID":"bd360e5abf4e93403d1fb23a27c946fed1b86d8d8863a98519338332ca726639","title":"Model routing with a decision model","url":"/docs/features/classifier-router-jev-strategy#model-routing-with-a-decision-model","content":"The classifier router has a jev strategy: one\ndecision-model round trip answers difficulty,\nrequired capabilities, risk and the model pick simultaneously, with a\ncalibrated confidence on each. This page is the strategy's own mechanics —\nclassifier-router.md covers the router as a whole.\n\nThe degradation contract. With no decision provider configured,\nresolveStrategy() resolves auto to heuristic, exactly as it always did.\nclassifyJev() itself never throws: any failure, timeout, or malformed answer\nfalls back to classifyHeuristic(). Configuring a decision provider (for example TYPESAFE_API_KEY or\nAI_GATEWAY_API_KEY) upgrades routing; it cannot make routing worse than before\nthe key existed. The same holds for PERPLEXITY_API_KEY, which is also the\nPerplexity text provider's key.","hierarchy":{"lvl0":"Features","lvl1":"Model routing with a decision model","lvl2":"Model routing with a decision model","lvl3":""}},
2628
- {"objectID":"a6b27c365174a477eafb6835b9ab5990c7434acd58cb21534f8419ea21c63dd5","title":"What goes into the state","url":"/docs/features/classifier-router-jev-strategy#what-goes-into-the-state","content":"classifyJev() sends the request truncated to 8000 characters, plus four\nsignals the caller already has on hand:\n\nAlongside it, one batch asks: difficulty (a choice over the five tiers),\nneeds_vision / needs_tools / needs_reasoning (boolean), risky\n(boolean), context (a score — see\nper-request context budget), and, only when the\npool has more than one member, model (a choice over the pool, rendered by\nthe catalogue). All of this rides in\none request. On TypeSafe that takes ~400ms, because latency is flat in question\ncount; on Perplexity each further question adds about 65 ms, up to 128 (see\nits guide). See\nthe batching rule.","hierarchy":{"lvl0":"Features","lvl1":"Model routing with a decision model","lvl2":"What goes into the state","lvl3":""}},
2626
+ {"objectID":"57703ab4e28ce3ad7f4c83d5d3e43f1b1d5541cbb2ac1d6f10f7a057046f0e1b","title":"Model routing with a decision model","url":"/docs/features/classifier-router-jev-strategy","content":"Model routing with a decision model\n\nThe classifier router has a jev strategy: one\ndecision-model round trip answers difficulty,\nrequired capabilities, risk and the model pick simultaneously, with a\ncalibrated confidence on each. This page is the strategy's own mechanics —\nclassifier-router.md covers the router as a whole.\n\nThe degradation contract. With no decision provider configured,\nresolveStrategy() resolves auto to heuristic, exactly as it always did.\nclassifyJev() itself never throws: any failure, timeout, or malformed answer\nfalls back to classifyHeuristic(). Configuring a decision provider (for example TYPESAFE_API_KEY or\nAI_GATEWAY_API_KEY) upgrades routing; it cannot make routing worse than before\nthe key existed. The same holds for PERPLEXITY_API_KEY, which is also the\nPerplexity text provider's key, and for CLOUDFLARE_API_KEY with\nCLOUDFLARE_ACCOUNT_ID, which are also the Cloudflare Workers AI text provider's\ntoken and account id.\n\nWhat goes into the state\n\nclassifyJev() sends the request truncated to 8000 characters, plus four\nsignals the caller already has on hand:\n\nAlongside it, one batch asks: difficulty (a choice over the five tiers),\nneeds_vision / needs_tools / needs_reasoning (boolean), risky\n(boolean), context (a score — see\nper-request context budget), and, only when the\npool has more than one member, model (a choice over the pool, rendered by\nthe catalogue). All of this rides in\none request. On TypeSafe that takes ~400ms, because latency is flat in question\ncount; on Perplexity each further question adds about 65 ms, up to 128 (see\nits guide); on\nCloudflare Clef a small request took 0.3 to 1.0 s on 2026-10-03 (see\nits guide). See\nthe batching rule.\n\nThe Cloudflare Clef endpoint ignores state text past about 2,048 tokens, and\nNeuroLink refuses a state it estimates at more than 1,500 tokens with\nmax_tokens_exceeded. A long request is therefore refused there, and the\nheuristic classifier's tier stands, as it does on any failure.\n\nThe difficulty rubric\n\nFive tiers, ordered easiest to hardest, worded about the shape of the work —\nthe model is never told a provider or model name:\n\n| Tier | Criterion |\n| ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- |\n| trivial | Mechanical and local: rename a symbol, fix a typo, add an import, run one named command, or answer something already stated. |\n| simple | A small localised change or a direct factual answer. One file, one obvious approach. |\n| moderate | Ordinary engineering: implement a well-specified change across a few files, write tests, fix a clearly described bug, review a small diff. |\n| hard | Deep reasoning: architecture and design, debugging a failure whose cause is unknown, security analysis, concurrency, cross-system refactors. |\n| expert | Frontier-level work: ambiguous requirements, novel design with no established pattern, or analysis where a wrong answer is costly and hard to detect. |\n\nAsymmetric confidence bars, and why they differ\n\nA verdict harder than the neutral tier (moderate) spends more if wrong; a\nverdict easier than neutral spends less if wrong. Those are not the same\nmistake:\nRouting a simple task to an expensive model wastes money. Cheap to be wrong\n about, so the bar to route up is low: minUpgradeConfidence defaults to\n 0.3.\nRouting a hard task to a weak model produces a wrong answer. Expensive to be\n wrong about, so the bar to route down is high: minDowngradeConfidence\n defaults to 0.6.\n\nBelow the applicable bar, the difficulty verdict is discarded and the heuristic\nclassifier's tier stands instead — not a downgraded jev answer, the ordinary\nzero-cost fallback.\n\nOne case skips both bars: a risky reading above 0.7 forces the tier to at\nleast hard, unconditionally. Risk can only ever raise the tier, never lower\none, and it is not itself gated by confidence.\n\nAsk about the act, not the subject\n\nThe risk question is not \"does this touch production, money, or credentials\":\n\nCarrying out this request would itself change production, move real money,\nexpose credentials, or alter data that cannot be restored. Writing or testing\ncode that deals with such things, without running it against the real system,\ndoes not count.\n\nThe naive phrasing was tried first and scored high on ordinary code that merely\nconcerns those things — \"add a refund endpoint that calls Stripe\" — which\nwould have escalated every such request to the most expensive tier. The second\nsentence is what separates writing the code from running it against something\nreal.\n\nThe model pick faces a bar too — but a different one\n\nWhen the pool has more than one member, jev is also asked to ch","hierarchy":{"lvl0":"Features","lvl1":"Model routing with a decision model","lvl2":"","lvl3":""}},
2627
+ {"objectID":"bd360e5abf4e93403d1fb23a27c946fed1b86d8d8863a98519338332ca726639","title":"Model routing with a decision model","url":"/docs/features/classifier-router-jev-strategy#model-routing-with-a-decision-model","content":"The classifier router has a jev strategy: one\ndecision-model round trip answers difficulty,\nrequired capabilities, risk and the model pick simultaneously, with a\ncalibrated confidence on each. This page is the strategy's own mechanics —\nclassifier-router.md covers the router as a whole.\n\nThe degradation contract. With no decision provider configured,\nresolveStrategy() resolves auto to heuristic, exactly as it always did.\nclassifyJev() itself never throws: any failure, timeout, or malformed answer\nfalls back to classifyHeuristic(). Configuring a decision provider (for example TYPESAFE_API_KEY or\nAI_GATEWAY_API_KEY) upgrades routing; it cannot make routing worse than before\nthe key existed. The same holds for PERPLEXITY_API_KEY, which is also the\nPerplexity text provider's key, and for CLOUDFLARE_API_KEY with\nCLOUDFLARE_ACCOUNT_ID, which are also the Cloudflare Workers AI text provider's\ntoken and account id.","hierarchy":{"lvl0":"Features","lvl1":"Model routing with a decision model","lvl2":"Model routing with a decision model","lvl3":""}},
2628
+ {"objectID":"a6b27c365174a477eafb6835b9ab5990c7434acd58cb21534f8419ea21c63dd5","title":"What goes into the state","url":"/docs/features/classifier-router-jev-strategy#what-goes-into-the-state","content":"classifyJev() sends the request truncated to 8000 characters, plus four\nsignals the caller already has on hand:\n\nAlongside it, one batch asks: difficulty (a choice over the five tiers),\nneeds_vision / needs_tools / needs_reasoning (boolean), risky\n(boolean), context (a score — see\nper-request context budget), and, only when the\npool has more than one member, model (a choice over the pool, rendered by\nthe catalogue). All of this rides in\none request. On TypeSafe that takes ~400ms, because latency is flat in question\ncount; on Perplexity each further question adds about 65 ms, up to 128 (see\nits guide); on\nCloudflare Clef a small request took 0.3 to 1.0 s on 2026-10-03 (see\nits guide). See\nthe batching rule.\n\nThe Cloudflare Clef endpoint ignores state text past about 2,048 tokens, and\nNeuroLink refuses a state it estimates at more than 1,500 tokens with\nmax_tokens_exceeded. A long request is therefore refused there, and the\nheuristic classifier's tier stands, as it does on any failure.","hierarchy":{"lvl0":"Features","lvl1":"Model routing with a decision model","lvl2":"What goes into the state","lvl3":""}},
2629
2629
  {"objectID":"ed3d90e1e698e522c2fb4cbd27b81d89a769c3d860bbbb85d004ef5cf1d1090d","title":"The difficulty rubric","url":"/docs/features/classifier-router-jev-strategy#the-difficulty-rubric","content":"Five tiers, ordered easiest to hardest, worded about the shape of the work —\nthe model is never told a provider or model name:\n\n| Tier | Criterion |\n| ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- |\n| trivial | Mechanical and local: rename a symbol, fix a typo, add an import, run one named command, or answer something already stated. |\n| simple | A small localised change or a direct factual answer. One file, one obvious approach. |\n| moderate | Ordinary engineering: implement a well-specified change across a few files, write tests, fix a clearly described bug, review a small diff. |\n| hard | Deep reasoning: architecture and design, debugging a failure whose cause is unknown, security analysis, concurrency, cross-system refactors. |\n| expert | Frontier-level work: ambiguous requirements, novel design with no established pattern, or analysis where a wrong answer is costly and hard to detect. |","hierarchy":{"lvl0":"Features","lvl1":"Model routing with a decision model","lvl2":"The difficulty rubric","lvl3":""}},
2630
2630
  {"objectID":"2bcf89adce5486bff7cc01f6916f9bf26080de8996ec2277df6e15eb3eb264e3","title":"Asymmetric confidence bars, and why they differ","url":"/docs/features/classifier-router-jev-strategy#asymmetric-confidence-bars-and-why-they-differ","content":"A verdict harder than the neutral tier (moderate) spends more if wrong; a\nverdict easier than neutral spends less if wrong. Those are not the same\nmistake:\nRouting a simple task to an expensive model wastes money. Cheap to be wrong\n about, so the bar to route up is low: minUpgradeConfidence defaults to\n 0.3.\nRouting a hard task to a weak model produces a wrong answer. Expensive to be\n wrong about, so the bar to route down is high: minDowngradeConfidence\n defaults to 0.6.\n\nBelow the applicable bar, the difficulty verdict is discarded and the heuristic\nclassifier's tier stands instead — not a downgraded jev answer, the ordinary\nzero-cost fallback.\n\nOne case skips both bars: a risky reading above 0.7 forces the tier to at\nleast hard, unconditionally. Risk can only ever raise the tier, never lower\none, and it is not itself gated by confidence.","hierarchy":{"lvl0":"Features","lvl1":"Model routing with a decision model","lvl2":"Asymmetric confidence bars, and why they differ","lvl3":""}},
2631
2631
  {"objectID":"516e7c128214404728015c7b49bc089992935102cb4baebbcdc3d8170078873c","title":"Ask about the act, not the subject","url":"/docs/features/classifier-router-jev-strategy#ask-about-the-act-not-the-subject","content":"The risk question is not \"does this touch production, money, or credentials\":\n\nCarrying out this request would itself change production, move real money,\nexpose credentials, or alter data that cannot be restored. Writing or testing\ncode that deals with such things, without running it against the real system,\ndoes not count.\n\nThe naive phrasing was tried first and scored high on ordinary code that merely\nconcerns those things — \"add a refund endpoint that calls Stripe\" — which\nwould have escalated every such request to the most expensive tier. The second\nsentence is what separates writing the code from running it against something\nreal.","hierarchy":{"lvl0":"Features","lvl1":"Model routing with a decision model","lvl2":"Ask about the act, not the subject","lvl3":""}},
2632
2632
  {"objectID":"d90c4af28403b7660202646864195bb2ff353a7f7a29ae72f01fd8aa997a39dd","title":"The model pick faces a bar too — but a different one","url":"/docs/features/classifier-router-jev-strategy#the-model-pick-faces-a-bar-too-but-a-different-one","content":"When the pool has more than one member, jev is also asked to choose directly:\n_\"Which of these models is the cheapest one that can still complete this\nrequest correctly?\"_ — over\nthe rendered candidate lines.\nThat pick is reported with its own confidence, separate from the difficulty\nconfidence, and the router decides which bar applies:\n\nPicking something costlier than what the difficulty tier would have picked\non its own risks only spending more than necessary, so it clears the low\nupgrade bar. Picking something cheaper risks handing the task to a model\nthat cannot do it, so it must clear the high downgrade bar. Agreeing with the\ntier needs no bar at all. If the pick fails its bar, it is dropped — logged as\n\"classifier pick dropped — below its bar\" — and the tier's own ranked list is\nused instead.\n\nThis is a genuinely different question from the difficulty asymmetry above:\nthat one asks whether the tier is trustworthy; this one asks whether the\nspecific model choice, once a tier is settled, is trustworthy — and the two\ncan point in opposite directions (a confident-enough \"hard\" verdict whose model\npick still misses its own, stricter bar).\n\nA pick that cannot physically hold the request is not gated at all — it is\ndropped outright, regardless of confidence, because that is a hard provider\nerror rather than a degraded answer: see\nthe catalogue's context-window filtering.\n\nThe llm strategy's picks are exempt from both bars: it reports no confidence\nfor its pick, so a bar would either always pass or always fail. Its picks are\nhonoured exactly as they were before jev existed — imposing a bar here would\nsilently change an unrelated, already-shipped strategy.","hierarchy":{"lvl0":"Features","lvl1":"Model routing with a decision model","lvl2":"The model pick faces a bar too — but a different one","lvl3":""}},
2633
2633
  {"objectID":"4553a84789531ca2e1bad075175d684e27e4ab3fcbb009062048cdf11765c2fb","title":"What this is bad at","url":"/docs/features/classifier-router-jev-strategy#what-this-is-bad-at","content":"It cannot tell you why. The reason field is a debug string built from\n the numbers, not an explanation the model gave. If you need an auditable\n rationale for a routing decision, this is the wrong tool.\nA close call still routes somewhere. A 0.29 upgrade verdict and a 0.61\n downgrade verdict are both one hundredth of a point from the bar, and both\n fall all the way back to the heuristic tier rather than to \"the second most\n likely tier.\" There is no partial credit.\nThe risk question is a single boolean. It cannot express \"risky, but\n only mildly\" — anything past 0.7 jumps straight to hard, whatever the\n actual severity.\nIt shares the base model's general limits. Literal reading, no\n arithmetic, and state relevance affecting accuracy all apply here exactly as\n described in what decide is bad at.\nThe pool still has to exist. jev chooses among what you declared (or\n what the catalogue built); it\n cannot invent a model you have no credentials for.","hierarchy":{"lvl0":"Features","lvl1":"Model routing with a decision model","lvl2":"What this is bad at","lvl3":""}},
2634
2634
  {"objectID":"ab4d25365369a722077741378e60be6e3ef36f612ed7a85340c93cf9d47223da","title":"See also","url":"/docs/features/classifier-router-jev-strategy#see-also","content":"The decide inference type\nClassifier Router\nThe model catalogue\nPer-request context budget","hierarchy":{"lvl0":"Features","lvl1":"Model routing with a decision model","lvl2":"See also","lvl3":""}},
2635
- {"objectID":"3f990fb1489f99d505b19941e00fd05b9368a5a91745c83666639aac85865343","title":"Classifier Router","url":"/docs/features/classifier-router","content":"Classifier Router\n\nStatus: Stable | Availability: SDK + CLI | Opt-in (disabled by default)\n\nOverview\n\nThe Classifier Router lets NeuroLink decide, per request, which model to use (and, optionally, which tools to expose) from a pool you declare — routing hard/complex tasks to more capable models and easy tasks to cheaper, faster ones. You give it a pool of models; it classifies the incoming prompt and switches to the best one transparently before the call runs.\n\nIt is opt-in (classifierRouter.enabled defaults to false), fails open (any classifier/selection error leaves the call exactly as it would have been), and is fully backward compatible — a default new NeuroLink() is unchanged.\n\nTypical use cases:\nCost optimization — send \"hi\" to a cheap model and a multi-step architecture question to a powerful one, automatically.\nLatency optimization — keep simple turns on fast models.\nCustom / self-hosted fleets — route across LiteLLM, OpenAI-compatible, or Ollama models that aren't in any registry.\nPer-difficulty tool scoping — expose fewer tools for trivial tasks.\n\nHow it works\n\nEach request flows through two stages:\nClassify — produce a difficulty bucket (trivial | simple | moderate | hard | expert) plus optional requiredCapabilities and tool hints. Four strategies:\nauto (default): resolves to jev when a decision provider is configured (TYPESAFE_API_KEY or AI_GATEWAY_API_KEY, LAYA_API_KEY with LAYA_BASE_URL, XOR_API_KEY with XOR_BASE_URL, or PERPLEXITY_API_KEY), and heuristic otherwise. Setting a key therefore upgrades routing with no code change; without one, behaviour is exactly as it was. PERPLEXITY_API_KEY is also the Perplexity text provider's key, so it counts here too. The gate reads the credentials given to the NeuroLink constructor and the environment, not credentials passed on a single call, so a per-call credentials.perplexityDecider does not influence it.\nheuristic: zero-cost keyword/length scoring of the prompt text. No LLM call, fully deterministic, provider-agnostic.\nllm: a cheap \"classifier model\" reads the prompt and returns a difficulty — and, when given your pool, picks a model directly by id.\njev: a decision model (the first configured decision provider, in the order TypeSafe's Jev, Laya, XOR, Perplexity) answers difficulty, required capabilities and the model pick in one request, with a confidence for each answer. On TypeSafe's Jev that request takes ~400 ms and the confidence is calibrated; the Perplexity guide gives its latency and the confidence it reports. Verdicts that miss the applicable confidence bar (minUpgradeConfidence 0.3 to route up, minDowngradeConfidence 0.6 to route down) fall through to the heuristic rather than acting on a guess. See the decide inference type.\nSelect — turn that into a concrete { provider, model, region } from your pool, optionally narrowing tools.\n\nThe router runs before the provider/model is constructed (it reuses the same pre-call seam as requestRouter). It is skipped when the caller pinned both provider and model, or when a modelPool is configured (the pool owns selection).\n\nQuick start (heuristic, SDK)\n\nDefining \"which model for which case\"\n\nThe router resolves a model using the first of these that applies:\n\n| # | Mechanism | How you define it | Best for |\n| --- | ---------------------- | ---------------------------------------------------------------------------------- | ------------------------------------- |\n| 1 | LLM direct pick | classifier: \"llm\" + a description on each pool member | Custom/registry-less models; smartest |\n| 2 | tierMap | Explicit difficulty → members map | Full, deterministic control |\n| 3 | Per-member tiers | tiers: [\"hard\",\"expert\"] on a member | Simple, explicit, generic |\n| 4 | Metadata scoring | cost / quality per member (declared, or auto-enriched from the model registry) | Known models with comparable metadata |\n\nWhen two members can't be separated (e.g. equal declared quality, or the model registry only knows both as \"high\" quality), the router keeps the declared pool order. For reliable hard-vs-easy separation, prefer mechanisms 1–3, or give members distinct quality values.\n\nMetadata scoring rules\ntrivial / simple → cheapest first (cost ascending)\nmoderate → best quality āˆ’ cost\nhard / expert → most capable first (quality descending)\n\nMembers may declare cost (relative, lower = cheaper) and quality (relative, higher = more capable). If omitted, NeuroLink tries to enrich them from its model registry; if the model is unknown (e.g. a custom LiteLLM endpoint), use mechanisms 1–3 instead.\n\nCustom & self-hosted models (LiteLLM, OpenAI-compatible, Ollama)\n\nThese models aren't in any registry, so define routing explicitly — both approaches are fully generic:\n\nHeurist","hierarchy":{"lvl0":"Features","lvl1":"Classifier Router","lvl2":"","lvl3":""}},
2635
+ {"objectID":"3f990fb1489f99d505b19941e00fd05b9368a5a91745c83666639aac85865343","title":"Classifier Router","url":"/docs/features/classifier-router","content":"Classifier Router\n\nStatus: Stable | Availability: SDK + CLI | Opt-in (disabled by default)\n\nOverview\n\nThe Classifier Router lets NeuroLink decide, per request, which model to use (and, optionally, which tools to expose) from a pool you declare — routing hard/complex tasks to more capable models and easy tasks to cheaper, faster ones. You give it a pool of models; it classifies the incoming prompt and switches to the best one transparently before the call runs.\n\nIt is opt-in (classifierRouter.enabled defaults to false), fails open (any classifier/selection error leaves the call exactly as it would have been), and is fully backward compatible — a default new NeuroLink() is unchanged.\n\nTypical use cases:\nCost optimization — send \"hi\" to a cheap model and a multi-step architecture question to a powerful one, automatically.\nLatency optimization — keep simple turns on fast models.\nCustom / self-hosted fleets — route across LiteLLM, OpenAI-compatible, or Ollama models that aren't in any registry.\nPer-difficulty tool scoping — expose fewer tools for trivial tasks.\n\nHow it works\n\nEach request flows through two stages:\nClassify — produce a difficulty bucket (trivial | simple | moderate | hard | expert) plus optional requiredCapabilities and tool hints. Four strategies:\nauto (default): resolves to jev when a decision provider is configured (TYPESAFE_API_KEY or AI_GATEWAY_API_KEY, LAYA_API_KEY with LAYA_BASE_URL, XOR_API_KEY with XOR_BASE_URL, PERPLEXITY_API_KEY, or CLOUDFLARE_API_KEY with CLOUDFLARE_ACCOUNT_ID), and heuristic otherwise. Setting a key therefore upgrades routing with no code change; without one, behaviour is exactly as it was. PERPLEXITY_API_KEY is also the Perplexity text provider's key, so it counts here too, and CLOUDFLARE_API_KEY with CLOUDFLARE_ACCOUNT_ID are also the Cloudflare Workers AI text provider's token and account id, so they count here as well. The gate reads the credentials given to the NeuroLink constructor and the environment, not credentials passed on a single call, so a per-call credentials.perplexityDecider or credentials.cloudflareClef does not influence it.\nheuristic: zero-cost keyword/length scoring of the prompt text. No LLM call, fully deterministic, provider-agnostic.\nllm: a cheap \"classifier model\" reads the prompt and returns a difficulty — and, when given your pool, picks a model directly by id.\njev: a decision model (the first configured decision provider, in the order TypeSafe's Jev, Laya, XOR, Perplexity, Cloudflare Clef) answers difficulty, required capabilities and the model pick in one request, with a confidence for each answer. On TypeSafe's Jev that request takes ~400 ms and the confidence is calibrated; the Perplexity guide gives its latency and the confidence it reports, and the Cloudflare Clef guide gives its latency and its state limit (about 2,048 tokens, so a long prompt is refused and the heuristic stands). Verdicts that miss the applicable confidence bar (minUpgradeConfidence 0.3 to route up, minDowngradeConfidence 0.6 to route down) fall through to the heuristic rather than acting on a guess. See the decide inference type.\nSelect — turn that into a concrete { provider, model, region } from your pool, optionally narrowing tools.\n\nThe router runs before the provider/model is constructed (it reuses the same pre-call seam as requestRouter). It is skipped when the caller pinned both provider and model, or when a modelPool is configured (the pool owns selection).\n\nQuick start (heuristic, SDK)\n\nDefining \"which model for which case\"\n\nThe router resolves a model using the first of these that applies:\n\n| # | Mechanism | How you define it | Best for |\n| --- | ---------------------- | ---------------------------------------------------------------------------------- | ------------------------------------- |\n| 1 | LLM direct pick | classifier: \"llm\" + a description on each pool member | Custom/registry-less models; smartest |\n| 2 | tierMap | Explicit difficulty → members map | Full, deterministic control |\n| 3 | Per-member tiers | tiers: [\"hard\",\"expert\"] on a member | Simple, explicit, generic |\n| 4 | Metadata scoring | cost / quality per member (declared, or auto-enriched from the model registry) | Known models with comparable metadata |\n\nWhen two members can't be separated (e.g. equal declared quality, or the model registry only knows both as \"high\" quality), the router keeps the declared pool order. For reliable hard-vs-easy separation, prefer mechanisms 1–3, or give members distinct quality values.\n\nMetadata scoring rules\ntrivial / simple → cheapest first (cost ascending)\nmoderate → best quality āˆ’ cost\nhard / expert → most capable first (quality descending)\n\nMembers may declare cost (relative, lower =","hierarchy":{"lvl0":"Features","lvl1":"Classifier Router","lvl2":"","lvl3":""}},
2636
2636
  {"objectID":"b8bb35b0db438959a0a557c68b22afa16480504b96128260a8947ba3c8a81f9f","title":"Classifier Router","url":"/docs/features/classifier-router#classifier-router","content":"Status: Stable | Availability: SDK + CLI | Opt-in (disabled by default)","hierarchy":{"lvl0":"Features","lvl1":"Classifier Router","lvl2":"Classifier Router","lvl3":""}},
2637
2637
  {"objectID":"543ddabeefd7754f6047758090a6a0ac6f517c98b9afb146645b9073ad75c5ee","title":"Overview","url":"/docs/features/classifier-router#overview","content":"The Classifier Router lets NeuroLink decide, per request, which model to use (and, optionally, which tools to expose) from a pool you declare — routing hard/complex tasks to more capable models and easy tasks to cheaper, faster ones. You give it a pool of models; it classifies the incoming prompt and switches to the best one transparently before the call runs.\n\nIt is opt-in (classifierRouter.enabled defaults to false), fails open (any classifier/selection error leaves the call exactly as it would have been), and is fully backward compatible — a default new NeuroLink() is unchanged.\n\nTypical use cases:\nCost optimization — send \"hi\" to a cheap model and a multi-step architecture question to a powerful one, automatically.\nLatency optimization — keep simple turns on fast models.\nCustom / self-hosted fleets — route across LiteLLM, OpenAI-compatible, or Ollama models that aren't in any registry.\nPer-difficulty tool scoping — expose fewer tools for trivial tasks.","hierarchy":{"lvl0":"Features","lvl1":"Classifier Router","lvl2":"Overview","lvl3":""}},
2638
- {"objectID":"abdbdbfcfd51c6b40ef9cdf848fa23aeb87e4e4513c71572f89ee754be41d834","title":"How it works","url":"/docs/features/classifier-router#how-it-works","content":"Each request flows through two stages:\nClassify — produce a difficulty bucket (trivial | simple | moderate | hard | expert) plus optional requiredCapabilities and tool hints. Four strategies:\nauto (default): resolves to jev when a decision provider is configured (TYPESAFE_API_KEY or AI_GATEWAY_API_KEY, LAYA_API_KEY with LAYA_BASE_URL, XOR_API_KEY with XOR_BASE_URL, or PERPLEXITY_API_KEY), and heuristic otherwise. Setting a key therefore upgrades routing with no code change; without one, behaviour is exactly as it was. PERPLEXITY_API_KEY is also the Perplexity text provider's key, so it counts here too. The gate reads the credentials given to the NeuroLink constructor and the environment, not credentials passed on a single call, so a per-call credentials.perplexityDecider does not influence it.\nheuristic: zero-cost keyword/length scoring of the prompt text. No LLM call, fully deterministic, provider-agnostic.\nllm: a cheap \"classifier model\" reads the prompt and returns a difficulty — and, when given your pool, picks a model directly by id.\njev: a decision model (the first configured decision provider, in the order TypeSafe's Jev, Laya, XOR, Perplexity) answers difficulty, required capabilities and the model pick in one request, with a confidence for each answer. On TypeSafe's Jev that request takes ~400 ms and the confidence is calibrated; the Perplexity guide gives its latency and the confidence it reports. Verdicts that miss the applicable confidence bar (minUpgradeConfidence 0.3 to route up, minDowngradeConfidence 0.6 to route down) fall through to the heuristic rather than acting on a guess. See the decide inference type.\nSelect — turn that into a concrete { provider, model, region } from your pool, optionally narrowing tools.\n\nThe router runs before the provider/model is constructed (it reuses the same pre-call seam as requestRouter). It is skipped when the caller pinned both provider and model, or when a modelPool is configured (the pool owns selection).","hierarchy":{"lvl0":"Features","lvl1":"Classifier Router","lvl2":"How it works","lvl3":""}},
2638
+ {"objectID":"abdbdbfcfd51c6b40ef9cdf848fa23aeb87e4e4513c71572f89ee754be41d834","title":"How it works","url":"/docs/features/classifier-router#how-it-works","content":"Each request flows through two stages:\nClassify — produce a difficulty bucket (trivial | simple | moderate | hard | expert) plus optional requiredCapabilities and tool hints. Four strategies:\nauto (default): resolves to jev when a decision provider is configured (TYPESAFE_API_KEY or AI_GATEWAY_API_KEY, LAYA_API_KEY with LAYA_BASE_URL, XOR_API_KEY with XOR_BASE_URL, PERPLEXITY_API_KEY, or CLOUDFLARE_API_KEY with CLOUDFLARE_ACCOUNT_ID), and heuristic otherwise. Setting a key therefore upgrades routing with no code change; without one, behaviour is exactly as it was. PERPLEXITY_API_KEY is also the Perplexity text provider's key, so it counts here too, and CLOUDFLARE_API_KEY with CLOUDFLARE_ACCOUNT_ID are also the Cloudflare Workers AI text provider's token and account id, so they count here as well. The gate reads the credentials given to the NeuroLink constructor and the environment, not credentials passed on a single call, so a per-call credentials.perplexityDecider or credentials.cloudflareClef does not influence it.\nheuristic: zero-cost keyword/length scoring of the prompt text. No LLM call, fully deterministic, provider-agnostic.\nllm: a cheap \"classifier model\" reads the prompt and returns a difficulty — and, when given your pool, picks a model directly by id.\njev: a decision model (the first configured decision provider, in the order TypeSafe's Jev, Laya, XOR, Perplexity, Cloudflare Clef) answers difficulty, required capabilities and the model pick in one request, with a confidence for each answer. On TypeSafe's Jev that request takes ~400 ms and the confidence is calibrated; the Perplexity guide gives its latency and the confidence it reports, and the Cloudflare Clef guide gives its latency and its state limit (about 2,048 tokens, so a long prompt is refused and the heuristic stands).","hierarchy":{"lvl0":"Features","lvl1":"Classifier Router","lvl2":"How it works","lvl3":""}},
2639
2639
  {"objectID":"08f69ffa1f39bf410a5c236efa5f21f51efc93034747ddab6d761f4b1615d7df","title":"Defining \"which model for which case\"","url":"/docs/features/classifier-router#defining-which-model-for-which-case","content":"The router resolves a model using the first of these that applies:\n\n| # | Mechanism | How you define it | Best for |\n| --- | ---------------------- | ---------------------------------------------------------------------------------- | ------------------------------------- |\n| 1 | LLM direct pick | classifier: \"llm\" + a description on each pool member | Custom/registry-less models; smartest |\n| 2 | tierMap | Explicit difficulty → members map | Full, deterministic control |\n| 3 | Per-member tiers | tiers: [\"hard\",\"expert\"] on a member | Simple, explicit, generic |\n| 4 | Metadata scoring | cost / quality per member (declared, or auto-enriched from the model registry) | Known models with comparable metadata |\n\nWhen two members can't be separated (e.g. equal declared quality, or the model registry only knows both as \"high\" quality), the router keeps the declared pool order. For reliable hard-vs-easy separation, prefer mechanisms 1–3, or give members distinct quality values.","hierarchy":{"lvl0":"Features","lvl1":"Classifier Router","lvl2":"Defining \"which model for which case\"","lvl3":""}},
2640
2640
  {"objectID":"bd4863bf711f00f17fb26ea1e60a541aef1f396afc1050ee47b9e29731c90f9a","title":"Metadata scoring rules","url":"/docs/features/classifier-router#metadata-scoring-rules","content":"trivial / simple → cheapest first (cost ascending)\nmoderate → best quality āˆ’ cost\nhard / expert → most capable first (quality descending)\n\nMembers may declare cost (relative, lower = cheaper) and quality (relative, higher = more capable). If omitted, NeuroLink tries to enrich them from its model registry; if the model is unknown (e.g. a custom LiteLLM endpoint), use mechanisms 1–3 instead.","hierarchy":{"lvl0":"Features","lvl1":"Classifier Router","lvl2":"Metadata scoring rules","lvl3":""}},
2641
2641
  {"objectID":"9e5f93906c1460117b27372537b30d578daec8b1d3fc38ecf1b0fb05c9944801","title":"Custom & self-hosted models (LiteLLM, OpenAI-compatible, Ollama)","url":"/docs/features/classifier-router#custom-self-hosted-models-litellm-openai-compatible-ollama","content":"These models aren't in any registry, so define routing explicitly — both approaches are fully generic:\n\nHeuristic + tiers (deterministic, no LLM cost):\n\nLLM picks per-prompt from plain-English descriptions (most flexible):\n\nThe classifier model is shown each candidate's id (defaults to provider/model) and description, and returns the best id for the prompt. An invalid or absent pick falls back to difficulty-based selection; any classifier failure falls back to the heuristic.","hierarchy":{"lvl0":"Features","lvl1":"Classifier Router","lvl2":"Custom & self-hosted models (LiteLLM, OpenAI-compatible, Ollama)","lvl3":""}},
@@ -2644,7 +2644,7 @@
2644
2644
  {"objectID":"cc0754af3d357fb93691ebfaad77343948f167866e856d3e941738c8293fa323","title":"Heuristic routing across a pool (inline JSON or a file path)","url":"/docs/features/classifier-router#heuristic-routing-across-a-pool-inline-json-or-a-file-path","content":"neurolink generate \"hi\" \\\n --classifier-router \\\n --classifier-pool '[{\"provider\":\"vertex\",\"model\":\"gemini-2.5-flash\",\"tiers\":[\"trivial\",\"simple\",\"moderate\"]},{\"provider\":\"vertex\",\"model\":\"gemini-2.5-pro\",\"tiers\":[\"hard\",\"expert\"]}]'","hierarchy":{"lvl0":"Features","lvl1":"Classifier Router","lvl2":"Heuristic routing across a pool (inline JSON or a file path)","lvl3":""}},
2645
2645
  {"objectID":"c1599b72e48c45d71d427eff367de7a998cb96e25063546d18d1d8a9a520a74d","title":"LLM classifier picks the model per prompt (descriptions drive the choice)","url":"/docs/features/classifier-router#llm-classifier-picks-the-model-per-prompt-descriptions-drive-the-choice","content":"neurolink generate \"Design a multi-region architecture\" \\\n --classifier-router \\\n --classifier-strategy llm \\\n --classifier-model-provider vertex --classifier-model-name gemini-2.5-flash \\\n --classifier-pool ./pool.json\n\n\n| Flag | Description |\n| --------------------------------------- | ------------------------------------------------------- |\n| --classifier-router | Enable the classifier router. |\n| --classifier-strategy | auto (default), heuristic, llm or jev. |\n| --classifier-min-upgrade-confidence | Confidence needed to route UP (jev; 0.3). |\n| --classifier-min-downgrade-confidence | Confidence needed to route DOWN (jev; 0.6). |\n| --classifier-model-provider | Provider for the LLM classifier model (strategy=llm). |\n| --classifier-model-name | Model name for the LLM classifier model. |\n| --classifier-model-region | Region for the LLM classifier model. |\n| --classifier-pool | JSON file path or inline JSON array of pool members. |\n| --classifier-timeout | LLM classifier hard timeout (ms). |\n\n> CLI flags cover the common case (strategy, classifier model, pool). For tierMap and toolDirectives`, use the SDK config.","hierarchy":{"lvl0":"Features","lvl1":"Classifier Router","lvl2":"LLM classifier picks the model per prompt (descriptions drive the choice)","lvl3":""}},
2646
2646
  {"objectID":"9fe3ff72b1728db15b0ad44eb4108926b4917c45feb4a7f6842cf5c247d300e3","title":"Precedence & interactions","url":"/docs/features/classifier-router#precedence-interactions","content":"Model selection resolves in this order: caller-pinned provider+model > classifierRouter > requestRouter > legacy enableOrchestration. The classifier marks the request so the downstream selectors stand down.\nmodelPool — when a modelPool is configured the classifier stands down (the pool owns selection); use one or the other for model choice.\ntoolRouting — the dedicated tool-routing feature still applies; classifier toolDirectives are additive.","hierarchy":{"lvl0":"Features","lvl1":"Classifier Router","lvl2":"Precedence & interactions","lvl3":""}},
2647
- {"objectID":"fac896bd6ef5ba1295df2e3b8249209a904d5a4b66780e388934069517ed6014","title":"Caveats","url":"/docs/features/classifier-router#caveats","content":"Registry quality is coarse. Auto-enrichment maps a model to a 3-bucket quality (high/medium/low), so two \"high\" models can't be separated on capability alone — declare quality/tiers/tierMap or use the LLM pick for reliable hard-vs-easy routing.\nLLM classifier latency/cost. The llm strategy adds one cheap call per uncached turn; prefer a small, fast, non-Gemini model and use heuristic where determinism matters.\nJev confidence is a gate, not a score. Unlike the llm strategy — whose self-reported confidence defaults to a hard-coded 0.7 when the model omits it — jev returns the confidence the decision provider reports (calibrated, on TypeSafe's Jev) rather than a constant, which is what makes minConfidence meaningful. The two bars differ because the mistakes cost differently: spending more on a wrong guess wastes money, spending less produces a wrong answer. Set either above 1 to force the heuristic while leaving the strategy configured.\nGemini tools + JSON schema. The classifier call uses a schema with tools disabled, so the Gemini exclusivity rule doesn't apply to it; when routing a tools + structured-output request, prefer a non-Gemini target model.\nPrompt privacy (llm and jev). Both strategies send a truncated copy of the prompt off-machine — to the classifier model, or to the configured decision provider — so the same data-handling and retention considerations as any provider call apply. heuristic keeps classification fully in-process (no prompt leaves your environment); prefer it where that matters, and note that auto selects jev as soon as a decision provider's key is present, including a PERPLEXITY_API_KEY set only for the Perplexity text provider.","hierarchy":{"lvl0":"Features","lvl1":"Classifier Router","lvl2":"Caveats","lvl3":""}},
2647
+ {"objectID":"fac896bd6ef5ba1295df2e3b8249209a904d5a4b66780e388934069517ed6014","title":"Caveats","url":"/docs/features/classifier-router#caveats","content":"Registry quality is coarse. Auto-enrichment maps a model to a 3-bucket quality (high/medium/low), so two \"high\" models can't be separated on capability alone — declare quality/tiers/tierMap or use the LLM pick for reliable hard-vs-easy routing.\nLLM classifier latency/cost. The llm strategy adds one cheap call per uncached turn; prefer a small, fast, non-Gemini model and use heuristic where determinism matters.\nJev confidence is a gate, not a score. Unlike the llm strategy — whose self-reported confidence defaults to a hard-coded 0.7 when the model omits it — jev returns the confidence the decision provider reports (calibrated, on TypeSafe's Jev) rather than a constant, which is what makes minConfidence meaningful. The two bars differ because the mistakes cost differently: spending more on a wrong guess wastes money, spending less produces a wrong answer. Set either above 1 to force the heuristic while leaving the strategy configured.\nGemini tools + JSON schema. The classifier call uses a schema with tools disabled, so the Gemini exclusivity rule doesn't apply to it; when routing a tools + structured-output request, prefer a non-Gemini target model.\nPrompt privacy (llm and jev). Both strategies send a truncated copy of the prompt off-machine — to the classifier model, or to the configured decision provider — so the same data-handling and retention considerations as any provider call apply. heuristic keeps classification fully in-process (no prompt leaves your environment); prefer it where that matters, and note that auto selects jev as soon as a decision provider's key is present, including a PERPLEXITY_API_KEY set only for the Perplexity text provider, or a CLOUDFLARE_API_KEY and CLOUDFLARE_ACCOUNT_ID set only for the Cloudflare Workers AI text provider.","hierarchy":{"lvl0":"Features","lvl1":"Classifier Router","lvl2":"Caveats","lvl3":""}},
2648
2648
  {"objectID":"44f6d5bc4c0e752129323d87508ca78b8f296e06d39a994f0b93706adff75033","title":"See also","url":"/docs/features/classifier-router#see-also","content":"Provider Orchestration & Model Pool\nProvider Fallback\nPer-Request Credentials\nModel routing with a decision model\nThe model catalogue","hierarchy":{"lvl0":"Features","lvl1":"Classifier Router","lvl2":"See also","lvl3":""}},
2649
2649
  {"objectID":"2c5c9a8afc6acb37fc955bbc5ed44588962efef98dbf7e9dd26126eb58d08241","title":"Claude Proxy Architecture","url":"/docs/features/claude-proxy-architecture","content":"Claude Proxy Architecture\nSystem Overview\n\nThe Claude proxy is a local HTTP server that sits between Claude Code and the Anthropic API. It provides multi-account rotation, automatic token refresh, rate-limit handling with exponential backoff, and optional model translation to non-Anthropic providers.\n\nTwo operational modes\n\n| Mode | When | What happens |\n| --------------- | ---------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |\n| Passthrough | Target provider is anthropic (or null) | The request body is forwarded byte-for-byte to api.anthropic.com via plain fetch() with client headers forwarded. No parsing, no tool injection, no SDK involvement. |\n| Translation | Target provider is anything else (e.g. vertex, openai) | The Claude-format request is parsed by parseClaudeRequest(), routed through ctx.neurolink.stream() / ctx.neurolink.generate(), and the NeuroLink response is serialized back to Claude SSE format via ClaudeStreamSerializer. |\n\nPassthrough exists because Claude Code sends complex bodies (multi-turn conversations, tool definitions, thinking blocks, context management betas) that would be lossy to parse and re-serialize. The proxy's job for Claude-to-Claude is purely auth and account management.\n\nHow it fits into NeuroLink\n\nThe proxy is started via the CLI (neurolink proxy start) and creates a Hono HTTP server. It registers routes from createClaudeProxyRoutes() and injects a live NeuroLink SDK instance into the request context for translation-mode and fallback paths. MCP initialization is explicitly skipped (NEUROLINK_SKIP_MCP=true) because tools come from Claude Code, not from MCP servers.\nRequest Lifecycle\n\nA complete request through the passthrough path:\nAccount Management\n\nAccount loading priority\n\nAccounts are loaded in the POST /v1/messages handler on every request (not cached across requests), in this order:\nTokenStore compound keys (anthropic:<label>) — The primary source. tokenStore.listProviders() returns all stored keys; those starting with anthropic: are loaded via tokenStore.loadTokens(key). Each yields { accessToken, refreshToken, expiresAt }.\nLegacy credentials file (~/.neurolink/anthropic-credentials.json) — Only checked when zero compound keys exist. Reads creds.oauth.accessToken directly from JSON.\nEnvironment variable (ANTHROPIC_API_KEY) — Only used when no OAuth accounts were found at all. Creates a single api_key-type account.\n\nAccount selection: strategy-driven with fill-first default\n\nThe request handler supports two real account-selection strategies:\nfill-first (default) — always begin with the current primary account and stay on it until it cools down or fails.\nround-robin — rotate the starting account on each request, then try the remaining accounts sequentially.\n\nExpired accounts are pruned at startup via tokenStore.pruneExpired() (one-time). Accounts that are persisted as disabled (via tokenStore.isDisabled()) are skipped. Expired tokens with a refresh token get one refresh attempt at startup; on failure, the account is disabled until re-authentication.\n\nThe CLI --strategy flag and the proxy config routing.strategy field both map directly to this account ordering logic. There are only two supported values today: fill-first and round-robin.\n\nPer-status cooldowns\n\n| HTTP Status | Cooldown | Behavior |\n| ------------------------------------ | ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- |\n| 429 (rate limit) | Exponential backoff (see below) | Continue to next account |\n| 401/402/403 (auth failure) | 5 minutes (AUTH_COOLDOWN_MS) | Attempt token refresh first (up to 5 retries); if all fail, cooldown and continue. After 15 consecutive refresh failures, account permanently disabled. |\n| 404 (not found) | None | Return error immediately (no failover) |\n| 5xx, 52x (t","hierarchy":{"lvl0":"Features","lvl1":"Claude Proxy Architecture","lvl2":"","lvl3":""}},
2650
2650
  {"objectID":"9992141969480e207aca925aec7f3d174e0dda96fa7b13097cf2b0d671cf7114","title":"1. System Overview","url":"/docs/features/claude-proxy-architecture#1-system-overview","content":"The Claude proxy is a local HTTP server that sits between Claude Code and the Anthropic API. It provides multi-account rotation, automatic token refresh, rate-limit handling with exponential backoff, and optional model translation to non-Anthropic providers.","hierarchy":{"lvl0":"Features","lvl1":"Claude Proxy Architecture","lvl2":"1. System Overview","lvl3":""}},
@@ -3240,7 +3240,7 @@
3240
3240
  {"objectID":"a18da32e2f63d235059c36cf45cff486c8f1dc0c4d4e55ee5f239dd89ede5d21","title":"BudgetCheckResult","url":"/docs/features/context-compaction#budgetcheckresult","content":"Returned by checkContextBudget().","hierarchy":{"lvl0":"Features","lvl1":"Context Compaction","lvl2":"BudgetCheckResult","lvl3":""}},
3241
3241
  {"objectID":"1a2b7ee1dace651a8300809eec11e00b29da084cbfe5b040e4760d58ea0b5879","title":"BudgetCheckParams","url":"/docs/features/context-compaction#budgetcheckparams","content":"Parameters for checkContextBudget().\n\nSource: src/lib/context/budgetChecker.ts:18-54","hierarchy":{"lvl0":"Features","lvl1":"Context Compaction","lvl2":"BudgetCheckParams","lvl3":""}},
3242
3242
  {"objectID":"4bad23fd2858cc71b5af8506276229485c7d0a9636cfddae5c5d9df400bc28b8","title":"The 5-Stage Pipeline","url":"/docs/features/context-compaction#the-5-stage-pipeline","content":"The ContextCompactor runs stages sequentially. Each stage only runs if the\nprevious stage didn't bring tokens below the target budget.\n\n| # | Stage | CompactionStage | Needs a decision model |\n| --- | ------------------------- | ----------------- | -------------------------------------- |\n| 0 | Relevance drop | relevance | yes — skipped entirely without one |\n| 1 | Tool output pruning | prune | no |\n| 2 | File read deduplication | deduplicate | no |\n| 3 | LLM summarization | summarize | no (its gate uses one) |\n| 4 | Sliding window truncation | truncate | no |\n\nresult.stagesUsed reports the ones that actually ran, in order.","hierarchy":{"lvl0":"Features","lvl1":"Context Compaction","lvl2":"The 5-Stage Pipeline","lvl3":""}},
3243
- {"objectID":"918fce938e3acd086621a86be0faaaecccf78059d8a0ef53b2dbc4fcd87347a1","title":"Stage 0: Relevance drop","url":"/docs/features/context-compaction#stage-0-relevance-drop","content":"File: src/lib/context/contextDecision.ts\n\nEverything below Stage 0 is chronological: the pipeline's only notion of\n\"droppable\" is \"old\". Stage 0 is the one stage that asks what a message is\nfor — one boolean per message (\"is this needed to answer the current\nrequest?\") in a single batch. On TypeSafe that costs the same for 200 messages\nas for one because decision latency is flat in question count; on Perplexity each\nfurther question adds about 65 ms and a request takes at most 128 (see\nits guide).\n\nIt is strictly additive. With no decision provider configured the stage\ndoes not run, stagesUsed omits relevance, and the pipeline behaves exactly\nas the four-stage one always did. It is also bounded by maxDropRatio and\nwalks oldest-first, so when the cap binds it spares the newest candidates —\nthe same recency assumption every other stage makes.","hierarchy":{"lvl0":"Features","lvl1":"Context Compaction","lvl2":"Stage 0: Relevance drop","lvl3":""}},
3243
+ {"objectID":"918fce938e3acd086621a86be0faaaecccf78059d8a0ef53b2dbc4fcd87347a1","title":"Stage 0: Relevance drop","url":"/docs/features/context-compaction#stage-0-relevance-drop","content":"File: src/lib/context/contextDecision.ts\n\nEverything below Stage 0 is chronological: the pipeline's only notion of\n\"droppable\" is \"old\". Stage 0 is the one stage that asks what a message is\nfor — one boolean per message (\"is this needed to answer the current\nrequest?\") in a single batch. On TypeSafe that costs the same for 200 messages\nas for one because decision latency is flat in question count; on Perplexity each\nfurther question adds about 65 ms and a request takes at most 128 (see\nits guide); on\nCloudflare Clef a request takes at most 64 questions, which tryDecide() splits\ninto batches of 64 (see its guide).\nThe Clef endpoint also ignores state text past about 2,048 tokens, and NeuroLink\nrefuses a state it estimates at more than 1,500 tokens, so a longer set of\neligible messages gets a refusal and the stage is skipped, as on any failure.\n\nIt is strictly additive. With no decision provider configured the stage\ndoes not run, stagesUsed omits relevance, and the pipeline behaves exactly\nas the four-stage one always did. It is also bounded by maxDropRatio and\nwalks oldest-first, so when the cap binds it spares the newest candidates —\nthe same recency assumption every other stage makes.","hierarchy":{"lvl0":"Features","lvl1":"Context Compaction","lvl2":"Stage 0: Relevance drop","lvl3":""}},
3244
3244
  {"objectID":"c1fa9b4e362150149aa7889fbd1374625457cdc5cba6f57fb15d82ef305fc492","title":"The summary gate","url":"/docs/features/context-compaction#the-summary-gate","content":"Stage 3 used to accept any non-empty string as a summary. When a decision\nmodel is configured, the generated summary is now checked first (\"does this\npreserve every decision and open question?\") and a rejected summary leaves the\nmessages untouched so a later stage can try instead. The rejection is recorded\non the span as compaction.stage3.summaryRejected, because a gate that\nsilently discarded work would be indistinguishable from one that never ran.\n\nRejection is deliberately rare: the gate exists to catch a summary that lost a\ndecision, not to second-guess wording.","hierarchy":{"lvl0":"Features","lvl1":"Context Compaction","lvl2":"The summary gate","lvl3":""}},
3245
3245
  {"objectID":"2a518830a84553879be50318dfc47b77a34bf2c61cb5bc21a443ad5a5fcba95b","title":"Stage 1: Tool Output Pruning","url":"/docs/features/context-compaction#stage-1-tool-output-pruning","content":"File: src/lib/context/stages/toolOutputPruner.ts\n\nWalks messages backwards, protecting the most recent tool outputs, and replaces older tool results with \"[Tool result cleared]\".\n\nPruneConfig:\n\n| Field | Type | Default | Description |\n| ---------------- | ---------- | ----------- | ----------------------------------------------------------- |\n| protectTokens | number | 40,000 | Token budget of recent tool outputs to protect from pruning |\n| minimumSavings | number | 20,000 | Minimum tokens that must be saved for pruning to be applied |\n| protectedTools | string[] | [\"skill\"] | Tool names that are never pruned |\n| provider | string | — | Provider name for token estimation multiplier |\n\nPruneResult:","hierarchy":{"lvl0":"Features","lvl1":"Context Compaction","lvl2":"Stage 1: Tool Output Pruning","lvl3":""}},
3246
3246
  {"objectID":"7c30f222482767920767c8a3dee1c3955d4b615a31c54b58851b467e409a9e64","title":"Stage 2: File Read Deduplication","url":"/docs/features/context-compaction#stage-2-file-read-deduplication","content":"File: src/lib/context/stages/fileReadDeduplicator.ts\n\nDetects multiple reads of the same file path. Keeps only the latest read, replaces earlier reads with \"[File <path> - refer to latest read below]\".\n\nDeduplicationResult:\n\nFile read detection uses the regex pattern: /(?:read|reading|read_file|readFile|Read file|cat)\\s+['\"]?([^\\s'\"\\n]+)/i\n\nA 30% savings threshold (DEDUP_THRESHOLD = 0.3) must be met for deduplication to be applied.","hierarchy":{"lvl0":"Features","lvl1":"Context Compaction","lvl2":"Stage 2: File Read Deduplication","lvl3":""}},
@@ -3351,24 +3351,25 @@
3351
3351
  {"objectID":"a78bcf758e0d2cbf2d2e9ad39bb04c851848618fb55718447904be31d350055b","title":"Limitations","url":"/docs/features/csv-support#limitations","content":"Max file size: 10MB by default (configurable)\nMax rows: 1000 by default (configurable)\nEncoding: auto-detected via BOM + chardet (UTF-8 / UTF-16 / Windows-1252 / Latin-1 …), or forced with encoding (#362)\nPer-row size: a single row is capped at 10MB to bound memory; larger rows fail fast with a clear error (#371)\nParse timeout: parsing is time-bounded (30s strings / 5min files); on timeout partial rows are returned with metadata.parseTimedOut (#379)\nRow shape: parsed rows are validated to be string-keyed objects with string values; a malformed row aborts the parse with [CSVProcessor] Invalid CSV row <n> (#384)\nToken limits: Large CSV files may exceed provider token limits\nStreaming: CSV content is parsed and formatted before sending (not streamed to LLM)","hierarchy":{"lvl0":"Features","lvl1":"CSV File Support","lvl2":"Limitations","lvl3":""}},
3352
3352
  {"objectID":"1e48831d1c4fc775af2633d0340bdeb3d11c96abc213cbdd814a53d14ce500df","title":"Related Features","url":"/docs/features/csv-support#related-features","content":"Office Documents: DOCX, PPTX, XLSX processing\nPDF Support: PDF document processing\nImage Support: Similar multimodal input for images\nFile Detection: Auto-detect file types with confidence scores\nMemory Efficient: Streaming parser for large files\nProvider Agnostic: Works across supported providers\nCLI Integration: Full CLI support with options","hierarchy":{"lvl0":"Features","lvl1":"CSV File Support","lvl2":"Related Features","lvl3":""}},
3353
3353
  {"objectID":"19734b9ffaeb8376dfcebd944730a4cedae77bc1e7ed769733b8b71b1758c421","title":"Summary","url":"/docs/features/csv-support#summary","content":"CSV support is multimodal input (like images)\nUse csvFiles array or files array (auto-detect)\nCustomize with csvOptions (maxRows, formatStyle, includeHeaders)\nWorks across supported providers (not just vision models)\nMemory efficient streaming parser\nCLI support with --csv, --file, --csv-max-rows, --csv-format\nOnly types exposed from package (not classes)","hierarchy":{"lvl0":"Features","lvl1":"CSV File Support","lvl2":"Summary","lvl3":""}},
3354
- {"objectID":"fce0b532226adb069bbeb5459dc5e4ec28037617ca805845768d918653ddbfd5","title":"The `decide` inference type","url":"/docs/features/decide-inference-type","content":"The decide inference type\n\nDeep-dive: generate, stream, decide: a third inference type for NeuroLink —\nwhy this shipped as a provider rather than a subsystem, the five call sites, and the four\ntransport bugs found by adding the Vercel AI Gateway (two of which produced plausible output).\n\nNeuroLink recognises three inference types. Two of them produce text:\n\n| Type | Call | Produces |\n| ------------ | ------------------------ | ------------------------------ |\n| generate | neurolink.generate() | text |\n| stream | neurolink.stream() | text, incrementally |\n| decide | neurolink.decide() | typed judgements — no text |\n\nA decision model takes one state plus a map of named, typed questions and\nreturns one typed answer per question, all evaluated in a single parallel pass.\nThere is no text anywhere in the response, so nothing has to be parsed back out\nof prose. The decide providers are TypeSafe's Jev, Convai Innovations'\nopen-weights Laya, Juspay's open-weights XOR, which also reads images and\nvideo, and Perplexity's hosted Decisions API, which also reads images.\n\nThis is not neurolink.evaluate(), which scores an\nalready-generated response with RAGAS scorers. Different feature, different\nword.\n\nThe three primitives\n\n| Type | Question | Answer fields |\n| --------- | ------------------------------ | ------------------------------------------------ |\n| boolean | Is this statement true? | probability (0–1) — no confidence |\n| choice | Which option from this set? | choice, probabilities, confidence |\n| score | Rate against an ordered rubric | score, legend, probabilities, confidence |\n\nAll three mix freely in one call.\n\nVocabulary note. TypeSafe calls the yes/no primitive a noul and answers\nit in a field of the same name. The Vercel AI SDK and Pydantic AI both renamed\nthat to boolean/probability when exposing it, and NeuroLink follows them —\nthe vendor's spelling is translated inside TypeSafeProvider, so another\ndecision provider slots in without changing any call site.\n\nConfidence is not probability\n\nprobabilities says what the model thinks. confidence says _whether you\nshould act on it_. On TypeSafe's Jev it is calibrated — derived from the\ndistribution, not self-reported — which is what makes it usable as a gate. Each\nprovider reports its own: Perplexity's is, in Perplexity's words, the model's own\ncertainty estimate and not the top probability, Perplexity does not call it\ncalibrated, and NeuroLink has not measured whether it is, so tune a threshold on\nyour own data (see the\nPerplexity guide).\n\nCalibration is a property of groups of answers, not a promise about any\none. Where a confidence is calibrated, answers scored 0.8 are right about 80%\nof the time across many answers. It does not mean a specific 0.8 answer is right.\n\nA boolean carries no confidence of its own. Use decisionBooleanConfidence(p)\n— distance from a coin flip, so 0.5 → 0 and 0/1 → 1. Note also that a boolean\nand an equivalent two-option choice are not guaranteed to agree, and\ncomplementary booleans do not reliably sum to 1, so a threshold tuned on one\nquestion shape does not transfer to another.\n\nEnabling it\n\nGet a key at console.typesafe.ai/keys.\n\nLaya, at a Laya server or a LiteLLM proxy route\n\nLaya has no built-in endpoint. The base URL and key can equally come from the\nconfig passed to the SDK, new NeuroLink({ credentials: { laya: { baseURL, apiKey } } }),\nor per call; config set there counts when NeuroLink picks the default decision\nprovider, exactly as the environment does.\n\nWhen a TypeSafe key (TYPESAFE_API_KEY or AI_GATEWAY_API_KEY) is configured\nalongside Laya, TypeSafe is the default; Laya runs where a caller names it\n(provider: \"laya\") or when neither TypeSafe key is set. Laya counts as\nconfigured only with both its key and its base URL. See the\nLaya provider guide.\n\nXOR, at a deployment or a LiteLLM proxy route\n\nXOR, Juspay's open-weights decision model, has no built-in endpoint either.\nNeuroLink calls <base URL>/v1/systemone, and a trailing /v1 on the base URL\nis accepted. The base URL and key can equally come from the config passed to the\nSDK, new NeuroLink({ credentials: { xor: { baseURL, apiKey } } }), or per\ncall; config set there counts when NeuroLink picks the default decision\nprovider, exactly as the environment does. On a LiteLLM proxy the key's team\nmust allow xor-1.1, otherwise the proxy answers 403 team_model_access_denied.\n\nBuilt-in features use the first configured decision provider in the order\nTypeSafe, Laya, XOR, Perplexity. XOR counts as configured only with both its key\nand its base URL, and runs where a caller names it (provider: \"xor\") or when\nneither TypeSafe nor Laya is configured. It reads images and video; see\nImages and video and the\nXOR provider guide.\n\nPerplexity, at its hosted endpoint\n\nPerplexity's pplx-decider-v1-27b i","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"","lvl3":""}},
3355
- {"objectID":"f91e7ef7473aa411ff8c0d3b04381325781477c1876408ffd5469dbaf2d9584c","title":"The decide inference type","url":"/docs/features/decide-inference-type#the-decide-inference-type","content":"Deep-dive: generate, stream, decide: a third inference type for NeuroLink —\nwhy this shipped as a provider rather than a subsystem, the five call sites, and the four\ntransport bugs found by adding the Vercel AI Gateway (two of which produced plausible output).\n\nNeuroLink recognises three inference types. Two of them produce text:\n\n| Type | Call | Produces |\n| ------------ | ------------------------ | ------------------------------ |\n| generate | neurolink.generate() | text |\n| stream | neurolink.stream() | text, incrementally |\n| decide | neurolink.decide() | typed judgements — no text |\n\nA decision model takes one state plus a map of named, typed questions and\nreturns one typed answer per question, all evaluated in a single parallel pass.\nThere is no text anywhere in the response, so nothing has to be parsed back out\nof prose. The decide providers are TypeSafe's Jev, Convai Innovations'\nopen-weights Laya, Juspay's open-weights XOR, which also reads images and\nvideo, and Perplexity's hosted Decisions API, which also reads images.\n\nThis is not neurolink.evaluate(), which scores an\nalready-generated response with RAGAS scorers. Different feature, different\nword.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"The decide inference type","lvl3":""}},
3354
+ {"objectID":"fce0b532226adb069bbeb5459dc5e4ec28037617ca805845768d918653ddbfd5","title":"The `decide` inference type","url":"/docs/features/decide-inference-type","content":"The decide inference type\n\nDeep-dive: generate, stream, decide: a third inference type for NeuroLink —\nwhy this shipped as a provider rather than a subsystem, the five call sites, and the four\ntransport bugs found by adding the Vercel AI Gateway (two of which produced plausible output).\n\nNeuroLink recognises three inference types. Two of them produce text:\n\n| Type | Call | Produces |\n| ------------ | ------------------------ | ------------------------------ |\n| generate | neurolink.generate() | text |\n| stream | neurolink.stream() | text, incrementally |\n| decide | neurolink.decide() | typed judgements — no text |\n\nA decision model takes one state plus a map of named, typed questions and\nreturns one typed answer per question, all evaluated in a single parallel pass.\nThere is no text anywhere in the response, so nothing has to be parsed back out\nof prose. The decide providers are TypeSafe's Jev, Convai Innovations'\nopen-weights Laya, Juspay's open-weights XOR, which also reads images and\nvideo, Perplexity's hosted Decisions API, which also reads images, and\nCloudflare's Clef models on Workers AI, which also read images.\n\nThis is not neurolink.evaluate(), which scores an\nalready-generated response with RAGAS scorers. Different feature, different\nword.\n\nThe three primitives\n\n| Type | Question | Answer fields |\n| --------- | ------------------------------ | ------------------------------------------------ |\n| boolean | Is this statement true? | probability (0–1) — no confidence |\n| choice | Which option from this set? | choice, probabilities, confidence |\n| score | Rate against an ordered rubric | score, legend, probabilities, confidence |\n\nAll three mix freely in one call.\n\nVocabulary note. TypeSafe calls the yes/no primitive a noul and answers\nit in a field of the same name. The Vercel AI SDK and Pydantic AI both renamed\nthat to boolean/probability when exposing it, and NeuroLink follows them —\nthe vendor's spelling is translated inside TypeSafeProvider, so another\ndecision provider slots in without changing any call site.\n\nConfidence is not probability\n\nprobabilities says what the model thinks. confidence says _whether you\nshould act on it_. On TypeSafe's Jev it is calibrated — derived from the\ndistribution, not self-reported — which is what makes it usable as a gate. Each\nprovider reports its own: Perplexity's is, in Perplexity's words, the model's own\ncertainty estimate and not the top probability, Perplexity does not call it\ncalibrated, and NeuroLink has not measured whether it is, so tune a threshold on\nyour own data (see the\nPerplexity guide).\n\nCalibration is a property of groups of answers, not a promise about any\none. Where a confidence is calibrated, answers scored 0.8 are right about 80%\nof the time across many answers. It does not mean a specific 0.8 answer is right.\n\nA boolean carries no confidence of its own. Use decisionBooleanConfidence(p)\n— distance from a coin flip, so 0.5 → 0 and 0/1 → 1. Note also that a boolean\nand an equivalent two-option choice are not guaranteed to agree, and\ncomplementary booleans do not reliably sum to 1, so a threshold tuned on one\nquestion shape does not transfer to another.\n\nEnabling it\n\nGet a key at console.typesafe.ai/keys.\n\nLaya, at a Laya server or a LiteLLM proxy route\n\nLaya has no built-in endpoint. The base URL and key can equally come from the\nconfig passed to the SDK, new NeuroLink({ credentials: { laya: { baseURL, apiKey } } }),\nor per call; config set there counts when NeuroLink picks the default decision\nprovider, exactly as the environment does.\n\nWhen a TypeSafe key (TYPESAFE_API_KEY or AI_GATEWAY_API_KEY) is configured\nalongside Laya, TypeSafe is the default; Laya runs where a caller names it\n(provider: \"laya\") or when neither TypeSafe key is set. Laya counts as\nconfigured only with both its key and its base URL. See the\nLaya provider guide.\n\nXOR, at a deployment or a LiteLLM proxy route\n\nXOR, Juspay's open-weights decision model, has no built-in endpoint either.\nNeuroLink calls <base URL>/v1/systemone, and a trailing /v1 on the base URL\nis accepted. The base URL and key can equally come from the config passed to the\nSDK, new NeuroLink({ credentials: { xor: { baseURL, apiKey } } }), or per\ncall; config set there counts when NeuroLink picks the default decision\nprovider, exactly as the environment does. On a LiteLLM proxy the key's team\nmust allow xor-1.1, otherwise the proxy answers 403 team_model_access_denied.\n\nBuilt-in features use the first configured decision provider in the order\nTypeSafe, Laya, XOR, Perplexity, Cloudflare Clef. XOR counts as configured only with both its key\nand its base URL, and runs where a caller names it (provider: \"xor\") or when\nneither TypeSafe nor Laya is configured. It reads images and video; see\nImages and video and the\nXOR provid","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"","lvl3":""}},
3355
+ {"objectID":"f91e7ef7473aa411ff8c0d3b04381325781477c1876408ffd5469dbaf2d9584c","title":"The decide inference type","url":"/docs/features/decide-inference-type#the-decide-inference-type","content":"Deep-dive: generate, stream, decide: a third inference type for NeuroLink —\nwhy this shipped as a provider rather than a subsystem, the five call sites, and the four\ntransport bugs found by adding the Vercel AI Gateway (two of which produced plausible output).\n\nNeuroLink recognises three inference types. Two of them produce text:\n\n| Type | Call | Produces |\n| ------------ | ------------------------ | ------------------------------ |\n| generate | neurolink.generate() | text |\n| stream | neurolink.stream() | text, incrementally |\n| decide | neurolink.decide() | typed judgements — no text |\n\nA decision model takes one state plus a map of named, typed questions and\nreturns one typed answer per question, all evaluated in a single parallel pass.\nThere is no text anywhere in the response, so nothing has to be parsed back out\nof prose. The decide providers are TypeSafe's Jev, Convai Innovations'\nopen-weights Laya, Juspay's open-weights XOR, which also reads images and\nvideo, Perplexity's hosted Decisions API, which also reads images, and\nCloudflare's Clef models on Workers AI, which also read images.\n\nThis is not neurolink.evaluate(), which scores an\nalready-generated response with RAGAS scorers. Different feature, different\nword.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"The decide inference type","lvl3":""}},
3356
3356
  {"objectID":"2c0f154b9ae29f0af1a097de9ccb435db1e8f5c35a4df8bb91526eadb7b5eb97","title":"The three primitives","url":"/docs/features/decide-inference-type#the-three-primitives","content":"| Type | Question | Answer fields |\n| --------- | ------------------------------ | ------------------------------------------------ |\n| boolean | Is this statement true? | probability (0–1) — no confidence |\n| choice | Which option from this set? | choice, probabilities, confidence |\n| score | Rate against an ordered rubric | score, legend, probabilities, confidence |\n\nAll three mix freely in one call.\n\nVocabulary note. TypeSafe calls the yes/no primitive a noul and answers\nit in a field of the same name. The Vercel AI SDK and Pydantic AI both renamed\nthat to boolean/probability when exposing it, and NeuroLink follows them —\nthe vendor's spelling is translated inside TypeSafeProvider, so another\ndecision provider slots in without changing any call site.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"The three primitives","lvl3":""}},
3357
3357
  {"objectID":"ff28050e2ddf3c21c910bc84ca32105968e1f7ffdbc3b19a983bb9adc6e9cc36","title":"Confidence is not probability","url":"/docs/features/decide-inference-type#confidence-is-not-probability","content":"probabilities says what the model thinks. confidence says _whether you\nshould act on it_. On TypeSafe's Jev it is calibrated — derived from the\ndistribution, not self-reported — which is what makes it usable as a gate. Each\nprovider reports its own: Perplexity's is, in Perplexity's words, the model's own\ncertainty estimate and not the top probability, Perplexity does not call it\ncalibrated, and NeuroLink has not measured whether it is, so tune a threshold on\nyour own data (see the\nPerplexity guide).\n\nCalibration is a property of groups of answers, not a promise about any\none. Where a confidence is calibrated, answers scored 0.8 are right about 80%\nof the time across many answers. It does not mean a specific 0.8 answer is right.\n\nA boolean carries no confidence of its own. Use decisionBooleanConfidence(p)\n— distance from a coin flip, so 0.5 → 0 and 0/1 → 1. Note also that a boolean\nand an equivalent two-option choice are not guaranteed to agree, and\ncomplementary booleans do not reliably sum to 1, so a threshold tuned on one\nquestion shape does not transfer to another.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"Confidence is not probability","lvl3":""}},
3358
3358
  {"objectID":"91bfa1275464235c40e00d8a9ac0805e9e181450cccb6235fabd43178bef7c02","title":"Enabling it","url":"/docs/features/decide-inference-type#enabling-it","content":"Get a key at console.typesafe.ai/keys.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"Enabling it","lvl3":""}},
3359
3359
  {"objectID":"71509585343bb3860ff000ad67405f6ed19fbc26f8eb973ac1fe58088af460b4","title":"Laya, at a Laya server or a LiteLLM proxy route","url":"/docs/features/decide-inference-type#laya-at-a-laya-server-or-a-litellm-proxy-route","content":"Laya has no built-in endpoint. The base URL and key can equally come from the\nconfig passed to the SDK, new NeuroLink({ credentials: { laya: { baseURL, apiKey } } }),\nor per call; config set there counts when NeuroLink picks the default decision\nprovider, exactly as the environment does.\n\nWhen a TypeSafe key (TYPESAFE_API_KEY or AI_GATEWAY_API_KEY) is configured\nalongside Laya, TypeSafe is the default; Laya runs where a caller names it\n(provider: \"laya\") or when neither TypeSafe key is set. Laya counts as\nconfigured only with both its key and its base URL. See the\nLaya provider guide.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"Laya, at a Laya server or a LiteLLM proxy route","lvl3":""}},
3360
- {"objectID":"19994573e52cd09183db8050eadaea8ece3e291ef40c3009a5c44dfa959bf99c","title":"XOR, at a deployment or a LiteLLM proxy route","url":"/docs/features/decide-inference-type#xor-at-a-deployment-or-a-litellm-proxy-route","content":"XOR, Juspay's open-weights decision model, has no built-in endpoint either.\nNeuroLink calls <base URL>/v1/systemone, and a trailing /v1 on the base URL\nis accepted. The base URL and key can equally come from the config passed to the\nSDK, new NeuroLink({ credentials: { xor: { baseURL, apiKey } } }), or per\ncall; config set there counts when NeuroLink picks the default decision\nprovider, exactly as the environment does. On a LiteLLM proxy the key's team\nmust allow xor-1.1, otherwise the proxy answers 403 team_model_access_denied.\n\nBuilt-in features use the first configured decision provider in the order\nTypeSafe, Laya, XOR, Perplexity. XOR counts as configured only with both its key\nand its base URL, and runs where a caller names it (provider: \"xor\") or when\nneither TypeSafe nor Laya is configured. It reads images and video; see\nImages and video and the\nXOR provider guide.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"XOR, at a deployment or a LiteLLM proxy route","lvl3":""}},
3361
- {"objectID":"dc8b9e9a2c66b1d2a60305a09a44263ceca755d40f04b6771e0a7c05a5237a64","title":"Perplexity, at its hosted endpoint","url":"/docs/features/decide-inference-type#perplexity-at-its-hosted-endpoint","content":"Perplexity's pplx-decider-v1-27b is a hosted API at a public endpoint, so a key\nalone configures it and there is no base URL to set. NeuroLink calls\n<base URL>/v1/decisions; a trailing /v1 on an override is accepted. The key\ncan equally come from the config passed to the SDK,\nnew NeuroLink({ credentials: { perplexityDecider: { apiKey } } }), or per call;\nconfig set there counts when a decide() call picks the default decision\nprovider, exactly as the environment does. The classifier router's auto gate is\nread differently: it looks at the constructor's credentials and the environment\nonly, so a credentials.perplexityDecider passed on one call does not influence\nit. The provider id is perplexity-decider: the perplexity provider is the\nSonar text provider, which serves generate() and stream() and not decide().\n\nPERPLEXITY_API_KEY is shared with that text provider. A host that set it\nonly to use Sonar has therefore also configured decide, and built-in features\nwill use Perplexity whenever none of TypeSafe, Laya or XOR is configured.\nPerplexity comes after XOR in descriptor order, so it never displaces one that\nis. Context compaction's relevance stage and summary gate need no opt-in of their\nown, so earlier conversation text starts going to Perplexity the first time a\nconversation outgrows its budget; model routing, tool routing and RAG planning do\nnothing unless enabled. To keep the shared key from activating decisions, pass\nthe text provider's key as credentials.perplexity instead of through the\nenvironment, and keep it out of .env too, because the SDK and the CLI load that\nfile into the environment, or configure one of TypeSafe, Laya and XOR, which then\nreceives those texts instead. The Perplexity provider guide lists\nwhat each consumer sends\nand the\nexact switches.\nIt reads images but no video; see Images and video.\n\nThe degradation contract.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"Perplexity, at its hosted endpoint","lvl3":""}},
3360
+ {"objectID":"19994573e52cd09183db8050eadaea8ece3e291ef40c3009a5c44dfa959bf99c","title":"XOR, at a deployment or a LiteLLM proxy route","url":"/docs/features/decide-inference-type#xor-at-a-deployment-or-a-litellm-proxy-route","content":"XOR, Juspay's open-weights decision model, has no built-in endpoint either.\nNeuroLink calls <base URL>/v1/systemone, and a trailing /v1 on the base URL\nis accepted. The base URL and key can equally come from the config passed to the\nSDK, new NeuroLink({ credentials: { xor: { baseURL, apiKey } } }), or per\ncall; config set there counts when NeuroLink picks the default decision\nprovider, exactly as the environment does. On a LiteLLM proxy the key's team\nmust allow xor-1.1, otherwise the proxy answers 403 team_model_access_denied.\n\nBuilt-in features use the first configured decision provider in the order\nTypeSafe, Laya, XOR, Perplexity, Cloudflare Clef. XOR counts as configured only with both its key\nand its base URL, and runs where a caller names it (provider: \"xor\") or when\nneither TypeSafe nor Laya is configured. It reads images and video; see\nImages and video and the\nXOR provider guide.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"XOR, at a deployment or a LiteLLM proxy route","lvl3":""}},
3361
+ {"objectID":"dc8b9e9a2c66b1d2a60305a09a44263ceca755d40f04b6771e0a7c05a5237a64","title":"Perplexity, at its hosted endpoint","url":"/docs/features/decide-inference-type#perplexity-at-its-hosted-endpoint","content":"Perplexity's pplx-decider-v1-27b is a hosted API at a public endpoint, so a key\nalone configures it and there is no base URL to set. NeuroLink calls\n<base URL>/v1/decisions; a trailing /v1 on an override is accepted. The key\ncan equally come from the config passed to the SDK,\nnew NeuroLink({ credentials: { perplexityDecider: { apiKey } } }), or per call;\nconfig set there counts when a decide() call picks the default decision\nprovider, exactly as the environment does. The classifier router's auto gate is\nread differently: it looks at the constructor's credentials and the environment\nonly, so a credentials.perplexityDecider passed on one call does not influence\nit. The provider id is perplexity-decider: the perplexity provider is the\nSonar text provider, which serves generate() and stream() and not decide().\n\nPERPLEXITY_API_KEY is shared with that text provider. A host that set it\nonly to use Sonar has therefore also configured decide, and built-in features\nwill use Perplexity whenever none of TypeSafe, Laya or XOR is configured.\nPerplexity comes after XOR in descriptor order, so it never displaces one that\nis. Context compaction's relevance stage and summary gate need no opt-in of their\nown, so earlier conversation text starts going to Perplexity the first time a\nconversation outgrows its budget; model routing, tool routing and RAG planning do\nnothing unless enabled. To keep the shared key from activating decisions, pass\nthe text provider's key as credentials.perplexity instead of through the\nenvironment, and keep it out of .env too, because the SDK and the CLI load that\nfile into the environment, or configure one of TypeSafe, Laya and XOR, which then\nreceives those texts instead. The Perplexity provider guide lists\nwhat each consumer sends\nand the\nexact switches.\nIt reads images but no video; see Images and video.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"Perplexity, at its hosted endpoint","lvl3":""}},
3362
+ {"objectID":"86d184ec9e501d4e94e973c833eaee71bb40744e683334ad36c497c3f625387f","title":"Cloudflare Clef, on Workers AI","url":"/docs/features/decide-inference-type#cloudflare-clef-on-workers-ai","content":"Cloudflare's clef and clef-flash are hosted on Workers AI at a public\nendpoint, so there is no base URL to set. The account id is part of the route, so\nClef counts as configured only with both the token and the account id. They can\nequally come from the config passed to the SDK,\nnew NeuroLink({ credentials: { cloudflareClef: { apiKey, accountId } } })\n(baseURL is optional), or per call; config set there counts when a decide()\ncall picks the default decision provider, exactly as the environment does. The\nclassifier router's auto gate is read differently: it looks at the\nconstructor's credentials and the environment only, so a\ncredentials.cloudflareClef passed on one call does not influence it. The\nprovider id is cloudflare-clef: the cloudflare provider is the Workers AI text\nprovider, which serves generate() and stream() and not decide(), and\ncredentials.cloudflare does not configure decide.\n\nThe Workers AI endpoint ignores state text past about 2,048 tokens.\nCloudflare documents a 64K window; text past the observed boundary is ignored\nwithout an error (hosted service or model: unknown). NeuroLink refuses a state it estimates at more than 1,500\ntokens with max_tokens_exceeded, which every built-in consumer treats as \"carry\non as before\"; see Limits and gotchas. Clef suits short\ndecisions.\n\nCLOUDFLARE_API_KEY and CLOUDFLARE_ACCOUNT_ID are shared with that text\nprovider. A host that set them only to use Workers AI text has therefore also\nconfigured decide, and built-in features will use Clef whenever none of\nTypeSafe, Laya, XOR or Perplexity is configured. Clef comes last in the order, so\nit never displaces one that is. No switch turns it off while the two variables\nare set. The Clef provider guide explains\nwhen NeuroLink uses it\nand\nhow the two providers share the token.\nIt reads images but no video; see Images and video.\n\nThe degradation contract.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"Cloudflare Clef, on Workers AI","lvl3":""}},
3362
3363
  {"objectID":"7990ffc08aca0afbeae4310cfdcd3758daf52da4d9edd1da90a82b683f4d7d41","title":"Two transports","url":"/docs/features/decide-inference-type#two-transports","content":"The same model is reachable two ways. Which one runs is decided once, in the\nconstructor:\n\n| | Direct | Vercel AI Gateway |\n| ------------------- | ------------------ | --------------------------------------------- |\n| Key | TYPESAFE_API_KEY | AI_GATEWAY_API_KEY |\n| Endpoint | api.typesafe.ai | ai-gateway.vercel.sh/v4/ai/evaluation-model |\n| Model named in | request body | ai-model-id header |\n| Question vocabulary | noul | boolean |\n| confidence | on each answer | on providerMetadata, not on the answer |\n| Billed by | TypeSafe | Vercel |\n\nHolding both keys keeps the direct transport, so the confidence figures a\nhost already sees do not shift underneath it when a second key appears. Both\ntransports report the vendor's calibrated confidence — the gateway simply puts\nit somewhere else, under providerMetadata.typesafe.confidence.<questionId>,\nleaving the answer objects without one. Set TYPESAFE_TRANSPORT=gateway (or\ncredentials.typesafe.transport) to override.\n\nāš ļø Read that field, not the distribution peak. It is tempting to take\nmax(probabilities) when an answer carries no confidence, and on a\nnear-certain answer the two agree. On an uncertain one they do not, and not by a\nlittle: a measured four-way choice returned probabilities\n{alpha 0.16, beta 0.28, gamma 0.33, delta 0.23} — a peak of 0.33 against a\nreported confidence of 0.10. That gap straddles the default\nminUpgradeConfidence of 0.3, so the derived number clears a bar the real one\nfails and a near-random pick gets acted on as a confident one. The peak stays as\nthe fallback when neither source reports a confidence, and it is genuinely a\ndifferent quantity: an even distribution over N options lands near 1/N, not 0.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"Two transports","lvl3":""}},
3363
- {"objectID":"e5df7db87d4cb31c67fcce769cf5a3573764cb358ed9ee4b41f7ce86cb3954b7","title":"Using it","url":"/docs/features/decide-inference-type#using-it","content":"readDecisionBoolean / readDecisionChoice / readDecisionScore validate at\nruntime and return undefined for a missing id or a mismatched type, so no\ncall site needs a type assertion.\n\nUse decide() instead of tryDecide() when you want the failure to surface;\nit throws a ProviderError whose cause carries a typed kind\n(authentication, rate_limit, max_tokens_exceeded, …).\n\nMore questions than a provider takes. A provider that caps the questions in\none request (Laya takes 64) refuses a longer map from decide() with\nmax_tokens_exceeded. tryDecide(), which every built-in consumer calls,\nsplits the map at the cap instead, keeps the questions in the order given, runs\nup to four batches at a time and joins the answers into one result: usage is\nsummed, the latency is the slowest batch's, and requestId lists the batches'\nids. A batch that fails costs only its own questions, which then have no answer\n(the read* helpers return undefined for them, and every built-in consumer\ncarries on without that decision); tryDecide() returns null only when every\nbatch failed. The state and any media are sent with each batch, so a split\nrequest costs the state's tokens once for each batch. A provider with no cap\n(TypeSafe, XOR) always gets one request.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"Using it","lvl3":""}},
3364
- {"objectID":"7232a490e9e72e01dbe1a12de9470a38ed66ef91aa076b8c8cc745770eb03545","title":"From the CLI","url":"/docs/features/decide-inference-type#from-the-cli","content":"The same primitive is available as neurolink decide [state], which calls\ndecide() (not tryDecide()) and prints one line per answer:\n\nTo ask about an image or a video, add --image <path> (repeatable) or\n--video <path>. XOR reads both, Perplexity reads images only, and TypeSafe and\nLaya read neither:\n\nSee the CLI command reference for the full flag\nlist, including --state-file, --questions-file, --image, --video and\n--format json.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"From the CLI","lvl3":""}},
3364
+ {"objectID":"e5df7db87d4cb31c67fcce769cf5a3573764cb358ed9ee4b41f7ce86cb3954b7","title":"Using it","url":"/docs/features/decide-inference-type#using-it","content":"readDecisionBoolean / readDecisionChoice / readDecisionScore validate at\nruntime and return undefined for a missing id or a mismatched type, so no\ncall site needs a type assertion.\n\nUse decide() instead of tryDecide() when you want the failure to surface;\nit throws a ProviderError whose cause carries a typed kind\n(authentication, rate_limit, max_tokens_exceeded, …).\n\nMore questions than a provider takes. A provider that caps the questions in\none request (Laya and Cloudflare Clef take 64) refuses a longer map from decide() with\nmax_tokens_exceeded. tryDecide(), which every built-in consumer calls,\nsplits the map at the cap instead, keeps the questions in the order given, runs\nup to four batches at a time and joins the answers into one result: usage is\nsummed, the latency is the slowest batch's, and requestId lists the batches'\nids. A batch that fails costs only its own questions, which then have no answer\n(the read* helpers return undefined for them, and every built-in consumer\ncarries on without that decision); tryDecide() returns null only when every\nbatch failed. The state and any media are sent with each batch, so a split\nrequest costs the state's tokens once for each batch. A provider with no cap\n(TypeSafe, XOR) always gets one request.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"Using it","lvl3":""}},
3365
+ {"objectID":"7232a490e9e72e01dbe1a12de9470a38ed66ef91aa076b8c8cc745770eb03545","title":"From the CLI","url":"/docs/features/decide-inference-type#from-the-cli","content":"The same primitive is available as neurolink decide [state], which calls\ndecide() (not tryDecide()) and prints one line per answer:\n\nTo ask about an image or a video, add --image <path> (repeatable) or\n--video <path>. XOR reads both, Perplexity and Cloudflare Clef read images only,\nand TypeSafe and Laya read neither:\n\nSee the CLI command reference for the full flag\nlist, including --state-file, --questions-file, --image, --video and\n--format json.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"From the CLI","lvl3":""}},
3365
3366
  {"objectID":"3b801c727cb399feb7390c3ad67aba9ea5f77581096bb852b7aafb87b70998aa","title":"A choice answer is also a ranking","url":"/docs/features/decide-inference-type#a-choice-answer-is-also-a-ranking","content":"readDecisionChoice returns ranked — every option sorted by probability,\nhighest first. One choice question over N options therefore ranks all N in a\nsingle request. This is the basis for picking from a large catalogue.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"A choice answer is also a ranking","lvl3":""}},
3366
- {"objectID":"e67682840a5ce9650dba2a1d437b0dfbba456cae60fa3f607127e209bf1ec3e7","title":"Images and video","url":"/docs/features/decide-inference-type#images-and-video","content":"A decision provider whose descriptor declares media limits reads images, and\npossibly one video, alongside state. XOR does, and so does Perplexity, which\ndeclares images only and no video; TypeSafe and Laya do not. Two optional fields\non the request carry them, for decide() and tryDecide() alike:\nimages — up to 8 images, the limit XOR and Perplexity each declare.\nvideo — one video, for a provider that declares video (XOR).\n\nEach image and the video takes the same three input forms:\na Buffer;\na local file path;\na data:image/…;base64, or data:video/…;base64, URL.\n\nNeuroLink identifies a Buffer or a file from its bytes, not its extension — PNG,\nJPEG, WebP and GIF images, and MP4, MOV and WebM video —\nand sends it as a data: URL. A data: URL you pass yourself must hold base64\nimage or video content, and is sent as given. Perplexity reads PNG, JPEG and\nWebP only, and refuses a GIF before any request.\n\nWhat is refused. Everything NeuroLink can check is refused before any\nrequest, as a non-retryable invalid_request:\nan http(s) URL — media is not fetched for you, so pass a Buffer, a file path\n or a data: URL;\na string that is not a file path or a data: URL, such as bare base64 or a\n URL with another scheme;\na missing file, a directory, or an empty Buffer or file;\nsomething that is not an image or a video;\nmore than 8 images;\na request body over the provider's limit: 8 MB for XOR, 32 MiB for Perplexity.\n The limit applies to the encoded body, so the base64 form counts. A file over\n it is refused from its size, before it is read;\na video sent to a provider that declares no video (Perplexity);\nan image Perplexity cannot read: one of another type, or over 2,048 tiles of\n 32 Ɨ 32 pixels, which the API would otherwise hold for about a minute before\n answering 504;\nany media sent to a provider that declares no media capability (TypeSafe,\n Laya). The message names the providers that accept media.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"Images and video","lvl3":""}},
3367
- {"objectID":"a6077cedd1ca3e03b757da9a2ef8c185a56d2f1032dd5ed43fe28bd71454836c","title":"The one rule: batch, never fan out","url":"/docs/features/decide-inference-type#the-one-rule-batch-never-fan-out","content":"This inverts the instinct you have from LLMs.\n\nOn TypeSafe, question count barely affects latency (measured against its live\nAPI):\n\n| questions | round trip | input tokens |\n| --------- | ---------- | ------------ |\n| 1 | 393 ms | 310 |\n| 10 | 390 ms | 481 |\n| 100 | 423 ms | 2 281 |\n| 400 | 465 ms | 8 581 |\n\nThese are TypeSafe's figures. perplexity-decider differs: about 0.3 to 1.1 s\nfor one question, about 65 ms for each further question (a request takes at\nmost 128), and an input time that grows faster than linearly with size: 3.9 s\nat 65,000 tokens, 10 s at 146,000 and 22.9 s at 251,000. See\nits latency section.\n\nOn TypeSafe, 400 questions cost ~70 ms more than one. Concurrent requests, by\ncontrast, queue: ten parallel calls take ~1.4 s wall with nine landing together\nat the end, while the server's own upstream time stays flat at 64–169 ms.\n\nSo on TypeSafe, 400 things in one request take ~465 ms; the same 400 as separate\nrequests take roughly a minute. Add every question you might need to the call\nyou are already making — speculative questions are nearly free on TypeSafe (on\nPerplexity each costs about 65 ms, up to 128 in a request), and a second round\ntrip is not.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"The one rule: batch, never fan out","lvl3":""}},
3368
- {"objectID":"0c1ce0b91d624674d6b4cc3b1f67fad986bcdfd379812b89e121784a94408ec3","title":"Model routing","url":"/docs/features/decide-inference-type#model-routing","content":"The classifier router gains a jev\nstrategy, and its default becomes auto — resolving to jev when a decision\nprovider is configured and heuristic when not.\n\nOne request asks for the difficulty tier, whether the task needs\nvision/tools/reasoning, whether carrying it out is risky, and which pool\nmember to use — all at once, in ~400 ms on TypeSafe's Jev.\n\n| | heuristic | llm | jev |\n| ---------------------- | ------------- | ------------------------------- | --------------------------------------------------- |\n| Added latency | 0 ms | ~1–8 s | ~400 ms on TypeSafe |\n| Cost per decision | none | a full LLM call | ~$0.00002 on TypeSafe |\n| Confidence | keyword score | self-reported (defaults to 0.7) | as the provider reports it (calibrated on TypeSafe) |\n| Picks a model directly | no | yes | yes |\n\nThe jev column is TypeSafe's figures; the\nPerplexity guide gives its\nlatency and the confidence it reports.\n\nThresholds are asymmetric, because the two mistakes do not cost the same:\nminUpgradeConfidence defaults to 0.3 (spending more on a wrong guess costs\nmoney) and minDowngradeConfidence to 0.6 (spending less on a wrong guess\nproduces a wrong answer).\n\nThe key upgrades routing; it does not switch routing on. The classifier\nrouter is still opt-in (classifierRouter.enabled) and still needs a pool,\nbecause NeuroLink cannot invent the set of models you are willing to route\nbetween. What the key changes is which classifier runs inside a router you\nalready enabled.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"Model routing","lvl3":""}},
3367
+ {"objectID":"e67682840a5ce9650dba2a1d437b0dfbba456cae60fa3f607127e209bf1ec3e7","title":"Images and video","url":"/docs/features/decide-inference-type#images-and-video","content":"A decision provider whose descriptor declares media limits reads images, and\npossibly one video, alongside state. XOR does, and so do Perplexity and\nCloudflare Clef, which declare images only and no video; TypeSafe and Laya do\nnot. Two optional fields on the request carry them, for decide() and\ntryDecide() alike:\nimages — up to 8 images, the limit XOR and Perplexity each declare; Cloudflare\n Clef takes up to 4.\nvideo — one video, for a provider that declares video (XOR).\n\nEach image and the video takes the same three input forms:\na Buffer;\na local file path;\na data:image/…;base64, or data:video/…;base64, URL.\n\nNeuroLink identifies a Buffer or a file from its bytes, not its extension — PNG,\nJPEG, WebP and GIF images, and MP4, MOV and WebM video —\nand sends it as a data: URL. A data: URL you pass yourself must hold base64\nimage or video content. Clef normalizes image/jpg to image/jpeg; other data URLs are sent as given. Perplexity and Cloudflare Clef read\nPNG, JPEG and WebP only, and refuse a GIF before any request.\n\nWhat is refused. Everything NeuroLink can check is refused before any\nrequest, as a non-retryable invalid_request:\nan http(s) URL — media is not fetched for you, so pass a Buffer, a file path\n or a data: URL;\na string that is not a file path or a data: URL, such as bare base64 or a\n URL with another scheme;\na missing file, a directory, or an empty Buffer or file;\nsomething that is not an image or a video;\nmore than 8 images, or more than 4 for Cloudflare Clef;\na request body over the provider's limit: 8 MB for XOR, 32 MiB for Perplexity,\n 256,000 bytes for Cloudflare Clef. The limit applies to the encoded body, so\n the base64 form counts, and a photo usually has to be resized to well under\n 190 KB before Clef will take it.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"Images and video","lvl3":""}},
3368
+ {"objectID":"a6077cedd1ca3e03b757da9a2ef8c185a56d2f1032dd5ed43fe28bd71454836c","title":"The one rule: batch, never fan out","url":"/docs/features/decide-inference-type#the-one-rule-batch-never-fan-out","content":"This inverts the instinct you have from LLMs.\n\nOn TypeSafe, question count barely affects latency (measured against its live\nAPI):\n\n| questions | round trip | input tokens |\n| --------- | ---------- | ------------ |\n| 1 | 393 ms | 310 |\n| 10 | 390 ms | 481 |\n| 100 | 423 ms | 2 281 |\n| 400 | 465 ms | 8 581 |\n\nThese are TypeSafe's figures. perplexity-decider differs: about 0.3 to 1.1 s\nfor one question, about 65 ms for each further question (a request takes at\nmost 128), and an input time that grows faster than linearly with size: 3.9 s\nat 65,000 tokens, 10 s at 146,000 and 22.9 s at 251,000. See\nits latency section.\ncloudflare-clef takes 0.3 to 1.0 s for a small request and 1.1 to 1.3 s for\n64 questions (measured 2026-10-03). A 64-question request with the same\nquestions and a shorter state took 1.5 s on clef-flash and 2.3 s on clef on 2026-10-04; see\nits latency section.\n\nOn TypeSafe, 400 questions cost ~70 ms more than one. Concurrent requests, by\ncontrast, queue: ten parallel calls take ~1.4 s wall with nine landing together\nat the end, while the server's own upstream time stays flat at 64–169 ms.\n\nSo on TypeSafe, 400 things in one request take ~465 ms; the same 400 as separate\nrequests take roughly a minute. Add every question you might need to the call\nyou are already making — speculative questions are nearly free on TypeSafe (on\nPerplexity each costs about 65 ms, up to 128 in a request; Cloudflare Clef takes\nup to 64), and a second round trip is not.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"The one rule: batch, never fan out","lvl3":""}},
3369
+ {"objectID":"0c1ce0b91d624674d6b4cc3b1f67fad986bcdfd379812b89e121784a94408ec3","title":"Model routing","url":"/docs/features/decide-inference-type#model-routing","content":"The classifier router gains a jev\nstrategy, and its default becomes auto — resolving to jev when a decision\nprovider is configured and heuristic when not.\n\nOne request asks for the difficulty tier, whether the task needs\nvision/tools/reasoning, whether carrying it out is risky, and which pool\nmember to use — all at once, in ~400 ms on TypeSafe's Jev.\n\n| | heuristic | llm | jev |\n| ---------------------- | ------------- | ------------------------------- | --------------------------------------------------- |\n| Added latency | 0 ms | ~1–8 s | ~400 ms on TypeSafe |\n| Cost per decision | none | a full LLM call | ~$0.00002 on TypeSafe |\n| Confidence | keyword score | self-reported (defaults to 0.7) | as the provider reports it (calibrated on TypeSafe) |\n| Picks a model directly | no | yes | yes |\n\nThe jev column is TypeSafe's figures; the\nPerplexity guide gives its\nlatency and the confidence it reports. The\nCloudflare Clef guide gives its\nlatency; because Clef reads only about 2,048 tokens of state, a long prompt is\nrefused there and routing falls back to the heuristic, as it does on any failure.\n\nThresholds are asymmetric, because the two mistakes do not cost the same:\nminUpgradeConfidence defaults to 0.3 (spending more on a wrong guess costs\nmoney) and minDowngradeConfidence to 0.6 (spending less on a wrong guess\nproduces a wrong answer).\n\nThe key upgrades routing; it does not switch routing on. The classifier\nrouter is still opt-in (classifierRouter.enabled) and still needs a pool,\nbecause NeuroLink cannot invent the set of models you are willing to route\nbetween.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"Model routing","lvl3":""}},
3369
3370
  {"objectID":"40f783ceca290c039cce7f0199ab5758d7a0719b0a91c335c1c6e490a72b1412","title":"The model catalogue","url":"/docs/features/decide-inference-type#the-model-catalogue","content":"Enabling classifierRouter.catalog widens the routable pool beyond what you\ndeclared by hand: candidates are built from the 64-model registry (7\nproviders — see the model catalogue\nfor which), intersected with the credentials this host actually holds, and\nranked by a deterministic formula whenever no decision provider is available\nto choose among them.\n\n| | Without catalog | With catalog.enabled |\n| -------------- | ------------------------ | ------------------------------------- |\n| Candidate pool | only the declared pool | declared pool plus registry matches |\n| Fallback pick | first pool member | tierScore()-ranked, tier-aware |\n| Cap | none needed | maxModels, default 120 |\n\nSee the model catalogue for how\ncandidates are filtered, rendered, and ranked.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"The model catalogue","lvl3":""}},
3370
3371
  {"objectID":"1362463e9ab1e8a456cb34a4e5ac2ec8493f18ddc2a8d277282de858c0420131","title":"Per-request context budget","url":"/docs/features/decide-inference-type#per-request-context-budget","content":"A per-request compactionThreshold option lowers the point at which history\ngets compacted, below the 0.8-of-window default. The jev strategy can fill\nit in automatically from a four-level scope rubric (current-message through\neverything) — and the mapping is a one-directional invariant: it can only\never lower the 0.8 default, never raise it, because over-filling a window is\nan unrecoverable provider error.\n\nSee per-request context budget for the\nrubric, the invariant, and how the threshold scales the compaction target.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"Per-request context budget","lvl3":""}},
3371
- {"objectID":"ac97e63916f0ffdf1d9befada357186f236753d40ea1f0e97bcab715e54fa7cd","title":"Relevance-driven compaction","url":"/docs/features/decide-inference-type#relevance-driven-compaction","content":"Before the existing positional compaction stages run, an optional Stage 0\nasks, per eligible message, whether the current request still needs it — at\n~400 ms for the whole batch regardless of message count (on TypeSafe;\nperplexity-decider differs, taking about 0.3 to 1.1 s for one question and about\n65 ms for each further one, plus an input time that grows faster than linearly\nwith size).\nOnly plain user/assistant text is eligible, the most recent messages are never\ntouched, and a message is dropped only on a confident \"no.\"\n\n| | Positional stages (1–4) | Stage 0 (relevance) |\n| ----------------- | ---------------------------------- | ------------------------------------------------------------ |\n| Basis for keeping | position (recency) | relevance to the current request |\n| Drop granularity | whole messages / summarized ranges | whole messages |\n| Runs when | always, once over budget | decision provider configured, request known, and over budget |\n\nSee relevance-driven compaction for\nthe eligibility rules, the drop cap, and the separate summary-quality gate on\nStage 3.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"Relevance-driven compaction","lvl3":""}},
3372
+ {"objectID":"ac97e63916f0ffdf1d9befada357186f236753d40ea1f0e97bcab715e54fa7cd","title":"Relevance-driven compaction","url":"/docs/features/decide-inference-type#relevance-driven-compaction","content":"Before the existing positional compaction stages run, an optional Stage 0\nasks, per eligible message, whether the current request still needs it — at\n~400 ms for the whole batch regardless of message count (on TypeSafe;\nperplexity-decider differs, taking about 0.3 to 1.1 s for one question and about\n65 ms for each further one, plus an input time that grows faster than linearly\nwith size; cloudflare-clef took 0.3 to 1.0 s for a small request on 2026-10-03, and its\nendpoint ignores state text past about 2,048 tokens).\nOnly plain user/assistant text is eligible, the most recent messages are never\ntouched, and a message is dropped only on a confident \"no.\"\n\n| | Positional stages (1–4) | Stage 0 (relevance) |\n| ----------------- | ---------------------------------- | ------------------------------------------------------------ |\n| Basis for keeping | position (recency) | relevance to the current request |\n| Drop granularity | whole messages / summarized ranges | whole messages |\n| Runs when | always, once over budget | decision provider configured, request known, and over budget |\n\nSee relevance-driven compaction for\nthe eligibility rules, the drop cap, and the separate summary-quality gate on\nStage 3.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"Relevance-driven compaction","lvl3":""}},
3372
3373
  {"objectID":"451d199b91f25696bed61dedb5ebc0465143dbf3bc350ca8a3943e622f9aedf8","title":"Tool / MCP routing","url":"/docs/features/decide-inference-type#tool-mcp-routing","content":"The shipped tool router asks a generative model for {servers: string[]} on\na 15-second budget — a shape that cannot express uncertainty. A decision\nmodel instead asks one yes/no question per server, answered with a probability,\nand a server is excluded only on a confident \"no\" (minDropConfidence default\n0.6), because dropping a needed server breaks the turn while keeping an unneeded\none only costs a few tokens.\n\nA measured wording change moved unrelated servers from a mean probability of\n0.31 (dropping 12 of 39 unneeded servers) to a mean of 0.03 (dropping 37 of\n39, with zero wrong drops) — seen in\ntool / MCP routing by decision model,\nwhich also covers the exact question shape and its size guards.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"Tool / MCP routing","lvl3":""}},
3373
3374
  {"objectID":"b35e1ab2390a7b132037c293ecaf436f9ff9498591ccdd47b58ee0a6a95e7616","title":"RAG retrieval planning","url":"/docs/features/decide-inference-type#rag-retrieval-planning","content":"RAGPipelineConfig.decide lets each RAG query get its own topK/hybrid/\ngraph/rerank plan instead of one fixed configuration for every query. An\nexplicit per-call QueryOptions field always wins over the plan, and a\ncapability the pipeline wasn't configured with can never be switched on by\nit.\n\nSee per-query RAG retrieval planning\nfor the breadth rubric and why its confidence bar is deliberately lower than\ntool routing's or compaction's.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"RAG retrieval planning","lvl3":""}},
3374
3375
  {"objectID":"6c0dd418b9f61219911960f90d1003a6c00c2e2cb81399ced3c885fd8209bc6e","title":"Limits and gotchas","url":"/docs/features/decide-inference-type#limits-and-gotchas","content":"Laya reads far less than Jev. About 768 tokens of state on\ntyped-decisions and multilingual, and 320 on english, auto or an\nunrecognised model name, against Jev's ~33,000. Laya's server does not refuse a\nlonger state; it answers from the first 1,024 tokens (512 on english). So the\nprovider estimates the state's size — counting non-Latin characters at a rate\nmeasured per checkpoint — and refuses anything over the limit locally with\nmax_tokens_exceeded, before any network call. The estimate errs toward\nrefusing.\n\nXOR shares one prefill between the state, the questions and any media.\nNeuroLink allows about 200,000 estimated tokens of state (about four characters\nper token for ASCII, one token per character for other scripts) and refuses more\nlocally with max_tokens_exceeded. That figure is a conservative default under\nthe deployment's 250,000-token prefill, which the questions and any images or\nvideo also draw on; it has not been measured against a live deployment. A\nrequest that passes the local check can therefore still be refused by the\nserver as too long. That arrives as max_tokens_exceeded (not retried) when\nthe status is 413 or the message says the context was too long, and otherwise as\nserver, retried once. XOR has no question cap in NeuroLink, and its server takes 2 to 255\noptions on a choice or score.\n\nPerplexity's input ceiling covers more than the state. Measured on a real\naccount in October 2026: the server reads under 262,144 input tokens per request,\ncounted over the state, the questions and the images, and refuses input of\n262,144 tokens or more with an explicit 400 instead of cutting the state off.\nImages are billed as input tokens, about one for each 32 Ɨ 32 tile plus 95 to\n103 more in the single-image requests measured, which looks like fixed request\noverhead and not a charge per image. They count toward the ceiling at one token\nfor each tile: eight 2,048-tile images added 16,384 to a request that was then\nrefused.","hierarchy":{"lvl0":"Features","lvl1":"The `decide` inference type","lvl2":"Limits and gotchas","lvl3":""}},
@@ -3570,9 +3571,9 @@
3570
3571
  {"objectID":"a6f80762fe08ca735f864be4541b6863065a0acfc5849344db2e40f38db28db3","title":"Verify file","url":"/docs/features/image-generation#verify-file","content":"file ./test-output/square.png","hierarchy":{"lvl0":"Features","lvl1":"Image Generation Streaming Guide","lvl2":"Verify file","lvl3":""}},
3571
3572
  {"objectID":"0de2f95f5125f67748b214f7ed674edcf44078cc415fa081f6edf46428d7c65f","title":"Output: ./test-output/square.png: PNG image data, 1024 x 1024, 8-bit/color RGB","url":"/docs/features/image-generation#output-test-outputsquarepng-png-image-data-1024-x-1024-8-bitcolor-rgb","content":"`","hierarchy":{"lvl0":"Features","lvl1":"Image Generation Streaming Guide","lvl2":"Output: ./test-output/square.png: PNG image data, 1024 x 1024, 8-bit/color RGB","lvl3":""}},
3572
3573
  {"objectID":"aa931e13d4db435aa9de30f72f599057c2977804e4316c3cb61f84d8e43fd97e","title":"Conclusion","url":"/docs/features/image-generation#conclusion","content":"NeuroLink's image generation streaming provides a unified interface for both text and image generation. The fake streaming approach ensures consistency while maintaining the benefits of streaming APIs. By following the patterns and examples in this guide, you can effectively integrate image generation into your applications.\n\nFor more information:\nAPI Reference\nProvider Comparison\nProvider Status Monitoring","hierarchy":{"lvl0":"Features","lvl1":"Image Generation Streaming Guide","lvl2":"Conclusion","lvl3":""}},
3573
- {"objectID":"e3366024d1a12d0c73fbcb1592927d2dcc0d92c12eaea661cf5164bb7217c32f","title":"Feature Guides","url":"/docs/features","content":"Feature Guides\n\nComprehensive guides for all NeuroLink features organized by category. Each guide includes setup, usage patterns, configuration, and troubleshooting.\n\nLatest Features (Q1 2026)\n\n| Feature | Description |\n| ---------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |\n| decide Inference Type | A third inference type alongside generate/stream: typed, calibrated boolean/choice/score judgments in one parallel pass — no text. ~400ms, ~$0.00002 per decision on TypeSafe; Perplexity adds about 65 ms for each further question, up to 128 (see its guide). Providers: TypeSafe Jev; Laya, open-weights and self-hosted; XOR, Juspay's open-weights model that also reads images and video; and Perplexity, a hosted API that also reads images. Fail-open: a no-op when no decision provider is configured (Perplexity's key is shared with its text provider). |\n| Model Routing with a Decision Model | One round trip answers difficulty, capabilities, risk, context scope and the model pick, with asymmetric confidence bars (0.3 up / 0.6 down). |\n| Model Catalogue | Ranks the 64-model registry as an addition to a host-declared pool, never a replacement. One choice question ranks all N candidates. |\n| Context Budget | A per-request compaction threshold derived from how much context the request actually needs. Only ever lowers the default, never raises it. ","hierarchy":{"lvl0":"Features","lvl1":"Feature Guides","lvl2":"","lvl3":""}},
3574
+ {"objectID":"e3366024d1a12d0c73fbcb1592927d2dcc0d92c12eaea661cf5164bb7217c32f","title":"Feature Guides","url":"/docs/features","content":"Feature Guides\n\nComprehensive guides for all NeuroLink features organized by category. Each guide includes setup, usage patterns, configuration, and troubleshooting.\n\nLatest Features (Q1 2026)\n\n| Feature | Description |\n| ---------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |\n| decide Inference Type | A third inference type alongside generate/stream: typed, calibrated boolean/choice/score judgments in one parallel pass — no text. ~400ms, ~$0.00002 per decision on TypeSafe; Perplexity adds about 65 ms for each further question, up to 128 (see its guide); Cloudflare Clef takes up to 64 questions and about 2,048 tokens of state (see its guide). Providers: TypeSafe Jev; Laya, open-weights and self-hosted; XOR, Juspay's open-weights model that also reads images and video; Perplexity, a hosted API that also reads images; and Cloudflare Clef, hosted on Workers AI, which also reads images. Fail-open: a no-op when no decision provider is configured (the credentials of Perplexity and Cloudflare Clef are shared with their text providers). |\n| Model Routing with a Decision Model | One round trip answers difficulty, capabilities, risk, context scope and the model pick, with asymmetric confidence bars (0.3 up / 0.6 down). |\n| Model Catalogue | Ranks the 64-model registry as an addition to a host-declared pool, never a replacement. One choice question ranks all N candidates. ","hierarchy":{"lvl0":"Features","lvl1":"Feature Guides","lvl2":"","lvl3":""}},
3574
3575
  {"objectID":"a477c76a577a17e7aeaf89736289a0a41734b4bcfe479a523cdd9e2051bd0ba7","title":"Feature Guides","url":"/docs/features#feature-guides","content":"Comprehensive guides for all NeuroLink features organized by category. Each guide includes setup, usage patterns, configuration, and troubleshooting.","hierarchy":{"lvl0":"Features","lvl1":"Feature Guides","lvl2":"Feature Guides","lvl3":""}},
3575
- {"objectID":"a43e81fa4a6c594f8b6b4cfc216db1356d75b3e4ebb54279dd2da7f69001d28f","title":"Latest Features (Q1 2026)","url":"/docs/features#latest-features-q1-2026","content":"| Feature | Description |\n| ---------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |\n| decide Inference Type | A third inference type alongside generate/stream: typed, calibrated boolean/choice/score judgments in one parallel pass — no text.","hierarchy":{"lvl0":"Features","lvl1":"Feature Guides","lvl2":"Latest Features (Q1 2026)","lvl3":""}},
3576
+ {"objectID":"a43e81fa4a6c594f8b6b4cfc216db1356d75b3e4ebb54279dd2da7f69001d28f","title":"Latest Features (Q1 2026)","url":"/docs/features#latest-features-q1-2026","content":"| Feature | Description |\n| ---------------------------------------------------------------------------------------- |","hierarchy":{"lvl0":"Features","lvl1":"Feature Guides","lvl2":"Latest Features (Q1 2026)","lvl3":""}},
3576
3577
  {"objectID":"84c604ce940f85fbe8aa8d94675e4c1dae80b22ede558d58102426598f6a7d06","title":"Core Features (shipped 2025)","url":"/docs/features#core-features-shipped-2025","content":"| Feature | Description |\n| -------------------------------------------------------- | ----------------------------------------------------------------------------------------------------- |\n| Image Generation | Generate images from text prompts using Gemini models via Vertex AI or Google AI Studio. |\n| Enterprise HITL | Production-ready HITL with approval workflows, confidence thresholds, and enterprise patterns. |\n| Interactive CLI | AI development environment with loop mode, session variables, and conversation memory. |\n| MCP Tools Showcase | Complete guide to 6 built-in tools and connecting external MCP servers across 6 categories. |\n| Human-in-the-Loop (HITL) | Pause AI tool execution for user approval before risky operations like file deletion or API calls. |\n| Guardrails Middleware | Content filtering, PII detection, and safety checks for AI outputs with zero configuration. |\n| Redis Conversation Export | Export complete session history as JSON for analytics, debugging, and compliance auditing. |\n| Context Compaction | Automatic conversation compression for long-running sessions to stay within token limits. |\n| LiteLLM Integration | Access 100+ AI models across a broad range of AI providers through unified LiteLLM routing interface. |\n| SageMaker Integration | Deploy and use custom-trained models on AWS SageMaker infrastructure with full control. |","hierarchy":{"lvl0":"Features","lvl1":"Feature Guides","lvl2":"Core Features (shipped 2025)","lvl3":""}},
3577
3578
  {"objectID":"f64aba3e564a8f56064839c33b533971468da990a9cc95d86ceb1e9d13223478","title":"Earlier Core Features (shipped Q3 2025)","url":"/docs/features#earlier-core-features-shipped-q3-2025","content":"| Feature | Description |\n| ------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------- |\n| Multimodal Chat Experiences | Stream text and images together with automatic provider fallbacks and format conversion. |\n| CSV File Support | Process CSV files for data analysis with automatic format conversion. Works across supported providers. |\n| PDF File Support | Process PDF documents for visual analysis and content extraction. Native provider support. |\n| Office Documents | Process DOCX, PPTX, XLSX files for document analysis. Native Bedrock, Vertex, Anthropic support. |\n| Auto Evaluation Engine | Automated quality scoring and metrics export for AI response validation using LLM-as-judge. |\n| CLI Loop Sessions | Persistent interactive mode with conversation memory and session state for prompt engineering. |\n| Regional Streaming Controls | Region-specific model deployment and routing for compliance and latency optimization. |\n| Provider Orchestration Brain | Adaptive provider and model selection with intelligent fallbacks based on task classification. |","hierarchy":{"lvl0":"Features","lvl1":"Feature Guides","lvl2":"Earlier Core Features (shipped Q3 2025)","lvl3":""}},
3578
3579
  {"objectID":"fbd682fc148477639d38ab175c38363d003a1464834e84b2a48d88087d001b91","title":"Platform Capabilities at a Glance","url":"/docs/features#platform-capabilities-at-a-glance","content":"| Category | Features | Documentation |\n| ------------------------ | ------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------- |\n| Provider unification | Provider integrations with automatic failover, cost-aware routing, providerFallback policy, modelChain config | Provider Setup |\n| Multimodal pipeline | Stream images + CSV data + PDF documents + Office files across providers with auto-detection for mixed file types. | Multimodal Guide, CSV Support, PDF Support, Office Docs |\n| Voice pipeline | TTS (6 providers) + STT (4 providers) + realtime APIs (OpenAI Realtime, Gemini Live) | TTS Guide, STT Guide, Realtime Services |\n| Quality & governance | Auto-evaluation engine (14 scorers), guardrails middleware, HITL workflows, audit logging | Auto Evaluation, Guardrails, HITL |\n| Memory & context | Per-user condensed memory (S3/Redis/SQLite), Redis session export, 5-stage context compaction | Conversation Memory, Memory, Redis Export |\n| CLI tooling | Loop sessions, setup wizard, config validation, Redis auto-detect, JSON output, TTS/STT flags | CLI Loop, CLI Commands |\n| Enterprise ops | Claude proxy, OTLP observability, OpenObserve dashboard, regional routing,","hierarchy":{"lvl0":"Features","lvl1":"Feature Guides","lvl2":"Platform Capabilities at a Glance","lvl3":""}},
@@ -4270,8 +4271,8 @@
4270
4271
  {"objectID":"e4334c89b485facc220fc359dcd520ec9e8464dff64701db51837e14c83a4ec2","title":"What an exposed proxy still reveals","url":"/docs/features/proxy-peer-sharing#what-an-exposed-proxy-still-reveals","content":"On the share listener — or on the main port with\nNEUROLINK_PROXY_REQUIRE_GRANT=1 — /v1/messages and the Codex and\nOpenAI-compatible routes all require a share token. Two endpoints stay open on\npurpose, because a tunnel and a peer both need a liveness probe:\n/health — status, readiness, version, uptime. No account data.\n/status — counters, health and routing state, with account identity\n redacted: labels become account-1, account-2, and the primary-account\n block is blanked. A caller holding the update-control token sees the real\n values.\n\nA loopback allowlist would not have worked here: cloudflared runs on the same\nmachine and connects to 127.0.0.1, so tunnelled traffic arrives from loopback\nexactly like the operator's own CLI does. Separating the two by listener is\nwhat makes the distinction real.","hierarchy":{"lvl0":"Features","lvl1":"Proxy Peer Sharing","lvl2":"What an exposed proxy still reveals","lvl3":""}},
4271
4272
  {"objectID":"752a4365e2038f8539a1c2e53cc1e34d7d1754ae660fec39639e158f92f4e347","title":"Files","url":"/docs/features/proxy-peer-sharing#files","content":"| Path | Owner | Contents |\n| -------------------------------------------- | -------- | ------------------------------------------------------------ |\n| ~/.neurolink/proxy-grants.json | lender | Grants, hashed tokens, policy, state, this node's public URL |\n| ~/.neurolink/proxy-share-ledger.json | lender | Coin spend and per-window buckets |\n| ~/.neurolink/proxy-share-audit.json | lender | Drift observations, streak, auto-pause marker |\n| ~/.neurolink/proxy-share-provisioning.json | lender | Outstanding split-PKCE challenges and single-use codes |\n| ~/.neurolink/proxy-peers.json | borrower | Peers, tokens, priorities, cooldowns, pending verifier |\n| ~/.neurolink/proxy-resident-grants.json | borrower | Leases governing credentials a lender provisioned here |\n| ~/.neurolink/proxy-share-receipts.json | lender | Signed receipts per grant, and the cumulative netted total |\n| ~/.neurolink/proxy-share-notes.json | issuer | Every coin note minted, and which have been redeemed |\n\nAll eight are 0o600 and written by atomic rename. Field-level detail is in the\nconfig reference.\n\nTokens are stored hashed on the lender's side. The raw token exists once, in\nthe output of share create — which is why share link cannot reprint one and\ntells you to rotate instead.","hierarchy":{"lvl0":"Features","lvl1":"Proxy Peer Sharing","lvl2":"Files","lvl3":""}},
4272
4273
  {"objectID":"c1a53699b19ffabe2eb7e0e20f27164c8900d82e292cf71b22843ae047c74382","title":"Scope","url":"/docs/features/proxy-peer-sharing#scope","content":"The gate covers every inbound proxy route, including the Codex and\nOpenAI-compatible surfaces. Account-level gates (reserve floor, window slice) and\ncoin settlement are implemented for the Anthropic engine; a borrowed request on\nanother engine is admitted or refused by the grant's request-level gates only.","hierarchy":{"lvl0":"Features","lvl1":"Proxy Peer Sharing","lvl2":"Scope","lvl3":""}},
4273
- {"objectID":"f34fbae8a587159df344af5151d04a122347f2fa0e0c0592548b0e9b6ea10529","title":"Per-query RAG retrieval planning","url":"/docs/features/rag-retrieval-planning","content":"Per-query RAG retrieval planning\n\nRAGPipeline.query() has always resolved four knobs — topK, hybrid,\ngraph, rerank — from config, with an optional per-call override on each.\nNothing inspected the query itself: \"what is the refund window?\" (one precise\npassage, worth matching by an exact phrase) and \"how does billing relate to\nentitlements?\" (many passages, relationships across documents) got the same\nplan. RAGPipelineConfig.decide lets a\ndecision model answer all four for\nthe query in hand, in one request (~400ms on TypeSafe; on Perplexity each further\nquestion adds about 65 ms, up to 128, see\nits guide).\n\nThe degradation contract. config.decide is optional and defaults to\nunset. Without it, query() behaves exactly as before — the configured\ndefaults and any explicit QueryOptions are all that determine topK,\nhybrid, graph and rerank. Setting plan: false on a call skips\nplanning for that one call even when decide is configured, without\ntouching anything else.\n\nāš ļø This is opt-in wiring, not automatic. Per-query planning lives on\nRAGPipeline. The rag: { files } shortcut on generate() / stream() does\nnot construct one, so that path keeps its fixed topK/hybrid/graph/\nrerank settings. To get planning you build the pipeline yourself and pass a\ndecide function, as below.\n\nWhat gets asked\n\nOne request always asks breadth — a score question rather than a raw\nnumber, because a decision model places a query on an ordered scale\nreliably and reads a digit string as text, not as a quantity to reason with:\n\n| Level | Criterion | topK multiplier |\n| ----- | ----------------------------------------------------------------------------------- | --------------- |\n| 0 | One specific fact, definition or value. A single passage answers it completely. | 0.5Ɨ |\n| 1 | A handful of related points — a procedure, a short comparison, one topic explained. | 1Ɨ |\n| 2 | Several distinct areas that each need their own supporting passage. | 1.5Ɨ |\n| 3 | A broad survey that needs evidence from across the whole corpus. | 2.5Ɨ |\n\nThe multiplier is applied to the pipeline's own configured defaultTopK and\nclamped to between 1 and 50. Below 0.5 confidence the breadth\nreading is dropped entirely and the configured topK stands untouched.\n\nhybrid, graph and rerank are each a plain boolean — and each is asked\nonly when the pipeline was actually configured with that capability.\nAsking about a knob nobody can act on would cost input tokens for nothing\nand invite the mistake of acting on it anyway:\nhybrid: \"This question contains exact terms that must be matched\n literally — an identifier, error code, file name, version number, API\n name, or a quoted phrase — rather than only a topic to match by meaning.\"\ngraph: \"Answering this requires connecting information that lives in\n separate documents, such as how two things relate, what depends on what,\n or tracing a chain across sources.\"\nrerank: \"This question is specific enough that the ORDER of the\n retrieved passages matters — a nearly-right passage would produce a wrong\n answer, so precision is worth an extra ranking pass.\"\n\nA deliberately lower bar than tool routing or compaction\n\nhybrid, graph and rerank are each read with the plain library default —\n0.5 probability, 0.4 confidence — not the stricter 0.6 confidence override\nthat tool routing and\nrelevance compaction both apply. This\nis a deliberate asymmetry, not an oversight: a wrong guess here is cheap (an\nextra ranking pass that didn't help, or a missed lexical match on an\notherwise-fine semantic result), where a wrong guess on a dropped tool\nserver or a dropped conversation message breaks the turn outright. The bar\nmatches the cost of being wrong.\n\nPrecedence: explicit always wins, capability is a hard ceiling\n\nAn explicit QueryOptions field always wins over the plan, per field —\nsetting hybrid: true on one call while letting topK be planned works\nexactly as written. And the plan can never turn on a capability the pipeline\nitself was not configured with: canHybrid/canGraph/canRerank gate\nwhether the question is even asked, so graph: true cannot appear in a plan\nfor a pipeline with no graph index. This is the same \"suggestion, not an\noverride of capability\" contract every other consumer of decide in this\ncodebase follows.\n\nWhat this is bad at\nBreadth is a rubric, not a real answer-length estimate. A level-3\n reading multiplies topK by 2.5Ɨ regardless of how large the corpus\n actually is — for a small collection that can mean requesting more\n passages than exist.\nThe three capability booleans don't interact. hybrid and rerank\n are decided independently even though a rerank pass changes how much a\n lexical-match boost from hybrid search actually matters; there's no joint\n reasoning about the combination, only three separate yes/no answers.\nIt only sees the query text. It","hierarchy":{"lvl0":"Features","lvl1":"Per-query RAG retrieval planning","lvl2":"","lvl3":""}},
4274
- {"objectID":"e670771db01b04c3e4226fc7e18c58fd9ce171ece9158b100f285cce608942ea","title":"Per-query RAG retrieval planning","url":"/docs/features/rag-retrieval-planning#per-query-rag-retrieval-planning","content":"RAGPipeline.query() has always resolved four knobs — topK, hybrid,\ngraph, rerank — from config, with an optional per-call override on each.\nNothing inspected the query itself: \"what is the refund window?\" (one precise\npassage, worth matching by an exact phrase) and \"how does billing relate to\nentitlements?\" (many passages, relationships across documents) got the same\nplan. RAGPipelineConfig.decide lets a\ndecision model answer all four for\nthe query in hand, in one request (~400ms on TypeSafe; on Perplexity each further\nquestion adds about 65 ms, up to 128, see\nits guide).\n\nThe degradation contract. config.decide is optional and defaults to\nunset. Without it, query() behaves exactly as before — the configured\ndefaults and any explicit QueryOptions are all that determine topK,\nhybrid, graph and rerank. Setting plan: false on a call skips\nplanning for that one call even when decide is configured, without\ntouching anything else.\n\nāš ļø This is opt-in wiring, not automatic. Per-query planning lives on\nRAGPipeline. The rag: { files } shortcut on generate() / stream() does\nnot construct one, so that path keeps its fixed topK/hybrid/graph/\nrerank settings. To get planning you build the pipeline yourself and pass a\ndecide function, as below.","hierarchy":{"lvl0":"Features","lvl1":"Per-query RAG retrieval planning","lvl2":"Per-query RAG retrieval planning","lvl3":""}},
4274
+ {"objectID":"f34fbae8a587159df344af5151d04a122347f2fa0e0c0592548b0e9b6ea10529","title":"Per-query RAG retrieval planning","url":"/docs/features/rag-retrieval-planning","content":"Per-query RAG retrieval planning\n\nRAGPipeline.query() has always resolved four knobs — topK, hybrid,\ngraph, rerank — from config, with an optional per-call override on each.\nNothing inspected the query itself: \"what is the refund window?\" (one precise\npassage, worth matching by an exact phrase) and \"how does billing relate to\nentitlements?\" (many passages, relationships across documents) got the same\nplan. RAGPipelineConfig.decide lets a\ndecision model answer all four for\nthe query in hand, in one request (~400ms on TypeSafe; on Perplexity each further\nquestion adds about 65 ms, up to 128, see\nits guide; on\nCloudflare Clef a small request took 0.3 to 1.0 s on 2026-10-03, see\nits guide).\n\nThe degradation contract. config.decide is optional and defaults to\nunset. Without it, query() behaves exactly as before — the configured\ndefaults and any explicit QueryOptions are all that determine topK,\nhybrid, graph and rerank. Setting plan: false on a call skips\nplanning for that one call even when decide is configured, without\ntouching anything else.\n\nāš ļø This is opt-in wiring, not automatic. Per-query planning lives on\nRAGPipeline. The rag: { files } shortcut on generate() / stream() does\nnot construct one, so that path keeps its fixed topK/hybrid/graph/\nrerank settings. To get planning you build the pipeline yourself and pass a\ndecide function, as below.\n\nWhat gets asked\n\nOne request always asks breadth — a score question rather than a raw\nnumber, because a decision model places a query on an ordered scale\nreliably and reads a digit string as text, not as a quantity to reason with:\n\n| Level | Criterion | topK multiplier |\n| ----- | ----------------------------------------------------------------------------------- | --------------- |\n| 0 | One specific fact, definition or value. A single passage answers it completely. | 0.5Ɨ |\n| 1 | A handful of related points — a procedure, a short comparison, one topic explained. | 1Ɨ |\n| 2 | Several distinct areas that each need their own supporting passage. | 1.5Ɨ |\n| 3 | A broad survey that needs evidence from across the whole corpus. | 2.5Ɨ |\n\nThe multiplier is applied to the pipeline's own configured defaultTopK and\nclamped to between 1 and 50. Below 0.5 confidence the breadth\nreading is dropped entirely and the configured topK stands untouched.\n\nhybrid, graph and rerank are each a plain boolean — and each is asked\nonly when the pipeline was actually configured with that capability.\nAsking about a knob nobody can act on would cost input tokens for nothing\nand invite the mistake of acting on it anyway:\nhybrid: \"This question contains exact terms that must be matched\n literally — an identifier, error code, file name, version number, API\n name, or a quoted phrase — rather than only a topic to match by meaning.\"\ngraph: \"Answering this requires connecting information that lives in\n separate documents, such as how two things relate, what depends on what,\n or tracing a chain across sources.\"\nrerank: \"This question is specific enough that the ORDER of the\n retrieved passages matters — a nearly-right passage would produce a wrong\n answer, so precision is worth an extra ranking pass.\"\n\nA deliberately lower bar than tool routing or compaction\n\nhybrid, graph and rerank are each read with the plain library default —\n0.5 probability, 0.4 confidence — not the stricter 0.6 confidence override\nthat tool routing and\nrelevance compaction both apply. This\nis a deliberate asymmetry, not an oversight: a wrong guess here is cheap (an\nextra ranking pass that didn't help, or a missed lexical match on an\notherwise-fine semantic result), where a wrong guess on a dropped tool\nserver or a dropped conversation message breaks the turn outright. The bar\nmatches the cost of being wrong.\n\nPrecedence: explicit always wins, capability is a hard ceiling\n\nAn explicit QueryOptions field always wins over the plan, per field —\nsetting hybrid: true on one call while letting topK be planned works\nexactly as written. And the plan can never turn on a capability the pipeline\nitself was not configured with: canHybrid/canGraph/canRerank gate\nwhether the question is even asked, so graph: true cannot appear in a plan\nfor a pipeline with no graph index. This is the same \"suggestion, not an\noverride of capability\" contract every other consumer of decide in this\ncodebase follows.\n\nWhat this is bad at\nBreadth is a rubric, not a real answer-length estimate. A level-3\n reading multiplies topK by 2.5Ɨ regardless of how large the corpus\n actually is — for a small collection that can mean requesting more\n passages than exist.\nThe three capability booleans don't interact. hybrid and rerank\n are decided independently even though a rerank pass changes how much a\n lexical-match boost from hybrid search actually matters; there's no joint\n reasoning about t","hierarchy":{"lvl0":"Features","lvl1":"Per-query RAG retrieval planning","lvl2":"","lvl3":""}},
4275
+ {"objectID":"e670771db01b04c3e4226fc7e18c58fd9ce171ece9158b100f285cce608942ea","title":"Per-query RAG retrieval planning","url":"/docs/features/rag-retrieval-planning#per-query-rag-retrieval-planning","content":"RAGPipeline.query() has always resolved four knobs — topK, hybrid,\ngraph, rerank — from config, with an optional per-call override on each.\nNothing inspected the query itself: \"what is the refund window?\" (one precise\npassage, worth matching by an exact phrase) and \"how does billing relate to\nentitlements?\" (many passages, relationships across documents) got the same\nplan. RAGPipelineConfig.decide lets a\ndecision model answer all four for\nthe query in hand, in one request (~400ms on TypeSafe; on Perplexity each further\nquestion adds about 65 ms, up to 128, see\nits guide; on\nCloudflare Clef a small request took 0.3 to 1.0 s on 2026-10-03, see\nits guide).\n\nThe degradation contract. config.decide is optional and defaults to\nunset. Without it, query() behaves exactly as before — the configured\ndefaults and any explicit QueryOptions are all that determine topK,\nhybrid, graph and rerank. Setting plan: false on a call skips\nplanning for that one call even when decide is configured, without\ntouching anything else.\n\nāš ļø This is opt-in wiring, not automatic. Per-query planning lives on\nRAGPipeline. The rag: { files } shortcut on generate() / stream() does\nnot construct one, so that path keeps its fixed topK/hybrid/graph/\nrerank settings. To get planning you build the pipeline yourself and pass a\ndecide function, as below.","hierarchy":{"lvl0":"Features","lvl1":"Per-query RAG retrieval planning","lvl2":"Per-query RAG retrieval planning","lvl3":""}},
4275
4276
  {"objectID":"e6334a89a1b2f75a81df5db11ab4234a96c5ffc4a240bdaee68eeeeef841b927","title":"What gets asked","url":"/docs/features/rag-retrieval-planning#what-gets-asked","content":"One request always asks breadth — a score question rather than a raw\nnumber, because a decision model places a query on an ordered scale\nreliably and reads a digit string as text, not as a quantity to reason with:\n\n| Level | Criterion | topK multiplier |\n| ----- | ----------------------------------------------------------------------------------- | --------------- |\n| 0 | One specific fact, definition or value. A single passage answers it completely. | 0.5Ɨ |\n| 1 | A handful of related points — a procedure, a short comparison, one topic explained. | 1Ɨ |\n| 2 | Several distinct areas that each need their own supporting passage. | 1.5Ɨ |\n| 3 | A broad survey that needs evidence from across the whole corpus. | 2.5Ɨ |\n\nThe multiplier is applied to the pipeline's own configured defaultTopK and\nclamped to between 1 and 50. Below 0.5 confidence the breadth\nreading is dropped entirely and the configured topK stands untouched.\n\nhybrid, graph and rerank are each a plain boolean — and each is asked\nonly when the pipeline was actually configured with that capability.\nAsking about a knob nobody can act on would cost input tokens for nothing\nand invite the mistake of acting on it anyway:\nhybrid: \"This question contains exact terms that must be matched\n literally — an identifier, error code, file name, version number, API\n name, or a quoted phrase — rather than only a topic to match by meaning.\"\ngraph: \"Answering this requires connecting information that lives in\n separate documents, such as how two things relate, what depends on what,\n or tracing a chain across sources.\"\nrerank: \"This question is specific enough that the ORDER of the\n retrieved passages matters — a nearly-right passage would produce a wrong\n answer, so precision is worth an extra ranking pass.\"","hierarchy":{"lvl0":"Features","lvl1":"Per-query RAG retrieval planning","lvl2":"What gets asked","lvl3":""}},
4276
4277
  {"objectID":"88e9accbffc38ab55b79e980a2607cec77993a9dd10d4f869f78cf8692ff9f60","title":"A deliberately lower bar than tool routing or compaction","url":"/docs/features/rag-retrieval-planning#a-deliberately-lower-bar-than-tool-routing-or-compaction","content":"hybrid, graph and rerank are each read with the plain library default —\n0.5 probability, 0.4 confidence — not the stricter 0.6 confidence override\nthat tool routing and\nrelevance compaction both apply. This\nis a deliberate asymmetry, not an oversight: a wrong guess here is cheap (an\nextra ranking pass that didn't help, or a missed lexical match on an\notherwise-fine semantic result), where a wrong guess on a dropped tool\nserver or a dropped conversation message breaks the turn outright. The bar\nmatches the cost of being wrong.","hierarchy":{"lvl0":"Features","lvl1":"Per-query RAG retrieval planning","lvl2":"A deliberately lower bar than tool routing or compaction","lvl3":""}},
4277
4278
  {"objectID":"7472123cec86c242b540be1ec22a4c60e0ba8015f8d7938189f049b648077f4e","title":"Precedence: explicit always wins, capability is a hard ceiling","url":"/docs/features/rag-retrieval-planning#precedence-explicit-always-wins-capability-is-a-hard-ceiling","content":"An explicit QueryOptions field always wins over the plan, per field —\nsetting hybrid: true on one call while letting topK be planned works\nexactly as written. And the plan can never turn on a capability the pipeline\nitself was not configured with: canHybrid/canGraph/canRerank gate\nwhether the question is even asked, so graph: true cannot appear in a plan\nfor a pipeline with no graph index. This is the same \"suggestion, not an\noverride of capability\" contract every other consumer of decide in this\ncodebase follows.","hierarchy":{"lvl0":"Features","lvl1":"Per-query RAG retrieval planning","lvl2":"Precedence: explicit always wins, capability is a hard ceiling","lvl3":""}},
@@ -4506,9 +4507,9 @@
4506
4507
  {"objectID":"5594cc075cbacf6979c75d68f20b45ccff0513a1e4c7087d756491900cbb1c4d","title":"Model Detection Utilities","url":"/docs/features/thinking-configuration#model-detection-utilities","content":"NeuroLink provides utilities to check thinking support:","hierarchy":{"lvl0":"Features","lvl1":"Extended Thinking Configuration","lvl2":"Model Detection Utilities","lvl3":""}},
4507
4508
  {"objectID":"559dd274be4e9bdfeabf680692f3830f92a5c13caefc42acfb819d9635f2dad7","title":"Important Notes","url":"/docs/features/thinking-configuration#important-notes","content":"Provider compatibility: Thinking configuration is provider-specific. Gemini uses thinkingLevel, Claude uses budgetTokens\nToken consumption: Extended thinking uses additional tokens beyond the response\nLatency impact: Higher thinking levels increase response time\nNot all models support thinking: Check supportsThinkingConfig() before enabling\nStreaming support: Thinking configuration works with both generate() and stream()","hierarchy":{"lvl0":"Features","lvl1":"Extended Thinking Configuration","lvl2":"Important Notes","lvl3":""}},
4508
4509
  {"objectID":"4bd3d62cf46ca16e0b9a52768c51ca4bced1f07f6c32f9a141c0cbfc599bcefb","title":"See Also","url":"/docs/features/thinking-configuration#see-also","content":"API Reference\nProvider Configuration\nStreaming","hierarchy":{"lvl0":"Features","lvl1":"Extended Thinking Configuration","lvl2":"See Also","lvl3":""}},
4509
- {"objectID":"ffa6f024a99609da8db45d9a02bc76178d7fac894f00af93b830f9b356e86f3c","title":"Tool / MCP routing by decision model","url":"/docs/features/tool-routing-decision-model","content":"Tool / MCP routing by decision model\n\nThe shipped tool router asks a generative model for {servers: string[]} on\na 15-second budget. That shape cannot express uncertainty — a server is in\nthe list or it is not, and the only recourse for a model that is unsure is to\ninclude it. A decision model instead\nasks one calibrated yes/no question per server, in a single round trip of\nabout 400ms and $0.00002 on TypeSafe (on Perplexity each further question adds\nabout 65 ms, up to 128 in a request; see\nits guide), and each\nanswer comes back with a real probability rather than a name that either did or\ndidn't make a list.\n\nThe degradation contract. selectServersByDecision() is used only when a\ndecision provider is configured; hosts never wire decideFn by hand — it is\nbound automatically wherever tool routing resolves, using the same\ntryDecide() that returns null on any failure or absent configuration. No\ndecision provider, a failed call, fewer than two candidate servers, or an\nanswer set that would drop nothing — any of these fall straight through to\nthe existing generative router, unchanged.\n\nOne question per server\n\nEach routable server gets its own boolean question, built from its\ndescription (or, absent one, its tool names):\n\nA server is excluded only on a confident false — minDropConfidence\ndefaults to 0.6. undefined (unanswered, malformed, or too close to a\ncoin flip) and a confident true both keep the server. This is the same\nasymmetry as the classifier's upgrade/downgrade bars: keeping an unneeded\nserver costs a few hundred tokens of tool definitions; dropping a needed one\nbreaks the turn outright, because the model can never call a tool it was\nnever shown and has no way to ask for it back.\n\nTwo size guards bound the request: at most 200 servers are asked about\nin one batch (MAX_SERVERS), and the query text sent as state is capped at\n10,000 characters. Servers past the cap are never asked about and are\ntherefore always kept — a server that was not offered to the model must\nnever be silently dropped by its own absence from the question set.\n\nThe wording that made this work: a measured A/B result\n\nThe first phrasing tried was the obvious one: \"answering this request will\nrequire calling at least one tool from this server,\" with false meaning\n\"this server is unrelated, OR the request needs no tool at all.\" Measured\nagainst a 10-request Ɨ 5-server labelled set, it separated correctly but\nweakly — unrelated servers averaged p = 0.31 and reached as high as\n0.80, so at the 0.6 drop bar only 12 of 39 unneeded servers were\nactually dropped.\n\nThree changes fixed it: naming the server explicitly, asking in the present\ntense about what carrying out the request involves rather than what\n\"will require,\" and splitting the bundled false criterion (which was\nreally two separate claims joined by \"or\") into one single claim. That\nmoved unrelated servers to a mean of p = 0.03 with a maximum of 0.35\n— 37 of 39 dropped at the same 0.6 bar, still with zero wrong drops.\n\nThe lesson generalises past this one question: a decision model reads\nliterally, and an \"or\" in a criterion is two questions wearing one coat. Each\nhalf of a compound criterion pulls the answer toward the middle whenever\nonly one half is true, which is exactly the muddy, hard-to-gate signal the\nfirst version produced.\n\nWhat this is bad at\nIt reasons about servers, not individual tools. The unit of decision is\n a whole MCP server; a server with twenty tools where the request needs one\n is kept or dropped as a unit, not tool-by-tool.\nThe description quality bounds the question quality. A server with no\n declared description falls back to a comma-joined list of its own tool\n names, which carries much less signal than a well-written one-line\n description — the wording fix above only helps once the server's own text\n is legible to a literal reader.\nA close call still resolves to \"keep.\" There is no partial exclusion;\n anything from a coin flip up to just under 0.6 confidence is treated\n identically to a confident true.\nIt shares the base model's general limits — literal reading, no\n arithmetic, accuracy sensitive to a noisy or oversized state — all\n described in\n what decide is bad at.\nThe query is untrusted input sent as state, and this module does not\n sanitize it. The blast radius is deliberately bounded instead: server ids\n are never read back off the wire (answers are matched by position, not by\n name), so the worst a crafted query can do is keep more already-registered\n servers than necessary — it cannot register a server that wasn't already\n configured.\n\nSee also\nThe decide inference type\nModel routing with a decision model\nRelevance-driven compaction","hierarchy":{"lvl0":"Features","lvl1":"Tool / MCP routing by decision model","lvl2":"","lvl3":""}},
4510
- {"objectID":"4b84d2a2ea7a8b34624599647a02792ace2fc30650e5bb5d3da1ea68e46fdad8","title":"Tool / MCP routing by decision model","url":"/docs/features/tool-routing-decision-model#tool-mcp-routing-by-decision-model","content":"The shipped tool router asks a generative model for {servers: string[]} on\na 15-second budget. That shape cannot express uncertainty — a server is in\nthe list or it is not, and the only recourse for a model that is unsure is to\ninclude it. A decision model instead\nasks one calibrated yes/no question per server, in a single round trip of\nabout 400ms and $0.00002 on TypeSafe (on Perplexity each further question adds\nabout 65 ms, up to 128 in a request; see\nits guide), and each\nanswer comes back with a real probability rather than a name that either did or\ndidn't make a list.\n\nThe degradation contract. selectServersByDecision() is used only when a\ndecision provider is configured; hosts never wire decideFn by hand — it is\nbound automatically wherever tool routing resolves, using the same\ntryDecide() that returns null on any failure or absent configuration. No\ndecision provider, a failed call, fewer than two candidate servers, or an\nanswer set that would drop nothing — any of these fall straight through to\nthe existing generative router, unchanged.","hierarchy":{"lvl0":"Features","lvl1":"Tool / MCP routing by decision model","lvl2":"Tool / MCP routing by decision model","lvl3":""}},
4511
- {"objectID":"0b9203bc3ddca1d61b578f4bea25de508f6a241da96f270ca4742163bd99884a","title":"One question per server","url":"/docs/features/tool-routing-decision-model#one-question-per-server","content":"Each routable server gets its own boolean question, built from its\ndescription (or, absent one, its tool names):\n\nA server is excluded only on a confident false — minDropConfidence\ndefaults to 0.6. undefined (unanswered, malformed, or too close to a\ncoin flip) and a confident true both keep the server. This is the same\nasymmetry as the classifier's upgrade/downgrade bars: keeping an unneeded\nserver costs a few hundred tokens of tool definitions; dropping a needed one\nbreaks the turn outright, because the model can never call a tool it was\nnever shown and has no way to ask for it back.\n\nTwo size guards bound the request: at most 200 servers are asked about\nin one batch (MAX_SERVERS), and the query text sent as state is capped at\n10,000 characters. Servers past the cap are never asked about and are\ntherefore always kept — a server that was not offered to the model must\nnever be silently dropped by its own absence from the question set.","hierarchy":{"lvl0":"Features","lvl1":"Tool / MCP routing by decision model","lvl2":"One question per server","lvl3":""}},
4510
+ {"objectID":"ffa6f024a99609da8db45d9a02bc76178d7fac894f00af93b830f9b356e86f3c","title":"Tool / MCP routing by decision model","url":"/docs/features/tool-routing-decision-model","content":"Tool / MCP routing by decision model\n\nThe shipped tool router asks a generative model for {servers: string[]} on\na 15-second budget. That shape cannot express uncertainty — a server is in\nthe list or it is not, and the only recourse for a model that is unsure is to\ninclude it. A decision model instead\nasks one calibrated yes/no question per server, in a single round trip of\nabout 400ms and $0.00002 on TypeSafe (on Perplexity each further question adds\nabout 65 ms, up to 128 in a request; see\nits guide; on\nCloudflare Clef a request takes up to 64 questions, and tryDecide() splits a\nlarger set into batches of 64; see\nits guide), and each\nanswer comes back with a real probability rather than a name that either did or\ndidn't make a list.\n\nThe degradation contract. selectServersByDecision() is used only when a\ndecision provider is configured; hosts never wire decideFn by hand — it is\nbound automatically wherever tool routing resolves, using the same\ntryDecide() that returns null on any failure or absent configuration. No\ndecision provider, a failed call, fewer than two candidate servers, or an\nanswer set that would drop nothing — any of these fall straight through to\nthe existing generative router, unchanged.\n\nOne question per server\n\nEach routable server gets its own boolean question, built from its\ndescription (or, absent one, its tool names):\n\nA server is excluded only on a confident false — minDropConfidence\ndefaults to 0.6. undefined (unanswered, malformed, or too close to a\ncoin flip) and a confident true both keep the server. This is the same\nasymmetry as the classifier's upgrade/downgrade bars: keeping an unneeded\nserver costs a few hundred tokens of tool definitions; dropping a needed one\nbreaks the turn outright, because the model can never call a tool it was\nnever shown and has no way to ask for it back.\n\nTwo size guards bound the request: at most 200 servers are asked about\nin one batch (MAX_SERVERS), and the query text sent as state is capped at\n10,000 characters. Servers past the cap are never asked about and are\ntherefore always kept — a server that was not offered to the model must\nnever be silently dropped by its own absence from the question set.\n\nThe Cloudflare Clef endpoint ignores state text past about 2,048 tokens, and\nNeuroLink refuses a state it estimates at more than 1,500 tokens with\nmax_tokens_exceeded. A query long enough to be estimated above that is refused,\nand routing falls through to the generative router, as it does on any failure.\n\nThe wording that made this work: a measured A/B result\n\nThe first phrasing tried was the obvious one: \"answering this request will\nrequire calling at least one tool from this server,\" with false meaning\n\"this server is unrelated, OR the request needs no tool at all.\" Measured\nagainst a 10-request Ɨ 5-server labelled set, it separated correctly but\nweakly — unrelated servers averaged p = 0.31 and reached as high as\n0.80, so at the 0.6 drop bar only 12 of 39 unneeded servers were\nactually dropped.\n\nThree changes fixed it: naming the server explicitly, asking in the present\ntense about what carrying out the request involves rather than what\n\"will require,\" and splitting the bundled false criterion (which was\nreally two separate claims joined by \"or\") into one single claim. That\nmoved unrelated servers to a mean of p = 0.03 with a maximum of 0.35\n— 37 of 39 dropped at the same 0.6 bar, still with zero wrong drops.\n\nThe lesson generalises past this one question: a decision model reads\nliterally, and an \"or\" in a criterion is two questions wearing one coat. Each\nhalf of a compound criterion pulls the answer toward the middle whenever\nonly one half is true, which is exactly the muddy, hard-to-gate signal the\nfirst version produced.\n\nWhat this is bad at\nIt reasons about servers, not individual tools. The unit of decision is\n a whole MCP server; a server with twenty tools where the request needs one\n is kept or dropped as a unit, not tool-by-tool.\nThe description quality bounds the question quality. A server with no\n declared description falls back to a comma-joined list of its own tool\n names, which carries much less signal than a well-written one-line\n description — the wording fix above only helps once the server's own text\n is legible to a literal reader.\nA close call still resolves to \"keep.\" There is no partial exclusion;\n anything from a coin flip up to just under 0.6 confidence is treated\n identically to a confident true.\nIt shares the base model's general limits — literal reading, no\n arithmetic, accuracy sensitive to a noisy or oversized state — all\n described in\n what decide is bad at.\nThe query is untrusted input sent as state, and this module does not\n sanitize it. The blast radius is deliberately bounded instead: server ids\n are never read back off the wire (answers are matched by position, not by\n name), so the worst a crafted query can do is keep more already-registered\n servers than necessary — it cannot register a server that ","hierarchy":{"lvl0":"Features","lvl1":"Tool / MCP routing by decision model","lvl2":"","lvl3":""}},
4511
+ {"objectID":"4b84d2a2ea7a8b34624599647a02792ace2fc30650e5bb5d3da1ea68e46fdad8","title":"Tool / MCP routing by decision model","url":"/docs/features/tool-routing-decision-model#tool-mcp-routing-by-decision-model","content":"The shipped tool router asks a generative model for {servers: string[]} on\na 15-second budget. That shape cannot express uncertainty — a server is in\nthe list or it is not, and the only recourse for a model that is unsure is to\ninclude it. A decision model instead\nasks one calibrated yes/no question per server, in a single round trip of\nabout 400ms and $0.00002 on TypeSafe (on Perplexity each further question adds\nabout 65 ms, up to 128 in a request; see\nits guide; on\nCloudflare Clef a request takes up to 64 questions, and tryDecide() splits a\nlarger set into batches of 64; see\nits guide), and each\nanswer comes back with a real probability rather than a name that either did or\ndidn't make a list.\n\nThe degradation contract. selectServersByDecision() is used only when a\ndecision provider is configured; hosts never wire decideFn by hand — it is\nbound automatically wherever tool routing resolves, using the same\ntryDecide() that returns null on any failure or absent configuration. No\ndecision provider, a failed call, fewer than two candidate servers, or an\nanswer set that would drop nothing — any of these fall straight through to\nthe existing generative router, unchanged.","hierarchy":{"lvl0":"Features","lvl1":"Tool / MCP routing by decision model","lvl2":"Tool / MCP routing by decision model","lvl3":""}},
4512
+ {"objectID":"0b9203bc3ddca1d61b578f4bea25de508f6a241da96f270ca4742163bd99884a","title":"One question per server","url":"/docs/features/tool-routing-decision-model#one-question-per-server","content":"Each routable server gets its own boolean question, built from its\ndescription (or, absent one, its tool names):\n\nA server is excluded only on a confident false — minDropConfidence\ndefaults to 0.6. undefined (unanswered, malformed, or too close to a\ncoin flip) and a confident true both keep the server. This is the same\nasymmetry as the classifier's upgrade/downgrade bars: keeping an unneeded\nserver costs a few hundred tokens of tool definitions; dropping a needed one\nbreaks the turn outright, because the model can never call a tool it was\nnever shown and has no way to ask for it back.\n\nTwo size guards bound the request: at most 200 servers are asked about\nin one batch (MAX_SERVERS), and the query text sent as state is capped at\n10,000 characters. Servers past the cap are never asked about and are\ntherefore always kept — a server that was not offered to the model must\nnever be silently dropped by its own absence from the question set.\n\nThe Cloudflare Clef endpoint ignores state text past about 2,048 tokens, and\nNeuroLink refuses a state it estimates at more than 1,500 tokens with\nmax_tokens_exceeded. A query long enough to be estimated above that is refused,\nand routing falls through to the generative router, as it does on any failure.","hierarchy":{"lvl0":"Features","lvl1":"Tool / MCP routing by decision model","lvl2":"One question per server","lvl3":""}},
4512
4513
  {"objectID":"19352c9bf82d208bbd76e4303e3e35ac7d4f7497562a4f1ee006d8cf1fae5f6b","title":"The wording that made this work: a measured A/B result","url":"/docs/features/tool-routing-decision-model#the-wording-that-made-this-work-a-measured-ab-result","content":"The first phrasing tried was the obvious one: \"answering this request will\nrequire calling at least one tool from this server,\" with false meaning\n\"this server is unrelated, OR the request needs no tool at all.\" Measured\nagainst a 10-request Ɨ 5-server labelled set, it separated correctly but\nweakly — unrelated servers averaged p = 0.31 and reached as high as\n0.80, so at the 0.6 drop bar only 12 of 39 unneeded servers were\nactually dropped.\n\nThree changes fixed it: naming the server explicitly, asking in the present\ntense about what carrying out the request involves rather than what\n\"will require,\" and splitting the bundled false criterion (which was\nreally two separate claims joined by \"or\") into one single claim. That\nmoved unrelated servers to a mean of p = 0.03 with a maximum of 0.35\n— 37 of 39 dropped at the same 0.6 bar, still with zero wrong drops.\n\nThe lesson generalises past this one question: a decision model reads\nliterally, and an \"or\" in a criterion is two questions wearing one coat. Each\nhalf of a compound criterion pulls the answer toward the middle whenever\nonly one half is true, which is exactly the muddy, hard-to-gate signal the\nfirst version produced.","hierarchy":{"lvl0":"Features","lvl1":"Tool / MCP routing by decision model","lvl2":"The wording that made this work: a measured A/B result","lvl3":""}},
4513
4514
  {"objectID":"9c9af1c1ff58051674501cb10c371fdcf50758c935a2dfa168df9719969bb7af","title":"What this is bad at","url":"/docs/features/tool-routing-decision-model#what-this-is-bad-at","content":"It reasons about servers, not individual tools. The unit of decision is\n a whole MCP server; a server with twenty tools where the request needs one\n is kept or dropped as a unit, not tool-by-tool.\nThe description quality bounds the question quality. A server with no\n declared description falls back to a comma-joined list of its own tool\n names, which carries much less signal than a well-written one-line\n description — the wording fix above only helps once the server's own text\n is legible to a literal reader.\nA close call still resolves to \"keep.\" There is no partial exclusion;\n anything from a coin flip up to just under 0.6 confidence is treated\n identically to a confident true.\nIt shares the base model's general limits — literal reading, no\n arithmetic, accuracy sensitive to a noisy or oversized state — all\n described in\n what decide is bad at.\nThe query is untrusted input sent as state, and this module does not\n sanitize it. The blast radius is deliberately bounded instead: server ids\n are never read back off the wire (answers are matched by position, not by\n name), so the worst a crafted query can do is keep more already-registered\n servers than necessary — it cannot register a server that wasn't already\n configured.","hierarchy":{"lvl0":"Features","lvl1":"Tool / MCP routing by decision model","lvl2":"What this is bad at","lvl3":""}},
4514
4515
  {"objectID":"b55a36dccb54a184ce24fa850942a8f5455cbe7f4391a23fc94f326aebeae19e","title":"See also","url":"/docs/features/tool-routing-decision-model#see-also","content":"The decide inference type\nModel routing with a decision model\nRelevance-driven compaction","hierarchy":{"lvl0":"Features","lvl1":"Tool / MCP routing by decision model","lvl2":"See also","lvl3":""}},
@@ -5601,7 +5602,26 @@
5601
5602
  {"objectID":"5862cb134695d7821bb2656e8d0c7c1ffee1ff3f295c36858f60fa939ba2facd","title":"Verification status","url":"/docs/getting-started/providers/chutes#verification-status","content":"Tier-2 onboarding normally requires a live-key capability sweep\n(docs/provider-integration/tiers/tier-2-catalog-entry.md,\n\"Live verification\"), gated by pnpm run verify:provider-onboarding. This\nentry was added under a credential-free onboarding pass instead: no\nChutes API key was ever created or used, and no authenticated or POST\nrequest was sent to any Chutes endpoint. Everything below is either a public\ndocument or a single unauthenticated GET.\n\n| Probe | Result |\n| ------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |\n| Roster | Unauthenticated GET https://llm.chutes.ai/v1/models, HTTP 200, 14 models, re-run 2026-09-28. /v1/models is itself documented by Chutes as a \"public catalog (no key required).\" |\n| Wire shape | Chutes' own homepage Python sample posts to https://llm.chutes.ai/v1/chat/completions with model + messages + Authorization: Bearer $CHUTES_API_KEY + stream: true — never executed here, only read. |\n| Tools / structured output | Confirmed from docs only: the \"Function Calling, Agents, and Tool Use\" guide shows OpenAI-shaped tools/tool_choice/tool_call_id, and the same guide shows a separate response_format: {\"type\": \"json_object\"} JSON-mode example.","hierarchy":{"lvl0":"Getting Started","lvl1":"Chutes Provider Guide","lvl2":"Verification status","lvl3":""}},
5602
5603
  {"objectID":"34c0be7e44128ded319ab20090fe5bfecd98725996a4350660f6dad0932a5601","title":"Troubleshooting","url":"/docs/getting-started/providers/chutes#troubleshooting","content":"| Symptom | Cause | Fix |\n| ---------------------------------------------------------- | ------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------- |\n| Requests hit an anonymous rate limit (HTTP 429) | No token, or an invalid token, was sent | Set CHUTES_API_KEY, and send it as Authorization: Bearer <key> — X-API-Key is documented as unsupported for inference calls |\n| Model not found | The roster changed since 2026-09-28 | Pick a current id from the unauthenticated GET /v1/models roster |\n| Structured output silently dropped when tools are used | structuredOutputWithTools is false — no combined probe was possible | This is intentional pending a live-key check, not a bug; NeuroLink's runtime conflict retry drops structured output and retries rather than losing the turn |\n| Unsure whether a specific model supports tools / JSON mode | capabilities.tools is \"model-dependent\" — two roster ids ship no supported_features at all | Check the model's entry in GET /v1/models, or consult the Models table above |","hierarchy":{"lvl0":"Getting Started","lvl1":"Chutes Provider Guide","lvl2":"Troubleshooting","lvl3":""}},
5603
5604
  {"objectID":"f743dd0fbfd3517546cecc2240a8513ad5a47938ceae45eec2ad4fe45aa06d9f","title":"See also","url":"/docs/getting-started/providers/chutes#see-also","content":"Provider setup overview\nAll providers\nTier-2 onboarding — how this provider's JSON becomes a working integration, and what live verification still needs to happen\nProvider feature compatibility","hierarchy":{"lvl0":"Getting Started","lvl1":"Chutes Provider Guide","lvl2":"See also","lvl3":""}},
5604
- {"objectID":"3803b4ecde578f946b8a4fe5139e3ee640d53281541c50578e8a5d47847d5157","title":"Cloudflare Workers AI Provider Guide","url":"/docs/getting-started/providers/cloudflare","content":"Cloudflare Workers AI Provider Guide\n\nOpen-model inference at the edge via Cloudflare Workers AI\n\nOverview\n\nCloudflare Workers AI\nserves Meta Llama, Mistral, and other open models from Cloudflare's\nglobal GPU cluster. NeuroLink talks to the OpenAI-compatible endpoint.\n\nKey Facts\nProtocol: OpenAI-compatible (/v1/chat/completions)\nDefault base URL: https://api.cloudflare.com/client/v4/accounts/{accountId}/ai/v1\nDefault model: @cf/meta/llama-3.3-70b-instruct-fp8-fast\nStreaming: Yes\nTool calling: Limited (model-dependent)\n\nQuick Start\nGet Credentials\n\nYou need both:\nA Cloudflare Account ID (Cloudflare dashboard → right sidebar)\nA Workers AI API token with the Workers AI Read & Write\n permission (Profile → API Tokens → Create Token)\nConfigure\nGenerate\n\nSupported Models (sample)\n\n| Model ID | Notes |\n| ------------------------------------------ | -------------- |\n| @cf/meta/llama-3.3-70b-instruct-fp8-fast | Default |\n| @cf/meta/llama-3.1-70b-instruct | Llama 3.1 70B |\n| @cf/meta/llama-3.1-8b-instruct | Fast tier |\n| @cf/meta/llama-3.2-11b-vision-instruct | Vision-capable |\n\nBrowse: https://developers.cloudflare.com/workers-ai/models\n\nCLI Usage\n\nProvider Aliases\n\n| Alias | Example |\n| ------------ | ----------------------- |\n| cloudflare | --provider cloudflare |\n| cf | --provider cf |\n\nConfiguration Reference\n\n| Environment Variable | Required | Default |\n| ----------------------- | -------- | ------------------------------------------ |\n| CLOUDFLARE_ACCOUNT_ID | Yes | — |\n| CLOUDFLARE_API_KEY | Yes | — |\n| CLOUDFLARE_MODEL | No | @cf/meta/llama-3.3-70b-instruct-fp8-fast |\n\nSee Also\nTogether AI Provider\nFireworks Provider","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Workers AI Provider Guide","lvl2":"","lvl3":""}},
5605
+ {"objectID":"1fb3d0f4a1510905b14436d78a4928a8e825b3dbdb324b6cb7c9a340e088d0f6","title":"Cloudflare Clef Provider Guide","url":"/docs/getting-started/providers/cloudflare-clef","content":"Cloudflare Clef Provider Guide\n\nA provider of decide — the same typed boolean / choice / score\nanswers as TypeSafe's Jev, Laya, XOR and\nPerplexity, from Cloudflare's Clef models on Workers\nAI, and it also reads images. It emits no text at all.\n\nThis is not the Cloudflare Workers AI text provider\n(cloudflare, which runs chat models). The two are separate providers that\nread the same CLOUDFLARE_API_KEY and CLOUDFLARE_ACCOUNT_ID; see\nOne token, two providers.\n\nRead Limits before you rely on it. Workers AI reads only about\nthe first 2,048 tokens of the state, far less than the 64K context Cloudflare\ndocuments, and ignores text past that point without an error (hosted service\nor model: unknown). NeuroLink refuses a state it\nestimates as longer than that rather than let a decision be made on text the\nmodel never saw.\n\nOverview\n\nClef is a decision model: it reads a state and a set of typed questions and\nreturns a probability for every allowed answer, with no free-form text and no\nreasoning tokens to wait for. Cloudflare publishes two sizes, clef (27B) and\nclef-flash (9B), and says it is releasing their weights under the Apache 2.0\nlicense. Both answered in about a second from a developer machine. Cloudflare's\npages call them available, and do not say whether they are generally available\nor in beta.\n\nKey Facts\n\n| | |\n| ----------------------- | ----------------------------------------------------------------------------------------------------------------------------------- |\n| Provider id | cloudflare-clef |\n| Models | clef (27B, the default) and clef-flash (9B). @cf/cloudflare/clef and @cf/cloudflare/clef-flash are accepted too |\n| Question types | boolean (sent as noul), choice (2 to 255 options), score (2 to 10 levels) |\n| Questions per request | 64 |\n| Images | up to 4 per request: PNG, JPEG or WebP. No video |\n| State | Workers AI reads about the first 2,048 tokens; NeuroLink refuses a state it estimates at more than 1,500 |\n| Request size | 256,000 bytes in NeuroLink, base64 image data included |\n| Endpoint | https://api.cloudflare.com/client/v4/accounts/<account id>/ai/run/@cf/cloudflare/<model> |\n| Credentials | an API token with Workers AI permission, and the account id |\n| Price | $0.24 per million input tokens for clef, $0.09 for clef-flash; no output price is listed (Cloudflare's Workers AI pricing page) |\n| Measured latency | 2026-10-03: 0.3 to 1.0 s for a small request; 64 questions: 1.1 s (clef-flash), 1.3 s (clef); 2026-10-04: 1.5 s / 2.3 s |\n| Default decide provider | last, after TypeSafe, Laya, XOR and Perplexity |\n| Verified live | 2026-10-03 and 2026-10-04, with a real account (see Limits) |\n\nQuick Start\nGet a token and your account id\n\nCreate an API token with the Workers AI: Read + Write permission at\ndash.cloudflare.com/profile/api-tokens, and copy your account id from\nthe dashboard URL or the \"Account ID\" panel.\nConfigure\n\nOr pass them in code, which wins over the environment:\n\nBoth the token and the account id are needed: the account id is part of the\nroute, so a token alone does not count as configured.\nUse it\n\nThe figures in the comments are what a live call to clef returned for this\nexample on 2026-10-03, rounded; clef-flash answered the same example with\n0.96, 0.82 and 0.50, 2.72. Pick the faster model for one call with\nmodel: \"clef-flash\".\n\nFrom the CLI\n\nImages\n\nPass up to four images beside the state, as a Buffer, a file path or a\ndata:image/…;base64, URL; an http(s) URL is refused. PNG, JPEG and WebP are\nread; image/jpg is accepted as a JPEG alias and sent as image/jpeg. Other formats are refused before they are sent.\nThey travel in their own images array, not inside the state. Cloudflare\n places them before the state.\nThe Workers AI API accepts no video. Its model page says it does, but the API refuses\n a video and has no field for one, so NeuroLink refuses a video too.\nPixels: Cloudflare documents 16 megapixels an image. 16.00 megapixe","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Clef Provider Guide","lvl2":"","lvl3":""}},
5606
+ {"objectID":"eebb9363ccd61130d0be9464df57f02fd61d00b6ff81821dd3162c030c0bc0db","title":"Cloudflare Clef Provider Guide","url":"/docs/getting-started/providers/cloudflare-clef#cloudflare-clef-provider-guide","content":"A provider of decide — the same typed boolean / choice / score\nanswers as TypeSafe's Jev, Laya, XOR and\nPerplexity, from Cloudflare's Clef models on Workers\nAI, and it also reads images. It emits no text at all.\n\nThis is not the Cloudflare Workers AI text provider\n(cloudflare, which runs chat models). The two are separate providers that\nread the same CLOUDFLARE_API_KEY and CLOUDFLARE_ACCOUNT_ID; see\nOne token, two providers.\n\nRead Limits before you rely on it. Workers AI reads only about\nthe first 2,048 tokens of the state, far less than the 64K context Cloudflare\ndocuments, and ignores text past that point without an error (hosted service\nor model: unknown). NeuroLink refuses a state it\nestimates as longer than that rather than let a decision be made on text the\nmodel never saw.","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Clef Provider Guide","lvl2":"Cloudflare Clef Provider Guide","lvl3":""}},
5607
+ {"objectID":"d3aa34515e7690cc133be051b57b7fda1e8fdd225ef0dd7a162d91ee2ffa91d8","title":"Overview","url":"/docs/getting-started/providers/cloudflare-clef#overview","content":"Clef is a decision model: it reads a state and a set of typed questions and\nreturns a probability for every allowed answer, with no free-form text and no\nreasoning tokens to wait for. Cloudflare publishes two sizes, clef (27B) and\nclef-flash (9B), and says it is releasing their weights under the Apache 2.0\nlicense. Both answered in about a second from a developer machine. Cloudflare's\npages call them available, and do not say whether they are generally available\nor in beta.","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Clef Provider Guide","lvl2":"Overview","lvl3":""}},
5608
+ {"objectID":"babb1678dd2f0f63c1db8cf3e5216f3ac6b22a00331fe49058519957948ca5fb","title":"Key Facts","url":"/docs/getting-started/providers/cloudflare-clef#key-facts","content":"| | |\n| ----------------------- | ----------------------------------------------------------------------------------------------------------------------------------- |\n| Provider id | cloudflare-clef |\n| Models | clef (27B, the default) and clef-flash (9B). @cf/cloudflare/clef and @cf/cloudflare/clef-flash are accepted too |\n| Question types | boolean (sent as noul), choice (2 to 255 options), score (2 to 10 levels) |\n| Questions per request | 64 |\n| Images | up to 4 per request: PNG, JPEG or WebP.","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Clef Provider Guide","lvl2":"Key Facts","lvl3":""}},
5609
+ {"objectID":"5f8fca781ea20214c667a17eb2e49fc41c5a9187b58b202f67758f2bfc581c4e","title":"1. Get a token and your account id","url":"/docs/getting-started/providers/cloudflare-clef#1-get-a-token-and-your-account-id","content":"Create an API token with the Workers AI: Read + Write permission at\ndash.cloudflare.com/profile/api-tokens, and copy your account id from\nthe dashboard URL or the \"Account ID\" panel.","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Clef Provider Guide","lvl2":"1. Get a token and your account id","lvl3":""}},
5610
+ {"objectID":"592a6f5702dc95c7cba68e1233dc74e8d60465eb77496a4313b4a096757237b3","title":"2. Configure","url":"/docs/getting-started/providers/cloudflare-clef#2-configure","content":"`bash","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Clef Provider Guide","lvl2":"2. Configure","lvl3":""}},
5611
+ {"objectID":"de4ffe6ae22db4a0719caae8e38b4a0a8a438224c94701a762c9e1f4cc209393","title":"Optional","url":"/docs/getting-started/providers/cloudflare-clef#optional","content":"typescript\n\nconst neurolink = new NeuroLink({\n credentials: {\n cloudflareClef: { apiKey: \"…\", accountId: \"…\" }, // baseURL is optional\n },\n});\n`\n\nBoth the token and the account id are needed: the account id is part of the\nroute, so a token alone does not count as configured.","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Clef Provider Guide","lvl2":"Optional","lvl3":""}},
5612
+ {"objectID":"0b26521b87bd72415fe58482c8628d4c3d04a49575ca2e89a55f29cf65834ccc","title":"3. Use it","url":"/docs/getting-started/providers/cloudflare-clef#3-use-it","content":"The figures in the comments are what a live call to clef returned for this\nexample on 2026-10-03, rounded; clef-flash answered the same example with\n0.96, 0.82 and 0.50, 2.72. Pick the faster model for one call with\nmodel: \"clef-flash\".","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Clef Provider Guide","lvl2":"3. Use it","lvl3":""}},
5613
+ {"objectID":"251d84b27d167ed1d24cad4c91fb45246bf057457ea226ad977f4e581af03d28","title":"Images","url":"/docs/getting-started/providers/cloudflare-clef#images","content":"Pass up to four images beside the state, as a Buffer, a file path or a\ndata:image/…;base64, URL; an http(s) URL is refused. PNG, JPEG and WebP are\nread; image/jpg is accepted as a JPEG alias and sent as image/jpeg. Other formats are refused before they are sent.\nThey travel in their own images array, not inside the state. Cloudflare\n places them before the state.\nThe Workers AI API accepts no video. Its model page says it does, but the API refuses\n a video and has no field for one, so NeuroLink refuses a video too.\nPixels: Cloudflare documents 16 megapixels an image. 16.00 megapixels were\n accepted (4096 Ɨ 3900 at 15.97 MP, 8000 Ɨ 2000 and 16000 Ɨ 1000 at 16.00 MP);\n 16.38 megapixels (4096 Ɨ 4000) were refused with a 422. The limit is on the\n pixel count, not on either side.\nCost: an image costs at most about 1,000 input tokens however large it is.\n A test request with one 512 Ɨ 512 image counted 389 input tokens in all, one\n with a 1,024 Ɨ 1,024 image counted 1,157, and larger images up to the pixel\n limit counted between 1,125 and 1,157.\nHow many: exactly four tiny PNG/JPEG/WebP images succeeded on both models\n on 2026-10-04; a fifth was refused with a 422. Four is also NeuroLink's limit.\nNeuroLink's size limit is 256,000 bytes for the encoded request body,\n including base64 image data. A 195 KB PNG encodes to about 260,000 characters\n and is refused locally. Resize or compress a photo to well under 190 KB first.\n The 2026-10-03 API probe accepted a 195 KB PNG and refused a 202 KB one.\n On 2026-10-04 the text-only request ceiling moved (see below); the current\n image-byte ceiling was not remeasured, so those image figures are historical.","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Clef Provider Guide","lvl2":"Images","lvl3":""}},
5614
+ {"objectID":"f98057fa646a175aee28fef3b525786c83a66b8ac8489feec28fd872866cf4ab","title":"When NeuroLink uses it","url":"/docs/getting-started/providers/cloudflare-clef#when-neurolink-uses-it","content":"Built-in features that call decide() pick the first decision provider that is\nconfigured, in this order: TypeSafe, Laya, XOR, Perplexity, then Clef. A host\nwith none of the first four, but with CLOUDFLARE_API_KEY and\nCLOUDFLARE_ACCOUNT_ID set for the Workers AI text provider, therefore has Clef\nas its default decision provider. Naming provider: \"cloudflare-clef\" reaches it\nwhatever else is configured.\n\nBecause of the state window below, a built-in feature whose state is longer than\nabout 1,500 estimated tokens gets a refusal from Clef, which each of them treats\nas \"carry on as before\". Clef suits short decisions: routing a request, a\nguardrail on one action, classifying one message or one picture.","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Clef Provider Guide","lvl2":"When NeuroLink uses it","lvl3":""}},
5615
+ {"objectID":"69f3db3e89a8fd474930cba6df7b01125fa2d9dd41aaadf8d8838423632ace30","title":"What is sent to Cloudflare","url":"/docs/getting-started/providers/cloudflare-clef#what-is-sent-to-cloudflare","content":"boolean is the SDK's name for the wire's noul. model is sent although the\npath already names it: Cloudflare's schema marks it required, and a body whose\nmodel differs from the path is refused. A request with no model was accepted\nlive and answered by clef-flash, so sending it matters only for following the\nschema. Cloudflare wraps every answer in its own\nenvelope, and NeuroLink reads the answers out of it:\n\nThe request id is the cf-ai-req-id response header; where that is absent (a\nrejected token never reaches the model) it is the edge's cf-ray.","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Clef Provider Guide","lvl2":"What is sent to Cloudflare","lvl3":""}},
5616
+ {"objectID":"37eb6d126d3af6ca22085fbf94025a8f23da15691c8f6070b946f5aaa27e2f10","title":"One token, two providers","url":"/docs/getting-started/providers/cloudflare-clef#one-token-two-providers","content":"cloudflare-clef and cloudflare read the same two variables, but each has its\nown credentials slice, so the two cannot be mixed up:\n\n| | cloudflare (text) | cloudflare-clef (decide) |\n| --------------- | --------------------------- | ---------------------------- |\n| Token | CLOUDFLARE_API_KEY | CLOUDFLARE_API_KEY |\n| Account id | CLOUDFLARE_ACCOUNT_ID | CLOUDFLARE_ACCOUNT_ID |\n| SDK credentials | credentials.cloudflare | credentials.cloudflareClef |\n| Serves | generate() and stream() | decide() only |\n\ncredentials.cloudflare does not configure decide. Setting the two\nvariables for the text provider also lets decide() use Clef when no other\ndecision provider is configured, and no switch turns that off while the two\nvariables are set in the environment. A host that wants the Workers AI text\nprovider without Clef can give the text provider its token and account id only\nthrough credentials.cloudflare, and leave the environment variables unset.","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Clef Provider Guide","lvl2":"One token, two providers","lvl3":""}},
5617
+ {"objectID":"fee62c91a128411e33fcad60cf048405bd86ff0f4015189d992b70d34fe1b86f","title":"Measured on a real account, October 2026","url":"/docs/getting-started/providers/cloudflare-clef#measured-on-a-real-account-october-2026","content":"The Workers AI endpoint ignores text past about 2,048 tokens, far below\nCloudflare's documented 64K context (hosted service or model: unknown). Request-size refusals also changed between two measurement dates.\nThe original measurements were on 2026-10-03; the bounded follow-up campaign\nused 133 probe calls on 2026-10-04, with one request at a time.\n\nWhich model: on both clef-flash and clef, the follow-up still read the\nfact at the 2026-10-03 clef-flash lower bound for logs, number lists, digit\narrays and compact JSON, and did not read it about 2.5% further on. English prose and\nrandom CJK had already matched on both models. It also checked the question,\nid, option, score-level, PNG/JPEG/WebP, pixel and four/five-image boundaries on\nclef, and the four-image and image/jpg cases on both models. Natural-script\nprose was measured on clef; a reworked many-key object probe matched on both.\nTypeScript, minified JSON and synthetic Devanagari/emoji cuts remain 9B-only\nmeasurements. Historical latency, billing and burst figures retain their dates.","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Clef Provider Guide","lvl2":"Measured on a real account, October 2026","lvl3":""}},
5618
+ {"objectID":"39c1f7ffec3b0df62b067ce9d778c3752d09cb8fdecce80b4e64028655048470","title":"What NeuroLink refuses before any request","url":"/docs/getting-started/providers/cloudflare-clef#what-neurolink-refuses-before-any-request","content":"| Refused | Limit | Error kind |\n| --------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------ | ------------------------------------- |\n| A state estimated over 1,500 tokens | digits 1 token each, punctuation 0.75, emoji 3, other non-ASCII text 1.5, letters about 4 characters a token | max_tokens_exceeded |\n| More than 64 questions in one decide() | 64 | max_tokens_exceeded |\n| A request body over 256,000 bytes | 256,000 | invalid_request |\n| More than 4 images, a video, or a format other than PNG, JPEG or WebP | | invalid_request |\n| A model name that is not a plain name, such as clef/../x | | invalid_request |\n| A missing token or account id, or a base URL that cannot work | | authentication or invalid_request |\n\nNeuroLink.tryDecide(), which every built-in feature calls, splits a question\nmap at 64 and runs up to four batches at once, so only a direct decide() meets\nthe 64-question refusal.\n\nWhy 1,500 and not 2,048.","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Clef Provider Guide","lvl2":"What NeuroLink refuses before any request","lvl3":""}},
5619
+ {"objectID":"bb0d1df2801529da652fbd15360879d9b2ccce53dfdf4d825a7a00f4e19d8d15","title":"What is not verified","url":"/docs/getting-started/providers/cloudflare-clef#what-is-not-verified","content":"The status Cloudflare returns for a token without Workers AI permission. No\n such token was available. A 403 is treated as an ordinary invalid_request,\n not authentication, so that if it does mean a missing permission, fixing it\n in the dashboard works at once, without restarting the process.\nAny 5xx. Cloudflare answered none during the probe, so the retry and\n classification of 500, 502, 503 and 504 rest on ordinary HTTP conventions.\nWhat happens to an account that has run out of credit.\nThe 27B TypeScript, minified-JSON and synthetic Devanagari/emoji cuts, and\n the image-byte ceiling after the observed request-size change; see the model\n coverage above.\nThe rate limit. Only a burst of 60 was tried.\nWhether Cloudflare's service is generally available or in beta.\nWhether the 2,048-token state window and the changing request-size limit belong\n to the hosted service or to the model itself. The open weights on Hugging Face\n were not run.\nWhether either limit changes. The live decide suite carries three canary cases\n (19.10, 19.10b and 19.11) that fail with an instruction to re-measure if it\n does; 19.11 checks both sides of the text-only request ceiling on clef-flash.","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Clef Provider Guide","lvl2":"What is not verified","lvl3":""}},
5620
+ {"objectID":"4214184c06ad6cee88e38838f55c38c5a554a78434494b5d765a1c120eb0e339","title":"Latency and the timeout","url":"/docs/getting-started/providers/cloudflare-clef#latency-and-the-timeout","content":"On 2026-10-03, a small request answered in 0.3 to 1.0 s from a developer\nmachine; 64 questions took 1.1 s on clef-flash and 1.3 s on clef; the\nslowest of 60 requests sent at once took 2.2 s. On 2026-10-04, a\n64-question request with the same questions and a shorter state took 1.5 s on clef-flash and 2.3 s on clef. The default timeout is 5 s. Pass timeoutMs to change it for\none call.","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Clef Provider Guide","lvl2":"Latency and the timeout","lvl3":""}},
5621
+ {"objectID":"c2d0b99e3c2760e14d8be1f81ba92c025ad96523cdeaa34dd002a748ae9658de","title":"Errors","url":"/docs/getting-started/providers/cloudflare-clef#errors","content":"Cloudflare answers with its own envelope, { \"success\": false, \"errors\": [{ \"code\",\n\"message\" }] }, and a refusal from the model nests a second envelope, with a\ntrailing request id, inside message. NeuroLink flattens both to one line and\ntakes the request id out of the text.\n\n| Status | Code | Seen for | Kind | Retried |\n| ------ | ----- | ----------------------------------------------------------------------------------------- | ----------------------------------- | ------------------------------ |\n| 401 | 10000 | a token Cloudflare does not know | authentication | no; trips the breaker |\n| 403 | | never seen; a token without Workers AI permission may draw it | invalid_request | no; does not trip the breaker |\n| 400 | 5006 | no questions; an unknown question type | invalid_request | no |\n| 400 | 7000 | a model path that does not exist (\"No route for that URI\") | invalid_request | no |\n| 400 | 6003 | a body that is not JSON | invalid_request | no |\n| 422 | 5012 | a validation failure: 65 questions, a fifth image, an unknown field, an empty instruction | invalid_request | no |\n| 413 | 5021 | a request past the estimate above | max_tokens_exceeded | no |\n| 429 | 3040 | \"Capacity temporarily exceeded\"; no Retry-After","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Clef Provider Guide","lvl2":"Errors","lvl3":""}},
5622
+ {"objectID":"ce05a3d9d2ab15b8867c8e721fbfa0f8e1248520e2413765fa7d70bd835e3f85","title":"Troubleshooting","url":"/docs/getting-started/providers/cloudflare-clef#troubleshooting","content":"\"requires the account id\" — set CLOUDFLARE_ACCOUNT_ID or\n credentials.cloudflareClef.accountId.\n\"The state is ~N tokens; Cloudflare's … model reads at most 1500\" — shorten\n the state, or send only the part the question is about. Cloudflare would have\n cut the rest without telling you.\n\"No route for that URI\" — the model name or the base URL is wrong. Clef is\n served as clef and clef-flash, and the base URL must end in /client/v4.\nA 429 \"Capacity temporarily exceeded\" — retried for you; if it keeps\n happening, send fewer requests at once.\nAn image is refused as \"The request is N bytes; Cloudflare accepts at most\n 256000\" — Cloudflare counts the base64 text of the image, so NeuroLink\n refuses it locally under the retained conservative cap. Resize or compress it\n to well under 190 KB; a 195 KB PNG, which Cloudflare accepted, is about 260,000\n characters in base64 and is refused here.\ndecide() uses Clef although you never chose it — you have the Workers AI\n token and account id set and no other decision provider; see\n When NeuroLink uses it.","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Clef Provider Guide","lvl2":"Troubleshooting","lvl3":""}},
5623
+ {"objectID":"64d2af754e23d6f77f9ba5991174048f9dc4e8434d0f9a633527785feff8158e","title":"See also","url":"/docs/getting-started/providers/cloudflare-clef#see-also","content":"The decide inference type\nTypeSafe (Jev) Provider Guide\nLaya Provider Guide\nXOR Provider Guide\nPerplexity Decisions Provider Guide\nCloudflare Workers AI text provider\nClef on Workers AI\nIntroducing Clef: our open-source decision models, and new RL fine-tuning platform (Cloudflare Blog)","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Clef Provider Guide","lvl2":"See also","lvl3":""}},
5624
+ {"objectID":"3803b4ecde578f946b8a4fe5139e3ee640d53281541c50578e8a5d47847d5157","title":"Cloudflare Workers AI Provider Guide","url":"/docs/getting-started/providers/cloudflare","content":"Cloudflare Workers AI Provider Guide\n\nOpen-model inference at the edge via Cloudflare Workers AI\n\nOverview\n\nCloudflare Workers AI\nserves Meta Llama, Mistral, and other open models from Cloudflare's\nglobal GPU cluster. NeuroLink talks to the OpenAI-compatible endpoint.\n\nKey Facts\nProtocol: OpenAI-compatible (/v1/chat/completions)\nDefault base URL: https://api.cloudflare.com/client/v4/accounts/{accountId}/ai/v1\nDefault model: @cf/meta/llama-3.3-70b-instruct-fp8-fast\nStreaming: Yes\nTool calling: Limited (model-dependent)\n\nQuick Start\nGet Credentials\n\nYou need both:\nA Cloudflare Account ID (Cloudflare dashboard → right sidebar)\nA Workers AI API token with the Workers AI Read & Write\n permission (Profile → API Tokens → Create Token)\nConfigure\nGenerate\n\nSupported Models (sample)\n\n| Model ID | Notes |\n| ------------------------------------------ | -------------- |\n| @cf/meta/llama-3.3-70b-instruct-fp8-fast | Default |\n| @cf/meta/llama-3.1-70b-instruct | Llama 3.1 70B |\n| @cf/meta/llama-3.1-8b-instruct | Fast tier |\n| @cf/meta/llama-3.2-11b-vision-instruct | Vision-capable |\n\nBrowse: https://developers.cloudflare.com/workers-ai/models\n\nCLI Usage\n\nProvider Aliases\n\n| Alias | Example |\n| ------------ | ----------------------- |\n| cloudflare | --provider cloudflare |\n| cf | --provider cf |\n\nConfiguration Reference\n\n| Environment Variable | Required | Default |\n| ----------------------- | -------- | ------------------------------------------ |\n| CLOUDFLARE_ACCOUNT_ID | Yes | — |\n| CLOUDFLARE_API_KEY | Yes | — |\n| CLOUDFLARE_MODEL | No | @cf/meta/llama-3.3-70b-instruct-fp8-fast |\n\nAlso configures decisions\n\nCLOUDFLARE_API_KEY and CLOUDFLARE_ACCOUNT_ID also configure a separate\nprovider, the Cloudflare Clef provider (cloudflare-clef),\nwhich serves decide() and emits no text. It is the last decision provider\nNeuroLink falls back to: built-in features that call decide() use the first one\nconfigured, in the order TypeSafe, Laya, XOR, Perplexity, then Clef. So a host\nthat set these two variables only for this text provider, and has no other\ndecision provider configured, has Clef as its default decision provider, and no\nswitch turns that off while they are set.\n\ncredentials.cloudflare does not configure decide; Clef reads its own\nslice, credentials.cloudflareClef. The text provider itself is unchanged: it\nstill serves generate() and stream() with the same variables, models and base\nURL. See\nOne token, two providers.\n\nSee Also\nCloudflare Clef Provider — typed decide() judgments on the same token and account id\nTogether AI Provider\nFireworks Provider","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Workers AI Provider Guide","lvl2":"","lvl3":""}},
5605
5625
  {"objectID":"aafa1a6f2adf2bacc5661cf9ccb309d1e32e300645f037c35a98217b31eb15df","title":"Cloudflare Workers AI Provider Guide","url":"/docs/getting-started/providers/cloudflare#cloudflare-workers-ai-provider-guide","content":"Open-model inference at the edge via Cloudflare Workers AI","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Workers AI Provider Guide","lvl2":"Cloudflare Workers AI Provider Guide","lvl3":""}},
5606
5626
  {"objectID":"25503e78c2d92fb7cee70f0a5e3045de4dbad7279ef1b5cb1a6c834a89c4d006","title":"Overview","url":"/docs/getting-started/providers/cloudflare#overview","content":"Cloudflare Workers AI\nserves Meta Llama, Mistral, and other open models from Cloudflare's\nglobal GPU cluster. NeuroLink talks to the OpenAI-compatible endpoint.","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Workers AI Provider Guide","lvl2":"Overview","lvl3":""}},
5607
5627
  {"objectID":"11f1e13f3ceb9f43bdaceb8669617ec44d6a62676514a7899f65463375681501","title":"Key Facts","url":"/docs/getting-started/providers/cloudflare#key-facts","content":"Protocol: OpenAI-compatible (/v1/chat/completions)\nDefault base URL: https://api.cloudflare.com/client/v4/accounts/{accountId}/ai/v1\nDefault model: @cf/meta/llama-3.3-70b-instruct-fp8-fast\nStreaming: Yes\nTool calling: Limited (model-dependent)","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Workers AI Provider Guide","lvl2":"Key Facts","lvl3":""}},
@@ -5609,7 +5629,8 @@
5609
5629
  {"objectID":"088e2cc83205e7cef6bcdc0633aee37efb875a215914292fe0d8472ad9956f7f","title":"Supported Models (sample)","url":"/docs/getting-started/providers/cloudflare#supported-models-sample","content":"| Model ID | Notes |\n| ------------------------------------------ | -------------- |\n| @cf/meta/llama-3.3-70b-instruct-fp8-fast | Default |\n| @cf/meta/llama-3.1-70b-instruct | Llama 3.1 70B |\n| @cf/meta/llama-3.1-8b-instruct | Fast tier |\n| @cf/meta/llama-3.2-11b-vision-instruct | Vision-capable |\n\nBrowse: https://developers.cloudflare.com/workers-ai/models","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Workers AI Provider Guide","lvl2":"Supported Models (sample)","lvl3":""}},
5610
5630
  {"objectID":"e1e8b466c22f042a3d3095b12f19a9fe24a96d6480b008310cfc189f341d45a5","title":"Provider Aliases","url":"/docs/getting-started/providers/cloudflare#provider-aliases","content":"| Alias | Example |\n| ------------ | ----------------------- |\n| cloudflare | --provider cloudflare |\n| cf | --provider cf |","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Workers AI Provider Guide","lvl2":"Provider Aliases","lvl3":""}},
5611
5631
  {"objectID":"62de16314951fdc73f9dd06c4c146e4c9202ddb9445bea480242fcac6cf0ab83","title":"Configuration Reference","url":"/docs/getting-started/providers/cloudflare#configuration-reference","content":"| Environment Variable | Required | Default |\n| ----------------------- | -------- | ------------------------------------------ |\n| CLOUDFLARE_ACCOUNT_ID | Yes | — |\n| CLOUDFLARE_API_KEY | Yes | — |\n| CLOUDFLARE_MODEL | No | @cf/meta/llama-3.3-70b-instruct-fp8-fast |","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Workers AI Provider Guide","lvl2":"Configuration Reference","lvl3":""}},
5612
- {"objectID":"c18937da302797e2d1727a8d5894df09424bcf4d107a5500e2a04ce958eb9673","title":"See Also","url":"/docs/getting-started/providers/cloudflare#see-also","content":"Together AI Provider\nFireworks Provider","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Workers AI Provider Guide","lvl2":"See Also","lvl3":""}},
5632
+ {"objectID":"a12a3191a8590f0031973af941fe6c2be7f7cf931536ea38903ef8e8bf9bb5c6","title":"Also configures decisions","url":"/docs/getting-started/providers/cloudflare#also-configures-decisions","content":"CLOUDFLARE_API_KEY and CLOUDFLARE_ACCOUNT_ID also configure a separate\nprovider, the Cloudflare Clef provider (cloudflare-clef),\nwhich serves decide() and emits no text. It is the last decision provider\nNeuroLink falls back to: built-in features that call decide() use the first one\nconfigured, in the order TypeSafe, Laya, XOR, Perplexity, then Clef. So a host\nthat set these two variables only for this text provider, and has no other\ndecision provider configured, has Clef as its default decision provider, and no\nswitch turns that off while they are set.\n\ncredentials.cloudflare does not configure decide; Clef reads its own\nslice, credentials.cloudflareClef. The text provider itself is unchanged: it\nstill serves generate() and stream() with the same variables, models and base\nURL. See\nOne token, two providers.","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Workers AI Provider Guide","lvl2":"Also configures decisions","lvl3":""}},
5633
+ {"objectID":"c18937da302797e2d1727a8d5894df09424bcf4d107a5500e2a04ce958eb9673","title":"See Also","url":"/docs/getting-started/providers/cloudflare#see-also","content":"Cloudflare Clef Provider — typed decide() judgments on the same token and account id\nTogether AI Provider\nFireworks Provider","hierarchy":{"lvl0":"Getting Started","lvl1":"Cloudflare Workers AI Provider Guide","lvl2":"See Also","lvl3":""}},
5613
5634
  {"objectID":"771db3571626e1c1eb3472bc9a61338f884160a5415735e50731de5e772d00f0","title":"Cohere Provider Guide","url":"/docs/getting-started/providers/cohere","content":"Cohere Provider Guide\n\nCommand R chat + Embed v3 embeddings via the Cohere API\n\nOverview\n\nCohere offers a production-grade chat\ncatalog (Command R / R+ / R7B) plus top-tier embeddings (Embed v3) and\nreranking (Rerank v3). NeuroLink wraps chat via the OpenAI-compatible\nendpoint and embeddings via the native /v2/embed endpoint.\n\nKey Facts\nProtocol: OpenAI-compatible chat at /compatibility/v1,\n native embed at /v2/embed\nDefault base URL: https://api.cohere.com/compatibility/v1\nDefault chat model: command-r-plus-08-2024\nDefault embed model: embed-english-v3.0\nStreaming: Yes\nTool calling: Yes (Command R / R+)\n\nQuick Start\nGet an API Key\n\nhttps://dashboard.cohere.com/api-keys\nConfigure\nGenerate Text\nGenerate Embeddings\n\nSupported Models\n\n| Model ID | Family | Notes |\n| ----------------------------- | ---------------- | ------------------------ |\n| command-r-plus-08-2024 | Chat (default) | Flagship |\n| command-r-08-2024 | Chat | Mid-tier |\n| command-r7b-12-2024 | Chat | Most compact |\n| command-a-reasoning-08-2025 | Reasoning | Reasoning traces |\n| embed-english-v3.0 | Embeddings (def) | 1024 dim, English |\n| embed-multilingual-v3.0 | Embeddings | 1024 dim, 100+ languages |\n\nCLI Usage\n\nConfiguration Reference\n\n| Environment Variable | Required | Default |\n| -------------------- | -------- | ----------------------------------------- |\n| COHERE_API_KEY | Yes | — |\n| COHERE_MODEL | No | command-r-plus-08-2024 |\n| COHERE_BASE_URL | No | https://api.cohere.com/compatibility/v1 |\n\nFeature Support Matrix\n\n| Feature | Support |\n| ----------------- | ------------ |\n| Text generation | Yes |\n| Streaming | Yes |\n| Tool calling | Yes |\n| Structured output | Yes |\n| Vision | No |\n| Embeddings | Yes (native) |\n| Reranking | Yes |\n\nTroubleshooting\nNot Found — the chosen model may not be available on your tier.\n Try command-r-plus-08-2024 or command-r-08-2024.\n\nSee Also\nVoyage Provider\nJina Provider","hierarchy":{"lvl0":"Getting Started","lvl1":"Cohere Provider Guide","lvl2":"","lvl3":""}},
5614
5635
  {"objectID":"765ab44d68aaf7fc0fe06a519c53fb7aad3106753d2a1b679c26293cce92e958","title":"Cohere Provider Guide","url":"/docs/getting-started/providers/cohere#cohere-provider-guide","content":"Command R chat + Embed v3 embeddings via the Cohere API","hierarchy":{"lvl0":"Getting Started","lvl1":"Cohere Provider Guide","lvl2":"Cohere Provider Guide","lvl3":""}},
5615
5636
  {"objectID":"8b1406677ad9cf4ecbc1b8fe434c6f9d9312515aad5ccc354e2f8a7299f395ec","title":"Overview","url":"/docs/getting-started/providers/cohere#overview","content":"Cohere offers a production-grade chat\ncatalog (Command R / R+ / R7B) plus top-tier embeddings (Embed v3) and\nreranking (Rerank v3). NeuroLink wraps chat via the OpenAI-compatible\nendpoint and embeddings via the native /v2/embed endpoint.","hierarchy":{"lvl0":"Getting Started","lvl1":"Cohere Provider Guide","lvl2":"Overview","lvl3":""}},
@@ -6154,11 +6175,12 @@
6154
6175
  {"objectID":"6ad36ba3488491bb913c7a77b86a278cc86509da88a054975bcded71f5c02130","title":"[OpenRouter](/docs/getting-started/providers/openrouter)","url":"/docs/getting-started/providers#openrouterdocsgetting-startedprovidersopenrouter","content":"300+ models from 60+ providers\n🌐 Single API across many AI providers (Anthropic, OpenAI, Google, Meta, etc.)\n⚔ Automatic failover and routing\nšŸ’° Competitive pricing with cost optimization\nšŸŽÆ Zero lock-in - switch models instantly\nšŸ“Š Usage tracking dashboard\nšŸ†“ Free models available\n\nSetup Guide →","hierarchy":{"lvl0":"Getting Started","lvl1":"AI Provider Guides","lvl2":"[OpenRouter](/docs/getting-started/providers/openrouter)","lvl3":""}},
6155
6176
  {"objectID":"733a2fa8ea4c547cdcdf502f89f003a2208c0e5828e36e7faaaa94e0903d3d49","title":"[OpenAI Compatible](/docs/getting-started/providers/openai-compatible)","url":"/docs/getting-started/providers#openai-compatibledocsgetting-startedprovidersopenai-compatible","content":"OpenRouter, vLLM, LocalAI, and more\n🌐 100+ models through OpenRouter\nšŸ’» Local deployment with vLLM\nšŸ”“ Self-hosted with LocalAI\nšŸ”„ Drop-in OpenAI replacement\n\nSetup Guide →","hierarchy":{"lvl0":"Getting Started","lvl1":"AI Provider Guides","lvl2":"[OpenAI Compatible](/docs/getting-started/providers/openai-compatible)","lvl3":""}},
6156
6177
  {"objectID":"fc415454bfde58f3634df544ef2adb88836c75fb1344b7042af7a9d6fdfa63d4","title":"[LiteLLM](/docs/getting-started/providers/litellm)","url":"/docs/getting-started/providers#litellmdocsgetting-startedproviderslitellm","content":"100+ providers through proxy\nšŸ”„ Unified API for 100+ providers\nšŸ“Š Load balancing and fallbacks\nšŸ’° Cost tracking\nšŸŽÆ Model routing\n\nSetup Guide →","hierarchy":{"lvl0":"Getting Started","lvl1":"AI Provider Guides","lvl2":"[LiteLLM](/docs/getting-started/providers/litellm)","lvl3":""}},
6157
- {"objectID":"6337b6964b5ebadd9610f93633594d912ee39f49c2ef5e1683214dce45a1aee3","title":"🧠 Decision-Only Providers","url":"/docs/getting-started/providers#decision-only-providers","content":"The providers that serve decide rather than generate/stream:\nTypeSafe, Laya, XOR and\nPerplexity. Each returns typed boolean/choice/score\nanswers and emits no text, so none appears in generation fallback chains or the\nhealth sweep.","hierarchy":{"lvl0":"Getting Started","lvl1":"AI Provider Guides","lvl2":"🧠 Decision-Only Providers","lvl3":""}},
6178
+ {"objectID":"6337b6964b5ebadd9610f93633594d912ee39f49c2ef5e1683214dce45a1aee3","title":"🧠 Decision-Only Providers","url":"/docs/getting-started/providers#decision-only-providers","content":"The providers that serve decide rather than generate/stream:\nTypeSafe, Laya, XOR,\nPerplexity and Cloudflare Clef. Each returns typed boolean/choice/score\nanswers and emits no text, so none appears in generation fallback chains or the\nhealth sweep.","hierarchy":{"lvl0":"Getting Started","lvl1":"AI Provider Guides","lvl2":"🧠 Decision-Only Providers","lvl3":""}},
6158
6179
  {"objectID":"50860d1f7a27dbe3ad7452ad7079ea437658db372b97081d0a509429369cacc6","title":"[TypeSafe (Jev)](/docs/getting-started/providers/typesafe)","url":"/docs/getting-started/providers#typesafe-jevdocsgetting-startedproviderstypesafe","content":"Typed, calibrated judgments instead of text\nšŸŽÆ boolean / choice / score answers, each with a calibrated confidence\n⚔ Latency flat in question count — 1 question ~393 ms, 400 questions ~465 ms\nšŸ’° ~$0.00002 per decision (~$0.042/M input, output billed at zero)\nšŸ”Œ Two transports: TypeSafe direct, or the Vercel AI Gateway\nšŸ›”ļø Fails open — with no decision provider configured, every consumer behaves exactly as before\nšŸ”‘ API key from console.typesafe.ai/keys\nšŸ”„ Aliases: jev, typesafe-ai\n\nSetup Guide →","hierarchy":{"lvl0":"Getting Started","lvl1":"AI Provider Guides","lvl2":"[TypeSafe (Jev)](/docs/getting-started/providers/typesafe)","lvl3":""}},
6159
6180
  {"objectID":"43b838d60f2794ce46361196d6fc8a708b3a301398dd55f2c8015d567aa21c2c","title":"[Laya](/docs/getting-started/providers/laya)","url":"/docs/getting-started/providers#layadocsgetting-startedproviderslaya","content":"Open-weights decision provider — the same typed boolean/choice/score answers as Jev, from Convai Innovations' Apache-2.0 checkpoints, at a Laya server or LiteLLM proxy route you configure\n🧭 Serves decide only; built-in features use it when its key and base URL are set and neither TYPESAFE_API_KEY nor AI_GATEWAY_API_KEY is\nšŸ“ Refuses more than ~768 tokens of state (320 on english) before any network call, since its encoders read only 1,024 (512)\nšŸ”Œ No built-in endpoint: LAYA_BASE_URL (or credentials.laya.baseURL) names a Laya server or a LiteLLM pass-through route to one\nšŸ”‘ LAYA_API_KEY is the key that endpoint accepts; on LiteLLM, the route must be in the key's Allowed Routes\n\nSetup Guide →","hierarchy":{"lvl0":"Getting Started","lvl1":"AI Provider Guides","lvl2":"[Laya](/docs/getting-started/providers/laya)","lvl3":""}},
6160
6181
  {"objectID":"613cbeca5e8fa0bbe765b8db9ec4ed2ff76a913dd45b8dab5e2de06a6e5f19ff","title":"[XOR](/docs/getting-started/providers/xor)","url":"/docs/getting-started/providers#xordocsgetting-startedprovidersxor","content":"Open-weights decision provider — the same typed boolean/choice/score answers as Jev, from xor-1.1, Juspay's Apache-2.0 model, at a deployment or LiteLLM proxy route you configure\n🧭 Serves decide only; built-in features use it when its key and base URL are set and no TypeSafe key, no AI_GATEWAY_API_KEY and no Laya key and base URL is\nšŸ–¼ļø Takes up to 8 images or one video with a decision, as a Buffer, a local file path or a data: URL; TypeSafe and Laya refuse media before any request, and Perplexity refuses a video\nšŸ“ Refuses more than about 200,000 estimated tokens of state before any network call\nšŸ”Œ No built-in endpoint: XOR_BASE_URL (or credentials.xor.baseURL) is the origin of a deployment or of a LiteLLM route to one; requests go to <base URL>/v1/systemone\nšŸ”‘ XOR_API_KEY is the key that endpoint accepts; on LiteLLM, the key's team must allow xor-1.1\nšŸ“¦ Weights and setup: huggingface.co/juspay/xor\n\nSetup Guide →","hierarchy":{"lvl0":"Getting Started","lvl1":"AI Provider Guides","lvl2":"[XOR](/docs/getting-started/providers/xor)","lvl3":""}},
6161
6182
  {"objectID":"6ab8cdd8152a0d7c9844d54b6d4eaccc905b60e79f4b28508d78d52c3cfe9470","title":"[Perplexity Decisions](/docs/getting-started/providers/perplexity-decider)","url":"/docs/getting-started/providers#perplexity-decisionsdocsgetting-startedprovidersperplexity-decider","content":"Hosted decision provider — the same typed boolean/choice/score answers as Jev, from Perplexity's pplx-decider-v1-27b, at https://api.perplexity.ai\n🧭 Serves decide only (provider id perplexity-decider, not the Sonar text provider perplexity); built-in features use it when it is configured and no TypeSafe key, no AI_GATEWAY_API_KEY, no Laya key and base URL and no XOR key and base URL is\nšŸ”‘ PERPLEXITY_API_KEY alone configures it, and it is the same key the Perplexity text provider reads, so a key set for Sonar also lets built-in features use it when none of TypeSafe, Laya or XOR is configured; the guide lists what each consumer then sends\nšŸ–¼ļø Takes up to 8 PNG, JPEG or WebP images with a decision; no video\nšŸ’° $0.04 per million input tokens (image tokens included), output free, as Perplexity documents it\nšŸ“ Refuses more than about 100,000 estimated tokens of state (NeuroLink's own window; the server's ceiling is 262,144 input tokens), or more than 128 questions in a decide() call (tryDecide() splits them), before any network call. The 128-question cap, the server's token ceiling and the latency were measured on a real account in October 2026\nšŸ”Œ Hosted endpoint, so no base URL is needed; PERPLEXITY_DECIDER_BASE_URL (or credentials.perplexityDecider.baseURL) can name another origin\n\nSetup Guide →","hierarchy":{"lvl0":"Getting Started","lvl1":"AI Provider Guides","lvl2":"[Perplexity Decisions](/docs/getting-started/providers/perplexity-decider)","lvl3":""}},
6183
+ {"objectID":"94b24c75db64d3f04ba04b2ee0f04ec13bc0509de0aea1fa1b3e5516dea73ea9","title":"[Cloudflare Clef](/docs/getting-started/providers/cloudflare-clef)","url":"/docs/getting-started/providers#cloudflare-clefdocsgetting-startedproviderscloudflare-clef","content":"Hosted decision provider on Workers AI — the same typed boolean/choice/score answers as Jev, from Cloudflare's clef (27B) and clef-flash (9B), at https://api.cloudflare.com/client/v4\n🧭 Serves decide only (provider id cloudflare-clef, not the Workers AI text provider cloudflare); built-in features use it only when it is configured and none of TypeSafe, Laya, XOR or Perplexity is — it comes last\nšŸ”‘ CLOUDFLARE_API_KEY (Workers AI permission) and CLOUDFLARE_ACCOUNT_ID configure it, and they are the same two variables the Cloudflare text provider reads, so a pair set for that provider also lets built-in features use it when no other decision provider is configured\nšŸ–¼ļø Takes up to 4 PNG, JPEG or WebP images with a decision; no video\nšŸ’° $0.24 per million input tokens for clef, $0.09 for clef-flash; no output price is listed on Cloudflare's pricing page\nšŸ“ The endpoint ignores state text past about 2,048 tokens (hosted service or model: unknown), far less than the 64K Cloudflare documents, and says nothing when it does; NeuroLink refuses more than about 1,500 estimated tokens (digits count a token each, punctuation 0.75, emoji 3), more than 64 questions in a decide() call (tryDecide() splits them) and a request over 256,000 bytes, base64 image data included, before any network call. The state cut and the request ceiling were measured on a real account in October 2026, with the follow-up text cuts and API limits checked on both models; see the guide for the tested shapes and dates\nšŸ”Œ Hosted endpoint, so no base URL is needed; CLOUDFLARE_CLEF_BASE_URL (or credentials.cloudflareClef.baseURL) can name another base\n\nSetup Guide →","hierarchy":{"lvl0":"Getting Started","lvl1":"AI Provider Guides","lvl2":"[Cloudflare Clef](/docs/getting-started/providers/cloudflare-clef)","lvl3":""}},
6162
6184
  {"objectID":"c98403894a115164f29a019817b7f39c93b59623a9854b4575ad06bc5f597320","title":"🧩 Additional Catalog Providers","url":"/docs/getting-started/providers#-additional-catalog-providers","content":"Every provider below is a Tier-2 catalog entry. Provider-specific\nbehavioural quirks may apply, so check the provider's catalog file. The whole\nintegration is one JSON file under src/lib/providers/catalog/. Each page is\ngenerated from that file, which is also what the CI onboarding gate reads.","hierarchy":{"lvl0":"Getting Started","lvl1":"AI Provider Guides","lvl2":"🧩 Additional Catalog Providers","lvl3":""}},
6163
6185
  {"objectID":"026138da0a559269ecd9f0b2ceb72d9039096e72e064142afb7bfb74a5cb7455","title":"[API Route](/docs/getting-started/providers/api-route)","url":"/docs/getting-started/providers#api-routedocsgetting-startedprovidersapi-route","content":"Claude Sonnet 4.6\nšŸ¤– 8 models; default claude-sonnet-4-6 (1M context)\nšŸ› ļø Native tool calling + structured output together\nšŸ’³ Free tier available\nāœ… Roster verified 2026-09-17 (authenticated GET /v1/models)\nšŸ”‘ API key from api-route.com\nšŸ”„ Aliases: apiroute\n\nSetup Guide →","hierarchy":{"lvl0":"Getting Started","lvl1":"AI Provider Guides","lvl2":"[API Route](/docs/getting-started/providers/api-route)","lvl3":""}},
6164
6186
  {"objectID":"85aa65949d959db077347ecf93a4c40ade9ebab4300d4d52c62c702f6929053b","title":"[Baseten](/docs/getting-started/providers/baseten)","url":"/docs/getting-started/providers#basetendocsgetting-startedprovidersbaseten","content":"GLM 5.3 Flash\nšŸ¤– 16 models; default zai-org/GLM-5.3-Flash (1M context)\nšŸ› ļø Native tool calling + structured output together\nšŸ’³ Free tier available\nāœ… Roster verified 2026-09-03 (authenticated GET /v1/models)\nšŸ”‘ API key from app.baseten.co\n\nSetup Guide →","hierarchy":{"lvl0":"Getting Started","lvl1":"AI Provider Guides","lvl2":"[Baseten](/docs/getting-started/providers/baseten)","lvl3":""}},
@@ -6245,18 +6267,18 @@
6245
6267
  {"objectID":"4e0264790a48f8469b5d25d609cab04fae1ad3ba2ccee40c8829cd8571c265c6","title":"Verification status","url":"/docs/getting-started/providers/koscompute#verification-status","content":"Tier-2 onboarding requires evidence before a provider is accepted, and\npnpm run verify:provider-onboarding gates it in CI. This is what the\ncatalog records for KosCompute — and, just as importantly, what it does\nnot yet record:\n\n| Probe | Result |\n| --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |\n| Roster | unauthenticated GET /v1/models, HTTP 200, 7 models, 2026-09-28. No API key was used or required for this call. |\n| Auth rejection | Not probed. Would require a POST with an invalid key, which this credential-free onboarding pass does not send.","hierarchy":{"lvl0":"Getting Started","lvl1":"KosCompute Provider Guide","lvl2":"Verification status","lvl3":""}},
6246
6268
  {"objectID":"3944867506991deef1761c0dab4d9fa3bc0efe0d53bfb3f9acb78437c41f3ecf","title":"Troubleshooting","url":"/docs/getting-started/providers/koscompute#troubleshooting","content":"| Symptom | Cause | Fix |\n| ------------------------------------ | ----------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- |\n| Invalid KosCompute API key | KOSCOMPUTE_API_KEY unset, wrong, expired, or lacks access | KosCompute's docs describe 401 as \"missing, invalid, expired, inactive, or conflicting credentials\" but publish no example error body or code string |\n| Model not found | Wrong model id, or the roster changed since 2026-09-28 | Pick a current id from an unauthenticated GET https://api.koscompute.com/v1/models |\n| 429 with concurrent_limit_exceeded | Per-key concurrency limit exceeded | Wait for an in-flight request to finish, or reduce parallel requests, before retrying |\n| No signup page found | Expected — KosCompute publishes no public signup or key-management page | Obtain a key out of band; there is no self-serve console to link to |","hierarchy":{"lvl0":"Getting Started","lvl1":"KosCompute Provider Guide","lvl2":"Troubleshooting","lvl3":""}},
6247
6269
  {"objectID":"2de12b046642ad2cef71f2aa84ab6de9163a41b51a046794c2d34738958d382d","title":"See also","url":"/docs/getting-started/providers/koscompute#see-also","content":"Provider setup overview\nAll providers\nTier-2 onboarding — how this provider's JSON becomes a working integration\nProvider feature compatibility","hierarchy":{"lvl0":"Getting Started","lvl1":"KosCompute Provider Guide","lvl2":"See also","lvl3":""}},
6248
- {"objectID":"6d919d5032d8bf60b8341d6b8ff90eeb0c7d2f87553ca6e6e96d08b988c7e6e9","title":"Laya Provider Guide","url":"/docs/getting-started/providers/laya","content":"Laya Provider Guide\n\nA provider of decide — the same typed boolean / choice / score\nanswers as TypeSafe's Jev, from an open-weights model you can\nrun yourself. It emits no text at all.\n\nOverview\n\nLaya is Convai Innovations' Apache-2.0 \"System One\" decision model. You send\none state plus named, typed questions; an encoder answers every question in a\nsingle forward pass and returns a probability distribution for each. NeuroLink\ncalls whatever base URL you configure — a Laya server, or a LiteLLM proxy with a\npass-through route to one. There is no built-in endpoint.\n\nKey Facts\n\n| | |\n| --------------------- | ------------------------------------------------------------------------------------------------- |\n| Inference type | decide only |\n| Checkpoints | typed-decisions (default, fine-tuned for typed decisions), multilingual, english, or auto |\n| Input window | 1,024 tokens (typed-decisions, multilingual); 512 (english) |\n| Questions per request | up to 64 |\n| Endpoint | <base URL>/predict; the base URL is required, with no default |\n| Precedence | used automatically when its key and base URL are set and no TypeSafe key is |\n\nQuick Start\nGet an endpoint and a key\n\nRun a Laya server (see below), or use a LiteLLM\nproxy with a pass-through route to one. On LiteLLM, create a virtual key and add\nthe route's /predict path (and /health, if you want to probe it) to the\nkey's Allowed Routes; a key without them is refused with a 403\nKey/team not allowed to access passthrough route.\nConfigure\n\nBoth the base URL and the key are required. Set them in the environment:\n\nor in the config passed to the SDK, exactly as for any other provider. Values\npassed per call override the constructor's, which override the environment:\nUse it\n\nFrom the CLI\n\nWhen NeuroLink uses it\n\nEvery built-in consumer of decide — model routing, relevance-driven\ncompaction, tool routing and RAG planning — asks for the default decision\nprovider. That is the first one that is configured — in the environment or in\nthe credentials passed to the SDK — in the order TypeSafe, Laya,\nXOR, Perplexity. TypeSafe has two keys,\nTYPESAFE_API_KEY and AI_GATEWAY_API_KEY (its Vercel AI Gateway route), and\neither one counts. Laya counts only with both its key and its base URL, and so\ndoes XOR. Perplexity counts with its key alone, and that key,\nPERPLEXITY_API_KEY, is shared with Perplexity's text provider. So:\nA TypeSafe key, plus Laya's key and base URL: built-in features use\n TypeSafe; Laya runs only where a caller asks for provider: \"laya\".\nOnly Laya's key and base URL: built-in features use Laya.\nLaya's key and base URL, plus XOR's: built-in features use Laya; XOR runs\n only where a caller asks for provider: \"xor\".\nA Laya key with no base URL: Laya is not configured. Built-in features\n ignore it, and provider: \"laya\" fails with Laya requires a base URL.\nLaya's key and base URL, plus a Perplexity key: built-in features use\n Laya; Perplexity runs only where a caller asks for\n provider: \"perplexity-decider\".\nNone of TypeSafe, Laya, XOR or Perplexity: everything behaves exactly as\n it did without a decision model.\n\nLimits\n\nThe window is small. Laya's encoders read 1,024 tokens, and 512 on the\nenglish checkpoint, against roughly 33,000 for Jev. Part of that is reserved\nfor each question and its options, so NeuroLink allows about 768 tokens of\nstate on typed-decisions and multilingual, and 320 on english, auto and\nany model name it does not recognise. Laya's server does not refuse a longer\nstate — it answers from the start of it and says nothing — so NeuroLink\nrefuses it instead, before any network call, with max_tokens_exceeded.\n\nThe size is an estimate, not Laya's tokenizer: about four characters per token\nfor ASCII text, and 1.5 tokens per character for other scripts (0.6 on\nmultilingual), calibrated against a live Laya 0.3.5 server. It errs toward refusing. Built-in\nconsumers treat a refusal as \"carry on as before\", which means long-prompt\nmodel routing usually falls back to the heuristic when Laya is the only\ndecision provider.\n\nMany options degrade it. A choice's options share a fixed token budget,\nso accuracy drops past about 20 options; some Laya servers also reject a\nquestion whose options overflow the budget, as invalid_request.\n\nPick the checkpoint deliberately. Laya's own benchmark puts its general\nenglish and multilingual checkpoints close to chance on typed decisions\nwithout fine-tuning; typed-decisions is the fine-tuned one, which is why it is\nthe default.\n\nErrors\n\n| Reply | Kind ","hierarchy":{"lvl0":"Getting Started","lvl1":"Laya Provider Guide","lvl2":"","lvl3":""}},
6270
+ {"objectID":"6d919d5032d8bf60b8341d6b8ff90eeb0c7d2f87553ca6e6e96d08b988c7e6e9","title":"Laya Provider Guide","url":"/docs/getting-started/providers/laya","content":"Laya Provider Guide\n\nA provider of decide — the same typed boolean / choice / score\nanswers as TypeSafe's Jev, from an open-weights model you can\nrun yourself. It emits no text at all.\n\nOverview\n\nLaya is Convai Innovations' Apache-2.0 \"System One\" decision model. You send\none state plus named, typed questions; an encoder answers every question in a\nsingle forward pass and returns a probability distribution for each. NeuroLink\ncalls whatever base URL you configure — a Laya server, or a LiteLLM proxy with a\npass-through route to one. There is no built-in endpoint.\n\nKey Facts\n\n| | |\n| --------------------- | ------------------------------------------------------------------------------------------------- |\n| Inference type | decide only |\n| Checkpoints | typed-decisions (default, fine-tuned for typed decisions), multilingual, english, or auto |\n| Input window | 1,024 tokens (typed-decisions, multilingual); 512 (english) |\n| Questions per request | up to 64 |\n| Endpoint | <base URL>/predict; the base URL is required, with no default |\n| Precedence | used automatically when its key and base URL are set and no TypeSafe key is |\n\nQuick Start\nGet an endpoint and a key\n\nRun a Laya server (see below), or use a LiteLLM\nproxy with a pass-through route to one. On LiteLLM, create a virtual key and add\nthe route's /predict path (and /health, if you want to probe it) to the\nkey's Allowed Routes; a key without them is refused with a 403\nKey/team not allowed to access passthrough route.\nConfigure\n\nBoth the base URL and the key are required. Set them in the environment:\n\nor in the config passed to the SDK, exactly as for any other provider. Values\npassed per call override the constructor's, which override the environment:\nUse it\n\nFrom the CLI\n\nWhen NeuroLink uses it\n\nEvery built-in consumer of decide — model routing, relevance-driven\ncompaction, tool routing and RAG planning — asks for the default decision\nprovider. That is the first one that is configured — in the environment or in\nthe credentials passed to the SDK — in the order TypeSafe, Laya,\nXOR, Perplexity, then\nCloudflare Clef. TypeSafe has two keys,\nTYPESAFE_API_KEY and AI_GATEWAY_API_KEY (its Vercel AI Gateway route), and\neither one counts. Laya counts only with both its key and its base URL, and so\ndoes XOR. Perplexity counts with its key alone, and that key,\nPERPLEXITY_API_KEY, is shared with Perplexity's text provider. Cloudflare Clef\ncounts only with both CLOUDFLARE_API_KEY and CLOUDFLARE_ACCOUNT_ID, the same\ntwo variables the Workers AI text provider reads. So:\nA TypeSafe key, plus Laya's key and base URL: built-in features use\n TypeSafe; Laya runs only where a caller asks for provider: \"laya\".\nOnly Laya's key and base URL: built-in features use Laya.\nLaya's key and base URL, plus XOR's: built-in features use Laya; XOR runs\n only where a caller asks for provider: \"xor\".\nA Laya key with no base URL: Laya is not configured. Built-in features\n ignore it, and provider: \"laya\" fails with Laya requires a base URL.\nLaya's key and base URL, plus a Perplexity key: built-in features use\n Laya; Perplexity runs only where a caller asks for\n provider: \"perplexity-decider\".\nLaya's key and base URL, plus CLOUDFLARE_API_KEY and\n CLOUDFLARE_ACCOUNT_ID: built-in features use Laya; Cloudflare Clef runs\n only where a caller asks for provider: \"cloudflare-clef\".\nNone of TypeSafe, Laya, XOR, Perplexity or Cloudflare Clef: everything\n behaves exactly as it did without a decision model.\n\nLimits\n\nThe window is small. Laya's encoders read 1,024 tokens, and 512 on the\nenglish checkpoint, against roughly 33,000 for Jev. Part of that is reserved\nfor each question and its options, so NeuroLink allows about 768 tokens of\nstate on typed-decisions and multilingual, and 320 on english, auto and\nany model name it does not recognise. Laya's server does not refuse a longer\nstate — it answers from the start of it and says nothing — so NeuroLink\nrefuses it instead, before any network call, with max_tokens_exceeded.\n\nThe size is an estimate, not Laya's tokenizer: about four characters per token\nfor ASCII text, and 1.5 tokens per character for other scripts (0.6 on\nmultilingual), calibrated against a live Laya 0.3.5 server. It errs toward refusing. Built-in\nconsumers treat a refusal as \"carry on as before\", which means long-prompt\nmodel routing usually falls back to the heuristic when Laya is the only\ndecision provider.\n\nMany options degrade it. A choice's options share a fixed token budget,\nso accuracy drops past about 20 options; some Laya servers also reject a\nquestion","hierarchy":{"lvl0":"Getting Started","lvl1":"Laya Provider Guide","lvl2":"","lvl3":""}},
6249
6271
  {"objectID":"c57a78b27b2b64f684ee60c36022d5b19a81323647a32687608b35530b4d3ed7","title":"Laya Provider Guide","url":"/docs/getting-started/providers/laya#laya-provider-guide","content":"A provider of decide — the same typed boolean / choice / score\nanswers as TypeSafe's Jev, from an open-weights model you can\nrun yourself. It emits no text at all.","hierarchy":{"lvl0":"Getting Started","lvl1":"Laya Provider Guide","lvl2":"Laya Provider Guide","lvl3":""}},
6250
6272
  {"objectID":"969ae83c9e8333da8d49838b0595db3588d8828b67611944cd46f08435f19d97","title":"Overview","url":"/docs/getting-started/providers/laya#overview","content":"Laya is Convai Innovations' Apache-2.0 \"System One\" decision model. You send\none state plus named, typed questions; an encoder answers every question in a\nsingle forward pass and returns a probability distribution for each. NeuroLink\ncalls whatever base URL you configure — a Laya server, or a LiteLLM proxy with a\npass-through route to one. There is no built-in endpoint.","hierarchy":{"lvl0":"Getting Started","lvl1":"Laya Provider Guide","lvl2":"Overview","lvl3":""}},
6251
6273
  {"objectID":"acf4d054e73fb41ef18db1c578cf1beb1117c3f43b24c6743ba30a4fe6569fb7","title":"Key Facts","url":"/docs/getting-started/providers/laya#key-facts","content":"| | |\n| --------------------- | ------------------------------------------------------------------------------------------------- |\n| Inference type | decide only |\n| Checkpoints | typed-decisions (default, fine-tuned for typed decisions), multilingual, english, or auto |\n| Input window | 1,024 tokens (typed-decisions, multilingual); 512 (english) |\n| Questions per request | up to 64 |\n| Endpoint | <base URL>/predict; the base URL is required, with no default |\n| Precedence | used automatically when its key and base URL are set and no TypeSafe key is |","hierarchy":{"lvl0":"Getting Started","lvl1":"Laya Provider Guide","lvl2":"Key Facts","lvl3":""}},
6252
6274
  {"objectID":"ae5970b1821e7485e9d952384e4d6749a99d486ea648241bcbd8a9d43579e79a","title":"1. Get an endpoint and a key","url":"/docs/getting-started/providers/laya#1-get-an-endpoint-and-a-key","content":"Run a Laya server (see below), or use a LiteLLM\nproxy with a pass-through route to one. On LiteLLM, create a virtual key and add\nthe route's /predict path (and /health, if you want to probe it) to the\nkey's Allowed Routes; a key without them is refused with a 403\nKey/team not allowed to access passthrough route.","hierarchy":{"lvl0":"Getting Started","lvl1":"Laya Provider Guide","lvl2":"1. Get an endpoint and a key","lvl3":""}},
6253
6275
  {"objectID":"662bbcdcfe03b7cb5abd93bb0088ff4bdcd8558ad74a79d70f1f7ab36023842c","title":"2. Configure","url":"/docs/getting-started/providers/laya#2-configure","content":"Both the base URL and the key are required. Set them in the environment:\n\nor in the config passed to the SDK, exactly as for any other provider. Values\npassed per call override the constructor's, which override the environment:","hierarchy":{"lvl0":"Getting Started","lvl1":"Laya Provider Guide","lvl2":"2. Configure","lvl3":""}},
6254
- {"objectID":"c99b86a6719e50f022a261b2688e6c849e03afaed0e6e398e9353d84cbf76835","title":"When NeuroLink uses it","url":"/docs/getting-started/providers/laya#when-neurolink-uses-it","content":"Every built-in consumer of decide — model routing, relevance-driven\ncompaction, tool routing and RAG planning — asks for the default decision\nprovider. That is the first one that is configured — in the environment or in\nthe credentials passed to the SDK — in the order TypeSafe, Laya,\nXOR, Perplexity. TypeSafe has two keys,\nTYPESAFE_API_KEY and AI_GATEWAY_API_KEY (its Vercel AI Gateway route), and\neither one counts. Laya counts only with both its key and its base URL, and so\ndoes XOR. Perplexity counts with its key alone, and that key,\nPERPLEXITY_API_KEY, is shared with Perplexity's text provider. So:\nA TypeSafe key, plus Laya's key and base URL: built-in features use\n TypeSafe; Laya runs only where a caller asks for provider: \"laya\".\nOnly Laya's key and base URL: built-in features use Laya.\nLaya's key and base URL, plus XOR's: built-in features use Laya; XOR runs\n only where a caller asks for provider: \"xor\".\nA Laya key with no base URL: Laya is not configured. Built-in features\n ignore it, and provider: \"laya\" fails with Laya requires a base URL.\nLaya's key and base URL, plus a Perplexity key: built-in features use\n Laya; Perplexity runs only where a caller asks for\n provider: \"perplexity-decider\".\nNone of TypeSafe, Laya, XOR or Perplexity: everything behaves exactly as\n it did without a decision model.","hierarchy":{"lvl0":"Getting Started","lvl1":"Laya Provider Guide","lvl2":"When NeuroLink uses it","lvl3":""}},
6276
+ {"objectID":"c99b86a6719e50f022a261b2688e6c849e03afaed0e6e398e9353d84cbf76835","title":"When NeuroLink uses it","url":"/docs/getting-started/providers/laya#when-neurolink-uses-it","content":"Every built-in consumer of decide — model routing, relevance-driven\ncompaction, tool routing and RAG planning — asks for the default decision\nprovider. That is the first one that is configured — in the environment or in\nthe credentials passed to the SDK — in the order TypeSafe, Laya,\nXOR, Perplexity, then\nCloudflare Clef. TypeSafe has two keys,\nTYPESAFE_API_KEY and AI_GATEWAY_API_KEY (its Vercel AI Gateway route), and\neither one counts. Laya counts only with both its key and its base URL, and so\ndoes XOR. Perplexity counts with its key alone, and that key,\nPERPLEXITY_API_KEY, is shared with Perplexity's text provider. Cloudflare Clef\ncounts only with both CLOUDFLARE_API_KEY and CLOUDFLARE_ACCOUNT_ID, the same\ntwo variables the Workers AI text provider reads. So:\nA TypeSafe key, plus Laya's key and base URL: built-in features use\n TypeSafe; Laya runs only where a caller asks for provider: \"laya\".\nOnly Laya's key and base URL: built-in features use Laya.\nLaya's key and base URL, plus XOR's: built-in features use Laya; XOR runs\n only where a caller asks for provider: \"xor\".\nA Laya key with no base URL: Laya is not configured. Built-in features\n ignore it, and provider: \"laya\" fails with Laya requires a base URL.\nLaya's key and base URL, plus a Perplexity key: built-in features use\n Laya; Perplexity runs only where a caller asks for\n provider: \"perplexity-decider\".\nLaya's key and base URL, plus CLOUDFLARE_API_KEY and\n CLOUDFLARE_ACCOUNT_ID: built-in features use Laya; Cloudflare Clef runs\n only where a caller asks for provider: \"cloudflare-clef\".\nNone of TypeSafe, Laya, XOR, Perplexity or Cloudflare Clef: everything\n behaves exactly as it did without a decision model.","hierarchy":{"lvl0":"Getting Started","lvl1":"Laya Provider Guide","lvl2":"When NeuroLink uses it","lvl3":""}},
6255
6277
  {"objectID":"adbc69b418b48f0bcaa2ae6e9f87f6bf8efc0dea1fe0d614cc12973156148d9f","title":"Limits","url":"/docs/getting-started/providers/laya#limits","content":"The window is small. Laya's encoders read 1,024 tokens, and 512 on the\nenglish checkpoint, against roughly 33,000 for Jev. Part of that is reserved\nfor each question and its options, so NeuroLink allows about 768 tokens of\nstate on typed-decisions and multilingual, and 320 on english, auto and\nany model name it does not recognise. Laya's server does not refuse a longer\nstate — it answers from the start of it and says nothing — so NeuroLink\nrefuses it instead, before any network call, with max_tokens_exceeded.\n\nThe size is an estimate, not Laya's tokenizer: about four characters per token\nfor ASCII text, and 1.5 tokens per character for other scripts (0.6 on\nmultilingual), calibrated against a live Laya 0.3.5 server. It errs toward refusing. Built-in\nconsumers treat a refusal as \"carry on as before\", which means long-prompt\nmodel routing usually falls back to the heuristic when Laya is the only\ndecision provider.\n\nMany options degrade it. A choice's options share a fixed token budget,\nso accuracy drops past about 20 options; some Laya servers also reject a\nquestion whose options overflow the budget, as invalid_request.\n\nPick the checkpoint deliberately. Laya's own benchmark puts its general\nenglish and multilingual checkpoints close to chance on typed decisions\nwithout fine-tuning; typed-decisions is the fine-tuned one, which is why it is\nthe default.","hierarchy":{"lvl0":"Getting Started","lvl1":"Laya Provider Guide","lvl2":"Limits","lvl3":""}},
6256
6278
  {"objectID":"141537feff8c79df3ecbca258b80136e36da0fdeff342aa93a971be020313215","title":"Errors","url":"/docs/getting-started/providers/laya#errors","content":"| Reply | Kind | Retried |\n| ----------------------------------------------- | ------------------------------------------------------- | ------- |\n| 401 / 403 from the proxy | authentication — the provider instance stops retrying | no |\n| 400 / 422 from Laya | invalid_request, with Laya's reason | no |\n| 413 from Laya | max_tokens_exceeded | no |\n| 429 | rate_limit | yes |\n| 503 | overloaded | yes |\n| other 5xx | server | yes |\n| more questions or state than Laya reads (local) | max_tokens_exceeded, with no network call | no |\n| no base URL configured (local) | invalid_request, with no network call | no |\n\nLiteLLM's 401 text echoes a masked copy of the rejected key and its hash; the\nprovider drops that part before the message reaches a log or an error.","hierarchy":{"lvl0":"Getting Started","lvl1":"Laya Provider Guide","lvl2":"Errors","lvl3":""}},
6257
6279
  {"objectID":"de358abc64cad163c1666dce7b06fc5f7b3637b9568e420ceb9269a526929d54","title":"Running your own Laya server","url":"/docs/getting-started/providers/laya#running-your-own-laya-server","content":"Laya's example server exposes the same /predict route. Point\nLAYA_BASE_URL at it:","hierarchy":{"lvl0":"Getting Started","lvl1":"Laya Provider Guide","lvl2":"Running your own Laya server","lvl3":""}},
6258
6280
  {"objectID":"e0dd67e8f79c7801d3e199dcd765a559f64a9d65d43cf1611026725a8ae1ee82","title":"Troubleshooting","url":"/docs/getting-started/providers/laya#troubleshooting","content":"Laya requires a base URL — set LAYA_BASE_URL, or pass\n credentials.laya.baseURL. There is no default endpoint.\nAuthentication failed — the proxy rejected the key. Check it in the\n LiteLLM Dashboard, including its Allowed Routes. The provider instance does\n not retry after a rejection.\nA 403 with error code: 1010 — Cloudflare in front of your proxy blocked\n the client by its signature. Node's fetch is accepted; a proxy or agent that\n rewrites the user agent may not be.\nexceeded the provider's token limit — the state is larger than Laya\n reads. Shorten it, or configure TypeSafe for long inputs.\nBuilt-in routing never uses Laya — a TypeSafe key is also set and takes\n precedence, or Laya has no base URL.","hierarchy":{"lvl0":"Getting Started","lvl1":"Laya Provider Guide","lvl2":"Troubleshooting","lvl3":""}},
6259
- {"objectID":"fb160b657869a262fd7735599ce2c2ffd49423f4a0f7c43f4ef5ee74aa4dcde5","title":"See also","url":"/docs/getting-started/providers/laya#see-also","content":"The decide inference type\nTypeSafe (Jev) Provider Guide\nXOR Provider Guide\nPerplexity Decisions Provider Guide\nLaya on GitHub","hierarchy":{"lvl0":"Getting Started","lvl1":"Laya Provider Guide","lvl2":"See also","lvl3":""}},
6281
+ {"objectID":"fb160b657869a262fd7735599ce2c2ffd49423f4a0f7c43f4ef5ee74aa4dcde5","title":"See also","url":"/docs/getting-started/providers/laya#see-also","content":"The decide inference type\nTypeSafe (Jev) Provider Guide\nXOR Provider Guide\nPerplexity Decisions Provider Guide\nCloudflare Clef Provider Guide\nLaya on GitHub","hierarchy":{"lvl0":"Getting Started","lvl1":"Laya Provider Guide","lvl2":"See also","lvl3":""}},
6260
6282
  {"objectID":"df16212357acdf1cff3460cb04e276abe1c702cc845767f49a4bc0cd2a92a763","title":"Lemonfox AI Provider Guide","url":"/docs/getting-started/providers/lemonfox-ai","content":"Lemonfox AI Provider Guide\n\nLemonfox AI is a Tier-2 catalog provider: its integration is one JSON file\n(src/lib/providers/catalog/lemonfox-ai.json) rather than hand-written code.\n\nVerification status: this entry is docs-verified only, not yet\nlive-verified. The vendor's GET https://api.lemonfox.ai/v1/models needs a\nkey (it answers HTTP 401 without one), so the two model ids come from the\nvendor's public chat API page, https://www.lemonfox.ai/apis/chat (retrieved\n2026-09-29), and the roster has not been checked. No account was created and no\nAPI key was used to build it. evidence.liveMatrix is null until someone\nruns the live capability matrix with a real key (see\nVerification status below).\n\nKey Facts\nProvider id: lemonfox-ai (alias lemonfox)\nProtocol: OpenAI-compatible (/chat/completions) — the chat API page says\n \"Our OpenAI-compatible API takes a list of messages as input and provides an\n AI-generated (assistant) message as output.\"\nBase URL: https://api.lemonfox.ai/v1\nDefault model: deepseek-v4-flash\nModels in catalog: 2, taken from the model table of\n https://www.lemonfox.ai/apis/chat\nStreaming: supported — the page's stream parameter reads \"incremental\n message updates are transmitted as server-sent events with data-only\n messages.\"\nTool calling: not declared (false) — the page's API Parameters section\n lists messages, model, max_tokens, stop, stream,\n frequency_penalty, presence_penalty, temperature and top_p\nTools while streaming: not declared (false)\nStructured output: not declared (false), same parameter list as above\nStructured output + tools together: not declared (false) — no combined\n probe was possible without credentials\nEmbeddings: not declared\nThinking: not declared\nBilling: free-with-card is the growth queue's classification, one of the\n three values the catalog schema allows; see Billing below for the\n vendor's own wording\nKey format: none declared\n\nQuick Start\nGet an API key\nVisit: https://lemonfox.ai/signup (the \"Get started\" and \"Start Your Free Trial\" links on https://www.lemonfox.ai/apis/chat both point to /signup); sign-up was not attempted while building this entry\nCreate an API key at https://lemonfox.ai/apis/keys — https://www.lemonfox.ai/apis/chat links \"create an API key\" to that page, and the API's 401 reply says \"You can obtain an API key from https://lemonfox.ai/apis/keys.\"\nBilling as the vendor states it: see Billing below\nSet LEMONFOX_AI_API_KEY in your .env file\nConfigure\nUse it\n\nPer-request credentials:\n\nBilling\n\nWhat the vendor's public pages show (retrieved 2026-09-29):\nhttps://www.lemonfox.ai/ — \"Try our APIs for 1 month for free.\"\nhttps://www.lemonfox.ai/apis/chat — \"First month for free!\" and \"Start Your\n Free Trial\"\nhttps://www.lemonfox.ai/terms — Lemon Fox GmbH uses Paddle as its Merchant of\n Record and offers \"subscription-based pricing and usage-based pricing\"\n\nThe entry records free-with-card, the growth queue's classification, because\nthe catalog schema has no unknown value; it is not a vendor statement of\nsign-up terms.\n\nModels\n\n| Model | Context | Price / 1M tokens | Notes |\n| ---------------------- | ------- | ----------------- | -------------------------------------------------------------------------------------------------------------------------------- |\n| deepseek-v4-flash ⭐ | 1M | $0.50 | NeuroLink default. Vendor: \"DeepSeek V4 Flash is our default model — fast, cost-effective, and suitable for most tasks.\" |\n| mimo-v2.5-pro | — | $1.25 | NeuroLink fallback. Vendor: \"MiMo V2.5 Pro is a powerful model from Xiaomi, recommended for complex reasoning and coding tasks.\" |\n\nModel ids, the 1M context figure and the two prices come from the model table\non https://www.lemonfox.ai/apis/chat (retrieved 2026-09-29), whose price column\nis headed \"Price / 1M tokens\". The table's context text for deepseek-v4-flash\nreads \"Supports a 1M token context window.\"; the catalog stores it as 1,000,000\n(NeuroLink's decimal reading of \"1M\"). mimo-v2.5-pro has no context window set.\nThe page also shows the model parameter as \"Specify the model ID to be used.\nDefault: deepseek-v4-flash. Also supported: mimo-v2.5-pro.\"\n\nmodels.defaultContextWindow (32,768) and models.defaultMaxOutputTokens\n(4,096) are placeholders the vendor does not publish, for model ids outside the\ncatalog; the page's max_tokens parameter reads \"integer, optional, default:\ninfinity\". pricingPerMTok is unset on both models because the catalog field\ntakes separate input and output figures and the table has the one \"Price / 1M\ntokens\" column.\n\nEach model's vision flag is false in the catalog, a placeholder for a\nboolean the schema requires; it is not a vendor statement. Each model's\nstatus is production, a value the schema requires.\n\nDeprecated models: a second table on the same page, under the heading\n\"Deprecated","hierarchy":{"lvl0":"Getting Started","lvl1":"Lemonfox AI Provider Guide","lvl2":"","lvl3":""}},
6261
6283
  {"objectID":"3fca5359ef393150cf2b043c562fa6d28cdd3d17be9b09b089efaf425beb7ef6","title":"Lemonfox AI Provider Guide","url":"/docs/getting-started/providers/lemonfox-ai#lemonfox-ai-provider-guide","content":"Lemonfox AI is a Tier-2 catalog provider: its integration is one JSON file\n(src/lib/providers/catalog/lemonfox-ai.json) rather than hand-written code.\n\nVerification status: this entry is docs-verified only, not yet\nlive-verified. The vendor's GET https://api.lemonfox.ai/v1/models needs a\nkey (it answers HTTP 401 without one), so the two model ids come from the\nvendor's public chat API page, https://www.lemonfox.ai/apis/chat (retrieved\n2026-09-29), and the roster has not been checked. No account was created and no\nAPI key was used to build it. evidence.liveMatrix is null until someone\nruns the live capability matrix with a real key (see\nVerification status below).","hierarchy":{"lvl0":"Getting Started","lvl1":"Lemonfox AI Provider Guide","lvl2":"Lemonfox AI Provider Guide","lvl3":""}},
6262
6284
  {"objectID":"58c275813c578b02ad4927ca1b96bc8b37a07c75c8829608a8098f8fd53256f1","title":"Key Facts","url":"/docs/getting-started/providers/lemonfox-ai#key-facts","content":"Provider id: lemonfox-ai (alias lemonfox)\nProtocol: OpenAI-compatible (/chat/completions) — the chat API page says\n \"Our OpenAI-compatible API takes a list of messages as input and provides an\n AI-generated (assistant) message as output.\"\nBase URL: https://api.lemonfox.ai/v1\nDefault model: deepseek-v4-flash\nModels in catalog: 2, taken from the model table of\n https://www.lemonfox.ai/apis/chat\nStreaming: supported — the page's stream parameter reads \"incremental\n message updates are transmitted as server-sent events with data-only\n messages.\"\nTool calling: not declared (false) — the page's API Parameters section\n lists messages, model, max_tokens, stop, stream,\n frequency_penalty, presence_penalty, temperature and top_p\nTools while streaming: not declared (false)\nStructured output: not declared (false), same parameter list as above\nStructured output + tools together: not declared (false) — no combined\n probe was possible without credentials\nEmbeddings: not declared\nThinking: not declared\nBilling: free-with-card is the growth queue's classification, one of the\n three values the catalog schema allows; see Billing below for the\n vendor's own wording\nKey format: none declared","hierarchy":{"lvl0":"Getting Started","lvl1":"Lemonfox AI Provider Guide","lvl2":"Key Facts","lvl3":""}},
@@ -6876,16 +6898,16 @@
6876
6898
  {"objectID":"106d04b8baec018f872ce608c3b0dd33f8b3d43ada85072831a2e9ab9c64c46b","title":"Verification status","url":"/docs/getting-started/providers/pareto-inference#verification-status","content":"Tier-2 onboarding requires evidence before a provider is accepted, and\npnpm run verify:provider-onboarding gates it in CI. This is what the\ncatalog records for Pareto Inference — and, just as importantly, what it\ndoes not yet record:\n\n| Probe | Result |\n| --------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |\n| Roster | unauthenticated GET /v1/models, HTTP 200, 1 model (z-ai/glm-5.3-flash), 2026-09-28. No API key was used or required for this call. |\n| Auth rejection | Not probed. Would require a POST with an invalid key, which this credential-free onboarding pass does not send. Pareto's docs describe 401 as \"the chat API key is missing, invalid, or no longer active,\" with no example error body. |\n| Live capability sweep | Not run. evidence.liveMatrix is null.","hierarchy":{"lvl0":"Getting Started","lvl1":"Pareto Inference Provider Guide","lvl2":"Verification status","lvl3":""}},
6877
6899
  {"objectID":"b844edaf09c93fb3089bbf0a71254c5039dedf111e2b59563321484496862aa4","title":"Troubleshooting","url":"/docs/getting-started/providers/pareto-inference#troubleshooting","content":"| Symptom | Cause | Fix |\n| --------------------------------------- | -------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- |\n| Invalid Pareto Inference API key | PARETO_INFERENCE_API_KEY unset, wrong, or expired | Pareto documents 401 as \"the chat API key is missing, invalid, or no longer active\" — check or replace the key in the dashboard |\n| 429 with credit_exhausted | Prepaid balance is empty | Buy more prepaid credits |\n| 429 with credit_insufficient | Held credits (reserved for max_tokens) are below what's needed | Lower max_tokens, or buy more credits |\n| 429 with model_capacity | The model itself is at capacity — not an account-level limit | Honor the Retry-After header and retry |\n| 502 mid-stream | The model request failed after the stream started | Retry the request; check whether a tool call already ran before retrying |\n| 503 | Temporary Pareto-side failure (e.g.","hierarchy":{"lvl0":"Getting Started","lvl1":"Pareto Inference Provider Guide","lvl2":"Troubleshooting","lvl3":""}},
6878
6900
  {"objectID":"57b8457eff592f229d3c0f0fb7c73148341ecffd72bb29c15d73f1471b1fd886","title":"See also","url":"/docs/getting-started/providers/pareto-inference#see-also","content":"Provider setup overview\nAll providers\nTier-2 onboarding — how this provider's JSON becomes a working integration\nProvider feature compatibility","hierarchy":{"lvl0":"Getting Started","lvl1":"Pareto Inference Provider Guide","lvl2":"See also","lvl3":""}},
6879
- {"objectID":"daf00919862e4bbed4b3fc1082bea17843a987e5e3970cefed2c95537b22cceb","title":"Perplexity Decisions Provider Guide","url":"/docs/getting-started/providers/perplexity-decider","content":"Perplexity Decisions Provider Guide\n\nA provider of decide — the same typed boolean / choice / score\nanswers as TypeSafe's Jev, Laya and XOR,\nfrom Perplexity's hosted Decisions API, and it also reads images. It emits no\ntext at all.\n\nThis is not the Perplexity text provider (perplexity, the\nSonar models), which serves generate() and stream(). The two share one API\nkey, which has a consequence worth reading before you set it. See\nWhat is sent to Perplexity and\nOne key, two providers.\n\nOverview\n\npplx-decider-v1-27b is Perplexity's decision model. You send one state\nplus named, typed questions, and optionally images. The model answers every\nquestion in a single batched pass and returns a typed answer for each, with a\nprobability instead of a sentence. NeuroLink calls Perplexity's public endpoint,\nso a key alone configures it: there is no base URL to set.\n\nKey Facts\n\n| | |\n| ------------------------ | ----------------------------------------------------------------------------------------------------------------- |\n| Provider id | perplexity-decider (no aliases) — distinct from perplexity, the Sonar text provider |\n| Inference type | decide only |\n| Model | pplx-decider-v1-27b; the API answers 400 to a missing or unknown model. PERPLEXITY_DECIDER_MODEL sets it |\n| Key | PERPLEXITY_API_KEY, shared with the Perplexity text provider, or credentials.perplexityDecider.apiKey |\n| Endpoint | https://api.perplexity.ai/v1/decisions; PERPLEXITY_DECIDER_BASE_URL can name another origin |\n| Media | images (PNG, JPEG or WebP), up to 8 per request; no video |\n| Questions per request | up to 128 |\n| Server input ceiling | under 262,144 tokens (state, questions and images); more is refused with an explicit 400, never cut off silently |\n| Images and that ceiling | billed as input tokens, and counted toward the ceiling at one token per 32 Ɨ 32 tile (measured) |\n| NeuroLink's state window | 100,000 estimated tokens: a deliberate local limit, not the server's |\n| Cost | $0.04 per million input tokens (image tokens included); output tokens are free. Perplexity's documented price |\n| Default timeout | 10 seconds plus 100 ms for each question (10.1 s for one, 22.8 s for 128); timeoutMs or --timeout replaces it |\n| Precedence | used automatically when TypeSafe, Laya and XOR are not configured |\n\nEach limit, token rate and latency in this guide is either measured on a real\naccount in October 2026 or taken from Perplexity's documentation, and says\nwhich where it appears. Measured: the 128-question and 8-image caps, the\n262,144-token input ceiling for the state, the questions and the images (nothing\nis cut off silently), characters per token for each kind of text, latency by\ninput size and by question count, what an image costs, which image sizes stall\n(the 2,048-tile rule, checked at its edge: see Images), and how the\nAPI answers a burst of requests. Documented and not tested: the 32 MiB request\nbody, the price, and the option and level counts. Perplexity documents a limit of\n10 requests per second for the account's tier (every organization, on every\nplan, with a token limit on large bursts), and a burst test on the one account\ntested saw the request limit act.\n\nQuick Start\nGet a key\n\nCreate one in the Perplexity console. Any\nPerplexity API key works.\nConfigure\n\nSet the key in the environment:\n\nor in the config passed to the SDK, exactly as for any other provider. Values\npassed per call override the constructor's, which override the environment:\n\ncredentials.perplexityDecider is its own slice. A key passed as\ncredentials.perplexity belongs to the text provider and does not configure\nthis one. PERPLEXITY_BASE_URL, which moves the text provider, is not read here;\nthe override for this provider is PERPLEXITY_DECIDER_BASE_URL, and a trailing\n/v1 on it is accepted.\nUse it\n\nNeuroLink's boolean question is sent to the API as its noul type, and the\nanswer is read back as a boolean. A boolean answer carries a probability and\nno confidence of its own; a choice or score answer carries the confidence the\nAPI reports, which Perplexity describes as the model's own certainty estimate,\nnot the top probability. Perplexity's API reference and quickstart do not call\nthat confidence calibrated, and NeuroLink did not measure whether it is, so tune\na t","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Decisions Provider Guide","lvl2":"","lvl3":""}},
6880
- {"objectID":"cd579839c9ecb7a84dee24ccd591e76f90eb1f05b384a35488c7603f72178d38","title":"Perplexity Decisions Provider Guide","url":"/docs/getting-started/providers/perplexity-decider#perplexity-decisions-provider-guide","content":"A provider of decide — the same typed boolean / choice / score\nanswers as TypeSafe's Jev, Laya and XOR,\nfrom Perplexity's hosted Decisions API, and it also reads images. It emits no\ntext at all.\n\nThis is not the Perplexity text provider (perplexity, the\nSonar models), which serves generate() and stream(). The two share one API\nkey, which has a consequence worth reading before you set it. See\nWhat is sent to Perplexity and\nOne key, two providers.","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Decisions Provider Guide","lvl2":"Perplexity Decisions Provider Guide","lvl3":""}},
6901
+ {"objectID":"daf00919862e4bbed4b3fc1082bea17843a987e5e3970cefed2c95537b22cceb","title":"Perplexity Decisions Provider Guide","url":"/docs/getting-started/providers/perplexity-decider","content":"Perplexity Decisions Provider Guide\n\nA provider of decide — the same typed boolean / choice / score\nanswers as TypeSafe's Jev, Laya, XOR and\nCloudflare Clef, from Perplexity's hosted Decisions API,\nand it also reads images. It emits no text at all.\n\nThis is not the Perplexity text provider (perplexity, the\nSonar models), which serves generate() and stream(). The two share one API\nkey, which has a consequence worth reading before you set it. See\nWhat is sent to Perplexity and\nOne key, two providers.\n\nOverview\n\npplx-decider-v1-27b is Perplexity's decision model. You send one state\nplus named, typed questions, and optionally images. The model answers every\nquestion in a single batched pass and returns a typed answer for each, with a\nprobability instead of a sentence. NeuroLink calls Perplexity's public endpoint,\nso a key alone configures it: there is no base URL to set.\n\nKey Facts\n\n| | |\n| ------------------------ | ----------------------------------------------------------------------------------------------------------------- |\n| Provider id | perplexity-decider (no aliases) — distinct from perplexity, the Sonar text provider |\n| Inference type | decide only |\n| Model | pplx-decider-v1-27b; the API answers 400 to a missing or unknown model. PERPLEXITY_DECIDER_MODEL sets it |\n| Key | PERPLEXITY_API_KEY, shared with the Perplexity text provider, or credentials.perplexityDecider.apiKey |\n| Endpoint | https://api.perplexity.ai/v1/decisions; PERPLEXITY_DECIDER_BASE_URL can name another origin |\n| Media | images (PNG, JPEG or WebP), up to 8 per request; no video |\n| Questions per request | up to 128 |\n| Server input ceiling | under 262,144 tokens (state, questions and images); more is refused with an explicit 400, never cut off silently |\n| Images and that ceiling | billed as input tokens, and counted toward the ceiling at one token per 32 Ɨ 32 tile (measured) |\n| NeuroLink's state window | 100,000 estimated tokens: a deliberate local limit, not the server's |\n| Cost | $0.04 per million input tokens (image tokens included); output tokens are free. Perplexity's documented price |\n| Default timeout | 10 seconds plus 100 ms for each question (10.1 s for one, 22.8 s for 128); timeoutMs or --timeout replaces it |\n| Precedence | used automatically when TypeSafe, Laya and XOR are not configured |\n\nEach limit, token rate and latency in this guide is either measured on a real\naccount in October 2026 or taken from Perplexity's documentation, and says\nwhich where it appears. Measured: the 128-question and 8-image caps, the\n262,144-token input ceiling for the state, the questions and the images (nothing\nis cut off silently), characters per token for each kind of text, latency by\ninput size and by question count, what an image costs, which image sizes stall\n(the 2,048-tile rule, checked at its edge: see Images), and how the\nAPI answers a burst of requests. Documented and not tested: the 32 MiB request\nbody, the price, and the option and level counts. Perplexity documents a limit of\n10 requests per second for the account's tier (every organization, on every\nplan, with a token limit on large bursts), and a burst test on the one account\ntested saw the request limit act.\n\nQuick Start\nGet a key\n\nCreate one in the Perplexity console. Any\nPerplexity API key works.\nConfigure\n\nSet the key in the environment:\n\nor in the config passed to the SDK, exactly as for any other provider. Values\npassed per call override the constructor's, which override the environment:\n\ncredentials.perplexityDecider is its own slice. A key passed as\ncredentials.perplexity belongs to the text provider and does not configure\nthis one. PERPLEXITY_BASE_URL, which moves the text provider, is not read here;\nthe override for this provider is PERPLEXITY_DECIDER_BASE_URL, and a trailing\n/v1 on it is accepted.\nUse it\n\nNeuroLink's boolean question is sent to the API as its noul type, and the\nanswer is read back as a boolean. A boolean answer carries a probability and\nno confidence of its own; a choice or score answer carries the confidence the\nAPI reports, which Perplexity describes as the model's own certainty estimate,\nnot the top probability. Perplexity's API reference and quickstart do not call\nthat confidence calibrated, and NeuroLink did not measure whether i","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Decisions Provider Guide","lvl2":"","lvl3":""}},
6902
+ {"objectID":"cd579839c9ecb7a84dee24ccd591e76f90eb1f05b384a35488c7603f72178d38","title":"Perplexity Decisions Provider Guide","url":"/docs/getting-started/providers/perplexity-decider#perplexity-decisions-provider-guide","content":"A provider of decide — the same typed boolean / choice / score\nanswers as TypeSafe's Jev, Laya, XOR and\nCloudflare Clef, from Perplexity's hosted Decisions API,\nand it also reads images. It emits no text at all.\n\nThis is not the Perplexity text provider (perplexity, the\nSonar models), which serves generate() and stream(). The two share one API\nkey, which has a consequence worth reading before you set it. See\nWhat is sent to Perplexity and\nOne key, two providers.","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Decisions Provider Guide","lvl2":"Perplexity Decisions Provider Guide","lvl3":""}},
6881
6903
  {"objectID":"0d2eec9390e1cc5078e5419138c190cc8dc524c5ebd8972a9c7afa9625dae43a","title":"Overview","url":"/docs/getting-started/providers/perplexity-decider#overview","content":"pplx-decider-v1-27b is Perplexity's decision model. You send one state\nplus named, typed questions, and optionally images. The model answers every\nquestion in a single batched pass and returns a typed answer for each, with a\nprobability instead of a sentence. NeuroLink calls Perplexity's public endpoint,\nso a key alone configures it: there is no base URL to set.","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Decisions Provider Guide","lvl2":"Overview","lvl3":""}},
6882
6904
  {"objectID":"c9690ba724ce4205d66ac4d75981884b735170b152767047afcc0d406d5e7c8d","title":"Key Facts","url":"/docs/getting-started/providers/perplexity-decider#key-facts","content":"| | |\n| ------------------------ | ----------------------------------------------------------------------------------------------------------------- |\n| Provider id | perplexity-decider (no aliases) — distinct from perplexity, the Sonar text provider |\n| Inference type | decide only |\n| Model | pplx-decider-v1-27b; the API answers 400 to a missing or unknown model. PERPLEXITY_DECIDER_MODEL sets it |\n| Key | PERPLEXITY_API_KEY, shared with the Perplexity text provider, or credentials.perplexityDecider.apiKey |\n| Endpoint | https://api.perplexity.ai/v1/decisions; PERPLEXITY_DECIDER_BASE_URL can name another origin |\n| Media | images (PNG, JPEG or WebP), up to 8 per request; no video |\n| Questions per request | up to 128 |\n| Server input ceiling | under 262,144 tokens (state, questions and images); more is refused with an explicit 400, never cut off silently |\n| Images and that ceiling | billed as input tokens, and counted toward the ceiling at one token per 32 Ɨ 32 tile (measured) |\n| NeuroLink's state window | 100,000 estimated tokens: a deliberate local limit, not the server's |\n| Cost | $0.04 per million input tokens (image tokens included); output tokens are free.","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Decisions Provider Guide","lvl2":"Key Facts","lvl3":""}},
6883
6905
  {"objectID":"467fb7efd250abc71afba6a9535ee164e671526dc06ff59407a827bda872b268","title":"1. Get a key","url":"/docs/getting-started/providers/perplexity-decider#1-get-a-key","content":"Create one in the Perplexity console. Any\nPerplexity API key works.","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Decisions Provider Guide","lvl2":"1. Get a key","lvl3":""}},
6884
6906
  {"objectID":"bd7509f708e016023be201690ca4cc963bc77551d7152b34c636ae628ea5e497","title":"2. Configure","url":"/docs/getting-started/providers/perplexity-decider#2-configure","content":"Set the key in the environment:\n\nor in the config passed to the SDK, exactly as for any other provider. Values\npassed per call override the constructor's, which override the environment:\n\ncredentials.perplexityDecider is its own slice. A key passed as\ncredentials.perplexity belongs to the text provider and does not configure\nthis one. PERPLEXITY_BASE_URL, which moves the text provider, is not read here;\nthe override for this provider is PERPLEXITY_DECIDER_BASE_URL, and a trailing\n/v1 on it is accepted.","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Decisions Provider Guide","lvl2":"2. Configure","lvl3":""}},
6885
6907
  {"objectID":"c3dbe3c712926dfb07b0d907370532df8bc3c98a5e81f4b46784a037ac0f1e0b","title":"3. Use it","url":"/docs/getting-started/providers/perplexity-decider#3-use-it","content":"NeuroLink's boolean question is sent to the API as its noul type, and the\nanswer is read back as a boolean. A boolean answer carries a probability and\nno confidence of its own; a choice or score answer carries the confidence the\nAPI reports, which Perplexity describes as the model's own certainty estimate,\nnot the top probability. Perplexity's API reference and quickstart do not call\nthat confidence calibrated, and NeuroLink did not measure whether it is, so tune\na threshold on it against your own labelled data. (The Hugging Face model card\nfor the open-weights model describes its output as calibrated probabilities.)\nPerplexity also notes that identical requests occasionally differ in the second\ndecimal place. Repeating a request returned the same numbers in NeuroLink's\nprobe, but leave some margin when you set a threshold anyway.","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Decisions Provider Guide","lvl2":"3. Use it","lvl3":""}},
6886
6908
  {"objectID":"6f3f0a67d6c85ce65d1d7d95a7016987ad0423b545c9f849bcde853edd0dd20f","title":"From the CLI","url":"/docs/getting-started/providers/perplexity-decider#from-the-cli","content":"The CLI reads the same environment variables:\n\nThe text output ends with a Cost: line, priced from the input tokens the API\nreports.","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Decisions Provider Guide","lvl2":"From the CLI","lvl3":""}},
6887
- {"objectID":"e1b9161659de68a10277e69ce53077f2aed241892d783054f01471fb793024bd","title":"Images","url":"/docs/getting-started/providers/perplexity-decider#images","content":"Perplexity reads images alongside the state. TypeSafe and Laya do not read\nmedia. XOR reads images and a video; Perplexity reads images only. In the probe\nthe model did read an image: asked red versus blue about one, it answered\ncorrectly at about 0.98 confidence.\n\nFrom the CLI, pass --image <path>, repeated for several:\n\nOn the wire the images go inside state: NeuroLink sends state as an array,\nyour state first and then one OpenAI-style image_url part per image, each a\nbase64 data URL. The API has no separate images field and rejects unknown\nfields. An image can also be the whole input (an image-only state was\naccepted in the probe); from the SDK, pass an empty string as state with\nimages. The CLI always needs a state.\n\nThe rules:\nUp to 8 images per request. Measured: 8 images with three questions were\n accepted, and 9 images with one question were refused with a 400 whose message\n words the cap as \"per question\". NeuroLink applies the 8 to the request as a\n whole.\nPNG, JPEG or WebP only. The type comes from the bytes, not the file\n extension. A GIF, or a data: URL of another type, is refused before any\n request, as a non-retryable invalid_request. The API refuses a GIF with a 400\n too.\nEach is a Buffer, a local file path or a data: URL. An http(s) URL is\n refused locally and never fetched. The API does not fetch URLs either: an\n https image URL returned a 400 in the probe.\nEach image may be at most 2,048 tiles of 32 Ɨ 32 pixels. Round the width\n and the height to the nearest multiple of 32 and keep (width / 32) Ɨ\n (height / 32) at or under 2,048: by that count 1440 Ɨ 1440 (2,025 tiles) and\n 2048 Ɨ 1024 (2,048) fit, and 1600 Ɨ 1310 (2,050) does not. Perplexity documents\n this rule, and it was checked at its edge. Fitting sizes were answered in\n about a second: 1920 Ɨ 1080 (2,040 tiles) in 1.1 seconds, 2048 Ɨ 1024 (2,048,\n exactly the cap) in 1.3 and 1450 Ɨ 1450 in 1.1.","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Decisions Provider Guide","lvl2":"Images","lvl3":""}},
6888
- {"objectID":"8ad813ca01099f441bbc2db40ee064d9a0294d269166c42a6f87fce2a1c1330f","title":"When NeuroLink uses it","url":"/docs/getting-started/providers/perplexity-decider#when-neurolink-uses-it","content":"Every built-in consumer of decide asks for the default decision provider. That\nis the first one that is configured, in the environment or in the credentials\npassed to the SDK, in the order TypeSafe, Laya, XOR, Perplexity. TypeSafe counts\nwith either of its keys, TYPESAFE_API_KEY or AI_GATEWAY_API_KEY. Laya and XOR\neach count only with both their key and their base URL. Perplexity counts with\nits key alone, a non-blank PERPLEXITY_API_KEY or\ncredentials.perplexityDecider.apiKey. A caller can always name it with\nprovider: \"perplexity-decider\". So:\nOnly a Perplexity key: built-in features use Perplexity.\nA Perplexity key, plus TypeSafe's key, or Laya's or XOR's key and base\n URL: built-in features use TypeSafe, Laya or XOR, in that order, whichever\n is configured. Perplexity runs only where a caller asks for\n provider: \"perplexity-decider\".\nA key for the text provider only: the same as the first case. The key is\n shared, so the text provider's key also configures this one. See\n One key, two providers.\nNone of TypeSafe, Laya, XOR or Perplexity configured: everything behaves\n exactly as it did without a decision model.\n\nThe classifier router's auto gate reads the credentials given to the\nNeuroLink constructor and the environment, not credentials passed on a single\ncall, so a per-call credentials.perplexityDecider does not influence it. The\nother consumers NeuroLink wires itself (compaction and tool routing) call\ntryDecide without credentials too.","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Decisions Provider Guide","lvl2":"When NeuroLink uses it","lvl3":""}},
6909
+ {"objectID":"e1b9161659de68a10277e69ce53077f2aed241892d783054f01471fb793024bd","title":"Images","url":"/docs/getting-started/providers/perplexity-decider#images","content":"Perplexity reads images alongside the state. TypeSafe and Laya do not read\nmedia. XOR reads images and a video; Perplexity reads images only, and so does\nCloudflare Clef, up to 4 of them. In the probe of\nPerplexity's model it did read an image: asked red versus blue about one, it\nanswered correctly at about 0.98 confidence.\n\nFrom the CLI, pass --image <path>, repeated for several:\n\nOn the wire the images go inside state: NeuroLink sends state as an array,\nyour state first and then one OpenAI-style image_url part per image, each a\nbase64 data URL. The API has no separate images field and rejects unknown\nfields. An image can also be the whole input (an image-only state was\naccepted in the probe); from the SDK, pass an empty string as state with\nimages. The CLI always needs a state.\n\nThe rules:\nUp to 8 images per request. Measured: 8 images with three questions were\n accepted, and 9 images with one question were refused with a 400 whose message\n words the cap as \"per question\". NeuroLink applies the 8 to the request as a\n whole.\nPNG, JPEG or WebP only. The type comes from the bytes, not the file\n extension. A GIF, or a data: URL of another type, is refused before any\n request, as a non-retryable invalid_request. The API refuses a GIF with a 400\n too.\nEach is a Buffer, a local file path or a data: URL. An http(s) URL is\n refused locally and never fetched. The API does not fetch URLs either: an\n https image URL returned a 400 in the probe.\nEach image may be at most 2,048 tiles of 32 Ɨ 32 pixels. Round the width\n and the height to the nearest multiple of 32 and keep (width / 32) Ɨ\n (height / 32) at or under 2,048: by that count 1440 Ɨ 1440 (2,025 tiles) and\n 2048 Ɨ 1024 (2,048) fit, and 1600 Ɨ 1310 (2,050) does not. Perplexity documents\n this rule, and it was checked at its edge. Fitting sizes were answered in\n about a second: 1920 Ɨ 1080 (2,040 tiles) in 1.1 seconds, 2048 Ɨ 1024 (2,048,\n exactly the cap) in 1.3 and 1450 Ɨ 1450 in 1.1.","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Decisions Provider Guide","lvl2":"Images","lvl3":""}},
6910
+ {"objectID":"8ad813ca01099f441bbc2db40ee064d9a0294d269166c42a6f87fce2a1c1330f","title":"When NeuroLink uses it","url":"/docs/getting-started/providers/perplexity-decider#when-neurolink-uses-it","content":"Every built-in consumer of decide asks for the default decision provider. That\nis the first one that is configured, in the environment or in the credentials\npassed to the SDK, in the order TypeSafe, Laya, XOR, Perplexity, then\nCloudflare Clef. TypeSafe counts\nwith either of its keys, TYPESAFE_API_KEY or AI_GATEWAY_API_KEY. Laya and XOR\neach count only with both their key and their base URL. Perplexity counts with\nits key alone, a non-blank PERPLEXITY_API_KEY or\ncredentials.perplexityDecider.apiKey. Cloudflare Clef counts only with both\nCLOUDFLARE_API_KEY and CLOUDFLARE_ACCOUNT_ID. A caller can always name\nPerplexity with provider: \"perplexity-decider\". So:\nOnly a Perplexity key: built-in features use Perplexity.\nA Perplexity key, plus TypeSafe's key, or Laya's or XOR's key and base\n URL: built-in features use TypeSafe, Laya or XOR, in that order, whichever\n is configured. Perplexity runs only where a caller asks for\n provider: \"perplexity-decider\".\nA Perplexity key, plus CLOUDFLARE_API_KEY and CLOUDFLARE_ACCOUNT_ID:\n built-in features use Perplexity; Cloudflare Clef runs only where a caller\n asks for provider: \"cloudflare-clef\".\nA key for the text provider only: the same as the first case. The key is\n shared, so the text provider's key also configures this one. See\n One key, two providers.\nNone of TypeSafe, Laya, XOR, Perplexity or Cloudflare Clef configured:\n everything behaves exactly as it did without a decision model.\n\nThe classifier router's auto gate reads the credentials given to the\nNeuroLink constructor and the environment, not credentials passed on a single\ncall, so a per-call credentials.perplexityDecider does not influence it. The\nother consumers NeuroLink wires itself (compaction and tool routing) call\ntryDecide without credentials too.","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Decisions Provider Guide","lvl2":"When NeuroLink uses it","lvl3":""}},
6889
6911
  {"objectID":"9545ac1f3da9bfc59774adf1978baa598ba4fc65670b5d0d4c96f7c10e2a1c3c","title":"What is sent to Perplexity","url":"/docs/getting-started/providers/perplexity-decider#what-is-sent-to-perplexity","content":"Perplexity receives the state and the questions of every decision made on\nyour behalf. The questions are NeuroLink's fixed wording, plus text from your own\nconversation where a consumer quotes it; the state is where your text goes.\nThis is what each built-in consumer sends and what turns it on, read from the\ncode:\nModel routing. One request behind the router, the model catalogue and the\n per-request context budget.\nSends: the first 8,000 characters of the prompt of the call being routed,\n plus whether tools are available, whether the request includes images, the\n estimated input tokens, the requested thinking level, whether a session is\n bound and how many prior messages the caller passed in. When the pool has\n more than one model, the questions carry the ids and descriptions of the\n models in it. The images, the session id and the conversation itself are not\n sent.\nTurned on by: classifierRouter.enabled with a pool: an opt-in. With\n the default classifier: \"auto\", a configured decision provider selects the\n decision strategy. classifier: \"heuristic\" or \"llm\" never calls a\n decision provider. A call that pins both provider and model is not\n routed.\nRelevance compaction. Stage 0 of context compaction.\nSends: the current request (first 4,000 characters) and, for each earlier\n user or assistant message that is plain text (not a tool call or result, a\n summary or a pinned skill), its role and its first 1,200 characters. That\n text travels twice, once in the state and once in the message's question. By\n default the six most recent messages are never sent\n (contextRelevance.protectRecent), and at most 300 are asked about.\nTurned on by: nothing beyond a configured decision provider. It runs\n inside context compaction, which starts when a conversation outgrows its\n context budget.\nSummary-quality gate. Guards the summarization stage of compaction.","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Decisions Provider Guide","lvl2":"What is sent to Perplexity","lvl3":""}},
6890
6912
  {"objectID":"01342235e59d6796c886025eea6c3dbdefc5f4420f47123db9b23eacc22955d3","title":"One key, two providers","url":"/docs/getting-started/providers/perplexity-decider#one-key-two-providers","content":"PERPLEXITY_API_KEY is the variable the\nPerplexity text provider reads to serve generate() and\nstream(). It is also the variable that configures this provider. Nothing else\nneeds to be set, so a host that set the key only to use Sonar has also\nconfigured decide:\nWith none of TypeSafe, Laya or XOR configured, the texts listed in\n What is sent to Perplexity go to Perplexity's\n Decisions API. Context compaction needs no other opt-in, so a host with long\n conversations starts sending conversation text the first time one outgrows its\n budget. Model routing, tool routing and RAG planning do so only where you have\n turned them on. No code changes: the key is the switch.\nIt is fail-open, and TypeSafe, Laya and XOR take precedence. If you have\n configured any of them, the shared key changes nothing for the built-in\n consumers, and a Perplexity failure leaves behaviour as it was.\nWith no decision provider configured at all, nothing is sent.\n\nAn unedited copy of .env.example does not set the key, because its\nPERPLEXITY_API_KEY line is commented out.\n\nA key in a .env file counts as a key in the environment. Importing the SDK\nloads the .env in the current working directory, and the CLI does the same at\nstart; when DOTENV_CONFIG_PATH is set, that file is loaded instead. A key that\nsits in .env therefore configures decide exactly as one exported in the\nshell does.\n\nTo keep a key that was set for Sonar from activating decisions, use one of\nthese:\nHand the key to the text provider through the SDK, not the environment, and\n keep it out of .env.\n new NeuroLink({ credentials: { perplexity: { apiKey } } }) configures the\n text provider and not this one, because NeuroLink looks for this provider's\n key in PERPLEXITY_API_KEY and in credentials.perplexityDecider, never in\n credentials.perplexity.","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Decisions Provider Guide","lvl2":"One key, two providers","lvl3":""}},
6891
6913
  {"objectID":"2c9da95e9403e04ac92a64727b9b098a3b1e626742b59f4ffebc6e24914ec316","title":"Measured on a real account, October 2026","url":"/docs/getting-started/providers/perplexity-decider#measured-on-a-real-account-october-2026","content":"NeuroLink probed the live API from an account limited to 10 requests per second.\nThese figures were observed, not taken from Perplexity's documentation:\n\n| What | Measured |\n| --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |\n| Questions per request | 128 accepted (128 answers in 8.9 s, 11,904 input tokens); 129 is refused with a 400. An extra minimal question adds about 92 input tokens, and a one-character state with one question bills 116.","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Decisions Provider Guide","lvl2":"Measured on a real account, October 2026","lvl3":""}},
@@ -6895,7 +6917,7 @@
6895
6917
  {"objectID":"e0c0e895220d617b9c97e379206986672c638c01de25fb9e0b52fa7c04bf0ca1","title":"Rate limits","url":"/docs/getting-started/providers/perplexity-decider#rate-limits","content":"The account measured allows 10 requests per second. A 429 is retried once, after\nthe wait its Retry-After header asks for. NeuroLink honours that header on any\nretried reply that carries it, and only as a whole number of seconds above zero\nor as an HTTP date in the future; anything else (zero, a negative number, a hex\nor exponent spelling, a fraction, a date already past) is ignored, and the\ndefault backoff of a quarter to half a second is used instead. A wait longer\nthan what is left of the attempt's timeout is not taken: the error is thrown at\nonce so a fail-open consumer is not held for the length of a cooldown, and a\ncaller's abort signal ends a wait immediately.","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Decisions Provider Guide","lvl2":"Rate limits","lvl3":""}},
6896
6918
  {"objectID":"f40b2da5feaf02bc4eaf5102d75c33ce6fc9f979f7106853002b142aad4fd67f","title":"Errors","url":"/docs/getting-started/providers/perplexity-decider#errors","content":"| Reply | Kind | Retried |\n| -------------------------------------------------------------------- | ------------------------------------------------------- | ------------------------------ |\n| 401 | authentication — the provider instance stops retrying | no |\n| 413 | max_tokens_exceeded (the body is over 32 MiB) | no |\n| 400 saying the input exceeds the model's maximum context length | max_tokens_exceeded | no |\n| any other 4xx (for example another 400, 403, 404) | invalid_request, with Perplexity's own message | no |\n| 429 | rate_limit | yes, once, after Retry-After |\n| 503 | overloaded | yes, once |\n| other 5xx (500, 502, 504) | server | yes, once |\n| no response within the timeout | timeout | yes, once |\n| the request fails before any reply (connection refused, DNS failure) | network | yes, once |\n| the caller's own signal aborts the request | network | no |\n| state or questions over the limit (local)","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Decisions Provider Guide","lvl2":"Errors","lvl3":""}},
6897
6919
  {"objectID":"a812289efc882316bb62c31aea533cd0a0e8e219eeea12289b4503450b4dd9a5","title":"Troubleshooting","url":"/docs/getting-started/providers/perplexity-decider#troubleshooting","content":"Perplexity requires an API key — set PERPLEXITY_API_KEY, or pass\n credentials.perplexityDecider.apiKey. A key passed as\n credentials.perplexity does not count.\nA 401, Invalid API key provided — Perplexity rejected the key, and it\n checks the key before anything else. The provider instance stops retrying\n after a rejection, so fix the key and construct a new one.\nA 400 naming the model — PERPLEXITY_DECIDER_MODEL, or a per-call\n model, names something other than pplx-decider-v1-27b.\nThe Perplexity base URL must not carry credentials… — the override has a\n user name, a password, a query string or a fragment. Set it to the origin only,\n or leave it unset for https://api.perplexity.ai.\nThe Perplexity base URL must start with https:// or http:// — the value\n parses as a URL with another scheme: an explicit one such as ftp://, or a host\n name and port with no scheme, such as localhost:8080, which reads as the\n scheme localhost:. Put https:// or http:// in front.\nThe Perplexity base URL is not a valid absolute URL — the value does not\n parse as a URL at all: a bare host name such as api.perplexity.ai (the likely\n mistake), an IP address and port such as 127.0.0.1:8080, a path, or a scheme\n with no host. Write the origin with its scheme.\nA refused base URL in the debug log — it is not there. A base URL that\n NeuroLink refuses (credentials, a query string, a fragment, a scheme other than\n http or https, or not a URL at all) is never written to the debug log: the line\n shows (invalid) in its place, because such a value can carry a secret.\nImage N is not a PNG, JPEG or WebP image — Perplexity reads only those\n three. Convert the image first.\nImage N is … pixels; Perplexity reads at most 2048 tiles — resize the\n image to fit, for example 1440 Ɨ 1440 or 2048 Ɨ 1024.\nA request with an image stalls until it times out, or fails as server with\n HTTP 504 when the timeout is longer than a minute — an image over the tile\n cap does this.","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Decisions Provider Guide","lvl2":"Troubleshooting","lvl3":""}},
6898
- {"objectID":"6d580e7ffb3867905a72879cdccbab76b5df280677e9d18b0029fd5bc5092025","title":"See also","url":"/docs/getting-started/providers/perplexity-decider#see-also","content":"The decide inference type\nTypeSafe (Jev) Provider Guide\nLaya Provider Guide\nXOR Provider Guide\nPerplexity text provider\nPerplexity Decisions API reference","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Decisions Provider Guide","lvl2":"See also","lvl3":""}},
6920
+ {"objectID":"6d580e7ffb3867905a72879cdccbab76b5df280677e9d18b0029fd5bc5092025","title":"See also","url":"/docs/getting-started/providers/perplexity-decider#see-also","content":"The decide inference type\nTypeSafe (Jev) Provider Guide\nLaya Provider Guide\nXOR Provider Guide\nCloudflare Clef Provider Guide\nPerplexity text provider\nPerplexity Decisions API reference","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Decisions Provider Guide","lvl2":"See also","lvl3":""}},
6899
6921
  {"objectID":"6d55ee7ec0f21369a70b421eaeaddf6b44fe88746beb28c252588d42d97048b7","title":"Perplexity Provider Guide","url":"/docs/getting-started/providers/perplexity","content":"Perplexity Provider Guide\n\nWeb-search-augmented generation via Perplexity's Sonar models\n\nOverview\n\nPerplexity hosts a family of LLMs (Sonar,\nSonar Pro, Sonar Reasoning) that pair the model with a live web search\nbackend; responses include citations to the documents the model relied on.\n\nKey Facts\nProtocol: OpenAI-compatible (/chat/completions)\nDefault base URL: https://api.perplexity.ai\nDefault model: sonar\nCitations: Returned via citations field on the response\nStreaming: Yes\nTool calling: No\n\nQuick Start\nGet an API Key\n\nhttps://www.perplexity.ai/settings/api\nConfigure\n\nPERPLEXITY_API_KEY also configures decisions. The same variable configures\nthe Perplexity Decisions provider (perplexity-decider,\nwhich serves decide()). So setting it for Sonar also lets NeuroLink's built-in\nfeatures send decision requests to Perplexity when no TypeSafe, Laya or XOR\nprovider is configured, and context compaction needs no other opt-in to do so.\nTo avoid that, pass the key through the SDK as credentials.perplexity.apiKey\ninstead of the environment, and keep it out of .env too: the SDK and the CLI\nload that file into the environment, so a PERPLEXITY_API_KEY line like the one\nabove would still configure decisions. (That the text provider reads this slice\nis taken from the code and has not been run against the live API.) See\nWhat is sent to Perplexity\nand One key, two providers.\nGenerate\n\nSupported Models\n\n| Model ID | Notes |\n| ----------------- | ----------------------------------- |\n| sonar | Default; fast search-augmented chat |\n| sonar-pro | Larger context, deeper search |\n| sonar-reasoning | Chain-of-thought reasoning |\n\nCLI Usage\n\nProvider Aliases\n\n| Alias | Example |\n| ------------ | ----------------------- |\n| perplexity | --provider perplexity |\n\nConfiguration Reference\n\n| Environment Variable | Required | Default |\n| --------------------- | -------- | --------------------------- |\n| PERPLEXITY_API_KEY | Yes | — |\n| PERPLEXITY_MODEL | No | sonar |\n| PERPLEXITY_BASE_URL | No | https://api.perplexity.ai |\n\nFeature Support Matrix\n\n| Feature | Support |\n| ----------------- | ------------ |\n| Text generation | Yes |\n| Streaming | Yes |\n| Tool calling | No |\n| Structured output | Limited |\n| Web search | Yes (native) |\n| Citations | Yes |\n\nSee Also\nPerplexity Decisions Provider — typed, calibrated decide() judgments on the same key\nAnthropic Provider — web-search via the web_search tool\nVertex Provider — Google search grounding","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Provider Guide","lvl2":"","lvl3":""}},
6900
6922
  {"objectID":"59d597c2561bea4977d9a4ae9a111235c2a042b1573e0fad7cf2147851b79e13","title":"Perplexity Provider Guide","url":"/docs/getting-started/providers/perplexity#perplexity-provider-guide","content":"Web-search-augmented generation via Perplexity's Sonar models","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Provider Guide","lvl2":"Perplexity Provider Guide","lvl3":""}},
6901
6923
  {"objectID":"755494265419901fd106a26a27844178614744eacb256af720b7acfdc700a22c","title":"Overview","url":"/docs/getting-started/providers/perplexity#overview","content":"Perplexity hosts a family of LLMs (Sonar,\nSonar Pro, Sonar Reasoning) that pair the model with a live web search\nbackend; responses include citations to the documents the model relied on.","hierarchy":{"lvl0":"Getting Started","lvl1":"Perplexity Provider Guide","lvl2":"Overview","lvl3":""}},
@@ -7213,19 +7235,19 @@
7213
7235
  {"objectID":"14c8755cfa45311a2b8645d3729d760f513b4e768f4cbe6d75e1dcaaa68c7e94","title":"Configuration Reference","url":"/docs/getting-started/providers/together-ai#configuration-reference","content":"| Environment Variable | Required | Default |\n| -------------------- | -------- | ----------------------------------------- |\n| TOGETHER_API_KEY | Yes | — |\n| TOGETHER_MODEL | No | meta-llama/Llama-3.3-70B-Instruct-Turbo |\n| TOGETHER_BASE_URL | No | https://api.together.xyz/v1 |","hierarchy":{"lvl0":"Getting Started","lvl1":"Together AI Provider Guide","lvl2":"Configuration Reference","lvl3":""}},
7214
7236
  {"objectID":"73dacb114377509534778a0d5d1d386e64fe4e71f75f590903d6a5c3dda075c6","title":"Feature Support Matrix","url":"/docs/getting-started/providers/together-ai#feature-support-matrix","content":"| Feature | Support |\n| ----------------- | ----------------- |\n| Text generation | Yes |\n| Streaming | Yes |\n| Tool calling | Yes (model-dep.) |\n| Structured output | Yes (model-dep.) |\n| Vision | Yes (Llama 3.2 V) |\n| Embeddings | Limited |","hierarchy":{"lvl0":"Getting Started","lvl1":"Together AI Provider Guide","lvl2":"Feature Support Matrix","lvl3":""}},
7215
7237
  {"objectID":"5a31d744d2d35af19cc87e27118e0d4057ab98d14fa7be6e9acb01dcbfc29f6a","title":"See Also","url":"/docs/getting-started/providers/together-ai#see-also","content":"Fireworks Provider\nGroq Provider","hierarchy":{"lvl0":"Getting Started","lvl1":"Together AI Provider Guide","lvl2":"See Also","lvl3":""}},
7216
- {"objectID":"63e8301588e0fd846272e6b041d27fbe6a754543f38d949f07ce7e6f181a0ced","title":"TypeSafe (Jev) Provider Guide","url":"/docs/getting-started/providers/typesafe","content":"TypeSafe (Jev) Provider Guide\n\nOne of the providers that serve decide rather than generate/stream —\nit returns typed, calibrated judgments and emits no text at all. The others are\nLaya and XOR, open-weights models served from an endpoint\nyou configure, and Perplexity, a hosted API. When more\nthan one is configured, TypeSafe is tried before the others.\n\nOverview\n\nTypeSafe's Jev is a \"System One\" model. You send one state plus a map of\nnamed, typed questions; it returns one typed answer per question, all evaluated\nin a single parallel pass. Nothing has to be parsed back out of prose, and every\nchoice/score answer carries a calibrated confidence rather than a\nself-reported one.\n\nBecause it emits no text, generate() and stream() are not available and\ngetAISDKModel() throws — the same shape Voyage and Jina already use for\nembedding-only providers. Its descriptor declares inferenceKinds: [\"decide\"],\nwhich keeps it out of auto-select and the health sweep, so those throws are\nunreachable in normal use.\n\nThis is not neurolink.evaluate(), which\nscores an already-generated response with RAGAS scorers. Different feature,\ndifferent word.\n\nKey Facts\nProvider id: typesafe (aliases: jev, typesafe-ai)\nInference kinds: decide only — one of the providers that do (the others\n are Laya, XOR and Perplexity);\n TypeSafe is tried before them when more than one is configured\nTool calling: none (toolSupport: \"none\") — a decision model calls nothing\nHealth check: env-only; it is never probed with a live generation\nDefault decide timeout: 5000 ms (timeouts.decideMs)\nLatency: flat in question count — 1 question ~393 ms, 400 questions\n ~465 ms. Concurrent requests queue instead, so batch every question into one\n call rather than fanning out.\nCost: ~$0.042 per million input tokens, output billed at zero — about\n $0.00002 per decision. Output tokens are reported — measured 21 for a\n single question, converging to ~17.5 per question in a batch of eight — they\n are simply not charged.\nAccuracy is the trade: 67.8% on TypeSafe's own 711-case benchmark against\n Opus 5's 73.1%. Right for decisions that are gated and reversible; wrong for\n final answers.\n\nQuick Start\nGet an API key\n\nCreate one at console.typesafe.ai/keys.\nConfigure\nUse it\n\ntryDecide() returns null on any failure. Use decide() when you want the\nfailure to surface; it throws a ProviderError whose cause carries a typed\nkind.\n\nFrom the CLI\n\nneurolink decide [state] calls decide() under the hood and reads\nTYPESAFE_API_KEY from the environment like every other CLI command. See the\nCLI command reference.\n\nThe degradation contract\n\nSetting the key is the entire switch for TypeSafe, and removing it is a\ncomplete undo while no other decision provider is configured. Every internal\nconsumer of decide fails open: with no decision provider configured, model\nrouting, context budgeting, relevance compaction, tool routing and RAG planning\nall behave exactly as they did before. Note that Perplexity's key is shared with\nits text provider, so a PERPLEXITY_API_KEY set for that provider also counts as\na configured decision provider; see\nOne key, two providers. There is no\nconfiguration in which a missing, invalid, slow or unreachable decision model\nchanges NeuroLink's observable behaviour.\n\nA credential the service does not accept disables that provider instance rather\nthan paying a round trip on every later call to be told so again.\n\nTwo transports\n\nThe same model is reachable two ways, and the choice is made once in the\nconstructor.\n\n| | Direct | Vercel AI Gateway |\n| ------------------- | ------------------- | --------------------------------------------- |\n| Key | TYPESAFE_API_KEY | AI_GATEWAY_API_KEY |\n| Endpoint | api.typesafe.ai | ai-gateway.vercel.sh/v4/ai/evaluation-model |\n| Endpoint override | TYPESAFE_BASE_URL | TYPESAFE_GATEWAY_URL |\n| Model named in | request body | ai-model-id header |\n| Question vocabulary | noul | boolean |\n| confidence | on each answer | on providerMetadata |\n| Billed by | TypeSafe | Vercel |\n\nHolding both keys keeps the direct transport, so the confidence figures a\nhost already sees do not shift underneath it when a second key appears. Force\none with TYPESAFE_TRANSPORT=direct|gateway or\ncredentials.typesafe.transport. Either endpoint can be moved without a\nrelease: credentials.typesafe.baseURL / TYPESAFE_BASE_URL for the direct\none, credentials.typesafe.gatewayURL / TYPESAFE_GATEWAY_URL for the\ngateway route.\n\nāš ļø The gateway refuses every request — free credits included — until the\nVercel team has a credit card on file, returning 403\ncustomer_verification_required. That is an account state, not a bad key, and it\narrives before the model id ","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"","lvl3":""}},
7217
- {"objectID":"7320dc5177f0e320fe7887f77c01c06519bb6c0ddeb8a8b38ecfa013a9363a1f","title":"TypeSafe (Jev) Provider Guide","url":"/docs/getting-started/providers/typesafe#typesafe-jev-provider-guide","content":"One of the providers that serve decide rather than generate/stream —\nit returns typed, calibrated judgments and emits no text at all. The others are\nLaya and XOR, open-weights models served from an endpoint\nyou configure, and Perplexity, a hosted API. When more\nthan one is configured, TypeSafe is tried before the others.","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"TypeSafe (Jev) Provider Guide","lvl3":""}},
7238
+ {"objectID":"63e8301588e0fd846272e6b041d27fbe6a754543f38d949f07ce7e6f181a0ced","title":"TypeSafe (Jev) Provider Guide","url":"/docs/getting-started/providers/typesafe","content":"TypeSafe (Jev) Provider Guide\n\nOne of the providers that serve decide rather than generate/stream —\nit returns typed, calibrated judgments and emits no text at all. The others are\nLaya and XOR, open-weights models served from an endpoint\nyou configure, and Perplexity and\nCloudflare Clef, hosted APIs. When more than one is\nconfigured, TypeSafe is tried before the others.\n\nOverview\n\nTypeSafe's Jev is a \"System One\" model. You send one state plus a map of\nnamed, typed questions; it returns one typed answer per question, all evaluated\nin a single parallel pass. Nothing has to be parsed back out of prose, and every\nchoice/score answer carries a calibrated confidence rather than a\nself-reported one.\n\nBecause it emits no text, generate() and stream() are not available and\ngetAISDKModel() throws — the same shape Voyage and Jina already use for\nembedding-only providers. Its descriptor declares inferenceKinds: [\"decide\"],\nwhich keeps it out of auto-select and the health sweep, so those throws are\nunreachable in normal use.\n\nThis is not neurolink.evaluate(), which\nscores an already-generated response with RAGAS scorers. Different feature,\ndifferent word.\n\nKey Facts\nProvider id: typesafe (aliases: jev, typesafe-ai)\nInference kinds: decide only — one of the providers that do (the others\n are Laya, XOR, Perplexity and\n Cloudflare Clef); TypeSafe is tried before them when\n more than one is configured\nTool calling: none (toolSupport: \"none\") — a decision model calls nothing\nHealth check: env-only; it is never probed with a live generation\nDefault decide timeout: 5000 ms (timeouts.decideMs)\nLatency: flat in question count — 1 question ~393 ms, 400 questions\n ~465 ms. Concurrent requests queue instead, so batch every question into one\n call rather than fanning out.\nCost: ~$0.042 per million input tokens, output billed at zero — about\n $0.00002 per decision. Output tokens are reported — measured 21 for a\n single question, converging to ~17.5 per question in a batch of eight — they\n are simply not charged.\nAccuracy is the trade: 67.8% on TypeSafe's own 711-case benchmark against\n Opus 5's 73.1%. Right for decisions that are gated and reversible; wrong for\n final answers.\n\nQuick Start\nGet an API key\n\nCreate one at console.typesafe.ai/keys.\nConfigure\nUse it\n\ntryDecide() returns null on any failure. Use decide() when you want the\nfailure to surface; it throws a ProviderError whose cause carries a typed\nkind.\n\nFrom the CLI\n\nneurolink decide [state] calls decide() under the hood and reads\nTYPESAFE_API_KEY from the environment like every other CLI command. See the\nCLI command reference.\n\nThe degradation contract\n\nSetting the key is the entire switch for TypeSafe, and removing it is a\ncomplete undo while no other decision provider is configured. Every internal\nconsumer of decide fails open: with no decision provider configured, model\nrouting, context budgeting, relevance compaction, tool routing and RAG planning\nall behave exactly as they did before. Note that Perplexity's key is shared with\nits text provider, so a PERPLEXITY_API_KEY set for that provider also counts as\na configured decision provider; see\nOne key, two providers. The same\ngoes for CLOUDFLARE_API_KEY with CLOUDFLARE_ACCOUNT_ID, which the Workers AI\ntext provider reads too: set together they configure Cloudflare Clef, the last\nfallback; see One token, two providers. There is no\nconfiguration in which a missing, invalid, slow or unreachable decision model\nchanges NeuroLink's observable behaviour.\n\nA credential the service does not accept disables that provider instance rather\nthan paying a round trip on every later call to be told so again.\n\nTwo transports\n\nThe same model is reachable two ways, and the choice is made once in the\nconstructor.\n\n| | Direct | Vercel AI Gateway |\n| ------------------- | ------------------- | --------------------------------------------- |\n| Key | TYPESAFE_API_KEY | AI_GATEWAY_API_KEY |\n| Endpoint | api.typesafe.ai | ai-gateway.vercel.sh/v4/ai/evaluation-model |\n| Endpoint override | TYPESAFE_BASE_URL | TYPESAFE_GATEWAY_URL |\n| Model named in | request body | ai-model-id header |\n| Question vocabulary | noul | boolean |\n| confidence | on each answer | on providerMetadata |\n| Billed by | TypeSafe | Vercel |\n\nHolding both keys keeps the direct transport, so the confidence figures a\nhost already sees do not shift underneath it when a second key appears. Force\none with TYPESAFE_TRANSPORT=direct|gateway or\ncredentials.typesafe.transport. Either endpoint can be moved without a\nrelease: credentials.typesafe.baseURL / TYPESAFE_BASE_URL for the direct\none, credentials.typesafe.gatewayURL / TYPESAFE_GATEWAY_URL for the\ngateway","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"","lvl3":""}},
7239
+ {"objectID":"7320dc5177f0e320fe7887f77c01c06519bb6c0ddeb8a8b38ecfa013a9363a1f","title":"TypeSafe (Jev) Provider Guide","url":"/docs/getting-started/providers/typesafe#typesafe-jev-provider-guide","content":"One of the providers that serve decide rather than generate/stream —\nit returns typed, calibrated judgments and emits no text at all. The others are\nLaya and XOR, open-weights models served from an endpoint\nyou configure, and Perplexity and\nCloudflare Clef, hosted APIs. When more than one is\nconfigured, TypeSafe is tried before the others.","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"TypeSafe (Jev) Provider Guide","lvl3":""}},
7218
7240
  {"objectID":"8cc640004e04379192695b8fcd3d49552ac0f4ebd763093bc610b4c93e823fee","title":"Overview","url":"/docs/getting-started/providers/typesafe#overview","content":"TypeSafe's Jev is a \"System One\" model. You send one state plus a map of\nnamed, typed questions; it returns one typed answer per question, all evaluated\nin a single parallel pass. Nothing has to be parsed back out of prose, and every\nchoice/score answer carries a calibrated confidence rather than a\nself-reported one.\n\nBecause it emits no text, generate() and stream() are not available and\ngetAISDKModel() throws — the same shape Voyage and Jina already use for\nembedding-only providers. Its descriptor declares inferenceKinds: [\"decide\"],\nwhich keeps it out of auto-select and the health sweep, so those throws are\nunreachable in normal use.\n\nThis is not neurolink.evaluate(), which\nscores an already-generated response with RAGAS scorers. Different feature,\ndifferent word.","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"Overview","lvl3":""}},
7219
- {"objectID":"4921fe7ae1c50c817be3b7742587c455084bee9d988015cf475edd5c3c196b84","title":"Key Facts","url":"/docs/getting-started/providers/typesafe#key-facts","content":"Provider id: typesafe (aliases: jev, typesafe-ai)\nInference kinds: decide only — one of the providers that do (the others\n are Laya, XOR and Perplexity);\n TypeSafe is tried before them when more than one is configured\nTool calling: none (toolSupport: \"none\") — a decision model calls nothing\nHealth check: env-only; it is never probed with a live generation\nDefault decide timeout: 5000 ms (timeouts.decideMs)\nLatency: flat in question count — 1 question ~393 ms, 400 questions\n ~465 ms. Concurrent requests queue instead, so batch every question into one\n call rather than fanning out.\nCost: ~$0.042 per million input tokens, output billed at zero — about\n $0.00002 per decision. Output tokens are reported — measured 21 for a\n single question, converging to ~17.5 per question in a batch of eight — they\n are simply not charged.\nAccuracy is the trade: 67.8% on TypeSafe's own 711-case benchmark against\n Opus 5's 73.1%. Right for decisions that are gated and reversible; wrong for\n final answers.","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"Key Facts","lvl3":""}},
7241
+ {"objectID":"4921fe7ae1c50c817be3b7742587c455084bee9d988015cf475edd5c3c196b84","title":"Key Facts","url":"/docs/getting-started/providers/typesafe#key-facts","content":"Provider id: typesafe (aliases: jev, typesafe-ai)\nInference kinds: decide only — one of the providers that do (the others\n are Laya, XOR, Perplexity and\n Cloudflare Clef); TypeSafe is tried before them when\n more than one is configured\nTool calling: none (toolSupport: \"none\") — a decision model calls nothing\nHealth check: env-only; it is never probed with a live generation\nDefault decide timeout: 5000 ms (timeouts.decideMs)\nLatency: flat in question count — 1 question ~393 ms, 400 questions\n ~465 ms. Concurrent requests queue instead, so batch every question into one\n call rather than fanning out.\nCost: ~$0.042 per million input tokens, output billed at zero — about\n $0.00002 per decision. Output tokens are reported — measured 21 for a\n single question, converging to ~17.5 per question in a batch of eight — they\n are simply not charged.\nAccuracy is the trade: 67.8% on TypeSafe's own 711-case benchmark against\n Opus 5's 73.1%. Right for decisions that are gated and reversible; wrong for\n final answers.","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"Key Facts","lvl3":""}},
7220
7242
  {"objectID":"f5f7265f7ee9ba7ead6f6016a4b9232258ffaf4256e544874d6bcca8323dbb04","title":"1. Get an API key","url":"/docs/getting-started/providers/typesafe#1-get-an-api-key","content":"Create one at console.typesafe.ai/keys.","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"1. Get an API key","lvl3":""}},
7221
7243
  {"objectID":"a7cfa6599d9458e8c4a9bfcef1bba07a4948ebe6b0f30a1352df07a713040d8e","title":"3. Use it","url":"/docs/getting-started/providers/typesafe#3-use-it","content":"tryDecide() returns null on any failure. Use decide() when you want the\nfailure to surface; it throws a ProviderError whose cause carries a typed\nkind.","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"3. Use it","lvl3":""}},
7222
7244
  {"objectID":"4ec7febf2956a104b0037472bab2e6c652e17df7f7cfdf69b73ebe2c5003f082","title":"From the CLI","url":"/docs/getting-started/providers/typesafe#from-the-cli","content":"neurolink decide [state] calls decide() under the hood and reads\nTYPESAFE_API_KEY from the environment like every other CLI command. See the\nCLI command reference.","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"From the CLI","lvl3":""}},
7223
- {"objectID":"98aafdb50791dab3469be7930897496cf4d2fb9dc57e42bacb2ee179ac0064f2","title":"The degradation contract","url":"/docs/getting-started/providers/typesafe#the-degradation-contract","content":"Setting the key is the entire switch for TypeSafe, and removing it is a\ncomplete undo while no other decision provider is configured. Every internal\nconsumer of decide fails open: with no decision provider configured, model\nrouting, context budgeting, relevance compaction, tool routing and RAG planning\nall behave exactly as they did before. Note that Perplexity's key is shared with\nits text provider, so a PERPLEXITY_API_KEY set for that provider also counts as\na configured decision provider; see\nOne key, two providers. There is no\nconfiguration in which a missing, invalid, slow or unreachable decision model\nchanges NeuroLink's observable behaviour.\n\nA credential the service does not accept disables that provider instance rather\nthan paying a round trip on every later call to be told so again.","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"The degradation contract","lvl3":""}},
7245
+ {"objectID":"98aafdb50791dab3469be7930897496cf4d2fb9dc57e42bacb2ee179ac0064f2","title":"The degradation contract","url":"/docs/getting-started/providers/typesafe#the-degradation-contract","content":"Setting the key is the entire switch for TypeSafe, and removing it is a\ncomplete undo while no other decision provider is configured. Every internal\nconsumer of decide fails open: with no decision provider configured, model\nrouting, context budgeting, relevance compaction, tool routing and RAG planning\nall behave exactly as they did before. Note that Perplexity's key is shared with\nits text provider, so a PERPLEXITY_API_KEY set for that provider also counts as\na configured decision provider; see\nOne key, two providers. The same\ngoes for CLOUDFLARE_API_KEY with CLOUDFLARE_ACCOUNT_ID, which the Workers AI\ntext provider reads too: set together they configure Cloudflare Clef, the last\nfallback; see One token, two providers. There is no\nconfiguration in which a missing, invalid, slow or unreachable decision model\nchanges NeuroLink's observable behaviour.\n\nA credential the service does not accept disables that provider instance rather\nthan paying a round trip on every later call to be told so again.","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"The degradation contract","lvl3":""}},
7224
7246
  {"objectID":"8fd6c7607ca461fdf5e7fb3b743b1761645c1ca9a2e666330b59614acb6455fa","title":"Two transports","url":"/docs/getting-started/providers/typesafe#two-transports","content":"The same model is reachable two ways, and the choice is made once in the\nconstructor.\n\n| | Direct | Vercel AI Gateway |\n| ------------------- | ------------------- | --------------------------------------------- |\n| Key | TYPESAFE_API_KEY | AI_GATEWAY_API_KEY |\n| Endpoint | api.typesafe.ai | ai-gateway.vercel.sh/v4/ai/evaluation-model |\n| Endpoint override | TYPESAFE_BASE_URL | TYPESAFE_GATEWAY_URL |\n| Model named in | request body | ai-model-id header |\n| Question vocabulary | noul | boolean |\n| confidence | on each answer | on providerMetadata |\n| Billed by | TypeSafe | Vercel |\n\nHolding both keys keeps the direct transport, so the confidence figures a\nhost already sees do not shift underneath it when a second key appears. Force\none with TYPESAFE_TRANSPORT=direct|gateway or\ncredentials.typesafe.transport. Either endpoint can be moved without a\nrelease: credentials.typesafe.baseURL / TYPESAFE_BASE_URL for the direct\none, credentials.typesafe.gatewayURL / TYPESAFE_GATEWAY_URL for the\ngateway route.\n\nāš ļø The gateway refuses every request — free credits included — until the\nVercel team has a credit card on file, returning 403\ncustomer_verification_required. That is an account state, not a bad key, and it\narrives before the model id is validated.\n\nFull detail, including the measured error table and why the distribution peak is\nnot a substitute for the reported confidence, is in\nThe decide inference type.","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"Two transports","lvl3":""}},
7225
7247
  {"objectID":"ac6585086d9cf682e7b1a66265cacfd1c4e961588384fe3883636c5deb66588c","title":"What NeuroLink uses it for","url":"/docs/getting-started/providers/typesafe#what-neurolink-uses-it-for","content":"| Area | What the decision replaces |\n| ---------------------------------------------------------------- | --------------------------------------------------------------- |\n| Model routing | difficulty + capabilities + risk + model pick in one round trip |\n| Model catalogue | one choice over the registry ranks all N candidates at once |\n| Context budget | a rubric-placed scope reading lowers the compaction threshold |\n| Relevance compaction | per-message keep/drop, plus a gate on the generated summary |\n| Tool / MCP routing | one boolean per server, replacing a 15s LLM call at ~400 ms |\n| RAG retrieval | per-query topK / hybrid / graph / rerank planning |","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"What NeuroLink uses it for","lvl3":""}},
7226
7248
  {"objectID":"04bf675bad2aeefe749c5229a396a476273c9b463d3b81a6767b9ca9357aa36d","title":"Limits and gotchas","url":"/docs/getting-started/providers/typesafe#limits-and-gotchas","content":"Two input ceilings, both enforced by the service: state plus the longest\n single question ā‰ˆ 33,000 tokens, and state plus all questions ā‰ˆ\n 64,000. Exceeding either returns max_tokens_exceeded with no message\n at all — the provider supplies a real sentence in its place.\nBatch, never fan out. Latency is flat in question count but concurrent\n requests queue, so a second round trip costs far more than a hundred extra\n questions.\nA boolean carries no confidence of its own. Use\n decisionBooleanConfidence(p) — distance from a coin flip, so 0.5 → 0 and\n 0/1 → 1.\n403 vs 401 are inverted on the direct API, and from TypeSafe's own docs: a\n missing Authorization header returns 403, an invalid key returns\nThe gateway does not share this quirk.","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"Limits and gotchas","lvl3":""}},
7227
7249
  {"objectID":"7d7786b2749554b22a763088bb2c011be0b62bdeb53ee881e9331cf8616dffa7","title":"Troubleshooting","url":"/docs/getting-started/providers/typesafe#troubleshooting","content":"| Symptom | Cause | Fix |\n| ------------------------------------ | --------------------------------------------------------------------- | ----------------------------------------------------------------------------- |\n| tryDecide() always returns null | No key, or the key was rejected once and the instance disabled itself | Check TYPESAFE_API_KEY; construct a new instance after fixing it |\n| max_tokens_exceeded | One of the two input ceilings | Shorten state, or split questions across calls — but prefer shrinking state |\n| 403 customer_verification_required | Gateway transport, no card on the Vercel team | Add a payment method, or use the direct transport |\n| Routing never changes | A modelPool is configured, which owns selection outright | See Provider Orchestration |","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"Troubleshooting","lvl3":""}},
7228
- {"objectID":"6116c23ca86f351b03ac4844a1d76d9de849fb522ab4a03eb37f9e75017e567c","title":"See also","url":"/docs/getting-started/providers/typesafe#see-also","content":"The decide inference type — the full reference\nModel routing with a decision model\nLaya Provider Guide\nXOR Provider Guide\nPerplexity Decisions Provider Guide\nProvider setup overview","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"See also","lvl3":""}},
7250
+ {"objectID":"6116c23ca86f351b03ac4844a1d76d9de849fb522ab4a03eb37f9e75017e567c","title":"See also","url":"/docs/getting-started/providers/typesafe#see-also","content":"The decide inference type — the full reference\nModel routing with a decision model\nLaya Provider Guide\nXOR Provider Guide\nPerplexity Decisions Provider Guide\nCloudflare Clef Provider Guide\nProvider setup overview","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"See also","lvl3":""}},
7229
7251
  {"objectID":"0db36dc931d978469a7ad4e6468f4b52f77cc7c0db36272c6cfe8ac11e411f42","title":"Umans AI Provider Guide","url":"/docs/getting-started/providers/umans-ai","content":"Umans AI Provider Guide\n\nUmans AI is a Tier-2 catalog provider: OpenAI-wire-compatible with no\nbehavioural quirks, so its entire integration is one JSON file\n(src/lib/providers/catalog/umans-ai.json) rather than hand-written code.\nThat file is the source of truth for everything on this page.\n\nVerification status: this entry is docs- and roster-verified, not yet\nlive-verified. Every field below comes from Umans AI's own public pages and\nunauthenticated GET requests — no account was created and no API key was\nused to build it. evidence.liveMatrix is null until someone runs the live\ncapability matrix with a real key (see\nVerification status below).\n\nKey Facts\nProvider id: umans-ai (alias umans)\nProtocol: OpenAI-compatible (/chat/completions) — the docs say \"Umans\n Code also implements the OpenAI Chat Completions API\"\n (docs). The same gateway also\n documents an Anthropic Messages route (POST /v1/messages) and a Responses\n route (POST /v1/responses); this entry uses only /chat/completions\nBase URL: https://api.code.umans.ai/v1\nDefault model: umans-coder — a routing alias, not a fixed model, see\n Models\nModels in catalog: 7 (curated from a 10-model live roster; three ids are\n left out on purpose — umans-deepseek-v4.1-flash-lab and\n umans-mimo-v2.6-pro-lab, which /v1/models/info describes as Labs\n experiments, \"temporary, not a permanent id\", with access \"seat-gated through\n the Labs page\", and umans-qwen3.6-35b-a3b, which it describes as a\n \"Technical alias for umans-flash\")\nStreaming: supported — the docs' OpenAI example sends \"stream\": true to\n POST /v1/chat/completions, and the reasoning section says reasoning \"streams\n live, token by token, alongside the answer\"\nTool calling: true — GET /v1/models/info reports supports_tools: true\n for every roster model. The docs show no tools request example for\n /v1/chat/completions, so this rests on that field alone and has not been\n exercised\nTools while streaming: not declared (false) — the docs show no\n tool-call request or response on this route, streaming or otherwise\nStructured output: not declared (false) — no page opened mentions\n response_format, json_schema or json_object\nStructured output + tools together: not declared (false) — no combined\n probe was run without credentials\nEmbeddings: not declared — the docs describe no embeddings route\nThinking: true — the docs' \"Reasoning & Extended Thinking\" section\n documents reasoning_effort on /v1/chat/completions and returns reasoning\n in a reasoning_content field, separate from content. Which levels a model\n accepts varies, see Models\nBilling: no-free-tier — a prepaid wallet billed per token, see\n Billing below\nKey format: the Umans pages conflict, so apiKeyFormat is null and\n NeuroLink does not validate a prefix. The docs' \"Any BYOK Tool\" section says\n the key \"starts with sk-\" and their examples write it as\n sk-your-umans-api-key, while the docs' wallet-summary example shows a key\n prefix of umans_7fKq and the homepage .env example shows umans_\n\nQuick Start\nGet an API key\nVisit: https://app.umans.ai/register and create an account with Google or an email and password (already have one? sign in at https://app.umans.ai/login). Organizations are provisioned by Umans: email contact@umans.ai (https://app.umans.ai/offers/code/docs/orgs)\nOpen https://app.umans.ai/billing, go to Dashboard -> API Keys and generate a key; the docs say it is shown only once, so copy it immediately\nPrepaid wallet, no free tier stated: top up before calling the API - the docs say a wallet with no credit (no paid top-up and no active promo grant) \"does not serve until it has some\", and promo credit serves only while it lasts (https://app.umans.ai/offers/code/docs)\nSet UMANS_AI_API_KEY in your .env file\nConfigure\nUse it\n\nPer-request credentials work as they do for every provider:\n\nBilling\n\nUmans AI's pricing page (https://app.umans.ai/pricing) says \"Pay per token. Top\nup, create a key, and go.\" and \"every request debits exactly what the model\nused\". The docs (https://app.umans.ai/offers/code/docs) say a wallet with no\ncredit at all — no paid top-up and no active promo grant — \"does not serve until\nit has some\"; promo credit \"serves while it lasts\" and then the wallet suspends\nuntil a paid top-up. A funded wallet that drains to its tier's floor (a small\nnegative allowance, $5 at Tiers 0-1) also suspends, and the docs call that \"a\nstop, not a 429\". No page states a free tier or standing free credits, so the\nentry records no-free-tier. The login page's \"Sign up for free\" link is about\ncreating an account; it is not a stated free allowance. The docs document\nGET /v1/usage and a wallet-summary route on app.umans.ai for checking\nbalance and spend. Both need a key; /v1/usage was called only without one,\nfor the 401 probe under Verification status.\n\nRequest limits sit on top of the balance. The docs give each wallet tier a\nrolling five-hour request window and a concurrency cap: Tier 0 (first top-up)\n2,000 requests and 4 in flight","hierarchy":{"lvl0":"Getting Started","lvl1":"Umans AI Provider Guide","lvl2":"","lvl3":""}},
7230
7252
  {"objectID":"ff607cfd31ffc8da0d09dd8b747b2825462342a8683ab3c6cd35251b2538b0ef","title":"Umans AI Provider Guide","url":"/docs/getting-started/providers/umans-ai#umans-ai-provider-guide","content":"Umans AI is a Tier-2 catalog provider: OpenAI-wire-compatible with no\nbehavioural quirks, so its entire integration is one JSON file\n(src/lib/providers/catalog/umans-ai.json) rather than hand-written code.\nThat file is the source of truth for everything on this page.\n\nVerification status: this entry is docs- and roster-verified, not yet\nlive-verified. Every field below comes from Umans AI's own public pages and\nunauthenticated GET requests — no account was created and no API key was\nused to build it. evidence.liveMatrix is null until someone runs the live\ncapability matrix with a real key (see\nVerification status below).","hierarchy":{"lvl0":"Getting Started","lvl1":"Umans AI Provider Guide","lvl2":"Umans AI Provider Guide","lvl3":""}},
7231
7253
  {"objectID":"b71e70c479a72c1b2c02508e7f57c9dc8f5bd366f7dab23c56b323d7dbb8e3ca","title":"Key Facts","url":"/docs/getting-started/providers/umans-ai#key-facts","content":"Provider id: umans-ai (alias umans)\nProtocol: OpenAI-compatible (/chat/completions) — the docs say \"Umans\n Code also implements the OpenAI Chat Completions API\"\n (docs). The same gateway also\n documents an Anthropic Messages route (POST /v1/messages) and a Responses\n route (POST /v1/responses); this entry uses only /chat/completions\nBase URL: https://api.code.umans.ai/v1\nDefault model: umans-coder — a routing alias, not a fixed model, see\n Models\nModels in catalog: 7 (curated from a 10-model live roster; three ids are\n left out on purpose — umans-deepseek-v4.1-flash-lab and\n umans-mimo-v2.6-pro-lab, which /v1/models/info describes as Labs\n experiments, \"temporary, not a permanent id\", with access \"seat-gated through\n the Labs page\", and umans-qwen3.6-35b-a3b, which it describes as a\n \"Technical alias for umans-flash\")\nStreaming: supported — the docs' OpenAI example sends \"stream\": true to\n POST /v1/chat/completions, and the reasoning section says reasoning \"streams\n live, token by token, alongside the answer\"\nTool calling: true — GET /v1/models/info reports supports_tools: true\n for every roster model. The docs show no tools request example for\n /v1/chat/completions, so this rests on that field alone and has not been\n exercised\nTools while streaming: not declared (false) — the docs show no\n tool-call request or response on this route, streaming or otherwise\nStructured output: not declared (false) — no page opened mentions\n response_format, json_schema or json_object\nStructured output + tools together: not declared (false) — no combined\n probe was run without credentials\nEmbeddings: not declared — the docs describe no embeddings route\nThinking: true — the docs' \"Reasoning & Extended Thinking\" section\n documents reasoning_effort on /v1/chat/completions and returns reasoning\n in a reasoning_content field, separate from content.","hierarchy":{"lvl0":"Getting Started","lvl1":"Umans AI Provider Guide","lvl2":"Key Facts","lvl3":""}},
@@ -7351,19 +7373,19 @@
7351
7373
  {"objectID":"7bc31a9a4e8ceaa9d384ce6953e41c6991344d5affeecf269a14248d769eb2f3","title":"\"xAI account has insufficient quota\"","url":"/docs/getting-started/providers/xai#xai-account-has-insufficient-quota","content":"Top up at https://console.x.ai/.","hierarchy":{"lvl0":"Getting Started","lvl1":"xAI Grok Provider Guide","lvl2":"\"xAI account has insufficient quota\"","lvl3":""}},
7352
7374
  {"objectID":"62c0c746ffa186a922276c344852f82e4ee3242f81c67b58a433defc4f4a4924","title":"\"Model not found\"","url":"/docs/getting-started/providers/xai#model-not-found","content":"Use one of the documented model IDs above. Custom fine-tunes are not\nexposed through the public API at this time.","hierarchy":{"lvl0":"Getting Started","lvl1":"xAI Grok Provider Guide","lvl2":"\"Model not found\"","lvl3":""}},
7353
7375
  {"objectID":"a55b6d892c11d17b5222d76096529452815cd15809b93c491a3c484756185910","title":"See Also","url":"/docs/getting-started/providers/xai#see-also","content":"Adding a new LLM provider — internal reference for the integration pattern this provider follows\nDeepSeek Provider — sibling OpenAI-compat provider with reasoning models\nGroq Provider — sibling OpenAI-compat provider with sub-100ms inference\n\nNeed Help? Open a GitHub Discussion or issue.","hierarchy":{"lvl0":"Getting Started","lvl1":"xAI Grok Provider Guide","lvl2":"See Also","lvl3":""}},
7354
- {"objectID":"03a891da2cc3529b3e4bfc042e5834454a779a2d3fa9389367ba4b3d05266871","title":"XOR Provider Guide","url":"/docs/getting-started/providers/xor","content":"XOR Provider Guide\n\nA provider of decide — the same typed boolean / choice /\nscore answers as TypeSafe's Jev and Laya, from\nJuspay's open-weights model, which also reads images and a video. It emits no\ntext at all.\n\nOverview\n\nXOR is Juspay's Apache-2.0 decision model, post-trained from Qwen3.6-35B-A3B;\nits model id is xor-1.1. You send one state plus named, typed questions, and\noptionally images or a video. XOR answers every question in a single batched\npass and returns a typed answer for each. NeuroLink calls whatever base URL you\nconfigure — a deployment of the model, or a LiteLLM proxy with a route to one.\nThere is no built-in endpoint. The weights and setup instructions are at\nhuggingface.co/juspay/xor.\n\nKey Facts\n\n| | |\n| --------------------- | -------------------------------------------------------------------- |\n| Inference type | decide only |\n| Model | xor-1.1 by default; XOR_MODEL sets the name your proxy uses |\n| Media | up to 8 images and one video; the whole request body at most 8 MB |\n| Input window | about 200,000 estimated tokens of state |\n| Questions per request | no cap in NeuroLink; a choice or score takes 2 to 255 options |\n| Endpoint | <base URL>/v1/systemone; the base URL is required, with no default |\n| Precedence | used automatically only when TypeSafe and Laya are not configured |\n\nQuick Start\nGet an endpoint and a key\n\nRun a deployment of XOR (see\nhuggingface.co/juspay/xor), or use a\nLiteLLM proxy with a route to one. NeuroLink posts to <base URL>/v1/systemone.\nThe base URL is the origin of the deployment or of the proxy route, and a\ntrailing /v1 is accepted. On a LiteLLM proxy the key's team must be allowed\nthe model you ask for (xor-1.1 by default); otherwise the proxy answers 403\nteam_model_access_denied. A proxy may serve XOR under a name of its own: set\nXOR_MODEL to that name.\nConfigure\n\nBoth the base URL and the key are required. Set them in the environment:\n\nor in the config passed to the SDK, exactly as for any other provider. Values\npassed per call override the constructor's, which override the environment:\nUse it\n\nFrom the CLI\n\nThe CLI reads the same environment variables:\n\nImages and video\n\nXOR reads images and one video alongside the state. TypeSafe and Laya do not\nread media, and Perplexity reads images but no video.\n\nFrom the CLI, pass --image <path> (repeatable) and --video <path>:\n\nThe rules:\nUp to 8 images and one video per request, in images and video.\nEach is a Buffer, a local file path or a data: URL. An http(s) URL is\n refused: NeuroLink does not fetch media for you.\nThe type comes from the bytes, not the file extension: PNG, JPEG, WebP\n and GIF images, and MP4, MOV and WebM video. NeuroLink\n sends a Buffer or a file to XOR as a data: URL. A data: URL you pass\n yourself must hold base64 image or video content, and is sent as given.\nThe whole request body may be at most 8 MB. The limit applies to the\n encoded body, so the base64 form counts, not the size of the files on disk. A\n file over the limit is refused from its size, before it is read.\nWhat can be checked locally is refused before any request, as a\n non-retryable invalid_request: a missing file, a directory, an empty Buffer\n or file, bytes that are not a recognised image or video, a remote URL or any\n other URL scheme, a string that is not a path or a data: URL, more than 8\n images, or a body over 8 MB. Bytes with the right signature that the model\n cannot decode are sent, and the server's refusal comes back as a server\n error, retried once.\nTypeSafe and Laya refuse media too, before any request, and the error\n names the providers that accept it. Perplexity refuses a video.\nImages and a video can be sent together (up to 8 images and one video),\n but the model does not reliably tell the two apart.\nThe result carries mediaBytes, the encoded size of the media sent. The\n decide spans carry decision.images.count and decision.media.bytes, and\n never any base64.\n\nWhen NeuroLink uses it\n\nEvery built-in consumer of decide — model routing, relevance-driven\ncompaction, tool routing and RAG planning — asks for the default decision\nprovider. That is the first one that is configured, in the environment or in the\ncredentials passed to the SDK, in the order TypeSafe, Laya, XOR,\nPerplexity. TypeSafe counts with either of its keys,\nTYPESAFE_API_KEY or AI_GATEWAY_API_KEY. Laya and XOR each count only with\nboth their key and their base URL. Perplexity counts with its key alone, which is\nshared with Perplexity's text provider. A caller can always name XOR with\nprovider: \"xor\". So:\nA TypeSafe key, or Laya's key and base URL, plus XOR's key and base URL:\n built-in features use TypeSafe or Laya, in that order; XOR runs only where a\n caller asks for prov","hierarchy":{"lvl0":"Getting Started","lvl1":"XOR Provider Guide","lvl2":"","lvl3":""}},
7376
+ {"objectID":"03a891da2cc3529b3e4bfc042e5834454a779a2d3fa9389367ba4b3d05266871","title":"XOR Provider Guide","url":"/docs/getting-started/providers/xor","content":"XOR Provider Guide\n\nA provider of decide — the same typed boolean / choice /\nscore answers as TypeSafe's Jev and Laya, from\nJuspay's open-weights model, which also reads images and a video. It emits no\ntext at all.\n\nOverview\n\nXOR is Juspay's Apache-2.0 decision model, post-trained from Qwen3.6-35B-A3B;\nits model id is xor-1.1. You send one state plus named, typed questions, and\noptionally images or a video. XOR answers every question in a single batched\npass and returns a typed answer for each. NeuroLink calls whatever base URL you\nconfigure — a deployment of the model, or a LiteLLM proxy with a route to one.\nThere is no built-in endpoint. The weights and setup instructions are at\nhuggingface.co/juspay/xor.\n\nKey Facts\n\n| | |\n| --------------------- | -------------------------------------------------------------------- |\n| Inference type | decide only |\n| Model | xor-1.1 by default; XOR_MODEL sets the name your proxy uses |\n| Media | up to 8 images and one video; the whole request body at most 8 MB |\n| Input window | about 200,000 estimated tokens of state |\n| Questions per request | no cap in NeuroLink; a choice or score takes 2 to 255 options |\n| Endpoint | <base URL>/v1/systemone; the base URL is required, with no default |\n| Precedence | used automatically only when TypeSafe and Laya are not configured |\n\nQuick Start\nGet an endpoint and a key\n\nRun a deployment of XOR (see\nhuggingface.co/juspay/xor), or use a\nLiteLLM proxy with a route to one. NeuroLink posts to <base URL>/v1/systemone.\nThe base URL is the origin of the deployment or of the proxy route, and a\ntrailing /v1 is accepted. On a LiteLLM proxy the key's team must be allowed\nthe model you ask for (xor-1.1 by default); otherwise the proxy answers 403\nteam_model_access_denied. A proxy may serve XOR under a name of its own: set\nXOR_MODEL to that name.\nConfigure\n\nBoth the base URL and the key are required. Set them in the environment:\n\nor in the config passed to the SDK, exactly as for any other provider. Values\npassed per call override the constructor's, which override the environment:\nUse it\n\nFrom the CLI\n\nThe CLI reads the same environment variables:\n\nImages and video\n\nXOR reads images and one video alongside the state. TypeSafe and Laya do not\nread media, and Perplexity and\nCloudflare Clef read images but no video.\n\nFrom the CLI, pass --image <path> (repeatable) and --video <path>:\n\nThe rules:\nUp to 8 images and one video per request, in images and video.\nEach is a Buffer, a local file path or a data: URL. An http(s) URL is\n refused: NeuroLink does not fetch media for you.\nThe type comes from the bytes, not the file extension: PNG, JPEG, WebP\n and GIF images, and MP4, MOV and WebM video. NeuroLink\n sends a Buffer or a file to XOR as a data: URL. A data: URL you pass\n yourself must hold base64 image or video content, and is sent as given.\nThe whole request body may be at most 8 MB. The limit applies to the\n encoded body, so the base64 form counts, not the size of the files on disk. A\n file over the limit is refused from its size, before it is read.\nWhat can be checked locally is refused before any request, as a\n non-retryable invalid_request: a missing file, a directory, an empty Buffer\n or file, bytes that are not a recognised image or video, a remote URL or any\n other URL scheme, a string that is not a path or a data: URL, more than 8\n images, or a body over 8 MB. Bytes with the right signature that the model\n cannot decode are sent, and the server's refusal comes back as a server\n error, retried once.\nTypeSafe and Laya refuse media too, before any request, and the error\n names the providers that accept it. Perplexity and Cloudflare Clef refuse a\n video.\nImages and a video can be sent together (up to 8 images and one video),\n but the model does not reliably tell the two apart.\nThe result carries mediaBytes, the encoded size of the media sent. The\n decide spans carry decision.images.count and decision.media.bytes, and\n never any base64.\n\nWhen NeuroLink uses it\n\nEvery built-in consumer of decide — model routing, relevance-driven\ncompaction, tool routing and RAG planning — asks for the default decision\nprovider. That is the first one that is configured, in the environment or in the\ncredentials passed to the SDK, in the order TypeSafe, Laya, XOR,\nPerplexity, then Cloudflare Clef.\nTypeSafe counts with either of its keys,\nTYPESAFE_API_KEY or AI_GATEWAY_API_KEY. Laya and XOR each count only with\nboth their key and their base URL. Perplexity counts with its key alone, which is\nshared with Perplexity's text provider. Cloudflare Clef counts only with both\nCLOUDFLARE_API_KEY and CLOUDFLARE_ACCOUNT_ID, the same two variables the\nWorkers AI text provider reads. A caller can always name","hierarchy":{"lvl0":"Getting Started","lvl1":"XOR Provider Guide","lvl2":"","lvl3":""}},
7355
7377
  {"objectID":"5ed0d5759a64d045ba6c8acb04b0b6e6fef6f00399b88df9847c053d52a8e70c","title":"XOR Provider Guide","url":"/docs/getting-started/providers/xor#xor-provider-guide","content":"A provider of decide — the same typed boolean / choice /\nscore answers as TypeSafe's Jev and Laya, from\nJuspay's open-weights model, which also reads images and a video. It emits no\ntext at all.","hierarchy":{"lvl0":"Getting Started","lvl1":"XOR Provider Guide","lvl2":"XOR Provider Guide","lvl3":""}},
7356
7378
  {"objectID":"62f1da4b03072e86ddddfcfed9098a9efa475d349ea26f88c901347eca7fc77a","title":"Overview","url":"/docs/getting-started/providers/xor#overview","content":"XOR is Juspay's Apache-2.0 decision model, post-trained from Qwen3.6-35B-A3B;\nits model id is xor-1.1. You send one state plus named, typed questions, and\noptionally images or a video. XOR answers every question in a single batched\npass and returns a typed answer for each. NeuroLink calls whatever base URL you\nconfigure — a deployment of the model, or a LiteLLM proxy with a route to one.\nThere is no built-in endpoint. The weights and setup instructions are at\nhuggingface.co/juspay/xor.","hierarchy":{"lvl0":"Getting Started","lvl1":"XOR Provider Guide","lvl2":"Overview","lvl3":""}},
7357
7379
  {"objectID":"2a68b1796658b1e42d1fb20868e2f28c9b9b13e9e83ebb2ac4411de95d15b8e3","title":"Key Facts","url":"/docs/getting-started/providers/xor#key-facts","content":"| | |\n| --------------------- | -------------------------------------------------------------------- |\n| Inference type | decide only |\n| Model | xor-1.1 by default; XOR_MODEL sets the name your proxy uses |\n| Media | up to 8 images and one video; the whole request body at most 8 MB |\n| Input window | about 200,000 estimated tokens of state |\n| Questions per request | no cap in NeuroLink; a choice or score takes 2 to 255 options |\n| Endpoint | <base URL>/v1/systemone; the base URL is required, with no default |\n| Precedence | used automatically only when TypeSafe and Laya are not configured |","hierarchy":{"lvl0":"Getting Started","lvl1":"XOR Provider Guide","lvl2":"Key Facts","lvl3":""}},
7358
7380
  {"objectID":"06cdd61cf369f5af93c8e03841be4ebddcc8e1fde819b193e35a9fabfa7db51a","title":"1. Get an endpoint and a key","url":"/docs/getting-started/providers/xor#1-get-an-endpoint-and-a-key","content":"Run a deployment of XOR (see\nhuggingface.co/juspay/xor), or use a\nLiteLLM proxy with a route to one. NeuroLink posts to <base URL>/v1/systemone.\nThe base URL is the origin of the deployment or of the proxy route, and a\ntrailing /v1 is accepted. On a LiteLLM proxy the key's team must be allowed\nthe model you ask for (xor-1.1 by default); otherwise the proxy answers 403\nteam_model_access_denied. A proxy may serve XOR under a name of its own: set\nXOR_MODEL to that name.","hierarchy":{"lvl0":"Getting Started","lvl1":"XOR Provider Guide","lvl2":"1. Get an endpoint and a key","lvl3":""}},
7359
7381
  {"objectID":"540db8a33720c64f9e3e59c312e3d3e2476799c726c6e5729729ecaf2bbefa05","title":"2. Configure","url":"/docs/getting-started/providers/xor#2-configure","content":"Both the base URL and the key are required. Set them in the environment:\n\nor in the config passed to the SDK, exactly as for any other provider. Values\npassed per call override the constructor's, which override the environment:","hierarchy":{"lvl0":"Getting Started","lvl1":"XOR Provider Guide","lvl2":"2. Configure","lvl3":""}},
7360
7382
  {"objectID":"3fd4afc9a2670046c937e06586a65b3b5eadb6de1560ee1e1a3abf5f05524876","title":"From the CLI","url":"/docs/getting-started/providers/xor#from-the-cli","content":"The CLI reads the same environment variables:","hierarchy":{"lvl0":"Getting Started","lvl1":"XOR Provider Guide","lvl2":"From the CLI","lvl3":""}},
7361
- {"objectID":"90a0f1f65868c0d5b68bb8fb5828a32f704d4ed3a164081275a6837b65a8e2f7","title":"Images and video","url":"/docs/getting-started/providers/xor#images-and-video","content":"XOR reads images and one video alongside the state. TypeSafe and Laya do not\nread media, and Perplexity reads images but no video.\n\nFrom the CLI, pass --image <path> (repeatable) and --video <path>:\n\nThe rules:\nUp to 8 images and one video per request, in images and video.\nEach is a Buffer, a local file path or a data: URL. An http(s) URL is\n refused: NeuroLink does not fetch media for you.\nThe type comes from the bytes, not the file extension: PNG, JPEG, WebP\n and GIF images, and MP4, MOV and WebM video. NeuroLink\n sends a Buffer or a file to XOR as a data: URL. A data: URL you pass\n yourself must hold base64 image or video content, and is sent as given.\nThe whole request body may be at most 8 MB. The limit applies to the\n encoded body, so the base64 form counts, not the size of the files on disk. A\n file over the limit is refused from its size, before it is read.\nWhat can be checked locally is refused before any request, as a\n non-retryable invalid_request: a missing file, a directory, an empty Buffer\n or file, bytes that are not a recognised image or video, a remote URL or any\n other URL scheme, a string that is not a path or a data: URL, more than 8\n images, or a body over 8 MB. Bytes with the right signature that the model\n cannot decode are sent, and the server's refusal comes back as a server\n error, retried once.\nTypeSafe and Laya refuse media too, before any request, and the error\n names the providers that accept it. Perplexity refuses a video.\nImages and a video can be sent together (up to 8 images and one video),\n but the model does not reliably tell the two apart.\nThe result carries mediaBytes, the encoded size of the media sent. The\n decide spans carry decision.images.count and decision.media.bytes, and\n never any base64.","hierarchy":{"lvl0":"Getting Started","lvl1":"XOR Provider Guide","lvl2":"Images and video","lvl3":""}},
7362
- {"objectID":"4fea1eacec247a5a95e9c0d017549cb5233da507c593a841ce0e392eb2d11e2a","title":"When NeuroLink uses it","url":"/docs/getting-started/providers/xor#when-neurolink-uses-it","content":"Every built-in consumer of decide — model routing, relevance-driven\ncompaction, tool routing and RAG planning — asks for the default decision\nprovider. That is the first one that is configured, in the environment or in the\ncredentials passed to the SDK, in the order TypeSafe, Laya, XOR,\nPerplexity. TypeSafe counts with either of its keys,\nTYPESAFE_API_KEY or AI_GATEWAY_API_KEY. Laya and XOR each count only with\nboth their key and their base URL. Perplexity counts with its key alone, which is\nshared with Perplexity's text provider. A caller can always name XOR with\nprovider: \"xor\". So:\nA TypeSafe key, or Laya's key and base URL, plus XOR's key and base URL:\n built-in features use TypeSafe or Laya, in that order; XOR runs only where a\n caller asks for provider: \"xor\".\nXOR's key and base URL, plus a Perplexity key: built-in features use XOR;\n Perplexity runs only where a caller asks for\n provider: \"perplexity-decider\".\nOnly XOR's key and base URL: built-in features use XOR.\nAn XOR key with no base URL: XOR is not configured. Built-in features\n ignore it, and provider: \"xor\" fails with XOR requires a base URL.\nNone of TypeSafe, Laya, XOR or Perplexity: everything behaves exactly as\n it did without a decision model.","hierarchy":{"lvl0":"Getting Started","lvl1":"XOR Provider Guide","lvl2":"When NeuroLink uses it","lvl3":""}},
7383
+ {"objectID":"90a0f1f65868c0d5b68bb8fb5828a32f704d4ed3a164081275a6837b65a8e2f7","title":"Images and video","url":"/docs/getting-started/providers/xor#images-and-video","content":"XOR reads images and one video alongside the state. TypeSafe and Laya do not\nread media, and Perplexity and\nCloudflare Clef read images but no video.\n\nFrom the CLI, pass --image <path> (repeatable) and --video <path>:\n\nThe rules:\nUp to 8 images and one video per request, in images and video.\nEach is a Buffer, a local file path or a data: URL. An http(s) URL is\n refused: NeuroLink does not fetch media for you.\nThe type comes from the bytes, not the file extension: PNG, JPEG, WebP\n and GIF images, and MP4, MOV and WebM video. NeuroLink\n sends a Buffer or a file to XOR as a data: URL. A data: URL you pass\n yourself must hold base64 image or video content, and is sent as given.\nThe whole request body may be at most 8 MB. The limit applies to the\n encoded body, so the base64 form counts, not the size of the files on disk. A\n file over the limit is refused from its size, before it is read.\nWhat can be checked locally is refused before any request, as a\n non-retryable invalid_request: a missing file, a directory, an empty Buffer\n or file, bytes that are not a recognised image or video, a remote URL or any\n other URL scheme, a string that is not a path or a data: URL, more than 8\n images, or a body over 8 MB. Bytes with the right signature that the model\n cannot decode are sent, and the server's refusal comes back as a server\n error, retried once.\nTypeSafe and Laya refuse media too, before any request, and the error\n names the providers that accept it. Perplexity and Cloudflare Clef refuse a\n video.\nImages and a video can be sent together (up to 8 images and one video),\n but the model does not reliably tell the two apart.\nThe result carries mediaBytes, the encoded size of the media sent. The\n decide spans carry decision.images.count and decision.media.bytes, and\n never any base64.","hierarchy":{"lvl0":"Getting Started","lvl1":"XOR Provider Guide","lvl2":"Images and video","lvl3":""}},
7384
+ {"objectID":"4fea1eacec247a5a95e9c0d017549cb5233da507c593a841ce0e392eb2d11e2a","title":"When NeuroLink uses it","url":"/docs/getting-started/providers/xor#when-neurolink-uses-it","content":"Every built-in consumer of decide — model routing, relevance-driven\ncompaction, tool routing and RAG planning — asks for the default decision\nprovider. That is the first one that is configured, in the environment or in the\ncredentials passed to the SDK, in the order TypeSafe, Laya, XOR,\nPerplexity, then Cloudflare Clef.\nTypeSafe counts with either of its keys,\nTYPESAFE_API_KEY or AI_GATEWAY_API_KEY. Laya and XOR each count only with\nboth their key and their base URL. Perplexity counts with its key alone, which is\nshared with Perplexity's text provider. Cloudflare Clef counts only with both\nCLOUDFLARE_API_KEY and CLOUDFLARE_ACCOUNT_ID, the same two variables the\nWorkers AI text provider reads. A caller can always name XOR with\nprovider: \"xor\". So:\nA TypeSafe key, or Laya's key and base URL, plus XOR's key and base URL:\n built-in features use TypeSafe or Laya, in that order; XOR runs only where a\n caller asks for provider: \"xor\".\nXOR's key and base URL, plus a Perplexity key: built-in features use XOR;\n Perplexity runs only where a caller asks for\n provider: \"perplexity-decider\".\nXOR's key and base URL, plus CLOUDFLARE_API_KEY and\n CLOUDFLARE_ACCOUNT_ID: built-in features use XOR; Cloudflare Clef runs\n only where a caller asks for provider: \"cloudflare-clef\".\nOnly XOR's key and base URL: built-in features use XOR.\nAn XOR key with no base URL: XOR is not configured. Built-in features\n ignore it, and provider: \"xor\" fails with XOR requires a base URL.\nNone of TypeSafe, Laya, XOR, Perplexity or Cloudflare Clef: everything\n behaves exactly as it did without a decision model.","hierarchy":{"lvl0":"Getting Started","lvl1":"XOR Provider Guide","lvl2":"When NeuroLink uses it","lvl3":""}},
7363
7385
  {"objectID":"20a784de9dff5b3afa1f5ee7c0c60e268a7274b7e0c2a447aca431642f8fcabc","title":"Limits","url":"/docs/getting-started/providers/xor#limits","content":"The window is large, but shared. NeuroLink allows about 200,000 estimated\ntokens of state, and refuses more before any network call, with\nmax_tokens_exceeded. The size is an estimate, not XOR's tokenizer: about four\ncharacters per token for ASCII text, and one token per character for other\nscripts, which errs toward refusing. The 200,000 figure is a conservative\ndefault under the deployment's 250,000-token prefill, and that prefill is shared\nby the state, the questions and any media.\n\nThe figures in this paragraph are measurements of one live LiteLLM route, not\nlimits of XOR: the model behind the route's name has not been confirmed as XOR,\nand its 262,144-token context length differs from the 250,000-token prefill\nabove. They were taken by sending requests to the route directly rather than\nthrough NeuroLink, which refuses a state of 1.2 million ASCII characters before\nany request, at about 315,000 estimated tokens. A state of about 215,000 tokens\n(1.2 million characters of English) was accepted, and one of about 430,000\ntokens was refused with a 500 whose text says the input is longer than the\nmodel's context length. The estimate is not exact in either direction. English\naveraged about 5.6 characters per token on the route, so the estimate\nover-counts it. Dense JSON averaged about 1.8, so the estimate under-counts it\nby more than half, and a dense state estimated under the limit can still be\nlonger than the server's context. A structured (non-string) state of 100,000\nChinese characters passed NeuroLink's local check and was refused by the\nserver, which counted it as 485,774 tokens.\n\nA state near the limit, or a lot of media, can therefore still be refused by the\nserver as too long. A 413, or a 5xx other than 503 whose message says the\ncontext was too long, arrives as max_tokens_exceeded and is not retried; every\nother reply is classified as in the Errors table below. Built-in consumers treat\na refusal as \"carry on as before\".\n\nA proxy may cap parallel requests.","hierarchy":{"lvl0":"Getting Started","lvl1":"XOR Provider Guide","lvl2":"Limits","lvl3":""}},
7364
7386
  {"objectID":"d3945a40a2095e352b8f5f71228a59bbcb1cb60823696dbe20f3fafffe6b7c29","title":"Errors","url":"/docs/getting-started/providers/xor#errors","content":"| Reply | Kind | Retried |\n| ------------------------------------------------------------ | --------------------------------------------------------- | --------- |\n| 401 | authentication — the provider instance stops retrying | no |\n| 403 / 402 | invalid_request — fixable; the instance is not disabled | no |\n| 413, or a 5xx other than 503 saying the context was exceeded | max_tokens_exceeded | no |\n| any other 4xx (for example 422, 404) | invalid_request | no |\n| 429 | rate_limit | yes, once |\n| 503 | overloaded | yes, once |\n| other 5xx | server | yes, once |\n| state over the window (local) | max_tokens_exceeded, with no network call | no |\n| unusable media (local) | invalid_request, with no network call | no |\n| no base URL configured (local) | invalid_request, with no network call | no |\n\nA 403 or 402 is deliberately not authentication. On a LiteLLM proxy a 403\nmeans the key's team does not allow the model you asked for\n(team_model_access_denied), and\na 402 means the team has no budget. An admin can fix either, so the provider\ninstance keeps working once it is fixed and is not disabled.","hierarchy":{"lvl0":"Getting Started","lvl1":"XOR Provider Guide","lvl2":"Errors","lvl3":""}},
7365
7387
  {"objectID":"37466cda852e51cad898009d283cae74e18a20774a3ca0fad9f146e3efb8a5a9","title":"Troubleshooting","url":"/docs/getting-started/providers/xor#troubleshooting","content":"XOR requires a base URL — set XOR_BASE_URL, or pass\n credentials.xor.baseURL. There is no default endpoint.\nThe XOR base URL must not carry credentials… — the base URL has a user\n name, a password, a query string or a fragment. Set it to the origin only, for\n example https://your-proxy.example.com, and give the key through\n XOR_API_KEY or credentials.xor.apiKey.\nThe XOR base URL must start with https:// or http:// — the value has no\n scheme (for example xor.internal:8080) or uses another scheme. Add\n https:// or http://.\nThe XOR base URL is not a valid absolute URL — the value cannot be\n parsed as a URL. Set it to an absolute origin.\nXOR requires an API key — set XOR_API_KEY, or pass\n credentials.xor.apiKey.\nA 401 — the endpoint rejected the key. The provider instance stops\n retrying after a rejection, so fix the key and construct a new one.\nA 403 with team_model_access_denied — the LiteLLM proxy's team for this\n key does not allow the model you asked for (XOR_MODEL, or xor-1.1 by\n default). Add the model to the team; the instance works again without a\n restart.\nmax_tokens_exceeded — the state is over about 200,000 estimated tokens,\n or the state, the questions and the media together are more than the\n deployment's prefill. Shorten the state, or send fewer or smaller images or a\n shorter video.\nImages seem to be ignored — a deployment started without\n OPENJEV_IMAGES=1 answers 200 and silently ignores images, and NeuroLink\n cannot detect that. Send a red image and a blue image and check that the two\n answers differ. Video is not affected.\nBuilt-in routing never uses XOR — a TypeSafe key, or Laya's key and base\n URL, is also set and takes precedence, or XOR has no base URL.","hierarchy":{"lvl0":"Getting Started","lvl1":"XOR Provider Guide","lvl2":"Troubleshooting","lvl3":""}},
7366
- {"objectID":"cdd6e7c83dfa23afa5cf0bb7d23081881600f72a5ac64ca56d383736c209f285","title":"See also","url":"/docs/getting-started/providers/xor#see-also","content":"TypeSafe (Jev) Provider Guide\nLaya Provider Guide\nPerplexity Decisions Provider Guide\nThe decide inference type\nXOR on Hugging Face","hierarchy":{"lvl0":"Getting Started","lvl1":"XOR Provider Guide","lvl2":"See also","lvl3":""}},
7388
+ {"objectID":"cdd6e7c83dfa23afa5cf0bb7d23081881600f72a5ac64ca56d383736c209f285","title":"See also","url":"/docs/getting-started/providers/xor#see-also","content":"TypeSafe (Jev) Provider Guide\nLaya Provider Guide\nPerplexity Decisions Provider Guide\nCloudflare Clef Provider Guide\nThe decide inference type\nXOR on Hugging Face","hierarchy":{"lvl0":"Getting Started","lvl1":"XOR Provider Guide","lvl2":"See also","lvl3":""}},
7367
7389
  {"objectID":"b62066defb7f4f21e315563cc367d3573fb4aacd3fa38685f76f1a3b110c0437","title":"Z.AI Provider Guide","url":"/docs/getting-started/providers/z-ai","content":"Z.AI Provider Guide\n\nZ.AI is a Tier-2 catalog provider: OpenAI-wire-compatible, so its entire\nintegration is one JSON file (src/lib/providers/catalog/z-ai.json) rather\nthan hand-written code. That file is the source of truth for everything on this\npage.\n\nVerification status: this entry is docs-verified only, not yet\nlive-verified. The vendor's GET https://api.z.ai/api/paas/v4/models needs\na key (it answers HTTP 401 without one), so the model ids come from Z.AI's\npublic docs and the roster has not been checked. No account was created and no\nAPI key was used to build this entry. evidence.liveMatrix is null until\nsomeone runs the live capability matrix with a real key (see\nVerification status below).\n\nKey Facts\nProvider id: z-ai (alias zai)\nProtocol: OpenAI-compatible (/chat/completions). The\n OpenAI Python SDK page\n says \"Z.AI provides interfaces compatible with OpenAI API\" and adds \"In some\n scenarios, there are still differences between Z.AI and OpenAI interfaces, but\n this does not affect overall compatibility.\"\nBase URL: https://api.z.ai/api/paas/v4 — the\n Introduction page calls it\n \"Z.ai Platform's general API endpoint\". Its warning says GLM Coding Plan users\n follow the Coding Plan tutorial to configure a dedicated endpoint, and the\n Coding Plan Endpoint Guide lists\n https://api.z.ai/api/coding/paas/v4 for OpenAI Chat Completions. This entry\n uses the general endpoint; Z_AI_BASE_URL overrides it.\nDefault model: glm-5.3 — the default of the model parameter in the\n Chat Completion reference\n and the model in its request examples\nModels in catalog: 12, curated from the 21 ids the Chat Completion\n reference lists (14 in its text-model enum, 7 in its vision-model enum). The\n nine ids left out to stay inside the 12-model cap are glm-4.5-x,\n glm-4.5-airx, glm-4.5-flash, glm-4-32b-0414-128k, glm-4.6v,\n glm-4.6v-flash, glm-4.6v-flashx, glm-4.5v and\n autoglm-phone-multilingual\nStreaming: supported — the\n Streaming Messages page\n documents stream=True with Server-Sent Events, and the Chat Completion\n reference says the Event Stream ends with a data: [DONE] message\nTool calling: model-dependent — the\n Function Calling page\n documents tools and a returned tool_calls array, and the Chat Completion\n reference says of tool_choice \"The default value is auto, and only auto is\n supported.\" For vision models it says tools are \"Only supported by\n GLM-5.3-Flash series, the GLM-4.6V series, and autoglm-phone-multilingual\"\nTools while streaming: supported (true) — the\n Thinking Mode page\n gives an Interleaved Thinking + Tool Calling example that sends tools with\n stream=True and reads tool_calls from the streamed deltas. The\n Tool Streaming Output page\n documents a separate tool_stream parameter, which this entry does not set.\n The Thinking Mode page also says \"thinking blocks should be explicitly\n preserved and returned together with the tool results.\", and its example\n appends an assistant message carrying reasoning_content and tool_calls\n before the tool message; this entry does not set\n quirks.replayReasoningContent\nStructured output: supported — the\n Structured Output page\n documents response_format set to {\"type\": \"json_object\"}. The Chat\n Completion reference lists text and json_object as the response_format\n type values, so the entry sets quirks.responseFormatDowngrade to send\n json_object for schema requests; generate({ schema }) still validates the\n result client-side. No request with json_schema was sent. The Chat\n Completion reference says of response_format \"Only text models support this\n field.\", and the\n GLM-5.3-Flash/FlashX page\n lists Structured Output among its Capabilities (\"Supports structured output\n formats such as JSON for seamless system integration.\")\nStructured output + tools together: not declared (false) — no combined\n probe was possible without credentials\nVision: glm-5.3-flash and glm-5.3-flashx — the\n GLM-5.3-Flash/FlashX page\n shows Input Modality \"Video / Image / Text / File\". The GLM-5.3 page says\n \"text-only inputs\", and the pages for the other nine catalog models show\n Input Modalities \"Text\"\nEmbeddings: not declared on this catalog entry\nThinking: not declared (false) — the entry sets none of the thinking\n or reasoning_effort request parameters. The Chat Completion reference\n documents both, and the\n GLM-5.3 page says \"GLM-5.3 always\n operates with reasoning enabled\"\nBilling: schema value free-tier, see Billing below\nKey format: none declared. The Introduction page shows the header\n Authorization: Bearer ZAI_API_KEY\n\nQuick Start\nGet an API key\nVisit: https://z.ai/manage-apikey/apikey-list — the Quick Start says to access the Z.AI Open Platform (https://z.ai/model-api) and \"Register or Login.\", then to \"Create an API Key\" in the API Keys management page and \"Copy your API Key for use.\"\nAuthentication is HTTP Bearer: the Introduction page (https://docs.z.ai/api-reference/introduction.md) shows the header Authorization: Bearer ZAIAPIKEY\n","hierarchy":{"lvl0":"Getting Started","lvl1":"Z.AI Provider Guide","lvl2":"","lvl3":""}},
7368
7390
  {"objectID":"ca9f10745365c71403f5bcba388684ac70aeb8f48c3d7b3fac85b832a4eead37","title":"Z.AI Provider Guide","url":"/docs/getting-started/providers/z-ai#zai-provider-guide","content":"Z.AI is a Tier-2 catalog provider: OpenAI-wire-compatible, so its entire\nintegration is one JSON file (src/lib/providers/catalog/z-ai.json) rather\nthan hand-written code. That file is the source of truth for everything on this\npage.\n\nVerification status: this entry is docs-verified only, not yet\nlive-verified. The vendor's GET https://api.z.ai/api/paas/v4/models needs\na key (it answers HTTP 401 without one), so the model ids come from Z.AI's\npublic docs and the roster has not been checked. No account was created and no\nAPI key was used to build this entry. evidence.liveMatrix is null until\nsomeone runs the live capability matrix with a real key (see\nVerification status below).","hierarchy":{"lvl0":"Getting Started","lvl1":"Z.AI Provider Guide","lvl2":"Z.AI Provider Guide","lvl3":""}},
7369
7391
  {"objectID":"d509ecad3c62c73f008d7e2f88de10168c19308f90c63101973f3ae4983a7318","title":"Key Facts","url":"/docs/getting-started/providers/z-ai#key-facts","content":"Provider id: z-ai (alias zai)\nProtocol: OpenAI-compatible (/chat/completions). The\n OpenAI Python SDK page\n says \"Z.AI provides interfaces compatible with OpenAI API\" and adds \"In some\n scenarios, there are still differences between Z.AI and OpenAI interfaces, but\n this does not affect overall compatibility.\"\nBase URL: https://api.z.ai/api/paas/v4 — the\n Introduction page calls it\n \"Z.ai Platform's general API endpoint\". Its warning says GLM Coding Plan users\n follow the Coding Plan tutorial to configure a dedicated endpoint, and the\n Coding Plan Endpoint Guide lists\n https://api.z.ai/api/coding/paas/v4 for OpenAI Chat Completions. This entry\n uses the general endpoint; Z_AI_BASE_URL overrides it.\nDefault model: glm-5.3 — the default of the model parameter in the\n Chat Completion reference\n and the model in its request examples\nModels in catalog: 12, curated from the 21 ids the Chat Completion\n reference lists (14 in its text-model enum, 7 in its vision-model enum). The\n nine ids left out to stay inside the 12-model cap are glm-4.5-x,\n glm-4.5-airx, glm-4.5-flash, glm-4-32b-0414-128k, glm-4.6v,\n glm-4.6v-flash, glm-4.6v-flashx, glm-4.5v and\n autoglm-phone-multilingual\nStreaming: supported — the\n Streaming Messages page\n documents stream=True with Server-Sent Events, and the Chat Completion\n reference says the Event Stream ends with a data: [DONE] message\nTool calling: model-dependent — the\n Function Calling page\n documents tools and a returned tool_calls array, and the Chat Completion\n reference says of tool_choice \"The default value is auto, and only auto is\n supported.\" For vision models it says tools are \"Only supported by\n GLM-5.3-Flash series, the GLM-4.6V series, and autoglm-phone-multilingual\"\nTools while streaming: supported (true) — the\n Thinking Mode page\n gives an Interleaved Thinking + Tool Calling example that sends tools with\n stream=True and reads tool_calls from the streamed deltas.","hierarchy":{"lvl0":"Getting Started","lvl1":"Z.AI Provider Guide","lvl2":"Key Facts","lvl3":""}},
@@ -8588,7 +8610,7 @@
8588
8610
  {"objectID":"95f574ed3e7979eee5be188ccd16efb2639e30e8258e55366d7ccd344bbc3126","title":"Related Documentation","url":"/docs/implementation-guides/14-rag-document-processing#related-documentation","content":"Vector Store Integrations\nEvaluation and Scoring\nMaster Implementation Guide","hierarchy":{"lvl0":"Implementation Guides","lvl1":"RAG Document Processing - Implementation Guide","lvl2":"Related Documentation","lvl3":""}},
8589
8611
  {"objectID":"60877eba17c1fe5c9fda2100a737f42fddb5c4c7083297e61a66884dbe5c326e","title":"NeuroLink","url":"/docs/","content":"🧠 NeuroLink\n The Pipe Layer of an AI Nervous System\n Provider Neurons Across Major AI Vendors | 3 Inference Types (generate Ā· stream Ā· decide) | Voice (TTS/STT/Realtime) | 58+ MCP Servers | HITL Security | Redis Persistence\n\nNeuroLink is the pipe layer of an AI nervous system: one interface connecting provider neurons — major AI vendors and local runtimes — to the applications that consume them. Built-in tooling and an opinionated factory architecture mean adding a new provider, or a new capability, never touches application code. NeuroLink ships as both a TypeScript SDK and a professional CLI so teams can build, operate, and iterate on AI features quickly.\n\n🧠 What is NeuroLink?\n\nNeuroLink is the pipe layer of an AI nervous system. Providers — OpenAI, Anthropic, Google, AWS, Azure, DeepSeek, NVIDIA NIM, local runtimes like Ollama and llama.cpp, and dozens more — are the neurons: each generates a different kind of intelligence, at a different cost and latency. NeuroLink is the vascular layer that carries that intelligence, as a stream, to the applications that consume it, across three inference types: generate and stream produce text, decide produces a calibrated boolean/choice/score judgment instead.\n\nExtracted from production systems at Juspay, NeuroLink provides a practical, TypeScript-first way to plug any application into that nervous system. Switch which neuron answers a request with a single parameter change — any provider you're building with, or any provider you add.\n\nWhy NeuroLink? Three genuine inference types, not one dressed up three ways — generate and stream produce text, while decide returns a typed, calibrated judgment (boolean / choice / score) with no text at all, for the routing and gating decisions the other two were never meant to make. Every neuron plugs into the same pipe. Switch providers with a single parameter change, leverage 64+ built-in tools and MCP servers, deploy with confidence using enterprise features like Redis memory and multi-provider failover, and optimize costs automatically with intelligent routing. Use it via our professional CLI or TypeScript SDK—whichever fits your workflow.\n\nWhere we're headed: We're building for the future of AI—edge-first execution and continuous streaming architectures that make AI practically free and universally available. Read our vision →\n\nGet Started in \\ Observability Guide\nServer Adapters -- Deploy NeuroLink as an HTTP API server with your framework of choice (Hono, Express, Fastify, Koa). Full CLI support with serve and server commands for foreground/background modes, route management, and OpenAPI generation. -> Server Adapters Guide\nTitle Generation Events -- Emit real-time events when conversation titles are auto-generated. Listen to conversation:titleGenerated for session tracking. -> Conversation Memory Guide\nCustom Title Prompts -- Customize conversation title generation with NEUROLINK_TITLE_PROMPT environment variable. Use ${userMessage} placeholder for dynamic prompts. -> Conversation Memory Guide\nVideo Generation -- Transform images into 8-second videos with synchronized audio using Google Veo 3.1 via Vertex AI. Supports 720p/1080p resolutions, portrait/landscape aspect ratios. -> Video Generation Guide\nImage Generation -- Generate images from text prompts using Gemini models via Vertex AI or Google AI Studio. Supports streaming mode with automatic file saving. -> Image Generation Guide\nHTTP/Streamable HTTP Transport for MCP -- Connect to remote MCP servers via HTTP with authentication headers, retry logic, and rate limiting. -> HTTP Transport Guide\nClaude Subscription (OAuth) Support -- Use your Claude Pro/Max/Team subscription with NeuroLink via OAuth authentication, no API key required. -> Subscription Guide\nGemini 3 Preview Support - Full support for gemini-3-flash-preview and gemini-3-pro-preview with extended thinking capabilities\nStructured Output with Zod Schemas -- Type-safe JSON generation with automatic validation using schema + output.format: \"json\" in generate(). -> Structured Output Guide\nCSV File Support -- Attach CSV files to prompts for AI-powered data analysis with auto-detection. -> CSV Guide\nPDF File Support -- Process PDF documents with native visual analysis for Vertex AI, Anthropic, Bedrock, AI Studio. -> PDF Guide\n50+ File Types -- Process Excel, Word, RTF, JSON, YAML, XML, HTML, SVG, Markdown, and 50+ code languages with intelligent content extraction. -> File Processors Guide\nLiteLLM Integration -- Access 100+ AI models across a broad range of AI providers through unified interface. -> Setup Guide\nSageMaker Integration -- Deploy and use custom trained models on AWS infrastructure. -> Setup Guide\nOpenRouter Integration -- Access 300+ models from OpenAI, Anthropic, Google, Meta, and more through a single unified API. -> Setup Guide\nHuman-in-the-loop workflows -- Pause generation for user approval/input before tool execution. -> HITL Guide\nGuardrails middleware -- Block PII, profanity, and unsafe cont","hierarchy":{"lvl0":"Docs","lvl1":"NeuroLink","lvl2":"","lvl3":""}},
8590
8612
  {"objectID":"7bd402837d03dbfbb9586674bc914cfc6ee7691e33540eeabe6555363697375e","title":"🧠 What is NeuroLink?","url":"/docs/#-what-is-neurolink","content":"NeuroLink is the pipe layer of an AI nervous system. Providers — OpenAI, Anthropic, Google, AWS, Azure, DeepSeek, NVIDIA NIM, local runtimes like Ollama and llama.cpp, and dozens more — are the neurons: each generates a different kind of intelligence, at a different cost and latency. NeuroLink is the vascular layer that carries that intelligence, as a stream, to the applications that consume it, across three inference types: generate and stream produce text, decide produces a calibrated boolean/choice/score judgment instead.\n\nExtracted from production systems at Juspay, NeuroLink provides a practical, TypeScript-first way to plug any application into that nervous system. Switch which neuron answers a request with a single parameter change — any provider you're building with, or any provider you add.\n\nWhy NeuroLink? Three genuine inference types, not one dressed up three ways — generate and stream produce text, while decide returns a typed, calibrated judgment (boolean / choice / score) with no text at all, for the routing and gating decisions the other two were never meant to make. Every neuron plugs into the same pipe. Switch providers with a single parameter change, leverage 64+ built-in tools and MCP servers, deploy with confidence using enterprise features like Redis memory and multi-provider failover, and optimize costs automatically with intelligent routing. Use it via our professional CLI or TypeScript SDK—whichever fits your workflow.\n\nWhere we're headed: We're building for the future of AI—edge-first execution and continuous streaming architectures that make AI practically free and universally available. Read our vision →\n\nGet Started in \\<5 Minutes →","hierarchy":{"lvl0":"Docs","lvl1":"NeuroLink","lvl2":"🧠 What is NeuroLink?","lvl3":""}},
8591
- {"objectID":"293082e7b04b9b0eb186f10fec40eefc1996d8eab9becff9e2b428731642291f","title":"What's New (Q1 2026)","url":"/docs/#whats-new-q1-2026","content":"| Feature | Version | Description | Guide |\n| ---------------------------------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |","hierarchy":{"lvl0":"Docs","lvl1":"NeuroLink","lvl2":"What's New (Q1 2026)","lvl3":""}},
8613
+ {"objectID":"293082e7b04b9b0eb186f10fec40eefc1996d8eab9becff9e2b428731642291f","title":"What's New (Q1 2026)","url":"/docs/#whats-new-q1-2026","content":"| Feature | Version | Description | Guide |\n| ---------------------------------- | ------- |","hierarchy":{"lvl0":"Docs","lvl1":"NeuroLink","lvl2":"What's New (Q1 2026)","lvl3":""}},
8592
8614
  {"objectID":"3e5ecc512d348470d92717ef33a2ff5067ffdd99c7ea212613df7986db133390","title":"Enterprise Security: Human-in-the-Loop (HITL)","url":"/docs/#enterprise-security-human-in-the-loop-hitl","content":"NeuroLink includes a HITL (Human-in-the-Loop) system for regulated industries and high-stakes AI operations:\n\n| Capability | Description | Use Case |\n| --------------------------- | ----------------------------------------------------------------------- | ------------------------------------------ |\n| Tool Approval Workflows | Require human approval before AI executes sensitive tools | Financial transactions, data modifications |\n| Output Validation | Route AI outputs through human review pipelines | Medical diagnosis, legal documents |\n| Confidence Thresholds | Automatically trigger human review below confidence level | Critical business decisions |\n| Complete Audit Trail | Audit logging to support your compliance program (HIPAA / SOC 2 / GDPR) | Regulated industries |\n\nEnterprise HITL Guide | Quick Start","hierarchy":{"lvl0":"Docs","lvl1":"NeuroLink","lvl2":"Enterprise Security: Human-in-the-Loop (HITL)","lvl3":""}},
8593
8615
  {"objectID":"c062d596f0f8940391d038563717712409b6055fdf70d92ab51cb35caf672663","title":"Get Started in Two Steps","url":"/docs/#get-started-in-two-steps","content":"`bash","hierarchy":{"lvl0":"Docs","lvl1":"NeuroLink","lvl2":"Get Started in Two Steps","lvl3":""}},
8594
8616
  {"objectID":"e9e7167316a23ea5fd5db23894b2e52405ba61f5ab57cc986d155490dd325674","title":"1. Run the interactive setup wizard (select providers, validate keys)","url":"/docs/#1-run-the-interactive-setup-wizard-select-providers-validate-keys","content":"pnpm dlx @juspay/neurolink setup","hierarchy":{"lvl0":"Docs","lvl1":"NeuroLink","lvl2":"1. Run the interactive setup wizard (select providers, validate keys)","lvl3":""}},
@@ -9650,13 +9672,13 @@
9650
9672
  {"objectID":"a68c16599df99491c1e91d0857d3b92f4bd78a3caa88b6279381ff1ca665926f","title":"Tier 2 (JSON catalog) providers no longer use a manifest here","url":"/docs/provider-integration/manifests/README#tier-2-json-catalog-providers-no-longer-use-a-manifest-here","content":"As of the provider-JSON-catalog refactor, Tier 2 providers are declared\nentirely in src/lib/providers/catalog/<id>.json, validated by the zod\nschema in src/lib/providers/catalog/schema.ts. That file's evidence\nobject — rosterVerified, addedInPR, and optionally authProbe,\nbillingProbe, liveMatrix — carries the same onboarding evidence a\nmanifest used to hold, so a separate manifest file would just duplicate\nit. tools/verify-provider-onboarding.ts reflects this: for any provider\nwith a matching src/lib/providers/catalog/<id>.json file, the gate\nchecks that the JSON file exists, parses via the real zod schema, and\n(via that same successful parse, since both fields are non-optional in\nthe schema) carries evidence.rosterVerified and evidence.addedInPR.\n\ncerebras.json and sambanova.json — the two manifests that used to\nlive in this directory — were removed for this reason: both providers\nare now JSON-catalog entries, and their onboarding evidence lives in\nsrc/lib/providers/catalog/cerebras.json and\nsrc/lib/providers/catalog/sambanova.json respectively.","hierarchy":{"lvl0":"Provider Integration","lvl1":"Provider Manifests","lvl2":"Tier 2 (JSON catalog) providers no longer use a manifest here","lvl3":""}},
9651
9673
  {"objectID":"00fb88ca42034f14d0a87eea1d06b38350fc2cb9b5bed8bb3f544f10b3cd99ab","title":"Tier 3/4 (hand-written) providers still use a manifest here","url":"/docs/provider-integration/manifests/README#tier-34-hand-written-providers-still-use-a-manifest-here","content":"A provider onboarded outside the JSON catalog — a custom adapter (Tier 3)\nor fully custom integration (Tier 4) — has no catalog JSON file, so\ntools/verify-provider-onboarding.ts falls back to its original\nfour-check flow for it, including a manifest at\ndocs/provider-integration/manifests/<name>.json. The shape below still\napplies to those providers.\n\nThe block is annotated JSONC for documentation purposes only — the\n// comments and trailing comma explain each field but are not valid\nJSON. A real <provider>.json manifest file must be strict JSON: no\ncomments, no trailing commas.","hierarchy":{"lvl0":"Provider Integration","lvl1":"Provider Manifests","lvl2":"Tier 3/4 (hand-written) providers still use a manifest here","lvl3":""}},
9652
9674
  {"objectID":"4384ec258a233739ea35584f79662d5111be0b8defd38f75a6ab8068bf3a7e30","title":"How it's checked","url":"/docs/provider-integration/manifests/README#how-its-checked","content":"pnpm run verify:provider-onboarding (tools/verify-provider-onboarding.ts)\nfails a PR that introduces a new AIProviderName member without matching\nonboarding evidence: a valid catalog JSON entry for Tier 2 providers (see\nabove), or a structurally valid manifest here for Tier 3/4 providers. A\nmanifest is structurally valid when it is a JSON object whose provider\nmatches the file name and which carries provider, tier, addedInPR,\naddedDate, filesTouched (an array of strings), mockedContractSection\nand manualTestStatus, plus tier4Justification when tier is 4. The gate\ndoes not retroactively require either for providers that predate the gate\n— see that tool's LEGACY_PROVIDERS list.","hierarchy":{"lvl0":"Provider Integration","lvl1":"Provider Manifests","lvl2":"How it's checked","lvl3":""}},
9653
- {"objectID":"c533e90a3190ff6ad955d7e0c25b0b6370956144cb2b59f1c07b4ffdb1b1e452","title":"Provider Descriptor Migration Ledger","url":"/docs/provider-integration/migration-ledger","content":"Provider Descriptor Migration Ledger\n\nInventory taken at origin/release @ 2cefa3ae4115f817f75a415b6bc70fc3ecaed2d3. TypeSafe was added afterwards (#1761) and is included below; the 25-entry count matched origin/release @ f536fd091.\n\nmistral, huggingface and deepseek have since migrated to the JSON catalog and been removed from HAND_DESCRIPTORS (#1781) — see Migrated below — and laya, xor and perplexity-decider (decision-only, like TypeSafe) were added as hand descriptors afterwards. Net effect: 25 minus the 3 migrated plus 3 (laya, xor, perplexity-decider) leaves 25 entries currently in HAND_DESCRIPTORS (src/lib/factories/providerDescriptors.ts). The category counts below (Shared adapter needed / Must remain class / Must remain core class) cover exactly those 25; the 3 migrated providers are recorded separately as done and no longer count toward \"what remains.\"\n\nThis ledger records, per provider, why it is (or isn't) a JSON-catalog migration candidate, so \"add a provider\" work doesn't re-litigate the same analysis per PR.\n\nEvery verdict allows one thing regardless of class: the static descriptor metadata (name, aliases, default model, credential env var names, setup URL) can always move into a class-backed JSON record. \"Must remain class\" means the execution — the inference loop (generate/stream/decide, per the provider's inferenceKinds), auth, media pipelines — cannot be reduced to declarative catalog data; it does not mean the provider is exempt from descriptor consolidation.\n\nMigrated (3)\n\nAll three ran the same class-removal path this ledger recommended below: resolve the JSON/descriptor data conflicts, then delete the hand descriptor and hand-written subclass so the provider is fully catalog-derived.\n\n| Provider | Note |\n| ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |\n| mistral | Moved onto catalog/mistral.json. defaultModel now derives from the new optional models.registryDefaultModel field when present (only Mistral sets it), and the setupUrl/key-format/timeout/priority conflicts this ledger flagged were resolved in the JSON rather than carried forward. |\n| huggingface | Moved onto catalog/huggingface.json, which gained wire.apiKeyFallbackEnvVars (HF_TOKEN alongside HUGGINGFACE_API_KEY) and capabilities.tools: \"model-dependent\" — the two schema gaps this ledger flagged below before the extension landed. |\n| deepseek | Moved onto catalog/deepseek.json, using the new quirks.responseFormatDowngrade: \"json-schema-to-json-object\" field (DeepSeek 400s on json_schema) instead of an executable hook — the named-quirk approach this ledger recommended. It also sets quirks.replayReasoningContent: true. |\n\nShared adapter needed (9)\n\nThree adapter families, not nine one-off migrations:\n\nLocal-runtime adapter (ollama, lm-studio, llamacpp) — one OpenAI-compatible local-runtime adapter with optional auth, model discovery/probe, actionable transport errors, timeout policy, optional embeddings. Vendor-specific pull/diagnostic guidance becomes adapter data.\n\nEmbedding-only adapter (voyage, jina, cohere) — one embedding-only provider adapter parameterized by path, request shape, batch size, response-extraction/index policy, unsupported-surface messages. jina extends it with an optional rerank operation profile (don't force reranking into the boolean capability shape). cohere composes this with a generic OpenAI chat adapter (it has both /compatibility/v1 chat and native /v2/embed); preserve its batch limit and body-field quirks in the embedding profile.\n\nImage-generation adapter (stability, ideogram, recraft) — one image-generation adapter family: explicit multipart vs JSON request profile, base64-vs-URL response profile, custom auth-header support, model-path mapping as data. SSRF and bounded-read policy remain shared mandatory behavior, not per-vendor opt-outs.\n\nMust remain class (15)\n\n| Provider | Why |\n| -------------------- | ---------------------------------------------------------------------------------","hierarchy":{"lvl0":"Provider Integration","lvl1":"Provider Descriptor Migration Ledger","lvl2":"","lvl3":""}},
9654
- {"objectID":"b5d36740e9c9cb41c6ccb01db0a9fb1f9c0ac8ac59f1084b9a92332dd79c3e4d","title":"Provider Descriptor Migration Ledger","url":"/docs/provider-integration/migration-ledger#provider-descriptor-migration-ledger","content":"Inventory taken at origin/release @ 2cefa3ae4115f817f75a415b6bc70fc3ecaed2d3. TypeSafe was added afterwards (#1761) and is included below; the 25-entry count matched origin/release @ f536fd091.\n\nmistral, huggingface and deepseek have since migrated to the JSON catalog and been removed from HAND_DESCRIPTORS (#1781) — see Migrated below — and laya, xor and perplexity-decider (decision-only, like TypeSafe) were added as hand descriptors afterwards. Net effect: 25 minus the 3 migrated plus 3 (laya, xor, perplexity-decider) leaves 25 entries currently in HAND_DESCRIPTORS (src/lib/factories/providerDescriptors.ts). The category counts below (Shared adapter needed / Must remain class / Must remain core class) cover exactly those 25; the 3 migrated providers are recorded separately as done and no longer count toward \"what remains.\"\n\nThis ledger records, per provider, why it is (or isn't) a JSON-catalog migration candidate, so \"add a provider\" work doesn't re-litigate the same analysis per PR.\n\nEvery verdict allows one thing regardless of class: the static descriptor metadata (name, aliases, default model, credential env var names, setup URL) can always move into a class-backed JSON record. \"Must remain class\" means the execution — the inference loop (generate/stream/decide, per the provider's inferenceKinds), auth, media pipelines — cannot be reduced to declarative catalog data; it does not mean the provider is exempt from descriptor consolidation.","hierarchy":{"lvl0":"Provider Integration","lvl1":"Provider Descriptor Migration Ledger","lvl2":"Provider Descriptor Migration Ledger","lvl3":""}},
9675
+ {"objectID":"c533e90a3190ff6ad955d7e0c25b0b6370956144cb2b59f1c07b4ffdb1b1e452","title":"Provider Descriptor Migration Ledger","url":"/docs/provider-integration/migration-ledger","content":"Provider Descriptor Migration Ledger\n\nInventory taken at origin/release @ 2cefa3ae4115f817f75a415b6bc70fc3ecaed2d3. TypeSafe was added afterwards (#1761) and is included below; the 25-entry count matched origin/release @ f536fd091.\n\nmistral, huggingface and deepseek have since migrated to the JSON catalog and been removed from HAND_DESCRIPTORS (#1781) — see Migrated below — and laya, xor, perplexity-decider and cloudflare-clef (decision-only, like TypeSafe) were added as hand descriptors afterwards. Net effect: 25 minus the 3 migrated plus 4 (laya, xor, perplexity-decider, cloudflare-clef) leaves 26 entries currently in HAND_DESCRIPTORS (src/lib/factories/providerDescriptors.ts). The category counts below (Shared adapter needed / Must remain class / Must remain core class) cover exactly those 26; the 3 migrated providers are recorded separately as done and no longer count toward \"what remains.\"\n\nThis ledger records, per provider, why it is (or isn't) a JSON-catalog migration candidate, so \"add a provider\" work doesn't re-litigate the same analysis per PR.\n\nEvery verdict allows one thing regardless of class: the static descriptor metadata (name, aliases, default model, credential env var names, setup URL) can always move into a class-backed JSON record. \"Must remain class\" means the execution — the inference loop (generate/stream/decide, per the provider's inferenceKinds), auth, media pipelines — cannot be reduced to declarative catalog data; it does not mean the provider is exempt from descriptor consolidation.\n\nMigrated (3)\n\nAll three ran the same class-removal path this ledger recommended below: resolve the JSON/descriptor data conflicts, then delete the hand descriptor and hand-written subclass so the provider is fully catalog-derived.\n\n| Provider | Note |\n| ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |\n| mistral | Moved onto catalog/mistral.json. defaultModel now derives from the new optional models.registryDefaultModel field when present (only Mistral sets it), and the setupUrl/key-format/timeout/priority conflicts this ledger flagged were resolved in the JSON rather than carried forward. |\n| huggingface | Moved onto catalog/huggingface.json, which gained wire.apiKeyFallbackEnvVars (HF_TOKEN alongside HUGGINGFACE_API_KEY) and capabilities.tools: \"model-dependent\" — the two schema gaps this ledger flagged below before the extension landed. |\n| deepseek | Moved onto catalog/deepseek.json, using the new quirks.responseFormatDowngrade: \"json-schema-to-json-object\" field (DeepSeek 400s on json_schema) instead of an executable hook — the named-quirk approach this ledger recommended. It also sets quirks.replayReasoningContent: true. |\n\nShared adapter needed (9)\n\nThree adapter families, not nine one-off migrations:\n\nLocal-runtime adapter (ollama, lm-studio, llamacpp) — one OpenAI-compatible local-runtime adapter with optional auth, model discovery/probe, actionable transport errors, timeout policy, optional embeddings. Vendor-specific pull/diagnostic guidance becomes adapter data.\n\nEmbedding-only adapter (voyage, jina, cohere) — one embedding-only provider adapter parameterized by path, request shape, batch size, response-extraction/index policy, unsupported-surface messages. jina extends it with an optional rerank operation profile (don't force reranking into the boolean capability shape). cohere composes this with a generic OpenAI chat adapter (it has both /compatibility/v1 chat and native /v2/embed); preserve its batch limit and body-field quirks in the embedding profile.\n\nImage-generation adapter (stability, ideogram, recraft) — one image-generation adapter family: explicit multipart vs JSON request profile, base64-vs-URL response profile, custom auth-header support, model-path mapping as data. SSRF and bounded-read policy remain shared mandatory behavior, not per-vendor opt-outs.\n\nMust remain class (16)\n\n| Provider | Why |\n| -------------------- | -----------------------------------------------","hierarchy":{"lvl0":"Provider Integration","lvl1":"Provider Descriptor Migration Ledger","lvl2":"","lvl3":""}},
9676
+ {"objectID":"b5d36740e9c9cb41c6ccb01db0a9fb1f9c0ac8ac59f1084b9a92332dd79c3e4d","title":"Provider Descriptor Migration Ledger","url":"/docs/provider-integration/migration-ledger#provider-descriptor-migration-ledger","content":"Inventory taken at origin/release @ 2cefa3ae4115f817f75a415b6bc70fc3ecaed2d3. TypeSafe was added afterwards (#1761) and is included below; the 25-entry count matched origin/release @ f536fd091.\n\nmistral, huggingface and deepseek have since migrated to the JSON catalog and been removed from HAND_DESCRIPTORS (#1781) — see Migrated below — and laya, xor, perplexity-decider and cloudflare-clef (decision-only, like TypeSafe) were added as hand descriptors afterwards. Net effect: 25 minus the 3 migrated plus 4 (laya, xor, perplexity-decider, cloudflare-clef) leaves 26 entries currently in HAND_DESCRIPTORS (src/lib/factories/providerDescriptors.ts). The category counts below (Shared adapter needed / Must remain class / Must remain core class) cover exactly those 26; the 3 migrated providers are recorded separately as done and no longer count toward \"what remains.\"\n\nThis ledger records, per provider, why it is (or isn't) a JSON-catalog migration candidate, so \"add a provider\" work doesn't re-litigate the same analysis per PR.\n\nEvery verdict allows one thing regardless of class: the static descriptor metadata (name, aliases, default model, credential env var names, setup URL) can always move into a class-backed JSON record. \"Must remain class\" means the execution — the inference loop (generate/stream/decide, per the provider's inferenceKinds), auth, media pipelines — cannot be reduced to declarative catalog data; it does not mean the provider is exempt from descriptor consolidation.","hierarchy":{"lvl0":"Provider Integration","lvl1":"Provider Descriptor Migration Ledger","lvl2":"Provider Descriptor Migration Ledger","lvl3":""}},
9655
9677
  {"objectID":"fb3b171594dd4553aaf04fd301572eb1db1e459c15541867bd93a67ece8cfadb","title":"Migrated (3)","url":"/docs/provider-integration/migration-ledger#migrated-3","content":"All three ran the same class-removal path this ledger recommended below: resolve the JSON/descriptor data conflicts, then delete the hand descriptor and hand-written subclass so the provider is fully catalog-derived.\n\n| Provider | Note |\n| ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |\n| mistral | Moved onto catalog/mistral.json. defaultModel now derives from the new optional models.registryDefaultModel field when present (only Mistral sets it), and the setupUrl/key-format/timeout/priority conflicts this ledger flagged were resolved in the JSON rather than carried forward. |\n| huggingface | Moved onto catalog/huggingface.json, which gained wire.apiKeyFallbackEnvVars (HF_TOKEN alongside HUGGINGFACE_API_KEY) and capabilities.tools: \"model-dependent\" — the two schema gaps this ledger flagged below before the extension landed. |\n| deepseek | Moved onto catalog/deepseek.json, using the new quirks.responseFormatDowngrade: \"json-schema-to-json-object\" field (DeepSeek 400s on json_schema) instead of an executable hook — the named-quirk approach this ledger recommended. It also sets quirks.replayReasoningContent: true. |","hierarchy":{"lvl0":"Provider Integration","lvl1":"Provider Descriptor Migration Ledger","lvl2":"Migrated (3)","lvl3":""}},
9656
9678
  {"objectID":"cbc4193e5e9ffb4c77dd167fe137bc62cfbab5adbe808e814a374670bef12175","title":"Shared adapter needed (9)","url":"/docs/provider-integration/migration-ledger#shared-adapter-needed-9","content":"Three adapter families, not nine one-off migrations:\n\nLocal-runtime adapter (ollama, lm-studio, llamacpp) — one OpenAI-compatible local-runtime adapter with optional auth, model discovery/probe, actionable transport errors, timeout policy, optional embeddings. Vendor-specific pull/diagnostic guidance becomes adapter data.\n\nEmbedding-only adapter (voyage, jina, cohere) — one embedding-only provider adapter parameterized by path, request shape, batch size, response-extraction/index policy, unsupported-surface messages. jina extends it with an optional rerank operation profile (don't force reranking into the boolean capability shape). cohere composes this with a generic OpenAI chat adapter (it has both /compatibility/v1 chat and native /v2/embed); preserve its batch limit and body-field quirks in the embedding profile.\n\nImage-generation adapter (stability, ideogram, recraft) — one image-generation adapter family: explicit multipart vs JSON request profile, base64-vs-URL response profile, custom auth-header support, model-path mapping as data. SSRF and bounded-read policy remain shared mandatory behavior, not per-vendor opt-outs.","hierarchy":{"lvl0":"Provider Integration","lvl1":"Provider Descriptor Migration Ledger","lvl2":"Shared adapter needed (9)","lvl3":""}},
9657
- {"objectID":"68866693ac3b73ea496e4c12c1b7d8fe2de069553e5d38707095387d8f39faba","title":"Must remain class (15)","url":"/docs/provider-integration/migration-ledger#must-remain-class-15","content":"| Provider | Why |\n| -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |\n| bedrock | AWS SigV4/Converse protocol, custom loop, not OpenAI-compatible. Tier 4. |\n| openai | Chat portion is catalog-shaped, but embeddings, image generation, and OpenAI-specific telemetry/errors are executable behavior. |\n| vertex | Two native protocols (Google GenAI + Anthropic-on-Vertex) and multiple modality pipelines behind one provider identity.","hierarchy":{"lvl0":"Provider Integration","lvl1":"Provider Descriptor Migration Ledger","lvl2":"Must remain class (15)","lvl3":""}},
9679
+ {"objectID":"b9d3b5b409f7d1545aa51c86379108fea5e57bbc3b19cb9f4aa352a00e7bc23b","title":"Must remain class (16)","url":"/docs/provider-integration/migration-ledger#must-remain-class-16","content":"| Provider | Why |\n| -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |\n| bedrock | AWS SigV4/Converse protocol, custom loop, not OpenAI-compatible. Tier 4. |\n| openai | Chat portion is catalog-shaped, but embeddings, image generation, and OpenAI-specific telemetry/errors are executable behavior. |\n| vertex | Two native protocols (Google GenAI + Anthropic-on-Vertex) and multiple modality pipelines behind one provider identity.","hierarchy":{"lvl0":"Provider Integration","lvl1":"Provider Descriptor Migration Ledger","lvl2":"Must remain class (16)","lvl3":""}},
9658
9680
  {"objectID":"a4f69e09308f075830f2efa7c1b0126236c5ec0fc12bfe08be21764a627e47a5","title":"Must remain core class (1)","url":"/docs/provider-integration/migration-ledger#must-remain-core-class-1","content":"| Provider | Why |\n| ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |\n| openai-compatible | This is the generic protocol surface itself, not a direct vendor. It should consume shared descriptor/config primitives but must never become a fixed-vendor JSON row — collapsing it into the catalog would conflate \"the adapter\" with \"an entry in the adapter's catalog.\" |","hierarchy":{"lvl0":"Provider Integration","lvl1":"Provider Descriptor Migration Ledger","lvl2":"Must remain core class (1)","lvl3":""}},
9659
- {"objectID":"235850085d9102413140920b9d5e5e2ff6113d007c865f5d78012af80fbf1843","title":"Suggested execution order","url":"/docs/provider-integration/migration-ledger#suggested-execution-order","content":"The former steps 1–2 below (mistral; huggingface, deepseek) are complete — see Migrated above. What remains:\nThree shared adapters (local-runtime, embedding-only, image-generation) — each unlocks 3 providers at once; build once, migrate three.\ngoogle-ai descriptor-only JSON metadata — proof of concept for moving static descriptor data out of a \"must remain class\" provider without touching its execution.\nDirect-vendor \"must remain class\" providers (bedrock, openai, vertex, anthropic, azure, sagemaker, nvidia-nim) get descriptor-only JSON metadata migrations, execution untouched — lower priority, cosmetic consolidation only.\nAggregators (litellm, openrouter), the core adapter (openai-compatible), and the decision-only providers (typesafe, laya, xor, perplexity-decider) are excluded from the first-class-count migration priority entirely — the first three add no direct-vendor coverage however they're implemented, and typesafe/laya/xor/perplexity-decider are decide-only providers with no generate/stream surface to migrate at all.","hierarchy":{"lvl0":"Provider Integration","lvl1":"Provider Descriptor Migration Ledger","lvl2":"Suggested execution order","lvl3":""}},
9681
+ {"objectID":"235850085d9102413140920b9d5e5e2ff6113d007c865f5d78012af80fbf1843","title":"Suggested execution order","url":"/docs/provider-integration/migration-ledger#suggested-execution-order","content":"The former steps 1–2 below (mistral; huggingface, deepseek) are complete — see Migrated above. What remains:\nThree shared adapters (local-runtime, embedding-only, image-generation) — each unlocks 3 providers at once; build once, migrate three.\ngoogle-ai descriptor-only JSON metadata — proof of concept for moving static descriptor data out of a \"must remain class\" provider without touching its execution.\nDirect-vendor \"must remain class\" providers (bedrock, openai, vertex, anthropic, azure, sagemaker, nvidia-nim) get descriptor-only JSON metadata migrations, execution untouched — lower priority, cosmetic consolidation only.\nAggregators (litellm, openrouter), the core adapter (openai-compatible), and the decision-only providers (typesafe, laya, xor, perplexity-decider, cloudflare-clef) are excluded from the first-class-count migration priority entirely — the first three add no direct-vendor coverage however they're implemented, and typesafe/laya/xor/perplexity-decider/cloudflare-clef are decide-only providers with no generate/stream surface to migrate at all.","hierarchy":{"lvl0":"Provider Integration","lvl1":"Provider Descriptor Migration Ledger","lvl2":"Suggested execution order","lvl3":""}},
9660
9682
  {"objectID":"6122b2e83358a45c64da87c1ca80a2d579cdfb087a6d9dc58b1e02805296b14a","title":"OpenAI-Compatible Provider Catalog","url":"/docs/provider-integration/openai-compat-catalog","content":"OpenAI-Compatible Provider Catalog\n\nEvery OpenAI-compatible provider in the catalog (80 as of 2026-09-29; the\ndirectory is the list) is registered from one JSON file each, under\nsrc/lib/providers/catalog/<id>.json, and served by one generic class,\nConfiguredOpenAICompatProvider\n(src/lib/providers/configuredOpenAICompat.ts). Adding another provider to\nthis family means adding one JSON file — no subclass, no registry edit, no\nhand-written enum member, no test edit.\n\nThe JSON is the source of truth\n\nOPENAI_COMPAT_CATALOG still exists and keeps its name and element type,\nbut it is now built by the loader (src/lib/providers/catalog/loader.ts)\nfrom the JSON files rather than hand-written. Two consumers read the JSON:\nCodegen (pnpm run codegen:catalog) writes the compile-time\n artifacts into marked regions — the AIProviderName member, the\n <Name>Models enum, the NeurolinkCredentials key — plus the generated\n index. Pre-commit and CI fail on stale output.\nThe loader builds the runtime entry; the descriptor, config options,\n context windows, pricing, vision map and model-choice tables all derive\n from it, as do the provider test suites' rows and counts.\n\nEach file is validated by a zod schema (src/lib/providers/catalog/schema.ts)\nwith a mirrored provider-catalog.schema.json for editor squiggles.\nProbe evidence (roster/auth/billing dates, live-matrix result, PR URL)\nlives in the file's evidence block — the old\ndocs/provider-integration/manifests/<id>.json files were folded into it.\n\nField-by-field reference and the escape hatches:\ntiers/tier-2-catalog-entry.md. Design rationale and the approved rulings:\ndocs/superpowers/plans/2026-08-28-provider-json-catalog-spec.md.\n\nWhen a provider belongs in the catalog\n\nA provider belongs in the JSON catalog if it needs only:\na credential (API key, optionally an extra field like Cloudflare's account id)\na base URL (static default + optional env override, or computed from an\n extra credential field)\na default/fallback model\nerror-message classification (auth / rate-limit / invalid-model / generic)\na named, closed catalog quirk. DeepSeek is the worked example: it 400s on\n json_schema structured-output requests, so deepseek.json sets\n quirks.responseFormatDowngrade: \"json-schema-to-json-object\" and the\n generic ConfiguredOpenAICompatProvider downgrades to json_object before\n sending — no subclass. It also sets quirks.replayReasoningContent: true:\n DeepSeek documents that each assistant turn's reasoning_content must go\n back on later requests once tools are in play, so the shared message\n converter and the streaming tool loop send it — for this quirk only, since\n strict OpenAI-compatible backends reject the unknown field.\na wire-proven capability such as\n capabilities.structuredOutputWithTools. The generic provider suppresses\n response_format when tools are attached by default; an explicit true\n keeps it on the same generate() or stream() request. Set this only after\n a combined tools-plus-schema request returned successfully. Separate tool\n and structured-output probes are not evidence for the combined capability.\n A stale opt-in is still protected by the runtime conflict retry, which drops\n structured output and retries rather than losing the turn.\n\nThe currently opted-in providers are Baseten, DeepSeek, Fireworks AI, GMI\nCloud, Inception Labs, io.net Intelligence, Novita AI, Together AI, Upstage and\nxAI. API Route remains opted out because its evidence verifies the features\nseparately, not together in one request. Mistral remains opted out because its\ncombined probes returned 429 twice, not a successful capability response.\n\nWhen a provider needs a dedicated subclass instead\n\nAzure OpenAI is deliberately not in the catalog because it overrides real\nrequest-shaping behavior that no named catalog quirk expresses:\nAzure OpenAI (src/lib/providers/azureOpenai.ts) overrides four hooks:\n getChatCompletionsURL (deployment-name URL routing across two Azure\n endpoint schemes), getAuthHeaders (Azure's api-key header instead of\n Authorization: Bearer), adjustRequestBody (renames max_tokens to\n max_completion_tokens for o-series/gpt-5+ deployments), and\n suppressResponseFormatWithTools (Azure supports both at once).\n\nIf a future provider needs any hook beyond the 3 mandatory ones\n(getProviderName, getDefaultModel, formatProviderError) or the 2\npurely-declarative optional ones (getFallbackModelName,\ngetFallbackModels) and that hook is not already a named catalog quirk, it\nneeds either a new closed quirk (the DeepSeek route) or a dedicated subclass\n(the Azure OpenAI route).\n\nError-message fidelity\n\nEach entry's errorRules is a direct, order-preserving translation of its\noriginal subclass's formatProviderError if/else ladder into rule data\n(status code and/or case-insensitive pattern), classified via\nclassifyProviderError()\n(src/lib/utils/errorClassifier.ts). Every bespoke message string is\npreserved verbatim — including xAI's \"top up your account\" quota URL and\nGroq's d","hierarchy":{"lvl0":"Provider Integration","lvl1":"OpenAI-Compatible Provider Catalog","lvl2":"","lvl3":""}},
9661
9683
  {"objectID":"9a91a5b4a32a9f1e89c6a28f762c3594c4a39d96d32cc9e6e1fa0f5cb3f609e4","title":"OpenAI-Compatible Provider Catalog","url":"/docs/provider-integration/openai-compat-catalog#openai-compatible-provider-catalog","content":"Every OpenAI-compatible provider in the catalog (80 as of 2026-09-29; the\ndirectory is the list) is registered from one JSON file each, under\nsrc/lib/providers/catalog/<id>.json, and served by one generic class,\nConfiguredOpenAICompatProvider\n(src/lib/providers/configuredOpenAICompat.ts). Adding another provider to\nthis family means adding one JSON file — no subclass, no registry edit, no\nhand-written enum member, no test edit.","hierarchy":{"lvl0":"Provider Integration","lvl1":"OpenAI-Compatible Provider Catalog","lvl2":"OpenAI-Compatible Provider Catalog","lvl3":""}},
9662
9684
  {"objectID":"16e60ee66171009a6638b5110cc8693f90c1a2cfde796c6932c3512d7f7421fb","title":"The JSON is the source of truth","url":"/docs/provider-integration/openai-compat-catalog#the-json-is-the-source-of-truth","content":"OPENAI_COMPAT_CATALOG still exists and keeps its name and element type,\nbut it is now built by the loader (src/lib/providers/catalog/loader.ts)\nfrom the JSON files rather than hand-written. Two consumers read the JSON:\nCodegen (pnpm run codegen:catalog) writes the compile-time\n artifacts into marked regions — the AIProviderName member, the\n <Name>Models enum, the NeurolinkCredentials key — plus the generated\n index. Pre-commit and CI fail on stale output.\nThe loader builds the runtime entry; the descriptor, config options,\n context windows, pricing, vision map and model-choice tables all derive\n from it, as do the provider test suites' rows and counts.\n\nEach file is validated by a zod schema (src/lib/providers/catalog/schema.ts)\nwith a mirrored provider-catalog.schema.json for editor squiggles.\nProbe evidence (roster/auth/billing dates, live-matrix result, PR URL)\nlives in the file's evidence block — the old\ndocs/provider-integration/manifests/<id>.json files were folded into it.\n\nField-by-field reference and the escape hatches:\ntiers/tier-2-catalog-entry.md. Design rationale and the approved rulings:\ndocs/superpowers/plans/2026-08-28-provider-json-catalog-spec.md.","hierarchy":{"lvl0":"Provider Integration","lvl1":"OpenAI-Compatible Provider Catalog","lvl2":"The JSON is the source of truth","lvl3":""}},
@@ -9886,10 +9908,10 @@
9886
9908
  {"objectID":"4578b51cac195eb59c8ac256ef9f2491197666719997b358b74e2b57b5a414fb","title":"TTS / Realtime Errors","url":"/docs/reference/error-codes#tts-realtime-errors","content":"TTSError and RealtimeError carry provider-specific messages (e.g.\nsynthesis failure, WebSocket disconnect, function-call failure). Inspect\nerror.cause for the underlying provider error.","hierarchy":{"lvl0":"Reference","lvl1":"Error Code Reference","lvl2":"TTS / Realtime Errors","lvl3":""}},
9887
9909
  {"objectID":"3e6c6ffd80fa23ecf3686bacf0db97fb3e7e6f430e82b2523745c88e257b6c45","title":"Common Triggers","url":"/docs/reference/error-codes#common-triggers","content":"STT_INVALID_AUDIO_FORMAT is thrown by STTProcessor.transcribe()\n when options.format doesn't appear in the provider's\n getSupportedFormats() list. The CLI infers the format from the\n --input-audio file extension; the SDK requires you to pass it explicitly.\n Fix: either convert the audio to a supported format, or use a different\n STT provider. See docs/getting-started/providers/azure-speech.md for the\n Azure-MP3 case.\nSTT_AUDIO_TOO_LONG is thrown when the buffer exceeds the per-call\n maxAudioBytes limit. Default is 25 MB (matches Whisper's documented\n ceiling). Override via stt.maxAudioBytes.","hierarchy":{"lvl0":"Reference","lvl1":"Error Code Reference","lvl2":"Common Triggers","lvl3":""}},
9888
9910
  {"objectID":"df1592c1b3ccdbed96fbbc998d2facc217e2652c2a3cbb2d630ed64f1c85fdb9","title":"Related Documentation","url":"/docs/reference/error-codes#related-documentation","content":"Troubleshooting Guide - Common issues and solutions\nConfiguration Reference - Environment variables and settings\nFAQ - Frequently asked questions\nProvider Feature Compatibility - Provider capabilities matrix","hierarchy":{"lvl0":"Reference","lvl1":"Error Code Reference","lvl2":"Related Documentation","lvl3":""}},
9889
- {"objectID":"4dccb885a61a05030e1eeb0b2413af7062a282bcce434d742561026c5af1d015","title":"Frequently Asked Questions","url":"/docs/reference/faq","content":"Frequently Asked Questions\n\nCommon questions and answers about NeuroLink usage, configuration, and troubleshooting.\n\nšŸš€ Getting Started\n\nQ: What is NeuroLink?\n\nA: NeuroLink is an enterprise AI development platform that provides unified access to multiple AI providers (OpenAI, Google AI, Anthropic, AWS Bedrock, etc.) through a single SDK and CLI. It includes built-in tools, analytics, evaluation capabilities, and supports the Model Context Protocol (MCP) for extended functionality.\n\nQ: Which AI providers does NeuroLink support?\n\nA: NeuroLink ships AI providers for text generation, streaming, and decision-making — plus separate provider systems for voice and media generation. The text and multimodal providers include:\nOpenAI (GPT-4o, GPT-4.1, o3, o4-mini)\nGoogle AI Studio (Gemini 3 Flash/Pro, Gemini 2.5 Pro/Flash)\nGoogle Vertex AI (Gemini 3, Claude via Vertex)\nAnthropic (Claude Opus 4.7, Sonnet 4.6, 4.5 Opus/Sonnet/Haiku)\nAWS Bedrock (Claude, Titan, Nova models)\nAzure OpenAI (GPT models)\nHugging Face (Open source models)\nOllama (Local AI models)\nMistral AI (Mistral models)\nLiteLLM (100+ models via proxy)\nAWS SageMaker (Custom endpoints)\nOpenAI-compatible (Any OpenAI-API-compatible endpoint)\nOpenRouter (300+ models via OpenRouter)\nDeepSeek (DeepSeek V3, R1)\nNVIDIA NIM (Llama 3.3 70B, 400+ catalog models)\nLM Studio (Local models loaded in LM Studio)\nllama.cpp (Local GGUF models via llama-server)\nGroq, Cerebras, SambaNova, Together AI, Fireworks AI, Perplexity, Cloudflare Workers AI, xAI, Baseten, GMI Cloud, Inception Labs, io.net Intelligence, Mancer, Upstage, API Route (zero-quirk OpenAI-wire-compatible catalog providers)\nCohere (chat, plus embed() and reranking)\nVoyage AI, Jina AI (embedding and/or reranking only — no chat completions)\nTypeSafe Jev, Laya, XOR, Perplexity Decisions (decision-only — serve decide(), not generate()/stream(); tried in the order TypeSafe, Laya, XOR, Perplexity when several are configured)\n\nSee Provider Setup for the complete roster with setup guides.\n\nVoice providers (a separate system from the providers above):\nOpenAI TTS (TTS-1, TTS-1-HD, GPT-4o Audio)\nElevenLabs (Multilingual v2, Turbo v2.5, Flash v2.5)\nDeepgram (Nova-3, Nova-2, Enhanced — STT)\nAzure Speech (Azure Cognitive Services TTS + STT)\nGoogle TTS / STT (Google Cloud Speech)\nWhisper (OpenAI Whisper — STT)\nFish Audio (TTS)\nCartesia (TTS)\nOpenAI Realtime + Gemini Live (realtime voice APIs)\n\nMedia generation providers (image / video / music / avatar) — Kling, Runway, Replicate, Beatoven, Lyria, D-ID, HeyGen. See Media Generation for the full list.\n\nQ: Do I need to install anything?\n\nA: No installation required! You can use NeuroLink directly with npx:\n\nFor frequent use, you can install globally: npm install -g @juspay/neurolink\n\nšŸ”§ Configuration\n\nQ: How do I set up API keys?\n\nA: Create a .env file in your project directory:\n\nNeuroLink automatically loads these environment variables.\n\nQ: Can I use NeuroLink behind a corporate proxy?\n\nA: Yes! NeuroLink automatically detects and uses corporate proxy settings:\n\nNo additional configuration needed.\n\nQ: How do I configure multiple environments (dev/staging/prod)?\n\nA: Use environment-specific .env files:\n\nšŸŽÆ Usage\n\nQ: What's the difference between CLI and SDK?\n\nA:\n\n| Feature | CLI | SDK |\n| -------------------- | ---------------------------- | ------------------------- |\n| Best for | Scripts, automation, testing | Applications, integration |\n| Installation | None required (npx) | npm install required |\n| Output | Text, JSON | Native JavaScript objects |\n| Batch processing | Built-in batch command | Manual implementation |\n| Learning curve | Low | Medium |\n\nQ: How do I choose the best provider for my use case?\n\nA: NeuroLink can auto-select the best provider, or you can choose based on:\nSpeed: Google AI (fastest responses)\nCoding: Anthropic Claude (best for code analysis)\nCreative: OpenAI (best for creative content)\nCost: Google AI Studio (free tier available)\nEnterprise: AWS Bedrock or Azure OpenAI\n\nQ: Can I use multiple providers in the same application?\n\nA: Yes! You can specify different providers for different requests:\n\nšŸ” Troubleshooting\n\nQ: Why am I getting \"API key not found\" errors?\n\nA: Common solutions:\nCheck .env file exists and is in the correct directory\nVerify file format: No spaces around = signs\nCheck file permissions: .env file should be readable\nVerify key format: Keys should start with provider-specific prefixes\n\nQ: Provider status shows \"Authentication failed\" - what should I do?\n\nA:\nVerify API key is correct and hasn't expired\nCheck account status - ensure billing is set up if required\nTest API key manually:\nCheck regional restrictions - some providers have geographic limitations\n\nQ: AWS Bedrock shows \"Not Authorized\" - how do I fix this?\n\nA: AWS Bedrock requires additional setup:\nRequest model acce","hierarchy":{"lvl0":"Reference","lvl1":"Frequently Asked Questions","lvl2":"","lvl3":""}},
9911
+ {"objectID":"4dccb885a61a05030e1eeb0b2413af7062a282bcce434d742561026c5af1d015","title":"Frequently Asked Questions","url":"/docs/reference/faq","content":"Frequently Asked Questions\n\nCommon questions and answers about NeuroLink usage, configuration, and troubleshooting.\n\nšŸš€ Getting Started\n\nQ: What is NeuroLink?\n\nA: NeuroLink is an enterprise AI development platform that provides unified access to multiple AI providers (OpenAI, Google AI, Anthropic, AWS Bedrock, etc.) through a single SDK and CLI. It includes built-in tools, analytics, evaluation capabilities, and supports the Model Context Protocol (MCP) for extended functionality.\n\nQ: Which AI providers does NeuroLink support?\n\nA: NeuroLink ships AI providers for text generation, streaming, and decision-making — plus separate provider systems for voice and media generation. The text and multimodal providers include:\nOpenAI (GPT-4o, GPT-4.1, o3, o4-mini)\nGoogle AI Studio (Gemini 3 Flash/Pro, Gemini 2.5 Pro/Flash)\nGoogle Vertex AI (Gemini 3, Claude via Vertex)\nAnthropic (Claude Opus 4.7, Sonnet 4.6, 4.5 Opus/Sonnet/Haiku)\nAWS Bedrock (Claude, Titan, Nova models)\nAzure OpenAI (GPT models)\nHugging Face (Open source models)\nOllama (Local AI models)\nMistral AI (Mistral models)\nLiteLLM (100+ models via proxy)\nAWS SageMaker (Custom endpoints)\nOpenAI-compatible (Any OpenAI-API-compatible endpoint)\nOpenRouter (300+ models via OpenRouter)\nDeepSeek (DeepSeek V3, R1)\nNVIDIA NIM (Llama 3.3 70B, 400+ catalog models)\nLM Studio (Local models loaded in LM Studio)\nllama.cpp (Local GGUF models via llama-server)\nGroq, Cerebras, SambaNova, Together AI, Fireworks AI, Perplexity, Cloudflare Workers AI, xAI, Baseten, GMI Cloud, Inception Labs, io.net Intelligence, Mancer, Upstage, API Route (zero-quirk OpenAI-wire-compatible catalog providers)\nCohere (chat, plus embed() and reranking)\nVoyage AI, Jina AI (embedding and/or reranking only — no chat completions)\nTypeSafe Jev, Laya, XOR, Perplexity Decisions, Cloudflare Clef (decision-only — serve decide(), not generate()/stream(); tried in the order TypeSafe, Laya, XOR, Perplexity, Cloudflare Clef when several are configured; Clef is configured by the same CLOUDFLARE_API_KEY and CLOUDFLARE_ACCOUNT_ID as the Cloudflare Workers AI text provider)\n\nSee Provider Setup for the complete roster with setup guides.\n\nVoice providers (a separate system from the providers above):\nOpenAI TTS (TTS-1, TTS-1-HD, GPT-4o Audio)\nElevenLabs (Multilingual v2, Turbo v2.5, Flash v2.5)\nDeepgram (Nova-3, Nova-2, Enhanced — STT)\nAzure Speech (Azure Cognitive Services TTS + STT)\nGoogle TTS / STT (Google Cloud Speech)\nWhisper (OpenAI Whisper — STT)\nFish Audio (TTS)\nCartesia (TTS)\nOpenAI Realtime + Gemini Live (realtime voice APIs)\n\nMedia generation providers (image / video / music / avatar) — Kling, Runway, Replicate, Beatoven, Lyria, D-ID, HeyGen. See Media Generation for the full list.\n\nQ: Do I need to install anything?\n\nA: No installation required! You can use NeuroLink directly with npx:\n\nFor frequent use, you can install globally: npm install -g @juspay/neurolink\n\nšŸ”§ Configuration\n\nQ: How do I set up API keys?\n\nA: Create a .env file in your project directory:\n\nNeuroLink automatically loads these environment variables.\n\nQ: Can I use NeuroLink behind a corporate proxy?\n\nA: Yes! NeuroLink automatically detects and uses corporate proxy settings:\n\nNo additional configuration needed.\n\nQ: How do I configure multiple environments (dev/staging/prod)?\n\nA: Use environment-specific .env files:\n\nšŸŽÆ Usage\n\nQ: What's the difference between CLI and SDK?\n\nA:\n\n| Feature | CLI | SDK |\n| -------------------- | ---------------------------- | ------------------------- |\n| Best for | Scripts, automation, testing | Applications, integration |\n| Installation | None required (npx) | npm install required |\n| Output | Text, JSON | Native JavaScript objects |\n| Batch processing | Built-in batch command | Manual implementation |\n| Learning curve | Low | Medium |\n\nQ: How do I choose the best provider for my use case?\n\nA: NeuroLink can auto-select the best provider, or you can choose based on:\nSpeed: Google AI (fastest responses)\nCoding: Anthropic Claude (best for code analysis)\nCreative: OpenAI (best for creative content)\nCost: Google AI Studio (free tier available)\nEnterprise: AWS Bedrock or Azure OpenAI\n\nQ: Can I use multiple providers in the same application?\n\nA: Yes! You can specify different providers for different requests:\n\nšŸ” Troubleshooting\n\nQ: Why am I getting \"API key not found\" errors?\n\nA: Common solutions:\nCheck .env file exists and is in the correct directory\nVerify file format: No spaces around = signs\nCheck file permissions: .env file should be readable\nVerify key format: Keys should start with provider-specific prefixes\n\nQ: Provider status shows \"Authentication failed\" - what should I do?\n\nA:\nVerify API key is correct and hasn't expired\nCheck account status - ensure billing is set up if required\nTest API key manually:\nCheck regional restrictions - some provi","hierarchy":{"lvl0":"Reference","lvl1":"Frequently Asked Questions","lvl2":"","lvl3":""}},
9890
9912
  {"objectID":"430e71707f7db7f9866d7ab0b23f14f769c4812334161d853920bf4358ad8c17","title":"Frequently Asked Questions","url":"/docs/reference/faq#frequently-asked-questions","content":"Common questions and answers about NeuroLink usage, configuration, and troubleshooting.","hierarchy":{"lvl0":"Reference","lvl1":"Frequently Asked Questions","lvl2":"Frequently Asked Questions","lvl3":""}},
9891
9913
  {"objectID":"e9218a0cd45a1960da2da4d8d664f5a8b77494352527c7b4d1e6bc12c137400f","title":"Q: What is NeuroLink?","url":"/docs/reference/faq#q-what-is-neurolink","content":"A: NeuroLink is an enterprise AI development platform that provides unified access to multiple AI providers (OpenAI, Google AI, Anthropic, AWS Bedrock, etc.) through a single SDK and CLI. It includes built-in tools, analytics, evaluation capabilities, and supports the Model Context Protocol (MCP) for extended functionality.","hierarchy":{"lvl0":"Reference","lvl1":"Frequently Asked Questions","lvl2":"Q: What is NeuroLink?","lvl3":""}},
9892
- {"objectID":"c32f5558167b5817bde6577d51c68bbce793c3118b1192fdbf6c6e5958d3faff","title":"Q: Which AI providers does NeuroLink support?","url":"/docs/reference/faq#q-which-ai-providers-does-neurolink-support","content":"A: NeuroLink ships AI providers for text generation, streaming, and decision-making — plus separate provider systems for voice and media generation. The text and multimodal providers include:\nOpenAI (GPT-4o, GPT-4.1, o3, o4-mini)\nGoogle AI Studio (Gemini 3 Flash/Pro, Gemini 2.5 Pro/Flash)\nGoogle Vertex AI (Gemini 3, Claude via Vertex)\nAnthropic (Claude Opus 4.7, Sonnet 4.6, 4.5 Opus/Sonnet/Haiku)\nAWS Bedrock (Claude, Titan, Nova models)\nAzure OpenAI (GPT models)\nHugging Face (Open source models)\nOllama (Local AI models)\nMistral AI (Mistral models)\nLiteLLM (100+ models via proxy)\nAWS SageMaker (Custom endpoints)\nOpenAI-compatible (Any OpenAI-API-compatible endpoint)\nOpenRouter (300+ models via OpenRouter)\nDeepSeek (DeepSeek V3, R1)\nNVIDIA NIM (Llama 3.3 70B, 400+ catalog models)\nLM Studio (Local models loaded in LM Studio)\nllama.cpp (Local GGUF models via llama-server)\nGroq, Cerebras, SambaNova, Together AI, Fireworks AI, Perplexity, Cloudflare Workers AI, xAI, Baseten, GMI Cloud, Inception Labs, io.net Intelligence, Mancer, Upstage, API Route (zero-quirk OpenAI-wire-compatible catalog providers)\nCohere (chat, plus embed() and reranking)\nVoyage AI, Jina AI (embedding and/or reranking only — no chat completions)\nTypeSafe Jev, Laya, XOR, Perplexity Decisions (decision-only — serve decide(), not generate()/stream(); tried in the order TypeSafe, Laya, XOR, Perplexity when several are configured)\n\nSee Provider Setup for the complete roster with setup guides.","hierarchy":{"lvl0":"Reference","lvl1":"Frequently Asked Questions","lvl2":"Q: Which AI providers does NeuroLink support?","lvl3":""}},
9914
+ {"objectID":"c32f5558167b5817bde6577d51c68bbce793c3118b1192fdbf6c6e5958d3faff","title":"Q: Which AI providers does NeuroLink support?","url":"/docs/reference/faq#q-which-ai-providers-does-neurolink-support","content":"A: NeuroLink ships AI providers for text generation, streaming, and decision-making — plus separate provider systems for voice and media generation. The text and multimodal providers include:\nOpenAI (GPT-4o, GPT-4.1, o3, o4-mini)\nGoogle AI Studio (Gemini 3 Flash/Pro, Gemini 2.5 Pro/Flash)\nGoogle Vertex AI (Gemini 3, Claude via Vertex)\nAnthropic (Claude Opus 4.7, Sonnet 4.6, 4.5 Opus/Sonnet/Haiku)\nAWS Bedrock (Claude, Titan, Nova models)\nAzure OpenAI (GPT models)\nHugging Face (Open source models)\nOllama (Local AI models)\nMistral AI (Mistral models)\nLiteLLM (100+ models via proxy)\nAWS SageMaker (Custom endpoints)\nOpenAI-compatible (Any OpenAI-API-compatible endpoint)\nOpenRouter (300+ models via OpenRouter)\nDeepSeek (DeepSeek V3, R1)\nNVIDIA NIM (Llama 3.3 70B, 400+ catalog models)\nLM Studio (Local models loaded in LM Studio)\nllama.cpp (Local GGUF models via llama-server)\nGroq, Cerebras, SambaNova, Together AI, Fireworks AI, Perplexity, Cloudflare Workers AI, xAI, Baseten, GMI Cloud, Inception Labs, io.net Intelligence, Mancer, Upstage, API Route (zero-quirk OpenAI-wire-compatible catalog providers)\nCohere (chat, plus embed() and reranking)\nVoyage AI, Jina AI (embedding and/or reranking only — no chat completions)\nTypeSafe Jev, Laya, XOR, Perplexity Decisions, Cloudflare Clef (decision-only — serve decide(), not generate()/stream(); tried in the order TypeSafe, Laya, XOR, Perplexity, Cloudflare Clef when several are configured; Clef is configured by the same CLOUDFLARE_API_KEY and CLOUDFLARE_ACCOUNT_ID as the Cloudflare Workers AI text provider)\n\nSee Provider Setup for the complete roster with setup guides.","hierarchy":{"lvl0":"Reference","lvl1":"Frequently Asked Questions","lvl2":"Q: Which AI providers does NeuroLink support?","lvl3":""}},
9893
9915
  {"objectID":"c7da7431322332ad0e9e6439acb89a8a9f6af428d15d84c94adeb886ea4fdaab","title":"Q: Do I need to install anything?","url":"/docs/reference/faq#q-do-i-need-to-install-anything","content":"A: No installation required! You can use NeuroLink directly with npx:\n\nFor frequent use, you can install globally: npm install -g @juspay/neurolink","hierarchy":{"lvl0":"Reference","lvl1":"Frequently Asked Questions","lvl2":"Q: Do I need to install anything?","lvl3":""}},
9894
9916
  {"objectID":"379b5c908063649275668a696f71c0e87923dd0d011f221a19a43f69a483acea","title":"Q: How do I set up API keys?","url":"/docs/reference/faq#q-how-do-i-set-up-api-keys","content":"A: Create a .env file in your project directory:\n\n`bash","hierarchy":{"lvl0":"Reference","lvl1":"Frequently Asked Questions","lvl2":"Q: How do I set up API keys?","lvl3":""}},
9895
9917
  {"objectID":"07c128fc6b5bcea9361ff3423a97e6a5a4e85fb14571f4a5d14bcb11eb389105","title":".env file","url":"/docs/reference/faq#env-file","content":"OPENAIAPIKEY=\"sk-your-openai-key\"\nGOOGLEAIAPI_KEY=\"AIza-your-google-ai-key\"\nANTHROPICAPIKEY=\"sk-ant-your-anthropic-key\"","hierarchy":{"lvl0":"Reference","lvl1":"Frequently Asked Questions","lvl2":".env file","lvl3":""}},
@@ -11019,8 +11041,8 @@
11019
11041
  {"objectID":"1a0933649d2e2a9e068e520585c510992a787eb886f8661e6da36cf83ce7e3eb","title":"Multiple inputs","url":"/docs/skills/neurolink-guide/multimodal#multiple-inputs","content":"neurolink generate \"Compare\" --image ./a.png --image ./b.png --pdf ./docs.pdf\n`","hierarchy":{"lvl0":"Skills","lvl1":"NeuroLink Multimodal Support","lvl2":"Multiple inputs","lvl3":""}},
11020
11042
  {"objectID":"e2858a465b5d2bc8ef1db9834500597eb845bf3b40f3a0ba4b29e54033a4b23b","title":"File Size Considerations","url":"/docs/skills/neurolink-guide/multimodal#file-size-considerations","content":"| File Type | Recommended Max | Notes |\n| --------- | --------------- | -------------------- |\n| Images | 20MB | Resized if larger |\n| PDFs | 50MB | Page limit may apply |\n| CSV | 10MB | Use maxRows option |\n| Code | 100KB | Split large files |","hierarchy":{"lvl0":"Skills","lvl1":"NeuroLink Multimodal Support","lvl2":"File Size Considerations","lvl3":""}},
11021
11043
  {"objectID":"8d5a67779513f44569761e7ea2fa399de1aafd682a08e0b14c08850f0c223cb2","title":"Next Steps","url":"/docs/skills/neurolink-guide/multimodal#next-steps","content":"MCP tools - Add external tools\nRAG integration - Document-grounded generation\nProviders - Configure vision-capable providers","hierarchy":{"lvl0":"Skills","lvl1":"NeuroLink Multimodal Support","lvl2":"Next Steps","lvl3":""}},
11022
- {"objectID":"d9e19daaa21e9512ad71f1187055c474d4b51da0c78b56a197286c6403126b9a","title":"NeuroLink Provider Configuration","url":"/docs/skills/neurolink-guide/providers","content":"NeuroLink Provider Configuration\n\nNeuroLink supports many AI providers through a unified API. This page highlights the most commonly-configured text providers — see the README provider table and the Provider Capabilities Audit for the full matrix, including newer text providers (DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, plus the Tier-2 catalog providers — Groq, Cerebras, SambaNova, Together AI, Fireworks AI, Perplexity, Cloudflare Workers AI, xAI, and more), the decision-only TypeSafe Jev, Laya, XOR and Perplexity Decisions providers (serve decide(), not generate()/stream(); tried in the order TypeSafe, Laya, XOR, Perplexity when several are configured), and voice providers (OpenAI TTS, ElevenLabs, Deepgram, Azure Speech, Whisper, Fish Audio, Cartesia, OpenAI Realtime, Gemini Live).\n\nCommon Providers\n\n| Provider | Enum Name | Aliases | Default Model |\n| ---------------- | -------------- | -------------- | --------------------------------------- |\n| OpenAI | openai | gpt, chatgpt | gpt-4o |\n| Anthropic | anthropic | claude | claude-3-5-sonnet-20241022 |\n| Google AI Studio | google-ai | gemini, google | gemini-2.5-flash |\n| Google Vertex AI | vertex | google-vertex | gemini-2.5-flash |\n| AWS Bedrock | bedrock | aws-bedrock | anthropic.claude-3-sonnet-20240229-v1:0 |\n| Azure OpenAI | azure-openai | azure | gpt-4o |\n| Mistral AI | mistral | - | mistral-large |\n| Ollama | ollama | - | llama3 |\n| LiteLLM | litellm | - | varies |\n| AWS SageMaker | sagemaker | - | custom |\n| Hugging Face | hugging-face | hf | varies |\n| OpenRouter | openrouter | - | varies |\n| Gateway | gateway | - | varies |\n\nOpenAI\n\nAvailable Models:\ngpt-4o - Latest GPT-4 Omni\ngpt-4o-mini - Faster, cheaper\ngpt-4-turbo - GPT-4 Turbo\no1 - Reasoning model\no1-mini - Smaller reasoning model\n\nAnthropic\n\nAvailable Models:\nclaude-3-5-sonnet-20241022 - Latest Sonnet\nclaude-3-7-sonnet-20250219 - Claude 3.7 Sonnet\nclaude-3-opus-20240229 - Most capable\nclaude-3-haiku-20240307 - Fastest\n\nExtended Thinking:\n\nGoogle AI Studio\n\nAvailable Models:\ngemini-2.5-flash - Fast and capable\ngemini-2.5-pro - Most capable\ngemini-2.0-flash - Previous generation\ngemini-3-flash-preview - Preview of Gemini 3\n\nGoogle Vertex AI\n\nAvailable Models:\ngemini-3-flash - Latest Gemini 3\ngemini-3-pro - Most capable Gemini 3\ngemini-2.5-flash - Fast\ngemini-2.5-pro - Previous gen capable\n\nExtended Thinking (Gemini 3):\n\nAWS Bedrock\n\nAvailable Models:\nanthropic.claude-3-sonnet-20240229-v1:0\nanthropic.claude-3-haiku-20240307-v1:0\nanthropic.claude-3-opus-20240229-v1:0\namazon.titan-text-express-v1\namazon.nova-pro-v1:0\nmeta.llama3-70b-instruct-v1:0\n\nAzure OpenAI\n\nMistral AI\n\nAvailable Models:\nmistral-large-latest - Most capable\nmistral-small-latest - Fast\ncodestral-latest - Code specialized\nministral-8b-latest - Small\n\nOllama (Local)\n\nSetup:\n\nAvailable Models:\nllama3 - Meta Llama 3\nllama3:70b - Larger Llama 3\nmistral - Mistral 7B\ncodellama - Code specialized\nphi3 - Microsoft Phi-3\n\nLiteLLM\n\nAWS SageMaker\n\nHugging Face\n\nOpenRouter\n\nProvider Fallback\n\nConfigure automatic fallback to another provider:\n\nCheck Provider Status\n\nProvider-Specific Options\n\nTemperature and Sampling\n\nSystem Prompts\n\nVision-Capable Models\n\nNot all models support image inputs:\n\n| Provider | Vision Models |\n| --------- | --------------------- |\n| OpenAI | gpt-4o, gpt-4-turbo |\n| Anthropic | All Claude 3 models |\n| Vertex | Gemini 2.5+, Gemini 3 |\n| Google AI | Gemini 2.5+, Gemini 3 |\n| Bedrock | Claude 3 models |\n\nNext Steps\nMultimodal inputs - Work with images and documents\nMCP tools - Add external tools\nRAG integration - Document-grounded generation","hierarchy":{"lvl0":"Skills","lvl1":"NeuroLink Provider Configuration","lvl2":"","lvl3":""}},
11023
- {"objectID":"b66a3f5004ec2abbc42277973845a5485480e16a3ea54ea9f31e83978a9ee521","title":"NeuroLink Provider Configuration","url":"/docs/skills/neurolink-guide/providers#neurolink-provider-configuration","content":"NeuroLink supports many AI providers through a unified API. This page highlights the most commonly-configured text providers — see the README provider table and the Provider Capabilities Audit for the full matrix, including newer text providers (DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, plus the Tier-2 catalog providers — Groq, Cerebras, SambaNova, Together AI, Fireworks AI, Perplexity, Cloudflare Workers AI, xAI, and more), the decision-only TypeSafe Jev, Laya, XOR and Perplexity Decisions providers (serve decide(), not generate()/stream(); tried in the order TypeSafe, Laya, XOR, Perplexity when several are configured), and voice providers (OpenAI TTS, ElevenLabs, Deepgram, Azure Speech, Whisper, Fish Audio, Cartesia, OpenAI Realtime, Gemini Live).","hierarchy":{"lvl0":"Skills","lvl1":"NeuroLink Provider Configuration","lvl2":"NeuroLink Provider Configuration","lvl3":""}},
11044
+ {"objectID":"d9e19daaa21e9512ad71f1187055c474d4b51da0c78b56a197286c6403126b9a","title":"NeuroLink Provider Configuration","url":"/docs/skills/neurolink-guide/providers","content":"NeuroLink Provider Configuration\n\nNeuroLink supports many AI providers through a unified API. This page highlights the most commonly-configured text providers — see the README provider table and the Provider Capabilities Audit for the full matrix, including newer text providers (DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, plus the Tier-2 catalog providers — Groq, Cerebras, SambaNova, Together AI, Fireworks AI, Perplexity, Cloudflare Workers AI, xAI, and more), the decision-only TypeSafe Jev, Laya, XOR, Perplexity Decisions and Cloudflare Clef providers (serve decide(), not generate()/stream(); tried in the order TypeSafe, Laya, XOR, Perplexity, Cloudflare Clef when several are configured; Clef reads the same CLOUDFLARE_API_KEY and CLOUDFLARE_ACCOUNT_ID as the Cloudflare Workers AI text provider), and voice providers (OpenAI TTS, ElevenLabs, Deepgram, Azure Speech, Whisper, Fish Audio, Cartesia, OpenAI Realtime, Gemini Live).\n\nCommon Providers\n\n| Provider | Enum Name | Aliases | Default Model |\n| ---------------- | -------------- | -------------- | --------------------------------------- |\n| OpenAI | openai | gpt, chatgpt | gpt-4o |\n| Anthropic | anthropic | claude | claude-3-5-sonnet-20241022 |\n| Google AI Studio | google-ai | gemini, google | gemini-2.5-flash |\n| Google Vertex AI | vertex | google-vertex | gemini-2.5-flash |\n| AWS Bedrock | bedrock | aws-bedrock | anthropic.claude-3-sonnet-20240229-v1:0 |\n| Azure OpenAI | azure-openai | azure | gpt-4o |\n| Mistral AI | mistral | - | mistral-large |\n| Ollama | ollama | - | llama3 |\n| LiteLLM | litellm | - | varies |\n| AWS SageMaker | sagemaker | - | custom |\n| Hugging Face | hugging-face | hf | varies |\n| OpenRouter | openrouter | - | varies |\n| Gateway | gateway | - | varies |\n\nOpenAI\n\nAvailable Models:\ngpt-4o - Latest GPT-4 Omni\ngpt-4o-mini - Faster, cheaper\ngpt-4-turbo - GPT-4 Turbo\no1 - Reasoning model\no1-mini - Smaller reasoning model\n\nAnthropic\n\nAvailable Models:\nclaude-3-5-sonnet-20241022 - Latest Sonnet\nclaude-3-7-sonnet-20250219 - Claude 3.7 Sonnet\nclaude-3-opus-20240229 - Most capable\nclaude-3-haiku-20240307 - Fastest\n\nExtended Thinking:\n\nGoogle AI Studio\n\nAvailable Models:\ngemini-2.5-flash - Fast and capable\ngemini-2.5-pro - Most capable\ngemini-2.0-flash - Previous generation\ngemini-3-flash-preview - Preview of Gemini 3\n\nGoogle Vertex AI\n\nAvailable Models:\ngemini-3-flash - Latest Gemini 3\ngemini-3-pro - Most capable Gemini 3\ngemini-2.5-flash - Fast\ngemini-2.5-pro - Previous gen capable\n\nExtended Thinking (Gemini 3):\n\nAWS Bedrock\n\nAvailable Models:\nanthropic.claude-3-sonnet-20240229-v1:0\nanthropic.claude-3-haiku-20240307-v1:0\nanthropic.claude-3-opus-20240229-v1:0\namazon.titan-text-express-v1\namazon.nova-pro-v1:0\nmeta.llama3-70b-instruct-v1:0\n\nAzure OpenAI\n\nMistral AI\n\nAvailable Models:\nmistral-large-latest - Most capable\nmistral-small-latest - Fast\ncodestral-latest - Code specialized\nministral-8b-latest - Small\n\nOllama (Local)\n\nSetup:\n\nAvailable Models:\nllama3 - Meta Llama 3\nllama3:70b - Larger Llama 3\nmistral - Mistral 7B\ncodellama - Code specialized\nphi3 - Microsoft Phi-3\n\nLiteLLM\n\nAWS SageMaker\n\nHugging Face\n\nOpenRouter\n\nProvider Fallback\n\nConfigure automatic fallback to another provider:\n\nCheck Provider Status\n\nProvider-Specific Options\n\nTemperature and Sampling\n\nSystem Prompts\n\nVision-Capable Models\n\nNot all models support image inputs:\n\n| Provider | Vision Models |\n| --------- | --------------------- |\n| OpenAI | gpt-4o, gpt-4-turbo |\n| Anthropic | All Claude 3 models |\n| Vertex | Gemini 2.5+, Gemini 3 |\n| Google AI | Gemini 2.5+, Gemini 3 |\n| Bedrock | Claude 3 models |\n\nNext Steps\nMultimodal inputs - Work with images and documents\nMCP tools - Add external tools\nRAG integration - Document-grounded generation","hierarchy":{"lvl0":"Skills","lvl1":"NeuroLink Provider Configuration","lvl2":"","lvl3":""}},
11045
+ {"objectID":"b66a3f5004ec2abbc42277973845a5485480e16a3ea54ea9f31e83978a9ee521","title":"NeuroLink Provider Configuration","url":"/docs/skills/neurolink-guide/providers#neurolink-provider-configuration","content":"NeuroLink supports many AI providers through a unified API. This page highlights the most commonly-configured text providers — see the README provider table and the Provider Capabilities Audit for the full matrix, including newer text providers (DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, plus the Tier-2 catalog providers — Groq, Cerebras, SambaNova, Together AI, Fireworks AI, Perplexity, Cloudflare Workers AI, xAI, and more), the decision-only TypeSafe Jev, Laya, XOR, Perplexity Decisions and Cloudflare Clef providers (serve decide(), not generate()/stream(); tried in the order TypeSafe, Laya, XOR, Perplexity, Cloudflare Clef when several are configured; Clef reads the same CLOUDFLARE_API_KEY and CLOUDFLARE_ACCOUNT_ID as the Cloudflare Workers AI text provider), and voice providers (OpenAI TTS, ElevenLabs, Deepgram, Azure Speech, Whisper, Fish Audio, Cartesia, OpenAI Realtime, Gemini Live).","hierarchy":{"lvl0":"Skills","lvl1":"NeuroLink Provider Configuration","lvl2":"NeuroLink Provider Configuration","lvl3":""}},
11024
11046
  {"objectID":"ca77e6781279a1e849f17c5edd76b2cc111546faa201202d3dc11820334170ce","title":"Common Providers","url":"/docs/skills/neurolink-guide/providers#common-providers","content":"| Provider | Enum Name | Aliases | Default Model |\n| ---------------- | -------------- | -------------- | --------------------------------------- |\n| OpenAI | openai | gpt, chatgpt | gpt-4o |\n| Anthropic | anthropic | claude | claude-3-5-sonnet-20241022 |\n| Google AI Studio | google-ai | gemini, google | gemini-2.5-flash |\n| Google Vertex AI | vertex | google-vertex | gemini-2.5-flash |\n| AWS Bedrock | bedrock | aws-bedrock | anthropic.claude-3-sonnet-20240229-v1:0 |\n| Azure OpenAI | azure-openai | azure | gpt-4o |\n| Mistral AI | mistral | - | mistral-large |\n| Ollama | ollama | - | llama3 |\n| LiteLLM | litellm | - | varies |\n| AWS SageMaker | sagemaker | - | custom |\n| Hugging Face | hugging-face | hf | varies |\n| OpenRouter | openrouter | - | varies |\n| Gateway | gateway | - | varies |","hierarchy":{"lvl0":"Skills","lvl1":"NeuroLink Provider Configuration","lvl2":"Common Providers","lvl3":""}},
11025
11047
  {"objectID":"9c5370e9c625cd036ee1c3c8ba7c0bea9a5b6b68ca08ccd9c710d3e65429d843","title":"OpenAI","url":"/docs/skills/neurolink-guide/providers#openai","content":"`bash","hierarchy":{"lvl0":"Skills","lvl1":"NeuroLink Provider Configuration","lvl2":"OpenAI","lvl3":""}},
11026
11048
  {"objectID":"fcc21bd062d4e7ab394bd0ad8a1257ec9ed367c205f68dab30479e6758c5cf20","title":"Environment","url":"/docs/skills/neurolink-guide/providers#environment","content":"OPENAIAPIKEY=sk-...\nOPENAIORGID=org-... # Optional\nOPENAIBASEURL=... # Optional, for proxies\ntypescript\nconst result = await neurolink.generate({\n input: { text: \"Hello\" },\n provider: \"openai\",\n model: \"gpt-4o\", // or gpt-4o-mini, gpt-4-turbo, o1, o1-mini\n});\n\n\n**Available Models:**\n\n- gpt-4o - Latest GPT-4 Omni\n- gpt-4o-mini - Faster, cheaper\n- gpt-4-turbo - GPT-4 Turbo\n- o1 - Reasoning model\n- o1-mini` - Smaller reasoning model","hierarchy":{"lvl0":"Skills","lvl1":"NeuroLink Provider Configuration","lvl2":"Environment","lvl3":""}},