@madgagarin/pi-agentrouter 2.0.0 → 2.1.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (3) hide show
  1. package/README.md +33 -32
  2. package/index.ts +746 -27
  3. package/package.json +4 -2
package/README.md CHANGED
@@ -5,9 +5,9 @@
5
5
  [![Pi Plugin](https://img.shields.io/badge/Pi-Extension-purple.svg)](https://pi.dev)
6
6
  [![AgentRouter Gateway](https://img.shields.io/badge/Gateway-agentrouter.org-orange.svg)](https://agentrouter.org)
7
7
 
8
- Use **GPT-5.6 Sol**, **Claude Opus 5**, **Claude Opus 4.8**, **DeepSeek V4 Flash**, and **GLM 5.3** in your [Pi Coding Agent](https://pi.dev) using a single API key from [AgentRouter](https://agentrouter.org).
8
+ Use **GPT-6 Astra**, **GPT-5.6 Sol**, **Claude Opus 5**, **Claude Opus 4.8**, and **DeepSeek V4 Flash** in your [Pi Coding Agent](https://pi.dev) using a single API key from [AgentRouter](https://agentrouter.org).
9
9
 
10
- > šŸŽ **Free Trial Credits:** New to AgentRouter? Get up to **$175 in free credits** (including a **+$50 bonus**) to test GPT-5.6, Claude Opus 5, and DeepSeek V4 — no credit card needed. That's enough for **over 80,000,000 tokens** on DeepSeek V4!
10
+ > šŸŽ **Free Trial Credits:** New to AgentRouter? Get up to **$175 in free credits** (including a **+$50 bonus**) to test GPT-6 Astra, Claude Opus 5, and DeepSeek V4 — no credit card needed. That's enough for **millions of tokens** on DeepSeek V4!
11
11
  > šŸ‘‰ **[Claim your free trial credits on AgentRouter.org →](https://agentrouter.org/register?aff=34dc)**
12
12
 
13
13
  ---
@@ -34,48 +34,48 @@ pi install npm:@madgagarin/pi-agentrouter
34
34
 
35
35
  ---
36
36
 
37
- ## Why use this plugin?
37
+ ## Features
38
38
 
39
- [AgentRouter](https://agentrouter.org) ([agentrouter.org](https://agentrouter.org)) provides affordable unified access to frontier LLMs, but using raw OpenAI/Anthropic proxy configurations in Pi often runs into edge cases: Cloudflare WAF checks, rate-limit bursts from parallel subagents, prompt caching cache misses, and role naming conflicts.
40
-
41
- This plugin fixes all of that automatically:
42
-
43
- - **Live USD Pricing in Pi:** Pulls current rates from [agentrouter.org/api/pricing](https://agentrouter.org/api/pricing) on startup so Pi's built-in cost tracking shows your exact spend in dollars.
44
- - **Live Quota Monitor (`/agentrouter check`):** Quick health check that pings all models to see if daily batch quotas are open, and displays your monthly usage in USD.
45
- - **1M Context Windows:** All models are configured with their full 1,048,576 token context limits and native reasoning/adaptive thinking.
46
- - **Cross-Process Rate Pacing:** Uses a shared lock file (`~/.pi/agent/.agentrouter-pacing`) so background subagents (`pi-subagents`) and the main chat won't trip 429 rate limits.
47
- - **Zero 400 & 401 Errors:** Handles canonical `pi-code` prompt header placement for WAF authorization and automatically normalizes OpenAI `developer` roles to `system`.
48
- - **High Cache Hit Rates (>80%):** Preserves session affinity headers and disables destructive prompt rewriting on AgentRouter routes.
39
+ - **DeepSeek Multi-Turn Tool Calling:** Seamlessly flattens multi-turn tool history and preserves `reasoning_content` across tool execution turns, completely preventing upstream gateway 400 thinking mode errors.
40
+ - **Model Synchronization:** Automatically registers and adds active models to `enabledModels` in `settings.json` for quick selection via `Ctrl+P`.
41
+ - **Schema Sanitization:** Automatically normalizes tool definitions (e.g. converting `required: null` to empty arrays) for strict OpenAI schema validation compatibility.
42
+ - **WAF Diagnostics & Safe Redaction:** Intercepts upstream blocks and safely redacts older messages while preserving thinking placeholders for reasoning models.
43
+ - **Isolated Credential Storage:** Manages API keys exclusively within `agentrouter-*` provider namespaces in `auth.json` without modifying default third-party provider keys.
44
+ - **Live Pricing & Quota Probing:** Fetches current rates from the gateway API on startup and provides `/agentrouter check` to probe model availability and track usage.
45
+ - **Subagent Rate Pacing:** Uses a file-based lock (`~/.pi/agent/.agentrouter-pacing`) across concurrent subagents to prevent 429 rate limit errors.
46
+ - **Prompt Caching Compatibility:** Preserves affinity headers and pricing for prompt caching reuse.
49
47
 
50
48
  ---
51
49
 
52
50
  ## Models & Pricing
53
51
 
54
- Rates are pulled directly from the [agentrouter.org](https://agentrouter.org) gateway API ($2.00 / 1M tokens base unit):
52
+ Rates are fetched from the [agentrouter.org](https://agentrouter.org) gateway API ($2.00 / 1M tokens base unit):
55
53
 
56
- | Model | Provider | Context | Output | Reasoning | Input / 1M | Output / 1M | Quota Policy |
57
- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- |
58
- | `deepseek-v4-flash` | `agentrouter-openai` | 1M | 64K | Yes | $2.00 | $6.00 | Unlimited |
59
- | `glm-5.3` | `agentrouter-openai` | 1M | 128K | Yes | $3.00 | $12.00 | Unlimited |
60
- | `gpt-5.6-sol` | `agentrouter-openai` | 1M | 128K | Yes | $3.00 | $15.00 | Daily batch drops |
61
- | `claude-opus-5` | `agentrouter-clode` | 1M | 64K | Yes (Adaptive) | $8.00 | $40.00 | Daily batch drops |
54
+ | Model | Provider | Context | Output | Reasoning | Input / 1M | Output / 1M | Cache Read / 1M | Quota Policy |
55
+ | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- |
56
+ | `deepseek-v4-flash` | `agentrouter-openai` | 1M | 64K | Yes | $4.00 | $12.00 | $2.00 | Unlimited |
57
+ | `glm-5.3` | `agentrouter-openai` | 1M | 128K | Yes | $3.00 | $12.00 | - | Unlimited |
58
+ | `gpt-6-astra` | `agentrouter-openai` | 1M | 128K | Yes | $3.00 | $15.00 | - | Daily batch drops |
59
+ | `gpt-5.6-sol` | `agentrouter-openai` | 1M | 128K | Yes | $3.00 | $15.00 | - | Daily batch drops |
60
+ | `claude-opus-5` | `agentrouter-clode` | 1M | 64K | Yes (Adaptive) | $6.00 | $30.00 | Daily batch drops |
62
61
  | `claude-opus-4-8` | `agentrouter-clode` | 1M | 64K | Yes (Adaptive) | $8.00 | $40.00 | Daily batch drops |
63
62
 
64
- *Note: Claude and GPT models are released in daily batches on AgentRouter. If you hit a 402, run `/agentrouter check` to verify, and switch to `deepseek-v4-flash` or `glm-5.3` for unlimited coding.*
63
+ *Note: Claude and GPT models are released in daily batches on AgentRouter. When a batch is exhausted (HTTP 402), use `/agentrouter check` to monitor status or switch to `deepseek-v4-flash` / `glm-5.3` for unrestricted usage.*
65
64
 
66
65
  ---
67
66
 
68
67
  ## In-Chat Commands
69
68
 
70
- | Command | What it does |
69
+ | Command | Description |
71
70
  | :--- | :--- |
72
- | `/agentrouter` | Overview of active model, current monthly USD spend, package order, and pacing delay. |
73
- | `/agentrouter check` | Live preflight probe of all model quotas (200 OK vs 402) and monthly usage. |
74
- | `/agentrouter pricing` | Fetches and prints the latest official pricing table from [agentrouter.org](https://agentrouter.org). |
75
- | `/agentrouter key <key>` | Sets API key and syncs it across `agentrouter.json` and Pi's `auth.json`. |
76
- | `/agentrouter pacing <ms>` | Sets delay between consecutive requests (default: `3500` ms). |
77
- | `/agentrouter fix-order` | Places this plugin above `pi-cache-optimizer` in `settings.json` if needed. |
78
- | `/compact` | Compresses chat history safely without 401 authorization drops. |
71
+ | `/agentrouter` | Show active model, current monthly spend, extension ordering, and pacing delay. |
72
+ | `/agentrouter check` | Probe model availability (200 OK vs 402) and display monthly usage. |
73
+ | `/agentrouter pricing` | Display current pricing table from [agentrouter.org](https://agentrouter.org). |
74
+ | `/agentrouter sync` | Sync active flagship models into `enabledModels` in `settings.json`. |
75
+ | `/agentrouter key <key>` | Set API key and store it in `agentrouter.json` and `auth.json`. |
76
+ | `/agentrouter pacing <ms>` | Configure delay between requests (default: `3500` ms). |
77
+ | `/agentrouter fix-order` | Reorder extension before `pi-cache-optimizer` in `settings.json` if necessary. |
78
+ | `/compact` | Compact conversation history while preserving required gateway headers. |
79
79
 
80
80
  ---
81
81
 
@@ -83,7 +83,7 @@ Rates are pulled directly from the [agentrouter.org](https://agentrouter.org) ga
83
83
 
84
84
  | Shortcut | Action |
85
85
  | :--- | :--- |
86
- | `Ctrl + P` | Cycle to next model (`deepseek-v4-flash` āž” `glm-5.3` āž” `gpt-5.6-sol` āž” `claude-opus-5` āž” `claude-opus-4-8`) |
86
+ | `Ctrl + P` | Cycle to next model (`deepseek-v4-flash` āž” `gpt-6-astra` āž” `gpt-5.6-sol` āž” `claude-opus-5` āž” `claude-opus-4-8`) |
87
87
  | `Shift + Ctrl + P` | Cycle to previous model |
88
88
  | `Shift + Tab` | Toggle reasoning depth (`off` āž” `minimal` āž” `low` āž” `medium` āž” `high`) |
89
89
  | `Ctrl + T` | Toggle reasoning block visibility |
@@ -93,7 +93,7 @@ Rates are pulled directly from the [agentrouter.org](https://agentrouter.org) ga
93
93
 
94
94
  ## Recommended `settings.json`
95
95
 
96
- Add this to `~/.pi/agent/settings.json` for convenient model switching:
96
+ Add this to `~/.pi/agent/settings.json` for model switching:
97
97
 
98
98
  ```json
99
99
  {
@@ -102,7 +102,7 @@ Add this to `~/.pi/agent/settings.json` for convenient model switching:
102
102
  "defaultThinkingLevel": "low",
103
103
  "enabledModels": [
104
104
  "agentrouter-openai/deepseek-v4-flash",
105
- "agentrouter-openai/glm-5.3",
105
+ "agentrouter-openai/gpt-6-astra",
106
106
  "agentrouter-openai/gpt-5.6-sol",
107
107
  "agentrouter-clode/claude-opus-5",
108
108
  "agentrouter-clode/claude-opus-4-8"
@@ -128,3 +128,4 @@ AgentRouter requires the base `pi-code` prompt signature for authentication. If
128
128
  ## License
129
129
 
130
130
  MIT Ā© [madgagarin](https://github.com/madgagarin)
131
+
package/index.ts CHANGED
@@ -41,6 +41,16 @@ export interface ApiPricingModel {
41
41
  }
42
42
 
43
43
  export const KNOWN_MODEL_SPECS: Record<string, ModelSpec> = {
44
+ "gpt-6-astra": {
45
+ id: "gpt-6-astra",
46
+ name: "gpt-6-astra",
47
+ providerType: "openai",
48
+ contextWindow: 1048576,
49
+ maxTokens: 131072,
50
+ reasoning: true,
51
+ compat: { sendSessionAffinityHeaders: true },
52
+ cost: { input: 3.0 / 1_000_000, output: 15.0 / 1_000_000, cacheRead: 0, cacheWrite: 0 },
53
+ },
44
54
  "deepseek-v4-flash": {
45
55
  id: "deepseek-v4-flash",
46
56
  name: "deepseek-v4-flash",
@@ -49,7 +59,7 @@ export const KNOWN_MODEL_SPECS: Record<string, ModelSpec> = {
49
59
  maxTokens: 65536,
50
60
  reasoning: true,
51
61
  compat: { sendSessionAffinityHeaders: true },
52
- cost: { input: 2.0 / 1_000_000, output: 6.0 / 1_000_000, cacheRead: 0, cacheWrite: 0 },
62
+ cost: { input: 4.0 / 1_000_000, output: 12.0 / 1_000_000, cacheRead: 2.0 / 1_000_000, cacheWrite: 0 },
53
63
  },
54
64
  "deepseek-v4f": {
55
65
  id: "deepseek-v4f",
@@ -59,7 +69,7 @@ export const KNOWN_MODEL_SPECS: Record<string, ModelSpec> = {
59
69
  maxTokens: 65536,
60
70
  reasoning: true,
61
71
  compat: { sendSessionAffinityHeaders: true },
62
- cost: { input: 2.0 / 1_000_000, output: 6.0 / 1_000_000, cacheRead: 0, cacheWrite: 0 },
72
+ cost: { input: 4.0 / 1_000_000, output: 12.0 / 1_000_000, cacheRead: 2.0 / 1_000_000, cacheWrite: 0 },
63
73
  },
64
74
  "glm-5.3": {
65
75
  id: "glm-5.3",
@@ -129,7 +139,7 @@ export const KNOWN_MODEL_SPECS: Record<string, ModelSpec> = {
129
139
  sendSessionAffinityHeaders: true,
130
140
  supportsEagerToolInputStreaming: false,
131
141
  },
132
- cost: { input: 8.0 / 1_000_000, output: 40.0 / 1_000_000, cacheRead: 0, cacheWrite: 0 },
142
+ cost: { input: 6.0 / 1_000_000, output: 30.0 / 1_000_000, cacheRead: 0, cacheWrite: 0 },
133
143
  },
134
144
  "claude-opus-4-7": {
135
145
  id: "claude-opus-4-7",
@@ -197,7 +207,6 @@ export function saveConfig(cfg: AgentRouterConfig): void {
197
207
  }
198
208
  auth["agentrouter-openai"] = { type: "api_key", key: cfg.apiKey };
199
209
  auth["agentrouter-clode"] = { type: "api_key", key: cfg.apiKey };
200
- auth["anthropic"] = { type: "api_key", key: cfg.apiKey };
201
210
  fs.writeFileSync(authPath, JSON.stringify(auth, null, 2), "utf-8");
202
211
  }
203
212
  } catch {}
@@ -272,6 +281,36 @@ export function fixPackagePriorityInSettings(): boolean {
272
281
  return false;
273
282
  }
274
283
 
284
+ export const FLAGSHIP_MODELS: string[] = [
285
+ "agentrouter-openai/deepseek-v4-flash",
286
+ "agentrouter-openai/gpt-6-astra",
287
+ "agentrouter-openai/gpt-5.6-sol",
288
+ "agentrouter-clode/claude-opus-5",
289
+ "agentrouter-clode/claude-opus-4-8",
290
+ ];
291
+
292
+ export function syncEnabledModelsInSettings(): { added: string[]; count: number } {
293
+ try {
294
+ if (!fs.existsSync(SETTINGS_FILE)) return { added: [], count: 0 };
295
+ const settings = JSON.parse(fs.readFileSync(SETTINGS_FILE, "utf-8"));
296
+ if (!Array.isArray(settings.enabledModels)) return { added: [], count: 0 };
297
+
298
+ const added: string[] = [];
299
+ for (const m of FLAGSHIP_MODELS) {
300
+ if (!settings.enabledModels.includes(m)) {
301
+ settings.enabledModels.push(m);
302
+ added.push(m);
303
+ }
304
+ }
305
+ if (added.length > 0) {
306
+ fs.writeFileSync(SETTINGS_FILE, JSON.stringify(settings, null, 2), "utf-8");
307
+ }
308
+ return { added, count: settings.enabledModels.length };
309
+ } catch {
310
+ return { added: [], count: 0 };
311
+ }
312
+ }
313
+
275
314
  export function getLastRequestEndTime(): number {
276
315
  try {
277
316
  if (fs.existsSync(PACING_FILE)) {
@@ -338,10 +377,611 @@ export function enforceCanonicalRootPrompt(systemPrompt: string | any[] | undefi
338
377
  return systemPrompt;
339
378
  }
340
379
 
380
+ export function isDeepSeekRequest(event: any, ctx: any): boolean {
381
+ const modelId = (
382
+ (event as any)?.model?.id ||
383
+ (ctx as any)?.model?.id ||
384
+ (event as any)?.payload?.model ||
385
+ ""
386
+ ).toLowerCase();
387
+ return modelId.includes("deepseek");
388
+ }
389
+ export const WAF_BLOCK_RE = /sensitive[_ ]words?[_ ]detected|content-blocked/i;
390
+ export const SENSITIVE_WORDS_RE = /sensitive[_ ]words?[_ ]detected/i;
391
+ export const REDACTED_NOTE = "[Message withheld by local policy]";
392
+
393
+ export function cleanJsonSchemaObject(schema: any): void {
394
+ if (!schema || typeof schema !== "object") return;
395
+
396
+ if (schema.type === "object" || schema.properties) {
397
+ if (schema.required === null || schema.required === undefined || !Array.isArray(schema.required)) {
398
+ schema.required = [];
399
+ }
400
+ }
401
+
402
+ if (schema.properties && typeof schema.properties === "object") {
403
+ for (const [propName, propDef] of Object.entries(schema.properties)) {
404
+ if (propDef && typeof propDef === "object") {
405
+ cleanJsonSchemaObject(propDef);
406
+ }
407
+ }
408
+ }
409
+
410
+ if (schema.items) {
411
+ if (typeof schema.items === "object") {
412
+ cleanJsonSchemaObject(schema.items);
413
+ } else if (schema.items === null) {
414
+ delete schema.items;
415
+ }
416
+ }
417
+ }
418
+
419
+ export function sanitizeOpenAiTools(tools: any[]): void {
420
+ if (!Array.isArray(tools)) return;
421
+ for (const tool of tools) {
422
+ if (!tool || typeof tool !== "object") continue;
423
+ const fn = tool.function || tool;
424
+ if (fn.parameters && typeof fn.parameters === "object") {
425
+ cleanJsonSchemaObject(fn.parameters);
426
+ }
427
+ }
428
+ }
429
+
430
+ export function isRecord(v: unknown): v is Record<string, unknown> {
431
+ return typeof v === "object" && v !== null && !Array.isArray(v);
432
+ }
433
+
434
+ export const redactSet = new Set<string>();
435
+ export let escalatePending = false;
436
+ export let isSensitiveBlock = false;
437
+ export let exhausted = false;
438
+ export let sessionAnchor: string | null = null;
439
+ export let wafNotified = false;
440
+
441
+ export function resetPoisonRedactionState(): void {
442
+ redactSet.clear();
443
+ escalatePending = false;
444
+ isSensitiveBlock = false;
445
+ exhausted = false;
446
+ sessionAnchor = null;
447
+ wafNotified = false;
448
+ }
449
+
450
+ export function triggerEscalation(sensitive: boolean = false): void {
451
+ escalatePending = true;
452
+ isSensitiveBlock = sensitive;
453
+ }
454
+
455
+
456
+ export function fingerprintOf(msg: Record<string, unknown>): string {
457
+ const tc = Array.isArray(msg.tool_calls) ? JSON.stringify(msg.tool_calls).slice(0, 80) : "";
458
+ if (msg.type === "function_call" || msg.type === "function_call_output") {
459
+ return `${msg.type}:${(msg as any).call_id}:${JSON.stringify((msg as any).arguments ?? (msg as any).output ?? "").slice(0, 160)}`;
460
+ }
461
+ return `${msg.role ?? msg.type}:${JSON.stringify(msg.content ?? "").slice(0, 160)}:${tc}`;
462
+ }
463
+
464
+ export function firstUserAnchor(messages: unknown[]): string {
465
+ for (const m of messages) {
466
+ if (isRecord(m) && isHumanUserMessage(m as Record<string, unknown>)) return fingerprintOf(m);
467
+ }
468
+ return String(messages.length);
469
+ }
470
+
471
+ export function isTextBlock(block: Record<string, unknown>): boolean {
472
+ return block.type === "text" || block.type === "input_text" || block.type === "output_text";
473
+ }
474
+
475
+ export function isHumanUserMessage(msg: Record<string, unknown>): boolean {
476
+ if (msg.role !== "user") return false;
477
+ if (typeof msg.content === "string") return true;
478
+ if (Array.isArray(msg.content)) {
479
+ const hasToolResult = msg.content.some((b) => isRecord(b) && b.type === "tool_result");
480
+ const hasHumanText = msg.content.some(
481
+ (b) =>
482
+ isRecord(b) &&
483
+ (b.type === "text" || b.type === "input_text") &&
484
+ typeof b.text === "string" &&
485
+ b.text.trim().length > 0 &&
486
+ b.text !== REDACTED_NOTE
487
+ );
488
+ if (hasToolResult && !hasHumanText) return false;
489
+ return true;
490
+ }
491
+ return false;
492
+ }
493
+
494
+ export function lastHumanUserIndex(messages: unknown[]): number {
495
+ for (let i = messages.length - 1; i >= 0; i--) {
496
+ if (isRecord(messages[i]) && isHumanUserMessage(messages[i] as Record<string, unknown>)) {
497
+ return i;
498
+ }
499
+ }
500
+ return -1;
501
+ }
502
+
503
+ export function hasRedactableText(content: unknown, msg?: Record<string, unknown>): boolean {
504
+ if (msg) {
505
+ if (msg.details && (typeof msg.details === "string" || Object.keys(msg.details).length > 0)) return true;
506
+ if (typeof (msg as any).reasoning_content === "string" && (msg as any).reasoning_content.length > 0) return true;
507
+ if (typeof (msg as any).thinking === "string" && (msg as any).thinking.length > 0) return true;
508
+ }
509
+ if (msg?.type === "function_call") return typeof (msg as any).arguments === "string" && (msg as any).arguments !== "{}";
510
+ if (msg?.type === "function_call_output") return hasRedactableText((msg as any).output);
511
+ if (typeof content === "string") return content.length > 0 && content !== REDACTED_NOTE;
512
+ if (Array.isArray(content)) {
513
+ for (const b of content) {
514
+ if (!isRecord(b)) continue;
515
+ if (b.type === "thinking" || b.type === "reasoning") return true;
516
+ if (isTextBlock(b) && typeof b.text === "string" && b.text.length > 0 && b.text !== REDACTED_NOTE) return true;
517
+ if (b.type === "tool_result") {
518
+ if (typeof b.content === "string" && b.content.length > 0 && b.content !== REDACTED_NOTE) return true;
519
+ if (Array.isArray(b.content)) {
520
+ for (const c of b.content) {
521
+ if (isRecord(c) && isTextBlock(c) && typeof c.text === "string" && c.text.length > 0 && c.text !== REDACTED_NOTE) {
522
+ return true;
523
+ }
524
+ }
525
+ }
526
+ }
527
+ if (b.type === "tool_use" && isRecord(b.input) && Object.keys(b.input).length > 0) {
528
+ return true;
529
+ }
530
+ if (b.type === "toolCall" && isRecord(b.arguments) && Object.keys(b.arguments).length > 0) {
531
+ return true;
532
+ }
533
+ }
534
+ }
535
+ if (msg && Array.isArray(msg.tool_calls) && msg.tool_calls.length > 0) {
536
+ for (const tc of msg.tool_calls) {
537
+ if (isRecord(tc) && isRecord(tc.function) && typeof tc.function.arguments === "string" && tc.function.arguments !== "{}") {
538
+ return true;
539
+ }
540
+ }
541
+ }
542
+ return false;
543
+ }
544
+
545
+ export function redactBlocks(content: unknown): void {
546
+ if (!Array.isArray(content)) return;
547
+ for (let idx = content.length - 1; idx >= 0; idx--) {
548
+ const block = content[idx];
549
+ if (!isRecord(block)) continue;
550
+
551
+ // For thinking / reasoning blocks, keep a harmless placeholder rather than deleting,
552
+ // because thinking models (like DeepSeek) strictly require thinking blocks to be present in multi-turn history
553
+ if (block.type === "thinking" || block.type === "reasoning") {
554
+ if (typeof (block as any).thinking === "string") {
555
+ (block as any).thinking = "Thinking...";
556
+ }
557
+ if (typeof (block as any).text === "string") {
558
+ (block as any).text = "Thinking...";
559
+ }
560
+ continue;
561
+ }
562
+
563
+ if (isTextBlock(block) && typeof block.text === "string") {
564
+ block.text = REDACTED_NOTE;
565
+ } else if (block.type === "tool_result") {
566
+ if (typeof block.content === "string") {
567
+ block.content = REDACTED_NOTE;
568
+ } else if (Array.isArray(block.content)) {
569
+ for (const b of block.content) {
570
+ if (isRecord(b) && isTextBlock(b) && typeof b.text === "string") {
571
+ b.text = REDACTED_NOTE;
572
+ }
573
+ }
574
+ }
575
+ } else if (block.type === "tool_use") {
576
+ block.input = {};
577
+ } else if (block.type === "toolCall") {
578
+ block.arguments = {};
579
+ }
580
+ }
581
+ }
582
+
583
+ export function isHideable(msg: Record<string, unknown>): boolean {
584
+ return (
585
+ msg.role === "user" ||
586
+ msg.role === "assistant" ||
587
+ msg.role === "tool" ||
588
+ msg.role === "toolResult" ||
589
+ msg.type === "function_call" ||
590
+ msg.type === "function_call_output"
591
+ );
592
+ }
593
+
594
+ export function redactMessageAt(messages: unknown[], i: number): void {
595
+ const msg = messages[i];
596
+ if (!isRecord(msg)) return;
597
+
598
+ // Clear tool result details / raw diffs / attachments
599
+ if ("details" in msg) delete msg.details;
600
+ // DeepSeek and reasoning models require reasoning_content/thinking to exist in thinking mode.
601
+ // Never delete them completely; keep a sanitized minimal placeholder.
602
+ if ("reasoning_content" in msg) {
603
+ (msg as any).reasoning_content = "Thinking...";
604
+ }
605
+ if ("thinking" in msg) {
606
+ (msg as any).thinking = "Thinking...";
607
+ }
608
+
609
+ if (msg.type === "function_call") {
610
+ (msg as any).arguments = "{}";
611
+ return;
612
+ }
613
+ if (msg.type === "function_call_output") {
614
+ if (typeof (msg as any).output === "string") (msg as any).output = REDACTED_NOTE;
615
+ else redactBlocks((msg as any).output);
616
+ return;
617
+ }
618
+ if (typeof msg.content === "string") {
619
+ msg.content = REDACTED_NOTE;
620
+ } else if (Array.isArray(msg.content)) {
621
+ redactBlocks(msg.content);
622
+ if (msg.content.length === 0 && !Array.isArray(msg.tool_calls)) {
623
+ msg.content = [{ type: "text", text: REDACTED_NOTE }];
624
+ }
625
+ } else if (msg.role === "tool" || msg.role === "toolResult" || msg.role === "assistant") {
626
+ msg.content = REDACTED_NOTE;
627
+ }
628
+ if (Array.isArray(msg.tool_calls)) {
629
+ for (const tc of msg.tool_calls) {
630
+ if (isRecord(tc) && isRecord(tc.function) && typeof tc.function.arguments === "string") {
631
+ tc.function.arguments = "{}";
632
+ }
633
+ }
634
+ }
635
+ }
636
+
637
+ export function applyPoisonRedaction(payload: Record<string, unknown>): void {
638
+ const messages = Array.isArray(payload.messages) ? payload.messages : (payload as any).input;
639
+ if (!Array.isArray(messages) || messages.length === 0) return;
640
+
641
+ // Prune failed assistant messages carrying error status/text in-place so dead error turns do not linger
642
+ for (let i = messages.length - 1; i >= 0; i--) {
643
+ const m = messages[i];
644
+ if (!isRecord(m)) continue;
645
+ if (m.role === "assistant") {
646
+ if (m.stopReason === "error") {
647
+ messages.splice(i, 1);
648
+ } else if (typeof m.content === "string" && WAF_BLOCK_RE.test(m.content)) {
649
+ messages.splice(i, 1);
650
+ }
651
+ }
652
+ }
653
+
654
+ const fps = messages.map((m) => (isRecord(m) ? fingerprintOf(m) : ""));
655
+
656
+ const anchor = firstUserAnchor(messages);
657
+ if (anchor !== sessionAnchor) {
658
+ const firstContact = sessionAnchor === null;
659
+ sessionAnchor = anchor;
660
+ if (!firstContact) {
661
+ redactSet.clear();
662
+ escalatePending = false;
663
+ isSensitiveBlock = false;
664
+ exhausted = false;
665
+ wafNotified = false;
666
+ }
667
+ }
668
+
669
+ if (redactSet.size > 0 && !fps.some((fp) => fp && redactSet.has(fp))) {
670
+ redactSet.clear();
671
+ }
672
+
673
+ const lastHumanUser = lastHumanUserIndex(messages);
674
+
675
+ for (let i = 0; i < messages.length; i++) {
676
+ if (i === lastHumanUser) continue;
677
+ if (!isRecord(messages[i]) || !isHideable(messages[i] as Record<string, unknown>)) continue;
678
+ if (redactSet.has(fps[i])) redactMessageAt(messages, i);
679
+ }
680
+
681
+ if (escalatePending) {
682
+ escalatePending = false;
683
+ const sensitive = isSensitiveBlock;
684
+ isSensitiveBlock = false;
685
+
686
+ if (sensitive) {
687
+ // Sensitive words detected: neutralize ALL turns (older user, tool, assistant, active tool results) except active human user
688
+ for (let i = 0; i < messages.length; i++) {
689
+ if (i === lastHumanUser) continue;
690
+ if (!isRecord(messages[i])) continue;
691
+ const m = messages[i] as Record<string, unknown>;
692
+ if (!isHideable(m)) continue;
693
+ redactSet.add(fps[i]);
694
+ if (hasRedactableText(m.content, m)) {
695
+ redactMessageAt(messages, i);
696
+ }
697
+ }
698
+ } else {
699
+ // Language ratio block: Stage 1 older user turns, Stage 2 assistant & tool
700
+ let anyRedacted = false;
701
+ for (let i = 0; i < messages.length; i++) {
702
+ if (i === lastHumanUser) continue;
703
+ if (!isRecord(messages[i])) continue;
704
+ const m = messages[i] as Record<string, unknown>;
705
+ if (m.role !== "user") continue;
706
+ if (redactSet.has(fps[i])) continue;
707
+ redactSet.add(fps[i]);
708
+ if (hasRedactableText(m.content, m)) {
709
+ redactMessageAt(messages, i);
710
+ anyRedacted = true;
711
+ }
712
+ }
713
+ if (!anyRedacted) {
714
+ for (let i = 0; i < messages.length; i++) {
715
+ if (i === lastHumanUser) continue;
716
+ if (!isRecord(messages[i])) continue;
717
+ const m = messages[i] as Record<string, unknown>;
718
+ if (!isHideable(m) || m.role === "user") continue;
719
+ if (redactSet.has(fps[i])) continue;
720
+ redactSet.add(fps[i]);
721
+ if (hasRedactableText(m.content, m)) {
722
+ redactMessageAt(messages, i);
723
+ anyRedacted = true;
724
+ }
725
+ }
726
+ }
727
+ }
728
+ }
729
+
730
+ // Exhausted if there are no more hideable messages left to redact other than lastHumanUser
731
+ exhausted = true;
732
+ for (let i = 0; i < messages.length; i++) {
733
+ if (i === lastHumanUser) continue;
734
+ if (!isRecord(messages[i]) || !isHideable(messages[i] as Record<string, unknown>)) continue;
735
+ if (!redactSet.has(fps[i])) {
736
+ exhausted = false;
737
+ break;
738
+ }
739
+ }
740
+ }
741
+
742
+ export function normalizeMessagesForAgentRouter(messages: any[], isDeepSeek: boolean = false): void {
743
+ if (!Array.isArray(messages)) return;
744
+
745
+ // Prune poisoned/failed assistant turns globally across all models (e.g. prior 402 quota failure, network errors, empty turns)
746
+ for (let i = messages.length - 1; i >= 0; i--) {
747
+ const msg = messages[i];
748
+ if (!msg || typeof msg !== "object") continue;
749
+ if (msg.role === "assistant") {
750
+ const hasToolCalls = Array.isArray(msg.tool_calls) && msg.tool_calls.length > 0;
751
+ const hasToolUseBlocks =
752
+ Array.isArray(msg.content) &&
753
+ msg.content.some((b: any) => b && typeof b === "object" && (b.type === "tool_use" || b.type === "tool_call"));
754
+ const hasTools = hasToolCalls || hasToolUseBlocks;
755
+ const isEmptyContent =
756
+ !msg.content ||
757
+ msg.content === "" ||
758
+ (Array.isArray(msg.content) && (
759
+ msg.content.length === 0 ||
760
+ msg.content.every((b: any) => b && typeof b === "object" && (b.type === "text" || b.type === "input_text") && (!b.text || b.text.trim() === ""))
761
+ ));
762
+
763
+ if (
764
+ msg.stopReason === "error" ||
765
+ (msg.stopReason === "aborted" && isEmptyContent) ||
766
+ (typeof msg.content === "string" && WAF_BLOCK_RE.test(msg.content)) ||
767
+ (isEmptyContent && !hasTools && !msg.reasoning_content)
768
+ ) {
769
+ messages.splice(i, 1);
770
+ }
771
+ }
772
+ }
773
+
774
+
775
+ for (const msg of messages) {
776
+ if (!msg || typeof msg !== "object") continue;
777
+ if (msg.role === "developer") {
778
+ msg.role = "system";
779
+ }
780
+ if (msg.role === "assistant") {
781
+ let extractedThinking: string | undefined;
782
+
783
+ if (Array.isArray(msg.content)) {
784
+ const thinkingParts: string[] = [];
785
+ const nonThinkingParts: any[] = [];
786
+
787
+ for (const part of msg.content) {
788
+ if (part && typeof part === "object" && (part.type === "thinking" || part.type === "reasoning")) {
789
+ const text = part.thinking || part.text;
790
+ if (text) thinkingParts.push(text);
791
+ } else {
792
+ nonThinkingParts.push(part);
793
+ }
794
+ }
795
+
796
+ if (thinkingParts.length > 0) {
797
+ extractedThinking = thinkingParts.join("\n");
798
+ }
799
+
800
+ if (nonThinkingParts.length === 0) {
801
+ msg.content = "";
802
+ } else if (nonThinkingParts.length === 1 && nonThinkingParts[0].type === "text") {
803
+ msg.content = nonThinkingParts[0].text;
804
+ } else {
805
+ msg.content = nonThinkingParts;
806
+ }
807
+ }
808
+
809
+ if (extractedThinking && !msg.reasoning_content) {
810
+ msg.reasoning_content = extractedThinking;
811
+ }
812
+
813
+ // If assistant executed tool calls, AgentRouter DeepSeek proxy strictly requires reasoning_content
814
+ if (
815
+ Array.isArray(msg.tool_calls) &&
816
+ msg.tool_calls.length > 0 &&
817
+ (!msg.reasoning_content || (typeof msg.reasoning_content === "string" && !msg.reasoning_content.trim()))
818
+ ) {
819
+ msg.reasoning_content = "Executing tools...";
820
+ }
821
+
822
+ // If DeepSeek model and content is still empty without tool calls, provide fallback text
823
+ if (isDeepSeek && (!msg.content || msg.content === "") && (!Array.isArray(msg.tool_calls) || msg.tool_calls.length === 0)) {
824
+ msg.content = "[Interrupted]";
825
+ if (!msg.reasoning_content) {
826
+ msg.reasoning_content = "Interrupted response";
827
+ }
828
+ }
829
+ }
830
+
831
+ if (isDeepSeek) {
832
+ // Flatten past assistant tool_calls into text and convert tool roles into user turns.
833
+ // This prevents AgentRouter's upstream Anthropic gateway from rejecting the request with:
834
+ // "400: The `content[].thinking` in the thinking mode must be passed back to the API."
835
+ if (msg.role === "assistant" && Array.isArray(msg.tool_calls) && msg.tool_calls.length > 0) {
836
+ let contentStr = typeof msg.content === "string" ? msg.content : "";
837
+ for (const tc of msg.tool_calls) {
838
+ const fnName = tc?.function?.name || tc?.name || "tool";
839
+ const fnArgs = tc?.function?.arguments || "{}";
840
+ contentStr += (contentStr ? "\n" : "") + `[Tool Call]: ${fnName}(${fnArgs})`;
841
+ }
842
+ msg.content = contentStr;
843
+ delete msg.tool_calls;
844
+ if (!msg.reasoning_content || (typeof msg.reasoning_content === "string" && !msg.reasoning_content.trim())) {
845
+ msg.reasoning_content = "Executing tools...";
846
+ }
847
+ }
848
+
849
+ if (msg.role === "tool" || msg.role === "toolResult") {
850
+ const text = typeof msg.content === "string" ? msg.content : "";
851
+ const toolName = (msg as any).toolName || (msg as any).name || "tool";
852
+ msg.role = "user";
853
+ msg.content = `[Tool Result for ${toolName}]:\n${text}`;
854
+ delete msg.tool_call_id;
855
+ delete (msg as any).toolCallId;
856
+ }
857
+
858
+ }
859
+ }
860
+ }
861
+
862
+ export function cleanupDeepSeekDuplicates(): {
863
+ cleanedSettings: boolean;
864
+ cleanedCache: boolean;
865
+ cleanedModelsJson: boolean;
866
+ } {
867
+ let cleanedSettings = false;
868
+ let cleanedCache = false;
869
+ let cleanedModelsJson = false;
870
+
871
+ // 1. Clean settings.json
872
+ try {
873
+ if (fs.existsSync(SETTINGS_FILE)) {
874
+ const settings = JSON.parse(fs.readFileSync(SETTINGS_FILE, "utf-8"));
875
+ let modified = false;
876
+
877
+ if (Array.isArray(settings.enabledModels)) {
878
+ const clodeIdx = settings.enabledModels.indexOf("agentrouter-clode/deepseek-v4-flash");
879
+ if (clodeIdx !== -1) {
880
+ settings.enabledModels.splice(clodeIdx, 1);
881
+ modified = true;
882
+ }
883
+ // Deduplicate enabledModels
884
+ const seen = new Set<string>();
885
+ const deduped: string[] = [];
886
+ for (const m of settings.enabledModels) {
887
+ if (!seen.has(m)) {
888
+ seen.add(m);
889
+ deduped.push(m);
890
+ } else {
891
+ modified = true;
892
+ }
893
+ }
894
+ settings.enabledModels = deduped;
895
+ }
896
+
897
+ if (settings.subagents && typeof settings.subagents === "object" && settings.subagents.agentOverrides) {
898
+ for (const override of Object.values(settings.subagents.agentOverrides)) {
899
+ if ((override as any)?.model === "agentrouter-clode/deepseek-v4-flash") {
900
+ (override as any).model = "agentrouter-openai/deepseek-v4-flash";
901
+ modified = true;
902
+ }
903
+ }
904
+ }
905
+
906
+ if (settings.defaultProvider === "agentrouter-clode" && settings.defaultModel === "deepseek-v4-flash") {
907
+ settings.defaultProvider = "agentrouter-openai";
908
+ modified = true;
909
+ }
910
+
911
+ if (modified) {
912
+ fs.writeFileSync(SETTINGS_FILE, JSON.stringify(settings, null, 2), "utf-8");
913
+ cleanedSettings = true;
914
+ }
915
+ }
916
+ } catch {}
917
+
918
+ // 2. Clean MODELS_CACHE_FILE (.agentrouter-models-cache.json)
919
+ try {
920
+ if (fs.existsSync(MODELS_CACHE_FILE)) {
921
+ const cacheData = JSON.parse(fs.readFileSync(MODELS_CACHE_FILE, "utf-8"));
922
+ if (Array.isArray(cacheData)) {
923
+ let modified = false;
924
+ for (const item of cacheData) {
925
+ if (item && item.model_name === "deepseek-v4-flash") {
926
+ if (Array.isArray(item.supported_endpoint_types) && item.supported_endpoint_types.includes("anthropic")) {
927
+ item.supported_endpoint_types = item.supported_endpoint_types.filter((t: string) => t !== "anthropic");
928
+ modified = true;
929
+ }
930
+ }
931
+ }
932
+ if (modified) {
933
+ fs.writeFileSync(MODELS_CACHE_FILE, JSON.stringify(cacheData, null, 2), "utf-8");
934
+ cleanedCache = true;
935
+ }
936
+ }
937
+ }
938
+ } catch {}
939
+
940
+ // 3. Clean models.json (if present)
941
+ try {
942
+ const modelsJsonPath = path.join(process.env.HOME || "", ".pi/agent/models.json");
943
+ if (fs.existsSync(modelsJsonPath)) {
944
+ const modelsJson = JSON.parse(fs.readFileSync(modelsJsonPath, "utf-8"));
945
+ let modified = false;
946
+ if (modelsJson?.providers?.["agentrouter-clode"]?.models) {
947
+ const origLen = modelsJson.providers["agentrouter-clode"].models.length;
948
+ modelsJson.providers["agentrouter-clode"].models = modelsJson.providers["agentrouter-clode"].models.filter(
949
+ (m: any) => (typeof m === "string" ? m !== "deepseek-v4-flash" : m?.id !== "deepseek-v4-flash")
950
+ );
951
+ if (modelsJson.providers["agentrouter-clode"].models.length !== origLen) {
952
+ modified = true;
953
+ }
954
+ }
955
+ if (modified) {
956
+ fs.writeFileSync(modelsJsonPath, JSON.stringify(modelsJson, null, 2), "utf-8");
957
+ cleanedModelsJson = true;
958
+ }
959
+ }
960
+ } catch {}
961
+
962
+ return { cleanedSettings, cleanedCache, cleanedModelsJson };
963
+ }
964
+
965
+ let lastWarnedErrorTimestamp = 0;
966
+ export function checkAndNotifyContentBlocked(errMessage: string | undefined, ctx: any): void {
967
+ if (!errMessage) return;
968
+ const now = Date.now();
969
+ if (now - lastWarnedErrorTimestamp < 2000) return;
970
+ if (errMessage.includes("content-blocked")) {
971
+ lastWarnedErrorTimestamp = now;
972
+ const tip = "[AgentRouter] Request was blocked by upstream gateway content filter (content-blocked).";
973
+ if (ctx?.hasUI) {
974
+ ctx.ui.notify(tip, "warning");
975
+ } else {
976
+ console.warn(`\nāš ļø ${tip}\n`);
977
+ }
978
+ }
979
+ }
980
+
341
981
  export async function fetchLivePricing(): Promise<ApiPricingModel[] | null> {
342
982
  try {
343
983
  const res = await fetch("https://agentrouter.org/api/pricing", {
344
- headers: { "User-Agent": "pi-code" },
984
+
345
985
  });
346
986
  if (!res.ok) return null;
347
987
  const data = await res.json();
@@ -365,8 +1005,7 @@ export async function fetchTokenUsage(apiKey: string): Promise<number | null> {
365
1005
  {
366
1006
  headers: {
367
1007
  Authorization: `Bearer ${apiKey}`,
368
- "User-Agent": "pi-code",
369
- },
1008
+ },
370
1009
  }
371
1010
  );
372
1011
  if (!res.ok) return null;
@@ -392,12 +1031,12 @@ export async function probeModelQuota(modelId: string, apiKey: string, isAnthrop
392
1031
  "Content-Type": "application/json",
393
1032
  "x-api-key": apiKey,
394
1033
  "anthropic-version": "2023-06-01",
395
- "User-Agent": "pi-code",
1034
+ "User-Agent": getPiUserAgent(),
396
1035
  }
397
1036
  : {
398
1037
  "Content-Type": "application/json",
399
1038
  Authorization: `Bearer ${apiKey}`,
400
- "User-Agent": "pi-code",
1039
+ "User-Agent": getPiUserAgent(),
401
1040
  };
402
1041
 
403
1042
  const body = isAnthropic
@@ -427,10 +1066,12 @@ export async function probeModelQuota(modelId: string, apiKey: string, isAnthrop
427
1066
  return { model: modelId, status: "READY", code: res.status };
428
1067
  }
429
1068
 
430
- const data = await res.json().catch(() => ({}));
431
- const msg = data.error?.message || data.message || "";
432
-
433
- if (res.status === 402 || msg.toLowerCase().includes("quota") || msg.toLowerCase().includes("exhausted")) {
1069
+ if (
1070
+ res.status === 402 ||
1071
+ msg.toLowerCase().includes("quota") ||
1072
+ msg.toLowerCase().includes("exhausted") ||
1073
+ msg.toLowerCase().includes("budget pool")
1074
+ ) {
434
1075
  return { model: modelId, status: "QUOTA_EXHAUSTED", code: 402, message: msg };
435
1076
  }
436
1077
  if (res.status === 403) {
@@ -572,6 +1213,7 @@ export default function (pi: ExtensionAPI) {
572
1213
  return newModels;
573
1214
  }
574
1215
 
1216
+ cleanupDeepSeekDuplicates();
575
1217
  const cachedPricing = loadCachedPricing();
576
1218
  registerAgentRouterProviders(currentApiKey, cachedPricing);
577
1219
 
@@ -598,16 +1240,19 @@ export default function (pi: ExtensionAPI) {
598
1240
 
599
1241
  const payload = event.payload;
600
1242
  if (payload) {
1243
+ const isDeepSeek = isDeepSeekRequest(event, ctx);
1244
+
601
1245
  if (payload.system !== undefined) {
602
1246
  payload.system = enforceCanonicalRootPrompt(payload.system);
603
1247
  }
604
- if (Array.isArray(payload.messages) && payload.messages.length > 0) {
605
- for (const msg of payload.messages) {
606
- if (msg && msg.role === "developer") {
607
- msg.role = "system";
608
- }
609
- }
610
- const firstMsg = payload.messages[0];
1248
+
1249
+ applyPoisonRedaction(payload);
1250
+
1251
+ const messages = Array.isArray(payload.messages) ? payload.messages : payload.input;
1252
+ if (Array.isArray(messages) && messages.length > 0) {
1253
+ normalizeMessagesForAgentRouter(messages, isDeepSeek);
1254
+
1255
+ const firstMsg = messages[0];
611
1256
  if (firstMsg && (firstMsg.role === "system" || firstMsg.role === "developer")) {
612
1257
  firstMsg.role = "system";
613
1258
  if (typeof firstMsg.content === "string") {
@@ -617,16 +1262,73 @@ export default function (pi: ExtensionAPI) {
617
1262
  }
618
1263
  }
619
1264
  }
1265
+ if (Array.isArray(payload.tools) && payload.tools.length > 0) {
1266
+ sanitizeOpenAiTools(payload.tools);
1267
+ }
620
1268
  }
621
1269
  }
622
1270
  return undefined;
623
1271
  });
624
1272
 
625
- pi.on("message_end", async (_event, ctx) => {
626
- const model = ctx?.model;
627
- if (isAgentRouter(model?.provider, (model as any)?.baseUrl)) {
1273
+ pi.on("message_end", async (event, ctx) => {
1274
+ const message = (event as any)?.message;
1275
+ if (!message) return;
1276
+
1277
+ if (message.role === "assistant" && message.stopReason !== "error") {
1278
+ wafNotified = false;
1279
+ isSensitiveBlock = false;
1280
+ }
1281
+
1282
+ const provider = message.provider ?? ctx?.model?.provider;
1283
+ const baseUrl = (message as any)?.baseUrl ?? (ctx?.model as any)?.baseUrl;
1284
+ if (isAgentRouter(provider, baseUrl)) {
628
1285
  setLastRequestEndTime(Date.now());
629
1286
  }
1287
+
1288
+ if (message.role !== "assistant" || message.stopReason !== "error") return;
1289
+ if (!isAgentRouter(provider, baseUrl)) return;
1290
+
1291
+ const errorMessage = message.errorMessage ?? "";
1292
+ if (!WAF_BLOCK_RE.test(errorMessage)) return;
1293
+
1294
+ escalatePending = true;
1295
+ if (SENSITIVE_WORDS_RE.test(errorMessage)) {
1296
+ isSensitiveBlock = true;
1297
+ }
1298
+
1299
+ if (exhausted) {
1300
+ if (!wafNotified) {
1301
+ wafNotified = true;
1302
+ const warning =
1303
+ "AgentRouter content filter keeps blocking even with earlier messages hidden. " +
1304
+ "Your latest message is likely the trigger — please rephrase or split it.";
1305
+ if (ctx?.hasUI) {
1306
+ ctx.ui.notify(warning, "warning");
1307
+ } else {
1308
+ console.warn(`\nāš ļø ${warning}\n`);
1309
+ }
1310
+ }
1311
+ return;
1312
+ }
1313
+
1314
+ if (!wafNotified) {
1315
+ wafNotified = true;
1316
+ const note = isSensitiveBlock
1317
+ ? "Sensitive words detected in agent activity. Neutralizing previous leftovers so subsequent chats can proceed safely."
1318
+ : "AgentRouter content filter blocked the request. Retrying automatically with earlier messages hidden.";
1319
+ if (ctx?.hasUI) {
1320
+ ctx.ui.notify(note, "warning");
1321
+ } else {
1322
+ console.warn(`\nāš ļø ${note}\n`);
1323
+ }
1324
+ }
1325
+
1326
+ return {
1327
+ message: {
1328
+ ...message,
1329
+ errorMessage: `${errorMessage} (provider returned error — retrying with earlier messages hidden)`,
1330
+ },
1331
+ };
630
1332
  });
631
1333
 
632
1334
  pi.on("turn_end", async (_event, ctx) => {
@@ -645,13 +1347,18 @@ export default function (pi: ExtensionAPI) {
645
1347
 
646
1348
  pi.on("session_start", async (_event, ctx) => {
647
1349
  currentApiKey = getEffectiveApiKey();
1350
+ resetPoisonRedactionState();
1351
+ cleanupDeepSeekDuplicates();
648
1352
  updatePromptRewriteEnvForModel(ctx.model);
649
1353
  setLastRequestEndTime(Date.now());
1354
+ syncEnabledModelsInSettings();
650
1355
 
651
1356
  fetchLivePricing().then((livePricing) => {
652
1357
  if (livePricing) {
653
1358
  saveCachedPricing(livePricing);
654
1359
  const newModels = registerAgentRouterProviders(currentApiKey, livePricing);
1360
+ cleanupDeepSeekDuplicates();
1361
+ syncEnabledModelsInSettings();
655
1362
  if (newModels.length > 0 && ctx.hasUI) {
656
1363
  ctx.ui.notify(
657
1364
  `[AgentRouter] Discovered new models on gateway: ${newModels.join(", ")}.\n` +
@@ -798,10 +1505,21 @@ export default function (pi: ExtensionAPI) {
798
1505
  currentApiKey = cleanKey;
799
1506
  saveConfig({ apiKey: cleanKey, minIntervalMs });
800
1507
  registerAgentRouterProviders(cleanKey, loadCachedPricing());
1508
+ syncEnabledModelsInSettings();
801
1509
  ctx.ui.notify("AgentRouter API key updated successfully for all models.", "info");
802
1510
  return;
803
1511
  }
804
1512
 
1513
+ if (action === "sync" || action === "enable-models") {
1514
+ const res = syncEnabledModelsInSettings();
1515
+ if (res.added.length > 0) {
1516
+ ctx.ui.notify(`[AgentRouter] Added to enabledModels: ${res.added.join(", ")} (total: ${res.count}).`, "info");
1517
+ } else {
1518
+ ctx.ui.notify(`[AgentRouter] All flagship models already enabled in settings.json (total: ${res.count}).`, "info");
1519
+ }
1520
+ return;
1521
+ }
1522
+
805
1523
  if (action === "fix-order" || action === "order") {
806
1524
  const order = getPackageOrderState();
807
1525
  if (!order.needsFix) {
@@ -893,10 +1611,10 @@ export default function (pi: ExtensionAPI) {
893
1611
  }))
894
1612
  : [
895
1613
  { id: "deepseek-v4-flash", isAnthropic: false },
896
- { id: "glm-5.3", isAnthropic: false },
1614
+ { id: "gpt-6-astra", isAnthropic: false },
897
1615
  { id: "gpt-5.6-sol", isAnthropic: false },
898
- { id: "claude-opus-4-8", isAnthropic: true },
899
1616
  { id: "claude-opus-5", isAnthropic: true },
1617
+ { id: "claude-opus-4-8", isAnthropic: true },
900
1618
  ];
901
1619
 
902
1620
  const probeResults = await Promise.all(
@@ -924,7 +1642,7 @@ export default function (pi: ExtensionAPI) {
924
1642
  }
925
1643
 
926
1644
  report +=
927
- `\nTip: Claude and GPT models use daily batch quotas. If exhausted, switch to DeepSeek V4 or GLM 5.3 which have unlimited availability.`;
1645
+ `\nTip: Claude and GPT models use daily batch quotas. If exhausted, switch to DeepSeek V4 Flash which has unlimited availability.`;
928
1646
 
929
1647
  ctx.ui.notify(report, "info");
930
1648
  return;
@@ -948,7 +1666,7 @@ export default function (pi: ExtensionAPI) {
948
1666
  }
949
1667
 
950
1668
  ctx.ui.notify(
951
- `[AgentRouter Plugin v2.0.0]\n` +
1669
+ `[AgentRouter Plugin v2.1.2]\n` +
952
1670
  `- Active model: ${activeModel?.id || "none"} (${isAR ? "AgentRouter [yes]" : "Other Provider"})\n` +
953
1671
  `- Package Priority: ${priorityStatus}\n` +
954
1672
  `- API Key: ${maskedKey}\n` +
@@ -956,6 +1674,7 @@ export default function (pi: ExtensionAPI) {
956
1674
  `- Commands:\n` +
957
1675
  ` /agentrouter check (probe live batch quotas & spending)\n` +
958
1676
  ` /agentrouter pricing (fetch live pricing table $/1M)\n` +
1677
+ ` /agentrouter sync (sync enabledModels in settings.json)\n` +
959
1678
  ` /agentrouter key <key> (update API key)\n` +
960
1679
  ` /agentrouter pacing <ms> (adjust rate limit delay)\n` +
961
1680
  ` /agentrouter fix-order (move plugin above pi-cache-optimizer)`,
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "@madgagarin/pi-agentrouter",
3
- "version": "2.0.0",
4
- "description": "Official Pi Coding Agent extension for AgentRouter (agentrouter.org). Connects GPT-5.6 Sol, Claude Opus 5, DeepSeek V4 Flash, and GLM 5.3 with live USD pricing, batch quota health probe, prompt caching, and WAF protection.",
3
+ "version": "2.1.2",
4
+ "description": "Official Pi Coding Agent extension for AgentRouter (agentrouter.org). Connects GPT-6 Astra, GPT-5.6 Sol, Claude Opus 5, DeepSeek V4 Flash, and GLM 5.3 with live USD pricing, auto-sync settings, batch quota probe, prompt caching, and WAF protection.",
5
5
  "publishConfig": {
6
6
  "access": "public"
7
7
  },
@@ -15,6 +15,8 @@
15
15
  "agentrouter-org",
16
16
  "agent-router",
17
17
  "pi-agentrouter",
18
+ "gpt-6-astra",
19
+ "gpt-6",
18
20
  "gpt-5.6-sol",
19
21
  "claude-opus-5",
20
22
  "claude-opus-4-8",