pi-zro-provider 1.3.11 → 1.3.12
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +5 -3
- package/models.json +29 -0
- package/package.json +2 -2
- package/patch.json +20 -0
package/README.md
CHANGED
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
**3+ models through [Zro](https://zro.moonmath.ai)**
|
|
6
6
|
|
|
7
|
-
_GLM-5.2, Kimi K3, and DeepSeek V4 Flash — run through the Zro inference endpoint with [pi](https://github.com/earendil-works/pi-coding-agent)._
|
|
7
|
+
_GLM-5.2, Kimi K3, and DeepSeek V4 / V4.1 Flash — run through the Zro inference endpoint with [pi](https://github.com/earendil-works/pi-coding-agent)._
|
|
8
8
|
|
|
9
9
|
[](https://github.com/earendil-works/pi-coding-agent)
|
|
10
10
|
[](./LICENSE)
|
|
@@ -15,8 +15,8 @@ _GLM-5.2, Kimi K3, and DeepSeek V4 Flash — run through the Zro inference endpo
|
|
|
15
15
|
|
|
16
16
|
## Features
|
|
17
17
|
|
|
18
|
-
- **
|
|
19
|
-
- **Native reasoning effort** — every model ships an endpoint-verified `piLevel` → level-id map (`glm-5.2`, `glm-5.3-flash`, `deepseek-v4-flash-0731`: full `off`–`max` ladder; `kimi-k3`: `low`/`high`/`max` — the proxy rejects `medium` there), so `/thinking low` etc. maps exactly to what the Zro proxy accepts. Quirks: `glm-5.3-flash` silently disables thinking at `low`, so `minimal`/`low` map to `minimal`; `glm-5.3` mirrors flash's full ladder — Z.ai officially removed `none` for the 5.3 family, but the proxy accepts it and streams cleanly (occasionally a short chain-of-thought preamble surfaces in `content` before the answer), and `xhigh`/`max` both send `max`
|
|
18
|
+
- **8+ AI Models** — GLM-5.2 (default), Kimi K3, DeepSeek V4 Flash, and DeepSeek V4.1 Flash, straight from Zro's production catalog
|
|
19
|
+
- **Native reasoning effort** — every model ships an endpoint-verified `piLevel` → level-id map (`glm-5.2`, `glm-5.3-flash`, `deepseek-v4-flash-0731`: full `off`–`max` ladder; `deepseek-v4.1-flash`: `low`/`high`/`max` — catalog has no `none`, and `xhigh`/`max` both send `max`; `kimi-k3`: `low`/`high`/`max` — the proxy rejects `medium` there), so `/thinking low` etc. maps exactly to what the Zro proxy accepts. Quirks: `glm-5.3-flash` silently disables thinking at `low`, so `minimal`/`low` map to `minimal`; `glm-5.3` mirrors flash's full ladder — Z.ai officially removed `none` for the 5.3 family, but the proxy accepts it and streams cleanly (occasionally a short chain-of-thought preamble surfaces in `content` before the answer), and `xhigh`/`max` both send `max`
|
|
20
20
|
- **OpenAI-compatible API** at `https://zro.moonmath.ai/v1`
|
|
21
21
|
- **Official catalog sync** from Zro's `/api/cli/models` endpoint — same one `zro models` uses
|
|
22
22
|
- **Zro login reuse** — if you've run `zro login`, the extension picks up `~/.config/zro/credentials.json` automatically (no duplicated keys)
|
|
@@ -28,6 +28,7 @@ _GLM-5.2, Kimi K3, and DeepSeek V4 Flash — run through the Zro inference endpo
|
|
|
28
28
|
|-------|------|---------|------------|------------|-------------|
|
|
29
29
|
| Auto | Text | 1.0M | 131K | — | — |
|
|
30
30
|
| DeepSeek V4 Flash | Text | 1.0M | 384K | $0.14 | $0.28 |
|
|
31
|
+
| DeepSeek V4.1 Flash | Text + Image | 1.0M | 384K | $0.14 | $0.28 |
|
|
31
32
|
| Dolly 1 | Text | 1.0M | 64K | — | — |
|
|
32
33
|
| GLM-5.2 | Text | 524K | 64K | $1.10 | $4.00 |
|
|
33
34
|
| GLM-5.3 | Text | 1.0M | 131K | $1.40 | $4.40 |
|
|
@@ -109,6 +110,7 @@ The available levels are model-specific and come straight from Zro's catalog:
|
|
|
109
110
|
| `glm-5.3` | `off`, `minimal`, `low`, `medium`, `high`, `xhigh`, `max` | `none`, `minimal`, `minimal`, `medium`, `high`, `max`, `max` |
|
|
110
111
|
| `kimi-k3` | `low`, `high`, `max` | `low`, `high`, `max` |
|
|
111
112
|
| `deepseek-v4-flash-0731` | `off`, `low`, `high`, `max` | `none`, `low`, `high`, `max` |
|
|
113
|
+
| `deepseek-v4.1-flash` | `low`, `high`, `xhigh`, `max` | `low`, `high`, `max`, `max` |
|
|
112
114
|
|
|
113
115
|
Levels map one-to-one through `thinkingLevelMap`, so `off` sends Zro's `none` token instead of silently dropping the field.
|
|
114
116
|
|
package/models.json
CHANGED
|
@@ -58,6 +58,35 @@
|
|
|
58
58
|
"maxTokensField": "max_tokens"
|
|
59
59
|
}
|
|
60
60
|
},
|
|
61
|
+
{
|
|
62
|
+
"id": "deepseek-v4.1-flash",
|
|
63
|
+
"name": "DeepSeek V4.1 Flash",
|
|
64
|
+
"reasoning": true,
|
|
65
|
+
"thinkingLevelMap": {
|
|
66
|
+
"off": null,
|
|
67
|
+
"minimal": null,
|
|
68
|
+
"low": "low",
|
|
69
|
+
"medium": null,
|
|
70
|
+
"high": "high",
|
|
71
|
+
"xhigh": "max"
|
|
72
|
+
},
|
|
73
|
+
"input": [
|
|
74
|
+
"text"
|
|
75
|
+
],
|
|
76
|
+
"cost": {
|
|
77
|
+
"input": 0,
|
|
78
|
+
"output": 0,
|
|
79
|
+
"cacheRead": 0,
|
|
80
|
+
"cacheWrite": 0
|
|
81
|
+
},
|
|
82
|
+
"contextWindow": 1048576,
|
|
83
|
+
"maxTokens": 384000,
|
|
84
|
+
"compat": {
|
|
85
|
+
"supportsDeveloperRole": false,
|
|
86
|
+
"supportsReasoningEffort": true,
|
|
87
|
+
"maxTokensField": "max_tokens"
|
|
88
|
+
}
|
|
89
|
+
},
|
|
61
90
|
{
|
|
62
91
|
"id": "dolly1",
|
|
63
92
|
"name": "Dolly 1",
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-zro-provider",
|
|
3
|
-
"version": "1.3.
|
|
4
|
-
"description": "Zro provider extension for pi - GLM-5.2, GLM-5.3 Flash, Kimi K3 & DeepSeek V4 Flash through the Zro inference endpoint",
|
|
3
|
+
"version": "1.3.12",
|
|
4
|
+
"description": "Zro provider extension for pi - GLM-5.2, GLM-5.3 Flash, Kimi K3, DeepSeek V4 Flash & DeepSeek V4.1 Flash through the Zro inference endpoint",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "index.ts",
|
|
7
7
|
"scripts": {
|
package/patch.json
CHANGED
|
@@ -15,6 +15,26 @@
|
|
|
15
15
|
"cacheRead": 0.0028
|
|
16
16
|
}
|
|
17
17
|
},
|
|
18
|
+
"deepseek-v4.1-flash": {
|
|
19
|
+
"input": [
|
|
20
|
+
"text",
|
|
21
|
+
"image"
|
|
22
|
+
],
|
|
23
|
+
"thinkingLevelMap": {
|
|
24
|
+
"off": null,
|
|
25
|
+
"minimal": null,
|
|
26
|
+
"low": "low",
|
|
27
|
+
"medium": null,
|
|
28
|
+
"high": "high",
|
|
29
|
+
"xhigh": "max",
|
|
30
|
+
"max": "max"
|
|
31
|
+
},
|
|
32
|
+
"cost": {
|
|
33
|
+
"input": 0.14,
|
|
34
|
+
"output": 0.28,
|
|
35
|
+
"cacheRead": 0.0028
|
|
36
|
+
}
|
|
37
|
+
},
|
|
18
38
|
"kimi-k3": {
|
|
19
39
|
"input": [
|
|
20
40
|
"text",
|