@doitian/dsh-provider-aliyun 0.1.0 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +126 -27
- package/lib/adapter.js +49 -6
- package/lib/catalog.js +110 -41
- package/lib/discovery.js +431 -0
- package/lib/index.js +121 -15
- package/lib/metadata.json +927 -0
- package/lib/models.js +168 -2
- package/package.json +8 -2
package/README.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Aliyun **DashScope (Bailian)** as a model provider for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness).
|
|
4
4
|
|
|
5
|
-
The plugin registers one provider route, `aliyun`, and **owns it outright** — endpoint, credential reference, protocol, and the
|
|
5
|
+
The plugin registers one provider route, `aliyun`, and **owns it outright** — endpoint, credential reference, protocol, and the model catalog. The catalog is *live*: the route asks the configured endpoint which models it serves (`GET {baseURL}/models`), so a model Aliyun publishes tomorrow is selectable without a package release. A shipped fallback list covers the time before the first listing arrives, and supplies the capacities a listing does not disclose.
|
|
6
6
|
|
|
7
7
|
```yaml
|
|
8
8
|
provider: aliyun
|
|
@@ -17,17 +17,19 @@ This package takes the other route. It registers its own route on the LLM seam a
|
|
|
17
17
|
|
|
18
18
|
| | Config-only route | This plugin |
|
|
19
19
|
|---|---|---|
|
|
20
|
-
| Where the model list
|
|
21
|
-
| Adding a
|
|
20
|
+
| Where the model list comes from | your `cordis.patch.yml` | the endpoint, re-read as it ages |
|
|
21
|
+
| Adding a model Aliyun starts serving | edit your profile | nothing — the next listing already has it |
|
|
22
|
+
| Capacities for a model nobody listed | you write them | shipped for the fallback ids, conservative defaults otherwise |
|
|
22
23
|
| Can a profile patch shadow it | — | no |
|
|
23
|
-
| Your profile config | the whole provider block | nothing |
|
|
24
|
+
| Your profile config | the whole provider block | nothing, or just a `baseURL` |
|
|
24
25
|
|
|
25
|
-
The trade-off is honest: the plugin depends on
|
|
26
|
+
The trade-off is honest: the plugin depends on internal shapes of the adapter and the discovery contract it uses (see [How it works](#how-it-works)), which a config-only route does not.
|
|
26
27
|
|
|
27
28
|
## Requirements
|
|
28
29
|
|
|
29
30
|
- DeepSeek Harness `0.2.0-rc.2` with the `@deepseek-ai/dsh-base` bundle, which mounts the LLM seam this plugin registers on.
|
|
30
31
|
- An Aliyun DashScope / Bailian API key.
|
|
32
|
+
- **An endpoint that answers `GET {baseURL}/models`.** The workspace hosts and the public compatible-mode endpoints do; a deployment that does not still works — the route just advertises the fallback catalog, and you can hand-list models with `models` in configuration instead.
|
|
31
33
|
- **No `aliyun` route configured through `llm-pi-ai`.** Two adapters cannot declare the same route: mounting this plugin while a profile still configures one fails with `configurable provider "aliyun" is already declared`. Remove that block from your profile patch first — that is the whole point of the plugin.
|
|
32
34
|
|
|
33
35
|
## Install
|
|
@@ -62,43 +64,100 @@ Nothing is needed to make the route exist, and no key is stored in this package
|
|
|
62
64
|
apiKeyEnv: ALIYUN_API_KEY # the default
|
|
63
65
|
```
|
|
64
66
|
|
|
65
|
-
|
|
67
|
+
Store it under `refs:` in `$DSH_HOME/.credentials.yaml`:
|
|
66
68
|
|
|
67
|
-
|
|
69
|
+
```yaml
|
|
70
|
+
refs:
|
|
71
|
+
ALIYUN_API_KEY: sk-…
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
The store watches that file and reloads it on change, so a key added while DSH is running takes effect on the next request — and on the next model listing, which is what turns the fallback catalog into the endpoint's own. Or export `ALIYUN_API_KEY` in the environment that launches DSH; that layer wins over the file and is reported read-only. Until the key resolves, selecting an Aliyun model fails immediately with `MISSING_CREDENTIAL`, before any network I/O.
|
|
75
|
+
|
|
76
|
+
> **The Models page has no field for this row.** A configuration surface renders a per-provider editor only for the `llm-deepseek` and `llm-pi-ai` namespaces; this plugin's own namespace shows the row, its missing-credential dot, and the hint that other fields live in `cordis.patch.yml`. That hint is right — the key goes in the file above, which is the same store the page writes to for the routes it does edit.
|
|
68
77
|
|
|
69
78
|
### Endpoints
|
|
70
79
|
|
|
71
80
|
| Region | Endpoint |
|
|
72
81
|
|---|---|
|
|
73
|
-
| Mainland China (default) | `https://dashscope.aliyuncs.com/compatible-mode/v1` |
|
|
82
|
+
| Mainland China, public (default) | `https://dashscope.aliyuncs.com/compatible-mode/v1` |
|
|
74
83
|
| International | `https://dashscope-intl.aliyuncs.com/compatible-mode/v1` |
|
|
84
|
+
| Bailian *workspace* | `https://llm-<workspace-id>.<region>.maas.aliyuncs.com/compatible-mode/v1` |
|
|
75
85
|
|
|
76
|
-
|
|
86
|
+
A workspace endpoint is the per-workspace host the console hands out; keys are often issued against one, and the workspace id is yours, so it can only come from configuration:
|
|
77
87
|
|
|
78
88
|
```yaml
|
|
79
89
|
- id: llm-aliyun
|
|
80
90
|
name: '@doitian/dsh-provider-aliyun'
|
|
81
91
|
config:
|
|
82
|
-
baseURL: https://
|
|
92
|
+
baseURL: https://llm-<workspace-id>.cn-beijing.maas.aliyuncs.com/compatible-mode/v1
|
|
83
93
|
```
|
|
84
94
|
|
|
95
|
+
Keep the endpoint and the key from the same product: a workspace key is not interchangeable with the public compatible-mode endpoints, and a *Token Plan* key belongs to `token-plan.<region>.maas.aliyuncs.com` instead.
|
|
96
|
+
|
|
85
97
|
## Models
|
|
86
98
|
|
|
87
|
-
`
|
|
99
|
+
The route advertises what the endpoint serves. On mount — and again whenever a listing goes stale or a credential is stored — the plugin calls `GET {baseURL}/models` and adopts the ids it finds, in the endpoint's own order. Within what the filter admits, membership follows that listing, removals included: a model Aliyun retires stops being selectable without a package release, and one it starts serving appears the same way.
|
|
100
|
+
|
|
101
|
+
Three things a listing does not do, and what happens instead:
|
|
102
|
+
|
|
103
|
+
| | Behaviour |
|
|
104
|
+
|---|---|
|
|
105
|
+
| Capacities | A listing discloses ids, not context windows. An id the fallback catalog names keeps its written facts; every other id gets the conservative defaults (`discovery.contextWindow`, `discovery.maxTokens`, `discovery.input`), because the harness trusts `contextWindow` when it compacts history and a generic number beats an invented precise one. |
|
|
106
|
+
| Failure | A failed read changes nothing: the last good listing stays in effect, and before the first one the fallback catalog does. A listing that names nothing counts as a failure — `{"data":[]}` says nothing about what an endpoint serves. |
|
|
107
|
+
| Latency | Only a model-list read waits for the endpoint, and only for `discovery.waitMs`. Every other read merely schedules a refresh, so a picker opened during a slow fetch shows what is already known. |
|
|
108
|
+
|
|
109
|
+
The shipped fallback list is three models: `qwen3.8-max`, `deepseek-v4.1-flash`, `glm-5.3`. It is *also* the metadata table, which is the whole trick — anything worth naming here gets your numbers.
|
|
110
|
+
|
|
111
|
+
**A listing is a catalogue, not a chat menu.** A workspace endpoint answers with everything the account can reach — on the workspace this plugin was developed against, 262 entries covering speech synthesis, speech recognition, image generation, embeddings, rerankers, realtime duplex models, and every legacy generation back to `qwen-7b-chat`. The listing discloses `id`, `object`, `created`, and `owned_by`, so *what an id is* cannot be read from the listing at all. Two things decide it instead, and the default uses both: name patterns from configuration, and the shipped metadata snapshot, which states outright whether a model reads and answers text. On that endpoint the three modes come out as:
|
|
88
112
|
|
|
89
|
-
|
|
113
|
+
| `discovery.filter` | Advertised | What admits an id |
|
|
114
|
+
|---|---|---|
|
|
115
|
+
| `chat` (default) | 79 of 262 | the include patterns, or the snapshot knowing it as a chat model |
|
|
116
|
+
| `patterns` | 44 of 262 | the include patterns alone — the tight picker |
|
|
117
|
+
| `all` | 175 of 262 | everything the listing names, minus `exclude` |
|
|
118
|
+
|
|
119
|
+
### Model facts
|
|
120
|
+
|
|
121
|
+
The listing carries no capacities, and the harness trusts `contextWindow` when it decides to compact history, so facts come from `lib/metadata.json`: a generated snapshot of [models.dev](https://models.dev)'s `alibaba-cn` and `alibaba` providers, committed here so nothing at run time depends on a third party. It covers 98 models with real context windows, output caps, accepted modalities, and whether a model can think — `qwen3.8-max` at 1M/131072, `qwen3-vl-plus` at 262144/32768 with image input, `kimi-k3`, `glm-5.3`, and so on.
|
|
122
|
+
|
|
123
|
+
```bash
|
|
124
|
+
npm run generate:metadata # refresh lib/metadata.json from models.dev
|
|
125
|
+
```
|
|
126
|
+
|
|
127
|
+
Precedence is fixed, most authoritative first: an entry in `models` is used exactly as written, then the snapshot, then the conservative defaults below. Membership stays the endpoint's business — the snapshot only admits a model the endpoint actually listed, and a model newer than the snapshot is still advertised when a pattern matches it, just with default numbers. Regenerating is a normal code change: `npm test` fails if the snapshot and the shipped catalog disagree about a model they both name.
|
|
128
|
+
|
|
129
|
+
**Verify capacities against your endpoint when you upgrade.** A wrong `contextWindow` is the failure that hurts.
|
|
90
130
|
|
|
91
|
-
|
|
131
|
+
### Discovery configuration
|
|
132
|
+
|
|
133
|
+
```yaml
|
|
134
|
+
- id: llm-aliyun
|
|
135
|
+
name: '@doitian/dsh-provider-aliyun'
|
|
136
|
+
config:
|
|
137
|
+
discovery:
|
|
138
|
+
enabled: true # false pins the route to `models`
|
|
139
|
+
filter: chat # chat | patterns | all, as the table above
|
|
140
|
+
include: # chat families to keep, as case-insensitive regular expressions
|
|
141
|
+
- '^qwen3\.8-(max|flash|omni-flash)$'
|
|
142
|
+
exclude: # dropped in every mode, even when a pattern matched
|
|
143
|
+
- '-realtime$'
|
|
144
|
+
ttlMs: 600000 # how long a listing stays fresh; 0 asks on every read
|
|
145
|
+
waitMs: 2000 # longest a model list waits for a cold listing
|
|
146
|
+
timeoutMs: 15000 # idle bound on one listing request
|
|
147
|
+
contextWindow: 131072 # capacity for an id neither `models` nor the snapshot names
|
|
148
|
+
maxTokens: 32768
|
|
149
|
+
input: [text]
|
|
150
|
+
```
|
|
92
151
|
|
|
93
|
-
|
|
152
|
+
Two rules keep the filter from hiding something you asked for. An id named in `models` is advertised whatever the patterns and the snapshot say — naming a model is a stronger statement than either — and a pattern that does not compile is ignored with a warning instead of failing the mount. `include` and `exclude` replace the shipped defaults, so copy the ones you want to keep from `lib/catalog.js`. `exclude` is also the way to trim what the snapshot admits, e.g. `- '^siliconflow/'` to drop one vendor's aliases of models that are already listed canonically.
|
|
94
153
|
|
|
95
|
-
|
|
154
|
+
Note what a discovered model gets for modalities: whatever the snapshot publishes, and `discovery.input` for everything else. A model neither source knows is advertised text-only, because neither a listing nor a guess says it accepts images; naming it in `models` settles the question.
|
|
96
155
|
|
|
97
|
-
|
|
156
|
+
After a failed attempt the endpoint is left alone for a minute, so an unreachable host cannot make every read slow; storing or rotating the key retries immediately.
|
|
98
157
|
|
|
99
|
-
###
|
|
158
|
+
### Pinning the catalog
|
|
100
159
|
|
|
101
|
-
`models`
|
|
160
|
+
`models` decides the fallback membership *and* the metadata, and a profile can replace it without forking:
|
|
102
161
|
|
|
103
162
|
```yaml
|
|
104
163
|
- id: llm-aliyun
|
|
@@ -114,7 +173,7 @@ Edit `lib/catalog.js` and publish. Consumers get the new list with `pnpm update`
|
|
|
114
173
|
thinkingLevelMap: { low: low, medium: medium, high: high }
|
|
115
174
|
```
|
|
116
175
|
|
|
117
|
-
|
|
176
|
+
With `discovery.enabled: false` this list is the entire route.
|
|
118
177
|
|
|
119
178
|
### Thinking
|
|
120
179
|
|
|
@@ -122,6 +181,8 @@ Aliyun turns thinking on with a boolean `enable_thinking` rather than a reasonin
|
|
|
122
181
|
|
|
123
182
|
`thinkingLevelMap` on each catalog entry declares which levels a model offers and the wire spelling of each. A level left out is not offered; `off` is offered unless it is named in the map.
|
|
124
183
|
|
|
184
|
+
A discovered model the fallback does not name is declared non-reasoning: a level map is a claim about a model, and claiming thinking it may not have would offer levels the endpoint can refuse. Name the model in `models` to give it levels.
|
|
185
|
+
|
|
125
186
|
## Uninstall
|
|
126
187
|
|
|
127
188
|
Remove it from `dsh.profile.bundles` (or delete the row on the Plugins page), then:
|
|
@@ -130,7 +191,7 @@ Remove it from `dsh.profile.bundles` (or delete the row on the Plugins page), th
|
|
|
130
191
|
dsh plugin --profile <name> remove @doitian/dsh-provider-aliyun
|
|
131
192
|
```
|
|
132
193
|
|
|
133
|
-
The credential in `$DSH_HOME/.credentials.yaml` is left untouched.
|
|
194
|
+
The credential in `$DSH_HOME/.credentials.yaml` is left untouched.
|
|
134
195
|
|
|
135
196
|
## How it works
|
|
136
197
|
|
|
@@ -148,11 +209,13 @@ and the patch inserts this package's own plugin entry:
|
|
|
148
209
|
name: '@doitian/dsh-provider-aliyun'
|
|
149
210
|
```
|
|
150
211
|
|
|
151
|
-
On mount, `lib/index.js` registers
|
|
212
|
+
On mount, `lib/index.js` registers three things on the LLM seam: the `aliyun` route with an adapter, a configurable-provider directory entry — that is the row on the Models page, not an API-key field (see [Configure the API key](#configure-the-api-key)) — and a model-discovery offer for the entry's settings namespace, so a configuration surface can interrogate this route's endpoint with a draft credential.
|
|
152
213
|
|
|
153
214
|
The adapter is `PiAiAdapter`, exported by `@deepseek-ai/dsh-llm-pi-ai`. It owns the hard part — harness history into pi-ai context, pi-ai events into harness stream chunks, image budgets, replay metadata, idle watchdogs — and reusing it is what keeps this package a catalog plus a few dozen lines instead of a second adapter implementation.
|
|
154
215
|
|
|
155
|
-
**
|
|
216
|
+
**How the advertised catalog moves.** `lib/discovery.js` reads the endpoint's listing and holds one `ids` array; `lib/models.js` merges it with the fallback entries; `lib/adapter.js` swaps the profile *map* the adapter memoizes its snapshot by identity. An operation already in flight keeps the snapshot it started with, and the next read sees the new catalog — which is also why a refresh never interrupts a stream. The harness's session catalog calls `listModels()` per provider on every read, so the picker shows whatever the route advertises at that moment. Only `listModels` waits for a cold listing, and that wait lives in a two-line subclass rather than a reimplementation of the adapter.
|
|
217
|
+
|
|
218
|
+
**The one thing to know if you maintain this.** `PiAiAdapter` is driven by a route's *resolved profile*, a shape its package does not export as a type; `lib/adapter.js` reproduces it and documents every field it must carry. Model metadata is not read from that profile — it comes from the pi-ai model descriptors built out of the catalog, which is why the catalog can live here at all. A DSH upgrade that starts reading a new profile field breaks this plugin. `test/adapter.test.mjs` drives the real published adapter against the factory, so that break shows up as a failing test rather than as a broken route in someone's profile. The discovery contract is a second such seam: `LlmModelDiscoveryRequest`, `registerModelDiscovery`, and the `attributionHeaders()` requirement on every provider HTTP request.
|
|
156
219
|
|
|
157
220
|
The engine is pinned to the harness's own pi-ai: the dependency is `^0.87.1`, which resolves to exactly the `0.87.1` the harness installs, so both share one copy.
|
|
158
221
|
|
|
@@ -160,13 +223,26 @@ The engine is pinned to the harness's own pi-ai: the dependency is `^0.87.1`, wh
|
|
|
160
223
|
|
|
161
224
|
```bash
|
|
162
225
|
npm install
|
|
163
|
-
npm run validate # manifest, patch wiring, catalog
|
|
164
|
-
npm test # drives the real PiAiAdapter
|
|
226
|
+
npm run validate # manifest, patch wiring, fallback catalog, discovery defaults
|
|
227
|
+
npm test # drives the real PiAiAdapter, the merge, and the listing state machine
|
|
228
|
+
```
|
|
229
|
+
|
|
230
|
+
`scripts/validate.mjs` catches what would otherwise fail silently: a tarball that drops the patch, a patch that names the wrong package, a duplicate model id, a non-integer capacity, a level map pi-ai would read as offering nothing, a discovery default that would produce a model the adapter cannot dispatch.
|
|
231
|
+
|
|
232
|
+
`npm test` needs the peer packages installed — that is the point, since it exercises the real adapter rather than a stub. A live request is out of scope: that needs a real key and endpoint. Discovery is tested with an injected `fetch`, so what is covered is every decision *around* the request — the shapes it reads, what a failure leaves in place, and how long a read may wait.
|
|
233
|
+
|
|
234
|
+
### Testing a build against the desktop app
|
|
235
|
+
|
|
236
|
+
```bash
|
|
237
|
+
npm run link:desktop # point the desktop profile at this checkout, then verify the copy
|
|
238
|
+
npm run probe # what the picker will show, read from the live endpoint
|
|
165
239
|
```
|
|
166
240
|
|
|
167
|
-
|
|
241
|
+
Then **restart DSH**: plugin modules are loaded once at startup, and the loader does not watch them.
|
|
242
|
+
|
|
243
|
+
`link:desktop` exists because a local dependency is not a live view of the worktree. pnpm hard-links a `file:` package into the profile, so an edit that *replaces* a file — `git checkout`, or any editor that saves atomically — leaves the installed copy on the previous bytes, and a plain `install`, `--force`, or `update` will not re-link it ("Already up to date", even when the directory is deleted). Re-adding the dependency is what forces a fresh resolution, and the script verifies the result byte for byte instead of assuming it worked.
|
|
168
244
|
|
|
169
|
-
|
|
245
|
+
Two things to expect while testing this way. The Plugins page may rewrite the dependency back to a registry range — that is a normal package install, and `link:desktop` switches it back. And a *published* release is the durable alternative: the app then installs it like any other plugin, and nothing local is involved.
|
|
170
246
|
|
|
171
247
|
## Publishing
|
|
172
248
|
|
|
@@ -185,6 +261,25 @@ npm trust github @doitian/dsh-provider-aliyun \
|
|
|
185
261
|
|
|
186
262
|
`--allow-publish` is required: it grants the `CREATE_PACKAGE` permission, and the command refuses without it or `--allow-stage-publish`. It infers `owner/repo` from `package.json` when `--repo` is omitted, warns when the two disagree, requires 2FA, and prompts for an OTP. Add `--dry-run` first to see exactly what it would create without committing it.
|
|
187
263
|
|
|
264
|
+
**A bypass-2FA token cannot do this step.** npm is retiring tokens that bypass 2FA, and the
|
|
265
|
+
restriction is asymmetric: such a token still *publishes*, but the registry refuses it for
|
|
266
|
+
publisher management with
|
|
267
|
+
|
|
268
|
+
```
|
|
269
|
+
Granular access tokens that bypass two-factor authentication may not perform this action.
|
|
270
|
+
```
|
|
271
|
+
|
|
272
|
+
So the bootstrap publish can come from a bypass token, while attaching the publisher needs an
|
|
273
|
+
interactive `npm login` session or the website form. Reading the config back is refused for
|
|
274
|
+
the same reason, so the release itself is the practical check.
|
|
275
|
+
|
|
276
|
+
**Account two-factor authentication does not reach the workflow.** The publish job authenticates
|
|
277
|
+
with an OIDC token, so no one-time password is involved and enabling 2FA on the npm account
|
|
278
|
+
cannot break a release — v0.1.1 was published exactly this way. What 2FA does gate is the
|
|
279
|
+
one-time publisher setup above: an account with 2FA can run `npm trust`, which is the step a
|
|
280
|
+
2FA-bypassing token is refused. Publishing by hand from a terminal is the only flow that now
|
|
281
|
+
asks for an OTP.
|
|
282
|
+
|
|
188
283
|
`--environment` is deliberately omitted, matching the publish job, which declares no `environment:`.
|
|
189
284
|
|
|
190
285
|
If the registry reports the package as missing, a trusted publisher cannot be attached to a name that does not exist yet: publish `0.1.0` once by hand, then run the command above and let the workflow own every release after that.
|
|
@@ -192,11 +287,15 @@ If the registry reports the package as missing, a trusted publisher cannot be at
|
|
|
192
287
|
Then release:
|
|
193
288
|
|
|
194
289
|
```bash
|
|
195
|
-
npm version patch
|
|
290
|
+
npm version patch # or skip it when the version is already committed in package.json
|
|
196
291
|
git push --follow-tags
|
|
197
292
|
gh release create v0.1.1 --generate-notes
|
|
198
293
|
```
|
|
199
294
|
|
|
295
|
+
`npm version` writes the version, commits it, and tags it in one step. When the version you mean
|
|
296
|
+
to release is already committed by hand — a feature release that bumped `minor` itself — tag that
|
|
297
|
+
commit instead: `git tag v0.2.0 && git push --follow-tags`, then create the release for it.
|
|
298
|
+
|
|
200
299
|
The publish job requires the release tag to match `package.json` (`v0.1.1` ↔ `0.1.1`) and re-runs validation and tests before uploading.
|
|
201
300
|
|
|
202
301
|
Two things that will bite when editing the workflow:
|
package/lib/adapter.js
CHANGED
|
@@ -24,6 +24,12 @@
|
|
|
24
24
|
* from the profile. It is read from the pi-ai model descriptors inside
|
|
25
25
|
* `piProvider`, which is why the catalog can live in this package.
|
|
26
26
|
*
|
|
27
|
+
* The catalog it carries is not fixed at mount. Discovery ({@link ./discovery.js})
|
|
28
|
+
* swaps the advertised entries as the endpoint's listing moves, and the swap is
|
|
29
|
+
* a new profile *map* so the adapter's identity-based memoization notices it:
|
|
30
|
+
* every operation already captured keeps the snapshot it started with, and the
|
|
31
|
+
* next read sees the new catalog.
|
|
32
|
+
*
|
|
27
33
|
* The risk this carries: a DSH upgrade that starts reading a new profile field
|
|
28
34
|
* breaks this plugin. `test/adapter.test.mjs` drives the real `PiAiAdapter`
|
|
29
35
|
* against this factory so that such a break shows up as a failing test rather
|
|
@@ -111,23 +117,60 @@ export function buildAliyunProfile(options) {
|
|
|
111
117
|
* @param {string} options.displayName - label for selectors.
|
|
112
118
|
* @param {string} options.apiKeyEnv - credential reference to resolve per request.
|
|
113
119
|
* @param {string} options.baseURL - endpoint for every model on the route.
|
|
114
|
-
* @param {readonly import('./catalog.js').AliyunModel[]} options.models - the
|
|
120
|
+
* @param {readonly import('./catalog.js').AliyunModel[]} options.models - the catalog to advertise first.
|
|
115
121
|
* @param {string} [options.reasoning] - route default thinking level.
|
|
116
122
|
* @param {Record<string, string>} [options.headers] - extra request headers.
|
|
117
123
|
* @param {(provider: string, profile: object) => Promise<string | undefined>} options.resolveApiKey - credential resolver.
|
|
118
124
|
* @param {() => object | undefined} [options.resolveAttachments] - durable attachment service accessor.
|
|
119
125
|
* @param {(attachments: object, ref: object) => unknown} [options.resolveImageAccess] - image access resolver.
|
|
120
126
|
* @param {(event: object) => void} [options.onReplayDegrade] - replay degradation reporter.
|
|
121
|
-
* @
|
|
127
|
+
* @param {() => void} [options.onRead] - called before every operation reads the profile map.
|
|
128
|
+
* @param {() => Promise<void>} [options.settle] - awaited before a model list is answered.
|
|
129
|
+
* @returns {{ adapter: PiAiAdapter, setCatalog: (models: readonly import('./catalog.js').AliyunModel[]) => void }} the adapter to register, and the way to move its catalog.
|
|
122
130
|
*/
|
|
123
131
|
export function createAliyunAdapter(options) {
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
|
|
132
|
+
/** The profile map the adapter memoizes by identity; a catalog swap replaces it. */
|
|
133
|
+
let profiles = profilesOver(options.models)
|
|
134
|
+
|
|
135
|
+
/** One route's profile for one catalog, as a fresh map. */
|
|
136
|
+
function profilesOver(models) {
|
|
137
|
+
return new Map([[ROUTE, buildAliyunProfile({ ...options, models })]])
|
|
138
|
+
}
|
|
139
|
+
|
|
140
|
+
/**
|
|
141
|
+
* `PiAiAdapter` answers `listModels` from its memoized snapshot, and that
|
|
142
|
+
* answer is what the picker reads. Waiting here — and nowhere else — is what
|
|
143
|
+
* lets a picker opened on a cold route show the endpoint's own list instead of
|
|
144
|
+
* the fallback one, while every other read stays off the network.
|
|
145
|
+
*/
|
|
146
|
+
class AliyunAdapter extends PiAiAdapter {
|
|
147
|
+
async listModels(provider) {
|
|
148
|
+
await options.settle?.()
|
|
149
|
+
return super.listModels(provider)
|
|
150
|
+
}
|
|
151
|
+
}
|
|
152
|
+
|
|
153
|
+
const adapter = new AliyunAdapter({
|
|
154
|
+
// Every operation asks for the profile map, so a stale listing is noticed
|
|
155
|
+
// and refreshed on the way past, without any of those operations waiting.
|
|
156
|
+
profiles: () => {
|
|
157
|
+
options.onRead?.()
|
|
158
|
+
return profiles
|
|
159
|
+
},
|
|
128
160
|
resolveApiKey: options.resolveApiKey,
|
|
129
161
|
resolveAttachments: options.resolveAttachments,
|
|
130
162
|
resolveImageAccess: options.resolveImageAccess,
|
|
131
163
|
onReplayDegrade: options.onReplayDegrade,
|
|
132
164
|
})
|
|
165
|
+
|
|
166
|
+
return {
|
|
167
|
+
adapter,
|
|
168
|
+
/**
|
|
169
|
+
* Advertise `models` from the next read on.
|
|
170
|
+
* @param {readonly import('./catalog.js').AliyunModel[]} models - the entries the route now serves.
|
|
171
|
+
*/
|
|
172
|
+
setCatalog(models) {
|
|
173
|
+
profiles = profilesOver(models)
|
|
174
|
+
},
|
|
175
|
+
}
|
|
133
176
|
}
|
package/lib/catalog.js
CHANGED
|
@@ -1,18 +1,26 @@
|
|
|
1
1
|
/**
|
|
2
|
-
* The Aliyun DashScope (Bailian)
|
|
2
|
+
* The Aliyun DashScope (Bailian) catalog this plugin ships.
|
|
3
3
|
*
|
|
4
|
-
* This file
|
|
5
|
-
*
|
|
6
|
-
*
|
|
4
|
+
* This file does two jobs, and the split is the point:
|
|
5
|
+
*
|
|
6
|
+
* - **Fallback membership.** {@link ./discovery.js} asks the configured endpoint
|
|
7
|
+
* which models it serves. Until that answer arrives — no credential stored
|
|
8
|
+
* yet, an endpoint that serves no listing, a network that cannot reach it —
|
|
9
|
+
* these entries are the whole route, so the picker is never empty.
|
|
10
|
+
* - **Metadata.** A listing discloses ids and nothing else, and the harness
|
|
11
|
+
* trusts `contextWindow` when it decides to compact history. An id named here
|
|
12
|
+
* carries the real capacities; an id discovered and *not* named here gets the
|
|
13
|
+
* conservative defaults from configuration, because an obviously generic
|
|
14
|
+
* capacity beats an invented precise one.
|
|
7
15
|
*
|
|
8
16
|
* It is plain data on purpose. {@link ./models.js} turns each entry into the
|
|
9
17
|
* pi-ai model descriptor the adapter dispatches with, filling in the route,
|
|
10
18
|
* endpoint, protocol, and compatibility switches from configuration — so an
|
|
11
19
|
* entry here states only what is true of the *model*, never of the deployment.
|
|
12
20
|
*
|
|
13
|
-
*
|
|
14
|
-
*
|
|
15
|
-
* failure that hurts
|
|
21
|
+
* These are the models this route is used with day to day. Verify the capacities
|
|
22
|
+
* against your own endpoint when you upgrade; a wrong `contextWindow` is the
|
|
23
|
+
* failure that hurts.
|
|
16
24
|
*/
|
|
17
25
|
|
|
18
26
|
/** Route name every request selects with `GenerateOptions.provider`. */
|
|
@@ -27,6 +35,15 @@ export const DEFAULT_BASE_URL = 'https://dashscope.aliyuncs.com/compatible-mode/
|
|
|
27
35
|
/** International Model Studio endpoint. */
|
|
28
36
|
export const INTL_BASE_URL = 'https://dashscope-intl.aliyuncs.com/compatible-mode/v1'
|
|
29
37
|
|
|
38
|
+
/**
|
|
39
|
+
* The shape of a Bailian *workspace* endpoint — the per-workspace host the
|
|
40
|
+
* console hands out, and the spelling most keys are issued against.
|
|
41
|
+
*
|
|
42
|
+
* An example to recognize, not a default this package can ship: the workspace id
|
|
43
|
+
* is the deployment's own.
|
|
44
|
+
*/
|
|
45
|
+
export const WORKSPACE_BASE_URL_EXAMPLE = 'https://llm-<workspace-id>.cn-beijing.maas.aliyuncs.com/compatible-mode/v1'
|
|
46
|
+
|
|
30
47
|
/** Credential reference the route resolves through the harness credential seam. */
|
|
31
48
|
export const DEFAULT_API_KEY_ENV = 'ALIYUN_API_KEY'
|
|
32
49
|
|
|
@@ -49,6 +66,81 @@ export const QWEN_COMPAT = Object.freeze({
|
|
|
49
66
|
supportsStrictMode: true,
|
|
50
67
|
})
|
|
51
68
|
|
|
69
|
+
/**
|
|
70
|
+
* What a discovered model gets when no entry below names it.
|
|
71
|
+
*
|
|
72
|
+
* Context is conservative on purpose: compacting early costs tokens, compacting
|
|
73
|
+
* late costs the request. Output sits comfortably under every cap this endpoint
|
|
74
|
+
* publishes, so a request is never refused for asking for more than a model
|
|
75
|
+
* allows.
|
|
76
|
+
*/
|
|
77
|
+
export const DISCOVERY_DEFAULTS = Object.freeze({
|
|
78
|
+
contextWindow: 131_072,
|
|
79
|
+
maxTokens: 32_768,
|
|
80
|
+
input: Object.freeze(['text']),
|
|
81
|
+
})
|
|
82
|
+
|
|
83
|
+
/** How long a listing stays fresh before a read asks the endpoint again. */
|
|
84
|
+
export const DISCOVERY_TTL_MS = 600_000
|
|
85
|
+
|
|
86
|
+
/** Longest a cold or stale model-list read waits before answering from what it has. */
|
|
87
|
+
export const DISCOVERY_WAIT_MS = 2_000
|
|
88
|
+
|
|
89
|
+
/**
|
|
90
|
+
* The chat families this route advertises by default, as case-insensitive
|
|
91
|
+
* regular expressions.
|
|
92
|
+
*
|
|
93
|
+
* A DashScope listing discloses `id`, `object`, `created`, and `owned_by` — no
|
|
94
|
+
* modality, no task type — so what an id *is* can only be read from its name.
|
|
95
|
+
* That matters because a workspace endpoint lists everything the account can
|
|
96
|
+
* reach: speech synthesis, speech recognition, image generation, embeddings,
|
|
97
|
+
* rerankers, realtime duplex models, and every legacy generation back to
|
|
98
|
+
* `qwen-7b-chat`. Advertising all of it makes the picker unusable and offers
|
|
99
|
+
* models that cannot answer a chat request at all.
|
|
100
|
+
*
|
|
101
|
+
* The patterns are anchored, which is also what drops vendor-prefixed aliases
|
|
102
|
+
* (`ZHIPU/GLM-5.3`, `vanchin/deepseek-v4.1-flash`) whose canonical id is already
|
|
103
|
+
* matched — the two prefixed exceptions are families with no canonical spelling
|
|
104
|
+
* on this endpoint.
|
|
105
|
+
*
|
|
106
|
+
* This is a *default*, not a rule. `discovery.filter: all` advertises the whole
|
|
107
|
+
* listing, `discovery.include` adds families, `discovery.exclude` drops ids, and
|
|
108
|
+
* an id named in `models` is advertised whatever these say.
|
|
109
|
+
*/
|
|
110
|
+
export const CHAT_MODEL_PATTERNS = Object.freeze([
|
|
111
|
+
'^qwen3\\.8-(max|flash|omni-flash)$',
|
|
112
|
+
'^qwen3\\.7-(max|plus|flash)$',
|
|
113
|
+
'^qwen3\\.6-(plus|flash)$',
|
|
114
|
+
'^qwen3\\.5-(plus|flash|omni-plus|omni-flash)$',
|
|
115
|
+
// Open-weight releases carry their size in the id, which is how they are told
|
|
116
|
+
// apart from the hosted aliases of the same generation.
|
|
117
|
+
'^qwen3\\.[5-8]-\\d+b(-a\\d+b)?$',
|
|
118
|
+
'^qwen3-(max|omni-flash)$',
|
|
119
|
+
'^qwen3-coder-(plus|flash)$',
|
|
120
|
+
'^deepseek-(v4\\.1-flash|v4-pro|v4-flash|v3(\\.[12])?|r1)$',
|
|
121
|
+
'^glm-(4\\.7|5(\\.\\d+)?(-prime)?)$',
|
|
122
|
+
'^kimi-k[23](\\.\\d+)?(-code|-thinking)?$',
|
|
123
|
+
'^MiniMax-M[23](\\.\\d+)?$',
|
|
124
|
+
'^stepfun/step-[0-9.]+(-flash)?$',
|
|
125
|
+
'^xiaomi/mimo-[a-z0-9.]+$',
|
|
126
|
+
])
|
|
127
|
+
|
|
128
|
+
/**
|
|
129
|
+
* Names to drop even when an include pattern matched, because these words are
|
|
130
|
+
* how this endpoint spells "not a plain chat completion model": a snapshot of a
|
|
131
|
+
* model this list already carries, an unreleased preview, or a streaming-first
|
|
132
|
+
* variant.
|
|
133
|
+
*/
|
|
134
|
+
export const NON_CHAT_MODEL_PATTERNS = Object.freeze([
|
|
135
|
+
'-realtime$',
|
|
136
|
+
'-preview$',
|
|
137
|
+
'-\\d{4}-\\d{2}-\\d{2}$',
|
|
138
|
+
'embedding',
|
|
139
|
+
'rerank',
|
|
140
|
+
'-highspeed$',
|
|
141
|
+
'-flashx$',
|
|
142
|
+
])
|
|
143
|
+
|
|
52
144
|
/**
|
|
53
145
|
* One shipped model.
|
|
54
146
|
*
|
|
@@ -61,11 +153,13 @@ export const QWEN_COMPAT = Object.freeze({
|
|
|
61
153
|
* @property {boolean} reasoning - whether the model can think.
|
|
62
154
|
* @property {Record<string, string>} [thinkingLevelMap] - selectable thinking
|
|
63
155
|
* levels and the wire spelling each one sends. A level left out is not
|
|
64
|
-
* offered; `off` is offered unless it is named here.
|
|
156
|
+
* offered; `off` is offered unless it is named here. This endpoint takes a
|
|
157
|
+
* boolean rather than an effort, so what the map really decides is the *set*
|
|
158
|
+
* of levels a selector offers.
|
|
65
159
|
*/
|
|
66
160
|
|
|
67
161
|
/** @type {readonly AliyunModel[]} */
|
|
68
|
-
export const
|
|
162
|
+
export const FALLBACK_MODELS = Object.freeze([
|
|
69
163
|
{
|
|
70
164
|
id: 'qwen3.8-max',
|
|
71
165
|
name: 'Qwen3.8 Max',
|
|
@@ -76,48 +170,23 @@ export const ALIYUN_MODELS = Object.freeze([
|
|
|
76
170
|
thinkingLevelMap: { low: 'low', medium: 'medium', high: 'high' },
|
|
77
171
|
},
|
|
78
172
|
{
|
|
79
|
-
|
|
80
|
-
|
|
173
|
+
// Aliyun publishes 384K output for this model, and the Aliyun catalog inside
|
|
174
|
+
// the pinned pi-ai carries the same figure.
|
|
175
|
+
id: 'deepseek-v4.1-flash',
|
|
176
|
+
name: 'DeepSeek V4.1 Flash',
|
|
81
177
|
contextWindow: 1_000_000,
|
|
82
|
-
maxTokens:
|
|
178
|
+
maxTokens: 384_000,
|
|
83
179
|
input: ['text', 'image'],
|
|
84
180
|
reasoning: true,
|
|
85
181
|
thinkingLevelMap: { low: 'low', medium: 'medium', high: 'high' },
|
|
86
182
|
},
|
|
87
183
|
{
|
|
88
|
-
id: '
|
|
89
|
-
name: '
|
|
184
|
+
id: 'glm-5.3',
|
|
185
|
+
name: 'GLM-5.3',
|
|
90
186
|
contextWindow: 1_000_000,
|
|
91
187
|
maxTokens: 131_072,
|
|
92
188
|
input: ['text'],
|
|
93
189
|
reasoning: true,
|
|
94
190
|
thinkingLevelMap: { low: 'low', medium: 'medium', high: 'high' },
|
|
95
191
|
},
|
|
96
|
-
{
|
|
97
|
-
id: 'qwen3.7-plus',
|
|
98
|
-
name: 'Qwen3.7 Plus',
|
|
99
|
-
contextWindow: 1_000_000,
|
|
100
|
-
maxTokens: 65_536,
|
|
101
|
-
input: ['text', 'image'],
|
|
102
|
-
reasoning: true,
|
|
103
|
-
thinkingLevelMap: { low: 'low', medium: 'medium', high: 'high' },
|
|
104
|
-
},
|
|
105
|
-
{
|
|
106
|
-
id: 'qwen3.6-plus',
|
|
107
|
-
name: 'Qwen3.6 Plus',
|
|
108
|
-
contextWindow: 1_000_000,
|
|
109
|
-
maxTokens: 65_536,
|
|
110
|
-
input: ['text', 'image'],
|
|
111
|
-
reasoning: true,
|
|
112
|
-
thinkingLevelMap: { low: 'low', medium: 'medium', high: 'high' },
|
|
113
|
-
},
|
|
114
|
-
{
|
|
115
|
-
id: 'qwen3.6-flash',
|
|
116
|
-
name: 'Qwen3.6 Flash',
|
|
117
|
-
contextWindow: 1_000_000,
|
|
118
|
-
maxTokens: 65_536,
|
|
119
|
-
input: ['text', 'image'],
|
|
120
|
-
reasoning: true,
|
|
121
|
-
thinkingLevelMap: { low: 'low', medium: 'medium', high: 'high' },
|
|
122
|
-
},
|
|
123
192
|
])
|