@doitian/dsh-provider-aliyun 0.1.0 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -2,7 +2,7 @@
2
2
 
3
3
  Aliyun **DashScope (Bailian)** as a model provider for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness).
4
4
 
5
- The plugin registers one provider route, `aliyun`, and **owns it outright** — endpoint, credential reference, protocol, and the Qwen model catalog. The catalog is a file in this package, so the model list moves forward with `pnpm update` and never has to be written into a profile.
5
+ The plugin registers one provider route, `aliyun`, and **owns it outright** — endpoint, credential reference, protocol, and the model catalog. The catalog is *live*: the route asks the configured endpoint which models it serves (`GET {baseURL}/models`), so a model Aliyun publishes tomorrow is selectable without a package release. A shipped fallback list covers the time before the first listing arrives, and supplies the capacities a listing does not disclose.
6
6
 
7
7
  ```yaml
8
8
  provider: aliyun
@@ -17,17 +17,19 @@ This package takes the other route. It registers its own route on the LLM seam a
17
17
 
18
18
  | | Config-only route | This plugin |
19
19
  |---|---|---|
20
- | Where the model list lives | your `cordis.patch.yml` | `lib/catalog.js` in this package |
21
- | Adding a new Qwen model | edit your profile | `pnpm update` |
20
+ | Where the model list comes from | your `cordis.patch.yml` | the endpoint, re-read as it ages |
21
+ | Adding a model Aliyun starts serving | edit your profile | nothing — the next listing already has it |
22
+ | Capacities for a model nobody listed | you write them | shipped for the fallback ids, conservative defaults otherwise |
22
23
  | Can a profile patch shadow it | — | no |
23
- | Your profile config | the whole provider block | nothing |
24
+ | Your profile config | the whole provider block | nothing, or just a `baseURL` |
24
25
 
25
- The trade-off is honest: the plugin depends on an internal shape of the adapter it reuses (see [How it works](#how-it-works)), which a config-only route does not.
26
+ The trade-off is honest: the plugin depends on internal shapes of the adapter and the discovery contract it uses (see [How it works](#how-it-works)), which a config-only route does not.
26
27
 
27
28
  ## Requirements
28
29
 
29
30
  - DeepSeek Harness `0.2.0-rc.2` with the `@deepseek-ai/dsh-base` bundle, which mounts the LLM seam this plugin registers on.
30
31
  - An Aliyun DashScope / Bailian API key.
32
+ - **An endpoint that answers `GET {baseURL}/models`.** The workspace hosts and the public compatible-mode endpoints do; a deployment that does not still works — the route just advertises the fallback catalog, and you can hand-list models with `models` in configuration instead.
31
33
  - **No `aliyun` route configured through `llm-pi-ai`.** Two adapters cannot declare the same route: mounting this plugin while a profile still configures one fails with `configurable provider "aliyun" is already declared`. Remove that block from your profile patch first — that is the whole point of the plugin.
32
34
 
33
35
  ## Install
@@ -62,43 +64,100 @@ Nothing is needed to make the route exist, and no key is stored in this package
62
64
  apiKeyEnv: ALIYUN_API_KEY # the default
63
65
  ```
64
66
 
65
- **Settings → Models → Aliyun DashScope** → paste the key into the **API key** field. It is written write-only to `$DSH_HOME/.credentials.yaml` under `ALIYUN_API_KEY`, and never read back.
67
+ Store it under `refs:` in `$DSH_HOME/.credentials.yaml`:
66
68
 
67
- Or export `ALIYUN_API_KEY` in the environment that launches DSH. Until the key resolves, selecting an Aliyun model fails immediately with `MISSING_CREDENTIAL`, before any network I/O.
69
+ ```yaml
70
+ refs:
71
+ ALIYUN_API_KEY: sk-…
72
+ ```
73
+
74
+ The store watches that file and reloads it on change, so a key added while DSH is running takes effect on the next request — and on the next model listing, which is what turns the fallback catalog into the endpoint's own. Or export `ALIYUN_API_KEY` in the environment that launches DSH; that layer wins over the file and is reported read-only. Until the key resolves, selecting an Aliyun model fails immediately with `MISSING_CREDENTIAL`, before any network I/O.
75
+
76
+ > **The Models page has no field for this row.** A configuration surface renders a per-provider editor only for the `llm-deepseek` and `llm-pi-ai` namespaces; this plugin's own namespace shows the row, its missing-credential dot, and the hint that other fields live in `cordis.patch.yml`. That hint is right — the key goes in the file above, which is the same store the page writes to for the routes it does edit.
68
77
 
69
78
  ### Endpoints
70
79
 
71
80
  | Region | Endpoint |
72
81
  |---|---|
73
- | Mainland China (default) | `https://dashscope.aliyuncs.com/compatible-mode/v1` |
82
+ | Mainland China, public (default) | `https://dashscope.aliyuncs.com/compatible-mode/v1` |
74
83
  | International | `https://dashscope-intl.aliyuncs.com/compatible-mode/v1` |
84
+ | Bailian *workspace* | `https://llm-<workspace-id>.<region>.maas.aliyuncs.com/compatible-mode/v1` |
75
85
 
76
- Set it in the Models page, or by adding `baseURL` to this entry's `config`:
86
+ A workspace endpoint is the per-workspace host the console hands out; keys are often issued against one, and the workspace id is yours, so it can only come from configuration:
77
87
 
78
88
  ```yaml
79
89
  - id: llm-aliyun
80
90
  name: '@doitian/dsh-provider-aliyun'
81
91
  config:
82
- baseURL: https://dashscope-intl.aliyuncs.com/compatible-mode/v1
92
+ baseURL: https://llm-<workspace-id>.cn-beijing.maas.aliyuncs.com/compatible-mode/v1
83
93
  ```
84
94
 
95
+ Keep the endpoint and the key from the same product: a workspace key is not interchangeable with the public compatible-mode endpoints, and a *Token Plan* key belongs to `token-plan.<region>.maas.aliyuncs.com` instead.
96
+
85
97
  ## Models
86
98
 
87
- `lib/catalog.js` is the catalog. It ships six Qwen models:
99
+ The route advertises what the endpoint serves. On mount — and again whenever a listing goes stale or a credential is stored — the plugin calls `GET {baseURL}/models` and adopts the ids it finds, in the endpoint's own order. Within what the filter admits, membership follows that listing, removals included: a model Aliyun retires stops being selectable without a package release, and one it starts serving appears the same way.
100
+
101
+ Three things a listing does not do, and what happens instead:
102
+
103
+ | | Behaviour |
104
+ |---|---|
105
+ | Capacities | A listing discloses ids, not context windows. An id the fallback catalog names keeps its written facts; every other id gets the conservative defaults (`discovery.contextWindow`, `discovery.maxTokens`, `discovery.input`), because the harness trusts `contextWindow` when it compacts history and a generic number beats an invented precise one. |
106
+ | Failure | A failed read changes nothing: the last good listing stays in effect, and before the first one the fallback catalog does. A listing that names nothing counts as a failure — `{"data":[]}` says nothing about what an endpoint serves. |
107
+ | Latency | Only a model-list read waits for the endpoint, and only for `discovery.waitMs`. Every other read merely schedules a refresh, so a picker opened during a slow fetch shows what is already known. |
108
+
109
+ The shipped fallback list is three models: `qwen3.8-max`, `deepseek-v4.1-flash`, `glm-5.3`. It is *also* the metadata table, which is the whole trick — anything worth naming here gets your numbers.
110
+
111
+ **A listing is a catalogue, not a chat menu.** A workspace endpoint answers with everything the account can reach — on the workspace this plugin was developed against, 262 entries covering speech synthesis, speech recognition, image generation, embeddings, rerankers, realtime duplex models, and every legacy generation back to `qwen-7b-chat`. The listing discloses `id`, `object`, `created`, and `owned_by`, so *what an id is* cannot be read from the listing at all. Two things decide it instead, and the default uses both: name patterns from configuration, and the shipped metadata snapshot, which states outright whether a model reads and answers text. On that endpoint the three modes come out as:
88
112
 
89
- `qwen3.8-max`, `qwen3.8-flash`, `qwen3.7-max`, `qwen3.7-plus`, `qwen3.6-plus`, `qwen3.6-flash`
113
+ | `discovery.filter` | Advertised | What admits an id |
114
+ |---|---|---|
115
+ | `chat` (default) | 79 of 262 | the include patterns, or the snapshot knowing it as a chat model |
116
+ | `patterns` | 44 of 262 | the include patterns alone — the tight picker |
117
+ | `all` | 175 of 262 | everything the listing names, minus `exclude` |
118
+
119
+ ### Model facts
120
+
121
+ The listing carries no capacities, and the harness trusts `contextWindow` when it decides to compact history, so facts come from `lib/metadata.json`: a generated snapshot of [models.dev](https://models.dev)'s `alibaba-cn` and `alibaba` providers, committed here so nothing at run time depends on a third party. It covers 98 models with real context windows, output caps, accepted modalities, and whether a model can think — `qwen3.8-max` at 1M/131072, `qwen3-vl-plus` at 262144/32768 with image input, `kimi-k3`, `glm-5.3`, and so on.
122
+
123
+ ```bash
124
+ npm run generate:metadata # refresh lib/metadata.json from models.dev
125
+ ```
126
+
127
+ Precedence is fixed, most authoritative first: an entry in `models` is used exactly as written, then the snapshot, then the conservative defaults below. Membership stays the endpoint's business — the snapshot only admits a model the endpoint actually listed, and a model newer than the snapshot is still advertised when a pattern matches it, just with default numbers. Regenerating is a normal code change: `npm test` fails if the snapshot and the shipped catalog disagree about a model they both name.
128
+
129
+ **Verify capacities against your endpoint when you upgrade.** A wrong `contextWindow` is the failure that hurts.
90
130
 
91
- Capacities and thinking levels mirror the facts Aliyun publishes for its Qwen lineup. A catalog entry states only what is true of the *model* — id, display name, context window, output cap, accepted modalities, and thinking levels. The route, endpoint, protocol, and compatibility switches are filled in from configuration by `lib/models.js`, so updating the catalog never means touching pi-ai plumbing.
131
+ ### Discovery configuration
132
+
133
+ ```yaml
134
+ - id: llm-aliyun
135
+ name: '@doitian/dsh-provider-aliyun'
136
+ config:
137
+ discovery:
138
+ enabled: true # false pins the route to `models`
139
+ filter: chat # chat | patterns | all, as the table above
140
+ include: # chat families to keep, as case-insensitive regular expressions
141
+ - '^qwen3\.8-(max|flash|omni-flash)$'
142
+ exclude: # dropped in every mode, even when a pattern matched
143
+ - '-realtime$'
144
+ ttlMs: 600000 # how long a listing stays fresh; 0 asks on every read
145
+ waitMs: 2000 # longest a model list waits for a cold listing
146
+ timeoutMs: 15000 # idle bound on one listing request
147
+ contextWindow: 131072 # capacity for an id neither `models` nor the snapshot names
148
+ maxTokens: 32768
149
+ input: [text]
150
+ ```
92
151
 
93
- **Verify capacities against your endpoint when you upgrade.** A wrong `contextWindow` is the failure that hurts, because the harness trusts it when it decides to compact history.
152
+ Two rules keep the filter from hiding something you asked for. An id named in `models` is advertised whatever the patterns and the snapshot say — naming a model is a stronger statement than either — and a pattern that does not compile is ignored with a warning instead of failing the mount. `include` and `exclude` replace the shipped defaults, so copy the ones you want to keep from `lib/catalog.js`. `exclude` is also the way to trim what the snapshot admits, e.g. `- '^siliconflow/'` to drop one vendor's aliases of models that are already listed canonically.
94
153
 
95
- ### Updating the catalog
154
+ Note what a discovered model gets for modalities: whatever the snapshot publishes, and `discovery.input` for everything else. A model neither source knows is advertised text-only, because neither a listing nor a guess says it accepts images; naming it in `models` settles the question.
96
155
 
97
- Edit `lib/catalog.js` and publish. Consumers get the new list with `pnpm update`.
156
+ After a failed attempt the endpoint is left alone for a minute, so an unreachable host cannot make every read slow; storing or rotating the key retries immediately.
98
157
 
99
- ### Overriding models for one profile
158
+ ### Pinning the catalog
100
159
 
101
- `models` is a normal configuration field whose default is the shipped catalog, so a profile can replace it without forking:
160
+ `models` decides the fallback membership *and* the metadata, and a profile can replace it without forking:
102
161
 
103
162
  ```yaml
104
163
  - id: llm-aliyun
@@ -114,7 +173,7 @@ Edit `lib/catalog.js` and publish. Consumers get the new list with `pnpm update`
114
173
  thinkingLevelMap: { low: low, medium: medium, high: high }
115
174
  ```
116
175
 
117
- The Models page edits the same field, showing the shipped catalog as inherited rows until the first edit materializes an override.
176
+ With `discovery.enabled: false` this list is the entire route.
118
177
 
119
178
  ### Thinking
120
179
 
@@ -122,6 +181,8 @@ Aliyun turns thinking on with a boolean `enable_thinking` rather than a reasonin
122
181
 
123
182
  `thinkingLevelMap` on each catalog entry declares which levels a model offers and the wire spelling of each. A level left out is not offered; `off` is offered unless it is named in the map.
124
183
 
184
+ A discovered model the fallback does not name is declared non-reasoning: a level map is a claim about a model, and claiming thinking it may not have would offer levels the endpoint can refuse. Name the model in `models` to give it levels.
185
+
125
186
  ## Uninstall
126
187
 
127
188
  Remove it from `dsh.profile.bundles` (or delete the row on the Plugins page), then:
@@ -130,7 +191,7 @@ Remove it from `dsh.profile.bundles` (or delete the row on the Plugins page), th
130
191
  dsh plugin --profile <name> remove @doitian/dsh-provider-aliyun
131
192
  ```
132
193
 
133
- The credential in `$DSH_HOME/.credentials.yaml` is left untouched. Deleting the route on the Models page removes it only when its reference is exactly the page-derived `ALIYUN_API_KEY`; a custom reference is retained deliberately, because the page cannot prove it owns it.
194
+ The credential in `$DSH_HOME/.credentials.yaml` is left untouched.
134
195
 
135
196
  ## How it works
136
197
 
@@ -148,11 +209,13 @@ and the patch inserts this package's own plugin entry:
148
209
  name: '@doitian/dsh-provider-aliyun'
149
210
  ```
150
211
 
151
- On mount, `lib/index.js` registers two things on the LLM seam: the `aliyun` route with an adapter, and a configurable-provider directory entry — the latter is what gives the route a row, and an API-key field, on the Models page.
212
+ On mount, `lib/index.js` registers three things on the LLM seam: the `aliyun` route with an adapter, a configurable-provider directory entry — that is the row on the Models page, not an API-key field (see [Configure the API key](#configure-the-api-key)) — and a model-discovery offer for the entry's settings namespace, so a configuration surface can interrogate this route's endpoint with a draft credential.
152
213
 
153
214
  The adapter is `PiAiAdapter`, exported by `@deepseek-ai/dsh-llm-pi-ai`. It owns the hard part — harness history into pi-ai context, pi-ai events into harness stream chunks, image budgets, replay metadata, idle watchdogs — and reusing it is what keeps this package a catalog plus a few dozen lines instead of a second adapter implementation.
154
215
 
155
- **The one thing to know if you maintain this.** `PiAiAdapter` is driven by a route's *resolved profile*, a shape its package does not export as a type; `lib/adapter.js` reproduces it and documents every field it must carry. Model metadata is not read from that profile — it comes from the pi-ai model descriptors built out of the catalog, which is why the catalog can live here at all. A DSH upgrade that starts reading a new profile field breaks this plugin. `test/adapter.test.mjs` drives the real published adapter against the factory, so that break shows up as a failing test rather than as a broken route in someone's profile.
216
+ **How the advertised catalog moves.** `lib/discovery.js` reads the endpoint's listing and holds one `ids` array; `lib/models.js` merges it with the fallback entries; `lib/adapter.js` swaps the profile *map* the adapter memoizes its snapshot by identity. An operation already in flight keeps the snapshot it started with, and the next read sees the new catalog — which is also why a refresh never interrupts a stream. The harness's session catalog calls `listModels()` per provider on every read, so the picker shows whatever the route advertises at that moment. Only `listModels` waits for a cold listing, and that wait lives in a two-line subclass rather than a reimplementation of the adapter.
217
+
218
+ **The one thing to know if you maintain this.** `PiAiAdapter` is driven by a route's *resolved profile*, a shape its package does not export as a type; `lib/adapter.js` reproduces it and documents every field it must carry. Model metadata is not read from that profile — it comes from the pi-ai model descriptors built out of the catalog, which is why the catalog can live here at all. A DSH upgrade that starts reading a new profile field breaks this plugin. `test/adapter.test.mjs` drives the real published adapter against the factory, so that break shows up as a failing test rather than as a broken route in someone's profile. The discovery contract is a second such seam: `LlmModelDiscoveryRequest`, `registerModelDiscovery`, and the `attributionHeaders()` requirement on every provider HTTP request.
156
219
 
157
220
  The engine is pinned to the harness's own pi-ai: the dependency is `^0.87.1`, which resolves to exactly the `0.87.1` the harness installs, so both share one copy.
158
221
 
@@ -160,13 +223,26 @@ The engine is pinned to the harness's own pi-ai: the dependency is `^0.87.1`, wh
160
223
 
161
224
  ```bash
162
225
  npm install
163
- npm run validate # manifest, patch wiring, catalog integrity
164
- npm test # drives the real PiAiAdapter against the catalog
226
+ npm run validate # manifest, patch wiring, fallback catalog, discovery defaults
227
+ npm test # drives the real PiAiAdapter, the merge, and the listing state machine
228
+ ```
229
+
230
+ `scripts/validate.mjs` catches what would otherwise fail silently: a tarball that drops the patch, a patch that names the wrong package, a duplicate model id, a non-integer capacity, a level map pi-ai would read as offering nothing, a discovery default that would produce a model the adapter cannot dispatch.
231
+
232
+ `npm test` needs the peer packages installed — that is the point, since it exercises the real adapter rather than a stub. A live request is out of scope: that needs a real key and endpoint. Discovery is tested with an injected `fetch`, so what is covered is every decision *around* the request — the shapes it reads, what a failure leaves in place, and how long a read may wait.
233
+
234
+ ### Testing a build against the desktop app
235
+
236
+ ```bash
237
+ npm run link:desktop # point the desktop profile at this checkout, then verify the copy
238
+ npm run probe # what the picker will show, read from the live endpoint
165
239
  ```
166
240
 
167
- `scripts/validate.mjs` catches what would otherwise fail silently: a tarball that drops the patch, a patch that names the wrong package, a duplicate model id, a non-integer capacity, a level map pi-ai would read as offering nothing.
241
+ Then **restart DSH**: plugin modules are loaded once at startup, and the loader does not watch them.
242
+
243
+ `link:desktop` exists because a local dependency is not a live view of the worktree. pnpm hard-links a `file:` package into the profile, so an edit that *replaces* a file — `git checkout`, or any editor that saves atomically — leaves the installed copy on the previous bytes, and a plain `install`, `--force`, or `update` will not re-link it ("Already up to date", even when the directory is deleted). Re-adding the dependency is what forces a fresh resolution, and the script verifies the result byte for byte instead of assuming it worked.
168
244
 
169
- `npm test` needs the peer packages installed — that is the point, since it exercises the real adapter rather than a stub. A live request is out of scope: that needs a real key and endpoint.
245
+ Two things to expect while testing this way. The Plugins page may rewrite the dependency back to a registry range — that is a normal package install, and `link:desktop` switches it back. And a *published* release is the durable alternative: the app then installs it like any other plugin, and nothing local is involved.
170
246
 
171
247
  ## Publishing
172
248
 
@@ -185,6 +261,25 @@ npm trust github @doitian/dsh-provider-aliyun \
185
261
 
186
262
  `--allow-publish` is required: it grants the `CREATE_PACKAGE` permission, and the command refuses without it or `--allow-stage-publish`. It infers `owner/repo` from `package.json` when `--repo` is omitted, warns when the two disagree, requires 2FA, and prompts for an OTP. Add `--dry-run` first to see exactly what it would create without committing it.
187
263
 
264
+ **A bypass-2FA token cannot do this step.** npm is retiring tokens that bypass 2FA, and the
265
+ restriction is asymmetric: such a token still *publishes*, but the registry refuses it for
266
+ publisher management with
267
+
268
+ ```
269
+ Granular access tokens that bypass two-factor authentication may not perform this action.
270
+ ```
271
+
272
+ So the bootstrap publish can come from a bypass token, while attaching the publisher needs an
273
+ interactive `npm login` session or the website form. Reading the config back is refused for
274
+ the same reason, so the release itself is the practical check.
275
+
276
+ **Account two-factor authentication does not reach the workflow.** The publish job authenticates
277
+ with an OIDC token, so no one-time password is involved and enabling 2FA on the npm account
278
+ cannot break a release — v0.1.1 was published exactly this way. What 2FA does gate is the
279
+ one-time publisher setup above: an account with 2FA can run `npm trust`, which is the step a
280
+ 2FA-bypassing token is refused. Publishing by hand from a terminal is the only flow that now
281
+ asks for an OTP.
282
+
188
283
  `--environment` is deliberately omitted, matching the publish job, which declares no `environment:`.
189
284
 
190
285
  If the registry reports the package as missing, a trusted publisher cannot be attached to a name that does not exist yet: publish `0.1.0` once by hand, then run the command above and let the workflow own every release after that.
@@ -192,11 +287,15 @@ If the registry reports the package as missing, a trusted publisher cannot be at
192
287
  Then release:
193
288
 
194
289
  ```bash
195
- npm version patch
290
+ npm version patch # or skip it when the version is already committed in package.json
196
291
  git push --follow-tags
197
292
  gh release create v0.1.1 --generate-notes
198
293
  ```
199
294
 
295
+ `npm version` writes the version, commits it, and tags it in one step. When the version you mean
296
+ to release is already committed by hand — a feature release that bumped `minor` itself — tag that
297
+ commit instead: `git tag v0.2.0 && git push --follow-tags`, then create the release for it.
298
+
200
299
  The publish job requires the release tag to match `package.json` (`v0.1.1` ↔ `0.1.1`) and re-runs validation and tests before uploading.
201
300
 
202
301
  Two things that will bite when editing the workflow:
package/lib/adapter.js CHANGED
@@ -24,6 +24,12 @@
24
24
  * from the profile. It is read from the pi-ai model descriptors inside
25
25
  * `piProvider`, which is why the catalog can live in this package.
26
26
  *
27
+ * The catalog it carries is not fixed at mount. Discovery ({@link ./discovery.js})
28
+ * swaps the advertised entries as the endpoint's listing moves, and the swap is
29
+ * a new profile *map* so the adapter's identity-based memoization notices it:
30
+ * every operation already captured keeps the snapshot it started with, and the
31
+ * next read sees the new catalog.
32
+ *
27
33
  * The risk this carries: a DSH upgrade that starts reading a new profile field
28
34
  * breaks this plugin. `test/adapter.test.mjs` drives the real `PiAiAdapter`
29
35
  * against this factory so that such a break shows up as a failing test rather
@@ -111,23 +117,60 @@ export function buildAliyunProfile(options) {
111
117
  * @param {string} options.displayName - label for selectors.
112
118
  * @param {string} options.apiKeyEnv - credential reference to resolve per request.
113
119
  * @param {string} options.baseURL - endpoint for every model on the route.
114
- * @param {readonly import('./catalog.js').AliyunModel[]} options.models - the shipped catalog.
120
+ * @param {readonly import('./catalog.js').AliyunModel[]} options.models - the catalog to advertise first.
115
121
  * @param {string} [options.reasoning] - route default thinking level.
116
122
  * @param {Record<string, string>} [options.headers] - extra request headers.
117
123
  * @param {(provider: string, profile: object) => Promise<string | undefined>} options.resolveApiKey - credential resolver.
118
124
  * @param {() => object | undefined} [options.resolveAttachments] - durable attachment service accessor.
119
125
  * @param {(attachments: object, ref: object) => unknown} [options.resolveImageAccess] - image access resolver.
120
126
  * @param {(event: object) => void} [options.onReplayDegrade] - replay degradation reporter.
121
- * @returns {PiAiAdapter} the adapter to register on the LLM seam.
127
+ * @param {() => void} [options.onRead] - called before every operation reads the profile map.
128
+ * @param {() => Promise<void>} [options.settle] - awaited before a model list is answered.
129
+ * @returns {{ adapter: PiAiAdapter, setCatalog: (models: readonly import('./catalog.js').AliyunModel[]) => void }} the adapter to register, and the way to move its catalog.
122
130
  */
123
131
  export function createAliyunAdapter(options) {
124
- const profile = buildAliyunProfile(options)
125
- const profiles = () => new Map([[ROUTE, profile]])
126
- return new PiAiAdapter({
127
- profiles,
132
+ /** The profile map the adapter memoizes by identity; a catalog swap replaces it. */
133
+ let profiles = profilesOver(options.models)
134
+
135
+ /** One route's profile for one catalog, as a fresh map. */
136
+ function profilesOver(models) {
137
+ return new Map([[ROUTE, buildAliyunProfile({ ...options, models })]])
138
+ }
139
+
140
+ /**
141
+ * `PiAiAdapter` answers `listModels` from its memoized snapshot, and that
142
+ * answer is what the picker reads. Waiting here — and nowhere else — is what
143
+ * lets a picker opened on a cold route show the endpoint's own list instead of
144
+ * the fallback one, while every other read stays off the network.
145
+ */
146
+ class AliyunAdapter extends PiAiAdapter {
147
+ async listModels(provider) {
148
+ await options.settle?.()
149
+ return super.listModels(provider)
150
+ }
151
+ }
152
+
153
+ const adapter = new AliyunAdapter({
154
+ // Every operation asks for the profile map, so a stale listing is noticed
155
+ // and refreshed on the way past, without any of those operations waiting.
156
+ profiles: () => {
157
+ options.onRead?.()
158
+ return profiles
159
+ },
128
160
  resolveApiKey: options.resolveApiKey,
129
161
  resolveAttachments: options.resolveAttachments,
130
162
  resolveImageAccess: options.resolveImageAccess,
131
163
  onReplayDegrade: options.onReplayDegrade,
132
164
  })
165
+
166
+ return {
167
+ adapter,
168
+ /**
169
+ * Advertise `models` from the next read on.
170
+ * @param {readonly import('./catalog.js').AliyunModel[]} models - the entries the route now serves.
171
+ */
172
+ setCatalog(models) {
173
+ profiles = profilesOver(models)
174
+ },
175
+ }
133
176
  }
package/lib/catalog.js CHANGED
@@ -1,18 +1,26 @@
1
1
  /**
2
- * The Aliyun DashScope (Bailian) model catalog this plugin ships.
2
+ * The Aliyun DashScope (Bailian) catalog this plugin ships.
3
3
  *
4
- * This file is the whole point of the package: the model list lives here, in
5
- * the plugin, so `pnpm update` refreshes it. Nothing about it needs to be
6
- * written into a profile, and no profile patch can shadow it.
4
+ * This file does two jobs, and the split is the point:
5
+ *
6
+ * - **Fallback membership.** {@link ./discovery.js} asks the configured endpoint
7
+ * which models it serves. Until that answer arrives — no credential stored
8
+ * yet, an endpoint that serves no listing, a network that cannot reach it —
9
+ * these entries are the whole route, so the picker is never empty.
10
+ * - **Metadata.** A listing discloses ids and nothing else, and the harness
11
+ * trusts `contextWindow` when it decides to compact history. An id named here
12
+ * carries the real capacities; an id discovered and *not* named here gets the
13
+ * conservative defaults from configuration, because an obviously generic
14
+ * capacity beats an invented precise one.
7
15
  *
8
16
  * It is plain data on purpose. {@link ./models.js} turns each entry into the
9
17
  * pi-ai model descriptor the adapter dispatches with, filling in the route,
10
18
  * endpoint, protocol, and compatibility switches from configuration — so an
11
19
  * entry here states only what is true of the *model*, never of the deployment.
12
20
  *
13
- * The entries mirror the facts Aliyun publishes for its Qwen lineup. Verify
14
- * them against your endpoint when you upgrade; a wrong `contextWindow` is the
15
- * failure that hurts, because the harness trusts it when it compacts history.
21
+ * These are the models this route is used with day to day. Verify the capacities
22
+ * against your own endpoint when you upgrade; a wrong `contextWindow` is the
23
+ * failure that hurts.
16
24
  */
17
25
 
18
26
  /** Route name every request selects with `GenerateOptions.provider`. */
@@ -27,6 +35,15 @@ export const DEFAULT_BASE_URL = 'https://dashscope.aliyuncs.com/compatible-mode/
27
35
  /** International Model Studio endpoint. */
28
36
  export const INTL_BASE_URL = 'https://dashscope-intl.aliyuncs.com/compatible-mode/v1'
29
37
 
38
+ /**
39
+ * The shape of a Bailian *workspace* endpoint — the per-workspace host the
40
+ * console hands out, and the spelling most keys are issued against.
41
+ *
42
+ * An example to recognize, not a default this package can ship: the workspace id
43
+ * is the deployment's own.
44
+ */
45
+ export const WORKSPACE_BASE_URL_EXAMPLE = 'https://llm-<workspace-id>.cn-beijing.maas.aliyuncs.com/compatible-mode/v1'
46
+
30
47
  /** Credential reference the route resolves through the harness credential seam. */
31
48
  export const DEFAULT_API_KEY_ENV = 'ALIYUN_API_KEY'
32
49
 
@@ -49,6 +66,81 @@ export const QWEN_COMPAT = Object.freeze({
49
66
  supportsStrictMode: true,
50
67
  })
51
68
 
69
+ /**
70
+ * What a discovered model gets when no entry below names it.
71
+ *
72
+ * Context is conservative on purpose: compacting early costs tokens, compacting
73
+ * late costs the request. Output sits comfortably under every cap this endpoint
74
+ * publishes, so a request is never refused for asking for more than a model
75
+ * allows.
76
+ */
77
+ export const DISCOVERY_DEFAULTS = Object.freeze({
78
+ contextWindow: 131_072,
79
+ maxTokens: 32_768,
80
+ input: Object.freeze(['text']),
81
+ })
82
+
83
+ /** How long a listing stays fresh before a read asks the endpoint again. */
84
+ export const DISCOVERY_TTL_MS = 600_000
85
+
86
+ /** Longest a cold or stale model-list read waits before answering from what it has. */
87
+ export const DISCOVERY_WAIT_MS = 2_000
88
+
89
+ /**
90
+ * The chat families this route advertises by default, as case-insensitive
91
+ * regular expressions.
92
+ *
93
+ * A DashScope listing discloses `id`, `object`, `created`, and `owned_by` — no
94
+ * modality, no task type — so what an id *is* can only be read from its name.
95
+ * That matters because a workspace endpoint lists everything the account can
96
+ * reach: speech synthesis, speech recognition, image generation, embeddings,
97
+ * rerankers, realtime duplex models, and every legacy generation back to
98
+ * `qwen-7b-chat`. Advertising all of it makes the picker unusable and offers
99
+ * models that cannot answer a chat request at all.
100
+ *
101
+ * The patterns are anchored, which is also what drops vendor-prefixed aliases
102
+ * (`ZHIPU/GLM-5.3`, `vanchin/deepseek-v4.1-flash`) whose canonical id is already
103
+ * matched — the two prefixed exceptions are families with no canonical spelling
104
+ * on this endpoint.
105
+ *
106
+ * This is a *default*, not a rule. `discovery.filter: all` advertises the whole
107
+ * listing, `discovery.include` adds families, `discovery.exclude` drops ids, and
108
+ * an id named in `models` is advertised whatever these say.
109
+ */
110
+ export const CHAT_MODEL_PATTERNS = Object.freeze([
111
+ '^qwen3\\.8-(max|flash|omni-flash)$',
112
+ '^qwen3\\.7-(max|plus|flash)$',
113
+ '^qwen3\\.6-(plus|flash)$',
114
+ '^qwen3\\.5-(plus|flash|omni-plus|omni-flash)$',
115
+ // Open-weight releases carry their size in the id, which is how they are told
116
+ // apart from the hosted aliases of the same generation.
117
+ '^qwen3\\.[5-8]-\\d+b(-a\\d+b)?$',
118
+ '^qwen3-(max|omni-flash)$',
119
+ '^qwen3-coder-(plus|flash)$',
120
+ '^deepseek-(v4\\.1-flash|v4-pro|v4-flash|v3(\\.[12])?|r1)$',
121
+ '^glm-(4\\.7|5(\\.\\d+)?(-prime)?)$',
122
+ '^kimi-k[23](\\.\\d+)?(-code|-thinking)?$',
123
+ '^MiniMax-M[23](\\.\\d+)?$',
124
+ '^stepfun/step-[0-9.]+(-flash)?$',
125
+ '^xiaomi/mimo-[a-z0-9.]+$',
126
+ ])
127
+
128
+ /**
129
+ * Names to drop even when an include pattern matched, because these words are
130
+ * how this endpoint spells "not a plain chat completion model": a snapshot of a
131
+ * model this list already carries, an unreleased preview, or a streaming-first
132
+ * variant.
133
+ */
134
+ export const NON_CHAT_MODEL_PATTERNS = Object.freeze([
135
+ '-realtime$',
136
+ '-preview$',
137
+ '-\\d{4}-\\d{2}-\\d{2}$',
138
+ 'embedding',
139
+ 'rerank',
140
+ '-highspeed$',
141
+ '-flashx$',
142
+ ])
143
+
52
144
  /**
53
145
  * One shipped model.
54
146
  *
@@ -61,11 +153,13 @@ export const QWEN_COMPAT = Object.freeze({
61
153
  * @property {boolean} reasoning - whether the model can think.
62
154
  * @property {Record<string, string>} [thinkingLevelMap] - selectable thinking
63
155
  * levels and the wire spelling each one sends. A level left out is not
64
- * offered; `off` is offered unless it is named here.
156
+ * offered; `off` is offered unless it is named here. This endpoint takes a
157
+ * boolean rather than an effort, so what the map really decides is the *set*
158
+ * of levels a selector offers.
65
159
  */
66
160
 
67
161
  /** @type {readonly AliyunModel[]} */
68
- export const ALIYUN_MODELS = Object.freeze([
162
+ export const FALLBACK_MODELS = Object.freeze([
69
163
  {
70
164
  id: 'qwen3.8-max',
71
165
  name: 'Qwen3.8 Max',
@@ -76,48 +170,23 @@ export const ALIYUN_MODELS = Object.freeze([
76
170
  thinkingLevelMap: { low: 'low', medium: 'medium', high: 'high' },
77
171
  },
78
172
  {
79
- id: 'qwen3.8-flash',
80
- name: 'Qwen3.8 Flash',
173
+ // Aliyun publishes 384K output for this model, and the Aliyun catalog inside
174
+ // the pinned pi-ai carries the same figure.
175
+ id: 'deepseek-v4.1-flash',
176
+ name: 'DeepSeek V4.1 Flash',
81
177
  contextWindow: 1_000_000,
82
- maxTokens: 131_072,
178
+ maxTokens: 384_000,
83
179
  input: ['text', 'image'],
84
180
  reasoning: true,
85
181
  thinkingLevelMap: { low: 'low', medium: 'medium', high: 'high' },
86
182
  },
87
183
  {
88
- id: 'qwen3.7-max',
89
- name: 'Qwen3.7 Max',
184
+ id: 'glm-5.3',
185
+ name: 'GLM-5.3',
90
186
  contextWindow: 1_000_000,
91
187
  maxTokens: 131_072,
92
188
  input: ['text'],
93
189
  reasoning: true,
94
190
  thinkingLevelMap: { low: 'low', medium: 'medium', high: 'high' },
95
191
  },
96
- {
97
- id: 'qwen3.7-plus',
98
- name: 'Qwen3.7 Plus',
99
- contextWindow: 1_000_000,
100
- maxTokens: 65_536,
101
- input: ['text', 'image'],
102
- reasoning: true,
103
- thinkingLevelMap: { low: 'low', medium: 'medium', high: 'high' },
104
- },
105
- {
106
- id: 'qwen3.6-plus',
107
- name: 'Qwen3.6 Plus',
108
- contextWindow: 1_000_000,
109
- maxTokens: 65_536,
110
- input: ['text', 'image'],
111
- reasoning: true,
112
- thinkingLevelMap: { low: 'low', medium: 'medium', high: 'high' },
113
- },
114
- {
115
- id: 'qwen3.6-flash',
116
- name: 'Qwen3.6 Flash',
117
- contextWindow: 1_000_000,
118
- maxTokens: 65_536,
119
- input: ['text', 'image'],
120
- reasoning: true,
121
- thinkingLevelMap: { low: 'low', medium: 'medium', high: 'high' },
122
- },
123
192
  ])