@leaf233/dsh-llm-rate-limiter 0.1.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md ADDED
@@ -0,0 +1,33 @@
1
+ # Changelog
2
+
3
+ All notable changes to this project will be documented in this file.
4
+
5
+ The format is based on [Keep a Changelog](https://keepachangelog.com/).
6
+
7
+ ## [0.1.1] - 2026-09-12
8
+
9
+ ### Changed
10
+ - **Package renamed to `@leaf233/dsh-llm-rate-limiter`** (npm scope) and published to npm; the internal plugin identity (`llm-rate-limiter`), settings namespace, and runtime behavior are unchanged
11
+ - `cordis.patch.yml` `name` and the browser client bundle id now follow the scoped package name (required by DSH module resolution)
12
+ - npm badge and npm install instructions added to README
13
+
14
+ ### Fixed (since the `v0.1.0` git tag)
15
+ - CI: `pnpm/action-setup` now precedes `actions/setup-node` (pnpm was missing from PATH)
16
+ - CI: `pnpm-lock.yaml` committed, so `--frozen-lockfile` installs work
17
+ - Removed private `@deepseek-ai/dsh-llm` from `peerDependencies` (unresolvable for consumers)
18
+
19
+ ## [0.1.0] - 2026-09-12
20
+
21
+ ### Added
22
+ - Token Bucket strategy with configurable `burstSize`, `refillRate`, `maxConcurrent`
23
+ - Sliding Window strategy with `windowMs`, `maxRpm`, `maxConcurrent`
24
+ - Per-model rate limit overrides (`models["provider/model"]`)
25
+ - Queue mode: throttled requests wait and are released when slots open
26
+ - Reject mode: throttled requests fail with `RATE_LIMIT` code (compatible with `dsh-llm-retry`)
27
+ - Interactive GUI in DSH Settings → Plugins (collapsible PluginCard)
28
+ - `maxRpm` → `refillRate` auto-derivation for token-bucket when only `maxRpm` is set
29
+ - Hot-reload: settings changes take effect immediately
30
+ - `maxQueueWaitMs` timeout for queue mode
31
+ - 19 unit tests covering concurrency, burst, queue, abort, and hot-reload
32
+ - E2E test (`test-3rpm.mjs`) verifying 3 rpm limit
33
+ - DSH version compatibility analysis (COMPATIBILITY.md)
@@ -0,0 +1,133 @@
1
+ # dsh-llm-rate-limiter — DSH 版本兼容性检测报告
2
+
3
+ > 检测日期: 2026-07-28
4
+ > 本地 DSH 版本: **0.1.2-rc.1**
5
+ > 测试方式: 源码分析(npm registry 在沙箱中被拦截,无法查询其他版本)
6
+
7
+ ---
8
+
9
+ ## 一、依赖 API 清单与存在性验证
10
+
11
+ ### Host 端(lib/index.js)— 6 个 API 依赖
12
+
13
+ | # | API | 所属包 | 本地存在 | 稳定性评估 |
14
+ |---|-----|--------|----------|------------|
15
+ | 1 | `ctx.settings.register(ns, schema, { base })` | dsh-settings | ✅ 3处 | 核心 API,SettingsService 的公共接口 |
16
+ | 2 | `scope.get()` | dsh-settings (返回值) | ✅ `registration.resolved` | 简单属性读取,极低风险 |
17
+ | 3 | `scope.watch(callback)` → `unwatch()` | dsh-settings (返回值) | ✅ watcher Set 管理 | 标准 observer 模式,极低风险 |
18
+ | 4 | `ctx.on("llm/stream", async function*(options, next))` | dsh-llm + cordis | ✅ 3处 | waterfall 中间件,LLM 调用的核心拦截点 |
19
+ | 5 | `ctx.effect(() => cleanup, label)` | cordis | ✅ | Cordis 核心生命周期 API |
20
+ | 6 | `ctx.logger?.("rate-limiter").info/warn(...)` | cordis | ✅ | 日志 API,如不存在用 `?.` 降级 |
21
+
22
+ ### Client 端(lib/client.js)— 5 个 API 依赖
23
+
24
+ | # | API | 所属包 | 本地存在 | 稳定性评估 |
25
+ |---|-----|--------|----------|------------|
26
+ | 1 | `window.__ModuleLoader__.load({ id, factory })` | DSH web shell | ✅ 1处 | 客户端模块加载器,所有 client 插件都用 |
27
+ | 2 | `settingsScope.get?.()` / `getSnapshot?.()` | dsh-client-ui-settings | ✅ 各1处 | get() 来自 host scope, getSnapshot 来自 client controller; 双兼容 |
28
+ | 3 | `settingsScope.subscribe(listener)` | dsh-client-ui-settings | ✅ 1处 | SettingsScopeController.subscribe |
29
+ | 4 | `settingsScope.set(field, value)` | dsh-client-ui-settings | ✅ 2处 | SettingsScopeController.set (单字段写) |
30
+ | 5 | `settingsScope.mutate([{ op, path, value }])` | dsh-client-ui-settings | ✅ 1处 | SettingsScopeController.mutate (嵌套操作) |
31
+ | 6 | `ctx.slots.inject("settings.plugin.item", ...)` | dsh-client-ui-settings-plugins | ✅ 6处 | 插件设置卡片注册 slot |
32
+
33
+ ### Schema 依赖 — 1 个包
34
+
35
+ | 包 | 本地版本 | 用到的 API |
36
+ |----|---------|-----------|
37
+ | @deepseek-ai/schemastery | 3.18.2+ | `Schema.object()`, `Schema.union()`, `Schema.const()`, `Schema.dict()`, `.default()`, `.int()`, `.min()`, `.description()` |
38
+
39
+ ---
40
+
41
+ ## 二、版本矩阵评估
42
+
43
+ ### 已知版本
44
+
45
+ | 组件 | 本地安装版本 | peerDep 要求 | DSH 用法 |
46
+ |------|------------|-------------|---------|
47
+ | @deepseek-ai/dsh | **0.1.2-rc.1** | — | 安装的 DSH 主包 |
48
+ | @deepseek-ai/cordis | **4.0.2** | ≥4.0.2 | DSH 所有插件统一用 ^4.0.2 |
49
+ | @deepseek-ai/schemastery | 3.18.2+ | ≥3.18.0 | DSH 用 ^3.18.2 |
50
+ | @deepseek-ai/dsh-settings | **0.1.2-rc.1** | — | 提供 `ctx.settings.register()` |
51
+ | @deepseek-ai/dsh-llm | **0.1.2-rc.1** | — | 提供 `llm/stream` waterfall |
52
+ | @deepseek-ai/dsh-llm-retry | **0.1.2-rc.1** | — | 本插件的参考实现 |
53
+ | @deepseek-ai/dsh-client-ui-settings | **0.1.2-rc.1** | — | 提供 `settingsScope` 服务 |
54
+ | @deepseek-ai/dsh-client-ui-settings-plugins | **0.1.2-rc.1** | — | 提供 `settings.plugin.item` slot |
55
+
56
+ ### 兼容性矩阵(推断)
57
+
58
+ | DSH 版本 | 预期兼容 | 风险点 |
59
+ |----------|---------|--------|
60
+ | 0.1.0 ~ 0.1.2 | ✅ 应兼容 | RC 阶段 API 趋于稳定,settings.register / llm/stream 已是核心 |
61
+ | 0.1.3+ (同 minor) | ✅ 应兼容 | 遵循 semver,接口不大改 |
62
+ | 0.2.x (minor 升级) | ⚠️ 需验证 | 可能新增/重命名 settings 参数、llm/stream 签名变体 |
63
+ | 1.0+ (major) | ⚠️ 必须重测 | 核心 API 可能重构(cordis 升级、settings 层重构) |
64
+
65
+ ---
66
+
67
+ ## 三、关键风险点分析
68
+
69
+ ### 🔴 高风险
70
+
71
+ | 风险 | 说明 | 影响 | 缓解措施 |
72
+ |------|------|------|----------|
73
+ | **`llm/stream` waterfall 签名变更** | DSH 当前签名: `(options: GenerateOptions, next) => AsyncIterable`,如果未来增加参数或改变 options 结构 | 插件读不到 `provider`/`model` 字段 | 用可选链 `options.provider ?? "unknown"` 降级 |
74
+ | **cordis major 升级** | cordis 是 DSH 的运行时内核,major 版本会改变插件生命周期 | apply/signature/effect 全部受影响 | peerDep 已锁定 `≥4.0.2`,major 升级时必须更新 |
75
+ | **0.1.x 是 RC 阶段** | 预发布版本的 API 不保证向后兼容 | 任何 patch 版本都可能引入 breaking change | 紧跟 DSH 版本更新 |
76
+
77
+ ### 🟡 中等风险
78
+
79
+ | 风险 | 说明 | 影响 | 缓解措施 |
80
+ |------|------|------|----------|
81
+ | **`settingsScope` API 名字变更** | Service 字符串 `"settingsScope"` 硬编码在 `SettingsScopeBinder` 构造函数中 | Client 端服务注入失败 | 该 Service 名是 UI 基础设施,改名成本极高,短期低概率 |
82
+ | **`settings.plugin.item` slot 名变更** | 插件设置页的 slot 注册点 | 卡片不会显示 | slot 名已被多处硬编码引用(Bash、AgentLoop 等),改名需全量迁移 |
83
+ | **schemastery 3.x → 4.x** | 新 major 可能改 `Schema.union/const` 等 API | 配置 Schema 编译失败 | peerDep 锁 ≥3.18.0,major 时必须适配 |
84
+
85
+ ### 🟢 低风险
86
+
87
+ | 风险 | 说明 |
88
+ |------|------|
89
+ | `ctx.on("llm/stream")` 事件名变更 | LLM waterfall 是 LLM 模块的核心公开接口 |
90
+ | `scope.get()` / `scope.watch()` 签名变更 | 简单 getter/observer,改动概率极低 |
91
+ | `settingsScope.set(field, value)` 签名变更 | 标准 setter,多个 UI 卡片都在用 |
92
+
93
+ ---
94
+
95
+ ## 四、已验证的兼容性事实
96
+
97
+ ### API 表面稳定性证据
98
+
99
+ 1. **`llm/stream` waterfall 被 2 处引用**:`dsh-llm/lib/index.js` 的 `stream()` 和 `invariant.js` 的验证层,形成双重稳定约束
100
+ 2. **`settings.register` 返回的 scope**:host 侧返回 `{ get, watch, update, replace }`,结构明确且简洁
101
+ 3. **`settingsScope.bind`** 由 `dsh-client-ui-settings` 提供,client 端 scope 返回 `{ set, unset, mutate, subscribe, getSnapshot }`,被 `dsh-client-ui-settings-plugins`、`dsh-client-ui-settings-models` 等多个官方 UI 包使用
102
+ 4. **`dsh-llm-retry` 作为参考**:同样使用 `ctx.on` 事件 + `ctx.effect` 清理,peerDep 了 `@deepseek-ai/cordis: ^4.0.2`,与本插件策略一致
103
+ 5. **cordis ^4.0.2 被所有 DSH 插件统一引用**:`dsh-llm`、`dsh-llm-retry`、`dsh-client-ui-settings` 都声明同一个范围
104
+
105
+ ---
106
+
107
+ ## 五、兼容性加固措施(已应用)
108
+
109
+ | # | 措施 | 位置 |
110
+ |---|------|------|
111
+ | 1 | client.js 的 `settingsScope.get?.()` / `.getSnapshot?.()` 双兼容 | 读取初始值 |
112
+ | 2 | client.js 的 `watch` / `subscribe` 双兼容(typeof 检测) | 监听配置变化 |
113
+ | 3 | host.js 的 `ctx.logger?.()` 可选链 | 日志降级 |
114
+ | 4 | peerDep 使用 `≥4.0.2` 范围(允许 minor 升级) | package.json |
115
+ | 5 | schemastery peerDep 使用 `≥3.18.0`(允许 minor 升级) | package.json |
116
+ | 6 | 终端 chunk 使用 DSH 标准格式 `{ type: "finish", reason: { kind, failure } }` | 与 adapterFailureChunk 一致 |
117
+
118
+ ---
119
+
120
+ ## 六、建议
121
+
122
+ 1. **当前 DSH 0.1.2-rc.1**:插件完全兼容,所有 API 已验证
123
+ 2. **发布时建议声明 peerDep**:
124
+ - `@deepseek-ai/cordis`: `>=4.0.2`
125
+ - `@deepseek-ai/dsh-llm`: `>=0.1.0`(因为我们只用 `llm/stream` 事件)
126
+ - `@deepseek-ai/schemastery`: `>=3.18.0`
127
+ 3. **DSH 升级到 0.2.x+ 时必须回归测试**以下清单:
128
+ - [ ] `llm/stream` 事件签名是否变化
129
+ - [ ] `options.provider` / `options.model` 字段是否存在
130
+ - [ ] `settings.register` 返回值结构是否变化
131
+ - [ ] `settingsScope.set/mutate` 接口是否变化
132
+ - [ ] `settings.plugin.item` slot 是否仍可注入
133
+ 4. **如果 cordis 升级到 5.0+**:整个 `apply(ctx)` 接口、`ctx.effect()`、`ctx.on()` 签名可能重写,需要全面适配
package/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 dsh-llm-rate-limiter contributors
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
package/README.md ADDED
@@ -0,0 +1,190 @@
1
+ # dsh-llm-rate-limiter
2
+
3
+ [![npm](https://img.shields.io/npm/v/@leaf233/dsh-llm-rate-limiter.svg)](https://www.npmjs.com/package/@leaf233/dsh-llm-rate-limiter)
4
+ [![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](LICENSE)
5
+ [![DSH 0.1.x](https://img.shields.io/badge/DSH-0.1.x-brightgreen.svg)](COMPATIBILITY.md)
6
+ [![Cordis 4.x](https://img.shields.io/badge/Cordis-%3E%3D4.0.2-brightgreen.svg)](package.json)
7
+ [![Tests](https://img.shields.io/badge/tests-19%2F19%20passing-brightgreen.svg)](test-strategies.mjs)
8
+
9
+ Per-model LLM call rate limiter for [DeepSeek Harness](https://github.com/deepseek-ai/dsh) with queue/reject support and interactive GUI configuration.
10
+
11
+ ---
12
+
13
+ ## Features
14
+
15
+ - **Per-model rate limiting** — independent concurrency, RPM, and burst limits for each `provider/model`
16
+ - **Two algorithms** — Token Bucket (allows bursts) or Sliding Window (smooth, strict RPM)
17
+ - **Queue mode** — throttled requests wait in queue and are released when a slot opens
18
+ - **Reject mode** — throttled requests fail immediately (integrates with `dsh-llm-retry` for auto-backoff)
19
+ - **Interactive GUI** — collapsible card in DSH Settings → Plugins → Configurable
20
+ - **Hot-reload** — settings changes take effect immediately, no restart needed
21
+ - **Every request checked** — intercepts `llm/stream` waterfall, covering every LLM call in every agent turn
22
+
23
+ ---
24
+
25
+ ## Installation
26
+
27
+ ### Option 1: npm (recommended)
28
+
29
+ ```bash
30
+ dsh plugin add <your-profile> @leaf233/dsh-llm-rate-limiter
31
+ # or, inside the profile directory:
32
+ pnpm add @leaf233/dsh-llm-rate-limiter
33
+ ```
34
+
35
+ ### Option 2: local path (development)
36
+
37
+ ```bash
38
+ dsh plugin add <your-profile> ./path/to/dsh-llm-rate-limiter
39
+ # or
40
+ dsh plugin add ./path/to/dsh-llm-rate-limiter # default profile
41
+ ```
42
+
43
+ > The plugin must be added as a dependency in the profile's `package.json`.
44
+ > The bundle entry (`cordis.patch.yml`) is auto-detected by `reconcilePlugins`.
45
+
46
+ ### Option 3: from GitHub
47
+
48
+ ```bash
49
+ dsh plugin add <your-profile> github:Leafyezi233/dsh-llm-rate-limiter
50
+ ```
51
+
52
+ > ⚠️ **Important**: Git-hosted plugins are blocked by pnpm's `allowBuilds` restriction on first install. If the install fails, check the error message for the exact key pnpm suggests, then add it to your profile's `pnpm-workspace.yaml`:
53
+ >
54
+ > ```yaml
55
+ > pnpm:
56
+ > allowBuilds:
57
+ > - '@leaf233/dsh-llm-rate-limiter'
58
+ > ```
59
+ >
60
+ > Then re-run the install command.
61
+
62
+ ---
63
+
64
+ ## Configuration
65
+
66
+ ### Via GUI
67
+
68
+ 1. Open DSH Web UI (`dsh web`)
69
+ 2. Go to **Settings → Plugins**
70
+ 3. Find **⚙ LLM 调用限速** card — click to expand
71
+ 4. Configure defaults, per-model overrides, and throttle behavior
72
+
73
+ ### Via file
74
+
75
+ Edit the profile's `settings.yaml` or use the GUI — changes are persisted to the DSH settings store:
76
+
77
+ ```yaml
78
+ llm-rate-limiter:
79
+ enabled: true
80
+ strategy: token-bucket # "token-bucket" | "sliding-window"
81
+ defaults:
82
+ maxConcurrent: 5
83
+ maxRpm: 60
84
+ burstSize: 10 # token-bucket only
85
+ refillRate: 1 # token-bucket only (tokens/sec)
86
+ models:
87
+ "deepseek/deepseek-chat":
88
+ maxConcurrent: 8
89
+ maxRpm: 120
90
+ "openai/gpt-4o":
91
+ maxConcurrent: 2
92
+ maxRpm: 10
93
+ burstSize: 3
94
+ "anthropic/claude-3-5-sonnet":
95
+ enabled: false # skip rate limiting for this model
96
+ onThrottled: queue # "queue" | "reject"
97
+ maxQueueWaitMs: 60000
98
+ ```
99
+
100
+ ---
101
+
102
+ ## Settings Reference
103
+
104
+ | Field | Default | Description |
105
+ |-------|---------|-------------|
106
+ | `enabled` | `true` | Global on/off switch. When off, zero overhead bypass. |
107
+ | `strategy` | `"token-bucket"` | `"token-bucket"` (allows bursts) or `"sliding-window"` (smooth, strict RPM) |
108
+ | `defaults.maxConcurrent` | `5` | Max simultaneous requests per model |
109
+ | `defaults.maxRpm` | `60` | Max requests per minute per model |
110
+ | `defaults.burstSize` | `10` | Token bucket capacity — how many requests can burst at once |
111
+ | `defaults.refillRate` | `1` | Tokens refilled per second (token-bucket). Auto-derived from `maxRpm / 60` if not set. |
112
+ | `models.<key>.maxConcurrent` | — | Per-model concurrency override |
113
+ | `models.<key>.maxRpm` | — | Per-model RPM override |
114
+ | `models.<key>.burstSize` | — | Per-model burst capacity override |
115
+ | `models.<key>.refillRate` | — | Per-model refill rate override |
116
+ | `models.<key>.enabled` | — | Set `false` to skip rate limiting for this specific model |
117
+ | `onThrottled` | `"queue"` | What happens when a request hits the limit: `"queue"` (wait) or `"reject"` (fail immediately) |
118
+ | `maxQueueWaitMs` | `60000` | Max time (ms) a request waits in queue before being rejected |
119
+
120
+ > **Note:** When a model overrides `maxRpm` without explicitly setting `refillRate`, the refill rate is automatically derived as `maxRpm / 60` (tokens per second). This ensures "set maxRpm=3" actually limits to 3 requests per minute.
121
+
122
+ ---
123
+
124
+ ## Algorithm Comparison
125
+
126
+ | | Token Bucket | Sliding Window |
127
+ |---|---|---|
128
+ | **Burst** | Yes (controlled by `burstSize`) | No — strictly smooth |
129
+ | **Recovery** | Tokens refill at `refillRate`/sec | Window slides continuously |
130
+ | **Best for** | Tolerating request spikes | APIs with hard per-minute limits |
131
+ | **GUI label** | 令牌桶 (Token Bucket) | 滑动窗口 (Sliding Window) |
132
+
133
+ ---
134
+
135
+ ## How It Works
136
+
137
+ ```
138
+ Agent Turn
139
+ → LLM Call (e.g. deepseek/deepseek-chat)
140
+ → ctx.on("llm/stream") interceptor
141
+ → Resolve rate limiter for this provider/model
142
+ → Token bucket: has tokens + concurrency room?
143
+ → If YES: consume token, acquire slot, forward to API
144
+ → If NO (reject mode): return RATE_LIMIT error immediately
145
+ → If NO (queue mode): park in waiters[], wait for token refill
146
+ → Request completes → release slot → drain waiting requests
147
+ → dsh-llm-retry catches RATE_LIMIT → exponential backoff → retry
148
+ ```
149
+
150
+ ---
151
+
152
+ ## Development
153
+
154
+ ```bash
155
+ # Clone
156
+ git clone https://github.com/Leafyezi233/dsh-llm-rate-limiter.git
157
+ cd dsh-llm-rate-limiter
158
+
159
+ # Install deps
160
+ pnpm install
161
+
162
+ # Run tests (19 tests)
163
+ node test-strategies.mjs
164
+
165
+ # Run E2E rate-limit test
166
+ node test-3rpm.mjs
167
+
168
+ # Install into a DSH profile for testing
169
+ dsh plugin add <your-profile> .
170
+ ```
171
+
172
+ The plugin uses a live symlink when installed via `link:` — edits to `lib/` take effect on browser hard-refresh (`Ctrl+Shift+R`) without reinstalling.
173
+
174
+ ---
175
+
176
+ ## Compatibility
177
+
178
+ | DSH Version | Status | Notes |
179
+ |-------------|--------|-------|
180
+ | 0.1.x (RC) | ✅ Tested | Verified against 0.1.2-rc.1, cordis 4.0.2 |
181
+ | 0.2.x | ⚠️ Untested | May need API adjustments |
182
+ | Cordis 5+ | ⚠️ Untested | Major version change likely requires rewrite |
183
+
184
+ See [COMPATIBILITY.md](COMPATIBILITY.md) for detailed API dependency analysis.
185
+
186
+ ---
187
+
188
+ ## License
189
+
190
+ [MIT](LICENSE)
@@ -0,0 +1,3 @@
1
+ - insert:
2
+ - id: llm-rate-limiter
3
+ name: '@leaf233/dsh-llm-rate-limiter'
package/lib/client.js ADDED
@@ -0,0 +1,233 @@
1
+ /* eslint-disable */
2
+ /**
3
+ * dsh-llm-rate-limiter — Client (browser) entry point.
4
+ *
5
+ * Registers a "Rate Limiter" card in the Plugins → Configurable settings page.
6
+ * The card uses a collapsible header (PluginCard pattern) with a chevron toggle.
7
+ *
8
+ * Browser format: CJS-style factory registered via window.__ModuleLoader__.load
9
+ * (loaded as a plain <script>, so NO top-level ESM `export` allowed).
10
+ *
11
+ * @module dsh-llm-rate-limiter/client
12
+ */
13
+ window.__ModuleLoader__.load({
14
+ id: "@leaf233/dsh-llm-rate-limiter",
15
+ factory: (require) => {
16
+ var module = { exports: {} };
17
+ var exports = module.exports;
18
+ Object.defineProperty(exports, Symbol.toStringTag, { value: "Module" });
19
+
20
+ // ── React imports (browser module system provides them) ──────────────
21
+ let react, jsx, jsxs;
22
+ react = require("react");
23
+ ({ jsx, jsxs } = require("react/jsx-runtime"));
24
+
25
+ // ── Shared constants ──
26
+ const name = "llm-rate-limiter";
27
+ const inject = ["slots", "settingsScope"];
28
+
29
+ /* ── CSS-in-JS (design tokens, injected once) ────────────────────── */
30
+ (() => {
31
+ const styles = [
32
+ ".rlr-card{border:.5px solid var(--dsw-alias-border-l4);background:var(--dsw-alias-bg-layer-3);border-radius:16px;list-style:none;transition:border-color .16s,background .16s;margin:0}",
33
+ ".rlr-card:hover{border-color:var(--dsw-alias-label-dimmed)}",
34
+ ".rlr-cardOpen{background:var(--dsw-alias-bg-layer-2);border-color:var(--dsw-alias-label-dimmed)}",
35
+ ".rlr-header{appearance:none;width:100%;font:inherit;color:inherit;text-align:left;cursor:pointer;background:0 0;border:0;border-radius:12px;align-items:center;gap:12px;padding:14px 16px;display:flex}",
36
+ ".rlr-header:focus-visible{outline:2px solid var(--dsw-alias-brand-primary);outline-offset:-2px}",
37
+ ".rlr-headText{flex-direction:column;flex:1;gap:4px;min-width:0;display:flex}",
38
+ ".rlr-name{color:var(--dsw-alias-label-primary);font-size:15px;font-weight:600;line-height:1.4}",
39
+ ".rlr-description{color:var(--dsw-alias-label-tertiary);font-size:13px;line-height:1.5}",
40
+ ".rlr-chevron{color:var(--dsw-alias-label-tertiary);flex:none;transition:transform .16s;font-size:12px}",
41
+ ".rlr-chevronOpen{transform:rotate(180deg)}",
42
+ ".rlr-body{border-top:.5px solid var(--dsw-alias-border-l2);margin:0 16px;padding:8px 0 8px}",
43
+ ".rlr-status{color:var(--dsw-alias-label-tertiary);margin:8px 0 4px;font-size:12px;line-height:1.5}",
44
+ ".rlr-row{display:flex;align-items:center;gap:10px;flex-wrap:wrap;margin:6px 0}",
45
+ ".rlr-label{min-width:140px;font-size:13px;color:var(--dsw-alias-label-secondary)}",
46
+ ".rlr-input{width:80px;padding:4px 8px;border-radius:6px;border:1px solid var(--dsw-alias-border-l4);background:var(--dsw-alias-surface-s2);color:var(--dsw-alias-label-primary);font-size:13px}",
47
+ ".rlr-select{padding:4px 8px;border-radius:6px;border:1px solid var(--dsw-alias-border-l4);background:var(--dsw-alias-surface-s2);color:var(--dsw-alias-label-primary);font-size:13px}",
48
+ ".rlr-btn{height:30px;padding:0 12px;border-radius:15px;border:none;cursor:pointer;font-size:13px;font-weight:500;color:var(--dsw-alias-label-primary-foreground);background:var(--dsh-alias-button-primary-fill,#4f46e5)}",
49
+ ".rlr-btnSec{height:26px;padding:0 10px;border-radius:13px;border:none;cursor:pointer;font-size:12px;background:var(--dsw-alias-button-secondary-fill);color:var(--dsw-alias-label-primary)}",
50
+ ".rlr-danger{height:26px;padding:0 8px;border-radius:13px;border:none;cursor:pointer;font-size:11px;background:var(--dsw-alias-state-error-primary);color:#fff}",
51
+ ".rlr-modelRow{display:flex;align-items:center;justify-content:space-between;padding:8px 12px;border-radius:12px;border:1px solid var(--dsw-alias-border-l4);gap:8px;margin:4px 0}",
52
+ ".rlr-modelName{font-size:13px;font-weight:500;font-family:monospace}",
53
+ ".rlr-modelMeta{display:flex;gap:12px;font-size:12px;color:var(--dsw-alias-label-tertiary)}",
54
+ ".rlr-section{display:flex;flex-direction:column;gap:4px;margin-top:2px}",
55
+ ".rlr-sectHead{margin:10px 0 0;font-size:14px;font-weight:500}",
56
+ ].join("\n");
57
+ const tagId = "@dsh-llm-rate-limiter/rate-limiter-card.css";
58
+ if (typeof document !== "undefined" && document.querySelector("style[data-plugin-css=" + JSON.stringify(tagId) + "]") === null) {
59
+ const tag = document.createElement("style");
60
+ tag.dataset.plugin = "dsh-llm-rate-limiter";
61
+ tag.dataset.pluginCss = tagId;
62
+ tag.textContent = styles;
63
+ document.head.appendChild(tag);
64
+ }
65
+ })();
66
+
67
+ /* ── helpers ─────────────────────────────────────────────────────── */
68
+ function readNum(cfg, key, field) {
69
+ return cfg.models?.[key]?.[field] ?? cfg.defaults?.[field] ?? 0;
70
+ }
71
+ function setNested(scope, path, value) {
72
+ scope.mutate([{ op: "set", path, value }]);
73
+ }
74
+ function unsetNested(scope, path) {
75
+ scope.mutate([{ op: "unset", path }]);
76
+ }
77
+
78
+ /* ── React component: collapsible card ───────────────────────────── */
79
+ function RateLimiterCard({ settingsScope }) {
80
+ const [cfg, setCfg] = react.useState(() => {
81
+ const snap = settingsScope.getSnapshot?.();
82
+ return snap?.status === "ready" && snap.value ? { ...snap.value } : {};
83
+ });
84
+
85
+ react.useEffect(() => {
86
+ const unsub = settingsScope.subscribe?.(() => {
87
+ const snap = settingsScope.getSnapshot();
88
+ if (snap?.status === "ready" && snap.value) setCfg({ ...snap.value });
89
+ });
90
+ return () => (typeof unsub === "function" ? unsub() : undefined);
91
+ }, [settingsScope]);
92
+
93
+ const modelKeys = Object.keys(cfg.models ?? {});
94
+
95
+ function addModel() {
96
+ const key = prompt('添加模型限速 (格式: provider/model)\n例如: deepseek/deepseek-chat');
97
+ if (!key) return;
98
+ setNested(settingsScope, ["models", key], {});
99
+ }
100
+ function removeModel(key) {
101
+ unsetNested(settingsScope, ["models", key]);
102
+ }
103
+ function setModelField(key, field, value) {
104
+ const path = ["models", key, field];
105
+ if (value === undefined || value === "") unsetNested(settingsScope, path);
106
+ else setNested(settingsScope, path, value);
107
+ }
108
+ function setDefault(field, value) {
109
+ setNested(settingsScope, ["defaults", field], value);
110
+ }
111
+ function setTop(field, value) {
112
+ settingsScope.set(field, value);
113
+ }
114
+
115
+ const [open, setOpen] = react.useState(false);
116
+
117
+ return jsxs("li", { className: "rlr-card" + (open ? " rlr-cardOpen" : ""), children: [
118
+ jsxs("button", {
119
+ type: "button",
120
+ className: "rlr-header",
121
+ "aria-expanded": open,
122
+ onClick: () => setOpen(!open),
123
+ children: [
124
+ jsxs("span", { className: "rlr-headText", children: [
125
+ jsx("span", { className: "rlr-name", children: "⚙ LLM 调用限速" }),
126
+ jsx("span", { className: "rlr-description", children: "按模型限制 LLM 请求频率,超出限制的请求排队等待" }),
127
+ ] }),
128
+ jsx("span", { className: "rlr-chevron" + (open ? " rlr-chevronOpen" : ""), children: "▾" }),
129
+ ]
130
+ }),
131
+ open && jsxs("div", { className: "rlr-body", children: [
132
+ /* ── 启用开关 ── */
133
+ jsxs("div", { className: "rlr-row", children: [
134
+ jsx("label", { children: [
135
+ jsx("input", { type: "checkbox", checked: cfg.enabled ?? true, onChange: (e) => setTop("enabled", e.target.checked) }),
136
+ " 启用限速",
137
+ ] }),
138
+ ] }),
139
+ /* ── 默认配置 ── */
140
+ jsxs("div", { className: "rlr-section", children: [
141
+ jsx("div", { className: "rlr-sectHead", children: "默认配置" }),
142
+ jsxs("div", { className: "rlr-row", children: [
143
+ jsx("span", { className: "rlr-label", children: "策略" }),
144
+ jsxs("select", { className: "rlr-select", value: cfg.strategy ?? "token-bucket", onChange: (e) => setTop("strategy", e.target.value), children: [
145
+ jsx("option", { value: "token-bucket", children: "令牌桶 (Token Bucket)" }),
146
+ jsx("option", { value: "sliding-window", children: "滑动窗口 (Sliding Window)" }),
147
+ ] }),
148
+ ] }),
149
+ jsxs("div", { className: "rlr-row", children: [
150
+ jsx("span", { className: "rlr-label", children: "最大并发数" }),
151
+ jsx("input", { type: "number", min: 1, className: "rlr-input", value: cfg.defaults?.maxConcurrent ?? 5, onChange: (e) => setDefault("maxConcurrent", Number(e.target.value) || 1) }),
152
+ ] }),
153
+ jsxs("div", { className: "rlr-row", children: [
154
+ jsx("span", { className: "rlr-label", children: "每分钟最大请求数" }),
155
+ jsx("input", { type: "number", min: 1, className: "rlr-input", value: cfg.defaults?.maxRpm ?? 60, onChange: (e) => setDefault("maxRpm", Number(e.target.value) || 1) }),
156
+ ] }),
157
+ cfg.strategy !== "sliding-window" && jsxs(react.Fragment, { children: [
158
+ jsxs("div", { className: "rlr-row", children: [
159
+ jsx("span", { className: "rlr-label", children: "突发容量" }),
160
+ jsx("input", { type: "number", min: 1, className: "rlr-input", value: cfg.defaults?.burstSize ?? 10, onChange: (e) => setDefault("burstSize", Number(e.target.value) || 1) }),
161
+ ] }),
162
+ jsxs("div", { className: "rlr-row", children: [
163
+ jsx("span", { className: "rlr-label", children: "补充速率 (个/秒)" }),
164
+ jsx("input", { type: "number", min: 0.1, step: 0.1, className: "rlr-input", value: cfg.defaults?.refillRate ?? 1, onChange: (e) => setDefault("refillRate", Number(e.target.value) || 0.1) }),
165
+ ] }),
166
+ ] }),
167
+ ] }),
168
+ /* ── 模型专属配置 ── */
169
+ jsxs("div", { className: "rlr-section", children: [
170
+ jsxs("div", { style: { display: "flex", justifyContent: "space-between", alignItems: "center" }, children: [
171
+ jsx("div", { className: "rlr-sectHead", children: "模型专属配置" }),
172
+ jsx("button", { className: "rlr-btn", onClick: addModel, children: "+ 添加模型" }),
173
+ ] }),
174
+ modelKeys.length === 0 && jsx("p", { className: "rlr-status", children: "未配置模型专属限速,所有模型使用默认值。" }),
175
+ modelKeys.map((key) => jsxs("div", { className: "rlr-modelRow", key, children: [
176
+ jsxs("div", { style: { display: "flex", flexDirection: "column", gap: 4, flex: 1 }, children: [
177
+ jsx("span", { className: "rlr-modelName", children: key }),
178
+ jsxs("div", { className: "rlr-modelMeta", children: [
179
+ jsx("span", { children: `并发: ${readNum(cfg, key, "maxConcurrent")}` }),
180
+ jsx("span", { children: `RPM: ${readNum(cfg, key, "maxRpm")}` }),
181
+ cfg.strategy !== "sliding-window" && jsx("span", { children: `突发: ${readNum(cfg, key, "burstSize")}` }),
182
+ ] }),
183
+ jsxs("div", { style: { display: "flex", alignItems: "center", gap: 6, marginTop: 4 }, children: [
184
+ jsx("input", { type: "number", min: 1, className: "rlr-input", style: { width: 60 }, title: "并发", placeholder: "并发", value: cfg.models?.[key]?.maxConcurrent ?? "", onChange: (e) => setModelField(key, "maxConcurrent", e.target.value === "" ? undefined : Number(e.target.value) || 1) }),
185
+ jsx("input", { type: "number", min: 1, className: "rlr-input", style: { width: 60 }, title: "RPM", placeholder: "RPM", value: cfg.models?.[key]?.maxRpm ?? "", onChange: (e) => setModelField(key, "maxRpm", e.target.value === "" ? undefined : Number(e.target.value) || 1) }),
186
+ cfg.strategy !== "sliding-window" && jsx("input", { type: "number", min: 1, className: "rlr-input", style: { width: 60 }, title: "突发", placeholder: "突发", value: cfg.models?.[key]?.burstSize ?? "", onChange: (e) => setModelField(key, "burstSize", e.target.value === "" ? undefined : Number(e.target.value) || 1) }),
187
+ jsx("button", { className: "rlr-btnSec", onClick: () => { const v = cfg.models?.[key]?.enabled; setModelField(key, "enabled", v === false ? undefined : false); }, children: cfg.models?.[key]?.enabled === false ? "已禁用" : "禁用" }),
188
+ ] }),
189
+ ] }),
190
+ jsx("button", { className: "rlr-danger", onClick: () => removeModel(key), title: "删除", children: "✕" }),
191
+ ] })),
192
+ ] }),
193
+ /* ── 被限速时的行为 ── */
194
+ jsxs("div", { className: "rlr-section", children: [
195
+ jsx("div", { className: "rlr-sectHead", children: "被限速时的行为" }),
196
+ jsxs("div", { className: "rlr-row", children: [
197
+ jsx("label", { style: { display: "flex", alignItems: "center", gap: 4, cursor: "pointer" }, children: [
198
+ jsx("input", { type: "radio", name: "rl-on-throttle", checked: cfg.onThrottled !== "reject", onChange: () => setTop("onThrottled", "queue") }),
199
+ "排队等待",
200
+ ] }),
201
+ jsx("label", { style: { display: "flex", alignItems: "center", gap: 4, cursor: "pointer" }, children: [
202
+ jsx("input", { type: "radio", name: "rl-on-throttle", checked: cfg.onThrottled === "reject", onChange: () => setTop("onThrottled", "reject") }),
203
+ "拒绝请求",
204
+ ] }),
205
+ ] }),
206
+ cfg.onThrottled !== "reject" && jsxs("div", { className: "rlr-row", children: [
207
+ jsx("span", { className: "rlr-label", children: "最大排队等待 (ms)" }),
208
+ jsx("input", { type: "number", min: 1000, step: 1000, className: "rlr-input", value: cfg.maxQueueWaitMs ?? 60000, onChange: (e) => setTop("maxQueueWaitMs", Number(e.target.value) || 60000) }),
209
+ ] }),
210
+ ] }),
211
+ ] }),
212
+ ] });
213
+ }
214
+
215
+ /* ── Plugin registration ── */
216
+ function apply(ctx) {
217
+ const scope = ctx.settingsScope.bind({ namespace: name });
218
+ ctx.slots.inject("settings.plugin.item", function* () {
219
+ yield ctx.slots.register({
220
+ name: "settings.plugin.item",
221
+ key: name,
222
+ locale: name,
223
+ inject: () => ({ settingsScope: scope }),
224
+ }, RateLimiterCard);
225
+ });
226
+ }
227
+
228
+ exports.apply = apply;
229
+ exports.inject = inject;
230
+ exports.name = name;
231
+ return module.exports;
232
+ }
233
+ });
package/lib/index.js ADDED
@@ -0,0 +1,203 @@
1
+ /**
2
+ * dsh-llm-rate-limiter — Host (server-side) entry point.
3
+ *
4
+ * Intercepts the `llm/stream` waterfall and enforces per-model rate limits.
5
+ * Configuration lives in the "llm-rate-limiter" settings namespace, editable
6
+ * from the browser GUI.
7
+ *
8
+ * The `llm/stream` waterfall signature is:
9
+ * (options: GenerateOptions, next: () => AsyncIterable<StreamChunk>) => AsyncIterable<StreamChunk>
10
+ *
11
+ * where `options` always has `.provider` (string) and `.model` (string).
12
+ *
13
+ * The settings service is injected dynamically (like dshmarket's
14
+ * installMarketSettings): `ctx.inject(['settings'], (scopedCtx) => ...)`, so a
15
+ * host without a settings provider simply runs without the GUI wiring — the
16
+ * composed entry stays as configured. Inherited from the DSH house pattern.
17
+ *
18
+ * @module dsh-llm-rate-limiter
19
+ */
20
+
21
+ import { TokenBucketStrategy } from "./strategies/token-bucket.js";
22
+ import { SlidingWindowStrategy } from "./strategies/sliding-window.js";
23
+ import { RateLimiterConfig } from "./types/config.js";
24
+
25
+ const name = "llm-rate-limiter";
26
+ const SETTINGS_NS = "llm-rate-limiter";
27
+
28
+ /* ── Helpers ─────────────────────────────────────────────────────────── */
29
+
30
+ /**
31
+ * Merge defaults + per-model overrides into a flat config object that both
32
+ * strategy constructors understand.
33
+ */
34
+ function resolveModelConfig(cfg, key) {
35
+ const o = cfg.models?.[key] ?? {};
36
+ const d = cfg.defaults ?? {};
37
+ const maxRpm = o.maxRpm ?? d.maxRpm ?? 60;
38
+ // When a model overrides maxRpm but not refillRate, derive refillRate from
39
+ // the model's maxRpm — this is the most common case ("set maxRpm=3").
40
+ // When model doesn't override maxRpm, use the defaults' refillRate as-is.
41
+ const refillRate = o.refillRate ?? (o.maxRpm != null ? o.maxRpm / 60 : d.refillRate ?? (maxRpm / 60));
42
+ return {
43
+ maxConcurrent: o.maxConcurrent ?? d.maxConcurrent ?? 5,
44
+ maxRpm,
45
+ burstSize: o.burstSize ?? d.burstSize ?? 10,
46
+ refillRate,
47
+ windowMs: d.windowMs ?? 60_000,
48
+ };
49
+ }
50
+
51
+ /**
52
+ * Ensure a rate-limiting strategy instance exists (or is updated) for `key`.
53
+ */
54
+ function ensureLimiter(limiters, key, cfg) {
55
+ const merged = resolveModelConfig(cfg, key);
56
+ const existing = limiters.get(key);
57
+ if (existing) {
58
+ existing.update(merged);
59
+ return existing;
60
+ }
61
+ const strategy = cfg.strategy === "sliding-window"
62
+ ? new SlidingWindowStrategy(merged)
63
+ : new TokenBucketStrategy(merged);
64
+ limiters.set(key, strategy);
65
+ return strategy;
66
+ }
67
+
68
+ /**
69
+ * Build a terminal StreamChunk for rate-limiting failures.
70
+ * Uses `kind: "aborted"` when the caller's AbortSignal fired, so the agent
71
+ * loop treats it as a cancellation rather than a retryable error.
72
+ * Otherwise uses `kind: "error"` with `RATE_LIMIT` code, which the
73
+ * `dsh-llm-retry` plugin can catch and back off.
74
+ */
75
+ function terminalChunk(modelKey, signal) {
76
+ const aborted = signal?.aborted === true;
77
+ return {
78
+ type: "finish",
79
+ reason: aborted
80
+ ? { kind: "aborted", failure: { code: "RATE_LIMIT", message: `Rate limit: ${modelKey} (aborted)` } }
81
+ : { kind: "error", failure: { code: "RATE_LIMIT", message: `Rate limit exceeded for ${modelKey}` } },
82
+ };
83
+ }
84
+
85
+ /* ── Plugin entry ────────────────────────────────────────────────────── */
86
+
87
+ /**
88
+ * @param {import("@deepseek-ai/cordis").Context} ctx
89
+ */
90
+ function apply(ctx) {
91
+ const logger = ctx.logger?.("rate-limiter");
92
+
93
+ // Shared mutable config state. Starts with built-in defaults; once the
94
+ // settings service is available the watcher replaces it with the live
95
+ // user-configured value on every change.
96
+ const defaultBase = {
97
+ enabled: true,
98
+ strategy: "token-bucket",
99
+ defaults: { maxConcurrent: 5, maxRpm: 60, burstSize: 10, refillRate: 1 },
100
+ models: {},
101
+ onThrottled: "queue",
102
+ maxQueueWaitMs: 60_000,
103
+ };
104
+ let cfg = { ...defaultBase };
105
+
106
+ /** @type {Map<string, TokenBucketStrategy | SlidingWindowStrategy>} */
107
+ const limiters = new Map();
108
+
109
+ // ── Settings wiring (dynamic inject) ───────────────────────
110
+ // Wrapped in ctx.inject so the plugin works on older DSH versions that
111
+ // don't have a settings service at all — the callback simply never runs.
112
+ ctx.inject(["settings"], (scopedCtx) => {
113
+ const scope = scopedCtx.settings.register(SETTINGS_NS, RateLimiterConfig, {
114
+ base: defaultBase,
115
+ });
116
+
117
+ logger?.info("rate limiter settings namespace registered (%s)", SETTINGS_NS);
118
+
119
+ // Replace cfg with live settings value on every change.
120
+ cfg = scope.get();
121
+ scope.watch((c) => { cfg = c; });
122
+
123
+ // Dispose limiters on settings teardown.
124
+ scopedCtx.effect(() => () => {
125
+ logger?.info("rate limiter settings scope tearing down (%s)", SETTINGS_NS);
126
+ for (const limiter of limiters.values()) limiter.dispose();
127
+ limiters.clear();
128
+ }, "rate-limiter: settings scope teardown");
129
+ });
130
+
131
+ // ── LLM stream interceptor ─────────────────────────────────
132
+ // ctx.on('llm/stream') doesn't need the settings service — it's a plain
133
+ // waterfall listener that reads the shared `cfg` variable (initialised
134
+ // with defaults; replaced by the settings watcher when available).
135
+ const disposeLlmListener = ctx.on("llm/stream", async function* rateLimitInterceptor(options, next) {
136
+ if (!cfg.enabled) {
137
+ yield* next();
138
+ return;
139
+ }
140
+
141
+ const modelKey = `${options.provider}/${options.model}`;
142
+
143
+ // Per-model opt-out: `enabled: false` skips this model entirely.
144
+ if (cfg.models?.[modelKey]?.enabled === false) {
145
+ yield* next();
146
+ return;
147
+ }
148
+
149
+ const limiter = ensureLimiter(limiters, modelKey, cfg);
150
+
151
+ // ── Phase 1: acquire ────────────────────────────────────────
152
+ let slotTaken = false;
153
+ if (cfg.onThrottled === "reject") {
154
+ // Non-blocking check — if at capacity, yield a terminal chunk immediately.
155
+ // acquireNonBlocking deducts a token AND takes a concurrent slot, so
156
+ // reject mode still enforces the frequency limit, not just concurrency.
157
+ const permit = limiter.acquireNonBlocking();
158
+ if (!permit.granted) {
159
+ logger?.warn("model %C rejected — rate limit reached", modelKey);
160
+ yield terminalChunk(modelKey, options.signal);
161
+ return;
162
+ }
163
+ // permit granted → slot already taken by acquireNonBlocking; skip acquireSlot.
164
+ slotTaken = true;
165
+ } else {
166
+ // Queue mode — park until a slot opens.
167
+ const permit = await limiter.acquire({
168
+ signal: options.signal,
169
+ timeoutMs: cfg.maxQueueWaitMs ?? 60_000,
170
+ });
171
+ if (!permit.granted) {
172
+ logger?.warn(
173
+ "model %C %s after %d ms",
174
+ modelKey, options.signal?.aborted ? "aborted" : "queue timeout", permit.waitMs,
175
+ );
176
+ yield terminalChunk(modelKey, options.signal);
177
+ return;
178
+ }
179
+ // Queue mode: acquire() does NOT take a concurrent slot.
180
+ // We must call acquireSlot() before the actual call.
181
+ slotTaken = false;
182
+ }
183
+
184
+ // ── Phase 2: run the real LLM call ─────────────────────────
185
+ if (!slotTaken) limiter.acquireSlot();
186
+ try {
187
+ yield* next();
188
+ } finally {
189
+ limiter.releaseSlot();
190
+ }
191
+ });
192
+
193
+ // ── Teardown ──
194
+ ctx.effect(() => () => {
195
+ disposeLlmListener();
196
+ for (const limiter of limiters.values()) limiter.dispose();
197
+ limiters.clear();
198
+ }, "rate-limiter: dispose llm/stream listener");
199
+
200
+ logger?.info("rate limiter active — llm/stream interceptor registered");
201
+ }
202
+
203
+ export { apply, name, SETTINGS_NS };
@@ -0,0 +1,148 @@
1
+ /**
2
+ * Sliding-Window rate limiter strategy.
3
+ *
4
+ * Counts requests inside a rolling window of `windowMs` milliseconds.
5
+ * The window slides continuously: each request's timestamp is recorded,
6
+ * and the count is the number of timestamps inside the window.
7
+ *
8
+ * Concurrency is controlled separately (same as TokenBucketStrategy).
9
+ *
10
+ * @module dsh-llm-rate-limiter/strategies/sliding-window
11
+ */
12
+
13
+ export class SlidingWindowStrategy {
14
+ /** @type {{ windowMs: number, maxRpm: number, maxConcurrent: number }} */
15
+ #cfg;
16
+ #concurrent = 0;
17
+ /** @type {number[]} — timestamps (Date.now()) inside the window */
18
+ #timestamps = [];
19
+ /** @type {Array<{ resolve: (r: { granted: boolean }) => void, abort: () => void }>} */
20
+ #waiters = [];
21
+ #timer;
22
+
23
+ /**
24
+ * @param {object} config — `{ windowMs, maxRpm, maxConcurrent }`
25
+ */
26
+ constructor(config) {
27
+ this.#cfg = { ...config };
28
+ // Prune stale timestamps every second.
29
+ this.#timer = setInterval(() => this.#prune(), 1_000);
30
+ }
31
+
32
+ /* ── public API ─────────────────────────────────────── */
33
+
34
+ acquire({ signal, timeoutMs }) {
35
+ const t0 = Date.now();
36
+
37
+ if (signal?.aborted) return Promise.resolve({ granted: false });
38
+
39
+ // Fast path: concurrency AND window capacity both available.
40
+ if (this.#concurrent < this.#cfg.maxConcurrent && this.#count() < this.#cfg.maxRpm) {
41
+ this.#timestamps.push(Date.now());
42
+ return Promise.resolve({ granted: true });
43
+ }
44
+
45
+ if (this.#cfg.maxConcurrent <= 0) return Promise.resolve({ granted: false });
46
+
47
+ return new Promise((resolve) => {
48
+ const onDone = (result) => {
49
+ clearTimeout(timer);
50
+ signal?.removeEventListener("abort", onAbort);
51
+ const idx = this.#waiters.indexOf(entry);
52
+ if (idx >= 0) this.#waiters.splice(idx, 1);
53
+ resolve({ ...result, waitMs: Date.now() - t0 });
54
+ };
55
+
56
+ const onAbort = () => onDone({ granted: false });
57
+ const timer = setTimeout(() => onDone({ granted: false }), timeoutMs);
58
+
59
+ const entry = {
60
+ resolve: () => onDone({ granted: true }),
61
+ abort: onAbort,
62
+ };
63
+
64
+ this.#waiters.push(entry);
65
+ signal?.addEventListener("abort", onAbort, { once: true });
66
+ });
67
+ }
68
+
69
+ /**
70
+ * Non-blocking acquisition: record the request in the window and take an
71
+ * implicit concurrent slot if both are available right now, without parking
72
+ * in the queue.
73
+ *
74
+ * Used by "reject" throttle mode, which must not wait, but must still count
75
+ * the request so the rate limit is enforced.
76
+ *
77
+ * @returns {{ granted: boolean }} — granted means the request was counted
78
+ * and a concurrent slot taken; the caller MUST pair it with `releaseSlot()`.
79
+ */
80
+ acquireNonBlocking() {
81
+ if (this.#concurrent < this.#cfg.maxConcurrent && this.#count() < this.#cfg.maxRpm) {
82
+ this.#timestamps.push(Date.now());
83
+ this.#concurrent += 1;
84
+ return { granted: true };
85
+ }
86
+ return { granted: false };
87
+ }
88
+
89
+ acquireSlot() {
90
+ this.#concurrent += 1;
91
+ }
92
+
93
+ releaseSlot() {
94
+ this.#concurrent = Math.max(0, this.#concurrent - 1);
95
+ this.#drain();
96
+ }
97
+
98
+ getStatus() {
99
+ return {
100
+ strategy: "sliding-window",
101
+ windowMs: this.#cfg.windowMs,
102
+ countInWindow: this.#count(),
103
+ maxRpm: this.#cfg.maxRpm,
104
+ concurrent: this.#concurrent,
105
+ maxConcurrent: this.#cfg.maxConcurrent,
106
+ queued: this.#waiters.length,
107
+ };
108
+ }
109
+
110
+ update(config) {
111
+ Object.assign(this.#cfg, config);
112
+ }
113
+
114
+ dispose() {
115
+ clearInterval(this.#timer);
116
+ for (const w of this.#waiters) w.abort();
117
+ this.#waiters = [];
118
+ }
119
+
120
+ /* ── internals ─────────────────────────────────────── */
121
+
122
+ /** Number of timestamps inside the current window. */
123
+ #count() {
124
+ const cutoff = Date.now() - this.#cfg.windowMs;
125
+ // timestamps are pushed in order, so find the first that's still inside.
126
+ let i = 0;
127
+ while (i < this.#timestamps.length && this.#timestamps[i] <= cutoff) i += 1;
128
+ return this.#timestamps.length - i;
129
+ }
130
+
131
+ /** Remove timestamps that fell out of the window. */
132
+ #prune() {
133
+ const cutoff = Date.now() - this.#cfg.windowMs;
134
+ while (this.#timestamps.length > 0 && this.#timestamps[0] <= cutoff) {
135
+ this.#timestamps.shift();
136
+ }
137
+ this.#drain();
138
+ }
139
+
140
+ /** Drain waiters as long as capacity is available. */
141
+ #drain() {
142
+ while (this.#waiters.length > 0 && this.#count() < this.#cfg.maxRpm && this.#concurrent < this.#cfg.maxConcurrent) {
143
+ this.#timestamps.push(Date.now());
144
+ const waiter = this.#waiters.shift();
145
+ waiter?.resolve();
146
+ }
147
+ }
148
+ }
@@ -0,0 +1,168 @@
1
+ /**
2
+ * Token-Bucket rate limiter strategy.
3
+ *
4
+ * Each model gets its own bucket. A bucket has:
5
+ * - `burstSize` — maximum tokens the bucket can hold
6
+ * - `refillRate` — tokens added per second
7
+ * - a concurrent-slot counter (separate from tokens, so burst + concurrency are orthogonal)
8
+ *
9
+ * When the bucket is empty the request is queued (a promise resolver parked in
10
+ * `waiters[]`). As tokens refill, waiters are drained in FIFO order.
11
+ *
12
+ * @module dsh-llm-rate-limiter/strategies/token-bucket
13
+ */
14
+
15
+ export class TokenBucketStrategy {
16
+ /** @type {{ burstSize: number, refillRate: number, maxConcurrent: number }} */
17
+ #cfg;
18
+ #tokens;
19
+ #lastRefill;
20
+ #concurrent = 0;
21
+ /** @type {Array<{ resolve: (r: { granted: boolean }) => void, abort: () => void }>} */
22
+ #waiters = [];
23
+ #timer;
24
+
25
+ /**
26
+ * @param {object} config — `{ burstSize, refillRate, maxConcurrent }`
27
+ */
28
+ constructor(config) {
29
+ this.#cfg = { ...config };
30
+ this.#tokens = config.burstSize;
31
+ this.#lastRefill = Date.now();
32
+ this.#startRefill();
33
+ }
34
+
35
+ /* ── public API ─────────────────────────────────────── */
36
+
37
+ /**
38
+ * Wait until a request is allowed.
39
+ * Resolves `{ granted: true }` immediately or after the queue drains.
40
+ * Resolves `{ granted: false }` on abort or timeout (timeout is the caller's
41
+ * responsibility — it races this promise against a timer).
42
+ *
43
+ * @param {{ signal?: AbortSignal, timeoutMs: number }} opts
44
+ * @returns {Promise<{ granted: boolean, waitMs?: number }>}
45
+ */
46
+ acquire({ signal, timeoutMs }) {
47
+ const t0 = Date.now();
48
+
49
+ // Already aborted — fail fast.
50
+ if (signal?.aborted) {
51
+ return Promise.resolve({ granted: false });
52
+ }
53
+
54
+ // Fast path: concurrency room AND a token available.
55
+ if (this.#concurrent < this.#cfg.maxConcurrent && this.#tokens >= 1) {
56
+ this.#tokens -= 1;
57
+ return Promise.resolve({ granted: true });
58
+ }
59
+
60
+ // Slow path: park in the waiter queue.
61
+ if (this.#cfg.maxConcurrent <= 0) {
62
+ return Promise.resolve({ granted: false });
63
+ }
64
+
65
+ return new Promise((resolve) => {
66
+ const onDone = (result) => {
67
+ clearTimeout(timer);
68
+ signal?.removeEventListener("abort", onAbort);
69
+ const idx = this.#waiters.indexOf(entry);
70
+ if (idx >= 0) this.#waiters.splice(idx, 1);
71
+ resolve({ ...result, waitMs: Date.now() - t0 });
72
+ };
73
+
74
+ const onAbort = () => onDone({ granted: false });
75
+
76
+ const timer = setTimeout(() => onDone({ granted: false }), timeoutMs);
77
+
78
+ const entry = {
79
+ resolve: () => {
80
+ // Token already deducted by _drain.
81
+ onDone({ granted: true });
82
+ },
83
+ abort: onAbort,
84
+ };
85
+
86
+ this.#waiters.push(entry);
87
+ signal?.addEventListener("abort", onAbort, { once: true });
88
+ });
89
+ }
90
+
91
+ /**
92
+ * Non-blocking acquisition: consume a token and an implicit concurrent slot
93
+ * if both are available right now, without parking in the queue.
94
+ *
95
+ * Used by "reject" throttle mode, which must not wait, but must still deduct
96
+ * tokens so the rate limit is enforced (otherwise only the concurrency cap
97
+ * would bite).
98
+ *
99
+ * @returns {{ granted: boolean }} — granted means a token was consumed and a
100
+ * concurrent slot taken; the caller MUST pair it with `releaseSlot()`.
101
+ */
102
+ acquireNonBlocking() {
103
+ if (this.#concurrent < this.#cfg.maxConcurrent && this.#tokens >= 1) {
104
+ this.#tokens -= 1;
105
+ this.#concurrent += 1;
106
+ return { granted: true };
107
+ }
108
+ return { granted: false };
109
+ }
110
+
111
+ /** Mark one concurrent slot as in-use. */
112
+ acquireSlot() {
113
+ this.#concurrent += 1;
114
+ }
115
+
116
+ /** Release one concurrent slot and try to drain waiters. */
117
+ releaseSlot() {
118
+ this.#concurrent = Math.max(0, this.#concurrent - 1);
119
+ this.#drain();
120
+ }
121
+
122
+ /** Live snapshot for the GUI. */
123
+ getStatus() {
124
+ return {
125
+ strategy: "token-bucket",
126
+ tokens: Math.round(this.#tokens * 100) / 100,
127
+ burstSize: this.#cfg.burstSize,
128
+ refillRate: this.#cfg.refillRate,
129
+ concurrent: this.#concurrent,
130
+ maxConcurrent: this.#cfg.maxConcurrent,
131
+ queued: this.#waiters.length,
132
+ };
133
+ }
134
+
135
+ /** Replace config at runtime (hot-reload). */
136
+ update(config) {
137
+ Object.assign(this.#cfg, config);
138
+ }
139
+
140
+ /** Dispose — release all waiters and stop the refill timer. */
141
+ dispose() {
142
+ clearInterval(this.#timer);
143
+ for (const w of this.#waiters) w.abort();
144
+ this.#waiters = [];
145
+ }
146
+
147
+ /* ── internals ─────────────────────────────────────── */
148
+
149
+ #startRefill() {
150
+ // Refill every 50 ms for smooth granularity.
151
+ this.#timer = setInterval(() => {
152
+ const now = Date.now();
153
+ const elapsed = (now - this.#lastRefill) / 1000;
154
+ this.#tokens = Math.min(this.#cfg.burstSize, this.#tokens + elapsed * this.#cfg.refillRate);
155
+ this.#lastRefill = now;
156
+ this.#drain();
157
+ }, 50);
158
+ }
159
+
160
+ /** Drain waiters as long as we have tokens and concurrency room. */
161
+ #drain() {
162
+ while (this.#waiters.length > 0 && this.#tokens >= 1 && this.#concurrent < this.#cfg.maxConcurrent) {
163
+ this.#tokens -= 1;
164
+ const waiter = this.#waiters.shift();
165
+ waiter?.resolve();
166
+ }
167
+ }
168
+ }
@@ -0,0 +1,53 @@
1
+ /**
2
+ * Rate limiter configuration schema.
3
+ *
4
+ * The settings namespace is "llm-rate-limiter". Users override values through
5
+ * the DSH settings document; the host merges them on top of `base`.
6
+ *
7
+ * @module dsh-llm-rate-limiter/types/config
8
+ */
9
+
10
+ import Schema from "@deepseek-ai/schemastery";
11
+
12
+ /**
13
+ * Per-model overrides — key is `"provider/model"` (e.g. `"deepseek/deepseek-chat"`).
14
+ * Every field is optional; omitted fields inherit from `defaults`.
15
+ */
16
+ const ModelOverride = Schema.object({
17
+ maxConcurrent: Schema.number().step(1).min(1).description("Max concurrent requests for this model"),
18
+ maxRpm: Schema.number().step(1).min(1).description("Max requests per minute for this model"),
19
+ burstSize: Schema.number().step(1).min(1).description("Burst capacity for this model (token-bucket)"),
20
+ refillRate: Schema.number().min(0.1).description("Tokens refilled per second (token-bucket)"),
21
+ enabled: Schema.boolean().description("Enable/disable rate limiting for this model"),
22
+ });
23
+
24
+ export const RateLimiterConfig = Schema.object({
25
+ /** Master switch */
26
+ enabled: Schema.boolean().default(true).description("Enable the rate limiter globally"),
27
+
28
+ /** Default strategy */
29
+ strategy: Schema.union([
30
+ Schema.const("token-bucket").description("Token Bucket — allows bursts up to burstSize"),
31
+ Schema.const("sliding-window").description("Sliding Window — smooth, fixed requests-per-minute"),
32
+ ]).default("token-bucket").description("Rate limiting algorithm"),
33
+
34
+ /** Global defaults — applied to every model unless overridden */
35
+ defaults: Schema.object({
36
+ maxConcurrent: Schema.number().step(1).min(1).default(5).description("Max concurrent requests"),
37
+ maxRpm: Schema.number().step(1).min(1).default(60).description("Max requests per minute"),
38
+ burstSize: Schema.number().step(1).min(1).default(10).description("Burst capacity (token-bucket)"),
39
+ refillRate: Schema.number().min(0.1).default(1).description("Tokens per second (token-bucket)"),
40
+ }).default({}).description("Default limits applied to all models"),
41
+
42
+ /** Per-model overrides — key is "provider/model" */
43
+ models: Schema.dict(ModelOverride).description("Per-model rate limit overrides"),
44
+
45
+ /** Queue behaviour when a request is throttled */
46
+ onThrottled: Schema.union([
47
+ Schema.const("queue").description("Wait in queue until a slot opens"),
48
+ Schema.const("reject").description("Immediately return an error"),
49
+ ]).default("queue").description("What happens when a request exceeds the limit"),
50
+
51
+ /** Maximum time a request waits in queue before being rejected */
52
+ maxQueueWaitMs: Schema.number().step(1).min(1000).default(60_000).description("Max queue wait (ms)"),
53
+ }).description("LLM call rate limiter settings");
@@ -0,0 +1 @@
1
+ export { RateLimiterConfig } from "./config.js";
package/package.json ADDED
@@ -0,0 +1,60 @@
1
+ {
2
+ "name": "@leaf233/dsh-llm-rate-limiter",
3
+ "version": "0.1.1",
4
+ "description": "Per-model LLM call rate limiter for DeepSeek Harness with queue support",
5
+ "type": "module",
6
+ "main": "lib/index.js",
7
+ "exports": {
8
+ ".": "./lib/index.js",
9
+ "./client": "./lib/client.js"
10
+ },
11
+ "dsh": {
12
+ "bundle": {
13
+ "patch": "./cordis.patch.yml"
14
+ },
15
+ "client": {
16
+ "inject": [
17
+ "@deepseek-ai/dsh-client-ui-settings"
18
+ ],
19
+ "platform": "web"
20
+ }
21
+ },
22
+ "scripts": {
23
+ "build": "echo 'no build needed'",
24
+ "test": "node test-strategies.mjs"
25
+ },
26
+ "repository": {
27
+ "type": "git",
28
+ "url": "https://github.com/Leafyezi233/dsh-llm-rate-limiter.git"
29
+ },
30
+ "bugs": {
31
+ "url": "https://github.com/Leafyezi233/dsh-llm-rate-limiter/issues"
32
+ },
33
+ "homepage": "https://github.com/Leafyezi233/dsh-llm-rate-limiter#readme",
34
+ "keywords": [
35
+ "dsh",
36
+ "deepseek-harness",
37
+ "rate-limiter",
38
+ "llm",
39
+ "token-bucket",
40
+ "sliding-window",
41
+ "cordis-plugin"
42
+ ],
43
+ "license": "MIT",
44
+ "peerDependencies": {
45
+ "@deepseek-ai/cordis": ">=4.0.0",
46
+ "@deepseek-ai/schemastery": ">=3.18.0"
47
+ },
48
+ "dependencies": {
49
+ "@deepseek-ai/schemastery": "^3.18.0"
50
+ },
51
+ "files": [
52
+ "lib/",
53
+ "cordis.patch.yml",
54
+ "package.json",
55
+ "README.md",
56
+ "CHANGELOG.md",
57
+ "COMPATIBILITY.md",
58
+ "LICENSE"
59
+ ]
60
+ }