@goodandready/dsh-cron 0.2.35 → 0.2.37

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -186,8 +186,11 @@ Every task picks its own runtime; non-LLM runtimes need no model and consume no
186
186
  ### 7. Cost Control: Fallback Model & Burn Guard
187
187
  * **Fallback Model** — A task can run on the cheap model by default and still finish on the strong one: set `fallbackModel` (and optionally `fallbackProvider`) and a failed run — `error` or `timeout` — is retried **once** on that model before the ordinary retry backoff applies. History records which model produced the result and whether the fallback was used, usage and cost of both attempts are summed, and the `{model}` template variable renders the model that finished the run. Only agent-mediated tasks (`llm`, `skill`, `workflow`) can use a fallback.
188
188
  * **Token & Cost Burn Guard** — Prevent runaway spending by configuring per-task limits: `costLimitUsd` (lifetime spend limit in USD), `dailyCostLimitUsd` (rolling 24-hour spend limit in USD), and `tokenLimit` (lifetime token limit). If a task exceeds any threshold, execution is halted, the task is automatically paused with `pausedReason` (`cost_limit_exceeded`, `daily_cost_limit_exceeded`, or `token_limit_exceeded`), and an alert notification is dispatched across all active channels.
189
+ * **Event-Driven Token & Cost Extraction** — Reads real token usage from DSH session stream events (`assistant/message`, `assistant/chunk` usage, and `assistant/attempt`), accounting for uncached input, cached reads, output tokens, and paid failed attempts across retries.
190
+ * **Rolling 24-Hour Cost Ledger** — Daily burn guard (`dailyCostLimitUsd`) maintains an independent rolling 24h cost ledger per task, preserved in `store.json`. Expenses remain fully counted across history archive rotations (beyond 100 runs) and survive daemon restarts.
189
191
 
190
192
  ### 8. Session Integration & Permissions
193
+ * **Real Assistant Output & Terminal Status Extraction** — Isolates current turn session events, extracts the actual assistant message text (excluding previous turn history in persistent sessions), and validates terminal turn status (`turn/end` errors or interruptions) to guarantee accurate reporting, chaining, and structured action execution.
191
194
  * **Per-task permission presets** — `default`, `read-only`, `workspace-write`, or `full` are applied to the task's agent session before the prompt runs.
192
195
  * **Session auto-archive** — isolated cron sessions are archived after each run (best-effort) so they do not clutter the chat list.
193
196
  * **History → session navigation** — every LLM run records its session; open it straight from the run history entry.
@@ -234,12 +237,15 @@ Prevent rogue processes from stacking concurrent duplicate executions:
234
237
  * **`skip`** (default): drops the overlapping run and records a `skipped` entry in the run history.
235
238
  * **`queue`**: queues the next execution and starts it as soon as the active job completes.
236
239
  * **`replace`**: aborts the active run via `AbortController` and launches a fresh execution.
240
+ * **Croner Overlap Policy Delegation** — Scheduled cron ticks fire without upstream suppression (`Croner protect: false`), allowing `skip` (with history logging), `queue` (delayed execution), and `replace` (clean abort) to govern recurring cron ticks and manual triggers consistently.
241
+ * **Context-Preserving Queues** — Both the global concurrency queue and task overlap queue retain full immutable execution context (`chainDepth`, `prevOutput`, `prevTaskId`, `prevStatus`, `prevCostUsd`), preventing data loss under load.
237
242
 
238
243
  If the daemon was offline at a scheduled time, the run is recorded as `missed` on startup, so gaps in the history stay visible.
239
244
 
240
245
  ### 14. Heartbeat Monitoring (#16-style dead man's switch)
241
246
  * Set `heartbeatUrl` and `heartbeatIntervalSec` in the plugin settings and the scheduler pings that URL on schedule — an external monitor alerts when the pings stop.
242
247
  * A built-in `GET /dsh-cron/heartbeat` endpoint reports liveness, active task count and the last run time for your own watchdogs.
248
+ * **Missed Heartbeat Alerts in `onlyOnFailure` Mode** — Missed heartbeats (`status === 'missed'`) are recognized as active failure events across all channel delivery predicates (`channels.js`, `telegram.js`, `integrations.js`), guaranteeing that Telegram, Discord, Webhook, and Kanban integrations alert immediately when an external process fails to report.
243
249
 
244
250
  ### 15. Declarative Jobs From the Profile Config (#50)
245
251
  Long-lived operational jobs can be declared in the profile configuration instead of being recreated by hand in the UI. The config file owns the jobs it declares: at every plugin start they are created or updated, and a job that disappears from the file is removed.
@@ -289,6 +295,7 @@ Set the token as the plugin setting `apiToken` (masked like every secret). Auth
289
295
  | `POST` | `/dsh-cron/api/tasks/:id/run` | Force an immediate run |
290
296
 
291
297
  The operations reuse the panel handlers, so the `x-dsh-cron-confirm: script` gate for code-executing types and the `409` refusals for config-owned tasks behave exactly as in the UI.
298
+ * **Bearer-Guarded Remote Access** — Remote clients can create and update tasks, ping heartbeats (`/dsh-cron/api/heartbeat/:id`), and preview schedules (`/dsh-cron/api/schedule/preview`) with valid `Authorization: Bearer <token>` credentials without loopback IP restrictions.
292
299
 
293
300
  ```bash
294
301
  BASE="http://127.0.0.1:3080"
@@ -382,6 +389,8 @@ Developer-facing, no behaviour change. `parseScheduleExpression` was split into
382
389
  - **Priority Queues & Concurrency Pools**: When the concurrent run limit is reached, queued tasks are ordered by `priority` (1 = highest, 10 = lowest) to ensure critical system alerts execute ahead of bulk background jobs.
383
390
  - **Self-Healing Runbooks & Auto-Diagnosis**: Failed tasks automatically execute an optional compensatory `selfHealingCommand` (e.g. system service restart or temp cleanup). Model failures can trigger `autoDiagnose: true` to append an instant root-cause diagnosis.
384
391
  - **Web UI Archive & Pipeline Visualization**: Interactive archive drawer with pagination and full log viewing; visual indicators for `➜ onSuccess` and `↳ onFailure` task connections.
392
+ - **Shell & Disk Pre-flight Gate Hardening**: Supports both `command` and `shell` types with fail-closed semantics for invalid syntax, unknown types, or inspection failures. Disk space checks support exact path and threshold syntax (e.g. `/:10%`, `/data:500MB`).
393
+ - **Complete Reliability Field Persistence**: HTTP POST task creation/editing and `cron_create_task` tool fully preserve and store all 13 reliability and session configuration fields (`agentPreset`, `targetSessionId`, `targetSessionReset`, `onSuccess`, `onFailure`, `heartbeatIntervalSeconds`, `gracePeriodSeconds`, `preflightType`, `preflightTarget`, `priority`, `concurrencyGroup`, `selfHealingCommand`, `autoDiagnose`, `fallbackProvider`, `fallbackModel`, `silentRule`, `inspectOnFailure`).
385
394
 
386
395
  ---
387
396
 
package/README.ru.md CHANGED
@@ -186,8 +186,11 @@ cron({
186
186
  ### 7. Экономия: fallback-модель и защита бюджета (Burn Guard)
187
187
  * **Fallback-модель** — задача может идти на дешёвой модели по умолчанию и всё же завершиться на сильной: задайте `fallbackModel` (и при необходимости `fallbackProvider`), и сбойный запуск (`error` или `timeout`) один раз повторится на этой модели, прежде чем включится обычный retry с задержкой. В истории видно, какая модель произвела результат и был ли использован fallback; расход и стоимость обеих попыток суммируются; переменная шаблона `{model}` подставляет модель, завершившую запуск. Fallback доступен только агентским типам (`llm`, `skill`, `workflow`).
188
188
  * **Защита бюджета токенов и расходов (Burn Guard)** — предотвращение неконтролируемых трат через индивидуальные лимиты задачи: `costLimitUsd` (общий лимит расходов в USD), `dailyCostLimitUsd` (суточный лимит за последние 24 часа в USD) и `tokenLimit` (лимит суммарных токенов). При превышении любого порога выполнение прекращается, задача автоматически переводится в паузу с фиксацией `pausedReason` (`cost_limit_exceeded`, `daily_cost_limit_exceeded`, `token_limit_exceeded`), а во все активные каналы отправляется тревожное оповещение.
189
+ * **Учёт токенов и расходов на событиях сессии** — считывание фактического расхода токенов напрямую из событий потока сессии DSH (`assistant/message`, `assistant/chunk` с типом usage, `assistant/attempt`), с разделением некэшированного ввода, чтений кэша, генерации и оплаченных сбойных попыток провайдера при повторах.
190
+ * **Скользящий суточный реестр (24h Cost Ledger)** — суточный лимит `dailyCostLimitUsd` опирается на независимый реестр расходов задачи за последние 24 часа, сохраняемый в `store.json`. Расходы сохраняются при ротации архива истории (свыше 100 запусков) и восстанавливаются после перезапуска службы.
189
191
 
190
192
  ### 8. Интеграция сессий и права
193
+ * **Извлечение реального ответа ассистента и статуса хода** — выделение событий только текущего хода (без подмешивания истории прошлых ходов долговременных сессий), сборка итогового текста сообщения ассистента и валидация признака завершения (`turn/end` с ошибкой или прерыванием) для безошибочной классификации сбоев, цепочек задач и выполнения директив.
191
194
  * **Permission-пресеты на задачу** — `default`, `read-only`, `workspace-write` или `full` применяются к сессии агента перед запуском промпта.
192
195
  * **Автоархивация сессий** — изолированные cron-сессии архивируются после запуска (best-effort), не засоряя список чатов.
193
196
  * **История → сессия** — каждый LLM-запуск хранит свою сессию; открыть диалог можно прямо из записи истории.
@@ -226,6 +229,8 @@ cron({
226
229
  * **`skip`** (по умолчанию): накладывающийся запуск отбрасывается, в истории появляется запись `skipped`;
227
230
  * **`queue`**: следующий запуск ставится в очередь и стартует по завершении активного;
228
231
  * **`replace`**: активный запуск прерывается через `AbortController`, запускается свежий.
232
+ * **Прямая передача тиков расписания политикам наложения** — тики Croner поступают без подавления на уровне планировщика (`Croner protect: false`), гарантируя одинаковую работу политик `skip` (с записью в историю), `queue` (отложенный запуск) и `replace` (чистый abort) как для cron-расписания, так и для ручных вызовов.
233
+ * **Сохранение контекста очереди** — глобальная очередь конкурентности и очередь наложения сохраняют полный контекст выполнения (`chainDepth`, `prevOutput`, `prevTaskId`, `prevStatus`, `prevCostUsd`), предотвращая потерю параметров цепочек под нагрузкой.
229
234
 
230
235
  Если сервис был выключен в момент планового запуска, при старте в истории появится запись `missed` — пробелы в истории остаются видимыми.
231
236
 
@@ -251,6 +256,8 @@ cron({
251
256
  - **Очереди с приоритетами**: При достижении лимита параллелизма задачи упорядочиваются по полю `priority` (1 — наивысший, 10 — низший).
252
257
  - **Команды самоисцеления и авто-диагностика (Self-Healing)**: Автоматический запуск компенсирующей команды `selfHealingCommand` при падении задачи; опция `autoDiagnose` для генерации AI-диагностики причин сбоя.
253
258
  - **Интерактивный архив логов в UI**: Модальное окно просмотра истории с пагинацией и полным выводом логов, визуальные ссылки конвейеров `➜ onSuccess` и `↳ onFailure`.
259
+ - **Безопасные preflight-проверки shell и диска**: Поддержка типов проверок `command` и `shell` с принципом fail-closed при синтаксических ошибках или сбоях проверки. Проверка диска поддерживает явные пути и пороговые значения (например, `/:10%`, `/data:500MB`).
260
+ - **Полное сохранение параметров надёжности**: Создание и редактирование задач через HTTP POST и агентский инструмент `cron_create_task` полностью сохраняют все 13 полей надёжности и управления сессиями (`agentPreset`, `targetSessionId`, `targetSessionReset`, `onSuccess`, `onFailure`, `heartbeatIntervalSeconds`, `gracePeriodSeconds`, `preflightType`, `preflightTarget`, `priority`, `concurrencyGroup`, `selfHealingCommand`, `autoDiagnose`, `fallbackProvider`, `fallbackModel`, `silentRule`, `inspectOnFailure`).
254
261
 
255
262
  ---
256
263
 
@@ -373,6 +380,7 @@ dsh-cron:
373
380
  ### 14. Мониторинг heartbeat (dead man's switch)
374
381
  * Задайте `heartbeatUrl` и `heartbeatIntervalSec` в настройках плагина — планировщик будет пинговать этот адрес по расписанию, и внешний монитор сообщит, когда пинги прекратятся.
375
382
  * Встроенный эндпоинт `GET /dsh-cron/heartbeat` сообщает живость, число активных задач и время последнего запуска для ваших собственных сторожей.
383
+ * **Оповещения о пропущенных heartbeat в режиме `onlyOnFailure`** — Пропущенные heartbeat (`status === 'missed'`) распознаются как полноценное состояние сбоя во всех фильтрах отправки (`channels.js`, `telegram.js`, `integrations.js`), гарантируя своевременную отправку тревог в Telegram, Discord, Webhook и Kanban при отсутствии отчётов от внешних служб.
376
384
 
377
385
  ### 15. Задачи из конфига профиля (#50)
378
386
  Долгоживущие эксплуатационные задачи можно объявлять в конфиге профиля, а не пересоздавать руками в интерфейсе. Владелец объявленных задач — файл конфига: при каждом старте плагина они создаются или обновляются, а задача, исчезнувшая из файла, удаляется.
@@ -422,6 +430,7 @@ dsh-cron:
422
430
  | `POST` | `/dsh-cron/api/tasks/:id/run` | Принудительный немедленный запуск |
423
431
 
424
432
  Операции переиспользуют обработчики панели, поэтому гейт `x-dsh-cron-confirm: script` для код-исполняющих типов и отказ `409` для конфиг-задач действуют здесь так же, как в UI.
433
+ * **Удалённый доступ по Bearer-токену** — Внешние клиенты с любого разрешённого IP могут создавать и обновлять задачи, отправлять heartbeat-сигналы (`/dsh-cron/api/heartbeat/:id`) и запрашивать предпросмотр расписания (`/dsh-cron/api/schedule/preview`) с валидным `Authorization: Bearer <token>`.
425
434
 
426
435
  ```bash
427
436
  BASE="http://127.0.0.1:3080"
package/README.zh.md CHANGED
@@ -186,8 +186,11 @@ cron({
186
186
  ### 7. 成本控制:回退模型与支出保护(Burn Guard)
187
187
  * **回退模型** —— 任务可以默认使用便宜模型,失败时改用更强模型完成:设置 `fallbackModel`(可选 `fallbackProvider`),失败(`error` 或 `timeout`)的运行会在该模型上重试一次,之后才进入常规重试退避。历史记录会标明最终产出结果的模型以及是否使用了回退,两次尝试的用量与成本都会累计,模板变量 `{model}` 渲染完成运行的模型。回退仅适用于智能体类型(`llm`、`skill`、`workflow`)。
188
188
  * **Token 与成本支出保护(Burn Guard)** —— 为任务配置严格预算上限:`costLimitUsd`(总支出美元上限)、`dailyCostLimitUsd`(24小时滚动支出上限)和 `tokenLimit`(Token总数上限)。一旦达到任一阈值,任务将自动暂停并记录 `pausedReason`(`cost_limit_exceeded`、`daily_cost_limit_exceeded` 或 `token_limit_exceeded`),同时向所有配置的通知渠道发送报警通知。
189
+ * **基于会话事件的 Token 与成本统计** —— 直接从 DSH 会话流式事件(`assistant/message`、`assistant/chunk` usage、`assistant/attempt`)中提取实际 Token 消耗,精确计入未命中输入、命中缓存读取、生成输出及重试过程中已计费的失败尝试。
190
+ * **滑动 24 小时成本账本(24h Cost Ledger)** —— 日耗保护(`dailyCostLimitUsd`)在 `store.json` 中为每个任务维护独立的滑动 24 小时成本记录。即使历史记录超过 100 条触发归档,24 小时内的所有花费依然完整保留并能跨服务重启持续生效。
189
191
 
190
192
  ### 8. 会话集成与权限
193
+ * **真实智能体输出与终端状态捕获** —— 隔离当前轮次的会话事件,提取最终真实的智能体回复文本(持久会话中自动排除以往历史轮次),并严格校验轮次终止状态(带有错误或中断的 `turn/end`),确保任务失败能准确反映到执行历史、链路调用及结构化指令中。
191
194
  * **按任务的权限预设** —— `default`、`read-only`、`workspace-write` 或 `full` 在提示词执行前应用于任务会话。
192
195
  * **会话自动归档** —— 隔离的 cron 会话在运行后自动归档(尽力而为),不干扰聊天列表。
193
196
  * **历史 → 会话** —— 每次 LLM 运行都会记录会话,可直接从历史记录打开对话。
@@ -226,12 +229,15 @@ cron({
226
229
  * **`skip`**(默认):丢弃重叠的运行,在历史中记录 `skipped`;
227
230
  * **`queue`**:将下一次运行排队,当前任务完成后自动开始;
228
231
  * **`replace`**:通过 `AbortController` 中止当前运行并启动新的执行。
232
+ * **调度器重叠策略直达** —— Croner 定时触发不再在上游被静默抑制(`Croner protect: false`),确保定时触发的重叠事件能够完整传递至调度器,严格执行 `skip`(记录历史)、`queue`(延迟排队)与 `replace`(优雅终止)。
233
+ * **队列上下文与链路深度延续** —— 全局并发限制队列与任务重叠队列均完整保留不可变执行参数(`chainDepth`、`prevOutput`、`prevTaskId`、`prevStatus`、`prevCostUsd`),防止高负载或排队时任务链路数据丢失。
229
234
 
230
235
  如果守护进程在计划时刻处于离线状态,启动时该次运行会被记录为 `missed`,历史空档始终可见。
231
236
 
232
237
  ### 14. 心跳监控(Dead man's switch)
233
238
  * 在插件设置中配置 `heartbeatUrl` 与 `heartbeatIntervalSec`,调度器会按间隔 GET 该地址 —— 外部监控可在心跳停止时告警。
234
239
  * 内置 `GET /dsh-cron/heartbeat` 端点返回存活状态、活跃任务数与最近运行时间,便于自建看门狗。
240
+ * **`onlyOnFailure` 模式支持心跳超时告警** — 遗漏心跳(`status === 'missed'`)在各通知渠道过滤判定(`channels.js`、`telegram.js`、`integrations.js`)中被统一视作失败状态,确保外部守护进程中断时 Telegram、Discord、Webhook 与看板即刻告警。
235
241
 
236
242
  ### 15. 来自配置的声明式任务(#50)
237
243
  长期运行的任务可以直接声明在配置文件里,而无需在界面中手工重建。配置文件拥有这些任务:每次插件启动时会创建或更新它们,从文件中消失的任务会被删除。
@@ -281,6 +287,7 @@ dsh-cron:
281
287
  | `POST` | `/dsh-cron/api/tasks/:id/run` | 强制执行一次 |
282
288
 
283
289
  这些操作复用面板处理器,因此对会执行代码类型的 `x-dsh-cron-confirm: script` 门禁以及对配置任务的 `409` 拒绝与 UI 完全一致。
290
+ * **Bearer Token 远程访问支持** — 远程客户端凭有效 `Authorization: Bearer <token>` 凭据可直接创建/修改任务、打卡心跳(`/dsh-cron/api/heartbeat/:id`)及预览调度(`/dsh-cron/api/schedule/preview`),解除非本机限制。
284
291
 
285
292
  ```bash
286
293
  BASE="http://127.0.0.1:3080"
@@ -363,7 +370,7 @@ bash deploy.sh verify [exact-version]
363
370
  ### 22. 自动化、任务链与可观测性包(v0.2.10,#137)
364
371
  - **Telegram 双向交互控制**:任务通知附带内嵌操作按钮(`🚀 立即运行`、`⏸️ 暂停/恢复`、`📋 最新日志`)。由 `POST /dsh-cron/telegram/webhook` 处理,严格鉴权 Chat ID 并调用 `answerCallbackQuery` 反馈。
365
372
  - **任务链上下文与动态变量插值**:配置 `onSuccess` 与 `onFailure` 下游触发器。父任务的执行结果与元数据自动传递给子任务,在 Shell 任务中提供 `$DSH_PREV_OUTPUT`、`$DSH_PREV_TASK_ID`、`$DSH_PREV_STATUS` 环境变量,在 LLM Prompt 中支持 `{{prev.output}}`(或 `{{prevOutput}}`)、`{{prev.taskId}}`、`{{prev.status}}` 占位符插值。Prompt 额外支持动态运行时时间与元数据变量:`{{date}}`、`{{time}}`、`{{datetime}}`、`{{timestamp}}`、`{{year}}`、`{{month}}`、`{{day}}`、`{{taskId}}`、`{{taskName}}`、`{{runCount}}`。内置最大 5 级深度递归防护,杜绝死循环。
366
- - **模型结构化动作指令**:自主分析任务可输出 JSON 指令触发级联任务(`trigger_task`)、定向告警(`notify`)或创建 Issue。受 `llmActionsEnabled: false` 严格保护。
373
+ - **模型结构化动作指令**:自主分析任务可输出 JSON 指令触发级联任务(`trigger_task`)、定向告警(`notify`)或创建 Issue。受 `llmActionsEnabled: false` 严格保护。`trigger_task` 指令共享全局链路深度上限(`chainDepth < 4`,最大 5 层调用),彻底阻断自调用死循环与 A → B → A 循环递归,拒绝原因完整记入历史与操作日志中。
367
374
  - **历史归档与延迟洞察**:REST 接口 `GET /dsh-cron/tasks/:id/archive`(支持分页)与 `GET /dsh-cron/tasks/:id/stats`;UI 任务卡片展示耗时彩色徽章(<5s 绿,<30s 黄,≥30s 红)。
368
375
  - **Prometheus 监控增强**:`/dsh-cron/metrics` 导出当前活动并发量 `dsh_cron_concurrent_running`、各任务 Token 计数器及成本预估指标。
369
376
 
@@ -374,6 +381,8 @@ bash deploy.sh verify [exact-version]
374
381
  - **优先级队列与并发池(Priority Queues)**:并发满载时,等待队列严格依据任务 `priority`(1 最高,10 最低)调度。
375
382
  - **自愈脚本与 AI 根因诊断(Self-Healing)**:任务失败后自动执行补偿指令 `selfHealingCommand`(例如重启服务或清理临时空间);`autoDiagnose` 自动生成 AI 故障根因摘要。
376
383
  - **UI 交互式归档与管道全景**:支持分页浏览任务历史运行全量输出,直观展示 `➜ 成功触发` 与 `↳ 失败触发` 关联关系。
384
+ - **Shell 与磁盘前置检查强化**:前置检查统一支持 `command` 与 `shell` 类型,未知类型或异常格式一律安全关闭(fail-closed);磁盘检查支持完整路径与阈值(如 `/:10%`、`/data:500MB`)。
385
+ - **可靠性与会话字段完整持久化**:HTTP POST 任务创建/修改接口与 `cron_create_task` 智能体工具完整保留并存储全部 13 个可靠性与会话高级字段(`agentPreset`、`targetSessionId`、`targetSessionReset`、`onSuccess`、`onFailure`、`heartbeatIntervalSeconds`、`gracePeriodSeconds`、`preflightType`、`preflightTarget`、`priority`、`concurrencyGroup`、`selfHealingCommand`、`autoDiagnose`、`fallbackProvider`、`fallbackModel`、`silentRule`、`inspectOnFailure`)。
377
386
 
378
387
  ---
379
388
 
@@ -3,12 +3,12 @@ import { sendJson, readBody, rejectCrossOrigin } from './http-utils.js';
3
3
 
4
4
  const NOT_ALLOWED = { ok: false, error: 'Method not allowed' };
5
5
 
6
- export function handleHeartbeatPing({ store, req, res, taskId }) {
6
+ export function handleHeartbeatPing({ store, req, res, taskId, apiToken: explicitToken }) {
7
7
  if (req.method !== 'POST') {
8
8
  sendJson(res, 405, NOT_ALLOWED);
9
9
  return;
10
10
  }
11
- const apiToken = store?.getSettings?.()?.apiToken;
11
+ const apiToken = explicitToken || store?.getSettings?.()?.apiToken;
12
12
  if (rejectCrossOrigin(req, res, { apiToken })) return;
13
13
 
14
14
  const task = store.get(taskId);
@@ -24,8 +24,8 @@ export function handleHeartbeatPing({ store, req, res, taskId }) {
24
24
  sendJson(res, 200, result);
25
25
  }
26
26
 
27
- export async function handleSchedulePreview({ req, res, url, store }) {
28
- const apiToken = store?.getSettings?.()?.apiToken;
27
+ export async function handleSchedulePreview({ req, res, url, store, apiToken: explicitToken }) {
28
+ const apiToken = explicitToken || store?.getSettings?.()?.apiToken;
29
29
  if (rejectCrossOrigin(req, res, { apiToken })) return;
30
30
 
31
31
  let schedule = '';
package/lib/api.js CHANGED
@@ -484,6 +484,19 @@ function mergeExecutionFields(current, body, parsed) {
484
484
  dailyCostLimitUsd: body.dailyCostLimitUsd !== undefined ? (body.dailyCostLimitUsd === null ? null : (Number(body.dailyCostLimitUsd) > 0 ? Number(body.dailyCostLimitUsd) : null)) : (current ? current.dailyCostLimitUsd : undefined),
485
485
  tokenLimit: body.tokenLimit !== undefined ? (body.tokenLimit === null ? null : (Number(body.tokenLimit) > 0 ? Number(body.tokenLimit) : null)) : (current ? current.tokenLimit : undefined),
486
486
  pausedReason: body.pausedReason !== undefined ? (body.pausedReason ? String(body.pausedReason) : null) : (current ? current.pausedReason : undefined),
487
+ agentPreset: body.agentPreset !== undefined ? String(body.agentPreset).trim() : (current ? (current.agentPreset || '') : ''),
488
+ targetSessionId: body.targetSessionId !== undefined ? String(body.targetSessionId).trim() : (current ? (current.targetSessionId || '') : ''),
489
+ targetSessionReset: body.targetSessionReset !== undefined ? String(body.targetSessionReset).trim() : (current ? (current.targetSessionReset || 'never') : 'never'),
490
+ onSuccess: body.onSuccess !== undefined ? String(body.onSuccess).trim() : (current ? (current.onSuccess || '') : ''),
491
+ onFailure: body.onFailure !== undefined ? String(body.onFailure).trim() : (current ? (current.onFailure || '') : ''),
492
+ heartbeatIntervalSeconds: body.heartbeatIntervalSeconds !== undefined ? Number(body.heartbeatIntervalSeconds) : (current ? (Number(current.heartbeatIntervalSeconds) || 0) : 0),
493
+ gracePeriodSeconds: body.gracePeriodSeconds !== undefined ? Number(body.gracePeriodSeconds) : (current ? (Number(current.gracePeriodSeconds) || 300) : 300),
494
+ preflightType: body.preflightType !== undefined ? String(body.preflightType).trim().toLowerCase() : (current ? (current.preflightType || 'none') : 'none'),
495
+ preflightTarget: body.preflightTarget !== undefined ? String(body.preflightTarget).trim() : (current ? (current.preflightTarget || '') : ''),
496
+ priority: body.priority !== undefined ? (Number(body.priority) || 5) : (current ? (Number(current.priority) || 5) : 5),
497
+ concurrencyGroup: body.concurrencyGroup !== undefined ? String(body.concurrencyGroup).trim() : (current ? (current.concurrencyGroup || 'default') : 'default'),
498
+ selfHealingCommand: body.selfHealingCommand !== undefined ? String(body.selfHealingCommand).trim() : (current ? (current.selfHealingCommand || '') : ''),
499
+ autoDiagnose: body.autoDiagnose !== undefined ? Boolean(body.autoDiagnose) : (current ? Boolean(current.autoDiagnose) : false),
487
500
  };
488
501
  }
489
502
 
@@ -495,14 +508,14 @@ function mergeExecutionFields(current, body, parsed) {
495
508
  * and reuses the handlers above, so the confirmation gate for code-executing
496
509
  * tasks and the config-owned refusals are the same code, not a second copy.
497
510
  */
498
- export async function handleExternalTaskRequest({ store, scheduler, req, res, url }) {
511
+ export async function handleExternalTaskRequest({ store, scheduler, req, res, url, apiToken }) {
499
512
  const parts = url.pathname.split('/').filter(Boolean); // dsh-cron, api, [tasks|heartbeat|schedule], :id?, :action?
500
513
  if (parts[2] === 'heartbeat' && parts[3]) {
501
- handleHeartbeatPing({ store, req, res, taskId: parts[3] });
514
+ handleHeartbeatPing({ store, req, res, taskId: parts[3], apiToken });
502
515
  return;
503
516
  }
504
517
  if (parts[2] === 'schedule' && parts[3] === 'preview') {
505
- await handleSchedulePreview({ req, res, url });
518
+ await handleSchedulePreview({ req, res, url, store, apiToken });
506
519
  return;
507
520
  }
508
521
  const id = parts[3];
@@ -517,7 +530,7 @@ export async function handleExternalTaskRequest({ store, scheduler, req, res, ur
517
530
  return;
518
531
  }
519
532
  if (req.method === 'POST') {
520
- await createOrUpdateTask({ store, scheduler, req, res });
533
+ await createOrUpdateTask({ store, scheduler, req, res, apiToken });
521
534
  return;
522
535
  }
523
536
  sendJson(res, 405, NOT_ALLOWED);
@@ -525,7 +538,7 @@ export async function handleExternalTaskRequest({ store, scheduler, req, res, ur
525
538
  }
526
539
 
527
540
  if (req.method === 'POST') {
528
- const handled = await handleItemPost({ store, scheduler, req, res, url, id, action });
541
+ const handled = await handleItemPost({ store, scheduler, req, res, url, id, action, apiToken });
529
542
  if (handled) return;
530
543
  }
531
544
  if (req.method === 'GET' && !action) {
package/lib/burn-guard.js CHANGED
@@ -14,10 +14,20 @@ import { bestEffort } from './best-effort.js';
14
14
  * @param {number} now
15
15
  * @returns {number}
16
16
  */
17
- export function calculateRollingCost(runs = [], windowMs = 86400000, now = Date.now()) {
17
+ export function calculateRollingCost(runsOrTask = [], windowMs = 86400000, now = Date.now()) {
18
18
  const cutoff = now - windowMs;
19
+ let items = [];
20
+ if (Array.isArray(runsOrTask)) {
21
+ items = runsOrTask;
22
+ } else if (runsOrTask && typeof runsOrTask === 'object') {
23
+ if (Array.isArray(runsOrTask.costLedger) && runsOrTask.costLedger.length > 0) {
24
+ items = runsOrTask.costLedger;
25
+ } else if (Array.isArray(runsOrTask.runs)) {
26
+ items = runsOrTask.runs;
27
+ }
28
+ }
19
29
  let total = 0;
20
- for (const r of runs) {
30
+ for (const r of items) {
21
31
  if (r && typeof r.at === 'number' && r.at >= cutoff) {
22
32
  total += Number(r.costUsd) || 0;
23
33
  }
@@ -32,14 +42,28 @@ export function calculateRollingCost(runs = [], windowMs = 86400000, now = Date.
32
42
  * @param {number} now
33
43
  * @returns {number}
34
44
  */
35
- export function calculateRollingTokens(runs = [], windowMs = 86400000, now = Date.now()) {
45
+ export function calculateRollingTokens(runsOrTask = [], windowMs = 86400000, now = Date.now()) {
36
46
  const cutoff = now - windowMs;
47
+ let items = [];
48
+ if (Array.isArray(runsOrTask)) {
49
+ items = runsOrTask;
50
+ } else if (runsOrTask && typeof runsOrTask === 'object') {
51
+ if (Array.isArray(runsOrTask.costLedger) && runsOrTask.costLedger.length > 0) {
52
+ items = runsOrTask.costLedger;
53
+ } else if (Array.isArray(runsOrTask.runs)) {
54
+ items = runsOrTask.runs;
55
+ }
56
+ }
37
57
  let total = 0;
38
- for (const r of runs) {
58
+ for (const r of items) {
39
59
  if (r && typeof r.at === 'number' && r.at >= cutoff) {
40
- const u = r.usage;
41
- const tokens = (u?.inputTokens || 0) + (u?.outputTokens || 0) + (u?.cacheReadTokens || 0);
42
- total += tokens;
60
+ if (typeof r.tokens === 'number') {
61
+ total += r.tokens;
62
+ } else {
63
+ const u = r.usage;
64
+ const tokens = (u?.inputTokens || 0) + (u?.outputTokens || 0) + (u?.cacheReadTokens || 0);
65
+ total += tokens;
66
+ }
43
67
  }
44
68
  }
45
69
  return total;
@@ -70,10 +94,11 @@ export function checkTaskBudgetLimits(task, runs = [], now = Date.now()) {
70
94
  }
71
95
  }
72
96
 
73
- // 2. Rolling 24h daily cost limit in USD
97
+ // 2. Rolling 24h daily cost limit in USD (#221)
74
98
  const dailyCostLimit = Number(task.dailyCostLimitUsd);
75
99
  if (dailyCostLimit > 0) {
76
- const dailyCost = calculateRollingCost(runs, 86400000, now);
100
+ const source = (Array.isArray(task.costLedger) && task.costLedger.length > 0) ? task.costLedger : runs;
101
+ const dailyCost = calculateRollingCost(source, 86400000, now);
77
102
  if (dailyCost >= dailyCostLimit) {
78
103
  return {
79
104
  exceeded: true,
package/lib/channels.js CHANGED
@@ -15,7 +15,8 @@ export const CHANNEL_IDS = ['telegram', 'kanban', 'discord', 'slack', 'ntfy', 'b
15
15
 
16
16
  // CHANNEL_LABELS removed (#155) — localized via lib/client.js
17
17
 
18
- const isFailed = (status) => status === 'error' || status === 'timeout';
18
+ export const isFailureStatus = (status) => status === 'error' || status === 'timeout' || status === 'missed';
19
+ export const isFailed = (status) => isFailureStatus(status);
19
20
 
20
21
  /** The message text for a channel: custom template, else built-in defaults. */
21
22
  function messageTextFor(channelId, task, runInfo, settings = {}) {
package/lib/cron-tool.js CHANGED
@@ -251,6 +251,60 @@ export const cronToolParameters = {
251
251
  type: 'number',
252
252
  description: 'Maximum allowed token usage before auto-pausing',
253
253
  },
254
+ agentPreset: {
255
+ type: 'string',
256
+ description: 'Agent preset name applied to the task session',
257
+ },
258
+ targetSessionId: {
259
+ type: 'string',
260
+ description: 'Persistent session ID for context continuity',
261
+ },
262
+ targetSessionReset: {
263
+ type: 'string',
264
+ enum: ['never', 'daily', 'weekly'],
265
+ description: 'Context rotation policy for target session: never | daily | weekly',
266
+ },
267
+ onSuccess: {
268
+ type: 'string',
269
+ description: 'Task ID to trigger on successful run',
270
+ },
271
+ onFailure: {
272
+ type: 'string',
273
+ description: 'Task ID to trigger on failed run',
274
+ },
275
+ heartbeatIntervalSeconds: {
276
+ type: 'number',
277
+ description: 'Expected heartbeat interval in seconds (0 = disabled)',
278
+ },
279
+ gracePeriodSeconds: {
280
+ type: 'number',
281
+ description: 'Grace period before alerting on missed heartbeat',
282
+ },
283
+ preflightType: {
284
+ type: 'string',
285
+ enum: ['none', 'http', 'shell', 'command', 'disk'],
286
+ description: 'Pre-flight health check type: none | http | shell | disk',
287
+ },
288
+ preflightTarget: {
289
+ type: 'string',
290
+ description: 'Pre-flight check target: URL, shell command, or path:minPercent / path:minMb',
291
+ },
292
+ concurrencyGroup: {
293
+ type: 'string',
294
+ description: 'Named concurrency group to serialize related tasks',
295
+ },
296
+ priority: {
297
+ type: 'number',
298
+ description: 'Task priority (1 = highest, 10 = lowest, default 5)',
299
+ },
300
+ selfHealingCommand: {
301
+ type: 'string',
302
+ description: 'Shell command executed to recover from run failure',
303
+ },
304
+ autoDiagnose: {
305
+ type: 'boolean',
306
+ description: 'Automatically run AI failure diagnosis and store in runbook',
307
+ },
254
308
  };
255
309
 
256
310
  export const cronToolOutput = {
@@ -66,7 +66,7 @@ export function createExternalApiHandler({ store, scheduler, getToken }) {
66
66
  // for example) surfaces as a thrown error from the shared handlers, and an
67
67
  // unhandled rejection here would answer nothing at all.
68
68
  try {
69
- await handleExternalTaskRequest({ store, scheduler, req, res, url });
69
+ await handleExternalTaskRequest({ store, scheduler, req, res, url, apiToken: expected });
70
70
  } catch (err) {
71
71
  sendJson(res, (err && err.statusCode) || 500, { ok: false, error: (err && err.message) || String(err) });
72
72
  }
@@ -91,7 +91,7 @@ export function shouldCreateKanbanCard(task, runInfo) {
91
91
  if (mode === 'none') return false;
92
92
  if (mode === 'always') return true;
93
93
  if (mode === 'on_failure') {
94
- return runInfo.status === 'error' || runInfo.status === 'timeout';
94
+ return runInfo.status === 'error' || runInfo.status === 'timeout' || runInfo.status === 'missed';
95
95
  }
96
96
  return false;
97
97
  }
@@ -65,7 +65,7 @@ function collectDirectives(obj, list) {
65
65
  /**
66
66
  * Execute parsed LLM action directives if enabled in plugin settings (#137).
67
67
  */
68
- export async function executeLlmActionDirectives({ directives, task, runInfo, scheduler, store, settings = {}, fetchFn = globalThis.fetch }) {
68
+ export async function executeLlmActionDirectives({ directives, task, runInfo, scheduler, store, settings = {}, fetchFn = globalThis.fetch, options = {} }) {
69
69
  if (!Array.isArray(directives) || directives.length === 0) return [];
70
70
  if (!settings.llmActionsEnabled) {
71
71
  return [];
@@ -78,11 +78,22 @@ export async function executeLlmActionDirectives({ directives, task, runInfo, sc
78
78
  if (type === 'trigger_task' || type === 'run_task') {
79
79
  const targetId = dir.taskId || dir.payload?.taskId;
80
80
  if (targetId && typeof scheduler.runNow === 'function') {
81
+ const currentDepth = options?.chainDepth || 0;
82
+ if (currentDepth >= 4) {
83
+ const limitMsg = `Structured action trigger_task recursion limit reached (depth ${currentDepth}). Target "${targetId}" not triggered.`;
84
+ logger.warn(`[dsh-cron] ${limitMsg}`);
85
+ results.push({ type, targetId, status: 'rejected', error: limitMsg, depth: currentDepth });
86
+ continue;
87
+ }
81
88
  await scheduler.runNow(targetId, {
89
+ chainDepth: currentDepth + 1,
82
90
  prevOutput: runInfo?.output || '',
83
91
  prevTaskId: task?.id || '',
92
+ prevStatus: runInfo?.status || 'success',
93
+ prevCostUsd: runInfo?.costUsd || 0,
94
+ prevDurationMs: runInfo?.durationMs || 0,
84
95
  });
85
- results.push({ type, targetId, status: 'triggered' });
96
+ results.push({ type, targetId, status: 'triggered', depth: currentDepth + 1 });
86
97
  }
87
98
  } else if (type === 'notify') {
88
99
  const msg = dir.message || dir.text || dir.payload?.message;
package/lib/runner.js CHANGED
@@ -324,7 +324,7 @@ export class SessionRunner {
324
324
 
325
325
  /** Create the session, drive one turn, and always clean up. */
326
326
  async _executeAgentTurn(task, prep) {
327
- let handle = null;
327
+ let handle = prep?.handle || null;
328
328
  let isResumed = false;
329
329
  try {
330
330
  // Resolve agent preset (#GH-1, #GH-2): scheduled llm sessions must join an agent preset
@@ -397,6 +397,13 @@ export class SessionRunner {
397
397
 
398
398
  this._applyPermissionPreset(task, handle);
399
399
  await prep.raceAbort(handle.agent.whenIdle());
400
+
401
+ // 1. Capture starting sequence before sending user message (#227, #228)
402
+ const session = handle.agent.session;
403
+ const startSeq = (session && typeof session.seq === 'number')
404
+ ? session.seq
405
+ : (Array.isArray(session?.log) ? session.log.length : 0);
406
+
400
407
  handle.agent.followup(prep.createUserMessage({
401
408
  content: [{ type: 'text', text: prep.agentPrompt }],
402
409
  source: { kind: 'plugin:dsh-cron', form: 'cron-execute' },
@@ -405,12 +412,73 @@ export class SessionRunner {
405
412
  if (typeof handle.agent.whenIdle === 'function') {
406
413
  await prep.raceAbort(handle.agent.whenIdle());
407
414
  }
408
- const usage = this._extractUsage(handle);
409
- const sessionNote = isResumed
410
- ? `resumed: ${prep.targetSessionId}`
411
- : `session: ${handle.agent.session?.id || 'cron'}`;
415
+
416
+ // 2. Snapshot events from this turn only (#227, #228)
417
+ let turnEvents = [];
418
+ if (typeof session?.snapshotEvents === 'function') {
419
+ turnEvents = session.snapshotEvents(startSeq) || [];
420
+ } else if (Array.isArray(session?.log)) {
421
+ turnEvents = session.log.slice(startSeq);
422
+ } else if (Array.isArray(handle.agent?.log)) {
423
+ turnEvents = handle.agent.log.slice(startSeq);
424
+ }
425
+
426
+ // 3. Terminal turn failure check (#227)
427
+ for (const ev of turnEvents) {
428
+ if (!ev) continue;
429
+ if (ev.type === 'turn/end' || ev.type === 'turn/error' || ev.type === 'turn/failure') {
430
+ const reason = ev.data?.reason || ev.reason;
431
+ const reasonKind = (reason && typeof reason === 'object') ? reason.kind : reason;
432
+ if (['error', 'interrupted', 'failed', 'aborted'].includes(reasonKind) || ev.data?.error || ev.error) {
433
+ const errMsg = ev.data?.error?.message || ev.error?.message || reason?.message || (typeof reason === 'string' ? reason : 'Agent turn ended with error');
434
+ const err = new Error(errMsg);
435
+ err.reason = reason;
436
+ err.turnFailed = true;
437
+ throw err;
438
+ }
439
+ }
440
+ }
441
+
442
+ // 4. Extract assistant message text from turn events (#227)
443
+ const messageTexts = [];
444
+ const chunkTexts = [];
445
+ for (const ev of turnEvents) {
446
+ if (!ev) continue;
447
+ if (ev.type === 'assistant/message') {
448
+ const msg = ev.data?.message || ev.data;
449
+ if (msg) {
450
+ if (typeof msg.content === 'string' && msg.content.trim()) {
451
+ messageTexts.push(msg.content.trim());
452
+ } else if (Array.isArray(msg.content)) {
453
+ const parts = msg.content
454
+ .filter((p) => p && p.type === 'text' && typeof p.text === 'string')
455
+ .map((p) => p.text)
456
+ .join('');
457
+ if (parts.trim()) messageTexts.push(parts.trim());
458
+ }
459
+ }
460
+ } else if (ev.type === 'assistant/chunk' && ev.data?.chunk) {
461
+ const chunk = ev.data.chunk;
462
+ if (chunk.type === 'text' && typeof chunk.text === 'string') {
463
+ chunkTexts.push(chunk.text);
464
+ } else if (chunk.type === 'text-delta' && typeof chunk.delta === 'string') {
465
+ chunkTexts.push(chunk.delta);
466
+ }
467
+ }
468
+ }
469
+ let outputText = messageTexts.length > 0 ? messageTexts.join('\n') : chunkTexts.join('');
470
+
471
+ // Fallback if no assistant message text was extracted from turn events
472
+ if (!outputText.trim()) {
473
+ const sessionNote = isResumed
474
+ ? `resumed: ${prep.targetSessionId}`
475
+ : `session: ${handle.agent.session?.id || 'cron'}`;
476
+ outputText = `[dsh-cron] Agent finished turn for task "${task.title}" (${sessionNote})`;
477
+ }
478
+
479
+ const usage = this._extractUsage(handle, turnEvents);
412
480
  return {
413
- output: `[dsh-cron] Agent finished turn for task "${task.title}" (${sessionNote})`,
481
+ output: outputText,
414
482
  usage,
415
483
  costUsd: estimateTokenCost(prep.model, usage),
416
484
  sessionId: handle.agent.session?.id || null,
@@ -433,17 +501,95 @@ export class SessionRunner {
433
501
  }
434
502
  }
435
503
 
436
- /** Token usage reported by the session, when the core exposes it. */
437
- _extractUsage(handle) {
504
+ /** Token usage reported by session events or the handle (#228). */
505
+ _extractUsage(handle, turnEvents = []) {
438
506
  const usage = { inputTokens: 0, outputTokens: 0, cacheReadTokens: 0 };
439
- bestEffort('extract-session-usage', () => {
440
- const sessionUsage = handle.agent.session?.usage || handle.agent.usage;
441
- if (sessionUsage) {
442
- usage.inputTokens = sessionUsage.inputTokens || sessionUsage.promptTokens || 0;
443
- usage.outputTokens = sessionUsage.outputTokens || sessionUsage.completionTokens || 0;
444
- usage.cacheReadTokens = sessionUsage.cacheReadTokens || sessionUsage.cachedTokens || 0;
507
+ let hasEventUsage = false;
508
+
509
+ bestEffort('extract-session-events-usage', () => {
510
+ const stepUsage = new Map();
511
+ const failedAttempts = [];
512
+
513
+ for (const ev of turnEvents) {
514
+ if (!ev || !ev.type) continue;
515
+
516
+ // assistant/chunk with usage type
517
+ if (ev.type === 'assistant/chunk' && ev.data?.chunk?.type === 'usage' && ev.data.chunk.usage) {
518
+ const u = ev.data.chunk.usage;
519
+ const key = `${ev.data.turn || 0}:${ev.data.step || 0}`;
520
+ if (!stepUsage.has(key)) {
521
+ stepUsage.set(key, {
522
+ inputTokens: Number(u.inputTokens || u.promptTokens || u.uncachedInputTokens) || 0,
523
+ outputTokens: Number(u.outputTokens || u.completionTokens) || 0,
524
+ cacheReadTokens: Number(u.cacheReadTokens || u.cachedTokens) || 0,
525
+ });
526
+ hasEventUsage = true;
527
+ }
528
+ }
529
+
530
+ // assistant/message finalized usage overwrites intermediate chunk usage for this step
531
+ if (ev.type === 'assistant/message') {
532
+ const u = ev.data?.usage || ev.data?.message?.usage;
533
+ if (u) {
534
+ const key = `${ev.data.turn || 0}:${ev.data.step || 0}`;
535
+ stepUsage.set(key, {
536
+ inputTokens: Number(u.inputTokens || u.promptTokens || u.uncachedInputTokens) || 0,
537
+ outputTokens: Number(u.outputTokens || u.completionTokens) || 0,
538
+ cacheReadTokens: Number(u.cacheReadTokens || u.cachedTokens) || 0,
539
+ });
540
+ hasEventUsage = true;
541
+ }
542
+ }
543
+
544
+ // assistant/attempt usage (e.g. paid failed attempts or retry attempts)
545
+ if (ev.type === 'assistant/attempt' && ev.data?.usage) {
546
+ const u = ev.data.usage;
547
+ if (ev.data.failed || ev.data.error || ev.data.status === 'failed') {
548
+ failedAttempts.push({
549
+ inputTokens: Number(u.inputTokens || u.promptTokens || u.uncachedInputTokens) || 0,
550
+ outputTokens: Number(u.outputTokens || u.completionTokens) || 0,
551
+ cacheReadTokens: Number(u.cacheReadTokens || u.cachedTokens) || 0,
552
+ });
553
+ hasEventUsage = true;
554
+ } else {
555
+ const key = `${ev.data.turn || 0}:${ev.data.step || 0}`;
556
+ if (!stepUsage.has(key)) {
557
+ stepUsage.set(key, {
558
+ inputTokens: Number(u.inputTokens || u.promptTokens || u.uncachedInputTokens) || 0,
559
+ outputTokens: Number(u.outputTokens || u.completionTokens) || 0,
560
+ cacheReadTokens: Number(u.cacheReadTokens || u.cachedTokens) || 0,
561
+ });
562
+ hasEventUsage = true;
563
+ }
564
+ }
565
+ }
566
+ }
567
+
568
+ if (hasEventUsage) {
569
+ for (const item of stepUsage.values()) {
570
+ usage.inputTokens += item.inputTokens;
571
+ usage.outputTokens += item.outputTokens;
572
+ usage.cacheReadTokens += item.cacheReadTokens;
573
+ }
574
+ for (const item of failedAttempts) {
575
+ usage.inputTokens += item.inputTokens;
576
+ usage.outputTokens += item.outputTokens;
577
+ usage.cacheReadTokens += item.cacheReadTokens;
578
+ }
445
579
  }
446
580
  });
581
+
582
+ if (!hasEventUsage) {
583
+ bestEffort('extract-session-usage-fallback', () => {
584
+ const sessionUsage = handle?.agent?.session?.usage || handle?.agent?.usage;
585
+ if (sessionUsage) {
586
+ usage.inputTokens = Number(sessionUsage.inputTokens || sessionUsage.promptTokens) || 0;
587
+ usage.outputTokens = Number(sessionUsage.outputTokens || sessionUsage.completionTokens) || 0;
588
+ usage.cacheReadTokens = Number(sessionUsage.cacheReadTokens || sessionUsage.cachedTokens) || 0;
589
+ }
590
+ });
591
+ }
592
+
447
593
  return usage;
448
594
  }
449
595
 
@@ -24,12 +24,21 @@ export function addUsage(a, b) {
24
24
  * Execute pre-flight check gate before running task body.
25
25
  */
26
26
  export async function executePreflight(task) {
27
- const type = String(task.preflightType || 'none').toLowerCase();
27
+ const type = String(task.preflightType || 'none').trim().toLowerCase();
28
28
  const target = String(task.preflightTarget || '').trim();
29
- if (!type || type === 'none' || !target) {
29
+ if (!type || type === 'none') {
30
30
  return { ok: true };
31
31
  }
32
32
 
33
+ const KNOWN_PREFLIGHT_TYPES = ['http', 'command', 'shell', 'disk'];
34
+ if (!KNOWN_PREFLIGHT_TYPES.includes(type)) {
35
+ return { ok: false, reason: `Unknown preflight type "${type}"` };
36
+ }
37
+
38
+ if (!target) {
39
+ return { ok: false, reason: `Preflight target is required for type "${type}"` };
40
+ }
41
+
33
42
  if (type === 'http') {
34
43
  try {
35
44
  const controller = new AbortController();
@@ -43,7 +52,7 @@ export async function executePreflight(task) {
43
52
  }
44
53
  }
45
54
 
46
- if (type === 'command') {
55
+ if (type === 'command' || type === 'shell') {
47
56
  return new Promise((resolve) => {
48
57
  exec(target, { timeout: 5000 }, (err, stdout, stderr) => {
49
58
  if (err) {
@@ -56,23 +65,49 @@ export async function executePreflight(task) {
56
65
  }
57
66
 
58
67
  if (type === 'disk') {
59
- const requiredMb = Number(target) || 100;
68
+ let diskPath = task.cwd || process.cwd();
69
+ let requirement = target;
70
+
71
+ // Check for "path:requirement", taking care not to split Windows drive letters (e.g. C:\:10%)
72
+ const lastColon = target.lastIndexOf(':');
73
+ if (lastColon > 0) {
74
+ const candidatePath = target.slice(0, lastColon).trim();
75
+ const candidateReq = target.slice(lastColon + 1).trim();
76
+ if (candidateReq) {
77
+ diskPath = candidatePath;
78
+ requirement = candidateReq;
79
+ }
80
+ }
81
+
82
+ const isPercent = requirement.endsWith('%');
83
+ const numericVal = parseFloat(requirement);
84
+ if (isNaN(numericVal) || numericVal <= 0) {
85
+ return { ok: false, reason: `Invalid preflight disk target requirement: "${requirement}"` };
86
+ }
87
+
60
88
  try {
61
- if (typeof fs.statfsSync === 'function') {
62
- const stats = fs.statfsSync(process.cwd());
89
+ if (typeof fs.statfsSync !== 'function') {
90
+ return { ok: false, reason: 'statfsSync is not available on this platform' };
91
+ }
92
+ const stats = fs.statfsSync(diskPath);
93
+ if (isPercent) {
94
+ const freePercent = stats.blocks > 0 ? (stats.bavail / stats.blocks) * 100 : 0;
95
+ if (freePercent < numericVal) {
96
+ return { ok: false, reason: `Free disk space on "${diskPath}" (${freePercent.toFixed(1)}%) is below required ${numericVal}%` };
97
+ }
98
+ } else {
63
99
  const freeMb = Math.round((stats.bavail * stats.bsize) / (1024 * 1024));
64
- if (freeMb < requiredMb) {
65
- return { ok: false, reason: `Free disk space (${freeMb} MB) is below required ${requiredMb} MB` };
100
+ if (freeMb < numericVal) {
101
+ return { ok: false, reason: `Free disk space on "${diskPath}" (${freeMb} MB) is below required ${numericVal} MB` };
66
102
  }
67
103
  }
68
104
  return { ok: true };
69
105
  } catch (err) {
70
- // Fail-open if statfsSync is unavailable or throws
71
- return { ok: true };
106
+ return { ok: false, reason: `Preflight disk check failed: filesystem "${diskPath}" not accessible (${err.message})` };
72
107
  }
73
108
  }
74
109
 
75
- return { ok: true };
110
+ return { ok: false, reason: `Unsupported preflight type "${type}"` };
76
111
  }
77
112
 
78
113
  /** Only agent-mediated types use a model, and only a failed run falls back. */
@@ -141,12 +176,17 @@ export async function executeWithFallback(scheduler, task, signal, options = {})
141
176
  /**
142
177
  * Begin a task run with concurrency and overlap policy checks.
143
178
  */
144
- export function beginRun(scheduler, task, taskId) {
179
+ export function beginRun(scheduler, task, taskId, options = {}) {
145
180
  if (scheduler.isStopped) return null;
146
181
  if (!scheduler.running.has(taskId) && scheduler.maxConcurrent > 0 && scheduler.running.size >= scheduler.maxConcurrent) {
147
182
  if (task.overlapPolicy === 'queue') {
148
183
  logger.info(`[dsh-cron] Task "${task.title}" (${taskId}) queued by concurrency limit (${scheduler.maxConcurrent})`);
149
- scheduler.queue.push({ taskId, priority: task.priority !== undefined ? task.priority : 5, queuedAt: Date.now() });
184
+ scheduler.queue.push({
185
+ taskId,
186
+ options: { ...options },
187
+ priority: task.priority !== undefined ? task.priority : 5,
188
+ queuedAt: Date.now(),
189
+ });
150
190
  scheduler.queue.sort((a, b) => (a.priority - b.priority) || (a.queuedAt - b.queuedAt));
151
191
  return null;
152
192
  }
@@ -169,14 +209,23 @@ export function beginRun(scheduler, task, taskId) {
169
209
  }
170
210
  if (overlapPolicy === 'queue') {
171
211
  logger.info(`[dsh-cron] Task "${task.title}" (${taskId}) is already running, queueing next run (overlapPolicy: queue)`);
172
- active.queueCount = (active.queueCount || 0) + 1;
212
+ active.queuedRuns = active.queuedRuns || [];
213
+ active.queuedRuns.push({ options: { ...options }, queuedAt: Date.now() });
214
+ active.queueCount = active.queuedRuns.length;
173
215
  return null;
174
216
  }
175
217
  }
176
218
 
177
219
  const controller = new AbortController();
178
220
  const runId = 'run_' + Math.random().toString(36).slice(2, 9) + '_' + Date.now();
179
- const currentRun = { id: runId, controller, startedAt: Date.now(), queueCount: 0 };
221
+ const pendingOverlap = Array.isArray(options?.pendingOverlapQueue) ? options.pendingOverlapQueue : [];
222
+ const currentRun = {
223
+ id: runId,
224
+ controller,
225
+ startedAt: Date.now(),
226
+ queuedRuns: pendingOverlap,
227
+ queueCount: pendingOverlap.length,
228
+ };
180
229
  scheduler.running.set(taskId, currentRun);
181
230
  return currentRun;
182
231
  }
@@ -221,7 +270,7 @@ export async function executeTask(scheduler, taskId, options = {}) {
221
270
  };
222
271
  }
223
272
 
224
- const currentRun = scheduler.beginRun(task, taskId);
273
+ const currentRun = scheduler.beginRun(task, taskId, options);
225
274
  if (currentRun === null) return null;
226
275
 
227
276
  const preflight = await executePreflight(task);
@@ -391,9 +440,15 @@ export function finishRun(scheduler, taskSnapshot, taskId, status, queuedCount,
391
440
  }
392
441
 
393
442
  if (queuedCount > 0 && !scheduler.isStopped) {
443
+ const queuedRuns = Array.isArray(meta?.queuedRuns) ? [...meta.queuedRuns] : [];
444
+ const nextItem = queuedRuns.shift();
445
+ const nextOptions = {
446
+ ...(nextItem?.options || {}),
447
+ ...(queuedRuns.length > 0 ? { pendingOverlapQueue: queuedRuns } : {}),
448
+ };
394
449
  setImmediate(() => {
395
450
  if (!scheduler.isStopped) {
396
- scheduler.runTask(taskId);
451
+ scheduler.runTask(taskId, nextOptions);
397
452
  }
398
453
  });
399
454
  }
@@ -462,10 +517,15 @@ export async function handleCompleteRun(scheduler, task, taskId, currentRun, out
462
517
  const finishedRun = scheduler.running.get(taskId);
463
518
  const isCurrentRun = (finishedRun === currentRun);
464
519
  if (isCurrentRun) scheduler.running.delete(taskId);
520
+ const queuedRuns = isCurrentRun ? (finishedRun?.queuedRuns || []) : [];
465
521
  const queuedCount = isCurrentRun ? (finishedRun?.queueCount || 0) : 0;
466
522
 
467
523
  if (!silent.skipped) await scheduler.deliverNotifications(task, runInfo);
468
- scheduler.finishRun(task, taskId, outcome.status, queuedCount, { isCurrentRun, runId: currentRun?.id });
524
+ scheduler.finishRun(task, taskId, outcome.status, queuedCount, {
525
+ isCurrentRun,
526
+ runId: currentRun?.id,
527
+ queuedRuns,
528
+ });
469
529
 
470
530
  if (scheduler.isStopped) return;
471
531
 
@@ -479,7 +539,7 @@ export async function handleCompleteRun(scheduler, task, taskId, currentRun, out
479
539
  if (queuedTask && queuedTask.status === 'active' && !scheduler.isStopped) {
480
540
  setImmediate(() => {
481
541
  if (!scheduler.isStopped) {
482
- scheduler.runTask(nextItem.taskId).catch((runErr) => {
542
+ scheduler.runTask(nextItem.taskId, nextItem.options || {}).catch((runErr) => {
483
543
  bestEffort('drain-queued-task', () => {}, scheduler.logger);
484
544
  });
485
545
  }
@@ -495,14 +555,18 @@ export async function handleCompleteRun(scheduler, task, taskId, currentRun, out
495
555
  if (settings.llmActionsEnabled) {
496
556
  const directives = parseLlmActionDirectives(outcome.output);
497
557
  if (directives.length > 0) {
498
- await executeLlmActionDirectives({
558
+ const actionResults = await executeLlmActionDirectives({
499
559
  directives,
500
560
  task,
501
561
  runInfo,
502
562
  scheduler,
503
563
  store: scheduler.store,
504
564
  settings,
565
+ options,
505
566
  });
567
+ if (actionResults && actionResults.length > 0) {
568
+ runInfo.actionResults = actionResults;
569
+ }
506
570
  }
507
571
  }
508
572
  } catch (actErr) {
package/lib/scheduler.js CHANGED
@@ -393,7 +393,6 @@ export class TaskScheduler {
393
393
  scheduleCron(task, parsed) {
394
394
  const timezone = task.timezone || this.defaultTimezone || undefined;
395
395
  const job = new Cron(parsed.cronPattern, {
396
- protect: true,
397
396
  sloppyRanges: true,
398
397
  catch: (err) => logger.error(`[dsh-cron] scheduled run of "${task.title}" (${task.id}) failed:`, (err && err.message) || err),
399
398
  ...(timezone ? { timezone } : {}),
@@ -426,8 +425,8 @@ export class TaskScheduler {
426
425
  }
427
426
  }
428
427
 
429
- beginRun(task, taskId) {
430
- return beginRun(this, task, taskId);
428
+ beginRun(task, taskId, options = {}) {
429
+ return beginRun(this, task, taskId, options);
431
430
  }
432
431
 
433
432
  countRun(status) {
package/lib/store.js CHANGED
@@ -53,7 +53,22 @@ export class TaskStore {
53
53
  this.tasks.clear();
54
54
  this.history.clear();
55
55
  if (Array.isArray(data.tasks)) {
56
+ const now = Date.now();
57
+ const cutoff24h = now - 86400000;
56
58
  for (const t of data.tasks) {
59
+ if (Array.isArray(t.costLedger) && t.costLedger.length > 0) {
60
+ t.costLedger = t.costLedger.filter(e => e && typeof e.at === 'number' && e.at >= cutoff24h);
61
+ } else {
62
+ // Seed from active history if present (#221)
63
+ const hist = (data.history && Array.isArray(data.history[t.id])) ? data.history[t.id] : [];
64
+ t.costLedger = hist
65
+ .filter(r => r && typeof r.at === 'number' && r.at >= cutoff24h)
66
+ .map(r => ({
67
+ at: r.at,
68
+ costUsd: Number(r.costUsd) || 0,
69
+ tokens: (r.usage?.inputTokens || 0) + (r.usage?.outputTokens || 0) + (r.usage?.cacheReadTokens || 0),
70
+ }));
71
+ }
57
72
  this.tasks.set(t.id, t);
58
73
  }
59
74
  }
@@ -424,6 +439,18 @@ export class TaskStore {
424
439
  task.totalCostUsd = Number(((task.totalCostUsd || 0) + costUsd).toFixed(6));
425
440
  task.updatedAt = Date.now();
426
441
 
442
+ // Rolling 24h cost ledger preserved across archive rotation and restarts (#221)
443
+ if (!Array.isArray(task.costLedger)) {
444
+ task.costLedger = [];
445
+ }
446
+ task.costLedger.push({
447
+ at: task.lastRunAt,
448
+ costUsd,
449
+ tokens: runTokens,
450
+ });
451
+ const cutoff24h = Date.now() - 86400000;
452
+ task.costLedger = task.costLedger.filter(e => e && typeof e.at === 'number' && e.at >= cutoff24h);
453
+
427
454
  const runs = this.history.get(id) || [];
428
455
  runs.unshift({
429
456
  id: 'run_' + Math.random().toString(36).slice(2, 9),
@@ -18,6 +18,23 @@ export function executeCreateTask(store, scheduler, args) {
18
18
  status: 'active',
19
19
  provider: args.provider,
20
20
  model: args.model,
21
+ fallbackProvider: args.fallbackProvider ? String(args.fallbackProvider).trim() : undefined,
22
+ fallbackModel: args.fallbackModel ? String(args.fallbackModel).trim() : undefined,
23
+ silentRule: args.silentRule !== undefined ? String(args.silentRule) : '',
24
+ inspectOnFailure: args.inspectOnFailure !== undefined ? Boolean(args.inspectOnFailure) : false,
25
+ agentPreset: args.agentPreset ? String(args.agentPreset).trim() : '',
26
+ targetSessionId: args.targetSessionId ? String(args.targetSessionId).trim() : '',
27
+ targetSessionReset: args.targetSessionReset ? String(args.targetSessionReset).trim() : 'never',
28
+ onSuccess: args.onSuccess ? String(args.onSuccess).trim() : '',
29
+ onFailure: args.onFailure ? String(args.onFailure).trim() : '',
30
+ heartbeatIntervalSeconds: Number(args.heartbeatIntervalSeconds) || 0,
31
+ gracePeriodSeconds: Number(args.gracePeriodSeconds) || 300,
32
+ preflightType: args.preflightType ? String(args.preflightType).trim().toLowerCase() : 'none',
33
+ preflightTarget: args.preflightTarget ? String(args.preflightTarget).trim() : '',
34
+ concurrencyGroup: args.concurrencyGroup ? String(args.concurrencyGroup).trim() : 'default',
35
+ priority: Number(args.priority) || 5,
36
+ selfHealingCommand: args.selfHealingCommand ? String(args.selfHealingCommand).trim() : '',
37
+ autoDiagnose: Boolean(args.autoDiagnose),
21
38
  notifyTelegram: Boolean(args.notifyTelegram),
22
39
  onlyOnFailure: Boolean(args.onlyOnFailure),
23
40
  timeoutSeconds: Number(args.timeoutSeconds) || 1800,
package/lib/task-patch.js CHANGED
@@ -27,8 +27,8 @@ export function isCodeExecutionIntroduced(current, body) {
27
27
  if (!current || !body) return false;
28
28
  if (body.type !== undefined && isCodeTypeSwitch(current, body.type)) return true;
29
29
 
30
- const hadPreflight = Boolean(current.preflightCommand && String(current.preflightCommand).trim()) || current.preflightType === 'command';
31
- const hasPreflight = Boolean(body.preflightCommand && String(body.preflightCommand).trim()) || body.preflightType === 'command';
30
+ const hadPreflight = Boolean(current.preflightCommand && String(current.preflightCommand).trim()) || current.preflightType === 'command' || current.preflightType === 'shell';
31
+ const hasPreflight = Boolean(body.preflightCommand && String(body.preflightCommand).trim()) || body.preflightType === 'command' || body.preflightType === 'shell';
32
32
  if (hasPreflight && !hadPreflight) return true;
33
33
 
34
34
  const hadSelfHealing = Boolean(current.selfHealingCommand && String(current.selfHealingCommand).trim());
@@ -228,7 +228,7 @@ export function isCodeExecutingTask(task) {
228
228
  const CODE_TYPES = ['script', 'node', 'python', 'ssh', 'docker'];
229
229
  if (CODE_TYPES.includes(normalizeTaskType(task.type))) return true;
230
230
  if (task.command && String(task.command).trim()) return true;
231
- if (task.preflightType === 'command' || (task.preflightCommand && String(task.preflightCommand).trim())) return true;
231
+ if (task.preflightType === 'command' || task.preflightType === 'shell' || (task.preflightCommand && String(task.preflightCommand).trim())) return true;
232
232
  if (task.selfHealingCommand && String(task.selfHealingCommand).trim()) return true;
233
233
  if (task.preflightScript && String(task.preflightScript).trim()) return true;
234
234
  return false;
package/lib/telegram.js CHANGED
@@ -210,7 +210,7 @@ export function shouldNotifyTask(task, runInfo, globalSettings = {}) {
210
210
  const isEnabled = task.notifyTelegram ?? globalSettings.notifyTelegram ?? false;
211
211
  if (!isEnabled) return false;
212
212
 
213
- const failed = runInfo.status === 'error' || runInfo.status === 'timeout';
213
+ const failed = runInfo.status === 'error' || runInfo.status === 'timeout' || runInfo.status === 'missed';
214
214
  const onlyOnFail = task.onlyOnFailure ?? globalSettings.onlyOnFailure ?? false;
215
215
  if (onlyOnFail && !failed) {
216
216
  return false;
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@goodandready/dsh-cron",
3
- "version": "0.2.35",
3
+ "version": "0.2.37",
4
4
  "description": "Background automation runner for DSH: isolated agent runs, script/HTTP/SSH/Docker runtimes, cost guard, notifications, heartbeats.",
5
5
  "type": "module",
6
6
  "main": "lib/index.js",