@lmzhen/dsh-evolution-review 0.3.81 → 0.3.82
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +4 -3
- package/lib/index.js +158 -35
- package/lib/types/session-state.d.ts +36 -0
- package/package.json +11 -11
package/README.md
CHANGED
|
@@ -31,19 +31,20 @@ Independent of request-prefix construction. This package does not alter the asse
|
|
|
31
31
|
- When the `evolution-state` service is not mounted, the memory/skill cadence state is not persisted and every turn restarts from a clean `{ turnsSinceMemory: 0, turnsSinceSkill: 0 }` baseline — the review schedule is stateless and re-decided each turn rather than accumulating across the conversation. The loss is surfaced once per process as a logger warning at the first turn/end.
|
|
32
32
|
- Read-before-write can see the review subagent's own `skill` reads only when the subagent backend exposes `localAgent` (the in-process driver does; out-of-process backends such as ACP and the CLI providers set `localAgent: undefined`). With a remote backend the subagent's reads are invisible, so a plan item patching a skill the subagent itself loaded is dropped as "unread" — the review then falls back to the parent session's reads only. Documented rather than worked around: recovering the child read set needs a `SubagentLike` contract change (v14 P2-6).
|
|
33
33
|
- **The default `reviewMode: 'inject'` produces no plan ledger.** The review runs in the parent session and emits no `evolution/plan-applied` event, so `evolution-activity`'s `activity.json` and the `evolution-replay` leaderboard never grow in this mode — the only production emit point sits inside the subagent path. Set `reviewMode: 'subagent'` on the `evolution-policy` row when the plan audit trail is required; the plugin states this once at load instead of leaving an empty ledger to be read as "no reviews happened" (v37 S2.2, plan decision (c)).
|
|
34
|
-
- **`.pinned` protection in the `'inject'` channel rides a family-internal session mark.** The review prompt marks the parent session, `tool-skill-manage` reads that mark and resolves both origin surfaces (approval + library) to `background_review`, and the next REAL user message (`source.kind === 'user'`) clears it — plugin notices, including the review prompt itself, do not. Two bounds follow from the platform's inject contract (no driver wake; a prompt is dropped on cancel/dispose and may be missed with an already-claimed batch): with `reviewWakeInject: false`, or on a host without `followup`, (a) a pending prompt may never execute while the session stays idle, and (b) when the user's next message wakes the session, the mark is cleared before that pending prompt reaches the model — its writes are then attributed `foreground` and the pinned guard does not cover them. The mark is in-memory, so a restart drops the window (v37 S2.2, plan decision (b)).
|
|
34
|
+
- **`.pinned` protection in the `'inject'` channel rides a family-internal session mark.** The review prompt marks the parent session, `tool-skill-manage` reads that mark and resolves both origin surfaces (approval + library) to `background_review`, and the next REAL user message (`source.kind === 'user'`) clears it — plugin notices, including the review prompt itself, do not. Two bounds follow from the platform's inject contract (no driver wake; a prompt is dropped on cancel/dispose and may be missed with an already-claimed batch): with `reviewWakeInject: false`, or on a host without `followup`, (a) a pending prompt may never execute while the session stays idle, and (b) when the user's next message wakes the session, the mark is cleared before that pending prompt reaches the model — its writes are then attributed `foreground` and the pinned guard does not cover them. The mark is in-memory, so a restart drops the window (v37 S2.2, plan decision (b)). A prompt delivered while a HUMAN message is already queued is not marked at all: the platform claims that human turn first, so a session-level window would attribute the user's own writes to `background_review` (S2-8, FLOW1-3) — the plugin reads the pre-claim queue (`agent.inbox`) because the platform exposes no claim identity, and warns once per mount when it withholds the mark.
|
|
35
35
|
|
|
36
36
|
## Configuration
|
|
37
37
|
|
|
38
38
|
`reviewProvider` selects the LLM provider for review subagents. When omitted, the subagent inherits the deployment default route instead of a hardcoded provider name. Model selection stays on the policy (`memoryReviewModel` / `skillReviewModel`).
|
|
39
39
|
|
|
40
|
-
`reviewTimeoutMs` bounds each review subagent run (an `AbortSignal.timeout`; `0` aborts immediately). The former `executionTimeoutMs` declaration was removed in v14 (nothing read it, so it was configuration that did nothing); use `reviewTimeoutMs`.
|
|
40
|
+
`reviewTimeoutMs` bounds each review subagent run (an `AbortSignal.timeout`; `0` aborts immediately) and the write leg that executes the plan. The deadline does not stop an abandoned write leg: an op that lands after it is reported to the model as its own notice (S2-10, FLOW1-6), while the `evolution/plan-applied` record stays the deadline snapshot, so late writes still have no plan-applied entry. The former `executionTimeoutMs` declaration was removed in v14 (nothing read it, so it was configuration that did nothing); use `reviewTimeoutMs`.
|
|
41
41
|
|
|
42
42
|
### Review delivery contract (0.3.38-0.3.42)
|
|
43
43
|
|
|
44
44
|
- **Both channels execute at conversation end only** (a `turn/end` with `reason.kind === 'completed'`): a cadence threshold fire mid-task merely latches the kind — no subagent spawn, no inject. The flush runs BEFORE the latch block (the completing turn may itself be a threshold-firing turn). `reviewMode` selects how the flush delivers: **`'inject'` is the default since 0.3.74**, `'subagent'` is an explicit opt-in. Rationale: the parent session already holds a warm prefix cache, so an injected prompt costs the new tokens only, while a spawned child re-prefills its own system prompt plus a redacted, re-serialized conversation digest (`buildReviewRequest`) under a **different model** (`skillReviewModel`/`memoryReviewModel`) — no prefix is shared with the parent, so the whole child input is paid at full price. `'subagent'` remains the choice for deployments that want the parent context kept clean (a review's skill reads and plan do not join the parent thread) or a dedicated review model; the subagent-only knobs (`reviewProvider`, `reviewTimeoutMs`, `reviewMaxDepth`, `reviewToolAllow`, the review models) are inert in inject mode. The explicit `'inject'` mode's historical "immediate on threshold" contract was superseded in 0.3.39 — both modes are end-of-conversation.
|
|
45
45
|
- **`skillReviewTrigger`** (default `'cadence'`): the cadence channel is **always on** (one end-of-conversation review from the cadence latch per task segment); the flag gates **only the completion channel** — `'cadence'` disables it, `'completion'` enables it (cadence still fires), `'both'` enables it on top of the always-on cadence. At one boundary a turn is served by exactly one review: the cadence flush runs first and returns, so `'both'` never double-sends a second task-complete prompt at the same boundary (V10-13).
|
|
46
|
-
- **`reviewWakeInject`** (default `true`): deliveries use `agent.followup` (next-turn + wake — the model starts processing immediately) instead of the non-waking `agent.inject` (which waits for the next driver wake). The host falls back to `inject` when it has no followup or the option is `false`. **The wake primitive is always called ON the agent instance** — the platform's `Agent.followup`/`inject` are prototype methods that call `this.send(...)`, so extracting one into a local and calling the detached reference throws (0.3.73: that throw was caught and logged while the cadence reset still ran, silently consuming every segment's review from 2026-09-07). A refused delivery now returns `false` and the caller keeps its latch and counters, so the review retries at the next completed boundary instead of vanishing; rule N13b in `packages/scripts/verify-arch-guards.mjs` pins the call form mechanically (comments and string literals are masked, and the detector self-tests at startup). The woken turn's own cadence fire is suppressed once (an injected review prompt alone must not re-trigger a review under `interval=1`); a restart clears the queue, so the loop cannot survive it.
|
|
46
|
+
- **`reviewWakeInject`** (default `true`): deliveries use `agent.followup` (next-turn + wake — the model starts processing immediately) instead of the non-waking `agent.inject` (which waits for the next driver wake). The host falls back to `inject` when it has no followup or the option is `false`. **The wake primitive is always called ON the agent instance** — the platform's `Agent.followup`/`inject` are prototype methods that call `this.send(...)`, so extracting one into a local and calling the detached reference throws (0.3.73: that throw was caught and logged while the cadence reset still ran, silently consuming every segment's review from 2026-09-07). A refused delivery now returns `false` and the caller keeps its latch and counters, so the review retries at the next completed boundary instead of vanishing; rule N13b in `packages/scripts/verify-arch-guards.mjs` pins the call form mechanically (comments and string literals are masked, and the detector self-tests at startup). The woken turn's own cadence fire is suppressed once (an injected review prompt alone must not re-trigger a review under `interval=1`); the suppression is bound to the turn the delivery WOKE — a busy-period turn that started before the delivery keeps its own cadence and cannot consume it (S2-9, FLOW1-5: the platform's `turn/start` carries only `{ turn }`, so ordering is the binding identity); a restart clears the queue, so the loop cannot survive it.
|
|
47
|
+
- **In-flight triggers coalesce, and the drain settles at delivery**: a trigger that arrives while a review is in flight is queued as ONE entry per (session, kind) — last trigger wins — so one window can never deliver two prompts for the same segment, and the next completed boundary starts a fresh run once the window closes (the settle wait is bounded; a handle the platform never settles after its abort is abandoned and reported as `evolution/review-error`). A drain delivery that fails restores that session's cadence latch (and, on the completion channel, its `completionInjected` flag), so the segment's review retries at the next completed boundary instead of vanishing with no trace (S2-7, FLOW1-2/1-4).
|
|
47
48
|
- **Counting window = injection-to-injection**: the `turnsSinceMemory`/`turnsSinceSkill` counters are monotonic across threshold fires (`resetOnFire: false`) and are zeroed at the flush delivery — a continued conversation starts a fresh segment from the injection. A threshold fire on the completing turn is caught by the flush (`pendingKind = latch ?? kind`). All deliveries (review prompt AND result notices) share the same waking channel; a failed counter-reset persist warns once per session (a stateful reload may re-deliver).
|
|
48
49
|
|
|
49
50
|
**Runtime invariant:** No companion is published. The platform auto-assembles nothing and the family mounts no `<pkg>/invariant` cordis row, so a companion here would never execute (v37 S2.1 / I-3).
|
package/lib/index.js
CHANGED
|
@@ -4,11 +4,63 @@ import { createUserMessage } from "@deepseek-ai/dsh-llm";
|
|
|
4
4
|
import { SessionId } from "@deepseek-ai/dsh-session";
|
|
5
5
|
import { COMPLETION_SKILL_REVIEW_PROMPT, DEFAULT_MAX_OPS_PER_PLAN, DEFAULT_MEMORY_CHAR_LIMIT, DEFAULT_MEMORY_REVIEW_MODEL, DEFAULT_REVIEW_CONTEXT_MESSAGES, DEFAULT_REVIEW_MEMORY_INTERVAL, DEFAULT_REVIEW_MESSAGE_CHARS, DEFAULT_REVIEW_SKILL_INTERVAL, DEFAULT_REVIEW_TIMEOUT_MS, DEFAULT_SKILL_CONTENT_CHARS, DEFAULT_SKILL_LIMITS, DEFAULT_SKILL_REVIEW_COMPLETION_MIN_TOOL_CALLS, DEFAULT_SKILL_REVIEW_MODEL, DEFAULT_SKILL_REVIEW_TRIGGER, DEFAULT_SUBSTANTIVE_MIN_AGENT_CHARS, DEFAULT_SUBSTANTIVE_MIN_TOOL_CALLS, DEFAULT_SUBSTANTIVE_MIN_USER_CHARS, DEFAULT_USER_CHAR_LIMIT, MAX_TIMER_DELAY_MS, PROMPT_BUNDLE, advanceReview, assertSkillsRootAliasRetired, clampedNumber, clearReviewChannel, contentHash, evolutionIoAdapter, foldToolDispatches, foldTurn, markReviewChannel, newSkillLibrary, readDispatchSignal, redactSecrets, resolveOrigins, resolveRootConfig, reviewPrompt, sessionAudited, skillReadNameOf, sweepReviewChannelSessions, verifyPromptBundle } from "@lmzhen/dsh-evolution-core";
|
|
6
6
|
import { validateEvolutionPlan } from "@lmzhen/dsh-evolution-plan-validator";
|
|
7
|
+
//#region lib/types/session-state.js
|
|
8
|
+
/**
|
|
9
|
+
* A registry of one plugin instance's session-keyed collections.
|
|
10
|
+
*
|
|
11
|
+
* P1-9 (S2-2): the review plugin kept eight per-session maps and cleared them
|
|
12
|
+
* from a hand-written list inside its dispose hook. The list drifted — the
|
|
13
|
+
* 0.3.38-0.3.42 additions (pendingCadenceReviews, pendingCadenceWarned,
|
|
14
|
+
* skipNextCadenceFire, cadenceResetWarned) were missing from it until V7-16 —
|
|
15
|
+
* because nothing tied a declaration to its cleanup. Registering at the
|
|
16
|
+
* declaration site removes the list: a collection clears because it was
|
|
17
|
+
* registered, and tests/session-state-ownership.spec.ts fails when a new
|
|
18
|
+
* `new Map<SessionId, …>` skips the registry.
|
|
19
|
+
*
|
|
20
|
+
* Deliberately package-local: evolution-review is the only current consumer
|
|
21
|
+
* (commands holds no per-session state, and state-json's warn dedupe sets are
|
|
22
|
+
* process-lifetime by design). Promote it to evolution-core when a second
|
|
23
|
+
* package needs the same registration discipline.
|
|
24
|
+
*/
|
|
25
|
+
var SessionScopedState = class {
|
|
26
|
+
registered = [];
|
|
27
|
+
/**
|
|
28
|
+
* Register a collection and return it, so declaration and registration stay
|
|
29
|
+
* one statement.
|
|
30
|
+
*
|
|
31
|
+
* @param key - the declaring variable name; also the diagnostic name.
|
|
32
|
+
* @param collection - any collection exposing clear().
|
|
33
|
+
* @returns the same collection.
|
|
34
|
+
*/
|
|
35
|
+
add(key, collection) {
|
|
36
|
+
if (this.registered.some((entry) => entry.key === key)) throw new Error(`dsh-evolution-review: session state "${key}" is registered twice`);
|
|
37
|
+
this.registered.push({
|
|
38
|
+
key,
|
|
39
|
+
collection
|
|
40
|
+
});
|
|
41
|
+
return collection;
|
|
42
|
+
}
|
|
43
|
+
/** Registered keys in registration order (diagnostics and the ownership spec). */
|
|
44
|
+
keys() {
|
|
45
|
+
return this.registered.map((entry) => entry.key);
|
|
46
|
+
}
|
|
47
|
+
/** Clear every registered collection. Idempotent; safe on a partially used mount. */
|
|
48
|
+
dispose() {
|
|
49
|
+
for (const entry of this.registered) entry.collection.clear();
|
|
50
|
+
}
|
|
51
|
+
};
|
|
52
|
+
//#endregion
|
|
7
53
|
//#region lib/types/index.js
|
|
8
54
|
/**
|
|
9
55
|
* Background review orchestration: signal gate → one-shot subagent → trusted plan execution.
|
|
10
56
|
* @module @lmzhen/dsh-evolution-review
|
|
11
57
|
*/
|
|
58
|
+
/** S2-6 (FLOW1-1): how long a subagent handle may keep the in-flight window
|
|
59
|
+
* open AFTER its own review timeout, before the review abandons it. Caps the
|
|
60
|
+
* wait at the review timeout itself so a short (test) budget stays short. */
|
|
61
|
+
const REVIEW_SETTLE_MARGIN_MS = 5e3;
|
|
62
|
+
/** Error name marking the S2-6 watchdog expiry (see the catch in trySubagentReview). */
|
|
63
|
+
const REVIEW_SETTLE_TIMEOUT = "ReviewSettleTimeout";
|
|
12
64
|
const name = "evolution-review";
|
|
13
65
|
const inject = ["agents"];
|
|
14
66
|
const Config = z.object({
|
|
@@ -61,6 +113,16 @@ function isSameKindPending(message, summary) {
|
|
|
61
113
|
const source = message.source;
|
|
62
114
|
return source?.kind === "plugin" && source.plugin === "dsh-evolution-review" && source.form === "notice" && source.summary === summary;
|
|
63
115
|
}
|
|
116
|
+
/**
|
|
117
|
+
* v43 audit (P1-4 / FLOW1): the cadence summary MUST name the review kind.
|
|
118
|
+
* The three cadence deliveries used to share the bare `auto-review` summary
|
|
119
|
+
* while their prompts differ by kind (memory / skill / combined), so
|
|
120
|
+
* `isSameKindPending` matched ACROSS kinds: a pending memory prompt was replaced
|
|
121
|
+
* in place by a skill prompt — one kind silently lost, while the caller still
|
|
122
|
+
* consumed the cadence latch and the deferred drain still emitted
|
|
123
|
+
* `evolution/review-scheduled` as if both had been delivered.
|
|
124
|
+
*/
|
|
125
|
+
const cadenceSummary = (kind) => `auto-review:${kind}`;
|
|
64
126
|
function clampReviewConfig(rawConfig, ctx) {
|
|
65
127
|
const clamped = [];
|
|
66
128
|
const field = (name, value, fallback, min, max) => {
|
|
@@ -94,14 +156,17 @@ function apply(ctx, rawConfig = {}) {
|
|
|
94
156
|
const config = clampReviewConfig(rawConfig, ctx);
|
|
95
157
|
assertSkillsRootAliasRetired(rawConfig);
|
|
96
158
|
const rootConfig = resolveRootConfig(rawConfig);
|
|
97
|
-
const
|
|
159
|
+
const sessionState = new SessionScopedState();
|
|
160
|
+
const turnStarts = sessionState.add("turnStarts", /* @__PURE__ */ new Map());
|
|
98
161
|
let statelessReviewStateWarned = false;
|
|
99
|
-
|
|
100
|
-
const
|
|
101
|
-
const
|
|
102
|
-
const
|
|
103
|
-
const
|
|
104
|
-
const
|
|
162
|
+
let channelMarkSuppressedWarned = false;
|
|
163
|
+
const cumulativeToolCalls = sessionState.add("cumulativeToolCalls", /* @__PURE__ */ new Map());
|
|
164
|
+
const completionInjected = sessionState.add("completionInjected", /* @__PURE__ */ new Set());
|
|
165
|
+
const pendingCadenceReviews = sessionState.add("pendingCadenceReviews", /* @__PURE__ */ new Map());
|
|
166
|
+
const pendingCadenceWarned = sessionState.add("pendingCadenceWarned", /* @__PURE__ */ new Set());
|
|
167
|
+
const skipNextCadenceFire = sessionState.add("skipNextCadenceFire", /* @__PURE__ */ new Map());
|
|
168
|
+
const lastTurnStart = sessionState.add("lastTurnStart", /* @__PURE__ */ new Map());
|
|
169
|
+
const cadenceResetWarned = sessionState.add("cadenceResetWarned", /* @__PURE__ */ new Set());
|
|
105
170
|
let reviewInFlight = false;
|
|
106
171
|
const policy = () => ctx.get("evolutionPolicy")?.get();
|
|
107
172
|
const schemaDefaults = Config["~standard"].validate({}).value;
|
|
@@ -122,8 +187,8 @@ function apply(ctx, rawConfig = {}) {
|
|
|
122
187
|
warnShadowed();
|
|
123
188
|
});
|
|
124
189
|
}
|
|
125
|
-
if ((policy()?.reviewMode ?? config.reviewMode) === "inject") ctx.logger.warn("dsh-evolution-review: reviewMode \"inject\" (default) runs the review in the parent session and emits NO evolution/plan-applied ledger entry — evolution-activity and evolution-replay stay empty in this mode (see
|
|
126
|
-
const reviewStateLocks = /* @__PURE__ */ new Map();
|
|
190
|
+
if ((policy()?.reviewMode ?? config.reviewMode) === "inject") ctx.logger.warn("dsh-evolution-review: reviewMode \"inject\" (default) runs the review in the parent session and emits NO evolution/plan-applied ledger entry — evolution-activity and evolution-replay stay empty in this mode (see the evolution-review README, \"Known Limitations and Deferred Work\"). Set reviewMode: \"subagent\" on the evolution-policy row to keep the audited plan path.");
|
|
191
|
+
const reviewStateLocks = sessionState.add("reviewStateLocks", /* @__PURE__ */ new Map());
|
|
127
192
|
async function withReviewStateLock(id, task) {
|
|
128
193
|
const next = (reviewStateLocks.get(id) ?? Promise.resolve()).catch(() => {}).then(task);
|
|
129
194
|
reviewStateLocks.set(id, next);
|
|
@@ -135,7 +200,10 @@ function apply(ctx, rawConfig = {}) {
|
|
|
135
200
|
}
|
|
136
201
|
ctx.on("session/event", (session, event) => {
|
|
137
202
|
if (!sessionAudited(ctx, session.id, config.sessionScoped)) return;
|
|
138
|
-
if (event.type === "turn/start" && session.header.origin !== "subagent")
|
|
203
|
+
if (event.type === "turn/start" && session.header.origin !== "subagent") {
|
|
204
|
+
turnStarts.set(session.id, session.seq - 1);
|
|
205
|
+
lastTurnStart.set(session.id, event.data.turn);
|
|
206
|
+
}
|
|
139
207
|
if (event.type === "user/message") {
|
|
140
208
|
if ((event.data?.source)?.kind === "user") clearReviewChannel(session.id);
|
|
141
209
|
return;
|
|
@@ -179,7 +247,8 @@ function apply(ctx, rawConfig = {}) {
|
|
|
179
247
|
ctx.logger.warn("dsh-evolution-review: evolution-state service not mounted — memory/skill review cadence is not persisted and resets every turn (see README Known Limitations).");
|
|
180
248
|
}
|
|
181
249
|
const snapshot = policy();
|
|
182
|
-
const
|
|
250
|
+
const suppression = skipNextCadenceFire.get(session.id);
|
|
251
|
+
const skipFire = suppression !== void 0 && event.data.turn > suppression.afterTurn;
|
|
183
252
|
if (skipFire) skipNextCadenceFire.delete(session.id);
|
|
184
253
|
let state = {
|
|
185
254
|
turnsSinceMemory: 0,
|
|
@@ -211,7 +280,7 @@ function apply(ctx, rawConfig = {}) {
|
|
|
211
280
|
if (pendingKind !== void 0) {
|
|
212
281
|
pendingCadenceReviews.delete(session.id);
|
|
213
282
|
if ((policy()?.reviewMode ?? config.reviewMode) === "inject") {
|
|
214
|
-
if (!deliverMessage(agent, reviewPrompt(pendingKind),
|
|
283
|
+
if (!deliverMessage(agent, reviewPrompt(pendingKind), cadenceSummary(pendingKind), true)) {
|
|
215
284
|
pendingCadenceReviews.set(session.id, pendingKind);
|
|
216
285
|
return;
|
|
217
286
|
}
|
|
@@ -245,7 +314,7 @@ function apply(ctx, rawConfig = {}) {
|
|
|
245
314
|
pendingCadenceReviews.set(session.id, pendingKind);
|
|
246
315
|
return;
|
|
247
316
|
} else if (reviewOutcome !== "deferred") {
|
|
248
|
-
if (!deliverMessage(agent, reviewPrompt(pendingKind),
|
|
317
|
+
if (!deliverMessage(agent, reviewPrompt(pendingKind), cadenceSummary(pendingKind), true)) {
|
|
249
318
|
pendingCadenceReviews.set(session.id, pendingKind);
|
|
250
319
|
return;
|
|
251
320
|
}
|
|
@@ -339,6 +408,31 @@ function apply(ctx, rawConfig = {}) {
|
|
|
339
408
|
* prototype method), the caller's catch demoted that to a console warning,
|
|
340
409
|
* and the cadence reset ran anyway — the segment's review was consumed with
|
|
341
410
|
* nothing queued (silent no-delivery window: 2026-09-07 → 0.3.73). */
|
|
411
|
+
/**
|
|
412
|
+
* S2-8 (FLOW1-3): is a HUMAN message already queued ahead of ours? The
|
|
413
|
+
* review-channel mark is a session-level WINDOW, not a claim on one message:
|
|
414
|
+
* if a real user message is pending when the review prompt is delivered, the
|
|
415
|
+
* platform claims that human turn first and the mark would be in effect for
|
|
416
|
+
* the user's own writes — attributing them `background_review` (the pinned /
|
|
417
|
+
* hermes-managed protection would cover human edits). The platform exposes no
|
|
418
|
+
* claim identity to the plugin, so the pre-claim queue is the identity that is
|
|
419
|
+
* readable; this is the plan's documented fallback, and it only ever WITHHOLDS
|
|
420
|
+
* the mark (a host without `inbox` behaves exactly as before).
|
|
421
|
+
*/
|
|
422
|
+
const humanQueuedAhead = (inbox) => [...inbox?.nextTurn ?? [], ...inbox?.nextStep ?? []].some((row) => row.source?.kind === "user");
|
|
423
|
+
/** Mark the delivered review prompt's session — unless human input is queued
|
|
424
|
+
* ahead of it (see humanQueuedAhead). Says so once per mount: an unmarked
|
|
425
|
+
* window silently downgrades that review's writes to foreground attribution. */
|
|
426
|
+
const markReviewChannelForDelivery = (agent, inbox) => {
|
|
427
|
+
if (humanQueuedAhead(inbox)) {
|
|
428
|
+
if (!channelMarkSuppressedWarned) {
|
|
429
|
+
channelMarkSuppressedWarned = true;
|
|
430
|
+
ctx.logger.warn("dsh-evolution-review: the review prompt was delivered BEHIND queued human input — the review-channel mark stays unset for this window, so this review’s writes are attributed foreground and the .pinned/.hermes-managed protection does not cover them");
|
|
431
|
+
}
|
|
432
|
+
return;
|
|
433
|
+
}
|
|
434
|
+
markReviewChannel(agent.session.id);
|
|
435
|
+
};
|
|
342
436
|
const deliverMessage = (agent, text, summary, reviewPrompt = false) => {
|
|
343
437
|
const message = createUserMessage({
|
|
344
438
|
content: [{
|
|
@@ -356,7 +450,7 @@ function apply(ctx, rawConfig = {}) {
|
|
|
356
450
|
if (inbox !== void 0 && typeof inbox.replace === "function") try {
|
|
357
451
|
const superseded = [...inbox.nextTurn ?? [], ...inbox.nextStep ?? []].find((row) => isSameKindPending(row, summary));
|
|
358
452
|
if (superseded !== void 0 && inbox.replace(superseded.id, message)) {
|
|
359
|
-
if (reviewPrompt)
|
|
453
|
+
if (reviewPrompt) markReviewChannelForDelivery(agent, inbox);
|
|
360
454
|
return true;
|
|
361
455
|
}
|
|
362
456
|
} catch (error) {
|
|
@@ -366,15 +460,31 @@ function apply(ctx, rawConfig = {}) {
|
|
|
366
460
|
try {
|
|
367
461
|
if (config.reviewWakeInject && typeof wake.followup === "function") {
|
|
368
462
|
wake.followup(message);
|
|
369
|
-
skipNextCadenceFire.set(agent.session.id,
|
|
463
|
+
skipNextCadenceFire.set(agent.session.id, { afterTurn: lastTurnStart.get(agent.session.id) ?? -1 });
|
|
370
464
|
} else agent.inject(message);
|
|
371
|
-
if (reviewPrompt)
|
|
465
|
+
if (reviewPrompt) markReviewChannelForDelivery(agent, inbox);
|
|
372
466
|
return true;
|
|
373
467
|
} catch (error) {
|
|
374
468
|
ctx.logger.warn(`dsh-evolution-review: review delivery failed (${error instanceof Error ? error.message : String(error)}) — nothing was queued; the review is NOT consumed and retries at the next completed boundary`);
|
|
375
469
|
return false;
|
|
376
470
|
}
|
|
377
471
|
};
|
|
472
|
+
const settleBudgetMs = () => Math.min(config.reviewTimeoutMs, REVIEW_SETTLE_MARGIN_MS);
|
|
473
|
+
const withSettleWatchdog = (promise, label) => new Promise((resolve, reject) => {
|
|
474
|
+
const ms = settleBudgetMs();
|
|
475
|
+
const timer = setTimeout(() => {
|
|
476
|
+
const error = /* @__PURE__ */ new Error(`dsh-evolution-review: ${label} did not settle within ${ms}ms of the review timeout — abandoning the handle`);
|
|
477
|
+
error.name = REVIEW_SETTLE_TIMEOUT;
|
|
478
|
+
reject(error);
|
|
479
|
+
}, ms);
|
|
480
|
+
promise.then((value) => {
|
|
481
|
+
clearTimeout(timer);
|
|
482
|
+
resolve(value);
|
|
483
|
+
}, (error) => {
|
|
484
|
+
clearTimeout(timer);
|
|
485
|
+
reject(error instanceof Error ? error : new Error(String(error)));
|
|
486
|
+
});
|
|
487
|
+
});
|
|
378
488
|
const withTimeout = (promise, ms, label) => new Promise((resolve, reject) => {
|
|
379
489
|
const timer = setTimeout(() => {
|
|
380
490
|
reject(/* @__PURE__ */ new Error(`dsh-evolution-review: ${label} timed out after ${ms}ms`));
|
|
@@ -417,16 +527,19 @@ function apply(ctx, rawConfig = {}) {
|
|
|
417
527
|
const subagents = ctx.get("subagents");
|
|
418
528
|
if (!subagents) return false;
|
|
419
529
|
if (reviewInFlight) {
|
|
420
|
-
|
|
421
|
-
|
|
530
|
+
const existing = deferredFallbackReviews.findIndex((entry) => entry.sessionId === session.id && entry.kind === kind);
|
|
531
|
+
if (existing >= 0 || deferredFallbackReviews.length < DEFERRED_REVIEW_CAP) {
|
|
532
|
+
const entry = {
|
|
422
533
|
agent,
|
|
423
534
|
sessionId: session.id,
|
|
424
535
|
kind,
|
|
425
536
|
prompt: reviewPrompt(kind),
|
|
426
|
-
label:
|
|
537
|
+
label: cadenceSummary(kind),
|
|
427
538
|
channel: "inject",
|
|
428
539
|
counts: signal
|
|
429
|
-
}
|
|
540
|
+
};
|
|
541
|
+
if (existing >= 0) deferredFallbackReviews[existing] = entry;
|
|
542
|
+
else deferredFallbackReviews.push(entry);
|
|
430
543
|
return "deferred";
|
|
431
544
|
}
|
|
432
545
|
ctx.logger.warn(`dsh-evolution-review: deferred-review queue at cap (${DEFERRED_REVIEW_CAP}) — dropping one fallback review prompt`);
|
|
@@ -463,7 +576,7 @@ function apply(ctx, rawConfig = {}) {
|
|
|
463
576
|
outputSchema: REVIEW_OUTPUT_SCHEMA
|
|
464
577
|
});
|
|
465
578
|
try {
|
|
466
|
-
const result = await run.result;
|
|
579
|
+
const result = await withSettleWatchdog(run.result, "subagent review result");
|
|
467
580
|
if (!result.structured) {
|
|
468
581
|
try {
|
|
469
582
|
ctx.emit("evolution/review-error", { sessionId: session.id });
|
|
@@ -509,10 +622,21 @@ function apply(ctx, rawConfig = {}) {
|
|
|
509
622
|
}
|
|
510
623
|
};
|
|
511
624
|
const landed = [];
|
|
625
|
+
let deadlineMissed = false;
|
|
626
|
+
const lateLandings = [];
|
|
627
|
+
const noteLanded = (action) => {
|
|
628
|
+
landed.push(action);
|
|
629
|
+
if (!deadlineMissed) return;
|
|
630
|
+
lateLandings.push(action);
|
|
631
|
+
const before = landed.length - lateLandings.length;
|
|
632
|
+
const earlier = before > 0 ? ` (${before} other op(s) had landed before the deadline)` : "";
|
|
633
|
+
deliverMessage(agent, `\u{1F4BE} Self-improvement review: ${action} ${lateLandings.length === 1 ? "landed AFTER the review timed out" : `landed after the timeout (${lateLandings.length} late so far)`}${earlier} \u2014 reconcile before re-requesting the same change.`, "self-improvement review late landing");
|
|
634
|
+
};
|
|
512
635
|
let executed;
|
|
513
636
|
try {
|
|
514
|
-
executed = await withTimeout(executePlan(validation.accepted, session,
|
|
637
|
+
executed = await withTimeout(executePlan(validation.accepted, session, noteLanded, preRunHashes), config.reviewTimeoutMs, "review plan execution");
|
|
515
638
|
} catch (error) {
|
|
639
|
+
deadlineMissed = true;
|
|
516
640
|
emitApplied({
|
|
517
641
|
actions: landed,
|
|
518
642
|
executionError: `execution timed out after ${config.reviewTimeoutMs}ms`
|
|
@@ -543,14 +667,19 @@ function apply(ctx, rawConfig = {}) {
|
|
|
543
667
|
return true;
|
|
544
668
|
} finally {
|
|
545
669
|
try {
|
|
546
|
-
await run.dispose();
|
|
670
|
+
await withSettleWatchdog(run.dispose(), "subagent dispose");
|
|
547
671
|
} catch (disposeError) {
|
|
548
672
|
ctx.logger.warn(`dsh-evolution-review: subagent dispose failed: ${disposeError instanceof Error ? disposeError.message : String(disposeError)}`);
|
|
549
673
|
}
|
|
550
674
|
}
|
|
551
675
|
} catch (error) {
|
|
552
676
|
ctx.logger.warn(`dsh-evolution-review: subagent review failed: ${error instanceof Error ? error.message : String(error)}`);
|
|
553
|
-
if (error instanceof Error && error.
|
|
677
|
+
if (error instanceof Error && error.name === REVIEW_SETTLE_TIMEOUT) try {
|
|
678
|
+
ctx.emit("evolution/review-error", { sessionId: session.id });
|
|
679
|
+
} catch (emitError) {
|
|
680
|
+
ctx.logger.warn(`dsh-evolution-review: review-error emit failed: ${emitError instanceof Error ? emitError.message : String(emitError)}`);
|
|
681
|
+
}
|
|
682
|
+
if (error instanceof Error && error.message.includes("plan execution timed out")) ctx.logger.warn("dsh-evolution-review: plan execution abandoned on timeout — late writes ARE reported to the model as they land (S2-10) but still have NO plan-applied record and race the fallback inject; inspect the skill tree and usage sidecar");
|
|
554
683
|
return false;
|
|
555
684
|
} finally {
|
|
556
685
|
reviewInFlight = false;
|
|
@@ -558,6 +687,7 @@ function apply(ctx, rawConfig = {}) {
|
|
|
558
687
|
for (const { agent: waitingAgent, sessionId: entrySession, kind: waitingKind, prompt, label, channel, counts: entryCounts } of deferred) {
|
|
559
688
|
if (!deliverMessage(waitingAgent, prompt, label, true)) {
|
|
560
689
|
if (channel === "completion") completionInjected.delete(entrySession);
|
|
690
|
+
else pendingCadenceReviews.set(entrySession, waitingKind);
|
|
561
691
|
continue;
|
|
562
692
|
}
|
|
563
693
|
try {
|
|
@@ -757,12 +887,12 @@ function apply(ctx, rawConfig = {}) {
|
|
|
757
887
|
if (op.action === "edit" || op.action === "update") {
|
|
758
888
|
const updated = await library.update(name, op.content ?? "", origin, anchorOf(op));
|
|
759
889
|
if (updated.stale === true) return staleRefusal(updated, name, op.file_path);
|
|
760
|
-
if (updated.ok) await ctx.get("skillUsage")?.record?.(name, "patch").catch(() => {});
|
|
890
|
+
if (updated.ok && updated.noop !== true) await ctx.get("skillUsage")?.record?.(name, "patch").catch(() => {});
|
|
761
891
|
return updated;
|
|
762
892
|
}
|
|
763
893
|
if (op.action === "patch") {
|
|
764
894
|
const patched = await library.patch(name, op.old_string ?? "", op.new_string ?? "", op.file_path ?? "", op.replace_all === true, origin);
|
|
765
|
-
if (patched.ok) await ctx.get("skillUsage")?.record?.(name, "patch").catch(() => {});
|
|
895
|
+
if (patched.ok && patched.noop !== true) await ctx.get("skillUsage")?.record?.(name, "patch").catch(() => {});
|
|
766
896
|
return patched;
|
|
767
897
|
}
|
|
768
898
|
if (op.action === "delete") {
|
|
@@ -784,7 +914,7 @@ function apply(ctx, rawConfig = {}) {
|
|
|
784
914
|
}
|
|
785
915
|
const removedSupport = await library.removeSupportFile(name, op.file_path ?? "", origin, anchor);
|
|
786
916
|
if (removedSupport.stale === true) return staleRefusal(removedSupport, name, op.file_path);
|
|
787
|
-
if (removedSupport.ok) await ctx.get("skillUsage")?.record?.(name, "patch").catch(() => {});
|
|
917
|
+
if (removedSupport.ok && removedSupport.noop !== true) await ctx.get("skillUsage")?.record?.(name, "patch").catch(() => {});
|
|
788
918
|
return removedSupport;
|
|
789
919
|
}
|
|
790
920
|
if (op.action === "restructure") {
|
|
@@ -801,14 +931,7 @@ function apply(ctx, rawConfig = {}) {
|
|
|
801
931
|
}
|
|
802
932
|
}
|
|
803
933
|
ctx.effect(() => () => {
|
|
804
|
-
|
|
805
|
-
cumulativeToolCalls.clear();
|
|
806
|
-
completionInjected.clear();
|
|
807
|
-
pendingCadenceReviews.clear();
|
|
808
|
-
pendingCadenceWarned.clear();
|
|
809
|
-
skipNextCadenceFire.clear();
|
|
810
|
-
cadenceResetWarned.clear();
|
|
811
|
-
reviewStateLocks.clear();
|
|
934
|
+
sessionState.dispose();
|
|
812
935
|
}, "dsh-evolution-review.cleanup");
|
|
813
936
|
}
|
|
814
937
|
/** Completion-channel decision: task finished normally AND the session is proven long. */
|
|
@@ -0,0 +1,36 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* A registry of one plugin instance's session-keyed collections.
|
|
3
|
+
*
|
|
4
|
+
* P1-9 (S2-2): the review plugin kept eight per-session maps and cleared them
|
|
5
|
+
* from a hand-written list inside its dispose hook. The list drifted — the
|
|
6
|
+
* 0.3.38-0.3.42 additions (pendingCadenceReviews, pendingCadenceWarned,
|
|
7
|
+
* skipNextCadenceFire, cadenceResetWarned) were missing from it until V7-16 —
|
|
8
|
+
* because nothing tied a declaration to its cleanup. Registering at the
|
|
9
|
+
* declaration site removes the list: a collection clears because it was
|
|
10
|
+
* registered, and tests/session-state-ownership.spec.ts fails when a new
|
|
11
|
+
* `new Map<SessionId, …>` skips the registry.
|
|
12
|
+
*
|
|
13
|
+
* Deliberately package-local: evolution-review is the only current consumer
|
|
14
|
+
* (commands holds no per-session state, and state-json's warn dedupe sets are
|
|
15
|
+
* process-lifetime by design). Promote it to evolution-core when a second
|
|
16
|
+
* package needs the same registration discipline.
|
|
17
|
+
*/
|
|
18
|
+
export declare class SessionScopedState {
|
|
19
|
+
private readonly registered;
|
|
20
|
+
/**
|
|
21
|
+
* Register a collection and return it, so declaration and registration stay
|
|
22
|
+
* one statement.
|
|
23
|
+
*
|
|
24
|
+
* @param key - the declaring variable name; also the diagnostic name.
|
|
25
|
+
* @param collection - any collection exposing clear().
|
|
26
|
+
* @returns the same collection.
|
|
27
|
+
*/
|
|
28
|
+
add<T extends {
|
|
29
|
+
clear(): void;
|
|
30
|
+
}>(key: string, collection: T): T;
|
|
31
|
+
/** Registered keys in registration order (diagnostics and the ownership spec). */
|
|
32
|
+
keys(): string[];
|
|
33
|
+
/** Clear every registered collection. Idempotent; safe on a partially used mount. */
|
|
34
|
+
dispose(): void;
|
|
35
|
+
}
|
|
36
|
+
//# sourceMappingURL=session-state.d.ts.map
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@lmzhen/dsh-evolution-review",
|
|
3
3
|
"description": "Background review orchestration (community build)",
|
|
4
|
-
"version": "0.3.
|
|
4
|
+
"version": "0.3.82",
|
|
5
5
|
"publishConfig": {
|
|
6
6
|
"access": "public"
|
|
7
7
|
},
|
|
@@ -27,9 +27,9 @@
|
|
|
27
27
|
"license": "MIT",
|
|
28
28
|
"dependencies": {
|
|
29
29
|
"@deepseek-ai/schemastery": "^3.18.1",
|
|
30
|
-
"@lmzhen/dsh-evolution-approval": "^0.3.
|
|
31
|
-
"@lmzhen/dsh-evolution-core": "^0.3.
|
|
32
|
-
"@lmzhen/dsh-evolution-plan-validator": "^0.3.
|
|
30
|
+
"@lmzhen/dsh-evolution-approval": "^0.3.82",
|
|
31
|
+
"@lmzhen/dsh-evolution-core": "^0.3.82",
|
|
32
|
+
"@lmzhen/dsh-evolution-plan-validator": "^0.3.82"
|
|
33
33
|
},
|
|
34
34
|
"peerDependencies": {
|
|
35
35
|
"@deepseek-ai/cordis": "^4.0.1",
|
|
@@ -37,8 +37,8 @@
|
|
|
37
37
|
"@deepseek-ai/dsh-llm": "^0.1.5-rc.2",
|
|
38
38
|
"@deepseek-ai/dsh-session": "^0.1.5-rc.2",
|
|
39
39
|
"@deepseek-ai/dsh-tools": "^0.1.5-rc.2",
|
|
40
|
-
"@lmzhen/dsh-evolution-state": "^0.3.
|
|
41
|
-
"@lmzhen/dsh-evolution-policy": "^0.3.
|
|
40
|
+
"@lmzhen/dsh-evolution-state": "^0.3.82",
|
|
41
|
+
"@lmzhen/dsh-evolution-policy": "^0.3.82"
|
|
42
42
|
},
|
|
43
43
|
"devDependencies": {
|
|
44
44
|
"@deepseek-ai/dsh-agent": "^0.1.5-rc.2",
|
|
@@ -48,10 +48,10 @@
|
|
|
48
48
|
"@deepseek-ai/dsh-session-persistence": "^0.1.5-rc.2",
|
|
49
49
|
"@deepseek-ai/dsh-session-persistence-jsonl": "^0.1.5-rc.2",
|
|
50
50
|
"@deepseek-ai/dsh-tools": "^0.1.5-rc.2",
|
|
51
|
-
"@lmzhen/dsh-evolution-approval": "^0.3.
|
|
52
|
-
"@lmzhen/dsh-evolution-core": "^0.3.
|
|
53
|
-
"@lmzhen/dsh-evolution-curator": "^0.3.
|
|
54
|
-
"@lmzhen/dsh-evolution-plan-validator": "^0.3.
|
|
55
|
-
"@lmzhen/dsh-evolution-state": "^0.3.
|
|
51
|
+
"@lmzhen/dsh-evolution-approval": "^0.3.82",
|
|
52
|
+
"@lmzhen/dsh-evolution-core": "^0.3.82",
|
|
53
|
+
"@lmzhen/dsh-evolution-curator": "^0.3.82",
|
|
54
|
+
"@lmzhen/dsh-evolution-plan-validator": "^0.3.82",
|
|
55
|
+
"@lmzhen/dsh-evolution-state": "^0.3.82"
|
|
56
56
|
}
|
|
57
57
|
}
|