@lmzhen/dsh-evolution-review 0.8.0 → 0.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -17,12 +17,13 @@ own, and the delivery mode lives on the `evolution-policy` row, not here.
17
17
  - Review subagents are spawned with the plain `skill` tool only (`reviewToolAllow` default and the host/preset config both = `[skill]`: the DSH tool catalog has no `skill_search`/`skill_load` discovery pair, so the Hermes-lineage Anchored Standard `skill_search`/`skill_load` allow-list does not exist here).
18
18
  - Review subagents run as `spawn` children on the deployment default preset rather than inheriting the parent agent's composition (`fork`): a fork child is always promoted by the Anchored Standard bootstrap and its narrowed resident catalog would drop the plain `skill` tool from the review allow-list.
19
19
  - The review request text is redacted for credential-shaped patterns before it reaches the subagent, but redaction is pattern-based and best-effort, not a security boundary.
20
- - Read-before-write tracks only reads through the `skill` tool. `skill_manage` has no per-skill read action (`list`/`review` are whole-library, not targeted at one name), so a skill that was only listed via `skill_manage` is not marked as read: a background review may still reject a patch to it until it is actually loaded.
20
+ - Read-before-write (the rule is `evolution-core`'s `skill-reads.ts`, shared with the tool path's admission gate) tracks only reads through the `skill` tool. `skill_manage` has no per-skill read action (`list`/`review` are whole-library, not targeted at one name), so a skill that was only listed via `skill_manage` is not marked as read: a background review may still reject a patch to it until it is actually loaded.
21
21
  - The completion-channel counters (`cumulativeToolCalls` / `completionInjected`) are in-memory only. A process restart resets them, which is accepted behavior: the completion review is a one-per-session post-task adaptation and a restart is treated as a fresh conversation boundary. The cadence state (`turnsSinceMemory` / `turnsSinceSkill`) is persisted via `ReviewState` and survives restart: bounded by `REVIEW_STATE_SESSION_CAP` (500, seam constant): the least-recently-active sessions are pruned on save, so a very old session restarting resumes from a fresh cadence baseline rather than an unbounded store.
22
22
  - `evolution/review-scheduled` and `evolution/review-error` are emitted for platform/user wiring only: this family has no in-repo production `ctx.on` consumer for them. They are declared externally owned (the platform side wires consumption), which matches the `EXEMPT_ORPHANS` set in `scripts/verify-event-pairing.mjs`.
23
23
  - When the `evolution-state` service is not mounted, the memory/skill cadence state is not persisted and every turn restarts from a clean `{ turnsSinceMemory: 0, turnsSinceSkill: 0 }` baseline: the review schedule is stateless and re-decided each turn rather than accumulating across the conversation. The loss is surfaced once per process as a logger warning at the first turn/end.
24
24
  - Read-before-write can see the review subagent's own `skill` reads only when the subagent backend exposes `localAgent` (the in-process driver does; out-of-process backends such as ACP and the CLI providers set `localAgent: undefined`). With a remote backend the subagent's reads are invisible, so a plan item patching a skill the subagent itself loaded is dropped as "unread": the review then falls back to the parent session's reads only.
25
25
  - **The default `reviewMode: 'inject'` produces no plan ledger.** The review runs in the parent session and emits no `evolution/plan-applied` event, so `evolution-activity`'s `activity.json` and the `evolution-replay` leaderboard never grow in this mode: the only production emit point sits inside the subagent path. Set `reviewMode: 'subagent'` on the `evolution-policy` row when the plan audit trail is required; the plugin states this once at load instead of leaving an empty ledger to be read as "no reviews happened" (v37 S2.2, plan decision (c)).
26
+ - **Evidence that cites only bookkeeping frames is REPORTED, never refused (phase 1).** `EVIDENCE_CLASS` (`evolution-plan-validator`) counts the ops whose entire `evidence` list points at turn/step boundary frames — a seq the range rule accepts but that carries no content — and this plugin logs one `dsh-evolution-review:` warning per plan; the count rides `evolution/plan-applied` as `evidenceClassReports` and appears in no model-visible text. Nothing is dropped: `EVIDENCE_RANGE` remains the only evidence gate, and a session log whose frames carry no `seq` leaves the report silent (`undefined` is not an empty index) instead of reporting every op.
26
27
  - **`.pinned` protection in the `'inject'` channel rides a family-internal session mark.** The review prompt marks the parent session, `tool-skill-manage` reads that mark and resolves both origin surfaces (approval + library) to `background_review`, and the next REAL user message (`source.kind === 'user'`) clears it (plugin notices, including the review prompt itself, do not). Two bounds follow from the platform's inject contract (no driver wake; a prompt is dropped on cancel/dispose and may be missed with an already-claimed batch): with `reviewWakeInject: false`, or on a host without `followup`, (a) a pending prompt may never execute while the session stays idle, and (b) when the user's next message wakes the session, the mark is cleared before that pending prompt reaches the model (its writes are then attributed `foreground` and the pinned guard does not cover them). A prompt delivered while a HUMAN message is already queued is not marked at all: the platform claims that human turn first, so a session-level window would attribute the user's own writes to `background_review` (S2-8, FLOW1-3): the plugin reads the pre-claim queue (`agent.inbox`) because the platform exposes no claim identity, and warns once per mount when it withholds the mark.
27
28
 
28
29
  ## Configuration
package/lib/index.js CHANGED
@@ -2,7 +2,7 @@ import { createHash, randomUUID } from "node:crypto";
2
2
  import z from "@deepseek-ai/schemastery";
3
3
  import { createUserMessage } from "@deepseek-ai/dsh-llm";
4
4
  import { SessionId } from "@deepseek-ai/dsh-session";
5
- import { COMPLETION_SKILL_REVIEW_PROMPT, DEFAULT_MAX_OPS_PER_PLAN, DEFAULT_MEMORY_CHAR_LIMIT, DEFAULT_MEMORY_REVIEW_MODEL, DEFAULT_REVIEW_CONTEXT_MESSAGES, DEFAULT_REVIEW_MEMORY_INTERVAL, DEFAULT_REVIEW_MESSAGE_CHARS, DEFAULT_REVIEW_SKILL_INTERVAL, DEFAULT_REVIEW_TIMEOUT_MS, DEFAULT_SKILL_CONTENT_CHARS, DEFAULT_SKILL_LIMITS, DEFAULT_SKILL_REVIEW_COMPLETION_MIN_TOOL_CALLS, DEFAULT_SKILL_REVIEW_MODEL, DEFAULT_SKILL_REVIEW_TRIGGER, DEFAULT_SUBSTANTIVE_MIN_AGENT_CHARS, DEFAULT_SUBSTANTIVE_MIN_TOOL_CALLS, DEFAULT_SUBSTANTIVE_MIN_USER_CHARS, DEFAULT_USER_CHAR_LIMIT, MAX_TIMER_DELAY_MS, PROMPT_BUNDLE, advanceReview, assertSkillsRootAliasRetired, clampedNumber, clearReviewChannel, contentHash, evolutionIoAdapter, foldToolDispatches, foldTurn, installParamSection, markReviewChannel, newSkillLibrary, paramNamespace, policyStageLimits, readDispatchSignal, readNumberParam, redactSecrets, resolveOrigins, resolveRootConfig, reviewPrompt, sessionAudited, skillReadNameOf, sweepReviewChannelSessions, verifyPromptBundle } from "@lmzhen/dsh-evolution-core";
5
+ import { COMPLETION_SKILL_REVIEW_PROMPT, DEFAULT_MAX_OPS_PER_PLAN, DEFAULT_MEMORY_CHAR_LIMIT, DEFAULT_MEMORY_REVIEW_MODEL, DEFAULT_REVIEW_CONTEXT_MESSAGES, DEFAULT_REVIEW_MEMORY_INTERVAL, DEFAULT_REVIEW_MESSAGE_CHARS, DEFAULT_REVIEW_SKILL_INTERVAL, DEFAULT_REVIEW_TIMEOUT_MS, DEFAULT_SKILL_CONTENT_CHARS, DEFAULT_SKILL_LIMITS, DEFAULT_SKILL_REVIEW_COMPLETION_MIN_TOOL_CALLS, DEFAULT_SKILL_REVIEW_MODEL, DEFAULT_SKILL_REVIEW_TRIGGER, DEFAULT_SUBSTANTIVE_MIN_AGENT_CHARS, DEFAULT_SUBSTANTIVE_MIN_TOOL_CALLS, DEFAULT_SUBSTANTIVE_MIN_USER_CHARS, DEFAULT_USER_CHAR_LIMIT, MAX_TIMER_DELAY_MS, PROMPT_BUNDLE, advanceReview, assertSkillsRootAliasRetired, clampedNumber, clearReviewChannel, collectReadSkillNames, contentHash, evidenceKindIndex, evolutionIoAdapter, filterUnreadSkillOps, filterUnreadSkillOps as filterUnreadSkillOps$1, foldTurn, installParamSection, markReviewChannel, newSkillLibrary, paramNamespace, policyStageLimits, readDispatchSignal, readNumberParam, redactSecrets, resolveOrigins, resolveRootConfig, reviewPrompt, sessionAudited, sweepReviewChannelSessions, verifyPromptBundle } from "@lmzhen/dsh-evolution-core";
6
6
  import { validateEvolutionPlan } from "@lmzhen/dsh-evolution-plan-validator";
7
7
  //#region lib/types/session-state.js
8
8
  /**
@@ -634,6 +634,7 @@ function apply(ctx, rawConfig = {}) {
634
634
  }
635
635
  const preRunHashes = await treeSkillHashes();
636
636
  const sessionSeqAtPlanTime = session.seq - 1;
637
+ const substantiveEvidenceSeqs = evidenceKindIndex(session.snapshotEvents());
637
638
  const run = await subagents.start("spawn", {
638
639
  label: "dsh-evolution-review",
639
640
  prompt: [{
@@ -661,19 +662,21 @@ function apply(ctx, rawConfig = {}) {
661
662
  ctx.logger.warn(`dsh-evolution-review: review subagent returned no structured plan${stopDetail !== void 0 ? ` (stopReason=${stopDetail}${diagDetail !== void 0 ? `; diagnostic=${diagDetail}` : ""})` : ""}`);
662
663
  return false;
663
664
  }
664
- const childReads = run.localAgent ? collectReadSkillNames(run.localAgent.session) : /* @__PURE__ */ new Set();
665
+ const childReads = run.localAgent ? collectReadSkillNames(run.localAgent.session.snapshotEvents()) : /* @__PURE__ */ new Set();
665
666
  const plan = result.structured;
666
667
  const policyFingerprint = fingerprintPolicy(snapshot);
667
668
  const validation = validateEvolutionPlan(plan, {
668
669
  sessionSeq: sessionSeqAtPlanTime,
670
+ substantiveEvidenceSeqs,
669
671
  maxOpsPerPlan: snapshot?.maxOpsPerPlan ?? DEFAULT_MAX_OPS_PER_PLAN,
670
672
  protectedSkillNames: new Set(snapshot?.protectedSkillNames ?? []),
671
673
  maxMemoryChars: snapshot?.memoryChars ?? DEFAULT_MEMORY_CHAR_LIMIT,
672
674
  maxUserChars: snapshot?.userChars ?? DEFAULT_USER_CHAR_LIMIT,
673
675
  maxSkillContentChars: snapshot?.skillContentChars ?? DEFAULT_SKILL_CONTENT_CHARS
674
676
  });
677
+ if (validation.reports.length > 0) ctx.logger.warn(`dsh-evolution-review: ${validation.reports.length} op(s) cite only bookkeeping frames as evidence: ${validation.reports.map((report) => report.reason).join("; ")}`);
675
678
  const acceptedSkillOps = validation.accepted.skillOps ?? [];
676
- const skippedUnread = filterUnreadSkillOps(acceptedSkillOps, new Set([...collectReadSkillNames(session), ...childReads]));
679
+ const skippedUnread = filterUnreadSkillOps$1(acceptedSkillOps, new Set([...collectReadSkillNames(session.snapshotEvents()), ...childReads]));
677
680
  const evidenceQuotes = [...validation.accepted.memoryOps ?? [], ...acceptedSkillOps].reduce((total, op) => total + (Array.isArray(op.evidence) ? op.evidence.length : 0), 0);
678
681
  const emitApplied = (report) => {
679
682
  try {
@@ -685,6 +688,7 @@ function apply(ctx, rawConfig = {}) {
685
688
  skillApplied: report.actions.filter((action) => action.startsWith("Skill ")).length,
686
689
  rejectedOps: validation.rejected.length,
687
690
  ...skippedUnread > 0 ? { skippedUnread } : {},
691
+ ...validation.reports.length > 0 ? { evidenceClassReports: validation.reports.length } : {},
688
692
  executionFailures: report.failedOps?.length ?? 0,
689
693
  ...report.executionError !== void 0 ? { executionError: report.executionError } : {},
690
694
  evidenceQuotes,
@@ -1054,29 +1058,6 @@ function staleRefusal(result, name, filePath) {
1054
1058
  message: filePath === void 0 ? `Skill "${name}" changed since this plan was produced — the full-content update was refused as stale. Re-read the skill and produce a fresh plan.` : `Support file "${filePath}" of "${name}" changed since this plan was produced — the staged file operation was refused as stale. Re-read the skill tree and produce a fresh plan.`
1055
1059
  };
1056
1060
  }
1057
- /**
1058
- * v37 P7a: the read-before-write credit now comes from evolution-core's
1059
- * `tool-dispatch` module — the ONE reader of the platform's dispatch event
1060
- * types, and the ONE authority on which tool reads a skill.
1061
- *
1062
- * v32 REV-06(a) is preserved by the normalizer: a skill counts as READ only
1063
- * when it did not fail, so a failed/timeout read still cannot pass the
1064
- * read-before-write gate and let the review blind-overwrite content the model
1065
- * never saw. What changed is the vocabulary the gate listens to: matching
1066
- * `tool/call` here meant every PTC session (`tool/ptc-dispatch*`) collected an
1067
- * EMPTY set, so `filterUnreadSkillOps` dropped every mutating op the model had
1068
- * legitimately read first — and nothing reported the loss.
1069
- * @param session - the session whose log is folded.
1070
- * @returns the skill names this session read through a non-failed dispatch.
1071
- */
1072
- function collectReadSkillNames(session) {
1073
- const names = /* @__PURE__ */ new Set();
1074
- for (const dispatch of foldToolDispatches(session.snapshotEvents())) {
1075
- const name = skillReadNameOf(dispatch);
1076
- if (name !== void 0) names.add(name);
1077
- }
1078
- return names;
1079
- }
1080
1061
  /** Map/set size that triggers a dead-session counter sweep (bounded, not a hard cap). */
1081
1062
  const COUNTER_SWEEP_THRESHOLD = 128;
1082
1063
  /**
@@ -1094,35 +1075,6 @@ function sweepDeadSessionEntries(entries, isAlive) {
1094
1075
  }
1095
1076
  return removed;
1096
1077
  }
1097
- /**
1098
- * Drop mutating ops whose target was not read this session, in place.
1099
- * Create is exempt (no read required to author a new skill). Covers the same
1100
- * mutating surface Hermes guards (edit/patch/write_file/remove_file), so a
1101
- * background review cannot blind-touch support files or edits of skills it
1102
- * never loaded. Returns the count of dropped ops so the plan event can report
1103
- * them as rejected.
1104
- */
1105
- function filterUnreadSkillOps(ops, readNames) {
1106
- const READ_REQUIRED = [
1107
- "edit",
1108
- "update",
1109
- "patch",
1110
- "delete",
1111
- "write_file",
1112
- "remove_file",
1113
- "restructure"
1114
- ];
1115
- let dropped = 0;
1116
- for (let index = ops.length - 1; index >= 0; index -= 1) {
1117
- const op = ops[index];
1118
- if (!op) continue;
1119
- if (op.name && READ_REQUIRED.includes(op.action ?? "patch") && !readNames.has(op.name)) {
1120
- ops.splice(index, 1);
1121
- dropped += 1;
1122
- }
1123
- }
1124
- return dropped;
1125
- }
1126
1078
  function fingerprintPolicy(snapshot) {
1127
1079
  try {
1128
1080
  return createHash("sha256").update(JSON.stringify(snapshot)).digest("hex").slice(0, 12);
@@ -6,6 +6,7 @@ import type { Context } from '@deepseek-ai/cordis';
6
6
  import z from '@deepseek-ai/schemastery';
7
7
  import type { Session } from '@deepseek-ai/dsh-session';
8
8
  import { type ReviewKind } from '@lmzhen/dsh-evolution-core';
9
+ export { filterUnreadSkillOps } from '@lmzhen/dsh-evolution-core';
9
10
  export declare const name = "evolution-review";
10
11
  export declare const inject: string[];
11
12
  export interface Config {
@@ -26,6 +27,8 @@ export interface Config {
26
27
  /** Shadowed by the policy snapshot in every shipped composition — configure
27
28
  * `reviewMemoryInterval` on the `evolution-policy` row instead (v37 P2-24).
28
29
  * Deprecated alias (G0/S0.3): still readable, refused by writes; removed 0.7.0. */
30
+ /** Review cadence in TOOL CALLS (signals.ts: a turn advances by its tool-call count, minimum 1,
31
+ * or by 1 when the turn itself carried the memory signal). */
29
32
  memoryInterval?: number;
30
33
  /** Shadowed by the policy snapshot in every shipped composition — configure
31
34
  * `reviewSkillInterval` on the `evolution-policy` row instead. Deprecated
@@ -133,9 +136,11 @@ export declare function clampReviewConfig(rawConfig: Config, ctx: Context): Clam
133
136
  * parameter ids from the registry, so the settings document, the params output,
134
137
  * the doctor report and the cards all spell one name. */
135
138
  export interface ReviewSettings {
136
- /** Activity units between skill-review injections. */
139
+ /** Tool calls between skill-review injections (a turn with no tool call counts as one; a turn
140
+ * that itself used a skill advances the counter by 1). */
137
141
  reviewSkillInterval: number;
138
- /** Activity units between memory-review injections. */
142
+ /** Tool calls between memory-review injections (a turn with no tool call counts as one; a turn
143
+ * that itself touched memory advances the counter by 1). */
139
144
  reviewMemoryInterval: number;
140
145
  /** Which channel may inject a skill review. */
141
146
  skillReviewTrigger: 'cadence' | 'completion' | 'both';
@@ -164,18 +169,6 @@ export declare function shouldCompletionReview(reason: {
164
169
  * number of removed entries.
165
170
  */
166
171
  export declare function sweepDeadSessionEntries<K>(entries: Map<K, unknown> | Set<K>, isAlive: (id: K) => boolean): number;
167
- /**
168
- * Drop mutating ops whose target was not read this session, in place.
169
- * Create is exempt (no read required to author a new skill). Covers the same
170
- * mutating surface Hermes guards (edit/patch/write_file/remove_file), so a
171
- * background review cannot blind-touch support files or edits of skills it
172
- * never loaded. Returns the count of dropped ops so the plan event can report
173
- * them as rejected.
174
- */
175
- export declare function filterUnreadSkillOps(ops: Array<{
176
- action?: string;
177
- name?: string;
178
- }>, readNames: ReadonlySet<string>): number;
179
172
  /**
180
173
  * V10-10 (P2-11) / A1 (audit P1-1): render one `[result]` evidence line from
181
174
  * a tool-result event payload. The former read (`data.output`) targeted a
@@ -204,5 +197,4 @@ export declare function buildReviewRequest(session: Session, kind: ReviewKind, s
204
197
  userChars: number;
205
198
  assistantChars: number;
206
199
  }, maxMessages: number, maxMessageChars: number): string;
207
- export {};
208
200
  //# sourceMappingURL=index.d.ts.map
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "@lmzhen/dsh-evolution-review",
3
3
  "description": "Background review orchestration (community build)",
4
- "version": "0.8.0",
4
+ "version": "0.10.0",
5
5
  "publishConfig": {
6
6
  "access": "public"
7
7
  },
@@ -27,9 +27,9 @@
27
27
  "license": "MIT",
28
28
  "dependencies": {
29
29
  "@deepseek-ai/schemastery": "^3.18.1",
30
- "@lmzhen/dsh-evolution-approval": "^0.8.0",
31
- "@lmzhen/dsh-evolution-core": "^0.8.0",
32
- "@lmzhen/dsh-evolution-plan-validator": "^0.8.0"
30
+ "@lmzhen/dsh-evolution-approval": "^0.10.0",
31
+ "@lmzhen/dsh-evolution-core": "^0.10.0",
32
+ "@lmzhen/dsh-evolution-plan-validator": "^0.10.0"
33
33
  },
34
34
  "peerDependencies": {
35
35
  "@deepseek-ai/cordis": "^4.0.1",
@@ -37,8 +37,8 @@
37
37
  "@deepseek-ai/dsh-llm": "^0.1.5-rc.2",
38
38
  "@deepseek-ai/dsh-session": "^0.1.5-rc.2",
39
39
  "@deepseek-ai/dsh-tools": "^0.1.5-rc.2",
40
- "@lmzhen/dsh-evolution-state": "^0.8.0",
41
- "@lmzhen/dsh-evolution-policy": "^0.8.0"
40
+ "@lmzhen/dsh-evolution-state": "^0.10.0",
41
+ "@lmzhen/dsh-evolution-policy": "^0.10.0"
42
42
  },
43
43
  "devDependencies": {
44
44
  "@deepseek-ai/dsh-agent": "^0.1.5-rc.2",
@@ -48,10 +48,10 @@
48
48
  "@deepseek-ai/dsh-session-persistence": "^0.1.5-rc.2",
49
49
  "@deepseek-ai/dsh-session-persistence-jsonl": "^0.1.5-rc.2",
50
50
  "@deepseek-ai/dsh-tools": "^0.1.5-rc.2",
51
- "@lmzhen/dsh-evolution-approval": "^0.8.0",
52
- "@lmzhen/dsh-evolution-core": "^0.8.0",
53
- "@lmzhen/dsh-evolution-curator": "^0.8.0",
54
- "@lmzhen/dsh-evolution-plan-validator": "^0.8.0",
55
- "@lmzhen/dsh-evolution-state": "^0.8.0"
51
+ "@lmzhen/dsh-evolution-approval": "^0.10.0",
52
+ "@lmzhen/dsh-evolution-core": "^0.10.0",
53
+ "@lmzhen/dsh-evolution-curator": "^0.10.0",
54
+ "@lmzhen/dsh-evolution-plan-validator": "^0.10.0",
55
+ "@lmzhen/dsh-evolution-state": "^0.10.0"
56
56
  }
57
57
  }