specrails-desktop 2.44.1 → 2.44.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,70 @@
1
+ # Verification handoff and framework activation regression
2
+
3
+ The September 10, 2026 implementation incident ended with an agent reporting
4
+ success while Core reported `implementation=complete`, `validation=blocked`,
5
+ `archive=done`, and `delivery=pending-host`. The repository's 127 tests passed;
6
+ inspection rejected the receipt with `Verification environment changed: npm`.
7
+
8
+ A second Windows execution reproduced the same terminal contradiction after
9
+ the coordinator reported successful unit and browser checks, completed archive,
10
+ and skipped host-owned shipping/CI phases. It refreshed verification and moved
11
+ archive execution into the coordinator's session to avoid handoff drift, but
12
+ the host still reported blocked validation. This corroborates the cross-process
13
+ failure pattern; the transcript alone does not identify the differing variables
14
+ or establish a cloud-synced project path as the cause. Runtime diagnostics now
15
+ report added/removed environment key names without values. If only values differ,
16
+ the aggregate hash cannot identify the individual key.
17
+
18
+ Three independent problems made recovery unreliable:
19
+
20
+ 1. The installed app bundled Core 5.2.2, but `framework/current` still resolved to
21
+ 5.2.1. Desktop materialized the providers in its project catalog, omitting Kimi
22
+ already represented in `current`. Core correctly refused to activate the
23
+ incomplete destination. The session-identity fix in 5.2.2 was never active.
24
+ 2. Verification included agent-session metadata and automatically generated
25
+ untracked memory in its evidence. A process handoff or a final memory write
26
+ could invalidate verification after a successful check. Ignoring metadata in
27
+ the fingerprint alone also left those unrecorded inputs visible to tests.
28
+ 3. A successful provider process could become a successful implementation step
29
+ without a final deterministic check of Core's state.
30
+
31
+ ## Required invariants
32
+
33
+ - Framework materialization must satisfy the same provider union as Core's
34
+ activation gate: requested providers, the shared registry, and provider
35
+ directories or completion stamps represented in `current`. Failure leaves the
36
+ previous version active.
37
+ - Core removes known session/launcher metadata from verification subprocesses
38
+ as well as from their environment fingerprint. Explicit command overrides are
39
+ retained and bound without storing their values. Application configuration,
40
+ PATH, NODE_OPTIONS, custom variables and SPECRAILS configuration still count.
41
+ Windows environment names are normalized without changing POSIX semantics.
42
+ - Only untracked files under known provider `agent-memory/` roots are excluded
43
+ as generated notes. Tracked memory, settings, skills and other source inputs
44
+ remain part of the candidate.
45
+ - Old environment-policy receipts require a fresh full verification. Do not
46
+ edit receipts or phase records to turn an old failure into success.
47
+ - Desktop checks installed Core after Implement/Batch steps, for both interactive
48
+ and one-shot execution. Blocked validation, incomplete phases, an invalid full
49
+ receipt, a failed status call or mismatched run identity prevents success.
50
+ `pending-host` describes delivery ownership and does not invalidate otherwise
51
+ verified implementation work.
52
+
53
+ ## Regression coverage and rollout
54
+
55
+ Core's `pipeline-state.test.ts` exercises real subprocess handoffs, archive
56
+ completion followed by a host inspection, explicit overrides, environment/input
57
+ changes, generated versus tracked memory and Windows environment normalization.
58
+ Desktop's `loop-core-completion.test.ts` runs the completion-status subprocess
59
+ through the loop engine and verifies the stored result. Framework tests cover
60
+ the omitted-provider activation failure and preservation of the previous version.
61
+ An offline test with the actual 5.2.1 and bundled 5.2.2 packages also validated
62
+ the provider-union update in an isolated home.
63
+
64
+ Activating the existing 5.2.2 bundle repairs the missed previous update. The
65
+ additional environment, memory, completion and materialization changes require
66
+ new Core and Desktop releases. These changes have been ported onto Core 5.2.2
67
+ and Desktop 2.44.1 for review. Never replace a newer installed runtime with an
68
+ older checkout build, and verify both
69
+ the selected package version and the actual `framework/current` target after
70
+ activation. Refresh copied workspace files through the normal reseed lifecycle.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "specrails-desktop",
3
- "version": "2.44.1",
3
+ "version": "2.44.2",
4
4
  "license": "MIT",
5
5
  "repository": {
6
6
  "type": "git",
@@ -3,7 +3,7 @@ var __importDefault = (this && this.__importDefault) || function (mod) {
3
3
  return (mod && mod.__esModule) ? mod : { "default": mod };
4
4
  };
5
5
  Object.defineProperty(exports, "__esModule", { value: true });
6
- exports.WINDOWS_BACKGROUND_BOOTSTRAP = void 0;
6
+ exports.WINDOWS_BACKGROUND_BOOTSTRAP = exports.WINDOWS_JOB_PREPARATION_TIMEOUT_MS = void 0;
7
7
  exports.spawnWindowsBackgroundBootstrap = spawnWindowsBackgroundBootstrap;
8
8
  const child_process_1 = require("child_process");
9
9
  const node_net_1 = require("node:net");
@@ -14,6 +14,9 @@ const node_crypto_1 = require("node:crypto");
14
14
  const path_resolver_1 = require("./path-resolver");
15
15
  const win_spawn_1 = require("./util/win-spawn");
16
16
  const windows_job_supervisor_1 = require("./windows-job-supervisor");
17
+ // PowerShell startup and Add-Type compilation can exceed 15s on cold Windows
18
+ // ARM runners. Keep preparation bounded without admitting any application code.
19
+ exports.WINDOWS_JOB_PREPARATION_TIMEOUT_MS = 60_000;
17
20
  // Keep the owned Windows root alive until its creation identity is captured.
18
21
  // User commands never enter argv and cannot start before the parent opens stdin.
19
22
  exports.WINDOWS_BACKGROUND_BOOTSTRAP = String.raw `
@@ -130,12 +133,12 @@ function spawnWindowsBackgroundBootstrap(command, cwd) {
130
133
  }
131
134
  catch { /* no child */ } });
132
135
  const startupTimer = setTimeout(() => {
133
- fail(new Error('Windows job containment preparation timed out; no application was admitted.'));
136
+ fail(new Error(`Windows job containment preparation timed out after ${exports.WINDOWS_JOB_PREPARATION_TIMEOUT_MS}ms (${socket ? 'control pipe connected, awaiting job assignment' : 'awaiting supervisor control pipe'}); no application was admitted.`));
134
137
  try {
135
138
  child?.kill();
136
139
  }
137
140
  catch { /* already gone */ }
138
- }, 15_000);
141
+ }, exports.WINDOWS_JOB_PREPARATION_TIMEOUT_MS);
139
142
  startupTimer.unref?.();
140
143
  const cleanup = () => {
141
144
  clearTimeout(startupTimer);
@@ -2,6 +2,7 @@
2
2
  Object.defineProperty(exports, "__esModule", { value: true });
3
3
  exports.prepareCoreExecution = prepareCoreExecution;
4
4
  exports.coreVerificationContext = coreVerificationContext;
5
+ exports.checkCoreCompletion = checkCoreCompletion;
5
6
  const node_child_process_1 = require("node:child_process");
6
7
  const node_fs_1 = require("node:fs");
7
8
  const node_path_1 = require("node:path");
@@ -110,3 +111,47 @@ function coreVerificationContext(contextPath, cwd, env, runId) {
110
111
  return '';
111
112
  }
112
113
  }
114
+ /** A successful provider turn is not a completed implementation. Ask Core to
115
+ * revalidate the journal on disk before Desktop accepts the implementation step,
116
+ * including when an agent claims PASS after its final receipt went stale. */
117
+ function checkCoreCompletion(contextPath, cwd, env, runId) {
118
+ const helper = (0, node_path_1.join)(cwd, '.specrails', 'runtime', 'pipeline.mjs');
119
+ if (!(0, node_fs_1.existsSync)(helper))
120
+ return { valid: false, reason: 'Core completion cannot be validated: pipeline runtime is missing' };
121
+ try {
122
+ const stdout = (0, node_child_process_1.execFileSync)((0, path_resolver_1.resolveBundledNodeExe)() ?? process.execPath, [helper, 'status', '--context', contextPath], {
123
+ cwd, env, encoding: 'utf8', timeout: 15_000, maxBuffer: 4 * 1024 * 1024, stdio: ['ignore', 'pipe', 'pipe'],
124
+ });
125
+ const status = JSON.parse(stdout);
126
+ if (status.schemaVersion !== 1 || status.runId !== runId)
127
+ return { valid: false, reason: 'Core completion returned an invalid or mismatched run identity' };
128
+ const reasons = [];
129
+ // Core 5.1 exposes phases/verification; newer runtimes also provide an
130
+ // explicit completion verdict. Never override a newer blocked verdict with
131
+ // older phase statuses or an agent-authored summary.
132
+ if (status.completion) {
133
+ for (const [key, expected] of [['implementation', 'complete'], ['archive', 'done']]) {
134
+ if (status.completion[key] !== expected)
135
+ reasons.push(`${key}=${status.completion[key] ?? 'missing'}`);
136
+ }
137
+ if (!['verified', 'with-exceptions'].includes(status.completion.validation ?? ''))
138
+ reasons.push(`validation=${status.completion.validation ?? 'missing'}`);
139
+ if (reasons.length)
140
+ reasons.push(...(status.completion.reasons ?? []));
141
+ }
142
+ for (const phase of ['architect', 'developer', 'reviewer', 'archive']) {
143
+ if (status.phases?.[phase]?.status !== 'done')
144
+ reasons.push(`${phase} phase is incomplete`);
145
+ }
146
+ if (status.resumePhase && ['architect', 'developer', 'reviewer', 'archive'].includes(status.resumePhase))
147
+ reasons.push(`resume from ${status.resumePhase}`);
148
+ const receipt = status.verification?.receipt;
149
+ if (status.verification?.valid !== true || receipt?.kind !== 'full' || !receipt.commands?.length || receipt.commands.some(command => command.exitCode !== 0)) {
150
+ reasons.push(...(status.verification?.reasons?.length ? status.verification.reasons : ['No valid full verification receipt']));
151
+ }
152
+ return reasons.length ? { valid: false, reason: `Core completion blocked: ${[...new Set(reasons)].join('; ')}` } : { valid: true };
153
+ }
154
+ catch {
155
+ return { valid: false, reason: 'Core completion cannot be validated: pipeline status failed or returned invalid JSON' };
156
+ }
157
+ }
@@ -44,6 +44,28 @@ function realpathOrSelf(p) {
44
44
  return p;
45
45
  }
46
46
  }
47
+ /** Match Core's swap-current provider requirement: a project removed from the
48
+ * Desktop catalog can still have live links through current. A subtree OR its
49
+ * completion stamp is enough to require that provider in the next version. */
50
+ function requiredFrameworkProviders(fwDir, requested, home) {
51
+ const directories = {
52
+ claude: '.claude', codex: '.codex', gemini: '.gemini', kimi: '.kimi-code',
53
+ };
54
+ const required = new Set(requested);
55
+ for (const entry of Object.values((0, artifact_registry_1.readRegistryOrEmpty)(home).projects)) {
56
+ for (const provider of [...(entry.providers ?? []), entry.primaryProvider]) {
57
+ if (Object.hasOwn(directories, provider))
58
+ required.add(provider);
59
+ }
60
+ }
61
+ for (const [provider, directory] of Object.entries(directories)) {
62
+ if (fs_1.default.existsSync(path_1.default.join(fwDir, 'current', directory))
63
+ || fs_1.default.existsSync(path_1.default.join(fwDir, 'current', `.framework-stamp${directory}.json`))) {
64
+ required.add(provider);
65
+ }
66
+ }
67
+ return [...required];
68
+ }
47
69
  class FrameworkManager {
48
70
  home;
49
71
  broadcast;
@@ -101,7 +123,7 @@ class FrameworkManager {
101
123
  fs_1.default.mkdirSync(fwDir, { recursive: true });
102
124
  const errors = [];
103
125
  const done = [];
104
- for (const provider of dedupe(providers)) {
126
+ for (const provider of requiredFrameworkProviders(fwDir, providers, this.home)) {
105
127
  const res = (0, child_process_1.spawnSync)(nodeInterpreter(), [
106
128
  cli,
107
129
  'install-framework',
@@ -357,6 +379,3 @@ class FrameworkManager {
357
379
  }
358
380
  }
359
381
  exports.FrameworkManager = FrameworkManager;
360
- function dedupe(values) {
361
- return Array.from(new Set(values));
362
- }
@@ -489,6 +489,12 @@ function createLoopExecutors(opts = {}) {
489
489
  ...(effectiveIdleTimeoutMs > 0 ? { idleTimeoutMs: effectiveIdleTimeoutMs } : {}),
490
490
  };
491
491
  },
492
+ validateCoreCompletion({ coreRun, provider, model, effort, profileName, cwd, repoDir, executionManifest }) {
493
+ const baseStepEnv = withProfileEnv(aiStepEnv(resolveEnv(), repoDir, executionManifest), provider, profileName);
494
+ const core = (0, core_execution_1.prepareCoreExecution)({ run: coreRun, cwd, repoDir, manifest: executionManifest, env: baseStepEnv });
495
+ const env = (0, runtime_1.buildProviderEnv)((0, providers_1.getAdapter)(provider), { prompt: '', model, reasoning_effort: effort }, core.env);
496
+ return (0, core_execution_1.checkCoreCompletion)(core.contextPath, cwd, env, coreRun.runId);
497
+ },
492
498
  async runShell({ command, cwd, repoDir, onLine, onSpawn, timeoutMs }) {
493
499
  const env = resolveEnv();
494
500
  return runShellCommand(command, repoDir ?? cwd, repoDir ? { ...env, SPECRAILS_REPO_DIR: repoDir } : env, timeoutMs ?? SHELL_TIMEOUT_MS, onLine, onSpawn);
@@ -1058,6 +1058,8 @@ class LoopRunManager {
1058
1058
  // --yes`, codex `$implement #<id> --yes`) — then resolve `{{spec.*}}`
1059
1059
  // data tokens and finally `{{const:*}}` library constants.
1060
1060
  const rawTemplate = String(node.data?.prompt ?? '');
1061
+ const requiresCoreCompletion = /\{\{\s*cmd:(?:implement|batch)\s*\}\}/.test(rawTemplate)
1062
+ || /^\s*(?:\/specrails:|\/skill:specrails-|\$)(?:implement|batch-implement)(?:\s|$)/.test(rawTemplate);
1061
1063
  const expanded = (0, loop_constants_1.resolveConstants)(resolveRunVars((0, loop_graph_1.interpolateSpec)((0, loop_command_catalog_1.expandCommands)(rawTemplate, { provider: nodeProvider, ticketIds: req.spec?.ticketIds, specId: req.spec?.id }), req.spec), runVars), constMap);
1062
1064
  const base = [(0, multi_repo_execution_store_1.executionManifestPrompt)(req.executionManifest), withReviewContinuationContext(expanded, rawTemplate, req.spec)].filter(Boolean).join('\n\n');
1063
1065
  // Inject the cross-iteration history only when there's no live session
@@ -1220,6 +1222,19 @@ class LoopRunManager {
1220
1222
  settled = true;
1221
1223
  break;
1222
1224
  }
1225
+ if (!res.failed && !zeroWork && !this._cancelled.has(runId)
1226
+ && requiresCoreCompletion
1227
+ && this.executors.validateCoreCompletion) {
1228
+ const completion = this.executors.validateCoreCompletion({
1229
+ coreRun: { runId, spec: req.spec, goal: req.spec ? undefined : JSON.stringify({ loop: req.loopName, graph: req.graph }), repositoryId: req.repositoryId },
1230
+ provider: nodeProvider, model: nodeModel, effort: nodeEffort, profileName: req.profileName,
1231
+ cwd: req.cwd, repoDir: req.repoDir, executionManifest: req.executionManifest,
1232
+ });
1233
+ if (!completion.valid) {
1234
+ res = { ...res, failed: true, errorText: completion.reason ?? 'Core implementation is incomplete' };
1235
+ logLine(`✖ ${res.errorText}`, 'stderr');
1236
+ }
1237
+ }
1223
1238
  const blockedReason = aiStepBlockedReason(res.text);
1224
1239
  const requiresVerification = node.data?.requireVerificationPass === true
1225
1240
  || /\{\{cmd:(?:verify|revision-verify|opsx:verify)\}\}/.test(rawTemplate);
@@ -1613,8 +1628,6 @@ class LoopRunManager {
1613
1628
  catch (err) {
1614
1629
  console.error(`[loop] staged accounting reconciliation failed for ${runId}:`, err);
1615
1630
  }
1616
- console.log(`[loop] settle run=${runId} outcome=${outcome} iterations=${iteration} ` +
1617
- (usageTelemetryAvailable ? `cost=$${totalCost.toFixed(4)}` : 'usage=unavailable'));
1618
1631
  // Execution and acceptance are separate facts. Persist the runtime's own
1619
1632
  // terminal snapshot alongside counters so history does not depend on prose.
1620
1633
  let coreCompletion = null;
@@ -1625,6 +1638,8 @@ class LoopRunManager {
1625
1638
  const executionOutcome = outcome;
1626
1639
  if (outcome === 'success' && coreCompletion && (coreCompletion.completion.implementation !== 'complete' || !['verified', 'with-exceptions'].includes(coreCompletion.completion.validation)))
1627
1640
  outcome = 'blocked';
1641
+ console.log(`[loop] settle run=${runId} outcome=${outcome} iterations=${iteration} ` +
1642
+ (usageTelemetryAvailable ? `cost=$${totalCost.toFixed(4)}` : 'usage=unavailable'));
1628
1643
  emitRunEvent('loop_completion', {
1629
1644
  version: 1, execution: executionOutcome, steps: stepNum, deciderEvaluations: iteration,
1630
1645
  turns: finalJobUsage.numTurns, costUsd: usageTelemetryAvailable ? totalCost : null,
@@ -1642,7 +1657,7 @@ class LoopRunManager {
1642
1657
  // `≥` when any cost-bearing step ended unpriced (timeout/crash) — the figure
1643
1658
  // is a lower bound, not exact. Providers without usage telemetry get an
1644
1659
  // explicit unavailable marker, never a fabricated "$0.0000".
1645
- logLine(`\n■ Loop execution finished: ${executionOutcome} — ${stepNum} step${stepNum === 1 ? '' : 's'}, ${iteration} decider evaluation${iteration === 1 ? '' : 's'}, ${finalJobUsage.numTurns ?? 'unknown'} agent turns, ` +
1660
+ logLine(`\n■ Loop execution finished: ${outcome} — ${stepNum} step${stepNum === 1 ? '' : 's'}, ${iteration} decider evaluation${iteration === 1 ? '' : 's'}, ${finalJobUsage.numTurns ?? 'unknown'} agent turns, ` +
1646
1661
  (usageTelemetryAvailable
1647
1662
  ? `${costUncertain ? '≥ ' : ''}$${totalCost.toFixed(4)}`
1648
1663
  : 'usage/cost unavailable'));