specrails-desktop 2.44.1 → 2.44.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/docs/internals/verification-handoff-regression.md +70 -0
- package/package.json +1 -1
- package/server/dist/background-windows-bootstrap.js +6 -3
- package/server/dist/core-execution.js +45 -0
- package/server/dist/framework-manager.js +23 -4
- package/server/dist/loop-executors.js +6 -0
- package/server/dist/loop-run-manager.js +18 -3
|
@@ -0,0 +1,70 @@
|
|
|
1
|
+
# Verification handoff and framework activation regression
|
|
2
|
+
|
|
3
|
+
The September 10, 2026 implementation incident ended with an agent reporting
|
|
4
|
+
success while Core reported `implementation=complete`, `validation=blocked`,
|
|
5
|
+
`archive=done`, and `delivery=pending-host`. The repository's 127 tests passed;
|
|
6
|
+
inspection rejected the receipt with `Verification environment changed: npm`.
|
|
7
|
+
|
|
8
|
+
A second Windows execution reproduced the same terminal contradiction after
|
|
9
|
+
the coordinator reported successful unit and browser checks, completed archive,
|
|
10
|
+
and skipped host-owned shipping/CI phases. It refreshed verification and moved
|
|
11
|
+
archive execution into the coordinator's session to avoid handoff drift, but
|
|
12
|
+
the host still reported blocked validation. This corroborates the cross-process
|
|
13
|
+
failure pattern; the transcript alone does not identify the differing variables
|
|
14
|
+
or establish a cloud-synced project path as the cause. Runtime diagnostics now
|
|
15
|
+
report added/removed environment key names without values. If only values differ,
|
|
16
|
+
the aggregate hash cannot identify the individual key.
|
|
17
|
+
|
|
18
|
+
Three independent problems made recovery unreliable:
|
|
19
|
+
|
|
20
|
+
1. The installed app bundled Core 5.2.2, but `framework/current` still resolved to
|
|
21
|
+
5.2.1. Desktop materialized the providers in its project catalog, omitting Kimi
|
|
22
|
+
already represented in `current`. Core correctly refused to activate the
|
|
23
|
+
incomplete destination. The session-identity fix in 5.2.2 was never active.
|
|
24
|
+
2. Verification included agent-session metadata and automatically generated
|
|
25
|
+
untracked memory in its evidence. A process handoff or a final memory write
|
|
26
|
+
could invalidate verification after a successful check. Ignoring metadata in
|
|
27
|
+
the fingerprint alone also left those unrecorded inputs visible to tests.
|
|
28
|
+
3. A successful provider process could become a successful implementation step
|
|
29
|
+
without a final deterministic check of Core's state.
|
|
30
|
+
|
|
31
|
+
## Required invariants
|
|
32
|
+
|
|
33
|
+
- Framework materialization must satisfy the same provider union as Core's
|
|
34
|
+
activation gate: requested providers, the shared registry, and provider
|
|
35
|
+
directories or completion stamps represented in `current`. Failure leaves the
|
|
36
|
+
previous version active.
|
|
37
|
+
- Core removes known session/launcher metadata from verification subprocesses
|
|
38
|
+
as well as from their environment fingerprint. Explicit command overrides are
|
|
39
|
+
retained and bound without storing their values. Application configuration,
|
|
40
|
+
PATH, NODE_OPTIONS, custom variables and SPECRAILS configuration still count.
|
|
41
|
+
Windows environment names are normalized without changing POSIX semantics.
|
|
42
|
+
- Only untracked files under known provider `agent-memory/` roots are excluded
|
|
43
|
+
as generated notes. Tracked memory, settings, skills and other source inputs
|
|
44
|
+
remain part of the candidate.
|
|
45
|
+
- Old environment-policy receipts require a fresh full verification. Do not
|
|
46
|
+
edit receipts or phase records to turn an old failure into success.
|
|
47
|
+
- Desktop checks installed Core after Implement/Batch steps, for both interactive
|
|
48
|
+
and one-shot execution. Blocked validation, incomplete phases, an invalid full
|
|
49
|
+
receipt, a failed status call or mismatched run identity prevents success.
|
|
50
|
+
`pending-host` describes delivery ownership and does not invalidate otherwise
|
|
51
|
+
verified implementation work.
|
|
52
|
+
|
|
53
|
+
## Regression coverage and rollout
|
|
54
|
+
|
|
55
|
+
Core's `pipeline-state.test.ts` exercises real subprocess handoffs, archive
|
|
56
|
+
completion followed by a host inspection, explicit overrides, environment/input
|
|
57
|
+
changes, generated versus tracked memory and Windows environment normalization.
|
|
58
|
+
Desktop's `loop-core-completion.test.ts` runs the completion-status subprocess
|
|
59
|
+
through the loop engine and verifies the stored result. Framework tests cover
|
|
60
|
+
the omitted-provider activation failure and preservation of the previous version.
|
|
61
|
+
An offline test with the actual 5.2.1 and bundled 5.2.2 packages also validated
|
|
62
|
+
the provider-union update in an isolated home.
|
|
63
|
+
|
|
64
|
+
Activating the existing 5.2.2 bundle repairs the missed previous update. The
|
|
65
|
+
additional environment, memory, completion and materialization changes require
|
|
66
|
+
new Core and Desktop releases. These changes have been ported onto Core 5.2.2
|
|
67
|
+
and Desktop 2.44.1 for review. Never replace a newer installed runtime with an
|
|
68
|
+
older checkout build, and verify both
|
|
69
|
+
the selected package version and the actual `framework/current` target after
|
|
70
|
+
activation. Refresh copied workspace files through the normal reseed lifecycle.
|
package/package.json
CHANGED
|
@@ -3,7 +3,7 @@ var __importDefault = (this && this.__importDefault) || function (mod) {
|
|
|
3
3
|
return (mod && mod.__esModule) ? mod : { "default": mod };
|
|
4
4
|
};
|
|
5
5
|
Object.defineProperty(exports, "__esModule", { value: true });
|
|
6
|
-
exports.WINDOWS_BACKGROUND_BOOTSTRAP = void 0;
|
|
6
|
+
exports.WINDOWS_BACKGROUND_BOOTSTRAP = exports.WINDOWS_JOB_PREPARATION_TIMEOUT_MS = void 0;
|
|
7
7
|
exports.spawnWindowsBackgroundBootstrap = spawnWindowsBackgroundBootstrap;
|
|
8
8
|
const child_process_1 = require("child_process");
|
|
9
9
|
const node_net_1 = require("node:net");
|
|
@@ -14,6 +14,9 @@ const node_crypto_1 = require("node:crypto");
|
|
|
14
14
|
const path_resolver_1 = require("./path-resolver");
|
|
15
15
|
const win_spawn_1 = require("./util/win-spawn");
|
|
16
16
|
const windows_job_supervisor_1 = require("./windows-job-supervisor");
|
|
17
|
+
// PowerShell startup and Add-Type compilation can exceed 15s on cold Windows
|
|
18
|
+
// ARM runners. Keep preparation bounded without admitting any application code.
|
|
19
|
+
exports.WINDOWS_JOB_PREPARATION_TIMEOUT_MS = 60_000;
|
|
17
20
|
// Keep the owned Windows root alive until its creation identity is captured.
|
|
18
21
|
// User commands never enter argv and cannot start before the parent opens stdin.
|
|
19
22
|
exports.WINDOWS_BACKGROUND_BOOTSTRAP = String.raw `
|
|
@@ -130,12 +133,12 @@ function spawnWindowsBackgroundBootstrap(command, cwd) {
|
|
|
130
133
|
}
|
|
131
134
|
catch { /* no child */ } });
|
|
132
135
|
const startupTimer = setTimeout(() => {
|
|
133
|
-
fail(new Error(
|
|
136
|
+
fail(new Error(`Windows job containment preparation timed out after ${exports.WINDOWS_JOB_PREPARATION_TIMEOUT_MS}ms (${socket ? 'control pipe connected, awaiting job assignment' : 'awaiting supervisor control pipe'}); no application was admitted.`));
|
|
134
137
|
try {
|
|
135
138
|
child?.kill();
|
|
136
139
|
}
|
|
137
140
|
catch { /* already gone */ }
|
|
138
|
-
},
|
|
141
|
+
}, exports.WINDOWS_JOB_PREPARATION_TIMEOUT_MS);
|
|
139
142
|
startupTimer.unref?.();
|
|
140
143
|
const cleanup = () => {
|
|
141
144
|
clearTimeout(startupTimer);
|
|
@@ -2,6 +2,7 @@
|
|
|
2
2
|
Object.defineProperty(exports, "__esModule", { value: true });
|
|
3
3
|
exports.prepareCoreExecution = prepareCoreExecution;
|
|
4
4
|
exports.coreVerificationContext = coreVerificationContext;
|
|
5
|
+
exports.checkCoreCompletion = checkCoreCompletion;
|
|
5
6
|
const node_child_process_1 = require("node:child_process");
|
|
6
7
|
const node_fs_1 = require("node:fs");
|
|
7
8
|
const node_path_1 = require("node:path");
|
|
@@ -110,3 +111,47 @@ function coreVerificationContext(contextPath, cwd, env, runId) {
|
|
|
110
111
|
return '';
|
|
111
112
|
}
|
|
112
113
|
}
|
|
114
|
+
/** A successful provider turn is not a completed implementation. Ask Core to
|
|
115
|
+
* revalidate the journal on disk before Desktop accepts the implementation step,
|
|
116
|
+
* including when an agent claims PASS after its final receipt went stale. */
|
|
117
|
+
function checkCoreCompletion(contextPath, cwd, env, runId) {
|
|
118
|
+
const helper = (0, node_path_1.join)(cwd, '.specrails', 'runtime', 'pipeline.mjs');
|
|
119
|
+
if (!(0, node_fs_1.existsSync)(helper))
|
|
120
|
+
return { valid: false, reason: 'Core completion cannot be validated: pipeline runtime is missing' };
|
|
121
|
+
try {
|
|
122
|
+
const stdout = (0, node_child_process_1.execFileSync)((0, path_resolver_1.resolveBundledNodeExe)() ?? process.execPath, [helper, 'status', '--context', contextPath], {
|
|
123
|
+
cwd, env, encoding: 'utf8', timeout: 15_000, maxBuffer: 4 * 1024 * 1024, stdio: ['ignore', 'pipe', 'pipe'],
|
|
124
|
+
});
|
|
125
|
+
const status = JSON.parse(stdout);
|
|
126
|
+
if (status.schemaVersion !== 1 || status.runId !== runId)
|
|
127
|
+
return { valid: false, reason: 'Core completion returned an invalid or mismatched run identity' };
|
|
128
|
+
const reasons = [];
|
|
129
|
+
// Core 5.1 exposes phases/verification; newer runtimes also provide an
|
|
130
|
+
// explicit completion verdict. Never override a newer blocked verdict with
|
|
131
|
+
// older phase statuses or an agent-authored summary.
|
|
132
|
+
if (status.completion) {
|
|
133
|
+
for (const [key, expected] of [['implementation', 'complete'], ['archive', 'done']]) {
|
|
134
|
+
if (status.completion[key] !== expected)
|
|
135
|
+
reasons.push(`${key}=${status.completion[key] ?? 'missing'}`);
|
|
136
|
+
}
|
|
137
|
+
if (!['verified', 'with-exceptions'].includes(status.completion.validation ?? ''))
|
|
138
|
+
reasons.push(`validation=${status.completion.validation ?? 'missing'}`);
|
|
139
|
+
if (reasons.length)
|
|
140
|
+
reasons.push(...(status.completion.reasons ?? []));
|
|
141
|
+
}
|
|
142
|
+
for (const phase of ['architect', 'developer', 'reviewer', 'archive']) {
|
|
143
|
+
if (status.phases?.[phase]?.status !== 'done')
|
|
144
|
+
reasons.push(`${phase} phase is incomplete`);
|
|
145
|
+
}
|
|
146
|
+
if (status.resumePhase && ['architect', 'developer', 'reviewer', 'archive'].includes(status.resumePhase))
|
|
147
|
+
reasons.push(`resume from ${status.resumePhase}`);
|
|
148
|
+
const receipt = status.verification?.receipt;
|
|
149
|
+
if (status.verification?.valid !== true || receipt?.kind !== 'full' || !receipt.commands?.length || receipt.commands.some(command => command.exitCode !== 0)) {
|
|
150
|
+
reasons.push(...(status.verification?.reasons?.length ? status.verification.reasons : ['No valid full verification receipt']));
|
|
151
|
+
}
|
|
152
|
+
return reasons.length ? { valid: false, reason: `Core completion blocked: ${[...new Set(reasons)].join('; ')}` } : { valid: true };
|
|
153
|
+
}
|
|
154
|
+
catch {
|
|
155
|
+
return { valid: false, reason: 'Core completion cannot be validated: pipeline status failed or returned invalid JSON' };
|
|
156
|
+
}
|
|
157
|
+
}
|
|
@@ -44,6 +44,28 @@ function realpathOrSelf(p) {
|
|
|
44
44
|
return p;
|
|
45
45
|
}
|
|
46
46
|
}
|
|
47
|
+
/** Match Core's swap-current provider requirement: a project removed from the
|
|
48
|
+
* Desktop catalog can still have live links through current. A subtree OR its
|
|
49
|
+
* completion stamp is enough to require that provider in the next version. */
|
|
50
|
+
function requiredFrameworkProviders(fwDir, requested, home) {
|
|
51
|
+
const directories = {
|
|
52
|
+
claude: '.claude', codex: '.codex', gemini: '.gemini', kimi: '.kimi-code',
|
|
53
|
+
};
|
|
54
|
+
const required = new Set(requested);
|
|
55
|
+
for (const entry of Object.values((0, artifact_registry_1.readRegistryOrEmpty)(home).projects)) {
|
|
56
|
+
for (const provider of [...(entry.providers ?? []), entry.primaryProvider]) {
|
|
57
|
+
if (Object.hasOwn(directories, provider))
|
|
58
|
+
required.add(provider);
|
|
59
|
+
}
|
|
60
|
+
}
|
|
61
|
+
for (const [provider, directory] of Object.entries(directories)) {
|
|
62
|
+
if (fs_1.default.existsSync(path_1.default.join(fwDir, 'current', directory))
|
|
63
|
+
|| fs_1.default.existsSync(path_1.default.join(fwDir, 'current', `.framework-stamp${directory}.json`))) {
|
|
64
|
+
required.add(provider);
|
|
65
|
+
}
|
|
66
|
+
}
|
|
67
|
+
return [...required];
|
|
68
|
+
}
|
|
47
69
|
class FrameworkManager {
|
|
48
70
|
home;
|
|
49
71
|
broadcast;
|
|
@@ -101,7 +123,7 @@ class FrameworkManager {
|
|
|
101
123
|
fs_1.default.mkdirSync(fwDir, { recursive: true });
|
|
102
124
|
const errors = [];
|
|
103
125
|
const done = [];
|
|
104
|
-
for (const provider of
|
|
126
|
+
for (const provider of requiredFrameworkProviders(fwDir, providers, this.home)) {
|
|
105
127
|
const res = (0, child_process_1.spawnSync)(nodeInterpreter(), [
|
|
106
128
|
cli,
|
|
107
129
|
'install-framework',
|
|
@@ -357,6 +379,3 @@ class FrameworkManager {
|
|
|
357
379
|
}
|
|
358
380
|
}
|
|
359
381
|
exports.FrameworkManager = FrameworkManager;
|
|
360
|
-
function dedupe(values) {
|
|
361
|
-
return Array.from(new Set(values));
|
|
362
|
-
}
|
|
@@ -489,6 +489,12 @@ function createLoopExecutors(opts = {}) {
|
|
|
489
489
|
...(effectiveIdleTimeoutMs > 0 ? { idleTimeoutMs: effectiveIdleTimeoutMs } : {}),
|
|
490
490
|
};
|
|
491
491
|
},
|
|
492
|
+
validateCoreCompletion({ coreRun, provider, model, effort, profileName, cwd, repoDir, executionManifest }) {
|
|
493
|
+
const baseStepEnv = withProfileEnv(aiStepEnv(resolveEnv(), repoDir, executionManifest), provider, profileName);
|
|
494
|
+
const core = (0, core_execution_1.prepareCoreExecution)({ run: coreRun, cwd, repoDir, manifest: executionManifest, env: baseStepEnv });
|
|
495
|
+
const env = (0, runtime_1.buildProviderEnv)((0, providers_1.getAdapter)(provider), { prompt: '', model, reasoning_effort: effort }, core.env);
|
|
496
|
+
return (0, core_execution_1.checkCoreCompletion)(core.contextPath, cwd, env, coreRun.runId);
|
|
497
|
+
},
|
|
492
498
|
async runShell({ command, cwd, repoDir, onLine, onSpawn, timeoutMs }) {
|
|
493
499
|
const env = resolveEnv();
|
|
494
500
|
return runShellCommand(command, repoDir ?? cwd, repoDir ? { ...env, SPECRAILS_REPO_DIR: repoDir } : env, timeoutMs ?? SHELL_TIMEOUT_MS, onLine, onSpawn);
|
|
@@ -1058,6 +1058,8 @@ class LoopRunManager {
|
|
|
1058
1058
|
// --yes`, codex `$implement #<id> --yes`) — then resolve `{{spec.*}}`
|
|
1059
1059
|
// data tokens and finally `{{const:*}}` library constants.
|
|
1060
1060
|
const rawTemplate = String(node.data?.prompt ?? '');
|
|
1061
|
+
const requiresCoreCompletion = /\{\{\s*cmd:(?:implement|batch)\s*\}\}/.test(rawTemplate)
|
|
1062
|
+
|| /^\s*(?:\/specrails:|\/skill:specrails-|\$)(?:implement|batch-implement)(?:\s|$)/.test(rawTemplate);
|
|
1061
1063
|
const expanded = (0, loop_constants_1.resolveConstants)(resolveRunVars((0, loop_graph_1.interpolateSpec)((0, loop_command_catalog_1.expandCommands)(rawTemplate, { provider: nodeProvider, ticketIds: req.spec?.ticketIds, specId: req.spec?.id }), req.spec), runVars), constMap);
|
|
1062
1064
|
const base = [(0, multi_repo_execution_store_1.executionManifestPrompt)(req.executionManifest), withReviewContinuationContext(expanded, rawTemplate, req.spec)].filter(Boolean).join('\n\n');
|
|
1063
1065
|
// Inject the cross-iteration history only when there's no live session
|
|
@@ -1220,6 +1222,19 @@ class LoopRunManager {
|
|
|
1220
1222
|
settled = true;
|
|
1221
1223
|
break;
|
|
1222
1224
|
}
|
|
1225
|
+
if (!res.failed && !zeroWork && !this._cancelled.has(runId)
|
|
1226
|
+
&& requiresCoreCompletion
|
|
1227
|
+
&& this.executors.validateCoreCompletion) {
|
|
1228
|
+
const completion = this.executors.validateCoreCompletion({
|
|
1229
|
+
coreRun: { runId, spec: req.spec, goal: req.spec ? undefined : JSON.stringify({ loop: req.loopName, graph: req.graph }), repositoryId: req.repositoryId },
|
|
1230
|
+
provider: nodeProvider, model: nodeModel, effort: nodeEffort, profileName: req.profileName,
|
|
1231
|
+
cwd: req.cwd, repoDir: req.repoDir, executionManifest: req.executionManifest,
|
|
1232
|
+
});
|
|
1233
|
+
if (!completion.valid) {
|
|
1234
|
+
res = { ...res, failed: true, errorText: completion.reason ?? 'Core implementation is incomplete' };
|
|
1235
|
+
logLine(`✖ ${res.errorText}`, 'stderr');
|
|
1236
|
+
}
|
|
1237
|
+
}
|
|
1223
1238
|
const blockedReason = aiStepBlockedReason(res.text);
|
|
1224
1239
|
const requiresVerification = node.data?.requireVerificationPass === true
|
|
1225
1240
|
|| /\{\{cmd:(?:verify|revision-verify|opsx:verify)\}\}/.test(rawTemplate);
|
|
@@ -1613,8 +1628,6 @@ class LoopRunManager {
|
|
|
1613
1628
|
catch (err) {
|
|
1614
1629
|
console.error(`[loop] staged accounting reconciliation failed for ${runId}:`, err);
|
|
1615
1630
|
}
|
|
1616
|
-
console.log(`[loop] settle run=${runId} outcome=${outcome} iterations=${iteration} ` +
|
|
1617
|
-
(usageTelemetryAvailable ? `cost=$${totalCost.toFixed(4)}` : 'usage=unavailable'));
|
|
1618
1631
|
// Execution and acceptance are separate facts. Persist the runtime's own
|
|
1619
1632
|
// terminal snapshot alongside counters so history does not depend on prose.
|
|
1620
1633
|
let coreCompletion = null;
|
|
@@ -1625,6 +1638,8 @@ class LoopRunManager {
|
|
|
1625
1638
|
const executionOutcome = outcome;
|
|
1626
1639
|
if (outcome === 'success' && coreCompletion && (coreCompletion.completion.implementation !== 'complete' || !['verified', 'with-exceptions'].includes(coreCompletion.completion.validation)))
|
|
1627
1640
|
outcome = 'blocked';
|
|
1641
|
+
console.log(`[loop] settle run=${runId} outcome=${outcome} iterations=${iteration} ` +
|
|
1642
|
+
(usageTelemetryAvailable ? `cost=$${totalCost.toFixed(4)}` : 'usage=unavailable'));
|
|
1628
1643
|
emitRunEvent('loop_completion', {
|
|
1629
1644
|
version: 1, execution: executionOutcome, steps: stepNum, deciderEvaluations: iteration,
|
|
1630
1645
|
turns: finalJobUsage.numTurns, costUsd: usageTelemetryAvailable ? totalCost : null,
|
|
@@ -1642,7 +1657,7 @@ class LoopRunManager {
|
|
|
1642
1657
|
// `≥` when any cost-bearing step ended unpriced (timeout/crash) — the figure
|
|
1643
1658
|
// is a lower bound, not exact. Providers without usage telemetry get an
|
|
1644
1659
|
// explicit unavailable marker, never a fabricated "$0.0000".
|
|
1645
|
-
logLine(`\n■ Loop execution finished: ${
|
|
1660
|
+
logLine(`\n■ Loop execution finished: ${outcome} — ${stepNum} step${stepNum === 1 ? '' : 's'}, ${iteration} decider evaluation${iteration === 1 ? '' : 's'}, ${finalJobUsage.numTurns ?? 'unknown'} agent turns, ` +
|
|
1646
1661
|
(usageTelemetryAvailable
|
|
1647
1662
|
? `${costUncertain ? '≥ ' : ''}$${totalCost.toFixed(4)}`
|
|
1648
1663
|
: 'usage/cost unavailable'));
|