infinity-harness 2.6.4 → 2.6.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +58 -0
- package/extensions/infinity-harness/index.ts +54 -36
- package/harness/skills/deep-research.md +38 -0
- package/package.json +1 -1
- package/src/core/config.ts +1 -0
- package/src/core/gates.ts +15 -7
- package/src/core/init.ts +6 -0
- package/src/core/phases.ts +54 -27
- package/src/core/types.ts +6 -0
- package/src/intake.ts +17 -8
- package/src/loop.ts +19 -4
- package/src/remote.ts +1 -1
- package/src/scheduler.ts +112 -9
- package/src/ui/dashboard.ts +91 -6
- package/src/ui/widget.ts +18 -0
- package/src/ui/wizard.ts +18 -1
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,64 @@ All notable changes to this project are documented here.
|
|
|
4
4
|
Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/); versions follow
|
|
5
5
|
[Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
6
6
|
|
|
7
|
+
## [2.6.6] — 2026-08-28
|
|
8
|
+
|
|
9
|
+
Every phase shows its tracked work, tabs isolate it, and autopilot actually drives.
|
|
10
|
+
|
|
11
|
+
### Fixed
|
|
12
|
+
|
|
13
|
+
- **All phases now show the damn breakdown you asked for (`wdisplay` = truth).** Every enabled phase owns tracked work: `RESEARCH` gets its 3/5/7 tasks+subtasks (by depth) even when its doc gate already passes, `DEFINE/PLAN` get starters when their gate `feature-criteria` fails, visibility is the wizard's `display` policy (`goal/sprint/feature/task/subtask` levels) respected everywhere — the widget lane and the dashboard still render `... N above/below` windowed.
|
|
14
|
+
|
|
15
|
+
- **Dashboard tabs: Goal(s) ▸ Phase(s).** Top nav filters the live plan without losing state: Goal tabs appear only when `goals.length >1` (if single goal, only Phase tabs), default = `activeTask`'s goal/phase otherwise `currentPhase`, persisted via `localStorage`, survives poll-refresh. Each feature card gets `data-phase` so phase tabs hide `phase-research` under `BUILD`, etc. `display.levels.goal/sprint/feature` still hide the tier chrome.
|
|
16
|
+
|
|
17
|
+
- **Autopilot actually auto-pilots (any phase).** `infinity_validate` was allowlisted to `research/define/plan`; now any `phaseModes[phase]==autopilot` PASS auto-advances, seeds the next phase (`ensurePhaseSeeded`), arms `spawnWorkers` scoped to `currentPhase`, and the brief still drives the main session. Copilot still parks on `needsApproval==true`. Fixes: `RESEARCH complete → advanced research → DEFINE` now lands on a `DEFINE 2/2` seeded plan, not `0/0`.
|
|
18
|
+
|
|
19
|
+
- **Seeding is per-phase, not per-first-feature.** `seedPhaseIfEmpty` previously appended to `list.features[0]` or first `phase==` match; with mixed phases it coalesced `research`+`define` into one `phase-research` feature. Now it reuses only `phase==cur` or creates `phase-<name>`, so `phase-research(5)` and `phase-define(2)` are separate cards and per-phase `computeProgress(phase)` tallies correctly. `src/loop.ts` special-cased `research` to always show (FAIL gated would leave it doc-only) and respects `approvalRejection` fingerprint (no seeding while a rejection stands, so `no-progress → give up` still fires).
|
|
20
|
+
|
|
21
|
+
- **Seeding never dirties a converged walk or an approval rejection.** `decideNext` skips seeding when `approvalRejection` present (lets `rejection-unaddressed` escalate), and only seeds `define/plan/...` when their gate probes `overall==false`. Research exception: `feature-criteria` `no real features planned yet` (seeded scaffolding excluded) so define triggers correctly. The existing `convergence` walk `define→ship` still passes because its synthetic satisfiable gates probe PASS and seeded scaffolds are skipped.
|
|
22
|
+
|
|
23
|
+
- **Widget dashboard URL: always visible.** Previously `renderWidget` only emitted `Dashboard: http://…` when `remoteServer` was live; now it always renders a hint `Dashboard: /infinity:dashboard → http://127.0.0.1:PORT` even without a server (OSC 8 hyperlink when live). You asked this twice last time and it was only half-wired.
|
|
24
|
+
|
|
25
|
+
- **Blink that blinks.** `dashboard.css` `cardPulse` 1.2s whisper → 1s `cardPulse + phasePulse + textPulse` with `4px` ring + `16px` glow, active goal/sprint/feature boxes now use `font-weight:800` + stronger `taskBlink` (`2.5px` outline, `0.85s`). `prefers-reduced-motion` still disables.
|
|
26
|
+
|
|
27
|
+
### Verified
|
|
28
|
+
|
|
29
|
+
- `tsc --noEmit` clean, `34/34` unit, `pipeline/extension/convergence/dashboard/widget/stops/coldstart/escalation/goal/realpi` all pass in isolation and `npm run e2e` (live skipped) green; `bakr_test_2.6.5 define` retro-seeded to visible `RESEARCH(5)+DEFINE(2)` on dashboard (goal tabs hidden when one goal, phase tabs live).
|
|
30
|
+
|
|
31
|
+
## [2.6.5] — 2026-08-27
|
|
32
|
+
|
|
33
|
+
Routing is honest about handoff, research knows how deep to go, and the human reads the same wiring the runner does.
|
|
34
|
+
|
|
35
|
+
### Added
|
|
36
|
+
|
|
37
|
+
- **Research depth tiers.** Wizard asks `How deep should research go?` only when `research` is in the pipeline — `standard` (Deep 3 tasks ≥5 sources ~800), `deep` (Very Deep 5 tasks ≥7 sources ~1800), `comprehensive` (Literature Review 7 tasks ≥15 annotated ~5000). `harness/config.json: researchDepth` persists it, `STARTER_TASKS_BY_DEPTH` feeds `seedPhaseIfEmpty`, `checkResearchDoc` enforces the right floor, and `harness/skills/deep-research.md` codifies the `process` skill per depth. `src/intake.ts` surfaces `Research <depth>` in the intake summary.
|
|
38
|
+
|
|
39
|
+
- **`harness/skills/deep-research.md`.** Process skill for the RESEARCH phase — depth-gated task/proof table, prior-art citation mandate, falsification experiment.
|
|
40
|
+
|
|
41
|
+
- **Dashboard actually shows its URL and the handoff→model contract.** `WidgetState.dashboardUrl` + `WidgetState.handoffModelNote` are populated from the live `remoteServer.url` and from `handoffModelNote(handoff)`; terminal header renders an `OSC 8` clickable `Dashboard: …` line, web dashboard renders `dash-url` + `handoff-note` under the masthead.
|
|
42
|
+
|
|
43
|
+
### Changed
|
|
44
|
+
|
|
45
|
+
- **Model routing follows handoff (Option A: hardest wins in the bucket).** `src/scheduler.ts: effectiveDifficultyForTask(task, handoff, list)` — `task/subtask` keeps own difficulty (subtasks inherit parent), `phase/feature/sprint/goal/off` collapses the bucket to its hardest (`easy < moderate < difficult`). `spawnWorkers` and `extensions/infinity-harness:index.ts` (`applyRouting` + `routingSummaryForBrief` + brief line) all read `config.session.handoff` each call; `src/intake.ts: HANDOFF_QUESTION` help text now states `Model per …` per choice. `src/handoff.ts` wording updated to match.
|
|
46
|
+
|
|
47
|
+
- **Seeding + config defaults.** `src/core/init.ts` writes `researchDepth` (`deep` when research enabled), `src/core/config.ts` defaults it, `src/core/phases.ts` returns the depth-appropriate starter set via `loadConfig` (no raw file read), `src/remote.ts` keeps the read-only dashboard read via `require` deferred import so pack audit stays green (`42` modules reachable).
|
|
48
|
+
|
|
49
|
+
### Fixed
|
|
50
|
+
|
|
51
|
+
- **Wizard ate its own answer.** Real-pi `research-first` and `build one` → research handoff paths previously timed out because depth answer matched the question title not option text and consumed the wrong line; e2e now answers depth with regex on option (`/Very Deep/`), research gate fixture length raised to `40×` so `deep = 1800` passes, and both realpi `wizard` + `custom` + `coldstart` `a custom workflow …` are green.
|
|
52
|
+
|
|
53
|
+
- **Second reuse init lost its saved workflow.** `coldstart-reuse` answered with a bare `Client work \(yours\)` line that no longer matched after research re-enabled depth intake; now answered `Very Deep` so the workflow survives.
|
|
54
|
+
|
|
55
|
+
- **Gate fixture regression.** `research.md` fixture at `20×` was `1152` chars — short of the new `deep 1800` floor — caused `the whole pipeline …` stall; corrected to `40×` (`2280`) so `research: pass` remains single-shot.
|
|
56
|
+
|
|
57
|
+
- **`src/intake.ts` CRLF drift.** Full-file rewrite was display-only (line endings); fixed to LF and only the 25-line surgical change retained.
|
|
58
|
+
|
|
59
|
+
- **Extension ESM require shim.** Prior `require("../../src/scheduler.ts")` in `widgetStateFor` broke ESM and never surfaced `handoffModelNote`; inlined pure map instead.
|
|
60
|
+
|
|
61
|
+
### Verified
|
|
62
|
+
|
|
63
|
+
- `tsc --noEmit` clean, `34/34` unit, `15/15` e2e (`coldstart 12/12`, `realpi 10/10`), `42` modules reachable via `package` scenario. `tests/skills.test.ts` bumps shipped skill count `28 → 29`, `tests/intake.test.ts` asserts `How deep …` + `Research deep` summary.
|
|
64
|
+
|
|
7
65
|
## [2.6.4] — 2026-08-25
|
|
8
66
|
|
|
9
67
|
`feature-criteria` now ignores seeded `phase-*` scaffolding so `DEFINE` does not pass on the
|
|
@@ -168,10 +168,24 @@ export default function (pi: ExtensionAPI): void {
|
|
|
168
168
|
|
|
169
169
|
// -- widget ---------------------------------------------------------------
|
|
170
170
|
|
|
171
|
+
const handoffNoteFor = (h: import("../../src/core/types.ts").HandoffGranularity): string | null => {
|
|
172
|
+
const map: Record<string, string> = {
|
|
173
|
+
off: "Model per run (off/goal) — the whole run shares its hardest model; finer per-task routing requires task/subtask handoff",
|
|
174
|
+
goal: "Model per run (off/goal) — the whole run shares its hardest model; finer per-task routing requires task/subtask handoff",
|
|
175
|
+
phase: "Model per phase — tasks & subtasks in a phase share the hardest model in that phase",
|
|
176
|
+
sprint: "Model per sprint — tasks & subtasks in a sprint share the hardest model in that sprint",
|
|
177
|
+
feature: "Model per feature — tasks & subtasks in a feature share the hardest model in that feature",
|
|
178
|
+
task: "Model per task — subtasks share their parent task's model",
|
|
179
|
+
subtask: "Model per subtask — each subtask may use its own model (needs subtask difficulty)",
|
|
180
|
+
};
|
|
181
|
+
return map[h] ?? null;
|
|
182
|
+
};
|
|
183
|
+
|
|
171
184
|
const widgetStateFor = (dir: string): WidgetState | null => {
|
|
172
185
|
try {
|
|
173
186
|
const { list } = loadFeatureList(dir);
|
|
174
187
|
const { config } = loadConfig(dir);
|
|
188
|
+
const handoffModelNote: string | null = handoffNoteFor((config.session?.handoff as import("../../src/core/types.ts").HandoffGranularity) ?? "task");
|
|
175
189
|
const spent = escalationSummary(dir);
|
|
176
190
|
const loop = readJsonSafe<{ escalations?: { strategy: string }[] } | null>(
|
|
177
191
|
loopStatePath(dir),
|
|
@@ -184,6 +198,8 @@ export default function (pi: ExtensionAPI): void {
|
|
|
184
198
|
return {
|
|
185
199
|
list,
|
|
186
200
|
view,
|
|
201
|
+
dashboardUrl: remoteServer?.url ?? null,
|
|
202
|
+
handoffModelNote,
|
|
187
203
|
sessions: run?.sessions ?? null,
|
|
188
204
|
intake: typeof config.intake?.brief === "string" ? config.intake.brief : null,
|
|
189
205
|
awaitingApproval: config.awaitingApproval ?? null,
|
|
@@ -301,10 +317,15 @@ export default function (pi: ExtensionAPI): void {
|
|
|
301
317
|
const applyRouting = async (ctx: ExtensionContext, dir: string, source: string): Promise<void> => {
|
|
302
318
|
try {
|
|
303
319
|
const { resolveModel, resolveThinking } = await import("../../src/modelRouter.ts");
|
|
320
|
+
const { effectiveDifficultyForTask } = await import("../../src/scheduler.ts");
|
|
304
321
|
const { nextActionableTask, findFeature } = await import("../../src/core/featureList.ts");
|
|
305
322
|
const { loadFeatureList: loadList } = await import("../../src/core/featureList.ts");
|
|
306
323
|
const list = loadList(dir).list;
|
|
307
324
|
const task = nextActionableTask(list);
|
|
325
|
+
const cfg = loadConfig(dir).config;
|
|
326
|
+
const handoff = (cfg.session?.handoff ?? "task") as import("../../src/core/types.ts").HandoffGranularity;
|
|
327
|
+
// Effective difficulty honors handoff bucket: phase/feature/sprint tasks share hardest in that bucket.
|
|
328
|
+
const effDiff = task && list ? (effectiveDifficultyForTask(task as import("../../src/core/featureList.ts").FlatTask, handoff, list) ?? (task as { difficulty?: string }).difficulty) : (task as { difficulty?: string } | null)?.difficulty;
|
|
308
329
|
// Resolve against task/parent feature/sprint difficulty; fall through to default when no actionable.
|
|
309
330
|
const feature = task ? findFeature(list, task.featureId) ?? undefined : undefined;
|
|
310
331
|
const sprint = feature?.sprintId ? (list.sprints ?? []).find((s) => s.id === feature.sprintId) ?? undefined : undefined;
|
|
@@ -312,7 +333,7 @@ export default function (pi: ExtensionAPI): void {
|
|
|
312
333
|
type T = NonNullable<ReturnType<typeof nextActionableTask>>;
|
|
313
334
|
const routedModel = resolveModel({
|
|
314
335
|
projectDir: dir,
|
|
315
|
-
task: (
|
|
336
|
+
task: effDiff || (feature as { difficulty?: string } | undefined)?.difficulty ? ({ difficulty: effDiff as string | undefined, modelHint: (task as T | undefined)?.modelHint, id: task?.id, key: (task as T | undefined)?.compositeKey ?? (task as T | undefined)?.key } as never) : undefined,
|
|
316
337
|
feature: feature as never,
|
|
317
338
|
sprint: sprint as never,
|
|
318
339
|
phase: loadConfig(dir).config.currentPhase ?? undefined,
|
|
@@ -320,7 +341,7 @@ export default function (pi: ExtensionAPI): void {
|
|
|
320
341
|
});
|
|
321
342
|
const routedThinking = resolveThinking({
|
|
322
343
|
projectDir: dir,
|
|
323
|
-
task: (task as T | null | undefined) ? ({ difficulty: (task as T | undefined)?.difficulty, id: task?.id, key: (task as T | undefined)?.compositeKey ?? (task as T | undefined)?.key } as never) : undefined,
|
|
344
|
+
task: effDiff ? ({ difficulty: effDiff as string } as never) : (task as T | null | undefined) ? ({ difficulty: (task as T | undefined)?.difficulty, id: task?.id, key: (task as T | undefined)?.compositeKey ?? (task as T | undefined)?.key } as never) : undefined,
|
|
324
345
|
feature: feature as never,
|
|
325
346
|
sprint: sprint as never,
|
|
326
347
|
});
|
|
@@ -378,15 +399,19 @@ export default function (pi: ExtensionAPI): void {
|
|
|
378
399
|
try {
|
|
379
400
|
const { nextActionableTask, findFeature } = await import("../../src/core/featureList.ts");
|
|
380
401
|
const { resolveModel, resolveThinking } = await import("../../src/modelRouter.ts");
|
|
402
|
+
const { effectiveDifficultyForTask } = await import("../../src/scheduler.ts");
|
|
381
403
|
const { list } = loadFeatureList(dir);
|
|
382
404
|
const task = nextActionableTask(list);
|
|
383
405
|
if (!task) return null;
|
|
406
|
+
const cfg = loadConfig(dir).config;
|
|
407
|
+
const handoff = (cfg.session?.handoff ?? "task") as import("../../src/core/types.ts").HandoffGranularity;
|
|
408
|
+
const effDiff = effectiveDifficultyForTask(task as import("../../src/core/featureList.ts").FlatTask, handoff, list) ?? (task as { difficulty?: string }).difficulty;
|
|
384
409
|
const feature = findFeature(list, task.featureId) ?? undefined;
|
|
385
410
|
const sprint = feature?.sprintId ? (list.sprints ?? []).find((s) => s.id === feature.sprintId) ?? undefined : undefined;
|
|
386
411
|
// eslint-disable-next-line @typescript-eslint/no-explicit-any
|
|
387
|
-
const m = resolveModel({ projectDir: dir, task: ({ difficulty:
|
|
412
|
+
const m = resolveModel({ projectDir: dir, task: ({ difficulty: effDiff as string | undefined, modelHint: (task as any).modelHint, id: task.id, key: (task as any).compositeKey ?? (task as any).key } as never), feature: feature as never, sprint: sprint as never, phase: loadConfig(dir).config.currentPhase ?? undefined });
|
|
388
413
|
// eslint-disable-next-line @typescript-eslint/no-explicit-any
|
|
389
|
-
const th = resolveThinking({ projectDir: dir, task: ({ difficulty:
|
|
414
|
+
const th = resolveThinking({ projectDir: dir, task: ({ difficulty: effDiff as string | undefined } as never), feature: feature as never, sprint: sprint as never });
|
|
390
415
|
if (!m || !m.trim()) return null;
|
|
391
416
|
return `Routing: ${task.compositeKey} → ${m}${th ? ` · thinking ${th}` : ""}`;
|
|
392
417
|
} catch { return null; }
|
|
@@ -1200,40 +1225,33 @@ export default function (pi: ExtensionAPI): void {
|
|
|
1200
1225
|
const lines = gate.checks
|
|
1201
1226
|
.map((c) => `${c.advisory ? "·" : c.pass ? "+" : "x"} ${c.name}: ${c.detail}`)
|
|
1202
1227
|
.join("\n");
|
|
1203
|
-
//
|
|
1204
|
-
//
|
|
1205
|
-
//
|
|
1206
|
-
//
|
|
1207
|
-
// phases that require real work (tests, coverage, clean tree) to have
|
|
1208
|
-
// genuinely passed on the *next* phase's gate as well — otherwise a
|
|
1209
|
-
// single infinity_validate hops build→verify→review.
|
|
1210
|
-
// Only auto-advance doc/process phases whose gate is purely content (
|
|
1211
|
-
// research, define, plan). BUILD and later require explicit validation.
|
|
1228
|
+
// Autopilot means auto-pilot: when the gate passes and the current
|
|
1229
|
+
// phase's phaseMode is autopilot, advance immediately (any phase). The
|
|
1230
|
+
// old allowlist stalled real autopilot builds after RESEARCH → DEFINE.
|
|
1231
|
+
// Copilot still parks via needsApproval check below.
|
|
1212
1232
|
if (gate.overall && !params?.feature && !params?.task) {
|
|
1213
|
-
|
|
1214
|
-
|
|
1215
|
-
|
|
1216
|
-
|
|
1217
|
-
const
|
|
1218
|
-
|
|
1219
|
-
|
|
1220
|
-
|
|
1221
|
-
|
|
1222
|
-
|
|
1223
|
-
|
|
1224
|
-
|
|
1225
|
-
|
|
1226
|
-
|
|
1227
|
-
|
|
1228
|
-
|
|
1229
|
-
|
|
1230
|
-
|
|
1231
|
-
|
|
1232
|
-
};
|
|
1233
|
-
}
|
|
1233
|
+
try {
|
|
1234
|
+
const { needsApproval } = await import("../../src/approval.ts");
|
|
1235
|
+
const fresh = loadConfig(dir).config;
|
|
1236
|
+
if (!needsApproval(fresh, fresh.currentPhase)) {
|
|
1237
|
+
const { advancePhase, ensurePhaseSeeded } = await import("../../src/core/phases.ts");
|
|
1238
|
+
const moved = await advancePhase(dir);
|
|
1239
|
+
if (moved.ok && moved.to) {
|
|
1240
|
+
try { ensurePhaseSeeded(dir, moved.to); } catch {}
|
|
1241
|
+
refreshWidget(ctx as ExtensionContext);
|
|
1242
|
+
const brief = await briefText(dir);
|
|
1243
|
+
return {
|
|
1244
|
+
content: [
|
|
1245
|
+
{
|
|
1246
|
+
type: "text",
|
|
1247
|
+
text: `Gate PASS on ${gate.phase} → advanced ${moved.from} → ${moved.to}\n${lines}\n\n${brief}`,
|
|
1248
|
+
},
|
|
1249
|
+
],
|
|
1250
|
+
details: { ...gate, advanced: moved } as unknown as typeof gate,
|
|
1251
|
+
};
|
|
1234
1252
|
}
|
|
1235
|
-
}
|
|
1236
|
-
}
|
|
1253
|
+
}
|
|
1254
|
+
} catch {}
|
|
1237
1255
|
}
|
|
1238
1256
|
return {
|
|
1239
1257
|
content: [
|
|
@@ -0,0 +1,38 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: deep-research
|
|
3
|
+
description: Deep prior-art sweep, synthesis and tradeoff analysis — literature-review level when needed
|
|
4
|
+
tags: [research, literature, prior-art, synthesis, constraints, tradeoffs, recommendation, falsification]
|
|
5
|
+
when: research depth is standard/deep/comprehensive and the question needs more than a web search
|
|
6
|
+
phases: [research]
|
|
7
|
+
kind: process
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
# Deep research
|
|
11
|
+
|
|
12
|
+
Use when `config.researchDepth` is set — `standard` (Deep), `deep` (Very Deep) or `comprehensive` (Literature Review). The wizard only asks when `research` is in the pipeline.
|
|
13
|
+
|
|
14
|
+
## Depth — what the harness expects
|
|
15
|
+
|
|
16
|
+
| Depth | Tasks | Sources | Gate | What you deliver |
|
|
17
|
+
|-------|-------|---------|------|-----------------|
|
|
18
|
+
| `standard` (Deep) | 3 | ≥5 primary, all with URLs | ~800 chars | comparison table, constraints table, ≥3 options |
|
|
19
|
+
| `deep` (Very Deep) | 5 | ≥7 primary + gap analysis | ~1800 chars | above + competitive matrix, cost/risk model, risk register |
|
|
20
|
+
| `comprehensive` (Literature Review) | 7 | ≥15 annotated bibliography | ~5000 chars | above + benchmarks on a toy case, synthesis, gap analysis, ADR |
|
|
21
|
+
|
|
22
|
+
Deeper = longer `harness/docs/RESEARCH.md`. The gate reads `config.researchDepth` and enforces the char floor.
|
|
23
|
+
|
|
24
|
+
## Process
|
|
25
|
+
|
|
26
|
+
1. Collect primary sources only — official docs, specs, first-party APIs, papers with DOIs, postmortems. Every claim needs a URL or citation.
|
|
27
|
+
2. Fill the constraints table `given vs inferred` (inferred = question for DEFINE).
|
|
28
|
+
3. Benchmark or reason about ≥1 approach on a minimal case where possible (comprehensive must).
|
|
29
|
+
4. Lay out genuine options (standard ≥3) with architecture sketch, cost, risk, team & time. A list of one is a decision wearing a disguise.
|
|
30
|
+
5. Write an ADR: recommendation + what would have to be true for it to be wrong (falsification + experiment design).
|
|
31
|
+
6. Risk register: known unknowns vs unknown unknowns, mitigations.
|
|
32
|
+
7. Ranked open questions for DEFINE (standard ≥5, deep ≥8, comprehensive ≥12) + glossary delta for `DOMAIN.md`.
|
|
33
|
+
|
|
34
|
+
## Anti-patterns
|
|
35
|
+
|
|
36
|
+
- No open questions — you did not look hard enough.
|
|
37
|
+
- One option presented as inevitable.
|
|
38
|
+
- Findings with no source, stated as confidently as sourced ones.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "infinity-harness",
|
|
3
|
-
"version": "2.6.
|
|
3
|
+
"version": "2.6.6",
|
|
4
4
|
"description": "A pi agent extension that runs a gated build pipeline unattended \u2014 enforces phases, validates with deterministic gates, and keeps working for hours or days without losing the plan.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"keywords": [
|
package/src/core/config.ts
CHANGED
|
@@ -50,6 +50,7 @@ export function defaultConfig(): HarnessConfig {
|
|
|
50
50
|
},
|
|
51
51
|
phases: { enabled: [...DEFAULT_ENABLED_PHASES] },
|
|
52
52
|
roles: { strict: false },
|
|
53
|
+
researchDepth: "deep" as import("./types.ts").ResearchDepth,
|
|
53
54
|
session: { handoff: "task", contextThreshold: 0.6, carryNotes: true },
|
|
54
55
|
execution: { parallelAt: "task", maxWorkers: 3 },
|
|
55
56
|
approvals: { research: false, define: false, plan: false },
|
package/src/core/gates.ts
CHANGED
|
@@ -231,8 +231,15 @@ async function checkChangelog({ targetDir }: Ctx): Promise<CheckResult> {
|
|
|
231
231
|
* bar deliberately — the gate judges that work happened, the human judges
|
|
232
232
|
* whether it was any good.
|
|
233
233
|
*/
|
|
234
|
-
async function checkResearchDoc({ targetDir }: Ctx): Promise<CheckResult> {
|
|
235
|
-
|
|
234
|
+
async function checkResearchDoc({ targetDir, config }: Ctx): Promise<CheckResult> {
|
|
235
|
+
// Depth-dependent threshold — Literature needs a real review, Standard is the old 400 baseline.
|
|
236
|
+
const depth = (config as { researchDepth?: string }).researchDepth as string | undefined;
|
|
237
|
+
const minChars = depth === "comprehensive" ? 5000 : depth === "deep" ? 1800 : 800;
|
|
238
|
+
// Note: depth="deep" is the default, 800 ≈ old 400 doubled but Standard is a true lite mode.
|
|
239
|
+
// Comprehensive still passes if Standard doc exists — depth is about work produced, not gate strictness,
|
|
240
|
+
// but the line below makes the harness actually demand the depth the wizard promised.
|
|
241
|
+
const effectiveMin = depth ? minChars : 400;
|
|
242
|
+
return docCheck("research-doc", P.researchPath(targetDir), effectiveMin, "harness/docs/RESEARCH.md");
|
|
236
243
|
}
|
|
237
244
|
|
|
238
245
|
async function checkArchitectureDoc({ targetDir }: Ctx): Promise<CheckResult> {
|
|
@@ -329,8 +336,10 @@ async function checkNoEmptyDirs({ targetDir }: Ctx): Promise<CheckResult> {
|
|
|
329
336
|
async function checkFeatureCriteria({ targetDir }: Ctx): Promise<CheckResult> {
|
|
330
337
|
const { list } = loadFeatureList(targetDir);
|
|
331
338
|
// Seeded starter features (phase-*) are scaffolding, not real features — ignore for criteria gate.
|
|
339
|
+
// Also: if no non-scaffold features exist yet, DEFINE has not been planned and must not PASS on a
|
|
340
|
+
// single phase-* dummy. Treat absent as FAIL with a distinct detail so callers seeding define know to act.
|
|
332
341
|
const features = (list.features ?? []).filter((f) => !(f as { phase?: string }).phase);
|
|
333
|
-
if (features.length === 0) return fail("feature-criteria", "no features planned yet");
|
|
342
|
+
if (features.length === 0) return fail("feature-criteria", "no real features planned yet — run DEFINE");
|
|
334
343
|
const missing = features.filter((f) => !(f.criteria ?? []).length).map((f) => f.id);
|
|
335
344
|
return missing.length === 0
|
|
336
345
|
? pass("feature-criteria", `${features.length} feature(s) have criteria`)
|
|
@@ -374,10 +383,9 @@ type Check = (ctx: Ctx) => Promise<CheckResult>;
|
|
|
374
383
|
|
|
375
384
|
const PHASE_CHECKS: Record<Phase, Check[]> = {
|
|
376
385
|
init: [checkGitRepo, checkConfigExists],
|
|
377
|
-
//
|
|
378
|
-
//
|
|
379
|
-
//
|
|
380
|
-
// still shows ... subtasks and progress as the real signal; BUILD keeps tasksComplete blocking.
|
|
386
|
+
// tasksComplete is advisory on seeded-phase gates only when seeded work exists — otherwise the
|
|
387
|
+
// converge walk would freeze on the seeded define/plan tasks that test never completes. Once BUILD
|
|
388
|
+
// has tasksComplete remains blocking.
|
|
381
389
|
research: [checkResearchDoc],
|
|
382
390
|
define: [checkFeatureCriteria, checkSkillsLoad],
|
|
383
391
|
plan: [checkFeatureCriteria, checkTasksPlanned],
|
package/src/core/init.ts
CHANGED
|
@@ -146,6 +146,7 @@ function pythonCommands(targetDir: string): ProjectCommands {
|
|
|
146
146
|
export type InitOptions = {
|
|
147
147
|
stack?: StackId;
|
|
148
148
|
mode?: "copilot" | "autopilot";
|
|
149
|
+
researchDepth?: import("./types.ts").ResearchDepth;
|
|
149
150
|
phases?: Phase[];
|
|
150
151
|
commands?: Partial<ProjectCommands>;
|
|
151
152
|
/** Re-scaffold missing files in a project that already has a config. */
|
|
@@ -213,6 +214,11 @@ export function initHarness(targetDir: string, options: InitOptions = {}): InitR
|
|
|
213
214
|
const phase = phases[0] ?? "define";
|
|
214
215
|
|
|
215
216
|
const config = defaultConfig();
|
|
217
|
+
if (options.researchDepth && (options.researchDepth === "standard" || options.researchDepth === "deep" || options.researchDepth === "comprehensive")) {
|
|
218
|
+
config.researchDepth = options.researchDepth;
|
|
219
|
+
} else if (phases.includes("research")) {
|
|
220
|
+
config.researchDepth = "deep";
|
|
221
|
+
}
|
|
216
222
|
config.stack = stack.id === "unknown" ? null : stack.id;
|
|
217
223
|
config.mode = options.mode ?? "copilot";
|
|
218
224
|
config.phases = { enabled: phases };
|
package/src/core/phases.ts
CHANGED
|
@@ -140,27 +140,37 @@ export type StarterTask = {
|
|
|
140
140
|
difficulty: "easy" | "moderate" | "difficult";
|
|
141
141
|
subtasks?: string[];
|
|
142
142
|
};
|
|
143
|
-
|
|
144
|
-
|
|
145
|
-
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
152
|
-
|
|
153
|
-
|
|
154
|
-
|
|
155
|
-
|
|
156
|
-
},
|
|
157
|
-
{
|
|
158
|
-
id: "research/r3",
|
|
159
|
-
description: "Recommend one option, falsification condition, and open questions for DEFINE",
|
|
160
|
-
difficulty: "moderate",
|
|
161
|
-
subtasks: ["recommendation + falsification", "open questions list"],
|
|
162
|
-
},
|
|
143
|
+
/** Depth: Standard(3 tasks/9 subtasks) < Deep(5/15) < Comprehensive(10/30+). Default deep. */
|
|
144
|
+
export type ResearchDepth = "standard" | "deep" | "comprehensive";
|
|
145
|
+
export const STARTER_TASKS_BY_DEPTH: Record<ResearchDepth, StarterTask[]> = {
|
|
146
|
+
standard: [
|
|
147
|
+
{ id: "research/r1", description: "Collect prior art — 3 primary sources with URLs, what exists, where it stops", difficulty: "moderate", subtasks: ["source 1 + URL + summary", "source 2 + URL + summary", "source 3 + URL + summary"] },
|
|
148
|
+
{ id: "research/r2", description: "Name constraints (given vs inferred) and lay out 2+ options with costs", difficulty: "moderate", subtasks: ["constraints given vs inferred table", "option A cost/benefit", "option B cost/benefit"] },
|
|
149
|
+
{ id: "research/r3", description: "Recommend one option, falsification condition, and open questions for DEFINE", difficulty: "moderate", subtasks: ["recommendation + falsification", "open questions list (≥5)"] },
|
|
150
|
+
],
|
|
151
|
+
deep: [
|
|
152
|
+
{ id: "research/r1", description: "Prior art: ≥5 primary sources with URLs (docs/specs/repos), what each gets right and where it stops", difficulty: "moderate", subtasks: ["sources 1-3 + URLs + summaries", "sources 4-5 + URLs + gap analysis", "comparison table: feature × prior art"] },
|
|
153
|
+
{ id: "research/r2", description: "Constraints & domain model: given vs inferred, glossary terms, actors & data", difficulty: "moderate", subtasks: ["constraints given vs inferred (table)", "domain glossary + actors", "data & platform constraints"] },
|
|
154
|
+
{ id: "research/r3", description: "Options: ≥3 genuine alternatives with architecture, cost, risk and trade-offs", difficulty: "difficult", subtasks: ["option A: arch + cost + risk", "option B: arch + cost + risk", "option C / hybrid + trade-off matrix"] },
|
|
155
|
+
{ id: "research/r4", description: "Recommendation with falsification: what must be true, what would prove it wrong", difficulty: "moderate", subtasks: ["recommendation + rationale", "falsification condition + experiment"] },
|
|
156
|
+
{ id: "research/r5", description: "Open questions for DEFINE: ranked questions only a human can answer", difficulty: "easy", subtasks: ["open questions (≥8) ranked", "DEFINE interview agenda"] },
|
|
163
157
|
],
|
|
158
|
+
comprehensive: [
|
|
159
|
+
{ id: "research/r1", description: "Literature sweep: ≥15 primary sources — papers, RFCs, repos, postmortems — annotated", difficulty: "difficult", subtasks: ["sources 1-5 annotated", "sources 6-10 annotated", "sources 11-15 annotated", "citation map + gaps in literature"] },
|
|
160
|
+
{ id: "research/r2", description: "Domain & constraints synthesis: glossary, actors, data, regulatory & platform limits", difficulty: "difficult", subtasks: ["constraints given vs inferred (full table)", "domain glossary + bounded contexts", "actors, data flows & invariants"] },
|
|
161
|
+
{ id: "research/r3", description: "Benchmark prior work: reproduce or reason about 3+ approaches on a toy case", difficulty: "difficult", subtasks: ["approach A benchmark", "approach B benchmark", "approach C benchmark + comparison matrix"] },
|
|
162
|
+
{ id: "research/r4", description: "Architecture options: ≥3 with diagrams, cost model, risk register, team & time", difficulty: "difficult", subtasks: ["option A: diagram + cost + risk", "option B: diagram + cost + risk", "option C: diagram + cost + risk", "trade-off matrix + decision criteria"] },
|
|
163
|
+
{ id: "research/r5", description: "Recommendation as a decision record + what falsifies it", difficulty: "moderate", subtasks: ["ADR: recommendation + alternatives rejected", "falsification condition + experiment design"] },
|
|
164
|
+
{ id: "research/r6", description: "Risk & unknowns register: known unknowns, unknown unknowns, mitigations", difficulty: "moderate", subtasks: ["risk register", "mitigations + owners", "open unknowns vs knowns"] },
|
|
165
|
+
{ id: "research/r7", description: "DEFINE handoff: ranked open questions (≥12) + interview agenda + glossary delta", difficulty: "easy", subtasks: ["open questions (≥12) ranked", "DEFINE interview agenda", "glossary delta for DOMAIN.md"] },
|
|
166
|
+
],
|
|
167
|
+
};
|
|
168
|
+
// Back-compat: default deep
|
|
169
|
+
const DEFAULT_RESEARCH_DEPTH: ResearchDepth = "deep";
|
|
170
|
+
export const STARTER_TASKS: Record<string, StarterTask[]> = {
|
|
171
|
+
get research(): StarterTask[] { return STARTER_TASKS_BY_DEPTH[DEFAULT_RESEARCH_DEPTH]; },
|
|
172
|
+
set research(v: StarterTask[]) { (STARTER_TASKS_BY_DEPTH as Record<string, StarterTask[]>)[DEFAULT_RESEARCH_DEPTH] = v; },
|
|
173
|
+
|
|
164
174
|
define: [
|
|
165
175
|
{
|
|
166
176
|
id: "define/d1",
|
|
@@ -197,19 +207,36 @@ export function isPhaseDone(dir: string, phase: import("./types.ts").Phase): boo
|
|
|
197
207
|
return tasks.length > 0 && tasks.every((t) => t.status === "complete");
|
|
198
208
|
}
|
|
199
209
|
|
|
210
|
+
function startersForPhase(dir: string, phase: string): StarterTask[] {
|
|
211
|
+
if (phase !== "research") return STARTER_TASKS[phase] ?? [];
|
|
212
|
+
try {
|
|
213
|
+
const { config } = loadConfig(dir);
|
|
214
|
+
const depth = (config as { researchDepth?: ResearchDepth }).researchDepth;
|
|
215
|
+
if (depth && STARTER_TASKS_BY_DEPTH[depth]) return STARTER_TASKS_BY_DEPTH[depth];
|
|
216
|
+
} catch {}
|
|
217
|
+
return STARTER_TASKS_BY_DEPTH[DEFAULT_RESEARCH_DEPTH] ?? [];
|
|
218
|
+
}
|
|
219
|
+
|
|
220
|
+
export function ensurePhaseSeeded(dir: string, phase: import("./types.ts").Phase): boolean {
|
|
221
|
+
try {
|
|
222
|
+
const { list } = loadFeatureList(dir);
|
|
223
|
+
if (tasksForPhase(list, phase).length > 0) return false;
|
|
224
|
+
const r = seedPhaseIfEmpty(dir, phase);
|
|
225
|
+
return r.seeded;
|
|
226
|
+
} catch { return false; }
|
|
227
|
+
}
|
|
228
|
+
|
|
200
229
|
export function seedPhaseIfEmpty(dir: string, phase: import("./types.ts").Phase): { seeded: boolean; error: string | null } {
|
|
201
|
-
const seeded =
|
|
230
|
+
const seeded = startersForPhase(dir, phase);
|
|
202
231
|
if (seeded.length === 0) return { seeded: false, error: null };
|
|
203
232
|
try {
|
|
204
233
|
const { list } = loadFeatureList(dir);
|
|
205
234
|
const existing = tasksForPhase(list, phase);
|
|
206
235
|
if (existing.length > 0) return { seeded: false, error: null };
|
|
207
|
-
//
|
|
208
|
-
|
|
209
|
-
|
|
210
|
-
|
|
211
|
-
({ id: `phase-${phase}`, name: phase.toUpperCase(), tasks: [] } as unknown as typeof list.features[number])
|
|
212
|
-
);
|
|
236
|
+
// Each phase gets its own feature (phase-<name>) so per-phase tabs/groups isolate correctly.
|
|
237
|
+
// Never append a "define" task to a "research" feature — the feature.phase is the grouping key.
|
|
238
|
+
const match = list.features.find((f) => (f as { phase?: string }).phase === phase);
|
|
239
|
+
const feature: typeof list.features[number] = match ?? ({ id: `phase-${phase}`, name: phase.toUpperCase(), tasks: [] } as unknown as typeof list.features[number]);
|
|
213
240
|
if (!list.features.includes(feature as any)) {
|
|
214
241
|
(feature as { phase?: string }).phase = phase;
|
|
215
242
|
list.features.push(feature as any);
|
package/src/core/types.ts
CHANGED
|
@@ -244,6 +244,10 @@ export type DisplayPolicy = {
|
|
|
244
244
|
taskWindow: number;
|
|
245
245
|
};
|
|
246
246
|
|
|
247
|
+
/** How deep the research phase goes, when enabled. Only asked when research is in the pipeline. */
|
|
248
|
+
export type ResearchDepth = "standard" | "deep" | "comprehensive";
|
|
249
|
+
export const RESEARCH_DEPTHS: readonly ResearchDepth[] = ["standard", "deep", "comprehensive"] as const;
|
|
250
|
+
|
|
247
251
|
/** What the start-up wizard settled, so it is never asked twice. */
|
|
248
252
|
export type IntakeState = {
|
|
249
253
|
/** True once the wizard has run to completion for this project. */
|
|
@@ -257,6 +261,8 @@ export type IntakeState = {
|
|
|
257
261
|
export type HarnessConfig = {
|
|
258
262
|
version: string;
|
|
259
263
|
stack: string | null;
|
|
264
|
+
/** Research depth — only meaningful when research is in phases.enabled. */
|
|
265
|
+
researchDepth?: ResearchDepth;
|
|
260
266
|
mode: "copilot" | "autopilot";
|
|
261
267
|
currentPhase: Phase | null;
|
|
262
268
|
currentRole: Role | null;
|
package/src/intake.ts
CHANGED
|
@@ -44,9 +44,13 @@ export type Mode = "copilot" | "autopilot";
|
|
|
44
44
|
export const INTAKE_STEPS = ["workflow", "brief", "handoff", "display"] as const;
|
|
45
45
|
export type IntakeStep = (typeof INTAKE_STEPS)[number];
|
|
46
46
|
|
|
47
|
+
export type ResearchDepth = import("./core/types.ts").ResearchDepth;
|
|
48
|
+
|
|
47
49
|
export type IntakeAnswers = {
|
|
48
50
|
/** The chosen workflow: a built-in, one they saved, or one they just built. */
|
|
49
51
|
workflow: Workflow;
|
|
52
|
+
/** Research depth — only when research is in the pipeline. */
|
|
53
|
+
researchDepth?: ResearchDepth;
|
|
50
54
|
/** What the human wants built, in their words. Empty is allowed but warned about. */
|
|
51
55
|
brief: string;
|
|
52
56
|
/** Session handoff policy. Defaults to a fresh session per phase. */
|
|
@@ -71,6 +75,7 @@ export type IntakePlan = {
|
|
|
71
75
|
/** Derived: "copilot" when the run stops for the human anywhere, else "autopilot". */
|
|
72
76
|
mode: Mode;
|
|
73
77
|
workflow: { id: string; name: string };
|
|
78
|
+
researchDepth?: ResearchDepth;
|
|
74
79
|
brief: string | null;
|
|
75
80
|
phases: Phase[];
|
|
76
81
|
phaseModes: PhaseModes;
|
|
@@ -155,9 +160,11 @@ export function planIntake(answers: IntakeAnswers): IntakePlan {
|
|
|
155
160
|
);
|
|
156
161
|
}
|
|
157
162
|
|
|
163
|
+
const _researchDepth: ResearchDepth | undefined = (phases.includes("research" as Phase) ? ((answers.researchDepth as ResearchDepth) ?? "deep") : undefined) as ResearchDepth | undefined;
|
|
158
164
|
return {
|
|
159
165
|
mode,
|
|
160
166
|
workflow: { id: workflow.id, name: workflow.name },
|
|
167
|
+
researchDepth: _researchDepth,
|
|
161
168
|
brief,
|
|
162
169
|
phases,
|
|
163
170
|
phaseModes,
|
|
@@ -170,7 +177,7 @@ export function planIntake(answers: IntakeAnswers): IntakePlan {
|
|
|
170
177
|
execution: { parallelAt, maxWorkers },
|
|
171
178
|
display,
|
|
172
179
|
router: answers.router,
|
|
173
|
-
summary: summarize(workflow, phases, phaseModes, session, display, brief),
|
|
180
|
+
summary: summarize(workflow, phases, phaseModes, session, display, brief, _researchDepth),
|
|
174
181
|
warnings,
|
|
175
182
|
};
|
|
176
183
|
}
|
|
@@ -182,6 +189,7 @@ function summarize(
|
|
|
182
189
|
session: SessionPolicy,
|
|
183
190
|
display: DisplayPolicy,
|
|
184
191
|
brief: string | null,
|
|
192
|
+
researchDepth?: ResearchDepth | undefined,
|
|
185
193
|
): string {
|
|
186
194
|
const signed = phases.filter((p) => modes[p] === "copilot");
|
|
187
195
|
const L: string[] = [];
|
|
@@ -200,6 +208,7 @@ function summarize(
|
|
|
200
208
|
}`,
|
|
201
209
|
);
|
|
202
210
|
L.push(`Display ${display.preset}`);
|
|
211
|
+
if (researchDepth) L.push(`Research ${researchDepth}`);
|
|
203
212
|
L.push(`Goal ${brief ?? "(none yet — you will be asked first thing)"}`);
|
|
204
213
|
return L.join("\n");
|
|
205
214
|
}
|
|
@@ -237,37 +246,37 @@ export const HANDOFF_QUESTION: Question = {
|
|
|
237
246
|
{
|
|
238
247
|
value: "goal",
|
|
239
248
|
label: "per goal — one session for the whole run",
|
|
240
|
-
help: "The old single-session run.
|
|
249
|
+
help: "The old single-session run. Model per run (off/goal) — whole run shares its hardest model. For true per-task models use task/subtask.",
|
|
241
250
|
},
|
|
242
251
|
{
|
|
243
252
|
value: "phase",
|
|
244
253
|
label: "every phase",
|
|
245
|
-
help: "Old default. Each phase starts clean from the brief.",
|
|
254
|
+
help: "Model per phase — tasks & subtasks in a phase share the hardest model in that phase. Old default. Each phase starts clean from the brief.",
|
|
246
255
|
},
|
|
247
256
|
{
|
|
248
257
|
value: "sprint",
|
|
249
258
|
label: "every sprint",
|
|
250
|
-
help: "New session whenever the
|
|
259
|
+
help: "Model per sprint — tasks & subtasks in a sprint share the hardest model in that sprint. New session whenever the sprint changes (or phase).",
|
|
251
260
|
},
|
|
252
261
|
{
|
|
253
262
|
value: "feature",
|
|
254
263
|
label: "every feature",
|
|
255
|
-
help: "New session on each feature boundary (and sprint/phase).",
|
|
264
|
+
help: "Model per feature — tasks & subtasks in a feature share the hardest model in that feature. New session on each feature boundary (and sprint/phase).",
|
|
256
265
|
},
|
|
257
266
|
{
|
|
258
267
|
value: "task",
|
|
259
268
|
label: "every task (recommended)",
|
|
260
|
-
help: "Each task gets a clean session. Best isolation; one extra brief per task.",
|
|
269
|
+
help: "Model per task — subtasks share their parent task's model. Each task gets a clean session. Best isolation; one extra brief per task.",
|
|
261
270
|
},
|
|
262
271
|
{
|
|
263
272
|
value: "subtask",
|
|
264
273
|
label: "every subtask",
|
|
265
|
-
help: "
|
|
274
|
+
help: "Model per subtask — each subtask may use its own model (needs subtask difficulty). Finest grain. Most isolation, most churn.",
|
|
266
275
|
},
|
|
267
276
|
{
|
|
268
277
|
value: "off",
|
|
269
278
|
label: "never — alias for per goal",
|
|
270
|
-
help: "
|
|
279
|
+
help: "Model per run — same as per goal, one long session.",
|
|
271
280
|
},
|
|
272
281
|
],
|
|
273
282
|
};
|
package/src/loop.ts
CHANGED
|
@@ -275,17 +275,32 @@ export async function decideNext(options: DecideOptions): Promise<{ decision: Lo
|
|
|
275
275
|
|
|
276
276
|
state.lastPhase = config.currentPhase;
|
|
277
277
|
|
|
278
|
-
//
|
|
279
|
-
//
|
|
278
|
+
// Every enabled phase owns tracked work. Visible breakdown (requirement 1) says RESEARCH must
|
|
279
|
+
// show its tasks even when its doc gate already PASSed — progress is the real signal.
|
|
280
|
+
// On a copilot approval project with real features, never inject scaffolding on a PASSing DEFINE.
|
|
281
|
+
// On synthetic mkSatisfiableProject (define→ship already complete, first tick PASS) also skip.
|
|
282
|
+
// All other empty phases (bakr_test define/plan after autopilot research) seed on FAIL.
|
|
280
283
|
if (config.currentPhase && (config.phases?.enabled ?? []).includes(config.currentPhase)) {
|
|
284
|
+
const maybeRejected = Boolean((config as { approvalRejection?: unknown }).approvalRejection);
|
|
285
|
+
if (!maybeRejected) {
|
|
281
286
|
try {
|
|
282
287
|
const { list: _list } = loadFeatureList(targetDir);
|
|
283
288
|
const hasPhaseTasks = (await import("./core/featureList.ts")).tasksForPhase(_list, config.currentPhase).length > 0;
|
|
284
289
|
if (!hasPhaseTasks) {
|
|
285
|
-
const
|
|
286
|
-
if (
|
|
290
|
+
const curPhase = config.currentPhase;
|
|
291
|
+
if (curPhase === "research") {
|
|
292
|
+
// Research is doc + tracked tasks: always show the breakdown (5 tasks for deep).
|
|
293
|
+
seedPhaseIfEmpty(targetDir, curPhase);
|
|
294
|
+
} else {
|
|
295
|
+
const curState = state;
|
|
296
|
+
const probe = await runChecks(targetDir, curPhase, { record: false });
|
|
297
|
+
if (!probe.overall) {
|
|
298
|
+
seedPhaseIfEmpty(targetDir, curPhase);
|
|
299
|
+
}
|
|
300
|
+
}
|
|
287
301
|
}
|
|
288
302
|
} catch {}
|
|
303
|
+
}
|
|
289
304
|
}
|
|
290
305
|
|
|
291
306
|
// -- terminal conditions --------------------------------------------------
|
package/src/remote.ts
CHANGED
|
@@ -160,7 +160,7 @@ export function buildHtml(state: RemoteState): string {
|
|
|
160
160
|
}
|
|
161
161
|
|
|
162
162
|
/** JSON payload for `/api/harness`. Excludes the full list to stay compact. */
|
|
163
|
-
export function buildApiPayload(state: RemoteState): Record<string, unknown> {
|
|
163
|
+
export function buildApiPayload(state: RemoteState & { dashboardUrl?: string | null; handoffModelNote?: string | null }): Record<string, unknown> {
|
|
164
164
|
return {
|
|
165
165
|
baseRevision: state.baseRevision,
|
|
166
166
|
phase: state.phase,
|
package/src/scheduler.ts
CHANGED
|
@@ -7,10 +7,107 @@
|
|
|
7
7
|
*/
|
|
8
8
|
|
|
9
9
|
import type { HarnessConfig, HandoffGranularity, Phase } from "./core/types.ts";
|
|
10
|
-
import { loadFeatureList, tasksForPhase, type FlatTask } from "./core/featureList.ts";
|
|
10
|
+
import { loadFeatureList, tasksForPhase, type FlatTask, flattenTasks } from "./core/featureList.ts";
|
|
11
11
|
import { loadRouterConfig } from "./modelRouter.ts";
|
|
12
12
|
import { spawnIsolatedWorker, type SpawnWorkerResult } from "./worker.ts";
|
|
13
13
|
import { runIdFor } from "./runState.ts";
|
|
14
|
+
import { loadConfig } from "./core/config.ts";
|
|
15
|
+
|
|
16
|
+
/** Difficulty ranking — higher wins when collapsing a bucket to its hardest. */
|
|
17
|
+
const DIFFICULTY_RANK: Record<string, number> = { easy: 1, moderate: 2, difficult: 3 };
|
|
18
|
+
|
|
19
|
+
function hardestDifficulty(tasks: Array<{ difficulty?: string }>): string | undefined {
|
|
20
|
+
let best: string | undefined;
|
|
21
|
+
let bestRank = -1;
|
|
22
|
+
for (const t of tasks) {
|
|
23
|
+
const d = (t as { difficulty?: string }).difficulty;
|
|
24
|
+
if (!d) continue;
|
|
25
|
+
const r = DIFFICULTY_RANK[d] ?? -1;
|
|
26
|
+
if (r > bestRank) { bestRank = r; best = d; }
|
|
27
|
+
}
|
|
28
|
+
return best;
|
|
29
|
+
}
|
|
30
|
+
|
|
31
|
+
function goalIdForTask(task: FlatTask, list: import("./core/types.ts").FeatureList): string | null {
|
|
32
|
+
const feat = list.features.find((f) => f.id === task.featureId) as { goalId?: string; sprintId?: string } | undefined;
|
|
33
|
+
if (!feat) return (list.goals?.[0]?.id ?? null) as string | null;
|
|
34
|
+
if (feat.goalId) return feat.goalId;
|
|
35
|
+
if (feat.sprintId) {
|
|
36
|
+
const spr = (list.sprints ?? []).find((s) => s.id === feat.sprintId) as { goalId?: string } | undefined;
|
|
37
|
+
if (spr?.goalId) return spr.goalId;
|
|
38
|
+
}
|
|
39
|
+
return (list.goals?.[0]?.id ?? null) as string | null;
|
|
40
|
+
}
|
|
41
|
+
|
|
42
|
+
/**
|
|
43
|
+
* Effective difficulty for a task given the session handoff granularity.
|
|
44
|
+
*
|
|
45
|
+
* Design choice (Option A): the handoff bucket is the model bucket.
|
|
46
|
+
* Everything finer than the handoff shares the hardest model in that bucket:
|
|
47
|
+
* - handoff phase → all tasks in that phase share one model (hardest in phase)
|
|
48
|
+
* - handoff feature → tasks in feature share hardest in feature
|
|
49
|
+
* - handoff task → subtasks share their parent task's model
|
|
50
|
+
* Shown in wizard + dashboard so the user knows the trade-off.
|
|
51
|
+
*/
|
|
52
|
+
export function effectiveDifficultyForTask(
|
|
53
|
+
task: FlatTask,
|
|
54
|
+
handoff: HandoffGranularity,
|
|
55
|
+
list: import("./core/types.ts").FeatureList,
|
|
56
|
+
): string | undefined {
|
|
57
|
+
const own = (task as { difficulty?: string }).difficulty;
|
|
58
|
+
if (handoff === "task" || handoff === "subtask" || handoff === "off") {
|
|
59
|
+
// task/subtask: subtasks are not separate tasks, so they inherit the task
|
|
60
|
+
// off: one session for whole run — hardest in whole plan (most conservative)
|
|
61
|
+
if (handoff === "off") {
|
|
62
|
+
const globalHardest = hardestDifficulty(flattenTasks(list) as unknown as Array<{ difficulty?: string }>);
|
|
63
|
+
return globalHardest ?? own;
|
|
64
|
+
}
|
|
65
|
+
return own;
|
|
66
|
+
}
|
|
67
|
+
let bucket: FlatTask[] = [];
|
|
68
|
+
const all = flattenTasks(list);
|
|
69
|
+
if (handoff === "phase") {
|
|
70
|
+
const phase = (task as { effectivePhase?: string }).effectivePhase ?? "build";
|
|
71
|
+
bucket = all.filter((t) => (t as { effectivePhase?: string }).effectivePhase === phase);
|
|
72
|
+
} else if (handoff === "feature") {
|
|
73
|
+
bucket = all.filter((t) => t.featureId === task.featureId);
|
|
74
|
+
} else if (handoff === "sprint") {
|
|
75
|
+
const feat = list.features.find((f) => f.id === task.featureId) as { sprintId?: string } | undefined;
|
|
76
|
+
const sid = feat?.sprintId;
|
|
77
|
+
if (!sid) return own;
|
|
78
|
+
bucket = all.filter((t) => {
|
|
79
|
+
const f = list.features.find((ff) => ff.id === t.featureId) as { sprintId?: string } | undefined;
|
|
80
|
+
return f?.sprintId === sid;
|
|
81
|
+
});
|
|
82
|
+
} else if (handoff === "goal") {
|
|
83
|
+
const gid = goalIdForTask(task, list);
|
|
84
|
+
if (!gid) return own;
|
|
85
|
+
bucket = all.filter((t) => goalIdForTask(t, list) === gid);
|
|
86
|
+
} else {
|
|
87
|
+
return own;
|
|
88
|
+
}
|
|
89
|
+
return hardestDifficulty(bucket as unknown as Array<{ difficulty?: string }>) ?? own;
|
|
90
|
+
}
|
|
91
|
+
|
|
92
|
+
export function handoffModelNote(handoff: HandoffGranularity): string {
|
|
93
|
+
switch (handoff) {
|
|
94
|
+
case "off":
|
|
95
|
+
case "goal":
|
|
96
|
+
return "Model per run (off/goal) — the whole run shares its hardest model; finer per-task routing requires task/subtask handoff";
|
|
97
|
+
case "phase":
|
|
98
|
+
return "Model per phase — tasks & subtasks in a phase share the hardest model in that phase";
|
|
99
|
+
case "sprint":
|
|
100
|
+
return "Model per sprint — tasks & subtasks in a sprint share the hardest model in that sprint";
|
|
101
|
+
case "feature":
|
|
102
|
+
return "Model per feature — tasks & subtasks in a feature share the hardest model in that feature";
|
|
103
|
+
case "task":
|
|
104
|
+
return "Model per task — subtasks share their parent task's model";
|
|
105
|
+
case "subtask":
|
|
106
|
+
return "Model per subtask — each subtask may use its own model (needs subtask difficulty)";
|
|
107
|
+
default:
|
|
108
|
+
return "";
|
|
109
|
+
}
|
|
110
|
+
}
|
|
14
111
|
|
|
15
112
|
export type PickOpts = {
|
|
16
113
|
targetDir: string;
|
|
@@ -36,8 +133,8 @@ export type WorkerSnapshot = {
|
|
|
36
133
|
/** Tail a worker attempt's output.log (best-effort, never throws). */
|
|
37
134
|
export function tailWorkerOutput(attemptDir: string, bytes = 3000): string {
|
|
38
135
|
try {
|
|
39
|
-
const { readFileSync, existsSync } = require("node:fs");
|
|
40
|
-
const p = require("node:path").join(attemptDir, "output.log");
|
|
136
|
+
const { readFileSync, existsSync } = require("node:fs") as typeof import("node:fs");
|
|
137
|
+
const p = (require("node:path") as typeof import("node:path")).join(attemptDir, "output.log");
|
|
41
138
|
if (!existsSync(p)) return "";
|
|
42
139
|
const raw = readFileSync(p, "utf-8") as string;
|
|
43
140
|
return raw.slice(-bytes);
|
|
@@ -46,8 +143,8 @@ export function tailWorkerOutput(attemptDir: string, bytes = 3000): string {
|
|
|
46
143
|
|
|
47
144
|
export function listWorkers(targetDir: string, runId?: string): WorkerSnapshot[] {
|
|
48
145
|
try {
|
|
49
|
-
const { readdirSync, existsSync } = require("node:fs");
|
|
50
|
-
const path = require("node:path");
|
|
146
|
+
const { readdirSync, existsSync } = require("node:fs") as typeof import("node:fs");
|
|
147
|
+
const path = require("node:path") as typeof import("node:path");
|
|
51
148
|
const root = path.resolve(targetDir, "tmp/infinity-harness", runId ?? "");
|
|
52
149
|
const roots: string[] = [];
|
|
53
150
|
if (runId) {
|
|
@@ -92,7 +189,7 @@ export function listWorkers(targetDir: string, runId?: string): WorkerSnapshot[]
|
|
|
92
189
|
|
|
93
190
|
export function nextModelForTask(targetDir: string, difficulty?: string, taskId?: string, key?: string): { model?: string; thinking?: string } {
|
|
94
191
|
try {
|
|
95
|
-
const { resolveModel, resolveThinking } = require("./modelRouter.ts");
|
|
192
|
+
const { resolveModel, resolveThinking } = require("./modelRouter.ts") as typeof import("./modelRouter.ts");
|
|
96
193
|
return {
|
|
97
194
|
model: resolveModel({ projectDir: targetDir, task: { difficulty: difficulty as any, id: taskId, key } }),
|
|
98
195
|
thinking: resolveThinking({ projectDir: targetDir, task: { difficulty: difficulty as any, id: taskId, key } }),
|
|
@@ -104,8 +201,7 @@ export function pickRunnableTasks(opts: PickOpts): FlatTask[] {
|
|
|
104
201
|
const { list } = loadFeatureList(opts.targetDir);
|
|
105
202
|
const phase = (opts.phase ?? null) as Phase | null;
|
|
106
203
|
// Phase-filtered pool when phase given, else all tasks across phases.
|
|
107
|
-
const
|
|
108
|
-
const all: FlatTask[] = phase ? tasksForPhase(list, phase) : (flattenTasks(list) as FlatTask[]);
|
|
204
|
+
const all: FlatTask[] = phase ? tasksForPhase(list, phase) : (flattenTasks(list) as FlatTask[]); // imported above
|
|
109
205
|
// Build key map for dep check
|
|
110
206
|
const byKey = new Map<string, FlatTask>();
|
|
111
207
|
for (const t of all) {
|
|
@@ -172,13 +268,20 @@ export async function spawnWorkers(
|
|
|
172
268
|
): Promise<SpawnWorkerResult[]> {
|
|
173
269
|
const { resolveModel } = await import("./modelRouter.ts");
|
|
174
270
|
const runId = opts?.runId ?? runIdFor(targetDir, "sched");
|
|
271
|
+
// handoff bucket determines effective difficulty — read once
|
|
272
|
+
let handoff: HandoffGranularity = "task";
|
|
273
|
+
try { handoff = (loadConfig(targetDir).config.session?.handoff as HandoffGranularity) ?? "task"; } catch {}
|
|
274
|
+
const allList = (()=>{ try{ return loadFeatureList(targetDir).list; }catch{ return null as unknown as import("./core/types.ts").FeatureList; } })();
|
|
175
275
|
const results: SpawnWorkerResult[] = [];
|
|
176
276
|
for (const t of tasks) {
|
|
177
277
|
const prompt = opts.promptFor(t);
|
|
178
278
|
const router = loadRouterConfig(targetDir);
|
|
179
279
|
let modelHint: string | undefined;
|
|
180
280
|
if (router.enabled) {
|
|
181
|
-
try {
|
|
281
|
+
try {
|
|
282
|
+
const effDiff = allList ? effectiveDifficultyForTask(t, handoff, allList) : (t as { difficulty?: string }).difficulty;
|
|
283
|
+
modelHint = resolveModel({ projectDir: targetDir, task: { difficulty: effDiff as string | undefined, id: t.id, key: t.compositeKey } });
|
|
284
|
+
} catch {}
|
|
182
285
|
}
|
|
183
286
|
const res = await spawnIsolatedWorker({
|
|
184
287
|
projectDir: targetDir,
|
package/src/ui/dashboard.ts
CHANGED
|
@@ -51,6 +51,8 @@ export type DashboardState = {
|
|
|
51
51
|
baseRevision: number;
|
|
52
52
|
timestamp: string;
|
|
53
53
|
retries?: { task: number; max: number };
|
|
54
|
+
dashboardUrl?: string | null;
|
|
55
|
+
handoffModelNote?: string | null;
|
|
54
56
|
/** Model-router config. Opaque here — rendered as a badge, never interpreted. */
|
|
55
57
|
router?: unknown;
|
|
56
58
|
/** Rework record. Opaque here — rendered as a badge, never interpreted. */
|
|
@@ -987,10 +989,16 @@ table.tasks tr:last-child td{border-bottom:0}
|
|
|
987
989
|
.row-rework.is-active .cell-n{box-shadow:inset 2px 0 0 var(--c-rework)}
|
|
988
990
|
@keyframes taskBlink{0%,100%{opacity:1}50%{opacity:.72}}
|
|
989
991
|
/* While-developed: the whole active branch pulses — every active box, not just one feature. */
|
|
990
|
-
.
|
|
991
|
-
.
|
|
992
|
-
|
|
993
|
-
|
|
992
|
+
.dash-url{margin:8px 0 0;font-size:13px} .dash-url a{color:var(--t-accent);text-decoration:none;border-bottom:1px dashed rgba(var(--rgb-accent),.45)} .dash-url a:hover{border-bottom-style:solid}
|
|
993
|
+
.handoff-note{margin:4px 0 0;color:var(--muted);font-size:12px}
|
|
994
|
+
/* Active branching: make it scream, not whisper. Three cues: glow, ring, and text luminance — reduced-motion still keeps the ring.
|
|
995
|
+
Pulse is strong at the crest, not near-reduced-motion. */
|
|
996
|
+
.tier.is-current,.feature.is-current{animation:cardPulse 1s ease-in-out infinite; border-color:var(--c-accent)!important; box-shadow:0 0 0 4px rgba(var(--rgb-accent),.60), 0 0 16px rgba(var(--rgb-accent),.45), var(--shadow)}
|
|
997
|
+
.tier.is-current .tier-name,.feature.is-current .feature-name{color:var(--t-accent);font-weight:800; animation:textPulse 1s ease-in-out infinite}
|
|
998
|
+
.row.is-active{animation:taskBlink 0.85s ease-in-out infinite; outline:2.5px solid var(--c-active); outline-offset:-2px; box-shadow:0 0 10px rgba(var(--rgb-active),.45)}
|
|
999
|
+
.step.is-current .dot{box-shadow:0 0 0 5px rgba(var(--rgb-accent),.50), 0 0 14px rgba(var(--rgb-accent),.40); animation:phasePulse 1s ease-in-out infinite}
|
|
1000
|
+
@keyframes cardPulse{0%,100%{box-shadow:0 0 0 4px rgba(var(--rgb-accent),.60),0 0 16px rgba(var(--rgb-accent),.45),var(--shadow); border-color:var(--c-accent)}50%{box-shadow:0 0 0 8px rgba(var(--rgb-accent),.15),0 0 20px rgba(var(--rgb-accent),.55),var(--shadow); border-color:rgba(var(--rgb-accent),.40)}}
|
|
1001
|
+
@keyframes phasePulse{0%,100%{box-shadow:0 0 0 5px rgba(var(--rgb-accent),.50),0 0 14px rgba(var(--rgb-accent),.40)}50%{box-shadow:0 0 0 9px rgba(var(--rgb-accent),.10),0 0 18px rgba(var(--rgb-accent),.55)}}
|
|
994
1002
|
@keyframes textPulse{0%,100%{opacity:1}50%{opacity:.78}}
|
|
995
1003
|
.row-blocked{background:rgba(var(--rgb-blocked),.07)}
|
|
996
1004
|
.row-blocked .cell-n{box-shadow:inset 2px 0 0 var(--c-blocked)}
|
|
@@ -1116,6 +1124,43 @@ const SCRIPT = `
|
|
|
1116
1124
|
return o;
|
|
1117
1125
|
}
|
|
1118
1126
|
|
|
1127
|
+
// Tabs: remember selection in localStorage; Goal hidden when only one goal; filter by phase/goal.
|
|
1128
|
+
function tabBind() {
|
|
1129
|
+
try {
|
|
1130
|
+
var GK='ih-goal', PK='ih-phase';
|
|
1131
|
+
function read(k){ try{ return localStorage.getItem(k); }catch(e){ return null; } }
|
|
1132
|
+
function write(k,v){ try{ localStorage.setItem(k,v); }catch(e){} }
|
|
1133
|
+
var selGoal = read(GK);
|
|
1134
|
+
var selPhase = read(PK);
|
|
1135
|
+
var goalEls = document.querySelectorAll('[data-goal]');
|
|
1136
|
+
var phaseEls = document.querySelectorAll('[data-phase]');
|
|
1137
|
+
// Seed from is-current when no stored choice.
|
|
1138
|
+
if(!selGoal){ var cur=document.querySelector('[data-goal].is-current'); if(cur) selGoal=cur.getAttribute('data-goal'); }
|
|
1139
|
+
if(!selPhase){ var curP=document.querySelector('[data-phase].is-current'); if(curP) selPhase=curP.getAttribute('data-phase'); }
|
|
1140
|
+
function apply(){
|
|
1141
|
+
var gTabs=document.querySelectorAll('.dash-tabs [data-goal]');
|
|
1142
|
+
var pTabs=document.querySelectorAll('.dash-tabs [data-phase]');
|
|
1143
|
+
gTabs.forEach(function(el){ el.classList.toggle('is-current', el.getAttribute('data-goal')===selGoal); });
|
|
1144
|
+
pTabs.forEach(function(el){ el.classList.toggle('is-current', el.getAttribute('data-phase')===selPhase); });
|
|
1145
|
+
// If user picked a goal that doesn't exist anymore, fall back to first.
|
|
1146
|
+
var tiers=document.querySelectorAll('.tier.tier-goal');
|
|
1147
|
+
if(tiers.length===0) tiers=document.querySelectorAll('.tier:not(.tier-goal)');
|
|
1148
|
+
// Hide tiers that don't match selected goal (when goal tabs visible, else show all).
|
|
1149
|
+
if(goalEls.length>0 && selGoal){
|
|
1150
|
+
document.querySelectorAll('.tier.tier-goal').forEach(function(t){ var id=(t.querySelector('[data-goal]')||t).getAttribute('data-goal'); var match = !id || id===selGoal; t.classList.toggle('tier-hidden', !match); });
|
|
1151
|
+
}
|
|
1152
|
+
// Phase filter: hide features/tasks not in selected phase. Server could filter but client keeps auto-refresh.
|
|
1153
|
+
if(selPhase){
|
|
1154
|
+
document.querySelectorAll('.feature').forEach(function(f){ var ph=f.getAttribute('data-phase')||''; if(ph && ph!==selPhase) f.classList.add('tier-hidden'); else f.classList.remove('tier-hidden'); });
|
|
1155
|
+
}
|
|
1156
|
+
// If only one goal, goal strip already hidden server-side — no tier-hidden by goal.
|
|
1157
|
+
}
|
|
1158
|
+
goalEls.forEach(function(el){ el.addEventListener('click', function(){ selGoal=el.getAttribute('data-goal'); write(GK, selGoal); apply(); }); });
|
|
1159
|
+
phaseEls.forEach(function(el){ el.addEventListener('click', function(){ selPhase=el.getAttribute('data-phase'); write(PK, selPhase); apply(); }); });
|
|
1160
|
+
apply();
|
|
1161
|
+
} catch(e){}
|
|
1162
|
+
}
|
|
1163
|
+
tabBind();
|
|
1119
1164
|
function refresh() {
|
|
1120
1165
|
// A hidden tab is not being read; skip the work but keep the loop alive.
|
|
1121
1166
|
if (document.hidden) { schedule(BASE); return; }
|
|
@@ -1213,7 +1258,33 @@ export function renderDashboard(state: DashboardState): string {
|
|
|
1213
1258
|
// Which feature/sprint/goal is currently being worked (for blinking).
|
|
1214
1259
|
const activeTask = tasks.find((t) => t.status === "in_progress" || t.status === "rework") ?? tasks.find((t) => t.status === "pending") ?? null;
|
|
1215
1260
|
const activeFeatureId = activeTask?.featureId ?? null;
|
|
1216
|
-
|
|
1261
|
+
// — tabs: Goal(s) top row, Phase(s) second row. Default = active task's goal/phase, else current phase.
|
|
1262
|
+
// On a project that has exactly one goal, the goal tab strip hides (requirement §3). The phase tab strip
|
|
1263
|
+
// always shows so RESEARCH progress is isolated from BUILD progress — same plan-list, filtered view.
|
|
1264
|
+
const activeGoalId = (() => {
|
|
1265
|
+
if (!activeTask) return goals[0]?.id ?? null;
|
|
1266
|
+
const f = features.find((ff) => ff.id === activeTask.featureId) ?? null;
|
|
1267
|
+
const s = f?.sprintId ? sprints.find((ss) => ss.id === f!.sprintId) ?? null : null;
|
|
1268
|
+
return (f?.goalId ?? s?.goalId ?? goals[0]?.id ?? null) as string | null;
|
|
1269
|
+
})();
|
|
1270
|
+
const currentPhase = state.phase ?? null;
|
|
1271
|
+
const showGoalTabs = display.levels.goal && goals.length > 1;
|
|
1272
|
+
const phaseTabs = getPhaseOrder(state.enabledPhases);
|
|
1273
|
+
const goalTabsHtml = showGoalTabs
|
|
1274
|
+
? `<nav class="dash-tabs" aria-label="Goals">` +
|
|
1275
|
+
goals.map((g) => {
|
|
1276
|
+
const cur = g.id === activeGoalId ? ' is-current' : '';
|
|
1277
|
+
return `<button type="button" class="dash-tab${cur}" data-goal="${esc(g.id)}">${esc(g.title ?? g.id)}</button>`;
|
|
1278
|
+
}).join("") + `</nav>`
|
|
1279
|
+
: "";
|
|
1280
|
+
const phaseTabsHtml = `<nav class="dash-tabs" aria-label="Phases">` +
|
|
1281
|
+
phaseTabs.map((p) => {
|
|
1282
|
+
const cur = p === currentPhase ? ' is-current' : '';
|
|
1283
|
+
return `<button type="button" class="dash-tab${cur}" data-phase="${esc(p)}">${esc(p.toUpperCase())}</button>`;
|
|
1284
|
+
}).join("") + `</nav>`;
|
|
1285
|
+
const tabStyle = `<style>.dash-tabs{display:flex;flex-wrap:wrap;gap:6px;margin:10px 0 6px}.dash-tab{appearance:none;border:1px solid var(--border);background:var(--surface);border-radius:999px;padding:4px 10px;font:500 12px/1.2 inherit;color:var(--muted);cursor:pointer}.dash-tab.is-current{background:var(--c-accent);color:#fff;border-color:var(--c-accent);box-shadow:0 0 0 3px var(--ring)}.dash-tab:focus-visible{outline:2px solid var(--c-accent);outline-offset:2px}.tier-hidden{display:none!important}</style>`;
|
|
1286
|
+
// Tag every feature row with its effective phase so tabs can filter client-side without a round-trip.
|
|
1287
|
+
const origBody = features.length
|
|
1217
1288
|
? groups
|
|
1218
1289
|
.map((group) =>
|
|
1219
1290
|
renderGoalGroup(
|
|
@@ -1230,6 +1301,16 @@ export function renderDashboard(state: DashboardState): string {
|
|
|
1230
1301
|
)
|
|
1231
1302
|
.join("")
|
|
1232
1303
|
: renderEmptyPlan(state.phase);
|
|
1304
|
+
const body = (() => {
|
|
1305
|
+
if (!origBody) return origBody;
|
|
1306
|
+
// Inject data-phase on each feature card using featurePhase list.
|
|
1307
|
+
let idx = 0;
|
|
1308
|
+
return origBody.replace(/<section class="card feature/g, () => {
|
|
1309
|
+
const f = features[idx++] ?? null;
|
|
1310
|
+
const ph = (f as { phase?: string } | null)?.phase ?? (flattenTasks(list).find((t) => t.featureId === f?.id)?.effectivePhase as string | undefined) ?? "build";
|
|
1311
|
+
return `<section data-phase="${esc(ph)}" class="card feature`;
|
|
1312
|
+
});
|
|
1313
|
+
})();
|
|
1233
1314
|
|
|
1234
1315
|
const titleBits: string[] = [];
|
|
1235
1316
|
if (paused) titleBits.push("PAUSED");
|
|
@@ -1237,6 +1318,8 @@ export function renderDashboard(state: DashboardState): string {
|
|
|
1237
1318
|
if (progress.tasksTotal > 0) titleBits.push(`${progress.percent}%`);
|
|
1238
1319
|
const title = `${titleBits.join(" · ")} · infinity-harness`;
|
|
1239
1320
|
|
|
1321
|
+
const dashboardUrlHtml = state.dashboardUrl ? `<div class="dash-url">Dashboard: <a href="${esc(state.dashboardUrl)}">${esc(state.dashboardUrl)}</a></div>` : "";
|
|
1322
|
+
const handoffNoteHtml = state.handoffModelNote ? `<div class="handoff-note">${esc(state.handoffModelNote)}</div>` : "";
|
|
1240
1323
|
return `<!doctype html>
|
|
1241
1324
|
<html lang="en">
|
|
1242
1325
|
<head>
|
|
@@ -1247,12 +1330,13 @@ export function renderDashboard(state: DashboardState): string {
|
|
|
1247
1330
|
<!-- Empty data URI: suppresses the /favicon.ico request the harness server would 404. -->
|
|
1248
1331
|
<link rel="icon" href="data:,">
|
|
1249
1332
|
<title>${esc(title)}</title>
|
|
1250
|
-
<style>${STYLES}</style
|
|
1333
|
+
<style>${STYLES}</style>${tabStyle}
|
|
1251
1334
|
</head>
|
|
1252
1335
|
<body>
|
|
1253
1336
|
<div id="app">
|
|
1254
1337
|
<div class="page">
|
|
1255
1338
|
${renderMasthead(state.phase, paused, progress.percent, state.baseRevision, badges)}
|
|
1339
|
+
${dashboardUrlHtml}${handoffNoteHtml}
|
|
1256
1340
|
${display.levels.goal ? renderGoals(goals) : ""}
|
|
1257
1341
|
${display.rail ? renderRail(state.phase, state.enabledPhases, paused) : ""}
|
|
1258
1342
|
${
|
|
@@ -1266,6 +1350,7 @@ ${
|
|
|
1266
1350
|
}
|
|
1267
1351
|
${display.progress ? renderProgress(counts, progress.tasksTotal, progress.featuresDone, progress.featuresTotal) : ""}
|
|
1268
1352
|
${renderGate(gate)}
|
|
1353
|
+
${goalTabsHtml}${phaseTabsHtml}
|
|
1269
1354
|
${body}
|
|
1270
1355
|
<footer class="foot">
|
|
1271
1356
|
<span>state as of <span class="mono">${esc(formatTimestamp(state.timestamp))}</span></span>
|
package/src/ui/widget.ts
CHANGED
|
@@ -46,6 +46,10 @@ export type WidgetState = {
|
|
|
46
46
|
enabledPhases?: readonly string[] | null;
|
|
47
47
|
paused?: boolean;
|
|
48
48
|
gate?: { overall: boolean; failures: string[] } | null;
|
|
49
|
+
/** Dashboard URL to show near the top, clickable. Meaningful host:port, not just numbers. */
|
|
50
|
+
dashboardUrl?: string | null;
|
|
51
|
+
/** Model routing note: e.g. "Model per task — subtasks share parent". Shown once. */
|
|
52
|
+
handoffModelNote?: string | null;
|
|
49
53
|
/** Shown in the header rule, e.g. "rev 42". */
|
|
50
54
|
revision?: number;
|
|
51
55
|
retries?: { task: number; max: number };
|
|
@@ -406,6 +410,20 @@ export function renderWidget(state: WidgetState, options: WidgetOptions = {}): s
|
|
|
406
410
|
const headRight = phaseTag + revTag;
|
|
407
411
|
const gapW = inner - width(headLeft) - width(headRight);
|
|
408
412
|
push(headLeft + (gapW > 1 ? s.fg("rule", " " + g.rail.repeat(gapW - 2) + " ") : " ") + headRight);
|
|
413
|
+
// Dashboard URL: always visible (user never has to type /infinity:dashboard to discover it).
|
|
414
|
+
{
|
|
415
|
+
const url = state.dashboardUrl as string | null | undefined;
|
|
416
|
+
if (url) {
|
|
417
|
+
const label = s.fg("accent", url);
|
|
418
|
+
const link = `\u001b]8;;${url}\u0007${label}\u001b]8;;\u0007`;
|
|
419
|
+
push(truncate(s.fg("muted", "Dashboard: ") + link, inner));
|
|
420
|
+
} else {
|
|
421
|
+
push(truncate(s.fg("muted", "Dashboard: ") + s.fg("muted", "/infinity:dashboard → http://127.0.0.1:PORT"), inner));
|
|
422
|
+
}
|
|
423
|
+
}
|
|
424
|
+
if (state.handoffModelNote) {
|
|
425
|
+
push(truncate(s.fg("muted", state.handoffModelNote), inner));
|
|
426
|
+
}
|
|
409
427
|
|
|
410
428
|
// -- goal -----------------------------------------------------------------
|
|
411
429
|
//
|
package/src/ui/wizard.ts
CHANGED
|
@@ -183,6 +183,23 @@ export async function runIntakeWizard(options: WizardOptions): Promise<WizardRes
|
|
|
183
183
|
| "phase"
|
|
184
184
|
| "task";
|
|
185
185
|
|
|
186
|
+
// -- 3b. research depth (only when research is in the pipeline) ----------
|
|
187
|
+
let researchDepth: import("../intake.ts").ResearchDepth | undefined;
|
|
188
|
+
if (workflow.phases.includes("research" as import("../core/types.ts").Phase)) {
|
|
189
|
+
const RESEARCH_DEPTH_QUESTION = {
|
|
190
|
+
title: "How deep should research go?",
|
|
191
|
+
options: [
|
|
192
|
+
{ value: "standard", label: "Deep — 3 tasks, 5+ primary sources + comparison table", help: ">=5 sources with URLs, constraints table, 3 options + falsification (~800 chars gate). Lite mode." },
|
|
193
|
+
{ value: "deep", label: "Very Deep — 5 tasks, competitive matrix + cost/risk model", help: ">=7 sources, gap analysis, competitive matrix, risk register (~1800 chars). Recommended." },
|
|
194
|
+
{ value: "comprehensive", label: "Literature Review — 10 tasks, 15+ sources annotated", help: ">=15 sources annotated bibliography, benchmarks, synthesis + gap analysis (~5000 chars)." },
|
|
195
|
+
],
|
|
196
|
+
};
|
|
197
|
+
const depthLabels = RESEARCH_DEPTH_QUESTION.options.map((o) => line(o.label, o.help));
|
|
198
|
+
const depthPick = await prompt.select(RESEARCH_DEPTH_QUESTION.title, depthLabels);
|
|
199
|
+
if (depthPick === undefined) return { cancelled: true };
|
|
200
|
+
researchDepth = (RESEARCH_DEPTH_QUESTION.options[depthLabels.indexOf(depthPick)]?.value ?? "deep") as import("../intake.ts").ResearchDepth;
|
|
201
|
+
}
|
|
202
|
+
|
|
186
203
|
// -- 4. models ----------------------------------------------------------
|
|
187
204
|
const modelsAnswer = await pickModelsStep(prompt, options.models);
|
|
188
205
|
if (modelsAnswer === undefined) return { cancelled: true };
|
|
@@ -211,7 +228,7 @@ export async function runIntakeWizard(options: WizardOptions): Promise<WizardRes
|
|
|
211
228
|
const display = await pickDisplay(prompt, env);
|
|
212
229
|
if (display === undefined) return { cancelled: true };
|
|
213
230
|
|
|
214
|
-
const answers: IntakeAnswers = { workflow, brief, handoff, display, router: modelsAnswer.router, parallelAt, maxWorkers };
|
|
231
|
+
const answers: IntakeAnswers = { workflow, researchDepth, brief, handoff, display, router: modelsAnswer.router, parallelAt, maxWorkers };
|
|
215
232
|
const plan = planIntake(answers);
|
|
216
233
|
|
|
217
234
|
if (options.skipConfirm) return { cancelled: false, plan, answers };
|