@pennixrv/trellis 0.7.0-beta.8 → 0.7.0-beta.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,9 @@
1
+ {
2
+ "version": "0.7.0-beta.9",
3
+ "description": "Pennix v0.7 beta planning and evidence governance",
4
+ "breaking": false,
5
+ "recommendMigrate": false,
6
+ "changelog": "**Features:**\n- Require complex analysis to use normal planning and a sealed decision chain.\n- Add evidence units and schema-v2 subnode reports with coordinator-only acceptance.",
7
+ "migrations": [],
8
+ "notes": "Run npm install -g @pennixrv/trellis@beta then trellis update. No --migrate required."
9
+ }
@@ -50,12 +50,15 @@ and its status; the Codex supervisor projects that reply into the durable
50
50
  Channel message and `done` events. The durable JSON file carries the
51
51
  reviewable result.
52
52
 
53
- The subnode copies identity, scope, and lens from `brief.json`. The minimal
53
+ The subnode copies identity, scope, and lens from `brief.json`. Schema version 2
54
+ requires one `scope_assessment` entry for every brief scope item, structured
55
+ findings with evidence references, and typed uncertainties or corrections when
56
+ present. Schema version 1 is retired and the validator rejects it. The minimal
54
57
  complete report is:
55
58
 
56
59
  ```json
57
60
  {
58
- "schema_version": 1,
61
+ "schema_version": 2,
59
62
  "task_id": "task-id-from-brief",
60
63
  "work_id": "work-id-from-brief",
61
64
  "subnode_id": "subnode-id-from-brief",
@@ -63,6 +66,14 @@ complete report is:
63
66
  "status": "complete",
64
67
  "scope": ["exact scope copied from brief"],
65
68
  "lens": "exact lens copied from brief",
69
+ "scope_assessment": [
70
+ {
71
+ "scope": "exact scope item from brief",
72
+ "status": "covered",
73
+ "conclusion": "What this scope establishes.",
74
+ "evidence_ids": ["stable-evidence-id"]
75
+ }
76
+ ],
66
77
  "evidence": [
67
78
  {
68
79
  "id": "stable-evidence-id",
@@ -70,12 +81,30 @@ complete report is:
70
81
  "summary": "What this independently reviewable evidence establishes."
71
82
  }
72
83
  ],
73
- "findings": ["Bounded conclusion."],
84
+ "findings": [
85
+ {
86
+ "id": "finding-id",
87
+ "conclusion": "Bounded conclusion.",
88
+ "evidence_ids": ["stable-evidence-id"]
89
+ }
90
+ ],
74
91
  "uncertainties": [],
75
92
  "corrections": []
76
93
  }
77
94
  ```
78
95
 
96
+ Append a checkpoint marker to `worklog.md` when a material unit is complete or
97
+ blocked. It is a bounded recovery projection, not coordinator acceptance:
98
+
99
+ ```text
100
+ <!-- trellis-checkpoint: {"id":"checkpoint-1","covered_scope":["exact scope item"],"evidence_ids":["stable-evidence-id"],"conclusion_or_blocker":"Current conclusion or blocker.","unknowns":[],"safe_resume_point":"Next safe action."} -->
101
+ ```
102
+
103
+ The validator returns `review_concern` for missing or incomplete checkpoint
104
+ coverage and for inconclusive scope assessments so the coordinator can inspect
105
+ them. Identity, schema, path, and malformed-structure failures remain hard
106
+ errors. A concern is never automatic acceptance or rejection.
107
+
79
108
  For `blocked`, `incomplete`, or `error`, include the same base fields plus a
80
109
  `completed_scope` list (empty when no assigned scope started) and a non-empty
81
110
  `blocker` string. Never use
@@ -24,13 +24,13 @@ Shows the Phase Index (Plan / Execute / Finish) with routing + skill mapping.
24
24
 
25
25
  `get_context.py` shows the active task's `status` field. Route by `status` + artifact presence. This command replaces the user needing to remember the Trellis flow; it does not itself approve implementation.
26
26
 
27
- - `status=planning` + `task.json.meta.delivery_mode = "analysis_only"` → complete the PRD's bounded evidence work, verify its acceptance criteria and no-change boundary, then commit task artifacts and archive directly. Do not run `task.py start`; a protected-target change requires a separate change-bearing task.
27
+ - `status=planning` + `task.json.meta.delivery_mode = "analysis_only"` → first confirm the task still satisfies the bounded evidence-only eligibility rule; then complete the PRD's evidence work, verify its acceptance criteria and no-change boundary, and archive directly. Do not run `task.py start`; a protected-target change requires a separate change-bearing task.
28
28
  - `status=planning` + no `prd.md` → **1.1** (load `trellis-brainstorm`)
29
29
  - `status=planning` + a recorded `decision-needed` or unsealed decision chain → return to the planning frontier and load `pennix-decision-gates` when independent material questions can be batched.
30
30
  - `status=in_progress` + a material unresolved decision → record the reason and run `task.py replan <task> "<reason>"`; do not ask a native question during implementation.
31
31
  - `status=planning` + `prd.md` only → decide whether the task is lightweight or complex. Lightweight can move to **1.4** review; complex returns to **1.1** to add `design.md` + `implement.md`.
32
32
  - `status=planning` + complex artifacts complete + sub-agent jsonl not curated (empty, or only a legacy `_example` placeholder row) → **1.3**
33
- - `status=planning` + required artifacts complete + required jsonl curated or inline mode → **1.4** (ask for start review; only run `task.py start` after user confirms)
33
+ - `status=planning` + required artifacts complete + required jsonl curated or inline mode → run the Planning Seal closure pass, then **1.4** (ask for start review; only run `task.py start` after user confirms)
34
34
  - `status=in_progress` + implementation not started → **2.1**
35
35
  - `status=in_progress` + implementation done, not yet checked → **2.2**
36
36
  - `status=in_progress` + check passed → **3.3** (spec update) → **3.4** (commit)
@@ -12,6 +12,8 @@ While any user-owned product, scope, UX, compatibility, risk, or acceptance deci
12
12
 
13
13
  When `task.json.meta.delivery_mode = "analysis_only"` exactly and the PRD names a bounded evidence deliverable plus a no-change boundary for product source, runtime configuration, deployment, credentials, and external systems, task-creation consent authorizes that evidence work. Do not require a second planning approval or run `task.py start`: perform the declared research, audit, or design work while status remains `planning`, record the evidence, verify acceptance criteria and the boundary, commit task artifacts, and archive directly. If the evidence recommends a protected-target change, record it and create a separate change-bearing task before doing it.
14
14
 
15
+ This exception is eligible only for a bounded evidence deliverable with no material user decision, design or implementation plan, cross-owner coordination, security or deployment change, release or credential action, or protected downstream task. Calling work "research", deferring source edits, or working in an audit/root repository does not make it analysis-only. If any of those conditions apply, use the normal complex planning and implementation-approval path.
16
+
15
17
  All other tasks follow the planning and implementation approval gates below.
16
18
 
17
19
  ## Non-Negotiable Evidence Rule
@@ -24,6 +26,10 @@ Do not ask the user to confirm facts that the repository can answer. Ask only fo
24
26
 
25
27
  Repository evidence establishes current behavior and technical constraints. The user's intended behavior, feature scope boundaries, and UX preferences are never answerable by repository evidence alone, even when an existing pattern exists; existing patterns are options and recommendation evidence, not decisions.
26
28
 
29
+ ## Evidence Units For Read-Heavy Work
30
+
31
+ When research, audit, review, or investigation is too large to leave one independently useful conclusion in the current bounded session, split it into evidence units. Each unit must have one question or scope, a minimal evidence range, a destination artifact, and a stop condition; write its facts, conclusion or blocker, unknowns, and recovery point before starting another unit. Size units so one normal context window can finish and persist one useful result; do not promise an exact token or time limit. Routine navigation and transient tool output do not need an artifact. Create a child task only when the unit has an independent owner, lifecycle, and acceptance contract.
32
+
27
33
  ---
28
34
 
29
35
  Use this skill during Phase 1 planning to turn the user's request into clear requirements and planning artifacts.
@@ -54,10 +60,10 @@ Use a concise title from the user's request. Both the title and `--description`
54
60
  - product intent still needed from the user
55
61
  - scope or risk decisions still needed from the user
56
62
  - likely out-of-scope items
57
- 4. If user-owned decisions remain, calculate the independent frontier. Use `pennix-decision-gates` for a bounded batch when two or more independent material decisions are ready; otherwise ask the single highest-value question. Include recommendation and trade-off, then stop the turn for native input.
58
- 5. After each answer batch, update `prd.md`, record the selected decisions, recheck evidence and conflicts, and repeat from step 2. Do not create a second Trellis lifecycle for the same decision chain.
63
+ 4. If user-owned decisions remain, calculate the independent frontier. Use `pennix-decision-gates` for a bounded batch when two or more independent material decisions are ready; otherwise ask the single highest-value question. Include recommendation and trade-off. Yield only while the answer is unavailable.
64
+ 5. When the host returns the current continuation's answer, immediately persist it in `prd.md` or the decision artifact, recheck evidence and conflicts, recalculate the frontier, and continue the same planning loop. Do not create a second Trellis lifecycle for the same decision chain. Stop only for a new unresolved frontier, a real capability or authority block, or a final sealed summary awaiting implementation approval.
59
65
  6. When no user-owned decision remains, create or update `design.md` and `implement.md` for complex tasks.
60
- 7. Run the requirement convergence gate, then the PRD convergence pass.
66
+ 7. Run the requirement convergence gate, then the PRD convergence pass. Finish with one Planning Seal closure pass.
61
67
  8. Present the final planning summary and stop. Do not run `task.py start` or edit product code in the same turn.
62
68
  9. Only a subsequent user message that explicitly approves the latest planning summary authorizes `task.py start` and implementation. If implementation reveals a material unresolved decision, record `decision-needed`, run `task.py replan <task> "<reason>"`, and return through this planning flow; do not open a popup during implementation.
63
69
 
@@ -140,6 +146,8 @@ Lightweight tasks may omit `design.md` and `implement.md`; they may not skip evi
140
146
 
141
147
  The final planning summary must show Goal, In Scope, Out of Scope, Acceptance Criteria, Key Decisions, relevant Risks or Deferred Items, and artifact status.
142
148
 
149
+ The Planning Seal closure pass must reconcile `task.json`, `prd.md`, `design.md`, `implement.md`, research, decision records, and manifests; verify the actual modification targets and branches, ordered dependencies and release steps, validation and rollback, dynamic-fact dispositions and replan triggers, and that every material decision has an owner and a fixed outcome. Remove static ambiguity before implementation: no `TBD`, `TODO`, `decision-needed`, unowned option, unspecified branch, open implementation path, validation gap, or conditional acceptance may remain. A material discovery invalidates the seal and returns to planning; implementation may consume only a sealed plan.
150
+
143
151
  ## Artifact Rules
144
152
 
145
153
  `prd.md` records requirements and acceptance:
@@ -12,6 +12,10 @@ For every non-trivial task, the user must respond at least once after the initia
12
12
 
13
13
  While any user-owned product, scope, UX, compatibility, risk, or acceptance decision remains unresolved, keep the task in planning. First inventory evidence and decision dependencies. If at least two independent material decisions remain and `pennix-decision-gates` is available, delegate one bounded batch of up to three frontier questions; otherwise ask the single highest-value question. Do not edit product code, dispatch implementation, or run `task.py start` until the decision chain is sealed.
14
14
 
15
+ ## Analysis-Only Exception
16
+
17
+ When `task.json.meta.delivery_mode = "analysis_only"` exactly and the PRD names a bounded evidence deliverable plus a no-change boundary for product source, runtime configuration, deployment, credentials, and external systems, task-creation consent authorizes that evidence work. Keep status `planning`, record and verify the evidence, commit task artifacts, and archive directly; do not run `task.py start` or wait for a second implementation approval. This exception is eligible only when there is no material user decision, design or implementation plan, cross-owner coordination, security or deployment change, release or credential action, or protected downstream task. Otherwise use normal complex planning.
18
+
15
19
  ## Non-Negotiable Evidence Rule
16
20
 
17
21
  If a question can be answered by exploring the codebase, explore the codebase instead.
@@ -22,6 +26,10 @@ Do not ask the user to confirm facts that the repository can answer. Ask only fo
22
26
 
23
27
  Repository evidence establishes current behavior and technical constraints. The user's intended behavior, feature scope boundaries, and UX preferences are never answerable by repository evidence alone, even when an existing pattern exists; existing patterns are options and recommendation evidence, not decisions.
24
28
 
29
+ ## Evidence Units For Read-Heavy Work
30
+
31
+ When research, audit, review, or investigation is too large for one independently useful conclusion in the current session, split it into evidence units. Each unit has one question or scope, a minimal evidence range, a destination artifact, and a stop condition; persist facts, conclusion or blocker, unknowns, and a recovery point before starting another unit. Size each unit for one normal context window without promising an exact token or time limit.
32
+
25
33
  ---
26
34
 
27
35
  Use this skill during Phase 1 planning to turn the user's request into clear requirements and planning artifacts.
@@ -52,10 +60,10 @@ Use a concise title from the user's request. Both the title and `--description`
52
60
  - product intent still needed from the user
53
61
  - scope or risk decisions still needed from the user
54
62
  - likely out-of-scope items
55
- 4. If user-owned decisions remain, calculate the independent frontier. Use `pennix-decision-gates` for a bounded batch when two or more independent material decisions are ready; otherwise ask the single highest-value question. Include recommendation and trade-off, then stop the turn for native input.
56
- 5. After each answer batch, update `prd.md`, record the selected decisions, recheck evidence and conflicts, and repeat from step 2. Do not create a second Trellis lifecycle for the same decision chain.
63
+ 4. If user-owned decisions remain, calculate the independent frontier. Use `pennix-decision-gates` for a bounded batch when two or more independent material decisions are ready; otherwise ask the single highest-value question. Include recommendation and trade-off. Yield only while the answer is unavailable.
64
+ 5. When the host returns the current continuation's answer, persist it in `prd.md` or the decision artifact, recheck evidence and conflicts, recalculate the frontier, and continue the same planning loop. Stop only for a new unresolved frontier, a real capability or authority block, or a final sealed summary awaiting implementation approval.
57
65
  6. When no user-owned decision remains, create or update `design.md` and `implement.md` for complex tasks.
58
- 7. Run the requirement convergence gate, then the PRD convergence pass.
66
+ 7. Run the requirement convergence gate, then the PRD convergence pass. Finish with one Planning Seal closure pass.
59
67
  8. Present the final planning summary and stop. Do not run `task.py start` or edit product code in the same turn.
60
68
  9. Only a subsequent user message that explicitly approves the latest planning summary authorizes `task.py start` and implementation. If implementation reveals a material unresolved decision, record `decision-needed`, run `task.py replan <task> "<reason>"`, and return through this planning flow; do not open a popup during implementation.
61
69
 
@@ -95,6 +103,8 @@ Lightweight tasks may omit `design.md` and `implement.md`; they may not skip evi
95
103
 
96
104
  The final planning summary must show Goal, In Scope, Out of Scope, Acceptance Criteria, Key Decisions, relevant Risks or Deferred Items, and artifact status.
97
105
 
106
+ The Planning Seal closure pass reconciles `task.json`, `prd.md`, `design.md`, `implement.md`, research, decision records, and manifests; verifies targets, branches, dependencies, release, validation, rollback, dynamic-fact dispositions, and replan triggers; and fixes every material decision to an owner and outcome. No `TBD`, `TODO`, `decision-needed`, unowned option, unspecified branch, open implementation path, validation gap, or conditional acceptance may remain. Any material discovery invalidates the seal and returns to planning.
107
+
98
108
  ## Artifact Rules
99
109
 
100
110
  `prd.md` records requirements and acceptance:
@@ -49,10 +49,17 @@ sandbox claim.
49
49
  4. Write `report.json` only when you are ready to stop. Its status is one of
50
50
  `complete`, `blocked`, `incomplete`, or `error`; it is always
51
51
  **pending coordinator review**, never accepted/rejected/deferred.
52
- 5. A complete report includes independently checkable evidence. A non-complete
53
- report explains completed scope and the blocker or error. Include the exact
54
- identity, scope, and lens required by the artifact helper.
55
- 6. Finish with one short final assistant reply that states the status and report
52
+ 5. Use report schema version 2. Include one `scope_assessment` for each brief
53
+ scope item, structured findings with `id`, `conclusion`, and `evidence_ids`,
54
+ and typed `uncertainties` or `corrections` when present. A complete report
55
+ includes independently checkable evidence; a non-complete report explains
56
+ completed scope and the blocker or error. Include the exact identity, scope,
57
+ and lens required by the artifact helper.
58
+ 6. Append a checkpoint marker to `worklog.md` after each material unit using
59
+ the exact `trellis-checkpoint` JSON fields documented by `subnode-work`.
60
+ Missing or incomplete checkpoint coverage becomes a coordinator review
61
+ concern; it does not become acceptance.
62
+ 7. Finish with one short final assistant reply that states the status and report
56
63
  path. Do not place the report JSON in the reply and do not run
57
64
  `trellis channel send`: the supervisor routes this final reply into the
58
65
  durable Channel message and `done` events.
@@ -22,10 +22,12 @@ from common.task_utils import is_within_tasks_dir, resolve_task_dir
22
22
  from common.tasks import load_task
23
23
 
24
24
 
25
- SCHEMA_VERSION = 1
25
+ SCHEMA_VERSION = 2
26
26
  MAX_DRAFT_BYTES = 64 * 1024
27
27
  MAX_REPORT_BYTES = 128 * 1024
28
+ MAX_WORKLOG_BYTES = 128 * 1024
28
29
  ID_RE = re.compile(r"^[A-Za-z0-9][A-Za-z0-9._-]{0,63}$")
30
+ CHECKPOINT_RE = re.compile(r"^<!-- trellis-checkpoint: (?P<payload>\{.*\}) -->$", re.MULTILINE)
29
31
  SECRET_PATTERNS = (
30
32
  re.compile(r"-----BEGIN [A-Z ]*PRIVATE KEY-----"),
31
33
  re.compile(r"\bsk-[A-Za-z0-9]{20,}\b"),
@@ -305,7 +307,7 @@ def _validate_brief_file(
305
307
  return brief, raw, brief_path, node_dir, task_dir
306
308
 
307
309
 
308
- def _validate_evidence(value: Any) -> set[tuple[str, str]]:
310
+ def _validate_evidence(value: Any) -> set[str]:
309
311
  if not isinstance(value, list):
310
312
  _fail("report.evidence must be a list")
311
313
  identities: set[tuple[str, str]] = set()
@@ -319,13 +321,100 @@ def _validate_evidence(value: Any) -> set[tuple[str, str]]:
319
321
  if identity in identities:
320
322
  _fail("report.evidence contains a duplicate evidence identity")
321
323
  identities.add(identity)
322
- return identities
324
+ return {evidence_id for evidence_id, _locator in identities}
325
+
326
+
327
+ def _validate_id_list(value: Any, field: str) -> list[str]:
328
+ if not isinstance(value, list):
329
+ _fail(f"{field} must be a list")
330
+ result = []
331
+ for index, item in enumerate(value):
332
+ result.append(_require_text(item, f"{field}[{index}]", max_len=128))
333
+ if len(result) != len(set(result)):
334
+ _fail(f"{field} must not contain duplicates")
335
+ return result
336
+
337
+
338
+ def _validate_typed_notes(value: Any, field: str) -> None:
339
+ if not isinstance(value, list):
340
+ _fail(f"report.{field} must be a list")
341
+ for index, item in enumerate(value):
342
+ if not isinstance(item, dict):
343
+ _fail(f"report.{field}[{index}] must be an object")
344
+ _require_id(item.get("id"), f"report.{field}[{index}].id")
345
+ _require_text(item.get("type"), f"report.{field}[{index}].type", max_len=64)
346
+ _require_text(item.get("detail"), f"report.{field}[{index}].detail")
347
+ _validate_id_list(item.get("evidence_ids", []), f"report.{field}[{index}].evidence_ids")
348
+
349
+
350
+ def _validate_checkpoint(node_dir: Path, evidence_ids: set[str], scope: list[str]) -> list[str]:
351
+ worklog_path = node_dir / "worklog.md"
352
+ if worklog_path.is_symlink():
353
+ _fail(f"refusing symlinked worklog: {worklog_path}")
354
+ try:
355
+ raw = worklog_path.read_bytes()
356
+ except FileNotFoundError:
357
+ _fail(f"worklog does not exist: {worklog_path}")
358
+ except OSError as exc:
359
+ _fail(f"could not read worklog {worklog_path}: {exc}")
360
+ if len(raw) > MAX_WORKLOG_BYTES:
361
+ _fail(f"worklog exceeds the {MAX_WORKLOG_BYTES} byte limit: {worklog_path}")
362
+ _reject_obvious_secrets(raw, "worklog")
363
+ try:
364
+ text = raw.decode("utf-8")
365
+ except UnicodeDecodeError:
366
+ _fail(f"worklog is not UTF-8: {worklog_path}")
367
+ matches = list(CHECKPOINT_RE.finditer(text))
368
+ if not matches:
369
+ return ["missing_worklog_checkpoint"]
370
+ concerns: list[str] = []
371
+ checkpoint_ids: set[str] = set()
372
+ for match in matches:
373
+ try:
374
+ checkpoint = json.loads(match.group("payload"))
375
+ except json.JSONDecodeError:
376
+ concerns.append("malformed_worklog_checkpoint")
377
+ continue
378
+ if not isinstance(checkpoint, dict):
379
+ concerns.append("malformed_worklog_checkpoint")
380
+ continue
381
+ try:
382
+ checkpoint_id = _require_id(checkpoint.get("id"), "worklog checkpoint.id")
383
+ if checkpoint_id in checkpoint_ids:
384
+ concerns.append("duplicate_worklog_checkpoint")
385
+ checkpoint_ids.add(checkpoint_id)
386
+ covered_scope = _require_text_list(
387
+ checkpoint.get("covered_scope"),
388
+ "worklog checkpoint.covered_scope",
389
+ allow_empty=True,
390
+ )
391
+ for item in covered_scope:
392
+ if item not in scope:
393
+ concerns.append("checkpoint_scope_outside_brief")
394
+ checkpoint_evidence = _validate_id_list(
395
+ checkpoint.get("evidence_ids"), "worklog checkpoint.evidence_ids"
396
+ )
397
+ if any(item not in evidence_ids for item in checkpoint_evidence):
398
+ concerns.append("checkpoint_evidence_unresolved")
399
+ _require_text(
400
+ checkpoint.get("conclusion_or_blocker"),
401
+ "worklog checkpoint.conclusion_or_blocker",
402
+ )
403
+ _require_text_list(
404
+ checkpoint.get("unknowns"),
405
+ "worklog checkpoint.unknowns",
406
+ allow_empty=True,
407
+ )
408
+ _require_text(checkpoint.get("safe_resume_point"), "worklog checkpoint.safe_resume_point")
409
+ except ArtifactError:
410
+ concerns.append("malformed_worklog_checkpoint")
411
+ return sorted(set(concerns))
323
412
 
324
413
 
325
414
  def _validate_report_data(
326
415
  report: dict[str, Any],
327
416
  brief: dict[str, Any],
328
- ) -> set[tuple[str, str]]:
417
+ ) -> tuple[set[str], list[str]]:
329
418
  if report.get("schema_version") != SCHEMA_VERSION:
330
419
  _fail(f"report.schema_version must be {SCHEMA_VERSION}")
331
420
  for field in ("task_id", "work_id", "subnode_id", "role_id"):
@@ -339,11 +428,52 @@ def _validate_report_data(
339
428
  if status not in {"complete", "blocked", "incomplete", "error"}:
340
429
  _fail("report.status must be complete, blocked, incomplete, or error")
341
430
  evidence = _validate_evidence(report.get("evidence"))
342
- for field in ("findings", "uncertainties", "corrections"):
343
- if not isinstance(report.get(field), list):
344
- _fail(f"report.{field} must be a list")
431
+ assessment = report.get("scope_assessment")
432
+ if not isinstance(assessment, list) or len(assessment) != len(brief["scope"]):
433
+ _fail("report.scope_assessment must contain one item for every brief scope item")
434
+ concerns: list[str] = []
435
+ for index, item in enumerate(assessment):
436
+ if not isinstance(item, dict):
437
+ _fail(f"report.scope_assessment[{index}] must be an object")
438
+ if item.get("scope") != brief["scope"][index]:
439
+ _fail(f"report.scope_assessment[{index}].scope does not match brief.scope")
440
+ if item.get("status") not in {"covered", "inconclusive", "not-started"}:
441
+ _fail(f"report.scope_assessment[{index}].status is invalid")
442
+ _require_text(item.get("conclusion"), f"report.scope_assessment[{index}].conclusion")
443
+ assessment_evidence = _validate_id_list(
444
+ item.get("evidence_ids"), f"report.scope_assessment[{index}].evidence_ids"
445
+ )
446
+ if any(value not in evidence for value in assessment_evidence):
447
+ concerns.append("scope_assessment_evidence_unresolved")
448
+ if item["status"] != "covered":
449
+ concerns.append("incomplete_scope_coverage")
450
+ if item["status"] == "covered" and not assessment_evidence:
451
+ concerns.append("covered_scope_without_evidence")
452
+ findings = report.get("findings")
453
+ if not isinstance(findings, list):
454
+ _fail("report.findings must be a list")
455
+ finding_ids: set[str] = set()
456
+ for index, item in enumerate(findings):
457
+ if not isinstance(item, dict):
458
+ _fail(f"report.findings[{index}] must be an object")
459
+ finding_id = _require_id(item.get("id"), f"report.findings[{index}].id")
460
+ if finding_id in finding_ids:
461
+ _fail("report.findings contains duplicate ids")
462
+ finding_ids.add(finding_id)
463
+ _require_text(item.get("conclusion"), f"report.findings[{index}].conclusion")
464
+ finding_evidence = _validate_id_list(
465
+ item.get("evidence_ids"), f"report.findings[{index}].evidence_ids"
466
+ )
467
+ if not finding_evidence:
468
+ concerns.append("finding_without_evidence")
469
+ if any(value not in evidence for value in finding_evidence):
470
+ concerns.append("finding_evidence_unresolved")
471
+ _validate_typed_notes(report.get("uncertainties"), "uncertainties")
472
+ _validate_typed_notes(report.get("corrections"), "corrections")
345
473
  if status == "complete" and not evidence:
346
474
  _fail("a complete report requires at least one evidence item")
475
+ if status == "complete" and not findings:
476
+ _fail("a complete report requires at least one finding")
347
477
  if status != "complete":
348
478
  _require_text_list(
349
479
  report.get("completed_scope"),
@@ -351,13 +481,13 @@ def _validate_report_data(
351
481
  allow_empty=True,
352
482
  )
353
483
  _require_text(report.get("blocker"), "report.blocker")
354
- return evidence
484
+ return evidence, sorted(set(concerns))
355
485
 
356
486
 
357
487
  def _validate_report_file(
358
488
  value: str,
359
489
  repo_root: Path,
360
- ) -> tuple[dict[str, Any], set[tuple[str, str]], dict[str, Any], Path]:
490
+ ) -> tuple[dict[str, Any], set[str], dict[str, Any], Path, list[str]]:
361
491
  report_path, node_dir, _task_dir = _resolve_artifact_file(value, "report.json", repo_root)
362
492
  brief_path = node_dir / "brief.json"
363
493
  brief, _brief_raw, _brief_path, _brief_node, _brief_task = _validate_brief_file(
@@ -368,8 +498,9 @@ def _validate_report_file(
368
498
  _fail("brief.report_path does not point to the report being validated")
369
499
  report, raw = _read_json_object(report_path, MAX_REPORT_BYTES, "report")
370
500
  _reject_obvious_secrets(raw, "report")
371
- evidence = _validate_report_data(report, brief)
372
- return report, evidence, brief, node_dir
501
+ evidence, concerns = _validate_report_data(report, brief)
502
+ concerns.extend(_validate_checkpoint(node_dir, evidence, brief["scope"]))
503
+ return report, evidence, brief, node_dir, sorted(set(concerns))
373
504
 
374
505
 
375
506
  def _validate_disposition_data(
@@ -466,13 +597,20 @@ def _init(args: argparse.Namespace) -> None:
466
597
 
467
598
  def _validate(args: argparse.Namespace) -> None:
468
599
  repo_root = get_repo_root()
469
- report, _evidence, _brief, _node_dir_value = _validate_report_file(args.report, repo_root)
600
+ report, _evidence, _brief, _node_dir_value, concerns = _validate_report_file(args.report, repo_root)
601
+ if concerns:
602
+ print(json.dumps({
603
+ "status": "review_concern",
604
+ "subnode_id": report["subnode_id"],
605
+ "concerns": concerns,
606
+ }))
607
+ return
470
608
  print(f"Validated pending-review report: {report['subnode_id']}")
471
609
 
472
610
 
473
611
  def _disposition(args: argparse.Namespace) -> None:
474
612
  repo_root = get_repo_root()
475
- report, _evidence, brief, node_dir = _validate_report_file(args.report, repo_root)
613
+ report, _evidence, brief, node_dir, _concerns = _validate_report_file(args.report, repo_root)
476
614
  disposition_path = node_dir / "disposition.json"
477
615
  if disposition_path.exists() or disposition_path.is_symlink():
478
616
  _fail(f"disposition already exists and cannot be replaced: {disposition_path}")
@@ -514,10 +652,10 @@ def _validate_counter(args: argparse.Namespace) -> None:
514
652
  primary_dir = repo_root / primary_dir
515
653
  if not counter_dir.is_absolute():
516
654
  counter_dir = repo_root / counter_dir
517
- primary_report, primary_evidence, primary_brief, primary_node = _validate_report_file(
655
+ primary_report, primary_evidence, primary_brief, primary_node, _primary_concerns = _validate_report_file(
518
656
  str(primary_dir / "report.json"), repo_root
519
657
  )
520
- counter_report, counter_evidence, counter_brief, counter_node = _validate_report_file(
658
+ counter_report, counter_evidence, counter_brief, counter_node, _counter_concerns = _validate_report_file(
521
659
  str(counter_dir / "report.json"), repo_root
522
660
  )
523
661
  if primary_node == counter_node:
@@ -167,6 +167,8 @@ Phase 3: Finish → verify, update spec, commit, and wrap up
167
167
 
168
168
  An analysis-only task is eligible only when `task.json.meta.delivery_mode = "analysis_only"` exactly and its `prd.md` names the evidence deliverable plus a no-change boundary for product source, runtime configuration, deployment, credentials, and external systems. Task creation consent authorizes that bounded evidence work, not protected-target changes.
169
169
 
170
+ The exception is invalid when the work still needs a material user decision, design or implementation plan, cross-owner coordination, security or deployment change, release or credential action, or a protected downstream task. "Research" and a deferred source edit do not override this classification; use normal complex planning when any condition applies.
171
+
170
172
  Keep an eligible analysis-only task in `planning`: write its research, audit, or design evidence; verify its acceptance criteria and boundary; commit task artifacts; then archive directly. Do not run `task.py start`, configure implementation context, or wait for a second implementation approval. If the evidence recommends a protected-target change, record the recommendation and create a separate change-bearing task before doing it.
171
173
 
172
174
  ### Planning Artifacts
@@ -177,6 +179,8 @@ Keep an eligible analysis-only task in `planning`: write its research, audit, or
177
179
  - `implement.jsonl` / `check.jsonl` — spec and research manifests for sub-agent context. They do not replace `implement.md`.
178
180
  - Lightweight tasks may be PRD-only. Complex tasks must have `prd.md`, `design.md`, and `implement.md` before `task.py start`.
179
181
 
182
+ For read-heavy research, audit, review, or investigation, split work that cannot persist one useful conclusion in the current bounded session into evidence units. Each unit records one question or scope, minimal evidence, destination, stop condition, conclusion or blocker, unknowns, and recovery point before the next unit begins. Size for one normal context window, without an exact token or time promise; routine navigation needs no record.
183
+
180
184
  ### Parent / Child Task Trees
181
185
 
182
186
  Use a parent task when one user request contains several independently verifiable deliverables. The parent task owns the source requirement set, the task map, cross-child acceptance criteria, and final integration review; it normally should not be the implementation target unless it also has direct work.
@@ -230,7 +234,7 @@ Preserve existing task fields and artifacts. If the correct status cannot be det
230
234
  [workflow-state:planning]
231
235
  Load `trellis-brainstorm`; stay in planning.
232
236
  If `task.json.meta.delivery_mode = "analysis_only"` exactly, complete the declared evidence work now. Do not wait for a start review or run `task.py start`; when the PRD boundary and acceptance evidence pass, commit task artifacts and archive directly. A protected-target recommendation requires a separate change-bearing task.
233
- Lightweight: `prd.md` can be enough. Complex: finish `prd.md`, `design.md`, and `implement.md`; ask for review before `task.py start`. If `decision-needed` items or an unsealed decision graph remain, load `pennix-decision-gates`, batch only independent frontier questions, and stay in planning.
237
+ Lightweight: `prd.md` can be enough. Complex: finish `prd.md`, `design.md`, and `implement.md`; run the Planning Seal closure pass before asking for review. If `decision-needed` items or an unsealed decision graph remain, load `pennix-decision-gates`, batch only independent frontier questions, and stay in planning. Answers returned by the current continuation must be persisted and fed back into the same planning loop.
234
238
  Multi-deliverable scope: consider a parent task plus independently verifiable child tasks; dependencies must be written in child artifacts, not implied by tree position.
235
239
  Sub-agent mode: curate `implement.jsonl` and `check.jsonl` as spec/research manifests before start.
236
240
  [/workflow-state:planning]
@@ -244,7 +248,7 @@ Sub-agent mode: curate `implement.jsonl` and `check.jsonl` as spec/research mani
244
248
  [workflow-state:planning-inline]
245
249
  Load `trellis-brainstorm`; stay in planning.
246
250
  If `task.json.meta.delivery_mode = "analysis_only"` exactly, complete the declared evidence work now. Do not wait for a start review or run `task.py start`; when the PRD boundary and acceptance evidence pass, commit task artifacts and archive directly. A protected-target recommendation requires a separate change-bearing task.
247
- Lightweight: `prd.md` can be enough. Complex: finish `prd.md`, `design.md`, and `implement.md`; ask for review before `task.py start`. If `decision-needed` items or an unsealed decision graph remain, load `pennix-decision-gates`, batch only independent frontier questions, and stay in planning.
251
+ Lightweight: `prd.md` can be enough. Complex: finish `prd.md`, `design.md`, and `implement.md`; run the Planning Seal closure pass before asking for review. If `decision-needed` items or an unsealed decision graph remain, load `pennix-decision-gates`, batch only independent frontier questions, and stay in planning. Answers returned by the current continuation must be persisted and fed back into the same planning loop.
248
252
  Multi-deliverable scope: consider a parent task plus independently verifiable child tasks; dependencies must be written in child artifacts, not implied by tree position.
249
253
  Inline mode: skip jsonl curation; Phase 2 reads artifacts/specs via `trellis-before-dev`.
250
254
  [/workflow-state:planning-inline]
@@ -389,6 +393,8 @@ The brainstorm skill will guide you to:
389
393
  - Split large scopes into a parent task plus child tasks when the deliverables can be verified independently
390
394
  - Keep `prd.md` focused on requirements and acceptance criteria
391
395
  - For complex tasks, produce `design.md` and `implement.md` before implementation starts
396
+ - For read-heavy work, split long investigation into evidence units and persist each unit's conclusion or recovery point before continuing
397
+ - Before review or `task.py start`, run the Planning Seal closure pass across all task artifacts and lock targets, branches, dependencies, release, validation, rollback, dynamic-fact handling, and every material decision
392
398
 
393
399
  When considering a parent/child split:
394
400
  - Use a parent task when one request contains several independently verifiable deliverables.
@@ -490,7 +496,7 @@ Skip this step. Context is loaded directly by the `trellis-before-dev` skill in
490
496
 
491
497
  This step applies only to change-bearing tasks. An eligible analysis-only task stays in `planning` and, after its evidence work is complete, continues directly to Phase 3.3 without running `task.py start`.
492
498
 
493
- After artifact review, flip the task status to `in_progress`:
499
+ After the Planning Seal closure pass and artifact review, flip the task status to `in_progress`:
494
500
 
495
501
  ```bash
496
502
  python3 ./.trellis/scripts/task.py start <task-dir>
@@ -513,6 +519,7 @@ If `task.py start` errors with a session-identity message (no context key from h
513
519
  | `research/` has artifacts (complex tasks) | recommended |
514
520
  | `design.md` exists (complex tasks) | ✅ |
515
521
  | `implement.md` exists (complex tasks) | ✅ |
522
+ | Planning Seal closure pass recorded; static decisions and implementation paths are locked | ✅ |
516
523
 
517
524
  [Claude Code, Cursor, OpenCode, codex-sub-agent, Kiro, Gemini, Qoder, CodeBuddy, Copilot, Droid, Pi, Oh My Pi, ZCode, Snow, Reasonix, Trae, Grok, Kimi Code, DeepSeek Harness]
518
525
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@pennixrv/trellis",
3
- "version": "0.7.0-beta.8",
3
+ "version": "0.7.0-beta.9",
4
4
  "description": "AI capabilities grow like ivy — Trellis provides the structure to guide them along a disciplined path",
5
5
  "type": "module",
6
6
  "main": "./dist/index.js",
@@ -34,7 +34,7 @@
34
34
  "inquirer": "^9.3.7",
35
35
  "undici": "^6.21.0",
36
36
  "zod": "^4.4.2",
37
- "@pennixrv/trellis-core": "0.7.0-beta.8"
37
+ "@pennixrv/trellis-core": "0.7.0-beta.9"
38
38
  },
39
39
  "devDependencies": {
40
40
  "@eslint/js": "^9.18.0",