@pennixrv/trellis 0.7.0-beta.7 → 0.7.0-beta.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,9 @@
1
+ {
2
+ "version": "0.7.0-beta.8",
3
+ "description": "Pennix v0.7 beta decision-chain replanning",
4
+ "breaking": false,
5
+ "recommendMigrate": false,
6
+ "changelog": "**Features:**\n- Add an auditable same-task replan transition for material implementation-time decisions.\n- Preserve every planning interval in `trellis mem --phase brainstorm`.",
7
+ "migrations": [],
8
+ "notes": "Run `npm install -g @pennixrv/trellis@beta` then `trellis update`. No `--migrate` required."
9
+ }
@@ -0,0 +1,9 @@
1
+ {
2
+ "version": "0.7.0-beta.9",
3
+ "description": "Pennix v0.7 beta planning and evidence governance",
4
+ "breaking": false,
5
+ "recommendMigrate": false,
6
+ "changelog": "**Features:**\n- Require complex analysis to use normal planning and a sealed decision chain.\n- Add evidence units and schema-v2 subnode reports with coordinator-only acceptance.",
7
+ "migrations": [],
8
+ "notes": "Run npm install -g @pennixrv/trellis@beta then trellis update. No --migrate required."
9
+ }
@@ -50,12 +50,15 @@ and its status; the Codex supervisor projects that reply into the durable
50
50
  Channel message and `done` events. The durable JSON file carries the
51
51
  reviewable result.
52
52
 
53
- The subnode copies identity, scope, and lens from `brief.json`. The minimal
53
+ The subnode copies identity, scope, and lens from `brief.json`. Schema version 2
54
+ requires one `scope_assessment` entry for every brief scope item, structured
55
+ findings with evidence references, and typed uncertainties or corrections when
56
+ present. Schema version 1 is retired and the validator rejects it. The minimal
54
57
  complete report is:
55
58
 
56
59
  ```json
57
60
  {
58
- "schema_version": 1,
61
+ "schema_version": 2,
59
62
  "task_id": "task-id-from-brief",
60
63
  "work_id": "work-id-from-brief",
61
64
  "subnode_id": "subnode-id-from-brief",
@@ -63,6 +66,14 @@ complete report is:
63
66
  "status": "complete",
64
67
  "scope": ["exact scope copied from brief"],
65
68
  "lens": "exact lens copied from brief",
69
+ "scope_assessment": [
70
+ {
71
+ "scope": "exact scope item from brief",
72
+ "status": "covered",
73
+ "conclusion": "What this scope establishes.",
74
+ "evidence_ids": ["stable-evidence-id"]
75
+ }
76
+ ],
66
77
  "evidence": [
67
78
  {
68
79
  "id": "stable-evidence-id",
@@ -70,12 +81,30 @@ complete report is:
70
81
  "summary": "What this independently reviewable evidence establishes."
71
82
  }
72
83
  ],
73
- "findings": ["Bounded conclusion."],
84
+ "findings": [
85
+ {
86
+ "id": "finding-id",
87
+ "conclusion": "Bounded conclusion.",
88
+ "evidence_ids": ["stable-evidence-id"]
89
+ }
90
+ ],
74
91
  "uncertainties": [],
75
92
  "corrections": []
76
93
  }
77
94
  ```
78
95
 
96
+ Append a checkpoint marker to `worklog.md` when a material unit is complete or
97
+ blocked. It is a bounded recovery projection, not coordinator acceptance:
98
+
99
+ ```text
100
+ <!-- trellis-checkpoint: {"id":"checkpoint-1","covered_scope":["exact scope item"],"evidence_ids":["stable-evidence-id"],"conclusion_or_blocker":"Current conclusion or blocker.","unknowns":[],"safe_resume_point":"Next safe action."} -->
101
+ ```
102
+
103
+ The validator returns `review_concern` for missing or incomplete checkpoint
104
+ coverage and for inconclusive scope assessments so the coordinator can inspect
105
+ them. Identity, schema, path, and malformed-structure failures remain hard
106
+ errors. A concern is never automatic acceptance or rejection.
107
+
79
108
  For `blocked`, `incomplete`, or `error`, include the same base fields plus a
80
109
  `completed_scope` list (empty when no assigned scope started) and a non-empty
81
110
  `blocker` string. Never use
@@ -17,6 +17,7 @@ Task lifecycle includes creation, start, context configuration, finish, archive,
17
17
  | --- | --- |
18
18
  | Automatically sync an external system after task creation | `hooks.after_create` in `.trellis/config.yaml`. |
19
19
  | Automatically update status after task start | `hooks.after_start` in `.trellis/config.yaml`. |
20
+ | Return an in-progress task to planning | `hooks.after_replan` in `.trellis/config.yaml`. |
20
21
  | Run a script after task finish | `hooks.after_finish` in `.trellis/config.yaml`. |
21
22
  | Clean external resources after archive | `hooks.after_archive` in `.trellis/config.yaml`. |
22
23
  | Change default task fields | `.trellis/scripts/common/task_store.py`. |
@@ -68,7 +68,7 @@ trellis mem list --cwd <project-path>
68
68
  trellis mem projects # → list active project cwds, then narrow
69
69
  ```
70
70
 
71
- Phase slicing (`--phase brainstorm|implement|all`) cuts the session at `task.py create` and `task.py start` boundaries. For a finish-work review of the current task, `--phase brainstorm` recovers the planning discussion and `--phase implement` recovers the execution loop. Default is `all`.
71
+ Phase slicing (`--phase brainstorm|implement|all`) cuts the session at `task.py create`/`replan` and `task.py start` boundaries. For a finish-work review of the current task, `--phase brainstorm` recovers the planning discussion and `--phase implement` recovers the execution loop. Default is `all`.
72
72
 
73
73
  ## Triggering patterns
74
74
 
@@ -23,7 +23,7 @@ Full flag reference for the five subcommands. Pin this as the authoritative sour
23
23
  | `--cwd <path>` | list / search | Force a specific project cwd instead of inferring from where you are. |
24
24
  | `--limit N` | list / search | Cap output rows. Default `50`. |
25
25
  | `--grep KW` | extract / context | Filter turns by keyword. Multi-token AND when whitespace-separated. |
26
- | `--phase brainstorm\|implement\|all` | extract | Slice session by Trellis task boundaries. `brainstorm` = `[task.py create, task.py start)`. `implement` = turns outside brainstorm windows. Default `all`. |
26
+ | `--phase brainstorm\|implement\|all` | extract | Slice session by Trellis task boundaries. `brainstorm` = `[task.py create/replan, task.py start)`. `implement` = turns outside brainstorm windows. Default `all`. |
27
27
  | `--turns N` | context | Number of hit turns to return. Default `3`. |
28
28
  | `--around N` | context | Surrounding turns to include per hit. Default `1`. |
29
29
  | `--max-chars N` | context | Total character budget. Default `6000` (~1500 tokens). |
@@ -24,11 +24,13 @@ Shows the Phase Index (Plan / Execute / Finish) with routing + skill mapping.
24
24
 
25
25
  `get_context.py` shows the active task's `status` field. Route by `status` + artifact presence. This command replaces the user needing to remember the Trellis flow; it does not itself approve implementation.
26
26
 
27
- - `status=planning` + `task.json.meta.delivery_mode = "analysis_only"` → complete the PRD's bounded evidence work, verify its acceptance criteria and no-change boundary, then commit task artifacts and archive directly. Do not run `task.py start`; a protected-target change requires a separate change-bearing task.
27
+ - `status=planning` + `task.json.meta.delivery_mode = "analysis_only"` → first confirm the task still satisfies the bounded evidence-only eligibility rule; then complete the PRD's evidence work, verify its acceptance criteria and no-change boundary, and archive directly. Do not run `task.py start`; a protected-target change requires a separate change-bearing task.
28
28
  - `status=planning` + no `prd.md` → **1.1** (load `trellis-brainstorm`)
29
+ - `status=planning` + a recorded `decision-needed` or unsealed decision chain → return to the planning frontier and load `pennix-decision-gates` when independent material questions can be batched.
30
+ - `status=in_progress` + a material unresolved decision → record the reason and run `task.py replan <task> "<reason>"`; do not ask a native question during implementation.
29
31
  - `status=planning` + `prd.md` only → decide whether the task is lightweight or complex. Lightweight can move to **1.4** review; complex returns to **1.1** to add `design.md` + `implement.md`.
30
32
  - `status=planning` + complex artifacts complete + sub-agent jsonl not curated (empty, or only a legacy `_example` placeholder row) → **1.3**
31
- - `status=planning` + required artifacts complete + required jsonl curated or inline mode → **1.4** (ask for start review; only run `task.py start` after user confirms)
33
+ - `status=planning` + required artifacts complete + required jsonl curated or inline mode → run the Planning Seal closure pass, then **1.4** (ask for start review; only run `task.py start` after user confirms)
32
34
  - `status=in_progress` + implementation not started → **2.1**
33
35
  - `status=in_progress` + implementation done, not yet checked → **2.2**
34
36
  - `status=in_progress` + check passed → **3.3** (spec update) → **3.4** (commit)
@@ -6,12 +6,14 @@ A request to build, implement, fix, refactor, or "go ahead" is not approval to l
6
6
 
7
7
  For every non-trivial task, the user must respond at least once after the initial request before implementation begins. If no clarification is needed, that response must approve the final planning summary described below.
8
8
 
9
- While any user-owned product, scope, UX, compatibility, risk, or acceptance decision remains unresolved, end the turn with exactly one highest-value question. Do not edit product code, dispatch implementation, or run `task.py start`.
9
+ While any user-owned product, scope, UX, compatibility, risk, or acceptance decision remains unresolved, keep the task in planning. First inventory evidence and decision dependencies. If at least two independent material decisions remain and `pennix-decision-gates` is available, delegate one bounded batch of up to three frontier questions; otherwise ask the single highest-value question. Do not edit product code, dispatch implementation, or run `task.py start` until the decision chain is sealed.
10
10
 
11
11
  ## Analysis-Only Exception
12
12
 
13
13
  When `task.json.meta.delivery_mode = "analysis_only"` exactly and the PRD names a bounded evidence deliverable plus a no-change boundary for product source, runtime configuration, deployment, credentials, and external systems, task-creation consent authorizes that evidence work. Do not require a second planning approval or run `task.py start`: perform the declared research, audit, or design work while status remains `planning`, record the evidence, verify acceptance criteria and the boundary, commit task artifacts, and archive directly. If the evidence recommends a protected-target change, record it and create a separate change-bearing task before doing it.
14
14
 
15
+ This exception is eligible only for a bounded evidence deliverable with no material user decision, design or implementation plan, cross-owner coordination, security or deployment change, release or credential action, or protected downstream task. Calling work "research", deferring source edits, or working in an audit/root repository does not make it analysis-only. If any of those conditions apply, use the normal complex planning and implementation-approval path.
16
+
15
17
  All other tasks follow the planning and implementation approval gates below.
16
18
 
17
19
  ## Non-Negotiable Evidence Rule
@@ -24,6 +26,10 @@ Do not ask the user to confirm facts that the repository can answer. Ask only fo
24
26
 
25
27
  Repository evidence establishes current behavior and technical constraints. The user's intended behavior, feature scope boundaries, and UX preferences are never answerable by repository evidence alone, even when an existing pattern exists; existing patterns are options and recommendation evidence, not decisions.
26
28
 
29
+ ## Evidence Units For Read-Heavy Work
30
+
31
+ When research, audit, review, or investigation is too large to leave one independently useful conclusion in the current bounded session, split it into evidence units. Each unit must have one question or scope, a minimal evidence range, a destination artifact, and a stop condition; write its facts, conclusion or blocker, unknowns, and recovery point before starting another unit. Size units so one normal context window can finish and persist one useful result; do not promise an exact token or time limit. Routine navigation and transient tool output do not need an artifact. Create a child task only when the unit has an independent owner, lifecycle, and acceptance contract.
32
+
27
33
  ---
28
34
 
29
35
  Use this skill during Phase 1 planning to turn the user's request into clear requirements and planning artifacts.
@@ -54,18 +60,18 @@ Use a concise title from the user's request. Both the title and `--description`
54
60
  - product intent still needed from the user
55
61
  - scope or risk decisions still needed from the user
56
62
  - likely out-of-scope items
57
- 4. If a user-owned decision remains, ask the single highest-value question, include your recommendation and trade-off, then stop. Do not perform implementation work in the same turn.
58
- 5. After each user answer, update `prd.md`, recompute the decision inventory, and repeat from step 2.
63
+ 4. If user-owned decisions remain, calculate the independent frontier. Use `pennix-decision-gates` for a bounded batch when two or more independent material decisions are ready; otherwise ask the single highest-value question. Include recommendation and trade-off. Yield only while the answer is unavailable.
64
+ 5. When the host returns the current continuation's answer, immediately persist it in `prd.md` or the decision artifact, recheck evidence and conflicts, recalculate the frontier, and continue the same planning loop. Do not create a second Trellis lifecycle for the same decision chain. Stop only for a new unresolved frontier, a real capability or authority block, or a final sealed summary awaiting implementation approval.
59
65
  6. When no user-owned decision remains, create or update `design.md` and `implement.md` for complex tasks.
60
- 7. Run the requirement convergence gate, then the PRD convergence pass.
66
+ 7. Run the requirement convergence gate, then the PRD convergence pass. Finish with one Planning Seal closure pass.
61
67
  8. Present the final planning summary and stop. Do not run `task.py start` or edit product code in the same turn.
62
- 9. Only a subsequent user message that explicitly approves the latest planning summary authorizes `task.py start` and implementation. If the artifacts change materially after approval, repeat the final review.
68
+ 9. Only a subsequent user message that explicitly approves the latest planning summary authorizes `task.py start` and implementation. If implementation reveals a material unresolved decision, record `decision-needed`, run `task.py replan <task> "<reason>"`, and return through this planning flow; do not open a popup during implementation.
63
69
 
64
70
  Do not invent a project-specific product/spec hierarchy. If the repository already has product, domain, or spec docs, use them. If it does not, proceed with the evidence that exists.
65
71
 
66
72
  ## Question Rules
67
73
 
68
- Ask only one question per message.
74
+ Ask one bounded batch per message: include up to three independent material frontier questions. Ask exactly one question only when it is the sole remaining material decision or later decisions depend on its answer.
69
75
 
70
76
  Each question must include:
71
77
 
@@ -140,6 +146,8 @@ Lightweight tasks may omit `design.md` and `implement.md`; they may not skip evi
140
146
 
141
147
  The final planning summary must show Goal, In Scope, Out of Scope, Acceptance Criteria, Key Decisions, relevant Risks or Deferred Items, and artifact status.
142
148
 
149
+ The Planning Seal closure pass must reconcile `task.json`, `prd.md`, `design.md`, `implement.md`, research, decision records, and manifests; verify the actual modification targets and branches, ordered dependencies and release steps, validation and rollback, dynamic-fact dispositions and replan triggers, and that every material decision has an owner and a fixed outcome. Remove static ambiguity before implementation: no `TBD`, `TODO`, `decision-needed`, unowned option, unspecified branch, open implementation path, validation gap, or conditional acceptance may remain. A material discovery invalidates the seal and returns to planning; implementation may consume only a sealed plan.
150
+
143
151
  ## Artifact Rules
144
152
 
145
153
  `prd.md` records requirements and acceptance:
@@ -10,7 +10,11 @@ A request to build, implement, fix, refactor, or "go ahead" is not approval to l
10
10
 
11
11
  For every non-trivial task, the user must respond at least once after the initial request before implementation begins. If no clarification is needed, that response must approve the final planning summary described below.
12
12
 
13
- While any user-owned product, scope, UX, compatibility, risk, or acceptance decision remains unresolved, end the turn with exactly one highest-value question. Do not edit product code, dispatch implementation, or run `task.py start`.
13
+ While any user-owned product, scope, UX, compatibility, risk, or acceptance decision remains unresolved, keep the task in planning. First inventory evidence and decision dependencies. If at least two independent material decisions remain and `pennix-decision-gates` is available, delegate one bounded batch of up to three frontier questions; otherwise ask the single highest-value question. Do not edit product code, dispatch implementation, or run `task.py start` until the decision chain is sealed.
14
+
15
+ ## Analysis-Only Exception
16
+
17
+ When `task.json.meta.delivery_mode = "analysis_only"` exactly and the PRD names a bounded evidence deliverable plus a no-change boundary for product source, runtime configuration, deployment, credentials, and external systems, task-creation consent authorizes that evidence work. Keep status `planning`, record and verify the evidence, commit task artifacts, and archive directly; do not run `task.py start` or wait for a second implementation approval. This exception is eligible only when there is no material user decision, design or implementation plan, cross-owner coordination, security or deployment change, release or credential action, or protected downstream task. Otherwise use normal complex planning.
14
18
 
15
19
  ## Non-Negotiable Evidence Rule
16
20
 
@@ -22,6 +26,10 @@ Do not ask the user to confirm facts that the repository can answer. Ask only fo
22
26
 
23
27
  Repository evidence establishes current behavior and technical constraints. The user's intended behavior, feature scope boundaries, and UX preferences are never answerable by repository evidence alone, even when an existing pattern exists; existing patterns are options and recommendation evidence, not decisions.
24
28
 
29
+ ## Evidence Units For Read-Heavy Work
30
+
31
+ When research, audit, review, or investigation is too large for one independently useful conclusion in the current session, split it into evidence units. Each unit has one question or scope, a minimal evidence range, a destination artifact, and a stop condition; persist facts, conclusion or blocker, unknowns, and a recovery point before starting another unit. Size each unit for one normal context window without promising an exact token or time limit.
32
+
25
33
  ---
26
34
 
27
35
  Use this skill during Phase 1 planning to turn the user's request into clear requirements and planning artifacts.
@@ -52,18 +60,18 @@ Use a concise title from the user's request. Both the title and `--description`
52
60
  - product intent still needed from the user
53
61
  - scope or risk decisions still needed from the user
54
62
  - likely out-of-scope items
55
- 4. If a user-owned decision remains, ask the single highest-value question, include your recommendation and trade-off, then stop. Do not perform implementation work in the same turn.
56
- 5. After each user answer, update `prd.md`, recompute the decision inventory, and repeat from step 2.
63
+ 4. If user-owned decisions remain, calculate the independent frontier. Use `pennix-decision-gates` for a bounded batch when two or more independent material decisions are ready; otherwise ask the single highest-value question. Include recommendation and trade-off. Yield only while the answer is unavailable.
64
+ 5. When the host returns the current continuation's answer, persist it in `prd.md` or the decision artifact, recheck evidence and conflicts, recalculate the frontier, and continue the same planning loop. Stop only for a new unresolved frontier, a real capability or authority block, or a final sealed summary awaiting implementation approval.
57
65
  6. When no user-owned decision remains, create or update `design.md` and `implement.md` for complex tasks.
58
- 7. Run the requirement convergence gate, then the PRD convergence pass.
66
+ 7. Run the requirement convergence gate, then the PRD convergence pass. Finish with one Planning Seal closure pass.
59
67
  8. Present the final planning summary and stop. Do not run `task.py start` or edit product code in the same turn.
60
- 9. Only a subsequent user message that explicitly approves the latest planning summary authorizes `task.py start` and implementation. If the artifacts change materially after approval, repeat the final review.
68
+ 9. Only a subsequent user message that explicitly approves the latest planning summary authorizes `task.py start` and implementation. If implementation reveals a material unresolved decision, record `decision-needed`, run `task.py replan <task> "<reason>"`, and return through this planning flow; do not open a popup during implementation.
61
69
 
62
70
  Do not invent a project-specific product/spec hierarchy. If the repository already has product, domain, or spec docs, use them. If it does not, proceed with the evidence that exists.
63
71
 
64
72
  ## Question Rules
65
73
 
66
- Ask only one question per message.
74
+ Ask one bounded batch per message: include up to three independent material frontier questions. Ask exactly one question only when it is the sole remaining material decision or later decisions depend on its answer.
67
75
 
68
76
  Each question must include:
69
77
 
@@ -95,6 +103,8 @@ Lightweight tasks may omit `design.md` and `implement.md`; they may not skip evi
95
103
 
96
104
  The final planning summary must show Goal, In Scope, Out of Scope, Acceptance Criteria, Key Decisions, relevant Risks or Deferred Items, and artifact status.
97
105
 
106
+ The Planning Seal closure pass reconciles `task.json`, `prd.md`, `design.md`, `implement.md`, research, decision records, and manifests; verifies targets, branches, dependencies, release, validation, rollback, dynamic-fact dispositions, and replan triggers; and fixes every material decision to an owner and outcome. No `TBD`, `TODO`, `decision-needed`, unowned option, unspecified branch, open implementation path, validation gap, or conditional acceptance may remain. Any material discovery invalidates the seal and returns to planning.
107
+
98
108
  ## Artifact Rules
99
109
 
100
110
  `prd.md` records requirements and acceptance:
@@ -4,7 +4,7 @@
4
4
  No researched platform exports a session id into its shell tool's child
5
5
  process, but every hook-capable one puts that id on hook stdin. So the hook
6
6
  that fires just before a shell command writes a short-lived runtime ticket
7
- whenever the pending command calls `task.py start/current/finish`, and the
7
+ whenever the pending command calls `task.py start/replan/current/finish`, and the
8
8
  task script consumes it when it has no native session environment.
9
9
 
10
10
  Registered on whichever pre-shell event the host provides — Cursor's
@@ -35,7 +35,7 @@ if callable(_stdin_reconfigure):
35
35
  DIR_WORKFLOW = ".trellis"
36
36
  DIR_RUNTIME = ".runtime"
37
37
  DIR_SHELL_TICKETS = "shell-tickets"
38
- SESSION_SUBCOMMANDS = {"start", "current", "finish"}
38
+ SESSION_SUBCOMMANDS = {"start", "replan", "current", "finish"}
39
39
  TICKET_TTL_SECONDS = 30
40
40
  CONTEXT_IDENTITY_KEYS = (
41
41
  "session_id",
@@ -177,7 +177,7 @@ def _extract_task_subcommands(command: str) -> list[dict[str, str]]:
177
177
  if name not in SESSION_SUBCOMMANDS:
178
178
  continue
179
179
  item = {"name": name}
180
- if name == "start" and index + 2 < len(tokens):
180
+ if name in {"start", "replan"} and index + 2 < len(tokens):
181
181
  item["task_ref"] = tokens[index + 2]
182
182
  subcommands.append(item)
183
183
  return subcommands
@@ -49,10 +49,17 @@ sandbox claim.
49
49
  4. Write `report.json` only when you are ready to stop. Its status is one of
50
50
  `complete`, `blocked`, `incomplete`, or `error`; it is always
51
51
  **pending coordinator review**, never accepted/rejected/deferred.
52
- 5. A complete report includes independently checkable evidence. A non-complete
53
- report explains completed scope and the blocker or error. Include the exact
54
- identity, scope, and lens required by the artifact helper.
55
- 6. Finish with one short final assistant reply that states the status and report
52
+ 5. Use report schema version 2. Include one `scope_assessment` for each brief
53
+ scope item, structured findings with `id`, `conclusion`, and `evidence_ids`,
54
+ and typed `uncertainties` or `corrections` when present. A complete report
55
+ includes independently checkable evidence; a non-complete report explains
56
+ completed scope and the blocker or error. Include the exact identity, scope,
57
+ and lens required by the artifact helper.
58
+ 6. Append a checkpoint marker to `worklog.md` after each material unit using
59
+ the exact `trellis-checkpoint` JSON fields documented by `subnode-work`.
60
+ Missing or incomplete checkpoint coverage becomes a coordinator review
61
+ concern; it does not become acceptance.
62
+ 7. Finish with one short final assistant reply that states the status and report
56
63
  path. Do not place the report JSON in the reply and do not run
57
64
  `trellis channel send`: the supervisor routes this final reply into the
58
65
  durable Channel message and `done` events.
@@ -45,6 +45,8 @@ max_journal_lines: 2000
45
45
  # - "echo 'Task created'"
46
46
  # after_start:
47
47
  # - "echo 'Task started'"
48
+ # after_replan:
49
+ # - "echo 'Task returned to planning'"
48
50
  # after_finish:
49
51
  # - "echo 'Task finished'"
50
52
  # after_archive:
@@ -33,7 +33,7 @@ DIR_SHELL_TICKETS = "shell-tickets"
33
33
  # platform that works today.
34
34
  DIR_LEGACY_CURSOR_SHELL_TICKETS = "cursor-shell"
35
35
  SHELL_TICKET_TTL_SECONDS = 30
36
- TASK_SESSION_COMMANDS = {"start", "current", "finish"}
36
+ TASK_SESSION_COMMANDS = {"start", "replan", "current", "finish"}
37
37
 
38
38
  _SESSION_KEYS = ("session_id", "sessionId", "sessionID")
39
39
  _CONVERSATION_KEYS = ("conversation_id", "conversationId", "conversationID")
@@ -425,7 +425,7 @@ def _pending_ticket_matches_args(ticket: dict[str, Any], repo_root: Path) -> boo
425
425
  continue
426
426
  if _string_value(subcommand.get("name")) != command_name:
427
427
  continue
428
- if command_name != "start":
428
+ if command_name not in {"start", "replan"}:
429
429
  return True
430
430
  task_ref = args[1] if len(args) > 1 else None
431
431
  if _task_refs_match(_string_value(subcommand.get("task_ref")), task_ref, repo_root):
@@ -22,10 +22,12 @@ from common.task_utils import is_within_tasks_dir, resolve_task_dir
22
22
  from common.tasks import load_task
23
23
 
24
24
 
25
- SCHEMA_VERSION = 1
25
+ SCHEMA_VERSION = 2
26
26
  MAX_DRAFT_BYTES = 64 * 1024
27
27
  MAX_REPORT_BYTES = 128 * 1024
28
+ MAX_WORKLOG_BYTES = 128 * 1024
28
29
  ID_RE = re.compile(r"^[A-Za-z0-9][A-Za-z0-9._-]{0,63}$")
30
+ CHECKPOINT_RE = re.compile(r"^<!-- trellis-checkpoint: (?P<payload>\{.*\}) -->$", re.MULTILINE)
29
31
  SECRET_PATTERNS = (
30
32
  re.compile(r"-----BEGIN [A-Z ]*PRIVATE KEY-----"),
31
33
  re.compile(r"\bsk-[A-Za-z0-9]{20,}\b"),
@@ -305,7 +307,7 @@ def _validate_brief_file(
305
307
  return brief, raw, brief_path, node_dir, task_dir
306
308
 
307
309
 
308
- def _validate_evidence(value: Any) -> set[tuple[str, str]]:
310
+ def _validate_evidence(value: Any) -> set[str]:
309
311
  if not isinstance(value, list):
310
312
  _fail("report.evidence must be a list")
311
313
  identities: set[tuple[str, str]] = set()
@@ -319,13 +321,100 @@ def _validate_evidence(value: Any) -> set[tuple[str, str]]:
319
321
  if identity in identities:
320
322
  _fail("report.evidence contains a duplicate evidence identity")
321
323
  identities.add(identity)
322
- return identities
324
+ return {evidence_id for evidence_id, _locator in identities}
325
+
326
+
327
+ def _validate_id_list(value: Any, field: str) -> list[str]:
328
+ if not isinstance(value, list):
329
+ _fail(f"{field} must be a list")
330
+ result = []
331
+ for index, item in enumerate(value):
332
+ result.append(_require_text(item, f"{field}[{index}]", max_len=128))
333
+ if len(result) != len(set(result)):
334
+ _fail(f"{field} must not contain duplicates")
335
+ return result
336
+
337
+
338
+ def _validate_typed_notes(value: Any, field: str) -> None:
339
+ if not isinstance(value, list):
340
+ _fail(f"report.{field} must be a list")
341
+ for index, item in enumerate(value):
342
+ if not isinstance(item, dict):
343
+ _fail(f"report.{field}[{index}] must be an object")
344
+ _require_id(item.get("id"), f"report.{field}[{index}].id")
345
+ _require_text(item.get("type"), f"report.{field}[{index}].type", max_len=64)
346
+ _require_text(item.get("detail"), f"report.{field}[{index}].detail")
347
+ _validate_id_list(item.get("evidence_ids", []), f"report.{field}[{index}].evidence_ids")
348
+
349
+
350
+ def _validate_checkpoint(node_dir: Path, evidence_ids: set[str], scope: list[str]) -> list[str]:
351
+ worklog_path = node_dir / "worklog.md"
352
+ if worklog_path.is_symlink():
353
+ _fail(f"refusing symlinked worklog: {worklog_path}")
354
+ try:
355
+ raw = worklog_path.read_bytes()
356
+ except FileNotFoundError:
357
+ _fail(f"worklog does not exist: {worklog_path}")
358
+ except OSError as exc:
359
+ _fail(f"could not read worklog {worklog_path}: {exc}")
360
+ if len(raw) > MAX_WORKLOG_BYTES:
361
+ _fail(f"worklog exceeds the {MAX_WORKLOG_BYTES} byte limit: {worklog_path}")
362
+ _reject_obvious_secrets(raw, "worklog")
363
+ try:
364
+ text = raw.decode("utf-8")
365
+ except UnicodeDecodeError:
366
+ _fail(f"worklog is not UTF-8: {worklog_path}")
367
+ matches = list(CHECKPOINT_RE.finditer(text))
368
+ if not matches:
369
+ return ["missing_worklog_checkpoint"]
370
+ concerns: list[str] = []
371
+ checkpoint_ids: set[str] = set()
372
+ for match in matches:
373
+ try:
374
+ checkpoint = json.loads(match.group("payload"))
375
+ except json.JSONDecodeError:
376
+ concerns.append("malformed_worklog_checkpoint")
377
+ continue
378
+ if not isinstance(checkpoint, dict):
379
+ concerns.append("malformed_worklog_checkpoint")
380
+ continue
381
+ try:
382
+ checkpoint_id = _require_id(checkpoint.get("id"), "worklog checkpoint.id")
383
+ if checkpoint_id in checkpoint_ids:
384
+ concerns.append("duplicate_worklog_checkpoint")
385
+ checkpoint_ids.add(checkpoint_id)
386
+ covered_scope = _require_text_list(
387
+ checkpoint.get("covered_scope"),
388
+ "worklog checkpoint.covered_scope",
389
+ allow_empty=True,
390
+ )
391
+ for item in covered_scope:
392
+ if item not in scope:
393
+ concerns.append("checkpoint_scope_outside_brief")
394
+ checkpoint_evidence = _validate_id_list(
395
+ checkpoint.get("evidence_ids"), "worklog checkpoint.evidence_ids"
396
+ )
397
+ if any(item not in evidence_ids for item in checkpoint_evidence):
398
+ concerns.append("checkpoint_evidence_unresolved")
399
+ _require_text(
400
+ checkpoint.get("conclusion_or_blocker"),
401
+ "worklog checkpoint.conclusion_or_blocker",
402
+ )
403
+ _require_text_list(
404
+ checkpoint.get("unknowns"),
405
+ "worklog checkpoint.unknowns",
406
+ allow_empty=True,
407
+ )
408
+ _require_text(checkpoint.get("safe_resume_point"), "worklog checkpoint.safe_resume_point")
409
+ except ArtifactError:
410
+ concerns.append("malformed_worklog_checkpoint")
411
+ return sorted(set(concerns))
323
412
 
324
413
 
325
414
  def _validate_report_data(
326
415
  report: dict[str, Any],
327
416
  brief: dict[str, Any],
328
- ) -> set[tuple[str, str]]:
417
+ ) -> tuple[set[str], list[str]]:
329
418
  if report.get("schema_version") != SCHEMA_VERSION:
330
419
  _fail(f"report.schema_version must be {SCHEMA_VERSION}")
331
420
  for field in ("task_id", "work_id", "subnode_id", "role_id"):
@@ -339,11 +428,52 @@ def _validate_report_data(
339
428
  if status not in {"complete", "blocked", "incomplete", "error"}:
340
429
  _fail("report.status must be complete, blocked, incomplete, or error")
341
430
  evidence = _validate_evidence(report.get("evidence"))
342
- for field in ("findings", "uncertainties", "corrections"):
343
- if not isinstance(report.get(field), list):
344
- _fail(f"report.{field} must be a list")
431
+ assessment = report.get("scope_assessment")
432
+ if not isinstance(assessment, list) or len(assessment) != len(brief["scope"]):
433
+ _fail("report.scope_assessment must contain one item for every brief scope item")
434
+ concerns: list[str] = []
435
+ for index, item in enumerate(assessment):
436
+ if not isinstance(item, dict):
437
+ _fail(f"report.scope_assessment[{index}] must be an object")
438
+ if item.get("scope") != brief["scope"][index]:
439
+ _fail(f"report.scope_assessment[{index}].scope does not match brief.scope")
440
+ if item.get("status") not in {"covered", "inconclusive", "not-started"}:
441
+ _fail(f"report.scope_assessment[{index}].status is invalid")
442
+ _require_text(item.get("conclusion"), f"report.scope_assessment[{index}].conclusion")
443
+ assessment_evidence = _validate_id_list(
444
+ item.get("evidence_ids"), f"report.scope_assessment[{index}].evidence_ids"
445
+ )
446
+ if any(value not in evidence for value in assessment_evidence):
447
+ concerns.append("scope_assessment_evidence_unresolved")
448
+ if item["status"] != "covered":
449
+ concerns.append("incomplete_scope_coverage")
450
+ if item["status"] == "covered" and not assessment_evidence:
451
+ concerns.append("covered_scope_without_evidence")
452
+ findings = report.get("findings")
453
+ if not isinstance(findings, list):
454
+ _fail("report.findings must be a list")
455
+ finding_ids: set[str] = set()
456
+ for index, item in enumerate(findings):
457
+ if not isinstance(item, dict):
458
+ _fail(f"report.findings[{index}] must be an object")
459
+ finding_id = _require_id(item.get("id"), f"report.findings[{index}].id")
460
+ if finding_id in finding_ids:
461
+ _fail("report.findings contains duplicate ids")
462
+ finding_ids.add(finding_id)
463
+ _require_text(item.get("conclusion"), f"report.findings[{index}].conclusion")
464
+ finding_evidence = _validate_id_list(
465
+ item.get("evidence_ids"), f"report.findings[{index}].evidence_ids"
466
+ )
467
+ if not finding_evidence:
468
+ concerns.append("finding_without_evidence")
469
+ if any(value not in evidence for value in finding_evidence):
470
+ concerns.append("finding_evidence_unresolved")
471
+ _validate_typed_notes(report.get("uncertainties"), "uncertainties")
472
+ _validate_typed_notes(report.get("corrections"), "corrections")
345
473
  if status == "complete" and not evidence:
346
474
  _fail("a complete report requires at least one evidence item")
475
+ if status == "complete" and not findings:
476
+ _fail("a complete report requires at least one finding")
347
477
  if status != "complete":
348
478
  _require_text_list(
349
479
  report.get("completed_scope"),
@@ -351,13 +481,13 @@ def _validate_report_data(
351
481
  allow_empty=True,
352
482
  )
353
483
  _require_text(report.get("blocker"), "report.blocker")
354
- return evidence
484
+ return evidence, sorted(set(concerns))
355
485
 
356
486
 
357
487
  def _validate_report_file(
358
488
  value: str,
359
489
  repo_root: Path,
360
- ) -> tuple[dict[str, Any], set[tuple[str, str]], dict[str, Any], Path]:
490
+ ) -> tuple[dict[str, Any], set[str], dict[str, Any], Path, list[str]]:
361
491
  report_path, node_dir, _task_dir = _resolve_artifact_file(value, "report.json", repo_root)
362
492
  brief_path = node_dir / "brief.json"
363
493
  brief, _brief_raw, _brief_path, _brief_node, _brief_task = _validate_brief_file(
@@ -368,8 +498,9 @@ def _validate_report_file(
368
498
  _fail("brief.report_path does not point to the report being validated")
369
499
  report, raw = _read_json_object(report_path, MAX_REPORT_BYTES, "report")
370
500
  _reject_obvious_secrets(raw, "report")
371
- evidence = _validate_report_data(report, brief)
372
- return report, evidence, brief, node_dir
501
+ evidence, concerns = _validate_report_data(report, brief)
502
+ concerns.extend(_validate_checkpoint(node_dir, evidence, brief["scope"]))
503
+ return report, evidence, brief, node_dir, sorted(set(concerns))
373
504
 
374
505
 
375
506
  def _validate_disposition_data(
@@ -466,13 +597,20 @@ def _init(args: argparse.Namespace) -> None:
466
597
 
467
598
  def _validate(args: argparse.Namespace) -> None:
468
599
  repo_root = get_repo_root()
469
- report, _evidence, _brief, _node_dir_value = _validate_report_file(args.report, repo_root)
600
+ report, _evidence, _brief, _node_dir_value, concerns = _validate_report_file(args.report, repo_root)
601
+ if concerns:
602
+ print(json.dumps({
603
+ "status": "review_concern",
604
+ "subnode_id": report["subnode_id"],
605
+ "concerns": concerns,
606
+ }))
607
+ return
470
608
  print(f"Validated pending-review report: {report['subnode_id']}")
471
609
 
472
610
 
473
611
  def _disposition(args: argparse.Namespace) -> None:
474
612
  repo_root = get_repo_root()
475
- report, _evidence, brief, node_dir = _validate_report_file(args.report, repo_root)
613
+ report, _evidence, brief, node_dir, _concerns = _validate_report_file(args.report, repo_root)
476
614
  disposition_path = node_dir / "disposition.json"
477
615
  if disposition_path.exists() or disposition_path.is_symlink():
478
616
  _fail(f"disposition already exists and cannot be replaced: {disposition_path}")
@@ -514,10 +652,10 @@ def _validate_counter(args: argparse.Namespace) -> None:
514
652
  primary_dir = repo_root / primary_dir
515
653
  if not counter_dir.is_absolute():
516
654
  counter_dir = repo_root / counter_dir
517
- primary_report, primary_evidence, primary_brief, primary_node = _validate_report_file(
655
+ primary_report, primary_evidence, primary_brief, primary_node, _primary_concerns = _validate_report_file(
518
656
  str(primary_dir / "report.json"), repo_root
519
657
  )
520
- counter_report, counter_evidence, counter_brief, counter_node = _validate_report_file(
658
+ counter_report, counter_evidence, counter_brief, counter_node, _counter_concerns = _validate_report_file(
521
659
  str(counter_dir / "report.json"), repo_root
522
660
  )
523
661
  if primary_node == counter_node:
@@ -9,6 +9,7 @@ Usage:
9
9
  python3 task.py validate <dir> # Validate jsonl files
10
10
  python3 task.py list-context <dir> # List jsonl entries
11
11
  python3 task.py start <dir> # Set active task, record current branch
12
+ python3 task.py replan <dir> "<reason>" # Return an in-progress task to planning
12
13
  python3 task.py current [--source] [--json] # Show active task
13
14
  python3 task.py ownership <operation> ... # Manage formal handoff task ownership
14
15
  python3 task.py finish # Clear active task
@@ -29,7 +30,9 @@ from __future__ import annotations
29
30
 
30
31
  import argparse
31
32
  import json
33
+ import os
32
34
  import sys
35
+ from datetime import datetime, timezone
33
36
  from pathlib import Path
34
37
 
35
38
  from common.log import Colors, colored
@@ -301,6 +304,74 @@ def cmd_start(args: argparse.Namespace) -> int:
301
304
  return 1
302
305
 
303
306
 
307
+ def cmd_replan(args: argparse.Namespace) -> int:
308
+ """Return an in-progress task to planning without losing its binding."""
309
+ repo_root = get_repo_root()
310
+ full_path = resolve_task_dir(args.dir, repo_root)
311
+ if full_path is None or not full_path.is_dir():
312
+ print(colored(f"Error: Task not found: {args.dir}", Colors.RED), file=sys.stderr)
313
+ return 1
314
+
315
+ try:
316
+ assert_task_mutation_allowed(repo_root, full_path)
317
+ except (OwnershipError, OSError) as exc:
318
+ print(colored(f"Error: {exc}", Colors.RED), file=sys.stderr)
319
+ return 2
320
+
321
+ reason = " ".join(args.reason).strip()
322
+ if not reason:
323
+ print(colored("Error: replan reason must not be empty", Colors.RED), file=sys.stderr)
324
+ return 1
325
+
326
+ task_json_path = full_path / FILE_TASK_JSON
327
+ data, read_reason = read_json_checked(task_json_path)
328
+ if data is None:
329
+ problem, hint = describe_json_read_failure(task_json_path, read_reason)
330
+ print(colored(f"Error: {problem}", Colors.RED), file=sys.stderr)
331
+ print(hint, file=sys.stderr)
332
+ return 1
333
+ if data.get("status") != "in_progress":
334
+ print(
335
+ colored(
336
+ f"Error: replan requires status=in_progress (found {data.get('status')!r})",
337
+ Colors.RED,
338
+ ),
339
+ file=sys.stderr,
340
+ )
341
+ return 1
342
+
343
+ event = {
344
+ "timestamp": datetime.now(timezone.utc).isoformat().replace("+00:00", "Z"),
345
+ "task": full_path.relative_to(repo_root).as_posix(),
346
+ "from_status": "in_progress",
347
+ "reason": reason,
348
+ "branch": data.get("branch"),
349
+ }
350
+ context_key = resolve_context_key()
351
+ if context_key:
352
+ event["session"] = context_key
353
+
354
+ replans_path = full_path / "replans.jsonl"
355
+ try:
356
+ with replans_path.open("a", encoding="utf-8") as handle:
357
+ handle.write(json.dumps(event, ensure_ascii=False) + "\n")
358
+ handle.flush()
359
+ os.fsync(handle.fileno())
360
+ except OSError as exc:
361
+ print(colored(f"Error: could not record replan: {exc}", Colors.RED), file=sys.stderr)
362
+ return 1
363
+
364
+ data["status"] = "planning"
365
+ if not write_json(task_json_path, data):
366
+ print(colored("Error: replan event recorded but task status was not changed", Colors.RED), file=sys.stderr)
367
+ return 1
368
+
369
+ print(colored(f"✓ Task returned to planning: {full_path.relative_to(repo_root)}", Colors.GREEN))
370
+ print("Reason:", reason)
371
+ run_task_hooks("after_replan", task_json_path, repo_root)
372
+ return 0
373
+
374
+
304
375
  def cmd_finish(args: argparse.Namespace) -> int:
305
376
  """Clear active task."""
306
377
  repo_root = get_repo_root()
@@ -900,6 +971,11 @@ def main() -> int:
900
971
  help="Start even when implement.jsonl / check.jsonl have no curated entries",
901
972
  )
902
973
 
974
+ # replan
975
+ p_replan = subparsers.add_parser("replan", help="Return an in-progress task to planning")
976
+ p_replan.add_argument("dir", help="Task directory")
977
+ p_replan.add_argument("reason", nargs="+", help="Material reason for returning to planning")
978
+
903
979
  # current
904
980
  p_current = subparsers.add_parser("current", help="Show active task")
905
981
  p_current.add_argument("--source", action="store_true",
@@ -1047,6 +1123,7 @@ def main() -> int:
1047
1123
  "validate": cmd_validate,
1048
1124
  "list-context": cmd_list_context,
1049
1125
  "start": cmd_start,
1126
+ "replan": cmd_replan,
1050
1127
  "current": cmd_current,
1051
1128
  "continuity": cmd_continuity,
1052
1129
  "ownership": cmd_ownership,
@@ -45,6 +45,7 @@ Every task has its own directory under `.trellis/tasks/{MM-DD-name}/` holding `t
45
45
  # Task lifecycle
46
46
  python3 ./.trellis/scripts/task.py create "<title>" [--slug <name>] [--parent <dir>]
47
47
  python3 ./.trellis/scripts/task.py start <name> # set active task (session-scoped when available)
48
+ python3 ./.trellis/scripts/task.py replan <name> "<reason>" # return material ambiguity to planning
48
49
  python3 ./.trellis/scripts/task.py current --source # show active task and source
49
50
  python3 ./.trellis/scripts/task.py finish # clear active task (triggers after_finish hooks)
50
51
  python3 ./.trellis/scripts/task.py archive <name> # move to archive/{year-month}/
@@ -166,6 +167,8 @@ Phase 3: Finish → verify, update spec, commit, and wrap up
166
167
 
167
168
  An analysis-only task is eligible only when `task.json.meta.delivery_mode = "analysis_only"` exactly and its `prd.md` names the evidence deliverable plus a no-change boundary for product source, runtime configuration, deployment, credentials, and external systems. Task creation consent authorizes that bounded evidence work, not protected-target changes.
168
169
 
170
+ The exception is invalid when the work still needs a material user decision, design or implementation plan, cross-owner coordination, security or deployment change, release or credential action, or a protected downstream task. "Research" and a deferred source edit do not override this classification; use normal complex planning when any condition applies.
171
+
169
172
  Keep an eligible analysis-only task in `planning`: write its research, audit, or design evidence; verify its acceptance criteria and boundary; commit task artifacts; then archive directly. Do not run `task.py start`, configure implementation context, or wait for a second implementation approval. If the evidence recommends a protected-target change, record the recommendation and create a separate change-bearing task before doing it.
170
173
 
171
174
  ### Planning Artifacts
@@ -176,6 +179,8 @@ Keep an eligible analysis-only task in `planning`: write its research, audit, or
176
179
  - `implement.jsonl` / `check.jsonl` — spec and research manifests for sub-agent context. They do not replace `implement.md`.
177
180
  - Lightweight tasks may be PRD-only. Complex tasks must have `prd.md`, `design.md`, and `implement.md` before `task.py start`.
178
181
 
182
+ For read-heavy research, audit, review, or investigation, split work that cannot persist one useful conclusion in the current bounded session into evidence units. Each unit records one question or scope, minimal evidence, destination, stop condition, conclusion or blocker, unknowns, and recovery point before the next unit begins. Size for one normal context window, without an exact token or time promise; routine navigation needs no record.
183
+
179
184
  ### Parent / Child Task Trees
180
185
 
181
186
  Use a parent task when one user request contains several independently verifiable deliverables. The parent task owns the source requirement set, the task map, cross-child acceptance criteria, and final integration review; it normally should not be the implementation target unless it also has direct work.
@@ -229,7 +234,7 @@ Preserve existing task fields and artifacts. If the correct status cannot be det
229
234
  [workflow-state:planning]
230
235
  Load `trellis-brainstorm`; stay in planning.
231
236
  If `task.json.meta.delivery_mode = "analysis_only"` exactly, complete the declared evidence work now. Do not wait for a start review or run `task.py start`; when the PRD boundary and acceptance evidence pass, commit task artifacts and archive directly. A protected-target recommendation requires a separate change-bearing task.
232
- Lightweight: `prd.md` can be enough. Complex: finish `prd.md`, `design.md`, and `implement.md`; ask for review before `task.py start`.
237
+ Lightweight: `prd.md` can be enough. Complex: finish `prd.md`, `design.md`, and `implement.md`; run the Planning Seal closure pass before asking for review. If `decision-needed` items or an unsealed decision graph remain, load `pennix-decision-gates`, batch only independent frontier questions, and stay in planning. Answers returned by the current continuation must be persisted and fed back into the same planning loop.
233
238
  Multi-deliverable scope: consider a parent task plus independently verifiable child tasks; dependencies must be written in child artifacts, not implied by tree position.
234
239
  Sub-agent mode: curate `implement.jsonl` and `check.jsonl` as spec/research manifests before start.
235
240
  [/workflow-state:planning]
@@ -243,7 +248,7 @@ Sub-agent mode: curate `implement.jsonl` and `check.jsonl` as spec/research mani
243
248
  [workflow-state:planning-inline]
244
249
  Load `trellis-brainstorm`; stay in planning.
245
250
  If `task.json.meta.delivery_mode = "analysis_only"` exactly, complete the declared evidence work now. Do not wait for a start review or run `task.py start`; when the PRD boundary and acceptance evidence pass, commit task artifacts and archive directly. A protected-target recommendation requires a separate change-bearing task.
246
- Lightweight: `prd.md` can be enough. Complex: finish `prd.md`, `design.md`, and `implement.md`; ask for review before `task.py start`.
251
+ Lightweight: `prd.md` can be enough. Complex: finish `prd.md`, `design.md`, and `implement.md`; run the Planning Seal closure pass before asking for review. If `decision-needed` items or an unsealed decision graph remain, load `pennix-decision-gates`, batch only independent frontier questions, and stay in planning. Answers returned by the current continuation must be persisted and fed back into the same planning loop.
247
252
  Multi-deliverable scope: consider a parent task plus independently verifiable child tasks; dependencies must be written in child artifacts, not implied by tree position.
248
253
  Inline mode: skip jsonl curation; Phase 2 reads artifacts/specs via `trellis-before-dev`.
249
254
  [/workflow-state:planning-inline]
@@ -262,6 +267,7 @@ Inline mode: skip jsonl curation; Phase 2 reads artifacts/specs via `trellis-bef
262
267
  Sub-agent dispatch protocol applies to all platforms and all sub-agents, including native Codex `SubagentStart` context injection with child-side pull fallback, class-2 Gemini/Qoder/Copilot/Reasonix/Trae/Grok/Kimi Code/DeepSeek Harness, hook-backed ZCode/Snow, and `trellis-research`: every dispatch prompt starts with `Active task: <task path from task.py current>` before role-specific instructions. On Grok Build, use `spawn_subagent` with `subagent_type` set to the Trellis agent name (e.g. `trellis-implement`). On Kimi Code, dispatch the built-in `coder` / `explore` sub-agent with the matching `.kimi-code/skills/trellis-<role>/SKILL.md` instructions. On DeepSeek Harness, tell the child to load the matching `.dsh/skills/trellis-agent-<role>/SKILL.md` exactly once, then choose the synchronization path by capability. If the optional companion plugin exposes `trellis_wait`, dispatch `subagent` in its default continuable background mode, continue independent work, and call `trellis_wait` once per dependent child id when a dependent gate is next; each call returns only after DSH has queued that child's native settlement notice. Without `trellis_wait`, dispatch each child with `run_in_background: false` from the outset so no dependent gate can overtake it. Never simulate waiting with shell sleep, polling loops, `job_output`, repeated `list_agents`, or another long-running command, and never leave a background child without an event-driven wait path.
263
268
 
264
269
  [workflow-state:in_progress]
270
+ If implementation discovers a material unresolved decision, record `decision-needed`, run `task.py replan <task> "<reason>"`, and return to the planning frontier; do not ask a native question during implementation.
265
271
  Tools: `trellis-implement` / `trellis-research` name sub-agent roles dispatched through your platform's sub-agent mechanism, not skills the main session loads itself (on Claude Code: use the Task/Agent tool, never the Skill tool). `trellis-update-spec` is a skill. `trellis-check` exists as both; prefer the Agent/role form when verifying after code changes.
266
272
  On DeepSeek Harness, role instructions ship as collision-free `trellis-agent-implement` / `trellis-agent-check` / `trellis-agent-research` skills under `.dsh/skills/`. The main session must not load them itself: tell the child to load the matching role skill exactly once. If `trellis_wait` is available, use the default background mode, do independent work, then call `trellis_wait` once per dependent child id and consume each native settlement notice before entering the dependent gate. If it is unavailable, dispatch every child with `run_in_background: false` from the outset. Do not poll, sleep, or start a background child without an event-driven wait path.
267
273
  Flow: `trellis-implement` -> `trellis-check` -> `trellis-update-spec` -> commit (Phase 3.4) -> `/trellis:finish-work`.
@@ -276,6 +282,7 @@ Dispatch prompt starts with `Active task: <task path from task.py current>`. Rea
276
282
 
277
283
  [workflow-state:in_progress-inline]
278
284
  Flow: `trellis-before-dev` -> edit -> `trellis-check` -> validation -> `trellis-update-spec` -> commit (Phase 3.4) -> `/trellis:finish-work`.
285
+ If implementation discovers a material unresolved decision, record `decision-needed`, run `task.py replan <task> "<reason>"`, and return to the planning frontier; native questions are planning-only.
279
286
  Do not dispatch implement/check sub-agents in inline mode.
280
287
  Read context: `prd.md` -> `design.md if present` -> `implement.md if present`, plus relevant spec/research loaded by skills.
281
288
  [/workflow-state:in_progress-inline]
@@ -379,13 +386,15 @@ Load the `trellis-brainstorm` skill and explore requirements interactively with
379
386
  For an eligible analysis-only task, converge the PRD boundary, then perform the declared research, audit, or design work in this phase. It does not need a start review or `task.py start`; after acceptance evidence is recorded, continue directly to Phase 3.3.
380
387
 
381
388
  The brainstorm skill will guide you to:
382
- - Ask one question at a time
389
+ - Batch up to three independent material frontier questions; ask one only when dependency ordering requires it
383
390
  - Prefer researching over asking the user
384
391
  - Prefer offering options over open-ended questions
385
392
  - Update `prd.md` immediately after each user answer
386
393
  - Split large scopes into a parent task plus child tasks when the deliverables can be verified independently
387
394
  - Keep `prd.md` focused on requirements and acceptance criteria
388
395
  - For complex tasks, produce `design.md` and `implement.md` before implementation starts
396
+ - For read-heavy work, split long investigation into evidence units and persist each unit's conclusion or recovery point before continuing
397
+ - Before review or `task.py start`, run the Planning Seal closure pass across all task artifacts and lock targets, branches, dependencies, release, validation, rollback, dynamic-fact handling, and every material decision
389
398
 
390
399
  When considering a parent/child split:
391
400
  - Use a parent task when one request contains several independently verifiable deliverables.
@@ -487,7 +496,7 @@ Skip this step. Context is loaded directly by the `trellis-before-dev` skill in
487
496
 
488
497
  This step applies only to change-bearing tasks. An eligible analysis-only task stays in `planning` and, after its evidence work is complete, continues directly to Phase 3.3 without running `task.py start`.
489
498
 
490
- After artifact review, flip the task status to `in_progress`:
499
+ After the Planning Seal closure pass and artifact review, flip the task status to `in_progress`:
491
500
 
492
501
  ```bash
493
502
  python3 ./.trellis/scripts/task.py start <task-dir>
@@ -510,6 +519,7 @@ If `task.py start` errors with a session-identity message (no context key from h
510
519
  | `research/` has artifacts (complex tasks) | recommended |
511
520
  | `design.md` exists (complex tasks) | ✅ |
512
521
  | `implement.md` exists (complex tasks) | ✅ |
522
+ | Planning Seal closure pass recorded; static decisions and implementation paths are locked | ✅ |
513
523
 
514
524
  [Claude Code, Cursor, OpenCode, codex-sub-agent, Kiro, Gemini, Qoder, CodeBuddy, Copilot, Droid, Pi, Oh My Pi, ZCode, Snow, Reasonix, Trae, Grok, Kimi Code, DeepSeek Harness]
515
525
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@pennixrv/trellis",
3
- "version": "0.7.0-beta.7",
3
+ "version": "0.7.0-beta.9",
4
4
  "description": "AI capabilities grow like ivy — Trellis provides the structure to guide them along a disciplined path",
5
5
  "type": "module",
6
6
  "main": "./dist/index.js",
@@ -34,7 +34,7 @@
34
34
  "inquirer": "^9.3.7",
35
35
  "undici": "^6.21.0",
36
36
  "zod": "^4.4.2",
37
- "@pennixrv/trellis-core": "0.7.0-beta.7"
37
+ "@pennixrv/trellis-core": "0.7.0-beta.9"
38
38
  },
39
39
  "devDependencies": {
40
40
  "@eslint/js": "^9.18.0",