@pennixrv/trellis 0.7.0-beta.8 → 0.7.0-beta.9
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/migrations/manifests/0.7.0-beta.9.json +9 -0
- package/dist/templates/common/bundled-skills/trellis-channel/references/subnode-work.md +32 -3
- package/dist/templates/common/commands/continue.md +2 -2
- package/dist/templates/common/skills/brainstorm.md +11 -3
- package/dist/templates/copilot/prompts/brainstorm.prompt.md +13 -3
- package/dist/templates/trellis/agents/subnode.md +11 -4
- package/dist/templates/trellis/scripts/subnode_artifact.py +153 -15
- package/dist/templates/trellis/workflow.md +10 -3
- package/package.json +2 -2
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
{
|
|
2
|
+
"version": "0.7.0-beta.9",
|
|
3
|
+
"description": "Pennix v0.7 beta planning and evidence governance",
|
|
4
|
+
"breaking": false,
|
|
5
|
+
"recommendMigrate": false,
|
|
6
|
+
"changelog": "**Features:**\n- Require complex analysis to use normal planning and a sealed decision chain.\n- Add evidence units and schema-v2 subnode reports with coordinator-only acceptance.",
|
|
7
|
+
"migrations": [],
|
|
8
|
+
"notes": "Run npm install -g @pennixrv/trellis@beta then trellis update. No --migrate required."
|
|
9
|
+
}
|
|
@@ -50,12 +50,15 @@ and its status; the Codex supervisor projects that reply into the durable
|
|
|
50
50
|
Channel message and `done` events. The durable JSON file carries the
|
|
51
51
|
reviewable result.
|
|
52
52
|
|
|
53
|
-
The subnode copies identity, scope, and lens from `brief.json`.
|
|
53
|
+
The subnode copies identity, scope, and lens from `brief.json`. Schema version 2
|
|
54
|
+
requires one `scope_assessment` entry for every brief scope item, structured
|
|
55
|
+
findings with evidence references, and typed uncertainties or corrections when
|
|
56
|
+
present. Schema version 1 is retired and the validator rejects it. The minimal
|
|
54
57
|
complete report is:
|
|
55
58
|
|
|
56
59
|
```json
|
|
57
60
|
{
|
|
58
|
-
"schema_version":
|
|
61
|
+
"schema_version": 2,
|
|
59
62
|
"task_id": "task-id-from-brief",
|
|
60
63
|
"work_id": "work-id-from-brief",
|
|
61
64
|
"subnode_id": "subnode-id-from-brief",
|
|
@@ -63,6 +66,14 @@ complete report is:
|
|
|
63
66
|
"status": "complete",
|
|
64
67
|
"scope": ["exact scope copied from brief"],
|
|
65
68
|
"lens": "exact lens copied from brief",
|
|
69
|
+
"scope_assessment": [
|
|
70
|
+
{
|
|
71
|
+
"scope": "exact scope item from brief",
|
|
72
|
+
"status": "covered",
|
|
73
|
+
"conclusion": "What this scope establishes.",
|
|
74
|
+
"evidence_ids": ["stable-evidence-id"]
|
|
75
|
+
}
|
|
76
|
+
],
|
|
66
77
|
"evidence": [
|
|
67
78
|
{
|
|
68
79
|
"id": "stable-evidence-id",
|
|
@@ -70,12 +81,30 @@ complete report is:
|
|
|
70
81
|
"summary": "What this independently reviewable evidence establishes."
|
|
71
82
|
}
|
|
72
83
|
],
|
|
73
|
-
"findings": [
|
|
84
|
+
"findings": [
|
|
85
|
+
{
|
|
86
|
+
"id": "finding-id",
|
|
87
|
+
"conclusion": "Bounded conclusion.",
|
|
88
|
+
"evidence_ids": ["stable-evidence-id"]
|
|
89
|
+
}
|
|
90
|
+
],
|
|
74
91
|
"uncertainties": [],
|
|
75
92
|
"corrections": []
|
|
76
93
|
}
|
|
77
94
|
```
|
|
78
95
|
|
|
96
|
+
Append a checkpoint marker to `worklog.md` when a material unit is complete or
|
|
97
|
+
blocked. It is a bounded recovery projection, not coordinator acceptance:
|
|
98
|
+
|
|
99
|
+
```text
|
|
100
|
+
<!-- trellis-checkpoint: {"id":"checkpoint-1","covered_scope":["exact scope item"],"evidence_ids":["stable-evidence-id"],"conclusion_or_blocker":"Current conclusion or blocker.","unknowns":[],"safe_resume_point":"Next safe action."} -->
|
|
101
|
+
```
|
|
102
|
+
|
|
103
|
+
The validator returns `review_concern` for missing or incomplete checkpoint
|
|
104
|
+
coverage and for inconclusive scope assessments so the coordinator can inspect
|
|
105
|
+
them. Identity, schema, path, and malformed-structure failures remain hard
|
|
106
|
+
errors. A concern is never automatic acceptance or rejection.
|
|
107
|
+
|
|
79
108
|
For `blocked`, `incomplete`, or `error`, include the same base fields plus a
|
|
80
109
|
`completed_scope` list (empty when no assigned scope started) and a non-empty
|
|
81
110
|
`blocker` string. Never use
|
|
@@ -24,13 +24,13 @@ Shows the Phase Index (Plan / Execute / Finish) with routing + skill mapping.
|
|
|
24
24
|
|
|
25
25
|
`get_context.py` shows the active task's `status` field. Route by `status` + artifact presence. This command replaces the user needing to remember the Trellis flow; it does not itself approve implementation.
|
|
26
26
|
|
|
27
|
-
- `status=planning` + `task.json.meta.delivery_mode = "analysis_only"` → complete the PRD's
|
|
27
|
+
- `status=planning` + `task.json.meta.delivery_mode = "analysis_only"` → first confirm the task still satisfies the bounded evidence-only eligibility rule; then complete the PRD's evidence work, verify its acceptance criteria and no-change boundary, and archive directly. Do not run `task.py start`; a protected-target change requires a separate change-bearing task.
|
|
28
28
|
- `status=planning` + no `prd.md` → **1.1** (load `trellis-brainstorm`)
|
|
29
29
|
- `status=planning` + a recorded `decision-needed` or unsealed decision chain → return to the planning frontier and load `pennix-decision-gates` when independent material questions can be batched.
|
|
30
30
|
- `status=in_progress` + a material unresolved decision → record the reason and run `task.py replan <task> "<reason>"`; do not ask a native question during implementation.
|
|
31
31
|
- `status=planning` + `prd.md` only → decide whether the task is lightweight or complex. Lightweight can move to **1.4** review; complex returns to **1.1** to add `design.md` + `implement.md`.
|
|
32
32
|
- `status=planning` + complex artifacts complete + sub-agent jsonl not curated (empty, or only a legacy `_example` placeholder row) → **1.3**
|
|
33
|
-
- `status=planning` + required artifacts complete + required jsonl curated or inline mode → **1.4** (ask for start review; only run `task.py start` after user confirms)
|
|
33
|
+
- `status=planning` + required artifacts complete + required jsonl curated or inline mode → run the Planning Seal closure pass, then **1.4** (ask for start review; only run `task.py start` after user confirms)
|
|
34
34
|
- `status=in_progress` + implementation not started → **2.1**
|
|
35
35
|
- `status=in_progress` + implementation done, not yet checked → **2.2**
|
|
36
36
|
- `status=in_progress` + check passed → **3.3** (spec update) → **3.4** (commit)
|
|
@@ -12,6 +12,8 @@ While any user-owned product, scope, UX, compatibility, risk, or acceptance deci
|
|
|
12
12
|
|
|
13
13
|
When `task.json.meta.delivery_mode = "analysis_only"` exactly and the PRD names a bounded evidence deliverable plus a no-change boundary for product source, runtime configuration, deployment, credentials, and external systems, task-creation consent authorizes that evidence work. Do not require a second planning approval or run `task.py start`: perform the declared research, audit, or design work while status remains `planning`, record the evidence, verify acceptance criteria and the boundary, commit task artifacts, and archive directly. If the evidence recommends a protected-target change, record it and create a separate change-bearing task before doing it.
|
|
14
14
|
|
|
15
|
+
This exception is eligible only for a bounded evidence deliverable with no material user decision, design or implementation plan, cross-owner coordination, security or deployment change, release or credential action, or protected downstream task. Calling work "research", deferring source edits, or working in an audit/root repository does not make it analysis-only. If any of those conditions apply, use the normal complex planning and implementation-approval path.
|
|
16
|
+
|
|
15
17
|
All other tasks follow the planning and implementation approval gates below.
|
|
16
18
|
|
|
17
19
|
## Non-Negotiable Evidence Rule
|
|
@@ -24,6 +26,10 @@ Do not ask the user to confirm facts that the repository can answer. Ask only fo
|
|
|
24
26
|
|
|
25
27
|
Repository evidence establishes current behavior and technical constraints. The user's intended behavior, feature scope boundaries, and UX preferences are never answerable by repository evidence alone, even when an existing pattern exists; existing patterns are options and recommendation evidence, not decisions.
|
|
26
28
|
|
|
29
|
+
## Evidence Units For Read-Heavy Work
|
|
30
|
+
|
|
31
|
+
When research, audit, review, or investigation is too large to leave one independently useful conclusion in the current bounded session, split it into evidence units. Each unit must have one question or scope, a minimal evidence range, a destination artifact, and a stop condition; write its facts, conclusion or blocker, unknowns, and recovery point before starting another unit. Size units so one normal context window can finish and persist one useful result; do not promise an exact token or time limit. Routine navigation and transient tool output do not need an artifact. Create a child task only when the unit has an independent owner, lifecycle, and acceptance contract.
|
|
32
|
+
|
|
27
33
|
---
|
|
28
34
|
|
|
29
35
|
Use this skill during Phase 1 planning to turn the user's request into clear requirements and planning artifacts.
|
|
@@ -54,10 +60,10 @@ Use a concise title from the user's request. Both the title and `--description`
|
|
|
54
60
|
- product intent still needed from the user
|
|
55
61
|
- scope or risk decisions still needed from the user
|
|
56
62
|
- likely out-of-scope items
|
|
57
|
-
4. If user-owned decisions remain, calculate the independent frontier. Use `pennix-decision-gates` for a bounded batch when two or more independent material decisions are ready; otherwise ask the single highest-value question. Include recommendation and trade-off
|
|
58
|
-
5.
|
|
63
|
+
4. If user-owned decisions remain, calculate the independent frontier. Use `pennix-decision-gates` for a bounded batch when two or more independent material decisions are ready; otherwise ask the single highest-value question. Include recommendation and trade-off. Yield only while the answer is unavailable.
|
|
64
|
+
5. When the host returns the current continuation's answer, immediately persist it in `prd.md` or the decision artifact, recheck evidence and conflicts, recalculate the frontier, and continue the same planning loop. Do not create a second Trellis lifecycle for the same decision chain. Stop only for a new unresolved frontier, a real capability or authority block, or a final sealed summary awaiting implementation approval.
|
|
59
65
|
6. When no user-owned decision remains, create or update `design.md` and `implement.md` for complex tasks.
|
|
60
|
-
7. Run the requirement convergence gate, then the PRD convergence pass.
|
|
66
|
+
7. Run the requirement convergence gate, then the PRD convergence pass. Finish with one Planning Seal closure pass.
|
|
61
67
|
8. Present the final planning summary and stop. Do not run `task.py start` or edit product code in the same turn.
|
|
62
68
|
9. Only a subsequent user message that explicitly approves the latest planning summary authorizes `task.py start` and implementation. If implementation reveals a material unresolved decision, record `decision-needed`, run `task.py replan <task> "<reason>"`, and return through this planning flow; do not open a popup during implementation.
|
|
63
69
|
|
|
@@ -140,6 +146,8 @@ Lightweight tasks may omit `design.md` and `implement.md`; they may not skip evi
|
|
|
140
146
|
|
|
141
147
|
The final planning summary must show Goal, In Scope, Out of Scope, Acceptance Criteria, Key Decisions, relevant Risks or Deferred Items, and artifact status.
|
|
142
148
|
|
|
149
|
+
The Planning Seal closure pass must reconcile `task.json`, `prd.md`, `design.md`, `implement.md`, research, decision records, and manifests; verify the actual modification targets and branches, ordered dependencies and release steps, validation and rollback, dynamic-fact dispositions and replan triggers, and that every material decision has an owner and a fixed outcome. Remove static ambiguity before implementation: no `TBD`, `TODO`, `decision-needed`, unowned option, unspecified branch, open implementation path, validation gap, or conditional acceptance may remain. A material discovery invalidates the seal and returns to planning; implementation may consume only a sealed plan.
|
|
150
|
+
|
|
143
151
|
## Artifact Rules
|
|
144
152
|
|
|
145
153
|
`prd.md` records requirements and acceptance:
|
|
@@ -12,6 +12,10 @@ For every non-trivial task, the user must respond at least once after the initia
|
|
|
12
12
|
|
|
13
13
|
While any user-owned product, scope, UX, compatibility, risk, or acceptance decision remains unresolved, keep the task in planning. First inventory evidence and decision dependencies. If at least two independent material decisions remain and `pennix-decision-gates` is available, delegate one bounded batch of up to three frontier questions; otherwise ask the single highest-value question. Do not edit product code, dispatch implementation, or run `task.py start` until the decision chain is sealed.
|
|
14
14
|
|
|
15
|
+
## Analysis-Only Exception
|
|
16
|
+
|
|
17
|
+
When `task.json.meta.delivery_mode = "analysis_only"` exactly and the PRD names a bounded evidence deliverable plus a no-change boundary for product source, runtime configuration, deployment, credentials, and external systems, task-creation consent authorizes that evidence work. Keep status `planning`, record and verify the evidence, commit task artifacts, and archive directly; do not run `task.py start` or wait for a second implementation approval. This exception is eligible only when there is no material user decision, design or implementation plan, cross-owner coordination, security or deployment change, release or credential action, or protected downstream task. Otherwise use normal complex planning.
|
|
18
|
+
|
|
15
19
|
## Non-Negotiable Evidence Rule
|
|
16
20
|
|
|
17
21
|
If a question can be answered by exploring the codebase, explore the codebase instead.
|
|
@@ -22,6 +26,10 @@ Do not ask the user to confirm facts that the repository can answer. Ask only fo
|
|
|
22
26
|
|
|
23
27
|
Repository evidence establishes current behavior and technical constraints. The user's intended behavior, feature scope boundaries, and UX preferences are never answerable by repository evidence alone, even when an existing pattern exists; existing patterns are options and recommendation evidence, not decisions.
|
|
24
28
|
|
|
29
|
+
## Evidence Units For Read-Heavy Work
|
|
30
|
+
|
|
31
|
+
When research, audit, review, or investigation is too large for one independently useful conclusion in the current session, split it into evidence units. Each unit has one question or scope, a minimal evidence range, a destination artifact, and a stop condition; persist facts, conclusion or blocker, unknowns, and a recovery point before starting another unit. Size each unit for one normal context window without promising an exact token or time limit.
|
|
32
|
+
|
|
25
33
|
---
|
|
26
34
|
|
|
27
35
|
Use this skill during Phase 1 planning to turn the user's request into clear requirements and planning artifacts.
|
|
@@ -52,10 +60,10 @@ Use a concise title from the user's request. Both the title and `--description`
|
|
|
52
60
|
- product intent still needed from the user
|
|
53
61
|
- scope or risk decisions still needed from the user
|
|
54
62
|
- likely out-of-scope items
|
|
55
|
-
4. If user-owned decisions remain, calculate the independent frontier. Use `pennix-decision-gates` for a bounded batch when two or more independent material decisions are ready; otherwise ask the single highest-value question. Include recommendation and trade-off
|
|
56
|
-
5.
|
|
63
|
+
4. If user-owned decisions remain, calculate the independent frontier. Use `pennix-decision-gates` for a bounded batch when two or more independent material decisions are ready; otherwise ask the single highest-value question. Include recommendation and trade-off. Yield only while the answer is unavailable.
|
|
64
|
+
5. When the host returns the current continuation's answer, persist it in `prd.md` or the decision artifact, recheck evidence and conflicts, recalculate the frontier, and continue the same planning loop. Stop only for a new unresolved frontier, a real capability or authority block, or a final sealed summary awaiting implementation approval.
|
|
57
65
|
6. When no user-owned decision remains, create or update `design.md` and `implement.md` for complex tasks.
|
|
58
|
-
7. Run the requirement convergence gate, then the PRD convergence pass.
|
|
66
|
+
7. Run the requirement convergence gate, then the PRD convergence pass. Finish with one Planning Seal closure pass.
|
|
59
67
|
8. Present the final planning summary and stop. Do not run `task.py start` or edit product code in the same turn.
|
|
60
68
|
9. Only a subsequent user message that explicitly approves the latest planning summary authorizes `task.py start` and implementation. If implementation reveals a material unresolved decision, record `decision-needed`, run `task.py replan <task> "<reason>"`, and return through this planning flow; do not open a popup during implementation.
|
|
61
69
|
|
|
@@ -95,6 +103,8 @@ Lightweight tasks may omit `design.md` and `implement.md`; they may not skip evi
|
|
|
95
103
|
|
|
96
104
|
The final planning summary must show Goal, In Scope, Out of Scope, Acceptance Criteria, Key Decisions, relevant Risks or Deferred Items, and artifact status.
|
|
97
105
|
|
|
106
|
+
The Planning Seal closure pass reconciles `task.json`, `prd.md`, `design.md`, `implement.md`, research, decision records, and manifests; verifies targets, branches, dependencies, release, validation, rollback, dynamic-fact dispositions, and replan triggers; and fixes every material decision to an owner and outcome. No `TBD`, `TODO`, `decision-needed`, unowned option, unspecified branch, open implementation path, validation gap, or conditional acceptance may remain. Any material discovery invalidates the seal and returns to planning.
|
|
107
|
+
|
|
98
108
|
## Artifact Rules
|
|
99
109
|
|
|
100
110
|
`prd.md` records requirements and acceptance:
|
|
@@ -49,10 +49,17 @@ sandbox claim.
|
|
|
49
49
|
4. Write `report.json` only when you are ready to stop. Its status is one of
|
|
50
50
|
`complete`, `blocked`, `incomplete`, or `error`; it is always
|
|
51
51
|
**pending coordinator review**, never accepted/rejected/deferred.
|
|
52
|
-
5.
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
52
|
+
5. Use report schema version 2. Include one `scope_assessment` for each brief
|
|
53
|
+
scope item, structured findings with `id`, `conclusion`, and `evidence_ids`,
|
|
54
|
+
and typed `uncertainties` or `corrections` when present. A complete report
|
|
55
|
+
includes independently checkable evidence; a non-complete report explains
|
|
56
|
+
completed scope and the blocker or error. Include the exact identity, scope,
|
|
57
|
+
and lens required by the artifact helper.
|
|
58
|
+
6. Append a checkpoint marker to `worklog.md` after each material unit using
|
|
59
|
+
the exact `trellis-checkpoint` JSON fields documented by `subnode-work`.
|
|
60
|
+
Missing or incomplete checkpoint coverage becomes a coordinator review
|
|
61
|
+
concern; it does not become acceptance.
|
|
62
|
+
7. Finish with one short final assistant reply that states the status and report
|
|
56
63
|
path. Do not place the report JSON in the reply and do not run
|
|
57
64
|
`trellis channel send`: the supervisor routes this final reply into the
|
|
58
65
|
durable Channel message and `done` events.
|
|
@@ -22,10 +22,12 @@ from common.task_utils import is_within_tasks_dir, resolve_task_dir
|
|
|
22
22
|
from common.tasks import load_task
|
|
23
23
|
|
|
24
24
|
|
|
25
|
-
SCHEMA_VERSION =
|
|
25
|
+
SCHEMA_VERSION = 2
|
|
26
26
|
MAX_DRAFT_BYTES = 64 * 1024
|
|
27
27
|
MAX_REPORT_BYTES = 128 * 1024
|
|
28
|
+
MAX_WORKLOG_BYTES = 128 * 1024
|
|
28
29
|
ID_RE = re.compile(r"^[A-Za-z0-9][A-Za-z0-9._-]{0,63}$")
|
|
30
|
+
CHECKPOINT_RE = re.compile(r"^<!-- trellis-checkpoint: (?P<payload>\{.*\}) -->$", re.MULTILINE)
|
|
29
31
|
SECRET_PATTERNS = (
|
|
30
32
|
re.compile(r"-----BEGIN [A-Z ]*PRIVATE KEY-----"),
|
|
31
33
|
re.compile(r"\bsk-[A-Za-z0-9]{20,}\b"),
|
|
@@ -305,7 +307,7 @@ def _validate_brief_file(
|
|
|
305
307
|
return brief, raw, brief_path, node_dir, task_dir
|
|
306
308
|
|
|
307
309
|
|
|
308
|
-
def _validate_evidence(value: Any) -> set[
|
|
310
|
+
def _validate_evidence(value: Any) -> set[str]:
|
|
309
311
|
if not isinstance(value, list):
|
|
310
312
|
_fail("report.evidence must be a list")
|
|
311
313
|
identities: set[tuple[str, str]] = set()
|
|
@@ -319,13 +321,100 @@ def _validate_evidence(value: Any) -> set[tuple[str, str]]:
|
|
|
319
321
|
if identity in identities:
|
|
320
322
|
_fail("report.evidence contains a duplicate evidence identity")
|
|
321
323
|
identities.add(identity)
|
|
322
|
-
return identities
|
|
324
|
+
return {evidence_id for evidence_id, _locator in identities}
|
|
325
|
+
|
|
326
|
+
|
|
327
|
+
def _validate_id_list(value: Any, field: str) -> list[str]:
|
|
328
|
+
if not isinstance(value, list):
|
|
329
|
+
_fail(f"{field} must be a list")
|
|
330
|
+
result = []
|
|
331
|
+
for index, item in enumerate(value):
|
|
332
|
+
result.append(_require_text(item, f"{field}[{index}]", max_len=128))
|
|
333
|
+
if len(result) != len(set(result)):
|
|
334
|
+
_fail(f"{field} must not contain duplicates")
|
|
335
|
+
return result
|
|
336
|
+
|
|
337
|
+
|
|
338
|
+
def _validate_typed_notes(value: Any, field: str) -> None:
|
|
339
|
+
if not isinstance(value, list):
|
|
340
|
+
_fail(f"report.{field} must be a list")
|
|
341
|
+
for index, item in enumerate(value):
|
|
342
|
+
if not isinstance(item, dict):
|
|
343
|
+
_fail(f"report.{field}[{index}] must be an object")
|
|
344
|
+
_require_id(item.get("id"), f"report.{field}[{index}].id")
|
|
345
|
+
_require_text(item.get("type"), f"report.{field}[{index}].type", max_len=64)
|
|
346
|
+
_require_text(item.get("detail"), f"report.{field}[{index}].detail")
|
|
347
|
+
_validate_id_list(item.get("evidence_ids", []), f"report.{field}[{index}].evidence_ids")
|
|
348
|
+
|
|
349
|
+
|
|
350
|
+
def _validate_checkpoint(node_dir: Path, evidence_ids: set[str], scope: list[str]) -> list[str]:
|
|
351
|
+
worklog_path = node_dir / "worklog.md"
|
|
352
|
+
if worklog_path.is_symlink():
|
|
353
|
+
_fail(f"refusing symlinked worklog: {worklog_path}")
|
|
354
|
+
try:
|
|
355
|
+
raw = worklog_path.read_bytes()
|
|
356
|
+
except FileNotFoundError:
|
|
357
|
+
_fail(f"worklog does not exist: {worklog_path}")
|
|
358
|
+
except OSError as exc:
|
|
359
|
+
_fail(f"could not read worklog {worklog_path}: {exc}")
|
|
360
|
+
if len(raw) > MAX_WORKLOG_BYTES:
|
|
361
|
+
_fail(f"worklog exceeds the {MAX_WORKLOG_BYTES} byte limit: {worklog_path}")
|
|
362
|
+
_reject_obvious_secrets(raw, "worklog")
|
|
363
|
+
try:
|
|
364
|
+
text = raw.decode("utf-8")
|
|
365
|
+
except UnicodeDecodeError:
|
|
366
|
+
_fail(f"worklog is not UTF-8: {worklog_path}")
|
|
367
|
+
matches = list(CHECKPOINT_RE.finditer(text))
|
|
368
|
+
if not matches:
|
|
369
|
+
return ["missing_worklog_checkpoint"]
|
|
370
|
+
concerns: list[str] = []
|
|
371
|
+
checkpoint_ids: set[str] = set()
|
|
372
|
+
for match in matches:
|
|
373
|
+
try:
|
|
374
|
+
checkpoint = json.loads(match.group("payload"))
|
|
375
|
+
except json.JSONDecodeError:
|
|
376
|
+
concerns.append("malformed_worklog_checkpoint")
|
|
377
|
+
continue
|
|
378
|
+
if not isinstance(checkpoint, dict):
|
|
379
|
+
concerns.append("malformed_worklog_checkpoint")
|
|
380
|
+
continue
|
|
381
|
+
try:
|
|
382
|
+
checkpoint_id = _require_id(checkpoint.get("id"), "worklog checkpoint.id")
|
|
383
|
+
if checkpoint_id in checkpoint_ids:
|
|
384
|
+
concerns.append("duplicate_worklog_checkpoint")
|
|
385
|
+
checkpoint_ids.add(checkpoint_id)
|
|
386
|
+
covered_scope = _require_text_list(
|
|
387
|
+
checkpoint.get("covered_scope"),
|
|
388
|
+
"worklog checkpoint.covered_scope",
|
|
389
|
+
allow_empty=True,
|
|
390
|
+
)
|
|
391
|
+
for item in covered_scope:
|
|
392
|
+
if item not in scope:
|
|
393
|
+
concerns.append("checkpoint_scope_outside_brief")
|
|
394
|
+
checkpoint_evidence = _validate_id_list(
|
|
395
|
+
checkpoint.get("evidence_ids"), "worklog checkpoint.evidence_ids"
|
|
396
|
+
)
|
|
397
|
+
if any(item not in evidence_ids for item in checkpoint_evidence):
|
|
398
|
+
concerns.append("checkpoint_evidence_unresolved")
|
|
399
|
+
_require_text(
|
|
400
|
+
checkpoint.get("conclusion_or_blocker"),
|
|
401
|
+
"worklog checkpoint.conclusion_or_blocker",
|
|
402
|
+
)
|
|
403
|
+
_require_text_list(
|
|
404
|
+
checkpoint.get("unknowns"),
|
|
405
|
+
"worklog checkpoint.unknowns",
|
|
406
|
+
allow_empty=True,
|
|
407
|
+
)
|
|
408
|
+
_require_text(checkpoint.get("safe_resume_point"), "worklog checkpoint.safe_resume_point")
|
|
409
|
+
except ArtifactError:
|
|
410
|
+
concerns.append("malformed_worklog_checkpoint")
|
|
411
|
+
return sorted(set(concerns))
|
|
323
412
|
|
|
324
413
|
|
|
325
414
|
def _validate_report_data(
|
|
326
415
|
report: dict[str, Any],
|
|
327
416
|
brief: dict[str, Any],
|
|
328
|
-
) -> set[
|
|
417
|
+
) -> tuple[set[str], list[str]]:
|
|
329
418
|
if report.get("schema_version") != SCHEMA_VERSION:
|
|
330
419
|
_fail(f"report.schema_version must be {SCHEMA_VERSION}")
|
|
331
420
|
for field in ("task_id", "work_id", "subnode_id", "role_id"):
|
|
@@ -339,11 +428,52 @@ def _validate_report_data(
|
|
|
339
428
|
if status not in {"complete", "blocked", "incomplete", "error"}:
|
|
340
429
|
_fail("report.status must be complete, blocked, incomplete, or error")
|
|
341
430
|
evidence = _validate_evidence(report.get("evidence"))
|
|
342
|
-
|
|
343
|
-
|
|
344
|
-
|
|
431
|
+
assessment = report.get("scope_assessment")
|
|
432
|
+
if not isinstance(assessment, list) or len(assessment) != len(brief["scope"]):
|
|
433
|
+
_fail("report.scope_assessment must contain one item for every brief scope item")
|
|
434
|
+
concerns: list[str] = []
|
|
435
|
+
for index, item in enumerate(assessment):
|
|
436
|
+
if not isinstance(item, dict):
|
|
437
|
+
_fail(f"report.scope_assessment[{index}] must be an object")
|
|
438
|
+
if item.get("scope") != brief["scope"][index]:
|
|
439
|
+
_fail(f"report.scope_assessment[{index}].scope does not match brief.scope")
|
|
440
|
+
if item.get("status") not in {"covered", "inconclusive", "not-started"}:
|
|
441
|
+
_fail(f"report.scope_assessment[{index}].status is invalid")
|
|
442
|
+
_require_text(item.get("conclusion"), f"report.scope_assessment[{index}].conclusion")
|
|
443
|
+
assessment_evidence = _validate_id_list(
|
|
444
|
+
item.get("evidence_ids"), f"report.scope_assessment[{index}].evidence_ids"
|
|
445
|
+
)
|
|
446
|
+
if any(value not in evidence for value in assessment_evidence):
|
|
447
|
+
concerns.append("scope_assessment_evidence_unresolved")
|
|
448
|
+
if item["status"] != "covered":
|
|
449
|
+
concerns.append("incomplete_scope_coverage")
|
|
450
|
+
if item["status"] == "covered" and not assessment_evidence:
|
|
451
|
+
concerns.append("covered_scope_without_evidence")
|
|
452
|
+
findings = report.get("findings")
|
|
453
|
+
if not isinstance(findings, list):
|
|
454
|
+
_fail("report.findings must be a list")
|
|
455
|
+
finding_ids: set[str] = set()
|
|
456
|
+
for index, item in enumerate(findings):
|
|
457
|
+
if not isinstance(item, dict):
|
|
458
|
+
_fail(f"report.findings[{index}] must be an object")
|
|
459
|
+
finding_id = _require_id(item.get("id"), f"report.findings[{index}].id")
|
|
460
|
+
if finding_id in finding_ids:
|
|
461
|
+
_fail("report.findings contains duplicate ids")
|
|
462
|
+
finding_ids.add(finding_id)
|
|
463
|
+
_require_text(item.get("conclusion"), f"report.findings[{index}].conclusion")
|
|
464
|
+
finding_evidence = _validate_id_list(
|
|
465
|
+
item.get("evidence_ids"), f"report.findings[{index}].evidence_ids"
|
|
466
|
+
)
|
|
467
|
+
if not finding_evidence:
|
|
468
|
+
concerns.append("finding_without_evidence")
|
|
469
|
+
if any(value not in evidence for value in finding_evidence):
|
|
470
|
+
concerns.append("finding_evidence_unresolved")
|
|
471
|
+
_validate_typed_notes(report.get("uncertainties"), "uncertainties")
|
|
472
|
+
_validate_typed_notes(report.get("corrections"), "corrections")
|
|
345
473
|
if status == "complete" and not evidence:
|
|
346
474
|
_fail("a complete report requires at least one evidence item")
|
|
475
|
+
if status == "complete" and not findings:
|
|
476
|
+
_fail("a complete report requires at least one finding")
|
|
347
477
|
if status != "complete":
|
|
348
478
|
_require_text_list(
|
|
349
479
|
report.get("completed_scope"),
|
|
@@ -351,13 +481,13 @@ def _validate_report_data(
|
|
|
351
481
|
allow_empty=True,
|
|
352
482
|
)
|
|
353
483
|
_require_text(report.get("blocker"), "report.blocker")
|
|
354
|
-
return evidence
|
|
484
|
+
return evidence, sorted(set(concerns))
|
|
355
485
|
|
|
356
486
|
|
|
357
487
|
def _validate_report_file(
|
|
358
488
|
value: str,
|
|
359
489
|
repo_root: Path,
|
|
360
|
-
) -> tuple[dict[str, Any], set[
|
|
490
|
+
) -> tuple[dict[str, Any], set[str], dict[str, Any], Path, list[str]]:
|
|
361
491
|
report_path, node_dir, _task_dir = _resolve_artifact_file(value, "report.json", repo_root)
|
|
362
492
|
brief_path = node_dir / "brief.json"
|
|
363
493
|
brief, _brief_raw, _brief_path, _brief_node, _brief_task = _validate_brief_file(
|
|
@@ -368,8 +498,9 @@ def _validate_report_file(
|
|
|
368
498
|
_fail("brief.report_path does not point to the report being validated")
|
|
369
499
|
report, raw = _read_json_object(report_path, MAX_REPORT_BYTES, "report")
|
|
370
500
|
_reject_obvious_secrets(raw, "report")
|
|
371
|
-
evidence = _validate_report_data(report, brief)
|
|
372
|
-
|
|
501
|
+
evidence, concerns = _validate_report_data(report, brief)
|
|
502
|
+
concerns.extend(_validate_checkpoint(node_dir, evidence, brief["scope"]))
|
|
503
|
+
return report, evidence, brief, node_dir, sorted(set(concerns))
|
|
373
504
|
|
|
374
505
|
|
|
375
506
|
def _validate_disposition_data(
|
|
@@ -466,13 +597,20 @@ def _init(args: argparse.Namespace) -> None:
|
|
|
466
597
|
|
|
467
598
|
def _validate(args: argparse.Namespace) -> None:
|
|
468
599
|
repo_root = get_repo_root()
|
|
469
|
-
report, _evidence, _brief, _node_dir_value = _validate_report_file(args.report, repo_root)
|
|
600
|
+
report, _evidence, _brief, _node_dir_value, concerns = _validate_report_file(args.report, repo_root)
|
|
601
|
+
if concerns:
|
|
602
|
+
print(json.dumps({
|
|
603
|
+
"status": "review_concern",
|
|
604
|
+
"subnode_id": report["subnode_id"],
|
|
605
|
+
"concerns": concerns,
|
|
606
|
+
}))
|
|
607
|
+
return
|
|
470
608
|
print(f"Validated pending-review report: {report['subnode_id']}")
|
|
471
609
|
|
|
472
610
|
|
|
473
611
|
def _disposition(args: argparse.Namespace) -> None:
|
|
474
612
|
repo_root = get_repo_root()
|
|
475
|
-
report, _evidence, brief, node_dir = _validate_report_file(args.report, repo_root)
|
|
613
|
+
report, _evidence, brief, node_dir, _concerns = _validate_report_file(args.report, repo_root)
|
|
476
614
|
disposition_path = node_dir / "disposition.json"
|
|
477
615
|
if disposition_path.exists() or disposition_path.is_symlink():
|
|
478
616
|
_fail(f"disposition already exists and cannot be replaced: {disposition_path}")
|
|
@@ -514,10 +652,10 @@ def _validate_counter(args: argparse.Namespace) -> None:
|
|
|
514
652
|
primary_dir = repo_root / primary_dir
|
|
515
653
|
if not counter_dir.is_absolute():
|
|
516
654
|
counter_dir = repo_root / counter_dir
|
|
517
|
-
primary_report, primary_evidence, primary_brief, primary_node = _validate_report_file(
|
|
655
|
+
primary_report, primary_evidence, primary_brief, primary_node, _primary_concerns = _validate_report_file(
|
|
518
656
|
str(primary_dir / "report.json"), repo_root
|
|
519
657
|
)
|
|
520
|
-
counter_report, counter_evidence, counter_brief, counter_node = _validate_report_file(
|
|
658
|
+
counter_report, counter_evidence, counter_brief, counter_node, _counter_concerns = _validate_report_file(
|
|
521
659
|
str(counter_dir / "report.json"), repo_root
|
|
522
660
|
)
|
|
523
661
|
if primary_node == counter_node:
|
|
@@ -167,6 +167,8 @@ Phase 3: Finish → verify, update spec, commit, and wrap up
|
|
|
167
167
|
|
|
168
168
|
An analysis-only task is eligible only when `task.json.meta.delivery_mode = "analysis_only"` exactly and its `prd.md` names the evidence deliverable plus a no-change boundary for product source, runtime configuration, deployment, credentials, and external systems. Task creation consent authorizes that bounded evidence work, not protected-target changes.
|
|
169
169
|
|
|
170
|
+
The exception is invalid when the work still needs a material user decision, design or implementation plan, cross-owner coordination, security or deployment change, release or credential action, or a protected downstream task. "Research" and a deferred source edit do not override this classification; use normal complex planning when any condition applies.
|
|
171
|
+
|
|
170
172
|
Keep an eligible analysis-only task in `planning`: write its research, audit, or design evidence; verify its acceptance criteria and boundary; commit task artifacts; then archive directly. Do not run `task.py start`, configure implementation context, or wait for a second implementation approval. If the evidence recommends a protected-target change, record the recommendation and create a separate change-bearing task before doing it.
|
|
171
173
|
|
|
172
174
|
### Planning Artifacts
|
|
@@ -177,6 +179,8 @@ Keep an eligible analysis-only task in `planning`: write its research, audit, or
|
|
|
177
179
|
- `implement.jsonl` / `check.jsonl` — spec and research manifests for sub-agent context. They do not replace `implement.md`.
|
|
178
180
|
- Lightweight tasks may be PRD-only. Complex tasks must have `prd.md`, `design.md`, and `implement.md` before `task.py start`.
|
|
179
181
|
|
|
182
|
+
For read-heavy research, audit, review, or investigation, split work that cannot persist one useful conclusion in the current bounded session into evidence units. Each unit records one question or scope, minimal evidence, destination, stop condition, conclusion or blocker, unknowns, and recovery point before the next unit begins. Size for one normal context window, without an exact token or time promise; routine navigation needs no record.
|
|
183
|
+
|
|
180
184
|
### Parent / Child Task Trees
|
|
181
185
|
|
|
182
186
|
Use a parent task when one user request contains several independently verifiable deliverables. The parent task owns the source requirement set, the task map, cross-child acceptance criteria, and final integration review; it normally should not be the implementation target unless it also has direct work.
|
|
@@ -230,7 +234,7 @@ Preserve existing task fields and artifacts. If the correct status cannot be det
|
|
|
230
234
|
[workflow-state:planning]
|
|
231
235
|
Load `trellis-brainstorm`; stay in planning.
|
|
232
236
|
If `task.json.meta.delivery_mode = "analysis_only"` exactly, complete the declared evidence work now. Do not wait for a start review or run `task.py start`; when the PRD boundary and acceptance evidence pass, commit task artifacts and archive directly. A protected-target recommendation requires a separate change-bearing task.
|
|
233
|
-
Lightweight: `prd.md` can be enough. Complex: finish `prd.md`, `design.md`, and `implement.md`;
|
|
237
|
+
Lightweight: `prd.md` can be enough. Complex: finish `prd.md`, `design.md`, and `implement.md`; run the Planning Seal closure pass before asking for review. If `decision-needed` items or an unsealed decision graph remain, load `pennix-decision-gates`, batch only independent frontier questions, and stay in planning. Answers returned by the current continuation must be persisted and fed back into the same planning loop.
|
|
234
238
|
Multi-deliverable scope: consider a parent task plus independently verifiable child tasks; dependencies must be written in child artifacts, not implied by tree position.
|
|
235
239
|
Sub-agent mode: curate `implement.jsonl` and `check.jsonl` as spec/research manifests before start.
|
|
236
240
|
[/workflow-state:planning]
|
|
@@ -244,7 +248,7 @@ Sub-agent mode: curate `implement.jsonl` and `check.jsonl` as spec/research mani
|
|
|
244
248
|
[workflow-state:planning-inline]
|
|
245
249
|
Load `trellis-brainstorm`; stay in planning.
|
|
246
250
|
If `task.json.meta.delivery_mode = "analysis_only"` exactly, complete the declared evidence work now. Do not wait for a start review or run `task.py start`; when the PRD boundary and acceptance evidence pass, commit task artifacts and archive directly. A protected-target recommendation requires a separate change-bearing task.
|
|
247
|
-
Lightweight: `prd.md` can be enough. Complex: finish `prd.md`, `design.md`, and `implement.md`;
|
|
251
|
+
Lightweight: `prd.md` can be enough. Complex: finish `prd.md`, `design.md`, and `implement.md`; run the Planning Seal closure pass before asking for review. If `decision-needed` items or an unsealed decision graph remain, load `pennix-decision-gates`, batch only independent frontier questions, and stay in planning. Answers returned by the current continuation must be persisted and fed back into the same planning loop.
|
|
248
252
|
Multi-deliverable scope: consider a parent task plus independently verifiable child tasks; dependencies must be written in child artifacts, not implied by tree position.
|
|
249
253
|
Inline mode: skip jsonl curation; Phase 2 reads artifacts/specs via `trellis-before-dev`.
|
|
250
254
|
[/workflow-state:planning-inline]
|
|
@@ -389,6 +393,8 @@ The brainstorm skill will guide you to:
|
|
|
389
393
|
- Split large scopes into a parent task plus child tasks when the deliverables can be verified independently
|
|
390
394
|
- Keep `prd.md` focused on requirements and acceptance criteria
|
|
391
395
|
- For complex tasks, produce `design.md` and `implement.md` before implementation starts
|
|
396
|
+
- For read-heavy work, split long investigation into evidence units and persist each unit's conclusion or recovery point before continuing
|
|
397
|
+
- Before review or `task.py start`, run the Planning Seal closure pass across all task artifacts and lock targets, branches, dependencies, release, validation, rollback, dynamic-fact handling, and every material decision
|
|
392
398
|
|
|
393
399
|
When considering a parent/child split:
|
|
394
400
|
- Use a parent task when one request contains several independently verifiable deliverables.
|
|
@@ -490,7 +496,7 @@ Skip this step. Context is loaded directly by the `trellis-before-dev` skill in
|
|
|
490
496
|
|
|
491
497
|
This step applies only to change-bearing tasks. An eligible analysis-only task stays in `planning` and, after its evidence work is complete, continues directly to Phase 3.3 without running `task.py start`.
|
|
492
498
|
|
|
493
|
-
After artifact review, flip the task status to `in_progress`:
|
|
499
|
+
After the Planning Seal closure pass and artifact review, flip the task status to `in_progress`:
|
|
494
500
|
|
|
495
501
|
```bash
|
|
496
502
|
python3 ./.trellis/scripts/task.py start <task-dir>
|
|
@@ -513,6 +519,7 @@ If `task.py start` errors with a session-identity message (no context key from h
|
|
|
513
519
|
| `research/` has artifacts (complex tasks) | recommended |
|
|
514
520
|
| `design.md` exists (complex tasks) | ✅ |
|
|
515
521
|
| `implement.md` exists (complex tasks) | ✅ |
|
|
522
|
+
| Planning Seal closure pass recorded; static decisions and implementation paths are locked | ✅ |
|
|
516
523
|
|
|
517
524
|
[Claude Code, Cursor, OpenCode, codex-sub-agent, Kiro, Gemini, Qoder, CodeBuddy, Copilot, Droid, Pi, Oh My Pi, ZCode, Snow, Reasonix, Trae, Grok, Kimi Code, DeepSeek Harness]
|
|
518
525
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@pennixrv/trellis",
|
|
3
|
-
"version": "0.7.0-beta.
|
|
3
|
+
"version": "0.7.0-beta.9",
|
|
4
4
|
"description": "AI capabilities grow like ivy — Trellis provides the structure to guide them along a disciplined path",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "./dist/index.js",
|
|
@@ -34,7 +34,7 @@
|
|
|
34
34
|
"inquirer": "^9.3.7",
|
|
35
35
|
"undici": "^6.21.0",
|
|
36
36
|
"zod": "^4.4.2",
|
|
37
|
-
"@pennixrv/trellis-core": "0.7.0-beta.
|
|
37
|
+
"@pennixrv/trellis-core": "0.7.0-beta.9"
|
|
38
38
|
},
|
|
39
39
|
"devDependencies": {
|
|
40
40
|
"@eslint/js": "^9.18.0",
|