@garygentry/feature-forge 0.2.1 → 0.2.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (46) hide show
  1. package/README.md +1 -1
  2. package/adapters/claude/references/forge-config-schema.json +20 -3
  3. package/adapters/claude/references/shared-conventions.md +3 -0
  4. package/adapters/claude/scripts/forge-init.sh +7 -1
  5. package/adapters/claude/scripts/forge-session.py +491 -0
  6. package/adapters/claude/skills/forge/SKILL.md +43 -1
  7. package/adapters/claude/skills/forge-5-loop/SKILL.md +2 -1
  8. package/adapters/claude/skills/forge-5-loop/references/runner-contract.md +41 -15
  9. package/adapters/claude/skills/forge-init/SKILL.md +3 -0
  10. package/adapters/codex/references/forge-config-schema.json +20 -3
  11. package/adapters/codex/references/shared-conventions.md +3 -0
  12. package/adapters/codex/scripts/forge-init.sh +7 -1
  13. package/adapters/codex/scripts/forge-session.py +491 -0
  14. package/adapters/codex/skills/forge/SKILL.md +43 -1
  15. package/adapters/codex/skills/forge-5-loop/SKILL.md +2 -1
  16. package/adapters/codex/skills/forge-5-loop/references/runner-contract.md +41 -15
  17. package/adapters/codex/skills/forge-init/SKILL.md +3 -0
  18. package/adapters/copilot/references/forge-config-schema.json +20 -3
  19. package/adapters/copilot/references/shared-conventions.md +3 -0
  20. package/adapters/copilot/scripts/forge-init.sh +7 -1
  21. package/adapters/copilot/scripts/forge-session.py +491 -0
  22. package/adapters/copilot/skills/forge/forge.md +43 -1
  23. package/adapters/copilot/skills/forge-5-loop/forge-5-loop.md +2 -1
  24. package/adapters/copilot/skills/forge-5-loop/references/runner-contract.md +41 -15
  25. package/adapters/copilot/skills/forge-init/forge-init.md +3 -0
  26. package/adapters/cursor/references/forge-config-schema.json +20 -3
  27. package/adapters/cursor/references/shared-conventions.md +3 -0
  28. package/adapters/cursor/scripts/forge-init.sh +7 -1
  29. package/adapters/cursor/scripts/forge-session.py +491 -0
  30. package/adapters/cursor/skills/forge/forge.mdc +43 -1
  31. package/adapters/cursor/skills/forge-5-loop/forge-5-loop.mdc +2 -1
  32. package/adapters/cursor/skills/forge-5-loop/references/runner-contract.md +41 -15
  33. package/adapters/cursor/skills/forge-init/forge-init.mdc +3 -0
  34. package/adapters/gemini/references/forge-config-schema.json +20 -3
  35. package/adapters/gemini/references/shared-conventions.md +3 -0
  36. package/adapters/gemini/scripts/forge-init.sh +7 -1
  37. package/adapters/gemini/scripts/forge-session.py +491 -0
  38. package/adapters/gemini/skills/forge/forge.md +43 -1
  39. package/adapters/gemini/skills/forge-5-loop/forge-5-loop.md +2 -1
  40. package/adapters/gemini/skills/forge-5-loop/references/runner-contract.md +41 -15
  41. package/adapters/gemini/skills/forge-init/forge-init.md +3 -0
  42. package/dist/manifest.d.ts +1 -1
  43. package/dist/rauf.d.ts +4 -4
  44. package/dist/rauf.js +3 -3
  45. package/dist/types.d.ts +1 -1
  46. package/package.json +1 -1
package/README.md CHANGED
@@ -71,7 +71,7 @@ Claude Code users can alternatively install via the plugin marketplace:
71
71
 
72
72
  The default loop runner is [**rauf**](https://github.com/garygentry/rauf), published as
73
73
  [`@garygentry/rauf`](https://www.npmjs.com/package/@garygentry/rauf). The installer runs a
74
- read-only resolvability preflight on the pin (`@garygentry/rauf@0.8.1`) and records it; pass
74
+ read-only resolvability preflight on the pin (`@garygentry/rauf@0.11.0`) and records it; pass
75
75
  `--skip-rauf` to defer the check (e.g. offline installs). Install the rauf CLI itself with
76
76
  `npx @garygentry/rauf` or its
77
77
  [binary script](https://github.com/garygentry/rauf#install). See the
@@ -56,6 +56,23 @@
56
56
  "minimum": 1,
57
57
  "description": "Multiplier applied to pending backlog item count to calculate loop iterations. Higher values allow more retries. Default: 1.5 (e.g., 10 items = 15 iterations)."
58
58
  },
59
+ "autoInvokeNextStage": {
60
+ "type": "boolean",
61
+ "default": true,
62
+ "description": "When true (default), the /feature-forge:forge navigator auto-invokes the next pipeline stage via the Skill tool after the user confirms it, instead of only printing the command to copy. Set false to keep the old copy-paste behavior (the navigator suggests the command but never launches it). Ignored on non-Claude hosts, which always fall back to printing the command."
63
+ },
64
+ "contextWindowTokens": {
65
+ "type": ["integer", "null"],
66
+ "default": null,
67
+ "description": "Context window size (tokens) used by the navigator's context-usage check to compute how full the current session is. Null (default) lets the helper infer from the session model and fall back to 200000 — and if observed usage already exceeds 200000 it auto-bumps the assumed window to 1000000 (proof a 1M-beta window is active). Set this explicitly to your model's window (e.g. 1000000 for a 1M-context model) for accurate percentages below 200000 too, since 1M cannot be detected from the transcript until usage crosses 200000."
68
+ },
69
+ "contextWarnThreshold": {
70
+ "type": "number",
71
+ "default": 0.7,
72
+ "minimum": 0,
73
+ "maximum": 1,
74
+ "description": "Fraction of the context window (0-1) past which the navigator recommends starting the next stage in a clean session rather than continuing. Default: 0.7."
75
+ },
59
76
  "workspaces": {
60
77
  "type": "array",
61
78
  "description": "Monorepo members. Absent for single-package projects.",
@@ -94,7 +111,7 @@
94
111
  "eventStreamCommand": {
95
112
  "type": "string",
96
113
  "default": "{bin} loop run . --backlog {backlogDir} --iterations {iterations} --ndjson",
97
- "description": "PREFERRED launch command for forge-5. Same as runCommand but emits one machine-readable JSON event per stdout line (NDJSON): item_completed / item_blocked / needs_human / signal_parsed / loop_completed / loop_error / loop_cancelled / llm_stuck_warning, each with {type, timestamp, projectPath} plus payload (a circuit-breaker halt surfaces as loop_error). forge-5 redirects this stdout to {backlogDir}/{stateDir}/events.ndjson and arms a Monitor on it for live, structured supervision. If a runner cannot emit NDJSON, omit this field — forge-5 falls back to runCommand + tailing the human log."
114
+ "description": "Stdout NDJSON launch command for a runner that does NOT persist its own event file — same as runCommand but emits one machine-readable JSON event per stdout line: item_completed / item_blocked / needs_human / signal_parsed / loop_completed / loop_error / loop_cancelled / llm_stuck_warning, each with {type, timestamp, projectPath} plus payload (a circuit-breaker halt surfaces as loop_error). NOTE: rauf (the default runner) ALREADY persists {stateDir}/events.ndjson natively and rotates it per run, so forge-5 launches the plain runCommand and monitors that native file — it does NOT use this field, and must NOT redirect --ndjson into {stateDir} (redundant, and it collides with the runner's own writer / archive rotation). This field is only for a stdout-only runner with no native event file; forge-5 then redirects its stdout to a file OUTSIDE {stateDir} and monitors that. Omit it entirely for a runner that self-persists or cannot emit NDJSON."
98
115
  },
99
116
  "validateCommand": {
100
117
  "type": "string",
@@ -173,8 +190,8 @@
173
190
  },
174
191
  "installHint": {
175
192
  "type": "string",
176
- "default": "Provision rauf for a multi-agent setup with the cross-agent installer: `npx @garygentry/feature-forge install` (records the pinned @garygentry/rauf@0.8.1 default). Or install/upgrade just the rauf CLI: `npx @garygentry/rauf@0.8.1 --version`, or `curl -fsSL https://raw.githubusercontent.com/garygentry/rauf/main/scripts/install-binary.sh | bash`.",
177
- "description": "Shown when the runner BINARY is missing or too old (version gate fails, minRunnerVersion floor) — how to obtain/upgrade the CLI itself. Names two distinct binary-provisioning paths: (1) the cross-agent installer (`npx @garygentry/feature-forge install`, the multi-agent provisioning path that pins @garygentry/rauf@0.8.1), and (2) the direct rauf-CLI install/upgrade one-liner. Distinct from setupHint (which installs per-project artifacts); a version-gate failure is ALWAYS this hint, never setupHint."
193
+ "default": "Provision rauf for a multi-agent setup with the cross-agent installer: `npx @garygentry/feature-forge install` (records the pinned @garygentry/rauf@0.11.0 default). Or install/upgrade just the rauf CLI: `npx @garygentry/rauf@0.11.0 --version`, or `curl -fsSL https://raw.githubusercontent.com/garygentry/rauf/main/scripts/install-binary.sh | bash`.",
194
+ "description": "Shown when the runner BINARY is missing or too old (version gate fails, minRunnerVersion floor) — how to obtain/upgrade the CLI itself. Names two distinct binary-provisioning paths: (1) the cross-agent installer (`npx @garygentry/feature-forge install`, the multi-agent provisioning path that pins @garygentry/rauf@0.11.0), and (2) the direct rauf-CLI install/upgrade one-liner. Distinct from setupHint (which installs per-project artifacts); a version-gate failure is ALWAYS this hint, never setupHint."
178
195
  },
179
196
  "schemaVersion": {
180
197
  "type": "string",
@@ -68,6 +68,9 @@ Extract these config values (use defaults if not present):
68
68
  - `branchPerFeature` (default: true)
69
69
  - `branchPrefix` (default: `forge/`)
70
70
  - `loopIterationMultiplier` (default: `1.5`)
71
+ - `autoInvokeNextStage` (default: `true` — the `/feature-forge:forge` navigator auto-invokes the next stage via the `Skill` tool after the user confirms; `false` keeps copy-paste behavior. Navigator-only.)
72
+ - `contextWindowTokens` (default: `null` — context window used by the navigator's context-usage check; `null` infers from the session model and falls back to 200000. Set to the model's window, e.g. `1000000` on a 1M model. Navigator-only.)
73
+ - `contextWarnThreshold` (default: `0.7` — fraction of the window past which the navigator recommends a clean session. Navigator-only.)
71
74
  - `loopRunner` (optional object — the loop runner to drive; **defaults to rauf** when absent, with every command templated. See `references/forge-config-schema.json` and `references/ralph-loop-contract.md`.)
72
75
 
73
76
  ## Feature Directory Resolution
@@ -21,7 +21,10 @@ cat > "$CONFIG_FILE" << 'EOF'
21
21
  "stack": null,
22
22
  "typeCheckCommand": null,
23
23
  "testCommand": null,
24
- "loopIterationMultiplier": 1.5
24
+ "loopIterationMultiplier": 1.5,
25
+ "autoInvokeNextStage": true,
26
+ "contextWindowTokens": null,
27
+ "contextWarnThreshold": 0.7
25
28
  }
26
29
  EOF
27
30
 
@@ -37,6 +40,9 @@ echo " stack: null (auto-detected during forge-2-tech)"
37
40
  echo " typeCheckCommand: null (auto-detected during forge-2-tech)"
38
41
  echo " testCommand: null (auto-detected during forge-2-tech)"
39
42
  echo " loopIterationMultiplier: 1.5 (multiplier for loop iterations)"
43
+ echo " autoInvokeNextStage: true (navigator auto-starts the next stage after you confirm)"
44
+ echo " contextWindowTokens: null (infer; set to 1000000 on a 1M-context model)"
45
+ echo " contextWarnThreshold: 0.7 (suggest a clean session past this fraction of the window)"
40
46
  echo ""
41
47
  echo "The loop runner defaults to rauf. To target a different ralph-style runner,"
42
48
  echo "add a \"loopRunner\" block (see references/forge-config-schema.json)."
@@ -0,0 +1,491 @@
1
+ #!/usr/bin/env python3
2
+ """Session-aware navigation helpers for the feature-forge pipeline navigator.
3
+
4
+ Two read-only subcommands that drive the usability features of the `/forge`
5
+ root navigator:
6
+
7
+ python3 forge-session.py rank-features [--specs-dir DIR] [--json]
8
+ python3 forge-session.py context-usage [--config FILE] [--window N] \
9
+ [--threshold F] [--json]
10
+
11
+ `rank-features` scans the specs tree for feature-shaped directories (those that
12
+ directly contain a `.pipeline-state.json`, in both the flat
13
+ `{specsDir}/{feature}/` and nested `{specsDir}/{epic}/{feature}/` layouts) and
14
+ reports the **active** ones ordered by `updatedAt` descending, so the navigator
15
+ can offer the most-recently-touched feature as the recency default. Each row
16
+ carries the next actionable stage + its slash command, derived from the single
17
+ ordered stage map below.
18
+
19
+ `context-usage` reads the live Claude Code session transcript (the most-recently
20
+ modified `*.jsonl` under `~/.claude/projects/<cwd-slug>/`), sums the last
21
+ assistant message's token usage, and compares it to the context window so the
22
+ navigator can recommend a clean session before the next stage. It is best-effort
23
+ and degrades gracefully: when no transcript or usage is found (a non-Claude host,
24
+ or a fresh session) it reports `{"available": false}` and still exits 0, so the
25
+ caller simply omits the context advice.
26
+
27
+ 3.10 baseline, Google-style docstrings, full type annotations, stdlib only —
28
+ matching the conventions of `scripts/epic-manifest.py`.
29
+
30
+ Exit codes:
31
+ 0 = ok (including an empty feature list or unavailable context usage)
32
+ 2 = usage error or unreadable I/O
33
+ """
34
+
35
+ from __future__ import annotations
36
+
37
+ import argparse
38
+ import json
39
+ import sys
40
+ from datetime import datetime
41
+ from pathlib import Path
42
+ from typing import Final, TypedDict
43
+
44
+
45
+ # --------------------------------------------------------------------------- #
46
+ # Constants
47
+ # --------------------------------------------------------------------------- #
48
+
49
+ #: A directory is "feature-shaped" iff it directly contains this file.
50
+ PIPELINE_STATE_FILENAME: Final = ".pipeline-state.json"
51
+ #: Epic roots hold this (and no .pipeline-state.json) — never a feature.
52
+ MANIFEST_FILENAME: Final = "epic-manifest.json"
53
+
54
+ #: The ordered production stages. This is the ONE place stage order lives.
55
+ PRODUCTION_STAGES: Final[tuple[str, ...]] = (
56
+ "forge-1-prd",
57
+ "forge-2-tech",
58
+ "forge-3-specs",
59
+ "forge-4-backlog",
60
+ "forge-5-loop",
61
+ "forge-6-docs",
62
+ )
63
+
64
+ #: Production stage -> the verify token its findings file uses, and the
65
+ #: `forge-verify-<token>` key its state lives under. forge-6-docs has no verify.
66
+ VERIFY_TOKEN_BY_STAGE: Final[dict[str, str]] = {
67
+ "forge-1-prd": "prd",
68
+ "forge-2-tech": "tech",
69
+ "forge-3-specs": "specs",
70
+ "forge-4-backlog": "backlog",
71
+ "forge-5-loop": "impl",
72
+ }
73
+
74
+ #: A production stage status that counts as "done" for next-stage selection.
75
+ _DONE_STATUS: Final = "complete"
76
+ #: Verify statuses that count as "resolved" (no outstanding verify needed).
77
+ _VERIFY_RESOLVED: Final = frozenset({"passed", "findings-applied", "skipped"})
78
+
79
+ #: Default context window when the model can't be inferred and config is silent.
80
+ _DEFAULT_WINDOW: Final = 200_000
81
+ #: Window for 1M-context models (model id carries a `[1m]` / `-1m` marker).
82
+ _WIDE_WINDOW: Final = 1_000_000
83
+ #: Default fraction of the window past which a clean session is recommended.
84
+ _DEFAULT_THRESHOLD: Final = 0.7
85
+
86
+
87
+ # --------------------------------------------------------------------------- #
88
+ # Types
89
+ # --------------------------------------------------------------------------- #
90
+
91
+
92
+ class FeatureRow(TypedDict):
93
+ """One active feature, ranked by recency, with its next actionable step."""
94
+
95
+ name: str
96
+ epic: str | None
97
+ currentStage: str
98
+ branch: str | None
99
+ updatedAt: str | None
100
+ complete: bool
101
+ nextStage: str | None
102
+ nextCommand: str | None
103
+ verifyPending: bool
104
+ verifyCommand: str | None
105
+
106
+
107
+ class UsageError(Exception):
108
+ """A usage or I/O failure that must exit 2."""
109
+
110
+
111
+ # --------------------------------------------------------------------------- #
112
+ # Feature scanning & ranking
113
+ # --------------------------------------------------------------------------- #
114
+
115
+
116
+ def _read_state(state_path: Path) -> dict:
117
+ """Read a `.pipeline-state.json`, tolerating missing/corrupt files.
118
+
119
+ A missing, unreadable, or unparseable state downgrades to ``{}`` rather than
120
+ crashing the scan — the navigator simply treats that feature as not-started.
121
+ """
122
+ try:
123
+ parsed = json.loads(state_path.read_text(encoding="utf-8"))
124
+ except (OSError, json.JSONDecodeError):
125
+ return {}
126
+ return parsed if isinstance(parsed, dict) else {}
127
+
128
+
129
+ def _scan_features(specs_dir: Path) -> list[tuple[str, str | None, dict]]:
130
+ """Find every feature-shaped dir under the specs tree (flat + nested).
131
+
132
+ Descends exactly one level below each top-level dir (never deeper), matching
133
+ ``epic-manifest.py``'s feature-shaped-dir bound.
134
+
135
+ Args:
136
+ specs_dir: The configured specs directory.
137
+
138
+ Returns:
139
+ A list of ``(feature_name, epic_name_or_None, state_dict)`` tuples. The
140
+ epic name is the parent dir name for a nested member, ``None`` for a flat
141
+ feature.
142
+ """
143
+ if not specs_dir.is_dir():
144
+ return []
145
+ out: list[tuple[str, str | None, dict]] = []
146
+ for top in sorted(p for p in specs_dir.iterdir() if p.is_dir()):
147
+ flat_state = top / PIPELINE_STATE_FILENAME
148
+ if flat_state.is_file():
149
+ out.append((top.name, None, _read_state(flat_state)))
150
+ # Descend one level for nested epic members (skip the epic root itself).
151
+ for child in sorted(p for p in top.iterdir() if p.is_dir()):
152
+ nested_state = child / PIPELINE_STATE_FILENAME
153
+ if nested_state.is_file():
154
+ out.append((child.name, top.name, _read_state(nested_state)))
155
+ return out
156
+
157
+
158
+ def _stage_status(state: dict, stage: str) -> str | None:
159
+ """Return the recorded status of a stage, or None if absent."""
160
+ stages = state.get("stages")
161
+ if not isinstance(stages, dict):
162
+ return None
163
+ entry = stages.get(stage)
164
+ if not isinstance(entry, dict):
165
+ return None
166
+ status = entry.get("status")
167
+ return status if isinstance(status, str) else None
168
+
169
+
170
+ def next_stage(state: dict) -> str | None:
171
+ """Return the first production stage that is not yet complete (the next step).
172
+
173
+ Walks ``PRODUCTION_STAGES`` in order and returns the first whose recorded
174
+ status is not ``complete`` (a missing/pending/in-progress/stale stage all
175
+ count as "not done"). Returns ``None`` when every production stage is
176
+ complete (nothing left to run).
177
+ """
178
+ for stage in PRODUCTION_STAGES:
179
+ if _stage_status(state, stage) != _DONE_STATUS:
180
+ return stage
181
+ return None
182
+
183
+
184
+ def pending_verify(state: dict) -> str | None:
185
+ """Return the production stage whose verify is outstanding, if any.
186
+
187
+ The most recently completed production stage whose corresponding
188
+ ``forge-verify-*`` is neither resolved (passed/findings-applied/skipped) nor
189
+ already run. Surfaced so the navigator can offer "verify before continuing"
190
+ as an alternative to advancing. Returns ``None`` when nothing needs verify.
191
+ """
192
+ for stage in reversed(PRODUCTION_STAGES):
193
+ if _stage_status(state, stage) != _DONE_STATUS:
194
+ continue
195
+ token = VERIFY_TOKEN_BY_STAGE.get(stage)
196
+ if token is None:
197
+ continue # forge-6-docs has no verify step
198
+ verify_status = _stage_status(state, f"forge-verify-{token}")
199
+ if verify_status not in _VERIFY_RESOLVED:
200
+ return stage
201
+ return None # most-recent complete stage is already verified
202
+ return None
203
+
204
+
205
+ def _parse_ts(value: str | None) -> datetime | None:
206
+ """Parse an ISO-8601 timestamp (tolerating a trailing 'Z'), else None."""
207
+ if not isinstance(value, str):
208
+ return None
209
+ try:
210
+ return datetime.fromisoformat(value.replace("Z", "+00:00"))
211
+ except ValueError:
212
+ return None
213
+
214
+
215
+ def build_rows(specs_dir: Path) -> list[FeatureRow]:
216
+ """Build the recency-ranked active-feature rows (the rank-features payload).
217
+
218
+ Active features (``pipelineStatus == "active"``, the default when absent) are
219
+ sorted by ``updatedAt`` descending — most recently touched first — so the
220
+ navigator's recency default is row 0.
221
+ """
222
+ rows: list[FeatureRow] = []
223
+ for name, epic, state in _scan_features(specs_dir):
224
+ status = state.get("pipelineStatus", "active")
225
+ if status != "active":
226
+ continue
227
+ nxt = next_stage(state)
228
+ verify_stage = pending_verify(state)
229
+ branch = state.get("branch")
230
+ updated = state.get("updatedAt")
231
+ rows.append({
232
+ "name": name,
233
+ "epic": epic,
234
+ "currentStage": state.get("currentStage") or (nxt or "complete"),
235
+ "branch": branch if isinstance(branch, str) else None,
236
+ "updatedAt": updated if isinstance(updated, str) else None,
237
+ "complete": nxt is None,
238
+ "nextStage": nxt,
239
+ "nextCommand": f"/feature-forge:{nxt} {name}" if nxt else None,
240
+ "verifyPending": verify_stage is not None,
241
+ "verifyCommand": f"/feature-forge:forge-verify {name}" if verify_stage else None,
242
+ })
243
+ # Sort by updatedAt desc; rows without a parseable timestamp sort last.
244
+ rows.sort(
245
+ key=lambda r: (_parse_ts(r["updatedAt"]) or datetime.min.replace(tzinfo=None)),
246
+ reverse=True,
247
+ )
248
+ return rows
249
+
250
+
251
+ def _counts(specs_dir: Path) -> dict[str, int]:
252
+ """Tally active/paused/abandoned pipelines across the specs tree."""
253
+ tally = {"active": 0, "paused": 0, "abandoned": 0}
254
+ for _name, _epic, state in _scan_features(specs_dir):
255
+ status = state.get("pipelineStatus", "active")
256
+ if status in tally:
257
+ tally[status] += 1
258
+ return tally
259
+
260
+
261
+ # --------------------------------------------------------------------------- #
262
+ # Context-window usage
263
+ # --------------------------------------------------------------------------- #
264
+
265
+
266
+ def _cwd_slug(cwd: Path) -> str:
267
+ """Map a working directory to its Claude Code project-dir slug.
268
+
269
+ Claude Code names the per-project transcript dir by replacing path
270
+ separators (and dots) in the absolute cwd with hyphens, e.g.
271
+ ``/home/u/proj`` -> ``-home-u-proj``.
272
+ """
273
+ return str(cwd.resolve()).replace("/", "-").replace(".", "-")
274
+
275
+
276
+ def _latest_transcript(cwd: Path) -> Path | None:
277
+ """Return the most-recently-modified transcript JSONL for this cwd, if any."""
278
+ project_dir = Path.home() / ".claude" / "projects" / _cwd_slug(cwd)
279
+ if not project_dir.is_dir():
280
+ return None
281
+ transcripts = [p for p in project_dir.glob("*.jsonl") if p.is_file()]
282
+ if not transcripts:
283
+ return None
284
+ return max(transcripts, key=lambda p: p.stat().st_mtime)
285
+
286
+
287
+ def _last_usage(transcript: Path) -> tuple[int, str | None] | None:
288
+ """Scan a transcript from the end for the last `usage` record.
289
+
290
+ Returns ``(token_total, model_id)`` where the total sums
291
+ ``input_tokens + cache_creation_input_tokens + cache_read_input_tokens +
292
+ output_tokens`` of the most recent message carrying a usage object — i.e. the
293
+ current context occupancy. Returns ``None`` if no usable record is found.
294
+ """
295
+ try:
296
+ lines = transcript.read_text(encoding="utf-8").splitlines()
297
+ except OSError:
298
+ return None
299
+ for line in reversed(lines):
300
+ line = line.strip()
301
+ if not line or '"usage"' not in line:
302
+ continue
303
+ try:
304
+ record = json.loads(line)
305
+ except json.JSONDecodeError:
306
+ continue
307
+ message = record.get("message")
308
+ usage = message.get("usage") if isinstance(message, dict) else record.get("usage")
309
+ if not isinstance(usage, dict):
310
+ continue
311
+ total = (
312
+ int(usage.get("input_tokens", 0) or 0)
313
+ + int(usage.get("cache_creation_input_tokens", 0) or 0)
314
+ + int(usage.get("cache_read_input_tokens", 0) or 0)
315
+ + int(usage.get("output_tokens", 0) or 0)
316
+ )
317
+ if total <= 0:
318
+ continue
319
+ model = message.get("model") if isinstance(message, dict) else record.get("model")
320
+ return total, (model if isinstance(model, str) else None)
321
+ return None
322
+
323
+
324
+ def _infer_window(model: str | None) -> int:
325
+ """Infer the context window from a model id (1M-context markers -> wide)."""
326
+ if model and ("[1m]" in model.lower() or "-1m" in model.lower()):
327
+ return _WIDE_WINDOW
328
+ return _DEFAULT_WINDOW
329
+
330
+
331
+ def _config_value(config_path: Path, key: str):
332
+ """Read a single key from forge.config.json, or None if absent/unreadable."""
333
+ try:
334
+ config = json.loads(config_path.read_text(encoding="utf-8"))
335
+ except (OSError, json.JSONDecodeError):
336
+ return None
337
+ return config.get(key) if isinstance(config, dict) else None
338
+
339
+
340
+ def context_usage(
341
+ config_path: Path,
342
+ window_override: int | None,
343
+ threshold_override: float | None,
344
+ ) -> dict:
345
+ """Compute live context-window occupancy for the current session.
346
+
347
+ Window precedence: ``--window`` > config ``contextWindowTokens`` > inferred
348
+ from the transcript's model id > ``_DEFAULT_WINDOW``. When inferring (no
349
+ override, no config) and the observed token total already exceeds the default
350
+ window, the window is auto-bumped to ``_WIDE_WINDOW`` — observed tokens above
351
+ 200k prove a wider (1M-beta) window is active, so this corrects the reading
352
+ without ever under-reporting a genuine 200k session. Threshold precedence:
353
+ ``--threshold`` > config ``contextWarnThreshold`` > ``_DEFAULT_THRESHOLD``.
354
+
355
+ Returns a dict with ``available: True`` and ``{tokens, windowTokens, pct,
356
+ overThreshold, recommendation, model}`` when usage is found, or
357
+ ``{available: False, reason}`` otherwise. Never raises for a missing
358
+ transcript — that is the expected non-Claude / fresh-session path.
359
+ """
360
+ threshold = threshold_override
361
+ if threshold is None:
362
+ cfg_threshold = _config_value(config_path, "contextWarnThreshold")
363
+ threshold = (
364
+ float(cfg_threshold)
365
+ if isinstance(cfg_threshold, (int, float))
366
+ else _DEFAULT_THRESHOLD
367
+ )
368
+
369
+ transcript = _latest_transcript(Path.cwd())
370
+ if transcript is None:
371
+ return {"available": False, "reason": "no session transcript found"}
372
+ found = _last_usage(transcript)
373
+ if found is None:
374
+ return {"available": False, "reason": "no usage record in transcript"}
375
+ tokens, model = found
376
+
377
+ window = window_override
378
+ if window is None or window <= 0:
379
+ cfg_window = _config_value(config_path, "contextWindowTokens")
380
+ if isinstance(cfg_window, int) and cfg_window > 0:
381
+ window = cfg_window
382
+ else:
383
+ # Inferring (no override, no config). Start from the model marker /
384
+ # conservative default, then auto-bump: observed tokens above the
385
+ # default window PROVE a wider window is active (a 200k session can
386
+ # never exceed 200k), so widen to 1M rather than report a nonsensical
387
+ # >100%. Never under-reports a real 200k session, which can't trip it.
388
+ window = _infer_window(model)
389
+ if tokens > window:
390
+ window = _WIDE_WINDOW
391
+
392
+ pct = round(tokens / window, 4)
393
+ over = pct >= threshold
394
+ if over:
395
+ recommendation = "clean-session"
396
+ else:
397
+ recommendation = "continue"
398
+ return {
399
+ "available": True,
400
+ "tokens": tokens,
401
+ "windowTokens": window,
402
+ "pct": pct,
403
+ "threshold": threshold,
404
+ "overThreshold": over,
405
+ "recommendation": recommendation,
406
+ "model": model,
407
+ }
408
+
409
+
410
+ # --------------------------------------------------------------------------- #
411
+ # CLI dispatch
412
+ # --------------------------------------------------------------------------- #
413
+
414
+
415
+ def _print_rank_table(rows: list[FeatureRow], counts: dict[str, int]) -> None:
416
+ """Print a human-readable recency-ranked feature list."""
417
+ print(
418
+ f"Active: {counts['active']} "
419
+ f"(paused: {counts['paused']}, abandoned: {counts['abandoned']})"
420
+ )
421
+ if not rows:
422
+ print(" (no active feature pipelines)")
423
+ return
424
+ for idx, row in enumerate(rows):
425
+ marker = "→" if idx == 0 else " "
426
+ label = row["name"] + (f" [{row['epic']}]" if row["epic"] else "")
427
+ nxt = row["nextCommand"] or "complete"
428
+ print(f" {marker} {label}: {row['currentStage']} — next: {nxt}")
429
+ if row["verifyPending"]:
430
+ print(f" (verify available: {row['verifyCommand']})")
431
+
432
+
433
+ def _print_context(usage: dict) -> None:
434
+ """Print a one-line human-readable context-usage summary."""
435
+ if not usage.get("available"):
436
+ print(f"context usage: unavailable ({usage.get('reason', 'unknown')})")
437
+ return
438
+ pct = round(usage["pct"] * 100, 1)
439
+ flag = " — over threshold, clean session recommended" if usage["overThreshold"] else ""
440
+ print(
441
+ f"context: {usage['tokens']:,} / {usage['windowTokens']:,} tokens "
442
+ f"(~{pct}%){flag}"
443
+ )
444
+
445
+
446
+ def main() -> int:
447
+ parser = argparse.ArgumentParser(prog="forge-session.py", description=__doc__)
448
+ sub = parser.add_subparsers(dest="cmd", required=True)
449
+
450
+ p_rank = sub.add_parser("rank-features", help="Rank active features by recency")
451
+ p_rank.add_argument("--specs-dir", default="./specs", help="Specs directory")
452
+ p_rank.add_argument("--json", action="store_true", dest="json_output")
453
+
454
+ p_ctx = sub.add_parser("context-usage", help="Report live context-window usage")
455
+ p_ctx.add_argument("--config", default="./forge.config.json", help="forge.config.json path")
456
+ p_ctx.add_argument("--window", type=int, default=None, help="Override context window size")
457
+ p_ctx.add_argument("--threshold", type=float, default=None, help="Override warn fraction (0-1)")
458
+ p_ctx.add_argument("--json", action="store_true", dest="json_output")
459
+
460
+ args = parser.parse_args()
461
+
462
+ try:
463
+ if args.cmd == "rank-features":
464
+ specs_dir = Path(args.specs_dir)
465
+ rows = build_rows(specs_dir)
466
+ counts = _counts(specs_dir)
467
+ if args.json_output:
468
+ print(json.dumps({"active": rows, "counts": counts}, indent=2, ensure_ascii=False))
469
+ else:
470
+ _print_rank_table(rows, counts)
471
+ return 0
472
+
473
+ if args.cmd == "context-usage":
474
+ usage = context_usage(Path(args.config), args.window, args.threshold)
475
+ if args.json_output:
476
+ print(json.dumps(usage, indent=2, ensure_ascii=False))
477
+ else:
478
+ _print_context(usage)
479
+ return 0
480
+
481
+ raise UsageError(f"unknown command: {args.cmd}")
482
+ except UsageError as exc:
483
+ print(f"Error: {exc}", file=sys.stderr)
484
+ return 2
485
+ except OSError as exc:
486
+ print(f"Error: {exc}", file=sys.stderr)
487
+ return 2
488
+
489
+
490
+ if __name__ == "__main__":
491
+ sys.exit(main())
@@ -35,7 +35,14 @@ python3 "$R/scripts/epic-manifest.py" render-status "{epic}" --specs-dir "{specs
35
35
  ```
36
36
  and show one rollup line: `{epic} — {complete}/{total} complete, next: {nextCommand}`.
37
37
  2. **Standalone features below.** Scan the remaining `{specsDir}/*/` that directly contain a `.pipeline-state.json` **without** an `epic` back-pointer. A nested member's `.pipeline-state.json` is **attributed to its epic (Tier 1), never listed as a standalone feature**.
38
- - Within this standalone tier the existing logic still applies: if exactly one active (non-complete) standalone pipeline exists, show its dashboard; if multiple exist, list them with a one-line summary each and use `AskUserQuestion` to ask which to focus on.
38
+ - **Rank by recency.** Run the recency ranker so the most-recently-touched active feature is the default — the user rarely has to type a name (especially on mobile after a `/clear`):
39
+ ```bash
40
+ R="$(for d in "$HOME"/.claude/skills/feature-forge "$HOME"/.claude/plugins/*/feature-forge "$HOME"/.agents/skills/feature-forge ./.agents/skills/feature-forge; do [ -x "$d/scripts/forge-root.sh" ] && exec "$d/scripts/forge-root.sh"; done)"
41
+ [ -n "$R" ] || { echo "feature-forge: cannot locate plugin root" >&2; exit 1; }
42
+ python3 "$R/scripts/forge-session.py" rank-features --specs-dir "{specsDir}" --json
43
+ ```
44
+ This returns `{active: [...], counts: {...}}` with active features sorted by `updatedAt` **descending** (row 0 is the most recent). Each row carries `currentStage`, `nextStage`, `nextCommand`, `verifyPending`, and `verifyCommand` (the single source of stage order). The `active` list excludes nested epic members surfaced in Tier 1 — but the ranker scans them too, so ignore rows whose `epic` is non-null here (they belong to the epic rollup).
45
+ - **Pick the feature:** if exactly one active standalone pipeline exists, show its dashboard. If multiple exist, use `AskUserQuestion` — **list the most-recently-updated first, labeled `(recommended)`**, each option's description showing its `currentStage` and a relative age ("updated 2h ago"). Always include a free-form escape ("A different feature / something else") so the user is never boxed in. Then render the chosen feature's dashboard.
39
46
 
40
47
  If no epics and no standalone features exist, say: "No active feature pipelines found. Start one with `/feature-forge:forge-1-prd <feature-name>` or group several with `/feature-forge:forge-0-epic <epic-name>`."
41
48
 
@@ -78,6 +85,35 @@ Use these status indicators:
78
85
  - ⏭️ = verification skipped (user chose to proceed without verifying)
79
86
  - ⚠️ = stale (built against an older version of an upstream artifact)
80
87
 
88
+ ### 3b. Drive to the Next Stage
89
+
90
+ After rendering a **per-feature** dashboard for an **active** pipeline (skip this for paused/abandoned pipelines and for the Epic Dashboard), don't stop at a text suggestion — actively offer to start the next stage. This removes the copy-paste-after-`/clear` chore that makes long, multi-stage runs painful (especially on mobile).
91
+
92
+ **1. Read the next step.** From the `rank-features --json` output (above), find this feature's row and read its `nextStage`, `nextCommand`, `verifyPending`, and `verifyCommand`. If the feature is not in the `active` list (paused/abandoned), or `nextStage` is `null` (every production stage complete), skip the drive prompt — instead congratulate the user and, if `forge-6-docs` has not run, offer it; otherwise note the pipeline is complete.
93
+
94
+ **2. Check the context window.** Run the context-usage helper so you can advise whether to continue here or start the next stage in a fresh session:
95
+ ```bash
96
+ R="$(for d in "$HOME"/.claude/skills/feature-forge "$HOME"/.claude/plugins/*/feature-forge "$HOME"/.agents/skills/feature-forge ./.agents/skills/feature-forge; do [ -x "$d/scripts/forge-root.sh" ] && exec "$d/scripts/forge-root.sh"; done)"
97
+ [ -n "$R" ] || { echo "feature-forge: cannot locate plugin root" >&2; exit 1; }
98
+ python3 "$R/scripts/forge-session.py" context-usage --json
99
+ ```
100
+ - `{"available": true, ...}` → note `pct` (e.g. "context ~68% full") and `overThreshold`. Window/threshold come from `contextWindowTokens` / `contextWarnThreshold` in `forge.config.json` (the helper defaults to a 200k window and 0.7 threshold, and auto-bumps the assumed window to 1M once observed usage exceeds 200k; **on a 1M-context model set `contextWindowTokens: 1000000` so the percentage is accurate below 200k too** — 1M can't be detected from the transcript until usage crosses 200k).
101
+ - `{"available": false, ...}` → omit context advice silently (non-Claude host, or a fresh session with no transcript). Never treat this as an error.
102
+
103
+ **3. Offer the next step via `AskUserQuestion`.** Output the dashboard + a one-line context note as text, then ask (per the Decision Support protocol in `references/shared-conventions.md`). Options, in this order:
104
+ - **Start `{nextStage}` now** — recommended **when context is healthy** (`overThreshold` false or context unavailable).
105
+ - **Start in a clean session** — recommended-**first** instead **when `overThreshold` is true**. The work survives a clear because all state is on disk: instruct the user to `/clear`, then re-run `/feature-forge:forge {feature}` (or run `{nextCommand}` directly) in the fresh session. Note plainly that you cannot `/clear` for them.
106
+ - **Verify `{stage}` first** — include **only when `verifyPending` is true**; selecting it runs `{verifyCommand}`.
107
+ - **Pick a different stage** — free-form escape to any stage or other action.
108
+
109
+ **4. Act on the choice.**
110
+ - **Start now** → if `autoInvokeNextStage` is true (default) **and** the `Skill` tool is available, invoke the chosen stage **via the `Skill` tool** in this same session (e.g. `skill: "feature-forge:forge-3-specs"`, `args: "{feature}"`) — no retyping, no paste. If `autoInvokeNextStage` is false, or the `Skill` tool is unavailable (a non-Claude host), fall back to printing `{nextCommand}` prominently for the user to run.
111
+ - **Clean session** → give the exact next command and the `/clear`-then-re-run instruction; do not invoke anything.
112
+ - **Verify** → invoke `feature-forge:forge-verify` via the `Skill` tool (or print `{verifyCommand}` on a non-Claude host).
113
+ - **Different stage** → honor the free-form request.
114
+
115
+ This applies whether the feature was named explicitly (`/feature-forge:forge {feature}`) or resolved from the recency default.
116
+
81
117
  ### Epic Dashboard
82
118
 
83
119
  When the named argument is an epic (`{specsDir}/{name}/epic-manifest.json` exists), render the epic dashboard instead of a per-feature one. Run:
@@ -143,6 +179,12 @@ Support these sub-commands for pipeline lifecycle management:
143
179
  - `/feature-forge:forge pause {feature}` — Set `pipelineStatus` to `"paused"`. Do NOT modify `currentStage` or any stage statuses. The pipeline freezes exactly as-is. Show a confirmation.
144
180
  - `/feature-forge:forge resume {feature}` — Set `pipelineStatus` back to `"active"`. Calculate how long the feature was paused (from `updatedAt` to now). If paused for more than 24 hours, show a hint: "This feature was paused for {duration}. Session context may have been lost — consider re-running `/feature-forge:forge-{currentStage} {feature}` to rebuild context."
145
181
  - `/feature-forge:forge abandon {feature}` — Set `pipelineStatus` to `"abandoned"`. Use `AskUserQuestion` to confirm first, and state what's reversible: abandoning does not delete artifacts and can be undone with `/feature-forge:forge resume {feature}`, so the cost is low — but if the user really means "stop and discard," point out that `pause` is the better choice when they're only setting it aside. Offer **Abandon** · **Pause instead** · **Cancel**.
182
+ - `/feature-forge:forge run [{feature}]` — **Opt-in auto-advance.** Drive the feature through consecutive stages in one session instead of confirming each boundary. This is a convenience wrapper over **3b. Drive to the Next Stage** — same stage order, same context gate — just looped:
183
+ 1. Resolve the feature (if omitted, use the recency default from `rank-features`; if multiple are equally plausible, ask once via `AskUserQuestion`).
184
+ 2. **Before each stage,** run `forge-session.py context-usage`. If `overThreshold` is true, **stop** and recommend a clean session (give the exact `{nextCommand}` and the `/clear`-then-re-run instruction) — never auto-`/clear`.
185
+ 3. Otherwise invoke the next stage's skill via the `Skill` tool, let it run to its natural stopping point, then re-read state and continue from step 2.
186
+ 4. **Stop conditions:** the next stage is an interview/decision point that calls `AskUserQuestion` (PRD and tech inherently pause for input — let them); `nextStage` is `null` (pipeline complete); context over threshold; or a stage signals it needs human input / is blocked. Always report where the loop stopped and why.
187
+ Per-stage confirmation (3b) remains the default — `run` is only used when the user explicitly asks to "run" / "drive" / "auto-advance" the pipeline. On a non-Claude host where the `Skill` tool is unavailable, fall back to printing the ordered list of commands to run.
146
188
 
147
189
  **Epic lifecycle.** When the argument names an **epic** (`{specsDir}/{name}/epic-manifest.json` exists), `pause` / `resume` / `abandon` operate on the epic manifest, not a `.pipeline-state.json`:
148
190