nccgs 1.3.0 → 2.0.0-rc.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/agents/nccgs-adversarial-reviewer.md +17 -6
- package/.claude/agents/nccgs-engine-programmer.md +2 -2
- package/.claude/agents/nccgs-fast-implementer.md +30 -0
- package/.claude/agents/nccgs-game-designer.md +26 -0
- package/.claude/agents/nccgs-gameplay-programmer.md +3 -4
- package/.claude/agents/nccgs-qa-engineer.md +2 -2
- package/.claude/agents/nccgs-task-scout.md +24 -0
- package/.claude/agents/nccgs-ui-programmer.md +2 -2
- package/.claude/agents/nccgs-unity-build-specialist.md +2 -2
- package/.claude/agents/nccgs-unity-implementer.md +7 -6
- package/.claude/agents/nccgs-unity-systems-specialist.md +2 -2
- package/.claude/agents/nccgs-unity-ui-specialist.md +2 -2
- package/.claude/agents/nccgs-verification-engineer.md +2 -2
- package/.claude/nccgs/MEMORY.md +59 -0
- package/.claude/nccgs/THIRD_PARTY_NOTICES.md +16 -0
- package/.claude/nccgs/VERSION +1 -1
- package/.claude/nccgs/agent-library/nccgs-adversarial-reviewer.md +16 -5
- package/.claude/nccgs/agent-library/nccgs-fast-implementer.md +28 -0
- package/.claude/nccgs/agent-library/nccgs-game-designer.md +1 -2
- package/.claude/nccgs/agent-library/nccgs-gameplay-programmer.md +1 -2
- package/.claude/nccgs/agent-library/nccgs-task-scout.md +22 -0
- package/.claude/nccgs/agent-library/nccgs-unity-implementer.md +5 -4
- package/.claude/nccgs/constitution.md +44 -12
- package/.claude/nccgs/hooks/agent-audit.mjs +2 -2
- package/.claude/nccgs/hooks/closure-guard.mjs +31 -8
- package/.claude/nccgs/hooks/common.mjs +2 -1
- package/.claude/nccgs/hooks/runtime-observation.mjs +2 -2
- package/.claude/nccgs/hooks/session-start.mjs +11 -4
- package/.claude/nccgs/packages/ponytail/LICENSE +21 -0
- package/.claude/nccgs/packages/ponytail/NOTICE.md +32 -0
- package/.claude/nccgs/packages/ponytail/README.md +46 -0
- package/.claude/nccgs/packages/ponytail/complexity-review.md +47 -0
- package/.claude/nccgs/packages/ponytail/engineering.md +68 -0
- package/.claude/nccgs/packages/ponytail/manifest.json +24 -0
- package/.claude/nccgs/pipeline.json +92 -8
- package/.claude/nccgs/protocols/agent-contract.md +50 -1
- package/.claude/nccgs/protocols/astra-claude.md +46 -12
- package/.claude/nccgs/protocols/bounded-repair.md +118 -0
- package/.claude/nccgs/protocols/context-packets.md +5 -0
- package/.claude/nccgs/protocols/delivery-acceptance.md +105 -0
- package/.claude/nccgs/protocols/design-lifecycle.md +139 -0
- package/.claude/nccgs/protocols/evidence.md +11 -1
- package/.claude/nccgs/protocols/model-routing.md +44 -42
- package/.claude/nccgs/protocols/orchestration.md +46 -5
- package/.claude/nccgs/protocols/review-handoff.md +65 -0
- package/.claude/nccgs/protocols/standalone-gates.md +273 -0
- package/.claude/nccgs/protocols/standalone-pilot.md +144 -0
- package/.claude/nccgs/protocols/standalone.md +187 -0
- package/.claude/nccgs/protocols/task-autonomy.md +55 -0
- package/.claude/nccgs/protocols/task-gates.md +73 -7
- package/.claude/nccgs/routing.generated.md +25 -7
- package/.claude/nccgs/settings.fragment.json +153 -141
- package/.claude/nccgs/studio.json +135 -101
- package/.claude/nccgs/tools/autonomy.mjs +27 -0
- package/.claude/nccgs/tools/command-process.mjs +44 -0
- package/.claude/nccgs/tools/compile-policy.mjs +34 -4
- package/.claude/nccgs/tools/configure-models.mjs +13 -5
- package/.claude/nccgs/tools/criteria.mjs +121 -0
- package/.claude/nccgs/tools/delivery-budget.mjs +52 -0
- package/.claude/nccgs/tools/design-readiness.mjs +158 -0
- package/.claude/nccgs/tools/evidence.mjs +153 -0
- package/.claude/nccgs/tools/gates.mjs +151 -24
- package/.claude/nccgs/tools/inputs.mjs +73 -0
- package/.claude/nccgs/tools/lifecycle.mjs +521 -0
- package/.claude/nccgs/tools/native-completion.mjs +21 -0
- package/.claude/nccgs/tools/orchestration.mjs +430 -0
- package/.claude/nccgs/tools/paths.mjs +60 -0
- package/.claude/nccgs/tools/policy.mjs +58 -13
- package/.claude/nccgs/tools/process-lock.mjs +74 -0
- package/.claude/nccgs/tools/repair-packet.mjs +87 -0
- package/.claude/nccgs/tools/review-handoff.mjs +32 -0
- package/.claude/nccgs/tools/runtime.mjs +33 -7
- package/.claude/nccgs/tools/task.mjs +62 -31
- package/.claude/nccgs/tools/tool-guard.mjs +117 -0
- package/.claude/nccgs/tools/verification-jobs.mjs +161 -0
- package/.claude/nccgs/tools/verification-worker.mjs +71 -0
- package/.claude/nccgs/tools/windows-job.cs +96 -0
- package/.claude/nccgs/workflow-catalog.json +286 -26
- package/.claude/skills/accessibility-review/SKILL.md +12 -5
- package/.claude/skills/architecture-decision/SKILL.md +12 -5
- package/.claude/skills/asset-audit/SKILL.md +12 -5
- package/.claude/skills/audit/SKILL.md +12 -5
- package/.claude/skills/balance-review/SKILL.md +12 -5
- package/.claude/skills/bug-triage/SKILL.md +12 -5
- package/.claude/skills/bug-triage/references/astra.md +12 -5
- package/.claude/skills/closure/SKILL.md +12 -5
- package/.claude/skills/compatibility-review/SKILL.md +12 -5
- package/.claude/skills/context-pack/SKILL.md +12 -5
- package/.claude/skills/dependency-review/SKILL.md +12 -5
- package/.claude/skills/design/SKILL.md +14 -5
- package/.claude/skills/design-review/SKILL.md +14 -5
- package/.claude/skills/evidence-review/SKILL.md +12 -5
- package/.claude/skills/hotfix/SKILL.md +12 -5
- package/.claude/skills/hotfix/references/astra.md +12 -5
- package/.claude/skills/incident-recovery/SKILL.md +12 -5
- package/.claude/skills/incident-recovery/references/astra.md +12 -5
- package/.claude/skills/localize-game/SKILL.md +12 -5
- package/.claude/skills/localize-game/references/astra.md +12 -5
- package/.claude/skills/migrate-project/SKILL.md +12 -5
- package/.claude/skills/migrate-project/references/astra.md +12 -5
- package/.claude/skills/migrate-project/references/procedure.md +2 -1
- package/.claude/skills/milestone-review/SKILL.md +12 -5
- package/.claude/skills/nccgs-code-review/SKILL.md +12 -5
- package/.claude/skills/nccgs-code-review/references/astra.md +8 -2
- package/.claude/skills/performance-audit/SKILL.md +12 -5
- package/.claude/skills/plan-feature/SKILL.md +12 -5
- package/.claude/skills/playtest/SKILL.md +12 -5
- package/.claude/skills/playtest/references/astra.md +16 -5
- package/.claude/skills/project-stage/SKILL.md +12 -5
- package/.claude/skills/prototype-feature/SKILL.md +12 -5
- package/.claude/skills/prototype-feature/references/astra.md +16 -5
- package/.claude/skills/qa-plan/SKILL.md +12 -5
- package/.claude/skills/release/SKILL.md +12 -5
- package/.claude/skills/release/references/astra.md +12 -5
- package/.claude/skills/release-readiness/SKILL.md +12 -5
- package/.claude/skills/retrospective/SKILL.md +12 -5
- package/.claude/skills/review/SKILL.md +12 -5
- package/.claude/skills/review/references/astra.md +8 -2
- package/.claude/skills/security-audit/SKILL.md +12 -5
- package/.claude/skills/sprint-plan/SKILL.md +12 -5
- package/.claude/skills/status/SKILL.md +12 -5
- package/.claude/skills/status/references/astra.md +6 -0
- package/.claude/skills/story-readiness/SKILL.md +14 -5
- package/.claude/skills/test/SKILL.md +12 -5
- package/.claude/skills/test/references/astra.md +12 -5
- package/.claude/skills/ui-review/SKILL.md +12 -5
- package/.claude/skills/work/SKILL.md +12 -5
- package/.claude/skills/work/references/astra.md +17 -5
- package/CHANGELOG.md +42 -0
- package/CLAUDE.md +2 -0
- package/MEMORY.md +55 -0
- package/README.md +132 -180
- package/THIRD_PARTY_NOTICES.md +15 -0
- package/UPGRADING.md +91 -69
- package/VERSION +1 -1
- package/docs/ARCHITECTURE.md +19 -5
- package/docs/BOUNDED-REPAIR.md +11 -0
- package/docs/HUONG-DAN-MIGRATE-VA-SU-DUNG.md +46 -34
- package/docs/NCCGS-1.3.md +10 -122
- package/docs/NCCGS-1.4.1.md +149 -0
- package/docs/NCCGS-1.4.md +184 -0
- package/docs/NCCGS-1.5-IMPLEMENTATION-PLAN.md +378 -0
- package/docs/NCCGS-1.5-PILOT.md +130 -0
- package/docs/NCCGS-1.5.md +237 -0
- package/docs/PHASE-1-OPERATIONS.md +137 -0
- package/docs/PHASE-2-OPERATIONS.md +183 -0
- package/docs/PROJECT-POLICY.md +21 -9
- package/docs/RELEASE-STATUS.md +42 -0
- package/docs/WORKFLOWS.md +42 -24
- package/package.json +7 -3
- package/scaffold/.nccgs/autonomy.json +4 -0
- package/scaffold/.nccgs/execution.json +10 -0
- package/scaffold/.nccgs/inputs.json +6 -0
- package/scaffold/.nccgs/project.yaml +4 -2
- package/scaffold/.nccgs/templates/agent-handoff.md +2 -0
- package/scaffold/.nccgs/templates/design-contract.json +34 -0
- package/scaffold/.nccgs/templates/feature-contract.md +21 -5
- package/scaffold/.nccgs/templates/implementation-report.md +4 -1
- package/scaffold/.nccgs/templates/repair-packet.json +14 -0
- package/scripts/benchmark-snapshot.mjs +205 -0
- package/scripts/benchmark-v15.mjs +173 -0
- package/scripts/cli.mjs +129 -2
- package/scripts/doctor.mjs +128 -22
- package/scripts/install.mjs +96 -34
- package/scripts/run.mjs +72 -0
- package/scripts/supervise.mjs +124 -0
- package/scripts/validate.mjs +45 -4
- package/tests/autonomy.test.mjs +113 -0
- package/tests/benchmark-v15.test.mjs +65 -0
- package/tests/design-readiness.test.mjs +200 -0
- package/tests/evidence.test.mjs +186 -0
- package/tests/framework.test.mjs +333 -16
- package/tests/gates.test.mjs +313 -4
- package/tests/lifecycle-v15.test.mjs +300 -0
- package/tests/orchestration-v15.test.mjs +545 -0
- package/tests/phase1.test.mjs +324 -0
- package/tests/phase2.test.mjs +250 -0
- package/tests/policy-v15.test.mjs +220 -0
- package/tests/ponytail-package.test.mjs +75 -0
- package/tests/runtime-v15.test.mjs +624 -0
|
@@ -0,0 +1,187 @@
|
|
|
1
|
+
# Standalone Claude Code execution — 2.0 RC
|
|
2
|
+
|
|
3
|
+
Read `.claude/nccgs/MEMORY.md`. NCCGS means New Claude Code Game Studio. Only
|
|
4
|
+
real Claude Code execution supplies NCCGS orchestration evidence. Codex maintenance,
|
|
5
|
+
synthetic tests, installed files and requested model aliases do not prove it.
|
|
6
|
+
|
|
7
|
+
Read [task autonomy](task-autonomy.md). The task authorizes necessary shell commands,
|
|
8
|
+
reports and outside-project operations without asking again. Start with `nccgs run`
|
|
9
|
+
to activate native permissions; execution mode and evidence gates remain separate.
|
|
10
|
+
|
|
11
|
+
## Select obligations before models
|
|
12
|
+
|
|
13
|
+
`.nccgs/execution.json` selects standalone or compatibility, external review
|
|
14
|
+
`off|advisory|required`, acceptance, observed-effort requirements and worker limit.
|
|
15
|
+
`.nccgs/project.yaml` selects the model profile independently. Missing execution
|
|
16
|
+
configuration means compatibility; never silently relax existing gates. New
|
|
17
|
+
installs use standalone and `claude-standalone`. Astra is an optional reviewer.
|
|
18
|
+
Policy and input hashes are pinned per task. Changed inputs require an explicit
|
|
19
|
+
refresh/reclassification with invalidation. Changed execution policy requires
|
|
20
|
+
restoring the pinned configuration for current work or explicitly scoped new work;
|
|
21
|
+
preserve prior history and spent budgets. No prior approval carries over.
|
|
22
|
+
|
|
23
|
+
## Intake and mode
|
|
24
|
+
|
|
25
|
+
For `/design`, `/design-review` and `/story-readiness`, also follow the
|
|
26
|
+
[design lifecycle](design-lifecycle.md). Before dependent feature implementation,
|
|
27
|
+
use its reviewed design handoff. New installs require this through execution policy;
|
|
28
|
+
a self-declared readiness value alone is insufficient. `task design-validate`
|
|
29
|
+
checks document structure; the independent designer/reviewer process owns semantics.
|
|
30
|
+
|
|
31
|
+
Use [delivery acceptance](delivery-acceptance.md) to distinguish implementing the
|
|
32
|
+
brief from bot performance and product-specific balance research. Keep research
|
|
33
|
+
questions outside delivery gates unless the task itself explicitly requests that
|
|
34
|
+
research deliverable. Preserve fixed rules and concrete acceptance in the brief.
|
|
35
|
+
|
|
36
|
+
Use `implement` for requested changes, `audit` for examination, and `plan` for
|
|
37
|
+
design/planning. `/audit`, `/review`, diagnostic and specialist audit/review commands
|
|
38
|
+
default to audit; `/design`, `/plan-feature`, `/qa-plan`, `/sprint-plan` default to
|
|
39
|
+
plan. `/status` reads state without dispatch. Other entrypoints follow the user's
|
|
40
|
+
requested outcome; never turn a review into a fix without authorization.
|
|
41
|
+
|
|
42
|
+
Write a bounded intake JSON under `.nccgs/evidence/intake/` with `objective`,
|
|
43
|
+
`mode`, `scope` (`local|feature|cross-system`), `systems`, `target`
|
|
44
|
+
(`fix|learning-prototype|alpha-ready|production-change`), `readiness`
|
|
45
|
+
(`ready|needs-inspection|needs-product-decision`), acceptance criteria,
|
|
46
|
+
`requiredChecks`, `requiredArtifacts`, and `manualGates`. Audit/plan outputs belong
|
|
47
|
+
only to explicitly listed `analysisArtifacts` under `.nccgs/evidence/analysis/`.
|
|
48
|
+
Use `riskSignals` for discovered compatibility, security and other critical risks.
|
|
49
|
+
Do not request irrelevant tests or multiple reviewers merely to fill a template.
|
|
50
|
+
Map automatic criteria to JSON measurement assertions, including assistance/rule
|
|
51
|
+
changes when natural gameplay is required. A check name or helper exit 0 alone is
|
|
52
|
+
not acceptance. Mark only a small, actual delivery prerequisite `checkpoint: "early"`;
|
|
53
|
+
do not make feature expansion depend on bot survival, victory or balance. Put late
|
|
54
|
+
integration checks in final acceptance when independent work can proceed.
|
|
55
|
+
Read `task acceptance-status`; preserve
|
|
56
|
+
FAIL/UNVERIFIED and missing evidence in partial reports. Exact fields and examples
|
|
57
|
+
are in [phase-2 operations](../../../docs/PHASE-2-OPERATIONS.md).
|
|
58
|
+
|
|
59
|
+
Initialize and bind through the CLI:
|
|
60
|
+
|
|
61
|
+
```text
|
|
62
|
+
node .claude/nccgs/tools/task.mjs init --id TASK --intake .nccgs/evidence/intake/TASK.json
|
|
63
|
+
node .claude/nccgs/tools/task.mjs bind --id TASK --session SESSION_ID
|
|
64
|
+
node .claude/nccgs/tools/task.mjs inspect --id TASK
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
The rule-based classifier chooses a conservative tier with reasons and evidence;
|
|
68
|
+
it is not a semantic oracle. Inspect the result. Unknown scope stays STANDARD;
|
|
69
|
+
save/migration/authority/commerce/cross-system signals require CONTROLLED. Any
|
|
70
|
+
override needs evidence and a reason. Reclassify when facts change, preserving
|
|
71
|
+
spent review rounds, dispatches and history. Substantial product ambiguity blocks
|
|
72
|
+
dependent work; independent inspection may proceed.
|
|
73
|
+
|
|
74
|
+
## Reserve, execute, verify, review
|
|
75
|
+
|
|
76
|
+
The main Claude Code session coordinates. Reserve every worker before invoking
|
|
77
|
+
Agent; do not delegate recursively. A reservation request contains a stable
|
|
78
|
+
`requestKey`, `kind` (`implementation|analysis|review`), known `agent`, concrete
|
|
79
|
+
`reason`, and explicit `writePaths` (empty for analysis/review). Save requests under
|
|
80
|
+
`.nccgs/evidence/requests/`. Use the returned reservation ID once:
|
|
81
|
+
|
|
82
|
+
```text
|
|
83
|
+
node .claude/nccgs/tools/task.mjs reserve --id TASK --request .nccgs/evidence/requests/TASK-work.json
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
Prefix the Agent prompt with `[NCCGS-DISPATCH:RETURNED_ID]`. Include the task's
|
|
87
|
+
contract, route, owned paths, acceptance criteria and expected bounded handoff.
|
|
88
|
+
The pre-tool hook claims the reservation using the actual session/tool-use ID.
|
|
89
|
+
Use foreground Agent calls for this preview. Launch distinct roles sequentially
|
|
90
|
+
through their start hooks; ambiguous start identities must not be guessed.
|
|
91
|
+
Background Agent/Task calls are rejected.
|
|
92
|
+
`nccgs run` also sets `CLAUDE_CODE_DISABLE_BACKGROUND_TASKS=1` to prevent native
|
|
93
|
+
auto-backgrounding; an omitted `run_in_background` flag alone is insufficient.
|
|
94
|
+
Launch one Agent call per coordinator message and wait for its final result before
|
|
95
|
+
starting another. Do not mutate a dispatch from uncertain to running to bypass this.
|
|
96
|
+
The environment setting also disables native background Bash tasks; use managed
|
|
97
|
+
verification jobs or explicitly tracked bounded collectors for long-running checks.
|
|
98
|
+
See https://code.claude.com/docs/en/env-vars for the native setting.
|
|
99
|
+
For required handoff documents, declare
|
|
100
|
+
fresh `handoffPaths` under `.nccgs/evidence/analysis/` in the reservation. Invocation
|
|
101
|
+
completion with missing declared artifacts cannot satisfy a work stage; undeclared
|
|
102
|
+
handoffs are unassessed, not automatically complete.
|
|
103
|
+
Use the risk-selected implementation model/effort and configured review model.
|
|
104
|
+
Aliases are requests; only resolved Claude model observations count as evidence.
|
|
105
|
+
|
|
106
|
+
At most two implementation workers may run with disjoint path ownership. Review
|
|
107
|
+
and analysis hold a stable source snapshot and cannot overlap writers. The review
|
|
108
|
+
slot is unique per task. Failed/cancelled attempts still consume quota, including
|
|
109
|
+
ancestor quota for child tasks. Do not split tasks to mint budget.
|
|
110
|
+
New implementation budgets protect initial review plus one repair, read-only QA and
|
|
111
|
+
re-review within the total cap. Use `purpose: "repair"` only against recorded negative
|
|
112
|
+
feedback; `purpose: "qa"` is a read-only QA/verification dispatch. Ordinary delivery
|
|
113
|
+
cannot consume the protected tail. Inspect `dispatch-status.deliveryBudget` and stop
|
|
114
|
+
expansion before it runs out; more review rounds do not create more dispatch slots.
|
|
115
|
+
|
|
116
|
+
Fresh repair reservations require a pinned `workPacket`, 1–4 exact existing source
|
|
117
|
+
files and one fresh handoff path. Choose the correction before reserving; do not
|
|
118
|
+
send an open-ended tuning investigation. `task dispatch-prompt --id TASK --dispatch ID`
|
|
119
|
+
returns the bounded assignment; the launch hook injects the same prompt. Repair
|
|
120
|
+
workers have at most 12 work tool calls (6 reads), 3 separate handoff writes and a
|
|
121
|
+
10-minute tool-admission deadline. Shell/MCP/web/delegation are blocked for these
|
|
122
|
+
workers. The coordinator runs declared checks after return. See
|
|
123
|
+
[bounded repair](bounded-repair.md) for packet fields and limitations.
|
|
124
|
+
These limits do not add native turns or reset dispatch quota.
|
|
125
|
+
|
|
126
|
+
Use targeted objective checks appropriate to the request. The coordinator's
|
|
127
|
+
`task verify` owns the command and captures a receipt on the first execution. For
|
|
128
|
+
long commands use `--request-key ATTEMPT --background`; retain the returned job ID
|
|
129
|
+
and use `job-status` or bounded `job-wait`. Reuse its receipt for QA and closure;
|
|
130
|
+
do not rerun successful checks just to capture evidence. A helper's arguments after
|
|
131
|
+
`--` belong to the helper, including its own `--project`. Audit/plan/review
|
|
132
|
+
workers use read tools and declared analysis artifacts; arbitrary shell and unknown
|
|
133
|
+
MCP tools are blocked. If a required analysis check needs command execution,
|
|
134
|
+
provide externally run evidence or start a separately authorized implementation
|
|
135
|
+
task. Do not loosen mode simply to bypass the guard.
|
|
136
|
+
|
|
137
|
+
Record implementation/analysis with its completed `dispatchId`, current snapshot,
|
|
138
|
+
checks, artifacts and captured runtime observation. Record independent review with
|
|
139
|
+
a fresh completed review dispatch and explicit verdict/findings. The same dispatch
|
|
140
|
+
cannot be consumed twice. Read `task inspect` and the task-gates schema documentation
|
|
141
|
+
for exact payload fields in [standalone gates](standalone-gates.md). Missing observations block the relevant stage; never
|
|
142
|
+
author runtime evidence to make it pass.
|
|
143
|
+
|
|
144
|
+
## Report and stop
|
|
145
|
+
|
|
146
|
+
Always permit an honest report outcome: `success`, `partial`, `blocked` or `failed`.
|
|
147
|
+
Keep unresolved findings, absent tests, missing measurements and required decisions
|
|
148
|
+
visible. Reporting does not imply DONE. Successful closure checks the final source,
|
|
149
|
+
contract, verification, independent review, required artifacts and configured
|
|
150
|
+
external/manual/acceptance gates. Stop once these gates pass; repeat work only for
|
|
151
|
+
changed inputs, new findings or failed evidence.
|
|
152
|
+
|
|
153
|
+
External review is file based. `off` adds no dependency. `advisory` retains findings
|
|
154
|
+
without gating closure. `required` gates closure with a receipt tied to task,
|
|
155
|
+
contract and source snapshot. External actors import their own decisions outside
|
|
156
|
+
Claude. NCCGS does not launch Astra or send messages to another service.
|
|
157
|
+
|
|
158
|
+
## Resume and evidence limits
|
|
159
|
+
|
|
160
|
+
Use `checkpoint`, `resume`, `dispatch-status` and `telemetry`; do not edit task or
|
|
161
|
+
ledger JSON. A background launch, missing callback or crashed worker retains its
|
|
162
|
+
slot. Only correlated final provider evidence or external reconciliation releases
|
|
163
|
+
it. No lease timeout refunds attempts. New local locks can recover only after the
|
|
164
|
+
owner and all managed command descendants are verified dead, with an audit entry.
|
|
165
|
+
Legacy/malformed/remote locks and uncertain spawn windows still need inspection.
|
|
166
|
+
Restore missing/corrupt ledger
|
|
167
|
+
history from a known checkpoint; never initialize a replacement to reset budgets.
|
|
168
|
+
|
|
169
|
+
Native hooks observe resolved models and selected usage fields when Claude Code
|
|
170
|
+
actually provides them. They do not prove effort when no effort field is available.
|
|
171
|
+
Unknown values remain null with coverage; they are never zero-cost measurements.
|
|
172
|
+
Time/token/cost summaries cover observed worker attempts, not unmeasured coordinator
|
|
173
|
+
work, human waiting, account-wide spend or provider billing. Monetary/duration
|
|
174
|
+
limits with missing measurements fail closed before another dispatch, but cannot
|
|
175
|
+
interrupt a provider request already running. Hooks are workflow controls, not a
|
|
176
|
+
hostile-code filesystem sandbox: shell/MCP writers still require independent diff
|
|
177
|
+
review and objective checks.
|
|
178
|
+
|
|
179
|
+
The optional external `nccgs run --supervise` launcher waits for managed jobs and
|
|
180
|
+
performs bounded same-session continuation with cumulative cost tracking. It cannot
|
|
181
|
+
invent missing native usage or release uncertain provider dispatches. A supervised
|
|
182
|
+
session may yield while that external owner waits; an unsupervised premature Stop
|
|
183
|
+
receives a bounded reminder to collect pending jobs. DONE still rejects pending jobs.
|
|
184
|
+
|
|
185
|
+
This preview is ready for a fresh Claude Code pilot after framework fixture tests.
|
|
186
|
+
It is not a completed real-runtime benchmark. Contaminated Vector Survivors CP0
|
|
187
|
+
must not be scored or presented as an independent NCCGS implementation.
|
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
# Task autonomy
|
|
2
|
+
|
|
3
|
+
An NCCGS assignment preauthorizes the ordinary actions needed to finish that task:
|
|
4
|
+
implementation, testing, tool execution, Bash/PowerShell, report generation,
|
|
5
|
+
temporary files and required operations outside the Unity/project directory.
|
|
6
|
+
Do not ask again merely because a command writes a report, uses shell redirection,
|
|
7
|
+
changes directory, invokes an installed tool elsewhere, or accesses an authorized
|
|
8
|
+
external task output. Choose routine implementation details and continue to a
|
|
9
|
+
tested result and an honest report. Ask for missing requirements only when they
|
|
10
|
+
cannot be resolved from the task; do not disguise a repeated permission request
|
|
11
|
+
as clarification.
|
|
12
|
+
|
|
13
|
+
## Activate the native session
|
|
14
|
+
|
|
15
|
+
The external user/operator starts `nccgs run --project PATH -- [Claude arguments]`.
|
|
16
|
+
With `.nccgs/autonomy.json` mode `full`, the launcher passes Claude Code's supported
|
|
17
|
+
`--permission-mode bypassPermissions` and inherits installed tools and model routes.
|
|
18
|
+
It does not install a restrictive Bash allowlist. The native mode cannot reliably
|
|
19
|
+
be activated by a project `permissions.defaultMode` setting in current Claude Code.
|
|
20
|
+
An already-running session must be resumed through the launcher by its external
|
|
21
|
+
operator; do not nest another Claude CLI to change your own permissions.
|
|
22
|
+
|
|
23
|
+
New installations select `full`. Updates preserve the recorded setting, and older
|
|
24
|
+
installations without this file retain `guarded` until the owner requests
|
|
25
|
+
`update --autonomy full`. `guarded` uses native default permissions. Installation
|
|
26
|
+
or configuration alone is not evidence that a session executed in the selected mode.
|
|
27
|
+
|
|
28
|
+
## Scope and evidence
|
|
29
|
+
|
|
30
|
+
Task authority covers the assigned objective, including its external paths. Use
|
|
31
|
+
the existing reserved implementation worker for shell operations outside the
|
|
32
|
+
project; no extra user approval is needed for the location. Declare its intended
|
|
33
|
+
external outputs in the dispatch prompt and handoff. Do not overlap other writers.
|
|
34
|
+
Audit/plan/review source restrictions, dispatch budgets and independent review stay
|
|
35
|
+
effective. The coordinator can record reports through the task CLI, including
|
|
36
|
+
quoted semicolons and multiline summary text. When outside the project, use the
|
|
37
|
+
absolute task CLI path and `--project` pointing to the original project.
|
|
38
|
+
|
|
39
|
+
Hooks remain anchored to the starting project so external commands do not move
|
|
40
|
+
the ledger, bypass gates or lose provider observations. External files are not
|
|
41
|
+
automatically included in the project source snapshot: copy required evidence
|
|
42
|
+
into `.nccgs/evidence/` and record its source path/hash. Do not claim external
|
|
43
|
+
source coverage or performance merely because the shell command returned zero.
|
|
44
|
+
|
|
45
|
+
This mode does not disable destructive-action guards, vendored-skill protection,
|
|
46
|
+
immutable evidence, external sign-off requirements, or DONE checks. It does not
|
|
47
|
+
turn a missing test into PASS or authorize an unrelated change. Explicit deny/ask
|
|
48
|
+
rules and organization, OS and sandbox restrictions can still block an operation.
|
|
49
|
+
Report the actual restriction; do not claim unrestricted machine access or try
|
|
50
|
+
another command solely to evade a rejection. No user/global Claude settings are
|
|
51
|
+
rewritten by the launcher.
|
|
52
|
+
|
|
53
|
+
Wait for launched verification jobs and collect their actual result before ending
|
|
54
|
+
a print-mode session. Task autonomy does not fix the provider's background-job
|
|
55
|
+
lifetime or make a pending job complete.
|
|
@@ -1,5 +1,9 @@
|
|
|
1
1
|
# Executable task gates
|
|
2
2
|
|
|
3
|
+
This document specifies compatibility/schema-1 gates. For standalone/schema-2
|
|
4
|
+
contracts and payloads use [the standalone gate guide](standalone-gates.md) and
|
|
5
|
+
[standalone execution](standalone.md). Existing tasks retain their original gates.
|
|
6
|
+
|
|
3
7
|
Run the installed tool as `node .claude/nccgs/tools/task.mjs COMMAND ...`, or use
|
|
4
8
|
`nccgs task COMMAND ...`. Paths in evidence are project-relative POSIX paths.
|
|
5
9
|
|
|
@@ -9,17 +13,24 @@ Run the installed tool as `node .claude/nccgs/tools/task.mjs COMMAND ...`, or us
|
|
|
9
13
|
to the hash of the reviewed project.yaml. A missing/stale mapping blocks task init.
|
|
10
14
|
Optional `--requirements FILE.json` adds the same arrays for this task. These
|
|
11
15
|
manifests are hashed; changing obligations invalidates the current task.
|
|
12
|
-
2. `task init --id feature-1 --plan .nccgs/features/feature-1.md
|
|
16
|
+
2. `task init --id feature-1 --plan .nccgs/features/feature-1.md --risk STANDARD`
|
|
17
|
+
creates a task. Risk is FAST, STANDARD, or CONTROLLED; omission defaults to
|
|
18
|
+
STANDARD. The task pins `risk`, `implementationRoute` (class/model/effort), and
|
|
19
|
+
`reviewBudget` (`used`, `limit`, `extensions`) from `contract.taskTiers`.
|
|
13
20
|
3. `task bind --id feature-1 --session SESSION_ID` connects hook telemetry to it.
|
|
14
21
|
4. `task snapshot` prints the current source fingerprint. It covers files, names,
|
|
15
|
-
additions, and deletions, including untracked files.
|
|
16
|
-
|
|
17
|
-
are
|
|
22
|
+
additions, and deletions, including untracked files. Top-level generated roots
|
|
23
|
+
in pipeline.json and explicit Build/Builds output subdirectories from
|
|
24
|
+
`.nccgs/inputs.json` are excluded. Requirements, decisions and registered canon
|
|
25
|
+
are pinned separately and included in the snapshot. The plan and evidence files
|
|
26
|
+
retain their own hashes. Never place source or canon in an excluded directory.
|
|
18
27
|
5. `task ref --file .nccgs/evidence/check.log` prints a `{path, sha256}` reference.
|
|
19
28
|
6. `task record --id feature-1 --stage implementation --evidence .nccgs/evidence/implementation.json`
|
|
20
29
|
records implementation with `snapshot`, `runtime` (artifact reference), and
|
|
21
30
|
nonempty `checks: [{name, status: "PASS", evidence: {path, sha256}}]`.
|
|
22
|
-
7.
|
|
31
|
+
7. Run `task check --id feature-1 --target review` before invoking Fable; proceed
|
|
32
|
+
only when allowed. The bound-task pre-agent guard also checks readiness and
|
|
33
|
+
remaining rounds. Record `review` with `snapshot`, `runtime`, `verdict`, `criteriaMet`, `findings`
|
|
23
34
|
(each has severity/status), and `evidence` referencing the independent review.
|
|
24
35
|
Only PASS with criteriaMet true and no unresolved BLOCKER/HIGH unlocks reporting.
|
|
25
36
|
Corrections require fresh implementation evidence and independent review.
|
|
@@ -38,12 +49,50 @@ Run the installed tool as `node .claude/nccgs/tools/task.mjs COMMAND ...`, or us
|
|
|
38
49
|
11. `task check --id feature-1 --target done` and `task close --id feature-1` enforce
|
|
39
50
|
the entire chain. Changed source, plan, policy, or evidence invalidates old gates.
|
|
40
51
|
|
|
41
|
-
|
|
52
|
+
Initialize once; remediation stays on the same task. Recording an upstream stage
|
|
53
|
+
clears every downstream stage and preserves history.
|
|
54
|
+
It does not reset the review-round counter.
|
|
42
55
|
The `check` command is read-only. A failed check exits nonzero and names the blocker.
|
|
43
56
|
Stop-hook DONE checks are supplemental; task close is the canonical closing action.
|
|
44
57
|
Waiting for Astra or the user is a valid stopping point, not a reason to keep an
|
|
45
58
|
agent running. Do not write task JSON or external approvals through Write/Edit.
|
|
46
59
|
|
|
60
|
+
## Review rounds and explicit extension
|
|
61
|
+
|
|
62
|
+
The default limits are FAST 2, STANDARD 3, and CONTROLLED 3 recorded review rounds.
|
|
63
|
+
Each valid review record consumes one round, including CHANGES_REQUIRED or BLOCKED.
|
|
64
|
+
Malformed records or invalid runtime evidence are rejected without consuming a
|
|
65
|
+
round. A PASS at the limit can proceed; the limit blocks the next review, not an
|
|
66
|
+
already valid result. `task check` reports used, limit, and remaining rounds.
|
|
67
|
+
|
|
68
|
+
Keep only one Fable review in flight per task; do not launch concurrent duplicate
|
|
69
|
+
reviewers. The pre-dispatch guard checks the current recorded budget but does not
|
|
70
|
+
reserve a round. Enforced limits count accepted review records, not parallel model
|
|
71
|
+
invocations, tokens, or cost. The single-review-in-flight rule is orchestration
|
|
72
|
+
policy, not an atomic execution cap.
|
|
73
|
+
|
|
74
|
+
When no usable PASS remains and the budget is exhausted, stop with the current
|
|
75
|
+
evidence, unresolved findings, and next decision. Do not silently reset the task,
|
|
76
|
+
lower its risk, waive findings, or treat exhaustion as success. Re-recording
|
|
77
|
+
implementation preserves the counter even though it clears stale review stages.
|
|
78
|
+
|
|
79
|
+
After an explicit decision, an external operator can run:
|
|
80
|
+
|
|
81
|
+
```text
|
|
82
|
+
task extend-review --id feature-1 --reason "Approved another round to verify the save migration fix" --rounds 1
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
Use the command prefix described above. `--reason` must be nonempty; `--rounds`
|
|
86
|
+
must be a positive integer no greater than that tier's base limit per extension.
|
|
87
|
+
The command refuses a Claude Code session environment and records the timestamp,
|
|
88
|
+
reason, added rounds, old/new limits, and history event. An extension authorizes
|
|
89
|
+
more review attempts, not new product scope or acceptance. It does not supply PASS.
|
|
90
|
+
|
|
91
|
+
Context and elapsed-time budgets are advisory planning limits. Review-round counts
|
|
92
|
+
are enforced by the task tool. Older tasks without risk, route, or review budget
|
|
93
|
+
must be replaced with a new task and fresh stage evidence; old sign-offs do not
|
|
94
|
+
silently carry over. Preserve the old record as history.
|
|
95
|
+
|
|
47
96
|
## Runtime observations
|
|
48
97
|
|
|
49
98
|
Agent lifecycle and Agent-tool hooks capture available native model telemetry,
|
|
@@ -58,7 +107,24 @@ taskId, snapshot, capturedAt, sessionId, agentId, observedModel, observedEffort,
|
|
|
58
107
|
modelsUsed (if swapped), status `completed`, and provenance with kind `provider-log`
|
|
59
108
|
or `operator-observation` and a nonempty reference. It must describe the actual run,
|
|
60
109
|
not settings or the agent's requested model. The tool hashes the source and emits
|
|
61
|
-
an observation reference for the stage record.
|
|
110
|
+
an observation reference for the stage record. Both observation and raw bytes are
|
|
111
|
+
stored under `.nccgs/evidence/runtime/` and must accompany a clean project copy.
|
|
112
|
+
Unknown effort remains blocked. Use `task migrate-evidence --id ID` to copy existing
|
|
113
|
+
legacy runtime evidence without deleting originals or changing run metadata.
|
|
114
|
+
|
|
115
|
+
## Input freshness and bounded verification
|
|
116
|
+
|
|
117
|
+
Before init, map all governing inputs in `.nccgs/inputs.json`; requirements.yaml
|
|
118
|
+
and the decisions tree are always pinned. JSON mapping is explicit: arbitrary
|
|
119
|
+
YAML and Markdown links are not automatically followed. Do not register generated
|
|
120
|
+
task/evidence folders as canonical inputs. Configure per-check `timeoutMs` and
|
|
121
|
+
explicit build output subdirectories in the same manifest.
|
|
122
|
+
|
|
123
|
+
After an approved input change, or for a 1.4 task lacking input pins, run
|
|
124
|
+
`task refresh-inputs --id ID --reason TEXT`. This archives the old state, clears
|
|
125
|
+
stages, retains review consumption, and requires fresh observations/checks and
|
|
126
|
+
approvals. It does not itself approve changed requirements. Closed legacy records
|
|
127
|
+
remain historical. See the installed 1.4.1 migration guide in the package docs.
|
|
62
128
|
|
|
63
129
|
Local files are an auditable workflow boundary, not cryptographic identity or a
|
|
64
130
|
defense against a person who can rewrite the tool and all evidence. Astra/PO
|
|
@@ -1,11 +1,29 @@
|
|
|
1
1
|
# Generated routing — edit pipeline.json, then compile-policy.mjs
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
NCCGS runs in Claude Code. New installations use standalone execution; model aliases describe requested routes, not observed provider versions. Model selection in project.yaml and execution obligations in execution.json are separate policies.
|
|
4
|
+
|
|
5
|
+
| Stage | Owner/model | Effort | Maximum review rounds |
|
|
6
|
+
|---|---|---|---|
|
|
7
|
+
| Intake, plan and documents | Claude Code | Routed by task | — |
|
|
8
|
+
| FAST implementation | sonnet | medium | 2 |
|
|
9
|
+
| STANDARD implementation | opus | medium | 3 |
|
|
10
|
+
| CONTROLLED implementation | opus | high | 3 |
|
|
11
|
+
| Independent review | opus | high | Shared task budget |
|
|
12
|
+
| External review | Off by default; advisory or required when configured | External evidence | — |
|
|
13
|
+
| Acceptance | Contract-defined Product Owner decision | Explicit when required | — |
|
|
14
|
+
|
|
15
|
+
FAST: nccgs-fast-implementer. STANDARD: the relevant domain programmer. CONTROLLED: nccgs-unity-implementer with the affected specialists. Use nccgs-task-scout only for a bounded missing fact; it does not replace review.
|
|
16
|
+
|
|
17
|
+
A task-wide review budget survives implementation retries and continuation. Exhaustion stops the loop; a reasoned extension is required. Report is available for blocked/partial work and is not approval. Missing runtime effort remains unknown unless supplied by observed evidence.
|
|
18
|
+
|
|
19
|
+
Existing installations without execution.json retain compatibility execution: Astra plan/document ownership, required external check and Product Owner acceptance. Legacy astra-claude routes remain unchanged:
|
|
20
|
+
|
|
21
|
+
| Risk | Model | Effort |
|
|
4
22
|
|---|---|---|
|
|
5
|
-
|
|
|
6
|
-
|
|
|
7
|
-
|
|
|
8
|
-
|
|
9
|
-
|
|
23
|
+
| FAST | sonnet | medium |
|
|
24
|
+
| STANDARD | claude-opus-5-5 | medium |
|
|
25
|
+
| CONTROLLED | claude-opus-5-5 | high |
|
|
26
|
+
|
|
27
|
+
Legacy independent review: fable/high. Updating never converts historical approval into a current approval. Explicit migration preserves tasks and consumed budgets.
|
|
10
28
|
|
|
11
|
-
Default active agents:
|
|
29
|
+
Default active agents: 13. Full library: 47.
|