code-factory-2-forge 0.6.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (35) hide show
  1. code_factory_2_forge-0.6.0/LICENSE-APACHE +18 -0
  2. code_factory_2_forge-0.6.0/LICENSE-MIT +21 -0
  3. code_factory_2_forge-0.6.0/NOTICE +6 -0
  4. code_factory_2_forge-0.6.0/PKG-INFO +308 -0
  5. code_factory_2_forge-0.6.0/README.md +293 -0
  6. code_factory_2_forge-0.6.0/code_factory_2_forge.egg-info/PKG-INFO +308 -0
  7. code_factory_2_forge-0.6.0/code_factory_2_forge.egg-info/SOURCES.txt +33 -0
  8. code_factory_2_forge-0.6.0/code_factory_2_forge.egg-info/dependency_links.txt +1 -0
  9. code_factory_2_forge-0.6.0/code_factory_2_forge.egg-info/entry_points.txt +2 -0
  10. code_factory_2_forge-0.6.0/code_factory_2_forge.egg-info/requires.txt +4 -0
  11. code_factory_2_forge-0.6.0/code_factory_2_forge.egg-info/top_level.txt +1 -0
  12. code_factory_2_forge-0.6.0/forgeline/__init__.py +14 -0
  13. code_factory_2_forge-0.6.0/forgeline/adapters.py +56 -0
  14. code_factory_2_forge-0.6.0/forgeline/attribution.py +71 -0
  15. code_factory_2_forge-0.6.0/forgeline/cli.py +101 -0
  16. code_factory_2_forge-0.6.0/forgeline/demo.py +83 -0
  17. code_factory_2_forge-0.6.0/forgeline/demo_learning.py +44 -0
  18. code_factory_2_forge-0.6.0/forgeline/gates/__init__.py +4 -0
  19. code_factory_2_forge-0.6.0/forgeline/gates/adversary.py +45 -0
  20. code_factory_2_forge-0.6.0/forgeline/gates/judge.py +26 -0
  21. code_factory_2_forge-0.6.0/forgeline/gates/qa_audit.py +134 -0
  22. code_factory_2_forge-0.6.0/forgeline/gates/reverse_classical.py +118 -0
  23. code_factory_2_forge-0.6.0/forgeline/gates/runtime_smoke.py +194 -0
  24. code_factory_2_forge-0.6.0/forgeline/gates/skill_check.py +14 -0
  25. code_factory_2_forge-0.6.0/forgeline/intent_thread.py +76 -0
  26. code_factory_2_forge-0.6.0/forgeline/learning.py +116 -0
  27. code_factory_2_forge-0.6.0/forgeline/orchestrator.py +263 -0
  28. code_factory_2_forge-0.6.0/forgeline/refinement.py +76 -0
  29. code_factory_2_forge-0.6.0/forgeline/run_store.py +43 -0
  30. code_factory_2_forge-0.6.0/forgeline/skill_memory.py +52 -0
  31. code_factory_2_forge-0.6.0/forgeline/ssat.py +93 -0
  32. code_factory_2_forge-0.6.0/forgeline/states.py +41 -0
  33. code_factory_2_forge-0.6.0/pyproject.toml +27 -0
  34. code_factory_2_forge-0.6.0/setup.cfg +4 -0
  35. code_factory_2_forge-0.6.0/tests/test_forgeline.py +565 -0
@@ -0,0 +1,18 @@
1
+ Apache License
2
+ Version 2.0, January 2004
3
+ http://www.apache.org/licenses/
4
+
5
+ Licensed under the Apache License, Version 2.0 (the "License");
6
+ you may not use this file except in compliance with the License.
7
+ You may obtain a copy of the License at
8
+
9
+ http://www.apache.org/licenses/LICENSE-2.0
10
+
11
+ Unless required by applicable law or agreed to in writing, software
12
+ distributed under the License is distributed on an "AS IS" BASIS,
13
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
14
+ See the License for the specific language governing permissions and
15
+ limitations under the License.
16
+
17
+ Full license text: http://www.apache.org/licenses/LICENSE-2.0.txt
18
+ Copyright 2026 WizeMe.APP
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 WizeMe.APP
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,6 @@
1
+ ForgeLine — autonomous software factory outer loop
2
+ Copyright 2026 WizeMe.APP
3
+
4
+ Dual-licensed under either of Apache License 2.0 or MIT license at your option.
5
+ Concepts drawn from the Agentic SDLC / Spec-Driven Engineering literature
6
+ (SSAT, adversarial review, architecture-as-CI-gate, skill memory).
@@ -0,0 +1,308 @@
1
+ Metadata-Version: 2.4
2
+ Name: code-factory-2-forge
3
+ Version: 0.6.0
4
+ Summary: ForgeLine — the autonomous outer loop for AI software factories. A CLI-backed state machine that drives intent -> spec -> plan -> code -> adversarial review -> ship, orchestrating SpecLine (spec governance) and Harness Software Factory (compiled decisions), with self-improving skills and architecture-as-a-CI-gate.
5
+ License-Expression: MIT OR Apache-2.0
6
+ Requires-Python: >=3.11
7
+ Description-Content-Type: text/markdown
8
+ License-File: LICENSE-APACHE
9
+ License-File: LICENSE-MIT
10
+ License-File: NOTICE
11
+ Requires-Dist: PyYAML>=6.0
12
+ Provides-Extra: dev
13
+ Requires-Dist: pytest>=8.0; extra == "dev"
14
+ Dynamic: license-file
15
+
16
+ # ForgeLine 🔨
17
+
18
+ **The autonomous outer loop for AI software factories.** ForgeLine is the tier
19
+ above spec-writing and code-generation: a **CLI-backed state machine** that
20
+ carries a feature from vague intent to shipped code through confidence gates,
21
+ runs a generate → **adversarial review** → refine loop, enforces
22
+ **architecture as a hard CI gate**, and gets smarter every run via an evolving
23
+ skill memory.
24
+
25
+ It orchestrates its siblings — [SpecLine](../specline) (spec governance) and
26
+ [Harness Software Factory](../harness-factory) (compiled decisions) — into one
27
+ pipeline you drive from Claude Code or Codex.
28
+
29
+ > Stop writing the code. Build — and supervise — the machine that writes it.
30
+
31
+ ## Workflow at a glance
32
+
33
+ ```mermaid
34
+ flowchart LR
35
+ A["Feature intent"] --> B["Expand use cases"]
36
+ B --> C["Human confidence gate"]
37
+ C --> D["Architect with SSAT"]
38
+ D --> E["Scaffold from architecture"]
39
+ E --> F["Fill implementation"]
40
+ F --> G["Adversarial review"]
41
+ G -->|"findings"| H["Refine and record lesson"]
42
+ H --> F
43
+ G -->|"pass"| I["Architecture CI gate"]
44
+ I --> J["Reverse-classical test verification"]
45
+ J --> K["Runtime smoke gate"]
46
+ K --> L["Ship with receipt"]
47
+ ```
48
+
49
+ ```
50
+ INTENT ─► EXPAND ─► ARCHITECT(SSAT) ─► SCAFFOLD ─► FILL ─┐
51
+ ▲ │
52
+ │ ┌── grumpy adversary + judge + arch-erosion ──┤
53
+ └── refine┤ (records a skill lesson) │ pass
54
+ └──────────────◄─────────────────────────────┘
55
+ │
56
+ ARCH-CI-GATE ─► SHIP
57
+ ```
58
+
59
+ ## Install (any OS, 60 seconds, no API keys)
60
+
61
+ **One command, works everywhere:**
62
+ ```bash
63
+ python install.py # Windows / macOS / Linux
64
+ ```
65
+ or double-click `install.sh` (macOS/Linux) / `install.bat` (Windows), or:
66
+ ```bash
67
+ pip install -e ".[dev]"
68
+ ```
69
+
70
+ Then:
71
+ ```bash
72
+ forge demo # 60-sec: watch it catch its own bad output, learn, ship
73
+ forge init
74
+ forge agent claude # wire your agent (see full list below)
75
+ forge status <feature> # the state machine names the ONE next action
76
+ pytest -q # 29 tests
77
+ ```
78
+
79
+ ## Works with every major coding agent
80
+
81
+ One command wires the entry-point skill wherever your agent reads it:
82
+
83
+ ```bash
84
+ forge agent claude # CLAUDE.md + .claude/skills/forge.md
85
+ forge agent codex # AGENTS.md
86
+ forge agent opencode # AGENTS.md + .opencode/forge.md
87
+ forge agent cursor # .cursorrules + .cursor/rules/forge.md
88
+ forge agent aider # CONVENTIONS.md
89
+ forge agent gemini # GEMINI.md
90
+ forge agent windsurf # .windsurfrules
91
+ forge agent generic # AGENT.md (unknown tools fall back to AGENTS.md)
92
+ ```
93
+ The contract is plain text; any harness that reads a project file can run the factory.
94
+
95
+ ## What makes it more than a wrapper
96
+
97
+ **A real state machine with confidence gates.** Eight states, legal-transition
98
+ enforcement (`E_ILLEGAL_TRANSITION`), and two **human confidence gates**
99
+ (use-case signoff, architecture signoff) where a person must approve before
100
+ agents proceed. Every transition is a hash-sealed receipt on disk — so the
101
+ loop survives context resets (the disk is the truth).
102
+
103
+ **SSAT — architecture as code, not a document.** A Semantic Software
104
+ Architecture Tree (YAML) declares modules, signatures, allowed dependency
105
+ edges, and invariants. The *same artifact* generates the scaffold AND serves
106
+ as the CI gate: `check_erosion()` detects signature drift (`E_SIG_DRIFT`),
107
+ illegal dependencies (`E_ILLEGAL_DEP`), and invariant violations
108
+ (`E_INVARIANT`) — structural erosion caught before merge.
109
+
110
+ **The grumpy adversary.** A review agent that *assumes your code is broken and
111
+ insecure* and makes the generator prove otherwise. Executable heuristics catch
112
+ eval/exec/shell-injection/hard-coded-secrets/bare-excepts, and it refuses to
113
+ pass any code that ships without tests (`A_NO_PROOF` — "prove it works"). No
114
+ LLM required to be useful; an LLM adversary layers behind the same interface.
115
+
116
+ **A self-improving skill flywheel.** Every gate failure records a structured
117
+ lesson to `skills/lessons.jsonl`. Lessons are injected into the next attempt's
118
+ context, and lessons seen ≥3 times **graduate into hard constraints**
119
+ (conventions-into-constraints) — promoted straight into SSAT invariants. The
120
+ factory literally learns your team's rules from its own mistakes.
121
+
122
+ **Decision handoff to the factory.** Specs carrying a decision table route to
123
+ HSF for one-time compilation into gated, deterministic code — the outer loop
124
+ knows the difference between *tissue* (agents write it) and *decisions* (the
125
+ factory compiles them, never improvised twice).
126
+
127
+ ## The entry-point skill (Claude Code / Codex)
128
+
129
+ `forge agent claude` writes `CLAUDE.md` + `.claude/skills/forge.md`;
130
+ `forge agent codex` writes `AGENTS.md`. The contract turns any agent into a
131
+ disciplined factory worker: read the skill, run `forge status`, do the one
132
+ named phase, run its gate, repeat — context resetting between phases. The
133
+ agent never free-codes; it advances a state machine. That's the whole point.
134
+
135
+ ## Commands
136
+
137
+ ```
138
+ forge init scaffold the factory
139
+ forge agent claude|codex wire the entry-point skill
140
+ forge status <feature> current state + the ONE next action
141
+ forge expand <feature> draft use cases (→ human gate)
142
+ forge architect <feature> <ssat> generate scaffold from architecture-as-code
143
+ forge review <feature> <ssat> judge + grumpy adversary + arch erosion (refine loop)
144
+ forge arch-gate <feature> <ssat> architecture CI gate
145
+ forge verify-tests <feature> <ssat> prove smoke checks fail on generated stubs
146
+ forge smoke <feature> runtime behavior gate
147
+ forge ship <feature> seal it
148
+ forge handoff <feature> <spec> route decision tables to HSF
149
+ forge lessons show the skill memory + promotable constraints
150
+ forge demo the 60-second story
151
+ ```
152
+
153
+ ## v0.3 — the Intent Thread (PRD → production traceability)
154
+
155
+ ForgeLine now consumes SpecLine's sealed **Intent Envelope** and verifies the
156
+ FINAL shipped code against the ORIGINAL rationalized intent — not the plan, not
157
+ the drifted spec, but the intent that was sealed at the plan gate. This closes
158
+ the last translation-loss gap: *does the shipped thing actually satisfy what we
159
+ rationalized we wanted?*
160
+
161
+ Every assumption SpecLine surfaced (auth exists, currency is single-source,
162
+ dependencies can fail) becomes a **checkable obligation**. The `ship` gate
163
+ blocks if the code shows no handling for an assumption the intent depended on,
164
+ and the sealed intent hash proves the intent wasn't quietly swapped underneath
165
+ the build. Full PRD-to-production traceability, enforced — not documented.
166
+
167
+ ## v0.2 — deeper QA + recursive learning
168
+
169
+ ForgeLine now audits quality quantitatively and **learns from its own runs**:
170
+
171
+ **Deep QA audit** (`forge qa`) grades every build on coverage-intent (do tests
172
+ actually call the functions?), cyclomatic complexity, a scored security surface
173
+ (eval/exec/shell/secrets/weak-crypto), and documentation — a composite A–F grade
174
+ that gates shipping. A pretty build with untested, over-complex, or insecure
175
+ code cannot pass.
176
+
177
+ **Recursive learning kernel** (`forge policy`, `forge demo-learning`) closes the
178
+ loop the skill memory only started:
179
+ - **observe** — a failure is recorded (as before)
180
+ - **promote** — a failure seen ≥3× becomes an *enforced active constraint*
181
+ - **validate** — when that constraint catches the same failure again, it's marked
182
+ effective; the policy tracks prevention counts
183
+ - **self-prune** — a promoted rule that never fires again goes to *probation*, so
184
+ the policy doesn't ossify
185
+
186
+ The factory's own run history becomes its QA policy, and that policy is measured,
187
+ not assumed. `forge demo-learning` shows the full observe→promote→validate→prune
188
+ cycle in 15 seconds.
189
+
190
+ **Escalating refine loop** — the review loop now gets stricter each attempt
191
+ (normal → elevated "fix all findings" → final "human review required") instead
192
+ of a flat retry cap.
193
+
194
+ ## The three-repo factory
195
+
196
+ | Repo | Tier | Owns |
197
+ |---|---|---|
198
+ | **ForgeLine** | outer loop | intent→ship state machine, adversarial gates, skill flywheel, arch-as-CI |
199
+ | **SpecLine** | spec governance | EARS specs, atomic task packets, token-lean context, intent-drift guard |
200
+ | **HSF** | decision compiler | ordered business rules → gated deterministic code, zero tokens/decision |
201
+
202
+ Same doctrine at every tier: gate everything, receipts or it didn't happen,
203
+ compile what shouldn't be reasoned twice.
204
+
205
+ ## License
206
+
207
+ Dual-licensed under either **Apache-2.0 OR MIT** at your option — the
208
+ permissive standard for broad adoption. Pick whichever your project prefers.
209
+
210
+ ---
211
+
212
+ ## v0.4 — Runtime Smoke Gate (behavior-by-inspection)
213
+
214
+ ForgeLine's earlier gates all verify code against *specifications* — the judge
215
+ checks consistency, the QA audit grades static quality, the intent thread proves
216
+ the shipped code honors the sealed envelope. That is correctness *by construction*.
217
+
218
+ None of it answers the question a per-PR preview deployment answers: **does the
219
+ built thing actually RUN and behave correctly?** A change can pass every static
220
+ gate and still crash on import or produce the wrong output. v0.4 closes that gap
221
+ at solo-builder scale — no ephemeral per-PR environments required.
222
+
223
+ ### New state + gate
224
+
225
+ The SDLC gains a `SMOKED` state between `ARCH_GATED` and `SHIPPED`:
226
+
227
+ ```
228
+ … → reviewed → arch_gated → smoked → shipped
229
+ ```
230
+
231
+ `forge smoke <feature>` runs every behavioral check declared in
232
+ `smoke/<feature>.json` in an **isolated subprocess with a timeout**, and blocks
233
+ ship on any runtime failure. `ship` now refuses to run until the smoke gate has
234
+ passed — you cannot ship unverified runtime behavior.
235
+
236
+ ### The smoke manifest
237
+
238
+ `smoke/<feature>.json` makes "correct runtime behavior" a reviewed artifact:
239
+
240
+ ```json
241
+ {
242
+ "checks": [
243
+ {
244
+ "name": "formatter_runtime",
245
+ "kind": "python",
246
+ "run": "from slices.notifier.formatter import format_message\nassert format_message({'kind':'ping','text':'hi'}) == 'ping: hi'\nprint('OK')",
247
+ "expect_exit": 0,
248
+ "expect_stdout": "OK",
249
+ "timeout_s": 15
250
+ }
251
+ ]
252
+ }
253
+ ```
254
+
255
+ `kind` is `python` (runs a snippet) or `command` (runs a shell command). Each
256
+ check asserts an exit code and, optionally, a stdout substring.
257
+
258
+ ### Fail-closed by design
259
+
260
+ - **No manifest** → BLOCK. You cannot ship runtime behavior you never verified.
261
+ - **Empty manifest** → BLOCK. A manifest with zero checks ships nothing verified.
262
+ - **Any check fails / times out / crashes** → BLOCK, with the failing check named
263
+ and the last lines of its output captured as a receipt.
264
+
265
+ ### Why this and not full preview deployments
266
+
267
+ Per-PR ephemeral environments are real DevOps infrastructure (provisioning,
268
+ isolation, teardown, secrets, cost) that pays off at *team* scale with many
269
+ parallel PRs and human reviewers. The smoke gate captures the core value —
270
+ *verify the artifact runs and behaves before merge* — deterministically, at the
271
+ cost of a subprocess. When team scale justifies it, the manifest model extends
272
+ naturally to spinning real environments; the behavioral contract is already written.
273
+ ## v0.6 - Reverse-Classical test verification
274
+
275
+ ForgeLine now refuses tests that prove nothing.
276
+
277
+ `forge verify-tests <feature> <ssat.yaml>` regenerates the SSAT scaffold into an
278
+ isolated temp root and runs every behavioral smoke check against those empty
279
+ stubs. Each behavioral check must fail on the stub before ForgeLine trusts it
280
+ against the real implementation. A check that passes against an empty stub is
281
+ classified as `HOLLOW_TEST` and blocks the feature.
282
+
283
+ The ordering is now:
284
+
285
+ ```text
286
+ ... -> reviewed -> arch_gated -> tests_verified -> smoked -> shipped
287
+ ```
288
+
289
+ Structural checks such as imports may declare `"must_fail_on_stub": false` in
290
+ `smoke/<feature>.json`, but the field defaults to `true`. Omission never buys
291
+ leniency, and a manifest where every check is exempt blocks as
292
+ `HOLLOW_MANIFEST`.
293
+
294
+ This gate is deterministic: no model, no git snapshot, no new runtime service.
295
+ It reuses the same SSAT scaffold generator as the normal `SCAFFOLDED` state, so
296
+ the mutant is the real generated stub.
297
+
298
+ ## Failure attribution and refinement
299
+
300
+ ForgeLine 0.5 reports review, architecture, QA, smoke, and intent failures at
301
+ their smallest actionable unit. `forge qa --root .` includes function-level
302
+ metrics and attribution in its JSON output.
303
+
304
+ The refinement engine accepts exactly one proposed edit at a time. Structural
305
+ edits precede configuration and parameter changes. An edit is retained only
306
+ when its targeted stage improves and no other stage regresses; rejected edits
307
+ are reverted and written with before/after rates to
308
+ `.forge/rejection_ledger.jsonl`. Two consecutive non-wins stop the loop.
@@ -0,0 +1,293 @@
1
+ # ForgeLine 🔨
2
+
3
+ **The autonomous outer loop for AI software factories.** ForgeLine is the tier
4
+ above spec-writing and code-generation: a **CLI-backed state machine** that
5
+ carries a feature from vague intent to shipped code through confidence gates,
6
+ runs a generate → **adversarial review** → refine loop, enforces
7
+ **architecture as a hard CI gate**, and gets smarter every run via an evolving
8
+ skill memory.
9
+
10
+ It orchestrates its siblings — [SpecLine](../specline) (spec governance) and
11
+ [Harness Software Factory](../harness-factory) (compiled decisions) — into one
12
+ pipeline you drive from Claude Code or Codex.
13
+
14
+ > Stop writing the code. Build — and supervise — the machine that writes it.
15
+
16
+ ## Workflow at a glance
17
+
18
+ ```mermaid
19
+ flowchart LR
20
+ A["Feature intent"] --> B["Expand use cases"]
21
+ B --> C["Human confidence gate"]
22
+ C --> D["Architect with SSAT"]
23
+ D --> E["Scaffold from architecture"]
24
+ E --> F["Fill implementation"]
25
+ F --> G["Adversarial review"]
26
+ G -->|"findings"| H["Refine and record lesson"]
27
+ H --> F
28
+ G -->|"pass"| I["Architecture CI gate"]
29
+ I --> J["Reverse-classical test verification"]
30
+ J --> K["Runtime smoke gate"]
31
+ K --> L["Ship with receipt"]
32
+ ```
33
+
34
+ ```
35
+ INTENT ─► EXPAND ─► ARCHITECT(SSAT) ─► SCAFFOLD ─► FILL ─┐
36
+ ▲ │
37
+ │ ┌── grumpy adversary + judge + arch-erosion ──┤
38
+ └── refine┤ (records a skill lesson) │ pass
39
+ └──────────────◄─────────────────────────────┘
40
+ │
41
+ ARCH-CI-GATE ─► SHIP
42
+ ```
43
+
44
+ ## Install (any OS, 60 seconds, no API keys)
45
+
46
+ **One command, works everywhere:**
47
+ ```bash
48
+ python install.py # Windows / macOS / Linux
49
+ ```
50
+ or double-click `install.sh` (macOS/Linux) / `install.bat` (Windows), or:
51
+ ```bash
52
+ pip install -e ".[dev]"
53
+ ```
54
+
55
+ Then:
56
+ ```bash
57
+ forge demo # 60-sec: watch it catch its own bad output, learn, ship
58
+ forge init
59
+ forge agent claude # wire your agent (see full list below)
60
+ forge status <feature> # the state machine names the ONE next action
61
+ pytest -q # 29 tests
62
+ ```
63
+
64
+ ## Works with every major coding agent
65
+
66
+ One command wires the entry-point skill wherever your agent reads it:
67
+
68
+ ```bash
69
+ forge agent claude # CLAUDE.md + .claude/skills/forge.md
70
+ forge agent codex # AGENTS.md
71
+ forge agent opencode # AGENTS.md + .opencode/forge.md
72
+ forge agent cursor # .cursorrules + .cursor/rules/forge.md
73
+ forge agent aider # CONVENTIONS.md
74
+ forge agent gemini # GEMINI.md
75
+ forge agent windsurf # .windsurfrules
76
+ forge agent generic # AGENT.md (unknown tools fall back to AGENTS.md)
77
+ ```
78
+ The contract is plain text; any harness that reads a project file can run the factory.
79
+
80
+ ## What makes it more than a wrapper
81
+
82
+ **A real state machine with confidence gates.** Eight states, legal-transition
83
+ enforcement (`E_ILLEGAL_TRANSITION`), and two **human confidence gates**
84
+ (use-case signoff, architecture signoff) where a person must approve before
85
+ agents proceed. Every transition is a hash-sealed receipt on disk — so the
86
+ loop survives context resets (the disk is the truth).
87
+
88
+ **SSAT — architecture as code, not a document.** A Semantic Software
89
+ Architecture Tree (YAML) declares modules, signatures, allowed dependency
90
+ edges, and invariants. The *same artifact* generates the scaffold AND serves
91
+ as the CI gate: `check_erosion()` detects signature drift (`E_SIG_DRIFT`),
92
+ illegal dependencies (`E_ILLEGAL_DEP`), and invariant violations
93
+ (`E_INVARIANT`) — structural erosion caught before merge.
94
+
95
+ **The grumpy adversary.** A review agent that *assumes your code is broken and
96
+ insecure* and makes the generator prove otherwise. Executable heuristics catch
97
+ eval/exec/shell-injection/hard-coded-secrets/bare-excepts, and it refuses to
98
+ pass any code that ships without tests (`A_NO_PROOF` — "prove it works"). No
99
+ LLM required to be useful; an LLM adversary layers behind the same interface.
100
+
101
+ **A self-improving skill flywheel.** Every gate failure records a structured
102
+ lesson to `skills/lessons.jsonl`. Lessons are injected into the next attempt's
103
+ context, and lessons seen ≥3 times **graduate into hard constraints**
104
+ (conventions-into-constraints) — promoted straight into SSAT invariants. The
105
+ factory literally learns your team's rules from its own mistakes.
106
+
107
+ **Decision handoff to the factory.** Specs carrying a decision table route to
108
+ HSF for one-time compilation into gated, deterministic code — the outer loop
109
+ knows the difference between *tissue* (agents write it) and *decisions* (the
110
+ factory compiles them, never improvised twice).
111
+
112
+ ## The entry-point skill (Claude Code / Codex)
113
+
114
+ `forge agent claude` writes `CLAUDE.md` + `.claude/skills/forge.md`;
115
+ `forge agent codex` writes `AGENTS.md`. The contract turns any agent into a
116
+ disciplined factory worker: read the skill, run `forge status`, do the one
117
+ named phase, run its gate, repeat — context resetting between phases. The
118
+ agent never free-codes; it advances a state machine. That's the whole point.
119
+
120
+ ## Commands
121
+
122
+ ```
123
+ forge init scaffold the factory
124
+ forge agent claude|codex wire the entry-point skill
125
+ forge status <feature> current state + the ONE next action
126
+ forge expand <feature> draft use cases (→ human gate)
127
+ forge architect <feature> <ssat> generate scaffold from architecture-as-code
128
+ forge review <feature> <ssat> judge + grumpy adversary + arch erosion (refine loop)
129
+ forge arch-gate <feature> <ssat> architecture CI gate
130
+ forge verify-tests <feature> <ssat> prove smoke checks fail on generated stubs
131
+ forge smoke <feature> runtime behavior gate
132
+ forge ship <feature> seal it
133
+ forge handoff <feature> <spec> route decision tables to HSF
134
+ forge lessons show the skill memory + promotable constraints
135
+ forge demo the 60-second story
136
+ ```
137
+
138
+ ## v0.3 — the Intent Thread (PRD → production traceability)
139
+
140
+ ForgeLine now consumes SpecLine's sealed **Intent Envelope** and verifies the
141
+ FINAL shipped code against the ORIGINAL rationalized intent — not the plan, not
142
+ the drifted spec, but the intent that was sealed at the plan gate. This closes
143
+ the last translation-loss gap: *does the shipped thing actually satisfy what we
144
+ rationalized we wanted?*
145
+
146
+ Every assumption SpecLine surfaced (auth exists, currency is single-source,
147
+ dependencies can fail) becomes a **checkable obligation**. The `ship` gate
148
+ blocks if the code shows no handling for an assumption the intent depended on,
149
+ and the sealed intent hash proves the intent wasn't quietly swapped underneath
150
+ the build. Full PRD-to-production traceability, enforced — not documented.
151
+
152
+ ## v0.2 — deeper QA + recursive learning
153
+
154
+ ForgeLine now audits quality quantitatively and **learns from its own runs**:
155
+
156
+ **Deep QA audit** (`forge qa`) grades every build on coverage-intent (do tests
157
+ actually call the functions?), cyclomatic complexity, a scored security surface
158
+ (eval/exec/shell/secrets/weak-crypto), and documentation — a composite A–F grade
159
+ that gates shipping. A pretty build with untested, over-complex, or insecure
160
+ code cannot pass.
161
+
162
+ **Recursive learning kernel** (`forge policy`, `forge demo-learning`) closes the
163
+ loop the skill memory only started:
164
+ - **observe** — a failure is recorded (as before)
165
+ - **promote** — a failure seen ≥3× becomes an *enforced active constraint*
166
+ - **validate** — when that constraint catches the same failure again, it's marked
167
+ effective; the policy tracks prevention counts
168
+ - **self-prune** — a promoted rule that never fires again goes to *probation*, so
169
+ the policy doesn't ossify
170
+
171
+ The factory's own run history becomes its QA policy, and that policy is measured,
172
+ not assumed. `forge demo-learning` shows the full observe→promote→validate→prune
173
+ cycle in 15 seconds.
174
+
175
+ **Escalating refine loop** — the review loop now gets stricter each attempt
176
+ (normal → elevated "fix all findings" → final "human review required") instead
177
+ of a flat retry cap.
178
+
179
+ ## The three-repo factory
180
+
181
+ | Repo | Tier | Owns |
182
+ |---|---|---|
183
+ | **ForgeLine** | outer loop | intent→ship state machine, adversarial gates, skill flywheel, arch-as-CI |
184
+ | **SpecLine** | spec governance | EARS specs, atomic task packets, token-lean context, intent-drift guard |
185
+ | **HSF** | decision compiler | ordered business rules → gated deterministic code, zero tokens/decision |
186
+
187
+ Same doctrine at every tier: gate everything, receipts or it didn't happen,
188
+ compile what shouldn't be reasoned twice.
189
+
190
+ ## License
191
+
192
+ Dual-licensed under either **Apache-2.0 OR MIT** at your option — the
193
+ permissive standard for broad adoption. Pick whichever your project prefers.
194
+
195
+ ---
196
+
197
+ ## v0.4 — Runtime Smoke Gate (behavior-by-inspection)
198
+
199
+ ForgeLine's earlier gates all verify code against *specifications* — the judge
200
+ checks consistency, the QA audit grades static quality, the intent thread proves
201
+ the shipped code honors the sealed envelope. That is correctness *by construction*.
202
+
203
+ None of it answers the question a per-PR preview deployment answers: **does the
204
+ built thing actually RUN and behave correctly?** A change can pass every static
205
+ gate and still crash on import or produce the wrong output. v0.4 closes that gap
206
+ at solo-builder scale — no ephemeral per-PR environments required.
207
+
208
+ ### New state + gate
209
+
210
+ The SDLC gains a `SMOKED` state between `ARCH_GATED` and `SHIPPED`:
211
+
212
+ ```
213
+ … → reviewed → arch_gated → smoked → shipped
214
+ ```
215
+
216
+ `forge smoke <feature>` runs every behavioral check declared in
217
+ `smoke/<feature>.json` in an **isolated subprocess with a timeout**, and blocks
218
+ ship on any runtime failure. `ship` now refuses to run until the smoke gate has
219
+ passed — you cannot ship unverified runtime behavior.
220
+
221
+ ### The smoke manifest
222
+
223
+ `smoke/<feature>.json` makes "correct runtime behavior" a reviewed artifact:
224
+
225
+ ```json
226
+ {
227
+ "checks": [
228
+ {
229
+ "name": "formatter_runtime",
230
+ "kind": "python",
231
+ "run": "from slices.notifier.formatter import format_message\nassert format_message({'kind':'ping','text':'hi'}) == 'ping: hi'\nprint('OK')",
232
+ "expect_exit": 0,
233
+ "expect_stdout": "OK",
234
+ "timeout_s": 15
235
+ }
236
+ ]
237
+ }
238
+ ```
239
+
240
+ `kind` is `python` (runs a snippet) or `command` (runs a shell command). Each
241
+ check asserts an exit code and, optionally, a stdout substring.
242
+
243
+ ### Fail-closed by design
244
+
245
+ - **No manifest** → BLOCK. You cannot ship runtime behavior you never verified.
246
+ - **Empty manifest** → BLOCK. A manifest with zero checks ships nothing verified.
247
+ - **Any check fails / times out / crashes** → BLOCK, with the failing check named
248
+ and the last lines of its output captured as a receipt.
249
+
250
+ ### Why this and not full preview deployments
251
+
252
+ Per-PR ephemeral environments are real DevOps infrastructure (provisioning,
253
+ isolation, teardown, secrets, cost) that pays off at *team* scale with many
254
+ parallel PRs and human reviewers. The smoke gate captures the core value —
255
+ *verify the artifact runs and behaves before merge* — deterministically, at the
256
+ cost of a subprocess. When team scale justifies it, the manifest model extends
257
+ naturally to spinning real environments; the behavioral contract is already written.
258
+ ## v0.6 - Reverse-Classical test verification
259
+
260
+ ForgeLine now refuses tests that prove nothing.
261
+
262
+ `forge verify-tests <feature> <ssat.yaml>` regenerates the SSAT scaffold into an
263
+ isolated temp root and runs every behavioral smoke check against those empty
264
+ stubs. Each behavioral check must fail on the stub before ForgeLine trusts it
265
+ against the real implementation. A check that passes against an empty stub is
266
+ classified as `HOLLOW_TEST` and blocks the feature.
267
+
268
+ The ordering is now:
269
+
270
+ ```text
271
+ ... -> reviewed -> arch_gated -> tests_verified -> smoked -> shipped
272
+ ```
273
+
274
+ Structural checks such as imports may declare `"must_fail_on_stub": false` in
275
+ `smoke/<feature>.json`, but the field defaults to `true`. Omission never buys
276
+ leniency, and a manifest where every check is exempt blocks as
277
+ `HOLLOW_MANIFEST`.
278
+
279
+ This gate is deterministic: no model, no git snapshot, no new runtime service.
280
+ It reuses the same SSAT scaffold generator as the normal `SCAFFOLDED` state, so
281
+ the mutant is the real generated stub.
282
+
283
+ ## Failure attribution and refinement
284
+
285
+ ForgeLine 0.5 reports review, architecture, QA, smoke, and intent failures at
286
+ their smallest actionable unit. `forge qa --root .` includes function-level
287
+ metrics and attribution in its JSON output.
288
+
289
+ The refinement engine accepts exactly one proposed edit at a time. Structural
290
+ edits precede configuration and parameter changes. An edit is retained only
291
+ when its targeted stage improves and no other stage regresses; rejected edits
292
+ are reverted and written with before/after rates to
293
+ `.forge/rejection_ledger.jsonl`. Two consecutive non-wins stop the loop.