code-factory-2-forge 0.6.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- code_factory_2_forge-0.6.0/LICENSE-APACHE +18 -0
- code_factory_2_forge-0.6.0/LICENSE-MIT +21 -0
- code_factory_2_forge-0.6.0/NOTICE +6 -0
- code_factory_2_forge-0.6.0/PKG-INFO +308 -0
- code_factory_2_forge-0.6.0/README.md +293 -0
- code_factory_2_forge-0.6.0/code_factory_2_forge.egg-info/PKG-INFO +308 -0
- code_factory_2_forge-0.6.0/code_factory_2_forge.egg-info/SOURCES.txt +33 -0
- code_factory_2_forge-0.6.0/code_factory_2_forge.egg-info/dependency_links.txt +1 -0
- code_factory_2_forge-0.6.0/code_factory_2_forge.egg-info/entry_points.txt +2 -0
- code_factory_2_forge-0.6.0/code_factory_2_forge.egg-info/requires.txt +4 -0
- code_factory_2_forge-0.6.0/code_factory_2_forge.egg-info/top_level.txt +1 -0
- code_factory_2_forge-0.6.0/forgeline/__init__.py +14 -0
- code_factory_2_forge-0.6.0/forgeline/adapters.py +56 -0
- code_factory_2_forge-0.6.0/forgeline/attribution.py +71 -0
- code_factory_2_forge-0.6.0/forgeline/cli.py +101 -0
- code_factory_2_forge-0.6.0/forgeline/demo.py +83 -0
- code_factory_2_forge-0.6.0/forgeline/demo_learning.py +44 -0
- code_factory_2_forge-0.6.0/forgeline/gates/__init__.py +4 -0
- code_factory_2_forge-0.6.0/forgeline/gates/adversary.py +45 -0
- code_factory_2_forge-0.6.0/forgeline/gates/judge.py +26 -0
- code_factory_2_forge-0.6.0/forgeline/gates/qa_audit.py +134 -0
- code_factory_2_forge-0.6.0/forgeline/gates/reverse_classical.py +118 -0
- code_factory_2_forge-0.6.0/forgeline/gates/runtime_smoke.py +194 -0
- code_factory_2_forge-0.6.0/forgeline/gates/skill_check.py +14 -0
- code_factory_2_forge-0.6.0/forgeline/intent_thread.py +76 -0
- code_factory_2_forge-0.6.0/forgeline/learning.py +116 -0
- code_factory_2_forge-0.6.0/forgeline/orchestrator.py +263 -0
- code_factory_2_forge-0.6.0/forgeline/refinement.py +76 -0
- code_factory_2_forge-0.6.0/forgeline/run_store.py +43 -0
- code_factory_2_forge-0.6.0/forgeline/skill_memory.py +52 -0
- code_factory_2_forge-0.6.0/forgeline/ssat.py +93 -0
- code_factory_2_forge-0.6.0/forgeline/states.py +41 -0
- code_factory_2_forge-0.6.0/pyproject.toml +27 -0
- code_factory_2_forge-0.6.0/setup.cfg +4 -0
- code_factory_2_forge-0.6.0/tests/test_forgeline.py +565 -0
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
Apache License
|
|
2
|
+
Version 2.0, January 2004
|
|
3
|
+
http://www.apache.org/licenses/
|
|
4
|
+
|
|
5
|
+
Licensed under the Apache License, Version 2.0 (the "License");
|
|
6
|
+
you may not use this file except in compliance with the License.
|
|
7
|
+
You may obtain a copy of the License at
|
|
8
|
+
|
|
9
|
+
http://www.apache.org/licenses/LICENSE-2.0
|
|
10
|
+
|
|
11
|
+
Unless required by applicable law or agreed to in writing, software
|
|
12
|
+
distributed under the License is distributed on an "AS IS" BASIS,
|
|
13
|
+
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
14
|
+
See the License for the specific language governing permissions and
|
|
15
|
+
limitations under the License.
|
|
16
|
+
|
|
17
|
+
Full license text: http://www.apache.org/licenses/LICENSE-2.0.txt
|
|
18
|
+
Copyright 2026 WizeMe.APP
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 WizeMe.APP
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,6 @@
|
|
|
1
|
+
ForgeLine — autonomous software factory outer loop
|
|
2
|
+
Copyright 2026 WizeMe.APP
|
|
3
|
+
|
|
4
|
+
Dual-licensed under either of Apache License 2.0 or MIT license at your option.
|
|
5
|
+
Concepts drawn from the Agentic SDLC / Spec-Driven Engineering literature
|
|
6
|
+
(SSAT, adversarial review, architecture-as-CI-gate, skill memory).
|
|
@@ -0,0 +1,308 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: code-factory-2-forge
|
|
3
|
+
Version: 0.6.0
|
|
4
|
+
Summary: ForgeLine — the autonomous outer loop for AI software factories. A CLI-backed state machine that drives intent -> spec -> plan -> code -> adversarial review -> ship, orchestrating SpecLine (spec governance) and Harness Software Factory (compiled decisions), with self-improving skills and architecture-as-a-CI-gate.
|
|
5
|
+
License-Expression: MIT OR Apache-2.0
|
|
6
|
+
Requires-Python: >=3.11
|
|
7
|
+
Description-Content-Type: text/markdown
|
|
8
|
+
License-File: LICENSE-APACHE
|
|
9
|
+
License-File: LICENSE-MIT
|
|
10
|
+
License-File: NOTICE
|
|
11
|
+
Requires-Dist: PyYAML>=6.0
|
|
12
|
+
Provides-Extra: dev
|
|
13
|
+
Requires-Dist: pytest>=8.0; extra == "dev"
|
|
14
|
+
Dynamic: license-file
|
|
15
|
+
|
|
16
|
+
# ForgeLine 🔨
|
|
17
|
+
|
|
18
|
+
**The autonomous outer loop for AI software factories.** ForgeLine is the tier
|
|
19
|
+
above spec-writing and code-generation: a **CLI-backed state machine** that
|
|
20
|
+
carries a feature from vague intent to shipped code through confidence gates,
|
|
21
|
+
runs a generate → **adversarial review** → refine loop, enforces
|
|
22
|
+
**architecture as a hard CI gate**, and gets smarter every run via an evolving
|
|
23
|
+
skill memory.
|
|
24
|
+
|
|
25
|
+
It orchestrates its siblings — [SpecLine](../specline) (spec governance) and
|
|
26
|
+
[Harness Software Factory](../harness-factory) (compiled decisions) — into one
|
|
27
|
+
pipeline you drive from Claude Code or Codex.
|
|
28
|
+
|
|
29
|
+
> Stop writing the code. Build — and supervise — the machine that writes it.
|
|
30
|
+
|
|
31
|
+
## Workflow at a glance
|
|
32
|
+
|
|
33
|
+
```mermaid
|
|
34
|
+
flowchart LR
|
|
35
|
+
A["Feature intent"] --> B["Expand use cases"]
|
|
36
|
+
B --> C["Human confidence gate"]
|
|
37
|
+
C --> D["Architect with SSAT"]
|
|
38
|
+
D --> E["Scaffold from architecture"]
|
|
39
|
+
E --> F["Fill implementation"]
|
|
40
|
+
F --> G["Adversarial review"]
|
|
41
|
+
G -->|"findings"| H["Refine and record lesson"]
|
|
42
|
+
H --> F
|
|
43
|
+
G -->|"pass"| I["Architecture CI gate"]
|
|
44
|
+
I --> J["Reverse-classical test verification"]
|
|
45
|
+
J --> K["Runtime smoke gate"]
|
|
46
|
+
K --> L["Ship with receipt"]
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
```
|
|
50
|
+
INTENT ─► EXPAND ─► ARCHITECT(SSAT) ─► SCAFFOLD ─► FILL ─┐
|
|
51
|
+
▲ │
|
|
52
|
+
│ ┌── grumpy adversary + judge + arch-erosion ──┤
|
|
53
|
+
└── refine┤ (records a skill lesson) │ pass
|
|
54
|
+
└──────────────◄─────────────────────────────┘
|
|
55
|
+
│
|
|
56
|
+
ARCH-CI-GATE ─► SHIP
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
## Install (any OS, 60 seconds, no API keys)
|
|
60
|
+
|
|
61
|
+
**One command, works everywhere:**
|
|
62
|
+
```bash
|
|
63
|
+
python install.py # Windows / macOS / Linux
|
|
64
|
+
```
|
|
65
|
+
or double-click `install.sh` (macOS/Linux) / `install.bat` (Windows), or:
|
|
66
|
+
```bash
|
|
67
|
+
pip install -e ".[dev]"
|
|
68
|
+
```
|
|
69
|
+
|
|
70
|
+
Then:
|
|
71
|
+
```bash
|
|
72
|
+
forge demo # 60-sec: watch it catch its own bad output, learn, ship
|
|
73
|
+
forge init
|
|
74
|
+
forge agent claude # wire your agent (see full list below)
|
|
75
|
+
forge status <feature> # the state machine names the ONE next action
|
|
76
|
+
pytest -q # 29 tests
|
|
77
|
+
```
|
|
78
|
+
|
|
79
|
+
## Works with every major coding agent
|
|
80
|
+
|
|
81
|
+
One command wires the entry-point skill wherever your agent reads it:
|
|
82
|
+
|
|
83
|
+
```bash
|
|
84
|
+
forge agent claude # CLAUDE.md + .claude/skills/forge.md
|
|
85
|
+
forge agent codex # AGENTS.md
|
|
86
|
+
forge agent opencode # AGENTS.md + .opencode/forge.md
|
|
87
|
+
forge agent cursor # .cursorrules + .cursor/rules/forge.md
|
|
88
|
+
forge agent aider # CONVENTIONS.md
|
|
89
|
+
forge agent gemini # GEMINI.md
|
|
90
|
+
forge agent windsurf # .windsurfrules
|
|
91
|
+
forge agent generic # AGENT.md (unknown tools fall back to AGENTS.md)
|
|
92
|
+
```
|
|
93
|
+
The contract is plain text; any harness that reads a project file can run the factory.
|
|
94
|
+
|
|
95
|
+
## What makes it more than a wrapper
|
|
96
|
+
|
|
97
|
+
**A real state machine with confidence gates.** Eight states, legal-transition
|
|
98
|
+
enforcement (`E_ILLEGAL_TRANSITION`), and two **human confidence gates**
|
|
99
|
+
(use-case signoff, architecture signoff) where a person must approve before
|
|
100
|
+
agents proceed. Every transition is a hash-sealed receipt on disk — so the
|
|
101
|
+
loop survives context resets (the disk is the truth).
|
|
102
|
+
|
|
103
|
+
**SSAT — architecture as code, not a document.** A Semantic Software
|
|
104
|
+
Architecture Tree (YAML) declares modules, signatures, allowed dependency
|
|
105
|
+
edges, and invariants. The *same artifact* generates the scaffold AND serves
|
|
106
|
+
as the CI gate: `check_erosion()` detects signature drift (`E_SIG_DRIFT`),
|
|
107
|
+
illegal dependencies (`E_ILLEGAL_DEP`), and invariant violations
|
|
108
|
+
(`E_INVARIANT`) — structural erosion caught before merge.
|
|
109
|
+
|
|
110
|
+
**The grumpy adversary.** A review agent that *assumes your code is broken and
|
|
111
|
+
insecure* and makes the generator prove otherwise. Executable heuristics catch
|
|
112
|
+
eval/exec/shell-injection/hard-coded-secrets/bare-excepts, and it refuses to
|
|
113
|
+
pass any code that ships without tests (`A_NO_PROOF` — "prove it works"). No
|
|
114
|
+
LLM required to be useful; an LLM adversary layers behind the same interface.
|
|
115
|
+
|
|
116
|
+
**A self-improving skill flywheel.** Every gate failure records a structured
|
|
117
|
+
lesson to `skills/lessons.jsonl`. Lessons are injected into the next attempt's
|
|
118
|
+
context, and lessons seen ≥3 times **graduate into hard constraints**
|
|
119
|
+
(conventions-into-constraints) — promoted straight into SSAT invariants. The
|
|
120
|
+
factory literally learns your team's rules from its own mistakes.
|
|
121
|
+
|
|
122
|
+
**Decision handoff to the factory.** Specs carrying a decision table route to
|
|
123
|
+
HSF for one-time compilation into gated, deterministic code — the outer loop
|
|
124
|
+
knows the difference between *tissue* (agents write it) and *decisions* (the
|
|
125
|
+
factory compiles them, never improvised twice).
|
|
126
|
+
|
|
127
|
+
## The entry-point skill (Claude Code / Codex)
|
|
128
|
+
|
|
129
|
+
`forge agent claude` writes `CLAUDE.md` + `.claude/skills/forge.md`;
|
|
130
|
+
`forge agent codex` writes `AGENTS.md`. The contract turns any agent into a
|
|
131
|
+
disciplined factory worker: read the skill, run `forge status`, do the one
|
|
132
|
+
named phase, run its gate, repeat — context resetting between phases. The
|
|
133
|
+
agent never free-codes; it advances a state machine. That's the whole point.
|
|
134
|
+
|
|
135
|
+
## Commands
|
|
136
|
+
|
|
137
|
+
```
|
|
138
|
+
forge init scaffold the factory
|
|
139
|
+
forge agent claude|codex wire the entry-point skill
|
|
140
|
+
forge status <feature> current state + the ONE next action
|
|
141
|
+
forge expand <feature> draft use cases (→ human gate)
|
|
142
|
+
forge architect <feature> <ssat> generate scaffold from architecture-as-code
|
|
143
|
+
forge review <feature> <ssat> judge + grumpy adversary + arch erosion (refine loop)
|
|
144
|
+
forge arch-gate <feature> <ssat> architecture CI gate
|
|
145
|
+
forge verify-tests <feature> <ssat> prove smoke checks fail on generated stubs
|
|
146
|
+
forge smoke <feature> runtime behavior gate
|
|
147
|
+
forge ship <feature> seal it
|
|
148
|
+
forge handoff <feature> <spec> route decision tables to HSF
|
|
149
|
+
forge lessons show the skill memory + promotable constraints
|
|
150
|
+
forge demo the 60-second story
|
|
151
|
+
```
|
|
152
|
+
|
|
153
|
+
## v0.3 — the Intent Thread (PRD → production traceability)
|
|
154
|
+
|
|
155
|
+
ForgeLine now consumes SpecLine's sealed **Intent Envelope** and verifies the
|
|
156
|
+
FINAL shipped code against the ORIGINAL rationalized intent — not the plan, not
|
|
157
|
+
the drifted spec, but the intent that was sealed at the plan gate. This closes
|
|
158
|
+
the last translation-loss gap: *does the shipped thing actually satisfy what we
|
|
159
|
+
rationalized we wanted?*
|
|
160
|
+
|
|
161
|
+
Every assumption SpecLine surfaced (auth exists, currency is single-source,
|
|
162
|
+
dependencies can fail) becomes a **checkable obligation**. The `ship` gate
|
|
163
|
+
blocks if the code shows no handling for an assumption the intent depended on,
|
|
164
|
+
and the sealed intent hash proves the intent wasn't quietly swapped underneath
|
|
165
|
+
the build. Full PRD-to-production traceability, enforced — not documented.
|
|
166
|
+
|
|
167
|
+
## v0.2 — deeper QA + recursive learning
|
|
168
|
+
|
|
169
|
+
ForgeLine now audits quality quantitatively and **learns from its own runs**:
|
|
170
|
+
|
|
171
|
+
**Deep QA audit** (`forge qa`) grades every build on coverage-intent (do tests
|
|
172
|
+
actually call the functions?), cyclomatic complexity, a scored security surface
|
|
173
|
+
(eval/exec/shell/secrets/weak-crypto), and documentation — a composite A–F grade
|
|
174
|
+
that gates shipping. A pretty build with untested, over-complex, or insecure
|
|
175
|
+
code cannot pass.
|
|
176
|
+
|
|
177
|
+
**Recursive learning kernel** (`forge policy`, `forge demo-learning`) closes the
|
|
178
|
+
loop the skill memory only started:
|
|
179
|
+
- **observe** — a failure is recorded (as before)
|
|
180
|
+
- **promote** — a failure seen ≥3× becomes an *enforced active constraint*
|
|
181
|
+
- **validate** — when that constraint catches the same failure again, it's marked
|
|
182
|
+
effective; the policy tracks prevention counts
|
|
183
|
+
- **self-prune** — a promoted rule that never fires again goes to *probation*, so
|
|
184
|
+
the policy doesn't ossify
|
|
185
|
+
|
|
186
|
+
The factory's own run history becomes its QA policy, and that policy is measured,
|
|
187
|
+
not assumed. `forge demo-learning` shows the full observe→promote→validate→prune
|
|
188
|
+
cycle in 15 seconds.
|
|
189
|
+
|
|
190
|
+
**Escalating refine loop** — the review loop now gets stricter each attempt
|
|
191
|
+
(normal → elevated "fix all findings" → final "human review required") instead
|
|
192
|
+
of a flat retry cap.
|
|
193
|
+
|
|
194
|
+
## The three-repo factory
|
|
195
|
+
|
|
196
|
+
| Repo | Tier | Owns |
|
|
197
|
+
|---|---|---|
|
|
198
|
+
| **ForgeLine** | outer loop | intent→ship state machine, adversarial gates, skill flywheel, arch-as-CI |
|
|
199
|
+
| **SpecLine** | spec governance | EARS specs, atomic task packets, token-lean context, intent-drift guard |
|
|
200
|
+
| **HSF** | decision compiler | ordered business rules → gated deterministic code, zero tokens/decision |
|
|
201
|
+
|
|
202
|
+
Same doctrine at every tier: gate everything, receipts or it didn't happen,
|
|
203
|
+
compile what shouldn't be reasoned twice.
|
|
204
|
+
|
|
205
|
+
## License
|
|
206
|
+
|
|
207
|
+
Dual-licensed under either **Apache-2.0 OR MIT** at your option — the
|
|
208
|
+
permissive standard for broad adoption. Pick whichever your project prefers.
|
|
209
|
+
|
|
210
|
+
---
|
|
211
|
+
|
|
212
|
+
## v0.4 — Runtime Smoke Gate (behavior-by-inspection)
|
|
213
|
+
|
|
214
|
+
ForgeLine's earlier gates all verify code against *specifications* — the judge
|
|
215
|
+
checks consistency, the QA audit grades static quality, the intent thread proves
|
|
216
|
+
the shipped code honors the sealed envelope. That is correctness *by construction*.
|
|
217
|
+
|
|
218
|
+
None of it answers the question a per-PR preview deployment answers: **does the
|
|
219
|
+
built thing actually RUN and behave correctly?** A change can pass every static
|
|
220
|
+
gate and still crash on import or produce the wrong output. v0.4 closes that gap
|
|
221
|
+
at solo-builder scale — no ephemeral per-PR environments required.
|
|
222
|
+
|
|
223
|
+
### New state + gate
|
|
224
|
+
|
|
225
|
+
The SDLC gains a `SMOKED` state between `ARCH_GATED` and `SHIPPED`:
|
|
226
|
+
|
|
227
|
+
```
|
|
228
|
+
… → reviewed → arch_gated → smoked → shipped
|
|
229
|
+
```
|
|
230
|
+
|
|
231
|
+
`forge smoke <feature>` runs every behavioral check declared in
|
|
232
|
+
`smoke/<feature>.json` in an **isolated subprocess with a timeout**, and blocks
|
|
233
|
+
ship on any runtime failure. `ship` now refuses to run until the smoke gate has
|
|
234
|
+
passed — you cannot ship unverified runtime behavior.
|
|
235
|
+
|
|
236
|
+
### The smoke manifest
|
|
237
|
+
|
|
238
|
+
`smoke/<feature>.json` makes "correct runtime behavior" a reviewed artifact:
|
|
239
|
+
|
|
240
|
+
```json
|
|
241
|
+
{
|
|
242
|
+
"checks": [
|
|
243
|
+
{
|
|
244
|
+
"name": "formatter_runtime",
|
|
245
|
+
"kind": "python",
|
|
246
|
+
"run": "from slices.notifier.formatter import format_message\nassert format_message({'kind':'ping','text':'hi'}) == 'ping: hi'\nprint('OK')",
|
|
247
|
+
"expect_exit": 0,
|
|
248
|
+
"expect_stdout": "OK",
|
|
249
|
+
"timeout_s": 15
|
|
250
|
+
}
|
|
251
|
+
]
|
|
252
|
+
}
|
|
253
|
+
```
|
|
254
|
+
|
|
255
|
+
`kind` is `python` (runs a snippet) or `command` (runs a shell command). Each
|
|
256
|
+
check asserts an exit code and, optionally, a stdout substring.
|
|
257
|
+
|
|
258
|
+
### Fail-closed by design
|
|
259
|
+
|
|
260
|
+
- **No manifest** → BLOCK. You cannot ship runtime behavior you never verified.
|
|
261
|
+
- **Empty manifest** → BLOCK. A manifest with zero checks ships nothing verified.
|
|
262
|
+
- **Any check fails / times out / crashes** → BLOCK, with the failing check named
|
|
263
|
+
and the last lines of its output captured as a receipt.
|
|
264
|
+
|
|
265
|
+
### Why this and not full preview deployments
|
|
266
|
+
|
|
267
|
+
Per-PR ephemeral environments are real DevOps infrastructure (provisioning,
|
|
268
|
+
isolation, teardown, secrets, cost) that pays off at *team* scale with many
|
|
269
|
+
parallel PRs and human reviewers. The smoke gate captures the core value —
|
|
270
|
+
*verify the artifact runs and behaves before merge* — deterministically, at the
|
|
271
|
+
cost of a subprocess. When team scale justifies it, the manifest model extends
|
|
272
|
+
naturally to spinning real environments; the behavioral contract is already written.
|
|
273
|
+
## v0.6 - Reverse-Classical test verification
|
|
274
|
+
|
|
275
|
+
ForgeLine now refuses tests that prove nothing.
|
|
276
|
+
|
|
277
|
+
`forge verify-tests <feature> <ssat.yaml>` regenerates the SSAT scaffold into an
|
|
278
|
+
isolated temp root and runs every behavioral smoke check against those empty
|
|
279
|
+
stubs. Each behavioral check must fail on the stub before ForgeLine trusts it
|
|
280
|
+
against the real implementation. A check that passes against an empty stub is
|
|
281
|
+
classified as `HOLLOW_TEST` and blocks the feature.
|
|
282
|
+
|
|
283
|
+
The ordering is now:
|
|
284
|
+
|
|
285
|
+
```text
|
|
286
|
+
... -> reviewed -> arch_gated -> tests_verified -> smoked -> shipped
|
|
287
|
+
```
|
|
288
|
+
|
|
289
|
+
Structural checks such as imports may declare `"must_fail_on_stub": false` in
|
|
290
|
+
`smoke/<feature>.json`, but the field defaults to `true`. Omission never buys
|
|
291
|
+
leniency, and a manifest where every check is exempt blocks as
|
|
292
|
+
`HOLLOW_MANIFEST`.
|
|
293
|
+
|
|
294
|
+
This gate is deterministic: no model, no git snapshot, no new runtime service.
|
|
295
|
+
It reuses the same SSAT scaffold generator as the normal `SCAFFOLDED` state, so
|
|
296
|
+
the mutant is the real generated stub.
|
|
297
|
+
|
|
298
|
+
## Failure attribution and refinement
|
|
299
|
+
|
|
300
|
+
ForgeLine 0.5 reports review, architecture, QA, smoke, and intent failures at
|
|
301
|
+
their smallest actionable unit. `forge qa --root .` includes function-level
|
|
302
|
+
metrics and attribution in its JSON output.
|
|
303
|
+
|
|
304
|
+
The refinement engine accepts exactly one proposed edit at a time. Structural
|
|
305
|
+
edits precede configuration and parameter changes. An edit is retained only
|
|
306
|
+
when its targeted stage improves and no other stage regresses; rejected edits
|
|
307
|
+
are reverted and written with before/after rates to
|
|
308
|
+
`.forge/rejection_ledger.jsonl`. Two consecutive non-wins stop the loop.
|
|
@@ -0,0 +1,293 @@
|
|
|
1
|
+
# ForgeLine 🔨
|
|
2
|
+
|
|
3
|
+
**The autonomous outer loop for AI software factories.** ForgeLine is the tier
|
|
4
|
+
above spec-writing and code-generation: a **CLI-backed state machine** that
|
|
5
|
+
carries a feature from vague intent to shipped code through confidence gates,
|
|
6
|
+
runs a generate → **adversarial review** → refine loop, enforces
|
|
7
|
+
**architecture as a hard CI gate**, and gets smarter every run via an evolving
|
|
8
|
+
skill memory.
|
|
9
|
+
|
|
10
|
+
It orchestrates its siblings — [SpecLine](../specline) (spec governance) and
|
|
11
|
+
[Harness Software Factory](../harness-factory) (compiled decisions) — into one
|
|
12
|
+
pipeline you drive from Claude Code or Codex.
|
|
13
|
+
|
|
14
|
+
> Stop writing the code. Build — and supervise — the machine that writes it.
|
|
15
|
+
|
|
16
|
+
## Workflow at a glance
|
|
17
|
+
|
|
18
|
+
```mermaid
|
|
19
|
+
flowchart LR
|
|
20
|
+
A["Feature intent"] --> B["Expand use cases"]
|
|
21
|
+
B --> C["Human confidence gate"]
|
|
22
|
+
C --> D["Architect with SSAT"]
|
|
23
|
+
D --> E["Scaffold from architecture"]
|
|
24
|
+
E --> F["Fill implementation"]
|
|
25
|
+
F --> G["Adversarial review"]
|
|
26
|
+
G -->|"findings"| H["Refine and record lesson"]
|
|
27
|
+
H --> F
|
|
28
|
+
G -->|"pass"| I["Architecture CI gate"]
|
|
29
|
+
I --> J["Reverse-classical test verification"]
|
|
30
|
+
J --> K["Runtime smoke gate"]
|
|
31
|
+
K --> L["Ship with receipt"]
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
```
|
|
35
|
+
INTENT ─► EXPAND ─► ARCHITECT(SSAT) ─► SCAFFOLD ─► FILL ─┐
|
|
36
|
+
▲ │
|
|
37
|
+
│ ┌── grumpy adversary + judge + arch-erosion ──┤
|
|
38
|
+
└── refine┤ (records a skill lesson) │ pass
|
|
39
|
+
└──────────────◄─────────────────────────────┘
|
|
40
|
+
│
|
|
41
|
+
ARCH-CI-GATE ─► SHIP
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
## Install (any OS, 60 seconds, no API keys)
|
|
45
|
+
|
|
46
|
+
**One command, works everywhere:**
|
|
47
|
+
```bash
|
|
48
|
+
python install.py # Windows / macOS / Linux
|
|
49
|
+
```
|
|
50
|
+
or double-click `install.sh` (macOS/Linux) / `install.bat` (Windows), or:
|
|
51
|
+
```bash
|
|
52
|
+
pip install -e ".[dev]"
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
Then:
|
|
56
|
+
```bash
|
|
57
|
+
forge demo # 60-sec: watch it catch its own bad output, learn, ship
|
|
58
|
+
forge init
|
|
59
|
+
forge agent claude # wire your agent (see full list below)
|
|
60
|
+
forge status <feature> # the state machine names the ONE next action
|
|
61
|
+
pytest -q # 29 tests
|
|
62
|
+
```
|
|
63
|
+
|
|
64
|
+
## Works with every major coding agent
|
|
65
|
+
|
|
66
|
+
One command wires the entry-point skill wherever your agent reads it:
|
|
67
|
+
|
|
68
|
+
```bash
|
|
69
|
+
forge agent claude # CLAUDE.md + .claude/skills/forge.md
|
|
70
|
+
forge agent codex # AGENTS.md
|
|
71
|
+
forge agent opencode # AGENTS.md + .opencode/forge.md
|
|
72
|
+
forge agent cursor # .cursorrules + .cursor/rules/forge.md
|
|
73
|
+
forge agent aider # CONVENTIONS.md
|
|
74
|
+
forge agent gemini # GEMINI.md
|
|
75
|
+
forge agent windsurf # .windsurfrules
|
|
76
|
+
forge agent generic # AGENT.md (unknown tools fall back to AGENTS.md)
|
|
77
|
+
```
|
|
78
|
+
The contract is plain text; any harness that reads a project file can run the factory.
|
|
79
|
+
|
|
80
|
+
## What makes it more than a wrapper
|
|
81
|
+
|
|
82
|
+
**A real state machine with confidence gates.** Eight states, legal-transition
|
|
83
|
+
enforcement (`E_ILLEGAL_TRANSITION`), and two **human confidence gates**
|
|
84
|
+
(use-case signoff, architecture signoff) where a person must approve before
|
|
85
|
+
agents proceed. Every transition is a hash-sealed receipt on disk — so the
|
|
86
|
+
loop survives context resets (the disk is the truth).
|
|
87
|
+
|
|
88
|
+
**SSAT — architecture as code, not a document.** A Semantic Software
|
|
89
|
+
Architecture Tree (YAML) declares modules, signatures, allowed dependency
|
|
90
|
+
edges, and invariants. The *same artifact* generates the scaffold AND serves
|
|
91
|
+
as the CI gate: `check_erosion()` detects signature drift (`E_SIG_DRIFT`),
|
|
92
|
+
illegal dependencies (`E_ILLEGAL_DEP`), and invariant violations
|
|
93
|
+
(`E_INVARIANT`) — structural erosion caught before merge.
|
|
94
|
+
|
|
95
|
+
**The grumpy adversary.** A review agent that *assumes your code is broken and
|
|
96
|
+
insecure* and makes the generator prove otherwise. Executable heuristics catch
|
|
97
|
+
eval/exec/shell-injection/hard-coded-secrets/bare-excepts, and it refuses to
|
|
98
|
+
pass any code that ships without tests (`A_NO_PROOF` — "prove it works"). No
|
|
99
|
+
LLM required to be useful; an LLM adversary layers behind the same interface.
|
|
100
|
+
|
|
101
|
+
**A self-improving skill flywheel.** Every gate failure records a structured
|
|
102
|
+
lesson to `skills/lessons.jsonl`. Lessons are injected into the next attempt's
|
|
103
|
+
context, and lessons seen ≥3 times **graduate into hard constraints**
|
|
104
|
+
(conventions-into-constraints) — promoted straight into SSAT invariants. The
|
|
105
|
+
factory literally learns your team's rules from its own mistakes.
|
|
106
|
+
|
|
107
|
+
**Decision handoff to the factory.** Specs carrying a decision table route to
|
|
108
|
+
HSF for one-time compilation into gated, deterministic code — the outer loop
|
|
109
|
+
knows the difference between *tissue* (agents write it) and *decisions* (the
|
|
110
|
+
factory compiles them, never improvised twice).
|
|
111
|
+
|
|
112
|
+
## The entry-point skill (Claude Code / Codex)
|
|
113
|
+
|
|
114
|
+
`forge agent claude` writes `CLAUDE.md` + `.claude/skills/forge.md`;
|
|
115
|
+
`forge agent codex` writes `AGENTS.md`. The contract turns any agent into a
|
|
116
|
+
disciplined factory worker: read the skill, run `forge status`, do the one
|
|
117
|
+
named phase, run its gate, repeat — context resetting between phases. The
|
|
118
|
+
agent never free-codes; it advances a state machine. That's the whole point.
|
|
119
|
+
|
|
120
|
+
## Commands
|
|
121
|
+
|
|
122
|
+
```
|
|
123
|
+
forge init scaffold the factory
|
|
124
|
+
forge agent claude|codex wire the entry-point skill
|
|
125
|
+
forge status <feature> current state + the ONE next action
|
|
126
|
+
forge expand <feature> draft use cases (→ human gate)
|
|
127
|
+
forge architect <feature> <ssat> generate scaffold from architecture-as-code
|
|
128
|
+
forge review <feature> <ssat> judge + grumpy adversary + arch erosion (refine loop)
|
|
129
|
+
forge arch-gate <feature> <ssat> architecture CI gate
|
|
130
|
+
forge verify-tests <feature> <ssat> prove smoke checks fail on generated stubs
|
|
131
|
+
forge smoke <feature> runtime behavior gate
|
|
132
|
+
forge ship <feature> seal it
|
|
133
|
+
forge handoff <feature> <spec> route decision tables to HSF
|
|
134
|
+
forge lessons show the skill memory + promotable constraints
|
|
135
|
+
forge demo the 60-second story
|
|
136
|
+
```
|
|
137
|
+
|
|
138
|
+
## v0.3 — the Intent Thread (PRD → production traceability)
|
|
139
|
+
|
|
140
|
+
ForgeLine now consumes SpecLine's sealed **Intent Envelope** and verifies the
|
|
141
|
+
FINAL shipped code against the ORIGINAL rationalized intent — not the plan, not
|
|
142
|
+
the drifted spec, but the intent that was sealed at the plan gate. This closes
|
|
143
|
+
the last translation-loss gap: *does the shipped thing actually satisfy what we
|
|
144
|
+
rationalized we wanted?*
|
|
145
|
+
|
|
146
|
+
Every assumption SpecLine surfaced (auth exists, currency is single-source,
|
|
147
|
+
dependencies can fail) becomes a **checkable obligation**. The `ship` gate
|
|
148
|
+
blocks if the code shows no handling for an assumption the intent depended on,
|
|
149
|
+
and the sealed intent hash proves the intent wasn't quietly swapped underneath
|
|
150
|
+
the build. Full PRD-to-production traceability, enforced — not documented.
|
|
151
|
+
|
|
152
|
+
## v0.2 — deeper QA + recursive learning
|
|
153
|
+
|
|
154
|
+
ForgeLine now audits quality quantitatively and **learns from its own runs**:
|
|
155
|
+
|
|
156
|
+
**Deep QA audit** (`forge qa`) grades every build on coverage-intent (do tests
|
|
157
|
+
actually call the functions?), cyclomatic complexity, a scored security surface
|
|
158
|
+
(eval/exec/shell/secrets/weak-crypto), and documentation — a composite A–F grade
|
|
159
|
+
that gates shipping. A pretty build with untested, over-complex, or insecure
|
|
160
|
+
code cannot pass.
|
|
161
|
+
|
|
162
|
+
**Recursive learning kernel** (`forge policy`, `forge demo-learning`) closes the
|
|
163
|
+
loop the skill memory only started:
|
|
164
|
+
- **observe** — a failure is recorded (as before)
|
|
165
|
+
- **promote** — a failure seen ≥3× becomes an *enforced active constraint*
|
|
166
|
+
- **validate** — when that constraint catches the same failure again, it's marked
|
|
167
|
+
effective; the policy tracks prevention counts
|
|
168
|
+
- **self-prune** — a promoted rule that never fires again goes to *probation*, so
|
|
169
|
+
the policy doesn't ossify
|
|
170
|
+
|
|
171
|
+
The factory's own run history becomes its QA policy, and that policy is measured,
|
|
172
|
+
not assumed. `forge demo-learning` shows the full observe→promote→validate→prune
|
|
173
|
+
cycle in 15 seconds.
|
|
174
|
+
|
|
175
|
+
**Escalating refine loop** — the review loop now gets stricter each attempt
|
|
176
|
+
(normal → elevated "fix all findings" → final "human review required") instead
|
|
177
|
+
of a flat retry cap.
|
|
178
|
+
|
|
179
|
+
## The three-repo factory
|
|
180
|
+
|
|
181
|
+
| Repo | Tier | Owns |
|
|
182
|
+
|---|---|---|
|
|
183
|
+
| **ForgeLine** | outer loop | intent→ship state machine, adversarial gates, skill flywheel, arch-as-CI |
|
|
184
|
+
| **SpecLine** | spec governance | EARS specs, atomic task packets, token-lean context, intent-drift guard |
|
|
185
|
+
| **HSF** | decision compiler | ordered business rules → gated deterministic code, zero tokens/decision |
|
|
186
|
+
|
|
187
|
+
Same doctrine at every tier: gate everything, receipts or it didn't happen,
|
|
188
|
+
compile what shouldn't be reasoned twice.
|
|
189
|
+
|
|
190
|
+
## License
|
|
191
|
+
|
|
192
|
+
Dual-licensed under either **Apache-2.0 OR MIT** at your option — the
|
|
193
|
+
permissive standard for broad adoption. Pick whichever your project prefers.
|
|
194
|
+
|
|
195
|
+
---
|
|
196
|
+
|
|
197
|
+
## v0.4 — Runtime Smoke Gate (behavior-by-inspection)
|
|
198
|
+
|
|
199
|
+
ForgeLine's earlier gates all verify code against *specifications* — the judge
|
|
200
|
+
checks consistency, the QA audit grades static quality, the intent thread proves
|
|
201
|
+
the shipped code honors the sealed envelope. That is correctness *by construction*.
|
|
202
|
+
|
|
203
|
+
None of it answers the question a per-PR preview deployment answers: **does the
|
|
204
|
+
built thing actually RUN and behave correctly?** A change can pass every static
|
|
205
|
+
gate and still crash on import or produce the wrong output. v0.4 closes that gap
|
|
206
|
+
at solo-builder scale — no ephemeral per-PR environments required.
|
|
207
|
+
|
|
208
|
+
### New state + gate
|
|
209
|
+
|
|
210
|
+
The SDLC gains a `SMOKED` state between `ARCH_GATED` and `SHIPPED`:
|
|
211
|
+
|
|
212
|
+
```
|
|
213
|
+
… → reviewed → arch_gated → smoked → shipped
|
|
214
|
+
```
|
|
215
|
+
|
|
216
|
+
`forge smoke <feature>` runs every behavioral check declared in
|
|
217
|
+
`smoke/<feature>.json` in an **isolated subprocess with a timeout**, and blocks
|
|
218
|
+
ship on any runtime failure. `ship` now refuses to run until the smoke gate has
|
|
219
|
+
passed — you cannot ship unverified runtime behavior.
|
|
220
|
+
|
|
221
|
+
### The smoke manifest
|
|
222
|
+
|
|
223
|
+
`smoke/<feature>.json` makes "correct runtime behavior" a reviewed artifact:
|
|
224
|
+
|
|
225
|
+
```json
|
|
226
|
+
{
|
|
227
|
+
"checks": [
|
|
228
|
+
{
|
|
229
|
+
"name": "formatter_runtime",
|
|
230
|
+
"kind": "python",
|
|
231
|
+
"run": "from slices.notifier.formatter import format_message\nassert format_message({'kind':'ping','text':'hi'}) == 'ping: hi'\nprint('OK')",
|
|
232
|
+
"expect_exit": 0,
|
|
233
|
+
"expect_stdout": "OK",
|
|
234
|
+
"timeout_s": 15
|
|
235
|
+
}
|
|
236
|
+
]
|
|
237
|
+
}
|
|
238
|
+
```
|
|
239
|
+
|
|
240
|
+
`kind` is `python` (runs a snippet) or `command` (runs a shell command). Each
|
|
241
|
+
check asserts an exit code and, optionally, a stdout substring.
|
|
242
|
+
|
|
243
|
+
### Fail-closed by design
|
|
244
|
+
|
|
245
|
+
- **No manifest** → BLOCK. You cannot ship runtime behavior you never verified.
|
|
246
|
+
- **Empty manifest** → BLOCK. A manifest with zero checks ships nothing verified.
|
|
247
|
+
- **Any check fails / times out / crashes** → BLOCK, with the failing check named
|
|
248
|
+
and the last lines of its output captured as a receipt.
|
|
249
|
+
|
|
250
|
+
### Why this and not full preview deployments
|
|
251
|
+
|
|
252
|
+
Per-PR ephemeral environments are real DevOps infrastructure (provisioning,
|
|
253
|
+
isolation, teardown, secrets, cost) that pays off at *team* scale with many
|
|
254
|
+
parallel PRs and human reviewers. The smoke gate captures the core value —
|
|
255
|
+
*verify the artifact runs and behaves before merge* — deterministically, at the
|
|
256
|
+
cost of a subprocess. When team scale justifies it, the manifest model extends
|
|
257
|
+
naturally to spinning real environments; the behavioral contract is already written.
|
|
258
|
+
## v0.6 - Reverse-Classical test verification
|
|
259
|
+
|
|
260
|
+
ForgeLine now refuses tests that prove nothing.
|
|
261
|
+
|
|
262
|
+
`forge verify-tests <feature> <ssat.yaml>` regenerates the SSAT scaffold into an
|
|
263
|
+
isolated temp root and runs every behavioral smoke check against those empty
|
|
264
|
+
stubs. Each behavioral check must fail on the stub before ForgeLine trusts it
|
|
265
|
+
against the real implementation. A check that passes against an empty stub is
|
|
266
|
+
classified as `HOLLOW_TEST` and blocks the feature.
|
|
267
|
+
|
|
268
|
+
The ordering is now:
|
|
269
|
+
|
|
270
|
+
```text
|
|
271
|
+
... -> reviewed -> arch_gated -> tests_verified -> smoked -> shipped
|
|
272
|
+
```
|
|
273
|
+
|
|
274
|
+
Structural checks such as imports may declare `"must_fail_on_stub": false` in
|
|
275
|
+
`smoke/<feature>.json`, but the field defaults to `true`. Omission never buys
|
|
276
|
+
leniency, and a manifest where every check is exempt blocks as
|
|
277
|
+
`HOLLOW_MANIFEST`.
|
|
278
|
+
|
|
279
|
+
This gate is deterministic: no model, no git snapshot, no new runtime service.
|
|
280
|
+
It reuses the same SSAT scaffold generator as the normal `SCAFFOLDED` state, so
|
|
281
|
+
the mutant is the real generated stub.
|
|
282
|
+
|
|
283
|
+
## Failure attribution and refinement
|
|
284
|
+
|
|
285
|
+
ForgeLine 0.5 reports review, architecture, QA, smoke, and intent failures at
|
|
286
|
+
their smallest actionable unit. `forge qa --root .` includes function-level
|
|
287
|
+
metrics and attribution in its JSON output.
|
|
288
|
+
|
|
289
|
+
The refinement engine accepts exactly one proposed edit at a time. Structural
|
|
290
|
+
edits precede configuration and parameter changes. An edit is retained only
|
|
291
|
+
when its targeted stage improves and no other stage regresses; rejected edits
|
|
292
|
+
are reverted and written with before/after rates to
|
|
293
|
+
`.forge/rejection_ledger.jsonl`. Two consecutive non-wins stop the loop.
|