software-factory 0.0.1__py3-none-any.whl
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- software_factory/SKILL.md +182 -0
- software_factory/__init__.py +8 -0
- software_factory/__main__.py +1246 -0
- software_factory/backend.py +255 -0
- software_factory/config.py +205 -0
- software_factory/engine.py +241 -0
- software_factory/judge.py +299 -0
- software_factory/pipeline.py +237 -0
- software_factory/runners/claude.yaml +25 -0
- software_factory/runners/shell.yaml +10 -0
- software_factory/runners/typesafe.yaml +5 -0
- software_factory/runners.py +141 -0
- software_factory/steps.py +330 -0
- software_factory-0.0.1.dist-info/METADATA +296 -0
- software_factory-0.0.1.dist-info/RECORD +18 -0
- software_factory-0.0.1.dist-info/WHEEL +4 -0
- software_factory-0.0.1.dist-info/entry_points.txt +2 -0
- software_factory-0.0.1.dist-info/licenses/LICENSE +21 -0
|
@@ -0,0 +1,182 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: sf
|
|
3
|
+
description: Drive the `sf` CLI in this repo - submit a request to the software factory, run the queue, see what is in flight, read a run back with replay, answer a request that is parked as needs_human, or write a pipeline YAML. Use whenever the user mentions the factory, `sf submit`, `sf run`, a request id like 7, a parked or needs_human request, or a pipeline stage/verdict.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Using the factory
|
|
7
|
+
|
|
8
|
+
A **request** (a sentence of intent) is worked through a **pipeline** (YAML you write) one
|
|
9
|
+
**stage** at a time. Each stage names the **runner** that performs it — an agent CLI with a
|
|
10
|
+
role prompt, a shell command, or a typed judgment. All of them return a verdict; the verdict
|
|
11
|
+
picks the next stage.
|
|
12
|
+
|
|
13
|
+
One global installation (`~/.sf`) serves every repo on the machine, like one Docker
|
|
14
|
+
daemon serves every image. Requests outlive your shell: submit here, ask from anywhere.
|
|
15
|
+
|
|
16
|
+
## Preflight
|
|
17
|
+
|
|
18
|
+
- `sf config` — settings in effect, known pipelines, repos in flight. Run it first if
|
|
19
|
+
anything looks misconfigured. `sf init` creates `~/.sf` and its directories. **No
|
|
20
|
+
pipeline ships** - if none exists yet, one has to be written before anything can run.
|
|
21
|
+
- **Submit from inside the repo the work belongs to**, and it must be a git repo with at
|
|
22
|
+
least one commit — a worktree needs a `HEAD`, so `submit` refuses without one.
|
|
23
|
+
- The absolute repo path is recorded at submit time, so **submit and run must agree on it**.
|
|
24
|
+
`--repo <path>` overrides the current directory; whatever it resolves to is what the run
|
|
25
|
+
must be able to reach.
|
|
26
|
+
|
|
27
|
+
## Commands
|
|
28
|
+
|
|
29
|
+
| Command | What it does |
|
|
30
|
+
|---|---|
|
|
31
|
+
| `sf submit "<request>" --name <label>` | queue a request; prints its id. `--file <path>` (`-` for stdin) for a long request, `--pipeline <name>`, `--repo <path>`, `--run` to work it immediately |
|
|
32
|
+
| `sf run` | work the queue, **blocking until it is done**. `--detach` leaves an engine working in the background instead, `--repo` limits to this repo, `<ids>` to only these, `--concurrency N` |
|
|
33
|
+
| `sf` / `sf status` | every request across every repo — the `docker ps` of the factory. `-t`/`--tabular` is the one-line-per-request table |
|
|
34
|
+
| any command | **you get JSON**: output is JSON whenever stdout is not a terminal, which it never is for you. `--json` says so explicitly; `SF_OUTPUT=human` gets the text form |
|
|
35
|
+
| `sf show <id>` | the raw item JSON (pipe to `jq`) |
|
|
36
|
+
| `sf replay <id>` | play the run back: every step, its route, cost, artifacts. `--step N` opens one up, `--artifact <name>` picks one out, `--json` is the flattened timeline |
|
|
37
|
+
| `sf runners` | every runner a step can `uses:`, whether its binary is on PATH, and what it supports |
|
|
38
|
+
| `sf run <id> --note "..."` | answer a parked request; resumes from `item.stage` with your note in context |
|
|
39
|
+
| `sf run <id> --stage <stage> --note "..."` | reject: re-enter at an earlier stage |
|
|
40
|
+
| `sf cancel <ids>` | stop them now; `sf run <id>` picks one back up |
|
|
41
|
+
| `sf delete <ids>` | drop the items, their artifacts and their worktrees. `--force` if one is running |
|
|
42
|
+
| `sf prune` | housekeeping: delete finished requests in bulk — every `done` one by default. `--status <s>` repeatable (`failed`, `cancelled`, `queued`, `needs_human`), `--repo`, `--older-than 7d`, `--dry-run` to see what would go, `--yes` to skip the confirmation. Never touches a `running` request, and takes each one's `sf/<id>` branch with it |
|
|
43
|
+
| `sf doctor` | one line per check of this installation: git, the `claude` CLI and its token, `SF_HOME`, the pipelines. Exits non-zero if something essential is missing — run it when a command behaves oddly |
|
|
44
|
+
|
|
45
|
+
## The loop
|
|
46
|
+
|
|
47
|
+
1. `sf submit "<request>" --name <label>` from inside the repo → note the id.
|
|
48
|
+
2. `sf run` (or `sf run <id>`).
|
|
49
|
+
3. Poll `sf` until the row is `done`, `failed` or `needs_human`.
|
|
50
|
+
4. `sf replay <id>` to read what happened. The work is the diff on branch `sf/<id>`
|
|
51
|
+
in its own worktree — `<worktrees>/<repo name>/<id>`, `~/.sf/worktrees/...` by default,
|
|
52
|
+
and printed as `.workspace` by `sf show <id>`. Not in the repo you submitted from.
|
|
53
|
+
|
|
54
|
+
`sf run` **blocks** until the queue is worked, exiting non-zero when nothing reached
|
|
55
|
+
`done` — so its exit code is a real answer and can gate CI. `--detach` is the opt-in form
|
|
56
|
+
that returns at once: after that one its exit code says only that the engine started, so
|
|
57
|
+
poll `sf` rather than reporting anything as done.
|
|
58
|
+
|
|
59
|
+
`sf run <id>` is also the recovery path for a request left `running` by a killed engine.
|
|
60
|
+
|
|
61
|
+
## When a request is parked (`needs_human`)
|
|
62
|
+
|
|
63
|
+
`needs_human` means an agent stopped and asked a question, or a `review: true` gate is
|
|
64
|
+
waiting. **That is a question for the user, not for you to answer.**
|
|
65
|
+
|
|
66
|
+
```bash
|
|
67
|
+
sf # find the parked request and its one-line reason
|
|
68
|
+
sf replay 1 # the route, and which step HALTED
|
|
69
|
+
sf replay 1 --step 4 # the agent's full reply — the actual question
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
Relay the question to the user, then act on their answer:
|
|
73
|
+
|
|
74
|
+
```bash
|
|
75
|
+
sf run 1 # approve a review gate as-is
|
|
76
|
+
sf run 1 --note "approved, that key is a test fixture"
|
|
77
|
+
sf run 1 --stage plan --note "requirement changed, redo"
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
Where the note lands depends on **why** it parked — the three cases differ:
|
|
81
|
+
|
|
82
|
+
- **An agent asked** (`human` verdict, no parseable verdict, no route for the verdict, or a
|
|
83
|
+
judge step that could not reach a decision): nothing was routed, so
|
|
84
|
+
`sf run <id> --note "..."` re-runs **that same stage** with the note in context. It does
|
|
85
|
+
not count as a rework pass.
|
|
86
|
+
- **`max_passes` exhausted**: the backwards edge was already taken and the pass already counted,
|
|
87
|
+
so the stage has advanced to the rework target. `sf run <id>` runs *that* target (e.g.
|
|
88
|
+
`code`), not the reviewer that failed it. And `passes` stays above `max_passes`, so the next
|
|
89
|
+
backwards edge parks it again immediately — the way out is `--stage` plus a change of scope,
|
|
90
|
+
not another `--note`.
|
|
91
|
+
- **A `review: true` gate**: the step's verdict was routed *before* parking, so the stage has
|
|
92
|
+
already advanced. `sf run <id>` resumes at the **next** stage and the note goes forward
|
|
93
|
+
into it — the gated step does not re-run.
|
|
94
|
+
|
|
95
|
+
To force an earlier stage to run again in either case, name it: `--stage <name>`. `--note` and
|
|
96
|
+
`--stage` both require an id; a bare `sf run` deliberately ignores parked requests.
|
|
97
|
+
|
|
98
|
+
## The verdict contract
|
|
99
|
+
|
|
100
|
+
Every agent prompt is suffixed with a request for a closing JSON object:
|
|
101
|
+
|
|
102
|
+
```json
|
|
103
|
+
{"verdict": "pass" | "fail" | "human", "notes": "..."}
|
|
104
|
+
```
|
|
105
|
+
|
|
106
|
+
`pass` and `fail` route via the step's edges. `human` halts the request where it stands.
|
|
107
|
+
**No parseable verdict also parks it** — the factory never guesses. The engine appends that
|
|
108
|
+
block to every agent prompt at run time, so a prompt file under the pipeline's own `prompts/` must not
|
|
109
|
+
restate it, and must not tell the agent to end its reply with anything else.
|
|
110
|
+
|
|
111
|
+
## Writing a pipeline
|
|
112
|
+
|
|
113
|
+
`.sf/pipelines/<name>.yaml` in the repo wins over `~/.sf/pipelines/<name>.yaml`,
|
|
114
|
+
and that is the whole search path - nothing ships. A step says who performs it
|
|
115
|
+
(`uses:`) and what to hand them (`with:`).
|
|
116
|
+
|
|
117
|
+
Shipped: `dev` (plan, code, test, review, commit), `quick` (code, commit — no plan, no
|
|
118
|
+
tests, no review) for requests small enough to state exactly (`--pipeline quick`), and
|
|
119
|
+
`judged` (a `uses: typesafe` gate in front of the reviewer; needs `TYPESAFE_API_KEY`).
|
|
120
|
+
|
|
121
|
+
```yaml
|
|
122
|
+
name: dev
|
|
123
|
+
max_passes: 3 # rework round-trips before a human is asked
|
|
124
|
+
start: plan # defaults to the first step
|
|
125
|
+
|
|
126
|
+
steps:
|
|
127
|
+
plan:
|
|
128
|
+
name: Write the plan # optional label, shown in status and replay
|
|
129
|
+
uses: claude # which runner; omitted, the configured default
|
|
130
|
+
with:
|
|
131
|
+
prompt: prompts/plan.md # path relative to this file
|
|
132
|
+
# effort: low # low|medium|high|xhigh|max
|
|
133
|
+
output: plan.md # what later steps can read
|
|
134
|
+
next: code
|
|
135
|
+
# review: true # park for a human after this step
|
|
136
|
+
gate:
|
|
137
|
+
uses: typesafe # typed questions + thresholds instead of an agent
|
|
138
|
+
with:
|
|
139
|
+
questions: prompts/review-gate.yaml
|
|
140
|
+
on: { pass: commit, fail: code } # parks for a human without TYPESAFE_API_KEY
|
|
141
|
+
# resume: false # start cold; the default resumes the previous stage's
|
|
142
|
+
# session, so the line is one conversation and earlier
|
|
143
|
+
# context is not billed again. A stage name picks that
|
|
144
|
+
# stage's own last session.
|
|
145
|
+
test:
|
|
146
|
+
uses: shell
|
|
147
|
+
with:
|
|
148
|
+
run: "pytest -q" # exit 0 = pass
|
|
149
|
+
output: test.log
|
|
150
|
+
on: { pass: review, fail: code }
|
|
151
|
+
review:
|
|
152
|
+
uses: claude
|
|
153
|
+
with:
|
|
154
|
+
prompt: prompts/review.md
|
|
155
|
+
input: [plan.md, test.log] # injected as `# Input: <name>` sections
|
|
156
|
+
on: { pass: done, fail: code }
|
|
157
|
+
```
|
|
158
|
+
|
|
159
|
+
Rules the loader enforces:
|
|
160
|
+
|
|
161
|
+
- A step needs **at least one** of `next:`/`on:`, and its `with:` must carry what the runner
|
|
162
|
+
its `uses:` names requires — `prompt:` for an agent, `run:` for `shell`, `questions:` for
|
|
163
|
+
`typesafe`. An unknown runner, or a `with:` key it does not take, fails at load.
|
|
164
|
+
- **Declaration order is the rank.** An edge to a later step is the happy path; an edge to an
|
|
165
|
+
equal-or-earlier step is *rework* — it carries notes back and increments `passes`. Past
|
|
166
|
+
`max_passes` the request parks instead of looping forever.
|
|
167
|
+
- An `input:` name no step produces is a wiring bug and is rejected at load, as are unknown
|
|
168
|
+
step keys and any name escaping `_FACTORY/`.
|
|
169
|
+
- A missing input is fine: on the first pass the reviewer has not run yet.
|
|
170
|
+
- `done` is the implicit terminal stage.
|
|
171
|
+
|
|
172
|
+
## Gotchas
|
|
173
|
+
|
|
174
|
+
- `_FACTORY/` in a worktree is factory bookkeeping, not the deliverable. It ignores itself.
|
|
175
|
+
- Never fabricate a status. Read it from `sf` / `sf show`.
|
|
176
|
+
- `--backend` takes a path or `file://`; any other scheme is an error today.
|
|
177
|
+
- A step that fails hard: `sf replay <id> --step N` prints its stderr and exit code. To
|
|
178
|
+
rerun it by hand, `--artifact command.txt` for an agent step, `--artifact input.sh` for a
|
|
179
|
+
`uses: shell` one. Ask for an artifact the step never wrote and replay prints nothing.
|
|
180
|
+
|
|
181
|
+
Depth lives in the software-factory repo's own `README.md` (YAML and command reference) and
|
|
182
|
+
its `docs/` directory (why it works this way).
|
|
@@ -0,0 +1,8 @@
|
|
|
1
|
+
"""The factory package. Its version is the installed distribution's, never a second copy."""
|
|
2
|
+
|
|
3
|
+
from importlib.metadata import PackageNotFoundError, version
|
|
4
|
+
|
|
5
|
+
try:
|
|
6
|
+
__version__ = version("software-factory")
|
|
7
|
+
except PackageNotFoundError: # a source tree nobody installed - `python -m factory`
|
|
8
|
+
__version__ = "0+source"
|