githeri 0.2.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- githeri-0.2.0/PKG-INFO +689 -0
- githeri-0.2.0/README.md +660 -0
- githeri-0.2.0/pyproject.toml +54 -0
- githeri-0.2.0/setup.cfg +4 -0
- githeri-0.2.0/src/githeri/__init__.py +3 -0
- githeri-0.2.0/src/githeri/assets/system_prompt_specgen.txt +1020 -0
- githeri-0.2.0/src/githeri/cli.py +371 -0
- githeri-0.2.0/src/githeri/client.py +74 -0
- githeri-0.2.0/src/githeri/config.py +71 -0
- githeri-0.2.0/src/githeri/core/__init__.py +1 -0
- githeri-0.2.0/src/githeri/core/spec_engine.py +197 -0
- githeri-0.2.0/src/githeri/core/spec_validator.py +638 -0
- githeri-0.2.0/src/githeri/core/synthesizer.py +640 -0
- githeri-0.2.0/src/githeri/core/validator.py +201 -0
- githeri-0.2.0/src/githeri.egg-info/PKG-INFO +689 -0
- githeri-0.2.0/src/githeri.egg-info/SOURCES.txt +28 -0
- githeri-0.2.0/src/githeri.egg-info/dependency_links.txt +1 -0
- githeri-0.2.0/src/githeri.egg-info/entry_points.txt +2 -0
- githeri-0.2.0/src/githeri.egg-info/requires.txt +6 -0
- githeri-0.2.0/src/githeri.egg-info/top_level.txt +2 -0
- githeri-0.2.0/tests/test_cli_and_billing.py +125 -0
- githeri-0.2.0/tests/test_cli_commands.py +113 -0
- githeri-0.2.0/tests/test_contract_verifier.py +119 -0
- githeri-0.2.0/tests/test_dependency_graph.py +33 -0
- githeri-0.2.0/tests/test_lmstudio_grg.py +11 -0
- githeri-0.2.0/tests/test_ollama.py +14 -0
- githeri-0.2.0/tests/test_orchestrator.py +48 -0
- githeri-0.2.0/tests/test_regex.py +20 -0
- githeri-0.2.0/tests/test_session.py +16 -0
- githeri-0.2.0/tests/test_validate.py +13 -0
githeri-0.2.0/PKG-INFO
ADDED
|
@@ -0,0 +1,689 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: githeri
|
|
3
|
+
Version: 0.2.0
|
|
4
|
+
Summary: Contract Engine for AI Coding Agents. Deterministic specs, test plans, and drift detection.
|
|
5
|
+
Author-email: Karakana Labs <dev@karakanalabs.com>
|
|
6
|
+
License-Expression: Apache-2.0
|
|
7
|
+
Project-URL: Homepage, https://githeri.com
|
|
8
|
+
Project-URL: Documentation, https://githeri.com/docs
|
|
9
|
+
Project-URL: Repository, https://github.com/karakana-labs/githeri
|
|
10
|
+
Project-URL: Issues, https://github.com/karakana-labs/githeri/issues
|
|
11
|
+
Keywords: ai,agents,spec,fastapi,testing,openapi,contract-testing
|
|
12
|
+
Classifier: Development Status :: 4 - Beta
|
|
13
|
+
Classifier: Intended Audience :: Developers
|
|
14
|
+
Classifier: Programming Language :: Python :: 3
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
18
|
+
Classifier: Topic :: Software Development :: Build Tools
|
|
19
|
+
Classifier: Topic :: Software Development :: Testing
|
|
20
|
+
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
|
|
21
|
+
Requires-Python: >=3.10
|
|
22
|
+
Description-Content-Type: text/markdown
|
|
23
|
+
Requires-Dist: typer>=0.12.0
|
|
24
|
+
Requires-Dist: rich>=13.0.0
|
|
25
|
+
Requires-Dist: requests>=2.28.0
|
|
26
|
+
Requires-Dist: pyyaml>=6.0
|
|
27
|
+
Requires-Dist: pydantic>=2.0.0
|
|
28
|
+
Requires-Dist: python-dotenv>=1.0.0
|
|
29
|
+
|
|
30
|
+
# githeri
|
|
31
|
+
|
|
32
|
+
**Spec-Forge**: a pipeline that trains an AI to bridge the gap between a human's natural-language feature request and a machine-executable specification that feeds the COMMAND_RUNWAY skill.
|
|
33
|
+
|
|
34
|
+
The central thesis: most time in AI-assisted development is lost before the AI and the human agree on what to build. Githeri attacks that by producing validated, structured YAML specs — which then feed into COMMAND_RUNWAY runbooks that an executor agent follows verbatim.
|
|
35
|
+
|
|
36
|
+
## Pipeline
|
|
37
|
+
|
|
38
|
+
```text
|
|
39
|
+
Human natural-language request
|
|
40
|
+
│
|
|
41
|
+
� ▼
|
|
42
|
+
Spec-Forge (LLM + validator)
|
|
43
|
+
Generates:
|
|
44
|
+
data/training_data.jsonl (valid prompt + spec_yaml pairs)
|
|
45
|
+
data/failed_specs.jsonl (invalid specs, saved for analysis)
|
|
46
|
+
│
|
|
47
|
+
� ▼
|
|
48
|
+
Runbook Scorer (standalone, post-generation)
|
|
49
|
+
Scores:
|
|
50
|
+
5 categories — Intent, Preconditions, Structure, Testability, Coverage
|
|
51
|
+
Hard gate: missing Inspect/Create/Verify = 0.0
|
|
52
|
+
│
|
|
53
|
+
� ▼
|
|
54
|
+
Human Review
|
|
55
|
+
"L1-L4 look good. Approve."
|
|
56
|
+
│
|
|
57
|
+
� ▼
|
|
58
|
+
COMMAND_RUNWAY Skill
|
|
59
|
+
Consumes:
|
|
60
|
+
a validated spec (single-feature YAML)
|
|
61
|
+
Produces:
|
|
62
|
+
COMMAND_RUNWAY.md
|
|
63
|
+
• ordered implementation plan
|
|
64
|
+
• exact file paths
|
|
65
|
+
• code modifications
|
|
66
|
+
• test skeletons
|
|
67
|
+
• verification commands (translated from spec local_goals)
|
|
68
|
+
• rollback guidance
|
|
69
|
+
• completion criteria
|
|
70
|
+
│
|
|
71
|
+
� ▼
|
|
72
|
+
GRG Executor (with COMMAND_RUNWAY pattern integration)
|
|
73
|
+
Consumes:
|
|
74
|
+
a validated spec OR COMMAND_RUNWAY plan JSON
|
|
75
|
+
Produces (all under foreign/ directory):
|
|
76
|
+
• implementation source files
|
|
77
|
+
• test files
|
|
78
|
+
• RUNBOOK.md (human-readable execution log with GRG scores)
|
|
79
|
+
• RUNBOOK.json (machine-readable execution data)
|
|
80
|
+
• automatic ruff check --fix on generated code
|
|
81
|
+
│
|
|
82
|
+
� ▼
|
|
83
|
+
Completed Feature
|
|
84
|
+
Outputs:
|
|
85
|
+
• implementation complete
|
|
86
|
+
• all tests passing
|
|
87
|
+
• OpenAPI updated
|
|
88
|
+
• documentation synchronized
|
|
89
|
+
• human notified
|
|
90
|
+
```
|
|
91
|
+
|
|
92
|
+
## Spec Enrichment (IMPROVE_SPEC)
|
|
93
|
+
|
|
94
|
+
Every generated spec now includes optional **enrichment fields** that make specs machine-executable:
|
|
95
|
+
|
|
96
|
+
| Field | Location | Purpose |
|
|
97
|
+
|-------|----------|---------|
|
|
98
|
+
| `business_rules` | top-level | Invariants & formulas (e.g., "JWT Secret: 256-bit random, rotated quarterly") |
|
|
99
|
+
| `test_fixtures` | top-level | Seed data & setup commands (e.g., `.venv/bin/python scripts/seed_admin.py`) |
|
|
100
|
+
| `environment` | top-level | Required packages + env vars (e.g., `pyyaml>=6.0`, `JWT_SECRET`) |
|
|
101
|
+
| `global_verification` | top-level | Post-execution gate commands (e.g., `pytest tests/`, `bandit -r src/`) |
|
|
102
|
+
| `blueprint` | per-goal | **Required for `type: create`** (≥100 chars). Code-level outline: class signatures, route decorators, SQLAlchemy models, business logic steps |
|
|
103
|
+
| `acceptance_criteria` | per-goal | List of `{test, steps}` — executable test cases in pseudo-code |
|
|
104
|
+
| `type` | per-goal | `create` \| `update` \| `delete` \| `inspect` \| `verify` — drives runbook stage classification |
|
|
105
|
+
|
|
106
|
+
These fields are validated by `scripts/validator.py` and consumed by downstream generators (plan, runbook, scorer).
|
|
107
|
+
|
|
108
|
+
## Runbook Scoring System
|
|
109
|
+
|
|
110
|
+
Every generated spec is scored against runbook-readiness criteria (see `docs/scoring_spec.md`). Scoring is decoupled from generation — specs are saved first, then scored in a separate pass via `make score`.
|
|
111
|
+
|
|
112
|
+
The scorer (`scripts/runbook_scorer.py`) evaluates five weighted categories:
|
|
113
|
+
|
|
114
|
+
| Category | Weight | Key Checks |
|
|
115
|
+
|----------|--------|------------|
|
|
116
|
+
| **Intent & Goals** | 20% | Summary present, goals have descriptions, endpoint tasks have HTTP verification |
|
|
117
|
+
| **Preconditions** | 15% | `depends_on` references valid globals/stages, CLI tools declared in context |
|
|
118
|
+
| **Command Runway Structure** | 30% | **Hard gate**: must have Inspect (file_exists/read CLI), Create/Modify (build CLI), Verify (HTTP/test CLI). Stage order: Inspect → Create → Verify |
|
|
119
|
+
| **Verification Testability** | 25% | Concrete commands, explicit assertions (status/exit_code/content), reproducible URLs |
|
|
120
|
+
| **Completion Coverage** | 10% | Prompt-mentioned status codes, tests, OpenAPI updates reflected in spec |
|
|
121
|
+
|
|
122
|
+
**Hard gate**: If any of the three runway stages (Inspect, Create/Modify, Verify) is missing, the spec scores **0.0** and is not runbook-ready.
|
|
123
|
+
|
|
124
|
+
The scorer now **honors explicit `type` on goals** — a goal with `type: create` counts as Create/Modify even if its verification is `file_exists` (executor will generate the file). Similarly `type: inspect` and `type: verify` map directly to stages.
|
|
125
|
+
|
|
126
|
+
## Quick Start
|
|
127
|
+
|
|
128
|
+
### 1. Setup
|
|
129
|
+
|
|
130
|
+
```bash
|
|
131
|
+
git clone <repo-url> && cd githeri
|
|
132
|
+
python3.11 -m venv .venv && .venv/bin/pip install -r requirements.txt
|
|
133
|
+
|
|
134
|
+
# For local generation (Ollama):
|
|
135
|
+
ollama pull specgen:latest
|
|
136
|
+
ollama pull qwen2.5-coder:7b-instruct # base model (fallback + code execution)
|
|
137
|
+
|
|
138
|
+
# For cloud generation (NVIDIA NIM - host any model like minimaxai/minimax-m3):
|
|
139
|
+
export NVIDIA_API_KEY=your_nvidia_api_key
|
|
140
|
+
|
|
141
|
+
# For LM Studio local execution (GRG pipeline):
|
|
142
|
+
# In LM Studio: enable "OpenAI Compatible Server" on port 1234
|
|
143
|
+
# Load qwen2.5-coder-14b-instruct-uncensored model
|
|
144
|
+
|
|
145
|
+
# For training: install on a GPU machine
|
|
146
|
+
pip install "unsloth[colab-new] @ git+https://github.com/unslothai/unsloth.git"
|
|
147
|
+
```
|
|
148
|
+
|
|
149
|
+
### 2. Generate specs
|
|
150
|
+
|
|
151
|
+
```bash
|
|
152
|
+
# Generate a validated spec from a fresh NL prompt (Ollama, specgen:latest)
|
|
153
|
+
make spec PROMPT="Add a POST /register endpoint that accepts email and password"
|
|
154
|
+
|
|
155
|
+
# Use NVIDIA NIM (default model: minimaxai/minimax-m3, base URL: https://integrate.api.nvidia.com/v1)
|
|
156
|
+
make spec PROMPT="Add a PATCH /users/{id}/settings endpoint" PROVIDER=nvidia API_KEY=$NVIDIA_API_KEY
|
|
157
|
+
|
|
158
|
+
# Use any OpenAI-compatible endpoint
|
|
159
|
+
make spec PROMPT="..." PROVIDER=openai-compat BASE_URL=https://api.fireworks.ai/inference/v1 API_KEY=... MODEL=accounts/nvidia/models/nemotron-3-ultra
|
|
160
|
+
|
|
161
|
+
# Generate N specs from random seed prompts (475 prompts across 21 categories)
|
|
162
|
+
make generate N=10
|
|
163
|
+
make generate N=100
|
|
164
|
+
|
|
165
|
+
# Generate ALL seed prompts in sequential order (full corpus generation)
|
|
166
|
+
make generate N=all
|
|
167
|
+
make generate-all
|
|
168
|
+
|
|
169
|
+
# Generate 10 random specs (explicit alias, good for short test runs)
|
|
170
|
+
make generate N=random
|
|
171
|
+
make generate-random
|
|
172
|
+
|
|
173
|
+
# Validate all specs in the corpus
|
|
174
|
+
make validate
|
|
175
|
+
|
|
176
|
+
# Score all specs against runbook criteria (standalone, post-generation)
|
|
177
|
+
make score
|
|
178
|
+
make score-failed # score invalid specs in data/failed_specs.jsonl
|
|
179
|
+
|
|
180
|
+
# Generate COMMAND_RUNWAY plan from a validated spec (two-step: prompt → LLM)
|
|
181
|
+
make plan-from-spec SPEC=data/training_data.jsonl#8
|
|
182
|
+
# Default: PLAN_MODEL=specgen:latest, PLAN_PROVIDER=ollama, OUTPUT_DIR=./out
|
|
183
|
+
# Output: ./out/PLAN.md — use PLAN_MODEL=qwen2.5-coder:7b-instruct for stronger planning
|
|
184
|
+
```
|
|
185
|
+
|
|
186
|
+
### 3. Convert + fine-tune + export
|
|
187
|
+
|
|
188
|
+
```bash
|
|
189
|
+
# Convert training data to chat format (filters by runbook score)
|
|
190
|
+
make convert-chat # default MIN_SCORE=0.75
|
|
191
|
+
make convert-chat MIN_SCORE=0.9 # only high-quality specs
|
|
192
|
+
|
|
193
|
+
# GRG Agent code training (generates verified code with real GRG scores)
|
|
194
|
+
make generate-code # cycles through all 475 SEED_PROMPTS
|
|
195
|
+
make convert-code-chat # converts passing examples to chat format
|
|
196
|
+
|
|
197
|
+
# LoRA fine-tuning (requires Unsloth + GPU with 8GB VRAM)
|
|
198
|
+
make train # outputs models/qwen2.5-coder-7b-specforge/
|
|
199
|
+
|
|
200
|
+
# Merge adapter + export GGUF for Ollama
|
|
201
|
+
make merge # outputs models/qwen2.5-coder-7b-specforge-gguf/
|
|
202
|
+
|
|
203
|
+
# Evaluate fine-tuned vs base model on held-out prompts
|
|
204
|
+
make eval-model # saves data/eval_results.json
|
|
205
|
+
|
|
206
|
+
# Register fine-tuned model in Ollama
|
|
207
|
+
ollama create specgen -f models/specgen/Modelfile
|
|
208
|
+
```
|
|
209
|
+
|
|
210
|
+
### 4. Upload to HuggingFace Hub
|
|
211
|
+
|
|
212
|
+
Model weights are stored in git (not LFS). The upload sends standard model files to HF Hub.
|
|
213
|
+
|
|
214
|
+
```bash
|
|
215
|
+
# Set up auth
|
|
216
|
+
echo "HF_TOKEN=hf_your_token_here" > .env
|
|
217
|
+
|
|
218
|
+
# Upload
|
|
219
|
+
make upload-hf REPO=githeri/specgen
|
|
220
|
+
make upload-hf REPO=githeri/specgen PRIVATE=1 # private repo
|
|
221
|
+
```
|
|
222
|
+
|
|
223
|
+
### 5. Install skill for agent use
|
|
224
|
+
|
|
225
|
+
The `skills/spec-forge/` directory is a self-contained Spec-Forge skill. Install it into your Hermes skills directory to use it from any session.
|
|
226
|
+
|
|
227
|
+
```bash
|
|
228
|
+
# Install to ~/.hermes/skills/ (default Hermes skills directory)
|
|
229
|
+
make install-skill # copies skills/spec-forge/ to ~/.hermes/skills/spec-forge/
|
|
230
|
+
|
|
231
|
+
# Install to a custom path (e.g. new server, project-local skills dir)
|
|
232
|
+
make install-skill-to-server HERMES_SKILLS_DIR=/path/to/hermes/skills
|
|
233
|
+
|
|
234
|
+
# Remove
|
|
235
|
+
make uninstall-skill
|
|
236
|
+
```
|
|
237
|
+
|
|
238
|
+
After install, load the skill in any Hermes session:
|
|
239
|
+
|
|
240
|
+
```bash
|
|
241
|
+
skill_view(name='spec-forge')
|
|
242
|
+
```
|
|
243
|
+
|
|
244
|
+
The skill provides: `make spec PROMPT="..."` (fresh NL → validated spec), `make spec-and-plan PROMPT="..."` (spec + plan prompt), and the bundled validator (`scripts/validator.py`) + plan assembler (`scripts/plan_from_spec.py`).
|
|
245
|
+
|
|
246
|
+
### 6. Plan from existing specs
|
|
247
|
+
|
|
248
|
+
Generate a COMMAND_RUNWAY plan from any validated spec in the corpus (two-step: extract + LLM plan generation):
|
|
249
|
+
|
|
250
|
+
```bash
|
|
251
|
+
# Default: specgen:latest, ollama, ./out/PLAN.md
|
|
252
|
+
make plan-from-spec SPEC=data/training_data.jsonl#8
|
|
253
|
+
|
|
254
|
+
# Stronger planning model (if available locally)
|
|
255
|
+
PLAN_MODEL=qwen2.5-coder:7b-instruct make plan-from-spec SPEC=data/training_data.jsonl#8
|
|
256
|
+
```
|
|
257
|
+
|
|
258
|
+
### 7. Interactive coding agent (`scripts/githeri.py`)
|
|
259
|
+
|
|
260
|
+
The project root ships a real interactive coding agent (like Codex/Claude Code). It has a persistent conversation, file operations, command execution, and model/provider switching inside the app.
|
|
261
|
+
|
|
262
|
+
```bash
|
|
263
|
+
# Interactive mode — start a conversation
|
|
264
|
+
.venv/bin/python scripts/githeri.py
|
|
265
|
+
|
|
266
|
+
# In-session commands:
|
|
267
|
+
# /model <name> - Change model (e.g. /model qwen2.5-coder:7b-instruct)
|
|
268
|
+
# /provider <name> - Change provider (e.g. /provider ollama)
|
|
269
|
+
# /clear - Clear conversation
|
|
270
|
+
# /help - Show help
|
|
271
|
+
# /quit, /exit - Exit
|
|
272
|
+
```
|
|
273
|
+
|
|
274
|
+
```bash
|
|
275
|
+
# Non-interactive: generate spec from NL prompt via specgen, then execute it
|
|
276
|
+
# Only two knobs: --prompt and exec model via env vars. Docker + spec model hardcoded.
|
|
277
|
+
GITHERI_MODEL=qwen2.5-coder:7b-instruct \
|
|
278
|
+
GITHERI_PROVIDER=ollama \
|
|
279
|
+
.venv/bin/python scripts/githeri.py \
|
|
280
|
+
--prompt "Add a PATCH endpoint to update user displayName and bio"
|
|
281
|
+
|
|
282
|
+
# Execute an existing spec.yaml (skip spec generation)
|
|
283
|
+
.venv/bin/python scripts/githeri.py --spec output/my-feature/spec.yaml
|
|
284
|
+
|
|
285
|
+
# Switch exec model — same prompt, different execution model
|
|
286
|
+
GITHERI_MODEL=qwen2.5-coder-14b-instruct-uncensored \
|
|
287
|
+
GITHERI_PROVIDER=lmstudio \
|
|
288
|
+
.venv/bin/python scripts/githeri.py --prompt "Add a /health endpoint"
|
|
289
|
+
|
|
290
|
+
# Overrides (rare):
|
|
291
|
+
# --no-docker - Disable Docker (not recommended)
|
|
292
|
+
# --workdir <path> - Custom workdir (default: output/sandbox)
|
|
293
|
+
# --docker-image <img> - Custom Docker image (default: githeri-sandbox)
|
|
294
|
+
# --server - Start server container after agent loop
|
|
295
|
+
# --test-cmd "<cmd>" - Run test command inside container via docker exec
|
|
296
|
+
```
|
|
297
|
+
|
|
298
|
+
**Defaults (hardcoded, never change between tests):**
|
|
299
|
+
- `--docker: True` — every command runs in throwaway `githeri-sandbox` container
|
|
300
|
+
- `--workdir: output/sandbox` — all file ops jailed here
|
|
301
|
+
- `--spec-model: specgen:latest` — hardcoded (fine-tuned for spec generation)
|
|
302
|
+
- `--spec-provider: ollama` — hardcoded
|
|
303
|
+
- Exec model/provider: `GITHERI_MODEL` / `GITHERI_PROVIDER` env vars only
|
|
304
|
+
|
|
305
|
+
**Pipeline flow:** 1) specgen:latest (Ollama) → NL prompt → spec.yaml; 2) exec model reads spec, creates files, installs deps, runs verifications — all inside Docker; 3) optional server container; 4) optional test command via docker exec.
|
|
306
|
+
|
|
307
|
+
See [docs/SPECGEN_HARNESS.md](docs/SPECGEN_HARNESS.md) for full architecture, troubleshooting, and output structure.
|
|
308
|
+
|
|
309
|
+
## Provider Configuration
|
|
310
|
+
|
|
311
|
+
All generation targets accept provider overrides:
|
|
312
|
+
|
|
313
|
+
```bash
|
|
314
|
+
# NVIDIA NIM (default: minimaxai/minimax-m3 at integrate.api.nvidia.com/v1)
|
|
315
|
+
make spec PROMPT="..." PROVIDER=nvidia API_KEY=$NVIDIA_API_KEY
|
|
316
|
+
|
|
317
|
+
# OpenAI GPT-4o
|
|
318
|
+
make spec PROMPT="..." PROVIDER=openai API_KEY=$OPENAI_API_KEY MODEL=gpt-4o
|
|
319
|
+
|
|
320
|
+
# Anthropic Claude
|
|
321
|
+
make spec PROMPT="..." PROVIDER=anthropic API_KEY=$ANTHROPIC_API_KEY MODEL=claude-3-5-sonnet-20241022
|
|
322
|
+
|
|
323
|
+
# LM Studio (local, OpenAI-compatible on port 1234)
|
|
324
|
+
make spec PROMPT="..." PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"
|
|
325
|
+
|
|
326
|
+
# Any OpenAI-compatible endpoint (Fireworks, Together, vLLM, self-hosted NIM)
|
|
327
|
+
make spec PROMPT="..." PROVIDER=openai-compat BASE_URL=https://api.fireworks.ai/inference/v1 API_KEY=... MODEL=accounts/nvidia/models/nemotron-3-ultra
|
|
328
|
+
|
|
329
|
+
# Override sampling
|
|
330
|
+
make generate N=5 PROVIDER=nvidia TEMPERATURE=0.3 MAX_TOKENS=4096
|
|
331
|
+
|
|
332
|
+
# NVIDIA NIM models with longer cold start: use bigger TIMEOUT
|
|
333
|
+
make generate N=1 PROVIDER=nvidia TIMEOUT=600
|
|
334
|
+
```
|
|
335
|
+
|
|
336
|
+
Environment variable fallbacks:
|
|
337
|
+
- `PROVIDER` → `ollama` (default)
|
|
338
|
+
- `MODEL` → provider-specific default
|
|
339
|
+
- `API_KEY` → `OPENAI_API_KEY`, `NVIDIA_API_KEY`, `ANTHROPIC_API_KEY`
|
|
340
|
+
- `BASE_URL` → `OPENAI_BASE_URL`, `NVIDIA_BASE_URL` (default `https://integrate.api.nvidia.com/v1`)
|
|
341
|
+
- `TEMPERATURE` → `0.2`
|
|
342
|
+
- `MAX_TOKENS` → `2048`
|
|
343
|
+
|
|
344
|
+
## Autonomous Execution System
|
|
345
|
+
|
|
346
|
+
The autonomous execution system takes a natural language prompt and produces a working implementation without human intervention in the loop. It uses the specgen pipeline to generate a validated spec, then uses the Command Runway skills to generate a plan and runbook, and finally executes the runbook using the GRG executor (with self-healing capabilities via GRG quality gates).
|
|
347
|
+
|
|
348
|
+
For detailed instructions, see [docs/SPRINTS_AUTONOMOUS.md](docs/SPRINTS_AUTONOMOUS.md).
|
|
349
|
+
|
|
350
|
+
### Quick Reference
|
|
351
|
+
|
|
352
|
+
```bash
|
|
353
|
+
# Basic local execution
|
|
354
|
+
.venv/bin/python scripts/spexecutor.py --prompt "Add a POST /notifications endpoint"
|
|
355
|
+
|
|
356
|
+
# Docker-isolated execution
|
|
357
|
+
.venv/bin/python scripts/spexecutor.py --prompt "Add user authentication" --docker
|
|
358
|
+
|
|
359
|
+
# Using Hermes as the execution backend
|
|
360
|
+
.venv/bin/python scripts/spexecutor.py --prompt "Implement rate limiting" --executor hermes
|
|
361
|
+
|
|
362
|
+
# Makefile: generate spec and plan
|
|
363
|
+
make spec-and-plan PROMPT="Add file upload endpoint"
|
|
364
|
+
|
|
365
|
+
# Makefile: full autonomous cycle (spec -> plan -> runbook -> execute -> report)
|
|
366
|
+
make autonomous-cycle SPEC=specs/test-endpoint.yaml
|
|
367
|
+
|
|
368
|
+
# GRG isolated execution (recommended for clean workspace)
|
|
369
|
+
make grg-full PROMPT="Add a POST /webhook endpoint that validates signature" PROVIDER=ollama
|
|
370
|
+
# Outputs: foreign/src/, foreign/tests/, foreign/RUNBOOK.md, foreign/RUNBOOK.json
|
|
371
|
+
```
|
|
372
|
+
|
|
373
|
+
## Recent Enhancements
|
|
374
|
+
|
|
375
|
+
### 2026-08-29 (This Session)
|
|
376
|
+
|
|
377
|
+
#### COMMAND_RUNWAY Plan Generation from Existing Specs (`make plan-from-spec`)
|
|
378
|
+
|
|
379
|
+
Added `make plan-from-spec SPEC=data/training_data.jsonl#<index>` — generates a full
|
|
380
|
+
COMMAND_RUNWAY plan from any validated spec in the corpus, in two steps:
|
|
381
|
+
|
|
382
|
+
1. **Extract** — load spec_yaml from JSONL by index (0-based) via `scripts/plan_from_spec.py`
|
|
383
|
+
2. **Generate** — build plan prompt + call LLM (`scripts/plan_from_spec_file.py` → `generate_plan()`)
|
|
384
|
+
|
|
385
|
+
Default: `PLAN_MODEL=specgen:latest`, `PLAN_PROVIDER=ollama`, `OUTPUT_DIR=./out`.
|
|
386
|
+
Output: `./out/PLAN.md` — full COMMAND_RUNWAY document with Feature, Target Environment,
|
|
387
|
+
and ordered Execution Stages (Objective, Verification, Completion Condition per stage).
|
|
388
|
+
|
|
389
|
+
Override for stronger plans: `PLAN_MODEL=qwen2.5-coder:7b-instruct PLAN_PROVIDER=ollama`.
|
|
390
|
+
|
|
391
|
+
Tested: 168-line plan with 3 execution stages generated for session-notes spec
|
|
392
|
+
(4176 chars, specgen:latest, ~60s). Requires `specgen:latest` (or any Ollama model)
|
|
393
|
+
to be pulled locally.
|
|
394
|
+
|
|
395
|
+
This unblocks the Sprint 9 workflow (prompt synthesis → bulk spec generation →
|
|
396
|
+
bulk plan generation → consolidated COMMAND_RUNWAY document) by providing the
|
|
397
|
+
per-spec plan generation step as a reusable Makefile target.
|
|
398
|
+
|
|
399
|
+
#### Scorer Fix: Accept `verification.type: content` as `file_exists` Alias
|
|
400
|
+
|
|
401
|
+
Both `qwen2.5-coder:7b-instruct` and `specgen:latest` emit `verification.type: "content"`
|
|
402
|
+
with `path` + `expect.content_contains` — structurally identical to `file_exists` but
|
|
403
|
+
not recognized by the scorer's canonical vocab. Fixed `scripts/runbook_scorer.py`:
|
|
404
|
+
|
|
405
|
+
- Added `"content"` to `REQUIRED_EXPECT_KEYS` (same keys as `file_exists`)
|
|
406
|
+
- `_check_inspect()` now treats `content` as inspect (read-only check)
|
|
407
|
+
- `_check_verify()` treats `content` with content_contains as verify evidence
|
|
408
|
+
- Verification testability section handles `content` same as `file_exists`
|
|
409
|
+
|
|
410
|
+
Before: avg score 0.104, 1/9 above 0.75.
|
|
411
|
+
After: avg score 0.972, 9/9 above 0.75.
|
|
412
|
+
|
|
413
|
+
#### Sprint 9 — Prompt Synthesis & Decomposition Pipeline (Planning)
|
|
414
|
+
|
|
415
|
+
Created `docs/SPRINTS.md` Sprint 9 (9a-9d) — four sub-sprints to accept raw unfocused
|
|
416
|
+
user text end-to-end:
|
|
417
|
+
|
|
418
|
+
- **9a** `scripts/prompt_synthesizer.py` — raw text → decomposed clean prompts (JSONL)
|
|
419
|
+
- **9b** `scripts/bulk_generate.py` — batch spec generation from prompt file
|
|
420
|
+
- **9c** `scripts/bulk_plan.py` — loop over specs → plans → consolidated doc
|
|
421
|
+
- **9d** `make text-to-plan TEXT="..."` — single target: synthesize → generate → score → plan
|
|
422
|
+
|
|
423
|
+
Independent of Sprints 5-8. Default models: specgen:latest for generation + planning.
|
|
424
|
+
|
|
425
|
+
#### Sprints 5, 6, 7 Marked Complete
|
|
426
|
+
|
|
427
|
+
- **Sprint 5** (Model provider switch) — already built: run_pipeline.py supports 7 providers
|
|
428
|
+
- **Sprint 6** (Skill bundling/install) — `install-skill` target body added to Makefile
|
|
429
|
+
- **Sprint 7** (Fine-tuning pipeline) — train*.py, merge*.py, eval*.py all exist; Makefile targets wired
|
|
430
|
+
|
|
431
|
+
`scripts/githeri.py` now supports a full end-to-end pipeline from natural language to tested code. **Spec generation is always handled by `specgen:latest` on Ollama** (finetuned for structured YAML output). **All execution and tests run inside a Docker sandbox by default** — the container is throwaway, so the host `.venv`/`.git` are never polluted across multiple test runs.
|
|
432
|
+
|
|
433
|
+
```bash
|
|
434
|
+
# Minimal invocation — only the prompt and exec model vary between tests.
|
|
435
|
+
# Everything else (docker, workdir, spec model/provider) is defaulted.
|
|
436
|
+
GITHERI_MODEL=qwen2.5-coder:7b-instruct \
|
|
437
|
+
GITHERI_PROVIDER=ollama \
|
|
438
|
+
.venv/bin/python scripts/githeri.py \
|
|
439
|
+
--prompt "Add a PATCH endpoint to update user displayName and bio"
|
|
440
|
+
```
|
|
441
|
+
|
|
442
|
+
**Defaults (set once, never touched between tests):**
|
|
443
|
+
|
|
444
|
+
- --docker: True — every run_command runs in throwaway githeri-sandbox container
|
|
445
|
+
- --workdir: output/sandbox — all file ops jailed here; created under repo root
|
|
446
|
+
- --spec-model: specgen:latest — **hardcoded** (finetuned for spec generation only)
|
|
447
|
+
- --spec-provider: ollama — **hardcoded**
|
|
448
|
+
- Exec model/provider: env vars GITHERI_MODEL / GITHERI_PROVIDER — only things you change between tests
|
|
449
|
+
- --docker-image: githeri-sandbox
|
|
450
|
+
|
|
451
|
+
**Only two knobs per test:** --prompt (or --spec) and the exec model via env vars. No flags for docker, workdir, or spec model — they're locked in.
|
|
452
|
+
|
|
453
|
+
```bash
|
|
454
|
+
# Switch exec model — same prompt, same defaults, different model
|
|
455
|
+
GITHERI_MODEL=qwen2.5-coder:7b-instruct \
|
|
456
|
+
GITHERI_PROVIDER=ollama \
|
|
457
|
+
.venv/bin/python scripts/githeri.py --prompt "Add a /health endpoint"
|
|
458
|
+
|
|
459
|
+
# Different provider for execution (specgen stays ollama/specgen)
|
|
460
|
+
GITHERI_MODEL=qwen2.5-14b-instruct-latest:2 \
|
|
461
|
+
GITHERI_PROVIDER=lmstudio \
|
|
462
|
+
.venv/bin/python scripts/githeri.py --prompt "Add a /health endpoint"
|
|
463
|
+
|
|
464
|
+
# Execute an existing spec.yaml instead of generating one
|
|
465
|
+
.venv/bin/python scripts/githeri.py --spec output/my-feature/spec.yaml
|
|
466
|
+
```
|
|
467
|
+
|
|
468
|
+
**Overrides (rare):**
|
|
469
|
+
|
|
470
|
+
```bash
|
|
471
|
+
# Disable docker (e.g. when Docker daemon is down) — not recommended for tests
|
|
472
|
+
.venv/bin/python scripts/githeri.py --prompt "..." --no-docker
|
|
473
|
+
|
|
474
|
+
# Custom workdir
|
|
475
|
+
.venv/bin/python scripts/githeri.py --prompt "..." --workdir output/custom-sandbox
|
|
476
|
+
|
|
477
|
+
# Custom docker image
|
|
478
|
+
.venv/bin/python scripts/githeri.py --prompt "..." --docker-image my-sandbox
|
|
479
|
+
```
|
|
480
|
+
|
|
481
|
+
**Pipeline flow:**
|
|
482
|
+
1. Specgen LLM (specgen:latest / Ollama) converts NL prompt → spec.yaml (with entrypoint, packages, goals)
|
|
483
|
+
2. Execution model (set via env vars) reads spec, creates files in output/sandbox, installs deps, runs verifications — all inside Docker
|
|
484
|
+
3. Server container starts from entrypoint.start_command (if --server)
|
|
485
|
+
4. Test command runs inside container via docker exec (if --test-cmd)
|
|
486
|
+
|
|
487
|
+
**Environment isolation (critical):**
|
|
488
|
+
- Docker is the default and recommended mode — every `run_command` (including `pip install`) runs in a throwaway container. The host `.venv` is never touched by spec-generated dependencies.
|
|
489
|
+
- `--no-docker` disables this isolation. When used, pip installs and code execution happen in the host environment (the project `.venv`), which can pollute it with spec-generated packages (Flask, FastAPI, etc.). Use `--no-docker` only when you understand this risk.
|
|
490
|
+
- If you must run without Docker, create a separate virtualenv for the workdir:
|
|
491
|
+
```bash
|
|
492
|
+
python -m venv /tmp/githeri-isolated
|
|
493
|
+
source /tmp/githeri-isolated/bin/activate
|
|
494
|
+
pip install -r requirements.txt # only githeri's own deps
|
|
495
|
+
GITHERI_MODEL=... GITHERI_PROVIDER=... \
|
|
496
|
+
scripts/githeri.py --prompt "..." --no-docker --workdir /tmp/githeri-workdir
|
|
497
|
+
```
|
|
498
|
+
|
|
499
|
+
**Model separation:** Spec generation uses specgen:latest (trained for structured YAML output, hardcoded). Execution uses whatever you set via GITHERI_MODEL / GITHERI_PROVIDER. Override via env vars only — the --spec-model / --spec-provider flags are suppressed (backwards-compatible but ignored).
|
|
500
|
+
|
|
501
|
+
**Server lifecycle:** After the agent loop completes, githeri parses entrypoint.start_command from the generated spec, starts a named Docker container (docker run -d), and runs --test-cmd inside it. Container stays running for manual inspection: docker exec <name> bash.
|
|
502
|
+
|
|
503
|
+
See [docs/SPECGEN_HARNESS.md](docs/SPECGEN_HARNESS.md) for full architecture, troubleshooting, and output structure.
|
|
504
|
+
|
|
505
|
+
### 2026-08-15 (This Session)
|
|
506
|
+
|
|
507
|
+
#### GRG Agent Skill — Hermes Native Integration
|
|
508
|
+
The GRG agent skill is now installed as a native Hermes skill (`~/.hermes/skills/autonomous-ai-agents/grg_agent/`) with full multi-provider support:
|
|
509
|
+
|
|
510
|
+
1. **Skill Installation** — Copied from `skills/grg_agent/` and installed via editable pip install
|
|
511
|
+
2. **Lightweight GRG Dependency** — Uses local `grg-0.1.0-py3-none-any.whl` wheel (no karakana dependency) providing `AlphaMomentumTracker` and `compute_structural_alpha`
|
|
512
|
+
3. **Hermes Proxy Support** — Skill accepts `provider` argument: `ollama` | `hermes` | `auto` — uses Hermes's configured providers (NVIDIA Nemotron, Nous Portal, xAI Grok, etc.) instead of local models
|
|
513
|
+
4. **Direct GRG Agent Execution** — `grg:execute` command now calls `self.agent.solve()` directly instead of legacy `run_pipeline.py` subprocess
|
|
514
|
+
5. **Make Target Integration** — `scripts/grg_make_spec.py` updated to use the skill with provider argument: `make grg-spec PROMPT="..." PROVIDER=ollama|hermes|auto`
|
|
515
|
+
6. **Project Virtual Environment** — Runs in project's own `.venv/` (not external karakana venv)
|
|
516
|
+
|
|
517
|
+
**Key Benefit**: You can now use cloud models via Hermes proxy (`hermes proxy start`) instead of relying on locally installed Ollama models. The skill routes through Hermes's provider config which supports NVIDIA Nemotron, Nous Portal, xAI Grok, and any OpenAI-compatible endpoint.
|
|
518
|
+
|
|
519
|
+
#### LM Studio Local Model — First End-to-End Working Pipeline
|
|
520
|
+
**LM Studio** with **`qwen2.5-coder-14b-instruct-uncensored`** is the first local model to complete the full GRG pipeline end-to-end:
|
|
521
|
+
|
|
522
|
+
```bash
|
|
523
|
+
# Full pipeline from NL prompt → validated spec → plan → execution → runbook
|
|
524
|
+
make grg-full PROMPT="Implement a FastAPI POST /api/health-check endpoint..." \
|
|
525
|
+
PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"
|
|
526
|
+
```
|
|
527
|
+
|
|
528
|
+
**What works:**
|
|
529
|
+
- **GRG Agent solving** — Generates implementation code with GRG quality gates (composite scoring, diversity control, convergence detection)
|
|
530
|
+
- **Code verification** — Execution-based verification (syntax + runtime) in isolated temp files
|
|
531
|
+
- **Multi-strategy generation** — Standard, decompose, test_first, refine strategies with adaptive temperature
|
|
532
|
+
- **Health-check endpoint example** — Generated FastAPI code with SQLAlchemy DB check + Redis cache check, returns 200/503
|
|
533
|
+
- **All artifacts isolated in `foreign/`** — Clean workspace separation
|
|
534
|
+
|
|
535
|
+
**Configuration for LM Studio:**
|
|
536
|
+
```bash
|
|
537
|
+
# In LM Studio: enable "OpenAI Compatible Server" on port 1234
|
|
538
|
+
# Load qwen2.5-coder-14b-instruct-uncensored model
|
|
539
|
+
# Then run:
|
|
540
|
+
make grg-spec PROMPT="..." PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"
|
|
541
|
+
```
|
|
542
|
+
|
|
543
|
+
**Why this model works:**
|
|
544
|
+
- 14B parameter coder model fine-tuned for code generation
|
|
545
|
+
- Uncensored variant removes alignment filters that can interfere with code structure
|
|
546
|
+
- OpenAI-compatible API in LM Studio works with GRG's `OllamaClient` (custom `api_key` support)
|
|
547
|
+
- Sufficient context window for spec + plan generation tasks
|
|
548
|
+
- Produces valid imports, proper error handling, and correct HTTP status codes
|
|
549
|
+
|
|
550
|
+
### 2026-07-30 (Latest)
|
|
551
|
+
|
|
552
|
+
#### THIRD_IMPROVE_SPEC — 7 Pipeline Fixes
|
|
553
|
+
After analyzing a 10-spec batch run, identified 7 recurring failure patterns and fixed all of them:
|
|
554
|
+
|
|
555
|
+
1. **Minimal structural skeleton in SYSTEM_PROMPT** — eliminated top-level field confusion
|
|
556
|
+
2. **task_id validator check** — rejects L11/G18-style IDs, requires descriptive slug like `jwt-auth-login`
|
|
557
|
+
3. **Verification types & expect keys table** — explicit vocabulary reduces hallucinated keys
|
|
558
|
+
4. **Placeholder ban strengthened** — concrete examples in prompt (`JWT_SECRET: "test-secret-..."`, `DATABASE_URL: "postgresql://user:***@..."`) + validator hint
|
|
559
|
+
5. **Helpful error hints** — "These fields belong inside a goal under `local_goals`, not at the spec root"
|
|
560
|
+
6. **depends_on validator hints** — rejects L/G refs, redirects to task_ids/stage names
|
|
561
|
+
7. **Structural acceptance_criteria template** — local goal field, not top-level
|
|
562
|
+
|
|
563
|
+
All 90 tests passing.
|
|
564
|
+
|
|
565
|
+
#### Generation Mode Aliases
|
|
566
|
+
Added `make generate N=random` and `make generate N=all` flags plus `generate-random` / `generate-all` aliases for short tests and full corpus runs. 475 seed prompts across 21 categories available.
|
|
567
|
+
|
|
568
|
+
#### Model Default Reverted
|
|
569
|
+
`OLLAMA_MODEL` reverted to `qwen2.5-coder:7b-instruct` (was regressed to `qwen3.5-4b-128k:latest`). Better structure compliance after prompt fixes.
|
|
570
|
+
|
|
571
|
+
### 2026-07-30 (Earlier)
|
|
572
|
+
|
|
573
|
+
#### Spec Enrichment (IMPROVE_SPEC)
|
|
574
|
+
Added five top-level enrichment fields (`business_rules`, `test_fixtures`, `environment`, `global_verification`) and three goal-level fields (`blueprint`, `acceptance_criteria`, `type`). All validated and scored.
|
|
575
|
+
|
|
576
|
+
#### Multi-Provider LLM Support
|
|
577
|
+
`run_pipeline.py` now supports Ollama, OpenAI, Anthropic, NVIDIA (via Together AI), and any OpenAI-compatible endpoint. Configured via `--provider` CLI arg or Makefile variables.
|
|
578
|
+
|
|
579
|
+
#### Runbook Scorer Stage Detection
|
|
580
|
+
Scorer now honors explicit `type: create|inspect|verify` on goals, fixing false "missing stage" penalties for `file_exists` verification on CREATE goals.
|
|
581
|
+
|
|
582
|
+
#### Validator Hardening
|
|
583
|
+
- Guarded against `expect` being a string instead of dict (prevents `AttributeError: 'str' object has no attribute 'get'`)
|
|
584
|
+
- Near-duplicate detection now handles malformed `expect` blocks defensively
|
|
585
|
+
- All 90 tests pass
|
|
586
|
+
|
|
587
|
+
## Model Compatibility Matrix
|
|
588
|
+
|
|
589
|
+
| Model | Context | Tools | Speed (M1 16GB) | Spec Gen | Recommended Use |
|
|
590
|
+
|-------|---------|-------|-----------------|----------|-----------------|
|
|
591
|
+
| **specgen:latest** | 128K | Yes | Fast | **Best first-attempt pass rate (fine-tuned)** | **Default local (Ollama) — spec generation** |
|
|
592
|
+
|| qwen2.5-coder:7b-instruct | 32K | Yes | Fast | Works (best structure compliance) | Fallback / base model for training |
|
|
593
|
+
|| qwen2.5-coder-14b-instruct-uncensored | 32K | Yes | Medium | **Works end-to-end (GRG pipeline)** | **LM Studio local** |
|
|
594
|
+
|| qwen3.5-4b-128k | 128K | Yes | Fast | Works | Larger context fallback |
|
|
595
|
+
|| qwen3.5-9b-code:128k | 128K | Yes | Slow | Excellent | Higher quality if time allows |
|
|
596
|
+
|| deepseek-r1:7b | 128K | TBD | Fast | YAML syntax errors | Not recommended for spec gen |
|
|
597
|
+
|| **Nemotron 3 Ultra** | 128K | Yes | Fast | Excellent | **Cloud via NVIDIA NIM (`--provider nvidia`)** |
|
|
598
|
+
|
|
599
|
+
**Current recommendation**:
|
|
600
|
+
- **Spec generation (default)** → `specgen:latest` on Ollama (fine-tuned for structured YAML output, best first-attempt pass rate)
|
|
601
|
+
- **Bulk corpus** → `specgen:latest` on Ollama (fine-tuned, best first-attempt pass rate)
|
|
602
|
+
- **Specific features (GRG pipeline)** → `qwen2.5-coder-14b-instruct-uncensored` on LM Studio with `make grg-full` (100% success via execution verification)
|
|
603
|
+
- **Cloud** → `minimaxai/minimax-m3` via NVIDIA NIM (`--provider nvidia`)
|
|
604
|
+
|
|
605
|
+
## Training Data Generation Strategies
|
|
606
|
+
|
|
607
|
+
Based on empirical testing (2026-08-16), here are the reliable approaches:
|
|
608
|
+
|
|
609
|
+
### 1. Bulk Corpus Generation (Recommended for Training Data)
|
|
610
|
+
```bash
|
|
611
|
+
# Fast, decent success rate, fine-tuned for this pipeline
|
|
612
|
+
make generate N=100 PROVIDER=ollama MODEL="specgen:latest" BASE_URL="http://localhost:11434" TEMPERATURE=0.2 MAX_TOKENS=4096
|
|
613
|
+
```
|
|
614
|
+
- **Success rate**: ~70%+ (fine-tuned specgen:latest)
|
|
615
|
+
- **Speed**: ~50-100s/spec
|
|
616
|
+
- **Best for**: Generating large training corpora quickly
|
|
617
|
+
|
|
618
|
+
### 2. High-Quality Individual Specs (GRG Pipeline)
|
|
619
|
+
```bash
|
|
620
|
+
# Best for specific features - multi-strategy + execution verification
|
|
621
|
+
make grg-full PROMPT="Your feature" PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"
|
|
622
|
+
```
|
|
623
|
+
- **Success rate**: ~100% (execution-verified)
|
|
624
|
+
- **Speed**: ~260s/spec
|
|
625
|
+
- **Best for**: Critical features requiring guaranteed working specs
|
|
626
|
+
|
|
627
|
+
### 3. GRG Agent Code Training (NEW — Real GRG Scores + Verified Code)
|
|
628
|
+
```bash
|
|
629
|
+
# Generates verified code with real GRG composite scores (not dummy 0.5)
|
|
630
|
+
# Uses qwen2.5-coder:7b-instruct on Ollama for real logprobs
|
|
631
|
+
make generate-code # cycles through all 475 SEED_PROMPTS
|
|
632
|
+
make convert-code-chat # converts passing examples to chat format
|
|
633
|
+
```
|
|
634
|
+
- **Success rate**: Variable (depends on prompt complexity)
|
|
635
|
+
- **Speed**: ~30-60s/prompt (optimized: 1 strategy, 1 candidate)
|
|
636
|
+
- **Output**: `data/training_data_code.jsonl` + `data/training_data_code_chat.jsonl`
|
|
637
|
+
- **Key difference**: Produces **executable code** with **real GRG composite scores** (0.47-0.49) because the model provides logprobs
|
|
638
|
+
- **Verification**: Syntax check + execution test (import + basic run)
|
|
639
|
+
- **Best for**: Fine-tuning code generation models with GRG quality signals
|
|
640
|
+
|
|
641
|
+
**Configuration** (in `scripts/generate_code_training_fast.py`):
|
|
642
|
+
```python
|
|
643
|
+
skill = create_skill(config={
|
|
644
|
+
'llm_provider': 'ollama',
|
|
645
|
+
'ollama_base_url': 'http://127.0.0.1:11434/v1',
|
|
646
|
+
'ollama_default_model': 'qwen2.5-coder:7b-instruct',
|
|
647
|
+
'max_iterations': 2, # Must be >=2 for convergence check
|
|
648
|
+
'temperature': 0.3,
|
|
649
|
+
'top_p': 0.9,
|
|
650
|
+
'max_tokens': 1024,
|
|
651
|
+
'candidates_per_strategy': 1, # Speed optimization
|
|
652
|
+
'max_strategies': 1, # Speed optimization
|
|
653
|
+
})
|
|
654
|
+
```
|
|
655
|
+
|
|
656
|
+
### 4. Hybrid Workflow (Best of Both)
|
|
657
|
+
```bash
|
|
658
|
+
# 1. Generate bulk corpus with specgen
|
|
659
|
+
make generate N=100 PROVIDER=ollama MODEL="specgen:latest" BASE_URL="http://localhost:11434"
|
|
660
|
+
|
|
661
|
+
# 2. Score and filter high-quality specs
|
|
662
|
+
make score MIN_SCORE=0.75
|
|
663
|
+
|
|
664
|
+
# 3. Re-generate failed critical specs with GRG pipeline
|
|
665
|
+
make grg-full PROMPT="..." PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"
|
|
666
|
+
|
|
667
|
+
# 4. Generate verified code for fine-tuning
|
|
668
|
+
make generate-code
|
|
669
|
+
make convert-code-chat
|
|
670
|
+
```
|
|
671
|
+
|
|
672
|
+
### Observed Failure Patterns (specgen:latest)
|
|
673
|
+
The remaining failures (if any) are primarily:
|
|
674
|
+
1. **Near-duplicate HTTP verifications** - multiple goals hitting same endpoint with same method
|
|
675
|
+
2. **Missing `acceptance_criteria`** for CREATE goals
|
|
676
|
+
3. **Placeholder values** in headers (e.g., `Authorization: *** ***`)
|
|
677
|
+
4. **YAML block mapping errors** - CLI verification indentation issues
|
|
678
|
+
|
|
679
|
+
## For More Information
|
|
680
|
+
|
|
681
|
+
- [docs/SPRINTS_AUTONOMOUS.md](docs/SPRINTS_AUTONOMOUS.md) — sprint breakdown, model experiments, decisions, next steps
|
|
682
|
+
- [docs/scoring_spec.md](docs/scoring_spec.md) — runbook scoring specification
|
|
683
|
+
- [docs/IMPROVE_SPEC.md](docs/IMPROVE_SPEC.md) — spec enrichment field specification (v1)
|
|
684
|
+
- [docs/SECOND_IMPROVE_SPEC.md](docs/SECOND_IMPROVE_SPEC.md) — enrichment field enforcement (v2)
|
|
685
|
+
- [docs/THIRD_IMPROVE_SPEC.txt](docs/THIRD_IMPROVE_SPEC.txt) — 7 pipeline failure patterns + fixes (v3)
|
|
686
|
+
- [MODEL_CARD.md](MODEL_CARD.md) — model card (uploaded to HF Hub)
|
|
687
|
+
- [skills/spec-forge/SKILL.md](skills/spec-forge/SKILL.md) — Spec-Forge skill reference
|
|
688
|
+
- [skills/command-runway-pattern/SKILL.md](skills/command-runway-pattern/SKILL.md) — Command Runway pattern skill
|
|
689
|
+
- [skills/grg_agent/SKILL.md](skills/grg_agent/SKILL.md) — GRG Agent skill reference
|