githeri 0.2.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
githeri-0.2.0/PKG-INFO ADDED
@@ -0,0 +1,689 @@
1
+ Metadata-Version: 2.4
2
+ Name: githeri
3
+ Version: 0.2.0
4
+ Summary: Contract Engine for AI Coding Agents. Deterministic specs, test plans, and drift detection.
5
+ Author-email: Karakana Labs <dev@karakanalabs.com>
6
+ License-Expression: Apache-2.0
7
+ Project-URL: Homepage, https://githeri.com
8
+ Project-URL: Documentation, https://githeri.com/docs
9
+ Project-URL: Repository, https://github.com/karakana-labs/githeri
10
+ Project-URL: Issues, https://github.com/karakana-labs/githeri/issues
11
+ Keywords: ai,agents,spec,fastapi,testing,openapi,contract-testing
12
+ Classifier: Development Status :: 4 - Beta
13
+ Classifier: Intended Audience :: Developers
14
+ Classifier: Programming Language :: Python :: 3
15
+ Classifier: Programming Language :: Python :: 3.10
16
+ Classifier: Programming Language :: Python :: 3.11
17
+ Classifier: Programming Language :: Python :: 3.12
18
+ Classifier: Topic :: Software Development :: Build Tools
19
+ Classifier: Topic :: Software Development :: Testing
20
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
21
+ Requires-Python: >=3.10
22
+ Description-Content-Type: text/markdown
23
+ Requires-Dist: typer>=0.12.0
24
+ Requires-Dist: rich>=13.0.0
25
+ Requires-Dist: requests>=2.28.0
26
+ Requires-Dist: pyyaml>=6.0
27
+ Requires-Dist: pydantic>=2.0.0
28
+ Requires-Dist: python-dotenv>=1.0.0
29
+
30
+ # githeri
31
+
32
+ **Spec-Forge**: a pipeline that trains an AI to bridge the gap between a human's natural-language feature request and a machine-executable specification that feeds the COMMAND_RUNWAY skill.
33
+
34
+ The central thesis: most time in AI-assisted development is lost before the AI and the human agree on what to build. Githeri attacks that by producing validated, structured YAML specs — which then feed into COMMAND_RUNWAY runbooks that an executor agent follows verbatim.
35
+
36
+ ## Pipeline
37
+
38
+ ```text
39
+ Human natural-language request
40
+ │
41
+ � ▼
42
+ Spec-Forge (LLM + validator)
43
+ Generates:
44
+ data/training_data.jsonl (valid prompt + spec_yaml pairs)
45
+ data/failed_specs.jsonl (invalid specs, saved for analysis)
46
+ │
47
+ � ▼
48
+ Runbook Scorer (standalone, post-generation)
49
+ Scores:
50
+ 5 categories — Intent, Preconditions, Structure, Testability, Coverage
51
+ Hard gate: missing Inspect/Create/Verify = 0.0
52
+ │
53
+ � ▼
54
+ Human Review
55
+ "L1-L4 look good. Approve."
56
+ │
57
+ � ▼
58
+ COMMAND_RUNWAY Skill
59
+ Consumes:
60
+ a validated spec (single-feature YAML)
61
+ Produces:
62
+ COMMAND_RUNWAY.md
63
+ • ordered implementation plan
64
+ • exact file paths
65
+ • code modifications
66
+ • test skeletons
67
+ • verification commands (translated from spec local_goals)
68
+ • rollback guidance
69
+ • completion criteria
70
+ │
71
+ � ▼
72
+ GRG Executor (with COMMAND_RUNWAY pattern integration)
73
+ Consumes:
74
+ a validated spec OR COMMAND_RUNWAY plan JSON
75
+ Produces (all under foreign/ directory):
76
+ • implementation source files
77
+ • test files
78
+ • RUNBOOK.md (human-readable execution log with GRG scores)
79
+ • RUNBOOK.json (machine-readable execution data)
80
+ • automatic ruff check --fix on generated code
81
+ │
82
+ � ▼
83
+ Completed Feature
84
+ Outputs:
85
+ • implementation complete
86
+ • all tests passing
87
+ • OpenAPI updated
88
+ • documentation synchronized
89
+ • human notified
90
+ ```
91
+
92
+ ## Spec Enrichment (IMPROVE_SPEC)
93
+
94
+ Every generated spec now includes optional **enrichment fields** that make specs machine-executable:
95
+
96
+ | Field | Location | Purpose |
97
+ |-------|----------|---------|
98
+ | `business_rules` | top-level | Invariants & formulas (e.g., "JWT Secret: 256-bit random, rotated quarterly") |
99
+ | `test_fixtures` | top-level | Seed data & setup commands (e.g., `.venv/bin/python scripts/seed_admin.py`) |
100
+ | `environment` | top-level | Required packages + env vars (e.g., `pyyaml>=6.0`, `JWT_SECRET`) |
101
+ | `global_verification` | top-level | Post-execution gate commands (e.g., `pytest tests/`, `bandit -r src/`) |
102
+ | `blueprint` | per-goal | **Required for `type: create`** (≥100 chars). Code-level outline: class signatures, route decorators, SQLAlchemy models, business logic steps |
103
+ | `acceptance_criteria` | per-goal | List of `{test, steps}` — executable test cases in pseudo-code |
104
+ | `type` | per-goal | `create` \| `update` \| `delete` \| `inspect` \| `verify` — drives runbook stage classification |
105
+
106
+ These fields are validated by `scripts/validator.py` and consumed by downstream generators (plan, runbook, scorer).
107
+
108
+ ## Runbook Scoring System
109
+
110
+ Every generated spec is scored against runbook-readiness criteria (see `docs/scoring_spec.md`). Scoring is decoupled from generation — specs are saved first, then scored in a separate pass via `make score`.
111
+
112
+ The scorer (`scripts/runbook_scorer.py`) evaluates five weighted categories:
113
+
114
+ | Category | Weight | Key Checks |
115
+ |----------|--------|------------|
116
+ | **Intent & Goals** | 20% | Summary present, goals have descriptions, endpoint tasks have HTTP verification |
117
+ | **Preconditions** | 15% | `depends_on` references valid globals/stages, CLI tools declared in context |
118
+ | **Command Runway Structure** | 30% | **Hard gate**: must have Inspect (file_exists/read CLI), Create/Modify (build CLI), Verify (HTTP/test CLI). Stage order: Inspect → Create → Verify |
119
+ | **Verification Testability** | 25% | Concrete commands, explicit assertions (status/exit_code/content), reproducible URLs |
120
+ | **Completion Coverage** | 10% | Prompt-mentioned status codes, tests, OpenAPI updates reflected in spec |
121
+
122
+ **Hard gate**: If any of the three runway stages (Inspect, Create/Modify, Verify) is missing, the spec scores **0.0** and is not runbook-ready.
123
+
124
+ The scorer now **honors explicit `type` on goals** — a goal with `type: create` counts as Create/Modify even if its verification is `file_exists` (executor will generate the file). Similarly `type: inspect` and `type: verify` map directly to stages.
125
+
126
+ ## Quick Start
127
+
128
+ ### 1. Setup
129
+
130
+ ```bash
131
+ git clone <repo-url> && cd githeri
132
+ python3.11 -m venv .venv && .venv/bin/pip install -r requirements.txt
133
+
134
+ # For local generation (Ollama):
135
+ ollama pull specgen:latest
136
+ ollama pull qwen2.5-coder:7b-instruct # base model (fallback + code execution)
137
+
138
+ # For cloud generation (NVIDIA NIM - host any model like minimaxai/minimax-m3):
139
+ export NVIDIA_API_KEY=your_nvidia_api_key
140
+
141
+ # For LM Studio local execution (GRG pipeline):
142
+ # In LM Studio: enable "OpenAI Compatible Server" on port 1234
143
+ # Load qwen2.5-coder-14b-instruct-uncensored model
144
+
145
+ # For training: install on a GPU machine
146
+ pip install "unsloth[colab-new] @ git+https://github.com/unslothai/unsloth.git"
147
+ ```
148
+
149
+ ### 2. Generate specs
150
+
151
+ ```bash
152
+ # Generate a validated spec from a fresh NL prompt (Ollama, specgen:latest)
153
+ make spec PROMPT="Add a POST /register endpoint that accepts email and password"
154
+
155
+ # Use NVIDIA NIM (default model: minimaxai/minimax-m3, base URL: https://integrate.api.nvidia.com/v1)
156
+ make spec PROMPT="Add a PATCH /users/{id}/settings endpoint" PROVIDER=nvidia API_KEY=$NVIDIA_API_KEY
157
+
158
+ # Use any OpenAI-compatible endpoint
159
+ make spec PROMPT="..." PROVIDER=openai-compat BASE_URL=https://api.fireworks.ai/inference/v1 API_KEY=... MODEL=accounts/nvidia/models/nemotron-3-ultra
160
+
161
+ # Generate N specs from random seed prompts (475 prompts across 21 categories)
162
+ make generate N=10
163
+ make generate N=100
164
+
165
+ # Generate ALL seed prompts in sequential order (full corpus generation)
166
+ make generate N=all
167
+ make generate-all
168
+
169
+ # Generate 10 random specs (explicit alias, good for short test runs)
170
+ make generate N=random
171
+ make generate-random
172
+
173
+ # Validate all specs in the corpus
174
+ make validate
175
+
176
+ # Score all specs against runbook criteria (standalone, post-generation)
177
+ make score
178
+ make score-failed # score invalid specs in data/failed_specs.jsonl
179
+
180
+ # Generate COMMAND_RUNWAY plan from a validated spec (two-step: prompt → LLM)
181
+ make plan-from-spec SPEC=data/training_data.jsonl#8
182
+ # Default: PLAN_MODEL=specgen:latest, PLAN_PROVIDER=ollama, OUTPUT_DIR=./out
183
+ # Output: ./out/PLAN.md — use PLAN_MODEL=qwen2.5-coder:7b-instruct for stronger planning
184
+ ```
185
+
186
+ ### 3. Convert + fine-tune + export
187
+
188
+ ```bash
189
+ # Convert training data to chat format (filters by runbook score)
190
+ make convert-chat # default MIN_SCORE=0.75
191
+ make convert-chat MIN_SCORE=0.9 # only high-quality specs
192
+
193
+ # GRG Agent code training (generates verified code with real GRG scores)
194
+ make generate-code # cycles through all 475 SEED_PROMPTS
195
+ make convert-code-chat # converts passing examples to chat format
196
+
197
+ # LoRA fine-tuning (requires Unsloth + GPU with 8GB VRAM)
198
+ make train # outputs models/qwen2.5-coder-7b-specforge/
199
+
200
+ # Merge adapter + export GGUF for Ollama
201
+ make merge # outputs models/qwen2.5-coder-7b-specforge-gguf/
202
+
203
+ # Evaluate fine-tuned vs base model on held-out prompts
204
+ make eval-model # saves data/eval_results.json
205
+
206
+ # Register fine-tuned model in Ollama
207
+ ollama create specgen -f models/specgen/Modelfile
208
+ ```
209
+
210
+ ### 4. Upload to HuggingFace Hub
211
+
212
+ Model weights are stored in git (not LFS). The upload sends standard model files to HF Hub.
213
+
214
+ ```bash
215
+ # Set up auth
216
+ echo "HF_TOKEN=hf_your_token_here" > .env
217
+
218
+ # Upload
219
+ make upload-hf REPO=githeri/specgen
220
+ make upload-hf REPO=githeri/specgen PRIVATE=1 # private repo
221
+ ```
222
+
223
+ ### 5. Install skill for agent use
224
+
225
+ The `skills/spec-forge/` directory is a self-contained Spec-Forge skill. Install it into your Hermes skills directory to use it from any session.
226
+
227
+ ```bash
228
+ # Install to ~/.hermes/skills/ (default Hermes skills directory)
229
+ make install-skill # copies skills/spec-forge/ to ~/.hermes/skills/spec-forge/
230
+
231
+ # Install to a custom path (e.g. new server, project-local skills dir)
232
+ make install-skill-to-server HERMES_SKILLS_DIR=/path/to/hermes/skills
233
+
234
+ # Remove
235
+ make uninstall-skill
236
+ ```
237
+
238
+ After install, load the skill in any Hermes session:
239
+
240
+ ```bash
241
+ skill_view(name='spec-forge')
242
+ ```
243
+
244
+ The skill provides: `make spec PROMPT="..."` (fresh NL → validated spec), `make spec-and-plan PROMPT="..."` (spec + plan prompt), and the bundled validator (`scripts/validator.py`) + plan assembler (`scripts/plan_from_spec.py`).
245
+
246
+ ### 6. Plan from existing specs
247
+
248
+ Generate a COMMAND_RUNWAY plan from any validated spec in the corpus (two-step: extract + LLM plan generation):
249
+
250
+ ```bash
251
+ # Default: specgen:latest, ollama, ./out/PLAN.md
252
+ make plan-from-spec SPEC=data/training_data.jsonl#8
253
+
254
+ # Stronger planning model (if available locally)
255
+ PLAN_MODEL=qwen2.5-coder:7b-instruct make plan-from-spec SPEC=data/training_data.jsonl#8
256
+ ```
257
+
258
+ ### 7. Interactive coding agent (`scripts/githeri.py`)
259
+
260
+ The project root ships a real interactive coding agent (like Codex/Claude Code). It has a persistent conversation, file operations, command execution, and model/provider switching inside the app.
261
+
262
+ ```bash
263
+ # Interactive mode — start a conversation
264
+ .venv/bin/python scripts/githeri.py
265
+
266
+ # In-session commands:
267
+ # /model <name> - Change model (e.g. /model qwen2.5-coder:7b-instruct)
268
+ # /provider <name> - Change provider (e.g. /provider ollama)
269
+ # /clear - Clear conversation
270
+ # /help - Show help
271
+ # /quit, /exit - Exit
272
+ ```
273
+
274
+ ```bash
275
+ # Non-interactive: generate spec from NL prompt via specgen, then execute it
276
+ # Only two knobs: --prompt and exec model via env vars. Docker + spec model hardcoded.
277
+ GITHERI_MODEL=qwen2.5-coder:7b-instruct \
278
+ GITHERI_PROVIDER=ollama \
279
+ .venv/bin/python scripts/githeri.py \
280
+ --prompt "Add a PATCH endpoint to update user displayName and bio"
281
+
282
+ # Execute an existing spec.yaml (skip spec generation)
283
+ .venv/bin/python scripts/githeri.py --spec output/my-feature/spec.yaml
284
+
285
+ # Switch exec model — same prompt, different execution model
286
+ GITHERI_MODEL=qwen2.5-coder-14b-instruct-uncensored \
287
+ GITHERI_PROVIDER=lmstudio \
288
+ .venv/bin/python scripts/githeri.py --prompt "Add a /health endpoint"
289
+
290
+ # Overrides (rare):
291
+ # --no-docker - Disable Docker (not recommended)
292
+ # --workdir <path> - Custom workdir (default: output/sandbox)
293
+ # --docker-image <img> - Custom Docker image (default: githeri-sandbox)
294
+ # --server - Start server container after agent loop
295
+ # --test-cmd "<cmd>" - Run test command inside container via docker exec
296
+ ```
297
+
298
+ **Defaults (hardcoded, never change between tests):**
299
+ - `--docker: True` — every command runs in throwaway `githeri-sandbox` container
300
+ - `--workdir: output/sandbox` — all file ops jailed here
301
+ - `--spec-model: specgen:latest` — hardcoded (fine-tuned for spec generation)
302
+ - `--spec-provider: ollama` — hardcoded
303
+ - Exec model/provider: `GITHERI_MODEL` / `GITHERI_PROVIDER` env vars only
304
+
305
+ **Pipeline flow:** 1) specgen:latest (Ollama) → NL prompt → spec.yaml; 2) exec model reads spec, creates files, installs deps, runs verifications — all inside Docker; 3) optional server container; 4) optional test command via docker exec.
306
+
307
+ See [docs/SPECGEN_HARNESS.md](docs/SPECGEN_HARNESS.md) for full architecture, troubleshooting, and output structure.
308
+
309
+ ## Provider Configuration
310
+
311
+ All generation targets accept provider overrides:
312
+
313
+ ```bash
314
+ # NVIDIA NIM (default: minimaxai/minimax-m3 at integrate.api.nvidia.com/v1)
315
+ make spec PROMPT="..." PROVIDER=nvidia API_KEY=$NVIDIA_API_KEY
316
+
317
+ # OpenAI GPT-4o
318
+ make spec PROMPT="..." PROVIDER=openai API_KEY=$OPENAI_API_KEY MODEL=gpt-4o
319
+
320
+ # Anthropic Claude
321
+ make spec PROMPT="..." PROVIDER=anthropic API_KEY=$ANTHROPIC_API_KEY MODEL=claude-3-5-sonnet-20241022
322
+
323
+ # LM Studio (local, OpenAI-compatible on port 1234)
324
+ make spec PROMPT="..." PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"
325
+
326
+ # Any OpenAI-compatible endpoint (Fireworks, Together, vLLM, self-hosted NIM)
327
+ make spec PROMPT="..." PROVIDER=openai-compat BASE_URL=https://api.fireworks.ai/inference/v1 API_KEY=... MODEL=accounts/nvidia/models/nemotron-3-ultra
328
+
329
+ # Override sampling
330
+ make generate N=5 PROVIDER=nvidia TEMPERATURE=0.3 MAX_TOKENS=4096
331
+
332
+ # NVIDIA NIM models with longer cold start: use bigger TIMEOUT
333
+ make generate N=1 PROVIDER=nvidia TIMEOUT=600
334
+ ```
335
+
336
+ Environment variable fallbacks:
337
+ - `PROVIDER` → `ollama` (default)
338
+ - `MODEL` → provider-specific default
339
+ - `API_KEY` → `OPENAI_API_KEY`, `NVIDIA_API_KEY`, `ANTHROPIC_API_KEY`
340
+ - `BASE_URL` → `OPENAI_BASE_URL`, `NVIDIA_BASE_URL` (default `https://integrate.api.nvidia.com/v1`)
341
+ - `TEMPERATURE` → `0.2`
342
+ - `MAX_TOKENS` → `2048`
343
+
344
+ ## Autonomous Execution System
345
+
346
+ The autonomous execution system takes a natural language prompt and produces a working implementation without human intervention in the loop. It uses the specgen pipeline to generate a validated spec, then uses the Command Runway skills to generate a plan and runbook, and finally executes the runbook using the GRG executor (with self-healing capabilities via GRG quality gates).
347
+
348
+ For detailed instructions, see [docs/SPRINTS_AUTONOMOUS.md](docs/SPRINTS_AUTONOMOUS.md).
349
+
350
+ ### Quick Reference
351
+
352
+ ```bash
353
+ # Basic local execution
354
+ .venv/bin/python scripts/spexecutor.py --prompt "Add a POST /notifications endpoint"
355
+
356
+ # Docker-isolated execution
357
+ .venv/bin/python scripts/spexecutor.py --prompt "Add user authentication" --docker
358
+
359
+ # Using Hermes as the execution backend
360
+ .venv/bin/python scripts/spexecutor.py --prompt "Implement rate limiting" --executor hermes
361
+
362
+ # Makefile: generate spec and plan
363
+ make spec-and-plan PROMPT="Add file upload endpoint"
364
+
365
+ # Makefile: full autonomous cycle (spec -> plan -> runbook -> execute -> report)
366
+ make autonomous-cycle SPEC=specs/test-endpoint.yaml
367
+
368
+ # GRG isolated execution (recommended for clean workspace)
369
+ make grg-full PROMPT="Add a POST /webhook endpoint that validates signature" PROVIDER=ollama
370
+ # Outputs: foreign/src/, foreign/tests/, foreign/RUNBOOK.md, foreign/RUNBOOK.json
371
+ ```
372
+
373
+ ## Recent Enhancements
374
+
375
+ ### 2026-08-29 (This Session)
376
+
377
+ #### COMMAND_RUNWAY Plan Generation from Existing Specs (`make plan-from-spec`)
378
+
379
+ Added `make plan-from-spec SPEC=data/training_data.jsonl#<index>` — generates a full
380
+ COMMAND_RUNWAY plan from any validated spec in the corpus, in two steps:
381
+
382
+ 1. **Extract** — load spec_yaml from JSONL by index (0-based) via `scripts/plan_from_spec.py`
383
+ 2. **Generate** — build plan prompt + call LLM (`scripts/plan_from_spec_file.py` → `generate_plan()`)
384
+
385
+ Default: `PLAN_MODEL=specgen:latest`, `PLAN_PROVIDER=ollama`, `OUTPUT_DIR=./out`.
386
+ Output: `./out/PLAN.md` — full COMMAND_RUNWAY document with Feature, Target Environment,
387
+ and ordered Execution Stages (Objective, Verification, Completion Condition per stage).
388
+
389
+ Override for stronger plans: `PLAN_MODEL=qwen2.5-coder:7b-instruct PLAN_PROVIDER=ollama`.
390
+
391
+ Tested: 168-line plan with 3 execution stages generated for session-notes spec
392
+ (4176 chars, specgen:latest, ~60s). Requires `specgen:latest` (or any Ollama model)
393
+ to be pulled locally.
394
+
395
+ This unblocks the Sprint 9 workflow (prompt synthesis → bulk spec generation →
396
+ bulk plan generation → consolidated COMMAND_RUNWAY document) by providing the
397
+ per-spec plan generation step as a reusable Makefile target.
398
+
399
+ #### Scorer Fix: Accept `verification.type: content` as `file_exists` Alias
400
+
401
+ Both `qwen2.5-coder:7b-instruct` and `specgen:latest` emit `verification.type: "content"`
402
+ with `path` + `expect.content_contains` — structurally identical to `file_exists` but
403
+ not recognized by the scorer's canonical vocab. Fixed `scripts/runbook_scorer.py`:
404
+
405
+ - Added `"content"` to `REQUIRED_EXPECT_KEYS` (same keys as `file_exists`)
406
+ - `_check_inspect()` now treats `content` as inspect (read-only check)
407
+ - `_check_verify()` treats `content` with content_contains as verify evidence
408
+ - Verification testability section handles `content` same as `file_exists`
409
+
410
+ Before: avg score 0.104, 1/9 above 0.75.
411
+ After: avg score 0.972, 9/9 above 0.75.
412
+
413
+ #### Sprint 9 — Prompt Synthesis & Decomposition Pipeline (Planning)
414
+
415
+ Created `docs/SPRINTS.md` Sprint 9 (9a-9d) — four sub-sprints to accept raw unfocused
416
+ user text end-to-end:
417
+
418
+ - **9a** `scripts/prompt_synthesizer.py` — raw text → decomposed clean prompts (JSONL)
419
+ - **9b** `scripts/bulk_generate.py` — batch spec generation from prompt file
420
+ - **9c** `scripts/bulk_plan.py` — loop over specs → plans → consolidated doc
421
+ - **9d** `make text-to-plan TEXT="..."` — single target: synthesize → generate → score → plan
422
+
423
+ Independent of Sprints 5-8. Default models: specgen:latest for generation + planning.
424
+
425
+ #### Sprints 5, 6, 7 Marked Complete
426
+
427
+ - **Sprint 5** (Model provider switch) — already built: run_pipeline.py supports 7 providers
428
+ - **Sprint 6** (Skill bundling/install) — `install-skill` target body added to Makefile
429
+ - **Sprint 7** (Fine-tuning pipeline) — train*.py, merge*.py, eval*.py all exist; Makefile targets wired
430
+
431
+ `scripts/githeri.py` now supports a full end-to-end pipeline from natural language to tested code. **Spec generation is always handled by `specgen:latest` on Ollama** (finetuned for structured YAML output). **All execution and tests run inside a Docker sandbox by default** — the container is throwaway, so the host `.venv`/`.git` are never polluted across multiple test runs.
432
+
433
+ ```bash
434
+ # Minimal invocation — only the prompt and exec model vary between tests.
435
+ # Everything else (docker, workdir, spec model/provider) is defaulted.
436
+ GITHERI_MODEL=qwen2.5-coder:7b-instruct \
437
+ GITHERI_PROVIDER=ollama \
438
+ .venv/bin/python scripts/githeri.py \
439
+ --prompt "Add a PATCH endpoint to update user displayName and bio"
440
+ ```
441
+
442
+ **Defaults (set once, never touched between tests):**
443
+
444
+ - --docker: True — every run_command runs in throwaway githeri-sandbox container
445
+ - --workdir: output/sandbox — all file ops jailed here; created under repo root
446
+ - --spec-model: specgen:latest — **hardcoded** (finetuned for spec generation only)
447
+ - --spec-provider: ollama — **hardcoded**
448
+ - Exec model/provider: env vars GITHERI_MODEL / GITHERI_PROVIDER — only things you change between tests
449
+ - --docker-image: githeri-sandbox
450
+
451
+ **Only two knobs per test:** --prompt (or --spec) and the exec model via env vars. No flags for docker, workdir, or spec model — they're locked in.
452
+
453
+ ```bash
454
+ # Switch exec model — same prompt, same defaults, different model
455
+ GITHERI_MODEL=qwen2.5-coder:7b-instruct \
456
+ GITHERI_PROVIDER=ollama \
457
+ .venv/bin/python scripts/githeri.py --prompt "Add a /health endpoint"
458
+
459
+ # Different provider for execution (specgen stays ollama/specgen)
460
+ GITHERI_MODEL=qwen2.5-14b-instruct-latest:2 \
461
+ GITHERI_PROVIDER=lmstudio \
462
+ .venv/bin/python scripts/githeri.py --prompt "Add a /health endpoint"
463
+
464
+ # Execute an existing spec.yaml instead of generating one
465
+ .venv/bin/python scripts/githeri.py --spec output/my-feature/spec.yaml
466
+ ```
467
+
468
+ **Overrides (rare):**
469
+
470
+ ```bash
471
+ # Disable docker (e.g. when Docker daemon is down) — not recommended for tests
472
+ .venv/bin/python scripts/githeri.py --prompt "..." --no-docker
473
+
474
+ # Custom workdir
475
+ .venv/bin/python scripts/githeri.py --prompt "..." --workdir output/custom-sandbox
476
+
477
+ # Custom docker image
478
+ .venv/bin/python scripts/githeri.py --prompt "..." --docker-image my-sandbox
479
+ ```
480
+
481
+ **Pipeline flow:**
482
+ 1. Specgen LLM (specgen:latest / Ollama) converts NL prompt → spec.yaml (with entrypoint, packages, goals)
483
+ 2. Execution model (set via env vars) reads spec, creates files in output/sandbox, installs deps, runs verifications — all inside Docker
484
+ 3. Server container starts from entrypoint.start_command (if --server)
485
+ 4. Test command runs inside container via docker exec (if --test-cmd)
486
+
487
+ **Environment isolation (critical):**
488
+ - Docker is the default and recommended mode — every `run_command` (including `pip install`) runs in a throwaway container. The host `.venv` is never touched by spec-generated dependencies.
489
+ - `--no-docker` disables this isolation. When used, pip installs and code execution happen in the host environment (the project `.venv`), which can pollute it with spec-generated packages (Flask, FastAPI, etc.). Use `--no-docker` only when you understand this risk.
490
+ - If you must run without Docker, create a separate virtualenv for the workdir:
491
+ ```bash
492
+ python -m venv /tmp/githeri-isolated
493
+ source /tmp/githeri-isolated/bin/activate
494
+ pip install -r requirements.txt # only githeri's own deps
495
+ GITHERI_MODEL=... GITHERI_PROVIDER=... \
496
+ scripts/githeri.py --prompt "..." --no-docker --workdir /tmp/githeri-workdir
497
+ ```
498
+
499
+ **Model separation:** Spec generation uses specgen:latest (trained for structured YAML output, hardcoded). Execution uses whatever you set via GITHERI_MODEL / GITHERI_PROVIDER. Override via env vars only — the --spec-model / --spec-provider flags are suppressed (backwards-compatible but ignored).
500
+
501
+ **Server lifecycle:** After the agent loop completes, githeri parses entrypoint.start_command from the generated spec, starts a named Docker container (docker run -d), and runs --test-cmd inside it. Container stays running for manual inspection: docker exec <name> bash.
502
+
503
+ See [docs/SPECGEN_HARNESS.md](docs/SPECGEN_HARNESS.md) for full architecture, troubleshooting, and output structure.
504
+
505
+ ### 2026-08-15 (This Session)
506
+
507
+ #### GRG Agent Skill — Hermes Native Integration
508
+ The GRG agent skill is now installed as a native Hermes skill (`~/.hermes/skills/autonomous-ai-agents/grg_agent/`) with full multi-provider support:
509
+
510
+ 1. **Skill Installation** — Copied from `skills/grg_agent/` and installed via editable pip install
511
+ 2. **Lightweight GRG Dependency** — Uses local `grg-0.1.0-py3-none-any.whl` wheel (no karakana dependency) providing `AlphaMomentumTracker` and `compute_structural_alpha`
512
+ 3. **Hermes Proxy Support** — Skill accepts `provider` argument: `ollama` | `hermes` | `auto` — uses Hermes's configured providers (NVIDIA Nemotron, Nous Portal, xAI Grok, etc.) instead of local models
513
+ 4. **Direct GRG Agent Execution** — `grg:execute` command now calls `self.agent.solve()` directly instead of legacy `run_pipeline.py` subprocess
514
+ 5. **Make Target Integration** — `scripts/grg_make_spec.py` updated to use the skill with provider argument: `make grg-spec PROMPT="..." PROVIDER=ollama|hermes|auto`
515
+ 6. **Project Virtual Environment** — Runs in project's own `.venv/` (not external karakana venv)
516
+
517
+ **Key Benefit**: You can now use cloud models via Hermes proxy (`hermes proxy start`) instead of relying on locally installed Ollama models. The skill routes through Hermes's provider config which supports NVIDIA Nemotron, Nous Portal, xAI Grok, and any OpenAI-compatible endpoint.
518
+
519
+ #### LM Studio Local Model — First End-to-End Working Pipeline
520
+ **LM Studio** with **`qwen2.5-coder-14b-instruct-uncensored`** is the first local model to complete the full GRG pipeline end-to-end:
521
+
522
+ ```bash
523
+ # Full pipeline from NL prompt → validated spec → plan → execution → runbook
524
+ make grg-full PROMPT="Implement a FastAPI POST /api/health-check endpoint..." \
525
+ PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"
526
+ ```
527
+
528
+ **What works:**
529
+ - **GRG Agent solving** — Generates implementation code with GRG quality gates (composite scoring, diversity control, convergence detection)
530
+ - **Code verification** — Execution-based verification (syntax + runtime) in isolated temp files
531
+ - **Multi-strategy generation** — Standard, decompose, test_first, refine strategies with adaptive temperature
532
+ - **Health-check endpoint example** — Generated FastAPI code with SQLAlchemy DB check + Redis cache check, returns 200/503
533
+ - **All artifacts isolated in `foreign/`** — Clean workspace separation
534
+
535
+ **Configuration for LM Studio:**
536
+ ```bash
537
+ # In LM Studio: enable "OpenAI Compatible Server" on port 1234
538
+ # Load qwen2.5-coder-14b-instruct-uncensored model
539
+ # Then run:
540
+ make grg-spec PROMPT="..." PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"
541
+ ```
542
+
543
+ **Why this model works:**
544
+ - 14B parameter coder model fine-tuned for code generation
545
+ - Uncensored variant removes alignment filters that can interfere with code structure
546
+ - OpenAI-compatible API in LM Studio works with GRG's `OllamaClient` (custom `api_key` support)
547
+ - Sufficient context window for spec + plan generation tasks
548
+ - Produces valid imports, proper error handling, and correct HTTP status codes
549
+
550
+ ### 2026-07-30 (Latest)
551
+
552
+ #### THIRD_IMPROVE_SPEC — 7 Pipeline Fixes
553
+ After analyzing a 10-spec batch run, identified 7 recurring failure patterns and fixed all of them:
554
+
555
+ 1. **Minimal structural skeleton in SYSTEM_PROMPT** — eliminated top-level field confusion
556
+ 2. **task_id validator check** — rejects L11/G18-style IDs, requires descriptive slug like `jwt-auth-login`
557
+ 3. **Verification types & expect keys table** — explicit vocabulary reduces hallucinated keys
558
+ 4. **Placeholder ban strengthened** — concrete examples in prompt (`JWT_SECRET: "test-secret-..."`, `DATABASE_URL: "postgresql://user:***@..."`) + validator hint
559
+ 5. **Helpful error hints** — "These fields belong inside a goal under `local_goals`, not at the spec root"
560
+ 6. **depends_on validator hints** — rejects L/G refs, redirects to task_ids/stage names
561
+ 7. **Structural acceptance_criteria template** — local goal field, not top-level
562
+
563
+ All 90 tests passing.
564
+
565
+ #### Generation Mode Aliases
566
+ Added `make generate N=random` and `make generate N=all` flags plus `generate-random` / `generate-all` aliases for short tests and full corpus runs. 475 seed prompts across 21 categories available.
567
+
568
+ #### Model Default Reverted
569
+ `OLLAMA_MODEL` reverted to `qwen2.5-coder:7b-instruct` (was regressed to `qwen3.5-4b-128k:latest`). Better structure compliance after prompt fixes.
570
+
571
+ ### 2026-07-30 (Earlier)
572
+
573
+ #### Spec Enrichment (IMPROVE_SPEC)
574
+ Added five top-level enrichment fields (`business_rules`, `test_fixtures`, `environment`, `global_verification`) and three goal-level fields (`blueprint`, `acceptance_criteria`, `type`). All validated and scored.
575
+
576
+ #### Multi-Provider LLM Support
577
+ `run_pipeline.py` now supports Ollama, OpenAI, Anthropic, NVIDIA (via Together AI), and any OpenAI-compatible endpoint. Configured via `--provider` CLI arg or Makefile variables.
578
+
579
+ #### Runbook Scorer Stage Detection
580
+ Scorer now honors explicit `type: create|inspect|verify` on goals, fixing false "missing stage" penalties for `file_exists` verification on CREATE goals.
581
+
582
+ #### Validator Hardening
583
+ - Guarded against `expect` being a string instead of dict (prevents `AttributeError: 'str' object has no attribute 'get'`)
584
+ - Near-duplicate detection now handles malformed `expect` blocks defensively
585
+ - All 90 tests pass
586
+
587
+ ## Model Compatibility Matrix
588
+
589
+ | Model | Context | Tools | Speed (M1 16GB) | Spec Gen | Recommended Use |
590
+ |-------|---------|-------|-----------------|----------|-----------------|
591
+ | **specgen:latest** | 128K | Yes | Fast | **Best first-attempt pass rate (fine-tuned)** | **Default local (Ollama) — spec generation** |
592
+ || qwen2.5-coder:7b-instruct | 32K | Yes | Fast | Works (best structure compliance) | Fallback / base model for training |
593
+ || qwen2.5-coder-14b-instruct-uncensored | 32K | Yes | Medium | **Works end-to-end (GRG pipeline)** | **LM Studio local** |
594
+ || qwen3.5-4b-128k | 128K | Yes | Fast | Works | Larger context fallback |
595
+ || qwen3.5-9b-code:128k | 128K | Yes | Slow | Excellent | Higher quality if time allows |
596
+ || deepseek-r1:7b | 128K | TBD | Fast | YAML syntax errors | Not recommended for spec gen |
597
+ || **Nemotron 3 Ultra** | 128K | Yes | Fast | Excellent | **Cloud via NVIDIA NIM (`--provider nvidia`)** |
598
+
599
+ **Current recommendation**:
600
+ - **Spec generation (default)** → `specgen:latest` on Ollama (fine-tuned for structured YAML output, best first-attempt pass rate)
601
+ - **Bulk corpus** → `specgen:latest` on Ollama (fine-tuned, best first-attempt pass rate)
602
+ - **Specific features (GRG pipeline)** → `qwen2.5-coder-14b-instruct-uncensored` on LM Studio with `make grg-full` (100% success via execution verification)
603
+ - **Cloud** → `minimaxai/minimax-m3` via NVIDIA NIM (`--provider nvidia`)
604
+
605
+ ## Training Data Generation Strategies
606
+
607
+ Based on empirical testing (2026-08-16), here are the reliable approaches:
608
+
609
+ ### 1. Bulk Corpus Generation (Recommended for Training Data)
610
+ ```bash
611
+ # Fast, decent success rate, fine-tuned for this pipeline
612
+ make generate N=100 PROVIDER=ollama MODEL="specgen:latest" BASE_URL="http://localhost:11434" TEMPERATURE=0.2 MAX_TOKENS=4096
613
+ ```
614
+ - **Success rate**: ~70%+ (fine-tuned specgen:latest)
615
+ - **Speed**: ~50-100s/spec
616
+ - **Best for**: Generating large training corpora quickly
617
+
618
+ ### 2. High-Quality Individual Specs (GRG Pipeline)
619
+ ```bash
620
+ # Best for specific features - multi-strategy + execution verification
621
+ make grg-full PROMPT="Your feature" PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"
622
+ ```
623
+ - **Success rate**: ~100% (execution-verified)
624
+ - **Speed**: ~260s/spec
625
+ - **Best for**: Critical features requiring guaranteed working specs
626
+
627
+ ### 3. GRG Agent Code Training (NEW — Real GRG Scores + Verified Code)
628
+ ```bash
629
+ # Generates verified code with real GRG composite scores (not dummy 0.5)
630
+ # Uses qwen2.5-coder:7b-instruct on Ollama for real logprobs
631
+ make generate-code # cycles through all 475 SEED_PROMPTS
632
+ make convert-code-chat # converts passing examples to chat format
633
+ ```
634
+ - **Success rate**: Variable (depends on prompt complexity)
635
+ - **Speed**: ~30-60s/prompt (optimized: 1 strategy, 1 candidate)
636
+ - **Output**: `data/training_data_code.jsonl` + `data/training_data_code_chat.jsonl`
637
+ - **Key difference**: Produces **executable code** with **real GRG composite scores** (0.47-0.49) because the model provides logprobs
638
+ - **Verification**: Syntax check + execution test (import + basic run)
639
+ - **Best for**: Fine-tuning code generation models with GRG quality signals
640
+
641
+ **Configuration** (in `scripts/generate_code_training_fast.py`):
642
+ ```python
643
+ skill = create_skill(config={
644
+ 'llm_provider': 'ollama',
645
+ 'ollama_base_url': 'http://127.0.0.1:11434/v1',
646
+ 'ollama_default_model': 'qwen2.5-coder:7b-instruct',
647
+ 'max_iterations': 2, # Must be >=2 for convergence check
648
+ 'temperature': 0.3,
649
+ 'top_p': 0.9,
650
+ 'max_tokens': 1024,
651
+ 'candidates_per_strategy': 1, # Speed optimization
652
+ 'max_strategies': 1, # Speed optimization
653
+ })
654
+ ```
655
+
656
+ ### 4. Hybrid Workflow (Best of Both)
657
+ ```bash
658
+ # 1. Generate bulk corpus with specgen
659
+ make generate N=100 PROVIDER=ollama MODEL="specgen:latest" BASE_URL="http://localhost:11434"
660
+
661
+ # 2. Score and filter high-quality specs
662
+ make score MIN_SCORE=0.75
663
+
664
+ # 3. Re-generate failed critical specs with GRG pipeline
665
+ make grg-full PROMPT="..." PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"
666
+
667
+ # 4. Generate verified code for fine-tuning
668
+ make generate-code
669
+ make convert-code-chat
670
+ ```
671
+
672
+ ### Observed Failure Patterns (specgen:latest)
673
+ The remaining failures (if any) are primarily:
674
+ 1. **Near-duplicate HTTP verifications** - multiple goals hitting same endpoint with same method
675
+ 2. **Missing `acceptance_criteria`** for CREATE goals
676
+ 3. **Placeholder values** in headers (e.g., `Authorization: *** ***`)
677
+ 4. **YAML block mapping errors** - CLI verification indentation issues
678
+
679
+ ## For More Information
680
+
681
+ - [docs/SPRINTS_AUTONOMOUS.md](docs/SPRINTS_AUTONOMOUS.md) — sprint breakdown, model experiments, decisions, next steps
682
+ - [docs/scoring_spec.md](docs/scoring_spec.md) — runbook scoring specification
683
+ - [docs/IMPROVE_SPEC.md](docs/IMPROVE_SPEC.md) — spec enrichment field specification (v1)
684
+ - [docs/SECOND_IMPROVE_SPEC.md](docs/SECOND_IMPROVE_SPEC.md) — enrichment field enforcement (v2)
685
+ - [docs/THIRD_IMPROVE_SPEC.txt](docs/THIRD_IMPROVE_SPEC.txt) — 7 pipeline failure patterns + fixes (v3)
686
+ - [MODEL_CARD.md](MODEL_CARD.md) — model card (uploaded to HF Hub)
687
+ - [skills/spec-forge/SKILL.md](skills/spec-forge/SKILL.md) — Spec-Forge skill reference
688
+ - [skills/command-runway-pattern/SKILL.md](skills/command-runway-pattern/SKILL.md) — Command Runway pattern skill
689
+ - [skills/grg_agent/SKILL.md](skills/grg_agent/SKILL.md) — GRG Agent skill reference