@pmelab/gtd 18.1.0 → 19.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "gtd",
3
- "version": "18.1.0",
3
+ "version": "19.0.0",
4
4
  "description": "Drive the gtd loop inside Claude Code: agent turns run as subagents, gtd decides every beat",
5
5
  "author": {
6
6
  "name": "Philipp Melab"
package/README.md CHANGED
@@ -40,6 +40,10 @@ the read side to its agent. `read` is not a security boundary unless your driver
40
40
  enforces it at the OS level. See
41
41
  [Configuration](https://github.com/pmelab/gtd/blob/main/docs/configuration.md#file-access).
42
42
 
43
+ > **`fastTestCommand` is required, with no fallback to `testCommand`.** Set it
44
+ > (everything but e2e) under `env:` in `.gtdrc` or as `GTD_FASTTESTCOMMAND`; gtd
45
+ > stops at the start gate, writing `.gtd/SETUP.md`, until it is.
46
+
43
47
  > **A repository's `gtd.config.ts` is code, and gtd runs it.** A custom workflow
44
48
  > is a TypeScript module, and every gtd command that looks at workflow state —
45
49
  > `gtd next` and `gtd lsp` included, not just `gtd land` — evaluates it. Treat
@@ -355,7 +359,7 @@ lists the open ones (editor-only — the phone UI does not show them).
355
359
 
356
360
  One built-in workflow drives all of that. From where you sit, it has four
357
361
  moments — everything between them runs without you, with the judged exceptions
358
- noted in steps 2, 3, and 4 below.
362
+ noted in steps 3 and 4 below.
359
363
 
360
364
  1. **You sketch.** Change anything, or write the idea into `.gtd/TODO.md`. Rough
361
365
  is fine; it is treated as a sketch, not as work.
@@ -368,25 +372,22 @@ noted in steps 2, 3, and 4 below.
368
372
  the same gate again. Close a thread by replying with a conclusion or deleting
369
373
  it. While a thread's last entry is the agent's, moving on is refused. Leave
370
374
  the file untouched and start the loop to accept the plan as-is, unanswered
371
- questions and all. One point along this phase is judged rather than always
372
- asking you outright:
373
- - Before the how-it-should-be-built pass starts: does this plan actually need
374
- one? A confident no skips it — and the review it would have raised — going
375
- straight from your answers to a single built package, with no technical
376
- plan shown to you at all.
377
-
378
- The reference driver answers this judgment itself (`gtd judge run`, auto
379
- selection); if that fails it shows you the message and stops, same as any
380
- other question. **The `llm` provider's `p` is self-reported by the model, not
381
- a measured probability, so a confidently wrong haiku verdict can skip a
382
- question you would have asked.**
375
+ questions and all. Every plan gets the technical pass: its document has four
376
+ sections, in order — `## Interfaces`, `## Call Stacks`, `## E2E Scenarios`,
377
+ `## Unit Tests` — behind a leading `## Open Questions` when there are any.
378
+ `## E2E Scenarios` is never empty: it holds the scenarios, or, when nothing
379
+ user-visible changes, the line `No e2e change.` with a one-line reason.
383
380
 
384
381
  3. **You wait.** The work is split into packages and built one at a time, each
385
- one checked against your test suite and fixed until it passes, then reviewed
386
- against its own spec before moving on. Three points along that loop are
387
- judged rather than always asking you outright — each stops and hands you a
388
- verdict to make (`gtd judge answer`, or land with a clean tree to accept the
389
- conservative default, which never skips work;
382
+ one starting from the unit tests it declares (a build turn missing one is
383
+ refused), checked against the fast suite and fixed until it passes before
384
+ moving on. A package that rewords a frozen `.feature` step stops at a wording
385
+ gate: accept the change, or reject it and the original is restored. After the
386
+ last package a full run (e2e included) has its own fix loop before the
387
+ quality lap. Two points along the process are judged rather than always
388
+ asking you outright — each stops and hands you a verdict to make
389
+ (`gtd judge answer`, or land with a clean tree to accept the conservative
390
+ default, which never skips work;
390
391
  `gtd judge run --provider fixed --answers <path>` — or the
391
392
  `GTD_JUDGE_ANSWERS` env var, inline JSON — answers one from a file, piped
392
393
  between `gtd judge --json` and `gtd judge answer`;
@@ -403,17 +404,15 @@ noted in steps 2, 3, and 4 below.
403
404
  stdout when they cannot answer every question):
404
405
  - Every red round after the first: was the failure identical, new, or
405
406
  progress?
406
- - Before spending a review turn on a package: does the code already satisfy
407
- each of its requirements?
408
- - After a review turn raises concerns: would each one actually violate the
409
- spec if left unaddressed, or is it a nit?
407
+ - After you review: is each of your notes an edit, a question, a nit or
408
+ praise?
410
409
 
411
410
  The reference driver answers these itself (`gtd judge run`, auto selection:
412
411
  jev when `TYPESAFE_API_KEY` is set, else `llm` via `claude`, default model
413
412
  haiku, `--model <name>` overrides); if that fails it shows you the message
414
413
  and stops. **The `llm` provider's `p` is self-reported by the model, not a
415
- measured probability, so a confidently wrong verdict can clear the 0.9/0.7
416
- floors and skip a gate unattended.**
414
+ measured probability, so a confidently wrong verdict can clear the 0.7 floor
415
+ and skip a gate unattended.**
417
416
 
418
417
  Once the last package is built, the whole change goes through a qualitative
419
418
  review lap before you see anything: six lenses, one turn each, in order —
@@ -421,8 +420,7 @@ noted in steps 2, 3, and 4 below.
421
420
  `conventions`, `spec-challenge`. Each traces the change from its own angle
422
421
  and records every finding, blocking or not; one fix turn then fixes ALL of
423
422
  them once, with no re-review after the fix. A clean turn means approval only
424
- when that lens found nothing at all. The per-package review above only judges
425
- that package against its own spec; this lap is where code quality is looked
423
+ when that lens found nothing at all. This lap is where code quality is looked
426
424
  at, and every round pays for it. It never replaces step 4 — your review stays
427
425
  the final gate, and nothing here skips it. The `gtd --entry fix-precheck`
428
426
  side door (below) repairs a red baseline through this same lap.
@@ -438,10 +436,10 @@ noted in steps 2, 3, and 4 below.
438
436
 
439
437
  4. **You review.** You get a review document listing what changed and what to
440
438
  look at. Before you see it, an automatic risk-fix pass
441
- (`build.review.fix-risks`) runs. Any risk the reviewer names (a note opening
442
- with `Risk:`) is fixed first, the suite kept green, and the review rewritten
443
- — once per review round, so a risk the rewrite still names reaches you
444
- unfixed; risk: a fix lands with no check that the risk was real. Tick the
439
+ (`build.review.fix.risks.fixing`) runs. Any risk the reviewer names (a note
440
+ opening with `Risk:`) is fixed first, the suite kept green, and the review
441
+ rewritten — once per review round, so a risk the rewrite still names reaches
442
+ you unfixed; risk: a fix lands with no check that the risk was real. Tick the
445
443
  boxes to approve, or write what is wrong. Approving ends the process;
446
444
  feedback is judged note by note, each as `edit`, `question`, `nit` or
447
445
  `praise`. An `edit` sends the process back to step 2 for a fresh plan — it