@pmelab/gtd 15.9.0 → 16.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -26,7 +26,10 @@ for the settings most projects tune.
26
26
  Also install all three skill sources the bundled workflow names in its prompts —
27
27
  none is optional. See
28
28
  [Setup](https://github.com/pmelab/gtd/blob/main/docs/setup.md) for the list, the
29
- install flags, and the never-install-the-plugin caveat.
29
+ install flags, and the never-install-the-plugin caveat. Each bundled agent step
30
+ has its own addressable `.gtdrc` `skills:` entry — see
31
+ [Configuration](https://github.com/pmelab/gtd/blob/main/docs/configuration.md#the-skills-key)
32
+ to repoint one to a set your own harness has instead.
30
33
 
31
34
  > **A repository's `gtd.config.ts` is code, and gtd runs it.** A custom workflow
32
35
  > is a TypeScript module, and every gtd command that looks at workflow state —
@@ -335,9 +338,11 @@ noted in steps 2, 3, and 4 below.
335
338
  straight from your answers to a single built package, with no technical
336
339
  plan shown to you at all.
337
340
 
338
- A driver built only to run this loop (not to answer judgments) still handles
339
- every one of these correctly: it shows you the message and stops, same as any
340
- other question.
341
+ The reference driver answers this judgment itself (`gtd judge run`, auto
342
+ selection); if that fails it shows you the message and stops, same as any
343
+ other question. **The `llm` provider's `p` is self-reported by the model, not
344
+ a measured probability, so a confidently wrong haiku verdict can skip a
345
+ question you would have asked.**
341
346
 
342
347
  3. **You wait.** The work is split into packages and built one at a time, each
343
348
  one checked against your test suite and fixed until it passes, then reviewed
@@ -349,8 +354,13 @@ noted in steps 2, 3, and 4 below.
349
354
  `GTD_JUDGE_ANSWERS` env var, inline JSON — answers one from a file, piped
350
355
  between `gtd judge --json` and `gtd judge answer`;
351
356
  `gtd judge run --provider jev` asks TypeSafe's Jev instead, with the key in
352
- `TYPESAFE_API_KEY` and `JEV_BASE_URL` optionally overriding the endpoint, and
353
- exits 1 with nothing on stdout when it cannot answer every question):
357
+ `TYPESAFE_API_KEY` and `JEV_BASE_URL` optionally overriding the endpoint;
358
+ `gtd judge run --provider llm` asks the `claude` CLI on `PATH` (Claude Code
359
+ users only; model `haiku`, `--model <name>` overrides) and reuses its login.
360
+ With no `--provider`, `gtd judge run` picks jev when `TYPESAFE_API_KEY` is
361
+ set and non-empty and no `--model` is given, else llm — never falling back
362
+ across them; `--model` with `--provider fixed`/`jev` is refused. jev and llm
363
+ exit 1 with nothing on stdout when they cannot answer every question):
354
364
  - Every red round after the first: was the failure identical, new, or
355
365
  progress?
356
366
  - Before spending a review turn on a package: does the code already satisfy
@@ -358,9 +368,12 @@ noted in steps 2, 3, and 4 below.
358
368
  - After a review turn raises concerns: would each one actually violate the
359
369
  spec if left unaddressed, or is it a nit?
360
370
 
361
- A driver built only to run this loop (not to answer judgments) still handles
362
- every one of these correctly: it shows you the message and stops, same as any
363
- other question.
371
+ The reference driver answers these itself (`gtd judge run`, auto selection:
372
+ jev when `TYPESAFE_API_KEY` is set, else `llm` via `claude`, default model
373
+ haiku, `--model <name>` overrides); if that fails it shows you the message
374
+ and stops. **The `llm` provider's `p` is self-reported by the model, not a
375
+ measured probability, so a confidently wrong verdict can clear the 0.9/0.7
376
+ floors and skip a gate unattended.**
364
377
 
365
378
  Once the last package is built, the whole change goes through a qualitative
366
379
  review lap before you see anything: one configured skill per turn, each