@pmelab/gtd 15.9.0 → 15.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -335,9 +335,11 @@ noted in steps 2, 3, and 4 below.
335
335
  straight from your answers to a single built package, with no technical
336
336
  plan shown to you at all.
337
337
 
338
- A driver built only to run this loop (not to answer judgments) still handles
339
- every one of these correctly: it shows you the message and stops, same as any
340
- other question.
338
+ The reference driver answers this judgment itself (`gtd judge run`, auto
339
+ selection); if that fails it shows you the message and stops, same as any
340
+ other question. **The `llm` provider's `p` is self-reported by the model, not
341
+ a measured probability, so a confidently wrong haiku verdict can skip a
342
+ question you would have asked.**
341
343
 
342
344
  3. **You wait.** The work is split into packages and built one at a time, each
343
345
  one checked against your test suite and fixed until it passes, then reviewed
@@ -349,8 +351,13 @@ noted in steps 2, 3, and 4 below.
349
351
  `GTD_JUDGE_ANSWERS` env var, inline JSON — answers one from a file, piped
350
352
  between `gtd judge --json` and `gtd judge answer`;
351
353
  `gtd judge run --provider jev` asks TypeSafe's Jev instead, with the key in
352
- `TYPESAFE_API_KEY` and `JEV_BASE_URL` optionally overriding the endpoint, and
353
- exits 1 with nothing on stdout when it cannot answer every question):
354
+ `TYPESAFE_API_KEY` and `JEV_BASE_URL` optionally overriding the endpoint;
355
+ `gtd judge run --provider llm` asks the `claude` CLI on `PATH` (Claude Code
356
+ users only; model `haiku`, `--model <name>` overrides) and reuses its login.
357
+ With no `--provider`, `gtd judge run` picks jev when `TYPESAFE_API_KEY` is
358
+ set and non-empty and no `--model` is given, else llm — never falling back
359
+ across them; `--model` with `--provider fixed`/`jev` is refused. jev and llm
360
+ exit 1 with nothing on stdout when they cannot answer every question):
354
361
  - Every red round after the first: was the failure identical, new, or
355
362
  progress?
356
363
  - Before spending a review turn on a package: does the code already satisfy
@@ -358,9 +365,12 @@ noted in steps 2, 3, and 4 below.
358
365
  - After a review turn raises concerns: would each one actually violate the
359
366
  spec if left unaddressed, or is it a nit?
360
367
 
361
- A driver built only to run this loop (not to answer judgments) still handles
362
- every one of these correctly: it shows you the message and stops, same as any
363
- other question.
368
+ The reference driver answers these itself (`gtd judge run`, auto selection:
369
+ jev when `TYPESAFE_API_KEY` is set, else `llm` via `claude`, default model
370
+ haiku, `--model <name>` overrides); if that fails it shows you the message
371
+ and stops. **The `llm` provider's `p` is self-reported by the model, not a
372
+ measured probability, so a confidently wrong verdict can clear the 0.9/0.7
373
+ floors and skip a gate unattended.**
364
374
 
365
375
  Once the last package is built, the whole change goes through a qualitative
366
376
  review lap before you see anything: one configured skill per turn, each