@pmelab/gtd 15.8.0 → 15.10.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +31 -14
- package/dist/gtd.bundle.mjs +613 -274
- package/package.json +1 -1
- package/src/workflows/prose.ts +5 -4
- package/src/workflows/review.test.ts +247 -0
- package/src/workflows/review.ts +125 -47
- package/src/workflows/steps.test.ts +39 -0
- package/src/workflows/steps.ts +38 -6
- package/src/workflows/text.fixture.ts +40 -42
- package/src/workflows/text.ts +91 -21
package/README.md
CHANGED
|
@@ -335,9 +335,11 @@ noted in steps 2, 3, and 4 below.
|
|
|
335
335
|
straight from your answers to a single built package, with no technical
|
|
336
336
|
plan shown to you at all.
|
|
337
337
|
|
|
338
|
-
|
|
339
|
-
|
|
340
|
-
other question.
|
|
338
|
+
The reference driver answers this judgment itself (`gtd judge run`, auto
|
|
339
|
+
selection); if that fails it shows you the message and stops, same as any
|
|
340
|
+
other question. **The `llm` provider's `p` is self-reported by the model, not
|
|
341
|
+
a measured probability, so a confidently wrong haiku verdict can skip a
|
|
342
|
+
question you would have asked.**
|
|
341
343
|
|
|
342
344
|
3. **You wait.** The work is split into packages and built one at a time, each
|
|
343
345
|
one checked against your test suite and fixed until it passes, then reviewed
|
|
@@ -349,8 +351,13 @@ noted in steps 2, 3, and 4 below.
|
|
|
349
351
|
`GTD_JUDGE_ANSWERS` env var, inline JSON — answers one from a file, piped
|
|
350
352
|
between `gtd judge --json` and `gtd judge answer`;
|
|
351
353
|
`gtd judge run --provider jev` asks TypeSafe's Jev instead, with the key in
|
|
352
|
-
`TYPESAFE_API_KEY` and `JEV_BASE_URL` optionally overriding the endpoint
|
|
353
|
-
|
|
354
|
+
`TYPESAFE_API_KEY` and `JEV_BASE_URL` optionally overriding the endpoint;
|
|
355
|
+
`gtd judge run --provider llm` asks the `claude` CLI on `PATH` (Claude Code
|
|
356
|
+
users only; model `haiku`, `--model <name>` overrides) and reuses its login.
|
|
357
|
+
With no `--provider`, `gtd judge run` picks jev when `TYPESAFE_API_KEY` is
|
|
358
|
+
set and non-empty and no `--model` is given, else llm — never falling back
|
|
359
|
+
across them; `--model` with `--provider fixed`/`jev` is refused. jev and llm
|
|
360
|
+
exit 1 with nothing on stdout when they cannot answer every question):
|
|
354
361
|
- Every red round after the first: was the failure identical, new, or
|
|
355
362
|
progress?
|
|
356
363
|
- Before spending a review turn on a package: does the code already satisfy
|
|
@@ -358,9 +365,12 @@ noted in steps 2, 3, and 4 below.
|
|
|
358
365
|
- After a review turn raises concerns: would each one actually violate the
|
|
359
366
|
spec if left unaddressed, or is it a nit?
|
|
360
367
|
|
|
361
|
-
|
|
362
|
-
|
|
363
|
-
|
|
368
|
+
The reference driver answers these itself (`gtd judge run`, auto selection:
|
|
369
|
+
jev when `TYPESAFE_API_KEY` is set, else `llm` via `claude`, default model
|
|
370
|
+
haiku, `--model <name>` overrides); if that fails it shows you the message
|
|
371
|
+
and stops. **The `llm` provider's `p` is self-reported by the model, not a
|
|
372
|
+
measured probability, so a confidently wrong verdict can clear the 0.9/0.7
|
|
373
|
+
floors and skip a gate unattended.**
|
|
364
374
|
|
|
365
375
|
Once the last package is built, the whole change goes through a qualitative
|
|
366
376
|
review lap before you see anything: one configured skill per turn, each
|
|
@@ -383,12 +393,19 @@ noted in steps 2, 3, and 4 below.
|
|
|
383
393
|
|
|
384
394
|
4. **You review.** You get a review document listing what changed and what to
|
|
385
395
|
look at. Tick the boxes to approve, or write what is wrong. Approving ends
|
|
386
|
-
the process; feedback
|
|
387
|
-
|
|
388
|
-
|
|
389
|
-
|
|
390
|
-
|
|
391
|
-
|
|
396
|
+
the process; feedback is judged note by note, each as `edit`, `question`,
|
|
397
|
+
`nit` or `praise`. An `edit` sends the process back to step 2 for a fresh
|
|
398
|
+
plan — it never patches over a design you rejected. A `question` is answered
|
|
399
|
+
inline under your note and the process stops at the review again — no new
|
|
400
|
+
plan. A `nit` is fixed in one batch, the suite must go green (a red one gets
|
|
401
|
+
fix turns first), then a fresh review of the change stops at the review
|
|
402
|
+
again. `praise` is dropped; a round of only praise signs off. When a round
|
|
403
|
+
mixes them, questions are answered and nits fixed first, then the edits are
|
|
404
|
+
planned — risk: the planning lap may redo nit fixes it touches. A hand-edit
|
|
405
|
+
to code always plans a lap. Only a confident non-`edit` verdict skips the
|
|
406
|
+
replan — a note whose evidence was cut, or that got no verdict, counts as
|
|
407
|
+
`edit`. The same `gtd judge answer` / conservative-default shape as step 3's
|
|
408
|
+
own judged points.
|
|
392
409
|
|
|
393
410
|
You never talk to it. Every exchange is a file in `.gtd/` that you edit in your
|
|
394
411
|
own editor, and every answer you give is a commit. Your test suite is the gate
|