ll-skills 2.0.2 → 3.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (115) hide show
  1. package/CHANGELOG.md +21 -0
  2. package/README.md +42 -20
  3. package/agents/ll-executor.md +1 -0
  4. package/assets/preamble.md +29 -35
  5. package/bin/install.js +4 -1
  6. package/hooks/ll-precompact.js +29 -1
  7. package/hooks/ll-skills-check-update.js +6 -6
  8. package/hooks/ll-state.js +30 -2
  9. package/package.json +3 -2
  10. package/scripts/evals/README.md +57 -0
  11. package/scripts/evals/cases/auto-dry-run/assert.sh +35 -0
  12. package/scripts/evals/cases/auto-dry-run/case.json +8 -0
  13. package/scripts/evals/cases/auto-dry-run/prompt.txt +1 -0
  14. package/scripts/evals/cases/auto-empty-repo/assert.sh +25 -0
  15. package/scripts/evals/cases/auto-empty-repo/case.json +8 -0
  16. package/scripts/evals/cases/auto-empty-repo/fixture/.gitkeep +0 -0
  17. package/scripts/evals/cases/auto-empty-repo/prompt.txt +1 -0
  18. package/scripts/evals/cases/decide-final-round/assert.sh +32 -0
  19. package/scripts/evals/cases/decide-final-round/case.json +8 -0
  20. package/scripts/evals/cases/decide-final-round/fixture/README.md +3 -0
  21. package/scripts/evals/cases/decide-final-round/prompt.txt +1 -0
  22. package/scripts/evals/cases/executor-block/assert.sh +33 -0
  23. package/scripts/evals/cases/executor-block/case.json +8 -0
  24. package/scripts/evals/cases/executor-block/prompt.txt +14 -0
  25. package/scripts/evals/cases/goal-autonomous/assert.sh +35 -0
  26. package/scripts/evals/cases/goal-autonomous/case.json +8 -0
  27. package/scripts/evals/cases/goal-autonomous/fixture/PLAN.md +42 -0
  28. package/scripts/evals/cases/goal-autonomous/fixture/PROGRESS.md +20 -0
  29. package/scripts/evals/cases/goal-autonomous/fixture/ROADMAP.md +29 -0
  30. package/scripts/evals/cases/goal-autonomous/fixture/package.json +8 -0
  31. package/scripts/evals/cases/goal-autonomous/fixture/src/money.js +6 -0
  32. package/scripts/evals/cases/goal-autonomous/fixture/test/reconcile.test.js +8 -0
  33. package/scripts/evals/cases/goal-autonomous/prompt.txt +1 -0
  34. package/scripts/evals/cases/implement-review-gate/assert.sh +35 -0
  35. package/scripts/evals/cases/implement-review-gate/case.json +8 -0
  36. package/scripts/evals/cases/implement-review-gate/prompt.txt +1 -0
  37. package/scripts/evals/cases/implement-stops-at-next/assert.sh +39 -0
  38. package/scripts/evals/cases/implement-stops-at-next/case.json +9 -0
  39. package/scripts/evals/cases/implement-stops-at-next/prompt.txt +1 -0
  40. package/scripts/evals/cases/preamble-no-ritual/assert.sh +17 -0
  41. package/scripts/evals/cases/preamble-no-ritual/case.json +8 -0
  42. package/scripts/evals/cases/preamble-no-ritual/fixture/README.md +3 -0
  43. package/scripts/evals/cases/preamble-no-ritual/fixture/src/a.ts +3 -0
  44. package/scripts/evals/cases/preamble-no-ritual/prompt.txt +1 -0
  45. package/scripts/evals/cases/router-execute/assert.sh +12 -0
  46. package/scripts/evals/cases/router-execute/case.json +8 -0
  47. package/scripts/evals/cases/router-execute/prompt.txt +1 -0
  48. package/scripts/evals/cases/router-research/assert.sh +11 -0
  49. package/scripts/evals/cases/router-research/case.json +8 -0
  50. package/scripts/evals/cases/router-research/fixture/README.md +3 -0
  51. package/scripts/evals/cases/router-research/prompt.txt +1 -0
  52. package/scripts/evals/cases/router-small/assert.sh +21 -0
  53. package/scripts/evals/cases/router-small/case.json +8 -0
  54. package/scripts/evals/cases/router-small/fixture/README.md +17 -0
  55. package/scripts/evals/cases/router-small/prompt.txt +1 -0
  56. package/scripts/evals/cases/scout-no-plan/assert.sh +41 -0
  57. package/scripts/evals/cases/scout-no-plan/case.json +8 -0
  58. package/scripts/evals/cases/scout-no-plan/prompt.txt +8 -0
  59. package/scripts/evals/cases/verifier-weakened-test/assert.sh +19 -0
  60. package/scripts/evals/cases/verifier-weakened-test/case.json +8 -0
  61. package/scripts/evals/cases/verifier-weakened-test/prompt.txt +13 -0
  62. package/scripts/evals/cases/verifier-weakened-test/setup.sh +19 -0
  63. package/scripts/evals/fixtures/manual-contract/out.json +29 -0
  64. package/scripts/evals/fixtures/manual-contract/out.txt +5 -0
  65. package/scripts/evals/fixtures/manual-contract/with-skill.json +46 -0
  66. package/scripts/evals/lib/assert.sh +107 -0
  67. package/scripts/evals/lib/extract.js +73 -0
  68. package/scripts/evals/run.sh +369 -0
  69. package/scripts/fixtures/auto-closed/PLAN.md +5 -0
  70. package/scripts/fixtures/auto-closed/PROGRESS.md +20 -0
  71. package/scripts/fixtures/auto-closed/ROADMAP.md +6 -0
  72. package/scripts/fixtures/auto-closed/docs/DELIVERY.md +3 -0
  73. package/scripts/fixtures/auto-decisions/decisions/DEC-0001-taken-alone.md +13 -0
  74. package/scripts/fixtures/auto-decisions/decisions/DEC-0002-owner.md +13 -0
  75. package/scripts/fixtures/auto-noroadmap/PLAN.md +20 -0
  76. package/scripts/fixtures/auto-noroadmap/PROGRESS.md +11 -0
  77. package/scripts/fixtures/auto-verify-next/PLAN.md +5 -0
  78. package/scripts/fixtures/auto-verify-next/PROGRESS.md +18 -0
  79. package/scripts/fixtures/auto-verify-next/ROADMAP.md +5 -0
  80. package/scripts/fixtures/auto-verify-next/phases/01/PLAN.md +6 -0
  81. package/scripts/fixtures/evals-auto/auto-dry-run/pass.txt +18 -0
  82. package/scripts/fixtures/evals-auto/auto-empty-repo/pass.txt +2 -0
  83. package/scripts/fixtures/evals-auto/goal-autonomous/pass.txt +29 -0
  84. package/scripts/fixtures/lint-bad/folded-description/SKILL.md +13 -0
  85. package/scripts/fixtures/lint-bad/model-invocation-false/SKILL.md +10 -0
  86. package/scripts/fixtures/next-bad/skills/ll-bad/SKILL.md +30 -0
  87. package/scripts/fixtures/next-good/skills/ll-good/SKILL.md +26 -0
  88. package/scripts/fixtures/project/PROGRESS.md +4 -0
  89. package/scripts/lint-contract.cjs +495 -0
  90. package/scripts/lint-prompts.sh +396 -0
  91. package/scripts/ll-tools.js +465 -447
  92. package/scripts/smoke-test.sh +380 -1
  93. package/skills/ll-auto/SKILL.md +74 -0
  94. package/skills/ll-auto/references/run.md +75 -0
  95. package/skills/ll-auto/references/stages.md +66 -0
  96. package/skills/ll-auto/scripts/ll-auto.js +345 -0
  97. package/skills/ll-brainstorm/SKILL.md +5 -4
  98. package/skills/ll-brainstorm/references/decision-policy.md +3 -0
  99. package/skills/ll-close/SKILL.md +5 -5
  100. package/skills/ll-close/references/delivery.md +3 -1
  101. package/skills/ll-decide/SKILL.md +13 -11
  102. package/skills/ll-decide/references/decision-policy.md +3 -0
  103. package/skills/ll-decide/references/interview.md +10 -0
  104. package/skills/ll-decide/references/plan-skeleton.md +14 -14
  105. package/skills/ll-decide/references/premise-gate.md +7 -0
  106. package/skills/ll-goal/SKILL.md +22 -4
  107. package/skills/ll-goal/references/goal-template.md +57 -0
  108. package/skills/ll-implement/SKILL.md +4 -2
  109. package/skills/ll-implement/references/decision-policy.md +3 -0
  110. package/skills/ll-oncall/SKILL.md +3 -2
  111. package/skills/ll-refine/SKILL.md +3 -2
  112. package/skills/ll-research/SKILL.md +3 -2
  113. package/skills/ll-resume/SKILL.md +4 -3
  114. package/skills/ll-update/SKILL.md +6 -1
  115. package/skills/ll-verify/SKILL.md +2 -1
@@ -1,10 +1,29 @@
1
1
  #!/usr/bin/env bash
2
2
  # Smoke test do ll-skills: gate estático, hooks, helper, instalador, preâmbulo,
3
3
  # poda de skills antigas e uninstall. Tudo num CLAUDE_CONFIG_DIR isolado.
4
+ # Uso: smoke-test.sh [--only <seção>] (sem argumento roda tudo)
5
+ # As seções são os blocos numerados (1, 2, 3, 4, 4b, 4c, 4d, 5..10), lint-orquestrador,
6
+ # goal-autonomo, evals-auto, lint-scratch e no-talk; os blocos numerados montam estado uns para
7
+ # os outros, então --only puxa o pré-requisito de quem não roda sozinho (ver prereq abaixo).
4
8
  set -euo pipefail
5
9
  cd "$(dirname "$0")/.."
6
10
  ROOT="$PWD"
7
11
 
12
+ ONLY=""
13
+ while [ $# -gt 0 ]; do
14
+ case "$1" in
15
+ --only) ONLY="${2:-}"; shift 2 ;;
16
+ *) echo "uso: smoke-test.sh [--only <seção>]" >&2; exit 2 ;;
17
+ esac
18
+ done
19
+ # Pré-requisitos: a seção 4 usa o $EMPTY que a 2 monta e o $FIX que a 3 monta, a 6 lê o
20
+ # CLAUDE_CONFIG_DIR e o log que a 5 monta, e a 8 desinstala o que a 5 instalou e o preâmbulo que a 6
21
+ # escreve. --only puxa essas seções antes da pedida.
22
+ prereq() { case "$1" in 4) echo "2 3" ;; 6) echo "5" ;; 8) echo "5 6" ;; *) echo "" ;; esac; }
23
+ RUN=""
24
+ [ -z "$ONLY" ] || RUN=" $(prereq "$ONLY") $ONLY "
25
+ section() { [ -z "$ONLY" ] || case "$RUN" in *" $1 "*) return 0 ;; *) return 1 ;; esac; }
26
+
8
27
  TMP="$(mktemp -d)"
9
28
  trap 'rm -rf "$TMP"' EXIT
10
29
 
@@ -12,11 +31,13 @@ N=0
12
31
  check() { N=$((N + 1)); if ! eval "$2"; then echo "FALHOU: $1"; exit 1; fi; }
13
32
 
14
33
  HELPER="node $ROOT/scripts/ll-tools.js"
34
+ AUTO="node $ROOT/skills/ll-auto/scripts/ll-auto.js"
15
35
  region() { awk '/^\/\/ <ll-shared:state>$/,/^\/\/ <\/ll-shared:state>$/' "$1"; }
16
36
 
17
37
  # ---------------------------------------------------------------------------
18
38
  # 1. gate estático
19
39
  # ---------------------------------------------------------------------------
40
+ if section 1; then
20
41
  for f in bin/*.js hooks/*.js scripts/*.js; do node --check "$f"; done
21
42
 
22
43
  check "helper ≤700 linhas" '[ "$(wc -l < scripts/ll-tools.js)" -le 700 ]'
@@ -30,10 +51,12 @@ region scripts/ll-tools.js > "$TMP/shared-tools.txt"
30
51
  region hooks/ll-state.js > "$TMP/shared-hook.txt"
31
52
  check "região ll-shared:state existe" '[ -s "$TMP/shared-tools.txt" ] && [ -s "$TMP/shared-hook.txt" ]'
32
53
  check "região ll-shared:state idêntica" 'cmp -s "$TMP/shared-tools.txt" "$TMP/shared-hook.txt"'
54
+ fi
33
55
 
34
56
  # ---------------------------------------------------------------------------
35
57
  # 2. hooks silenciosos fora de um projeto
36
58
  # ---------------------------------------------------------------------------
59
+ if section 2; then
37
60
  EMPTY="$TMP/empty"
38
61
  mkdir -p "$EMPTY"
39
62
  cp -R scripts/fixtures/empty/. "$EMPTY/"
@@ -56,10 +79,12 @@ for h in ll-state.js ll-precompact.js; do
56
79
  check "$h <300 ms no vazio" '[ "$(cat "$TMP/hook.ms")" -lt 300 ]'
57
80
  done
58
81
  check "hooks não criaram arquivos no vazio" '[ "$(find "$EMPTY" | wc -l)" -eq "$BEFORE_FILES" ]'
82
+ fi
59
83
 
60
84
  # ---------------------------------------------------------------------------
61
85
  # 3. hooks no fixture
62
86
  # ---------------------------------------------------------------------------
87
+ if section 3; then
63
88
  FIX="$TMP/fix"
64
89
  bash scripts/fixtures/git-history.sh "$FIX" > /dev/null
65
90
 
@@ -88,10 +113,12 @@ check "precompact avisa a sessão" 'grep -q "compaction marked in PROGRESS.md
88
113
  check "precompact inseriu 1 linha" '[ "$(grep -c "^- \[compaction" "$FIX/PROGRESS.md")" -eq "$((C0 + 1))" ]'
89
114
  check "linha nova fica acima do epílogo" \
90
115
  'grep -n "^## Epilogue" "$FIX/PROGRESS.md" | head -1 | cut -d: -f1 | xargs -I{} sh -c "sed -n \"\$(({} - 2))p\" \"$FIX/PROGRESS.md\"" | grep -q "^- \[compaction .* manual "'
116
+ fi
91
117
 
92
118
  # ---------------------------------------------------------------------------
93
119
  # 4. os 12 comandos do helper
94
120
  # ---------------------------------------------------------------------------
121
+ if section 4; then
95
122
  cd "$FIX"
96
123
  $HELPER waves phases/07/PLAN.md --json > "$TMP/waves.json"
97
124
  check "waves: 4 ondas" 'grep -q "\"wave\":4" "$TMP/waves.json"'
@@ -103,14 +130,24 @@ $HELPER tdd-gate M1 --json > "$TMP/tdd1.json"
103
130
  check "tdd-gate M1 passa" 'grep -q "\"tdd\":\"pass\"" "$TMP/tdd1.json"'
104
131
  check "tdd-gate M1 ignora M10" '! grep -q "M10" "$TMP/tdd1.json"'
105
132
  check "tdd-gate M2 reprova" '$HELPER tdd-gate M2 --json | grep -q "\"tdd\":\"fail\""'
133
+ check "tdd-gate acha o id entre posicionais" '$HELPER tdd-gate phases/07/PLAN.md M1 --json | grep -q "\"tdd\":\"pass\""'
106
134
 
107
135
  check "spot-check M1 passa" '$HELPER spot-check M1 --files src/a.ts,test/a.test.ts --json | grep -q "\"verdict\":\"pass\""'
108
136
  check "state: fase 07, 1/3" '$HELPER state --json | grep -q "\"total\":3,\"passed\":1"'
109
137
  check "ledger: 1 FRESH 1 STALE 1 UNKNOWN" '$HELPER ledger --json | grep -q "\"FRESH\":1,\"STALE\":1,\"UNKNOWN\":1"'
110
138
  check "epilogue 07 aponta a fase 8" '$HELPER epilogue 07 --json | grep -q "\"next_command\":\"ll-implement 8\""'
111
139
  check "phase-stats: 5 dias com trabalho" '$HELPER phase-stats --json | grep -q "\"days_with_work\":5"'
140
+ check "phase-stats: fase 07 com questions/owner_prompts" \
141
+ '$HELPER phase-stats --json | grep -q "\"phase\":\"07\"" && $HELPER phase-stats --json | grep -q "\"questions\":3" && $HELPER phase-stats --json | grep -q "\"owner_prompts\":2"'
142
+ check "phase-stats: targets presente" '$HELPER phase-stats --json | grep -q "\"targets\""'
143
+ mkdir -p "$TMP/ep" && awk '/^## Epilogue — phase 07/{print "## Epilogue — phase 06 — 2026-09-08"; print "passed: M1 · left: none"; print ""} {print}' PROGRESS.md | sed 's/· amendments 1 · verification:/· amendments 1 (M3) · verification:/' > "$TMP/ep/PROGRESS.md"
144
+ $HELPER phase-stats --cwd "$TMP/ep" --json > "$TMP/ep.json"
145
+ check "phase-stats: epílogo sem linha de contagem não rouba a da fase seguinte" 'grep -q "\"phase\":\"06\",\"count_line\":false" "$TMP/ep.json"'
146
+ check "phase-stats: parêntese após amendments é aceito" 'grep -q "\"amendments\":1,\"verdict\":\"APPROVED\",\"verification\":\"phases/07/VERIFICATION.md\"" "$TMP/ep.json"'
147
+ check "epilogue 07 humano traz targets" '$HELPER epilogue 07 | grep -q "targets:"'
112
148
  check "plan-lint reprova o PLAN" '$HELPER plan-lint phases/07/PLAN.md --json | grep -q "\"verdict\":\"fail\""'
113
149
  check "plan-lint acusa acceptance-missing em M6" '$HELPER plan-lint phases/07/PLAN.md --json | grep -q "acceptance-missing"'
150
+ check "plan-lint aceita o número da fase" '$HELPER plan-lint 7 --json | grep -q "\"verdict\":\"fail\"" && $HELPER waves 07 --json | grep -q "\"waves\""'
114
151
 
115
152
  cp PROGRESS.md "$TMP/progress-before"
116
153
  $HELPER passes M2 true --json > "$TMP/passes.json"
@@ -118,6 +155,9 @@ check "passes M2 true grava" 'grep -q "\"passes\":true" "$TMP/passes.js
118
155
  check "passes muda exatamente 2 linhas" \
119
156
  '[ "$(diff "$TMP/progress-before" PROGRESS.md | grep -c "^[<>]")" -eq 2 ]'
120
157
  check "passes M9 cria a linha" '$HELPER passes M9 true --json | grep -q "\"created\":true"'
158
+ mkdir -p "$TMP/emptymap" && sed 's/^milestones:$/milestones: {}/' "$TMP/progress-before" | awk '/^milestones: \{\}$/{print; skip=1; next} skip && /^ (M[0-9]+|G-[0-9]+):/{next} {skip=0; print}' > "$TMP/emptymap/PROGRESS.md"
159
+ check "passes com milestones: {} cria a entrada" \
160
+ '$HELPER passes M1 true --cwd "$TMP/emptymap" --json | grep -q "\"created\":true" && grep -q "^milestones:$" "$TMP/emptymap/PROGRESS.md" && grep -q "^ M1: {" "$TMP/emptymap/PROGRESS.md"'
121
161
 
122
162
  check "dec-reserve 2 → 0042/0043" '$HELPER dec-reserve 2 --json | grep -q "\"DEC-0042\",\"DEC-0043\""'
123
163
  check "dec-reserve 2 → 0044/0045" '$HELPER dec-reserve 2 --json | grep -q "\"DEC-0044\",\"DEC-0045\""'
@@ -145,10 +185,12 @@ for cmd in waves tdd-gate spot-check state ledger backlog-reconcile epilogue pha
145
185
  "$HELPER $cmd M1 > \"$TMP/r.out\" 2>/dev/null; [ \$? -eq 0 ] && grep -q '\"ok\":false' \"$TMP/r.out\""
146
186
  done
147
187
  cd "$ROOT"
188
+ fi
148
189
 
149
190
  # ---------------------------------------------------------------------------
150
191
  # 4b. nested state root (PROGRESS.md/phases/ under docs/state/, sources at git top)
151
192
  # ---------------------------------------------------------------------------
193
+ if section 4b; then
152
194
  NEST="$TMP/nest"
153
195
  bash scripts/fixtures/git-history.sh "$NEST" > /dev/null
154
196
  mkdir -p "$NEST/docs/state"
@@ -166,6 +208,19 @@ check "nested state: git_top aponta ao topo do git" \
166
208
  check "nested ledger: FRESH/STALE/UNKNOWN iguais ao topo" \
167
209
  '$HELPER ledger --cwd "$NEST/docs/state" --json | grep -q "\"FRESH\":1,\"STALE\":1,\"UNKNOWN\":1"'
168
210
 
211
+ printf '{"cwd":"%s","source":"compact"}' "$NEST" | node hooks/ll-state.js > "$TMP/nest-hook.json"
212
+ check "nested hook ll-state acha o PROGRESS em docs/state" 'grep -q "post-compaction" "$TMP/nest-hook.json" && grep -q "milestones [0-9]/3" "$TMP/nest-hook.json"'
213
+ printf '{"cwd":"%s","trigger":"auto"}' "$NEST" | node hooks/ll-precompact.js > "$TMP/nest-pre.out"
214
+ check "nested hook precompact carimba docs/state/PROGRESS.md" 'grep -q "compaction marked" "$TMP/nest-pre.out" && grep -q "^- \[compaction " "$NEST/docs/state/PROGRESS.md"'
215
+ mkdir -p "$NEST/docs/other" && printf '# other progress\n\nno state block here\n' > "$NEST/docs/other/PROGRESS.md" && touch "$NEST/docs/other/PROGRESS.md"
216
+ printf '{"cwd":"%s","source":"compact"}' "$NEST" | node hooks/ll-state.js > "$TMP/nest-hook2.json"
217
+ check "nested hook: entre vários PROGRESS.md escolhe o que tem bloco ll-state" 'grep -q "milestones [0-9]/3" "$TMP/nest-hook2.json"'
218
+ rm -rf "$NEST/docs/other"
219
+ mkdir -p "$NEST/docs/fixtures/project" && cp "$NEST/docs/state/PROGRESS.md" "$NEST/docs/fixtures/project/PROGRESS.md" && mv "$NEST/docs/state" "$TMP/state-aside"
220
+ check "hook ignora PROGRESS.md dentro de fixtures/" '[ -z "$(printf "{\"cwd\":\"%s\",\"source\":\"startup\"}" "$NEST" | node hooks/ll-state.js)" ]'
221
+ mv "$TMP/state-aside" "$NEST/docs/state"
222
+ rm -rf "$NEST/docs/fixtures"
223
+
169
224
  $HELPER backlog-reconcile --run --cwd "$NEST/docs/state" --json > "$TMP/nest-backlog.json"
170
225
  check "nested backlog-reconcile fecha B-014 (comando roda no topo do git)" \
171
226
  'grep -q "\"closed\":\[\"B-014\"\]" "$TMP/nest-backlog.json"'
@@ -181,10 +236,96 @@ check "state do topo reporta git_top correto" \
181
236
  $HELPER spot-check M1 --files src/a.ts,test/a.test.ts --cwd "$NEST" --json > "$TMP/nest-top-spot.json"
182
237
  check "spot-check do topo acha a raiz deslocada" \
183
238
  'grep -q "\"verdict\":\"pass\"" "$TMP/nest-top-spot.json"'
239
+ fi
240
+
241
+ # ---------------------------------------------------------------------------
242
+ # 4c. helper ll-auto: detect
243
+ # ---------------------------------------------------------------------------
244
+ if section 4c; then
245
+ # ids e status exatos de um detect --json: renomear um status ou perder um estágio reprova
246
+ auto_ids() { node -e 'const o=JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"));console.log(o.stages.map((s)=>s.id).join(","))' "$1"; }
247
+ auto_status_set() { node -e 'const o=JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"));console.log([...new Set(o.stages.map((s)=>s.status))].sort().join(","))' "$1"; }
248
+ check "ll-auto detect: fixture project → decide done, 07 half" \
249
+ '$AUTO detect --cwd "$ROOT/scripts/fixtures/project" --json > "$TMP/auto-project.json" \
250
+ && grep -q "\"id\":\"decide\",\"status\":\"done\"" "$TMP/auto-project.json" \
251
+ && grep -q "\"id\":\"phase-07\",\"status\":\"half\"" "$TMP/auto-project.json"'
252
+ check "ll-auto detect: fixture empty → os 4 estágios exatos, todos todo" \
253
+ '$AUTO detect --cwd "$ROOT/scripts/fixtures/empty" --json > "$TMP/auto-empty.json" \
254
+ && [ "$(auto_ids "$TMP/auto-empty.json")" = "research,brainstorm,decide,close" ] \
255
+ && [ "$(auto_status_set "$TMP/auto-empty.json")" = "todo" ]'
256
+ check "ll-auto detect: sem ROADMAP.md as fases saem da tabela do §8 do PLAN" \
257
+ '$AUTO detect --cwd "$ROOT/scripts/fixtures/auto-noroadmap" --json > "$TMP/auto-noroadmap.json" \
258
+ && [ "$(auto_ids "$TMP/auto-noroadmap.json")" = "research,brainstorm,decide,phase-01,phase-02,verify-01,verify-02,close" ] \
259
+ && grep -q "\"id\":\"phase-01\",\"status\":\"done\"" "$TMP/auto-noroadmap.json" \
260
+ && grep -q "\"id\":\"phase-02\",\"status\":\"todo\"" "$TMP/auto-noroadmap.json"'
261
+ check "ll-auto detect: DELIVERY.md + epílogo da última fase → close done" \
262
+ '$AUTO detect --cwd "$ROOT/scripts/fixtures/auto-closed" --json > "$TMP/auto-closed.json" \
263
+ && grep -q "\"id\":\"close\",\"status\":\"done\",\"evidence\":\"docs/DELIVERY.md + epilogue for phase 02\"" "$TMP/auto-closed.json" \
264
+ && grep -q "\"id\":\"phase-02\",\"status\":\"done\"" "$TMP/auto-closed.json" \
265
+ && [ "$(auto_status_set "$TMP/auto-closed.json")" = "done,todo" ]'
266
+ check "ll-auto detect: diretório inexistente → ok:false e exit 0" \
267
+ '$AUTO detect --cwd /nonexistent --json > "$TMP/auto-nodir.json" 2>/dev/null; [ $? -eq 0 ] \
268
+ && grep -q "\"ok\":false" "$TMP/auto-nodir.json" && grep -q "not a directory" "$TMP/auto-nodir.json"'
269
+ check "ll-auto detect --json é JSON válido" \
270
+ '$AUTO detect --cwd "$ROOT/scripts/fixtures/project" --json \
271
+ | node -e "JSON.parse(require(\"fs\").readFileSync(0,\"utf8\"))"'
272
+ fi
273
+
274
+ # ---------------------------------------------------------------------------
275
+ # 4d. helper ll-auto: roteiro, next-cmd, report, auto-md
276
+ # ---------------------------------------------------------------------------
277
+ if section 4d; then
278
+ check "ll-auto roteiro: --verify all → 07, verify-07, 08, verify-08, close" \
279
+ '$AUTO roteiro --cwd "$ROOT/scripts/fixtures/project" --flags "--verify all" --json > "$TMP/auto-verify-all.json" \
280
+ && grep -q "\"stage\":\"phase-07\".*\"stage\":\"verify-07\".*\"stage\":\"phase-08\".*\"stage\":\"verify-08\".*\"stage\":\"close\"" "$TMP/auto-verify-all.json" \
281
+ && ! grep -qE "\"stage\":\"(research|brainstorm|decide|phase-05|phase-06)\"" "$TMP/auto-verify-all.json"'
282
+ check "ll-auto roteiro: --only 8 corta o close e as outras fases" \
283
+ '$AUTO roteiro --cwd "$ROOT/scripts/fixtures/project" --flags "--only 8" --json > "$TMP/auto-only8.json" \
284
+ && grep -q "\"stage\":\"phase-08\",\"command\":\"ll-implement 08 --no-talk\"" "$TMP/auto-only8.json" \
285
+ && ! grep -qE "\"stage\":\"(phase-07|close)\"" "$TMP/auto-only8.json"'
286
+ check "ll-auto roteiro: repo vazio sem objetivo pede o objetivo" \
287
+ '$AUTO roteiro --cwd "$ROOT/scripts/fixtures/empty" --flags "" --json > "$TMP/auto-vazio.json" \
288
+ && grep -q "\"needs_objective\":true" "$TMP/auto-vazio.json" \
289
+ && grep -q "\"roteiro\":\[\]" "$TMP/auto-vazio.json"'
290
+ check "ll-auto roteiro: repo vazio com objetivo → research, brainstorm, decide, close" \
291
+ '$AUTO roteiro --cwd "$ROOT/scripts/fixtures/empty" --objective "um objetivo" \
292
+ --flags "--research --brainstorm" --json > "$TMP/auto-vazio2.json" \
293
+ && grep -q "\"stage\":\"research\".*\"stage\":\"brainstorm\".*\"stage\":\"decide\".*\"stage\":\"close\"" "$TMP/auto-vazio2.json" \
294
+ && grep -q "\"needs_objective\":false" "$TMP/auto-vazio2.json"'
295
+ check "ll-auto roteiro: --interactive --redo phase-05 --pause-at 8" \
296
+ '$AUTO roteiro --cwd "$ROOT/scripts/fixtures/project" \
297
+ --flags "--interactive --redo phase-05 --pause-at 8" --json > "$TMP/auto-redo.json" \
298
+ && grep -q "\"stage\":\"phase-05\",\"command\":\"ll-implement 05\",\"status\":\"done\"" "$TMP/auto-redo.json" \
299
+ && ! grep -q "ll-implement 05 --no-talk" "$TMP/auto-redo.json" \
300
+ && grep -q "\"stage\":\"phase-08\",\"command\":\"ll-implement 08\",\"status\":\"todo\",\"pause_after\":true" "$TMP/auto-redo.json"'
301
+ check "ll-auto roteiro: epílogo que pede ll-verify 01 insere verify-01 sem --verify all" \
302
+ '$AUTO roteiro --cwd "$ROOT/scripts/fixtures/auto-verify-next" --flags "" --json > "$TMP/auto-vnext.json" \
303
+ && grep -q "\"stage\":\"phase-01\",\"command\":\"ll-implement 01 --no-talk\".*\"stage\":\"verify-01\",\"command\":\"ll-verify 01\"" "$TMP/auto-vnext.json" \
304
+ && $AUTO roteiro --cwd "$ROOT/scripts/fixtures/project" --flags "" --json > "$TMP/auto-noverify.json" \
305
+ && ! grep -q "\"stage\":\"verify-" "$TMP/auto-noverify.json"'
306
+ check "ll-auto next-cmd: o epílogo da fixture aponta ll-implement 8" \
307
+ '[ "$($AUTO next-cmd "$ROOT/scripts/fixtures/project/PROGRESS.md")" = "ll-implement 8" ]'
308
+ check "ll-auto next-cmd: arquivo sem ▶ Next → saída vazia, exit 0" \
309
+ '[ -z "$($AUTO next-cmd "$ROOT/scripts/fixtures/project/ROADMAP.md")" ]'
310
+ check "ll-auto report: só o DEC com a marca de decidido sozinho" \
311
+ '$AUTO report --cwd "$ROOT/scripts/fixtures/auto-decisions" --json > "$TMP/auto-report.json" \
312
+ && grep -q "DEC-0001-taken-alone.md" "$TMP/auto-report.json" \
313
+ && ! grep -q "DEC-0002-owner.md" "$TMP/auto-report.json"'
314
+ check "ll-auto auto-md: as cinco seções, o objetivo e as flags" \
315
+ '$AUTO auto-md --cwd "$ROOT/scripts/fixtures/project" --objective "um objetivo" --flags "--only 8" \
316
+ > "$TMP/auto-md.txt" \
317
+ && grep -q "^## Objective" "$TMP/auto-md.txt" && grep -q "^## Flags" "$TMP/auto-md.txt" \
318
+ && grep -q "^## Roteiro" "$TMP/auto-md.txt" && grep -q "^## Decisions taken alone" "$TMP/auto-md.txt" \
319
+ && grep -q "^## Log" "$TMP/auto-md.txt" \
320
+ && grep -q "^um objetivo$" "$TMP/auto-md.txt" && grep -q -e "^--only 8$" "$TMP/auto-md.txt" \
321
+ && grep -q "^| # | stage | command | status | evidence |$" "$TMP/auto-md.txt" \
322
+ && grep -q "^| 1 | phase-08 | ll-implement 08 --no-talk | todo |" "$TMP/auto-md.txt"'
323
+ fi
184
324
 
185
325
  # ---------------------------------------------------------------------------
186
326
  # 5. instalador
187
327
  # ---------------------------------------------------------------------------
328
+ if section 5; then
188
329
  export CLAUDE_CONFIG_DIR="$TMP/cfg" XDG_CACHE_HOME="$TMP/cache"
189
330
  mkdir -p "$CLAUDE_CONFIG_DIR"
190
331
 
@@ -208,7 +349,7 @@ echo '{"version":"0.0.1","files":{"skills/ll-antiga/SKILL.md":"a","skills/alheia
208
349
 
209
350
  node bin/install.js < /dev/null > "$TMP/install1.log"
210
351
 
211
- check "11 skills ll-* instaladas" '[ "$(ls "$CLAUDE_CONFIG_DIR/skills" | grep -c "^ll-")" -eq 11 ]'
352
+ check "12 skills ll-* instaladas" '[ "$(ls "$CLAUDE_CONFIG_DIR/skills" | grep -c "^ll-")" -eq 12 ]'
212
353
  for s in ll-brainstorm ll-research ll-decide ll-goal ll-implement ll-verify ll-close ll-resume ll-refine ll-oncall ll-update; do
213
354
  check "skill $s instalada" '[ -f "$CLAUDE_CONFIG_DIR/skills/'"$s"'/SKILL.md" ]'
214
355
  done
@@ -240,6 +381,11 @@ for s in ll-implement ll-verify ll-close; do
240
381
  '[ "$(sha256sum "$CLAUDE_CONFIG_DIR/skills/'"$s"'/scripts/ll-tools.js" | cut -d" " -f1)" = "$SHA_SRC" ]'
241
382
  done
242
383
 
384
+ SHA_AUTO_SRC="$(sha256sum skills/ll-auto/scripts/ll-auto.js | cut -d" " -f1)"
385
+ check "ll-auto.js instalado executável" '[ -x "$CLAUDE_CONFIG_DIR/skills/ll-auto/scripts/ll-auto.js" ]'
386
+ check "ll-auto.js instalado com sha idêntico" \
387
+ '[ "$(sha256sum "$CLAUDE_CONFIG_DIR/skills/ll-auto/scripts/ll-auto.js" | cut -d" " -f1)" = "$SHA_AUTO_SRC" ]'
388
+
243
389
  cp "$CLAUDE_CONFIG_DIR/settings.json" "$TMP/s1.json"
244
390
  node bin/install.js < /dev/null > /dev/null
245
391
  check "segunda instalação sem diff no settings" 'cmp -s "$CLAUDE_CONFIG_DIR/settings.json" "$TMP/s1.json"'
@@ -260,10 +406,12 @@ mkdir -p "$XDG_CACHE_HOME/ll-skills"
260
406
  V="$(cat "$CLAUDE_CONFIG_DIR/ll-skills/VERSION")"
261
407
  printf '{"checked":%s,"installed":"%s","latest":"99.0.0","update_available":true,"method":"npm"}' "$(date +%s)" "$V" > "$XDG_CACHE_HOME/ll-skills/update-check.json"
262
408
  check "leitor avisa versão nova" 'node "$CLAUDE_CONFIG_DIR/hooks/ll-skills-check-update.js" | grep -q "/ll-update"'
409
+ fi
263
410
 
264
411
  # ---------------------------------------------------------------------------
265
412
  # 6. preâmbulo
266
413
  # ---------------------------------------------------------------------------
414
+ if section 6; then
267
415
  check "não-TTY sem --yes não escreve CLAUDE.md" '[ ! -e "$CLAUDE_CONFIG_DIR/CLAUDE.md" ]'
268
416
  check "não-TTY imprime o diff e a dica --yes" 'grep -q -- "--yes" "$TMP/install1.log"'
269
417
 
@@ -284,10 +432,12 @@ sed -i 's/^## Skills$/## Skills MEXIDO A MAO/' "$CMD"
284
432
  node bin/install.js --yes < /dev/null > /dev/null
285
433
  check "corpo alterado à mão é restaurado" 'cmp -s "$CMD" "$TMP/claudemd-1"'
286
434
  check "backup .ll-skills.bak existe" '[ -f "$CMD.ll-skills.bak" ]'
435
+ fi
287
436
 
288
437
  # ---------------------------------------------------------------------------
289
438
  # 7. poda das skills 1.x (com e sem manifesto)
290
439
  # ---------------------------------------------------------------------------
440
+ if section 7; then
291
441
  LEGACY_SKILLS="ll-atualizar ll-decidir-antes ll-desarmar ll-orquestrar ll-pesquisar ll-pesquisar-mercado ll-verificar-entrega ll-voltar-do-futuro"
292
442
  CFG_L="$TMP/cfg-legacy"
293
443
  seed_legacy() {
@@ -318,10 +468,12 @@ assert_pruned "com manifesto"
318
468
  seed_legacy # agora o manifesto atual (2.x) não lista nada disso: só KNOWN_LEGACY resolve
319
469
  CLAUDE_CONFIG_DIR="$CFG_L" node bin/install.js --no-preamble < /dev/null > /dev/null
320
470
  assert_pruned "sem manifesto"
471
+ fi
321
472
 
322
473
  # ---------------------------------------------------------------------------
323
474
  # 8. uninstall
324
475
  # ---------------------------------------------------------------------------
476
+ if section 8; then
325
477
  node bin/install.js --uninstall < /dev/null > /dev/null
326
478
  check "uninstall removeu skills" '[ -z "$(ls "$CLAUDE_CONFIG_DIR/skills" | grep "^ll-" || true)" ]'
327
479
  check "uninstall removeu as cópias do helper" '[ ! -e "$CLAUDE_CONFIG_DIR/skills/ll-implement" ]'
@@ -331,10 +483,237 @@ check "uninstall removeu o check-update" '! grep -q "ll-skills-check-update" "$C
331
483
  check "uninstall preservou alheios" 'grep -q "echo alheio" "$CLAUDE_CONFIG_DIR/settings.json" && [ -f "$CLAUDE_CONFIG_DIR/skills/alheia/SKILL.md" ]'
332
484
  check "uninstall removeu o preâmbulo" '! grep -q "ll-skills:preamble" "$CMD"'
333
485
  check "uninstall preservou as sentinelas" 'grep -q "SENTINEL-TOP" "$CMD" && grep -q "SENTINEL-BOTTOM" "$CMD"'
486
+ fi
334
487
 
488
+ # ---------------------------------------------------------------------------
489
+ # 9. lint dos prompts (uma checagem por regra de scripts/lint-prompts.sh)
490
+ # ---------------------------------------------------------------------------
491
+ if section 9; then
492
+ lint() { # lint <n>: roda uma regra e só imprime a saída quando ela falha
493
+ bash "$ROOT/scripts/lint-prompts.sh" --rule "$1" > "$TMP/lint-$1.txt" 2>&1 \
494
+ || { cat "$TMP/lint-$1.txt"; return 1; }
495
+ }
496
+
497
+ check "lint 1: frontmatter das skills" 'lint 1'
498
+ check "lint 2: frontmatter dos agentes" 'lint 2'
499
+ check "lint 3: tetos de linhas" 'lint 3'
500
+ printf '[{"unpackedSize":999999999}]' > "$TMP/pack-list-big.json"
501
+ check "lint 3: honra LL_PACK_JSON na forma lista (npm ≤ 11) e reprova no teto sintético" \
502
+ '! LL_PACK_JSON="$TMP/pack-list-big.json" bash "$ROOT/scripts/lint-prompts.sh" --rule 3 > "$TMP/lint-pack-list.out" 2>&1; \
503
+ grep -q "unpackedSize 999999999 bytes, ceiling 921600" "$TMP/lint-pack-list.out"'
504
+ printf '{"ll-skills":{"unpackedSize":999999999}}' > "$TMP/pack-object-big.json"
505
+ check "lint 3: honra LL_PACK_JSON na forma objeto (npm 12) e reprova no teto sintético" \
506
+ '! LL_PACK_JSON="$TMP/pack-object-big.json" bash "$ROOT/scripts/lint-prompts.sh" --rule 3 > "$TMP/lint-pack-object.out" 2>&1; \
507
+ grep -q "unpackedSize 999999999 bytes, ceiling 921600" "$TMP/lint-pack-object.out"'
508
+ check "lint 4: forma das SKILL.md" 'lint 4'
509
+ check "lint 5: strings proibidas" 'lint 5'
510
+ check "lint 6: cópias idênticas" 'lint 6'
511
+ check "lint 7: idioma" 'lint 7'
512
+ check "lint 8: lista privada" 'lint 8'
513
+ fi
514
+
515
+ # 10. lint de contrato entre as peças (uma checagem por regra de scripts/lint-contract.cjs)
516
+ if section 10; then
517
+ contract() { node "$ROOT/scripts/lint-contract.cjs" --rule "$1" >/dev/null 2>&1; }
518
+ check "contrato 1: references citadas existem e são citadas" 'contract 1'
519
+ check "contrato 2: comandos do helper definidos e citados" 'contract 2'
520
+ check "contrato 3: agentes, modelos e campos do brief" 'contract 3'
521
+ check "contrato 4: vocabulário de veredito e estados" 'contract 4'
522
+ check "contrato 5: arquivos de estado nas tabelas de entregáveis" 'contract 5'
523
+ check "contrato 6: alvos do ▶ Next existem" 'contract 6'
524
+ check "contrato 7: instalador × pacote" 'contract 7'
525
+
526
+ # regra 6 contra as fixtures da gramática do ▶ Next (bad falha, good passa, raiz sem skills falha)
527
+ contract_root() { node "$ROOT/scripts/lint-contract.cjs" --rule 6 --root "$1" >/dev/null 2>&1; }
528
+ contract_fails() { node "$ROOT/scripts/lint-contract.cjs" --rule 6 --root "$1" 2>/dev/null | grep -c "^FAIL 6 " || true; }
529
+ check "contrato 6: next-bad falha" '! contract_root scripts/fixtures/next-bad'
530
+ # a contagem é fixa: afrouxar uma das formas rejeitadas reprova aqui, não só o exit
531
+ check "contrato 6: next-bad com 6 FAIL" '[ "$(contract_fails scripts/fixtures/next-bad)" -eq 6 ]'
532
+ check "contrato 6: duas skills diferentes num ▶ Next fora de parêntese reprovam" \
533
+ 'node "$ROOT/scripts/lint-contract.cjs" --rule 6 --root scripts/fixtures/next-bad > "$TMP/next-bad.out" 2>/dev/null; \
534
+ grep -q "lists alternatives outside a parenthetical —ll-bad or ll-nope" "$TMP/next-bad.out"'
535
+ check "contrato 6: next-good passa" 'contract_root scripts/fixtures/next-good'
536
+ check "contrato 6: raiz sem skills falha" '! contract_root scripts/fixtures/empty'
335
537
 
336
538
  # --- optional private word list (never shipped): LL_FORBIDDEN_FILE=<path> enables the check
337
539
  if [ -n "${LL_FORBIDDEN_FILE:-}" ] && [ -f "$LL_FORBIDDEN_FILE" ]; then
338
540
  check "no private terms in tracked files" '! git ls-files | grep -v "^\.gitignore$" | xargs grep -n -i -w -E -f "$LL_FORBIDDEN_FILE" 2>/dev/null | grep -q .'
339
541
  fi
542
+ fi
543
+
544
+ # ---------------------------------------------------------------------------
545
+ # lint-orquestrador. as exceções de ll-auto valem só para ll-auto
546
+ # ---------------------------------------------------------------------------
547
+ if section lint-orquestrador; then
548
+ cd "$ROOT"
549
+ SCR="$TMP/lint-scratch"
550
+ mkdir -p "$SCR"
551
+ git ls-files -z | xargs -0 cp --parents -t "$SCR"
552
+ git -C "$SCR" init -q
553
+ git -C "$SCR" add -A
554
+ mkdir -p "$SCR/skills/ll-fake"
555
+ cat > "$SCR/skills/ll-fake/SKILL.md" <<'FAKE'
556
+ ---
557
+ name: ll-fake
558
+ description: Pretends to be an orchestrator so the lint can prove the ll-auto exceptions do not leak to any other skill on the tree.
559
+ argument-hint: "[--flag]"
560
+ disable-model-invocation: true
561
+ allowed-tools: Bash(${CLAUDE_SKILL_DIR}/scripts/ll-auto.js *)
562
+ ---
563
+
564
+ # Fake
565
+
566
+ ## Deliverables
567
+
568
+ | File | Role | Mutability |
569
+ |---|---|---|
570
+ | `docs/FAKE.md` | nothing | never |
571
+
572
+ ## Flow
573
+
574
+ /ll-research x
575
+
576
+ ## Completion criterion
577
+
578
+ ▶ Next — /clear, then ll-resume
579
+ FAKE
580
+ git -C "$SCR" add -A
581
+ lint_scr() { bash "$SCR/scripts/lint-prompts.sh" --rule "$1" > "$TMP/scr-rule$1.out" 2>&1; }
582
+ lint_real() { bash "$ROOT/scripts/lint-prompts.sh" --rule "$1" >/dev/null 2>&1; }
583
+
584
+ check "orquestrador: skills/ll-auto/SKILL.md existe" '[ -f "$ROOT/skills/ll-auto/SKILL.md" ]'
585
+ check "orquestrador: regra 1 reprova o allowed-tools do ll-auto em outra skill" \
586
+ '! lint_scr 1 && grep -q "^FAIL 1 .*skills/ll-fake/SKILL.md: allowed-tools" "$TMP/scr-rule1.out"'
587
+ check "orquestrador: regra 5 reprova a linha /ll- em outra skill" \
588
+ '! lint_scr 5 && grep -q "^FAIL 5 .*skills/ll-fake/SKILL.md: line .* invokes a skill as a command" "$TMP/scr-rule5.out"'
589
+ check "orquestrador: regras 1 e 5 passam na árvore real (com ll-auto)" \
590
+ 'lint_real 1 && lint_real 5'
591
+ fi
592
+
593
+ # ---------------------------------------------------------------------------
594
+ # goal-autonomo. o modo --autonomous do ll-goal: hint, template e exemplo
595
+ # ---------------------------------------------------------------------------
596
+ if section goal-autonomo; then
597
+ cd "$ROOT"
598
+ GOAL_SKILL="skills/ll-goal/SKILL.md"
599
+ GOAL_TPL="skills/ll-goal/references/goal-template.md"
600
+ # primeiro bloco cercado sob "## Autonomous example", sem as cercas
601
+ goal_example() { awk '/^## Autonomous example$/{f=1;next} f&&/^```/{if(b){exit}b=1;next} f&&b' "$GOAL_TPL"; }
602
+
603
+ check "ll-goal argument-hint aceita --autonomous" \
604
+ 'sed -n "/^argument-hint:/p" "$GOAL_SKILL" | grep -q -- "--autonomous"'
605
+ check "goal-template cita ll-auto --auto-decision" \
606
+ '[ "$(grep -c "ll-auto --auto-decision" "$GOAL_TPL")" -ge 1 ]'
607
+ check "exemplo autônomo entre 500 e 4000 bytes" \
608
+ 'n=$(goal_example | wc -c); [ "$n" -ge 500 ] && [ "$n" -le 4000 ]'
609
+ check "goal-template ≤ 150 linhas" \
610
+ '[ "$(wc -l < "$GOAL_TPL")" -le 150 ]'
611
+ fi
612
+
613
+ # ---------------------------------------------------------------------------
614
+ # evals-auto. os asserts dos casos autônomos provados offline, sem nenhuma chamada paga
615
+ # ---------------------------------------------------------------------------
616
+ if section evals-auto; then
617
+ cd "$ROOT"
618
+ EV="$TMP/evals-auto"
619
+ mkdir -p "$EV"
620
+ # a resposta errada: não carrega nenhuma linha que os asserts exigem
621
+ printf 'I ran nothing and wrote nothing.\n' > "$EV/wrong.txt"
622
+
623
+ # a captura limpa: um evento assistant de texto (para o no_tool_use varrer algo) e o result.
624
+ CAP_CLEAN='[{"type":"assistant","message":{"role":"assistant","content":[{"type":"text","text":"Rodei o dry run e nao escrevi nada."}]}},{"type":"result","subtype":"success","is_error":false,"result":""}]'
625
+ # a mesma captura com uma chamada de skill por tool_use: o assert tem de reprovar.
626
+ CAP_SKILL='[{"type":"assistant","message":{"role":"assistant","content":[{"type":"tool_use","id":"t1","name":"Skill","input":{"command":"ll-implement"}}]}},{"type":"result","subtype":"success","is_error":false,"result":""}]'
627
+
628
+ # uma árvore de trabalho limpa por chamada (os asserts leem `git status --porcelain` dela)
629
+ auto_assert() { # auto_assert <caso> <arquivo de resposta> [captura]
630
+ local w="$EV/$1"
631
+ rm -rf "$w"; mkdir -p "$w"
632
+ git -C "$w" init -q -b main
633
+ printf '%s' "${3:-$CAP_CLEAN}" > "$w/out.json"
634
+ bash "$ROOT/scripts/evals/cases/$1/assert.sh" "$w" "$w/out.json" "$2" > "$EV/$1.log" 2>&1
635
+ }
636
+
637
+ check "a captura offline tem ao menos um evento assistant" \
638
+ 'node -e "const a=JSON.parse(process.argv[1]);process.exit(a.filter(e=>e.type===\"assistant\").length?0:1)" "$CAP_CLEAN"'
639
+
640
+ check "assert auto-dry-run aceita a resposta boa" \
641
+ 'auto_assert auto-dry-run "$ROOT/scripts/fixtures/evals-auto/auto-dry-run/pass.txt"'
642
+ check "assert auto-dry-run rejeita resposta errada" \
643
+ '! auto_assert auto-dry-run "$EV/wrong.txt"'
644
+ check "assert auto-empty-repo aceita a resposta boa" \
645
+ 'auto_assert auto-empty-repo "$ROOT/scripts/fixtures/evals-auto/auto-empty-repo/pass.txt"'
646
+ check "assert auto-empty-repo rejeita resposta errada" \
647
+ '! auto_assert auto-empty-repo "$EV/wrong.txt"'
648
+ # o no_tool_use não é vazio: uma captura com tool_use Skill reprova o mesmo assert
649
+ check "assert auto-dry-run rejeita captura com tool_use Skill" \
650
+ '! auto_assert auto-dry-run "$ROOT/scripts/fixtures/evals-auto/auto-dry-run/pass.txt" "$CAP_SKILL"'
651
+ check "assert auto-empty-repo rejeita captura com tool_use Skill" \
652
+ '! auto_assert auto-empty-repo "$ROOT/scripts/fixtures/evals-auto/auto-empty-repo/pass.txt" "$CAP_SKILL"'
653
+
654
+ # goal-autonomous: o assert espera docs/GOAL.md já no lugar (um `ll-goal --autonomous` real
655
+ # escreve e comita o arquivo antes de imprimir o texto colável), não só o out.json do harness.
656
+ goal_autonomous_assert() { # goal_autonomous_assert <arquivo de resposta> [captura]
657
+ local w="$EV/goal-autonomous"
658
+ rm -rf "$w"; mkdir -p "$w/docs"
659
+ git -C "$w" init -q -b main
660
+ printf 'mode: autonomous\nphase: all\n' > "$w/docs/GOAL.md"
661
+ git -C "$w" add -A && git -C "$w" -c user.email=t@t -c user.name=t commit -q -m "docs: GOAL.md"
662
+ printf '%s' "${2:-$CAP_CLEAN}" > "$w/out.json"
663
+ bash "$ROOT/scripts/evals/cases/goal-autonomous/assert.sh" "$w" "$w/out.json" "$1" > "$EV/goal-autonomous.log" 2>&1
664
+ }
665
+ check "assert goal-autonomous aceita a resposta boa" \
666
+ 'goal_autonomous_assert "$ROOT/scripts/fixtures/evals-auto/goal-autonomous/pass.txt"'
667
+ check "assert goal-autonomous rejeita resposta vazia" \
668
+ '! goal_autonomous_assert "$EV/wrong.txt"'
669
+ echo "evals-auto: goal-autonomous pair asserted"
670
+ fi
671
+
672
+ # ---------------------------------------------------------------------------
673
+ # lint-scratch. a regra 1 do lint-prompts contra SKILL.md ruins, numa árvore de rascunho
674
+ # ---------------------------------------------------------------------------
675
+ if section lint-scratch; then
676
+ cd "$ROOT"
677
+ LSCR="$TMP/lint-bad-scratch"
678
+ mkdir -p "$LSCR"
679
+ git ls-files -z | xargs -0 cp --parents -t "$LSCR"
680
+ git -C "$LSCR" init -q
681
+ mkdir -p "$LSCR/skills/ll-fake"
682
+ # lint_bad <caso> <trecho esperado no FAIL>: planta a fixture como skills/ll-fake e roda a regra 1
683
+ lint_bad() {
684
+ cp "$ROOT/scripts/fixtures/lint-bad/$1/SKILL.md" "$LSCR/skills/ll-fake/SKILL.md"
685
+ git -C "$LSCR" add -A
686
+ bash "$LSCR/scripts/lint-prompts.sh" --rule 1 > "$TMP/lint-bad-$1.out" 2>&1 && return 1
687
+ grep -q "^FAIL 1 .*skills/ll-fake/SKILL.md: $2" "$TMP/lint-bad-$1.out"
688
+ }
689
+ check "lint 1: descrição dobrada (>) em mais de uma linha reprova" \
690
+ 'lint_bad folded-description "description is not one line"'
691
+ check "lint 1: disable-model-invocation: false reprova" \
692
+ 'lint_bad model-invocation-false "disable-model-invocation: true missing"'
693
+ # sem a fixture a mesma árvore passa: o FAIL vem da SKILL.md ruim, não da cópia
694
+ rm -rf "$LSCR/skills/ll-fake"
695
+ git -C "$LSCR" add -A
696
+ check "lint 1: a árvore de rascunho sem a fixture passa" \
697
+ 'bash "$LSCR/scripts/lint-prompts.sh" --rule 1 > "$TMP/lint-bad-clean.out" 2>&1'
698
+ fi
699
+
700
+ # ---------------------------------------------------------------------------
701
+ # no-talk. o modo silencioso continua declarado no argument-hint de quem o implementa
702
+ # ---------------------------------------------------------------------------
703
+ if section no-talk; then
704
+ cd "$ROOT"
705
+ hint_has_no_talk() { # hint_has_no_talk <SKILL.md>
706
+ sed -n '/^argument-hint:/p' "$1" > "$TMP/hint.txt"
707
+ grep -q -- "--no-talk" "$TMP/hint.txt"
708
+ }
709
+ for s in ll-decide ll-close; do
710
+ check "argument-hint de $s declara --no-talk" 'hint_has_no_talk "skills/'"$s"'/SKILL.md"'
711
+ done
712
+ # a checagem não é vazia: sem a flag na linha, ela reprova
713
+ NT="$TMP/no-talk"
714
+ mkdir -p "$NT"
715
+ sed 's/ \[--no-talk\]//' skills/ll-decide/SKILL.md > "$NT/SKILL.md"
716
+ check "a checagem reprova quando --no-talk some do hint" '! hint_has_no_talk "$NT/SKILL.md"'
717
+ fi
718
+
340
719
  echo "smoke test OK — $N checks"
@@ -0,0 +1,74 @@
1
+ ---
2
+ name: ll-auto
3
+ description: Drives the whole cycle from the state on disk — research, brainstorm, decide, phases, verifications, close — following each stage's own skill in place with the flags it was given, and reports at the end every decision it took alone.
4
+ argument-hint: "[\"<objective>\"] [--research] [--brainstorm] [--interactive] [--auto-decision] [--pause-at <stage|N>] [--from N] [--to N] [--only N] [--verify all] [--redo <stage>] [--dry-run] [--resume]"
5
+ disable-model-invocation: true
6
+ allowed-tools: Bash(${CLAUDE_SKILL_DIR}/scripts/ll-auto.js *)
7
+ ---
8
+
9
+ # Auto
10
+
11
+ Current state: !`${CLAUDE_SKILL_DIR}/scripts/ll-auto.js detect`
12
+
13
+ One invocation, the cycle the repository still owes: the helper reads the state from disk, the flags cut it into a roteiro, and this session walks the roteiro stage by stage, following each stage's own skill as a workflow file. `docs/AUTO.md` carries the run: objective, flags, the roteiro table and one log line per transition, so a compaction or a stop never loses the position. The run starts from the middle when the middle is what exists.
14
+
15
+ Four boundaries hold for the whole run:
16
+ - **The Skill tool is never called.** A stage runs by reading `${CLAUDE_CONFIG_DIR:-$HOME/.claude}/skills/<skill>/SKILL.md` with `Read` (falling back to `skills/<skill>/SKILL.md` in the package tree when this repository is the one running) and following its Flow in place with the stage's arguments. Every skill is locked, this one included; the lock is not worked around, it is honoured by reading the instructions instead of invoking them.
17
+ - **Only this skill follows another skill.** The exception exists because the owner typed the command, and it lasts only for this invocation. No stage started here starts a further `ll-auto`.
18
+ - **No spending ceiling, ever.** Never stop, throttle or shorten a stage to save tokens or money, never announce a budget, never ask whether to continue for cost. The run ends when the roteiro ends, a pause flag says so, or nothing can move without the owner.
19
+ - **No question of its own.** The run never opens `AskUserQuestion`. The only questions the owner sees are the ones a stage asks while running under `--interactive`. When something is missing, the run prints the command that would complete it and stops.
20
+
21
+ Reply to the owner in Portuguese; every file written is in English.
22
+
23
+ ## Deliverables
24
+
25
+ | File | Role | Mutability |
26
+ |---|---|---|
27
+ | `docs/AUTO.md` | the run: objective, flags, roteiro table, decisions taken alone, log | rewritten by this skill on every transition; the roteiro rows only through `auto-md` and the log append-only |
28
+ | `docs/research-*/SUMMARY.md`, `docs/decide/`, `PLAN.md`, `ROADMAP.md`, `PROGRESS.md`, `phases/NN/`, `decisions/`, `BACKLOG.md`, `VERIFICATION.md`, `docs/DELIVERY.md` | whatever each stage writes, by the rules of that stage's own skill | owned by the stage skill; this skill writes none of them directly |
29
+
30
+ ## Flow
31
+
32
+ Helper: `${CLAUDE_SKILL_DIR}/scripts/ll-auto.js`, written `ll-auto.js` below. Read commands exit 0 and answer `{"ok":false,"reason":…}` on a broken input. Commands: `detect`, `roteiro`, `next-cmd`, `report`, `auto-md`.
33
+
34
+ ### 0. State, flags and roteiro
35
+
36
+ `ll-auto.js detect --json` gives the stage table above; `ll-auto.js roteiro --flags "<the arguments>" --objective "<objective>" --json` turns it into the ordered list of stages to run, with a command and a status per entry. `references/stages.md` holds the detection rules, the command per stage and the flag table.
37
+
38
+ Three stops before any work:
39
+ - `needs_objective: true` (no research, no `docs/decide/OPENING.md`, no `PLAN.md`, and no objective in the arguments) — print exactly two lines and stop, with no question: `Nada encontrado neste repositório: sem pesquisa, OPENING.md nem PLAN.md.` and the command to complete, `/ll-auto "<objetivo>" [--research] [--brainstorm]`.
40
+ - `--dry-run` — print the stage table from `detect` (one line per stage, as the helper prints it) and the roteiro table, then stop, before writing `docs/AUTO.md`.
41
+ - an empty roteiro with an objective present — say the cycle has nothing left and hand over.
42
+
43
+ `--resume`, or a plain invocation in a repository that already has `docs/AUTO.md`, reads the flags from that file's `## Flags` section and continues from the first row that is not `done`.
44
+
45
+ ### 1. Open the run
46
+
47
+ Write `docs/AUTO.md` with `ll-auto.js auto-md --objective "<objective>" --flags "<the arguments>"`, redirected to the file. Print the roteiro table to the owner in one screen: this is the plan of the run and the last moment before it starts.
48
+
49
+ ### 2. One stage at a time
50
+
51
+ For every roteiro entry, in order:
52
+ 1. Mark the row `running` in `docs/AUTO.md` and append the log line `<date> <stage> running`.
53
+ 2. Read that stage's `SKILL.md` at the path of boundary 1 and follow its Flow in place, with the entry's command as its arguments. The stage owns its files, its questions and its own ▶ Next line; that line is read, not pasted to the owner, and never starts another skill from here.
54
+ 3. Run `ll-auto.js detect --json` again. Mark the row `done` when the stage's own status turned `done`, `half` when it did not, and append the log line with the evidence the helper gives.
55
+ 4. List the `decisions/*.md` in state `WAITING` and apply `references/run.md`: without `--auto-decision` skip the stages that depend on them and stop when nothing else can run; with `--auto-decision` resolve each to its recommended option and continue.
56
+ 5. When the entry carries `pause_after: true` (`--pause-at`), stop here with `▶ Next — /clear, then ll-auto --resume`.
57
+
58
+ A stage that fails inside its own skill ends the run at that row, marked with its status and its last line in the log; the roteiro below it stays `todo`.
59
+
60
+ ### 3. Close the run
61
+
62
+ `ll-auto.js report --json` lists every `decisions/*.md` carrying `[decided by absence — revisable]`. Print that list as the end-of-run block, one line per file with its title, and write the same block into the `## Decisions taken alone` section of `docs/AUTO.md`. Then the count line and the handover. The end block, verbatim, is in `references/run.md`.
63
+
64
+ ## Completion criterion
65
+
66
+ Every roteiro row is `done`, `skipped` or `waiting` in `docs/AUTO.md`, each with a log line carrying the evidence `detect` gave after the stage; the decisions taken alone are listed both in the conversation and in the file; nothing was decided outside `--auto-decision`; no question was asked by this skill. What the run did not reach is named with the command that would close it.
67
+ The run's last count line, printed and written to `docs/AUTO.md`:
68
+ `stages done X/Y · decisions taken alone N · waiting for the owner K · flags: <the arguments>`
69
+ ▶ Next — /clear, then ll-resume
70
+
71
+ ## References
72
+
73
+ - `references/stages.md` — the detection table, the command per stage and the flag table; step 0.
74
+ - `references/run.md` — the `docs/AUTO.md` template, the WAITING handling and the end-of-run block; steps 1, 2 and 3.