@walwal-harness/cli 2.5.0 → 3.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/assets/templates/AGENTS.md.template +31 -0
- package/assets/templates/HARNESS.md +180 -6
- package/assets/templates/config.json +182 -10
- package/assets/templates/evaluation-functional.md.template +119 -0
- package/assets/templates/evaluation-visual.md.template +151 -0
- package/assets/templates/memory.md +20 -0
- package/assets/templates/progress.json.template +6 -0
- package/package.json +5 -2
- package/scripts/harness-next.sh +254 -1
- package/scripts/harness-session-start.sh +15 -12
- package/scripts/harness-statusline.sh +93 -0
- package/scripts/harness-user-prompt-submit.sh +31 -72
- package/scripts/lib/harness-guardrail.sh +143 -0
|
@@ -0,0 +1,151 @@
|
|
|
1
|
+
# Evaluation — Visual (Sprint {{SPRINT}})
|
|
2
|
+
|
|
3
|
+
> Generated by: evaluator-visual
|
|
4
|
+
> Date: {{DATE}}
|
|
5
|
+
> Verdict: **PENDING**
|
|
6
|
+
|
|
7
|
+
## 1. Scoring Rubric
|
|
8
|
+
|
|
9
|
+
> **Threshold: 2.80 / 3.00 = PASS**
|
|
10
|
+
> Score 2.79 이하는 이유를 불문하고 FAIL.
|
|
11
|
+
> 모든 항목에 Evidence(스크린샷 경로 또는 DOM 스냅샷)를 반드시 기입. Evidence 없는 Score는 0점 처리.
|
|
12
|
+
|
|
13
|
+
| # | Criterion | Weight | Score (0-3) | Evidence |
|
|
14
|
+
|---|-----------|--------|-------------|----------|
|
|
15
|
+
| V1 | 레이아웃 정확성 — 모든 페이지가 디자인 의도대로 배치, 겹침/넘침/잘림 없음 | 20% | | |
|
|
16
|
+
| V2 | 반응형 적합성 — 최소 3개 뷰포트(mobile 375px, tablet 768px, desktop 1280px)에서 정상 렌더링 | 20% | | |
|
|
17
|
+
| V3 | 접근성 (a11y) — WCAG 2.1 AA 준수, axe-core 위반 0건, 키보드 네비게이션 가능, ARIA 라벨 존재 | 20% | | |
|
|
18
|
+
| V4 | 시각적 일관성 — 색상/타이포/간격이 디자인 시스템과 일치, AI슬롭 패턴(그라데이션 남용, 무의미한 아이콘, 과도한 그림자) 없음 | 20% | | |
|
|
19
|
+
| V5 | 인터랙션 피드백 — 로딩 상태, 에러 상태, 빈 상태, 호버/포커스 상태 모두 구현됨 | 20% | | |
|
|
20
|
+
|
|
21
|
+
### Scoring Guide (엄격 적용)
|
|
22
|
+
|
|
23
|
+
| Score | 의미 | 기준 |
|
|
24
|
+
|-------|------|------|
|
|
25
|
+
| 3 | 완벽 | 모든 뷰포트, 모든 상태에서 결함 없음. 프로덕션 출시 가능 수준 |
|
|
26
|
+
| 2 | 충족 | 핵심 화면 정상이나 경미한 시각 이슈 1-2건 (1px 오차, 미사용 상태 아이콘 누락 등) |
|
|
27
|
+
| 1 | 부분 충족 | 특정 뷰포트에서 레이아웃 깨짐, 또는 접근성 위반 3건 이상 |
|
|
28
|
+
| 0 | 미충족 | 렌더링 실패, 빈 화면, 또는 핵심 UI 요소 누락 |
|
|
29
|
+
|
|
30
|
+
### Anti-Rubber-Stamping Check
|
|
31
|
+
|
|
32
|
+
Score를 기입하기 전에 반드시 아래 질문에 답하라:
|
|
33
|
+
- [ ] "이 페이지를 실제로 3개 뷰포트에서 스크린샷을 찍어 확인했는가?" (추측 = 0점)
|
|
34
|
+
- [ ] "axe-core 또는 접근성 검사를 실제로 실행했는가?" (미실행 = V3 최대 1점)
|
|
35
|
+
- [ ] "이전 Sprint 스크린샷과 비교하여 의도치 않은 시각 변경이 없는지 확인했는가?"
|
|
36
|
+
- [ ] "내가 이 UI를 직접 만들었다면, 이 수준으로 디자인 리뷰를 통과시키겠는가?" (NO = FAIL)
|
|
37
|
+
|
|
38
|
+
## 2. Screenshot Matrix
|
|
39
|
+
|
|
40
|
+
> 모든 페이지 x 3개 뷰포트의 스크린샷. 누락 셀이 있으면 해당 페이지 Score 산정 불가.
|
|
41
|
+
|
|
42
|
+
| Page | Mobile (375px) | Tablet (768px) | Desktop (1280px) | Issues |
|
|
43
|
+
|------|---------------|----------------|-------------------|--------|
|
|
44
|
+
| | | | | |
|
|
45
|
+
|
|
46
|
+
**Screenshot Coverage**: 0/0 pages x 3 viewports = 0/0
|
|
47
|
+
|
|
48
|
+
## 3. Accessibility Audit
|
|
49
|
+
|
|
50
|
+
> axe-core 또는 Playwright accessibility snapshot 결과.
|
|
51
|
+
|
|
52
|
+
| Page | Violations | Critical | Serious | Moderate | Minor |
|
|
53
|
+
|------|-----------|----------|---------|----------|-------|
|
|
54
|
+
| | | | | | |
|
|
55
|
+
|
|
56
|
+
**a11y Summary**: 0 total violations across 0 pages
|
|
57
|
+
**Critical 또는 Serious 위반이 1건이라도 있으면 V3 = 0**
|
|
58
|
+
|
|
59
|
+
## 4. Interaction States Checklist
|
|
60
|
+
|
|
61
|
+
> 각 핵심 컴포넌트의 상태별 구현 확인.
|
|
62
|
+
|
|
63
|
+
| Component | Default | Loading | Error | Empty | Hover | Focus | Disabled |
|
|
64
|
+
|-----------|---------|---------|-------|-------|-------|-------|----------|
|
|
65
|
+
| | | | | | | | |
|
|
66
|
+
|
|
67
|
+
**상태 커버리지**: 0/0 components x states
|
|
68
|
+
|
|
69
|
+
## 5. AI Slop Detection
|
|
70
|
+
|
|
71
|
+
> AI 생성 UI의 흔한 저품질 패턴 감지.
|
|
72
|
+
|
|
73
|
+
| Pattern | Detected? | Location | Severity |
|
|
74
|
+
|---------|-----------|----------|----------|
|
|
75
|
+
| 무의미한 그라데이션 | | | |
|
|
76
|
+
| 과도한 그림자/블러 | | | |
|
|
77
|
+
| 장식용 아이콘 남발 | | | |
|
|
78
|
+
| 불필요한 애니메이션 | | | |
|
|
79
|
+
| 일관성 없는 간격/정렬 | | | |
|
|
80
|
+
| placeholder 텍스트 미교체 | | | |
|
|
81
|
+
| 깨진 이미지/아이콘 | | | |
|
|
82
|
+
|
|
83
|
+
**AI Slop이 2건 이상 감지되면 V4 = 최대 1점**
|
|
84
|
+
|
|
85
|
+
## 6. Regression Visual Diff
|
|
86
|
+
|
|
87
|
+
> 이전 Sprint 스크린샷과 현재 비교. 의도치 않은 시각 변경 감지.
|
|
88
|
+
> Sprint 1이면 "N/A - 첫 스프린트"로 기입.
|
|
89
|
+
|
|
90
|
+
| Page | Viewport | Change Detected | Intentional? | Notes |
|
|
91
|
+
|------|----------|----------------|-------------|-------|
|
|
92
|
+
| | | | | |
|
|
93
|
+
|
|
94
|
+
**의도치 않은 시각 변경이 1건이라도 있으면 Verdict = FAIL**
|
|
95
|
+
|
|
96
|
+
## 7. Cross-Validation with Functional
|
|
97
|
+
|
|
98
|
+
> evaluation-functional.md의 Cross-Validation Data와 교차 검증.
|
|
99
|
+
|
|
100
|
+
| Check | Functional Result | Visual Result | Consistent? |
|
|
101
|
+
|-------|------------------|---------------|-------------|
|
|
102
|
+
| 페이지 존재 여부 | endpoints_tested → pages | 실제 렌더링 확인 | |
|
|
103
|
+
| 폼 제출 결과 | API 응답 status | UI 피드백 메시지 | |
|
|
104
|
+
| 에러 표시 | 에러 응답 코드 | 에러 UI 렌더링 | |
|
|
105
|
+
|
|
106
|
+
**불일치가 1건이라도 있으면 관련 Feature의 해당 항목 재검증 필요 → CONDITIONAL FAIL**
|
|
107
|
+
|
|
108
|
+
```json
|
|
109
|
+
{
|
|
110
|
+
"evaluator": "visual",
|
|
111
|
+
"sprint": {{SPRINT}},
|
|
112
|
+
"verdict": "PENDING",
|
|
113
|
+
"total_score": 0.00,
|
|
114
|
+
"threshold": 2.80,
|
|
115
|
+
"criteria_scores": {
|
|
116
|
+
"V1_layout": 0,
|
|
117
|
+
"V2_responsive": 0,
|
|
118
|
+
"V3_accessibility": 0,
|
|
119
|
+
"V4_consistency": 0,
|
|
120
|
+
"V5_interaction_states": 0
|
|
121
|
+
},
|
|
122
|
+
"screenshot_coverage": 0.0,
|
|
123
|
+
"a11y_violations": { "critical": 0, "serious": 0, "moderate": 0, "minor": 0 },
|
|
124
|
+
"ai_slop_count": 0,
|
|
125
|
+
"regression_visual_changes": 0,
|
|
126
|
+
"cross_validation_inconsistencies": 0,
|
|
127
|
+
"pages_tested": []
|
|
128
|
+
}
|
|
129
|
+
```
|
|
130
|
+
|
|
131
|
+
## 8. Final Verdict
|
|
132
|
+
|
|
133
|
+
| Metric | Value |
|
|
134
|
+
|--------|-------|
|
|
135
|
+
| Weighted Score | 0.00 / 3.00 |
|
|
136
|
+
| Threshold | 2.80 |
|
|
137
|
+
| Screenshot Coverage | 0% |
|
|
138
|
+
| a11y Critical+Serious | 0 |
|
|
139
|
+
| AI Slop Count | 0 |
|
|
140
|
+
| Regression Visual Changes | 0 |
|
|
141
|
+
| Cross-Validation Inconsistencies | 0 |
|
|
142
|
+
| **Verdict** | **PENDING** |
|
|
143
|
+
|
|
144
|
+
### Verdict Rules (위반 불가)
|
|
145
|
+
1. Weighted Score < 2.80 → **FAIL**
|
|
146
|
+
2. Screenshot 누락 페이지 존재 → 해당 페이지 관련 항목 Score = 0으로 재계산
|
|
147
|
+
3. a11y Critical 또는 Serious > 0 → V3 = 0으로 재계산
|
|
148
|
+
4. AI Slop >= 2 → V4 = 최대 1점으로 재계산
|
|
149
|
+
5. Regression 의도치 않은 변경 > 0 → **FAIL** (점수 무관)
|
|
150
|
+
6. Cross-Validation 불일치 > 0 → **CONDITIONAL FAIL** (해당 Feature 재검증)
|
|
151
|
+
7. Evidence 누락 항목 존재 → 해당 항목 Score = 0으로 재계산
|
|
@@ -3,4 +3,24 @@
|
|
|
3
3
|
> Dispatcher가 관리. **모든 에이전트**는 세션 시작 시 이 파일을 읽고 학습된 규칙을 따릅니다.
|
|
4
4
|
> gotchas(에이전트별 실수 기록)와 달리, 이 파일은 **프로젝트 전체에 적용되는 구조적 교훈**을 담습니다.
|
|
5
5
|
|
|
6
|
+
## 항목 형식
|
|
7
|
+
|
|
8
|
+
```markdown
|
|
9
|
+
### [M-NNN] 간결한 제목
|
|
10
|
+
- **Date**: YYYY-MM-DD
|
|
11
|
+
- **Status**: unverified | verified
|
|
12
|
+
- **TTL**: YYYY-MM-DD (기본 +60일. 만료 후 Planner가 리뷰)
|
|
13
|
+
- **Lesson**: 교훈 내용
|
|
14
|
+
- **Context**: 발견 배경
|
|
15
|
+
- **Applies to**: 적용 대상 에이전트/상황
|
|
16
|
+
```
|
|
17
|
+
|
|
18
|
+
## 오염 방어 규칙
|
|
19
|
+
|
|
20
|
+
- 새 항목은 반드시 `unverified` 상태로 시작
|
|
21
|
+
- Planner가 스프린트 리뷰 시 유효성 확인 후 `verified` 승격
|
|
22
|
+
- TTL 만료 항목은 Planner가 리뷰: 갱신 또는 삭제
|
|
23
|
+
- 환각(hallucination) 의심 항목: 코드/git 이력으로 검증 불가하면 즉시 삭제
|
|
24
|
+
- 항목이 15개 초과 시 가장 오래된 unverified 항목부터 정리
|
|
25
|
+
|
|
6
26
|
<!-- 항목이 추가되면 아래에 기록됩니다 -->
|
|
@@ -17,6 +17,12 @@
|
|
|
17
17
|
"message": null,
|
|
18
18
|
"retry_target": null
|
|
19
19
|
},
|
|
20
|
+
"artifacts": {
|
|
21
|
+
"plan.md": { "status": "pending", "updated_by": null, "updated_at": null },
|
|
22
|
+
"feature-list.json": { "status": "pending", "updated_by": null, "updated_at": null },
|
|
23
|
+
"api-contract.json": { "status": "pending", "updated_by": null, "updated_at": null },
|
|
24
|
+
"sprint-contract.md": { "status": "pending", "updated_by": null, "updated_at": null }
|
|
25
|
+
},
|
|
20
26
|
"updated_at": "{{DATE}}",
|
|
21
27
|
"history": []
|
|
22
28
|
}
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@walwal-harness/cli",
|
|
3
|
-
"version": "2.
|
|
3
|
+
"version": "3.2.0",
|
|
4
4
|
"description": "Production harness for AI agent engineering — Planner, Generator(BE/FE), Evaluator(Func/Visual), optional Brainstormer (requirements refinement). Supports React and Flutter FE stacks.",
|
|
5
5
|
"bin": {
|
|
6
6
|
"walwal-harness": "bin/init.js"
|
|
@@ -34,5 +34,8 @@
|
|
|
34
34
|
"scripts/",
|
|
35
35
|
"assets/",
|
|
36
36
|
"gotchas/"
|
|
37
|
-
]
|
|
37
|
+
],
|
|
38
|
+
"dependencies": {
|
|
39
|
+
"@walwal-harness/cli": "^3.2.0"
|
|
40
|
+
}
|
|
38
41
|
}
|
package/scripts/harness-next.sh
CHANGED
|
@@ -6,6 +6,7 @@ set -e
|
|
|
6
6
|
|
|
7
7
|
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
|
8
8
|
source "$SCRIPT_DIR/lib/harness-render-progress.sh"
|
|
9
|
+
source "$SCRIPT_DIR/lib/harness-guardrail.sh"
|
|
9
10
|
|
|
10
11
|
# ─────────────────────────────────────────
|
|
11
12
|
# Resolve project root
|
|
@@ -71,6 +72,90 @@ substitute_fe_stack() {
|
|
|
71
72
|
echo "$sub"
|
|
72
73
|
}
|
|
73
74
|
|
|
75
|
+
# ─────────────────────────────────────────
|
|
76
|
+
# Pre-Eval Gate — deterministic checks before Evaluator
|
|
77
|
+
# Generator가 완료되고 다음이 Evaluator일 때, lint/type/test를 먼저 실행.
|
|
78
|
+
# 실패 시 Evaluator를 건너뛰고 Generator로 리라우팅.
|
|
79
|
+
# ─────────────────────────────────────────
|
|
80
|
+
run_pre_eval_gate() {
|
|
81
|
+
local next="$1"
|
|
82
|
+
local gate_enabled
|
|
83
|
+
gate_enabled=$(jq -r '.flow.pre_eval_gate.enabled // false' "$CONFIG" 2>/dev/null)
|
|
84
|
+
if [ "$gate_enabled" != "true" ]; then return 0; fi
|
|
85
|
+
|
|
86
|
+
# Evaluator 에이전트인지 확인
|
|
87
|
+
case "$next" in
|
|
88
|
+
evaluator-*) ;;
|
|
89
|
+
*) return 0 ;;
|
|
90
|
+
esac
|
|
91
|
+
|
|
92
|
+
local timeout
|
|
93
|
+
timeout=$(jq -r '.flow.pre_eval_gate.timeout_seconds // 120' "$CONFIG" 2>/dev/null)
|
|
94
|
+
|
|
95
|
+
# 실패 위치 결정 (backend or frontend)
|
|
96
|
+
local location="backend"
|
|
97
|
+
local checks_key="backend_checks"
|
|
98
|
+
if [ "$current_agent" = "generator-frontend" ] || [ "$current_agent" = "generator-frontend-flutter" ]; then
|
|
99
|
+
location="frontend"
|
|
100
|
+
checks_key="frontend_checks"
|
|
101
|
+
fi
|
|
102
|
+
|
|
103
|
+
local -a checks
|
|
104
|
+
mapfile -t checks < <(jq -r ".flow.pre_eval_gate.${checks_key}[]" "$CONFIG" 2>/dev/null)
|
|
105
|
+
|
|
106
|
+
if [ ${#checks[@]} -eq 0 ]; then return 0; fi
|
|
107
|
+
|
|
108
|
+
echo ""
|
|
109
|
+
echo " ── Pre-Eval Gate ──────────────────────"
|
|
110
|
+
local all_pass=true
|
|
111
|
+
local fail_log=""
|
|
112
|
+
|
|
113
|
+
for cmd in "${checks[@]}"; do
|
|
114
|
+
printf " %-40s " "$cmd"
|
|
115
|
+
local output
|
|
116
|
+
if output=$(cd "$PROJECT_ROOT" && timeout "${timeout}s" bash -c "$cmd" 2>&1); then
|
|
117
|
+
echo "✓"
|
|
118
|
+
else
|
|
119
|
+
echo "✗"
|
|
120
|
+
all_pass=false
|
|
121
|
+
fail_log+="[FAIL] $cmd"$'\n'"$output"$'\n\n'
|
|
122
|
+
fi
|
|
123
|
+
done
|
|
124
|
+
|
|
125
|
+
if [ "$all_pass" = true ]; then
|
|
126
|
+
echo " Gate: PASS — proceeding to $next"
|
|
127
|
+
echo ""
|
|
128
|
+
return 0
|
|
129
|
+
else
|
|
130
|
+
echo ""
|
|
131
|
+
echo " Gate: FAIL — rerouting to $current_agent"
|
|
132
|
+
echo ""
|
|
133
|
+
|
|
134
|
+
# progress.json 업데이트: 실패 기록 + Generator로 리라우팅
|
|
135
|
+
local new_retry=$((retry_count + 1))
|
|
136
|
+
local fail_summary
|
|
137
|
+
fail_summary=$(echo "$fail_log" | head -20)
|
|
138
|
+
|
|
139
|
+
jq --arg agent "$current_agent" \
|
|
140
|
+
--arg loc "$location" \
|
|
141
|
+
--arg msg "Pre-eval gate failed: $fail_summary" \
|
|
142
|
+
--arg target "$current_agent" \
|
|
143
|
+
--argjson retry "$new_retry" \
|
|
144
|
+
'.sprint.status = "failed" |
|
|
145
|
+
.sprint.retry_count = $retry |
|
|
146
|
+
.agent_status = "failed" |
|
|
147
|
+
.next_agent = $target |
|
|
148
|
+
.failure.agent = $agent |
|
|
149
|
+
.failure.location = $loc |
|
|
150
|
+
.failure.message = $msg |
|
|
151
|
+
.failure.retry_target = $target' "$PROGRESS" > "${PROGRESS}.tmp" && mv "${PROGRESS}.tmp" "$PROGRESS"
|
|
152
|
+
|
|
153
|
+
# next_agent를 Generator로 덮어쓰기
|
|
154
|
+
next_agent="$current_agent"
|
|
155
|
+
return 1
|
|
156
|
+
fi
|
|
157
|
+
}
|
|
158
|
+
|
|
74
159
|
# ─────────────────────────────────────────
|
|
75
160
|
# Determine next agent
|
|
76
161
|
# ─────────────────────────────────────────
|
|
@@ -138,6 +223,71 @@ if [ "$agent_status" = "completed" ] && [ "$next_agent" = "null" ]; then
|
|
|
138
223
|
next_agent=$(compute_next_agent "$current_agent" "$agent_status")
|
|
139
224
|
fi
|
|
140
225
|
|
|
226
|
+
# ─────────────────────────────────────────
|
|
227
|
+
# Artifact Prerequisites — 선행 아티팩트 상태 검증
|
|
228
|
+
# ─────────────────────────────────────────
|
|
229
|
+
verify_artifact_prerequisites() {
|
|
230
|
+
local target_agent="$1"
|
|
231
|
+
local states_order='["pending","draft","reviewed","approved"]'
|
|
232
|
+
|
|
233
|
+
# 에이전트의 prerequisites 가져오기
|
|
234
|
+
local prereqs
|
|
235
|
+
prereqs=$(jq -r ".artifacts.prerequisites[\"${target_agent}\"] // empty" "$CONFIG" 2>/dev/null)
|
|
236
|
+
if [ -z "$prereqs" ] || [ "$prereqs" = "null" ]; then return 0; fi
|
|
237
|
+
|
|
238
|
+
local all_met=true
|
|
239
|
+
echo ""
|
|
240
|
+
echo " ── Artifact Prerequisites ─────────────"
|
|
241
|
+
|
|
242
|
+
# prereqs의 각 키(artifact명)를 순회
|
|
243
|
+
while IFS='=' read -r artifact required_status; do
|
|
244
|
+
artifact=$(echo "$artifact" | tr -d '"' | tr -d ' ')
|
|
245
|
+
required_status=$(echo "$required_status" | tr -d '"' | tr -d ' ')
|
|
246
|
+
if [ -z "$artifact" ]; then continue; fi
|
|
247
|
+
|
|
248
|
+
local current_status
|
|
249
|
+
current_status=$(jq -r ".artifacts[\"${artifact}\"].status // \"pending\"" "$PROGRESS" 2>/dev/null)
|
|
250
|
+
|
|
251
|
+
# 상태 순서 비교
|
|
252
|
+
local required_idx current_idx
|
|
253
|
+
required_idx=$(echo "$states_order" | jq "index(\"$required_status\") // 0")
|
|
254
|
+
current_idx=$(echo "$states_order" | jq "index(\"$current_status\") // 0")
|
|
255
|
+
|
|
256
|
+
if [ "$current_idx" -ge "$required_idx" ]; then
|
|
257
|
+
printf " ✓ %-25s %s (required: %s)\n" "$artifact" "$current_status" "$required_status"
|
|
258
|
+
else
|
|
259
|
+
printf " ✗ %-25s %s (required: %s)\n" "$artifact" "$current_status" "$required_status"
|
|
260
|
+
all_met=false
|
|
261
|
+
fi
|
|
262
|
+
done < <(jq -r ".artifacts.prerequisites[\"${target_agent}\"] | to_entries[] | \"\(.key)=\(.value)\"" "$CONFIG" 2>/dev/null)
|
|
263
|
+
|
|
264
|
+
echo ""
|
|
265
|
+
|
|
266
|
+
if [ "$all_met" = true ]; then
|
|
267
|
+
echo " Prerequisites: PASS"
|
|
268
|
+
else
|
|
269
|
+
echo " Prerequisites: FAIL — preceding agent must complete artifacts first"
|
|
270
|
+
fi
|
|
271
|
+
echo ""
|
|
272
|
+
|
|
273
|
+
[ "$all_met" = true ] && return 0 || return 1
|
|
274
|
+
}
|
|
275
|
+
|
|
276
|
+
# ─────────────────────────────────────────
|
|
277
|
+
# Runtime Guardrail — 파일 소유권 검증
|
|
278
|
+
# ─────────────────────────────────────────
|
|
279
|
+
if [ "$agent_status" = "completed" ]; then
|
|
280
|
+
verify_file_ownership "$PROJECT_ROOT" || true
|
|
281
|
+
fi
|
|
282
|
+
|
|
283
|
+
# ─────────────────────────────────────────
|
|
284
|
+
# Run Pre-Eval Gate (if applicable)
|
|
285
|
+
# ─────────────────────────────────────────
|
|
286
|
+
if [ "$agent_status" = "completed" ]; then
|
|
287
|
+
run_pre_eval_gate "$next_agent" || true
|
|
288
|
+
verify_artifact_prerequisites "$next_agent" || true
|
|
289
|
+
fi
|
|
290
|
+
|
|
141
291
|
# ─────────────────────────────────────────
|
|
142
292
|
# Render progress
|
|
143
293
|
# ─────────────────────────────────────────
|
|
@@ -149,6 +299,18 @@ echo ""
|
|
|
149
299
|
# Generate next-prompt.txt
|
|
150
300
|
# ─────────────────────────────────────────
|
|
151
301
|
if [ "$next_agent" != "null" ] && [ "$next_agent" != "archive" ] && [ "$agent_status" != "blocked" ]; then
|
|
302
|
+
# ── Escalation check: 3회 실패 시 Planner에게 scope 축소 요청 ──
|
|
303
|
+
escalate_after=$(jq -r '.flow.escalate_to_planner_after // 3' "$CONFIG" 2>/dev/null || echo 3)
|
|
304
|
+
if [ "$retry_count" -ge "$escalate_after" ] && [ "$sprint_status" = "failed" ] && [ "$next_agent" != "planner" ]; then
|
|
305
|
+
echo " ⚠ Escalation: ${retry_count}회 실패 — Planner에게 scope 축소/접근 변경 요청"
|
|
306
|
+
next_agent="planner"
|
|
307
|
+
# progress.json에 에스컬레이션 기록
|
|
308
|
+
jq --arg msg "Escalated after ${retry_count} failures. Planner must review scope or approach." \
|
|
309
|
+
'.next_agent = "planner" |
|
|
310
|
+
.failure.message = $msg |
|
|
311
|
+
.failure.retry_target = "planner"' "$PROGRESS" > "${PROGRESS}.tmp" && mv "${PROGRESS}.tmp" "$PROGRESS"
|
|
312
|
+
fi
|
|
313
|
+
|
|
152
314
|
# Build prompt
|
|
153
315
|
prompt="/harness-${next_agent} 를 실행하세요."
|
|
154
316
|
|
|
@@ -162,14 +324,105 @@ if [ "$next_agent" != "null" ] && [ "$next_agent" != "archive" ] && [ "$agent_st
|
|
|
162
324
|
|
|
163
325
|
prompt+=$'\n'".harness/progress.json을 읽고 현재 상태를 확인하세요."
|
|
164
326
|
|
|
165
|
-
# Add failure context if retrying
|
|
327
|
+
# Add failure context if retrying (include previous failure summary)
|
|
166
328
|
failure_msg=$(jq -r '.failure.message // empty' "$PROGRESS")
|
|
167
329
|
if [ -n "$failure_msg" ] && [ "$failure_msg" != "null" ]; then
|
|
168
330
|
prompt+=$'\n\n'"이전 실패 사유: ${failure_msg}"
|
|
331
|
+
prompt+=$'\n'"같은 접근을 반복하지 말고, 실패 원인을 분석한 후 다른 전략으로 시도하세요."
|
|
169
332
|
fi
|
|
170
333
|
|
|
171
334
|
echo "$prompt" > "$NEXT_PROMPT"
|
|
172
335
|
|
|
336
|
+
# ── Generate structured handoff.json ──
|
|
337
|
+
HANDOFF="$PROJECT_ROOT/.harness/handoff.json"
|
|
338
|
+
FEATURE_LIST="$PROJECT_ROOT/.harness/actions/feature-list.json"
|
|
339
|
+
|
|
340
|
+
# Collect available artifacts
|
|
341
|
+
local -a artifacts_ready=()
|
|
342
|
+
for f in plan.md feature-list.json api-contract.json sprint-contract.md evaluation-functional.md evaluation-visual.md; do
|
|
343
|
+
if [ -f "$PROJECT_ROOT/.harness/actions/$f" ]; then
|
|
344
|
+
artifacts_ready+=("$f")
|
|
345
|
+
fi
|
|
346
|
+
done
|
|
347
|
+
local artifacts_json
|
|
348
|
+
artifacts_json=$(printf '%s\n' "${artifacts_ready[@]}" | jq -R . | jq -s .)
|
|
349
|
+
|
|
350
|
+
# Collect focus features (incomplete ones)
|
|
351
|
+
local focus_features="[]"
|
|
352
|
+
if [ -f "$FEATURE_LIST" ]; then
|
|
353
|
+
focus_features=$(jq '[.features[]? | select(.passes == null or (.passes | length) == 0 or ((.passes // []) | map(select(. == "evaluator-functional")) | length == 0)) | .id] | .[0:5]' "$FEATURE_LIST" 2>/dev/null || echo "[]")
|
|
354
|
+
fi
|
|
355
|
+
|
|
356
|
+
# ── Regression data: collect previous sprint's passed AC from archive ──
|
|
357
|
+
local regression_source="null"
|
|
358
|
+
local prev_sprint=$((sprint_num - 1))
|
|
359
|
+
local prev_archive="$PROJECT_ROOT/.harness/archive/sprint-$(printf '%03d' $prev_sprint)"
|
|
360
|
+
if [ "$prev_sprint" -ge 1 ] && [ -d "$prev_archive" ]; then
|
|
361
|
+
if [ -f "$prev_archive/feature-list.json" ]; then
|
|
362
|
+
regression_source=$(jq '{
|
|
363
|
+
sprint: '"$prev_sprint"',
|
|
364
|
+
passed_features: [.features[]? | select((.passes // []) | map(select(. == "evaluator-functional")) | length > 0) | {id, name, acceptance_criteria}],
|
|
365
|
+
archive_path: "'"$prev_archive"'"
|
|
366
|
+
}' "$prev_archive/feature-list.json" 2>/dev/null || echo "null")
|
|
367
|
+
fi
|
|
368
|
+
fi
|
|
369
|
+
|
|
370
|
+
# ── Eval-specific scoring config ──
|
|
371
|
+
local eval_config="null"
|
|
372
|
+
local cross_validation_data="null"
|
|
373
|
+
case "$next_agent" in
|
|
374
|
+
evaluator-*)
|
|
375
|
+
eval_config=$(jq '{
|
|
376
|
+
pass_threshold: .evaluation.scoring.pass_threshold,
|
|
377
|
+
scale: .evaluation.scoring.scale,
|
|
378
|
+
verdict_rules: .evaluation.scoring.verdict_rules,
|
|
379
|
+
regression_enabled: .evaluation.regression.enabled,
|
|
380
|
+
cross_validation_enabled: .evaluation.cross_validation.enabled,
|
|
381
|
+
adversarial_rules: .agents["'"$next_agent"'"].adversarial_rules.rules,
|
|
382
|
+
forbidden: .agents["'"$next_agent"'"].adversarial_rules.forbidden
|
|
383
|
+
}' "$CONFIG" 2>/dev/null || echo "null")
|
|
384
|
+
;;
|
|
385
|
+
esac
|
|
386
|
+
|
|
387
|
+
# ── Cross-Validation: evaluator-visual이면 functional 결과 파싱 ──
|
|
388
|
+
if [ "$next_agent" = "evaluator-visual" ]; then
|
|
389
|
+
local func_eval="$PROJECT_ROOT/.harness/actions/evaluation-functional.md"
|
|
390
|
+
if [ -f "$func_eval" ]; then
|
|
391
|
+
# evaluation-functional.md 내 JSON 코드블록에서 Cross-Validation Data 추출
|
|
392
|
+
cross_validation_data=$(sed -n '/```json/,/```/p' "$func_eval" | tail -n +2 | head -n -1 | jq 'select(.evaluator == "functional")' 2>/dev/null || echo "null")
|
|
393
|
+
fi
|
|
394
|
+
fi
|
|
395
|
+
|
|
396
|
+
# Build handoff.json
|
|
397
|
+
jq -n \
|
|
398
|
+
--arg from "${current_agent:-dispatcher}" \
|
|
399
|
+
--arg to "$next_agent" \
|
|
400
|
+
--argjson sprint "$sprint_num" \
|
|
401
|
+
--argjson retry "$retry_count" \
|
|
402
|
+
--arg status "$sprint_status" \
|
|
403
|
+
--arg failure_msg "${failure_msg:-}" \
|
|
404
|
+
--argjson artifacts "$artifacts_json" \
|
|
405
|
+
--argjson focus "$focus_features" \
|
|
406
|
+
--argjson regression "$regression_source" \
|
|
407
|
+
--argjson eval_config "$eval_config" \
|
|
408
|
+
--argjson cross_val "$cross_validation_data" \
|
|
409
|
+
--arg timestamp "$(date -u +%Y-%m-%dT%H:%M:%SZ)" \
|
|
410
|
+
'{
|
|
411
|
+
from: $from,
|
|
412
|
+
to: $to,
|
|
413
|
+
sprint: $sprint,
|
|
414
|
+
retry_count: $retry,
|
|
415
|
+
sprint_status: $status,
|
|
416
|
+
failure_context: (if $failure_msg != "" then $failure_msg else null end),
|
|
417
|
+
artifacts_ready: $artifacts,
|
|
418
|
+
focus_features: $focus,
|
|
419
|
+
regression: $regression,
|
|
420
|
+
eval_config: $eval_config,
|
|
421
|
+
cross_validation_from_functional: $cross_val,
|
|
422
|
+
warnings: [],
|
|
423
|
+
timestamp: $timestamp
|
|
424
|
+
}' > "$HANDOFF"
|
|
425
|
+
|
|
173
426
|
elif [ "$next_agent" = "archive" ]; then
|
|
174
427
|
cat > "$NEXT_PROMPT" <<'PROMPT'
|
|
175
428
|
Sprint 문서를 아카이브 하세요.
|
|
@@ -1,31 +1,34 @@
|
|
|
1
1
|
#!/bin/bash
|
|
2
|
-
# harness-session-start.sh — SessionStart 훅
|
|
3
|
-
#
|
|
4
|
-
# .claude/settings.json의 SessionStart 훅으로 등록된다.
|
|
2
|
+
# harness-session-start.sh — SessionStart 훅 (compact)
|
|
3
|
+
# statusline이 상시 상태를 표시하므로, 여기서는 핵심 안내만 출력.
|
|
5
4
|
|
|
6
5
|
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
|
7
6
|
LIB="$SCRIPT_DIR/lib/harness-render-progress.sh"
|
|
8
7
|
|
|
9
|
-
# lib이 없으면 silent exit (훅이므로 에러 출력하지 않음)
|
|
10
8
|
if [ ! -f "$LIB" ]; then exit 0; fi
|
|
11
9
|
source "$LIB"
|
|
12
|
-
|
|
13
|
-
# jq 없으면 silent exit
|
|
14
10
|
command -v jq &>/dev/null || exit 0
|
|
15
11
|
|
|
16
|
-
# .harness/ 찾기
|
|
17
12
|
PROJECT_ROOT="$(resolve_harness_root "." 2>/dev/null)" || exit 0
|
|
18
|
-
|
|
19
13
|
PROGRESS="$PROJECT_ROOT/.harness/progress.json"
|
|
20
14
|
[ -f "$PROGRESS" ] || exit 0
|
|
21
15
|
|
|
22
|
-
# init 상태면 간단 안내만
|
|
23
16
|
sprint_status=$(jq -r '.sprint.status // "init"' "$PROGRESS" 2>/dev/null)
|
|
17
|
+
current_agent=$(jq -r '.current_agent // "none"' "$PROGRESS" 2>/dev/null)
|
|
18
|
+
next_agent=$(jq -r '.next_agent // "none"' "$PROGRESS" 2>/dev/null)
|
|
19
|
+
agent_status=$(jq -r '.agent_status // "pending"' "$PROGRESS" 2>/dev/null)
|
|
20
|
+
|
|
21
|
+
# init 상태: 간단 안내
|
|
24
22
|
if [ "$sprint_status" = "init" ]; then
|
|
25
23
|
echo "# Harness ready — say \"하네스 엔지니어링 시작\" or /harness-dispatcher"
|
|
26
24
|
exit 0
|
|
27
25
|
fi
|
|
28
26
|
|
|
29
|
-
#
|
|
30
|
-
|
|
31
|
-
|
|
27
|
+
# 활성 세션: 다음 액션만 안내 (상세 프로그래스는 statusline에서 상시 표시)
|
|
28
|
+
if [ "$agent_status" = "blocked" ]; then
|
|
29
|
+
echo "# Harness BLOCKED — user intervention required. Run: bash scripts/harness-next.sh"
|
|
30
|
+
elif [ "$next_agent" != "none" ] && [ "$next_agent" != "null" ]; then
|
|
31
|
+
echo "# Harness: next → /harness-${next_agent}"
|
|
32
|
+
elif [ "$current_agent" != "none" ] && [ "$current_agent" != "null" ]; then
|
|
33
|
+
echo "# Harness: ${current_agent} [${agent_status}]"
|
|
34
|
+
fi
|
|
@@ -0,0 +1,93 @@
|
|
|
1
|
+
#!/bin/bash
|
|
2
|
+
# harness-statusline.sh — Claude Code statusline hook
|
|
3
|
+
# 터미널 하단에 항상 고정되는 1줄 compact 상태 표시.
|
|
4
|
+
# stdin: Claude Code JSON payload (model, context_window, cost 등)
|
|
5
|
+
# stdout: 상태 문자열 (Claude Code가 터미널 하단에 렌더링)
|
|
6
|
+
|
|
7
|
+
# Read Claude Code session data from stdin
|
|
8
|
+
input=$(cat)
|
|
9
|
+
|
|
10
|
+
# Extract Claude Code built-in data
|
|
11
|
+
context_pct=$(echo "$input" | jq -r '.context_window.used_percentage // 0' 2>/dev/null | cut -d. -f1)
|
|
12
|
+
cost=$(echo "$input" | jq -r '.cost.total_cost_usd // 0' 2>/dev/null)
|
|
13
|
+
|
|
14
|
+
# Resolve project root (walk up from cwd)
|
|
15
|
+
CWD=$(echo "$input" | jq -r '.workspace.current_dir // empty' 2>/dev/null)
|
|
16
|
+
if [ -z "$CWD" ]; then CWD="$PWD"; fi
|
|
17
|
+
|
|
18
|
+
PROJECT_ROOT="$CWD"
|
|
19
|
+
while [ "$PROJECT_ROOT" != "/" ]; do
|
|
20
|
+
if [ -d "$PROJECT_ROOT/.harness" ]; then break; fi
|
|
21
|
+
PROJECT_ROOT="$(dirname "$PROJECT_ROOT")"
|
|
22
|
+
done
|
|
23
|
+
|
|
24
|
+
PROGRESS="$PROJECT_ROOT/.harness/progress.json"
|
|
25
|
+
FEATURE_LIST="$PROJECT_ROOT/.harness/actions/feature-list.json"
|
|
26
|
+
PIPELINE_JSON="$PROJECT_ROOT/.harness/actions/pipeline.json"
|
|
27
|
+
|
|
28
|
+
# No harness → minimal status
|
|
29
|
+
if [ ! -f "$PROGRESS" ]; then
|
|
30
|
+
echo "harness: not initialized | ctx ${context_pct}%"
|
|
31
|
+
exit 0
|
|
32
|
+
fi
|
|
33
|
+
|
|
34
|
+
# Read harness state
|
|
35
|
+
sprint_num=$(jq -r '.sprint.number // 0' "$PROGRESS" 2>/dev/null)
|
|
36
|
+
sprint_status=$(jq -r '.sprint.status // "init"' "$PROGRESS" 2>/dev/null)
|
|
37
|
+
pipeline=$(jq -r '.pipeline // "?"' "$PROGRESS" 2>/dev/null)
|
|
38
|
+
current_agent=$(jq -r '.current_agent // "none"' "$PROGRESS" 2>/dev/null)
|
|
39
|
+
agent_status=$(jq -r '.agent_status // "pending"' "$PROGRESS" 2>/dev/null)
|
|
40
|
+
next_agent=$(jq -r '.next_agent // "none"' "$PROGRESS" 2>/dev/null)
|
|
41
|
+
retry_count=$(jq -r '.sprint.retry_count // 0' "$PROGRESS" 2>/dev/null)
|
|
42
|
+
|
|
43
|
+
# Pipeline short name
|
|
44
|
+
case "$pipeline" in
|
|
45
|
+
FULLSTACK) pl="FULL" ;;
|
|
46
|
+
FE-ONLY) pl="FE" ;;
|
|
47
|
+
BE-ONLY) pl="BE" ;;
|
|
48
|
+
null|"?") pl="?" ;;
|
|
49
|
+
*) pl="$pipeline" ;;
|
|
50
|
+
esac
|
|
51
|
+
|
|
52
|
+
# Agent short name (strip harness prefix)
|
|
53
|
+
agent_short="${current_agent#generator-}"
|
|
54
|
+
agent_short="${agent_short#evaluator-}"
|
|
55
|
+
if [ "$current_agent" = "none" ] || [ "$current_agent" = "null" ]; then
|
|
56
|
+
agent_short="$next_agent"
|
|
57
|
+
if [ "$agent_short" = "none" ] || [ "$agent_short" = "null" ]; then
|
|
58
|
+
agent_short="idle"
|
|
59
|
+
fi
|
|
60
|
+
fi
|
|
61
|
+
|
|
62
|
+
# Feature progress
|
|
63
|
+
total_features=0
|
|
64
|
+
completed_features=0
|
|
65
|
+
if [ -f "$FEATURE_LIST" ]; then
|
|
66
|
+
total_features=$(jq '.features | length' "$FEATURE_LIST" 2>/dev/null || echo 0)
|
|
67
|
+
completed_features=$(jq '[.features[]? | select(
|
|
68
|
+
(.passes // []) | (
|
|
69
|
+
(map(select(. == "evaluator-functional")) | length > 0) and
|
|
70
|
+
(map(select(. == "evaluator-visual")) | length > 0)
|
|
71
|
+
)
|
|
72
|
+
)] | length' "$FEATURE_LIST" 2>/dev/null || echo 0)
|
|
73
|
+
fi
|
|
74
|
+
|
|
75
|
+
# Status indicator
|
|
76
|
+
status_icon=""
|
|
77
|
+
case "$agent_status" in
|
|
78
|
+
running) status_icon=">" ;;
|
|
79
|
+
completed) status_icon="v" ;;
|
|
80
|
+
failed) status_icon="x" ;;
|
|
81
|
+
blocked) status_icon="!" ;;
|
|
82
|
+
*) status_icon="-" ;;
|
|
83
|
+
esac
|
|
84
|
+
|
|
85
|
+
# Retry indicator
|
|
86
|
+
retry_str=""
|
|
87
|
+
if [ "$retry_count" -gt 0 ]; then
|
|
88
|
+
retry_str=" R${retry_count}"
|
|
89
|
+
fi
|
|
90
|
+
|
|
91
|
+
# Build compact status line
|
|
92
|
+
# Format: [S1] FULL | >backend | 2/5 feat | R0 | ctx 45% | $1.23
|
|
93
|
+
echo "[S${sprint_num}] ${pl} | ${status_icon}${agent_short}${retry_str} | ${completed_features}/${total_features} feat | ctx ${context_pct}% | \$${cost}"
|