@walwal-harness/cli 2.4.0 → 3.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,119 @@
1
+ # Evaluation — Functional (Sprint {{SPRINT}})
2
+
3
+ > Generated by: evaluator-functional
4
+ > Date: {{DATE}}
5
+ > Verdict: **PENDING**
6
+
7
+ ## 1. Scoring Rubric
8
+
9
+ > **Threshold: 2.80 / 3.00 = PASS**
10
+ > Score 2.79 이하는 이유를 불문하고 FAIL.
11
+ > 모든 항목에 Evidence(근거)를 반드시 기입. Evidence 없는 Score는 0점 처리.
12
+
13
+ | # | Criterion | Weight | Score (0-3) | Evidence |
14
+ |---|-----------|--------|-------------|----------|
15
+ | R1 | API Contract 준수 — 모든 엔드포인트가 api-contract.json과 일치 (method, path, request/response schema, status codes) | 25% | | |
16
+ | R2 | Acceptance Criteria 통과 — feature-list.json의 모든 AC 항목 실행 결과 | 25% | | |
17
+ | R3 | 부정 테스트 (Negative Testing) — 각 엔드포인트별 최소 2개 비정상 시나리오 (빈 입력, 잘못된 타입, 미인증, 권한 없음, 중복, 초과값) | 20% | | |
18
+ | R4 | E2E 시나리오 통과 — Playwright로 사용자 플로우 전체 재현 (회원가입→로그인→핵심기능→로그아웃 등) | 15% | | |
19
+ | R5 | 에러 핸들링 & 엣지케이스 — 서버 에러 응답 형식 일관성, DB 제약조건 위반 처리, 동시성 충돌 | 15% | | |
20
+
21
+ ### Scoring Guide (엄격 적용)
22
+
23
+ | Score | 의미 | 기준 |
24
+ |-------|------|------|
25
+ | 3 | 완벽 | 모든 케이스 통과, 추가 검증 불필요, 엣지케이스까지 커버 |
26
+ | 2 | 충족 | 핵심 기능 동작하나 경미한 이슈 1-2건 존재 (로그 누락, 응답 필드 오타 등) |
27
+ | 1 | 부분 충족 | 핵심 기능은 동작하나 명확한 수정 필요 (잘못된 status code, 누락된 validation 등) |
28
+ | 0 | 미충족 | 미구현, 서버 에러, 또는 contract 불일치 |
29
+
30
+ ### Anti-Rubber-Stamping Check
31
+
32
+ Score를 기입하기 전에 반드시 아래 질문에 답하라:
33
+ - [ ] "이 엔드포인트에 실제로 요청을 보내고 응답을 확인했는가?" (추측 = 0점)
34
+ - [ ] "정상 케이스뿐 아니라 비정상 케이스도 테스트했는가?" (정상만 = Score 최대 2)
35
+ - [ ] "api-contract.json과 실제 응답을 문자열 수준으로 비교했는가?"
36
+
37
+ ## 2. Acceptance Criteria Checklist
38
+
39
+ > feature-list.json의 acceptance_criteria를 기계적으로 실행한 결과.
40
+ > 각 AC에 대해 PASS/FAIL + 실제 결과를 기록.
41
+
42
+ | Feature | AC ID | Description | Type | Expected | Actual | Result |
43
+ |---------|-------|-------------|------|----------|--------|--------|
44
+ | | | | | | | |
45
+
46
+ **AC Summary**: 0/0 passed (0%)
47
+ **AC 전체 통과가 아니면 R2는 최대 Score 1**
48
+
49
+ ## 3. Negative Test Results
50
+
51
+ > 각 엔드포인트별 최소 2개 부정 테스트. 미실시 엔드포인트가 있으면 R3 = 0.
52
+
53
+ | Endpoint | Test Case | Expected | Actual | Result |
54
+ |----------|-----------|----------|--------|--------|
55
+ | | | | | |
56
+
57
+ ## 4. E2E Scenario Log
58
+
59
+ > Playwright 실행 로그. 스크린샷 경로 포함.
60
+
61
+ ```
62
+ (Playwright 실행 결과를 여기에 붙여넣기)
63
+ ```
64
+
65
+ ## 5. Regression Check
66
+
67
+ > 이전 Sprint에서 PASS된 기능 재검증 결과.
68
+ > Sprint 1이면 "N/A - 첫 스프린트"로 기입.
69
+
70
+ | Sprint | Feature | AC ID | Result | Notes |
71
+ |--------|---------|-------|--------|-------|
72
+ | | | | | |
73
+
74
+ **Regression Summary**: 0/0 passed
75
+ **회귀 실패가 1건이라도 있으면 전체 Verdict = FAIL (신규 기능 점수 무관)**
76
+
77
+ ## 6. Cross-Validation Data
78
+
79
+ > evaluator-visual이 참조할 수 있도록 기계 판독 가능한 결과 요약.
80
+ > 이 섹션은 반드시 아래 JSON 코드블록 형식을 유지할 것.
81
+
82
+ ```json
83
+ {
84
+ "evaluator": "functional",
85
+ "sprint": {{SPRINT}},
86
+ "verdict": "PENDING",
87
+ "total_score": 0.00,
88
+ "threshold": 2.80,
89
+ "criteria_scores": {
90
+ "R1_api_contract": 0,
91
+ "R2_acceptance_criteria": 0,
92
+ "R3_negative_testing": 0,
93
+ "R4_e2e_scenario": 0,
94
+ "R5_error_handling": 0
95
+ },
96
+ "ac_pass_rate": 0.0,
97
+ "negative_test_count": 0,
98
+ "regression_failures": 0,
99
+ "endpoints_tested": [],
100
+ "pages_tested": []
101
+ }
102
+ ```
103
+
104
+ ## 7. Final Verdict
105
+
106
+ | Metric | Value |
107
+ |--------|-------|
108
+ | Weighted Score | 0.00 / 3.00 |
109
+ | Threshold | 2.80 |
110
+ | AC Pass Rate | 0% |
111
+ | Regression Failures | 0 |
112
+ | **Verdict** | **PENDING** |
113
+
114
+ ### Verdict Rules (위반 불가)
115
+ 1. Weighted Score < 2.80 → **FAIL**
116
+ 2. AC Pass Rate < 100% → **FAIL** (부분 통과 불인정)
117
+ 3. Regression Failures > 0 → **FAIL** (신규 점수 무관)
118
+ 4. Evidence 누락 항목 존재 → 해당 항목 Score = 0으로 재계산
119
+ 5. 부정 테스트 미실시 엔드포인트 존재 → R3 = 0으로 재계산
@@ -0,0 +1,151 @@
1
+ # Evaluation — Visual (Sprint {{SPRINT}})
2
+
3
+ > Generated by: evaluator-visual
4
+ > Date: {{DATE}}
5
+ > Verdict: **PENDING**
6
+
7
+ ## 1. Scoring Rubric
8
+
9
+ > **Threshold: 2.80 / 3.00 = PASS**
10
+ > Score 2.79 이하는 이유를 불문하고 FAIL.
11
+ > 모든 항목에 Evidence(스크린샷 경로 또는 DOM 스냅샷)를 반드시 기입. Evidence 없는 Score는 0점 처리.
12
+
13
+ | # | Criterion | Weight | Score (0-3) | Evidence |
14
+ |---|-----------|--------|-------------|----------|
15
+ | V1 | 레이아웃 정확성 — 모든 페이지가 디자인 의도대로 배치, 겹침/넘침/잘림 없음 | 20% | | |
16
+ | V2 | 반응형 적합성 — 최소 3개 뷰포트(mobile 375px, tablet 768px, desktop 1280px)에서 정상 렌더링 | 20% | | |
17
+ | V3 | 접근성 (a11y) — WCAG 2.1 AA 준수, axe-core 위반 0건, 키보드 네비게이션 가능, ARIA 라벨 존재 | 20% | | |
18
+ | V4 | 시각적 일관성 — 색상/타이포/간격이 디자인 시스템과 일치, AI슬롭 패턴(그라데이션 남용, 무의미한 아이콘, 과도한 그림자) 없음 | 20% | | |
19
+ | V5 | 인터랙션 피드백 — 로딩 상태, 에러 상태, 빈 상태, 호버/포커스 상태 모두 구현됨 | 20% | | |
20
+
21
+ ### Scoring Guide (엄격 적용)
22
+
23
+ | Score | 의미 | 기준 |
24
+ |-------|------|------|
25
+ | 3 | 완벽 | 모든 뷰포트, 모든 상태에서 결함 없음. 프로덕션 출시 가능 수준 |
26
+ | 2 | 충족 | 핵심 화면 정상이나 경미한 시각 이슈 1-2건 (1px 오차, 미사용 상태 아이콘 누락 등) |
27
+ | 1 | 부분 충족 | 특정 뷰포트에서 레이아웃 깨짐, 또는 접근성 위반 3건 이상 |
28
+ | 0 | 미충족 | 렌더링 실패, 빈 화면, 또는 핵심 UI 요소 누락 |
29
+
30
+ ### Anti-Rubber-Stamping Check
31
+
32
+ Score를 기입하기 전에 반드시 아래 질문에 답하라:
33
+ - [ ] "이 페이지를 실제로 3개 뷰포트에서 스크린샷을 찍어 확인했는가?" (추측 = 0점)
34
+ - [ ] "axe-core 또는 접근성 검사를 실제로 실행했는가?" (미실행 = V3 최대 1점)
35
+ - [ ] "이전 Sprint 스크린샷과 비교하여 의도치 않은 시각 변경이 없는지 확인했는가?"
36
+ - [ ] "내가 이 UI를 직접 만들었다면, 이 수준으로 디자인 리뷰를 통과시키겠는가?" (NO = FAIL)
37
+
38
+ ## 2. Screenshot Matrix
39
+
40
+ > 모든 페이지 x 3개 뷰포트의 스크린샷. 누락 셀이 있으면 해당 페이지 Score 산정 불가.
41
+
42
+ | Page | Mobile (375px) | Tablet (768px) | Desktop (1280px) | Issues |
43
+ |------|---------------|----------------|-------------------|--------|
44
+ | | | | | |
45
+
46
+ **Screenshot Coverage**: 0/0 pages x 3 viewports = 0/0
47
+
48
+ ## 3. Accessibility Audit
49
+
50
+ > axe-core 또는 Playwright accessibility snapshot 결과.
51
+
52
+ | Page | Violations | Critical | Serious | Moderate | Minor |
53
+ |------|-----------|----------|---------|----------|-------|
54
+ | | | | | | |
55
+
56
+ **a11y Summary**: 0 total violations across 0 pages
57
+ **Critical 또는 Serious 위반이 1건이라도 있으면 V3 = 0**
58
+
59
+ ## 4. Interaction States Checklist
60
+
61
+ > 각 핵심 컴포넌트의 상태별 구현 확인.
62
+
63
+ | Component | Default | Loading | Error | Empty | Hover | Focus | Disabled |
64
+ |-----------|---------|---------|-------|-------|-------|-------|----------|
65
+ | | | | | | | | |
66
+
67
+ **상태 커버리지**: 0/0 components x states
68
+
69
+ ## 5. AI Slop Detection
70
+
71
+ > AI 생성 UI의 흔한 저품질 패턴 감지.
72
+
73
+ | Pattern | Detected? | Location | Severity |
74
+ |---------|-----------|----------|----------|
75
+ | 무의미한 그라데이션 | | | |
76
+ | 과도한 그림자/블러 | | | |
77
+ | 장식용 아이콘 남발 | | | |
78
+ | 불필요한 애니메이션 | | | |
79
+ | 일관성 없는 간격/정렬 | | | |
80
+ | placeholder 텍스트 미교체 | | | |
81
+ | 깨진 이미지/아이콘 | | | |
82
+
83
+ **AI Slop이 2건 이상 감지되면 V4 = 최대 1점**
84
+
85
+ ## 6. Regression Visual Diff
86
+
87
+ > 이전 Sprint 스크린샷과 현재 비교. 의도치 않은 시각 변경 감지.
88
+ > Sprint 1이면 "N/A - 첫 스프린트"로 기입.
89
+
90
+ | Page | Viewport | Change Detected | Intentional? | Notes |
91
+ |------|----------|----------------|-------------|-------|
92
+ | | | | | |
93
+
94
+ **의도치 않은 시각 변경이 1건이라도 있으면 Verdict = FAIL**
95
+
96
+ ## 7. Cross-Validation with Functional
97
+
98
+ > evaluation-functional.md의 Cross-Validation Data와 교차 검증.
99
+
100
+ | Check | Functional Result | Visual Result | Consistent? |
101
+ |-------|------------------|---------------|-------------|
102
+ | 페이지 존재 여부 | endpoints_tested → pages | 실제 렌더링 확인 | |
103
+ | 폼 제출 결과 | API 응답 status | UI 피드백 메시지 | |
104
+ | 에러 표시 | 에러 응답 코드 | 에러 UI 렌더링 | |
105
+
106
+ **불일치가 1건이라도 있으면 관련 Feature의 해당 항목 재검증 필요 → CONDITIONAL FAIL**
107
+
108
+ ```json
109
+ {
110
+ "evaluator": "visual",
111
+ "sprint": {{SPRINT}},
112
+ "verdict": "PENDING",
113
+ "total_score": 0.00,
114
+ "threshold": 2.80,
115
+ "criteria_scores": {
116
+ "V1_layout": 0,
117
+ "V2_responsive": 0,
118
+ "V3_accessibility": 0,
119
+ "V4_consistency": 0,
120
+ "V5_interaction_states": 0
121
+ },
122
+ "screenshot_coverage": 0.0,
123
+ "a11y_violations": { "critical": 0, "serious": 0, "moderate": 0, "minor": 0 },
124
+ "ai_slop_count": 0,
125
+ "regression_visual_changes": 0,
126
+ "cross_validation_inconsistencies": 0,
127
+ "pages_tested": []
128
+ }
129
+ ```
130
+
131
+ ## 8. Final Verdict
132
+
133
+ | Metric | Value |
134
+ |--------|-------|
135
+ | Weighted Score | 0.00 / 3.00 |
136
+ | Threshold | 2.80 |
137
+ | Screenshot Coverage | 0% |
138
+ | a11y Critical+Serious | 0 |
139
+ | AI Slop Count | 0 |
140
+ | Regression Visual Changes | 0 |
141
+ | Cross-Validation Inconsistencies | 0 |
142
+ | **Verdict** | **PENDING** |
143
+
144
+ ### Verdict Rules (위반 불가)
145
+ 1. Weighted Score < 2.80 → **FAIL**
146
+ 2. Screenshot 누락 페이지 존재 → 해당 페이지 관련 항목 Score = 0으로 재계산
147
+ 3. a11y Critical 또는 Serious > 0 → V3 = 0으로 재계산
148
+ 4. AI Slop >= 2 → V4 = 최대 1점으로 재계산
149
+ 5. Regression 의도치 않은 변경 > 0 → **FAIL** (점수 무관)
150
+ 6. Cross-Validation 불일치 > 0 → **CONDITIONAL FAIL** (해당 Feature 재검증)
151
+ 7. Evidence 누락 항목 존재 → 해당 항목 Score = 0으로 재계산
@@ -3,4 +3,24 @@
3
3
  > Dispatcher가 관리. **모든 에이전트**는 세션 시작 시 이 파일을 읽고 학습된 규칙을 따릅니다.
4
4
  > gotchas(에이전트별 실수 기록)와 달리, 이 파일은 **프로젝트 전체에 적용되는 구조적 교훈**을 담습니다.
5
5
 
6
+ ## 항목 형식
7
+
8
+ ```markdown
9
+ ### [M-NNN] 간결한 제목
10
+ - **Date**: YYYY-MM-DD
11
+ - **Status**: unverified | verified
12
+ - **TTL**: YYYY-MM-DD (기본 +60일. 만료 후 Planner가 리뷰)
13
+ - **Lesson**: 교훈 내용
14
+ - **Context**: 발견 배경
15
+ - **Applies to**: 적용 대상 에이전트/상황
16
+ ```
17
+
18
+ ## 오염 방어 규칙
19
+
20
+ - 새 항목은 반드시 `unverified` 상태로 시작
21
+ - Planner가 스프린트 리뷰 시 유효성 확인 후 `verified` 승격
22
+ - TTL 만료 항목은 Planner가 리뷰: 갱신 또는 삭제
23
+ - 환각(hallucination) 의심 항목: 코드/git 이력으로 검증 불가하면 즉시 삭제
24
+ - 항목이 15개 초과 시 가장 오래된 unverified 항목부터 정리
25
+
6
26
  <!-- 항목이 추가되면 아래에 기록됩니다 -->
@@ -17,6 +17,12 @@
17
17
  "message": null,
18
18
  "retry_target": null
19
19
  },
20
+ "artifacts": {
21
+ "plan.md": { "status": "pending", "updated_by": null, "updated_at": null },
22
+ "feature-list.json": { "status": "pending", "updated_by": null, "updated_at": null },
23
+ "api-contract.json": { "status": "pending", "updated_by": null, "updated_at": null },
24
+ "sprint-contract.md": { "status": "pending", "updated_by": null, "updated_at": null }
25
+ },
20
26
  "updated_at": "{{DATE}}",
21
27
  "history": []
22
28
  }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@walwal-harness/cli",
3
- "version": "2.4.0",
3
+ "version": "3.2.0",
4
4
  "description": "Production harness for AI agent engineering — Planner, Generator(BE/FE), Evaluator(Func/Visual), optional Brainstormer (requirements refinement). Supports React and Flutter FE stacks.",
5
5
  "bin": {
6
6
  "walwal-harness": "bin/init.js"
@@ -34,5 +34,8 @@
34
34
  "scripts/",
35
35
  "assets/",
36
36
  "gotchas/"
37
- ]
37
+ ],
38
+ "dependencies": {
39
+ "@walwal-harness/cli": "^3.2.0"
40
+ }
38
41
  }
@@ -6,6 +6,7 @@ set -e
6
6
 
7
7
  SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
8
8
  source "$SCRIPT_DIR/lib/harness-render-progress.sh"
9
+ source "$SCRIPT_DIR/lib/harness-guardrail.sh"
9
10
 
10
11
  # ─────────────────────────────────────────
11
12
  # Resolve project root
@@ -41,17 +42,24 @@ next_agent=$(jq -r '.next_agent // "null"' "$PROGRESS")
41
42
  retry_count=$(jq -r '.sprint.retry_count // 0' "$PROGRESS")
42
43
  max_retries=$(jq -r '.flow.max_retries_per_sprint // 10' "$CONFIG" 2>/dev/null || echo 10)
43
44
 
44
- # fe_stack 치환 (Flutter 지원) — pipeline.json 에서 읽음
45
+ # fe_stack + fe_target 치환 (Flutter Web/Mobile/Desktop 지원) — pipeline.json 에서 읽음
45
46
  fe_stack="react"
47
+ fe_target="web"
46
48
  if [ -f "$PIPELINE_JSON" ]; then
47
49
  fe_stack=$(jq -r '.fe_stack // "react"' "$PIPELINE_JSON" 2>/dev/null || echo "react")
50
+ fe_target=$(jq -r '.fe_target // empty' "$PIPELINE_JSON" 2>/dev/null || true)
51
+ if [ -z "$fe_target" ]; then
52
+ # pipeline.json 에 fe_target 미지정 시 config.json 의 _default_target 사용
53
+ fe_target=$(jq -r ".flow.pipeline_selection.fe_stack_substitution.${fe_stack}._default_target // \"web\"" "$CONFIG" 2>/dev/null || echo "web")
54
+ fi
48
55
  fi
49
56
 
50
57
  # ─────────────────────────────────────────
51
- # fe_stack 치환 헬퍼
52
- # pipeline_selection.pipelines에서 읽은 에이전트명을 fe_stack에 따라 치환
58
+ # fe_stack + fe_target 치환 헬퍼
59
+ # pipeline_selection.pipelines 에서 읽은 에이전트명을 fe_stack/fe_target 에 따라 치환
53
60
  # - react: 그대로
54
- # - flutter: generator-frontend → generator-frontend-flutter 등, __skip__ 은 건너뜀
61
+ # - flutter+web: generator-frontend → generator-frontend-flutter 만 치환, eval 은 그대로 (Playwright 사용 가능)
62
+ # - flutter+mobile/desktop: eval 도 정적 분석용으로 치환, evaluator-visual 은 __skip__
55
63
  # ─────────────────────────────────────────
56
64
  substitute_fe_stack() {
57
65
  local agent="$1"
@@ -60,10 +68,94 @@ substitute_fe_stack() {
60
68
  return
61
69
  fi
62
70
  local sub
63
- sub=$(jq -r ".flow.pipeline_selection.fe_stack_substitution.flutter[\"${agent}\"] // \"${agent}\"" "$CONFIG" 2>/dev/null)
71
+ sub=$(jq -r ".flow.pipeline_selection.fe_stack_substitution.${fe_stack}.by_target[\"${fe_target}\"][\"${agent}\"] // \"${agent}\"" "$CONFIG" 2>/dev/null)
64
72
  echo "$sub"
65
73
  }
66
74
 
75
+ # ─────────────────────────────────────────
76
+ # Pre-Eval Gate — deterministic checks before Evaluator
77
+ # Generator가 완료되고 다음이 Evaluator일 때, lint/type/test를 먼저 실행.
78
+ # 실패 시 Evaluator를 건너뛰고 Generator로 리라우팅.
79
+ # ─────────────────────────────────────────
80
+ run_pre_eval_gate() {
81
+ local next="$1"
82
+ local gate_enabled
83
+ gate_enabled=$(jq -r '.flow.pre_eval_gate.enabled // false' "$CONFIG" 2>/dev/null)
84
+ if [ "$gate_enabled" != "true" ]; then return 0; fi
85
+
86
+ # Evaluator 에이전트인지 확인
87
+ case "$next" in
88
+ evaluator-*) ;;
89
+ *) return 0 ;;
90
+ esac
91
+
92
+ local timeout
93
+ timeout=$(jq -r '.flow.pre_eval_gate.timeout_seconds // 120' "$CONFIG" 2>/dev/null)
94
+
95
+ # 실패 위치 결정 (backend or frontend)
96
+ local location="backend"
97
+ local checks_key="backend_checks"
98
+ if [ "$current_agent" = "generator-frontend" ] || [ "$current_agent" = "generator-frontend-flutter" ]; then
99
+ location="frontend"
100
+ checks_key="frontend_checks"
101
+ fi
102
+
103
+ local -a checks
104
+ mapfile -t checks < <(jq -r ".flow.pre_eval_gate.${checks_key}[]" "$CONFIG" 2>/dev/null)
105
+
106
+ if [ ${#checks[@]} -eq 0 ]; then return 0; fi
107
+
108
+ echo ""
109
+ echo " ── Pre-Eval Gate ──────────────────────"
110
+ local all_pass=true
111
+ local fail_log=""
112
+
113
+ for cmd in "${checks[@]}"; do
114
+ printf " %-40s " "$cmd"
115
+ local output
116
+ if output=$(cd "$PROJECT_ROOT" && timeout "${timeout}s" bash -c "$cmd" 2>&1); then
117
+ echo "✓"
118
+ else
119
+ echo "✗"
120
+ all_pass=false
121
+ fail_log+="[FAIL] $cmd"$'\n'"$output"$'\n\n'
122
+ fi
123
+ done
124
+
125
+ if [ "$all_pass" = true ]; then
126
+ echo " Gate: PASS — proceeding to $next"
127
+ echo ""
128
+ return 0
129
+ else
130
+ echo ""
131
+ echo " Gate: FAIL — rerouting to $current_agent"
132
+ echo ""
133
+
134
+ # progress.json 업데이트: 실패 기록 + Generator로 리라우팅
135
+ local new_retry=$((retry_count + 1))
136
+ local fail_summary
137
+ fail_summary=$(echo "$fail_log" | head -20)
138
+
139
+ jq --arg agent "$current_agent" \
140
+ --arg loc "$location" \
141
+ --arg msg "Pre-eval gate failed: $fail_summary" \
142
+ --arg target "$current_agent" \
143
+ --argjson retry "$new_retry" \
144
+ '.sprint.status = "failed" |
145
+ .sprint.retry_count = $retry |
146
+ .agent_status = "failed" |
147
+ .next_agent = $target |
148
+ .failure.agent = $agent |
149
+ .failure.location = $loc |
150
+ .failure.message = $msg |
151
+ .failure.retry_target = $target' "$PROGRESS" > "${PROGRESS}.tmp" && mv "${PROGRESS}.tmp" "$PROGRESS"
152
+
153
+ # next_agent를 Generator로 덮어쓰기
154
+ next_agent="$current_agent"
155
+ return 1
156
+ fi
157
+ }
158
+
67
159
  # ─────────────────────────────────────────
68
160
  # Determine next agent
69
161
  # ─────────────────────────────────────────
@@ -131,6 +223,71 @@ if [ "$agent_status" = "completed" ] && [ "$next_agent" = "null" ]; then
131
223
  next_agent=$(compute_next_agent "$current_agent" "$agent_status")
132
224
  fi
133
225
 
226
+ # ─────────────────────────────────────────
227
+ # Artifact Prerequisites — 선행 아티팩트 상태 검증
228
+ # ─────────────────────────────────────────
229
+ verify_artifact_prerequisites() {
230
+ local target_agent="$1"
231
+ local states_order='["pending","draft","reviewed","approved"]'
232
+
233
+ # 에이전트의 prerequisites 가져오기
234
+ local prereqs
235
+ prereqs=$(jq -r ".artifacts.prerequisites[\"${target_agent}\"] // empty" "$CONFIG" 2>/dev/null)
236
+ if [ -z "$prereqs" ] || [ "$prereqs" = "null" ]; then return 0; fi
237
+
238
+ local all_met=true
239
+ echo ""
240
+ echo " ── Artifact Prerequisites ─────────────"
241
+
242
+ # prereqs의 각 키(artifact명)를 순회
243
+ while IFS='=' read -r artifact required_status; do
244
+ artifact=$(echo "$artifact" | tr -d '"' | tr -d ' ')
245
+ required_status=$(echo "$required_status" | tr -d '"' | tr -d ' ')
246
+ if [ -z "$artifact" ]; then continue; fi
247
+
248
+ local current_status
249
+ current_status=$(jq -r ".artifacts[\"${artifact}\"].status // \"pending\"" "$PROGRESS" 2>/dev/null)
250
+
251
+ # 상태 순서 비교
252
+ local required_idx current_idx
253
+ required_idx=$(echo "$states_order" | jq "index(\"$required_status\") // 0")
254
+ current_idx=$(echo "$states_order" | jq "index(\"$current_status\") // 0")
255
+
256
+ if [ "$current_idx" -ge "$required_idx" ]; then
257
+ printf " ✓ %-25s %s (required: %s)\n" "$artifact" "$current_status" "$required_status"
258
+ else
259
+ printf " ✗ %-25s %s (required: %s)\n" "$artifact" "$current_status" "$required_status"
260
+ all_met=false
261
+ fi
262
+ done < <(jq -r ".artifacts.prerequisites[\"${target_agent}\"] | to_entries[] | \"\(.key)=\(.value)\"" "$CONFIG" 2>/dev/null)
263
+
264
+ echo ""
265
+
266
+ if [ "$all_met" = true ]; then
267
+ echo " Prerequisites: PASS"
268
+ else
269
+ echo " Prerequisites: FAIL — preceding agent must complete artifacts first"
270
+ fi
271
+ echo ""
272
+
273
+ [ "$all_met" = true ] && return 0 || return 1
274
+ }
275
+
276
+ # ─────────────────────────────────────────
277
+ # Runtime Guardrail — 파일 소유권 검증
278
+ # ─────────────────────────────────────────
279
+ if [ "$agent_status" = "completed" ]; then
280
+ verify_file_ownership "$PROJECT_ROOT" || true
281
+ fi
282
+
283
+ # ─────────────────────────────────────────
284
+ # Run Pre-Eval Gate (if applicable)
285
+ # ─────────────────────────────────────────
286
+ if [ "$agent_status" = "completed" ]; then
287
+ run_pre_eval_gate "$next_agent" || true
288
+ verify_artifact_prerequisites "$next_agent" || true
289
+ fi
290
+
134
291
  # ─────────────────────────────────────────
135
292
  # Render progress
136
293
  # ─────────────────────────────────────────
@@ -142,6 +299,18 @@ echo ""
142
299
  # Generate next-prompt.txt
143
300
  # ─────────────────────────────────────────
144
301
  if [ "$next_agent" != "null" ] && [ "$next_agent" != "archive" ] && [ "$agent_status" != "blocked" ]; then
302
+ # ── Escalation check: 3회 실패 시 Planner에게 scope 축소 요청 ──
303
+ escalate_after=$(jq -r '.flow.escalate_to_planner_after // 3' "$CONFIG" 2>/dev/null || echo 3)
304
+ if [ "$retry_count" -ge "$escalate_after" ] && [ "$sprint_status" = "failed" ] && [ "$next_agent" != "planner" ]; then
305
+ echo " ⚠ Escalation: ${retry_count}회 실패 — Planner에게 scope 축소/접근 변경 요청"
306
+ next_agent="planner"
307
+ # progress.json에 에스컬레이션 기록
308
+ jq --arg msg "Escalated after ${retry_count} failures. Planner must review scope or approach." \
309
+ '.next_agent = "planner" |
310
+ .failure.message = $msg |
311
+ .failure.retry_target = "planner"' "$PROGRESS" > "${PROGRESS}.tmp" && mv "${PROGRESS}.tmp" "$PROGRESS"
312
+ fi
313
+
145
314
  # Build prompt
146
315
  prompt="/harness-${next_agent} 를 실행하세요."
147
316
 
@@ -155,14 +324,105 @@ if [ "$next_agent" != "null" ] && [ "$next_agent" != "archive" ] && [ "$agent_st
155
324
 
156
325
  prompt+=$'\n'".harness/progress.json을 읽고 현재 상태를 확인하세요."
157
326
 
158
- # Add failure context if retrying
327
+ # Add failure context if retrying (include previous failure summary)
159
328
  failure_msg=$(jq -r '.failure.message // empty' "$PROGRESS")
160
329
  if [ -n "$failure_msg" ] && [ "$failure_msg" != "null" ]; then
161
330
  prompt+=$'\n\n'"이전 실패 사유: ${failure_msg}"
331
+ prompt+=$'\n'"같은 접근을 반복하지 말고, 실패 원인을 분석한 후 다른 전략으로 시도하세요."
162
332
  fi
163
333
 
164
334
  echo "$prompt" > "$NEXT_PROMPT"
165
335
 
336
+ # ── Generate structured handoff.json ──
337
+ HANDOFF="$PROJECT_ROOT/.harness/handoff.json"
338
+ FEATURE_LIST="$PROJECT_ROOT/.harness/actions/feature-list.json"
339
+
340
+ # Collect available artifacts
341
+ local -a artifacts_ready=()
342
+ for f in plan.md feature-list.json api-contract.json sprint-contract.md evaluation-functional.md evaluation-visual.md; do
343
+ if [ -f "$PROJECT_ROOT/.harness/actions/$f" ]; then
344
+ artifacts_ready+=("$f")
345
+ fi
346
+ done
347
+ local artifacts_json
348
+ artifacts_json=$(printf '%s\n' "${artifacts_ready[@]}" | jq -R . | jq -s .)
349
+
350
+ # Collect focus features (incomplete ones)
351
+ local focus_features="[]"
352
+ if [ -f "$FEATURE_LIST" ]; then
353
+ focus_features=$(jq '[.features[]? | select(.passes == null or (.passes | length) == 0 or ((.passes // []) | map(select(. == "evaluator-functional")) | length == 0)) | .id] | .[0:5]' "$FEATURE_LIST" 2>/dev/null || echo "[]")
354
+ fi
355
+
356
+ # ── Regression data: collect previous sprint's passed AC from archive ──
357
+ local regression_source="null"
358
+ local prev_sprint=$((sprint_num - 1))
359
+ local prev_archive="$PROJECT_ROOT/.harness/archive/sprint-$(printf '%03d' $prev_sprint)"
360
+ if [ "$prev_sprint" -ge 1 ] && [ -d "$prev_archive" ]; then
361
+ if [ -f "$prev_archive/feature-list.json" ]; then
362
+ regression_source=$(jq '{
363
+ sprint: '"$prev_sprint"',
364
+ passed_features: [.features[]? | select((.passes // []) | map(select(. == "evaluator-functional")) | length > 0) | {id, name, acceptance_criteria}],
365
+ archive_path: "'"$prev_archive"'"
366
+ }' "$prev_archive/feature-list.json" 2>/dev/null || echo "null")
367
+ fi
368
+ fi
369
+
370
+ # ── Eval-specific scoring config ──
371
+ local eval_config="null"
372
+ local cross_validation_data="null"
373
+ case "$next_agent" in
374
+ evaluator-*)
375
+ eval_config=$(jq '{
376
+ pass_threshold: .evaluation.scoring.pass_threshold,
377
+ scale: .evaluation.scoring.scale,
378
+ verdict_rules: .evaluation.scoring.verdict_rules,
379
+ regression_enabled: .evaluation.regression.enabled,
380
+ cross_validation_enabled: .evaluation.cross_validation.enabled,
381
+ adversarial_rules: .agents["'"$next_agent"'"].adversarial_rules.rules,
382
+ forbidden: .agents["'"$next_agent"'"].adversarial_rules.forbidden
383
+ }' "$CONFIG" 2>/dev/null || echo "null")
384
+ ;;
385
+ esac
386
+
387
+ # ── Cross-Validation: evaluator-visual이면 functional 결과 파싱 ──
388
+ if [ "$next_agent" = "evaluator-visual" ]; then
389
+ local func_eval="$PROJECT_ROOT/.harness/actions/evaluation-functional.md"
390
+ if [ -f "$func_eval" ]; then
391
+ # evaluation-functional.md 내 JSON 코드블록에서 Cross-Validation Data 추출
392
+ cross_validation_data=$(sed -n '/```json/,/```/p' "$func_eval" | tail -n +2 | head -n -1 | jq 'select(.evaluator == "functional")' 2>/dev/null || echo "null")
393
+ fi
394
+ fi
395
+
396
+ # Build handoff.json
397
+ jq -n \
398
+ --arg from "${current_agent:-dispatcher}" \
399
+ --arg to "$next_agent" \
400
+ --argjson sprint "$sprint_num" \
401
+ --argjson retry "$retry_count" \
402
+ --arg status "$sprint_status" \
403
+ --arg failure_msg "${failure_msg:-}" \
404
+ --argjson artifacts "$artifacts_json" \
405
+ --argjson focus "$focus_features" \
406
+ --argjson regression "$regression_source" \
407
+ --argjson eval_config "$eval_config" \
408
+ --argjson cross_val "$cross_validation_data" \
409
+ --arg timestamp "$(date -u +%Y-%m-%dT%H:%M:%SZ)" \
410
+ '{
411
+ from: $from,
412
+ to: $to,
413
+ sprint: $sprint,
414
+ retry_count: $retry,
415
+ sprint_status: $status,
416
+ failure_context: (if $failure_msg != "" then $failure_msg else null end),
417
+ artifacts_ready: $artifacts,
418
+ focus_features: $focus,
419
+ regression: $regression,
420
+ eval_config: $eval_config,
421
+ cross_validation_from_functional: $cross_val,
422
+ warnings: [],
423
+ timestamp: $timestamp
424
+ }' > "$HANDOFF"
425
+
166
426
  elif [ "$next_agent" = "archive" ]; then
167
427
  cat > "$NEXT_PROMPT" <<'PROMPT'
168
428
  Sprint 문서를 아카이브 하세요.