@frankzhang2026/opencode-android-orchestrator 1.0.4 → 1.0.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,7 +1,7 @@
1
1
  # Troubleshooting
2
2
 
3
3
  Use this guide for
4
- `@frankzhang2026/opencode-android-orchestrator@1.0.4`.
4
+ `@frankzhang2026/opencode-android-orchestrator@1.0.5`.
5
5
 
6
6
  ## Start with read-only evidence
7
7
 
@@ -11,7 +11,7 @@ From the repository root, capture:
11
11
  git status --short --branch
12
12
  git rev-parse HEAD
13
13
  opencode --version
14
- npx @frankzhang2026/opencode-android-orchestrator@1.0.4 doctor . --json
14
+ npx @frankzhang2026/opencode-android-orchestrator@1.0.5 doctor . --json
15
15
  ```
16
16
 
17
17
  If installation never completed, doctor will correctly report a missing or
@@ -39,9 +39,9 @@ command-scoped override:
39
39
 
40
40
  ```sh
41
41
  npm --registry=https://registry.npmjs.org/ view \
42
- @frankzhang2026/opencode-android-orchestrator@1.0.4 version
42
+ @frankzhang2026/opencode-android-orchestrator@1.0.5 version
43
43
  npx --yes --registry=https://registry.npmjs.org/ \
44
- @frankzhang2026/opencode-android-orchestrator@1.0.4 upgrade . --json
44
+ @frankzhang2026/opencode-android-orchestrator@1.0.5 upgrade . --json
45
45
  ```
46
46
 
47
47
  This leaves the company's saved npm configuration unchanged. Use the option
@@ -71,7 +71,7 @@ Git-backed Superpowers plugin at runtime.
71
71
  | Invalid `--long-command-timeout-ms` | The value is not an integer from `120000` through `7200000`. | Use the `1800000` ms default or pass an intentional bounded value to `init`/`upgrade`; do not edit the generated config directly. |
72
72
  | Android SDK failure | No valid explicit SDK, `ANDROID_HOME`, `ANDROID_SDK_ROOT`, or `local.properties` `sdk.dir` was found. | Configure one real SDK root containing `platforms/` and `build-tools/`. Do not publish `local.properties`. |
73
73
  | Missing `git`, `jq`, `rg`, `shasum`, or Java | Required deterministic command is unavailable on `PATH`. | Install or restore the missing command, record its version, and rerun the read-only checks. |
74
- | `Bundled Orchestrator skill is unavailable` | The installed `1.0.4` package is incomplete, damaged, or loaded from an unsupported partial copy. | Reinstall the exact package, inspect its `resources/third-party/superpowers-v6.2.0/skills/` entries, restart OpenCode, and rerun `opencode debug skill`. Do not add an external Superpowers plugin as a fallback. |
74
+ | `Bundled Orchestrator skill is unavailable` | The installed `1.0.5` package is incomplete, damaged, or loaded from an unsupported partial copy. | Reinstall the exact package, inspect its `resources/third-party/superpowers-v6.2.0/skills/` entries, restart OpenCode, and rerun `opencode debug skill`. Do not add an external Superpowers plugin as a fallback. |
75
75
  | `current process does not own this task queue execution` immediately after Coder start on 1.0.1 | OpenCode created the tool shell in a separate process group, so 1.0.1 rejected a legitimate Worker descendant. | Upgrade to 1.0.2 or later, restart OpenCode, then use the approved resume or abort workflow for the retained task. Do not edit the queue or lease files. |
76
76
  | The exact Superpowers v6.2.0 plugin remains after upgrade | That entry existed in the verified pre-install OpenCode file and is therefore user-owned. | Leave it in place or remove it as a separate reviewed configuration change. Upgrade only removes the old Orchestrator-managed entry. |
77
77
 
@@ -88,7 +88,7 @@ silence of `./gradlew tasks --all --console=plain | rg ...` in a large build.
88
88
  For an existing installation, run:
89
89
 
90
90
  ```sh
91
- npx @frankzhang2026/opencode-android-orchestrator@1.0.4 upgrade . \
91
+ npx @frankzhang2026/opencode-android-orchestrator@1.0.5 upgrade . \
92
92
  --refresh-gradle-discovery
93
93
  ```
94
94
 
@@ -97,7 +97,7 @@ least `1800000` milliseconds. A higher timeout already supplied by the caller
97
97
  is preserved; unrelated Bash commands are unchanged. To configure one hour,
98
98
  run `upgrade . --long-command-timeout-ms 3600000` on a healthy installation.
99
99
  If a command still reports `120000 ms`, confirm that the project manifest and
100
- OpenCode plugin reference are both `1.0.4`, restart the OpenCode session so the
100
+ OpenCode plugin reference are both `1.0.5`, restart the OpenCode session so the
101
101
  plugin reloads, and rerun doctor before attempting recovery.
102
102
 
103
103
  After installation, inspect OpenCode discovery separately:
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@frankzhang2026/opencode-android-orchestrator",
3
- "version": "1.0.4",
3
+ "version": "1.0.5",
4
4
  "description": "Reusable OpenCode orchestration for Android projects",
5
5
  "license": "MIT",
6
6
  "author": "frankzhang2026",
@@ -109,6 +109,15 @@ literally. Do not infer missing requirements and do not ask questions during a
109
109
  non-interactive run. If anything is ambiguous or blocked, stop and report the exact
110
110
  reason; the deterministic scripts own state transitions.
111
111
 
112
+ For schema V3 tasks, keep production code unchanged while adding the approved
113
+ tests, then call `./scripts/automation/record-red.sh <TASK-ID>` with no model-
114
+ chosen failure text. The script checks every declared case. A familiar exception
115
+ name in a log is not sufficient RED. Fix a test-only preparation error only when
116
+ the contract remains unchanged and its preparation budget allows it; rerun the
117
+ preflight afterwards. Contract contradictions must be blocked with the reported
118
+ case IDs and measured results. Never delete evidence or weaken, skip or reclassify
119
+ a test to pass the preflight.
120
+
112
121
  You may edit only paths allowed both by this agent and by the task contract.
113
122
  Treat `.automation-worktree-allowlist` and the status JSON's
114
123
  `runtime.effectiveWorktreeAllowlist` paths as human-owned local state: never
@@ -62,6 +62,16 @@ approval. After approval, assemble a complete plan and contract in memory and
62
62
  call `android_orchestrator_intake` with action `draft`; do not create files in the
63
63
  product checkout. Preserve the snapshot's target branch and planningHead.
64
64
 
65
+ Use schema V3 structured verification. Give every verification case a stable ID,
66
+ its one-based acceptance-criterion reference, its source, and its exact test
67
+ identity. Classify preserved behavior as `before: pass`, changed behavior as
68
+ `before: fail`, and an uncertain old boundary as `before: observe`. Do not infer
69
+ exact serializer, parser, locale, date or framework output from declarations or
70
+ memory. For preserved behavior, use an existing trusted test or explicitly mark
71
+ the value for baseline capture. Plan examples are implementation guidance and
72
+ must not add requirements beyond the contract. Check acceptance criteria,
73
+ non-goals and verification expectations for contradictions before drafting.
74
+
65
75
  When the draft does not specify a workspace or commit policy, use the values in
66
76
  `automation/config.json`; new installations configure `inPlaceExclusive` and
67
77
  `humanApproval`. A user may explicitly override the commit policy to
@@ -52,7 +52,13 @@ it with `./scripts/automation/block-task.sh <TASK-ID> <reason>` before stopping.
52
52
 
53
53
  4. On the initial coding cycle, add or change the smallest behavior test
54
54
  permitted by `allowedPaths`.
55
- 5. If RED evidence does not already exist, capture a genuine RED result with:
55
+ 5. If RED evidence does not already exist, capture a genuine RED result. For a
56
+ schema V3 contract use:
57
+
58
+ `./scripts/automation/record-red.sh <TASK-ID>`
59
+
60
+ It checks every declared preserved, changed and observed case against fresh
61
+ structured output. For a legacy schema V1/V2 contract use:
56
62
 
57
63
  `./scripts/automation/record-red.sh <TASK-ID> <expected-failure-text> -- <test-filter>`
58
64
 
@@ -38,6 +38,10 @@ you review.
38
38
  5. Check each acceptance criterion against observable behavior. Inspect for
39
39
  regression risk, missing edge cases, out-of-scope changes, test deletion,
40
40
  ignored tests, relaxed assertions, and implementation-shaped tests.
41
+ For schema V3, also verify every structured case source and identity, that
42
+ preserved cases passed before implementation, that changed cases failed only
43
+ for their declared reason, and that RED contains no undeclared failure. Treat
44
+ the Planner and Coder summaries as claims; use the bound preflight evidence.
41
45
  6. Decide independently:
42
46
 
43
47
  - approve only when the diff is correct and evidence is sufficient;
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "$schema": "https://json-schema.org/draft/2020-12/schema",
3
- "$id": "urn:frankzhang2026:opencode-android-orchestrator:task-contract:v2",
3
+ "$id": "urn:frankzhang2026:opencode-android-orchestrator:task-contract:v3",
4
4
  "title": "Scheduled coding task contract",
5
5
  "type": "object",
6
6
  "additionalProperties": false,
@@ -14,6 +14,63 @@
14
14
  "type": "string",
15
15
  "minLength": 1,
16
16
  "pattern": "^[A-Za-z0-9_.#$*-]+$"
17
+ },
18
+ "expectedFailure": {
19
+ "type": "object",
20
+ "additionalProperties": false,
21
+ "required": ["type", "origin"],
22
+ "properties": {
23
+ "type": { "type": "string", "minLength": 1 },
24
+ "messageIncludes": { "type": "string", "minLength": 3 },
25
+ "origin": { "type": "string", "minLength": 12 }
26
+ }
27
+ },
28
+ "verificationCase": {
29
+ "type": "object",
30
+ "additionalProperties": false,
31
+ "required": ["id", "criterion", "intent", "before", "after", "source", "test"],
32
+ "properties": {
33
+ "id": { "type": "string", "pattern": "^[A-Z][A-Z0-9-]{2,63}$" },
34
+ "criterion": { "type": "integer", "minimum": 1 },
35
+ "intent": { "enum": ["preserve", "change", "observe"] },
36
+ "before": { "enum": ["pass", "fail", "observe"] },
37
+ "after": { "const": "pass" },
38
+ "source": { "enum": ["userRequirement", "existingTest", "baselineCapture", "measuredFact"] },
39
+ "test": {
40
+ "type": "object",
41
+ "additionalProperties": false,
42
+ "required": ["target", "className", "name"],
43
+ "properties": {
44
+ "target": { "type": "integer", "minimum": 0 },
45
+ "className": { "type": "string", "minLength": 1 },
46
+ "name": { "type": "string", "minLength": 1 }
47
+ }
48
+ },
49
+ "expectedFailure": { "$ref": "#/$defs/expectedFailure" }
50
+ },
51
+ "allOf": [
52
+ {
53
+ "if": { "properties": { "intent": { "const": "preserve" } } },
54
+ "then": {
55
+ "properties": {
56
+ "before": { "const": "pass" },
57
+ "source": { "enum": ["existingTest", "baselineCapture", "measuredFact"] }
58
+ },
59
+ "not": { "required": ["expectedFailure"] }
60
+ }
61
+ },
62
+ {
63
+ "if": { "properties": { "intent": { "const": "change" } } },
64
+ "then": {
65
+ "properties": { "before": { "const": "fail" }, "source": { "const": "userRequirement" } },
66
+ "required": ["expectedFailure"]
67
+ }
68
+ },
69
+ {
70
+ "if": { "properties": { "intent": { "const": "observe" } } },
71
+ "then": { "properties": { "before": { "const": "observe" } } }
72
+ }
73
+ ]
17
74
  }
18
75
  },
19
76
  "required": [
@@ -32,10 +89,11 @@
32
89
  "nonGoals",
33
90
  "targetTests",
34
91
  "deviceTestsRequired",
35
- "testPolicy"
92
+ "testPolicy",
93
+ "verification"
36
94
  ],
37
95
  "properties": {
38
- "schemaVersion": { "const": 2 },
96
+ "schemaVersion": { "const": 3 },
39
97
  "id": { "type": "string", "pattern": "^TASK-[A-Z0-9-]+$" },
40
98
  "title": { "type": "string", "minLength": 1 },
41
99
  "designApproved": { "const": true },
@@ -87,6 +145,20 @@
87
145
  },
88
146
  "deviceTestsRequired": { "type": "boolean" },
89
147
  "testPolicy": { "enum": ["required", "not-required"] },
90
- "testPolicyReason": { "type": "string" }
148
+ "testPolicyReason": { "type": "string" },
149
+ "verification": {
150
+ "type": "object",
151
+ "additionalProperties": false,
152
+ "required": ["version", "maxPreparationFixes", "cases"],
153
+ "properties": {
154
+ "version": { "const": 1 },
155
+ "maxPreparationFixes": { "type": "integer", "minimum": 0, "maximum": 1 },
156
+ "cases": {
157
+ "type": "array",
158
+ "minItems": 1,
159
+ "items": { "$ref": "#/$defs/verificationCase" }
160
+ }
161
+ }
162
+ }
91
163
  }
92
164
  }
@@ -1,5 +1,5 @@
1
1
  {
2
- "schemaVersion": 2,
2
+ "schemaVersion": 3,
3
3
  "id": "TASK-EXAMPLE-001",
4
4
  "title": "Replace with one small, observable behavior change",
5
5
  "designApproved": true,
@@ -53,5 +53,28 @@
53
53
  ],
54
54
  "deviceTestsRequired": false,
55
55
  "testPolicy": "required",
56
- "testPolicyReason": "Behavior changes require a focused regression test"
56
+ "testPolicyReason": "Behavior changes require a focused regression test",
57
+ "verification": {
58
+ "version": 1,
59
+ "maxPreparationFixes": 1,
60
+ "cases": [
61
+ {
62
+ "id": "FOCUSED-BEHAVIOR",
63
+ "criterion": 1,
64
+ "intent": "change",
65
+ "before": "fail",
66
+ "after": "pass",
67
+ "source": "userRequirement",
68
+ "test": {
69
+ "target": 0,
70
+ "className": "ReplaceWithFocusedTest",
71
+ "name": "replace with observable behavior"
72
+ },
73
+ "expectedFailure": {
74
+ "type": "java.lang.AssertionError",
75
+ "origin": "The assertion at the approved behavior call"
76
+ }
77
+ }
78
+ ]
79
+ }
57
80
  }
@@ -26,6 +26,16 @@ red_exit_code="$(jq -er '.exitCode' "$red_file")"
26
26
  review_verification_exit_code="$(jq -er '.verificationExitCode' "$review_file")"
27
27
  [[ "$red_exit_code" -ne 0 ]] || automation_die "RED evidence does not contain a failing test result"
28
28
  [[ "$review_verification_exit_code" -eq 0 ]] || automation_die "independent review verification did not pass"
29
+ structured_red=null
30
+ if [[ "$(jq -er '.schemaVersion' "$contract")" == "3" ]]; then
31
+ preflight_file="$evidence_dir/test-preflight.json"
32
+ manifest_file="$evidence_dir/test-manifest.json"
33
+ [[ -f "$preflight_file" && -f "$manifest_file" ]] || automation_die "structured RED acceptance evidence is incomplete"
34
+ [[ "$(jq -er '.valid' "$preflight_file")" == "true" ]] || automation_die "structured RED preflight was not valid"
35
+ [[ "$(jq -er '.preflightSha256' "$red_file")" == "$(automation_file_sha256 "$preflight_file")" ]] || automation_die "structured RED preflight changed"
36
+ [[ "$(jq -er '.manifestSha256' "$red_file")" == "$(automation_file_sha256 "$manifest_file")" ]] || automation_die "structured RED manifest changed"
37
+ structured_red="$(jq -c '{valid, reasonCode, summary, cases: [.cases[] | {id, criterion, intent, expectedBefore, test, valid}]}' "$preflight_file")"
38
+ fi
29
39
 
30
40
  recorded_task_root="$(automation_workspace_task_root "$workspace_file")"
31
41
  workspace_strategy="$(automation_workspace_strategy "$workspace_file")"
@@ -90,6 +100,7 @@ jq -n \
90
100
  --argjson acceptanceCriteria "$(jq -c '.acceptanceCriteria' "$contract")" \
91
101
  --argjson nonGoals "$(jq -c '.nonGoals' "$contract")" \
92
102
  --argjson targetTests "$(jq -c '.targetTests' "$contract")" \
103
+ --argjson structuredRed "$structured_red" \
93
104
  '{taskId: $taskId, title: $title, state: $state,
94
105
  generatedAt: $generatedAt, originalBranch: $originalBranch,
95
106
  originalHeadBeforeContract: $originalHeadBeforeContract,
@@ -110,6 +121,7 @@ jq -n \
110
121
  baselineRecorded: true,
111
122
  redRecorded: true,
112
123
  redExitCode: $redExitCode,
124
+ structuredRed: $structuredRed,
113
125
  qualityGate: "PASSED",
114
126
  gateAttempts: $gateAttempts,
115
127
  codingCycle: $codingCycle,
@@ -566,6 +566,114 @@ automation_run_focused_test() {
566
566
  )
567
567
  }
568
568
 
569
+ # Run one focused target with fresh Test execution and machine-readable case
570
+ # results. Test assertion failures are collected instead of failing Gradle so
571
+ # the caller can distinguish approved RED from build and fixture failures.
572
+ automation_run_classified_focused_test() {
573
+ local task="$1"
574
+ local filter="$2"
575
+ local root="$3"
576
+ local result_file="$4"
577
+ local log_file="$5"
578
+ local init_file status
579
+
580
+ automation_validate_config || return 1
581
+ automation_validate_gradle_task "$task" || return 1
582
+ automation_validate_test_filter "$filter" || return 1
583
+ jq -e --arg task "$task" \
584
+ '.gradleVerification.focusedTestTasks | index($task) != null' \
585
+ "$AUTOMATION_CONFIG" >/dev/null || {
586
+ automation_die "focused Gradle task is not allowed by automation/config.json: $task"
587
+ return 1
588
+ }
589
+
590
+ mkdir -p "$AUTOMATION_RUNTIME_ROOT/gradle" "$(dirname "$result_file")" "$(dirname "$log_file")"
591
+ init_file="$(mktemp "$AUTOMATION_RUNTIME_ROOT/gradle/classified-tests.XXXXXX")"
592
+ : > "$result_file"
593
+ cat > "$init_file" <<'GRADLE'
594
+ import groovy.json.JsonOutput
595
+ import org.gradle.api.tasks.testing.Test
596
+ import org.gradle.api.tasks.testing.TestDescriptor
597
+ import org.gradle.api.tasks.testing.TestListener
598
+ import org.gradle.api.tasks.testing.TestResult
599
+
600
+ def outputPath = System.getProperty('orchestrator.caseResultFile')
601
+ if (outputPath == null || outputPath.isEmpty()) {
602
+ throw new GradleException('orchestrator.caseResultFile is required')
603
+ }
604
+ def outputFile = new File(outputPath)
605
+ def appendResult = { Map value ->
606
+ synchronized (gradle) {
607
+ outputFile << JsonOutput.toJson(value) << System.lineSeparator()
608
+ }
609
+ }
610
+
611
+ gradle.allprojects { project ->
612
+ project.tasks.withType(Test).configureEach { testTask ->
613
+ outputs.upToDateWhen { false }
614
+ outputs.doNotCacheIf('Orchestrator requires fresh classified test execution') { true }
615
+ ignoreFailures = true
616
+ failFast = false
617
+ if (testTask.hasProperty('dryRun')) testTask.dryRun = false
618
+ addTestListener(new TestListener() {
619
+ void beforeSuite(TestDescriptor descriptor) {}
620
+ void beforeTest(TestDescriptor descriptor) {}
621
+ void afterTest(TestDescriptor descriptor, TestResult result) {
622
+ def failure = result.exceptions == null || result.exceptions.isEmpty() ? null : result.exceptions[0]
623
+ appendResult([
624
+ kind: 'case', taskPath: testTask.path,
625
+ className: descriptor.className ?: '', name: descriptor.name ?: '',
626
+ result: result.resultType.toString(),
627
+ exceptionType: failure == null ? null : failure.class.name,
628
+ exceptionMessage: failure == null ? null : (failure.message ?: '')
629
+ ])
630
+ }
631
+ void afterSuite(TestDescriptor descriptor, TestResult result) {
632
+ if (descriptor.parent == null) {
633
+ appendResult([
634
+ kind: 'suite', taskPath: testTask.path,
635
+ tests: result.testCount, failures: result.failedTestCount,
636
+ skipped: result.skippedTestCount,
637
+ result: result.resultType.toString()
638
+ ])
639
+ }
640
+ }
641
+ })
642
+ }
643
+ }
644
+ GRADLE
645
+
646
+ set +e
647
+ (cd "$root" && ./gradlew "$task" --tests "$filter" \
648
+ --no-configuration-cache --console=plain --init-script "$init_file" \
649
+ "-Dorchestrator.caseResultFile=$result_file") 2>&1 | tee "$log_file"
650
+ status=${PIPESTATUS[0]}
651
+ set -e
652
+ rm -f "$init_file"
653
+ return "$status"
654
+ }
655
+
656
+ automation_test_diff_sha() {
657
+ local task_id="$1"
658
+ local root="${2:-$AUTOMATION_ROOT}"
659
+ local path tracked=0
660
+ {
661
+ while IFS= read -r path; do
662
+ [[ -n "$path" ]] || continue
663
+ if automation_array_matches_path "$AUTOMATION_CONFIG" '.androidProject.testPaths' "$path"; then
664
+ tracked=1
665
+ if git -C "$root" ls-files --error-unmatch -- "$path" >/dev/null 2>&1; then
666
+ git -C "$root" diff --binary --no-renames HEAD -- "$path"
667
+ else
668
+ printf 'UNTRACKED %s\0' "$path"
669
+ git -C "$root" hash-object -- "$path"
670
+ fi
671
+ fi
672
+ done < <(automation_product_changed_paths_at "$task_id" "$root")
673
+ [[ "$tracked" == "1" ]] || printf 'NO-TEST-CHANGES'
674
+ } | shasum -a 256 | awk '{print $1}'
675
+ }
676
+
569
677
  automation_require_approval() {
570
678
  local kind="$1"
571
679
  local supplied="$2"
@@ -6,51 +6,189 @@ SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
6
6
  source "$SCRIPT_DIR/lib.sh"
7
7
 
8
8
  task_id="${1:-}"
9
- expected="${2:-}"
10
- separator="${3:-}"
11
- filter="${4:-}"
12
-
13
- if [[ -z "$task_id" || -z "$expected" || "$separator" != "--" || -z "$filter" || "$#" -ne 4 ]]; then
14
- printf 'Usage: %s TASK-ID EXPECTED-FAILURE-TEXT -- TEST-FILTER\n' "$0" >&2
15
- exit 2
16
- fi
17
-
9
+ [[ -n "$task_id" ]] || { printf 'Usage: %s TASK-ID [LEGACY-EXPECTED -- LEGACY-FILTER]\n' "$0" >&2; exit 2; }
18
10
  automation_validate_task_id "$task_id"
19
11
  automation_require_queue_execution "$task_id"
20
- automation_validate_test_filter "$filter"
21
- [[ ${#expected} -ge 3 ]] || automation_die "expected failure text is too short"
22
12
  [[ "$(automation_read_state "$task_id")" == "CODING" ]] || automation_die "$task_id is not CODING"
23
13
  "$SCRIPT_DIR/validate-contract.sh" "$task_id"
24
14
 
25
15
  contract="$(automation_contract_path "$task_id")"
26
- target_count="$(jq -er --arg filter "$filter" '[.targetTests[] | select(.filter == $filter)] | length' "$contract")"
27
- [[ "$target_count" -eq 1 ]] || automation_die "test filter must identify exactly one declared contract target: $filter"
28
- gradle_task="$(jq -er --arg filter "$filter" '.targetTests[] | select(.filter == $filter) | .gradleTask' "$contract")"
29
-
16
+ schema_version="$(jq -er '.schemaVersion' "$contract")"
30
17
  evidence_dir="$(automation_evidence_path "$task_id")"
31
18
  mkdir -p "$evidence_dir"
32
- red_log="$evidence_dir/red.log"
33
19
  red_meta="$evidence_dir/red.json"
34
20
  [[ ! -e "$red_meta" ]] || automation_die "RED evidence already exists for $task_id"
35
- started_at="$(automation_now)"
36
21
 
37
- set +e
38
- automation_run_focused_test "$gradle_task" "$filter" "$AUTOMATION_ROOT" 2>&1 | tee "$red_log"
39
- red_status=${PIPESTATUS[0]}
40
- set -e
22
+ # Preserve the exact V1/V2 behavior for already approved contracts.
23
+ if [[ "$schema_version" != "3" ]]; then
24
+ expected="${2:-}"
25
+ separator="${3:-}"
26
+ filter="${4:-}"
27
+ if [[ -z "$expected" || "$separator" != "--" || -z "$filter" || "$#" -ne 4 ]]; then
28
+ printf 'Usage: %s TASK-ID EXPECTED-FAILURE-TEXT -- TEST-FILTER\n' "$0" >&2
29
+ exit 2
30
+ fi
31
+ automation_validate_test_filter "$filter"
32
+ [[ ${#expected} -ge 3 ]] || automation_die "expected failure text is too short"
33
+ target_count="$(jq -er --arg filter "$filter" '[.targetTests[] | select(.filter == $filter)] | length' "$contract")"
34
+ [[ "$target_count" -eq 1 ]] || automation_die "test filter must identify exactly one declared contract target: $filter"
35
+ gradle_task="$(jq -er --arg filter "$filter" '.targetTests[] | select(.filter == $filter) | .gradleTask' "$contract")"
36
+ red_log="$evidence_dir/red.log"
37
+ started_at="$(automation_now)"
38
+ set +e
39
+ automation_run_focused_test "$gradle_task" "$filter" "$AUTOMATION_ROOT" 2>&1 | tee "$red_log"
40
+ red_status=${PIPESTATUS[0]}
41
+ set -e
42
+ [[ "$red_status" -ne 0 ]] || automation_die "RED capture failed: focused test passed before implementation"
43
+ rg -F "$expected" "$red_log" >/dev/null || automation_die "RED output does not contain the expected failure text"
44
+ jq -n --arg taskId "$task_id" --arg startedAt "$started_at" --arg finishedAt "$(automation_now)" \
45
+ --arg expectedFailure "$expected" --arg gradleTask "$gradle_task" --arg testFilter "$filter" \
46
+ --argjson exitCode "$red_status" \
47
+ '{taskId: $taskId, startedAt: $startedAt, finishedAt: $finishedAt,
48
+ command: ["./gradlew", $gradleTask, "--tests", $testFilter],
49
+ expectedFailure: $expectedFailure, exitCode: $exitCode}' | automation_record_json "$red_meta"
50
+ automation_info "$task_id legacy RED evidence recorded"
51
+ exit 0
52
+ fi
53
+
54
+ [[ "$#" -eq 1 ]] || { printf 'Schema V3 usage: %s TASK-ID\n' "$0" >&2; exit 2; }
55
+ "$SCRIPT_DIR/scope-gate.sh" "$task_id" >/dev/null
41
56
 
42
- [[ "$red_status" -ne 0 ]] || automation_die "RED capture failed: focused test passed before implementation"
43
- rg -F "$expected" "$red_log" >/dev/null || automation_die "RED output does not contain the expected failure text"
57
+ # RED must be captured before any production change. Planning artifacts are
58
+ # excluded; test sources and fixtures are the only permitted product changes.
59
+ while IFS= read -r changed_path; do
60
+ [[ -n "$changed_path" ]] || continue
61
+ if ! automation_array_matches_path "$AUTOMATION_CONFIG" '.androidProject.testPaths' "$changed_path"; then
62
+ automation_die "RED preflight requires unchanged production code; non-test path changed: $changed_path"
63
+ fi
64
+ done < <(automation_product_changed_paths_at "$task_id" "$AUTOMATION_ROOT")
65
+
66
+ previous_attempt=0
67
+ preflight_meta="$evidence_dir/test-preflight.json"
68
+ if [[ -f "$preflight_meta" ]]; then
69
+ previous_attempt="$(jq -er '.attempt' "$preflight_meta")"
70
+ fi
71
+ attempt=$((previous_attempt + 1))
72
+ max_fixes="$(jq -er '.verification.maxPreparationFixes' "$contract")"
73
+ if [[ "$attempt" -gt $((max_fixes + 1)) ]]; then
74
+ automation_die "test preparation retry budget exhausted; revise the contract or abort the task"
75
+ fi
76
+
77
+ attempt_id="$(printf '%03d' "$attempt")"
78
+ attempt_dir="$evidence_dir/attempts/red-preflight-$attempt_id"
79
+ mkdir -p "$attempt_dir"
80
+ observed_file="$attempt_dir/observed.jsonl"
81
+ : > "$observed_file"
82
+ execution_failures='[]'
83
+ started_at="$(automation_now)"
84
+ target_index=0
85
+ while IFS=$'\t' read -r gradle_task filter; do
86
+ result_file="$attempt_dir/target-$target_index.jsonl"
87
+ log_file="$attempt_dir/target-$target_index.log"
88
+ set +e
89
+ automation_run_classified_focused_test "$gradle_task" "$filter" "$AUTOMATION_ROOT" "$result_file" "$log_file"
90
+ gradle_status=$?
91
+ set -e
92
+ if [[ "$gradle_status" -ne 0 ]]; then
93
+ execution_failures="$(jq -nc --argjson current "$execution_failures" --argjson target "$target_index" \
94
+ --arg task "$gradle_task" --arg filter "$filter" --argjson exitCode "$gradle_status" \
95
+ '$current + [{target: $target, gradleTask: $task, filter: $filter, exitCode: $exitCode}]')"
96
+ fi
97
+ if [[ -s "$result_file" ]]; then
98
+ while IFS= read -r result_line; do
99
+ if ! jq -e . >/dev/null 2>&1 <<< "$result_line"; then
100
+ execution_failures="$(jq -nc --argjson current "$execution_failures" --argjson target "$target_index" \
101
+ '$current + [{target: $target, reason: "invalid structured test output"}]')"
102
+ continue
103
+ fi
104
+ jq -c --argjson target "$target_index" '. + {target: $target}' <<< "$result_line" >> "$observed_file"
105
+ done < "$result_file"
106
+ else
107
+ execution_failures="$(jq -nc --argjson current "$execution_failures" --argjson target "$target_index" \
108
+ '$current + [{target: $target, reason: "no structured test results"}]')"
109
+ fi
110
+ target_index=$((target_index + 1))
111
+ done < <(jq -r '.targetTests[] | [.gradleTask, .filter] | @tsv' "$contract")
112
+
113
+ contract_sha="$(automation_file_sha256 "$contract")"
114
+ baseline_head="$(jq -er '.head' "$evidence_dir/baseline.json")"
115
+ test_diff_sha="$(automation_test_diff_sha "$task_id" "$AUTOMATION_ROOT")"
116
+ finished_at="$(automation_now)"
44
117
 
45
118
  jq -n \
119
+ --slurpfile contract "$contract" \
120
+ --slurpfile observed "$observed_file" \
46
121
  --arg taskId "$task_id" \
47
122
  --arg startedAt "$started_at" \
48
- --arg finishedAt "$(automation_now)" \
49
- --arg expectedFailure "$expected" \
50
- --arg gradleTask "$gradle_task" \
51
- --arg testFilter "$filter" \
52
- --argjson exitCode "$red_status" \
53
- '{taskId: $taskId, startedAt: $startedAt, finishedAt: $finishedAt, command: ["./gradlew", $gradleTask, "--tests", $testFilter], expectedFailure: $expectedFailure, exitCode: $exitCode}' \
54
- | automation_record_json "$red_meta"
55
-
56
- automation_info "$task_id RED evidence recorded"
123
+ --arg finishedAt "$finished_at" \
124
+ --arg contractSha256 "$contract_sha" \
125
+ --arg baselineHead "$baseline_head" \
126
+ --arg testDiffSha256 "$test_diff_sha" \
127
+ --argjson attempt "$attempt" \
128
+ --argjson executionFailures "$execution_failures" '
129
+ def failure_matches($decl; $actual):
130
+ ($actual.result == "FAILURE") and
131
+ ($actual.exceptionType == $decl.expectedFailure.type) and
132
+ (($decl.expectedFailure.messageIncludes // "") as $fragment |
133
+ ($fragment == "" or (($actual.exceptionMessage // "") | contains($fragment))));
134
+ ($contract[0].verification.cases | map(
135
+ . as $decl |
136
+ [$observed[] | select(.kind == "case" and .target == $decl.test.target and
137
+ .className == $decl.test.className and .name == $decl.test.name)] as $matches |
138
+ (if ($matches | length) == 1 then $matches[0] else null end) as $actual |
139
+ (if ($matches | length) != 1 then false
140
+ elif $decl.before == "pass" then $actual.result == "SUCCESS"
141
+ elif $decl.before == "fail" then failure_matches($decl; $actual)
142
+ elif $actual.result == "SUCCESS" then true
143
+ elif ($decl | has("expectedFailure")) then failure_matches($decl; $actual)
144
+ else false end) as $valid |
145
+ {id: $decl.id, criterion: $decl.criterion, intent: $decl.intent,
146
+ expectedBefore: $decl.before, test: $decl.test, expectedFailure: ($decl.expectedFailure // null),
147
+ matches: ($matches | length), actual: $actual, valid: $valid}
148
+ )) as $cases |
149
+ [$observed[] | select(.kind == "case") as $actual |
150
+ select(any($contract[0].verification.cases[];
151
+ .test.target == $actual.target and .test.className == $actual.className and .test.name == $actual.name) | not)] as $undeclared |
152
+ [$observed[] | select(.kind == "suite")] as $suites |
153
+ ($executionFailures | length == 0 and
154
+ ($suites | length) >= ($contract[0].targetTests | length) and
155
+ ($cases | all(.valid)) and ($undeclared | length == 0) and
156
+ ($cases | any(.intent == "change" and .actual.result == "FAILURE"))) as $valid |
157
+ {taskId: $taskId, attempt: $attempt, startedAt: $startedAt, finishedAt: $finishedAt,
158
+ contractSha256: $contractSha256, baselineHead: $baselineHead,
159
+ testDiffSha256: $testDiffSha256, valid: $valid,
160
+ reasonCode: (if $valid then "VALID_RED"
161
+ elif ($executionFailures | length) > 0 then "EXECUTION_FAILURE"
162
+ elif ($undeclared | length) > 0 then "UNDECLARED_TEST_RESULT"
163
+ else "CASE_EXPECTATION_MISMATCH" end),
164
+ summary: {declared: ($cases | length), valid: ([$cases[] | select(.valid)] | length),
165
+ invalid: ([$cases[] | select(.valid | not)] | length),
166
+ expectedRed: ([$cases[] | select(.intent == "change" and .actual.result == "FAILURE")] | length),
167
+ undeclared: ($undeclared | length)},
168
+ cases: $cases, undeclaredCases: $undeclared, suites: $suites,
169
+ executionFailures: $executionFailures}' | automation_record_json "$preflight_meta"
170
+
171
+ cp "$preflight_meta" "$attempt_dir/evaluation.json"
172
+ if [[ "$(jq -r '.valid' "$preflight_meta")" != "true" ]]; then
173
+ reason="$(jq -r '.reasonCode' "$preflight_meta")"
174
+ invalid="$(jq -r '[.cases[] | select(.valid | not) | .id] | join(", ")' "$preflight_meta")"
175
+ automation_die "RED preflight rejected ($reason); invalid cases: ${invalid:-none}; evidence: $preflight_meta"
176
+ fi
177
+
178
+ jq '{taskId, contractSha256, baselineHead, testDiffSha256,
179
+ cases: [.cases[] | {id, criterion, intent, test, actual}]}' "$preflight_meta" \
180
+ | automation_record_json "$evidence_dir/test-manifest.json"
181
+ preflight_sha="$(automation_file_sha256 "$preflight_meta")"
182
+ manifest_sha="$(automation_file_sha256 "$evidence_dir/test-manifest.json")"
183
+ jq -n --arg taskId "$task_id" --arg startedAt "$started_at" --arg finishedAt "$finished_at" \
184
+ --arg contractSha256 "$contract_sha" --arg baselineHead "$baseline_head" \
185
+ --arg testDiffSha256 "$test_diff_sha" --arg preflightSha256 "$preflight_sha" \
186
+ --arg manifestSha256 "$manifest_sha" --arg attemptPath "attempts/red-preflight-$attempt_id" \
187
+ --argjson attempt "$attempt" \
188
+ '{taskId: $taskId, schemaVersion: 3, startedAt: $startedAt, finishedAt: $finishedAt,
189
+ contractSha256: $contractSha256, baselineHead: $baselineHead,
190
+ testDiffSha256: $testDiffSha256, preflightSha256: $preflightSha256,
191
+ manifestSha256: $manifestSha256, attempt: $attempt, attemptPath: $attemptPath,
192
+ structuredCasesVerified: true, exitCode: 1}' | automation_record_json "$red_meta"
193
+
194
+ automation_info "$task_id structured RED evidence recorded"