npm - oh-my-opencode - Versions diffs - 4.3.1 → 4.5.0 - Mend

oh-my-opencode 4.3.1 → 4.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.

Files changed (222) hide show

package/.opencode/skills/work-with-pr/SKILL.md ADDED Viewed

@@ -0,0 +1,360 @@
+---
+name: work-with-pr
+description: "Full PR lifecycle: git worktree → implement → atomic commits → PR creation → verification loop (CI + review-work + Cubic approval) → merge. Keeps iterating until ALL gates pass and PR is merged. Worktree auto-cleanup after merge. Use whenever implementation work needs to land as a PR. Triggers: 'create a PR', 'implement and PR', 'work on this and make a PR', 'implement issue', 'land this as a PR', 'work-with-pr', 'PR workflow', 'implement end to end', even when user just says 'implement X' if the context implies PR delivery."
+---
+# Work With PR — Full PR Lifecycle
+You are executing a complete PR lifecycle: from isolated worktree setup through implementation, PR creation, and an unbounded verification loop until the PR is merged. The loop has three gates — CI, review-work, and Cubic — and you keep fixing and pushing until all three pass simultaneously.
+<architecture>
+```
+Phase 0: Setup         → Branch + worktree in sibling directory
+Phase 1: Implement     → Do the work, atomic commits
+Phase 2: PR Creation   → Push, create PR targeting dev
+Phase 3: Verify Loop   → Unbounded iteration until ALL gates pass:
+  ├─ Gate A: CI         → gh pr checks (bun test, typecheck, build)
+  ├─ Gate B: review-work → 5-agent parallel review
+  └─ Gate C: Cubic      → cubic-dev-ai[bot] "No issues found"
+Phase 4: Merge         → Squash merge, worktree cleanup
+```
+</architecture>
+---
+## Phase 0: Setup
+Create an isolated worktree so the user's main working directory stays clean. This matters because the user may have uncommitted work, and checking out a branch would destroy it.
+<setup>
+### 1. Resolve repository context
+```bash
+REPO=$(gh repo view --json nameWithOwner -q .nameWithOwner)
+REPO_NAME=$(basename "$PWD")
+BASE_BRANCH="dev"  # CI blocks PRs to master
+```
+### 2. Create branch
+If user provides a branch name, use it. Otherwise, derive from the task:
+```bash
+# Auto-generate: feature/short-description or fix/short-description
+BRANCH_NAME="feature/$(echo "$TASK_SUMMARY" | tr '[:upper:] ' '[:lower:]-' | head -c 50)"
+git fetch origin "$BASE_BRANCH"
+git branch "$BRANCH_NAME" "origin/$BASE_BRANCH"
+```
+### 3. Create worktree
+Place worktrees as siblings to the repo — not inside it. This avoids git nested repo issues and keeps the working tree clean.
+```bash
+WORKTREE_PATH="../${REPO_NAME}-wt/${BRANCH_NAME}"
+mkdir -p "$(dirname "$WORKTREE_PATH")"
+git worktree add "$WORKTREE_PATH" "$BRANCH_NAME"
+```
+### 4. Set working context
+All subsequent work happens inside the worktree. Install dependencies if needed:
+```bash
+cd "$WORKTREE_PATH"
+# If bun project:
+[ -f "bun.lock" ] && bun install
+```
+</setup>
+---
+## Phase 1: Implement
+Do the actual implementation work inside the worktree. The agent using this skill does the work directly — no subagent delegation for the implementation itself.
+**Scope discipline**: For bug fixes, stay minimal. Fix the bug, add a test for it, done. Do not refactor surrounding code, add config options, or "improve" things that aren't broken. The verification loop will catch regressions — trust the process.
+<implementation>
+### Commit strategy
+Use the git-master skill's atomic commit principles. The reason for atomic commits: if CI fails on one change, you can isolate and fix it without unwinding everything.
+```
+3+ files changed  → 2+ commits minimum
+5+ files changed  → 3+ commits minimum
+10+ files changed → 5+ commits minimum
+```
+Each commit should pair implementation with its tests. Load `git-master` skill when committing:
+```
+task(category="quick", load_skills=["git-master"], prompt="Commit the changes atomically following git-master conventions. Repository is at {WORKTREE_PATH}.")
+```
+### Pre-push local validation
+Before pushing, run the same checks CI will run. Catching failures locally saves a full CI round-trip (~3-5 min):
+```bash
+bun run typecheck
+bun test
+bun run build
+```
+Fix any failures before pushing. Each fix-commit cycle should be atomic.
+</implementation>
+---
+## Phase 2: PR Creation
+<pr_creation>
+### Push and create PR
+```bash
+git push -u origin "$BRANCH_NAME"
+```
+Create the PR using the project's template structure:
+```bash
+gh pr create \
+  --base "$BASE_BRANCH" \
+  --head "$BRANCH_NAME" \
+  --title "$PR_TITLE" \
+  --body "$(cat <<'EOF'
+## Summary
+[1-3 sentences describing what this PR does and why]
+## Changes
+[Bullet list of key changes]
+## Testing
+- `bun run typecheck` ✅
+- `bun test` ✅
+- `bun run build` ✅
+## Related Issues
+[Link to issue if applicable]
+EOF
+)"
+```
+Capture the PR number:
+```bash
+PR_NUMBER=$(gh pr view --json number -q .number)
+```
+</pr_creation>
+---
+## Phase 3: Verification Loop
+This is the core of the skill. Three gates must ALL pass for the PR to be ready. The loop has no iteration cap — keep going until done. Gate ordering is intentional: CI is cheapest/fastest, review-work is most thorough, Cubic is external and asynchronous.
+<verify_loop>
+```
+while true:
+  1. Wait for CI          → Gate A
+  2. If CI fails          → read logs, fix, commit, push, continue
+  3. Run review-work      → Gate B
+  4. If review fails      → fix blocking issues, commit, push, continue
+  5. Check Cubic          → Gate C
+  6. If Cubic has issues   → fix issues, commit, push, continue
+  7. All three pass       → break
+```
+### Gate A: CI Checks
+CI is the fastest feedback loop. Wait for it to complete, then parse results.
+```bash
+# Wait for checks to start (GitHub needs a moment after push)
+# Then watch for completion
+gh pr checks "$PR_NUMBER" --watch --fail-fast
+```
+**On failure**: Get the failed run logs to understand what broke:
+```bash
+# Find the failed run
+RUN_ID=$(gh run list --branch "$BRANCH_NAME" --status failure --json databaseId --jq '.[0].databaseId')
+# Get failed job logs
+gh run view "$RUN_ID" --log-failed
+```
+Read the logs, fix the issue, commit atomically, push, and re-enter the loop.
+### Gate B: review-work
+The review-work skill launches 5 parallel sub-agents (goal verification, QA, code quality, security, context mining). All 5 must pass.
+Invoke review-work after CI passes — there's no point reviewing code that doesn't build:
+```
+task(
+  category="unspecified-high",
+  load_skills=["review-work"],
+  run_in_background=false,
+  description="Post-implementation review of PR changes",
+  prompt="Review the implementation work on branch {BRANCH_NAME}. The worktree is at {WORKTREE_PATH}. Goal: {ORIGINAL_GOAL}. Constraints: {CONSTRAINTS}. Run command: bun run dev (or as appropriate)."
+)
+```
+**On failure**: review-work reports blocking issues with specific files and line numbers. Fix each blocking issue, commit, push, and re-enter the loop from Gate A (since code changed, CI must re-run).
+### Gate C: Cubic Approval
+Cubic (`cubic-dev-ai[bot]`) is an automated review bot that comments on PRs. It does NOT use GitHub's APPROVED review state — instead it posts comments with issue counts and confidence scores.
+**Approval signal**: The latest Cubic comment contains `**No issues found**` and confidence `**5/5**`.
+**Issue signal**: The comment lists issues with file-level detail.
+```bash
+# Get the latest Cubic review
+CUBIC_REVIEW=$(gh api "repos/${REPO}/pulls/${PR_NUMBER}/reviews" \
+  --jq '[.[] | select(.user.login == "cubic-dev-ai[bot]")] | last | .body')
+# Check if approved
+if echo "$CUBIC_REVIEW" | grep -q "No issues found"; then
+  echo "Cubic: APPROVED"
+else
+  echo "Cubic: ISSUES FOUND"
+  echo "$CUBIC_REVIEW"
+fi
+```
+**On issues**: Cubic's review body contains structured issue descriptions. Parse them, determine which are valid (some may be false positives), fix the valid ones, commit, push, re-enter from Gate A.
+Cubic reviews are triggered automatically on PR updates. After pushing a fix, wait for the new review to appear before checking again. Use `gh api` polling with a conditional loop:
+```bash
+# Wait for new Cubic review after push
+PUSH_TIME=$(date -u +%Y-%m-%dT%H:%M:%SZ)
+while true; do
+  LATEST_REVIEW_TIME=$(gh api "repos/${REPO}/pulls/${PR_NUMBER}/reviews" \
+    --jq '[.[] | select(.user.login == "cubic-dev-ai[bot]")] | last | .submitted_at')
+  if [[ "$LATEST_REVIEW_TIME" > "$PUSH_TIME" ]]; then
+    break
+  fi
+  # Use gh api call itself as the delay mechanism — each call takes ~1-2s
+  # For longer waits, use: timeout 30 gh pr checks "$PR_NUMBER" --watch 2>/dev/null || true
+done
+```
+### Iteration discipline
+Each iteration through the loop:
+1. Fix ONLY the issues identified by the failing gate
+2. Commit atomically (one logical fix per commit)
+3. Push
+4. Re-enter from Gate A (code changed → full re-verification)
+Avoid the temptation to "improve" unrelated code during fix iterations. Scope creep in the fix loop makes debugging harder and can introduce new failures.
+</verify_loop>
+---
+## Phase 4: Merge & Cleanup
+Once all three gates pass:
+<merge_cleanup>
+### Merge the PR
+```bash
+# Squash merge to keep history clean
+gh pr merge "$PR_NUMBER" --squash --delete-branch
+```
+### Sync .omo state back to main repo
+Before removing the worktree, copy `.omo/` state back. When `.omo/` is gitignored, files written there during worktree execution are not committed or merged — they would be lost on worktree removal.
+```bash
+# Sync .omo state from worktree to main repo (preserves task state, plans, notepads)
+if [ -d "$WORKTREE_PATH/.omo" ]; then
+  mkdir -p "$ORIGINAL_DIR/.omo"
+  cp -r "$WORKTREE_PATH/.omo/"* "$ORIGINAL_DIR/.omo/" 2>/dev/null || true
+fi
+```
+### Clean up the worktree
+The worktree served its purpose — remove it to avoid disk bloat:
+```bash
+cd "$ORIGINAL_DIR"  # Return to original working directory
+git worktree remove "$WORKTREE_PATH"
+# Prune any stale worktree references
+git worktree prune
+```
+### Report completion
+Summarize what happened:
+```
+## PR Merged ✅
+- **PR**: #{PR_NUMBER} — {PR_TITLE}
+- **Branch**: {BRANCH_NAME} → {BASE_BRANCH}
+- **Iterations**: {N} verification loops
+- **Gates passed**: CI ✅ | review-work ✅ | Cubic ✅
+- **Worktree**: cleaned up
+```
+</merge_cleanup>
+---
+## Failure Recovery
+<failure_recovery>
+If you hit an unrecoverable error (e.g., merge conflict with base branch, infrastructure failure):
+1. **Do NOT delete the worktree** — the user may want to inspect or continue manually
+2. Report what happened, what was attempted, and where things stand
+3. Include the worktree path so the user can resume
+For merge conflicts:
+```bash
+cd "$WORKTREE_PATH"
+git fetch origin "$BASE_BRANCH"
+git rebase "origin/$BASE_BRANCH"
+# Resolve conflicts, then continue the loop
+```
+</failure_recovery>
+---
+## Anti-Patterns
+| Violation | Why it fails | Severity |
+|-----------|-------------|----------|
+| Working in main worktree instead of isolated worktree | Pollutes user's working directory, may destroy uncommitted work | CRITICAL |
+| Pushing directly to dev/master | Bypasses review entirely | CRITICAL |
+| Skipping CI gate after code changes | review-work and Cubic may pass on stale code | CRITICAL |
+| Fixing unrelated code during verification loop | Scope creep causes new failures | HIGH |
+| Deleting worktree on failure | User loses ability to inspect/resume | HIGH |
+| Ignoring Cubic false positives without justification | Cubic issues should be evaluated, not blindly dismissed | MEDIUM |
+| Giant single commits | Harder to isolate failures, violates git-master principles | MEDIUM |
+| Not running local checks before push | Wastes CI time on obvious failures | MEDIUM |

package/.opencode/skills/work-with-pr-workspace/evals/evals.json ADDED Viewed

@@ -0,0 +1,76 @@
+{
+  "skill_name": "work-with-pr",
+  "evals": [
+    {
+      "id": 1,
+      "prompt": "I need to add a `max_background_agents` config option to oh-my-opencode that limits how many background agents can run simultaneously. It should be in the plugin config schema with a default of 5. Add validation and make sure the background manager respects it. Create a PR for this.",
+      "expected_output": "Agent creates worktree, implements config option with schema validation, adds tests, creates PR, iterates through verification gates until merged",
+      "files": [],
+      "assertions": [
+        {"id": "worktree-isolation", "text": "Plan uses git worktree in a sibling directory (not main working directory)"},
+        {"id": "branch-from-dev", "text": "Branch is created from origin/dev (not master/main)"},
+        {"id": "atomic-commits", "text": "Plan specifies multiple atomic commits for multi-file changes"},
+        {"id": "local-validation", "text": "Runs bun run typecheck, bun test, and bun run build before pushing"},
+        {"id": "pr-targets-dev", "text": "PR is created targeting dev branch (not master)"},
+        {"id": "three-gates", "text": "Verification loop includes all 3 gates: CI, review-work, and Cubic"},
+        {"id": "gate-ordering", "text": "Gates are checked in order: CI first, then review-work, then Cubic"},
+        {"id": "cubic-check-method", "text": "Cubic check uses gh api to check cubic-dev-ai[bot] reviews for 'No issues found'"},
+        {"id": "worktree-cleanup", "text": "Plan includes worktree cleanup after merge"},
+        {"id": "real-file-references", "text": "Code changes reference actual files in the codebase (config schema, background manager)"}
+      ]
+    },
+    {
+      "id": 2,
+      "prompt": "The atlas hook has a bug where it crashes when boulder.json is missing the worktree_path field. Fix it and land the fix as a PR. Make sure CI passes.",
+      "expected_output": "Agent creates worktree for the fix branch, adds null check and test for missing worktree_path, creates PR, iterates verification loop",
+      "files": [],
+      "assertions": [
+        {"id": "worktree-isolation", "text": "Plan uses git worktree in a sibling directory"},
+        {"id": "minimal-fix", "text": "Fix is minimal — adds null check, doesn't refactor unrelated code"},
+        {"id": "test-added", "text": "Test case added for the missing worktree_path scenario"},
+        {"id": "three-gates", "text": "Verification loop includes all 3 gates: CI, review-work, Cubic"},
+        {"id": "real-atlas-files", "text": "References actual atlas hook files in src/hooks/atlas/"},
+        {"id": "fix-branch-naming", "text": "Branch name follows fix/ prefix convention"}
+      ]
+    },
+    {
+      "id": 3,
+      "prompt": "Refactor src/tools/delegate-task/constants.ts to split DEFAULT_CATEGORIES and CATEGORY_MODEL_REQUIREMENTS into separate files. Keep backward compatibility with the barrel export. Make a PR.",
+      "expected_output": "Agent creates worktree, splits file with atomic commits, ensures imports still work via barrel, creates PR, runs through all gates",
+      "files": [],
+      "assertions": [
+        {"id": "worktree-isolation", "text": "Plan uses git worktree in a sibling directory"},
+        {"id": "multiple-atomic-commits", "text": "Uses 2+ commits for the multi-file refactor"},
+        {"id": "barrel-export", "text": "Maintains backward compatibility via barrel re-export in constants.ts or index.ts"},
+        {"id": "three-gates", "text": "Verification loop includes all 3 gates"},
+        {"id": "real-constants-file", "text": "References actual src/tools/delegate-task/constants.ts file and its exports"}
+      ]
+    },
+    {
+      "id": 4,
+      "prompt": "implement issue #100 - we need to add a new built-in MCP for arxiv paper search. just the basic search endpoint, nothing fancy. pr it",
+      "expected_output": "Agent creates worktree, implements arxiv MCP following existing MCP patterns (websearch, context7, grep_app), creates PR with proper template, verification loop runs",
+      "files": [],
+      "assertions": [
+        {"id": "worktree-isolation", "text": "Plan uses git worktree in a sibling directory"},
+        {"id": "follows-mcp-pattern", "text": "New MCP follows existing pattern from src/mcp/ (websearch, context7, grep_app)"},
+        {"id": "three-gates", "text": "Verification loop includes all 3 gates"},
+        {"id": "pr-targets-dev", "text": "PR targets dev branch"},
+        {"id": "local-validation", "text": "Runs local checks before pushing"}
+      ]
+    },
+    {
+      "id": 5,
+      "prompt": "The comment-checker hook is too aggressive - it's flagging legitimate comments that happen to contain 'Note:' as AI slop. Relax the regex pattern and add test cases for the false positives. Work on a separate branch and make a PR.",
+      "expected_output": "Agent creates worktree, fixes regex, adds specific test cases for false positive scenarios, creates PR, all three gates pass",
+      "files": [],
+      "assertions": [
+        {"id": "worktree-isolation", "text": "Plan uses git worktree in a sibling directory"},
+        {"id": "real-comment-checker-files", "text": "References actual comment-checker hook files in the codebase"},
+        {"id": "regression-tests", "text": "Adds test cases specifically for 'Note:' false positive scenarios"},
+        {"id": "three-gates", "text": "Verification loop includes all 3 gates"},
+        {"id": "minimal-change", "text": "Only modifies regex and adds tests — no unrelated changes"}
+      ]
+    }
+  ]
+}

package/.opencode/skills/work-with-pr-workspace/iteration-1/benchmark.json ADDED Viewed

@@ -0,0 +1,138 @@
+{
+  "skill_name": "work-with-pr",
+  "iteration": 1,
+  "summary": {
+    "with_skill": {
+      "pass_rate": 0.968,
+      "mean_duration_seconds": 340.2,
+      "stddev_duration_seconds": 169.3
+    },
+    "without_skill": {
+      "pass_rate": 0.516,
+      "mean_duration_seconds": 303.0,
+      "stddev_duration_seconds": 77.8
+    },
+    "delta": {
+      "pass_rate": 0.452,
+      "mean_duration_seconds": 37.2,
+      "stddev_duration_seconds": 91.5
+    }
+  },
+  "evals": [
+    {
+      "eval_name": "happy-path-feature-config-option",
+      "with_skill": {
+        "pass_rate": 1.0,
+        "passed": 10,
+        "total": 10,
+        "duration_seconds": 292,
+        "failed_assertions": []
+      },
+      "without_skill": {
+        "pass_rate": 0.4,
+        "passed": 4,
+        "total": 10,
+        "duration_seconds": 365,
+        "failed_assertions": [
+          {"assertion": "Plan uses git worktree in a sibling directory", "reason": "Uses git checkout -b, no worktree isolation"},
+          {"assertion": "Plan specifies multiple atomic commits for multi-file changes", "reason": "Steps listed sequentially but no atomic commit strategy mentioned"},
+          {"assertion": "Verification loop includes all 3 gates: CI, review-work, and Cubic", "reason": "Only mentions CI pipeline in step 6. No review-work or Cubic."},
+          {"assertion": "Gates are checked in order: CI first, then review-work, then Cubic", "reason": "No gate ordering - only CI mentioned"},
+          {"assertion": "Cubic check uses gh api to check cubic-dev-ai[bot] reviews", "reason": "No mention of Cubic at all"},
+          {"assertion": "Plan includes worktree cleanup after merge", "reason": "No worktree used, no cleanup needed"}
+        ]
+      }
+    },
+    {
+      "eval_name": "bugfix-atlas-null-check",
+      "with_skill": {
+        "pass_rate": 1.0,
+        "passed": 6,
+        "total": 6,
+        "duration_seconds": 506,
+        "failed_assertions": []
+      },
+      "without_skill": {
+        "pass_rate": 0.667,
+        "passed": 4,
+        "total": 6,
+        "duration_seconds": 325,
+        "failed_assertions": [
+          {"assertion": "Plan uses git worktree in a sibling directory", "reason": "No worktree. Steps go directly to creating branch and modifying files."},
+          {"assertion": "Verification loop includes all 3 gates", "reason": "Only mentions CI pipeline (step 5). No review-work or Cubic."}
+        ]
+      }
+    },
+    {
+      "eval_name": "refactor-split-constants",
+      "with_skill": {
+        "pass_rate": 1.0,
+        "passed": 5,
+        "total": 5,
+        "duration_seconds": 181,
+        "failed_assertions": []
+      },
+      "without_skill": {
+        "pass_rate": 0.4,
+        "passed": 2,
+        "total": 5,
+        "duration_seconds": 229,
+        "failed_assertions": [
+          {"assertion": "Plan uses git worktree in a sibling directory", "reason": "git checkout -b only, no worktree"},
+          {"assertion": "Uses 2+ commits for the multi-file refactor", "reason": "Single atomic commit: 'refactor: split delegate-task constants and category model requirements'"},
+          {"assertion": "Verification loop includes all 3 gates", "reason": "Only mentions typecheck/test/build. No review-work or Cubic."}
+        ]
+      }
+    },
+    {
+      "eval_name": "new-mcp-arxiv-casual",
+      "with_skill": {
+        "pass_rate": 1.0,
+        "passed": 5,
+        "total": 5,
+        "duration_seconds": 152,
+        "failed_assertions": []
+      },
+      "without_skill": {
+        "pass_rate": 0.6,
+        "passed": 3,
+        "total": 5,
+        "duration_seconds": 197,
+        "failed_assertions": [
+          {"assertion": "Verification loop includes all 3 gates", "reason": "Only mentions bun test/typecheck/build. No review-work or Cubic."}
+        ]
+      }
+    },
+    {
+      "eval_name": "regex-fix-false-positive",
+      "with_skill": {
+        "pass_rate": 0.8,
+        "passed": 4,
+        "total": 5,
+        "duration_seconds": 570,
+        "failed_assertions": [
+          {"assertion": "Only modifies regex and adds tests — no unrelated changes", "reason": "Also proposes config schema change (exclude_patterns) and Go binary update — goes beyond minimal fix"}
+        ]
+      },
+      "without_skill": {
+        "pass_rate": 0.6,
+        "passed": 3,
+        "total": 5,
+        "duration_seconds": 399,
+        "failed_assertions": [
+          {"assertion": "Plan uses git worktree in a sibling directory", "reason": "git checkout -b, no worktree"},
+          {"assertion": "Verification loop includes all 3 gates", "reason": "Only bun test and typecheck. No review-work or Cubic."}
+        ]
+      }
+    }
+  ],
+  "analyst_observations": [
+    "Three-gates assertion (CI + review-work + Cubic) is the strongest discriminator: 5/5 with-skill vs 0/5 without-skill. Without the skill, agents never know about Cubic or review-work gates.",
+    "Worktree isolation is nearly as discriminating (5/5 vs 1/5). One without-skill run (eval-4) independently chose worktree, suggesting some agents already know worktree patterns, but the skill makes it consistent.",
+    "The skill's only failure (eval-5 minimal-change) reveals a potential over-engineering tendency: the skill-guided agent proposed config schema changes and Go binary updates for what should have been a minimal regex fix. Consider adding explicit guidance for fix-type tasks to stay minimal.",
+    "Duration tradeoff: with-skill is 12% slower on average (340s vs 303s), driven mainly by eval-2 (bugfix) and eval-5 (regex fix) where the skill's thorough verification planning adds overhead. For eval-1 and eval-3-4, with-skill was actually faster.",
+    "Without-skill duration has lower variance (stddev 78s vs 169s), suggesting the skill introduces more variable execution paths depending on task complexity.",
+    "Non-discriminating assertions: 'References actual files', 'PR targets dev', 'Runs local checks' — these pass regardless of skill. They validate baseline agent competence, not skill value. Consider removing or downweighting in future iterations.",
+    "Atomic commits assertion discriminates moderately (2/2 with-skill tested vs 0/2 without-skill tested). Without the skill, agents default to single commits even for multi-file refactors."
+  ]
+}

package/.opencode/skills/work-with-pr-workspace/iteration-1/benchmark.md ADDED Viewed

@@ -0,0 +1,42 @@
+# Benchmark: work-with-pr (Iteration 1)
+## Summary
+| Metric | With Skill | Without Skill | Delta |
+|--------|-----------|---------------|-------|
+| Pass Rate | 96.8% (30/31) | 51.6% (16/31) | +45.2% |
+| Mean Duration | 340.2s | 303.0s | +37.2s |
+| Duration Stddev | 169.3s | 77.8s | +91.5s |
+## Per-Eval Breakdown
+| Eval | With Skill | Without Skill | Delta |
+|------|-----------|---------------|-------|
+| happy-path-feature-config-option | 100% (10/10) | 40% (4/10) | +60% |
+| bugfix-atlas-null-check | 100% (6/6) | 67% (4/6) | +33% |
+| refactor-split-constants | 100% (5/5) | 40% (2/5) | +60% |
+| new-mcp-arxiv-casual | 100% (5/5) | 60% (3/5) | +40% |
+| regex-fix-false-positive | 80% (4/5) | 60% (3/5) | +20% |
+## Key Discriminators
+- **three-gates** (CI + review-work + Cubic): 5/5 vs 0/5 — strongest signal
+- **worktree-isolation**: 5/5 vs 1/5
+- **atomic-commits**: 2/2 vs 0/2
+- **cubic-check-method**: 1/1 vs 0/1
+## Non-Discriminating Assertions
+- References actual files: passes in both conditions
+- PR targets dev: passes in both conditions
+- Runs local checks before pushing: passes in both conditions
+## Only With-Skill Failure
+- **eval-5 minimal-change**: Skill-guided agent proposed config schema changes and Go binary update for a minimal regex fix. The skill may encourage over-engineering in fix scenarios.
+## Analyst Notes
+- The skill adds most value for procedural knowledge (verification gates, worktree workflow) that agents cannot infer from codebase alone.
+- Duration cost is modest (+12%) and acceptable given the +45% pass rate improvement.
+- Consider adding explicit "fix-type tasks: stay minimal" guidance in iteration 2.

package/.opencode/skills/work-with-pr-workspace/iteration-1/eval-1/eval_metadata.json ADDED Viewed

@@ -0,0 +1,57 @@
+{
+  "eval_id": 1,
+  "eval_name": "happy-path-feature-config-option",
+  "prompt": "I need to add a `max_background_agents` config option to oh-my-opencode that limits how many background agents can run simultaneously. It should be in the plugin config schema with a default of 5. Add validation and make sure the background manager respects it. Create a PR for this.",
+  "assertions": [
+    {
+      "id": "worktree-isolation",
+      "text": "Plan uses git worktree in a sibling directory (not main working directory)",
+      "type": "manual"
+    },
+    {
+      "id": "branch-from-dev",
+      "text": "Branch is created from origin/dev (not master/main)",
+      "type": "manual"
+    },
+    {
+      "id": "atomic-commits",
+      "text": "Plan specifies multiple atomic commits for multi-file changes",
+      "type": "manual"
+    },
+    {
+      "id": "local-validation",
+      "text": "Runs bun run typecheck, bun test, and bun run build before pushing",
+      "type": "manual"
+    },
+    {
+      "id": "pr-targets-dev",
+      "text": "PR is created targeting dev branch (not master)",
+      "type": "manual"
+    },
+    {
+      "id": "three-gates",
+      "text": "Verification loop includes all 3 gates: CI, review-work, and Cubic",
+      "type": "manual"
+    },
+    {
+      "id": "gate-ordering",
+      "text": "Gates are checked in order: CI first, then review-work, then Cubic",
+      "type": "manual"
+    },
+    {
+      "id": "cubic-check-method",
+      "text": "Cubic check uses gh api to check cubic-dev-ai[bot] reviews for 'No issues found'",
+      "type": "manual"
+    },
+    {
+      "id": "worktree-cleanup",
+      "text": "Plan includes worktree cleanup after merge",
+      "type": "manual"
+    },
+    {
+      "id": "real-file-references",
+      "text": "Code changes reference actual files in the codebase (config schema, background manager)",
+      "type": "manual"
+    }
+  ]
+}