@caddis/cli 0.0.0 → 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (153) hide show
  1. package/LICENSE +21 -0
  2. package/README.md +150 -1
  3. package/bundles/antigravity-plugin/agents/code-reviewer.md +47 -0
  4. package/bundles/antigravity-plugin/agents/preflight.md +53 -0
  5. package/bundles/antigravity-plugin/guard_agy.py +338 -0
  6. package/bundles/antigravity-plugin/hooks.json +36 -0
  7. package/bundles/antigravity-plugin/mcp_config.json +8 -0
  8. package/bundles/antigravity-plugin/mcp_ping_server.py +55 -0
  9. package/bundles/antigravity-plugin/plugin.json +5 -0
  10. package/bundles/antigravity-plugin/session_end_agy.py +57 -0
  11. package/bundles/antigravity-plugin/skills/_registry.md +115 -0
  12. package/bundles/antigravity-plugin/skills/add-rules/SKILL.md +45 -0
  13. package/bundles/antigravity-plugin/skills/api-design/SKILL.md +249 -0
  14. package/bundles/antigravity-plugin/skills/backend-development/SKILL.md +305 -0
  15. package/bundles/antigravity-plugin/skills/best-practices/SKILL.md +500 -0
  16. package/bundles/antigravity-plugin/skills/best-practices/agents/best-practices-referencer.md +263 -0
  17. package/bundles/antigravity-plugin/skills/best-practices/agents/codebase-context-builder.md +326 -0
  18. package/bundles/antigravity-plugin/skills/best-practices/agents/task-intent-analyzer.md +245 -0
  19. package/bundles/antigravity-plugin/skills/best-practices/references/anti-patterns.md +571 -0
  20. package/bundles/antigravity-plugin/skills/best-practices/references/before-after-examples.md +1114 -0
  21. package/bundles/antigravity-plugin/skills/best-practices/references/best-practices-guide.md +513 -0
  22. package/bundles/antigravity-plugin/skills/best-practices/references/common-workflows.md +692 -0
  23. package/bundles/antigravity-plugin/skills/best-practices/references/prompt-patterns.md +547 -0
  24. package/bundles/antigravity-plugin/skills/brainstorming/SKILL.md +57 -0
  25. package/bundles/antigravity-plugin/skills/ci-cd-pipeline/SKILL.md +315 -0
  26. package/bundles/antigravity-plugin/skills/code-documentation/SKILL.md +271 -0
  27. package/bundles/antigravity-plugin/skills/code-review/SKILL.md +122 -0
  28. package/bundles/antigravity-plugin/skills/codebase-audit/SKILL.md +204 -0
  29. package/bundles/antigravity-plugin/skills/context-curator/SKILL.md +157 -0
  30. package/bundles/antigravity-plugin/skills/cross-review/SKILL.md +40 -0
  31. package/bundles/antigravity-plugin/skills/css-architecture/SKILL.md +305 -0
  32. package/bundles/antigravity-plugin/skills/css-architecture/references/RESPONSIVE-DESIGN.md +604 -0
  33. package/bundles/antigravity-plugin/skills/database-design/SKILL.md +177 -0
  34. package/bundles/antigravity-plugin/skills/db-diagram/SKILL.md +148 -0
  35. package/bundles/antigravity-plugin/skills/db-diagram/scripts/sql_to_graph.py +1212 -0
  36. package/bundles/antigravity-plugin/skills/digress/SKILL.md +61 -0
  37. package/bundles/antigravity-plugin/skills/draw-io/SKILL.md +162 -0
  38. package/bundles/antigravity-plugin/skills/draw-io/references/aws-icons.md +677 -0
  39. package/bundles/antigravity-plugin/skills/draw-io/references/layout-guidelines.md +142 -0
  40. package/bundles/antigravity-plugin/skills/draw-io/references/troubleshooting.md +118 -0
  41. package/bundles/antigravity-plugin/skills/draw-io/references/workflows.md +103 -0
  42. package/bundles/antigravity-plugin/skills/draw-io/scripts/convert-drawio-to-png.sh +25 -0
  43. package/bundles/antigravity-plugin/skills/draw-io/scripts/find_aws_icon.py +79 -0
  44. package/bundles/antigravity-plugin/skills/error-handling/SKILL.md +260 -0
  45. package/bundles/antigravity-plugin/skills/excalidraw-db/SKILL.md +38 -0
  46. package/bundles/antigravity-plugin/skills/fastapi-dev/SKILL.md +300 -0
  47. package/bundles/antigravity-plugin/skills/feature-plan/SKILL.md +198 -0
  48. package/bundles/antigravity-plugin/skills/frontend-design/SKILL.md +193 -0
  49. package/bundles/antigravity-plugin/skills/frontend-design/references/my-tech-stack.md +127 -0
  50. package/bundles/antigravity-plugin/skills/gh-cli/SKILL.md +195 -0
  51. package/bundles/antigravity-plugin/skills/git-commit/SKILL.md +285 -0
  52. package/bundles/antigravity-plugin/skills/golden-plan/SKILL.md +577 -0
  53. package/bundles/antigravity-plugin/skills/handoff/SKILL.md +100 -0
  54. package/bundles/antigravity-plugin/skills/implement/SKILL.md +114 -0
  55. package/bundles/antigravity-plugin/skills/javascript-typescript/SKILL.md +142 -0
  56. package/bundles/antigravity-plugin/skills/kb/SKILL.md +60 -0
  57. package/bundles/antigravity-plugin/skills/mermaid-db/SKILL.md +33 -0
  58. package/bundles/antigravity-plugin/skills/mermaid-diagrams/SKILL.md +236 -0
  59. package/bundles/antigravity-plugin/skills/mermaid-diagrams/references/ENHANCEMENTS.md +264 -0
  60. package/bundles/antigravity-plugin/skills/mermaid-diagrams/references/MERMAID-SUMMARY.md +137 -0
  61. package/bundles/antigravity-plugin/skills/mermaid-diagrams/references/advanced-features.md +556 -0
  62. package/bundles/antigravity-plugin/skills/mermaid-diagrams/references/architecture-diagrams.md +192 -0
  63. package/bundles/antigravity-plugin/skills/mermaid-diagrams/references/c4-diagrams.md +410 -0
  64. package/bundles/antigravity-plugin/skills/mermaid-diagrams/references/class-diagrams.md +361 -0
  65. package/bundles/antigravity-plugin/skills/mermaid-diagrams/references/erd-diagrams.md +510 -0
  66. package/bundles/antigravity-plugin/skills/mermaid-diagrams/references/flowcharts.md +450 -0
  67. package/bundles/antigravity-plugin/skills/mermaid-diagrams/references/sequence-diagrams.md +394 -0
  68. package/bundles/antigravity-plugin/skills/mermaid-diagrams/references/troubleshooting.md +335 -0
  69. package/bundles/antigravity-plugin/skills/mermaid-diagrams/references/workflows.md +418 -0
  70. package/bundles/antigravity-plugin/skills/migrate-dir/SKILL.md +68 -0
  71. package/bundles/antigravity-plugin/skills/mockup/SKILL.md +242 -0
  72. package/bundles/antigravity-plugin/skills/particle-art/SKILL.md +243 -0
  73. package/bundles/antigravity-plugin/skills/particle-art/references/canvas-utils.ts +171 -0
  74. package/bundles/antigravity-plugin/skills/particle-art/references/dot-field.template.tsx +203 -0
  75. package/bundles/antigravity-plugin/skills/particle-art/references/flow-field.template.tsx +263 -0
  76. package/bundles/antigravity-plugin/skills/particle-art/references/node-shape.template.tsx +261 -0
  77. package/bundles/antigravity-plugin/skills/particle-art/references/shape-sampler.ts +281 -0
  78. package/bundles/antigravity-plugin/skills/particle-art/references/stipple-morph.template.tsx +167 -0
  79. package/bundles/antigravity-plugin/skills/particle-art/references/stipple.template.tsx +175 -0
  80. package/bundles/antigravity-plugin/skills/particle-art/references/trail-ghost.template.tsx +266 -0
  81. package/bundles/antigravity-plugin/skills/particle-art/references/usage-examples.md +320 -0
  82. package/bundles/antigravity-plugin/skills/playwright/API_REFERENCE.md +653 -0
  83. package/bundles/antigravity-plugin/skills/playwright/SKILL.md +454 -0
  84. package/bundles/antigravity-plugin/skills/playwright/lib/helpers.js +441 -0
  85. package/bundles/antigravity-plugin/skills/playwright/package.json +26 -0
  86. package/bundles/antigravity-plugin/skills/playwright/run.js +228 -0
  87. package/bundles/antigravity-plugin/skills/prd/SKILL.md +107 -0
  88. package/bundles/antigravity-plugin/skills/preflight/SKILL.md +435 -0
  89. package/bundles/antigravity-plugin/skills/python/SKILL.md +388 -0
  90. package/bundles/antigravity-plugin/skills/react-best-practices/SKILL.md +269 -0
  91. package/bundles/antigravity-plugin/skills/react-dev/README.md +404 -0
  92. package/bundles/antigravity-plugin/skills/react-dev/SKILL.md +459 -0
  93. package/bundles/antigravity-plugin/skills/react-dev/examples/generic-components.md +579 -0
  94. package/bundles/antigravity-plugin/skills/react-dev/examples/server-components.md +579 -0
  95. package/bundles/antigravity-plugin/skills/react-dev/references/event-handlers.md +574 -0
  96. package/bundles/antigravity-plugin/skills/react-dev/references/hooks.md +456 -0
  97. package/bundles/antigravity-plugin/skills/react-dev/references/react-19-patterns.md +638 -0
  98. package/bundles/antigravity-plugin/skills/react-dev/references/react-router.md +1002 -0
  99. package/bundles/antigravity-plugin/skills/react-dev/references/tanstack-router.md +587 -0
  100. package/bundles/antigravity-plugin/skills/refactoring/SKILL.md +486 -0
  101. package/bundles/antigravity-plugin/skills/resume/SKILL.md +36 -0
  102. package/bundles/antigravity-plugin/skills/security-review/SKILL.md +196 -0
  103. package/bundles/antigravity-plugin/skills/setup-project-ai/SKILL.md +61 -0
  104. package/bundles/antigravity-plugin/skills/ship/SKILL.md +107 -0
  105. package/bundles/antigravity-plugin/skills/ship-merge/SKILL.md +103 -0
  106. package/bundles/antigravity-plugin/skills/ship-pr/SKILL.md +102 -0
  107. package/bundles/antigravity-plugin/skills/skill-creator/LICENSE.txt +202 -0
  108. package/bundles/antigravity-plugin/skills/skill-creator/SKILL.md +491 -0
  109. package/bundles/antigravity-plugin/skills/skill-creator/agents/analyzer.md +274 -0
  110. package/bundles/antigravity-plugin/skills/skill-creator/agents/comparator.md +202 -0
  111. package/bundles/antigravity-plugin/skills/skill-creator/agents/grader.md +223 -0
  112. package/bundles/antigravity-plugin/skills/skill-creator/assets/eval_review.html +146 -0
  113. package/bundles/antigravity-plugin/skills/skill-creator/eval-viewer/generate_review.py +471 -0
  114. package/bundles/antigravity-plugin/skills/skill-creator/eval-viewer/viewer.html +1325 -0
  115. package/bundles/antigravity-plugin/skills/skill-creator/references/schemas.md +430 -0
  116. package/bundles/antigravity-plugin/skills/skill-creator/scripts/__init__.py +0 -0
  117. package/bundles/antigravity-plugin/skills/skill-creator/scripts/aggregate_benchmark.py +401 -0
  118. package/bundles/antigravity-plugin/skills/skill-creator/scripts/generate_report.py +326 -0
  119. package/bundles/antigravity-plugin/skills/skill-creator/scripts/improve_description.py +247 -0
  120. package/bundles/antigravity-plugin/skills/skill-creator/scripts/package_skill.py +136 -0
  121. package/bundles/antigravity-plugin/skills/skill-creator/scripts/quick_validate.py +103 -0
  122. package/bundles/antigravity-plugin/skills/skill-creator/scripts/run_eval.py +310 -0
  123. package/bundles/antigravity-plugin/skills/skill-creator/scripts/run_loop.py +328 -0
  124. package/bundles/antigravity-plugin/skills/skill-creator/scripts/utils.py +47 -0
  125. package/bundles/antigravity-plugin/skills/sql/SKILL.md +321 -0
  126. package/bundles/antigravity-plugin/skills/tdd/SKILL.md +37 -0
  127. package/bundles/antigravity-plugin/skills/tdd-workflow/SKILL.md +188 -0
  128. package/bundles/antigravity-plugin/skills/technical-writing/SKILL.md +286 -0
  129. package/bundles/antigravity-plugin/skills/test-strategy/SKILL.md +155 -0
  130. package/bundles/antigravity-plugin/skills/ui-brief/SKILL.md +84 -0
  131. package/bundles/antigravity-plugin/skills/ui-review/SKILL.md +176 -0
  132. package/bundles/antigravity-plugin/skills/ui-review/references/framework-fixes.md +471 -0
  133. package/bundles/antigravity-plugin/skills/ui-review/references/visual-checklist.md +236 -0
  134. package/bundles/antigravity-plugin/skills/usage-review/SKILL.md +77 -0
  135. package/bundles/antigravity-plugin/skills/use-model/SKILL.md +64 -0
  136. package/bundles/antigravity-plugin/skills/using-git-worktrees/SKILL.md +217 -0
  137. package/bundles/antigravity-plugin/skills/version/SKILL.md +18 -0
  138. package/bundles/antigravity-plugin/skills/warm-editorial-ui/DESIGN_TOKENS.md +487 -0
  139. package/bundles/antigravity-plugin/skills/warm-editorial-ui/IMPLEMENTATION_GUIDE.md +177 -0
  140. package/bundles/antigravity-plugin/skills/warm-editorial-ui/SKILL.md +732 -0
  141. package/bundles/antigravity-plugin/skills/webapp-testing/LICENSE.txt +202 -0
  142. package/bundles/antigravity-plugin/skills/webapp-testing/SKILL.md +97 -0
  143. package/bundles/antigravity-plugin/skills/webapp-testing/examples/console_logging.py +35 -0
  144. package/bundles/antigravity-plugin/skills/webapp-testing/examples/element_discovery.py +40 -0
  145. package/bundles/antigravity-plugin/skills/webapp-testing/examples/static_html_automation.py +33 -0
  146. package/bundles/antigravity-plugin/skills/webapp-testing/scripts/with_server.py +106 -0
  147. package/bundles/antigravity-plugin/skills/windows-deployment/SKILL.md +880 -0
  148. package/bundles/antigravity-plugin/skills/writing-plans/SKILL.md +384 -0
  149. package/bundles/antigravity-plugin/statusline-command-agy.sh +91 -0
  150. package/bundles/antigravity-plugin/warm_start_agy.py +149 -0
  151. package/bundles/manifest.json +6 -0
  152. package/dist/cli.js +5363 -0
  153. package/package.json +61 -4
@@ -0,0 +1,328 @@
1
+ #!/usr/bin/env python3
2
+ """Run the eval + improve loop until all pass or max iterations reached.
3
+
4
+ Combines run_eval.py and improve_description.py in a loop, tracking history
5
+ and returning the best description found. Supports train/test split to prevent
6
+ overfitting.
7
+ """
8
+
9
+ import argparse
10
+ import json
11
+ import random
12
+ import sys
13
+ import tempfile
14
+ import time
15
+ import webbrowser
16
+ from pathlib import Path
17
+
18
+ from scripts.generate_report import generate_html
19
+ from scripts.improve_description import improve_description
20
+ from scripts.run_eval import find_project_root, run_eval
21
+ from scripts.utils import parse_skill_md
22
+
23
+
24
+ def split_eval_set(eval_set: list[dict], holdout: float, seed: int = 42) -> tuple[list[dict], list[dict]]:
25
+ """Split eval set into train and test sets, stratified by should_trigger."""
26
+ random.seed(seed)
27
+
28
+ # Separate by should_trigger
29
+ trigger = [e for e in eval_set if e["should_trigger"]]
30
+ no_trigger = [e for e in eval_set if not e["should_trigger"]]
31
+
32
+ # Shuffle each group
33
+ random.shuffle(trigger)
34
+ random.shuffle(no_trigger)
35
+
36
+ # Calculate split points
37
+ n_trigger_test = max(1, int(len(trigger) * holdout))
38
+ n_no_trigger_test = max(1, int(len(no_trigger) * holdout))
39
+
40
+ # Split
41
+ test_set = trigger[:n_trigger_test] + no_trigger[:n_no_trigger_test]
42
+ train_set = trigger[n_trigger_test:] + no_trigger[n_no_trigger_test:]
43
+
44
+ return train_set, test_set
45
+
46
+
47
+ def run_loop(
48
+ eval_set: list[dict],
49
+ skill_path: Path,
50
+ description_override: str | None,
51
+ num_workers: int,
52
+ timeout: int,
53
+ max_iterations: int,
54
+ runs_per_query: int,
55
+ trigger_threshold: float,
56
+ holdout: float,
57
+ model: str,
58
+ verbose: bool,
59
+ live_report_path: Path | None = None,
60
+ log_dir: Path | None = None,
61
+ ) -> dict:
62
+ """Run the eval + improvement loop."""
63
+ project_root = find_project_root()
64
+ name, original_description, content = parse_skill_md(skill_path)
65
+ current_description = description_override or original_description
66
+
67
+ # Split into train/test if holdout > 0
68
+ if holdout > 0:
69
+ train_set, test_set = split_eval_set(eval_set, holdout)
70
+ if verbose:
71
+ print(f"Split: {len(train_set)} train, {len(test_set)} test (holdout={holdout})", file=sys.stderr)
72
+ else:
73
+ train_set = eval_set
74
+ test_set = []
75
+
76
+ history = []
77
+ exit_reason = "unknown"
78
+
79
+ for iteration in range(1, max_iterations + 1):
80
+ if verbose:
81
+ print(f"\n{'='*60}", file=sys.stderr)
82
+ print(f"Iteration {iteration}/{max_iterations}", file=sys.stderr)
83
+ print(f"Description: {current_description}", file=sys.stderr)
84
+ print(f"{'='*60}", file=sys.stderr)
85
+
86
+ # Evaluate train + test together in one batch for parallelism
87
+ all_queries = train_set + test_set
88
+ t0 = time.time()
89
+ all_results = run_eval(
90
+ eval_set=all_queries,
91
+ skill_name=name,
92
+ description=current_description,
93
+ num_workers=num_workers,
94
+ timeout=timeout,
95
+ project_root=project_root,
96
+ runs_per_query=runs_per_query,
97
+ trigger_threshold=trigger_threshold,
98
+ model=model,
99
+ )
100
+ eval_elapsed = time.time() - t0
101
+
102
+ # Split results back into train/test by matching queries
103
+ train_queries_set = {q["query"] for q in train_set}
104
+ train_result_list = [r for r in all_results["results"] if r["query"] in train_queries_set]
105
+ test_result_list = [r for r in all_results["results"] if r["query"] not in train_queries_set]
106
+
107
+ train_passed = sum(1 for r in train_result_list if r["pass"])
108
+ train_total = len(train_result_list)
109
+ train_summary = {"passed": train_passed, "failed": train_total - train_passed, "total": train_total}
110
+ train_results = {"results": train_result_list, "summary": train_summary}
111
+
112
+ if test_set:
113
+ test_passed = sum(1 for r in test_result_list if r["pass"])
114
+ test_total = len(test_result_list)
115
+ test_summary = {"passed": test_passed, "failed": test_total - test_passed, "total": test_total}
116
+ test_results = {"results": test_result_list, "summary": test_summary}
117
+ else:
118
+ test_results = None
119
+ test_summary = None
120
+
121
+ history.append({
122
+ "iteration": iteration,
123
+ "description": current_description,
124
+ "train_passed": train_summary["passed"],
125
+ "train_failed": train_summary["failed"],
126
+ "train_total": train_summary["total"],
127
+ "train_results": train_results["results"],
128
+ "test_passed": test_summary["passed"] if test_summary else None,
129
+ "test_failed": test_summary["failed"] if test_summary else None,
130
+ "test_total": test_summary["total"] if test_summary else None,
131
+ "test_results": test_results["results"] if test_results else None,
132
+ # For backward compat with report generator
133
+ "passed": train_summary["passed"],
134
+ "failed": train_summary["failed"],
135
+ "total": train_summary["total"],
136
+ "results": train_results["results"],
137
+ })
138
+
139
+ # Write live report if path provided
140
+ if live_report_path:
141
+ partial_output = {
142
+ "original_description": original_description,
143
+ "best_description": current_description,
144
+ "best_score": "in progress",
145
+ "iterations_run": len(history),
146
+ "holdout": holdout,
147
+ "train_size": len(train_set),
148
+ "test_size": len(test_set),
149
+ "history": history,
150
+ }
151
+ live_report_path.write_text(generate_html(partial_output, auto_refresh=True, skill_name=name))
152
+
153
+ if verbose:
154
+ def print_eval_stats(label, results, elapsed):
155
+ pos = [r for r in results if r["should_trigger"]]
156
+ neg = [r for r in results if not r["should_trigger"]]
157
+ tp = sum(r["triggers"] for r in pos)
158
+ pos_runs = sum(r["runs"] for r in pos)
159
+ fn = pos_runs - tp
160
+ fp = sum(r["triggers"] for r in neg)
161
+ neg_runs = sum(r["runs"] for r in neg)
162
+ tn = neg_runs - fp
163
+ total = tp + tn + fp + fn
164
+ precision = tp / (tp + fp) if (tp + fp) > 0 else 1.0
165
+ recall = tp / (tp + fn) if (tp + fn) > 0 else 1.0
166
+ accuracy = (tp + tn) / total if total > 0 else 0.0
167
+ print(f"{label}: {tp+tn}/{total} correct, precision={precision:.0%} recall={recall:.0%} accuracy={accuracy:.0%} ({elapsed:.1f}s)", file=sys.stderr)
168
+ for r in results:
169
+ status = "PASS" if r["pass"] else "FAIL"
170
+ rate_str = f"{r['triggers']}/{r['runs']}"
171
+ print(f" [{status}] rate={rate_str} expected={r['should_trigger']}: {r['query'][:60]}", file=sys.stderr)
172
+
173
+ print_eval_stats("Train", train_results["results"], eval_elapsed)
174
+ if test_summary:
175
+ print_eval_stats("Test ", test_results["results"], 0)
176
+
177
+ if train_summary["failed"] == 0:
178
+ exit_reason = f"all_passed (iteration {iteration})"
179
+ if verbose:
180
+ print(f"\nAll train queries passed on iteration {iteration}!", file=sys.stderr)
181
+ break
182
+
183
+ if iteration == max_iterations:
184
+ exit_reason = f"max_iterations ({max_iterations})"
185
+ if verbose:
186
+ print(f"\nMax iterations reached ({max_iterations}).", file=sys.stderr)
187
+ break
188
+
189
+ # Improve the description based on train results
190
+ if verbose:
191
+ print(f"\nImproving description...", file=sys.stderr)
192
+
193
+ t0 = time.time()
194
+ # Strip test scores from history so improvement model can't see them
195
+ blinded_history = [
196
+ {k: v for k, v in h.items() if not k.startswith("test_")}
197
+ for h in history
198
+ ]
199
+ new_description = improve_description(
200
+ skill_name=name,
201
+ skill_content=content,
202
+ current_description=current_description,
203
+ eval_results=train_results,
204
+ history=blinded_history,
205
+ model=model,
206
+ log_dir=log_dir,
207
+ iteration=iteration,
208
+ )
209
+ improve_elapsed = time.time() - t0
210
+
211
+ if verbose:
212
+ print(f"Proposed ({improve_elapsed:.1f}s): {new_description}", file=sys.stderr)
213
+
214
+ current_description = new_description
215
+
216
+ # Find the best iteration by TEST score (or train if no test set)
217
+ if test_set:
218
+ best = max(history, key=lambda h: h["test_passed"] or 0)
219
+ best_score = f"{best['test_passed']}/{best['test_total']}"
220
+ else:
221
+ best = max(history, key=lambda h: h["train_passed"])
222
+ best_score = f"{best['train_passed']}/{best['train_total']}"
223
+
224
+ if verbose:
225
+ print(f"\nExit reason: {exit_reason}", file=sys.stderr)
226
+ print(f"Best score: {best_score} (iteration {best['iteration']})", file=sys.stderr)
227
+
228
+ return {
229
+ "exit_reason": exit_reason,
230
+ "original_description": original_description,
231
+ "best_description": best["description"],
232
+ "best_score": best_score,
233
+ "best_train_score": f"{best['train_passed']}/{best['train_total']}",
234
+ "best_test_score": f"{best['test_passed']}/{best['test_total']}" if test_set else None,
235
+ "final_description": current_description,
236
+ "iterations_run": len(history),
237
+ "holdout": holdout,
238
+ "train_size": len(train_set),
239
+ "test_size": len(test_set),
240
+ "history": history,
241
+ }
242
+
243
+
244
+ def main():
245
+ parser = argparse.ArgumentParser(description="Run eval + improve loop")
246
+ parser.add_argument("--eval-set", required=True, help="Path to eval set JSON file")
247
+ parser.add_argument("--skill-path", required=True, help="Path to skill directory")
248
+ parser.add_argument("--description", default=None, help="Override starting description")
249
+ parser.add_argument("--num-workers", type=int, default=10, help="Number of parallel workers")
250
+ parser.add_argument("--timeout", type=int, default=30, help="Timeout per query in seconds")
251
+ parser.add_argument("--max-iterations", type=int, default=5, help="Max improvement iterations")
252
+ parser.add_argument("--runs-per-query", type=int, default=3, help="Number of runs per query")
253
+ parser.add_argument("--trigger-threshold", type=float, default=0.5, help="Trigger rate threshold")
254
+ parser.add_argument("--holdout", type=float, default=0.4, help="Fraction of eval set to hold out for testing (0 to disable)")
255
+ parser.add_argument("--model", required=True, help="Model for improvement")
256
+ parser.add_argument("--verbose", action="store_true", help="Print progress to stderr")
257
+ parser.add_argument("--report", default="auto", help="Generate HTML report at this path (default: 'auto' for temp file, 'none' to disable)")
258
+ parser.add_argument("--results-dir", default=None, help="Save all outputs (results.json, report.html, log.txt) to a timestamped subdirectory here")
259
+ args = parser.parse_args()
260
+
261
+ eval_set = json.loads(Path(args.eval_set).read_text())
262
+ skill_path = Path(args.skill_path)
263
+
264
+ if not (skill_path / "SKILL.md").exists():
265
+ print(f"Error: No SKILL.md found at {skill_path}", file=sys.stderr)
266
+ sys.exit(1)
267
+
268
+ name, _, _ = parse_skill_md(skill_path)
269
+
270
+ # Set up live report path
271
+ if args.report != "none":
272
+ if args.report == "auto":
273
+ timestamp = time.strftime("%Y%m%d_%H%M%S")
274
+ live_report_path = Path(tempfile.gettempdir()) / f"skill_description_report_{skill_path.name}_{timestamp}.html"
275
+ else:
276
+ live_report_path = Path(args.report)
277
+ # Open the report immediately so the user can watch
278
+ live_report_path.write_text("<html><body><h1>Starting optimization loop...</h1><meta http-equiv='refresh' content='5'></body></html>")
279
+ webbrowser.open(str(live_report_path))
280
+ else:
281
+ live_report_path = None
282
+
283
+ # Determine output directory (create before run_loop so logs can be written)
284
+ if args.results_dir:
285
+ timestamp = time.strftime("%Y-%m-%d_%H%M%S")
286
+ results_dir = Path(args.results_dir) / timestamp
287
+ results_dir.mkdir(parents=True, exist_ok=True)
288
+ else:
289
+ results_dir = None
290
+
291
+ log_dir = results_dir / "logs" if results_dir else None
292
+
293
+ output = run_loop(
294
+ eval_set=eval_set,
295
+ skill_path=skill_path,
296
+ description_override=args.description,
297
+ num_workers=args.num_workers,
298
+ timeout=args.timeout,
299
+ max_iterations=args.max_iterations,
300
+ runs_per_query=args.runs_per_query,
301
+ trigger_threshold=args.trigger_threshold,
302
+ holdout=args.holdout,
303
+ model=args.model,
304
+ verbose=args.verbose,
305
+ live_report_path=live_report_path,
306
+ log_dir=log_dir,
307
+ )
308
+
309
+ # Save JSON output
310
+ json_output = json.dumps(output, indent=2)
311
+ print(json_output)
312
+ if results_dir:
313
+ (results_dir / "results.json").write_text(json_output)
314
+
315
+ # Write final HTML report (without auto-refresh)
316
+ if live_report_path:
317
+ live_report_path.write_text(generate_html(output, auto_refresh=False, skill_name=name))
318
+ print(f"\nReport: {live_report_path}", file=sys.stderr)
319
+
320
+ if results_dir and live_report_path:
321
+ (results_dir / "report.html").write_text(generate_html(output, auto_refresh=False, skill_name=name))
322
+
323
+ if results_dir:
324
+ print(f"Results saved to: {results_dir}", file=sys.stderr)
325
+
326
+
327
+ if __name__ == "__main__":
328
+ main()
@@ -0,0 +1,47 @@
1
+ """Shared utilities for skill-creator scripts."""
2
+
3
+ from pathlib import Path
4
+
5
+
6
+
7
+ def parse_skill_md(skill_path: Path) -> tuple[str, str, str]:
8
+ """Parse a SKILL.md file, returning (name, description, full_content)."""
9
+ content = (skill_path / "SKILL.md").read_text()
10
+ lines = content.split("\n")
11
+
12
+ if lines[0].strip() != "---":
13
+ raise ValueError("SKILL.md missing frontmatter (no opening ---)")
14
+
15
+ end_idx = None
16
+ for i, line in enumerate(lines[1:], start=1):
17
+ if line.strip() == "---":
18
+ end_idx = i
19
+ break
20
+
21
+ if end_idx is None:
22
+ raise ValueError("SKILL.md missing frontmatter (no closing ---)")
23
+
24
+ name = ""
25
+ description = ""
26
+ frontmatter_lines = lines[1:end_idx]
27
+ i = 0
28
+ while i < len(frontmatter_lines):
29
+ line = frontmatter_lines[i]
30
+ if line.startswith("name:"):
31
+ name = line[len("name:"):].strip().strip('"').strip("'")
32
+ elif line.startswith("description:"):
33
+ value = line[len("description:"):].strip()
34
+ # Handle YAML multiline indicators (>, |, >-, |-)
35
+ if value in (">", "|", ">-", "|-"):
36
+ continuation_lines: list[str] = []
37
+ i += 1
38
+ while i < len(frontmatter_lines) and (frontmatter_lines[i].startswith(" ") or frontmatter_lines[i].startswith("\t")):
39
+ continuation_lines.append(frontmatter_lines[i].strip())
40
+ i += 1
41
+ description = " ".join(continuation_lines)
42
+ continue
43
+ else:
44
+ description = value.strip('"').strip("'")
45
+ i += 1
46
+
47
+ return name, description, content
@@ -0,0 +1,321 @@
1
+ ---
2
+ name: sql
3
+ description: Write high-quality, optimized SQL with best practices for performance, NULL handling, security, and readability. Database-agnostic patterns with dialect-specific notes.
4
+ ---
5
+
6
+ # SQL Expert Skill
7
+
8
+ Write high-quality, optimized SQL for enterprise environments. Covers performance, edge cases, and real-world gotchas.
9
+
10
+ ## When to Use
11
+
12
+ - Writing new SQL queries
13
+ - Optimizing existing queries
14
+ - Handling large datasets (100K+ rows)
15
+ - Creating stored procedures or views
16
+ - Debugging slow queries
17
+
18
+ ---
19
+
20
+ ## SQL Quality Standards
21
+
22
+ ### 1. Always Include Comments
23
+
24
+ ```sql
25
+ -- Get open orders with SLA breach risk
26
+ -- Filters: Last 7 days, excludes test data
27
+ -- Expected rows: ~500-2000
28
+ SELECT
29
+ OrderID,
30
+ CustomerName,
31
+ DATEDIFF(HOUR, CreatedAt, GETDATE()) AS hours_open
32
+ FROM Orders
33
+ WHERE CreatedAt >= DATEADD(DAY, -7, GETDATE())
34
+ AND Status = 'Open'
35
+ ORDER BY hours_open DESC;
36
+ ```
37
+
38
+ ### 2. Readable Formatting
39
+
40
+ ```sql
41
+ SELECT
42
+ o.OrderID,
43
+ o.CustomerName,
44
+ o.Status,
45
+ COUNT(i.ItemID) AS item_count
46
+ FROM Orders o
47
+ LEFT JOIN OrderItems i ON o.OrderID = i.OrderID
48
+ WHERE o.Status = 'Open'
49
+ GROUP BY o.OrderID, o.CustomerName, o.Status
50
+ HAVING COUNT(i.ItemID) > 0
51
+ ORDER BY item_count DESC;
52
+ ```
53
+
54
+ ---
55
+
56
+ ## Large Dataset Patterns
57
+
58
+ ### Pagination
59
+
60
+ ```sql
61
+ -- OFFSET-FETCH (SQL Server 2012+ / PostgreSQL)
62
+ SELECT OrderID, CustomerName, Status
63
+ FROM Orders
64
+ WHERE Status = 'Open'
65
+ ORDER BY CreatedAt DESC
66
+ OFFSET 50 ROWS FETCH NEXT 25 ROWS ONLY;
67
+
68
+ -- Keyset pagination (faster for deep pages)
69
+ SELECT TOP 25 OrderID, CustomerName, CreatedAt
70
+ FROM Orders
71
+ WHERE CreatedAt < @lastSeenDate AND Status = 'Open'
72
+ ORDER BY CreatedAt DESC;
73
+ ```
74
+
75
+ ### Batching Large Operations
76
+
77
+ ```sql
78
+ DECLARE @BatchSize INT = 5000;
79
+ DECLARE @RowsAffected INT = 1;
80
+
81
+ WHILE @RowsAffected > 0
82
+ BEGIN
83
+ UPDATE TOP (@BatchSize) Orders
84
+ SET ProcessedFlag = 1
85
+ WHERE Status = 'Closed' AND ProcessedFlag = 0;
86
+ SET @RowsAffected = @@ROWCOUNT;
87
+ WAITFOR DELAY '00:00:00.100';
88
+ END
89
+ ```
90
+
91
+ ---
92
+
93
+ ## NULL Handling
94
+
95
+ ```sql
96
+ -- Use IS NULL (not = NULL)
97
+ SELECT * FROM Orders WHERE Category IS NULL;
98
+
99
+ -- ISNULL (SQL Server) vs COALESCE (ANSI standard)
100
+ SELECT ISNULL(Category, 'Unknown') FROM Orders;
101
+ SELECT COALESCE(Category, SubCategory, 'Unknown') FROM Orders;
102
+
103
+ -- NULL in aggregations: COUNT(*) counts all, COUNT(col) excludes NULLs
104
+ SELECT COUNT(*) AS total, COUNT(Category) AS non_null FROM Orders;
105
+ SELECT ISNULL(SUM(Amount), 0) AS total FROM Orders;
106
+ ```
107
+
108
+ ---
109
+
110
+ ## Performance Killers to Avoid
111
+
112
+ 1. **Implicit type conversions** -- match parameter types to column types
113
+ 2. **Functions on indexed columns** -- rewrite `WHERE YEAR(date) = 2026` as range
114
+ 3. **OR conditions** -- consider UNION ALL for index use
115
+ 4. **SELECT DISTINCT on large sets** -- use GROUP BY or EXISTS
116
+ 5. **UNION vs UNION ALL** -- use UNION ALL when duplicates are OK
117
+
118
+ ---
119
+
120
+ ## Window Functions
121
+
122
+ ```sql
123
+ -- Running totals
124
+ SELECT OrderID, CreatedAt,
125
+ COUNT(*) OVER (ORDER BY CreatedAt) AS running_count,
126
+ SUM(Amount) OVER (ORDER BY CreatedAt) AS running_total
127
+ FROM Orders;
128
+
129
+ -- Latest record per group
130
+ WITH ranked AS (
131
+ SELECT *, ROW_NUMBER() OVER (PARTITION BY CustomerID ORDER BY CreatedAt DESC) AS rn
132
+ FROM Orders
133
+ )
134
+ SELECT * FROM ranked WHERE rn = 1;
135
+ ```
136
+
137
+ ---
138
+
139
+ ## Error Handling
140
+
141
+ ```sql
142
+ BEGIN TRY
143
+ BEGIN TRANSACTION;
144
+ UPDATE Orders SET Status = 'Processed' WHERE OrderID = @OrderID;
145
+ IF @@ROWCOUNT = 0 THROW 50001, 'Order not found', 1;
146
+ COMMIT;
147
+ END TRY
148
+ BEGIN CATCH
149
+ IF @@TRANCOUNT > 0 ROLLBACK;
150
+ DECLARE @ErrorMsg NVARCHAR(4000) = ERROR_MESSAGE();
151
+ RAISERROR(@ErrorMsg, ERROR_SEVERITY(), ERROR_STATE());
152
+ END CATCH;
153
+ ```
154
+
155
+ ---
156
+
157
+ ## Security (CRITICAL)
158
+
159
+ ```sql
160
+ -- NEVER concatenate user input
161
+ -- Always use parameterized queries
162
+ EXEC sp_executesql N'SELECT * FROM Orders WHERE OrderID = @id',
163
+ N'@id UNIQUEIDENTIFIER', @id = @orderId;
164
+ ```
165
+
166
+ ```python
167
+ # Python: always parameterized
168
+ query = "SELECT * FROM Orders WHERE Status = ?"
169
+ result = adapter.fetch_dataframe(query, (user_input,))
170
+ ```
171
+
172
+ ---
173
+
174
+ ## EXISTS vs IN vs JOIN
175
+
176
+ ```sql
177
+ -- EXISTS: Best for "does any match exist?" - stops at first match
178
+ SELECT * FROM Customers c
179
+ WHERE EXISTS (SELECT 1 FROM Orders WHERE CustomerID = c.CustomerID);
180
+
181
+ -- IN: Good for small lists, can be slow with NULLs in subquery
182
+ SELECT * FROM Orders WHERE CustomerID IN ('CUST001', 'CUST002', 'CUST003');
183
+
184
+ -- ⚠️ IN with subquery containing NULLs can return unexpected results
185
+ SELECT * FROM Orders
186
+ WHERE CustomerID IN (SELECT CustomerID FROM FlaggedCustomers); -- If NULL in subquery, watch out!
187
+
188
+ -- JOIN: Best when you need columns from both tables
189
+ SELECT c.*, o.OrderID
190
+ FROM Customers c
191
+ INNER JOIN Orders o ON c.CustomerID = o.CustomerID;
192
+ ```
193
+
194
+ ---
195
+
196
+ ## Concurrency & Locking
197
+
198
+ ### Read Patterns
199
+
200
+ ```sql
201
+ -- For reports/dashboards (OK with slightly stale data)
202
+ SELECT * FROM Orders WITH (NOLOCK) -- Dirty reads, but no blocking
203
+ WHERE Status = 'Open';
204
+
205
+ -- ⚠️ NOLOCK risks:
206
+ -- - Can read uncommitted data (that might roll back)
207
+ -- - Can read rows twice or skip rows during scans
208
+ -- Only use for approximate counts, dashboards, reports
209
+
210
+ -- For accurate reads that shouldn't block
211
+ SET TRANSACTION ISOLATION LEVEL READ COMMITTED SNAPSHOT;
212
+ -- Requires database setting: ALTER DATABASE [DB] SET READ_COMMITTED_SNAPSHOT ON;
213
+ ```
214
+
215
+ ### Deadlock Prevention
216
+
217
+ ```sql
218
+ -- ✅ Always access tables in same order across all processes
219
+ -- ✅ Keep transactions short
220
+
221
+ BEGIN TRANSACTION;
222
+ UPDATE Orders SET Status = 'Processed' WHERE OrderID = @id;
223
+ COMMIT;
224
+ ```
225
+
226
+ ---
227
+
228
+ ## Temp Tables vs Table Variables
229
+
230
+ ```sql
231
+ -- TEMP TABLES: Better for large datasets (has statistics)
232
+ CREATE TABLE #TempOrders (
233
+ OrderID VARCHAR(50),
234
+ CustomerID VARCHAR(50),
235
+ INDEX IX_Temp_CustomerID (CustomerID)
236
+ );
237
+ INSERT INTO #TempOrders SELECT OrderID, CustomerID FROM Orders WHERE Status = 'Open';
238
+ SELECT * FROM #TempOrders t JOIN Customers c ON t.CustomerID = c.CustomerID;
239
+ DROP TABLE #TempOrders;
240
+
241
+ -- TABLE VARIABLES: Better for small datasets (<1000 rows)
242
+ DECLARE @Orders TABLE (OrderID VARCHAR(50), CustomerID VARCHAR(50));
243
+ INSERT INTO @Orders SELECT OrderID, CustomerID FROM Orders WHERE Status = 'Closed';
244
+ -- ⚠️ Table variables assume 1 row for optimization - bad for large sets
245
+ ```
246
+
247
+ ---
248
+
249
+ ## Date/Time Gotchas
250
+
251
+ ```sql
252
+ -- ❌ BAD: String comparison for dates (locale issues)
253
+ SELECT * FROM Orders WHERE CreatedAt > '01/02/2026'; -- Jan 2 or Feb 1?
254
+
255
+ -- ✅ GOOD: Use unambiguous ISO format
256
+ SELECT * FROM Orders WHERE CreatedAt > '2026-01-02';
257
+
258
+ -- Date range queries (include full day)
259
+ -- ❌ WRONG: BETWEEN misses end-of-day records
260
+ SELECT * FROM Orders WHERE CreatedAt BETWEEN '2026-01-01' AND '2026-01-31';
261
+
262
+ -- ✅ CORRECT: Use < next day
263
+ SELECT * FROM Orders
264
+ WHERE CreatedAt >= '2026-01-01' AND CreatedAt < '2026-02-01';
265
+
266
+ -- Get date part only (for grouping)
267
+ SELECT CAST(CreatedAt AS DATE) AS created_date, COUNT(*) AS daily_count
268
+ FROM Orders GROUP BY CAST(CreatedAt AS DATE);
269
+ ```
270
+
271
+ ---
272
+
273
+ ## Query Complexity Guidelines
274
+
275
+ | Guideline | Threshold | If Exceeded |
276
+ |-----------|-----------|-------------|
277
+ | Max JOINs | 3-4 | Create VIEW or break into CTEs |
278
+ | Max subqueries | 2 levels | Use CTEs instead |
279
+ | Max CASE statements | 3-4 | Move logic to application layer |
280
+ | Max columns | 15-20 | Consider if all are needed |
281
+
282
+ ### CTEs for Complex Logic
283
+
284
+ ```sql
285
+ WITH open_orders AS (
286
+ SELECT OrderID, CustomerID, CreatedAt FROM Orders WHERE Status = 'Open'
287
+ ),
288
+ order_metrics AS (
289
+ SELECT o.OrderID, COUNT(i.ItemID) AS item_count, MAX(i.UpdatedAt) AS last_update
290
+ FROM open_orders o
291
+ LEFT JOIN OrderItems i ON o.OrderID = i.OrderID
292
+ GROUP BY o.OrderID
293
+ )
294
+ SELECT * FROM order_metrics WHERE item_count > 5;
295
+ ```
296
+
297
+ ---
298
+
299
+ ## Execution Plan Red Flags
300
+
301
+ | Warning | Meaning | Fix |
302
+ |---------|---------|-----|
303
+ | Table Scan | No index used | Add index or fix WHERE clause |
304
+ | Key Lookup | Extra I/O for non-covered columns | Add columns to index or use covering index |
305
+ | CONVERT_IMPLICIT | Type mismatch causing conversion | Match parameter types to column types |
306
+ | Sort (high cost) | Sorting large dataset | Add index that matches ORDER BY |
307
+ | Hash Match | Large join, possibly missing index | Add indexes on join columns |
308
+ | Parallelism | Query is complex/large | May be OK, but check if needed |
309
+
310
+ ---
311
+
312
+ ## Performance Checklist
313
+
314
+ - [ ] Only selecting needed columns (no SELECT *)
315
+ - [ ] Filtering in WHERE, not application code
316
+ - [ ] JOINs on indexed columns
317
+ - [ ] No functions on indexed columns in WHERE
318
+ - [ ] Types match (no implicit conversions)
319
+ - [ ] Parameterized (no string concatenation)
320
+ - [ ] Pagination for large result sets
321
+ - [ ] Comments explain purpose and expected volume