oh-my-opencode 5.0.0-beta.26 → 5.0.0-beta.28
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/cli/config-manager/parse-opencode-config-file.d.ts +2 -1
- package/dist/cli/doctor/checks/system-plugin.d.ts +4 -3
- package/dist/cli/index.js +59 -45
- package/dist/cli-node/index.js +82 -45
- package/dist/hooks/todo-continuation-enforcer/types.d.ts +1 -0
- package/dist/index.js +53 -28
- package/dist/shared/index.d.ts +1 -0
- package/dist/shared/legacy-plugin-warning.d.ts +2 -1
- package/dist/shared/plugin-entry-migrator.d.ts +4 -3
- package/dist/shared/plugin-entry-shape.d.ts +6 -0
- package/dist/skills/data-scientist/SKILL.md +99 -239
- package/dist/skills/data-scientist/references/execution-surfaces.md +91 -0
- package/dist/skills/data-scientist/references/placement.md +74 -0
- package/dist/skills/data-scientist/references/polars-lane.md +95 -0
- package/dist/skills/data-scientist/references/uv-setup.md +1 -1
- package/dist/skills/data-scientist/references/visualization.md +64 -0
- package/dist/skills/data-scientist/scripts/ensure-js-deps.sh +28 -0
- package/dist/skills/data-scientist/scripts/ensure-py-deps.sh +37 -0
- package/dist/skills/ulw-research/SKILL.md +3 -1
- package/dist/tui.d.ts +12 -0
- package/dist/tui.js +3 -3
- package/package.json +15 -15
- package/packages/lsp-core/src/request-context.test.ts +29 -0
- package/packages/lsp-core/src/request-context.ts +1 -1
- package/packages/lsp-daemon/dist/cli.js +1 -1
- package/packages/lsp-daemon/dist/client.js +1 -1
- package/packages/lsp-daemon/dist/index.js +1 -1
- package/packages/lsp-tools-mcp/dist/cli.js +1 -1
- package/packages/lsp-tools-mcp/dist/mcp.js +1 -1
- package/packages/lsp-tools-mcp/dist/request-context.js +1 -1
- package/packages/lsp-tools-mcp/dist/tools.js +1 -1
- package/packages/omo-codex/plugin/.codex-plugin/plugin.json +1 -1
- package/packages/omo-codex/plugin/components/bootstrap/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/bootstrap/package.json +1 -1
- package/packages/omo-codex/plugin/components/codegraph/package.json +1 -1
- package/packages/omo-codex/plugin/components/comment-checker/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/comment-checker/package.json +1 -1
- package/packages/omo-codex/plugin/components/git-bash/hooks/hooks.json +2 -2
- package/packages/omo-codex/plugin/components/git-bash/package.json +1 -1
- package/packages/omo-codex/plugin/components/lazycodex-executor-verify/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/lazycodex-executor-verify/package.json +1 -1
- package/packages/omo-codex/plugin/components/lsp/dist/.omo-runtime-manifest.json +2 -2
- package/packages/omo-codex/plugin/components/lsp/hooks/hooks.json +2 -2
- package/packages/omo-codex/plugin/components/lsp/package.json +1 -1
- package/packages/omo-codex/plugin/components/rules/hooks/hooks.json +4 -4
- package/packages/omo-codex/plugin/components/rules/package.json +1 -1
- package/packages/omo-codex/plugin/components/teammode/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/teammode/package.json +1 -1
- package/packages/omo-codex/plugin/components/telemetry/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/telemetry/package.json +1 -1
- package/packages/omo-codex/plugin/components/ultrawork/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/ultrawork/package.json +1 -1
- package/packages/omo-codex/plugin/components/ulw-execute-continuation/hooks/hooks.json +2 -2
- package/packages/omo-codex/plugin/components/ulw-execute-continuation/package.json +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/hooks/hooks.json +4 -4
- package/packages/omo-codex/plugin/components/ulw-loop/package.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-compact-resetting-git-bash-mcp-reminder.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-compact-resetting-lsp-diagnostics-cache.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-compact-resetting-project-rule-cache.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-checking-codegraph-init-guidance.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-checking-comments.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-checking-lsp-diagnostics.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-checking-thread-title-hygiene.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-matching-project-rules.json +1 -1
- package/packages/omo-codex/plugin/hooks/pre-tool-use-enforcing-unlimited-goal-budget.json +1 -1
- package/packages/omo-codex/plugin/hooks/pre-tool-use-guarding-ulw-loop-spawns.json +1 -1
- package/packages/omo-codex/plugin/hooks/pre-tool-use-recommending-git-bash-mcp.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-checking-auto-update.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-checking-bootstrap-provisioning.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-checking-codegraph-bootstrap.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-loading-project-rules.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-recording-session-telemetry.json +1 -1
- package/packages/omo-codex/plugin/hooks/stop-checking-ulw-execute-continuation.json +1 -1
- package/packages/omo-codex/plugin/hooks/stop-checking-ulw-loop-resume.json +1 -1
- package/packages/omo-codex/plugin/hooks/subagent-stop-checking-ulw-execute-continuation.json +1 -1
- package/packages/omo-codex/plugin/hooks/subagent-stop-verifying-lazycodex-executor-evidence.json +1 -1
- package/packages/omo-codex/plugin/hooks/user-prompt-submit-checking-ultrawork-trigger.json +1 -1
- package/packages/omo-codex/plugin/hooks/user-prompt-submit-checking-ulw-loop-steering.json +1 -1
- package/packages/omo-codex/plugin/hooks/user-prompt-submit-loading-project-rules.json +1 -1
- package/packages/omo-codex/plugin/package-lock.json +13 -13
- package/packages/omo-codex/plugin/package.json +1 -1
- package/packages/omo-codex/plugin/skills/data-scientist/SKILL.md +99 -239
- package/packages/omo-codex/plugin/skills/data-scientist/references/execution-surfaces.md +91 -0
- package/packages/omo-codex/plugin/skills/data-scientist/references/placement.md +74 -0
- package/packages/omo-codex/plugin/skills/data-scientist/references/polars-lane.md +95 -0
- package/packages/omo-codex/plugin/skills/data-scientist/references/uv-setup.md +1 -1
- package/packages/omo-codex/plugin/skills/data-scientist/references/visualization.md +64 -0
- package/packages/omo-codex/plugin/skills/data-scientist/scripts/ensure-js-deps.sh +28 -0
- package/packages/omo-codex/plugin/skills/data-scientist/scripts/ensure-py-deps.sh +37 -0
- package/packages/omo-codex/plugin/skills/ulw-research/SKILL.md +3 -1
- package/packages/omo-codex/scripts/install-dist/install-local.mjs +2 -2
- package/packages/shared-skills/skills/data-scientist/SKILL.md +99 -239
- package/packages/shared-skills/skills/data-scientist/references/execution-surfaces.md +91 -0
- package/packages/shared-skills/skills/data-scientist/references/placement.md +74 -0
- package/packages/shared-skills/skills/data-scientist/references/polars-lane.md +95 -0
- package/packages/shared-skills/skills/data-scientist/references/uv-setup.md +1 -1
- package/packages/shared-skills/skills/data-scientist/references/visualization.md +64 -0
- package/packages/shared-skills/skills/data-scientist/scripts/ensure-js-deps.sh +28 -0
- package/packages/shared-skills/skills/data-scientist/scripts/ensure-py-deps.sh +37 -0
- package/packages/shared-skills/skills/ulw-research/SKILL.md +3 -1
- package/dist/skills/data-scientist/references/common-scenarios.md +0 -176
- package/dist/skills/data-scientist/references/execution-templates.md +0 -197
- package/dist/skills/data-scientist/references/integration-patterns.md +0 -153
- package/dist/skills/data-scientist/references/performance-benchmarks.md +0 -37
- package/packages/omo-codex/plugin/skills/data-scientist/references/common-scenarios.md +0 -176
- package/packages/omo-codex/plugin/skills/data-scientist/references/execution-templates.md +0 -197
- package/packages/omo-codex/plugin/skills/data-scientist/references/integration-patterns.md +0 -153
- package/packages/omo-codex/plugin/skills/data-scientist/references/performance-benchmarks.md +0 -37
- package/packages/shared-skills/skills/data-scientist/references/common-scenarios.md +0 -176
- package/packages/shared-skills/skills/data-scientist/references/execution-templates.md +0 -197
- package/packages/shared-skills/skills/data-scientist/references/integration-patterns.md +0 -153
- package/packages/shared-skills/skills/data-scientist/references/performance-benchmarks.md +0 -37
|
@@ -0,0 +1,95 @@
|
|
|
1
|
+
# Polars lane (resident Python kernel)
|
|
2
|
+
|
|
3
|
+
When the work is DataFrame-shaped, Polars is the right engine — and it runs in the resident
|
|
4
|
+
Python kernel by default. Kernels rarely ship polars/pyarrow preinstalled, so inject them
|
|
5
|
+
once per session; the install lands in a user cache keyed to the kernel's interpreter, and
|
|
6
|
+
the interpreter itself is never mutated (run with the skill directory as cwd, or spell out
|
|
7
|
+
the script's absolute path):
|
|
8
|
+
|
|
9
|
+
```python
|
|
10
|
+
import subprocess, sys
|
|
11
|
+
site = subprocess.run(["bash", "scripts/ensure-py-deps.sh", sys.executable],
|
|
12
|
+
capture_output=True, text=True, check=True).stdout.strip()
|
|
13
|
+
sys.path.insert(0, site)
|
|
14
|
+
import polars as pl
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
After this, Polars lives across cells like every other resident engine: lazy frames,
|
|
18
|
+
intermediate results, and the DuckDB handoff all persist with no per-call process cost.
|
|
19
|
+
|
|
20
|
+
## When Polars wins over DuckDB SQL
|
|
21
|
+
|
|
22
|
+
- Expression-chain transforms: many derived columns, per-column conditional logic, string
|
|
23
|
+
pipelines — `with_columns` chains read and optimize better than nested SQL SELECTs.
|
|
24
|
+
- Reshapes: `unpivot`/`pivot` beat SQL gymnastics.
|
|
25
|
+
- Larger-than-RAM pipelines: the streaming engine executes lazy plans in chunks.
|
|
26
|
+
- Window-heavy feature engineering with `over()`.
|
|
27
|
+
|
|
28
|
+
SQL-shaped work (joins, aggregation, ad-hoc questions) stays in DuckDB; mixed pipelines hand
|
|
29
|
+
off zero-copy (below) instead of forcing one engine to do everything.
|
|
30
|
+
|
|
31
|
+
## Current API (1.x) — older spellings fail or warn
|
|
32
|
+
|
|
33
|
+
Training data is full of the pre-1.0 API. Current names:
|
|
34
|
+
|
|
35
|
+
| Use | Not |
|
|
36
|
+
| --- | --- |
|
|
37
|
+
| `pl.scan_csv` / `pl.scan_parquet` + `.collect()` | eager `read_*` on big files |
|
|
38
|
+
| `.group_by(...)` | `.groupby(...)` |
|
|
39
|
+
| `pl.len()` | `pl.count()` |
|
|
40
|
+
| `.collect(engine="streaming")` | `.collect(streaming=True)` |
|
|
41
|
+
| `.unpivot(...)` | `.melt(...)` |
|
|
42
|
+
|
|
43
|
+
Lazy first: `scan_*` builds a plan, pushes filters and projections down to the file read, and
|
|
44
|
+
executes once at `.collect()`. Eager `read_*` is for small files mutated interactively.
|
|
45
|
+
|
|
46
|
+
```python
|
|
47
|
+
out = (pl.scan_csv("data.csv")
|
|
48
|
+
.filter(pl.col("value") > 100)
|
|
49
|
+
.group_by("category")
|
|
50
|
+
.agg(pl.col("value").sum().alias("total"), pl.len().alias("n"))
|
|
51
|
+
.sort("total", descending=True)
|
|
52
|
+
.collect())
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
## Zero-copy handoff with DuckDB
|
|
56
|
+
|
|
57
|
+
Both engines speak Arrow, so mixed pipelines pay no serialization cost — all in-kernel:
|
|
58
|
+
|
|
59
|
+
```python
|
|
60
|
+
import duckdb
|
|
61
|
+
df = duckdb.sql("SELECT * FROM 'orders.csv' o JOIN 'items.csv' i USING (id)").pl()
|
|
62
|
+
shaped = df.with_columns((pl.col("qty") * pl.col("price")).alias("rev"))
|
|
63
|
+
duckdb.register("shaped", shaped)
|
|
64
|
+
out = duckdb.sql("SELECT category, SUM(rev) AS total FROM shaped GROUP BY 1").pl()
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
- `.pl()` requires pyarrow — the injection above provides it; without it, it raises
|
|
68
|
+
`ModuleNotFoundError`.
|
|
69
|
+
- Never `.df()`: it requires pandas (raising without it), and pandas is banned and absent.
|
|
70
|
+
|
|
71
|
+
## Streaming past RAM
|
|
72
|
+
|
|
73
|
+
```python
|
|
74
|
+
out = (pl.scan_parquet("huge.parquet")
|
|
75
|
+
.filter(pl.col("status") == "active")
|
|
76
|
+
.group_by("region").agg(pl.len())
|
|
77
|
+
.collect(engine="streaming"))
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
Streaming executes lazy plans only — keep the plan lazy end-to-end, with no intermediate
|
|
81
|
+
`.collect()` breaking it into eager pieces.
|
|
82
|
+
|
|
83
|
+
## Kernel-less fallback (uv one-shot)
|
|
84
|
+
|
|
85
|
+
On a harness with no persistent kernel, the same code runs as one-shots — batch several
|
|
86
|
+
questions per process, since each invocation pays spawn plus imports:
|
|
87
|
+
|
|
88
|
+
```bash
|
|
89
|
+
uv run --with duckdb --with polars --with pyarrow python -c "
|
|
90
|
+
import duckdb
|
|
91
|
+
import polars as pl
|
|
92
|
+
df = duckdb.sql(\"SELECT * FROM 'data.csv'\").pl()
|
|
93
|
+
print(df.group_by('category').agg(pl.len()).sort('category'))
|
|
94
|
+
"
|
|
95
|
+
```
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# uv Setup — Per-Platform
|
|
2
2
|
|
|
3
|
-
This skill
|
|
3
|
+
This skill's uv lane (kernel-less harnesses, isolated one-shots) runs through `uv run --with ...`, and `scripts/ensure-py-deps.sh` uses uv as its installer. If `uv --version` fails, set uv up with the automated scripts or the manual commands below, then verify.
|
|
4
4
|
|
|
5
5
|
## Automated (recommended)
|
|
6
6
|
|
|
@@ -0,0 +1,64 @@
|
|
|
1
|
+
# Visualization
|
|
2
|
+
|
|
3
|
+
A chart exists to answer a question at a glance. Render it with matplotlib (resident in most
|
|
4
|
+
Python kernels; `uv run --with matplotlib` otherwise), then look at it before delivering —
|
|
5
|
+
a chart nobody inspected is not evidence.
|
|
6
|
+
|
|
7
|
+
## When to chart
|
|
8
|
+
|
|
9
|
+
Chart when the user asked for one, and default to charting when the answer is a shape prose
|
|
10
|
+
cannot carry: a trend over time, a distribution, a comparison across many categories, a
|
|
11
|
+
relationship between variables. Skip the chart when a number or a five-row table answers the
|
|
12
|
+
question — decoration dilutes the answer.
|
|
13
|
+
|
|
14
|
+
## Chart type follows the question
|
|
15
|
+
|
|
16
|
+
| Question shape | Chart |
|
|
17
|
+
| --- | --- |
|
|
18
|
+
| How did X change over time? | line, datetime x-axis |
|
|
19
|
+
| Which categories are biggest? | horizontal bar, sorted by value |
|
|
20
|
+
| How is X distributed? | histogram (tune bin count) or box plot per group |
|
|
21
|
+
| Is X related to Y? | scatter; add a trend line only when it aids the eye |
|
|
22
|
+
| Composition of a whole? | stacked or 100% bar — pie only for four or fewer slices |
|
|
23
|
+
| Many series over time? | small multiples over one spaghetti chart |
|
|
24
|
+
|
|
25
|
+
## Quality bar — every chart
|
|
26
|
+
|
|
27
|
+
- Title states the finding ("Seoul overtook Busan in March"), not the dataset name.
|
|
28
|
+
- Axis labels carry units. Tick density stays readable: `fig.autofmt_xdate()` for dates,
|
|
29
|
+
rotate or abbreviate long category names.
|
|
30
|
+
- Size for the medium: inline chat reads well around `figsize=(10, 6)` at default dpi;
|
|
31
|
+
documents want `dpi=150` or more at export.
|
|
32
|
+
- `tight_layout()` (or `constrained_layout=True`) before saving — clipped labels are the
|
|
33
|
+
most common chart defect.
|
|
34
|
+
- Few series: label lines directly, or keep the legend inside empty plot space. Many
|
|
35
|
+
series: gray the context, color only the series that answers the question.
|
|
36
|
+
- The default color cycle is fine; avoid rainbow palettes and 3D. Sort categorical bars by
|
|
37
|
+
value, never alphabetically.
|
|
38
|
+
|
|
39
|
+
## CJK and other non-Latin text
|
|
40
|
+
|
|
41
|
+
Matplotlib's default font renders CJK as empty boxes (tofu). Set a fallback before plotting
|
|
42
|
+
whenever any label or title contains CJK:
|
|
43
|
+
|
|
44
|
+
```python
|
|
45
|
+
import platform
|
|
46
|
+
import matplotlib
|
|
47
|
+
cjk = {"Darwin": "AppleGothic", "Windows": "Malgun Gothic"}.get(platform.system(), "Noto Sans CJK KR")
|
|
48
|
+
matplotlib.rcParams["font.family"] = [cjk, "DejaVu Sans"]
|
|
49
|
+
matplotlib.rcParams["axes.unicode_minus"] = False # keeps the minus sign rendering
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
## Output contract
|
|
53
|
+
|
|
54
|
+
1. Save a PNG next to the work: `plt.savefig(path, dpi=150, bbox_inches="tight")`.
|
|
55
|
+
2. Also render inline when the surface displays rich output (kernels usually do).
|
|
56
|
+
3. Report the file path together with the answer.
|
|
57
|
+
|
|
58
|
+
## Visual QA — mandatory
|
|
59
|
+
|
|
60
|
+
Open the produced image — kernel display, or the harness's image-reading surface — and
|
|
61
|
+
check four things: labels readable and unclipped, no tofu or mojibake, nothing overlapping,
|
|
62
|
+
and the chart actually shows the finding the title claims. A failed check means fix and
|
|
63
|
+
re-render, not ship with a caveat. This one pass catches nearly every chart defect;
|
|
64
|
+
skipping it is how tofu titles reach users.
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
#!/usr/bin/env bash
|
|
2
|
+
# Install @duckdb/node-api into a user-level cache (outside any repo) and print the
|
|
3
|
+
# absolute import path as the ONLY stdout line. Idempotent: re-runs reuse the install.
|
|
4
|
+
set -euo pipefail
|
|
5
|
+
|
|
6
|
+
log() { printf '[ensure-js-deps] %s\n' "$*" >&2; }
|
|
7
|
+
|
|
8
|
+
CACHE_DIR="${OMO_DATA_SCIENTIST_CACHE:-$HOME/.cache/omo-data-scientist}"
|
|
9
|
+
IMPORT_PATH="$CACHE_DIR/node_modules/@duckdb/node-api/lib/index.js"
|
|
10
|
+
|
|
11
|
+
if ! command -v bun >/dev/null 2>&1; then
|
|
12
|
+
log "bun is required (https://bun.sh); install it, or use the uv lane instead."
|
|
13
|
+
exit 1
|
|
14
|
+
fi
|
|
15
|
+
|
|
16
|
+
if [ ! -f "$IMPORT_PATH" ]; then
|
|
17
|
+
log "installing @duckdb/node-api into $CACHE_DIR"
|
|
18
|
+
mkdir -p "$CACHE_DIR"
|
|
19
|
+
[ -f "$CACHE_DIR/package.json" ] || printf '{"name":"omo-data-scientist-cache","private":true}\n' > "$CACHE_DIR/package.json"
|
|
20
|
+
(cd "$CACHE_DIR" && bun add @duckdb/node-api 1>&2)
|
|
21
|
+
fi
|
|
22
|
+
|
|
23
|
+
if [ ! -f "$IMPORT_PATH" ]; then
|
|
24
|
+
log "install finished but $IMPORT_PATH is missing; inspect $CACHE_DIR"
|
|
25
|
+
exit 1
|
|
26
|
+
fi
|
|
27
|
+
|
|
28
|
+
printf '%s\n' "$IMPORT_PATH"
|
|
@@ -0,0 +1,37 @@
|
|
|
1
|
+
#!/usr/bin/env bash
|
|
2
|
+
# Install polars + pyarrow for a given Python interpreter into a user-level cache
|
|
3
|
+
# (never mutating the interpreter itself) and print the site directory as the ONLY
|
|
4
|
+
# stdout line. Idempotent: re-runs reuse the install. arg1 = python executable
|
|
5
|
+
# (default: python3); pass the kernel's sys.executable for kernel use.
|
|
6
|
+
set -euo pipefail
|
|
7
|
+
|
|
8
|
+
log() { printf '[ensure-py-deps] %s\n' "$*" >&2; }
|
|
9
|
+
|
|
10
|
+
PYTHON_BIN="${1:-python3}"
|
|
11
|
+
|
|
12
|
+
if ! command -v uv >/dev/null 2>&1; then
|
|
13
|
+
log "uv is required (see references/uv-setup.md); install it first."
|
|
14
|
+
exit 1
|
|
15
|
+
fi
|
|
16
|
+
if ! command -v "$PYTHON_BIN" >/dev/null 2>&1 && [ ! -x "$PYTHON_BIN" ]; then
|
|
17
|
+
log "python executable not found: $PYTHON_BIN"
|
|
18
|
+
exit 1
|
|
19
|
+
fi
|
|
20
|
+
|
|
21
|
+
TAG="$("$PYTHON_BIN" -c 'import sys; print(f"cp{sys.version_info[0]}{sys.version_info[1]}")')" \
|
|
22
|
+
|| { log "not a working python interpreter: $PYTHON_BIN"; exit 1; }
|
|
23
|
+
CACHE_DIR="${OMO_DATA_SCIENTIST_CACHE:-$HOME/.cache/omo-data-scientist}"
|
|
24
|
+
SITE_DIR="$CACHE_DIR/py-$TAG"
|
|
25
|
+
|
|
26
|
+
if [ ! -d "$SITE_DIR/polars" ] || [ ! -d "$SITE_DIR/pyarrow" ]; then
|
|
27
|
+
log "installing polars + pyarrow for $TAG into $SITE_DIR"
|
|
28
|
+
mkdir -p "$SITE_DIR"
|
|
29
|
+
uv pip install --python "$PYTHON_BIN" --target "$SITE_DIR" polars pyarrow 1>&2
|
|
30
|
+
fi
|
|
31
|
+
|
|
32
|
+
if [ ! -d "$SITE_DIR/polars" ]; then
|
|
33
|
+
log "install finished but $SITE_DIR/polars is missing; inspect $SITE_DIR"
|
|
34
|
+
exit 1
|
|
35
|
+
fi
|
|
36
|
+
|
|
37
|
+
printf '%s\n' "$SITE_DIR"
|
|
@@ -161,7 +161,9 @@ Record the answer in the journal; Phase 5 opens by turning it into `design-spec.
|
|
|
161
161
|
|
|
162
162
|
## Phase 1 — Saturation wave
|
|
163
163
|
|
|
164
|
-
|
|
164
|
+
**When the user asked for MASS research, the wave is sized by the topic's angles, not by the roster ceiling.** "mass ulw research", "mulw research", "ulw mass research" — in any language — order over-collection that a team of 8 cannot produce. Where the harness has a dependency-graph surface (the `mass-ulw` skill and its `references/planning.md`), run collection through it instead: a 60+ node opening wave covering every angle the topic has, routed across the whole difficulty ladder in one graph — mechanical sweeps and per-item harvest batches on the cheapest tier, judgment-shaped angles a step up, cross-territory angles a step above that, and the deepest tier reserved for genuinely hairy contradictions. Each wave's EXPAND leads define the next wave's nodes, and the synthesis reduces through several parallel architect-tier nodes into ONE architect-tier reducer (the strongest reasoning tier substitutes wherever no architect tier exists). Absent that surface, hold the same shape with background workers in batched waves. Everything else in this skill still binds: the format gate, the journal, the claim graph, the convergence rules, and the delivery gates.
|
|
165
|
+
|
|
166
|
+
Otherwise launch the entire first wave in one turn — every axis at once, as team members if you formed a team, else as background workers. Sequential launches and "start with one and see" defeat the mode.
|
|
165
167
|
|
|
166
168
|
Scaling floor — more angles always justify more workers:
|
|
167
169
|
|
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
#!/usr/bin/env node
|
|
2
|
-
// omo-codex-install:819e258f7485f9ccdcbb64a54ac0097bf81cd8e2980d7f907017f224b1f99b98:
|
|
2
|
+
// omo-codex-install:819e258f7485f9ccdcbb64a54ac0097bf81cd8e2980d7f907017f224b1f99b98:c7cf2301eb62bcb863c4e5844198424f235d9a88b5621bbf03436d1ef66c3491
|
|
3
3
|
var __defProp = Object.defineProperty;
|
|
4
4
|
var __returnValue = (v) => v;
|
|
5
5
|
function __exportSetter(name, newValue) {
|
|
@@ -7977,7 +7977,7 @@ var package_default;
|
|
|
7977
7977
|
var init_package = __esm(() => {
|
|
7978
7978
|
package_default = {
|
|
7979
7979
|
name: "@oh-my-opencode/omo-codex",
|
|
7980
|
-
version: "5.0.0-beta.
|
|
7980
|
+
version: "5.0.0-beta.28",
|
|
7981
7981
|
type: "module",
|
|
7982
7982
|
private: true,
|
|
7983
7983
|
description: "Codex harness adapter for oh-my-openagent. Vendored Codex plugin namespace (omo) + TypeScript installer + telemetry.",
|
|
@@ -1,243 +1,103 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: data-scientist
|
|
3
|
-
description: "Expert data processing
|
|
3
|
+
description: "Expert data processing with a hybrid engine strategy: resident-kernel engines first - DuckDB plus a resident Python stack (Polars/numpy/matplotlib) in persistent js/py eval kernels where the harness has them, bun/uv one-shots elsewhere - and per-action placement judgment (in-memory vs streaming vs remote-in-place). Triggers: 'analyze the data', 'what is in this CSV/parquet/json', 'summarize this', 'group by', 'filter rows', 'sort by', 'join these files', 'merge datasets', 'time series trend', 'compare yesterday and today', 'distribution/histogram', 'correlation', 'clean duplicates', 'handle missing values', 'dataset larger than RAM', 'SQL query on files', 'DataFrame operations', 'chart/plot this data', DuckDB vs Polars selection, quick data exploration CLI. NOT for plain text/code inspection, configs, or tiny inline math."
|
|
4
4
|
---
|
|
5
5
|
|
|
6
|
-
# Data Scientist:
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
-
|
|
61
|
-
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
##
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
import duckdb
|
|
105
|
-
# Query file directly - no memory load
|
|
106
|
-
result = duckdb.sql("""
|
|
107
|
-
SELECT category, SUM(amount) as total
|
|
108
|
-
FROM 'data.csv'
|
|
109
|
-
GROUP BY category
|
|
110
|
-
""").pl() # .pl() -> Polars via Arrow. Requires pyarrow. Never .df() (pandas).
|
|
111
|
-
```
|
|
112
|
-
|
|
113
|
-
### Polars Lazy Evaluation
|
|
114
|
-
|
|
115
|
-
```python
|
|
116
|
-
import polars as pl
|
|
117
|
-
# Lazy scan - optimizes and executes once
|
|
118
|
-
result = (
|
|
119
|
-
pl.scan_csv('data.csv')
|
|
120
|
-
.filter(pl.col('value') > 100)
|
|
121
|
-
.sort('value', descending=True)
|
|
122
|
-
.collect()
|
|
123
|
-
)
|
|
124
|
-
```
|
|
125
|
-
|
|
126
|
-
### Zero-Copy DuckDB → Polars
|
|
127
|
-
|
|
128
|
-
```python
|
|
129
|
-
import duckdb
|
|
130
|
-
# Direct conversion via Arrow (pyarrow required in the package set)
|
|
131
|
-
df_polars = duckdb.sql("SELECT * FROM 'data.csv'").pl()
|
|
132
|
-
```
|
|
133
|
-
|
|
134
|
-
### Hybrid Approach
|
|
135
|
-
|
|
136
|
-
```python
|
|
137
|
-
import duckdb
|
|
138
|
-
import polars as pl
|
|
139
|
-
|
|
140
|
-
# Phase 1: DuckDB for joins
|
|
141
|
-
joined = duckdb.sql(
|
|
142
|
-
"SELECT * FROM 'orders.csv' o "
|
|
143
|
-
"JOIN 'customers.csv' c ON o.customer_id = c.customer_id"
|
|
144
|
-
).pl()
|
|
145
|
-
|
|
146
|
-
# Phase 2: Polars for filtering
|
|
147
|
-
filtered = joined.filter(pl.col('amount') > 100)
|
|
148
|
-
|
|
149
|
-
# Phase 3: Back to DuckDB for aggregation
|
|
150
|
-
duckdb.register('filtered_data', filtered)
|
|
151
|
-
final = duckdb.sql('SELECT category, SUM(amount) FROM filtered_data GROUP BY category').pl()
|
|
152
|
-
```
|
|
153
|
-
|
|
154
|
-
## Quick Query CLI
|
|
155
|
-
|
|
156
|
-
For ad-hoc data exploration, use the built-in query runner:
|
|
157
|
-
|
|
158
|
-
```bash
|
|
159
|
-
# SQL query (uses DuckDB)
|
|
160
|
-
uv run scripts/quick-query.py data.csv "SELECT category, COUNT(*) FROM data GROUP BY category"
|
|
161
|
-
|
|
162
|
-
# Filter expression — Polars SQL syntax, e.g. "amount > 100" (NOT Python: never passes through eval)
|
|
163
|
-
uv run scripts/quick-query.py data.csv --filter "amount > 100"
|
|
164
|
-
|
|
165
|
-
# Auto-describe (schema + stats)
|
|
166
|
-
uv run scripts/quick-query.py data.parquet --describe
|
|
167
|
-
```
|
|
168
|
-
|
|
169
|
-
Supports CSV, Parquet, JSON, NDJSON. Cross-platform (macOS, Linux, Windows). Excel files are not read directly — export to CSV or Parquet first.
|
|
170
|
-
|
|
171
|
-
## Reference Documentation
|
|
172
|
-
|
|
173
|
-
For detailed guidance, consult these reference files:
|
|
174
|
-
|
|
175
|
-
- **Environment setup per platform**: See [uv-setup.md](references/uv-setup.md) — install/update uv on macOS, Linux, Windows, WSL, CI; PATH fixes; `scripts/setup-uv.sh` / `scripts/setup-uv.ps1` automate it.
|
|
176
|
-
- **Performance benchmarks and operation detection**: See [performance-benchmarks.md](references/performance-benchmarks.md)
|
|
177
|
-
- **Integration patterns and best practices**: See [integration-patterns.md](references/integration-patterns.md)
|
|
178
|
-
- **Execution templates**: See [execution-templates.md](references/execution-templates.md)
|
|
179
|
-
- **Common scenarios**: See [common-scenarios.md](references/common-scenarios.md)
|
|
180
|
-
|
|
181
|
-
## Quality Assurance Process
|
|
182
|
-
|
|
183
|
-
### Before Execution
|
|
184
|
-
1. **Analyze request** → Detect operation types (filter, join, aggregate, etc.)
|
|
185
|
-
2. **Select optimal tool** → Apply decision tree based on detected operations
|
|
186
|
-
3. **Verify approach** → Confirm tool selection matches the benchmark heuristics
|
|
187
|
-
4. **Check package list** → Ensure numpy AND pyarrow are included
|
|
188
|
-
|
|
189
|
-
### During Execution
|
|
190
|
-
1. **Use lazy evaluation** when possible (Polars `scan_*`, DuckDB direct queries)
|
|
191
|
-
2. **Monitor for errors** and have fallback strategy ready
|
|
192
|
-
3. **Provide progress updates** for long operations
|
|
193
|
-
|
|
194
|
-
### After Execution
|
|
195
|
-
1. **Report performance** → Show processing time and row counts
|
|
196
|
-
2. **Validate results** → Confirm output matches expectations
|
|
197
|
-
3. **Document tool choice** → Explain why specific tool was selected
|
|
198
|
-
|
|
199
|
-
## Activation Context
|
|
200
|
-
|
|
201
|
-
**Automatic activation triggers:**
|
|
202
|
-
|
|
203
|
-
### Exploratory Questions
|
|
204
|
-
- "Analyze the data" / "What's in the data" / "What's in this file"
|
|
205
|
-
- "Show me the data" / "Take a look at this file" / "Check the file contents"
|
|
206
|
-
|
|
207
|
-
### Temporal/Historical Analysis
|
|
208
|
-
- "What happened in the past N days?" / "How's last week's data?"
|
|
209
|
-
- "What's the trend for the last 30 days?" / "Compare yesterday and today"
|
|
210
|
-
|
|
211
|
-
### Aggregation/Summary Requests
|
|
212
|
-
- "Summarize this" / "What's the total?" / "What's the average?"
|
|
213
|
-
- "Show by category" / "Show statistics" / "How many?"
|
|
214
|
-
|
|
215
|
-
### Filtering/Search Patterns
|
|
216
|
-
- "Show only above 100" / "Find specific conditions" / "Top 10"
|
|
217
|
-
|
|
218
|
-
### Comparison/Correlation
|
|
219
|
-
- "Compare A and B" / "What's the difference?" / "Is there a correlation?" / "Merge two files"
|
|
220
|
-
|
|
221
|
-
### Transformation/Cleaning
|
|
222
|
-
- "Clean this up" / "Remove duplicates" / "Handle missing values" / "Convert format"
|
|
223
|
-
|
|
224
|
-
### Technical Patterns
|
|
225
|
-
- Working with CSV, Parquet, JSON, NDJSON, or `.duckdb` files
|
|
226
|
-
- File paths ending in `.csv`, `.parquet`, `.json`, `.jsonl`, `.ndjson`, `.tsv`, `.duckdb`
|
|
227
|
-
- Requests involving calculations or aggregations
|
|
228
|
-
- Joining, filtering, sorting, or transforming datasets
|
|
229
|
-
- Processing large datasets that may exceed memory
|
|
230
|
-
- Comparing or analyzing data from multiple sources
|
|
231
|
-
- Performance-critical data operations
|
|
232
|
-
- SQL queries or DataFrame operations mentioned
|
|
233
|
-
|
|
234
|
-
### When NOT to Activate
|
|
235
|
-
- Simple file reading for text/code inspection (use the harness's file-read surface)
|
|
236
|
-
- Non-data files (images, videos, binaries)
|
|
237
|
-
- Configuration files (YAML, TOML, JSON configs) unless specifically for data analysis
|
|
238
|
-
- Small inline calculations (run them directly)
|
|
239
|
-
- Excel files — convert to CSV/Parquet first
|
|
240
|
-
|
|
241
|
-
---
|
|
242
|
-
|
|
243
|
-
**Core execution principle:** Always apply intelligent tool selection based on operation characteristics, never use pandas, and always include numpy and pyarrow in the execution environment.
|
|
6
|
+
# Data Scientist: Hybrid-Engine Data Processing
|
|
7
|
+
|
|
8
|
+
Answer data questions through the cheapest engine and surface that can prove the answer, and
|
|
9
|
+
decide where the computation should live before touching the data.
|
|
10
|
+
|
|
11
|
+
## Execution surfaces: resident kernel first
|
|
12
|
+
|
|
13
|
+
A persistent REPL/eval kernel (many harnesses expose one for JavaScript and Python) is the
|
|
14
|
+
default surface. Reason: each one-shot process pays roughly a second of spawn-plus-import
|
|
15
|
+
overhead and re-scans the input file, while a resident connection amortizes both — after a
|
|
16
|
+
one-time load, repeat queries return in milliseconds. Exploration is repeat queries, so this
|
|
17
|
+
difference dominates the session.
|
|
18
|
+
|
|
19
|
+
1. **JavaScript kernel (Bun)**: run `scripts/ensure-js-deps.sh` once; it prints the absolute
|
|
20
|
+
import path for `@duckdb/node-api`. Dynamic-import it, connect once, query across cells.
|
|
21
|
+
2. **Python kernel**: the default surface for Python work. duckdb/numpy/matplotlib are
|
|
22
|
+
typically resident; Polars and pyarrow come from `scripts/ensure-py-deps.sh`, which
|
|
23
|
+
installs them once into a user cache keyed to the kernel's interpreter —
|
|
24
|
+
`sys.path.insert` the printed directory and import. The interpreter itself is never
|
|
25
|
+
mutated.
|
|
26
|
+
3. **uv lane** (`uv run --with ...`): isolation for a heavy or crash-prone one-shot that
|
|
27
|
+
should not take the kernel down.
|
|
28
|
+
4. **No kernel** (plain-shell harness): the same engines as one-shots — `bun -e` for
|
|
29
|
+
DuckDB-js, `uv run python -c` for the Python stack — batching several questions per
|
|
30
|
+
process.
|
|
31
|
+
|
|
32
|
+
Per-surface patterns and pitfalls: read `references/execution-surfaces.md` before first use.
|
|
33
|
+
|
|
34
|
+
## Engine selection
|
|
35
|
+
|
|
36
|
+
- **DuckDB** for SQL-shaped work: direct file queries, joins, aggregation, subqueries,
|
|
37
|
+
window functions. It queries CSV/Parquet/JSON in place without loading, spills to disk
|
|
38
|
+
past its memory limit, and reads remote files with the same syntax.
|
|
39
|
+
- **Polars** when the pipeline is DataFrame-shaped: expression-chain transforms, reshapes,
|
|
40
|
+
streaming datasets past RAM — resident in the Python kernel via `ensure-py-deps.sh`.
|
|
41
|
+
Read `references/polars-lane.md` — the current 1.x API differs from widely-memorized
|
|
42
|
+
older spellings.
|
|
43
|
+
- **numpy** when numeric work goes beyond SQL/DataFrame aggregation: statistical tests,
|
|
44
|
+
linear algebra, FFT, random sampling.
|
|
45
|
+
- **matplotlib** for every chart — read `references/visualization.md` first; it carries the
|
|
46
|
+
quality bar and a mandatory visual check.
|
|
47
|
+
|
|
48
|
+
Performance folklore ("X is Nx faster at filtering") varies with data shape, cardinality,
|
|
49
|
+
and hardware. When the engine choice materially matters, measure on the actual data instead
|
|
50
|
+
of trusting remembered multipliers.
|
|
51
|
+
|
|
52
|
+
## Placement: decide where the computation lives
|
|
53
|
+
|
|
54
|
+
Probe before you compute — one cell: file size, free RAM, and (when unclear) a row count via
|
|
55
|
+
a direct scan. Then place the work:
|
|
56
|
+
|
|
57
|
+
- **Load into memory** when the working set stays within roughly a quarter of free RAM AND
|
|
58
|
+
the session will run repeated queries: `CREATE TABLE t AS SELECT ...` (or a collected
|
|
59
|
+
DataFrame) once, then iterate. One scan up front converts every later query from a file
|
|
60
|
+
re-scan into milliseconds.
|
|
61
|
+
- **Query in place / stream** when the question is single-pass, or the data exceeds RAM:
|
|
62
|
+
DuckDB reads files directly (`FROM 'data.csv'`); past RAM, cap DuckDB's memory and let it
|
|
63
|
+
spill, or use Polars' streaming engine in the Python kernel. NEVER load a larger-than-RAM
|
|
64
|
+
dataset fully into memory — swapping stalls the whole machine, while streaming merely
|
|
65
|
+
takes longer.
|
|
66
|
+
- **Query remotely, in place** when the data lives elsewhere: DuckDB reads http(s)/S3
|
|
67
|
+
Parquet and CSV with projection and predicate pushdown, so fetch the columns and rows the
|
|
68
|
+
question needs, never the whole file. When data sits on another machine you can execute
|
|
69
|
+
on, ship the query to the data and return the small result. Rule: result much smaller
|
|
70
|
+
than data — move the query; repeated local iteration planned — move a pruned copy of the
|
|
71
|
+
data once.
|
|
72
|
+
|
|
73
|
+
Sizing heuristics and recipes: `references/placement.md`.
|
|
74
|
+
|
|
75
|
+
## Hard rules
|
|
76
|
+
|
|
77
|
+
- **NEVER use pandas.** DuckDB and Polars beat it decisively on every workload this skill
|
|
78
|
+
covers, and the environments this skill assumes do not ship it — `.df()` on a DuckDB
|
|
79
|
+
result raises unless pandas is installed; convert with `.pl()` via Arrow instead.
|
|
80
|
+
- Excel files are not read directly: export to CSV or Parquet first.
|
|
81
|
+
|
|
82
|
+
## Output contract
|
|
83
|
+
|
|
84
|
+
Answer the question; report row counts and timing for anything heavy; then stop — no bonus
|
|
85
|
+
charts, no extra exploration passes beyond what the question needed. Chart when asked, or
|
|
86
|
+
when the answer is a shape (trend, distribution, comparison) that prose cannot carry — then
|
|
87
|
+
follow `references/visualization.md` including its visual QA step.
|
|
88
|
+
|
|
89
|
+
## References
|
|
90
|
+
|
|
91
|
+
| Read | When |
|
|
92
|
+
| --- | --- |
|
|
93
|
+
| `references/execution-surfaces.md` | before the first query on any surface: kernel patterns, one-shot recipes, escalation rules |
|
|
94
|
+
| `references/polars-lane.md` | DataFrame-shaped pipeline or data past RAM: current API, Arrow handoff, package sets |
|
|
95
|
+
| `references/placement.md` | before heavy or remote work: sizing probe, memory limits, remote reads |
|
|
96
|
+
| `references/visualization.md` | before any chart: type selection, quality bar, CJK fonts, visual QA |
|
|
97
|
+
| `references/uv-setup.md` | uv missing or broken on this machine |
|
|
98
|
+
|
|
99
|
+
## CLI fallback
|
|
100
|
+
|
|
101
|
+
When no kernel or REPL surface exists, `uv run scripts/quick-query.py <file> [SQL]`
|
|
102
|
+
(`--filter <polars-sql-expr>`, `--describe`) answers ad-hoc questions with zero code.
|
|
103
|
+
Supports CSV, Parquet, JSON, NDJSON.
|