oh-my-opencode 5.0.0-beta.26 → 5.0.0-beta.28
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/cli/config-manager/parse-opencode-config-file.d.ts +2 -1
- package/dist/cli/doctor/checks/system-plugin.d.ts +4 -3
- package/dist/cli/index.js +59 -45
- package/dist/cli-node/index.js +82 -45
- package/dist/hooks/todo-continuation-enforcer/types.d.ts +1 -0
- package/dist/index.js +53 -28
- package/dist/shared/index.d.ts +1 -0
- package/dist/shared/legacy-plugin-warning.d.ts +2 -1
- package/dist/shared/plugin-entry-migrator.d.ts +4 -3
- package/dist/shared/plugin-entry-shape.d.ts +6 -0
- package/dist/skills/data-scientist/SKILL.md +99 -239
- package/dist/skills/data-scientist/references/execution-surfaces.md +91 -0
- package/dist/skills/data-scientist/references/placement.md +74 -0
- package/dist/skills/data-scientist/references/polars-lane.md +95 -0
- package/dist/skills/data-scientist/references/uv-setup.md +1 -1
- package/dist/skills/data-scientist/references/visualization.md +64 -0
- package/dist/skills/data-scientist/scripts/ensure-js-deps.sh +28 -0
- package/dist/skills/data-scientist/scripts/ensure-py-deps.sh +37 -0
- package/dist/skills/ulw-research/SKILL.md +3 -1
- package/dist/tui.d.ts +12 -0
- package/dist/tui.js +3 -3
- package/package.json +15 -15
- package/packages/lsp-core/src/request-context.test.ts +29 -0
- package/packages/lsp-core/src/request-context.ts +1 -1
- package/packages/lsp-daemon/dist/cli.js +1 -1
- package/packages/lsp-daemon/dist/client.js +1 -1
- package/packages/lsp-daemon/dist/index.js +1 -1
- package/packages/lsp-tools-mcp/dist/cli.js +1 -1
- package/packages/lsp-tools-mcp/dist/mcp.js +1 -1
- package/packages/lsp-tools-mcp/dist/request-context.js +1 -1
- package/packages/lsp-tools-mcp/dist/tools.js +1 -1
- package/packages/omo-codex/plugin/.codex-plugin/plugin.json +1 -1
- package/packages/omo-codex/plugin/components/bootstrap/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/bootstrap/package.json +1 -1
- package/packages/omo-codex/plugin/components/codegraph/package.json +1 -1
- package/packages/omo-codex/plugin/components/comment-checker/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/comment-checker/package.json +1 -1
- package/packages/omo-codex/plugin/components/git-bash/hooks/hooks.json +2 -2
- package/packages/omo-codex/plugin/components/git-bash/package.json +1 -1
- package/packages/omo-codex/plugin/components/lazycodex-executor-verify/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/lazycodex-executor-verify/package.json +1 -1
- package/packages/omo-codex/plugin/components/lsp/dist/.omo-runtime-manifest.json +2 -2
- package/packages/omo-codex/plugin/components/lsp/hooks/hooks.json +2 -2
- package/packages/omo-codex/plugin/components/lsp/package.json +1 -1
- package/packages/omo-codex/plugin/components/rules/hooks/hooks.json +4 -4
- package/packages/omo-codex/plugin/components/rules/package.json +1 -1
- package/packages/omo-codex/plugin/components/teammode/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/teammode/package.json +1 -1
- package/packages/omo-codex/plugin/components/telemetry/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/telemetry/package.json +1 -1
- package/packages/omo-codex/plugin/components/ultrawork/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/ultrawork/package.json +1 -1
- package/packages/omo-codex/plugin/components/ulw-execute-continuation/hooks/hooks.json +2 -2
- package/packages/omo-codex/plugin/components/ulw-execute-continuation/package.json +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/hooks/hooks.json +4 -4
- package/packages/omo-codex/plugin/components/ulw-loop/package.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-compact-resetting-git-bash-mcp-reminder.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-compact-resetting-lsp-diagnostics-cache.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-compact-resetting-project-rule-cache.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-checking-codegraph-init-guidance.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-checking-comments.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-checking-lsp-diagnostics.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-checking-thread-title-hygiene.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-matching-project-rules.json +1 -1
- package/packages/omo-codex/plugin/hooks/pre-tool-use-enforcing-unlimited-goal-budget.json +1 -1
- package/packages/omo-codex/plugin/hooks/pre-tool-use-guarding-ulw-loop-spawns.json +1 -1
- package/packages/omo-codex/plugin/hooks/pre-tool-use-recommending-git-bash-mcp.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-checking-auto-update.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-checking-bootstrap-provisioning.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-checking-codegraph-bootstrap.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-loading-project-rules.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-recording-session-telemetry.json +1 -1
- package/packages/omo-codex/plugin/hooks/stop-checking-ulw-execute-continuation.json +1 -1
- package/packages/omo-codex/plugin/hooks/stop-checking-ulw-loop-resume.json +1 -1
- package/packages/omo-codex/plugin/hooks/subagent-stop-checking-ulw-execute-continuation.json +1 -1
- package/packages/omo-codex/plugin/hooks/subagent-stop-verifying-lazycodex-executor-evidence.json +1 -1
- package/packages/omo-codex/plugin/hooks/user-prompt-submit-checking-ultrawork-trigger.json +1 -1
- package/packages/omo-codex/plugin/hooks/user-prompt-submit-checking-ulw-loop-steering.json +1 -1
- package/packages/omo-codex/plugin/hooks/user-prompt-submit-loading-project-rules.json +1 -1
- package/packages/omo-codex/plugin/package-lock.json +13 -13
- package/packages/omo-codex/plugin/package.json +1 -1
- package/packages/omo-codex/plugin/skills/data-scientist/SKILL.md +99 -239
- package/packages/omo-codex/plugin/skills/data-scientist/references/execution-surfaces.md +91 -0
- package/packages/omo-codex/plugin/skills/data-scientist/references/placement.md +74 -0
- package/packages/omo-codex/plugin/skills/data-scientist/references/polars-lane.md +95 -0
- package/packages/omo-codex/plugin/skills/data-scientist/references/uv-setup.md +1 -1
- package/packages/omo-codex/plugin/skills/data-scientist/references/visualization.md +64 -0
- package/packages/omo-codex/plugin/skills/data-scientist/scripts/ensure-js-deps.sh +28 -0
- package/packages/omo-codex/plugin/skills/data-scientist/scripts/ensure-py-deps.sh +37 -0
- package/packages/omo-codex/plugin/skills/ulw-research/SKILL.md +3 -1
- package/packages/omo-codex/scripts/install-dist/install-local.mjs +2 -2
- package/packages/shared-skills/skills/data-scientist/SKILL.md +99 -239
- package/packages/shared-skills/skills/data-scientist/references/execution-surfaces.md +91 -0
- package/packages/shared-skills/skills/data-scientist/references/placement.md +74 -0
- package/packages/shared-skills/skills/data-scientist/references/polars-lane.md +95 -0
- package/packages/shared-skills/skills/data-scientist/references/uv-setup.md +1 -1
- package/packages/shared-skills/skills/data-scientist/references/visualization.md +64 -0
- package/packages/shared-skills/skills/data-scientist/scripts/ensure-js-deps.sh +28 -0
- package/packages/shared-skills/skills/data-scientist/scripts/ensure-py-deps.sh +37 -0
- package/packages/shared-skills/skills/ulw-research/SKILL.md +3 -1
- package/dist/skills/data-scientist/references/common-scenarios.md +0 -176
- package/dist/skills/data-scientist/references/execution-templates.md +0 -197
- package/dist/skills/data-scientist/references/integration-patterns.md +0 -153
- package/dist/skills/data-scientist/references/performance-benchmarks.md +0 -37
- package/packages/omo-codex/plugin/skills/data-scientist/references/common-scenarios.md +0 -176
- package/packages/omo-codex/plugin/skills/data-scientist/references/execution-templates.md +0 -197
- package/packages/omo-codex/plugin/skills/data-scientist/references/integration-patterns.md +0 -153
- package/packages/omo-codex/plugin/skills/data-scientist/references/performance-benchmarks.md +0 -37
- package/packages/shared-skills/skills/data-scientist/references/common-scenarios.md +0 -176
- package/packages/shared-skills/skills/data-scientist/references/execution-templates.md +0 -197
- package/packages/shared-skills/skills/data-scientist/references/integration-patterns.md +0 -153
- package/packages/shared-skills/skills/data-scientist/references/performance-benchmarks.md +0 -37
|
@@ -1,12 +1,12 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@sisyphuslabs/omo-codex-plugin",
|
|
3
|
-
"version": "5.0.0-beta.
|
|
3
|
+
"version": "5.0.0-beta.28",
|
|
4
4
|
"lockfileVersion": 3,
|
|
5
5
|
"requires": true,
|
|
6
6
|
"packages": {
|
|
7
7
|
"": {
|
|
8
8
|
"name": "@sisyphuslabs/omo-codex-plugin",
|
|
9
|
-
"version": "5.0.0-beta.
|
|
9
|
+
"version": "5.0.0-beta.28",
|
|
10
10
|
"workspaces": [
|
|
11
11
|
"components/codegraph",
|
|
12
12
|
"components/comment-checker",
|
|
@@ -101,7 +101,7 @@
|
|
|
101
101
|
},
|
|
102
102
|
"components/codegraph": {
|
|
103
103
|
"name": "@sisyphuslabs/codex-codegraph",
|
|
104
|
-
"version": "5.0.0-beta.
|
|
104
|
+
"version": "5.0.0-beta.28",
|
|
105
105
|
"bin": {
|
|
106
106
|
"omo-codegraph": "dist/cli.js"
|
|
107
107
|
},
|
|
@@ -120,7 +120,7 @@
|
|
|
120
120
|
},
|
|
121
121
|
"components/comment-checker": {
|
|
122
122
|
"name": "@code-yeongyu/codex-comment-checker",
|
|
123
|
-
"version": "5.0.0-beta.
|
|
123
|
+
"version": "5.0.0-beta.28",
|
|
124
124
|
"license": "MIT",
|
|
125
125
|
"bin": {
|
|
126
126
|
"omo-comment-checker": "dist/cli.js"
|
|
@@ -141,7 +141,7 @@
|
|
|
141
141
|
},
|
|
142
142
|
"components/git-bash": {
|
|
143
143
|
"name": "@sisyphuslabs/codex-git-bash-hook",
|
|
144
|
-
"version": "5.0.0-beta.
|
|
144
|
+
"version": "5.0.0-beta.28",
|
|
145
145
|
"bin": {
|
|
146
146
|
"omo-git-bash-hook": "dist/cli.js"
|
|
147
147
|
},
|
|
@@ -155,7 +155,7 @@
|
|
|
155
155
|
},
|
|
156
156
|
"components/lazycodex-executor-verify": {
|
|
157
157
|
"name": "@code-yeongyu/codex-lazycodex-executor-verify",
|
|
158
|
-
"version": "5.0.0-beta.
|
|
158
|
+
"version": "5.0.0-beta.28",
|
|
159
159
|
"license": "MIT",
|
|
160
160
|
"bin": {
|
|
161
161
|
"lazycodex-executor-verify": "dist/cli.js"
|
|
@@ -172,7 +172,7 @@
|
|
|
172
172
|
},
|
|
173
173
|
"components/lsp": {
|
|
174
174
|
"name": "@code-yeongyu/codex-lsp",
|
|
175
|
-
"version": "5.0.0-beta.
|
|
175
|
+
"version": "5.0.0-beta.28",
|
|
176
176
|
"license": "MIT",
|
|
177
177
|
"dependencies": {
|
|
178
178
|
"@code-yeongyu/lsp-daemon": "file:../../../../lsp-daemon",
|
|
@@ -193,7 +193,7 @@
|
|
|
193
193
|
},
|
|
194
194
|
"components/rules": {
|
|
195
195
|
"name": "@code-yeongyu/codex-rules",
|
|
196
|
-
"version": "5.0.0-beta.
|
|
196
|
+
"version": "5.0.0-beta.28",
|
|
197
197
|
"license": "MIT",
|
|
198
198
|
"dependencies": {
|
|
199
199
|
"picomatch": "^4.0.7"
|
|
@@ -215,7 +215,7 @@
|
|
|
215
215
|
},
|
|
216
216
|
"components/teammode": {
|
|
217
217
|
"name": "@sisyphuslabs/codex-teammode",
|
|
218
|
-
"version": "5.0.0-beta.
|
|
218
|
+
"version": "5.0.0-beta.28",
|
|
219
219
|
"devDependencies": {
|
|
220
220
|
"@types/node": "^26.2.0",
|
|
221
221
|
"bun-types": "^1.4.0",
|
|
@@ -228,7 +228,7 @@
|
|
|
228
228
|
},
|
|
229
229
|
"components/telemetry": {
|
|
230
230
|
"name": "@code-yeongyu/codex-telemetry",
|
|
231
|
-
"version": "5.0.0-beta.
|
|
231
|
+
"version": "5.0.0-beta.28",
|
|
232
232
|
"license": "MIT",
|
|
233
233
|
"bin": {
|
|
234
234
|
"omo-telemetry": "dist/cli.js"
|
|
@@ -246,7 +246,7 @@
|
|
|
246
246
|
},
|
|
247
247
|
"components/ultrawork": {
|
|
248
248
|
"name": "@code-yeongyu/codex-ultrawork",
|
|
249
|
-
"version": "5.0.0-beta.
|
|
249
|
+
"version": "5.0.0-beta.28",
|
|
250
250
|
"license": "MIT",
|
|
251
251
|
"bin": {
|
|
252
252
|
"omo-ultrawork": "dist/cli.js"
|
|
@@ -264,7 +264,7 @@
|
|
|
264
264
|
},
|
|
265
265
|
"components/ulw-execute-continuation": {
|
|
266
266
|
"name": "@code-yeongyu/codex-ulw-execute-continuation",
|
|
267
|
-
"version": "5.0.0-beta.
|
|
267
|
+
"version": "5.0.0-beta.28",
|
|
268
268
|
"license": "MIT",
|
|
269
269
|
"bin": {
|
|
270
270
|
"omo-ulw-execute-continuation": "dist/cli.js"
|
|
@@ -281,7 +281,7 @@
|
|
|
281
281
|
},
|
|
282
282
|
"components/ulw-loop": {
|
|
283
283
|
"name": "@code-yeongyu/codex-ulw-loop",
|
|
284
|
-
"version": "5.0.0-beta.
|
|
284
|
+
"version": "5.0.0-beta.28",
|
|
285
285
|
"license": "MIT",
|
|
286
286
|
"bin": {
|
|
287
287
|
"omo-ulw-loop": "dist/cli.js",
|
|
@@ -1,243 +1,103 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: data-scientist
|
|
3
|
-
description: "Expert data processing
|
|
3
|
+
description: "Expert data processing with a hybrid engine strategy: resident-kernel engines first - DuckDB plus a resident Python stack (Polars/numpy/matplotlib) in persistent js/py eval kernels where the harness has them, bun/uv one-shots elsewhere - and per-action placement judgment (in-memory vs streaming vs remote-in-place). Triggers: 'analyze the data', 'what is in this CSV/parquet/json', 'summarize this', 'group by', 'filter rows', 'sort by', 'join these files', 'merge datasets', 'time series trend', 'compare yesterday and today', 'distribution/histogram', 'correlation', 'clean duplicates', 'handle missing values', 'dataset larger than RAM', 'SQL query on files', 'DataFrame operations', 'chart/plot this data', DuckDB vs Polars selection, quick data exploration CLI. NOT for plain text/code inspection, configs, or tiny inline math."
|
|
4
4
|
---
|
|
5
5
|
|
|
6
|
-
# Data Scientist:
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
-
|
|
61
|
-
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
##
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
import duckdb
|
|
105
|
-
# Query file directly - no memory load
|
|
106
|
-
result = duckdb.sql("""
|
|
107
|
-
SELECT category, SUM(amount) as total
|
|
108
|
-
FROM 'data.csv'
|
|
109
|
-
GROUP BY category
|
|
110
|
-
""").pl() # .pl() -> Polars via Arrow. Requires pyarrow. Never .df() (pandas).
|
|
111
|
-
```
|
|
112
|
-
|
|
113
|
-
### Polars Lazy Evaluation
|
|
114
|
-
|
|
115
|
-
```python
|
|
116
|
-
import polars as pl
|
|
117
|
-
# Lazy scan - optimizes and executes once
|
|
118
|
-
result = (
|
|
119
|
-
pl.scan_csv('data.csv')
|
|
120
|
-
.filter(pl.col('value') > 100)
|
|
121
|
-
.sort('value', descending=True)
|
|
122
|
-
.collect()
|
|
123
|
-
)
|
|
124
|
-
```
|
|
125
|
-
|
|
126
|
-
### Zero-Copy DuckDB → Polars
|
|
127
|
-
|
|
128
|
-
```python
|
|
129
|
-
import duckdb
|
|
130
|
-
# Direct conversion via Arrow (pyarrow required in the package set)
|
|
131
|
-
df_polars = duckdb.sql("SELECT * FROM 'data.csv'").pl()
|
|
132
|
-
```
|
|
133
|
-
|
|
134
|
-
### Hybrid Approach
|
|
135
|
-
|
|
136
|
-
```python
|
|
137
|
-
import duckdb
|
|
138
|
-
import polars as pl
|
|
139
|
-
|
|
140
|
-
# Phase 1: DuckDB for joins
|
|
141
|
-
joined = duckdb.sql(
|
|
142
|
-
"SELECT * FROM 'orders.csv' o "
|
|
143
|
-
"JOIN 'customers.csv' c ON o.customer_id = c.customer_id"
|
|
144
|
-
).pl()
|
|
145
|
-
|
|
146
|
-
# Phase 2: Polars for filtering
|
|
147
|
-
filtered = joined.filter(pl.col('amount') > 100)
|
|
148
|
-
|
|
149
|
-
# Phase 3: Back to DuckDB for aggregation
|
|
150
|
-
duckdb.register('filtered_data', filtered)
|
|
151
|
-
final = duckdb.sql('SELECT category, SUM(amount) FROM filtered_data GROUP BY category').pl()
|
|
152
|
-
```
|
|
153
|
-
|
|
154
|
-
## Quick Query CLI
|
|
155
|
-
|
|
156
|
-
For ad-hoc data exploration, use the built-in query runner:
|
|
157
|
-
|
|
158
|
-
```bash
|
|
159
|
-
# SQL query (uses DuckDB)
|
|
160
|
-
uv run scripts/quick-query.py data.csv "SELECT category, COUNT(*) FROM data GROUP BY category"
|
|
161
|
-
|
|
162
|
-
# Filter expression — Polars SQL syntax, e.g. "amount > 100" (NOT Python: never passes through eval)
|
|
163
|
-
uv run scripts/quick-query.py data.csv --filter "amount > 100"
|
|
164
|
-
|
|
165
|
-
# Auto-describe (schema + stats)
|
|
166
|
-
uv run scripts/quick-query.py data.parquet --describe
|
|
167
|
-
```
|
|
168
|
-
|
|
169
|
-
Supports CSV, Parquet, JSON, NDJSON. Cross-platform (macOS, Linux, Windows). Excel files are not read directly — export to CSV or Parquet first.
|
|
170
|
-
|
|
171
|
-
## Reference Documentation
|
|
172
|
-
|
|
173
|
-
For detailed guidance, consult these reference files:
|
|
174
|
-
|
|
175
|
-
- **Environment setup per platform**: See [uv-setup.md](references/uv-setup.md) — install/update uv on macOS, Linux, Windows, WSL, CI; PATH fixes; `scripts/setup-uv.sh` / `scripts/setup-uv.ps1` automate it.
|
|
176
|
-
- **Performance benchmarks and operation detection**: See [performance-benchmarks.md](references/performance-benchmarks.md)
|
|
177
|
-
- **Integration patterns and best practices**: See [integration-patterns.md](references/integration-patterns.md)
|
|
178
|
-
- **Execution templates**: See [execution-templates.md](references/execution-templates.md)
|
|
179
|
-
- **Common scenarios**: See [common-scenarios.md](references/common-scenarios.md)
|
|
180
|
-
|
|
181
|
-
## Quality Assurance Process
|
|
182
|
-
|
|
183
|
-
### Before Execution
|
|
184
|
-
1. **Analyze request** → Detect operation types (filter, join, aggregate, etc.)
|
|
185
|
-
2. **Select optimal tool** → Apply decision tree based on detected operations
|
|
186
|
-
3. **Verify approach** → Confirm tool selection matches the benchmark heuristics
|
|
187
|
-
4. **Check package list** → Ensure numpy AND pyarrow are included
|
|
188
|
-
|
|
189
|
-
### During Execution
|
|
190
|
-
1. **Use lazy evaluation** when possible (Polars `scan_*`, DuckDB direct queries)
|
|
191
|
-
2. **Monitor for errors** and have fallback strategy ready
|
|
192
|
-
3. **Provide progress updates** for long operations
|
|
193
|
-
|
|
194
|
-
### After Execution
|
|
195
|
-
1. **Report performance** → Show processing time and row counts
|
|
196
|
-
2. **Validate results** → Confirm output matches expectations
|
|
197
|
-
3. **Document tool choice** → Explain why specific tool was selected
|
|
198
|
-
|
|
199
|
-
## Activation Context
|
|
200
|
-
|
|
201
|
-
**Automatic activation triggers:**
|
|
202
|
-
|
|
203
|
-
### Exploratory Questions
|
|
204
|
-
- "Analyze the data" / "What's in the data" / "What's in this file"
|
|
205
|
-
- "Show me the data" / "Take a look at this file" / "Check the file contents"
|
|
206
|
-
|
|
207
|
-
### Temporal/Historical Analysis
|
|
208
|
-
- "What happened in the past N days?" / "How's last week's data?"
|
|
209
|
-
- "What's the trend for the last 30 days?" / "Compare yesterday and today"
|
|
210
|
-
|
|
211
|
-
### Aggregation/Summary Requests
|
|
212
|
-
- "Summarize this" / "What's the total?" / "What's the average?"
|
|
213
|
-
- "Show by category" / "Show statistics" / "How many?"
|
|
214
|
-
|
|
215
|
-
### Filtering/Search Patterns
|
|
216
|
-
- "Show only above 100" / "Find specific conditions" / "Top 10"
|
|
217
|
-
|
|
218
|
-
### Comparison/Correlation
|
|
219
|
-
- "Compare A and B" / "What's the difference?" / "Is there a correlation?" / "Merge two files"
|
|
220
|
-
|
|
221
|
-
### Transformation/Cleaning
|
|
222
|
-
- "Clean this up" / "Remove duplicates" / "Handle missing values" / "Convert format"
|
|
223
|
-
|
|
224
|
-
### Technical Patterns
|
|
225
|
-
- Working with CSV, Parquet, JSON, NDJSON, or `.duckdb` files
|
|
226
|
-
- File paths ending in `.csv`, `.parquet`, `.json`, `.jsonl`, `.ndjson`, `.tsv`, `.duckdb`
|
|
227
|
-
- Requests involving calculations or aggregations
|
|
228
|
-
- Joining, filtering, sorting, or transforming datasets
|
|
229
|
-
- Processing large datasets that may exceed memory
|
|
230
|
-
- Comparing or analyzing data from multiple sources
|
|
231
|
-
- Performance-critical data operations
|
|
232
|
-
- SQL queries or DataFrame operations mentioned
|
|
233
|
-
|
|
234
|
-
### When NOT to Activate
|
|
235
|
-
- Simple file reading for text/code inspection (use the harness's file-read surface)
|
|
236
|
-
- Non-data files (images, videos, binaries)
|
|
237
|
-
- Configuration files (YAML, TOML, JSON configs) unless specifically for data analysis
|
|
238
|
-
- Small inline calculations (run them directly)
|
|
239
|
-
- Excel files — convert to CSV/Parquet first
|
|
240
|
-
|
|
241
|
-
---
|
|
242
|
-
|
|
243
|
-
**Core execution principle:** Always apply intelligent tool selection based on operation characteristics, never use pandas, and always include numpy and pyarrow in the execution environment.
|
|
6
|
+
# Data Scientist: Hybrid-Engine Data Processing
|
|
7
|
+
|
|
8
|
+
Answer data questions through the cheapest engine and surface that can prove the answer, and
|
|
9
|
+
decide where the computation should live before touching the data.
|
|
10
|
+
|
|
11
|
+
## Execution surfaces: resident kernel first
|
|
12
|
+
|
|
13
|
+
A persistent REPL/eval kernel (many harnesses expose one for JavaScript and Python) is the
|
|
14
|
+
default surface. Reason: each one-shot process pays roughly a second of spawn-plus-import
|
|
15
|
+
overhead and re-scans the input file, while a resident connection amortizes both — after a
|
|
16
|
+
one-time load, repeat queries return in milliseconds. Exploration is repeat queries, so this
|
|
17
|
+
difference dominates the session.
|
|
18
|
+
|
|
19
|
+
1. **JavaScript kernel (Bun)**: run `scripts/ensure-js-deps.sh` once; it prints the absolute
|
|
20
|
+
import path for `@duckdb/node-api`. Dynamic-import it, connect once, query across cells.
|
|
21
|
+
2. **Python kernel**: the default surface for Python work. duckdb/numpy/matplotlib are
|
|
22
|
+
typically resident; Polars and pyarrow come from `scripts/ensure-py-deps.sh`, which
|
|
23
|
+
installs them once into a user cache keyed to the kernel's interpreter —
|
|
24
|
+
`sys.path.insert` the printed directory and import. The interpreter itself is never
|
|
25
|
+
mutated.
|
|
26
|
+
3. **uv lane** (`uv run --with ...`): isolation for a heavy or crash-prone one-shot that
|
|
27
|
+
should not take the kernel down.
|
|
28
|
+
4. **No kernel** (plain-shell harness): the same engines as one-shots — `bun -e` for
|
|
29
|
+
DuckDB-js, `uv run python -c` for the Python stack — batching several questions per
|
|
30
|
+
process.
|
|
31
|
+
|
|
32
|
+
Per-surface patterns and pitfalls: read `references/execution-surfaces.md` before first use.
|
|
33
|
+
|
|
34
|
+
## Engine selection
|
|
35
|
+
|
|
36
|
+
- **DuckDB** for SQL-shaped work: direct file queries, joins, aggregation, subqueries,
|
|
37
|
+
window functions. It queries CSV/Parquet/JSON in place without loading, spills to disk
|
|
38
|
+
past its memory limit, and reads remote files with the same syntax.
|
|
39
|
+
- **Polars** when the pipeline is DataFrame-shaped: expression-chain transforms, reshapes,
|
|
40
|
+
streaming datasets past RAM — resident in the Python kernel via `ensure-py-deps.sh`.
|
|
41
|
+
Read `references/polars-lane.md` — the current 1.x API differs from widely-memorized
|
|
42
|
+
older spellings.
|
|
43
|
+
- **numpy** when numeric work goes beyond SQL/DataFrame aggregation: statistical tests,
|
|
44
|
+
linear algebra, FFT, random sampling.
|
|
45
|
+
- **matplotlib** for every chart — read `references/visualization.md` first; it carries the
|
|
46
|
+
quality bar and a mandatory visual check.
|
|
47
|
+
|
|
48
|
+
Performance folklore ("X is Nx faster at filtering") varies with data shape, cardinality,
|
|
49
|
+
and hardware. When the engine choice materially matters, measure on the actual data instead
|
|
50
|
+
of trusting remembered multipliers.
|
|
51
|
+
|
|
52
|
+
## Placement: decide where the computation lives
|
|
53
|
+
|
|
54
|
+
Probe before you compute — one cell: file size, free RAM, and (when unclear) a row count via
|
|
55
|
+
a direct scan. Then place the work:
|
|
56
|
+
|
|
57
|
+
- **Load into memory** when the working set stays within roughly a quarter of free RAM AND
|
|
58
|
+
the session will run repeated queries: `CREATE TABLE t AS SELECT ...` (or a collected
|
|
59
|
+
DataFrame) once, then iterate. One scan up front converts every later query from a file
|
|
60
|
+
re-scan into milliseconds.
|
|
61
|
+
- **Query in place / stream** when the question is single-pass, or the data exceeds RAM:
|
|
62
|
+
DuckDB reads files directly (`FROM 'data.csv'`); past RAM, cap DuckDB's memory and let it
|
|
63
|
+
spill, or use Polars' streaming engine in the Python kernel. NEVER load a larger-than-RAM
|
|
64
|
+
dataset fully into memory — swapping stalls the whole machine, while streaming merely
|
|
65
|
+
takes longer.
|
|
66
|
+
- **Query remotely, in place** when the data lives elsewhere: DuckDB reads http(s)/S3
|
|
67
|
+
Parquet and CSV with projection and predicate pushdown, so fetch the columns and rows the
|
|
68
|
+
question needs, never the whole file. When data sits on another machine you can execute
|
|
69
|
+
on, ship the query to the data and return the small result. Rule: result much smaller
|
|
70
|
+
than data — move the query; repeated local iteration planned — move a pruned copy of the
|
|
71
|
+
data once.
|
|
72
|
+
|
|
73
|
+
Sizing heuristics and recipes: `references/placement.md`.
|
|
74
|
+
|
|
75
|
+
## Hard rules
|
|
76
|
+
|
|
77
|
+
- **NEVER use pandas.** DuckDB and Polars beat it decisively on every workload this skill
|
|
78
|
+
covers, and the environments this skill assumes do not ship it — `.df()` on a DuckDB
|
|
79
|
+
result raises unless pandas is installed; convert with `.pl()` via Arrow instead.
|
|
80
|
+
- Excel files are not read directly: export to CSV or Parquet first.
|
|
81
|
+
|
|
82
|
+
## Output contract
|
|
83
|
+
|
|
84
|
+
Answer the question; report row counts and timing for anything heavy; then stop — no bonus
|
|
85
|
+
charts, no extra exploration passes beyond what the question needed. Chart when asked, or
|
|
86
|
+
when the answer is a shape (trend, distribution, comparison) that prose cannot carry — then
|
|
87
|
+
follow `references/visualization.md` including its visual QA step.
|
|
88
|
+
|
|
89
|
+
## References
|
|
90
|
+
|
|
91
|
+
| Read | When |
|
|
92
|
+
| --- | --- |
|
|
93
|
+
| `references/execution-surfaces.md` | before the first query on any surface: kernel patterns, one-shot recipes, escalation rules |
|
|
94
|
+
| `references/polars-lane.md` | DataFrame-shaped pipeline or data past RAM: current API, Arrow handoff, package sets |
|
|
95
|
+
| `references/placement.md` | before heavy or remote work: sizing probe, memory limits, remote reads |
|
|
96
|
+
| `references/visualization.md` | before any chart: type selection, quality bar, CJK fonts, visual QA |
|
|
97
|
+
| `references/uv-setup.md` | uv missing or broken on this machine |
|
|
98
|
+
|
|
99
|
+
## CLI fallback
|
|
100
|
+
|
|
101
|
+
When no kernel or REPL surface exists, `uv run scripts/quick-query.py <file> [SQL]`
|
|
102
|
+
(`--filter <polars-sql-expr>`, `--describe`) answers ad-hoc questions with zero code.
|
|
103
|
+
Supports CSV, Parquet, JSON, NDJSON.
|
|
@@ -0,0 +1,91 @@
|
|
|
1
|
+
# Execution surfaces
|
|
2
|
+
|
|
3
|
+
How to run the engines on each surface, and when to escalate between them.
|
|
4
|
+
|
|
5
|
+
## Persistent kernel, JavaScript (Bun)
|
|
6
|
+
|
|
7
|
+
One-time setup per machine — the bundled script installs `@duckdb/node-api` into a user-level
|
|
8
|
+
cache outside any repo and prints the absolute import path (its only stdout line):
|
|
9
|
+
|
|
10
|
+
```bash
|
|
11
|
+
bash scripts/ensure-js-deps.sh # run from the skill directory
|
|
12
|
+
```
|
|
13
|
+
|
|
14
|
+
In the kernel — top-level `require` may not exist, dynamic import always works:
|
|
15
|
+
|
|
16
|
+
```js
|
|
17
|
+
const { DuckDBInstance } = await import("<printed path>");
|
|
18
|
+
const db = await DuckDBInstance.create(":memory:");
|
|
19
|
+
const conn = await db.connect();
|
|
20
|
+
const reader = await conn.runAndReadAll("SELECT category, SUM(v) AS total FROM 'data.csv' GROUP BY 1");
|
|
21
|
+
reader.getRowObjects(); // array of plain row objects
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
- The connection and any tables created live across cells — connect once per session, reuse.
|
|
25
|
+
- COUNT/SUM over integer columns return BigInt; convert (`Number(x)` or `String(x)`) before
|
|
26
|
+
`JSON.stringify`, which throws on BigInt.
|
|
27
|
+
- Bun builtins cover ingest gaps with zero installs: `Bun.JSONL.parse`, `Bun.JSON5.parse`,
|
|
28
|
+
`Bun.XML.parse`, `Bun.TOML.parse`, `Bun.Archive` for tarballs.
|
|
29
|
+
- nodejs-polars is NOT part of this skill's toolkit: its API lags the Python release by
|
|
30
|
+
major versions (option objects that work in Python throw napi type errors). Polars work
|
|
31
|
+
belongs to the Python kernel (below).
|
|
32
|
+
|
|
33
|
+
## Persistent kernel, Python (the default Python surface)
|
|
34
|
+
|
|
35
|
+
duckdb, numpy, and matplotlib are typically resident — import and use them directly.
|
|
36
|
+
Polars and pyarrow rarely ship with a kernel, so inject them once per session (run from the
|
|
37
|
+
skill directory; the script installs on first use, then just prints the path):
|
|
38
|
+
|
|
39
|
+
```python
|
|
40
|
+
import subprocess, sys
|
|
41
|
+
site = subprocess.run(["bash", "scripts/ensure-py-deps.sh", sys.executable],
|
|
42
|
+
capture_output=True, text=True, check=True).stdout.strip()
|
|
43
|
+
sys.path.insert(0, site)
|
|
44
|
+
import polars as pl
|
|
45
|
+
import pyarrow
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
- The install goes to a user cache keyed to the kernel's interpreter version; the
|
|
49
|
+
interpreter itself is never mutated (it is frequently an externally-managed system
|
|
50
|
+
Python, and mutating it breaks other tools).
|
|
51
|
+
- After injection the whole Python stack is resident: `duckdb.sql(...).pl()` hands off via
|
|
52
|
+
Arrow, `duckdb.register(name, df)` goes the other way, and Polars lazy pipelines run
|
|
53
|
+
in-kernel across cells.
|
|
54
|
+
- `duckdb.sql("SELECT ... FROM 'data.csv'")` queries files in place; without the injection,
|
|
55
|
+
keep results in DuckDB or fetch plain Python values (`.fetchall()`).
|
|
56
|
+
- matplotlib figures render natively in kernels that display rich output; also save a PNG so
|
|
57
|
+
the artifact survives the session.
|
|
58
|
+
|
|
59
|
+
## uv lane (fallback and isolation)
|
|
60
|
+
|
|
61
|
+
```bash
|
|
62
|
+
uv run --with duckdb --with polars --with pyarrow --with numpy python -c "<code>"
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
- Reach for it when there is no kernel, or when a heavy, crash-prone one-shot should not
|
|
66
|
+
run inside (and possibly take down) the kernel.
|
|
67
|
+
- Include exactly the packages the code imports, plus pyarrow whenever `.pl()` is used.
|
|
68
|
+
- Each invocation pays process spawn plus imports (roughly 0.3s warm) and re-reads its inputs —
|
|
69
|
+
fine for one-shots, wasteful for exploration loops.
|
|
70
|
+
- Past a few lines, a temp file beats `-c` quoting: write the script, `uv run script.py`.
|
|
71
|
+
|
|
72
|
+
## No kernel at all
|
|
73
|
+
|
|
74
|
+
Same engines, one process per batch of questions:
|
|
75
|
+
|
|
76
|
+
```bash
|
|
77
|
+
bun -e '<the JavaScript kernel pattern above>' # DuckDB via @duckdb/node-api
|
|
78
|
+
uv run --with duckdb python -c "<sql via duckdb.sql>" # DuckDB via Python
|
|
79
|
+
uv run scripts/quick-query.py data.csv "SELECT ..." # zero-code CLI fallback
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
## Escalation rules
|
|
83
|
+
|
|
84
|
+
Start on the resident kernel. Move a step down when a concrete need appears:
|
|
85
|
+
|
|
86
|
+
- polars/pyarrow missing from the kernel — inject via `ensure-py-deps.sh` (above), not a
|
|
87
|
+
uv one-shot.
|
|
88
|
+
- Crash-prone or memory-hungry one-shot that should not take the kernel down — uv lane.
|
|
89
|
+
- No kernel on this harness — one-shot recipes above.
|
|
90
|
+
- Data lives remotely or exceeds local RAM — read `placement.md` and move the query, not
|
|
91
|
+
the data.
|
|
@@ -0,0 +1,74 @@
|
|
|
1
|
+
# Placement: where should this computation live?
|
|
2
|
+
|
|
3
|
+
Decide before touching the data. Wrong placement wastes minutes (re-scanning a file queried
|
|
4
|
+
ten times) or kills the machine (loading a dataset larger than RAM and swapping).
|
|
5
|
+
|
|
6
|
+
## The probe (run first, once)
|
|
7
|
+
|
|
8
|
+
Three facts, one cell or script:
|
|
9
|
+
|
|
10
|
+
```python
|
|
11
|
+
import os, shutil, subprocess, sys
|
|
12
|
+
size = os.path.getsize("data.csv") # bytes on disk
|
|
13
|
+
disk_free = shutil.disk_usage(".").free # spill headroom
|
|
14
|
+
if sys.platform == "darwin":
|
|
15
|
+
ram = int(subprocess.run(["sysctl", "-n", "hw.memsize"], capture_output=True, text=True).stdout)
|
|
16
|
+
else:
|
|
17
|
+
ram = os.sysconf("SC_PAGE_SIZE") * os.sysconf("SC_PHYS_PAGES")
|
|
18
|
+
# row estimate without loading (DuckDB streams the scan):
|
|
19
|
+
# duckdb.sql("SELECT count(*) FROM 'data.csv'")
|
|
20
|
+
```
|
|
21
|
+
|
|
22
|
+
CSV typically expands 2-5x in memory (string columns dominate); Parquet expands less
|
|
23
|
+
predictably — compressed columns can inflate 10x. Estimate the working set from the
|
|
24
|
+
decompressed size of the columns the question actually touches, not the file size.
|
|
25
|
+
|
|
26
|
+
## In memory — load once, iterate
|
|
27
|
+
|
|
28
|
+
When the working set stays within roughly 25% of free RAM AND the session will run repeated
|
|
29
|
+
queries: load once (`CREATE TABLE t AS SELECT ...` in DuckDB, or a collected DataFrame),
|
|
30
|
+
then iterate. One scan up front converts every later query from a file re-scan into
|
|
31
|
+
milliseconds. Prune at load time — select only the needed columns, filter obvious dross —
|
|
32
|
+
so the resident table is the working set, not the raw file.
|
|
33
|
+
|
|
34
|
+
## In place / streaming — single pass, or bigger than RAM
|
|
35
|
+
|
|
36
|
+
- Single-pass questions: query the file directly (`FROM 'data.csv'`). Loading first is pure
|
|
37
|
+
waste.
|
|
38
|
+
- Bigger than RAM, SQL-shaped: cap DuckDB and let it spill —
|
|
39
|
+
|
|
40
|
+
```sql
|
|
41
|
+
SET memory_limit = '4GB';
|
|
42
|
+
SET temp_directory = '/tmp/duckdb_spill';
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
Aggregations, sorts, and window functions run out-of-core: slower, but bounded.
|
|
46
|
+
- Bigger than RAM, DataFrame-shaped: Polars streaming (`collect(engine="streaming")` on a
|
|
47
|
+
lazy plan) in the resident kernel — or a uv one-shot on kernel-less harnesses.
|
|
48
|
+
- Manual chunked loops (read N rows, process, repeat) are the last resort — the engines'
|
|
49
|
+
own out-of-core paths are faster and simpler than hand-rolled chunking.
|
|
50
|
+
|
|
51
|
+
## Remote, in place — move the query to the data
|
|
52
|
+
|
|
53
|
+
- Files behind http(s)/S3: DuckDB's httpfs extension reads Parquet and CSV remotely with
|
|
54
|
+
projection and predicate pushdown —
|
|
55
|
+
|
|
56
|
+
```sql
|
|
57
|
+
INSTALL httpfs; LOAD httpfs; -- one-time per environment
|
|
58
|
+
SELECT region, SUM(amount) FROM 'https://example.com/sales.parquet'
|
|
59
|
+
WHERE sale_date >= '2026-01-01' GROUP BY region;
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
Only matching row groups and referenced columns cross the network, not the file.
|
|
63
|
+
- Data on another machine you can execute on (a remote worker with more RAM, a box closer
|
|
64
|
+
to the data): run the query there and return the aggregate. A group-by result is
|
|
65
|
+
kilobytes; the source is gigabytes.
|
|
66
|
+
- Decision rule: result much smaller than data — move the query. Repeated local iteration
|
|
67
|
+
on one slice — move a pruned copy of that slice once, then work locally in memory.
|
|
68
|
+
|
|
69
|
+
## Hardware notes
|
|
70
|
+
|
|
71
|
+
- Both engines parallelize across all cores by default; leave that alone except on shared
|
|
72
|
+
machines (`SET threads = N` in DuckDB, `POLARS_MAX_THREADS` for Polars).
|
|
73
|
+
- Sustained swapping is the failure mode to avoid on memory-tight machines: when the probe
|
|
74
|
+
says the working set is close to free RAM, choose streaming, not hope.
|