oh-my-opencode 4.19.3 → 4.19.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/command/get-unpublished-changes.md +2 -0
- package/.agents/command/omomomo.md +1 -1
- package/.agents/command/publish.md +7 -0
- package/.agents/skills/get-unpublished-changes/SKILL.md +2 -0
- package/.agents/skills/hyperplan/SKILL.md +3 -3
- package/.agents/skills/omomomo/SKILL.md +1 -1
- package/.agents/skills/publish/SKILL.md +7 -0
- package/.opencode/command/get-unpublished-changes.md +2 -0
- package/.opencode/command/omomomo.md +1 -1
- package/.opencode/command/publish.md +7 -0
- package/.opencode/skills/hyperplan/SKILL.md +3 -3
- package/dist/agents/sisyphus-junior/agent.d.ts +1 -1
- package/dist/cli/doctor/checks/deprecated-reasoning-keys.d.ts +2 -0
- package/dist/cli/index.js +1625 -1368
- package/dist/cli-node/index.js +1625 -1368
- package/dist/config/schema/agent-overrides.d.ts +800 -0
- package/dist/config/schema/categories.d.ts +132 -0
- package/dist/config/schema/fallback-models.d.ts +50 -0
- package/dist/config/schema/oh-my-opencode-config.d.ts +818 -2
- package/dist/config-migration/index.d.ts +1 -0
- package/dist/config-migration/migration-plans.d.ts +1 -0
- package/dist/config-migration/reasoning-unification.d.ts +3 -0
- package/dist/features/team-mode/tools/lifecycle-test-fixture.d.ts +2 -0
- package/dist/hooks/codegraph-bootstrap/command-runner.d.ts +1 -0
- package/dist/hooks/model-fallback/next-fallback.d.ts +1 -0
- package/dist/hooks/runtime-fallback/constants.d.ts +1 -1
- package/dist/index.js +3558 -3241
- package/dist/oh-my-opencode.schema.json +1672 -8
- package/dist/plugin-handlers/prometheus-agent-config-builder.d.ts +1 -0
- package/dist/shared/agent-variant.d.ts +11 -0
- package/dist/shared/session-prompt-params-helpers.d.ts +6 -1
- package/dist/skills/coding-agent-sessions/SKILL.md +4 -3
- package/dist/skills/coding-agent-sessions/references/all-platforms.md +3 -1
- package/dist/skills/coding-agent-sessions/scripts/agent_sessions/aside_scanner.py +140 -0
- package/dist/skills/coding-agent-sessions/scripts/agent_sessions/scanners.py +3 -0
- package/dist/skills/data-scientist/SKILL.md +243 -0
- package/dist/skills/data-scientist/references/common-scenarios.md +176 -0
- package/dist/skills/data-scientist/references/execution-templates.md +197 -0
- package/dist/skills/data-scientist/references/integration-patterns.md +153 -0
- package/dist/skills/data-scientist/references/performance-benchmarks.md +37 -0
- package/dist/skills/data-scientist/references/uv-setup.md +78 -0
- package/dist/skills/data-scientist/scripts/quick-query.py +111 -0
- package/dist/skills/data-scientist/scripts/setup-uv.ps1 +53 -0
- package/dist/skills/data-scientist/scripts/setup-uv.sh +60 -0
- package/dist/skills/debugging/SKILL.md +1 -1
- package/dist/skills/programming/SKILL.md +1 -2
- package/dist/skills/ulw-plan/SKILL.md +1 -1
- package/dist/skills/ulw-research/SKILL.md +122 -11
- package/dist/tools/delegate-task/builtin-categories.d.ts +1 -0
- package/dist/tools/delegate-task/builtin-category-definition.d.ts +1 -0
- package/dist/tools/delegate-task/constants.d.ts +1 -1
- package/dist/tui.js +974 -1090
- package/package.json +13 -13
- package/packages/lsp-core/src/lsp/client-diagnostics-freshness.integration.test.ts +2 -2
- package/packages/omo-codex/plugin/.codex-plugin/plugin.json +1 -1
- package/packages/omo-codex/plugin/components/bootstrap/dist/cli.js +21 -3
- package/packages/omo-codex/plugin/components/bootstrap/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/bootstrap/package.json +1 -1
- package/packages/omo-codex/plugin/components/codegraph/dist/cli.js +279 -58
- package/packages/omo-codex/plugin/components/codegraph/dist/serve.js +230 -41
- package/packages/omo-codex/plugin/components/codegraph/package.json +1 -1
- package/packages/omo-codex/plugin/components/codegraph/test/hook.test.ts +6 -0
- package/packages/omo-codex/plugin/components/comment-checker/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/comment-checker/package.json +1 -1
- package/packages/omo-codex/plugin/components/git-bash/hooks/hooks.json +2 -2
- package/packages/omo-codex/plugin/components/git-bash/package.json +1 -1
- package/packages/omo-codex/plugin/components/lazycodex-executor-verify/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/lazycodex-executor-verify/package.json +1 -1
- package/packages/omo-codex/plugin/components/lsp/dist/.omo-runtime-manifest.json +2 -2
- package/packages/omo-codex/plugin/components/lsp/hooks/hooks.json +2 -2
- package/packages/omo-codex/plugin/components/lsp/package.json +1 -1
- package/packages/omo-codex/plugin/components/rules/dist/cli.js +1 -1
- package/packages/omo-codex/plugin/components/rules/hooks/hooks.json +4 -4
- package/packages/omo-codex/plugin/components/rules/package.json +1 -1
- package/packages/omo-codex/plugin/components/rules/src/post-compact-budget.ts +1 -1
- package/packages/omo-codex/plugin/components/start-work-continuation/hooks/hooks.json +2 -2
- package/packages/omo-codex/plugin/components/start-work-continuation/package.json +1 -1
- package/packages/omo-codex/plugin/components/teammode/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/teammode/package.json +1 -1
- package/packages/omo-codex/plugin/components/telemetry/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/telemetry/package.json +1 -1
- package/packages/omo-codex/plugin/components/ultrawork/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/ultrawork/package.json +1 -1
- package/packages/omo-codex/plugin/components/ultrawork/skills/ulw-plan/SKILL.md +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/hooks/hooks.json +4 -4
- package/packages/omo-codex/plugin/components/ulw-loop/package.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-compact-resetting-git-bash-mcp-reminder.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-compact-resetting-lsp-diagnostics-cache.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-compact-resetting-project-rule-cache.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-checking-codegraph-init-guidance.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-checking-comments.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-checking-lsp-diagnostics.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-checking-thread-title-hygiene.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-matching-project-rules.json +1 -1
- package/packages/omo-codex/plugin/hooks/pre-tool-use-enforcing-unlimited-goal-budget.json +1 -1
- package/packages/omo-codex/plugin/hooks/pre-tool-use-guarding-ulw-loop-spawns.json +1 -1
- package/packages/omo-codex/plugin/hooks/pre-tool-use-recommending-git-bash-mcp.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-checking-auto-update.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-checking-bootstrap-provisioning.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-checking-codegraph-bootstrap.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-loading-project-rules.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-recording-session-telemetry.json +1 -1
- package/packages/omo-codex/plugin/hooks/stop-checking-start-work-continuation.json +1 -1
- package/packages/omo-codex/plugin/hooks/stop-checking-ulw-loop-resume.json +1 -1
- package/packages/omo-codex/plugin/hooks/subagent-stop-checking-start-work-continuation.json +1 -1
- package/packages/omo-codex/plugin/hooks/subagent-stop-verifying-lazycodex-executor-evidence.json +1 -1
- package/packages/omo-codex/plugin/hooks/user-prompt-submit-checking-ultrawork-trigger.json +1 -1
- package/packages/omo-codex/plugin/hooks/user-prompt-submit-checking-ulw-loop-steering.json +1 -1
- package/packages/omo-codex/plugin/hooks/user-prompt-submit-loading-project-rules.json +1 -1
- package/packages/omo-codex/plugin/package-lock.json +13 -13
- package/packages/omo-codex/plugin/package.json +1 -1
- package/packages/omo-codex/plugin/scripts/sync-skills.mjs +6 -0
- package/packages/omo-codex/plugin/skills/coding-agent-sessions/SKILL.md +4 -3
- package/packages/omo-codex/plugin/skills/coding-agent-sessions/references/all-platforms.md +3 -1
- package/packages/omo-codex/plugin/skills/coding-agent-sessions/scripts/agent_sessions/aside_scanner.py +140 -0
- package/packages/omo-codex/plugin/skills/coding-agent-sessions/scripts/agent_sessions/scanners.py +3 -0
- package/packages/omo-codex/plugin/skills/data-scientist/SKILL.md +243 -0
- package/packages/omo-codex/plugin/skills/data-scientist/agents/openai.yaml +2 -0
- package/packages/omo-codex/plugin/skills/data-scientist/references/common-scenarios.md +176 -0
- package/packages/omo-codex/plugin/skills/data-scientist/references/execution-templates.md +197 -0
- package/packages/omo-codex/plugin/skills/data-scientist/references/integration-patterns.md +153 -0
- package/packages/omo-codex/plugin/skills/data-scientist/references/performance-benchmarks.md +37 -0
- package/packages/omo-codex/plugin/skills/data-scientist/references/uv-setup.md +78 -0
- package/packages/omo-codex/plugin/skills/data-scientist/scripts/quick-query.py +111 -0
- package/packages/omo-codex/plugin/skills/data-scientist/scripts/setup-uv.ps1 +53 -0
- package/packages/omo-codex/plugin/skills/data-scientist/scripts/setup-uv.sh +60 -0
- package/packages/omo-codex/plugin/skills/debugging/SKILL.md +1 -1
- package/packages/omo-codex/plugin/skills/programming/SKILL.md +1 -2
- package/packages/omo-codex/plugin/skills/ulw-plan/SKILL.md +1 -1
- package/packages/omo-codex/plugin/skills/ulw-research/SKILL.md +121 -11
- package/packages/omo-codex/plugin/test/sync-skills-test-support.mjs +7 -0
- package/packages/omo-codex/plugin/test/sync-skills.test.mjs +12 -0
- package/packages/omo-codex/scripts/install-dist/install-local.mjs +24 -5
- package/packages/shared-skills/skills/coding-agent-sessions/SKILL.md +4 -3
- package/packages/shared-skills/skills/coding-agent-sessions/references/all-platforms.md +3 -1
- package/packages/shared-skills/skills/coding-agent-sessions/scripts/agent_sessions/aside_scanner.py +140 -0
- package/packages/shared-skills/skills/coding-agent-sessions/scripts/agent_sessions/scanners.py +3 -0
- package/packages/shared-skills/skills/data-scientist/SKILL.md +243 -0
- package/packages/shared-skills/skills/data-scientist/references/common-scenarios.md +176 -0
- package/packages/shared-skills/skills/data-scientist/references/execution-templates.md +197 -0
- package/packages/shared-skills/skills/data-scientist/references/integration-patterns.md +153 -0
- package/packages/shared-skills/skills/data-scientist/references/performance-benchmarks.md +37 -0
- package/packages/shared-skills/skills/data-scientist/references/uv-setup.md +78 -0
- package/packages/shared-skills/skills/data-scientist/scripts/quick-query.py +111 -0
- package/packages/shared-skills/skills/data-scientist/scripts/setup-uv.ps1 +53 -0
- package/packages/shared-skills/skills/data-scientist/scripts/setup-uv.sh +60 -0
- package/packages/shared-skills/skills/debugging/SKILL.md +1 -1
- package/packages/shared-skills/skills/programming/SKILL.md +1 -2
- package/packages/shared-skills/skills/ulw-plan/SKILL.md +1 -1
- package/packages/shared-skills/skills/ulw-research/SKILL.md +122 -11
|
@@ -0,0 +1,176 @@
|
|
|
1
|
+
# Common Scenarios with Tool Selection
|
|
2
|
+
|
|
3
|
+
## Scenario 1: Simple Calculation
|
|
4
|
+
|
|
5
|
+
**Decision: Python (always)**
|
|
6
|
+
|
|
7
|
+
```python
|
|
8
|
+
uv run --with numpy python -c "
|
|
9
|
+
import numpy as np
|
|
10
|
+
result = np.sum([1, 2, 3, 4, 5])
|
|
11
|
+
print(f'Result: {result}')
|
|
12
|
+
"
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
## Scenario 2: CSV Quick Analysis
|
|
16
|
+
|
|
17
|
+
**Decision: DuckDB (direct query, no memory load)**
|
|
18
|
+
|
|
19
|
+
```python
|
|
20
|
+
uv run --with numpy --with duckdb python -c "
|
|
21
|
+
import duckdb
|
|
22
|
+
result = duckdb.sql('''
|
|
23
|
+
SELECT * FROM 'data.csv'
|
|
24
|
+
LIMIT 10
|
|
25
|
+
''').pl()
|
|
26
|
+
print(result)
|
|
27
|
+
"
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
## Scenario 3: Filter + Sort on Large Dataset
|
|
31
|
+
|
|
32
|
+
**Decision: Polars (128x faster filtering, 12x faster sorting)**
|
|
33
|
+
|
|
34
|
+
```python
|
|
35
|
+
uv run --with numpy --with polars python -c "
|
|
36
|
+
import polars as pl
|
|
37
|
+
result = (
|
|
38
|
+
pl.scan_csv('large.csv')
|
|
39
|
+
.filter(pl.col('value') > 1000)
|
|
40
|
+
.sort('value', descending=True)
|
|
41
|
+
.head(100)
|
|
42
|
+
.collect()
|
|
43
|
+
)
|
|
44
|
+
print(result)
|
|
45
|
+
"
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
## Scenario 4: Multi-Table Join + Aggregation
|
|
49
|
+
|
|
50
|
+
**Decision: DuckDB (best for joins and aggregations)**
|
|
51
|
+
|
|
52
|
+
```python
|
|
53
|
+
uv run --with numpy --with duckdb python -c "
|
|
54
|
+
import duckdb
|
|
55
|
+
result = duckdb.sql('''
|
|
56
|
+
SELECT
|
|
57
|
+
a.category,
|
|
58
|
+
COUNT(*) as count,
|
|
59
|
+
SUM(b.amount) as total
|
|
60
|
+
FROM 'table1.csv' a
|
|
61
|
+
JOIN 'table2.csv' b ON a.id = b.id
|
|
62
|
+
GROUP BY a.category
|
|
63
|
+
ORDER BY total DESC
|
|
64
|
+
''').pl()
|
|
65
|
+
print(result)
|
|
66
|
+
"
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
## Scenario 5: Data Exploration
|
|
70
|
+
|
|
71
|
+
**Decision: DuckDB for quick exploration**
|
|
72
|
+
|
|
73
|
+
```python
|
|
74
|
+
uv run --with numpy --with duckdb python -c "
|
|
75
|
+
import duckdb
|
|
76
|
+
|
|
77
|
+
# Show first few rows
|
|
78
|
+
print('**Sample Data**')
|
|
79
|
+
print(duckdb.sql('SELECT * FROM \"data.csv\" LIMIT 5').pl())
|
|
80
|
+
|
|
81
|
+
# Show summary statistics
|
|
82
|
+
print('\\n**Summary Statistics**')
|
|
83
|
+
print(duckdb.sql('DESCRIBE SELECT * FROM \"data.csv\"').pl())
|
|
84
|
+
|
|
85
|
+
# Show row count
|
|
86
|
+
print('\\n**Row Count**')
|
|
87
|
+
print(duckdb.sql('SELECT COUNT(*) as total_rows FROM \"data.csv\"').pl())
|
|
88
|
+
"
|
|
89
|
+
```
|
|
90
|
+
|
|
91
|
+
## Scenario 6: Time-Series Analysis
|
|
92
|
+
|
|
93
|
+
**Decision: DuckDB for aggregation + matplotlib for visualization**
|
|
94
|
+
|
|
95
|
+
```python
|
|
96
|
+
uv run --with numpy --with duckdb --with pyarrow --with matplotlib python -c "
|
|
97
|
+
import duckdb
|
|
98
|
+
import matplotlib.pyplot as plt
|
|
99
|
+
|
|
100
|
+
# Aggregate by date
|
|
101
|
+
result = duckdb.sql('''
|
|
102
|
+
SELECT
|
|
103
|
+
DATE_TRUNC('day', timestamp) as date,
|
|
104
|
+
COUNT(*) as count,
|
|
105
|
+
AVG(value) as avg_value
|
|
106
|
+
FROM 'timeseries.csv'
|
|
107
|
+
GROUP BY date
|
|
108
|
+
ORDER BY date
|
|
109
|
+
''').pl()
|
|
110
|
+
|
|
111
|
+
# Plot
|
|
112
|
+
plt.figure(figsize=(12, 6))
|
|
113
|
+
plt.subplot(2, 1, 1)
|
|
114
|
+
plt.plot(result['date'], result['count'])
|
|
115
|
+
plt.title('Daily Count')
|
|
116
|
+
|
|
117
|
+
plt.subplot(2, 1, 2)
|
|
118
|
+
plt.plot(result['date'], result['avg_value'])
|
|
119
|
+
plt.title('Daily Average Value')
|
|
120
|
+
|
|
121
|
+
plt.tight_layout()
|
|
122
|
+
plt.savefig('timeseries.png')
|
|
123
|
+
print('Saved to timeseries.png')
|
|
124
|
+
"
|
|
125
|
+
```
|
|
126
|
+
|
|
127
|
+
## Scenario 7: Complex Transformation
|
|
128
|
+
|
|
129
|
+
**Decision: Polars for efficient transformations**
|
|
130
|
+
|
|
131
|
+
```python
|
|
132
|
+
uv run --with numpy --with polars python -c "
|
|
133
|
+
import polars as pl
|
|
134
|
+
|
|
135
|
+
result = (
|
|
136
|
+
pl.scan_csv('data.csv')
|
|
137
|
+
.with_columns([
|
|
138
|
+
# Create new calculated columns
|
|
139
|
+
(pl.col('price') * pl.col('quantity')).alias('total'),
|
|
140
|
+
pl.col('date').str.strptime(pl.Date, '%Y-%m-%d').alias('parsed_date'),
|
|
141
|
+
pl.col('name').str.to_uppercase().alias('upper_name'),
|
|
142
|
+
])
|
|
143
|
+
.filter(pl.col('total') > 100)
|
|
144
|
+
.select(['parsed_date', 'upper_name', 'total'])
|
|
145
|
+
.collect()
|
|
146
|
+
)
|
|
147
|
+
|
|
148
|
+
print(result)
|
|
149
|
+
"
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
## Scenario 8: Large File Processing
|
|
153
|
+
|
|
154
|
+
**Decision: Polars streaming mode**
|
|
155
|
+
|
|
156
|
+
```python
|
|
157
|
+
uv run --with numpy --with polars python -c "
|
|
158
|
+
import polars as pl
|
|
159
|
+
|
|
160
|
+
# Process file larger than RAM
|
|
161
|
+
result = (
|
|
162
|
+
pl.scan_csv('huge_file.csv')
|
|
163
|
+
.filter(pl.col('active') == True)
|
|
164
|
+
.groupby('category')
|
|
165
|
+
.agg([
|
|
166
|
+
pl.count().alias('count'),
|
|
167
|
+
pl.sum('amount').alias('total'),
|
|
168
|
+
pl.mean('amount').alias('average'),
|
|
169
|
+
])
|
|
170
|
+
.collect(streaming=True) # Streaming mode
|
|
171
|
+
)
|
|
172
|
+
|
|
173
|
+
print(result)
|
|
174
|
+
print(f'\\nProcessed {result[\"count\"].sum():,} rows')
|
|
175
|
+
"
|
|
176
|
+
```
|
|
@@ -0,0 +1,197 @@
|
|
|
1
|
+
# Execution Templates
|
|
2
|
+
|
|
3
|
+
## Template 1: DuckDB for Simple/Complex SQL Queries
|
|
4
|
+
|
|
5
|
+
**Use when:**
|
|
6
|
+
- Simple aggregation on single file
|
|
7
|
+
- Complex multi-table joins
|
|
8
|
+
- Heavy GROUP BY operations
|
|
9
|
+
- Window functions with complex SQL logic
|
|
10
|
+
- Ad-hoc exploration queries
|
|
11
|
+
|
|
12
|
+
```python
|
|
13
|
+
uv run --with numpy --with duckdb python -c "
|
|
14
|
+
import duckdb
|
|
15
|
+
|
|
16
|
+
# Simple query - direct file access
|
|
17
|
+
result = duckdb.sql('''
|
|
18
|
+
SELECT
|
|
19
|
+
category,
|
|
20
|
+
COUNT(*) as count,
|
|
21
|
+
AVG(amount) as avg_amount,
|
|
22
|
+
SUM(amount) as total
|
|
23
|
+
FROM 'data.csv'
|
|
24
|
+
WHERE date >= '2024-01-01'
|
|
25
|
+
GROUP BY category
|
|
26
|
+
ORDER BY total DESC
|
|
27
|
+
LIMIT 10
|
|
28
|
+
''').pl()
|
|
29
|
+
|
|
30
|
+
print('**Results**')
|
|
31
|
+
print(result)
|
|
32
|
+
print(f'\\nProcessed {len(result)} categories')
|
|
33
|
+
"
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
## Template 2: Polars for Filtering & Sorting
|
|
37
|
+
|
|
38
|
+
**Use when:**
|
|
39
|
+
- Primary operation is filtering large dataset
|
|
40
|
+
- Sorting required
|
|
41
|
+
- Chain transformations
|
|
42
|
+
- Memory-efficient processing needed
|
|
43
|
+
|
|
44
|
+
```python
|
|
45
|
+
uv run --with numpy --with polars python -c "
|
|
46
|
+
import polars as pl
|
|
47
|
+
|
|
48
|
+
# Lazy evaluation for optimal performance
|
|
49
|
+
result = (
|
|
50
|
+
pl.scan_csv('data.csv') # Lazy scan
|
|
51
|
+
.filter(
|
|
52
|
+
(pl.col('amount') > 1000) &
|
|
53
|
+
(pl.col('status') == 'active') &
|
|
54
|
+
(pl.col('date') >= '2024-01-01')
|
|
55
|
+
)
|
|
56
|
+
.sort('amount', descending=True)
|
|
57
|
+
.head(100)
|
|
58
|
+
.collect() # Execute optimized plan
|
|
59
|
+
)
|
|
60
|
+
|
|
61
|
+
print('**Filtered and Sorted Results**')
|
|
62
|
+
print(result)
|
|
63
|
+
print(f'\\nFound {len(result)} matching rows')
|
|
64
|
+
"
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
## Template 3: Hybrid Approach (Best of Both)
|
|
68
|
+
|
|
69
|
+
**Use when:**
|
|
70
|
+
- Need joins AND heavy filtering
|
|
71
|
+
- Complex SQL followed by transformations
|
|
72
|
+
- Optimize different operation stages
|
|
73
|
+
|
|
74
|
+
```python
|
|
75
|
+
uv run --with numpy --with duckdb --with polars --with pyarrow python -c "
|
|
76
|
+
import duckdb
|
|
77
|
+
import polars as pl
|
|
78
|
+
|
|
79
|
+
print('Phase 1: DuckDB for complex join (3x faster)')
|
|
80
|
+
# DuckDB excels at joins
|
|
81
|
+
joined = duckdb.sql('''
|
|
82
|
+
SELECT
|
|
83
|
+
o.order_id,
|
|
84
|
+
o.amount,
|
|
85
|
+
c.customer_id,
|
|
86
|
+
c.region,
|
|
87
|
+
p.category
|
|
88
|
+
FROM 'orders.csv' o
|
|
89
|
+
JOIN 'customers.csv' c ON o.customer_id = c.customer_id
|
|
90
|
+
JOIN 'products.csv' p ON o.product_id = p.product_id
|
|
91
|
+
WHERE o.date >= '2024-01-01'
|
|
92
|
+
''').pl() # Convert to Polars
|
|
93
|
+
|
|
94
|
+
print(f'Joined {len(joined):,} rows')
|
|
95
|
+
|
|
96
|
+
print('\\nPhase 2: Polars for ultra-fast filtering (128x faster)')
|
|
97
|
+
# Polars excels at filtering
|
|
98
|
+
filtered = (
|
|
99
|
+
joined
|
|
100
|
+
.filter(
|
|
101
|
+
(pl.col('amount') > 100) &
|
|
102
|
+
(pl.col('region').is_in(['North', 'South', 'East']))
|
|
103
|
+
)
|
|
104
|
+
.with_columns([
|
|
105
|
+
(pl.col('amount') * 1.1).alias('amount_with_tax')
|
|
106
|
+
])
|
|
107
|
+
)
|
|
108
|
+
|
|
109
|
+
print(f'Filtered to {len(filtered):,} rows')
|
|
110
|
+
|
|
111
|
+
print('\\nPhase 3: DuckDB for final aggregation (4x faster)')
|
|
112
|
+
# Back to DuckDB for aggregation
|
|
113
|
+
duckdb.register('filtered_data', filtered)
|
|
114
|
+
final = duckdb.sql('''
|
|
115
|
+
SELECT
|
|
116
|
+
region,
|
|
117
|
+
category,
|
|
118
|
+
COUNT(DISTINCT customer_id) as customers,
|
|
119
|
+
SUM(amount_with_tax) as total_revenue,
|
|
120
|
+
AVG(amount_with_tax) as avg_transaction
|
|
121
|
+
FROM filtered_data
|
|
122
|
+
GROUP BY region, category
|
|
123
|
+
HAVING total_revenue > 10000
|
|
124
|
+
ORDER BY total_revenue DESC
|
|
125
|
+
''').pl()
|
|
126
|
+
|
|
127
|
+
print('\\n**Final Results**')
|
|
128
|
+
print(final)
|
|
129
|
+
"
|
|
130
|
+
```
|
|
131
|
+
|
|
132
|
+
## Template 4: Polars Streaming for Large Files
|
|
133
|
+
|
|
134
|
+
**Use when:**
|
|
135
|
+
- Dataset larger than available RAM
|
|
136
|
+
- Need to process data in batches
|
|
137
|
+
- Memory constraints
|
|
138
|
+
|
|
139
|
+
```python
|
|
140
|
+
uv run --with numpy --with polars python -c "
|
|
141
|
+
import polars as pl
|
|
142
|
+
|
|
143
|
+
# Streaming mode - processes data in chunks
|
|
144
|
+
result = (
|
|
145
|
+
pl.scan_csv('huge_file.csv')
|
|
146
|
+
.filter(pl.col('status') == 'active')
|
|
147
|
+
.with_columns([
|
|
148
|
+
(pl.col('amount') * 1.1).alias('adjusted_amount')
|
|
149
|
+
])
|
|
150
|
+
.groupby('category')
|
|
151
|
+
.agg([
|
|
152
|
+
pl.sum('adjusted_amount').alias('total'),
|
|
153
|
+
pl.count().alias('count')
|
|
154
|
+
])
|
|
155
|
+
.collect(streaming=True) # Streaming mode for large data
|
|
156
|
+
)
|
|
157
|
+
|
|
158
|
+
print('**Streaming Results**')
|
|
159
|
+
print(result)
|
|
160
|
+
print(f'\\nProcessed {result[\"count\"].sum():,} total rows')
|
|
161
|
+
"
|
|
162
|
+
```
|
|
163
|
+
|
|
164
|
+
## Template 5: Visualization with Matplotlib
|
|
165
|
+
|
|
166
|
+
**Use when:**
|
|
167
|
+
- User requests charts, graphs, or plots
|
|
168
|
+
- Exploratory data analysis (EDA)
|
|
169
|
+
- Time-series or distribution analysis
|
|
170
|
+
|
|
171
|
+
```python
|
|
172
|
+
uv run --with numpy --with duckdb --with polars --with pyarrow --with matplotlib python -c "
|
|
173
|
+
import duckdb
|
|
174
|
+
import matplotlib.pyplot as plt
|
|
175
|
+
|
|
176
|
+
# Query data
|
|
177
|
+
result = duckdb.sql('''
|
|
178
|
+
SELECT
|
|
179
|
+
date,
|
|
180
|
+
SUM(amount) as total
|
|
181
|
+
FROM 'data.csv'
|
|
182
|
+
GROUP BY date
|
|
183
|
+
ORDER BY date
|
|
184
|
+
''').pl()
|
|
185
|
+
|
|
186
|
+
# Create visualization
|
|
187
|
+
plt.figure(figsize=(10, 6))
|
|
188
|
+
plt.plot(result['date'], result['total'])
|
|
189
|
+
plt.xlabel('Date')
|
|
190
|
+
plt.ylabel('Total Amount')
|
|
191
|
+
plt.title('Daily Total Trends')
|
|
192
|
+
plt.xticks(rotation=45)
|
|
193
|
+
plt.tight_layout()
|
|
194
|
+
plt.savefig('output.png')
|
|
195
|
+
print('Chart saved to output.png')
|
|
196
|
+
"
|
|
197
|
+
```
|
|
@@ -0,0 +1,153 @@
|
|
|
1
|
+
# Integration Patterns: DuckDB ↔ Polars
|
|
2
|
+
|
|
3
|
+
## Zero-Copy Conversions (FASTEST)
|
|
4
|
+
|
|
5
|
+
### DuckDB → Polars (Recommended)
|
|
6
|
+
|
|
7
|
+
```python
|
|
8
|
+
import duckdb
|
|
9
|
+
import polars as pl
|
|
10
|
+
|
|
11
|
+
# Direct conversion with .pl() - zero-copy via Arrow
|
|
12
|
+
df_polars = duckdb.sql("""
|
|
13
|
+
SELECT * FROM 'data.parquet'
|
|
14
|
+
WHERE amount > 100
|
|
15
|
+
""").pl() # Returns Polars DataFrame directly
|
|
16
|
+
|
|
17
|
+
# Lazy version for large datasets
|
|
18
|
+
lazy_df = duckdb.sql("SELECT * FROM 'data.parquet'").pl(lazy=True)
|
|
19
|
+
result = lazy_df.filter(pl.col('status') == 'active').collect()
|
|
20
|
+
```
|
|
21
|
+
|
|
22
|
+
### Polars → DuckDB (Direct Reference)
|
|
23
|
+
|
|
24
|
+
```python
|
|
25
|
+
import duckdb
|
|
26
|
+
import polars as pl
|
|
27
|
+
|
|
28
|
+
# DuckDB can query Polars DataFrames directly by name
|
|
29
|
+
df = pl.read_parquet('data.parquet')
|
|
30
|
+
|
|
31
|
+
result = duckdb.sql("""
|
|
32
|
+
SELECT category, SUM(amount) as total
|
|
33
|
+
FROM df
|
|
34
|
+
GROUP BY category
|
|
35
|
+
ORDER BY total DESC
|
|
36
|
+
""").pl() # Query df directly, return as Polars
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
### Via Arrow (When Needed)
|
|
40
|
+
|
|
41
|
+
```python
|
|
42
|
+
# Polars → Arrow → DuckDB
|
|
43
|
+
df_polars = pl.read_csv('data.csv')
|
|
44
|
+
duckdb.register('my_table', df_polars.to_arrow())
|
|
45
|
+
|
|
46
|
+
# DuckDB → Arrow → Polars
|
|
47
|
+
arrow_table = duckdb.sql("SELECT * FROM data").arrow()
|
|
48
|
+
df_polars = pl.from_arrow(arrow_table)
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
## WRONG vs RIGHT Patterns
|
|
52
|
+
|
|
53
|
+
### ❌ NEVER - Using Pandas
|
|
54
|
+
|
|
55
|
+
```python
|
|
56
|
+
# FORBIDDEN - decisively slower than Polars/DuckDB on every operation!
|
|
57
|
+
import pandas as pd
|
|
58
|
+
df = pd.read_csv('data.csv')
|
|
59
|
+
result = df.groupby('category')['amount'].sum()
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
### ✅ CORRECT - DuckDB for simple aggregation query
|
|
63
|
+
|
|
64
|
+
```python
|
|
65
|
+
import duckdb
|
|
66
|
+
# Direct file query - no memory loading!
|
|
67
|
+
result = duckdb.sql("""
|
|
68
|
+
SELECT category, SUM(amount) as total
|
|
69
|
+
FROM 'data.csv'
|
|
70
|
+
GROUP BY category
|
|
71
|
+
""").pl() # Fast, memory-efficient
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
### ❌ NEVER - Loading file before DuckDB query
|
|
75
|
+
|
|
76
|
+
```python
|
|
77
|
+
# WRONG - Unnecessary memory usage
|
|
78
|
+
import polars as pl
|
|
79
|
+
import duckdb
|
|
80
|
+
df = pl.read_csv('data.csv') # Loads entire file
|
|
81
|
+
result = duckdb.sql("SELECT * FROM df WHERE amount > 100").pl()
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
### ✅ CORRECT - Let DuckDB query directly
|
|
85
|
+
|
|
86
|
+
```python
|
|
87
|
+
import duckdb
|
|
88
|
+
# DuckDB queries file directly - much faster!
|
|
89
|
+
result = duckdb.sql("""
|
|
90
|
+
SELECT * FROM 'data.csv'
|
|
91
|
+
WHERE amount > 100
|
|
92
|
+
""").pl()
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
### ❌ NEVER - Eager evaluation in Polars
|
|
96
|
+
|
|
97
|
+
```python
|
|
98
|
+
# WRONG - Loads everything immediately
|
|
99
|
+
import polars as pl
|
|
100
|
+
df = pl.read_csv('large_data.csv') # Eager load
|
|
101
|
+
filtered = df.filter(pl.col('value') > 100)
|
|
102
|
+
```
|
|
103
|
+
|
|
104
|
+
### ✅ CORRECT - Lazy evaluation
|
|
105
|
+
|
|
106
|
+
```python
|
|
107
|
+
import polars as pl
|
|
108
|
+
# Lazy - builds query plan, optimizes, executes once
|
|
109
|
+
df = pl.scan_csv('large_data.csv') # Lazy
|
|
110
|
+
result = (
|
|
111
|
+
df
|
|
112
|
+
.filter(pl.col('value') > 100)
|
|
113
|
+
.groupby('category')
|
|
114
|
+
.agg(pl.sum('value'))
|
|
115
|
+
.collect() # Execute optimized plan
|
|
116
|
+
)
|
|
117
|
+
```
|
|
118
|
+
|
|
119
|
+
### ❌ NEVER - Unnecessary Conversions
|
|
120
|
+
|
|
121
|
+
```python
|
|
122
|
+
# WASTEFUL (DuckDB → Pandas → Polars)
|
|
123
|
+
import duckdb, pandas as pd, polars as pl
|
|
124
|
+
df_pd = duckdb.sql("SELECT * FROM 'data.csv'").df() # requires pandas - the skill never ships it
|
|
125
|
+
df_pl = pl.from_pandas(df_pd)
|
|
126
|
+
```
|
|
127
|
+
|
|
128
|
+
### ✅ CORRECT - Direct conversion
|
|
129
|
+
|
|
130
|
+
```python
|
|
131
|
+
# DIRECT (DuckDB → Polars via Arrow)
|
|
132
|
+
import duckdb
|
|
133
|
+
df_pl = duckdb.sql("SELECT * FROM 'data.csv'").pl()
|
|
134
|
+
```
|
|
135
|
+
|
|
136
|
+
### ❌ NEVER - Wrong tool for heavy filtering
|
|
137
|
+
|
|
138
|
+
```python
|
|
139
|
+
# SLOW (DuckDB not optimal for filtering)
|
|
140
|
+
import duckdb
|
|
141
|
+
result = duckdb.sql("""
|
|
142
|
+
SELECT * FROM 'huge.csv'
|
|
143
|
+
WHERE complex_filter = true
|
|
144
|
+
""").pl()
|
|
145
|
+
```
|
|
146
|
+
|
|
147
|
+
### ✅ CORRECT - Use Polars for filtering
|
|
148
|
+
|
|
149
|
+
```python
|
|
150
|
+
# FAST (Polars 128x faster for filtering)
|
|
151
|
+
import polars as pl
|
|
152
|
+
result = pl.scan_csv('huge.csv').filter(pl.col('complex_filter')).collect()
|
|
153
|
+
```
|
|
@@ -0,0 +1,37 @@
|
|
|
1
|
+
# Performance Benchmarks
|
|
2
|
+
|
|
3
|
+
Routing heuristics distilled from public 2024-2025 results in the [db-benchmark suite](https://github.com/h2oai/db-benchmark) (rendered at [h2oai.github.io/db-benchmark](https://h2oai.github.io/db-benchmark/)) and from the vendors' own documentation ([duckdb.org](https://duckdb.org/), [pola.rs](https://pola.rs/)).
|
|
4
|
+
|
|
5
|
+
**Read these as routing guidance, not guarantees.** Multipliers vary with dataset size, cardinality, data types, and hardware. When the choice materially matters, measure on the actual data.
|
|
6
|
+
|
|
7
|
+
## Operation Performance Comparison
|
|
8
|
+
|
|
9
|
+
| Operation Type | Winner | Indicative advantage | When to Use |
|
|
10
|
+
|---------------------|-------------|----------------------|------------------------------------------------|
|
|
11
|
+
| **Filtering** | **Polars** | often the fastest by a wide margin (SIMD, predicate pushdown) | Single-table filters |
|
|
12
|
+
| **Sorting** | **Polars** | typically the fastest | ORDER BY operations, ranking |
|
|
13
|
+
| **Joins** | **DuckDB** | typically faster; richer join types | Multi-table joins, especially complex joins |
|
|
14
|
+
| **Aggregations** | **DuckDB** | typically faster on large datasets | GROUP BY, complex aggregations |
|
|
15
|
+
| **Window Funcs** | **Polars** | typically faster | RANK, LAG/LEAD, running totals |
|
|
16
|
+
| **Transformations** | **Polars** | typically faster | Pivot, melt, string operations |
|
|
17
|
+
| **Direct Query** | **DuckDB** | avoids loading to memory entirely | Ad-hoc exploration without loading to memory |
|
|
18
|
+
| **Streaming** | **Polars** | handles datasets larger than RAM | Datasets larger than RAM |
|
|
19
|
+
|
|
20
|
+
Both tools are decisively faster than pandas on 1M+ row workloads, which is why pandas is banned outright in this skill.
|
|
21
|
+
|
|
22
|
+
## Operation Detection Keywords
|
|
23
|
+
|
|
24
|
+
Automatically detect operation types from user requests:
|
|
25
|
+
|
|
26
|
+
- **Filter**: "where", "filter", "condition", "select rows", "find"
|
|
27
|
+
- **Sort**: "sort", "order", "top", "bottom", "rank"
|
|
28
|
+
- **Join**: "join", "merge", "combine tables"
|
|
29
|
+
- **Aggregate**: "group", "sum", "avg", "count", "mean", "total", "aggregate"
|
|
30
|
+
- **Transform**: "pivot", "melt", "reshape", "string operations", "clean", "transform"
|
|
31
|
+
- **Window**: "running", "cumulative", "lag", "lead", "rank", "partition"
|
|
32
|
+
|
|
33
|
+
## Benchmarking Notes
|
|
34
|
+
|
|
35
|
+
- Public suites exercise typical analytical workloads (1M-100M rows).
|
|
36
|
+
- Performance advantages are approximate and vary by dataset characteristics.
|
|
37
|
+
- Real-world performance depends on data types, cardinality, and hardware.
|
|
@@ -0,0 +1,78 @@
|
|
|
1
|
+
# uv Setup — Per-Platform
|
|
2
|
+
|
|
3
|
+
This skill runs every data operation through `uv run --with ...`. If `uv --version` fails, set uv up with the automated scripts or the manual commands below, then verify.
|
|
4
|
+
|
|
5
|
+
## Automated (recommended)
|
|
6
|
+
|
|
7
|
+
| Platform | Command |
|
|
8
|
+
|---|---|
|
|
9
|
+
| macOS / Linux / WSL / Git Bash | `bash scripts/setup-uv.sh` |
|
|
10
|
+
| Windows (PowerShell) | `powershell -ExecutionPolicy Bypass -File scripts/setup-uv.ps1` |
|
|
11
|
+
|
|
12
|
+
Both scripts: detect OS + architecture → install uv to the latest release when missing → upgrade it when present (`uv self update`) → make it resolvable for the current shell → verify with `uv --version`. They are idempotent — safe to re-run any time.
|
|
13
|
+
|
|
14
|
+
## Manual install per platform
|
|
15
|
+
|
|
16
|
+
### macOS
|
|
17
|
+
|
|
18
|
+
```bash
|
|
19
|
+
curl -LsSf https://astral.sh/uv/install.sh | sh # official installer → ~/.local/bin/uv
|
|
20
|
+
# or, with Homebrew:
|
|
21
|
+
brew install uv
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
### Linux (x86_64 / aarch64)
|
|
25
|
+
|
|
26
|
+
```bash
|
|
27
|
+
curl -LsSf https://astral.sh/uv/install.sh | sh # official installer → ~/.local/bin/uv
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
The installer detects glibc vs musl and downloads the right static binary. On minimal containers, ensure `curl` (or `wget`) exists; `wget -qO- https://astral.sh/uv/install.sh | sh` is the fallback.
|
|
31
|
+
|
|
32
|
+
### Windows (native)
|
|
33
|
+
|
|
34
|
+
```powershell
|
|
35
|
+
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"
|
|
36
|
+
# or, with winget:
|
|
37
|
+
winget install --id=astral-sh.uv -e
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
### Windows Subsystem for Linux / Git Bash
|
|
41
|
+
|
|
42
|
+
Use the Linux/macOS installer inside the Unix shell, not the PowerShell installer:
|
|
43
|
+
|
|
44
|
+
```bash
|
|
45
|
+
curl -LsSf https://astral.sh/uv/install.sh | sh
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
### CI
|
|
49
|
+
|
|
50
|
+
```yaml
|
|
51
|
+
# GitHub Actions
|
|
52
|
+
- uses: astral-sh/setup-uv@v5
|
|
53
|
+
# or plain shell anywhere:
|
|
54
|
+
- run: curl -LsSf https://astral.sh/uv/install.sh | sh
|
|
55
|
+
```
|
|
56
|
+
|
|
57
|
+
## PATH notes
|
|
58
|
+
|
|
59
|
+
- The official installers put the binary in `~/.local/bin` (Unix) or `%USERPROFILE%\.local\bin` (Windows). New shells get it automatically on most setups; an already-open shell needs `export PATH="$HOME/.local/bin:$PATH"` (Unix) or `$env:Path = "$env:USERPROFILE\.local\bin;$env:Path"` (PowerShell) once.
|
|
60
|
+
- Homebrew and winget install into their own prefixes that are already on PATH.
|
|
61
|
+
|
|
62
|
+
## Upgrade to latest
|
|
63
|
+
|
|
64
|
+
```bash
|
|
65
|
+
uv self update
|
|
66
|
+
```
|
|
67
|
+
|
|
68
|
+
`uv self update` only works for official-installer binaries; Homebrew/winget installs upgrade through their package manager (`brew upgrade uv` / `winget upgrade astral-sh.uv`). The setup scripts handle this automatically.
|
|
69
|
+
|
|
70
|
+
## Verify
|
|
71
|
+
|
|
72
|
+
```bash
|
|
73
|
+
uv --version
|
|
74
|
+
```
|
|
75
|
+
|
|
76
|
+
## Offline / air-gapped
|
|
77
|
+
|
|
78
|
+
Download the matching archive from the uv GitHub releases page, extract it, and put the `uv` binary anywhere on PATH. `uv run --with <pkg>` still needs network for first-time package resolution unless a mirror is configured via `UV_INDEX_URL`.
|