@ngockhoale/ukit 1.6.7 → 2.0.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (59) hide show
  1. package/CHANGELOG.md +26 -0
  2. package/manifests/platform.full.yaml +24 -0
  3. package/package.json +2 -1
  4. package/scripts/skill/audit-skill.mjs +39 -0
  5. package/src/cli/commands/doctor.js +22 -2
  6. package/src/cli/commands/memory.js +76 -1
  7. package/src/core/memory/store.js +125 -1
  8. package/src/core/skillProfile.js +45 -0
  9. package/src/skill/auditSkill.js +99 -0
  10. package/templates/.claude/agents/code-reviewer.md +51 -7
  11. package/templates/.claude/agents/handoff-planner.md +18 -2
  12. package/templates/.claude/skills/canvas-design/SKILL.md +2 -20
  13. package/templates/.claude/skills/canvas-design/philosophy-examples.md +23 -0
  14. package/templates/.claude/skills/debugging-toolkit/SKILL.md +2 -30
  15. package/templates/.claude/skills/debugging-toolkit/reference-tables.md +33 -0
  16. package/templates/.claude/skills/docs-manager/SKILL.md +7 -249
  17. package/templates/.claude/skills/docs-manager/conventions-and-examples.md +221 -0
  18. package/templates/.claude/skills/docx/SKILL.md +3 -34
  19. package/templates/.claude/skills/docx/redlining-reference.md +34 -0
  20. package/templates/.claude/skills/duraone/SKILL.md +12 -16
  21. package/templates/.claude/skills/executing-plans/SKILL.md +31 -19
  22. package/templates/.claude/skills/file-organizer/SKILL.md +2 -170
  23. package/templates/.claude/skills/file-organizer/examples-and-practices.md +173 -0
  24. package/templates/.claude/skills/pdf/SKILL.md +1 -62
  25. package/templates/.claude/skills/pdf/reference.md +65 -0
  26. package/templates/.claude/skills/pdf-processing-pro/SKILL.md +2 -73
  27. package/templates/.claude/skills/pdf-processing-pro/workflows-and-troubleshooting.md +80 -0
  28. package/templates/.claude/skills/pptx/SKILL.md +14 -286
  29. package/templates/.claude/skills/pptx/design-references.md +81 -0
  30. package/templates/.claude/skills/pptx/template-replacement-reference.md +150 -0
  31. package/templates/.claude/skills/pptx/utilities.md +62 -0
  32. package/templates/.claude/skills/project-learning/SKILL.md +32 -0
  33. package/templates/.claude/skills/root-cause-tracing/SKILL.md +2 -35
  34. package/templates/.claude/skills/root-cause-tracing/diagrams.md +44 -0
  35. package/templates/.claude/skills/sharing-skills/SKILL.md +1 -41
  36. package/templates/.claude/skills/sharing-skills/complete-example.md +41 -0
  37. package/templates/.claude/skills/skill-quality/SKILL.md +37 -0
  38. package/templates/.claude/skills/skill-quality/pressure-scenario-template.md +20 -0
  39. package/templates/.claude/skills/skill-quality/rationalization-table-template.md +15 -0
  40. package/templates/.claude/skills/skill-quality/trigger-accuracy-template.md +32 -0
  41. package/templates/.claude/skills/sql-optimization-patterns/SKILL.md +13 -440
  42. package/templates/.claude/skills/sql-optimization-patterns/references/advanced-techniques.md +128 -0
  43. package/templates/.claude/skills/sql-optimization-patterns/references/core-concepts.md +112 -0
  44. package/templates/.claude/skills/sql-optimization-patterns/references/query-patterns.md +204 -0
  45. package/templates/.claude/skills/subagent-driven-development/SKILL.md +4 -51
  46. package/templates/.claude/skills/subagent-driven-development/example-workflow.md +40 -0
  47. package/templates/.claude/skills/systematic-debugging/SKILL.md +2 -28
  48. package/templates/.claude/skills/systematic-debugging/reference-tables.md +33 -0
  49. package/templates/.claude/skills/test-driven-development/SKILL.md +2 -51
  50. package/templates/.claude/skills/test-driven-development/reference-tables.md +56 -0
  51. package/templates/.claude/skills/testing-anti-patterns/SKILL.md +1 -10
  52. package/templates/.claude/skills/testing-anti-patterns/reference-tables.md +14 -0
  53. package/templates/.claude/skills/verification-before-completion/SKILL.md +1 -31
  54. package/templates/.claude/skills/verification-before-completion/key-patterns.md +33 -0
  55. package/templates/.claude/ukit/index/extract-image.mjs +10 -3
  56. package/templates/CLAUDE.md +4 -0
  57. package/src/core/memory/index.js +0 -2
  58. package/src/core/router/index.js +0 -2
  59. package/src/core/validation/index.js +0 -2
@@ -20,397 +20,25 @@ Transform slow database queries into lightning-fast operations through systemati
20
20
 
21
21
  ## Core Concepts
22
22
 
23
- ### 1. Query Execution Plans (EXPLAIN)
23
+ Full syntax examples for all three live in [`references/core-concepts.md`](references/core-concepts.md):
24
24
 
25
- Understanding EXPLAIN output is fundamental to optimization.
26
-
27
- **PostgreSQL EXPLAIN:**
28
- ```sql
29
- -- Basic explain
30
- EXPLAIN SELECT * FROM users WHERE email = 'user@example.com';
31
-
32
- -- With actual execution stats
33
- EXPLAIN ANALYZE
34
- SELECT * FROM users WHERE email = 'user@example.com';
35
-
36
- -- Verbose output with more details
37
- EXPLAIN (ANALYZE, BUFFERS, VERBOSE)
38
- SELECT u.*, o.order_total
39
- FROM users u
40
- JOIN orders o ON u.id = o.user_id
41
- WHERE u.created_at > NOW() - INTERVAL '30 days';
42
- ```
43
-
44
- **Key Metrics to Watch:**
45
- - **Seq Scan**: Full table scan (usually slow for large tables)
46
- - **Index Scan**: Using index (good)
47
- - **Index Only Scan**: Using index without touching table (best)
48
- - **Nested Loop**: Join method (okay for small datasets)
49
- - **Hash Join**: Join method (good for larger datasets)
50
- - **Merge Join**: Join method (good for sorted data)
51
- - **Cost**: Estimated query cost (lower is better)
52
- - **Rows**: Estimated rows returned
53
- - **Actual Time**: Real execution time
54
-
55
- ### 2. Index Strategies
56
-
57
- Indexes are the most powerful optimization tool.
58
-
59
- **Index Types:**
60
- - **B-Tree**: Default, good for equality and range queries
61
- - **Hash**: Only for equality (=) comparisons
62
- - **GIN**: Full-text search, array queries, JSONB
63
- - **GiST**: Geometric data, full-text search
64
- - **BRIN**: Block Range INdex for very large tables with correlation
65
-
66
- ```sql
67
- -- Standard B-Tree index
68
- CREATE INDEX idx_users_email ON users(email);
69
-
70
- -- Composite index (order matters!)
71
- CREATE INDEX idx_orders_user_status ON orders(user_id, status);
72
-
73
- -- Partial index (index subset of rows)
74
- CREATE INDEX idx_active_users ON users(email)
75
- WHERE status = 'active';
76
-
77
- -- Expression index
78
- CREATE INDEX idx_users_lower_email ON users(LOWER(email));
79
-
80
- -- Covering index (include additional columns)
81
- CREATE INDEX idx_users_email_covering ON users(email)
82
- INCLUDE (name, created_at);
83
-
84
- -- Full-text search index
85
- CREATE INDEX idx_posts_search ON posts
86
- USING GIN(to_tsvector('english', title || ' ' || body));
87
-
88
- -- JSONB index
89
- CREATE INDEX idx_metadata ON events USING GIN(metadata);
90
- ```
91
-
92
- ### 3. Query Optimization Patterns
93
-
94
- **Avoid SELECT \*:**
95
- ```sql
96
- -- Bad: Fetches unnecessary columns
97
- SELECT * FROM users WHERE id = 123;
98
-
99
- -- Good: Fetch only what you need
100
- SELECT id, email, name FROM users WHERE id = 123;
101
- ```
102
-
103
- **Use WHERE Clause Efficiently:**
104
- ```sql
105
- -- Bad: Function prevents index usage
106
- SELECT * FROM users WHERE LOWER(email) = 'user@example.com';
107
-
108
- -- Good: Create functional index or use exact match
109
- CREATE INDEX idx_users_email_lower ON users(LOWER(email));
110
- -- Then:
111
- SELECT * FROM users WHERE LOWER(email) = 'user@example.com';
112
-
113
- -- Or store normalized data
114
- SELECT * FROM users WHERE email = 'user@example.com';
115
- ```
116
-
117
- **Optimize JOINs:**
118
- ```sql
119
- -- Bad: Cartesian product then filter
120
- SELECT u.name, o.total
121
- FROM users u, orders o
122
- WHERE u.id = o.user_id AND u.created_at > '2024-01-01';
123
-
124
- -- Good: Filter before join
125
- SELECT u.name, o.total
126
- FROM users u
127
- JOIN orders o ON u.id = o.user_id
128
- WHERE u.created_at > '2024-01-01';
129
-
130
- -- Better: Filter both tables
131
- SELECT u.name, o.total
132
- FROM (SELECT * FROM users WHERE created_at > '2024-01-01') u
133
- JOIN orders o ON u.id = o.user_id;
134
- ```
25
+ 1. **Query Execution Plans (EXPLAIN)** — read Seq Scan/Index Scan/Index Only Scan, join method, cost, and actual-time metrics before optimizing anything.
26
+ 2. **Index Strategies** — pick the right index type (B-Tree, Hash, GIN, GiST, BRIN) and form (composite, partial, expression, covering) for the query shape.
27
+ 3. **Basic Query Rewrites** — avoid `SELECT *`, avoid functions on indexed columns in `WHERE` (or add a functional index), filter before JOINing.
135
28
 
136
29
  ## Optimization Patterns
137
30
 
138
- ### Pattern 1: Eliminate N+1 Queries
139
-
140
- **Problem: N+1 Query Anti-Pattern**
141
- ```python
142
- # Bad: Executes N+1 queries
143
- users = db.query("SELECT * FROM users LIMIT 10")
144
- for user in users:
145
- orders = db.query("SELECT * FROM orders WHERE user_id = ?", user.id)
146
- # Process orders
147
- ```
148
-
149
- **Solution: Use JOINs or Batch Loading**
150
- ```sql
151
- -- Solution 1: JOIN
152
- SELECT
153
- u.id, u.name,
154
- o.id as order_id, o.total
155
- FROM users u
156
- LEFT JOIN orders o ON u.id = o.user_id
157
- WHERE u.id IN (1, 2, 3, 4, 5);
158
-
159
- -- Solution 2: Batch query
160
- SELECT * FROM orders
161
- WHERE user_id IN (1, 2, 3, 4, 5);
162
- ```
163
-
164
- ```python
165
- # Good: Single query with JOIN or batch load
166
- # Using JOIN
167
- results = db.query("""
168
- SELECT u.id, u.name, o.id as order_id, o.total
169
- FROM users u
170
- LEFT JOIN orders o ON u.id = o.user_id
171
- WHERE u.id IN (1, 2, 3, 4, 5)
172
- """)
173
-
174
- # Or batch load
175
- users = db.query("SELECT * FROM users LIMIT 10")
176
- user_ids = [u.id for u in users]
177
- orders = db.query(
178
- "SELECT * FROM orders WHERE user_id IN (?)",
179
- user_ids
180
- )
181
- # Group orders by user_id
182
- orders_by_user = {}
183
- for order in orders:
184
- orders_by_user.setdefault(order.user_id, []).append(order)
185
- ```
186
-
187
- ### Pattern 2: Optimize Pagination
188
-
189
- **Bad: OFFSET on Large Tables**
190
- ```sql
191
- -- Slow for large offsets
192
- SELECT * FROM users
193
- ORDER BY created_at DESC
194
- LIMIT 20 OFFSET 100000; -- Very slow!
195
- ```
196
-
197
- **Good: Cursor-Based Pagination**
198
- ```sql
199
- -- Much faster: Use cursor (last seen ID)
200
- SELECT * FROM users
201
- WHERE created_at < '2024-01-15 10:30:00' -- Last cursor
202
- ORDER BY created_at DESC
203
- LIMIT 20;
204
-
205
- -- With composite sorting
206
- SELECT * FROM users
207
- WHERE (created_at, id) < ('2024-01-15 10:30:00', 12345)
208
- ORDER BY created_at DESC, id DESC
209
- LIMIT 20;
210
-
211
- -- Requires index
212
- CREATE INDEX idx_users_cursor ON users(created_at DESC, id DESC);
213
- ```
214
-
215
- ### Pattern 3: Aggregate Efficiently
216
-
217
- **Optimize COUNT Queries:**
218
- ```sql
219
- -- Bad: Counts all rows
220
- SELECT COUNT(*) FROM orders; -- Slow on large tables
221
-
222
- -- Good: Use estimates for approximate counts
223
- SELECT reltuples::bigint AS estimate
224
- FROM pg_class
225
- WHERE relname = 'orders';
226
-
227
- -- Good: Filter before counting
228
- SELECT COUNT(*) FROM orders
229
- WHERE created_at > NOW() - INTERVAL '7 days';
230
-
231
- -- Better: Use index-only scan
232
- CREATE INDEX idx_orders_created ON orders(created_at);
233
- SELECT COUNT(*) FROM orders
234
- WHERE created_at > NOW() - INTERVAL '7 days';
235
- ```
31
+ Five recurring patterns, each with full before/after SQL in [`references/query-patterns.md`](references/query-patterns.md):
236
32
 
237
- **Optimize GROUP BY:**
238
- ```sql
239
- -- Bad: Group by then filter
240
- SELECT user_id, COUNT(*) as order_count
241
- FROM orders
242
- GROUP BY user_id
243
- HAVING COUNT(*) > 10;
244
-
245
- -- Better: Filter first, then group (if possible)
246
- SELECT user_id, COUNT(*) as order_count
247
- FROM orders
248
- WHERE status = 'completed'
249
- GROUP BY user_id
250
- HAVING COUNT(*) > 10;
251
-
252
- -- Best: Use covering index
253
- CREATE INDEX idx_orders_user_status ON orders(user_id, status);
254
- ```
255
-
256
- ### Pattern 4: Subquery Optimization
257
-
258
- **Transform Correlated Subqueries:**
259
- ```sql
260
- -- Bad: Correlated subquery (runs for each row)
261
- SELECT u.name, u.email,
262
- (SELECT COUNT(*) FROM orders o WHERE o.user_id = u.id) as order_count
263
- FROM users u;
264
-
265
- -- Good: JOIN with aggregation
266
- SELECT u.name, u.email, COUNT(o.id) as order_count
267
- FROM users u
268
- LEFT JOIN orders o ON o.user_id = u.id
269
- GROUP BY u.id, u.name, u.email;
270
-
271
- -- Better: Use window functions
272
- SELECT DISTINCT ON (u.id)
273
- u.name, u.email,
274
- COUNT(o.id) OVER (PARTITION BY u.id) as order_count
275
- FROM users u
276
- LEFT JOIN orders o ON o.user_id = u.id;
277
- ```
278
-
279
- **Use CTEs for Clarity:**
280
- ```sql
281
- -- Using Common Table Expressions
282
- WITH recent_users AS (
283
- SELECT id, name, email
284
- FROM users
285
- WHERE created_at > NOW() - INTERVAL '30 days'
286
- ),
287
- user_order_counts AS (
288
- SELECT user_id, COUNT(*) as order_count
289
- FROM orders
290
- WHERE created_at > NOW() - INTERVAL '30 days'
291
- GROUP BY user_id
292
- )
293
- SELECT ru.name, ru.email, COALESCE(uoc.order_count, 0) as orders
294
- FROM recent_users ru
295
- LEFT JOIN user_order_counts uoc ON ru.id = uoc.user_id;
296
- ```
297
-
298
- ### Pattern 5: Batch Operations
299
-
300
- **Batch INSERT:**
301
- ```sql
302
- -- Bad: Multiple individual inserts
303
- INSERT INTO users (name, email) VALUES ('Alice', 'alice@example.com');
304
- INSERT INTO users (name, email) VALUES ('Bob', 'bob@example.com');
305
- INSERT INTO users (name, email) VALUES ('Carol', 'carol@example.com');
306
-
307
- -- Good: Batch insert
308
- INSERT INTO users (name, email) VALUES
309
- ('Alice', 'alice@example.com'),
310
- ('Bob', 'bob@example.com'),
311
- ('Carol', 'carol@example.com');
312
-
313
- -- Better: Use COPY for bulk inserts (PostgreSQL)
314
- COPY users (name, email) FROM '/tmp/users.csv' CSV HEADER;
315
- ```
316
-
317
- **Batch UPDATE:**
318
- ```sql
319
- -- Bad: Update in loop
320
- UPDATE users SET status = 'active' WHERE id = 1;
321
- UPDATE users SET status = 'active' WHERE id = 2;
322
- -- ... repeat for many IDs
323
-
324
- -- Good: Single UPDATE with IN clause
325
- UPDATE users
326
- SET status = 'active'
327
- WHERE id IN (1, 2, 3, 4, 5, ...);
328
-
329
- -- Better: Use temporary table for large batches
330
- CREATE TEMP TABLE temp_user_updates (id INT, new_status VARCHAR);
331
- INSERT INTO temp_user_updates VALUES (1, 'active'), (2, 'active'), ...;
332
-
333
- UPDATE users u
334
- SET status = t.new_status
335
- FROM temp_user_updates t
336
- WHERE u.id = t.id;
337
- ```
33
+ 1. **Eliminate N+1 Queries** — replace a per-row query loop with a single JOIN or batch `IN (...)` load.
34
+ 2. **Optimize Pagination** — replace `OFFSET` on large tables with cursor-based pagination (`WHERE created_at < :cursor ... LIMIT n`), backed by a matching index.
35
+ 3. **Aggregate Efficiently** — filter before `COUNT`/`GROUP BY`, use estimates (`pg_class.reltuples`) for approximate counts, and back frequent aggregations with a covering index.
36
+ 4. **Subquery Optimization** — replace correlated subqueries with JOIN + aggregation or window functions; use CTEs for clarity.
37
+ 5. **Batch Operations** — batch `INSERT`/`UPDATE` (or `COPY`/temp-table joins for large batches) instead of looping single-row statements.
338
38
 
339
39
  ## Advanced Techniques
340
40
 
341
- ### Materialized Views
342
-
343
- Pre-compute expensive queries.
344
-
345
- ```sql
346
- -- Create materialized view
347
- CREATE MATERIALIZED VIEW user_order_summary AS
348
- SELECT
349
- u.id,
350
- u.name,
351
- COUNT(o.id) as total_orders,
352
- SUM(o.total) as total_spent,
353
- MAX(o.created_at) as last_order_date
354
- FROM users u
355
- LEFT JOIN orders o ON u.id = o.user_id
356
- GROUP BY u.id, u.name;
357
-
358
- -- Add index to materialized view
359
- CREATE INDEX idx_user_summary_spent ON user_order_summary(total_spent DESC);
360
-
361
- -- Refresh materialized view
362
- REFRESH MATERIALIZED VIEW user_order_summary;
363
-
364
- -- Concurrent refresh (PostgreSQL)
365
- REFRESH MATERIALIZED VIEW CONCURRENTLY user_order_summary;
366
-
367
- -- Query materialized view (very fast)
368
- SELECT * FROM user_order_summary
369
- WHERE total_spent > 1000
370
- ORDER BY total_spent DESC;
371
- ```
372
-
373
- ### Partitioning
374
-
375
- Split large tables for better performance.
376
-
377
- ```sql
378
- -- Range partitioning by date (PostgreSQL)
379
- CREATE TABLE orders (
380
- id SERIAL,
381
- user_id INT,
382
- total DECIMAL,
383
- created_at TIMESTAMP
384
- ) PARTITION BY RANGE (created_at);
385
-
386
- -- Create partitions
387
- CREATE TABLE orders_2024_q1 PARTITION OF orders
388
- FOR VALUES FROM ('2024-01-01') TO ('2024-04-01');
389
-
390
- CREATE TABLE orders_2024_q2 PARTITION OF orders
391
- FOR VALUES FROM ('2024-04-01') TO ('2024-07-01');
392
-
393
- -- Queries automatically use appropriate partition
394
- SELECT * FROM orders
395
- WHERE created_at BETWEEN '2024-02-01' AND '2024-02-28';
396
- -- Only scans orders_2024_q1 partition
397
- ```
398
-
399
- ### Query Hints and Optimization
400
-
401
- ```sql
402
- -- Force index usage (MySQL)
403
- SELECT * FROM users
404
- USE INDEX (idx_users_email)
405
- WHERE email = 'user@example.com';
406
-
407
- -- Parallel query (PostgreSQL)
408
- SET max_parallel_workers_per_gather = 4;
409
- SELECT * FROM large_table WHERE condition;
410
-
411
- -- Join hints (PostgreSQL)
412
- SET enable_nestloop = OFF; -- Force hash or merge join
413
- ```
41
+ Materialized views, table partitioning, and query hints (MySQL `USE INDEX`, PostgreSQL parallel workers/join hints) are covered with full examples in [`references/advanced-techniques.md`](references/advanced-techniques.md#materialized-views).
414
42
 
415
43
  ## Best Practices
416
44
 
@@ -421,21 +49,7 @@ SET enable_nestloop = OFF; -- Force hash or merge join
421
49
  5. **Normalize Thoughtfully**: Balance normalization vs performance
422
50
  6. **Cache Frequently Accessed Data**: Use application-level caching
423
51
  7. **Connection Pooling**: Reuse database connections
424
- 8. **Regular Maintenance**: VACUUM, ANALYZE, rebuild indexes
425
-
426
- ```sql
427
- -- Update statistics
428
- ANALYZE users;
429
- ANALYZE VERBOSE orders;
430
-
431
- -- Vacuum (PostgreSQL)
432
- VACUUM ANALYZE users;
433
- VACUUM FULL users; -- Reclaim space (locks table)
434
-
435
- -- Reindex
436
- REINDEX INDEX idx_users_email;
437
- REINDEX TABLE users;
438
- ```
52
+ 8. **Regular Maintenance**: VACUUM, ANALYZE, rebuild indexes — commands in [`references/advanced-techniques.md`](references/advanced-techniques.md#maintenance)
439
53
 
440
54
  ## Common Pitfalls
441
55
 
@@ -449,45 +63,4 @@ REINDEX TABLE users;
449
63
 
450
64
  ## Monitoring Queries
451
65
 
452
- ```sql
453
- -- Find slow queries (PostgreSQL)
454
- SELECT query, calls, total_time, mean_time
455
- FROM pg_stat_statements
456
- ORDER BY mean_time DESC
457
- LIMIT 10;
458
-
459
- -- Find missing indexes (PostgreSQL)
460
- SELECT
461
- schemaname,
462
- tablename,
463
- seq_scan,
464
- seq_tup_read,
465
- idx_scan,
466
- seq_tup_read / seq_scan AS avg_seq_tup_read
467
- FROM pg_stat_user_tables
468
- WHERE seq_scan > 0
469
- ORDER BY seq_tup_read DESC
470
- LIMIT 10;
471
-
472
- -- Find unused indexes (PostgreSQL)
473
- SELECT
474
- schemaname,
475
- tablename,
476
- indexname,
477
- idx_scan,
478
- idx_tup_read,
479
- idx_tup_fetch
480
- FROM pg_stat_user_indexes
481
- WHERE idx_scan = 0
482
- ORDER BY pg_relation_size(indexrelid) DESC;
483
- ```
484
-
485
- ## Resources
486
-
487
- - **references/postgres-optimization-guide.md**: PostgreSQL-specific optimization
488
- - **references/mysql-optimization-guide.md**: MySQL/MariaDB optimization
489
- - **references/query-plan-analysis.md**: Deep dive into EXPLAIN plans
490
- - **assets/index-strategy-checklist.md**: When and how to create indexes
491
- - **assets/query-optimization-checklist.md**: Step-by-step optimization guide
492
- - **scripts/analyze-slow-queries.sql**: Identify slow queries in your database
493
- - **scripts/index-recommendations.sql**: Generate index recommendations
66
+ Ready-to-run PostgreSQL queries for slow queries, missing indexes, and unused indexes: [`references/advanced-techniques.md`](references/advanced-techniques.md#monitoring-queries).
@@ -0,0 +1,128 @@
1
+ # Advanced SQL Optimization Techniques
2
+
3
+ Deeper material referenced from the main skill: materialized views, partitioning, query hints, maintenance, and monitoring queries.
4
+
5
+ ## Materialized Views
6
+
7
+ Pre-compute expensive queries.
8
+
9
+ ```sql
10
+ -- Create materialized view
11
+ CREATE MATERIALIZED VIEW user_order_summary AS
12
+ SELECT
13
+ u.id,
14
+ u.name,
15
+ COUNT(o.id) as total_orders,
16
+ SUM(o.total) as total_spent,
17
+ MAX(o.created_at) as last_order_date
18
+ FROM users u
19
+ LEFT JOIN orders o ON u.id = o.user_id
20
+ GROUP BY u.id, u.name;
21
+
22
+ -- Add index to materialized view
23
+ CREATE INDEX idx_user_summary_spent ON user_order_summary(total_spent DESC);
24
+
25
+ -- Refresh materialized view
26
+ REFRESH MATERIALIZED VIEW user_order_summary;
27
+
28
+ -- Concurrent refresh (PostgreSQL)
29
+ REFRESH MATERIALIZED VIEW CONCURRENTLY user_order_summary;
30
+
31
+ -- Query materialized view (very fast)
32
+ SELECT * FROM user_order_summary
33
+ WHERE total_spent > 1000
34
+ ORDER BY total_spent DESC;
35
+ ```
36
+
37
+ ## Partitioning
38
+
39
+ Split large tables for better performance.
40
+
41
+ ```sql
42
+ -- Range partitioning by date (PostgreSQL)
43
+ CREATE TABLE orders (
44
+ id SERIAL,
45
+ user_id INT,
46
+ total DECIMAL,
47
+ created_at TIMESTAMP
48
+ ) PARTITION BY RANGE (created_at);
49
+
50
+ -- Create partitions
51
+ CREATE TABLE orders_2024_q1 PARTITION OF orders
52
+ FOR VALUES FROM ('2024-01-01') TO ('2024-04-01');
53
+
54
+ CREATE TABLE orders_2024_q2 PARTITION OF orders
55
+ FOR VALUES FROM ('2024-04-01') TO ('2024-07-01');
56
+
57
+ -- Queries automatically use appropriate partition
58
+ SELECT * FROM orders
59
+ WHERE created_at BETWEEN '2024-02-01' AND '2024-02-28';
60
+ -- Only scans orders_2024_q1 partition
61
+ ```
62
+
63
+ ## Query Hints and Optimization
64
+
65
+ ```sql
66
+ -- Force index usage (MySQL)
67
+ SELECT * FROM users
68
+ USE INDEX (idx_users_email)
69
+ WHERE email = 'user@example.com';
70
+
71
+ -- Parallel query (PostgreSQL)
72
+ SET max_parallel_workers_per_gather = 4;
73
+ SELECT * FROM large_table WHERE condition;
74
+
75
+ -- Join hints (PostgreSQL)
76
+ SET enable_nestloop = OFF; -- Force hash or merge join
77
+ ```
78
+
79
+ ## Maintenance
80
+
81
+ ```sql
82
+ -- Update statistics
83
+ ANALYZE users;
84
+ ANALYZE VERBOSE orders;
85
+
86
+ -- Vacuum (PostgreSQL)
87
+ VACUUM ANALYZE users;
88
+ VACUUM FULL users; -- Reclaim space (locks table)
89
+
90
+ -- Reindex
91
+ REINDEX INDEX idx_users_email;
92
+ REINDEX TABLE users;
93
+ ```
94
+
95
+ ## Monitoring Queries
96
+
97
+ ```sql
98
+ -- Find slow queries (PostgreSQL)
99
+ SELECT query, calls, total_time, mean_time
100
+ FROM pg_stat_statements
101
+ ORDER BY mean_time DESC
102
+ LIMIT 10;
103
+
104
+ -- Find missing indexes (PostgreSQL)
105
+ SELECT
106
+ schemaname,
107
+ tablename,
108
+ seq_scan,
109
+ seq_tup_read,
110
+ idx_scan,
111
+ seq_tup_read / seq_scan AS avg_seq_tup_read
112
+ FROM pg_stat_user_tables
113
+ WHERE seq_scan > 0
114
+ ORDER BY seq_tup_read DESC
115
+ LIMIT 10;
116
+
117
+ -- Find unused indexes (PostgreSQL)
118
+ SELECT
119
+ schemaname,
120
+ tablename,
121
+ indexname,
122
+ idx_scan,
123
+ idx_tup_read,
124
+ idx_tup_fetch
125
+ FROM pg_stat_user_indexes
126
+ WHERE idx_scan = 0
127
+ ORDER BY pg_relation_size(indexrelid) DESC;
128
+ ```
@@ -0,0 +1,112 @@
1
+ # SQL Optimization Core Concepts (Detailed)
2
+
3
+ Full syntax examples for EXPLAIN, index types, and basic query rewrites summarized in the main skill.
4
+
5
+ ## Query Execution Plans (EXPLAIN)
6
+
7
+ **PostgreSQL EXPLAIN:**
8
+ ```sql
9
+ -- Basic explain
10
+ EXPLAIN SELECT * FROM users WHERE email = 'user@example.com';
11
+
12
+ -- With actual execution stats
13
+ EXPLAIN ANALYZE
14
+ SELECT * FROM users WHERE email = 'user@example.com';
15
+
16
+ -- Verbose output with more details
17
+ EXPLAIN (ANALYZE, BUFFERS, VERBOSE)
18
+ SELECT u.*, o.order_total
19
+ FROM users u
20
+ JOIN orders o ON u.id = o.user_id
21
+ WHERE u.created_at > NOW() - INTERVAL '30 days';
22
+ ```
23
+
24
+ **Key Metrics to Watch:**
25
+ - **Seq Scan**: Full table scan (usually slow for large tables)
26
+ - **Index Scan**: Using index (good)
27
+ - **Index Only Scan**: Using index without touching table (best)
28
+ - **Nested Loop**: Join method (okay for small datasets)
29
+ - **Hash Join**: Join method (good for larger datasets)
30
+ - **Merge Join**: Join method (good for sorted data)
31
+ - **Cost**: Estimated query cost (lower is better)
32
+ - **Rows**: Estimated rows returned
33
+ - **Actual Time**: Real execution time
34
+
35
+ ## Index Strategies
36
+
37
+ **Index Types:**
38
+ - **B-Tree**: Default, good for equality and range queries
39
+ - **Hash**: Only for equality (=) comparisons
40
+ - **GIN**: Full-text search, array queries, JSONB
41
+ - **GiST**: Geometric data, full-text search
42
+ - **BRIN**: Block Range INdex for very large tables with correlation
43
+
44
+ ```sql
45
+ -- Standard B-Tree index
46
+ CREATE INDEX idx_users_email ON users(email);
47
+
48
+ -- Composite index (order matters!)
49
+ CREATE INDEX idx_orders_user_status ON orders(user_id, status);
50
+
51
+ -- Partial index (index subset of rows)
52
+ CREATE INDEX idx_active_users ON users(email)
53
+ WHERE status = 'active';
54
+
55
+ -- Expression index
56
+ CREATE INDEX idx_users_lower_email ON users(LOWER(email));
57
+
58
+ -- Covering index (include additional columns)
59
+ CREATE INDEX idx_users_email_covering ON users(email)
60
+ INCLUDE (name, created_at);
61
+
62
+ -- Full-text search index
63
+ CREATE INDEX idx_posts_search ON posts
64
+ USING GIN(to_tsvector('english', title || ' ' || body));
65
+
66
+ -- JSONB index
67
+ CREATE INDEX idx_metadata ON events USING GIN(metadata);
68
+ ```
69
+
70
+ ## Basic Query Rewrites
71
+
72
+ **Avoid SELECT \*:**
73
+ ```sql
74
+ -- Bad: Fetches unnecessary columns
75
+ SELECT * FROM users WHERE id = 123;
76
+
77
+ -- Good: Fetch only what you need
78
+ SELECT id, email, name FROM users WHERE id = 123;
79
+ ```
80
+
81
+ **Use WHERE Clause Efficiently:**
82
+ ```sql
83
+ -- Bad: Function prevents index usage
84
+ SELECT * FROM users WHERE LOWER(email) = 'user@example.com';
85
+
86
+ -- Good: Create functional index or use exact match
87
+ CREATE INDEX idx_users_email_lower ON users(LOWER(email));
88
+ -- Then:
89
+ SELECT * FROM users WHERE LOWER(email) = 'user@example.com';
90
+
91
+ -- Or store normalized data
92
+ SELECT * FROM users WHERE email = 'user@example.com';
93
+ ```
94
+
95
+ **Optimize JOINs:**
96
+ ```sql
97
+ -- Bad: Cartesian product then filter
98
+ SELECT u.name, o.total
99
+ FROM users u, orders o
100
+ WHERE u.id = o.user_id AND u.created_at > '2024-01-01';
101
+
102
+ -- Good: Filter before join
103
+ SELECT u.name, o.total
104
+ FROM users u
105
+ JOIN orders o ON u.id = o.user_id
106
+ WHERE u.created_at > '2024-01-01';
107
+
108
+ -- Better: Filter both tables
109
+ SELECT u.name, o.total
110
+ FROM (SELECT * FROM users WHERE created_at > '2024-01-01') u
111
+ JOIN orders o ON u.id = o.user_id;
112
+ ```