@1aboveio/skills 0.14.0 → 0.16.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (54) hide show
  1. package/README.md +8 -2
  2. package/package.json +1 -1
  3. package/runtime/skills/distribution/generated/recipes.json +178 -34
  4. package/runtime/skills/distribution/scripts/bundles.mjs +11 -3
  5. package/runtime/skills/engineering/engineering-runtime/scripts/workflow-policy.mjs +1 -1
  6. package/skills/{backend → data-science}/pyspark/SKILL.md +42 -40
  7. package/skills/{backend → data-science}/pyspark/references/diagnosis-and-profiling.md +28 -13
  8. package/skills/data-science/pyspark/references/production-validation.md +130 -0
  9. package/skills/data-science/pyspark/references/reconciliation.md +38 -0
  10. package/skills/{backend → data-science}/pyspark/references/transformation-design.md +30 -2
  11. package/skills/engineering/engineering-runtime/coherence/workflow.json +14 -14
  12. package/skills/engineering/engineering-runtime/scripts/workflow-policy.mjs +1 -1
  13. package/skills/engineering/resolve-issues/generated/workflow-repair-policy.json +11 -11
  14. package/skills/engineering/resolve-issues/scripts/run-state.mjs +1 -1
  15. package/skills/payment/fraud-analysis/LICENSE +3 -0
  16. package/skills/payment/fraud-analysis/SKILL.md +113 -0
  17. package/skills/payment/fraud-analysis/evals/evals.json +40 -0
  18. package/skills/payment/fraud-analysis/references/archetypes/authorized-payment-scam.md +41 -0
  19. package/skills/payment/fraud-analysis/references/archetypes/first-party-fraud.md +44 -0
  20. package/skills/payment/fraud-analysis/references/archetypes/third-party-fraud.md +27 -0
  21. package/skills/payment/fraud-analysis/references/contexts/bank-transfer.md +24 -0
  22. package/skills/payment/fraud-analysis/references/contexts/card-payment.md +30 -0
  23. package/skills/payment/fraud-analysis/references/contexts/payment-collection.md +20 -0
  24. package/skills/payment/fraud-analysis/references/contexts/payout.md +20 -0
  25. package/skills/payment/fraud-analysis/references/feature-engineering.md +158 -0
  26. package/skills/payment/fraud-analysis/references/mechanisms/account-takeover.md +36 -0
  27. package/skills/payment/fraud-analysis/references/report-rationale.md +45 -0
  28. package/skills/payment/fraud-analysis/references/report-template.md +190 -0
  29. package/skills/payment/fraud-analysis/references/review-checklist.md +175 -0
  30. package/skills/payment/fraud-analysis/references/taxonomy.md +79 -0
  31. package/skills/payment/fraud-analysis/references/terminology.md +108 -0
  32. package/skills/payment/fraud-analysis/references/workflow.md +175 -0
  33. package/skills/payment/payment-analysis/LICENSE +3 -0
  34. package/skills/payment/payment-analysis/SKILL.md +127 -0
  35. package/skills/payment/payment-analysis/references/auth-rate-actions.md +30 -0
  36. package/skills/payment/payment-analysis/references/chargebacks.md +88 -0
  37. package/skills/payment/payment-analysis/references/event-layers.md +79 -0
  38. package/skills/payment/payment-analysis/references/fx.md +59 -0
  39. package/skills/payment/payment-analysis/references/journey.md +78 -0
  40. package/skills/payment/payment-analysis/references/metrics.md +62 -0
  41. package/skills/payment/payment-analysis/references/report-template.md +98 -0
  42. package/skills/payment/payment-analysis/references/terminology.md +85 -0
  43. package/skills/payment/payment-analysis/references/visualization.md +47 -0
  44. package/skills/backend/pyspark/references/parity-testing.md +0 -83
  45. package/skills/backend/pyspark/references/production-validation.md +0 -166
  46. /package/skills/{backend → data-science}/airflow-dag-develop/LICENSE +0 -0
  47. /package/skills/{backend → data-science}/airflow-dag-develop/SKILL.md +0 -0
  48. /package/skills/{backend → data-science}/pyspark/LICENSE +0 -0
  49. /package/skills/{backend → data-science}/pyspark/assets/templates/etl.py +0 -0
  50. /package/skills/{backend → data-science}/pyspark/assets/templates/utils/__init__.py +0 -0
  51. /package/skills/{backend → data-science}/pyspark/assets/templates/utils/hudi_metadata.py +0 -0
  52. /package/skills/{backend → data-science}/pyspark/references/etl-contract.md +0 -0
  53. /package/skills/{backend → data-science}/pyspark/references/velocity-feature-calculation.md +0 -0
  54. /package/skills/{backend → data-science}/pyspark/scripts/spark_eventlog_summary.py +0 -0
package/README.md CHANGED
@@ -67,7 +67,9 @@ planned same-named skill. The coordinator still owns group and new-member consen
67
67
 
68
68
  - `engineering-workflow` — Engineering Workflow: e2e-test, ensure-coverage, implement-and-pr, resolve-issues, resolve-release, review-pr, smoke, engineering-runtime
69
69
  - `delivery-infrastructure` — Delivery Infrastructure: cloud-build, cloud-debug, cloud-deploy, docker, google-cloud, mergify, podman
70
- - `backend` — Backend: airflow-dag-develop, app-debug, pyspark, python-backend
70
+ - `backend` — Backend: app-debug, python-backend
71
+ - `data-science` — Data Science: airflow-dag-develop, pyspark
72
+ - `payment` — Payment: fraud-analysis, payment-analysis
71
73
  - `fullstack` — Fullstack: better-auth, monorepo, nextjs-fullstack, prisma-setup, shadcn, zod-v4
72
74
 
73
75
  Internal dependencies, installed with the groups above and not selectable on their own: `harness-runtime`.
@@ -119,7 +121,11 @@ npx skills@1.5.22 remove --global --agent universal claude-code --skill e2e-test
119
121
  # Delivery Infrastructure
120
122
  npx skills@1.5.22 remove --global --agent universal claude-code --skill cloud-build cloud-debug cloud-deploy docker google-cloud mergify podman
121
123
  # Backend
122
- npx skills@1.5.22 remove --global --agent universal claude-code --skill airflow-dag-develop app-debug pyspark python-backend
124
+ npx skills@1.5.22 remove --global --agent universal claude-code --skill app-debug python-backend
125
+ # Data Science
126
+ npx skills@1.5.22 remove --global --agent universal claude-code --skill airflow-dag-develop pyspark
127
+ # Payment
128
+ npx skills@1.5.22 remove --global --agent universal claude-code --skill fraud-analysis payment-analysis
123
129
  # Fullstack
124
130
  npx skills@1.5.22 remove --global --agent universal claude-code --skill better-auth monorepo nextjs-fullstack prisma-setup shadcn zod-v4
125
131
  ```
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@1aboveio/skills",
3
- "version": "0.14.0",
3
+ "version": "0.16.1",
4
4
  "description": "Install the 1AboveIO first-party skill groups through the native Skills CLI.",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -8,7 +8,7 @@
8
8
  "id": "first-party",
9
9
  "type": "first-party",
10
10
  "package": "@1aboveio/skills",
11
- "version": "0.14.0",
11
+ "version": "0.16.1",
12
12
  "updatePolicy": "pinned-npm-version",
13
13
  "members": [
14
14
  {
@@ -102,26 +102,38 @@
102
102
  "licenseResponsibility": "1aboveio-mit"
103
103
  },
104
104
  {
105
- "installName": "airflow-dag-develop",
106
- "path": "skills/backend/airflow-dag-develop",
105
+ "installName": "app-debug",
106
+ "path": "skills/backend/app-debug",
107
107
  "role": "skill",
108
108
  "licenseResponsibility": "1aboveio-mit"
109
109
  },
110
110
  {
111
- "installName": "app-debug",
112
- "path": "skills/backend/app-debug",
111
+ "installName": "python-backend",
112
+ "path": "skills/backend/python-backend",
113
+ "role": "skill",
114
+ "licenseResponsibility": "1aboveio-mit"
115
+ },
116
+ {
117
+ "installName": "airflow-dag-develop",
118
+ "path": "skills/data-science/airflow-dag-develop",
113
119
  "role": "skill",
114
120
  "licenseResponsibility": "1aboveio-mit"
115
121
  },
116
122
  {
117
123
  "installName": "pyspark",
118
- "path": "skills/backend/pyspark",
124
+ "path": "skills/data-science/pyspark",
119
125
  "role": "skill",
120
126
  "licenseResponsibility": "1aboveio-mit"
121
127
  },
122
128
  {
123
- "installName": "python-backend",
124
- "path": "skills/backend/python-backend",
129
+ "installName": "fraud-analysis",
130
+ "path": "skills/payment/fraud-analysis",
131
+ "role": "skill",
132
+ "licenseResponsibility": "1aboveio-mit"
133
+ },
134
+ {
135
+ "installName": "payment-analysis",
136
+ "path": "skills/payment/payment-analysis",
125
137
  "role": "skill",
126
138
  "licenseResponsibility": "1aboveio-mit"
127
139
  },
@@ -184,7 +196,7 @@
184
196
  "sourceId": "first-party",
185
197
  "sourceType": "first-party",
186
198
  "package": "@1aboveio/skills",
187
- "version": "0.14.0",
199
+ "version": "0.16.1",
188
200
  "installPath": null,
189
201
  "members": [
190
202
  "harness-runtime",
@@ -218,7 +230,7 @@
218
230
  "sourceId": "first-party",
219
231
  "sourceType": "first-party",
220
232
  "package": "@1aboveio/skills",
221
- "version": "0.14.0",
233
+ "version": "0.16.1",
222
234
  "installPath": null,
223
235
  "members": [
224
236
  "harness-runtime",
@@ -262,7 +274,7 @@
262
274
  "sourceId": "first-party",
263
275
  "sourceType": "first-party",
264
276
  "package": "@1aboveio/skills",
265
- "version": "0.14.0",
277
+ "version": "0.16.1",
266
278
  "installPath": null,
267
279
  "members": [
268
280
  "harness-runtime",
@@ -295,7 +307,7 @@
295
307
  "sourceId": "first-party",
296
308
  "sourceType": "first-party",
297
309
  "package": "@1aboveio/skills",
298
- "version": "0.14.0",
310
+ "version": "0.16.1",
299
311
  "installPath": null,
300
312
  "members": [
301
313
  "harness-runtime",
@@ -338,19 +350,17 @@
338
350
  "sourceId": "first-party",
339
351
  "sourceType": "first-party",
340
352
  "package": "@1aboveio/skills",
341
- "version": "0.14.0",
353
+ "version": "0.16.1",
342
354
  "installPath": null,
343
355
  "members": [
344
356
  "harness-runtime",
345
- "airflow-dag-develop",
346
357
  "app-debug",
347
- "pyspark",
348
358
  "python-backend"
349
359
  ],
350
360
  "commands": [
351
361
  {
352
362
  "transport": "local",
353
- "command": "npx skills@1.5.22 add . --global --agent universal claude-code --skill harness-runtime airflow-dag-develop app-debug pyspark python-backend"
363
+ "command": "npx skills@1.5.22 add . --global --agent universal claude-code --skill harness-runtime app-debug python-backend"
354
364
  }
355
365
  ],
356
366
  "onFailure": {
@@ -368,19 +378,149 @@
368
378
  "sourceId": "first-party",
369
379
  "sourceType": "first-party",
370
380
  "package": "@1aboveio/skills",
371
- "version": "0.14.0",
381
+ "version": "0.16.1",
372
382
  "installPath": null,
373
383
  "members": [
374
384
  "harness-runtime",
375
- "airflow-dag-develop",
376
385
  "app-debug",
377
- "pyspark",
378
386
  "python-backend"
379
387
  ],
380
388
  "commands": [
381
389
  {
382
390
  "transport": "local",
383
- "command": "npx skills@1.5.22 add . --global --agent universal claude-code --skill harness-runtime airflow-dag-develop app-debug pyspark python-backend"
391
+ "command": "npx skills@1.5.22 add . --global --agent universal claude-code --skill harness-runtime app-debug python-backend"
392
+ }
393
+ ],
394
+ "onFailure": {
395
+ "action": "stop",
396
+ "message": "Source first-party fetch failed; stop this recipe before running another source."
397
+ }
398
+ }
399
+ ]
400
+ },
401
+ "uninstall": {
402
+ "command": "npx skills@1.5.22 remove --global --agent universal claude-code --skill app-debug python-backend",
403
+ "removesDependencies": false
404
+ }
405
+ },
406
+ {
407
+ "id": "data-science",
408
+ "displayName": "Data Science",
409
+ "dependsOn": [
410
+ "harness-runtime"
411
+ ],
412
+ "install": {
413
+ "stopOnFailure": true,
414
+ "steps": [
415
+ {
416
+ "sourceId": "first-party",
417
+ "sourceType": "first-party",
418
+ "package": "@1aboveio/skills",
419
+ "version": "0.16.1",
420
+ "installPath": null,
421
+ "members": [
422
+ "harness-runtime",
423
+ "airflow-dag-develop",
424
+ "pyspark"
425
+ ],
426
+ "commands": [
427
+ {
428
+ "transport": "local",
429
+ "command": "npx skills@1.5.22 add . --global --agent universal claude-code --skill harness-runtime airflow-dag-develop pyspark"
430
+ }
431
+ ],
432
+ "onFailure": {
433
+ "action": "stop",
434
+ "message": "Source first-party fetch failed; stop this recipe before running another source."
435
+ }
436
+ }
437
+ ]
438
+ },
439
+ "update": {
440
+ "mode": "rerun-install-recipe",
441
+ "stopOnFailure": true,
442
+ "steps": [
443
+ {
444
+ "sourceId": "first-party",
445
+ "sourceType": "first-party",
446
+ "package": "@1aboveio/skills",
447
+ "version": "0.16.1",
448
+ "installPath": null,
449
+ "members": [
450
+ "harness-runtime",
451
+ "airflow-dag-develop",
452
+ "pyspark"
453
+ ],
454
+ "commands": [
455
+ {
456
+ "transport": "local",
457
+ "command": "npx skills@1.5.22 add . --global --agent universal claude-code --skill harness-runtime airflow-dag-develop pyspark"
458
+ }
459
+ ],
460
+ "onFailure": {
461
+ "action": "stop",
462
+ "message": "Source first-party fetch failed; stop this recipe before running another source."
463
+ }
464
+ }
465
+ ]
466
+ },
467
+ "uninstall": {
468
+ "command": "npx skills@1.5.22 remove --global --agent universal claude-code --skill airflow-dag-develop pyspark",
469
+ "removesDependencies": false
470
+ }
471
+ },
472
+ {
473
+ "id": "payment",
474
+ "displayName": "Payment",
475
+ "dependsOn": [
476
+ "harness-runtime"
477
+ ],
478
+ "install": {
479
+ "stopOnFailure": true,
480
+ "steps": [
481
+ {
482
+ "sourceId": "first-party",
483
+ "sourceType": "first-party",
484
+ "package": "@1aboveio/skills",
485
+ "version": "0.16.1",
486
+ "installPath": null,
487
+ "members": [
488
+ "harness-runtime",
489
+ "fraud-analysis",
490
+ "payment-analysis"
491
+ ],
492
+ "commands": [
493
+ {
494
+ "transport": "local",
495
+ "command": "npx skills@1.5.22 add . --global --agent universal claude-code --skill harness-runtime fraud-analysis payment-analysis"
496
+ }
497
+ ],
498
+ "onFailure": {
499
+ "action": "stop",
500
+ "message": "Source first-party fetch failed; stop this recipe before running another source."
501
+ }
502
+ }
503
+ ]
504
+ },
505
+ "update": {
506
+ "mode": "rerun-install-recipe",
507
+ "stopOnFailure": true,
508
+ "steps": [
509
+ {
510
+ "sourceId": "first-party",
511
+ "sourceType": "first-party",
512
+ "package": "@1aboveio/skills",
513
+ "version": "0.16.1",
514
+ "installPath": null,
515
+ "members": [
516
+ "harness-runtime",
517
+ "fraud-analysis",
518
+ "payment-analysis"
519
+ ],
520
+ "commands": [
521
+ {
522
+ "transport": "local",
523
+ "command": "npx skills@1.5.22 add . --global --agent universal claude-code --skill harness-runtime fraud-analysis payment-analysis"
384
524
  }
385
525
  ],
386
526
  "onFailure": {
@@ -391,7 +531,7 @@
391
531
  ]
392
532
  },
393
533
  "uninstall": {
394
- "command": "npx skills@1.5.22 remove --global --agent universal claude-code --skill airflow-dag-develop app-debug pyspark python-backend",
534
+ "command": "npx skills@1.5.22 remove --global --agent universal claude-code --skill fraud-analysis payment-analysis",
395
535
  "removesDependencies": false
396
536
  }
397
537
  },
@@ -408,7 +548,7 @@
408
548
  "sourceId": "first-party",
409
549
  "sourceType": "first-party",
410
550
  "package": "@1aboveio/skills",
411
- "version": "0.14.0",
551
+ "version": "0.16.1",
412
552
  "installPath": null,
413
553
  "members": [
414
554
  "harness-runtime",
@@ -440,7 +580,7 @@
440
580
  "sourceId": "first-party",
441
581
  "sourceType": "first-party",
442
582
  "package": "@1aboveio/skills",
443
- "version": "0.14.0",
583
+ "version": "0.16.1",
444
584
  "installPath": null,
445
585
  "members": [
446
586
  "harness-runtime",
@@ -480,7 +620,7 @@
480
620
  "sourceId": "first-party",
481
621
  "sourceType": "first-party",
482
622
  "package": "@1aboveio/skills",
483
- "version": "0.14.0",
623
+ "version": "0.16.1",
484
624
  "installPath": null,
485
625
  "members": [
486
626
  "harness-runtime"
@@ -506,7 +646,7 @@
506
646
  "sourceId": "first-party",
507
647
  "sourceType": "first-party",
508
648
  "package": "@1aboveio/skills",
509
- "version": "0.14.0",
649
+ "version": "0.16.1",
510
650
  "installPath": null,
511
651
  "members": [
512
652
  "harness-runtime"
@@ -538,7 +678,7 @@
538
678
  "sourceId": "first-party",
539
679
  "sourceType": "first-party",
540
680
  "package": "@1aboveio/skills",
541
- "version": "0.14.0",
681
+ "version": "0.16.1",
542
682
  "installPath": null,
543
683
  "members": [
544
684
  "harness-runtime",
@@ -557,10 +697,12 @@
557
697
  "google-cloud",
558
698
  "mergify",
559
699
  "podman",
560
- "airflow-dag-develop",
561
700
  "app-debug",
562
- "pyspark",
563
701
  "python-backend",
702
+ "airflow-dag-develop",
703
+ "pyspark",
704
+ "fraud-analysis",
705
+ "payment-analysis",
564
706
  "better-auth",
565
707
  "monorepo",
566
708
  "nextjs-fullstack",
@@ -571,7 +713,7 @@
571
713
  "commands": [
572
714
  {
573
715
  "transport": "local",
574
- "command": "npx skills@1.5.22 add . --global --agent universal claude-code --skill harness-runtime e2e-test ensure-coverage implement-and-pr resolve-issues resolve-release review-pr smoke engineering-runtime cloud-build cloud-debug cloud-deploy docker google-cloud mergify podman airflow-dag-develop app-debug pyspark python-backend better-auth monorepo nextjs-fullstack prisma-setup shadcn zod-v4 && npm --prefix \"$(realpath \"$HOME/.agents/skills/engineering-runtime/scripts\")\" ci --ignore-scripts"
716
+ "command": "npx skills@1.5.22 add . --global --agent universal claude-code --skill harness-runtime e2e-test ensure-coverage implement-and-pr resolve-issues resolve-release review-pr smoke engineering-runtime cloud-build cloud-debug cloud-deploy docker google-cloud mergify podman app-debug python-backend airflow-dag-develop pyspark fraud-analysis payment-analysis better-auth monorepo nextjs-fullstack prisma-setup shadcn zod-v4 && npm --prefix \"$(realpath \"$HOME/.agents/skills/engineering-runtime/scripts\")\" ci --ignore-scripts"
575
717
  }
576
718
  ],
577
719
  "onFailure": {
@@ -589,7 +731,7 @@
589
731
  "sourceId": "first-party",
590
732
  "sourceType": "first-party",
591
733
  "package": "@1aboveio/skills",
592
- "version": "0.14.0",
734
+ "version": "0.16.1",
593
735
  "installPath": null,
594
736
  "members": [
595
737
  "harness-runtime",
@@ -608,10 +750,12 @@
608
750
  "google-cloud",
609
751
  "mergify",
610
752
  "podman",
611
- "airflow-dag-develop",
612
753
  "app-debug",
613
- "pyspark",
614
754
  "python-backend",
755
+ "airflow-dag-develop",
756
+ "pyspark",
757
+ "fraud-analysis",
758
+ "payment-analysis",
615
759
  "better-auth",
616
760
  "monorepo",
617
761
  "nextjs-fullstack",
@@ -622,7 +766,7 @@
622
766
  "commands": [
623
767
  {
624
768
  "transport": "local",
625
- "command": "npx skills@1.5.22 add . --global --agent universal claude-code --skill harness-runtime e2e-test ensure-coverage implement-and-pr resolve-issues resolve-release review-pr smoke engineering-runtime cloud-build cloud-debug cloud-deploy docker google-cloud mergify podman airflow-dag-develop app-debug pyspark python-backend better-auth monorepo nextjs-fullstack prisma-setup shadcn zod-v4 && npm --prefix \"$(realpath \"$HOME/.agents/skills/engineering-runtime/scripts\")\" ci --ignore-scripts"
769
+ "command": "npx skills@1.5.22 add . --global --agent universal claude-code --skill harness-runtime e2e-test ensure-coverage implement-and-pr resolve-issues resolve-release review-pr smoke engineering-runtime cloud-build cloud-debug cloud-deploy docker google-cloud mergify podman app-debug python-backend airflow-dag-develop pyspark fraud-analysis payment-analysis better-auth monorepo nextjs-fullstack prisma-setup shadcn zod-v4 && npm --prefix \"$(realpath \"$HOME/.agents/skills/engineering-runtime/scripts\")\" ci --ignore-scripts"
626
770
  }
627
771
  ],
628
772
  "onFailure": {
@@ -633,7 +777,7 @@
633
777
  ]
634
778
  },
635
779
  "uninstall": {
636
- "command": "npx skills@1.5.22 remove --global --agent universal claude-code --skill e2e-test ensure-coverage implement-and-pr resolve-issues resolve-release review-pr smoke engineering-runtime cloud-build cloud-debug cloud-deploy docker google-cloud mergify podman airflow-dag-develop app-debug pyspark python-backend better-auth monorepo nextjs-fullstack prisma-setup shadcn zod-v4 harness-runtime"
780
+ "command": "npx skills@1.5.22 remove --global --agent universal claude-code --skill e2e-test ensure-coverage implement-and-pr resolve-issues resolve-release review-pr smoke engineering-runtime cloud-build cloud-debug cloud-deploy docker google-cloud mergify podman app-debug python-backend airflow-dag-develop pyspark fraud-analysis payment-analysis better-auth monorepo nextjs-fullstack prisma-setup shadcn zod-v4 harness-runtime"
637
781
  }
638
782
  }
639
783
  }
@@ -38,6 +38,10 @@ const COMMAND_TRANSPORTS = ['local', 'https'];
38
38
  // is never presented as something a user picks. The picker, `list`, the plan and the generated
39
39
  // package README all derive their answer from here rather than each re-listing the ids.
40
40
  export const SUPPORT_GROUPS = new Set(['harness-runtime']);
41
+ // External toolkits stay in the manifest for standalone docs/canary, but the coordinator — checkout
42
+ // and the public package alike — does not offer or install them.
43
+ export const EXTERNAL_ONLY_GROUPS = new Set(['matt-pocock-toolkit']);
44
+ const HIDDEN_GROUPS = new Set([...SUPPORT_GROUPS, ...EXTERNAL_ONLY_GROUPS]);
41
45
  const ACTIONS = new Set(['install', 'update']);
42
46
 
43
47
  // Selection memory (ADR 0002/0004 amendment, epic #1124): the remembered user-facing group
@@ -113,7 +117,10 @@ function writeSelectionState(state, env = process.env) {
113
117
  // user-selected) are dropped with a warning; memory never resurrects them and never falls back to
114
118
  // filesystem inference. Nothing left is an error, not an empty plan.
115
119
  function rememberedGroupIds(state, catalog, stderr) {
116
- const known = new Set(catalog.map((bundle) => bundle.id));
120
+ const known = new Set([
121
+ ...catalog.map((bundle) => bundle.id),
122
+ ...EXTERNAL_ONLY_GROUPS,
123
+ ]);
117
124
  const kept = [];
118
125
  const dropped = [];
119
126
  for (const group of state.groups) {
@@ -256,7 +263,7 @@ function assertLocalRoot(localRoot) {
256
263
 
257
264
  export function bundleCatalog(transport = {}) {
258
265
  return readRecipes(resolveBundleTransport(transport).recipesPath).groups
259
- .filter((group) => !SUPPORT_GROUPS.has(group.id))
266
+ .filter((group) => !HIDDEN_GROUPS.has(group.id))
260
267
  .map((group) => ({
261
268
  id: group.id,
262
269
  displayName: group.displayName,
@@ -744,7 +751,8 @@ async function consentNewMembers(candidates, { yes, stdin, stdout, stderr }) {
744
751
 
745
752
  function selectionIncludesTdd(groupIds, catalog) {
746
753
  const selected = new Set(groupIds);
747
- return catalog.some((bundle) => selected.has(bundle.id) && bundle.members.includes('tdd'));
754
+ return catalog.some((bundle) => selected.has(bundle.id) && bundle.members.includes('tdd'))
755
+ || groupIds.includes('matt-pocock-toolkit');
748
756
  }
749
757
 
750
758
  // Dependencies are injectable so the coordinator seam can be tested black-box: a stubbed native
@@ -97,7 +97,7 @@ export const WORKFLOW_TRUSTED_SOURCES = deepFreeze({
97
97
  id: 'first-party',
98
98
  type: 'first-party',
99
99
  package: '@1aboveio/skills',
100
- version: '0.14.0',
100
+ version: '0.16.1',
101
101
  },
102
102
  matt: {
103
103
  id: 'matt-pocock',
@@ -1,17 +1,27 @@
1
1
  ---
2
2
  name: pyspark
3
- description: Implement, debug, review, and optimize Apache Spark or PySpark batch jobs, ETL pipelines, feature calculations, and Hudi/warehouse writes. Use whenever the task involves Spark DataFrames or Spark SQL, joins, windows, grouped state, UDFs, shuffle, skew, spill, slow stages, driver/executor OOM, Spark Connect, spark-submit, Hudi commits, backfills, bidirectional parity tests, validate-only canaries, or Spark event-log profiling — including when the user only mentions a warehouse job timeout, empty Hudi write, or performance regression on an EMR/Dataproc job.
3
+ description: >
4
+ Implement, debug, review, and optimize Apache Spark or PySpark batch jobs,
5
+ ETL pipelines, feature calculations, and Hudi/warehouse writes. Follow
6
+ contract → transform → materialize → local test → production validation
7
+ (one-month canary, optional profile if that canary exceeds 5 minutes without
8
+ resource constraints, authorized write, cluster reconciliation). Use whenever
9
+ the task involves Spark DataFrames or Spark SQL, joins, windows, grouped
10
+ state, UDFs, shuffle, skew, spill, slow stages, driver/executor OOM, Spark
11
+ Connect, spark-submit, Hudi commits, backfills, validate-only canaries,
12
+ cluster reconciliation, or Spark event-log profiling — including warehouse
13
+ job timeouts, empty Hudi writes, or EMR/Dataproc performance regressions.
4
14
  ---
5
15
 
6
16
  # PySpark
7
17
 
8
- Use this skill as the workflow and routing index. Read the referenced playbooks
9
- that match the task before editing code or running a production-scale job.
18
+ Use this skill as the workflow and routing index. Read the playbook for the
19
+ step you are on before editing code or submitting to the cluster.
10
20
 
11
21
  ## Workflow
12
22
 
13
23
  ```text
14
- contract -> transform -> materialize -> test -> CI/CD
24
+ contract -> transform -> materialize -> test -> production validation
15
25
  ```
16
26
 
17
27
  Follow all five steps in order.
@@ -60,48 +70,40 @@ windows or FIFO paths, also read
60
70
  - Keep the validate-only guard inside the job and before every writer. Do not
61
71
  add `.count()`, `.cache()`, or `.persist()` to warm execution.
62
72
 
63
- Read [Production Validation](references/production-validation.md).
64
-
65
73
  ### 4. Test
66
74
 
67
- 1. **Local test:** run pure state, tiny real-Spark, plan-shape, focused repository,
68
- and bidirectional parity tests. Cover boundaries, ties, nulls, arithmetic,
69
- identity, and precedence.
70
- 2. **Cluster canary test:** run the production entry point and source scope with
71
- `--validate-only`. Require zero blockers, an explicit publication skip, zero
72
- Spark output records/bytes, and readable driver/event logs.
73
- 3. **Profiling:** measure an isolated baseline, parse the canary event log, and
74
- compare both runs on the same source scope and resource profile. Record
75
- wall/stage/task time, cardinality, shuffle, spill, skew, output, and target
76
- plan patterns.
75
+ Run **local** tests only. Do not treat a laptop Spark session as cluster
76
+ evidence.
77
77
 
78
- ```bash
79
- spark-submit <production-options> job.py --validate-only
80
- set -o pipefail
81
- hdfs dfs -cat "$EVENT_LOG" | \
82
- python {baseDir}/scripts/spark_eventlog_summary.py > canary-summary.json
83
- ```
78
+ - Pure state, tiny real-Spark, plan-shape, and focused repository tests.
79
+ - Cover boundaries, ties, nulls, arithmetic, identity, and precedence.
84
80
 
85
- Reject missing evidence, incomparable scope, unexplained regressions, or a
86
- missed performance gate. Keep validate-only time separate from write time.
81
+ ### 5. Production validation
87
82
 
88
- Read [Parity Testing](references/parity-testing.md),
89
- [Diagnosis And Profiling](references/diagnosis-and-profiling.md), and
90
- [Production Validation](references/production-validation.md).
83
+ Run the production entry point on the cluster against a **small sample**,
84
+ typically **one month** of source data. Local tests do not substitute.
91
85
 
92
- ### 5. CI/CD
86
+ 1. **Canary:** `spark-submit` with `--validate-only` on that sample and the
87
+ production runtime. Require zero blockers, an explicit publication skip,
88
+ zero Spark output records/bytes, and readable driver/event logs.
89
+ 2. **Profiling (optional):** do this only when the small-sample canary wall
90
+ time exceeds **5 minutes** **and** the run was not resource-constrained
91
+ (queued, under-provisioned executors, competing warehouse work). Otherwise
92
+ skip. When required, follow
93
+ [Diagnosis And Profiling](references/diagnosis-and-profiling.md).
94
+ 3. **Write (authorized):** re-check admission, then publish the same pinned
95
+ sample. Do not present canary time as write performance.
96
+ 4. **Reconciliation:** independently prove the published sample. Follow
97
+ [Reconciliation](references/reconciliation.md).
98
+
99
+ ```bash
100
+ spark-submit <production-options> job.py --validate-only
101
+ ```
93
102
 
94
- 1. Open a PR with required-behavior, test, canary, profiling, and risk evidence.
95
- 2. Run required CI and independent review; fix failures.
96
- 3. Enqueue through the configured merge queue and merge without bypassing gates.
97
- 4. Deploy only the merged commit through the normal deployment path.
98
- 5. Run write-enabled production validation only with explicit authorization;
99
- reconcile counts, identity, schema, coverage, blockers, readability, write
100
- mode, and commit evidence independently.
101
- 6. Restore scheduler/process ownership and report compute/validation,
102
- publication, and end-to-end time separately.
103
+ Restore scheduler/process ownership after the run.
103
104
 
104
- Read [Production Validation](references/production-validation.md).
105
+ Read [Production Validation](references/production-validation.md) for admission,
106
+ canary evidence, and write rules.
105
107
 
106
108
  ## Reference Index
107
109
 
@@ -111,6 +113,6 @@ Read [Production Validation](references/production-validation.md).
111
113
  | Diagnose or profile | [Diagnosis And Profiling](references/diagnosis-and-profiling.md) |
112
114
  | Design transformations | [Transformation Design](references/transformation-design.md) |
113
115
  | Optimize temporal/FIFO state | [Velocity Feature Calculation](references/velocity-feature-calculation.md) |
114
- | Prove parity | [Parity Testing](references/parity-testing.md) |
115
- | Run canaries or writes | [Production Validation](references/production-validation.md) |
116
+ | Reconcile published sample vs previous path | [Reconciliation](references/reconciliation.md) |
117
+ | Run production validation (canary, optional profile, write, reconciliation) | [Production Validation](references/production-validation.md) |
116
118
  | Summarize an event log | `scripts/spark_eventlog_summary.py` |
@@ -1,9 +1,24 @@
1
1
  # Diagnosis And Profiling
2
2
 
3
3
  Use this reference to decide whether a Spark job is fast enough, identify the
4
- largest performance gap, and determine whether a fix is ready to ship. Focus on
5
- the required evidence and outcome; choose environment-specific commands and
6
- APIs from the available runtime.
4
+ largest performance gap, and determine whether a canary is ready for write and
5
+ reconciliation. Focus on evidence and outcome; choose environment-specific
6
+ commands from the available runtime.
7
+
8
+ ## When Profiling Is Required
9
+
10
+ Profiling is optional during production validation. It is **required** only when
11
+ the small-sample canary (typically one month) wall time exceeds **5 minutes**
12
+ and the run was **not** resource-constrained (queued, under-provisioned
13
+ executors, competing warehouse work). Then parse the canary event log and
14
+ compare an isolated baseline on the same sample and resource profile. Keep
15
+ validate-only time separate from write time.
16
+
17
+ ```bash
18
+ set -o pipefail
19
+ hdfs dfs -cat "$EVENT_LOG" | \
20
+ python {baseDir}/scripts/spark_eventlog_summary.py > canary-summary.json
21
+ ```
7
22
 
8
23
  ## Performance Target
9
24
 
@@ -114,13 +129,12 @@ settings merely because the slow work becomes visible at a later action.
114
129
  3. **Make the fix:** reduce the responsible logical or physical work without
115
130
  tuning parameters, adding resources, or weakening output identity, ordering,
116
131
  boundaries, arithmetic, history, validation, or write behavior.
117
- 4. **Validate locally:** prove required behavior, bidirectional parity, boundary
118
- cases, and the intended physical-plan change with focused tests.
119
- 5. **Canary and profile in the cluster:** run the production entry point in
120
- validate-only mode on the comparable isolated snapshot; retain logs, prove
121
- zero writes, and compare the same metrics with the baseline.
122
- 6. **Ship:** open and review the PR, pass CI, enqueue and merge, deploy the merged
123
- commit, then run authorized production validation and reconciliation.
132
+ 4. **Validate locally:** prove required behavior, boundary cases, and the
133
+ intended physical-plan change with focused tests. Do not treat local Spark as
134
+ cluster parity.
135
+ 5. **Production validation:** one-month (typical) sample: canary, optional
136
+ profile (only if that canary exceeds 5 minutes without resource constraints),
137
+ authorized write, then reconciliation.
124
138
 
125
139
  Repeat the loop when the target is still missed or the dominant gap moves.
126
140
 
@@ -132,9 +146,10 @@ Repeat the loop when the target is still missed or the dominant gap moves.
132
146
  confirmation of the same runtime, resources, and isolation conditions.
133
147
  - **Performance:** wall time, dominant stage, maximum task, cardinality,
134
148
  shuffle/spill, logic change, and pass/fail against the 5-minute target.
135
- - **Correctness:** local parity/plan results and canary blocker/output status.
136
- - **Delivery:** deployed hash, publication/reconciliation result when applicable,
137
- and remaining limitations.
149
+ - **Correctness:** local plan/boundary results, canary blocker/output status,
150
+ and cluster reconciliation (aggregates + bidirectional parity).
151
+ - **Delivery:** cluster application IDs, published hash/commit, reconciliation
152
+ result when a write ran, and remaining limitations.
138
153
 
139
154
  Keep evidence aggregate-only. Never print credentials or entity-level rows.
140
155