@1aboveio/skills 0.14.0 → 0.16.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (54) hide show
  1. package/README.md +8 -2
  2. package/package.json +1 -1
  3. package/runtime/skills/distribution/generated/recipes.json +178 -34
  4. package/runtime/skills/distribution/scripts/bundles.mjs +11 -3
  5. package/runtime/skills/engineering/engineering-runtime/scripts/workflow-policy.mjs +1 -1
  6. package/skills/{backend → data-science}/pyspark/SKILL.md +42 -40
  7. package/skills/{backend → data-science}/pyspark/references/diagnosis-and-profiling.md +28 -13
  8. package/skills/data-science/pyspark/references/production-validation.md +130 -0
  9. package/skills/data-science/pyspark/references/reconciliation.md +38 -0
  10. package/skills/{backend → data-science}/pyspark/references/transformation-design.md +30 -2
  11. package/skills/engineering/engineering-runtime/coherence/workflow.json +14 -14
  12. package/skills/engineering/engineering-runtime/scripts/workflow-policy.mjs +1 -1
  13. package/skills/engineering/resolve-issues/generated/workflow-repair-policy.json +11 -11
  14. package/skills/engineering/resolve-issues/scripts/run-state.mjs +1 -1
  15. package/skills/payment/fraud-analysis/LICENSE +3 -0
  16. package/skills/payment/fraud-analysis/SKILL.md +113 -0
  17. package/skills/payment/fraud-analysis/evals/evals.json +40 -0
  18. package/skills/payment/fraud-analysis/references/archetypes/authorized-payment-scam.md +41 -0
  19. package/skills/payment/fraud-analysis/references/archetypes/first-party-fraud.md +44 -0
  20. package/skills/payment/fraud-analysis/references/archetypes/third-party-fraud.md +27 -0
  21. package/skills/payment/fraud-analysis/references/contexts/bank-transfer.md +24 -0
  22. package/skills/payment/fraud-analysis/references/contexts/card-payment.md +30 -0
  23. package/skills/payment/fraud-analysis/references/contexts/payment-collection.md +20 -0
  24. package/skills/payment/fraud-analysis/references/contexts/payout.md +20 -0
  25. package/skills/payment/fraud-analysis/references/feature-engineering.md +158 -0
  26. package/skills/payment/fraud-analysis/references/mechanisms/account-takeover.md +36 -0
  27. package/skills/payment/fraud-analysis/references/report-rationale.md +45 -0
  28. package/skills/payment/fraud-analysis/references/report-template.md +190 -0
  29. package/skills/payment/fraud-analysis/references/review-checklist.md +175 -0
  30. package/skills/payment/fraud-analysis/references/taxonomy.md +79 -0
  31. package/skills/payment/fraud-analysis/references/terminology.md +108 -0
  32. package/skills/payment/fraud-analysis/references/workflow.md +175 -0
  33. package/skills/payment/payment-analysis/LICENSE +3 -0
  34. package/skills/payment/payment-analysis/SKILL.md +127 -0
  35. package/skills/payment/payment-analysis/references/auth-rate-actions.md +30 -0
  36. package/skills/payment/payment-analysis/references/chargebacks.md +88 -0
  37. package/skills/payment/payment-analysis/references/event-layers.md +79 -0
  38. package/skills/payment/payment-analysis/references/fx.md +59 -0
  39. package/skills/payment/payment-analysis/references/journey.md +78 -0
  40. package/skills/payment/payment-analysis/references/metrics.md +62 -0
  41. package/skills/payment/payment-analysis/references/report-template.md +98 -0
  42. package/skills/payment/payment-analysis/references/terminology.md +85 -0
  43. package/skills/payment/payment-analysis/references/visualization.md +47 -0
  44. package/skills/backend/pyspark/references/parity-testing.md +0 -83
  45. package/skills/backend/pyspark/references/production-validation.md +0 -166
  46. /package/skills/{backend → data-science}/airflow-dag-develop/LICENSE +0 -0
  47. /package/skills/{backend → data-science}/airflow-dag-develop/SKILL.md +0 -0
  48. /package/skills/{backend → data-science}/pyspark/LICENSE +0 -0
  49. /package/skills/{backend → data-science}/pyspark/assets/templates/etl.py +0 -0
  50. /package/skills/{backend → data-science}/pyspark/assets/templates/utils/__init__.py +0 -0
  51. /package/skills/{backend → data-science}/pyspark/assets/templates/utils/hudi_metadata.py +0 -0
  52. /package/skills/{backend → data-science}/pyspark/references/etl-contract.md +0 -0
  53. /package/skills/{backend → data-science}/pyspark/references/velocity-feature-calculation.md +0 -0
  54. /package/skills/{backend → data-science}/pyspark/scripts/spark_eventlog_summary.py +0 -0
@@ -1,83 +0,0 @@
1
- # Parity Testing
2
-
3
- Use this playbook for semantics-preserving Spark rewrites.
4
-
5
- ## Keep The Old Path As A Bounded Oracle
6
-
7
- Retain the previous implementation only in tests while parity is being proven.
8
- Use small fixtures that are dense enough to expose the old complexity but
9
- bounded enough to run reliably.
10
-
11
- ## Compare Full Rows In Both Directions
12
-
13
- ```python
14
- expected = old_build(source).select(*OUTPUT_COLUMNS)
15
- actual = new_build(source).select(*OUTPUT_COLUMNS)
16
-
17
- assert expected.exceptAll(actual).count() == 0
18
- assert actual.exceptAll(expected).count() == 0
19
- ```
20
-
21
- Counts alone cannot detect wrong values, swapped identities, or duplicate rows.
22
- Use `exceptAll`, not `subtract`, when duplicate multiplicity matters.
23
-
24
- ## Add Literal Boundary Assertions
25
-
26
- Parity can be self-confirming if oracle and candidate share the same mistake.
27
- Assert critical rules directly:
28
-
29
- ```python
30
- by_id = {row.event_id: row for row in actual.collect()}
31
- assert by_id["at-24h"].count_24h == 2 # inclusive lower boundary
32
- assert by_id["after-24h"].count_24h == 1 # one unit outside
33
- assert by_id["same-time-b"].count_1h == 2 # stable tie ordering
34
- ```
35
-
36
- Keep collected fixtures synthetic and non-sensitive.
37
-
38
- ## Minimum Fixture Matrix
39
-
40
- Cover:
41
-
42
- - empty and single-event groups;
43
- - exact lower/upper boundaries and one unit beyond;
44
- - same-timestamp events with every tie-breaker;
45
- - null key, event time, amount, FX, and counterparty;
46
- - both directions and relevant statuses;
47
- - repeated and distinct counterparties;
48
- - state open, partial close, full close, eviction, and re-entry;
49
- - floating-point add/remove residue;
50
- - one dense hot key;
51
- - duplicate and missing output identities.
52
-
53
- ## Test The Physical Shape
54
-
55
- Assert the expensive relation is absent without snapshotting the whole plan:
56
-
57
- ```python
58
- plan = actual._jdf.queryExecution().executedPlan().toString().lower()
59
- assert not ("join" in plan and "event_time" in plan and "interval" in plan)
60
- ```
61
-
62
- Prefer a targeted helper that counts known pair-join nodes. Also assert the
63
- expected grouped-sweep or built-in operator is present when that guards the
64
- optimization.
65
-
66
- ## Mutation-Prove Important Tests
67
-
68
- Temporarily introduce representative faults and confirm tests fail:
69
-
70
- - change `< cutoff` to `<= cutoff`;
71
- - remove a tie-breaker;
72
- - use ordinary rolling addition/subtraction;
73
- - reintroduce the pair join;
74
- - omit a downstream-required column.
75
-
76
- Record red/green evidence for high-risk money, date, and identity semantics.
77
-
78
- ## Match Test Cost To Ownership
79
-
80
- Behavioral Spark tests belong on files that own formulas, source scope,
81
- history, or state transitions. Unrelated upstream changes may use a focused
82
- required-column check instead of rerunning an expensive full-history suite.
83
- Do not demote behavioral coverage for a semantic change.
@@ -1,166 +0,0 @@
1
- # Production Validation
2
-
3
- Use this playbook before production-scale profiling, canaries, backfills, or
4
- writes.
5
-
6
- ## Fail-Closed Admission
7
-
8
- Do not launch until all applicable facts are recorded and valid:
9
-
10
- - exact commit/module hash;
11
- - production Spark, Python, Java, and table-format runtimes;
12
- - source snapshot or date boundary;
13
- - scheduler pause/ownership state;
14
- - zero competing warehouse work;
15
- - explicit validate-only or write-enabled mode;
16
- - durable writable log and event-log destination;
17
- - restart/recovery plan for writes.
18
-
19
- Classify existing processes and working-tree changes before touching them. Do
20
- not kill or overwrite another operator's work.
21
-
22
- ## Run A Validate-Only Canary Before The Writer
23
-
24
- Implement the guard in the Spark job, not only in a wrapper. Keep each stage
25
- literal:
26
-
27
- ```python
28
- import json
29
-
30
- from pyspark.sql import functions as F
31
-
32
-
33
- def validate_candidate(candidate):
34
- # first() is a Spark action: it executes the lazy transform plan.
35
- metrics = candidate.agg(
36
- F.count("*").alias("row_count"),
37
- F.countDistinct("event_id").alias("distinct_event_count"),
38
- F.sum(
39
- F.when(F.col("event_id").isNull(), 1).otherwise(0)
40
- ).alias("null_event_id_count"),
41
- F.sum(
42
- F.when(F.col("event_count_1h") < 0, 1).otherwise(0)
43
- ).alias("invalid_event_count_1h"),
44
- F.count("amount_usd_1h").alias("non_null_amount_usd_1h_count"),
45
- ).first().asDict()
46
-
47
- blockers = []
48
- if metrics["row_count"] != metrics["distinct_event_count"]:
49
- blockers.append("event_id is not unique")
50
- if metrics["null_event_id_count"]:
51
- blockers.append("event_id contains nulls")
52
- if metrics["invalid_event_count_1h"]:
53
- blockers.append("event_count_1h contains negative values")
54
- if blockers:
55
- raise RuntimeError(f"validation blocked: {blockers}")
56
- return metrics
57
-
58
-
59
- def run(validate_only, source_snapshot):
60
- # Bind every action to the same immutable source version or date boundary.
61
- source = extract(source_snapshot)
62
- candidate = transform(source) # build a lazy DataFrame plan
63
-
64
- metrics = validate_candidate(candidate) # execute validation aggregates
65
- print(json.dumps(metrics, sort_keys=True))
66
-
67
- if validate_only: # canary stops before any writer call
68
- print("validate-only: skipping publication")
69
- return
70
-
71
- publish(candidate) # execute the writer from the same source
72
- reconcile_published_scope(metrics) # independently read back aggregates
73
- ```
74
-
75
- Validation and publication are separate Spark actions, so Spark may recompute
76
- the candidate for publication. Pin every source read to the same immutable
77
- snapshot, table version, or exact date boundary. If the source cannot be pinned,
78
- the validation result does not prove what the writer will publish; stop and
79
- redesign the run.
80
-
81
- Adapt identity and blocker checks to the table's required behavior. Reference
82
- every derived output in at least one validation aggregate; otherwise Catalyst
83
- may prune an unreferenced calculation and the canary will not exercise the full
84
- candidate. Validation must use aggregate evidence and fail before publication.
85
- The canary exercises the production-shaped extract, transform, and validation
86
- path without calling the writer.
87
-
88
- ## Avoid Explicit Cache And Persistence
89
-
90
- Do not add `.cache()` or `.persist()` as a performance fix. They add storage,
91
- serialization, eviction, and cleanup concerns without improving the logical
92
- work. Prefer a better transformation shape and a pinned source snapshot.
93
-
94
- Treat existing persistence as separate operational behavior: do not expand it
95
- while optimizing the job, and remove it when repository tests and same-snapshot
96
- evidence show it is unnecessary.
97
-
98
- ## Launch Through The Production Entry Point
99
-
100
- Use the same `spark-submit` path and driver/executor runtimes as production:
101
-
102
- ```bash
103
- export PYSPARK_DRIVER_PYTHON="$DRIVER_PYTHON"
104
- export PYSPARK_PYTHON="$EXECUTOR_PYTHON"
105
-
106
- spark-submit \
107
- --name "$CANARY_NAME" \
108
- --conf spark.pyspark.driver.python="$DRIVER_PYTHON" \
109
- --conf spark.pyspark.python="$EXECUTOR_PYTHON" \
110
- --conf spark.eventLog.enabled=true \
111
- job.py --validate-only
112
- ```
113
-
114
- Do not replace deployed scheduler files with an unmerged staged module. Stage
115
- an isolated artifact or deploy merged code through the normal path.
116
-
117
- ## Prove The Canary Did Not Write
118
-
119
- Require all three forms of evidence:
120
-
121
- 1. Validation completed with zero blockers.
122
- 2. The terminal log says publication was skipped.
123
- 3. The event log reports zero Spark output records and bytes.
124
-
125
- Parse retained logs with upstream failures visible:
126
-
127
- ```bash
128
- set -o pipefail
129
- hdfs dfs -cat "$EVENT_LOG" | \
130
- python {baseDir}/scripts/spark_eventlog_summary.py > canary-summary.json
131
- ```
132
-
133
- Record application ID, shell/event-log wall, stage/task metrics, plan nodes,
134
- shuffle, spill, output, code hash, and snapshot denominator.
135
-
136
- ## Run Write Verification Separately
137
-
138
- Run write-enabled verification only after merge, deployment, successful
139
- validate-only evidence, and explicit authorization. Re-check admission just
140
- before launch.
141
-
142
- Verify row count, distinct identity count, date/partition coverage, commit
143
- success, schema, consumer compatibility, validation blockers, write mode,
144
- single-writer constraints, and downstream readability.
145
-
146
- Report compute/validation, publication, and total time separately. Never
147
- present validate-only time as end-to-end write performance.
148
-
149
- If a write run is cancelled, prove from logs/source that publication was not
150
- reached. YARN `KILLED` alone does not prove no commit occurred.
151
-
152
- ## Keep Evidence Aggregate-Only
153
-
154
- Safe evidence includes counts, null rates, distinct counts, min/max dates,
155
- hashes, stage metrics, plan-node counts, and commit IDs. Do not print
156
- credentials, entity rows, account identifiers, or per-entity values.
157
-
158
- ## Restore Ownership
159
-
160
- After the run:
161
-
162
- 1. Confirm application and wrapper processes are terminal.
163
- 2. Remove agent-owned watchers and temporary processes.
164
- 3. Restore scheduler pause state or ownership.
165
- 4. Separate unrelated downstream failures from the optimized job.
166
- 5. Publish evidence with limitations and unmeasured surfaces.