@1aboveio/skills 0.14.0 → 0.16.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +8 -2
- package/package.json +1 -1
- package/runtime/skills/distribution/generated/recipes.json +178 -34
- package/runtime/skills/distribution/scripts/bundles.mjs +11 -3
- package/runtime/skills/engineering/engineering-runtime/scripts/workflow-policy.mjs +1 -1
- package/skills/{backend → data-science}/pyspark/SKILL.md +42 -40
- package/skills/{backend → data-science}/pyspark/references/diagnosis-and-profiling.md +28 -13
- package/skills/data-science/pyspark/references/production-validation.md +130 -0
- package/skills/data-science/pyspark/references/reconciliation.md +38 -0
- package/skills/{backend → data-science}/pyspark/references/transformation-design.md +30 -2
- package/skills/engineering/engineering-runtime/coherence/workflow.json +14 -14
- package/skills/engineering/engineering-runtime/scripts/workflow-policy.mjs +1 -1
- package/skills/engineering/resolve-issues/generated/workflow-repair-policy.json +11 -11
- package/skills/engineering/resolve-issues/scripts/run-state.mjs +1 -1
- package/skills/payment/fraud-analysis/LICENSE +3 -0
- package/skills/payment/fraud-analysis/SKILL.md +113 -0
- package/skills/payment/fraud-analysis/evals/evals.json +40 -0
- package/skills/payment/fraud-analysis/references/archetypes/authorized-payment-scam.md +41 -0
- package/skills/payment/fraud-analysis/references/archetypes/first-party-fraud.md +44 -0
- package/skills/payment/fraud-analysis/references/archetypes/third-party-fraud.md +27 -0
- package/skills/payment/fraud-analysis/references/contexts/bank-transfer.md +24 -0
- package/skills/payment/fraud-analysis/references/contexts/card-payment.md +30 -0
- package/skills/payment/fraud-analysis/references/contexts/payment-collection.md +20 -0
- package/skills/payment/fraud-analysis/references/contexts/payout.md +20 -0
- package/skills/payment/fraud-analysis/references/feature-engineering.md +158 -0
- package/skills/payment/fraud-analysis/references/mechanisms/account-takeover.md +36 -0
- package/skills/payment/fraud-analysis/references/report-rationale.md +45 -0
- package/skills/payment/fraud-analysis/references/report-template.md +190 -0
- package/skills/payment/fraud-analysis/references/review-checklist.md +175 -0
- package/skills/payment/fraud-analysis/references/taxonomy.md +79 -0
- package/skills/payment/fraud-analysis/references/terminology.md +108 -0
- package/skills/payment/fraud-analysis/references/workflow.md +175 -0
- package/skills/payment/payment-analysis/LICENSE +3 -0
- package/skills/payment/payment-analysis/SKILL.md +127 -0
- package/skills/payment/payment-analysis/references/auth-rate-actions.md +30 -0
- package/skills/payment/payment-analysis/references/chargebacks.md +88 -0
- package/skills/payment/payment-analysis/references/event-layers.md +79 -0
- package/skills/payment/payment-analysis/references/fx.md +59 -0
- package/skills/payment/payment-analysis/references/journey.md +78 -0
- package/skills/payment/payment-analysis/references/metrics.md +62 -0
- package/skills/payment/payment-analysis/references/report-template.md +98 -0
- package/skills/payment/payment-analysis/references/terminology.md +85 -0
- package/skills/payment/payment-analysis/references/visualization.md +47 -0
- package/skills/backend/pyspark/references/parity-testing.md +0 -83
- package/skills/backend/pyspark/references/production-validation.md +0 -166
- /package/skills/{backend → data-science}/airflow-dag-develop/LICENSE +0 -0
- /package/skills/{backend → data-science}/airflow-dag-develop/SKILL.md +0 -0
- /package/skills/{backend → data-science}/pyspark/LICENSE +0 -0
- /package/skills/{backend → data-science}/pyspark/assets/templates/etl.py +0 -0
- /package/skills/{backend → data-science}/pyspark/assets/templates/utils/__init__.py +0 -0
- /package/skills/{backend → data-science}/pyspark/assets/templates/utils/hudi_metadata.py +0 -0
- /package/skills/{backend → data-science}/pyspark/references/etl-contract.md +0 -0
- /package/skills/{backend → data-science}/pyspark/references/velocity-feature-calculation.md +0 -0
- /package/skills/{backend → data-science}/pyspark/scripts/spark_eventlog_summary.py +0 -0
|
@@ -1,83 +0,0 @@
|
|
|
1
|
-
# Parity Testing
|
|
2
|
-
|
|
3
|
-
Use this playbook for semantics-preserving Spark rewrites.
|
|
4
|
-
|
|
5
|
-
## Keep The Old Path As A Bounded Oracle
|
|
6
|
-
|
|
7
|
-
Retain the previous implementation only in tests while parity is being proven.
|
|
8
|
-
Use small fixtures that are dense enough to expose the old complexity but
|
|
9
|
-
bounded enough to run reliably.
|
|
10
|
-
|
|
11
|
-
## Compare Full Rows In Both Directions
|
|
12
|
-
|
|
13
|
-
```python
|
|
14
|
-
expected = old_build(source).select(*OUTPUT_COLUMNS)
|
|
15
|
-
actual = new_build(source).select(*OUTPUT_COLUMNS)
|
|
16
|
-
|
|
17
|
-
assert expected.exceptAll(actual).count() == 0
|
|
18
|
-
assert actual.exceptAll(expected).count() == 0
|
|
19
|
-
```
|
|
20
|
-
|
|
21
|
-
Counts alone cannot detect wrong values, swapped identities, or duplicate rows.
|
|
22
|
-
Use `exceptAll`, not `subtract`, when duplicate multiplicity matters.
|
|
23
|
-
|
|
24
|
-
## Add Literal Boundary Assertions
|
|
25
|
-
|
|
26
|
-
Parity can be self-confirming if oracle and candidate share the same mistake.
|
|
27
|
-
Assert critical rules directly:
|
|
28
|
-
|
|
29
|
-
```python
|
|
30
|
-
by_id = {row.event_id: row for row in actual.collect()}
|
|
31
|
-
assert by_id["at-24h"].count_24h == 2 # inclusive lower boundary
|
|
32
|
-
assert by_id["after-24h"].count_24h == 1 # one unit outside
|
|
33
|
-
assert by_id["same-time-b"].count_1h == 2 # stable tie ordering
|
|
34
|
-
```
|
|
35
|
-
|
|
36
|
-
Keep collected fixtures synthetic and non-sensitive.
|
|
37
|
-
|
|
38
|
-
## Minimum Fixture Matrix
|
|
39
|
-
|
|
40
|
-
Cover:
|
|
41
|
-
|
|
42
|
-
- empty and single-event groups;
|
|
43
|
-
- exact lower/upper boundaries and one unit beyond;
|
|
44
|
-
- same-timestamp events with every tie-breaker;
|
|
45
|
-
- null key, event time, amount, FX, and counterparty;
|
|
46
|
-
- both directions and relevant statuses;
|
|
47
|
-
- repeated and distinct counterparties;
|
|
48
|
-
- state open, partial close, full close, eviction, and re-entry;
|
|
49
|
-
- floating-point add/remove residue;
|
|
50
|
-
- one dense hot key;
|
|
51
|
-
- duplicate and missing output identities.
|
|
52
|
-
|
|
53
|
-
## Test The Physical Shape
|
|
54
|
-
|
|
55
|
-
Assert the expensive relation is absent without snapshotting the whole plan:
|
|
56
|
-
|
|
57
|
-
```python
|
|
58
|
-
plan = actual._jdf.queryExecution().executedPlan().toString().lower()
|
|
59
|
-
assert not ("join" in plan and "event_time" in plan and "interval" in plan)
|
|
60
|
-
```
|
|
61
|
-
|
|
62
|
-
Prefer a targeted helper that counts known pair-join nodes. Also assert the
|
|
63
|
-
expected grouped-sweep or built-in operator is present when that guards the
|
|
64
|
-
optimization.
|
|
65
|
-
|
|
66
|
-
## Mutation-Prove Important Tests
|
|
67
|
-
|
|
68
|
-
Temporarily introduce representative faults and confirm tests fail:
|
|
69
|
-
|
|
70
|
-
- change `< cutoff` to `<= cutoff`;
|
|
71
|
-
- remove a tie-breaker;
|
|
72
|
-
- use ordinary rolling addition/subtraction;
|
|
73
|
-
- reintroduce the pair join;
|
|
74
|
-
- omit a downstream-required column.
|
|
75
|
-
|
|
76
|
-
Record red/green evidence for high-risk money, date, and identity semantics.
|
|
77
|
-
|
|
78
|
-
## Match Test Cost To Ownership
|
|
79
|
-
|
|
80
|
-
Behavioral Spark tests belong on files that own formulas, source scope,
|
|
81
|
-
history, or state transitions. Unrelated upstream changes may use a focused
|
|
82
|
-
required-column check instead of rerunning an expensive full-history suite.
|
|
83
|
-
Do not demote behavioral coverage for a semantic change.
|
|
@@ -1,166 +0,0 @@
|
|
|
1
|
-
# Production Validation
|
|
2
|
-
|
|
3
|
-
Use this playbook before production-scale profiling, canaries, backfills, or
|
|
4
|
-
writes.
|
|
5
|
-
|
|
6
|
-
## Fail-Closed Admission
|
|
7
|
-
|
|
8
|
-
Do not launch until all applicable facts are recorded and valid:
|
|
9
|
-
|
|
10
|
-
- exact commit/module hash;
|
|
11
|
-
- production Spark, Python, Java, and table-format runtimes;
|
|
12
|
-
- source snapshot or date boundary;
|
|
13
|
-
- scheduler pause/ownership state;
|
|
14
|
-
- zero competing warehouse work;
|
|
15
|
-
- explicit validate-only or write-enabled mode;
|
|
16
|
-
- durable writable log and event-log destination;
|
|
17
|
-
- restart/recovery plan for writes.
|
|
18
|
-
|
|
19
|
-
Classify existing processes and working-tree changes before touching them. Do
|
|
20
|
-
not kill or overwrite another operator's work.
|
|
21
|
-
|
|
22
|
-
## Run A Validate-Only Canary Before The Writer
|
|
23
|
-
|
|
24
|
-
Implement the guard in the Spark job, not only in a wrapper. Keep each stage
|
|
25
|
-
literal:
|
|
26
|
-
|
|
27
|
-
```python
|
|
28
|
-
import json
|
|
29
|
-
|
|
30
|
-
from pyspark.sql import functions as F
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
def validate_candidate(candidate):
|
|
34
|
-
# first() is a Spark action: it executes the lazy transform plan.
|
|
35
|
-
metrics = candidate.agg(
|
|
36
|
-
F.count("*").alias("row_count"),
|
|
37
|
-
F.countDistinct("event_id").alias("distinct_event_count"),
|
|
38
|
-
F.sum(
|
|
39
|
-
F.when(F.col("event_id").isNull(), 1).otherwise(0)
|
|
40
|
-
).alias("null_event_id_count"),
|
|
41
|
-
F.sum(
|
|
42
|
-
F.when(F.col("event_count_1h") < 0, 1).otherwise(0)
|
|
43
|
-
).alias("invalid_event_count_1h"),
|
|
44
|
-
F.count("amount_usd_1h").alias("non_null_amount_usd_1h_count"),
|
|
45
|
-
).first().asDict()
|
|
46
|
-
|
|
47
|
-
blockers = []
|
|
48
|
-
if metrics["row_count"] != metrics["distinct_event_count"]:
|
|
49
|
-
blockers.append("event_id is not unique")
|
|
50
|
-
if metrics["null_event_id_count"]:
|
|
51
|
-
blockers.append("event_id contains nulls")
|
|
52
|
-
if metrics["invalid_event_count_1h"]:
|
|
53
|
-
blockers.append("event_count_1h contains negative values")
|
|
54
|
-
if blockers:
|
|
55
|
-
raise RuntimeError(f"validation blocked: {blockers}")
|
|
56
|
-
return metrics
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
def run(validate_only, source_snapshot):
|
|
60
|
-
# Bind every action to the same immutable source version or date boundary.
|
|
61
|
-
source = extract(source_snapshot)
|
|
62
|
-
candidate = transform(source) # build a lazy DataFrame plan
|
|
63
|
-
|
|
64
|
-
metrics = validate_candidate(candidate) # execute validation aggregates
|
|
65
|
-
print(json.dumps(metrics, sort_keys=True))
|
|
66
|
-
|
|
67
|
-
if validate_only: # canary stops before any writer call
|
|
68
|
-
print("validate-only: skipping publication")
|
|
69
|
-
return
|
|
70
|
-
|
|
71
|
-
publish(candidate) # execute the writer from the same source
|
|
72
|
-
reconcile_published_scope(metrics) # independently read back aggregates
|
|
73
|
-
```
|
|
74
|
-
|
|
75
|
-
Validation and publication are separate Spark actions, so Spark may recompute
|
|
76
|
-
the candidate for publication. Pin every source read to the same immutable
|
|
77
|
-
snapshot, table version, or exact date boundary. If the source cannot be pinned,
|
|
78
|
-
the validation result does not prove what the writer will publish; stop and
|
|
79
|
-
redesign the run.
|
|
80
|
-
|
|
81
|
-
Adapt identity and blocker checks to the table's required behavior. Reference
|
|
82
|
-
every derived output in at least one validation aggregate; otherwise Catalyst
|
|
83
|
-
may prune an unreferenced calculation and the canary will not exercise the full
|
|
84
|
-
candidate. Validation must use aggregate evidence and fail before publication.
|
|
85
|
-
The canary exercises the production-shaped extract, transform, and validation
|
|
86
|
-
path without calling the writer.
|
|
87
|
-
|
|
88
|
-
## Avoid Explicit Cache And Persistence
|
|
89
|
-
|
|
90
|
-
Do not add `.cache()` or `.persist()` as a performance fix. They add storage,
|
|
91
|
-
serialization, eviction, and cleanup concerns without improving the logical
|
|
92
|
-
work. Prefer a better transformation shape and a pinned source snapshot.
|
|
93
|
-
|
|
94
|
-
Treat existing persistence as separate operational behavior: do not expand it
|
|
95
|
-
while optimizing the job, and remove it when repository tests and same-snapshot
|
|
96
|
-
evidence show it is unnecessary.
|
|
97
|
-
|
|
98
|
-
## Launch Through The Production Entry Point
|
|
99
|
-
|
|
100
|
-
Use the same `spark-submit` path and driver/executor runtimes as production:
|
|
101
|
-
|
|
102
|
-
```bash
|
|
103
|
-
export PYSPARK_DRIVER_PYTHON="$DRIVER_PYTHON"
|
|
104
|
-
export PYSPARK_PYTHON="$EXECUTOR_PYTHON"
|
|
105
|
-
|
|
106
|
-
spark-submit \
|
|
107
|
-
--name "$CANARY_NAME" \
|
|
108
|
-
--conf spark.pyspark.driver.python="$DRIVER_PYTHON" \
|
|
109
|
-
--conf spark.pyspark.python="$EXECUTOR_PYTHON" \
|
|
110
|
-
--conf spark.eventLog.enabled=true \
|
|
111
|
-
job.py --validate-only
|
|
112
|
-
```
|
|
113
|
-
|
|
114
|
-
Do not replace deployed scheduler files with an unmerged staged module. Stage
|
|
115
|
-
an isolated artifact or deploy merged code through the normal path.
|
|
116
|
-
|
|
117
|
-
## Prove The Canary Did Not Write
|
|
118
|
-
|
|
119
|
-
Require all three forms of evidence:
|
|
120
|
-
|
|
121
|
-
1. Validation completed with zero blockers.
|
|
122
|
-
2. The terminal log says publication was skipped.
|
|
123
|
-
3. The event log reports zero Spark output records and bytes.
|
|
124
|
-
|
|
125
|
-
Parse retained logs with upstream failures visible:
|
|
126
|
-
|
|
127
|
-
```bash
|
|
128
|
-
set -o pipefail
|
|
129
|
-
hdfs dfs -cat "$EVENT_LOG" | \
|
|
130
|
-
python {baseDir}/scripts/spark_eventlog_summary.py > canary-summary.json
|
|
131
|
-
```
|
|
132
|
-
|
|
133
|
-
Record application ID, shell/event-log wall, stage/task metrics, plan nodes,
|
|
134
|
-
shuffle, spill, output, code hash, and snapshot denominator.
|
|
135
|
-
|
|
136
|
-
## Run Write Verification Separately
|
|
137
|
-
|
|
138
|
-
Run write-enabled verification only after merge, deployment, successful
|
|
139
|
-
validate-only evidence, and explicit authorization. Re-check admission just
|
|
140
|
-
before launch.
|
|
141
|
-
|
|
142
|
-
Verify row count, distinct identity count, date/partition coverage, commit
|
|
143
|
-
success, schema, consumer compatibility, validation blockers, write mode,
|
|
144
|
-
single-writer constraints, and downstream readability.
|
|
145
|
-
|
|
146
|
-
Report compute/validation, publication, and total time separately. Never
|
|
147
|
-
present validate-only time as end-to-end write performance.
|
|
148
|
-
|
|
149
|
-
If a write run is cancelled, prove from logs/source that publication was not
|
|
150
|
-
reached. YARN `KILLED` alone does not prove no commit occurred.
|
|
151
|
-
|
|
152
|
-
## Keep Evidence Aggregate-Only
|
|
153
|
-
|
|
154
|
-
Safe evidence includes counts, null rates, distinct counts, min/max dates,
|
|
155
|
-
hashes, stage metrics, plan-node counts, and commit IDs. Do not print
|
|
156
|
-
credentials, entity rows, account identifiers, or per-entity values.
|
|
157
|
-
|
|
158
|
-
## Restore Ownership
|
|
159
|
-
|
|
160
|
-
After the run:
|
|
161
|
-
|
|
162
|
-
1. Confirm application and wrapper processes are terminal.
|
|
163
|
-
2. Remove agent-owned watchers and temporary processes.
|
|
164
|
-
3. Restore scheduler pause state or ownership.
|
|
165
|
-
4. Separate unrelated downstream failures from the optimized job.
|
|
166
|
-
5. Publish evidence with limitations and unmeasured surfaces.
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|