@1aboveio/skills 0.15.0 → 0.17.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +8 -2
- package/package.json +1 -1
- package/runtime/skills/distribution/generated/recipes.json +178 -34
- package/runtime/skills/distribution/scripts/bundles.mjs +11 -3
- package/runtime/skills/engineering/engineering-runtime/scripts/workflow-policy.mjs +1 -1
- package/skills/data-science/pyspark/SKILL.md +126 -0
- package/skills/{backend → data-science}/pyspark/assets/templates/etl.py +51 -0
- package/skills/{backend → data-science}/pyspark/references/diagnosis-and-profiling.md +38 -14
- package/skills/{backend → data-science}/pyspark/references/etl-contract.md +19 -0
- package/skills/data-science/pyspark/references/production-validation.md +131 -0
- package/skills/data-science/pyspark/references/reconciliation.md +38 -0
- package/skills/{backend → data-science}/pyspark/references/transformation-design.md +30 -2
- package/skills/engineering/engineering-runtime/coherence/workflow.json +14 -14
- package/skills/engineering/engineering-runtime/scripts/workflow-policy.mjs +1 -1
- package/skills/engineering/resolve-issues/generated/workflow-repair-policy.json +11 -11
- package/skills/engineering/resolve-issues/scripts/run-state.mjs +1 -1
- package/skills/payment/fraud-analysis/LICENSE +3 -0
- package/skills/payment/fraud-analysis/SKILL.md +113 -0
- package/skills/payment/fraud-analysis/evals/evals.json +40 -0
- package/skills/payment/fraud-analysis/references/archetypes/authorized-payment-scam.md +41 -0
- package/skills/payment/fraud-analysis/references/archetypes/first-party-fraud.md +44 -0
- package/skills/payment/fraud-analysis/references/archetypes/third-party-fraud.md +27 -0
- package/skills/payment/fraud-analysis/references/contexts/bank-transfer.md +24 -0
- package/skills/payment/fraud-analysis/references/contexts/card-payment.md +30 -0
- package/skills/payment/fraud-analysis/references/contexts/payment-collection.md +20 -0
- package/skills/payment/fraud-analysis/references/contexts/payout.md +20 -0
- package/skills/payment/fraud-analysis/references/feature-engineering.md +158 -0
- package/skills/payment/fraud-analysis/references/mechanisms/account-takeover.md +36 -0
- package/skills/payment/fraud-analysis/references/report-rationale.md +45 -0
- package/skills/payment/fraud-analysis/references/report-template.md +190 -0
- package/skills/payment/fraud-analysis/references/review-checklist.md +175 -0
- package/skills/payment/fraud-analysis/references/taxonomy.md +79 -0
- package/skills/payment/fraud-analysis/references/terminology.md +108 -0
- package/skills/payment/fraud-analysis/references/workflow.md +175 -0
- package/skills/payment/payment-analysis/LICENSE +3 -0
- package/skills/payment/payment-analysis/SKILL.md +127 -0
- package/skills/payment/payment-analysis/references/auth-rate-actions.md +30 -0
- package/skills/payment/payment-analysis/references/chargebacks.md +88 -0
- package/skills/payment/payment-analysis/references/event-layers.md +79 -0
- package/skills/payment/payment-analysis/references/fx.md +59 -0
- package/skills/payment/payment-analysis/references/journey.md +78 -0
- package/skills/payment/payment-analysis/references/metrics.md +62 -0
- package/skills/payment/payment-analysis/references/report-template.md +98 -0
- package/skills/payment/payment-analysis/references/terminology.md +85 -0
- package/skills/payment/payment-analysis/references/visualization.md +47 -0
- package/skills/backend/pyspark/SKILL.md +0 -116
- package/skills/backend/pyspark/references/parity-testing.md +0 -83
- package/skills/backend/pyspark/references/production-validation.md +0 -166
- /package/skills/{backend → data-science}/airflow-dag-develop/LICENSE +0 -0
- /package/skills/{backend → data-science}/airflow-dag-develop/SKILL.md +0 -0
- /package/skills/{backend → data-science}/pyspark/LICENSE +0 -0
- /package/skills/{backend → data-science}/pyspark/assets/templates/utils/__init__.py +0 -0
- /package/skills/{backend → data-science}/pyspark/assets/templates/utils/hudi_metadata.py +0 -0
- /package/skills/{backend → data-science}/pyspark/references/velocity-feature-calculation.md +0 -0
- /package/skills/{backend → data-science}/pyspark/scripts/spark_eventlog_summary.py +0 -0
|
@@ -0,0 +1,62 @@
|
|
|
1
|
+
# Metrics and denominators
|
|
2
|
+
|
|
3
|
+
Compute with explicit denominators. Prefer count and amount side by side.
|
|
4
|
+
Amounts are USD per [fx.md](fx.md) before any sum or share.
|
|
5
|
+
|
|
6
|
+
## Authorization
|
|
7
|
+
|
|
8
|
+
| KPI | Formula | Notes |
|
|
9
|
+
|---|---|---|
|
|
10
|
+
| Attempts | count of auth rows in window | |
|
|
11
|
+
| Approvals | count where response = approval | |
|
|
12
|
+
| Auth rate | approvals / attempts | Not capture rate |
|
|
13
|
+
| Approved amount | sum(amount \| approval) | USD after [fx.md](fx.md) |
|
|
14
|
+
| Declines | attempts − approvals | Or hard-only if stated |
|
|
15
|
+
| Decline rate | declines / attempts | |
|
|
16
|
+
| Reason share (of declines) | reason_count / declines | Default for mix tables |
|
|
17
|
+
| Reason incidence (of attempts) | reason_count / attempts | Label clearly if used |
|
|
18
|
+
| Country attempt share | country_attempts / attempts | |
|
|
19
|
+
| Country amount share | country_amount / total_amount | Often diverges from attempt share |
|
|
20
|
+
| Country auth rate | country_approvals / country_attempts | Require min-n (e.g. ≥30) for rankings |
|
|
21
|
+
|
|
22
|
+
Trend grains: daily (`Timestamp` date) and weekly (state week rule, e.g. W-SUN).
|
|
23
|
+
|
|
24
|
+
Card product: `Payment method` × `Card type` when both exist.
|
|
25
|
+
|
|
26
|
+
## Settlement
|
|
27
|
+
|
|
28
|
+
| KPI | Formula | Notes |
|
|
29
|
+
|---|---|---|
|
|
30
|
+
| Sale count / amount | Settlement type = Sale | |
|
|
31
|
+
| Refund count / amount | Settlement type = Refund | Show sign or absolute + label |
|
|
32
|
+
| Net settled amount | sales + refunds (refunds negative) | |
|
|
33
|
+
| Refund rate (count) | refunds / sales | |
|
|
34
|
+
| Refund rate (amount) | abs(refund amount) / sale amount | |
|
|
35
|
+
| Status mix | by Settlement status | |
|
|
36
|
+
|
|
37
|
+
Do not equalize settled sales to approved auth amount without a timing caveat.
|
|
38
|
+
|
|
39
|
+
## Chargebacks
|
|
40
|
+
|
|
41
|
+
Monthly **拒付率** by brand is [chargebacks.md](chargebacks.md): Visa
|
|
42
|
+
same-month sales, Mastercard previous-month sales, skip empty brand-months,
|
|
43
|
+
no blended Visa+Mastercard ratio. The table below is mix / match, not that
|
|
44
|
+
programme rate.
|
|
45
|
+
|
|
46
|
+
| KPI | Formula | Notes |
|
|
47
|
+
|---|---|---|
|
|
48
|
+
| CB count / dispute amount | as filed | Amounts USD after [fx.md](fx.md); use abs for magnitude if signed |
|
|
49
|
+
| Reason mix | by Reason category / code | Fraud ≠ all CBs |
|
|
50
|
+
| Match rate to auths | matched / CB rows | State join key |
|
|
51
|
+
| Visa monthly 拒付率 | CB_Visa(M) / Sales_Visa(M) | Received month; not VAMP unless TC40 in file |
|
|
52
|
+
| Mastercard monthly 拒付率 | CB_MC(M) / Sales_MC(M−1) | Skip M if M−1 sales missing |
|
|
53
|
+
|
|
54
|
+
Lag: chargeback date is later than original auth. Programme rates use
|
|
55
|
+
**received** month, not origination vintage.
|
|
56
|
+
|
|
57
|
+
## Ranking and “Other”
|
|
58
|
+
|
|
59
|
+
- Rank countries by attempts unless the user asks for amount.
|
|
60
|
+
- Roll the long tail into `Other` after top 10–15; show Other share.
|
|
61
|
+
- For decline-by-country matrices, use top countries by attempts and top reasons
|
|
62
|
+
globally; collapse remaining reasons to `Other declines`.
|
|
@@ -0,0 +1,98 @@
|
|
|
1
|
+
# Report template
|
|
2
|
+
|
|
3
|
+
Save:
|
|
4
|
+
|
|
5
|
+
```text
|
|
6
|
+
reports/payment-analysis/<scope>_overview.md
|
|
7
|
+
reports/payment-analysis/<scope>_overview.zh.md
|
|
8
|
+
reports/payment-analysis/<scope>_overview.zh.html # preferred when charts
|
|
9
|
+
```
|
|
10
|
+
|
|
11
|
+
Follow [visualization.md](visualization.md), [journey.md](journey.md),
|
|
12
|
+
[terminology.md](terminology.md), and [fx.md](fx.md) for amount currency. Omit
|
|
13
|
+
topic sections whose layer is missing.
|
|
14
|
+
No Markdown bold in report body. After writing, grep banned strings in
|
|
15
|
+
terminology.md.
|
|
16
|
+
|
|
17
|
+
## Required structure
|
|
18
|
+
|
|
19
|
+
English Markdown headings on the left; Chinese HTML / `.zh.md` on the right.
|
|
20
|
+
|
|
21
|
+
```markdown
|
|
22
|
+
# Payment Overview — <window> # 支付业务概览 — <window>
|
|
23
|
+
|
|
24
|
+
## 1. Executive summary # 1. 执行摘要
|
|
25
|
+
## 2. Top statistics # 2. 核心指标
|
|
26
|
+
## 3. Payment journey (Sankey) # 3. 支付旅程(Sankey)
|
|
27
|
+
## 4. Topics # 4. 专题
|
|
28
|
+
### 4.1 Overall transaction volume # 4.1 整体交易量
|
|
29
|
+
### 4.2 Week-on-week trend # 4.2 周趋势
|
|
30
|
+
### 4.3 Authorization rate # 4.3 授权成功率
|
|
31
|
+
### 4.4 Contribution by BIN country # 4.4 BIN 国家贡献
|
|
32
|
+
### 4.5 Decline by reason code # 4.5 授权拒绝原因贡献
|
|
33
|
+
### 4.6 Decline by BIN country # 4.6 各国授权拒绝结构
|
|
34
|
+
### 4.7 Settlements (when present) # 4.7 清算
|
|
35
|
+
### 4.8 Chargebacks (optional) # 4.8 拒付
|
|
36
|
+
## 5. Operational impact # 5. 运营影响
|
|
37
|
+
## 6. Appendix # 6. 附录
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
HTML uses the Chinese headings. KPI labels, legends, and Sankey node names use
|
|
41
|
+
the glossary in terminology.md (`授权尝试`, `授权成功`, `授权拒绝`,
|
|
42
|
+
`发卡行国家`, `差错`, `产品服务质量`, `Do not honor`).
|
|
43
|
+
|
|
44
|
+
## Chapter requirements
|
|
45
|
+
|
|
46
|
+
### 1. Executive summary / 执行摘要
|
|
47
|
+
|
|
48
|
+
≤5 numbered findings: volume, trend, auth rate, country concentration, main
|
|
49
|
+
decline or settlement/CB driver.
|
|
50
|
+
|
|
51
|
+
### 2. Top statistics / 核心指标
|
|
52
|
+
|
|
53
|
+
KPI cards / table from auth + settlement: attempts, approvals, auth rate,
|
|
54
|
+
attempt amount, approved amount, sale count/amount, refund count/amount, net
|
|
55
|
+
settled, distinct PAN/BIN/countries, portfolio constants (MID, MCC, provider).
|
|
56
|
+
|
|
57
|
+
Chinese labels: 授权尝试, 授权成功, 授权成功率, 授权成功金额, 清算销售, 退款,
|
|
58
|
+
净清算, 发卡行国家.
|
|
59
|
+
|
|
60
|
+
### 3. Payment journey (Sankey) / 支付旅程
|
|
61
|
+
|
|
62
|
+
Count Sankey per [journey.md](journey.md). Caption lists skipped checkpoints
|
|
63
|
+
(e.g. risk decision / 3DS when authentication fields are empty).
|
|
64
|
+
|
|
65
|
+
Collapsed nodes when 3DS/risk are absent:
|
|
66
|
+
|
|
67
|
+
`开始 → 授权 → 授权成功 / 授权拒绝 → 已清算销售 / 授权成功未清算`
|
|
68
|
+
|
|
69
|
+
### 4. Topics / 专题
|
|
70
|
+
|
|
71
|
+
Each topic: short lede + chart(s) + supporting table.
|
|
72
|
+
|
|
73
|
+
| Topic | Minimum viz |
|
|
74
|
+
|---|---|
|
|
75
|
+
| 4.1 Volume | count + amount bars (dual axis) |
|
|
76
|
+
| 4.2 WoW trend | weekly count + amount bars; auth-rate line |
|
|
77
|
+
| 4.3 Auth rate | rate line by day and/or segment |
|
|
78
|
+
| 4.4 BIN country | count + amount bars dual axis; rate line |
|
|
79
|
+
| 4.5 Decline reason | count bars |
|
|
80
|
+
| 4.6 Decline by country | stacked count bars |
|
|
81
|
+
| 4.7 Settlements | sale/refund amount bars; country contribution dual axis |
|
|
82
|
+
| 4.8 Chargebacks | reason-category bars; Visa / Mastercard monthly 拒付率 line per [chargebacks.md](chargebacks.md) (note under the rate chart; skip empty brand-months) |
|
|
83
|
+
|
|
84
|
+
Use both authorization and settlement data when both files are present.
|
|
85
|
+
|
|
86
|
+
### 5. Operational impact / 运营影响
|
|
87
|
+
|
|
88
|
+
Numbered operational takeaways. Do not title this 运营含义. Do not invent
|
|
89
|
+
supervised fraud rules. Choose from
|
|
90
|
+
[auth-rate-actions.md](auth-rate-actions.md) only when the extract supports
|
|
91
|
+
the trigger (SCA → 3DS; high 疑似欺诈 country → 风控拦截 or 3DS; 软性拒绝 →
|
|
92
|
+
retry).
|
|
93
|
+
|
|
94
|
+
### 6. Appendix / 附录
|
|
95
|
+
|
|
96
|
+
Definitions (from terminology.md) and caveats (FX per [fx.md](fx.md),
|
|
97
|
+
issuer country ≠ ship-to, auth rate ≠ capture, journey skips,
|
|
98
|
+
chargeback rates are network-style estimates).
|
|
@@ -0,0 +1,85 @@
|
|
|
1
|
+
# Terminology
|
|
2
|
+
|
|
3
|
+
Use these terms consistently in payment-analysis reports. Read this file before
|
|
4
|
+
writing English or Chinese copy. After writing, run the proofread step in
|
|
5
|
+
`SKILL.md`.
|
|
6
|
+
|
|
7
|
+
## Canonical glossary
|
|
8
|
+
|
|
9
|
+
| English | Chinese | Do not write |
|
|
10
|
+
|---|---|---|
|
|
11
|
+
| Attempt | 授权尝试 | 单独的「尝试」;交易;conversion |
|
|
12
|
+
| Approval | 授权成功 | 批准;成功(易与清算混淆) |
|
|
13
|
+
| Auth rate | 授权成功率 | 转化率;单独的「成功率」(首次出现时) |
|
|
14
|
+
| Decline | 授权拒绝 | 全部叫欺诈 |
|
|
15
|
+
| Hard / soft / reject | 硬性拒绝 / 软性拒绝 / 风控拦截 | hard 拦截(指拒绝类型时) |
|
|
16
|
+
| Reason / reason code | 拒绝原因 / 原因码 | |
|
|
17
|
+
| Insufficient funds | 余额不足 | 余额 |
|
|
18
|
+
| Suspected fraud | 疑似欺诈 | 已确认欺诈;正文里写 D102(表内 Pazien 代码可保留) |
|
|
19
|
+
| Invalid expiration date | 有效期无效 | |
|
|
20
|
+
| Security code invalid | 安全码无效 | |
|
|
21
|
+
| Policy reasons | 政策原因 | |
|
|
22
|
+
| Do not honor | Do not honor | 拒绝承兑 |
|
|
23
|
+
| BIN / IIN | BIN / IIN | |
|
|
24
|
+
| BIN country | 发卡行国家 / BIN 国家 | 客户国家;收货国;发卡国 |
|
|
25
|
+
| Settlement | 清算 | |
|
|
26
|
+
| Sale / Refund | 销售 / 退款 | 把销售金额写成清算金额 |
|
|
27
|
+
| Chargeback / dispute | 拒付 / 争议 | 与清算混用 |
|
|
28
|
+
| Clerical | 差错 | 文书;文书错误 |
|
|
29
|
+
| Quality | 产品服务质量 | 质量;质量争议 |
|
|
30
|
+
| Incidence | 发生率(必须写出分母) | |
|
|
31
|
+
| GMV (loose) | 销售金额合计(说明是否含退款) | |
|
|
32
|
+
| Week-on-week trend | 周趋势 | 周环比(除非单元格是环比百分比) |
|
|
33
|
+
| Operational impact | 运营影响 | 运营含义 |
|
|
34
|
+
| Ticket size | 客单价 | |
|
|
35
|
+
| Debit / credit | 借记卡 / 贷记卡 | 借记;贷记 |
|
|
36
|
+
| Payment journey | 支付旅程 | 支付旅程 Sankey;Sankey 旅程(标题/结构行) |
|
|
37
|
+
| This analysis window | 当前分析时段内 | 本窗口 |
|
|
38
|
+
| SCA | SCA | 强客户认证(可在首次括注后只用 SCA) |
|
|
39
|
+
| Authentication / 3DS | 鉴权 / 认证 / 3DS | 把 3DS 叫转化 |
|
|
40
|
+
| Visa / Mastercard 拒付率 | Visa / Mastercard 拒付率 | VAMP ratio(无 TC40 时);ECM(从本抽取断言商户在项目中) |
|
|
41
|
+
| Amounts in USD | 金额均为 USD | 未做汇率折算(当全部已是 USD) |
|
|
42
|
+
| Count-heavy / amount-light | 授权尝试笔数多但金额小 | 走廊;笔数负担型走廊 |
|
|
43
|
+
|
|
44
|
+
Keep proper nouns in source form when useful: Visa, MasterCard, BIN, 3DS,
|
|
45
|
+
Sankey (in chart captions only), Order ID, MCC, Worldpay, Pazien, MID.
|
|
46
|
+
Do not use Pazien codes (e.g. D102) in prose; use 疑似欺诈 / suspected fraud.
|
|
47
|
+
Tables may still show the code column.
|
|
48
|
+
|
|
49
|
+
## Axis and UI exceptions
|
|
50
|
+
|
|
51
|
+
- Rate axis may say `成功率 %` if the legend or lede already says 授权成功率.
|
|
52
|
+
Amount axes still use the full metric plus currency
|
|
53
|
+
([visualization.md](visualization.md)).
|
|
54
|
+
- Week-over-week *percentage* columns may say 笔数周环比 / 金额周环比.
|
|
55
|
+
- Product rankings that exclude tiny samples: write
|
|
56
|
+
`在授权尝试超过 50 笔的卡种里` (or the n actually used), never
|
|
57
|
+
`有规模的最弱卡种` or `弱卡种`.
|
|
58
|
+
- If every amount is USD, write `金额均为 USD`. Do not add `未做汇率折算`.
|
|
59
|
+
Only mention missing FX conversion when mixed currencies were actually summed
|
|
60
|
+
or dropped.
|
|
61
|
+
- Chart amount fields append the currency: `授权拒绝金额 USD`.
|
|
62
|
+
- HTML footer: data-source filenames only. No skill name, no Chart.js / ECharts.
|
|
63
|
+
- Do not call issuer country “customer country”.
|
|
64
|
+
- Do not call auth rate “conversion” unless contrasting checkout vs issuer
|
|
65
|
+
response, and then say so.
|
|
66
|
+
|
|
67
|
+
## Banned strings (grep after writing)
|
|
68
|
+
|
|
69
|
+
Chinese: `走廊`, `杠杆`, `弱转化`, `转化率`, `赋能`, `抓手`, `闭环`, `打法`,
|
|
70
|
+
`运营含义`, `本窗口`, `弱卡种`, `有规模的最弱`, `强客户认证`, `拒绝承兑`,
|
|
71
|
+
`文书错误`, `质量争议`, `发卡国`
|
|
72
|
+
|
|
73
|
+
English: `auth drag`, `conversion problem`, `count-heavy`, `amount-light`,
|
|
74
|
+
`Protect US amount`, `Unlike currencies`
|
|
75
|
+
|
|
76
|
+
If a hit is a legitimate exception (e.g. 周环比 as a table header for WoW %,
|
|
77
|
+
or D102 in a reason-code table), leave it and do not expand the ban.
|
|
78
|
+
|
|
79
|
+
## Proofread
|
|
80
|
+
|
|
81
|
+
1. Apply the glossary to headings, KPI labels, chart legends, Sankey skip
|
|
82
|
+
chips, and body copy.
|
|
83
|
+
2. Grep the report files for the banned strings.
|
|
84
|
+
3. Replace hits using this table. Do not change numbers, formulas, or file
|
|
85
|
+
paths.
|
|
@@ -0,0 +1,47 @@
|
|
|
1
|
+
# Visualization rules
|
|
2
|
+
|
|
3
|
+
Preferred deliverable: self-contained HTML with Chart.js (KPI charts) and
|
|
4
|
+
ECharts (Sankey). Markdown tables remain required beside charts.
|
|
5
|
+
|
|
6
|
+
## Mark encoding
|
|
7
|
+
|
|
8
|
+
| Metric kind | Chart mark | Notes |
|
|
9
|
+
|---|---|---|
|
|
10
|
+
| Count | Bar | Vertical or horizontal |
|
|
11
|
+
| Amount | Bar | Never encode amount as a line |
|
|
12
|
+
| Rate | Line | Percent on its own axis |
|
|
13
|
+
|
|
14
|
+
## Multi-metric charts
|
|
15
|
+
|
|
16
|
+
1. Count and amount in the same chart → both are **bars** on **different axes**
|
|
17
|
+
(left = count, right = amount).
|
|
18
|
+
2. Rate with count and/or amount → rate is a **line** on a dedicated % axis.
|
|
19
|
+
3. Rate alone → line chart (points allowed). No bars for a lone rate series.
|
|
20
|
+
4. Composition of one total → doughnut or stacked bars; not for time trends.
|
|
21
|
+
|
|
22
|
+
## Axis hygiene
|
|
23
|
+
|
|
24
|
+
- Label every axis with the full metric name. Amount axes/legends include
|
|
25
|
+
currency (`授权尝试金额 USD`, `授权拒绝金额 USD`, `清算销售金额 USD`,
|
|
26
|
+
`拒付金额 USD`). Do not use bare `笔数` / `金额` / `%`.
|
|
27
|
+
- Only one axis draws the main grid on dual-axis charts.
|
|
28
|
+
- Stacked bars = mix within a group (e.g. decline reasons by country), never rate.
|
|
29
|
+
- Sort categories by a stated order (usually attempts desc).
|
|
30
|
+
|
|
31
|
+
## Required charts (HTML)
|
|
32
|
+
|
|
33
|
+
1. Payment journey Sankey — see [journey.md](journey.md)
|
|
34
|
+
2. Volume trend (WoW / daily): count bar + amount bar (dual axis) + auth-rate line
|
|
35
|
+
3. Authorization rate by segment (card product and/or country): rate as line
|
|
36
|
+
4. BIN-country contribution: count bar + amount bar (dual axis)
|
|
37
|
+
5. Decline by reason: count bars (+ amount bars/dual axis if shown)
|
|
38
|
+
6. Decline mix by BIN country: stacked count bars
|
|
39
|
+
|
|
40
|
+
## Anti-patterns
|
|
41
|
+
|
|
42
|
+
- Line for amount or count
|
|
43
|
+
- Bar for a standalone rate series
|
|
44
|
+
- Count and amount on one axis without dual-axis scaling
|
|
45
|
+
- Mixing native EUR (etc.) with USD in one amount series; convert first ([fx.md](fx.md))
|
|
46
|
+
- Plotting a 0% chargeback-rate point for a brand-month with no CBs ([chargebacks.md](chargebacks.md))
|
|
47
|
+
- Inventing Sankey checkpoints not present in the data
|
|
@@ -1,116 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: pyspark
|
|
3
|
-
description: Implement, debug, review, and optimize Apache Spark or PySpark batch jobs, ETL pipelines, feature calculations, and Hudi/warehouse writes. Use whenever the task involves Spark DataFrames or Spark SQL, joins, windows, grouped state, UDFs, shuffle, skew, spill, slow stages, driver/executor OOM, Spark Connect, spark-submit, Hudi commits, backfills, bidirectional parity tests, validate-only canaries, or Spark event-log profiling — including when the user only mentions a warehouse job timeout, empty Hudi write, or performance regression on an EMR/Dataproc job.
|
|
4
|
-
---
|
|
5
|
-
|
|
6
|
-
# PySpark
|
|
7
|
-
|
|
8
|
-
Use this skill as the workflow and routing index. Read the referenced playbooks
|
|
9
|
-
that match the task before editing code or running a production-scale job.
|
|
10
|
-
|
|
11
|
-
## Workflow
|
|
12
|
-
|
|
13
|
-
```text
|
|
14
|
-
contract -> transform -> materialize -> test -> CI/CD
|
|
15
|
-
```
|
|
16
|
-
|
|
17
|
-
Follow all five steps in order.
|
|
18
|
-
|
|
19
|
-
### 1. Contract
|
|
20
|
-
|
|
21
|
-
- **Input:** declare all parameters, defaults, validation, source snapshot/date
|
|
22
|
-
scope, overlapping-run behavior, idempotency/single-writer rules, and exact
|
|
23
|
-
Python and PySpark versions.
|
|
24
|
-
- **Output:** declare plain Spark versus Hudi table type, output identity/schema,
|
|
25
|
-
partition columns, record key, deterministic order or precombine field, and
|
|
26
|
-
append/upsert/overwrite write mode.
|
|
27
|
-
- **Conventions:** declare native Spark versus Spark Connect and full rebuild
|
|
28
|
-
versus incremental processing, including the incremental watermark,
|
|
29
|
-
lookback/checkpoint, and restart boundary when applicable.
|
|
30
|
-
|
|
31
|
-
Read [ETL Contract](references/etl-contract.md). When the repository has no ETL
|
|
32
|
-
base module, adapt [`assets/templates/etl.py`](assets/templates/etl.py).
|
|
33
|
-
|
|
34
|
-
### 2. Transform
|
|
35
|
-
|
|
36
|
-
- Make `transform` accept DataFrames and return the complete candidate without
|
|
37
|
-
invoking validation or a writer.
|
|
38
|
-
- Estimate rows after joins/explodes/groups and measure maximum rows per key.
|
|
39
|
-
Fix algorithmic expansion and hot keys before tuning Spark configuration.
|
|
40
|
-
- Prefer built-in expressions and windows. Use exact bucketing or one sorted
|
|
41
|
-
grouped sweep only for state that built-ins cannot express cleanly.
|
|
42
|
-
- Preserve complete row identity, ordering, output columns, and required
|
|
43
|
-
behavior.
|
|
44
|
-
- Avoid explicit `.cache()` and `.persist()` by default.
|
|
45
|
-
|
|
46
|
-
Read [Transformation Design](references/transformation-design.md). For temporal
|
|
47
|
-
windows or FIFO paths, also read
|
|
48
|
-
[Velocity Feature Calculation](references/velocity-feature-calculation.md).
|
|
49
|
-
|
|
50
|
-
### 3. Materialize
|
|
51
|
-
|
|
52
|
-
- Treat materialization as the Spark action that executes the lazy candidate,
|
|
53
|
-
not as caching.
|
|
54
|
-
- Use one aggregate validation action before every writer. Reference every
|
|
55
|
-
derived output so Catalyst cannot prune untested calculations.
|
|
56
|
-
- Check row count, distinct/null/duplicate identities, domain blockers, and
|
|
57
|
-
required date/partition scope. Fail closed on any blocker.
|
|
58
|
-
- Bind validation and publication to the same immutable source snapshot because
|
|
59
|
-
Spark may recompute the candidate.
|
|
60
|
-
- Keep the validate-only guard inside the job and before every writer. Do not
|
|
61
|
-
add `.count()`, `.cache()`, or `.persist()` to warm execution.
|
|
62
|
-
|
|
63
|
-
Read [Production Validation](references/production-validation.md).
|
|
64
|
-
|
|
65
|
-
### 4. Test
|
|
66
|
-
|
|
67
|
-
1. **Local test:** run pure state, tiny real-Spark, plan-shape, focused repository,
|
|
68
|
-
and bidirectional parity tests. Cover boundaries, ties, nulls, arithmetic,
|
|
69
|
-
identity, and precedence.
|
|
70
|
-
2. **Cluster canary test:** run the production entry point and source scope with
|
|
71
|
-
`--validate-only`. Require zero blockers, an explicit publication skip, zero
|
|
72
|
-
Spark output records/bytes, and readable driver/event logs.
|
|
73
|
-
3. **Profiling:** measure an isolated baseline, parse the canary event log, and
|
|
74
|
-
compare both runs on the same source scope and resource profile. Record
|
|
75
|
-
wall/stage/task time, cardinality, shuffle, spill, skew, output, and target
|
|
76
|
-
plan patterns.
|
|
77
|
-
|
|
78
|
-
```bash
|
|
79
|
-
spark-submit <production-options> job.py --validate-only
|
|
80
|
-
set -o pipefail
|
|
81
|
-
hdfs dfs -cat "$EVENT_LOG" | \
|
|
82
|
-
python {baseDir}/scripts/spark_eventlog_summary.py > canary-summary.json
|
|
83
|
-
```
|
|
84
|
-
|
|
85
|
-
Reject missing evidence, incomparable scope, unexplained regressions, or a
|
|
86
|
-
missed performance gate. Keep validate-only time separate from write time.
|
|
87
|
-
|
|
88
|
-
Read [Parity Testing](references/parity-testing.md),
|
|
89
|
-
[Diagnosis And Profiling](references/diagnosis-and-profiling.md), and
|
|
90
|
-
[Production Validation](references/production-validation.md).
|
|
91
|
-
|
|
92
|
-
### 5. CI/CD
|
|
93
|
-
|
|
94
|
-
1. Open a PR with required-behavior, test, canary, profiling, and risk evidence.
|
|
95
|
-
2. Run required CI and independent review; fix failures.
|
|
96
|
-
3. Enqueue through the configured merge queue and merge without bypassing gates.
|
|
97
|
-
4. Deploy only the merged commit through the normal deployment path.
|
|
98
|
-
5. Run write-enabled production validation only with explicit authorization;
|
|
99
|
-
reconcile counts, identity, schema, coverage, blockers, readability, write
|
|
100
|
-
mode, and commit evidence independently.
|
|
101
|
-
6. Restore scheduler/process ownership and report compute/validation,
|
|
102
|
-
publication, and end-to-end time separately.
|
|
103
|
-
|
|
104
|
-
Read [Production Validation](references/production-validation.md).
|
|
105
|
-
|
|
106
|
-
## Reference Index
|
|
107
|
-
|
|
108
|
-
| Need | Read or run |
|
|
109
|
-
|---|---|
|
|
110
|
-
| Define or review an ETL job contract | [ETL Contract](references/etl-contract.md) and `assets/templates/etl.py` |
|
|
111
|
-
| Diagnose or profile | [Diagnosis And Profiling](references/diagnosis-and-profiling.md) |
|
|
112
|
-
| Design transformations | [Transformation Design](references/transformation-design.md) |
|
|
113
|
-
| Optimize temporal/FIFO state | [Velocity Feature Calculation](references/velocity-feature-calculation.md) |
|
|
114
|
-
| Prove parity | [Parity Testing](references/parity-testing.md) |
|
|
115
|
-
| Run canaries or writes | [Production Validation](references/production-validation.md) |
|
|
116
|
-
| Summarize an event log | `scripts/spark_eventlog_summary.py` |
|
|
@@ -1,83 +0,0 @@
|
|
|
1
|
-
# Parity Testing
|
|
2
|
-
|
|
3
|
-
Use this playbook for semantics-preserving Spark rewrites.
|
|
4
|
-
|
|
5
|
-
## Keep The Old Path As A Bounded Oracle
|
|
6
|
-
|
|
7
|
-
Retain the previous implementation only in tests while parity is being proven.
|
|
8
|
-
Use small fixtures that are dense enough to expose the old complexity but
|
|
9
|
-
bounded enough to run reliably.
|
|
10
|
-
|
|
11
|
-
## Compare Full Rows In Both Directions
|
|
12
|
-
|
|
13
|
-
```python
|
|
14
|
-
expected = old_build(source).select(*OUTPUT_COLUMNS)
|
|
15
|
-
actual = new_build(source).select(*OUTPUT_COLUMNS)
|
|
16
|
-
|
|
17
|
-
assert expected.exceptAll(actual).count() == 0
|
|
18
|
-
assert actual.exceptAll(expected).count() == 0
|
|
19
|
-
```
|
|
20
|
-
|
|
21
|
-
Counts alone cannot detect wrong values, swapped identities, or duplicate rows.
|
|
22
|
-
Use `exceptAll`, not `subtract`, when duplicate multiplicity matters.
|
|
23
|
-
|
|
24
|
-
## Add Literal Boundary Assertions
|
|
25
|
-
|
|
26
|
-
Parity can be self-confirming if oracle and candidate share the same mistake.
|
|
27
|
-
Assert critical rules directly:
|
|
28
|
-
|
|
29
|
-
```python
|
|
30
|
-
by_id = {row.event_id: row for row in actual.collect()}
|
|
31
|
-
assert by_id["at-24h"].count_24h == 2 # inclusive lower boundary
|
|
32
|
-
assert by_id["after-24h"].count_24h == 1 # one unit outside
|
|
33
|
-
assert by_id["same-time-b"].count_1h == 2 # stable tie ordering
|
|
34
|
-
```
|
|
35
|
-
|
|
36
|
-
Keep collected fixtures synthetic and non-sensitive.
|
|
37
|
-
|
|
38
|
-
## Minimum Fixture Matrix
|
|
39
|
-
|
|
40
|
-
Cover:
|
|
41
|
-
|
|
42
|
-
- empty and single-event groups;
|
|
43
|
-
- exact lower/upper boundaries and one unit beyond;
|
|
44
|
-
- same-timestamp events with every tie-breaker;
|
|
45
|
-
- null key, event time, amount, FX, and counterparty;
|
|
46
|
-
- both directions and relevant statuses;
|
|
47
|
-
- repeated and distinct counterparties;
|
|
48
|
-
- state open, partial close, full close, eviction, and re-entry;
|
|
49
|
-
- floating-point add/remove residue;
|
|
50
|
-
- one dense hot key;
|
|
51
|
-
- duplicate and missing output identities.
|
|
52
|
-
|
|
53
|
-
## Test The Physical Shape
|
|
54
|
-
|
|
55
|
-
Assert the expensive relation is absent without snapshotting the whole plan:
|
|
56
|
-
|
|
57
|
-
```python
|
|
58
|
-
plan = actual._jdf.queryExecution().executedPlan().toString().lower()
|
|
59
|
-
assert not ("join" in plan and "event_time" in plan and "interval" in plan)
|
|
60
|
-
```
|
|
61
|
-
|
|
62
|
-
Prefer a targeted helper that counts known pair-join nodes. Also assert the
|
|
63
|
-
expected grouped-sweep or built-in operator is present when that guards the
|
|
64
|
-
optimization.
|
|
65
|
-
|
|
66
|
-
## Mutation-Prove Important Tests
|
|
67
|
-
|
|
68
|
-
Temporarily introduce representative faults and confirm tests fail:
|
|
69
|
-
|
|
70
|
-
- change `< cutoff` to `<= cutoff`;
|
|
71
|
-
- remove a tie-breaker;
|
|
72
|
-
- use ordinary rolling addition/subtraction;
|
|
73
|
-
- reintroduce the pair join;
|
|
74
|
-
- omit a downstream-required column.
|
|
75
|
-
|
|
76
|
-
Record red/green evidence for high-risk money, date, and identity semantics.
|
|
77
|
-
|
|
78
|
-
## Match Test Cost To Ownership
|
|
79
|
-
|
|
80
|
-
Behavioral Spark tests belong on files that own formulas, source scope,
|
|
81
|
-
history, or state transitions. Unrelated upstream changes may use a focused
|
|
82
|
-
required-column check instead of rerunning an expensive full-history suite.
|
|
83
|
-
Do not demote behavioral coverage for a semantic change.
|
|
@@ -1,166 +0,0 @@
|
|
|
1
|
-
# Production Validation
|
|
2
|
-
|
|
3
|
-
Use this playbook before production-scale profiling, canaries, backfills, or
|
|
4
|
-
writes.
|
|
5
|
-
|
|
6
|
-
## Fail-Closed Admission
|
|
7
|
-
|
|
8
|
-
Do not launch until all applicable facts are recorded and valid:
|
|
9
|
-
|
|
10
|
-
- exact commit/module hash;
|
|
11
|
-
- production Spark, Python, Java, and table-format runtimes;
|
|
12
|
-
- source snapshot or date boundary;
|
|
13
|
-
- scheduler pause/ownership state;
|
|
14
|
-
- zero competing warehouse work;
|
|
15
|
-
- explicit validate-only or write-enabled mode;
|
|
16
|
-
- durable writable log and event-log destination;
|
|
17
|
-
- restart/recovery plan for writes.
|
|
18
|
-
|
|
19
|
-
Classify existing processes and working-tree changes before touching them. Do
|
|
20
|
-
not kill or overwrite another operator's work.
|
|
21
|
-
|
|
22
|
-
## Run A Validate-Only Canary Before The Writer
|
|
23
|
-
|
|
24
|
-
Implement the guard in the Spark job, not only in a wrapper. Keep each stage
|
|
25
|
-
literal:
|
|
26
|
-
|
|
27
|
-
```python
|
|
28
|
-
import json
|
|
29
|
-
|
|
30
|
-
from pyspark.sql import functions as F
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
def validate_candidate(candidate):
|
|
34
|
-
# first() is a Spark action: it executes the lazy transform plan.
|
|
35
|
-
metrics = candidate.agg(
|
|
36
|
-
F.count("*").alias("row_count"),
|
|
37
|
-
F.countDistinct("event_id").alias("distinct_event_count"),
|
|
38
|
-
F.sum(
|
|
39
|
-
F.when(F.col("event_id").isNull(), 1).otherwise(0)
|
|
40
|
-
).alias("null_event_id_count"),
|
|
41
|
-
F.sum(
|
|
42
|
-
F.when(F.col("event_count_1h") < 0, 1).otherwise(0)
|
|
43
|
-
).alias("invalid_event_count_1h"),
|
|
44
|
-
F.count("amount_usd_1h").alias("non_null_amount_usd_1h_count"),
|
|
45
|
-
).first().asDict()
|
|
46
|
-
|
|
47
|
-
blockers = []
|
|
48
|
-
if metrics["row_count"] != metrics["distinct_event_count"]:
|
|
49
|
-
blockers.append("event_id is not unique")
|
|
50
|
-
if metrics["null_event_id_count"]:
|
|
51
|
-
blockers.append("event_id contains nulls")
|
|
52
|
-
if metrics["invalid_event_count_1h"]:
|
|
53
|
-
blockers.append("event_count_1h contains negative values")
|
|
54
|
-
if blockers:
|
|
55
|
-
raise RuntimeError(f"validation blocked: {blockers}")
|
|
56
|
-
return metrics
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
def run(validate_only, source_snapshot):
|
|
60
|
-
# Bind every action to the same immutable source version or date boundary.
|
|
61
|
-
source = extract(source_snapshot)
|
|
62
|
-
candidate = transform(source) # build a lazy DataFrame plan
|
|
63
|
-
|
|
64
|
-
metrics = validate_candidate(candidate) # execute validation aggregates
|
|
65
|
-
print(json.dumps(metrics, sort_keys=True))
|
|
66
|
-
|
|
67
|
-
if validate_only: # canary stops before any writer call
|
|
68
|
-
print("validate-only: skipping publication")
|
|
69
|
-
return
|
|
70
|
-
|
|
71
|
-
publish(candidate) # execute the writer from the same source
|
|
72
|
-
reconcile_published_scope(metrics) # independently read back aggregates
|
|
73
|
-
```
|
|
74
|
-
|
|
75
|
-
Validation and publication are separate Spark actions, so Spark may recompute
|
|
76
|
-
the candidate for publication. Pin every source read to the same immutable
|
|
77
|
-
snapshot, table version, or exact date boundary. If the source cannot be pinned,
|
|
78
|
-
the validation result does not prove what the writer will publish; stop and
|
|
79
|
-
redesign the run.
|
|
80
|
-
|
|
81
|
-
Adapt identity and blocker checks to the table's required behavior. Reference
|
|
82
|
-
every derived output in at least one validation aggregate; otherwise Catalyst
|
|
83
|
-
may prune an unreferenced calculation and the canary will not exercise the full
|
|
84
|
-
candidate. Validation must use aggregate evidence and fail before publication.
|
|
85
|
-
The canary exercises the production-shaped extract, transform, and validation
|
|
86
|
-
path without calling the writer.
|
|
87
|
-
|
|
88
|
-
## Avoid Explicit Cache And Persistence
|
|
89
|
-
|
|
90
|
-
Do not add `.cache()` or `.persist()` as a performance fix. They add storage,
|
|
91
|
-
serialization, eviction, and cleanup concerns without improving the logical
|
|
92
|
-
work. Prefer a better transformation shape and a pinned source snapshot.
|
|
93
|
-
|
|
94
|
-
Treat existing persistence as separate operational behavior: do not expand it
|
|
95
|
-
while optimizing the job, and remove it when repository tests and same-snapshot
|
|
96
|
-
evidence show it is unnecessary.
|
|
97
|
-
|
|
98
|
-
## Launch Through The Production Entry Point
|
|
99
|
-
|
|
100
|
-
Use the same `spark-submit` path and driver/executor runtimes as production:
|
|
101
|
-
|
|
102
|
-
```bash
|
|
103
|
-
export PYSPARK_DRIVER_PYTHON="$DRIVER_PYTHON"
|
|
104
|
-
export PYSPARK_PYTHON="$EXECUTOR_PYTHON"
|
|
105
|
-
|
|
106
|
-
spark-submit \
|
|
107
|
-
--name "$CANARY_NAME" \
|
|
108
|
-
--conf spark.pyspark.driver.python="$DRIVER_PYTHON" \
|
|
109
|
-
--conf spark.pyspark.python="$EXECUTOR_PYTHON" \
|
|
110
|
-
--conf spark.eventLog.enabled=true \
|
|
111
|
-
job.py --validate-only
|
|
112
|
-
```
|
|
113
|
-
|
|
114
|
-
Do not replace deployed scheduler files with an unmerged staged module. Stage
|
|
115
|
-
an isolated artifact or deploy merged code through the normal path.
|
|
116
|
-
|
|
117
|
-
## Prove The Canary Did Not Write
|
|
118
|
-
|
|
119
|
-
Require all three forms of evidence:
|
|
120
|
-
|
|
121
|
-
1. Validation completed with zero blockers.
|
|
122
|
-
2. The terminal log says publication was skipped.
|
|
123
|
-
3. The event log reports zero Spark output records and bytes.
|
|
124
|
-
|
|
125
|
-
Parse retained logs with upstream failures visible:
|
|
126
|
-
|
|
127
|
-
```bash
|
|
128
|
-
set -o pipefail
|
|
129
|
-
hdfs dfs -cat "$EVENT_LOG" | \
|
|
130
|
-
python {baseDir}/scripts/spark_eventlog_summary.py > canary-summary.json
|
|
131
|
-
```
|
|
132
|
-
|
|
133
|
-
Record application ID, shell/event-log wall, stage/task metrics, plan nodes,
|
|
134
|
-
shuffle, spill, output, code hash, and snapshot denominator.
|
|
135
|
-
|
|
136
|
-
## Run Write Verification Separately
|
|
137
|
-
|
|
138
|
-
Run write-enabled verification only after merge, deployment, successful
|
|
139
|
-
validate-only evidence, and explicit authorization. Re-check admission just
|
|
140
|
-
before launch.
|
|
141
|
-
|
|
142
|
-
Verify row count, distinct identity count, date/partition coverage, commit
|
|
143
|
-
success, schema, consumer compatibility, validation blockers, write mode,
|
|
144
|
-
single-writer constraints, and downstream readability.
|
|
145
|
-
|
|
146
|
-
Report compute/validation, publication, and total time separately. Never
|
|
147
|
-
present validate-only time as end-to-end write performance.
|
|
148
|
-
|
|
149
|
-
If a write run is cancelled, prove from logs/source that publication was not
|
|
150
|
-
reached. YARN `KILLED` alone does not prove no commit occurred.
|
|
151
|
-
|
|
152
|
-
## Keep Evidence Aggregate-Only
|
|
153
|
-
|
|
154
|
-
Safe evidence includes counts, null rates, distinct counts, min/max dates,
|
|
155
|
-
hashes, stage metrics, plan-node counts, and commit IDs. Do not print
|
|
156
|
-
credentials, entity rows, account identifiers, or per-entity values.
|
|
157
|
-
|
|
158
|
-
## Restore Ownership
|
|
159
|
-
|
|
160
|
-
After the run:
|
|
161
|
-
|
|
162
|
-
1. Confirm application and wrapper processes are terminal.
|
|
163
|
-
2. Remove agent-owned watchers and temporary processes.
|
|
164
|
-
3. Restore scheduler pause state or ownership.
|
|
165
|
-
4. Separate unrelated downstream failures from the optimized job.
|
|
166
|
-
5. Publish evidence with limitations and unmeasured surfaces.
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|