@1aboveio/skills 0.15.0 → 0.17.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (55) hide show
  1. package/README.md +8 -2
  2. package/package.json +1 -1
  3. package/runtime/skills/distribution/generated/recipes.json +178 -34
  4. package/runtime/skills/distribution/scripts/bundles.mjs +11 -3
  5. package/runtime/skills/engineering/engineering-runtime/scripts/workflow-policy.mjs +1 -1
  6. package/skills/data-science/pyspark/SKILL.md +126 -0
  7. package/skills/{backend → data-science}/pyspark/assets/templates/etl.py +51 -0
  8. package/skills/{backend → data-science}/pyspark/references/diagnosis-and-profiling.md +38 -14
  9. package/skills/{backend → data-science}/pyspark/references/etl-contract.md +19 -0
  10. package/skills/data-science/pyspark/references/production-validation.md +131 -0
  11. package/skills/data-science/pyspark/references/reconciliation.md +38 -0
  12. package/skills/{backend → data-science}/pyspark/references/transformation-design.md +30 -2
  13. package/skills/engineering/engineering-runtime/coherence/workflow.json +14 -14
  14. package/skills/engineering/engineering-runtime/scripts/workflow-policy.mjs +1 -1
  15. package/skills/engineering/resolve-issues/generated/workflow-repair-policy.json +11 -11
  16. package/skills/engineering/resolve-issues/scripts/run-state.mjs +1 -1
  17. package/skills/payment/fraud-analysis/LICENSE +3 -0
  18. package/skills/payment/fraud-analysis/SKILL.md +113 -0
  19. package/skills/payment/fraud-analysis/evals/evals.json +40 -0
  20. package/skills/payment/fraud-analysis/references/archetypes/authorized-payment-scam.md +41 -0
  21. package/skills/payment/fraud-analysis/references/archetypes/first-party-fraud.md +44 -0
  22. package/skills/payment/fraud-analysis/references/archetypes/third-party-fraud.md +27 -0
  23. package/skills/payment/fraud-analysis/references/contexts/bank-transfer.md +24 -0
  24. package/skills/payment/fraud-analysis/references/contexts/card-payment.md +30 -0
  25. package/skills/payment/fraud-analysis/references/contexts/payment-collection.md +20 -0
  26. package/skills/payment/fraud-analysis/references/contexts/payout.md +20 -0
  27. package/skills/payment/fraud-analysis/references/feature-engineering.md +158 -0
  28. package/skills/payment/fraud-analysis/references/mechanisms/account-takeover.md +36 -0
  29. package/skills/payment/fraud-analysis/references/report-rationale.md +45 -0
  30. package/skills/payment/fraud-analysis/references/report-template.md +190 -0
  31. package/skills/payment/fraud-analysis/references/review-checklist.md +175 -0
  32. package/skills/payment/fraud-analysis/references/taxonomy.md +79 -0
  33. package/skills/payment/fraud-analysis/references/terminology.md +108 -0
  34. package/skills/payment/fraud-analysis/references/workflow.md +175 -0
  35. package/skills/payment/payment-analysis/LICENSE +3 -0
  36. package/skills/payment/payment-analysis/SKILL.md +127 -0
  37. package/skills/payment/payment-analysis/references/auth-rate-actions.md +30 -0
  38. package/skills/payment/payment-analysis/references/chargebacks.md +88 -0
  39. package/skills/payment/payment-analysis/references/event-layers.md +79 -0
  40. package/skills/payment/payment-analysis/references/fx.md +59 -0
  41. package/skills/payment/payment-analysis/references/journey.md +78 -0
  42. package/skills/payment/payment-analysis/references/metrics.md +62 -0
  43. package/skills/payment/payment-analysis/references/report-template.md +98 -0
  44. package/skills/payment/payment-analysis/references/terminology.md +85 -0
  45. package/skills/payment/payment-analysis/references/visualization.md +47 -0
  46. package/skills/backend/pyspark/SKILL.md +0 -116
  47. package/skills/backend/pyspark/references/parity-testing.md +0 -83
  48. package/skills/backend/pyspark/references/production-validation.md +0 -166
  49. /package/skills/{backend → data-science}/airflow-dag-develop/LICENSE +0 -0
  50. /package/skills/{backend → data-science}/airflow-dag-develop/SKILL.md +0 -0
  51. /package/skills/{backend → data-science}/pyspark/LICENSE +0 -0
  52. /package/skills/{backend → data-science}/pyspark/assets/templates/utils/__init__.py +0 -0
  53. /package/skills/{backend → data-science}/pyspark/assets/templates/utils/hudi_metadata.py +0 -0
  54. /package/skills/{backend → data-science}/pyspark/references/velocity-feature-calculation.md +0 -0
  55. /package/skills/{backend → data-science}/pyspark/scripts/spark_eventlog_summary.py +0 -0
@@ -0,0 +1,62 @@
1
+ # Metrics and denominators
2
+
3
+ Compute with explicit denominators. Prefer count and amount side by side.
4
+ Amounts are USD per [fx.md](fx.md) before any sum or share.
5
+
6
+ ## Authorization
7
+
8
+ | KPI | Formula | Notes |
9
+ |---|---|---|
10
+ | Attempts | count of auth rows in window | |
11
+ | Approvals | count where response = approval | |
12
+ | Auth rate | approvals / attempts | Not capture rate |
13
+ | Approved amount | sum(amount \| approval) | USD after [fx.md](fx.md) |
14
+ | Declines | attempts − approvals | Or hard-only if stated |
15
+ | Decline rate | declines / attempts | |
16
+ | Reason share (of declines) | reason_count / declines | Default for mix tables |
17
+ | Reason incidence (of attempts) | reason_count / attempts | Label clearly if used |
18
+ | Country attempt share | country_attempts / attempts | |
19
+ | Country amount share | country_amount / total_amount | Often diverges from attempt share |
20
+ | Country auth rate | country_approvals / country_attempts | Require min-n (e.g. ≥30) for rankings |
21
+
22
+ Trend grains: daily (`Timestamp` date) and weekly (state week rule, e.g. W-SUN).
23
+
24
+ Card product: `Payment method` × `Card type` when both exist.
25
+
26
+ ## Settlement
27
+
28
+ | KPI | Formula | Notes |
29
+ |---|---|---|
30
+ | Sale count / amount | Settlement type = Sale | |
31
+ | Refund count / amount | Settlement type = Refund | Show sign or absolute + label |
32
+ | Net settled amount | sales + refunds (refunds negative) | |
33
+ | Refund rate (count) | refunds / sales | |
34
+ | Refund rate (amount) | abs(refund amount) / sale amount | |
35
+ | Status mix | by Settlement status | |
36
+
37
+ Do not equalize settled sales to approved auth amount without a timing caveat.
38
+
39
+ ## Chargebacks
40
+
41
+ Monthly **拒付率** by brand is [chargebacks.md](chargebacks.md): Visa
42
+ same-month sales, Mastercard previous-month sales, skip empty brand-months,
43
+ no blended Visa+Mastercard ratio. The table below is mix / match, not that
44
+ programme rate.
45
+
46
+ | KPI | Formula | Notes |
47
+ |---|---|---|
48
+ | CB count / dispute amount | as filed | Amounts USD after [fx.md](fx.md); use abs for magnitude if signed |
49
+ | Reason mix | by Reason category / code | Fraud ≠ all CBs |
50
+ | Match rate to auths | matched / CB rows | State join key |
51
+ | Visa monthly 拒付率 | CB_Visa(M) / Sales_Visa(M) | Received month; not VAMP unless TC40 in file |
52
+ | Mastercard monthly 拒付率 | CB_MC(M) / Sales_MC(M−1) | Skip M if M−1 sales missing |
53
+
54
+ Lag: chargeback date is later than original auth. Programme rates use
55
+ **received** month, not origination vintage.
56
+
57
+ ## Ranking and “Other”
58
+
59
+ - Rank countries by attempts unless the user asks for amount.
60
+ - Roll the long tail into `Other` after top 10–15; show Other share.
61
+ - For decline-by-country matrices, use top countries by attempts and top reasons
62
+ globally; collapse remaining reasons to `Other declines`.
@@ -0,0 +1,98 @@
1
+ # Report template
2
+
3
+ Save:
4
+
5
+ ```text
6
+ reports/payment-analysis/<scope>_overview.md
7
+ reports/payment-analysis/<scope>_overview.zh.md
8
+ reports/payment-analysis/<scope>_overview.zh.html # preferred when charts
9
+ ```
10
+
11
+ Follow [visualization.md](visualization.md), [journey.md](journey.md),
12
+ [terminology.md](terminology.md), and [fx.md](fx.md) for amount currency. Omit
13
+ topic sections whose layer is missing.
14
+ No Markdown bold in report body. After writing, grep banned strings in
15
+ terminology.md.
16
+
17
+ ## Required structure
18
+
19
+ English Markdown headings on the left; Chinese HTML / `.zh.md` on the right.
20
+
21
+ ```markdown
22
+ # Payment Overview — <window> # 支付业务概览 — <window>
23
+
24
+ ## 1. Executive summary # 1. 执行摘要
25
+ ## 2. Top statistics # 2. 核心指标
26
+ ## 3. Payment journey (Sankey) # 3. 支付旅程(Sankey)
27
+ ## 4. Topics # 4. 专题
28
+ ### 4.1 Overall transaction volume # 4.1 整体交易量
29
+ ### 4.2 Week-on-week trend # 4.2 周趋势
30
+ ### 4.3 Authorization rate # 4.3 授权成功率
31
+ ### 4.4 Contribution by BIN country # 4.4 BIN 国家贡献
32
+ ### 4.5 Decline by reason code # 4.5 授权拒绝原因贡献
33
+ ### 4.6 Decline by BIN country # 4.6 各国授权拒绝结构
34
+ ### 4.7 Settlements (when present) # 4.7 清算
35
+ ### 4.8 Chargebacks (optional) # 4.8 拒付
36
+ ## 5. Operational impact # 5. 运营影响
37
+ ## 6. Appendix # 6. 附录
38
+ ```
39
+
40
+ HTML uses the Chinese headings. KPI labels, legends, and Sankey node names use
41
+ the glossary in terminology.md (`授权尝试`, `授权成功`, `授权拒绝`,
42
+ `发卡行国家`, `差错`, `产品服务质量`, `Do not honor`).
43
+
44
+ ## Chapter requirements
45
+
46
+ ### 1. Executive summary / 执行摘要
47
+
48
+ ≤5 numbered findings: volume, trend, auth rate, country concentration, main
49
+ decline or settlement/CB driver.
50
+
51
+ ### 2. Top statistics / 核心指标
52
+
53
+ KPI cards / table from auth + settlement: attempts, approvals, auth rate,
54
+ attempt amount, approved amount, sale count/amount, refund count/amount, net
55
+ settled, distinct PAN/BIN/countries, portfolio constants (MID, MCC, provider).
56
+
57
+ Chinese labels: 授权尝试, 授权成功, 授权成功率, 授权成功金额, 清算销售, 退款,
58
+ 净清算, 发卡行国家.
59
+
60
+ ### 3. Payment journey (Sankey) / 支付旅程
61
+
62
+ Count Sankey per [journey.md](journey.md). Caption lists skipped checkpoints
63
+ (e.g. risk decision / 3DS when authentication fields are empty).
64
+
65
+ Collapsed nodes when 3DS/risk are absent:
66
+
67
+ `开始 → 授权 → 授权成功 / 授权拒绝 → 已清算销售 / 授权成功未清算`
68
+
69
+ ### 4. Topics / 专题
70
+
71
+ Each topic: short lede + chart(s) + supporting table.
72
+
73
+ | Topic | Minimum viz |
74
+ |---|---|
75
+ | 4.1 Volume | count + amount bars (dual axis) |
76
+ | 4.2 WoW trend | weekly count + amount bars; auth-rate line |
77
+ | 4.3 Auth rate | rate line by day and/or segment |
78
+ | 4.4 BIN country | count + amount bars dual axis; rate line |
79
+ | 4.5 Decline reason | count bars |
80
+ | 4.6 Decline by country | stacked count bars |
81
+ | 4.7 Settlements | sale/refund amount bars; country contribution dual axis |
82
+ | 4.8 Chargebacks | reason-category bars; Visa / Mastercard monthly 拒付率 line per [chargebacks.md](chargebacks.md) (note under the rate chart; skip empty brand-months) |
83
+
84
+ Use both authorization and settlement data when both files are present.
85
+
86
+ ### 5. Operational impact / 运营影响
87
+
88
+ Numbered operational takeaways. Do not title this 运营含义. Do not invent
89
+ supervised fraud rules. Choose from
90
+ [auth-rate-actions.md](auth-rate-actions.md) only when the extract supports
91
+ the trigger (SCA → 3DS; high 疑似欺诈 country → 风控拦截 or 3DS; 软性拒绝 →
92
+ retry).
93
+
94
+ ### 6. Appendix / 附录
95
+
96
+ Definitions (from terminology.md) and caveats (FX per [fx.md](fx.md),
97
+ issuer country ≠ ship-to, auth rate ≠ capture, journey skips,
98
+ chargeback rates are network-style estimates).
@@ -0,0 +1,85 @@
1
+ # Terminology
2
+
3
+ Use these terms consistently in payment-analysis reports. Read this file before
4
+ writing English or Chinese copy. After writing, run the proofread step in
5
+ `SKILL.md`.
6
+
7
+ ## Canonical glossary
8
+
9
+ | English | Chinese | Do not write |
10
+ |---|---|---|
11
+ | Attempt | 授权尝试 | 单独的「尝试」;交易;conversion |
12
+ | Approval | 授权成功 | 批准;成功(易与清算混淆) |
13
+ | Auth rate | 授权成功率 | 转化率;单独的「成功率」(首次出现时) |
14
+ | Decline | 授权拒绝 | 全部叫欺诈 |
15
+ | Hard / soft / reject | 硬性拒绝 / 软性拒绝 / 风控拦截 | hard 拦截(指拒绝类型时) |
16
+ | Reason / reason code | 拒绝原因 / 原因码 | |
17
+ | Insufficient funds | 余额不足 | 余额 |
18
+ | Suspected fraud | 疑似欺诈 | 已确认欺诈;正文里写 D102(表内 Pazien 代码可保留) |
19
+ | Invalid expiration date | 有效期无效 | |
20
+ | Security code invalid | 安全码无效 | |
21
+ | Policy reasons | 政策原因 | |
22
+ | Do not honor | Do not honor | 拒绝承兑 |
23
+ | BIN / IIN | BIN / IIN | |
24
+ | BIN country | 发卡行国家 / BIN 国家 | 客户国家;收货国;发卡国 |
25
+ | Settlement | 清算 | |
26
+ | Sale / Refund | 销售 / 退款 | 把销售金额写成清算金额 |
27
+ | Chargeback / dispute | 拒付 / 争议 | 与清算混用 |
28
+ | Clerical | 差错 | 文书;文书错误 |
29
+ | Quality | 产品服务质量 | 质量;质量争议 |
30
+ | Incidence | 发生率(必须写出分母) | |
31
+ | GMV (loose) | 销售金额合计(说明是否含退款) | |
32
+ | Week-on-week trend | 周趋势 | 周环比(除非单元格是环比百分比) |
33
+ | Operational impact | 运营影响 | 运营含义 |
34
+ | Ticket size | 客单价 | |
35
+ | Debit / credit | 借记卡 / 贷记卡 | 借记;贷记 |
36
+ | Payment journey | 支付旅程 | 支付旅程 Sankey;Sankey 旅程(标题/结构行) |
37
+ | This analysis window | 当前分析时段内 | 本窗口 |
38
+ | SCA | SCA | 强客户认证(可在首次括注后只用 SCA) |
39
+ | Authentication / 3DS | 鉴权 / 认证 / 3DS | 把 3DS 叫转化 |
40
+ | Visa / Mastercard 拒付率 | Visa / Mastercard 拒付率 | VAMP ratio(无 TC40 时);ECM(从本抽取断言商户在项目中) |
41
+ | Amounts in USD | 金额均为 USD | 未做汇率折算(当全部已是 USD) |
42
+ | Count-heavy / amount-light | 授权尝试笔数多但金额小 | 走廊;笔数负担型走廊 |
43
+
44
+ Keep proper nouns in source form when useful: Visa, MasterCard, BIN, 3DS,
45
+ Sankey (in chart captions only), Order ID, MCC, Worldpay, Pazien, MID.
46
+ Do not use Pazien codes (e.g. D102) in prose; use 疑似欺诈 / suspected fraud.
47
+ Tables may still show the code column.
48
+
49
+ ## Axis and UI exceptions
50
+
51
+ - Rate axis may say `成功率 %` if the legend or lede already says 授权成功率.
52
+ Amount axes still use the full metric plus currency
53
+ ([visualization.md](visualization.md)).
54
+ - Week-over-week *percentage* columns may say 笔数周环比 / 金额周环比.
55
+ - Product rankings that exclude tiny samples: write
56
+ `在授权尝试超过 50 笔的卡种里` (or the n actually used), never
57
+ `有规模的最弱卡种` or `弱卡种`.
58
+ - If every amount is USD, write `金额均为 USD`. Do not add `未做汇率折算`.
59
+ Only mention missing FX conversion when mixed currencies were actually summed
60
+ or dropped.
61
+ - Chart amount fields append the currency: `授权拒绝金额 USD`.
62
+ - HTML footer: data-source filenames only. No skill name, no Chart.js / ECharts.
63
+ - Do not call issuer country “customer country”.
64
+ - Do not call auth rate “conversion” unless contrasting checkout vs issuer
65
+ response, and then say so.
66
+
67
+ ## Banned strings (grep after writing)
68
+
69
+ Chinese: `走廊`, `杠杆`, `弱转化`, `转化率`, `赋能`, `抓手`, `闭环`, `打法`,
70
+ `运营含义`, `本窗口`, `弱卡种`, `有规模的最弱`, `强客户认证`, `拒绝承兑`,
71
+ `文书错误`, `质量争议`, `发卡国`
72
+
73
+ English: `auth drag`, `conversion problem`, `count-heavy`, `amount-light`,
74
+ `Protect US amount`, `Unlike currencies`
75
+
76
+ If a hit is a legitimate exception (e.g. 周环比 as a table header for WoW %,
77
+ or D102 in a reason-code table), leave it and do not expand the ban.
78
+
79
+ ## Proofread
80
+
81
+ 1. Apply the glossary to headings, KPI labels, chart legends, Sankey skip
82
+ chips, and body copy.
83
+ 2. Grep the report files for the banned strings.
84
+ 3. Replace hits using this table. Do not change numbers, formulas, or file
85
+ paths.
@@ -0,0 +1,47 @@
1
+ # Visualization rules
2
+
3
+ Preferred deliverable: self-contained HTML with Chart.js (KPI charts) and
4
+ ECharts (Sankey). Markdown tables remain required beside charts.
5
+
6
+ ## Mark encoding
7
+
8
+ | Metric kind | Chart mark | Notes |
9
+ |---|---|---|
10
+ | Count | Bar | Vertical or horizontal |
11
+ | Amount | Bar | Never encode amount as a line |
12
+ | Rate | Line | Percent on its own axis |
13
+
14
+ ## Multi-metric charts
15
+
16
+ 1. Count and amount in the same chart → both are **bars** on **different axes**
17
+ (left = count, right = amount).
18
+ 2. Rate with count and/or amount → rate is a **line** on a dedicated % axis.
19
+ 3. Rate alone → line chart (points allowed). No bars for a lone rate series.
20
+ 4. Composition of one total → doughnut or stacked bars; not for time trends.
21
+
22
+ ## Axis hygiene
23
+
24
+ - Label every axis with the full metric name. Amount axes/legends include
25
+ currency (`授权尝试金额 USD`, `授权拒绝金额 USD`, `清算销售金额 USD`,
26
+ `拒付金额 USD`). Do not use bare `笔数` / `金额` / `%`.
27
+ - Only one axis draws the main grid on dual-axis charts.
28
+ - Stacked bars = mix within a group (e.g. decline reasons by country), never rate.
29
+ - Sort categories by a stated order (usually attempts desc).
30
+
31
+ ## Required charts (HTML)
32
+
33
+ 1. Payment journey Sankey — see [journey.md](journey.md)
34
+ 2. Volume trend (WoW / daily): count bar + amount bar (dual axis) + auth-rate line
35
+ 3. Authorization rate by segment (card product and/or country): rate as line
36
+ 4. BIN-country contribution: count bar + amount bar (dual axis)
37
+ 5. Decline by reason: count bars (+ amount bars/dual axis if shown)
38
+ 6. Decline mix by BIN country: stacked count bars
39
+
40
+ ## Anti-patterns
41
+
42
+ - Line for amount or count
43
+ - Bar for a standalone rate series
44
+ - Count and amount on one axis without dual-axis scaling
45
+ - Mixing native EUR (etc.) with USD in one amount series; convert first ([fx.md](fx.md))
46
+ - Plotting a 0% chargeback-rate point for a brand-month with no CBs ([chargebacks.md](chargebacks.md))
47
+ - Inventing Sankey checkpoints not present in the data
@@ -1,116 +0,0 @@
1
- ---
2
- name: pyspark
3
- description: Implement, debug, review, and optimize Apache Spark or PySpark batch jobs, ETL pipelines, feature calculations, and Hudi/warehouse writes. Use whenever the task involves Spark DataFrames or Spark SQL, joins, windows, grouped state, UDFs, shuffle, skew, spill, slow stages, driver/executor OOM, Spark Connect, spark-submit, Hudi commits, backfills, bidirectional parity tests, validate-only canaries, or Spark event-log profiling — including when the user only mentions a warehouse job timeout, empty Hudi write, or performance regression on an EMR/Dataproc job.
4
- ---
5
-
6
- # PySpark
7
-
8
- Use this skill as the workflow and routing index. Read the referenced playbooks
9
- that match the task before editing code or running a production-scale job.
10
-
11
- ## Workflow
12
-
13
- ```text
14
- contract -> transform -> materialize -> test -> CI/CD
15
- ```
16
-
17
- Follow all five steps in order.
18
-
19
- ### 1. Contract
20
-
21
- - **Input:** declare all parameters, defaults, validation, source snapshot/date
22
- scope, overlapping-run behavior, idempotency/single-writer rules, and exact
23
- Python and PySpark versions.
24
- - **Output:** declare plain Spark versus Hudi table type, output identity/schema,
25
- partition columns, record key, deterministic order or precombine field, and
26
- append/upsert/overwrite write mode.
27
- - **Conventions:** declare native Spark versus Spark Connect and full rebuild
28
- versus incremental processing, including the incremental watermark,
29
- lookback/checkpoint, and restart boundary when applicable.
30
-
31
- Read [ETL Contract](references/etl-contract.md). When the repository has no ETL
32
- base module, adapt [`assets/templates/etl.py`](assets/templates/etl.py).
33
-
34
- ### 2. Transform
35
-
36
- - Make `transform` accept DataFrames and return the complete candidate without
37
- invoking validation or a writer.
38
- - Estimate rows after joins/explodes/groups and measure maximum rows per key.
39
- Fix algorithmic expansion and hot keys before tuning Spark configuration.
40
- - Prefer built-in expressions and windows. Use exact bucketing or one sorted
41
- grouped sweep only for state that built-ins cannot express cleanly.
42
- - Preserve complete row identity, ordering, output columns, and required
43
- behavior.
44
- - Avoid explicit `.cache()` and `.persist()` by default.
45
-
46
- Read [Transformation Design](references/transformation-design.md). For temporal
47
- windows or FIFO paths, also read
48
- [Velocity Feature Calculation](references/velocity-feature-calculation.md).
49
-
50
- ### 3. Materialize
51
-
52
- - Treat materialization as the Spark action that executes the lazy candidate,
53
- not as caching.
54
- - Use one aggregate validation action before every writer. Reference every
55
- derived output so Catalyst cannot prune untested calculations.
56
- - Check row count, distinct/null/duplicate identities, domain blockers, and
57
- required date/partition scope. Fail closed on any blocker.
58
- - Bind validation and publication to the same immutable source snapshot because
59
- Spark may recompute the candidate.
60
- - Keep the validate-only guard inside the job and before every writer. Do not
61
- add `.count()`, `.cache()`, or `.persist()` to warm execution.
62
-
63
- Read [Production Validation](references/production-validation.md).
64
-
65
- ### 4. Test
66
-
67
- 1. **Local test:** run pure state, tiny real-Spark, plan-shape, focused repository,
68
- and bidirectional parity tests. Cover boundaries, ties, nulls, arithmetic,
69
- identity, and precedence.
70
- 2. **Cluster canary test:** run the production entry point and source scope with
71
- `--validate-only`. Require zero blockers, an explicit publication skip, zero
72
- Spark output records/bytes, and readable driver/event logs.
73
- 3. **Profiling:** measure an isolated baseline, parse the canary event log, and
74
- compare both runs on the same source scope and resource profile. Record
75
- wall/stage/task time, cardinality, shuffle, spill, skew, output, and target
76
- plan patterns.
77
-
78
- ```bash
79
- spark-submit <production-options> job.py --validate-only
80
- set -o pipefail
81
- hdfs dfs -cat "$EVENT_LOG" | \
82
- python {baseDir}/scripts/spark_eventlog_summary.py > canary-summary.json
83
- ```
84
-
85
- Reject missing evidence, incomparable scope, unexplained regressions, or a
86
- missed performance gate. Keep validate-only time separate from write time.
87
-
88
- Read [Parity Testing](references/parity-testing.md),
89
- [Diagnosis And Profiling](references/diagnosis-and-profiling.md), and
90
- [Production Validation](references/production-validation.md).
91
-
92
- ### 5. CI/CD
93
-
94
- 1. Open a PR with required-behavior, test, canary, profiling, and risk evidence.
95
- 2. Run required CI and independent review; fix failures.
96
- 3. Enqueue through the configured merge queue and merge without bypassing gates.
97
- 4. Deploy only the merged commit through the normal deployment path.
98
- 5. Run write-enabled production validation only with explicit authorization;
99
- reconcile counts, identity, schema, coverage, blockers, readability, write
100
- mode, and commit evidence independently.
101
- 6. Restore scheduler/process ownership and report compute/validation,
102
- publication, and end-to-end time separately.
103
-
104
- Read [Production Validation](references/production-validation.md).
105
-
106
- ## Reference Index
107
-
108
- | Need | Read or run |
109
- |---|---|
110
- | Define or review an ETL job contract | [ETL Contract](references/etl-contract.md) and `assets/templates/etl.py` |
111
- | Diagnose or profile | [Diagnosis And Profiling](references/diagnosis-and-profiling.md) |
112
- | Design transformations | [Transformation Design](references/transformation-design.md) |
113
- | Optimize temporal/FIFO state | [Velocity Feature Calculation](references/velocity-feature-calculation.md) |
114
- | Prove parity | [Parity Testing](references/parity-testing.md) |
115
- | Run canaries or writes | [Production Validation](references/production-validation.md) |
116
- | Summarize an event log | `scripts/spark_eventlog_summary.py` |
@@ -1,83 +0,0 @@
1
- # Parity Testing
2
-
3
- Use this playbook for semantics-preserving Spark rewrites.
4
-
5
- ## Keep The Old Path As A Bounded Oracle
6
-
7
- Retain the previous implementation only in tests while parity is being proven.
8
- Use small fixtures that are dense enough to expose the old complexity but
9
- bounded enough to run reliably.
10
-
11
- ## Compare Full Rows In Both Directions
12
-
13
- ```python
14
- expected = old_build(source).select(*OUTPUT_COLUMNS)
15
- actual = new_build(source).select(*OUTPUT_COLUMNS)
16
-
17
- assert expected.exceptAll(actual).count() == 0
18
- assert actual.exceptAll(expected).count() == 0
19
- ```
20
-
21
- Counts alone cannot detect wrong values, swapped identities, or duplicate rows.
22
- Use `exceptAll`, not `subtract`, when duplicate multiplicity matters.
23
-
24
- ## Add Literal Boundary Assertions
25
-
26
- Parity can be self-confirming if oracle and candidate share the same mistake.
27
- Assert critical rules directly:
28
-
29
- ```python
30
- by_id = {row.event_id: row for row in actual.collect()}
31
- assert by_id["at-24h"].count_24h == 2 # inclusive lower boundary
32
- assert by_id["after-24h"].count_24h == 1 # one unit outside
33
- assert by_id["same-time-b"].count_1h == 2 # stable tie ordering
34
- ```
35
-
36
- Keep collected fixtures synthetic and non-sensitive.
37
-
38
- ## Minimum Fixture Matrix
39
-
40
- Cover:
41
-
42
- - empty and single-event groups;
43
- - exact lower/upper boundaries and one unit beyond;
44
- - same-timestamp events with every tie-breaker;
45
- - null key, event time, amount, FX, and counterparty;
46
- - both directions and relevant statuses;
47
- - repeated and distinct counterparties;
48
- - state open, partial close, full close, eviction, and re-entry;
49
- - floating-point add/remove residue;
50
- - one dense hot key;
51
- - duplicate and missing output identities.
52
-
53
- ## Test The Physical Shape
54
-
55
- Assert the expensive relation is absent without snapshotting the whole plan:
56
-
57
- ```python
58
- plan = actual._jdf.queryExecution().executedPlan().toString().lower()
59
- assert not ("join" in plan and "event_time" in plan and "interval" in plan)
60
- ```
61
-
62
- Prefer a targeted helper that counts known pair-join nodes. Also assert the
63
- expected grouped-sweep or built-in operator is present when that guards the
64
- optimization.
65
-
66
- ## Mutation-Prove Important Tests
67
-
68
- Temporarily introduce representative faults and confirm tests fail:
69
-
70
- - change `< cutoff` to `<= cutoff`;
71
- - remove a tie-breaker;
72
- - use ordinary rolling addition/subtraction;
73
- - reintroduce the pair join;
74
- - omit a downstream-required column.
75
-
76
- Record red/green evidence for high-risk money, date, and identity semantics.
77
-
78
- ## Match Test Cost To Ownership
79
-
80
- Behavioral Spark tests belong on files that own formulas, source scope,
81
- history, or state transitions. Unrelated upstream changes may use a focused
82
- required-column check instead of rerunning an expensive full-history suite.
83
- Do not demote behavioral coverage for a semantic change.
@@ -1,166 +0,0 @@
1
- # Production Validation
2
-
3
- Use this playbook before production-scale profiling, canaries, backfills, or
4
- writes.
5
-
6
- ## Fail-Closed Admission
7
-
8
- Do not launch until all applicable facts are recorded and valid:
9
-
10
- - exact commit/module hash;
11
- - production Spark, Python, Java, and table-format runtimes;
12
- - source snapshot or date boundary;
13
- - scheduler pause/ownership state;
14
- - zero competing warehouse work;
15
- - explicit validate-only or write-enabled mode;
16
- - durable writable log and event-log destination;
17
- - restart/recovery plan for writes.
18
-
19
- Classify existing processes and working-tree changes before touching them. Do
20
- not kill or overwrite another operator's work.
21
-
22
- ## Run A Validate-Only Canary Before The Writer
23
-
24
- Implement the guard in the Spark job, not only in a wrapper. Keep each stage
25
- literal:
26
-
27
- ```python
28
- import json
29
-
30
- from pyspark.sql import functions as F
31
-
32
-
33
- def validate_candidate(candidate):
34
- # first() is a Spark action: it executes the lazy transform plan.
35
- metrics = candidate.agg(
36
- F.count("*").alias("row_count"),
37
- F.countDistinct("event_id").alias("distinct_event_count"),
38
- F.sum(
39
- F.when(F.col("event_id").isNull(), 1).otherwise(0)
40
- ).alias("null_event_id_count"),
41
- F.sum(
42
- F.when(F.col("event_count_1h") < 0, 1).otherwise(0)
43
- ).alias("invalid_event_count_1h"),
44
- F.count("amount_usd_1h").alias("non_null_amount_usd_1h_count"),
45
- ).first().asDict()
46
-
47
- blockers = []
48
- if metrics["row_count"] != metrics["distinct_event_count"]:
49
- blockers.append("event_id is not unique")
50
- if metrics["null_event_id_count"]:
51
- blockers.append("event_id contains nulls")
52
- if metrics["invalid_event_count_1h"]:
53
- blockers.append("event_count_1h contains negative values")
54
- if blockers:
55
- raise RuntimeError(f"validation blocked: {blockers}")
56
- return metrics
57
-
58
-
59
- def run(validate_only, source_snapshot):
60
- # Bind every action to the same immutable source version or date boundary.
61
- source = extract(source_snapshot)
62
- candidate = transform(source) # build a lazy DataFrame plan
63
-
64
- metrics = validate_candidate(candidate) # execute validation aggregates
65
- print(json.dumps(metrics, sort_keys=True))
66
-
67
- if validate_only: # canary stops before any writer call
68
- print("validate-only: skipping publication")
69
- return
70
-
71
- publish(candidate) # execute the writer from the same source
72
- reconcile_published_scope(metrics) # independently read back aggregates
73
- ```
74
-
75
- Validation and publication are separate Spark actions, so Spark may recompute
76
- the candidate for publication. Pin every source read to the same immutable
77
- snapshot, table version, or exact date boundary. If the source cannot be pinned,
78
- the validation result does not prove what the writer will publish; stop and
79
- redesign the run.
80
-
81
- Adapt identity and blocker checks to the table's required behavior. Reference
82
- every derived output in at least one validation aggregate; otherwise Catalyst
83
- may prune an unreferenced calculation and the canary will not exercise the full
84
- candidate. Validation must use aggregate evidence and fail before publication.
85
- The canary exercises the production-shaped extract, transform, and validation
86
- path without calling the writer.
87
-
88
- ## Avoid Explicit Cache And Persistence
89
-
90
- Do not add `.cache()` or `.persist()` as a performance fix. They add storage,
91
- serialization, eviction, and cleanup concerns without improving the logical
92
- work. Prefer a better transformation shape and a pinned source snapshot.
93
-
94
- Treat existing persistence as separate operational behavior: do not expand it
95
- while optimizing the job, and remove it when repository tests and same-snapshot
96
- evidence show it is unnecessary.
97
-
98
- ## Launch Through The Production Entry Point
99
-
100
- Use the same `spark-submit` path and driver/executor runtimes as production:
101
-
102
- ```bash
103
- export PYSPARK_DRIVER_PYTHON="$DRIVER_PYTHON"
104
- export PYSPARK_PYTHON="$EXECUTOR_PYTHON"
105
-
106
- spark-submit \
107
- --name "$CANARY_NAME" \
108
- --conf spark.pyspark.driver.python="$DRIVER_PYTHON" \
109
- --conf spark.pyspark.python="$EXECUTOR_PYTHON" \
110
- --conf spark.eventLog.enabled=true \
111
- job.py --validate-only
112
- ```
113
-
114
- Do not replace deployed scheduler files with an unmerged staged module. Stage
115
- an isolated artifact or deploy merged code through the normal path.
116
-
117
- ## Prove The Canary Did Not Write
118
-
119
- Require all three forms of evidence:
120
-
121
- 1. Validation completed with zero blockers.
122
- 2. The terminal log says publication was skipped.
123
- 3. The event log reports zero Spark output records and bytes.
124
-
125
- Parse retained logs with upstream failures visible:
126
-
127
- ```bash
128
- set -o pipefail
129
- hdfs dfs -cat "$EVENT_LOG" | \
130
- python {baseDir}/scripts/spark_eventlog_summary.py > canary-summary.json
131
- ```
132
-
133
- Record application ID, shell/event-log wall, stage/task metrics, plan nodes,
134
- shuffle, spill, output, code hash, and snapshot denominator.
135
-
136
- ## Run Write Verification Separately
137
-
138
- Run write-enabled verification only after merge, deployment, successful
139
- validate-only evidence, and explicit authorization. Re-check admission just
140
- before launch.
141
-
142
- Verify row count, distinct identity count, date/partition coverage, commit
143
- success, schema, consumer compatibility, validation blockers, write mode,
144
- single-writer constraints, and downstream readability.
145
-
146
- Report compute/validation, publication, and total time separately. Never
147
- present validate-only time as end-to-end write performance.
148
-
149
- If a write run is cancelled, prove from logs/source that publication was not
150
- reached. YARN `KILLED` alone does not prove no commit occurred.
151
-
152
- ## Keep Evidence Aggregate-Only
153
-
154
- Safe evidence includes counts, null rates, distinct counts, min/max dates,
155
- hashes, stage metrics, plan-node counts, and commit IDs. Do not print
156
- credentials, entity rows, account identifiers, or per-entity values.
157
-
158
- ## Restore Ownership
159
-
160
- After the run:
161
-
162
- 1. Confirm application and wrapper processes are terminal.
163
- 2. Remove agent-owned watchers and temporary processes.
164
- 3. Restore scheduler pause state or ownership.
165
- 4. Separate unrelated downstream failures from the optimized job.
166
- 5. Publish evidence with limitations and unmeasured surfaces.