@thinkingai/ae-cli 6.1.17 → 6.1.18
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/{auth-2WTQOP77.js → auth-GBMV6TEJ.js} +3 -3
- package/dist/{auth-77BUFLGC.js → auth-ROB2EDYV.js} +12 -13
- package/dist/{capability-72DTW5M2.js → capability-DKMYUTLC.js} +50 -13
- package/dist/{capability-PJHNI4GJ.js → capability-HYVVPG25.js} +49 -12
- package/dist/chunk-3FY3RJ26.js +293 -0
- package/dist/{chunk-PTE56QPL.js → chunk-3KWQYGYI.js} +4 -0
- package/dist/{chunk-UOUS37JQ.js → chunk-4XXOWOTA.js} +3 -3
- package/dist/{chunk-GT46FPXN.js → chunk-BYYS3ANB.js} +17 -8
- package/dist/{chunk-YA6SMTXG.js → chunk-EFH4XWYC.js} +3 -3
- package/dist/{chunk-4SGZG4XY.js → chunk-J2DEBMRF.js} +9 -7
- package/dist/{chunk-LYVNONC4.js → chunk-JHENBQ5B.js} +35 -0
- package/dist/{chunk-VR3LCBHW.js → chunk-JRJY5DMJ.js} +5 -5
- package/dist/{chunk-ILIU36SU.js → chunk-OMPRXM3V.js} +3 -3
- package/dist/{chunk-VPKZ7I72.js → chunk-QNOLN2LJ.js} +2 -2
- package/dist/{chunk-SAU3QFIQ.js → chunk-QZ3AS4KK.js} +3 -3
- package/dist/chunk-RJDU7NYP.js +1198 -0
- package/dist/{chunk-4KVPKXFX.js → chunk-RNAALWJK.js} +2 -2
- package/dist/{chunk-C4MGVGJW.js → chunk-SERWF6G5.js} +1 -1
- package/dist/{chunk-RGKJGKT7.js → chunk-Y3LOALAV.js} +5 -5
- package/dist/{chunk-P3FGXJTU.js → chunk-ZQ47LWTI.js} +4 -4
- package/dist/chunk-ZQKDZXDO.js +317 -0
- package/dist/{client-TKG4WBHN.js → client-L2YDMHQ6.js} +5 -4
- package/dist/{community-report-client-FI4LNVYS.js → community-report-client-C7WDGET3.js} +1 -2
- package/dist/{config-RE6CMGPK.js → config-BMYZX2UE.js} +7 -6
- package/dist/{data-integration-XQYB4X4F.js → data-integration-QEKDWQDY.js} +1601 -192
- package/dist/index.js +131 -1241
- package/dist/{local-data-upload-client-BWHSUQQK.js → local-data-upload-client-4YYHSYD6.js} +1 -2
- package/dist/{memory-CHRU2F7W.js → memory-3ORCR7JH.js} +6 -6
- package/dist/{memory-YK33G4T7.js → memory-I2WXDTV2.js} +5 -5
- package/dist/{metadata-XXR34N5P.js → metadata-I4C2EWUN.js} +10 -10
- package/dist/{metadata-UILXHBWF.js → metadata-VUOQJE26.js} +9 -9
- package/dist/{model-NR3JHFSJ.js → model-HLHIEFMU.js} +5 -5
- package/dist/{model-K3KLWIW6.js → model-UGRDX4MW.js} +6 -6
- package/dist/personal-semantic-preference-LIPACBDX.js +239 -0
- package/dist/personal-semantic-preference-OEISBRHM.js +239 -0
- package/dist/project-semantic-FFPWFPIW.js +1114 -0
- package/dist/project-semantic-RT3R2VQD.js +1114 -0
- package/dist/{sync-FCKOVWWS.js → sync-HKIOZXQE.js} +6 -6
- package/dist/{sync-DAVKYVMW.js → sync-TFHU2UTG.js} +7 -7
- package/dist/{te-agent-HLW4VTQK.js → te-agent-BR6VDBNX.js} +9 -8
- package/dist/{te-agent-4BKBODMF.js → te-agent-VLYOV7S4.js} +8 -7
- package/dist/{te-analysis-ZMNGOVNW.js → te-analysis-4YGQL5RC.js} +437 -38
- package/dist/{te-analysis-O6DCO6BS.js → te-analysis-7VUNUYWZ.js} +436 -37
- package/dist/{te-community-HLC43QKH.js → te-community-5DMNKJWY.js} +5 -5
- package/dist/{te-community-6HPBWJUZ.js → te-community-ISDQWJU7.js} +6 -6
- package/dist/{te-dataops-HDRUXY4K.js → te-dataops-6P5IKWNJ.js} +8 -7
- package/dist/{te-dataops-EJP56W3K.js → te-dataops-CVULXNVB.js} +7 -6
- package/dist/{te-engage-RAK5PESW.js → te-engage-KZPR5R22.js} +9 -9
- package/dist/{te-engage-FGBGQ4IY.js → te-engage-N5WI32H6.js} +8 -8
- package/dist/{te-experiment-SO5MPDMJ.js → te-experiment-6BITX4RD.js} +226 -8
- package/dist/{te-experiment-VZF7BT6G.js → te-experiment-UVR4HLND.js} +227 -9
- package/dist/{te-kb-SQCLHG6X.js → te-kb-RCLSSH2Q.js} +7 -7
- package/dist/{te-system-Z77IKZFN.js → te-system-FXITO2JG.js} +5 -5
- package/dist/{te-system-YARIK4S5.js → te-system-K2GYMCTB.js} +6 -6
- package/dist/{te-team-EFKWYKMK.js → te-team-ADOC2ROP.js} +6 -6
- package/dist/{update-OGPSZM5A.js → update-YCYCKJOO.js} +7 -6
- package/package.json +1 -1
- package/skills/ae-agent/SKILL.md +3 -4
- package/skills/ae-agent/references/edit-skill.md +3 -0
- package/skills/ae-agent/references/get-skill-content.md +1 -1
- package/skills/ae-agent/references/rescan-skills.md +15 -13
- package/skills/ae-agent/references/upload-skill.md +7 -4
- package/skills/ae-analysis/SKILL.md +45 -4
- package/skills/ae-analysis/metadata_resolution.md +38 -4
- package/skills/ae-analysis/references/analysis_data_retrieval.md +29 -0
- package/skills/ae-analysis/references/asset_authentication_export.md +22 -0
- package/skills/ae-analysis/references/asset_authentication_list.md +18 -14
- package/skills/ae-analysis/references/asset_authentication_update.md +29 -14
- package/skills/ae-analysis/references/command_index.md +17 -9
- package/skills/ae-analysis/references/dashboard_get.md +18 -1
- package/skills/ae-analysis/references/dashboard_update.md +3 -0
- package/skills/ae-analysis/references/personal_semantic_preference_add.md +23 -0
- package/skills/ae-analysis/references/personal_semantic_preference_delete.md +17 -0
- package/skills/ae-analysis/references/personal_semantic_preference_get.md +19 -0
- package/skills/ae-analysis/references/personal_semantic_preference_list.md +21 -0
- package/skills/ae-analysis/references/personal_semantic_preference_update.md +19 -0
- package/skills/ae-data-integration/SKILL.md +23 -4
- package/skills/ae-data-integration/references/custom-layer.md +93 -0
- package/skills/ae-data-integration/references/error-handling.md +92 -0
- package/skills/ae-data-integration/references/handoff.md +77 -18
- package/skills/ae-data-integration/references/local-analysis.md +1 -1
- package/skills/ae-data-integration/references/reuse.md +9 -5
- package/skills/ae-data-integration/references/sink-upload.md +1 -1
- package/skills/ae-data-integration/references/source-inspect.md +18 -12
- package/skills/ae-data-integration/references/tracking-plan.md +7 -5
- package/skills/ae-data-integration/references/transform.md +10 -10
- package/skills/ae-data-integration/references/ue-mapping.md +33 -11
- package/skills/ae-engage/references/build-task-save-guide.md +9 -0
- package/skills/ae-engage/references/save-task.md +82 -0
- package/skills/ae-experiment/SKILL.md +8 -2
- package/skills/ae-experiment/references/manage_feature_whitelist.md +66 -0
- package/skills/ae-experiment/references/manage_guardrail_metrics.md +26 -0
- package/skills/ae-experiment/references/save_experiment.md +1 -1
- package/skills/ae-kb/SKILL.md +1 -1
- package/skills/ae-project-semantic/SKILL.md +193 -0
- package/skills/ae-project-semantic/references/query-routing-v5.md +165 -0
- package/skills/ae-project-semantic/references/recommendation-quality.md +68 -0
- package/dist/chunk-QGM4M3NI.js +0 -37
- package/dist/chunk-ZZUOD757.js +0 -598
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
# personal-semantic-preference add
|
|
2
|
+
|
|
3
|
+
Add one personal semantic preference for the authenticated user in one project.
|
|
4
|
+
|
|
5
|
+
Use this when the user explicitly asks to save a personal semantic preference, or when the current project task contains an explicit stable statement, correction, or confirmation that should become a reusable current-user preference. A current-user working definition remains eligible even when the same content may benefit other users. A second "save" confirmation is not required after that evidence gate is met, unless the target meaning is ambiguous.
|
|
6
|
+
|
|
7
|
+
Command:
|
|
8
|
+
|
|
9
|
+
```bash
|
|
10
|
+
ae-cli personal-semantic-preference add --project-id <project_id> --context-type <context_type> --title <title> --summary <summary> --content <content> [--keywords '["keyword"]'] [--resource-refs '[{"resource_type":"report","resource_key":"101","display_name":"Revenue daily report"}]'] [--fresh-until-at "yyyy-MM-dd HH:mm:ss"] [--request-id <id>]
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
Use `preference` for durable interpretation/output preferences, `asset_context` for durable wording or intent bound to exact assets, `experience` for confirmed reusable work methods, and `background` for stable personal context. `--resource-refs` is required and non-empty only for `asset_context`; for every other type it must be absent or empty.
|
|
14
|
+
|
|
15
|
+
`--resource-refs` accepts 1 to 50 ordered objects. Each object contains exactly `resource_type`, string `resource_key`, and `display_name`; `(resource_type, resource_key)` must be unique. `resource_type` is generic lower snake_case rather than a report-only enum, so events, properties, metrics, tags, clusters, reports, dashboards, data tables, and later asset types share the same shape. Array order is the user's intended priority.
|
|
16
|
+
|
|
17
|
+
`--request-id` is an idempotency key; omit it for ordinary interactive use and the CLI will generate one.
|
|
18
|
+
|
|
19
|
+
This command creates only a current-user preference; it never creates or approves a project semantic. Do not reject an otherwise valid personal preference merely because the same content may benefit other users, and do not imply that the saved preference is shared authority. Keep project-candidate recommendation separate: after personal capture, ask whether to recommend broadly reusable content as a project semantic candidate, and submit nothing without that choice.
|
|
20
|
+
|
|
21
|
+
Store only the present working preference. Do not append future governance or lifecycle instructions. Do not use this command for company knowledge, standalone metadata facts, reports, dashboards, transient task details, one-off analysis results, or automatic stale/expired preference handling.
|
|
22
|
+
|
|
23
|
+
Output is the gateway envelope. `data.preference` contains the full created preference, current revision, and complete ordered `resource_refs`. A separate readback is unnecessary unless a later step needs to refresh the record.
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
# personal-semantic-preference delete
|
|
2
|
+
|
|
3
|
+
Soft-delete one personal semantic preference using optimistic locking.
|
|
4
|
+
|
|
5
|
+
Use this only after explicit user confirmation to remove a saved personal preference.
|
|
6
|
+
|
|
7
|
+
Command:
|
|
8
|
+
|
|
9
|
+
```bash
|
|
10
|
+
ae-cli personal-semantic-preference delete --project-id <project_id> --id <preference_id> --expected-revision <revision> [--request-id <id>] --yes
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
Read the current item first and pass its revision as `--expected-revision`. This is a high-risk write; use `--dry-run` when the target is ambiguous.
|
|
14
|
+
|
|
15
|
+
Do not use this command for automatic stale/expired preference handling. Stale or low-freshness personal preferences are hidden by list filtering and backend maintenance.
|
|
16
|
+
|
|
17
|
+
Output is the gateway envelope. `data.deleted` identifies the deleted preference and status. After deletion, it should no longer appear in `personal-semantic-preference list`.
|
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
# personal-semantic-preference get
|
|
2
|
+
|
|
3
|
+
Get one personal semantic preference by ID.
|
|
4
|
+
|
|
5
|
+
Use this command after `personal-semantic-preference list` identifies a likely current-user preference. Pass `--mark-used` when the preference is adopted for the answer, query path, or as the matched target for an update.
|
|
6
|
+
|
|
7
|
+
Command:
|
|
8
|
+
|
|
9
|
+
```bash
|
|
10
|
+
ae-cli personal-semantic-preference get --project-id <project_id> --id <preference_id> [--mark-used]
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
Input uses `project_id`, `id`, and optional `mark_used`. `id` must be the exact `preference_<id>` value returned by list/add.
|
|
14
|
+
|
|
15
|
+
Do not use this command as a keyword search, project semantics lookup, or asset catalog lookup. Do not call it repeatedly for every catalog row. Do not use `--mark-used` for a candidate that turns out not to match the user's intent or is only inspected and then rejected.
|
|
16
|
+
|
|
17
|
+
When a published project semantic conflicts with the personal item, the project semantic is the formal definition. Fetch and mark the personal item only when it materially affects the response, such as an explicitly requested non-formal alternative; never silently use it to override the project semantic.
|
|
18
|
+
|
|
19
|
+
Output is the gateway envelope. `data.preference` contains the full personal preference, including content, complete ordered `resource_refs`, and revision. When `--mark-used` is set, the backend increments `heat_count` and updates `last_used_at` for that record.
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
# personal-semantic-preference list
|
|
2
|
+
|
|
3
|
+
List the authenticated user's active personal semantic preference catalog for one project.
|
|
4
|
+
|
|
5
|
+
Use this command once per host, authenticated user, project, and conversation after resolving the project. Keep the returned directory in conversation context for later questions where user-specific wording or preferences may change interpretation.
|
|
6
|
+
|
|
7
|
+
Command:
|
|
8
|
+
|
|
9
|
+
```bash
|
|
10
|
+
ae-cli personal-semantic-preference list --project-id <project_id>
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
Input uses `project_id` only. Do not add pagination parameters: the backend returns a compact catalog intended for agent context.
|
|
14
|
+
|
|
15
|
+
Do not use this command for project semantics, shared knowledge, metadata catalogs, report/dashboard lists, or complete asset discovery. It only returns the current user's personal semantic preferences in the current project.
|
|
16
|
+
|
|
17
|
+
Output is the gateway envelope. `data.items[]` contains only `id`, `context_type`, `title`, truncated `summary`, limited `keywords`, `resource_ref_count`, distinct `resource_types`, and `revision`; it deliberately omits content, full asset references, heat, and timestamps. `data.returned_count` is at most 200, `data.truncated` says whether entries were omitted, and `data.selection_policy` is `HOT_160_PLUS_RECENT_40`: up to 160 highest-heat items plus up to 40 recently changed items not already selected. The backend may return fewer items to keep the data payload within 64 KiB.
|
|
18
|
+
|
|
19
|
+
If one returned item is actually adopted to interpret the user's request, call `ae-cli personal-semantic-preference get --project-id <project_id> --id <preference_id> --mark-used` before using its full content. Do not mark an item used when it was only inspected or rejected.
|
|
20
|
+
|
|
21
|
+
Compare a likely match with the published project semantic catalog. Published project semantics remain the formal project-wide authority; personal items provide current-user defaults and working interpretations. If they conflict, use the project semantic for the formal result, explicitly disclose the personal difference, and do not mark the personal item used unless the user explicitly adopts it as a labeled alternative.
|
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
# personal-semantic-preference update
|
|
2
|
+
|
|
3
|
+
Update one personal semantic preference using optimistic locking.
|
|
4
|
+
|
|
5
|
+
Use this when the user explicitly asks to change an existing personal preference, or when the current project task provides an explicit stable correction or confirmation that should replace a matching saved current-user preference. A current-user working definition remains eligible even when the same content may benefit other users. First read the matched current item with `personal-semantic-preference get --mark-used` and pass the returned revision as `--expected-revision`.
|
|
6
|
+
|
|
7
|
+
Command:
|
|
8
|
+
|
|
9
|
+
```bash
|
|
10
|
+
ae-cli personal-semantic-preference update --project-id <project_id> --id <preference_id> --expected-revision <revision> --context-type <context_type> --title <title> --summary <summary> --content <content> [--keywords '["keyword"]'] [--resource-refs '[{"resource_type":"event","resource_key":"$login","display_name":"Login event"}]'] [--fresh-until-at "yyyy-MM-dd HH:mm:ss"] [--request-id <id>]
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
The update replaces all editable fields. Include the full intended `title`, `summary`, `content`, keywords, and asset bindings rather than a partial patch. `asset_context` requires 1 to 50 complete ordered `resource_refs`; other context types cannot contain non-empty refs.
|
|
14
|
+
|
|
15
|
+
This command updates only the current-user preference; it does not update, approve, or publish project semantics. Do not block the update merely because the same content may benefit other users, and do not imply that the saved preference is shared authority. Keep project-candidate recommendation separate and ask before submitting it.
|
|
16
|
+
|
|
17
|
+
Keep future governance and lifecycle handling out of the stored content. Do not use this command for shared knowledge, standalone metadata facts, reports, dashboards, transient task details, one-off analysis results, or automatic stale/expired preference handling.
|
|
18
|
+
|
|
19
|
+
Output is the gateway envelope. `data.preference` contains the full updated preference and new revision. If the revision is stale, read the latest record and ask the user how to merge rather than overwriting silently.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: ae-data-integration
|
|
3
|
-
description: "Bring local CSV, TSV, TXT, JSON, JSONL (NDJSON), XLS, and XLSX files into AE end-to-end: identify the source's business meaning, generate and confirm a tracking plan, transform rows into UE records, and upload. Also supports privacy-preserving local analysis and handing a small file to AE Agent. Use whenever a user wants to import offline/local data into AE or analyze a file without uploading it."
|
|
3
|
+
description: "Bring local CSV, TSV, TXT, JSON, JSONL (NDJSON), XLS, and XLSX files into AE end-to-end: identify the source's business meaning, generate and confirm a tracking plan, transform rows into UE records, and upload. Also supports privacy-preserving local analysis and handing a small file to AE Agent. Use whenever a user wants to import offline/local data into AE or analyze a file without uploading it. Trigger words: 本地数据导入 / 离线数据 / 数据文件 / 文件导入 / 文件上报 / CSV 导入 / Excel 导入 / TSV 导入 / JSON 导入 / 导入到 AE / 导入到 ThinkingData / local data import / import local file."
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# AE Data Integration
|
|
@@ -23,16 +23,35 @@ Two entrances lead here: the AE Agent dialog (attach / plus-button upload) and `
|
|
|
23
23
|
- If a batch times out or loses the network, treat that batch as unknown. Stop. Ask the user to verify receiver/AE data before the user chooses `--resume-from`; never resume automatically.
|
|
24
24
|
- Local analysis stays local. AE Agent attachment is a separate, confirmed branch with a 50 MB per-file limit.
|
|
25
25
|
|
|
26
|
+
## When to use / When NOT to use
|
|
27
|
+
|
|
28
|
+
Use this skill when the user wants to bring a **local data file** (CSV/TSV/TXT/JSON/JSONL/XLS/XLSX) into AE, or analyze such a file locally without uploading.
|
|
29
|
+
|
|
30
|
+
| User intent | Use this instead |
|
|
31
|
+
| --- | --- |
|
|
32
|
+
| How to integrate the SDK / tracking code / LogBus2 config / reporting-error triage (usage Q&A, no local file) | ae-data-integration-helper |
|
|
33
|
+
| Database / datasource direct sync (MySQL, DataX, data warehouse, data-dev platform) | ae-dataops |
|
|
34
|
+
| Community content (posts / comments / chat / WeCom groups) insight or submission | ae-community |
|
|
35
|
+
| Generate / upload a project-level tracking plan (source material is PRD / chat / template / code; deliverable is a real platform tracking plan) | ae-generate-tracking-plan |
|
|
36
|
+
| Upload documents / URLs to a knowledge base | ae-kb |
|
|
37
|
+
| Reports / dashboards / queries / governance on data already in AE | ae-analysis |
|
|
38
|
+
|
|
39
|
+
This skill also produces a tracking-plan draft (`source_type: data`) as a governance prerequisite; that draft is an input to ae-generate-tracking-plan, not a substitute for its five-phase platform plan.
|
|
40
|
+
|
|
26
41
|
## Workflow
|
|
27
42
|
|
|
28
43
|
Walk the four submodules in order. Each submodule is its own reference; follow it and come back here for the next step.
|
|
29
44
|
|
|
30
45
|
1. **Source — business identification.** Read [references/source-inspect.md](references/source-inspect.md). Profile every file fully, infer its business meaning using business-doc / user-prompt priors, then pick a branch via [references/ue-routing.md](references/ue-routing.md).
|
|
31
|
-
2. **Reuse check.** If
|
|
32
|
-
3. **Tracking plan.** Read [references/tracking-plan.md](references/tracking-plan.md).
|
|
46
|
+
2. **Reuse check.** If the profile is `ue_eligible`, read [references/reuse.md](references/reuse.md) and match the recommended mapping against the handoff index. `reuse` searches the current directory's `.ae-cli/data-integration/` upward, then `~/.ae-cli/data-integration/`, so a package written elsewhere is still found. A match proposes a frozen package; after one explicit confirmation, run the returned `transform.mjs` command and jump to Sink (step 5). No match → continue.
|
|
47
|
+
3. **Tracking plan.** Read [references/tracking-plan.md](references/tracking-plan.md). The plan is generated from the mapping (`plan --mapping`), so confirm the recommended mapping's key system fields with the user first — `mode`, `#account_id`/`#distinct_id`, `#time` + timezone, `#event_name`, `#ip`/`#uuid` (see [references/transform.md](references/transform.md) steps 1–5) — then generate the event/property plan and get a single explicit confirmation from the user before touching data. The plan is a separate, required deliverable from the transform mapping: a user who supplies a column→field mapping directly has **not** completed this step, so build the plan from the confirmed mapping anyway. `user_set` still requires a plan (no events; every property becomes a user property). This step runs for **every** file: a second or later file merges its new events and properties into the existing project plan (tracking-plan.md step 4) — an existing plan is never a reason to skip it.
|
|
33
48
|
4. **Transform.** Read [references/transform.md](references/transform.md). Map columns to AE system fields and properties, convert, and quarantine dirty rows per [references/ue-mapping.md](references/ue-mapping.md).
|
|
34
49
|
5. **Sink — upload.** Read [references/sink-upload.md](references/sink-upload.md). Resolve the destination, dry-run, confirm, then upload per [references/sync-json-upload.md](references/sync-json-upload.md). `receiver_accepted` is not persistence: after a ~1-minute ingestion delay, verify the data landed with ae-cli (`tracking live-data list` / `tracking ingest summary` / `tracking ingest-error list`) rather than telling the user to check the console.
|
|
35
|
-
6. **Handoff.** Read [references/handoff.md](references/handoff.md). Export the reusable package (frozen
|
|
50
|
+
6. **Handoff.** Read [references/handoff.md](references/handoff.md). Export the reusable package (pipeline descriptor + frozen mappings + stage executors + docs) and a shareable zip; in the completion response, state the absolute zip path, the package directory, and the one-line way to run the next same-shape file.
|
|
51
|
+
|
|
52
|
+
## Error handling
|
|
53
|
+
|
|
54
|
+
When a step fails, classify the failure before acting — see [references/error-handling.md](references/error-handling.md). A quarantined row, a ragged line, and a disk-full are three different problems with three different responses: match on the error `code`, never retry a parse failure by guessing the encoding, and never report a program failure as a data problem.
|
|
36
55
|
|
|
37
56
|
## Local analysis branch
|
|
38
57
|
|
|
@@ -0,0 +1,93 @@
|
|
|
1
|
+
# Project custom layer (verify + salvage overlays)
|
|
2
|
+
|
|
3
|
+
The handoff package is deliberately **standard flow only**: it ships the generic
|
|
4
|
+
stage executors and never embeds environment-specific workarounds such as SQL
|
|
5
|
+
direct queries or multi-round salvage loops. Some projects need those — a hard
|
|
6
|
+
per-event persistence judge, or a repeated salvage loop for a dirty source. Add
|
|
7
|
+
them as a **custom layer next to the package**, never by editing the generic
|
|
8
|
+
package itself.
|
|
9
|
+
|
|
10
|
+
## When to add a layer
|
|
11
|
+
|
|
12
|
+
- **Hard persistence judge.** `bin/verify.py` is a soft check: it computes the
|
|
13
|
+
submit window and expected counts from the local UE output, then shows the
|
|
14
|
+
`tracking ingest summary` payload before and after for comparison. It does not
|
|
15
|
+
parse that payload into per-event counts — the capability's `data` shape is
|
|
16
|
+
server-defined, and a shared project's window delta cannot be attributed to one
|
|
17
|
+
import. When the project requires an automated per-event ✓/✗ verdict, overlay a
|
|
18
|
+
SQL layer that reads the project's event/user tables directly.
|
|
19
|
+
- **Multi-round salvage.** `bin/run.sh` already prints a single-round salvage
|
|
20
|
+
hint when `invalid.rows.jsonl` is non-empty. For sources that fail in layers,
|
|
21
|
+
wrap that command in a loop that re-feeds each round's quarantine into the next
|
|
22
|
+
`--salvage-from` until nothing fails or the user stops.
|
|
23
|
+
|
|
24
|
+
## Rules
|
|
25
|
+
|
|
26
|
+
1. Put overlays in a sibling directory (e.g. `custom/`) under the package root,
|
|
27
|
+
and reference `../pipeline.json` / `../index.json` — never modify the frozen
|
|
28
|
+
mappings, `bin/`, or the generated docs.
|
|
29
|
+
2. Keep every existing confirmation gate. A custom layer may add gates; it must
|
|
30
|
+
not remove the `--confirm` upload gate or the plan/shape gates.
|
|
31
|
+
3. Keep secrets in `.local/target.env`. Never write APPID, endpoints, tokens, or
|
|
32
|
+
raw data values into an overlay file.
|
|
33
|
+
4. An overlay is project-specific. Do not copy it into another project's package
|
|
34
|
+
without re-confirming the environment assumptions (table names, APPID, host).
|
|
35
|
+
|
|
36
|
+
## Skeleton 1 — SQL verify (hard per-event judge)
|
|
37
|
+
|
|
38
|
+
The reference deployment queries the ingested tables directly. **Table and
|
|
39
|
+
column names vary by deployment** (`v_event_1` / `v_user_1` and `$part_date` /
|
|
40
|
+
`$part_event` are examples only) — confirm them against the project's receiver
|
|
41
|
+
schema before use, and keep the query read-only.
|
|
42
|
+
|
|
43
|
+
```bash
|
|
44
|
+
# custom/verify.sh — hard judge; requires the SQL capability to be available.
|
|
45
|
+
# usage: custom/verify.sh <run-dir> [--baseline|--check]
|
|
46
|
+
# Reuses bin/verify.py's window/count computation idea, but answers the question
|
|
47
|
+
# "did exactly these events land?" with a direct count over the event table.
|
|
48
|
+
```
|
|
49
|
+
|
|
50
|
+
The query shape (adapt table/column names):
|
|
51
|
+
|
|
52
|
+
```text
|
|
53
|
+
select <event_column>, count(*)
|
|
54
|
+
from <event_table>
|
|
55
|
+
where <date_partition_column> between '<window-start>' and '<window-end>'
|
|
56
|
+
group by 1
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
Compare the per-event delta (after upload minus before upload) against the
|
|
60
|
+
expected counts from `valid.ue.jsonl`. User-profile rows are overwrite-writes,
|
|
61
|
+
so a zero user-table delta is normal; compare event rows per event, not totals.
|
|
62
|
+
|
|
63
|
+
Keep this layer read-only and non-blocking: report the verdict, never retransmit
|
|
64
|
+
automatically on a mismatch — first check `tracking ingest-error list` for the
|
|
65
|
+
silent-drop reason.
|
|
66
|
+
|
|
67
|
+
## Skeleton 2 — multi-round salvage loop
|
|
68
|
+
|
|
69
|
+
`bin/run.sh` prints the one-round hint. Wrap it into a loop that re-processes each
|
|
70
|
+
round's quarantine against the fixed mapping until clean or the user stops:
|
|
71
|
+
|
|
72
|
+
```bash
|
|
73
|
+
# custom/salvage.sh <run-dir> — re-process quarantined rows round by round.
|
|
74
|
+
# Each round's valid.ue.jsonl is disjoint from earlier rounds', so upload each
|
|
75
|
+
# round independently (same --confirm and --allow-clean-subset gates).
|
|
76
|
+
round=1
|
|
77
|
+
while [ -s "<run-dir>/invalid.rows.jsonl" ]; do
|
|
78
|
+
ae-cli data-integration convert \
|
|
79
|
+
--input-file '<same-source>' \
|
|
80
|
+
--mapping '<fixed-mapping.json>' \
|
|
81
|
+
--salvage-from "<run-dir>/invalid.rows.jsonl" \
|
|
82
|
+
--output-dir "<run-dir>-salvage-$round"
|
|
83
|
+
round=$((round + 1))
|
|
84
|
+
# stop condition is the user's: a row that still fails after a fix is a real
|
|
85
|
+
# data defect, not a code bug.
|
|
86
|
+
done
|
|
87
|
+
```
|
|
88
|
+
|
|
89
|
+
See [transform.md](transform.md) for the `--salvage-from` semantics (single-file
|
|
90
|
+
only; the source must be the same file that produced the quarantine) and
|
|
91
|
+
[handoff.md](handoff.md) for the package the overlay attaches to. Persistence
|
|
92
|
+
verification always follows [sink-upload.md](sink-upload.md): `receiver_accepted`
|
|
93
|
+
is not durability.
|
|
@@ -0,0 +1,92 @@
|
|
|
1
|
+
# Error handling
|
|
2
|
+
|
|
3
|
+
Failures in the data-integration pipeline split into three classes. Each class has a distinct
|
|
4
|
+
response; an agent that mixes them up wastes retries and misleads the user.
|
|
5
|
+
|
|
6
|
+
| Class | What it is | Response |
|
|
7
|
+
| --- | --- | --- |
|
|
8
|
+
| Abnormal data | Rows or fields that do not meet UE rules | Quarantine the row or skip the field, report counts, salvage |
|
|
9
|
+
| File parsing exception | The source cannot be read as its detected format | Surface the parse error, do not retry by guessing; fix the file or format |
|
|
10
|
+
| Program execution exception | The program itself failed (disk, permissions, output dir, mapping mismatch) | Surface the exact error, fix the environment/inputs, re-run |
|
|
11
|
+
|
|
12
|
+
Every error leaves the CLI as the standard envelope `{ ok, data, error: { type, code, message, hint } }`
|
|
13
|
+
(JSON on stdout; progress and warnings on stderr). The `code` is the stable identifier an agent
|
|
14
|
+
branches on; `hint` is human-readable guidance.
|
|
15
|
+
|
|
16
|
+
## Class 1 — Abnormal data
|
|
17
|
+
|
|
18
|
+
Happens inside `convert`. Two severities:
|
|
19
|
+
|
|
20
|
+
- **Whole-row quarantine.** Any error on a row — identity, time, event name, record type, or a
|
|
21
|
+
single property (type coercion, size/limit) — drops the entire row. The row is written to
|
|
22
|
+
`invalid.rows.jsonl` with its error codes and counted in `manifest.output.invalid_records`; the
|
|
23
|
+
manifest becomes `blocked`. Row-level codes: `MISSING_USER_ID`, `USER_ID_TOO_LONG`,
|
|
24
|
+
`INVALID_RECORD_TYPE`, `INVALID_TIME`, `TIME_OUT_OF_RANGE`, `INVALID_EVENT_NAME`,
|
|
25
|
+
`INVALID_ZONE_OFFSET`, `PROPERTY_TYPE_CONFLICT`, `PROPERTY_LIMIT_EXCEEDED`.
|
|
26
|
+
- **Field skip.** A `#ip` or `#uuid` value that violates its spec drops only that field and keeps
|
|
27
|
+
the row. Counts land in `manifest.output.skipped_fields`; private/LAN IPs are kept but counted in
|
|
28
|
+
`manifest.output.lan_ip_records`. Field-skip codes: `INVALID_IP`, `INVALID_UUID`. Never quarantined.
|
|
29
|
+
|
|
30
|
+
Responses:
|
|
31
|
+
|
|
32
|
+
- A blocked manifest is not a failure to retry blindly. Fix the mapping, then re-run `convert
|
|
33
|
+
--salvage-from <invalid.rows.jsonl>` to re-process only the quarantined rows (never re-send the
|
|
34
|
+
rows that already passed). Repeat the salvage loop until no rows fail or the user stops.
|
|
35
|
+
- An empty source (zero data rows) blocks with the reason `The source contained no data rows.` —
|
|
36
|
+
a distinct, clearer message than a failed validation run. Do not re-run the same command on the
|
|
37
|
+
same empty file expecting a different result.
|
|
38
|
+
- Ragged delimited rows (a CSV/TSV row whose column count differs from the header) are tolerated:
|
|
39
|
+
extra fields are dropped and missing fields are treated as empty, and a stderr warning reports
|
|
40
|
+
how many rows were ragged. A row that becomes missing its identity/time because a field was
|
|
41
|
+
absent is then quarantined normally. Treat the warning as a data-quality signal for the user.
|
|
42
|
+
|
|
43
|
+
## Class 2 — File parsing exceptions
|
|
44
|
+
|
|
45
|
+
The source cannot be parsed as its detected format. The agent's job is to report precisely and let
|
|
46
|
+
the user decide — never to guess an encoding or structure and retry silently.
|
|
47
|
+
|
|
48
|
+
| Code | Meaning | What to do |
|
|
49
|
+
| --- | --- | --- |
|
|
50
|
+
| `LOCAL_DATA_INPUT_NOT_FOUND` | Path is not a readable file | Ask for the correct path |
|
|
51
|
+
| `LOCAL_DATA_FILE_TOO_LARGE` | XLS over 1 GB | Convert to XLSX or split the workbook |
|
|
52
|
+
| `LOCAL_DATA_INPUT_INVALID` | Generic parse failure (malformed CSV/TSV/JSON/XLS) | Verify encoding and structure, then retry without changing the source |
|
|
53
|
+
| `LOCAL_DATA_JSONL_INVALID` | A JSONL line is not valid JSON (`location.record` names the line) | Point at the offending record |
|
|
54
|
+
| `LOCAL_DATA_JSON_ROOT_INVALID` | JSON root is not an object or array | Check the file's top-level shape |
|
|
55
|
+
| `LOCAL_DATA_XLSX_INVALID` | Workbook metadata missing / no readable sheets / worksheet entry missing | Re-export the workbook |
|
|
56
|
+
| `LOCAL_DATA_SET_NOT_FOUND` / `LOCAL_DATA_SET_REQUIRED` | Sheet or JSON Path not found / ambiguous | Ask which `--data-set` to use |
|
|
57
|
+
|
|
58
|
+
A parse error is never a reason to change the mapping or the tracking plan. Report the code and
|
|
59
|
+
hint verbatim, and ask the user to fix the file (re-export, re-encode, or split).
|
|
60
|
+
|
|
61
|
+
## Class 3 — Program execution exceptions
|
|
62
|
+
|
|
63
|
+
The program itself failed; the source data is usually fine. These must never be mislabeled as
|
|
64
|
+
parse errors: a per-row callback failure (for example a disk that filled up mid-convert) propagates
|
|
65
|
+
as itself, not as `LOCAL_DATA_INPUT_INVALID`.
|
|
66
|
+
|
|
67
|
+
| Code | Meaning | What to do |
|
|
68
|
+
| --- | --- | --- |
|
|
69
|
+
| `LOCAL_DATA_OUTPUT_NOT_EMPTY` | The output directory must be new or empty | Point convert at a fresh `<run-id>` directory |
|
|
70
|
+
| `LOCAL_DATA_SOURCE_CHANGED` / `LOCAL_DATA_SOURCE_FORMAT_CHANGED` | The source no longer matches the mapping fingerprint/format | Re-run inspect and review a new mapping |
|
|
71
|
+
| `LOCAL_DATA_MAPPING_INVALID` / `LOCAL_DATA_MAPPING_INVALID_JSON` / `LOCAL_DATA_MAPPING_NOT_FOUND` | The mapping cannot be read or validated | Re-read the mapping reference, fix the mapping file |
|
|
72
|
+
| `LOCAL_DATA_SALVAGE_INVALID` / `LOCAL_DATA_SALVAGE_EMPTY` / `LOCAL_DATA_SALVAGE_NO_MATCH` | The salvage file is not a valid quarantine file, is empty, or lists no rows from this source | Point at the correct `invalid.rows.jsonl` from the same source |
|
|
73
|
+
| `LOCAL_DATA_TYPE_CONFLICTS_UNRESOLVED` / `LOCAL_DATA_TYPE_RESOLUTIONS_INVALID` | Cross-file column type conflicts need explicit resolutions | Build `--type-resolutions` |
|
|
74
|
+
| `LOCAL_DATA_PLAN_INVALID_EVENT_NAME` / `LOCAL_DATA_PLAN_EVENT_NAMES_REQUIRED` / `LOCAL_DATA_PLAN_INVALID_LANG` | Tracking-plan draft inputs are invalid | Fix the event name(s) or `--lang` |
|
|
75
|
+
| `LOCAL_DATA_HANDOFF_INDEX_INVALID` / `LOCAL_DATA_HANDOFF_PLAN_NOT_FOUND` / `LOCAL_DATA_HANDOFF_PLAN_INVALID` | Handoff index/plan cannot be read | Point at a valid handoff directory/plan file |
|
|
76
|
+
| `LOCAL_DATA_ENDPOINT_INVALID` / `LOCAL_DATA_APPID_INVALID` / `LOCAL_DATA_BATCH_SIZE_INVALID` / `LOCAL_DATA_COMPRESS_INVALID` / `LOCAL_DATA_RESUME_INVALID` / `LOCAL_DATA_RESUME_OUT_OF_RANGE` | Upload arguments are invalid | Fix the flag before uploading |
|
|
77
|
+
| `LOCAL_DATA_MANIFEST_INVALID` / `LOCAL_DATA_UE_FILE_INVALID` / `LOCAL_DATA_UE_FILE_NOT_FOUND` / `LOCAL_DATA_UE_FILE_CHANGED` / `LOCAL_DATA_UE_COUNT_MISMATCH` / `LOCAL_DATA_MANIFEST_FILE_MISMATCH` | Upload preconditions fail | Re-check the manifest and UE file pairing |
|
|
78
|
+
| `LOCAL_DATA_CLEAN_SUBSET_CONFIRMATION_REQUIRED` | Uploading from a blocked manifest needs a separate clean-subset decision | Confirm the subset explicitly before `--allow-clean-subset` |
|
|
79
|
+
|
|
80
|
+
Write failures (disk full `ENOSPC`, permission denied `EACCES`) do not hang the command: the
|
|
81
|
+
output streams fail with a clear `Failed to write "<path>"` message telling the user to check disk
|
|
82
|
+
space and directory permissions. Fix the environment, then re-run into a fresh output directory.
|
|
83
|
+
|
|
84
|
+
## Cross-cutting rules
|
|
85
|
+
|
|
86
|
+
- Classify first, act second. Match on `code`, not on message text.
|
|
87
|
+
- A parse error is not a data problem; a program error is not a parse error. Do not conflate them
|
|
88
|
+
when reporting back to the user.
|
|
89
|
+
- Never retry a parse failure by guessing the encoding, delimiter, or header layout. Show the
|
|
90
|
+
error and ask.
|
|
91
|
+
- After any fix, re-run from the start of the step that failed; do not resume a half-written run
|
|
92
|
+
or reuse a partially populated output directory.
|
|
@@ -1,43 +1,102 @@
|
|
|
1
1
|
# Handoff (reusable package)
|
|
2
2
|
|
|
3
|
-
After a successful run, export a **handoff package** so the next file of the same shape skips the full pipeline. The package freezes the transform logic (the confirmed
|
|
3
|
+
After a successful run, export a **handoff package** so the next file of the same shape skips the full pipeline. The package is a DataX-style pipeline — a declarative `source → transform → sink` descriptor plus generic stage executors that dispatch to `ae-cli data-integration` subcommands. It freezes the transform logic (the confirmed mappings) and the tracking-plan reference, and ships a `bin/` directory that a human or agent can run directly. It never re-freezes raw data or upload secrets.
|
|
4
4
|
|
|
5
5
|
## When
|
|
6
6
|
|
|
7
|
-
Run handoff after Transform — or after Sink — once the
|
|
7
|
+
Run handoff after Transform — or after Sink — once the mappings are confirmed. It is local-only and idempotent: re-running it for the same table refreshes the package in place.
|
|
8
8
|
|
|
9
9
|
## CLI
|
|
10
10
|
|
|
11
11
|
```
|
|
12
|
-
ae-cli data-integration handoff --mapping <mapping> [--plan-file <draft.json>] [--out-dir .ae-data-integration]
|
|
12
|
+
ae-cli data-integration handoff --mapping <mapping> [--mapping <mapping2> ...] [--plan-file <draft.json>] [--pushurl <url>] [--project-id <id>] [--out-dir .ae-cli/data-integration]
|
|
13
13
|
```
|
|
14
14
|
|
|
15
|
-
- `--mapping` —
|
|
16
|
-
- `--plan-file` — optional tracking-plan `draft.json` to reference inside
|
|
17
|
-
- `--
|
|
18
|
-
- `--
|
|
15
|
+
- `--mapping` — one or more confirmed `ae-data-integration-mapping/v1` mappings (the frozen transform logic, including `value_mapping` and `flatten_rules`). Repeat it for multi-sheet workbooks: one mapping per sheet.
|
|
16
|
+
- `--plan-file` — optional tracking-plan `draft.json` to reference inside each mapping directory (`plan.json`).
|
|
17
|
+
- `--pushurl` — optional receiver base URL to record as the reuse upload target (the sink endpoint is `pushurl` + `/sync_json`). Record it when the next same-shape file will most likely land at the same receiver.
|
|
18
|
+
- `--project-id` — optional numeric destination project ID to record; `bin/upload.sh` derives the APPID from it via `project info get` at upload time.
|
|
19
|
+
- `--out-dir` — handoff root. Default `.ae-cli/data-integration/` (project workspace; travels with the project).
|
|
20
|
+
- `--dry-run` previews the fingerprints, target, file list, and zip path without writing.
|
|
19
21
|
|
|
20
22
|
## Package layout
|
|
21
23
|
|
|
22
24
|
```
|
|
23
|
-
.ae-data-integration/
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
25
|
+
.ae-cli/data-integration/ ← handoff root (= out-dir, project workspace)
|
|
26
|
+
pipeline.json ← source → transform → sink descriptor (ae-data-integration-pipeline/v1)
|
|
27
|
+
index.json ← shared structure-fingerprint index (reuse detection)
|
|
28
|
+
shape.json ← per-mapping column baseline (shape gate)
|
|
29
|
+
<fingerprint[:16]>/ ← one directory per mapping
|
|
30
|
+
mapping.json ← frozen mapping
|
|
31
|
+
transform.mjs ← node transform.mjs <new-file> [<output-dir>]
|
|
32
|
+
plan.json ← optional tracking-plan reference
|
|
33
|
+
bin/ ← generic stage executors (read pipeline.json)
|
|
34
|
+
run.sh ← source → transform → plan (never uploads)
|
|
35
|
+
upload.sh ← sink (dry-run by default; --confirm uploads; resolves recorded target)
|
|
36
|
+
bind_mapping.py ← shape check + rebind sha256/data_set to the new file
|
|
37
|
+
summarize.py ← valid/quarantined counts
|
|
38
|
+
plan_check.py ← tracking-plan coverage gate (exit 3 on new events/properties)
|
|
39
|
+
verify.py ← soft persistence check (submit window vs ingest summary)
|
|
40
|
+
resolve_appid.py ← APPID derivation helper (project info get)
|
|
41
|
+
README.md / RUNBOOK.md ← how to run, the four gates, persistence verification
|
|
42
|
+
.local/target.env.example ← destination template (no real secrets)
|
|
43
|
+
.gitignore ← inbox/ runs/ .local/target.env
|
|
44
|
+
inbox/ runs/ ← daily input / per-run outputs
|
|
29
45
|
```
|
|
30
46
|
|
|
31
|
-
|
|
47
|
+
A shareable archive is written next to the package root: `<parent>/ae-data-integration-handoff-<fingerprint[:8]>.zip`. The zip carries only this handoff round — the frozen mappings just written, a scoped `index.json` (just this round's entries), and the generic executors/docs — not the accumulated history, which stays in `.ae-cli/data-integration/` for reuse matching.
|
|
48
|
+
|
|
49
|
+
## Pipeline descriptor
|
|
50
|
+
|
|
51
|
+
`pipeline.json` declares the three stages and their types. Only `source: local_file` and `sink: restful_sync_json` are implemented today; the `type` fields reserve logbus / datax / mysql for later phases. `bin/run.sh` and `bin/upload.sh` read the descriptor and dispatch each stage by its `type` to `ae-cli data-integration inspect / convert / upload` — they are executors, not a second runtime engine.
|
|
52
|
+
|
|
53
|
+
## Recorded destination
|
|
54
|
+
|
|
55
|
+
`pipeline.json` → `sink.params` records `pushurl` and `project_id` when the handoff
|
|
56
|
+
was run with those flags. Reuse defaults to that target, but **`bin/upload.sh`
|
|
57
|
+
never sends without `--confirm`**, so the operator re-confirms the address and
|
|
58
|
+
project on every reuse. Resolution order at upload time:
|
|
59
|
+
|
|
60
|
+
- endpoint: recorded `pushurl` (+ `/sync_json`), else `AE_ENDPOINT`.
|
|
61
|
+
- APPID: `AE_APPID`, else derived via `ae-cli project info get --project-id <id>`
|
|
62
|
+
(see `bin/resolve_appid.py`; set `AE_APPID` when that payload lacks `appid`).
|
|
63
|
+
- project id: recorded `project_id`, else `AE_PROJECT_ID`.
|
|
64
|
+
|
|
65
|
+
`project info get` returns `data.appid` at the top level (verified against the AE
|
|
66
|
+
demo host); `bin/resolve_appid.py` reads that exact field and prints it, falling
|
|
67
|
+
back to `AE_APPID` when the field is absent or not a non-empty string.
|
|
68
|
+
|
|
69
|
+
## Project custom layers
|
|
70
|
+
|
|
71
|
+
For projects that need a hard per-event SQL judge or a multi-round salvage loop
|
|
72
|
+
(the reference package's environment-specific workarounds), overlay a custom layer
|
|
73
|
+
next to the package instead of editing it — see [custom-layer.md](custom-layer.md).
|
|
74
|
+
|
|
75
|
+
## Four confirmation gates
|
|
76
|
+
|
|
77
|
+
The RUNBOOK and the scripts enforce four gates. The first two run automatically; the last two always need human confirmation:
|
|
78
|
+
|
|
79
|
+
1. **Shape gate** — `bind_mapping.py` compares the new file's column set against `shape.json` and fails fast on a mismatch. A changed shape means the frozen logic was never reviewed for it: re-run the full pipeline.
|
|
80
|
+
2. **Transform** — `ae-cli data-integration convert` per mapping; quarantined rows go to `invalid.rows.jsonl`, never silently dropped.
|
|
81
|
+
3. **Tracking-plan gate** — `plan_check.py` verifies every produced event/property already exists in `plan.json`; new ones exit 3 and must be merged into the project plan first.
|
|
82
|
+
4. **Sink gate** — `upload.sh` is dry-run by default; `--confirm` is the explicit upload decision.
|
|
32
83
|
|
|
33
84
|
## Structure fingerprint and index
|
|
34
85
|
|
|
35
|
-
Every handoff records one entry in `index.json` (`ae-data-integration-index/v1`) keyed by a **structure fingerprint** — a SHA-256 over the table shape: columns (
|
|
86
|
+
Every handoff records one entry per mapping in `index.json` (`ae-data-integration-index/v1`) keyed by a **structure fingerprint** — a SHA-256 over the table shape: the raw source columns (by name, reconstructed so flatten, exclude, and account-vs-distinct decisions don't move it), the format, and the event model (`mode`). Business logic (`value_mapping`, transforms, `time_format`, `flatten_rules`, `exclude_columns`, the fixed `default_event_name`, and the system-field assignments) is excluded, so re-handing off the same table with new business rules refreshes the existing entry instead of forking a new one.
|
|
87
|
+
|
|
88
|
+
Reuse matching is its own step — see [references/reuse.md](reuse.md).
|
|
36
89
|
|
|
37
|
-
|
|
90
|
+
## Completion response
|
|
91
|
+
|
|
92
|
+
After handoff succeeds, state the **absolute zip path**, the package directory, and the one-line way to run the next same-shape file:
|
|
93
|
+
|
|
94
|
+
```
|
|
95
|
+
cd <out-dir> && bin/run.sh <new-file> # then bin/upload.sh runs/<run-id> --confirm
|
|
96
|
+
```
|
|
38
97
|
|
|
39
98
|
## Safety rules
|
|
40
99
|
|
|
41
|
-
- Treat the
|
|
42
|
-
- Never write APPID,
|
|
100
|
+
- Treat the mappings, plans, generated artifacts, the `.ae-cli/data-integration/` directory, and the zip as sensitive.
|
|
101
|
+
- Never write APPID, tokens, or raw data values into a handoff package. The package records at most a destination `pushurl` and `project_id`; uploads still require an explicit, confirmed `upload` call, so the operator re-confirms the address and project each time.
|
|
43
102
|
- Do not invent a mapping or plan. Handoff only packages what the user already confirmed.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Local analysis
|
|
2
2
|
|
|
3
|
-
Keep the source on the local machine. Generated scripts and reports belong under `.ae-cli/data-integration/<run-id>/` with restrictive permissions.
|
|
3
|
+
Keep the source on the local machine. Generated scripts and reports belong under `.ae-cli/data-integration/runs/<run-id>/` with restrictive permissions.
|
|
4
4
|
Set the directory to `0700` and generated scripts/reports to `0600`.
|
|
5
5
|
|
|
6
6
|
## Default report when no question is supplied
|
|
@@ -4,21 +4,25 @@ Before walking the full pipeline for a new file, check whether a **handoff packa
|
|
|
4
4
|
|
|
5
5
|
## When
|
|
6
6
|
|
|
7
|
-
Run reuse right after `inspect`, only when the profile is `ue_eligible` and a `.ae-data-integration/index.json` exists. It is read-only and makes no writes.
|
|
7
|
+
Run reuse right after `inspect`, only when the profile is `ue_eligible` and a `.ae-cli/data-integration/index.json` exists. It is read-only and makes no writes.
|
|
8
8
|
|
|
9
9
|
## CLI
|
|
10
10
|
|
|
11
11
|
```
|
|
12
|
-
ae-cli data-integration reuse --mapping <recommended_mapping> [--out-dir .ae-data-integration]
|
|
12
|
+
ae-cli data-integration reuse --mapping <recommended_mapping> [--out-dir .ae-cli/data-integration]
|
|
13
13
|
```
|
|
14
14
|
|
|
15
15
|
- `--mapping` — the candidate mapping, typically `inspect`'s `recommended_mapping` (the new file's structure, inferred exactly as Source would infer it).
|
|
16
|
-
- `--out-dir` — handoff root.
|
|
16
|
+
- `--out-dir` — handoff root. Optional. Without it, `reuse` searches in order: the current directory's `.ae-cli/data-integration/`, then each parent directory upward, then `~/.ae-cli/data-integration/` as a global fallback — so a package written in another directory or another session is still reachable.
|
|
17
17
|
- `--dry-run` previews the fingerprint and match verdict without reading frozen packages.
|
|
18
18
|
|
|
19
|
+
## Search paths
|
|
20
|
+
|
|
21
|
+
When `--out-dir` is omitted, the result carries `searched_paths` — the exact `index.json` paths probed, in order. A `matched: false` result with a non-empty `searched_paths` means all of them were checked; a missing global fallback (no `$HOME`) simply omits it from the list.
|
|
22
|
+
|
|
19
23
|
## Matching
|
|
20
24
|
|
|
21
|
-
The command computes the same **structure fingerprint** the handoff index is keyed on — columns (
|
|
25
|
+
The command computes the same **structure fingerprint** the handoff index is keyed on — the raw source columns (by name, reconstructed so flatten, exclude, and account-vs-distinct decisions don't move it), the format, and the event model — and looks it up in `index.json`. Business logic (the frozen event name, `value_mapping`, transforms, `time_format`, `flatten_rules`, `exclude_columns`, and the system-field assignments) is not part of the key, so a same-shape file with different content or a different file name still matches.
|
|
22
26
|
|
|
23
27
|
## Result
|
|
24
28
|
|
|
@@ -26,7 +30,7 @@ The command computes the same **structure fingerprint** the handoff index is key
|
|
|
26
30
|
- **Match** (`matched: true`): the result carries the matched package — `mapping_file`, optional `plan_file`, the frozen `default_event_name` (so the user sees which event name will be reused), and a `run` command:
|
|
27
31
|
|
|
28
32
|
```
|
|
29
|
-
node .ae-data-integration/<fingerprint[:16]>/transform.mjs <new-input-file> [<output-dir>]
|
|
33
|
+
node .ae-cli/data-integration/<fingerprint[:16]>/transform.mjs <new-input-file> [<output-dir>]
|
|
30
34
|
```
|
|
31
35
|
|
|
32
36
|
## Confirmation gate
|
|
@@ -31,7 +31,7 @@ ae-cli data-integration upload \
|
|
|
31
31
|
--dry-run
|
|
32
32
|
```
|
|
33
33
|
|
|
34
|
-
Show the masked target, project, file fingerprint, record count, quarantined count, batch count, and persistence limitation. Re-state the system-field mapping first (`#type`/mode, `#account_id`, `#distinct_id`, `#time` + source timezone, `#event_name`), then re-list the final property mapping for every file or sheet being uploaded — source column → target AE name → type, grouped into event properties (`track`) and user properties (profile modes) — never a counts-only summary. Wait for explicit confirmation. Execute the same command without `--dry-run` only after confirmation. For a blocked manifest, first show the quarantine statistics and separately ask whether the user accepts uploading only valid rows; add `--allow-clean-subset` only after a clear yes.
|
|
34
|
+
Show the masked target, project, file fingerprint, record count, quarantined count, batch count, and persistence limitation. Re-state the system-field mapping first (`#type`/mode, `#account_id`, `#distinct_id`, `#time` + source timezone, `#event_name`), then re-list the final property mapping for every file or sheet being uploaded — source column → target AE name → type (+ `display_name`/`desc` when set), grouped into event properties (`track`) and user properties (profile modes), with each event's attached properties listed — never a counts-only summary. For a multi-sheet workbook, group by sheet. Wait for explicit confirmation. Execute the same command without `--dry-run` only after confirmation. For a blocked manifest, first show the quarantine statistics and separately ask whether the user accepts uploading only valid rows; add `--allow-clean-subset` only after a clear yes.
|
|
35
35
|
|
|
36
36
|
`status=receiver_accepted` means receiver acceptance only, not durable storage. Say that persistence remains unverified. Never report success on this status alone.
|
|
37
37
|
|
|
@@ -39,7 +39,9 @@ If `selection_required=true`, show only the Sheet/JSON Path candidates and ask t
|
|
|
39
39
|
ae-cli data-integration inspect --input-file '<path>' --data-set '<candidate-id>' --source-timezone '<iana-timezone>'
|
|
40
40
|
```
|
|
41
41
|
|
|
42
|
-
Report row/column counts, field types, missing/unique/time-parse ratios, UE eligibility, mapping confidence, and warnings. Samples are bounded (up to 5 distinct, truncated) — summarize, never paste them. ID-like columns (`id`, `*_id`, `*_key`, `*_code`, `*_no`, `*_num`) stay `string` even when every value is numeric; JSON-encoded object/array values inside CSV cells are recognized as `object`/`list`, not `string`. Read [UE routing](ue-routing.md) before choosing a branch.
|
|
42
|
+
Report row/column counts, field types, missing/unique/time-parse ratios, UE eligibility, mapping confidence, and warnings. Samples are bounded (up to 5 distinct, truncated) — summarize, never paste them. ID-like columns (`id`, `*_id`, `*_key`, `*_code`, `*_no`, `*_num`) stay `string` even when every value is numeric; JSON-encoded object/array values inside CSV cells are recognized as `object`/`list`, not `string`. IP- and UUID-named columns are additionally checked against their value specs: inspect warns how many non-empty values are invalid IPv4/IPv6, private/LAN IPs, or non-UUID strings, so the user can decide whether to map them as `ip_field`/`uuid_field`. Read [UE routing](ue-routing.md) before choosing a branch.
|
|
43
|
+
|
|
44
|
+
A stderr `Warning: … column count different from the header row …` means the CSV/TSV has ragged rows (extra fields dropped, missing fields treated as empty); report it as a data-quality signal. For how every pipeline failure — abnormal data, parse errors, and program errors — is classified and handled, see [error handling](error-handling.md).
|
|
43
45
|
|
|
44
46
|
## Advanced input
|
|
45
47
|
|
|
@@ -49,20 +51,24 @@ Report row/column counts, field types, missing/unique/time-parse ratios, UE elig
|
|
|
49
51
|
- **Excel sheets** — `--merge-sheets` streams every worksheet in file order instead of a single selected sheet; otherwise ask which sheet/`--data-set` to use. Inspect also reports `header_consistency` (`all_same` or `different`) across a workbook's sheets, with `header_details` listing each sheet's header row when they differ; prefer `--merge-sheets` only when headers match.
|
|
50
52
|
- **Multi-file type conflicts** — when the same column has different inferred types across files, present each conflict and resolve with `--type-resolutions` on `convert` (see [transform](transform.md)).
|
|
51
53
|
|
|
52
|
-
## Nested flattening (NDJSON/JSON records and JSON-encoded CSV/TSV cells)
|
|
54
|
+
## Nested flattening (NDJSON/JSON records and JSON-encoded CSV/TSV/Excel cells)
|
|
53
55
|
|
|
54
|
-
Nested data is
|
|
56
|
+
Nested data is analyzed per field, and the recommended mapping encodes one decision per container: keep whole, flatten one level, or flatten to the leaves. This flow is agent-driven and non-interactive: never pipe pre-filled answers into any prompt, and never silently pick a depth for a node.
|
|
55
57
|
|
|
56
58
|
1. **Locate the nested structure.**
|
|
57
59
|
- NDJSON/JSON: read the top-level `nested_tree` from the inspect result (record-root paths).
|
|
58
|
-
- CSV/TSV: a JSON-encoded
|
|
60
|
+
- CSV/TSV/Excel: a JSON-encoded column carries its own tree at `columns[].nested_tree` (paths relative to that cell). Object cells list child keys; array cells (`items`-style) expose an `elementKind` and, for object arrays, the union of the element fields.
|
|
59
61
|
Object nodes list child keys, array nodes carry an `elementKind`, primitive leaves carry an inferred type and bounded samples. Summarize node kinds, never paste samples.
|
|
60
|
-
2. **
|
|
61
|
-
|
|
62
|
-
-
|
|
63
|
-
-
|
|
64
|
-
-
|
|
65
|
-
|
|
62
|
+
2. **Review the recommended mapping's per-field decision, then confirm adjustments.**
|
|
63
|
+
The recommended mapping already encodes, per container, a data-driven depth decision — the depth is not uniform across fields:
|
|
64
|
+
- **Keep whole** — a single-level object (every child scalar or a scalar array) is declared `type: 'object'` with `transform: 'json'`, plus one `parent.child` sub-property per child; a single-level object array is declared `type: 'array_row'` with one `parent.child` sub-property per element field. The parent entry carries the native value and conversion emits it once; each child is a plan-only declaration (scalar or `list`) that never reads a column of its own.
|
|
65
|
+
- **Flatten one level / to the leaves** — an object that itself contains an object or object array cannot be kept whole; it collapses one level and each child is decided independently: scalar children become flat properties materialized through `flatten_rules`, and nested objects/arrays recurse until a single-level container is reached (which is then kept whole).
|
|
66
|
+
- **Scalar array** — `["a","b"]` stays `list` (→ AE `array_string`), a leaf with no children.
|
|
67
|
+
- **Nested element field** — an element field that is itself an object/array stays inside the array data and is not declared as a sub-property (array-element flattening is not supported); the mapping warns so it is never silently dropped.
|
|
68
|
+
Present the resulting property list as a table and default to it; ask the user to confirm or adjust only where you disagree with a node's inferred decision (business-entity names with stable scalar children vs generic containers like `payload`/`data`/`config` vs nesting deeper than one level with no clear intent).
|
|
69
|
+
3. **Record overrides in `flatten_rules` (`{ "out_column": "dot.path" }`) and `exclude_columns`.**
|
|
70
|
+
The recommended mapping already carries the flatten rules for its collapsed levels. When you override it:
|
|
66
71
|
- NDJSON/JSON: the path is from the record root (`user_info.name`).
|
|
67
|
-
- CSV/TSV: the path is `<column>.<cell-relative path>` (`user_profile.name`), and add the source column to `exclude_columns` so the whole object is not also mapped.
|
|
68
|
-
|
|
72
|
+
- CSV/TSV/Excel: the path is `<column>.<cell-relative path>` (`user_profile.name`), and add the source column to `exclude_columns` so the whole object is not also mapped.
|
|
73
|
+
- **`flatten_rules` only materializes the out column in the row — it does not emit it.** For every out column, also add a `properties` entry with `source` set to that out-column name (plus `target`/`type`/`desc`); otherwise the flattened value is silently dropped from the output record.
|
|
74
|
+
A leaf path becomes a snake_case out column when not explicitly named (`user_info.address.geo.lat` → `user_info_address_geo_lat`; cell-relative rules prefix the column: `user_profile.level` → `user_profile_level`). A string that looks like a number (phone, zip, ID) is kept as a string unless the user says otherwise. Containers kept whole are declared `type: 'object'`/`'array_row'`/`'list'` **with `transform: 'json'`** so conversion restores the native structure; a kept whole `object`/`array_row` also declares its scalar children as dotted `parent.child` sub-properties in `properties`.
|